Isabelle Augenstein

dblp:93/11424 · DBLP profile ↗
← Back
80ranked-venue papers
8as first author
50since 2021 · last 2026
0000-0003-1562-7909ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 72 · 5 first-author · 47 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CUB: Benchmarking Context Utilisation Techniques for Language Models
abstract
Incorporating external knowledge is crucial for knowledge-intensive tasks, such as question answering and fact checking. However, language models (LMs) may ignore relevant information that contradicts outdated parametric memory or be distracted by irrelevant contexts. While many context utilisation manipulation techniques (CMTs) have recently been proposed to alleviate these issues, few have seen systematic comparison. In this paper, we develop CUB (Context Utilisation Benchmark) - the first comprehensive benchmark designed to help diagnose CMTs under diverse noisy context conditions within retrieval-augmented generation (RAG). With this benchmark, we conduct the most extensive evaluation to date of seven state-of-the-art methods, representative of the main categories of CMTs, across three diverse datasets and tasks, applied to 11 LMs. Our findings expose critical gaps in current CMT evaluation practices, demonstrating the need for holistic testing. We reveal that most existing CMTs struggle to handle the full spectrum of context types encountered in real-world RAG scenarios. We also find that many CMTs display inflated performance on simple synthesised datasets, compared to more realistic datasets with naturally occurring samples.
Lovisa Hagström, Youna Kim, Haeun Yu, Sang-goo Lee, Richard Johansson, Hyunsoo Cho, Isabelle Augenstein
ACL (1)7
2026 Stress Testing Factual Consistency Metrics for Long-Document Summarization
abstract
Evaluating the factual consistency of abstractive text summarization remains a significant challenge, particularly for long documents, where conventional metrics struggle with input length limitations and long-range dependencies.In this work, we systematically evaluate the reliability of six widely used reference-free factuality metrics, originally proposed for shortform summarization, in the long-document setting.We probe metric robustness through seven factuality-preserving perturbations applied to summaries, namely paraphrasing, simplification, synonym replacement, logically equivalent negations, vocabulary reduction, compression, and source text insertion, and further analyze their sensitivity to retrieval context and claim information density.Across three longform benchmark datasets spanning science fiction, legal, and scientific domains, our results reveal that existing short-form metrics produce inconsistent scores for semantically equivalent summaries and exhibit declining reliability for information-dense claims whose content is semantically similar to many parts of the source document.While expanding the retrieval context improves stability in some domains, no metric consistently maintains factual alignment under long-context conditions.Finally, our results highlight concrete directions for improving factuality evaluation, including multi-span reasoning, context-aware calibration, and training on meaning-preserving variations to enhance robustness in long-form summarization. 1
Zain Muhammad Mujahid, Dustin Wright 0001, Isabelle Augenstein
ACL (1)3
2026 Explaining Sources of Uncertainty in Automated Fact-Checking
abstract
Human-AI collaboration in knowledgeintensive tasks such as fact-checking requires understanding model uncertainty in multi-document reasoning amid conflicting/agreeing evidence.Yet, existing methods only express uncertainty as numbers or hedges without revealing which evidence conflicts cause the uncertainty, leaving users unable to resolve disagreements.We present CLUE (Conflict-&Agreement-aware Language-model Uncertainty Explanations), a plug-and-play white-box framework that, to our knowledge, is the first to generate natural-language explanations of model uncertainty grounded in conflicting/agreeing evidence.CLUE (i) identifies span-level claim-evidence and inter-evidence relations that signal conflict or agreement without supervision, and (ii) uses these relations to steer explanation generation, articulating how they drive the model's uncertainty.Across three language models and two fact-checking datasets, CLUE produces explanations that more faithfully track model uncertainty and better align with the model's fact-checking decisions than span-agnostic explanation prompting; human raters also judge them more helpful, more informative, less redundant, and more logically consistent with the input.By explicitly tying uncertainty to evidence conflicts and agreements, CLUE supports practical fact-checking and other tasks that require reasoning over complex, conflicting information. * Equal contribution.Claim: Scientific data has shown that cats can be infected with SARS-CoV-2 and can spread it to other cats. ModelOutput: Supports ✅ Model Certainty: 73% [...] there is a possibility of spreading SARS-CoV-2 through domestic pets Evidence 1 [...] no further transmission events to other animals or persons Evidence 2 Automated claim verification Span interactions for model uncertainty Claim: Scientific data has shown that cats can be infected with SARS-CoV-2 and can spread it to other cats.
Greta Warren, Irina Shklovski, Isabelle Augenstein
ACL (1)4
2025 A Reality Check on Context Utilisation for Retrieval-Augmented Generation
abstract
Retrieval-augmented generation (RAG) helps address the limitations of parametric knowledge embedded within a language model (LM). In real world settings, retrieved information can vary in complexity, yet most investigations of LM utilisation of context has been limited to synthetic text. We introduce DRUID (Dataset of Retrieved Unreliable, Insufficient and Difficult-to-understand contexts) with real-world queries and contexts manually annotated for stance. The dataset is based on the prototypical task of automated claim verification, for which automated retrieval of real-world evidence is crucial. We compare DRUID to synthetic datasets (CounterFact, ConflictQA) and find that artificial datasets often fail to represent the complexity and diversity of realistically retrieved context. We show that synthetic datasets exaggerate context characteristics rare in real retrieved data, which leads to inflated context utilisation results, as measured by our novel ACU score. Moreover, while previous work has mainly focused on singleton context characteristics to explain context utilisation, correlations between singleton context properties and ACU on DRUID are surprisingly small compared to other properties related to context source. Overall, our work underscores the need for real-world aligned context utilisation studies to represent and improve performance in real-world RAG settings.
Lovisa Hagström, Sara Marjanovic, Haeun Yu, Arnav Arora, Christina Lioma, Maria Maistro, Pepa Atanasova, Isabelle Augenstein
ACL (1)8
2025 Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking
abstract
The pervasiveness of large language models and generative AI in online media has amplified the need for effective automated fact-checking to assist fact-checkers in tackling the increasing volume and sophistication of misinformation. The complex nature of fact-checking demands that automated fact-checking systems provide explanations that enable fact-checkers to scrutinise their outputs. However, it is unclear how these explanations should align with the decision-making and reasoning processes of fact-checkers to be effectively integrated into their workflows. Through semi-structured interviews with fact-checking professionals, we bridge this gap by: (i) providing an account of how fact-checkers assess evidence, make decisions, and explain their processes; (ii) examining how fact-checkers use automated tools in practice; and (iii) identifying fact-checker explanation requirements for automated fact-checking tools. The findings show unmet explanation needs and identify important criteria for replicable fact-checking explanations that trace the model's reasoning path, reference specific evidence, and highlight uncertainty and information gaps.
Greta Warren, Irina Shklovski, Isabelle Augenstein
CHI3
2025 SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages
abstract
Question Answering (QA) datasets have been instrumental in developing and evaluating Large Language Model (LLM) capabilities. However, such datasets are scarce for languages other than English due to the cost and difficulties of collection and manual annotation. This means that producing novel models and measuring the performance of multilingual LLMs in low-resource languages is challenging. To mitigate this, we propose SynDARin, a method for generating and validating QA datasets for low-resoucre languages. We utilize parallel content mining to obtain human-curated paragraphs between English and the target language. We use the English data as context to generate synthetic multiple-choice (MC) question-answer pairs, which are automatically translated and further validated for quality. Combining these with their designated non-English human-curated paragraphs form the final QA dataset. The method allows to maintain content quality, reduces the likelihood of factual errors, and circumvents the need for costly annotation. To test the method, we created a QA dataset with 1.2K samples for the Armenian language. The human evaluation shows that 98% of the generated English data maintains quality and diversity in the question types and topics, while the translation validation pipeline can filter out ~70% of data with poor quality. We use the dataset to benchmark state-of-the-art LLMs, showing their inability to achieve human accuracy with some model performances closer to random chance. This shows that the generated dataset is non-trivial and can be used to evaluate reasoning capabilities in low-resource language.
Gayane Ghazaryan, Erik Arakelyan, Isabelle Augenstein, Pasquale Minervini
COLING3
2025 FLARE: Faithful Logic-Aided Reasoning and Exploration
abstract
Modern Question Answering (QA) and Reasoning approaches with Large Language Models (LLMs) commonly use Chain-of-Thought (CoT) prompting but struggle with generating outputs faithful to their intermediate reasoning chains.While neuro-symbolic methods like Faithful CoT (F-CoT) offer higher faithfulness through external solvers, they require codespecialized models and struggle with ambiguous tasks.We introduce Faithful Logic-Aided Reasoning and Exploration (FLARE), which uses LLMs to plan solutions, formalize queries into logic programs, and simulate code execution through multi-hop search without external solvers.Our method achieves SOTA results on 7 out of 9 diverse reasoning benchmarks and 3 out of 3 logic inference benchmarks while enabling measurement of reasoning faithfulness.We demonstrate that model faithfulness correlates with performance and that successful reasoning traces show an 18.1% increase in unique emergent facts, 8.6% higher overlap between code-defined and execution-trace relations, and 3.6% reduction in unused relations.
Erik Arakelyan, Pasquale Minervini, Patrick S. H. Lewis, Patrick Verga, Isabelle Augenstein
EMNLP5
2025 Multi-Modal Framing Analysis of News
abstract
Automated frame analysis of political communication is a popular task in computational social science that is used to study how authors select aspects of a topic to frame its reception.So far, such studies have been narrow, in that they use a fixed set of pre-defined frames and focus only on the text, ignoring the visual contexts in which those texts appear.Especially for framing in the news, this leaves out valuable information about editorial choices, which include not just the written article but also accompanying photographs.To overcome such limitations, we present a method for conducting multi-modal, multi-label framing analysis at scale using large (vision-) language models.Grounding our work in framing theory, we extract latent meaning embedded in images used to convey a certain point and contrast that to the text by comparing the respective frames used.We also identify highly partisan framing of topics with issue-specific frame analysis found in prior qualitative work.We demonstrate a method for doing scalable integrative framing analysis of both text and image in news, providing a more complete picture for understanding media bias.
Arnav Arora, Srishti Yadav, Maria Antoniak, Serge J. Belongie, Isabelle Augenstein
EMNLP5
2025 Explainability and Interpretability of Multilingual Large Language Models: A Survey
abstract
Multilingual large language models (MLLMs) demonstrate state-of-the-art capabilities across diverse cross-lingual and multilingual tasks.Their complex internal mechanisms, however, often lack transparency, posing significant challenges in elucidating their internal processing of multilingualism, cross-lingual transfer dynamics and handling of language-specific features.This paper addresses this critical gap by presenting a survey of current explainability and interpretability methods specifically for MLLMs.To our knowledge, it is the first comprehensive review of its kind.Existing literature is categorised according to the explainability techniques employed, the multilingual tasks addressed, the languages investigated and available resources.The survey further identifies key challenges, distils core findings and outlines promising avenues for future research within this rapidly evolving domain.
Lucas Resck, Isabelle Augenstein, Anna Korhonen
EMNLP2
2025 Unstructured Evidence Attribution for Long Context Query Focused Summarization
abstract
Large language models (LLMs) are capable of generating coherent summaries from very long contexts given a user query, and extracting and citing evidence spans helps improve the trustworthiness of these summaries.Whereas previous work has focused on evidence citation with fixed levels of granularity (e.g.sentence, paragraph, document, etc.), we propose to extract unstructured (i.e., spans of any length) evidence in order to acquire more relevant and consistent evidence than in the fixed granularity case.We show how existing systems struggle to copy and properly cite unstructured evidence, which also tends to be "lost-in-the-middle".To help models perform this task, we create the Summaries with Unstructured Evidence Text dataset (SUnsET), a synthetic dataset generated using a novel pipeline, which can be used as training supervision for unstructured evidence summarization.We demonstrate across 5 LLMs and 4 datasets spanning human written, synthetic, single, and multi-document settings that LLMs adapted with SUnsET generate more relevant and factually consistent evidence with their summaries, extract evidence from more diverse locations in their context, and can generate more relevant and consistent summaries than baselines with no fine-tuning and fixed granularity evidence.We release SUnsET and our generation code to the public.1
Dustin Wright 0001, Zain Muhammad Mujahid, Lu Wang 0008, Isabelle Augenstein, David Jurgens
EMNLP4
2025 Graph-Guided Textual Explanation Generation Framework
abstract
Natural language explanations (NLEs) are commonly used to provide plausible free-text explanations of a model's reasoning about its predictions.However, recent work has questioned their faithfulness, as they may not accurately reflect the model's internal reasoning process regarding its predicted answer.In contrast, highlight explanations-input fragments critical for the model's predicted answers-exhibit measurable faithfulness.Building on this foundation, we propose G-TEx, a Graph-Guided Textual Explanation Generation framework designed to enhance the faithfulness of NLEs.Specifically, highlight explanations are first extracted as faithful cues reflecting the model's reasoning logic toward answer prediction.They are subsequently encoded through a graph neural network layer to guide the NLE generation, which aligns the generated explanations with the model's underlying reasoning toward the predicted answer.Experiments on both encoder-decoder and decoder-only models across three reasoning datasets demonstrate that G-TEx improves NLE faithfulness by up to 12.18% compared to baseline methods.Additionally, G-TEx generates NLEs with greater semantic and lexical similarity to human-written ones.Human evaluations show that G-TEx can decrease redundant content and enhance the overall quality of NLEs.Our work presents a novel method for explicitly guiding NLE generation to enhance faithfulness, serving as a foundation for addressing broader criteria in NLE and generated text.
Shuzhou Yuan, Ran Zhang 0013, Michael Färber 0001, Steffen Eger, Pepa Atanasova, Isabelle Augenstein
EMNLP7
2025 Investigating Human Values in Online Communities
abstract
Nadav Borenstein, Arnav Arora, Lucie-Aimée Kaffee, Isabelle Augenstein. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Nadav Borenstein, Arnav Arora, Lucie-Aimée Kaffee, Isabelle Augenstein
NAACL (Long Papers)4
2025 Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations
abstract
Yong Cao, Haijiang Liu, Arnav Arora, Isabelle Augenstein, Paul Röttger, Daniel Hershcovich. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yong Cao 0001, Arnav Arora, Isabelle Augenstein, Paul Röttger, Daniel Hershcovich
NAACL (Long Papers)4
2025 Measuring and Benchmarking Large Language Models' Capabilities to Generate Persuasive Language
abstract
Amalie Brogaard Pauli, Isabelle Augenstein, Ira Assent. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Amalie Brogaard Pauli, Isabelle Augenstein, Ira Assent
NAACL (Long Papers)2
2025 Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
abstract
Jingyi Sun, Pepa Atanasova, Isabelle Augenstein. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Pepa Atanasova, Isabelle Augenstein
NAACL (Long Papers)3
2025 Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking
abstract
Evaluating the truthfulness of online content is critical for combating misinformation. This study examines the efficiency and effectiveness of crowdsourced truthfulness assessments through a comparative analysis of two approaches: one involving full-length webpages as evidence for each claim, and another using summaries for each evidence document generated with an LLM. Using an A/B testing setting, we engage a diverse pool of participants tasked with evaluating the truthfulness of statements under these conditions.
Kevin Roitero, Dustin Wright 0001, Michael Soprano, Isabelle Augenstein, Stefano Mizzaro
SIGIR4
2024 What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular Languages
abstract
Nadav Borenstein, Anej Svete, Robin Chan, Josef Valvoda, Franz Nowak, Isabelle Augenstein, Eleanor Chodroff, Ryan Cotterell. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Nadav Borenstein, Anej Svete, Robin Shing Moon Chan, Josef Valvoda, Franz Nowak, Isabelle Augenstein, Eleanor Chodroff, Ryan Cotterell
ACL (1)6
2024 Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
abstract
Language Models (LMs) acquire parametric knowledge from their training process, embedding it within their weights.The increasing scalability of LMs, however, poses significant challenges for understanding a model's inner workings and further for updating or correcting this embedded knowledge without the significant cost of retraining.This underscores the importance of unveiling exactly what knowledge is stored and its association with specific model components.Instance Attribution (IA) and Neuron Attribution (NA) offer insights into this training-acquired knowledge, though they have not been compared systematically.Our study introduces a novel evaluation framework to quantify and compare the knowledge revealed by IA and NA.To align the results of the methods we introduce the attribution method NA-Instances to apply NA for retrieving influential training instances, and IA-Neurons to discover important neurons of influential instances discovered by IA.We further propose a comprehensive list of faithfulness tests to evaluate the comprehensiveness and sufficiency of the explanations provided by both methods.Through extensive experiments and analysis, we demonstrate that NA generally reveals more diverse and comprehensive information regarding the LM's parametric knowledge compared to IA.Nevertheless, IA provides unique and valuable insights into the LM's parametric knowledge, which are not revealed by NA.Our findings further suggest the potential of a synergistic approach of combining the diverse findings of IA and NA for a more holistic understanding of an LM's parametric knowledge.
Haeun Yu, Pepa Atanasova, Isabelle Augenstein
ACL (1)3
2024 Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models
abstract
Recent studies of the emergent capabilities of transformer-based Natural Language Understanding (NLU) models have indicated that they have an understanding of lexical and compositional semantics.We provide evidence that suggests these claims should be taken with a grain of salt: we find that state-of-the-art Natural Language Inference (NLI) models are sensitive towards minor semantics preserving surface-form variations, which lead to sizable inconsistent model decisions during inference.Notably, this behaviour differs from valid and in-depth comprehension of compositional semantics, however does neither emerge when evaluating model accuracy on standard benchmarks nor when probing for syntactic, monotonic, and logically robust reasoning.We propose a novel framework to measure the extent of semantic sensitivity.To this end, we evaluate NLI models on adversarially generated examples containing minor semantics-preserving surface-form input noise.This is achieved using conditional text generation, with the explicit condition that the NLI model predicts the relationship between the original and adversarial inputs as a symmetric equivalence entailment.We systematically study the effects of the phenomenon across NLI models for in-and outof domain settings.Our experiments show that semantic sensitivity causes performance degradations of 12.92% and 23.71% average over in-and out-of-domain settings, respectively.We further perform ablation studies, analysing this phenomenon across models, datasets, and variations in inference and show that semantic sensitivity can lead to major inconsistency within model predictions.
Erik Arakelyan, Zhaoqi Liu, Isabelle Augenstein
EACL (1)3
2024 Social Bias Probing: Fairness Benchmarking for Language Models
abstract
While the impact of social biases in language models has been recognized, prior methods for bias evaluation have been limited to binary association tests on small datasets, limiting our understanding of bias complexities. This paper proposes a novel framework for probing language models for social biases by assessing disparate treatment, which involves treating individuals differently according to their affiliation with a sensitive demographic group. We curate SOFA, a large-scale benchmark designed to address the limitations of existing fairness collections. SOFA expands the analysis beyond the binary comparison of stereotypical versus anti-stereotypical identities to include a diverse range of identities and stereotypes. Comparing our methodology with existing benchmarks, we reveal that biases within language models are more nuanced than acknowledged, indicating a broader scope of encoded biases than previously recognized. Benchmarking LMs on SOFA, we expose how identities expressing different religions lead to the most pronounced disparate treatments across all models. Finally, our findings indicate that real-life adversities faced by various groups such as women and people with disabilities are mirrored in the behavior of these models.
Marta Marchiori Manerba, Karolina Stanczak, Riccardo Guidotti, Isabelle Augenstein
EMNLP4
2024 Can Transformers Learn n-gram Language Models?
abstract
Much theoretical work has described the ability of transformers to represent formal languages.However, linking theoretical results to empirical performance is not straightforward due to the complex interplay between the architecture, the learning algorithm, and training data.To test whether theoretical lower bounds imply learnability of formal languages, we turn to recent work relating transformers to n-gram language models (LMs).We study transformers' ability to learn random n-gram LMs of two kinds: ones with arbitrary next-symbol probabilities and ones where those are defined with shared parameters.We find that classic estimation techniques for n-gram LMs such as add-λ smoothing outperform transformers on the former, while transformers perform better on the latter, outperforming methods specifically designed to learn n-gram LMs.github.com/rycolab/learning-ngrams
Anej Svete, Nadav Borenstein, Mike Zhou, Isabelle Augenstein, Ryan Cotterell
EMNLP4
2023 A Latent-Variable Model for Intrinsic Probing
abstract
The success of pre-trained contextualized representations has prompted researchers to analyze them for the presence of linguistic information. Indeed, it is natural to assume that these pre-trained representations do encode some level of linguistic knowledge as they have brought about large empirical improvements on a wide variety of NLP tasks, which suggests they are learning true linguistic generalization. In this work, we focus on intrinsic probing, an analysis technique where the goal is not only to identify whether a representation encodes a linguistic attribute but also to pinpoint where this attribute is encoded. We propose a novel latent-variable formulation for constructing intrinsic probes and derive a tractable variational approximation to the log-likelihood. Our results show that our model is versatile and yields tighter mutual information estimates than two intrinsic probes previously proposed in the literature. Finally, we find empirical evidence that pre-trained representations develop a cross-lingually entangled notion of morphosyntax.
Karolina Stanczak, Lucas Torroba Hennigen, Adina Williams, Ryan Cotterell, Isabelle Augenstein
AAAI5
2023 Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection
abstract
Stance Detection is concerned with identifying the attitudes expressed by an author towards a target of interest. This task spans a variety of domains ranging from social media opinion identification to detecting the stance for a legal claim. However, the framing of the task varies within these domains, in terms of the data collection protocol, the label dictionary and the number of available annotations. Furthermore, these stance annotations are significantly imbalanced on a per-topic and inter-topic basis. These make multi-domain stance detection a challenging task, requiring standardization and domain adaptation. To overcome this challenge, we propose Topic Efficient StancE Detection (TESTED), consisting of a topic-guided diversity sampling technique and a contrastive objective that is used for fine-tuning a stance classifier. We evaluate the method on an existing benchmark of 16 datasets with in-domain, i.e. all topics seen and out-of-domain, i.e. unseen topics, experiments. The results show that our method outperforms the state-of-the-art with an average of 3.5 F1 points increase in-domain, and is more generalizable with an averaged increase of 10.2 F1 on out-of-domain evaluation while using ≤ 10% of the training data. We show that our sampling technique mitigates both inter- and per-topic class imbalances. Finally, our analysis demonstrates that the contrastive learning objective allows the model a more pronounced segmentation of samples with varying labels.
Erik Arakelyan, Arnav Arora, Isabelle Augenstein
ACL (1)3
2023 Multilingual Event Extraction from Historical Newspaper Adverts
abstract
NLP methods can aid historians in analyzing textual materials in greater volumes than manually feasible.Developing such methods poses substantial challenges though.First, acquiring large, annotated historical datasets is difficult, as only domain experts can reliably label them.Second, most available off-the-shelf NLP models are trained on modern language texts, rendering them significantly less effective when applied to historical corpora.This is particularly problematic for less well studied tasks, and for languages other than English.This paper addresses these challenges while focusing on the under-explored task of event extraction from a novel domain of historical texts.We introduce a new multilingual dataset in English, French, and Dutch composed of newspaper ads from the early modern colonial period reporting on enslaved people who liberated themselves from enslavement.We find that: 1) even with scarce annotated data, it is possible to achieve surprisingly good results by formulating the problem as an extractive QA task and leveraging existing datasets and models for modern languages; and 2) cross-lingual low-resource learning for historical languages is highly challenging, and machine translation of the historical datasets to the considered target languages is, in practice, often the best-performing solution.
Nadav Borenstein, Natalia da Silva Perez, Isabelle Augenstein
ACL (1)3
2023 PHD: Pixel-Based Language Modeling of Historical Documents
abstract
The digitisation of historical documents has provided historians with unprecedented research opportunities. Yet, the conventional approach to analysing historical documents involves converting them from images to text using OCR, a process that overlooks the potential benefits of treating them as images and introduces high levels of noise. To bridge this gap, we take advantage of recent advancements in pixel-based language models trained to reconstruct masked patches of pixels instead of predicting token distributions. Due to the scarcity of real historical scans, we propose a novel method for generating synthetic scans to resemble real historical documents. We then pre-train our model, PHD, on a combination of synthetic scans and real historical newspapers from the 1700-1900 period. Through our experiments, we demonstrate that PHD exhibits high proficiency in reconstructing masked image patches and provide evidence of our model's noteworthy language understanding capabilities. Notably, we successfully apply our model to a historical QA task, highlighting its usefulness in this domain.
Nadav Borenstein, Phillip Rust, Desmond Elliott, Isabelle Augenstein
EMNLP4
2023 Explaining Interactions Between Text Spans
abstract
Reasoning over spans of tokens from different parts of the input is essential for natural language understanding (NLU) tasks such as fact-checking (FC), machine reading comprehension (MRC) or natural language inference (NLI).However, existing highlight-based explanations primarily focus on identifying individual important tokens or interactions only between adjacent tokens or tuples of tokens.Most notably, there is a lack of annotations capturing the human decision-making process w.r.t. the necessary interactions for informed decision-making in such tasks.To bridge this gap, we introduce SpanEx, a multi-annotator dataset of human span interaction explanations for two NLU tasks: NLI and FC.We then investigate the decision-making processes of multiple fine-tuned large language models in terms of the employed connections between spans in separate parts of the input and compare them to the human reasoning processes.Finally, we present a novel community detection based unsupervised method to extract such interaction explanations from a model's inner workings.1
Sagnik Ray Choudhury, Pepa Atanasova, Isabelle Augenstein
EMNLP3
2023 Why Should This Article Be Deleted? Transparent Stance Detection in Multilingual Wikipedia Editor Discussions
abstract
The moderation of content on online platforms is usually non-transparent.On Wikipedia, however, this discussion is carried out publicly and editors are encouraged to use the content moderation policies as explanations for making moderation decisions.Currently, only a few comments explicitly mention those policies -20% of the English ones, but as few as 2% of the German and Turkish comments.To aid in this process of understanding how content is moderated, we construct a novel multilingual dataset of Wikipedia editor discussions along with their reasoning in three languages.The dataset contains the stances of the editors (keep, delete, merge, comment), along with the stated reason, and a content moderation policy, for each edit decision.We demonstrate that stance and corresponding reason (policy) can be predicted jointly with a high degree of accuracy, adding transparency to the decision-making process.We release both our joint prediction models and the multilingual content moderation dataset for further research on automated transparent content moderation.
Lucie-Aimée Kaffee, Arnav Arora, Isabelle Augenstein
EMNLP3
2023 People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
abstract
NLP models are used in a variety of critical social computing tasks, such as detecting sexist, racist, or otherwise hateful content.Therefore, it is imperative that these models are robust to spurious features.Past work has attempted to tackle such spurious features using training data augmentation, including Counterfactually Augmented Data (CADs).CADs introduce minimal changes to existing training data points and flip their labels; training on them may reduce model dependency on spurious features.However, manually generating CADs can be time-consuming and expensive.Hence in this work, we assess if this task can be automated using generative NLP models.We automatically generate CADs using Polyjuice, Chat-GPT, and Flan-T5, and evaluate their usefulness in improving model robustness compared to manually-generated CADs.By testing both model performance on multiple out-of-domain test sets and individual data point efficacy, our results show that while manual CADs are still the most effective, CADs generated by Chat-GPT come a close second.One key reason for the lower performance of automated methods is that the changes they introduce are often insufficient to flip the original label. 1Warning: This paper has instances of hateful and sexist language to serve as examples.
Indira Sen, Dennis Assenmacher, Mattia Samory, Isabelle Augenstein, Wil M. P. van der Aalst, Claudia Wagner 0001
EMNLP4
2023 Adapting Neural Link Predictors for Data-Efficient Complex Query Answering
abstract
Answering complex queries on incomplete knowledge graphs is a challenging task where a model needs to answer complex logical queries in the presence of missing knowledge. Prior work in the literature has proposed to address this problem by designing architectures trained end-to-end for the complex query answering task with a reasoning process that is hard to interpret while requiring data and resource-intensive training. Other lines of research have proposed re-using simple neural link predictors to answer complex queries, reducing the amount of training data by orders of magnitude while providing interpretable answers. The neural link predictor used in such approaches is not explicitly optimised for the complex query answering task, implying that its scores are not calibrated to interact together. We propose to address these problems via CQD$^{\mathcal{A}}$, a parameter-efficient score \emph{adaptation} model optimised to re-calibrate neural link prediction scores for the complex query answering task. While the neural link predictor is frozen, the adaptation component -- which only increases the number of model parameters by $0.03\%$ -- is trained on the downstream complex query answering task. Furthermore, the calibration component enables us to support reasoning over queries that include atomic negations, which was previously impossible with link predictors. In our experiments, CQD$^{\mathcal{A}}$ produces significantly more accurate results than current state-of-the-art methods, improving from $34.4$ to $35.1$ Mean Reciprocal Rank values averaged across all datasets and query types while using $\leq 30\%$ of the available training query types. We further show that CQD$^{\mathcal{A}}$ is data-efficient, achieving competitive results with only $1\%$ of the complex training queries and robust in out-of-domain evaluations. Source code and datasets are available at https://github.com/EdinburghNLP/adaptive-cqd.
Erik Arakelyan, Pasquale Minervini, Daniel Daza, Michael Cochez, Isabelle Augenstein
NeurIPS5
2023 Preface: Special issue on NLP approaches to offensive content online
abstract
We are delighted to present the Special Issue on NLP Approaches to Offensive Content Online published in the Journal of Natural Language Engineering issue 29.6. We are happy to have received a total of 26 submissions to the special issue evidencing the interest of the NLP community in this topic. Our guest editorial board comprised of international experts in the field has worked hard to review all submissions over multiple rounds of peer review. Ultimately, we accepted nine articles to appear in this special issue.
Marcos Zampieri, Isabelle Augenstein, Siddharth Krishnan, Joshua Melton, Preslav Nakov
Nat. Lang. Eng.2
2022 Diagnostics-Guided Explanation Generation
abstract
Explanations shed light on a machine learning model's rationales and can aid in identifying deficiencies in its reasoning process. Explanation generation models are typically trained in a supervised way given human explanations. When such annotations are not available, explanations are often selected as those portions of the input that maximise a downstream task's performance, which corresponds to optimising an explanation's Faithfulness to a given model. Faithfulness is one of several so-called diagnostic properties, which prior work has identified as useful for gauging the quality of an explanation without requiring annotations. Other diagnostic properties are Data Consistency, which measures how similar explanations are for similar input instances, and Confidence Indication, which shows whether the explanation reflects the confidence of the model. In this work, we show how to directly optimise for these diagnostic properties when training a model to generate sentence-level explanations, which markedly improves explanation quality, agreement with human rationales, and downstream task performance on three complex reasoning tasks.
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle Augenstein
AAAI4
2022 Few-Shot Cross-Lingual Stance Detection with Sentiment-Based Pre-training
abstract
The goal of stance detection is to determine the viewpoint expressed in a piece of text towards a target. These viewpoints or contexts are often expressed in many different languages depending on the user and the platform, which can be a local news outlet, a social media platform, a news forum, etc. Most research on stance detection, however, has been limited to working with a single language and on a few limited targets, with little work on cross-lingual stance detection. Moreover, non-English sources of labelled data are often scarce and present additional challenges. Recently, large multilingual language models have substantially improved the performance on many non-English tasks, especially such with a limited number of examples. This highlights the importance of model pre-training and its ability to learn from few examples. In this paper, we present the most comprehensive study of cross-lingual stance detection to date: we experiment with 15 diverse datasets in 12 languages from 6 language families, and with 6 low-resource evaluation settings each. For our experiments, we build on pattern-exploiting training (PET), proposing the addition of a novel label encoder to simplify the verbalisation procedure. We further propose sentiment-based generation of stance data for pre-training, which shows sizeable improvement of more than 6% F1 absolute in few-shot learning settings compared to several strong baselines.
Momchil Hardalov, Arnav Arora, Preslav Nakov, Isabelle Augenstein
AAAI4
2022 Generating Scientific Claims for Zero-Shot Scientific Fact Checking
abstract
Dustin Wright, David Wadden, Kyle Lo, Bailey Kuehl, Arman Cohan, Isabelle Augenstein, Lucy Wang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Dustin Wright 0001, Dave Wadden, Kyle Lo, Bailey Kuehl, Arman Cohan, Isabelle Augenstein, Lucy Lu Wang
ACL (1)6
2022 Can Edge Probing Tests Reveal Linguistic Knowledge in QA Models?
abstract
There have been many efforts to try to understand what grammatical knowledge (e.g., ability to understand the part of speech of a token) is encoded in large pre-trained language models (LM). This is done through ‘Edge Probing’ (EP) tests: supervised classification tasks to predict the grammatical properties of a span (whether it has a particular part of speech) using only the token representations coming from the LM encoder. However, most NLP applications fine-tune these LM encoders for specific tasks. Here, we ask: if an LM is fine-tuned, does the encoding of linguistic information in it change, as measured by EP tests? Specifically, we focus on the task of Question Answering (QA) and conduct experiments on multiple datasets. We find that EP test results do not change significantly when the fine-tuned model performs well or in adversarial situations where the model is forced to learn wrong correlations. From a similar finding, some recent papers conclude that fine-tuning does not change linguistic knowledge in encoders but they do not provide an explanation. We find that EP models are susceptible to exploiting spurious correlations in the EP datasets. When this dataset bias is corrected, we do see an improvement in the EP test results as expected.
Sagnik Ray Choudhury, Nikita Bhutani, Isabelle Augenstein
COLING3
2022 Machine Reading, Fast and Slow: When Do Models "Understand" Language?
abstract
Two of the most fundamental issues in Natural Language Understanding (NLU) at present are: (a) how it can established whether deep learning-based models score highly on NLU benchmarks for the ”right” reasons; and (b) what those reasons would even be. We investigate the behavior of reading comprehension models with respect to two linguistic ”skills”: coreference resolution and comparison. We propose a definition for the reasoning steps expected from a system that would be ”reading slowly”, and compare that with the behavior of five models of the BERT family of various sizes, observed through saliency scores and counterfactual explanations. We find that for comparison (but not coreference) the systems based on larger encoders are more likely to rely on the ”right” information, but even they struggle with generalization, suggesting that they still learn specific lexical patterns rather than the general principles of comparison.
Sagnik Ray Choudhury, Anna Rogers, Isabelle Augenstein
COLING3
2022 Multi3Generation: Multitask, Multilingual, Multimodal Language Generation
abstract
This paper presents the Multitask, Multilingual, Multimodal Language Generation COST Action – Multi3Generation (CA18231), an interdisciplinary network of research groups working on different aspects of language generation. This “meta-paper” will serve as reference for citations of the Action in future publications. It presents the objectives, challenges and a the links for the achieved outcomes.
Anabela Barreiro, José Guilherme Camargo de Souza, Albert Gatt, Mehul Bhatt, Elena Lloret, Aykut Erdem, Dimitra Gkatzia, Helena Moniz, Irene Russo, Fábio N. Kepler, Iacer Calixto, Marcin Paprzycki, François Portet, Isabelle Augenstein, Mirela Alhasani
EAMT14
2022 Modeling Information Change in Science Communication with Semantically Matched Paraphrases
abstract
Whether the media faithfully communicate scientific information has long been a core issue to the science community.Automatically identifying paraphrased scientific findings could enable large-scale tracking and analysis of information changes in the science communication process, but this requires systems to understand the similarity between scientific information across multiple domains.To this end, we present the SCIENTIFIC PARAPHRASE AND INFOR-MATION CHANGE DATASET (SPICED), the first paraphrase dataset of scientific findings annotated for degree of information change.SPICED contains 6,000 scientific finding pairs extracted from news stories, social media discussions, and full texts of original papers.We demonstrate that SPICED poses a challenging task and that models trained on SPICED improve downstream performance on evidence retrieval for fact checking of real-world scientific claims.Finally, we show that models trained on SPICED can reveal large-scale trends in the degrees to which people and organizations faithfully communicate new scientific findings.Data, code, and pre-trained models are available at http://www.copenlu.com/publication/ 2022_emnlp_wright/.
Dustin Wright 0001, Jiaxin Pei, David Jurgens, Isabelle Augenstein
EMNLP4
2022 Neighborhood Contrastive Learning for Scientific Document Representations with Citation Embeddings
abstract
Learning scientific document representations can be substantially improved through contrastive learning objectives, where the challenge lies in creating positive and negative training samples that encode the desired similarity semantics.Prior work relies on discrete citation relations to generate contrast samples.However, discrete citations enforce a hard cutoff to similarity.This is counter-intuitive to similarity-based learning and ignores that scientific papers can be very similar despite lacking a direct citation -a core problem of finding related research.Instead, we use controlled nearest neighbor sampling over citation graph embeddings for contrastive learning.This control allows us to learn continuous similarity, to sample hard-to-learn negatives and positives, and also to avoid collisions between negative and positive samples by controlling the sampling margin between them.The resulting method SciNCL outperforms the state-of-theart on the SciDocs benchmark.Furthermore, we demonstrate that it can train (or tune) language models sample-efficiently and that it can be combined with recent training-efficient methods.Perhaps surprisingly, even training a general-domain language model this way outperforms baselines pretrained in-domain.
Malte Ostendorff, Nils Rethmeier, Isabelle Augenstein, Bela Gipp, Georg Rehm
EMNLP3
2022 Counterfactually Augmented Data and Unintended Bias: The Case of Sexism and Hate Speech Detection
abstract
Indira Sen, Mattia Samory, Claudia Wagner, Isabelle Augenstein. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Indira Sen, Mattia Samory, Claudia Wagner 0001, Isabelle Augenstein
NAACL-HLT4
2022 Same Neurons, Different Languages: Probing Morphosyntax in Multilingual Pre-trained Models
abstract
Karolina Stanczak, Edoardo Ponti, Lucas Torroba Hennigen, Ryan Cotterell, Isabelle Augenstein. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Karolina Stanczak, Edoardo Maria Ponti, Lucas Torroba Hennigen, Ryan Cotterell, Isabelle Augenstein
NAACL-HLT5
2022 TempEL: Linking Dynamically Evolving and Newly Emerging Entities
abstract
In our continuously evolving world, entities change over time and new, previously non-existing or unknown, entities appear. We study how this evolutionary scenario impacts the performance on a well established entity linking (EL) task. For that study, we introduce TempEL, an entity linking dataset that consists of time-stratified English Wikipedia snapshots from 2013 to 2022, from which we collect both anchor mentions of entities, and these target entities’ descriptions. By capturing such temporal aspects, our newly introduced TempEL resource contrasts with currently existing entity linking datasets, which are composed of fixed mentions linked to a single static version of a target Knowledge Base (e.g., Wikipedia 2010 for CoNLL-AIDA). Indeed, for each of our collected temporal snapshots, TempEL contains links to entities that are continual, i.e., occur in all of the years, as well as completely new entities that appear for the first time at some point. Thus, we enable to quantify the performance of current state-of-the-art EL models for: (i) entities that are subject to changes over time in their Knowledge Base descriptions as well as their mentions’ contexts, and (ii) newly created entities that were previously non-existing (e.g., at the time the EL model was trained). Our experimental results show that in terms of temporal performance degradation, (i) continual entities suffer a decrease of up to 3.1% EL accuracy, while (ii) for new entities this accuracy drop is up to 17.9%. This highlights the challenge of the introduced TempEL dataset and opens new research prospects in the area of time-evolving entity disambiguation.
Klim Zaporojets, Lucie-Aimée Kaffee, Johannes Deleu, Thomas Demeester, Chris Develder, Isabelle Augenstein
NeurIPS6
2022 Joint emotion label space modeling for affect lexica
Luna De Bruyne, Pepa Atanasova, Isabelle Augenstein
Comput. Speech Lang.3
2022 Fact Checking with Insufficient Evidence
abstract
Abstract Automating the fact checking (FC) process relies on information obtained from external sources. In this work, we posit that it is crucial for FC models to make veracity predictions only when there is sufficient evidence and otherwise indicate when it is not enough. To this end, we are the first to study what information FC models consider sufficient by introducing a novel task and advancing it with three main contributions. First, we conduct an in-depth empirical analysis of the task with a new fluency-preserving method for omitting information from the evidence at the constituent and sentence level. We identify when models consider the remaining evidence (in)sufficient for FC, based on three trained models with different Transformer architectures and three FC datasets. Second, we ask annotators whether the omitted evidence was important for FC, resulting in a novel diagnostic dataset, SufficientFacts1, for FC with omitted evidence. We find that models are least successful in detecting missing evidence when adverbial modifiers are omitted (21% accuracy), whereas it is easiest for omitted date modifiers (63% accuracy). Finally, we propose a novel data augmentation strategy for contrastive self-learning of missing evidence by employing the proposed omission method combined with tri-training. It improves performance for Evidence Sufficiency Prediction by up to 17.8 F1 score, which in turn improves FC performance by up to 2.6 F1 score.
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle Augenstein
Trans. Assoc. Comput. Linguistics4
2022 A Neighborhood Framework for Resource-Lean Content Flagging
abstract
Abstract We propose a novel framework for cross- lingual content flagging with limited target- language data, which significantly outperforms prior work in terms of predictive performance. The framework is based on a nearest-neighbor architecture. It is a modern instantiation of the vanilla k-nearest neighbor model, as we use Transformer representations in all its components. Our framework can adapt to new source- language instances, without the need to be retrained from scratch. Unlike prior work on neighborhood-based approaches, we encode the neighborhood information based on query– neighbor interactions. We propose two encoding schemes and we show their effectiveness using both qualitative and quantitative analysis. Our evaluation results on eight languages from two different datasets for abusive language detection show sizable improvements of up to 9.5 F1 points absolute (for Italian) over strong baselines. On average, we achieve 3.6 absolute F1 points of improvement for the three languages in the Jigsaw Multilingual dataset and 2.14 points for the WUL dataset.
Sheikh Muhammad Sarwar, Dimitrina Zlatkova, Momchil Hardalov, Yoan Dinkov, Isabelle Augenstein, Preslav Nakov
Trans. Assoc. Comput. Linguistics5
2021 Does Typological Blinding Impede Cross-Lingual Sharing?
abstract
Bridging the performance gap between highand low-resource languages has been the focus of much previous work.Typological features from databases such as the World Atlas of Language Structures (WALS) are a prime candidate for this, as such data exists even for very low-resource languages.However, previous work has only found minor benefits from using typological information.Our hypothesis is that a model trained in a cross-lingual setting will pick up on typological cues from the input data, thus overshadowing the utility of explicitly using such features.We verify this hypothesis by blinding a model to typological information, and investigate how cross-lingual sharing and performance is impacted.Our model is based on a cross-lingual architecture in which the latent weights governing the sharing between languages is learnt during training.We show that (i) preventing this model from exploiting typology severely reduces performance, while a control experiment reaffirms that (ii) encouraging sharing according to typology somewhat improves performance.
Johannes Bjerva, Isabelle Augenstein
EACL2
2021 Semi-Supervised Exaggeration Detection of Health Science Press Releases
abstract
Public trust in science depends on honest and factual communication of scientific papers.However, recent studies have demonstrated a tendency of news media to misrepresent scientific papers by exaggerating their findings.Given this, we present a formalization of and study into the problem of exaggeration detection in science communication.While there are an abundance of scientific papers and popular media articles written about them, very rarely do the articles include a direct link to the original paper, making data collection challenging.We address this by curating a set of labeled press release/abstract pairs from existing expert annotated studies on exaggeration in press releases of scientific papers suitable for benchmarking the performance of machine learning models on the task.Using limited data from this and previous studies on exaggeration detection in science, we introduce MT-PET, a multi-task version of Pattern Exploiting Training (PET), which leverages knowledge from complementary clozestyle QA tasks to improve few-shot learning.We demonstrate that MT-PET outperforms PET and supervised learning both when data is limited, as well as when there is an abundance of data for the main task. 1
Dustin Wright 0001, Isabelle Augenstein
EMNLP (1)2
2021 Cross-Domain Label-Adaptive Stance Detection
abstract
Stance detection concerns the classification of a writer's viewpoint towards a target.There are different task variants, e.g., stance of a tweet vs. a full article, or stance with respect to a claim vs. an (implicit) topic.Moreover, task definitions vary, which includes the label inventory, the data collection, and the annotation protocol.All these aspects hinder cross-domain studies, as they require changes to standard domain adaptation approaches.In this paper, we perform an in-depth analysis of 16 stance detection datasets, and we explore the possibility for cross-domain learning from them.Moreover, we propose an end-to-end unsupervised framework for outof-domain prediction of unseen, user-defined labels.In particular, we combine domain adaptation techniques such as mixture of experts and domain-adversarial training with label embeddings, and we demonstrate sizable performance gains over strong baselines, both (i) indomain, i.e., for seen targets, and (ii) out-ofdomain, i.e., for unseen targets.Finally, we perform an exhaustive analysis of the crossdomain results, and we highlight the important factors influencing the model performance.
Momchil Hardalov, Arnav Arora, Preslav Nakov, Isabelle Augenstein
EMNLP (1)4
2021 How Does Counterfactually Augmented Data Impact Models for Social Computing Constructs?
abstract
As NLP models are increasingly deployed in socially situated settings such as online abusive content detection, it is crucial to ensure that these models are robust.One way of improving model robustness is to generate counterfactually augmented data (CAD) for training models that can better learn to distinguish between core features and data artifacts.While models trained on this type of data have shown promising out-of-domain generalizability, it is still unclear what the sources of such improvements are.We investigate the benefits of CAD for social NLP models by focusing on three social computing constructs -sentiment, sexism, and hate speech.Assessing the performance of models trained with and without CAD across different types of datasets, we find that while models trained on CAD show lower in-domain performance, they generalize better out-of-domain.We unpack this apparent discrepancy using machine explanations and find that CAD reduces model reliance on spurious features.Leveraging a novel typology of CAD to analyze their relationship with model performance, we find that CAD which acts on the construct directly or a diverse set of CAD leads to higher performance.
Indira Sen, Mattia Samory, Fabian Flöck, Claudia Wagner 0001, Isabelle Augenstein
EMNLP (1)5
2021 Multi-Hop Fact Checking of Political Claims
abstract
Recent work has proposed multi-hop models and datasets for studying complex natural language reasoning. One notable task requiring multi-hop reasoning is fact checking, where a set of connected evidence pieces leads to the final verdict of a claim. However, existing datasets either do not provide annotations for gold evidence pages, or the only dataset which does (FEVER) mostly consists of claims which can be fact-checked with simple reasoning and is constructed artificially. Here, we study more complex claim verification of naturally occurring claims with multiple hops over interconnected evidence chunks. We: 1) construct a small annotated dataset, PolitiHop, of evidence sentences for claim verification; 2) compare it to existing multi-hop datasets; and 3) study how to transfer knowledge from more extensive in- and out-of-domain resources to PolitiHop. We find that the task is complex and achieve the best performance with an architecture that specifically models reasoning over evidence pieces in combination with in-domain transfer learning.
Wojciech Ostrowski, Arnav Arora, Pepa Atanasova, Isabelle Augenstein
IJCAI4
2021 Time-aware evidence ranking for fact-checking
abstract
Truth can vary over time. Fact-checking decisions on claim veracity should therefore take into account temporal information of both the claim and supporting or refuting evidence. In this work, we investigate the hypothesis that the timestamp of a Web page is crucial to how it should be ranked for a given claim. We delineate four temporal ranking methods that constrain evidence ranking differently and simulate hypothesis-specific evidence rankings given the evidence timestamps as gold standard. Evidence ranking in three fact-checking models is ultimately optimized using a learning-to-rank loss function. Our study reveals that time-aware evidence ranking not only surpasses relevance assumptions based purely on semantic similarity or position in a search results list, but also improves veracity predictions of time-sensitive claims in particular.
Liesbeth Allein, Isabelle Augenstein, Marie-Francine Moens
J. Web Semant.2
2020 Back to the Future - Temporal Adaptation of Text Representations
abstract
Language evolves over time in many ways relevant to natural language processing tasks. For example, recent occurrences of tokens 'BERT' and 'ELMO' in publications refer to neural network architectures rather than persons. This type of temporal signal is typically overlooked, but is important if one aims to deploy a machine learning model over an extended period of time. In particular, language evolution causes data drift between time-steps in sequential decision-making tasks. Examples of such tasks include prediction of paper acceptance for yearly conferences (regular intervals) or author stance prediction for rumours on Twitter (irregular intervals). Inspired by successes in computer vision, we tackle data drift by sequentially aligning learned representations. We evaluate on three challenging tasks varying in terms of time-scales, linguistic units, and domains. These tasks show our method outperforming several strong baselines, including using all available data. We argue that, due to its low computational expense, sequential alignment is a practical solution to dealing with language evolution.
Johannes Bjerva, Wouter M. Kouw, Isabelle Augenstein
AAAI3
2020 2kenize: Tying Subword Sequences for Chinese Script Conversion
abstract
Simplified Chinese to Traditional Chinese character conversion is a common preprocessing step in Chinese NLP.Despite this, current approaches have insufficient performance because they do not take into account that a simplified Chinese character can correspond to multiple traditional characters.Here, we propose a model that can disambiguate between mappings and convert between the two scripts.The model is based on subword segmentation, two language models, as well as a method for mapping between subword sequences.We further construct benchmark datasets for topic classification and script conversion.Our proposed method outperforms previous Chinese Character conversion approaches by 6 points in accuracy.These results are further confirmed in a downstream application, where 2kenize is used to convert pretraining dataset for topic classification.An error analysis reveals that our method's particular strengths are in dealing with code mixing and named entities.The code and dataset is available at https: //github.com/pranav-ust/2kenize
A Pranav 0001, Isabelle Augenstein
ACL2
2020 Generating Fact Checking Explanations
abstract
Most existing work on automated fact checking is concerned with predicting the veracity of claims based on metadata, social network spread, language used in claims, and, more recently, evidence supporting or denying claims.A crucial piece of the puzzle that is still missing is to understand how to automate the most elaborate part of the process -generating justifications for verdicts on claims.This paper provides the first study of how these explanations can be generated automatically based on available claim context, and how this task can be modelled jointly with veracity prediction.Our results indicate that optimising both objectives at the same time, rather than training them separately, improves the performance of a fact checking system.The results of a manual evaluation further suggest that the informativeness, coverage and overall quality of the generated explanations are also improved in the multi-task model.Claim: The last major oil spill from a drilling accident in America happened over 40 years ago in 1969.Ruling Comments: (...) The last major oil spill from a drilling accident in America happened over 40 years ago in 1969.(...) The largest in volume was the Santa Barbara spill of 1969 referenced by Murdock and Johnson, in which an estimated 100,000 barrels of oil spilled into the Pacific Ocean, according to the API.The Santa Barbara spill was so big it ranked seventh among the 10 largest oil spills caused by marine well blowouts in the world, the report states.Two other U.S. spills, both in 1970, rank eighth and 10th.Fourteen marine blowouts have taken place in the U.S. between 1969 and 2007.Six of them took place after 1990 and spilled a total of nearly 13,700 barrels.(...) We interviewed three scientists who said that the impact of a spill has little to do with its volume.Scientists have proven that spills far smaller than Santa Barbara's have been devastating.Justification: While the nation's largest oil well blowout did take place in 1969, it's not factually correct to call it the "last major oil spill".First of all, two of the largest blowouts in the world took place in the U. S. the following year.More importantly, experts agree that spills far smaller in volume to the 1969 disaster have been devastating.From a scientific perspective, Johnson's decision to single out the 1969 blowout as the last "major" one makes no sense.
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle Augenstein
ACL4
2020 A Diagnostic Study of Explainability Techniques for Text Classification
abstract
Recent developments in machine learning have introduced models that approach human performance at the cost of increased architectural complexity.Efforts to make the rationales behind the models' predictions transparent have inspired an abundance of new explainability techniques.Provided with an already trained model, they compute saliency scores for the words of an input instance.However, there exists no definitive guide on (i) how to choose such a technique given a particular application task and model architecture, and (ii) the benefits and drawbacks of using each such technique.In this paper, we develop a comprehensive list of diagnostic properties for evaluating existing explainability techniques.We then employ the proposed list to compare a set of diverse explainability techniques on downstream text classification tasks and neural network architectures.We also compare the saliency scores assigned by the explainability techniques with human annotations of salient input regions to find relations between a model's performance and the agreement of its rationales with human ones.Overall, we find that the gradient-based explanations perform best across tasks and model architectures, and we present further insights into the properties of the reviewed explainability techniques.
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle Augenstein
EMNLP (1)4
2020 Generating Label Cohesive and Well-Formed Adversarial Claims
abstract
Adversarial attacks reveal important vulnerabilities and flaws of trained models.One potent type of attack are universal adversarial triggers, which are individual n-grams that, when appended to instances of a class under attack, can trick a model into predicting a target class.However, for inference tasks such as fact checking, these triggers often inadvertently invert the meaning of instances they are inserted in.In addition, such attacks produce semantically nonsensical inputs, as they simply concatenate triggers to existing samples.Here, we investigate how to generate adversarial attacks against fact checking systems that preserve the ground truth meaning and are semantically valid.We extend the HotFlip attack algorithm used for universal trigger generation by jointly minimizing the target class loss of a fact checking model and the entailment class loss of an auxiliary natural language inference model.We then train a conditional language model to generate semantically valid statements, which include the found universal triggers.We find that the generated attacks maintain the directionality and semantic validity of the claim better than previous work.
Pepa Atanasova, Dustin Wright 0001, Isabelle Augenstein
EMNLP (1)3
2020 SubjQA: A Dataset for Subjectivity and Review Comprehension
abstract
Subjectivity is the expression of internal opinions or beliefs which cannot be objectively observed or verified, and has been shown to be important for sentiment analysis and wordsense disambiguation.Furthermore, subjectivity is an important aspect of user-generated data.In spite of this, subjectivity has not been investigated in contexts where such data is widespread, such as in question answering (QA).We develop a new dataset which allows us to investigate this relationship.We find that subjectivity is an important feature in the case of QA, albeit with more intricate interactions between subjectivity and QA performance than found in previous work on sentiment analysis.For instance, a subjective question may or may not be associated with a subjective answer.We release an English QA dataset (SUBJQA) based on customer reviews, containing subjectivity annotations for questions and answer spans across 6 domains.
Johannes Bjerva, Nikita Bhutani, Behzad Golshan, Wang Chiew Tan, Isabelle Augenstein
EMNLP (1)5
2020 Zero-Shot Cross-Lingual Transfer with Meta Learning
abstract
Learning what to share between tasks has become a topic of great importance, as strategic sharing of knowledge has been shown to improve downstream task performance.This is particularly important for multilingual applications, as most languages in the world are under-resourced.Here, we consider the setting of training models on multiple different languages at the same time, when little or no data is available for languages other than English.We show that this challenging setup can be approached using meta-learning: in addition to training a source language model, another model learns to select which training instances are the most beneficial to the first.We experiment using standard supervised, zero-shot cross-lingual, as well as fewshot cross-lingual settings for different natural language understanding tasks (natural language inference, question answering).Our extensive experimental setup demonstrates the consistent effectiveness of meta-learning for a total of 15 languages.We improve upon the state-of-the-art for zero-shot and few-shot NLI (on MultiNLI and XNLI) and QA (on the MLQA dataset).A comprehensive error analysis indicates that the correlation of typological features between languages can partly explain when parameter sharing learned via meta-learning is beneficial.
Farhad Nooralahzadeh, Giannis Bekoulis, Johannes Bjerva, Isabelle Augenstein
EMNLP (1)4
2020 Transformer Based Multi-Source Domain Adaptation
abstract
In practical machine learning settings, the data on which a model must make predictions often come from a different distribution than the data it was trained on.Here, we investigate the problem of unsupervised multi-source domain adaptation, where a model is trained on labelled data from multiple source domains and must make predictions on a domain for which no labelled data has been seen.Prior work with CNNs and RNNs has demonstrated the benefit of mixture of experts, where the predictions of multiple domain expert classifiers are combined; as well as domain adversarial training, to induce a domain agnostic representation space.Inspired by this, we investigate how such methods can be effectively applied to large pretrained transformer models.We find that domain adversarial training has an effect on the learned representations of these models while having little effect on their performance, suggesting that large transformer-based models are already relatively robust across domains.Additionally, we show that mixture of experts leads to significant performance improvements by comparing several variants of mixing functions, including one novel mixture based on attention.Finally, we demonstrate that the predictions of large pretrained transformer based domain experts are highly homogenous, making it challenging to learn effective functions for mixing their predictions.
Dustin Wright 0001, Isabelle Augenstein
EMNLP (1)2
2020 TX-Ray: Quantifying and Explaining Model-Knowledge Transfer in (Un-)Supervised NLP
abstract
While state-of-the-art NLP explainability (XAI) methods focus on explaining per-sample decisions in supervised end or probing tasks, this is insufficient to explain and quantify model knowledge transfer during (un-)supervised training. Thus, for TX-Ray, we modify the established computer vision explainability principle of ‘visualizing preferred inputs of neurons’ to make it usable for both NLP and for transfer analysis. This allows one to analyze, track and quantify how self- or supervised NLP models first build knowledge abstractions in pretraining (1), andthen transfer abstractions to a new domain (2), or adapt them during supervised finetuning (3) – see Fig. 1. TX-Ray expresses neurons as feature preference distributions to quantify fine-grained knowledge transfer or adaptation and guide human analysis. We find that, similar to Lottery Ticket based pruning, TX-Ray based pruning can improve test set generalization and that it can reveal how early stages of self-supervision automatically learn linguistic abstractions like parts-of-speech.
Nils Rethmeier, Vageesh Kumar Saxena, Isabelle Augenstein
UAI3
2019 Latent Multi-Task Architecture Learning
abstract
Multi-task learning (MTL) allows deep neural networks to learn from related tasks by sharing parameters with other networks. In practice, however, MTL involves searching an enormous space of possible parameter sharing architectures to find (a) the layers or subspaces that benefit from sharing, (b) the appropriate amount of sharing, and (c) the appropriate relative weights of the different task losses. Recent work has addressed each of the above problems in isolation. In this work we present an approach that learns a latent multi-task architecture that jointly addresses (a)–(c). We present experiments on synthetic data and data from OntoNotes 5.0, including four different tasks and seven different domains. Our extension consistently outperforms previous approaches to learning latent architectures for multi-task problems and achieves up to 15% average error reductions over common approaches to MTL.
Sebastian Ruder, Joachim Bingel, Isabelle Augenstein, Anders Søgaard
AAAI3
2019 Uncovering Probabilistic Implications in Typological Knowledge Bases
abstract
The study of linguistic typology is rooted in the implications we find between linguistic features, such as the fact that languages with object-verb word ordering tend to have postpositions. Uncovering such implications typically amounts to time-consuming manual processing by trained and experienced linguists, which potentially leaves key linguistic universals unexplored. In this paper, we present a computational model which successfully identifies known universals, including Greenberg universals, but also uncovers new ones, worthy of further linguistic investigation. Our approach outperforms baselines previously used for this problem, as well as a strong baseline from knowledge base population.
Johannes Bjerva, Yova Kementchedjhieva, Ryan Cotterell, Isabelle Augenstein
ACL (1)4
2019 Unsupervised Discovery of Gendered Language through Latent-Variable Modeling
abstract
Studying the ways in which language is gendered has long been an area of interest in sociolinguistics.Studies have explored, for example, the speech of male and female characters in film and the language used to describe male and female politicians.In this paper, we aim not to merely study this phenomenon qualitatively, but instead to quantify the degree to which the language used to describe men and women is different and, moreover, different in a positive or negative way.To that end, we introduce a generative latent-variable model that jointly represents adjective (or verb) choice, with its sentiment, given the natural gender of a head (or dependent) noun.We find that there are significant differences between descriptions of male and female nouns and that these differences align with common gender stereotypes: Positive adjectives used to describe women are more often related to their bodies than adjectives used to describe men.
Alexander Miserlis Hoyle, Lawrence Wolf-Sonkin, Hanna M. Wallach, Isabelle Augenstein, Ryan Cotterell
ACL (1)4
2019 MultiFC: A Real-World Multi-Domain Dataset for Evidence-Based Fact Checking of Claims
abstract
Isabelle Augenstein, Christina Lioma, Dongsheng Wang, Lucas Chaves Lima, Casper Hansen, Christian Hansen, Jakob Grue Simonsen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Isabelle Augenstein, Christina Lioma, Dongsheng Wang 0005, Lucas Chaves Lima, Casper Hansen, Christian Hansen 0004, Jakob Grue Simonsen
EMNLP/IJCNLP (1)1
2019 What Do Language Representations Really Represent?
abstract
A neural language model trained on a text corpus can be used to induce distributed representations of words, such that similar words end up with similar representations. If the corpus is multilingual, the same model can be used to learn distributed representations of languages, such that similar languages end up with similar representations. We show that this holds even when the multilingual corpus has been translated into English, by picking up the faint signal left by the source languages. However, just as it is a thorny problem to separate semantic from syntactic similarity in word representations, it is not obvious what type of similarity is captured by language representations. We investigate correlations and causal relationships between language representations learned from translations on one hand, and genetic, geographical, and several levels of structural similarity between languages on the other. Of these, structural similarity is found to correlate most strongly with language representation similarity, whereas genetic relationships—a convenient benchmark used for evaluation in previous work—appears to be a confounding factor. Apart from implications about translation effects, we see this more generally as a case where NLP and linguistic typology can interact and benefit one another.
Johannes Bjerva, Robert Östling, Maria Han Veiga, Jörg Tiedemann, Isabelle Augenstein
Comput. Linguistics5
2018 A strong baseline for question relevancy ranking
abstract
The best systems at the SemEval-16 and SemEval-17 community question answering shared tasks -a task that amounts to question relevancy ranking -involve complex pipelines and manual feature engineering.Despite this, many of these still fail at beating the IR baseline, i.e., the rankings provided by Google's search engine.We present a strong baseline for question relevancy ranking by training a simple multi-task feed forward network on a bag of 14 distance measures for the input question pair.This baseline model, which is fast to train and uses only language-independent features, outperforms the best shared task systems on the task of retrieving relevant previously asked questions.
Ana Valeria González-Garduño, Isabelle Augenstein, Anders Søgaard
EMNLP2
2018 Parameter sharing between dependency parsers for related languages
abstract
Previous work has suggested that parameter sharing between transition-based neural dependency parsers for related languages can lead to better performance, but there is no consensus on what parameters to share.We present an evaluation of 27 different parameter sharing strategies across 10 languages, representing five pairs of related languages, each pair from a different language family.We find that sharing transition classifier parameters always helps, whereas the usefulness of sharing word and/or character LSTM parameters varies.Based on this result, we propose an architecture where the transition classifier is shared, and the sharing of word and character parameters is controlled by a parameter that can be tuned on validation data.This model is linguistically motivated and obtains significant improvements over a mono-lingually trained baseline.We also find that sharing transition classifier parameters helps when training a parser on unrelated language pairs, but we find that, in the case of unrelated languages, sharing too many parameters does not help.
Miryam de Lhoneux, Johannes Bjerva, Isabelle Augenstein, Anders Søgaard
EMNLP3
2018 Multi-Task Learning of Pairwise Sequence Classification Tasks over Disparate Label Spaces
abstract
Isabelle Augenstein, Sebastian Ruder, Anders Søgaard. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Isabelle Augenstein, Sebastian Ruder, Anders Søgaard
NAACL-HLT1
2018 From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings
abstract
Johannes Bjerva, Isabelle Augenstein. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Johannes Bjerva, Isabelle Augenstein
NAACL-HLT2
2018 Discourse-aware rumour stance classification in social media using sequential classifiers
Arkaitz Zubiaga, Elena Kochkina, Maria Liakata, Rob Procter, Michal Lukasik, Kalina Bontcheva, Trevor Cohn, Isabelle Augenstein
Inf. Process. Manag.8
2017 A Supervised Approach to Extractive Summarisation of Scientific Papers
abstract
Automatic summarisation is a popular approach to reduce a document to its main arguments.Recent research in the area has focused on neural approaches to summarisation, which can be very data-hungry.However, few large datasets exist and none for the traditionally popular domain of scientific publications, which opens up challenging research avenues centered on encoding large, complex documents.In this paper, we introduce a new dataset for summarisation of computer science publications by exploiting a large resource of author provided summaries and show straightforward ways of extending it further.We develop models on the dataset making use of both neural sentence encoding and traditionally used summarisation features and show that models which encode sentences as well as their local and global context perform best, significantly outperforming well-established baseline methods.
Ed Collins, Isabelle Augenstein, Sebastian Riedel 0001
CoNLL2
2017 Generalisation in named entity recognition: A quantitative analysis
abstract
Named Entity Recognition (NER) is a key NLP task, which is all the more challenging on Web and user-generated content with their diverse and continuously changing language. This paper aims to quantify how this diversity impacts state-of-the-art NER methods, by measuring named entity (NE) and context variability, feature sparsity, and their effects on precision and recall. In particular, our findings indicate that NER approaches struggle to generalise in diverse genres with limited training data. Unseen NEs, in particular, play an important role, which have a higher incidence in diverse genres such as social media than in more regular genres such as newswire. Coupled with a higher incidence of unseen features more generally and the lack of large training corpora, this leads to significantly lower F 1 scores for diverse genres as compared to more regular ones. We also find that leading systems rely heavily on surface forms found in training data, having problems generalising beyond these, and offer explanations for this observation.
Isabelle Augenstein, Leon Derczynski, Kalina Bontcheva
Comput. Speech Lang.1
2016 Stance Detection with Bidirectional Conditional Encoding
abstract
Stance detection is the task of classifying the attitude expressed in a text towards a target such as Hillary Clinton to be "positive", negative" or "neutral". Previous work has assumed that either the target is mentioned in the text or that training data for every target is given. This paper considers the more challenging version of this task, where targets are not always mentioned and no training data is available for the test targets. We experiment with conditional LSTM encoding, which builds a representation of the tweet that is dependent on the target, and demonstrate that it outperforms encoding the tweet and the target independently. Performance is improved further when the conditional model is augmented with bidirectional encoding. We evaluate our approach on the SemEval 2016 Task 6 Twitter Stance Detection corpus achieving performance second best only to a system trained on semi-automatically labelled tweets for the test target. When such weak supervision is added, our approach achieves state-of-the-art results.
Isabelle Augenstein, Tim Rocktäschel, Andreas Vlachos 0001, Kalina Bontcheva
EMNLP1
2016 Numerically Grounded Language Models for Semantic Error Correction
abstract
Semantic error detection and correction is an important task for applications such as fact checking, speech-to-text or grammatical error correction.Current approaches generally focus on relatively shallow semantics and do not account for numeric quantities.Our approach uses language models grounded in numbers within the text.Such groundings are easily achieved for recurrent neural language model architectures, which can be further conditioned on incomplete background knowledge bases.Our evaluation on clinical reports shows that numerical grounding improves perplexity by 33% and F1 for semantic error correction by 5 points when compared to ungrounded approaches.Conditioning on a knowledge base yields further improvements.
Georgios Spithourakis, Isabelle Augenstein, Sebastian Riedel 0001
EMNLP2
2016 Monolingual Social Media Datasets for Detecting Contradiction and Entailment
Piroska Lendvai, Isabelle Augenstein, Kalina Bontcheva, Thierry Declerck
LREC2
2015 Extracting Relations between Non-Standard Entities using Distant Supervision and Imitation Learning
abstract
Distantly supervised approaches have become popular in recent years as they allow training relation extractors without textbound annotation, using instead known relations from a knowledge base and a large textual corpus from an appropriate domain.While state of the art distant supervision approaches use off-theshelf named entity recognition and classification (NERC) systems to identify relation arguments, discrepancies in domain or genre between the data used for NERC training and the intended domain for the relation extractor can lead to low performance.This is particularly problematic for "non-standard" named entities such as album which would fall into the MISC category.We propose to ameliorate this issue by jointly training the named entity classifier and the relation extractor using imitation learning which reduces structured prediction learning to classification learning.We further experiment with Web features different features and compare against using two off-the-shelf supervised NERC systems, Stanford NER and FIGER, for named entity classification.Our experiments show that imitation learning improves average precision by 4 points over an one-stage classification model, while removing Web features results in a 6 points reduction.Compared to using FIGER and Stanford NER, average precision is 10 points and 19 points higher with our imitation learning approach.
Isabelle Augenstein, Andreas Vlachos 0001, Diana Maynard
EMNLP1
2014 Relation Extraction from the Web Using Distant Supervision
Isabelle Augenstein, Diana Maynard, Fabio Ciravegna
EKAW1
2014 Joint Information Extraction from the Web Using Linked Data
Isabelle Augenstein
ISWC (2)1
2013 Unsupervised wrapper induction using linked data
abstract
This work explores the usage of Linked Data for Web scale Information Extraction and shows encouraging results on the task of Wrapper Induction. We propose a simple knowledge based method which is (i) highly flexible with respect to different domains and (ii) does not require any training material, but exploits Linked Data as background knowledge source to build essential learning resources. The major contribution of this work is a study of how Linked Data - an imprecise, redundant and large-scale knowledge resource - can be used to support Web scale Information Extraction in an effective and efficient way and identify the challenges involved. We show that, for domains that are covered, Linked Data serve as a powerful knowledge resource for Information Extraction. Experiments on a publicly available dataset demonstrate that, under certain conditions, this simple unsupervised approach can achieve competitive results against some complex state of the art that always depends on training data.
Anna Lisa Gentile, Ziqi Zhang 0001, Isabelle Augenstein, Fabio Ciravegna
K-CAP3
2013 Statistical Knowledge Patterns: Identifying Synonymous Relations in Large Linked Datasets
Ziqi Zhang 0001, Anna Lisa Gentile, Eva Blomqvist, Isabelle Augenstein, Fabio Ciravegna
ISWC (1)4
2012 LODifier: Generating Linked Data from Unstructured Text
Isabelle Augenstein, Sebastian Padó, Sebastian Rudolph
ESWC1