VLDB 2026 Research / reviewers in the wild / expert
Pepa Atanasova
dblp:224/2054
· DBLP profile ↗
16ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0002-0023-2616ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Reality Check on Context Utilisation for Retrieval-Augmented GenerationabstractRetrieval-augmented generation (RAG) helps address the limitations of parametric knowledge embedded within a language model (LM). In real world settings, retrieved information can vary in complexity, yet most investigations of LM utilisation of context has been limited to synthetic text. We introduce DRUID (Dataset of Retrieved Unreliable, Insufficient and Difficult-to-understand contexts) with real-world queries and contexts manually annotated for stance. The dataset is based on the prototypical task of automated claim verification, for which automated retrieval of real-world evidence is crucial. We compare DRUID to synthetic datasets (CounterFact, ConflictQA) and find that artificial datasets often fail to represent the complexity and diversity of realistically retrieved context. We show that synthetic datasets exaggerate context characteristics rare in real retrieved data, which leads to inflated context utilisation results, as measured by our novel ACU score. Moreover, while previous work has mainly focused on singleton context characteristics to explain context utilisation, correlations between singleton context properties and ACU on DRUID are surprisingly small compared to other properties related to context source. Overall, our work underscores the need for real-world aligned context utilisation studies to represent and improve performance in real-world RAG settings. Lovisa Hagström, Sara Marjanovic, Haeun Yu, Arnav Arora, Christina Lioma, Maria Maistro, Pepa Atanasova, Isabelle Augenstein |
ACL (1) | 7 |
| 2025 | Self-Critique and Refinement for Faithful Natural Language ExplanationsabstractWith the rapid development of Large Language Models (LLMs), Natural Language Explanations (NLEs) have become increasingly important for understanding model predictions.However, these explanations often fail to faithfully represent the model's actual reasoning process.While existing work has demonstrated that LLMs can self-critique and refine their initial outputs for various tasks, this capability remains unexplored for improving explanation faithfulness.To address this gap, we introduce Self-critique and Refinement for Natural Language Explanations (SR-NLE), a framework that enables models to improve the faithfulness of their own explanations -specifically, posthoc NLEs -through an iterative critique and refinement process without external supervision.Our framework leverages different feedback mechanisms to guide the refinement process, including natural language self-feedback and, notably, a novel feedback approach based on feature attribution that highlights important input words.Our experiments across three datasets and four state-of-the-art LLMs demonstrate that SR-NLE significantly reduces unfaithfulness rates, with our best method achieving an average unfaithfulness rate of 36.02%, compared to 54.81% for baseline -an absolute reduction of 18.79%.These findings reveal that the investigated LLMs can indeed refine their explanations to better reflect their actual reasoning process, requiring only appropriate guidance through feedback without additional training or fine-tuning.Our code is available at https://github.com/ymwangv/SR-NLE. Identify the logical relationship between premise and hypothesis.Premise: A man in a red shirt is playing guitar on stage.Hypothesis: A man is performing music.Answer: Entailment Please, provide an explanation for your answer.Explanation: The man is wearing a shirt.Please, provide the 2 most important words for the prediction.Feedback with 2 most important words: playing, performing Please, refine your explanation based on the most important words for the prediction.Refined explanation: The man is playing guitar, so he is performing music. Pepa Atanasova |
EMNLP | 2 |
| 2025 | Graph-Guided Textual Explanation Generation FrameworkabstractNatural language explanations (NLEs) are commonly used to provide plausible free-text explanations of a model's reasoning about its predictions.However, recent work has questioned their faithfulness, as they may not accurately reflect the model's internal reasoning process regarding its predicted answer.In contrast, highlight explanations-input fragments critical for the model's predicted answers-exhibit measurable faithfulness.Building on this foundation, we propose G-TEx, a Graph-Guided Textual Explanation Generation framework designed to enhance the faithfulness of NLEs.Specifically, highlight explanations are first extracted as faithful cues reflecting the model's reasoning logic toward answer prediction.They are subsequently encoded through a graph neural network layer to guide the NLE generation, which aligns the generated explanations with the model's underlying reasoning toward the predicted answer.Experiments on both encoder-decoder and decoder-only models across three reasoning datasets demonstrate that G-TEx improves NLE faithfulness by up to 12.18% compared to baseline methods.Additionally, G-TEx generates NLEs with greater semantic and lexical similarity to human-written ones.Human evaluations show that G-TEx can decrease redundant content and enhance the overall quality of NLEs.Our work presents a novel method for explicitly guiding NLE generation to enhance faithfulness, serving as a foundation for addressing broader criteria in NLE and generated text. Shuzhou Yuan, Ran Zhang 0013, Michael Färber 0001, Steffen Eger, Pepa Atanasova, Isabelle Augenstein |
EMNLP | 6 |
| 2025 | Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation FrameworkabstractJingyi Sun, Pepa Atanasova, Isabelle Augenstein. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Pepa Atanasova, Isabelle Augenstein |
NAACL (Long Papers) | 2 |
| 2024 | Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution MethodsabstractLanguage Models (LMs) acquire parametric knowledge from their training process, embedding it within their weights.The increasing scalability of LMs, however, poses significant challenges for understanding a model's inner workings and further for updating or correcting this embedded knowledge without the significant cost of retraining.This underscores the importance of unveiling exactly what knowledge is stored and its association with specific model components.Instance Attribution (IA) and Neuron Attribution (NA) offer insights into this training-acquired knowledge, though they have not been compared systematically.Our study introduces a novel evaluation framework to quantify and compare the knowledge revealed by IA and NA.To align the results of the methods we introduce the attribution method NA-Instances to apply NA for retrieving influential training instances, and IA-Neurons to discover important neurons of influential instances discovered by IA.We further propose a comprehensive list of faithfulness tests to evaluate the comprehensiveness and sufficiency of the explanations provided by both methods.Through extensive experiments and analysis, we demonstrate that NA generally reveals more diverse and comprehensive information regarding the LM's parametric knowledge compared to IA.Nevertheless, IA provides unique and valuable insights into the LM's parametric knowledge, which are not revealed by NA.Our findings further suggest the potential of a synergistic approach of combining the diverse findings of IA and NA for a more holistic understanding of an LM's parametric knowledge. Haeun Yu, Pepa Atanasova, Isabelle Augenstein |
ACL (1) | 2 |
| 2023 | bgGLUE: A Bulgarian General Language Understanding Evaluation BenchmarkabstractMomchil Hardalov, Pepa Atanasova, Todor Mihaylov, Galia Angelova, Kiril Simov, Petya Osenova, Veselin Stoyanov, Ivan Koychev, Preslav Nakov, Dragomir Radev. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Momchil Hardalov, Pepa Atanasova, Todor Mihaylov, Galia Angelova, Kiril Ivanov Simov, Petya Osenova, Veselin Stoyanov, Ivan Koychev, Preslav Nakov, Dragomir R. Radev |
ACL (1) | 2 |
| 2023 | Explaining Interactions Between Text SpansabstractReasoning over spans of tokens from different parts of the input is essential for natural language understanding (NLU) tasks such as fact-checking (FC), machine reading comprehension (MRC) or natural language inference (NLI).However, existing highlight-based explanations primarily focus on identifying individual important tokens or interactions only between adjacent tokens or tuples of tokens.Most notably, there is a lack of annotations capturing the human decision-making process w.r.t. the necessary interactions for informed decision-making in such tasks.To bridge this gap, we introduce SpanEx, a multi-annotator dataset of human span interaction explanations for two NLU tasks: NLI and FC.We then investigate the decision-making processes of multiple fine-tuned large language models in terms of the employed connections between spans in separate parts of the input and compare them to the human reasoning processes.Finally, we present a novel community detection based unsupervised method to extract such interaction explanations from a model's inner workings.1 Sagnik Ray Choudhury, Pepa Atanasova, Isabelle Augenstein |
EMNLP | 2 |
| 2022 | Diagnostics-Guided Explanation GenerationabstractExplanations shed light on a machine learning model's rationales and can aid in identifying deficiencies in its reasoning process. Explanation generation models are typically trained in a supervised way given human explanations. When such annotations are not available, explanations are often selected as those portions of the input that maximise a downstream task's performance, which corresponds to optimising an explanation's Faithfulness to a given model. Faithfulness is one of several so-called diagnostic properties, which prior work has identified as useful for gauging the quality of an explanation without requiring annotations. Other diagnostic properties are Data Consistency, which measures how similar explanations are for similar input instances, and Confidence Indication, which shows whether the explanation reflects the confidence of the model. In this work, we show how to directly optimise for these diagnostic properties when training a model to generate sentence-level explanations, which markedly improves explanation quality, agreement with human rationales, and downstream task performance on three complex reasoning tasks. Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle Augenstein |
AAAI | 1 |
| 2022 | Joint emotion label space modeling for affect lexica
Luna De Bruyne, Pepa Atanasova, Isabelle Augenstein |
Comput. Speech Lang. | 2 |
| 2022 | Fact Checking with Insufficient EvidenceabstractAbstract Automating the fact checking (FC) process relies on information obtained from external sources. In this work, we posit that it is crucial for FC models to make veracity predictions only when there is sufficient evidence and otherwise indicate when it is not enough. To this end, we are the first to study what information FC models consider sufficient by introducing a novel task and advancing it with three main contributions. First, we conduct an in-depth empirical analysis of the task with a new fluency-preserving method for omitting information from the evidence at the constituent and sentence level. We identify when models consider the remaining evidence (in)sufficient for FC, based on three trained models with different Transformer architectures and three FC datasets. Second, we ask annotators whether the omitted evidence was important for FC, resulting in a novel diagnostic dataset, SufficientFacts1, for FC with omitted evidence. We find that models are least successful in detecting missing evidence when adverbial modifiers are omitted (21% accuracy), whereas it is easiest for omitted date modifiers (63% accuracy). Finally, we propose a novel data augmentation strategy for contrastive self-learning of missing evidence by employing the proposed omission method combined with tri-training. It improves performance for Evidence Sufficiency Prediction by up to 17.8 F1 score, which in turn improves FC performance by up to 2.6 F1 score. Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle Augenstein |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | Multi-Hop Fact Checking of Political ClaimsabstractRecent work has proposed multi-hop models and datasets for studying complex natural language reasoning. One notable task requiring multi-hop reasoning is fact checking, where a set of connected evidence pieces leads to the final verdict of a claim. However, existing datasets either do not provide annotations for gold evidence pages, or the only dataset which does (FEVER) mostly consists of claims which can be fact-checked with simple reasoning and is constructed artificially. Here, we study more complex claim verification of naturally occurring claims with multiple hops over interconnected evidence chunks. We: 1) construct a small annotated dataset, PolitiHop, of evidence sentences for claim verification; 2) compare it to existing multi-hop datasets; and 3) study how to transfer knowledge from more extensive in- and out-of-domain resources to PolitiHop. We find that the task is complex and achieve the best performance with an architecture that specifically models reasoning over evidence pieces in combination with in-domain transfer learning. Wojciech Ostrowski, Arnav Arora, Pepa Atanasova, Isabelle Augenstein |
IJCAI | 3 |
| 2020 | Generating Fact Checking ExplanationsabstractMost existing work on automated fact checking is concerned with predicting the veracity of claims based on metadata, social network spread, language used in claims, and, more recently, evidence supporting or denying claims.A crucial piece of the puzzle that is still missing is to understand how to automate the most elaborate part of the process -generating justifications for verdicts on claims.This paper provides the first study of how these explanations can be generated automatically based on available claim context, and how this task can be modelled jointly with veracity prediction.Our results indicate that optimising both objectives at the same time, rather than training them separately, improves the performance of a fact checking system.The results of a manual evaluation further suggest that the informativeness, coverage and overall quality of the generated explanations are also improved in the multi-task model.Claim: The last major oil spill from a drilling accident in America happened over 40 years ago in 1969.Ruling Comments: (...) The last major oil spill from a drilling accident in America happened over 40 years ago in 1969.(...) The largest in volume was the Santa Barbara spill of 1969 referenced by Murdock and Johnson, in which an estimated 100,000 barrels of oil spilled into the Pacific Ocean, according to the API.The Santa Barbara spill was so big it ranked seventh among the 10 largest oil spills caused by marine well blowouts in the world, the report states.Two other U.S. spills, both in 1970, rank eighth and 10th.Fourteen marine blowouts have taken place in the U.S. between 1969 and 2007.Six of them took place after 1990 and spilled a total of nearly 13,700 barrels.(...) We interviewed three scientists who said that the impact of a spill has little to do with its volume.Scientists have proven that spills far smaller than Santa Barbara's have been devastating.Justification: While the nation's largest oil well blowout did take place in 1969, it's not factually correct to call it the "last major oil spill".First of all, two of the largest blowouts in the world took place in the U. S. the following year.More importantly, experts agree that spills far smaller in volume to the 1969 disaster have been devastating.From a scientific perspective, Johnson's decision to single out the 1969 blowout as the last "major" one makes no sense. Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle Augenstein |
ACL | 1 |
| 2020 | A Diagnostic Study of Explainability Techniques for Text ClassificationabstractRecent developments in machine learning have introduced models that approach human performance at the cost of increased architectural complexity.Efforts to make the rationales behind the models' predictions transparent have inspired an abundance of new explainability techniques.Provided with an already trained model, they compute saliency scores for the words of an input instance.However, there exists no definitive guide on (i) how to choose such a technique given a particular application task and model architecture, and (ii) the benefits and drawbacks of using each such technique.In this paper, we develop a comprehensive list of diagnostic properties for evaluating existing explainability techniques.We then employ the proposed list to compare a set of diverse explainability techniques on downstream text classification tasks and neural network architectures.We also compare the saliency scores assigned by the explainability techniques with human annotations of salient input regions to find relations between a model's performance and the agreement of its rationales with human ones.Overall, we find that the gradient-based explanations perform best across tasks and model architectures, and we present further insights into the properties of the reviewed explainability techniques. Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle Augenstein |
EMNLP (1) | 1 |
| 2020 | Generating Label Cohesive and Well-Formed Adversarial ClaimsabstractAdversarial attacks reveal important vulnerabilities and flaws of trained models.One potent type of attack are universal adversarial triggers, which are individual n-grams that, when appended to instances of a class under attack, can trick a model into predicting a target class.However, for inference tasks such as fact checking, these triggers often inadvertently invert the meaning of instances they are inserted in.In addition, such attacks produce semantically nonsensical inputs, as they simply concatenate triggers to existing samples.Here, we investigate how to generate adversarial attacks against fact checking systems that preserve the ground truth meaning and are semantically valid.We extend the HotFlip attack algorithm used for universal trigger generation by jointly minimizing the target class loss of a fact checking model and the entailment class loss of an auxiliary natural language inference model.We then train a conditional language model to generate semantically valid statements, which include the found universal triggers.We find that the generated attacks maintain the directionality and semantic validity of the claim better than previous work. Pepa Atanasova, Dustin Wright 0001, Isabelle Augenstein |
EMNLP (1) | 1 |
| 2019 | CheckThat! at CLEF 2019: Automatic Identification and Verification of Claims
Tamer Elsayed, Preslav Nakov, Alberto Barrón-Cedeño, Maram Hasanain, Reem Suwaileh, Giovanni Da San Martino, Pepa Atanasova |
ECIR (2) | 7 |
| 2019 | Evaluating Variable-Length Multiple-Option Lists in Chatbots and Mobile SearchabstractIn recent years, the proliferation of smart mobile devices has lead to the gradual integration of search functionality within mobile platforms. This has created an incentive to move away from the "ten blue links" metaphor, as mobile users are less likely to click on them, expecting to get the answer directly from the snippets. In turn, this has revived the interest in Question Answering. Then, along came chatbots, conversational systems, and messaging platforms, where the user needs could be better served with the system asking follow-up questions in order to better understand the user's intent. While typically a user would expect a single response at any utterance, a system could also return multiple options for the user to select from, based on different system understandings of the user's intent. However, this possibility should not be overused, as this practice could confuse and/or annoy the user. How to produce good variable-length lists, given the conflicting objectives of staying short while maximizing the likelihood of having a correct answer included in the list, is an underexplored problem. It is also unclear how to evaluate a system that tries to do that. Here we aim to bridge this gap. In particular, we define some necessary and some optional properties that an evaluation measure fit for this purpose should have. We further show that existing evaluation measures from the IR tradition are not entirely suitable for this setup, and we propose novel evaluation measures that address it satisfactorily. Pepa Atanasova, Georgi Karadzhov, Yasen Kiprov, Preslav Nakov, Fabrizio Sebastiani 0001 |
SIGIR | 1 |