Nafise Sadat Moosavi

dblp:127/3977 · DBLP profile ↗
← Back
30ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-8332-307XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 6 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rethinking the Idiomaticity Decomposability Hypothesis: Evidence from Distributional Learning
abstract
Maggie Mi, Golzar Atefi, Atsuki Yamaguchi, Felix Gers, Aline Villavicencio, Nafise Sadat Moosavi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Maggie Mi, Golzar Atefi, Atsuki Yamaguchi, Felix A. Gers, Aline Villavicencio, Nafise Sadat Moosavi
ACL (1)6
2026 No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding
abstract
Abstract Understanding culture requires reasoning across context, tradition, and implicit social knowledge, far beyond recalling isolated facts. Yet most culturally focused question answering (QA) benchmarks rely on singlehop questions, which may allow models to exploit shallow cues rather than demonstrate genuine cultural reasoning. In this work, we introduce ID-MoCQA, the first large-scale multi-hop QA dataset for assessing the cultural understanding of large language models (LLMs), grounded in Indonesian traditions and available in both English and Indonesian. We present a new framework that systematically transforms single-hop cultural questions into multi-hop reasoning chains spanning six clue types (e.g., commonsense, temporal, geographical). Our multi-stage validation pipeline, combining expert review and LLM-as-a-judge filtering, ensures high-quality question-answer pairs. Our evaluation across state-of-the-art models reveals substantial gaps in cultural reasoning, particularly in tasks requiring nuanced inference. ID-MoCQA provides a challenging and essential benchmark for advancing the cultural competency of LLMs.1
Vynska Amalia Permadi, Xingwei Tan, Nafise Sadat Moosavi, Nikolaos Aletras
Trans. Assoc. Comput. Linguistics3
2025 Rolling the DICE on Idiomaticity: How LLMs Fail to Grasp Context
abstract
Human processing of idioms heavily depends on interpreting the surrounding context in which they appear.While large language models (LLMs) have achieved impressive performance on idiomaticity detection benchmarks, this success may be driven by reasoning shortcuts present in existing datasets.To address this, we introduce a novel, controlled contrastive dataset (DICE) specifically designed to assess whether LLMs can effectively leverage context to disambiguate idiomatic meanings.Furthermore, we investigate the influence of collocational frequency and sentence probability-proxies for human processing known to affect idiom resolution-on model performance.Our results show that LLMs frequently fail to resolve idiomaticity when it depends on contextual understanding, and they perform better on sentences deemed more likely by the model.Additionally, idiom frequency influences performance but does not guarantee accurate interpretation.Our findings emphasize the limitations of current models in grasping contextual meaning and highlight the need for more context-sensitive evaluation.
Maggie Mi, Aline Villavicencio, Nafise Sadat Moosavi
ACL (1)3
2025 How to Leverage Digit Embeddings to Represent Numbers?
abstract
Within numerical reasoning, understanding numbers themselves is still a challenge for existing language models. Simple generalisations, such as solving 100+200 instead of 1+2, can substantially affect model performance (Sivakumar and Moosavi, 2023). Among various techniques, character-level embeddings of numbers have emerged as a promising approach to improve number representation. However, this method has limitations as it leaves the task of aggregating digit representations to the model, which lacks direct supervision for this process. In this paper, we explore the use of mathematical priors to compute aggregated digit embeddings and explicitly incorporate these aggregates into transformer models. This can be achieved either by adding a special token to the input embeddings or by introducing an additional loss function to enhance correct predictions. We evaluate the effectiveness of incorporating this explicit aggregation, analysing its strengths and shortcomings, and discuss future directions to better benefit from this approach. Our methods, while simple, are compatible with any pretrained model, easy to implement, and have been made publicly available.
Jasivan Alex Sivakumar, Nafise Sadat Moosavi
COLING2
2025 From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors
abstract
Language models often struggle with idiomatic, figurative, or context-sensitive inputs, not because they produce flawed outputs, but because they misinterpret the input from the outset.We propose an input-only method for anticipating such failures using token-level likelihood features inspired by surprisal and the Uniform Information Density hypothesis.These features capture localized uncertainty in input comprehension and outperform standard baselines across five linguistically challenging datasets.We show that span-localized features improve error detection for larger models, while smaller models benefit from global patterns.Our method requires no access to outputs or hidden activations, offering a lightweight and generalizable approach to pre-generation error prediction. https://github.com/mi-m1/input_ perceptionHow is the expression "caught between a rock and a hard place" used in the following sentence?Literally or figuratively? LLM AnswerOutput-level Error Detection External Knowledge BaseLog Probabilities LLM as a Judge Suddenly she was caught between a rock and a hard place.'Suddenly', 'Ġshe', 'Ġwas', 'Ġcaught', 'Ġbetween', 'Ġa', 'Ġrock', 'Ġand', 'Ġa', 'Ġhard', 'Ġplace' Error: 81% Correct:19% Prompt: How is the expression "caught between a rock and a hard place" used in the following sentence?Literally or figuratively?Sentence: Suddenly she was caught between a rock and a hard place.
Maggie Mi, Aline Villavicencio, Nafise Sadat Moosavi
EMNLP3
2025 Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language
abstract
Dehumanization, i.e., denying human qualities to individuals or groups, is a particularly harmful form of hate speech that can normalize violence against marginalized communities.Despite advances in NLP for detecting general hate speech, approaches to identifying dehumanizing language remain limited due to scarce annotated data and the subtle nature of such expressions.In this work, we systematically evaluate four state-of-the-art large language models (LLMs) -Claude, GPT, Mistral, and Qwenfor dehumanization detection.Our results show that only one model-Claude-achieves strong performance (over 80% F 1 ) under an optimized configuration, while others, despite their capabilities, perform only moderately.Performance drops further when distinguishing dehumanization from related hate types such as derogation.We also identify systematic disparities across target groups: models tend to over-predict dehumanization for some identities (e.g., Gay men), while under-identifying it for others (e.g., Refugees).These findings motivate the need for systematic, group-level evaluation when applying pretrained language models to dehumanization detection tasks.
Hamidreza Saffari, Mohammadamin Shafiei, Hezhao Zhang, Lasana T. Harris, Nafise Sadat Moosavi
EMNLP5
2024 Spanning the Spectrum of Hatred Detection: A Persian Multi-Label Hate Speech Dataset with Annotator Rationales
abstract
With the alarming rise of hate speech in online communities, the demand for effective NLP models to identify instances of offensive language has reached a critical point. However, the development of such models heavily relies on the availability of annotated datasets, which are scarce, particularly for less-studied languages. To bridge this gap for the Persian language, we present a novel dataset specifically tailored to multi-label hate speech detection. Our dataset, called Phate, consists of an extensive collection of over seven thousand manually-annotated Persian tweets, offering a rich resource for training and evaluating hate speech detection models on this language. Notably, each annotation in our dataset specifies the targeted group of hate speech and includes a span of the tweet which elucidates the rationale behind the assigned label. The incorporation of these information expands the potential applications of our dataset, facilitating the detection of targeted online harm or allowing the benchmark to serve research on interpretability of hate speech detection models. The dataset, annotation guideline, and all associated codes are accessible at https://github.com/Zahra-D/Phate.
Zahra Delbari, Nafise Sadat Moosavi, Mohammad Taher Pilehvar
AAAI2
2024 Universal Anaphora: The First Three Years
abstract
The aim of the Universal Anaphora initiative is to push forward the state of the art in anaphora and anaphora resolution by expanding the aspects of anaphoric interpretation which are or can be reliably annotated in anaphoric corpora, producing unified standards to annotate and encode these annotations, delivering datasets encoded according to these standards, and developing methods for evaluating models that carry out this type of interpretation. Although several papers on aspects of the initiative have appeared, no overall description of the initiative’s goals, proposals and achievements has been published yet except as an online draft. This paper aims to fill this gap, as well as to discuss its progress so far.
Massimo Poesio, Maciej Ogrodniczuk, Vincent Ng 0001, Sameer Pradhan, Juntao Yu, Nafise Sadat Moosavi, Silviu Paun, Amir Zeldes, Anna Nedoluzhko, Michal Novák 0001, Martin Popel, Zdenek Zabokrtský, Daniel Zeman
LREC/COLING6
2024 Blending Social Interaction Realms: Harmonizing Online and Offline Interactions through Augmented Reality
Guanxuan Jiang, Yuyang Wang 0002, Yue Li 0023, Nafise Sadat Moosavi, Pan Hui 0001
VINCI4
2023 FERMAT: An Alternative to Accuracy for Numerical Reasoning
abstract
While pre-trained language models achieve impressive performance on various NLP benchmarks, they still struggle with tasks that require numerical reasoning.Recent advances in improving numerical reasoning are mostly achieved using very large language models that contain billions of parameters and are not accessible to everyone.In addition, numerical reasoning is measured using a single score on existing datasets.As a result, we do not have a clear understanding of the strengths and shortcomings of existing models on different numerical reasoning aspects and therefore, potential ways to improve them apart from scaling them up.Inspired by CheckList (Ribeiro et al., 2020), we introduce a multi-view evaluation set for numerical reasoning in English, called FERMAT.Instead of reporting a single score on a whole dataset, FERMAT evaluates models on various key numerical reasoning aspects such as number understanding, mathematical operations, and training dependency.Apart from providing a comprehensive evaluation of models on different numerical reasoning aspects, FERMAT enables a systematic and automated generation of an arbitrarily large training or evaluation set for each aspect.The datasets and codes are publicly available to generate further multi-view data for ulterior tasks and languages.
Jasivan Alex Sivakumar, Nafise Sadat Moosavi
ACL (1)2
2023 Lessons Learned from a Citizen Science Project for Natural Language Processing
abstract
Jan-Christoph Klie, Ji-Ung Lee, Kevin Stowe, Gözde Şahin, Nafise Sadat Moosavi, Luke Bates, Dominic Petrak, Richard Eckart De Castilho, Iryna Gurevych. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Jan-Christoph Klie, Ji-Ung Lee, Kevin Stowe, Gözde Gül Sahin, Nafise Sadat Moosavi, Luke Bates, Dominic Petrak, Richard Eckart de Castilho, Iryna Gurevych
EACL5
2023 Learning From Free-Text Human Feedback - Collect New Datasets Or Extend Existing Ones?
abstract
Learning from free-text human feedback is essential for dialog systems, but annotated data is scarce and usually covers only a small fraction of error types known in conversational AI.Instead of collecting and annotating new datasets from scratch, recent advances in synthetic dialog generation could be used to augment existing dialog datasets with the necessary annotations.However, to assess the feasibility of such an effort, it is important to know the types and frequency of free-text human feedback included in these datasets.In this work, we investigate this question for a variety of commonly used dialog datasets, including MultiWoZ, SGD, BABI, PersonaChat, Wizardsof-Wikipedia, and the human-bot split of the Self-Feeding Chatbot.Using our observations, we derive new taxonomies for the annotation of free-text human feedback in dialogs and investigate the impact of including such data in response generation for three SOTA language generation models, including GPT-2, LLAMA, and Flan-T5.Our findings provide new insights into the composition of the datasets examined, including error types, user response types, and the relations between them 1 .
Dominic Petrak, Nafise Sadat Moosavi, Nikolai Rozanov, Iryna Gurevych
EMNLP2
2022 Layer or Representation Space: What Makes BERT-based Evaluation Metrics Robust?
abstract
The evaluation of recent embedding-based evaluation metrics for text generation is primarily based on measuring their correlation with human evaluations on standard benchmarks. However, these benchmarks are mostly from similar domains to those used for pretraining word embeddings. This raises concerns about the (lack of) generalization of embedding-based metrics to new and noisy domains that contain a different vocabulary than the pretraining data. In this paper, we examine the robustness of BERTScore, one of the most popular embedding-based metrics for text generation. We show that (a) an embedding-based metric that has the highest correlation with human evaluations on a standard benchmark can have the lowest correlation if the amount of input noise or unknown tokens increases, (b) taking embeddings from the first layer of pretrained models improves the robustness of all metrics, and (c) the highest robustness is achieved when using character-level embeddings, instead of token-based embeddings, from the first layer of the pretrained model.
Doan Nam Long Vu, Nafise Sadat Moosavi, Steffen Eger
COLING2
2022 The Universal Anaphora Scorer
abstract
The aim of the Universal Anaphora initiative is to push forward the state of the art in anaphora and anaphora resolution by expanding the aspects of anaphoric interpretation which are or can be reliably annotated in anaphoric corpora, producing unified standards to annotate and encode these annotations, deliver datasets encoded according to these standards, and developing methods for evaluating models carrying out this type of interpretation. Such expansion of the scope of anaphora resolution requires a comparable expansion of the scope of the scorers used to evaluate this work. In this paper, we introduce an extended version of the Reference Coreference Scorer (Pradhan et al., 2014) that can be used to evaluate the extended range of anaphoric interpretation included in the current Universal Anaphora proposal. The UA scorer supports the evaluation of identity anaphora resolution and of bridging reference resolution, for which scorers already existed but not integrated in a single package. It also supports the evaluation of split antecedent anaphora and discourse deixis, for which no tools existed. The proposed approach to the evaluation of split antecedent anaphora is entirely novel; the proposed approach to the evaluation of discourse deixis leverages the encoding of discourse deixis proposed in Universal Anaphora to enable the use for discourse deixis of the same metrics already used for identity anaphora. The scorer was tested in the recent CODI-CRAC 2021 Shared Task on Anaphora Resolution in Dialogues.
Juntao Yu, Sopan Khosla, Nafise Sadat Moosavi, Silviu Paun, Sameer Pradhan, Massimo Poesio
LREC3
2022 Adaptable Adapters
abstract
Nafise Moosavi, Quentin Delfosse, Kristian Kersting, Iryna Gurevych. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Nafise Sadat Moosavi, Quentin Delfosse, Kristian Kersting, Iryna Gurevych
NAACL-HLT1
2022 Falsesum: Generating Document-level NLI Examples for Recognizing Factual Inconsistency in Summarization
abstract
Prasetya Utama, Joshua Bambrick, Nafise Moosavi, Iryna Gurevych. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Prasetya Ajie Utama, Joshua Bambrick, Nafise Sadat Moosavi, Iryna Gurevych
NAACL-HLT3
2021 Coreference Reasoning in Machine Reading Comprehension
abstract
Mingzhu Wu, Nafise Sadat Moosavi, Dan Roth, Iryna Gurevych. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Mingzhu Wu, Nafise Sadat Moosavi, Dan Roth 0001, Iryna Gurevych
ACL/IJCNLP (1)2
2021 Avoiding Inference Heuristics in Few-shot Prompt-based Finetuning
abstract
This is a repository copy of Avoiding inference heuristics in few-shot prompt-based finetuning.
Prasetya Ajie Utama, Nafise Sadat Moosavi, Victor Sanh, Iryna Gurevych
EMNLP (1)2
2021 Stay Together: A System for Single and Split-antecedent Anaphora Resolution
abstract
Juntao Yu, Nafise Sadat Moosavi, Silviu Paun, Massimo Poesio. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Juntao Yu, Nafise Sadat Moosavi, Silviu Paun, Massimo Poesio
NAACL-HLT2
2020 Mind the Trade-off: Debiasing NLU Models without Degrading the In-distribution Performance
abstract
Models for natural language understanding (NLU) tasks often rely on the idiosyncratic biases of the dataset, which make them brittle against test cases outside the training distribution.Recently, several proposed debiasing methods are shown to be very effective in improving out-of-distribution performance.However, their improvements come at the expense of performance drop when models are evaluated on the in-distribution data, which contain examples with higher diversity.This seemingly inevitable trade-off may not tell us much about the changes in the reasoning and understanding capabilities of the resulting models on broader types of examples beyond the small subset represented in the outof-distribution data.In this paper, we address this trade-off by introducing a novel debiasing method, called confidence regularization, which discourage models from exploiting biases while enabling them to receive enough incentive to learn from all the training examples.We evaluate our method on three NLU tasks and show that, in contrast to its predecessors, it improves the performance on out-of-distribution datasets (e.g., 7pp gain on HANS dataset) while maintaining the original in-distribution accuracy.1
Prasetya Ajie Utama, Nafise Sadat Moosavi, Iryna Gurevych
ACL2
2020 Free the Plural: Unrestricted Split-Antecedent Anaphora Resolution
abstract
Now that the performance of coreference resolvers on the simpler forms of anaphoric reference has greatly improved, more attention is devoted to more complex aspects of anaphora.One limitation of virtually all coreference resolution models is the focus on single-antecedent anaphors.Plural anaphors with multiple antecedents-so-called split-antecedent anaphors (as in John met Mary.They went to the movies)-have not been widely studied, because they are not annotated in ONTONOTES and are relatively infrequent in other corpora.In this paper, we introduce the first model for unrestricted resolution of split-antecedent anaphors.We start with a strong baseline enhanced by BERT embeddings, and show that we can substantially improve its performance by addressing the sparsity issue.To do this, we experiment with auxiliary corpora where split-antecedent anaphors were annotated by the crowd, and with transfer learning models using element-of bridging references and single-antecedent coreference as auxiliary tasks.Evaluation on the gold annotated ARRAU corpus shows that the out best model uses a combination of three auxiliary corpora achieved F1 scores of 70% and 43.6% when evaluated in a lenient and strict setting, respectively, i.e., 11 and 21 percentage points gain when compared with our baseline.
Juntao Yu, Nafise Sadat Moosavi, Silviu Paun, Massimo Poesio
COLING2
2020 Towards Debiasing NLU Models from Unknown Biases
abstract
NLU models often exploit biases to achieve high dataset-specific performance without properly learning the intended task.Recently proposed debiasing methods are shown to be effective in mitigating this tendency.However, these methods rely on a major assumption that the types of bias should be known a-priori, which limits their application to many NLU tasks and datasets.In this work, we present the first step to bridge this gap by introducing a self-debiasing framework that prevents models from mainly utilizing biases without knowing them in advance.The proposed framework is general and complementary to the existing debiasing methods.We show that it allows these existing methods to retain the improvement on the challenge datasets (i.e., sets of examples designed to expose models' reliance on biases) without specifically targeting certain biases.Furthermore, the evaluation suggests that applying the framework results in improved overall robustness. 1
Prasetya Ajie Utama, Nafise Sadat Moosavi, Iryna Gurevych
EMNLP (1)2
2019 COALA: A Neural Coverage-Based Approach for Long Answer Selection with Small Data
abstract
Current neural network based community question answering (cQA) systems fall short of (1) properly handling long answers which are common in cQA; (2) performing under small data conditions, where a large amount of training data is unavailable—i.e., for some domains in English and even more so for a huge number of datasets in other languages; and (3) benefiting from syntactic information in the model—e.g., to differentiate between identical lexemes with different syntactic roles. In this paper, we propose COALA, an answer selection approach that (a) selects appropriate long answers due to an effective comparison of all question-answer aspects, (b) has the ability to generalize from a small number of training examples, and (c) makes use of the information about syntactic roles of words. We show that our approach outperforms existing answer selection models by a large margin on six cQA datasets from different domains. Furthermore, we report the best results on the passage retrieval benchmark WikiPassageQA.
Andreas Rücklé, Nafise Sadat Moosavi, Iryna Gurevych
AAAI2
2019 Using Automatically Extracted Minimum Spans to Disentangle Coreference Evaluation from Boundary Detection
abstract
This is a repository copy of Using automatically extracted minimum spans to disentangle coreference evaluation from boundary detection.
Nafise Sadat Moosavi, Leo Born, Massimo Poesio, Michael Strube 0001
ACL (1)1
2019 Neural Duplicate Question Detection without Labeled Training Data
abstract
Andreas Rücklé, Nafise Sadat Moosavi, Iryna Gurevych. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Andreas Rücklé, Nafise Sadat Moosavi, Iryna Gurevych
EMNLP/IJCNLP (1)2
2018 Using Linguistic Features to Improve the Generalization Capability of Neural Coreference Resolvers
abstract
This is a repository copy of Using linguistic features to improve the generalization capability of neural coreference resolvers.
Nafise Sadat Moosavi, Michael Strube 0001
EMNLP1
2017 Revisiting Selectional Preferences for Coreference Resolution
abstract
Selectional preferences have long been claimed to be essential for coreference resolution.However, they are mainly modeled only implicitly by current coreference resolvers.We propose a dependencybased embedding model of selectional preferences which allows fine-grained compatibility judgments with high coverage.We show that the incorporation of our model improves coreference resolution performance on the CoNLL dataset, matching the state-of-the-art results of a more complex system.However, it comes with a cost that makes it debatable how worthwhile such improvements are.
Benjamin Heinzerling, Nafise Sadat Moosavi, Michael Strube 0001
EMNLP2
2016 Which Coreference Evaluation Metric Do You Trust? A Proposal for a Link-based Entity Aware Metric
abstract
This is a repository copy of Which coreference evaluation metric do you trust?A proposal for a link-based entity aware metric.
Nafise Sadat Moosavi, Michael Strube 0001
ACL (1)1
2016 Search Space Pruning: A Simple Solution for Better Coreference Resolvers
abstract
This is a repository copy of Search space pruning: a simple solution for better coreference resolvers.
Nafise Sadat Moosavi, Michael Strube 0001
HLT-NAACL1
2014 Unsupervised Coreference Resolution by Utilizing the Most Informative Relations
Nafise Sadat Moosavi, Michael Strube 0001
COLING1