Fahime Same

dblp:280/0239 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-3413-3739ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021
YearPublicationVenuePosition
2025 Do My Eyes Deceive Me? A Survey of Human Evaluations of Hallucinations in NLG
abstract
Hallucinations are one of the most pressing challenges for large language models (LLMs). While numerous methods have been proposed to detect and mitigate them automatically, human evaluation continues to serve as the gold standard. However, these human evaluations of hallucinations show substantial variation in definitions, terminology, and evaluation practices. In this paper, we survey 64 studies involving human evaluation of hallucination published between 2019 and 2024, to investigate how hallucinations are currently defined and assessed. Our analysis reveals a lack of consistency in definitions and exposes several concerning methodological shortcomings. Crucial details, such as evaluation guidelines, user interface design, inter-annotator agreement metrics, and annotator demographics, are frequently under-reported or omitted altogether.
Patrícia Schmidtová, Eduardo Calò, Simone Balloccu, Dimitra Gkatzia, Rudali Huidrom, Mateusz Lango, Fahime Same, Vilém Zouhar, Saad Mahamood, Ondrej Dusek
INLG7
2025 Analysing Reference Production of Large Language Models
abstract
This study investigates how large language models (LLMs) produce referring expressions (REs) and to what extent their behaviour aligns with human patterns. We evaluate LLM performance in two settings: slot filling, %KvD the conventional task of referring expression generation, where REs are generated within a fixed context, and language generation, where REs are analysed within fully generated texts. Using the WebNLG corpus, we assess how well LLMs capture human variation in reference production and analyse their behaviour by examining the influence of several factors known to affect human reference production, including referential form, syntactic position, recency, and discourse status. Our findings show that (1) task framing significantly affects LLMs’ reference production; (2) while LLMs are sensitive to some of these factors, their referential behaviour consistently diverges from human use; and (3) larger model size does not necessarily yield more human-like variation. These results underscore key limitations in current LLMs’ ability to replicate human referential choices.
Chengzhao Wu, Guanyi Chen, Fahime Same, Tingting He 0003
INLG3
2024 Intrinsic Task-based Evaluation for Referring Expression Generation
abstract
Recently, a human evaluation study of Referring Expression Generation (REG) models had an unexpected conclusion: on WEBNLG, Referring Expressions (REs) generated by the state-of-the-art neural models were not only indistinguishable from the REs in WEBNLG but also from the REs generated by a simple rulebased system.Here, we argue that this limitation could stem from the use of a purely ratings-based human evaluation (which is a common practice in Natural Language Generation).To investigate these issues, we propose an intrinsic task-based evaluation for REG models, in which, in addition to rating the quality of REs, participants were asked to accomplish two meta-level tasks.One of these tasks concerns the referential success of each RE; the other task asks participants to suggest a better alternative for each RE.The outcomes suggest that, in comparison to previous evaluations, the new evaluation protocol assesses the performance of each REG model more comprehensively and makes the participants' ratings more reliable and discriminable.
Guanyi Chen, Fahime Same, Kees van Deemter
ACL (1)2
2024 Experimental versus In-Corpus Variation in Referring Expression Choice
abstract
In this paper, we compare the results of three studies. The first explored feature-conditioned distributions of referring expression (RE) forms in the original corpus from which the contexts were taken. The second is a crowdsourcing study in which we asked participants to express entities within a pre-existing context, given fully specified referents. The third study replicates the crowdsourcing experiment using Large Language Models (LLMs). We evaluate how well the corpus itself can model the variation found when multiple informants (either human participants or LLMs) choose REs in the same contexts. We measure the similarity of the conditional distributions of form categories using the Jensen-Shannon Divergence metric and Description Length metric. We find that the experimental methodology introduces substantial noise, but by taking this noise into account, we can model the variation captured from the corpus and RE form choices made during experiments. Furthermore, we compared the three conditional distributions over the corpus, the human experimental results, and the GPT models. Against our expectations, the divergence is greatest between the corpus and the GPT model.
T. Mark Ellison, Fahime Same
LREC/COLING2
2024 Generating Hotel Highlights from Unstructured Text using LLMs
abstract
We describe our implementation and evaluation of the Hotel Highlights system which has been deployed live by trivago.This system leverages a large language model (LLM) to generate a set of highlights from accommodation descriptions and reviews, enabling travellers to quickly understand its unique aspects.In this paper, we discuss our motivation for building this system and the human evaluation we conducted, comparing the generated highlights against the source input to assess the degree of hallucinations and/or contradictions present.Finally, we outline the lessons learned and the improvements needed.
Srinivas Ramesh Kamath, Fahime Same, Saad Mahamood
INLG2
2023 Models of reference production: How do they withstand the test of time?
abstract
In recent years, many NLP studies have focused solely on performance improvement.In this work, we focus on the linguistic and scientific aspects of NLP.We use the task of generating referring expressions in context (REG-incontext) as a case study and start our analysis from GREC, a comprehensive set of shared tasks in English that addressed this topic over a decade ago.We ask what the performance of models would be if we assessed them (1) on more realistic datasets, and (2) using more advanced methods.We test the models using different evaluation metrics and feature selection experiments.We conclude that GREC can no longer be regarded as offering a reliable assessment of models' ability to mimic human reference production, because the results are highly impacted by the choice of corpus and evaluation metrics.Our results also suggest that pre-trained language models are less dependent on the choice of corpus than classic Machine Learning models, and therefore make more robust class predictions.
Fahime Same, Guanyi Chen, Kees van Deemter
INLG1
2023 Neural referential form selection: Generalisability and interpretability
abstract
In recent years, a range of Neural Referring Expression Generation (REG) systems have been built and they have often achieved encouraging results. However, these models are often thought to lack transparency and generality. Firstly, it is hard to understand what these neural REG models can learn and to compare their performance with existing linguistic theories. Secondly, it is unclear whether they can generalise to data in different text genres and different languages. To answer these questions, we propose to focus on a sub-task of REG: Referential Form Selection (RFS). We introduce the task of RFS and a series of neural RFS models built on state-of-the-art neural REG models. To address the issue of interpretability, we probe these RFS models using probing classifiers that consider information known to impact the human choice of Referential Forms. To address the issue of generalisability, we assess the performance of RFS models on multiple datasets in multiple genres and two different languages, namely, English and Chinese.
Guanyi Chen, Fahime Same, Kees van Deemter
Comput. Speech Lang.2
2022 Non-neural Models Matter: a Re-evaluation of Neural Referring Expression Generation Systems
abstract
In recent years, neural models have often outperformed rule-based and classic Machine Learning approaches in NLG.These classic approaches are now often disregarded, for example when new neural models are evaluated.We argue that they should not be overlooked, since for some tasks, well-designed non-neural approaches achieve better performance than neural ones.In this paper, the task of generating referring expressions in linguistic context is used as an example.We examined two very different English datasets (WEBNLG and WSJ), and evaluated each algorithm using both automatic and human evaluations.Overall, the results of these evaluations suggest that rule-based systems with simple rule sets achieve on-par or better performance on both datasets compared to state-of-the-art neural REG systems.In the case of the more realistic dataset, WSJ, a machine learning-based system with well-designed linguistic features performed best.We hope that our work can encourage researchers to consider non-neural models in future.* Equal contribution.Order determined by swapping the order in Chen et al. (2021).
Fahime Same, Guanyi Chen, Kees van Deemter
ACL (1)1
2022 Constructing Distributions of Variation in Referring Expression Type from Corpora for Model Evaluation
abstract
The generation of referring expressions (REs) is a non-deterministic task. However, the algorithms for the generation of REs are standardly evaluated against corpora of written texts which include only one RE per each reference. Our goal in this work is firstly to reproduce one of the few studies taking the distributional nature of the RE generation into account. We add to this work, by introducing a method for exploring variation in human RE choice on the basis of longitudinal corpora - substantial corpora with a single human judgement (in the process of composition) per RE. We focus on the prediction of RE types, proper name, description and pronoun. We compare evaluations made against distributions over these types with evaluations made against parallel human judgements. Our results show agreement in the evaluation of learning algorithms against distributions constructed from parallel human evaluations and from longitudinal data.
T. Mark Ellison, Fahime Same
LREC2
2021 What can Neural Referential Form Selectors Learn?
abstract
Despite achieving encouraging results, neural Referring Expression Generation models are often thought to lack transparency.We probed neural Referential Form Selection (RFS) models to find out to what extent the linguistic features influencing the RE form are learnt and captured by state-of-the-art RFS models.The results of 8 probing tasks show that all the defined features were learnt to some extent.The probing tasks pertaining to referential status and syntactic position exhibited the highest performance.The lowest performance was achieved by the probing models designed to predict discourse structure properties beyond the sentence level.
Guanyi Chen, Fahime Same, Kees van Deemter
INLG2
2020 A Linguistic Perspective on Reference: Choosing a Feature Set for Generating Referring Expressions in Context
abstract
This paper reports on a structured evaluation of feature-based Machine Learning algorithms for selecting the form of a referring expression in discourse context.Based on this evaluation, we selected seven feature sets from the literature, amounting to 65 distinct linguistic features.The features were then grouped into 9 broad classes.After building Random Forest models, we used Feature Importance Ranking and Sequential Forward Search methods to assess the "importance" of the features.Combining the results of the two methods, we propose a consensus feature set.The 6 features in our consensus set come from 4 different classes, namely grammatical role, inherent features of the referent, antecedent form and recency.
Fahime Same, Kees van Deemter
COLING1