VLDB 2026 Research / reviewers in the wild / expert
Fahime Same
dblp:280/0239
· DBLP profile ↗
11ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-3413-3739ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Do My Eyes Deceive Me? A Survey of Human Evaluations of Hallucinations in NLGabstractHallucinations are one of the most pressing challenges for large language models (LLMs). While numerous methods have been proposed to detect and mitigate them automatically, human evaluation continues to serve as the gold standard. However, these human evaluations of hallucinations show substantial variation in definitions, terminology, and evaluation practices. In this paper, we survey 64 studies involving human evaluation of hallucination published between 2019 and 2024, to investigate how hallucinations are currently defined and assessed. Our analysis reveals a lack of consistency in definitions and exposes several concerning methodological shortcomings. Crucial details, such as evaluation guidelines, user interface design, inter-annotator agreement metrics, and annotator demographics, are frequently under-reported or omitted altogether. Patrícia Schmidtová, Eduardo Calò, Simone Balloccu, Dimitra Gkatzia, Rudali Huidrom, Mateusz Lango, Fahime Same, Vilém Zouhar, Saad Mahamood, Ondrej Dusek |
INLG | 7 |
| 2025 | Analysing Reference Production of Large Language ModelsabstractThis study investigates how large language models (LLMs) produce referring expressions (REs) and to what extent their behaviour aligns with human patterns. We evaluate LLM performance in two settings: slot filling, %KvD the conventional task of referring expression generation, where REs are generated within a fixed context, and language generation, where REs are analysed within fully generated texts. Using the WebNLG corpus, we assess how well LLMs capture human variation in reference production and analyse their behaviour by examining the influence of several factors known to affect human reference production, including referential form, syntactic position, recency, and discourse status. Our findings show that (1) task framing significantly affects LLMs’ reference production; (2) while LLMs are sensitive to some of these factors, their referential behaviour consistently diverges from human use; and (3) larger model size does not necessarily yield more human-like variation. These results underscore key limitations in current LLMs’ ability to replicate human referential choices. Chengzhao Wu, Guanyi Chen, Fahime Same, Tingting He 0003 |
INLG | 3 |
| 2024 | Intrinsic Task-based Evaluation for Referring Expression GenerationabstractRecently, a human evaluation study of Referring Expression Generation (REG) models had an unexpected conclusion: on WEBNLG, Referring Expressions (REs) generated by the state-of-the-art neural models were not only indistinguishable from the REs in WEBNLG but also from the REs generated by a simple rulebased system.Here, we argue that this limitation could stem from the use of a purely ratings-based human evaluation (which is a common practice in Natural Language Generation).To investigate these issues, we propose an intrinsic task-based evaluation for REG models, in which, in addition to rating the quality of REs, participants were asked to accomplish two meta-level tasks.One of these tasks concerns the referential success of each RE; the other task asks participants to suggest a better alternative for each RE.The outcomes suggest that, in comparison to previous evaluations, the new evaluation protocol assesses the performance of each REG model more comprehensively and makes the participants' ratings more reliable and discriminable. Guanyi Chen, Fahime Same, Kees van Deemter |
ACL (1) | 2 |
| 2024 | Experimental versus In-Corpus Variation in Referring Expression ChoiceabstractIn this paper, we compare the results of three studies. The first explored feature-conditioned distributions of referring expression (RE) forms in the original corpus from which the contexts were taken. The second is a crowdsourcing study in which we asked participants to express entities within a pre-existing context, given fully specified referents. The third study replicates the crowdsourcing experiment using Large Language Models (LLMs). We evaluate how well the corpus itself can model the variation found when multiple informants (either human participants or LLMs) choose REs in the same contexts. We measure the similarity of the conditional distributions of form categories using the Jensen-Shannon Divergence metric and Description Length metric. We find that the experimental methodology introduces substantial noise, but by taking this noise into account, we can model the variation captured from the corpus and RE form choices made during experiments. Furthermore, we compared the three conditional distributions over the corpus, the human experimental results, and the GPT models. Against our expectations, the divergence is greatest between the corpus and the GPT model. T. Mark Ellison, Fahime Same |
LREC/COLING | 2 |
| 2024 | Generating Hotel Highlights from Unstructured Text using LLMsabstractWe describe our implementation and evaluation of the Hotel Highlights system which has been deployed live by trivago.This system leverages a large language model (LLM) to generate a set of highlights from accommodation descriptions and reviews, enabling travellers to quickly understand its unique aspects.In this paper, we discuss our motivation for building this system and the human evaluation we conducted, comparing the generated highlights against the source input to assess the degree of hallucinations and/or contradictions present.Finally, we outline the lessons learned and the improvements needed. Srinivas Ramesh Kamath, Fahime Same, Saad Mahamood |
INLG | 2 |
| 2023 | Models of reference production: How do they withstand the test of time?abstractIn recent years, many NLP studies have focused solely on performance improvement.In this work, we focus on the linguistic and scientific aspects of NLP.We use the task of generating referring expressions in context (REG-incontext) as a case study and start our analysis from GREC, a comprehensive set of shared tasks in English that addressed this topic over a decade ago.We ask what the performance of models would be if we assessed them (1) on more realistic datasets, and (2) using more advanced methods.We test the models using different evaluation metrics and feature selection experiments.We conclude that GREC can no longer be regarded as offering a reliable assessment of models' ability to mimic human reference production, because the results are highly impacted by the choice of corpus and evaluation metrics.Our results also suggest that pre-trained language models are less dependent on the choice of corpus than classic Machine Learning models, and therefore make more robust class predictions. Fahime Same, Guanyi Chen, Kees van Deemter |
INLG | 1 |
| 2023 | Neural referential form selection: Generalisability and interpretabilityabstractIn recent years, a range of Neural Referring Expression Generation (REG) systems have been built and they have often achieved encouraging results. However, these models are often thought to lack transparency and generality. Firstly, it is hard to understand what these neural REG models can learn and to compare their performance with existing linguistic theories. Secondly, it is unclear whether they can generalise to data in different text genres and different languages. To answer these questions, we propose to focus on a sub-task of REG: Referential Form Selection (RFS). We introduce the task of RFS and a series of neural RFS models built on state-of-the-art neural REG models. To address the issue of interpretability, we probe these RFS models using probing classifiers that consider information known to impact the human choice of Referential Forms. To address the issue of generalisability, we assess the performance of RFS models on multiple datasets in multiple genres and two different languages, namely, English and Chinese. Guanyi Chen, Fahime Same, Kees van Deemter |
Comput. Speech Lang. | 2 |
| 2022 | Non-neural Models Matter: a Re-evaluation of Neural Referring Expression Generation SystemsabstractIn recent years, neural models have often outperformed rule-based and classic Machine Learning approaches in NLG.These classic approaches are now often disregarded, for example when new neural models are evaluated.We argue that they should not be overlooked, since for some tasks, well-designed non-neural approaches achieve better performance than neural ones.In this paper, the task of generating referring expressions in linguistic context is used as an example.We examined two very different English datasets (WEBNLG and WSJ), and evaluated each algorithm using both automatic and human evaluations.Overall, the results of these evaluations suggest that rule-based systems with simple rule sets achieve on-par or better performance on both datasets compared to state-of-the-art neural REG systems.In the case of the more realistic dataset, WSJ, a machine learning-based system with well-designed linguistic features performed best.We hope that our work can encourage researchers to consider non-neural models in future.* Equal contribution.Order determined by swapping the order in Chen et al. (2021). Fahime Same, Guanyi Chen, Kees van Deemter |
ACL (1) | 1 |
| 2022 | Constructing Distributions of Variation in Referring Expression Type from Corpora for Model EvaluationabstractThe generation of referring expressions (REs) is a non-deterministic task. However, the algorithms for the generation of REs are standardly evaluated against corpora of written texts which include only one RE per each reference. Our goal in this work is firstly to reproduce one of the few studies taking the distributional nature of the RE generation into account. We add to this work, by introducing a method for exploring variation in human RE choice on the basis of longitudinal corpora - substantial corpora with a single human judgement (in the process of composition) per RE. We focus on the prediction of RE types, proper name, description and pronoun. We compare evaluations made against distributions over these types with evaluations made against parallel human judgements. Our results show agreement in the evaluation of learning algorithms against distributions constructed from parallel human evaluations and from longitudinal data. T. Mark Ellison, Fahime Same |
LREC | 2 |
| 2021 | What can Neural Referential Form Selectors Learn?abstractDespite achieving encouraging results, neural Referring Expression Generation models are often thought to lack transparency.We probed neural Referential Form Selection (RFS) models to find out to what extent the linguistic features influencing the RE form are learnt and captured by state-of-the-art RFS models.The results of 8 probing tasks show that all the defined features were learnt to some extent.The probing tasks pertaining to referential status and syntactic position exhibited the highest performance.The lowest performance was achieved by the probing models designed to predict discourse structure properties beyond the sentence level. Guanyi Chen, Fahime Same, Kees van Deemter |
INLG | 2 |
| 2020 | A Linguistic Perspective on Reference: Choosing a Feature Set for Generating Referring Expressions in ContextabstractThis paper reports on a structured evaluation of feature-based Machine Learning algorithms for selecting the form of a referring expression in discourse context.Based on this evaluation, we selected seven feature sets from the literature, amounting to 65 distinct linguistic features.The features were then grouped into 9 broad classes.After building Random Forest models, we used Feature Importance Ranking and Sequential Forward Search methods to assess the "importance" of the features.Combining the results of the two methods, we propose a consensus feature set.The 6 features in our consensus set come from 4 different classes, namely grammatical role, inherent features of the referent, antecedent form and recency. Fahime Same, Kees van Deemter |
COLING | 1 |