Dimitra Gkatzia

dblp:148/4568 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0001-8568-7806ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 5 first-author · 14 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 ContrastSkill: Task-Oriented Contrastive Intermediate Training for Skill Extraction
Aleksander Bielinski, David Brazier 0001, Dimitra Gkatzia, Alistair Lawson
DATA (1)3
2025 Do My Eyes Deceive Me? A Survey of Human Evaluations of Hallucinations in NLG
abstract
Hallucinations are one of the most pressing challenges for large language models (LLMs). While numerous methods have been proposed to detect and mitigate them automatically, human evaluation continues to serve as the gold standard. However, these human evaluations of hallucinations show substantial variation in definitions, terminology, and evaluation practices. In this paper, we survey 64 studies involving human evaluation of hallucination published between 2019 and 2024, to investigate how hallucinations are currently defined and assessed. Our analysis reveals a lack of consistency in definitions and exposes several concerning methodological shortcomings. Crucial details, such as evaluation guidelines, user interface design, inter-annotator agreement metrics, and annotator demographics, are frequently under-reported or omitted altogether.
Patrícia Schmidtová, Eduardo Calò, Simone Balloccu, Dimitra Gkatzia, Rudali Huidrom, Mateusz Lango, Fahime Same, Vilém Zouhar, Saad Mahamood, Ondrej Dusek
INLG4
2025 You Are What You Write: Author re-identification privacy attacks in the era of pre-trained language models
abstract
The widespread use of pre-trained language models has revolutionised knowledge transfer in natural language processing tasks. However, there is a concern regarding potential breaches of user trust due to the risk of re-identification attacks, where malicious users could extract Personally Identifiable Information (PII) from other datasets. To assess the extent of extractable personal information on popular pre-trained models, we conduct the first wide coverage evaluation and comparison of state-of-the-art privacy-preserving algorithms on a large multi-lingual dataset for sentiment analysis annotated with demographic information (including location, age, and gender). Our results suggest a link between model complexity, pre-training data volume, and the efficacy of privacy-preserving embeddings. We found that privacy-preserving methods demonstrate greater effectiveness when applied to larger and more complex models, with improvements exceeding > 20 % over non-private baselines. Additionally, we observe that local differential privacy imposes serious performance penalties of ≈ 20 % in our test setting, which can be mitigated using hybrid or metric-DP techniques. • Compares privacy-preservation techniques across popular language models. • Analysis demonstrates significant reduction in relative attack success. • Uses a multilingual corpus, illustrating results on diverse European languages. • Includes both high-, medium-, and lower-resource languages. • To our knowledge, this is the first study to extend this analysis beyond English. • Expands on previous work (Plant et al., 2021) with MDP and CGT techniques.
Richard Plant, Valerio Giuffrida, Dimitra Gkatzia
Comput. Speech Lang.3
2025 From documents to dialogue: Context matters in common sense-enhanced task-based dialogue grounded in documents
abstract
Humans can engage in a conversation to collaborate on multi-step tasks and divert briefly to complete essential sub-tasks, such as asking for confirmation or clarification, before resuming the overall task. This communication is necessary as some knowledge in instructional documents can be implicit rather than grounded in the dialogue, meaning that people must rely on their own and others’ knowledge for problem-solving. We often attribute this capability to common sense , i.e., the assumption that interlocutors perceive behaviours , temporality , context , space and object properties in a similar way. To explore the significance of emulating such problem-solving capabilities, we developed a novel hybrid document-grounded dialogue system (DGDS) called ChefBot 1 1 https://github.com/NapierNLP/CiViL . leveraging the contextual understanding of a pre-trained language model and the structuring of a sequence-to-sequence model trained on a series of commonsense knowledge databases. In a human evaluation, the hybrid system proved more effective in capturing object knowledge (utility, appearance, storage, relationships, handling) and contextual knowledge (understanding of events and situations) compared to a rule-based baseline. A key finding of this paper is demonstrating how inferring context from different document sources enhances the dialogue by allowing richer and more fluid interaction. To our knowledge, this research is innovative in its scope as the first effort to model task-based dialogue grounded in commonsense knowledge across multiple documents. • We demonstrate though human evaluation the significant of modelling and generating information from a commonsense knowledge database during tasks. • We selected a cooking task for our project as it has lots of potential areas for dialogue complexity. • Our results show that generating new concepts form a commonsense knowledge database using an PTLM improved user understanding.
Carl Strathearn, Yanchao Yu, Dimitra Gkatzia
Expert Syst. Appl.3
2024 Exploring the impact of data representation on neural data-to-text generation
abstract
A relatively under-explored area in research on neural natural language generation is the impact of the data representation on text quality.Here we report experiments on two leading input representations for data-to-text generation: attribute-value pairs and Resource Description Framework (RDF) triples.Evaluating the performance of encoder-decoder seq2seq models as well as recent large language models (LLMs) with both automated metrics and human evaluation, we find that the input representation does not seem to have a large impact on the performance of either purpose-built seq2seq models or LLMs.Finally, we present an error analysis of the texts generated by the LLMs and provide some insights into where these models fail.
David M. Howcroft, Lewis N. Watson, Olesia Nedopas, Dimitra Gkatzia
INLG4
2024 Automatic Metrics in Natural Language Generation: A survey of Current Evaluation Practices
abstract
Patricia Schmidtova, Saad Mahamood, Simone Balloccu, Ondrej Dusek, Albert Gatt, Dimitra Gkatzia, David M. Howcroft, Ondrej Platek, Adarsa Sivaprasad. Proceedings of the 17th International Natural Language Generation Conference. 2024.
Patrícia Schmidtová, Saad Mahamood, Simone Balloccu, Ondrej Dusek, Albert Gatt, Dimitra Gkatzia, David M. Howcroft, Ondrej Plátek, Adarsa Sivaprasad
INLG6
2024 Automated Human-Readable Label Generation in Open Intent Discovery
abstract
The correct determination of user intent is key in dialog systems. However, an intent classifier often requires a large, labelled training dataset to identify a set of known intents. The creation of such a dataset is a complex and time-consuming task which usually involves humans applying clustering tools to unlabelled data, analysing the results, and creating human-readable labels for each cluster. While many Open Intent Discovery works tackle the problem of discovering clusters of common intent, few generate a human-readable label that can be used to make decisions in downstream systems. To address this, we introduce a novel candidate label extraction method then evaluate six combinations of candidate extraction and label selection methods on three datasets. We find that our extraction method produces more detailed labels than the alternatives and that high quality intent labels can be generated from unlabelled data without resorting to applying costly pre-trained language models.
Grant Anderson, Emma Hart, Dimitra Gkatzia, Ian Beaver
INTERSPEECH3
2024 An Open Intent Discovery Evaluation Framework
abstract
In the development of dialog systems the discovery of the set of target intents to identify is a crucial first step that is often overlooked.Most intent detection works assume that a labelled dataset already exists, however creating these datasets is no trivial task and usually requires humans to manually analyse, decide on intent labels and tag accordingly.The field of Open Intent Discovery (OID) addresses this problem by automating the process of grouping utterances and providing the user with the discovered intents.Our OID framework allows for the user to choose from a range of different techniques for each step in the discovery process, including the ability to extend previous works with a human-readable label generation stage.We also provide an analysis of the relationship between dataset features and optimal combination of techniques for each step to help others choose without having to explore every possible combination for their unlabelled data.
Grant Anderson, Emma Hart, Dimitra Gkatzia, Ian Beaver
SIGDIAL3
2023 Building a dual dataset of text- and image-grounded conversations and summarisation in Gàidhlig (Scottish Gaelic)
abstract
Gàidhlig (Scottish Gaelic; gd) is spoken by about 57k people in Scotland, 1 but remains an under-resourced language with respect to natural language processing in general and natural language generation (NLG) in particular.To address this gap, we developed the first datasets for Scottish Gaelic NLG, collecting both conversational and summarisation data in a single setting.Our task setup involves dialogues between a pair of proficient speakers discussing museum exhibits, grounding the conversation in images and texts.Then, each interlocutor summarises the dialogue resulting in a secondary dialogue summarisation dataset.This paper presents the dialogue and summarisation corpora, as well as the software used for data collection.The dialogue dataset consists of 43 conversations (13.7k words) and 61 summaries (2.0k words). 2
David M. Howcroft, William Lamb, Anna Groundwater, Dimitra Gkatzia
INLG4
2022 Multi3Generation: Multitask, Multilingual, Multimodal Language Generation
abstract
This paper presents the Multitask, Multilingual, Multimodal Language Generation COST Action – Multi3Generation (CA18231), an interdisciplinary network of research groups working on different aspects of language generation. This “meta-paper” will serve as reference for citations of the Action in future publications. It presents the objectives, challenges and a the links for the achieved outcomes.
Anabela Barreiro, José Guilherme Camargo de Souza, Albert Gatt, Mehul Bhatt, Elena Lloret, Aykut Erdem, Dimitra Gkatzia, Helena Moniz, Irene Russo, Fábio N. Kepler, Iacer Calixto, Marcin Paprzycki, François Portet, Isabelle Augenstein, Mirela Alhasani
EAMT7
2021 CAPE: Context-Aware Private Embeddings for Private Language Learning
abstract
Neural language models have contributed to state-of-the-art results in a number of downstream applications including sentiment analysis, intent classification and others.However, obtaining text representations or embeddings using these models risks encoding personally identifiable information learned from language and context cues that may lead to privacy leaks.To ameliorate this issue, we propose Context-Aware Private Embeddings (CAPE), a novel approach which combines differential privacy and adversarial learning to preserve privacy during training of embeddings.Specifically, CAPE firstly applies calibrated noise through differential privacy to maintain the privacy of text representations by preserving the encoded semantic links while obscuring sensitive information.Next, CAPE employs an adversarial training regime that obscures identified private variables.Experimental results demonstrate that our proposed approach is more effective in reducing private information leakage than either single intervention, with approximately a 3% reduction in attacker performance compared to the best-performing current method.
Richard Plant, Dimitra Gkatzia, Valerio Giuffrida
EMNLP (1)2
2021 Underreporting of errors in NLG output, and what to do about it
abstract
Emiel van Miltenburg, Miruna Clinciu, Ondřej Dušek, Dimitra Gkatzia, Stephanie Inglis, Leo Leppänen, Saad Mahamood, Emma Manning, Stephanie Schoch, Craig Thomson, Luou Wen. Proceedings of the 14th International Conference on Natural Language Generation. 2021.
Emiel van Miltenburg, Miruna-Adriana Clinciu, Ondrej Dusek, Dimitra Gkatzia, Stephanie Inglis, Leo Leppänen, Saad Mahamood, Emma Manning, Stephanie Schoch, Craig Thomson, Luou Wen
INLG4
2021 Chefbot: A Novel Framework for the Generation of Commonsense-enhanced Responses for Task-based Dialogue Systems
abstract
Conversational systems aim to generate responses that are accurate, relevant and engaging, either through utilising neural end-to-end models or through slot filling.Human-tohuman conversations are enhanced by not only the latest utterance of the interlocutor, but also by recalling and referring to relevant information about concepts/objects covered in the conversation so far.Such information may contain recent referred concepts, commonsense knowledge and more.A concrete scenario of such dialogues is the cooking scenario, i.e. when an artificial agent (personal assistant, robot, chatbot) and a human converse about a recipe.We will demo a novel system for commonsense enhanced response generation in the scenario of cooking, where the conversational system is able to not only provide directions for cooking step-by-step, but also display commonsense capabilities such as offering explanations on object use and recommending replacements of ingredients.
Carl Strathearn, Dimitra Gkatzia
INLG2
2021 Generating unambiguous and diverse referring expressions
Nikolaos Panagiaris, Emma Hart, Dimitra Gkatzia
Comput. Speech Lang.3
2020 Twenty Years of Confusion in Human Evaluation: NLG Needs Evaluation Sheets and Standardised Definitions
abstract
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, Verena Rieser. Proceedings of the 13th International Conference on Natural Language Generation. 2020.
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, Verena Rieser
INLG4
2020 Improving the Naturalness and Diversity of Referring Expression Generation models using Minimum Risk Training
abstract
In this paper we consider the problem of optimizing neural Referring Expression Generation (REG) models with sequence level objectives.Recently reinforcement learning (RL) techniques have been adopted to train deep end-to-end systems to directly optimize sequence-level objectives.However, there are two issues associated with RL training: (1) effectively applying RL is challenging, and (2) the generated sentences lack in diversity and naturalness due to deficiencies in the generated word distribution, smaller vocabulary size, and repetitiveness of frequent words or phrases.To alleviate these issues, we propose a novel strategy for training REG models, using minimum risk training (MRT) with maximum likelihood estimation (MLE) and we show that our approach outperforms RL w.r.t naturalness and diversity of the output.Specifically, our approach achieves an increase in CIDEr scores between 23%-57% in two datasets.We further demonstrate the robustness of the proposed method through a detailed comparison with different REG models.
Nikolaos Panagiaris, Emma Hart, Dimitra Gkatzia
INLG3
2017 Improving the Naturalness and Expressivity of Language Generation for Spanish
abstract
We present a flexible Natural Language Generation approach for Spanish, focused on the surface realisation stage, which integrates an inflection module in order to improve the naturalness and expressivity of the generated language. This inflection module inflects the verbs using an ensemble of trainable algorithms whereas the other types of words (e.g. nouns, determiners, etc) are inflected using hand-crafted rules. We show that our approach achieves 2% higher accuracy than two state-of-art inflection generation approaches. Furthermore, our proposed approach also predicts an extra feature: the inflection of the imperative mood, which was not taken into account by previous work. We also present a user evaluation, where we demonstrate that the proposed method significantly improves the perceived naturalness of the generated language.
Cristina Barros, Dimitra Gkatzia, Elena Lloret
INLG2
2016 How to talk to strangers: Generating medical reports for first-time users
abstract
We propose a novel approach for handling first-time users in the context of automatic report generation from time-series data in the health domain. Handling first-time users is a common problem for Natural Language Generation (NLG) and interactive systems in general - the system cannot adapt to users without prior interaction or user knowledge. In this paper, we propose a novel framework for generating medical reports for first-time users, using multi-objective optimisation (MOO) to account for the preferences of multiple possible user types, where the content preferences of potential users are modelled as objective functions. Our proposed approach outperforms two meaningful baselines in an evaluation with prospective users, yielding large (= .79) and medium (= .46) effect sizes respectively.
Dimitra Gkatzia, Verena Rieser, Oliver Lemon
FUZZ-IEEE1
2016 The REAL Corpus: A Crowd-Sourced Corpus of Human Generated and Evaluated Spatial References to Real-World Urban Scenes
Phil J. Bartie, William A. Mackaness, Dimitra Gkatzia, Verena Rieser
LREC3
2015 From the Virtual to the RealWorld: Referring to Objects in Real-World Spatial Scenes
abstract
Predicting the success of referring expressions (RE) is vital for real-world applications such as navigation systems.Traditionally, research has focused on studying Referring Expression Generation (REG) in virtual, controlled environments.In this paper, we describe a novel study of spatial references from real scenes rather than virtual.First, we investigate how humans describe objects in open, uncontrolled scenarios and compare our findings to those reported in virtual environments.We show that REs in real-world scenarios differ significantly to those in virtual worlds.Second, we propose a novel approach to quantifying image complexity when complete annotations are not present (e.g.due to poor object recognition capabitlities), and third, we present a model for success prediction of REs for objects in real scenes.Finally, we discuss implications for Natural Language Generation (NLG) systems and future directions.
Dimitra Gkatzia, Verena Rieser, Phil J. Bartie, William A. Mackaness
EMNLP1
2015 Exploratory Navigation for Runners Through Geographic Area Classification with Crowd-Sourced Data
abstract
Navigation when running is exploratory, characterised by both starting and ending in the same location, and iteratively foraging the environment to find areas with the most suitable running conditions. Runners do not wish to be explicitly directed, or refer to navigation aids that cause them to stop running, such as maps. Such undirected navigation is also common in other 'on-foot' scenarios, but how to support it is under-investigated. We contribute a novel method that uses crowd-sourced venue databases to rate a geographical area on its suitability to run in using linear regression. Our regression model is able to accurately predict the suitability of an area to run in (Pearson r=0.74) with a low mean error (RMSE=1.0). We outline how our method can support runners, and can be applied to other undirected navigation scenarios.
David K. McGookin, Dimitra Gkatzia, Helen Hastie
MobileHCI2
2014 Comparing Multi-label Classification with Reinforcement Learning for Summarisation of Time-series Data
abstract
We present a novel approach for automatic report generation from time-series data, in the context of student feedback generation. Our proposed methodology treats content selection as a multi-label (ML)classification problem, which takes as input time-series data and outputs a set of templates, while capturing the dependencies between selected templates. We show that this method generates output closer to the feedback that lecturers actually generated, achieving 3.5% higher accuracy and 15% higher F-score than multiple simple classifiers that keep a history of selected templates. Furthermore, we compare a ML classifier with a Reinforcement Learning (RL) approach in simulation and using ratings from real student users. We show that the different methods have different benefits, with ML being moreaccurate for predicting what was seen in the training data, whereas RL is more exploratory and slightly preferred by the students.
Dimitra Gkatzia, Helen Hastie, Oliver Lemon
ACL (1)1
2014 Finding middle ground? Multi-objective Natural Language Generation from time-series data
abstract
A Natural Language Generation (NLG) system is able to generate text from nonlinguistic data, ideally personalising the content to a user’s specific needs. In some cases, however, there are multiple stakeholders with their own individual goals, needs and preferences. In this paper, we explore the feasibility of combining the preferences of two different user groups, lecturers and students, when generatingsummaries in the context of student feedback generation. The preferences of each user group are modelled as a multivariateoptimisation function, therefore the task of generation is seen as a multi-objective (MO) optimisation task, where the two functions are combined into one. This initial study shows that treating the preferences of each user group equally smooths the weights of the MO function, in a way that preferred content of the user groups isnot presented in the generated summary.
Dimitra Gkatzia, Helen Hastie, Oliver Lemon
EACL1
2014 Multi-adaptive Natural Language Generation using Principal Component Regression
abstract
We present FeedbackGen, a system that uses a multi-adaptive approach to Natural Language Generation. With the term ‘multi-adaptive’, we refer to a system that is able to adapt its content to different user groups simultaneously, in our case adapting to both lecturers and students. We present a novel approach to student feedback generation, which simultaneously takes into account the preferences of lecturers and students when determining the content to be conveyed in a feedback summary. In this framework, we utilise knowledge derived from ratings on feedback summaries by extracting the most relevant features using Principal Component Regression (PCR) analysis. We then model a reward function that is used for training a Reinforcement Learning agent. Our results with students suggest that, from the students’ perspective, such an approach can generate more preferable summaries than a purely lecturer-adapted approach.
Dimitra Gkatzia, Helen Hastie, Oliver Lemon
INLG1