EDBT 2026 Demo / reviewers in the wild / expert
Albert Gatt
dblp:38/3390
· DBLP profile ↗
67ranked-venue papers
16as first author
21since 2021 · last 2027
0000-0001-6388-8244ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 61 · 16 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | TABERTA: A Language Model for Dataset Discovery
Enas Khwaileh, Leonard Traeger, Albert Gatt, Yannis Velegrakis |
EDBT | 3 |
| 2026 | Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTaskabstractCollaborative dialogue relies on participants incrementally establishing common ground, yet in asymmetric settings they may believe they agree while referring to different entities. We introduce a perspectivist annotation scheme for the HCRC MapTask corpus (Anderson et al., 1991) that separately captures speaker and addressee grounded interpretations for each reference expression, enabling us to trace how understanding emerges, diverges, and repairs over time. Using a scheme-constrained LLM annotation pipeline, we obtain 13k annotated reference expressions with reliability estimates and analyze the resulting understanding states. The results show that full misunderstandings are rare once lexical variants are unified, but multiplicity discrepancies systematically induce divergences, revealing how apparent grounding can mask referential misalignment. Our framework provides both a resource and an analytic lens for studying grounded misunderstanding and for evaluating (V)LLMs' capacity to model perspective-dependent grounding in collaborative dialogue. Nan Li 0087, Albert Gatt, Massimo Poesio |
LREC | 2 |
| 2026 | Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric DialogueabstractIn collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be established through interaction. We investigate whether vision-language models (VLMs) can distinguish what could be shared from what has been shared between dialogue participants through grounding. We formulate this as an interpretation-matching task on 13,077 annotated reference expressions from HCRC MapTask dialogues, and evaluate VLMs under systematically controlled manipulations of dialogue context and map-information access. Our results show that providing authentic map images improves overall performance but shifts models toward over-predicting alignment. Textual descriptions of the same map content reproduce this bias, while non-informative images suppress alignment predictions entirely, indicating that the bias is driven by task-relevant map content, not the visual channel. This improvement comes at the cost of degraded accuracy on non-aligned cases. Calibration analysis and reference-chain tracking further suggest that models rely on static referential cues on the maps rather than tracking how grounding unfolds through dialogue history. We observe these patterns most clearly in Qwen3-VL-8B-Instruct and, to varying degrees, in four additional models from two architecture families. In models that exhibit the bias, map content, whether presented visually or textually, is treated as evidence of mutual understanding, conflating potential with established common ground. Nan Li 0087, Albert Gatt, Massimo Poesio |
SIGDIAL | 2 |
| 2026 | Breaking the Script: Do Role-Playing Agents Maintain Goal Alignment under Distraction?abstractAs large language models (LLMs) are increasingly deployed as role-playing agents in educational and professional training simulations, their susceptibility to user-induced distraction threatens their pedagogical utility. We formalise goal-competing distraction as a controlled evaluation paradigm and introduce a simulation framework that captures both immediate reactions and multi-turn trajectories under targeted distraction, using an LLM-based user simulator. Building upon this framework, we evaluate agent behaviour across three models: Gemini-2.0-Flash, Llama-3.3-70B-Instruct, and Llama-3.1-8B-Instruct. Our findings reveal a critical trade-off dependent on model scale. While larger models tend to remain socially responsive and more frequently engage with distractor topics, the smaller model shows rigid goal adherence by resisting and rejecting distraction. Although redirection is the most common initial response, subsequent trajectories diverge substantially. The inclusion of explicit dialogue state demonstrates model-dependent effects, acting as a stabilising anchor for smaller models but providing limited benefit for larger ones. These results suggest that maintaining goal alignment in role-playing agents requires explicitly managing the trade-off between conversational responsiveness and goal adherence. Dongxu Lu, Albert Gatt, Johan Jeuring |
SIGDIAL | 2 |
| 2026 | Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential CommunicationabstractAbstract To what degree and under what conditions do VLMs rely on scene context when generating references to objects? To address this question, we introduce the Common Objects Out-of-Context (COOCo) dataset and conduct experiments on several VLMs under different degrees of scene–object congruency and noise. We find that models leverage scene context adaptively, depending on scene-object semantic relatedness and noise level. Based on these consistent trends across models, we turn to the question of how VLM attention patterns change as a function of target-scene semantic fit, and to what degree these patterns are predictive of categorisation accuracy. We find that successful object categorisation is associated with increased mid-layer attention to the target. We also find a non-monotonic dependency on semantic fit, with attention dropping at moderate fit and increasing for both low and high fit. These results suggest that VLMs dynamically balance local and contextual information for reference generation. Dataset and code are available here: https://github.com/cs-nlp-uu/scenereg. Filippo Merlo, Ece Takmaz, Albert Gatt |
Trans. Assoc. Comput. Linguistics | 4 |
| 2025 | Disentangling the Roles of Representation and Selection in Data PruningabstractData pruning, selecting small but impactful subsets, offers a promising way to efficiently scale NLP model training. However, existing methods often involve many different design choices, which have not been systematically studied. This limits future developments. In this work, we decompose data pruning into two key components: the data representation and the selection algorithm, and we systematically analyze their influence on the selection of instances. Our theoretical and empirical results highlight the crucial role of representations: better representations, e.g., training gradients, generally lead to a better selection of instances, regardless of the chosen selection algorithm. Furthermore, different selection algorithms excel in different settings, and none consistently outperforms the others. Moreover, the selection algorithms do not always align with their intended objectives: for example, algorithms designed for the same objective can select drastically different instances, highlighting the need for careful evaluation. Yupei Du, Yingjin Song, Hugh Mee Wong, Daniil Ignatev, Albert Gatt, Dong Nguyen 0002 |
ACL (1) | 5 |
| 2025 | CV-Probes: Studying the interplay of lexical and world knowledge in visually grounded verb understanding
Ivana Benová, Michal Gregor, Albert Gatt |
CogSci | 3 |
| 2025 | FTFT: Efficient and Robust Fine-Tuning by Transferring Training DynamicsabstractDespite the massive success of fine-tuning Pre-trained Language Models (PLMs), they remain susceptible to out-of-distribution input. Dataset cartography is a simple yet effective dual-model approach that improves the robustness of fine-tuned PLMs. It involves fine-tuning a model on the original training set (i.e. reference model), selecting a subset of important training instances based on the training dynamics, % of the reference model, and fine-tuning again only on these selected examples (i.e. main model). However, this approach requires fine-tuning the same model twice, which is computationally expensive for large PLMs. In this paper, we show that 1) training dynamics are highly transferable across model sizes and pre-training methods, and that 2) fine-tuning main models using these selected training instances achieves higher training efficiency than empirical risk minimization (ERM). Building on these observations, we propose a novel fine-tuning approach: Fine-Tuning by transFerring Training dynamics (FTFT). Compared with dataset cartography, FTFT uses more efficient reference models and aggressive early stopping. FTFT achieves robustness improvements over ERM while lowering the training cost by up to ~50% Yupei Du, Albert Gatt, Dong Nguyen 0002 |
COLING | 2 |
| 2025 | Incorporating Formulaicness in the Automatic Evaluation of Naturalness: A Case Study in Logic-to-Text GenerationabstractData-to-text natural language generation (NLG) models may produce outputs that closely mirror the structure of their input. We introduce formulaicness as a measure of the output-to-input structural resemblance, proposing it as an enhancement for reference-less naturalness evaluation. Focusing on logic-to-text generation, we construct a dataset and train a regressor to predict formulaicness scores. We collect human judgments on naturalness and examine how incorporating formulaicness into existing metrics affects alignment with these judgments. Eduardo Calò, Guanyi Chen, Elias Stengel-Eskin, Albert Gatt, Kees van Deemter |
INLG | 4 |
| 2025 | References Matter: Investigating the Impact of Reference Set Variation on Summarization EvaluationabstractHuman language production exhibits remarkable richness and variation, reflecting diverse communication styles and intents. However, this variation is often overlooked in summarization evaluation. While having multiple reference summaries is known to improve correlation with human judgments, the impact of the reference set on reference-based metrics has not been systematically investigated. This work examines the sensitivity of widely used reference-based metrics in relation to the choice of reference sets, analyzing three diverse multi-reference summarization datasets: SummEval, GUMSum, and DUC2004. We demonstrate that many popular metrics exhibit significant instability. This instability is particularly concerning for n-gram-based metrics like ROUGE, where model rankings vary depending on the reference sets, undermining the reliability of model comparisons. We also collect human judgments on LLM outputs for genre-diverse data and examine their correlation with metrics to supplement existing findings beyond newswire summaries, finding weak-to-no correlation. Taken together, we recommend incorporating reference set variation into summarization evaluation to enhance consistency alongside correlation with human judgments, especially when evaluating LLMs. Silvia Casola, Yang Janet Liu, Siyao Peng, Oliver Kraus, Albert Gatt, Barbara Plank |
INLG | 5 |
| 2025 | Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play DialoguesabstractEvaluating large language models (LLMs) in long-form, knowledge-grounded role-play dialogues remains challenging. This study compares LLM-generated and human-authored responses in multi-turn professional training simulations through human evaluation (N = 38) and automated LLM-as-a-judge assessment. Human evaluation revealed significant degradation in LLM-generated response quality across turns, particularly in naturalness, context maintenance and overall quality, while human-authored responses progressively improved. In line with this finding, participants also indicated a consistent preference for human-authored dialogue. These human judgements were validated by our automated LLM-as-a-judge evaluation, where GEMINI 2.0 FLASH achieved strong alignment with human evaluators on both zero-shot pairwise preference and stochastic 6-shot construct ratings, confirming the widening quality gap between LLM and human responses over time. Our work contributes a multi-turn benchmark exposing LLM degradation in knowledge-grounded role-play dialogues and provides a validated hybrid evaluation framework to guide the reliable integration of LLMs in training simulations. Dongxu Lu, Johan Jeuring, Albert Gatt |
INLG | 3 |
| 2024 | A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic InferencesabstractThe reasoning abilities of Large Language Models (LLMs) are becoming a central focus of study in NLP.In this paper, we consider the case of syllogistic reasoning, an area of deductive reasoning studied extensively in logic and cognitive psychology.Previous research has shown that pre-trained LLMs exhibit reasoning biases, such as content effects, avoid answering that no conclusion follows, display human-like difficulties, and struggle with multi-step reasoning.We contribute to this research line by systematically investigating the effects of chainof-thought reasoning, in-context learning (ICL), and supervised fine-tuning (SFT) on syllogistic reasoning, considering syllogisms with conclusions that support or violate world knowledge, as well as ones with multiple premises.Crucially, we go beyond the standard focus on accuracy, with an in-depth analysis of the conclusions generated by the models.Our results suggest that the behavior of pre-trained LLMs can be explained by heuristics studied in cognitive science and that both ICL and SFT improve model performance on valid inferences, although only the latter mitigates most reasoning biases without harming model consistency. Leonardo Bertolazzi, Albert Gatt, Raffaella Bernardi |
EMNLP | 2 |
| 2024 | ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language ModelsabstractWith the ever-increasing popularity of pretrained Video-Language Models (VidLMs), there is a pressing need to develop robust evaluation methodologies that delve deeper into their visio-linguistic capabilities. To address this challenge, we present ViLMA (Video Language Model Assessment), a task-agnostic benchmark that places the assessment of fine-grained capabilities of these models on a firm footing. Task-based evaluations, while valuable, fail to capture the complexities and specific temporal aspects of moving images that VidLMs need to process. Through carefully curated counterfactuals, ViLMA offers a controlled evaluation suite that sheds light on the true potential of these models, as well as their performance gaps compared to human-level understanding. ViLMA also includes proficiency tests, which assess basic capabilities deemed essential to solving the main counterfactual tests. We show that current VidLMs’ grounding abilities are no better than those of vision-language models which use static images. This is especially striking once the performance on proficiency tests is factored in. Our benchmark serves as a catalyst for future research on VidLMs, helping to highlight areas that still need to be explored. Ilker Kesen, Andrea Pedrotti, Mustafa Dogan, Michele Cafagna, Emre Can Acikgoz, Letitia Parcalabescu, Iacer Calixto, Anette Frank, Albert Gatt, Aykut Erdem, Erkut Erdem |
ICLR | 9 |
| 2024 | Automatic Metrics in Natural Language Generation: A survey of Current Evaluation PracticesabstractPatricia Schmidtova, Saad Mahamood, Simone Balloccu, Ondrej Dusek, Albert Gatt, Dimitra Gkatzia, David M. Howcroft, Ondrej Platek, Adarsa Sivaprasad. Proceedings of the 17th International Natural Language Generation Conference. 2024. Patrícia Schmidtová, Saad Mahamood, Simone Balloccu, Ondrej Dusek, Albert Gatt, Dimitra Gkatzia, David M. Howcroft, Ondrej Plátek, Adarsa Sivaprasad |
INLG | 5 |
| 2024 | Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive LearningabstractVisual storytelling systems generate multisentence stories from image sequences.In this task, capturing contextual information and bridging visual variation bring additional challenges.We propose a simple yet effective framework that leverages the generalization capabilities of pretrained foundation models, only training a lightweight vision-language mapping network to connect modalities, while incorporating context to enhance coherence.We introduce a multimodal contrastive objective that also improves visual relevance and story informativeness.Extensive experimental results, across both automatic metrics and human evaluations, demonstrate that the stories generated by our framework are diverse, coherent, informative, and interesting. Yingjin Song, Denis Paperno, Albert Gatt |
INLG | 3 |
| 2023 | HL Dataset: Visually-grounded Description of Scenes, Actions and RationalesabstractCurrent captioning datasets focus on object-centric captions, describing the visible objects in the image, often ending up stating the obvious (for humans), e.g. "people eating food in a park". Although these datasets are useful to evaluate the ability of Vision & Language models to recognize and describe visual content, they do not support controlled experiments involving model testing or fine-tuning, with more high-level captions, which humans find easy and natural to produce. For example, people often describe images based on the type of scene they depict ("people at a holiday resort") and the actions they perform ("people having a picnic"). Such concepts are based on personal experience and contribute to forming common sense assumptions. We present the High-Level Dataset, a dataset extending 14997 images from the COCO dataset, aligned with a new set of 134,973 human-annotated (high-level) captions collected along three axes: scenes, actions and rationales. We further extend this dataset with confidence scores collected from an independent set of readers, as well as a set of narrative captions generated synthetically, by combining each of the three axes. We describe this dataset and analyse it extensively. We also present baseline results for the High-Level Captioning task. Michele Cafagna, Kees van Deemter, Albert Gatt |
INLG | 3 |
| 2022 | VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic PhenomenaabstractLetitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank, Iacer Calixto, Albert Gatt. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Letitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank, Iacer Calixto, Albert Gatt |
ACL (1) | 6 |
| 2022 | Multi3Generation: Multitask, Multilingual, Multimodal Language GenerationabstractThis paper presents the Multitask, Multilingual, Multimodal Language Generation COST Action – Multi3Generation (CA18231), an interdisciplinary network of research groups working on different aspects of language generation. This “meta-paper” will serve as reference for citations of the Action in future publications. It presents the objectives, challenges and a the links for the achieved outcomes. Anabela Barreiro, José Guilherme Camargo de Souza, Albert Gatt, Mehul Bhatt, Elena Lloret, Aykut Erdem, Dimitra Gkatzia, Helena Moniz, Irene Russo, Fábio N. Kepler, Iacer Calixto, Marcin Paprzycki, François Portet, Isabelle Augenstein, Mirela Alhasani |
EAMT | 3 |
| 2022 | Measuring Model Understandability by means of Shapley Additive ExplanationsabstractIn this work we link the understandability of machine learning models to the complexity of their SHapley Additive exPlanations (SHAP). Thanks to this reframing we introduce two novel metrics for understandability: SHAP Length and SHAP Interaction Length. These are model-agnostic, efficient, intuitive and theoretically grounded metrics that are anchored in well-established game-theoretic and psychological principles. We show how these metrics resonate with other model-specific ones and how they can enable a fairer comparison of epistemically different models in the context of Explainable Artificial Intelligence. In particular, we quantitatively explore the understandability-performance tradeoff of different models which are applied to both classification and regression problems. Reported results suggest the value of the new metrics in the context of automated machine learning and multi-objective optimisation. Ettore Mariotti, Jose Maria Alonso-Moral, Albert Gatt |
FUZZ-IEEE | 3 |
| 2022 | Neural Natural Language Generation: A Survey on Multilinguality, Multimodality, Controllability and LearningabstractDeveloping artificial learning systems that can understand and generate natural language has been one of the long-standing goals of artificial intelligence. Recent decades have witnessed an impressive progress on both of these problems, giving rise to a new family of approaches. Especially, the advances in deep learning over the past couple of years have led to neural approaches to natural language generation (NLG). These methods combine generative language learning techniques with neural-networks based frameworks. With a wide range of applications in natural language processing, neural NLG (NNLG) is a new and fast growing field of research. In this state-of-the-art report, we investigate the recent developments and applications of NNLG in its full extent from a multidimensional view, covering critical perspectives such as multimodality, multilinguality, controllability and learning strategies. We summarize the fundamental building blocks of NNLG approaches from these aspects and provide detailed reviews of commonly used preprocessing steps and basic neural architectures. This report also focuses on the seminal applications of these NNLG models such as machine translation, description generation, automatic speech recognition, abstractive summarization, text simplification, question answering and generation, and dialogue generation. Finally, we conclude with a thorough discussion of the described frameworks by pointing out some open research directions. Erkut Erdem, Menekse Kuyu, Semih Yagcioglu, Anette Frank, Letitia Parcalabescu, Barbara Plank, Andrii Babii, Oleksii Turuta, Aykut Erdem, Iacer Calixto, Elena Lloret, Elena Apostol, Ciprian-Octavian Truica, Branislava Sandrih, Sanda Martincic-Ipsic, Gábor Berend, Albert Gatt, Grazina Korvel |
J. Artif. Intell. Res. | 17 |
| 2021 | Human evaluation of automatically generated text: Current trends and best practice guidelinesabstractCurrently, there is little agreement as to how Natural Language Generation (NLG) systems should be evaluated, with a particularly high degree of variation in the way that human evaluation is carried out. This paper provides an overview of how (mostly intrinsic) human evaluation is currently conducted and presents a set of best practices, grounded in the literature. These best practices are also linked to the stages that researchers go through when conducting an evaluation research (planning stage; execution and release stage), and the specific steps in these stages. With this paper, we hope to contribute to the quality and consistency of human evaluations in NLG. Chris van der Lee, Albert Gatt, Emiel van Miltenburg, Emiel Krahmer |
Comput. Speech Lang. | 2 |
| 2020 | Gradations of Error Severity in Automatic Image DescriptionsabstractEarlier research has shown that evaluation metrics based on textual similarity (e.g., BLEU, CIDEr, Meteor) do not correlate well with human evaluation scores for automatically generated text.We carried out an experiment with Chinese speakers, where we systematically manipulated image descriptions to contain different kinds of errors.Because our manipulated descriptions form minimal pairs with the reference descriptions, we are able to assess the impact of different kinds of errors on the perceived quality of the descriptions.Our results show that different kinds of errors elicit significantly different evaluation scores, even though all erroneous descriptions differ in only one character from the reference descriptions.Evaluation metrics based solely on textual similarity are unable to capture these differences, which (at least partially) explains their poor correlation with human judgments.Our work provides the foundations for future work, where we aim to understand why different errors are seen as more or less severe. Emiel van Miltenburg, Wei-Ting Lu, Emiel Krahmer, Albert Gatt, Guanyi Chen, Kees van Deemter |
INLG | 4 |
| 2020 | Automatic Removal of Identifying Information in Official EU Languages for Public Administrations: The MAPA ProjectabstractThe European MAPA (Multilingual Anonymisation for Public Administrations) project aims at developing an open-source solution for automatic de-identification of medical and legal documents. We introduce here the context, partners and aims of the project, and report on preliminary results. Lucie Gianola, Eriks Ajausks, Victoria Arranz, Chomicha Bendahman, Laurent Bié, Claudia Borg, Aleix Cerdà-i-Cucó, Khalid Choukri, Montse Cuadros, Ona de Gibert Bonet, Hans Degroote, Elena Edelman, Thierry Etchegoyhen, Ángela Franco Torres, Mercedes García Hernandez, Aitor García-Pablos, Albert Gatt, Cyril Grouin, Manuel Herranz, Alejandro Kohan, Thomas Lavergne, Maite Melero, Patrick Paroubek, Mickaël Rigault, Mike Rosner, Roberts Rozis, Lonneke van der Plas, Rinalds Viksna, Pierre Zweigenbaum |
JURIX | 17 |
| 2020 | Annotating for Hate Speech: The MaNeCo Corpus and Some Input from Critical Discourse AnalysisabstractThis paper presents a novel scheme for the annotation of hate speech in corpora of Web 2.0 commentary. The proposed scheme is motivated by the critical analysis of posts made in reaction to news reports on the Mediterranean migration crisis and LGBTIQ+ matters in Malta, which was conducted under the auspices of the EU-funded C.O.N.T.A.C.T. project. Based on the realisation that hate speech is not a clear-cut category to begin with, appears to belong to a continuum of discriminatory discourse and is often realised through the use of indirect linguistic means, it is argued that annotation schemes for its detection should refrain from directly including the label ‘hate speech,’ as different annotators might have different thresholds as to what constitutes hate speech and what not. In view of this, we propose a multi-layer annotation scheme, which is pilot-tested against a binary ±hate speech classification and appears to yield higher inter-annotator agreement. Motivating the postulation of our scheme, we then present the MaNeCo corpus on which it will eventually be used; a substantial corpus of on-line newspaper comments spanning 10 years. Stavros Assimakopoulos, Rebecca Vella Muskat, Lonneke van der Plas, Albert Gatt |
LREC | 4 |
| 2020 | MASRI-HEADSET: A Maltese Corpus for Speech RecognitionabstractMaltese, the national language of Malta, is spoken by approximately 500,000 people. Speech processing for Maltese is still in its early stages of development. In this paper, we present the first spoken Maltese corpus designed purposely for Automatic Speech Recognition (ASR). The MASRI-HEADSET corpus was developed by the MASRI project at the University of Malta. It consists of 8 hours of speech paired with text, recorded by using short text snippets in a laboratory environment. The speakers were recruited from different geographical locations all over the Maltese islands, and were roughly evenly distributed by gender. This paper also presents some initial results achieved in baseline experiments for Maltese ASR using Sphinx and Kaldi. The MASRI HEADSET Corpus is publicly available for research/academic purposes. Carlos Daniel Hernandez Mena, Albert Gatt, Andrea DeMarco, Claudia Borg, Lonneke van der Plas, Amanda Muscat, Ian Padovani |
LREC | 2 |
| 2019 | You Write like You Eat: Stylistic Variation as a Predictor of Social StratificationabstractInspired by Labov's seminal work on stylistic variation as a function of social stratification, we develop and compare neural models that predict a person's presumed socio-economic status, obtained through distant supervision, from their writing style on social media.The focus of our work is on identifying the most important stylistic parameters to predict socioeconomic group.In particular, we show the effectiveness of morpho-syntactic features as stylistic predictors of socio-economic group, in contrast to lexical features, which are good predictors of topic. Angelo Basile, Albert Gatt, Malvina Nissim |
ACL (1) | 2 |
| 2019 | Visually grounded generation of entailments from premisesabstractNatural Language Inference (NLI) is the task of determining the semantic relationship between a premise and a hypothesis.In this paper, we focus on the generation of hypotheses from premises in a multimodal setting, to generate a sentence (hypothesis) given an image and/or its description (premise) as the input.The main goals of this paper are (a) to investigate whether it is reasonable to frame NLI as a generation task; and (b) to consider the degree to which grounding textual premises in visual information is beneficial to generation.We compare different neural architectures, showing through automatic and human evaluation that entailments can indeed be generated successfully.We also show that multimodal models outperform unimodal models in this task, albeit marginally. Somayeh Jafaritazehjani, Albert Gatt, Marc Tanti |
INLG | 2 |
| 2019 | Best practices for the human evaluation of automatically generated textabstractCurrently, there is little agreement as to how Natural Language Generation (NLG) systems should be evaluated, with a particularly high degree of variation in the way that human evaluation is carried out.This paper provides an overview of how human evaluation is currently conducted, and presents a set of best practices, grounded in the literature.With this paper, we hope to contribute to the quality and consistency of human evaluations in NLG. Chris van der Lee, Albert Gatt, Emiel van Miltenburg, Sander Wubben, Emiel Krahmer |
INLG | 2 |
| 2018 | Grounded Textual EntailmentabstractCapturing semantic relations between sentences, such as entailment, is a long-standing challenge for computational semantics. Logic-based models analyse entailment in terms of possible worlds (interpretations, or situations) where a premise P entails a hypothesis H iff in all worlds where P is true, H is also true. Statistical models view this relationship probabilistically, addressing it in terms of whether a human would likely infer H from P. In this paper, we wish to bridge these two perspectives, by arguing for a visually-grounded version of the Textual Entailment task. Specifically, we ask whether models can perform better if, in addition to P and H, there is also an image (corresponding to the relevant “world” or “situation”). We use a multimodal version of the SNLI dataset (Bowman et al., 2015) and we compare “blind” and visually-augmented models of textual entailment. We show that visual information is beneficial, but we also conduct an in-depth error analysis that reveals that current multimodal models are not performing “grounding” in an optimal fashion. Hoa Trong Vu, Claudio Greco 0002, Aliia Erofeeva, Somayeh Jafaritazehjan, Guido Linders, Marc Tanti, Alberto Testoni, Raffaella Bernardi, Albert Gatt |
COLING | 9 |
| 2018 | Specificity measures and referenceabstractIn this paper we study empirically the validity of measures of referential success for referring expressions involving gradual properties.More specifically, we study the ability of several measures of referential success to predict the success of a user in choosing the right object, given a referring expression.Experimental results indicate that certain fuzzy measures of success are able to predict human accuracy in reference resolution.Such measures are therefore suitable for the estimation of the success or otherwise of a referring expression produced by a generation algorithm, especially in case the properties in a domain cannot be assumed to have crisp denotations. Albert Gatt, Nicolás Marín, Gustavo Rivas-Gervilla, Daniel Sánchez 0001 |
INLG | 1 |
| 2018 | Meteorologists and Students: A resource for language grounding of geographical descriptorsabstractWe present a data resource which can be useful for research purposes on language grounding tasks in the context of geographical referring expression generation.The resource is composed of two data sets that encompass 25 different geographical descriptors and a set of associated graphical representations, drawn as polygons on a map by two groups of human subjects: teenage students and expert meteorologists. Alejandro Ramos-Soto, Ehud Reiter, Kees van Deemter, Jose Maria Alonso-Moral, Albert Gatt |
INLG | 5 |
| 2018 | Face2Text: Collecting an Annotated Image Description Corpus for the Generation of Rich Face Descriptions
Albert Gatt, Marc Tanti, Adrian Muscat, Patrizia Paggio, Reuben A. Farrugia, Claudia Borg, Kenneth P. Camilleri, Mike Rosner, Lonneke van der Plas |
LREC | 1 |
| 2018 | Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluationabstractThis paper surveys the current state of the art in Natural Language Generation (NLG), defined as the task of generating text or speech from non-linguistic input. A survey of NLG is timely in view of the changes that the field has undergone over the past two decades, especially in relation to new (usually data-driven) methods, as well as new applications of NLG technology. This survey therefore aims to (a) give an up-to-date synthesis of research on the core tasks in NLG and the architectures adopted in which such tasks are organised; (b) highlight a number of recent research topics that have arisen partly as a result of growing synergies between NLG and other areas of artificial intelligence; (c) draw attention to the challenges in NLG evaluation, relating them to similar challenges faced in other areas of NLP, with an emphasis on different evaluation methods and the relationships between them. Albert Gatt, Emiel Krahmer |
J. Artif. Intell. Res. | 1 |
| 2018 | Where to put the image in an image caption generatorabstractAbstract When a recurrent neural network (RNN) language model is used for caption generation, the image information can be fed to the neural network either by directly incorporating it in the RNN – conditioning the language model by ‘injecting’ image features – or in a layer following the RNN – conditioning the language model by ‘merging’ image features. While both options are attested in the literature, there is as yet no systematic comparison between the two. In this paper, we empirically show that it is not especially detrimental to performance whether one architecture is used or another. The merge architecture does have practical advantages, as conditioning by merging allows the RNN’s hidden state vector to shrink in size by up to four times. Our results suggest that the visual and linguistic modalities for caption generation need not be jointly encoded by the RNN as that yields large, memory-intensive models with few tangible advantages in performance; rather, the multimodal integration should be delayed to a subsequent stage. Marc Tanti, Albert Gatt, Kenneth P. Camilleri |
Nat. Lang. Eng. | 2 |
| 2017 | An empirical approach for modeling fuzzy geographical descriptorsabstractWe present a novel heuristic approach that defines fuzzy geographical descriptors using data gathered from a survey with human subjects. The participants were asked to provide graphical interpretations of the descriptors `north' and `south' for the Galician region (Spain). Based on these interpretations, our approach builds fuzzy descriptors that are able to compute membership degrees for geographical locations. We evaluated our approach in terms of efficiency and precision. The fuzzy descriptors are meant to be used as the cornerstones of a geographical referring expression generation algorithm that is able to linguistically characterize geographical locations and regions. This work is also part of a general research effort that intends to establish a methodology which reunites the empirical studies traditionally practiced in data-to-text and the use of fuzzy sets to model imprecision and vagueness in words and expressions for text generation purposes. Alejandro Ramos-Soto, Jose Maria Alonso-Moral, Ehud Reiter, Kees van Deemter, Albert Gatt |
FUZZ-IEEE | 5 |
| 2017 | What is the Role of Recurrent Neural Networks (RNNs) in an Image Caption Generator?abstractIn neural image captioning systems, a recurrent neural network (RNN) is typically viewed as the primary 'generation' component.This view suggests that the image features should be 'injected' into the RNN.This is in fact the dominant view in the literature.Alternatively, the RNN can instead be viewed as only encoding the previously generated words.This view suggests that the RNN should only be used to encode linguistic features and that only the final representation should be 'merged' with the image features at a later stage.This paper compares these two architectures.We find that, in general, late merging outperforms injection, suggesting that RNNs are better viewed as encoders, rather than generators. Marc Tanti, Albert Gatt, Kenneth P. Camilleri |
INLG | 2 |
| 2016 | Viewing time affects overspecification: Evidence for two strategies of attribute selection during reference production
Ruud Koolen, Albert Gatt, Roger P. G. van Gompel, Emiel Krahmer, Kees van Deemter |
CogSci | 2 |
| 2016 | The Role of Graduality for Referring Expression Generation in Visual Scenes
Albert Gatt, Nicolás Marín, François Portet, Daniel Sánchez 0001 |
IPMU (1) | 1 |
| 2016 | Reasoning About Partial ContractsabstractNatural language techniques have been employed in attempts to automatically translate legal texts, and specifically contracts, into formal models that allow automatic reasoning. However, such techniques suffer from incomplete coverage, typically resulting in parts of the text being left uninterpreted, and which, in turn, may result in the formal models failing to identify potential problems due to these unknown parts. In this paper we present a formal approach to deal with partiality, by syntactically and semantically permitting unknown subcontracts in an action-based deontic logic, with accompanying formal analysis techniques to enable reasoning under incomplete knowledge. Shaun Azzopardi, Albert Gatt, Gordon J. Pace |
JURIX | 2 |
| 2016 | Multilingual generation of uncertain temporal expressions from data: A study of a possibilistic formalism and its consistency with human subjective evaluations
Albert Gatt, François Portet |
Fuzzy Sets Syst. | 1 |
| 2014 | Symposium: The Role of Alternatives in Pragmatic Inference
Judith Degen, Noah D. Goodman, Roni Katzir, David Barner, Albert Gatt |
CogSci | 5 |
| 2014 | Learning when to point: A data-driven approach
Albert Gatt, Patrizia Paggio |
COLING | 1 |
| 2014 | Crowd-sourcing evaluation of automatically acquired, morphologically related word groupings
Claudia Borg, Albert Gatt |
LREC | 2 |
| 2013 | Workshop Proposal: PRE-CogSci 2013: Bridging the gap between cognitive and computational approaches to reference
Albert Gatt, Roger P. G. van Gompel, Ellen Gurman Bard, Emiel Krahmer, Kees van Deemter |
CogSci | 1 |
| 2013 | Production of referring expressions: Preference trumps discrimination
Albert Gatt, Emiel Krahmer, Roger P. G. van Gompel, Kees van Deemter |
CogSci | 1 |
| 2012 | Does domain size impact speech onset time during reference production?
Albert Gatt, Roger P. G. van Gompel, Emiel Krahmer, Kees van Deemter |
CogSci | 1 |
| 2012 | A Repository of Data and Evaluation Resources for Natural Language Generation
Anya Belz, Albert Gatt |
LREC | 2 |
| 2012 | Incorporating an Error Corpus into a Spellchecker for Maltese
Mike Rosner, Albert Gatt, Andrew Attard, Jan Joachimsen |
LREC | 2 |
| 2012 | Automatic generation of natural language nursing shift summaries in neonatal intensive care: BT-Nurse
Jim Hunter, Yvonne Freer, Albert Gatt, Ehud Reiter, Somayajulu Sripada, Cindy Sykes |
Artif. Intell. Medicine | 3 |
| 2011 | PRE-CogSci 2011 - Bridging the gap between computational, empirical and theoretical approaches to reference
Kees van Deemter, Albert Gatt, Roger P. G. van Gompel, Emiel Krahmer |
CogSci | 2 |
| 2011 | Attribute preference and priming in reference production: Experimental evidence and computational modeling
Albert Gatt, Martijn Goudbeek, Emiel Krahmer |
CogSci | 1 |
| 2011 | BT-Nurse: computer generation of natural language shift summaries from complex heterogeneous medical dataabstractThe BT-Nurse system uses data-to-text technology to automatically generate a natural language nursing shift summary in a neonatal intensive care unit (NICU). The summary is solely based on data held in an electronic patient record system, no additional data-entry is required. BT-Nurse was tested for two months in the Royal Infirmary of Edinburgh NICU. Nurses were asked to rate the understandability, accuracy, and helpfulness of the computer-generated summaries; they were also asked for free-text comments about the summaries. The nurses found the majority of the summaries to be understandable, accurate, and helpful (p<0.001 for all measures). However, nurses also pointed out many deficiencies, especially with regard to extra content they wanted to see in the computer-generated summaries. In conclusion, natural language NICU shift summaries can be automatically generated from an electronic patient record, but our proof-of-concept software needs considerable additional development work before it can be deployed. Jim Hunter, Yvonne Freer, Albert Gatt, Ehud Reiter, Somayajulu Sripada, Cindy Sykes, Dave Westwater |
J. Am. Medical Informatics Assoc. | 3 |
| 2010 | Generation Challenges 2010 Preface
Anya Belz, Albert Gatt, Alexander Koller |
INLG | 2 |
| 2010 | Textual Properties and Task-based Evaluation: Investigating the Role of Surface Properties, Structure and Content
Albert Gatt, François Portet |
INLG | 1 |
| 2009 | Automatic generation of textual summaries from neonatal intensive care data
François Portet, Ehud Reiter, Albert Gatt, Jim Hunter, Somayajulu Sripada, Yvonne Freer, Cindy Sykes |
Artif. Intell. | 3 |
| 2008 | Summarising Complex ICU Data in Natural Language
Jim Hunter, Yvonne Freer, Albert Gatt, Robert H. Logie, Neil McIntosh, Marian van der Meulen, François Portet, Ehud Reiter, Somayajulu Sripada, Cindy Sykes |
AMIA | 3 |
| 2008 | Using Natural Language Generation Technology to Improve Information Flows in Intensive Care UnitsabstractIn the drive to improve patient safety, patients in modern intensive care units are closely monitored with the generation of very large volumes of data. Unless the data are further processed, it is difficult for medical and nursing staff to assimilate what is important. It has been demonstrated that data summarization in natural language has the potential to improve clinical decision making; we have implemented and evaluated a prototype system which generates such textual summaries automatically. Our evaluation of the computer generated summaries showed that the decisions made by medical and nursing staff after reading the summaries were as good as those made after viewing the currently available graphical presentations with the same information content. Since our automatically generated textual summaries can be improved by including additional content and expert knowledge, they promise to enhance information exchange between the medical and nursing staff, particularly when integrated with the currently available graphical presentations. The main feature of this technology is that it brings together a diverse set of techniques such as medical signal analysis, knowledge based reasoning, medical ontology and natural language generation. In this paper we discuss the main components of our approach with a critical analysis of their strengths and limitations and present options for improvement to address these limitations. Jim Hunter, Albert Gatt, François Portet, Ehud Reiter, Somayajulu Sripada |
ECAI | 2 |
| 2008 | REG Challenge Preface
Anya Belz, Albert Gatt |
INLG | 2 |
| 2008 | The GREC Challenge 2008: Overview and Evaluation Results
Anya Belz, Eric Kow, Jette Viethen, Albert Gatt |
INLG | 4 |
| 2008 | Attribute Selection for Referring Expression Generation: New Algorithms and Evaluation Methods
Albert Gatt, Anya Belz |
INLG | 1 |
| 2008 | The TUNA Challenge 2008: Overview and Evaluation Results
Albert Gatt, Anya Belz, Eric Kow |
INLG | 1 |
| 2008 | The Importance of Narrative and Other Lessons from an Evaluation of an NLG System that Summarises Clinical Data
Ehud Reiter, Albert Gatt, François Portet, Marian van der Meulen |
INLG | 2 |
| 2007 | Incremental Generation of Plural Descriptions: Similarity and Partitioning
Albert Gatt, Kees van Deemter |
EMNLP-CoNLL | 1 |
| 2006 | Conceptual Coherence in the Generation of Referring Expressions
Albert Gatt, Kees van Deemter |
ACL | 1 |
| 2006 | Structuring Knowledge for Reference Generation: A Clustering Algorithm
Albert Gatt |
EACL | 1 |
| 2006 | Building a Semantically Transparent Corpus for the Generation of Referring Expressions
Kees van Deemter, Ielka van der Sluis, Albert Gatt |
INLG | 3 |
| 2006 | In pursuit of satisfaction and the prevention of embarrassment: affective state in group recommender systems
Judith Masthoff, Albert Gatt |
User Model. User Adapt. Interact. | 2 |