EDBT 2026 Demo / reviewers in the wild / expert
Ehud Reiter
dblp:21/3710
· DBLP profile ↗
87ranked-venue papers
24as first author
18since 2021 · last 2025
0000-0002-7548-9504ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 78 · 24 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8Human-computer interaction and ubiquitous computing · 6Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scalability of Bayesian Network Structure Elicitation with Large Language Models: a Novel Methodology and Comparative AnalysisabstractIn this work, we propose a novel method for Bayesian Networks (BNs) structure elicitation that is based on the initialization of several LLMs with different experiences, independently querying them to create a structure of the BN, and further obtaining the final structure by majority voting. We compare the method with one alternative method on various widely and not widely known BNs of different sizes and study the scalability of both methods on them. We also propose an approach to check the contamination of BNs in LLM, which shows that some widely known BNs are inapplicable for testing the LLM usage for BNs structure elicitation. We also show that some BNs may be inapplicable for such experiments because their node names are indistinguishable. The experiments on the other BNs show that our method performs better than the existing method with one of the three studied LLMs; however, the performance of both methods significantly decreases with the increase in BN size. Nikolay Babakov, Ehud Reiter, Alberto Bugarín Diz |
COLING | 2 |
| 2025 | When LLMs Can't Help: Real-World Evaluation of LLMs in NutritionabstractThe increasing trust in large language models (LLMs), especially in the form of chatbots, is often undermined by the lack of their extrinsic evaluation. This holds particularly true in nutrition, where randomised controlled trials (RCTs) are the gold standard, and experts demand them for evidence-based deployment. LLMs have shown promising results in this field, but these are limited to intrinsic setups. We address this gap by running the first RCT involving LLMs for nutrition. We augment a rule-based chatbot with two LLM-based features: (1) message rephrasing for conversational variety and engagement, and (2) nutritional counselling through a fine-tuned model. In our seven-week RCT (n=81), we compare chatbot variants with and without LLM integration. We measure effects on dietary outcome, emotional well-being, and engagement. Despite our LLM-based features performing well in intrinsic evaluation, we find that they did not yield consistent benefits in real-world deployment. These results highlight critical gaps between intrinsic evaluations and real-world impact, emphasising the need for interdisciplinary, human-centred approaches. Karen Jia-Hui Li, Simone Balloccu, Ondrej Dusek, Ehud Reiter |
INLG | 4 |
| 2025 | Input Matters: Evaluating Input Structure's Impact on LLM Summaries of Sports Play-by-PlayabstractA major concern when deploying LLMs in accuracy-critical domains such as sports reporting is that the generated text may not faithfully reflect the input data. We quantify how input structure affects hallucinations and other factual errors in LLM-generated summaries of NBA play-by-play data, across three formats: row-structured, JSON and unstructured. We manually annotated 3,312 factual errors across 180 game summaries produced by two models, Llama-3.1-70B and Qwen2.5-72B. Input structure has a strong effect: JSON input reduces error rates by 69% for Llama and 65% for Qwen compared to unstructured input, while row-structured input reduces errors by 54% for Llama and 51% for Qwen. A two-way repeated-measures ANOVA shows that input structure accounts for over 80% of the variance in error rates, with Tukey HSD post hoc tests confirming statistically significant differences between all input formats. Barkavi Sundararajan, Somayajulu Sripada, Ehud Reiter |
INLG | 3 |
| 2025 | Reusability of Bayesian Networks case studies: a surveyabstractAbstract Bayesian Networks (BNs) are probabilistic graphical models used to represent variables and their conditional dependencies, making them highly valuable in a wide range of fields, such as radiology, agriculture, neuroscience, construction management, medicine, and engineering systems, among many others. Despite their widespread application, the reusability of BNs presented in papers that describe their application to real-world tasks has not been thoroughly examined. In this paper, we perform a structured survey on the reusability of BNs using the PRISMA methodology, analyzing 147 papers from various domains. Our results indicate that only 18% of the papers provide sufficient information to enable the reusability of the described BNs. This creates significant challenges for other researchers attempting to reuse these models, especially since many BNs are developed using expert knowledge elicitation. Additionally, direct requests to authors for reusable BNs yielded positive results in only 12% of cases. These findings underscore the importance of improving reusability and reproducibility practices within the BN research community, a need that is equally relevant across the broader field of Artificial Intelligence. Nikolay Babakov, Adarsa Sivaprasad, Ehud Reiter, Alberto Bugarín Diz |
Appl. Intell. | 3 |
| 2025 | The role of natural language processing in improving cancer care: A scoping review with narrative synthesisabstractOBJECTIVES: To review studies of Natural Language Processing (NLP) systems that assist in cancer care, explore use cases and summarise current research progress. METHODS: A scoping review, searching six databases (1) MEDLINE, (2) Embase, (3) IEEE Xplore, (4) ACM Digital Library, (5) Web of Science, and (6) ACL Anthology. Studies were included that reported NLP systems that had been used to improve cancer management by patients or clinicians. Studies were synthesised descriptively and using content analysis. RESULTS: Twenty-nine studies were included. Studies mainly applied NLP in mixed cancer types (n = 10, 34.48 %) and breast cancer (n = 8, 27.59 %). NLP was used in four main ways: (1) to support patient education and self-management; (2) to improve efficiency in clinical care by summarising, extracting, and categorising data, and supporting record-keeping; (3) to support prevention and early detection of patient problems or cancer recurrence; and (4) to improve cancer treatment by supporting clinicians to make evidence-based treatment decisions. Studies highlighted a wide variety of use cases for NLP technologies in cancer care. However, few technologies have been evaluated within clinical settings, none have been evaluated against clinical outcomes, and none have been implemented into clinical care. CONCLUSION: NLP has the potential to improve cancer care via several mechanisms, including information extraction and classification, which could enable automation and personalization of care processes. Additionally, NLP tools such as chatbots show promise in improving patient communication and support. However, there are deficiencies in the evaluation and clinical integration challenges. Interdisciplinary collaboration between computer scientists and clinicians will be essential if NLP technologies are to fulfil their potential to improve patient experience and outcomes. Registered Protocol: https://doi.org/10.17605/OSF.IO/G9DSR. Mengxuan Sun, Ehud Reiter, Lisa Duncan, Rosalind Adam |
Artif. Intell. Medicine | 2 |
| 2024 | Improving Factual Accuracy of Neural Table-to-Text Output by Addressing Input Problems in ToTToabstractBarkavi Sundararajan, Yaji Sripada, Ehud Reiter. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Barkavi Sundararajan, Somayajulu Sripada, Ehud Reiter |
NAACL-HLT | 3 |
| 2024 | Common Flaws in Running Human Evaluation Experiments in NLPabstractAbstract While conducting a coordinated set of repeat runs of human evaluation experiments in NLP, we discovered flaws in every single experiment we selected for inclusion via a systematic process. In this squib, we describe the types of flaws we discovered, which include coding errors (e.g., loading the wrong system outputs to evaluate), failure to follow standard scientific practice (e.g., ad hoc exclusion of participants and responses), and mistakes in reported numerical results (e.g., reported numbers not matching experimental data). If these problems are widespread, it would have worrying implications for the rigor of NLP evaluation experiments as currently conducted. We discuss what researchers can do to reduce the occurrence of such flaws, including pre-registration, better code development practices, increased testing and piloting, and post-publication addressing of errors. Craig Thomson, Ehud Reiter, Anya Belz |
Comput. Linguistics | 2 |
| 2023 | Are Experts Needed? On Human Evaluation of Counselling Reflection GenerationabstractZixiu Wu, Simone Balloccu, Ehud Reiter, Rim Helaoui, Diego Reforgiato Recupero, Daniele Riboni. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zixiu Wu, Simone Balloccu, Ehud Reiter, Rim Helaoui, Diego Reforgiato Recupero, Daniele Riboni |
ACL (1) | 3 |
| 2023 | Enhancing factualness and controllability of Data-to-Text Generation via data Views and constraintsabstractNeural data-to-text systems lack the control and factual accuracy required to generate useful and insightful summaries of multidimensional data.We propose a solution in the form of data views, where each view describes an entity and its attributes along specific dimensions.A sequence of views can then be used as a high-level schema for document planning, with the neural model handling the complexities of micro-planning and surface realization.We show that our view-based system retains factual accuracy while offering high-level control of output that can be tailored based on user preference or other norms within the domain. Craig Thomson, Clément Rebuffel, Ehud Reiter, Laure Soulier, Somayajulu Sripada, Patrick Gallinari |
INLG | 3 |
| 2023 | Influence of context on users' views about explanations for decision-tree predictionsabstractWe consider the influence of two types of contextual information, background information available to users and users’ goals , on users’ views and preferences regarding textual explanations generated for the outcomes predicted by Decision Trees (DTs). To investigate the influence of background information, we generate contrastive explanations that address potential conflicts between aspects of DT predictions and plausible expectations licensed by background information. We define four types of conflicts, operationalize their identification, and specify explanatory schemas that address them. To investigate the influence of users’ goals, we employ an interactive setting where given a goal and an initial explanation for a predicted outcome, users select follow-up questions, and assess the explanations that answer these questions. Here, we offer algorithms to generate explanations that address six types of follow-up questions. The main result from both user studies is that explanations which have a contrastive aspect about a predicted class are generally preferred by users. In addition, the results from the first study indicate that these explanations are deemed especially valuable when users expectations differ from predicted outcomes; and the results from the second study indicate that contrastive explanations which describe how to change a predicted outcome are particularly well regarded in terms of helping users achieve this goal, and they are also popular in terms of helping users achieve other goals. Sameen Maruf, Ingrid Zukerman, Ehud Reiter, Gholamreza Haffari |
Comput. Speech Lang. | 3 |
| 2023 | Evaluating factual accuracy in complex data-to-text
Craig Thomson, Ehud Reiter, Barkavi Sundararajan |
Comput. Speech Lang. | 2 |
| 2022 | Human Evaluation and Correlation with Automatic Metrics in Consultation Note GenerationabstractFrancesco Moramarco, Alex Papadopoulos Korfiatis, Mark Perera, Damir Juric, Jack Flann, Ehud Reiter, Anya Belz, Aleksandar Savkov. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Francesco Moramarco, Alex Papadopoulos-Korfiatis, Mark Perera, Damir Juric, Jack Flann, Ehud Reiter, Anya Belz, Aleksandar Savkov |
ACL (1) | 6 |
| 2022 | Anno-MI: A Dataset of Expert-Annotated Counselling DialoguesabstractResearch on natural language processing for counselling dialogue analysis has seen substantial development in recent years, but access to this area remains extremely limited due to the lack of publicly available expert-annotated therapy conversations. In this work, we introduce AnnoMI, the first publicly and freely accessible dataset of professionally transcribed and expert-annotated therapy dialogues. It consists of 133 conversations that demonstrate high- and low-quality motivational interviewing (MI), an effective counselling technique, and the annotations by domain experts cover key MI attributes. We detail the data collection process including dialogue selection, transcription and annotation. We also present analyses of AnnoMI and discuss its potential applications. Zixiu Wu, Simone Balloccu, Vivek Kumar 0007, Rim Helaoui, Ehud Reiter, Diego Reforgiato Recupero, Daniele Riboni |
ICASSP | 5 |
| 2022 | User-Driven Research of Medical Note Generation SoftwareabstractTom Knoll, Francesco Moramarco, Alex Papadopoulos Korfiatis, Rachel Young, Claudia Ruffini, Mark Perera, Christian Perstl, Ehud Reiter, Anya Belz, Aleksandar Savkov. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Tom Knoll, Francesco Moramarco, Alex Papadopoulos-Korfiatis, Rachel Young, Claudia Ruffini, Mark Perera, Christian Perstl, Ehud Reiter, Anya Belz, Aleksandar Savkov |
NAACL-HLT | 8 |
| 2021 | A Systematic Review of Reproducibility Research in Natural Language ProcessingabstractAgainst the background of what has been termed a reproducibility crisis in science, the NLP field is becoming increasingly interested in, and conscientious about, the reproducibility of its results.The past few years have seen an impressive range of new initiatives, events and active research in the area.However, the field is far from reaching a consensus about how reproducibility should be defined, measured and addressed, with diversity of views currently increasing rather than converging.With this focused contribution, we aim to provide a wideangle, and as near as possible complete, snapshot of current work on reproducibility in NLP, delineating differences and similarities, and providing pointers to common denominators. Anya Belz, Anastasia Shimorina, Ehud Reiter |
EACL | 4 |
| 2021 | The ReproGen Shared Task on Reproducibility of Human Evaluations in NLG: Overview and ResultsabstractThe NLP field has recently seen a substantial increase in work related to reproducibility of results, and more generally in recognition of the importance of having shared definitions and practices relating to evaluation.Much of the work on reproducibility has so far focused on metric scores, with reproducibility of human evaluation results receiving far less attention.As part of a research programme designed to develop theory and practice of reproducibility assessment in NLP, we organised the first shared task on reproducibility of human evaluations, ReproGen 2021.This paper describes the shared task in detail, summarises results from each of the reproduction studies submitted, and provides further comparative analysis of the results.Out of nine initial team registrations, we received submissions from four teams.Meta-analysis of the four reproduction studies revealed varying degrees of reproducibility, and allowed very tentative first conclusions about what types of evaluation tend to have better reproducibility. Anya Belz, Anastasia Shimorina, Ehud Reiter |
INLG | 4 |
| 2021 | Explaining Decision-Tree Predictions by Addressing Potential Conflicts between Predictions and Plausible ExpectationsabstractWe offer an approach to explain Decision Tree (DT) predictions by addressing potential conflicts between aspects of these predictions and plausible expectations licensed by background information.We define four types of conflicts, operationalize their identification, and specify explanatory schemas that address them.Our human evaluation focused on the effect of explanations on users' understanding of a DT's reasoning and their willingness to act on its predictions.The results show that (1) explanations that address potential conflicts are considered at least as good as baseline explanations that just follow a DT path; and (2) the conflictbased explanations are deemed especially valuable when users' expectations disagree with the DT's predictions. Sameen Maruf, Ingrid Zukerman, Ehud Reiter, Gholamreza Haffari |
INLG | 3 |
| 2021 | Generation Challenges: Results of the Accuracy Evaluation Shared TaskabstractThe Shared Task on Evaluating Accuracy focused on techniques (both manual and automatic) for evaluating the factual accuracy of texts produced by neural NLG systems, in a sports-reporting domain.Four teams submitted evaluation techniques for this task, using very different approaches and techniques.The best-performing submissions did encouragingly well at this difficult task.However, all automatic submissions struggled to detect factual errors which are semantically or pragmatically complex (for example, based on incorrect computation or inference). Craig Thomson, Ehud Reiter |
INLG | 2 |
| 2020 | Arabic NLG Language FunctionsabstractThe Arabic language has very limited supports from NLG researchers.In this paper, we explain the challenges of the core grammar, provide a lexical resource, and implement the first language functions for the Arabic language.We did a human evaluation to evaluate our functions in generating sentences from the NADA Corpus. Wael Abed, Ehud Reiter |
INLG | 2 |
| 2020 | ReproGen: Proposal for a Shared Task on Reproducibility of Human Evaluations in NLGabstractAcross NLP, a growing body of work is looking at the issue of reproducibility.However, replicability of human evaluation experiments and reproducibility of their results is currently under-addressed, and this is of particular concern for NLG where human evaluations are the norm.This paper outlines our ideas for a shared task on reproducibility of human evaluations in NLG which aims (i) to shed light on the extent to which past NLG evaluations have been replicable and reproducible, and (ii) to draw conclusions regarding how evaluations can be designed and reported to increase replicability and reproducibility.If the task is run over several years, we hope to be able to document an overall increase in levels of replicability and reproducibility over time. Anya Belz, Anastasia Shimorina, Ehud Reiter |
INLG | 4 |
| 2020 | Shared Task on Evaluating AccuracyabstractWe propose a shared task on methodologies and algorithms for evaluating the accuracy of generated texts, specifically summaries of basketball games produced from basketball box score and other game data.We welcome submissions based on protocols for human evaluation, automatic metrics, as well as combinations of human evaluations and metrics. Ehud Reiter, Craig Thomson |
INLG | 1 |
| 2020 | A Gold Standard Methodology for Evaluating Accuracy in Data-To-Text SystemsabstractMost Natural Language Generation systems need to produce accurate texts.We propose a methodology for high-quality human evaluation of the accuracy of generated texts, which is intended to serve as a gold-standard for accuracy evaluations of data-to-text systems.We use our methodology to evaluate the accuracy of computer generated basketball summaries.We then show how our gold standard evaluation can be used to validate automated metrics. Craig Thomson, Ehud Reiter |
INLG | 2 |
| 2018 | Generating Summaries of Sets of Consumer Products: Learning from ExperimentsabstractWe explored the task of creating a textual summary describing a large set of objects characterised by a small number of features using an e-commerce dataset.When a set of consumer products is large and varied, it can be difficult for a consumer to understand how the products in the set differ; consequently, it can be challenging to choose the most suitable product from the set.To assist consumers, we generated high-level summaries of product sets.Two generation algorithms are presented, discussed, and evaluated with human users.Our evaluation results suggest a positive contribution to consumers' understanding of the domain. Kittipitch Kuptavanich, Ehud Reiter, Kees van Deemter, Advaith Siddharthan |
INLG | 2 |
| 2018 | Meteorologists and Students: A resource for language grounding of geographical descriptorsabstractWe present a data resource which can be useful for research purposes on language grounding tasks in the context of geographical referring expression generation.The resource is composed of two data sets that encompass 25 different geographical descriptors and a set of associated graphical representations, drawn as polygons on a map by two groups of human subjects: teenage students and expert meteorologists. Alejandro Ramos-Soto, Ehud Reiter, Kees van Deemter, Jose Maria Alonso-Moral, Albert Gatt |
INLG | 2 |
| 2018 | Comprehension Driven Document Planning in Natural Language Generation SystemsabstractThis paper proposes an approach to NLG system design which focuses on generating output text which can be more easily processed by the reader.Ways in which cognitive theory might be combined with existing NLG techniques are discussed and two simple experiments in content ordering are presented. Craig Thomson, Ehud Reiter, Somayajulu Sripada |
INLG | 2 |
| 2018 | A Structured Review of the Validity of BLEUabstractThe BLEU metric has been widely used in NLP for over 15 years to evaluate NLP systems, especially in machine translation and natural language generation. I present a structured review of the evidence on whether BLEU is a valid evaluation technique—in other words, whether BLEU scores correlate with real-world utility and user-satisfaction of NLP systems; this review covers 284 correlations reported in 34 papers. Overall, the evidence supports using BLEU for diagnostic evaluation of MT systems (which is what it was originally proposed for), but does not support using BLEU outside of MT, for evaluation of individual texts, or for scientific hypothesis testing. Ehud Reiter |
Comput. Linguistics | 1 |
| 2018 | SaferDrive: An NLG-based behaviour change support system for driversabstractAbstract Despite the long history of Natural Language Generation (NLG) research, the potential for influencing real world behaviour through automatically generated texts has not received much attention. In this paper, we presentSaferDrive, a behaviour change support system that uses NLG and telematic data in order to create weekly textual feedback for automobile drivers, which is delivered through a smartphone application. Usage-based car insurances use sensors to track driver behaviour. Although the data collected by such insurances could provide detailed feedback about the driving style, they are typically withheld from the driver and used only to calculate insurance premiums.SaferDriveinstead provides detailed textual feedback about the driving style, with the intent to help drivers improve their driving habits. We evaluate the system with real drivers and report that the textual feedback generated by our system does have a positive influence on driving habits, especially with regard to speeding. Daniel Braun 0003, Ehud Reiter, Advaith Siddharthan |
Nat. Lang. Eng. | 2 |
| 2017 | An exploratory study on the benefits of using natural language for explaining fuzzy rule-based systemsabstractThis paper presents an empirical research. It focuses on testing empirically the benefits of providing users, in a specific domain, with textual interpretation of the fuzzy inferences carried out by a fuzzy classifier for a given selection of samples. The hypothesis to test is as follows: “Users understand easier the decision made by a fuzzy system when they are provided with a textual interpretation of the fuzzy inference mechanism that the system carried out”. This hypothesis was successfully tested in a web survey. The application domain was leaf classification. The fuzzy classifiers were built with the GUAJE fuzzy modeling open source software which is aimed at generating interpretable fuzzy systems. The textual interpretation was handmade by an expert who followed the guidelines of the Natural Language Generation approach proposed by Reiter and Dale. Reported results encourage us to go on with a series of additional experiments devoted to deeply explore how Natural Language Generation techniques can contribute to facilitate the understanding of fuzzy systems. Jose Maria Alonso-Moral, Alejandro Ramos-Soto, Ehud Reiter, Kees van Deemter |
FUZZ-IEEE | 3 |
| 2017 | An empirical approach for modeling fuzzy geographical descriptorsabstractWe present a novel heuristic approach that defines fuzzy geographical descriptors using data gathered from a survey with human subjects. The participants were asked to provide graphical interpretations of the descriptors `north' and `south' for the Galician region (Spain). Based on these interpretations, our approach builds fuzzy descriptors that are able to compute membership degrees for geographical locations. We evaluated our approach in terms of efficiency and precision. The fuzzy descriptors are meant to be used as the cornerstones of a geographical referring expression generation algorithm that is able to linguistically characterize geographical locations and regions. This work is also part of a general research effort that intends to establish a methodology which reunites the empirical studies traditionally practiced in data-to-text and the use of fuzzy sets to model imprecision and vagueness in words and expressions for text generation purposes. Alejandro Ramos-Soto, Jose Maria Alonso-Moral, Ehud Reiter, Kees van Deemter, Albert Gatt |
FUZZ-IEEE | 3 |
| 2017 | Textually Summarising Incomplete DataabstractMany data-to-text NLG systems work with data sets which are incomplete, ie some of the data is missing.We have worked with data journalists to understand how they describe incomplete data, and are building NLG algorithms based on these insights.A pilot evaluation showed mixed results, and highlighted several areas where we need to improve our system. Stephanie Inglis, Ehud Reiter, Somayajulu Sripada |
INLG | 2 |
| 2017 | A Commercial Perspective on ReferenceabstractI briefly describe some of the commercial work which Arria NLG is doing in referring expression algorithms, and highlight differences between what is commercially important (at least to Arria) and the NLG research literature.Arria's focus is on high-quality algorithms for types of reference which are important in its systems.These algorithms need to be parametrisable for different genres and domains, usable in hybrid systems which include some canned text, and support variation. Ehud Reiter |
INLG | 1 |
| 2016 | Natural language generation and fuzzy sets: An exploratory study on geographical referring expression generationabstractWe explore how the problem of uncertainty and imprecision in natural language generation (NLG) could be addressed through the use of fuzzy sets. We propose bringing together standard empirical procedures for knowledge acquisition in NLG and computing with words/perceptions related techniques (with a special focus on linguistic description of data) to address an open challenge in NLG: the generation of geographical referring expressions. Following this methodology, we present an exploratory experiment which provides some insights about how human subjects refer to geographical expressions and discuss how the obtained results might relate to the use of fuzzy sets. Alejandro Ramos-Soto, Nava Tintarev, Rodrigo de Oliveira, Ehud Reiter, Kees van Deemter |
FUZZ-IEEE | 4 |
| 2016 | Absolute and Relative Properties in Geographic Referring ExpressionsabstractThis paper discusses the importance of computing relative properties and not just retrieving absolute properties when generating geographic referring expressions such as "northern France".We describe an algorithm that computes spatial properties at run-time by means of spatial operations such as intersecting and analyzing parts of wholes.The evaluation of the algorithm suggests that part-whole relations are key in geographic expressions. Rodrigo de Oliveira, Somayajulu Sripada, Ehud Reiter |
INLG | 3 |
| 2016 | Personal storytelling: Using Natural Language Generation for children with complex communication needs, in the wild
Nava Tintarev, Ehud Reiter, Rolf Black, Annalu Waller, Joseph Reddington |
Int. J. Hum. Comput. Stud. | 2 |
| 2014 | Generating Annotated Graphs using the NLG Pipeline ArchitectureabstractThe Arria NLG Engine has been extended to generate annotated graphs: data graphs that contain computer-generated textual annotations to explain phenomena in those graphs. These graphs are generated along-side text-only data summaries. 1 Saad Mahamood, William Bradshaw, Ehud Reiter |
INLG | 3 |
| 2014 | Providing Adaptive Health Updates Across the Personal Social NetworkabstractThis article presents research conducted to establish how information is shared across the personal social network in the sensitive context of a health crisis. We worked with parents of very sick babies who were cared for in a hospital's Neonatal Unit (NNU). Through a combination of interviews, a focus group, and surveys, we developed a user model of the information that parents wanted to share, and how they adapted this information to individual recipients. We then developed a prototype software tool which created adaptive updates for members of the parents' social network. The updates contained summaries of large volumes of complex medical data about the baby, nonmedical information about the parents, and practical information about the hospital. Updates were automatically adapted to individual members of parents' social networks, based on our user model. The tool was evaluated in a large NNU in the United Kingdom with parents of babies who were currently being cared for in the unit. We found that parents adapted the information that they shared about themselves and their babies based on the emotional proximity of their network members. They gave most detail to those who were emotionally closest to them and least to those who were less close. Parents also adapted information content to the recipient's tendency to worry and empathize. Two adaptive strategies were deployed by parents, (a) benign deceit—not telling the whole truth—and (b) promotion of empathetic members of the social network to a higher level of emotional proximity, so that they were given more information. We generated a number of directions for future work, and issues to consider around designing adaptive mediated communications systems for sensitive contexts. These include the potential to generalize our model to other medical contexts and considerations to apply when deliberately designing deceit into adaptive systems. Wendy Moncur, Judith Masthoff, Ehud Reiter, Yvonne Freer, Hien Nguyen 0002 |
Hum. Comput. Interact. | 3 |
| 2013 | Typicality and Object Reference
Margaret Mitchell, Ehud Reiter, Kees van Deemter |
CogSci | 2 |
| 2013 | Generating Expressions that Refer to Visible Objects
Margaret Mitchell, Kees van Deemter, Ehud Reiter |
HLT-NAACL | 3 |
| 2012 | Working with Clinicians to Improve a Patient-Information NLG System
Saad Mahamood, Ehud Reiter |
INLG | 2 |
| 2012 | Automatic generation of natural language nursing shift summaries in neonatal intensive care: BT-Nurse
Jim Hunter, Yvonne Freer, Albert Gatt, Ehud Reiter, Somayajulu Sripada, Cindy Sykes |
Artif. Intell. Medicine | 4 |
| 2012 | Supporting Personal Narrative for Children with Complex Communication NeedsabstractChildren with complex communication needs who use voice output communication aids seldom engage in extended conversation. The “How was School today...?” system has been designed to enable such children to talk about their school day. The system uses data-to-text technology to generate narratives from sensor data. Observations, interviews and prototyping were used to ensure that stakeholders were involved in the design of the system. Evaluations with three children showed that the prototype system, which automatically generates utterances, has the potential to support disabled individuals to participate better in interactive conversation. Analysis of a conversational transcript and observations indicate that the children were able to access relevant conversation and had more control in the conversation in comparison to their usual interactions where control lay mainly with the speaking partner. Further research to develop an improved, more rugged system that supports users with different levels of language ability is now underway. Rolf Black, Annalu Waller, Ross Turner, Ehud Reiter |
ACM Trans. Comput. Hum. Interact. | 4 |
| 2011 | A mobile phone based personal narrative systemabstractCurrently available commercial Augmentative and Alternative Communication (AAC) technology makes little use of computing power to improve the access to words and phrases for personal narrative, an essential part of social interaction. In this paper, we describe the development and evaluation of a mobile phone application to enable data collection for a personal narrative system for children with severe speech and physical impairments (SSPI). Based on user feedback from the previous project "How was School today?" we developed a modular system where school staff can use a mobile phone to track interaction with people and objects and user location at school. The phone also allows taking digital photographs and recording voice message sets by both school staff and parents/carers at home. These sets can be played back by the child for immediate narrative sharing similar to established AAC device interaction using sequential voice recorders. The mobile phone sends all the gathered data to a remote server. The data can then be used for automatic narrative generation on the child's PC based communication aid. Early results from the ongoing evaluation of the application in a special school with two participants and school staff show that staff were able to track interactions, record voice messages and take photographs. Location tracking was less successful, but was supplemented by timetable information. The participating children were able to play back voice messages and show photographs on the mobile phone for interactive narrative sharing using both direct and switch activated playback options. Rolf Black, Annalu Waller, Nava Tintarev, Ehud Reiter, Joseph Reddington |
ASSETS | 4 |
| 2011 | On the Use of Size Modifiers When Referring to Visible Objects
Margaret Mitchell, Kees van Deemter, Ehud Reiter |
CogSci | 3 |
| 2011 | BT-Nurse: computer generation of natural language shift summaries from complex heterogeneous medical dataabstractThe BT-Nurse system uses data-to-text technology to automatically generate a natural language nursing shift summary in a neonatal intensive care unit (NICU). The summary is solely based on data held in an electronic patient record system, no additional data-entry is required. BT-Nurse was tested for two months in the Royal Infirmary of Edinburgh NICU. Nurses were asked to rate the understandability, accuracy, and helpfulness of the computer-generated summaries; they were also asked for free-text comments about the summaries. The nurses found the majority of the summaries to be understandable, accurate, and helpful (p<0.001 for all measures). However, nurses also pointed out many deficiencies, especially with regard to extra content they wanted to see in the computer-generated summaries. In conclusion, natural language NICU shift summaries can be automatically generated from an electronic patient record, but our proof-of-concept software needs considerable additional development work before it can be deployed. Jim Hunter, Yvonne Freer, Albert Gatt, Ehud Reiter, Somayajulu Sripada, Cindy Sykes, Dave Westwater |
J. Am. Medical Informatics Assoc. | 4 |
| 2010 | Natural Reference to Objects in a Visual Domain
Margaret Mitchell, Kees van Deemter, Ehud Reiter |
INLG | 3 |
| 2010 | Modeling the socially intelligent communication of health information to a patient's personal social networkabstractThis study examined how emotional proximity and gender affect people's information requirements when someone that they know is chronically or critically ill. In an online study, participants were asked what information they would want to receive about members of their social network in three categories: someone who was very close, someone who was not so close, and someone who was not close at all. Our results show that the information that people want can be predicted from their gender and emotional proximity to the network member. The closer the relationship with the patient, the more information people want. Women want more information than men. We propose a model for the socially intelligent communication of health information across the social network, and discuss areas for its application. Wendy Moncur, Ehud Reiter, Judith Masthoff, Alex Carmichael |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2009 | Automatic generation of textual summaries from neonatal intensive care data
François Portet, Ehud Reiter, Albert Gatt, Jim Hunter, Somayajulu Sripada, Yvonne Freer, Cindy Sykes |
Artif. Intell. | 2 |
| 2009 | An Investigation into the Validity of Some Metrics for Automatically Evaluating Natural Language Generation SystemsabstractThere is growing interest in using automatically computed corpus-based evaluation metrics to evaluate Natural Language Generation (NLG) systems, because these are often considerably cheaper than the human-based evaluations which have traditionally been used in NLG. We review previous work on NLG evaluation and on validation of automatic metrics in NLP, and then present the results of two studies of how well some metrics which are popular in other areas of NLP (notably BLEU and ROUGE) correlate with human judgments in the domain of computer-generated weather forecasts. Our results suggest that, at least in this domain, metrics may provide a useful measure of language quality, although the evidence for this is not as strong as we would ideally like to see; however, they do not provide a useful measure of content quality. We also discuss a number of caveats which must be kept in mind when interpreting this and other validation studies. Ehud Reiter, Anya Belz |
Comput. Linguistics | 1 |
| 2008 | Summarising Complex ICU Data in Natural Language
Jim Hunter, Yvonne Freer, Albert Gatt, Robert H. Logie, Neil McIntosh, Marian van der Meulen, François Portet, Ehud Reiter, Somayajulu Sripada, Cindy Sykes |
AMIA | 8 |
| 2008 | Neonatal Intensive Care Information for Parents - An Affective ApproachabstractBased upon qualitative work done with former Neonatal Intensive Care Unit parents, we propose a potential user model to estimate the level of stress/anxiety that a parent is experiencing and how information given to such parents should be adjusted to meet their informational and emotional needs. Saad Mahamood, Ehud Reiter, Chris Mellish |
CBMS | 2 |
| 2008 | What Do You Want to Know? Investigating the Information Requirements of Patient SupportersabstractThere is a vast amount of data associated with any one patient. It is challenging for medical staff to understand all this data. It is even harder for a lay person, who may not even know what medical terms mean. The research project BabyTalk-Clan aims to create personalized summaries of data for a lay audience. It uses sensitive, highly-detailed clinical data relating to a patient. This includes medication given, test results, notes made by medical staff, and continuous physiological signals such as heart rate. We took a qualitative approach to knowledge acquisition for user requirements. Using interviews and a focus group within a Grounded Theory methodology, we discovered that most lay users want only a very high-level summary of the baby's state. What lay users do want is information about how the parents are coping, and what support they need. Findings were cross-validated through a questionnaire. Wendy Moncur, Judith Masthoff, Ehud Reiter |
CBMS | 3 |
| 2008 | Using Natural Language Generation Technology to Improve Information Flows in Intensive Care UnitsabstractIn the drive to improve patient safety, patients in modern intensive care units are closely monitored with the generation of very large volumes of data. Unless the data are further processed, it is difficult for medical and nursing staff to assimilate what is important. It has been demonstrated that data summarization in natural language has the potential to improve clinical decision making; we have implemented and evaluated a prototype system which generates such textual summaries automatically. Our evaluation of the computer generated summaries showed that the decisions made by medical and nursing staff after reading the summaries were as good as those made after viewing the currently available graphical presentations with the same information content. Since our automatically generated textual summaries can be improved by including additional content and expert knowledge, they promise to enhance information exchange between the medical and nursing staff, particularly when integrated with the currently available graphical presentations. The main feature of this technology is that it brings together a diverse set of techniques such as medical signal analysis, knowledge based reasoning, medical ontology and natural language generation. In this paper we discuss the main components of our approach with a critical analysis of their strengths and limitations and present options for improvement to address these limitations. Jim Hunter, Albert Gatt, François Portet, Ehud Reiter, Somayajulu Sripada |
ECAI | 4 |
| 2008 | The Importance of Narrative and Other Lessons from an Evaluation of an NLG System that Summarises Clinical Data
Ehud Reiter, Albert Gatt, François Portet, Marian van der Meulen |
INLG | 1 |
| 2008 | Using Spatial Reference Frames to Generate Grounded Textual Summaries of Georeferenced Data
Ross Turner, Somayajulu Sripada, Ehud Reiter, Ian Davy |
INLG | 3 |
| 2008 | Generating basic skills reports for low-skilled readersabstractAbstract We describe SkillSum, a Natural Language Generation (NLG) system that generates a personalised feedback report for someone who has just completed a screening assessment of their basic literacy and numeracy skills. Because many SkillSum users have limited literacy, the generated reports must be easily comprehended by people with limited reading skills; this is the most novel aspect of SkillSum, and the focus of this paper. We used two approaches to maximise readability. First, for determining content and structure (document planning), we did not explicitly model readability, but rather followed a pragmatic approach of repeatedly revising content and structure following pilot experiments and interviews with domain experts. Second, for choosing linguistic expressions (microplanning), we attempted to formulate explicitly the choices that enhanced readability, using a constraints approach and preference rules; our constraints were based on corpus analysis and our preference rules were based on psycholinguistic findings. Evaluation of the SkillSum system was twofold: it compared the usefulness of NLG technology to that of canned text output, and it assessed the effectiveness of the readability model. Results showed that NLG was more effective than canned text at enhancing users' knowledge of their skills, and also suggested that the empirical ‘revise based on experiments and interviews’ approach made a substantial contribution to readability as well as our explicit psycholinguistically inspired models of readability choices. Sandra Williams, Ehud Reiter |
Nat. Lang. Eng. | 2 |
| 2007 | Automatic Generation of Textual Summaries from Neonatal Intensive Care Data
François Portet, Ehud Reiter, Jim Hunter, Somayajulu Sripada |
AIME | 2 |
| 2007 | The Shrinking Horizons of Computational LinguisticsabstractThe ProblemUnderstanding language is one of the great challenges of science, and languagerelated technology is one of the great opportunities of Information Technology.Consequently, many different kinds of researchers work on language issues.Within the computer science community, language is studied by the "ACL community," by which I mean researchers who regularly publish in Association for Computational Linguistics (ACL) venues, such as the journal Computational Linguistics and ACL conferences.But language-related research is also carried out by researchers in other areas of computer science, including knowledge representation, cognitive modeling, vision and robotics, and human-computer interaction communities.Additionally, there are even more people outside computer science who study language, including linguists, psycholinguists, philosophers, and sociolinguists.This is fine; understanding language and developing language technology are huge problems, and it is very useful to have many research communities from diverse backgrounds working on language.This will be especially true if the different research communities are aware of each other, so they can share insights, observations, problems, and so forth.Unfortunately, my impression is that the ACL community is much less interested in research with other language-related research communities than it used to be.This impression is mostly based on discussions I have had with researchers who are on the border between ACL and another language-research community.Several such people have told me that whereas ten years ago they occasionally submitted papers to ACL venues and attended ACL conferences, now they do not bother, because they believe that the ACL community has no interest in their research.In attempt to quantify this insight, I have analyzed citations from papers published in Computational Linguistics in 1995 and in 2005.Specifically, I extracted all citations from Computational Linguistics (CL) articles (excluding book reviews) in these years to journal papers.I then classified the cited journal papers into one of the categories shown in Table 1; whenever possible this classification was based on the subject category assigned by ISI Journal Citation Reports (JCR) to the cited journal.For example, a citation of a paper in Cognitive Science would count as a psychology citation, since ISI JCR classifies Cognitive Science as "Psychology, Experimental."I counted citations myself, rather than relying on ISI JCR's count, as there were some mistakes in JCR's counting.I also created my own "other NLP and speech" classification (that is, references to speech and NLP Ehud Reiter |
Comput. Linguistics | 1 |
| 2007 | Choosing the content of textual summaries of large time-series data setsabstractNatural Language Generation (NLG) can be used to generate textual summaries of numeric data sets. In this paper we develop an architecture for generating short (a few sentences) summaries of large (100KB or more) time-series data sets. The architecture integrates pattern recognition, pattern abstraction, selection of the most significant patterns, microplanning (especially aggregation), and realisation. We also describe and evaluate SumTime-Turbine, a prototype system which uses this architecture to generate textualsummaries of sensor data from gas turbines. Ehud Reiter, Jim Hunter, Chris Mellish |
Nat. Lang. Eng. | 2 |
| 2006 | Comparing Automatic and Human Evaluation of NLG Systems
Anya Belz, Ehud Reiter |
EACL | 2 |
| 2006 | Generating Spatio-Temporal Descriptions in Pollen Forecasts
Ross Turner, Somayajulu Sripada, Ehud Reiter, Ian P. Davy |
EACL | 3 |
| 2006 | GENEVAL: A Proposal for Shared-task Evaluation in NLG
Ehud Reiter, Anya Belz |
INLG | 1 |
| 2005 | Evaluating an NLG System using Post-Editing
Somayajulu Sripada, Ehud Reiter, Lezan Hawizy |
IJCAI | 2 |
| 2005 | Appropriate Microplanning Choices for Low-Skilled Readers
Sandra Williams, Ehud Reiter |
IJCAI | 2 |
| 2005 | Choosing words in computer-generated weather forecasts
Ehud Reiter, Somayajulu Sripada, Jim Hunter, Ian Davy |
Artif. Intell. | 1 |
| 2005 | Connecting language to the world
Deb Roy, Ehud Reiter |
Artif. Intell. | 2 |
| 2004 | Lessons from Deploying NLG Technology for Marine Weather Forecast Text Generation
Somayajulu Sripada, Ehud Reiter, Ian Davy, Kristian Nilssen |
ECAI | 2 |
| 2004 | Contextual Influences on Near-Synonym Choice
Ehud Reiter, Somayajulu Sripada |
INLG | 1 |
| 2003 | Summarizing Neonatal Time Series Data
Somayajulu Sripada, Ehud Reiter, Jim Hunter |
EACL | 2 |
| 2003 | SumTime-Turbine: A Knowledge-Based System to Communicate Gas Turbine Time-Series Data
Ehud Reiter, Jim Hunter, Somayajulu Sripada |
IEA/AIE | 2 |
| 2003 | Generating English summaries of time series data using the Gricean maximsabstractWe are developing technology for generating English textual summaries of time-series data, in three domains: weather forecasts, gas-turbine sensor readings, and hospital intensive care data. Our weather-forecast generator is currently operational and being used daily by a meteorological company. We generate summaries in three steps: (a) selecting the most important trends and patterns to communicate; (b) mapping these patterns onto words and phrases; and (c) generating actual texts based on these words and phrases. In this paper we focus on the first step, (a), selecting the information to communicate, and describe how we perform this using modified versions of standard data analysis algorithms such as segmentation. The modifications arose out of empirical work with users and domain experts, and in fact can all be regarded as applications of the Gricean maxims of Quality, Quantity, Relevance, and Manner, which describe how a cooperative speaker should behave in order to help a hearer correctly interpret a text. The Gricean maxims are perhaps a key element of adapting data analysis algorithms for effective communication of information to human users, and should be considered by other researchers interested in communicating data to human users. Somayajulu Sripada, Ehud Reiter, Jim Hunter |
KDD | 2 |
| 2003 | Lessons from a failure: Generating tailored smoking cessation letters
Ehud Reiter, Roma Robertson, Liesl Osman |
Artif. Intell. | 1 |
| 2003 | Acquiring Correct Knowledge for Natural Language GenerationabstractNatural language generation (NLG) systems are computer software systems that produce texts in English and other human languages, often from non-linguistic input data. NLG systems, like most AI systems, need substantial amounts of knowledge. However, our experience in two NLG projects suggests that it is difficult to acquire correct knowledge for NLG systems; indeed, every knowledge acquisition (KA) technique we tried had significant problems. In general terms, these problems were due to the complexity, novelty, and poorly understood nature of the tasks our systems attempted, and were worsened by the fact that people write so differently. This meant in particular that corpus-based KA approaches suffered because it was impossible to assemble a sizable corpus of high-quality consistent manually written texts in our domains; and structured expert-oriented KA techniques suffered because experts disagreed and because we could not get enough information about special and unusual cases to build robust systems. We believe that such problems are likely to affect many other NLG systems as well. In the long term, we hope that new KA techniques may emerge to help NLG system builders. In the shorter term, we believe that understanding how individual KA techniques can fail, and using a mixture of different KA techniques with different strengths and weaknesses, can help developers acquire NLG knowledge that is mostly correct. Ehud Reiter, Somayajulu Sripada, Roma Robertson |
J. Artif. Intell. Res. | 1 |
| 2002 | Should Corpora Texts Be Gold Standards for NLG?
Ehud Reiter, Somayajulu Sripada |
INLG | 1 |
| 2002 | The Spoken Language Translator - M. Rayner, D. Carter, P. Bouillon, V. Digalakis, M. Wirn (Eds.), Studies in Natural Language Processing Series, Cambridge University Press, Cambridge, UK, 2000, ISBN 0521770777
Ehud Reiter |
Artif. Intell. Medicine | 1 |
| 2002 | Human Variation and Lexical ChoiceabstractMuch natural language processing research implicitly assumes that word meanings are fixed in a language community, but in fact there is good evidence that different people probably associate slightly different meanings with words. We summarize some evidence for this claim from the literature and from an ongoing research project, and discuss its implications for natural language generation, especially for lexical choice, that is, choosing appropriate words for a generated text. Ehud Reiter, Somayajulu Sripada |
Comput. Linguistics | 1 |
| 2001 | Using a Randomised Controlled Clinical Trial to Evaluate an NLG SystemabstractThe STOP system, which generates personalised smoking-cessation letters, was evaluated by a randomised controlled clinical trial. We believe this is the largest and perhaps most rigorous task effectiveness evaluation ever performed on an NLG system. The detailed results of the clinical trial have been presented elsewhere, in the medical literature. In this paper we discuss the clinical trial itself: its structure and cost, what we did and did not learn from it (especially considering that the trial showed that STOP was not effective), and how it compares to other NLG evaluation techniques. Ehud Reiter, Roma Robertson, A. Scott Lennox, Liesl Osman |
ACL | 1 |
| 2000 | Knowledge Acquisition for Natural Language GenerationabstractWe describe the knowledge acquisition (KA) techniques used to build the STOP system, especially sorting and think-aloud protocols. That is, we describe the ways in which we interacted with domain experts to determine appropriate user categories, schemas, detailed content rules, and so forth for STOP. Informal evaluations of these techniques suggest that they had some benefit, but perhaps were most successful as a source of insight and hypotheses, and should ideally have been supplemented by other techniques when deciding on the specific rules and knowledge incorporated into STOP. Ehud Reiter, Roma Robertson, Liesl Osman |
INLG | 1 |
| 2000 | Pipelines and Size ConstraintsabstractSome types of documents need to meet size constraints, such as fitting into a limited number of pages. This can be a difficult constraint to enforce in a pipelined natural language generation (NLG) system, because size is mostly determined by content decisions, which usually are made at the beginning of the pipeline, but size cannot be accurately measured until the document has been completely processed by the NLG system. I present experimental data on the performance of single-solution pipeline, multiple-solution pipeline, and revision-based variants of the STOP system (which produces personalized smoking-cessation leaflets) in meeting a size constraint. This shows that a multiple-solution pipeline does much better than a single-solution pipeline, and that a revision-based system does best of all. Ehud Reiter |
Comput. Linguistics | 1 |
| 1997 | Building applied natural language generation systemsabstractIn this article, we give an overview of Natural Language Generation (NLG) from an applied system-building perspective. The article includes a discussion of when NLG techniques should be used; suggestions for carrying out requirements analyses; and a description of the basic NLG tasks of content determination, discourse planning, sentence aggregation, lexicalization, referring expression generation, and linguistic realisation. Throughout, the emphasis is on established techniques that can be used to build simple but practical working systems now. We also provide pointers to techniques in the literature that are appropriate for more complicated scenarios. Ehud Reiter, Robert Dale |
Nat. Lang. Eng. | 1 |
| 1994 | Has a Consensus NL Generation Architecture Appeared, and is it Psycholinguistically Plausible?
Ehud Reiter |
INLG | 1 |
| 1993 | Using Classification as a Programming Language
Chris Mellish, Ehud Reiter |
IJCAI | 2 |
| 1993 | Optimizing the Costs and Benefits of Natural Language Generation
Ehud Reiter, Chris Mellish |
IJCAI | 1 |
| 1992 | Using Classification to Generate TextabstractThe IDAS natural-language generation system uses a KL-ONE type classifier to perform content determination, surface realisation, and part of text planning. Generation-by-classification allows IDAS to use a single representation and reasoning component for both domain and linguistic knowledge, which is difficult for systems based on unification or systemic generation techniques. Ehud Reiter, Chris Mellish |
ACL | 1 |
| 1992 | A Fast Algorithm for the Generation of Referring Expressions
Ehud Reiter, Robert Dale |
COLING | 1 |
| 1990 | Avoiding Unwanted Conversational Implicatures in Text and Graphics
Joseph Marks, Ehud Reiter |
AAAI | 2 |
| 1990 | The Computational Complexity of Avoiding Conversational ImplicaturesabstractReferring expressions and other object descriptions should be maximal under the Local Brevity, No Unnecessary Components, and Lexical Preference preference rules; otherwise, they may lead hearers to infer unwanted conversational implicatures. These preference rules can be incorporated into a polynomial time generation algorithm, while some alternative formalizations of conversational implicature make the generation task NP-Hard. Ehud Reiter |
ACL | 1 |
| 1990 | A New Model for Lexical Choice for Open-Class Words
Ehud Reiter |
INLG | 1 |