EDBT 2026 Demo / reviewers in the wild / expert
Anya Belz
dblp:212/6084 · also Anja Belz
· DBLP profile ↗
50ranked-venue papers
26as first author
16since 2021 · last 2025
0000-0002-0552-8096ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 26 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Query-driven Document-level Scientific Evidence Extraction from Biomedical StudiesabstractMassimiliano Pronesti, Joao H Bettencourt-Silva, Paul Flanagan, Alessandra Pascale, Oisín Redmond, Anya Belz, Yufang Hou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Massimiliano Pronesti, Joao H. Bettencourt-Silva, Paul Flanagan, Alessandra Pascale, Oisin Redmond, Anya Belz, Yufang Hou 0001 |
ACL (1) | 6 |
| 2025 | Enhancing Study-Level Inference from Clinical Trial Papers via Reinforcement Learning-Based Numeric ReasoningabstractSystematic reviews in medicine play a critical role in evidence-based decision-making by aggregating findings from multiple studies.A central bottleneck in automating this process is extracting numeric evidence and determining study-level conclusions for specific outcomes and comparisons.Prior work has framed this problem as a textual inference task by retrieving relevant content fragments and inferring conclusions from them.However, such approaches often rely on shallow textual cues and fail to capture the underlying numeric reasoning behind expert assessments.In this work, we conceptualise the problem as one of quantitative reasoning.Rather than inferring conclusions from surface text, we extract structured numerical evidence (e.g., event counts or standard deviations) and apply domain knowledge informed logic to derive outcome-specific conclusions.We develop a numeric reasoning system composed of a numeric data extraction model and an effect estimate component, enabling more accurate and interpretable inference aligned with the domain expert principles.We train the numeric data extraction model using different strategies, including supervised fine-tuning (SFT), and reinforcement learning (RL) with a new value reward model.When evaluated on the COCHRANEFOREST benchmark, our best-performing approach -using RL to train a small-scale number extraction modelyields up to a 21% absolute improvement in F1 score over retrieval-based systems and outperforms general-purpose LLMs of over 400B parameters by up to 9%.Our results demonstrate the promise of reasoning-driven approaches for automating systematic evidence synthesis. Massimiliano Pronesti, Michela Lorandi, Paul Flanagan, Oisin Redmond, Anya Belz, Yufang Hou 0001 |
EMNLP | 5 |
| 2025 | Assessing Semantic Consistency in Data-to-Text Generation: A Meta-Evaluation of Textual, Semantic and Model-Based MetricsabstractEnsuring semantic consistency between semantic-triple inputs and generated text is crucial in data‐to‐text generation, but continues to pose challenges both during generation and in evaluation. In order to assess how accurately semantic consistency can currently be assessed, we meta-evaluate 29 different evaluation methods in terms of their ability to predict human semantic-consistency ratings. The evaluation methods include embeddings‐based, overlap‐based, and edit‐distance metrics, as well as learned regressors and a prompted ‘LLM‐as‐judge’ protocol. We meta-evaluate on two datasets: the WebNLG 2017 human evaluation dataset, and a newly created WebNLG-style dataset that none of the methods can have seen during training. We find that none of the traditional textual similarity metrics or the pre-Transformer model-based metrics are suitable for the task of semantic consistency assessment. LLM-based methods perform well on the whole, but best correlations with human judgments still lag behind those seen in other text generation tasks. Rudali Huidrom, Michela Lorandi, Simon Mille, Craig Thomson, Anya Belz |
INLG | 5 |
| 2025 | Scaling Up Data-to-Text Generation to Longer Sequences: A New Dataset and Benchmark Results for Generation from Large Triple SetsabstractThe ability of LLMs to write coherent, faithful long texts from structured data inputs remains relatively uncharted, in part because nearly all public data-to-text datasets contain only short input-output pairs. To address these gaps, we benchmark six LLMs, a rule‐based system and human-written texts on a new long-input dataset in English and Irish via LLM-based evaluation. We find substantial differences between models and languages. Chinonso Cynthia Osuji, Simon Mille, Ornait O'Connell, Thiago Castro Ferreira, Anya Belz, Brian Davis 0001 |
INLG | 5 |
| 2024 | Differences in Semantic Errors Made by Different Types of Data-to-text SystemsabstractIn this paper, we investigate how different semantic, or content-related, errors made by different types of data-to-text systems differ in terms of number and type.In total, we examine 15 systems: three rule-based and 12 neural systems including two large language models without training or fine-tuning.All systems were tested on the English WebNLG dataset version 3.0.We use a semantic error taxonomy and the brat annotation tool to obtain wordspan error annotations on a sample of system outputs.The annotations enable us to establish how many semantic errors different (types of) systems make and what specific types of errors they make, and thus to get an overall understanding of semantic strengths and weaknesses among various types of NLG systems.Among our main findings, we observe that symbolic (rule and template-based) systems make fewer semantic errors overall, non-LLM neural systems have better fluency and data coverage, but make more semantic errors, while LLM-based systems require improvement particularly in addressing superfluous. Rudali Huidrom, Anya Belz, Michela Lorandi |
INLG | 2 |
| 2024 | (Mostly) Automatic Experiment Execution for Human Evaluations of NLP SystemsabstractHuman evaluation is widely considered the most reliable form of evaluation in NLP, but recent research has shown it to be riddled with mistakes, often as a result of manual execution of tasks.This paper argues that such mistakes could be avoided if we were to automate, as much as is practical, the process of performing experiments for human evaluation of NLP systems.We provide a simple methodology that can improve both the transparency and reproducibility of experiments.We show how the sequence of component processes of a human evaluation can be defined in advance, facilitating full or partial automation, detailed preregistration of the process, and research transparency and repeatability. Craig Thomson, Anya Belz |
INLG | 2 |
| 2024 | Common Flaws in Running Human Evaluation Experiments in NLPabstractAbstract While conducting a coordinated set of repeat runs of human evaluation experiments in NLP, we discovered flaws in every single experiment we selected for inclusion via a systematic process. In this squib, we describe the types of flaws we discovered, which include coding errors (e.g., loading the wrong system outputs to evaluate), failure to follow standard scientific practice (e.g., ad hoc exclusion of participants and responses), and mistakes in reported numerical results (e.g., reported numbers not matching experimental data). If these problems are widespread, it would have worrying implications for the rigor of NLP evaluation experiments as currently conducted. We discuss what researchers can do to reduce the occurrence of such flaws, including pre-registration, better code development practices, increased testing and piloting, and post-publication addressing of errors. Craig Thomson, Ehud Reiter, Anya Belz |
Comput. Linguistics | 3 |
| 2023 | Mod-D2T: A Multi-layer Dataset for Modular Data-to-Text GenerationabstractRule-based text generators lack the coverage and fluency of their neural counterparts, but have two big advantages over them: (i) they are entirely controllable and do not hallucinate; and (ii) they can fully explain how an output was generated from an input.In this paper we leverage these two advantages to create large and reliable synthetic datasets with multiple human-intelligible intermediate representations.We present the Modular Data-to-Text (Mod-D2T) Dataset which incorporates ten intermediate-level representations between input triple sets and output text; the mappings from one level to the next can broadly be interpreted as the traditional modular tasks of an NLG pipeline.We describe the Mod-D2T dataset, evaluate its quality via manual validation and discuss its applications and limitations. Simon Mille, François Lareau, Stamatia Dasiopoulou, Anya Belz |
INLG | 4 |
| 2022 | Quantified Reproducibility Assessment of NLP ResultsabstractThis paper describes and tests a method for carrying out quantified reproducibility assessment (QRA) that is based on concepts and definitions from metrology.QRA produces a single score estimating the degree of reproducibility of a given system and evaluation measure, on the basis of the scores from, and differences between, different reproductions.We test QRA on 18 system and evaluation measure combinations (involving diverse NLP tasks and types of evaluation), for each of which we have the original results and one to seven reproduction results.The proposed QRA method produces degree-of-reproducibility scores that are comparable across multiple reproductions not only of the same, but of different original studies.We find that the proposed method facilitates insights into causes of variation between reproductions, and allows conclusions to be drawn about what changes to system and/or evaluation design might lead to improved reproducibility. Anya Belz, Maja Popovic, Simon Mille |
ACL (1) | 1 |
| 2022 | Human Evaluation and Correlation with Automatic Metrics in Consultation Note GenerationabstractFrancesco Moramarco, Alex Papadopoulos Korfiatis, Mark Perera, Damir Juric, Jack Flann, Ehud Reiter, Anya Belz, Aleksandar Savkov. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Francesco Moramarco, Alex Papadopoulos-Korfiatis, Mark Perera, Damir Juric, Jack Flann, Ehud Reiter, Anya Belz, Aleksandar Savkov |
ACL (1) | 7 |
| 2022 | User-Driven Research of Medical Note Generation SoftwareabstractTom Knoll, Francesco Moramarco, Alex Papadopoulos Korfiatis, Rachel Young, Claudia Ruffini, Mark Perera, Christian Perstl, Ehud Reiter, Anya Belz, Aleksandar Savkov. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Tom Knoll, Francesco Moramarco, Alex Papadopoulos-Korfiatis, Rachel Young, Claudia Ruffini, Mark Perera, Christian Perstl, Ehud Reiter, Anya Belz, Aleksandar Savkov |
NAACL-HLT | 9 |
| 2022 | A Metrological Perspective on Reproducibility in NLPabstractAbstract Reproducibility has become an increasingly debated topic in NLP and ML over recent years, but so far, no commonly accepted definitions of even basic terms or concepts have emerged. The range of different definitions proposed within NLP/ML not only do not agree with each other, they are also not aligned with standard scientific definitions. This article examines the standard definitions of repeatability and reproducibility provided by the meta-science of metrology, and explores what they imply in terms of how to assess reproducibility, and what adopting them would mean for reproducibility assessment in NLP/ML. It turns out the standard definitions lead directly to a method for assessing reproducibility in quantified terms that renders results from reproduction studies comparable across multiple reproductions of the same original study, as well as reproductions of different original studies. The article considers where this method sits in relation to other aspects of NLP work one might wish to assess in the context of reproducibility. Anya Belz |
Comput. Linguistics | 1 |
| 2021 | A Systematic Review of Reproducibility Research in Natural Language ProcessingabstractAgainst the background of what has been termed a reproducibility crisis in science, the NLP field is becoming increasingly interested in, and conscientious about, the reproducibility of its results.The past few years have seen an impressive range of new initiatives, events and active research in the area.However, the field is far from reaching a consensus about how reproducibility should be defined, measured and addressed, with diversity of views currently increasing rather than converging.With this focused contribution, we aim to provide a wideangle, and as near as possible complete, snapshot of current work on reproducibility in NLP, delineating differences and similarities, and providing pointers to common denominators. Anya Belz, Anastasia Shimorina, Ehud Reiter |
EACL | 1 |
| 2021 | The ReproGen Shared Task on Reproducibility of Human Evaluations in NLG: Overview and ResultsabstractThe NLP field has recently seen a substantial increase in work related to reproducibility of results, and more generally in recognition of the importance of having shared definitions and practices relating to evaluation.Much of the work on reproducibility has so far focused on metric scores, with reproducibility of human evaluation results receiving far less attention.As part of a research programme designed to develop theory and practice of reproducibility assessment in NLP, we organised the first shared task on reproducibility of human evaluations, ReproGen 2021.This paper describes the shared task in detail, summarises results from each of the reproduction studies submitted, and provides further comparative analysis of the results.Out of nine initial team registrations, we received submissions from four teams.Meta-analysis of the four reproduction studies revealed varying degrees of reproducibility, and allowed very tentative first conclusions about what types of evaluation tend to have better reproducibility. Anya Belz, Anastasia Shimorina, Ehud Reiter |
INLG | 1 |
| 2021 | Another PASS: A Reproduction Study of the Human Evaluation of a Football Report Generation SystemabstractThis paper reports results from a reproduction study in which we repeated the human evaluation of the PASS Dutch-language football report generation system (van der Lee et al., 2017).The work was carried out as part of the ReproGen Shared Task on Reproducibility of Human Evaluations in NLG, in Track A (Paper 1).We aimed to repeat the original study exactly, with the main difference that a different set of evaluators was used.We describe the study design, present the results from the original and the reproduction study, and then compare and analyse the differences between the two sets of results.For the two 'headline' results of average Fluency and Clarity, we find that in both studies, the system was rated more highly for Clarity than for Fluency, and Clarity had higher standard deviation.Clarity and Fluency ratings were higher, and their standard deviations lower, in the reproduction study than in the original study by substantial margins.Clarity had a higher degree of reproducibility than Fluency, as measured by the coefficient of variation.Data and code are publicly available.1 Simon Mille, Thiago Castro Ferreira, Anya Belz, Brian Davis 0001 |
INLG | 3 |
| 2021 | A Reproduction Study of an Annotation-based Human Evaluation of MT OutputsabstractIn this paper we report our reproduction study of the Croatian part of an annotation-based human evaluation of machine-translated user reviews (Popović, 2020).The work was carried out as part of the ReproGen Shared Task on Reproducibility of Human Evaluation in NLG.Our aim was to repeat the original study exactly, except for using a different set of evaluators.We describe the experimental design, characterise differences between original and reproduction study, and present the results from each study, along with analysis of the similarity between them.For the six main evaluation results of Major/Minor/All Comprehension error rates and Major/Minor/All Adequacy error rates, we find that (i) 4/6 system rankings are the same in both studies, (ii) the relative differences between systems are replicated well for Major Comprehension and Adequacy (Pearson's > 0.9), but not for the corresponding Minor error rates (Pearson's 0.36 for Adequacy, 0.67 for Comprehension), and (iii) the individual system scores for both types of Minor error rates had a higher degree of reproducibility than the corresponding Major error rates.We also examine inter-annotator agreement and compare the annotations obtained in the original and reproduction studies. Maja Popovic, Anya Belz |
INLG | 2 |
| 2020 | ReproGen: Proposal for a Shared Task on Reproducibility of Human Evaluations in NLGabstractAcross NLP, a growing body of work is looking at the issue of reproducibility.However, replicability of human evaluation experiments and reproducibility of their results is currently under-addressed, and this is of particular concern for NLG where human evaluations are the norm.This paper outlines our ideas for a shared task on reproducibility of human evaluations in NLG which aims (i) to shed light on the extent to which past NLG evaluations have been replicable and reproducible, and (ii) to draw conclusions regarding how evaluations can be designed and reported to increase replicability and reproducibility.If the task is run over several years, we hope to be able to document an overall increase in levels of replicability and reproducibility over time. Anya Belz, Anastasia Shimorina, Ehud Reiter |
INLG | 1 |
| 2020 | Disentangling the Properties of Human Evaluation Methods: A Classification System to Support Comparability, Meta-Evaluation and Reproducibility TestingabstractCurrent standards for designing and reporting human evaluations in NLP mean it is generally unclear which evaluations are comparable and can be expected to yield similar results when applied to the same system outputs.This has serious implications for reproducibility testing and meta-evaluation, in particular given that human evaluation is considered the gold standard against which the trustworthiness of automatic metrics is gauged.Using examples from NLG, we propose a classification system for evaluations based on disentangling (i) what is being evaluated (which aspect of quality), and (ii) how it is evaluated in specific (a) evaluation modes and (b) experimental designs.We show that this approach provides a basis for determining comparability, hence for comparison of evaluations across papers, meta-evaluation experiments, reproducibility testing. Anya Belz, Simon Mille, David M. Howcroft |
INLG | 1 |
| 2020 | Twenty Years of Confusion in Human Evaluation: NLG Needs Evaluation Sheets and Standardised DefinitionsabstractDavid M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, Verena Rieser. Proceedings of the 13th International Conference on Natural Language Generation. 2020. David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, Verena Rieser |
INLG | 2 |
| 2018 | SpatialVOC2K: A Multilingual Dataset of Images with Annotations and Features for Spatial Relations between ObjectsabstractWe present SpatialVOC2K, the first multilingual image dataset with spatial relation annotations and object features for imageto-text generation, built using 2,026 images from the PASCAL VOC2008 dataset.The dataset incorporates (i) the labelled object bounding boxes from VOC2008, (ii) geometrical, language and depth features for each object, and (iii) for each pair of objects in both orders, (a) the single best preposition and (b) the set of possible prepositions in the given language that describe the spatial relationship between the two objects.Compared to previous versions of the dataset, we have roughly doubled the size for French, and completely reannotated as well as increased the size of the English portion, providing single best prepositions for English for the first time.Furthermore, we have added explicit 3D depth features for objects.We are releasing our dataset for free reuse, along with evaluation tools to enable comparative evaluation. Anya Belz, Adrian Muscat, Pierre Anguill, Mouhamadou Sow, Gaetan Vincent, Yassine Zinessabah |
INLG | 1 |
| 2018 | Adding the Third Dimension to Spatial Relation Detection in 2D ImagesabstractDetection of spatial relations between objects in images is currently a popular subject in image description research.A range of different language and geometric object features have been used in this context, but methods have not so far used explicit information about the third dimension (depth), except when manually added to annotations.The lack of such information hampers detection of spatial relations that are inherently 3D.In this paper, we use a fully automatic method for creating a depth map of an image and derive several different object-level depth features from it which we add to an existing feature set to test the effect on spatial relation detection.We show that performance increases are obtained from adding depth features in all scenarios tested. Brandon Birmingham, Adrian Muscat, Anya Belz |
INLG | 3 |
| 2018 | Underspecified Universal Dependency Structures as Inputs for Multilingual Surface RealisationabstractIn this paper, we present the datasets used in the Shallow and Deep Tracks of the First Multilingual Surface Realisation Shared Task (SR'18).For the Shallow Track, data in ten languages has been released: Arabic, Czech, Dutch, English, Finnish, French, Italian, Portuguese, Russian and Spanish.For the Deep Track, data in three languages is made available: English, French and Spanish.We describe in detail how the datasets were derived from the Universal Dependencies V2.0, and report on an evaluation of the Deep Track input quality.In addition, we examine the motivation for, and likely usefulness of, deriving NLG inputs from annotations in resources originally developed for Natural Language Understanding (NLU), and assess whether the resulting inputs supply enough information of the right kind for the final stage in the NLG process. Simon Mille, Anya Belz, Bernd Bohnet, Leo Wanner |
INLG | 2 |
| 2018 | From image to language and back againabstractWork in computer vision and natural language processing involving images and text has been experiencing explosive growth over the past decade, with a particular boost coming from the neural network revolution. The present volume brings together five research articles from several different corners of the area: multilingual multimodal image description (Franket al.), multimodal machine translation (Madhyasthaet al., Franket al.), image caption generation (Madhyasthaet al., Tantiet al.), visual scene understanding (Silbereret al.), and multimodal learning of high-level attributes (Sorodocet al.). In this article, we touch upon all of these topics as we review work involving images and text under the three main headings of image description (Section 2), visually grounded referring expression generation (REG) and comprehension (Section 3), and visual question answering (VQA) (Section 4). Anya Belz, Tamara L. Berg, Licheng Yu |
Nat. Lang. Eng. | 1 |
| 2017 | Shared Task Proposal: Multilingual Surface Realization Using Universal Dependency TreesabstractWe propose a shared task on multilingual Surface Realization, i.e., on mapping unordered and uninflected universal dependency trees to correctly ordered and inflected sentences in a number of languages.A second deeper input will be available in which, in addition, functional words, fine-grained PoS and morphological information will be removed from the input trees.The first shared task on Surface Realization was carried out in 2011 with a similar setup, with a focus on English.We think that it is time for relaunching such a shared task effort in view of the arrival of Universal Dependencies annotated treebanks for a large number of languages on the one hand, and the increasing dominance of Deep Learning, which proved to be a game changer for NLP, on the other hand. Simon Mille, Bernd Bohnet, Leo Wanner, Anya Belz |
INLG | 4 |
| 2016 | Effect of Data Annotation, Feature Selection and Model Choice on Spatial Description Generation in FrenchabstractIn this paper, we look at automatic generation of spatial descriptions in French, more particularly, selecting a spatial preposition for a pair of objects in an image.Our focus is on assessing the effect on accuracy of (i) increasing data set size, (ii) removing synonyms from the set of prepositions used for annotation, (iii) optimising feature sets, and (iv) training on best prepositions only vs. training on all acceptable prepositions.We describe a new data set where each object pair in each image is annotated with the best and all acceptable prepositions that describe the spatial relationship between the two objects.We report results for three new methods for this task, and find that the best, 75% Accuracy, is 25 points higher than our previous best result for this task. Anya Belz, Adrian Muscat, Brandon Birmingham, Jessie Levacher, Julie Pain, Adam Quinquenel |
INLG | 1 |
| 2014 | A Comparative Evaluation Methodology for NLG in Interactive Systems
Helen Hastie, Anya Belz |
LREC | 2 |
| 2012 | The Surface Realisation Task: Recent Developments and Future Plans
Anya Belz, Bernd Bohnet, Simon Mille, Leo Wanner, Michael White 0001 |
INLG | 1 |
| 2012 | A Repository of Data and Evaluation Resources for Natural Language Generation
Anya Belz, Albert Gatt |
LREC | 1 |
| 2012 | LG-Eval: A Toolkit for Creating Online Language Evaluation Experiments
Eric Kow, Anya Belz |
LREC | 2 |
| 2010 | Generation Challenges 2010 Preface
Anya Belz, Albert Gatt, Alexander Koller |
INLG | 1 |
| 2010 | Comparing Rating Scales and Preference Judgements in Language Evaluation
Anya Belz, Eric Kow |
INLG | 1 |
| 2010 | Extracting Parallel Fragments from Comparable Corpora for Data-to-text Generation
Anya Belz, Eric Kow |
INLG | 1 |
| 2010 | The GREC Challenges 2010: Overview and Evaluation Results
Anya Belz, Eric Kow |
INLG | 1 |
| 2010 | Finding Common Ground: Towards a Surface Realisation Shared Task
Anya Belz, Mike White, Josef van Genabith, Deirdre Hogan, Amanda Stent |
INLG | 1 |
| 2010 | A Game-based Approach to Transcribing Images of Text
Khalil Dahab, Anya Belz |
LREC | 2 |
| 2009 | That's Nice ... What Can You Do With It?abstractA regular fixture on the mid 1990s international research seminar circuit was the "billion-neuron artificial brain" talk.The idea behind this project was simple: in order to create artificial intelligence, what was needed first of all was a very large artificial brain; if a big enough set of interconnected modules of neurons could be implemented, then it would be possible to evolve mammalian-level behavior with current computationalneuron technology.The talk included progress reports on the current size of the artificial brain, its structure, "update rate," and power consumption, and explained how intelligent behavior was going to develop by mechanisms simulating biological evolution.What the talk didn't mention was what kind of functionality the team had so far managed to evolve, and so the first comment at the end of the talk was inevitably "nice work, but have you actually done anything with the brain yet?" 1 In human language technology (HLT) research, we currently report a range of evaluation scores that measure and assess various aspects of systems, in particular the similarity of their outputs to samples of human language or to human-produced goldstandard annotations, but are we leaving ourselves open to the same question as the billion-neuron artificial brain researchers? Shrinking HorizonsHLT evaluation has a long history.Spärck Jones's Information Retrieval Experiment (1981) already had two decades of IR evaluation history to look back on.It provides a fairly comprehensive snapshot of HLT evaluation at the time, as much of HLT evaluation research was in the field of IR.One thing that is striking from today's perspective is the rich diversity of evaluation paradigms-user-oriented and developer-oriented, intrinsic and extrinsic 2 -that were being investigated and discussed on an equal footing Anya Belz |
Comput. Linguistics | 1 |
| 2009 | An Investigation into the Validity of Some Metrics for Automatically Evaluating Natural Language Generation SystemsabstractThere is growing interest in using automatically computed corpus-based evaluation metrics to evaluate Natural Language Generation (NLG) systems, because these are often considerably cheaper than the human-based evaluations which have traditionally been used in NLG. We review previous work on NLG evaluation and on validation of automatic metrics in NLP, and then present the results of two studies of how well some metrics which are popular in other areas of NLP (notably BLEU and ROUGE) correlate with human judgments in the domain of computer-generated weather forecasts. Our results suggest that, at least in this domain, metrics may provide a useful measure of language quality, although the evidence for this is not as strong as we would ideally like to see; however, they do not provide a useful measure of content quality. We also discuss a number of caveats which must be kept in mind when interpreting this and other validation studies. Ehud Reiter, Anya Belz |
Comput. Linguistics | 2 |
| 2008 | REG Challenge Preface
Anya Belz, Albert Gatt |
INLG | 1 |
| 2008 | The GREC Challenge 2008: Overview and Evaluation Results
Anya Belz, Eric Kow, Jette Viethen, Albert Gatt |
INLG | 1 |
| 2008 | Attribute Selection for Referring Expression Generation: New Algorithms and Evaluation Methods
Albert Gatt, Anya Belz |
INLG | 2 |
| 2008 | The TUNA Challenge 2008: Overview and Evaluation Results
Albert Gatt, Anya Belz, Eric Kow |
INLG | 2 |
| 2008 | Automatic generation of weather forecast texts using comprehensive probabilistic generation-space modelsabstractAbstract Two important recent trends in natural language generation are (i) probabilistic techniques and (ii) comprehensive approaches that move away from traditional strictly modular and sequential models. This paper reports experiments in whichpcru– a generation framework that combines probabilistic generation methodology with a comprehensive model of the generation space – was used to semi-automatically create five different versions of a weather forecast generator. The generators were evaluated in terms of output quality, development time and computational efficiency against (i) human forecasters, (ii) a traditional handcrafted pipelinednlgsystem and (iii) ahalogen-style statistical generator. The most striking result is that despite acquiring all decision-making abilities automatically, the bestpcrugenerators produce outputs of high enough quality to be scored more highly by human judges than forecasts written by experts. Anya Belz |
Nat. Lang. Eng. | 1 |
| 2007 | Probabilistic Generation of Weather Forecast Texts
Anya Belz |
HLT-NAACL | 1 |
| 2006 | Comparing Automatic and Human Evaluation of NLG Systems
Anya Belz, Ehud Reiter |
EACL | 1 |
| 2006 | Introduction to the INLG'06 Special Session on Sharing Data and Comparative Evaluation
Anya Belz, Robert Dale |
INLG | 1 |
| 2006 | Shared-Task Evaluations in HLT: Lessons for NLG
Anya Belz, Adam Kilgarriff |
INLG | 1 |
| 2006 | GENEVAL: A Proposal for Shared-task Evaluation in NLG
Ehud Reiter, Anya Belz |
INLG | 2 |
| 2002 | Learning Grammars for Different Parsing Tasks by Partition Search
Anya Belz |
COLING | 1 |
| 2002 | PILLS: Multilingual generation of medical information documents with overlapping content
Nadjet Bouayad-Agha, Richard Power, Donia Scott, Anya Belz |
LREC | 4 |
| 1998 | A Few English Words Can Help Improve Your Russian
Anya Belz |
ECAI | 1 |