Anya Belz

dblp:212/6084 · also Anja Belz · DBLP profile ↗
← Back
50ranked-venue papers
26as first author
16since 2021 · last 2025
0000-0002-0552-8096ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 26 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 Query-driven Document-level Scientific Evidence Extraction from Biomedical Studies
abstract
Massimiliano Pronesti, Joao H Bettencourt-Silva, Paul Flanagan, Alessandra Pascale, Oisín Redmond, Anya Belz, Yufang Hou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Massimiliano Pronesti, Joao H. Bettencourt-Silva, Paul Flanagan, Alessandra Pascale, Oisin Redmond, Anya Belz, Yufang Hou 0001
ACL (1)6
2025 Enhancing Study-Level Inference from Clinical Trial Papers via Reinforcement Learning-Based Numeric Reasoning
abstract
Systematic reviews in medicine play a critical role in evidence-based decision-making by aggregating findings from multiple studies.A central bottleneck in automating this process is extracting numeric evidence and determining study-level conclusions for specific outcomes and comparisons.Prior work has framed this problem as a textual inference task by retrieving relevant content fragments and inferring conclusions from them.However, such approaches often rely on shallow textual cues and fail to capture the underlying numeric reasoning behind expert assessments.In this work, we conceptualise the problem as one of quantitative reasoning.Rather than inferring conclusions from surface text, we extract structured numerical evidence (e.g., event counts or standard deviations) and apply domain knowledge informed logic to derive outcome-specific conclusions.We develop a numeric reasoning system composed of a numeric data extraction model and an effect estimate component, enabling more accurate and interpretable inference aligned with the domain expert principles.We train the numeric data extraction model using different strategies, including supervised fine-tuning (SFT), and reinforcement learning (RL) with a new value reward model.When evaluated on the COCHRANEFOREST benchmark, our best-performing approach -using RL to train a small-scale number extraction modelyields up to a 21% absolute improvement in F1 score over retrieval-based systems and outperforms general-purpose LLMs of over 400B parameters by up to 9%.Our results demonstrate the promise of reasoning-driven approaches for automating systematic evidence synthesis.
Massimiliano Pronesti, Michela Lorandi, Paul Flanagan, Oisin Redmond, Anya Belz, Yufang Hou 0001
EMNLP5
2025 Assessing Semantic Consistency in Data-to-Text Generation: A Meta-Evaluation of Textual, Semantic and Model-Based Metrics
abstract
Ensuring semantic consistency between semantic-triple inputs and generated text is crucial in data‐to‐text generation, but continues to pose challenges both during generation and in evaluation. In order to assess how accurately semantic consistency can currently be assessed, we meta-evaluate 29 different evaluation methods in terms of their ability to predict human semantic-consistency ratings. The evaluation methods include embeddings‐based, overlap‐based, and edit‐distance metrics, as well as learned regressors and a prompted ‘LLM‐as‐judge’ protocol. We meta-evaluate on two datasets: the WebNLG 2017 human evaluation dataset, and a newly created WebNLG-style dataset that none of the methods can have seen during training. We find that none of the traditional textual similarity metrics or the pre-Transformer model-based metrics are suitable for the task of semantic consistency assessment. LLM-based methods perform well on the whole, but best correlations with human judgments still lag behind those seen in other text generation tasks.
Rudali Huidrom, Michela Lorandi, Simon Mille, Craig Thomson, Anya Belz
INLG5
2025 Scaling Up Data-to-Text Generation to Longer Sequences: A New Dataset and Benchmark Results for Generation from Large Triple Sets
abstract
The ability of LLMs to write coherent, faithful long texts from structured data inputs remains relatively uncharted, in part because nearly all public data-to-text datasets contain only short input-output pairs. To address these gaps, we benchmark six LLMs, a rule‐based system and human-written texts on a new long-input dataset in English and Irish via LLM-based evaluation. We find substantial differences between models and languages.
Chinonso Cynthia Osuji, Simon Mille, Ornait O'Connell, Thiago Castro Ferreira, Anya Belz, Brian Davis 0001
INLG5
2024 Differences in Semantic Errors Made by Different Types of Data-to-text Systems
abstract
In this paper, we investigate how different semantic, or content-related, errors made by different types of data-to-text systems differ in terms of number and type.In total, we examine 15 systems: three rule-based and 12 neural systems including two large language models without training or fine-tuning.All systems were tested on the English WebNLG dataset version 3.0.We use a semantic error taxonomy and the brat annotation tool to obtain wordspan error annotations on a sample of system outputs.The annotations enable us to establish how many semantic errors different (types of) systems make and what specific types of errors they make, and thus to get an overall understanding of semantic strengths and weaknesses among various types of NLG systems.Among our main findings, we observe that symbolic (rule and template-based) systems make fewer semantic errors overall, non-LLM neural systems have better fluency and data coverage, but make more semantic errors, while LLM-based systems require improvement particularly in addressing superfluous.
Rudali Huidrom, Anya Belz, Michela Lorandi
INLG2
2024 (Mostly) Automatic Experiment Execution for Human Evaluations of NLP Systems
abstract
Human evaluation is widely considered the most reliable form of evaluation in NLP, but recent research has shown it to be riddled with mistakes, often as a result of manual execution of tasks.This paper argues that such mistakes could be avoided if we were to automate, as much as is practical, the process of performing experiments for human evaluation of NLP systems.We provide a simple methodology that can improve both the transparency and reproducibility of experiments.We show how the sequence of component processes of a human evaluation can be defined in advance, facilitating full or partial automation, detailed preregistration of the process, and research transparency and repeatability.
Craig Thomson, Anya Belz
INLG2
2024 Common Flaws in Running Human Evaluation Experiments in NLP
abstract
Abstract While conducting a coordinated set of repeat runs of human evaluation experiments in NLP, we discovered flaws in every single experiment we selected for inclusion via a systematic process. In this squib, we describe the types of flaws we discovered, which include coding errors (e.g., loading the wrong system outputs to evaluate), failure to follow standard scientific practice (e.g., ad hoc exclusion of participants and responses), and mistakes in reported numerical results (e.g., reported numbers not matching experimental data). If these problems are widespread, it would have worrying implications for the rigor of NLP evaluation experiments as currently conducted. We discuss what researchers can do to reduce the occurrence of such flaws, including pre-registration, better code development practices, increased testing and piloting, and post-publication addressing of errors.
Craig Thomson, Ehud Reiter, Anya Belz
Comput. Linguistics3
2023 Mod-D2T: A Multi-layer Dataset for Modular Data-to-Text Generation
abstract
Rule-based text generators lack the coverage and fluency of their neural counterparts, but have two big advantages over them: (i) they are entirely controllable and do not hallucinate; and (ii) they can fully explain how an output was generated from an input.In this paper we leverage these two advantages to create large and reliable synthetic datasets with multiple human-intelligible intermediate representations.We present the Modular Data-to-Text (Mod-D2T) Dataset which incorporates ten intermediate-level representations between input triple sets and output text; the mappings from one level to the next can broadly be interpreted as the traditional modular tasks of an NLG pipeline.We describe the Mod-D2T dataset, evaluate its quality via manual validation and discuss its applications and limitations.
Simon Mille, François Lareau, Stamatia Dasiopoulou, Anya Belz
INLG4
2022 Quantified Reproducibility Assessment of NLP Results
abstract
This paper describes and tests a method for carrying out quantified reproducibility assessment (QRA) that is based on concepts and definitions from metrology.QRA produces a single score estimating the degree of reproducibility of a given system and evaluation measure, on the basis of the scores from, and differences between, different reproductions.We test QRA on 18 system and evaluation measure combinations (involving diverse NLP tasks and types of evaluation), for each of which we have the original results and one to seven reproduction results.The proposed QRA method produces degree-of-reproducibility scores that are comparable across multiple reproductions not only of the same, but of different original studies.We find that the proposed method facilitates insights into causes of variation between reproductions, and allows conclusions to be drawn about what changes to system and/or evaluation design might lead to improved reproducibility.
Anya Belz, Maja Popovic, Simon Mille
ACL (1)1
2022 Human Evaluation and Correlation with Automatic Metrics in Consultation Note Generation
abstract
Francesco Moramarco, Alex Papadopoulos Korfiatis, Mark Perera, Damir Juric, Jack Flann, Ehud Reiter, Anya Belz, Aleksandar Savkov. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Francesco Moramarco, Alex Papadopoulos-Korfiatis, Mark Perera, Damir Juric, Jack Flann, Ehud Reiter, Anya Belz, Aleksandar Savkov
ACL (1)7
2022 User-Driven Research of Medical Note Generation Software
abstract
Tom Knoll, Francesco Moramarco, Alex Papadopoulos Korfiatis, Rachel Young, Claudia Ruffini, Mark Perera, Christian Perstl, Ehud Reiter, Anya Belz, Aleksandar Savkov. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Tom Knoll, Francesco Moramarco, Alex Papadopoulos-Korfiatis, Rachel Young, Claudia Ruffini, Mark Perera, Christian Perstl, Ehud Reiter, Anya Belz, Aleksandar Savkov
NAACL-HLT9
2022 A Metrological Perspective on Reproducibility in NLP
abstract
Abstract Reproducibility has become an increasingly debated topic in NLP and ML over recent years, but so far, no commonly accepted definitions of even basic terms or concepts have emerged. The range of different definitions proposed within NLP/ML not only do not agree with each other, they are also not aligned with standard scientific definitions. This article examines the standard definitions of repeatability and reproducibility provided by the meta-science of metrology, and explores what they imply in terms of how to assess reproducibility, and what adopting them would mean for reproducibility assessment in NLP/ML. It turns out the standard definitions lead directly to a method for assessing reproducibility in quantified terms that renders results from reproduction studies comparable across multiple reproductions of the same original study, as well as reproductions of different original studies. The article considers where this method sits in relation to other aspects of NLP work one might wish to assess in the context of reproducibility.
Anya Belz
Comput. Linguistics1
2021 A Systematic Review of Reproducibility Research in Natural Language Processing
abstract
Against the background of what has been termed a reproducibility crisis in science, the NLP field is becoming increasingly interested in, and conscientious about, the reproducibility of its results.The past few years have seen an impressive range of new initiatives, events and active research in the area.However, the field is far from reaching a consensus about how reproducibility should be defined, measured and addressed, with diversity of views currently increasing rather than converging.With this focused contribution, we aim to provide a wideangle, and as near as possible complete, snapshot of current work on reproducibility in NLP, delineating differences and similarities, and providing pointers to common denominators.
Anya Belz, Anastasia Shimorina, Ehud Reiter
EACL1
2021 The ReproGen Shared Task on Reproducibility of Human Evaluations in NLG: Overview and Results
abstract
The NLP field has recently seen a substantial increase in work related to reproducibility of results, and more generally in recognition of the importance of having shared definitions and practices relating to evaluation.Much of the work on reproducibility has so far focused on metric scores, with reproducibility of human evaluation results receiving far less attention.As part of a research programme designed to develop theory and practice of reproducibility assessment in NLP, we organised the first shared task on reproducibility of human evaluations, ReproGen 2021.This paper describes the shared task in detail, summarises results from each of the reproduction studies submitted, and provides further comparative analysis of the results.Out of nine initial team registrations, we received submissions from four teams.Meta-analysis of the four reproduction studies revealed varying degrees of reproducibility, and allowed very tentative first conclusions about what types of evaluation tend to have better reproducibility.
Anya Belz, Anastasia Shimorina, Ehud Reiter
INLG1
2021 Another PASS: A Reproduction Study of the Human Evaluation of a Football Report Generation System
abstract
This paper reports results from a reproduction study in which we repeated the human evaluation of the PASS Dutch-language football report generation system (van der Lee et al., 2017).The work was carried out as part of the ReproGen Shared Task on Reproducibility of Human Evaluations in NLG, in Track A (Paper 1).We aimed to repeat the original study exactly, with the main difference that a different set of evaluators was used.We describe the study design, present the results from the original and the reproduction study, and then compare and analyse the differences between the two sets of results.For the two 'headline' results of average Fluency and Clarity, we find that in both studies, the system was rated more highly for Clarity than for Fluency, and Clarity had higher standard deviation.Clarity and Fluency ratings were higher, and their standard deviations lower, in the reproduction study than in the original study by substantial margins.Clarity had a higher degree of reproducibility than Fluency, as measured by the coefficient of variation.Data and code are publicly available.1
Simon Mille, Thiago Castro Ferreira, Anya Belz, Brian Davis 0001
INLG3
2021 A Reproduction Study of an Annotation-based Human Evaluation of MT Outputs
abstract
In this paper we report our reproduction study of the Croatian part of an annotation-based human evaluation of machine-translated user reviews (Popović, 2020).The work was carried out as part of the ReproGen Shared Task on Reproducibility of Human Evaluation in NLG.Our aim was to repeat the original study exactly, except for using a different set of evaluators.We describe the experimental design, characterise differences between original and reproduction study, and present the results from each study, along with analysis of the similarity between them.For the six main evaluation results of Major/Minor/All Comprehension error rates and Major/Minor/All Adequacy error rates, we find that (i) 4/6 system rankings are the same in both studies, (ii) the relative differences between systems are replicated well for Major Comprehension and Adequacy (Pearson's > 0.9), but not for the corresponding Minor error rates (Pearson's 0.36 for Adequacy, 0.67 for Comprehension), and (iii) the individual system scores for both types of Minor error rates had a higher degree of reproducibility than the corresponding Major error rates.We also examine inter-annotator agreement and compare the annotations obtained in the original and reproduction studies.
Maja Popovic, Anya Belz
INLG2
2020 ReproGen: Proposal for a Shared Task on Reproducibility of Human Evaluations in NLG
abstract
Across NLP, a growing body of work is looking at the issue of reproducibility.However, replicability of human evaluation experiments and reproducibility of their results is currently under-addressed, and this is of particular concern for NLG where human evaluations are the norm.This paper outlines our ideas for a shared task on reproducibility of human evaluations in NLG which aims (i) to shed light on the extent to which past NLG evaluations have been replicable and reproducible, and (ii) to draw conclusions regarding how evaluations can be designed and reported to increase replicability and reproducibility.If the task is run over several years, we hope to be able to document an overall increase in levels of replicability and reproducibility over time.
Anya Belz, Anastasia Shimorina, Ehud Reiter
INLG1
2020 Disentangling the Properties of Human Evaluation Methods: A Classification System to Support Comparability, Meta-Evaluation and Reproducibility Testing
abstract
Current standards for designing and reporting human evaluations in NLP mean it is generally unclear which evaluations are comparable and can be expected to yield similar results when applied to the same system outputs.This has serious implications for reproducibility testing and meta-evaluation, in particular given that human evaluation is considered the gold standard against which the trustworthiness of automatic metrics is gauged.Using examples from NLG, we propose a classification system for evaluations based on disentangling (i) what is being evaluated (which aspect of quality), and (ii) how it is evaluated in specific (a) evaluation modes and (b) experimental designs.We show that this approach provides a basis for determining comparability, hence for comparison of evaluations across papers, meta-evaluation experiments, reproducibility testing.
Anya Belz, Simon Mille, David M. Howcroft
INLG1
2020 Twenty Years of Confusion in Human Evaluation: NLG Needs Evaluation Sheets and Standardised Definitions
abstract
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, Verena Rieser. Proceedings of the 13th International Conference on Natural Language Generation. 2020.
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, Verena Rieser
INLG2
2018 SpatialVOC2K: A Multilingual Dataset of Images with Annotations and Features for Spatial Relations between Objects
abstract
We present SpatialVOC2K, the first multilingual image dataset with spatial relation annotations and object features for imageto-text generation, built using 2,026 images from the PASCAL VOC2008 dataset.The dataset incorporates (i) the labelled object bounding boxes from VOC2008, (ii) geometrical, language and depth features for each object, and (iii) for each pair of objects in both orders, (a) the single best preposition and (b) the set of possible prepositions in the given language that describe the spatial relationship between the two objects.Compared to previous versions of the dataset, we have roughly doubled the size for French, and completely reannotated as well as increased the size of the English portion, providing single best prepositions for English for the first time.Furthermore, we have added explicit 3D depth features for objects.We are releasing our dataset for free reuse, along with evaluation tools to enable comparative evaluation.
Anya Belz, Adrian Muscat, Pierre Anguill, Mouhamadou Sow, Gaetan Vincent, Yassine Zinessabah
INLG1
2018 Adding the Third Dimension to Spatial Relation Detection in 2D Images
abstract
Detection of spatial relations between objects in images is currently a popular subject in image description research.A range of different language and geometric object features have been used in this context, but methods have not so far used explicit information about the third dimension (depth), except when manually added to annotations.The lack of such information hampers detection of spatial relations that are inherently 3D.In this paper, we use a fully automatic method for creating a depth map of an image and derive several different object-level depth features from it which we add to an existing feature set to test the effect on spatial relation detection.We show that performance increases are obtained from adding depth features in all scenarios tested.
Brandon Birmingham, Adrian Muscat, Anya Belz
INLG3
2018 Underspecified Universal Dependency Structures as Inputs for Multilingual Surface Realisation
abstract
In this paper, we present the datasets used in the Shallow and Deep Tracks of the First Multilingual Surface Realisation Shared Task (SR'18).For the Shallow Track, data in ten languages has been released: Arabic, Czech, Dutch, English, Finnish, French, Italian, Portuguese, Russian and Spanish.For the Deep Track, data in three languages is made available: English, French and Spanish.We describe in detail how the datasets were derived from the Universal Dependencies V2.0, and report on an evaluation of the Deep Track input quality.In addition, we examine the motivation for, and likely usefulness of, deriving NLG inputs from annotations in resources originally developed for Natural Language Understanding (NLU), and assess whether the resulting inputs supply enough information of the right kind for the final stage in the NLG process.
Simon Mille, Anya Belz, Bernd Bohnet, Leo Wanner
INLG2
2018 From image to language and back again
abstract
Work in computer vision and natural language processing involving images and text has been experiencing explosive growth over the past decade, with a particular boost coming from the neural network revolution. The present volume brings together five research articles from several different corners of the area: multilingual multimodal image description (Franket al.), multimodal machine translation (Madhyasthaet al., Franket al.), image caption generation (Madhyasthaet al., Tantiet al.), visual scene understanding (Silbereret al.), and multimodal learning of high-level attributes (Sorodocet al.). In this article, we touch upon all of these topics as we review work involving images and text under the three main headings of image description (Section 2), visually grounded referring expression generation (REG) and comprehension (Section 3), and visual question answering (VQA) (Section 4).
Anya Belz, Tamara L. Berg, Licheng Yu
Nat. Lang. Eng.1
2017 Shared Task Proposal: Multilingual Surface Realization Using Universal Dependency Trees
abstract
We propose a shared task on multilingual Surface Realization, i.e., on mapping unordered and uninflected universal dependency trees to correctly ordered and inflected sentences in a number of languages.A second deeper input will be available in which, in addition, functional words, fine-grained PoS and morphological information will be removed from the input trees.The first shared task on Surface Realization was carried out in 2011 with a similar setup, with a focus on English.We think that it is time for relaunching such a shared task effort in view of the arrival of Universal Dependencies annotated treebanks for a large number of languages on the one hand, and the increasing dominance of Deep Learning, which proved to be a game changer for NLP, on the other hand.
Simon Mille, Bernd Bohnet, Leo Wanner, Anya Belz
INLG4
2016 Effect of Data Annotation, Feature Selection and Model Choice on Spatial Description Generation in French
abstract
In this paper, we look at automatic generation of spatial descriptions in French, more particularly, selecting a spatial preposition for a pair of objects in an image.Our focus is on assessing the effect on accuracy of (i) increasing data set size, (ii) removing synonyms from the set of prepositions used for annotation, (iii) optimising feature sets, and (iv) training on best prepositions only vs. training on all acceptable prepositions.We describe a new data set where each object pair in each image is annotated with the best and all acceptable prepositions that describe the spatial relationship between the two objects.We report results for three new methods for this task, and find that the best, 75% Accuracy, is 25 points higher than our previous best result for this task.
Anya Belz, Adrian Muscat, Brandon Birmingham, Jessie Levacher, Julie Pain, Adam Quinquenel
INLG1
2014 A Comparative Evaluation Methodology for NLG in Interactive Systems
Helen Hastie, Anya Belz
LREC2
2012 The Surface Realisation Task: Recent Developments and Future Plans
Anya Belz, Bernd Bohnet, Simon Mille, Leo Wanner, Michael White 0001
INLG1
2012 A Repository of Data and Evaluation Resources for Natural Language Generation
Anya Belz, Albert Gatt
LREC1
2012 LG-Eval: A Toolkit for Creating Online Language Evaluation Experiments
Eric Kow, Anya Belz
LREC2
2010 Generation Challenges 2010 Preface
Anya Belz, Albert Gatt, Alexander Koller
INLG1
2010 Comparing Rating Scales and Preference Judgements in Language Evaluation
Anya Belz, Eric Kow
INLG1
2010 Extracting Parallel Fragments from Comparable Corpora for Data-to-text Generation
Anya Belz, Eric Kow
INLG1
2010 The GREC Challenges 2010: Overview and Evaluation Results
Anya Belz, Eric Kow
INLG1
2010 Finding Common Ground: Towards a Surface Realisation Shared Task
Anya Belz, Mike White, Josef van Genabith, Deirdre Hogan, Amanda Stent
INLG1
2010 A Game-based Approach to Transcribing Images of Text
Khalil Dahab, Anya Belz
LREC2
2009 That's Nice ... What Can You Do With It?
abstract
A regular fixture on the mid 1990s international research seminar circuit was the "billion-neuron artificial brain" talk.The idea behind this project was simple: in order to create artificial intelligence, what was needed first of all was a very large artificial brain; if a big enough set of interconnected modules of neurons could be implemented, then it would be possible to evolve mammalian-level behavior with current computationalneuron technology.The talk included progress reports on the current size of the artificial brain, its structure, "update rate," and power consumption, and explained how intelligent behavior was going to develop by mechanisms simulating biological evolution.What the talk didn't mention was what kind of functionality the team had so far managed to evolve, and so the first comment at the end of the talk was inevitably "nice work, but have you actually done anything with the brain yet?" 1 In human language technology (HLT) research, we currently report a range of evaluation scores that measure and assess various aspects of systems, in particular the similarity of their outputs to samples of human language or to human-produced goldstandard annotations, but are we leaving ourselves open to the same question as the billion-neuron artificial brain researchers? Shrinking HorizonsHLT evaluation has a long history.Spärck Jones's Information Retrieval Experiment (1981) already had two decades of IR evaluation history to look back on.It provides a fairly comprehensive snapshot of HLT evaluation at the time, as much of HLT evaluation research was in the field of IR.One thing that is striking from today's perspective is the rich diversity of evaluation paradigms-user-oriented and developer-oriented, intrinsic and extrinsic 2 -that were being investigated and discussed on an equal footing
Anya Belz
Comput. Linguistics1
2009 An Investigation into the Validity of Some Metrics for Automatically Evaluating Natural Language Generation Systems
abstract
There is growing interest in using automatically computed corpus-based evaluation metrics to evaluate Natural Language Generation (NLG) systems, because these are often considerably cheaper than the human-based evaluations which have traditionally been used in NLG. We review previous work on NLG evaluation and on validation of automatic metrics in NLP, and then present the results of two studies of how well some metrics which are popular in other areas of NLP (notably BLEU and ROUGE) correlate with human judgments in the domain of computer-generated weather forecasts. Our results suggest that, at least in this domain, metrics may provide a useful measure of language quality, although the evidence for this is not as strong as we would ideally like to see; however, they do not provide a useful measure of content quality. We also discuss a number of caveats which must be kept in mind when interpreting this and other validation studies.
Ehud Reiter, Anya Belz
Comput. Linguistics2
2008 REG Challenge Preface
Anya Belz, Albert Gatt
INLG1
2008 The GREC Challenge 2008: Overview and Evaluation Results
Anya Belz, Eric Kow, Jette Viethen, Albert Gatt
INLG1
2008 Attribute Selection for Referring Expression Generation: New Algorithms and Evaluation Methods
Albert Gatt, Anya Belz
INLG2
2008 The TUNA Challenge 2008: Overview and Evaluation Results
Albert Gatt, Anya Belz, Eric Kow
INLG2
2008 Automatic generation of weather forecast texts using comprehensive probabilistic generation-space models
abstract
Abstract Two important recent trends in natural language generation are (i) probabilistic techniques and (ii) comprehensive approaches that move away from traditional strictly modular and sequential models. This paper reports experiments in whichpcru– a generation framework that combines probabilistic generation methodology with a comprehensive model of the generation space – was used to semi-automatically create five different versions of a weather forecast generator. The generators were evaluated in terms of output quality, development time and computational efficiency against (i) human forecasters, (ii) a traditional handcrafted pipelinednlgsystem and (iii) ahalogen-style statistical generator. The most striking result is that despite acquiring all decision-making abilities automatically, the bestpcrugenerators produce outputs of high enough quality to be scored more highly by human judges than forecasts written by experts.
Anya Belz
Nat. Lang. Eng.1
2007 Probabilistic Generation of Weather Forecast Texts
Anya Belz
HLT-NAACL1
2006 Comparing Automatic and Human Evaluation of NLG Systems
Anya Belz, Ehud Reiter
EACL1
2006 Introduction to the INLG'06 Special Session on Sharing Data and Comparative Evaluation
Anya Belz, Robert Dale
INLG1
2006 Shared-Task Evaluations in HLT: Lessons for NLG
Anya Belz, Adam Kilgarriff
INLG1
2006 GENEVAL: A Proposal for Shared-task Evaluation in NLG
Ehud Reiter, Anya Belz
INLG2
2002 Learning Grammars for Different Parsing Tasks by Partition Search
Anya Belz
COLING1
2002 PILLS: Multilingual generation of medical information documents with overlapping content
Nadjet Bouayad-Agha, Richard Power, Donia Scott, Anya Belz
LREC4
1998 A Few English Words Can Help Improve Your Russian
Anya Belz
ECAI1