EDBT 2026 Demo / reviewers in the wild / expert
Verena Rieser
dblp:75/5602
· DBLP profile ↗
64ranked-venue papers
14as first author
18since 2021 · last 2025
0000-0001-6117-4395ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 13 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Value Profiles for Encoding Human VariationabstractTaylor Sorensen, Pushkar Mishra, Roma Patel, Michael Henry Tessler, Michiel A. Bakker, Georgina Evans, Iason Gabriel, Noah Goodman, Verena Rieser. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Taylor Sorensen, Pushkar Mishra, Roma Patel, Michael Henry Tessler, Michiel A. Bakker, Georgina Evans, Iason Gabriel, Noah D. Goodman, Verena Rieser |
EMNLP | 9 |
| 2025 | Century: A Framework and Dataset for Evaluating Historical Contextualisation of Sensitive ImagesabstractHow do multi-modal generative models describe images of recent historical events and figures, whose legacies may be nuanced, multifaceted, or contested? This task necessitates not only accurate visual recognition, but also socio-cultural knowledge and cross-modal reasoning. To address this evaluation challenge, we introduce Century -- a novel dataset of sensitive historical images. This dataset consists of 1,500 images from recent history, created through an automated method combining knowledge graphs and language models with quality and diversity criteria created from the practices of museums and digital archives. We demonstrate through automated and human evaluation that this method produces a set of images that depict events and figures that are diverse across topics and represents all regions of the world.
We additionally propose an evaluation framework for evaluating the historical contextualisation capabilities along dimensions of accuracy, thoroughness, and objectivity. We demonstrate this approach by using Century to evaluate four foundation models, scoring performance using both automated and human evaluation. We find that historical contextualisation of sensitive images poses a significant challenge for modern multi-modal foundation models, and offer practical recommendations for how developers can use Century to evaluate improvements to models and applications. Canfer Akbulut, Kevin Robinson, Maribeth Rauh, Isabela Albuquerque, Olivia Wiles, Laura Weidinger, Verena Rieser, Yana Hasson, Nahema Marchal, Iason Gabriel, William Isaac 0001, Lisa Anne Hendricks |
ICLR | 7 |
| 2025 | Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image ModelsabstractCurrent text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralism in AI alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions to achieve this in T2I models. First, we introduce a novel dataset for Diverse Intersectional Visual Evaluation (DIVE) -- the first multimodal dataset for pluralistic alignment. It enables deep alignment to diverse safety perspectives through a large pool of demographically intersectional human raters who provided extensive feedback across 1000 prompts, with high replication, capturing nuanced safety perceptions. Second, we empirically confirm demographics as a crucial proxy for diverse viewpoints in this domain, revealing significant, context-dependent differences in harm perception that diverge from conventional evaluations. Finally, we discuss implications for building aligned T2I models, including efficient data collection strategies, LLM judgment capabilities, and model steerability towards diverse perspectives. This research offers foundational tools for more equitable and aligned T2I systems.Content Warning: The paper includes sensitive content that may be harmful. Charvi Rastogi, Tian Huey Teh, Pushkar Mishra, Roma Patel, Ding Wang 0006, Mark Diaz, Alicia Parrish, Aida Mostafazadeh Davani, Zoe Ashwood, Michela Paganini, Vinodkumar Prabhakaran, Verena Rieser, Lora Aroyo |
NeurIPS | 12 |
| 2024 | All Too Human? Mapping and Mitigating the Risk from Anthropomorphic AIabstractThe development of highly-capable conversational agents, underwritten by large language models, has the potential to shape user interaction with this technology in profound ways, particularly when the technology is anthropomorphic, or appears human-like. Although the effects of anthropomorphic AI are often benign, anthropomorphic design features also create new kinds of risk. For example, users may form emotional connections to human-like AI, creating the risk of infringing on user privacy and autonomy through over-reliance. To better understand the possible pitfalls of anthropomorphic AI systems, we make two contributions: first, we explore anthropomorphic features that have been embedded in interactive systems in the past, and leverage this precedent to highlight the current implications of anthropomorphic design. Second, we propose research directions for informing the ethical design of anthropomorphic AI. In advancing the responsible development of AI, we promote approaches to the ethical foresight, evaluation, and mitigation of harms arising from user interactions with anthropomorphic AI. Canfer Akbulut, Laura Weidinger, Arianna Manzini, Iason Gabriel, Verena Rieser |
AIES (1) | 5 |
| 2024 | Beyond Thumbs Up/Down: Untangling Challenges of Fine-Grained Feedback for Text-to-Image GenerationabstractHuman feedback plays a critical role in learning and refining reward models for text-to-image generation, but the optimal form the feedback should take for learning an accurate reward function has not been conclusively established. This paper investigates the effectiveness of fine-grained feedback which captures nuanced distinctions in image quality and prompt-alignment, compared to traditional coarse-grained feedback (for example, thumbs up/down or ranking between a set of options). While fine-grained feedback holds promise, particularly for systems catering to diverse societal preferences, we show that demonstrating its superiority to coarse-grained feedback is not automatic. Through experiments on real and synthetic preference data, we surface the complexities of building effective models due to the interplay of model choice, feedback type, and the alignment between human judgment and computational interpretation. We identify key challenges in eliciting and utilizing fine-grained feedback, prompting a reassessment of its assumed benefits and practicality. Our findings -- e.g., that fine-grained feedback can lead to worse models for a fixed budget, in some settings; however, in controlled settings with known attributes, fine grained rewards can indeed be more helpful -- call for careful consideration of feedback attributes and potentially beckon novel modeling approaches to appropriately unlock the potential value of fine-grained feedback in-the-wild. Katie Collins, Najoung Kim, Yonatan Bitton, Verena Rieser, Shayegan Omidshafiei, Yushi Hu, Sherol Chen, Senjuti Dutta, Minsuk Chang, Kimin Lee, Youwei Liang, Georgina Evans, Sahil Singla 0005, Gang Li 0021, Adrian Weller, Junfeng He, Deepak Ramachandran, Krishnamurthy Dvijotham |
AIES (1) | 4 |
| 2024 | Gaps in the Safety Evaluation of Generative AIabstractGenerative AI systems produce a range of ethical and social risks. Evaluation of these risks is a critical step on the path to ensuring the safety of these systems. However, evaluation requires the availability of validated and established measurement approaches and tools. In this paper, we provide an empirical review of the methods and tools that are available for evaluating known safety of generative AI systems to date. To this end, we review more than 200 safety-related evaluations that have been applied to generative AI systems. We categorise each evaluation along multiple axes to create a detailed snapshot of the safety evaluation landscape to date. We release this data for researchers and AI safety practitioners (https://bitly.ws/3hUzu). Analysing the current safety evaluation landscape reveals three systemic ”evaluation gaps”. First, a ”modality gap” emerges as few safety evaluations exist for non-text modalities. Second, a ”risk coverage gap” arises as evaluations for several ethical and social risks are simply lacking. Third, a ”context gap” arises as most safety evaluations are model-centric and fail to take into account the broader context in which AI systems operate. Devising next steps for safety practitioners based on these findings, we present tactical ”low-hanging fruit” steps towards closing the identified evaluation gaps and their limitations. We close by discussing the role and limitations of safety evaluation to ensure the safety of generative AI systems. Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Ramona Comanescu, Canfer Akbulut, Thomas S. Stepleton, Juan Mateos-Garcia, A. Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, William Isaac 0001, Laura Weidinger |
AIES (1) | 14 |
| 2024 | STAR: SocioTechnical Approach to Red Teaming Language ModelsabstractLaura Weidinger, John F J Mellor, Bernat Guillén Pegueroles, Nahema Marchal, Ravin Kumar, Kristian Lum, Canfer Akbulut, Mark Diaz, A. Stevie Bergman, Mikel D. Rodriguez, Verena Rieser, William Isaac. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Laura Weidinger, John Mellor, Bernat Guillen Pegueroles, Nahema Marchal, Kristian Lum, Canfer Akbulut, Mark Diaz, A. Stevie Bergman, Mikel Rodriguez, Verena Rieser, William Isaac 0001 |
EMNLP | 11 |
| 2023 | Mirages. On Anthropomorphism in Dialogue SystemsabstractAutomated dialogue or conversational systems are anthropomorphised by developers and personified by users.While a degree of anthropomorphism may be inevitable due to the choice of medium, conscious and unconscious design choices can guide users to personify such systems to varying degrees.Encouraging users to relate to automated systems as if they were human can lead to high risk scenarios caused by over-reliance on their outputs.As a result, natural language processing researchers have investigated the factors that induce personification and develop resources to mitigate such effects.However, these efforts are fragmented, and many aspects of anthropomorphism have yet to be explored.In this paper, we discuss the linguistic factors that contribute to the anthropomorphism of dialogue systems and the harms that can arise, including reinforcing gender stereotypes and notions of acceptable language.We recommend that future efforts towards developing dialogue systems take particular care in their design, development, release, and description; and attend to the many linguistic cues that can elicit personification by users. Gavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser, Zeerak Talat |
EMNLP | 4 |
| 2023 | Multitask Multimodal Prompted Training for Interactive Embodied Task CompletionabstractGeorgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage, Arash Eshghi, Ioannis Konstas, Verena Rieser, Oliver Lemon, Alessandro Suglia. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh 0001, Bhathiya Hemanthage, Arash Eshghi, Ioannis Konstas, Verena Rieser, Oliver Lemon, Alessandro Suglia |
EMNLP | 7 |
| 2023 | Quality-agnostic Image Captioning to Safely Assist People with Vision ImpairmentabstractAutomated image captioning has the potential to be a useful tool for people with vision impairments. Images taken by this user group are often noisy, which leads to incorrect and even unsafe model predictions. In this paper, we propose a quality-agnostic framework to improve the performance and robustness of image captioning models for visually impaired people. We address this problem from three angles: data, model, and evaluation. First, we show how data augmentation techniques for generating synthetic noise can address data sparsity in this domain. Second, we enhance the robustness of the model by expanding a state-of-the-art model to a dual network architecture, using the augmented data and leveraging different consistency losses. Our results demonstrate increased performance, e.g. an absolute improvement of 2.15 on CIDEr, compared to state-of-the-art image captioning networks, as well as increased robustness to noise with up to 3 points improvement on CIDEr in more noisy settings. Finally, we evaluate the prediction reliability using confidence calibration on images with different difficulty / noise levels, showing that our models perform more reliably in safety-critical situations. The improved model is part of an assisted living application, which we develop in partnership with the Royal National Institute of Blind People. Malvina Nikandrou, Jiali Jin, Verena Rieser |
IJCAI | 4 |
| 2023 | FurChat: An Embodied Conversational Agent using LLMs, Combining Open and Closed-Domain Dialogue with Facial ExpressionsabstractNeeraj Cherakara, Finny Varghese, Sheena Shabana, Nivan Nelson, Abhiram Karukayil, Rohith Kulothungan, Mohammed Afil Farhan, Birthe Nesset, Meriam Moujahid, Tanvi Dinkar, Verena Rieser, Oliver Lemon. Proceedings of the 24th Meeting of the Special Interest Group on Discourse and Dialogue. 2023. Neeraj Cherakara, Finny Varghese, Sheena Shabana, Nivan Nelson, Abhiram Karukayil, Rohith Kulothungan, Mohammed Afil Farhan, Birthe Nesset, Meriam Moujahid, Tanvi Dinkar, Verena Rieser, Oliver Lemon |
SIGDIAL | 11 |
| 2022 | SafetyKit: First Aid for Measuring Safety in Open-domain Conversational SystemsabstractEmily Dinan, Gavin Abercrombie, A. Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, Verena Rieser. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Emily Dinan, Gavin Abercrombie, A. Stevie Bergman, Shannon L. Spruit, Dirk Hovy, Y-Lan Boureau, Verena Rieser |
ACL (1) | 7 |
| 2022 | Guiding the Release of Safer E2E Conversational AI through Value Sensitive DesignabstractA. Stevie Bergman, Gavin Abercrombie, Shannon Spruit, Dirk Hovy, Emily Dinan, Y-Lan Boureau, Verena Rieser. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2022. A. Stevie Bergman, Gavin Abercrombie, Shannon L. Spruit, Dirk Hovy, Emily Dinan, Y-Lan Boureau, Verena Rieser |
SIGDIAL | 7 |
| 2022 | Demonstrating EMMA: Embodied MultiModal Agent for Language-guided Action Execution in 3D Simulated EnvironmentsabstractAlessandro Suglia, Bhathiya Hemanthage, Malvina Nikandrou, Georgios Pantazopoulos, Amit Parekh, Arash Eshghi, Claudio Greco, Ioannis Konstas, Oliver Lemon, Verena Rieser. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2022. Alessandro Suglia, Bhathiya Hemanthage, Malvina Nikandrou, Georgios Pantazopoulos, Amit Parekh 0001, Arash Eshghi, Claudio Greco 0002, Ioannis Konstas, Oliver Lemon, Verena Rieser |
SIGDIAL | 10 |
| 2021 | OTTers: One-turn Topic Transitions for Open-Domain DialogueabstractKarin Sevegnani, David M. Howcroft, Ioannis Konstas, Verena Rieser. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Karin Sevegnani, David M. Howcroft, Ioannis Konstas, Verena Rieser |
ACL/IJCNLP (1) | 4 |
| 2021 | AggGen: Ordering and Aggregating while GeneratingabstractXinnuo Xu, Ondřej Dušek, Verena Rieser, Ioannis Konstas. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xinnuo Xu, Ondrej Dusek, Verena Rieser, Ioannis Konstas |
ACL/IJCNLP (1) | 3 |
| 2021 | ConvAbuse: Data, Analysis, and Benchmarks for Nuanced Detection in Conversational AIabstractWe present the first English corpus study on abusive language towards three conversational AI systems gathered 'in the wild': an opendomain social bot, a rule-based chatbot, and a task-based system.To account for the complexity of the task, we take a more 'nuanced' approach where our ConvAI dataset reflects fine-grained notions of abuse, as well as views from multiple expert annotators.We find that the distribution of abuse is vastly different compared to other commonly used datasets, with more sexually tinted aggression towards the virtual persona of these systems.Finally, we report results from bench-marking existing models against this data.Unsurprisingly, we find that there is substantial room for improvement with F1 scores below 90%.Warning: This paper contains examples of language that some people may find offensive or upsetting. Amanda Cercas Curry, Gavin Abercrombie, Verena Rieser |
EMNLP (1) | 3 |
| 2021 | What happens if you treat ordinal ratings as interval data? Human evaluations in NLP are even more under-powered than you thinkabstractPrevious work has shown that human evaluations in NLP are notoriously under-powered.Here, we argue that there are two common factors which make this problem even worse: NLP studies usually (a) treat ordinal data as interval data and (b) operate under high variance settings while the differences they are hoping to detect are often subtle.We demonstrate through simulation that ordinal mixed effects models are better able to detect small differences between models, especially in high variance settings common in evaluations of generated texts.We release tools for researchers to conduct their own power analysis and test their assumptions.We also make recommendations for improving statistical power. David M. Howcroft, Verena Rieser |
EMNLP (1) | 2 |
| 2020 | History for Visual Dialog: Do we really need it?abstractVisual Dialog involves "understanding" the dialog history (what has been discussed previously) and the current question (what is asked), in addition to grounding information in the image, to generate the correct response.In this paper, we show that co-attention models which explicitly encode dialog history outperform models that don't, achieving state-ofthe-art performance (72 % NDCG on val set).However, we also expose shortcomings of the crowd-sourcing dataset collection procedure by showing that history is indeed only required for a small amount of the data and that the current evaluation metric encourages generic replies.To that end, we propose a challenging subset (VisDialConv) of the VisDial val set and provide a benchmark of 63% NDCG. Shubham Agarwal 0001, Trung Bui, Joon-Young Lee, Ioannis Konstas, Verena Rieser |
ACL | 5 |
| 2020 | Fact-based Content Weighting for Evaluating Abstractive Summarisationabstractive summarisation is notoriously hard to evaluate since standard word-overlap-based metrics are insufficient. We introduce a new evaluation metric which is based on fact-level content weighting, i.e. relating the facts of the document to the facts of the summary. We fol- low the assumption that a good summary will reflect all relevant facts, i.e. the ones present in the ground truth (human-generated refer- ence summary). We confirm this hypothe- sis by showing that our weightings are highly correlated to human perception and compare favourably to the recent manual highlight- based metric of Hardy et al. (2019). Xinnuo Xu, Ondrej Dusek, Verena Rieser, Ioannis Konstas |
ACL | 4 |
| 2020 | SLURP: A Spoken Language Understanding Resource PackageabstractSpoken Language Understanding infers semantic meaning directly from audio data, and thus promises to reduce error propagation and misunderstandings in end-user applications.However, publicly available SLU resources are limited.In this paper, we release SLURP, a new SLU package containing the following: (1) A new challenging dataset in English spanning 18 domains, which is substantially bigger and linguistically more diverse than existing datasets; (2) Competitive baselines based on state-of-the-art NLU and ASR systems; (3) A new transparent metric for entity labelling which enables a detailed error analysis for identifying potential areas of improvement.SLURP is available at Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, Verena Rieser |
EMNLP (1) | 4 |
| 2020 | Twenty Years of Confusion in Human Evaluation: NLG Needs Evaluation Sheets and Standardised DefinitionsabstractDavid M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, Verena Rieser. Proceedings of the 13th International Conference on Natural Language Generation. 2020. David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, Verena Rieser |
INLG | 10 |
| 2020 | Evaluating the state-of-the-art of End-to-End Natural Language Generation: The E2E NLG challengeabstractThis paper provides a comprehensive analysis of the first shared task on End-to-End Natural Language Generation (NLG) and identifies avenues for future research based on the results. This shared task aimed to assess whether recent end-to-end NLG systems can generate more complex output by learning from datasets containing higher lexical richness, syntactic complexity and diverse discourse phenomena. Introducing novel automatic and human metrics, we compare 62 systems submitted by 17 institutions, covering a wide range of approaches, including machine learning architectures – with the majority implementing sequence-to-sequence models (seq2seq) – as well as systems based on grammatical rules and templates. Seq2seq-based systems have demonstrated a great potential for NLG in the challenge. We find that seq2seq systems generally score high in terms of word-overlap metrics and human evaluations of naturalness – with the winning Slug system (Juraska et al., 2018) being seq2seq-based. However, vanilla seq2seq models often fail to correctly express a given meaning representation if they lack a strong semantic control mechanism applied during decoding. Moreover, seq2seq models can be outperformed by hand-engineered systems in terms of overall quality, as well as complexity, length and diversity of outputs. This research has influenced, inspired and motivated a number of recent studies outwith the original competition, which we also summarise as part of this paper. Ondrej Dusek, Jekaterina Novikova, Verena Rieser |
Comput. Speech Lang. | 3 |
| 2019 | Semantic Noise Matters for Neural Natural Language GenerationabstractNeural natural language generation (NNLG) systems are known for their pathological outputs, i.e. generating text which is unrelated to the input specification.In this paper, we show the impact of semantic noise on state-of-theart NNLG models which implement different semantic control mechanisms.We find that cleaned data can improve semantic correctness by up to 97%, while maintaining fluency.We also find that the most common error is omitting information, rather than hallucination. Ondrej Dusek, David M. Howcroft, Verena Rieser |
INLG | 3 |
| 2019 | Automatic Quality Estimation for Natural Language Generation: Ranting (Jointly Rating and Ranking)abstractWe present a recurrent neural network based system for automatic quality estimation of natural language generation (NLG) outputs, which jointly learns to assign numerical ratings to individual outputs and to provide pairwise rankings of two different outputs.The latter is trained using pairwise hinge loss over scores from two copies of the rating network.We use learning to rank and synthetic data to improve the quality of ratings assigned by our system: we synthesise training pairs of distorted system outputs and train the system to rank the less distorted one higher.This leads to a 12% increase in correlation with human ratings over the previous benchmark.We also establish the state of the art on the dataset of relative rankings from the E2E NLG Challenge (Dušek et al., 2019), where synthetic data lead to a 4% accuracy increase over the base model. Ondrej Dusek, Karin Sevegnani, Ioannis Konstas, Verena Rieser |
INLG | 4 |
| 2019 | Let's Chat! Can Virtual Agents learn how to have a Conversation?abstractIntelligent virtual agents frequently engage the user in conversation. The underlying technology - often referred to as spoken dialogue systems - have experienced a revolution over the past decade, moving from being completely handcrafted to using data-driven machine learning methods. Verena Rieser |
IVA | 1 |
| 2019 | A Crowd-based Evaluation of Abuse Response Strategies in Conversational AgentsabstractHow should conversational agents respond to verbal abuse through the user?To answer this question, we conduct a large-scale crowdsourced evaluation of abuse response strategies employed by current state-of-the-art systems.Our results show that some strategies, such as "polite refusal" score highly across the board, while for other strategies demographic factors, such as age, as well as the severity of the preceding abuse influence the user's perception of which response is appropriate.In addition, we find that most data-driven models lag behind rule-based or commercial systems in terms of their perceived appropriateness. Amanda Cercas Curry, Verena Rieser |
SIGdial | 2 |
| 2019 | User Evaluation of a Multi-dimensional Statistical Dialogue SystemabstractWe present the first complete spoken dialogue system driven by a multi-dimensional statistical dialogue manager.This framework has been shown to substantially reduce data needs by leveraging domain-independent dimensions, such as social obligations or feedback, which (as we show) can be transferred between domains.In this paper, we conduct a user study and show that the performance of a multi-dimensional system, which can be adapted from a source domain, is equivalent to that of a one-dimensional baseline, which can only be trained from scratch. Simon Keizer, Ondrej Dusek, Xingkun Liu, Verena Rieser |
SIGdial | 4 |
| 2018 | Better Conversations by Modeling, Filtering, and Optimizing for Coherence and DiversityabstractWe present three enhancements to existing encoder-decoder models for open-domain conversational agents, aimed at effectively modeling coherence and promoting output diversity: (1) We introduce a measure of coherence as the GloVe embedding similarity between the dialogue context and the generated response, (2) we filter our training corpora based on the measure of coherence to obtain topically coherent and lexically diverse context-response pairs, (3) we then train a response generator using a conditional variational autoencoder model that incorporates the measure of coherence as a latent variable and uses a context gate to guarantee topical consistency with the context and promote lexical diversity.Experiments on the OpenSubtitles corpus show a substantial improvement over competitive neural models in terms of BLEU score as well as metrics of coherence and diversity. Xinnuo Xu, Ondrej Dusek, Ioannis Konstas, Verena Rieser |
EMNLP | 4 |
| 2018 | Improving Context Modelling in Multimodal Dialogue GenerationabstractIn this work, we investigate the task of textual response generation in a multimodal task-oriented dialogue system.Our work is based on the recently released Multimodal Dialogue (MMD) dataset (Saha et al., 2017) in the fashion domain.We introduce a multimodal extension to the Hierarchical Recurrent Encoder-Decoder (HRED) model and show that this extension outperforms strong baselines in terms of text-based similarity metrics.We also showcase the shortcomings of current vision and language models by performing an error analysis on our system's output. Shubham Agarwal 0001, Ondrej Dusek, Ioannis Konstas, Verena Rieser |
INLG | 4 |
| 2018 | Findings of the E2E NLG ChallengeabstractThis paper summarises the experimental setup and results of the first shared task on end-to-end (E2E) natural language generation (NLG) in spoken dialogue systems.Recent end-to-end generation systems are promising since they reduce the need for data annotation.However, they are currently limited to small, delexicalised datasets.The E2E NLG shared task aims to assess whether these novel approaches can generate better-quality output by learning from a dataset containing higher lexical richness, syntactic complexity and diverse discourse phenomena.We compare 62 systems submitted by 17 institutions, covering a wide range of approaches, including machine learning architectures -with the majority implementing sequence-to-sequence models (seq2seq) -as well as systems based on grammatical rules and templates. Ondrej Dusek, Jekaterina Novikova, Verena Rieser |
INLG | 3 |
| 2017 | Why We Need New Evaluation Metrics for NLGabstractThe majority of NLG evaluation relies on automatic metrics, such as BLEU.In this paper, we motivate the need for novel, system-and data-independent automatic evaluation methods: We investigate a wide range of metrics, including state-of-the-art word-based and novel grammar-based ones, and demonstrate that they only weakly reflect human judgements of system outputs as generated by data-driven, end-to-end NLG.We also show that metric performance is data-and system-specific.Nevertheless, our results also suggest that automatic metrics perform reliably at system-level and can support system development by finding cases where a system performs poorly.4 https://github.com/glampouras/JLOLS_NLG 5 Note that we use lexicalised versions of SFHOTEL and SFREST and a partially lexicalised version of BAGEL, where proper names and place names are replaced by placeholders ("X"), in correspondence with the outputs generated by the MR: inform(name=X, area=X, pricerange=moderate, type=restaurant) Reference: "X is a moderately priced restaurant in X." Jekaterina Novikova, Ondrej Dusek, Amanda Cercas Curry, Verena Rieser |
EMNLP | 4 |
| 2017 | The E2E Dataset: New Challenges For End-to-End GenerationabstractThis paper describes the E2E data, a new dataset for training end-to-end, datadriven natural language generation systems in the restaurant domain, which is ten times bigger than existing, frequently used datasets in this area.The E2E dataset poses new challenges: (1) its human reference texts show more lexical richness and syntactic variation, including discourse phenomena; (2) generating from this set requires content selection.As such, learning from this dataset promises more natural, varied and less template-like system utterances.We also establish a baseline on this dataset, which illustrates some of the difficulties associated with this data. Jekaterina Novikova, Ondrej Dusek, Verena Rieser |
SIGDIAL Conference | 3 |
| 2016 | How to talk to strangers: Generating medical reports for first-time usersabstractWe propose a novel approach for handling first-time users in the context of automatic report generation from time-series data in the health domain. Handling first-time users is a common problem for Natural Language Generation (NLG) and interactive systems in general - the system cannot adapt to users without prior interaction or user knowledge. In this paper, we propose a novel framework for generating medical reports for first-time users, using multi-objective optimisation (MOO) to account for the preferences of multiple possible user types, where the content preferences of potential users are modelled as objective functions. Our proposed approach outperforms two meaningful baselines in an evaluation with prospective users, yielding large (= .79) and medium (= .46) effect sizes respectively. Dimitra Gkatzia, Verena Rieser, Oliver Lemon |
FUZZ-IEEE | 2 |
| 2016 | Crowd-sourcing NLG Data: Pictures Elicit Better DataabstractRecent advances in corpus-based Natural Language Generation (NLG) hold the promise of being easily portable across domains, but require costly training data, consisting of meaning representations (MRs) paired with Natural Language (NL) utterances.In this work, we propose a novel framework for crowdsourcing high quality NLG training data, using automatic quality control measures and evaluating different MRs with which to elicit data.We show that pictorial MRs result in better NL data being collected than logicbased MRs: utterances elicited by pictorial MRs are judged as significantly more natural, more informative, and better phrased, with a significant increase in average quality ratings (around 0.5 points on a 6-point scale), compared to using the logical MRs.As the MR becomes more complex, the benefits of pictorial stimuli increase.The collected data will be released as part of this submission. Jekaterina Novikova, Oliver Lemon, Verena Rieser |
INLG | 3 |
| 2016 | The aNALoGuE Challenge: Non Aligned Language GEneration
Jekaterina Novikova, Verena Rieser |
INLG | 2 |
| 2016 | The REAL Corpus: A Crowd-Sourced Corpus of Human Generated and Evaluated Spatial References to Real-World Urban Scenes
Phil J. Bartie, William A. Mackaness, Dimitra Gkatzia, Verena Rieser |
LREC | 4 |
| 2016 | Information density and overlap in spoken dialogue
Nina Dethlefs, Helen Hastie, Heriberto Cuayáhuitl, Yanchao Yu, Verena Rieser, Oliver Lemon |
Comput. Speech Lang. | 5 |
| 2015 | From the Virtual to the RealWorld: Referring to Objects in Real-World Spatial ScenesabstractPredicting the success of referring expressions (RE) is vital for real-world applications such as navigation systems.Traditionally, research has focused on studying Referring Expression Generation (REG) in virtual, controlled environments.In this paper, we describe a novel study of spatial references from real scenes rather than virtual.First, we investigate how humans describe objects in open, uncontrolled scenarios and compare our findings to those reported in virtual environments.We show that REs in real-world scenarios differ significantly to those in virtual worlds.Second, we propose a novel approach to quantifying image complexity when complete annotations are not present (e.g.due to poor object recognition capabitlities), and third, we present a model for success prediction of REs for objects in real scenes.Finally, we discuss implications for Natural Language Generation (NLG) systems and future directions. Dimitra Gkatzia, Verena Rieser, Phil J. Bartie, William A. Mackaness |
EMNLP | 2 |
| 2015 | Agent-based Modelling for Green Space Allocation in Urban Areas - Factors Influencing Agent Behaviour
Marta Vallejo, Verena Rieser, David W. Corne |
ICAART (1) | 2 |
| 2015 | Benchmarking Machine Translated Sentiment Analysis for Arabic TweetsabstractTraditional approaches to Sentiment Analysis (SA) rely on large annotated data sets or wide-coverage sentiment lexica, and as such often perform poorly on under-resourced languages.This paper presents empirical evidence of an efficient SA approach using freely available machine translation (MT) systems to translate Arabic tweets to English, which we then label for sentiment using a state-of-theart English SA system.We show that this approach significantly outperforms a number of standard approaches on a gold-standard heldout data set, and performs equally well compared to more cost-intense methods with 76% accuracy.This confirms MT-based SA as a cheap and effective alternative to building a fully fledged SA system when dealing with under-resourced languages. Eshrag Refaee, Verena Rieser |
HLT-NAACL | 2 |
| 2014 | Cluster-based Prediction of User Ratings for Stylistic Surface RealisationabstractSurface realisations typically depend on their target style and audience.A challenge in estimating a stylistic realiser from data is that humans vary significantly in their subjective perceptions of linguistic forms and styles, leading to almost no correlation between ratings of the same utterance.We address this problem in two steps.First, we estimate a mapping function between the linguistic features of a corpus of utterances and their human style ratings.Users are partitioned into clusters based on the similarity of their ratings, so that ratings for new utterances can be estimated, even for new, unknown users.In a second step, the estimated model is used to re-rank the outputs of a number of surface realisers to produce stylistically adaptive output.Results confirm that the generated styles are recognisable to human judges and that predictive models based on clusters of users lead to better rating predictions than models based on an average population of users. Nina Dethlefs, Heriberto Cuayáhuitl, Helen Hastie, Verena Rieser, Oliver Lemon |
EACL | 4 |
| 2014 | An Arabic Twitter Corpus for Subjectivity and Sentiment Analysis
Eshrag Refaee, Verena Rieser |
LREC | 2 |
| 2014 | The PARLANCE mobile application for interactive search in English and MandarinabstractHelen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Hugues Bouchard, Catherine Breslin, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gašić, James Henderson, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Tim Potter, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay, Boris Villazon-Terrazas, Majid Yazdani, Steve Young, Yanchao Yu. Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL). 2014. Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Hugues Bouchard, Catherine Breslin, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gasic, James Henderson 0001, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Tim Potter, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay, Boris Villazón-Terrazas, Majid Yazdani, Steve J. Young, Yanchao Yu |
SIGDIAL Conference | 15 |
| 2014 | Natural Language Generation as Incremental Planning Under Uncertainty: Adaptive Information Presentation for Statistical Dialogue SystemsabstractWe present and evaluate a novel approach to natural language generation (NLG) in statistical spoken dialogue systems (SDS) using a data-driven statistical optimization framework for incremental information presentation (IP), where there is a trade-off to be solved between presenting “enough" information to the user while keeping the utterances short and understandable. The trained IP model is adaptive to variation from the current generation context (e.g. a user and a non-deterministic sentence planner), and it incrementally adapts the IP policy at the turn level. Reinforcement learning is used to automatically optimize the IP policy with respect to a data-driven objective function. In a case study on presenting restaurant information, we show that an optimized IP strategy trained on Wizard-of-Oz data outperforms a baseline mimicking the wizard behavior in terms of total reward gained. The policy is then also tested with real users, and improves on a conventional hand-coded IP strategy used in a deployed SDS in terms of overall task success. The evaluation found that the trained IP strategy significantly improves dialogue task completion for real users, with up to a 8.2% increase in task success. This methodology also provides new insights into the nature of the IP problem, which has previously been treated as a module following dialogue management with no access to lower-level context features (e.g. from a surface realizer and/or speech synthesizer). Verena Rieser, Oliver Lemon, Simon Keizer |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2013 | Evolving Urbanisation Policies - Using a Statistical Model to Accelerate Optimisation over Agent-based Simulations
Marta Vallejo, David W. Corne, Verena Rieser |
ICAART (2) | 3 |
| 2013 | Demonstration of the PARLANCE system: a data-driven incremental, spoken dialogue system for interactive search
Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gasic, James Henderson 0001, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay |
SIGDIAL Conference | 12 |
| 2012 | Optimising Incremental Dialogue Decisions Using Information Density for Interactive Systems
Nina Dethlefs, Helen Hastie, Verena Rieser, Oliver Lemon |
EMNLP-CoNLL | 3 |
| 2012 | Optimising Incremental Generation for Spoken Dialogue Systems: Reducing the Need for Fillers
Nina Dethlefs, Helen Hastie, Verena Rieser, Oliver Lemon |
INLG | 3 |
| 2011 | Learning and Evaluation of Dialogue Strategies for New Applications: Empirical Methods for Optimization from Small Data SetsabstractWe present a new data-driven methodology for simulation-based dialogue strategy learning, which allows us to address several problems in the field of automatic optimization of dialogue strategies: learning effective dialogue strategies when no initial data or system exists, and determining a data-driven reward function. In addition, we evaluate the result with real users, and explore how results transfer between simulated and real interactions. We use Reinforcement Learning (RL) to learn multimodal dialogue strategies by interaction with a simulated environment which is “bootstrapped” from small amounts of Wizard-of-Oz (WOZ) data. This use of WOZ data allows data-driven development of optimal strategies for domains where no working prototype is available. Using simulation-based RL allows us to find optimal policies which are not (necessarily) present in the original data. Our results show that simulation-based RL significantly outperforms the average (human wizard) strategy as learned from the data by using Supervised Learning. The bootstrapped RL-based policy gains on average 50 times more reward when tested in simulation, and almost 18 times more reward when interacting with real users. Users also subjectively rate the RL-based policy on average 10% higher. We also show that results from simulated interaction do transfer to interaction with real users, and we explicitly evaluate the stability of the data-driven reward function. Verena Rieser, Oliver Lemon |
Comput. Linguistics | 1 |
| 2010 | Optimising Information Presentation for Spoken Dialogue Systems
Verena Rieser, Oliver Lemon, Xingkun Liu |
ACL | 1 |
| 2010 | Generation Under Uncertainty
Oliver Lemon, Srinivasan Janarthanam, Verena Rieser |
INLG | 3 |
| 2010 | Learning human multimodal dialogue strategiesabstractAbstract We investigate the use of different machine learning methods in combination with feature selection techniques to explore human multimodal dialogue strategies and the use of those strategies for automated dialogue systems. We learn policies from data collected in a Wizard-of-Oz study where different human ‘wizards’ decide whether to ask a clarification request in a multimodal manner or else to use speech alone. We first describe the data collection, the coding scheme and annotated corpus, and the validation of the multimodal annotations. We then show that there is a uniform multimodal dialogue strategy across wizards, which is based on multiple features in the dialogue context. These are generic features, available at runtime, which can be implemented in dialogue systems. Our prediction models (for human wizard behaviour) achieve a weighted f-score of 88.6 per cent (which is a 25.6 per cent improvement over the majority baseline). We interpret and discuss the learned strategy. We conclude that human wizard behaviour is not optimal for automatic dialogue systems, and argue for the use of automatic optimization methods, such as Reinforcement Learning. Throughout the investigation we also discuss the issues arising from using small initial Wizard-of-Oz data sets, and we show that feature engineering is an essential step when learning dialogue strategies from such limited data. Verena Rieser, Oliver Lemon |
Nat. Lang. Eng. | 1 |
| 2009 | Natural Language Generation as Planning Under Uncertainty for Spoken Dialogue Systems
Verena Rieser, Oliver Lemon |
EACL | 1 |
| 2009 | Predicting how it sounds: re-ranking dialogue prompts based on TTS quality for adaptive spoken dialogue systems
Cédric Boidin, Verena Rieser, Lonneke van der Plas, Oliver Lemon, Jonathan Chevelu |
INTERSPEECH | 2 |
| 2009 | Does this list contain what you were searching for? Learning adaptive dialogue strategies for interactive question answeringabstractAbstract Policy learning is an active topic in dialogue systems research, but it has not been explored in relation to interactive question answering (IQA). We take a first step in learning adaptive interaction policies for question answering : we address the question of how to acquire enough reliable query constraints, how many database results to present to the user and when to present them, given the competing trade-offs between the length of the answer list, the length of the interaction, the type of database and the noise in the communication channel. The operating conditions are reflected in an objective function which we use to derive a hand-coded threshold-based policy and rewards to train a reinforcement learning policy. The same objective function is used for evaluation. We show that we can learn strategies for this complex trade-off problem which perform significantly better than a variety of hand-coded policies, for a wide range of noise conditions, user types, types of DB and turn-penalties. Our policy learning framework thus covers a wide spectrum of operating conditions. The learned policies produce an averagerelativeincrease in reward of 86.78% over the hand-coded policies. In 93% of the cases the learned policies perform significantly better than the hand-coded ones (p< .001). Furthermore we show that the type of database has a significant effect on learning and we give qualitative descriptions of the learned IQA policies. Verena Rieser, Oliver Lemon |
Nat. Lang. Eng. | 1 |
| 2008 | Learning Effective Multimodal Dialogue Strategies from Wizard-of-Oz Data: Bootstrapping and Evaluation
Verena Rieser, Oliver Lemon |
ACL | 1 |
| 2008 | Automatic Learning and Evaluation of User-Centered Objective Functions for Dialogue System Optimisation
Verena Rieser, Oliver Lemon |
LREC | 1 |
| 2007 | Learning dialogue strategies for interactive database search
Verena Rieser, Oliver Lemon |
INTERSPEECH | 1 |
| 2006 | Using Machine Learning to Explore Human Multimodal Clarification Strategies
Verena Rieser, Oliver Lemon |
ACL | 1 |
| 2006 | Cluster-based user simulations for learning dialogue strategiesabstractAbstract Good dialogue strate gies in spok en dialogue systems help to en-sure and maintain mutual understanding and thus play a crucialrole in rob ust con versational interaction. W e focus on clariÞca-tion strate gies and build user simulations which are critical forreinforcement learning, which is a cheap and principled w ay toautomatically optimise dialogue management. In this paper wepresent a no vel cluster -based technique for building user simula-tions which sho w varying , but complete and consistent beha viourwith respect to real users. W e use this technique to build usersimulations and we also introduce the S U P E R evaluation metricwhich allo ws us to evaluate user simulations with respect to thesedesiderata. W e sho w that the cluster -based user simulation tech-nique performs signiÞcantly better (at P < 0.01 ) than decisionsmade using either the one most likely action or a random base-line. The cluster -based user simulations reduce the average errorof these other models by 53% and 34% respecti vely .Index T erms : spok en dialogue, user simulation, evaluation met-rics, reinforcement learning, dialogue strate gies Verena Rieser, Oliver Lemon |
INTERSPEECH | 1 |
| 2006 | The SAMMIE Corpus of Multimodal Dialogues with an MP3 Player
Ivana Kruijff-Korbayová, Tilman Becker, Nate Blaylock, Ciprian Gerstenberger, Michael Kaißer, Peter Poller, Verena Rieser, Jan Schehl |
LREC | 7 |
| 2006 | Using logistic Regression to Initialise Reinforcement-Learning-Based Dialogue SystemsabstractWe investigate the use of logistic regression (LR) to initialise reinforcement learning (RL)-based dialogue systems with models of human dialogue strategies. LR produces accurate predictions and performs feature selection. We illustrate this technique in exploring human multimodal clarification strategies, observed in a Wizard-of-Oz experiment. We use it to initialise an RL-based system with features which significantly influence human behaviour. We show that the strategy applied by the human wizards is sensitive to different dialogue contexts. Furthermore we show that for predicting clarification behaviour the logistic models improve over the baseline on average twice as much as the supervised learning techniques used in previous work. Verena Rieser, Oliver Lemon |
SLT | 1 |
| 2005 | Implications for Generating Clarification Requests in Task-Oriented DialoguesabstractClarification requests (CRs) in conversation ensure and maintain mutual understanding and thus play a crucial role in robust dialogue interaction. In this paper, we describe a corpus study of CRs in task-oriented dialogue and compare our findings to those reported in two prior studies. We find that CR behavior in task-oriented dialogue differs significantly from that in everyday conversation in a number of ways. Moreover, the dialogue type, the modality and the channel quality all influence the decision of when to clarify and at which level of the grounding process. Finally we identify form-function correlations which can inform the generation of CRs. Verena Rieser, Johanna D. Moore |
ACL | 1 |