Oliver Lemon

dblp:36/6352 · DBLP profile ↗
← Back
119ranked-venue papers
11as first author
19since 2021 · last 2025
0000-0001-9497-4743ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 102 · 9 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 3 first-authorHuman-computer interaction and ubiquitous computing · 18 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Systems, architecture and hardware · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Playpen: An Environment for Exploring Learning From Dialogue Game Feedback
abstract
Nicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Momentè, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, Raffaella Bernardi, Raquel Fernández, Alexander Koller, Oliver Lemon, David Schlangen, Mario Giulianelli, Alessandro Suglia. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Nicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Momentè, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, Raffaella Bernardi, Raquel Fernández, Alexander Koller, Oliver Lemon, David Schlangen, Mario Giulianelli, Alessandro Suglia
EMNLP13
2024 Clarifying Completions: Evaluating How LLMs Respond to Incomplete Questions
abstract
People understand and produce language incrementally on a word by word basis. This gives rise to many characteristic conversational phenomena including long mid-sentence pauses that are followed by incremental clarification requests (iCRs) intended to recover the rest of the truncated turn (see Fig. 1; (A), (B), (C)). The ability to generate iCRs is important in natural conversational AI systems, and crucial to their accessibility to users with memory impairment. In this paper, we collect, release and analyse SLUICE-CR: a large corpus of 3000 human produced iCRs. We then use this corpus to probe the incremental processing capability of a number of state of the art LLMs by evaluating the quality of the model’s generated iCRs in response to incomplete questions. Our evaluations show that the ability to generate contextually appropriate iCRs only emerges at larger LLM sizes, and only when prompted with example iCRs from our corpus. They also indicate that autoregressive LMs are, in principle, able to both understand and generate language incrementally.
Angus Addlesee, Oliver Lemon, Arash Eshghi
LREC/COLING2
2024 RECANTFormer: Referring Expression Comprehension with Varying Numbers of Targets
abstract
The Generalized Referring Expression Comprehension (GREC) task extends classic REC by generating image bounding boxes for objects referred to in natural language expressions, which may indicate zero, one, or multiple targets. This generalization enhances the practicality of REC models for diverse real-world applications. However, the presence of varying numbers of targets in samples makes GREC a more complex task, both in terms of training supervision and final prediction selection strategy. Addressing these challenges, we introduce RECANTFormer, a one-stage method for GREC that combines a decoder-free (encoder-only) transformer architecture with DETR-like Hungarian matching. Our approach consistently outperforms baselines by significant margins in three GREC datasets.
Bhathiya Hemanthage, Hakan Bilen, Phil J. Bartie, Christian Dondrup, Oliver Lemon
EMNLP5
2024 Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling
abstract
This study explores replacing Transformers in Visual Language Models (VLMs) with Mamba, a recent structured state space model (SSM) that demonstrates promising performance in sequence modeling.We test models up to 3B parameters under controlled conditions, showing that Mamba-based VLMs outperforms Transformers-based VLMs in captioning, question answering, and reading comprehension.However, we find that Transformers achieve greater performance in visual grounding and the performance gap widens with scale.We explore two hypotheses to explain this phenomenon: 1) the effect of task-agnostic visual encoding on the updates of the hidden states, and 2) the difficulty in performing visual grounding from the perspective of in-context multimodal retrieval.Our results indicate that a task-aware encoding yields minimal performance gains on grounding, however, Transformers significantly outperform Mamba at incontext multimodal retrieval.Overall, Mamba shows promising performance on tasks where the correct output relies on a summary of the image but struggles when retrieval of explicit information from the context is required 1 .
Georgios Pantazopoulos, Malvina Nikandrou, Alessandro Suglia, Oliver Lemon, Arash Eshghi
EMNLP4
2024 A Holistic Evaluation Methodology for Multi-Party Spoken Conversational Agents
abstract
While research in multi-party spoken conversation with intelligent embodied agents has made significant progress in sub-tasks like speaker identification and non-verbal cues, there’s a gap in fully autonomous applications users can directly interact with. This lack translates to the absence of a standard methodology for evaluating multi-party conversational speech agents that considers both task-based system performance and user experience.
Nancie Gunson, Angus Addlesee, Daniel Hernández García, Marta Romeo, Christian Dondrup, Oliver Lemon
IVA6
2024 Divide and Conquer: Rethinking Ambiguous Candidate Identification in Multimodal Dialogues with Pseudo-Labelling
abstract
Ambiguous Candidate Identification (ACI) in multimodal dialogue is the task of identifying all potential objects that a user's utterance could be referring to in a visual scene, in cases where the reference cannot be uniquely determined.End-to-end models are the dominant approach for this task, but have limited real-world applicability due to unrealistic inference-time assumptions such as requiring predefined catalogues of items.Focusing on a more generalized and realistic ACI setup, we demonstrate that a modular approach, which first emphasizes language-only reasoning over dialogue context before performing vision-language fusion, significantly outperforms end-to-end trained baselines.To mitigate the lack of annotations for training the language-only module (student), we propose a pseudo-labelling strategy with a prompted Large Language Model (LLM) as the teacher.
Bhathiya Hemanthage, Christian Dondrup, Hakan Bilen, Oliver Lemon
SIGDIAL4
2024 Visually Grounded Language Learning: A Review of Language Games, Datasets, Tasks, and Models
abstract
In recent years, several machine learning models have been proposed. They are trained with a language modelling objective on large-scale text-only data. With such pretraining, they can achieve impressive results on many Natural Language Understanding and Generation tasks. However, many facets of meaning cannot be learned by “listening to the radio” only. In the literature, many Vision+Language (V+L) tasks have been defined with the aim of creating models that can ground symbols in the visual modality. In this work, we provide a systematic literature review of several tasks and models proposed in the V+L field. We rely on Wittgenstein’s idea of ‘language games’ to categorise such tasks into 3 different families: 1) discriminative games, 2) generative games, and 3) interactive games. Our analysis of the literature provides evidence that future work should be focusing on interactive games where communication in Natural Language is important to resolve ambiguities about object referents and action plans and that physical embodiment is essential to understand the semantics of situations and events. Overall, these represent key requirements for developing grounded meanings in neural models.
Alessandro Suglia, Ioannis Konstas, Oliver Lemon
J. Artif. Intell. Res.3
2023 Multitask Multimodal Prompted Training for Interactive Embodied Task Completion
abstract
Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage, Arash Eshghi, Ioannis Konstas, Verena Rieser, Oliver Lemon, Alessandro Suglia. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh 0001, Bhathiya Hemanthage, Arash Eshghi, Ioannis Konstas, Verena Rieser, Oliver Lemon, Alessandro Suglia
EMNLP8
2023 Multi-party Goal Tracking with LLMs: Comparing Pre-training, Fine-tuning, and Prompt Engineering
abstract
Angus Addlesee, Weronika Sieińska, Nancie Gunson, Daniel Hernandez Garcia, Christian Dondrup, Oliver Lemon. Proceedings of the 24th Meeting of the Special Interest Group on Discourse and Dialogue. 2023.
Angus Addlesee, Weronika Sieinska, Nancie Gunson, Daniel Hernández García, Christian Dondrup, Oliver Lemon
SIGDIAL6
2023 FurChat: An Embodied Conversational Agent using LLMs, Combining Open and Closed-Domain Dialogue with Facial Expressions
abstract
Neeraj Cherakara, Finny Varghese, Sheena Shabana, Nivan Nelson, Abhiram Karukayil, Rohith Kulothungan, Mohammed Afil Farhan, Birthe Nesset, Meriam Moujahid, Tanvi Dinkar, Verena Rieser, Oliver Lemon. Proceedings of the 24th Meeting of the Special Interest Group on Discourse and Dialogue. 2023.
Neeraj Cherakara, Finny Varghese, Sheena Shabana, Nivan Nelson, Abhiram Karukayil, Rohith Kulothungan, Mohammed Afil Farhan, Birthe Nesset, Meriam Moujahid, Tanvi Dinkar, Verena Rieser, Oliver Lemon
SIGDIAL12
2022 Multi-party Interaction with a Robot Receptionist
abstract
We introduce a situated interactive robot receptionist that can coordinate turn-taking and handle multi-party engagement and dialogue in dynamic environments, where users might enter or leave the scene at any time. The objective is to create a multi-user engagement policy to manage turn-taking using the robot's gaze, head pose, and verbal communication as parameters and to analyse the participant's perception of the robot. Participant feedback on the system was collected using an online survey that allowed for a comparison of subjective feedback for 4 different interaction policies. The results confirm the hypothesis that a robot is perceived as more intelligent and conscious when it reacts using eye gaze or head pose, once a new user enters the scene. Furthermore, we find that robots need to use a combination of verbal and non-verbal cues to coordinate turn-taking, in order to be perceived as polite and aware of human social norms.
Meriam Moujahid, Helen Hastie, Oliver Lemon
HRI3
2022 Demonstration of a Robot Receptionist with Multi-party Situated Interaction
abstract
We present a demonstration of a Robot Receptionist: a situated interactive robot that can coordinate turn-taking and handle multi-party engagement and dialogue in dynamic environments, where users might enter or leave the scene at any time. We use a Furhat robot, which is highly expressive and can use verbal communication as well as non-verbal cues, such as facial expressions. The system demonstrated and described here is composed of several modules, including scene analysis, engagement policies, and a dialogue manager.
Meriam Moujahid, Bruce Wilson, Helen Hastie, Oliver Lemon
HRI4
2022 Developing a Social Conversational Robot for the Hospital waiting room
abstract
Possible applications for Social Robots in health-care settings, that could have a tremendous social impact in helping alleviating staff workload, are those of a patient-facing role such as robot receptionist, providing assistance to patients and visitors. Examples of functions that such robots would need to be able to execute are greeting visitors, reception check-in/out of patients, answering common questions they may have, showing them where to sit, helping them locate missing objects, providing directions to facilities, guiding them to different locations, etc. In this paper we describe current progress towards developing a multimodal conversational AI system integrated in a Social Conversational Robot (an ARI robot) that will act as a receptionist in a hospital waiting room. We present the developed architecture of the system and report on an initial experimental validation study carried out in laboratory conditions with the ARI robot.
Nancie Gunson, Daniel Hernández García, Weronika Sieinska, Christian Dondrup, Oliver Lemon
RO-MAN5
2022 A Visually-Aware Conversational Robot Receptionist
abstract
Nancie Gunson, Daniel Hernandez Garcia, Weronika Sieińska, Angus Addlesee, Christian Dondrup, Oliver Lemon, Jose L. Part, Yanchao Yu. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2022.
Nancie Gunson, Daniel Hernández García, Weronika Sieinska, Angus Addlesee, Christian Dondrup, Oliver Lemon, Jose L. Part, Yanchao Yu
SIGDIAL6
2022 Demonstrating EMMA: Embodied MultiModal Agent for Language-guided Action Execution in 3D Simulated Environments
abstract
Alessandro Suglia, Bhathiya Hemanthage, Malvina Nikandrou, Georgios Pantazopoulos, Amit Parekh, Arash Eshghi, Claudio Greco, Ioannis Konstas, Oliver Lemon, Verena Rieser. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2022.
Alessandro Suglia, Bhathiya Hemanthage, Malvina Nikandrou, Georgios Pantazopoulos, Amit Parekh 0001, Arash Eshghi, Claudio Greco 0002, Ioannis Konstas, Oliver Lemon, Verena Rieser
SIGDIAL9
2021 An Empirical Study on the Generalization Power of Neural Representations Learned via Visual Guessing Games
abstract
Alessandro Suglia, Yonatan Bisk, Ioannis Konstas, Antonio Vergari, Emanuele Bastianelli, Andrea Vanzo, Oliver Lemon. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Alessandro Suglia, Yonatan Bisk, Ioannis Konstas, Antonio Vergari, Emanuele Bastianelli, Andrea Vanzo, Oliver Lemon
EACL7
2021 Am I Allergic to This? Assisting Sight Impaired People in the Kitchen
abstract
Sight-Impaired People (SIP) need assistance with food packaging - which can contain safety-critical information. Therefore, Textual Visual Question Answering (VQA) could prove critical in increasing the independence of SIP [26, 35], yet has only seen recent attention. For instance, comprehending text within an image is necessary to determine: what type of soup is in a can, how long to cook a microwave meal, when a box of eggs will expire, and whether a meal contains an ingredient they are allergic to. This handful of examples relate to a kitchen setting — a particularly challenging area for SIP. We extended the existing Aye-saacvoice assistant prototype with this task and setting in mind. We developed textual VQA components to accurately understand what a user is asking, extract relevant text from images in an intelligent manner, and to provide Natural Language answers that build upon the context of previous questions. As our system is created to be assistive, we designed it with a particular focus on privacy, transparency, and controllability. These are vital objectives that existing systems do not cover. We found that our system outperformed other VQA systems on real food packaging questions asked by SIP from the VizWiz dataset [19].
Elisa Ramil Brick, Vanesa Caballero Alonso, Conor O'Brien, Sheron Tong, Emilie Tavernier, Amit Parekh 0001, Angus Addlesee, Oliver Lemon
ICMI8
2021 Combining Visual and Social Dialogue for Human-Robot Interaction
abstract
We will demonstrate a prototype multimodal conversational AI system that will act as a receptionist in a hospital waiting room, combining visually-grounded dialogue with social conversation. The system supports visual object conversation in the waiting room (e.g. looking for available seats or personal belongings), task-based dialogues regarding navigation and check-in procedures in the hospital, as well as access to the latest news, and a quiz game about coronavirus. The prototype system therefore demonstrates how to weave together a wide range of natural, daily conversations with end users that vary in complexity; from complex visual dialogue to chitchat and quiz games, to task-oriented domain-specific conversations. We are currently able to demonstrate the system via a web-based interface. It will soon be deployed on the ARI robot in a hospital waiting room.
Nancie Gunson, Daniel Hernández García, Jose L. Part, Yanchao Yu, Weronika Sieinska, Christian Dondrup, Oliver Lemon
ICMI7
2021 ViCA: Combining visual, Social, and Task-orientedconversational AI in a Healthcare Setting
abstract
Recent developments in computer vision and conversational systems have provided the AI community with novel perspectives towards improving the cognitive capabilities of engaging socially assistive robots. We show how to develop conversational skills for a hospital receptionist robot that incorporates social conversation based on visual information as well as task-based dialog. Fusing the traditional modular conversational system architecture with recent developments in computer vision and scene graph research, our agent (called ‘ViCA’) supports both visual question answering and social conversational capabilities based on the visual scene. In particular, our agent can provide guidance to users by locating visible objects in the room and can engage in social dialog using visual prompts, such as the user’s clothing or possessions. We conduct a comprehensive online evaluation study with 21 participants, showcasing that the ViCA system is perceived as both helpful and entertaining.
Georgios Pantazopoulos, Jeremy Bruyere, Malvina Nikandrou, Thibaud Boissier, Supun Hemanthage, Binha Kumar Sachish, Vidyul Shah, Christian Dondrup, Oliver Lemon
ICMI9
2020 CompGuessWhat?!: A Multi-task Evaluation Framework for Grounded Language Learning
abstract
Alessandro Suglia, Ioannis Konstas, Andrea Vanzo, Emanuele Bastianelli, Desmond Elliott, Stella Frank, Oliver Lemon. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Alessandro Suglia, Ioannis Konstas, Andrea Vanzo, Emanuele Bastianelli, Desmond Elliott, Stella Frank, Oliver Lemon
ACL7
2020 Imagining Grounded Conceptual Representations from Perceptual Information in Situated Guessing Games
abstract
Alessandro Suglia, Antonio Vergari, Ioannis Konstas, Yonatan Bisk, Emanuele Bastianelli, Andrea Vanzo, Oliver Lemon. Proceedings of the 28th International Conference on Computational Linguistics. 2020.
Alessandro Suglia, Antonio Vergari, Ioannis Konstas, Yonatan Bisk, Emanuele Bastianelli, Andrea Vanzo, Oliver Lemon
COLING7
2020 It's Good to Chat?: Evaluation and Design Guidelines for Combining Open-Domain Social Conversation with Task-Based Dialogue in Intelligent Buildings
abstract
We present and evaluate a deployed conversational AI system that acts as a host of a working public building on a university campus. The system combines open-domain social chat with task-based conversation regarding navigation in the building, live resource updates (e.g. available computers), and events in the building. We investigated the impact of open-domain social chat on task completion and user preferences by comparing the combined system with a task-only version. We find that there is no significant difference in task completion or several aspects of user preference between the two systems, but that users would be significantly happier to talk to the task-only system in the future. This suggests that the "walk-up" public setting and workplace nature of the environment creates a markedly different use case to the in-home, and more individual and private "companion/assistant" setting which is commonly assumed for systems like Alexa. We discuss the implications for the design of conversational systems in other public settings.
Nancie Gunson, Weronika Sieinska, Christopher Walsh, Christian Dondrup, Oliver Lemon
IVA5
2020 Conversational Agents for Intelligent Buildings
abstract
We will demonstrate a deployed conversational AI system that acts as a host of a smartbuilding on a university campus.The system combines open-domain social conversation with task-based conversation regarding navigation in the building, live resource updates (e.g.available computers) and events in the building.We are able to demonstrate the system on several platforms: Google Home devices, Android phones, and a Furhat robot.
Weronika Sieinska, Christian Dondrup, Nancie Gunson, Oliver Lemon
SIGdial4
2019 Data-Efficient Goal-Oriented Conversation with Dialogue Knowledge Transfer Networks
abstract
Igor Shalyminov, Sungjin Lee, Arash Eshghi, Oliver Lemon. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Igor Shalyminov, Arash Eshghi, Oliver Lemon
EMNLP/IJCNLP (1)4
2019 Towards a Robot Architecture for Situated Lifelong Object Learning
abstract
The ability to acquire knowledge incrementally and after deployment is of utmost importance for robots operating in the real world. Moreover, robots that have to operate alongside people need to be able to interact in a way that is intuitive for the users, e.g., by understanding and producing natural language. In this paper we present a first prototype of a robot architecture developed for situated lifelong object learning. The system is able to communicate with its users through natural language and perform object learning and recognition on the spot through situated interactions. In this first stage, we evaluate the system in terms of recognition accuracy which gives an indirect measure of the quality of the collected data with the proposed pipeline. Our results show that the robot can use this data for both learning and recognition with acceptable incremental performance. We also discuss limitations and steps that are necessary in order to improve performance as well as to shed some light on system usability.
Jose L. Part, Oliver Lemon
IROS2
2019 Few-Shot Dialogue Generation Without Annotated Data: A Transfer Learning Approach
abstract
Learning with minimal data is one of the key challenges in the development of practical, production-ready goal-oriented dialogue systems.In a real-world enterprise setting where dialogue systems are developed rapidly and are expected to work robustly for an evergrowing variety of domains, products, and scenarios, efficient learning from a limited number of examples becomes indispensable.In this paper, we introduce a technique to achieve state-of-the-art dialogue generation performance in a few-shot setup, without using any annotated data.We do this by leveraging background knowledge from a larger, more highly represented dialogue sourcenamely, the MetaLWOz dataset.We evaluate our model on the Stanford Multi-Domain Dialogue Dataset, consisting of human-human goal-oriented dialogues in in-car navigation, appointment scheduling, and weather information domains.We show that our few-shot approach achieves state-of-the art results on that dataset by consistently outperforming the previous best model in terms of BLEU and Entity F1 scores, while being more data-efficient by not requiring any data annotation.
Igor Shalyminov, Arash Eshghi, Oliver Lemon
SIGdial4
2019 Hierarchical Multi-Task Natural Language Understanding for Cross-domain Conversational AI: HERMIT NLU
abstract
We present a new neural architecture for widecoverage Natural Language Understanding in Spoken Dialogue Systems.We develop a hierarchical multi-task architecture, which delivers a multi-layer representation of sentence meaning (i.e., Dialogue Acts and Frame-like structures).The architecture is a hierarchy of self-attention mechanisms and BiLSTM encoders followed by CRF tagging layers.We describe a variety of experiments, showing that our approach obtains promising results on a dataset annotated with Dialogue Acts and Frame Semantics.Moreover, we demonstrate its applicability to a different, publicly available NLU dataset annotated with domainspecific intents and corresponding semantic roles, providing overall performance higher than state-of-the-art tools such as RASA, Dialogflow, LUIS, and Watson.For example, we show an average 4.45% improvement in entity tagging F-score over Rasa, Dialogflow and LUIS.
Andrea Vanzo, Emanuele Bastianelli, Oliver Lemon
SIGdial3
2018 Spoken Conversational AI in Video Games: Emotional Dialogue Management Increases User Engagement
abstract
In a traditional role-playing game (RPG) conversing with a Non-Playable Character (NPC) typically appears somewhat unrealistic and can break immersion and user engagement. In commercial games, the player usually selects one of several possible predefined conversation options which are displayed as text or labels on the screen, to progress the conversation. In contrast, we first present a spoken conversational interface, built using a state-of-the-art open-domain social conversational AI developed for the Amazon Alexa Challenge, which was modified for use in a video game. This system is designed to keep users engaged in the conversation -- which we measure by time taken speaking with the character. In particular, we use emotion detection and emotional dialogue management to enhance the conversational experience. We then evaluate the contribution of emotion detection and conversational responses in a spoken dialogue system for a role-playing video game. In order to do this, two prototypes of the same game were created: one system using sentiment analysis and emotional modelling and the other system that does not detect or react to emotions. Both systems use a spoken conversational AI system where the user can freely talk to a Non-Playable-Character using unconstrained speech input.
Jamie Fraser, Ioannis Papaioannou, Oliver Lemon
IVA3
2018 Blending Human and Artificial Intelligence to Support Autistic Children's Social Communication Skills
abstract
This article examines the educational efficacy of a learning environment in which children diagnosed with Autism Spectrum Conditions (ASC) engage in social interactions with an artificially intelligent (AI) virtual agent and where a human practitioner acts in support of the interactions. A multi-site intervention study in schools across the UK was conducted with 29 children with ASC and learning difficulties, aged 4--14 years old. For reasons related to data completeness and amount of exposure to the AI environment, data for 15 children was included in the analysis. The analysis revealed a significant increase in the proportion of social responses made by ASC children to human practitioners. The number of initiations made to human practitioners and to the virtual agent by the ASC children also increased numerically over the course of the sessions. However, due to large individual differences within the ASC group, this did not reach significance. Although no evidence of transfer to the real-world post-test was shown, anecdotal evidence of classroom transfer was reported. The work presented in this article offers an important contribution to the growing body of research in the context of AI technology design and use for autism intervention in real school contexts. Specifically, the work highlights key methodological challenges and opportunities in this area by leveraging interdisciplinary insights in a way that (i) bridges between educational interventions and intelligent technology design practices, (ii) considers the design of technology as well as the design of its use (context and procedures) on par with one another, and (iii) includes design contributions from different stakeholders, including children with and without ASC diagnosis, educational practitioners, and researchers.
Kaska Porayska-Pomsta, Alyssa Alcorn, Katerina Avramides, Sandra Beale, Sara Bernardini, Mary Ellen Foster, Christopher Frauenberger, Judith Good, Karen Guldberg, Wendy Keay-Bright, Lila Kossyvaki, Oliver Lemon, Marilena Mademtzi, Rachel Menzies, Helen Pain, Gnanathusharan Rajendran, Annalu Waller, Sam Wass, Tim J. Smith
ACM Trans. Comput. Hum. Interact.12
2017 Bootstrapping incremental dialogue systems from minimal data: the generalisation power of dialogue grammars
abstract
We investigate an end-to-end method for automatically inducing task-based dialogue systems from small amounts of unannotated dialogue data.It combines an incremental semantic grammar -Dynamic Syntax and Type Theory with Records (DS-TTR) -with Reinforcement Learning (RL), where language generation and dialogue management are a joint decision problem.The systems thus produced are incremental: dialogues are processed word-by-word, shown previously to be essential in supporting natural, spontaneous dialogue.We hypothesised that the rich linguistic knowledge within the grammar should enable a combinatorially large number of dialogue variations to be processed, even when trained on very few dialogues.Our experiments show that our model can process 74% of the Facebook AI bAbI dataset even when trained on only 0.13% of the data (5 dialogues).It can in addition process 65% of bAbI+, a corpus 1 we created by systematically adding incremental dialogue phenomena such as restarts and self-corrections to bAbI.We compare our model with a state-of-theart retrieval model, memn2n (Bordes et al., 2017).We find that, in terms of semantic accuracy, memn2n shows very poor robustness to the bAbI+ transformations even when trained on the full bAbI dataset.
Arash Eshghi, Igor Shalyminov, Oliver Lemon
EMNLP3
2017 Teaching Robots through Situated Interactive Dialogue and Visual Demonstrations
abstract
The ability to quickly adapt to new environments and incorporate new knowledge is of great importance for robots operating in unstructured environments and interacting with non-expert users. This paper reports on our current progress in tackling this problem. We propose the development of a framework for teaching robots to perform tasks using natural language instructions, visual demonstrations and interactive dialogue. Moreover, we present a module for learning objects incrementally and on-the-fly that would enable robots to ground referents in the natural language instructions and reason about the state of the world.
Jose L. Part, Oliver Lemon
IJCAI2
2017 Hybrid chat and task dialogue for more engaging HRI using reinforcement learning
abstract
Most of today's task-based spoken dialogue systems perform poorly if the user goal is not within the system's task domain. On the other hand, chatbots cannot perform tasks involving robot actions but are able to deal with unforeseen user input. To overcome the limitations of each of these separate approaches and be able to exploit their strengths, we present and evaluate a fully autonomous robotic system using a novel combination of task-based and chat-style dialogue in order to enhance the user experience with human-robot dialogue systems. We employ Reinforcement Learning (RL) to create a scalable and extensible approach to combining chat and task-based dialogue for multimodal systems. In an evaluation with real users, the combined system was rated as significantly more “pleasant” and better met the users' expectations in a hybrid task+chat condition, compared to the task-only condition, without suffering any significant loss in task completion.
Ioannis Papaioannou, Christian Dondrup, Jekaterina Novikova, Oliver Lemon
RO-MAN4
2017 VOILA: An Optimised Dialogue System for Interactively Learning Visually-Grounded Word Meanings (Demonstration System)
abstract
We present VOILA: an optimised, multimodal dialogue agent for interactive learning of visually grounded word meanings from a human user.VOILA is: (1) able to learn new visual categories interactively from users from scratch; (2) trained on real human-human dialogues in the same domain, and so is able to conduct natural spontaneous dialogue; (3) optimised to find the most effective trade-off between the accuracy of the visual categories it learns and the cost it incurs to users.VOILA is deployed on Furhat 1 , a humanlike, multi-modal robot head with backprojection of the face, and a graphical virtual character.
Yanchao Yu, Arash Eshghi, Oliver Lemon
SIGDIAL Conference3
2016 Transfer of Reinforcement Learning Negotiation Policies: From Bilateral to Multilateral Scenarios
abstract
Trading and negotiation dialogue capabilities have been identified as important in a variety of AI application areas. In prior work, it was shown how Reinforcement Learning (RL) agents in bilateral negotiations can learn to use manipulation in dialogue to deceive adversaries in non-cooperative trading games. In this paper we show that such trained policies can also be used effectively for multilateral negotiations, and can even outperform those which are trained in these multilateral environments. Ultimately, it is shown that training in simple bilateral environments (e.g. a generic version of “Catan”) may suffice for complex multilateral non-cooperative trading scenarios (e.g. the full version of Catan).
Ioannis Efstathiou, Oliver Lemon
ECAI2
2016 How to talk to strangers: Generating medical reports for first-time users
abstract
We propose a novel approach for handling first-time users in the context of automatic report generation from time-series data in the health domain. Handling first-time users is a common problem for Natural Language Generation (NLG) and interactive systems in general - the system cannot adapt to users without prior interaction or user knowledge. In this paper, we propose a novel framework for generating medical reports for first-time users, using multi-objective optimisation (MOO) to account for the preferences of multiple possible user types, where the content preferences of potential users are modelled as objective functions. Our proposed approach outperforms two meaningful baselines in an evaluation with prospective users, yielding large (= .79) and medium (= .46) effect sizes respectively.
Dimitra Gkatzia, Verena Rieser, Oliver Lemon
FUZZ-IEEE3
2016 Crowd-sourcing NLG Data: Pictures Elicit Better Data
abstract
Recent advances in corpus-based Natural Language Generation (NLG) hold the promise of being easily portable across domains, but require costly training data, consisting of meaning representations (MRs) paired with Natural Language (NL) utterances.In this work, we propose a novel framework for crowdsourcing high quality NLG training data, using automatic quality control measures and evaluating different MRs with which to elicit data.We show that pictorial MRs result in better NL data being collected than logicbased MRs: utterances elicited by pictorial MRs are judged as significantly more natural, more informative, and better phrased, with a significant increase in average quality ratings (around 0.5 points on a 6-point scale), compared to using the logical MRs.As the MR becomes more complex, the benefits of pictorial stimuli increase.The collected data will be released as part of this submission.
Jekaterina Novikova, Oliver Lemon, Verena Rieser
INLG2
2016 Incremental Generation of Visually Grounded Language in Situated Dialogue (demonstration system)
abstract
We present a multi-modal dialogue system for interactive learning of perceptually grounded word meanings from a human tutor (Yu et al., ).The system integrates an incremental, semantic, and bidirectional grammar framework -Dynamic Syntax and Type Theory with Records (DS-TTR 1 , (Eshghi et al., 2012;Kempson et al., 2001)) -with a set of visual classifiers that are learned throughout the interaction and which ground the semantic/contextual representations that it produces (c.f.Kennington & Schlangen (2015) where words, rather than semantic atoms, are grounded in visual classifiers).Our approach extends Dobnik et al. ( 2012) in integrating perception (vision in this case) and language within a single formal system: Type Theory with Records (TTR (Cooper, 2005)).The combination of deep semantic representations in TTR with an incremental grammar (Dynamic Syntax) allows for complex multi-turn dialogues to be parsed and generated (Eshghi et al., 2015).These include clarification interaction, corrections, ellipsis and utterance continuations (see e.g. the dialogue in Fig. 1).
Yanchao Yu, Arash Eshghi, Oliver Lemon
INLG3
2016 Training an adaptive dialogue policy for interactive learning of visually grounded word meanings
abstract
We present a multi-modal dialogue system for interactive learning of perceptually grounded word meanings from a human tutor.The system integrates an incremental, semantic parsing/generation framework -Dynamic Syntax and Type Theory with Records (DS-TTR) -with a set of visual classifiers that are learned throughout the interaction and which ground the meaning representations that it produces.We use this system in interaction with a simulated human tutor to study the effects of different dialogue policies and capabilities on accuracy of learned meanings, learning rates, and efforts/costs to the tutor.We show that the overall performance of the learning agent is affected by (1) who takes initiative in the dialogues;(2) the ability to express/use their confidence level about visual attributes; and (3) the ability to process elliptical and incrementally constructed dialogue turns.Ultimately, we train an adaptive dialogue policy which optimises the trade-off between classifier accuracy and tutoring costs.
Yanchao Yu, Arash Eshghi, Oliver Lemon
SIGDIAL Conference3
2016 Information density and overlap in spoken dialogue
Nina Dethlefs, Helen Hastie, Heriberto Cuayáhuitl, Yanchao Yu, Verena Rieser, Oliver Lemon
Comput. Speech Lang.6
2015 Learning Trading Negotiations Using Manually and Automatically Labelled Data
abstract
Strategic conversational agents often need to trade resources with their opponent conversants -- and trading strategically can lead to better results. While rule-based or supervised agents can be used for such a purpose, here we explore a learning approach based on automatically labelled examples from human players for automatic trading in the game of Settlers of Catan. Our experiments are based on data collected from human players trading in text-based natural language. We compare the performance of Bayes Nets, Conditional Random Fields, and Random Forests on the task of ranking trading offers, trained from both manually labelled and automatically labelled data. Our experimental results show that our best agent trained on automatic labels outperformed its counterpart trained on manual labels (with moderate annotator agreement) in terms of (a) predicting human trading negotiations better, and (b) winning more games.
Heriberto Cuayáhuitl, Simon Keizer, Oliver Lemon
ICTAI3
2015 Erratum to: Developing technology for autism: an interdisciplinary approach
Kaska Porayska-Pomsta, Christopher Frauenberger, Helen Pain, Gnanathusharan Rajendran, Tim J. Smith, Rachel Menzies, Mary Ellen Foster, Alyssa Alcorn, Sam Wass, Sara Bernardini, Katerina Avramides, Wendy Keay-Bright, Annalu Waller, Karen Guldberg, Judith Good, Oliver Lemon
Pers. Ubiquitous Comput.17
2014 Comparing Multi-label Classification with Reinforcement Learning for Summarisation of Time-series Data
abstract
We present a novel approach for automatic report generation from time-series data, in the context of student feedback generation. Our proposed methodology treats content selection as a multi-label (ML)classification problem, which takes as input time-series data and outputs a set of templates, while capturing the dependencies between selected templates. We show that this method generates output closer to the feedback that lecturers actually generated, achieving 3.5% higher accuracy and 15% higher F-score than multiple simple classifiers that keep a history of selected templates. Furthermore, we compare a ML classifier with a Reinforcement Learning (RL) approach in simulation and using ratings from real student users. We show that the different methods have different benefits, with ML being moreaccurate for predicting what was seen in the training data, whereas RL is more exploratory and slightly preferred by the students.
Dimitra Gkatzia, Helen Hastie, Oliver Lemon
ACL (1)3
2014 Cluster-based Prediction of User Ratings for Stylistic Surface Realisation
abstract
Surface realisations typically depend on their target style and audience.A challenge in estimating a stylistic realiser from data is that humans vary significantly in their subjective perceptions of linguistic forms and styles, leading to almost no correlation between ratings of the same utterance.We address this problem in two steps.First, we estimate a mapping function between the linguistic features of a corpus of utterances and their human style ratings.Users are partitioned into clusters based on the similarity of their ratings, so that ratings for new utterances can be estimated, even for new, unknown users.In a second step, the estimated model is used to re-rank the outputs of a number of surface realisers to produce stylistically adaptive output.Results confirm that the generated styles are recognisable to human judges and that predictive models based on clusters of users lead to better rating predictions than models based on an average population of users.
Nina Dethlefs, Heriberto Cuayáhuitl, Helen Hastie, Verena Rieser, Oliver Lemon
EACL5
2014 Finding middle ground? Multi-objective Natural Language Generation from time-series data
abstract
A Natural Language Generation (NLG) system is able to generate text from nonlinguistic data, ideally personalising the content to a user’s specific needs. In some cases, however, there are multiple stakeholders with their own individual goals, needs and preferences. In this paper, we explore the feasibility of combining the preferences of two different user groups, lecturers and students, when generatingsummaries in the context of student feedback generation. The preferences of each user group are modelled as a multivariateoptimisation function, therefore the task of generation is seen as a multi-objective (MO) optimisation task, where the two functions are combined into one. This initial study shows that treating the preferences of each user group equally smooths the weights of the MO function, in a way that preferred content of the user groups isnot presented in the generated summary.
Dimitra Gkatzia, Helen Hastie, Oliver Lemon
EACL3
2014 Learning non-cooperative behaviour for dialogue agents
abstract
Non-cooperative dialogue behaviour for artificial agents (e.g. deception and information hiding) has been identified as important in a variety of application areas, including education and health-care, but it has not yet been addressed using modern statistical approaches to dialogue agents. Deception has also been argued to be a requirement for high-order intentionality in AI. We develop and evaluate a statistical dialogue agent using Reinforcement Learning which learns to perform non-cooperative dialogue moves in order to complete its own objectives in a stochastic trading game with imperfect information. We show that, when given the ability to perform both cooperative and non-cooperative dialogue moves, such an agent can learn to bluff and to lie so as to win more games. For example, we show that a non-cooperative dialogue agent learns to win 10.5% more games than a strong rule-based adversary, when compared to an optimised agent which cannot perform non-cooperative moves. This work is the first to show how agents can learn to use dialogue in a non-cooperative way to meet their own goals.
Ioannis Efstathiou, Oliver Lemon
ECAI2
2014 Towards action selection under uncertainty for a socially aware robot bartender
abstract
We describe how the state representation of a socially aware robot is being extended to handle uncertainty. It incorporates the full range of information provided by the input sensors, including the confidence of all hypotheses. We also show how the Interaction Manager is being updated to make use of the extended representation.
Mary Ellen Foster, Simon Keizer, Oliver Lemon
HRI3
2014 Multi-adaptive Natural Language Generation using Principal Component Regression
abstract
We present FeedbackGen, a system that uses a multi-adaptive approach to Natural Language Generation. With the term ‘multi-adaptive’, we refer to a system that is able to adapt its content to different user groups simultaneously, in our case adapting to both lecturers and students. We present a novel approach to student feedback generation, which simultaneously takes into account the preferences of lecturers and students when determining the content to be conveyed in a feedback summary. In this framework, we utilise knowledge derived from ratings on feedback summaries by extracting the most relevant features using Principal Component Regression (PCR) analysis. We then model a reward function that is used for training a Reinforcement Learning agent. Our results with students suggest that, from the students’ perspective, such an approach can generate more preferable summaries than a purely lecturer-adapted approach.
Dimitra Gkatzia, Helen Hastie, Oliver Lemon
INLG3
2014 Handling uncertain input in multi-user human-robot interaction
abstract
In this paper we present results from a user evaluation of a robot bartender system which handles state uncertainty derived from speech input by using belief tracking and generating appropriate clarification questions. We present a combination of state estimation and action selection components in which state uncertainty is tracked and exploited, and compare it to a baseline version that uses standard speech recognition confidence score thresholds instead of belief tracking. The results suggest that users are served fewer incorrect drinks when the uncertainty is retained in the state.
Simon Keizer, Mary Ellen Foster, Andre Gaschler, Manuel Giuliani, Amy Isard, Oliver Lemon
RO-MAN6
2014 Evaluating a social multi-user interaction model using a Nao robot
abstract
This paper presents results from a user evaluation of a robot bartender system, which supports social engagement and interaction with multiple customers. The system is a Nao-based alternative version of an existing robot bartender developed in the JAMES project [1]. The Nao-based version has given us a local experimentation platform, allowing us to focus on social multi-user interaction rather than the robot technology of object manipulation. We will describe the design of the Nao-based system and discuss the differences with the original JAMES system. In a recent evaluation of the JAMES system with real users, a trained and a hand-coded version of the action selection policy were compared [2]. Here we present results from a similar comparative user evaluation on the Nao-based system, which confirm the conclusions of the previous experiment and provide further evidence in favour of the trained action selection mechanism. Task success was found to be almost 20% higher with the trained policy, with interaction times being about 10% shorter. Participants also rated the trained system as significantly more natural, more understanding, and better at providing appropriate attention.
Simon Keizer, Pantelis Kastoris, Mary Ellen Foster, Amol A. Deshmukh, Oliver Lemon
RO-MAN5
2014 Learning non-cooperative dialogue behaviours
abstract
Non-cooperative dialogue behaviour has been identified as important in a vari-ety of application areas, including educa-tion, military operations, video games and healthcare. However, it has not been ad-dressed using statistical approaches to di-alogue management, which have always been trained for co-operative dialogue. We develop and evaluate a statistical dia-logue agent which learns to perform non-cooperative dialogue moves in order to complete its own objectives in a stochas-tic trading game. We show that, when given the ability to perform both coopera-tive and non-cooperative dialogue moves, such an agent can learn to bluff and to lie so as to win games more often – against a variety of adversaries, and under var-ious conditions such as risking penalties for being caught in deception. For exam-ple, we show that a non-cooperative dia-logue agent can learn to win an additional 15.47 % of games against a strong rule-based adversary, when compared to an op-timised agent which cannot perform non-cooperative moves. This work is the first to show how an agent can learn to use non-cooperative dialogue to effectively meet its own goals. 1
Ioannis Efstathiou, Oliver Lemon
SIGDIAL Conference2
2014 The PARLANCE mobile application for interactive search in English and Mandarin
abstract
Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Hugues Bouchard, Catherine Breslin, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gašić, James Henderson, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Tim Potter, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay, Boris Villazon-Terrazas, Majid Yazdani, Steve Young, Yanchao Yu. Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL). 2014.
Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Hugues Bouchard, Catherine Breslin, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gasic, James Henderson 0001, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Tim Potter, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay, Boris Villazón-Terrazas, Majid Yazdani, Steve J. Young, Yanchao Yu
SIGDIAL Conference10
2014 Adaptive Generation in Dialogue Systems Using Dynamic User Modeling
abstract
We address the problem of dynamically modeling and adapting to unknown users in resource-scarce domains in the context of interactive spoken dialogue systems. As an example, we show how a system can learn to choose referring expressions to refer to domain entities for users with different levels of domain expertise, and whose domain knowledge is initially unknown to the system. We approach this problem using a three step process: collecting data using a Wizard-of-Oz method, building simulated users, and learning to model and adapt to users using Reinforcement Learning techniques. We show that by using only a small corpus of non-adaptive dialogues and user knowledge profiles it is possible to learn an adaptive user modeling policy using a sense-predict-adapt approach. Our evaluation results show that the learned user modeling and adaptation strategies performed better in terms of adaptation than some simple hand-coded baseline policies, with both simulated and real users. With real users, the learned policy produced around a 20% increase in adaptation in comparison to an adaptive hand-coded baseline. We also show that adaptation to users' domain knowledge results in improving task success (99.47% for the learned policy vs. 84.7% for a hand-coded baseline) and reducing dialogue time of the conversation (11% relative difference). We also compared the learned policy to a variety of carefully hand-crafted adaptive policies that employ the user knowledge profiles to adapt their choices of referring expressions throughout a conversation. We show that the learned policy generalises better to unseen user profiles than these hand-coded policies, while having comparable performance on known user profiles. We discuss the overall advantages of this method and how it can be extended to other levels of adaptation such as content selection and dialogue management, and to other domains where adapting to users' domain knowledge is useful, such as travel and healthcare.
Srinivasan Janarthanam, Oliver Lemon
Comput. Linguistics2
2014 Real user evaluation of a POMDP spoken dialogue system using automatic belief compression
Paul A. Crook, Simon Keizer, Wenshuo Tang, Oliver Lemon
Comput. Speech Lang.5
2014 Natural Language Generation as Incremental Planning Under Uncertainty: Adaptive Information Presentation for Statistical Dialogue Systems
abstract
We present and evaluate a novel approach to natural language generation (NLG) in statistical spoken dialogue systems (SDS) using a data-driven statistical optimization framework for incremental information presentation (IP), where there is a trade-off to be solved between presenting “enough" information to the user while keeping the utterances short and understandable. The trained IP model is adaptive to variation from the current generation context (e.g. a user and a non-deterministic sentence planner), and it incrementally adapts the IP policy at the turn level. Reinforcement learning is used to automatically optimize the IP policy with respect to a data-driven objective function. In a case study on presenting restaurant information, we show that an optimized IP strategy trained on Wizard-of-Oz data outperforms a baseline mimicking the wizard behavior in terms of total reward gained. The policy is then also tested with real users, and improves on a conventional hand-coded IP strategy used in a deployed SDS in terms of overall task success. The evaluation found that the trained IP strategy significantly improves dialogue task completion for real users, with up to a 8.2% increase in task success. This methodology also provides new insights into the nature of the IP problem, which has previously been treated as a module following dialogue management with no access to lower-level context features (e.g. from a surface realizer and/or speech synthesizer).
Verena Rieser, Oliver Lemon, Simon Keizer
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Machine Learning for Social Multiparty Human-Robot Interaction
abstract
We describe a variety of machine-learning techniques that are being applied to social multiuser human--robot interaction using a robot bartender in our scenario. We first present a data-driven approach to social state recognition based on supervised learning . We then describe an approach to social skills execution—that is, action selection for generating socially appropriate robot behavior—which is based on reinforcement learning , using a data-driven simulation of multiple users to train execution policies for social skills. Next, we describe how these components for social state recognition and skills execution have been integrated into an end-to-end robot bartender system, and we discuss the results of a user evaluation. Finally, we present an alternative unsupervised learning framework that combines social state recognition and social skills execution based on hierarchical Dirichlet processes and an infinite POMDP interaction manager. The models make use of data from both human--human interactions collected in a number of German bars and human--robot interactions recorded in the evaluation of an initial version of the system.
Simon Keizer, Mary Ellen Foster, Oliver Lemon
ACM Trans. Interact. Intell. Syst.4
2013 Conditional Random Fields for Responsive Surface Realisation using Global Features
Nina Dethlefs, Helen Hastie, Heriberto Cuayáhuitl, Oliver Lemon
ACL (1)4
2013 Evaluating a City Exploration Dialogue System with Integrated Question-Answering and Pedestrian Navigation
Srinivasan Janarthanam, Oliver Lemon, Phil J. Bartie, Tiphaine Dalmas, Anna Dickinson, Xingkun Liu, William A. Mackaness, Bonnie L. Webber
ACL (1)2
2013 Barge-in effects in Bayesian dialogue act recognition and simulation
abstract
Dialogue act recognition and simulation are traditionally considered separate processes. Here, we argue that both can be fruitfully treated as interleaved processes within the same probabilistic model, leading to a synchronous improvement of performance in both. To demonstrate this, we train multiple Bayes Nets that predict the timing and content of the next user utterance. A specific focus is on providing support for barge-ins. We describe experiments using the Let's Go data that show an improvement in classification accuracy (+5%) in Bayesian dialogue act recognition involving barge-ins using partial context compared to using full context. Our results also indicate that simulated dialogues with user barge-in are more realistic than simulations without barge-in events.
Heriberto Cuayáhuitl, Nina Dethlefs, Helen Hastie, Oliver Lemon
ASRU4
2013 Impact of ASR N-Best Information on Bayesian Dialogue Act Recognition
Heriberto Cuayáhuitl, Nina Dethlefs, Helen Hastie, Oliver Lemon
SIGDIAL Conference4
2013 Demonstration of the PARLANCE system: a data-driven incremental, spoken dialogue system for interactive search
Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gasic, James Henderson 0001, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay
SIGDIAL Conference8
2013 A Multithreaded Conversational Interface for Pedestrian Navigation and Question Answering
Srinivasan Janarthanam, Oliver Lemon, Xingkun Liu, Phil J. Bartie, William A. Mackaness, Tiphaine Dalmas
SIGDIAL Conference2
2013 Training and evaluation of an MDP model for social multi-user human-robot interaction
Simon Keizer, Mary Ellen Foster, Oliver Lemon, Andre Gaschler, Manuel Giuliani
SIGDIAL Conference3
2013 A Simple and Generic Belief Tracking Mechanism for the Dialog State Tracking Challenge: On the believability of observed information
Oliver Lemon
SIGDIAL Conference2
2012 A Statistical Spoken Dialogue System using Complex User Goals and Value Directed Compression
Paul A. Crook, Xingkun Liu, Oliver Lemon
EACL4
2012 Optimising Incremental Dialogue Decisions Using Information Density for Interactive Systems
Nina Dethlefs, Helen Hastie, Verena Rieser, Oliver Lemon
EMNLP-CoNLL4
2012 Optimising Incremental Generation for Spoken Dialogue Systems: Reducing the Need for Fillers
Nina Dethlefs, Helen Hastie, Verena Rieser, Oliver Lemon
INLG4
2012 Integrating Location, Visibility, and Question-Answering in a Spoken Dialogue System for Pedestrian City Exploration
Srinivasan Janarthanam, Oliver Lemon, Xingkun Liu, Phil J. Bartie, William A. Mackaness, Tiphaine Dalmas, Jana Götze
SIGDIAL Conference2
2012 A nonparametric Bayesian approach to learning multimodal interaction management
abstract
Managing multimodal interactions between humans and computer systems requires a combination of state estimation based on multiple observation streams, and optimisation of time-dependent action selection. Previous work using partially observable Markov decision processes (POMDPs) for multimodal interaction has focused on simple turn-based systems. However, state persistence and implicit state transitions are frequent in real-world multimodal interactions. These phenomena cannot be fully modelled using turn-based systems, where the timing of system actions is a non-trivial issue. In addition, in prior work the POMDP parameterisation has been either hand-coded or learned from labelled data, which requires significant domain-specific knowledge and is labor-consuming. We therefore propose a nonparametric Bayesian method to automatically infer the (distributional) representations of POMDP states for multimodal interactive systems, without using any domain knowledge. We develop an extended version of the infinite POMDP method, to better address state persistence, implicit transition, and timing issues observed in real data. The main contribution is a “sticky” infinite POMDP model that is biased towards self-transitions. The performance of the proposed unsupervised approach is evaluated based on both artificially synthesised data and a manually transcribed and annotated human-human interaction corpus. We show statistically significant improvements (e.g. in ability of the planner to recall human bartender actions) over a supervised POMDP method.
Oliver Lemon
SLT2
2012 Developing technology for autism: an interdisciplinary approach
Kaska Porayska-Pomsta, Christopher Frauenberger, Helen Pain, Gnanathusharan Rajendran, Tim J. Smith, Rachel Menzies, Mary Ellen Foster, Alyssa Alcorn, Sam Wass, Sara Bernardini, Katerina Avramides, Wendy Keay-Bright, Annalu Waller, Karen Guldberg, Judith Good, Oliver Lemon
Pers. Ubiquitous Comput.17
2011 Social Communication between Virtual Characters and Children with Autism
Alyssa Alcorn, Helen Pain, Gnanathusharan Rajendran, Tim J. Smith, Oliver Lemon, Kaska Porayska-Pomsta, Mary Ellen Foster, Katerina Avramides, Christopher Frauenberger, Sara Bernardini
AIED5
2011 Lossless Value Directed Compression of Complex User Goal States for Statistical Spoken Dialogue Systems
abstract
This paper presents initial results in the application of Value Directed Compression (VDC) to spoken dialogue management belief states for reasoning about complex user goals. On a small but realistic SDS problem VDC generates a lossless compression which achieves a 6-fold reduction in the number of dialogue states required by a Partially Observable Markov Decision Process (POMDP) dialogue manager (DM). Reducing the number of dialogue states reduces the computational power, memory, and storage requirements of the hardware used to deploy such POMDP SDSs, thus increasing the complexity of the systems which could theoretically be deployed. In addition, in the case when on-line reinforcement learning is used to learn the DM policy, it should lead to, in this case, a 6-fold reduction in policy learning time. These are the first automatic compression results that have been presented for POMDP SDS states which represent user goals as sets over possible domain objects.
Paul A. Crook, Oliver Lemon
INTERSPEECH2
2011 Spoken Dialog Challenge 2010: Comparison of Live and Control Test Results
Alan W. Black, Susanne Burger, Alistair Conkie, Helen Hastie, Simon Keizer, Oliver Lemon, Nicolas Merigaud, Gabriel Parent, Gabriel Schubiner, Blaise Thomson, Jason D. Williams, Kai Yu 0004, Steve J. Young, Maxine Eskénazi
SIGDIAL Conference6
2011 "The day after the day after tomorrow?" A machine learning approach to adaptive temporal expression generation: training and evaluation with real users
Srinivasan Janarthanam, Helen Hastie, Oliver Lemon, Xingkun Liu
SIGDIAL Conference3
2011 Learning and Evaluation of Dialogue Strategies for New Applications: Empirical Methods for Optimization from Small Data Sets
abstract
We present a new data-driven methodology for simulation-based dialogue strategy learning, which allows us to address several problems in the field of automatic optimization of dialogue strategies: learning effective dialogue strategies when no initial data or system exists, and determining a data-driven reward function. In addition, we evaluate the result with real users, and explore how results transfer between simulated and real interactions. We use Reinforcement Learning (RL) to learn multimodal dialogue strategies by interaction with a simulated environment which is “bootstrapped” from small amounts of Wizard-of-Oz (WOZ) data. This use of WOZ data allows data-driven development of optimal strategies for domains where no working prototype is available. Using simulation-based RL allows us to find optimal policies which are not (necessarily) present in the original data. Our results show that simulation-based RL significantly outperforms the average (human wizard) strategy as learned from the data by using Supervised Learning. The bootstrapped RL-based policy gains on average 50 times more reward when tested in simulation, and almost 18 times more reward when interacting with real users. Users also subjectively rate the RL-based policy on average 10% higher. We also show that results from simulated interaction do transfer to interaction with real users, and we explicitly evaluate the stability of the data-driven reward function.
Verena Rieser, Oliver Lemon
Comput. Linguistics2
2011 Learning what to say and how to say it: Joint optimisation of spoken dialogue management and natural language generation
Oliver Lemon
Comput. Speech Lang.1
2010 Learning to Adapt to Unknown Users: Referring Expression Generation in Spoken Dialogue Systems
Srinivasan Janarthanam, Oliver Lemon
ACL2
2010 Optimising Information Presentation for Spoken Dialogue Systems
Verena Rieser, Oliver Lemon, Xingkun Liu
ACL2
2010 Generation Under Uncertainty
Oliver Lemon, Srinivasan Janarthanam, Verena Rieser
INLG1
2010 Supporting children's social communication skills through interactive narratives with virtual characters
abstract
The development of social communication skills in children relies on multimodal aspects of communication such as gaze, facial expression, and gesture. We introduce a multimodal learning environment for social skills which uses computer vision to estimate the children's gaze direction, processes gestures from a large multi-touch screen, estimates in real time the affective state of the users, and generates interactive narratives with embodied virtual characters. We also describe how the structure underlying this system is currently being extended into a general framework for the development of interactive multimodal systems.
Mary Ellen Foster, Katerina Avramides, Sara Bernardini, Christopher Frauenberger, Oliver Lemon, Kaska Porayska-Pomsta
ACM Multimedia6
2010 Representing Uncertainty about Complex User Goals in Statistical Dialogue Systems
Paul A. Crook, Oliver Lemon
SIGDIAL Conference2
2010 Adaptive Referring Expression Generation in Spoken Dialogue Systems: Evaluation with Real Users
Srinivasan Janarthanam, Oliver Lemon
SIGDIAL Conference2
2010 "Let's Go, DUDE!" using the Spoken Dialogue Challenge to teach Spoken Dialogue development
abstract
Educational tools are essential in teaching the field of Spoken Dialogue Systems given the complexity and variety of disciplines involved. This paper describes DUDE, a Dialogue and Understanding Development Environment that enables researchers and students to efficiently create Information State Update (ISU) Spoken Dialogue Systems using large scale databases with minimal programming and grammar development. The experience of creating real Spoken Dialogue Systems that they can call through a VoiceXML platform, increases students' motivation and improves learning by grounding key concepts, including introducing students to ISU dialogue modelling.
Helen Hastie, Nicolas Merigaud, Xingkun Liu, Oliver Lemon
SLT4
2010 Evaluation of a hierarchical reinforcement learning spoken dialogue system
Heriberto Cuayáhuitl, Steve Renals, Oliver Lemon, Hiroshi Shimodaira
Comput. Speech Lang.3
2010 Learning human multimodal dialogue strategies
abstract
Abstract We investigate the use of different machine learning methods in combination with feature selection techniques to explore human multimodal dialogue strategies and the use of those strategies for automated dialogue systems. We learn policies from data collected in a Wizard-of-Oz study where different human ‘wizards’ decide whether to ask a clarification request in a multimodal manner or else to use speech alone. We first describe the data collection, the coding scheme and annotated corpus, and the validation of the multimodal annotations. We then show that there is a uniform multimodal dialogue strategy across wizards, which is based on multiple features in the dialogue context. These are generic features, available at runtime, which can be implemented in dialogue systems. Our prediction models (for human wizard behaviour) achieve a weighted f-score of 88.6 per cent (which is a 25.6 per cent improvement over the majority baseline). We interpret and discuss the learned strategy. We conclude that human wizard behaviour is not optimal for automatic dialogue systems, and argue for the use of automatic optimization methods, such as Reinforcement Learning. Throughout the investigation we also discuss the issues arising from using small initial Wizard-of-Oz data sets, and we show that feature engineering is an essential step when learning dialogue strategies from such limited data.
Verena Rieser, Oliver Lemon
Nat. Lang. Eng.2
2009 User Simulations for Context-Sensitive Speech Recognition in Spoken Dialogue Systems
Oliver Lemon, Ioannis Konstas
EACL1
2009 Natural Language Generation as Planning Under Uncertainty for Spoken Dialogue Systems
Verena Rieser, Oliver Lemon
EACL2
2009 Predicting how it sounds: re-ranking dialogue prompts based on TTS quality for adaptive spoken dialogue systems
Cédric Boidin, Verena Rieser, Lonneke van der Plas, Oliver Lemon, Jonathan Chevelu
INTERSPEECH4
2009 Automatic Generation of Information State Update Dialogue Systems that Dynamically Create Voice XML, as Demonstrated on the iPhone
Helen Hastie, Xingkun Liu, Oliver Lemon
SIGDIAL Conference3
2009 A Two-Tier User Simulation Model for Reinforcement Learning of Adaptive Referring Expression Generation Policies
Srinivasan Janarthanam, Oliver Lemon
SIGDIAL Conference2
2009 Automatic annotation of context and speech acts for dialogue corpora
abstract
Abstract Richly annotated dialogue corpora are essential for new research directions in statistical learning approaches to dialogue management, context-sensitive interpretation, and context-sensitive speech recognition. In particular, large dialogue corpora annotated with contextual information and speech acts are urgently required. We explore how existing dialogue corpora (usually consisting of utterance transcriptions) can be automatically processed to yield new corpora where dialogue context and speech acts are accurately represented. We present a conceptual and computational framework for generating such corpora. As an example, we present and evaluate an automatic annotation system which builds ‘Information State Update’ (ISU) representations of dialogue context for the Communicator (2000 and 2001) corpora of human–machine dialogues (2,331 dialogues). The purposes of this annotation are to generate corpora for reinforcement learning of dialogue policies, for building user simulations, for evaluating different dialogue strategies against a baseline, and for training models for context-dependent interpretation and speech recognition. The automatic annotation system parses system and user utterances into speech acts and builds up sequences of dialogue context representations using an ISU dialogue manager. We present the architecture of the automatic annotation system and a detailed example to illustrate how the system components interact to produce the annotations. We also evaluate the annotations, with respect to the task completion metrics of the original corpus and in comparison to hand-annotated data and annotations produced by a baseline automatic system. The automatic annotations perform well and largely outperform the baseline automatic annotations in all measures. The resulting annotated corpus has been used to train high-quality user simulations and to learn successful dialogue strategies. The final corpus will be made publicly available.
Kallirroi Georgila, Oliver Lemon, James Henderson 0001, Johanna D. Moore
Nat. Lang. Eng.2
2009 Does this list contain what you were searching for? Learning adaptive dialogue strategies for interactive question answering
abstract
Abstract Policy learning is an active topic in dialogue systems research, but it has not been explored in relation to interactive question answering (IQA). We take a first step in learning adaptive interaction policies for question answering : we address the question of how to acquire enough reliable query constraints, how many database results to present to the user and when to present them, given the competing trade-offs between the length of the answer list, the length of the interaction, the type of database and the noise in the communication channel. The operating conditions are reflected in an objective function which we use to derive a hand-coded threshold-based policy and rewards to train a reinforcement learning policy. The same objective function is used for evaluation. We show that we can learn strategies for this complex trade-off problem which perform significantly better than a variety of hand-coded policies, for a wide range of noise conditions, user types, types of DB and turn-penalties. Our policy learning framework thus covers a wide spectrum of operating conditions. The learned policies produce an averagerelativeincrease in reward of 86.78% over the hand-coded policies. In 93% of the cases the learned policies perform significantly better than the hand-coded ones (p< .001). Furthermore we show that the type of database has a significant effect on learning and we give qualitative descriptions of the learned IQA policies.
Verena Rieser, Oliver Lemon
Nat. Lang. Eng.2
2008 Learning Effective Multimodal Dialogue Strategies from Wizard-of-Oz Data: Bootstrapping and Evaluation
Verena Rieser, Oliver Lemon
ACL2
2008 Using dialogue acts to learn better repair strategies for spoken dialogue systems
abstract
Repair or error-recovery strategies are an important design issue in spoken dialogue systems (SDSs) - how to conduct the dialogue when there is no progress (e.g. due to repeated ASR errors). Nearly all current SDSs use hand-crafted repair rules, but a more robust approach is to use reinforcement learning (RL) for data-driven dialogue strategy learning. However, as well as usually being tested only in simulation, current RL approaches use small state spaces which do not contain linguistically motivated features such as "dialogue acts" (DAs). We show that a strategy learned with DA features outperforms hand-crafted and slot-status strategies when tested with real users (+9% average task completion, p < 0.05). We then explore how using DAs produces better repair strategies e.g. focus-switching. We show that DAs are useful in de.ciding both when to use a repair strategy, and which one to use.
Matthew Frampton, Oliver Lemon
ICASSP2
2008 Accurate statistical spoken language understanding from limited development resources
abstract
Robust spoken language understanding (SLU) is a key component of spoken dialogue systems. Recent statistical approaches to this problem require additional resources (e.g. gazetteers, grammars, syntactic treebanks) which are expensive and time-consuming to produce and maintain. However, simple datasets annotated only with slot-values are commonly used in dialogue systems development, and are easy to collect, automatically annotate, and update. We show that it is possible to reach state-of-the-art performance using minimal additional resources, by using Markov logic networks (MLNs). We also show that performance can be further improved by exploiting long distance dependencies between slot-values. For example, by representing such features in MLNs, but without using a gazetteer, we outperform the hidden vector state (HVS) model of He and Young 2006 (1.26% improvement, a 13% error reduction).
Iván V. Meza, Sebastian Riedel 0001, Oliver Lemon
ICASSP3
2008 Automatic Learning and Evaluation of User-Centered Objective Functions for Dialogue System Optimisation
Verena Rieser, Oliver Lemon
LREC2
2008 Hybrid Reinforcement/Supervised Learning of Dialogue Policies from Fixed Data Sets
abstract
We propose a method for learning dialogue management policies from a fixed data set. The method addresses the challenges posed by Information State Update (ISU)-based dialogue systems, which represent the state of a dialogue as a large set of features, resulting in a very large state space and a huge policy space. To address the problem that any fixed data set will only provide information about small portions of these state and policy spaces, we propose a hybrid model that combines reinforcement learning with supervised learning. The reinforcement learning is used to optimize a measure of dialogue reward, while the supervised learning is used to restrict the learned policy to the portions of these spaces for which we have data. We also use linear function approximation to address the need to generalize from a fixed amount of data to large state spaces. To demonstrate the effectiveness of this method on this challenging task, we trained this model on the COMMUNICATOR corpus, to which we have added annotations for user actions and Information States. When tested with a user simulation trained on a different part of the same data set, our hybrid model outperforms a pure supervised learning model and a pure reinforcement learning model. It also outperforms the hand-crafted systems on the COMMUNICATOR data, according to automatic evaluation measures, improving over the average COMMUNICATOR system policy by 10%. The proposed method will improve techniques for bootstrapping and automatic optimization of dialogue management policies from limited initial data sets.
James Henderson 0001, Oliver Lemon, Kallirroi Georgila
Comput. Linguistics2
2007 Hierarchical dialogue optimization using semi-Markov decision processes
abstract
This paper addresses the problem of dialogue optimization on large search spaces. For such a purpose, in this paper we propose to learn dialogue strategies using multiple Semi-Markov Decision Processes and hierarchical reinforcement learning. This approach factorizes state variables and actions in order to learn a hierarchy of policies. Our experiments are based on a simulated flight booking dialogue system and compare flat versus hierarchical reinforcement learning. Experimental results show that the proposed approach produced a dramatic search space reduction (99.36 than flat reinforcement learning with a very small loss in optimality (on average 0.3 system turns). Results also report that the learnt policies outperformed a hand-crafted one under three different conditions of ASR confidence levels. This approach is appealing to dialogue optimization due to faster learning, reusable subsolutions, and scalability to larger problems.
Heriberto Cuayáhuitl, Steve Renals, Oliver Lemon, Hiroshi Shimodaira
INTERSPEECH3
2007 Machine learning for spoken dialogue systems
abstract
During the last decade, research in the field of Spoken Dialogue Systems (SDS) has experienced increasing growth. However, the design and optimization of SDS is not only about combin- ing speech and language processing systems such as Automatic Speech Recognition (ASR), parsers, Natural Language Gener- ation (NLG), and Text-to-Speech (TTS) synthesis systems. It also requires the development of dialogue strategies taking at least into account the performances of these subsystems (and others), the nature of the task (e.g. form filling, tutoring, robot control, or database search/browsing), and the user's behaviour (e.g. cooperativeness, expertise). Due to the great variability of these factors, reuse of previous hand-crafted designs is also made very difficult. For these reasons, statistical machine learn- ing (ML) methods applied to automatic SDS optimization have been a leading research area for the last few years. In this paper, we provide a short review of the field and of recent advances.
Oliver Lemon, Olivier Pietquin
INTERSPEECH1
2007 Learning dialogue strategies for interactive database search
Verena Rieser, Oliver Lemon
INTERSPEECH2
2006 Learning More Effective Dialogue Strategies Using Limited Dialogue Move Features
abstract
We explore the use of restricted dialogue contexts in reinforcement learning (RL) of effective dialogue strategies for information seeking spoken dialogue systems (e.g. COMMUNICATOR (Walker et al., 2001)). The contexts we use are richer than previous research in this area, e.g. (Levin and Pieraccini, 1997; Scheffler and Young, 2001; Singh et al., 2002; Pietquin, 2004), which use only slot-based information, but are much less complex than the full dialogue "Information States" explored in (Henderson et al., 2005), for which tractabe learning is an issue. We explore how incrementally adding richer features allows learning of more effective dialogue strategies. We use 2 user simulations learned from COMMUNICATOR data (Walker et al., 2001; Georgila et al., 2005b) to explore the effects of different features on learned dialogue strategies. Our results show that adding the dialogue moves of the last system and user turns increases the average reward of the automatically learned strategies by 65.9% over the original (hand-coded) COMMUNICATOR systems, and by 7.8% over a baseline RL policy that uses only slot-status features. We show that the learned strategies exhibit an emergent "focus switching" strategy and effective use of the 'give help' action.
Matthew Frampton, Oliver Lemon
ACL2
2006 Using Machine Learning to Explore Human Multimodal Clarification Strategies
Verena Rieser, Oliver Lemon
ACL2
2006 An ISU Dialogue System Exhibiting Reinforcement Learning of Dialogue Policies: Generic Slot-Filling in the TALK In-car System
Oliver Lemon, Kallirroi Georgila, James Henderson 0001, Matthew N. Stuttle
EACL1
2006 DUDE: A Dialogue and Understanding Development Environment, Mapping Business Process Models to Information State Update Dialogue Systems
Oliver Lemon, Xingkun Liu
EACL1
2006 Learning multi-goal dialogue strategies using reinforcement learning with reduced state-action spaces
abstract
Learning dialogue strategies using the reinforcement learning framework is problematic due to its expensive computational cost. In this paper we propose an algorithm that reduces a state-action space to one which includes only valid state-actions. We performed experiments on full and reduced spaces using three systems (with 5, 9 and 20 slots) in the travel domain using a simulated environment. The task was to learn multi-goal dialogue strategies optimizing single and multiple confirmations. Average results using strategies learnt on reduced spaces reveal the following benefits against full spaces: 1) less computer memory (94 % reduction), 2) faster learning (93 % faster convergence) and better performance (8.4 % less time steps and 7.7 % higher reward). Index Terms: reinforcement learning, spoken dialogue systems. 1.
Heriberto Cuayáhuitl, Steve Renals, Oliver Lemon, Hiroshi Shimodaira
INTERSPEECH3
2006 User simulation for spoken dialogue systems: learning and evaluation
abstract
We propose the “advanced ” n-grams as a new technique for simulating user behaviour in spoken dialogue systems, and we compare it with two methods used in our prior work, i.e. linear feature combination and “normal ” n-grams. All methods operate on the intention level and can incorporate speech recognition and understanding errors. In the linear feature combination model user actions (lists of 〈 speech act, task 〉 pairs) are selected, based on features of the current dialogue state which encodes the whole history of the dialogue. The user simulation based on “normal ” n-grams treats a dialogue as a sequence of lists of 〈 speech act, task 〉 pairs. Here the length of the history considered is restricted by the order of the n-gram. The “advanced ” n-grams are a variation of the normal ngrams, where user actions are conditioned not only on speech acts and tasks but also on the current status of the tasks, i.e. whether
Kallirroi Georgila, James Henderson 0001, Oliver Lemon
INTERSPEECH3
2006 Cluster-based user simulations for learning dialogue strategies
abstract
Abstract Good dialogue strate gies in spok en dialogue systems help to en-sure and maintain mutual understanding and thus play a crucialrole in rob ust con versational interaction. W e focus on clariÞca-tion strate gies and build user simulations which are critical forreinforcement learning, which is a cheap and principled w ay toautomatically optimise dialogue management. In this paper wepresent a no vel cluster -based technique for building user simula-tions which sho w varying , but complete and consistent beha viourwith respect to real users. W e use this technique to build usersimulations and we also introduce the S U P E R evaluation metricwhich allo ws us to evaluate user simulations with respect to thesedesiderata. W e sho w that the cluster -based user simulation tech-nique performs signiÞcantly better (at P < 0.01 ) than decisionsmade using either the one most likely action or a random base-line. The cluster -based user simulations reduce the average errorof these other models by 53% and 34% respecti vely .Index T erms : spok en dialogue, user simulation, evaluation met-rics, reinforcement learning, dialogue strate gies
Verena Rieser, Oliver Lemon
INTERSPEECH2
2006 Evolving optimal inspectable strategies for spoken dialogue systems
Dave Toney, Johanna D. Moore, Oliver Lemon
HLT-NAACL3
2006 Let's Discoh: Collecting an Annotated Open Corpuswith Dialogue Acts and Reward signals for Natural Language Helpdesks
abstract
We motivate and explain the DlSCoH project, which uses a publicly deployed spoken dialogue system for conference services to collect a richly annotated corpus of mixed-initiative human- machine spoken dialogues. System users are able to call a phone number and learn about a conference, including paper submission, program, venue, accommodation options and costs, etc. The collected corpus is (1) usable for training, evaluating and comparing statistical models, (2) naturally spoken and task oriented, (3) extendible / generalizable, (4) collected using state-of-the-art research and commercial technology, (5) freely available to researchers. We explain the principles behind the dialogue context representations and reward signals collected by the system, as well as the overall system design, call types, and call flow. We also present results regarding the initial ASR models and spoken language understanding models. We expect the resulting corpora to be used in advanced dialogue research over the coming years.
Giovanni Andreani, Giuseppe Di Fabbrizio, Mazin Gilbert, Daniel Gillick, Dilek Hakkani-Tür, Oliver Lemon
SLT6
2006 Reinforcement Learning of Dialogue Strategies with Hierarchical Abstract Machines
abstract
In this paper we propose partially specified dialogue strategies for dialogue strategy optimization, where part of the strategy is specified deterministically and the rest optimized with reinforcement learning (RL). To do this we apply RL with hierarchical abstract machines (HAMs). We also propose to build simulated users using HAMs, incorporating a combination of hierarchical deterministic and probabilistic behaviour. We performed experiments using a single-goal flight booking dialogue system, and compare two dialogue strategies (deterministic and optimized) using three types of simulated user (novice, experienced and expert). Our results show that HAMs are promising for both dialogue optimization and simulation, and provide evidence that indeed partially specified dialogue strategies can outperform deterministic ones (on average 4.7 fewer system turns) with faster learning than the traditional RL framework.
Heriberto Cuayáhuitl, Steve Renals, Oliver Lemon, Hiroshi Shimodaira
SLT3
2006 Evaluating Effectiveness and Portability of Reinforcement Learned Dialogue Strategies with Real Users: the Talk Towninfo Evaluation
abstract
We report evaluation results for real users of a learnt dialogue management policy versus a hand-coded policy in the TALK project's "Townlnfo" tourist information system. The learnt policy, for filling and confirming information slots, was derived from COMMUNICATOR (flight-booking) data using reinforcement learning (RL) as described in [2], ported to the tourist information domain (using a general method that we propose here), and tested using 18 human users in 180 dialogues, who also used a state-of-the-art hand- coded dialogue policy embedded in an otherwise identical system. We found that users of the (ported) learned policy had an average gain in perceived task completion of 14.2% (from 67.6% to 81.8% at p < .03), that the hand-coded policy dialogues had on average 3.3 more system turns (p < .01), and that the user satisfaction results were comparable, even though the policy was learned for a different domain. Combining these in a dialogue reward score, we found a 14.4% increase for the learnt policy (a 23.8% relative increase, p < .03). These results are important because they show a) that results for real users are consistent with results for automatic evaluation [2] of learned policies using simulated users [3, 4], b) that a policy learned using linear function approximation over a very large policy space [2] is effective for real users, and c) that policies learned using data for one domain can be used successfully in other domains. We also present a qualitative discussion of the learnt policy.
Oliver Lemon, Kallirroi Georgila, James Henderson 0001
SLT1
2006 Using logistic Regression to Initialise Reinforcement-Learning-Based Dialogue Systems
abstract
We investigate the use of logistic regression (LR) to initialise reinforcement learning (RL)-based dialogue systems with models of human dialogue strategies. LR produces accurate predictions and performs feature selection. We illustrate this technique in exploring human multimodal clarification strategies, observed in a Wizard-of-Oz experiment. We use it to initialise an RL-based system with features which significantly influence human behaviour. We show that the strategy applied by the human wizards is sensitive to different dialogue contexts. Furthermore we show that for predicting clarification behaviour the logistic models improve over the baseline on average twice as much as the supervised learning techniques used in previous work.
Verena Rieser, Oliver Lemon
SLT2
2005 Learning user simulations for information state update dialogue systems
abstract
This paper describes and compares two methods for simulating user behaviour in spoken dialogue systems. User simulations are important for automatic dialogue strategy learning and the evaluation of competing strategies. Our methods are designed for use with "Information State Update" (ISU)-based dialogue systems. The first method is based on supervised learning using linear feature combination and a normalised exponential output function. The user is modelled as a stochastic process which selects user actions ( pairs) based on features of the current dialogue state, which encodes the whole history of the dialogue. The second method uses n-grams of speech act, task pairs, restricting the length of the history considered by the order of the n-gram. Both models were trained and evaluated on a subset of the COMMUNICATOR corpus, to which we added annotations for user actions and Information States. The model based on linear feature combination has a perplexity of 2.08 whereas the best n-gram (4-gram) has a perplexity of 3.58. Each one of the user models ran against a system policy trained on the same corpus with a method similar to the one used for our linear feature combination model. The quality of the simulated dialogues produced was then measured as a function of the filled slots, confirmed slots, and number of actions performed by the system in each dialogue. In this experiment both the linear feature combination model and the best n-grams (5-gram and 4-gram) produced similar quality simulated dialogues.
Kallirroi Georgila, James Henderson 0001, Oliver Lemon
INTERSPEECH3
2004 Combining Acoustic and Pragmatic Features to Predict Recognition Performance in Spoken Dialogue Systems
abstract
We use machine learners trained on a combination of acoustic confidence and pragmatic plausibility features computed from dialogue context to predict the accuracy of incoming n-best recognition hypotheses to a spoken dialogue system. Our best results show a 25% weighted f-score improvement over a baseline system that implements a "grammar-switching" approach to context-sensitive speech recognition.
Malte Gabsdil, Oliver Lemon
ACL2
2004 multithreaded context for robust conversational interfaces: Context-sensitive speech recognition and interpretation of corrective fragments
abstract
We focus on the issue of robustness of conversational interfaces that are flexible enough to allow natural "multithreaded" conversational flow. Our main advance is to use context-sensitive speech recognition in a general way, with a representation of dialogue context that is rich and flexible enough to support conversation about multiple interleaved topics, as well as the interpretation of corrective fragments. We explain, by use of worked examples, the use of our "Conversational Intelligence Architecture" (CIA) to represent conversational threads, and how each thread can be associated with a language model (LM) for more robust speech recognition. The CIA uses fine-grained dynamic representations of dialogue context, which supersede those used in finite-state or form-based dialogue managers. In an evaluation of a dialogue system built using this architecture we found that 87.9% of recognized utterances were recognized using a context-specific language model, resulting in an 11.5% reduction in the overall utterance recognition error rate, and a 13.4% reduction in concept error rate. Thus we show that by using context-sensitive recognition based on the predicted type of the user's next dialogue move, a more flexible dialogue system can also exhibit an improvement in speech recognition performance.
Oliver Lemon, Alexander Gruenstein
ACM Trans. Comput. Hum. Interact.1
2003 Targeted Help for Spoken Dialogue Systems
Beth Ann Hockey, Oliver Lemon, Ellen Campana, Laura M. Hiatt, Gregory Aist, James Hieronymus, John Dowding, Alexander Gruenstein
EACL2
2002 Language Resources for Multi-Modal Dialogue Systems
Oliver Lemon, Alexander Gruenstein
LREC1
2001 The WITAS multi-modal dialogue system I
Oliver Lemon, Anne Bracy, Alexander Gruenstein, Stanley Peters
INTERSPEECH1
2000 Constraint Matching for Diagram Design: Qualitative Visual Languages
Ana von Klopp Lemon, Oliver Lemon
Diagrams2
1996 Semantical Foundations of Spatial Logics
Oliver Lemon
KR1