Milica Gasic

dblp:27/7520 · DBLP profile ↗
← Back
91ranked-venue papers
13as first author
20since 2021 · last 2026
0000-0003-0318-9147ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 73 · 10 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 7 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Text-to-SQL Task-oriented Dialogue Ontology Construction
abstract
Abstract Large language models (LLMs) are widely used as general-purpose knowledge sources, but they rely on parametric knowledge, limiting explainability and trustworthiness. In task-oriented dialogue (TOD) systems, this separation is explicit, using an external database structured by an explicit ontology to ensure explainability and controllability. However, building such ontologies requires manual labels or supervised training. We introduce TeQoDO: a Text-to-SQL task-oriented Dialogue Ontology construction method. Here, an LLM autonomously builds a TOD ontology from scratch using only its inherent SQL programming capabilities combined with concepts from modular TOD systems provided in the prompt. We show that TeQoDO outperforms transfer learning approaches, and its constructed ontology is competitive on a downstream dialogue state tracking task. Ablation studies demonstrate the key role of modular TOD system concepts. TeQoDO also scales to allow construction of much larger ontologies, which we investigate on a Wikipedia and arXiv dataset. We view this as a step towards broader application of ontologies.1
Renato Vukovic, Carel van Niekerk, Michael Heck, Benjamin Matthias Ruppik, Hsien-Chin Lin, Shutong Feng, Nurul Lubis, Milica Gasic
Trans. Assoc. Comput. Linguistics8
2025 Learning from Noisy Labels via Self-Taught On-the-Fly Meta Loss Rescaling
abstract
Correct labels are indispensable for training effective machine learning models. However, creating high-quality labels is expensive, and even professionally labeled data contains errors and ambiguities. Filtering and denoising can be applied to curate labeled data prior to training, at the cost of additional processing and loss of information. An alternative is on-the-fly sample reweighting during the training process to decrease the negative impact of incorrect or ambiguous labels, but this typically requires clean seed data. In this work we propose unsupervised on-the-fly meta loss rescaling to reweight training samples. Crucially, we rely only on features provided by the model being trained, to learn a rescaling function in real time without knowledge of the true clean data distribution. We achieve this via a novel meta learning setup that samples validation data for the meta update directly from the noisy training corpus by employing the rescaling function being trained. Our proposed method consistently improves performance across various NLP tasks with minimal computational overhead. Further, we are among the first to attempt on-the-fly training data reweighting on the challenging task of dialogue modeling, where noisy and ambiguous labels are common. Our strategy is robust in the face of noisy and clean data, handles class imbalance, and prevents overfitting to noisy labels. Our self-taught loss rescaling improves as the model trains, showing the ability to keep learning from the model's own signals. As training progresses, the impact of correctly labeled data is scaled up, while the impact of wrongly labeled data is suppressed.
Michael Heck, Christian Geishauser, Nurul Lubis, Carel van Niekerk, Shutong Feng, Hsien-Chin Lin, Benjamin Matthias Ruppik, Renato Vukovic, Milica Gasic
AAAI9
2025 Less is More: Local Intrinsic Dimensions of Contextual Language Models
abstract
Understanding the internal mechanisms of large language models (LLMs) remains a challenging and complex endeavor. Even fundamental questions, such as how fine-tuning affects model behavior, often require extensive empirical evaluation. In this paper, we introduce a novel perspective based on the geometric properties of contextual latent embeddings to study the effects of training and fine-tuning. To that end, we measure the local dimensions of a contextual language model's latent space and analyze their shifts during training and fine-tuning. We show that the local dimensions provide insights into the model's training dynamics and generalization ability. Specifically, the mean of the local dimensions predicts when the model’s training capabilities are exhausted, as exemplified in a dialogue state tracking task, overfitting, as demonstrated in an emotion recognition task, and grokking, as illustrated with an arithmetic task. Furthermore, our experiments suggest a practical heuristic: reductions in the mean local dimension tend to accompany and predict subsequent performance gains. Through this exploration, we aim to provide practitioners with a deeper understanding of the implications of fine-tuning on embedding spaces, facilitating informed decisions when configuring models for specific applications. The results of this work contribute to the ongoing discourse on the interpretability, adaptability, and generalizability of LLMs by bridging the gap between intrinsic model mechanisms and geometric properties in the respective embeddings.
Benjamin Matthias Ruppik, Julius von Rohrscheidt, Carel van Niekerk, Michael Heck, Renato Vukovic, Shutong Feng, Hsien-Chin Lin, Nurul Lubis, Bastian Rieck, Marcus Zibrowius, Milica Gasic
NeurIPS11
2025 A Confidence-based Acquisition Model for Self-supervised Active Learning and Label Correction
abstract
Abstract Supervised neural approaches are hindered by their dependence on large, meticulously annotated datasets, a requirement that is particularly cumbersome for sequential tasks. The quality of annotations tends to deteriorate with the transition from expert-based to crowd-sourced labeling. To address these challenges, we present CAMEL (Confidence-based Acquisition Model for Efficient self-supervised active Learning), a pool-based active learning framework tailored to sequential multi-output problems. CAMEL possesses two core features: (1) it requires expert annotators to label only a fraction of a chosen sequence, and (2) it facilitates self-supervision for the remainder of the sequence. By deploying a label correction mechanism, CAMEL can also be utilized for data cleaning. We evaluate CAMEL on two sequential tasks, with a special emphasis on dialogue belief tracking, a task plagued by the constraints of limited and noisy datasets. Our experiments demonstrate that CAMEL significantly outperforms the baselines in terms of efficiency. Furthermore, the data corrections suggested by our method contribute to an overall improvement in the quality of the resulting datasets.1
Carel van Niekerk, Christian Geishauser, Michael Heck, Shutong Feng, Hsien-Chin Lin, Nurul Lubis, Benjamin Matthias Ruppik, Renato Vukovic, Milica Gasic
Trans. Assoc. Comput. Linguistics9
2024 Infusing Emotions into Task-oriented Dialogue Systems: Understanding, Management, and Generation
abstract
Shutong Feng, Hsien-chin Lin, Christian Geishauser, Nurul Lubis, Carel van Niekerk, Michael Heck, Benjamin Ruppik, Renato Vukovic, Milica Gašić. Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2024.
Shutong Feng, Hsien-Chin Lin, Christian Geishauser, Nurul Lubis, Carel van Niekerk, Michael Heck, Benjamin Matthias Ruppik, Renato Vukovic, Milica Gasic
SIGDIAL9
2024 Affect Recognition in Conversations Using Large Language Models
abstract
Affect recognition, encompassing emotions, moods, and feelings, plays a pivotal role in human communication.In the realm of conversational artificial intelligence, the ability to discern and respond to human affective cues is a critical factor for creating engaging and empathetic interactions.This study investigates the capacity of large language models (LLMs) to recognise human affect in conversations, with a focus on both open-domain chit-chat dialogues and task-oriented dialogues.Leveraging three diverse datasets, namely IEMOCAP (Busso et al., 2008), EmoWOZ (Feng et al., 2022), and DAIC-WOZ (Gratch et al., 2014), covering a spectrum of dialogues from casual conversations to clinical interviews, we evaluate and compare LLMs' performance in affect recognition.Our investigation explores the zero-shot and few-shot capabilities of LLMs through incontext learning as well as their model capacities through task-specific fine-tuning.Additionally, this study takes into account the potential impact of automatic speech recognition errors on LLM predictions.With this work, we aim to shed light on the extent to which LLMs can replicate human-like affect recognition capabilities in conversations.
Shutong Feng, Guangzhi Sun, Nurul Lubis, Wen Wu 0007, Chao Zhang 0031, Milica Gasic
SIGDIAL6
2024 Local Topology Measures of Contextual Language Model Latent Spaces with Applications to Dialogue Term Extraction
abstract
Benjamin Matthias Ruppik, Michael Heck, Carel van Niekerk, Renato Vukovic, Hsien-chin Lin, Shutong Feng, Marcus Zibrowius, Milica Gasic. Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2024.
Benjamin Matthias Ruppik, Michael Heck, Carel van Niekerk, Renato Vukovic, Hsien-Chin Lin, Shutong Feng, Marcus Zibrowius, Milica Gasic
SIGDIAL8
2024 Dialogue Ontology Relation Extraction via Constrained Chain-of-Thought Decoding
abstract
Renato Vukovic, David Arps, Carel van Niekerk, Benjamin Matthias Ruppik, Hsien-chin Lin, Michael Heck, Milica Gasic. Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2024.
Renato Vukovic, David Arps, Carel van Niekerk, Benjamin Matthias Ruppik, Hsien-Chin Lin, Michael Heck, Milica Gasic
SIGDIAL7
2024 Learning With an Open Horizon in Ever-Changing Dialogue Circumstances
abstract
Task-orienteddialogue systems aid users in achieving their goals for specific tasks, e.g., booking a hotel room or managing a schedule. The systems experience various changes during their lifetime such as new tasks emerging or varying user behaviours and task requests, which requires the ability of continually learning throughout their lifetime. Current dialogue systems either perform no continual learning or do it in an unrealistic way that mostly focuses on avoiding catastrophic forgetting. Unlike current dialogue systems, humans learn in such a way that it benefits their present and future, while adapting their behaviour to current circumstances. In order to equip dialogue systems with the capability of learning for the future, we propose the usage of lifetime return in the reinforcement learning (RL) objective of dialogue policies. Moreover, we enable dynamic adaptation of hyperparameters of the underlying RL algorithm used for training the dialogue policy by employing meta-gradient reinforcement learning. We furthermore propose a more general and challenging continual learning environment in order to approximate how dialogue systems can learn in the ever-changing real world. Extensive experiments demonstrate that lifetime return and meta-gradient RL lead to more robust and improved results in continuously changing circumstances. The results warrant further development of dialogue systems that evolve throughout their lifetime.
Christian Geishauser, Carel van Niekerk, Nurul Lubis, Hsien-Chin Lin, Michael Heck, Shutong Feng, Benjamin Matthias Ruppik, Renato Vukovic, Milica Gasic
IEEE ACM Trans. Audio Speech Lang. Process.9
2023 From Chatter to Matter: Addressing Critical Steps of Emotion Recognition Learning in Task-oriented Dialogue
abstract
Shutong Feng, Nurul Lubis, Benjamin Ruppik, Christian Geishauser, Michael Heck, Hsien-chin Lin, Carel van Niekerk, Renato Vukovic, Milica Gasic. Proceedings of the 24th Meeting of the Special Interest Group on Discourse and Dialogue. 2023.
Shutong Feng, Nurul Lubis, Benjamin Matthias Ruppik, Christian Geishauser, Michael Heck, Hsien-Chin Lin, Carel van Niekerk, Renato Vukovic, Milica Gasic
SIGDIAL9
2023 EmoUS: Simulating User Emotions in Task-Oriented Dialogues
abstract
Existing user simulators (USs) for task-oriented dialogue systems only model user behaviour on semantic and natural language levels without considering the user persona and emotions. Optimising dialogue systems with generic user policies, which cannot model diverse user behaviour driven by different emotional states, may result in a high drop-off rate when deployed in the real world. Thus, we present EmoUS, a user simulator that learns to simulate user emotions alongside user behaviour. EmoUS generates user emotions, semantic actions, and natural language responses based on the user goal, the dialogue history, and the user persona. By analysing what kind of system behaviour elicits what kind of user emotions, we show that EmoUS can be used as a probe to evaluate a variety of dialogue systems and in particular their effect on the user's emotional state. Developing such methods is important in the age of large language model chat-bots and rising ethical concerns.
Hsien-Chin Lin, Shutong Feng, Christian Geishauser, Nurul Lubis, Carel van Niekerk, Michael Heck, Benjamin Matthias Ruppik, Renato Vukovic, Milica Gasic
SIGIR9
2022 Dynamic Dialogue Policy for Continual Reinforcement Learning
abstract
Continual learning is one of the key components of human learning and a necessary requirement of artificial intelligence. As dialogue can potentially span infinitely many topics and tasks, a task-oriented dialogue system must have the capability to continually learn, dynamically adapting to new challenges while preserving the knowledge it already acquired. Despite the importance, continual reinforcement learning of the dialogue policy has remained largely unaddressed. The lack of a framework with training protocols, baseline models and suitable metrics, has so far hindered research in this direction. In this work we fill precisely this gap, enabling research in dialogue policy optimisation to go from static to dynamic learning. We provide a continual learning algorithm, baseline architectures and metrics for assessing continual learning models. Moreover, we propose the dynamic dialogue policy transformer (DDPT), a novel dynamic architecture that can integrate new knowledge seamlessly, is capable of handling large state spaces and obtains significant zero-shot performance when being exposed to unseen domains, without any growth in network parameter size. We validate the strengths of DDPT in simulation with two user simulators as well as with humans.
Christian Geishauser, Carel van Niekerk, Hsien-Chin Lin, Nurul Lubis, Michael Heck, Shutong Feng, Milica Gasic
COLING7
2022 EmoWOZ: A Large-Scale Corpus and Labelling Scheme for Emotion Recognition in Task-Oriented Dialogue Systems
abstract
The ability to recognise emotions lends a conversational artificial intelligence a human touch. While emotions in chit-chat dialogues have received substantial attention, emotions in task-oriented dialogues remain largely unaddressed. This is despite emotions and dialogue success having equally important roles in a natural system. Existing emotion-annotated task-oriented corpora are limited in size, label richness, and public availability, creating a bottleneck for downstream tasks. To lay a foundation for studies on emotions in task-oriented dialogues, we introduce EmoWOZ, a large-scale manually emotion-annotated corpus of task-oriented dialogues. EmoWOZ is based on MultiWOZ, a multi-domain task-oriented dialogue dataset. It contains more than 11K dialogues with more than 83K emotion annotations of user utterances. In addition to Wizard-of-Oz dialogues from MultiWOZ, we collect human-machine dialogues within the same set of domains to sufficiently cover the space of various emotions that can happen during the lifetime of a data-driven dialogue system. To the best of our knowledge, this is the first large-scale open-source corpus of its kind. We propose a novel emotion labelling scheme, which is tailored to task-oriented dialogues. We report a set of experimental results to show the usability of this corpus for emotion recognition and state tracking in task-oriented dialogues.
Shutong Feng, Nurul Lubis, Christian Geishauser, Hsien-Chin Lin, Michael Heck, Carel van Niekerk, Milica Gasic
LREC7
2022 GenTUS: Simulating User Behaviour and Language in Task-oriented Dialogues with Generative Transformers
abstract
Hsien-chin Lin, Christian Geishauser, Shutong Feng, Nurul Lubis, Carel van Niekerk, Michael Heck, Milica Gasic. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2022.
Hsien-Chin Lin, Christian Geishauser, Shutong Feng, Nurul Lubis, Carel van Niekerk, Michael Heck, Milica Gasic
SIGDIAL7
2022 Dialogue Evaluation with Offline Reinforcement Learning
abstract
Nurul Lubis, Christian Geishauser, Hsien-chin Lin, Carel van Niekerk, Michael Heck, Shutong Feng, Milica Gasic. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2022.
Nurul Lubis, Christian Geishauser, Hsien-Chin Lin, Carel van Niekerk, Michael Heck, Shutong Feng, Milica Gasic
SIGDIAL7
2022 Dialogue Term Extraction using Transfer Learning and Topological Data Analysis
abstract
Goal oriented dialogue systems were originally designed as a natural language interface to a fixed data-set of entities that users might inquire about, further described by domain, slots and values.As we move towards adaptable dialogue systems where knowledge about domains, slots and values may change, there is an increasing need to automatically extract these terms from raw dialogues or related nondialogue data on a large scale.In this paper, we take an important step in this direction by exploring different features that can enable systems to discover realizations of domains, slots and values in dialogues in a purely data-driven fashion.The features that we examine stem from word embeddings, language modelling features, as well as topological features of the word embedding space.To examine the utility of each feature set, we train a seed model based on the widely used MultiWOZ data-set.Then, we apply this model to a different corpus, the Schema-Guided Dialogue data-set.Our method outperforms the previously proposed approach that relies solely on word embeddings.We also demonstrate that each of the features is responsible for discovering different kinds of content.We believe our results warrant further research towards ontology induction, and continued harnessing of topological data analysis for dialogue and natural language processing research.
Renato Vukovic, Michael Heck, Benjamin Matthias Ruppik, Carel van Niekerk, Marcus Zibrowius, Milica Gasic
SIGDIAL6
2022 Robust Dialogue State Tracking with Weak Supervision and Sparse Data
abstract
Abstract Generalizing dialogue state tracking (DST) to new data is especially challenging due to the strong reliance on abundant and fine-grained supervision during training. Sample sparsity, distributional shift, and the occurrence of new concepts and topics frequently lead to severe performance degradation during inference. In this paper we propose a training strategy to build extractive DST models without the need for fine-grained manual span labels. Two novel input-level dropout methods mitigate the negative impact of sample sparsity. We propose a new model architecture with a unified encoder that supports value as well as slot independence by leveraging the attention mechanism. We combine the strengths of triple copy strategy DST and value matching to benefit from complementary predictions without violating the principle of ontology independence. Our experiments demonstrate that an extractive DST model can be trained without manual span labels. Our architecture and training strategies improve robustness towards sample sparsity, new concepts, and topics, leading to state-of-the-art performance on a range of benchmarks. We further highlight our model’s ability to effectively learn from non-dialogue data.
Michael Heck, Nurul Lubis, Carel van Niekerk, Shutong Feng, Christian Geishauser, Hsien-Chin Lin, Milica Gasic
Trans. Assoc. Comput. Linguistics7
2021 What does the User Want? Information Gain for Hierarchical Dialogue Policy Optimisation
abstract
The dialogue management component of a task-oriented dialogue system is typically optimised via reinforcement learning (RL). Optimisation via RL is highly susceptible to sample inefficiency and instability. The hierarchical approach called Feudal Dialogue Management takes a step towards more efficient learning by decomposing the action space. However, it still suffers from instability due to the reward only being provided at the end of the dialogue. We propose the usage of an intrinsic reward based on information gain to address this issue. Our proposed reward favours actions that resolve uncertainty or query the user whenever necessary. It enables the policy to learn how to retrieve the users' needs efficiently, which is an integral aspect in every task-oriented conversation. Our algorithm, which we call FeudalGain, achieves state-of-the-art results in most environments of the PyDial framework, outperforming much more complex approaches. We confirm the sample efficiency and stability of our algorithm through experiments in simulation and a human trial.
Christian Geishauser, Songbo Hu, Hsien-Chin Lin, Nurul Lubis, Michael Heck, Shutong Feng, Carel van Niekerk, Milica Gasic
ASRU8
2021 Uncertainty Measures in Neural Belief Tracking and the Effects on Dialogue Policy Performance
abstract
Carel van Niekerk, Andrey Malinin, Christian Geishauser, Michael Heck, Hsien-chin Lin, Nurul Lubis, Shutong Feng, Milica Gasic. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Carel van Niekerk, Andrey Malinin, Christian Geishauser, Michael Heck, Hsien-Chin Lin, Nurul Lubis, Shutong Feng, Milica Gasic
EMNLP (1)8
2021 Domain-independent User Simulation with Transformers for Task-oriented Dialogue Systems
abstract
Hsien-chin Lin, Nurul Lubis, Songbo Hu, Carel van Niekerk, Christian Geishauser, Michael Heck, Shutong Feng, Milica Gasic. Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2021.
Hsien-Chin Lin, Nurul Lubis, Songbo Hu, Carel van Niekerk, Christian Geishauser, Michael Heck, Shutong Feng, Milica Gasic
SIGDIAL8
2020 Out-of-Task Training for Dialog State Tracking Models
abstract
Dialog state tracking (DST) suffers from severe data sparsity.While many natural language processing (NLP) tasks benefit from transfer learning and multi-task learning, in dialog these methods are limited by the amount of available data and by the specificity of dialog applications.In this work, we successfully utilize non-dialog data from unrelated NLP tasks to train dialog state trackers.This opens the door to the abundance of unrelated NLP corpora to mitigate the data sparsity issue inherent to DST.
Michael Heck, Christian Geishauser, Hsien-Chin Lin, Nurul Lubis, Marco Moresi, Carel van Niekerk, Milica Gasic
COLING7
2020 LAVA: Latent Action Spaces via Variational Auto-encoding for Dialogue Policy Optimization
abstract
Reinforcement learning (RL) can enable task-oriented dialogue systems to steer the conversation towards successful task completion.In an end-to-end setting, a response can be constructed in a word-level sequential decision making process with the entire system vocabulary as action space.Policies trained in such a fashion do not require expert-defined action spaces, but they have to deal with large action spaces and long trajectories, making RL impractical.Using the latent space of a variational model as action space alleviates this problem.However, current approaches use an uninformed prior for training and optimize the latent distribution solely on the context.It is therefore unclear whether the latent representation truly encodes the characteristics of different actions.In this paper, we explore three ways of leveraging an auxiliary task to shape the latent variable distribution: via pre-training, to obtain an informed prior, and via multitask learning.We choose response auto-encoding as the auxiliary task, as this captures the generative factors of dialogue responses while requiring low computational cost and neither additional data nor labels.Our approach yields a more action-characterized latent representations which support end-to-end dialogue policy optimization and achieves state-of-the-art success rates.These results warrant a more wide-spread use of RL in end-to-end dialogue models.
Nurul Lubis, Christian Geishauser, Michael Heck, Hsien-Chin Lin, Marco Moresi, Carel van Niekerk, Milica Gasic
COLING7
2020 TripPy: A Triple Copy Strategy for Value Independent Neural Dialog State Tracking
abstract
Michael Heck, Carel van Niekerk, Nurul Lubis, Christian Geishauser, Hsien-Chin Lin, Marco Moresi, Milica Gasic. Proceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2020.
Michael Heck, Carel van Niekerk, Nurul Lubis, Christian Geishauser, Hsien-Chin Lin, Marco Moresi, Milica Gasic
SIGdial7
2019 Curiosity-driven Reinforcement Learning for Dialogue Management
abstract
In this paper we describe the use of curiosity rewards for dialogue policy learning of goal oriented dialogues via reinforcement learning. Using curiosity improves state-action space exploration and helps overcome reward sparsity. Additionally, for goal oriented dialogues it makes sense to perform inherently curious actions in order to gain knowledge about the user goal. We show that intrinsic curiosity rewards can replace random -greedy exploration and stabilize training. The best results are achieved when curiosity rewards are combined with -greedy exploration.
Paula Wesselmann, Yen-Chen Wu, Milica Gasic
ICASSP3
2019 Tree-Structured Semantic Encoder with Knowledge Sharing for Domain Adaptation in Natural Language Generation
abstract
Domain adaptation in natural language generation (NLG) remains challenging because of the high complexity of input semantics across domains and limited data of a target domain.This is particularly the case for dialogue systems, where we want to be able to seamlessly include new domains into the conversation.Therefore, it is crucial for generation models to share knowledge across domains for the effective adaptation from one domain to another.In this study, we exploit a tree-structured semantic encoder to capture the internal structure of complex semantic representations required for multi-domain dialogues in order to facilitate knowledge sharing across domains.In addition, a layer-wise attention mechanism between the tree encoder and the decoder is adopted to further improve the model's capability.The automatic evaluation results show that our model outperforms previous methods in terms of the BLEU score and the slot error rate, in particular when the adaptation data is limited.In subjective evaluation, human judges tend to prefer the sentences generated by our model, rating them more highly on informativeness and naturalness than other systems.
Bo-Hsiang Tseng, Pawel Budzianowski, Yen-Chen Wu, Milica Gasic
SIGdial4
2019 AgentGraph: Toward Universal Dialogue Management With Structured Deep Reinforcement Learning
abstract
Dialogue policy plays an important role in task-oriented spoken dialogue systems. It determines how to respond to users. The recently proposed deep reinforcement learning (DRL) approaches have been used for policy optimization. However, these deep models are still challenging for two reasons: first, many DRL-based policies are not sample efficient; and second, most models do not have the capability of policy transfer between different domains. In this paper, we propose a universal framework, AgentGraph, to tackle these two problems. The proposed AgentGraph is the combination of graph neural network (GNN) based architecture and DRL-based algorithm. It can be regarded as one of the multi-agent reinforcement learning approaches. Each agent corresponds to a node in a graph, which is defined according to the dialogue domain ontology. When making a decision, each agent can communicate with its neighbors on the graph. Under AgentGraph framework, we further propose dual GNN-based dialogue policy, which implicitly decomposes the decision in each turn into a high-level global decision and a low-level local decision. Experiments show that AgentGraph models significantly outperform traditional reinforcement learning approaches on most of the 18 tasks of the PyDial benchmark. Moreover, when transferred from the source task to a target task, these models not only have acceptable initial performance but also converge much faster on the target task.
Lu Chen 0002, Zhi Chen 0006, Bowen Tan, Sishan Long, Milica Gasic, Kai Yu 0004
IEEE ACM Trans. Audio Speech Lang. Process.5
2018 MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling
abstract
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, Milica Gašić. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Pawel Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, Milica Gasic
EMNLP7
2018 Policy Adaptation for Deep Reinforcement Learning-Based Dialogue Management
abstract
Policy optimization is the core part of statistical dialogue management. Deep reinforcement learning has been successfully used for dialogue policy optimization for a static pre-defined domain. However, when the domain changes dynamically, e.g. a new previously unseen concept (or slot) which can be then used as a database search constraint is added, or the policy for one domain is transferred to another domain, the dialogue state space and action sets both will change. Therefore, the model structures for different domains have to be different. This makes dialogue policy adaptation/transfer challenging. Here a multi -agent dialogue policy (MADP) is proposed to tackle these problems. MADP consists of some slot-dependent agents (S-Agents) and a slot-independent agent (G-Agent). S-Agents have shared parameters in addition to private parameters for each one. During policy transfer, the shared parameters in S-Agents and the parameters in G-Agent can be directly transferred to the agents in extended/new domain. Simulation experiments showed that MADP can significantly speed up the policy learning and facilitate policy adaptation.
Lu Chen 0002, Zhi Chen 0006, Bowen Tan, Milica Gasic, Kai Yu 0004
ICASSP5
2018 Benchmarking Uncertainty Estimates with Deep Reinforcement Learning for Dialogue Policy Optimisation
abstract
In statistical dialogue management, the dialogue manager learns a policy that maps a belief state to an action for the system to perform. Efficient exploration is key to successful policy optimisation. Current deep reinforcement learning methods are very promising but rely on ε-greedy exploration, thus subjecting the user to a random choice of action during learning. Alternative approaches such as Gaussian Process SARSA (GP-SARSA) estimate uncertainties and sample actions leading to better user experience, but on the expense of a greater computational complexity. This paper examines approaches to extract uncertainty estimates from deep Q-networks (DQN) in the context of dialogue management. We perform thorough analysis of Bayes-By-Backpropagation DQN (BBQN). In addition we examine dropout, its concrete variation, bootstrapped ensemble and a-divergences as other means to extract uncertainty estimates from DQN. We find that BBQN achieves faster convergence to an optimal policy than any other method, and reaches performance comparable to the state of the art, but without the high computational complexity of GP-SARSA.
Christopher Tegho, Pawel Budzianowski, Milica Gasic
ICASSP3
2018 Feudal Dialogue Management with Jointly Learned Feature Extractors
abstract
Reinforcement learning (RL) is a promising dialogue policy optimisation approach, but traditional RL algorithms fail to scale to large domains.Recently, Feudal Dialogue Management (FDM), has shown to increase the scalability to large domains by decomposing the dialogue management decision into two steps, making use of the domain ontology to abstract the dialogue state in each step.In order to abstract the state space, however, previous work on FDM relies on handcrafted feature functions.In this work, we show that these feature functions can be learned jointly with the policy model while obtaining similar performance, even outperforming the handcrafted features in several environments and domains.
Iñigo Casanueva, Pawel Budzianowski, Stefan Ultes, Florian Kreyssig, Bo-Hsiang Tseng, Yen-Chen Wu, Milica Gasic
SIGDIAL Conference7
2018 Neural User Simulation for Corpus-based Policy Optimisation of Spoken Dialogue Systems
abstract
User Simulators are one of the major tools that enable offline training of task-oriented dialogue systems.For this task the Agenda-Based User Simulator (ABUS) is often used.The ABUS is based on handcrafted rules and its output is in semantic form.Issues arise from both properties such as limited diversity and the inability to interface a text-level belief tracker.This paper introduces the Neural User Simulator (NUS) whose behaviour is learned from a corpus and which generates natural language, hence needing a less labelled dataset than simulators generating a semantic output.In comparison to much of the past work on this topic, which evaluates user simulators on corpus-based metrics, we use the NUS to train the policy of a reinforcement learning based Spoken Dialogue System.The NUS is compared to the ABUS by evaluating the policies that were trained using the simulators.Crossmodel evaluation is performed i.e. training on one simulator and testing on the other.Furthermore, the trained policies are tested on real users.In both evaluation tasks the NUS outperformed the ABUS.
Florian Kreyssig, Iñigo Casanueva, Pawel Budzianowski, Milica Gasic
SIGDIAL Conference4
2018 Variational Cross-domain Natural Language Generation for Spoken Dialogue Systems
abstract
Cross-domain natural language generation (NLG) is still a difficult task within spoken dialogue modelling.Given a semantic representation provided by the dialogue manager, the language generator should generate sentences that convey desired information.Traditional template-based generators can produce sentences with all necessary information, but these sentences are not sufficiently diverse.With RNN-based models, the diversity of the generated sentences can be high, however, in the process some information is lost.In this work, we improve an RNN-based generator by considering latent information at the sentence level during generation using the conditional variational autoencoder architecture.We demonstrate that our model outperforms the original RNN-based generator, while yielding highly diverse sentences.In addition, our model performs better when the training data is limited.
Bo-Hsiang Tseng, Florian Kreyssig, Pawel Budzianowski, Iñigo Casanueva, Yen-Chen Wu, Stefan Ultes, Milica Gasic
SIGDIAL Conference7
2018 Addressing Objects and Their Relations: The Conversational Entity Dialogue Model
abstract
Stefan Ultes, Paweł Budzianowski, Iñigo Casanueva, Lina M. Rojas-Barahona, Bo-Hsiang Tseng, Yen-Chen Wu, Steve Young, Milica Gašić. Proceedings of the 19th Annual SIGdial Meeting on Discourse and Dialogue. 2018.
Stefan Ultes, Pawel Budzianowski, Iñigo Casanueva, Lina Maria Rojas-Barahona, Bo-Hsiang Tseng, Yen-Chen Wu, Steve J. Young, Milica Gasic
SIGDIAL Conference8
2018 Reward estimation for dialogue policy optimisation
Pei-hao Su, Milica Gasic, Steve J. Young
Comput. Speech Lang.2
2018 Sample Efficient Deep Reinforcement Learning for Dialogue Systems With Large Action Spaces
abstract
In spoken dialogue systems, we aim to deploy artificial intelligence to build automated dialogue agents that can converse with humans. A part of this effort is the policy optimization task, which attempts to find a policy describing how to respond to humans, in the form of a function taking the current state of the dialogue and returning the response of the system. In this paper, we investigate deep reinforcement learning approaches to solve this problem. Particular attention is given to actor-critic methods, off-policy reinforcement learning with experience replay, and various methods aimed at reducing the bias and variance of estimators. When combined, these methods result in the previously proposed ACER algorithm that gave competitive results in gaming environments. These environments, however, are fully observable and have a relatively small action set so, in this paper, we examine the application of ACER to dialogue policy optimization. We show that this method beats the current state of the art in deep learning approaches for spoken dialogue systems. This not only leads to a more sample efficient algorithm that can train faster, but also allows us to apply the algorithm in more difficult environments than before. We thus experiment with learning in a very large action space, which has two orders of magnitude more actions than previously considered. We find that ACER trains significantly faster than the current state of the art.
Gellért Weisz, Pawel Budzianowski, Pei-hao Su, Milica Gasic
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 A Network-based End-to-End Trainable Task-oriented Dialogue System
abstract
Tsung-Hsien Wen, David Vandyke, Nikola Mrkšić, Milica Gašić, Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, Steve Young. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Tsung-Hsien Wen, David Vandyke, Nikola Mrksic, Milica Gasic, Lina Maria Rojas-Barahona, Pei-hao Su, Stefan Ultes, Steve J. Young
EACL (1)4
2017 Domain-Independent User Satisfaction Reward Estimation for Dialogue Policy Learning
abstract
Learning suitable and well-performing dialogue behaviour in statistical spoken dialogue systems has been in the focus of research for many years. While most work which is based on reinforcement learning employs an objective measure like task success for modelling the reward signal, we propose to use a reward based on user satisfaction. We will show in simulated experiments that a live user satisfaction estimation model may be applied resulting in higher estimated satisfaction whilst achieving similar success rates. Moreover, we will show that one satisfaction estimation model which has been trained on one domain may be applied in many other domains which cover a similar task. We will verify our findings by employing the model to one of the domains for learning a policy from real users and compare its performance to policies using the user satisfaction and task success acquired directly from the users as reward.
Stefan Ultes, Pawel Budzianowski, Iñigo Casanueva, Nikola Mrksic, Lina Maria Rojas-Barahona, Pei-hao Su, Tsung-Hsien Wen, Milica Gasic, Steve J. Young
INTERSPEECH8
2017 Sub-domain Modelling for Dialogue Management with Hierarchical Reinforcement Learning
abstract
Paweł Budzianowski, Stefan Ultes, Pei-Hao Su, Nikola Mrkšić, Tsung-Hsien Wen, Iñigo Casanueva, Lina M. Rojas-Barahona, Milica Gašić. Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue. 2017.
Pawel Budzianowski, Stefan Ultes, Pei-hao Su, Nikola Mrksic, Tsung-Hsien Wen, Iñigo Casanueva, Lina Maria Rojas-Barahona, Milica Gasic
SIGDIAL Conference8
2017 DialPort, Gone Live: An Update After A Year of Development
abstract
Kyusong Lee, Tiancheng Zhao, Yulun Du, Edward Cai, Allen Lu, Eli Pincus, David Traum, Stefan Ultes, Lina M. Rojas-Barahona, Milica Gasic, Steve Young, Maxine Eskenazi. Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue. 2017.
Kyusong Lee, Yulun Du, Edward Cai, Allen Lu, Eli Pincus, David R. Traum, Stefan Ultes, Lina Maria Rojas-Barahona, Milica Gasic, Steve J. Young, Maxine Eskénazi
SIGDIAL Conference10
2017 Sample-efficient Actor-Critic Reinforcement Learning with Supervised Data for Dialogue Management
abstract
Deep reinforcement learning (RL) methods have significant potential for dialogue policy optimisation.However, they suffer from a poor performance in the early stages of learning.This is especially problematic for on-line learning with real users.Two approaches are introduced to tackle this problem.Firstly, to speed up the learning process, two sampleefficient neural networks algorithms: trust region actor-critic with experience replay (TRACER) and episodic natural actorcritic with experience replay (eNACER) are presented.For TRACER, the trust region helps to control the learning step size and avoid catastrophic model changes.For eNACER, the natural gradient identifies the steepest ascent direction in policy space to speed up the convergence.Both models employ off-policy learning with experience replay to improve sampleefficiency.Secondly, to mitigate the cold start issue, a corpus of demonstration data is utilised to pre-train the models prior to on-line reinforcement learning.Combining these two approaches, we demonstrate a practical approach to learning deep RLbased dialogue policies and demonstrate their effectiveness in a task-oriented information seeking domain.
Pei-hao Su, Pawel Budzianowski, Stefan Ultes, Milica Gasic, Steve J. Young
SIGDIAL Conference4
2017 Reward-Balancing for Statistical Spoken Dialogue Systems using Multi-objective Reinforcement Learning
abstract
Stefan Ultes, Paweł Budzianowski, Iñigo Casanueva, Nikola Mrkšić, Lina M. Rojas-Barahona, Pei-Hao Su, Tsung-Hsien Wen, Milica Gašić, Steve Young. Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue. 2017.
Stefan Ultes, Pawel Budzianowski, Iñigo Casanueva, Nikola Mrksic, Lina Maria Rojas-Barahona, Pei-hao Su, Tsung-Hsien Wen, Milica Gasic, Steve J. Young
SIGDIAL Conference8
2017 Spoken language understanding and interaction: machine learning for human-like conversational systems
Milica Gasic, Dilek Hakkani-Tür, Asli Celikyilmaz
Comput. Speech Lang.1
2017 Dialogue manager domain adaptation using Gaussian process reinforcement learning
abstract
Spoken dialogue systems allow humans to interact with machines using natural speech. As such, they have many benefits. By using speech as the primary communication medium, a computer interface can facilitate swift, human-like acquisition of information. In recent years, speech interfaces have become ever more popular, as is evident from the rise of personal assistants such as Siri, Google Now, Cortana and Amazon Alexa. Recently, data-driven machine learning methods have been applied to dialogue modelling and the results achieved for limited-domain applications are comparable to or out-perform traditional approaches. Methods based on Gaussian processes are particularly effective as they enable good models to be estimated from limited training data. Furthermore, they provide an explicit estimate of the uncertainty which is particularly useful for reinforcement learning. This article explores the additional steps that are necessary to extend these methods to model multiple dialogue domains. We show that Gaussian process reinforcement learning is an elegant framework that naturally supports a range of methods, including prior knowledge, Bayesian committee machines and multi-agent learning, for facilitating extensible and adaptable dialogue systems.
Milica Gasic, Nikola Mrksic, Lina Maria Rojas-Barahona, Pei-hao Su, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, Steve J. Young
Comput. Speech Lang.1
2017 Semantic Specialization of Distributional Word Vector Spaces using Monolingual and Cross-Lingual Constraints
abstract
We present Attract-Repel, an algorithm for improving the semantic quality of word vectors by injecting constraints extracted from lexical resources. Attract-Repel facilitates the use of constraints from mono- and cross-lingual resources, yielding semantically specialized cross-lingual vector spaces. Our evaluation shows that the method can make use of existing cross-lingual lexicons to construct high-quality vector spaces for a plethora of different languages, facilitating semantic transfer from high- to lower-resource ones. The effectiveness of our approach is demonstrated with state-of-the-art results on semantic similarity datasets in six languages. We next show that Attract-Repel-specialized vectors boost performance in the downstream task of dialogue state tracking (DST) across multiple languages. Finally, we show that cross-lingual vector spaces produced by our algorithm facilitate the training of multilingual DST models, which brings further performance improvements.
Nikola Mrksic, Ivan Vulic, Diarmuid Ó Séaghdha, Ira Leviant, Roi Reichart, Milica Gasic, Anna Korhonen, Steve J. Young
Trans. Assoc. Comput. Linguistics6
2016 On-line Active Reward Learning for Policy Optimisation in Spoken Dialogue Systems
abstract
The ability to compute an accurate reward function is essential for optimising a dialogue policy via reinforcement learning. In real-world applications, using explicit user feedback as the reward signal is often unreliable and costly to collect. This problem can be mitigated if the user's intent is known in advance or data is available to pre-train a task success predictor off-line. In practice neither of these apply for most real world applications. Here we propose an on-line learning framework whereby the dialogue policy is jointly trained alongside the reward model via active learning with a Gaussian process model. This Gaussian process operates on a continuous space dialogue representation generated in an unsupervised fashion using a recurrent neural network encoder-decoder. The experimental results demonstrate that the proposed framework is able to significantly reduce data annotation costs and mitigate noisy user feedback in dialogue policy learning.
Pei-hao Su, Milica Gasic, Nikola Mrksic, Lina Maria Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, Steve J. Young
ACL (1)2
2016 Exploiting Sentence and Context Representations in Deep Neural Models for Spoken Language Understanding
abstract
This paper presents a deep learning architecture for the semantic decoder component of a Statistical Spoken Dialogue System. In a slot-filling dialogue, the semantic decoder predicts the dialogue act and a set of slot-value pairs from a set of n-best hypotheses returned by the Automatic Speech Recognition. Most current models for spoken language understanding assume (i) word-aligned semantic annotations as in sequence taggers and (ii) delexicalisation, or a mapping of input words to domain-specific concepts using heuristics that try to capture morphological variation but that do not scale to other domains nor to language variation (e.g., morphology, synonyms, paraphrasing ). In this work the semantic decoder is trained using unaligned semantic annotations and it uses distributed semantic representation learning to overcome the limitations of explicit delexicalisation. The proposed architecture uses a convolutional neural network for the sentence representation and a long-short term memory network for the context representation. Results are presented for the publicly available DSTC2 corpus and an In-car corpus which is similar to DSTC2 but has a significantly higher word error rate (WER).
Lina Maria Rojas-Barahona, Milica Gasic, Nikola Mrksic, Pei-hao Su, Stefan Ultes, Tsung-Hsien Wen, Steve J. Young
COLING2
2016 Conditional Generation and Snapshot Learning in Neural Dialogue Systems
abstract
Tsung-Hsien Wen, Milica Gašić, Nikola Mrkšić, Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, David Vandyke, Steve Young. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2016.
Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Lina Maria Rojas-Barahona, Pei-hao Su, Stefan Ultes, David Vandyke, Steve J. Young
EMNLP2
2016 Counter-fitting Word Vectors to Linguistic Constraints
abstract
Nikola Mrkšić, Diarmuid Ó Séaghdha, Blaise Thomson, Milica Gašić, Lina M. Rojas-Barahona, Pei-Hao Su, David Vandyke, Tsung-Hsien Wen, Steve Young. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Nikola Mrksic, Diarmuid Ó Séaghdha, Blaise Thomson, Milica Gasic, Lina Maria Rojas-Barahona, Pei-hao Su, David Vandyke, Tsung-Hsien Wen, Steve J. Young
HLT-NAACL4
2016 Multi-domain Neural Network Language Generation for Spoken Dialogue Systems
abstract
Tsung-Hsien Wen, Milica Gašić, Nikola Mrkšić, Lina M. Rojas-Barahona, Pei-Hao Su, David Vandyke, Steve Young. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Lina Maria Rojas-Barahona, Pei-hao Su, David Vandyke, Steve J. Young
HLT-NAACL2
2015 Policy committee for adaptation in multi-domain spoken dialogue systems
abstract
Moving from limited-domain dialogue systems to open domain dialogue systems raises a number of challenges. One of them is the ability of the system to utilise small amounts of data from disparate domains to build a dialogue manager policy. Previous work has focused on using data from different domains to adapt a generic policy to a specific domain. Inspired by Bayesian committee machines, this paper proposes the use of a committee of dialogue policies. The results show that such a model is particularly beneficial for adaptation in multi-domain dialogue systems. The use of this model significantly improves performance compared to a single policy baseline, as confirmed by the performed real-user trial. This is the first time a dialogue policy has been trained on multiple domains on-line in interaction with real users.
Milica Gasic, Nikola Mrksic, Pei-hao Su, David Vandyke, Tsung-Hsien Wen, Steve J. Young
ASRU1
2015 Multi-domain dialogue success classifiers for policy training
abstract
We propose a method for constructing dialogue success classifiers that are capable of making accurate predictions in domains unseen during training. Pooling and adaptation are also investigated for constructing multi-domain models when data is available in the new domain. This is achieved by reformulating the features input to the recurrent neural network models introduced in [1]. Importantly, on our task of main interest, this enables policy training in a new domain without the dialogue success classifier (which forms the reinforcement learning reward function) ever having seen data from that domain before. This occurs whilst incurring only a small reduction in performance relative to developing and using an in-domain dialogue success classifier. Finally, given the motivation with these dialogue success classifiers is to enable policy training with real users, we demonstrate that these initial policy training results obtained with a simulated user carry over to learning from paid human users.
David Vandyke, Pei-hao Su, Milica Gasic, Nikola Mrksic, Tsung-Hsien Wen, Steve J. Young
ASRU3
2015 Semantically Conditioned LSTM-based Natural Language Generation for Spoken Dialogue Systems
abstract
Natural language generation (NLG) is a critical component of spoken dialogue and it has a significant impact both on usability and perceived quality.Most NLG systems in common use employ rules and heuristics and tend to generate rigid and stylised responses without the natural variation of human language.They are also not easily scaled to systems covering multiple domains and languages.This paper presents a statistical language generator based on a semantically controlled Long Short-term Memory (LSTM) structure.The LSTM generator can learn from unaligned data by jointly optimising sentence planning and surface realisation using a simple cross entropy training criterion, and language variation can be easily achieved by sampling from output candidates.With fewer heuristics, an objective evaluation in two differing test domains showed the proposed method improved performance compared to previous methods.Human judges scored the LSTM system higher on informativeness and naturalness and overall preferred it to the other systems.
Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Pei-hao Su, David Vandyke, Steve J. Young
EMNLP2
2015 Distributed dialogue policies for multi-domain statistical dialogue management
abstract
Statistical dialogue systems offer the potential to reduce costs by learning policies automatically on-line, but are not designed to scale to large open-domains. This paper proposes a hierarchical distributed dialogue architecture in which policies are organised in a class hierarchy aligned to an underlying knowledge graph. This allows a system to be deployed using a modest amount of data to train a small set of generic policies. As further data is collected, generic policies can be adapted to give in-domain performance. Using Gaussian process-based reinforcement learning, it is shown that within this framework generic policies can be constructed which provide acceptable user performance, and better performance than can be obtained using under-trained domain specific policies. It is also shown that as sufficient in-domain data becomes available, it is possible to seamlessly improve performance, without subjecting users to unacceptable behaviour during the adaptation period and without limiting the final performance compared to policies trained from scratch.
Milica Gasic, Pirros Tsiakoulis, Steve J. Young
ICASSP1
2015 Learning from real users: rating dialogue success with neural networks for reinforcement learning in spoken dialogue systems
abstract
To train a statistical spoken dialogue system (SDS) it is essential that an accurate method for measuring task success is available.To date training has relied on presenting a task to either simulated or paid users and inferring the dialogue's success by observing whether this presented task was achieved or not.Our aim however is to be able to learn from real users acting under their own volition, in which case it is non-trivial to rate the success as any prior knowledge of the task is simply unavailable.User feedback may be utilised but has been found to be inconsistent.Hence, here we present two neural network models that evaluate a sequence of turn-level features to rate the success of a dialogue.Importantly these models make no use of any prior knowledge of the user's task.The models are trained on dialogues generated by a simulated user and the best model is then used to train a policy on-line which is shown to perform at least as well as a baseline system using prior knowledge of the user's task.We note that the models should also be of interest for evaluating SDS and for monitoring a dialogue in rule-based SDS.
Pei-hao Su, David Vandyke, Milica Gasic, Nikola Mrksic, Tsung-Hsien Wen, Steve J. Young
INTERSPEECH3
2015 Hyper-parameter Optimisation of Gaussian Process Reinforcement Learning for Statistical Dialogue Management
abstract
Gaussian processes reinforcement learning provides an appealing framework for training the dialogue policy as it takes into account correlations of the objective function given different dialogue belief states, which can significantly speed up the learning.These correlations are modelled by the kernel function which may depend on hyper-parameters.So far, for real-world dialogue systems the hyperparameters have been hand-tuned, relying on the designer to adjust the correlations, or simple non-parametrised kernel functions have been used instead.Here, we examine different kernel structures and show that it is possible to optimise the hyperparameters from data yielding improved performance of the resulting dialogue policy.We confirm this in a real user trial.
Lu Chen 0002, Pei-hao Su, Milica Gasic
SIGDIAL Conference3
2015 Reward Shaping with Recurrent Neural Networks for Speeding up On-Line Policy Learning in Spoken Dialogue Systems
abstract
Statistical spoken dialogue systems have the attractive property of being able to be optimised from data via interactions with real users.However in the reinforcement learning paradigm the dialogue manager (agent) often requires significant time to explore the state-action space to learn to behave in a desirable manner.This is a critical issue when the system is trained on-line with real users where learning costs are expensive.Reward shaping is one promising technique for addressing these concerns.Here we examine three recurrent neural network (RNN) approaches for providing reward shaping information in addition to the primary (task-orientated) environmental feedback.These RNNs are trained on returns from dialogues generated by a simulated user and attempt to diffuse the overall evaluation of the dialogue back down to the turn level to guide the agent towards good behaviour faster.In both simulated and real user scenarios these RNNs are shown to increase policy learning speed.Importantly, they do not require prior knowledge of the user's goal.
Pei-hao Su, David Vandyke, Milica Gasic, Nikola Mrksic, Tsung-Hsien Wen, Steve J. Young
SIGDIAL Conference3
2015 Stochastic Language Generation in Dialogue using Recurrent Neural Networks with Convolutional Sentence Reranking
abstract
The natural language generation (NLG) component of a spoken dialogue system (SDS) usually needs a substantial amount of handcrafting or a well-labeled dataset to be trained on.These limitations add significantly to development costs and make cross-domain, multi-lingual dialogue systems intractable.Moreover, human languages are context-aware.The most natural response should be directly learned from data rather than depending on predefined syntaxes or rules.This paper presents a statistical language generator based on a joint recurrent and convolutional neural network structure which can be trained on dialogue act-utterance pairs without any semantic alignments or predefined grammar trees.Objective metrics suggest that this new model outperforms previous methods under the same experimental conditions.Results of an evaluation by human judges indicate that it produces not only high quality but linguistically varied utterances which are preferred compared to n-gram and rule-based systems.
Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Pei-hao Su, David Vandyke, Steve J. Young
SIGDIAL Conference2
2014 Dialogue context sensitive HMM-based speech synthesis
abstract
The focus of this work is speech synthesis tailored to the needs of spoken dialogue systems. More specifically, the framework of HMM-based speech synthesis is utilized to train an emphatic voice that also considers dialogue context for decision tree state clustering. To achieve this, we designed and recorded a speech corpus comprising system prompts from human-computer interaction, as well as additional prompts for slot-level emphasis. This corpus, combined with a general purpose text-to-speech one, was used to train voices using a) baseline context features, b) additional emphasis features, and c) additional dialogue context features. Both emphasis and dialogue context features are extracted from the dialogue act semantic representation. The voices were evaluated in pairs for dialogue appropriateness using a preference listening test. The results show that the emphatic voice is preferred to the baseline when emphasis markup is present, while the dialogue context-sensitive voice is preferred to the plain emphatic one when no emphasis markup is present and preferable to the baseline in both cases. This demonstrates that including dialogue context features for decision tree state clustering significantly improves the quality of the synthetic voice for dialogue.
Pirros Tsiakoulis, Catherine Breslin, Milica Gasic, Matthew Henderson, Martin Szummer, Blaise Thomson, Steve J. Young
ICASSP3
2014 Incremental on-line adaptation of POMDP-based dialogue managers to extended domains
abstract
An important property of open domain spoken dialogue systems is their ability to deal with a set of new, previously unseen, concepts introduced in the conversation. The dialogue manager must then quickly learn how to talk about the new concepts using its knowledge of the existing concepts. It has previously been shown that a single new concept could be accommodated by mapping the kernel function of a Gaussian process to incorporate an additional concept into the domain of a statistical dialogue manager. Here we present an incremental scheme which enables the domain of a dialogue manager to be repeatedly extended by recursively specifying priors in Gaussian processes. We show that it is possible to effectively double the number of concepts understood by a system providing restaurant information using only 1000 adaptation dialogues with real users.
Milica Gasic, Pirros Tsiakoulis, Catherine Breslin, Matthew Henderson, Martin Szummer, Blaise Thomson, Steve J. Young
INTERSPEECH1
2014 Inverse reinforcement learning for micro-turn management
abstract
Existing spoken dialogue systems are typically not de-signed to provide natural interaction since they impose a strict turn-taking regime in which a dialogue consists of interleaved system and user turns. To allow more responsive and natural interaction, this paper describes a system in which turn-taking decisions are taken at a more fine-grained micro-turn level. A decision-theoretic approach is then applied to optimise turn-taking control. Inverse reinforcement learning is used to cap-ture the complex but natural behaviours from human-human di-alogues and optimise interaction without specifying a reward function manually. Using a corpus of human-human interac-tion, experiments show that IRL is able to learn an effective reward function which outperforms a comparable handcrafted policy. Index Terms: dialogue management, spoken dialogue systems, inverse reinforcement learning, Markov decision processes
Catherine Breslin, Pirros Tsiakoulis, Milica Gasic, Matthew Henderson, Steve J. Young
INTERSPEECH4
2014 Dialogue context sensitive speech synthesis using factorized decision trees
abstract
This paper extends our recent work on rich context utilization for expressive speech synthesis in spoken dialogue systems in which significant improvements to the appropriateness of HMM-based synthetic voices were achieved by introducing dialogue context into the decision tree state clustering stage. Continuing in this direction, this paper investigates the performance of dialogue context-sensitive voices in different domains. The Context Adaptive Training with Factorized Decision trees (FD-CAT) approach was used to train a dialogue context-sensitive synthetic voice which was then compared to a baseline system using the standard decision tree approach. Preference-based listening tests were conducted for two different domains. The first domain concerned restaurant information and had significant coverage in the training data, while the second dealing with appointment bookings had minimal coverage in the training data. No significant preference was found for any of the voices when tested in the restaurant domain whereas in the appointment booking domain, listeners showed a statistically significant preference for the adaptively trained voice. Index Terms: HMM-based expressive speech synthesis, dia-logue context-sensitive speech synthesis, context adaptive train-ing, factorized fecision trees 1.
Pirros Tsiakoulis, Catherine Breslin, Milica Gasic, Matthew Henderson, Steve J. Young
INTERSPEECH3
2014 The PARLANCE mobile application for interactive search in English and Mandarin
abstract
Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Hugues Bouchard, Catherine Breslin, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gašić, James Henderson, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Tim Potter, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay, Boris Villazon-Terrazas, Majid Yazdani, Steve Young, Yanchao Yu. Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL). 2014.
Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Hugues Bouchard, Catherine Breslin, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gasic, James Henderson 0001, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Tim Potter, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay, Boris Villazón-Terrazas, Majid Yazdani, Steve J. Young, Yanchao Yu
SIGDIAL Conference8
2014 The use of discriminative belief tracking in POMDP-based dialogue systems
abstract
Statistical spoken dialogue systems based on Partially Observable Markov Decision Processes (POMDPs) have been shown to be more robust to speech recognition errors by maintaining a belief distribution over multiple dialogue states and making policy decisions based on the entire distribution rather than the single most likely hypothesis. To date most POMDP-based systems have used generative trackers. However, concerns about modelling accuracy have created interest in discriminative methods, and recent results from the second Dialog State Tracking Challenge (DSTC2) have shown that discriminative trackers can significantly outperform generative models in terms of tracking accuracy. The aim of this paper is to investigate the extent to which these improvements translate into improved task completion rates when incorporated into a spoken dialogue system. To do this, the Recurrent Neural Network (RNN) tracker described by Henderson et al in DSTC2 was integrated into the Cambridge statistical dialogue system and compared with the existing generative Bayesian network tracker. Using a Gaussian Process (GP) based policy, the experimental results indicate that the system using the RNN tracker performs significantly better than the system with the original Bayesian network tracker.
Matthew Henderson, Milica Gasic, Pirros Tsiakoulis, Steve J. Young
SLT3
2014 Gaussian Processes for POMDP-Based Dialogue Manager Optimization
abstract
A partially observable Markov decision process (POMDP) has been proposed as a dialog model that enables automatic optimization of the dialog policy and provides robustness to speech understanding errors. Various approximations allow such a model to be used for building real-world dialog systems. However, they require a large number of dialogs to train the dialog policy and hence they typically rely on the availability of a user simulator. They also require significant designer effort to hand-craft the policy representation. We investigate the use of Gaussian processes (GPs) in policy modeling to overcome these problems. We show that GP policy optimization can be implemented for a real world POMDP dialog manager, and in particular: 1) we examine different formulations of a GP policy to minimize variability in the learning process; 2) we find that the use of GP increases the learning rate by an order of magnitude thereby allowing learning by direct interaction with human users; and 3) we demonstrate that designer effort can be substantially reduced by basing the policy directly on the full belief space thereby avoiding ad hoc feature space modeling. Overall, the GP approach represents an important step forward towards fully automatic dialog policy optimization in real world systems.
Milica Gasic, Steve J. Young
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Continuous asr for flexible incremental dialogue
abstract
Spoken dialogue systems provide a convenient way for users to interact with a machine using only speech. However, they often rely on a rigid turn taking regime in which a voice activity detection (VAD) module is used to determine when the user is speaking and decide when is an appropriate time for the system to respond. This paper investigates replacing the VAD and discrete utterance recogniser of a conventional turn-taking system with a continuously operating recogniser that is always listening, and using the recogniser 1-best path to guide turn taking. In this way, a flexible framework for incremental dialogue management is possible. Experimental results show that it is possible to remove the VAD component and successfully use the recogniser best path to identify user speech, with more robustness to noise, potentially smaller latency times, and a reduction in overall recognition error rate compared to using the conventional approach.
Catherine Breslin, Milica Gasic, Matthew Henderson, Martin Szummer, Blaise Thomson, Pirros Tsiakoulis, Steve J. Young
ICASSP2
2013 On-line policy optimisation of Bayesian spoken dialogue systems via human interaction
abstract
A partially observable Markov decision process has been proposed as a dialogue model that enables robustness to speech recognition errors and automatic policy optimisation using reinforcement learning (RL). However, conventional RL algorithms require a very large number of dialogues, necessitating a user simulator. Recently, Gaussian processes have been shown to substantially speed up the optimisation, making it possible to learn directly from interaction with human users. However, early studies have been limited to very low dimensional spaces and the learning has exhibited convergence problems. Here we investigate learning from human interaction using the Bayesian Update of Dialogue State system. This dynamic Bayesian network based system has an optimisation space covering more than one hundred features, allowing a wide range of behaviours to be learned. Using an improved policy model and a more robust reward function, we show that stable learning can be achieved that significantly outperforms a simulator trained policy.
Milica Gasic, Catherine Breslin, Matthew Henderson, Martin Szummer, Blaise Thomson, Pirros Tsiakoulis, Steve J. Young
ICASSP1
2013 POMDP-based dialogue manager adaptation to extended domains
Milica Gasic, Catherine Breslin, Matthew Henderson, Martin Szummer, Blaise Thomson, Pirros Tsiakoulis, Steve J. Young
SIGDIAL Conference1
2013 Demonstration of the PARLANCE system: a data-driven incremental, spoken dialogue system for interactive search
Helen Hastie, Marie-Aude Aufaure, Panos Alexopoulos, Heriberto Cuayáhuitl, Nina Dethlefs, Milica Gasic, James Henderson 0001, Oliver Lemon, Xingkun Liu, Peter Mika, Nesrine Ben Mustapha, Verena Rieser, Blaise Thomson, Pirros Tsiakoulis, Yves Vanrompay
SIGDIAL Conference6
2013 POMDP-Based Statistical Spoken Dialog Systems: A Review
abstract
Statistical dialog systems (SDSs) are motivated by the need for a data-driven framework that reduces the cost of laboriously handcrafting complex dialog managers and that provides robustness against the errors created by speech recognizers operating in noisy environments. By including an explicit Bayesian model of uncertainty and by optimizing the policy via a reward-driven process, partially observable Markov decision processes (POMDPs) provide such a framework. However, exact model representation and optimization is computationally intractable. Hence, the practical application of POMDP-based systems requires efficient algorithms and carefully constructed approximations. This review article provides an overview of the current state of the art in the development of POMDP-based spoken dialog systems.
Steve J. Young, Milica Gasic, Blaise Thomson, Jason D. Williams
Proc. IEEE2
2012 The Effect of Cognitive Load on a Statistical Dialogue System
Milica Gasic, Pirros Tsiakoulis, Matthew Henderson, Blaise Thomson, Kai Yu 0004, Eli Tzirkel, Steve J. Young
SIGDIAL Conference1
2012 Policy optimisation of POMDP-based dialogue systems without state space compression
abstract
The partially observable Markov decision process (POMDP) has been proposed as a dialogue model that enables automatic improvement of the dialogue policy and robustness to speech understanding errors. It requires, however, a large number of dialogues to train the dialogue policy. Gaussian processes (GP) have recently been applied to POMDP dialogue management optimisation showing an ability to substantially increase the speed of learning. Here, we investigate this further using the Bayesian Update of Dialogue State dialogue manager. We show that it is possible to apply Gaussian processes directly to the belief state, removing the need for a parametric policy representation. In addition, the resulting policy learns significantly faster while maintaining operational performance.
Milica Gasic, Matthew Henderson, Blaise Thomson, Pirros Tsiakoulis, Steve J. Young
SLT1
2012 Discriminative spoken language understanding using word confusion networks
abstract
Current commercial dialogue systems typically use hand-crafted grammars for Spoken Language Understanding (SLU) operating on the top one or two hypotheses output by the speech recogniser. These systems are expensive to develop and they suffer from significant degradation in performance when faced with recognition errors. This paper presents a robust method for SLU based on features extracted from the full posterior distribution of recognition hypotheses encoded in the form of word confusion networks. Following [1], the system uses SVM classifiers operating on n-gram features, trained on unaligned input/output pairs. Performance is evaluated on both an off-line corpus and on-line in a live user trial. It is shown that a statistical discriminative approach to SLU operating on the full posterior ASR output distribution can substantially improve performance both in terms of accuracy and overall dialogue reward. Furthermore, additional gains can be obtained by incorporating features from the previous system output.
Matthew Henderson, Milica Gasic, Blaise Thomson, Pirros Tsiakoulis, Kai Yu 0004, Steve J. Young
SLT2
2012 N-best error simulation for training spoken dialogue systems
abstract
A recent trend in spoken dialogue research is the use of reinforcement learning to train dialogue systems in a simulated environment. Past researchers have shown that the types of errors that are simulated can have a significant effect on simulated dialogue performance. Since modern systems typically receive an N-best list of possible user utterances, it is important to be able to simulate a full N-best list of hypotheses. This paper presents a new method for simulating such errors based on logistic regression, as well as a new method for simulating the structure of N-best lists of semantics and their probabilities, based on the Dirichlet distribution. Off-line evaluations show that the new Dirichlet model results in a much closer match to the receiver operating characteristics (ROC) of the live data. Experiments also show that the logistic model gives confusions that are closer to the type of confusions observed in live situations. The hope is that these new error models will be able to improve the resulting performance of trained dialogue systems.
Blaise Thomson, Milica Gasic, Matthew Henderson, Pirros Tsiakoulis, Steve J. Young
SLT2
2011 On-line policy optimisation of spoken dialogue systems via live interaction with human subjects
abstract
Statistical dialogue models have required a large number of dialogues to optimise the dialogue policy, relying on the use of a simulated user. This results in a mismatch between training and live conditions, and significant development costs for the simulator thereby mitigating many of the claimed benefits of such models. Recent work on Gaussian process reinforcement learning, has shown that learning can be substantially accelerated. This paper reports on an experiment to learn a policy for a real-world task directly from human interaction using rewards provided by users. It shows that a usable policy can be learnt in just a few hundred dialogues without needing a user simulator and, using a learning strategy that reduces the risk of taking bad actions. The paper also investigates adaptation behaviour when the system continues learning for several thousand dialogues and highlights the need for robustness to noisy rewards.
Milica Gasic, Filip Jurcícek, Blaise Thomson, Kai Yu 0004, Steve J. Young
ASRU1
2011 Uncertainty Management for On-Line Optimisation of a POMDP-Based Large-Scale Spoken Dialogue System
abstract
The optimization of dialogue policies using reinforcement learning (RL) is now an accepted part of the state of the art in spoken dialogue systems (SDS). Yet, it is still the case that the commonly used training algorithms for SDS require a large number of dialogues and hence most systems still rely on artificial data generated by a user simulator. Optimization is therefore performed off-line before releasing the system to real users. Gaussian Processes (GP) for RL have recently been applied to dialogue systems. One advantage of GP is that they compute an explicit measure of uncertainty in the value function estimates computed during learning. In this paper, a class of novel learning strategies is described which use uncertainty to control exploration on-line. Comparisons between several exploration schemes show that significant improvements to learning speed can be obtained and that rapid and safe online optimisation is possible, even on a complex task.
Lucie Daubigney, Milica Gasic, Senthilkumar Chandramohan, Matthieu Geist, Olivier Pietquin, Steve J. Young
INTERSPEECH2
2011 Real User Evaluation of Spoken Dialogue Systems Using Amazon Mechanical Turk
abstract
This paper describes a framework for evaluation of spoken dialogue systems. Typically, evaluation of dialogue systems is performed in a controlled test environment with carefully selected and instructed users. However, this approach is very demanding. An alternative is to recruit a large group of users who evaluate the dialogue systems in a remote setting under virtually no supervision. Crowdsourcing technology, for example Amazon Mechanical Turk (AMT), provides an efficient way of recruiting subjects. This paper describes an evaluation framework for spoken dialogue systems using AMT users and compares the obtained results with a recent trial in which the systems were tested by locally recruited users. The results suggest that the use of crowdsourcing technology is feasible and it can provide reliable results. Index Terms: crowdsourcing, spoken dialogue systems, evaluation 1.
Filip Jurcícek, Simon Keizer, Milica Gasic, François Mairesse, Blaise Thomson, Kai Yu 0004, Steve J. Young
INTERSPEECH3
2010 Phrase-Based Statistical Language Generation Using Graphical Models and Active Learning
François Mairesse, Milica Gasic, Filip Jurcícek, Simon Keizer, Blaise Thomson, Kai Yu 0004, Steve J. Young
ACL2
2010 Natural belief-critic: a reinforcement algorithm for parameter estimation in statistical spoken dialogue systems
abstract
This paper presents a novel algorithm for learning parameters in statistical dialogue systems which are modelled as Partially Observable Markov Decision Processes (POMDPs). The three main components of a POMDP dialogue manager are a dialogue model representing dialogue state information; a policy which selects the system’s responses based on the inferred state; and a reward function which specifies the desired behaviour of the system. Ideally both the model parameters and the policy would be designed to maximise the reward function. However, whilst there are many techniques available for learning the optimal policy, there are no good ways of learning the optimal model parameters that scale to real-world dialogue systems. The Natural Belief-Critic (NBC) algorithm presented in this paper is a policy gradient method which offers a solution to this problem. Based on observed rewards, the algorithm estimates the natural gradient of the expected reward. The resulting gradient is then used to adapt the prior distribution of the dialogue model parameters. The algorithm is evaluated on a spoken dialogue system in the tourist information domain. The experiments show that model parameters estimated to maximise the reward function result in significantly improved performance compared to the baseline handcrafted parameters.
Filip Jurcícek, Blaise Thomson, Simon Keizer, François Mairesse, Milica Gasic, Kai Yu 0004, Steve J. Young
INTERSPEECH5
2010 Gaussian Processes for Fast Policy Optimisation of POMDP-based Dialogue Managers
Milica Gasic, Filip Jurcícek, Simon Keizer, François Mairesse, Blaise Thomson, Kai Yu 0004, Steve J. Young
SIGDIAL Conference1
2010 Parameter estimation for agenda-based user simulation
Simon Keizer, Milica Gasic, Filip Jurcícek, François Mairesse, Blaise Thomson, Kai Yu 0004, Steve J. Young
SIGDIAL Conference2
2010 Parameter learning for POMDP spoken dialogue models
abstract
The partially observable Markov decision process (POMDP) provides a popular framework for modelling spoken dialogue. This paper describes how the expectation propagation algorithm (EP) can be used to learn the parameters of the POMDP user model. Various special probability factors applicable to this task are presented, which allow the parameters be to learned when the structure of the dialogue is complex. No annotations, neither the true dialogue state nor the true semantics of user utterances, are required. Parameters optimised using the proposed techniques are shown to improve the performance of both offline transcription experiments as well as simulated dialogue management performance.
Blaise Thomson, Filip Jurcícek, Milica Gasic, Simon Keizer, François Mairesse, Kai Yu 0004, Steve J. Young
SLT3
2010 Bayesian dialogue system for the Let's Go Spoken Dialogue Challenge
abstract
This paper describes how Bayesian updates of dialogue state can be used to build a bus information spoken dialogue system. The resulting system was deployed as part of the 2010 Spoken Dialogue Challenge. The purpose of this paper is to describe the system, and provide both simulated and human evaluations of its performance. In control tests by human users, the success rate of the system was 24.5% higher than the baseline Lets Go! system.
Blaise Thomson, Kai Yu 0004, Simon Keizer, Milica Gasic, Filip Jurcícek, François Mairesse, Steve J. Young
SLT4
2010 The Hidden Information State model: A practical framework for POMDP-based spoken dialogue management
Steve J. Young, Milica Gasic, Simon Keizer, François Mairesse, Jost Schatzmann, Blaise Thomson, Kai Yu 0004
Comput. Speech Lang.2
2009 Back-off action selection in summary space-based POMDP dialogue systems
abstract
This paper deals with the issue of invalid state-action pairs in the Partially Observable Markov Decision Process (POMDP) framework, with a focus on real-world tasks where the need for approximate solutions exacerbates this problem. In particular, when modelling dialogue as a POMDP, both the state and the action space must be reduced to smaller scale summary spaces in order to make learning tractable. However, since not all actions are valid in all states, the action proposed by the policy in summary space sometimes leads to an invalid action when mapped back to master space. Some form of back-off scheme must then be used to generate an alternative action. This paper demonstrates how the value function derived during reinforcement learning can be used to order back-off actions in an N-best list. Compared to a simple baseline back-off strategy and to a strategy that extends the summary space to minimise the occurrence of invalid actions, the proposed N-best action selection scheme is shown to be significantly more robust.
Milica Gasic, Fabrice Lefèvre, Filip Jurcícek, Simon Keizer, François Mairesse, Blaise Thomson, Kai Yu 0004, Steve J. Young
ASRU1
2009 Spoken language understanding from unaligned data using discriminative classification models
abstract
While data-driven methods for spoken language understanding reduce maintenance and portability costs compared with handcrafted parsers, the collection of word-level semantic annotations for training remains a time-consuming task. A recent line of research has focused on building generative models from unaligned semantic representations, using expectation-maximisation techniques to align semantic concepts. This paper presents an efficient, simple technique that parses a semantic tree by recursively calling discriminative semantic classification models. Results show that it outperforms methods based on the Hidden Vector State model and Markov Logic Networks, while performance is close to more complex grammar induction techniques. We also show that our method is robust to speech recognition errors, by improving over a handcrafted parser previously used for dialogue data collection.
François Mairesse, Milica Gasic, Filip Jurcícek, Simon Keizer, Blaise Thomson, Kai Yu 0004, Steve J. Young
ICASSP2
2009 Probablistic modelling of F0 in unvoiced regions in HMM based speech synthesis
abstract
HMM based synthesis has attracted great interest due to its compact and flexible modelling of spectral and prosodic parameters. In this approach, short term spectra, fundamental frequency (F0) and duration are simultaneously modelled by multi-stream HMMs. However, since F0 values in unvoiced regions are normally considered as undefined, it is difficult to use standard HMMs for F0 modelling. The currently preferred solution to this is to use a multi-space distribution HMM (MSDHMM) in which discrete distributions are used for modelling the voiced/unvoiced decision and continuous Gaussian distributions are used for modelling the F0 values within the voiced regions. However, the assumption of undefined unvoiced F0 regions and the special structure of the MSDHMM lead to limitations in the accurate modelling of F0 patterns. In this paper an alternative is explored whereby unvoiced F0 values are assumed to exist and are modelled within the standard HMM framework using a globally tied distribution (GTD). Subjective evaluations show that these regular HMMs with GTD can produce significant improvements in the naturalness of the synthesised speech compared to the MSDHMM, and furthermore, the method is insensitive to the exact method used for unvoiced F0 generation.
Kai Yu 0004, Tomoki Toda, Milica Gasic, Simon Keizer, François Mairesse, Blaise Thomson, Steve J. Young
ICASSP3
2009 Transformation-based learning for semantic parsing
abstract
This paper presents a semantic parser that transforms an initial semantic hypothesis into the correct semantics by applying an ordered list of transformation rules. These rules are learnt automatically from a training corpus with no prior linguistic knowledge and no alignment between words and semantic concepts. The learning algorithm produces a compact set of rules which enables the parser to be very efficient while retaining high accuracy. We show that this parser is competitive with respect to the state-of-the-art semantic parsers on the ATIS and TownInfo tasks. Index Terms: spoken language understanding, semantics, natural language processing, transformation-based learning
Filip Jurcícek, Milica Gasic, Simon Keizer, François Mairesse, Blaise Thomson, Kai Yu 0004, Steve J. Young
INTERSPEECH2
2009 k-Nearest Neighbor Monte-Carlo Control Algorithm for POMDP-Based Dialogue Systems
Fabrice Lefèvre, Milica Gasic, Filip Jurcícek, Simon Keizer, François Mairesse, Blaise Thomson, Kai Yu 0004, Steve J. Young
SIGDIAL Conference2
2008 User study of the Bayesian update of dialogue state approach to dialogue management
abstract
This paper presents the results of a comparative user evaluation of various approaches to dialogue management. The major con-tribution is a comparison of traditional systems against a system that uses a Bayesian Update of Dialogue State approach. This approach is based on the Partially Observable Markov Decision Process (POMDP), which has previously been shown to give improved robustness in simulation experiments. Results from this paper show that the benefits demonstrated in simulation ex-periments are also obtained when testing a live system with real users.
Blaise Thomson, Milica Gasic, Simon Keizer, François Mairesse, Jost Schatzmann, Kai Yu 0004, Steve J. Young
INTERSPEECH2
2008 Evaluating semantic-level confidence scores with multiple hypotheses
abstract
In any dialogue manager, confidence scores play a central role in ensuring robust operation. Recently, dialogue managers have attempted to exploit N-best lists of alternatives for the semantics rather than the single most likely interpretation. Each alternative in the N-best list must have an associated confidence score and it is very useful to be able to evaluate the utility of these scored lists independent of the application in which they are used. This paper adapts several traditional metrics for confidence scoring to the context of the N-best semantic hypotheses output by a speech understanding system. An alternative metric, called the Item-level Cross Entropy (ICE), is proposed and is shown to have good theoretical and experimental characteristics. As an example of the use of the metrics, various simple methods for assigning confidences are discussed and evaluated. Of all the metrics tested only the ICE metric provided a consistent monotonic ranking of the various systems.
Blaise Thomson, Kai Yu 0004, Milica Gasic, Simon Keizer, François Mairesse, Jost Schatzmann, Steve J. Young
INTERSPEECH3
2008 Modelling user behaviour in the HIS-POMDP dialogue manager
abstract
In the design of spoken dialogue systems that are robust to speech recognition and interpretation errors, modelling uncertainty is crucial. Recently, Partially Observable Markov Decision Processes (POMDPs) have been shown to provide a well-founded probabilistic framework for developing such systems. This paper reports on the design and evaluation of the user act model (UAM) as part of the Hidden Information State (HIS) POMDP dialogue manager. Within this system, the UAM represents the probability of a user producing a certain dialogue act, given the last system act and the dialogue state. Its design is domain-independent and founded on the notions of adjacency pairs and dialogue act preconditions. Experimental evaluation results on both simulated and real data show that the UAM plays a significant role in improving robustness, but it requires that the N-best lists of user act hypotheses and their confidence scores are of good quality.
Simon Keizer, Milica Gasic, François Mairesse, Blaise Thomson, Kai Yu 0004, Steve J. Young
SLT2