VLDB 2026 Research / reviewers in the wild / expert
Simon Keizer
dblp:80/6099
· DBLP profile ↗
37ranked-venue papers
9as first author
8since 2021 · last 2024
0000-0003-0173-8966ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 8 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
Mohan Li, Simon Keizer, Rama Sanand Doddipatla |
INTERSPEECH | 2 |
| 2024 | WHISMA: A Speech-LLM to Perform Zero-Shot Spoken Language UnderstandingabstractSpeech large language models (speech-LLMs) integrate speech and text-based foundation models to provide a unified framework for handling a wide range of downstream tasks. In this paper, we introduce WHISMA, a speech-LLM tailored for spoken language understanding (SLU) that demonstrates robust performance in various zero-shot settings. WHISMA combines the speech encoder from Whisper with the Llama-3 LLM, and is fine-tuned in a parameter-efficient manner on a comprehensive collection of SLU-related datasets. Our experiments show that WHISMA significantly improves the zero-shot slot filling performance on the SLURP benchmark, achieving a relative gain of 26.6% compared to the current state-of-the-art model. Furthermore, to evaluate WHISMA’s generalisation capabilities to unseen domains, we develop a new task-agnostic benchmark named SLU-GLUE. The evaluation results indicate that WHISMA outperforms an existing speech-LLM (Qwen-Audio) with a relative gain of 33.0%. Mohan Li, Cong-Thanh Do, Simon Keizer, Youmna Farag, Svetlana Stoyanchev, Rama Sanand Doddipatla |
SLT | 3 |
| 2024 | Entity Resolution in Situated Dialog With Unimodal and Multimodal TransformersabstractIn this work we address the entity resolution task for situated multimodal dialog investigating how a unimodal approach, which uses only textual information as input (representing visual attributes as text), compares to a multimodal system, which processes both text and visual information. We analyze two of the top performing models presented in the Tenth Dialog Systems Technology Challenge and propose modifications that enhance their performance on the multimodal coreference resolution task. We evaluate these approaches on in- and out-of-domain settings by training the models on the fashion domain and testing on the furniture domain, and vice-versa, to assess the generalizability of the models. Through systematic analysis, we show that while both systems achieve similar performance on in-domain scenarios, the multimodal system generalizes better to out-of-domain settings. A combination strategy of enhanced unimodal and multimodal systems achieves F1 = 0.80 (5% absolute gain compared to the best performing system). Finally, human performance on the same task is evaluated on a small subset, suggesting that the performance of the current automatic models is on par with people on this task. Alejandro Santorum Varela, Svetlana Stoyanchev, Simon Keizer, Rama Sanand Doddipatla, Kate M. Knill |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Combining Structured and Unstructured Knowledge in an Interactive Search Dialogue SystemabstractUsers of interactive search dialogue systems specify their preferences with natural language utterances.However, a schema-driven system is limited to handling the preferences that correspond to the predefined database content.In this work, we present a methodology for extending a schema-driven interactive search dialogue system with the ability to handle unconstrained user preferences.Using unsupervised semantic similarity metrics and text snippets associated with the search items, the system identifies suitable items for the user's unconstrained natural language query.In a crowd-sourced evaluation, the users were asked to chat with our extended restaurant search system.Based on objective metrics and subjective user ratings, we demonstrate the feasibility of using this unsupervised low latency approach to extend a schema-driven search dialogue system to handle unconstrained user preferences. Svetlana Stoyanchev, Suraj Pandey, Simon Keizer, Norbert Braunschweiler, Rama Sanand Doddipatla |
SIGDIAL | 3 |
| 2022 | Factors in Emotion Recognition With Deep Learning Models Using Speech and Text on Multiple CorporaabstractEmotion recognition performance of deep learning models is influenced by multiple factors such as acoustic condition, textual content, style of emotion expression (e.g. acted, natural), etc. In this paper, multiple factors are analysed by training and evaluating state-of-the-art deep learning models using the input modalities speech, text, and their combination across 6 emotional speech corpora. A novel deep learning model architecture is presented that further improves the state-of-the-art in multimodal emotion recognition with speech and text on the IEMOCAP corpus. Results from models trained on individual corpora show that combining speech and text improves performance only on corpora where the text of utterances varies across different emotions, while it reduced performance on corpora with fixed text expressed in different emotions, where the speech-only models performed better. Further, cross-corpus investigations are presented to understand the robustness to changing acoustic and textual content. Results show that models perform significantly better in matched conditions in particular single corpus models perform better than multi-corpus models, with the latter showing a tendency to be more robust to acoustic variations, while performance still depends on characteristics of both training corpora and test corpus. Norbert Braunschweiler, Rama Sanand Doddipatla, Simon Keizer, Svetlana Stoyanchev |
IEEE Signal Process. Lett. | 3 |
| 2021 | A Study on Cross-Corpus Speech Emotion Recognition and Data AugmentationabstractModels that can handle a wide range of speakers and acoustic conditions are essential in speech emotion recognition (SER). Often, these models tend to show mixed results when presented with speakers or acoustic conditions that were not visible during training. This paper investigates the impact of cross-corpus data complementation and data augmentation on the performance of SER models in matched (test-set from same corpus) and mismatched (test-set from different corpus) conditions. Investigations using six emotional speech corpora that include single and multiple speakers as well as variations in emotion style (acted, elicited, natural) and recording conditions are presented. Observations show that, as expected, models trained on single corpora perform best in matched conditions while performance decreases between 10-40% in mismatched conditions, depending on corpus specific features. Models trained on mixed corpora can be more stable in mismatched contexts, and the performance reductions range from 1 to 8% when compared with single corpus models in matched conditions. Data augmentation yields additional gains up to 4% and seem to benefit mismatched conditions more than matched ones. Norbert Braunschweiler, Rama Sanand Doddipatla, Simon Keizer, Svetlana Stoyanchev |
ASRU | 3 |
| 2021 | Dialogue Strategy Adaptation to New Action Sets Using Multi-Dimensional ModellingabstractA major bottleneck for building statistical spoken dialogue systems for new domains and applications is the need for large amounts of training data. To address this problem, we adopt the multi-dimensional approach to dialogue management and evaluate its potential for transfer learning. Specifically, we exploit pre-trained task-independent policies to speed up training for an extended task-specific action set, in which the single summary action for requesting a slot is replaced by multiple slot-specific request actions. Policy optimisation and evaluation experiments using an agenda-based user simulator show that with limited training data, much better performance levels can be achieved when using the proposed multi-dimensional adaptation method. We confirm this improvement in a crowd-sourced human user evaluation of our spoken dialogue system, comparing partially trained policies. The multi-dimensional system (with adaptation on limited training data in the target scenario) outperforms the one-dimensional baseline (without adaptation on the same amount of training data) by 7% perceived success rate. Simon Keizer, Norbert Braunschweiler, Svetlana Stoyanchev, Rama Sanand Doddipatla |
ASRU | 1 |
| 2021 | Action State Update Approach to Dialogue ManagementabstractUtterance interpretation is one of the main functions of a dialogue manager, which is the key component of a dialogue system. We propose the action state update approach (ASU) for utterance interpretation, featuring a statistically trained binary classifier used to detect dialogue state update actions in the text of a user utterance. Our goal is to interpret referring expressions in user input without a domain-specific natural language understanding component. For training the model, we use active learning to automatically select simulated training examples. With both user-simulated and interactive human evaluations, we show that the ASU approach successfully interprets user utterances in a dialogue system, including those with referring expressions. Svetlana Stoyanchev, Simon Keizer, Rama Sanand Doddipatla |
ICASSP | 2 |
| 2020 | The ISO Standard for Dialogue Act Annotation, Second EditionabstractISO standard 24617-2 for dialogue act annotation, established in 2012, has in the past few years been used both in corpus annotation and in the design of components for spoken and multimodal dialogue systems. This has brought some inaccuracies and undesirbale limitations of the standard to light, which are addressed in a proposed second edition. This second edition allows a more accurate annotation of dependence relations and rhetorical relations in dialogue. Following the ISO 24617-4 principles of semantic annotation, and borrowing ideas from EmotionML, a triple-layered plug-in mechanism is introduced which allows dialogue act descriptions to be enriched with information about their semantic content, about accompanying emotions, and other information, and allows the annotation scheme to be customised by adding application-specific dialogue act types. Harry Bunt, Volha Petukhova, Emer Gilmartin, Catherine Pelachaud, Alex Chengyu Fang, Simon Keizer, Laurent Prévot 0001 |
LREC | 6 |
| 2019 | User Evaluation of a Multi-dimensional Statistical Dialogue SystemabstractWe present the first complete spoken dialogue system driven by a multi-dimensional statistical dialogue manager.This framework has been shown to substantially reduce data needs by leveraging domain-independent dimensions, such as social obligations or feedback, which (as we show) can be transferred between domains.In this paper, we conduct a user study and show that the performance of a multi-dimensional system, which can be adapted from a source domain, is equivalent to that of a one-dimensional baseline, which can only be trained from scratch. Simon Keizer, Ondrej Dusek, Xingkun Liu, Verena Rieser |
SIGdial | 1 |
| 2015 | Learning Trading Negotiations Using Manually and Automatically Labelled DataabstractStrategic conversational agents often need to trade resources with their opponent conversants -- and trading strategically can lead to better results. While rule-based or supervised agents can be used for such a purpose, here we explore a learning approach based on automatically labelled examples from human players for automatic trading in the game of Settlers of Catan. Our experiments are based on data collected from human players trading in text-based natural language. We compare the performance of Bayes Nets, Conditional Random Fields, and Random Forests on the task of ranking trading offers, trained from both manually labelled and automatically labelled data. Our experimental results show that our best agent trained on automatic labels outperformed its counterpart trained on manual labels (with moderate annotator agreement) in terms of (a) predicting human trading negotiations better, and (b) winning more games. Heriberto Cuayáhuitl, Simon Keizer, Oliver Lemon |
ICTAI | 2 |
| 2014 | Towards action selection under uncertainty for a socially aware robot bartenderabstractWe describe how the state representation of a socially aware robot is being extended to handle uncertainty. It incorporates the full range of information provided by the input sensors, including the confidence of all hypotheses. We also show how the Interaction Manager is being updated to make use of the extended representation. Mary Ellen Foster, Simon Keizer, Oliver Lemon |
HRI | 2 |
| 2014 | Handling uncertain input in multi-user human-robot interactionabstractIn this paper we present results from a user evaluation of a robot bartender system which handles state uncertainty derived from speech input by using belief tracking and generating appropriate clarification questions. We present a combination of state estimation and action selection components in which state uncertainty is tracked and exploited, and compare it to a baseline version that uses standard speech recognition confidence score thresholds instead of belief tracking. The results suggest that users are served fewer incorrect drinks when the uncertainty is retained in the state. Simon Keizer, Mary Ellen Foster, Andre Gaschler, Manuel Giuliani, Amy Isard, Oliver Lemon |
RO-MAN | 1 |
| 2014 | Evaluating a social multi-user interaction model using a Nao robotabstractThis paper presents results from a user evaluation of a robot bartender system, which supports social engagement and interaction with multiple customers. The system is a Nao-based alternative version of an existing robot bartender developed in the JAMES project [1]. The Nao-based version has given us a local experimentation platform, allowing us to focus on social multi-user interaction rather than the robot technology of object manipulation. We will describe the design of the Nao-based system and discuss the differences with the original JAMES system. In a recent evaluation of the JAMES system with real users, a trained and a hand-coded version of the action selection policy were compared [2]. Here we present results from a similar comparative user evaluation on the Nao-based system, which confirm the conclusions of the previous experiment and provide further evidence in favour of the trained action selection mechanism. Task success was found to be almost 20% higher with the trained policy, with interaction times being about 10% shorter. Participants also rated the trained system as significantly more natural, more understanding, and better at providing appropriate attention. Simon Keizer, Pantelis Kastoris, Mary Ellen Foster, Amol A. Deshmukh, Oliver Lemon |
RO-MAN | 1 |
| 2014 | Real user evaluation of a POMDP spoken dialogue system using automatic belief compression
Paul A. Crook, Simon Keizer, Wenshuo Tang, Oliver Lemon |
Comput. Speech Lang. | 2 |
| 2014 | Natural Language Generation as Incremental Planning Under Uncertainty: Adaptive Information Presentation for Statistical Dialogue SystemsabstractWe present and evaluate a novel approach to natural language generation (NLG) in statistical spoken dialogue systems (SDS) using a data-driven statistical optimization framework for incremental information presentation (IP), where there is a trade-off to be solved between presenting “enough" information to the user while keeping the utterances short and understandable. The trained IP model is adaptive to variation from the current generation context (e.g. a user and a non-deterministic sentence planner), and it incrementally adapts the IP policy at the turn level. Reinforcement learning is used to automatically optimize the IP policy with respect to a data-driven objective function. In a case study on presenting restaurant information, we show that an optimized IP strategy trained on Wizard-of-Oz data outperforms a baseline mimicking the wizard behavior in terms of total reward gained. The policy is then also tested with real users, and improves on a conventional hand-coded IP strategy used in a deployed SDS in terms of overall task success. The evaluation found that the trained IP strategy significantly improves dialogue task completion for real users, with up to a 8.2% increase in task success. This methodology also provides new insights into the nature of the IP problem, which has previously been treated as a module following dialogue management with no access to lower-level context features (e.g. from a surface realizer and/or speech synthesizer). Verena Rieser, Oliver Lemon, Simon Keizer |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Machine Learning for Social Multiparty Human-Robot InteractionabstractWe describe a variety of machine-learning techniques that are being applied to social multiuser human--robot interaction using a robot bartender in our scenario. We first present a data-driven approach to social state recognition based on supervised learning . We then describe an approach to social skills execution—that is, action selection for generating socially appropriate robot behavior—which is based on reinforcement learning , using a data-driven simulation of multiple users to train execution policies for social skills. Next, we describe how these components for social state recognition and skills execution have been integrated into an end-to-end robot bartender system, and we discuss the results of a user evaluation. Finally, we present an alternative unsupervised learning framework that combines social state recognition and social skills execution based on hierarchical Dirichlet processes and an infinite POMDP interaction manager. The models make use of data from both human--human interactions collected in a number of German bars and human--robot interactions recorded in the evaluation of an initial version of the system. Simon Keizer, Mary Ellen Foster, Oliver Lemon |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2013 | Training and evaluation of an MDP model for social multi-user human-robot interaction
Simon Keizer, Mary Ellen Foster, Oliver Lemon, Andre Gaschler, Manuel Giuliani |
SIGDIAL Conference | 1 |
| 2011 | Real User Evaluation of Spoken Dialogue Systems Using Amazon Mechanical TurkabstractThis paper describes a framework for evaluation of spoken dialogue systems. Typically, evaluation of dialogue systems is performed in a controlled test environment with carefully selected and instructed users. However, this approach is very demanding. An alternative is to recruit a large group of users who evaluate the dialogue systems in a remote setting under virtually no supervision. Crowdsourcing technology, for example Amazon Mechanical Turk (AMT), provides an efficient way of recruiting subjects. This paper describes an evaluation framework for spoken dialogue systems using AMT users and compares the obtained results with a recent trial in which the systems were tested by locally recruited users. The results suggest that the use of crowdsourcing technology is feasible and it can provide reliable results. Index Terms: crowdsourcing, spoken dialogue systems, evaluation 1. Filip Jurcícek, Simon Keizer, Milica Gasic, François Mairesse, Blaise Thomson, Kai Yu 0004, Steve J. Young |
INTERSPEECH | 2 |
| 2011 | Spoken Dialog Challenge 2010: Comparison of Live and Control Test Results
Alan W. Black, Susanne Burger, Alistair Conkie, Helen Hastie, Simon Keizer, Oliver Lemon, Nicolas Merigaud, Gabriel Parent, Gabriel Schubiner, Blaise Thomson, Jason D. Williams, Kai Yu 0004, Steve J. Young, Maxine Eskénazi |
SIGDIAL Conference | 5 |
| 2010 | Phrase-Based Statistical Language Generation Using Graphical Models and Active Learning
François Mairesse, Milica Gasic, Filip Jurcícek, Simon Keizer, Blaise Thomson, Kai Yu 0004, Steve J. Young |
ACL | 4 |
| 2010 | Natural belief-critic: a reinforcement algorithm for parameter estimation in statistical spoken dialogue systemsabstractThis paper presents a novel algorithm for learning parameters in statistical dialogue systems which are modelled as Partially Observable Markov Decision Processes (POMDPs). The three main components of a POMDP dialogue manager are a dialogue model representing dialogue state information; a policy which selects the system’s responses based on the inferred state; and a reward function which specifies the desired behaviour of the system. Ideally both the model parameters and the policy would be designed to maximise the reward function. However, whilst there are many techniques available for learning the optimal policy, there are no good ways of learning the optimal model parameters that scale to real-world dialogue systems. The Natural Belief-Critic (NBC) algorithm presented in this paper is a policy gradient method which offers a solution to this problem. Based on observed rewards, the algorithm estimates the natural gradient of the expected reward. The resulting gradient is then used to adapt the prior distribution of the dialogue model parameters. The algorithm is evaluated on a spoken dialogue system in the tourist information domain. The experiments show that model parameters estimated to maximise the reward function result in significantly improved performance compared to the baseline handcrafted parameters. Filip Jurcícek, Blaise Thomson, Simon Keizer, François Mairesse, Milica Gasic, Kai Yu 0004, Steve J. Young |
INTERSPEECH | 3 |
| 2010 | Gaussian Processes for Fast Policy Optimisation of POMDP-based Dialogue Managers
Milica Gasic, Filip Jurcícek, Simon Keizer, François Mairesse, Blaise Thomson, Kai Yu 0004, Steve J. Young |
SIGDIAL Conference | 3 |
| 2010 | Parameter estimation for agenda-based user simulation
Simon Keizer, Milica Gasic, Filip Jurcícek, François Mairesse, Blaise Thomson, Kai Yu 0004, Steve J. Young |
SIGDIAL Conference | 1 |
| 2010 | Parameter learning for POMDP spoken dialogue modelsabstractThe partially observable Markov decision process (POMDP) provides a popular framework for modelling spoken dialogue. This paper describes how the expectation propagation algorithm (EP) can be used to learn the parameters of the POMDP user model. Various special probability factors applicable to this task are presented, which allow the parameters be to learned when the structure of the dialogue is complex. No annotations, neither the true dialogue state nor the true semantics of user utterances, are required. Parameters optimised using the proposed techniques are shown to improve the performance of both offline transcription experiments as well as simulated dialogue management performance. Blaise Thomson, Filip Jurcícek, Milica Gasic, Simon Keizer, François Mairesse, Kai Yu 0004, Steve J. Young |
SLT | 4 |
| 2010 | Bayesian dialogue system for the Let's Go Spoken Dialogue ChallengeabstractThis paper describes how Bayesian updates of dialogue state can be used to build a bus information spoken dialogue system. The resulting system was deployed as part of the 2010 Spoken Dialogue Challenge. The purpose of this paper is to describe the system, and provide both simulated and human evaluations of its performance. In control tests by human users, the success rate of the system was 24.5% higher than the baseline Lets Go! system. Blaise Thomson, Kai Yu 0004, Simon Keizer, Milica Gasic, Filip Jurcícek, François Mairesse, Steve J. Young |
SLT | 3 |
| 2010 | The Hidden Information State model: A practical framework for POMDP-based spoken dialogue management
Steve J. Young, Milica Gasic, Simon Keizer, François Mairesse, Jost Schatzmann, Blaise Thomson, Kai Yu 0004 |
Comput. Speech Lang. | 3 |
| 2009 | Back-off action selection in summary space-based POMDP dialogue systemsabstractThis paper deals with the issue of invalid state-action pairs in the Partially Observable Markov Decision Process (POMDP) framework, with a focus on real-world tasks where the need for approximate solutions exacerbates this problem. In particular, when modelling dialogue as a POMDP, both the state and the action space must be reduced to smaller scale summary spaces in order to make learning tractable. However, since not all actions are valid in all states, the action proposed by the policy in summary space sometimes leads to an invalid action when mapped back to master space. Some form of back-off scheme must then be used to generate an alternative action. This paper demonstrates how the value function derived during reinforcement learning can be used to order back-off actions in an N-best list. Compared to a simple baseline back-off strategy and to a strategy that extends the summary space to minimise the occurrence of invalid actions, the proposed N-best action selection scheme is shown to be significantly more robust. Milica Gasic, Fabrice Lefèvre, Filip Jurcícek, Simon Keizer, François Mairesse, Blaise Thomson, Kai Yu 0004, Steve J. Young |
ASRU | 4 |
| 2009 | Spoken language understanding from unaligned data using discriminative classification modelsabstractWhile data-driven methods for spoken language understanding reduce maintenance and portability costs compared with handcrafted parsers, the collection of word-level semantic annotations for training remains a time-consuming task. A recent line of research has focused on building generative models from unaligned semantic representations, using expectation-maximisation techniques to align semantic concepts. This paper presents an efficient, simple technique that parses a semantic tree by recursively calling discriminative semantic classification models. Results show that it outperforms methods based on the Hidden Vector State model and Markov Logic Networks, while performance is close to more complex grammar induction techniques. We also show that our method is robust to speech recognition errors, by improving over a handcrafted parser previously used for dialogue data collection. François Mairesse, Milica Gasic, Filip Jurcícek, Simon Keizer, Blaise Thomson, Kai Yu 0004, Steve J. Young |
ICASSP | 4 |
| 2009 | Probablistic modelling of F0 in unvoiced regions in HMM based speech synthesisabstractHMM based synthesis has attracted great interest due to its compact and flexible modelling of spectral and prosodic parameters. In this approach, short term spectra, fundamental frequency (F0) and duration are simultaneously modelled by multi-stream HMMs. However, since F0 values in unvoiced regions are normally considered as undefined, it is difficult to use standard HMMs for F0 modelling. The currently preferred solution to this is to use a multi-space distribution HMM (MSDHMM) in which discrete distributions are used for modelling the voiced/unvoiced decision and continuous Gaussian distributions are used for modelling the F0 values within the voiced regions. However, the assumption of undefined unvoiced F0 regions and the special structure of the MSDHMM lead to limitations in the accurate modelling of F0 patterns. In this paper an alternative is explored whereby unvoiced F0 values are assumed to exist and are modelled within the standard HMM framework using a globally tied distribution (GTD). Subjective evaluations show that these regular HMMs with GTD can produce significant improvements in the naturalness of the synthesised speech compared to the MSDHMM, and furthermore, the method is insensitive to the exact method used for unvoiced F0 generation. Kai Yu 0004, Tomoki Toda, Milica Gasic, Simon Keizer, François Mairesse, Blaise Thomson, Steve J. Young |
ICASSP | 4 |
| 2009 | Transformation-based learning for semantic parsingabstractThis paper presents a semantic parser that transforms an initial semantic hypothesis into the correct semantics by applying an ordered list of transformation rules. These rules are learnt automatically from a training corpus with no prior linguistic knowledge and no alignment between words and semantic concepts. The learning algorithm produces a compact set of rules which enables the parser to be very efficient while retaining high accuracy. We show that this parser is competitive with respect to the state-of-the-art semantic parsers on the ATIS and TownInfo tasks. Index Terms: spoken language understanding, semantics, natural language processing, transformation-based learning Filip Jurcícek, Milica Gasic, Simon Keizer, François Mairesse, Blaise Thomson, Kai Yu 0004, Steve J. Young |
INTERSPEECH | 3 |
| 2009 | k-Nearest Neighbor Monte-Carlo Control Algorithm for POMDP-Based Dialogue Systems
Fabrice Lefèvre, Milica Gasic, Filip Jurcícek, Simon Keizer, François Mairesse, Blaise Thomson, Kai Yu 0004, Steve J. Young |
SIGDIAL Conference | 4 |
| 2008 | User study of the Bayesian update of dialogue state approach to dialogue managementabstractThis paper presents the results of a comparative user evaluation of various approaches to dialogue management. The major con-tribution is a comparison of traditional systems against a system that uses a Bayesian Update of Dialogue State approach. This approach is based on the Partially Observable Markov Decision Process (POMDP), which has previously been shown to give improved robustness in simulation experiments. Results from this paper show that the benefits demonstrated in simulation ex-periments are also obtained when testing a live system with real users. Blaise Thomson, Milica Gasic, Simon Keizer, François Mairesse, Jost Schatzmann, Kai Yu 0004, Steve J. Young |
INTERSPEECH | 3 |
| 2008 | Evaluating semantic-level confidence scores with multiple hypothesesabstractIn any dialogue manager, confidence scores play a central role in ensuring robust operation. Recently, dialogue managers have attempted to exploit N-best lists of alternatives for the semantics rather than the single most likely interpretation. Each alternative in the N-best list must have an associated confidence score and it is very useful to be able to evaluate the utility of these scored lists independent of the application in which they are used. This paper adapts several traditional metrics for confidence scoring to the context of the N-best semantic hypotheses output by a speech understanding system. An alternative metric, called the Item-level Cross Entropy (ICE), is proposed and is shown to have good theoretical and experimental characteristics. As an example of the use of the metrics, various simple methods for assigning confidences are discussed and evaluated. Of all the metrics tested only the ICE metric provided a consistent monotonic ranking of the various systems. Blaise Thomson, Kai Yu 0004, Milica Gasic, Simon Keizer, François Mairesse, Jost Schatzmann, Steve J. Young |
INTERSPEECH | 4 |
| 2008 | Modelling user behaviour in the HIS-POMDP dialogue managerabstractIn the design of spoken dialogue systems that are robust to speech recognition and interpretation errors, modelling uncertainty is crucial. Recently, Partially Observable Markov Decision Processes (POMDPs) have been shown to provide a well-founded probabilistic framework for developing such systems. This paper reports on the design and evaluation of the user act model (UAM) as part of the Hidden Information State (HIS) POMDP dialogue manager. Within this system, the UAM represents the probability of a user producing a certain dialogue act, given the last system act and the dialogue state. Its design is domain-independent and founded on the notions of adjacency pairs and dialogue act preconditions. Experimental evaluation results on both simulated and real data show that the UAM plays a significant role in improving robustness, but it requires that the N-best lists of user act hypotheses and their confidence scores are of good quality. Simon Keizer, Milica Gasic, François Mairesse, Blaise Thomson, Kai Yu 0004, Steve J. Young |
SLT | 1 |
| 2007 | Dialogue act recognition under uncertainty using Bayesian networksabstractAbstract In this paper we discuss the task of dialogue act recognition as a part of interpreting user utterances in context. To deal with the uncertainty that is inherent in natural language processing in general and dialogue act recognition in particular we use machine learning techniques to train classifiers from corpus data. These classifiers make use of both lexical features of the (Dutch) keyboard-typed utterances in the corpus used, and context features in the form of dialogue acts of previous utterances. In particular, we consider probabilistic models in the form of Bayesian networks to be proposed as a more general framework for dealing with uncertainty in the dialogue modelling process. Simon Keizer, Rieks op den Akker |
Nat. Lang. Eng. | 1 |
| 2005 | From question answering to spoken dialogue: towards an information search assistant for interactive multimodal information extractionabstractThis paper gives an overview of issues related to extending simple question answering (QA) with dialogue capabilities, when designing a multimodal interactive information extraction system for a large, though restricted, domain. We present the way in which these issues are approached in the IMIX program. The IMIX demonstrator system, under development in this program, may be considered the most difficult case of QA, answering non-factoid questions in a large domain, and accepting speech input as well. We describe our approach to the addition of dialogue capabilities to this system. We will look at QA from a dialogue system perspective and from a HCI perspective, and consider the consequences of our choice of interaction metaphor, the `information search assistant'. Rieks op den Akker, Harry Bunt, Simon Keizer, Boris W. van Schooten |
INTERSPEECH | 3 |