Joakim Gustafson

dblp:28/6376 · DBLP profile ↗
← Back
92ranked-venue papers
11as first author
23since 2021 · last 2025
0000-0002-0397-6442ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 76 · 11 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 54 · 7 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 18 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 2
YearPublicationVenuePosition
2025 From Static to Dynamic: Enhancing AAC with Generative Imagery and Zero-Shot TTS
Juliana Francis, Joakim Gustafson, Éva Székely
INTERSPEECH2
2025 VoiceQualityVC: A Voice Conversion System for Studying the Perceptual Effects of Voice Quality in Speech
Harm Lameris, Joakim Gustafson, Éva Székely
INTERSPEECH2
2025 Towards Adaptable and Intelligible Speech Synthesis in Noisy Environments
Lubos Marcinek, Jonas Beskow, Joakim Gustafson
INTERSPEECH3
2025 Role of Reasoning in LLM Enjoyment Detection: Evaluation Across Conversational Levels for Human-Robot Interaction
abstract
User enjoyment is central to developing conversational AI systems that can recover from failures and maintain interest over time. However, existing approaches often struggle to detect subtle cues that reflect user experience. Large Language Models (LLMs) with reasoning capabilities have outperformed standard models on various other tasks, suggesting potential benefits for enjoyment detection. This study investigates whether models with reasoning capabilities outperform standard models when assessing enjoyment in a human-robot dialogue corpus at both turn and interaction levels. Results indicate that reasoning capabilities have complex, model-dependent effects rather than universal benefits. While performance was nearly identical at the interaction level (0.44 vs 0.43), reasoning models substantially outperformed at the turn level (0.42 vs 0.36). Notably, LLMs correlated better with users’ self-reported enjoyment metrics than human annotators, despite achieving lower accuracy against human consensus ratings. Analysis revealed distinctive error patterns: non-reasoning models showed bias toward positive ratings at the turn level, while both model types exhibited central tendency bias at the interaction level. These findings suggest that reasoning should be applied selectively based on model architecture and assessment context, with assessment granularity significantly influencing relative effectiveness.
Lubos Marcinek, Bahar Irfan, Gabriel Skantze, André Pereira 0001, Joakim Gustafson
SIGDIAL5
2024 The Role of Creaky Voice in Turn Taking and the Perception of Speaker Stance: Experiments Using Controllable TTS
abstract
Recent advancements in spontaneous text-to-speech (TTS) have enabled the realistic synthesis of creaky voice, a voice quality known for its diverse pragmatic and paralinguistic functions. In this study, we used synthesized creaky voice in perceptual tests, to explore how listeners without formal training perceive two distinct types of creaky voice. We annotated a spontaneous speech corpus using creaky voice detection tools and modified a neural TTS engine with a creaky phonation embedding to control the presence of creaky phonation in the synthesized speech. We performed an objective analysis using a creak detection tool which revealed significant differences in creaky phonation levels between the two creaky voice types and modal voice. Two subjective listening experiments were performed to investigate the effect of creaky voice on perceived certainty, valence, sarcasm, and turn finality. Participants rated non-positional creak as less certain, less positive, and more indicative of turn finality, while positional creak was rated significantly more turn final compared to modal phonation.
Harm Lameris, Éva Székely, Joakim Gustafson
LREC/COLING3
2024 Revisiting Three Text-to-Speech Synthesis Experiments with a Web-Based Audience Response System
abstract
In order to investigate the strengths and weaknesses of Audience Response System (ARS) in text-to-speech synthesis (TTS) evaluations, we revisit three previously published TTS studies and perform an ARS-based evaluation on the stimuli used in each study. The experiments are performed with a participant pool of 39 respondents, using a web-based tool that emulates an ARS experiment. The results of the first experiment confirms that ARS is highly useful for evaluating long and continuous stimuli, particularly if we wish for a diagnostic result rather than a single overall metric, while the second and third experiments highlight weaknesses in ARS with unsuitable materials as well as the importance of framing and instruction when conducting ARS-based evaluation.
Christina Tånnander, Jens Edlund, Joakim Gustafson
LREC/COLING3
2024 Multimodal User Enjoyment Detection in Human-Robot Conversation: The Power of Large Language Models
abstract
Enjoyment is a crucial yet complex indicator of positive user experience in Human-Robot Interaction (HRI). While manual enjoyment annotation is feasible, developing reliable automatic detection methods remains a challenge. This paper investigates a multimodal approach to automatic enjoyment annotation for HRI conversations, leveraging large language models (LLMs), visual, audio, and temporal cues. Our findings demonstrate that both text-only and multimodal LLMs with carefully designed prompts can achieve performance comparable to human annotators in detecting user enjoyment. Furthermore, results reveal a stronger alignment between LLM-based annotations and user self-reports of enjoyment compared to human annotators. While multimodal supervised learning techniques did not improve all of our performance metrics, they could successfully replicate human annotators and highlighted the importance of visual and audio cues in detecting subtle shifts in enjoyment. This research demonstrates the potential of LLMs for real-time enjoyment detection, paving the way for adaptive companion robots that can dynamically enhance user experiences.
André Pereira 0001, Lubos Marcinek, Jura Miniota, Sofia Thunberg, Erik Lagerstedt, Joakim Gustafson, Gabriel Skantze, Bahar Irfan
ICMI6
2024 ConnecTone: a modular AAC system prototype with contextual generative text prediction and style-adaptive conversational TTS
Juliana Francis, Éva Székely, Joakim Gustafson
INTERSPEECH3
2024 CreakVC: a voice conversion tool for modulating creaky voice
Harm Lameris, Joakim Gustafson, Éva Székely
INTERSPEECH2
2024 Contextual Interactive Evaluation of TTS Models in Dialogue Systems
Éva Székely, Joakim Gustafson
INTERSPEECH3
2023 Casual chatter or speaking up? Adjusting articulatory effort in generation of speech and animation for conversational characters
abstract
Embodied conversational agents and social robots need to be able to generate spontaneous behavior in order to be believable in social interactions. We present a system that can generate spontaneous speech with supporting lip movements. The conversational TTS voice is trained on a podcast corpus that has been prosodically tagged (f0, speaking rate and energy) and transcribed (including tokens for breathing, fillers and laughter). We introduce a speech animation algorithm where articulatory effort can be adjusted. The speech animation is driven by time-stamped phonemes obtained from the internal alignment attention map of the TTS system, and we use prominence estimates from the synthesised speech waveform to modulate the lip- and jaw movements accordingly.
Joakim Gustafson, Éva Székely, Simon Alexanderson, Jonas Beskow
FG1
2023 Prosody-Controllable Spontaneous TTS with Neural HMMS
abstract
Spontaneous speech has many affective and pragmatic functions that are interesting and challenging to model in TTS. However, the presence of reduced articulation, fillers, repetitions, and other disfluencies in spontaneous speech make the text and acoustics less aligned than in read speech, which is problematic for attention-based TTS. We propose a TTS architecture that can rapidly learn to speak from small and irregular datasets, while also reproducing the diversity of expressive phenomena present in spontaneous speech. Specifically, we add utterance-level prosody control to an existing neural HMM-based TTS system which is capable of stable, monotonic alignments for spontaneous speech. We objectively evaluate control accuracy and perform perceptual tests that demonstrate that prosody control does not degrade synthesis quality. To exemplify the power of combining prosody control and ecologically valid data for reproducing intricate spontaneous speech phenomena, we evaluate the system’s capability of synthesizing two types of creaky voice.
Harm Lameris, Shivam Mehta, Gustav Eje Henter, Joakim Gustafson, Éva Székely
ICASSP4
2023 Automatic Evaluation of Turn-taking Cues in Conversational Speech Synthesis
Erik Ekstedt, Éva Székely, Joakim Gustafson, Gabriel Skantze
INTERSPEECH4
2023 Pardon my disfluency: The impact of disfluency effects on the perception of speaker competence and confidence
Ambika Kirkland, Joakim Gustafson, Éva Székely
INTERSPEECH2
2023 Beyond Style: Synthesizing Speech with Pragmatic Functions
Harm Lameris, Joakim Gustafson, Éva Székely
INTERSPEECH2
2023 Prosody-controllable Gender-ambiguous Speech Synthesis: A Tool for Investigating Implicit Bias in Speech Perception
Éva Székely, Joakim Gustafson, Ilaria Torre 0002
INTERSPEECH2
2023 So-to-Speak: An Exploratory Platform for Investigating the Interplay between Style and Prosody in TTS
Éva Székely, Joakim Gustafson
INTERSPEECH3
2023 Generation of speech and facial animation with controllable articulatory effort for amusing conversational characters
abstract
Engaging embodied conversational agents need to generate expressive behavior in order to be believable in socializing interactions. We present a system that can generate spontaneous speech with supporting lip movements. The neural conversational TTS voice is trained on a multi-style speech corpus that has been prosodically tagged (pitch and speaking rate) and transcribed (including tokens for breathing, fillers and laughter). We introduce a speech animation algorithm where articulatory effort can be adjusted. The facial animation is driven by time-stamped phonemes and prominence estimates from the synthesised speech waveform to modulate the lip-and jaw movements accordingly. In objective evaluations we show that the system is able to generate speech and facial animation that vary in articulation effort. In subjective evaluations we compare our conversational TTS system's capability to deliver jokes with a commercial TTS. Both system succeeded equally good.
Joakim Gustafson, Éva Székely, Jonas Beskow
IVA1
2023 Hi robot, it's not what you say, it's how you say it
abstract
Many robots use their voice to communicate with people in spoken language but the voices commonly used for robots are often optimized for transactional interactions, rather than social ones. This can limit their ability to create engaging and natural interactions. To address this issue, we designed a spontaneous text-to-speech tool and used it to author natural and spontaneous robot speech. A crowdsourcing evaluation methodology is proposed to compare this type of speech to natural speech and state-of-the-art text-to-speech technology, both in disembodied and embodied form. We created speech samples in a naturalistic setting of people playing tabletop games and conducted a user study evaluating Naturalness, Intelligibility, Social Impression, Prosody, and Perceived Intelligence. The speech samples were chosen to represent three contexts that are common in tabletop games and the contexts were introduced to the participants that evaluated the speech samples. The study results show that the proposed evaluation methodology allowed for a robust analysis that successfully compared the different conditions. Moreover, the spontaneous voice met our target design goal of being perceived as more natural than a leading commercial text-to-speech.
Jura Miniota, Jonas Beskow, Joakim Gustafson, Éva Székely, André Pereira 0001
RO-MAN4
2022 Where's the uh, hesitation? The interplay between filled pause location, speech rate and fundamental frequency in perception of confidence
Ambika Kirkland, Harm Lameris, Éva Székely, Joakim Gustafson
INTERSPEECH4
2022 Evaluating Sampling-based Filler Insertion with Spontaneous TTS
abstract
Inserting fillers (such as “um”, “like”) to clean speech text has a rich history of study. One major application is to make dialogue systems sound more spontaneous. The ambiguity of filler occurrence and inter-speaker difference make both modeling and evaluation difficult. In this paper, we study sampling-based filler insertion, a simple yet unexplored approach to inserting fillers. We propose an objective score called Filler Perplexity (FPP). We build three models trained on two single-speaker spontaneous corpora, and evaluate them with FPP and perceptual tests. We implement two innovations in perceptual tests, (1) evaluating filler insertion on dialogue systems output, (2) synthesizing speech with neural spontaneous TTS engines. FPP proves to be useful in analysis but does not correlate well with perceptual MOS. Perceptual results show little difference between compared filler insertion models including with ground-truth, which may be due to the ambiguity of what is good filler insertion and a strong neural spontaneous TTS that produces natural speech irrespective of input. Results also show preference for filler-inserted speech synthesized with spontaneous TTS. The same test using TTS based on read speech obtains the opposite results, which shows the importance of using spontaneous TTS in evaluating filler insertions. Audio samples: www.speech.kth.se/tts-demos/LREC22
Joakim Gustafson, Éva Székely
LREC2
2021 A Systematic Cross-Corpus Analysis of Human Reactions to Robot Conversational Failures
abstract
In this paper, we analyze multimodal behavioral responses to robot failures across different tasks. Two multimodal datasets are examined in which humans interact with guided-task robots in task-oriented dialogues. In both datasets, the robots simulated failures of conversational breakdown and miscommunication typically observed in human-robot interactions. We closely examine human reactions to these failures looking at facial and acoustic features. Our analyses identify the significant behavioral features for automatic detection of such failures in interaction. We also examine human responses to different types of robot failures and if failures occurred early or late in the interaction cause variation in the responses. Our findings indicate that several nonverbal behaviors are consistently present in responses to robots’ failures, e.g., gaze and speech prosody, whereas, linguistic features appear to be task-dependent. We discuss how these findings may generalize to other tasks, and how autonomous robots may identify opportunities to detect and recover from failures in interactions with humans.
Dimosthenis Kontogiorgos, Minh Tran 0004, Joakim Gustafson, Mohammad Soleymani 0001
ICMI3
2021 Integrated Speech and Gesture Synthesis
abstract
Text-to-speech and co-speech gesture synthesis have until now been treated as separate areas by two different research communities, and applications merely stack the two technologies using a simple system-level pipeline. This can lead to modeling inefficiencies and may introduce inconsistencies that limit the achievable naturalness. We propose to instead synthesize the two modalities in a single model, a new problem we call integrated speech and gesture synthesis (ISG). We also propose a set of models modified from state-of-the-art neural speech-synthesis engines to achieve this goal. We evaluate the models in three carefully-designed user studies, two of which evaluate the synthesized speech and gesture in isolation, plus a combined study that evaluates the models like they will be used in real-world applications – speech and gesture presented together. The results show that participants rate one of the proposed integrated synthesis models as being as good as the state-of-the-art pipeline system we compare against, in all three tests. The model is able to achieve this with faster synthesis time and greatly reduced parameter count compared to the pipeline system, illustrating some of the potential benefits of treating speech and gesture synthesis together as a single, unified problem.
Simon Alexanderson, Joakim Gustafson, Jonas Beskow, Gustav Eje Henter, Éva Székely
ICMI3
2020 Embodiment Effects in Interactions with Failing Robots
abstract
The increasing use of robots in real-world applications will inevitably cause users to encounter more failures in interactions. While there is a longstanding effort in bringing human-likeness to robots, how robot embodiment affects users' perception of failures remains largely unexplored. In this paper, we extend prior work on robot failures by assessing the impact that embodiment and failure severity have on people's behaviours and their perception of robots. Our findings show that when using a smart-speaker embodiment, failures negatively affect users' intention to frequently interact with the device, however not when using a human-like robot embodiment. Additionally, users significantly rate the human-like robot higher in terms of perceived intelligence and social presence. Our results further suggest that in higher severity situations, human-likeness is distracting and detrimental to the interaction. Drawing on quantitative findings, we discuss benefits and drawbacks of embodiment in robot failures that occur in guided tasks.
Dimosthenis Kontogiorgos, Sanne van Waveren, Olle Wallberg, André Pereira 0001, Iolanda Leite, Joakim Gustafson
CHI6
2020 Behavioural Responses to Robot Conversational Failures
abstract
Humans and robots will increasingly collaborate in domestic environments which will cause users to encounter more failures in interactions. Robots should be able to infer conversational failures by detecting human users' behavioural and social signals. In this paper, we study and analyse these behavioural cues in response to robot conversational failures. Using a guided task corpus, where robot embodiment and time pressure are manipulated, we ask human annotators to estimate whether user affective states differ during various types of robot failures. We also train a random forest classifier to detect whether a robot failure has occurred and compare results to human annotator benchmarks. Our findings show that human-like robots augment users' reactions to failures, as shown in users' visual attention, in comparison to non-human-like smart-speaker embodiments. The results further suggest that speech behaviours are utilised more in responses to failures when non-human-like designs are present. This is particularly important to robot failure detection mechanisms that may need to consider the robot's physical design in its failure detection model.
Dimosthenis Kontogiorgos, André Pereira 0001, Boran Sahindal, Sanne van Waveren, Joakim Gustafson
HRI5
2020 Effects of Different Interaction Contexts when Evaluating Gaze Models in HRI
abstract
We previously introduced a responsive joint attention system that uses multimodal information from users engaged in a spatial reasoning task with a robot and communicates joint attention via the robot's gaze behavior. An initial evaluation of our system with adults showed it to improve users' perceptions of the robot's social presence. To investigate the repeatability of our prior findings across settings and populations, here we conducted two further studies employing the same gaze system with the same robot and task but in different contexts: evaluation of the system with external observers and evaluation with children. The external observer study suggests that third-person perspectives over videos of gaze manipulations can be used either as a manipulation check before committing to costly real-time experiments or to further establish previous findings. However, the replication of our original adults study with children in school did not confirm the effectiveness of our gaze manipulation, suggesting that different interaction contexts can affect the generalizability of results in human-robot interaction gaze studies.
André Pereira 0001, Catharine Oertel, Leonor Fermoselle, Joseph Mendelson, Joakim Gustafson
HRI5
2020 Breathing and Speech Planning in Spontaneous Speech Synthesis
abstract
Breathing and speech planning in spontaneous speech are coordinated processes, often exhibiting disfluent patterns. While synthetic speech is not subject to respiratory needs, integrating breath into synthesis has advantages for naturalness and recall. At the same time, a synthetic voice reproducing disfluent breathing patterns learned from the data can be problematic. To address this, we first propose training stochastic TTS on a corpus of overlapping breath-group bigrams, to take context into account. Next, we introduce an unsupervised automatic annotation of likely-disfluent breath events, through a product-of-experts model that combines the output of two breath- event predictors, each using complementary information and operating in opposite directions. This annotation enables creating an automatically-breathing spontaneous speech synthesiser with a more fluent breathing style. A subjective evaluation on two spoken genres (impromptu and rehearsed) found the proposed system to be preferred over the baseline approach treating all breath events the same.
Éva Székely, Gustav Eje Henter, Jonas Beskow, Joakim Gustafson
ICASSP4
2020 Chinese Whispers: A Multimodal Dataset for Embodied Language Grounding
abstract
In this paper, we introduce a multimodal dataset in which subjects are instructing each other how to assemble IKEA furniture. Using the concept of ‘Chinese Whispers’, an old children’s game, we employ a novel method to avoid implicit experimenter biases. We let subjects instruct each other on the nature of the task: the process of the furniture assembly. Uncertainty, hesitations, repairs and self-corrections are naturally introduced in the incremental process of establishing common ground. The corpus consists of 34 interactions, where each subject first assembles and then instructs. We collected speech, eye-gaze, pointing gestures, and object movements, as well as subjective interpretations of mutual understanding, collaboration and task recall. The corpus is of particular interest to researchers who are interested in multimodal signals in situated dialogue, especially in referential communication and the process of language grounding.
Dimosthenis Kontogiorgos, Elena Sibirtseva, Joakim Gustafson
LREC3
2020 Augmented Prompt Selection for Evaluation of Spontaneous Speech Synthesis
abstract
By definition, spontaneous speech is unscripted and created on the fly by the speaker. It is dramatically different from read speech, where the words are authored as text before they are spoken. Spontaneous speech is emergent and transient, whereas text read out loud is pre-planned. For this reason, it is unsuitable to evaluate the usability and appropriateness of spontaneous speech synthesis by having it read out written texts sampled from for example newspapers or books. Instead, we need to use transcriptions of speech as the target - something that is much less readily available. In this paper, we introduce Starmap, a tool allowing developers to select a varied, representative set of utterances from a spoken genre, to be used for evaluation of TTS for a given domain. The selection can be done from any speech recording, without the need for transcription. The tool uses interactive visualisation of prosodic features with t-SNE, along with a tree-based algorithm to guide the user through thousands of utterances and ensure coverage of a variety of prompts. A listening test has shown that with a selection of genre-specific utterances, it is possible to show significant differences across genres between two synthetic voices built from spontaneous speech.
Éva Székely, Jens Edlund, Joakim Gustafson
LREC3
2019 The Effects of Embodiment and Social Eye-Gaze in Conversational Agents
Dimosthenis Kontogiorgos, Gabriel Skantze, André Pereira 0001, Joakim Gustafson
CogSci4
2019 Casting to Corpus: Segmenting and Selecting Spontaneous Dialogue for Tts with a Cnn-lstm Speaker-dependent Breath Detector
abstract
This paper considers utilising breaths to create improved spontaneous-speech corpora for conversational text-to-speech from found audio recordings such as dialogue podcasts. Breaths are of interest since they relate to prosody and speech planning and are independent of language and transcription. Specifically, we propose a semi-supervised approach where a fraction of coarsely annotated data is used to train a convolutional and recurrent speaker-specific breath detector operating on spectrograms and zero-crossing rate. The classifier output is used to find target-speaker breath groups (audio segments delineated by breaths) and subsequently select those that constitute clean utterances appropriate for a synthesis corpus. An application to 11 hours of raw podcast audio extracts 1969 utterances (106 minutes), 87% of which are clean and correctly segmented. This outperforms a baseline that performs integrated VAD and speaker attribution without accounting for breaths.
Éva Székely, Gustav Eje Henter, Joakim Gustafson
ICASSP3
2019 Estimating Uncertainty in Task-Oriented Dialogue
abstract
Situated multimodal systems that instruct humans need to handle user uncertainties, as expressed in behaviour, and plan their actions accordingly. Speakers’ decision to reformulate or repair previous utterances depends greatly on the listeners’ signals of uncertainty. In this paper, we estimate uncertainty in a situated guided task, as leveraged in non-verbal cues expressed by the listener, and predict that the speaker will reformulate their utterance. We use a corpus where people instruct how to assemble furniture, and extract their multimodal features. While uncertainty is in cases verbally expressed, most instances are expressed non-verbally, which indicates the importance of multimodal approaches. In this work, we present a model for uncertainty estimation. Our findings indicate that uncertainty estimation from non-verbal cues works well, and can exceed human annotator performance when verbal features cannot be perceived.
Dimosthenis Kontogiorgos, André Pereira 0001, Joakim Gustafson
ICMI3
2019 Off the Cuff: Exploring Extemporaneous Speech Delivery with TTS
Éva Székely, Gustav Eje Henter, Jonas Beskow, Joakim Gustafson
INTERSPEECH4
2019 Spontaneous Conversational Speech Synthesis from Found Data
abstract
Synthesising spontaneous speech is a difficult task due to disfluencies, high variability and syntactic conventions different from those of written language. Using found data, as opposed to lab-rec ...
Éva Székely, Gustav Eje Henter, Jonas Beskow, Joakim Gustafson
INTERSPEECH4
2019 Responsive Joint Attention in Human-Robot Interaction
abstract
Joint attention has been shown to be not only crucial for human-human interaction but also human-robot interaction. Joint attention can help to make cooperation more efficient, support disambiguation in instances of uncertainty and make interactions appear more natural and familiar. In this paper, we present an autonomous gaze system that uses multimodal perception capabilities to model responsive joint attention mechanisms. We investigate the effects of our system on people's perception of a robot within a problem-solving task. Results from a user study suggest that responsive joint attention mechanisms evoke higher perceived feelings of social presence on scales that regard the direction of the robot's perception.
André Pereira 0001, Catharine Oertel, Leonor Fermoselle, Joe Mendelson, Joakim Gustafson
IROS5
2019 The Effects of Anthropomorphism and Non-verbal Social Behaviour in Virtual Assistants
abstract
The adoption of virtual assistants is growing at a rapid pace. However, these assistants are not optimised to simulate key social aspects of human conversational environments. Humans are intellectually biased toward social activity when facing anthropomorphic agents or when presented with subtle social cues. In this paper, we test whether humans respond the same way to assistants in guided tasks, when in different forms of embodiment and social behaviour. In a within-subject study (N=30), we asked subjects to engage in dialogue with a smart speaker and a social robot. We observed shifting of interactive behaviour, as shown in behavioural and subjective measures. Our findings indicate that it is not always favourable for agents to be anthropomorphised or to communicate with nonverbal cues. We found a trade-off between task performance and perceived sociability when controlling for anthropomorphism and social behaviour.
Dimosthenis Kontogiorgos, André Pereira 0001, Olle Andersson, Marco Koivisto, Elena Gonzalez Rabal, Ville Vartiainen, Joakim Gustafson
IVA7
2018 Interactive, Collaborative Robots: Challenges and Opportunities
abstract
Robotic technology has transformed manufacturing industry ever since the first industrial robot was put in use in the beginning of the 60s. The challenge of developing flexible solutions where production lines can be quickly re-planned, adapted and structured for new or slightly changed products is still an important open problem. Industrial robots today are still largely preprogrammed for their tasks, not able to detect errors in their own performance or to robustly interact with a complex environment and a human worker. The challenges are even more serious when it comes to various types of service robots. Full robot autonomy, including natural interaction, learning from and with human, safe and flexible performance for challenging tasks in unstructured environments will remain out of reach for the foreseeable future. In the envisioned future factory setups, home and office environments, humans and robots will share the same workspace and perform different object manipulation tasks in a collaborative manner. We discuss some of the major challenges of developing such systems and provide examples of the current state of the art.
Danica Kragic, Joakim Gustafson, Hakan Karaoguz, Patric Jensfelt, Robert Krug 0002
IJCAI2
2018 Crowdsourced Multimodal Corpora Collection Tool
Patrik Jonell, Catharine Oertel, Dimosthenis Kontogiorgos, Jonas Beskow, Joakim Gustafson
LREC5
2018 A Multimodal Corpus for Mutual Gaze and Joint Attention in Multiparty Situated Interaction
Dimosthenis Kontogiorgos, Vanya Avramova, Simon Alexanderson, Patrik Jonell, Catharine Oertel, Jonas Beskow, Gabriel Skantze, Joakim Gustafson
LREC8
2018 A Comparison of Visualisation Methods for Disambiguating Verbal Requests in Human-Robot Interaction
abstract
Picking up objects requested by a human user is a common task in human-robot interaction. When multiple objects match the user's verbal description, the robot needs to clarify which object the user is referring to before executing the action. Previous research has focused on perceiving user's multimodal behaviour to complement verbal commands or minimising the number of follow up questions to reduce task time. In this paper, we propose a system for reference disambiguation based on visualisation and compare three methods to disambiguate natural language instructions. In a controlled experiment with a YuMi robot, we investigated realtime augmentations of the workspace in three conditions - head-mounted display, projector, and a monitor as the baseline - using objective measures such as time and accuracy, and subjective measures like engagement, immersion, and display interference. Significant differences were found in accuracy and engagement between the conditions, but no differences were found in task time. Despite the higher error rates in the head-mounted display condition, participants found that modality more engaging than the other two, but overall showed preference for the projector condition over the monitor and head-mounted display conditions.
Elena Sibirtseva, Dimosthenis Kontogiorgos, Olov Nykvist, Hakan Karaoguz, Iolanda Leite, Joakim Gustafson, Danica Kragic
RO-MAN6
2017 Controlling Prominence Realisation in Parametric DNN-Based Speech Synthesis
abstract
This work aims to improve text-To-speech synthesis forWikipedia by advancing and implementing models of prosodic prominence. We propose a new system architecture with explicit prominence modeling a ...
Zofia Malisz, Harald Berthelsen, Jonas Beskow, Joakim Gustafson
INTERSPEECH4
2017 Crowd-Sourced Design of Artificial Attentive Listeners
abstract
Feedback generation is an important component of humanhuman communication. Humans can choose to signal support, understanding, agreement or also sceptiscism by means of feedback tokens. Many studie ...
Catharine Oertel, Patrik Jonell, Dimosthenis Kontogiorgos, Joseph Mendelson, Jonas Beskow, Joakim Gustafson
INTERSPEECH6
2017 Synthesising Uncertainty: The Interplay of Vocal Effort and Hesitation Disfluencies
abstract
As synthetic voices become more flexible, and conversational systems gain more potential to adapt to the environmental and social situation, the question needs to be examined, how different modific ...
Éva Székely, Joseph Mendelson, Joakim Gustafson
INTERSPEECH3
2017 Crowd-Powered Design of Virtual Attentive Listeners
Patrik Jonell, Catharine Oertel, Dimosthenis Kontogiorgos, Jonas Beskow, Joakim Gustafson
IVA5
2016 Towards building an attentive artificial listener: on the perception of attentiveness in audio-visual feedback tokens
abstract
Current dialogue systems typically lack a variation of audio-visual feedback tokens. Either they do not encompass feedback tokens at all, or only support a limited set of stereotypical functions. However, this does not mirror the subtleties of spontaneous conversations. If we want to be able to build an artificial listener, as a first step towards building an empathetic artificial agent, we also need to be able to synthesize more subtle audio-visual feedback tokens. In this study, we devised an array of monomodal and multimodal binary comparison perception tests and experiments to understand how different realisations of verbal and visual feedback tokens influence third-party perception of the degree of attentiveness. This allowed us to investigate i) which features (amplitude, frequency, duration...) of the visual feedback influences attentiveness perception; ii) whether visual or verbal backchannels are perceived to be more attentive iii) whether the fusion of unimodal tokens with low perceived attentiveness increases the degree of perceived attentiveness compared to unimodal tokens with high perceived attentiveness taken alone; iv) the automatic ranking of audio-visual feedback token in terms of conveyed degree of attentiveness.
Catharine Oertel, José Lopes 0001, Yu Yu 0003, Kenneth Alberto Funes Mora, Joakim Gustafson, Alan W. Black, Jean-Marc Odobez
ICMI5
2016 Towards Building an Attentive Artificial Listener: On the Perception of Attentiveness in Feedback Utterances
abstract
Towards Building an Attentive Artificial Listener: On the Perception of Attentiveness in Feedback Utterances
Catharine Oertel, Joakim Gustafson, Alan W. Black
INTERSPEECH2
2016 Hidden Resources ― Strategies to Acquire and Exploit Potential Spoken Language Resources in National Archives
Jens Edlund, Joakim Gustafson
LREC2
2015 Deciphering the Silent Participant: On the Use of Audio-Visual Cues for the Classification of Listener Categories in Group Discussions
abstract
Estimating a silent participant's degree of engagement and his role within a group discussion can be challenging, as there are no speech related cues available at the given time. Having this information available, however, can provide important insights into the dynamics of the group as a whole. In this paper, we study the classification of listeners into several categories (attentive listener, side participant and bystander). We devised a thin-sliced perception test where subjects were asked to assess listener roles and engagement levels in 15-second video-clips taken from a corpus of group interviews. Results show that humans are usually able to assess silent participant roles. Using the annotation to identify from a set of multimodal low-level features, such as past speaking activity, backchannels (both visual and verbal), as well as gaze patterns, we could identify the features which are able to distinguish between different listener categories. Moreover, the results show that many of the audio-visual effects observed on listeners in dyadic interactions, also hold for multi-party interactions. A preliminary classifier achieves an accuracy of 64 %.
Catharine Oertel, Kenneth Alberto Funes Mora, Joakim Gustafson, Jean-Marc Odobez
ICMI3
2015 Detecting repetitions in spoken dialogue systems using phonetic distances
abstract
This paper addresses the problem of automatic detection of re-peated turns in Spoken Dialogue Systems. Repetitions can be a symptom of problematic communication between users and systems. Such repetitions are often due to speech recognition errors, which in turn makes it hard to use speech recognition to detect repetitions. We present an approach to detect rep-etition using the phonetic distance to find the best alignment between turns in the same dialogue. The alignment score ob-tained is combined with different features to improve repeti-tion detection. To evaluate the method proposed we compare several alignment techniques from edit distance to DTW-based distance, previously used in Spoken-Term detection tasks. We also compare two different methods to compute the phonetic distance: the first one using the phoneme sequence, and the second one using the distance between the phone posterior vec-tors. Two different datasets were used in this evaluation: a bus-schedule information system (in English) and a call routing system (in Swedish). The results show that approaches using phoneme distances over-perform approaches using Levenshtein distances between ASR outputs for repetition detection. Index Terms: spoken dialogue systems, repetition detection, phonetic distance
José Lopes 0001, Giampiero Salvi, Gabriel Skantze, Alberto Abad, Joakim Gustafson, Fernando Batista, Raveesh Meena, Isabel Trancoso
INTERSPEECH5
2015 Automatic Detection of Miscommunication in Spoken Dialogue Systems
abstract
In this paper, we present a data-driven approach for detecting instances of miscommunication in dialogue system interactions.A range of generic features that are both automatically extractable and manually annotated were used to train two models for online detection and one for offline analysis.Online detection could be used to raise the error awareness of the system, whereas offline detection could be used by a system designer to identify potential flaws in the dialogue design.In experimental evaluations on system logs from three different dialogue systems that vary in their dialogue strategy, the proposed models performed substantially better than the majority class baseline models.
Raveesh Meena, José Lopes 0001, Gabriel Skantze, Joakim Gustafson
SIGDIAL Conference4
2014 Human-robot collaborative tutoring using multiparty multimodal spoken dialogue
abstract
In this paper, we describe a project that explores a novel experimental setup towards building a spoken, multi-modally rich, and human-like multiparty tutoring robot. A human-robot interaction setup is designed, and a human-human dialogue corpus is collected. The corpus targets the development of a dialogue system platform to study verbal and nonverbal tutoring strategies in multiparty spoken interactions with robots which are capable of spoken dialogue. The dialogue task is centered on two participants involved in a dialogue aiming to solve a card-ordering game. Along with the participants sits a tutor (robot) that helps the participants perform the task, and organizes and balances their interaction. Different multimodal signals captured and auto-synchronized by different audio-visual capture technologies, such as a microphone array, Kinects, and video cameras, were coupled with manual annotations. These are used build a situated model of the interaction based on the participants personalities, their state of attention, their conversational engagement and verbal dominance, and how that is correlated with the verbal and visual feed-back, turn-management, and conversation regulatory actions generated by the tutor. Driven by the analysis of the corpus, we will show also the detailed design methodologies for an affective, and multimodally rich dialogue system that allows the robot to measure incrementally the attention states, and the dominance for each participant, allowing the robot head Furhat to maintain a well-coordinated, balanced, and engaging conversation, that attempts to maximize the agreement and the contribution to solve the task.
Samer Al Moubayed, Jonas Beskow, Bajibabu Bollepalli, Joakim Gustafson, Ahmed Hussen Abdelaziz, Martin Johansson, Maria Koutsombogera, José Lopes 0001, Jekaterina Novikova, Catharine Oertel, Gabriel Skantze, Kalin Stefanov, Gül Varol
HRI4
2014 A comparative evaluation of vocoding techniques for HMM-based laughter synthesis
abstract
This paper presents an experimental comparison of various leading vocoders for the application of HMM-based laughter synthesis. Four vocoders, commonly used in HMM-based speech synthesis, are used in copy-synthesis and HMM-based synthesis of both male and female laughter. Subjective evaluations are conducted to assess the performance of the vocoders. The results show that all vocoders perform relatively well in copy-synthesis. In HMM-based laughter synthesis using original phonetic transcriptions, all synthesized laughter voices were significantly lower in quality than copy-synthesis, indicating a challenging task and room for improvements. Interestingly, two vocoders using rather simple and robust excitation modeling performed the best, indicating that robustness in speech parameter extraction and simple parameter representation in statistical modeling are key factors in successful laughter synthesis.
Bajibabu Bollepalli, Jérôme Urbain, Tuomo Raitio, Joakim Gustafson, Hüseyin Çakmak
ICASSP4
2014 Crowdsourcing Street-level Geographic Information Using a Spoken Dialogue System
abstract
We present a technique for crowdsourcing street-level geographic information using spoken natural language.In particular, we are interested in obtaining first-person-view information about what can be seen from different positions in the city.This information can then for example be used for pedestrian routing services.The approach has been tested in the lab using a fully implemented spoken dialogue system, and has shown promising results.
Raveesh Meena, Johan Boye, Gabriel Skantze, Joakim Gustafson
SIGDIAL Conference4
2014 Data-driven models for timing feedback responses in a Map Task dialogue system
Raveesh Meena, Gabriel Skantze, Joakim Gustafson
Comput. Speech Lang.3
2013 Analysis of gaze and speech patterns in three-party quiz game interaction
abstract
In order to understand and model the dynamics between interaction phenomena such as gaze and speech in face-to-face multiparty interaction between humans, we need large quantities of reliable, objective data of such interactions. To date, this type of data is in short supply. We present a data collection setup using automated, objective techniques in which we capture the gaze and speech patterns of triads deeply engaged in a high-stakes quiz game. The resulting corpus consists of five one-hour recordings, and is unique in that it makes use of three state-of-the-art gaze trackers (one per subject) in combination with a state-of-theart conical microphone array designed to capture roundtable meetings. Several video channels are also included. In this paper we present the obstacles we encountered and the possibilities afforded by a synchronised, reliable combination of large-scale multi-party speech and gaze data, and an overview of the first analyses of the data. Index Terms: multimodal corpus, multiparty dialogue, gaze patterns, multiparty gaze.
Samer Al Moubayed, Jens Edlund, Joakim Gustafson
INTERSPEECH3
2013 The Map Task Dialogue System: A Test-bed for Modelling Human-Like Dialogue
Raveesh Meena, Gabriel Skantze, Joakim Gustafson
SIGDIAL Conference3
2013 A Data-driven Model for Timing Feedback in a Map Task Dialogue System
Raveesh Meena, Gabriel Skantze, Joakim Gustafson
SIGDIAL Conference3
2013 Semi-supervised methods for exploring the acoustics of simple productive feedback
Daniel Neiberg, Giampiero Salvi, Joakim Gustafson
Speech Commun.3
2012 Multimodal multiparty social interaction with the furhat head
abstract
We will show in this demonstrator an advanced multimodal and multiparty spoken conversational system using Furhat, a robot head based on projected facial animation. Furhat is a human-like interface that utilizes facial animation for physical robot heads using back-projection. In the system, multimodality is enabled using speech and rich visual input signals such as multi-person real-time face tracking and microphone tracking. The demonstrator will showcase a system that is able to carry out social dialogue with multiple interlocutors simultaneously with rich output signals such as eye and head coordination, lips synchronized speech synthesis, and non-verbal facial gestures used to regulate fluent and expressive multiparty conversations.
Samer Al Moubayed, Gabriel Skantze, Jonas Beskow, Kalin Stefanov, Joakim Gustafson
ICMI5
2012 On the effect of the acoustic environment on the accuracy of perception of speaker orientation from auditory cues alone
abstract
The ability of people, and of machines, to determine the position of a sound source in a room is well studied. The related ability to determine the orientation of a directed sound source, on the other hand, is not, but the few studies there are show people to be surprisingly skilled at it. This has bearing for studies of face-to- face interaction and of embodied spoken dialogue systems, as sound source orientation of a speaker is connected to the head pose of the speaker, which is meaningful in a number of ways. The feature most often implicated for detection of sound source orientation is the inter-aural level difference - a feature which it is assumed is more easily exploited in anechoic chambers than in everyday surroundings. We expand here on our previous studies and compare detection of speaker orientation within and outside of the anechoic chamber. Our results show that listeners find the task easier, rather than harder, in everyday surroundings, which suggests that inter-aural level differences is not the only feature at play.
Jens Edlund, Mattias Heldner, Joakim Gustafson
INTERSPEECH3
2012 A Data-driven Approach to Understanding Spoken Route Directions in Human-Robot Dialogue
abstract
In this paper, we present a data-driven chunking parser for automatic interpretation of spoken route directions into a route graph that is useful for robot navigation. Different sets of features and machine learning algorithms are explored. The results indicate that our approach is robust to speech recognition errors. Index Terms: spoken language understanding, route directions, human-robot interaction
Raveesh Meena, Gabriel Skantze, Joakim Gustafson
INTERSPEECH3
2012 Gaze Patterns in Turn-Taking
abstract
Oertel C, Wlodarczak M, Edlund J, Wagner P, Gustafson J. Gaze patterns in turn-taking. In: 13th Annual Conference of the International Speech Communication Association 2012 (INTERSPEECH 2012). Red Hook, NY: Curran; 2013: 2243-2246.
Catharine Oertel, Marcin Wlodarczak, Jens Edlund, Petra Wagner, Joakim Gustafson
INTERSPEECH5
2011 Tracking Pitch Contours Using Minimum Jerk Trajectories
abstract
This paper proposes a fundamental frequency tracker, with the specific purpose of comparing the automatic estimates with pitch contours that are sketched by trained phoneticians. The method uses a frequency domain approach to estimate pitch tracks that form minimum jerk trajectories. This method tries to mimic motor movements of the hand made while sketch-ing. When the fundamental frequency tracked by the proposed method on the oral and laryngograph signals were compared using the MOCHA-TIMIT database, the correlation was 0.98 and the root mean squared error was 4.0 Hz, which was slightly better than a state-of-the-art pitch tracking algorithm included in the ESPS. We also demonstrate how the proposed algorithm could to be applied when comparing with sketches made by phoneticians for the variations in accent II among the Swedish dialects. Index Terms: pitch tracking, Constant-Q, Swedish accent II 1.
Daniel Neiberg, Gopal Ananthakrishnan, Joakim Gustafson
INTERSPEECH3
2011 Predicting Speaker Changes and Listener Responses with and without Eye-Contact
abstract
This paper compares turn-taking in terms of timing and prediction in human-human conversations under the conditions when participants has eye-contact versus when there is no eyecontact, as found in the HCRC Map Task corpus. By measuring between speaker intervals it was found that a larger proportion of speaker shifts occurred in overlap for the no eyecontact condition. For prediction we used prosodic and spectral features parametrized by time-varying length-invariant discrete cosine coefficients. With Gaussian Mixture Modeling and variations of classifier fusion schemes, we explored the task of predicting whether there is an upcoming speaker change (SC) or not (HOLD), at the end of an utterance (EOU) with a pause lag of 200 ms. The label SC was further split into LRs (listener responses, e.g. back-channels) and other TURN-SHIFTs. The prediction was found to be somewhat easier for the eye-contact condition, for which the average recall rates were
Daniel Neiberg, Joakim Gustafson
INTERSPEECH2
2011 A Dual Channel Coupled Decoder for Fillers and Feedback
abstract
This study presents a dual channel decoder capable of modeling cross-speaker dependencies for segmentation and classification of fillers and feedbacks in conversational speech found in the DEAL cor ...
Daniel Neiberg, Joakim Gustafson
INTERSPEECH2
2011 Enhanced visual scene understanding through human-robot dialog
abstract
We propose a novel human-robot-interaction framework for robust visual scene understanding. Without any a-priori knowledge about the objects, the task of the robot is to correctly enumerate how many of them are in the scene and segment them from the background. Our approach builds on top of state-of-the-art computer vision methods, generating object hypotheses through segmentation. This process is combined with a natural dialog system, thus including a `human in the loop' where, by exploiting the natural conversation of an advanced dialog system, the robot gains knowledge about ambiguous situations. We present an entropy-based system allowing the robot to detect the poorest object hypotheses and query the user for arbitration. Based on the information obtained from the human-robot dialog, the scene segmentation can be re-seeded and thereby improved. We present experimental results on real data that show an improved segmentation performance compared to segmentation without interaction.
Matthew Johnson-Roberson, Jeannette Bohg, Gabriel Skantze, Joakim Gustafson, Rolf Carlson, Babak Rasolzadeh, Danica Kragic
IROS4
2010 The prosody of Swedish conversational grunts
abstract
This paper explores conversational grunts in a face-to-face setting. The study investigates the prosody and turn-taking effect of fillers and feedback tokens that has been annotated for attitudes. The grunts were selected from the DEAL corpus and automatically annotated for their turn taking effect. A novel suprasegmental prosodic signal representation and contextual timing features are used for classification and visualization. Classification results using linear discriminant analysis, show that turn-initial feedback tokens lose some of their attitude-signaling prosodic cues compared to non-overlapping continuer feedback tokens. Turn taking effects can be predicted well over chance level, except Simultaneous Starts. However, feedback tokens before places where both speakers take the turn were more similar to feedback continuers than to turn initial feedback tokens.
Daniel Neiberg, Joakim Gustafson
INTERSPEECH2
2009 The MonAMI reminder: a spoken dialogue system for face-to-face interaction
abstract
We describe the MonAMI Reminder, a multimodal spoken dialogue system which can assist elderly and disabled people in organising and initiating their daily activities. Based on deep interviews with potential users, we have designed a calendar and reminder application which uses an innovative mix of an embodied conversational agent, digital pen and paper, and the web to meet the needs of those users as well as the current constraints of speech technology. We also explore the use of head pose tracking for interaction and attention control in human-computer face-to-face interaction.
Jonas Beskow, Jens Edlund, Björn Granström, Joakim Gustafson, Gabriel Skantze, Helena Tobiasson
INTERSPEECH4
2009 Eliciting Interactional Phenomena in Human-Human Dialogues
Joakim Gustafson, Miray Merkes
SIGDIAL Conference1
2009 Attention and Interaction Control in a Human-Human-Computer Dialogue Setting
Gabriel Skantze, Joakim Gustafson
SIGDIAL Conference2
2008 Innovative interfaces in MonAMI: the reminder
abstract
This demo paper presents an early version of the Reminder, a prototype ECA developed in the European project MonAMI, which aims at "mainstreaming accessibility in consumer goods and services, using advanced technologies to ensure equal access, independent living and participation for all". The Reminder helps users to plan activities and to remember what to do. The prototype merges mobile ECA technology with other, existing technologies: Google Calendar and a digital pen and paper. The solution allows users to continue using a paper calendar in the manner they are used to, whilst the ECA provides notifications on what has been written in the calendar. Users may ask questions such as "When was I supposed to meet Sara?" or "What's my schedule today?"
Jonas Beskow, Jens Edlund, Teodore Gjermani, Björn Granström, Joakim Gustafson, Oskar Jonsson, Gabriel Skantze, Helena Tobiasson
ICMI5
2008 What makes a good speaker? subject ratings, acoustic measurements and perceptual evaluations
abstract
This paper deals with subjective qualities and acoustic-prosodic features contributing to the impression of a good speaker. Subjects rated a variety of samples of political speech on a number of su ...
Eva Strangert, Joakim Gustafson
INTERSPEECH2
2008 Towards human-like spoken dialogue systems
Jens Edlund, Joakim Gustafson, Mattias Heldner, Anna Hjalmarsson
Speech Commun.2
2007 Children's convergence in referring expressions to graphical objects in a speech-enabled computer game
abstract
This paper describes an empirical study of children's spontaneous interactions with an animated character in a speech-enabled computer game. More specifically, it deals with convergence of referrin ...
Linda Bell, Joakim Gustafson
INTERSPEECH2
2006 Robust spoken language understanding in a computer game
Johan Boye, Joakim Gustafson, Mats Wirén
Speech Commun.2
2005 The Swedish NICE corpus - spoken dialogues between children and embodied characters in a computer game scenario
abstract
This article describes the collection and analysis of a Swedish database of spontaneous and unconstrained children-machine dialogues. The Swedish NICE corpus consists of spoken dialogues between children aged 8 to 15 and embodied fairytale characters in a computer game scenario. Compared to previously collected corpora of children's computer-directed speech, the Swedish NICE corpus contains extended interactions, including three-party conversation, in which the young users used spoken dialogue as the primary means of progression in the game.
Linda Bell, Johan Boye, Joakim Gustafson, Mattias Heldner, Anders Lindström, Mats Wirén
INTERSPEECH3
2005 Providing Computer Game Characters with Conversational Abilities
Joakim Gustafson, Johan Boye, Morgan Fredriksson, Lasse Johannesson, Jürgen Königsmann
IVA1
2003 Child and adult speaker adaptation during error resolution in a publicly available spoken dialogue system
abstract
This paper describes how speakers adapt their language during error resolution when interacting with the animated agent Pixie.A corpus of spontaneous human-computer interaction was collected at the Telecommunication museum in Stockholm, Sweden.Adult and children speakers were compared with respect to user behavior and strategies during error resolution.In this study, 16 adults and 16 children speakers were randomly selected from a corpus from almost 3.000 speakers.This sub-corpus was then analyzed in greater detail.Results indicate that adults and children use partly different strategies when their interactions with Pixie become problematic.Children tend to repeat the same utterance verbatim, altering certain phonetic features.Adults, on the other hand, often modify other aspects of their utterances such as lexicon and syntax.Results from the present study will be useful for constructing future spoken dialogue systems with improved error handling for adults as well as children.
Linda Bell, Joakim Gustafson
INTERSPEECH2
2002 Voice transformations for improving children²s speech recognition in a publicly available dialogue system
abstract
To be able to build acoustic models for children, that can beused in spoken dialogue systems, speech data has to be collected. Commercial recognizers available for Swedish are trained on adult speech, which makes them less suitable for children’s computer-directed speech. This paper describes some experiments with on-the-fly voice transformation of children’s speech. Two transformation methods were tested, one inspired by the Phase Vocoder algorithm and another by the Time-Domain Pitch-Synchronous Overlap-Add (TD-PSOLA)algorithm. The speech signal is transformed before being sent to the speech recognizer for adult speech. Our results show that this method reduces the error rates in the order of thirty to fortyfive percent for children users.
Joakim Gustafson, Kåre Sjölander
INTERSPEECH1
2000 A comparison of disfluency distribution in a unimodal and a multimodal speech interface
abstract
In this paper, we compare the distribution of disfluencies in two human--computer dialogue corpora. One corpus consists of unimodal travel booking dialogues, which were recorded over the telephone. In this unimodal system, all components except the speech recognition were authentic. The other corpus was collected using a semi-simulated multi-modal dialogue system with an animated talking agent and a clickable map. The aim of this paper is to analyze and discuss the effects of modality, task and interface design on the distribution and frequency of disfluencies in these two corpora.
Linda Bell, Robert Eklund, Joakim Gustafson
INTERSPEECH3
2000 Positive and negative user feedback in a spoken dialogue corpus
abstract
This paper examines feedback strategies in a Swedish corpus of multimodal human--computer interaction. The aim of the study is to investigate how users provide positive and negative feedback to a dialogue system and to discuss the function of these utterances in the dialogues. User feedback in the AdApt corpus was labeled and analyzed, and its distribution in the dialogues is discussed. The question of whether it is possible to utilize user feedback in future systems is considered. More specifically, we discuss how error handling in human--computer dialogue might be improved through greater knowledge of user feedback strategies. In the present corpus, almost all subjects used positive or negative feedback at least once during their interaction with the system. Our results indicate that some types of feedback more often occur in certain positions in the dialogue. Another observation is that there appear to be great individual variations in feedback strategies, so that certain subjects give feedback at almost every turn while others rarely or never respond to a spoken dialogue system in this manner. Finally, we discuss how feedback could be used to prevent problems in human--computer dialogue.
Linda Bell, Joakim Gustafson
INTERSPEECH2
2000 Adapt - a multimodal conversational dialogue system in an apartment domain
abstract
A general overview of the AdApt project and the research that is performed within the project is presented. In this project various aspects of human-computer interaction in a multimodal conversational dialogue systems are investigated. The project will also include studies on the integration of user/system/dialogue dependent speech recognition and multimodal speech synthesis. A domain in which multimodal interaction is highly useful has been chosen, namely, finding available apartments in Stockholm. A Wizard-of-Oz data collection within this domain is also described. 1.
Joakim Gustafson, Linda Bell, Jonas Beskow, Johan Boye, Rolf Carlson, Jens Edlund, Björn Granström, David House, Mats Wirén
INTERSPEECH1
2000 Speech technology on trial: Experiences from the August system
abstract
In this paper, the August spoken dialogue system is described. This experimental Swedish dialogue system, which featured an animated talking agent, was exposed to the general public during a trial period of six months. The construction of the system was partly motivated by the need to collect genuine speech data from people with little or no previous experience of spoken dialogue systems. A corpus of more than 10,000 utterances of spontaneous computer- directed speech was collected and empirical linguistic analyses were carried out. Acoustical, lexical and syntactical aspects of this data were examined. In particular, user behavior and user adaptation during error resolution were emphasized. Repetitive sequences in the database were analyzed in detail. Results suggest that computer-directed speech during error resolution is increased in duration, hyperarticulated and contains inserted pauses. Design decisions which may have influenced how the users behaved when they interacted with August are discussed and implications for the development of future systems are outlined.
Joakim Gustafson, Linda Bell
Nat. Lang. Eng.1
1999 Interaction with an animated agent in a spoken dialogue system
Linda Bell, Joakim Gustafson
EUROSPEECH2
1999 The august spoken dialogue system
abstract
This paper describes how a telephone-based application can perform a variety of tasks in a completely hands-free mode. The overall architecture of the speech component is multimodal in that each mode is tailored to a specific need of the interface. The various modes are described as well as the underlying core technology. To illustrate the effectiveness of the implementation, we present experimental results in American English, UK English and French on a variety of benchmarks, including live data collected during actual use of the system.
Joakim Gustafson, Nikolaj Lindberg, Magnus Lundeberg
EUROSPEECH1
1998 An educational dialogue system with a user controllable dialogue manager
abstract
We have developed an educational environment for a modular spoken dialogue system. The aim of the environment is to provide students, with different backgrounds, means to understand the behaviour of spoken dialogue systems. Focus in this paper is on dialogue and dialogue management. The dialogue is recorded in a dialogue tree whose nodes are dialogue objects. The dialogue objects model the constituents of the dialogue and consist of parameters for modelling dialogue structure, focus structure and a process description describing the actions of the dialogue system. Various dialogue system behaviours can be achieved by modifying these parameters. This is done using the educational environment, which is interactive and facilitates examination, expansion and modification of the dialogue object parameters and hence the system. The educational system has been used in a number of courses at various universities in Sweden. 1.
Joakim Gustafson, Patrik Elmberg, Rolf Carlson, Arne Jönsson
ICSLP1
1998 Web-based educational tools for speech technology
abstract
This paper describes the efforts at KTH in creating educational tools for speech technology. The demand for such tools is increasing with the advent of speech as a medium for man– machine communication. The world wide web was chosen as our platform in order to increase the usability and accessibility of our computer exercises. The aim was to provide dedicated educational software instead of exercises based on complex research tools. Currently, the set of exercises comprises basic speech analysis, multi-modal speech synthesis and spoken dialogue systems. Students access web pages in which the exercises have been embedded as applets. This makes it possible to use them in a classroom setting, as well as from the students’ home computers.
Kåre Sjölander, Jonas Beskow, Joakim Gustafson, Erland Lewin, Rolf Carlson, Björn Granström
ICSLP3
1997 How do system questions influence lexical choices in user answers?
abstract
This paper describes some studies on the effect of the system vocabulary on the lexical choices of the users. There are many theories about human-human dialogues that could be useful in the design of spoken dialoguesystems. This paper will give an overview of some of these theories and report the results from two experiments that examines one of these theories, namely lexical entrainment. The first experiment was a small Wizard of Oz-test that simulated a tourist informationsystem with a speech interface, and the second experiment simulated a system with speech recognition that controlled a questionnaire about peoples plans for their vacation. Both experiments show that the subjects mostly adapt their lexical choices to the system questions. Only in less than 5% of the cases did they use an alternative main verb in the answer. These results encourage us to investigate the possibility to add anadaptive language model in the speech recognizer in our dialogue system, where the probabilities for the words used in the system questions are increased.
Joakim Gustafson, Anette Larsson, Rolf Carlson, K. Hellman
EUROSPEECH1
1997 An integrated system for teaching spoken dialogue systems technology
Kåre Sjölander, Joakim Gustafson
EUROSPEECH2
1995 The waxholm application database
abstract
This paper describes an application database collected in Wizard-of-Oz experiments in a spoken dialogue system, WAXHOLM. The system provides information on boat traffic in the Stockholm archipelago. The database consists of utterance-length speech files, their corresponding transcriptions, and log files of the dialogue sessions. In addition to the spontaneous dialogue speech, the material also comprise recordings of phonetically balanced reference sentences uttered by all 66 subjects. In the paper the recording procedure is described as well as some characteristics of the speech data and the dialogue. INTRODUCTION WAXHOLM is a demonstrator spoken dialogue system in which we apply our research on speech recognition and speech synthesis. The system uses visual and auditory means to provide information on boat traffic, accomodation, etc. in the Stockholm archipelago [1], [2], [3]. The application has great similarities to the ATIS domain within the ARPA community, the Voyager system from...
J. Bertenstam, Mats Blomberg, Rolf Carlson, Kjell Elenius, Björn Granström, Joakim Gustafson, Sheri Hunnicutt, Jesper Högberg, Roger Lindell, Lennart Neovius, Lennart Nord, Antonio de Serpa-Leitao, Nikko Strom
EUROSPEECH6
1995 Using two-level morphology to transcribe Swedish names
abstract
Names are difficult to handle for normal letter-to-sound rules, since these usually are designed for ordinary words. The structure of Swedish names differ from ordinary words - but their multi-morphemic structure make them suitable to analyse with a morphological analyser.
Joakim Gustafson
EUROSPEECH1
1993 An experimental dialogue system: waxholm
abstract
Recently we have begun to build the basic tools for a generic speech-dialogue system, WAXHOLM. The main modules, their function and internal communication have been specified. The different components are connected through a computer network. A preliminary version of the system has been tested, using simplified versions of the modules. We will give a general overview of the system and describe some of the components in more detail. Application specific data are collected with the help of Wizard-of-Oz techniques. The dialogue system is used during the data collection and the wizard only replaces the speech-recognition module.
Mats Blomberg, Rolf Carlson, Kjell Elenius, Björn Granström, Joakim Gustafson, Sheri Hunnicutt, Roger Lindell, Lennart Neovius
EUROSPEECH5