EDBT 2026 Demo / reviewers in the wild / expert
Joakim Gustafson
dblp:28/6376
· DBLP profile ↗
92ranked-venue papers
11as first author
23since 2021 · last 2025
0000-0002-0397-6442ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 76 · 11 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 54 · 7 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 18 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Static to Dynamic: Enhancing AAC with Generative Imagery and Zero-Shot TTS
Juliana Francis, Joakim Gustafson, Éva Székely |
INTERSPEECH | 2 |
| 2025 | VoiceQualityVC: A Voice Conversion System for Studying the Perceptual Effects of Voice Quality in Speech
Harm Lameris, Joakim Gustafson, Éva Székely |
INTERSPEECH | 2 |
| 2025 | Towards Adaptable and Intelligible Speech Synthesis in Noisy Environments
Lubos Marcinek, Jonas Beskow, Joakim Gustafson |
INTERSPEECH | 3 |
| 2025 | Role of Reasoning in LLM Enjoyment Detection: Evaluation Across Conversational Levels for Human-Robot InteractionabstractUser enjoyment is central to developing conversational AI systems that can recover from failures and maintain interest over time. However, existing approaches often struggle to detect subtle cues that reflect user experience. Large Language Models (LLMs) with reasoning capabilities have outperformed standard models on various other tasks, suggesting potential benefits for enjoyment detection. This study investigates whether models with reasoning capabilities outperform standard models when assessing enjoyment in a human-robot dialogue corpus at both turn and interaction levels. Results indicate that reasoning capabilities have complex, model-dependent effects rather than universal benefits. While performance was nearly identical at the interaction level (0.44 vs 0.43), reasoning models substantially outperformed at the turn level (0.42 vs 0.36). Notably, LLMs correlated better with users’ self-reported enjoyment metrics than human annotators, despite achieving lower accuracy against human consensus ratings. Analysis revealed distinctive error patterns: non-reasoning models showed bias toward positive ratings at the turn level, while both model types exhibited central tendency bias at the interaction level. These findings suggest that reasoning should be applied selectively based on model architecture and assessment context, with assessment granularity significantly influencing relative effectiveness. Lubos Marcinek, Bahar Irfan, Gabriel Skantze, André Pereira 0001, Joakim Gustafson |
SIGDIAL | 5 |
| 2024 | The Role of Creaky Voice in Turn Taking and the Perception of Speaker Stance: Experiments Using Controllable TTSabstractRecent advancements in spontaneous text-to-speech (TTS) have enabled the realistic synthesis of creaky voice, a voice quality known for its diverse pragmatic and paralinguistic functions. In this study, we used synthesized creaky voice in perceptual tests, to explore how listeners without formal training perceive two distinct types of creaky voice. We annotated a spontaneous speech corpus using creaky voice detection tools and modified a neural TTS engine with a creaky phonation embedding to control the presence of creaky phonation in the synthesized speech. We performed an objective analysis using a creak detection tool which revealed significant differences in creaky phonation levels between the two creaky voice types and modal voice. Two subjective listening experiments were performed to investigate the effect of creaky voice on perceived certainty, valence, sarcasm, and turn finality. Participants rated non-positional creak as less certain, less positive, and more indicative of turn finality, while positional creak was rated significantly more turn final compared to modal phonation. Harm Lameris, Éva Székely, Joakim Gustafson |
LREC/COLING | 3 |
| 2024 | Revisiting Three Text-to-Speech Synthesis Experiments with a Web-Based Audience Response SystemabstractIn order to investigate the strengths and weaknesses of Audience Response System (ARS) in text-to-speech synthesis (TTS) evaluations, we revisit three previously published TTS studies and perform an ARS-based evaluation on the stimuli used in each study. The experiments are performed with a participant pool of 39 respondents, using a web-based tool that emulates an ARS experiment. The results of the first experiment confirms that ARS is highly useful for evaluating long and continuous stimuli, particularly if we wish for a diagnostic result rather than a single overall metric, while the second and third experiments highlight weaknesses in ARS with unsuitable materials as well as the importance of framing and instruction when conducting ARS-based evaluation. Christina Tånnander, Jens Edlund, Joakim Gustafson |
LREC/COLING | 3 |
| 2024 | Multimodal User Enjoyment Detection in Human-Robot Conversation: The Power of Large Language ModelsabstractEnjoyment is a crucial yet complex indicator of positive user experience in Human-Robot Interaction (HRI). While manual enjoyment annotation is feasible, developing reliable automatic detection methods remains a challenge. This paper investigates a multimodal approach to automatic enjoyment annotation for HRI conversations, leveraging large language models (LLMs), visual, audio, and temporal cues. Our findings demonstrate that both text-only and multimodal LLMs with carefully designed prompts can achieve performance comparable to human annotators in detecting user enjoyment. Furthermore, results reveal a stronger alignment between LLM-based annotations and user self-reports of enjoyment compared to human annotators. While multimodal supervised learning techniques did not improve all of our performance metrics, they could successfully replicate human annotators and highlighted the importance of visual and audio cues in detecting subtle shifts in enjoyment. This research demonstrates the potential of LLMs for real-time enjoyment detection, paving the way for adaptive companion robots that can dynamically enhance user experiences. André Pereira 0001, Lubos Marcinek, Jura Miniota, Sofia Thunberg, Erik Lagerstedt, Joakim Gustafson, Gabriel Skantze, Bahar Irfan |
ICMI | 6 |
| 2024 | ConnecTone: a modular AAC system prototype with contextual generative text prediction and style-adaptive conversational TTS
Juliana Francis, Éva Székely, Joakim Gustafson |
INTERSPEECH | 3 |
| 2024 | CreakVC: a voice conversion tool for modulating creaky voice
Harm Lameris, Joakim Gustafson, Éva Székely |
INTERSPEECH | 2 |
| 2024 | Contextual Interactive Evaluation of TTS Models in Dialogue Systems
Éva Székely, Joakim Gustafson |
INTERSPEECH | 3 |
| 2023 | Casual chatter or speaking up? Adjusting articulatory effort in generation of speech and animation for conversational charactersabstractEmbodied conversational agents and social robots need to be able to generate spontaneous behavior in order to be believable in social interactions. We present a system that can generate spontaneous speech with supporting lip movements. The conversational TTS voice is trained on a podcast corpus that has been prosodically tagged (f0, speaking rate and energy) and transcribed (including tokens for breathing, fillers and laughter). We introduce a speech animation algorithm where articulatory effort can be adjusted. The speech animation is driven by time-stamped phonemes obtained from the internal alignment attention map of the TTS system, and we use prominence estimates from the synthesised speech waveform to modulate the lip- and jaw movements accordingly. Joakim Gustafson, Éva Székely, Simon Alexanderson, Jonas Beskow |
FG | 1 |
| 2023 | Prosody-Controllable Spontaneous TTS with Neural HMMSabstractSpontaneous speech has many affective and pragmatic functions that are interesting and challenging to model in TTS. However, the presence of reduced articulation, fillers, repetitions, and other disfluencies in spontaneous speech make the text and acoustics less aligned than in read speech, which is problematic for attention-based TTS. We propose a TTS architecture that can rapidly learn to speak from small and irregular datasets, while also reproducing the diversity of expressive phenomena present in spontaneous speech. Specifically, we add utterance-level prosody control to an existing neural HMM-based TTS system which is capable of stable, monotonic alignments for spontaneous speech. We objectively evaluate control accuracy and perform perceptual tests that demonstrate that prosody control does not degrade synthesis quality. To exemplify the power of combining prosody control and ecologically valid data for reproducing intricate spontaneous speech phenomena, we evaluate the system’s capability of synthesizing two types of creaky voice. Harm Lameris, Shivam Mehta, Gustav Eje Henter, Joakim Gustafson, Éva Székely |
ICASSP | 4 |
| 2023 | Automatic Evaluation of Turn-taking Cues in Conversational Speech Synthesis
Erik Ekstedt, Éva Székely, Joakim Gustafson, Gabriel Skantze |
INTERSPEECH | 4 |
| 2023 | Pardon my disfluency: The impact of disfluency effects on the perception of speaker competence and confidence
Ambika Kirkland, Joakim Gustafson, Éva Székely |
INTERSPEECH | 2 |
| 2023 | Beyond Style: Synthesizing Speech with Pragmatic Functions
Harm Lameris, Joakim Gustafson, Éva Székely |
INTERSPEECH | 2 |
| 2023 | Prosody-controllable Gender-ambiguous Speech Synthesis: A Tool for Investigating Implicit Bias in Speech Perception
Éva Székely, Joakim Gustafson, Ilaria Torre 0002 |
INTERSPEECH | 2 |
| 2023 | So-to-Speak: An Exploratory Platform for Investigating the Interplay between Style and Prosody in TTS
Éva Székely, Joakim Gustafson |
INTERSPEECH | 3 |
| 2023 | Generation of speech and facial animation with controllable articulatory effort for amusing conversational charactersabstractEngaging embodied conversational agents need to generate expressive behavior in order to be believable in socializing interactions. We present a system that can generate spontaneous speech with supporting lip movements. The neural conversational TTS voice is trained on a multi-style speech corpus that has been prosodically tagged (pitch and speaking rate) and transcribed (including tokens for breathing, fillers and laughter). We introduce a speech animation algorithm where articulatory effort can be adjusted. The facial animation is driven by time-stamped phonemes and prominence estimates from the synthesised speech waveform to modulate the lip-and jaw movements accordingly. In objective evaluations we show that the system is able to generate speech and facial animation that vary in articulation effort. In subjective evaluations we compare our conversational TTS system's capability to deliver jokes with a commercial TTS. Both system succeeded equally good. Joakim Gustafson, Éva Székely, Jonas Beskow |
IVA | 1 |
| 2023 | Hi robot, it's not what you say, it's how you say itabstractMany robots use their voice to communicate with people in spoken language but the voices commonly used for robots are often optimized for transactional interactions, rather than social ones. This can limit their ability to create engaging and natural interactions. To address this issue, we designed a spontaneous text-to-speech tool and used it to author natural and spontaneous robot speech. A crowdsourcing evaluation methodology is proposed to compare this type of speech to natural speech and state-of-the-art text-to-speech technology, both in disembodied and embodied form. We created speech samples in a naturalistic setting of people playing tabletop games and conducted a user study evaluating Naturalness, Intelligibility, Social Impression, Prosody, and Perceived Intelligence. The speech samples were chosen to represent three contexts that are common in tabletop games and the contexts were introduced to the participants that evaluated the speech samples. The study results show that the proposed evaluation methodology allowed for a robust analysis that successfully compared the different conditions. Moreover, the spontaneous voice met our target design goal of being perceived as more natural than a leading commercial text-to-speech. Jura Miniota, Jonas Beskow, Joakim Gustafson, Éva Székely, André Pereira 0001 |
RO-MAN | 4 |
| 2022 | Where's the uh, hesitation? The interplay between filled pause location, speech rate and fundamental frequency in perception of confidence
Ambika Kirkland, Harm Lameris, Éva Székely, Joakim Gustafson |
INTERSPEECH | 4 |
| 2022 | Evaluating Sampling-based Filler Insertion with Spontaneous TTSabstractInserting fillers (such as “um”, “like”) to clean speech text has a rich history of study. One major application is to make dialogue systems sound more spontaneous. The ambiguity of filler occurrence and inter-speaker difference make both modeling and evaluation difficult. In this paper, we study sampling-based filler insertion, a simple yet unexplored approach to inserting fillers. We propose an objective score called Filler Perplexity (FPP). We build three models trained on two single-speaker spontaneous corpora, and evaluate them with FPP and perceptual tests. We implement two innovations in perceptual tests, (1) evaluating filler insertion on dialogue systems output, (2) synthesizing speech with neural spontaneous TTS engines. FPP proves to be useful in analysis but does not correlate well with perceptual MOS. Perceptual results show little difference between compared filler insertion models including with ground-truth, which may be due to the ambiguity of what is good filler insertion and a strong neural spontaneous TTS that produces natural speech irrespective of input. Results also show preference for filler-inserted speech synthesized with spontaneous TTS. The same test using TTS based on read speech obtains the opposite results, which shows the importance of using spontaneous TTS in evaluating filler insertions. Audio samples: www.speech.kth.se/tts-demos/LREC22 Joakim Gustafson, Éva Székely |
LREC | 2 |
| 2021 | A Systematic Cross-Corpus Analysis of Human Reactions to Robot Conversational FailuresabstractIn this paper, we analyze multimodal behavioral responses to robot failures across different tasks. Two multimodal datasets are examined in which humans interact with guided-task robots in task-oriented dialogues. In both datasets, the robots simulated failures of conversational breakdown and miscommunication typically observed in human-robot interactions. We closely examine human reactions to these failures looking at facial and acoustic features. Our analyses identify the significant behavioral features for automatic detection of such failures in interaction. We also examine human responses to different types of robot failures and if failures occurred early or late in the interaction cause variation in the responses. Our findings indicate that several nonverbal behaviors are consistently present in responses to robots’ failures, e.g., gaze and speech prosody, whereas, linguistic features appear to be task-dependent. We discuss how these findings may generalize to other tasks, and how autonomous robots may identify opportunities to detect and recover from failures in interactions with humans. Dimosthenis Kontogiorgos, Minh Tran 0004, Joakim Gustafson, Mohammad Soleymani 0001 |
ICMI | 3 |
| 2021 | Integrated Speech and Gesture SynthesisabstractText-to-speech and co-speech gesture synthesis have until now been treated as separate areas by two different research communities, and applications merely stack the two technologies using a simple system-level pipeline. This can lead to modeling inefficiencies and may introduce inconsistencies that limit the achievable naturalness. We propose to instead synthesize the two modalities in a single model, a new problem we call integrated speech and gesture synthesis (ISG). We also propose a set of models modified from state-of-the-art neural speech-synthesis engines to achieve this goal. We evaluate the models in three carefully-designed user studies, two of which evaluate the synthesized speech and gesture in isolation, plus a combined study that evaluates the models like they will be used in real-world applications – speech and gesture presented together. The results show that participants rate one of the proposed integrated synthesis models as being as good as the state-of-the-art pipeline system we compare against, in all three tests. The model is able to achieve this with faster synthesis time and greatly reduced parameter count compared to the pipeline system, illustrating some of the potential benefits of treating speech and gesture synthesis together as a single, unified problem. Simon Alexanderson, Joakim Gustafson, Jonas Beskow, Gustav Eje Henter, Éva Székely |
ICMI | 3 |
| 2020 | Embodiment Effects in Interactions with Failing RobotsabstractThe increasing use of robots in real-world applications will inevitably cause users to encounter more failures in interactions. While there is a longstanding effort in bringing human-likeness to robots, how robot embodiment affects users' perception of failures remains largely unexplored. In this paper, we extend prior work on robot failures by assessing the impact that embodiment and failure severity have on people's behaviours and their perception of robots. Our findings show that when using a smart-speaker embodiment, failures negatively affect users' intention to frequently interact with the device, however not when using a human-like robot embodiment. Additionally, users significantly rate the human-like robot higher in terms of perceived intelligence and social presence. Our results further suggest that in higher severity situations, human-likeness is distracting and detrimental to the interaction. Drawing on quantitative findings, we discuss benefits and drawbacks of embodiment in robot failures that occur in guided tasks. Dimosthenis Kontogiorgos, Sanne van Waveren, Olle Wallberg, André Pereira 0001, Iolanda Leite, Joakim Gustafson |
CHI | 6 |
| 2020 | Behavioural Responses to Robot Conversational FailuresabstractHumans and robots will increasingly collaborate in domestic environments which will cause users to encounter more failures in interactions. Robots should be able to infer conversational failures by detecting human users' behavioural and social signals. In this paper, we study and analyse these behavioural cues in response to robot conversational failures. Using a guided task corpus, where robot embodiment and time pressure are manipulated, we ask human annotators to estimate whether user affective states differ during various types of robot failures. We also train a random forest classifier to detect whether a robot failure has occurred and compare results to human annotator benchmarks. Our findings show that human-like robots augment users' reactions to failures, as shown in users' visual attention, in comparison to non-human-like smart-speaker embodiments. The results further suggest that speech behaviours are utilised more in responses to failures when non-human-like designs are present. This is particularly important to robot failure detection mechanisms that may need to consider the robot's physical design in its failure detection model. Dimosthenis Kontogiorgos, André Pereira 0001, Boran Sahindal, Sanne van Waveren, Joakim Gustafson |
HRI | 5 |
| 2020 | Effects of Different Interaction Contexts when Evaluating Gaze Models in HRIabstractWe previously introduced a responsive joint attention system that uses multimodal information from users engaged in a spatial reasoning task with a robot and communicates joint attention via the robot's gaze behavior. An initial evaluation of our system with adults showed it to improve users' perceptions of the robot's social presence. To investigate the repeatability of our prior findings across settings and populations, here we conducted two further studies employing the same gaze system with the same robot and task but in different contexts: evaluation of the system with external observers and evaluation with children. The external observer study suggests that third-person perspectives over videos of gaze manipulations can be used either as a manipulation check before committing to costly real-time experiments or to further establish previous findings. However, the replication of our original adults study with children in school did not confirm the effectiveness of our gaze manipulation, suggesting that different interaction contexts can affect the generalizability of results in human-robot interaction gaze studies. André Pereira 0001, Catharine Oertel, Leonor Fermoselle, Joseph Mendelson, Joakim Gustafson |
HRI | 5 |
| 2020 | Breathing and Speech Planning in Spontaneous Speech SynthesisabstractBreathing and speech planning in spontaneous speech are coordinated processes, often exhibiting disfluent patterns. While synthetic speech is not subject to respiratory needs, integrating breath into synthesis has advantages for naturalness and recall. At the same time, a synthetic voice reproducing disfluent breathing patterns learned from the data can be problematic. To address this, we first propose training stochastic TTS on a corpus of overlapping breath-group bigrams, to take context into account. Next, we introduce an unsupervised automatic annotation of likely-disfluent breath events, through a product-of-experts model that combines the output of two breath- event predictors, each using complementary information and operating in opposite directions. This annotation enables creating an automatically-breathing spontaneous speech synthesiser with a more fluent breathing style. A subjective evaluation on two spoken genres (impromptu and rehearsed) found the proposed system to be preferred over the baseline approach treating all breath events the same. Éva Székely, Gustav Eje Henter, Jonas Beskow, Joakim Gustafson |
ICASSP | 4 |
| 2020 | Chinese Whispers: A Multimodal Dataset for Embodied Language GroundingabstractIn this paper, we introduce a multimodal dataset in which subjects are instructing each other how to assemble IKEA furniture. Using the concept of ‘Chinese Whispers’, an old children’s game, we employ a novel method to avoid implicit experimenter biases. We let subjects instruct each other on the nature of the task: the process of the furniture assembly. Uncertainty, hesitations, repairs and self-corrections are naturally introduced in the incremental process of establishing common ground. The corpus consists of 34 interactions, where each subject first assembles and then instructs. We collected speech, eye-gaze, pointing gestures, and object movements, as well as subjective interpretations of mutual understanding, collaboration and task recall. The corpus is of particular interest to researchers who are interested in multimodal signals in situated dialogue, especially in referential communication and the process of language grounding. Dimosthenis Kontogiorgos, Elena Sibirtseva, Joakim Gustafson |
LREC | 3 |
| 2020 | Augmented Prompt Selection for Evaluation of Spontaneous Speech SynthesisabstractBy definition, spontaneous speech is unscripted and created on the fly by the speaker. It is dramatically different from read speech, where the words are authored as text before they are spoken. Spontaneous speech is emergent and transient, whereas text read out loud is pre-planned. For this reason, it is unsuitable to evaluate the usability and appropriateness of spontaneous speech synthesis by having it read out written texts sampled from for example newspapers or books. Instead, we need to use transcriptions of speech as the target - something that is much less readily available. In this paper, we introduce Starmap, a tool allowing developers to select a varied, representative set of utterances from a spoken genre, to be used for evaluation of TTS for a given domain. The selection can be done from any speech recording, without the need for transcription. The tool uses interactive visualisation of prosodic features with t-SNE, along with a tree-based algorithm to guide the user through thousands of utterances and ensure coverage of a variety of prompts. A listening test has shown that with a selection of genre-specific utterances, it is possible to show significant differences across genres between two synthetic voices built from spontaneous speech. Éva Székely, Jens Edlund, Joakim Gustafson |
LREC | 3 |
| 2019 | The Effects of Embodiment and Social Eye-Gaze in Conversational Agents
Dimosthenis Kontogiorgos, Gabriel Skantze, André Pereira 0001, Joakim Gustafson |
CogSci | 4 |
| 2019 | Casting to Corpus: Segmenting and Selecting Spontaneous Dialogue for Tts with a Cnn-lstm Speaker-dependent Breath DetectorabstractThis paper considers utilising breaths to create improved spontaneous-speech corpora for conversational text-to-speech from found audio recordings such as dialogue podcasts. Breaths are of interest since they relate to prosody and speech planning and are independent of language and transcription. Specifically, we propose a semi-supervised approach where a fraction of coarsely annotated data is used to train a convolutional and recurrent speaker-specific breath detector operating on spectrograms and zero-crossing rate. The classifier output is used to find target-speaker breath groups (audio segments delineated by breaths) and subsequently select those that constitute clean utterances appropriate for a synthesis corpus. An application to 11 hours of raw podcast audio extracts 1969 utterances (106 minutes), 87% of which are clean and correctly segmented. This outperforms a baseline that performs integrated VAD and speaker attribution without accounting for breaths. Éva Székely, Gustav Eje Henter, Joakim Gustafson |
ICASSP | 3 |
| 2019 | Estimating Uncertainty in Task-Oriented DialogueabstractSituated multimodal systems that instruct humans need to handle user uncertainties, as expressed in behaviour, and plan their actions accordingly. Speakers’ decision to reformulate or repair previous utterances depends greatly on the listeners’ signals of uncertainty. In this paper, we estimate uncertainty in a situated guided task, as leveraged in non-verbal cues expressed by the listener, and predict that the speaker will reformulate their utterance. We use a corpus where people instruct how to assemble furniture, and extract their multimodal features. While uncertainty is in cases verbally expressed, most instances are expressed non-verbally, which indicates the importance of multimodal approaches. In this work, we present a model for uncertainty estimation. Our findings indicate that uncertainty estimation from non-verbal cues works well, and can exceed human annotator performance when verbal features cannot be perceived. Dimosthenis Kontogiorgos, André Pereira 0001, Joakim Gustafson |
ICMI | 3 |
| 2019 | Off the Cuff: Exploring Extemporaneous Speech Delivery with TTS
Éva Székely, Gustav Eje Henter, Jonas Beskow, Joakim Gustafson |
INTERSPEECH | 4 |
| 2019 | Spontaneous Conversational Speech Synthesis from Found DataabstractSynthesising spontaneous speech is a difficult task due to disfluencies, high variability and syntactic conventions different from those of written language. Using found data, as opposed to lab-rec ... Éva Székely, Gustav Eje Henter, Jonas Beskow, Joakim Gustafson |
INTERSPEECH | 4 |
| 2019 | Responsive Joint Attention in Human-Robot InteractionabstractJoint attention has been shown to be not only crucial for human-human interaction but also human-robot interaction. Joint attention can help to make cooperation more efficient, support disambiguation in instances of uncertainty and make interactions appear more natural and familiar. In this paper, we present an autonomous gaze system that uses multimodal perception capabilities to model responsive joint attention mechanisms. We investigate the effects of our system on people's perception of a robot within a problem-solving task. Results from a user study suggest that responsive joint attention mechanisms evoke higher perceived feelings of social presence on scales that regard the direction of the robot's perception. André Pereira 0001, Catharine Oertel, Leonor Fermoselle, Joe Mendelson, Joakim Gustafson |
IROS | 5 |
| 2019 | The Effects of Anthropomorphism and Non-verbal Social Behaviour in Virtual AssistantsabstractThe adoption of virtual assistants is growing at a rapid pace. However, these assistants are not optimised to simulate key social aspects of human conversational environments. Humans are intellectually biased toward social activity when facing anthropomorphic agents or when presented with subtle social cues. In this paper, we test whether humans respond the same way to assistants in guided tasks, when in different forms of embodiment and social behaviour. In a within-subject study (N=30), we asked subjects to engage in dialogue with a smart speaker and a social robot. We observed shifting of interactive behaviour, as shown in behavioural and subjective measures. Our findings indicate that it is not always favourable for agents to be anthropomorphised or to communicate with nonverbal cues. We found a trade-off between task performance and perceived sociability when controlling for anthropomorphism and social behaviour. Dimosthenis Kontogiorgos, André Pereira 0001, Olle Andersson, Marco Koivisto, Elena Gonzalez Rabal, Ville Vartiainen, Joakim Gustafson |
IVA | 7 |
| 2018 | Interactive, Collaborative Robots: Challenges and OpportunitiesabstractRobotic technology has transformed manufacturing industry ever since the first industrial robot was put in use in the beginning of the 60s. The challenge of developing flexible solutions where production lines can be quickly re-planned, adapted and structured for new or slightly changed products is still an important open problem. Industrial robots today are still largely preprogrammed for their tasks, not able to detect errors in their own performance or to robustly interact with a complex environment and a human worker. The challenges are even more serious when it comes to various types of service robots. Full robot autonomy, including natural interaction, learning from and with human, safe and flexible performance for challenging tasks in unstructured environments will remain out of reach for the foreseeable future. In the envisioned future factory setups, home and office environments, humans and robots will share the same workspace and perform different object manipulation tasks in a collaborative manner. We discuss some of the major challenges of developing such systems and provide examples of the current state of the art. Danica Kragic, Joakim Gustafson, Hakan Karaoguz, Patric Jensfelt, Robert Krug 0002 |
IJCAI | 2 |
| 2018 | Crowdsourced Multimodal Corpora Collection Tool
Patrik Jonell, Catharine Oertel, Dimosthenis Kontogiorgos, Jonas Beskow, Joakim Gustafson |
LREC | 5 |
| 2018 | A Multimodal Corpus for Mutual Gaze and Joint Attention in Multiparty Situated Interaction
Dimosthenis Kontogiorgos, Vanya Avramova, Simon Alexanderson, Patrik Jonell, Catharine Oertel, Jonas Beskow, Gabriel Skantze, Joakim Gustafson |
LREC | 8 |
| 2018 | A Comparison of Visualisation Methods for Disambiguating Verbal Requests in Human-Robot InteractionabstractPicking up objects requested by a human user is a common task in human-robot interaction. When multiple objects match the user's verbal description, the robot needs to clarify which object the user is referring to before executing the action. Previous research has focused on perceiving user's multimodal behaviour to complement verbal commands or minimising the number of follow up questions to reduce task time. In this paper, we propose a system for reference disambiguation based on visualisation and compare three methods to disambiguate natural language instructions. In a controlled experiment with a YuMi robot, we investigated realtime augmentations of the workspace in three conditions - head-mounted display, projector, and a monitor as the baseline - using objective measures such as time and accuracy, and subjective measures like engagement, immersion, and display interference. Significant differences were found in accuracy and engagement between the conditions, but no differences were found in task time. Despite the higher error rates in the head-mounted display condition, participants found that modality more engaging than the other two, but overall showed preference for the projector condition over the monitor and head-mounted display conditions. Elena Sibirtseva, Dimosthenis Kontogiorgos, Olov Nykvist, Hakan Karaoguz, Iolanda Leite, Joakim Gustafson, Danica Kragic |
RO-MAN | 6 |
| 2017 | Controlling Prominence Realisation in Parametric DNN-Based Speech SynthesisabstractThis work aims to improve text-To-speech synthesis forWikipedia by advancing and implementing models of prosodic prominence. We propose a new system architecture with explicit prominence modeling a ... Zofia Malisz, Harald Berthelsen, Jonas Beskow, Joakim Gustafson |
INTERSPEECH | 4 |
| 2017 | Crowd-Sourced Design of Artificial Attentive ListenersabstractFeedback generation is an important component of humanhuman communication. Humans can choose to signal support, understanding, agreement or also sceptiscism by means of feedback tokens. Many studie ... Catharine Oertel, Patrik Jonell, Dimosthenis Kontogiorgos, Joseph Mendelson, Jonas Beskow, Joakim Gustafson |
INTERSPEECH | 6 |
| 2017 | Synthesising Uncertainty: The Interplay of Vocal Effort and Hesitation DisfluenciesabstractAs synthetic voices become more flexible, and conversational systems gain more potential to adapt to the environmental and social situation, the question needs to be examined, how different modific ... Éva Székely, Joseph Mendelson, Joakim Gustafson |
INTERSPEECH | 3 |
| 2017 | Crowd-Powered Design of Virtual Attentive Listeners
Patrik Jonell, Catharine Oertel, Dimosthenis Kontogiorgos, Jonas Beskow, Joakim Gustafson |
IVA | 5 |
| 2016 | Towards building an attentive artificial listener: on the perception of attentiveness in audio-visual feedback tokensabstractCurrent dialogue systems typically lack a variation of audio-visual feedback tokens. Either they do not encompass feedback tokens at all, or only support a limited set of stereotypical functions. However, this does not mirror the subtleties of spontaneous conversations. If we want to be able to build an artificial listener, as a first step towards building an empathetic artificial agent, we also need to be able to synthesize more subtle audio-visual feedback tokens. In this study, we devised an array of monomodal and multimodal binary comparison perception tests and experiments to understand how different realisations of verbal and visual feedback tokens influence third-party perception of the degree of attentiveness. This allowed us to investigate i) which features (amplitude, frequency, duration...) of the visual feedback influences attentiveness perception; ii) whether visual or verbal backchannels are perceived to be more attentive iii) whether the fusion of unimodal tokens with low perceived attentiveness increases the degree of perceived attentiveness compared to unimodal tokens with high perceived attentiveness taken alone; iv) the automatic ranking of audio-visual feedback token in terms of conveyed degree of attentiveness. Catharine Oertel, José Lopes 0001, Yu Yu 0003, Kenneth Alberto Funes Mora, Joakim Gustafson, Alan W. Black, Jean-Marc Odobez |
ICMI | 5 |
| 2016 | Towards Building an Attentive Artificial Listener: On the Perception of Attentiveness in Feedback UtterancesabstractTowards Building an Attentive Artificial Listener: On the Perception of Attentiveness in Feedback Utterances Catharine Oertel, Joakim Gustafson, Alan W. Black |
INTERSPEECH | 2 |
| 2016 | Hidden Resources ― Strategies to Acquire and Exploit Potential Spoken Language Resources in National Archives
Jens Edlund, Joakim Gustafson |
LREC | 2 |
| 2015 | Deciphering the Silent Participant: On the Use of Audio-Visual Cues for the Classification of Listener Categories in Group DiscussionsabstractEstimating a silent participant's degree of engagement and his role within a group discussion can be challenging, as there are no speech related cues available at the given time. Having this information available, however, can provide important insights into the dynamics of the group as a whole. In this paper, we study the classification of listeners into several categories (attentive listener, side participant and bystander). We devised a thin-sliced perception test where subjects were asked to assess listener roles and engagement levels in 15-second video-clips taken from a corpus of group interviews. Results show that humans are usually able to assess silent participant roles. Using the annotation to identify from a set of multimodal low-level features, such as past speaking activity, backchannels (both visual and verbal), as well as gaze patterns, we could identify the features which are able to distinguish between different listener categories. Moreover, the results show that many of the audio-visual effects observed on listeners in dyadic interactions, also hold for multi-party interactions. A preliminary classifier achieves an accuracy of 64 %. Catharine Oertel, Kenneth Alberto Funes Mora, Joakim Gustafson, Jean-Marc Odobez |
ICMI | 3 |
| 2015 | Detecting repetitions in spoken dialogue systems using phonetic distancesabstractThis paper addresses the problem of automatic detection of re-peated turns in Spoken Dialogue Systems. Repetitions can be a symptom of problematic communication between users and systems. Such repetitions are often due to speech recognition errors, which in turn makes it hard to use speech recognition to detect repetitions. We present an approach to detect rep-etition using the phonetic distance to find the best alignment between turns in the same dialogue. The alignment score ob-tained is combined with different features to improve repeti-tion detection. To evaluate the method proposed we compare several alignment techniques from edit distance to DTW-based distance, previously used in Spoken-Term detection tasks. We also compare two different methods to compute the phonetic distance: the first one using the phoneme sequence, and the second one using the distance between the phone posterior vec-tors. Two different datasets were used in this evaluation: a bus-schedule information system (in English) and a call routing system (in Swedish). The results show that approaches using phoneme distances over-perform approaches using Levenshtein distances between ASR outputs for repetition detection. Index Terms: spoken dialogue systems, repetition detection, phonetic distance José Lopes 0001, Giampiero Salvi, Gabriel Skantze, Alberto Abad, Joakim Gustafson, Fernando Batista, Raveesh Meena, Isabel Trancoso |
INTERSPEECH | 5 |
| 2015 | Automatic Detection of Miscommunication in Spoken Dialogue SystemsabstractIn this paper, we present a data-driven approach for detecting instances of miscommunication in dialogue system interactions.A range of generic features that are both automatically extractable and manually annotated were used to train two models for online detection and one for offline analysis.Online detection could be used to raise the error awareness of the system, whereas offline detection could be used by a system designer to identify potential flaws in the dialogue design.In experimental evaluations on system logs from three different dialogue systems that vary in their dialogue strategy, the proposed models performed substantially better than the majority class baseline models. Raveesh Meena, José Lopes 0001, Gabriel Skantze, Joakim Gustafson |
SIGDIAL Conference | 4 |
| 2014 | Human-robot collaborative tutoring using multiparty multimodal spoken dialogueabstractIn this paper, we describe a project that explores a novel experimental setup towards building a spoken, multi-modally rich, and human-like multiparty tutoring robot. A human-robot interaction setup is designed, and a human-human dialogue corpus is collected. The corpus targets the development of a dialogue system platform to study verbal and nonverbal tutoring strategies in multiparty spoken interactions with robots which are capable of spoken dialogue. The dialogue task is centered on two participants involved in a dialogue aiming to solve a card-ordering game. Along with the participants sits a tutor (robot) that helps the participants perform the task, and organizes and balances their interaction. Different multimodal signals captured and auto-synchronized by different audio-visual capture technologies, such as a microphone array, Kinects, and video cameras, were coupled with manual annotations. These are used build a situated model of the interaction based on the participants personalities, their state of attention, their conversational engagement and verbal dominance, and how that is correlated with the verbal and visual feed-back, turn-management, and conversation regulatory actions generated by the tutor. Driven by the analysis of the corpus, we will show also the detailed design methodologies for an affective, and multimodally rich dialogue system that allows the robot to measure incrementally the attention states, and the dominance for each participant, allowing the robot head Furhat to maintain a well-coordinated, balanced, and engaging conversation, that attempts to maximize the agreement and the contribution to solve the task. Samer Al Moubayed, Jonas Beskow, Bajibabu Bollepalli, Joakim Gustafson, Ahmed Hussen Abdelaziz, Martin Johansson, Maria Koutsombogera, José Lopes 0001, Jekaterina Novikova, Catharine Oertel, Gabriel Skantze, Kalin Stefanov, Gül Varol |
HRI | 4 |
| 2014 | A comparative evaluation of vocoding techniques for HMM-based laughter synthesisabstractThis paper presents an experimental comparison of various leading vocoders for the application of HMM-based laughter synthesis. Four vocoders, commonly used in HMM-based speech synthesis, are used in copy-synthesis and HMM-based synthesis of both male and female laughter. Subjective evaluations are conducted to assess the performance of the vocoders. The results show that all vocoders perform relatively well in copy-synthesis. In HMM-based laughter synthesis using original phonetic transcriptions, all synthesized laughter voices were significantly lower in quality than copy-synthesis, indicating a challenging task and room for improvements. Interestingly, two vocoders using rather simple and robust excitation modeling performed the best, indicating that robustness in speech parameter extraction and simple parameter representation in statistical modeling are key factors in successful laughter synthesis. Bajibabu Bollepalli, Jérôme Urbain, Tuomo Raitio, Joakim Gustafson, Hüseyin Çakmak |
ICASSP | 4 |
| 2014 | Crowdsourcing Street-level Geographic Information Using a Spoken Dialogue SystemabstractWe present a technique for crowdsourcing street-level geographic information using spoken natural language.In particular, we are interested in obtaining first-person-view information about what can be seen from different positions in the city.This information can then for example be used for pedestrian routing services.The approach has been tested in the lab using a fully implemented spoken dialogue system, and has shown promising results. Raveesh Meena, Johan Boye, Gabriel Skantze, Joakim Gustafson |
SIGDIAL Conference | 4 |
| 2014 | Data-driven models for timing feedback responses in a Map Task dialogue system
Raveesh Meena, Gabriel Skantze, Joakim Gustafson |
Comput. Speech Lang. | 3 |
| 2013 | Analysis of gaze and speech patterns in three-party quiz game interactionabstractIn order to understand and model the dynamics between interaction phenomena such as gaze and speech in face-to-face multiparty interaction between humans, we need large quantities of reliable, objective data of such interactions. To date, this type of data is in short supply. We present a data collection setup using automated, objective techniques in which we capture the gaze and speech patterns of triads deeply engaged in a high-stakes quiz game. The resulting corpus consists of five one-hour recordings, and is unique in that it makes use of three state-of-the-art gaze trackers (one per subject) in combination with a state-of-theart conical microphone array designed to capture roundtable meetings. Several video channels are also included. In this paper we present the obstacles we encountered and the possibilities afforded by a synchronised, reliable combination of large-scale multi-party speech and gaze data, and an overview of the first analyses of the data. Index Terms: multimodal corpus, multiparty dialogue, gaze patterns, multiparty gaze. Samer Al Moubayed, Jens Edlund, Joakim Gustafson |
INTERSPEECH | 3 |
| 2013 | The Map Task Dialogue System: A Test-bed for Modelling Human-Like Dialogue
Raveesh Meena, Gabriel Skantze, Joakim Gustafson |
SIGDIAL Conference | 3 |
| 2013 | A Data-driven Model for Timing Feedback in a Map Task Dialogue System
Raveesh Meena, Gabriel Skantze, Joakim Gustafson |
SIGDIAL Conference | 3 |
| 2013 | Semi-supervised methods for exploring the acoustics of simple productive feedback
Daniel Neiberg, Giampiero Salvi, Joakim Gustafson |
Speech Commun. | 3 |
| 2012 | Multimodal multiparty social interaction with the furhat headabstractWe will show in this demonstrator an advanced multimodal and multiparty spoken conversational system using Furhat, a robot head based on projected facial animation. Furhat is a human-like interface that utilizes facial animation for physical robot heads using back-projection. In the system, multimodality is enabled using speech and rich visual input signals such as multi-person real-time face tracking and microphone tracking. The demonstrator will showcase a system that is able to carry out social dialogue with multiple interlocutors simultaneously with rich output signals such as eye and head coordination, lips synchronized speech synthesis, and non-verbal facial gestures used to regulate fluent and expressive multiparty conversations. Samer Al Moubayed, Gabriel Skantze, Jonas Beskow, Kalin Stefanov, Joakim Gustafson |
ICMI | 5 |
| 2012 | On the effect of the acoustic environment on the accuracy of perception of speaker orientation from auditory cues aloneabstractThe ability of people, and of machines, to determine the position of a sound source in a room is well studied. The related ability to determine the orientation of a directed sound source, on the other hand, is not, but the few studies there are show people to be surprisingly skilled at it. This has bearing for studies of face-to- face interaction and of embodied spoken dialogue systems, as sound source orientation of a speaker is connected to the head pose of the speaker, which is meaningful in a number of ways. The feature most often implicated for detection of sound source orientation is the inter-aural level difference - a feature which it is assumed is more easily exploited in anechoic chambers than in everyday surroundings. We expand here on our previous studies and compare detection of speaker orientation within and outside of the anechoic chamber. Our results show that listeners find the task easier, rather than harder, in everyday surroundings, which suggests that inter-aural level differences is not the only feature at play. Jens Edlund, Mattias Heldner, Joakim Gustafson |
INTERSPEECH | 3 |
| 2012 | A Data-driven Approach to Understanding Spoken Route Directions in Human-Robot DialogueabstractIn this paper, we present a data-driven chunking parser for automatic interpretation of spoken route directions into a route graph that is useful for robot navigation. Different sets of features and machine learning algorithms are explored. The results indicate that our approach is robust to speech recognition errors. Index Terms: spoken language understanding, route directions, human-robot interaction Raveesh Meena, Gabriel Skantze, Joakim Gustafson |
INTERSPEECH | 3 |
| 2012 | Gaze Patterns in Turn-TakingabstractOertel C, Wlodarczak M, Edlund J, Wagner P, Gustafson J. Gaze patterns in turn-taking. In: 13th Annual Conference of the International Speech Communication Association 2012 (INTERSPEECH 2012). Red Hook, NY: Curran; 2013: 2243-2246. Catharine Oertel, Marcin Wlodarczak, Jens Edlund, Petra Wagner, Joakim Gustafson |
INTERSPEECH | 5 |
| 2011 | Tracking Pitch Contours Using Minimum Jerk TrajectoriesabstractThis paper proposes a fundamental frequency tracker, with the specific purpose of comparing the automatic estimates with pitch contours that are sketched by trained phoneticians. The method uses a frequency domain approach to estimate pitch tracks that form minimum jerk trajectories. This method tries to mimic motor movements of the hand made while sketch-ing. When the fundamental frequency tracked by the proposed method on the oral and laryngograph signals were compared using the MOCHA-TIMIT database, the correlation was 0.98 and the root mean squared error was 4.0 Hz, which was slightly better than a state-of-the-art pitch tracking algorithm included in the ESPS. We also demonstrate how the proposed algorithm could to be applied when comparing with sketches made by phoneticians for the variations in accent II among the Swedish dialects. Index Terms: pitch tracking, Constant-Q, Swedish accent II 1. Daniel Neiberg, Gopal Ananthakrishnan, Joakim Gustafson |
INTERSPEECH | 3 |
| 2011 | Predicting Speaker Changes and Listener Responses with and without Eye-ContactabstractThis paper compares turn-taking in terms of timing and prediction in human-human conversations under the conditions when participants has eye-contact versus when there is no eyecontact, as found in the HCRC Map Task corpus. By measuring between speaker intervals it was found that a larger proportion of speaker shifts occurred in overlap for the no eyecontact condition. For prediction we used prosodic and spectral features parametrized by time-varying length-invariant discrete cosine coefficients. With Gaussian Mixture Modeling and variations of classifier fusion schemes, we explored the task of predicting whether there is an upcoming speaker change (SC) or not (HOLD), at the end of an utterance (EOU) with a pause lag of 200 ms. The label SC was further split into LRs (listener responses, e.g. back-channels) and other TURN-SHIFTs. The prediction was found to be somewhat easier for the eye-contact condition, for which the average recall rates were Daniel Neiberg, Joakim Gustafson |
INTERSPEECH | 2 |
| 2011 | A Dual Channel Coupled Decoder for Fillers and FeedbackabstractThis study presents a dual channel decoder capable of modeling cross-speaker dependencies for segmentation and classification of fillers and feedbacks in conversational speech found in the DEAL cor ... Daniel Neiberg, Joakim Gustafson |
INTERSPEECH | 2 |
| 2011 | Enhanced visual scene understanding through human-robot dialogabstractWe propose a novel human-robot-interaction framework for robust visual scene understanding. Without any a-priori knowledge about the objects, the task of the robot is to correctly enumerate how many of them are in the scene and segment them from the background. Our approach builds on top of state-of-the-art computer vision methods, generating object hypotheses through segmentation. This process is combined with a natural dialog system, thus including a `human in the loop' where, by exploiting the natural conversation of an advanced dialog system, the robot gains knowledge about ambiguous situations. We present an entropy-based system allowing the robot to detect the poorest object hypotheses and query the user for arbitration. Based on the information obtained from the human-robot dialog, the scene segmentation can be re-seeded and thereby improved. We present experimental results on real data that show an improved segmentation performance compared to segmentation without interaction. Matthew Johnson-Roberson, Jeannette Bohg, Gabriel Skantze, Joakim Gustafson, Rolf Carlson, Babak Rasolzadeh, Danica Kragic |
IROS | 4 |
| 2010 | The prosody of Swedish conversational gruntsabstractThis paper explores conversational grunts in a face-to-face setting. The study investigates the prosody and turn-taking effect of fillers and feedback tokens that has been annotated for attitudes. The grunts were selected from the DEAL corpus and automatically annotated for their turn taking effect. A novel suprasegmental prosodic signal representation and contextual timing features are used for classification and visualization. Classification results using linear discriminant analysis, show that turn-initial feedback tokens lose some of their attitude-signaling prosodic cues compared to non-overlapping continuer feedback tokens. Turn taking effects can be predicted well over chance level, except Simultaneous Starts. However, feedback tokens before places where both speakers take the turn were more similar to feedback continuers than to turn initial feedback tokens. Daniel Neiberg, Joakim Gustafson |
INTERSPEECH | 2 |
| 2009 | The MonAMI reminder: a spoken dialogue system for face-to-face interactionabstractWe describe the MonAMI Reminder, a multimodal spoken dialogue system which can assist elderly and disabled people in organising and initiating their daily activities. Based on deep interviews with potential users, we have designed a calendar and reminder application which uses an innovative mix of an embodied conversational agent, digital pen and paper, and the web to meet the needs of those users as well as the current constraints of speech technology. We also explore the use of head pose tracking for interaction and attention control in human-computer face-to-face interaction. Jonas Beskow, Jens Edlund, Björn Granström, Joakim Gustafson, Gabriel Skantze, Helena Tobiasson |
INTERSPEECH | 4 |
| 2009 | Eliciting Interactional Phenomena in Human-Human Dialogues
Joakim Gustafson, Miray Merkes |
SIGDIAL Conference | 1 |
| 2009 | Attention and Interaction Control in a Human-Human-Computer Dialogue Setting
Gabriel Skantze, Joakim Gustafson |
SIGDIAL Conference | 2 |
| 2008 | Innovative interfaces in MonAMI: the reminderabstractThis demo paper presents an early version of the Reminder, a prototype ECA developed in the European project MonAMI, which aims at "mainstreaming accessibility in consumer goods and services, using advanced technologies to ensure equal access, independent living and participation for all". The Reminder helps users to plan activities and to remember what to do. The prototype merges mobile ECA technology with other, existing technologies: Google Calendar and a digital pen and paper. The solution allows users to continue using a paper calendar in the manner they are used to, whilst the ECA provides notifications on what has been written in the calendar. Users may ask questions such as "When was I supposed to meet Sara?" or "What's my schedule today?" Jonas Beskow, Jens Edlund, Teodore Gjermani, Björn Granström, Joakim Gustafson, Oskar Jonsson, Gabriel Skantze, Helena Tobiasson |
ICMI | 5 |
| 2008 | What makes a good speaker? subject ratings, acoustic measurements and perceptual evaluationsabstractThis paper deals with subjective qualities and acoustic-prosodic features contributing to the impression of a good speaker. Subjects rated a variety of samples of political speech on a number of su ... Eva Strangert, Joakim Gustafson |
INTERSPEECH | 2 |
| 2008 | Towards human-like spoken dialogue systems
Jens Edlund, Joakim Gustafson, Mattias Heldner, Anna Hjalmarsson |
Speech Commun. | 2 |
| 2007 | Children's convergence in referring expressions to graphical objects in a speech-enabled computer gameabstractThis paper describes an empirical study of children's spontaneous interactions with an animated character in a speech-enabled computer game. More specifically, it deals with convergence of referrin ... Linda Bell, Joakim Gustafson |
INTERSPEECH | 2 |
| 2006 | Robust spoken language understanding in a computer game
Johan Boye, Joakim Gustafson, Mats Wirén |
Speech Commun. | 2 |
| 2005 | The Swedish NICE corpus - spoken dialogues between children and embodied characters in a computer game scenarioabstractThis article describes the collection and analysis of a Swedish database of spontaneous and unconstrained children-machine dialogues. The Swedish NICE corpus consists of spoken dialogues between children aged 8 to 15 and embodied fairytale characters in a computer game scenario. Compared to previously collected corpora of children's computer-directed speech, the Swedish NICE corpus contains extended interactions, including three-party conversation, in which the young users used spoken dialogue as the primary means of progression in the game. Linda Bell, Johan Boye, Joakim Gustafson, Mattias Heldner, Anders Lindström, Mats Wirén |
INTERSPEECH | 3 |
| 2005 | Providing Computer Game Characters with Conversational Abilities
Joakim Gustafson, Johan Boye, Morgan Fredriksson, Lasse Johannesson, Jürgen Königsmann |
IVA | 1 |
| 2003 | Child and adult speaker adaptation during error resolution in a publicly available spoken dialogue systemabstractThis paper describes how speakers adapt their language during error resolution when interacting with the animated agent Pixie.A corpus of spontaneous human-computer interaction was collected at the Telecommunication museum in Stockholm, Sweden.Adult and children speakers were compared with respect to user behavior and strategies during error resolution.In this study, 16 adults and 16 children speakers were randomly selected from a corpus from almost 3.000 speakers.This sub-corpus was then analyzed in greater detail.Results indicate that adults and children use partly different strategies when their interactions with Pixie become problematic.Children tend to repeat the same utterance verbatim, altering certain phonetic features.Adults, on the other hand, often modify other aspects of their utterances such as lexicon and syntax.Results from the present study will be useful for constructing future spoken dialogue systems with improved error handling for adults as well as children. Linda Bell, Joakim Gustafson |
INTERSPEECH | 2 |
| 2002 | Voice transformations for improving children²s speech recognition in a publicly available dialogue systemabstractTo be able to build acoustic models for children, that can beused in spoken dialogue systems, speech data has to be collected. Commercial recognizers available for Swedish are trained on adult speech, which makes them less suitable for children’s computer-directed speech. This paper describes some experiments with on-the-fly voice transformation of children’s speech. Two transformation methods were tested, one inspired by the Phase Vocoder algorithm and another by the Time-Domain Pitch-Synchronous Overlap-Add (TD-PSOLA)algorithm. The speech signal is transformed before being sent to the speech recognizer for adult speech. Our results show that this method reduces the error rates in the order of thirty to fortyfive percent for children users. Joakim Gustafson, Kåre Sjölander |
INTERSPEECH | 1 |
| 2000 | A comparison of disfluency distribution in a unimodal and a multimodal speech interfaceabstractIn this paper, we compare the distribution of disfluencies in two human--computer dialogue corpora. One corpus consists of unimodal travel booking dialogues, which were recorded over the telephone. In this unimodal system, all components except the speech recognition were authentic. The other corpus was collected using a semi-simulated multi-modal dialogue system with an animated talking agent and a clickable map. The aim of this paper is to analyze and discuss the effects of modality, task and interface design on the distribution and frequency of disfluencies in these two corpora. Linda Bell, Robert Eklund, Joakim Gustafson |
INTERSPEECH | 3 |
| 2000 | Positive and negative user feedback in a spoken dialogue corpusabstractThis paper examines feedback strategies in a Swedish corpus of multimodal human--computer interaction. The aim of the study is to investigate how users provide positive and negative feedback to a dialogue system and to discuss the function of these utterances in the dialogues. User feedback in the AdApt corpus was labeled and analyzed, and its distribution in the dialogues is discussed. The question of whether it is possible to utilize user feedback in future systems is considered. More specifically, we discuss how error handling in human--computer dialogue might be improved through greater knowledge of user feedback strategies. In the present corpus, almost all subjects used positive or negative feedback at least once during their interaction with the system. Our results indicate that some types of feedback more often occur in certain positions in the dialogue. Another observation is that there appear to be great individual variations in feedback strategies, so that certain subjects give feedback at almost every turn while others rarely or never respond to a spoken dialogue system in this manner. Finally, we discuss how feedback could be used to prevent problems in human--computer dialogue. Linda Bell, Joakim Gustafson |
INTERSPEECH | 2 |
| 2000 | Adapt - a multimodal conversational dialogue system in an apartment domainabstractA general overview of the AdApt project and the research that is performed within the project is presented. In this project various aspects of human-computer interaction in a multimodal conversational dialogue systems are investigated. The project will also include studies on the integration of user/system/dialogue dependent speech recognition and multimodal speech synthesis. A domain in which multimodal interaction is highly useful has been chosen, namely, finding available apartments in Stockholm. A Wizard-of-Oz data collection within this domain is also described. 1. Joakim Gustafson, Linda Bell, Jonas Beskow, Johan Boye, Rolf Carlson, Jens Edlund, Björn Granström, David House, Mats Wirén |
INTERSPEECH | 1 |
| 2000 | Speech technology on trial: Experiences from the August systemabstractIn this paper, the August spoken dialogue system is described. This experimental Swedish dialogue system, which featured an animated talking agent, was exposed to the general public during a trial period of six months. The construction of the system was partly motivated by the need to collect genuine speech data from people with little or no previous experience of spoken dialogue systems. A corpus of more than 10,000 utterances of spontaneous computer- directed speech was collected and empirical linguistic analyses were carried out. Acoustical, lexical and syntactical aspects of this data were examined. In particular, user behavior and user adaptation during error resolution were emphasized. Repetitive sequences in the database were analyzed in detail. Results suggest that computer-directed speech during error resolution is increased in duration, hyperarticulated and contains inserted pauses. Design decisions which may have influenced how the users behaved when they interacted with August are discussed and implications for the development of future systems are outlined. Joakim Gustafson, Linda Bell |
Nat. Lang. Eng. | 1 |
| 1999 | Interaction with an animated agent in a spoken dialogue system
Linda Bell, Joakim Gustafson |
EUROSPEECH | 2 |
| 1999 | The august spoken dialogue systemabstractThis paper describes how a telephone-based application can perform a variety of tasks in a completely hands-free mode. The overall architecture of the speech component is multimodal in that each mode is tailored to a specific need of the interface. The various modes are described as well as the underlying core technology. To illustrate the effectiveness of the implementation, we present experimental results in American English, UK English and French on a variety of benchmarks, including live data collected during actual use of the system. Joakim Gustafson, Nikolaj Lindberg, Magnus Lundeberg |
EUROSPEECH | 1 |
| 1998 | An educational dialogue system with a user controllable dialogue managerabstractWe have developed an educational environment for a modular spoken dialogue system. The aim of the environment is to provide students, with different backgrounds, means to understand the behaviour of spoken dialogue systems. Focus in this paper is on dialogue and dialogue management. The dialogue is recorded in a dialogue tree whose nodes are dialogue objects. The dialogue objects model the constituents of the dialogue and consist of parameters for modelling dialogue structure, focus structure and a process description describing the actions of the dialogue system. Various dialogue system behaviours can be achieved by modifying these parameters. This is done using the educational environment, which is interactive and facilitates examination, expansion and modification of the dialogue object parameters and hence the system. The educational system has been used in a number of courses at various universities in Sweden. 1. Joakim Gustafson, Patrik Elmberg, Rolf Carlson, Arne Jönsson |
ICSLP | 1 |
| 1998 | Web-based educational tools for speech technologyabstractThis paper describes the efforts at KTH in creating educational tools for speech technology. The demand for such tools is increasing with the advent of speech as a medium for man– machine communication. The world wide web was chosen as our platform in order to increase the usability and accessibility of our computer exercises. The aim was to provide dedicated educational software instead of exercises based on complex research tools. Currently, the set of exercises comprises basic speech analysis, multi-modal speech synthesis and spoken dialogue systems. Students access web pages in which the exercises have been embedded as applets. This makes it possible to use them in a classroom setting, as well as from the students’ home computers. Kåre Sjölander, Jonas Beskow, Joakim Gustafson, Erland Lewin, Rolf Carlson, Björn Granström |
ICSLP | 3 |
| 1997 | How do system questions influence lexical choices in user answers?abstractThis paper describes some studies on the effect of the system vocabulary on the lexical choices of the users. There are many theories about human-human dialogues that could be useful in the design of spoken dialoguesystems. This paper will give an overview of some of these theories and report the results from two experiments that examines one of these theories, namely lexical entrainment. The first experiment was a small Wizard of Oz-test that simulated a tourist informationsystem with a speech interface, and the second experiment simulated a system with speech recognition that controlled a questionnaire about peoples plans for their vacation. Both experiments show that the subjects mostly adapt their lexical choices to the system questions. Only in less than 5% of the cases did they use an alternative main verb in the answer. These results encourage us to investigate the possibility to add anadaptive language model in the speech recognizer in our dialogue system, where the probabilities for the words used in the system questions are increased. Joakim Gustafson, Anette Larsson, Rolf Carlson, K. Hellman |
EUROSPEECH | 1 |
| 1997 | An integrated system for teaching spoken dialogue systems technology
Kåre Sjölander, Joakim Gustafson |
EUROSPEECH | 2 |
| 1995 | The waxholm application databaseabstractThis paper describes an application database collected in Wizard-of-Oz experiments in a spoken dialogue system, WAXHOLM. The system provides information on boat traffic in the Stockholm archipelago. The database consists of utterance-length speech files, their corresponding transcriptions, and log files of the dialogue sessions. In addition to the spontaneous dialogue speech, the material also comprise recordings of phonetically balanced reference sentences uttered by all 66 subjects. In the paper the recording procedure is described as well as some characteristics of the speech data and the dialogue. INTRODUCTION WAXHOLM is a demonstrator spoken dialogue system in which we apply our research on speech recognition and speech synthesis. The system uses visual and auditory means to provide information on boat traffic, accomodation, etc. in the Stockholm archipelago [1], [2], [3]. The application has great similarities to the ATIS domain within the ARPA community, the Voyager system from... J. Bertenstam, Mats Blomberg, Rolf Carlson, Kjell Elenius, Björn Granström, Joakim Gustafson, Sheri Hunnicutt, Jesper Högberg, Roger Lindell, Lennart Neovius, Lennart Nord, Antonio de Serpa-Leitao, Nikko Strom |
EUROSPEECH | 6 |
| 1995 | Using two-level morphology to transcribe Swedish namesabstractNames are difficult to handle for normal letter-to-sound rules, since these usually are designed for ordinary words. The structure of Swedish names differ from ordinary words - but their multi-morphemic structure make them suitable to analyse with a morphological analyser. Joakim Gustafson |
EUROSPEECH | 1 |
| 1993 | An experimental dialogue system: waxholmabstractRecently we have begun to build the basic tools for a generic speech-dialogue system, WAXHOLM. The main modules, their function and internal communication have been specified. The different components are connected through a computer network. A preliminary version of the system has been tested, using simplified versions of the modules. We will give a general overview of the system and describe some of the components in more detail. Application specific data are collected with the help of Wizard-of-Oz techniques. The dialogue system is used during the data collection and the wizard only replaces the speech-recognition module. Mats Blomberg, Rolf Carlson, Kjell Elenius, Björn Granström, Joakim Gustafson, Sheri Hunnicutt, Roger Lindell, Lennart Neovius |
EUROSPEECH | 5 |