Gabriel Skantze

dblp:54/3812 · DBLP profile ↗
← Back
93ranked-venue papers
20as first author
42since 2021 · last 2026
0000-0002-8579-1790ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 79 · 16 first-author · 36 since 2021Human-computer interaction and ubiquitous computing · 33 · 5 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Aligning Backchannel and Dialogue Context Representations via Contrastive LLM Fine-Tuning
abstract
Backchannels (e.g., 'yeah', 'mhm', and 'right') are short, non-interruptive feedback signals whose lexical form and prosody jointly convey pragmatic meaning.While prior computational research has largely focused on predicting backchannel timing, the relationship between lexico-prosodic form and meaning remains underexplored.We propose a two-stage framework: first, fine-tuning large language models on dialogue transcripts to derive rich contextual representations; and second, learning a joint embedding space for dialogue contexts and backchannel realizations.We evaluate alignment with human perception via triadic similarity judgments (prosodic and cross-lexical) and a context-backchannel suitability task.Our results demonstrate that the learned projections substantially improve context-backchannel retrieval compared to previous methods.In addition, they reveal that backchannel form is highly sensitive to extended conversational context and that the learned embeddings align more closely with human judgments than raw WavLM features.Pitch length: 89.1 Pitch range: 14.
Livia Qian, Gabriel Skantze
ACL (1)2
2026 Investigating the Representation of Backchannels and Fillers in Fine-tuned Language Models
abstract
Backchannels and fillers are important linguistic expressions in dialogue, but often treated as 'noise' to be bypassed in modern transformerbased language models (LMs).Here, we study how they are represented in LMs using three fine-tuning strategies on three dialogue corpora in English and Japanese, in which backchannels and fillers are both preserved and annotated.This allows us to investigate how fine-tuning can help LMs learn these representations.We first apply clustering analysis to the learnt representation of backchannels and fillers, and find increased silhouette scores in representations from fine-tuned models, which suggests that fine-tuning enables LMs to distinguish the nuanced semantic variation in different backchannel and filler use.We also employ natural language generation metrics and qualitative analyses to verify that utterances produced by fine-tuned LMs resemble those produced by humans more closely.Our findings suggest the potential for transforming general LMs into conversational LMs that can produce human-like language more adequately.
Yu Wang 0294, Leyi Lao, Langchu Huang, Gabriel Skantze, Yang Xu 0024, Hendrik Buschmeier
ACL (1)4
2026 Human-Robot Interaction Conversational User Enjoyment Scale (HRI CUES)
abstract
Understanding user enjoyment is crucial in human-robot interaction (HRI), as it can impact interaction quality and influence user acceptance and long-term engagement with robots, particularly in the context of conversations with social robots. However, current assessment methods rely solely on self-reported questionnaires, failing to capture interaction dynamics. This work introduces the Human-Robot Interaction Conversational User Enjoyment Scale (HRI CUES), a novel 5-point scale to assess user enjoyment from an external perspective (e.g.by an annotator) for conversations with a robot. The scale was developed through rigorous evaluations and discussions among three annotators with relevant expertise, using open-domain conversations with a companion robot that was powered by a large language model, and was applied to each conversation exchange (i.e.a robot-participant turn pair) alongside overall interaction. It was evaluated on 25 older adults' interactions with the companion robot, corresponding to 174 minutes of data, showing moderate to good alignment between annotators. Although the scale was developed and tested in the context of older adult interactions with a robot, its basis in general and non-task-specific indicators of enjoyment supports its broader applicability. The study further offers insights into understanding the nuances and challenges of assessing user enjoyment in robot interactions, and provides guidelines on applying the scale to other domains and populations. The dataset is available online.
Bahar Irfan, Jura Miniota, Sofia Thunberg, Erik Lagerstedt, Sanna Kuoppamäki, Gabriel Skantze, André Pereira 0001
IEEE Trans. Affect. Comput.6
2026 Robots as Hosts in Autonomous Buses: A Field Trial
abstract
In Autonomous Public Transport (APT), particularly with shuttle buses, passengers travel in smaller, more intimate vehicles—and in the future, such vehicles may operate without an authoritative driver or host. This setup may lead to potential safety concerns, as passengers are left alone together. Additionally, this future absence of a driver or host means that there is no one to address questions or uncertainties that may arise. One proposed solution is introducing a robot onboard the bus, serving a similar role to a human host. To explore this solution, an experiment was conducted in Barkarby, Stockholm, Sweden. Passengers, generally unfamiliar with APT or social robots, experienced two short rides on a bus equipped with either an embodied Furhat robot as the host or a disembodied voice agent in the ceiling. Data were collected from passenger-agent interactions, post-questionnaires, and semi-structured focus group interviews. Results indicate a division in passenger preferences, with some favoring the robot and others the voice assistant. Passengers asked more questions to the robot, suggesting a clearer affordance for interaction. While the questionnaires did not show significant differences, passenger behaviors indicated that they anthropomorphized the robot more. The interviews revealed that passengers felt more secure with a human operator and doubted the robot’s authority during incidents with aggressive passengers or accidents. Our findings show that social robots can help make autonomous buses feel more welcoming and interactive. Future APT systems have many design issues that need to be resolved before riders can find them safe and appropriate to use, and social robots can play a role in resolving such issues—both the ones we see today, and potentially ones that will appear in the future.
Agnes Johanna Axelsson, Bhavana Vaddadi, Cristian Bogdan, Deirdre Tobin, Gabriel Skantze
ACM Trans. Hum. Robot Interact.5
2025 Linguistic Anthropomorphism in Chatbots: Effects of Style, Topic, and Interaction on Users' Perceptions and Behaviors
abstract
Designing chatbots to use anthropomorphic language offers potential benefits but also entails risks, shaping users’ perceptions and behaviors in sometimes unexpected ways. To examine these effects, we conducted two online experiments using a 3 × 3 between-subjects design. In the first experiment (N = 530), participants read chatbot transcripts; in the second (N = 560), they interacted with the chatbot directly. We varied conversational style (machine-like, human-like formal, human-like casual) and topic (transactional, small-talk, sensitive) to assess users’ perceptions (anthropomorphism, security, competence, warmth, enjoyment, satisfaction, trust) and behaviors (self-disclosure, donation to charity). Overall, human-like styles increased perception of warmth, enjoyment, and trust, but primarily when the style matched the topic. Style–topic congruence enhanced satisfaction and trust, while mismatches reduced perceived competence and enjoyment. Notably, during live interaction, users disclosed more to machine-like chatbots, despite expressing greater preference and trust for human-like ones, suggesting that anthropomorphic cues, while improving perceptions, may increase concerns about social judgment or risk. Moreover, sensitive topics increased donations in the scripted experiment but not during live interaction, highlighting a gap between hypothetical and actual behavior. These findings underscore the nuanced impact of linguistic anthropomorphism and suggest that chatbot effectiveness depends less on maximizing human-likeness and more on aligning conversational style with task, topic, and mode of interaction.
Charlotte Stinkeste, Gabriel Skantze
HAI2
2025 Between You and Me: Ethics of Self-Disclosure in Human-Robot Interaction
abstract
As we move toward a future where robots are increasingly part of daily life, the privacy risks associated with interactions, particularly those relying on cloud-based large language models (LLMs), are becoming more pressing. Users may unknowingly share sensitive information in environments, such as homes or hospitals. To explore these risks, we conducted a study with 39 native English speakers using a Furhat robot with an integrated LLM. Participants discussed two moral dilemmas: (i) dishonesty, sharing personal stories of justified lying, and (ii) robot disobedience, discussing whether robots should disobey commands. On average, participants disclosed personal stories 45% of the time when asked in both scenarios. The main reason for non-disclosure was difficulty recalling examples quickly (33.3-56%), rather than reluctance to share (7.2-16%). However, most participants reported a lack of discomfort and concern about sharing personal information with the robot, indicating limited awareness of the privacy risks involved in such disclosures.
Bahar Irfan, Gabriel Skantze
HRI2
2025 Online Prediction of User Enjoyment in Human-Robot Dialogue with LLMs
abstract
Large Language Models (LLMs) allow social robots to engage in unconstrained open-domain dialogue, but often make mistakes when employed in real-world interactions, requiring adaptation of LLMs to specific conversational contexts. However, LLM adaptation techniques require a feedback signal, ideally for multiple alternative utterances. At the same time, human-robot dialogue data is scarce and research often relies on external annotators. A tool for automatic prediction of user enjoyment in human-robot dialogue is therefore needed. We investigate the possibility of predicting user enjoyment turn-by-turn using an LLM, giving it a proposed robot utterance within the dialogue context, but without access to user response. We compare this performance to the system's enjoyment ratings when user responses are available and to assessments by expert human annotators, in addition to self-reported user perceptions. We evaluate the proposed LLM predictor in a human-robot interaction (HRI) dataset with conversation transcripts of 25 older adults' 7-minute dialogues with a companion robot. Our results show that an LLM is capable of predicting user enjoyment, without loss of performance despite the lack of user response and even achieving performance similar to that of human expert annotators. Furthermore, results show that the system surpasses expert annotators in its correlation with the user's self-reported perceptions of the conversation. This work presents a tool to remove the reliance on external annotators for enjoyment evaluation and paves the way toward real-time adaptation in human-robot dialogue.
Ruben Janssens, André Pereira 0001, Gabriel Skantze, Bahar Irfan, Tony Belpaeme
HRI3
2025 Comparing Monolingual and Bilingual Social Robots as Conversational Practice Companions in Language Learning
abstract
This study explores the impact of monolingual and bilingual robots in Robot-Assisted Language Learning (RALL) for non-native Swedish learners. In a within-group design, 47 participants interacted with a social robot under two conditions: a monolingual robot that communicated exclusively in Swedish and a bilingual robot capable of switching between Swedish and English. Each participant engaged in multiple role-play scenarios designed to match their language proficiency levels, and their experiences were assessed through surveys and behavioral data. The results show that the bilingual robot was generally favored by participants, leading to a more relaxed, enjoyable experience. The perceived learning was improved at the end of the experiment regardless of the condition. These findings suggest that incorporating bilingual support in language-learning robots may enhance user engagement and effectiveness, particularly for lower-proficiency learners.
Alireza Mahmoudi Kamelabad, Elin Inoue, Gabriel Skantze
HRI3
2025 What Can You Say to a Robot? Capability Communication Leads to More Natural Conversations
abstract
When encountering a robot in the wild, it is not inherently clear to human users what the robot's capabilities are. When encountering misunderstandings or problems in spoken interaction, robots often just apologize and move on, without additional effort to make sure the user understands what happened. We set out to compare the effect of two speech based capability communication strategies (proactive, reactive) to a robot without such a strategy, in regard to the user's rating of and their behavior during the interaction. For this, we conducted an in-person user study with 120 participants who had three speech-based interactions with a social robot in a restaurant setting. Our results suggest that users preferred the robot communicating its capabilities proactively and adjusted their behavior in those interactions, using a more conversational interaction style while also enjoying the interaction more.
Merle M. Reimann, Koen V. Hindriks, Florian Kunneman, Catharine Oertel, Gabriel Skantze, Iolanda Leite
HRI5
2025 Applying General Turn-Taking Models to Conversational Human-Robot Interaction
abstract
Turn-taking is a fundamental aspect of conversation, but current Human-Robot Interaction (HRI) systems often rely on simplistic, silence-based models, leading to unnatural pauses and interruptions. This paper investigates, for the first time, the application of general turn-taking models, specifically TurnGPT and Voice Activity Projection (VAP), to improve conversational dynamics in HRI. These models are trained on human-human dialogue data using self-supervised learning objectives, without requiring domain-specific fine-tuning. We propose methods for using these models in tandem to predict when a robot should begin preparing responses, take turns, and handle potential interruptions. We evaluated the proposed system in a within-subject study against a traditional baseline system, using the Furhat robot with 39 adults in a conversational setting, in combination with a large language model for autonomous response generation. The results show that participants significantly prefer the proposed system, and it significantly reduces response delays and interruptions.
Gabriel Skantze, Bahar Irfan
HRI1
2025 Speech-to-Joy: Self-Supervised Features for Enjoyment Prediction in Human-Robot Conversation
Ricardo Santana, Bahar Irfan, Erik Lagerstedt, Gabriel Skantze, André Pereira 0001
ICMI4
2025 "Dyadosyncrasy", Idiosyncrasy and Demographic Factors in Turn-Taking
abstract
Turn-taking in dialogue follows universal constraints but also varies significantly. This study examines how demographic (sex, age, education) and individual factors shape turn-taking using a large dataset of US English conversations (Fisher). We analyze Transition Floor Offset (TFO) and find notable interspeaker variation. Sex and age have small but significant effects — female speakers and older individuals exhibit slightly shorter offsets — while education shows no effect. Lighter topics correlate with shorter TFOs. However, individual differences have a greater impact, driven by a strong idiosyncratic and an even stronger "dyadosyncratic" component — speakers in a dyad resemble each other more than they resemble themselves in different dyads. This suggests that the dyadic relationship and joint activity are the strongest determinants of TFO, outweighing demographic influences.
Julio Cesar Cavalcanti, Gabriel Skantze
INTERSPEECH2
2025 Representation of Perceived Prosodic Similarity of Conversational Feedback
Livia Qian, Carol Figueroa, Gabriel Skantze
INTERSPEECH3
2025 Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
abstract
Koji Inoue, Divesh Lala, Gabriel Skantze, Tatsuya Kawahara. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Koji Inoue, Divesh Lala, Gabriel Skantze, Tatsuya Kawahara
NAACL (Long Papers)3
2025 The Multilingual Student Support Robot
abstract
International students in UK universities often struggle with interactions in English, particularly on their first days in the country. We have developed a multilingual support robot tailored to their needs. To evaluate the performance of the robot, 60 international students asked the robot, in either English or their native language (Modern Standard Arabic or Mandarin Chinese), for support on topics including campus directions, local tax exemption, financial aid, and official documents. Overall, users preferred to use their native language when interacting with the support robot, and using their native language in the robot interaction also had a positive effect on their perception of the interaction itself.
Shaul Ashkenazi, Gabriel Skantze, Jane Stuart-Smith, Mary Ellen Foster
RO-MAN2
2025 Manners Matter: How Robot Politeness Influences Human Risk-Taking and Social Perception
abstract
Robots are no longer confined to factories; they now collaborate, assist, and engage with people in social environments, making their communication style a crucial factor in human-robot interaction. This study examines whether robots’ (im)politeness affects human risk-taking behavior and social perception, key factors in fostering effective human-robot interaction. In a between-subject experiment, sixty participants interacted with either a polite robot, employing politeness strategies, or a rude robot, using face-threatening acts. Risk-taking behavior was assessed through a button-pressing task that allowed participants to accumulate monetary rewards while risking total loss, and social perception was assessed with the Human-Robot Interaction Evaluation Scale (HRIES). Although politeness did not alter risk-taking, it significantly influenced perception: the polite robot was rated as more sociable and agentic, whereas the rude robot was seen as more disturbing; perceptions of animacy remained unchanged. An exploratory factor analysis refined the Agency scale, raising questions about how users conceptualize robotic autonomy. These findings confirm that politeness enhances social acceptance, but may not universally alter behavior. In social domains like customer service and healthcare, politeness is beneficial, but in financial decision-making contexts where outcomes carry real-world consequences, decision support may require additional strategies, such as assertiveness or personalized feedback.
Charlotte Stinkeste, Albin Wikström Kempe, Gabriel Skantze
RO-MAN3
2025 Role of Reasoning in LLM Enjoyment Detection: Evaluation Across Conversational Levels for Human-Robot Interaction
abstract
User enjoyment is central to developing conversational AI systems that can recover from failures and maintain interest over time. However, existing approaches often struggle to detect subtle cues that reflect user experience. Large Language Models (LLMs) with reasoning capabilities have outperformed standard models on various other tasks, suggesting potential benefits for enjoyment detection. This study investigates whether models with reasoning capabilities outperform standard models when assessing enjoyment in a human-robot dialogue corpus at both turn and interaction levels. Results indicate that reasoning capabilities have complex, model-dependent effects rather than universal benefits. While performance was nearly identical at the interaction level (0.44 vs 0.43), reasoning models substantially outperformed at the turn level (0.42 vs 0.36). Notably, LLMs correlated better with users’ self-reported enjoyment metrics than human annotators, despite achieving lower accuracy against human consensus ratings. Analysis revealed distinctive error patterns: non-reasoning models showed bias toward positive ratings at the turn level, while both model types exhibited central tendency bias at the interaction level. These findings suggest that reasoning should be applied selectively based on model architecture and assessment context, with assessment granularity significantly influencing relative effectiveness.
Lubos Marcinek, Bahar Irfan, Gabriel Skantze, André Pereira 0001, Joakim Gustafson
SIGDIAL3
2025 Using LLMs to Grade Clinical Reasoning for Medical Students in Virtual Patient Dialogues
abstract
This paper presents an evaluation of the use of large language models (LLMs) for grading clinical reasoning during rheumatology medical history virtual patient (VP) simulations. The study explores the feasibility of using state-of-the-art LLMs, including both general-purpose models, with various prompting strategies such as zero-shot, analysis-first, and chain-of-thought prompting, as well as reasoning models. The performance of these models in grading transcribed dialogues from VP simulations conducted on a Furhat robot was evaluated against human expert annotations. Human experts initially achieved a 65% inter-rater agreement, which resulted in a pooled Cohen’s Kappa of 0.71 and 82.3% correctness. The best LLM, o3-mini, achieved a pooled Kappa of 0.68 and 81.5% correctness, with response times under 30 seconds, compared to approximately 6 minutes for human grading. These results indicate the possibility that automatic assessments can approach human reliability under controlled simulation conditions while delivering time and cost efficiencies.
Jonathan Schiött, William Ivegren, Alexander Borg, Ioannis Parodis, Gabriel Skantze
SIGDIAL5
2024 Multilingual Turn-taking Prediction Using Voice Activity Projection
abstract
This paper investigates the application of voice activity projection (VAP), a predictive turn-taking model for spoken dialogue, on multilingual data, encompassing English, Mandarin, and Japanese. The VAP model continuously predicts the upcoming voice activities of participants in dyadic dialogue, leveraging a cross-attention Transformer to capture the dynamic interplay between participants. The results show that a monolingual VAP model trained on one language does not make good predictions when applied to other languages. However, a multilingual model, trained on all three languages, demonstrates predictive performance on par with monolingual models across all languages. Further analyses show that the multilingual model has learned to discern the language of the input signal. We also analyze the sensitivity to pitch, a prosodic cue that is thought to be important for turn-taking. Finally, we compare two different audio encoders, contrastive predictive coding (CPC) pre-trained on English, with a recent model based on multilingual wav2vec 2.0 (MMS).
Koji Inoue, Bing'er Jiang, Erik Ekstedt, Tatsuya Kawahara, Gabriel Skantze
LREC/COLING5
2024 Multimodal User Enjoyment Detection in Human-Robot Conversation: The Power of Large Language Models
abstract
Enjoyment is a crucial yet complex indicator of positive user experience in Human-Robot Interaction (HRI). While manual enjoyment annotation is feasible, developing reliable automatic detection methods remains a challenge. This paper investigates a multimodal approach to automatic enjoyment annotation for HRI conversations, leveraging large language models (LLMs), visual, audio, and temporal cues. Our findings demonstrate that both text-only and multimodal LLMs with carefully designed prompts can achieve performance comparable to human annotators in detecting user enjoyment. Furthermore, results reveal a stronger alignment between LLM-based annotations and user self-reports of enjoyment compared to human annotators. While multimodal supervised learning techniques did not improve all of our performance metrics, they could successfully replicate human annotators and highlighted the importance of visual and audio cues in detecting subtle shifts in enjoyment. This research demonstrates the potential of LLMs for real-time enjoyment detection, paving the way for adaptive companion robots that can dynamically enhance user experiences.
André Pereira 0001, Lubos Marcinek, Jura Miniota, Sofia Thunberg, Erik Lagerstedt, Joakim Gustafson, Gabriel Skantze, Bahar Irfan
ICMI7
2024 Referring Expression Generation in Visually Grounded Dialogue with Discourse-aware Comprehension Guiding
abstract
We propose an approach to referring expression generation (REG) in visually grounded dialogue that is meant to produce referring expressions (REs) that are both discriminative and discourse-appropriate. Our method constitutes a two-stage process. First, we model REG as a text- and image-conditioned next-token prediction task. REs are autoregressively generated based on their preceding linguistic context and a visual representation of the referent. Second, we propose the use of discourse-aware comprehension guiding as part of a generate-and-rerank strategy through which candidate REs generated with our REG model are reranked based on their discourse-dependent discriminatory power. Results from our human evaluation indicate that our proposed two-stage approach is effective in producing discriminative REs, with higher performance in terms of text-image retrieval accuracy for reranked REs compared to those generated using greedy decoding.
Bram Willemsen, Gabriel Skantze
INLG2
2024 Joint Learning of Context and Feedback Embeddings in Spoken Dialogue
Livia Qian, Gabriel Skantze
INTERSPEECH2
2024 Conformity and Trust in Multi-party vs. Individual Human-Robot Interaction
abstract
In this study, we explored how conformity and trust vary in adolescent students’ interactions with a social robot. Specifically, we compared how this was influenced by whether the participants had individual or multi-party interaction with robot and whether the robot was portrayed as an adult or a child through appearance and voice. Our experiment involved 75 Swedish middle school students participating in a card sorting game with the Furhat robot, where the objective was to discuss and reach an agreement on the card sequence. The data analysis focused firstly on the participants’ willingness to rearrange cards following the robot’s suggestions and secondly their post-session subjective trust in the robot’s advice. Results indicated that individuals interacting with the robot individually were more likely to conform to its suggestions than those interacting with it together with a peer. Individuals interacting alone with the robot also showed higher post-session trust levels than those in multi-party settings, indicating group size impacts robot trustworthiness perceptions. However, the robot’s perceived age did not affect the level of conformity. Exploratory analyses also showed that mutual understanding was lower in the multi-party setting, while the child robot condition improved user experience, highlighting the complex influence of group dynamics and robot portrayal on human-robot interactions in education.
Alireza Mahmoudi Kamelabad, Olov Engwall, Gabriel Skantze
IVA3
2024 Mhm... Yeah? Okay! Evaluating the Naturalness and Communicative Function of Synthesized Feedback Responses in Spoken Dialogue
abstract
To create conversational systems with humanlike listener behavior, generating short feedback responses (e.g., "mhm", "ah", "wow") appropriate for their context is crucial.These responses convey their communicative function through their lexical form and their prosodic realization.In this paper, we transplant the prosody of feedback responses from humanhuman U.S. English telephone conversations to a target speaker using two synthesis techniques (TTS and signal processing).Our evaluation focuses on perceived naturalness, contextual appropriateness and preservation of communicative function.Results indicate TTS-generated feedback were perceived as more natural than signal-processing-based feedback, with no significant difference in appropriateness.However, the TTS did not consistently convey the communicative function of the original feedback.
Carol Figueroa, Marcel de Korte, Magalie Ochs, Gabriel Skantze
SIGDIAL4
2023 Why is my Agent so Slow? Deploying Human-Like Conversational Turn-Taking
abstract
The emphasis on one-to-one speak/wait spoken conversational interaction with intelligent agents leads to long pauses between conversational turns, undermines the flow and naturalness of the interaction, and undermines the user experience. Despite ground breaking advances in the area of generating and understanding natural language with techniques such as LLMs, conversational interaction has remained relatively overlooked. In this workshop we will discuss and review the challenges, recent work and potential impact of improving conversational interaction with artificial systems. We hope to share experiences of poor human/system interaction, best practices with third party tools, and generate design guidance for the community.
Matthew P. Aylett, Éva Székely, Donald McMillan, Gabriel Skantze, Marta Romeo, Joel E. Fischer, Gisela Reyes-Cruz
HAI4
2023 Do You Follow?: A Fully Automated System for Adaptive Robot Presenters
abstract
An interesting application for social robots is to act as a presenter, for example as a museum guide. In this paper, we present a fully automated system architecture for building adaptive presentations for embodied agents. The presentation is generated from a knowledge graph, which is also used to track the grounding state of information, based on multimodal feedback from the user. We introduce a novel way to use large-scale language models (GPT-3 in our case) to lexicalise arbitrary knowledge graph triples, greatly simplifying the design of this aspect of the system. We also present an evaluation where 43 participants interacted with the system. The results show that users prefer the adaptive system and consider it more human-like and flexible than a static version of the same system, but only partial results are seen in their learning of the facts presented by the robot.
Agnes Axelsson, Gabriel Skantze
HRI2
2023 I Learn Better Alone!: Collaborative and Individual Word Learning With a Child and Adult Robot
abstract
The use of social robots as a tool for language learning has been studied quite extensively recently. Although their effectiveness and comparison with other technologies are well studied, the effects of the robot's appearance and the interaction setting have received less attention. As educational robots are envisioned to appear in household or school environments, it is important to investigate how their designed persona or interaction dynamics affect learning outcomes. In such environments, children may do the activities together or alone or perform them in the presence of an adult or another child. In this regard, we have identified two novel factors to investigate: the robot's perceived age (adult or child) and the number of learners interacting with the robot simultaneously (one or two). We designed an incidental word learning card game with the Furhat robot and ran a between-subject experiment with 75 middle school participants. We investigated the interactions and effects of children's word learning outcomes, speech activity, and perception of the robot's role. The results show that children who played alone with the robot had better word retention and anthropomorphized the robot more, compared to those who played in pairs. Furthermore, unlike previous findings from human-human interactions, children did not show different behaviors in the presence of a robot designed as an adult or a child. We discuss these factors in detail and make a novel contribution to the direct comparison of collaborative versus individual learning and the new concept of the robot's age.
Alireza Mahmoudi Kamelabad, Gabriel Skantze
HRI2
2023 Show & Tell: Voice Activity Projection and Turn-taking
Erik Ekstedt, Gabriel Skantze
INTERSPEECH2
2023 Automatic Evaluation of Turn-taking Cues in Conversational Speech Synthesis
Erik Ekstedt, Éva Székely, Joakim Gustafson, Gabriel Skantze
INTERSPEECH5
2023 The Open-domain Paradox for Chatbots: Common Ground as the Basis for Human-like Dialogue
abstract
There is a surge in interest in the development of open-domain chatbots, driven by the recent advancements of large language models.The "openness" of the dialogue is expected to be maximized by providing minimal information to the users about the common ground they can expect, including the presumed joint activity.However, evidence suggests that the effect is the opposite.Asking users to "just chat about anything" results in a very narrow form of dialogue, which we refer to as the open-domain paradox.In this position paper, we explain this paradox through the theory of common ground as the basis for human-like communication.Furthermore, we question the assumptions behind open-domain chatbots and identify paths forward for enabling common ground in human-computer dialogue.
Gabriel Skantze, A. Seza Dogruöz
SIGDIAL1
2023 Resolving References in Visually-Grounded Dialogue via Text Generation
abstract
Vision-language models (VLMs) have shown to be effective at image retrieval based on simple text queries, but text-image retrieval based on conversational input remains a challenge.Consequently, if we want to use VLMs for reference resolution in visually-grounded dialogue, the discourse processing capabilities of these models need to be augmented.To address this issue, we propose fine-tuning a causal large language model (LLM) to generate definite descriptions that summarize coreferential information found in the linguistic context of references.We then use a pretrained VLM to identify referents based on the generated descriptions, zero-shot.We evaluate our approach on a manually annotated dataset of visuallygrounded dialogues and achieve results that, on average, exceed the performance of the baselines we compare against.Furthermore, we find that using referent descriptions based on larger context windows has the potential to yield higher returns.
Bram Willemsen, Livia Qian, Gabriel Skantze
SIGDIAL3
2022 CreativeBot: a Creative Storyteller robot to stimulate creativity in children
abstract
We present the design and evaluation of a storytelling activity between children and an autonomous robot aiming at nurturing children’s creativity. We assessed whether a robot displaying creative behavior will positively impact children’s creativity skills in a storytelling context. We developed two models for the robot to engage in the storytelling activity: creative model, where the robot generates creative story ideas, and the non-creative model, where the robot generates non-creative story ideas. We also investigated whether the type of the storytelling interaction will have an impact on children’s creativity skills. We used two types of interaction: 1) Collaborative, where the child and the robot collaborate together by taking turns to tell a story. 2) Non-collaborative: where the robot first tells a story to the child and then asks the child to tell it another story. We conducted a between-subjects study with 103 children in four different conditions: Creative collaborative, Non-creative collaborative, Creative non-collaborative and Non-Creative non-collaborative. The children’s stories were evaluated according to the four standard creativity variables: fluency, flexibility, elaboration and originality. Results emphasized that children who interacted with a creative robot showed higher creativity during the interaction than children who interacted with a non-creative robot. Nevertheless, no significant effect of the type of the interaction was found on children’s creativity skills. Our findings are significant to the Child-Robot interaction (cHRI) community since they enrich the scientific understanding of the development of child-robot encounters for educational applications.
Maha Elgarf, Sahba Zojaji, Gabriel Skantze, Christopher Peters 0001
ICMI3
2022 Voice Activity Projection: Self-supervised Learning of Turn-taking Events
Erik Ekstedt, Gabriel Skantze
INTERSPEECH2
2022 Annotation of Communicative Functions of Short Feedback Tokens in Switchboard
abstract
There has been a lot of work on predicting the timing of feedback in conversational systems. However, there has been less focus on predicting the prosody and lexical form of feedback given their communicative function. Therefore, in this paper we present our preliminary annotations of the communicative functions of 1627 short feedback tokens from the Switchboard corpus and an analysis of their lexical realizations and prosodic characteristics. Since there is no standard scheme for annotating the communicative function of feedback we propose our own annotation scheme. Although our work is ongoing, our preliminary analysis revealed lexical tokens such as “yeah” are ambiguous and therefore lexical forms alone are not indicative of the function. Both the lexical form and prosodic characteristics need to be taken into account in order to predict the communicative function. We also found that feedback functions have distinguishable prosodic characteristics in terms of duration, mean pitch, pitch slope, and pitch range.
Carol Figueroa, Adaeze Adigwe, Magalie Ochs, Gabriel Skantze
LREC4
2022 Collecting Visually-Grounded Dialogue with A Game Of Sorts
abstract
An idealized, though simplistic, view of the referring expression production and grounding process in (situated) dialogue assumes that a speaker must merely appropriately specify their expression so that the target referent may be successfully identified by the addressee. However, referring in conversation is a collaborative process that cannot be aptly characterized as an exchange of minimally-specified referring expressions. Concerns have been raised regarding assumptions made by prior work on visually-grounded dialogue that reveal an oversimplified view of conversation and the referential process. We address these concerns by introducing a collaborative image ranking task, a grounded agreement game we call “A Game Of Sorts”. In our game, players are tasked with reaching agreement on how to rank a set of images given some sorting criterion through a largely unrestricted, role-symmetric dialogue. By putting emphasis on the argumentation in this mixed-initiative interaction, we collect discussions that involve the collaborative referential process. We describe results of a small-scale data collection experiment with the proposed task. All discussed materials, which includes the collected data, the codebase, and a containerized version of the application, are publicly available.
Bram Willemsen, Dmytro Kalpakchi, Gabriel Skantze
LREC3
2022 Knowing Where to Look: A Planning-based Architecture to Automate the Gaze Behavior of Social Robots
abstract
Gaze cues play an important role in human communication and are used to coordinate turn-taking and joint attention, as well as to regulate intimacy. In order to have fluent conversations with people, social robots need to exhibit humanlike gaze behavior. Previous Gaze Control Systems (GCS) in HRI have automated robot gaze using data-driven or heuristic approaches. However, these systems tend to be mainly reactive in nature. Planning the robot gaze ahead of time could help in achieving more realistic gaze behavior and better eye-head coordination. In this paper, we propose and implement a novel planning-based GCS. We evaluate our system in a comparative within-subjects user study (N=26) between a reactive system and our proposed system. The results show that the users preferred the proposed system and that it was significantly more interpretable and better at regulating intimacy.
Chinmaya Mishra, Gabriel Skantze
RO-MAN2
2022 How Much Does Prosody Help Turn-taking? Investigations using Voice Activity Projection Models
abstract
Turn-taking is a fundamental aspect of human communication and can be described as the ability to take turns, project upcoming turn shifts, and supply backchannels at appropriate locations throughout a conversation.In this work, we investigate the role of prosody in turntaking using the recently proposed Voice Activity Projection model, which incrementally models the upcoming speech activity of the interlocutors in a self-supervised manner, without relying on explicit annotation of turn-taking events, or the explicit modeling of prosodic features.Through manipulation of the speech signal, we investigate how these models implicitly utilize prosodic information.We show that these systems learn to utilize various prosodic aspects of speech both on aggregate quantitative metrics of long-form conversations and on single utterances specifically designed to depend on prosody.
Erik Ekstedt, Gabriel Skantze
SIGDIAL2
2022 CoLLIE: Continual Learning of Language Grounding from Language-Image Embeddings
abstract
This paper presents CoLLIE: a simple, yet effective model for continual learning of how language is grounded in vision. Given a pre-trained multimodal embedding model, where language and images are projected in the same semantic space (in this case CLIP by OpenAI), CoLLIE learns a transformation function that adjusts the language embeddings when needed to accommodate new language use. This is done by predicting the difference vector that needs to be applied, as well as a scaling factor for this vector, so that the adjustment is only applied when needed. Unlike traditional few-shot learning, the model does not just learn new classes and labels, but can also generalize to similar language use and leverage semantic compositionality. We verify the model’s performance on two different tasks of identifying the targets of referring expressions, where it has to learn new language use. The results show that the model can efficiently learn and generalize from only a few examples, with little interference with the model’s original zero-shot performance.
Gabriel Skantze, Bram Willemsen
J. Artif. Intell. Res.1
2021 Once Upon a Story: Can a Creative Storyteller Robot Stimulate Creativity in Children?
abstract
Creativity is a vital inherent human trait. In an attempt to stimulate children's creativity, we present the design and evaluation of an interaction between a child and a social robot in a storytelling context. Using a software interface, children were asked to collaboratively create a story with the robot. We conducted a study with 38 children in two conditions. In one condition, the children interacted with a robot exhibiting creative behavior while in the other condition, they interacted with a robot exhibiting non creative behavior. The robot's creativity was defined as verbal and performance creativity. The robot's creative and non creative behaviors were extracted from a previously collected data set and were validated in an online survey with 100 participants. Contrary to our initial hypothesis, children's creativity measures were not higher in the creative condition than in the non creative condition. Our results suggest that merely the robot's creative behavior is insufficient to stimulate creativity in children in a child robot interaction. We further discuss other design factors that may facilitate sparking creativity in children in similar settings in the future.
Maha Elgarf, Gabriel Skantze, Christopher Peters 0001
IVA2
2021 How "open" are the conversations with open-domain chatbots? A proposal for Speech Event based evaluation
abstract
Open-domain chatbots are supposed to converse freely with humans without being restricted to a topic, task or domain.However, the boundaries and/or contents of opendomain conversations are not clear.To clarify the boundaries of "openness", we conduct two studies: First, we classify the types of "speech events" encountered in a chatbot evaluation data set (i.e., Meena by Google) and find that these conversations mainly cover the "small talk" category and exclude the other speech event categories encountered in real life human-human communication.Second, we conduct a small-scale pilot study to generate online conversations covering a wider range of speech event categories between two humans vs. a human and a state-of-the-art chatbot (i.e., Blender by Facebook).A human evaluation of these generated conversations indicates a preference for human-human conversations, since the human-chatbot conversations lack coherence in most speech event categories.Based on these results, we suggest (a) using the term "small talk" instead of "opendomain" for the current chatbots which are not that "open" in terms of conversational abilities yet, and (b) revising the evaluation methods to test the chatbot conversations against other speech events.
A. Seza Dogruöz, Gabriel Skantze
SIGDIAL2
2021 Projection of Turn Completion in Incremental Spoken Dialogue Systems
abstract
The ability to take turns in a fluent way (i.e., without long response delays or frequent interruptions) is a fundamental aspect of any spoken dialog system.However, practical speech recognition services typically induce a long response delay, as it takes time before the processing of the user's utterance is complete.There is a considerable amount of research indicating that humans achieve fast response times by projecting what the interlocutor will say and estimating upcoming turn completions.In this work, we implement this mechanism in an incremental spoken dialog system, by using a language model that generates possible futures to project upcoming completion points.In theory, this could make the system more responsive, while still having access to semantic information not yet processed by the speech recognizer.We conduct a small study which indicates that this is a viable approach for practical dialog systems, and that this is a promising direction for future research.
Erik Ekstedt, Gabriel Skantze
SIGDIAL2
2021 Turn-taking in Conversational Systems and Human-Robot Interaction: A Review
abstract
The taking of turns is a fundamental aspect of dialogue. Since it is difficult to speak and listen at the same time, the participants need to coordinate who is currently speaking and when the next person can start to speak. Humans are very good at this coordination, and typically achieve fluent turn-taking with very small gaps and little overlap. Conversational systems (including voice assistants and social robots), on the other hand, typically have problems with frequent interruptions and long response delays, which has called for a substantial body of research on how to improve turn-taking in conversational systems. In this review article, we provide an overview of this research and give directions for future research. First, we provide a theoretical background of the linguistic research tradition on turn-taking and some of the fundamental concepts in theories of turn-taking. We also provide an extensive review of multi-modal cues (including verbal cues, prosody, breathing, gaze and gestures) that have been found to facilitate the coordination of turn-taking in human-human interaction, and which can be utilised for turn-taking in conversational systems. After this, we review work that has been done on modelling turn-taking, including end-of-turn detection, handling of user interruptions, generation of turn-taking cues, and multi-party human-robot interaction. Finally, we identify key areas where more research is needed to achieve fluent turn-taking in spoken interaction between man and machine.
Gabriel Skantze
Comput. Speech Lang.1
2020 Using knowledge graphs and behaviour trees for feedback-aware presentation agents
abstract
In this paper, we address the problem of how an interactive agent (such as a robot) can present information to an audience and adapt the presentation according to the feedback it receives. We extend a previous behaviour tree-based model to generate the presentation from a knowledge graph (Wikidata), which allows the agent to handle feedback incrementally, and adapt accordingly. Our main contribution is using this knowledge graph not just for generating the system's dialogue, but also as the structure through which short-term user modelling happens. In an experiment using simulated users and third-party observers, we show that referring expressions generated by the system are rated more highly when they adapt to the type of feedback given by the user, and when they are based on previously grounded information as opposed to new information.
Nils Axelsson, Gabriel Skantze
IVA2
2019 The Effects of Embodiment and Social Eye-Gaze in Conversational Agents
Dimosthenis Kontogiorgos, Gabriel Skantze, André Pereira 0001, Joakim Gustafson
CogSci2
2019 Fundamental Frequency Accommodation in Multi-Party Human-Robot Game Interactions: The Effect of Winning or Losing
abstract
In human-human interactions, the situational context plays a large role in the degree of speakers’ accommodation. In this paper, we investigate whether the degree of accommodation in a human-robot computer game is affected by (a) the duration of the interaction and (b) the success of the players in the game. 30 teams of two players played two card games with a conversational robot in which they had to find a correct order of five cards. After game 1, the players received the result of the game on a success scale from 1 (lowest success) to 5 (highest). Speakers’ fo accommodation was measured as the Euclidean distance between the human speakers and each human and the robot. Results revealed that (a) the duration of the game had no influence on the degree of fo accommodation and (b) the result of Game 1 correlated with the degree of fo accommodation in Game 2 (higher success equals lower Euclidean distance). We argue that game success is most likely considered as a sign of the success of players’ cooperation during the discussion, which leads to a higher accommodation behavior in speech.
Omnia Ibrahim, Gabriel Skantze, Sabine Stoll, Volker Dellwo
INTERSPEECH2
2019 Modelling Adaptive Presentations in Human-Robot Interaction using Behaviour Trees
abstract
In dialogue, speakers continuously adapt their speech to accommodate the listener, based on the feedback they receive.In this paper, we explore the modelling of such behaviours in the context of a robot presenting a painting.A Behaviour Tree is used to organise the behaviour on different levels, and allow the robot to adapt its behaviour in real-time; the tree organises engagement, joint attention, turn-taking, feedback and incremental speech processing.An initial implementation of the model is presented, and the system is evaluated in a user study, where the adaptive robot presenter is compared to a non-adaptive version.The adaptive version is found to be more engaging by the users, although no effects are found on the retention of the presented material.
Nils Axelsson, Gabriel Skantze
SIGdial2
2018 Using lexical alignment and referring ability to address data sparsity in situated dialog reference resolution
abstract
Referring to entities in situated dialog is a collaborative process, whereby interlocutors often expand, repair and/or replace referring expressions in an iterative process, converging on conceptual pacts of referring language use in doing so.Nevertheless, much work on exophoric reference resolution (i.e.resolution of references to entities outside of a given text) follows a literary model, whereby individual referring expressions are interpreted as unique identifiers of their referents given the state of the dialog the referring expression is initiated.In this paper, we address this collaborative nature to improve dialogic reference resolution in two ways: First, we trained a words-asclassifiers logistic regression model of word semantics and incrementally adapt the model to idiosyncratic language between dyad partners during evaluation of the dialog.We then used these semantic models to learn the general referring ability of each word, which is independent of referent features.These methods facilitate accurate automatic reference resolution in situated dialog without annotation of referring expressions, even with little background data.
Todd Shore, Gabriel Skantze
EMNLP2
2018 Multimodal Continuous Turn-Taking Prediction Using Multiscale RNNs
abstract
In human conversational interactions, turn-taking exchanges can be coordinated using cues from multiple modalities. To design spoken dialog systems that can conduct fluid interactions it is desirable to incorporate cues from separate modalities into turn-taking models. We propose that there is an appropriate temporal granularity at which modalities should be modeled. We design a multiscale RNN architecture to model modalities at separate timescales in a continuous manner. Our results show that modeling linguistic and acoustic features at separate temporal rates can be beneficial for turn-taking modeling. We also show that our approach can be used to incorporate gaze features into turn-taking models.
Matthew Roddy, Gabriel Skantze, Naomi Harte
ICMI2
2018 Investigating Speech Features for Continuous Turn-Taking Prediction Using LSTMs
abstract
For spoken dialog systems to conduct fluid conversational interactions with users, the systems must be sensitive to turn-taking cues produced by a user. Models should be designed so that effective decisions can be made as to when it is appropriate, or not, for the system to speak. Traditional end-of-turn models, where decisions are made at utterance end-points, are limited in their ability to model fast turn-switches and overlap. A more flexible approach is to model turn-taking in a continuous manner using RNNs, where the system predicts speech probability scores for discrete frames within a future window. The continuous predictions represent generalized turn-taking behaviors observed in the training data and can be applied to make decisions that are not just limited to end-of-turn detection. In this paper, we investigate optimal speech-related feature sets for making predictions at pauses and overlaps in conversation. We find that while traditional acoustic features perform well, part-of-speech features generally perform worse than word features. We show that our current models outperform previously reported baselines.
Matthew Roddy, Gabriel Skantze, Naomi Harte
INTERSPEECH2
2018 Effects of Posture and Embodiment on Social Distance in Human-Agent Interaction in Mixed Reality
abstract
Mixed reality offers new potentials for social interaction experiences with virtual agents. In addition, it can be used to experiment with the design of physical robots. However, while previous studies have investigated comfortable social distances between humans and artificial agents in real and virtual environments, there is little data with regards to mixed reality environments. In this paper, we conducted an experiment in which participants were asked to walk up to an agent to ask a question, in order to investigate the social distances maintained, as well as the subject's experience of the interaction. We manipulated both the embodiment of the agent (robot vs. human and virtual vs. physical) as well as closed vs. open posture of the agent. The virtual agent was displayed using a mixed reality headset. Our experiment involved 35 participants in a within-subject design. We show that, in the context of social interactions, mixed reality fares well against physical environments, and robots fare well against humans, barring a few technical challenges.
Theofronia Androulakaki, Alex Yuan Gao, Fangkai Yang, Himangshu Saikia, Christopher Peters 0001, Gabriel Skantze
IVA7
2018 A Multimodal Corpus for Mutual Gaze and Joint Attention in Multiparty Situated Interaction
Dimosthenis Kontogiorgos, Vanya Avramova, Simon Alexanderson, Patrik Jonell, Catharine Oertel, Jonas Beskow, Gabriel Skantze, Joakim Gustafson
LREC7
2018 KTH Tangrams: A Dataset for Research on Alignment and Conceptual Pacts in Task-Oriented Dialogue
Todd Shore, Theofronia Androulakaki, Gabriel Skantze
LREC3
2017 Predicting and Regulating Participation Equality in Human-robot Conversations: Effects of Age and Gender
abstract
In this paper, we investigate participation equality, in terms of speaking time, between users in multi-party human-robot conversations. We analyse a dataset where pairs of users (540 in total) interact with a conversational robot exhibited at a technical museum. The data encompass a wide range of different users in terms of age (adults/children) and gender (male/female), in different combinations. Overall, the analysis indicates that demographically heterogeneous pairs are more imbalanced, especially pairs of adults and children, where children are less prone to self-select in the turn-taking. The analysis also indicates that it is possible for the robot to reduce the imbalance by addressing the least dominant user and asking directed questions. However, for children to respond, it is important to seek mutual gaze and switch addressee often. Finally, we show that it is possible to predict the imbalance at an early stage in the interaction -- in order to increase the participation equality as early as possible -- and that knowledge about the users' age and gender helps in this prediction.
Gabriel Skantze
HRI1
2017 A Virtual Poster Presenter Using Mixed Reality
Vanya Avramova, Fangkai Yang, Christopher Peters 0001, Gabriel Skantze
IVA5
2017 A Psychotherapy Training Environment with Virtual Patients Implemented Using the Furhat Robot Platform
Robert Johansson, Gabriel Skantze, Arne Jönsson
IVA2
2017 Towards a General, Continuous Model of Turn-taking in Spoken Dialogue using LSTM Recurrent Neural Networks
abstract
Previous models of turn-taking have mostly been trained for specific turn-taking decisions, such as discriminating between turn shifts and turn retention in pauses.In this paper, we present a predictive, continuous model of turntaking using Long Short-Term Memory (LSTM) Recurrent Neural Networks (RNN).The model is trained on human-human dialogue data to predict upcoming speech activity in a future time window.We show how this general model can be applied to two different tasks that it was not specifically trained for.First, to predict whether a turn-shift will occur or not in pauses, where the model achieves a better performance than human observers, and better than results achieved with more traditional models.Second, to make a prediction at speech onset whether the utterance will be a short backchannel or a longer utterance.Finally, we show how the hidden layer in the network can be used as a feature vector for turntaking decisions in a human-robot interaction scenario.
Gabriel Skantze
SIGDIAL Conference1
2016 Root Cause Analysis of Miscommunication Hotspots in Spoken Dialogue Systems
abstract
A major challenge in Spoken Dialogue Systems (SDS) is the detection of problematic communication (hotspots), as well as the classification of these hotspots into different types (root cause analysi ...
Spiros Georgiladakis, Georgia Athanasopoulou, Raveesh Meena, José Lopes 0001, Arodami Chorianopoulou, Elisavet Palogiannidi, Elias Iosif, Gabriel Skantze, Alexandros Potamianos
INTERSPEECH8
2015 Exploring Turn-taking Cues in Multi-party Human-robot Discussions about Objects
abstract
In this paper, we present a dialog system that was exhibited at the Swedish National Museum of Science and Technology. Two visitors at a time could play a collaborative card sorting game together with the robot head Furhat, where the three players discuss the solution together. The cards are shown on a touch table between the players, thus constituting a target for joint attention. We describe how the system was implemented in order to manage turn-taking and attention to users and objects in the shared physical space. We also discuss how multi-modal redundancy (from speech, card movements and head pose) is exploited to maintain meaningful discussions, given that the system has to process conversational speech from both children and adults in a noisy environment. Finally, we present an analysis of 373 interactions, where we investigate the robustness of the system, to what extent the system's attention can shape the users' turn-taking behaviour, and how the system can produce multi-modal turn-taking signals (filled pauses, facial gestures, breath and gaze) to deal with processing delays in the system.
Gabriel Skantze, Martin Johansson, Jonas Beskow
ICMI1
2015 Detecting repetitions in spoken dialogue systems using phonetic distances
abstract
This paper addresses the problem of automatic detection of re-peated turns in Spoken Dialogue Systems. Repetitions can be a symptom of problematic communication between users and systems. Such repetitions are often due to speech recognition errors, which in turn makes it hard to use speech recognition to detect repetitions. We present an approach to detect rep-etition using the phonetic distance to find the best alignment between turns in the same dialogue. The alignment score ob-tained is combined with different features to improve repeti-tion detection. To evaluate the method proposed we compare several alignment techniques from edit distance to DTW-based distance, previously used in Spoken-Term detection tasks. We also compare two different methods to compute the phonetic distance: the first one using the phoneme sequence, and the second one using the distance between the phone posterior vec-tors. Two different datasets were used in this evaluation: a bus-schedule information system (in English) and a call routing system (in Swedish). The results show that approaches using phoneme distances over-perform approaches using Levenshtein distances between ASR outputs for repetition detection. Index Terms: spoken dialogue systems, repetition detection, phonetic distance
José Lopes 0001, Giampiero Salvi, Gabriel Skantze, Alberto Abad, Joakim Gustafson, Fernando Batista, Raveesh Meena, Isabel Trancoso
INTERSPEECH3
2015 A Collaborative Human-Robot Game as a Test-bed for Modelling Multi-party, Situated Interaction
Gabriel Skantze, Martin Johansson, Jonas Beskow
IVA1
2015 Opportunities and Obligations to Take Turns in Collaborative Multi-Party Human-Robot Interaction
abstract
In this paper we present a data-driven model for detecting opportunities and obligations for a robot to take turns in multi-party discussions about objects.The data used for the model was collected in a public setting, where the robot head Furhat played a collaborative card sorting game together with two users.The model makes a combined detection of addressee and turn-yielding cues, using multi-modal data from voice activity, syntax, prosody, head pose, movement of cards, and dialogue context.The best result for a binary decision is achieved when several modalities are combined, giving a weighted F 1 score of 0.876 on data from a previously unseen interaction, using only automatically extractable features.
Martin Johansson, Gabriel Skantze
SIGDIAL Conference2
2015 Automatic Detection of Miscommunication in Spoken Dialogue Systems
abstract
In this paper, we present a data-driven approach for detecting instances of miscommunication in dialogue system interactions.A range of generic features that are both automatically extractable and manually annotated were used to train two models for online detection and one for offline analysis.Online detection could be used to raise the error awareness of the system, whereas offline detection could be used by a system designer to identify potential flaws in the dialogue design.In experimental evaluations on system logs from three different dialogue systems that vary in their dialogue strategy, the proposed models performed substantially better than the majority class baseline models.
Raveesh Meena, José Lopes 0001, Gabriel Skantze, Joakim Gustafson
SIGDIAL Conference3
2015 Modelling situated human-robot interaction using IrisTK
abstract
In this demonstration we show how situated multi-party human-robot interaction can be modelled using the open source framework IrisTK.We will demonstrate the capabilities of IrisTK by showing an application where two users are playing a collaborative card sorting game together with the robot head Furhat, where the cards are shown on a touch table between the players.The application is interesting from a research perspective, as it involves both multi-party interaction, as well as joint attention to the objects under discussion.
Gabriel Skantze, Martin Johansson
SIGDIAL Conference1
2015 Introduction for Speech and language for interactive robots
Heriberto Cuayáhuitl, Kazunori Komatani, Gabriel Skantze
Comput. Speech Lang.3
2014 Human-robot collaborative tutoring using multiparty multimodal spoken dialogue
abstract
In this paper, we describe a project that explores a novel experimental setup towards building a spoken, multi-modally rich, and human-like multiparty tutoring robot. A human-robot interaction setup is designed, and a human-human dialogue corpus is collected. The corpus targets the development of a dialogue system platform to study verbal and nonverbal tutoring strategies in multiparty spoken interactions with robots which are capable of spoken dialogue. The dialogue task is centered on two participants involved in a dialogue aiming to solve a card-ordering game. Along with the participants sits a tutor (robot) that helps the participants perform the task, and organizes and balances their interaction. Different multimodal signals captured and auto-synchronized by different audio-visual capture technologies, such as a microphone array, Kinects, and video cameras, were coupled with manual annotations. These are used build a situated model of the interaction based on the participants personalities, their state of attention, their conversational engagement and verbal dominance, and how that is correlated with the verbal and visual feed-back, turn-management, and conversation regulatory actions generated by the tutor. Driven by the analysis of the corpus, we will show also the detailed design methodologies for an affective, and multimodally rich dialogue system that allows the robot to measure incrementally the attention states, and the dominance for each participant, allowing the robot head Furhat to maintain a well-coordinated, balanced, and engaging conversation, that attempts to maximize the agreement and the contribution to solve the task.
Samer Al Moubayed, Jonas Beskow, Bajibabu Bollepalli, Joakim Gustafson, Ahmed Hussen Abdelaziz, Martin Johansson, Maria Koutsombogera, José Lopes 0001, Jekaterina Novikova, Catharine Oertel, Gabriel Skantze, Kalin Stefanov, Gül Varol
HRI11
2014 Spontaneous spoken dialogues with the furhat human-like robot head
abstract
Furhat [1] is a robot head that deploys a back-projected animated face that is realistic and human-like in anatomy. Furhat relies on a state-of-the-art facial animation architecture allowing accurate synchronized lip movements with speech, and the control and generation of non-verbal gestures, eye movements and facial expressions.
Samer Al Moubayed, Jonas Beskow, Gabriel Skantze
HRI3
2014 UM3I 2014: International Workshop on Understanding and Modeling Multiparty, Multimodal Interactions
abstract
In this paper, we present a brief summary of the international workshop on Modeling Multiparty, Multimodal Interactions. The UM3I 2014 workshop is held in conjunction with the ICMI 2014 conference. The workshop will highlight recent developments and adopted methodologies in the analysis and modeling of multiparty and multimodal interactions, the design and implementation principles of related human-machine interfaces, as well as the identification of potential limitations and ways of overcoming them.
Samer Al Moubayed, Dan Bohus, Anna Esposito, Dirk Heylen, Maria Koutsombogera, Harris Papageorgiou, Gabriel Skantze
ICMI7
2014 Crowdsourcing Street-level Geographic Information Using a Spoken Dialogue System
abstract
We present a technique for crowdsourcing street-level geographic information using spoken natural language.In particular, we are interested in obtaining first-person-view information about what can be seen from different positions in the city.This information can then for example be used for pedestrian routing services.The approach has been tested in the lab using a fully implemented spoken dialogue system, and has shown promising results.
Raveesh Meena, Johan Boye, Gabriel Skantze, Joakim Gustafson
SIGDIAL Conference3
2014 Data-driven models for timing feedback responses in a Map Task dialogue system
Raveesh Meena, Gabriel Skantze, Joakim Gustafson
Comput. Speech Lang.2
2014 Turn-taking, feedback and joint attention in situated human-robot interaction
Gabriel Skantze, Anna Hjalmarsson, Catharine Oertel
Speech Commun.1
2013 The furhat social companion talking head
Samer Al Moubayed, Jonas Beskow, Gabriel Skantze
INTERSPEECH3
2013 User feedback in human-robot interaction: prosody, gaze and timing
abstract
This paper investigates forms and functions of user feedback in a map task dialogue between a human and a robot, where the robot is the instruction-giver and the human is the instruction-follower. First, we investigate how user acknowledgements in task-oriented dialogue signal whether an activity is about to be initiated or has been completed. The parameters analysed include the users ’ lexical and prosodic realisation as well as gaze direction and response timing. Second, we investigate the relation between these parameters and the perception of uncertainty.
Gabriel Skantze, Catharine Oertel, Anna Hjalmarsson
INTERSPEECH1
2013 The Map Task Dialogue System: A Test-bed for Modelling Human-Like Dialogue
Raveesh Meena, Gabriel Skantze, Joakim Gustafson
SIGDIAL Conference2
2013 A Data-driven Model for Timing Feedback in a Map Task Dialogue System
Raveesh Meena, Gabriel Skantze, Joakim Gustafson
SIGDIAL Conference2
2013 Exploring the effects of gaze and pauses in situated human-robot interaction
Gabriel Skantze, Anna Hjalmarsson, Catharine Oertel
SIGDIAL Conference1
2013 Towards incremental speech generation in conversational systems
Gabriel Skantze, Anna Hjalmarsson
Comput. Speech Lang.1
2012 Multimodal multiparty social interaction with the furhat head
abstract
We will show in this demonstrator an advanced multimodal and multiparty spoken conversational system using Furhat, a robot head based on projected facial animation. Furhat is a human-like interface that utilizes facial animation for physical robot heads using back-projection. In the system, multimodality is enabled using speech and rich visual input signals such as multi-person real-time face tracking and microphone tracking. The demonstrator will showcase a system that is able to carry out social dialogue with multiple interlocutors simultaneously with rich output signals such as eye and head coordination, lips synchronized speech synthesis, and non-verbal facial gestures used to regulate fluent and expressive multiparty conversations.
Samer Al Moubayed, Gabriel Skantze, Jonas Beskow, Kalin Stefanov, Joakim Gustafson
ICMI2
2012 IrisTK: a statechart-based toolkit for multi-party face-to-face interaction
abstract
In this paper, we present IrisTK - a toolkit for rapid development of real-time systems for multi-party face-to-face interaction. The toolkit consists of a message passing system, a set of modules for multi-modal input and output, and a dialog authoring language based on the notion of statecharts. The toolkit has been applied to a large scale study in a public museum setting, where the back-projected robot head Furhat interacted with the visitors in multi-party dialog.
Gabriel Skantze, Samer Al Moubayed
ICMI1
2012 A Data-driven Approach to Understanding Spoken Route Directions in Human-Robot Dialogue
abstract
In this paper, we present a data-driven chunking parser for automatic interpretation of spoken route directions into a route graph that is useful for robot navigation. Different sets of features and machine learning algorithms are explored. The results indicate that our approach is robust to speech recognition errors. Index Terms: spoken language understanding, route directions, human-robot interaction
Raveesh Meena, Gabriel Skantze, Joakim Gustafson
INTERSPEECH2
2012 Lip-Reading: Furhat Audio Visual Intelligibility of a Back Projected Animated Face
Samer Al Moubayed, Gabriel Skantze, Jonas Beskow
IVA2
2011 Enhanced visual scene understanding through human-robot dialog
abstract
We propose a novel human-robot-interaction framework for robust visual scene understanding. Without any a-priori knowledge about the objects, the task of the robot is to correctly enumerate how many of them are in the scene and segment them from the background. Our approach builds on top of state-of-the-art computer vision methods, generating object hypotheses through segmentation. This process is combined with a natural dialog system, thus including a `human in the loop' where, by exploiting the natural conversation of an advanced dialog system, the robot gains knowledge about ambiguous situations. We present an entropy-based system allowing the robot to detect the poorest object hypotheses and query the user for arbitration. Based on the information obtained from the human-robot dialog, the scene segmentation can be re-seeded and thereby improved. We present experimental results on real data that show an improved segmentation performance compared to segmentation without interaction.
Matthew Johnson-Roberson, Jeannette Bohg, Gabriel Skantze, Joakim Gustafson, Rolf Carlson, Babak Rasolzadeh, Danica Kragic
IROS3
2010 Middleware for Incremental Processing in Conversational Agents
David Schlangen, Timo Baumann, Hendrik Buschmeier, Okko Buß, Stefan Kopp, Gabriel Skantze, Ramin Yaghoubzadeh
SIGDIAL Conference6
2010 Towards Incremental Speech Generation in Dialogue Systems
Gabriel Skantze, Anna Hjalmarsson
SIGDIAL Conference1
2009 A General, Abstract Model of Incremental Dialogue Processing
David Schlangen, Gabriel Skantze
EACL2
2009 Incremental Dialogue Processing in a Micro-Domain
Gabriel Skantze, David Schlangen
EACL1
2009 The MonAMI reminder: a spoken dialogue system for face-to-face interaction
abstract
We describe the MonAMI Reminder, a multimodal spoken dialogue system which can assist elderly and disabled people in organising and initiating their daily activities. Based on deep interviews with potential users, we have designed a calendar and reminder application which uses an innovative mix of an embodied conversational agent, digital pen and paper, and the web to meet the needs of those users as well as the current constraints of speech technology. We also explore the use of head pose tracking for interaction and attention control in human-computer face-to-face interaction.
Jonas Beskow, Jens Edlund, Björn Granström, Joakim Gustafson, Gabriel Skantze, Helena Tobiasson
INTERSPEECH5
2009 Attention and Interaction Control in a Human-Human-Computer Dialogue Setting
Gabriel Skantze, Joakim Gustafson
SIGDIAL Conference1
2008 Innovative interfaces in MonAMI: the reminder
abstract
This demo paper presents an early version of the Reminder, a prototype ECA developed in the European project MonAMI, which aims at "mainstreaming accessibility in consumer goods and services, using advanced technologies to ensure equal access, independent living and participation for all". The Reminder helps users to plan activities and to remember what to do. The prototype merges mobile ECA technology with other, existing technologies: Google Calendar and a digital pen and paper. The solution allows users to continue using a paper calendar in the manner they are used to, whilst the ECA provides notifications on what has been written in the calendar. Users may ask questions such as "When was I supposed to meet Sara?" or "What's my schedule today?"
Jonas Beskow, Jens Edlund, Teodore Gjermani, Björn Granström, Joakim Gustafson, Oskar Jonsson, Gabriel Skantze, Helena Tobiasson
ICMI7
2006 User responses to prosodic variation in fragmentary grounding utterances in dialog
abstract
In a previous study we demonstrated that subjects could use prosodic features (primarily peak height and alignment) to make different interpretations of synthesized fragmentary grounding utterances. In the present study we test the hypothesis that subjects also change their behavior accordingly in a human-computer dialog setting. We report on an experiment in which subjects participate in a color-naming task in a Wizard-of-Oz controlled human-computer dialog in Swedish. The results show that two annotators were able to categorize the subjects ’ responses based on pragmatic meaning. Moreover, the subjects ’ response times differed significantly, depending on the prosodic features of the grounding fragment spoken by the system. Index terms: dialog systems, prosody, error handling 1.
Gabriel Skantze, David House, Jens Edlund
INTERSPEECH1
2005 The effects of prosodic features on the interpretation of clarification ellipses
abstract
In this paper, the effects of prosodic features on the interpretation of elliptical clarification requests in dialogue are studied. An experiment is presented where subjects were asked to listen to ...
Jens Edlund, David House, Gabriel Skantze
INTERSPEECH3
2005 Exploring human error recovery strategies: Implications for spoken dialogue systems
Gabriel Skantze
Speech Commun.1
2004 Higgins - a spoken dialogue system for investigating error handling techniques
abstract
In this paper, an overview of the Higgins project and the research within the project is presented. The project incorporates studies of error handling for spoken dialogue systems on several levels, from processing to dialogue level. A domain in which a range of different error types can be studied has been chosen: pedestrian navigation and guiding. Several data collections within Higgins have been analysed along with data from Higgins' predecessor, the AdApt system. The error handling research issues in the project are presented in light of these analyses.
Jens Edlund, Gabriel Skantze, Rolf Carlson
INTERSPEECH2
2002 Coordination of referring expressions in multimodal human-computer dialogue
abstract
This study examines coordination of referring expressions in multimodal human-computer dialogue, i.e. to what extent users' choices of referring expressions are affected by the referring expressions that the system is designed to use. An experiment was conducted, using a semi-automatic multimodal dialogue system for apartment seeking. The user and the system could refer to areas and apartments on an interactive map by means of speech and pointing gestures. Results indicate that the referring expressions of the system have great influence on the user's choice of referring expressions, both in terms of modality and linguistic content. From this follows a number of implications for the design of multimodal dialogue systems.
Gabriel Skantze
INTERSPEECH1