Catharine Oertel

dblp:24/10573 · DBLP profile ↗
← Back
42ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0002-8273-0132ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 27 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 23 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Reflecti-Mate: A Conversational Agent for Adaptive Decision-Making Support Through System 1 and System 2 Thinking
abstract
Making high-stakes personal decisions involves cognitive, emotional, and intuitive processes, and individuals differ in how they allocate attention across these modes. Integration of these processes has shown to benefit decision making. Yet, most current decision-support systems focus primarily on supporting cognitive aspects, rather than adapting to the individual’s thinking profile to support integration of different types of thoughts. In this study, we investigate an agent designed to encourage integration by adapting to the individual user’s thought patterns. We explore its effects on participants’ perceptions of the agent and their reflective behavior, in comparison with unaided pre-reflection and a baseline agent. In a between-subjects study (N = 128), our agent, which fostered broad and elaborated thinking, enabled more personalized reflective trajectories, elicited more integrative reflective language, and was perceived as providing stronger support for holistic reflection. In contrast, the baseline agent produced homogenized profiles dominated by cognitive language across participants.
Morita Tarvirdians, Senthil Chandrasegaran, Hayley Hung, Catholijn M. Jonker, Catharine Oertel
UMAP5
2026 From human teams to hybrid intelligence teams: identifying, characterizing, and evaluating foundational quality attributes
abstract
Hybrid Intelligence (HI) is an emerging paradigm in which artificial intelligence (AI) augments human intelligence. The current literature lacks systematic models that guide the design and evaluation of HI systems. Further, discussions around HI primarily focus on technology, neglecting the holistic human-AI ensemble. In this paper, we take the initial steps toward the development of a quality model for characterizing and evaluating HI systems from a human-AI teams perspective. We first conducted a study investigating the adequacy of properties commonly associated with effective human teams to describe HI. The study features the insights of 50 HI researchers, and shows that various human team properties, including boundedness, interdependence, competency, purposefulness, initiative, normativity, and effectiveness, are important for HI systems. Based on these results, we developed a quality model for HI teams composed of seven high-level quality attributes, further refined into 16 specific ones. To evaluate the relevance and understanding of the proposed attributes, we conducted a second empirical investigation by staging competitions in which participants used the quality model to develop and analyze HI usage scenarios. Our analysis of 48 collected scenarios, which we openly release, confirms the proposed attributes' relevance and highlights insights that emerge when designers consider the quality model in HI system design.
Davide Dell'Anna, Pradeep K. Murukannaiah, Mireia Yurrita, Bernd Dudzik, Davide Grossi, Catholijn M. Jonker, Catharine Oertel, Pinar Yolum
Auton. Agents Multi Agent Syst.7
2026 A music recommendation system for constructed music-evoked episodic memories (CoMEEMs)
abstract
Music is widely used in human–computer interaction (HCI) to enhance engagement, sustain attention, and support cognitive stimulation. Yet its potential for deliberate mood regulation, particularly through personalized memory recall, remains largely unexplored. Music-evoked autobiographical memories (MEAMs) are often elicited by well-known, favorite songs, yielding stronger mood effects than music without personal memory associations. However, songs can also trigger distressing memories, and will never capture all positive personal memories. Since happy personal memories can enhance mood, broader methods for retrieval are needed. To address this, we introduce Constructed Music-Evoked Episodic Memories (CoMEEMs), a framework linking chosen episodic memories to music. By creating a personalized song-memory database, CoMEEMs enable autonomous mood regulation and communication in interactive systems, integrating memory cues—such as people and places—alongside mood congruence, to help choose songs with high mood regulatory impact. In an experiment with 71 Dutch and French adults, participants described 87 positive memories and received song recommendations based on associated people and places, with and without mood matching. Results showed that song familiarity and genre were the strongest predictors of perceived fit, while valence, arousal, tempo, and lyrics played smaller roles. Mood congruence, especially in valence, significantly influenced song relevance. Participants emphasized the need for user input on emotional states and memory context. Based on these findings, we propose design guidelines to improve future music recommendation systems targeting memories.
Paul Raingeard de la Bletiere, Mark A. Neerincx, Rebecca Schaefer 0001, Catharine Oertel
Int. J. Hum. Comput. Stud.4
2026 Dynamics of Collective Group Affect: Group-Level Annotations and the Multimodal Modeling of Convergence and Divergence
abstract
Collaborating in a purposive group, whether face-to-face or virtually, involves continuously expressing emotions and interpreting those of other group members. As such, understanding group affect is essential to comprehending how groups interact and succeed in collaborative efforts. In this study, we move beyond individual-level affect and investigate group-level affect—a collective phenomenon that reflects the shared mood or emotions among group members at a particular moment. As the first in the literature, we gather annotations for group-level affective expressions in purposive group interactions using a fine-grained temporal approach (15 second windows) that also captures the inherent dynamics of this collective construct. To this end, we extensively train annotators and develop an annotation procedure specifically tuned to capture the entire scope of the group interaction from one interaction moment to the next. In addition, we model the ebb and flow of group affect by accounting for the underlying convergence (driven by emotional contagion) and divergence (resulting from emotional reactivity) of affective expressions among group members. To capture these interpersonal dynamics, we employ two approaches: (i) extracting synchrony-based handcrafted features from both audio and visual modalities, and (ii) introducing a novel, data-driven graph neural network to model interpersonal dynamics among group members. Our results highlight the advantages of the graph network over the handcrafted features in modeling group affect, while also emphasizing the importance of temporal modeling and incorporating multimodal cues. Additionally, our analysis of affective convergence and divergence reveals that groups tend to diverge in their social signals during neutral collective affect, while exhibiting convergence during more emotionally intense moments. These insights are drawn from comparative results across both modeling techniques.
Navin Raj Prabhu, Maria Tsfasman, Catharine Oertel, Timo Gerkmann, Nale Lehmann-Willenbrock
IEEE Trans. Affect. Comput.3
2025 Knowing Me, Knowing AU: How Should We Design Agent-Mediated Mimicry?
Agnes Johanna Axelsson, Weilun Chen, Deborah van Sinttruije, Iulia Lefter, Laurens Rook, Catholijn M. Jonker, Catharine Oertel
Conference on Designing Interactive Systems7
2025 JELAI: Integrating AI and Learning Analytics in Jupyter Notebooks
Manuel Valle Torre, Thom van der Velden, Marcus Specht, Catharine Oertel
AIED (6)4
2025 What Can You Say to a Robot? Capability Communication Leads to More Natural Conversations
abstract
When encountering a robot in the wild, it is not inherently clear to human users what the robot's capabilities are. When encountering misunderstandings or problems in spoken interaction, robots often just apologize and move on, without additional effort to make sure the user understands what happened. We set out to compare the effect of two speech based capability communication strategies (proactive, reactive) to a robot without such a strategy, in regard to the user's rating of and their behavior during the interaction. For this, we conducted an in-person user study with 120 participants who had three speech-based interactions with a social robot in a restaurant setting. Our results suggest that users preferred the robot communicating its capabilities proactively and adjusted their behavior in those interactions, using a more conversational interaction style while also enjoying the interaction more.
Merle M. Reimann, Koen V. Hindriks, Florian Kunneman, Catharine Oertel, Gabriel Skantze, Iolanda Leite
HRI4
2024 Memory with Meaning: Enabling Value-Centric Long-Term Human-Agent Dialogue
abstract
When a human makes a decision, an observer may want to understand the reasons and motivations behind the decision. This understanding is important when IVAs are involved in contextual decision-making or coaching practices. To address this challenge, we propose that an agent’s understanding of its user should include knowledge of the user’s underlying values. Humans prioritise different values – sometimes contradictory – in a manner that depends on the context. We present a method where the agent and user build the required context-sensitive value model together. We use Schwartz’s value theory, which places individuals’ values into ten categories. In a between-subject experiment, with three sessions on different days, we elicit user values by presenting them with moral dilemmas in different contexts on the first day, refine the model by asking users to argue about contradictions on the second day, and let them reflect on the model that they have built together with the system on the third day. We find that users exposed to a value-aware condition are more likely to agree with the robot’s representations of their values post-reflection than those in a baseline. Participants also prioritise different values depending on the context, agreeing with previous findings.
Tom Saveur, Agnes Johanna Axelsson, Franziska Burger, Mark A. Neerincx, Catharine Oertel
IVA5
2024 The Sequence Matters in Learning - A Systematic Literature Review
abstract
Describing and analysing learner behaviour using sequential data and analysis is becoming more and more popular in Learning Analytics. Nevertheless, we found a variety of definitions of learning sequences, as well as choices regarding data aggregation and the methods implemented for analysis. Furthermore, sequences are used to study different educational settings and serve as a base for various interventions. In this literature review, the authors aim to generate an overview of these aspects to describe the current state of using sequence analysis in educational support and learning analytics. The 74 included articles were selected based on the criteria that they conduct empirical research on an educational environment using sequences of learning actions as the main focus of their analysis. The results enable us to highlight different learning tasks where sequences are analysed, identify data mapping strategies for different types of sequence actions, differentiate techniques based on purpose and scope, and identify educational interventions based on the outcomes of sequence analysis.
Manuel Valle Torre, Catharine Oertel, Marcus Specht
LAK2
2024 Impact of Annotation Modality on Label Quality and Model Performance in the Automatic Assessment of Laughter In-the-Wild
abstract
Although laughter is known to be a multimodal signal, it is primarily annotated from audio. It is unclear how laughter labels may differ when annotated from modalities like video, which capture body movements and are relevant in in-the-wild studies. In this work we ask whether annotations of laughter are congruent across modalities, and compare the effect that labeling modality has on machine learning model performance. We compare annotations and models for laughter detection, intensity estimation, and segmentation, using a challenging in-the-wild conversational dataset with a variety of camera angles, noise conditions and voices. Our study with 48 annotators revealed evidence for incongruity in the perception of laughter and its intensity between modalities, mainly due to lower recall in the video condition. Our machine learning experiments compared the performance of modern unimodal and multi-modal models for different combinations of input modalities, training, and testing label modalities. In addition to the same input modalities rated by annotators (audio and video), we trained models with body acceleration inputs, robust to cross-contamination, occlusion and perspective differences. Our results show that performance of models with body movement inputs does not suffer when trained with video-acquired labels, despite their lower inter-rater agreement.
José Vargas Quiros, Laura Cabrera Quiros, Catharine Oertel, Hayley Hung
IEEE Trans. Affect. Comput.3
2024 A Survey on Dialogue Management in Human-robot Interaction
abstract
As social robots see increasing deployment within the general public, improving the interaction with those robots is essential. Spoken language offers an intuitive interface for the human–robot interaction (HRI), with dialogue management (DM) being a key component in those interactive systems. Yet, to overcome current challenges and manage smooth, informative, and engaging interaction, a more structural approach to combining HRI and DM is needed. In this systematic review, we analyze the current use of DM in HRI and focus on the type of dialogue manager used, its capabilities, evaluation methods, and the challenges specific to DM in HRI. We identify the challenges and current scientific frontier related to the DM approach, interaction domain, robot appearance, physical situatedness, and multimodality.
Merle M. Reimann, Florian Kunneman, Catharine Oertel, Koen V. Hindriks
ACM Trans. Hum. Robot Interact.3
2023 Predicting Interaction Quality Aspects Using Level-Based Scores for Conversational Agents
abstract
In order to improve human-agent interaction, it is essential to have good measures of interaction quality. We define interaction quality based on multiple aspects, including usability, likability and perceived conversation quality as subjective measures, and interaction length, completion rate and frequency of unrecognized utterances as objective measures. Determining necessary improvements to a conversational agent is a non-trivial task, because it is difficult to infer from an evaluation of the agent as a whole, which aspects of the agent need to be improved to raise the interaction quality. In this paper, we propose a scoring system for task-oriented conversational agents to predict aspects of interaction quality and to guide an iterative improvement process. Our scoring system does not provide a single score, but leverages structural features of the dialogue management approach and assigns a score on three levels: the utterance, dialogue move, and genre level. Using the agent's scores on separate levels to predict the interaction quality allows making targeted improvements to the conversational agent. In order to evaluate our scoring system, we apply it over the course of multiple crowdsourcing pilot studies, using a recipe recommendation agent. We evaluate the obtained scores in regard to their ability to predict selected objective and subjective interaction quality aspects, as well as their suitability for making informed decisions about necessary improvements.
Merle M. Reimann, Catharine Oertel, Florian Kunneman, Koen V. Hindriks
IVA2
2022 Towards creating a conversational memory for long-term meeting support: predicting memorable moments in multi-party conversations through eye-gaze
abstract
When working in a group, it is essential to understand each other’s viewpoints to increase group cohesion and meeting productivity. This can be challenging in teams: participants might be left misunderstood and the discussion could be going around in circles. To tackle this problem, previous research on group interactions has addressed topics such as dominance detection, group engagement, and group creativity. Conversational memory, however, remains a widely unexplored area in the field of multimodal analysis of group interaction. The ability to track what each participant or a group as a whole find memorable from each meeting would allow a system or agent to continuously optimise its strategy to help a team meet its goals. In the present paper, we therefore investigate what participants take away from each meeting and how it is reflected in group dynamics.As a first step toward such a system, we recorded a multimodal longitudinal meeting corpus (MEMO), which comprises a first-party annotation of what participants remember from a discussion and why they remember it. We investigated whether participants of group interactions encode what they remember non-verbally and whether we can use such non-verbal multimodal features to predict what groups are likely to remember automatically. We devise a coding scheme to cluster participants’ memorisation reasons into higher-level constructs. We find that low-level multimodal cues, such as gaze and speaker activity, can predict conversational memorability. We also find that non-verbal signals can indicate when a memorable moment starts and ends. We could predict four levels of conversational memorability with an average accuracy of 44 %. We also showed that reasons related to participants’ personal feelings and experiences are the most frequently mentioned grounds for remembering meeting segments.
Maria Tsfasman, Kristian Fenech, Morita Tarvirdians, András Lörincz, Catholijn M. Jonker, Catharine Oertel
ICMI6
2022 The need for a female perspective in designing agent-based negotiation support
abstract
This study investigates whether an agent-based Negotiation Training System (NTS) can teach women Strategic Empathy - a recently introduced negotiation strategy based on perspective taking - and whether this can improve their negotiation performance. Developed and tested through an interaction-based real-time experiment was a NTS that integrated instructions on how to utilize Strategic Empathy. Women in the experimental group showed significantly higher levels of perspective-taking compared to the control group, and their understanding and use of Strategic Empathy increased over time. Also, a significant positive effect was found of Strategic Empathy on women's self-efficacy. No significant positive effect was found of Strategic Empathy on persistence. The high cognitive load of the experiment and a lack of intrinsic motivation may have caused this finding. Overall, this work demonstrates the applicability of using NTS to teach Strategic Empathy, and its effectiveness for enhancing women's self-efficacy in salary negotiations.
Katja Bouman, Iulia Lefter, Laurens Rook, Catharine Oertel, Catholijn M. Jonker, Frances M. T. Brazier
IVA4
2022 Giving Social Robots a Conversational Memory for Motivational Experience Sharing
abstract
In ongoing and consecutive conversations with persons, a social robot has to determine which aspects to remember and how to address them in the conversation. In the health domain, important aspects concern the health-related goals, the experienced progress (expressed sentiment) and the ongoing motivation to pursue them. Despite the progress in speech technology and conversational agents, most social robots lack a memory for such experience sharing. This paper presents the design and evaluation of a conversational memory for personalized behavior change support conversations on healthy nutrition via memory-based motivational rephrasing. The main hypothesis is that referring to previous sessions improves motivation and goal attainment, particularly when references vary. In addition, the paper explores how far motivational rephrasing affects user’s perception of the conversational agent (the virtual Furhat). An experiment with 79 participants was conducted via Zoom, consisting of three conversation sessions. The results showed a significant increase in participants’ change in motivation when multiple references to previous sessions were provided.
Avinash Saravanan, Maria Tsfasman, Mark A. Neerincx, Catharine Oertel
RO-MAN4
2021 Insights on Group and Team Dynamics
abstract
We are organizing again the workshop on Interdisciplinary Insights into Group and Team Dynamics which is a joint effort between researchers in the the ICMI and INGRoup (Interdisciplinary Network for Group Research) communities. This workshop aims to provide a common destination for researchers to exchange ideas and collaborate. We have found in previous years that instigating interdisciplinary collaborations can be hard. The aim of this workshop is to sustain a joint community to foster continued cross-disciplinary exchange and mutual understanding.
Joseph A. Allen, Hayley Hung, Joann Keyton, Gabriel Murray, Catharine Oertel, Giovanna Varni
ICMI5
2020 Supporting Empathy Training Through Virtual Patients
Jennifer K. Olsen 0001, Catharine Oertel
AIED (2)2
2020 Effects of Different Interaction Contexts when Evaluating Gaze Models in HRI
abstract
We previously introduced a responsive joint attention system that uses multimodal information from users engaged in a spatial reasoning task with a robot and communicates joint attention via the robot's gaze behavior. An initial evaluation of our system with adults showed it to improve users' perceptions of the robot's social presence. To investigate the repeatability of our prior findings across settings and populations, here we conducted two further studies employing the same gaze system with the same robot and task but in different contexts: evaluation of the system with external observers and evaluation with children. The external observer study suggests that third-person perspectives over videos of gaze manipulations can be used either as a manipulation check before committing to costly real-time experiments or to further establish previous findings. However, the replication of our original adults study with children in school did not confirm the effectiveness of our gaze manipulation, suggesting that different interaction contexts can affect the generalizability of results in human-robot interaction gaze studies.
André Pereira 0001, Catharine Oertel, Leonor Fermoselle, Joseph Mendelson, Joakim Gustafson
HRI2
2020 Workshop on Interdisciplinary Insights into Group and Team Dynamics
abstract
There has been gathering momentum over the last 10 years in the study of group behavior in multimodal multiparty interactions. While many works in the computer science community focus on the analysis of individual or dyadic interactions, we believe that the study of groups adds an additional layer of complexity with respect to how humans cooperate and what outcomes can be achieved in these settings. Moreover, the development of technologies that can help to interpret and enhance group behaviours dynamically is still an emerging field. Social theories that accompany the study of groups dynamics are in their infancy and there is a need for more interdisciplinary dialogue between computer scientists and social scientists on this topic. This workshop has been organised to facilitate those discussions and strengthen the bonds between these overlapping research communities
Hayley Hung, Gabriel Murray, Giovanna Varni, Nale Lehmann-Willenbrock, Fabiola H. Gerpott, Catharine Oertel
ICMI6
2020 Towards Understanding the Effect of Voice on Human-Agent Negotiation
abstract
Virtual agents are increasingly being used for communication training such as public speaking-, job interviews-, as well as negotiation training. In these use-cases the agent is generally taking on the role of interviewer and its behaviour is altered according to the nonverbal cues of its human interlocutor. However, understanding how the agent's non-verbal cues influence human behaviour, perception or interactions outcomes is equally important. This contributes to appropriate behaviour generation in agents, but also to our understanding of the intricate interplay of non-verbal behaviours on human perception and interaction outcomes.
Joanna Mania, Fieke Miedema, Rose Browne, Joost Broekens, Catharine Oertel
IVA5
2019 BloomGraph: Graph-Based Exploration of Bouquet Designs for Florist Apprentices
Kevin Gonyop Kim, Catharine Oertel, Pierre Dillenbourg
EC-TEL2
2019 On the Use of Gaze as a Measure for Performance in a Visual Exploration Task
Catharine Oertel, Alessia Coppi, Jennifer K. Olsen 0001, Alberto A. P. Cattaneo, Pierre Dillenbourg
EC-TEL1
2019 Responsive Joint Attention in Human-Robot Interaction
abstract
Joint attention has been shown to be not only crucial for human-human interaction but also human-robot interaction. Joint attention can help to make cooperation more efficient, support disambiguation in instances of uncertainty and make interactions appear more natural and familiar. In this paper, we present an autonomous gaze system that uses multimodal perception capabilities to model responsive joint attention mechanisms. We investigate the effects of our system on people's perception of a robot within a problem-solving task. Results from a user study suggest that responsive joint attention mechanisms evoke higher perceived feelings of social presence on scales that regard the direction of the robot's perception.
André Pereira 0001, Catharine Oertel, Leonor Fermoselle, Joe Mendelson, Joakim Gustafson
IROS2
2018 Group Interaction Frontiers in Technology
abstract
Analysis of group interaction and team dynamics is an important topic in a wide variety of fields, owing to the amount of time that individuals typically spend in small groups for both professional and personal purposes, and given how crucial group cohesion and productivity are to the success of businesses and other organizations. This fact is attested by the rapid growth of fields such as People Analytics and Human Resource Analytics, which in turn have grown out of many decades of research in social psychology, organizational behaviour, computing, and network science, amongst other fields. The goal of this workshop is to bring together researchers from diverse fields related to group interaction, team dynamics, people analytics, multi-modal speech and language processing, social psychology, and organizational behaviour.
Gabriel Murray, Hayley Hung, Joann Keyton, Catherine Lai, Nale Lehmann-Willenbrock, Catharine Oertel
ICMI6
2018 Predicting Group Performance in Task-Based Interaction
abstract
We address the problem of automatically predicting group performance on a task, using multimodal features derived from the group conversation. These include acoustic features extracted from the speech signal, and linguistic features derived from the conversation transcripts. Because much work on social signal processing has focused on nonverbal features such as voice prosody and gestures, we explicitly investigate whether features of linguistic content are useful for predicting group performance. The conclusion is that the best-performing models utilize both linguistic and acoustic features, and that linguistic features alone can also yield good performance on this task. Because there is a relatively small amount of task data available, we present experimental approaches using domain adaptation and a simple data augmentation method, both of which yield drastic improvements in predictive performance, compared with a target-only model.
Gabriel Murray, Catharine Oertel
ICMI2
2018 FARMI: A FrAmework for Recording Multi-Modal Interactions
Patrik Jonell, Mattias Bystedt, Per Fallgren, Dimosthenis Kontogiorgos, José Lopes 0001, Zofia Malisz, Samuel Mascarenhas, Catharine Oertel, Eran Raveh, Todd Shore
LREC8
2018 Crowdsourced Multimodal Corpora Collection Tool
Patrik Jonell, Catharine Oertel, Dimosthenis Kontogiorgos, Jonas Beskow, Joakim Gustafson
LREC2
2018 A Multimodal Corpus for Mutual Gaze and Joint Attention in Multiparty Situated Interaction
Dimosthenis Kontogiorgos, Vanya Avramova, Simon Alexanderson, Patrik Jonell, Catharine Oertel, Jonas Beskow, Gabriel Skantze, Joakim Gustafson
LREC5
2017 Crowd-Sourced Design of Artificial Attentive Listeners
abstract
Feedback generation is an important component of humanhuman communication. Humans can choose to signal support, understanding, agreement or also sceptiscism by means of feedback tokens. Many studie ...
Catharine Oertel, Patrik Jonell, Dimosthenis Kontogiorgos, Joseph Mendelson, Jonas Beskow, Joakim Gustafson
INTERSPEECH1
2017 Crowd-Powered Design of Virtual Attentive Listeners
Patrik Jonell, Catharine Oertel, Dimosthenis Kontogiorgos, Jonas Beskow, Joakim Gustafson
IVA2
2016 Towards building an attentive artificial listener: on the perception of attentiveness in audio-visual feedback tokens
abstract
Current dialogue systems typically lack a variation of audio-visual feedback tokens. Either they do not encompass feedback tokens at all, or only support a limited set of stereotypical functions. However, this does not mirror the subtleties of spontaneous conversations. If we want to be able to build an artificial listener, as a first step towards building an empathetic artificial agent, we also need to be able to synthesize more subtle audio-visual feedback tokens. In this study, we devised an array of monomodal and multimodal binary comparison perception tests and experiments to understand how different realisations of verbal and visual feedback tokens influence third-party perception of the degree of attentiveness. This allowed us to investigate i) which features (amplitude, frequency, duration...) of the visual feedback influences attentiveness perception; ii) whether visual or verbal backchannels are perceived to be more attentive iii) whether the fusion of unimodal tokens with low perceived attentiveness increases the degree of perceived attentiveness compared to unimodal tokens with high perceived attentiveness taken alone; iv) the automatic ranking of audio-visual feedback token in terms of conveyed degree of attentiveness.
Catharine Oertel, José Lopes 0001, Yu Yu 0003, Kenneth Alberto Funes Mora, Joakim Gustafson, Alan W. Black, Jean-Marc Odobez
ICMI1
2016 Towards Building an Attentive Artificial Listener: On the Perception of Attentiveness in Feedback Utterances
abstract
Towards Building an Attentive Artificial Listener: On the Perception of Attentiveness in Feedback Utterances
Catharine Oertel, Joakim Gustafson, Alan W. Black
INTERSPEECH1
2015 Deciphering the Silent Participant: On the Use of Audio-Visual Cues for the Classification of Listener Categories in Group Discussions
abstract
Estimating a silent participant's degree of engagement and his role within a group discussion can be challenging, as there are no speech related cues available at the given time. Having this information available, however, can provide important insights into the dynamics of the group as a whole. In this paper, we study the classification of listeners into several categories (attentive listener, side participant and bystander). We devised a thin-sliced perception test where subjects were asked to assess listener roles and engagement levels in 15-second video-clips taken from a corpus of group interviews. Results show that humans are usually able to assess silent participant roles. Using the annotation to identify from a set of multimodal low-level features, such as past speaking activity, backchannels (both visual and verbal), as well as gaze patterns, we could identify the features which are able to distinguish between different listener categories. Moreover, the results show that many of the audio-visual effects observed on listeners in dyadic interactions, also hold for multi-party interactions. A preliminary classifier achieves an accuracy of 64 %.
Catharine Oertel, Kenneth Alberto Funes Mora, Joakim Gustafson, Jean-Marc Odobez
ICMI1
2014 Human-robot collaborative tutoring using multiparty multimodal spoken dialogue
abstract
In this paper, we describe a project that explores a novel experimental setup towards building a spoken, multi-modally rich, and human-like multiparty tutoring robot. A human-robot interaction setup is designed, and a human-human dialogue corpus is collected. The corpus targets the development of a dialogue system platform to study verbal and nonverbal tutoring strategies in multiparty spoken interactions with robots which are capable of spoken dialogue. The dialogue task is centered on two participants involved in a dialogue aiming to solve a card-ordering game. Along with the participants sits a tutor (robot) that helps the participants perform the task, and organizes and balances their interaction. Different multimodal signals captured and auto-synchronized by different audio-visual capture technologies, such as a microphone array, Kinects, and video cameras, were coupled with manual annotations. These are used build a situated model of the interaction based on the participants personalities, their state of attention, their conversational engagement and verbal dominance, and how that is correlated with the verbal and visual feed-back, turn-management, and conversation regulatory actions generated by the tutor. Driven by the analysis of the corpus, we will show also the detailed design methodologies for an affective, and multimodally rich dialogue system that allows the robot to measure incrementally the attention states, and the dominance for each participant, allowing the robot head Furhat to maintain a well-coordinated, balanced, and engaging conversation, that attempts to maximize the agreement and the contribution to solve the task.
Samer Al Moubayed, Jonas Beskow, Bajibabu Bollepalli, Joakim Gustafson, Ahmed Hussen Abdelaziz, Martin Johansson, Maria Koutsombogera, José Lopes 0001, Jekaterina Novikova, Catharine Oertel, Gabriel Skantze, Kalin Stefanov, Gül Varol
HRI10
2014 The Tutorbot Corpus ― A Corpus for Studying Tutoring Behaviour in Multiparty Face-to-Face Spoken Dialogue
Maria Koutsombogera, Samer Al Moubayed, Bajibabu Bollepalli, Ahmed Hussen Abdelaziz, Martin Johansson, José Lopes 0001, Jekaterina Novikova, Catharine Oertel, Kalin Stefanov, Gül Varol
LREC8
2014 Turn-taking, feedback and joint attention in situated human-robot interaction
Gabriel Skantze, Anna Hjalmarsson, Catharine Oertel
Speech Commun.3
2013 Towards developing a model for group involvement and individual engagement
abstract
This PhD project is concerned with the multi-modal modeling of conversational dynamics. In particular I focus on investigating how people organise themselves within a multiparty conversation. I am interested in identifying bonds between people, their individual engagement level in the conversation and how the engagement level of the individual person influences the perceived involvement of the whole group of people. To this end machine learning experiments are carried out and I am planning to build a conversational involvement module to be implemented in a dialogue system.
Catharine Oertel
ICMI1
2013 A gaze-based method for relating group involvement to individual engagement in multimodal multiparty dialogue
abstract
This paper is concerned with modelling individual engagement and group involvement as well as their relationship in an eight-party, mutimodal corpus. We propose a number of features (presence, entropy, symmetry and maxgaze) that summarise different aspects of eye-gaze patterns and allow us to describe individual as well as group behaviour in time. We use these features to define similarities between the subjects and we compare this information with the engagement rankings the subjects expressed at the end of each interactions about themselves and the other participants. We analyse how these features relate to four classes of group involvement and we build a classifier that is able to distinguish between those classes with 71\% of accuracy.
Catharine Oertel, Giampiero Salvi
ICMI1
2013 User feedback in human-robot interaction: prosody, gaze and timing
abstract
This paper investigates forms and functions of user feedback in a map task dialogue between a human and a robot, where the robot is the instruction-giver and the human is the instruction-follower. First, we investigate how user acknowledgements in task-oriented dialogue signal whether an activity is about to be initiated or has been completed. The parameters analysed include the users ’ lexical and prosodic realisation as well as gaze direction and response timing. Second, we investigate the relation between these parameters and the perception of uncertainty.
Gabriel Skantze, Catharine Oertel, Anna Hjalmarsson
INTERSPEECH2
2013 Exploring the effects of gaze and pauses in situated human-robot interaction
Gabriel Skantze, Anna Hjalmarsson, Catharine Oertel
SIGDIAL Conference3
2012 Gaze Patterns in Turn-Taking
abstract
Oertel C, Wlodarczak M, Edlund J, Wagner P, Gustafson J. Gaze patterns in turn-taking. In: 13th Annual Conference of the International Speech Communication Association 2012 (INTERSPEECH 2012). Red Hook, NY: Curran; 2013: 2243-2246.
Catharine Oertel, Marcin Wlodarczak, Jens Edlund, Petra Wagner, Joakim Gustafson
INTERSPEECH1
2011 On the Use of Multimodal Cues for the Prediction of Degrees of Involvement in Spontaneous Conversation
abstract
Quantifying the degree of involvement of a group of participants in a conversation is a task which humans accomplish every day, but it is something that, as of yet, machines are unable to do. In this study we first investigate the correlation between visual cues (gaze and blinking rate) and involvement. We then test the suitability of prosodic cues (acoustic model) as well as gaze and blinking (visual model) for the prediction of the degree of involvement by using a support vector machine (SVM). We also test whether the fusion of the acoustic and the visual model improves the prediction. We show that we are able to predict three classes of involvement with an reduction of error rate of 0.30 (accuracy =0.68).
Catharine Oertel, Stefan Scherer, Nick Campbell 0001
INTERSPEECH1