Elisabeth André

dblp:a/EAndre · DBLP profile ↗
← Back
263ranked-venue papers
24as first author
88since 2021 · last 2026
0000-0002-2367-162XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 155 · 9 first-author · 43 since 2021Artificial intelligence and machine learning · 115 · 14 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 65 · 4 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 29 · 17 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Security and privacy · 2
YearPublicationVenuePosition
2026 CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty
abstract
Existing benchmarks for Large Language Model (LLM) agents focus on task completion under idealistic settings but overlook reliability in real-world, user-facing applications.In domains, such as in-car voice assistants, users often issue incomplete or ambiguous requests, creating intrinsic uncertainty that agents must manage through dialogue, tool use, and policy adherence.We introduce CAR-bench, a benchmark for evaluating consistency, uncertainty handling, and capability awareness in multi-turn, tool-using LLM agents in an in-car assistant domain.The environment features an LLM-simulated user, domain policies, and 58 interconnected tools spanning navigation, productivity, charging, and vehicle control.Beyond standard task completion, CAR-bench introduces Hallucination tasks that test agents' limit-awareness under missing tools or information, and Disambiguation tasks that require resolving uncertainty through clarification or internal information gathering.Baseline results reveal large gaps between occasional and consistent success on all task types.Even frontier reasoning LLMs achieve less than 50% consistent pass rate on Disambiguation tasks due to premature actions, and frequently violate policies or fabricate information to satisfy user requests in Hallucination tasks, underscoring the need for more reliable and self-aware LLM agents in real-world settings.1
Johannes Kirmayr, Lukas Stappen, Elisabeth André
ACL (1)3
2026 Future Horizons in Human-AI Interaction: Joint Perspectives from HCI and AI
abstract
Recent advances in artificial intelligence (AI), including generative models and socially interactive agents, are reshaping the design of interactive systems and raising new challenges for Human–Computer Interaction (HCI). While AI research has traditionally focused on improving performance and scalability, HCI emphasizes usability, transparency, and the broader social implications of technology, all aspects that imply the application of proper human-centered design methods. Bridging these perspectives is increasingly critical as AI moves from backend functionality to a central role in user interaction. This paper reports on the workshop Future Horizons in Human–AI Interaction: Joint Perspectives from HCI and AI, held at AVI 2026. It synthesizes key insights on emerging challenges and opportunities in designing human-centered AI systems, emphasizing the importance of integrating perspectives from HCI and AI, to advance the design of interactive AI systems that are both technically robust and aligned with human values.
Elisabeth André, Cristina Conati, Shelly Levy-Tzedek, Maristella Matera, Micol Spitale
AVI1
2026 From 'Nice Try' to 'Nice Throw': Exploring Counterfactual Explanations as Corrective Feedback for Javelin Throwing
abstract
Providing athletes with feedback to refine their technique is key in sports coaching and is critical for improving performance and preventing injuries. However, access to expert coaching is often limited. In this paper, we explore a novel counterfactual-based feedback system as a complementary tool to expert coaching and conduct a small-scale user study to explore its perceived usability. Our approach uses an augmented GANterfactual framework, a modified CycleGAN architecture with a classifier-guided counterfactual loss, to synthesize plausible, actionable feedback. As a test bed for our approach, we use the complex motor task of javelin throwing, a sport that is characterized by high biomechanical demands and injury risk. As we are interested in the perceived usability of our approach, we conduct a user study with 21 sports students. The subjective feedback provided by participants of our user study shows that, while pose-based counterfactual feedback visualizations are appreciated by athletes, for some users they require too much domain-specific knowledge and are not “coach-like” enough. We find that athletes are looking for accompanying textual feedback, supporting recent research in the field of feedback generation for sports and motor learning.
Lennart Eing, Annika Stippler, Cristina Conati, Stefan Künzell, Elisabeth André, Silvan Mertes
AVI5
2026 "What Are You Doing?": Effects of Intermediate Feedback from Agentic LLM In-Car Assistants During Multi-Step Processing
abstract
Agentic AI assistants that autonomously perform multi-step tasks raise open questions for user experience: how should such systems communicate progress and reasoning during extended operations, especially in attention-critical contexts such as driving? We investigate feedback timing and verbosity from agentic LLM-based in-car assistants through a controlled, mixed-methods study (N=45) comparing planned steps and intermediate results feedback against silent operation with final-only response. Using a dual-task paradigm with an in-car voice assistant, we found that intermediate feedback significantly improved perceived speed, trust, and user experience while reducing task load - effects that held across varying task complexities and interaction contexts. Interviews further revealed user preferences for an adaptive approach: high initial transparency to establish trust, followed by progressively reducing verbosity as systems prove reliable, with adjustments based on task stakes and situational context. We translate our empirical findings into design implications for feedback timing and verbosity in agentic in-car assistants, balancing transparency and efficiency.
Johannes Kirmayr, Raphael Wennmacher, Khanh Huynh, Lukas Stappen, Elisabeth André, Florian Alt
CHI5
2026 Adaptive Sequencing in Interval Ear Training: A Multi-Armed Bandit Approach
Yasmine Elsadat, Anan Schütt, Hannes Ritschel, Elisabeth André
CSEDU (1)4
2026 Integrating LLM-based Explanations into Open-Ended Graph Practice Exercises for Increased Learning and Engagement
Ali Mahmoud Shokry, Anan Schütt, Elisabeth André
CSEDU (1)3
2026 "We Will Grow into the Age of Robots": A Participatory Interview Study for Service Robots and Their Value for Care
abstract
The aging population and chronic staff shortages are prompting care facilities to use service robots (SR) as part of daily care. There are high hopes for support with physical workloads and routine tasks, but adoption often stalls due to technical complexity, poor integration into workflows, and fears that "support" could become "replacement". We address this problem by viewing care as a value-driven practice rather than a list of tasks. In a participatory interview study with caregivers and care recipients in three facilities, based on value-sensitive design, we identified expectations, non-negotiable boundaries, and the values that should guide robot behavior. Participants identified credible roles for SRs in logistics, documentation, reminders, and guidance, but rejected intimate or safety-critical care tasks. Acceptance depends on value-oriented and fluid adaptivity. Robots should dynamically modulate initiative, proactivity, and interaction modality to maintain human attentiveness and warmth, sustain independence, support control over workload, and take legal safeguards into account. We contribute to this (1) with an empirically grounded overview of acceptable potentials and limitations guided by stakeholder values, (2) with a value-sensitive design-based framework for fluid adaptivity as a mechanism that operationalizes values in daily interaction, and (3) design requirements for user-centered, transparent, and context-sensitive SRs that reduce workload and create space for human care rather than replacing it.
Stina Klein, Shuyuan Shen, Elisabeth André, Matthias Kraus 0001
HRI3
2026 Optimizing Sequential Models through Temporal Landmark Selection and Normalization for Sign Language Recognition
Sergio Esteban Romero, Iván Martín-Fernández, Cristina Luna Jiménez, Manuel Gil-Martín, Fernando Fernández Martínez, Elisabeth André
ICAART (3)6
2026 Annotating Conversational Phases and Communication Techniques: A Corpus of German Teacher-Parent Counseling Conversations
Tobias Hallmen, Kathrin Gietl, Karoline Hillesheim, Annemarie Friedrich, Elisabeth André
LREC5
2026 Evaluation of Failure Communication Strategies for Trust Repair in Human-AI Collaboration
Stina Klein, Alexandru Wurm, Elisabeth André, Matthias Kraus 0001
LREC3
2026 MUDiC: A Dataset for Multi-User Dialogue and Collaboration in Chatbot Interaction
Nicolas Wagner 0001, Cristina Luna Jiménez, Elisabeth André, Wolfgang Minker, Stefan Ultes
LREC3
2025 Physiological and Cognitive Responses to Walking in Natural and Built Urban Environments
abstract
Walking in natural environments is widely recognized as an effective stress reduction strategy, often offering greater benefits than walking in built environments. We examined the physiological and cognitive responses to walking in urban forest versus urban built environments in summer and in winter. This study utilized continuous heart rate monitoring with a wearable chest sensor and 2-back cognitive tests. Higher increases in heart rate during the walk and slower post-walk recovery were observed for walks in built environments compared to those in the forest, in both seasons. However, the magnitudes varied between the seasons, emphasizing the contextual nature of restorative benefits. Improvements in the accuracy of the cognitive tests were observed during the forest walks in summer, but the results were less conclusive in winter. Despite these differences, walking in built environments still conferred well-being benefits, supporting stress reduction regardless of the environment or season.
Bhargavi Mahesh, Jauwairia Nasir, Stina Klein, Tobias Hallmen, Yekta Said Can, Jonathan Simon, Christoph Beck, Joachim Rathmann, Max Stocker, Lisa-Marie Falkenrodt, Elisabeth André
BSN11
2025 Live Link's Awakening of a Humorous Real-Time Character
abstract
Virtual characters require the real-time streaming of verbal and nonverbal behaviors for the expression of dynamically generated humor. In this paper, we present the Live Link Animator, a real-time solution for multimodal animation of Unreal Engine characters using individual blendshapes. We demonstrate the tool through an example interaction with a MetaHuman character and outline potential areas of application in the domain of virtual agent humor research.
Thomas Kiderle, Jauwairia Nasir, Georgiana Cristina Dobre, Carlos González Díaz, Elisabeth André, Hannes Ritschel
HAI5
2025 What do the Face and Voice Reveal? Investigating Trust Dynamics During Human-Robot Interaction
abstract
Existing research has shown that vocal and non-vocal human cues correlate with human trust and distrust behaviours, suggesting their potential to measure human trust in robots in real-time. However, there is a lack of research in Human-Robot Interaction that integrates vocal and non-vocal cues into a comprehensive model to measure trust. This paper aims to estimate human trust in robots by examining vocal and non-vocal cues differences between trust and distrust states across multiple sessions of collaborative game-based HRI with 40 participants. Our analysis revealed that vocal and non-vocal human cues can indeed predict trust in HRI, with certain facial expressions, facial movements, and pitch being significant factors. Random Forest classifier achieved the highest accuracy (84 %) in classifying trust states, with key features such as facial expressions (fear, angry), facial blendshapes (cheekSquintRight, jawRight), and vocal characteristics (Duration, Harmonicity std) being the most predictive of trust. These findings demonstrate the importance of combining vocal and non-vocal cues for accurate trust measurement and highlight the potential for real-time trust assessment in robotic systems.
Abdullah S. Alzahrani, Jauwairia Nasir, Ahmad Tayeb, Elisabeth André, Muneeb Imtiaz Ahmad
HRI4
2025 Lightweight Transformers for Isolated Sign Language Recognition
Cristina Luna Jiménez, Lennart Eing, Annalena Aicher, Fabrizio Nunnari, Elisabeth André
ICMI5
2025 CO-PARLEY: A Co-Regulative Socially Interactive Agent for Emotion Regulation Support
abstract
This demo presents CO-PARLEY, a mobile socially interactive agent designed to support individuals experiencing difficulties with emotion regulation.The system engages users through reciprocal coregulation, a dynamic and two-way process in which the agent and user mutually influence each other's emotional and physiological states.By combining verbal and nonverbal interaction with real-time physiological synchrony, CO-PARLEY fosters therapeutic alliance, trust, and emotional awareness.Built on a modular framework that integrates multimodal sensing, dialogue management, and adaptive behavior generation, the agent supports users in emotionally challenging moments.In general, this alliance can improve the effectiveness of psychoeducational and awareness exercises.
Mina Ameli, Chirag Bhuvaneshwara, Janet Wessler, Michael Dietz, Tanja Schneeberger, Elisabeth André, Patrick Gebhard
IVA6
2025 SIA-Lab: A Platform for Exploring Assistive and Supportive Socially Interactive Agents
abstract
This paper introduces SIA-Lab, a versatile and broadly applicable platform to advance socially interactive agents (SIAs).Unlike existing single-use case systems, SIA-Lab is a modular and scalable framework built on Android and mobile technologies, enabling rapid explorative and comparative studies between human-human and human-agent interactions, offering valuable insights into behavioral dynamics and user engagement.Effective in health-related applications, including screening, therapeutic assistance, and posttreatment care, it is equally adaptable to education and other settings.By integrating dialog management, affective modeling, and multimodal interaction analysis, SIA-Lab represents a comprehensive toolkit for researchers and practitioners to evaluate and design next-generation supportive technologies.
Mina Ameli, Tanja Schneeberger, Janet Wessler, Michael Dietz, Elisabeth André, Patrick Gebhard
IVA5
2025 Multimodal Generation of Contextualized Jokes for a Real-Time Virtual Character
abstract
Humor often serves as a catalyst for smoother interpersonal communication, enhancing interaction experience between individuals.While virtual characters can also gain from these benefits, implementing humor naturally in human-character interactions remains an open challenge.In this paper, we propose the Joking and Multimodally Amusing Real-Time Character (J-MARC) system, combining a photorealistic character with advanced large language model (LLM) techniques to contextualize jokes within small talk.In the real-time interaction, the character is able to present the jokes multimodally and to apply nonverbal behavior while listening.
Thomas Kiderle, Georgiana Cristina Dobre, Jauwairia Nasir, Carlos González Díaz, Hannes Ritschel, Stina Klein, Silvan Mertes, Elisabeth André
IVA8
2025 VoiceX as a Design Tool for Virtual Agents' Voices
abstract
Modern TTS systems are capable of creating highly realistic and natural-sounding speech, making them an important tool when designing virtual agents.While sounding highly realistic, the process of customizing such TTS voices remains a complex task, mostly requiring the expertise of specialists within the field.One reason for this is the utilization of deep learning models, which are characterized by their expansive, non-interpretable parameter spaces, restricting the feasibility of manual voice customization.In this paper, we present a novel human-in-the-loop paradigm based on an evolutionary algorithm for directly interacting with the parameter space of a neural TTS model.We integrated our approach into a user-friendly graphical user interface that allows users to efficiently create original voices.Those voices can then be used to equip virtual agents with highly customized TTS capabilities by using an open-source programming interface provided by us.Further, in a first pilot study, we show that VoiceX is an appropriate tool for creating individual, custom voices.
Daksitha Withanage, Florian Lingenfelser, Johanna Magdalena Kuch, Otto Grothe, Ruben Schlagowski, Elisabeth André, Silvan Mertes
IVA6
2025 MultiMediate '25: Cross-cultural Multi-domain Engagement Estimation
abstract
Estimating momentary conversational engagement is central to assistive, socially aware AI systems, yet models are typically trained and evaluated within a single domain, limiting real-world robustness. The MultiMediate '25 challenge advances engagement estimation to more challenging, cross-cultural, and multi-domain settings. Building on prior challenge editions, we expand beyond NOXI as the sole training source by introducing NOXI-J, a new multilingual corpus covering Japanese and Chinese interactions, enabling both training and evaluation in diverse linguistic contexts. Although NOXI-J conceptually extends NOXI, we treat it as a distinct domain because linguistic, cultural, capture, and annotation differences induce measurable distribution shifts. In this paper, we present new annotations, precomputed multi-modal features (visual, vocal, and verbal), baseline evaluations, and an analysis of the best performing challenge solutions. Beyond accuracy, we quantify fairness using Conditional Demographic Disparity for gender and language. Our baselines confirm strong in-domain performance (e.g., paralinguistic eGeMAPS and video-transformer features) and reveal notable cross-domain drops, underscoring the challenge of cultural, linguistic, and interactional shifts. Fairness analyses indicate generally small discrepancies for our baselines. We observe the largest disparities for the proposed challenge solutions on the Chinese language test set. All annotations, features, code, and leaderboards are made publicly available to foster sustained progress on robust and fair engagement estimation.
Daksitha Withanage, Marius Funk, Michal Balazia, Huajian Qiu, Shogo Okada, François Brémond, Jan Alexandersson, Andreas Bulling, Elisabeth André, Philipp Müller 0001
ACM Multimedia9
2025 REACT 2025: the Third Multiple Appropriate Facial Reaction Generation Challenge
abstract
In dyadic interactions, a broad spectrum of human facial reactions might be appropriate for responding to each human speaker behaviour. Following the successful organisation of the REACT 2023 and REACT 2024 challenges, we are proposing the REACT 2025 challenge encouraging the development and benchmarking of Machine Learning (ML) models that can be used to generate multiple appropriate, diverse, realistic and synchronised human-style facial reactions expressed by human listeners in response to an input stimulus (i.e., audio-visual behaviours expressed by their corresponding speakers). As a key of the challenge, we provide challenge participants with the first natural and large-scale multi-modal Multiple Appropriate Facial Reaction Generation (MAFRG) dataset (called MARS) recording 136 human-human dyadic interactions containing a total of 2856 interaction sessions covering five different topics. In addition, this paper also presents the challenge guidelines and the performance of our baselines on the two proposed sub-challenges: Offline MAFRG and Online MAFRG, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2025
Siyang Song, Micol Spitale, Xiangyu Kong 0001, Hengde Zhu, Cristina Palmero, Germán Barquero, Sergio Escalera, Michel F. Valstar, Mohamed Daoudi, Tobias Baur 0001, Fabien Ringeval, Andrew Howes 0001, Elisabeth André, Hatice Gunes
ACM Multimedia14
2025 Your Robot, My Voice: Enhancing Android Robot Likability through Personalization by Cloning the User's Voice
abstract
This study investigates whether personalized voice cloning can improve a robot’s likability compared to a design-congruent voice and a distinctly dissimilar voice. Participants interacted with a gender-ambiguous android robot in three different voice conditions. We compared: (1) a personalized voice clone based on the participant’s voice, (2) a design-congruent voice matching the robot’s appearance, and (3) a dissimilar voice, which differs from both the participant’s and the robot’s features.The cloned and design-congruent voices significantly increased likability compared to the dissimilar voice, while anthropomorphism and familiarity showed no significant differences across conditions. Most participants did not immediately recognize their cloned voice until informed that one of the voices was a clone. However, most of the participants were successful when asked to pick out their cloned voice from those used. We assume that voice personalization through similarity to the user improves likability even before the user is aware of this similarity.Our results show that personalized voice cloning is a simple alternative to other methods for the design of robotic voices. It significantly increases robot likability while requiring minimal user effort.
Johanna Magdalena Kuch, Marcel Heisler, Stina Klein, Silvan Mertes, Lennart Eing, Elisabeth André, Christian Becker-Asano
RO-MAN6
2025 On Speakers' Identities, Autism Self-Disclosures and LLM-Powered Robots
abstract
Dialogue agents become more engaging through recipient design, which needs user-specific information. However, a user’s identification with marginalized communities, such as migration or disability background, can elicit biased language. This study compares LLM responses to neurodivergent user personas with disclosed vs. masked neurodivergent identities. A dataset built from public Instagram comments was used to evaluate four open-source models on story generation, dialogue generation, and retrieval-augmented question answering. Our analyses show biases in user’s identity construction across all models and tasks. Binary classifiers trained on each model can distinguish between language generated for prompts with or without self-disclosures, with stronger biases linked to more explicit disclosures. Some models’ safety mechanisms result in denial of service behaviors. LLM’s recipient design to neurodivergent identities relies on stereotypes tied to neurodivergence.
Sviatlana Höhn, Fred Philippy, Elisabeth André
SIGDIAL3
2025 Application of Multimodal Self-Supervised Architectures for Daily Life Affect Recognition
abstract
The recognition of affects (an umbrella term including but not limited to emotions, mood, and stress) in daily life is crucial for maintaining mental well-being and preventing long-term health issues. Wearable devices, such as smart bands, can collect physiological data including heart rate variability, electrodermal activity, skin temperature, and acceleration facilitating daily life affect monitoring via machine learning models. However, accurately labeling this data for model evaluation is challenging in affective computing research, as individuals often provide subjective, inaccurate, or incomplete labels in their daily lives. This study introduces the adaptation of self-supervised learning architectures for multimodal daily life stress and emotion recognition tasks, focusing on self-representation and contrastive learning methods. By leveraging unlabeled multimodal physiological signals, we aim to alleviate the need for extensive labeled data and enhance model generalizability. Our research demonstrates that self-supervised learning can effectively learn meaningful representations from physiological data without explicit labels, offering a promising approach for developing robust affect recognition systems that can operate in dynamic and uncontrolled environments. This work represents a significant improvement in recognizing affects in the wild, with potential implications for personalized mental health support and timely interventions.
Yekta Said Can, Mohamed Benouis, Bhargavi Mahesh, Elisabeth André
IEEE Trans. Affect. Comput.4
2025 The ForDigitStress Dataset: A Multi-Modal Dataset for Automatic Stress Recognition
abstract
We present a multi-modal stress dataset that uses digital job interviews to induce stress. The dataset provides multi-modal data of 40 participants including audio, video (motion capturing, facial landmarks, eye tracking), as well as physiological information (photoplethysmography, electrodermal activity). In addition to that, the dataset contains time-continuous annotations for stress and occurred emotions (e.g., shame, anger, anxiety, and surprise). In order to establish a baseline, five different machine learning classifiers (Support Vector Machine, K-Nearest Neighbors, Random Forest, Feed-forward Neural Network, and Long-Short-Term Memory Network) have been trained and evaluated on the presented dataset for a binary stress classification task. The best-performing classifier has been a Long-Short-Term Memory Network, which achieved an accuracy of 91.7% and an F1-score of 90.2%. The ForDigitStress dataset is freely available to other researchers.
Alexander Heimerl, Pooja Prajod, Silvan Mertes, Tobias Baur 0001, Matthias Kraus 0001, Ailin Liu, Helen Risack, Nicolas Rohleder, Elisabeth André, Linda Becker
IEEE Trans. Affect. Comput.9
2025 Guest Editorial Extremely Low-Resource Autonomous Affective Learning
Xinzhou Xu, Björn W. Schuller, Elisabeth André, Erik Cambria
IEEE Trans. Affect. Comput.3
2025 GANonymization: A GAN-Based Face Anonymization Framework for Preserving Emotional Expressions
abstract
In recent years, the increasing availability of personal data has raised concerns regarding privacy and security. One of the critical processes to address these concerns is data anonymization, which aims to protect individual privacy and prevent the release of sensitive information. This research focuses on the importance of face anonymization. Therefore, we introduce GANonymization, a novel face anonymization framework with facial expression-preserving abilities. Our approach is based on a high-level representation of a face, which is synthesized into an anonymized version based on a generative adversarial network (GAN). The effectiveness of the approach was assessed by evaluating its performance in removing identifiable facial attributes to increase the anonymity of the given individual face. Additionally, the performance of preserving facial expressions was evaluated on several affect recognition datasets and outperformed the state-of-the-art methods in most categories. Finally, our approach was analyzed for its ability to remove various facial traits, such as jewelry, hair color, and multiple others. Here, it demonstrated reliable performance in removing these attributes. Our results suggest that GANonymization is a promising approach for anonymizing faces while preserving facial expressions.
Fabio Hellmann, Silvan Mertes, Mohamed Benouis, Alexander Hustinx, Tzung-Chien Hsieh, Cristina Conati, Peter M. Krawitz, Elisabeth André
ACM Trans. Multim. Comput. Commun. Appl.8
2024 The AffectToolbox: Affect Analysis for Everyone
abstract
In the field of affective computing, where research continually advances at a rapid pace, the demand for user-friendly tools has become increasingly apparent. In this paper, we present the AffectToolbox, a novel software system that aims to support researchers in developing affect-sensitive studies and prototypes. The proposed system addresses the challenges posed by existing frameworks, which often require profound programming knowledge and cater primarily to power-users or skilled developers. Aiming to facilitate ease of use, the AffectToolbox requires no programming knowledge and offers its functionality to reliably analyze the affective state of users through an accessible graphical user interface. The architecture encompasses a variety of models for emotion recognition on multiple affective channels and modalities, as well as an elaborate fusion system to merge multi-modal assessments into a unified result. The entire system is open-sourced and will be publicly available to ensure easy integration into more complex applications through a well-structured, Python-based code base - therefore marking a substantial contribution toward advancing affective computing research and fostering a more collaborative and inclusive environment within this interdisciplinary field.
Silvan Mertes, Dominik Schiller, Michael Dietz, Elisabeth André, Florian Lingenfelser
ACII4
2024 Recognizing Emotion Regulation Strategies from Human Behavior with Large Language Models
abstract
Human emotions are often not expressed directly, but regulated according to internal processes and social display rules. For affective computing systems, an understanding of how users regulate their emotions can be highly useful, for example to provide feedback in job interview training, or in psychotherapeutic scenarios. However, at present no method to automatically classify different emotion regulation strategies in a cross-user scenario exists. At the same time, recent studies showed that instruction-tuned Large Language Models (LLMs) can reach impressive performance across a variety of affect recognition tasks such as categorical emotion recognition or sentiment analysis. While these results are promising, it remains unclear to what extent the representational power of LLMs can be utilized in the more subtle task of classifying users' internal emotion regulation strategy. To close this gap, we make use of the recently introduced Deep corpus for modeling the social display of the emotion shame, where each point in time is annotated with one of seven different emotion regulation classes. We fine-tune Llama2-7B as well as the recently introduced Gemma model using Low-rank Optimization on prompts generated from different sources of information on the Deep corpus. These include verbal and nonverbal behavior, person factors, as well as the results of an indepth interview after the interaction. Our results show, that a fine-tuned Llama2-7B LLM is able to classify the utilized emotion regulation strategy with high accuracy (0.84) without needing access to data from post-interaction interviews. This represents a significant improvement over previous approaches based on Bayesian Networks and highlights the importance of modeling verbal behavior in emotion regulation.
Philipp Müller 0001, Alexander Heimerl, Sayed Muddashir Hossain, Lea Siegel, Jan Alexandersson, Patrick Gebhard, Elisabeth André, Tanja Schneeberger
ACII7
2024 Estimating Chess Puzzle Difficulty Without Past Game Records Using a Human Problem-Solving Inspired Neural Network Architecture
abstract
For chess players to sharpen their tactical skills effectively, they train on chess puzzles with a fitting difficulty level. This paper presents an approach to estimate the difficulty level of chess puzzles using a deep neural network. The proposed approach achieved second place in the IEEE BigData Cup 2024 competition: Predicting chess puzzle difficulty. For the design of our network architecture, we take inspiration from the human problem-solving process for chess puzzles. We train the model to predict the correct move as an auxiliary task to improve the training process. We also predict themes, which are patterns in chess puzzles as a second auxiliary task. Finally, we use the uncertainty in the position, i.e. how incorrect the model’s move prediction is, as a further input to guide the estimation of the puzzle difficulty.
Anan Schütt, Tobias Huber, Elisabeth André
IEEE Big Data3
2024 Giving Robots a Voice: Human-in-the-Loop Voice Creation and open-ended Labeling
abstract
Speech is a natural interface for humans to interact with robots. Yet, aligning a robot’s voice to its appearance is challenging due to the rich vocabulary of both modalities. Previous research has explored a few labels to describe robots and tested them on a limited number of robots and existing voices. Here, we develop a robot-voice creation tool followed by large-scale behavioral human experiments (N=2,505). First, participants collectively tune robotic voices to match 175 robot images using an adaptive human-in-the-loop pipeline. Then, participants describe their impression of the robot or their matched voice using another human-in-the-loop paradigm for open-ended labeling. The elicited taxonomy is then used to rate robot attributes and to predict the best voice for an unseen robot. We offer a web interface to aid engineers in customizing robot voices, demonstrating the synergy between cognitive science and machine learning for engineering tools.
Pol van Rijn, Silvan Mertes, Kathrin Janowski, Katharina Weitz, Nori Jacoby, Elisabeth André
CHI6
2024 Explaining It Your Way - Findings from a Co-Creative Design Workshop on Designing XAI Applications with AI End-Users from the Public Sector
abstract
Human-Centered AI prioritizes end-users’ needs like transparency and usability. This is vital for applications that affect people’s everyday lives, such as social assessment tasks in the public sector. This paper discusses our pioneering effort to involve public sector AI users in XAI application design through a co-creative workshop with unemployment consultants from Estonia. The workshop’s objectives were identifying user needs and creating novel XAI interfaces for the used AI system. As a result of our user-centered design approach, consultants were able to develop AI interface prototypes that would support them in creating success stories for their clients by getting detailed feedback and suggestions. We present a discussion on the value of co-creative design methods with end-users working in the public sector to improve AI application design and provide a summary of recommendations for practitioners and researchers working on AI systems in the public sector.
Katharina Weitz, Ruben Schlagowski, Elisabeth André, Maris Männiste, Ceenu George
CHI3
2024 Bridging Skills and Scenarios: Initial Steps Towards Using Faded Worked Examples as Personalized Exercises in Vocational Education
abstract
In this paper, we present a method for generating faded worked examples as personalized exercises aimed at bridging the gap between knowledge of theoretical concepts and their application in the real world, which is particularly important in vocational education. Previous works suggest that faded worked examples are effective learning material that can also adapt to learners of different levels. Yet, there is no formulated method for automatically generating faded worked examples personalized to different learners in real-time. We develop a method for generating faded worked examples from scenarios, changing the faded positions and degree of fading based on the targeted skills and the learner’s proficiency level. We evaluate our method through a user study involving 13 computer science students from a German university, who practice specific computer networking skills. The results indicate significant improvement in the targeted skill over the untargeted one, highlighting the potenti al of our approach in vocational education settings. Our study is an early but promising step towards the future of personalized learning, paving the way for further research in adaptive and personalized vocational training.
Torben Soennecken, Anan Schütt, Björn Petrak, Elisabeth André
CSEDU (1)4
2024 From a Social POV: The Impact of Point of View on Player Behavior, Engagement, and Experience in a Serious Social Simulation Game
abstract
Multiplayer games with social aspects vary widely regarding client design, e.g., point of view or camera perspective. While design paradigms usually arise from gold standards that are set by previously successful games in the industry, the impact of such paradigms is under-researched for games that serve as scientific instruments, e.g., to research social behavior. Intending to investigate how such games should be designed, we built two multiplayer clients with the same game logic, one using a first-person point of view, while the other includes a top-down camera perspective. Then, we conducted an online user study in which players tested these game clients in extensive multiplayer sessions. Analyzing speech time, in-game logs, questionnaires, and qualitative feedback, we look at the perspectives’ impact on player behavior, engagement, and game experience in a scientific or "serious games" context. In addition, we have made our designed game UNISON and both clients available as open source to facilitate future empirical social science research.
Ruben Schlagowski, Frederick Herget, Niklas Heimerl, Maximilian Hammerl, Tobias Huber, Pamina Zwolsky, Jan Gruca, Elisabeth André
FDG8
2024 REACT 2024: the Second Multiple Appropriate Facial Reaction Generation Challenge
abstract
In dyadic interactions, humans communicate their intentions and state of mind using verbal and non-verbal cues, where multiple different facial reactions might be appropriate in response to a specific speaker behaviour. Then, how to develop a machine learning (ML) model that can automatically generate multiple appropriate, diverse, realistic and synchronised human facial reactions from an previously unseen speaker behaviour is a challenging task. Following the successful organisation of the first REACT challenge (REACT 2023), this edition of the challenge (REACT 2024) employs a subset used by the previous challenge, which contains segmented 30-secs dyadic interaction clips originally recorded as part of the NOXI and RECOLA datasets, encouraging participants to develop and benchmark Machine Learning (ML) models that can generate multiple appropriate facial reactions (including facial image sequences and their attributes) given an input conversational partner's stimulus under various dyadic video conference scenarios. This paper presents: (i) the guidelines of the REACT 2024 challenge; (ii) the dataset utilized in the challenge; and (iii) the performance of the baseline systems on the two proposed sub-challenges: Offline Multiple Appropriate Facial Reaction Generation and Online Multiple Appropriate Facial Reaction Generation, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2024.
Siyang Song, Micol Spitale, Cristina Palmero, Germán Barquero, Hengde Zhu, Sergio Escalera, Michel F. Valstar, Tobias Baur 0001, Fabien Ringeval, Elisabeth André, Hatice Gunes
FG11
2024 Exploring the Impact of Non-Verbal Virtual Agent Behavior on User Engagement in Argumentative Dialogues
abstract
Engaging in discussions that involve diverse perspectives and exchanging arguments on a controversial issue is a natural way for humans to form opinions. In this process, the way arguments are presented plays a crucial role in determining how engaged users are, whether the interaction takes place solely among humans or within human-agent teams. This is of great importance as user engagement plays a crucial role in determining the success or failure of cooperative argumentative discussions. One main goal is to maintain the user’s motivation to participate in a reflective opinion-building process, even when addressing contradicting viewpoints. This work investigates how non-verbal agent behavior, specifically co-speech gestures, influences the user’s engagement and interest during an ongoing argumentative interaction. The results of a laboratory study conducted with 56 participants demonstrate that the agent’s co-speech gestures have a substantial impact on user engagement and interest and the overall perception of the system. Therefore, this research offers valuable insights for the design of future cooperative argumentative virtual agents.
Annalena Aicher, Yuki Matsuda 0001, Keiichi Yasumoto, Wolfgang Minker, Elisabeth André, Stefan Ultes
HAI5
2024 Beyond Pretend-Reality Dualism: Frame Analysis of LLM-powered Role Play with Social Agents
abstract
Role-playing activities offer opportunities for developing individuals’ creativity, communication, and problem-solving skills. Recent advances in large language models (LLM) facilitate fluent conversations with machines. To investigate benefits and pitfalls of LLMs in a relatively unexplored context of human-agent role-play as a culturally contextualised activity, a dataset of twelve human-agent interactions produced by two researchers with two state-of-the-art LLMs was annotated based on a frame analysis scheme from literature. The pilot study shows that human-agent play has a similar complexity as human-human play in which players maintain identities of themselves, external observers and play characters simultaneously going beyond the pretend-reality dualism. Results suggest that, while the LLMs can maintain and shift between roles, they play some roles better than others, and display cultural and gender stereotypes. Additionally, the coding scheme shows potential to help identify LLM outputs that require embodied enactment, and to be used for LLM bench-marking for role-play.
Sviatlana Höhn, Jauwairia Nasir, Daniel Tozadore, Ali Paikan, Pouyan Ziafati, Elisabeth André
HAI6
2024 Towards Automated Annotation of Infant-Caregiver Engagement Phases with Multimodal Foundation Models
abstract
Caregiver mental health disorders increase the risk of insecure infant attachment and can negatively impact multiple aspects of child development, including cognitive, emotional, and social growth. Infant-caregiver interactions contain subtle psychological and behavioral cues that reveal these adverse effects, underscoring the need for analytical methods to assess them effectively. The Face-to-Face-Still-Face (FFSF) paradigm is a key approach in psychological research for investigating these dynamics, and the Infant and Caregiver Engagement Phases revised German edition (ICEP-R) annotation scheme provides a structured framework for evaluating FFSF interactions. However, manual annotation is labor-intensive and limits scalability, thus hindering a deeper understanding of early developmental impairments. To address this, we developed a computational method that automates the annotation of caregiver-infant interactions using features extracted from audio-visual foundational models. Our approach was tested on 92 FFSF video sessions. Findings demonstrate that models based on bidirectional LSTM and linear classifiers show varying effectiveness depending on the role and feature modality. Specifically, bidirectional LSTM models generally perform better in predicting complex infant engagement phases across multimodal features, while linear models show competitive performance, particularly with unimodal feature encodings like Wav2Vec2-BERT. To support further research, we share our raw feature dataset annotated with ICEP-R labels, enabling broader refinement of computational methods in this area.
Daksitha Withanage, Dominik Schiller, Tobias Hallmen, Silvan Mertes, Tobias Baur 0001, Florian Lingenfelser, Mitho Müller, Lea Kaubisch, Corinna Reck, Elisabeth André
ICMI10
2024 Multilingual Dyadic Interaction Corpus NoXi+J: Toward Understanding Asian-European Non-verbal Cultural Characteristics and their Influences on Engagement
abstract
Non-verbal behavior is a central challenge in understanding the dynamics of a conversation and the affective states between interlocutors arising from the interaction. Although psychological research has demonstrated that non-verbal behaviors vary across cultures, limited computational analysis has been conducted to clarify these differences and assess their impact on engagement recognition. To gain a greater understanding of engagement and non-verbal behaviors among a wide range of cultures and language spheres, in this study we conduct a multilingual computational analysis of non-verbal features and investigate their role in engagement and engagement prediction. To achieve this goal, we first expanded the NoXi dataset, which contains interaction data from participants living in France, Germany, and the United Kingdom, by collecting session data of dyadic conversations in Japanese and Chinese, resulting in the enhanced dataset NoXi+J. Next, we extracted multimodal non-verbal features, including speech acoustics, facial expressions, backchanneling and gestures, via various pattern recognition techniques and algorithms. Then, we conducted a statistical analysis of listening behaviors and backchannel patterns to identify culturally dependent and independent features in each language and common features among multiple languages. These features were also correlated with the engagement shown by the interlocutors. Finally, we analyzed the influence of cultural differences in the input features of LSTM models trained to predict engagement for five language datasets. A SHAP analysis combined with transfer learning confirmed a considerable correlation between the importance of input features for a language set and the significant cultural characteristics analyzed.
Marius Funk, Shogo Okada, Elisabeth André
ICMI3
2024 Stressor Type Matters! - Exploring Factors Influencing Cross-Dataset Generalizability of Physiological Stress Detection
abstract
Automatic stress detection using heart rate variability (HRV) features has gained significant traction as it utilizes unobtrusive wearable sensors measuring signals like electrocardiogram (ECG) or blood volume pulse (BVP). However, detecting stress through such physiological signals presents a considerable challenge owing to the variations in recorded signals influenced by factors, such as perceived stress intensity and measurement devices. Consequently, stress detection models developed on one dataset may perform poorly on unseen data collected under different conditions. To address this challenge, this study explores the generalizability of machine learning models trained on HRV features for binary stress detection. Our goal extends beyond evaluating generalization performance; we aim to identify the characteristics of datasets that have the most significant influence on generalizability. We leverage four publicly available stress datasets (WESAD, SWELL-KW, ForDigitStress, VerBIO) that vary in at least one of the characteristics such as stress elicitation techniques, stress intensity, and sensor devices. Employing a cross-dataset evaluation approach, we explore which of these characteristics strongly influence model generalizability. Our findings reveal a crucial factor affecting model generalizability: primary stressor. Models achieved good performance across datasets when the primary stressor (e.g., social evaluation in our case) remains consistent. Factors like stress intensity or brand of the measurement device had minimal impact on cross-dataset performance. Based on our findings, we recommend matching the primary stressor when deploying HRV-based stress models in new environments. Although previous works have performed cross-dataset evaluation of stress models, this is the first study to systematically investigate the factors influencing the cross-dataset applicability of HRV-based stress models. Our insights are crucial for scenarios with limited data, where techniques like domain generalization and domain adaptation may not be applicable.
Pooja Prajod, Bhargavi Mahesh, Elisabeth André
ICMI3
2024 Relevant Irrelevance: Generating Alterfactual Explanations for Image Classifiers
Silvan Mertes, Tobias Huber, Christina Karle, Katharina Weitz, Ruben Schlagowski, Cristina Conati, Elisabeth André
IJCAI7
2024 A Gaze into Argumentative Chatbots: Exploring the Influence of Challenger Arguments on Reflection and Attention
abstract
A natural way to resolve different points of view and form opinions is through exchanging arguments and knowledge. Facing the vast amount of available information on the internet, people tend to focus on information consistent with their beliefs. To support a fair and unbiased opinion-building process, we propose an intelligent agent in the form of a chatbot that engages in a deliberative dialogue with a human. In contrast to persuasive systems, the chatbot aims to provide a diverse and representative overview - embedded in a conversation with the user. To account for a reflective and unbiased exploration of the topic, we enable the system to intervene if the user is too focused on their pre-existing opinion. To achieve that, the agent employs a metric to assess the user’s focus on challenger arguments.
Klaus Weber 0001, Natalie Hogh, Cristina Conati, Elisabeth André
IVA4
2024 Does Difficulty even Matter? Investigating Difficulty Adjustment and Practice Behavior in an Open-Ended Learning Task
abstract
Difficulty adjustment in practice exercises has been shown to be beneficial for learning. However, previous research has mostly investigated close-ended tasks, which do not offer the students multiple ways to reach a valid solution. Contrary to this, in order to learn in an open-ended learning task, students need to effectively explore the solution space as there are multiple ways to reach a solution. For this reason, the effects of difficulty adjustment could be different for open-ended tasks. To investigate this, as our first contribution, we compare different methods of difficulty adjustment in a user study conducted with 86 participants. Furthermore, as the practice behavior of the students is expected to influence how well the students learn, we additionally look at their practice behavior as a post-hoc analysis. Therefore, as a second contribution, we identify different types of practice behavior and how they link to students’ learning outcomes and subjective evaluation measures as well as explore the influence the difficulty adjustment methods have on the practice behaviors. Our results suggest the usefulness of taking into account the practice behavior in addition to only using the practice performance to inform adaptive intervention and difficulty adjustment methods.
Anan Schütt, Tobias Huber, Jauwairia Nasir, Cristina Conati, Elisabeth André
LAK5
2024 MultiMediate'24: Multi-Domain Engagement Estimation
abstract
Estimating the momentary level of participant's engagement is an important prerequisite for assistive systems that support human interactions. Previous work has addressed this task in within-domain evaluation scenarios, i.e. training and testing on the same dataset. This is in contrast to real-life scenarios where domain shifts between training and testing data frequently occur. With MultiMediate'24, we present the first challenge addressing multi-domain engagement estimation. As training data, we utilise the NOXI database of dyadic novice-expert interactions. In addition to within-domain test data, we add two new test domains. First, we introduce recordings following the NOXI protocol but covering languages that are not present in the NOXI training data. Second, we collected novel engagement annotations on the MPIIGroupInteraction dataset which consists of group discussions between three to four people. In this way, MultiMediate'24 evaluates the ability of approaches to generalise across factors such as language and cultural background, group size, task, and screen-mediated vs. face-to-face interaction. This paper describes the MultiMediate'24 challenge and presents baseline results. In addition, we discuss selected challenge solutions.
Philipp Müller 0001, Michal Balazia, Tobias Baur 0001, Michael Dietz, Alexander Heimerl, Anna Penzkofer, Dominik Schiller, François Brémond, Jan Alexandersson, Elisabeth André, Andreas Bulling
ACM Multimedia10
2024 Evaluating Gender Ambiguity, Novelty and Anthropomorphism in Humming and Talking Voices for Robots
abstract
This paper investigates the effects of gender neutralization on the perception of anthropomorphism, gender specificity, and novelty for human voices, comparing spoken and hummed voice modalities. We evaluated gender-neutralized and original voice samples in both spoken and hummed formats using an online survey. Our results confirm that gender-neutralizing filters effectively reduce perceived gender specificity in both modalities, supporting their use in creating gender-neutral voices for humanoid robots. Hummed voices were perceived as more anthropomorphic and less novel than spoken voices, suggesting that non-verbal sound modalities can enhance the human likeness of gender-neutral androids while maintaining gender ambiguity. The study contributes to HRI by highlighting the potential of humming to fulfill users’ expectations of interaction with android robots.
Johanna Magdalena Kuch, Jauwairia Nasir, Silvan Mertes, Ruben Schlagowski, Christian Becker-Asano, Elisabeth André
RO-MAN6
2024 Approximating facial expression effects on diagnostic accuracy via generative AI in medical genetics
abstract
Artificial intelligence (AI) is increasingly used in genomics research and practice, and generative AI has garnered significant recent attention. In clinical applications of generative AI, aspects of the underlying datasets can impact results, and confounders should be studied and mitigated. One example involves the facial expressions of people with genetic conditions. Stereotypically, Williams (WS) and Angelman (AS) syndromes are associated with a "happy" demeanor, including a smiling expression. Clinical geneticists may be more likely to identify these conditions in images of smiling individuals. To study the impact of facial expression, we analyzed publicly available facial images of approximately 3500 individuals with genetic conditions. Using a deep learning (DL) image classifier, we found that WS and AS images with non-smiling expressions had significantly lower prediction probabilities for the correct syndrome labels than those with smiling expressions. This was not seen for 22q11.2 deletion and Noonan syndromes, which are not associated with a smiling expression. To further explore the effect of facial expressions, we computationally altered the facial expressions for these images. We trained HyperStyle, a GAN-inversion technique compatible with StyleGAN2, to determine the vector representations of our images. Then, following the concept of InterfaceGAN, we edited these vectors to recreate the original images in a phenotypically accurate way but with a different facial expression. Through online surveys and an eye-tracking experiment, we examined how altered facial expressions affect the performance of human experts. We overall found that facial expression is associated with diagnostic accuracy variably in different genetic conditions.
Tanviben Patel, Amna A. Othman, Ömer Sümer, Fabio Hellmann, Peter M. Krawitz, Elisabeth André, Molly E. Ripper, Chris Fortney, Susan Persky, Cedrik Tekendo-Ngongang, Suzanna E. Ledgister Hanchard, Kendall A. Flaharty, Rebekah L. Waikel, Dat Duong, Benjamin D. Solomon
Bioinform.6
2024 COLD Fusion: Calibrated and Ordinal Latent Distribution Fusion for Uncertainty-Aware Multimodal Emotion Recognition
abstract
Automatically recognising apparent emotions from face and voice is hard, in part because of various sources of uncertainty, including in the input data and the labels used in a machine learning framework. This paper introduces an uncertainty-aware multimodal fusion approach that quantifies modality-wise aleatoric or data uncertainty towards emotion prediction. We propose a novel fusion framework, in which latent distributions over unimodal temporal context are learned by constraining their variance. These variance constraints, Calibration and Ordinal Ranking, are designed such that the variance estimated for a modality can represent how informative the temporal context of that modality is w.r.t. emotion recognition. When well-calibrated, modality-wise uncertainty scores indicate how much their corresponding predictions are likely to differ from the ground truth labels. Well-ranked uncertainty scores allow the ordinal ranking of different frames across different modalities. To jointly impose both these constraints, we propose a softmax distributional matching loss. Our evaluation on AVEC 2019 CES, CMU-MOSEI, and IEMOCAP datasets shows that the proposed multimodal fusion method not only improves the generalisation performance of emotion recognition models and their predictive uncertainty estimates, but also makes the models robust to novel noise patterns encountered at test time.
Mani Kumar Tellamekala, Shahin Amiriparian, Björn W. Schuller, Elisabeth André, Timo Giesbrecht, Michel F. Valstar
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 The Deep Method: Towards Computational Modeling of the Social Emotion Shame Driven by Theory, Introspection, and Social Signals
abstract
Understanding emotions is key to Affective Computing. Emotion recognition focuses on the communicative component of emotions encoded in social signals. This view alone is insufficient for a deeper understanding and computational representation of the internal, subjectively experienced component of emotions. This paper presents a cognition-based method calledDeepas a starting point for deeper computational modeling of the internal component of emotions.Deepincorporates an approach to query individual internal emotional experiences and to represent such information computationally. It combines social signals, verbalized introspection information, context information, and theory-driven knowledge. We apply theDeepmethod to the emotion of shame as an example and compare it to a typical emotion recognition model, highlighting the differences and advantages.
Tanja Schneeberger, Mirella Hladký, Ann-Kristin Thurner, Jana Volkert, Alexander Heimerl, Tobias Baur 0001, Elisabeth André, Patrick Gebhard
IEEE Trans. Affect. Comput.7
2024 Are 3D Face Shapes Expressive Enough for Recognising Continuous Emotions and Action Unit Intensities?
abstract
Recognising continuous emotions and action unit (AU) intensities from face videos, requires a spatial and temporal understanding of expression dynamics. Existing works primarily rely on 2D face appearance features to extract such dynamics. This work focuses on a promising alternative based on parametric 3D face alignment models, which disentangle different factors of variation, including expression-induced shape variations. We aim to understand how expressive 3D face shapes are in estimating valence-arousal and AU intensities compared to the state-of-the-art 2D appearance-based models. We benchmark five recent 3D face models: ExpNet, 3DDFA-V2, RingNet, DECA, and EMOCA. In valence-arousal estimation, expression features of 3D face models consistently surpassed previous works and yielded an average concordance correlation of. 745 and. 574 on SEWA and AVEC 2019 CES corpora, respectively. We also study how 3D face shapes performed on AU intensity estimation on BP4D and DISFA datasets, and report that 3D face features were on par with 2D appearance features in recognising AUs 4, 6, 10, 12, and 25, but not the entire set of AUs. To understand this discrepancy, we conduct a correspondence analysis between valence-arousal and AUs, which points out that accurate prediction of valence-arousal may require the knowledge of only a few AUs.
Mani Kumar Tellamekala, Ömer Sümer, Björn W. Schuller, Elisabeth André, Timo Giesbrecht, Michel F. Valstar
IEEE Trans. Affect. Comput.4
2023 Wish You Were Here: Mental and Physiological Effects of Remote Music Collaboration in Mixed Reality
abstract
With face-to-face music collaboration being severely limited during the recent pandemic, mixed reality technologies and their potential to provide musicians a feeling of "being there" with their musical partner can offer tremendous opportunities. In order to assess this potential, we conducted a laboratory study in which musicians made music together in real-time while simultaneously seeing their jamming partner’s mixed reality point cloud via a head-mounted display and compared mental effects such as flow, affect, and co-presence to an audio-only baseline. In addition, we tracked the musicians’ physiological signals and evaluated their features during times of self-reported flow. For users jamming in mixed reality, we observed a significant increase in co-presence. Regardless of the condition (mixed reality or audio-only), we observed an increase in positive affect after jamming remotely. Furthermore, we identified heart rate and HF/LF as promising features for classifying the flow state musicians experienced while making music together.
Ruben Schlagowski, Dariia Nazarenko, Yekta Said Can, Kunal Gupta, Silvan Mertes, Mark Billinghurst, Elisabeth André
CHI7
2023 Around the world in 60 words: A generative vocabulary test for online research
Pol van Rijn, Harin Lee, Raja Marjieh, Ilia Sucholutsky, Francesca Lanzarini, Elisabeth André, Nori Jacoby
CogSci7
2023 Fast Dynamic Difficulty Adjustment for Intelligent Tutoring Systems with Small Datasets
Anan Schütt, Tobias Huber, Ilhan Aslan, Elisabeth André
EDM4
2023 Social Signals as a Facilitator of Human-Robot Interaction
abstract
The automatic analysis and synthesis of social signals, including voice, gestures, and facial expressions, are pivotal for the advancement of next-generation interfaces, facilitating more intuitive and natural human-computer interactions with both robots and virtual agents. During my presentation, I will introduce computational methodologies for implementing socially interactive behaviors in artificial agents, with a specific focus on three essential components: Social Perception, Socially-Aware Behavior Synthesis, and Learning Socially-Aware Behaviors. In addition to discussing analytic methods grounded in cognitive and social science theories, I will explore empirical approaches that empower artificial agents to learn socially interactive behaviors from recordings of human-human interactions or real-life engagements with human interlocutors. I will also delve into the potential and challenges arising from neural behavior generation techniques, promising to elevate virtual agents and social robots to new levels of human-likeness. Throughout the presentation, I will offer practical insights and examples drawn from our work across various application fields. To benefit users, we need to extend our focus beyond technical solutions to encompass ethical, legal, and societal considerations.
Elisabeth André
HAI1
2023 The Influence of Avatar Interfaces on Argumentative Dialogues
abstract
Humans form opinions and justify different points of view by exchanging arguments and knowledge. Likewise to human-human interaction, the way arguments are presented influence the user's willingness to engage into a critical reflection. Especially when interacting with conversational agents the user's engagement and motivation are important factors and highly influence the success or failure of such a mixed team. To maintain the users' trust and satisfaction, the users' perception of the respective system is an important indicator. Thus, this work investigates the design of a cooperative argumentative dialogue system using a virtual avatar compared to a non-avatar interface by evaluating a crowdsourcing study conducted with 84 participants. The results indicate, that the avatar system is perceived as significantly more appealing and natural and thus, engaging which also influences the acceptance and perception of the quality of presented arguments. Furthermore, we found that the presence of the avatar often led to an increase in the anticipated level of conversational proficiency similar to that of a human interlocutor. Therefore, this work provides important insights for the design of future cooperative argumentative virtual avatar interfaces.
Annalena Aicher, Klaus Weber 0001, Elisabeth André, Wolfgang Minker, Stefan Ultes
IVA3
2023 Socially Interactive Agents as Cobot Avatars: Developing a Model to Support Flow Experiences and Weil-Being in the Workplace
abstract
This study evaluates a socially interactive agent to create an embodied cobot. It tests a real-time continuous emotional modeling method and an aligned transparent behavioral model, BASSF (boredom, anxiety, self-efficacy, self-compassion, flow). The BASSF model anticipates and counteracts counterproductive emotional experiences of operators working under stress with cobots on tedious tasks. The flow experience is represented in the three-dimensional pleasure, arousal, and dominance (PAD) space. The embodied covatar (cobot and avatar) is introduced to support flow experiences through emotion regulation guidance. The study tests the model's main theoretical assumptions about flow, dominance, self-efficacy, and boredom. Twenty participants worked on a task for an hour, assembling pieces in collaboration with the covatar. After the task, participants completed questionnaires on flow, their affective experience, and self-efficacy, and they were interviewed to understand their emotions and regulation during the task. The results suggest that the dominance dimension plays a vital role in task-related settings as it predicts the participants' self-efficacy and flow. However, the relationship between flow, pleasure, and arousal requires further investigation. Qualitative interview analysis revealed that participants regulated negative emotions, like boredom, also without support, but some strategies could negatively impact well-being and productivity, which aligns with theory.
Sebastian Beyrodt, Matteo Lavit Nicora, Fabrizio Nunnari, Lara Chehayeb, Pooja Prajod, Tanja Schneeberger, Elisabeth André, Matteo Malosio, Patrick Gebhard, Dimitra Tsovaltzi
IVA7
2023 Multimodal Irony for Virtual Characters
abstract
Humor is an important communicative skill in human interactions. Intelligent virtual agents can leverage it to increase their believability and overall interaction experience. In this paper, we focus on transferring and implementing existing multimodal irony markers from the literature to a photorealistic virtual character. The verbal content is generated dynamically by an irony generator. We demonstrate how the ironic turn can be augmented with prosodic and facial markers. An expressivity parameter allows us to manipulate the encoding of the irony style.
Thomas Kiderle, Hannes Ritschel, Silvan Mertes, Elisabeth André
IVA4
2023 The Affective Bar Piano
abstract
Music is a great way of supporting a story. It adds a new layer of affective information and as such substantially increases the listening experience in storytelling scenarios. However, in real-time settings, creating emotionally fitting music requires permanent adaptation to the story's mood. While methods to compose and modify music according to emotional states are widely explored, current research rarely uses those techniques in a real-time setting, where such accompanying background music still requires improvisation by human musicians. In this work, we introduce the Affective Bar Piano, a virtual agent that assesses the mood of a story in real time. At the same time, the agent adapts its play to mirror the sensed affect of a human storyteller. In the presented demonstration scenario, the virtual agent is embodied by a 3D piano character playing music in a Wild West saloon setting.
Hannes Ritschel, Silvan Mertes, Florian Lingenfelser, Thomas Kiderle, Elisabeth André
IVA5
2023 MultiMediate '23: Engagement Estimation and Bodily Behaviour Recognition in Social Interactions
abstract
Automatic analysis of human behaviour is a fundamental prerequisite for the creation of machines that can effectively interact with- and support humans in social interactions. In MultiMediate'23, we address two key human social behaviour analysis tasks for the first time in a controlled challenge: engagement estimation and bodily behaviour recognition in social interactions. This paper describes the MultiMediate'23 challenge and presents novel sets of annotations for both tasks. For engagement estimation we collected novel annotations on the NOvice eXpert Interaction (NOXI) database. For bodily behaviour recognition, we annotated test recordings of the MPIIGroupInteraction corpus with the BBSI annotation scheme. In addition, we present baseline results for both challenge tasks.
Philipp Müller 0001, Michal Balazia, Tobias Baur 0001, Michael Dietz, Alexander Heimerl, Dominik Schiller, Mohammed Guermal, Dominike Thomas, François Brémond, Jan Alexandersson, Elisabeth André, Andreas Bulling
ACM Multimedia11
2023 REACT2023: The First Multiple Appropriate Facial Reaction Generation Challenge
abstract
The Multiple Appropriate Facial Reaction Generation Challenge (REACT2023) is the first competition event focused on evaluating multimedia processing and machine learning techniques for generating human-appropriate facial reactions in various dyadic interaction scenarios, with all participants competing strictly under the same conditions. The goal of the challenge is to provide the first benchmark test set for multi-modal information processing and to foster collaboration among the audio, visual, and audio-visual behaviour analysis and behaviour generation (a.k.a generative AI) communities, to compare the relative merits of the approaches to automatic appropriate facial reaction generation under different spontaneous dyadic interaction conditions. This paper presents: (i) the novelties, contributions and guidelines of the REACT2023 challenge; (ii) the dataset utilized in the challenge; and (iii) the performance of the baseline systems on the two proposed sub-challenges: Offline Multiple Appropriate Facial Reaction Generation and Online Multiple Appropriate Facial Reaction Generation, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2023.
Siyang Song, Micol Spitale, Germán Barquero, Cristina Palmero, Sergio Escalera, Michel F. Valstar, Tobias Baur 0001, Fabien Ringeval, Elisabeth André, Hatice Gunes
ACM Multimedia10
2023 Improving Deep Facial Phenotyping for Ultra-rare Disorder Verification Using Model Ensembles
abstract
Rare genetic disorders affect more than 6% of the global population. Reaching a diagnosis is challenging because rare disorders are very diverse. Many disorders have recognizable facial features that are hints for clinicians to diagnose patients. Previous work, such as GestaltMatcher, utilized representation vectors produced by a DCNN similar to AlexNet to match patients in high-dimensional feature space to support "unseen" ultra-rare disorders. However, the architecture and dataset used for transfer learning in GestaltMatcher have become outdated. Moreover, a way to train the model for generating better representation vectors for unseen ultra-rare disorders has not yet been studied. Because of the overall scarcity of patients with ultra-rare disorders, it is infeasible to directly train a model on them. Therefore, we first analyzed the influence of replacing GestaltMatcher DCNN with a state-of-the-art face recognition approach, iResNet with ArcFace. Additionally, we experimented with different face recognition datasets for transfer learning. Furthermore, we proposed test-time augmentation, and model ensembles that mix general face verification models and models specific for verifying disorders to improve the disorder verification accuracy of unseen ultra-rare disorders. Our proposed ensemble model achieves state-of-the-art performance on both seen and unseen disorders. Code is available at github.com/igsb/GestaltMatcher-Arc.
Alexander Hustinx, Fabio Hellmann, Ömer Sümer, Behnam Javanmardi, Elisabeth André, Peter M. Krawitz, Tzung-Chien Hsieh
WACV5
2023 Approaches, Applications, and Challenges in Physiological Emotion Recognition - A Tutorial Overview
abstract
An automatic emotion recognition system can serve as a fundamental framework for various applications in daily life from monitoring emotional well-being to improving the quality of life through better emotion regulation. Understanding the process of emotion manifestation becomes crucial for building emotion recognition systems. An emotional experience results in changes not only in interpersonal behavior but also in physiological responses. Physiological signals are one of the most reliable means for recognizing emotions since individuals cannot consciously manipulate them for a long duration. These signals can be captured by medical-grade wearable devices, as well as commercial smart watches and smart bands. With the shift in research direction from laboratory to unrestricted daily life, commercial devices have been employed ubiquitously. However, this shift has introduced several challenges, such as low data quality, dependency on subjective self-reports, unlimited movement-related changes, and artifacts in physiological signals. This tutorial provides an overview of practical aspects of emotion recognition, such as experiment design, properties of different physiological modalities, existing datasets, suitable machine learning algorithms for physiological data, and several applications. It aims to provide the necessary psychological and physiological backgrounds through various emotion theories and the physiological manifestation of emotions, thereby laying a foundation for emotion recognition. Finally, the tutorial discusses open research directions and possible solutions.
Yekta Said Can, Bhargavi Mahesh, Elisabeth André
Proc. IEEE3
2023 An Overview of Affective Speech Synthesis and Conversion in the Deep Learning Era
abstract
Speech is the fundamental mode of human communication, and its synthesis has long been a core priority in human–computer interaction research. In recent years, machines have managed to master the art of generating speech that is understandable by humans. However, the linguistic content of an utterance encompasses only a part of its meaning. Affect, or expressivity, has the capacity to turn speech into a medium capable of conveying intimate thoughts, feelings, and emotions—aspects that are essential for engaging and naturalistic interpersonal communication. While the goal of imparting expressivity to synthesized utterances has so far remained elusive, following recent advances in text-to-speech synthesis, a paradigm shift is well under way in the fields of affective speech synthesis and conversion as well. Deep learning, as the technology that underlies most of the recent advances in artificial intelligence, is spearheading these efforts. In this overview, we outline ongoing trends and summarize state-of-the-art approaches in an attempt to provide a broad overview of this exciting field.
Andreas Triantafyllopoulos, Björn W. Schuller, Gökçe Iymen, Tevfik Metin Sezgin, Xiangheng He, Zijiang Yang 0007, Panagiotis Tzirakis, Shuo Liu 0012, Silvan Mertes, Elisabeth André, Ruibo Fu, Jianhua Tao 0001
Proc. IEEE10
2023 Editorial Transactions on Affective Computing-News on the Journal
abstract
I N THE past year, we continued the successful cooperation with the International Conference on Affective Computing and Intelligent Interaction (ACII).ACII 2022 was held as a hybrid event in Nara, Japan, from October 18th to October 21st, 2022.One highlight was a dedicated ACII session, "TAFFC Best Paper Presentation Awards," where the Best Papers published in 2021 were presented.The winners were selected from 82 papers published in issues 12(1)-12(4) based on a vote by the Associate Editors.We warmly congratulate the authors of the following papers (in no specific order):1)
Elisabeth André
IEEE Trans. Affect. Comput.1
2022 Generating Personalized Behavioral Feedback for a Virtual Job Interview Training System Through Adversarial Learning
Alexander Heimerl, Silvan Mertes, Tanja Schneeberger, Tobias Baur 0001, Ailin Liu, Linda Becker, Nicolas Rohleder, Patrick Gebhard, Elisabeth André
AIED (1)9
2022 Flow with the Beat! Human-Centered Design of Virtual Environments for Musical Creativity Support in VR
abstract
As previous studies have shown, the environment of creative people can have a significant impact on their creative process and thus on their creations. However, with the advent of digital tools such as virtual instruments and digital audio workstations, more and more creative work is digital and decoupled from the creator’s environment. Virtual Reality technologies open up new possibilities here, as creative tools can seamlessly merge with any virtual environment the user finds himself in. This paper reports on the human-centered design process of a VR application that aims at supporting the user’s individual needs to support their creativity while composing percussive beats in virtual environments. For this purpose, we derived factors that influence creativity from literature and conducted focus group interviews in order to learn how virtual environments and 3DUI can be designed for creativity support. In a subsequent laboratory study, we let users interact with a virtual step sequencer UI in virtual environments that were either customizable or fixed/unchangeable. By analyzing post-test ratings from music experts, self-report questionnaires, and user behavior data, we examined the effects of such customizable virtual environments on user creativity, user experience, flow, and subjective creativity support scales. While we did not observe a significant impact of this independent variable on user creativity, user experience or flow, we found that users had specific individual needs regarding their virtual surroundings and strongly preferred customizable virtual environments, even though the fixed virtual environment was designed to be creatively stimulating. We also observed consistently high flow and user experience ratings, which promote human-centered design of VR-based creativity support tools in a musical context.
Ruben Schlagowski, Fabian Wildgrube, Silvan Mertes, Ceenu George, Elisabeth André
Creativity & Cognition5
2022 On the Generalizability of ECG-based Stress Detection Models
abstract
Stress is prevalent in many aspects of everyday life including work, healthcare, and social interactions. Many works have studied handcrafted features from various bio-signals that are indicators of stress. Recently, deep learning models have also been proposed to detect stress. Typically, stress models are trained and validated on the same dataset, often involving one stressful scenario. However, it is not practical to collect stress data for every scenario. So, it is crucial to study the generalizability of these models and determine to what extent they can be used in other scenarios. In this paper, we explore the generalization capabilities of Electrocardiogram (ECG)-based deep learning models and models based on handcrafted ECG features, i.e., Heart Rate Variability (HRV) features. To this end, we train three HRV models and two deep learning models that use ECG signals as input. We use ECG signals from two popular stress datasets WESAD and SWELL-KW - differing in terms of stressors and recording devices. First, we evaluate the models using leave-one-subject-out (LOSO) cross-validation using training and validation samples from the same dataset. Next, we perform a cross-dataset validation of the models, that is, LOSO models trained on the WESAD dataset are validated using SWELL-KW samples and vice versa. While deep learning models achieve the best results on the same dataset, models based on HRV features considerably outperform them on data from a different dataset. This trend is observed for all the models on both datasets. Therefore, HRV models are a better choice for stress recognition in applications that are different from the dataset scenario. To the best of our knowledge, this is the first work to compare the cross-dataset generalizability between ECG-based deep learning models and HRV models.
Pooja Prajod, Elisabeth André
ICMLA2
2022 Local and Global Explanations of Agent Behavior: Integrating Strategy Summaries with Saliency Maps (Extended Abstract)
abstract
With advances in reinforcement learning (RL), agents are now being developed in high-stakes application domains such as healthcare and transportation. Explaining the behavior of these agents is challenging, as they act in large state spaces, and their decision-making can be affected by delayed rewards. In this paper, we explore a combination of explanations that attempt to convey the global behavior of the agent and local explanations which provide information regarding the agent's decision-making in a particular state. Specifically, we augment strategy summaries that demonstrate the agent's actions in a range of states with saliency maps highlighting the information it attends to. Our user study shows that intelligently choosing what states to include in the summary (global information) results in an improved analysis of the agents. We find mixed results with respect to augmenting summaries with saliency maps (local information).
Tobias Huber, Katharina Weitz, Elisabeth André, Ofra Amir
IJCAI3
2022 VoiceMe: Personalized voice generation in TTS
Pol van Rijn, Silvan Mertes, Dominik Schiller, Piotr Dura, Hubert Siuzdak, Peter M. C. Harrison, Elisabeth André, Nori Jacoby
INTERSPEECH7
2022 MultiMediate'22: Backchannel Detection and Agreement Estimation in Group Interactions
abstract
Backchannels, i.e. short interjections of the listener, serve important meta-conversational purposes like signifying attention or indicating agreement. Despite their key role, automatic analysis of backchannels in group interactions has been largely neglected so far. The MultiMediate challenge addresses, for the first time, the tasks of backchannel detection and agreement estimation from backchannels in group conversations. This paper describes the MultiMediate challenge and presents a novel set of annotations consisting of 7234 backchannel instances for the MPIIGroup Interaction dataset. Each backchannel was additionally annotated with the extent by which it expresses agreement towards the current speaker. In addition to a an analysis of the collected annotations, we present baseline results for both challenge tasks.
Philipp Müller 0001, Michael Dietz, Dominik Schiller, Dominike Thomas, Hali Lindsay, Patrick Gebhard, Elisabeth André, Andreas Bulling
ACM Multimedia7
2022 Using Explainable AI to Identify Differences Between Clinical and Experimental Pain Detection Models Based on Facial Expressions
Pooja Prajod, Tobias Huber, Elisabeth André
MMM (1)3
2022 Editorial: Transactions on Affective Computing - Another Year in the Shade of Covid-19
Elisabeth André
IEEE Trans. Affect. Comput.1
2022 Unraveling ML Models of Emotion With NOVA: Multi-Level Explainable AI for Non-Experts
abstract
In this article, we introduce a next-generation annotation tool calledNOVAfor emotional behaviour analysis, which implements a workflow that interactively incorporates the ‘human in the loop’. A main aspect of NOVA is the possibility of applying semi-supervised active learning where Machine Learning techniques are used already during the annotation process by giving the possibility to pre-label data automatically. Furthermore, NOVA implements recent eXplainable AI (XAI) techniques to provide users with both, a confidence value of the automatically predicted annotations, as well as visual explanations. We investigate how such techniques can assist non-experts in terms of trust, perceived self-efficacy, cognitive workload as well as creating correct mental models about the system by conducting a user study with 53 participants. The results show that NOVA can easily be used by non-experts and lead to a high computer self-efficacy. Furthermore, the results indicate that XAI visualisations help users to create more correct mental models about the machine learning system compared to the baseline condition. Nevertheless, we suggest that explanations in the field of AI have to be more focused on user-needs as well as on the classification task and the model they want to explain.
Alexander Heimerl, Katharina Weitz, Tobias Baur 0001, Elisabeth André
IEEE Trans. Affect. Comput.4
2021 Towards a Deeper Modeling of Emotions: The Deep Method and its Application on Shame
abstract
Understanding emotions is key to Affective Computing. Emotion recognition focuses on the communicative component of emotions encoded in social signals. This view alone is insufficient for deeper understanding and computational representation of the internal, subjectively experienced component of emotions. This paper presents the Deep method as a starting point for a deeper computational modeling of internal emotions. The method includes how to query individual internal emotional experiences, and it shows an approach to represent such information computationally. It combines social signals, verbalized introspection information, context information, and theory-driven knowledge. We apply the Deep method exemplary on the emotion shame and present a schematic dynamic Bayesian network for modeling it.
Tanja Schneeberger, Mirella Hladký, Ann-Kristin Thurner, Jana Volkert, Alexander Heimerl, Tobias Baur 0001, Elisabeth André, Patrick Gebhard
ACII7
2021 "It's our fault!": Insights Into Users' Understanding and Interaction With an Explanatory Collaborative Dialog System
abstract
Human-AI collaboration, a long standing goal in AI, refers to a partnership where a human and artificial intelligence work together towards a shared goal. Collaborative dialog allows human-AI teams to communicate and leverage strengths from both partners. To design collaborative dialog systems, it is important to understand what mental models users form about their AI-dialog partners, however, how users perceive these systems is not fully understood. In this study, we designed a novel, collaborative, communication-based puzzle game and explanatory dialog system. We created a public corpus from 117 conversations and post-surveys and used this to analyze what mental models users formed. Key takeaways include: Even when users were not engaged in the game, they perceived the AI-dialog partner as intelligent and likeable, implying they saw it as a partner separate from the game. This was further supported by users often overestimating the system’s abilities and projecting human-like attributes which led to miscommunications. We conclude that creating shared mental models between users and AI systems is important to achieving successful dialogs. We propose that our insights on mental models and miscommunication, the game, and our corpus provide useful tools for designing collaborative dialog systems.
Katharina Weitz, Lindsey Vanderlyn, Ngoc Thang Vu, Elisabeth André
CoNLL4
2021 "An Error Occurred!" - Trust Repair With Virtual Robot Using Levels of Mistake Explanation
abstract
Human-robot collaboration in industrial settings is an expanding research field in robotics. When working together, robot mistakes are an important factor to decrease trust and therefore interferes with cooperation. It is unclear whether explanations help to restore human-robot trust after a mistake. In our study, we investigate whether system explanations as a trust-repairing action after a robot makes a mistake in a collaborative task is helpful. Our pilot study revealed that users are more interested in solutions to errors than they are in just why the error happened. Therefore, in our main study, we evaluated three levels of mistake explanations (no explanation, explanation, and explanation with solution) after a robot in VR made a mistake in executing a shared objective. After testing with 30 participants we found that the robot making a mistake significantly affects trust toward the robot, compared to it completing the task successfully. While participants found the explanations helpful to trust or distrust the robot, the levels of the explanation did not lead to an increase in trust towards the robot after a mistake. In addition, we found no significant impact of explanations on self-efficacy and the emotional state of the participants. Our results show that explanations alone are not sufficient to increase human-computer trust after robot mistakes.
Kasper Hald, Katharina Weitz, Elisabeth André, Matthias Rehm
HAI3
2021 Socially Interactive Artificial Intelligence: Past, Present and Future
abstract
Socially interactive artificial agents are no longer mere fiction. For many, they are already part of everyday life. Due to technical advances in multimodal behavior analysis and synthesis, the asymmetry of communication between machines and humans is dissolving. Consequently, the interaction with robots and virtual characters has become more intuitive and natural, particularly for everyday users. Nevertheless, there is still some work to be done until artificial agents are able to smoothly interact with people over more extended periods in their homes and to cope with unforeseen situations.
Elisabeth André
ICMI1
2021 A Prototypical Network Approach for Evaluating Generated Emotional Speech
abstract
The collection of emotional speech data is a time-consuming and costly endeavour.Generative networks can be applied to augment the limited audio data artificially.However, it is challenging to evaluate generated audio for its similarity to source data, as current quantitative metrics are not necessarily suited to the audio domain.We explore the use of a prototypical network to evaluate four classes of generated emotional audio with this in mind.We first extract spectrogram images from WAVEGAN generated audio and other audio augmentation approaches, comparing similarity to the class prototype and diversity within the embedding space.Furthermore, we augment the source training set with each augmentation type and perform a classification to explore the generated audio plausibility.Results suggest that quality and diversity can be quantitatively observed with this approach.In the chosen context, we see that WAVEGAN generated data is recognisable as a source data class (F1-score 43.6 %), and the samples add similar diversity as unseen source data.This result leads to more plausible data for augmentation of the source training set -achieving up to 63.9 % F1 which is a 3.5 % improvement over the source data baseline.
Alice Baird, Silvan Mertes, Manuel Milling, Lukas Stappen, Thomas Wiest, Elisabeth André, Björn W. Schuller
Interspeech6
2021 Exploring Emotional Prototypes in a High Dimensional TTS Latent Space
abstract
Recent TTS systems are able to generate prosodically varied and realistic speech. However, it is unclear how this prosodic variation contributes to the perception of speakers' emotional states. Here we use the recent psychological paradigm 'Gibbs Sampling with People' to search the prosodic latent space in a trained GST Tacotron model to explore prototypes of emotional prosody. Participants are recruited online and collectively manipulate the latent space of the generative speech model in a sequentially adaptive way so that the stimulus presented to one group of participants is determined by the response of the previous groups. We demonstrate that (1) particular regions of the model's latent space are reliably associated with particular emotions, (2) the resulting emotional prototypes are well-recognized by a separate group of human raters, and (3) these emotional prototypes can be effectively transferred to new sentences. Collectively, these experiments demonstrate a novel approach to the understanding of emotional speech by providing a tool to explore the relation between the latent space of generative models and human semantics.
Pol van Rijn, Silvan Mertes, Dominik Schiller, Peter M. C. Harrison, Pauline Larrouy-Maestri, Elisabeth André, Nori Jacoby
Interspeech6
2021 Analysis by Synthesis: Using an Expressive TTS Model as Feature Extractor for Paralinguistic Speech Classification
abstract
Modeling adequate features of speech prosody is one key factor to good performance in affective speech classification.However, the distinction between the prosody that is induced by 'how' something is said (i.e., affective prosody) and the prosody that is induced by 'what' is being said (i.e., linguistic prosody) is neglected in state-of-the-art feature extraction systems.This results in high variability of the calculated feature values for different sentences that are spoken with the same affective intent, which might negatively impact the performance of the classification.While this distinction between different prosody types is mostly neglected in affective speech recognition, it is explicitly modeled in expressive speech synthesis to create controlled prosodic variation.In this work, we use the expressive Text-To-Speech model Global Style Token Tacotron to extract features for a speech analysis task.We show that the learned prosodic representations outperform state-of-the-art feature extraction systems in the exemplary use case of Escalation Level Classification.
Dominik Schiller, Silvan Mertes, Pol van Rijn, Elisabeth André
Interspeech4
2021 MultiMediate: Multi-modal Group Behaviour Analysis for Artificial Mediation
abstract
Artificial mediators are promising to support human group conversations but at present their abilities are limited by insufficient progress in group behaviour analysis. The MultiMediate challenge addresses, for the first time, two fundamental group behaviour analysis tasks in well-defined conditions: eye contact detection and next speaker prediction. For training and evaluation, MultiMediate makes use of the MPIIGroup Interaction dataset consisting of 22 three- to four-person discussions as well as of an unpublished test set of six additional discussions. This paper describes the MultiMediate challenge and presents the challenge dataset including novel fine-grained speaking annotations that were collected for the purpose of MultiMediate. Furthermore, we present baseline approaches and ablation studies for both challenge tasks
Philipp Müller 0001, Michael Dietz, Dominik Schiller, Dominike Thomas, Patrick Gebhard, Elisabeth André, Andreas Bulling
ACM Multimedia7
2021 Social Signals and Multimedia: Past, Present, Future
abstract
The rising popularity of Artificial Intelligence (AI) has brought considerable public interest as well faster and more direct transfer of research ideas into practice. One of the aspects of AI that still trails behind considerably is the role of machines in interpreting, enhancing, modeling, generating, and influencing social behavior. Such behavior is captured as social signals, usually by sensors recording multiple modalities, making it classic multimedia data. Such behavior can also be generated by an AI system when interacting with humans. Using AI techniques in combination with multimedia data can be used to pursue multiple goals, two of which are high-lighted here. First, supporting people during social interactions and helping them to fulfil their social needs either actively or passively.Second, improving our understanding of how people collaborate, build relationships, and process self identity. Despite the rise of fields such as Social Signal Processing, a similar panel organised at ACM Multimedia 2014, and an area on social and emotional signal sat the ACM MM since 2014, we argue that we have yet to truly fulfil the potential of the combining social signals and multimedia. This panel asks where we have come far enough and what remaining challenges there are in light of recent global events.
Hayley Hung, Cathal Gurrin, Martha A. Larson, Hatice Gunes, Fabien Ringeval, Elisabeth André, Louis-Philippe Morency
ACM Multimedia6
2021 A human-driven control architecture for promoting good mental health in collaborative robot scenarios
abstract
This paper introduces the control architecture of a platform aimed at promoting good mental health for workers interacting with collaborative robots (cobots). The platform aim is to render industrial production cells capable of automatically adapting their behavior in order to improve the operator’s quality of experience and level of engagement and to minimize his/her psychological strain. In order to achieve such a goal, an extremely rich and complex framework is required. Starting from the identification of the parameters that could influence the collaboration experience, the envisioned human- driven control structure is presented together with a detailed description of the components required to implement such an automated system. Future works will include proper tuning of control parameters with dedicated experimental sessions, together with the definition of organizational and technical guidelines for the design of a mental-health-friendly cobot-based manufacturing workplace.
Matteo Lavit Nicora, Elisabeth André, Daniel Berkmans, Claudia Carissoli, Tiziana D'Orazio, Antonella Delle Fave, Patrick Gebhard, Roberto Marani, Robert Mihai Mira, Luca Negri, Fabrizio Nunnari, Alberto Peña Fernández, Alessandro Scano, Gianluigi Reni, Matteo Malosio
RO-MAN2
2021 Do You Mind if I Pass Through? Studying the Appropriate Robot Behavior when Traversing two Conversing People in a Hallway Setting
abstract
Several works highlight how robots can navigate in a socially-aware manner by respecting and avoiding people’s personal spaces. But how should the robot act when there is no way around a group of persons? In this work, we explore this question by comparing three different ways to cross two conversing people in a hallway environment. In an online study with 135 participants, users rated the robot’s behavior on several items such as "social adequacy" or how "disturbing" it was. The three versions differ in the type of contact intention, i.e., no contact, nonverbal contact, and a combination of nonverbal and verbal contact. The results show that, on the one hand, users expect social behavior from the robot, so that they can anticipate its behavior, but on the other hand, they want it to be as little disruptive as possible.
Björn Petrak, Gundula Sopper, Katharina Weitz, Elisabeth André
RO-MAN4
2021 To Move or Not to Move? Social Acceptability of Robot Proxemics Behavior Depending on User Emotion
abstract
Various works show that proxemics occupies an important role in human-robot interaction and that appropriate proxemic interaction depends on many characteristics of humans and robots. However, there is none that shows the relationship between an emotional state expressed by a user and a proxemic reaction of the robot to it, in a social interaction between these interactants. In the current experiment (N = 82), we investigate this using an online study in which we examine which proxemic response (i.e., approaching, not moving, moving away) to a person’s expressed emotional state (i.e., anger, fear, disgust, surprise, sadness, joy) is perceived as appropriate. The quantitative and qualitative data collected suggests that the robot’s approach was considered appropriate for the expressed fear, sadness, and joy, whereas moving away was perceived as inappropriate in most scenarios. Further exploratory findings underline the importance of appropriate nonverbal behavior on the perception of the robot.
Björn Petrak, Julia G. Stapels, Katharina Weitz, Friederike Eyssel, Elisabeth André
RO-MAN5
2021 "Can you help me move this over there?": training children with ASD to joint action through tangible interaction and virtual agent
abstract
New technologies for autism focus on the training of either social skills or motor skills, but not both. Such a dichotomy omits a wide range of joint action tasks that require the coordination of two persons (e.g. moving a heavy furniture). The training of these physical tasks performed in dyad has great potential to foster inclusiveness while having an impact on both social and motor skills. In this paper, we present the design of a tangible and virtual interactive system for the training of children with Autism Spectrum Disorder (ASD) in performing joint actions. The proposed system is composed of a virtual character projected onto a surface on which a tangible object is magnetized: both the user and the virtual character hold the object, thus simulating a joint action. We report and discuss preliminary results of a field training study, which shows the potential of the interactive system.
Tom Giraud, Brian Ravenet, Chi Tai Dang, Jacqueline Nadel, Elise Prigent, Gael Poli, Elisabeth André, Jean-Claude Martin
TEI7
2021 Local and global explanations of agent behavior: Integrating strategy summaries with saliency maps
abstract
With advances in reinforcement learning (RL), agents are now being developed in high-stakes application domains such as healthcare and transportation. Explaining the behavior of these agents is challenging, as the environments in which they act have large state spaces, and their decision-making can be affected by delayed rewards, making it difficult to analyze their behavior. To address this problem, several approaches have been developed. Some approaches attempt to convey the global behavior of the agent, describing the actions it takes in different states. Other approaches devised local explanations which provide information regarding the agent's decision-making in a particular state. In this paper, we combine global and local explanation methods, and evaluate their joint and separate contributions, providing (to the best of our knowledge) the first user study of combined local and global explanations for RL agents. Specifically, we augment strategy summaries that extract important trajectories of states from simulations of the agent with saliency maps which show what information the agent attends to. Our results show that the choice of what states to include in the summary (global information) strongly affects people's understanding of agents: participants shown summaries that included important states significantly outperformed participants who were presented with agent behavior in a set of world-states that are likely to appear during gameplay. We find mixed results with respect to augmenting demonstrations with saliency maps (local information), as the addition of saliency maps, in the form of raw heat maps, did not significantly improve performance in most cases. However, we do find some evidence that saliency maps can help users better understand what information the agent relies on during its decision-making, suggesting avenues for future work that can further improve explanations of RL agents.
Tobias Huber, Katharina Weitz, Elisabeth André, Ofra Amir
Artif. Intell.3
2021 Editorial: Transactions on Affective Computing - Affective Computing in the Times of Pandemics
abstract
Presents the editorial for thiis issue of the publication.
Elisabeth André
IEEE Trans. Affect. Comput.1
2021 Multi-Modal Pain Intensity Recognition Based on the SenseEmotion Database
abstract
The subjective nature of pain makes it a very challenging phenomenon to assess. Most of the current pain assessment approaches rely on an individual’s ability to recognise and report an observed pain episode. However, pain perception and expression are affected by numerous factors ranging from personality traits to physical and psychological health state. Hence, several approaches have been proposed for the automatic recognition of pain intensity, based on measurable physiological and audiovisual parameters. In the current paper, an assessment of several fusion architectures for the development of a multi-modal pain intensity classification system is performed. The contribution of the presented work is two-fold: (1) 3 distinctive modalities consisting of audio, video and physiological channels are assessed and combined for the classification of several levels of pain elicitation. (2) An extensive assessment of several fusion strategies is carried out in order to design a classification architecture that improves the performance of the pain recognition system. The assessment is based on theSenseEmotion Databaseand experimental validation demonstrates the relevance of the multi-modal classification approach, which achieves classification rates of respectively$83.39\%$,$59.53\%$and$43.89\%$in a 2-class, 3-class and 4-class pain intensity classification task.
Patrick Thiam, Viktor Kessler, Mohammadreza Amirian, Peter Bellmann, Georg Layher, Yan Zhang 0054, Maria Velana, Sascha Gruss, Steffen Walter 0001, Harald C. Traue, Daniel Schork, Jonghwa Kim 0001, Elisabeth André, Heiko Neumann, Friedhelm Schwenker
IEEE Trans. Affect. Comput.13
2020 PiHearts: Resonating Experiences of Self and Others Enabled by a Tangible Somaesthetic Design
abstract
A human's heart beating can be sensed by sensors and displayed for others to see, hear, feel, and potentially "resonate'' with. Previous work in studying interaction designs with physiological data, such as a heart's pulse rate, have argued that feeding it back to the users may, for example support users' mindfulness and self-awareness during various everyday activities and ultimately support their health and wellbeing. Inspired by Somaesthetics as a discipline, we designed and explored multimodal displays, which enable experiencing heart beats as natural stimuli from oneself and others in social proximity. In this paper, we report on the design process of our design PiHearts and present qualitative results of a field study with 30 pairs of participants. Participants were asked to use PiHearts during watching short movies together and report their perceived experience in three different display conditions while watching movies. We found, for example that participants reported significant effects in experiencing sensory immersion when they received their own heart beats as stimuli compared to the condition without any heart beat display, and that feeling their partner's heart beats resulted in significant effects on social experience. We refer to resonance theory to motivate and discuss the results, highlighting the potential of how digitalization of heart beats as rhythmic natural stimuli may provide resonance in a modern society facing social acceleration.
Ilhan Aslan, Andreas Seiderer, Chi Tai Dang, Simon Rädler, Elisabeth André
ICMI5
2020 NOVA: A Tool for Explanatory Multimodal Behavior Analysis and Its Application to Psychotherapy
Tobias Baur 0001, Sina Clausen, Alexander Heimerl, Florian Lingenfelser, Wolfgang Lutz 0001, Elisabeth André
MMM (2)6
2020 An Evolutionary-based Generative Approach for Audio Data Augmentation
abstract
In this paper, we introduce a novel framework to augment raw audio data for machine learning classification tasks. For the first part of our framework, we employ a generative adversarial network (GAN) to create new variants of the audio samples that are already existing in our source dataset for the classification task. In the second step, we then utilize an evolutionary algorithm to search the input domain space of the previously trained GAN, with respect to predefined characteristics of the generated audio. This way we are able to generate audio in a controlled manner that contributes to an improvement in classification performance of the original task. To validate our approach, we chose to test it on the task of soundscape classification. We show that our approach leads to a substantial improvement in classification results when compared to a training routine without data augmentation and training with uncontrolled data augmentation with GANs.
Silvan Mertes, Alice Baird, Dominik Schiller, Björn W. Schuller, Elisabeth André
MMSP5
2020 Transactions on Affective Computing - Celebrating the 10th Year of Publication
abstract
Presents the editorial for this issue of the publication.
Elisabeth André
IEEE Trans. Affect. Comput.1
2020 A Generic Human-Machine Annotation Framework Based on Dynamic Cooperative Learning
abstract
The task of obtaining meaningful annotations is a tedious work, incurring considerable costs and time consumption. Dynamic active learning and cooperative learning are recently proposed approaches to reduce human effort of annotating data with subjective phenomena. In this paper, we introduce a novel generic annotation framework, with the aim to achieve the optimal tradeoff between label reliability and cost reduction by making efficient use of human and machine work force. To this end, we use dropout to assess model uncertainty and thereby to decide which instances can be automatically labeled by the machine and which ones require human inspection. In addition, we propose an early stopping criterion based on inter-rater agreement in order to focus human resources on those ambiguous instances that are difficult to label. In contrast to the existing algorithms, the new confidence measures are not only applicable to binary classification tasks but also regression problems. The proposed method is evaluated on the benchmark datasets for non-native English prosody estimation, provided in the Interspeech computational paralinguistics challenge. In the result, the novel dynamic cooperative learning algorithm yields 0.424 Spearman's correlation coefficient compared to 0.413 with passive learning, while reducing the amount of human annotations by 74%.
Yue Zhang 0014, Andrea Michi, Johannes Wagner 0001, Elisabeth André, Björn W. Schuller, Felix Weninger
IEEE Trans. Cybern.4
2019 NOVA - A tool for eXplainable Cooperative Machine Learning
abstract
In this paper, we introduce a next-generation annotation tool called NOVA, which implements a workflow that interactively incorporates the `human in the loop'. In particular, NOVA offers a collaborative annotation backend where multiple annotators join their workforce. A main aspect of NOVA is the possibility of applying semi-supervised active learning where Machine Learning techniques are used already during the annotation process by giving the possibility to pre-label data automatically. Furthermore, NOVA implements recent eXplainable AI (XAI) techniques to provide users with both, a confidence value of the automatically predicted annotations, as well as visual explanation. This way, annotators get to understand whether they can trust their ML models, or more annotated data is necessary.
Alexander Heimerl, Tobias Baur 0001, Florian Lingenfelser, Johannes Wagner 0001, Elisabeth André
ACII5
2019 Personalized Synthesis of Intentional and Emotional Non-Verbal Sounds for Social Robots
abstract
Non-verbal sounds are an essential communication channel for social robots. However, it requires expert knowledge to create and compose synthesizers, develop melodic structures or record samples which express a robot's internal intentions and emotions. This paper presents an approach for adapting a robot's timbre based on non-expert human comparative feedback in order to personalize the sonic interaction design to an individual user's preferences. An evolution strategy learns parameters of real-time sound synthesis for different intentions and emotions. Ultimately, the strategy aims to improve the perceived goodness of how well a specific melody's sound maps to a specific emotion or intention. In order to demonstrate the feasibility of the approach, we report on a user study with a robot, 6 exemplary melodies and 27 participants. Our study results show that the strategy indeed results in improved and preferred sound designs and that many participants are willing to apply such a process to improve their robots' expressivity.
Hannes Ritschel, Ilhan Aslan, Silvan Mertes, Andreas Seiderer, Elisabeth André
ACII5
2019 Socially-Aware User Interfaces: Can Genuine Sensitivity Be Learnt at all?
abstract
Recent years have initiated a paradigm shift from pure task-based human-machine interfaces towards socially-aware interaction. Advances in deep learning have led to anthropomorphic interfaces with robust sensing capabilities that come close to or even exceed human performance. In some cases, these interfaces may convey to humans the illusion of a sentient being that cares for them. At the same time, there is the risk that - at some point - these systems may have to reveal their lack of true comprehension of the situative context and the user’s needs with serious consequences to user trust. The talk will discuss challenges that arise when designing multimodal interfaces that hide the underlying complexity from the user, but still demonstrate a transparent and plausible behavior. It will argue for hybrid AI approaches that look beyond deep learning to encompass a theory of mind to obtain a better understanding of the rationale behind human behaviors.
Elisabeth André
ICMI1
2019 Creativity Support and Multimodal Pen-based Interaction
abstract
Creativity as a skill is associated with a potential to drive both productivity and psychological wellbeing. Since multimodality can foster cognitive ability, multimodal digital tools should also be ideal to support creativity as an essentially cognitive skill. In this paper, we explore this notion by presenting a multimodal pen-based interaction technique and studying how it supports creativity. The multimodal solution uses micro-controller-technology to augment a digital pen with RGB LEDs and a Leap Motion sensor to enable bimanual input. We report on a user study with 26 participants demonstrating that the multimodal technique is indeed perceived as supporting creativity significantly more than a baseline condition.
Ilhan Aslan, Katharina Weitz, Ruben Schlagowski, Simon Flutura, Susana Garcia Valesco, Marius Pfeil, Elisabeth André
ICMI7
2019 Relevance-Based Feature Masking: Improving Neural Network Based Whale Classification Through Explainable Artificial Intelligence
abstract
Underwater sounds provide essential information for marine researchers to study sea mammals.During long-term studies large amounts of sound signals are being recorded using hydrophones.To facilitate the time consuming process of manually evaluating the recorded data, computational systems are often employed.Recent approaches utilize Convolutional Neural Networks (CNNs) to analyze spectrograms extracted from the audio signal.In this paper we explore the potential of relevance analysis to enhance the performance of existing CNN approaches.For this purpose, we present a fusion system that utilizes intermediate outputs of three state of the art CNNs, which are fine tuned to recognize whale sounds in spectrograms.Hereby we use Explainable Artificial Intelligence (XAI) to asses the relevance of each feature within the obtained representations.Based on those relevance values, we create novel masking algorithms to extract significant subsets of respective representations.These subsets are used to train an ensemble of classification systems that are serving as input for the final fusion step.We observe that a classification system can benefit from the inclusion of Relevance-based Feature Masking in terms of improved performance and reduced input dimensionality.The presented work is part of the INTERSPEECH 2019 Computational Paralinguistics Challenge.
Dominik Schiller, Tobias Huber, Florian Lingenfelser, Michael Dietz, Andreas Seiderer, Elisabeth André
INTERSPEECH6
2019 Designing a Mobile Social and Vocational Reintegration Assistant for Burn-out Outpatient Treatment
abstract
Using Social Agents as health-care assistants or trainers is one focus area of IVA research. This paper presents a concept of our mobile Social Agent EmmA in the role of a vocational reintegration assistant for burn-out outpatient treatment. We follow a typical par- ticipatory design approach including experts and patients in order to address requirements from both sides. Since the success of such treatments is related to a patients emotion regulation capabilities, we employ a real-time social signal interpretation together with a computational simulation of emotion regulation that influences the agent's social behavior as well as the situational selection of verbal treatment strategies. Overall, our interdisciplinary approach sketches a novel integrative concept for Social Agents as assistants for burn-out patients.
Patrick Gebhard, Tanja Schneeberger, Michael Dietz, Elisabeth André, Nida ul Habib Bajwa
IVA4
2019 Designing the Impression of Social Agents' Real-time Interruption Handling
abstract
Human interaction partners can deal with interruptions and then resume the interaction. This ability should be emulated by social agents. How fast interruptions are handled might influence the overall impression of an agent. In this paper, we present the results of a user study on how a human dialog partner perceives the be- havior of a virtual agent handling verbal user interruptions with different reaction times. The study goes beyond typical perception experiments by preserving the real-time interaction experience. For the evaluation, we rely on a parametrizable parallelized computa- tional model that represents dialog flow, overlap detection, conflict recognition, and conflict handling in real-time. The evaluation re- sults show that the timing of the agent's interruption handling in interactive human-agent dialogues is related to different interper- sonal attitudes.
Patrick Gebhard, Tanja Schneeberger, Gregor Mehlmann, Tobias Baur 0001, Elisabeth André
IVA5
2019 "Do you trust me?": Increasing User-Trust by Integrating Virtual Agents in Explainable AI Interaction Design
abstract
While the research area of artificial intelligence benefited from increasingly sophisticated machine learning techniques in recent years, the resulting systems suffer from a loss of transparency and comprehensibility. This development led to an on-going resurgence of the research area of explainable artificial intelligence (XAI) which aims to reduce the opaqueness of those black-box-models. However, much of the current XAI-Research is focused on machine learning practitioners and engineers while omitting the specific needs of end-users. In this paper, we examine the impact of virtual agents within the field of XAI on the perceived trustworthiness of autonomous intelligent systems. To assess the practicality of this concept, we conducted a user study based on a simple speech recognition task. As a result of this experiment, we found significant evidence suggesting that the integration of virtual agents into XAI interaction design leads to an increase of trust in the autonomous intelligent system.
Katharina Weitz, Dominik Schiller, Ruben Schlagowski, Tobias Huber, Elisabeth André
IVA5
2019 Legal and Ethical Challenges in Multimedia Research
abstract
Multimedia research has now moved beyond laboratory experiments and is rapidly being deployed in real-life applications including advertisements, social interaction, search, security, automated driving, and healthcare. Hence, the developed algorithms now have a direct impact on the individuals using the abovementioned services and the society as a whole. While there is a huge potential to benefit the society using such technologies, there is also an urgent need to identify the checks and balances to ensure that the impact of such technologies is ethical and positive. This panel will bring together an array of experts who have experience collecting large-scale datasets, building multimedia algorithms, and deploying them in practical applications, as well as, a lawyer whose eyes have been on the fundamental rights at stake. They will lead a discussion on the ethics and lawfulness of dataset creation, licensing, privacy of individuals represented in the datasets, algorithmic transparency, algorithmic bias, explainability, and the implications of application deployment. Through an interactive process engaging the audience, the panel hopes to: increase the awareness of such concepts in the multimedia research community; initiate a discussion on community guidelines all for setting the future direction of conducting multimedia research in a lawful and ethical manner.
Vivek K. Singh 0001, Elisabeth André, Susanne Boll, Mireille Hildebrandt, David A. Shamma, Tat-Seng Chua
ACM Multimedia2
2019 Mouse, touch, or fich: comparing traditional input modalities to a novel pre-touch technique
abstract
Finger touch and mouse-based interaction are today's predominant modalities to interact with screen-based user interfaces. Related work suggests that new techniques interweaving pre-touch sensing and touch are useful future alternatives. In this paper, we introduce Fich, a novel pre-touch technique that augments conventional touch interfaces with tooltips and further "fingerover" effects, opening up the space in front of the screen for user interaction. To study Fich in-depth, we developed a Fich-enabled weather application and compared user experience and interface discovery ("serendipity") of Fich against the traditional input modalities Mouse and finger Touch in a user study with 42 subjects. We report on the results, implying Fich's user experience to be rated significantly higher in terms of hedonic quality and significantly lower in terms of pragmatic quality, as compared to traditional input modalities.
Lea Rieger, Ilhan Aslan, Christoph Anneser, Malte Sandstede, Felix Schwarzmeier, Björn Petrak, Elisabeth André
MUM8
2019 Let Me Show You Your New Home: Studying the Effect of Proxemic-awareness of Robots on Users' First Impressions
abstract
First impressions play an important part in social interactions, establishing the foundation of a person's opinion about their counterparts. Since interpersonal communication is essentially multimodal, people are judged during first encounters by both their verbal utterances and nonverbal behavior, such as how they utilize eye contact, body distance, and body orientation. In this paper, we argue that robots would provide better user experiences, including being perceived as more likable if they were able to make a good first impression when introduced to a new home. Moreover, we wanted to test if robots can improve their perceived impression by behaving in a proxemic-aware manner; i.e., by following established social norms, which prescribe, for example how far people should position themselves around other objects to improve the facilitation of social interactions. In order to test this hypothesis, we conducted a user study with 16 participants in a virtual reality setting, comparing the impression of two agents being introduced to their new homes by users. We found that the proxemic-aware agent was indeed perceived as significantly better considering multiple constructs, including perceived anthropomorphism and trustworthiness.
Björn Petrak, Katharina Weitz, Ilhan Aslan, Elisabeth André
RO-MAN4
2019 Of Smarthomes, IoT Plants, and Implicit Interaction Design
abstract
There seems to be a danger to carelessly replace routine tasks in homes through automation with IoT-technology. But since routines, such as watering houseplants also have positive influences on inhabitants' wellbeing, they should be transformed through carefully performed designs. To this end, an attempt to use technology for augmenting a set of houseplants' non-verbal communication capabilities is presented. First, we describe in detail how implicit interactions have been designed to support inhabitants in watering their plants through meaningful interactions. Then, we report on a field study with 24 participants, comparing two alternative design implementations based on contrasting embodied interaction technologies (i.e., augmented reality and embedded computing technology). The study results highlight shortcomings of today's smartphone mediated augmented reality compared to physical interface alternatives, considering measurements of perceived attractiveness and expected effects on determinants of wellbeing, and discusses potentials of combining both modalities for future solutions.
Björn Petrak, Ilhan Aslan, Chi Tai Dang, Elisabeth André
TEI4
2019 IEEE Transactions on Affective Computing-Entering the 10th Year of Publication
Elisabeth André
IEEE Trans. Affect. Comput.1
2019 Serious Games for Training Social Skills in Job Interviews
abstract
In this paper, we focus on experience-based role play with virtual agents to provide young adults at the risk of exclusion with social skill training. We present a scenario-based serious game simulation platform. It comes with a social signal interpretation component, a scripted and autonomous agent dialog and social interaction behavior model, and an engine for 3-D rendering of lifelike virtual social agents in a virtual environment. We show how two training systems developed on the basis of this simulation platform can be used to educate people in showing appropriate socioemotive reactions in job interviews. Furthermore, we give an overview of four conducted studies investigating the effect of the agents' portrayed personality and the appearance of the environment on the players' perception of the characters and the learning experience.
Patrick Gebhard, Tanja Schneeberger, Elisabeth André, Tobias Baur 0001, Ionut Damian, Gregor Mehlmann, Cornelius J. König, Markus Langer
IEEE Trans. Games3
2018 Progress to a VOCA with Prosodic Synthesised Speech
Jan-Oliver Wülfing, Elisabeth André
ICCHP (1)2
2018 Gazeover - Exploring the UX of Gaze-triggered Affordance Communication for GUI Elements
abstract
The user experience (UX) of graphical user interfaces (GUIs) often depends on how clearly visual designs communicate/signify "affordances", such as if an element on the screen can be pushed, dragged, or rotated. Especially for novice users figuring out the complexity of a new interface can be cumbersome. In the "past" era of mouse-based interaction mouseover effects were successfully utilized to trigger a variety of assistance, and help users in exploring interface elements without causing unintended interactions and associated negative experiences. Today's GUIs are increasingly designed for touch and lack a method similiar to mouseover to help (novice) users to get acquainted with interface elements. In order to address this issue, we have studied gazeover, as a technique for triggering "help or guidance" when a user's gaze is over an interactive element, which we believe is suitable for today's touch interfaces. We report on a user study comparing pragmatic and hedonic qualities of gazeover and mouseover, which showed significant higher ratings in hedonic quality for the gazeover technique. We conclude by discussing limitations and implications of our findings.
Ilhan Aslan, Michael Dietz, Elisabeth André
ICMI3
2018 Pen + Mid-Air Gestures: Eliciting Contextual Gestures
abstract
Combining mid-air gestures with pen input for bi-manual input on tablets has been reported as an alternative and attractive input technique in drawing applications. Previous work has also argued that mid-air gestural input can cause discomfort and arm fatigue over time, which can be addressed in a desktop setting by allowing users to gesture in alternative restful arm positions (e.g., elbow rests on desk). However, it is unclear if and how gesture preferences and gesture designs would be different for alternative arm positions. In order to inquire these research question we report on a user and choice based gesture elicitation study in which 10 participants designed gestures for different arm positions. We provide an in-depth qualitative analysis and detailed categorization of gestures, discussing commonalities and differences in the gesture sets based on a "think aloud" protocol, video recordings, and self-reports on user preferences.
Ilhan Aslan, Tabea Schmidt, Jens Woehrle, Lukas Vogel 0001, Elisabeth André
ICMI5
2018 EVA: A Multimodal Argumentative Dialogue System
abstract
This work introduces EVA, a multimodal argumentative Dialogue System that is capable of discussing controversial topics with the user. The interaction is structured as an argument game in which the user and the system select respective moves in order to convince their opponent. EVA's response is presented as a natural language utterance by a virtual agent that supports the respective content using characteristic gestures and mimic.
Niklas Rach, Klaus Weber 0001, Louisa Pragst, Elisabeth André, Wolfgang Minker, Stefan Ultes
ICMI4
2018 How to Shape the Humor of a Robot - Social Behavior Adaptation Based on Reinforcement Learning
abstract
A shared sense of humor can result in positive feelings associated with amusement, laughter, and moments of bonding. If robotic companions could acquire their human counterparts' sense of humor in an unobtrusive manner, they could improve their skills of engagement. In order to explore this assumption, we have developed a dynamic user modeling approach based on Reinforcement Learning, which allows a robot to analyze a person's reaction while it tells jokes and continuously adapts its sense of humor. We evaluated our approach in a test scenario with a Reeti robot acting as an entertainer and telling different types of jokes. The exemplary adaptation process is accomplished only by using the audience's vocal laughs and visual smiles, but no other form of explicit feedback. We report on results of a user study with 24 participants, comparing our approach to a baseline condition (with a non-learning version of the robot) and conclude by providing limitations and implications of our approach in detail.
Klaus Weber 0001, Hannes Ritschel, Ilhan Aslan, Florian Lingenfelser, Elisabeth André
ICMI5
2018 Deep Learning in Paralinguistic Recognition Tasks: Are Hand-crafted Features Still Relevant?
abstract
In the past, the performance of machine learning algorithms depended heavily on the representation of the data.Well-designed features therefore played a key role in speech and paralinguistic recognition tasks.Consequently, engineers have put a great deal of work into manually designing large and complex acoustic feature sets.With the emergence of Deep Neural Networks (DNNs), however, it is now possible to automatically infer higher abstractions from simple spectral representations or even learn directly from raw waveforms.This raises the question if (complex) hand-crafted features will still be needed in the future.We take this year's INTERSPEECH Computational Paralinguistic Challenge as an opportunity to approach this issue by means of two corpora -Atypical Affect and Crying.At first, we train a Recurrent Neural Network (RNN) to evaluate the performance of several hand-crafted feature sets of varying complexity.Afterwards, we make the network do the feature engineering all on its own by prefixing a stack of convolutional layers.Our results show that there is no clear winner (yet).This creates room to discuss chances and limits of either approach.
Johannes Wagner 0001, Dominik Schiller, Andreas Seiderer, Elisabeth André
INTERSPEECH4
2018 Decision-Theoretic Personality-Based Reasoning about Turn-Taking Conflicts
abstract
This paper outlines the use of an influence diagram for modeling turn-taking timing. In contrast to related works, our model focuses on an agent's personality and attitude towards the conversation partner. We also describe how this model is implemented in a first prototype application.
Kathrin Janowski, Elisabeth André
IVA2
2018 Providing Life-Style-Intervention to Improve Well-Being of Elderly People
Thomas Rist, Andreas Seiderer, Elisabeth André
ICEC3
2018 Exploring the User Experience of Proxemic Hand and Pen Input Above and Aside a Drawing Screen
abstract
Digital drawing experiences are not only fused by the flexibility of digital materials but also influenced by the availability of interaction space. In this paper, we first present a prototype, which implements a method to turn the (mid-air) space above and aside a drawing screen in a desktop setting dynamically into sensory space for gestural and spatial input. Then we report on a user study exploring how participants experience digital drawing when the additional interaction space above and aside a screen is exploited for exemplary proxemic input techniques for zooming and panning a drawing. Our results show that the new multimodal input techniques are perceived as significantly more attractive than a baseline drawing condition which only utilizes touch based input. We conclude by discussing implications and limitations of our findings and input above and aside a drawing screen in general.
Ilhan Aslan, Björn Petrak, Florian Müller 0011, Elisabeth André
MUM4
2018 Honeypot: A Socializing App to Promote Train Commuters' Wellbeing
abstract
The number of commuters has been increasing for many years and the negative effects on wellbeing are therefore affecting more and more people. Following a user centered design process that focuses on known wellbeing determinants, such as relatedness and empathy, we developed the Honeypot socializing app. The app allows commuters to find other travelers to chat with and meet in person to enhance their wellbeing through fostering meaningful and contextual social interactions. First, we describe the development of the idea and the design of the app. Then, we report on a field study with 16 participants, which we carried out on trains. The study results show that the app helps to get in contact with fellow travelers and that it has the potential to promote the wellbeing of commuters in the long term.
Christoph Anneser, Malte Sandstede, Lea Rieger, Adnan Alhomssi, Felix Schwarzmeier, Björn Petrak, Ilhan Aslan, Elisabeth André
MUM9
2018 Asynchronous and Event-Based Fusion Systems for Affect Recognition on Naturalistic Data in Comparison to Conventional Approaches
abstract
Throughout many present studies dealing with multi-modal fusion, decisions are synchronously forced for fixed time segments across all modalities. Varying success is reported, sometimes performance is worse than unimodal classification. Our goal is the synergistic exploitation of multimodality whilst implementing a real-time system for affect recognition in a naturalistic setting. Therefore we present a categorization of possible fusion strategies for affect recognition on continuous time frames of complete recording sessions and we evaluate multiple implementations from resulting categories. These involve conventional fusion strategies as well as novel approaches that incorporate the asynchronous nature of observed modalities. Some of the latter algorithms consider temporal alignments between modalities and observed frames by applying asynchronous neural networks that use memory blocks to model temporal dependencies. Others use an indirect approach that introduces events as an intermediate layer to accumulate evidence for the target class through all modalities. Recognition results gained on a naturalistic conversational corpus show a drop in recognition accuracy when moving from unimodal classification to synchronous multimodal fusion. However, with our proposed asynchronous and event-based fusion techniques we are able to raise the recognition system's accuracy by 7.83 percent compared to video analysis and 13.71 percent in comparison to common fusion strategies.
Florian Lingenfelser, Johannes Wagner 0001, Raymond Brueckner, Björn W. Schuller, Elisabeth André
IEEE Trans. Affect. Comput.6
2018 MyBrush: Brushing and Linking with Personal Agency
abstract
We extend the popular brushing and linking technique by incorporating personal agency in the interaction. We map existing research related to brushing and linking into a design space that deconstructs the interaction technique into three components: source (what is being brushed), link (the expression of relationship between source and target), and target (what is revealed as related to the source). Using this design space, we created MyBrush, a unified interface that offers personal agency over brushing and linking by giving people the flexibility to configure the source, link, and target of multiple brushes. The results of three focus groups demonstrate that people with different backgrounds leveraged personal agency in different ways, including performing complex tasks and showing links explicitly. We reflect on these results, paving the way for future research on the role of personal agency in information visualization.
Philipp Koytek, Charles Perin, Jo Vermeulen, Elisabeth André, Sheelagh Carpendale
IEEE Trans. Vis. Comput. Graph.4
2017 Pre-touch proxemics: moving the design space of touch targets from still graphics towards proxemic behaviors
abstract
Proxemic touch targets continuously change in relation to a user's hand in mid-air before a physical touch occurs. Previous work has, for example shown that expanding targets are capable to improve target acquisition performance on touch interfaces. However, it is unclear how proxemic touch targets influence user experience (UX) in a broader sense, including hedonic qualities. Towards closing this research gap the paper reports on two user studies. The first study is a qualitative study with five experts, providing in-depth insights on a variety of change-types (e.g., size, form, color) and how they, for example, influence perceived functional and aesthetic qualities of proxemic touch targets. A follow-up user study with 36 participants explores the UX of a proxemic touch target compared to non-proximal versions of the same target. The results highlight a positive significant effect of the proxemic design on both pragmatic and hedonic qualities.
Ilhan Aslan, Elisabeth André
ICMI2
2017 The NoXi database: multimodal recordings of mediated novice-expert interactions
abstract
We present a novel multi-lingual database of natural dyadic novice-expert interactions, named NoXi, featuring screen-mediated dyadic human interactions in the context of information exchange and retrieval. NoXi is designed to provide spontaneous interactions with emphasis on adaptive behaviors and unexpected situations (e.g. conversational interruptions). A rich set of audio-visual data, as well as continuous and discrete annotations are publicly available through a web interface. Descriptors include low level social signals (e.g. gestures, smiles), functional descriptors (e.g. turn-taking, dialogue acts) and interaction descriptors (e.g. engagement, interest, and fluidity).
Angelo Cafaro, Johannes Wagner 0001, Tobias Baur 0001, Soumia Dermouche, Mercedes Torres, Catherine Pelachaud, Elisabeth André, Michel F. Valstar
ICMI7
2017 Exploring Opportunistic Ambient Notifications in the Smart Home to Enhance Quality of Live
Andreas Seiderer, Chi Tai Dang, Elisabeth André
ICOST3
2017 Infected Phonemes: How a Cold Impairs Speech on a Phonetic Level
abstract
The realization of language through vocal sounds involves a complex interplay between the lungs, the vocal cords, and a series of resonant chambers (e.g.mouth and nasal cavities).Due to their connection to the outside world, these body parts are popular spots for viruses and bacteria to enter the human organism.Affected people may suffer from an upper respiratory tract infection (URTIC) and consequently their voice often sounds breathy, raspy or sniffly.In this paper, we investigate the audible effects of a cold on a phonetic level.Results on a German corpus show that the articulation of consonants is more impaired than that of vowels.Surprisingly, nasal sounds do not follow this trend in our experiments.We finally try to predict a speaker's health condition by fusing decisions we derive from single phonemes.The presented work is part of the INTER-SPEECH 2017 Computational Paralinguistics Challenge.
Johannes Wagner 0001, Thiago Fraga-Silva, Yvan Josse, Dominik Schiller, Andreas Seiderer, Elisabeth André
INTERSPEECH6
2017 Temporal Visualization of Energy Consumption Loads Using Time-Tone
abstract
Feedback plays an important role in assisting users to better understand their energy consumption behaviour. This is particularly true when users want to change their behaviour in order to reduce their energy consumption, and to manage their usage more effectively so as to avoid putting unnecessary load on energy providers. This paper presents the time-tone visualization, which aims to assist users by displaying variations in energy consumption by different categories of household devices over time, and their respective contributions to the total energy usage load. A user study conducted to compare time-tone against area-charts shows that although the two visualizations are comparable, time-tone is more effective for cases where there are large variations in energy usage loads.
Masood Masoodian, Ida Buchwald, Saturnino Luz, Elisabeth André
IV4
2017 Adapting a Robot's linguistic style based on socially-aware reinforcement learning
abstract
When looking at Socially Interactive Robots, adaptation to the user's preferences plays an important role in today's Human-Robot Interaction to keep interaction interesting and engaging over a long period of time. Findings indicate an increase in user engagement for robots with adaptive behavior and personality, but also that it depends on the task context whether a similar or opposing robot personality is preferred. We present an approach based on Reinforcement Learning, which gets its reward directly from social signals in real-time during the interaction, to quickly learn about and dynamically address individual human preferences. Our scenario involves a Reeti robot in the role of a story teller talking about the main characters in the novel “Alice's Adventures in Wonderland” by generating descriptions with varying degree of introversion/extraversion. After initial simulation results, an interactive prototype is presented which allows to explore the learning process adapting to the human interaction partner's engagement.
Hannes Ritschel, Tobias Baur 0001, Elisabeth André
RO-MAN3
2016 Measuring the impact of multimodal behavioural feedback loops on social interactions
abstract
In this paper we explore the concept of automatic behavioural feedback loops during social interactions. Behavioural feedback loops (BFL) are rapid processes which analyse the behaviour of the user in realtime and provide the user with live feedback on how to improve the behaviour quality. In this context, we implemented an open source software framework for designing, creating and executing BFL on Android powered mobile devices. To get a better understanding of the effects of BFL on face-to-face social interactions, we conducted a user study and compared between four different BFL types spanning three modalities: tactile, auditory and visual. For the study, the BFL have been designed to improve the users' perception of their speaking time in an effort to create more balanced group discussions. The study yielded valuable insights into the impact of BFL on conversations and how humans react to such systems.
Ionut Damian, Tobias Baur 0001, Elisabeth André
ICMI3
2016 Social signal processing for dummies
abstract
We introduce SSJ Creator, a modern Android GUI enabling users to design and execute social signal processing pipelines using nothing but their smartphones and without writing a single line of code. It is based on a modular Java-based social signal processing framework (SSJ), which is able to perform realtime multimodal behaviour analysis on Android devices using both device internal and external sensors.
Ionut Damian, Michael Dietz, Frank Gaibler, Elisabeth André
ICMI4
2016 MobileSSI: asynchronous fusion for social signal interpretation in the wild
abstract
Over the last years, mobile devices have become an integral part of people's everyday life. At the same time, they provide more and more computational power and memory capacity to perform complex calculations that formerly could only be accomplished with bulky desktop machines. These capabilities combined with the willingness of people to permanently carry them around open up completely new perspectives to the area of Social Signal Processing. To allow for an immediate analysis and interaction, real-time assessment is necessary. To exploit the benefits of multiple sensors, fusion algorithms are required that are able to cope with data loss in asynchronous data streams. In this paper we present MobileSSI, a port of the Social Signal Interpretation (SSI) framework to Android and embedded Linux platforms. We will test to what extent it is possible to run sophisticated synchronization and fusion mechanisms in an everyday mobile setting and compare the results with similar tasks in a laboratory environment.
Simon Flutura, Johannes Wagner 0001, Florian Lingenfelser, Andreas Seiderer, Elisabeth André
ICMI5
2016 Laughter detection in the wild: demonstrating a tool for mobile social signal processing and visualization
abstract
In this demo, we present MobileSSI, a flexible software framework for Android and embedded Linux platforms, that provides developers with tools to record, analyze and recognize human behavior in real-time on mobile devices. To illustrate the benefits of the framework for the analysis of social group dynamics in naturalistic mobile settings, we present a demonstrator for laughter recognition that was implemented with MobileSSI. The demonstrator makes use of smartphones for sensing and analyzing data and employs smartwatches and tablets for visualizing the results and providing user feedback. To enable communication within the resulting ecology of mobile devices, MobileSSI includes a web socket plugin.
Simon Flutura, Johannes Wagner 0001, Florian Lingenfelser, Andreas Seiderer, Elisabeth André
ICMI5
2016 Ask Alice: an artificial retrieval of information agent
abstract
We present a demonstration of the ARIA framework, a modular approach for rapid development of virtual humans for information retrieval that have linguistic, emotional, and social skills and a strong personality. We demonstrate the framework's capabilities in a scenario where `Alice in Wonderland', a popular English literature book, is embodied by a virtual human representing Alice. The user can engage in an information exchange dialogue, where Alice acts as the expert on the book, and the user as an interested novice. Besides speech recognition, sophisticated audio-visual behaviour analysis is used to inform the core agent dialogue module about the user's state and intentions, so that it can go beyond simple chat-bot dialogue. The behaviour generation module features a unique new capability of being able to deal gracefully with interruptions of the agent.
Michel F. Valstar, Tobias Baur 0001, Angelo Cafaro, Alexandru Ghitulescu, Blaise Potard, Johannes Wagner 0001, Elisabeth André, Laurent Durieu, Matthew P. Aylett, Soumia Dermouche, Catherine Pelachaud, Eduardo Coutinho, Björn W. Schuller, Yue Zhang 0014, Dirk Heylen, Mariët Theune, Jelte van Waterschoot
ICMI7
2016 Exploring Eye-Tracking-Based Detection of Visual Search for Elderly People
abstract
Visual search plays an important role in our daily lives and can be very frustrating whenever we cannot remember where we left objects, such as keys or wallets. This is especially true for elderly people, since they forget things more often and face this challenge very frequently. While there are several studies which analyze eye movements during visual search, none of them actually tries to detect whether a user is searching for something or not. However, this information is necessary to recognize when the user needs assistance. Therefore, we propose an eye-tracking-based multimodal approach in order to detect visual search and to support the user in that situation. Furthermore, we explore multiple strategies to inform the user of the desired object's location using a head mounted display. With the help of a prototypical implementation and evaluation of the acquired sensor data, we show that our method is feasible and capable of dealing with this challenge.
Michael Dietz, Daniel Schork, Elisabeth André
Intelligent Environments3
2016 MobileSSI - A Multi-modal Framework for Social Signal Interpretation on Mobile Devices
abstract
Over the last years, new generations of mobile devices have found their way into our pockets. They provide more and more computational power and memory capacity to perform complex calculations that formerly could only be accomplished with bulky desktop machines. Moreover, mobile devices are equipped with a range of sensors to capture people's motion, environmental sound etc. These capabilities combined with the willingness of people to permanently carry them around open up completely new ways of observing human behaviour no longer in laboratories, but "in the wild". However, the detection and analysis of social cues is still a challenging task and requires adequate tools to synchronise, process and analyse relevant signals. This may be the reason why many studies and applications focus on offline analysis and typically collect data over long periods of time and analyse them afterwards. To allow for immediate feedback, real-time assessment is necessary. In this paper, we present MobileSSI, a port of the Social Signal Interpretation (SSI) framework to Android and embedded Linux platforms. The framework supports the joint development of processing pipelines for the analysis of social signals on a desktop computer and mobile devices. Throughout the paper we report on challenges we had to face when porting SSI to a mobile context. Furthermore, we summarise first experiences with a real-life setting in a pub where we focused on the analysis of multimodal social group dynamics investigating laughter as a sign of enjoyment.
Simon Flutura, Johannes Wagner 0001, Florian Lingenfelser, Andreas Seiderer, Elisabeth André
Intelligent Environments5
2016 Socially-Sensitive Interfaces: From Offline Studies to Interactive Experiences
abstract
Recent years have initiated a paradigm shift from pure taskbased human-machine interfaces towards socially-sensitive interaction. In addition to what users explicitly say or gesture at, socially-sensitive interfaces are able to sense more subtle human cues, such as head postures and movements, to infer psychological user states, such as attention and affect, and also to enrich system responses with social signals. However, most approaches focus on offline analysis of previously recorded data limiting the investigation to prototypical behaviors in laboratory-like settings. In my presentation, I will focus on challenges that arise when integrating social signal processing techniques into interactive systems designed for real-world applications. From a technical perspective, this requires effective tools able to synchronize, process, and analyze relevant signals in online mode. From a user perspective, appropriate strategies need to be defined to respond to social signals at the right moment in time without disturbing the flow of interaction. I will discuss two interaction styles for socially-sensitive interfaces. In the area of information retrieval, the concept of empathic stimulation has been used to optimize the selection and presentation of data. The basic idea is to exploit sensory data on the users' emotional state to provide them with cues that inspire their curiosity during the data exploration task. In the domain of social coaching, the concept of social augmentation has been employed to give people ambient feedback on their behavior while being engaged in a social interaction. The presentation will be illustrated by examples from various national and international projects following these two interaction styles.
Elisabeth André
IUI1
2016 Investigating Politeness Strategies and Their Persuasiveness for a Robotic Elderly Assistant
Stephan Hammer, Birgit Lugrin, Sergey Bogomolov, Kathrin Janowski, Elisabeth André
PERSUASIVE5
2016 Exploring the Potential of Realtime Haptic Feedback during Social Interactions
abstract
We explore the use of haptic feedback to deliver supportive information during social interactions in realtime. In an exploratory user study, we investigated perceptual limitations of vibration patterns during a conversation between peers. The results from this study have then been used to develop a system for providing users with realtime information regarding the quality of their nonverbal behaviour while engaged in a public speech.
Ionut Damian, Elisabeth André
TEI2
2016 The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affective Computing
abstract
Work on voice sciences over recent decades has led to a proliferation of acoustic parameters that are used quite selectively and are not always extracted in a similar fashion. With many independent teams working in different research areas, shared standards become an essential safeguard to ensure compliance with state-of-the-art methods allowing appropriate comparison of results across studies and potential integration and combination of extraction and recognition systems. In this paper we propose a basic standard acoustic parameter set for various areas of automatic voice analysis, such as paralinguistic or clinical speech analysis. In contrast to a large brute-force parameter set, we present a minimalistic set of voice parameters here. These were selected based on a) their potential to index affective physiological changes in voice production, b) their proven value in former studies as well as their automatic extractability, and c) their theoretical significance. The set is intended to provide a common baseline for evaluation of future research and eliminate differences caused by varying parameter sets or even different implementations of the same parameters. Our implementation is publicly available with the openSMILE toolkit. Comparative evaluations of the proposed feature set and large baseline feature sets of INTERSPEECH challenges show a high performance of the proposed set in relation to its size.
Florian Eyben, Klaus R. Scherer, Björn W. Schuller, Johan Sundberg, Elisabeth André, Carlos Busso, Laurence Devillers, Julien Epps, Petri Laukka, Shri Narayanan, Khiet P. Truong
IEEE Trans. Affect. Comput.5
2015 The Belfast storytelling database: A spontaneous social interaction database with laughter focused annotation
abstract
To support the endeavor of creating intelligent interfaces between computers and humans the use of training materials based on realistic human-human interactions has been recognized as a crucial task. One of the effects of the creation of these databases is an increased realization of the importance of often overlooked social signals and behaviours in organizing and orchestrating our interactions. Laughter is one of these key social signals; its importance in maintaining the smooth flow of human interaction has only recently become apparent in the embodied conversational agent domain. In turn, these realizations require training data that focus on these key social signals. This paper presents a database that is well annotated and theoretically constructed with respect to understanding laughter as it is used within human social interaction. Its construction, motivation, annotation and availability are presented in detail in this paper.
Gary McKeown, William Curran, Johannes Wagner 0001, Florian Lingenfelser, Elisabeth André
ACII5
2015 Games are Better than Books: In-Situ Comparison of an Interactive Job Interview Game with Conventional Training
Ionut Damian, Tobias Baur 0001, Birgit Lugrin, Patrick Gebhard, Gregor Mehlmann, Elisabeth André
AIED6
2015 Augmenting Social Interactions: Realtime Behavioural Feedback using Social Signal Processing Techniques
abstract
Nonverbal and unconscious behaviour is an important component of daily human-human interaction. This is especially true in situations such as public speaking, job interviews or information sensitive conversations, where researchers have shown that an increased awareness of one's behaviour can improve the outcome of the interaction. With wearable technology, such as Google Glass, we now have the opportunity to augment social interactions and provide realtime feedback on one's behaviour in an unobtrusive way. In this paper we present Logue, a system that provides realtime feedback on the presenters' openness, body energy and speech rate during public speaking. The system analyses the user's nonverbal behaviour using social signal processing techniques and gives visual feedback on a head-mounted display. We conducted two user studies with a staged and a real presentation scenario which yielded that Logue's feedback was perceived helpful and had a positive impact on the speaker's performance.
Ionut Damian, Chiew Seng Sean Tan, Tobias Baur 0001, Johannes Schöning, Kris Luyten, Elisabeth André
CHI6
2015 Fostering Smart Energy Applications
Masood Masoodian, Elisabeth André, Thomas Rist
INTERACT (4)2
2015 Combining hierarchical classification with frequency weighting for the recognition of eating conditions
abstract
Though parents regularly remind their children not to do so, talking while eating is a typical everyday situation automatic speech analysis systems should be able to deal with.The Paralinguistic Eating Condition (EC) Challenge at INTERSPEECH 2015 sets the task to classify whether a speaker is eating or not, and if so, which type of food the speaker is currently tasting.The approach we follow in this paper is rather unusual: instead of suppressing the influence of noise to enhance the intelligibility of a spoken message, we try to emphasize the noisy parts of the spectrum to improve the recognition of food classes.To allow for a fine-grained adaption to the characteristic spectrum of single food types we adopt a hierarchical tree structure and decompose the classification task into a sequence of binary decisions.At each node we apply frequency-dependent weighting to tune the spectrum to the involved target classes.With our approach we are able to improve results in a 7-class recognition problem (6 types of food and no food) by more than 7% on the training set (using leave-one-eater-out cross validation) and 4% on the test set, respectively.
Johannes Wagner 0001, Andreas Seiderer, Florian Lingenfelser, Elisabeth André
INTERSPEECH4
2015 Visualization Support for Comparing Energy Consumption Data
abstract
Providing effective feedback can empower users to change their behaviour and take the necessary actions to reduce their energy consumption. The types of feedback that allow comparison of energy usage seem to be particularly valuable. This paper introduces the time-stack visualization, which has been designed to support comparisons of individual and collective energy usage data. It also describes a user study conducted to compare the effectiveness of time-stack against a similar visualization called time-pie. The results show that although the two visualizations are generally comparable in their effectiveness, users rate time-stack more favourably.
Masood Masoodian, Birgit Lugrin, René Bühling, Elisabeth André
IV4
2015 Context-Aware Automated Analysis and Annotation of Social Human-Agent Interactions
abstract
The outcome of interpersonal interactions depends not only on the contents that we communicate verbally, but also on nonverbal social signals. Because a lack of social skills is a common problem for a significant number of people, serious games and other training environments have recently become the focus of research. In this work, we present NovA ( No n v erbal behavior A nalyzer), a system that analyzes and facilitates the interpretation of social signals automatically in a bidirectional interaction with a conversational agent. It records data of interactions, detects relevant social cues, and creates descriptive statistics for the recorded data with respect to the agent's behavior and the context of the situation. This enhances the possibilities for researchers to automatically label corpora of human--agent interactions and to give users feedback on strengths and weaknesses of their social behavior.
Tobias Baur 0001, Gregor Mehlmann, Ionut Damian, Florian Lingenfelser, Johannes Wagner 0001, Birgit Lugrin, Elisabeth André, Patrick Gebhard
ACM Trans. Interact. Intell. Syst.7
2015 Trust-based decision-making for smart and adaptive environments
abstract
Abstract Smart environments are able to support users during their daily life. For example, smart energy systems can be used to support energy saving by controlling devices, such as lights or displays, depending on context information, such as the brightness in a room or the presence of users. However, proactive decisions should also match the users’ preferences to maintain the users’ trust in the system. Wrong decisions could negatively influence the users’ acceptance of a system and at worst could make them abandon the system. In this paper, a trust-based model, called User Trust Model (UTM), for automatic decision-making is proposed, which is based on Bayesian networks. The UTM’s construction, the initialization with empirical data gathered in an online survey, and its integration in an office setting are described. Furthermore, the results of a live study and a live survey analyzing the users’ experience and acceptance are presented.
Stephan Hammer, Michael Wissner, Elisabeth André
User Model. User Adapt. Interact.3
2014 Fostering smart energy applications through advanced visual interfaces
abstract
There is an increasing need for technology that assist people with more effective monitoring and management of their energy generation and consumption. In recent years a considerable number of research activities have resulted in a multitude of new ICT-supported tools and services for both the private energy consumer market, as well as for energy related business and industries (e.g., utility and grid companies, facility management, etc.). This workshop focuses on advanced interaction, interface, and visualization techniques for energy-related applications, tools, and services. It brings together researchers and practitioners from a diverse range of background, including interaction design, human-computer interaction, visualization, computer games, and other fields concerned with the development of advanced visual interfaces for smart energy applications.
Masood Masoodian, Elisabeth André, Saturnino Luz, Thomas Rist
AVI2
2014 Modeling Gaze Mechanisms for Grounding in HRI
abstract
Grounding is essential in human interaction and crucial for social robots collaborating with humans. Gaze plays versatile roles for establishing, maintaining and repairing the common ground. It is combined with parallel modalities and involved in several processes for behavior generation and recognition. We present a uniform modeling approach focusing on the multi-modal, parallel and bidirectional aspects of gaze and their interleaving with the dialog logic.
Gregor Mehlmann, Kathrin Janowski, Tobias Baur 0001, Markus Häring, Elisabeth André, Patrick Gebhard
ECAI5
2014 Would you like to play with me?: how robots' group membership and task features influence human-robot interaction
abstract
In the present experiment, we investigated how robots' social category membership and characteristics of an HRI task affect humans' evaluative and behavioral reactions toward robots. Participants (N = 38) played a card game together with two robots, one belonging to participants' social in-group and the other one being a social out-group member. Furthermore, participants were either asked to cooperate with the in- and to compete with the out-group robot (congruent condition), or they were asked to cooperate with the out-group robot while competing with the in-group robot (incongruent condition). The results largely support our hypotheses: Participants showed more positive evaluative reactions toward the in-group (vs. the out-group) robot and they anthropomorphized it more strongly, independent of the congruency or incongruence of the HRI. Moreover, if required, participants cooperated with both the in- and the out-group robot, whereas their cooperativeness was more pronounced toward the in-group robot. Finally, participants indicated more difficulties with the HRI in the incongruent vs. the congruent condition. The theoretical and practical implications of the findings are discussed.
Markus Häring, Dieta Kuchenbrandt, Elisabeth André
HRI3
2014 Exploring a Model of Gaze for Grounding in Multimodal HRI
abstract
Grounding is an important process that underlies all human interaction. Hence, it is crucial for building social robots that are expected to collaborate effectively with humans. Gaze behavior plays versatile roles in establishing, maintaining and repairing the common ground. Integrating all these roles in a computational dialog model is a complex task since gaze is generally combined with multiple parallel information modalities and involved in multiple processes for the generation and recognition of behavior. Going beyond related work, we present a modeling approach focusing on these multi-modal, parallel and bi-directional aspects of gaze that need to be considered for grounding and their interleaving with the dialog and task management. We illustrate and discuss the different roles of gaze as well as advantages and drawbacks of our modeling approach based on a first user study with a technically sophisticated shared workspace application with a social humanoid robot.
Gregor Mehlmann, Markus Häring, Kathrin Janowski, Tobias Baur 0001, Patrick Gebhard, Elisabeth André
ICMI6
2014 Exploring social augmentation concepts for public speaking using peripheral feedback and real-time behavior analysis
abstract
Non-verbal and unconscious behavior plays an important role for efficient human-to-human communication but are often undervalued when training people to become better communicators. This is particularly true for public speakers who need not only behave according to a social etiquette but do so while generating enthusiasm and interest for dozens if not hundreds of other persons. In this paper we propose the concept of social augmentation using wearable computing with the goal of giving users the ability to continuously monitor their performance as a communicator. To this end we explore interaction modalities and feedback mechanisms which would lend themselves to this task.
Ionut Damian, Chiew Seng Sean Tan, Tobias Baur 0001, Johannes Schöning, Kris Luyten, Elisabeth André
ISMAR6
2014 Simulating Deceptive Cues of Joy in Humanoid Robots
Birgit Lugrin, Markus Häring, Gasser Akila, Elisabeth André
IVA4
2014 Full Body Interaction with Virtual Characters in an Interactive Storytelling Scenario
Felix Kistler, Birgit Lugrin, Elisabeth André
IVA3
2014 An Event Driven Fusion Approach for Enjoyment Recognition in Real-time
abstract
Social signals and interpretation of carried information is of high importance in Human Computer Interaction. Often used for affect recognition, the cues within these signals are displayed in various modalities. Fusion of multi-modal signals is a natural and interesting way to improve automatic classification of emotions transported in social signals. Throughout most present studies, uni-modal affect recognition as well as multi-modal fusion, decisions are forced for fixed annotation segments across all modalities. In this paper, we investigate the less prevalent approach of event driven fusion, which indirectly accumulates asynchronous events in all modalities for final predictions. We present a fusion approach, handling short-timed events in a vector space, which is of special interest for real-time applications. We compare results of segmentation based uni-modal classification and fusion schemes to the event driven fusion approach. The evaluation is carried out via detection of enjoyment-episodes within the audiovisual Belfast Story-Telling Corpus.
Florian Lingenfelser, Johannes Wagner 0001, Elisabeth André, Gary McKeown, William Curran
ACM Multimedia3
2014 Trust-Based Decision-Making for Energy-Aware Device Management
Stephan Hammer, Michael Wissner, Elisabeth André
UMAP3
2014 Who's Afraid of Job Interviews? Definitely a Question for User Modelling
Kaska Porayska-Pomsta, Paola Rizzo, Ionut Damian, Tobias Baur 0001, Elisabeth André, Nicolas Sabouret, Hazaël Jones, Keith Anderson, Evi Chryssafidou
UMAP5
2014 Designing User-Character Dialog in Interactive Narratives: An Exploratory Experiment
abstract
Through interaction with the virtual environment and virtual characters, users are able to influence the storyline of many games. The design choice for the style of interactivity can thereby have a crucial influence on the user's experience. However, only a few approaches evaluate different interaction modalities for one system to investigate the impact of design choice on the users' experience. In this paper, we present an experimental approach in which we first reflect on design alternatives concerning a specific element of interactive narratives-user-character dialog-and then investigate user responses to different design options (round-based dialog versus continuous dialog). Results of an experimental evaluation study show that users tend to prefer continuous interaction in a soap-opera-like game environment using typed text input to communicate with virtual characters that act and react using speech output, although the recognition rate of user utterances of the continuous version was slightly worse compared to the round-based version.
Birgit Lugrin, Christoph Klimmt, Gregor Mehlmann, Elisabeth André, Christian Roth 0001
IEEE Trans. Comput. Intell. AI Games4
2013 The TARDIS Framework: Intelligent Virtual Agents for Social Coaching in Job Interviews
Keith Anderson, Elisabeth André, Tobias Baur 0001, Sara Bernardini, Mathieu Chollet, Evi Chryssafidou, Ionut Damian, Cathy Ennis, Arjan Egges, Patrick Gebhard, Hazaël Jones, Magalie Ochs, Catherine Pelachaud, Kaska Porayska-Pomsta, Paola Rizzo, Nicolas Sabouret
Advances in Computer Entertainment2
2013 Question Generation and Adaptation Using a Bayesian Network of the Learner's Achievements
Michael Wissner, Floris Linnebank, Jochem Liem, Bert Bredeweg, Elisabeth André
AIED5
2013 Traveller: Interacting with agents to deal with misunderstandings due to culture
Nick Degens, Gert Jan Hofstede, Samuel Mascarenhas, Ana Paiva 0001, André Silva 0001, Felix Kistler, Elisabeth André, Arvid Kappas, Ruth Aylett
FDG7
2013 User-Defined Body Gestures for an Interactive Storytelling Scenario
Felix Kistler, Elisabeth André
INTERACT (2)2
2013 Traveller: An Interactive Cultural Training System Controlled by User-Defined Body Gestures
Felix Kistler, Elisabeth André, Samuel Mascarenhas, André Silva 0001, Ana Paiva 0001, Nick Degens, Gert Jan Hofstede, Eva Krumhuber, Arvid Kappas, Ruth Aylett
INTERACT (4)2
2013 Using phonetic patterns for detecting social cues in natural conversations
abstract
Laughter and fillers like “uhm” and “ah” are social cues expressed in human speech. Detection and interpretation of such non-linguistic events can reveal important information about the speakers’ intensions and emotional state. The INTERSPEECH 2013 Social Signals Sub-Challenge sets the task to localize and classify laughter and fillers in the “SSPNet Vocalization Corpus” (SVC) based on acoustics. In the paper at hand we investigate phonetic patterns extracted from raw speech transcriptions obtained with the CMU Sphinx toolkit for speech recognition. Even though Sphinx was used out of the box and no dedicated training on the target classes was applied, we were able to successfully predict laughter and filler frames in the development set with ∼ 87% accuracy (unweighted average Area Under the Curve (AUC)). By accumulating our features with a set of standard features provided by the challenge organizers results increased above 92%. When applying the combined set to the test corpus we achieved 87.7% as highest score, which is 4.4% above the challenge baseline.
Johannes Wagner 0001, Florian Lingenfelser, Elisabeth André
INTERSPEECH3
2013 Motion capturing empowered interaction with a virtual agent in an Augmented Reality environment
abstract
We present an Augmented Reality (AR) system where we immerse the user's whole body in the virtual scene using a motion capturing (MoCap) suit. The goal is to allow for seamless interaction with the virtual content within the AR environment. We describe an evaluation study of a prototype application featuring an interactive scenario with a virtual agent. The scenario contains two conditions: in one, the agent has access to the full tracking data of the MoCap suit and therefore is aware of the exact actions of the user, while in the second condition, the agent does not get this information. We then report and discuss the differences we were able to detect regarding the users' perception of the interaction with the agent and give future research directions.
Ionut Damian, René Bühling, Felix Kistler, Mark Billinghurst, Mohammad Obaid, Elisabeth André
ISMAR6
2013 Time-Pie visualization: Providing Contextual Information for Energy Consumption Data
abstract
In recent years a growing number of information visualization systems have been developed to assist users with monitoring their energy consumption, with the hope of reducing energy use through more effective user-awareness. Most of these visualizations can be categorized into either some form of a time-series or pie chart, each with their own limitations. These visualization systems also often ignore incorporating contextual (e.g. weather, environmental) information which could assist users with better interpretation of their energy use information. In this paper we introduce the time-pie visualization technique, which combines the concepts of timeseries and pie charts, and allows the addition of contextual information to energy consumption data.
Masood Masoodian, Birgit Lugrin, René Bühling, Pavel Ermolin, Elisabeth André
IV5
2013 The social signal interpretation (SSI) framework: multimodal signal processing and recognition in real-time
abstract
Automatic detection and interpretation of social signals carried by voice, gestures, mimics, etc. will play a key-role for next-generation interfaces as it paves the way towards a more intuitive and natural human-computer interaction. The paper at hand introduces Social Signal Interpretation (SSI), a framework for real-time recognition of social signals. SSI supports a large range of sensor devices, filter and feature algorithms, as well as, machine learning and pattern recognition tools. It encourages developers to add new components using SSI's C++ API, but also addresses front end users by offering an XML interface to build pipelines with a text editor. SSI is freely available under GPL at http://openssi.net.
Johannes Wagner 0001, Florian Lingenfelser, Tobias Baur 0001, Ionut Damian, Felix Kistler, Elisabeth André
ACM Multimedia6
2013 Trust-based decision-making for the adaptation of public displays in changing social contexts
abstract
Public displays may adapt intelligently to the social context, tailoring information on the screen, for example, to the profiles of spectators, their gender or based on their mutual proximity. However, such adaptation decisions should on the one hand match user preferences and on the other maintain the user's trust in the system. A wrong decision can negatively influence the user's acceptance of a system, cause frustration and, as a result, make users abandon the system. In this paper, we propose a trust-based mechanism for automatic decision-making, which is based on Bayesian Networks. We present the process of network construction, initialization with empirical data, and validation. The validation demonstrates that the mechanism generates accurate decisions on adaptation which match user preferences and support user trust.
Ekaterina Kurdjokova, Michael Wissner, Stephan Hammer, Elisabeth André
PST4
2013 Investigating the influence of culture on proxemic behaviors for humanoid robots
abstract
In social robotics, the behavior of humanoid robots is intended to be designed in a way that they behave in a human-like manner and serve as natural interaction partners for human users. Several aspects of human behavior such as speech, gestures, eye-gaze as well as the personal and social background of the user need therefore to be considered. In this paper, we investigate interpersonal distance as a behavioral aspect that varies with the cultural background of the user. We present two studies that explore whether users of different cultures (Arabs and Germans) expect robots to behave similar to their own cultural background. The results of the first study reveal that Arabs and Germans have different expectations on the interpersonal distance between themselves and robots in a static setting. In the second study, we use the results of the first study to investigate the users' reactions on robots using the observed interpersonal distances themselves. Although the data of this dynamic setting is not conclusive, it suggests that users prefer robots that show behavior that has been observed for their own cultural background before.
Ghadeer Eresha, Markus Häring, Birgit Lugrin, Elisabeth André, Mohammad Obaid
RO-MAN4
2013 TabletopCars: interaction with active tangible remote controlled cars
abstract
In this paper, we report on the development of the competitive tangible tabletop game TabletopCars, which combines the virtual world with the physical world. We brought together micro scaled radio controlled cars as active tangibles with an interactive tabletop surface to realize the game. Furthermore, we included Microsoft Kinect depth sensing as an interaction mode for embedded and embodied interaction. Our aim was to investigate the possibilities that emerge through the augmentation capabilities of interactive tabletops for creating novel game concepts and the interaction modes that novel input devices facilitate. This work presents TabletopCars as a testbed for embedded and embodied interaction and describes the system in detail. Finally, we report on a preliminary user study where users controlled the active tangible micro scaled cars through hand gestures.
Chi Tai Dang, Elisabeth André
TEI2
2013 Modelling Users' Affect in Job Interviews: Technological Demo
Kaska Porayska-Pomsta, Keith Anderson, Ionut Damian, Tobias Baur 0001, Elisabeth André, Sara Bernardini, Paola Rizzo
UMAP5
2013 Investigating culture-related aspects of behavior for virtual characters
Birgit Lugrin, Elisabeth André, Matthias Rehm, Yukiko I. Nakano
Auton. Agents Multi Agent Syst.2
2013 Introduction to the special section on eye gaze and conversation
abstract
This editorial introduction first explains the origin of this special section. It then outlines how each of the two articles included sheds light on possibilities for conversational dialog systems to use eye gaze as a signal that reflects aspects of participation in the dialog: degree of engagement and turn taking behavior, respectively.
Elisabeth André, Joyce Y. Chai
ACM Trans. Interact. Intell. Syst.1
2013 Exploiting unconscious user signals in multimodal human-computer interaction
abstract
This article presents the idea of empathic stimulation that relies on the power and potential of unconsciously conveyed attentive and emotional information to facilitate human-machine interaction. Starting from a historical review of related work presented at past ACM Multimedia conferences, we discuss challenges that arise when exploiting unconscious human signals for empathic stimulation, such as the real-time analysis of psychological user states and the smooth adaptation of the human-machine interface based on this analysis. A classical application field that might benefit from the idea of unconscious human-computer interaction is the exploration of massive datasets.
Elisabeth André
ACM Trans. Multim. Comput. Commun. Appl.1
2012 City Pulse: Supporting Going-Out Activities with a Context-Aware Urban Display
Mohammad Obaid, Ekaterina Kurdyukova, Elisabeth André
Advances in Computer Entertainment3
2012 Is stereoscopic 3D a better choice for information representation in the car?
abstract
In modern cars users need to interact with safety and comfort functions, driver assistance systems, and infotainment devices. Basic requirements include the perception of the current status and of information items as well as the control of functions. Handling that myriad amount of information while driving requires an appropriate interaction design, structure and visualization of the data. This paper investigates potentials and limitations of stereoscopic 3D for visualizing an in-vehicle information system. We developed a spatial in-car visualization concept that exploits three dimensions for the system's output. Based on a prototype, that implements the central functionality of our concept, we evaluate the 3D representation. A laboratory study with 32 users indicates that stereoscopic 3D is the better choice as it improves the user experience, increases the attractiveness, and helps the user in recognizing the current state of the system. The study shows no significant differences between non-stereoscopic and stereoscopic representations in the users' workload. This indicates that stereoscopic visualizations have no negative impact on the primary driving task.
Nora Broy, Elisabeth André, Albrecht Schmidt 0001
AutomotiveUI2
2012 Modeling multimodal integration with event logic charts
abstract
In this paper we present a novel approach to the combined modeling of multimodal fusion and interaction management. The approach is based on a declarative multimodal event logic that allows the integration of inputs distributed over multiple modalities in accordance to spatial, temporal and semantic constraints. In conjunction with a visual state chart language, our approach supports the incremental parsing and fusion of inputs and a tight coupling with interaction management. The incremental and parallel parsing approach allows us to cope with concurrent continuous and discrete interactions and fusion on different levels of abstraction. The high-level visual and declarative modeling methods support rapid prototyping and iterative development of multimodal systems.
Gregor Mehlmann, Elisabeth André
ICMI2
2012 A Frame Pruning Approach for Paralinguistic Recognition Tasks
abstract
In conventional paralinguistic classification approaches, information gained by low level features is described over broad segments (like whole turns) via statistical functionals. This procedure presumes meaningful information to be embodied within the whole segment. This assumption may be misleading if distinctive cues within a sample are surrounded by non-meaningful information or noise. In this case it would surely be beneficial to keep only parts of the sample that are most relevant for the recognition task. In this paper we propose a novel cluster-based approach, which aims at identifying frames likely to carry distinctive information. Evaluation is done within the INTERSPEECH 2012 Speaker Trait Challenge. Results show that under certain configurations frame pruning in fact leads to an improvement in recognition accuracy. On the observed corpus most stable improvements were achieved at a frame drop of 4-8%. Index Terms: paralinguistic recognition, frame pruning, personality traits
Johannes Wagner 0001, Florian Lingenfelser, Elisabeth André
INTERSPEECH3
2012 Studying user-defined iPad gestures for interaction in multi-display environment
abstract
The paper investigates the iPad gestures that users naturally perform for data transfer. We examine the transfer between two iPads, iPad and a tabletop, and iPad and a public display. Three gesture modalities are investigated: multi-touch gestures, performed using iPad display, spatial gestures, performed by manipulating iPad in 3D space, and direct contact gestures, involving the physical contact of iPad and other device. We report on user choices of the modalities and gesture types, and derive critical points for the design of iPad gestures.
Ekaterina Kurdyukova, Matthias Redlin, Elisabeth André
IUI3
2012 Cultural Behaviors of Virtual Agents in an Augmented Reality Environment
Mohammad Obaid, Ionut Damian, Felix Kistler, Birgit Lugrin, Johannes Wagner 0001, Elisabeth André
IVA6
2012 Mobile augmented reality and adaptive art: a game-based motivation for energy saving
abstract
We present the design of an educational treasure hunt game that uses mobile Augmented Reality (AR) and adaptive virtual gardens to raise awareness on energy consumption ways. Within the game, AR is used to present 3D content that allows for visually understanding the energy consumption problems and their solutions. In addition, the players' performances are given in a form of a visual feedback as dynamic virtual gardens that change from poor to good status. Initial tests show that the presented game is highly appealing to players and motivated them to be aware of the presented problem using AR technologies and the visual effects of their virtual garden's health. Future work will focus on evaluating the learnability and engagement of players.
René Bühling, Mohammad Obaid, Stephan Hammer, Elisabeth André
MUM4
2012 Direct, bodily or mobile interaction?: comparing interaction techniques for personalized public displays
abstract
Interaction with personalized data on a large public display represents a sensitive scenario: first, users expose the fact of interaction in public; and, personalized data may be private. In this work we investigate how interaction design can support the user in such a scenario. Through experimentation, we compare three interaction techniques: direct, bodily, and mobile-based. We report on the users' preferences with the presented techniques at different interaction phases (identification, navigation, and collecting results). We analyze how user preferences in the personalized display scenario are similar or different to other scenarios, such as interaction with physical objects or non-personalized public displays. The analysis is summarized in a form of design recommendations that should be considered when designing for interaction with personalized public displays.
Ekaterina Kurdyukova, Mohammad Obaid, Elisabeth André
MUM3
2012 Introduction to the special issue on eye gaze in intelligent human-machine interaction
abstract
Given the recent advances in eye tracking technology and the availability of nonintrusive and high-performance eye tracking devices, there has never been a better time to explore new opportunities to incorporate eye gaze in intelligent and natural human-machine communication. In this special issue, we present six articles that cover various aspects of eye gaze in human-machine interaction, including applications of gaze tracking in human-machine interaction, techniques that recognize gaze gestures and render gaze behaviors, and the analysis of gaze behaviors in social interactions.
Elisabeth André, Joyce Y. Chai
ACM Trans. Interact. Intell. Syst.1
2011 Character Roles and Interaction in the DynaLearn Intelligent Learning Environment
Michael Wissner, Wouter Beek, Esther Lozano, Gregor Mehlmann, Floris Linnebank, Jochem Liem, Markus Häring, René Bühling, Jorge Gracia, Bert Bredeweg, Elisabeth André
AIED11
2011 A cooperative in-car game for heterogeneous players
abstract
Car rides are often perceived as dull by the passengers, especially children. Therefore, we aim to introduce a system fostering a collaborative and communicative experience in this environment. This paper presents the design for a game played together by all car-occupants, including the driver, according to their abilities and capacities. A fully implemented prototype of our system called nICE: nice In-Car Experience is evaluated under real world conditions in a user study with five families using a qualitative approach.
Nora Broy, Sebastian Goebl, Matheus Hauder, Thomas Kothmayr, Michael Kugler, Florian Reinhart, Martin Salfer, Kevin Schlieper, Elisabeth André
AutomotiveUI9
2011 Adaptive Art - A Shape Language Driven Approach to Communicate Dramaturgy and Mood
René Bühling, Emilie Brihi, Michael Wissner, Elisabeth André
ICIDS4
2011 Exploration of User Reactions to Different Dialog-Based Interaction Styles
Birgit Lugrin, Christoph Klimmt, Gregor Mehlmann, Elisabeth André, Christian Roth 0001
ICIDS4
2011 Full Body Gestures Enhancing a Game Book for Interactive Story Telling
Felix Kistler, Dominik Sollfrank, Nikolaus Bee, Elisabeth André
ICIDS4
2011 A systematic discussion of fusion techniques for multi-modal affect recognition tasks
abstract
Recently, automatic emotion recognition has been established as a major research topic in the area of human computer interaction (HCI). Since humans express emotions through various channels, a user's emotional state can naturally be perceived by combining emotional cues derived from all available modalities. Yet most effort has been put into single-channel emotion recognition, while only a few studies with focus on the fusion of multiple channels have been published. Even though most of these studies apply rather simple fusion strategies -- such as the sum or product rule -- some of the reported results show promising improvements compared to the single channels. Such results encourage investigations if there is further potential for enhancement if more sophisticated methods are incorporated. Therefore we apply a wide variety of possible fusion techniques such as feature fusion, decision level combination rules, meta-classification or hybrid-fusion. We carry out a systematic comparison of a total of 16 fusion methods on different corpora and compare results using a novel visualization technique. We find that multi-modal fusion is in almost any case at least on par with single channel classification, though homogeneous results within corpora point to interchangeability between concrete fusion schemes.
Florian Lingenfelser, Johannes Wagner 0001, Elisabeth André
ICMI3
2011 Modeling parallel state charts for multithreaded multimodal dialogues
abstract
In this paper, we present a modeling approach for the management of highly interactive, multithreaded and multimodal dialogues. Our approach enforces the separation of dialogue content and dialogue structure and is based on a statechart language enfolding concepts for hierarchy, concurrency, variable scoping and a detailed runtime history. These concepts facilitate the modeling of interactive dialogues with multiple virtual characters, autonomous and parallel behaviors, flexible interruption policies, context-sensitive interpretation of the user's discourse acts and coherent resumptions of dialogues. An interpreter allows the realtime visualization and modification of the model to allow a rapid prototyping and easy debugging. Our approach has successfully been used in applications and research projects as well as evaluated in field tests with non-expert authors. We present a demonstrator illustrating our concepts in a social game scenario.
Gregor Mehlmann, Birgit Lugrin, Elisabeth André
ICMI3
2011 Usage and Recognition of Finger Orientation for Multi-Touch Tabletop Interaction
Chi Tai Dang, Elisabeth André
INTERACT (3)2
2011 The Social Signal Interpretation Framework (SSI) for Real Time Signal Processing and Recognition
abstract
The construction of systems for recording, processing and recognising a human's social and affective signals is a challenging effort that includes numerous but necessary sub-tasks to be dealt with.In this article, we introduce our Social Signal Interpretation (SSI) tool, a framework dedicated to support the development of such systems.It provides a flexible architecture to construct pipelines to handle multiple modalities like audio or video and establishing on-and offline recognition tasks.The plug-in system of SSI encourages developers to integrate external code, while a XML interface allows anyone to write own applications with a simple text editor.Furthermore, data recording, annotation and classification can be done using a straightforward graphical user interface, allowing simple access to inexperienced users.
Johannes Wagner 0001, Florian Lingenfelser, Elisabeth André
INTERSPEECH3
2011 A Software Framework for Individualized Agent Behavior
Ionut Damian, Birgit Lugrin, Nikolaus Bee, Elisabeth André
IVA4
2011 Culture-Related Topic Selection in Small Talk Conversations across Germany and Japan
Birgit Lugrin, Yukiko I. Nakano, Afia Akhter Lipi, Matthias Rehm, Elisabeth André
IVA5
2011 Individualized Agent Interactions
Ionut Damian, Birgit Lugrin, Peter Huber, Nikolaus Bee, Elisabeth André
MIG5
2011 Creation and Evaluation of emotion expression with body movement, sound and eye color for humanoid robots
abstract
The ability to display emotions is a key feature in human communication and also for robots that are expected to interact with humans in social environments. For expressions based on Body Movement and other signals than facial expressions, like Sound, no common grounds have been established so far. Based on psychological research on human expression of emotions and perception of emotional stimuli we created eight different expressional designs for the emotions Anger, Sadness, Fear and Joy, consisting of Body Movements, Sounds and Eye Colors. In a large pre-test we evaluated the recognition ratios for the different expressional designs. In our main experiment we separated the expressional designs into their single cues (Body Movement, Sound, Eye Color) and evaluated their expressivity. The detailed view at the perception of our expressional cues, allowed us to evaluate the appropriateness of the stimuli, check our implementations for flaws and build a basis for systematical revision. Our analysis revealed that almost all Body Movements were appropriate for their target emotion and that some of our Sounds need a revision. Eye Colors could be identified as an unreliable component for emotional expression.
Markus Häring, Nikolaus Bee, Elisabeth André
RO-MAN3
2011 Planning Small Talk behavior with cultural influences for multiagent systems
Birgit Lugrin, Matthias Rehm, Elisabeth André
Comput. Speech Lang.3
2011 Exploring Fusion Methods for Multimodal Emotion Recognition with Missing Data
abstract
The study at hand aims at the development of a multimodal, ensemble-based system for emotion recognition. Special attention is given to a problem often neglected: missing data in one or more modalities. In offline evaluation the issue can be easily solved by excluding those parts of the corpus where one or more channels are corrupted or not suitable for evaluation. In real applications, however, we cannot neglect the challenge of missing data and have to find adequate ways to handle it. To address this, we do not expect examined data to be completely available at all time in our experiments. The presented system solves the problem at the multimodal fusion stage, so various ensemble techniques-covering established ones as well as rather novel emotion specific approaches-will be explained and enriched with strategies on how to compensate for temporarily unavailable modalities. We will compare and discuss advantages and drawbacks of fusion categories and extensive evaluation of mentioned techniques is carried out on the CALLAS Expressivity Corpus, featuring facial, vocal, and gestural modalities.
Johannes Wagner 0001, Elisabeth André, Florian Lingenfelser, Jonghwa Kim 0001
IEEE Trans. Affect. Comput.2
2010 Trustworthy Organic Computing Systems: Challenges and Perspectives
Jan-Philipp Steghöfer, Rolf Kiefhaber, Karin Bee, Yvonne Bernard, Lukas Klejnowski, Wolfgang Reif, Theo Ungerer, Elisabeth André, Jörg Hähner, Christian Müller-Schloer
ATC8
2010 Age and gender classification from speech using decision level fusion and ensemble based techniques
abstract
In this contribution to INTERSPEECH 2010 Paralinguistic Challenge we explore the capabilities of decision level fusion and ensemble based techniques for classification tasks on the provided AGENDER corpus.Ensemble members are generated by providing multiple feature sets generated by feature selection, and novel fusion methods (developed in order to give special support to under-represented classes) are applied for decision making.Results are compared to standard classification approaches and possible benefits are discussed.
Florian Lingenfelser, Johannes Wagner 0001, Thurid Vogt, Jonghwa Kim 0001, Elisabeth André
INTERSPEECH5
2010 Workshop: eye gaze in intelligent human machine interaction
abstract
This workshop brought researchers from academia and industry together to share recent advances and discuss research directions and opportunities for next generation of intelligent human machine interaction that incorporate eye gaze.
Elisabeth André, Joyce Y. Chai
IUI1
2010 Bossy or Wimpy: Expressing Social Dominance by Combining Gaze and Linguistic Behaviors
Nikolaus Bee, Colin Pollock, Elisabeth André, Marilyn A. Walker
IVA3
2010 Generating Culture-Specific Gestures for Virtual Agent Dialogs
Birgit Lugrin, Ionut Damian, Peter Huber, Matthias Rehm, Elisabeth André
IVA5
2010 Level of Detail Based Behavior Control for Virtual Characters
Felix Kistler, Michael Wissner, Elisabeth André
IVA3
2010 Multiple Agent Roles in an Adaptive Virtual Classroom Environment
Gregor Mehlmann, Markus Häring, René Bühling, Michael Wissner, Elisabeth André
IVA5
2010 Level of Detail AI for Virtual Characters in Games and Simulation
Michael Wissner, Felix Kistler, Elisabeth André
MIG3
2010 Trust-centered design for multi-display applications
abstract
This work describes a trust-centered user study that was conducted during the design process of a multi-display ubiquitous application. The objective of the study was to find out how the adaptation of the displays should be designed in order to protect user trust. The study was conducted in the form of focus group interviews; it investigated user attitude towards data and events that can be seen as trust-critical in a multi-display interaction scenario. The quantitative results, along with user comments and discussions, provide an interesting insight how the system should be adapted in order to preserve user trust.
Ekaterina Kurdyukova, Elisabeth André, Karin Bee
MoMM2
2010 Managing user trust for self-adaptive ubiquitous computing systems
abstract
Ubiquitous computing systems can cause serious problems for user trust. In particular if the system is self-adaptive and situations appear which are poorly self-explanatory. In this paper we aim at the trust management of adaptive systems. We present a user study that covers the correlation of trust dimensions and user feelings on user trust. As results of this study, a Bayesian Network is introduced that, at the design time and runtime of the system, provides knowledge about the interplay between a truster's disposition, system events and actions, trust dimensions, user trust and user response.
Karin Bee, Elisabeth André, Ekaterina Kurdyukova
MoMM2
2010 MED-StyleR: METABO diabetes-lifestyle recommender
abstract
Lifestyle plays an essential role in controlling diabetes and in both the prevention and management of diabetes. Many reports from clinical research support the theory that healthy eating and regular exercise are much more effective at managing diabetes than traditional medication. In this paper we introduce an innovative approach to the multimodal recommender system conceived in the EU METABO project. The most important feature of the METABO Diabetes-Lifestyle Recommender (MED-StyleR) is to generate highly personalized recommendations that satisfy medical prescriptions for patients' long-term health alongside and short-term preferences of patients in their daily lives.
Stephan Hammer, Jonghwa Kim 0001, Elisabeth André
RecSys3
2010 Exploring the usability of immersive interactive storytelling
abstract
The Entertainment potential of Virtual Reality is yet to be fully realised.In recent years, this potential has been described through the Holodeck™ metaphor, without however addressing the issue of content creation and gameplay.Recent progress in Interactive Narrative technology makes it possible to envision immersive systems.Yet, little is known about the usability of such systems or which paradigms should be adopted for gameplay and interaction.We report user experiments carried out with a fully immersive Interactive Narrative system based on a CAVE-like system, which explore two interactivity paradigms for user involvement (Actor and Ghost).Our results confirm the potential of immersive Interactive Narratives in terms of performance but also of user acceptance.
Jean-Luc Lugrin, Marc Cavazza, David Pizzi, Thurid Vogt, Elisabeth André
VRST5
2010 Editorial
abstract
This volume of Computer Animation and Virtual Worlds (CAVW) contains a selection of papers submitted to CASA 2010, the 23rd International Conference on Computer Animation and Social Agents. CASA is one of the premier international conferences in the field of computer animation and social agents, organized under the auspices of the Computer Graphics Society (CGS). It has been founded in 1988, and, over the last years, it has been organized in Europe: Geneva (2002, 2004, 2006), Hasselt (2007), Amsterdam (2009); in USA: Philadelphia (1998, 2000), New Jersey (2003); and in Asia: Seoul (2001, 2008), Hong Kong (2005). This year, CASA 2010 was organized in Saint-Malo, France from the 30th of May to the 2nd of June 2010. The organization was done by Bunraku, an INRIA Project-team in common with CNRS, INSA of Rennes, University of Rennes 1, and Ecole Normale Supérieure de Cachan. The CASA 2010 edition received 104 submissions from 27 countries and 6 continents. Each submission received at least 3 reviews and 32 among them were selected to appear in this special issue of Computer Animation and Virtual Worlds. We thank all the authors who have submitted their work to this conference allowing us to present this nice and diversified program. We thank also the International Program Committee members and the additional external reviewers for the time and energy they have invested in the reviewing process. A particular thank goes to the INRIA conference support team, especially to Edith Blin-Guyot and Steeve Tessier, for their support in organizing the conference, taking care of financial, material, and organizational matters. CASA 2010 has been sponsored by INRIA, by the IRIS European Network of Excellence (Integrating Research in Interactive Storytelling), by GDR IG (Groupement de Recherche Informatique Graphique), by Fondation Michel Métivier, by the Brittany Regional Council, by the University of Rennes 1, by the Ecole Normale Supérieure de Cachan, and by the Biometrics Company. The 32 papers presented in this special issue are divided into several categories: Cartoon and Sketch-based animation techniques, Stylized animation, Deformable models, Meshes, Physically based animation, Motion Analysis and Synthesis, Steering and Crowds, Facial expression, Social Agents, and finally Augmented Reality.
Stéphane Donikian, Elisabeth André, Shi-Min Hu 0001, Daniel Thalmann
Comput. Animat. Virtual Worlds2
2009 What Would You Do in Their Shoes? Experiencing Different Perspectives in an Interactive Drama for Multiple Users
Birgit Lugrin, Michael Boegler, Nikolaus Bee, Elisabeth André
ICIDS4
2009 Introducing Multiple Interaction Devices to Interactive Storytelling: Experiences from Practice
Ekaterina Kurdyukova, Elisabeth André, Karin Bee
ICIDS2
2009 Exploring the benefits of discretization of acoustic features for speech emotion recognition
abstract
We present a contribution to the Open Performance subchallenge of the INTERSPEECH 2009 Emotion Challenge.We evaluate the feature extraction and classifier of EmoVoice, our framework for real-time emotion recognition from voice on the challenge database and achieve competitive results.Furthermore, we explore the benefits of discretizing numeric acoustic features and find it beneficial in a multi-class task.
Thurid Vogt, Elisabeth André
INTERSPEECH2
2009 Simplified facial animation control utilizing novel input devices: a comparative study
abstract
Editing facial expressions of virtual characters is quite a complex task. The face is made up of many muscles, which are partly activated concurrently. Virtual faces with human expressiveness are usually designed with a limited amount of facial regulators. Such regulators are derived from the facial muscle parts that are concurrently activated. Common tools for editing such facial expressions use slider-based interfaces where only a single input at a time is possible. Novel input devices, such as gamepads or data gloves, which allow parallel editing, could not only speed up editing, but also simplify the composition of new facial expressions. We created a virtual face with 23 facial controls and connected it with a slider-based GUI, a gamepad, and a data glove. We first conducted a survey with professional graphics designers to find out how the latter two new input devices would be received in a commercial context. A second comparative study with 17 subjects was conducted to analyze the performance and quality of these two new input devices using subjective and objective measurements.
Nikolaus Bee, Bernhard Falk, Elisabeth André
IUI3
2009 Breaking the Ice in Human-Agent Communication: Eye-Gaze Based Initiation of Contact with an Embodied Conversational Agent
Nikolaus Bee, Elisabeth André, Susanne Tober
IVA2
2009 Studying multi-user settings for pervasive games
abstract
Whenever a pervasive game has to be developed for a group of children an appropriate multi-user setting has to be found.If the pervasive game does not support the children with an adequate multi-user setting, unintended situations can emerge, such as a single user can dominate the game while the other users are bored and disinterested.In our research we approach that problem by investigating various multiuser settings that are characterized by a different distribution of interaction devices.We describe three multi-user settings, a pervasive game which we used as a test bed, and a user study with 18 children to find out how the multiuser settings influence the children's social behaviour as expressed by the level of activity for all group members, the offtask behaviour and the level of task-related conversations.
Karin Bee, Elisabeth André
Mobile HCI2
2009 Introduction to the special issue on the Fourth German Conference on Multiagent System Technologies (MATES)
Klaus Fischer 0001, Ingo J. Timm, Elisabeth André
Auton. Agents Multi Agent Syst.3
2008 User-Centred Development of Mobile Interfaces to a Pervasive Computing Environment
abstract
A challenging issue for HCI is the development of usable mobile interfaces for interactions with a complex pervasive environment. We consider a need for interfaces which automatically adapt their interaction and presentation capabilities on the user's situational needs and expectations to decrease the complexity of the environment and increase the usability of the system. Therefore, a rule-set is required which gives knowledge on the mobile interface's adaptations as a consequence on a user's situations within the environment. This rule-set iteratively emerges within a user-centred development process by considering and testing each contextual situation of the user when interacting with the mobile interface. In this paper we describe an approach of a usage model for specifying each context of the user and the environment as well as the user's goals and mental model. Moreover, we describe our used user-centred process to develop the usage model and rule-set, practical experience in development of mobile interfaces, some guidelines and our planned future work.
Karin Bee, Elisabeth André
ACHI2
2008 Exploring emotions and multimodality in digitally augmented puppeteering
abstract
Recently, multimodal and affective technologies have been adopted to support expressive and engaging interaction, bringing up a plethora of new research questions. Among the challenges, two essential topics are 1) how to devise truly multimodal systems that can be used seamlessly for customized performance and content generation, and 2) how to utilize the tracking of emotional cues and respond to them in order to create affective interaction loops. We present PuppetWall, a multi-user, multimodal system intended for digitally augmented puppeteering. This application allows natural interaction to control puppets and manipulate playgrounds comprising background, props, and puppets. PuppetWall utilizes hand movement tracking, a multi-touch display and emotion speech recognition input for interfacing. Here we document the technical features of the system and an initial evaluation. The evaluation involved two professional actors and also aimed at exploring naturally emerging expressive speech categories. We conclude by summarizing challenges in tracking emotional cues from acoustic features and their relevance for the design of affective interactive systems.
Lassi A. Liikkanen, Giulio Jacucci, Eero Huvio, Toni Laitinen, Elisabeth André
AVI5
2008 Bi-channel sensor fusion for automatic sign language recognition
abstract
In this paper, we investigate the mutual-complementary functionality of accelerometer (ACC) and electromyogram (EMG) for recognizing seven word-level sign vocabularies in German sign language (GSL). Results are discussed for the single channels and for feature-level fusion for the bichannel sensor data. For the subject-dependent condition, this fusion method proves to be effective. Most relevant features for all subjects are extracted and their universal effectiveness is proven with a high average accuracy for the single subjects. Additionally, results are given for the subject-independent condition, where subjective differences do not allow for high recognition rates. Finally we discuss a problem of feature-level fusion caused by high disparity between accuracies of each single channel classification.
Jonghwa Kim 0001, Johannes Wagner 0001, Matthias Rehm, Elisabeth André
FG4
2008 The IRIS Network of Excellence: Integrating Research in Interactive Storytelling
Marc Cavazza, Stéphane Donikian, Marc Christie, Ulrike Spierling, Nicolas Szilas, Peter Vorderer, Tilo Hartmann, Christoph Klimmt, Elisabeth André, Ronan Champagnat, Paolo Petta, Patrick Olivier
ICIDS9
2008 EMG-based hand gesture recognition for realtime biosignal interfacing
abstract
In this paper the development of an electromyogram (EMG) based interface for hand gesture recognition is presented. To recognize control signs in the gestures, we used a single channel EMG sensor positioned on the inside of the forearm. In addition to common statistical features such as variance, mean value, and standard deviation, we also calculated features from the time and frequency domain including Fourier variance, region length, zerocrosses, occurrences, etc. For realizing real-time classification assuring acceptable recognition accuracy, we combined two simple linear classifiers (k-NN and Bayes) in decision level fusion. Overall, a recognition accuracy of 94% was achieved by using the combined classifier with a selected feature set. The performance of the interfacing system was evaluated through 40 test sessions with 30 subjects using an RC Car. Instead of using a remote control unit, the car was controlled by four different gestures performed with one hand. In addition, we conducted a study to investigate the controllability and ease of use of the interface and the employed gestures.
Jonghwa Kim 0001, Stephan Mastnik, Elisabeth André
IUI3
2008 Enculturating conversational interfaces by socio-cultural aspects of communication
abstract
The workshop is centered around three main research challenges: 1.) Computationally viable models of cultural aspects of conversations: Cultural norms and values penetrate all our communications and interactions by giving us heuristics how to behave and how to interpret the verbal and nonverbal behavior of others. To make such a notion like culture available for computation, we need a very specific theory of culture that takes its effects on communication and interaction into account.2.) Reliable empirical data on cultural/cross-cultural interaction: To realize technical systems that take cultural influences on behavior into account, precise data analysis on how this influence manifests itself is necessary. In the literature, this information is often given in very general forms without to the precise data on which the observations are based.3.) Enculturating conversational interfaces: Having identified cultural influences on verbal/nonverbal communicative behaviors, it remains to be shown how this can be applied to the development of human-computer interfaces, for instance in an interface reflecting cultural norms and values of communication.
Matthias Rehm, Elisabeth André, Yukiko I. Nakano, Toyoaki Nishida
IUI2
2008 Creating and Scripting Second Life Bots Using MPML3D
Birgit Lugrin, Helmut Prendinger, Elisabeth André, Mitsuru Ishizuka
IVA3
2008 Cross-Cultural Evaluations of Avatar Facial Expressions Designed by Western Designers
Tomoko Koda, Matthias Rehm, Elisabeth André
IVA3
2008 Culture-Specific First Meeting Encounters between Virtual Agents
Matthias Rehm, Yukiko I. Nakano, Elisabeth André, Toyoaki Nishida
IVA3
2008 EVAL - an evaluation component for mobile interfaces
abstract
The Eval Tool is a usability evaluation environment which can be used to evaluate users an their behaviour while they interact with their pervasive computing environment via a mobile phone interface. The tool supports audio-visual recordings of several users and automatically annotates and synchronizes them with context data emerging from the pervasive environment. The annotation of video material with contextual information is important for the analysis of user studies and the detection of usability issues. The EVAL Tool consists of a recorder and a analyzer component. Using the recorder component the capturing of the videos and the logging of the context can be controlled whereas the analyzer components helps to interpret the user study.
Karin Bee, Dennis Erdmann, Elisabeth André
Mobile HCI3
2008 E-tree: emotionally driven augmented reality art
abstract
In this paper, we describe an Augmented Reality Art installation, which reacts to user behaviour using Multimodal analysis of affective signals. The installation features a virtual tree, whose growth is influenced by the perceived emotional response from spectators. The system implements a 'magic mirror' paradigm (using a large-screen display or projection system) and is based on the ARToolkit with extended representations for scene graphs. The system relies on a PAD dimensional model of affect to support the fusion of different affective modalities, while also supporting the representation of affective responses that relate to aesthetic impressions. The influence of affective input on the visual component is achieved by mapping affective data to an L-System governing virtual tree behaviour. We have performed an early evaluation of the system, both from the technical perspective and in terms of user experience. Post-hoc questionnaires were generally consistent with data from multimodal affective processing, and users rated the overall experience as positive and enjoyable, regardless of how proactive they were in their interaction with the installation.
Stephen W. Gilroy, Marc Cavazza, Rémi Chaignon, Satu-Marja Mäkelä, Markus Niiranen, Elisabeth André, Thurid Vogt, Jérôme Urbain, Mark Billinghurst, Hartmut Seichter, Maurice Benayoun
ACM Multimedia6
2008 Xenakis: combining tangible interaction with probability-based musical composition
abstract
In this paper we present the table-based tangible interface application Xenakis which uses probability models in order to compose music in a way that can be strongly influenced by the user. Our musical sequencing application is based on a framework for tangible interfaces with an architecture that is strongly inspired by the model-view-controller pattern. In addition, we developed a hardware setup for tangible interfaces and used MatraX for tracking markers. The sequencer is the first implementation based on this framework. It allows users to create music simply by moving tangibles on the table. The graphics engine Horde3D is used to visualize the user-interaction and to show the relationships between the tangible objects on the table, creating an appealing audio-visual experience. An evaluation with 37 first time users was conducted in order to discover the strong and the weak points of such tangible user interfaces, especially in the context of our application.
Markus Bischof, Bettina Conradi, Peter Lachenmaier, Kai Linde, Max Meier, Philipp Pötzl, Elisabeth André
TEI7
2008 Emotion Recognition Based on Physiological Changes in Music Listening
abstract
Little attention has been paid so far to physiological signals for emotion recognition compared to audiovisual emotion channels such as facial expression or speech. This paper investigates the potential of physiological signals as reliable channels for emotion recognition. All essential stages of an automatic recognition system are discussed, from the recording of a physiological dataset to a feature-based multiclass classification. In order to collect a physiological dataset from multiple subjects over many weeks, we used a musical induction method which spontaneously leads subjects to real emotional states, without any deliberate lab setting. Four-channel biosensors were used to measure electromyogram, electrocardiogram, skin conductivity and respiration changes. A wide range of physiological features from various analysis domains, including time/frequency, entropy, geometric analysis, subband spectra, multiscale entropy, etc., is proposed in order to find the best emotion-relevant features and to correlate them with emotional states. The best features extracted are specified in detail and their effectiveness is proven by classification results. Classification of four musical emotions (positive/high arousal, negative/high arousal, negative/low arousal, positive/low arousal) is performed by using an extended linear discriminant analysis (pLDA). Furthermore, by exploiting a dichotomic property of the 2D emotion model, we develop a novel scheme of emotion-specific multilevel dichotomous classification (EMDC) and compare its performance with direct multiclass classification using the pLDA. Improved recognition accuracy of 95\% and 70\% for subject-dependent and subject-independent classification, respectively, is achieved by using the EMDC scheme.
Jonghwa Kim 0001, Elisabeth André
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 I Know What I Did Last Summer: Autobiographic Memory in Synthetic Characters
João Dias 0001, Wan Ching Ho, Thurid Vogt, Nathalie Beeckman, Ana Paiva 0001, Elisabeth André
ACII6
2007 Lexical Affect Sensing: Are Affect Dictionaries Necessary to Analyze Affect?
Alexander Osherenko, Elisabeth André
ACII2
2007 A Systematic Comparison of Different HMM Designs for Emotion Recognition from Acted and Spontaneous Speech
Johannes Wagner 0001, Thurid Vogt, Elisabeth André
ACII3
2007 Attentive Presentation Agents
Tobias Eichner, Helmut Prendinger, Elisabeth André, Mitsuru Ishizuka
IVA3
2007 Integrating a Virtual Agent into the Real World: The Virtual Anatomy Assistant Ritchie
Volker Wiendl, Klaus Dorfmüller-Ulhaas, Nicolas Schulz, Elisabeth André
IVA4
2006 MPML3D: A Reactive Framework for the Multimodal Presentation Markup Language
Michael Nischt, Helmut Prendinger, Elisabeth André, Mitsuru Ishizuka
IVA3
2006 A Plug-and-Play Framework for Theories of Social Group Dynamics
Matthias Rehm, Birgit Lugrin, Elisabeth André
IVA3
2006 Improving Automatic Emotion Recognition from Speech via Gender Differentiaion
Thurid Vogt, Elisabeth André
LREC2
2005 Cross-Cultural Evaluation of Politeness in Tactics for Pedagogical Agents
W. Lewis Johnson, Richard E. Mayer, Elisabeth André, Matthias Rehm
AIED3
2005 Comparing Feature Sets for Acted and Spontaneous Speech in View of Automatic Emotion Recognition
abstract
We present a data-mining experiment on feature selection for automatic emotion recognition. Starting from more than 1000 features derived from pitch, energy and MFCC time series, the most relevant features in respect to the data are selected from this set by removing correlated features. The features selected for acted and realistic emotions are analyzed and show significant differences. All features are computed automatically and we also contrast automatically with manually units of analysis. A higher degree of automation did not prove to be a disadvantage in terms of recognition accuracy
Thurid Vogt, Elisabeth André
ICME2
2005 From Physiological Signals to Emotions: Implementing and Comparing Selected Methods for Feature Extraction and Classification
abstract
Little attention has been paid so far to physiological signals for emotion recognition compared to audio-visual emotion channels, such as facial expressions or speech. In this paper, we discuss the most important stages of a fully implemented emotion recognition system including data analysis and classification. For collecting physiological signals in different affective states, we used a music induction method which elicits natural emotional reactions from the subject. Four-channel biosensors are used to obtain electromyogram, electrocardiogram, skin conductivity and respiration changes. After calculating a sufficient amount of features from the raw signals, several feature selection/reduction methods are tested to extract a new feature set consisting of the most significant features for improving classification performance. Three well-known classifiers, linear discriminant function, k-nearest neighbour and multilayer perceptron, are then used to perform supervised classification
Johannes Wagner 0001, Jonghwa Kim 0001, Elisabeth André
ICME3
2005 Integrating information from speech and physiological signals to achieve emotional sensitivity
abstract
Recently, there has been a significant amount of work on the recognition of emotions from speech and biosignals.Most approaches to emotion recognition so far concentrate on a single modality and do not take advantage of the fact that an integrated multimodal analysis may help to resolve ambiguities and compensate for errors.In this paper, we describe various methods for fusing physiological and voice data at the feature-level and the decision-level as well as a hybrid integration scheme.The results of the integrated recognition approach are then compared with the individual recognition results from each modality.
Jonghwa Kim 0001, Elisabeth André, Matthias Rehm, Thurid Vogt, Johannes Wagner 0001
INTERSPEECH2
2005 The Synthetic Character Ritchie: First Steps Towards a Virtual Companion for Mixed Reality
abstract
Unlike most existing work on traversable interfaces, we focus on the use of synthetic characters to accompany the user in mixed reality (MR) applications. We examine virtual companions as a promising means to design smooth transitions between different worlds and to avoid orientation problems. We propose a taxonomy to describe the spatial relationship between character and user which has an important impact on the style of interaction. To flexibly transfer user and character into different spaces, we have created a platform that supports the design of interfaces derived from the proposed taxonomy as well as transitions between them.
Klaus Dorfmüller-Ulhaas, Elisabeth André
ISMAR2
2005 Where Do They Look? Gaze Behaviors of Multiple Users Interacting with an Embodied Conversational Agent
Matthias Rehm, Elisabeth André
IVA2
2004 Workshop on Social and Emotional Intelligence in Learning Environments
Claude Frasson, Kaska Porayska-Pomsta, Cristina Conati, Guy Gouardères, W. Lewis Johnson, Helen Pain, Elisabeth André, Timothy W. Bickmore, Paul Brna, Isabel Fernández de Castro, Stefano A. Cerri, Cleide Jane Costa, James C. Lester, Christine L. Lisetti, Stacy Marsella, Jack Mostow, Roger Nkambou, Magalie Ochs, Ana Paiva 0001, Fábio Paraguaçu, Natalie K. Person, Rosalind W. Picard, Candace L. Sidner, Angel de Vicente
Intelligent Tutoring Systems7
2004 Exploiting emotions to disambiguate dialogue acts
abstract
This paper describes an attempt to reveal the user's intention from dialogue acts, thereby improving the effectiveness of natural interfaces to pedagogical agents. It focuses on cases where the intention is unclear from the dialogue context or utterance structure, but where the intention may still be identified using the emotional state of the user. The recognition of emotions is based on physiological user input. Our initial user study gave promising results that support our hypothesis that physiological evidence of emotions could be used to disambiguate dialogue acts. This paper presents our approach to the integration of natural language and emotions as well as our first empirical results, which may be used to endow interactive agents with emotional capabilities.
Wauter Bosma, Elisabeth André
IUI2
2004 A generate and sense approach to automated music composition
abstract
Nobody would deny that music may evoke deep and profound emotions. In this paper, we present a perceptual music composition system that aims at the controlled manipulation of a user's emotional state. In contrast to traditional composing techniques, the single components of a composition, such as melody, harmony, rhythm and instrumentation, are selected and combined in a user-specific manner without requiring the user to continuously provide comments on the music employing input devices, such as keyboard or mouse.
SunJung Kim, Elisabeth André
IUI2
2003 A flexible platform for building applications with life-like characters
abstract
During the last years, an increasing number of R&D projects has started to deploy life-like characters for presentation tasks in a diverse range of application areas, including, for example, E-Commerce, E-learning, and help systems. Depending on factors, such as the degree of interactivity and the number of the deployed characters, different architectures have been proposed for system implementation. In this contribution, we first analyse a number of existing user interfaces with presentation characters from an architectural point of view. We then introduce the MIAU platform and illustrate by means of illustrated generation examples how MIAU can be used for the realization of character applications with different conversational settings. Finally, we sketch a number of potential application fields for the MIAU platform
Thomas Rist, Elisabeth André, Stephan Baldes
IUI2
2003 Building applications with life-like characters: the MIAU platform
abstract
No abstract available.
Thomas Rist, Elisabeth André, Stephan Baldes
IUI2
2003 Editorial
Elisabeth André, Ana Paiva 0001
User Model. User Adapt. Interact.1
2001 Presenting through performing: on the use of multiple lifelike characters in knowledge-based presentation systems
Elisabeth André, Thomas Rist
Knowl. Based Syst.1
2000 Getting the Mobile Users in: Three Systems that Support Collaboration in an Environment with Heterogeneous Communication Devices
abstract
In this paper we present MapViews, Magic Lounge, and Call-Kiosk, three different but related systems that address the integration of mobile communication terminals into multi-user applications. MapViews is a test-bed to investigate how a small group of geographically dispersed users can jointly solve localization and route planning tasks while being equipped with different communication terminals. Magic Lounge is a virtual meeting space that provides a number of communication support services and allows its users to connect via heterogeneous devices. Finally, we sketch Call-Kiosk a system that is currently being designed for setting up a commercial information service for mobile clients. All three systems emphasize the high demand for automated design approaches which are able to generate information presentations that are tailored to the available presentation capabilities of particular target devices.
Thomas Rist, Patrick Brandmeier, Gerd Herzog, Elisabeth André
Advanced Visual Interfaces4
2000 Presenting through performing: on the use of multiple lifelike characters in knowledge-based presentation systems
abstract
In this paper, we investigate a new style for presenting information. We introduce the motion of presentation teams which — rather than addressing the user directly — convey information in the style of performances to be observed by him or her. The paper presents an approach to the automated generation of performances which has been tested in two different application scenarios, car sales dialogues and soccer commentary.
Elisabeth André, Thomas Rist
IUI1
1998 Guiding the User Through Dynamically Generated Hypermedia Presentations with a Life-like Character
abstract
Rapid growth of competition on the electronic market place, will generate the demand for new innovative communication styles with web users. In this paper, we develop an operational approach for the automated generation of hypermedia presentations. Unlike conventional hypermedia, we use a life-like presentation agent which presents the generated material, and guides the user through a dynamically expanding navigation space. The approach relies on a model that combines behavior planning for life-like characters with concepts from hypermedia authoring such as timeline structures and navigation graphs.
Elisabeth André, Thomas Rist, Jochen Müller 0001
IUI1
1998 Rocco: A RoboCup Soccer Commentator System
Dirk Voelz, Elisabeth André, Gerd Herzog, Thomas Rist
RoboCup2
1998 WebPersona: a lifelike presentation agent for the World-Wide Web
Elisabeth André, Thomas Rist, Jochen Müller 0001
Knowl. Based Syst.1
1997 Adding Animated Presentation Agents to the Interface
abstract
A growing number of research projects both in academia and industries have started to investigate the use of animated agents in the interface. Such agents, either based on real video, cartoon-style drawings or even model-based 3D graphics, are likely to become integral parts of future user interfaces. To be useful, however, interface agents have to be intelligent in the sense that they exhibit a reasonable behavior. In this paper, we present a system that uses a lifelike character, the so-called PPP Persona, to present multimedia material to the user. This material has been either automatically generated or fetched from the web and modified if necessary. The underlying approach is based on our previous work on multimedia presentation planning. This core approach is complemented by additional concepts, namely the temporal coordination of presentation acts and the consideration of the human-factors dimension of the added visual metaphor.
Thomas Rist, Elisabeth André, Jochen Müller 0001
IUI2
1997 Generating Multimedia Presentations for RoboCup Soccer Games
Elisabeth André, Gerd Herzog, Thomas Rist
RoboCup1
1995 WIP: From Multimedia to Intellimedia
Elisabeth André, Wolfgang Finkler, Winfried Graf, Karin Harbusch, Jochen Heinsohn, Anne Kilger, Bernhard Nebel, Hans-Jürgen Profitlich, Thomas Rist, Wolfgang Wahlster, Andreas Butz, Anthony Jameson
IJCAI1
1994 Referring To World Objects With Text And Pictures
Elisabeth André, Thomas Rist
COLING1
1993 Plan-Based Integration of Natural Language and Graphics Generation
Wolfgang Wahlster, Elisabeth André, Wolfgang Finkler, Hans-Jürgen Profitlich, Thomas Rist
Artif. Intell.2
1992 From Presentation Tasks to Pictures: Towards a Computational Approach to Graphics Design
Thomas Rist, Elisabeth André
ECAI2
1991 Designing Illustrated Texts: How Language Production Is Influenced By Graphics Generation
Wolfgang Wahlster, Elisabeth André, Winfried Graf, Thomas Rist
EACL2
1990 Towards a Plan-Based Synthesis of Illustrated Documents
Elisabeth André, Thomas Rist
ECAI1
1988 On the Simultaneous Interpretation of Real World Image Sequences and their Natural Language Description: The System Soccer
Elisabeth André, Gerd Herzog, Thomas Rist
ECAI1