EDBT 2026 Demo / reviewers in the wild / expert
Silvan Mertes
dblp:255/0008
· DBLP profile ↗
24ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0001-5230-5218ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 14 · 1 first-author · 13 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From 'Nice Try' to 'Nice Throw': Exploring Counterfactual Explanations as Corrective Feedback for Javelin ThrowingabstractProviding athletes with feedback to refine their technique is key in sports coaching and is critical for improving performance and preventing injuries. However, access to expert coaching is often limited. In this paper, we explore a novel counterfactual-based feedback system as a complementary tool to expert coaching and conduct a small-scale user study to explore its perceived usability. Our approach uses an augmented GANterfactual framework, a modified CycleGAN architecture with a classifier-guided counterfactual loss, to synthesize plausible, actionable feedback. As a test bed for our approach, we use the complex motor task of javelin throwing, a sport that is characterized by high biomechanical demands and injury risk. As we are interested in the perceived usability of our approach, we conduct a user study with 21 sports students. The subjective feedback provided by participants of our user study shows that, while pose-based counterfactual feedback visualizations are appreciated by athletes, for some users they require too much domain-specific knowledge and are not “coach-like” enough. We find that athletes are looking for accompanying textual feedback, supporting recent research in the field of feedback generation for sports and motor learning. Lennart Eing, Annika Stippler, Cristina Conati, Stefan Künzell, Elisabeth André, Silvan Mertes |
AVI | 6 |
| 2025 | Multimodal Generation of Contextualized Jokes for a Real-Time Virtual CharacterabstractHumor often serves as a catalyst for smoother interpersonal communication, enhancing interaction experience between individuals.While virtual characters can also gain from these benefits, implementing humor naturally in human-character interactions remains an open challenge.In this paper, we propose the Joking and Multimodally Amusing Real-Time Character (J-MARC) system, combining a photorealistic character with advanced large language model (LLM) techniques to contextualize jokes within small talk.In the real-time interaction, the character is able to present the jokes multimodally and to apply nonverbal behavior while listening. Thomas Kiderle, Georgiana Cristina Dobre, Jauwairia Nasir, Carlos González Díaz, Hannes Ritschel, Stina Klein, Silvan Mertes, Elisabeth André |
IVA | 7 |
| 2025 | VoiceX as a Design Tool for Virtual Agents' VoicesabstractModern TTS systems are capable of creating highly realistic and natural-sounding speech, making them an important tool when designing virtual agents.While sounding highly realistic, the process of customizing such TTS voices remains a complex task, mostly requiring the expertise of specialists within the field.One reason for this is the utilization of deep learning models, which are characterized by their expansive, non-interpretable parameter spaces, restricting the feasibility of manual voice customization.In this paper, we present a novel human-in-the-loop paradigm based on an evolutionary algorithm for directly interacting with the parameter space of a neural TTS model.We integrated our approach into a user-friendly graphical user interface that allows users to efficiently create original voices.Those voices can then be used to equip virtual agents with highly customized TTS capabilities by using an open-source programming interface provided by us.Further, in a first pilot study, we show that VoiceX is an appropriate tool for creating individual, custom voices. Daksitha Withanage, Florian Lingenfelser, Johanna Magdalena Kuch, Otto Grothe, Ruben Schlagowski, Elisabeth André, Silvan Mertes |
IVA | 7 |
| 2025 | Your Robot, My Voice: Enhancing Android Robot Likability through Personalization by Cloning the User's VoiceabstractThis study investigates whether personalized voice cloning can improve a robot’s likability compared to a design-congruent voice and a distinctly dissimilar voice. Participants interacted with a gender-ambiguous android robot in three different voice conditions. We compared: (1) a personalized voice clone based on the participant’s voice, (2) a design-congruent voice matching the robot’s appearance, and (3) a dissimilar voice, which differs from both the participant’s and the robot’s features.The cloned and design-congruent voices significantly increased likability compared to the dissimilar voice, while anthropomorphism and familiarity showed no significant differences across conditions. Most participants did not immediately recognize their cloned voice until informed that one of the voices was a clone. However, most of the participants were successful when asked to pick out their cloned voice from those used. We assume that voice personalization through similarity to the user improves likability even before the user is aware of this similarity.Our results show that personalized voice cloning is a simple alternative to other methods for the design of robotic voices. It significantly increases robot likability while requiring minimal user effort. Johanna Magdalena Kuch, Marcel Heisler, Stina Klein, Silvan Mertes, Lennart Eing, Elisabeth André, Christian Becker-Asano |
RO-MAN | 4 |
| 2025 | The ForDigitStress Dataset: A Multi-Modal Dataset for Automatic Stress RecognitionabstractWe present a multi-modal stress dataset that uses digital job interviews to induce stress. The dataset provides multi-modal data of 40 participants including audio, video (motion capturing, facial landmarks, eye tracking), as well as physiological information (photoplethysmography, electrodermal activity). In addition to that, the dataset contains time-continuous annotations for stress and occurred emotions (e.g., shame, anger, anxiety, and surprise). In order to establish a baseline, five different machine learning classifiers (Support Vector Machine, K-Nearest Neighbors, Random Forest, Feed-forward Neural Network, and Long-Short-Term Memory Network) have been trained and evaluated on the presented dataset for a binary stress classification task. The best-performing classifier has been a Long-Short-Term Memory Network, which achieved an accuracy of 91.7% and an F1-score of 90.2%. The ForDigitStress dataset is freely available to other researchers. Alexander Heimerl, Pooja Prajod, Silvan Mertes, Tobias Baur 0001, Matthias Kraus 0001, Ailin Liu, Helen Risack, Nicolas Rohleder, Elisabeth André, Linda Becker |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | GANonymization: A GAN-Based Face Anonymization Framework for Preserving Emotional ExpressionsabstractIn recent years, the increasing availability of personal data has raised concerns regarding privacy and security. One of the critical processes to address these concerns is data anonymization, which aims to protect individual privacy and prevent the release of sensitive information. This research focuses on the importance of face anonymization. Therefore, we introduce GANonymization, a novel face anonymization framework with facial expression-preserving abilities. Our approach is based on a high-level representation of a face, which is synthesized into an anonymized version based on a generative adversarial network (GAN). The effectiveness of the approach was assessed by evaluating its performance in removing identifiable facial attributes to increase the anonymity of the given individual face. Additionally, the performance of preserving facial expressions was evaluated on several affect recognition datasets and outperformed the state-of-the-art methods in most categories. Finally, our approach was analyzed for its ability to remove various facial traits, such as jewelry, hair color, and multiple others. Here, it demonstrated reliable performance in removing these attributes. Our results suggest that GANonymization is a promising approach for anonymizing faces while preserving facial expressions. Fabio Hellmann, Silvan Mertes, Mohamed Benouis, Alexander Hustinx, Tzung-Chien Hsieh, Cristina Conati, Peter M. Krawitz, Elisabeth André |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | The AffectToolbox: Affect Analysis for EveryoneabstractIn the field of affective computing, where research continually advances at a rapid pace, the demand for user-friendly tools has become increasingly apparent. In this paper, we present the AffectToolbox, a novel software system that aims to support researchers in developing affect-sensitive studies and prototypes. The proposed system addresses the challenges posed by existing frameworks, which often require profound programming knowledge and cater primarily to power-users or skilled developers. Aiming to facilitate ease of use, the AffectToolbox requires no programming knowledge and offers its functionality to reliably analyze the affective state of users through an accessible graphical user interface. The architecture encompasses a variety of models for emotion recognition on multiple affective channels and modalities, as well as an elaborate fusion system to merge multi-modal assessments into a unified result. The entire system is open-sourced and will be publicly available to ensure easy integration into more complex applications through a well-structured, Python-based code base - therefore marking a substantial contribution toward advancing affective computing research and fostering a more collaborative and inclusive environment within this interdisciplinary field. Silvan Mertes, Dominik Schiller, Michael Dietz, Elisabeth André, Florian Lingenfelser |
ACII | 1 |
| 2024 | Giving Robots a Voice: Human-in-the-Loop Voice Creation and open-ended LabelingabstractSpeech is a natural interface for humans to interact with robots. Yet, aligning a robot’s voice to its appearance is challenging due to the rich vocabulary of both modalities. Previous research has explored a few labels to describe robots and tested them on a limited number of robots and existing voices. Here, we develop a robot-voice creation tool followed by large-scale behavioral human experiments (N=2,505). First, participants collectively tune robotic voices to match 175 robot images using an adaptive human-in-the-loop pipeline. Then, participants describe their impression of the robot or their matched voice using another human-in-the-loop paradigm for open-ended labeling. The elicited taxonomy is then used to rate robot attributes and to predict the best voice for an unseen robot. We offer a web interface to aid engineers in customizing robot voices, demonstrating the synergy between cognitive science and machine learning for engineering tools. Pol van Rijn, Silvan Mertes, Kathrin Janowski, Katharina Weitz, Nori Jacoby, Elisabeth André |
CHI | 2 |
| 2024 | Towards Automated Annotation of Infant-Caregiver Engagement Phases with Multimodal Foundation ModelsabstractCaregiver mental health disorders increase the risk of insecure infant attachment and can negatively impact multiple aspects of child development, including cognitive, emotional, and social growth. Infant-caregiver interactions contain subtle psychological and behavioral cues that reveal these adverse effects, underscoring the need for analytical methods to assess them effectively. The Face-to-Face-Still-Face (FFSF) paradigm is a key approach in psychological research for investigating these dynamics, and the Infant and Caregiver Engagement Phases revised German edition (ICEP-R) annotation scheme provides a structured framework for evaluating FFSF interactions. However, manual annotation is labor-intensive and limits scalability, thus hindering a deeper understanding of early developmental impairments. To address this, we developed a computational method that automates the annotation of caregiver-infant interactions using features extracted from audio-visual foundational models. Our approach was tested on 92 FFSF video sessions. Findings demonstrate that models based on bidirectional LSTM and linear classifiers show varying effectiveness depending on the role and feature modality. Specifically, bidirectional LSTM models generally perform better in predicting complex infant engagement phases across multimodal features, while linear models show competitive performance, particularly with unimodal feature encodings like Wav2Vec2-BERT. To support further research, we share our raw feature dataset annotated with ICEP-R labels, enabling broader refinement of computational methods in this area. Daksitha Withanage, Dominik Schiller, Tobias Hallmen, Silvan Mertes, Tobias Baur 0001, Florian Lingenfelser, Mitho Müller, Lea Kaubisch, Corinna Reck, Elisabeth André |
ICMI | 4 |
| 2024 | Relevant Irrelevance: Generating Alterfactual Explanations for Image Classifiers
Silvan Mertes, Tobias Huber, Christina Karle, Katharina Weitz, Ruben Schlagowski, Cristina Conati, Elisabeth André |
IJCAI | 1 |
| 2024 | Evaluating Gender Ambiguity, Novelty and Anthropomorphism in Humming and Talking Voices for RobotsabstractThis paper investigates the effects of gender neutralization on the perception of anthropomorphism, gender specificity, and novelty for human voices, comparing spoken and hummed voice modalities. We evaluated gender-neutralized and original voice samples in both spoken and hummed formats using an online survey. Our results confirm that gender-neutralizing filters effectively reduce perceived gender specificity in both modalities, supporting their use in creating gender-neutral voices for humanoid robots. Hummed voices were perceived as more anthropomorphic and less novel than spoken voices, suggesting that non-verbal sound modalities can enhance the human likeness of gender-neutral androids while maintaining gender ambiguity. The study contributes to HRI by highlighting the potential of humming to fulfill users’ expectations of interaction with android robots. Johanna Magdalena Kuch, Jauwairia Nasir, Silvan Mertes, Ruben Schlagowski, Christian Becker-Asano, Elisabeth André |
RO-MAN | 3 |
| 2024 | The STOIC2021 COVID-19 AI challenge: Applying reusable training methodologies to private dataabstractChallenges drive the state-of-the-art of automated medical image analysis. The quantity of public training data that they provide can limit the performance of their solutions. Public access to the training methodology for these solutions remains absent. This study implements the Type Three (T3) challenge format, which allows for training solutions on private data and guarantees reusable training methodologies. With T3, challenge organizers train a codebase provided by the participants on sequestered training data. T3 was implemented in the STOIC2021 challenge, with the goal of predicting from a computed tomography (CT) scan whether subjects had a severe COVID-19 infection, defined as intubation or death within one month. STOIC2021 consisted of a Qualification phase, where participants developed challenge solutions using 2000 publicly available CT scans, and a Final phase, where participants submitted their training methodologies with which solutions were trained on CT scans of 9724 subjects. The organizers successfully trained six of the eight Final phase submissions. The submitted codebases for training and running inference were released publicly. The winning solution obtained an area under the receiver operating characteristic curve for discerning between severe and non-severe COVID-19 of 0.815. The Final phase solutions of all finalists improved upon their Qualification phase solutions. Luuk H. Boulogne, Julian Lorenz, Daniel Kienzle, Robin Schön, Katja Ludwig, Rainer Lienhart, Simon Jégou, Derik Shi, Mayug Maniparambil, Dominik Müller, Silvan Mertes, Niklas Schröter, Fabio Hellmann, Miriam Elia, Ine Dirks, Matías N. Bossa, Abel Díaz Berenguer, Tanmoy Mukherjee, Jef Vandemeulebroucke, Hichem Sahli, Nikos Deligiannis, Panagiotis Gonidakis, Ngoc Dung Huynh, Muhammad Imran Razzak, Mohamed Reda Bouadjenek, Mario Verdicchio, Pasquale Borrelli, Marco Aiello 0003, James A. Meakin, Alexander Lemm, Christoph Russ, Razvan Ionasec, Nikos Paragios, Bram van Ginneken, Marie-Pierre Revel |
Medical Image Anal. | 14 |
| 2023 | Wish You Were Here: Mental and Physiological Effects of Remote Music Collaboration in Mixed RealityabstractWith face-to-face music collaboration being severely limited during the recent pandemic, mixed reality technologies and their potential to provide musicians a feeling of "being there" with their musical partner can offer tremendous opportunities. In order to assess this potential, we conducted a laboratory study in which musicians made music together in real-time while simultaneously seeing their jamming partner’s mixed reality point cloud via a head-mounted display and compared mental effects such as flow, affect, and co-presence to an audio-only baseline. In addition, we tracked the musicians’ physiological signals and evaluated their features during times of self-reported flow. For users jamming in mixed reality, we observed a significant increase in co-presence. Regardless of the condition (mixed reality or audio-only), we observed an increase in positive affect after jamming remotely. Furthermore, we identified heart rate and HF/LF as promising features for classifying the flow state musicians experienced while making music together. Ruben Schlagowski, Dariia Nazarenko, Yekta Said Can, Kunal Gupta, Silvan Mertes, Mark Billinghurst, Elisabeth André |
CHI | 5 |
| 2023 | Multimodal Irony for Virtual CharactersabstractHumor is an important communicative skill in human interactions. Intelligent virtual agents can leverage it to increase their believability and overall interaction experience. In this paper, we focus on transferring and implementing existing multimodal irony markers from the literature to a photorealistic virtual character. The verbal content is generated dynamically by an irony generator. We demonstrate how the ironic turn can be augmented with prosodic and facial markers. An expressivity parameter allows us to manipulate the encoding of the irony style. Thomas Kiderle, Hannes Ritschel, Silvan Mertes, Elisabeth André |
IVA | 3 |
| 2023 | The Affective Bar PianoabstractMusic is a great way of supporting a story. It adds a new layer of affective information and as such substantially increases the listening experience in storytelling scenarios. However, in real-time settings, creating emotionally fitting music requires permanent adaptation to the story's mood. While methods to compose and modify music according to emotional states are widely explored, current research rarely uses those techniques in a real-time setting, where such accompanying background music still requires improvisation by human musicians. In this work, we introduce the Affective Bar Piano, a virtual agent that assesses the mood of a story in real time. At the same time, the agent adapts its play to mirror the sensed affect of a human storyteller. In the presented demonstration scenario, the virtual agent is embodied by a 3D piano character playing music in a Wild West saloon setting. Hannes Ritschel, Silvan Mertes, Florian Lingenfelser, Thomas Kiderle, Elisabeth André |
IVA | 2 |
| 2023 | An Overview of Affective Speech Synthesis and Conversion in the Deep Learning EraabstractSpeech is the fundamental mode of human communication, and its synthesis has long been a core priority in human–computer interaction research. In recent years, machines have managed to master the art of generating speech that is understandable by humans. However, the linguistic content of an utterance encompasses only a part of its meaning. Affect, or expressivity, has the capacity to turn speech into a medium capable of conveying intimate thoughts, feelings, and emotions—aspects that are essential for engaging and naturalistic interpersonal communication. While the goal of imparting expressivity to synthesized utterances has so far remained elusive, following recent advances in text-to-speech synthesis, a paradigm shift is well under way in the fields of affective speech synthesis and conversion as well. Deep learning, as the technology that underlies most of the recent advances in artificial intelligence, is spearheading these efforts. In this overview, we outline ongoing trends and summarize state-of-the-art approaches in an attempt to provide a broad overview of this exciting field. Andreas Triantafyllopoulos, Björn W. Schuller, Gökçe Iymen, Tevfik Metin Sezgin, Xiangheng He, Zijiang Yang 0007, Panagiotis Tzirakis, Shuo Liu 0012, Silvan Mertes, Elisabeth André, Ruibo Fu, Jianhua Tao 0001 |
Proc. IEEE | 9 |
| 2022 | Generating Personalized Behavioral Feedback for a Virtual Job Interview Training System Through Adversarial Learning
Alexander Heimerl, Silvan Mertes, Tanja Schneeberger, Tobias Baur 0001, Ailin Liu, Linda Becker, Nicolas Rohleder, Patrick Gebhard, Elisabeth André |
AIED (1) | 2 |
| 2022 | Flow with the Beat! Human-Centered Design of Virtual Environments for Musical Creativity Support in VRabstractAs previous studies have shown, the environment of creative people can have a significant impact on their creative process and thus on their creations. However, with the advent of digital tools such as virtual instruments and digital audio workstations, more and more creative work is digital and decoupled from the creator’s environment. Virtual Reality technologies open up new possibilities here, as creative tools can seamlessly merge with any virtual environment the user finds himself in. This paper reports on the human-centered design process of a VR application that aims at supporting the user’s individual needs to support their creativity while composing percussive beats in virtual environments. For this purpose, we derived factors that influence creativity from literature and conducted focus group interviews in order to learn how virtual environments and 3DUI can be designed for creativity support. In a subsequent laboratory study, we let users interact with a virtual step sequencer UI in virtual environments that were either customizable or fixed/unchangeable. By analyzing post-test ratings from music experts, self-report questionnaires, and user behavior data, we examined the effects of such customizable virtual environments on user creativity, user experience, flow, and subjective creativity support scales. While we did not observe a significant impact of this independent variable on user creativity, user experience or flow, we found that users had specific individual needs regarding their virtual surroundings and strongly preferred customizable virtual environments, even though the fixed virtual environment was designed to be creatively stimulating. We also observed consistently high flow and user experience ratings, which promote human-centered design of VR-based creativity support tools in a musical context. Ruben Schlagowski, Fabian Wildgrube, Silvan Mertes, Ceenu George, Elisabeth André |
Creativity & Cognition | 3 |
| 2022 | VoiceMe: Personalized voice generation in TTS
Pol van Rijn, Silvan Mertes, Dominik Schiller, Piotr Dura, Hubert Siuzdak, Peter M. C. Harrison, Elisabeth André, Nori Jacoby |
INTERSPEECH | 2 |
| 2021 | A Prototypical Network Approach for Evaluating Generated Emotional SpeechabstractThe collection of emotional speech data is a time-consuming and costly endeavour.Generative networks can be applied to augment the limited audio data artificially.However, it is challenging to evaluate generated audio for its similarity to source data, as current quantitative metrics are not necessarily suited to the audio domain.We explore the use of a prototypical network to evaluate four classes of generated emotional audio with this in mind.We first extract spectrogram images from WAVEGAN generated audio and other audio augmentation approaches, comparing similarity to the class prototype and diversity within the embedding space.Furthermore, we augment the source training set with each augmentation type and perform a classification to explore the generated audio plausibility.Results suggest that quality and diversity can be quantitatively observed with this approach.In the chosen context, we see that WAVEGAN generated data is recognisable as a source data class (F1-score 43.6 %), and the samples add similar diversity as unseen source data.This result leads to more plausible data for augmentation of the source training set -achieving up to 63.9 % F1 which is a 3.5 % improvement over the source data baseline. Alice Baird, Silvan Mertes, Manuel Milling, Lukas Stappen, Thomas Wiest, Elisabeth André, Björn W. Schuller |
Interspeech | 2 |
| 2021 | Exploring Emotional Prototypes in a High Dimensional TTS Latent SpaceabstractRecent TTS systems are able to generate prosodically varied and realistic speech. However, it is unclear how this prosodic variation contributes to the perception of speakers' emotional states. Here we use the recent psychological paradigm 'Gibbs Sampling with People' to search the prosodic latent space in a trained GST Tacotron model to explore prototypes of emotional prosody. Participants are recruited online and collectively manipulate the latent space of the generative speech model in a sequentially adaptive way so that the stimulus presented to one group of participants is determined by the response of the previous groups. We demonstrate that (1) particular regions of the model's latent space are reliably associated with particular emotions, (2) the resulting emotional prototypes are well-recognized by a separate group of human raters, and (3) these emotional prototypes can be effectively transferred to new sentences. Collectively, these experiments demonstrate a novel approach to the understanding of emotional speech by providing a tool to explore the relation between the latent space of generative models and human semantics. Pol van Rijn, Silvan Mertes, Dominik Schiller, Peter M. C. Harrison, Pauline Larrouy-Maestri, Elisabeth André, Nori Jacoby |
Interspeech | 2 |
| 2021 | Analysis by Synthesis: Using an Expressive TTS Model as Feature Extractor for Paralinguistic Speech ClassificationabstractModeling adequate features of speech prosody is one key factor to good performance in affective speech classification.However, the distinction between the prosody that is induced by 'how' something is said (i.e., affective prosody) and the prosody that is induced by 'what' is being said (i.e., linguistic prosody) is neglected in state-of-the-art feature extraction systems.This results in high variability of the calculated feature values for different sentences that are spoken with the same affective intent, which might negatively impact the performance of the classification.While this distinction between different prosody types is mostly neglected in affective speech recognition, it is explicitly modeled in expressive speech synthesis to create controlled prosodic variation.In this work, we use the expressive Text-To-Speech model Global Style Token Tacotron to extract features for a speech analysis task.We show that the learned prosodic representations outperform state-of-the-art feature extraction systems in the exemplary use case of Escalation Level Classification. Dominik Schiller, Silvan Mertes, Pol van Rijn, Elisabeth André |
Interspeech | 2 |
| 2020 | An Evolutionary-based Generative Approach for Audio Data AugmentationabstractIn this paper, we introduce a novel framework to augment raw audio data for machine learning classification tasks. For the first part of our framework, we employ a generative adversarial network (GAN) to create new variants of the audio samples that are already existing in our source dataset for the classification task. In the second step, we then utilize an evolutionary algorithm to search the input domain space of the previously trained GAN, with respect to predefined characteristics of the generated audio. This way we are able to generate audio in a controlled manner that contributes to an improvement in classification performance of the original task. To validate our approach, we chose to test it on the task of soundscape classification. We show that our approach leads to a substantial improvement in classification results when compared to a training routine without data augmentation and training with uncontrolled data augmentation with GANs. Silvan Mertes, Alice Baird, Dominik Schiller, Björn W. Schuller, Elisabeth André |
MMSP | 1 |
| 2019 | Personalized Synthesis of Intentional and Emotional Non-Verbal Sounds for Social RobotsabstractNon-verbal sounds are an essential communication channel for social robots. However, it requires expert knowledge to create and compose synthesizers, develop melodic structures or record samples which express a robot's internal intentions and emotions. This paper presents an approach for adapting a robot's timbre based on non-expert human comparative feedback in order to personalize the sonic interaction design to an individual user's preferences. An evolution strategy learns parameters of real-time sound synthesis for different intentions and emotions. Ultimately, the strategy aims to improve the perceived goodness of how well a specific melody's sound maps to a specific emotion or intention. In order to demonstrate the feasibility of the approach, we report on a user study with a robot, 6 exemplary melodies and 27 participants. Our study results show that the strategy indeed results in improved and preferred sound designs and that many participants are willing to apply such a process to improve their robots' expressivity. Hannes Ritschel, Ilhan Aslan, Silvan Mertes, Andreas Seiderer, Elisabeth André |
ACII | 3 |