Florian Lingenfelser

dblp:36/9230 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0003-1582-9850ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author
YearPublicationVenuePosition
2025 VoiceX as a Design Tool for Virtual Agents' Voices
abstract
Modern TTS systems are capable of creating highly realistic and natural-sounding speech, making them an important tool when designing virtual agents.While sounding highly realistic, the process of customizing such TTS voices remains a complex task, mostly requiring the expertise of specialists within the field.One reason for this is the utilization of deep learning models, which are characterized by their expansive, non-interpretable parameter spaces, restricting the feasibility of manual voice customization.In this paper, we present a novel human-in-the-loop paradigm based on an evolutionary algorithm for directly interacting with the parameter space of a neural TTS model.We integrated our approach into a user-friendly graphical user interface that allows users to efficiently create original voices.Those voices can then be used to equip virtual agents with highly customized TTS capabilities by using an open-source programming interface provided by us.Further, in a first pilot study, we show that VoiceX is an appropriate tool for creating individual, custom voices.
Daksitha Withanage, Florian Lingenfelser, Johanna Magdalena Kuch, Otto Grothe, Ruben Schlagowski, Elisabeth André, Silvan Mertes
IVA2
2024 The AffectToolbox: Affect Analysis for Everyone
abstract
In the field of affective computing, where research continually advances at a rapid pace, the demand for user-friendly tools has become increasingly apparent. In this paper, we present the AffectToolbox, a novel software system that aims to support researchers in developing affect-sensitive studies and prototypes. The proposed system addresses the challenges posed by existing frameworks, which often require profound programming knowledge and cater primarily to power-users or skilled developers. Aiming to facilitate ease of use, the AffectToolbox requires no programming knowledge and offers its functionality to reliably analyze the affective state of users through an accessible graphical user interface. The architecture encompasses a variety of models for emotion recognition on multiple affective channels and modalities, as well as an elaborate fusion system to merge multi-modal assessments into a unified result. The entire system is open-sourced and will be publicly available to ensure easy integration into more complex applications through a well-structured, Python-based code base - therefore marking a substantial contribution toward advancing affective computing research and fostering a more collaborative and inclusive environment within this interdisciplinary field.
Silvan Mertes, Dominik Schiller, Michael Dietz, Elisabeth André, Florian Lingenfelser
ACII5
2024 Towards Automated Annotation of Infant-Caregiver Engagement Phases with Multimodal Foundation Models
abstract
Caregiver mental health disorders increase the risk of insecure infant attachment and can negatively impact multiple aspects of child development, including cognitive, emotional, and social growth. Infant-caregiver interactions contain subtle psychological and behavioral cues that reveal these adverse effects, underscoring the need for analytical methods to assess them effectively. The Face-to-Face-Still-Face (FFSF) paradigm is a key approach in psychological research for investigating these dynamics, and the Infant and Caregiver Engagement Phases revised German edition (ICEP-R) annotation scheme provides a structured framework for evaluating FFSF interactions. However, manual annotation is labor-intensive and limits scalability, thus hindering a deeper understanding of early developmental impairments. To address this, we developed a computational method that automates the annotation of caregiver-infant interactions using features extracted from audio-visual foundational models. Our approach was tested on 92 FFSF video sessions. Findings demonstrate that models based on bidirectional LSTM and linear classifiers show varying effectiveness depending on the role and feature modality. Specifically, bidirectional LSTM models generally perform better in predicting complex infant engagement phases across multimodal features, while linear models show competitive performance, particularly with unimodal feature encodings like Wav2Vec2-BERT. To support further research, we share our raw feature dataset annotated with ICEP-R labels, enabling broader refinement of computational methods in this area.
Daksitha Withanage, Dominik Schiller, Tobias Hallmen, Silvan Mertes, Tobias Baur 0001, Florian Lingenfelser, Mitho Müller, Lea Kaubisch, Corinna Reck, Elisabeth André
ICMI6
2023 The Affective Bar Piano
abstract
Music is a great way of supporting a story. It adds a new layer of affective information and as such substantially increases the listening experience in storytelling scenarios. However, in real-time settings, creating emotionally fitting music requires permanent adaptation to the story's mood. While methods to compose and modify music according to emotional states are widely explored, current research rarely uses those techniques in a real-time setting, where such accompanying background music still requires improvisation by human musicians. In this work, we introduce the Affective Bar Piano, a virtual agent that assesses the mood of a story in real time. At the same time, the agent adapts its play to mirror the sensed affect of a human storyteller. In the presented demonstration scenario, the virtual agent is embodied by a 3D piano character playing music in a Wild West saloon setting.
Hannes Ritschel, Silvan Mertes, Florian Lingenfelser, Thomas Kiderle, Elisabeth André
IVA3
2020 NOVA: A Tool for Explanatory Multimodal Behavior Analysis and Its Application to Psychotherapy
Tobias Baur 0001, Sina Clausen, Alexander Heimerl, Florian Lingenfelser, Wolfgang Lutz 0001, Elisabeth André
MMM (2)4
2019 NOVA - A tool for eXplainable Cooperative Machine Learning
abstract
In this paper, we introduce a next-generation annotation tool called NOVA, which implements a workflow that interactively incorporates the `human in the loop'. In particular, NOVA offers a collaborative annotation backend where multiple annotators join their workforce. A main aspect of NOVA is the possibility of applying semi-supervised active learning where Machine Learning techniques are used already during the annotation process by giving the possibility to pre-label data automatically. Furthermore, NOVA implements recent eXplainable AI (XAI) techniques to provide users with both, a confidence value of the automatically predicted annotations, as well as visual explanation. This way, annotators get to understand whether they can trust their ML models, or more annotated data is necessary.
Alexander Heimerl, Tobias Baur 0001, Florian Lingenfelser, Johannes Wagner 0001, Elisabeth André
ACII3
2019 Relevance-Based Feature Masking: Improving Neural Network Based Whale Classification Through Explainable Artificial Intelligence
abstract
Underwater sounds provide essential information for marine researchers to study sea mammals.During long-term studies large amounts of sound signals are being recorded using hydrophones.To facilitate the time consuming process of manually evaluating the recorded data, computational systems are often employed.Recent approaches utilize Convolutional Neural Networks (CNNs) to analyze spectrograms extracted from the audio signal.In this paper we explore the potential of relevance analysis to enhance the performance of existing CNN approaches.For this purpose, we present a fusion system that utilizes intermediate outputs of three state of the art CNNs, which are fine tuned to recognize whale sounds in spectrograms.Hereby we use Explainable Artificial Intelligence (XAI) to asses the relevance of each feature within the obtained representations.Based on those relevance values, we create novel masking algorithms to extract significant subsets of respective representations.These subsets are used to train an ensemble of classification systems that are serving as input for the final fusion step.We observe that a classification system can benefit from the inclusion of Relevance-based Feature Masking in terms of improved performance and reduced input dimensionality.The presented work is part of the INTERSPEECH 2019 Computational Paralinguistics Challenge.
Dominik Schiller, Tobias Huber, Florian Lingenfelser, Michael Dietz, Andreas Seiderer, Elisabeth André
INTERSPEECH3
2018 How to Shape the Humor of a Robot - Social Behavior Adaptation Based on Reinforcement Learning
abstract
A shared sense of humor can result in positive feelings associated with amusement, laughter, and moments of bonding. If robotic companions could acquire their human counterparts' sense of humor in an unobtrusive manner, they could improve their skills of engagement. In order to explore this assumption, we have developed a dynamic user modeling approach based on Reinforcement Learning, which allows a robot to analyze a person's reaction while it tells jokes and continuously adapts its sense of humor. We evaluated our approach in a test scenario with a Reeti robot acting as an entertainer and telling different types of jokes. The exemplary adaptation process is accomplished only by using the audience's vocal laughs and visual smiles, but no other form of explicit feedback. We report on results of a user study with 24 participants, comparing our approach to a baseline condition (with a non-learning version of the robot) and conclude by providing limitations and implications of our approach in detail.
Klaus Weber 0001, Hannes Ritschel, Ilhan Aslan, Florian Lingenfelser, Elisabeth André
ICMI4
2018 Asynchronous and Event-Based Fusion Systems for Affect Recognition on Naturalistic Data in Comparison to Conventional Approaches
abstract
Throughout many present studies dealing with multi-modal fusion, decisions are synchronously forced for fixed time segments across all modalities. Varying success is reported, sometimes performance is worse than unimodal classification. Our goal is the synergistic exploitation of multimodality whilst implementing a real-time system for affect recognition in a naturalistic setting. Therefore we present a categorization of possible fusion strategies for affect recognition on continuous time frames of complete recording sessions and we evaluate multiple implementations from resulting categories. These involve conventional fusion strategies as well as novel approaches that incorporate the asynchronous nature of observed modalities. Some of the latter algorithms consider temporal alignments between modalities and observed frames by applying asynchronous neural networks that use memory blocks to model temporal dependencies. Others use an indirect approach that introduces events as an intermediate layer to accumulate evidence for the target class through all modalities. Recognition results gained on a naturalistic conversational corpus show a drop in recognition accuracy when moving from unimodal classification to synchronous multimodal fusion. However, with our proposed asynchronous and event-based fusion techniques we are able to raise the recognition system's accuracy by 7.83 percent compared to video analysis and 13.71 percent in comparison to common fusion strategies.
Florian Lingenfelser, Johannes Wagner 0001, Raymond Brueckner, Björn W. Schuller, Elisabeth André
IEEE Trans. Affect. Comput.1
2016 MobileSSI: asynchronous fusion for social signal interpretation in the wild
abstract
Over the last years, mobile devices have become an integral part of people's everyday life. At the same time, they provide more and more computational power and memory capacity to perform complex calculations that formerly could only be accomplished with bulky desktop machines. These capabilities combined with the willingness of people to permanently carry them around open up completely new perspectives to the area of Social Signal Processing. To allow for an immediate analysis and interaction, real-time assessment is necessary. To exploit the benefits of multiple sensors, fusion algorithms are required that are able to cope with data loss in asynchronous data streams. In this paper we present MobileSSI, a port of the Social Signal Interpretation (SSI) framework to Android and embedded Linux platforms. We will test to what extent it is possible to run sophisticated synchronization and fusion mechanisms in an everyday mobile setting and compare the results with similar tasks in a laboratory environment.
Simon Flutura, Johannes Wagner 0001, Florian Lingenfelser, Andreas Seiderer, Elisabeth André
ICMI3
2016 Laughter detection in the wild: demonstrating a tool for mobile social signal processing and visualization
abstract
In this demo, we present MobileSSI, a flexible software framework for Android and embedded Linux platforms, that provides developers with tools to record, analyze and recognize human behavior in real-time on mobile devices. To illustrate the benefits of the framework for the analysis of social group dynamics in naturalistic mobile settings, we present a demonstrator for laughter recognition that was implemented with MobileSSI. The demonstrator makes use of smartphones for sensing and analyzing data and employs smartwatches and tablets for visualizing the results and providing user feedback. To enable communication within the resulting ecology of mobile devices, MobileSSI includes a web socket plugin.
Simon Flutura, Johannes Wagner 0001, Florian Lingenfelser, Andreas Seiderer, Elisabeth André
ICMI3
2016 MobileSSI - A Multi-modal Framework for Social Signal Interpretation on Mobile Devices
abstract
Over the last years, new generations of mobile devices have found their way into our pockets. They provide more and more computational power and memory capacity to perform complex calculations that formerly could only be accomplished with bulky desktop machines. Moreover, mobile devices are equipped with a range of sensors to capture people's motion, environmental sound etc. These capabilities combined with the willingness of people to permanently carry them around open up completely new ways of observing human behaviour no longer in laboratories, but "in the wild". However, the detection and analysis of social cues is still a challenging task and requires adequate tools to synchronise, process and analyse relevant signals. This may be the reason why many studies and applications focus on offline analysis and typically collect data over long periods of time and analyse them afterwards. To allow for immediate feedback, real-time assessment is necessary. In this paper, we present MobileSSI, a port of the Social Signal Interpretation (SSI) framework to Android and embedded Linux platforms. The framework supports the joint development of processing pipelines for the analysis of social signals on a desktop computer and mobile devices. Throughout the paper we report on challenges we had to face when porting SSI to a mobile context. Furthermore, we summarise first experiences with a real-life setting in a pub where we focused on the analysis of multimodal social group dynamics investigating laughter as a sign of enjoyment.
Simon Flutura, Johannes Wagner 0001, Florian Lingenfelser, Andreas Seiderer, Elisabeth André
Intelligent Environments3
2015 The Belfast storytelling database: A spontaneous social interaction database with laughter focused annotation
abstract
To support the endeavor of creating intelligent interfaces between computers and humans the use of training materials based on realistic human-human interactions has been recognized as a crucial task. One of the effects of the creation of these databases is an increased realization of the importance of often overlooked social signals and behaviours in organizing and orchestrating our interactions. Laughter is one of these key social signals; its importance in maintaining the smooth flow of human interaction has only recently become apparent in the embodied conversational agent domain. In turn, these realizations require training data that focus on these key social signals. This paper presents a database that is well annotated and theoretically constructed with respect to understanding laughter as it is used within human social interaction. Its construction, motivation, annotation and availability are presented in detail in this paper.
Gary McKeown, William Curran, Johannes Wagner 0001, Florian Lingenfelser, Elisabeth André
ACII4
2015 Combining hierarchical classification with frequency weighting for the recognition of eating conditions
abstract
Though parents regularly remind their children not to do so, talking while eating is a typical everyday situation automatic speech analysis systems should be able to deal with.The Paralinguistic Eating Condition (EC) Challenge at INTERSPEECH 2015 sets the task to classify whether a speaker is eating or not, and if so, which type of food the speaker is currently tasting.The approach we follow in this paper is rather unusual: instead of suppressing the influence of noise to enhance the intelligibility of a spoken message, we try to emphasize the noisy parts of the spectrum to improve the recognition of food classes.To allow for a fine-grained adaption to the characteristic spectrum of single food types we adopt a hierarchical tree structure and decompose the classification task into a sequence of binary decisions.At each node we apply frequency-dependent weighting to tune the spectrum to the involved target classes.With our approach we are able to improve results in a 7-class recognition problem (6 types of food and no food) by more than 7% on the training set (using leave-one-eater-out cross validation) and 4% on the test set, respectively.
Johannes Wagner 0001, Andreas Seiderer, Florian Lingenfelser, Elisabeth André
INTERSPEECH3
2015 Context-Aware Automated Analysis and Annotation of Social Human-Agent Interactions
abstract
The outcome of interpersonal interactions depends not only on the contents that we communicate verbally, but also on nonverbal social signals. Because a lack of social skills is a common problem for a significant number of people, serious games and other training environments have recently become the focus of research. In this work, we present NovA ( No n v erbal behavior A nalyzer), a system that analyzes and facilitates the interpretation of social signals automatically in a bidirectional interaction with a conversational agent. It records data of interactions, detects relevant social cues, and creates descriptive statistics for the recorded data with respect to the agent's behavior and the context of the situation. This enhances the possibilities for researchers to automatically label corpora of human--agent interactions and to give users feedback on strengths and weaknesses of their social behavior.
Tobias Baur 0001, Gregor Mehlmann, Ionut Damian, Florian Lingenfelser, Johannes Wagner 0001, Birgit Lugrin, Elisabeth André, Patrick Gebhard
ACM Trans. Interact. Intell. Syst.4
2014 An Event Driven Fusion Approach for Enjoyment Recognition in Real-time
abstract
Social signals and interpretation of carried information is of high importance in Human Computer Interaction. Often used for affect recognition, the cues within these signals are displayed in various modalities. Fusion of multi-modal signals is a natural and interesting way to improve automatic classification of emotions transported in social signals. Throughout most present studies, uni-modal affect recognition as well as multi-modal fusion, decisions are forced for fixed annotation segments across all modalities. In this paper, we investigate the less prevalent approach of event driven fusion, which indirectly accumulates asynchronous events in all modalities for final predictions. We present a fusion approach, handling short-timed events in a vector space, which is of special interest for real-time applications. We compare results of segmentation based uni-modal classification and fusion schemes to the event driven fusion approach. The evaluation is carried out via detection of enjoyment-episodes within the audiovisual Belfast Story-Telling Corpus.
Florian Lingenfelser, Johannes Wagner 0001, Elisabeth André, Gary McKeown, William Curran
ACM Multimedia1
2013 Using phonetic patterns for detecting social cues in natural conversations
abstract
Laughter and fillers like “uhm” and “ah” are social cues expressed in human speech. Detection and interpretation of such non-linguistic events can reveal important information about the speakers’ intensions and emotional state. The INTERSPEECH 2013 Social Signals Sub-Challenge sets the task to localize and classify laughter and fillers in the “SSPNet Vocalization Corpus” (SVC) based on acoustics. In the paper at hand we investigate phonetic patterns extracted from raw speech transcriptions obtained with the CMU Sphinx toolkit for speech recognition. Even though Sphinx was used out of the box and no dedicated training on the target classes was applied, we were able to successfully predict laughter and filler frames in the development set with ∼ 87% accuracy (unweighted average Area Under the Curve (AUC)). By accumulating our features with a set of standard features provided by the challenge organizers results increased above 92%. When applying the combined set to the test corpus we achieved 87.7% as highest score, which is 4.4% above the challenge baseline.
Johannes Wagner 0001, Florian Lingenfelser, Elisabeth André
INTERSPEECH2
2013 The social signal interpretation (SSI) framework: multimodal signal processing and recognition in real-time
abstract
Automatic detection and interpretation of social signals carried by voice, gestures, mimics, etc. will play a key-role for next-generation interfaces as it paves the way towards a more intuitive and natural human-computer interaction. The paper at hand introduces Social Signal Interpretation (SSI), a framework for real-time recognition of social signals. SSI supports a large range of sensor devices, filter and feature algorithms, as well as, machine learning and pattern recognition tools. It encourages developers to add new components using SSI's C++ API, but also addresses front end users by offering an XML interface to build pipelines with a text editor. SSI is freely available under GPL at http://openssi.net.
Johannes Wagner 0001, Florian Lingenfelser, Tobias Baur 0001, Ionut Damian, Felix Kistler, Elisabeth André
ACM Multimedia2
2012 A Frame Pruning Approach for Paralinguistic Recognition Tasks
abstract
In conventional paralinguistic classification approaches, information gained by low level features is described over broad segments (like whole turns) via statistical functionals. This procedure presumes meaningful information to be embodied within the whole segment. This assumption may be misleading if distinctive cues within a sample are surrounded by non-meaningful information or noise. In this case it would surely be beneficial to keep only parts of the sample that are most relevant for the recognition task. In this paper we propose a novel cluster-based approach, which aims at identifying frames likely to carry distinctive information. Evaluation is done within the INTERSPEECH 2012 Speaker Trait Challenge. Results show that under certain configurations frame pruning in fact leads to an improvement in recognition accuracy. On the observed corpus most stable improvements were achieved at a frame drop of 4-8%. Index Terms: paralinguistic recognition, frame pruning, personality traits
Johannes Wagner 0001, Florian Lingenfelser, Elisabeth André
INTERSPEECH2
2011 A systematic discussion of fusion techniques for multi-modal affect recognition tasks
abstract
Recently, automatic emotion recognition has been established as a major research topic in the area of human computer interaction (HCI). Since humans express emotions through various channels, a user's emotional state can naturally be perceived by combining emotional cues derived from all available modalities. Yet most effort has been put into single-channel emotion recognition, while only a few studies with focus on the fusion of multiple channels have been published. Even though most of these studies apply rather simple fusion strategies -- such as the sum or product rule -- some of the reported results show promising improvements compared to the single channels. Such results encourage investigations if there is further potential for enhancement if more sophisticated methods are incorporated. Therefore we apply a wide variety of possible fusion techniques such as feature fusion, decision level combination rules, meta-classification or hybrid-fusion. We carry out a systematic comparison of a total of 16 fusion methods on different corpora and compare results using a novel visualization technique. We find that multi-modal fusion is in almost any case at least on par with single channel classification, though homogeneous results within corpora point to interchangeability between concrete fusion schemes.
Florian Lingenfelser, Johannes Wagner 0001, Elisabeth André
ICMI1
2011 The Social Signal Interpretation Framework (SSI) for Real Time Signal Processing and Recognition
abstract
The construction of systems for recording, processing and recognising a human's social and affective signals is a challenging effort that includes numerous but necessary sub-tasks to be dealt with.In this article, we introduce our Social Signal Interpretation (SSI) tool, a framework dedicated to support the development of such systems.It provides a flexible architecture to construct pipelines to handle multiple modalities like audio or video and establishing on-and offline recognition tasks.The plug-in system of SSI encourages developers to integrate external code, while a XML interface allows anyone to write own applications with a simple text editor.Furthermore, data recording, annotation and classification can be done using a straightforward graphical user interface, allowing simple access to inexperienced users.
Johannes Wagner 0001, Florian Lingenfelser, Elisabeth André
INTERSPEECH2
2011 Exploring Fusion Methods for Multimodal Emotion Recognition with Missing Data
abstract
The study at hand aims at the development of a multimodal, ensemble-based system for emotion recognition. Special attention is given to a problem often neglected: missing data in one or more modalities. In offline evaluation the issue can be easily solved by excluding those parts of the corpus where one or more channels are corrupted or not suitable for evaluation. In real applications, however, we cannot neglect the challenge of missing data and have to find adequate ways to handle it. To address this, we do not expect examined data to be completely available at all time in our experiments. The presented system solves the problem at the multimodal fusion stage, so various ensemble techniques-covering established ones as well as rather novel emotion specific approaches-will be explained and enriched with strategies on how to compensate for temporarily unavailable modalities. We will compare and discuss advantages and drawbacks of fusion categories and extensive evaluation of mentioned techniques is carried out on the CALLAS Expressivity Corpus, featuring facial, vocal, and gestural modalities.
Johannes Wagner 0001, Elisabeth André, Florian Lingenfelser, Jonghwa Kim 0001
IEEE Trans. Affect. Comput.3
2010 Age and gender classification from speech using decision level fusion and ensemble based techniques
abstract
In this contribution to INTERSPEECH 2010 Paralinguistic Challenge we explore the capabilities of decision level fusion and ensemble based techniques for classification tasks on the provided AGENDER corpus.Ensemble members are generated by providing multiple feature sets generated by feature selection, and novel fusion methods (developed in order to give special support to under-represented classes) are applied for decision making.Results are compared to standard classification approaches and possible benefits are discussed.
Florian Lingenfelser, Johannes Wagner 0001, Thurid Vogt, Jonghwa Kim 0001, Elisabeth André
INTERSPEECH1