VLDB 2026 Research / reviewers in the wild / expert
Dominik Schiller
dblp:192/9841
· DBLP profile ↗
14ranked-venue papers
2as first author
9since 2021 · last 2024
0000-0001-7364-5772ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | The AffectToolbox: Affect Analysis for EveryoneabstractIn the field of affective computing, where research continually advances at a rapid pace, the demand for user-friendly tools has become increasingly apparent. In this paper, we present the AffectToolbox, a novel software system that aims to support researchers in developing affect-sensitive studies and prototypes. The proposed system addresses the challenges posed by existing frameworks, which often require profound programming knowledge and cater primarily to power-users or skilled developers. Aiming to facilitate ease of use, the AffectToolbox requires no programming knowledge and offers its functionality to reliably analyze the affective state of users through an accessible graphical user interface. The architecture encompasses a variety of models for emotion recognition on multiple affective channels and modalities, as well as an elaborate fusion system to merge multi-modal assessments into a unified result. The entire system is open-sourced and will be publicly available to ensure easy integration into more complex applications through a well-structured, Python-based code base - therefore marking a substantial contribution toward advancing affective computing research and fostering a more collaborative and inclusive environment within this interdisciplinary field. Silvan Mertes, Dominik Schiller, Michael Dietz, Elisabeth André, Florian Lingenfelser |
ACII | 2 |
| 2024 | Towards Automated Annotation of Infant-Caregiver Engagement Phases with Multimodal Foundation ModelsabstractCaregiver mental health disorders increase the risk of insecure infant attachment and can negatively impact multiple aspects of child development, including cognitive, emotional, and social growth. Infant-caregiver interactions contain subtle psychological and behavioral cues that reveal these adverse effects, underscoring the need for analytical methods to assess them effectively. The Face-to-Face-Still-Face (FFSF) paradigm is a key approach in psychological research for investigating these dynamics, and the Infant and Caregiver Engagement Phases revised German edition (ICEP-R) annotation scheme provides a structured framework for evaluating FFSF interactions. However, manual annotation is labor-intensive and limits scalability, thus hindering a deeper understanding of early developmental impairments. To address this, we developed a computational method that automates the annotation of caregiver-infant interactions using features extracted from audio-visual foundational models. Our approach was tested on 92 FFSF video sessions. Findings demonstrate that models based on bidirectional LSTM and linear classifiers show varying effectiveness depending on the role and feature modality. Specifically, bidirectional LSTM models generally perform better in predicting complex infant engagement phases across multimodal features, while linear models show competitive performance, particularly with unimodal feature encodings like Wav2Vec2-BERT. To support further research, we share our raw feature dataset annotated with ICEP-R labels, enabling broader refinement of computational methods in this area. Daksitha Withanage, Dominik Schiller, Tobias Hallmen, Silvan Mertes, Tobias Baur 0001, Florian Lingenfelser, Mitho Müller, Lea Kaubisch, Corinna Reck, Elisabeth André |
ICMI | 2 |
| 2024 | MultiMediate'24: Multi-Domain Engagement EstimationabstractEstimating the momentary level of participant's engagement is an important prerequisite for assistive systems that support human interactions. Previous work has addressed this task in within-domain evaluation scenarios, i.e. training and testing on the same dataset. This is in contrast to real-life scenarios where domain shifts between training and testing data frequently occur. With MultiMediate'24, we present the first challenge addressing multi-domain engagement estimation. As training data, we utilise the NOXI database of dyadic novice-expert interactions. In addition to within-domain test data, we add two new test domains. First, we introduce recordings following the NOXI protocol but covering languages that are not present in the NOXI training data. Second, we collected novel engagement annotations on the MPIIGroupInteraction dataset which consists of group discussions between three to four people. In this way, MultiMediate'24 evaluates the ability of approaches to generalise across factors such as language and cultural background, group size, task, and screen-mediated vs. face-to-face interaction. This paper describes the MultiMediate'24 challenge and presents baseline results. In addition, we discuss selected challenge solutions. Philipp Müller 0001, Michal Balazia, Tobias Baur 0001, Michael Dietz, Alexander Heimerl, Anna Penzkofer, Dominik Schiller, François Brémond, Jan Alexandersson, Elisabeth André, Andreas Bulling |
ACM Multimedia | 7 |
| 2023 | MultiMediate '23: Engagement Estimation and Bodily Behaviour Recognition in Social InteractionsabstractAutomatic analysis of human behaviour is a fundamental prerequisite for the creation of machines that can effectively interact with- and support humans in social interactions. In MultiMediate'23, we address two key human social behaviour analysis tasks for the first time in a controlled challenge: engagement estimation and bodily behaviour recognition in social interactions. This paper describes the MultiMediate'23 challenge and presents novel sets of annotations for both tasks. For engagement estimation we collected novel annotations on the NOvice eXpert Interaction (NOXI) database. For bodily behaviour recognition, we annotated test recordings of the MPIIGroupInteraction corpus with the BBSI annotation scheme. In addition, we present baseline results for both challenge tasks. Philipp Müller 0001, Michal Balazia, Tobias Baur 0001, Michael Dietz, Alexander Heimerl, Dominik Schiller, Mohammed Guermal, Dominike Thomas, François Brémond, Jan Alexandersson, Elisabeth André, Andreas Bulling |
ACM Multimedia | 6 |
| 2022 | VoiceMe: Personalized voice generation in TTS
Pol van Rijn, Silvan Mertes, Dominik Schiller, Piotr Dura, Hubert Siuzdak, Peter M. C. Harrison, Elisabeth André, Nori Jacoby |
INTERSPEECH | 3 |
| 2022 | MultiMediate'22: Backchannel Detection and Agreement Estimation in Group InteractionsabstractBackchannels, i.e. short interjections of the listener, serve important meta-conversational purposes like signifying attention or indicating agreement. Despite their key role, automatic analysis of backchannels in group interactions has been largely neglected so far. The MultiMediate challenge addresses, for the first time, the tasks of backchannel detection and agreement estimation from backchannels in group conversations. This paper describes the MultiMediate challenge and presents a novel set of annotations consisting of 7234 backchannel instances for the MPIIGroup Interaction dataset. Each backchannel was additionally annotated with the extent by which it expresses agreement towards the current speaker. In addition to a an analysis of the collected annotations, we present baseline results for both challenge tasks. Philipp Müller 0001, Michael Dietz, Dominik Schiller, Dominike Thomas, Hali Lindsay, Patrick Gebhard, Elisabeth André, Andreas Bulling |
ACM Multimedia | 3 |
| 2021 | Exploring Emotional Prototypes in a High Dimensional TTS Latent SpaceabstractRecent TTS systems are able to generate prosodically varied and realistic speech. However, it is unclear how this prosodic variation contributes to the perception of speakers' emotional states. Here we use the recent psychological paradigm 'Gibbs Sampling with People' to search the prosodic latent space in a trained GST Tacotron model to explore prototypes of emotional prosody. Participants are recruited online and collectively manipulate the latent space of the generative speech model in a sequentially adaptive way so that the stimulus presented to one group of participants is determined by the response of the previous groups. We demonstrate that (1) particular regions of the model's latent space are reliably associated with particular emotions, (2) the resulting emotional prototypes are well-recognized by a separate group of human raters, and (3) these emotional prototypes can be effectively transferred to new sentences. Collectively, these experiments demonstrate a novel approach to the understanding of emotional speech by providing a tool to explore the relation between the latent space of generative models and human semantics. Pol van Rijn, Silvan Mertes, Dominik Schiller, Peter M. C. Harrison, Pauline Larrouy-Maestri, Elisabeth André, Nori Jacoby |
Interspeech | 3 |
| 2021 | Analysis by Synthesis: Using an Expressive TTS Model as Feature Extractor for Paralinguistic Speech ClassificationabstractModeling adequate features of speech prosody is one key factor to good performance in affective speech classification.However, the distinction between the prosody that is induced by 'how' something is said (i.e., affective prosody) and the prosody that is induced by 'what' is being said (i.e., linguistic prosody) is neglected in state-of-the-art feature extraction systems.This results in high variability of the calculated feature values for different sentences that are spoken with the same affective intent, which might negatively impact the performance of the classification.While this distinction between different prosody types is mostly neglected in affective speech recognition, it is explicitly modeled in expressive speech synthesis to create controlled prosodic variation.In this work, we use the expressive Text-To-Speech model Global Style Token Tacotron to extract features for a speech analysis task.We show that the learned prosodic representations outperform state-of-the-art feature extraction systems in the exemplary use case of Escalation Level Classification. Dominik Schiller, Silvan Mertes, Pol van Rijn, Elisabeth André |
Interspeech | 1 |
| 2021 | MultiMediate: Multi-modal Group Behaviour Analysis for Artificial MediationabstractArtificial mediators are promising to support human group conversations but at present their abilities are limited by insufficient progress in group behaviour analysis. The MultiMediate challenge addresses, for the first time, two fundamental group behaviour analysis tasks in well-defined conditions: eye contact detection and next speaker prediction. For training and evaluation, MultiMediate makes use of the MPIIGroup Interaction dataset consisting of 22 three- to four-person discussions as well as of an unpublished test set of six additional discussions. This paper describes the MultiMediate challenge and presents the challenge dataset including novel fine-grained speaking annotations that were collected for the purpose of MultiMediate. Furthermore, we present baseline approaches and ablation studies for both challenge tasks Philipp Müller 0001, Michael Dietz, Dominik Schiller, Dominike Thomas, Patrick Gebhard, Elisabeth André, Andreas Bulling |
ACM Multimedia | 3 |
| 2020 | An Evolutionary-based Generative Approach for Audio Data AugmentationabstractIn this paper, we introduce a novel framework to augment raw audio data for machine learning classification tasks. For the first part of our framework, we employ a generative adversarial network (GAN) to create new variants of the audio samples that are already existing in our source dataset for the classification task. In the second step, we then utilize an evolutionary algorithm to search the input domain space of the previously trained GAN, with respect to predefined characteristics of the generated audio. This way we are able to generate audio in a controlled manner that contributes to an improvement in classification performance of the original task. To validate our approach, we chose to test it on the task of soundscape classification. We show that our approach leads to a substantial improvement in classification results when compared to a training routine without data augmentation and training with uncontrolled data augmentation with GANs. Silvan Mertes, Alice Baird, Dominik Schiller, Björn W. Schuller, Elisabeth André |
MMSP | 3 |
| 2019 | Relevance-Based Feature Masking: Improving Neural Network Based Whale Classification Through Explainable Artificial IntelligenceabstractUnderwater sounds provide essential information for marine researchers to study sea mammals.During long-term studies large amounts of sound signals are being recorded using hydrophones.To facilitate the time consuming process of manually evaluating the recorded data, computational systems are often employed.Recent approaches utilize Convolutional Neural Networks (CNNs) to analyze spectrograms extracted from the audio signal.In this paper we explore the potential of relevance analysis to enhance the performance of existing CNN approaches.For this purpose, we present a fusion system that utilizes intermediate outputs of three state of the art CNNs, which are fine tuned to recognize whale sounds in spectrograms.Hereby we use Explainable Artificial Intelligence (XAI) to asses the relevance of each feature within the obtained representations.Based on those relevance values, we create novel masking algorithms to extract significant subsets of respective representations.These subsets are used to train an ensemble of classification systems that are serving as input for the final fusion step.We observe that a classification system can benefit from the inclusion of Relevance-based Feature Masking in terms of improved performance and reduced input dimensionality.The presented work is part of the INTERSPEECH 2019 Computational Paralinguistics Challenge. Dominik Schiller, Tobias Huber, Florian Lingenfelser, Michael Dietz, Andreas Seiderer, Elisabeth André |
INTERSPEECH | 1 |
| 2019 | "Do you trust me?": Increasing User-Trust by Integrating Virtual Agents in Explainable AI Interaction DesignabstractWhile the research area of artificial intelligence benefited from increasingly sophisticated machine learning techniques in recent years, the resulting systems suffer from a loss of transparency and comprehensibility. This development led to an on-going resurgence of the research area of explainable artificial intelligence (XAI) which aims to reduce the opaqueness of those black-box-models. However, much of the current XAI-Research is focused on machine learning practitioners and engineers while omitting the specific needs of end-users. In this paper, we examine the impact of virtual agents within the field of XAI on the perceived trustworthiness of autonomous intelligent systems. To assess the practicality of this concept, we conducted a user study based on a simple speech recognition task. As a result of this experiment, we found significant evidence suggesting that the integration of virtual agents into XAI interaction design leads to an increase of trust in the autonomous intelligent system. Katharina Weitz, Dominik Schiller, Ruben Schlagowski, Tobias Huber, Elisabeth André |
IVA | 2 |
| 2018 | Deep Learning in Paralinguistic Recognition Tasks: Are Hand-crafted Features Still Relevant?abstractIn the past, the performance of machine learning algorithms depended heavily on the representation of the data.Well-designed features therefore played a key role in speech and paralinguistic recognition tasks.Consequently, engineers have put a great deal of work into manually designing large and complex acoustic feature sets.With the emergence of Deep Neural Networks (DNNs), however, it is now possible to automatically infer higher abstractions from simple spectral representations or even learn directly from raw waveforms.This raises the question if (complex) hand-crafted features will still be needed in the future.We take this year's INTERSPEECH Computational Paralinguistic Challenge as an opportunity to approach this issue by means of two corpora -Atypical Affect and Crying.At first, we train a Recurrent Neural Network (RNN) to evaluate the performance of several hand-crafted feature sets of varying complexity.Afterwards, we make the network do the feature engineering all on its own by prefixing a stack of convolutional layers.Our results show that there is no clear winner (yet).This creates room to discuss chances and limits of either approach. Johannes Wagner 0001, Dominik Schiller, Andreas Seiderer, Elisabeth André |
INTERSPEECH | 2 |
| 2017 | Infected Phonemes: How a Cold Impairs Speech on a Phonetic LevelabstractThe realization of language through vocal sounds involves a complex interplay between the lungs, the vocal cords, and a series of resonant chambers (e.g.mouth and nasal cavities).Due to their connection to the outside world, these body parts are popular spots for viruses and bacteria to enter the human organism.Affected people may suffer from an upper respiratory tract infection (URTIC) and consequently their voice often sounds breathy, raspy or sniffly.In this paper, we investigate the audible effects of a cold on a phonetic level.Results on a German corpus show that the articulation of consonants is more impaired than that of vowels.Surprisingly, nasal sounds do not follow this trend in our experiments.We finally try to predict a speaker's health condition by fusing decisions we derive from single phonemes.The presented work is part of the INTER-SPEECH 2017 Computational Paralinguistics Challenge. Johannes Wagner 0001, Thiago Fraga-Silva, Yvan Josse, Dominik Schiller, Andreas Seiderer, Elisabeth André |
INTERSPEECH | 4 |