Ingo Siegert

dblp:67/8154 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0001-7447-7141ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorSecurity and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Avatar Motion Signatures: Evaluating Linkability of Expressive De-Identification
Fenja Schulz, Jan Marquenie, Carlos Franzreb, Tim Polzehl, Ingo Siegert, Sebastian Möller 0001
ICISSP (2)5
2026 Voices across Decades: A Multimodal Diachronic Corpus of German Bundestag Debates (GerParlDia-MM)
Ingo Siegert
LREC1
2025 Queer Waves: A German Speech Dataset Capturing Gender and Sexual Diversity from Podcasts and YouTube
Ingo Siegert, Jan Marquenie, Sven Grawunder
INTERSPEECH1
2024 AnonEmoFace: Emotion Preserving Facial Anonymization
Jan Hintz, Jacob Rühe, Ingo Siegert
ICISSP3
2024 Anonymising Elderly and Pathological Speech: Voice Conversion Using DDSP and Query-by-Example
abstract
Speech anonymisation aims to protect speaker identity by changing personal identifiers in speech while retaining linguistic content. Current methods fail to retain prosody and unique speech patterns found in elderly and pathological speech domains, which is essential for remote health monitoring. To address this gap, we propose a voice conversion-based method (DDSP-QbE) using differentiable digital signal processing and query-by-example. The proposed method, trained with novel losses, aids in disentangling linguistic, prosodic, and domain representations, enabling the model to adapt to uncommon speech patterns. Objective and subjective evaluations show that DDSP-QbE significantly outperforms the voice conversion state-of-the-art concerning intelligibility, prosody, and domain preservation across diverse datasets, pathologies, and speakers while maintaining quality and speaker anonymity. Experts validate domain preservation by analysing twelve clinically pertinent domain attributes.
Suhita Ghosh, Mélanie Jouaiti, Yamini Sinha, Tim Polzehl, Ingo Siegert, Sebastian Stober
INTERSPEECH6
2024 Challenges of German Speech Recognition: A Study on Multi-ethnolectal Speech Among Adolescents
Martha Schubert, Daniel Duran 0001, Ingo Siegert
INTERSPEECH3
2023 Emo-StarGAN: A Semi-Supervised Any-to-Many Non-Parallel Emotion-Preserving Voice Conversion
abstract
Speech anonymisation prevents misuse of spoken data by removing any personal identifier while preserving at least linguistic content. However, emotion preservation is crucial for natural human-computer interaction. The well-known voice conversion technique StarGANv2-VC achieves anonymisation but fails to preserve emotion. This work presents an any-to-many semi-supervised StarGANv2-VC variant trained on partially emotion-labelled non-parallel data. We propose emotion-aware losses computed on the emotion embeddings and acoustic features correlated to emotion. Additionally, we use an emotion classifier to provide direct emotion supervision. Objective and subjective evaluations show that the proposed approach significantly improves emotion preservation over the vanilla StarGANv2-VC. This considerable improvement is seen over diverse datasets, emotions, target speakers, and inter-group conversions without compromising intelligibility and anonymisation.
Suhita Ghosh, Yamini Sinha, Ingo Siegert, Tim Polzehl, Sebastian Stober
INTERSPEECH4
2021 Effects of Prosodic Variations on Accidental Triggers of a Commercial Voice Assistant
Ingo Siegert
Interspeech1
2020 An Analysis of the Applicability of VoiceXML as Basis for a Dialog Control Flow in Industrial Interaction Management
abstract
Industry 4.0 (I4.0) looks to enable intelligent production by connecting and evaluating data. The asset administration shell, the Industry 4.0 specification of a digital twin describes various concepts to realize this data exchange. One part of the asset administration shell is the I4.0-language, which intends to standardize complex interactions between machines by using interaction protocols. VoiceXML is a W3C-standard from the field of interactive voice response, which has been established for several years. In this paper we analyze to which extend a future I4.0 interaction manager could derive VoiceXML-concepts to the asset administration shell. For this purpose, parallels between VoiceXML and the I4.0-language are shown and compared by implementing a selected interaction protocol from the VDI 2193 using VoiceXML.
Felix Böhm, Ingo Siegert, Alexander Belyaev, Christian Diedrich
ETFA2
2020 "Alexa in the wild" - Collecting Unconstrained Conversations with a Modern Voice Assistant in a Public Environment
abstract
Datasets featuring modern voice assistants such as Alexa, Siri, Cortana and others allow an easy study of human-machine interactions. But data collections offering an unconstrained, unscripted public interaction are quite rare. Many studies so far have focused on private usage, short pre-defined task or specific domains. This contribution presents a dataset providing a large amount of unconstrained public interactions with a voice assistant. Up to now around 40 hours of device directed utterances were collected during a science exhibition touring through Germany. The data recording was part of an exhibit that engages visitors to interact with a commercial voice assistant system (Amazon’s ALEXA), but did not restrict them to a specific topic. A specifically developed quiz was starting point of the conversation, as the voice assistant was presented to the visitors as a possible joker for the quiz. But the visitors were not forced to solve the quiz with the help of the voice assistant and thus many visitors had an open conversation. The provided dataset – Voice Assistant Conversations in the wild (VACW) – includes the transcripts of both visitors requests and Alexa answers, identified topics and sessions as well as acoustic characteristics automatically extractable from the visitors’ audio files.
Ingo Siegert
LREC1
2020 Improving Automatic Speech Recognition Utilizing Audio-codecs for Data Augmentation
abstract
To train end-to-end automatic speech recognition models, it requires a large amount of labeled speech data. This goal is challenging for languages with fewer resources. In contrast to the commonly used feature level data augmentation, we propose to expand the training set by using different audio codecs at the data level. The augmentation method consists of using different audio codecs with changed bit rate, sampling rate, and bit depth. The change reassures variation in the input data without drastically affecting the audio quality. Besides, we can ensure that humans still perceive the audio, and any feature extraction is possible later. To demonstrate the general applicability of the proposed augmentation technique, we evaluated it in an end-to-end automatic speech recognition architecture in four languages. After applying the method, on the Amharic, Dutch, Slovenian, and Turkish datasets, we achieved a 1.57 average improvement in the character error rates (CER) without integrating language models. The result is comparable to the baseline result, showing CER improvement of 2.78, 1.25, 1.21, and 1.05 for each language. On the Amharic dataset, we reached a syllable error rate reduction of 6.12 compared to the baseline result.
Nirayo Hailu Gebreegziabher, Ingo Siegert, Andreas Nürnberger
MMSP2
2019 Three's a Crowd? Effects of a Second Human on Vocal Accommodation with a Voice Assistant
Eran Raveh, Ingo Siegert, Ingmar Steiner, Iona Gessinger, Bernd Möbius
INTERSPEECH2
2019 Cross-Corpus Data Augmentation for Acoustic Addressee Detection
abstract
Acoustic addressee detection (AD) is a modern paralinguistic and dialogue challenge that especially arises in voice assistants.In the present study, we distinguish addressees in two settings (a conversation between several people and a spoken dialogue system, and a conversation between several adults and a child) and introduce the first competitive baseline (unweighted average recall equals 0.891) for the Voice Assistant Conversation Corpus that models the first setting.We jointly solve both classification problems, using three models: a linear support vector machine dealing with acoustic functionals and two neural networks utilising raw waveforms alongside with acoustic low-level descriptors.We investigate how different corpora influence each other, applying the mixup approach to data augmentation.We also study the influence of various acoustic context lengths on AD.Two-second speech fragments turn out to be sufficient for reliable AD.Mixup is shown to be beneficial for merging acoustic data (extracted features but not raw waveforms) from different domains that allows us to reach a higher classification performance on human-machine AD and also for training a multipurpose neural network that is capable of solving both human-machine and adult-child AD problems.
Oleg Akhtiamov, Ingo Siegert, Alexey Karpov 0001, Wolfgang Minker
SIGdial2
2018 Using a PCA-based dataset similarity measure to improve cross-corpus emotion recognition
Ingo Siegert, Ronald Böck, Andreas Wendemuth
Comput. Speech Lang.1
2016 ERM4CT 2016: 2nd international workshop on emotion representations and modelling for companion systems (workshop summary)
abstract
In this paper the organisers present a brief overview of the 2nd International Workshop on Emotion Representations and Modelling for Companion Systems (ERM4CT). The ERM4CT 2016 Workshop is held in conjunction with the 18th ACM International Conference on Multimodal Interaction (ICMI 2016) taking place Tokyo, Japan. The ERM4CT is the follow-up of three previous workshops on emotion modelling for affective human-computer interaction and companion systems. Apart from its usual focus on emotion representations and models, this year's ERM4CT puts special emphasis on how to model adequate affective system behaviour. For the first time, this year's ERM4CT gave out a dataset, which all attendees could investigate to jointly discuss their findings.
Kim Hartmann, Ingo Siegert, Albert Ali Salah, Khiet P. Truong
ICMI2
2015 Exploring dataset similarities using PCA-based feature selection
abstract
In emotion recognition from speech, several well-established corpora are used to date for the development of classification engines. The data is annotated differently, and the community in the field uses a variety of feature extraction schemes. The aim of this paper is to investigate promising features for individual corpora and then compare the results for proposing optimal features across data sets, introducing a new ranking method. Further, this enables us to present a method for automatic identification of groups of corpora with similar characteristics. This answers an urgent question in classifier development, namely whether data from different corpora is similar enough to jointly be used as training material, overcoming shortage of material in matching domains. We compare the results of this method with manual groupings of corpora. We consider the established emotional speech corpora AVIC, ABC, DES, EMO-DB, ENTERFACE, SAL, SMARTKOM, SUSAS and VAM, however our approach is general.
Ingo Siegert, Ronald Böck, Andreas Wendemuth, Bogdan Vlasenko
ACII1
2014 Application of image processing methods to filled pauses detection from spontaneous speech
Dmytro Prylipko, Olga Egorow, Ingo Siegert, Andreas Wendemuth
INTERSPEECH3
2013 Annotation and Classification of Changes of Involvement in Group Conversation
abstract
The detection of involvement in a conversation is important to assess the level humans are participating in either a human-human or human-computer interaction. Especially, detecting changes in a group's involvement in a multi-party interaction is of interest to distinguish several constellations in the group itself. This information can further be used in situations where technical support of meetings is favoured, for instance, focusing a camera, switching microphones, etc. Moreover, this information could also help to improve the performance of technical systems applied in human-machine interaction. In this paper, we concentrate on video material given by the Table Talk corpus. Therefore, we introduce a way of annotating and classifying changes of involvement and discuss the reliability of the annotation. Further, we present classification results based on video features using Multi-Layer Networks.
Ronald Böck, Stefan Glüge, Ingo Siegert, Andreas Wendemuth
ACII3
2012 Towards Emotion and Affect Detection in the Multimodal LAST MINUTE Corpus
Jörg Frommer, Bernd Michaelis, Dietmar F. Rösner, Andreas Wendemuth, Rafael Friesen, Matthias Haase, Manuela Kunze, Rico Andrich, Julia Lange, Axel Panning, Ingo Siegert
LREC11
2011 ikannotate - A Tool for Labelling, Transcription, and Annotation of Emotionally Coloured Speech
Ronald Böck, Ingo Siegert, Matthias Haase, Julia Lange, Andreas Wendemuth
ACII (1)2
2011 Appropriate emotional labelling of non-acted speech using basic emotions, geneva emotion wheel and self assessment manikins
abstract
In emotion recognition from speech, a good transcription and annotation of given material is crucial. Moreover, the question of how to find good emotional labels for new data material is a basic issue. It is not only the question of which emotion labels to choose, it is also a matter of how labellers can cope with annotation methods. In this paper, we present our investigations for emotional labelling with three different methods (Basic Emotions, Geneva Emotion Wheel and Self Assessment Manikins) and compare them in terms of emotion coverage and usability. We show that emotion labels derived from Geneva Emotion Wheel or Self Assessment Manikins fulfill our requirements, but Basic Emotions are not feasible for emotion labelling from spontaneous speech.
Ingo Siegert, Ronald Böck, Bogdan Vlasenko, David Philippou-Hübner, Andreas Wendemuth
ICME1
2011 Vowels formants analysis allows straightforward detection of high arousal emotions
abstract
Recently, automatic emotion recognition from speech has achieved growing interest within the human-machine interaction research community. Most part of emotion recognition methods use context independent frame-level analysis or turn-level analysis. In this article, we introduce context dependent vowel level analysis applied for emotion classification. An average first formant value extracted on vowel level has been used as unidimensional acoustic feature vector. The Neyman-Pearson criterion has been used for classification purpose. Our classifier is able to detect high-arousal emotions with small error rates. Within our research we proved that the smallest emotional unit should be the vowel instead of the word. We find out that using vowel level analysis can be an important issue during developing a robust emotion classifier. Also, our research can be useful for developing robust affective speech recognition methods and high quality emotional speech synthesis systems.
Bogdan Vlasenko, David Philippou-Hübner, Dmytro Prylipko, Ronald Böck, Ingo Siegert, Andreas Wendemuth
ICME5
2010 Developing an Expressive Speech Labeling Tool Incorporating the Temporal Characteristics of Emotion
Stefan Scherer, Ingo Siegert, Lutz Bigalke, Sascha Meudt
LREC2