EDBT 2026 Demo / reviewers in the wild / expert
Peter Birkholz
dblp:80/1867
· DBLP profile ↗
72ranked-venue papers
22as first author
37since 2021 · last 2025
0000-0003-0167-8123ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 57 · 18 first-author · 27 since 2021Artificial intelligence and machine learning · 55 · 17 first-author · 28 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Precisely Controllable Neural Speech SynthesisabstractRecent advances in deep learning have significantly improved the quality of speech synthesis, yet these models often suffer from limited controllability and lack of interpretability due to their black-box nature. In contrast, articulatory speech synthesis offers fine-grained control and transparency by simulating sound production based on vocal tract geometry, though it struggles with naturalness and synthesis quality. To bridge these gaps, we propose a novel white-box approach that leverages synthetic articulatory trajectories for neural synthesis, ensuring a fully disentangled, interpretable, and controllable yet high-quality speech synthesis process. Utilizing the VocalTractLab articulatory synthesizer, our method allows the quality of its speech representation to be verified through physical simulation. The proposed system achieves state-of-the-art results in both articulatory and deep articulatory synthesis. To the best of our knowledge, this is the first work to synthesize highly intelligible speech from a purely synthetic articulatory latent representation. Paul Konstantin Krug, Peter Birkholz, Timo Stich |
ICASSP | 3 |
| 2025 | Exploring Antenna Placement Configurations with a Radar-based Silent Speech InterfaceabstractSilent Speech Interfaces (SSIs) are systems developed to reconstruct speech based on non-acoustic bio-signals generated by brain, muscular or articulatory activity. A broad variety of sensing methods have been employed in SSIs, but determining a single optimal sensor configuration is generally a challenge. This work focuses on investigating sensing configurations for an on-body radar-based SSI and evaluating them based on their phoneme recognition performance. Audio and radar data were recorded synchronously by 8 native German speakers. One emitting and two receiving antennas were placed on the cheeks of each speaker in three different configurations, which were evaluated by nested cross-validated Support Vector Machines (SVMs). With these sensing configurations, higher accuracies than in previous works with on-body radar-based SSIs were achieved for individual speakers and for a multi-speaker evaluation. This work succeeds to produce an antenna placement guideline for future recordings with similar SSIs. João Vítor Menezes, Martin Schütze, Petr Schaffer, Dirk Plettemeier, Peter Birkholz |
ICASSP | 5 |
| 2025 | Non-invasive Speaker-dependent Continuous Phoneme Recognition with a Radar-based Silent Speech InterfaceabstractSilent speech interfaces (SSIs) are systems that can enable speech communication in the partial or total absence of the acoustic speech signal. The research on SSIs includes command word recognition, speech synthesis and continuous phoneme/word recognition. The focus of this work is to perform continuous phoneme recognition with a radar-based SSI. Audio and radar data were recorded synchronously during continuous speech of one speaker. Six feature sets based on combinations of the magnitude spectrum, the phase spectrum, and the impulse response of the radar data were used as input to the recognition task and compared to each other and to a benchmark acoustic feature set based on the audio data. With a CNN-MLP followed by a Viterbi-based biphone phoneme decoder, the lowest mean phoneme error rate (PER) obtained by any radar-based feature set was 45.62%, which is 15% higher than the mean PER achieved by the benchmark acoustic feature set. The most common sources of errors were minimal pairs of consonants. The results of this work are promising and point towards future possible improvements for radar-based continuous phoneme recognition. João Vítor Menezes, Peter Steiner, Petr Schaffer, Dirk Plettemeier, Peter Birkholz |
ICASSP | 6 |
| 2025 | Influence of wall coverings of 3D-printed vocal tract models on measured transfer functionsabstractInternational audience Peter Birkholz, Dominik Schäfer, Patrick Häsner, Jihyeon Yun, Iris Kruppke, Rémi Blandin |
INTERSPEECH | 1 |
| 2025 | Evaluation of a model for sound radiation from the vocal tract wall
Peter Birkholz |
INTERSPEECH | 1 |
| 2025 | Multimodal Silent Recognition of Phonemes Using Radar and Optopalatographic Silent Speech Interfaces
João Menezes, Aubin Mouras, Arne-Lukas Fietkau, Dani Kazzy, Peter Birkholz |
INTERSPEECH | 5 |
| 2025 | Learnability of English diphthongs: One dynamic target vs. two static targets
Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Paul Konstantin Krug, Santitham Prom-on, Peter Birkholz, Yi Xu 0007 |
Speech Commun. | 6 |
| 2024 | Measurement and simulation of pressure losses due to airflow in vocal tract models
Peter Birkholz, Patrick Häsner |
INTERSPEECH | 1 |
| 2024 | A demonstrator for articulation-based command word recognition
João Vítor Possamai de Menezes, Arne-Lukas Fietkau, Tom Diener, Steffen Kürbis, Peter Birkholz |
INTERSPEECH | 5 |
| 2024 | Monophthong vocal tract shapes are sufficient for articulatory synthesis of German primary diphthongsabstractGerman primary diphthongs are conventionally transcribed using the same symbols used for some monophthong vowels. However, if the corresponding vocal tract shapes are used for articulatory synthesis, the results often sound unnatural. Furthermore, there is no clear consensus in the literature if diphthongs have monopthong constituents and if so, which ones. This study therefore analyzed a set of audio recordings from the reference speaker of the state-of-the-art articulatory synthesizer VocalTractLab to identify likely candidates for the monophthong constituents of the German primary diphthongs. We then evaluated these candidates in a listening experiment with naive listeners to determine a naturalness ranking of these candidates and specialized diphthong shapes. The results showed that the German primary diphthongs can indeed be synthesized with no significant loss in naturalness by replacing the specialized diphthong shapes for the initial and final segments by shapes also used for monopthong vowels. Simon Stone, Peter Birkholz |
Speech Commun. | 2 |
| 2024 | Articulatory Copy Synthesis Based on the Speech Synthesizer VocalTractLab and Convolutional Recurrent Neural NetworksabstractArticulatory copy synthesis (ACS) refers to the synthetic reproduction of natural utterances. The existing methods of ACS have the limitations of poor generalizability for unknown speakers, high computing costs, the lack of systematic evaluation, etc. Here we propose an ACS method based on the articulatory speech synthesizer VocalTractLab (VTL) and convolutional recurrent neural networks. We first created paired articulatory-acoustic samples using VTL, and then trained neural-network-based ACS models with acoustic features and articulatory trajectories as inputs and outputs, respectively. The basic approach for training relied on fully synthetic training data (and was later supplemented with natural speech and corresponding synthetic articulatory data). In addition, to represent as much of the articulatory and acoustic space as possible, the training samples were augmented by varying the phonation type, speaking effort, and the vocal tract length of the synthetic utterances. Furthermore, two regularization methods were proposed: one based on the smoothness loss of articulatory trajectories and another based on the acoustic loss between original and estimated acoustic features. For given new utterances of arbitrary length, the trained ACS models could estimate articulatory trajectories that were then fed into VTL to synthesize new speech. Experiments showed that our proposed ACS method achieved an average correlation coefficient of 0.983 between the reference and estimated VTL articulatory parameters for speaker-dependent German utterances. When applied to speaker-independent German, English, and Mandarin Chinese utterances, the copy-synthesized speech achieved recognition rates of 73.88%, 52.92%, and 52.41%, respectively, using the automatic speech recognizer Google Speech-to-Text. Yingming Gao, Peter Birkholz, Ya Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Non-Standard Echo State Networks for Video Door State MonitoringabstractIn recent years, Echo State Networks (ESNs), a special type of Recurrent Neural Networks (RNNs), have become increasingly established in the Machine Learning (ML) community due to their relatively simple initialization and training methods. Traditionally, the input and recurrent weights are generated randomly, with only the output weights being trained, typically using linear regression. However, recent publications have proposed alternative ways to initialize the weight matrices, e.g., by using more deterministic methods or data-driven approaches. This is the first work comparing different simple reservoir structures and an ESN with pre-trained input weights for the task of monitoring the state of a door using a surveillance camera in real-time. The results show that deterministic ESN structures perform better than the randomly initialized baseline, achieving a frame error rate of 2.62% vs. 2.93%. Peter Steiner, Azarakhsh Jalalvand, Peter Birkholz |
IJCNN | 3 |
| 2023 | Self-Supervised Solution to the Control Problem of Articulatory SynthesisabstractGiven an articulatory-to-acoustic forward model, it is a priori \nunknown how its motor control must be operated to achieve a \ndesired acoustic result. This control problem is a fundamental \nissue of articulatory speech synthesis and the cradle of acousticto-articulatory inversion, a discipline which attempts to address \nthe issue by the means of various methods. This work presents \nan end-to-end solution to the articulatory control problem, in \nwhich synthetic motor trajectories of Monte-Carlo-generated \nartificial speech are linked to input modalities (such as natural speech recordings or phoneme sequence input) via speakerindependent latent representations of a vector-quantized variational autoencoder. The proposed method is self-supervised and \nthus, in principle, synthesizer and speaker model independent. Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007 |
INTERSPEECH | 2 |
| 2023 | Exploring unsupervised pre-training for echo state networksabstractAbstract Echo State Networks (ESNs) are a special type of Recurrent Neural Networks (RNNs), in which the input and recurrent connections are traditionally generated randomly, and only the output weights are trained. However, recent publications have addressed the problem that a purely random initialization may not be ideal. Instead, a completely deterministic or data-driven initialized ESN structure was proposed. In this work, an unsupervised training methodology for the hidden components of an ESN is proposed. Motivated by traditional Hidden Markov Models (HMMs), which have been widely used for speech recognition for decades, we present an unsupervised pre-training method for the recurrent weights and bias weights of ESNs. This approach allows for using unlabeled data during the training procedure and shows superior results for continuous spoken phoneme recognition, as well as for a large variety of time-series classification datasets. Peter Steiner, Azarakhsh Jalalvand, Peter Birkholz |
Neural Comput. Appl. | 3 |
| 2023 | Simulating vocal learning of spoken language: Beyond imitationabstractComputational approaches have an important role to play in understanding the complex process of speech acquisition, in general, and have recently been popular in studies of vocal learning in particular. In this article we suggest that two significant problems associated with imitative vocal learning of spoken language, the speaker normalisation and phonological correspondence problems, can be addressed by linguistically grounded auditory perception. In particular, we show how the articulation of consonant–vowel syllables may be learnt from auditory percepts that can represent either individual utterances by speakers with different vocal tract characteristics or ideal phonetic realisations. The result is an optimisation-based implementation of vocal exploration – incorporating semantic, auditory, and articulatory signals – that can serve as a basis for simulating vocal learning beyond imitation. Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Lorna F. Halliday, Santitham Prom-on, Yi Xu 0007 |
Speech Commun. | 5 |
| 2023 | A Comparative Study of 3D and 1D Acoustic Simulations of the Higher Frequencies of SpeechabstractArticulatory synthesis generates speech sounds by simulating the physical phenomena involved in speech production. The accuracy of the physical modelling is expected to affect the naturalness of the synthesis: the more realistic the description is, the greater the naturalness is expected to be. In this work, the accuracy of acoustic wave propagation in the vocal tract was evaluated with two perceptual experiments. Sustained vowels generated using a one-dimensional acoustic model, a three-dimensional acoustic model and an artificial bandwidth extension algorithm (without a physical basis) were compared. Since the difference between the acoustic methods tested affects mainly the frequencies above 4 kHz, we ensured that the low frequency part of the stimuli, up to 4 kHz, was similar. Thus, the participants' responses were based only on the differences at high frequency. The first experiment was a pair comparison, in which the participants had to select the more natural sounding stimuli. In the second experiment, the participants had to rate the naturalness of the stimuli on a linear scale. The results confirmed that a more accurate physical modeling leads to greater naturalness. However, this was limited to the phonemes /o/ and /u/, for which transverse resonances in the anterior vocal tract may play an important role that only a 3D acoustic simulation can accurately represent. It was also found that male stimuli were perceived as significantly more natural than female ones. However, voice quality did not affect naturalness. Rémi Blandin, Simon Stone, Angélique Remacle, Vincent Didone, Peter Birkholz |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Artificial Vocal Learning Guided by Phoneme Recognition and Visual InformationabstractThis paper introduces a paradigm shift regarding vocal learning simulations, in which the communicative function of speech acquisition determines the learning process and intelligibility is considered the primary measure of learning success. Thereby, a novel approach for artificial vocal learning is presented that utilizes deep neural network-based phoneme recognition in order to calculate the speech acquisition objective function. This function guides a learning framework that involves the state-of-the-art articulatory speech synthesizer VocalTractLab as the motor-to-acoustic forward model. In this way, an extensive set of German phonemes, including most of the consonants and all stressed vowels, was produced successfully. The synthetic phonemes were rated as highly intelligible by human listeners. Furthermore, it is shown that visual speech information, such as lip and jaw movements, can be extracted from video recordings and be incorporated into the learning framework as an additional loss component during the optimization process. It was observed that this visual loss did not increase the overall intelligibility of phonemes. Instead, the visual loss acted as a regularization mechanism that facilitated the finding of more biologically plausible solutions in the articulatory domain. Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Cluster-Based Input Weight Initialization for Echo State NetworksabstractEcho state networks (ESNs) are a special type of recurrent neural networks (RNNs), in which the input and recurrent connections are traditionally generated randomly, and only the output weights are trained. Despite the recent success of ESNs in various tasks of audio, image, and radar recognition, we postulate that a purely random initialization is not the ideal way of initializing ESNs. The aim of this work is to propose an unsupervised initialization of the input connections using the K -means algorithm on the training data. We show that for a large variety of datasets, this initialization performs equivalently or superior than a randomly initialized ESN while needing significantly less reservoir neurons. Furthermore, we discuss that this approach provides the opportunity to estimate a suitable size of the reservoir based on prior knowledge about the data. Peter Steiner, Azarakhsh Jalalvand, Peter Birkholz |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Acoustic Comparison of Physical Vocal Tract Models with Hard and Soft WallsabstractThis study explored how the frequencies and bandwidths of the acoustic resonances of physical tube models of the vocal tract differ when they have hard versus soft walls. For each of 10 tube shapes representing different vowels, two physical models were made: one with rigid plastic walls, and one with soft silicone walls. For all models, the acoustic transfer functions were measured and the bandwidths and frequencies of the first three resonances were determined. For the models with soft walls, the resonance frequencies and bandwidths below 2 kHz were consistently higher than for the models with hard walls. These results confirm the general predictions of mathematical lumped-element models for the mechanical impedance of the vocal tract walls. The resonance bandwidths of the silicone models were well within the range of 50–90 Hz measured for humans, while they were too small for the rigid tubes below 2 kHz. Peter Birkholz, Patrick Häsner, Steffen Kürbis |
ICASSP | 1 |
| 2022 | Carina - A Corpus of Aligned German Read Speech Including AnnotationsabstractThis paper presents the semi-automatically created Corpus of Aligned Read Speech Including Annotations (CARInA), a speech corpus based on the German Spoken Wikipedia Corpus (GSWC). CARInA tokenizes, consolidates and organizes the vast, but rather unstructured material contained in GSWC. The contents are grouped by annotation completeness, and extended by canonic, morphosyntactic and prosodic annotations. The annotations are provided in BPF and TextGrid format. It contains 194 hours of speech material from 327 speakers, of which 124 hours are fully phonetically aligned and 30 hours are fully aligned at all annotation levels. CARInA is freely available1, designed to grow and improve over time, and suitable for large-scale speech analyses or machine learning tasks as illustrated by two examples shown in this paper. Hannes Kath, Simon Stone, Stefan Rapp, Peter Birkholz |
ICASSP | 4 |
| 2022 | A user-friendly headset for radar-based silent speech recognition
Pouriya Amini Digehsara, João Vítor Possamai de Menezes, Michael Bärhold, Petr Schaffer, Dirk Plettemeier, Peter Birkholz |
INTERSPEECH | 7 |
| 2022 | Relationship between the acoustic time intervals and tongue movements of German diphthongs
Arne-Lukas Fietkau, Simon Stone, Peter Birkholz |
INTERSPEECH | 3 |
| 2022 | Articulatory Synthesis for Data Augmentation in Phoneme RecognitionabstractWhile numerous studies on automatic speech recognition have been published in recent years describing data augmentation strategies based on time or frequency domain signal processing, few works exist on the artificial extensions of training data sets using purely synthetic speech data.In this work, the German KIEL corpus was augmented with synthetic data generated with the state-of-the-art articulatory synthesizer VOCALTRACT-LAB.It is shown that the additional synthetic data can lead to a significantly better performance in single-phoneme recognition in certain cases, while at the same time, the performance can also decrease in other cases, depending on the degree of acoustic naturalness of the synthetic phonemes.As a result, this work can potentially guide future studies to improve the quality of articulatory synthesis via the link between synthetic speech production and automatic speech recognition. Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007 |
INTERSPEECH | 2 |
| 2022 | Glottal inverse filtering based on articulatory synthesis and deep learning
Ingo Langheinrich, Simon Stone, Peter Birkholz |
INTERSPEECH | 4 |
| 2022 | An investigation of regression-based prediction of the femininity or masculinity in speech of transgender people
Leon Liebig, Alexander Mainka, Peter Birkholz |
INTERSPEECH | 4 |
| 2022 | Evaluation of different antenna types and positions in a stepped frequency continuous-wave radar-based silent speech interface
João Vítor Menezes, Pouriya Amini Digehsara, Marco Mütze, Michael Bärhold, Petr Schaffer, Dirk Plettemeier, Peter Birkholz |
INTERSPEECH | 8 |
| 2022 | Three-dimensional finite-difference time-domain acoustic analysis of simplified vocal tract shapes
Debasish Ray Mohapatra, Mario Fleischer, Victor Zappi, Peter Birkholz, Sidney S. Fels |
INTERSPEECH | 4 |
| 2022 | Exploration strategies for articulatory synthesis of complex syllable onsetsabstractHigh-quality articulatory speech synthesis has many potential applications in speech science and technology. However, developing appropriate mappings from linguistic specification to articulatory gestures is difficult and time consuming. In this paper we construct an optimisation-based framework as a first step towards learning these mappings without manual intervention. We demonstrate the production of CCV syllables and discuss the quality of the articulatory gestures with reference to coarticulation. Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Yi Xu 0007 |
INTERSPEECH | 5 |
| 2022 | Evoc-Learn - High quality simulation of early vocal learning
Yi Xu 0007, Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Peter Birkholz, Paul Konstantin Krug, Santitham Prom-on, Lorna F. Halliday |
INTERSPEECH | 5 |
| 2022 | PyRCN: A toolbox for exploration and application of Reservoir Computing Networks
Peter Steiner, Azarakhsh Jalalvand, Simon Stone, Peter Birkholz |
Eng. Appl. Artif. Intell. | 4 |
| 2022 | Articulatory Synthesis of Vocalized /r/ Allophones in GermanabstractArticulatory synthesis relies on precise, parametric vocal tract shapes to generate natural-sounding speech. In German, a particular challenge is the accurate synthesis of the vocalic /r/ allophones following vowels or in syllable coda position. Using established phonetic conventions, no satisfying results could be achieved so far, implying a possible shortcoming of these existing conventions. This study therefore analyzed a large number of natural recordings of the sounds in question from a single speaker to find the optimal number of target vocal tract shapes. Applying clustering techniques, the manifold of vocalic /r/ allophones could be reduced to two prototypical [ɐ] variants previously undescribed in the literature. As shown by a listening experiment, which of these two sounds was preferred in which context depended not only on the respective context vowel’s tenseness, but also on its openness and acoustic distance to other context vowels ending in the same [ɐ] variant. This indicates that the two different allophones might serve as a contrastive cue to help differentiate between otherwise similar context vowels. Simon Stone, Yingming Gao, Peter Birkholz |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Unsupervised Pretraining of Echo State Networks for Onset Detection
Peter Steiner, Azarakhsh Jalalvand, Peter Birkholz |
ICANN (5) | 3 |
| 2021 | Comparison of the Finite Element Method, the Multimodal Method and the Transmission-Line Model for the Computation of Vocal Tract Transfer FunctionsabstractInternational audience Rémi Blandin, Marc Arnela, Simon Félix, Jean-Baptiste Doc, Peter Birkholz |
Interspeech | 5 |
| 2021 | Articulatory Data Recorder: A Framework for Real-Time Articulatory Data Recording
Alexander Wilbrandt, Simon Stone, Peter Birkholz |
Interspeech | 3 |
| 2021 | Model-Based Exploration of Linking Between Vowel Articulatory Space and Acoustic SpaceabstractWhile the acoustic vowel space has been extensively studied in previous research, little is known about the high-dimensional articulatory space of vowels. The articulatory imaging techniques are limited to tracking only a few key articulators, leaving the rest of the articulators unmonitored. In the present study, we attempted to develop a detailed articulatory space obtained by training a 3D articulatory synthesizer to learn eleven British English vowels. An analysis-by-synthesis strategy was used to acoustically optimize vocal tract parameters that represent twenty articulatory dimensions. The results show that tongue height and retraction, larynx location and lip roundness are the most perceptually distinctive articulatory dimensions. Yet, even for these dimensions, there is a fair amount of articulatory overlap between vowels, unlike the fine-grained acoustic space. This method opens up the possibility of using modelling to investigate the link between speech production and perception. Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Paul Konstantin Krug, Santitham Prom-on, Peter Birkholz, Yi Xu 0007 |
Interspeech | 6 |
| 2021 | Acoustic and articulatory analysis and synthesis of shouted vowels
Yawen Xue, Michael Marxen, Masato Akagi, Peter Birkholz |
Comput. Speech Lang. | 4 |
| 2021 | Effects of the piriform fossae, transvelar acoustic coupling, and laryngeal wall vibration on the naturalness of articulatory speech synthesis
Peter Birkholz, Susanne Drechsel |
Speech Commun. | 1 |
| 2020 | Accounting for Microprosody in Modeling IntonationabstractIntonation models are often used for the generation of fundamental frequency (f0) contours in speech synthesis. Current intonation models only represent the intentional f0components that are related to the phonological structure of the utterance. However, natural speech also contains non-intentional microvariations of f0, which are usually not accounted for. Here, we derived models for two forms of microvariations: the drop in f0during voiced obstruents, and the increased f0at the onset of vowels following voiceless obstruents. These models were applied to remove the microvariations of f0in a database of natural speech before the f0contours were reproduced with the Target Approximation Model. The previously removed microvariations were then superimposed on the modeled f0contours. The resulting model f0contours were significantly more similar to the original (natural) f0contours than model contours that did not account for the mi-crovariations. This approach might improve f0modeling in future parametric speech synthesizers. Peter Birkholz |
ICASSP | 1 |
| 2020 | Cross-Speaker Silent-Speech Command Word Recognition Using Electro-Optical StomatographyabstractSpeech recognition based on articulatory movements instead of the acoustic signal is of growing interest in the community. In this work, we present the results of a study using a novel measurement technology called Electro-Optical Stomatography to capture speech movements and use the acquired data to recognize a number of command words. The performance of the recognition system was evaluated using two vocabularies (one with 30 and one with 10 words) and four speakers. The speaker-dependent results were up to the state-of-the-art with average word accuracies of 97% to 99.5%, while the speaker-independent results exceeded it with average word accuracies of approx. 56% to 62%. Simon Stone, Peter Birkholz |
ICASSP | 2 |
| 2020 | Prediction of Voicing and the F0 Contour from Electromagnetic Articulography Data for Articulation-to-Speech SynthesisabstractArticulation-to-speech synthesis based solely on supraglottal articulation requires some sort of intonation control. This paper examines to what extent the f0contour of an utterance can be predicted from such supraglottal articulation data. To that end, three groups of machine learning models (support vector machines, kernel ridge regression and neural networks) were trained and evaluated on the mngu0speech corpus containing synchronous articulatory and audio data. The best voiced/unvoiced/silence classification rates were achieved by a deep neural network with two hidden layers: 85.8 % with no look-ahead (important for on-line applications) and 86 % with a look-ahead of 50 ms. The best f0prediction model without look-ahead scored a root-mean-square error (RMSE) (when compared to the original f0contours) of 10.4 Hz using a neural network with one hidden layer, while the best prediction with a look-ahead of 50 ms was attained by kernel ridge regression and an RMSE of 10.3 Hz. The predicted f0contours were also subjectively evaluated in a listening test by manipulating the f0of the original speech files using PRAAT. The results are consistent with the objective evaluation. Simon Stone, Peter Birkholz |
ICASSP | 3 |
| 2020 | Feature Engineering and Stacked Echo State Networks for Musical Onset DetectionabstractIn music analysis, one of the most fundamental tasks is note onset detection - detecting the beginning of new note events. As the target function of onset detection is related to other tasks, such as beat tracking or tempo estimation, onset detection is the basis for such related tasks. Furthermore, it can help to improve Automatic Music Transcription (AMT). Typically, different approaches for onset detection follow a similar outline: An audio signal is transformed into an Onset Detection Function (ODF), which should have rather low values (i.e. close to zero) for most of the time but with pronounced peaks at onset times, which can then be extracted by applying peak picking algorithms on the ODF. In the recent years, several kinds of neural networks were used successfully to compute the ODF from feature vectors. Currently, Convolutional Neural Networks (CNNs) define the state of the art. In this paper, we build up on an alternative approach to obtain a ODF by Echo State Networks (ESNs), which have achieved comparable results to CNNs in several tasks, such as speech and image recognition. In contrast to the typical iterative training procedures of deep learning architectures, such as CNNs or networks consisting of Long-Short-Term Memory Cells (LSTMs), in ESNs only a very small part of the weights is easily trained in one shot using linear regression. By comparing the performance of several feature extraction methods, pre-processing steps and introducing a new way to stack ESNs, we expand our previous approach to achieve results that fall between a bidirectional LSTM network and a CNN with relative improvements of 1.8 % and -1.4 %, respectively. For the evaluation, we used exactly the same 8-fold cross validation setup as for the reference results. Peter Steiner, Azarakhsh Jalalvand, Simon Stone, Peter Birkholz |
ICPR | 4 |
| 2020 | An Investigation of the Target Approximation Model for Tone Modeling and Recognition in Continuous Mandarin SpeechabstractThe complex f0 variations in continuous speech make it rather difficult to perform automatic recognition of tones in a language like Mandarin Chinese.In this study, we tested the use of target approximation model (TAM) for continuous tone recognition on two datasets.TAM simulates f0 production from the articulatory point of view and so allow to discover the underlying pitch targets from the surface f0 contour.The f0 contour of each tone represented by 30 equidistant points in the first dataset was simulated by the TAM model.Using a support vector machine (SVM) to classify tones showed that, compared to the representation by 30 f0 values, the estimated three-dimensional TAM parameters had a comparable performance in characterizing tone patterns.The TAM model was further tested on the second dataset containing more complex tonal variations.With equal or a fewer number of features, the TAM parameters provided better performance than the coefficients of the cosine transform and a slightly worse performance than the statistical f0 parameters for tone recognition.Furthermore, we investigated bidirectional LSTM neural network for modelling the sequential tonal variations, which proved to be more powerful than the SVM classifier.The BLSTM system incorporating TAM and statistical f0 parameters achieved the best accuracy of 87.56%. Yingming Gao, Jinsong Zhang 0001, Peter Birkholz |
INTERSPEECH | 5 |
| 2020 | Finding Intelligible Consonant-Vowel Sounds Using High-Quality Articulatory SynthesisabstractIn this study, a state-of-the-art articulatory speech synthesiser was used as the basis for simulating the exploration of CV sounds imitating speech stimuli. By adopting a relevant kinematic model and systematically reducing the search space of consonant articulatory targets, intelligible CV sounds can be found. Derivative-free optimisation strategies were evaluated to speed up the process of exploring articulatory space and the possibility of using automatic speech recognition as a means of evaluating intelligibility was explored. Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Yi Xu 0007 |
INTERSPEECH | 5 |
| 2020 | Effect of articulatory and acoustic features on the intelligibility of speech in noise: An articulatory synthesis study
Thuan Van Ngo, Masato Akagi, Peter Birkholz |
Speech Commun. | 3 |
| 2019 | Perceptual Optimization of an Enhanced Geometric Vocal Fold Model for Articulatory Speech Synthesis
Peter Birkholz, Susanne Drechsel, Simon Stone |
INTERSPEECH | 1 |
| 2019 | Articulatory Copy Synthesis Based on a Genetic Algorithm
Yingming Gao, Simon Stone, Peter Birkholz |
INTERSPEECH | 3 |
| 2019 | How modeling entrance loss and flow separation in a two-mass model affects the oscillation and synthesis quality
Peter Birkholz, Daniel Pape |
Speech Commun. | 1 |
| 2018 | Non-Invasive Silent Phoneme Recognition Using Microwave SignalsabstractBesides the recognition of audible speech, there is currently an increasing interest in the recognition of silent speech, which has a range of novel applications. A major obstacle for a wide spread of silent-speech technology is the lack of measurement methods for speech movements that are convenient, non-invasive, portable, and robust at the same time. Therefore, as an alternative to established methods, we examined to what extent different phonemes can be discriminated from the electromagnetic transmission and reflection properties of the vocal tract. To this end, we attached two Vivaldi antennas on the cheek and below the chin of two subjects. While the subjects produced 25 phonemes in multiple phonetic contexts each, we measured the electromagnetic transmission spectra from one antenna to the other, and the reflection spectra for each antenna (radar), in a frequency band from 2-12 GHz. Two classification methods (k-nearest neighbors and linear discriminant analysis) were trained to predict the phoneme identity from the spectral data. With linear discriminant analysis, cross-validated phoneme recognition rates of 93% and 85% were achieved for the two subjects. Although these results are speaker- and session-dependent, they suggest that electromagnetic transmission and reflection measurements of the vocal tract have great potential for future silent-speech interfaces. Peter Birkholz, Simon Stone, Klaus Wolf, Dirk Plettemeier |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Construction and Evaluation of a Parametric One-Dimensional Vocal Tract ModelabstractArticulatory speech synthesis based on aero-acoustic simulations of the vocal tract is computationally expensive and, therefore, requires simple yet precise models. Modeling the onedimensional vocal tract area function directly instead of a higher dimensional vocal tract model is an efficient way to minimize the computational overhead of the simulations. In this paper, we propose a new parametric vocal tract model that is controlled by six points and capable of modeling a large variety of vocal tract shapes. We geometrically and perceptually evaluated the model on a set of 22 reference area functions corresponding to German vowels and consonants. The model was able to geometrically approximate the reference area functions with a minimum root-mean-square error of 0.302 cm2, a maximum error of 1.142 cm2, and a median error of 0.891 cm2. After optimizations, a perceptual evaluation of the synthesis using our model in combination with a state-of-the-art aero-acoustic simulation achieved a vowel recognition rate of 90.7% and a consonant recognition rate of 73.2%. Simon Stone, Michael Marxen, Peter Birkholz |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Does Posh English Sound Attractive?abstractPoshness refers to how much a British English speaker sounds upper class when they talk. Popular descriptions of posh English mostly focus on vocabulary, accent and phonology. This study tests the hypothesis that, as a social index, poshness is also manifested via phonetic properties known to encode vocal attractiveness. Specifically, posh English, because of its impression of being detached, authoritative and condescending, would more closely resemble an attractive male voice than an attractive female voice. In four experiments, we tested this hypothesis by acoustically manipulating Cambridge-Accented English utterances by a male and a female speaker through PSOLA resynthesis, and having native speakers of British English judge how posh or attractive each utterance sounds. The manipulated acoustic dimensions are formant dispersion, pitch shift and speech rate. Initial results from the first two experiments showed a trend in the hypothesized direction for the male speakers' utterances. But for the female utterances there was a ceiling effect due to the frequent alternation of speaker gender within the same test session. When the two speakers' utterances were separated by blocks in the third and fourth experiments, a clearer support for t he main hypothesis was found. Chengxia Wang, Cristiane Hsu, Peter Birkholz |
INTERSPEECH | 4 |
| 2017 | The Sound of Deception - What Makes a Speaker Credible?
Anne Schröder, Simon Stone, Peter Birkholz |
INTERSPEECH | 3 |
| 2017 | A Time-Warping Pitch Tracking Algorithm Considering Fast f0 Changes
Simon Stone, Peter Steiner, Peter Birkholz |
INTERSPEECH | 3 |
| 2017 | Manipulation of the prosodic features of vocal tract length, nasality and articulatory precision using articulatory synthesis
Peter Birkholz, Lucia Martin, Yi Xu 0007, Stefan Scherbaum, Christiane Neuschaefer-Rube |
Comput. Speech Lang. | 1 |
| 2016 | Towards Minimally Invasive Velar State Detection in Normal and Silent Speech
Peter Birkholz, Petko Bakardjiev, Steffen Kürbis, Rico Petrick |
INTERSPEECH | 1 |
| 2016 | Silent-Speech Command Word Recognition Using Electro-Optical Stomatography
Simon Stone, Peter Birkholz |
INTERSPEECH | 2 |
| 2015 | Optical sensor calibration for electro-optical stomatographyabstractWe are currently developing a technology called “electrooptical stomatography” to measure and visualize articulatory movements within the vocal tract using electrical contact sensors and optical proximity sensors. To measure tongue movements with the optical sensors in this system, a mapping between the raw sensor values and actual tongue positions has to be determined. This mapping is non-linear and different for every tongue and sensor. The lack of an accurate, reliable calibration method has so far prevented wide-spread use of optical measurements within the vocal tract. Here, we present a calibration method based on a multi-linear regression model that maps the sensor value at a single distance of 0mm to calibration values at 0, 5, 10, 15, 20, 25, and 30mm. The coefficients of the model are determined by a least-squares regression in 25 training data sets (recorded with 5 subjects and 5 sensors). Evaluation in a leave-one-out cross-validation and on five more data sets recorded with another, different subject on 5 additional sensors yields very good results with maximum median position errors close to 1mm. The calibration of the optical sensors can therefore be semi-automatically accomplished based on a single, easily obtainable measurement during direct tongue contact. Simon Preuß, Peter Birkholz |
INTERSPEECH | 2 |
| 2015 | Intervocalic fricative perception in European Portuguese: An articulatory synthesis study
Daniel Pape, Luis M. T. Jesus, Peter Birkholz |
Speech Commun. | 3 |
| 2014 | Tongue Contour Reconstruction from Optical and Electrical PalatographyabstractTongue shape reconstruction based on safe and convenient measurement techniques is of great interest for speech research and speech therapy. Two potentially useful and related measurement techniques for this purpose are electropalatography (EPG) and optopalatography (OPG). While EPG measures the time-varying contact pattern between the hard palate and the tongue, OPG measures distances between the two. Here, we examined the potential of EPG, OPG, and their combination for predicting the whole tongue contour using a multiple linear regression model. The model was trained and tested with tongue shapes and virtual sensor data obtained from Magnetic Resonance Images of sustained articulations of two speakers. When the model was trained and tested with the same speaker, the error of tongue contour reconstruction was significantly lower for predictions based on OPG data than for predictions based on EPG data. When the model was trained with one speaker and tested with the other, the error pattern was less consistent and the overall error was higher. Hence, especially OPG is well suited for tongue contour prediction, but an adaptation method is needed to transfer the model to a new speaker. Rizwan Mumtaz, Simon Preuß, Christiane Neuschaefer-Rube, Christiane Hey, Robert Sader, Peter Birkholz |
IEEE Signal Process. Lett. | 6 |
| 2013 | Real-time control of a 2d animation model of the vocal tract using optopalatographyabstractThis paper presents an animated 2D articulation model of the tongue and the lips for biofeedback applications. The model is controlled by real-time optopalatographic measurements of the positions of the upper lip and the tongue in the anterior oral cavity. The measurement system is an improvement on a previous prototype with increased spatial resolution and an enhanced close-range behavior. The posterior part of the tongue was added to the model by linear prediction. The prediction coefficients were determined and evaluated using a corpus of vocal tract traces of 25 sustained phonemes. The model represents the tongue motion and the lip opening physiologically plausible during articulation in real-time. Index Terms: optopalatography, glossometry, linear prediction, animated vocal tract model, biofeedback Simon Preuß, Christiane Neuschaefer-Rube, Peter Birkholz |
INTERSPEECH | 3 |
| 2013 | Training an articulatory synthesizer with continuous acoustic dataabstractThis paper reports preliminary results of our effort to address the acoustic-to-articulatory inversion problem. We tested an approach that simulates speech production acquisition as a distal learning task, with acoustic signals of natural utterances in the form of MFCC as input, VocalTractLab — a 3D articulatory synthesizer controlled by target approximation models as the learner, and stochastic gradient descent as the training method. The approach was tested on a number of natural utterances, and the results were highly encouraging. Santitham Prom-on, Peter Birkholz, Yi Xu 0007 |
INTERSPEECH | 2 |
| 2012 | Advances in combined electro-optical palatographyabstractThis paper describes the development of a device that combines the electropalatographic measurement of tongue-palate contact with optical distance sensing to measure the mid-sagittal contour of the tongue and the position of the lips. The device consists of a thin acrylic pseudopalate that contains both contact sensors and optical reflective sensors. Application areas are, for example, experimental phonetics, speech therapy, and silent speech interfaces. With regard to the latter, the prototype of the system was applied to the recognition of vowels from the sensor signals. It was shown that a classifier using the combined input data from both the contact sensors and the optical sensors had a higher recognition rate than classifiers based on only one type of sensory input. Peter Birkholz, Philippe Daechert, Christiane Neuschaefer-Rube |
INTERSPEECH | 1 |
| 2012 | Intrinsic velocity differences of lip and jaw movements: preliminary resultsabstractThe observed kinematics of speech movements are the result of both the control by the brain and the biomechanical properties of the peripheral speech apparatus. For many kinematic phenomena, it is not clear whether they are actively controlled or intrinsic to the biomechanical system. This pilot study investigated the movement of sensors on the lips and the jaw in cyclical vowel transitions at specific speaking rates to identify possible intrinsic differences in the velocities of the articulators. Thereby, the lower lip was found to be significantly faster in approaching its targets than the upper lip, the mouth corners, and the jaw. Furthermore, for the mouth corners, backward movements were significantly faster than forward movements. Peter Birkholz, Phil Hoole |
INTERSPEECH | 1 |
| 2011 | Synthesis of Breathy, Normal, and Pressed Phonation Using a Two-Mass Model with a Triangular GlottisabstractTwo-mass models of the vocal folds and their variants are valuable tools for voice synthesis and analysis, but are not able to produce breathy voice qualities. The produced voice qualities usually lie between normal and pressed. The reason for this property is that the mass elements are aligned parallel to the dorso-ventral axis. Thereby, the glottis always closes simultaneously along the entire length of the vocal folds. For breathy phonation, however, the closure happens rather gradual. This article introduces a modified two-mass model with mass elements that are inclined with respect to the dorso-ventral axis as a function of the degree of abduction. In this way, the closing phase of the glottis becomes progressively more gradual when the degree of abduction is increased. This model is able to produce the continuum of voice qualities from pressed over normal to breathy voices. Peter Birkholz, Bernd J. Kröger, Christiane Neuschaefer-Rube |
INTERSPEECH | 1 |
| 2011 | Combined Optical Distance Sensing and Electropalatography to Measure ArticulationabstractWe present the first prototype of a new optoelectronic instrument for the combined real-time measurement of the tongue contour in the mid-sagittal plane, the contact pattern between the tongue and the palate, and the position of the lips. The instrument consists of a thin acrylic pseudopalate with embedded contact sensors, as for electropalatography, and optical distance sensors to measure tongue-palate distances, as for glossometry. One additional distance sensor is located at the anterior side of the upper incisors to register the degree of opening and protrusion of the lips. Together, the sensors provide complementary information about the articulation of vowels and consonants, which was verified in initial experiments. The instrument offers new perspectives for the study of normal and disordered speech production, as well as for silent speech interfaces and speech prostheses for laryngectomees. Peter Birkholz, Christiane Neuschaefer-Rube |
INTERSPEECH | 1 |
| 2011 | Model-Based Reproduction of Articulatory Trajectories for Consonant-Vowel SequencesabstractWe present a novel quantitative model for the generation of articulatory trajectories based on the concept of sequential target approximation. The model was applied for the detailed reproduction of movements in repeated consonant-vowel syllables measured by electromagnetic articulography (EMA). The trajectories for the constrictor (lower lip, tongue tip, or tongue dorsum) and the jaw were reproduced. Thereby, we tested the following hypotheses about invariant properties of articulatory commands: (1) The target of the primary articulator for a consonant is invariant with respect to phonetic context, stress, and speaking rate. (2) Vowel targets are invariant with respect to speaking rate and stress. (3) The onsets of articulatory commands for the jaw and the constrictor are synchronized. Our results in terms of high-quality matches between observed and model-generated trajectories support these hypotheses. The findings of this study can be applied to the development of control models for articulatory speech synthesis. Peter Birkholz, Bernd J. Kröger, Christiane Neuschaefer-Rube |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Articulatory synthesis and perception of plosive-vowel syllables with virtual consonant targetsabstractVirtual articulatory targets are a concept to explain the different trajectories of primary and secondary articulators during consonant production, as well as the different places of the tongue-palate contact depending on the context vowel, for example in [igi] vs. [ugu]. The virtual targets for the tongue tip and the tongue body in apical and dorsal plosives are assumed to lie above the palate, and for bilabial consonants, the target is a negative degree of lip opening. In the present study, we discuss the concept of virtual targets and its application for articulatory speech synthesis. In particular, we examined how the location of virtual targets affects the acoustics and intelligibility of synthetic plosive-vowel syllables. It turned out that virtual targets that lie about 10 mm beyond the consonantal closure location allow a more precise reproduction of natural speech signals than virtual targets at a distance of about 1 mm. However, we found no effect on the intellegibility of the consonants. Peter Birkholz, Bernd J. Kröger, Christiane Neuschaefer-Rube |
INTERSPEECH | 1 |
| 2007 | Control of an articulatory speech synthesizer based on dynamic approximation of spatial articulatory targetsabstractWe present a novel approach to the generation of speech move-ments for an articulatory speech synthesizer. The movements of the articulators are modeled by dynamical third order linear sys-tems that respond to sequences of simple motor commands. The motor commands are derived automatically from a high level schedule for the input phonemes. The proposed model consid-ers velocity differences of the articulators and accounts for coar-ticulation between vowels and consonants. Preliminary tests of the model in the framework of an articulatory speech syn-thesizer indicate its potential to produce realistic speech move-ments and thereby to contribute to a higher quality of the syn-thesized speech. Peter Birkholz |
INTERSPEECH | 1 |
| 2007 | Articulatory synthesis of singingabstractA system for the synthesis of singing on the basis of an ar-ticulatory speech synthesizer is presented. To enable the synthe-sis of singing, the speech synthesizer was extended in many re-spects. Most importantly, a rule-based transformation of a mu-sical score into a gestural score for articulatory gestures was de-veloped. Furthermore, a pitch-dependent articulation of vowels was implemented. The results of these extensions are demon-strated by the synthesis of the canon “Dona nobis pacem”. The two voices in the canon were generated with the same under-lying articulatory models and the same musical score, the only difference being that their pitches differ by one octave. Index Terms: Articulatory singing synthesis 1. Peter Birkholz |
INTERSPEECH | 1 |
| 2007 | Simulation of Losses Due to Turbulence in the Time-Varying Vocal SystemabstractFlow separation in the vocal system at the outlet of a constriction causes turbulence and a fluid dynamic pressure loss. In articulatory synthesizers, the pressure drop associated with such a loss is usually assumed to be concentrated at one specific position near the constriction and is represented by a lumped nonlinear resistance to the flow. This paper highlights discontinuity problems of this simplified loss treatment when the constriction location changes during dynamic articulation. The discontinuities can manifest as undesirable acoustic artifacts in the synthetic speech signal that need to be avoided for high-quality articulatory synthesis. We present a solution to this problem based on a more realistic distributed consideration of fluid dynamic pressure changes. The proposed method was implemented in an articulatory synthesizer where it proved to prevent any acoustic artifacts Peter Birkholz, Dietmar Jackèl, Bernd J. Kröger |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Construction And Control Of A Three-Dimensional Vocal Tract ModelabstractWe present a novel 3D vocal tract model and a method to control the articulatory movements of the model. The vocal tract model consists of 7 wireframe meshes that represent the three dimensional surfaces of the articulators and the vocal tract walls. 23 parameters determine the shape of the meshes. The articulatory movements in terms of the parameter curves are generated from a gestural description of an utterance. The work presented here is an integral part of a complete articulatory speech synthesizer for high quality synthesis Peter Birkholz, Dietmar Jackèl, Bernd J. Kröger |
ICASSP (1) | 1 |
| 2006 | Modeling sensory-to-motor mappings using neural nets and a 3d articulatory speech synthesizerabstractA comprehensive neural model of speech motor control including a three dimensional articulatory speech synthesizer as a front-end device is described in detail in this paper.The training of the sensory-to-motor mappings -which can be interpreted as the prelinguistic phase of speech acquisition -is described in detail for quasi-static as well as for dynamic articulation. Bernd J. Kröger, Peter Birkholz, Jim Kannampuzha, Christiane Neuschaefer-Rube |
INTERSPEECH | 2 |
| 2004 | Influence of temporal discretization schemes on formant frequencies and bandwidths in time domain simulations of the vocal tract systemabstractA time domain simulation of acoustic propagation in the vocal tract requires the spatial and temporal discretization of the equations of motion and continuity. In the classic transmission line model of the vocal tract with lumped elements, the spatial discretization is provided by the piece-wise constant area function. The temporal finitedifference approximation of the differential equations can, however, vary from one implementation to the other (e.g., [4] vs. [5]). In this study, we have adopted a general finite-difference scheme that depends on a parameter θ where 0 ≤ θ ≤ 1. As special cases, this general method includes the trapezoid rule (θ = 0.5) as well as the implicit (θ = 1) and explicit (θ = 0) finite-difference schemes. We have examined how formant frequencies and bandwidths of simulated vowels are effected by the choice of θ. The experiments were conducted for the sampling rates of 44.1 kHz and 88.2 kHz and compared with the accurate and thus desirable frequencies and bandwidths measured in frequency domain simulations of the vocal tract. It can be shown that optimal values for θ are slightly above 0.5 depending on the sampling rate. Peter Birkholz, Dietmar Jackèl |
INTERSPEECH | 1 |