Paul Konstantin Krug

dblp:265/6571 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0001-8518-8142ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Precisely Controllable Neural Speech Synthesis
abstract
Recent advances in deep learning have significantly improved the quality of speech synthesis, yet these models often suffer from limited controllability and lack of interpretability due to their black-box nature. In contrast, articulatory speech synthesis offers fine-grained control and transparency by simulating sound production based on vocal tract geometry, though it struggles with naturalness and synthesis quality. To bridge these gaps, we propose a novel white-box approach that leverages synthetic articulatory trajectories for neural synthesis, ensuring a fully disentangled, interpretable, and controllable yet high-quality speech synthesis process. Utilizing the VocalTractLab articulatory synthesizer, our method allows the quality of its speech representation to be verified through physical simulation. The proposed system achieves state-of-the-art results in both articulatory and deep articulatory synthesis. To the best of our knowledge, this is the first work to synthesize highly intelligible speech from a purely synthetic articulatory latent representation.
Paul Konstantin Krug, Peter Birkholz, Timo Stich
ICASSP1
2025 Learnability of English diphthongs: One dynamic target vs. two static targets
Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Paul Konstantin Krug, Santitham Prom-on, Peter Birkholz, Yi Xu 0007
Speech Commun.4
2023 Self-Supervised Solution to the Control Problem of Articulatory Synthesis
abstract
Given an articulatory-to-acoustic forward model, it is a priori \nunknown how its motor control must be operated to achieve a \ndesired acoustic result. This control problem is a fundamental \nissue of articulatory speech synthesis and the cradle of acousticto-articulatory inversion, a discipline which attempts to address \nthe issue by the means of various methods. This work presents \nan end-to-end solution to the articulatory control problem, in \nwhich synthetic motor trajectories of Monte-Carlo-generated \nartificial speech are linked to input modalities (such as natural speech recordings or phoneme sequence input) via speakerindependent latent representations of a vector-quantized variational autoencoder. The proposed method is self-supervised and \nthus, in principle, synthesizer and speaker model independent.
Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007
INTERSPEECH1
2023 Simulating vocal learning of spoken language: Beyond imitation
abstract
Computational approaches have an important role to play in understanding the complex process of speech acquisition, in general, and have recently been popular in studies of vocal learning in particular. In this article we suggest that two significant problems associated with imitative vocal learning of spoken language, the speaker normalisation and phonological correspondence problems, can be addressed by linguistically grounded auditory perception. In particular, we show how the articulation of consonant–vowel syllables may be learnt from auditory percepts that can represent either individual utterances by speakers with different vocal tract characteristics or ideal phonetic realisations. The result is an optimisation-based implementation of vocal exploration – incorporating semantic, auditory, and articulatory signals – that can serve as a basis for simulating vocal learning beyond imitation.
Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Lorna F. Halliday, Santitham Prom-on, Yi Xu 0007
Speech Commun.4
2023 Artificial Vocal Learning Guided by Phoneme Recognition and Visual Information
abstract
This paper introduces a paradigm shift regarding vocal learning simulations, in which the communicative function of speech acquisition determines the learning process and intelligibility is considered the primary measure of learning success. Thereby, a novel approach for artificial vocal learning is presented that utilizes deep neural network-based phoneme recognition in order to calculate the speech acquisition objective function. This function guides a learning framework that involves the state-of-the-art articulatory speech synthesizer VocalTractLab as the motor-to-acoustic forward model. In this way, an extensive set of German phonemes, including most of the consonants and all stressed vowels, was produced successfully. The synthetic phonemes were rated as highly intelligible by human listeners. Furthermore, it is shown that visual speech information, such as lip and jaw movements, can be extracted from video recordings and be incorporated into the learning framework as an additional loss component during the optimization process. It was observed that this visual loss did not increase the overall intelligibility of phonemes. Instead, the visual loss acted as a regularization mechanism that facilitated the finding of more biologically plausible solutions in the articulatory domain.
Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 Articulatory Synthesis for Data Augmentation in Phoneme Recognition
abstract
While numerous studies on automatic speech recognition have been published in recent years describing data augmentation strategies based on time or frequency domain signal processing, few works exist on the artificial extensions of training data sets using purely synthetic speech data.In this work, the German KIEL corpus was augmented with synthetic data generated with the state-of-the-art articulatory synthesizer VOCALTRACT-LAB.It is shown that the additional synthetic data can lead to a significantly better performance in single-phoneme recognition in certain cases, while at the same time, the performance can also decrease in other cases, depending on the degree of acoustic naturalness of the synthetic phonemes.As a result, this work can potentially guide future studies to improve the quality of articulatory synthesis via the link between synthetic speech production and automatic speech recognition.
Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007
INTERSPEECH1
2022 Exploration strategies for articulatory synthesis of complex syllable onsets
abstract
High-quality articulatory speech synthesis has many potential applications in speech science and technology. However, developing appropriate mappings from linguistic specification to articulatory gestures is difficult and time consuming. In this paper we construct an optimisation-based framework as a first step towards learning these mappings without manual intervention. We demonstrate the production of CCV syllables and discuss the quality of the articulatory gestures with reference to coarticulation.
Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Yi Xu 0007
INTERSPEECH4
2022 Evoc-Learn - High quality simulation of early vocal learning
Yi Xu 0007, Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Peter Birkholz, Paul Konstantin Krug, Santitham Prom-on, Lorna F. Halliday
INTERSPEECH6
2021 Model-Based Exploration of Linking Between Vowel Articulatory Space and Acoustic Space
abstract
While the acoustic vowel space has been extensively studied in previous research, little is known about the high-dimensional articulatory space of vowels. The articulatory imaging techniques are limited to tracking only a few key articulators, leaving the rest of the articulators unmonitored. In the present study, we attempted to develop a detailed articulatory space obtained by training a 3D articulatory synthesizer to learn eleven British English vowels. An analysis-by-synthesis strategy was used to acoustically optimize vocal tract parameters that represent twenty articulatory dimensions. The results show that tongue height and retraction, larynx location and lip roundness are the most perceptually distinctive articulatory dimensions. Yet, even for these dimensions, there is a fair amount of articulatory overlap between vowels, unlike the fine-grained acoustic space. This method opens up the possibility of using modelling to investigate the link between speech production and perception.
Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Paul Konstantin Krug, Santitham Prom-on, Peter Birkholz, Yi Xu 0007
Interspeech4
2020 Finding Intelligible Consonant-Vowel Sounds Using High-Quality Articulatory Synthesis
abstract
In this study, a state-of-the-art articulatory speech synthesiser was used as the basis for simulating the exploration of CV sounds imitating speech stimuli. By adopting a relevant kinematic model and systematically reducing the search space of consonant articulatory targets, intelligible CV sounds can be found. Derivative-free optimisation strategies were evaluated to speed up the process of exploring articulatory space and the possibility of using automatic speech recognition as a means of evaluating intelligibility was explored.
Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Yi Xu 0007
INTERSPEECH4