Daniel R. van Niekerk

dblp:15/9231 · also Daniel Rudolph van Niekerk · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
8since 2021 · last 2025
0000-0002-7324-2751ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 first-author · 7 since 2021Artificial intelligence and machine learning · 13 · 7 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Learnability of English diphthongs: One dynamic target vs. two static targets
Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Paul Konstantin Krug, Santitham Prom-on, Peter Birkholz, Yi Xu 0007
Speech Commun.2
2023 Self-Supervised Solution to the Control Problem of Articulatory Synthesis
abstract
Given an articulatory-to-acoustic forward model, it is a priori \nunknown how its motor control must be operated to achieve a \ndesired acoustic result. This control problem is a fundamental \nissue of articulatory speech synthesis and the cradle of acousticto-articulatory inversion, a discipline which attempts to address \nthe issue by the means of various methods. This work presents \nan end-to-end solution to the articulatory control problem, in \nwhich synthetic motor trajectories of Monte-Carlo-generated \nartificial speech are linked to input modalities (such as natural speech recordings or phoneme sequence input) via speakerindependent latent representations of a vector-quantized variational autoencoder. The proposed method is self-supervised and \nthus, in principle, synthesizer and speaker model independent.
Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007
INTERSPEECH4
2023 Simulating vocal learning of spoken language: Beyond imitation
abstract
Computational approaches have an important role to play in understanding the complex process of speech acquisition, in general, and have recently been popular in studies of vocal learning in particular. In this article we suggest that two significant problems associated with imitative vocal learning of spoken language, the speaker normalisation and phonological correspondence problems, can be addressed by linguistically grounded auditory perception. In particular, we show how the articulation of consonant–vowel syllables may be learnt from auditory percepts that can represent either individual utterances by speakers with different vocal tract characteristics or ideal phonetic realisations. The result is an optimisation-based implementation of vocal exploration – incorporating semantic, auditory, and articulatory signals – that can serve as a basis for simulating vocal learning beyond imitation.
Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Lorna F. Halliday, Santitham Prom-on, Yi Xu 0007
Speech Commun.1
2023 Artificial Vocal Learning Guided by Phoneme Recognition and Visual Information
abstract
This paper introduces a paradigm shift regarding vocal learning simulations, in which the communicative function of speech acquisition determines the learning process and intelligibility is considered the primary measure of learning success. Thereby, a novel approach for artificial vocal learning is presented that utilizes deep neural network-based phoneme recognition in order to calculate the speech acquisition objective function. This function guides a learning framework that involves the state-of-the-art articulatory speech synthesizer VocalTractLab as the motor-to-acoustic forward model. In this way, an extensive set of German phonemes, including most of the consonants and all stressed vowels, was produced successfully. The synthetic phonemes were rated as highly intelligible by human listeners. Furthermore, it is shown that visual speech information, such as lip and jaw movements, can be extracted from video recordings and be incorporated into the learning framework as an additional loss component during the optimization process. It was observed that this visual loss did not increase the overall intelligibility of phonemes. Instead, the visual loss acted as a regularization mechanism that facilitated the finding of more biologically plausible solutions in the articulatory domain.
Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Articulatory Synthesis for Data Augmentation in Phoneme Recognition
abstract
While numerous studies on automatic speech recognition have been published in recent years describing data augmentation strategies based on time or frequency domain signal processing, few works exist on the artificial extensions of training data sets using purely synthetic speech data.In this work, the German KIEL corpus was augmented with synthetic data generated with the state-of-the-art articulatory synthesizer VOCALTRACT-LAB.It is shown that the additional synthetic data can lead to a significantly better performance in single-phoneme recognition in certain cases, while at the same time, the performance can also decrease in other cases, depending on the degree of acoustic naturalness of the synthetic phonemes.As a result, this work can potentially guide future studies to improve the quality of articulatory synthesis via the link between synthetic speech production and automatic speech recognition.
Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007
INTERSPEECH4
2022 Exploration strategies for articulatory synthesis of complex syllable onsets
abstract
High-quality articulatory speech synthesis has many potential applications in speech science and technology. However, developing appropriate mappings from linguistic specification to articulatory gestures is difficult and time consuming. In this paper we construct an optimisation-based framework as a first step towards learning these mappings without manual intervention. We demonstrate the production of CCV syllables and discuss the quality of the articulatory gestures with reference to coarticulation.
Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Yi Xu 0007
INTERSPEECH1
2022 Evoc-Learn - High quality simulation of early vocal learning
Yi Xu 0007, Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Peter Birkholz, Paul Konstantin Krug, Santitham Prom-on, Lorna F. Halliday
INTERSPEECH3
2021 Model-Based Exploration of Linking Between Vowel Articulatory Space and Acoustic Space
abstract
While the acoustic vowel space has been extensively studied in previous research, little is known about the high-dimensional articulatory space of vowels. The articulatory imaging techniques are limited to tracking only a few key articulators, leaving the rest of the articulators unmonitored. In the present study, we attempted to develop a detailed articulatory space obtained by training a 3D articulatory synthesizer to learn eleven British English vowels. An analysis-by-synthesis strategy was used to acoustically optimize vocal tract parameters that represent twenty articulatory dimensions. The results show that tongue height and retraction, larynx location and lip roundness are the most perceptually distinctive articulatory dimensions. Yet, even for these dimensions, there is a fair amount of articulatory overlap between vowels, unlike the fine-grained acoustic space. This method opens up the possibility of using modelling to investigate the link between speech production and perception.
Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Paul Konstantin Krug, Santitham Prom-on, Peter Birkholz, Yi Xu 0007
Interspeech2
2020 Finding Intelligible Consonant-Vowel Sounds Using High-Quality Articulatory Synthesis
abstract
In this study, a state-of-the-art articulatory speech synthesiser was used as the basis for simulating the exploration of CV sounds imitating speech stimuli. By adopting a relevant kinematic model and systematically reducing the search space of consonant articulatory targets, intelligible CV sounds can be found. Derivative-free optimisation strategies were evaluated to speed up the process of exploring articulatory space and the possibility of using automatic speech recognition as a means of evaluating intelligibility was explored.
Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Yi Xu 0007
INTERSPEECH1
2017 Rapid Development of TTS Corpora for Four South African Languages
abstract
This paper describes the development of text-to-speech corpora \nfor four South African languages. The approach followed investigated \nthe possibility of using low-cost methods including informal \nrecording environments and untrained volunteer speakers. \nThis objective and the additional future goal of expanding \nthe corpus to increase coverage of South Africa’s 11 official \nlanguages necessitated experimenting with multi-speaker and \ncode-switched data. The process and relevant observations are \ndetailed throughout. The latest version of the corpora are available \nfor download under an open-source license and will likely \nsee further development and refinement in future. \nIndex Terms: text-to-speech corpus, under-resourced languages
Daniel R. van Niekerk, Charl Johannes van Heerden, Marelie H. Davel, Neil Kleynhans, Oddur Kjartansson, Martin Jansche, Linne Ha
INTERSPEECH1
2016 AfriBooms: An Online Treebank for Afrikaans
Liesbeth Augustinus, Peter Dirix, Daniel R. van Niekerk, Ineke Schuurman, Vincent Vandeghinste, Frank Van Eynde, Gerhard B. Van Huyssteen
LREC3
2014 A target approximation intonation model for yorùbá TTS
abstract
A complete intonation model based on quantitative target approximation \nis described for Yor`ub´a text-to-speech (TTS) synthesis. \nThis model is evaluated analytically and perceptually \nand compared to a fundamental frequency (F0) model using \nthe standard HTS implementation. Analytical results suggest \nthat the proposed approach more efficiently models F0 contours \ngiven typical data constraints in under-resourced environments \nand perceptual results comparing the proposed model with HTS \nare encouraging.
Daniel R. van Niekerk, Etienne Barnard
INTERSPEECH1
2014 Predicting utterance pitch targets in Yorùbá for tone realisation in speech synthesis
Daniel R. van Niekerk, Etienne Barnard
Speech Commun.1
2013 Generating fundamental frequency contours for speech synthesis in yorùbá
abstract
We present methods for modelling and synthesising fundamental frequency (F0) contours suitable for application in textto-speech (TTS) synthesis of Yoruba (an African tone language). These methods are discussed and compared with a baseline approach using the HMM-based speech synthesis system HTS. Evaluation is done by comparing ten-fold cross validation squared errors on a small corpus of four speakers. We show that the proposed methods are relatively effective at modelling and generating F0 contours in this context, achieving lower error rates than the baseline. These results suggest that our methods will be useful for the generation of improved synthesis of tone in African languages, which has been a challenge to date.
Daniel R. van Niekerk, Etienne Barnard
INTERSPEECH1
2010 An intonation model for TTS in sepedi
abstract
We present an initial investigation into the acoustic realisation of tone in continuous utterances in Sepedi (a language in the Southern Bantu family). An analytic model for the generation of appropriate pitch contours given an utterance with linguistic tone specification is presented and evaluated. By comparing the model output to speech data from a small tone-marked corpus we conclude that the initial implementation presented here is capable of generating pitch contours exhibiting some realistic properties and identify a number of aspects that require further attention. Lastly, we present some initial perceptual results when integrating the proposed model into a Hidden Markov Model-based speech synthesis system. Index Terms: speech synthesis, tone languages, Sepedi
Daniel R. van Niekerk, Etienne Barnard
INTERSPEECH1
2009 Phonetic alignment for speech synthesis in under-resourced languages
abstract
The rapid development of concatenative speech synthesis systems in resource scarce languages requires an efficient and accurate solution with regard to automated phonetic alignment. However, in this context corpora are often minimally designed due to a lack of resources and expertise necessary for large scale development. Under these circumstances many techniques toward accurate segmentation are not feasible and it is unclear which approaches should be followed. In this paper we investigate this problem by evaluating alignment approaches and demonstrating how these approaches can be applied to limit manual interaction while achieving acceptable alignment accuracy with minimal ideal resources. Index Terms: speech synthesis, phonetic speech segmentation, resource scarce language
Daniel R. van Niekerk, Etienne Barnard
INTERSPEECH1