Benjamin van Niekerk

dblp:169/9023 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0001-9207-6309ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features
Kyle Janse van Rensburg, Benjamin van Niekerk, Herman Kamper
IEEE Signal Process. Lett.2
2025 Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
abstract
We look at the long-standing problem of segmenting unlabeled speech into word-like segments and clustering these into a lexicon. Several previous methods use a scoring model coupled with dynamic programming to find an optimal segmentation. Here we propose a much simpler strategy: we predict word boundaries using the dissimilarity between adjacent self-supervised features, then we cluster the predicted segments to construct a lexicon. For a fair comparison, we update the older ES-KMeans dynamic programming method with better features and boundary constraints. On the five-language ZeroSpeech benchmarks, our simple approach gives similar state-of-the-art results compared to the new ES-KMeans+ method, while being almost five times faster. Project webpage: https://s-malan.github.io/prom-seg-clus.
Simon Malan, Benjamin van Niekerk, Herman Kamper
ICASSP2
2025 LinearVC: Linear Transformations of Self-Supervised Features Through the Lens of Voice Conversion
Herman Kamper, Benjamin van Niekerk, Julian Zaidi, Marc-André Carbonneau
INTERSPEECH2
2024 Spoken-Term Discovery using Discrete Speech Units
Benjamin van Niekerk, Julian Zaidi, Marc-André Carbonneau, Herman Kamper
INTERSPEECH1
2023 Voice Conversion With Just Nearest Neighbors
Matthew Baas, Benjamin van Niekerk, Herman Kamper
INTERSPEECH2
2023 Visually grounded few-shot word acquisition with fewer shots
abstract
We propose a visually grounded speech model that acquires new words and their visual depictions from just a few wordimage example pairs.Given a set of test images and a spoken query, we ask the model which image depicts the query word.Previous work has simplified this problem by either using an artificial setting with digit word-image pairs or by using a large number of examples per class.We propose an approach that can work on natural word-image pairs but with less examples, i.e. fewer shots.Our approach involves using the given word-image example pairs to mine new unsupervised word-image training pairs from large collections of unlabelled speech and images.Additionally, we use a word-to-image attention mechanism to determine word-image similarity.With this new model, we achieve better performance with fewer shots than any existing approach.
Leanne Nortje, Benjamin van Niekerk, Herman Kamper
INTERSPEECH2
2023 Rhythm Modeling for Voice Conversion
abstract
Voice conversion aims to transform source speech into a different target voice. However, typical voice conversion systems do not account for rhythm, which is an important factor in the perception of speaker identity. To bridge this gap, we introduce Urhythmic—an unsupervised method for rhythm conversion that does not require parallel data or text transcriptions. Using self-supervised representations, we first divide source audio into segments approximating sonorants, obstruents, and silences. Then we model rhythm by estimating speaking rate or the duration distribution of each segment type. Finally, we match the target speaking rate or rhythm by time-stretching the speech segments. Experiments show that Urhythmic outperforms existing unsupervised methods in terms of quality and prosody.
Benjamin van Niekerk, Marc-André Carbonneau, Herman Kamper
IEEE Signal Process. Lett.1
2022 A Comparison of Discrete and Soft Speech Units for Improved Voice Conversion
abstract
The goal of voice conversion is to transform source speech into a target voice, keeping the content unchanged. In this paper, we focus on self-supervised representation learning for voice conversion. Specifically, we compare discrete and soft speech units as input features. We find that discrete representations effectively remove speaker information but discard some linguistic content – leading to mispronunciations. As a solution, we propose soft speech units learned by predicting a distribution over the discrete units. By modeling uncertainty, soft units capture more content information, improving the intelligibility and naturalness of converted speech.12
Benjamin van Niekerk, Marc-André Carbonneau, Julian Zaidi, Matthew Baas, Hugo Seuté, Herman Kamper
ICASSP1
2022 Daft-Exprt: Cross-Speaker Prosody Transfer on Any Text for Expressive Speech Synthesis
abstract
This paper presents Daft-Exprt, a multi-speaker acoustic model advancing the state-of-the-art for cross-speaker prosody transfer on any text.This is one of the most challenging, and rarely directly addressed, task in speech synthesis, especially for highly expressive data.Daft-Exprt uses FiLM conditioning layers to strategically inject different prosodic information in all parts of the architecture.The model explicitly encodes traditional low-level prosody features such as pitch, loudness and duration, but also higher level prosodic information that helps generating convincing voices in highly expressive styles.Speaker identity and prosodic information are disentangled through an adversarial training strategy that enables accurate prosody transfer across speakers.Experimental results show that Daft-Exprt significantly outperforms strong baselines on inter-text crossspeaker prosody transfer tasks, while yielding naturalness comparable to state-of-the-art expressive models.Moreover, results indicate that the model discards speaker identity information from the prosody representation, and consistently generate speech with the desired voice.We publicly release our code 1 and provide speech samples from our experiments 2 .
Julian Zaidi, Hugo Seuté, Benjamin van Niekerk, Marc-André Carbonneau
INTERSPEECH3
2021 Towards Unsupervised Phone and Word Segmentation Using Self-Supervised Vector-Quantized Neural Networks
abstract
We investigate segmenting and clustering speech into low-bitrate phone-like sequences without supervision. We specifically constrain pretrained self-supervised vector-quantized (VQ) neural networks so that blocks of contiguous feature vectors are assigned to the same code, thereby giving a variable-rate segmentation of the speech into discrete units. Two segmentation methods are considered. In the first, features are greedily merged until a prespecified number of segments are reached. The second uses dynamic programming to optimize a squared error with a penalty term to encourage fewer but longer segments. We show that these VQ segmentation methods can be used without alteration across a wide range of tasks: unsupervised phone segmentation, ABX phone discrimination, same-different word discrimination, and as inputs to a symbolic word segmentation algorithm. The penalized dynamic programming method generally performs best. While performance on individual tasks is only comparable to the state-of-the-art in some cases, in all tasks a reasonable competing approach is outperformed at a substantially lower bitrate.
Herman Kamper, Benjamin van Niekerk
Interspeech2
2021 Analyzing Speaker Information in Self-Supervised Models to Improve Zero-Resource Speech Processing
abstract
Contrastive predictive coding (CPC) aims to learn representations of speech by distinguishing future observations from a set of negative examples. Previous work has shown that linear classifiers trained on CPC features can accurately predict speaker and phone labels. However, it is unclear how the features actually capture speaker and phonetic information, and whether it is possible to normalize out the irrelevant details (depending on the downstream task). In this paper, we first show that the per-utterance mean of CPC features captures speaker information to a large extent. Concretely, we find that comparing means performs well on a speaker verification task. Next, probing experiments show that standardizing the features effectively removes speaker information. Based on this observation, we propose a speaker normalization step to improve acoustic unit discovery using K-means clustering of CPC features. Finally, we show that a language model trained on the resulting units achieves some of the best results in the ZeroSpeech2021~Challenge.
Benjamin van Niekerk, Leanne Nortje, Matthew Baas, Herman Kamper
Interspeech1
2020 Vector-Quantized Neural Networks for Acoustic Unit Discovery in the ZeroSpeech 2020 Challenge
abstract
In this paper, we explore vector quantization for acoustic unit discovery. Leveraging unlabelled data, we aim to learn discrete representations of speech that separate phonetic content from speaker-specific details. We propose two neural models to tackle this challenge - both use vector quantization to map continuous features to a finite set of codes. The first model is a type of vector-quantized variational autoencoder (VQ-VAE). The VQ-VAE encodes speech into a sequence of discrete units before reconstructing the audio waveform. Our second model combines vector quantization with contrastive predictive coding (VQ-CPC). The idea is to learn a representation of speech by predicting future acoustic units. We evaluate the models on English and Indonesian data for the ZeroSpeech 2020 challenge. In ABX phone discrimination tests, both models outperform all submissions to the 2019 and 2020 challenges, with a relative improvement of more than 30%. The models also perform competitively on a downstream voice conversion task. Of the two, VQ-CPC performs slightly better in general and is simpler and faster to train. Finally, probing experiments show that vector quantization is an effective bottleneck, forcing the models to discard speaker information.
Benjamin van Niekerk, Leanne Nortje, Herman Kamper
INTERSPEECH1
2020 If dropout limits trainable depth, does critical initialisation still matter? A large-scale statistical analysis on ReLU networks
Arnu Pretorius, Elan Van Biljon, Benjamin van Niekerk, Ryan Eloff, Matthew Reynard, Steven James 0001, Benjamin Rosman, Herman Kamper, Steve Kroon
Pattern Recognit. Lett.3
2019 Composing Value Functions in Reinforcement Learning
abstract
An important property for lifelong-learning agents is the ability to combine existing skills to solve new unseen tasks. In general, however, it is unclear how to compose existing skills in a principled manner. Under the assumption of deterministic dynamics, we prove that optimal value function composition can be achieved in entropy-regularised reinforcement learning (RL), and extend this result to the standard RL setting. Composition is demonstrated in a high-dimensional video game, where an agent with an existing library of skills is immediately able to solve new tasks without the need for further learning.
Benjamin van Niekerk, Steven James 0001, Adam Christopher Earle, Benjamin Rosman
ICML1
2019 Unsupervised Acoustic Unit Discovery for Speech Synthesis Using Discrete Latent-Variable Neural Networks
abstract
For our submission to the ZeroSpeech 2019 challenge, we apply discrete latent-variable neural networks to unlabelled speech and use the discovered units for speech synthesis. Unsupervised discrete subword modelling could be useful for studies of phonetic category learning in infants or in low-resource speech technology requiring symbolic input. We use an autoencoder (AE) architecture with intermediate discretisation. We decouple acoustic unit discovery from speaker modelling by conditioning the AE's decoder on the training speaker identity. At test time, unit discovery is performed on speech from an unseen speaker, followed by unit decoding conditioned on a known target speaker to obtain reconstructed filterbanks. This output is fed to a neural vocoder to synthesise speech in the target speaker's voice. For discretisation, categorical variational autoencoders (CatVAEs), vector-quantised VAEs (VQ-VAEs) and straight-through estimation are compared at different compression levels on two languages. Our final model uses convolutional encoding, VQ-VAE discretisation, deconvolutional decoding and an FFTNet vocoder. We show that decoupled speaker conditioning intrinsically improves discrete acoustic representations, yielding competitive synthesis quality compared to the challenge baseline.
Ryan Eloff, André Nortje, Benjamin van Niekerk, Avashna Govender, Leanne Nortje, Arnu Pretorius, Elan Van Biljon, Ewald van der Westhuizen, Lisa van Staden, Herman Kamper
INTERSPEECH3
2017 Online Constrained Model-based Reinforcement Learning
Benjamin van Niekerk, Andreas Damianou, Benjamin Rosman
UAI1