Branislav Gerazov

dblp:154/1475 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0003-2498-6831ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Privacy-Preserving Child Voice TTS for On-Device AAC: A Recording-to-Deployment Protocol
Silvio Marcello Pagliara, Branislav Gerazov, Vanesa Lazareva, Marija Markovska, Katerina Mavrou, Eleni Theodorou, Dimitar Taskovski, Danche Todorovska, Francesco Zanfardino, Antonio Spera, Ilaria Tatulli, Anna Rybinska, May Agius, Nefi Charalambous-Darden, Katarzyna Luszczak, Antonello Mura
ICCHP (1)2
2025 MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
Mohammed Salah Al-Radhi, Géza Németh, Branislav Gerazov
INTERSPEECH3
2025 Learnability of English diphthongs: One dynamic target vs. two static targets
Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Paul Konstantin Krug, Santitham Prom-on, Peter Birkholz, Yi Xu 0007
Speech Commun.3
2023 Self-Supervised Solution to the Control Problem of Articulatory Synthesis
abstract
Given an articulatory-to-acoustic forward model, it is a priori \nunknown how its motor control must be operated to achieve a \ndesired acoustic result. This control problem is a fundamental \nissue of articulatory speech synthesis and the cradle of acousticto-articulatory inversion, a discipline which attempts to address \nthe issue by the means of various methods. This work presents \nan end-to-end solution to the articulatory control problem, in \nwhich synthetic motor trajectories of Monte-Carlo-generated \nartificial speech are linked to input modalities (such as natural speech recordings or phoneme sequence input) via speakerindependent latent representations of a vector-quantized variational autoencoder. The proposed method is self-supervised and \nthus, in principle, synthesizer and speaker model independent.
Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007
INTERSPEECH3
2023 Simulating vocal learning of spoken language: Beyond imitation
abstract
Computational approaches have an important role to play in understanding the complex process of speech acquisition, in general, and have recently been popular in studies of vocal learning in particular. In this article we suggest that two significant problems associated with imitative vocal learning of spoken language, the speaker normalisation and phonological correspondence problems, can be addressed by linguistically grounded auditory perception. In particular, we show how the articulation of consonant–vowel syllables may be learnt from auditory percepts that can represent either individual utterances by speakers with different vocal tract characteristics or ideal phonetic realisations. The result is an optimisation-based implementation of vocal exploration – incorporating semantic, auditory, and articulatory signals – that can serve as a basis for simulating vocal learning beyond imitation.
Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Lorna F. Halliday, Santitham Prom-on, Yi Xu 0007
Speech Commun.3
2023 Artificial Vocal Learning Guided by Phoneme Recognition and Visual Information
abstract
This paper introduces a paradigm shift regarding vocal learning simulations, in which the communicative function of speech acquisition determines the learning process and intelligibility is considered the primary measure of learning success. Thereby, a novel approach for artificial vocal learning is presented that utilizes deep neural network-based phoneme recognition in order to calculate the speech acquisition objective function. This function guides a learning framework that involves the state-of-the-art articulatory speech synthesizer VocalTractLab as the motor-to-acoustic forward model. In this way, an extensive set of German phonemes, including most of the consonants and all stressed vowels, was produced successfully. The synthetic phonemes were rated as highly intelligible by human listeners. Furthermore, it is shown that visual speech information, such as lip and jaw movements, can be extracted from video recordings and be incorporated into the learning framework as an additional loss component during the optimization process. It was observed that this visual loss did not increase the overall intelligibility of phonemes. Instead, the visual loss acted as a regularization mechanism that facilitated the finding of more biologically plausible solutions in the articulatory domain.
Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Articulatory Synthesis for Data Augmentation in Phoneme Recognition
abstract
While numerous studies on automatic speech recognition have been published in recent years describing data augmentation strategies based on time or frequency domain signal processing, few works exist on the artificial extensions of training data sets using purely synthetic speech data.In this work, the German KIEL corpus was augmented with synthetic data generated with the state-of-the-art articulatory synthesizer VOCALTRACT-LAB.It is shown that the additional synthetic data can lead to a significantly better performance in single-phoneme recognition in certain cases, while at the same time, the performance can also decrease in other cases, depending on the degree of acoustic naturalness of the synthetic phonemes.As a result, this work can potentially guide future studies to improve the quality of articulatory synthesis via the link between synthetic speech production and automatic speech recognition.
Paul Konstantin Krug, Peter Birkholz, Branislav Gerazov, Daniel R. van Niekerk, Anqi Xu 0002, Yi Xu 0007
INTERSPEECH3
2022 Exploration strategies for articulatory synthesis of complex syllable onsets
abstract
High-quality articulatory speech synthesis has many potential applications in speech science and technology. However, developing appropriate mappings from linguistic specification to articulatory gestures is difficult and time consuming. In this paper we construct an optimisation-based framework as a first step towards learning these mappings without manual intervention. We demonstrate the production of CCV syllables and discuss the quality of the articulatory gestures with reference to coarticulation.
Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Yi Xu 0007
INTERSPEECH3
2022 Evoc-Learn - High quality simulation of early vocal learning
Yi Xu 0007, Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Peter Birkholz, Paul Konstantin Krug, Santitham Prom-on, Lorna F. Halliday
INTERSPEECH4
2021 ProsoBeast Prosody Annotation Tool
abstract
The labelling of speech corpora is a laborious and time-consuming process. The ProsoBeast Annotation Tool seeks to ease and accelerate this process by providing an interactive 2D representation of the prosodic landscape of the data, in which contours are distributed based on their similarity. This interactive map allows the user to inspect and label the utterances. The tool integrates several state-of-the-art methods for dimensionality reduction and feature embedding, including variational autoencoders. The user can use these to find a good representation for their data. In addition, as most of these methods are stochastic, each can be used to generate an unlimited number of different prosodic maps. The web app then allows the user to seamlessly switch between these alternative representations in the annotation process. Experiments with a sample prosodically rich dataset have shown that the tool manages to find good representations of varied data and is helpful both for annotation and label correction. The tool is released as free software for use by the community.
Branislav Gerazov, Michael Wagner 0019
Interspeech1
2021 Model-Based Exploration of Linking Between Vowel Articulatory Space and Acoustic Space
abstract
While the acoustic vowel space has been extensively studied in previous research, little is known about the high-dimensional articulatory space of vowels. The articulatory imaging techniques are limited to tracking only a few key articulators, leaving the rest of the articulators unmonitored. In the present study, we attempted to develop a detailed articulatory space obtained by training a 3D articulatory synthesizer to learn eleven British English vowels. An analysis-by-synthesis strategy was used to acoustically optimize vocal tract parameters that represent twenty articulatory dimensions. The results show that tongue height and retraction, larynx location and lip roundness are the most perceptually distinctive articulatory dimensions. Yet, even for these dimensions, there is a fair amount of articulatory overlap between vowels, unlike the fine-grained acoustic space. This method opens up the possibility of using modelling to investigate the link between speech production and perception.
Anqi Xu 0002, Daniel R. van Niekerk, Branislav Gerazov, Paul Konstantin Krug, Santitham Prom-on, Peter Birkholz, Yi Xu 0007
Interspeech3
2020 Finding Intelligible Consonant-Vowel Sounds Using High-Quality Articulatory Synthesis
abstract
In this study, a state-of-the-art articulatory speech synthesiser was used as the basis for simulating the exploration of CV sounds imitating speech stimuli. By adopting a relevant kinematic model and systematically reducing the search space of consonant articulatory targets, intelligible CV sounds can be found. Derivative-free optimisation strategies were evaluated to speed up the process of exploring articulatory space and the possibility of using automatic speech recognition as a means of evaluating intelligibility was explored.
Daniel R. van Niekerk, Anqi Xu 0002, Branislav Gerazov, Paul Konstantin Krug, Peter Birkholz, Yi Xu 0007
INTERSPEECH3
2018 A Weighted Superposition of Functional Contours Model for Modelling Contextual Prominence of Elementary Prosodic Contours
abstract
International audience
Branislav Gerazov, Gérard Bailly, Yi Xu 0007
INTERSPEECH1
2018 Intonation modelling using a muscle model and perceptually weighted matching pursuit
Pierre-Edouard Honnet, Branislav Gerazov, Aleksandar Gjoreski, Philip N. Garner
Speech Commun.2
2018 Prosodic stress detection for fixed stress languages using formal atom decomposition and a statistical hidden Markov hybrid
György Szaszák, Máté Ákos Tündik, Branislav Gerazov
Speech Commun.3
2015 Atom decomposition-based intonation modelling
abstract
Current statistical parametric text-to-speech (TTS) synthesis methods allow production of neutral speech with acceptable quality. However, prosody is often qualified as unsatisfactory and sounding too flat. In this paper, we address intonation modelling for TTS based on physiological aspects of prosody production. A set of gamma distribution shaped atoms is defined and then intonation decomposition is performed using a matching pursuit algorithm. Some preliminary experiments show that this model allows easy extraction of physiologically meaningful atoms that could be used to generate intonation in a TTS system.
Pierre-Edouard Honnet, Branislav Gerazov, Philip N. Garner
ICASSP2
2015 Weighted correlation based atom decomposition intonation modelling
abstract
Intonation modelling is an integral part of text-to-speech systems from their very beginnings. This has led to the proliferation of various intonation models, each with its own relative strengths and weaknesses. Only a few of these intonation models are based on physiology, despite the advantage that such models are language independent. We propose a new intonation model inspired by the physiology of intonation production, which is based on decomposing the F0 contour into elementary atoms. The model, named the Weighted Correlation Atom Decomposition model (WCAD), is a generalisation of the command response (CR) model and has the advantage of having a simple parameter extraction method. The decomposition process follows a matching pursuit approach based on using the perceptually relevant weighted correlation as a cost function. The results have affirmed the plausibility of using the WCAD model to model F0 contours across different languages and speakers. The results have also shown that the WCAD model has good comparative performance to the CR model, giving it practical importance.
Branislav Gerazov, Pierre-Edouard Honnet, Aleksandar Gjoreski, Philip N. Garner
INTERSPEECH1
2015 Kernel Power Flow Orientation Coefficients for Noise-Robust Speech Recognition
abstract
Noise-robustness has become a crucial parameter in Automatic Speech Recognition (ASR) systems today with their increased use in noise-filled real-world environments. One way to address this issue is to develop features that are innately noise-robust. The Kernel Power flow Orientation Coefficients (KPOCs) are a novel feature set based on spectro-temporal analysis that uses a bank of 2D kernels to extract the dominant orientation of the power flow at each point in the auditory spectrogram of the speech signal. The collection of dominant power flow orientation angles forms a novel representation of the speech signal named the Power flow Orientation Spectrogram (POS), which is innately resistant to the spectral masking introduced by the presence of noise and reverberation. This approach not only grants KPOC its noise robustness, but also keeps the number of output coefficients inherently small, thus eliminating the need of the feature dimensionality reduction otherwise necessary in the conventional the spectro-temporal approach. KPOCs performance has been evaluated on three experimental frameworks, and the results have shown that they outperform a number of well-known noise-robust features for average and low SNRs. The relative improvement in Word Recognition Accuracy (WRA) to the classic Mel Frequency Cepstral Coefficients (MFCCs) for the Aurora 2 task goes from 32% up to 190% for SNRs in the range from 10 down to - 5 dB. The experimental results also show that in clean training the performance of KPOC approaches that of the state-of-the-art noise-robust ASR frontends in all noise scenarios for small vocabulary ASR tasks.
Branislav Gerazov, Zoran A. Ivanovski
IEEE ACM Trans. Audio Speech Lang. Process.1