Esther Klabbers

dblp:22/3972 · also Esther Klabbers-Judd · DBLP profile ↗
← Back
28ranked-venue papers
9as first author
5since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 9 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 7 first-author · 5 since 2021
YearPublicationVenuePosition
2023 The Effects of Input Type and Pronunciation Dictionary Usage in Transfer Learning for Low-Resource Text-to-Speech
abstract
We compare phone labels and articulatory features as input for cross-lingual transfer learning in text-to-speech (TTS) for low-resource languages (LRLs). Experiments with FastSpeech 2 and the LRL West Frisian show that using articulatory features outperformed using phone labels in both intelligibility and naturalness. For LRLs without pronunciation dictionaries, we propose two novel approaches: a) using a massively multilingual model to convert grapheme-to-phone (G2P) in both training and synthesizing, and b) using a universal phone recognizer to create a makeshift dictionary. Results show that the G2P approach performs largely on par with using a ground-truth dictionary and the phone recognition approach, while performing generally worse, remains a viable option for LRLs less suitable for the G2P approach. Within each approach, using articulatory features as input outperforms using phone labels.
Phat Do, Matt Coler, Jelske Dijkstra, Esther Klabbers
INTERSPEECH4
2023 Resource-Efficient Fine-Tuning Strategies for Automatic MOS Prediction in Text-to-Speech for Low-Resource Languages
abstract
We train a MOS prediction model based on wav2vec 2.0 using the open-access data sets BVCC and SOMOS. Our test with neural TTS data in the low-resource language (LRL) West Frisian shows that pre-training on BVCC before fine-tuning on SOMOS leads to the best accuracy for both fine-tuned and zero-shot prediction. Further fine-tuning experiments show that using more than 30 percent of the total data does not lead to significant improvements. In addition, fine-tuning with data from a single listener shows promising system-level accuracy, supporting the viability of one-participant pilot tests. These findings can all assist the resource-conscious development of TTS for LRLs by progressing towards better zero-shot MOS prediction and informing the design of listening tests, especially in early-stage evaluation.
Phat Do, Matt Coler, Jelske Dijkstra, Esther Klabbers
INTERSPEECH4
2022 Strategies for developing a Conversational Speech Dataset for Text-To-Speech Synthesis
abstract
There have been many efforts to improve the quality of speech synthesis systems in conversational AI. Although state-of-the-art systems are capable of producing natural-sounding speech, the generated speech often lacks prosodic variation and is not always suited to the task. In this paper, we examine dialogue data collection methods to use as training data for our acoustic models. We collect speech using three different setups: (1) Random read-aloud sentences; (2) Performed dialogues; (3) Semi-Spontaneous dialogues. We analyze prosodic and textual properties of the data collected in these setups and make some recommendations to collect data for speech synthesis in conversational AI settings.
Adaeze Adigwe, Esther Klabbers
INTERSPEECH2
2022 Data-augmented cross-lingual synthesis in a teacher-student framework
abstract
Cross-lingual synthesis can be defined as the task of letting a speaker generate fluent synthetic speech in another language.This is a challenging task, and resulting speech can suffer from reduced naturalness, accented speech, and/or loss of essential voice characteristics.Previous research shows that many models appear to have insufficient generalization capabilities to perform well on every of these cross-lingual aspects.To overcome these generalization problems, we propose to apply the teacher-student paradigm to cross-lingual synthesis.While a teacher model is commonly used to produce teacher forced data, we propose to also use it to produce augmented data of unseen speaker-language pairs, where the aim is to retain essential speaker characteristics.Both sets of data are then used for student model training, which is trained to retain the naturalness and prosodic variation present in the teacher forced data, while learning the speaker identity from the augmented data.Some modifications to the student model are proposed to make the separation of teacher forced and augmented data more straightforward.Results show that the proposed approach improves the retention of speaker characteristics in the speech, while managing to retain high levels of naturalness and prosodic variation.
Marcel de Korte, Jaebok Kim, Aki Kunikoshi, Adaeze Adigwe, Esther Klabbers
INTERSPEECH5
2021 A Systematic Review and Analysis of Multilingual Data Strategies in Text-to-Speech for Low-Resource Languages
abstract
We provide a systematic review of past studies that use multilingual data for text-to-speech (TTS) of low-resource languages (LRLs). We focus on the strategies used by these studies for incorporating multilingual data and how they affect output speech quality. To investigate the difference in output quality between corresponding monolingual and multilingual models, we propose a novel measure to compare this difference across the included studies and their various evaluation metrics. This measure, called the Multilingual Model Effect (MLME), is found to be affected by: acoustic model architecture, the difference ratio of target language data between corresponding multilingual and monolingual experiments, the balance ratio of target language data to total data, and the amount of target language data used. These findings can act as reference for data strategies in future experiments with multilingual TTS models for LRLs. Language family classification, despite being widely used, is not found to be an effective criterion for selecting source languages.
Phat Do, Matt Coler, Jelske Dijkstra, Esther Klabbers
Interspeech4
2020 Efficient Neural Speech Synthesis for Low-Resource Languages Through Multilingual Modeling
abstract
Recent advances in neural TTS have led to models that can produce high-quality synthetic speech. However, these models typically require large amounts of training data, which can make it costly to produce a new voice with the desired quality. Although multi-speaker modeling can reduce the data requirements necessary for a new voice, this approach is usually not viable for many low-resource languages for which abundant multi-speaker data is not available. In this paper, we therefore investigated to what extent multilingual multi-speaker modeling can be an alternative to monolingual multi-speaker modeling, and explored how data from foreign languages may best be combined with low-resource language data. We found that multilingual modeling can increase the naturalness of low-resource language speech, showed that multilingual models can produce speech with a naturalness comparable to monolingual multi-speaker models, and saw that the target language naturalness was affected by the strategy used to add foreign language data.
Marcel de Korte, Jaebok Kim, Esther Klabbers
INTERSPEECH3
2020 BLISS: An Agent for Collecting Spoken Dialogue Data about Health and Well-being
abstract
An important objective in health-technology is the ability to gather information about people’s well-being. Structured interviews can be used to obtain this information, but are time-consuming and not scalable. Questionnaires provide an alternative way to extract such information, though typically lack depth. In this paper, we present our first prototype of the BLISS agent, an artificial intelligent agent which intends to automatically discover what makes people happy and healthy. The goal of Behaviour-based Language-Interactive Speaking Systems (BLISS) is to understand the motivations behind people’s happiness by conducting a personalized spoken dialogue based on a happiness model. We built our first prototype of the model to collect 55 spoken dialogues, in which the BLISS agent asked questions to users about their happiness and well-being. Apart from a description of the BLISS architecture, we also provide details about our dataset, which contains over 120 activities and 100 motivations and is made available for usage.
Jelte van Waterschoot, Iris Hendrickx, Esther Klabbers, Marcel de Korte, Helmer Strik, Catia Cucchiarini, Mariët Theune
LREC4
2014 A novel pitch decomposition method for the generalized linear alignment model
abstract
Superpositional models of intonation typically propose decomposing fundamental frequency (F0) contours into phrase curves and accent curves, aligned with phrases and left-headed feet, respectively. Extracting these component curves from F0contours without making undue assumptions is challenging. We propose a novel method for decomposing pitch curves, based on the assumption that accent curves can be described by combining skewed normal distributions and sigmoid functions. In contrast to an earlier pitch decomposition algorithm (“PRISM”), this allows for simple joint optimization of phrase and accent curve parameters, using fewer parameters. The proposed method was evaluated on three speech corpora containing: (1) synthetically generated pitch curves, (2) all-sonorant utterances, and (3) utterances containing both sonorant and non-sonorant speech sounds. The root weighted mean squared error is small, and, on the corpus for which comparable data are available, is significantly smaller than for PRISM.
Mahsa Sadat Elyasi Langarani, Esther Klabbers, Jan P. H. van Santen
ICASSP2
2012 Synthetic F0 Can Effectively Convey Speaker ID in Delexicalized Speech
abstract
We investigate the extent to which F0 can convey speaker ID in the absence of spectral, segmental, and durational information. We propose two methods of F0 synthesis based on the Linear Alignment Model (LAM) [2]: one parametric, the other corpusbased. Through a perceptual experiment, we show that F0 alone is able to convey information about speaker ID. We find that F0 synthesized with either LAM-based method conveys speaker ID almost as effectively as natural F0. Index Terms: F0, prosody, speech synthesis, speaker identity, recombinant synthesis
Eric Morley, Esther Klabbers, Jan P. H. van Santen, Alexander Kain, Seyed Hamidreza Mohammadi
INTERSPEECH2
2011 F0 range and peak alignment across speakers and emotions
abstract
We present an analysis of F0range and peak alignment in emotional speech from a heterogeneous group of speakers varying in age and gender. Both speaker and emotion had a strong effect on F0range. Despite these large changes in the F0trajectory, peak alignment was remarkably stable. Using the Linear Alignment Model (LAM), we show that the effects on alignment of emotion and speaker differences, al though statistically significant, are small. This stability results in a conclusion that peak alignment, unlike F0range, does not appear to carry much information about speaker identity or emotional state. The LAM is effective in that it explains 42% of the variance in peak location on average, and furthermore it predicts the time of F0peaks with an average RMS error of 12ms.
Eric Morley, Jan P. H. van Santen, Esther Klabbers, Alexander Kain
ICASSP3
2010 Evaluation of speaker mimic technology for personalizing SGD voices
abstract
In this paper, we demonstrate the use of state-of-the-art speech technology to transform speech from a source speaker to mimic a particular target speaker with the intention of providng personalized voices to users of Speech Generating Devices (SGDs). This speaker mimicry (SM) capability allows us to use highquality acoustic inventories from professional speakers and transform them to a different target speaker using a very limited set of sentences from that speaker. This technology targets future SGD users who still have a limited vocabulary or available previous recordings. The results of a perceptual study show that listeners can identify which SM voices most resemble their respective target voices. 1
Esther Klabbers, Alexander Kain, Jan P. H. van Santen
INTERSPEECH1
2007 Application of speech technology in a home based assessment kiosk for early detection of alzheimer's disease
abstract
Alzheimer’s disease, a degenerative disease that affects an es-timated 4.5 million people in the U.S., can be treated far more effectively when it is detected early. There are numerous chal-lenges to early detection. One is objectivity, since caretakers are often emotionally invested in the health of the patients, who may be their family members. Consistency of administration can also be an issue, especially where longitudinal results from different examiners are compared. Finally, the frequency of testing can be adversely affected by scheduling or cost con-straints for in-home psychometrician visits. The kiosk system described in this paper, currently deployed in homes around the country, uses speech technology to provide advantages that ad-dress these challenges. Index Terms: spoken language systems, in-home psychometric testing, Alzheimer’s disease.
Rachel Coulston, Esther Klabbers, Jacques de Villiers, John-Paul Hosom
INTERSPEECH2
2007 The Contribution of Various Sources of Spectral Mismatch to Audible Discontinuities in a Diphone Database
abstract
One of the major problems in concatenative synthesis is the occurrence of audible discontinuities between two successive concatenative units. Several studies have attempted to discover objective distance measures that predict the audibility of these discontinuities. In this paper, we investigate mid-vowel joins for three vowels with a range of post-vocalic consonant contexts typical for diphone databases. A first perceptual experiment uses a pairwise comparison procedure to find two subsets of unit combinations: Those with versus without audible discontinuities. A second perceptual experiment uses these two subsets in a procedure where formant resynthesis is used to manipulate three sources of discontinuity separately: formant frequencies, formant bandwidths, and overall energy. Results show mismatch in formant frequencies provides the largest contribution to audible discontinuity, followed by mismatch in overall energy
Esther Klabbers, Jan P. H. van Santen, Alexander Kain
IEEE Trans. Speech Audio Process.1
2005 Discontinuity detection in concatenated speech synthesis based on nonlinear speech analysis
abstract
An objective distance measure which is able to predict audible discontinuity in concatenated speech synthesis systems is very important. Previous works were primarily based on features estimated by linear and/or stationary models of speech. In this paper, we introduce two nonlinear approaches for the detection of discontinuity. The first method is based on a nonlinear harmonic model of speech while the second method is based on the demodulation of speech in an amplitude and a frequency component using the Teager energy operator. Fisher’s linear discriminant was used for the separation of signals with audible discontinuity from those perceived as continuous. When we combined the two methods using Fisher’s linear discriminant a detection rate of 56.5 % was achieved which is an 90 % improvement over previously published results on the same database. 1.
Yannis Pantazis, Yannis Stylianou, Esther Klabbers
INTERSPEECH3
2005 Synthesis of prosody using multi-level unit sequences
Jan P. H. van Santen, Alexander Kain, Esther Klabbers, Taniya Mishra
Speech Commun.3
2003 Control and prediction of the impact of pitch modification on synthetic speech quality
abstract
In order to use speech synthesis to generate highly expressive speech convincingly, the problem of poor prosody (both prediction and generation) needs to be overcome. In this paper we will show that with a simple annotation scheme using the notion of foot structure, we can more accurately predict the shape of local pitch contours. The assumption is that with a better selection mechanism we can reduce the amount of pitch modification required, thereby reducing speech degradation. In addition, we present a perceptual experiment that investigates the degradation introduced by pitch modification using the OGIresLPC algorithm. We correlated the weighted perceptual score with different pitch and delta pitch distances. The best combination of distance measures is able to explain 63% of the variance in the perceptual scores. Decreasing the pitch is shown to have a higher impact on perception than increasing the pitch.
Esther Klabbers, Jan P. H. van Santen
INTERSPEECH1
2003 Detection of list-type sentences
abstract
In this paper, we explore a text type based scheme of text analysis, through the specific problem of detecting the list text type. This is important because TTS systems that can generate the very distinct F0 contour of lists sound more natural. The presented list detection algorithm uses part-of-speech tags as input, and detects lists by computing the alignment costs of clauses in a sentence. The algorithm detects lists with 80 % accuracy. 1.
Taniya Mishra, Esther Klabbers, Jan P. H. van Santen
INTERSPEECH2
2003 Applications of computer generated expressive speech for communication disorders
abstract
This paper focuses on generation of expressive speech, specifically speech displaying vocal affect. Generating speech with vocal affect is important for diagnosis, research, and remediation for children with autism and developmental language disorders. However, because vocal affect involves many acoustic factors working together in complex ways, it is unlikely that we will be able to generate compelling vocal affect with traditional diphone synthesis. Instead, methods are needed that preserve as much of the original signals as possible. We describe an approach to concatenative synthesis that attempts to combine the naturalness of unit selection based synthesis with the ability of diphone based synthesis to handle unrestricted input domains. 1.
Jan P. H. van Santen, Lois M. Black, Gilead Cohen, Alexander Kain, Esther Klabbers, Taniya Mishra, Jacques de Villiers, Xiaochuan Niu
INTERSPEECH5
2003 On the computation of the Kullback-Leibler measure for spectral distances
abstract
Efficient algorithms for the exact and approximate computation of the symmetrical Kullback-Leibler (1998) measure for spectral distances are presented for linear predictive coding (LPC) spectra. A interpretation of this measure is given in terms of the poles of the spectra. The performances of the algorithms in terms of accuracy and computational complexity are assessed for the application of computing concatenation costs in unit-selection-based speech synthesis. With the same complexity and storage requirements, the exact method is superior in terms of accuracy.
Raymond N. J. Veldhuis, Esther Klabbers
IEEE Trans. Speech Audio Process.2
2001 Speech synthesis development made easy: the bonn open synthesis system
abstract
This paper describes a new open source architecture for unit-selection based speech synthesis called BOSS (Bonn Open Synthesis System). It is built up modularly, with communications between modules taking place in a fixed format. This makes the addition, deletion and substitution of modules very easy. The strict separation between data and algorithms allows for the simple creation of new speech corpora for different domains and languages. 1.
Esther Klabbers, Karlheinz Stöber, Raymond N. J. Veldhuis, Petra Wagner, Stefan Breuer
INTERSPEECH1
2001 From data to speech: a general approach
Mariët Theune, Esther Klabbers, Jan-Roelof de Pijper, Emiel Krahmer, Jan Odijk
Nat. Lang. Eng.2
2001 Reducing audible spectral discontinuities
abstract
A common problem in diphone synthesis is discussed, viz., the occurrence of audible discontinuities at diphone boundaries. Informal observations show that spectral mismatch is the most likely the clause of this phenomenon. We first set out to find an objective spectral measure for discontinuity. To this end, several spectral distance measures are related to the results of a listening experiment. Then, we studied the feasibility of extending the diphone database with context-sensitive diphones to reduce the occurrence of audible discontinuities. The number of additional diphones is limited by clustering consonant contexts that have a similar effect on the surrounding vowels on the basis of the best performing distance measure. A listening experiment has shown that the addition of these context-sensitive diphones significantly reduces the amount of audible discontinuities.
Esther Klabbers, Raymond N. J. Veldhuis
IEEE Trans. Speech Audio Process.1
2000 Predicting segmental durations for Dutch using the sums-of-products approach
abstract
This paper presents the results of a duration study performed for Dutch using the sums-of-products approach [5]. With a relatively small corpus of 297 sentences, a duration model could be constructed with an RMSE of 27 ms, which compares well to similar models for English, French and German. In an evaluation study the predicted durations of the duration model were compared to those predicted by a rule-based duration model. 1.
Esther Klabbers, Jan P. H. van Santen
INTERSPEECH1
2000 A solution to the reduction of concatenation artefacts in speech synthesis
abstract
One problem with speech synthesis impeding high quality is the occurrence of audible discontinuities at segment boundaries. Formant jumps across concatenation points suggest the problem to be due to spectral differences. The problem is most apparent in vowels and semi-vowels. We propose to reduce the number of audible discontinuities by adding context-sensitive diphones to the database. The number of additional diphones is limited by clustering contexts with similar spectral effects on the neighbouring vowels, using the Kullback-Leibler distance. A listening experiment has shown that the percentage of perceived discontinuities has significantly decreased. 1.
Esther Klabbers, Raymond N. J. Veldhuis, Kim Koppen
INTERSPEECH1
1998 System Demonstration Goalgetter: Generation Of Spoken Soccer Reports
Mariët Theune, Esther Klabbers
INLG2
1998 A generic algorithm for generating spoken monologues
abstract
The defining property of a Concept-to-Speech system is that it combines language and speech generation. Language generation converts the input concepts into natural language, which speech generation subsequently transforms into speech. Potentially, this leads to a more `natural sounding' output than can be achieved in a plain Text-to-Speech system, since the correct placement of pitch accents and intonational boundaries ---an important factor contributing to the `naturalness' of the generated speech--- is co-determined by syntactic and discourse information, which is typically available in the language generation module. In this paper, a generic algorithm for the generation of coherent spoken monologues is discussed, called D2S. Language generation is done by a module called LGM which is based on TAG-like syntactic structures with open slots, combined with conditions which determine when the syntactic structure can be used properly. A speech generation module (SGM) converts the output of the LGM into speech using either phrase-concatenation or diphone-synthesis.
Esther Klabbers, Emiel Krahmer, Mariët Theune
ICSLP1
1998 On the reduction of concatenation artefacts in diphone synthesis
abstract
One well-known problem with diphone concatenation is the occurrence of audible discontinuities at diphone boundaries, which are most prominent in vowels and semi-vowels. Significant formant jumps at certain boundaries suggest that the problem is of a spectral nature. We have examined this hypothesis by correlating the results of a listening experiment with spectral distances measured across diphone boundaries. The aim is to find a spectral distance measure that best predicts when discontinuities are audible in order to find out how the diphone database can best be extended with context-sensitive diphones. The results show that the KullbackLeibler measure is the best predictor. 1. INTRODUCTION Most speech synthesis systems available today are based on diphone concatenation. One well-known problem with diphone concatenation is the occurrence of audible discontinuities at diphone boundaries, which are most prominent in vowels and semi-vowels and are caused by contextual influences. Our ...
Esther Klabbers, Raymond N. J. Veldhuis
ICSLP1
1995 The Dutch polyphone corpus
Els den Os, T. I. Boogaart, Lou Boves, Esther Klabbers
EUROSPEECH4