R. J. J. H. van Son

dblp:34/4715 · also Rob J. J. H. van Son, Rob van Son · DBLP profile ↗
← Back
49ranked-venue papers
24as first author
8since 2021 · last 2025
0000-0001-6321-7635ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 23 first-author · 7 since 2021Artificial intelligence and machine learning · 41 · 21 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Relationship between objective and subjective perceptual measures of speech in individuals with head and neck cancer
abstract
Meaningful speech assessment is vital in clinical phonetics and therapy monitoring. This study examined the link between perceptual speech assessments and objective acoustic measures in a large head and neck cancer (HNC) dataset. Trained listeners provided ratings of intelligibility, articulation, voice quality, phonation, speech rate, nasality, and background noise on speech. Strong correlations were found between subjective intelligibility, articulation, and voice quality, likely due to a shared underlying cause of speech symptoms in our speaker population. Objective measures of intelligibility and speech rate aligned with their subjective counterpart. Our results suggest that a single intelligibility measure may be sufficient for the clinical monitoring of speakers treated for HNC using concomitant chemoradiation.
Bence Mark Halpern, Thomas Tienkamp, Teja Rebernik, R. J. J. H. van Son, Martijn Wieling 0001, Defne Abur, Tomoki Toda
INTERSPEECH4
2025 Articulatory clarity and variability before and after surgery for tongue cancer
abstract
Surgical treatment for tongue cancer can negatively affect the mobility and musculature of the tongue, which can influence articulatory clarity and variability. In this study, we investigated articulatory clarity through the vowel articulation index (VAI) and variability through vowel formant dispersion (VFD). Using a sentence reading task, we assessed 11 individuals pre and six months post tongue cancer surgery, alongside 11 sex- and age-matched typical speakers. Our results show that while the VAI was significantly smaller post-surgery compared to pre-surgery, there was no significant difference between patients and typical speakers at either time point. Post-surgery, speakers had higher VFD values for /i/ compared to pre-surgery and typical speakers, signalling higher variability. Taken together, our results suggest that while articulatory clarity remained within typical ranges following surgery for tongue cancer for the speakers in our study, articulatory variability increased.
Thomas Tienkamp, Fleur van Ast, Roos van der Veen, Teja Rebernik, Raoul Buurke, Nikki Hoekzema, Katharina Polsterer, Hedwig Sekeres, R. J. J. H. van Son, Martijn Wieling 0001, Max J. H. Witjes, Sebastiaan A. H. J. de Visscher, Defne Abur
INTERSPEECH9
2023 Improving Severity Preservation of Healthy-to-Pathological Voice Conversion With Global Style Tokens
abstract
In healthy-to-pathological voice conversion (H2P-VC), healthy speech is converted into pathological while preserving the identity. The paper improves on previous two-stage approach to H2 P-VC where (1) speech is created first with the appropriate severity, (2) then the speaker identity of the voice is converted while preserving the severity of the voice. Specifically, we propose improvements to (2) by using phonetic posteriorgrams (PPG) and global style tokens (GST). Furthermore, we present a new dataset that contains parallel recordings of pathological and healthy speakers with the same identity which allows more precise evaluation. Listening tests by expert listeners show that the framework preserves severity of the source sample, while modelling target speaker’s voice. We also show that (a) pathology impacts x-vectors but not all speaker information is lost, (b) choosing source speakers based on severity labels alone is insufficient.
Bence Mark Halpern, Wen-Chin Huang, Lester Phillip Violeta, R. J. J. H. van Son, Tomoki Toda
ASRU4
2023 Automatic evaluation of spontaneous oral cancer speech using ratings from naive listeners
abstract
In this paper, we build and compare multiple speech systems for the automatic evaluation of the severity of a speech impairment due to oral cancer, based on spontaneous speech. To be able to build and evaluate such systems, we collected a new spontaneous oral cancer speech corpus from YouTube consisting of 124 utterances rated by 100 non-expert listeners and one trained speech-language pathologist, which we made publicly available. We evaluated the systems in two scenarios: a scenario where transcriptions were available (reference-based) and a scenario where transcriptions might not be available (reference-free). The results of extensive experiments showed that (1) when transcriptions were available, the highest correlation with the human severity ratings was obtained using an automatic speech recognition (ASR) retrained with oral cancer speech. (2) When transcriptions were not available, the best results were achieved by a LASSO model using modulation spectrum features. (3) We found that naive listeners’ ratings are highly similar to the speech pathologist’s ratings for speech severity evaluation. (4) The use of binary labels led to lower correlations of the automatic methods with the human ratings than using severity scores.
Bence Mark Halpern, Siyuan Feng 0001, R. J. J. H. van Son, Michiel W. M. van den Brekel, Odette Scharenborg
Speech Commun.3
2022 Compensation in Verbal and Nonverbal Communication after Total Laryngectomy
abstract
Total laryngectomy is a major surgical procedure with life-changing consequences. As a result of the surgery, the upper and lower airways are disconnected, the natural voice is lost, and patients breathe through a tracheostoma in the neck. Tracheoesophageal speech is the most common speech rehabilitation technique. Due to the lack of air volume, and the amount of muscle tension in the esophagus, some patients may suffer from a hyper- or hypo-tonic voice, resulting in less intelligible speech. To communicate as intelligibly as possible, patients likely adapt their verbal and nonverbal communication to their physical disabilities. The current study aimed to explore the compensation techniques in verbal and nonverbal communication after total laryngectomy focusing on the complexity of grammar and the use of co-speech gestures. We analyzed previously obtained interviews of eight laryngectomized women on the syntactic complexity in speech and the use and type of co-speech gestures. Results were compared with analyses of productions by healthy controls. We found that laryngectomized women reduce the syntactic complexity of their speech, and use nonverbal gestures in their communication. Further research is needed with systematically obtained data and more suitable match-groups.
Marise Neijman, Femke Hof, Noelle Oosterom, Roland Pfau, Bertus van Rooy, R. J. J. H. van Son, Michiel W. M. van den Brekel
INTERSPEECH6
2022 Adjustable deterministic pseudonymization of speech
abstract
While public speech resources become increasingly available, there is a growing interest to preserve the privacy of the speakers, through methods that anonymize the speaker information from speech while preserving the spoken linguistic content. In this paper, a method for pseudonymization (reversible anonymization) of speech is presented, that allows to obfuscate the speaker identity in untranscribed running speech. The approach manipulates the spectro-temporal structure of the speech to simulate a different length and structure of the vocal tract by modifying the formant locations, as well as by altering the pitch and speaking rate. The method is deterministic and partially reversible, and the changes are adjustable on a continuous scale. The method has been evaluated in terms of (i) ABX listening experiments, and (ii) automatic speaker verification and speech recognition. ABX experimental results indicate that the speaker identifiability among forced choice pairs reduced from over 90% to less than 70% through pseudonymization, and that de-pseudonymization was partially effective. An evaluation on the VoicePrivacy 2020 challenge data showed that the proposed approach performs better than the signal processing based baseline method that uses McAdams coefficient and performs slightly worse than the neural source filtering based baseline method. Further analysis showed that the proposed approach: (i) is comparable to the neural source filtering baseline based method in terms of phone posterior feature based objective intelligibility measure, (ii) preserves formant tracks better than the McAdams based method, and (iii) preserves paralinguistic aspects such as dysarthria in several speakers.
S. Pavankumar Dubagunta, R. J. J. H. van Son, Mathew Magimai-Doss
Comput. Speech Lang.2
2022 Low-resource automatic speech recognition and error analyses of oral cancer speech
abstract
In this paper, we introduce a new corpus of oral cancer speech and present our study on the automatic recognition and analysis of oral cancer speech. A two-hour English oral cancer speech dataset is collected from YouTube. Formulated as a low-resource oral cancer ASR task, we investigate three acoustic modelling approaches that previously have worked well with low-resource scenarios using two different architectures; a hybrid architecture and a transformer-based end-to-end (E2E) model: (1) a retraining approach; (2) a speaker adaptation approach; and (3) a disentangled representation learning approach (only using the hybrid architecture). The approaches achieve a (1) 4.7% (hybrid) and 7.5% (E2E); (2) 7.7%; and (3) 2.0% absolute word error rate reduction, respectively, compared to a baseline system which is not trained on oral cancer speech. A detailed analysis of the speech recognition results shows that (1) plosives and certain vowels are the most difficult sounds to recognise in oral cancer speech — this problem is successfully alleviated by our proposed approaches; (3) however these sounds are also relatively poorly recognised in the case of healthy speech with the exception of/p/. (2) recognition performance of certain phonemes is strongly data-dependent; (4) In terms of the manner of articulation, E2E performs better with the exception of vowels — however, vowels have a large contribution to overall performance. As for the place of articulation, vowels, labiodentals, dentals and glottals are better captured by hybrid models, E2E is better on bilabial, alveolar, postalveolar, palatal and velar information. (5) Finally, our analysis provides some guidelines for selecting words that can be used as voice commands for ASR systems for oral cancer speakers.
Bence Mark Halpern, Siyuan Feng 0001, R. J. J. H. van Son, Michiel W. M. van den Brekel, Odette Scharenborg
Speech Commun.3
2021 Measuring Voice Quality Parameters After Speaker Pseudonymization
abstract
Collecting and sharing speech resources is important for progress in speech science and technology. Often, speech resources cannot be shared because of concerns over the privacy of the speakers, e.g., minors or people with medical conditions. Current technologies for pseudonymizing speech have only been tested on “standard” speech for which pseudonymization methods are evaluated on speaker identification risk, intelligibility, and naturalness. For many applications, the important characteristics are para-linguistic aspects of the speech, e.g., voice quality, emotion, or disease progression. Little information is available about the extent to which speaker pseudonymization methods preserve such paralinguistic information. The current study investigates how well voice quality parameters are preserved by an example speech pseudonymization application. Correlations prove to be high between original and pseudonymized recordings for seven acoustic parameters and a composite measure of dysphonia, the AVQI. Root mean square errors for these parameters were reasonably small. A linear mixed effect model shows a link between the difference between source and target speaker and the size of the absolute difference in the AVQI. It is argued that new measures of quality are needed for pseudonymized non-standard speech before wide-spread application of pseudonymized speech can be considered in research and clinical practise.
R. J. J. H. van Son
Interspeech1
2020 Detecting and Analysing Spontaneous Oral Cancer Speech in the Wild
abstract
Oral cancer speech is a disease which impacts more than half a million people worldwide every year. Analysis of oral cancer speech has so far focused on read speech. In this paper, we 1) present and 2) analyse a three-hour long spontaneous oral cancer speech dataset collected from YouTube. 3) We set baselines for an oral cancer speech detection task on this dataset. The analysis of these explainable machine learning baselines shows that sibilants and stop consonants are the most important indicators for spontaneous oral cancer speech detection.
Bence Mark Halpern, R. J. J. H. van Son, Michiel W. M. van den Brekel, Odette Scharenborg
INTERSPEECH2
2019 CNN-Based Phoneme Classifier from Vocal Tract MRI Learns Embedding Consistent with Articulatory Topology
abstract
Recent advances in real-time magnetic resonance imaging (rtMRI) of the vocal tract provides opportunities for studying human speech. This modality together with acquired speech may enable the mapping of articulatory configurations to acoustic features. In this study, we take the first step by training a deep learning model to classify 27 different phonemes from midsagittal MR images of the vocal tract.An American English database was used to train a convolutional neural network for classifying vowels (13 classes), consonants (14 classes) and all phonemes (27 classes) of 17 subjects. Classification top-1 accuracy of the test set for all phonemes was 57%. Erroranalysis showedvoiced and unvoiced sounds often being confused. Moreover, we performed principal component analysis on the network’s embedding and observed topological similarities between thenetwork learned representation and the vowel diagram.Saliency maps gaveinsight intothe anatomical regions most important for classification and show congruence with knownregions of articulatory importance.We demonstrate the feasibility for deep learning to distinguish between phonemes from MRI. Network analysis can be used to improve understanding of normal articulation and speech and, in the future, impaired speech. This study brings us a step closer to the articulatory-to-acoustic mapping from rtMRI.
Kicky G. van Leeuwen, Paula Bos, Stefano Trebeschi, Maarten J. A. van Alphen, Luuk Voskuilen, Ludi E. Smeele, Ferdinand van der Heijden, R. J. J. H. van Son
INTERSPEECH8
2018 Vowel Space as a Tool to Evaluate Articulation Problems
abstract
Treatment for oral tumors can lead to long term changes in the anatomy and physiology of the vocal tract and result in problems with articulation. There are currently no readily available automatic methods to evaluate changes in articulation. We developed a Praat script which plots and measures vowel space coverage. The script reproduces speaker specific vowel space use and speaking-style dependent vowel reduction in normal speech from a Dutch corpus. Speaker identity and speaking style explain more than 60% of the variance in the measured area of the vowel triangle. In recordings of patients treated for oral tumors, vowel space use before and after treatment is still significantly correlated. Articulation before and after treatment is evaluated in a listening experiment and from a maximal articulation speed task. Linear models can explain 50-75% of variance in perceptual ratings and relative articulation rate from values at previous recordings and vowel space measures.
R. J. J. H. van Son, Catherine Middag, Kris Demuynck
INTERSPEECH1
2016 Long-Term Stability of Tracheoesophageal Voices
abstract
Long-term voice outcomes of 13 tracheoesophageal speakers are assessed using speech samples that were recorded with at least 7 years in between. Intelligibility and voice quality are perceptually evaluated by 10 experienced speech and language pathologists. In addition, automatic speech evaluations are performed with tools from Ghent University. No significant group effect was found for changes in voice quality and intelligibility. The recordings showed a wide interspeaker variability. It is concluded that intelligibility and voice quality of tracheoesophageal voice is mostly stable over a period of 7 to 18 years.
Klaske E. van Sluis, Michiel W. M. van den Brekel, Frans J. M. Hilgers, R. J. J. H. van Son
INTERSPEECH4
2016 Computing scores of voice quality and speech intelligibility in tracheoesophageal speech for speech stimuli of varying lengths
Renee Peje Clapham, Jean-Pierre Martens, R. J. J. H. van Son, Frans J. M. Hilgers, Michiel W. M. van den Brekel, Catherine Middag
Comput. Speech Lang.3
2015 A Survey on perceived speaker traits: Personality, likability, pathology, and the first challenge
Björn W. Schuller, Stefan Steidl, Anton Batliner, Elmar Nöth, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son, Felix Weninger, Florian Eyben, Tobias Bocklet, Gelareh Mohammadi, Benjamin Weiss 0001
Comput. Speech Lang.7
2015 Introduction
Björn W. Schuller, Stefan Steidl, Anton Batliner, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son
Comput. Speech Lang.6
2014 Robust automatic intelligibility assessment techniques evaluated on speakers treated for head and neck cancer
Catherine Middag, Renee Peje Clapham, R. J. J. H. van Son, Jean-Pierre Martens
Comput. Speech Lang.3
2014 Developing automatic articulation, phonation and accent assessment techniques for speakers treated for advanced head and neck cancer
Renee Peje Clapham, Catherine Middag, Frans J. M. Hilgers, Jean-Pierre Martens, Michiel W. M. van den Brekel, R. J. J. H. van Son
Speech Commun.6
2013 Automatic tracheoesophageal voice typing using acoustic parameters
abstract
The acoustics of isolated vowels, e.g. of /a/, have in many studies been linked to pathological voice types, such as tracheoesophageal (TE) voice. To study the possibilities of objective and automatic classification of pathological TE voice types, the acoustic features of /a/ were quantified and subsequently classified using a suit of machine learning technologies. Best classification was achieved by using a voiced-voiceless measurement and the harmonics-to-noise ratio. Other common acoustic features were correlated to pathological type as well, but were less distinctive in classification. We conclude that for objective and automatic classification of TE voice pathology, voicing distinction and harmonics-to-noise ratio are most relevant.
Renee Peje Clapham, Corina J. van As-Brooks, Michiel W. M. van den Brekel, Frans J. M. Hilgers, R. J. J. H. van Son
INTERSPEECH5
2012 The INTERSPEECH 2012 Speaker Trait Challenge
abstract
LIDIAP
Björn W. Schuller, Stefan Steidl, Anton Batliner, Elmar Nöth, Alessandro Vinciarelli, Felix Burkhardt, R. J. J. H. van Son, Felix Weninger, Florian Eyben, Tobias Bocklet, Gelareh Mohammadi, Benjamin Weiss 0001
INTERSPEECH7
2012 NKI-CCRT Corpus - Speech Intelligibility Before and After Advanced Head and Neck Cancer Treated with Concomitant Chemoradiotherapy
Renee Peje Clapham, Lisette van der Molen, R. J. J. H. van Son, Michiel W. M. van den Brekel, Frans J. M. Hilgers
LREC3
2010 Manipulating treacheoesophageal speech
abstract
Speech therapy aiming at improving voice quality and speech intelligibility is often hampered by the lack of knowledge of the underlying deficits.One way to help speech therapists treating patients would be to supply synthetic benchmarks for pathological speech.These can be used to train therapists and evaluate and interpret automatic speech recognizers used for diagnosing pathological speech.Moreover, synthetic pathological speech can also be used to make expected therapy aims audible before treatment.In a listening experiment testing perceived intelligibility, three types of manipulations of tracheoesophageal speech were evaluated by experienced speech therapists.It was found that modeling the intensity contour of the voice source signal improved speech quality over plain analysis-synthesis.Replacing the voicing source with fully synthetic source periods decreased the perceived intelligibility markedly.Making the source fully periodic with a regular pitch had no effect on perceived intelligibility.Low quality speech benefitted more from manipulations, or deteriorated less, than high quality speech.
R. J. J. H. van Son, Irene Jacobi, Frans J. M. Hilgers
INTERSPEECH1
2008 The IFADV Corpus: a Free Dialog Video Corpus
R. J. J. H. van Son, Wieneke Wesseling, Eric Sanders, Henk van den Heuvel
LREC1
2007 Formal modelling of L1 and L2 perceptual learning: computational linguistics versus machine learning
abstract
Abstract ∗ In this paper, we evaluate the adequacy of two widely used machine learning algorithms and a computational linguistic proposal to model L2 perceptual development. The three proposals are, in order, Nearest Neighbor, Naive Bayesian and Stochastic OT and the Gradual Learning Algorithm. We compared the three models ’ outputs to those of Spanish learners of Dutch who were asked to categorize synthetic stimuli as one of the 12 Dutch vowels. The empirical results of the human learners show that L2 learners differ significantly from native listeners, but also that their perceptual spaces tend to become more native-like with L2 proficiency. The results of the simulations show that all three algorithms are able to model listeners ’ data to a certain extent but that Stochastic OT and the Gradual Learning Algorithm, i.e. the linguistic model, best reproduces L1 and L2 data. 1.
Paola Escudero, Jelle Kastelein, Klara A. Weiand, R. J. J. H. van Son
INTERSPEECH4
2007 Learning tone distinctions for Mandarin Chinese
abstract
We describe the SpeakGoodChinese system that supports beginning students of Mandarin Chinese to produce tones correctly
David Weenink, Guangqin Chen, Zongyan Chen, Stefan de Konink, Dennis Vierkant, Eveline van Hagen, R. J. J. H. van Son
INTERSPEECH7
2007 The influence of masking words on the prediction of TRPs in a shadowed dialog
abstract
It is well known that listeners can ignore disturbances in speech and rely on context to interpolate the message.This fact is used to determine the importance of individual words for projecting Transition Relevance Places, TRPs.Subjects were asked to shadow manipulated pre-recorded dialogs with minimal responses, saying 'ah' when they feel it is appropriate.In these dialogs, at random, of each utterance, either one of the last four words was replaced by white noise (masked condition), or no word was replaced (non masked condition).The reaction times were analyzed for effects of masked words.The presence of masked words, even prominent words, did not affect the response times of our subjects unless the very last word of the utterance was masked.This indicates that listeners are able to seamlessly interpolate the missing words and only need the identity of the last word to determine the exact position of the TRP.
Wieneke Wesseling, R. J. J. H. van Son, Louis C. W. Pols
INTERSPEECH2
2006 Prominent words as anchors for TRP projection
abstract
The effect of the position of the last accented word on the projection of TRPs was investigated with two RT experiments.Subjects were asked to respond with minimal responses to prerecorded dialogs and impoverished versions of these dialogs, containing either only intonation and pause information,hummed stimuli, or no periodic component at all, whispered stimuli.The distribution of these elicited response delays was comparable to that of natural turn switches.It is shown that the presence of non-prominent words before a TRP reduces the delays of elicited and natural responses alike, even in impoverished speech.This suggests that the presence of an prominent, informative, word starts the projection of a possible upcoming TRP.The availability of non-prominent, predictable, speech then allows listeners to improve their predictions of the exact timing of the TRP.
R. J. J. H. van Son, Wieneke Wesseling, Louis C. W. Pols
INTERSPEECH1
2006 On the sufficiency and redundancy of pitch for TRP projection
abstract
In two RT experiments, subjects were asked to respond with minimal responses to prerecorded dialogs and impoverished versions of these dialogs, containing either only intonation and pause information, hummed stimuli, or no periodic component at all, whispered stimuli. For the hummed, stimuli, response delays and, especially, variances were higher than the original recordings. Responses to mid-frequency pitch utterance-ends were significantly longer than responses to low pitch utterance-ends, suggesting that our subjects fell back to reacting to pauses when presented with hummed utterances ending in a mid-frequency tone. This suggests that, in contrast to low or high end-tones, intonation contours that end in a mid-frequency tone might not contain any useful information for predicting end-of-utterance TRPs. We conclude that just the intonation and pauses of a conversation contain sufficient information for projection of TRPs. However this information is measurably impoverished with respect to original to an extent that increases the “processing ” time by 10%. No difference was found between whispered and original speech. This lack of any effect of removing all periodic sound components from the speech signal indicates that in natural speech the pitch signal itself might be redundant for predicting TRPs. 1.
Wieneke Wesseling, R. J. J. H. van Son, Louis C. W. Pols
INTERSPEECH2
2005 Timing of experimentally elicited minimal responses as quantitative evidence for the use of intonation in projecting TRPs
abstract
In an RT experiment, subjects were asked to respond with minimal responses to prerecorded dialogs and a manipulated version of these dialogs that contained only intonation and pause information. Response delays and, especially, variances were higher to the impoverished, intonation only, stimuli than to the original recordings. It was also found that intonation only utterances ending in a mid-frequency pitch induced significantly longer response delays than utterances ending in a low pitch. These results are interpreted as evidence that just the intonation and pauses of a conversation already contain sufficient information to project end-of-utterance TRPs. However this information is measurably impoverished with respect to full speech to an extent that increases the “processing ” time by 10%. Our subjects seemed to fall back to reacting to pauses when presented with intonation only utterances ending in a mid-frequency tone. This suggests that, in contrast to low or high end-tones, intonation contours that end in a mid-frequency tone might not contain any useful information for predicting end-of-utterance TRPs. 1.
Wieneke Wesseling, R. J. J. H. van Son
INTERSPEECH2
2005 Note from the Guest Editors
Cecilia Odé, R. J. J. H. van Son
Speech Commun.2
2005 Duration and spectral balance of intervocalic consonants: A case for efficient communication
R. J. J. H. van Son, Jan P. H. van Santen
Speech Commun.1
2004 Frequency effects on vowel reduction in three typologically different languages (dutch, finish, Russian)
R. J. J. H. van Son, Olga Bolotova, Louis C. W. Pols, Mietta Lennes
INTERSPEECH1
2003 Information structure and efficiency in speech production
abstract
Speech is considered an efficient communication channel. This implies that the organization of utterances is such that more speaking effort is directed towards important parts than towards redundant parts. Based on a model of incremental word recognition, the importance of a segment is defined as its contribution to word-disambiguation. This importance is measured as the segmental information content, in bits. On a labeled Dutch speech corpus it is then shown that crucial aspects of the information structure of utterances partition thesegmental information content and explain 90% of the variance. Two measures of acoustical reduction, duration and spectral center of gravity, are correlated with the segmental information content in such a way that more important phonemes are less reduced. It is concluded that the organization of conventional information structure does indeed increase efficiency.
R. J. J. H. van Son, Louis C. W. Pols
INTERSPEECH1
2002 Evidence for efficiency in vowel production
abstract
Speaking is generally considered efficient in that less effort is spent articulating more redundant items.With efficient speech production, less reduction is expected in the pronunciation of phonemes that are more important (distinctive) for word identification.The importance of a single phoneme in word recognition can be quantified as the information (in bits) it adds to the preceding word onset to narrow down the lexical search.In our study, segmental information showed to correlate consistently with two measures of reduction: vowel duration and formant reduction.This correlation was found after accounting for speaker and vowel identity, speaking style, lexical stress, modeled prominence, and position of the syllable in the word.However, consistent correlations are only found in high-frequency words.Furthermore, the correlation is strongest in normal reading and weaker in spontaneous and anomalous read speech.Combined, these facts suggest that this type of efficiency in production might rely on retrieving stored words from memory.Efficiency in vowel production seems to be less or absent when words have to be assembled on-line.
R. J. J. H. van Son, Louis C. W. Pols
INTERSPEECH1
2001 The IFA corpus: a phonemically segmented dutch "open source" speech database
abstract
An open source database of hand-segmented Dutch speech was constructed with off-the-shelf software using speech from 8 speakers in a variety of speaking styles.For a total of 50,000 words, speech acquisition and preparation took around 3 person-weeks per speaker.Hand segmentation took 1,000 hours of labeling altogether.The asymptotic segmentation speed was about one word, or four boundaries, per minute.An evaluation showed that the Median Absolute Difference of the segment boundaries was 6 ms between labelers, and 4 ms within labelers.Label differences (substitutions, insertions, and deletions) were found in 8% of the segments between labelers and 5% within labelers.Compiled data are available in relational database format for querying with SQL.
R. J. J. H. van Son, Diana Binnenpoorte, Henk van den Heuvel, Louis C. W. Pols
INTERSPEECH1
2000 An acoustic profile of speech efficiency
R. J. J. H. van Son, Barbertje M. Streefkerk, Louis C. W. Pols
INTERSPEECH1
1999 Effects of stress and lexical structure on speech efficiency
R. J. J. H. van Son, Louis C. W. Pols
EUROSPEECH1
1999 An acoustic description of consonant reduction
R. J. J. H. van Son, Louis C. W. Pols
Speech Commun.1
1999 Perisegmental speech improves consonant and vowel identification
R. J. J. H. van Son, Louis C. W. Pols
Speech Commun.1
1998 Efficiency as an organizing principle of natural speech
abstract
A large part of the variation in natural speech appears along the dimensions of articulatory precision / perceptual distinctiveness. We propose that this variation is the result of an effort to communicate efficiently. Speaking i s considered efficient if the speech sound contains only the information needed to understand it. This efficiency is tested by means of a corpus of spontaneous and matched read speech, and syllable and word frequencies as measures of information content (12007 syllables, 8046 word forms, 1582 intervocalic consonants, and 2540 vowels). It is indeed found that the duration and spectral reduction of consonants and vowels correlate with the frequency of syllables and words in this corpus. Consonant intelligibility correlates with both the acoustic factors and the syllable and word frequencies. It i s concluded that the principle of efficient communication organizes at least some aspects of speech production.
R. J. J. H. van Son, Florien J. van Beinum, Louis C. W. Pols
ICSLP1
1997 The correlation between consonant identification and the amount of acoustic consonant reduction
abstract
Reduction causes changes in the acoustics of consonant realizations that affect their identification. In this study we try to identify some of the acoustic parameters that are correlated with this change in identification. Speaking style is used to manipulate the degree of reduction. Pairs of otherwise identical intervocalic consonants from read and spontaneous utterances are presented to subjects in an identification experiment. The resulting identification scores are correlated to five different acoustical measures that are affected by the amount of consonant reduction: Segmental duration, spectral Center of Gravity, intervocalic sound energy difference, intervocalic F 2 slope difference, and the amount of vowel reduction in the syllable kernel. The identification differences between the read and spontaneous realizations are compared with the differences in each of the acoustic measures. It showed that only segmental duration and the spectral Center of Gravity are significantly corre...
R. J. J. H. van Son, Louis C. W. Pols
EUROSPEECH1
1997 Strong interaction between factors influencing consonant duration
R. J. J. H. van Son, Jan P. H. van Santen
EUROSPEECH1
1996 An acoustic profile of consonant reduction
R. J. J. H. van Son, Louis C. W. Pols
ICSLP1
1995 A method to quantify the error distribution in confusion matrices
R. J. J. H. van Son
EUROSPEECH1
1995 The influence of local context on the identification of vowels and consonants
R. J. J. H. van Son, Louis C. W. Pols
EUROSPEECH1
1995 What does consonant reduction look like, if it exists?
R. J. J. H. van Son, Louis C. W. Pols
EUROSPEECH1
1993 Vowel identification as influenced by vowel duration and formant track shape
R. J. J. H. van Son, Louis C. W. Pols
EUROSPEECH1
1993 Acoustics and perception of dynamic vowel segments
Louis C. W. Pols, R. J. J. H. van Son
Speech Commun.2
1991 The influence of formant track shape on the perception of synthetic vowels
R. J. J. H. van Son, Louis C. W. Pols
EUROSPEECH1
1989 Comparing formant movements in fast and normal rate speech
R. J. J. H. van Son, Louis C. W. Pols
EUROSPEECH1