Roeland van Hout

dblp:98/8063 · also R. W. N. M. van Hout · DBLP profile ↗
← Back
28ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-8870-1631ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 9 since 2021
YearPublicationVenuePosition
2026 Predicting accentedness and comprehensibility through ASR scores and acoustic features
abstract
• Compared two automatic speech recognition models (TDNN and Whisper) for predicting accentedness and comprehensibility. • Used a data-driven method to select the most relevant acoustic features for accentedness and comprehensibility. • Used a linear mixed-effects model to incorporate speaker and utterance differences. • Combined segmental and suprasegmental features to better understand accentedness and comprehensibility of non-native speech. Accentedness and comprehensibility scales are widely used in measuring the oral proficiency of second language (L2) learners, including learners of English as a Second Language (ESL). In this paper, we focus on gaining a better understanding of the concepts of accentedness and comprehensibility by developing and applying automatic measures to ESL utterances produced by Indonesian learners. We extracted features both on the segmental and the suprasegmental (fundamental frequency, loudness, energy et al.) levels to investigate which features are actually related to expert judgments on accentedness and comprehensibility. Automatic Speech Recognition (ASR) pronunciation scores based on the traditional Kaldi Time Delay Neural Network (TDNN) model and on the End-to-End Whisper model were applied, and data-driven methods were used by combining acoustic features extracted by the Geneva Minimalistic Acoustic Parameter Set (eGeMAPS) and Praat. The experimental results showed that Whisper outperformed the Kaldi-TDNN model. The Whisper model gave the best results for predicting comprehensibility on the basis of phone distance, and the best results for predicting accentedness on the basis of grapheme distance. Combining segmental and suprasegmental features improved the results, yielding different feature rankings for comprehensibility and accentedness. In our final step of analysis, we included differences between utterances and learners as random effects in a mixed linear regression model. Exploiting these information sources yielded a substantial improvement in predicting both comprehensibility and accentedness.
Wenwei Dong, Catia Cucchiarini, Roeland van Hout, Helmer Strik
Comput. Speech Lang.3
2025 Evaluating Progress of CALL System Users on Accentedness and Comprehensibility: An Acoustic and ASR-Based Approach
abstract
Contains fulltext : 322930.pdf (Publisher’s version ) (Open Access)
Wenwei Dong, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH3
2025 Can ASR generate valid measures of child reading fluency?
Wieke Harmsen, Roeland van Hout, Catia Cucchiarini, Helmer Strik
INTERSPEECH2
2023 An ASR-enabled Reading Tutor: Investigating Feedback to Optimize Interaction for Learning to Read
abstract
Contains fulltext : 299814.pdf (Publisher’s version ) (Open Access)
Ferdy Hubers, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH4
2023 Assessing Intelligibility in Non-native Speech: Comparing Measures Obtained at Different Levels
abstract
Contains fulltext : 301527.pdf (Publisher’s version ) (Open Access)
Roeland van Hout, Catia Cucchiarini, Danielle Reuvekamp, Helmer Strik
INTERSPEECH2
2023 Measuring the intelligibility of dysarthric speech through automatic speech recognition in a pluricentric language
abstract
Speech intelligibility is an essential though complex construct for evaluating dysarthric speech. Various procedures can be used to measure speech intelligibility, most of which are based on subjective ratings assigned by experts. Since these procedures are subjective and laborious, automatic speech recognition (ASR) has been proposed to obtain objective metrics of intelligibility. Although promising results have been reported, ASR for dysarthric speech generally requires large amounts of data consisting of recorded and annotated speech. In the present study, we explored the possibility of using dysarthric speech resources from the dominant language variety to improve the performance of ASR systems on the dysarthric speech of the non-dominant variety of the same pluricentric language. Dutch is used as an example of a pluricentric language, with Netherlandic Dutch considered the dominant and Flemish Dutch the non-dominant variety. The performance of ASR is evaluated by using two types of intelligibility metrics: orthographic transcriptions and global intelligibility assessments, both obtained from experts. Overall, the results show that dysarthric speech data from the dominant language variety can contribute to improving automatic transcriptions and to developing objective, automatic global measures of speech intelligibility only when no data from the non-dominant variety are available for training ASR models.
Catia Cucchiarini, Roeland van Hout, Helmer Strik
Speech Commun.3
2022 The Effects of Implicit and Explicit Feedback in an ASR-based Reading Tutor for Dutch First-graders
abstract
Contains fulltext : 288876.pdf (Publisher’s version ) (Open Access)
Ferdy Hubers, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH4
2022 Automatic Speech Recognition and Pronunciation Error Detection of Dutch Non-native Speech: cumulating speech resources in a pluricentric language
abstract
The shortage of large-scale learners’ speech corpora and precise manual annotations are two major challenges for automatic L2 speech recognition and error detection in L2 speech, especially for non-dominant varieties of pluricentric languages. In these cases, collecting and annotating large non-native (L2 learner) corpora for all language varieties is often unattainable. In this study, we investigated ways of addressing these problems through conventional and transfer learning Deep Neural Network (DNN) based Automatic Speech Recognition (ASR) and ASR-based pronunciation error detection (PED) by cumulating Netherlandic Dutch and Flemish Dutch speech resources. First, we show that for ASR the baseline system can be improved by combining the Netherlandic Dutch and Flemish Dutch datasets. Next, through the knowledge learned from models trained on the Netherlandic Dutch data, the Flemish Dutch learners' ASR model can be further improved. In order to evaluate the performance of the PED algorithms in the absence of learner speech data with pronunciation error annotations, we introduced plausible pronunciation errors in the native corpora based on knowledge from Flemish learner speech, in order to simulate non-native speech errors. For PED we found that the results are much better for a GOP classifier trained on Flemish Dutch data than for one trained on Netherlandic Dutch data. PED produced worse results when the Netherlandic Dutch data were merged with the Flemish Dutch data, while for ASR, lower WERs were attained. Whether adding Netherlandic Dutch data to Flemish Dutch data is beneficial, thus seems to depend on the specific task the data are used for. We discuss these results, compare them to those of related research and suggest avenues for future research.
Catia Cucchiarini, Roeland van Hout, Helmer Strik
Speech Commun.3
2021 Speech Intelligibility of Dysarthric Speech: Human Scores and Acoustic-Phonetic Features
abstract
We investigated speech intelligibility in dysarthric and nondysarthric speakers as measured by two commonly used metrics, ratings through the Visual Analogue Scale (VAS) and word accuracy (AcW) through orthographic transcriptions.To gain a better understanding of how acoustic-phonetic correlates could be employed to obtain more objective measures of speech intelligibility and a better classification of dysarthric and non-dysarthric speakers, we studied the relation between these measures of intelligibility and some important acoustic-phonetic correlates.We found that the two intelligibility measures are related, but distinct, and that they might refer to different components of the intelligibility construct.The acoustic-phonetic features showed no difference in the mean values between the two speaker types at the utterance level, but more than half of them played a role in classifying the two speaker types.We computed an acoustic-phonetic probability index (API) at the speaker level.API is moderately correlated to VAS ratings but not correlated to AcW.In addition, API and VAS complement each other in classifying dysarthric and non-dysarthric speakers.This suggests that the intelligibility measures assigned by human raters and acoustic-phonetic features relate to different constructs of intelligibility.
Roeland van Hout, Fleur Boogmans, Mario Ganzeboom, Catia Cucchiarini, Helmer Strik
Interspeech2
2021 The effect of intermittent noise on lexically-guided perceptual learning in native and non-native listening
abstract
There is ample evidence that both native and non-native listeners deal with speech variation by quickly tuning into a speaker and adjusting their phonetic categories according to the speaker’s ambiguous pronunciation. This process is called lexically-guided perceptual learning. Moreover, the presence of noise in the speech signal has previously been shown to change the word competition process by increasing the number of candidate words competing for recognition and slowing down the recognition process. Given that reliable lexical information should be available quickly to induce lexically-guided perceptual learning and that word recognition is slowed down in the presence of noise, and especially so for non-native listeners, the present study investigated whether noise interferes with lexically-guided perceptual learning in native and non-native listening. Native English and Dutch listeners were exposed to a story in English in clean speech or with stretches of noise. All the /l/ and /ɹ/ sounds in the story were replaced with an ambiguous sound half-way between /l/ and /ɹ/. Although noise altered the pattern of responses for the non-native listeners in a subsequent phonetic categorization task, both native and non-native listeners demonstrated lexically-guided perceptual learning in both clean and noisy listening conditions. We argue that the robustness of perceptual learning in the presence of intermittent noise for both native and non-native listeners is additional evidence for the remarkable flexibility of native and non-native perceptual systems even in adverse listening conditions.
Polina Drozdova, Roeland van Hout, Sven L. Mattys, Odette Scharenborg
Speech Commun.2
2020 Analyzing Read Aloud Speech by Primary School Pupils: Insights for Research and Development
abstract
Contains fulltext : 228091.pdf (Publisher’s version ) (Open Access)
S. Limonard, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH3
2020 Towards a Comprehensive Assessment of Speech Intelligibility for Pathological Speech
abstract
Contains fulltext : 228265pub.pdf (Publisher’s version ) (Open Access)
Viviana Mendoza Ramos, Wieke Harmsen, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH5
2018 A Fast and Flexible Webinterface for Dialect Research in the Low Countries
Roeland van Hout, Nicoline van der Sijs, Erwin Komen, Henk van den Heuvel
LREC1
2016 Processing and Adaptation to Ambiguous Sounds during the Course of Perceptual Learning
abstract
Contains fulltext : 162019.pdf (Publisher’s version ) (Open Access)
Polina Drozdova, Roeland van Hout, Odette Scharenborg
INTERSPEECH2
2016 Does the Importance of Word-Initial and Word-Final Information Differ in Native versus Non-Native Spoken-Word Recognition?
abstract
Contains fulltext : 161987.pdf (Publisher’s version ) (Open Access)
Odette Scharenborg, Juul Coumans, Sofoklis Kakouros, Roeland van Hout
INTERSPEECH4
2016 Palabras: Crowdsourcing Transcriptions of L2 Speech
Eric Sanders, Pepi Burgos, Catia Cucchiarini, Roeland van Hout
LREC4
2015 Auris populi: crowdsourced native transcriptions of Dutch vowels spoken by adult Spanish learners
abstract
\n Contains fulltext :\n 145184.pdf (Publisher’s version ) (Open Access)\n
Pepi Burgos, Eric Sanders, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH4
2015 Confusability in L2 vowels: analyzing the role of different features
abstract
Contains fulltext : 150849.pdf (Publisher’s version ) (Open Access)
Matyas Jani, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH3
2015 Subjective accent strength perceptions are not only a function of objective accent strength. Evidence from Netherlandic Standard Dutch
Stefan Grondelaers, Roeland van Hout, Sander van der Harst
Speech Commun.2
2014 Dutch vowel production by Spanish learners: duration and spectral features
abstract
In this paper we present a study on Dutch vowel production by Spanish learners that was carried out within the framework of our research on Computer Assisted Pronunciation Training (CAPT). The aim of this study was to obtain detailed information on production of Dutch vowels by Spanish learners, which can be employed to develop effective CAPT programs for this specific target group. We collected speech from learners with varying proficiency levels (A1 - B2 of the CEFR), which was transcribed, segmented and acoustically analyzed. We present data on the frequency of pronunciation errors and on detailed analyses of duration and acoustic properties of the vocalic realizations. The results indicate that Spanish learners of Dutch have difficulties in realizing several Dutch vowel contrasts and that they differ from native speakers in the way they employ duration and spectral properties to realize these contrasts. We discuss these results in relation to those of previous studies on Dutch vowel perception by Spanish listeners and relate them to current theories on speech learning. Index Terms: L2 phonology acquisition, language learning, Computer Assisted Pronunciation Training (CAPT)
Pepi Burgos, Matyas Jani, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH4
2014 Non-native word recognition in noise: the role of word-initial and word-final information
abstract
Contains fulltext : 131628.pdf (Publisher’s version ) (Open Access)
Juul Coumans, Roeland van Hout, Odette Scharenborg
INTERSPEECH2
2014 Phoneme category retuning in a non-native language
abstract
Item does not contain fulltext
Polina Drozdova, Roeland van Hout, Odette Scharenborg
INTERSPEECH2
2014 ASR-based CALL systems and learner speech data: new resources and opportunities for research and development in second language learning
Catia Cucchiarini, Steve Bodnar, Bart Penning de Vries, Roeland van Hout, Helmer Strik
LREC4
2014 Vulnerability in Acquisition, Language Impairments in Dutch: Creating a VALID Data Archive
Jetske Klatter, Roeland van Hout, Henk van den Heuvel, Paula Fikkert, Anne Baker, Jan De Jong, Frank Wijnen, Eric Sanders, Paul Trilsbeek
LREC2
2013 Pronunciation errors by Spanish learners of Dutch: a data-driven study for ASR-based pronunciation training
abstract
In this paper we report on a study on pronunciation errors by Spanish learners of Dutch, which was aimed at obtaining information to develop a dedicated Computer Assisted Pronunciation Training (CAPT) program for this fixed language pair (Spanish L1, Dutch L2).The results of our study indicate, that, first, vowel errors are more frequent and variable than consonant mispronunciations.Second, Spanish natives appear to have problems with vowel length, vowel height, and front rounded vowels.Third, they tend to fall back on the pronunciation of their L1 vowels.
Pepi Burgos, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH3
2008 Evaluating the Relationship between Linguistic and Geographic Distances using a 3D Visualization
Folkert de Vriend, Jan Pieter Kunst, Louis ten Bosch, Charlotte Giesbers, Roeland van Hout
LREC5
2006 A Unified Structure for Dutch Dialect Dictionary Data
Folkert de Vriend, Lou Boves, Henk van den Heuvel, Roeland van Hout, Joep Kruijsen, Jos Swanenberg
LREC4
2001 A comparison between human vowel normalization strategies and acoustic vowel transformation techniques
abstract
Perceptual and acoustic representations of vowel data were compared directly to evaluate the perceptual relevance of several speaker normalization transformations. The acoustic representations consisted of raw F0 and formant data. The perceptual representations were obtained through an experimental procedure, with phonetically trained listeners as subjects. The raw acoustic data were transformed according to several normalization schemes. The perceptual and the acoustic representations were compared using regression techniques. A zscore-transformation of the raw data appeared to resemble the perceptual data.
Patti Adank, Roeland van Hout, Roel Smits
INTERSPEECH2