Catia Cucchiarini

dblp:28/1285 · DBLP profile ↗
← Back
84ranked-venue papers
20as first author
19since 2021 · last 2026
0000-0001-5908-0824ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 74 · 17 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 61 · 14 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment
abstract
Contains fulltext : 331543.pdf (Publisher’s version ) (Closed access)
Aditya Parikh, Cristian Tejedor García, Catia Cucchiarini, Helmer Strik
LREC3
2026 Predicting accentedness and comprehensibility through ASR scores and acoustic features
abstract
• Compared two automatic speech recognition models (TDNN and Whisper) for predicting accentedness and comprehensibility. • Used a data-driven method to select the most relevant acoustic features for accentedness and comprehensibility. • Used a linear mixed-effects model to incorporate speaker and utterance differences. • Combined segmental and suprasegmental features to better understand accentedness and comprehensibility of non-native speech. Accentedness and comprehensibility scales are widely used in measuring the oral proficiency of second language (L2) learners, including learners of English as a Second Language (ESL). In this paper, we focus on gaining a better understanding of the concepts of accentedness and comprehensibility by developing and applying automatic measures to ESL utterances produced by Indonesian learners. We extracted features both on the segmental and the suprasegmental (fundamental frequency, loudness, energy et al.) levels to investigate which features are actually related to expert judgments on accentedness and comprehensibility. Automatic Speech Recognition (ASR) pronunciation scores based on the traditional Kaldi Time Delay Neural Network (TDNN) model and on the End-to-End Whisper model were applied, and data-driven methods were used by combining acoustic features extracted by the Geneva Minimalistic Acoustic Parameter Set (eGeMAPS) and Praat. The experimental results showed that Whisper outperformed the Kaldi-TDNN model. The Whisper model gave the best results for predicting comprehensibility on the basis of phone distance, and the best results for predicting accentedness on the basis of grapheme distance. Combining segmental and suprasegmental features improved the results, yielding different feature rankings for comprehensibility and accentedness. In our final step of analysis, we included differences between utterances and learners as random effects in a mixed linear regression model. Exploiting these information sources yielded a substantial improvement in predicting both comprehensibility and accentedness.
Wenwei Dong, Catia Cucchiarini, Roeland van Hout, Helmer Strik
Comput. Speech Lang.2
2025 Evaluating Progress of CALL System Users on Accentedness and Comprehensibility: An Acoustic and ASR-Based Approach
abstract
Contains fulltext : 322930.pdf (Publisher’s version ) (Open Access)
Wenwei Dong, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH2
2025 Multitalker Babble in English Vowel Perception Training: A Comparison between Humans and Neural Models
abstract
Contains fulltext : 322924.pdf (Publisher’s version ) (Open Access)
Wenwei Dong, Alif Silpachai, Catia Cucchiarini, Helmer Strik
INTERSPEECH3
2025 Improving Child Speech Recognition and Reading Mistake Detection by Using Prompts
abstract
Automatic reading aloud evaluation can provide valuable support to teachers by enabling more efficient scoring of reading exercises. However, research on reading evaluation systems and applications remains limited. We present a novel multimodal approach that leverages audio and knowledge from text resources. In particular, we explored the potential of using Whisper and instruction-tuned large language models (LLMs) with prompts to improve transcriptions for child speech recognition, as well as their effectiveness in downstream reading mistake detection. Our results demonstrate the effectiveness of prompting Whisper and prompting LLM, compared to the baseline Whisper model without prompting. The best performing system achieved state-of-the-art recognition performance in Dutch child read speech, with a word error rate (WER) of 5.1%, improving the baseline WER of 9.4%. Furthermore, it significantly improved reading mistake detection, increasing the F1 score from 0.39 to 0.73.
Lingyun Gao, Cristian Tejedor García, Catia Cucchiarini, Helmer Strik
INTERSPEECH3
2025 Can ASR generate valid measures of child reading fluency?
Wieke Harmsen, Roeland van Hout, Catia Cucchiarini, Helmer Strik
INTERSPEECH3
2025 Evaluating Logit-Based GOP Scores for Mispronunciation Detection
abstract
Pronunciation assessment relies on goodness of pronunciation (GOP) scores, traditionally derived from softmax-based posterior probabilities. However, posterior probabilities may suffer from overconfidence and poor phoneme separation, limiting their effectiveness. This study compares logit-based GOP scores with probability-based GOP scores for mispronunciation detection. We conducted our experiment on two L2 English speech datasets spoken by Dutch and Mandarin speakers, assessing classification performance and correlation with human ratings. Logit-based methods outperform probability-based GOP in classification, but their effectiveness depends on dataset characteristics. The maximum logit GOP shows the strongest alignment with human perception, while a combination of different GOP scores balances probability and logit features. The findings suggest that hybrid GOP methods incorporating uncertainty modeling and phoneme-specific weighting improve pronunciation assessment.
Aditya Parikh, Cristian Tejedor García, Catia Cucchiarini, Helmer Strik
INTERSPEECH3
2025 Enhancing GOP in CTC-Based Mispronunciation Detection with Phonological Knowledge
abstract
Computer-Assisted Pronunciation Training (CAPT) systems employ automatic measures of pronunciation quality, such as the goodness of pronunciation (GOP) metric. GOP relies on forced alignments, which are prone to labeling and segmentation errors due to acoustic variability. While alignment-free methods address these challenges, they are computationally expensive and scale poorly with phoneme sequence length and inventory size. To enhance efficiency, we introduce a substitution-aware alignment-free GOP that restricts phoneme substitutions based on phoneme clusters and common learner errors. We evaluated our GOP on two L2 English speech datasets, one with child speech, My Pronunciation Coach (MPC), and SpeechOcean762, which includes child and adult speech. We compared RPS (restricted phoneme substitutions) and UPS (unrestricted phoneme substitutions) setups within alignment-free methods, which outperformed the baseline. We discuss our results and outline avenues for future research.
Aditya Parikh, Cristian Tejedor García, Catia Cucchiarini, Helmer Strik
INTERSPEECH3
2024 Reading Miscue Detection in Primary School through Automatic Speech Recognition
abstract
Contains fulltext : 309809.pdf (Publisher’s version ) (Open Access)
Lingyun Gao, Cristian Tejedor García, Helmer Strik, Catia Cucchiarini
INTERSPEECH4
2024 An introduction to pluricentric languages in speech science and technology
abstract
Pluricentric languages are languages that are spoken in at least two countries where they have an official function and thus develop national varieties with specific linguistic and pragmatic features. Presently 43 languages have been identified as belonging to this category, for instance, English, Spanish, German, Bengali, Hindi and Urdu. This article forms an introduction to the special issue “Pluricentric Languages in Speech Science and Technology” by giving an overview of current challenges with respect to the development of speech and language resources, annotation and analysis tools, as well as speech technology services for pluricentric languages. The article discusses potential solutions that come from cross-fertilization: on the one hand, how phonetic and linguistic knowledge may contribute to advancements in speech technology, and on the other, how speech technology may facilitate phonetic and linguistic studies on pluricentric languages. In our discussion, we include the research methods and findings of the eight research articles of this special issue and point towards promising paths for future research in the field.
Barbara Schuppler, Martine Adda-Decker, Catia Cucchiarini, Rudolf Muhr
Speech Commun.3
2023 An ASR-enabled Reading Tutor: Investigating Feedback to Optimize Interaction for Learning to Read
abstract
Contains fulltext : 299814.pdf (Publisher’s version ) (Open Access)
Ferdy Hubers, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH3
2023 Automatic Assessment of Oral Reading Accuracy for Reading Diagnostics
abstract
Contains fulltext : 295425.pdf (Publisher’s version ) (Open Access)
Bo Molenaar, Cristian Tejedor García, Catia Cucchiarini, Helmer Strik
INTERSPEECH3
2023 Assessing Intelligibility in Non-native Speech: Comparing Measures Obtained at Different Levels
abstract
Contains fulltext : 301527.pdf (Publisher’s version ) (Open Access)
Roeland van Hout, Catia Cucchiarini, Danielle Reuvekamp, Helmer Strik
INTERSPEECH3
2023 Measuring the intelligibility of dysarthric speech through automatic speech recognition in a pluricentric language
abstract
Speech intelligibility is an essential though complex construct for evaluating dysarthric speech. Various procedures can be used to measure speech intelligibility, most of which are based on subjective ratings assigned by experts. Since these procedures are subjective and laborious, automatic speech recognition (ASR) has been proposed to obtain objective metrics of intelligibility. Although promising results have been reported, ASR for dysarthric speech generally requires large amounts of data consisting of recorded and annotated speech. In the present study, we explored the possibility of using dysarthric speech resources from the dominant language variety to improve the performance of ASR systems on the dysarthric speech of the non-dominant variety of the same pluricentric language. Dutch is used as an example of a pluricentric language, with Netherlandic Dutch considered the dominant and Flemish Dutch the non-dominant variety. The performance of ASR is evaluated by using two types of intelligibility metrics: orthographic transcriptions and global intelligibility assessments, both obtained from experts. Overall, the results show that dysarthric speech data from the dominant language variety can contribute to improving automatic transcriptions and to developing objective, automatic global measures of speech intelligibility only when no data from the non-dominant variety are available for training ASR models.
Catia Cucchiarini, Roeland van Hout, Helmer Strik
Speech Commun.2
2022 The Effects of Implicit and Explicit Feedback in an ASR-based Reading Tutor for Dutch First-graders
abstract
Contains fulltext : 288876.pdf (Publisher’s version ) (Open Access)
Ferdy Hubers, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH3
2022 Using the LARA Little Prince to compare human and TTS audio quality
abstract
A popular idea in Computer Assisted Language Learning (CALL) is to use multimodal annotated texts, with annotations typically including embedded audio and translations, to support L2 learning through reading. An important question is how to create good quality audio, which can be done either through human recording or by a Text-To-Speech (TTS) engine. We may reasonably expect TTS to be quicker and easier, but human to be of higher quality. Here, we report a study using the open source LARA platform and ten languages. Samples of audio totalling about five minutes, representing the same four passages taken from LARA versions of Saint-Exupèry’s “Le petit prince”, were provided for each language in both human and TTS form; the passages were chosen to instantiate the 2x2 cross product of the conditions dialogue, not-dialogue and humour, not-humour. 251 subjects used a web form to compare human and TTS versions of each item and rate the voices as a whole. For the three languages where TTS did best, English, French and Irish, the evidence from this study and the previous one it extended suggest that TTS audio is now pedagogically adequate and roughly comparable with a non-professional human voice in terms of exemplifying correct pronunciation and prosody. It was however still judged substantially less natural and less pleasant to listen to. No clear evidence was found to support the hypothesis that dialogue and humour pose special problems for TTS. All data and software will be made freely available.
Elham Akhlaghi, Ingibjörg Iðha Auðhunardóttir, Anna Baczkowska, Branislav Bédi, Hakeem Beedar, Harald Berthelsen, Cathy Chua, Catia Cucchiarini, Hanieh Habibi, Ivana Horváthová, Junta Ikeda, Christèle Maizonniaux, Neasa Ní Chiaráin, Chadi Raheb, Manny Rayner, John Sloan, Nikos Tsourakis, Chunlin Yao
LREC8
2022 Multilingual Transfer Learning for Children Automatic Speech Recognition
abstract
Despite recent advances in automatic speech recognition (ASR), the recognition of children’s speech still remains a significant challenge. This is mainly due to the high acoustic variability and the limited amount of available training data. The latter problem is particularly evident in languages other than English, which are usually less-resourced. In the current paper, we address children ASR in a number of less-resourced languages by combining several small-sized children speech corpora from these languages. In particular, we address the following research question: Does a novel two-step training strategy in which multilingual learning is followed by language-specific transfer learning outperform conventional single language/task training for children speech, as well as multilingual and transfer learning alone? Based on previous experimental results with English, we hypothesize that multilingual learning provides a better generalization of the underlying characteristics of children’s speech. Our results provide a positive answer to our research question, by showing that using transfer learning on top of a multilingual model for an unseen language outperforms conventional single language-specific learning.
Thomas Rolland, Alberto Abad, Catia Cucchiarini, Helmer Strik
LREC3
2022 Automatic Speech Recognition and Pronunciation Error Detection of Dutch Non-native Speech: cumulating speech resources in a pluricentric language
abstract
The shortage of large-scale learners’ speech corpora and precise manual annotations are two major challenges for automatic L2 speech recognition and error detection in L2 speech, especially for non-dominant varieties of pluricentric languages. In these cases, collecting and annotating large non-native (L2 learner) corpora for all language varieties is often unattainable. In this study, we investigated ways of addressing these problems through conventional and transfer learning Deep Neural Network (DNN) based Automatic Speech Recognition (ASR) and ASR-based pronunciation error detection (PED) by cumulating Netherlandic Dutch and Flemish Dutch speech resources. First, we show that for ASR the baseline system can be improved by combining the Netherlandic Dutch and Flemish Dutch datasets. Next, through the knowledge learned from models trained on the Netherlandic Dutch data, the Flemish Dutch learners' ASR model can be further improved. In order to evaluate the performance of the PED algorithms in the absence of learner speech data with pronunciation error annotations, we introduced plausible pronunciation errors in the native corpora based on knowledge from Flemish learner speech, in order to simulate non-native speech errors. For PED we found that the results are much better for a GOP classifier trained on Flemish Dutch data than for one trained on Netherlandic Dutch data. PED produced worse results when the Netherlandic Dutch data were merged with the Flemish Dutch data, while for ASR, lower WERs were attained. Whether adding Netherlandic Dutch data to Flemish Dutch data is beneficial, thus seems to depend on the specific task the data are used for. We discuss these results, compare them to those of related research and suggest avenues for future research.
Catia Cucchiarini, Roeland van Hout, Helmer Strik
Speech Commun.2
2021 Speech Intelligibility of Dysarthric Speech: Human Scores and Acoustic-Phonetic Features
abstract
We investigated speech intelligibility in dysarthric and nondysarthric speakers as measured by two commonly used metrics, ratings through the Visual Analogue Scale (VAS) and word accuracy (AcW) through orthographic transcriptions.To gain a better understanding of how acoustic-phonetic correlates could be employed to obtain more objective measures of speech intelligibility and a better classification of dysarthric and non-dysarthric speakers, we studied the relation between these measures of intelligibility and some important acoustic-phonetic correlates.We found that the two intelligibility measures are related, but distinct, and that they might refer to different components of the intelligibility construct.The acoustic-phonetic features showed no difference in the mean values between the two speaker types at the utterance level, but more than half of them played a role in classifying the two speaker types.We computed an acoustic-phonetic probability index (API) at the speaker level.API is moderately correlated to VAS ratings but not correlated to AcW.In addition, API and VAS complement each other in classifying dysarthric and non-dysarthric speakers.This suggests that the intelligibility measures assigned by human raters and acoustic-phonetic features relate to different constructs of intelligibility.
Roeland van Hout, Fleur Boogmans, Mario Ganzeboom, Catia Cucchiarini, Helmer Strik
Interspeech5
2020 ASR-Based Evaluation and Feedback for Individualized Reading Practice
abstract
Contains fulltext : 230238.pdf (Publisher’s version ) (Open Access)
Ferdy Hubers, Catia Cucchiarini, Helmer Strik
INTERSPEECH3
2020 Mobile-Assisted Prosody Training for Limited English Proficiency: Learner Background and Speech Learning Pattern
abstract
Contains fulltext : 228191.pdf (Publisher’s version ) (Open Access)
Kevin Hirschi, Okim Kang, Catia Cucchiarini, John H. L. Hansen, Keelan Evanini, Helmer Strik
INTERSPEECH3
2020 Analyzing Read Aloud Speech by Primary School Pupils: Insights for Research and Development
abstract
Contains fulltext : 228091.pdf (Publisher’s version ) (Open Access)
S. Limonard, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH2
2020 Towards a Comprehensive Assessment of Speech Intelligibility for Pathological Speech
abstract
Contains fulltext : 228265pub.pdf (Publisher’s version ) (Open Access)
Viviana Mendoza Ramos, Wieke Harmsen, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH4
2020 Constructing Multimodal Language Learner Texts Using LARA: Experiences with Nine Languages
abstract
LARA (Learning and Reading Assistant) is an open source platform whose purpose is to support easy conversion of plain texts into multimodal online versions suitable for use by language learners. This involves semi-automatically tagging the text, adding other annotations and recording audio. The platform is suitable for creating texts in multiple languages via crowdsourcing techniques that can be used for teaching a language via reading and listening. We present results of initial experiments by various collaborators where we measure the time required to produce substantial LARA resources, up to the length of short novels, in Dutch, English, Farsi, French, German, Icelandic, Irish, Swedish and Turkish. The first results are encouraging. Although there are some startup problems, the conversion task seems manageable for the languages tested so far. The resulting enriched texts are posted online and are freely available in both source and compiled form.
Elham Akhlaghi, Branislav Bédi, Fatih Bektas, Harald Berthelsen, Matt Butterweck, Cathy Chua, Catia Cucchiarini, Gülsen Eryigit, Johanna Gerlach, Hanieh Habibi, Neasa Ní Chiaráin, Manny Rayner, Steinþór Steingrímsson, Helmer Strik
LREC7
2020 Dedicated Language Resources for Interdisciplinary Research on Multiword Expressions: Best Thing since Sliced Bread
abstract
Multiword expressions such as idioms (beat about the bush), collocations (plastic surgery) and lexical bundles (in the middle of) are challenging for disciplines like Natural Language Processing (NLP), psycholinguistics and second language acquisition, , due to their more or less fixed character. Idiomatic expressions are especially problematic, because they convey a figurative meaning that cannot always be inferred from the literal meanings of the component words. Researchers acknowledge that important properties that characterize idioms such as frequency of exposure, familiarity, transparency, and imageability, should be taken into account in research, but these are typically properties that rely on subjective judgments. This is probably one of the reasons why many studies that investigated idiomatic expressions collected limited information about idiom properties for very small numbers of idioms only. In this paper we report on cross-boundary work aimed at developing a set of tools and language resources that are considered crucial for this kind of multifaceted research. We discuss the results of our research and suggest possible avenues for future research
Ferdy Hubers, Catia Cucchiarini, Helmer Strik
LREC2
2020 BLISS: An Agent for Collecting Spoken Dialogue Data about Health and Well-being
abstract
An important objective in health-technology is the ability to gather information about people’s well-being. Structured interviews can be used to obtain this information, but are time-consuming and not scalable. Questionnaires provide an alternative way to extract such information, though typically lack depth. In this paper, we present our first prototype of the BLISS agent, an artificial intelligent agent which intends to automatically discover what makes people happy and healthy. The goal of Behaviour-based Language-Interactive Speaking Systems (BLISS) is to understand the motivations behind people’s happiness by conducting a personalized spoken dialogue based on a happiness model. We built our first prototype of the model to collect 55 spoken dialogues, in which the BLISS agent asked questions to users about their happiness and well-being. Apart from a description of the BLISS architecture, we also provide details about our dataset, which contains over 120 activities and 100 motivations and is made available for usage.
Jelte van Waterschoot, Iris Hendrickx, Esther Klabbers, Marcel de Korte, Helmer Strik, Catia Cucchiarini, Mariët Theune
LREC7
2017 Multi-Stage DNN Training for Automatic Recognition of Dysarthric Speech
abstract
10.21437/Interspeech.2017-303
Emre Yilmaz 0001, Mario Ganzeboom, Catia Cucchiarini, Helmer Strik
INTERSPEECH3
2016 Intelligibility of Disordered Speech: Global and Detailed Scores
abstract
Contains fulltext : 160745.pdf (Publisher’s version ) (Open Access)
Mario Ganzeboom, Marjoke Bakker, Catia Cucchiarini, Helmer Strik
INTERSPEECH3
2016 Combining Non-Pathological Data of Different Language Varieties to Improve DNN-HMM Performance on Pathological Speech
abstract
Contains fulltext : 160601pub.pdf (Publisher’s version ) (Open Access)
Emre Yilmaz 0001, Mario Ganzeboom, Catia Cucchiarini, Helmer Strik
INTERSPEECH3
2016 Palabras: Crowdsourcing Transcriptions of L2 Speech
Eric Sanders, Pepi Burgos, Catia Cucchiarini, Roeland van Hout
LREC3
2016 A Dutch Dysarthric Speech Database for Individualized Speech Therapy Research
Emre Yilmaz 0001, Mario Ganzeboom, Lilian Beijer, Catia Cucchiarini, Helmer Strik
LREC4
2015 Auris populi: crowdsourced native transcriptions of Dutch vowels spoken by adult Spanish learners
abstract
\n Contains fulltext :\n 145184.pdf (Publisher’s version ) (Open Access)\n
Pepi Burgos, Eric Sanders, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH3
2015 Confusability in L2 vowels: analyzing the role of different features
abstract
Contains fulltext : 150849.pdf (Publisher’s version ) (Open Access)
Matyas Jani, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH2
2014 Dutch vowel production by Spanish learners: duration and spectral features
abstract
In this paper we present a study on Dutch vowel production by Spanish learners that was carried out within the framework of our research on Computer Assisted Pronunciation Training (CAPT). The aim of this study was to obtain detailed information on production of Dutch vowels by Spanish learners, which can be employed to develop effective CAPT programs for this specific target group. We collected speech from learners with varying proficiency levels (A1 - B2 of the CEFR), which was transcribed, segmented and acoustically analyzed. We present data on the frequency of pronunciation errors and on detailed analyses of duration and acoustic properties of the vocalic realizations. The results indicate that Spanish learners of Dutch have difficulties in realizing several Dutch vowel contrasts and that they differ from native speakers in the way they employ duration and spectral properties to realize these contrasts. We discuss these results in relation to those of previous studies on Dutch vowel perception by Spanish listeners and relate them to current theories on speech learning. Index Terms: L2 phonology acquisition, language learning, Computer Assisted Pronunciation Training (CAPT)
Pepi Burgos, Matyas Jani, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH3
2014 ASR-based CALL systems and learner speech data: new resources and opportunities for research and development in second language learning
Catia Cucchiarini, Steve Bodnar, Bart Penning de Vries, Roeland van Hout, Helmer Strik
LREC1
2013 Pronunciation errors by Spanish learners of Dutch: a data-driven study for ASR-based pronunciation training
abstract
In this paper we report on a study on pronunciation errors by Spanish learners of Dutch, which was aimed at obtaining information to develop a dedicated Computer Assisted Pronunciation Training (CAPT) program for this fixed language pair (Spanish L1, Dutch L2).The results of our study indicate, that, first, vowel errors are more frequent and variable than consonant mispronunciations.Second, Spanish natives appear to have problems with vowel length, vowel height, and front rounded vowels.Third, they tend to fall back on the pronunciation of their L1 vowels.
Pepi Burgos, Catia Cucchiarini, Roeland van Hout, Helmer Strik
INTERSPEECH2
2013 L2 syntax acquisition: the effect of oral and written computer assisted practice
abstract
Contains fulltext : 116143.pdf (Publisher’s version ) (Open Access)
Polina Drozdova, Catia Cucchiarini, Helmer Strik
INTERSPEECH2
2012 Practice and feedback in L2 speaking: an evaluation of the DISCO CALL system
abstract
Contains fulltext : 101943.pdf (Publisher’s version ) (Open Access)
Catia Cucchiarini, Joost van Doremalen, Helmer Strik
INTERSPEECH1
2012 The DISCO ASR-based CALL system: practicing L2 oral skills and beyond
Helmer Strik, Jozef Colpaert, Joost van Doremalen, Catia Cucchiarini
LREC4
2011 Computer-assisted Grammar Practice for Oral Communication
Stephen Bodnar, Catia Cucchiarini, Helmer Strik
CSEDU (1)2
2011 Error Selection for ASR-Based English Pronunciation Training in 'My Pronunciation Coach'
abstract
In this paper we report on a study of pronunciation errors that was conducted within the framework of the project "My Pronunciation Coach", which is aimed at developing an ASRbased system for pronunciation training for learners of English with Dutch as their mother tongue.The aim of this study was to obtain quantitative data on the occurrence of pronunciation errors in Dutch English speech.We present the results of this study and compare them to those of previous investigations.Finally, we discuss the implications of these results for the development of My Pronunciation Coach.
Catia Cucchiarini, Henk van den Heuvel, Eric Sanders, Helmer Strik
INTERSPEECH1
2010 Using non-native error patterns to improve pronunciation verification
abstract
In this paper we show how a pronunciation quality measure can be improved by making use of information on frequent pronunciation errors made by non-native speakers. We propose a new measure, called weighted Goodness of Pronunciation (wGOP), and compare it to the much used GOP measure. We applied this measure to the task of discriminating correctly from incorrectly realized Dutch vowels produced by non-native speakers and observed a substantial increase in performance when sufficient training material is available. Index Terms: pronunciation error detection, computer-assisted language learning, confidence measures, weighted GOP
Joost van Doremalen, Catia Cucchiarini, Helmer Strik
INTERSPEECH2
2010 Human Language Technology and Communicative Disabilities: Requirements and Possibilities for the Future
Marina B. Ruiter, Toni C. M. Rietveld, Catia Cucchiarini, Emiel Krahmer, Helmer Strik
LREC3
2009 Automatic detection of vowel pronunciation errors using multiple information sources
abstract
Frequent pronunciation errors made by L2 learners of Dutch often concern vowel substitutions. To detect such pronunciation errors, ASR-based confidence measures (CMs) are generally used. In the current paper we compare and combine confidence measures with MFCCs and phonetic features. The results show that the best results are obtained by using MFCCs, then CMs, and finally phonetic features, and that substantial improvements can be obtained by combining different features.
Joost van Doremalen, Catia Cucchiarini, Helmer Strik
ASRU2
2009 Optimizing non-native speech recognition for CALL applications
abstract
We are developing a Computer Assisted Language Learning (CALL) system that gives feedback to grammar and pronunciation that makes use of Automatic Speech Recognition (ASR). However, good quality unconstrained non-native ASR is not yet feasible. Therefore, we use an approach in which we try to elicit constrained responses. The task in the current experiments is to select utterances from a list of responses. The results of our experiments show that significant improvements can be obtained by optimizing the language model and acoustic models. In this way we could reduce the utterance error rate from 29-26 % to
Joost van Doremalen, Helmer Strik, Catia Cucchiarini
INTERSPEECH3
2009 Oral proficiency training in Dutch L2: The contribution of ASR-based corrective feedback
Catia Cucchiarini, Ambra Neri, Helmer Strik
Speech Commun.1
2009 Comparing different approaches for automatic pronunciation error detection
Helmer Strik, Khiet P. Truong, Febe de Wet, Catia Cucchiarini
Speech Commun.4
2008 DISCO: development and integration of speech technology into courseware for language learning
abstract
Recent research has shown that a properly designed ASR-based CALL system (Dutch-CAPT) was capable of detecting pronunciation errors and of providing comprehensible feedback on pronunciation. Since pronunciation is not the only skill required for speaking a second language, we explored the possibility of extending the Dutch-CAPT approach to other aspects of speaking proficiency like morphology and syntax. In this paper we explain how a number of errors in morphology and syntax that are common in spoken Dutch L2 could be addressed in an ASR-based CALL system. Finally, we present our new project in which corrective feedback will be provided on all three aspects of spoken proficiency: pronunciation, morphology and syntax. Index Terms: pronunciation training, CALL, ASR, error detection.
Catia Cucchiarini, Joost van Doremalen, Helmer Strik
INTERSPEECH1
2008 Pronunciation reduction: how it relates to speech style, gender, and age
abstract
\n Contains fulltext :\n 68373.pdf (author's version ) (Open Access)\n
Helmer Strik, Joost van Doremalen, Catia Cucchiarini
INTERSPEECH3
2008 Recording Speech of Children, Non-Natives and Elderly People for HLT Applications: the JASMIN-CGN Corpus
Catia Cucchiarini, Joris Driesen, Hugo Van hamme, Eric Sanders
LREC1
2008 The Dutch-Flemish Comprehensive Approach to HLT Stimulation and Innovation: STEVIN, HLT Agency and beyond
Peter Spyns, Elisabeth D'Halleweyn, Catia Cucchiarini
LREC3
2007 ASR-based pronunciation training: scoring accuracy and pedagogical effectiveness of a system for dutch L2 learners
abstract
A system for providing Computer Assisted Pronunciation Training for Dutch was developed, Dutch-CAPT, which appeared to be effective in improving pronunciation quality of L2 learners of Dutch.In this paper we describe the architecture of the system paying particular attention to the rationale behind this system, to the performance of the error detection algorithm and its relationship to the pedagogical effectiveness of the corrective feedback provided Index Terms: Computer Assisted Pronunciation Training (CAPT), corrective feedback, pronunciation error detection, Goodness Of Pronunciation (GOP)
Catia Cucchiarini, Ambra Neri, Febe de Wet, Helmer Strik
INTERSPEECH1
2007 Comparing classifiers for pronunciation error detection
abstract
CITATION: Strik, H. et al. 2007. Comparing classifiers for pronunciation error detection. In Hamme, H. van; Son, R. van (ed.), Proceedings of Interspeech 2007, pp. 1837-1840.
Helmer Strik, Khiet P. Truong, Febe de Wet, Catia Cucchiarini
INTERSPEECH4
2006 ASR-based corrective feedback on pronunciation: does it really work?
abstract
We studied a group of immigrants who were following regular, teacher-fronted Dutch classes, and who were assigned to three groups using either a) Dutch CAPT, an ASR-based Computer Assisted Pronunciation Training (CAPT) system that provides feedback on a number of Dutch speech sounds that are problematic for L2 learners b) a CAPT system without feedback c) no CAPT system. Participants were tested before and after the training. The results show that the ASR-based feedback was effective in correcting the errors addressed in the training. 1.
Ambra Neri, Catia Cucchiarini, Helmer Strik
INTERSPEECH2
2006 JASMIN-CGN: Extension of the Spoken Dutch Corpus with Speech of Elderly People, Children and Non-natives in the Human-Machine Interaction Modality
Catia Cucchiarini, Hugo Van hamme, Olga van Herwijnen, Felix Smits
LREC1
2006 The Dutch-Flemish HLT Programme STEVIN: Essential Speech and Language Technology Resources
Elisabeth D'Halleweyn, Jan Odijk, Lisanne Teunissen, Catia Cucchiarini
LREC4
2005 Multiword expressions in spontaneous speech: do we really speak like that?
abstract
In this study, we examined the pronunciation characteristics of multiword expressions (MWEs). We first drew up an inventory of frequently occurring N-grams extracted from orthographic transcriptions of spontaneous speech contained in a large corpus of spoken Dutch. For about 10 % of these Ngrams phonetic transcriptions were available, which were examined. Our results show that the pronunciation of these Ngrams differed to a large extent from the canonical form. In order to determine whether this is a general characteristic of spontaneous speech or rather the effect of the specific status of these N-grams, we analyzed the pronunciations of the individual words composing the N-grams in two context conditions: 1) in the N-gram context and 2) in any other context. We found that words in N-grams do indeed have peculiar pronunciation patterns. This seems to suggest that these N-grams may be considered as MWEs that should therefore be treated as lexical entries with their own specific pronunciation variants in the pronunciation lexicons used for e.g. automatic speech recognition (ASR) and automatic phonetic transcription (APT). 1.
Helmer Strik, Diana Binnenpoorte, Catia Cucchiarini
INTERSPEECH3
2005 Automatic detection of frequent pronunciation errors made by L2-learners
abstract
Contains fulltext : 41035.pdf (Publisher’s version ) (Open Access)
Khiet P. Truong, Ambra Neri, Febe de Wet, Catia Cucchiarini, Helmer Strik
INTERSPEECH4
2005 Multiword expressions in spoken language: An exploratory study on pronunciation variation
Diana Binnenpoorte, Catia Cucchiarini, Lou Boves, Helmer Strik
Comput. Speech Lang.2
2004 Improving Automatic Phonetic Transcription of Spontaneous Speech Through Variant-Based Pronunciation Variation Modelling
Diana Binnenpoorte, Catia Cucchiarini, Helmer Strik, Lou Boves
LREC2
2004 The New Dutch-Flemish HLT Programme: a Concerted Effort to Stimulate the HLT Sector
Catia Cucchiarini, Elisabeth D'Halleweyn
LREC1
2003 A data-driven method for modeling pronunciation variation
Judith M. Kessens, Catia Cucchiarini, Helmer Strik
Speech Commun.2
2002 Validation and improvement of automatic phonetic transcriptions
abstract
\n Contains fulltext :\n 76483.pdf (author's version ) (Open Access)\n
Catia Cucchiarini, Diana Binnenpoorte
INTERSPEECH1
2002 Feedback in computer assisted pronunciation training: technology push or demand pull?
abstract
In this paper, we examine the type of feedback that currently available Computer Assisted Pronunciation Training (CAPT) systems provide, with a view to establishing whether this meets pedagogically sound requirements.W e show that many commercial systems tend to prefer technological novelties that do not always comply with pedagogical criteria and that despite the limitations of today's technology, it is possible to design CAPT systems that are more in line with learners' needs.
Ambra Neri, Catia Cucchiarini, Helmer Strik
INTERSPEECH2
2002 Dutch HLT resources: from BLARK to priority lists
abstract
\n Contains fulltext :\n 76208.pdf (author's version ) (Open Access)\n
Helmer Strik, Walter Daelemans, Diana Binnenpoorte, Janienke Sturm, Folkert de Vriend, Catia Cucchiarini
INTERSPEECH6
2002 A Field Survey for Establishing Priorities in the Development of HLT Resources for Dutch
Diana Binnenpoorte, Folkert de Vriend, Janienke Sturm, Walter Daelemans, Helmer Strik, Catia Cucchiarini
LREC6
2002 A Human Language Technologies Platform for the Dutch language: awareness, management maintenance and distribution
Catia Cucchiarini, Elisabeth D'Halleweyn, Lisanne Teunissen
LREC1
2001 Phonetic transcriptions in the spoken dutch corpus: how to combine efficiency and good transcription quality
abstract
The following full text is an author's version which may differ from the publisher's version.
Catia Cucchiarini, Diana Binnenpoorte, Simo M. A. Goddijn
INTERSPEECH1
2001 Comparing the performance of two CSRs: how to determine the significance level of the differences
abstract
When two CSRs are compared, it is important to test what the significance level of the difference is. For this purpose a metric and a statistical test are needed. In this paper we compare several combinations of a metric with a statistical test, in order to find a combination which is suitable for this task. Four combinations which are introduced in this paper appear to be suitable for this task.
Helmer Strik, Catia Cucchiarini, Judith M. Kessens
INTERSPEECH2
2000 A bottom-up method for obtaining information about pronunciation variation
abstract
\n Contains fulltext :\n 76197.pdf (author's version ) (Open Access)\n
Judith M. Kessens, Helmer Strik, Catia Cucchiarini
INTERSPEECH3
2000 L2 pronunciation quality in read and spontaneous speech
abstract
This paper describes two experiments aimed at exploring the relationship between objective properties of speech and perceived pronunciation quality in read and spontaneous speech, with a view to determining whether such quantitative measures can be used to develop objective pronunciation tests. Read and spontaneous speech of two groups of 60 learners of Dutch as a second language was scored for pronunciation quality by human raters and was analyzed by means of a continuous speech recognizer to calculate six quantitative measures of speech quality related to speech timing. The results show that quantitative, temporal measures of speech are strongly related to pronunciation quality, in both read and spontaneous speech, although not all variables suitable for measuring pronunciation quality in read speech are as effective in spontaneous speech. 1.
Helmer Strik, Catia Cucchiarini, Diana Binnenpoorte
INTERSPEECH2
2000 Comparing the recognition performance of CSRs: in search of an adequate metric and statistical significance test
abstract
In this paper a new measure of recognition accuracy is introduced which can be used when comparing the performance of two speech recognizers, to establish which is the better one. This metric combines the advantages of previous measures, but excludes their disadvantages. Essentially, the metric is an attempt to quantify the degree of recognition accuracy for each sentence, thus obtaining a more informative measure than either correct or incorrect, in such a way that the statistical significance of the observed differences can be tested. The advantages of our assessment method are illustrated on the basis of both artificial and real performance data of different recognizers. 1.
Helmer Strik, Catia Cucchiarini, Judith M. Kessens
INTERSPEECH2
2000 NL-Translex: Machine Translation for Dutch
Catia Cucchiarini, Johan Van Hoorde, Elisabeth D'Halleweyn
LREC1
2000 Different aspects of expert pronunciation quality ratings and their relation to scores produced by speech recognition algorithms
Catia Cucchiarini, Helmer Strik, Lou Boves
Speech Commun.1
2000 Erratum to: "Different aspects of expert pronunciation quality ratings and their relation to scores produced by speech recognition algorithms": [Speech Communication 30 (2000) 109-119]
Catia Cucchiarini, Helmer Strik, Lou Boves
Speech Commun.1
1999 Using likelihood ratios to perform utterance verification in automatic pronunciation assessment
abstract
The aim of our current research is to investigate the possibility of using likelihood ratios to perform utterance verification within the context of automatic oral proficiency assessment. The likelihood ratios under investigation have the appealing feature that they may be computed simply by using an off-theshelf automatic speech recognition system in two different recognition modes (forced and free phone) instead of using a system with specifically trained anti-models. We achieved 93% correct classification for 10 phonetically rich sentences uttered by 60 non-native language students. 1. INTRODUCTION The long-term goal of our research is to employ ASR technology in an automatic pronunciation test for Dutch as a second language. As a consequence of this aim we are not concerned with learners of Dutch with a specific mother tongue, but rather with a group of speakers who are highly varied in this respect. In this sense our situation is different from that of many studies on the use of...
Febe de Wet, Catia Cucchiarini, Helmer Strik, Lou Boves
EUROSPEECH2
1999 Modeling pronunciation variation for ASR: A survey of the literature
Helmer Strik, Catia Cucchiarini
Speech Commun.2
1998 Postvocalic /r/-deletion in standard dutch: how experimental phonology can profit from ASR technology
abstract
\n Contains fulltext :\n 76425.pdf ( ) (Open Access)\n
Catia Cucchiarini, Henk van den Heuvel
ICSLP1
1998 Quantitative assessment of second language learners' fluency: an automatic approach
abstract
\n Contains fulltext :\n 74998.pdf (author's version ) (Open Access)\n
Catia Cucchiarini, Helmer Strik, Lou Boves
ICSLP1
1998 Assessment of dutch pronunciation by means of automatic speech recognition technology
abstract
\n Contains fulltext :\n 75007.pdf (author's version ) (Open Access)\n
Catia Cucchiarini, Febe de Wet, Helmer Strik, Lou Boves
ICSLP1
1998 The selection of pronunciation variants: comparing the performance of man and machine
abstract
Dans cet article, les performances d'un outil de transcription automatique sont évaluées.L'outil de transcription est un reconnaisseur de parole continue (CSR) fonctionnant en mode de reconnaissance forcée.Pour l'évaluation les performances du CSR ont été comparées à celles de neuf auditeurs experts.La machine et l'humain ont effectué exactement la même tâche: décider si un segment était présent ou non dans 467 cas.Il s'est avéré que les performances du CSR étaient comparables à celle des experts.
Judith M. Kessens, Mirjam Wester, Catia Cucchiarini, Helmer Strik
ICSLP3
1997 Automatic assessment of foreign speakers' pronunciation of dutch
abstract
Item does not contain fulltext
Catia Cucchiarini, Lou Boves
EUROSPEECH1
1996 Localizing an automatic inquiry system for public transport information
abstract
This paper reports on the development o f a spoken dialogue system for providing information about public transport in the Netherlands.It is explained how a German prototype was adapted for Dutch.Emphasis is laid on the specific approach chosen to collect speech material that could be used to gradually improve the system.The pros and cons of this method are discussed.
Helmer Strik, Albert Russel, Henk van den Heuvel, Catia Cucchiarini, Lou Boves
ICSLP4
1992 Familiarity with the language transcribed and context as determinants of intratranscriber agreement
abstract
Item does not contain fulltext
Catia Cucchiarini, Renée van Bezooijen
ICSLP1