VLDB 2026 Research / reviewers in the wild / expert
Helmer Strik
dblp:68/1093
· DBLP profile ↗
113ranked-venue papers
20as first author
25since 2021 · last 2026
0000-0003-1722-3465ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 96 · 16 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 92 · 19 first-author · 20 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech AssessmentabstractContains fulltext : 331543.pdf (Publisher’s version ) (Closed access) Aditya Parikh, Cristian Tejedor García, Catia Cucchiarini, Helmer Strik |
LREC | 4 |
| 2026 | Predicting accentedness and comprehensibility through ASR scores and acoustic featuresabstract• Compared two automatic speech recognition models (TDNN and Whisper) for predicting accentedness and comprehensibility. • Used a data-driven method to select the most relevant acoustic features for accentedness and comprehensibility. • Used a linear mixed-effects model to incorporate speaker and utterance differences. • Combined segmental and suprasegmental features to better understand accentedness and comprehensibility of non-native speech. Accentedness and comprehensibility scales are widely used in measuring the oral proficiency of second language (L2) learners, including learners of English as a Second Language (ESL). In this paper, we focus on gaining a better understanding of the concepts of accentedness and comprehensibility by developing and applying automatic measures to ESL utterances produced by Indonesian learners. We extracted features both on the segmental and the suprasegmental (fundamental frequency, loudness, energy et al.) levels to investigate which features are actually related to expert judgments on accentedness and comprehensibility. Automatic Speech Recognition (ASR) pronunciation scores based on the traditional Kaldi Time Delay Neural Network (TDNN) model and on the End-to-End Whisper model were applied, and data-driven methods were used by combining acoustic features extracted by the Geneva Minimalistic Acoustic Parameter Set (eGeMAPS) and Praat. The experimental results showed that Whisper outperformed the Kaldi-TDNN model. The Whisper model gave the best results for predicting comprehensibility on the basis of phone distance, and the best results for predicting accentedness on the basis of grapheme distance. Combining segmental and suprasegmental features improved the results, yielding different feature rankings for comprehensibility and accentedness. In our final step of analysis, we included differences between utterances and learners as random effects in a mixed linear regression model. Exploiting these information sources yielded a substantial improvement in predicting both comprehensibility and accentedness. Wenwei Dong, Catia Cucchiarini, Roeland van Hout, Helmer Strik |
Comput. Speech Lang. | 4 |
| 2025 | Evaluating Progress of CALL System Users on Accentedness and Comprehensibility: An Acoustic and ASR-Based ApproachabstractContains fulltext : 322930.pdf (Publisher’s version ) (Open Access) Wenwei Dong, Catia Cucchiarini, Roeland van Hout, Helmer Strik |
INTERSPEECH | 4 |
| 2025 | Multitalker Babble in English Vowel Perception Training: A Comparison between Humans and Neural ModelsabstractContains fulltext : 322924.pdf (Publisher’s version ) (Open Access) Wenwei Dong, Alif Silpachai, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 4 |
| 2025 | Improving Child Speech Recognition and Reading Mistake Detection by Using PromptsabstractAutomatic reading aloud evaluation can provide valuable support to teachers by enabling more efficient scoring of reading exercises. However, research on reading evaluation systems and applications remains limited. We present a novel multimodal approach that leverages audio and knowledge from text resources. In particular, we explored the potential of using Whisper and instruction-tuned large language models (LLMs) with prompts to improve transcriptions for child speech recognition, as well as their effectiveness in downstream reading mistake detection. Our results demonstrate the effectiveness of prompting Whisper and prompting LLM, compared to the baseline Whisper model without prompting. The best performing system achieved state-of-the-art recognition performance in Dutch child read speech, with a word error rate (WER) of 5.1%, improving the baseline WER of 9.4%. Furthermore, it significantly improved reading mistake detection, increasing the F1 score from 0.39 to 0.73. Lingyun Gao, Cristian Tejedor García, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 4 |
| 2025 | Can ASR generate valid measures of child reading fluency?
Wieke Harmsen, Roeland van Hout, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 4 |
| 2025 | Evaluating Logit-Based GOP Scores for Mispronunciation DetectionabstractPronunciation assessment relies on goodness of pronunciation (GOP) scores, traditionally derived from softmax-based posterior probabilities. However, posterior probabilities may suffer from overconfidence and poor phoneme separation, limiting their effectiveness. This study compares logit-based GOP scores with probability-based GOP scores for mispronunciation detection. We conducted our experiment on two L2 English speech datasets spoken by Dutch and Mandarin speakers, assessing classification performance and correlation with human ratings. Logit-based methods outperform probability-based GOP in classification, but their effectiveness depends on dataset characteristics. The maximum logit GOP shows the strongest alignment with human perception, while a combination of different GOP scores balances probability and logit features. The findings suggest that hybrid GOP methods incorporating uncertainty modeling and phoneme-specific weighting improve pronunciation assessment. Aditya Parikh, Cristian Tejedor García, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 4 |
| 2025 | Enhancing GOP in CTC-Based Mispronunciation Detection with Phonological KnowledgeabstractComputer-Assisted Pronunciation Training (CAPT) systems employ automatic measures of pronunciation quality, such as the goodness of pronunciation (GOP) metric. GOP relies on forced alignments, which are prone to labeling and segmentation errors due to acoustic variability. While alignment-free methods address these challenges, they are computationally expensive and scale poorly with phoneme sequence length and inventory size. To enhance efficiency, we introduce a substitution-aware alignment-free GOP that restricts phoneme substitutions based on phoneme clusters and common learner errors. We evaluated our GOP on two L2 English speech datasets, one with child speech, My Pronunciation Coach (MPC), and SpeechOcean762, which includes child and adult speech. We compared RPS (restricted phoneme substitutions) and UPS (unrestricted phoneme substitutions) setups within alignment-free methods, which outperformed the baseline. We discuss our results and outline avenues for future research. Aditya Parikh, Cristian Tejedor García, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 4 |
| 2024 | Reading Miscue Detection in Primary School through Automatic Speech RecognitionabstractContains fulltext : 309809.pdf (Publisher’s version ) (Open Access) Lingyun Gao, Cristian Tejedor García, Helmer Strik, Catia Cucchiarini |
INTERSPEECH | 3 |
| 2023 | An ASR-enabled Reading Tutor: Investigating Feedback to Optimize Interaction for Learning to ReadabstractContains fulltext : 299814.pdf (Publisher’s version ) (Open Access) Ferdy Hubers, Catia Cucchiarini, Roeland van Hout, Helmer Strik |
INTERSPEECH | 5 |
| 2023 | Automatic assessments of dysarthric speech: the usability of acoustic-phonetic features
Loes van Bemmel, Chiara Pesenti, Helmer Strik |
INTERSPEECH | 4 |
| 2023 | Alzheimer Disease Classification through ASR-based Transcriptions: Exploring the Impact of Punctuation and PausesabstractContains fulltext : 295424.pdf (Publisher’s version ) (Open Access) Lucía Gómez-Zaragozá, Simone Wills, Cristian Tejedor García, Javier Marín-Morales, Mariano Alcañiz Raya, Helmer Strik |
INTERSPEECH | 6 |
| 2023 | Automatic Assessment of Oral Reading Accuracy for Reading DiagnosticsabstractContains fulltext : 295425.pdf (Publisher’s version ) (Open Access) Bo Molenaar, Cristian Tejedor García, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 4 |
| 2023 | Assessing Intelligibility in Non-native Speech: Comparing Measures Obtained at Different LevelsabstractContains fulltext : 301527.pdf (Publisher’s version ) (Open Access) Roeland van Hout, Catia Cucchiarini, Danielle Reuvekamp, Helmer Strik |
INTERSPEECH | 5 |
| 2023 | Measuring the intelligibility of dysarthric speech through automatic speech recognition in a pluricentric languageabstractSpeech intelligibility is an essential though complex construct for evaluating dysarthric speech. Various procedures can be used to measure speech intelligibility, most of which are based on subjective ratings assigned by experts. Since these procedures are subjective and laborious, automatic speech recognition (ASR) has been proposed to obtain objective metrics of intelligibility. Although promising results have been reported, ASR for dysarthric speech generally requires large amounts of data consisting of recorded and annotated speech. In the present study, we explored the possibility of using dysarthric speech resources from the dominant language variety to improve the performance of ASR systems on the dysarthric speech of the non-dominant variety of the same pluricentric language. Dutch is used as an example of a pluricentric language, with Netherlandic Dutch considered the dominant and Flemish Dutch the non-dominant variety. The performance of ASR is evaluated by using two types of intelligibility metrics: orthographic transcriptions and global intelligibility assessments, both obtained from experts. Overall, the results show that dysarthric speech data from the dominant language variety can contribute to improving automatic transcriptions and to developing objective, automatic global measures of speech intelligibility only when no data from the non-dominant variety are available for training ASR models. Catia Cucchiarini, Roeland van Hout, Helmer Strik |
Speech Commun. | 4 |
| 2022 | Detection of COPD Exacerbation from Speech: Comparison of Acoustic Features and Deep Learning Based Speech Breathing ModelsabstractRespiration is a primary process involved in speech production. We can often hear if a person has respiratory difficulty, thus making speech a good pathological indicator for respiratory conditions. This is more relevant to conditions like chronic obstructive pulmonary disease (COPD). Patients with COPD suffer from voice changes with respect to the healthy population. Medical professionals observe that the speech of COPD patients during stable periods differs from the speech during exacerbation. In this paper, we investigate this detection of COPD exacerbation from speech in three approaches: acoustic features identification using a statistical approach, low-level descriptive features with classification, and speech breathing models based on deep learning architectures to estimate the patients’ breathing rate. Our analysis indicates that each of these approaches indeed results in a clear distinction of speech during exacerbation and stable periods of COPD. Venkata Srikanth Nallanthighal, Aki Härmä, Helmer Strik |
ICASSP | 3 |
| 2022 | The Effects of Implicit and Explicit Feedback in an ASR-based Reading Tutor for Dutch First-gradersabstractContains fulltext : 288876.pdf (Publisher’s version ) (Open Access) Ferdy Hubers, Catia Cucchiarini, Roeland van Hout, Helmer Strik |
INTERSPEECH | 5 |
| 2022 | Improved ASR Performance for Dysarthric Speech Using Two-stage DataAugmentationabstractContains fulltext : 296885.pdf (Publisher’s version ) (Open Access) Chitralekha Bhat, Ashish Panda, Helmer Strik |
INTERSPEECH | 3 |
| 2022 | COVID-19 detection based on respiratory sensing from speechabstractCOVID-19 affects a person's respiratory health, which is manifested in the form of shortness of breath during speech. Recent work shows that it is possible to use deep learning techniques to sense the speaker's respiratory parameters from a speech signal directly. Thus respiratory parameters like speech breathing rate and tidal volume can be computed and compared using deep learning techniques to detect COVID-19 from speech recordings. In this paper, we compute respiratory parameters using our pre-trained deep learning-based speech breathing models and use them for detecting COVID-19 from speech. Apart from using speech breathing models, we perform acoustic features identification using a statistical approach and classification based on low-level descriptive features. Our analysis investigates the distinction of speech of a healthy person and COVID-19 affected person. Venkata Srikanth Nallanthighal, Aki Härmä, Helmer Strik |
INTERSPEECH | 3 |
| 2022 | Multilingual Transfer Learning for Children Automatic Speech RecognitionabstractDespite recent advances in automatic speech recognition (ASR), the recognition of children’s speech still remains a significant challenge. This is mainly due to the high acoustic variability and the limited amount of available training data. The latter problem is particularly evident in languages other than English, which are usually less-resourced. In the current paper, we address children ASR in a number of less-resourced languages by combining several small-sized children speech corpora from these languages. In particular, we address the following research question: Does a novel two-step training strategy in which multilingual learning is followed by language-specific transfer learning outperform conventional single language/task training for children speech, as well as multilingual and transfer learning alone? Based on previous experimental results with English, we hypothesize that multilingual learning provides a better generalization of the underlying characteristics of children’s speech. Our results provide a positive answer to our research question, by showing that using transfer learning on top of a multilingual model for an unseen language outperforms conventional single language-specific learning. Thomas Rolland, Alberto Abad, Catia Cucchiarini, Helmer Strik |
LREC | 4 |
| 2022 | Negation Detection in Dutch Spoken Human-Computer ConversationsabstractProper recognition and interpretation of negation signals in text or communication is crucial for any form of full natural language understanding. It is also essential for computational approaches to natural language processing. In this study we focus on negation detection in Dutch spoken human-computer conversations. Since there exists no Dutch (dialogue) corpus annotated for negation we have annotated a Dutch corpus sample to evaluate our method for automatic negation detection. We use transfer learning and trained NegBERT (an existing BERT implementation used for negation detection) on English data with multilingual BERT to detect negation in Dutch dialogues. Our results show that adding in-domain training material improves the results. We show that we can detect both negation cues and scope in Dutch dialogues with high precision and recall. We provide a detailed error analysis and discuss the effects of cross-lingual and cross-domain transfer learning on automatic negation detection. Tom Sweers, Iris Hendrickx, Helmer Strik |
LREC | 3 |
| 2022 | Automatic Speech Recognition and Pronunciation Error Detection of Dutch Non-native Speech: cumulating speech resources in a pluricentric languageabstractThe shortage of large-scale learners’ speech corpora and precise manual annotations are two major challenges for automatic L2 speech recognition and error detection in L2 speech, especially for non-dominant varieties of pluricentric languages. In these cases, collecting and annotating large non-native (L2 learner) corpora for all language varieties is often unattainable. In this study, we investigated ways of addressing these problems through conventional and transfer learning Deep Neural Network (DNN) based Automatic Speech Recognition (ASR) and ASR-based pronunciation error detection (PED) by cumulating Netherlandic Dutch and Flemish Dutch speech resources. First, we show that for ASR the baseline system can be improved by combining the Netherlandic Dutch and Flemish Dutch datasets. Next, through the knowledge learned from models trained on the Netherlandic Dutch data, the Flemish Dutch learners' ASR model can be further improved. In order to evaluate the performance of the PED algorithms in the absence of learner speech data with pronunciation error annotations, we introduced plausible pronunciation errors in the native corpora based on knowledge from Flemish learner speech, in order to simulate non-native speech errors. For PED we found that the results are much better for a GOP classifier trained on Flemish Dutch data than for one trained on Netherlandic Dutch data. PED produced worse results when the Netherlandic Dutch data were merged with the Flemish Dutch data, while for ASR, lower WERs were attained. Whether adding Netherlandic Dutch data to Flemish Dutch data is beneficial, thus seems to depend on the specific task the data are used for. We discuss these results, compare them to those of related research and suggest avenues for future research. Catia Cucchiarini, Roeland van Hout, Helmer Strik |
Speech Commun. | 4 |
| 2021 | On The Relationship Between Speech-Based Breathing Signal Prediction Evaluation Measures and Breathing Parameters EstimationabstractThe respiratory system is one of the major components of the speech production system. Any alteration in breathing can result in changes in speech. Specific breathing characteristics, such as breathing rate and tidal volume, can indicate a person’s pathological condition. More recently, neural network-based methods have started emerging for predicting the breathing signal from the speech signal. The neural networks are trained and evaluated with different objective measures, such as mean squared error (MSE) and Pearson’s correlation. This paper investigates whether there is a systematic relationship between the different objective measures used for training and evaluating the neural network models and the end-goal, i.e. estimation of breathing parameters such as, breathing rate and tidal volume. Our investigations on two different data sets with two different neural network-based approaches show that there is no clear systematic relationship. In other words, obtaining a high Pearson’s correlation on the evaluation set does not necessarily mean better breathing parameter estimation. Thus, indicating the need for developing other objective evaluation measures. Zohreh Mostaani, Venkata Srikanth Nallanthighal, Aki Härmä, Helmer Strik, Mathew Magimai-Doss |
ICASSP | 4 |
| 2021 | Speech Intelligibility of Dysarthric Speech: Human Scores and Acoustic-Phonetic FeaturesabstractWe investigated speech intelligibility in dysarthric and nondysarthric speakers as measured by two commonly used metrics, ratings through the Visual Analogue Scale (VAS) and word accuracy (AcW) through orthographic transcriptions.To gain a better understanding of how acoustic-phonetic correlates could be employed to obtain more objective measures of speech intelligibility and a better classification of dysarthric and non-dysarthric speakers, we studied the relation between these measures of intelligibility and some important acoustic-phonetic correlates.We found that the two intelligibility measures are related, but distinct, and that they might refer to different components of the intelligibility construct.The acoustic-phonetic features showed no difference in the mean values between the two speaker types at the utterance level, but more than half of them played a role in classifying the two speaker types.We computed an acoustic-phonetic probability index (API) at the speaker level.API is moderately correlated to VAS ratings but not correlated to AcW.In addition, API and VAS complement each other in classifying dysarthric and non-dysarthric speakers.This suggests that the intelligibility measures assigned by human raters and acoustic-phonetic features relate to different constructs of intelligibility. Roeland van Hout, Fleur Boogmans, Mario Ganzeboom, Catia Cucchiarini, Helmer Strik |
Interspeech | 6 |
| 2021 | Deep learning architectures for estimating breathing signal and respiratory parameters from speech recordingsabstractRespiration is an essential and primary mechanism for speech production. We first inhale and then produce speech while exhaling. When we run out of breath, we stop speaking and inhale. Though this process is involuntary, speech production involves a systematic outflow of air during exhalation characterized by linguistic content and prosodic factors of the utterance. Thus speech and respiration are closely related, and modeling this relationship makes sensing respiratory dynamics directly from the speech plausible, however is not well explored. In this article, we conduct a comprehensive study to explore techniques for sensing breathing signal and breathing parameters from speech using deep learning architectures and address the challenges involved in establishing the practical purpose of this technology. Estimating the breathing pattern from the speech would give us information about the respiratory parameters, thus enabling us to understand the respiratory health using one's speech. Venkata Srikanth Nallanthighal, Zohreh Mostaani, Aki Härmä, Helmer Strik, Mathew Magimai-Doss |
Neural Networks | 4 |
| 2020 | Speech Breathing Estimation Using Deep Learning MethodsabstractBreathing is the primary mechanism for maintaining the subglottal pressure for speech production. Speech can be seen as a systematic outflow of air during exhalation characterized by linguistic content and prosodic factors. Thus, sensing respiratory dynamics from the speech is plausible. In this paper, we explore techniques for sensing breathing from speech using deep learning architectures including multi-task learning approaches. Estimating the breathing pattern from the speech would give us information about the respiration rate, breathing capacity and thus enable us to understand the pathological condition of a person using one's speech. Training and evaluation of our model on our database of breathing signal and speech for 40 subjects yielded a sensitivity of 0.88 for breath event detection and 5.6 % error for breathing rate estimation. Venkata Srikanth Nallanthighal, Aki Härmä, Helmer Strik |
ICASSP | 3 |
| 2020 | ASR-Based Evaluation and Feedback for Individualized Reading PracticeabstractContains fulltext : 230238.pdf (Publisher’s version ) (Open Access) Ferdy Hubers, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 4 |
| 2020 | A Comparison of Acoustic and Linguistics Methodologies for Alzheimer's Dementia RecognitionabstractContains fulltext : 228158.pdf (Publisher’s version ) (Open Access) Nicholas Cummins, Yilin Pan, Zhao Ren, Julian Fritsch, Venkata Srikanth Nallanthighal, Heidi Christensen, Daniel Blackburn, Björn W. Schuller, Mathew Magimai-Doss, Helmer Strik, Aki Härmä |
INTERSPEECH | 10 |
| 2020 | Mobile-Assisted Prosody Training for Limited English Proficiency: Learner Background and Speech Learning PatternabstractContains fulltext : 228191.pdf (Publisher’s version ) (Open Access) Kevin Hirschi, Okim Kang, Catia Cucchiarini, John H. L. Hansen, Keelan Evanini, Helmer Strik |
INTERSPEECH | 6 |
| 2020 | Analyzing Read Aloud Speech by Primary School Pupils: Insights for Research and DevelopmentabstractContains fulltext : 228091.pdf (Publisher’s version ) (Open Access) S. Limonard, Catia Cucchiarini, Roeland van Hout, Helmer Strik |
INTERSPEECH | 4 |
| 2020 | Towards a Comprehensive Assessment of Speech Intelligibility for Pathological SpeechabstractContains fulltext : 228265pub.pdf (Publisher’s version ) (Open Access) Viviana Mendoza Ramos, Wieke Harmsen, Catia Cucchiarini, Roeland van Hout, Helmer Strik |
INTERSPEECH | 6 |
| 2020 | Constructing Multimodal Language Learner Texts Using LARA: Experiences with Nine LanguagesabstractLARA (Learning and Reading Assistant) is an open source platform whose purpose is to support easy conversion of plain texts into multimodal online versions suitable for use by language learners. This involves semi-automatically tagging the text, adding other annotations and recording audio. The platform is suitable for creating texts in multiple languages via crowdsourcing techniques that can be used for teaching a language via reading and listening. We present results of initial experiments by various collaborators where we measure the time required to produce substantial LARA resources, up to the length of short novels, in Dutch, English, Farsi, French, German, Icelandic, Irish, Swedish and Turkish. The first results are encouraging. Although there are some startup problems, the conversion task seems manageable for the languages tested so far. The resulting enriched texts are posted online and are freely available in both source and compiled form. Elham Akhlaghi, Branislav Bédi, Fatih Bektas, Harald Berthelsen, Matt Butterweck, Cathy Chua, Catia Cucchiarini, Gülsen Eryigit, Johanna Gerlach, Hanieh Habibi, Neasa Ní Chiaráin, Manny Rayner, Steinþór Steingrímsson, Helmer Strik |
LREC | 14 |
| 2020 | Dedicated Language Resources for Interdisciplinary Research on Multiword Expressions: Best Thing since Sliced BreadabstractMultiword expressions such as idioms (beat about the bush), collocations (plastic surgery) and lexical bundles (in the middle of) are challenging for disciplines like Natural Language Processing (NLP), psycholinguistics and second language acquisition, , due to their more or less fixed character. Idiomatic expressions are especially problematic, because they convey a figurative meaning that cannot always be inferred from the literal meanings of the component words. Researchers acknowledge that important properties that characterize idioms such as frequency of exposure, familiarity, transparency, and imageability, should be taken into account in research, but these are typically properties that rely on subjective judgments. This is probably one of the reasons why many studies that investigated idiomatic expressions collected limited information about idiom properties for very small numbers of idioms only. In this paper we report on cross-boundary work aimed at developing a set of tools and language resources that are considered crucial for this kind of multifaceted research. We discuss the results of our research and suggest possible avenues for future research Ferdy Hubers, Catia Cucchiarini, Helmer Strik |
LREC | 3 |
| 2020 | BLISS: An Agent for Collecting Spoken Dialogue Data about Health and Well-beingabstractAn important objective in health-technology is the ability to gather information about people’s well-being. Structured interviews can be used to obtain this information, but are time-consuming and not scalable. Questionnaires provide an alternative way to extract such information, though typically lack depth. In this paper, we present our first prototype of the BLISS agent, an artificial intelligent agent which intends to automatically discover what makes people happy and healthy. The goal of Behaviour-based Language-Interactive Speaking Systems (BLISS) is to understand the motivations behind people’s happiness by conducting a personalized spoken dialogue based on a happiness model. We built our first prototype of the model to collect 55 spoken dialogues, in which the BLISS agent asked questions to users about their happiness and well-being. Apart from a description of the BLISS architecture, we also provide details about our dataset, which contains over 120 activities and 100 motivations and is made available for usage. Jelte van Waterschoot, Iris Hendrickx, Esther Klabbers, Marcel de Korte, Helmer Strik, Catia Cucchiarini, Mariët Theune |
LREC | 6 |
| 2019 | Deep Sensing of Breathing Signal During Conversational SpeechabstractContains fulltext : 214126.pdf (Publisher’s version ) (Open Access) Venkata Srikanth Nallanthighal, Aki Härmä, Helmer Strik |
INTERSPEECH | 3 |
| 2018 | Overview of the 2018 Spoken CALL Shared TaskabstractContains fulltext : 199312.pdf (Publisher’s version ) (Open Access) Claudia Baur, Andrew Caines, Cathy Chua, Johanna Gerlach, Mengjie Qian 0001, Manny Rayner, Martin J. Russell, Helmer Strik, Xizi Wei |
INTERSPEECH | 8 |
| 2017 | Multi-Stage DNN Training for Automatic Recognition of Dysarthric Speechabstract10.21437/Interspeech.2017-303 Emre Yilmaz 0001, Mario Ganzeboom, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 4 |
| 2016 | Intelligibility of Disordered Speech: Global and Detailed ScoresabstractContains fulltext : 160745.pdf (Publisher’s version ) (Open Access) Mario Ganzeboom, Marjoke Bakker, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 4 |
| 2016 | Combining Non-Pathological Data of Different Language Varieties to Improve DNN-HMM Performance on Pathological SpeechabstractContains fulltext : 160601pub.pdf (Publisher’s version ) (Open Access) Emre Yilmaz 0001, Mario Ganzeboom, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 4 |
| 2016 | A Shared Task for Spoken CALL?
Claudia Baur, Johanna Gerlach, Manny Rayner, Martin J. Russell, Helmer Strik |
LREC | 5 |
| 2016 | A Dutch Dysarthric Speech Database for Individualized Speech Therapy Research
Emre Yilmaz 0001, Mario Ganzeboom, Lilian Beijer, Catia Cucchiarini, Helmer Strik |
LREC | 5 |
| 2015 | Auris populi: crowdsourced native transcriptions of Dutch vowels spoken by adult Spanish learnersabstract\n Contains fulltext :\n 145184.pdf (Publisher’s version ) (Open Access)\n Pepi Burgos, Eric Sanders, Catia Cucchiarini, Roeland van Hout, Helmer Strik |
INTERSPEECH | 5 |
| 2015 | Confusability in L2 vowels: analyzing the role of different featuresabstractContains fulltext : 150849.pdf (Publisher’s version ) (Open Access) Matyas Jani, Catia Cucchiarini, Roeland van Hout, Helmer Strik |
INTERSPEECH | 4 |
| 2014 | Dutch vowel production by Spanish learners: duration and spectral featuresabstractIn this paper we present a study on Dutch vowel production by Spanish learners that was carried out within the framework of our research on Computer Assisted Pronunciation Training (CAPT). The aim of this study was to obtain detailed information on production of Dutch vowels by Spanish learners, which can be employed to develop effective CAPT programs for this specific target group. We collected speech from learners with varying proficiency levels (A1 - B2 of the CEFR), which was transcribed, segmented and acoustically analyzed. We present data on the frequency of pronunciation errors and on detailed analyses of duration and acoustic properties of the vocalic realizations. The results indicate that Spanish learners of Dutch have difficulties in realizing several Dutch vowel contrasts and that they differ from native speakers in the way they employ duration and spectral properties to realize these contrasts. We discuss these results in relation to those of previous studies on Dutch vowel perception by Spanish listeners and relate them to current theories on speech learning. Index Terms: L2 phonology acquisition, language learning, Computer Assisted Pronunciation Training (CAPT) Pepi Burgos, Matyas Jani, Catia Cucchiarini, Roeland van Hout, Helmer Strik |
INTERSPEECH | 5 |
| 2014 | ASR-based CALL systems and learner speech data: new resources and opportunities for research and development in second language learning
Catia Cucchiarini, Steve Bodnar, Bart Penning de Vries, Roeland van Hout, Helmer Strik |
LREC | 5 |
| 2013 | Pronunciation errors by Spanish learners of Dutch: a data-driven study for ASR-based pronunciation trainingabstractIn this paper we report on a study on pronunciation errors by Spanish learners of Dutch, which was aimed at obtaining information to develop a dedicated Computer Assisted Pronunciation Training (CAPT) program for this fixed language pair (Spanish L1, Dutch L2).The results of our study indicate, that, first, vowel errors are more frequent and variable than consonant mispronunciations.Second, Spanish natives appear to have problems with vowel length, vowel height, and front rounded vowels.Third, they tend to fall back on the pronunciation of their L1 vowels. Pepi Burgos, Catia Cucchiarini, Roeland van Hout, Helmer Strik |
INTERSPEECH | 4 |
| 2013 | L2 syntax acquisition: the effect of oral and written computer assisted practiceabstractContains fulltext : 116143.pdf (Publisher’s version ) (Open Access) Polina Drozdova, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 3 |
| 2012 | The effect of domain and text type on text prediction quality
Suzan Verberne, Antal van den Bosch, Helmer Strik, Lou Boves |
EACL | 3 |
| 2012 | Practice and feedback in L2 speaking: an evaluation of the DISCO CALL systemabstractContains fulltext : 101943.pdf (Publisher’s version ) (Open Access) Catia Cucchiarini, Joost van Doremalen, Helmer Strik |
INTERSPEECH | 3 |
| 2012 | The DISCO ASR-based CALL system: practicing L2 oral skills and beyond
Helmer Strik, Jozef Colpaert, Joost van Doremalen, Catia Cucchiarini |
LREC | 1 |
| 2011 | Computer-assisted Grammar Practice for Oral Communication
Stephen Bodnar, Catia Cucchiarini, Helmer Strik |
CSEDU (1) | 3 |
| 2011 | Error Selection for ASR-Based English Pronunciation Training in 'My Pronunciation Coach'abstractIn this paper we report on a study of pronunciation errors that was conducted within the framework of the project "My Pronunciation Coach", which is aimed at developing an ASRbased system for pronunciation training for learners of English with Dutch as their mother tongue.The aim of this study was to obtain quantitative data on the occurrence of pronunciation errors in Dutch English speech.We present the results of this study and compare them to those of previous investigations.Finally, we discuss the implications of these results for the development of My Pronunciation Coach. Catia Cucchiarini, Henk van den Heuvel, Eric Sanders, Helmer Strik |
INTERSPEECH | 4 |
| 2010 | Using non-native error patterns to improve pronunciation verificationabstractIn this paper we show how a pronunciation quality measure can be improved by making use of information on frequent pronunciation errors made by non-native speakers. We propose a new measure, called weighted Goodness of Pronunciation (wGOP), and compare it to the much used GOP measure. We applied this measure to the task of discriminating correctly from incorrectly realized Dutch vowels produced by non-native speakers and observed a substantial increase in performance when sufficient training material is available. Index Terms: pronunciation error detection, computer-assisted language learning, confidence measures, weighted GOP Joost van Doremalen, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 3 |
| 2010 | Human Language Technology and Communicative Disabilities: Requirements and Possibilities for the Future
Marina B. Ruiter, Toni C. M. Rietveld, Catia Cucchiarini, Emiel Krahmer, Helmer Strik |
LREC | 5 |
| 2009 | Automatic detection of vowel pronunciation errors using multiple information sourcesabstractFrequent pronunciation errors made by L2 learners of Dutch often concern vowel substitutions. To detect such pronunciation errors, ASR-based confidence measures (CMs) are generally used. In the current paper we compare and combine confidence measures with MFCCs and phonetic features. The results show that the best results are obtained by using MFCCs, then CMs, and finally phonetic features, and that substantial improvements can be obtained by combining different features. Joost van Doremalen, Catia Cucchiarini, Helmer Strik |
ASRU | 3 |
| 2009 | Optimizing non-native speech recognition for CALL applicationsabstractWe are developing a Computer Assisted Language Learning (CALL) system that gives feedback to grammar and pronunciation that makes use of Automatic Speech Recognition (ASR). However, good quality unconstrained non-native ASR is not yet feasible. Therefore, we use an approach in which we try to elicit constrained responses. The task in the current experiments is to select utterances from a list of responses. The results of our experiments show that significant improvements can be obtained by optimizing the language model and acoustic models. In this way we could reduce the utterance error rate from 29-26 % to Joost van Doremalen, Helmer Strik, Catia Cucchiarini |
INTERSPEECH | 2 |
| 2009 | Functional data analysis as a tool for analyzing speech dynamics - a case study on the French word c'étaitabstractIn this paper we introduce Functional Data Analysis (FDA) as a tool for analyzing dynamic transitions in speech signals. FDA makes it possible to perform statistical analyses of sets of mathematical functions in the same way as classical multivariate analysis treats scalar measurement data. We illustrate the use of FDA with a reduction phenomenon affecting the French word c'était /setε/ 'it was', which can be reduced to [stε] in conversational speech. FDA reveals that the dynamics of the transition from [s] to [t] in fully reduced cases may still be different from the dynamics of [s] - [t] transitions in underlying /st/ clusters such as in the word stage. Michele Gubian, Francisco Torreira, Helmer Strik, Lou Boves |
INTERSPEECH | 3 |
| 2009 | Oral proficiency training in Dutch L2: The contribution of ASR-based corrective feedback
Catia Cucchiarini, Ambra Neri, Helmer Strik |
Speech Commun. | 3 |
| 2009 | Comparing different approaches for automatic pronunciation error detection
Helmer Strik, Khiet P. Truong, Febe de Wet, Catia Cucchiarini |
Speech Commun. | 1 |
| 2008 | DISCO: development and integration of speech technology into courseware for language learningabstractRecent research has shown that a properly designed ASR-based CALL system (Dutch-CAPT) was capable of detecting pronunciation errors and of providing comprehensible feedback on pronunciation. Since pronunciation is not the only skill required for speaking a second language, we explored the possibility of extending the Dutch-CAPT approach to other aspects of speaking proficiency like morphology and syntax. In this paper we explain how a number of errors in morphology and syntax that are common in spoken Dutch L2 could be addressed in an ASR-based CALL system. Finally, we present our new project in which corrective feedback will be provided on all three aspects of spoken proficiency: pronunciation, morphology and syntax. Index Terms: pronunciation training, CALL, ASR, error detection. Catia Cucchiarini, Joost van Doremalen, Helmer Strik |
INTERSPEECH | 3 |
| 2008 | Pronunciation reduction: how it relates to speech style, gender, and ageabstract\n Contains fulltext :\n 68373.pdf (author's version ) (Open Access)\n Helmer Strik, Joost van Doremalen, Catia Cucchiarini |
INTERSPEECH | 1 |
| 2007 | Segment deletion in spontaneous speech: a corpus study using mixed effects models with crossed random effectsabstract\n Contains fulltext :\n 44640.pdf (Publisher’s version ) (Open Access)\n Christophe Van Bael, R. Harald Baayen, Helmer Strik |
INTERSPEECH | 3 |
| 2007 | ASR-based pronunciation training: scoring accuracy and pedagogical effectiveness of a system for dutch L2 learnersabstractA system for providing Computer Assisted Pronunciation Training for Dutch was developed, Dutch-CAPT, which appeared to be effective in improving pronunciation quality of L2 learners of Dutch.In this paper we describe the architecture of the system paying particular attention to the rationale behind this system, to the performance of the error detection algorithm and its relationship to the pedagogical effectiveness of the corrective feedback provided Index Terms: Computer Assisted Pronunciation Training (CAPT), corrective feedback, pronunciation error detection, Goodness Of Pronunciation (GOP) Catia Cucchiarini, Ambra Neri, Febe de Wet, Helmer Strik |
INTERSPEECH | 4 |
| 2007 | Structure-based and template-based automatic speech recognition - comparing parametric and non-parametric approachesabstractThis paper provides an introductory tutorial for the Interspeech07 special session on “Structure-Based and Template-Based Automatic Speech Recognition”. The purpose of the special session is to bring together researchers who have special interest in novel techniques that are aimed at overcoming weaknesses of HMMs for acoustic modeling in speech recognition. Numerous such approaches have been taken over the past dozen years, which can be broadly classified into structured-based (parametric) and templatebased (non-parametric) ones. In this paper, we will provide an overview of both approaches, focusing on the incorporation of long-range temporal dependencies of the speech features and phonetic detail in speech recognition algorithms. We will provide a high-level survey on major existing work and systems using these two types of “beyond-HMM” frameworks. The contributed papers in this special session will elaborate further on the related topics. Index Terms: structure-based, template-based, automatic speech recognition Li Deng 0001, Helmer Strik |
INTERSPEECH | 2 |
| 2007 | Comparing classifiers for pronunciation error detectionabstractCITATION: Strik, H. et al. 2007. Comparing classifiers for pronunciation error detection. In Hamme, H. van; Son, R. van (ed.), Proceedings of Interspeech 2007, pp. 1837-1840. Helmer Strik, Khiet P. Truong, Febe de Wet, Catia Cucchiarini |
INTERSPEECH | 1 |
| 2007 | Automatic phonetic transcription of large speech corpora
Christophe Van Bael, Lou Boves, Henk van den Heuvel, Helmer Strik |
Comput. Speech Lang. | 4 |
| 2006 | Automatic phonetic transcription of large speech corpora: a comparative studyabstractMost large speech corpora are delivered with a lexicon that contains a canonical transcription of every word in the orthographic transcription.Such a lexicon can be used for generating a hypothetical 'canonical' phonetic transcription from the orthography.In addition, time and money permitting, some speech corpora are provided with a manually verified broad phonetic transcription of at least part of the material.Since the manual verification of phonetic transcriptions is time-consuming and expensive, we investigated whether existing automatic transcription procedures and combinations of such procedures can offer a quick and cheap alternative for the generation of phonetic transcriptions like the manually verified transcriptions delivered with large speech corpora.In our study, we used ten automatic transcription procedures to generate a broad phonetic transcription of well-prepared speech (readaloud texts) and spontaneous speech (telephone dialogues) from the Spoken Dutch Corpus.The performance was assessed in terms of the number and the nature of the discrepancies between the emerging phonetic transcriptions and the corresponding manually verified phonetic transcriptions delivered with the Spoken Dutch Corpus.The resulting automatic transcriptions appeared to be comparable to the manually verified transcriptions. Christophe Van Bael, Lou Boves, Henk van den Heuvel, Helmer Strik |
INTERSPEECH | 4 |
| 2006 | ASR-based corrective feedback on pronunciation: does it really work?abstractWe studied a group of immigrants who were following regular, teacher-fronted Dutch classes, and who were assigned to three groups using either a) Dutch CAPT, an ASR-based Computer Assisted Pronunciation Training (CAPT) system that provides feedback on a number of Dutch speech sounds that are problematic for L2 learners b) a CAPT system without feedback c) no CAPT system. Participants were tested before and after the training. The results show that the ASR-based feedback was effective in correcting the errors addressed in the training. 1. Ambra Neri, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 3 |
| 2005 | Multiword expressions in spontaneous speech: do we really speak like that?abstractIn this study, we examined the pronunciation characteristics of multiword expressions (MWEs). We first drew up an inventory of frequently occurring N-grams extracted from orthographic transcriptions of spontaneous speech contained in a large corpus of spoken Dutch. For about 10 % of these Ngrams phonetic transcriptions were available, which were examined. Our results show that the pronunciation of these Ngrams differed to a large extent from the canonical form. In order to determine whether this is a general characteristic of spontaneous speech or rather the effect of the specific status of these N-grams, we analyzed the pronunciations of the individual words composing the N-grams in two context conditions: 1) in the N-gram context and 2) in any other context. We found that words in N-grams do indeed have peculiar pronunciation patterns. This seems to suggest that these N-grams may be considered as MWEs that should therefore be treated as lexical entries with their own specific pronunciation variants in the pronunciation lexicons used for e.g. automatic speech recognition (ASR) and automatic phonetic transcription (APT). 1. Helmer Strik, Diana Binnenpoorte, Catia Cucchiarini |
INTERSPEECH | 1 |
| 2005 | Automatic detection of frequent pronunciation errors made by L2-learnersabstractContains fulltext : 41035.pdf (Publisher’s version ) (Open Access) Khiet P. Truong, Ambra Neri, Febe de Wet, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 5 |
| 2005 | Multiword expressions in spoken language: An exploratory study on pronunciation variation
Diana Binnenpoorte, Catia Cucchiarini, Lou Boves, Helmer Strik |
Comput. Speech Lang. | 4 |
| 2004 | Investigating speech style specific pronunciation variation in large spoken language corporaabstractIn the past, linguistic research was typically conducted on relatively small datasets that were specifically designed for the research at hand.Whereas to date many large spoken language corpora have become available, the usefulness of these corpora is still not fully established in linguistic research.The research reported on in this paper was conducted to illustrate the potential of large multi-purpose spoken language corpora for linguistic research.The possibility was investigated of identifying phonetic regularities in different speech styles.To this end, a datadriven study was conducted with a large multi-purpose spoken language corpus comprising a manually corrected broad phonetic transcription of the data.Our results show that speech style specific pronunciation processes can indeed be found in such a large corpus.This indicates that large multi purpose spoken language corpora can contribute to linguistic research, if only for the purpose of hypothesis generation and verification. Christophe Van Bael, Henk van den Heuvel, Helmer Strik |
INTERSPEECH | 3 |
| 2004 | Towards automatic word segmentation of dialect speechabstractThis paper is about the creation of a digital dialect database, and the focus is on automatic word segmentation.Automatic word segmentation has been studied by several research groups during the last two decades.However, the task we are faced with differs in several respects from previous ones.For instance, in our case we are dealing with recordings of interviews containing spontaneous dialect speech and 'enriched' (quasi-phonetic) orthographic transcriptions (instead of 'normal' orthographic transcriptions, which are usually available).Furthermore, the nature of the task requires that the word segmentation procedure can be adapted for each interview. Eric Sanders, Andrea Diersen, Willy Jongenburger, Helmer Strik |
INTERSPEECH | 4 |
| 2004 | On the Usefulness of Large Spoken Language Corpora for Linguistic Research
Christophe Van Bael, Helmer Strik, Henk van den Heuvel |
LREC | 2 |
| 2004 | Improving Automatic Phonetic Transcription of Spontaneous Speech Through Variant-Based Pronunciation Variation Modelling
Diana Binnenpoorte, Catia Cucchiarini, Helmer Strik, Lou Boves |
LREC | 3 |
| 2004 | On automatic phonetic transcription quality: lower word error rates do not guarantee better transcriptions
Judith M. Kessens, Helmer Strik |
Comput. Speech Lang. | 2 |
| 2003 | Validation of phonetic transcriptions based on recognition performanceabstractIn fundamental linguistic as well as in speech technology re search there is an increasing need for procedures to automat ically generate and validate phonetic transcriptions.Whereas much research has already focussed on the automatic genera tion o f phonetic transcriptions, far less attention has been paid to the validation o f such transcriptions.In the little research performed in this area, the estimation o f the quality o f (auto matically generated) phonetic transcriptions is typically based on the comparison between these transcriptions and a human made reference transcription.We believe, however, that the quality o f phonetic transcriptions should ideally be estimated with the application in which the transcriptions will be used in mind, provided that the application is known at validation time.The application focussed on in this paper is automatic speech recognition, the validation criterion is the word error rate.We achieved a higher accuracy with a recogniser trained on an automatically generated transcription than with a similar recogniser trained on a human-made transcription resembling a human-made reference transcription more.This indicates that the traditional validation approach may not always be the most optimal one. Christophe Van Bael, Diana Binnenpoorte, Helmer Strik, Henk van den Heuvel |
INTERSPEECH | 3 |
| 2003 | Automatic transcription of football commentaries in the MUMIS projectabstractThis paper describes experiments carried out to automatically transcribe football commentaries in Dutch, English and German for multimedia indexing. Our results show that the high levels of stadium noise in the material create a task that is extremely difficult for conventional ASR. The baseline WERs vary from 83% to 94% for the three languages investigated. Employing state-of-the-art noise robustness techniques leads to relative reductions of 9-10% WER. Application specific words such as players names are recognized correctly in about 50% of cases. Although this result is substantially better than the overall result, it is inadequate. Much better results can be obtained if the football commentaries are recorded separately from the stadium noise. This would make the automatic transcriptions more useful for multimedia indexing. Janienke Sturm, Judith M. Kessens, Mirjam Wester, Febe de Wet, Eric Sanders, Helmer Strik |
INTERSPEECH | 6 |
| 2003 | A data-driven method for modeling pronunciation variation
Judith M. Kessens, Catia Cucchiarini, Helmer Strik |
Speech Commun. | 3 |
| 2002 | Feedback in computer assisted pronunciation training: technology push or demand pull?abstractIn this paper, we examine the type of feedback that currently available Computer Assisted Pronunciation Training (CAPT) systems provide, with a view to establishing whether this meets pedagogically sound requirements.W e show that many commercial systems tend to prefer technological novelties that do not always comply with pedagogical criteria and that despite the limitations of today's technology, it is possible to design CAPT systems that are more in line with learners' needs. Ambra Neri, Catia Cucchiarini, Helmer Strik |
INTERSPEECH | 3 |
| 2002 | Automatic recognition of dutch dysarthric speech: a pilot studyabstractThis paper describes a feasibility study into automatic recognition o f Dutch dysarthric speech.Recognition experiments with speaker independent and speaker dependent models are compared, for tasks with different perplexities.The results show that speaker dependent speech recognition for dysarthric speakers is very well possible, even for higher perplexity tasks. Eric Sanders, Marina B. Ruiter, Lilian Beijer, Helmer Strik |
INTERSPEECH | 4 |
| 2002 | Dutch HLT resources: from BLARK to priority listsabstract\n Contains fulltext :\n 76208.pdf (author's version ) (Open Access)\n Helmer Strik, Walter Daelemans, Diana Binnenpoorte, Janienke Sturm, Folkert de Vriend, Catia Cucchiarini |
INTERSPEECH | 1 |
| 2002 | Goal-directed ASR in a multimedia indexing and searching environment (MUMIS)abstract\n Contains fulltext :\n 76207.pdf (author's version ) (Open Access)\n Mirjam Wester, Judith M. Kessens, Helmer Strik |
INTERSPEECH | 3 |
| 2002 | A Field Survey for Establishing Priorities in the Development of HLT Resources for Dutch
Diana Binnenpoorte, Folkert de Vriend, Janienke Sturm, Walter Daelemans, Helmer Strik, Catia Cucchiarini |
LREC | 5 |
| 2001 | Lower WERs do not guarantee better transcriptionsabstractThe goal of this paper is to investigate the effect of various properties of the CSR on automatic transcription. To this end, we used various versions of a continuous speech recognizer (CSR) to make automatic transcriptions. Our results show that changing certain properties of the CSR affects the resulting automatic transcriptions. The best results were obtained when `short' hidden Markov models (HMMs), and contextindependent HMMs were used. Furthermore, we found that minimizing the amount of contamination in the HMMs improves the quality of the automatic transcriptions. Another important result is that there does not appear to be a straightforward relation between word error rate (WER) and the transcription quality. In other words: A CSR with a lower WER does not always guarantee better transcriptions. Judith M. Kessens, Helmer Strik |
INTERSPEECH | 2 |
| 2001 | Comparing the performance of two CSRs: how to determine the significance level of the differencesabstractWhen two CSRs are compared, it is important to test what the significance level of the difference is. For this purpose a metric and a statistical test are needed. In this paper we compare several combinations of a metric with a statistical test, in order to find a combination which is suitable for this task. Four combinations which are introduced in this paper appear to be suitable for this task. Helmer Strik, Catia Cucchiarini, Judith M. Kessens |
INTERSPEECH | 1 |
| 2000 | A bottom-up method for obtaining information about pronunciation variationabstract\n Contains fulltext :\n 76197.pdf (author's version ) (Open Access)\n Judith M. Kessens, Helmer Strik, Catia Cucchiarini |
INTERSPEECH | 2 |
| 2000 | L2 pronunciation quality in read and spontaneous speechabstractThis paper describes two experiments aimed at exploring the relationship between objective properties of speech and perceived pronunciation quality in read and spontaneous speech, with a view to determining whether such quantitative measures can be used to develop objective pronunciation tests. Read and spontaneous speech of two groups of 60 learners of Dutch as a second language was scored for pronunciation quality by human raters and was analyzed by means of a continuous speech recognizer to calculate six quantitative measures of speech quality related to speech timing. The results show that quantitative, temporal measures of speech are strongly related to pronunciation quality, in both read and spontaneous speech, although not all variables suitable for measuring pronunciation quality in read speech are as effective in spontaneous speech. 1. Helmer Strik, Catia Cucchiarini, Diana Binnenpoorte |
INTERSPEECH | 1 |
| 2000 | Comparing the recognition performance of CSRs: in search of an adequate metric and statistical significance testabstractIn this paper a new measure of recognition accuracy is introduced which can be used when comparing the performance of two speech recognizers, to establish which is the better one. This metric combines the advantages of previous measures, but excludes their disadvantages. Essentially, the metric is an attempt to quantify the degree of recognition accuracy for each sentence, thus obtaining a more informative measure than either correct or incorrect, in such a way that the statistical significance of the observed differences can be tested. The advantages of our assessment method are illustrated on the basis of both artificial and real performance data of different recognizers. 1. Helmer Strik, Catia Cucchiarini, Judith M. Kessens |
INTERSPEECH | 1 |
| 2000 | Pronunciation variation in ASR: which variation to model?abstractThis paper describes how the performance of a continuous speech recognizer for Dutch has been improved by modeling within-word and cross-word pronunciation variation. A relative improvement of 8.8% in WER was found compared to baseline system performance. However, as WERs do not reveal the full effect of modeling pronunciation variation, we performed a detailed analysis of the differences in recognition results that occur due to modeling pronunciation variation and found that indeed a lot of the differences in recognition results are not reflected in the error rates. Furthermore, error analysis revealed that testing sets of variants in isolation does not predict their behavior in combination. However, these results appeared to be corpus dependent. Mirjam Wester, Judith M. Kessens, Helmer Strik |
INTERSPEECH | 3 |
| 2000 | Different aspects of expert pronunciation quality ratings and their relation to scores produced by speech recognition algorithms
Catia Cucchiarini, Helmer Strik, Lou Boves |
Speech Commun. | 2 |
| 2000 | Erratum to: "Different aspects of expert pronunciation quality ratings and their relation to scores produced by speech recognition algorithms": [Speech Communication 30 (2000) 109-119]
Catia Cucchiarini, Helmer Strik, Lou Boves |
Speech Commun. | 2 |
| 1999 | Using likelihood ratios to perform utterance verification in automatic pronunciation assessmentabstractThe aim of our current research is to investigate the possibility of using likelihood ratios to perform utterance verification within the context of automatic oral proficiency assessment. The likelihood ratios under investigation have the appealing feature that they may be computed simply by using an off-theshelf automatic speech recognition system in two different recognition modes (forced and free phone) instead of using a system with specifically trained anti-models. We achieved 93% correct classification for 10 phonetically rich sentences uttered by 60 non-native language students. 1. INTRODUCTION The long-term goal of our research is to employ ASR technology in an automatic pronunciation test for Dutch as a second language. As a consequence of this aim we are not concerned with learners of Dutch with a specific mother tongue, but rather with a group of speakers who are highly varied in this respect. In this sense our situation is different from that of many studies on the use of... Febe de Wet, Catia Cucchiarini, Helmer Strik, Lou Boves |
EUROSPEECH | 3 |
| 1999 | Improving the performance of a Dutch CSR by modeling within-word and cross-word pronunciation variation
Judith M. Kessens, Mirjam Wester, Helmer Strik |
Speech Commun. | 3 |
| 1999 | Editorial
Helmer Strik |
Speech Commun. | 1 |
| 1999 | Modeling pronunciation variation for ASR: A survey of the literature
Helmer Strik, Catia Cucchiarini |
Speech Commun. | 1 |
| 1998 | Quantitative assessment of second language learners' fluency: an automatic approachabstract\n Contains fulltext :\n 74998.pdf (author's version ) (Open Access)\n Catia Cucchiarini, Helmer Strik, Lou Boves |
ICSLP | 2 |
| 1998 | Assessment of dutch pronunciation by means of automatic speech recognition technologyabstract\n Contains fulltext :\n 75007.pdf (author's version ) (Open Access)\n Catia Cucchiarini, Febe de Wet, Helmer Strik, Lou Boves |
ICSLP | 3 |
| 1998 | The selection of pronunciation variants: comparing the performance of man and machineabstractDans cet article, les performances d'un outil de transcription automatique sont évaluées.L'outil de transcription est un reconnaisseur de parole continue (CSR) fonctionnant en mode de reconnaissance forcée.Pour l'évaluation les performances du CSR ont été comparées à celles de neuf auditeurs experts.La machine et l'humain ont effectué exactement la même tâche: décider si un segment était présent ou non dans 467 cas.Il s'est avéré que les performances du CSR étaient comparables à celle des experts. Judith M. Kessens, Mirjam Wester, Catia Cucchiarini, Helmer Strik |
ICSLP | 4 |
| 1998 | Modeling pronunciation variation for a dutch CSR: testing three methodsabstractThis paper describes how the performance of a continuous speech recognizer for Dutch has been improved by modeling pronunciation variation. We used three methods to model pronunciation variation. First, within-word variation was dealt with. Phonological rules were applied to the words in the lexicon, thus automatically generating pronunciation variants. Secondly, cross-word pronunciation variation was modeled using two different approaches. The first approach was to model cross-word processes by adding the variants as separate words to the lexicon and in the second approach this was done by using multi-words. For each of the methods, recognition experiments were carried out. A significant improvement was found for modeling within-word variation. Furthermore, modeling crossword processes using multi-words leads to significantly better results than modeling them using separate words in the lexicon. 1. INTRODUCTION The work reported on here concerns the Continuous Speech Recognition (CS... Mirjam Wester, Judith M. Kessens, Helmer Strik |
ICSLP | 3 |
| 1998 | Two automatic approaches for analyzing connected speech processes in dutchabstractThis paper describes two automatic approaches used to study connected speech processes (CSPs) in Dutch. The first approach was from a linguistic point of view- the top-down method. This method can be used for verification of hypotheses about CSPs. The second approach- the bottom-up method-uses a constrained phone recognizer to generate phone transcriptions. An alignment was carried out between the two transcriptions and a reference transcription. A comparison between the two methods showed that 68 % agreement was achieved on the CSPs. Although phone accuracy is only 63%, the bottom-up approach is useful for studying CSPs. From the data generated using the bottom-up method, indications of which CSPs are present in the material can be found. These indications can be used to generate hypotheses which can then be tested using the top-down method. 1. Mirjam Wester, Judith M. Kessens, Helmer Strik |
ICSLP | 3 |
| 1997 | The effect of low-pass filtering on estimated voice source parametersabstractVoice source parameters are often obtained by parametrizing glottal flow signals.However, before parametrization these glottal flow signals are usually lowpass filtered.As low-pass filtering changes the shape of the glottal pulses, it will also cause an error in the estimated voice source parameters.The present article presents results of our research on the effect of low-pass filtering on the estimated voice source parameters.We will first present an evaluation method which makes it possible to study the effect of low-pass filtering in detail.The evaluation results show that low-pass filtering leads to an error in all estimated voice source parameters.However, the magnitude of the errors differs for the various voice source parameters, and also depends on the estimation method used.We will show that the errors can be reduced substantially by choosing the appropriate estimation method. Helmer Strik |
EUROSPEECH | 1 |
| 1997 | Parabolic spectral parameter - A new method for quantification of the glottal flow
Paavo Alku, Helmer Strik, Erkki Vilkman |
Speech Commun. | 2 |
| 1996 | Localizing an automatic inquiry system for public transport informationabstractThis paper reports on the development o f a spoken dialogue system for providing information about public transport in the Netherlands.It is explained how a German prototype was adapted for Dutch.Emphasis is laid on the specific approach chosen to collect speech material that could be used to gradually improve the system.The pros and cons of this method are discussed. Helmer Strik, Albert Russel, Henk van den Heuvel, Catia Cucchiarini, Lou Boves |
ICSLP | 1 |
| 1995 | The relation between physiological signals and F0: a quantitative analysis methodabstractMeasurements were obtained of several physiological mechanisms which are known to be important in the control of fundamental frequency (F0). The data were analysed by means of a multiple regression analysis in which F0 is the criterion and the physiological signals are the predictors. Separate analyses were carried out for statements and questions, and for falling and rising F0. The results reveal no considerable differences in the control of F0 for the various datasets. 1. Helmer Strik |
EUROSPEECH | 1 |
| 1994 | Automatic estimation of voice source parametersabstractVoice source parameters can be estimated by fitting a voice source model to the glottal flow signal which is obtained by means of inverse filtering. In this paper we investigate the behaviour of the LF-model in a number of non-linear parameter estimation procedures. It is concluded that (1) the parameter estimates are robust against additive (white and narrow band) noise in the flow waveforms, (2) simplex search algorithms perform better than steepest descent algorithms, provided that (3) the LF-pulse is generated with an algorithm that treats all parameters as real numbers. 1. Helmer Strik, Lou Boves |
ICSLP | 1 |
| 1993 | Fitting a LF-model to inverse filter signalsabstractItem does not contain fulltext Helmer Strik, Bert Cranen, Lou Boves |
EUROSPEECH | 1 |
| 1992 | Comparing methods for automatic extraction of voice source parameters from continuous speechabstractItem does not contain fulltext Helmer Strik, Joop Jansen, Lou Boves |
ICSLP | 1 |
| 1992 | On the relation between voice source parameters and prosodic features in connected speech
Helmer Strik, Lou Boves |
Speech Commun. | 1 |
| 1991 | Optimizing lexical fast search in a large vocabulary isolated word speech recognition systemabstractItem does not contain fulltext H. Drexler, R. Roddeman, Lou Boves, Helmer Strik |
EUROSPEECH | 4 |
| 1991 | On the relation between voice source characteristics and prosodyabstractItem does not contain fulltext Helmer Strik, Lou Boves |
EUROSPEECH | 1 |
| 1990 | Extraction of control parameters for the voice source in a text-to-speech systemabstractIn order to derive voice source control rules from natural speech, the parameters of a source model must be derived from the acoustic signal. This is done by parameterizing the results of glottal inverse filtering. A number of different inverse filtering procedures and the ease with which their results can be parameterized are compared. It is shown that closed glottic interval covariance linear predictive coding is as powerful as more sophisticated techniques, because it is the only known method that can strictly be limited to the closed glottis interval.> Johan de Veth, Bert Cranen, Helmer Strik, Lou Boves |
ICASSP | 3 |
| 1989 | The fundamental frequency - subglottal pressure ratioabstractItem does not contain fulltext Helmer Strik, Lou Boves |
EUROSPEECH | 1 |