EDBT 2026 Demo / reviewers in the wild / expert
Louis ten Bosch
dblp:04/5667
· DBLP profile ↗
102ranked-venue papers
35as first author
14since 2021 · last 2025
0000-0002-0152-9024ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 90 · 34 first-author · 8 since 2021Artificial intelligence and machine learning · 87 · 30 first-author · 14 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dual-Objective Adversarial Disentanglement for Protecting Speech Data used for Diagnosing Parkinson's DiseaseabstractRecently, the challenge of protecting privacy-sensitive information in speech data has received growing attention. While many studies have explored the protection of speaker identity, the protection of individual speaker attributes, such as gender, has not been thoroughly investigated. In this paper, we propose a dual-objective approach to adversarial disentanglement that protects the gender attribute of the speaker in speech data used for the diagnosis of Parkinson's disease (PD). The approach combines an adversarial Gradient Reversal Layer (GRL) objective with a utility objective. Experiments on the PC-GITA and NeuroVoz PD speech datasets show that our approach can block the ability of a classifier to infer the gender of the speaker, while preserving the utility for diagnostic purposes. Our work contributes to speech privacy, but also to the understanding of gender for PD diagnosis. Mehtab Ur Rahman, Martha A. Larson, Louis ten Bosch, Cristian Tejedor García |
CBMI | 3 |
| 2025 | Word stress in self-supervised speech models: A cross-linguistic comparison
Martijn Bentum, Louis ten Bosch, Tomas O. Lentz |
INTERSPEECH | 2 |
| 2025 | Evaluating the Usefulness of Non-Diagnostic Speech Data for Developing Parkinson's Disease ClassifiersabstractContains fulltext : 322329.pdf (Publisher’s version ) (Open Access) Terry Yi Zhong, Esther Janse, Cristian Tejedor García, Louis ten Bosch, Martha A. Larson |
INTERSPEECH | 4 |
| 2024 | Ensembles of Hybrid and End-to-End Speech RecognitionabstractWe propose a method to combine the hybrid Kaldi-based Automatic Speech Recognition (ASR) system with the end-to-end wav2vec 2.0 XLS-R ASR using confidence measures. Our research is focused on the low-resource Irish language. Given the limited available open-source resources, neither the standalone hybrid ASR nor the end-to-end ASR system can achieve optimal performance. By applying the Recognizer Output Voting Error Reduction (ROVER) technique, we illustrate how ensemble learning could facilitate mutual error correction between both ASR systems. This paper outlines the strategies for merging the hybrid Kaldi ASR model and the end-to-end XLS-R model with the help of confidence scores. Although contemporary state-of-the-art end-to-end ASR models face challenges related to prediction overconfidence, we utilize Renyi’s entropy-based confidence approach, tuned with temperature scaling, to align it with the Kaldi ASR confidence. Although there was no significant difference in the Word Error Rate (WER) between the hybrid and end-to-end ASR, we could achieve a notable reduction in WER after ensembling through ROVER. This resulted in an almost 14% Word Error Rate Reduction (WERR) on our primary test set and an approximately 20% WERR on other noisy and imbalanced test data. Aditya Parikh, Louis ten Bosch, Henk van den Heuvel |
LREC/COLING | 2 |
| 2024 | SignON - a Co-creative Machine Translation for Sign and Spoken Languages (end-of-project results, contributions and lessons learned)abstractSignON, a 3-year Horizon 20202 project addressing the lack of technology and services for MT between sign languages (SLs) and spoken languages (SpLs) ended in December 2023. SignON was unprecedented. Not only it addressed the wider complexity of the aforementioned problem – from research and development of recognition, translation and synthesis, through development of easy-to-use mobile applications and a cloud-based framework to do the “heavy lifting” as well as to establishing ethical, privacy and inclusivenesspolicies and operation guidelines – but also engaged with the deaf and hard of hearing communities in an effective co-creation approach where these main stakeholders drove the development in the right direction and had the final say.Currently we are witnessing advances in natural language processing for SLs, including MT. SignON was one of the largest projects that contributed to this surge with 17 partners and more than 60 consortium members, working in parallel with other international and European initiatives, such as project EASIER and others. Dimitar Sht. Shterionov, Vincent Vandeghinste, Mirella De Sisto, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Andy Way, Josep Blat, Frankie Picron, Davy Van Landuyt, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Caro Brosens, Jorn Rijckaert, Víctor Ubieto Nogales, Bram Vanroy, Santiago Egea Gómez, Ineke Schuurman, Gorka Labaka, Adrián Núñez-Marcos, Irene Murtagh, Euan McGill, Horacio Saggion |
EAMT (2) | 13 |
| 2024 | The Processing of Stress in End-to-End Automatic Speech Recognition ModelsabstractListeners use stress to facilitate word recognition and speech segmentation. Classical ASR systems did not incorporate stress in their recognition process. In contrast, end-to-end ASR systems may use the information carried by stress. The present study shows that Wav2vec 2.0 is indeed sensitive to stress, and that this sensitivity is not a mere reflection of acoustic correlates of stress. Diagnostic classifiers of the CNN output reveal vowel-specific stress representations, that perform on par with acoustic features. Stress classifiers trained on transformer layers outperform classifiers based on acoustic correlates, but degrade when context is removed, showing that higher layers take the relative nature of stress into account. Results obtained by testing a stress classifier on a vowel it is not trained on, show that stress processing is to some extent abstract, i.e., the classifier does not simply detect a set of stressed vowel representations but rather, their common denominator Martijn Bentum, Louis ten Bosch, Tom Lentz |
INTERSPEECH | 2 |
| 2023 | SignON: Sign Language Translation. Progress and challengesabstractSignON (https://signon-project.eu/) is a Horizon 2020 project, running from 2021 until the end of 2023, which addresses the lack of technology and services for the automatic translation between sign languages (SLs) and spoken languages, through an inclusive, human-centric solution, hence contributing to the repertoire of communication media for deaf, hard of hearing (DHH) and hearing individuals. In this paper, we present an update of the status of the project, describing the approaches developed to address the challenges and peculiarities of SL machine translation (SLMT). Vincent Vandeghinste, Dimitar Sht. Shterionov, Mirella De Sisto, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Josep Blat, Frankie Picron, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Jorn Rijckaert, Bram Vanroy, Víctor Ubieto Nogales, Santiago Egea Gómez, Ineke Schuurman, Gorka Labaka, Adrián Núñez-Marcos, Irene Murtagh, Euan McGill, Horacio Saggion |
EAMT | 11 |
| 2023 | Exploring the Importance of Sign Language Phonology for a Deep Neural NetworkabstractWe conduct an initial investigation to gain insight into whether a deep neural network learns phonological aspects of sign language when classifying video recordings of isolated signs from a continuous signing scenario.We train a series of neural networks to distinguish pairs of signs in Dutch Sign Language, controlling the phonological difference between the signs in each pair.Our results suggest that the intrinsic dimension of the final hidden layer of a network is surprisingly insensitive to the phonological difference between the signs in a pair.However, the ability of the network to discriminate two signs shows a clear trend towards increasing with increasing phonological distinctiveness. Related WorkIsolated sign language recognition.Sign language recognition (SLR) is the problem of recognizing and identifying a particular sign in a video clip.In this paper, we study isolated SLR, also known as word-level SLR, which 1 https://github.com/JavierMartnz/MindTheLinguisticGap Javier Martinez Rodriguez, Martha A. Larson, Louis ten Bosch |
ESANN | 3 |
| 2023 | Phonemic competition in end-to-end ASR modelsabstractContains fulltext : 296407.pdf (Publisher’s version ) (Open Access) Louis ten Bosch, Martijn Bentum, Lou Boves |
INTERSPEECH | 1 |
| 2022 | Sign Language Translation: Ongoing Development, Challenges and Innovations in the SignON ProjectabstractThe SignON project (www.signon-project.eu) focuses on the research and development of a Sign Language (SL) translation mobile application and an open communications framework. SignON rectifies the lack of technology and services for the automatic translation between signed and spoken languages, through an inclusive, humancentric solution which facilitates communication between deaf, hard of hearing (DHH) and hearing individuals. We present an overview of the current status of the project, describing the milestones reached to date and the approaches that are being developed to address the challenges and peculiarities of Sign Language Machine Translation (SLMT). Dimitar Sht. Shterionov, Mirella De Sisto, Vincent Vandeghinste, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Josep Blat, Frankie Picron, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Jorn Rijckaert |
EAMT | 11 |
| 2022 | A Speech Recognizer for Frisian/Dutch Council MeetingsabstractWe developed a bilingual Frisian/Dutch speech recognizer for council meetings in Fryslân (the Netherlands). During these meetings both Frisian and Dutch are spoken, and code switching between both languages shows up frequently. The new speech recognizer is based on an existing speech recognizer for Frisian and Dutch named FAME!, which was trained and tested on historical radio broadcasts. Adapting a speech recognizer for the council meeting domain is challenging because of acoustic background noise, speaker overlap and the jargon typically used in council meetings. To train the new recognizer, we used the radio broadcast materials utilized for the development of the FAME! recognizer and added newly created manually transcribed audio recordings of council meetings from eleven Frisian municipalities, the Frisian provincial council and the Frisian water board. The council meeting recordings consist of 49 hours of speech, with 26 hours of Frisian speech and 23 hours of Dutch speech. Furthermore, from the same sources, we obtained texts in the domain of council meetings containing 11 million words; 1.1 million Frisian words and 9.9 million Dutch words. We describe the methods used to train the new recognizer, report the observed word error rates, and perform an error analysis on remaining errors. Martijn Bentum, Louis ten Bosch, Henk van den Heuvel, Simone Wills, Domenique van der Niet, Jelske Dijkstra, Hans Van de Velde |
LREC | 2 |
| 2021 | Word Competition: An Entropy-Based Approach in the DIANA Model of Human Word Comprehensionabstract\n Contains fulltext :\n 238372.pdf (Publisher’s version ) (Open Access)\n \n Contains fulltext :\n 238372pre.pdf (Author’s version preprint ) (Open Access)\n Louis ten Bosch, Lou Boves |
Interspeech | 1 |
| 2021 | Time-to-Event Models for Analyzing Reaction Time Sequencesabstract\n Contains fulltext :\n 238373.pdf (Publisher’s version ) (Open Access)\n Louis ten Bosch, Lou Boves |
Interspeech | 1 |
| 2021 | Models of Reaction Times in Auditory Lexical Decision: RTonset versus RToffsetabstract\n Contains fulltext :\n 238370.pdf (Publisher’s version ) (Open Access)\n Sophie Brand, Kimberley Mulder, Louis ten Bosch, Lou Boves |
Interspeech | 3 |
| 2020 | Comparing EEG Analyses with Different Epoch Alignments in an Auditory Lexical Decision Experimentabstract\n Contains fulltext :\n 228101.pdf (Publisher’s version ) (Open Access)\n Louis ten Bosch, Kimberley Mulder, Lou Boves |
INTERSPEECH | 1 |
| 2019 | Listening with Great Expectations: An Investigation of Word Form Anticipations in Naturalistic SpeechabstractThe event-related potential (ERP) component named phonological mismatch negativity (PMN) arises when listeners hear an unexpected word form in a spoken sentence [1]. The PMN is thought to reflect the mismatch between expected and perceived auditory speech input. In this paper, we use the PMN to test a central premise in the predictive coding framework [2], namely that the mismatch between prior expectations and sensory input is an important mechanism of perception. We test this with natural speech materials containing approximately 50,000 word tokens. The corresponding EEG-signal was recorded while participants (n = 48) listened to these materials. Following [3], we quantify the mismatch with two word probability distributions (WPD): a WPD based on preceding context, and a WPD that is additionally updated based on the incoming audio of the current word. We use the between-WPD cross entropy for each word in the utterances and show that a higher cross entropy correlates with a more negative PMN. Our results show that listeners anticipate auditory input while processing each word in naturalistic speech. Moreover, complementing previous research, we show that predictive language processing occurs across the whole probability spectrum. Martijn Bentum, Louis ten Bosch, Antal van den Bosch, Mirjam Ernestus |
INTERSPEECH | 2 |
| 2019 | Quantifying Expectation Modulation in Human Speech ProcessingabstractThe mismatch between top-down predicted and bottom-up perceptual input is an important mechanism of perception according to the predictive coding framework (Friston, [1]).In this paper we develop and validate a new information-theoretic measure that quantifies the mismatch between expected and observed auditory input during speech processing.We argue that such a mismatch measure is useful for the study of speech processing.To compute the mismatch measure, we use naturalistic speech materials containing approximately 50,000 word tokens.For each word token we first estimate the prior word probability distribution with the aid of statistical language modelling, and next use automatic speech recognition to update this word probability distribution based on the unfolding speech signal.We validate the mismatch measure with multiple analyses, and show that the auditory-based update improves the probability of the correct word and lowers the uncertainty of the word probability distribution.Based on these results, we argue that it is possible to explicitly estimate the mismatch between predicted and perceived speech input with the cross entropy between word expectations computed before and after an auditory update. Martijn Bentum, Louis ten Bosch, Antal van den Bosch, Mirjam Ernestus |
INTERSPEECH | 2 |
| 2019 | Analyzing Reaction Time and Error Sequences in Lexical Decision ExperimentsabstractReaction times (RTs) are used widely in psychological and\npsycholinguistic research as inexpensive measures of underlying\ncognitive processes. However, inferring cognitive processes\nfrom RTs is hampered by the fact that actual responses are the\nresult of multiple factors, many of which may not be related to\nthe process of interest. In lexical decision experiments, the use\nof RTs is further complicated by the fact that the response to\nsome stimuli is missing, and the fact that part of the responses\nare ’incorrect’.\nIn this paper we investigate the distribution of missing and incorrect\nresponses in the RT sequences of two large lexical decision\nexperiments. It appears that a substantial part of incorrect\nresponses cluster together. Then, we investigate the effect of\nclusters of incorrect responses on surrounding RTs.\nAlso, we extend previous research on methods for discovering\nand removing so-called local speed effects from RT sequences.\nFor this purpose, we show that a recently introduced graph based\nRT analysis method can help to better understand and\nanalyze RT sequences. Louis ten Bosch, Lou Boves, Kimberley Mulder |
INTERSPEECH | 1 |
| 2019 | Phase Synchronization Between EEG Signals as a Function of Differences Between Stimuli CharacteristicsabstractThe neural processing of speech leads to specific patterns in\nthe brain which can be measured as, e.g., EEG signals. When\nproperly aligned with the speech input and averaged over many\ntokens, the Event Related Potential (ERP) signal is able to\ndifferentiate specific contrasts between speech signals. Well known\neffects relate to the difference between expected and\nunexpected words, in particular in the N400, while effects in\nN100 and P200 are related to attention and acoustic onset effects.\nMost EEG studies deal with the amplitude of EEG signals\nover time, sidestepping the effect of phase and phase synchronization.\nThis paper investigates the relation between phase in\nthe EEG signals measured in an auditory lexical decision task\nby Dutch participants listening to full and reduced English word\nforms. We show that phase synchronization takes place across\nstimulus conditions, and that the so-called circular variance is\nnarrowly related to the type of contrast between stimuli. Louis ten Bosch, Kimberley Mulder, Lou Boves |
INTERSPEECH | 1 |
| 2019 | ERP Signal Analysis with Temporal Resolution Using a Time Window Bankabstract\n Contains fulltext :\n 208217.pdf (Publisher’s version ) (Open Access)\n Annika Nijveld, Louis ten Bosch, Mirjam Ernestus |
INTERSPEECH | 2 |
| 2018 | Information Encoding by Deep Neural Networks: What Can We Learn?abstractThe recent advent of deep learning techniques in speech tech-nology and in particular in automatic speech recognition hasyielded substantial performance improvements. This suggeststhat deep neural networks (DNNs) are able to capture structurein speech data that older methods for acoustic modeling, suchas Gaussian Mixture Models and shallow neural networks failto uncover. In image recognition it is possible to link repre-sentations on the first couple of layers in DNNs to structuralproperties of images, and to representations on early layers inthe visual cortex. This raises the question whether it is possi-ble to accomplish a similar feat with representations on DNNlayers when processing speech input. In this paper we presentthree different experiments in which we attempt to untanglehow DNNs encode speech signals, and to relate these repre-sentations to phonetic knowledge, with the aim to advance con-ventional phonetic concepts and to choose the topology of aDNNs more efficiently. Two experiments investigate represen-tations formed by auto-encoders. A third experiment investi-gates representations on convolutional layers that treat speechspectrograms as if they were images. The results lay the basisfor future experiments with recursive networks. Louis ten Bosch, Lou Boves |
INTERSPEECH | 1 |
| 2018 | Analyzing Reaction Time Sequences from Human Participants in Auditory ExperimentsabstractSequences of reaction times (RT) produced by participants in an experiment are not only influenced by the stimuli, but by many other factors as well, including fatigue, attention, experience, IQ, handedness, etc. These confounding factors result in longterm effects (such as a participant’s overall reaction capability) and in short- and medium-time fluctuations in RTs (often referred to as ‘local speed effects’). Because stimuli are usually presented in a random sequence different for each participant, local speed effects affect the underlying ‘true’ RTs of specific trials in different ways across participants. To be able to focus statistical analysis on the effects of the cognitive process under study, it is necessary to reduce the effect of confounding factors as much as possible. In this paper we propose and compare techniques and criteria for doing so, with focus on reducing (‘filtering’) the local speed effects. We show that filtering matters substantially for the significance analyses of predictors in linear mixed effect regression models. The performance of filtering is assessed by the average between-participant correlation between filtered RT sequences and by Akaike’s Information Criterion, an important measure of the goodness-of-fit of linear mixed effect regression models. Louis ten Bosch, Mirjam Ernestus, Lou Boves |
INTERSPEECH | 1 |
| 2018 | Analyzing EEG Signals in Auditory Speech Comprehension Using Temporal Response Functions and Generalized Additive ModelsabstractAnalyzing EEG signals recorded while participants are listening to continuous speech with the purpose of testing linguistic hypotheses is complicated by the fact that the signals simultaneously reflect exogenous acoustic excitation and endogenous linguistic processing. This makes it difficult to trace subtle differences that occur in mid-sentence position. We apply an analysis based on multivariate temporal response functions to uncover subtle mid-sentence effects. This approach is based on a per-stimulus estimate of the response of the neural system to speech input. Analyzing EEG signals predicted on the basis of the response functions might then bring to light conditionspecific differences in the filtered signals. We validate this approach by means of an analysis of EEG signals recorded with isolated word stimuli. Then, we apply the validated method to the analysis of the responses to the same words in the middle of meaningful sentences. Kimberley Mulder, Louis ten Bosch, Lou Boves |
INTERSPEECH | 2 |
| 2018 | Implementing DIANA to Model Isolated Auditory Word Recognition in EnglishabstractDIANA, an end-to-end computational model of spoken word recognition, was previously used to simulate auditory lexical decision experiments in Dutch.A single test conducted for North American English showed promising results as well.However, this simulation used a relatively small amount of data collected in the pilot phase of the Massive Auditory Lexical Decision (MALD) project.Additionally, already existing acoustic models were implemented.In this paper, we expand the analysis of MALD data by including a larger sample of both stimuli and participants.Acknowledging that most speech humans hear is conversational speech, we also test new acoustic models created using spontaneous speech corpora.Simulations successfully replicate expected trends in word competition and show plausible competitors as the signal unfolds, but acoustic model accuracy should be improved.Despite the number of responses per word being relatively small (never more than five), correlations between model estimates and participants' responses are moderate.Future directions in acoustic model training and simulating MALD data are discussed. Filip Nenadic, Louis ten Bosch, Benjamin V. Tucker |
INTERSPEECH | 2 |
| 2017 | The Recognition of Compounds: A Computational AccountabstractThis paper investigates the processes in comprehending spoken noun-noun compounds, using data from the BALDEY database. BALDEY contains lexicality judgments and reaction times (RTs) for Dutch stimuli for which also linguistic information is included. Two different approaches are combined. The first is based on regression by Dynamic Survival Analysis, which models decisions and RTs as a consequence of the fact that a cumulative density function exceeds some threshold. The parameters of that function are estimated from the observed RT data. The second approach is based on DIANA, a process-oriented computational model of human word comprehension, which simulates the comprehension process with the acoustic stimulus as input. DIANA gives the identity and the number of the word candidates that are activated at each 10 ms time step. Both approaches show how the processes involved in comprehending compounds change during a stimulus. Survival Analysis shows that the impact of word duration varies during the course of a stimulus. The density of word and non-word hypotheses in DIANA shows a corresponding pattern with different regimes. We show how the approaches complement each other, and discuss additional ways in which data and process models can be combined. Louis ten Bosch, Lou Boves, Mirjam Ernestus |
INTERSPEECH | 1 |
| 2016 | Combining Data-Oriented and Process-Oriented Approaches to Modeling Reaction Time DataabstractThis paper combines two different approaches to modeling reaction time data from lexical decision experiments, viz. a dataoriented statistical analysis by means of a linear mixed effects model, and a process-oriented computational model of human speech comprehension. The linear mixed effect model is implemented by lmer in R. As computational model we apply DIANA, an end-to-end computational model which aims at modeling the cognitive processes underlying speech comprehension. DIANA takes as input the speech signal, and provides as output the orthographic transcription of the stimulus, a word/non-word judgment and the associated reaction time. Previous studies have shown that DIANA shows good results for large-scale lexical decision experiments in Dutch and North-American English. We investigate whether predictors that appear significant in an lmer analysis and processes implemented in DIANA can be related and inform both approaches. Predictors such as ‘previous reaction time’ can be related to a process description; other predictors, such as ‘lexical neighborhood’ are hard-coded in lmer and emergent in DIANA. The analysis focuses on the interaction between subject variables and task variables in lmer, and the ways in which these interactions can be implemented in DIANA. Louis ten Bosch, Lou Boves, Mirjam Ernestus |
INTERSPEECH | 1 |
| 2016 | Analytical Assessment of Dual-Stream Merging for Noise-Robust ASRabstract\n Contains fulltext :\n 166330.pdf (Publisher’s version ) (Open Access)\n Louis ten Bosch, Bert Cranen |
INTERSPEECH | 1 |
| 2016 | Comparing Different Methods for Analyzing ERP SignalsabstractEvent-Related Potential (ERP) signals obtained from EEG recordings are widely used for studying cognitive processes in spoken language processing.The computation of ERPs involves averaging over multiple participants and multiple stimuli.Especially with speech stimuli, which also evoke substantial exogenous excitation, even averaging within conditions results in pooling many sources of variance.This raises questions about the statistical processing needed to uncover reliable differences between conditions.In this study we investigate differences between ERPs when participants listened to full and reduced pronunciations of verb forms in Dutch, in isolation and in mid-sentence position.Conventional statistical analysis uncovers some (but not all) differences between full and reduced forms in isolation, but not in mid-sentence position.In this paper, we show that linear mixed models (lmer) and generalized additive models (gam), which are able to account for participant-and stimulus-related variance may uncover more effects than conventional statistical models.However, depending on the complexity of the data, lmer and gam models may not be able to fit the data closely enough to warrant blind interpretation of the summary output.We discuss opportunities and threats of these approaches to analyzing ERP signals. Kimberley Mulder, Louis ten Bosch, Lou Boves |
INTERSPEECH | 2 |
| 2016 | Locally learning heterogeneous manifolds for phonetic classificationabstractMost state-of-the-art phone classifiers use the same features and decision criteria for all phones, despite the fact that different broad classes are characterized by different manners and place of articulation that result in different acoustic features. This paper uses manifold learning to address structure in the acoustic space. Previous approaches to dimensionality reduction based on manifold learning assumed that the acoustic space can be characterized by a uniform manifold structure. In this paper we relax this assumption by learning different manifold structures for broad phonetic classes. Because all known classifiers make confusions between broad classes, we designed a two-level classifier in which the top level consists of a number of partially overlapping broad classes. Since the resulting classifiers are not statistically independent, we propose a new method for fusing the classifiers. Experimental results show that our two-level classifier obtained slightly better results when broad-class specific manifolds were learned, compared to a uniform manifold. However, the accuracy is still considerably lower than what could be obtained with oracle knowledge about broad class membership. From this we infer that phones do not form compact clusters in acoustic space. Heyun Huang, Yang Liu 0007, Louis ten Bosch, Bert Cranen, Lou Boves |
Comput. Speech Lang. | 3 |
| 2016 | Human-inspired modulation frequency features for noise-robust ASRabstractThis paper investigates a computational model that combines a frontend based on an auditory model with an exemplar-based sparse coding procedure for estimating the posterior probabilities of sub-word units when processing noisified speech. Envelope modulation spectrogram (EMS) features are extracted using an auditory model which decomposes the envelopes of the outputs of a bank of gammatone filters into one lowpass and multiple bandpass components. Through a systematic analysis of the configuration of the modulation filterbank , we investigate how and why different configurations affect the posterior probabilities of sub-word units by measuring the recognition accuracy on a semantics-free speech recognition task. Our main finding is that representing speech signal dynamics by means of multiple bandpass filters typically improves recognition accuracy. This effect is particularly noticeable in very noisy conditions. In addition we find that to have maximum noise robustness, the bandpass filters should focus on low modulation frequencies . This reenforces our intuition that noise robustness can be increased by exploiting redundancy in those frequency channels which have long enough integration time not to suffer from envelope modulations that are solely due to noise. The ASR system we design based on these findings behaves more similar to human recognition of noisified digit strings than conventional ASR systems. Thanks to the relation between the modulation filterbank and procedures for computing dynamic acoustic features in conventional ASR systems, the finding can be used for improving the frontends in those systems. Sara Ahmadi, Bert Cranen, Lou Boves, Louis ten Bosch, Antal van den Bosch |
Speech Commun. | 4 |
| 2016 | Phone classification via manifold learning based dimensionality reduction algorithms
Heyun Huang, Louis ten Bosch, Bert Cranen, Lou Boves |
Speech Commun. | 2 |
| 2015 | Unconstrained Speech Segmentation using Deep Neural Networks
Van Zyl van Vuuren, Louis ten Bosch, Thomas Niesler |
ICPRAM (1) | 2 |
| 2015 | DIANA: towards computational modeling reaction times in lexical decision in north American EnglishabstractDIANA is an end-to-end computational model of speech processing, which takes as input the speech signal, and provides as output the orthographic transcription of the stimulus, a word/non-word judgment and the associated estimated reaction time. So far, the model has only been tested for Dutch. In this paper, we extend DIANA such that it can also process North American English. The model is tested by having it simulate human participants in a large scale North American English lexical decision experiment. The simulations show that DIANA can adequately approximate the reaction times of an average participant (r = 0.45). In addition, they indicate that DIANA does not yet adequately model the cognitive processes that take place after stimulus offset. Louis ten Bosch, Lou Boves, Benjamin V. Tucker, Mirjam Ernestus |
INTERSPEECH | 1 |
| 2014 | Comparing reaction time sequences from human participants and computational modelsabstractThis paper addresses the question how to compare reaction times computed by a computational model of speech comprehension with observed reaction times by participants. The question is based on the observation that reaction time sequences substantially differ per participant, which raises the issue of how exactly the model is to be assessed. Part of the variation in reaction time sequences is caused by the so-called local speed: the current reaction time correlates to some extent with a number of previous reaction times, due to slowly varying variations in attention, fatigue etc. This paper proposes a method, based on time series analysis, to filter the observed reaction times in order to separate the local speed effects. Results show that after such filtering the between-participant correlations increase as well as the average correlation between participant and model increases. The presented technique provides insights into relevant aspects that are to be taken into account when comparing reaction time sequences Louis ten Bosch, Mirjam Ernestus, Lou Boves |
INTERSPEECH | 1 |
| 2014 | Fusion of parametric and non-parametric approaches to noise-robust ASR
Jort F. Gemmeke, Bert Cranen, Louis ten Bosch, Lou Boves |
Speech Commun. | 4 |
| 2013 | Towards an end-to-end computational model of speech comprehension: simulating a lexical decision taskabstractThis paper describes a computational model of speech comprehension that takes the acoustic signal as input and predicts reaction times as observed in an auditory lexical decision task. By doing so, we explore a new generation of end-to-end computational models that are able to simulate the behaviour of human subjects participating in a psycholinguistic experiment. So far, nearly all computational models of speech comprehension do not start from the speech signal itself, but from abstract representations of the speech signal, while the few existing models that do start from the acoustic signal cannot directly model reaction times as obtained in comprehension experiments. The main functional components in our model are the perception stage, which is compatible with the psycholinguistic model Shortlist B and is implemented with techniques from automatic speech recognition, and the decision stage, which is based on the linear ballistic accumulation decision model. We successfully tested our model against data from 20 participants performing a largescale auditory lexical decision experiment. Analyses show that the model is a good predictor for the average judgment and reaction time for each word. Louis ten Bosch, Lou Boves, Mirjam Ernestus |
INTERSPEECH | 1 |
| 2013 | Quantifying cross-linguistic variation in grapheme-to-phoneme mappingabstractContains fulltext : 119399.pdf (Publisher’s version ) (Open Access) Martine Coene, Annemiek Hammer, Wojtek Kowalczyk, Louis ten Bosch, Bart Vaerenberg, Paul Govaerts |
INTERSPEECH | 4 |
| 2013 | Word identification using phonetic features: towards a method to support multivariate fMRI speech decodingabstractContains fulltext : 119391.pdf (Publisher’s version ) (Open Access) Tijl Grootswagers, Karen Dijkstra, Louis ten Bosch, Alex Brandmeyer, Makiko Sadakata |
INTERSPEECH | 3 |
| 2013 | Balancing word lists in speech audiometry through large spoken language corporaabstractContains fulltext : 119393.pdf (Publisher’s version ) (Open Access) Annemiek Hammer, Bart Vaerenberg, Wojtek Kowalczyk, Louis ten Bosch, Martine Coene, Paul Govaerts |
INTERSPEECH | 4 |
| 2013 | Training log-linear acoustic models in higher-order polynomial feature space for speech recognitionabstractThe use of higher-order polynomial acoustic features can improve the performance of automatic speech recognition.However, the dimensionality of the polynomial representation can be prohibitively large, making the training of acoustic models using polynomial features for large vocabulary ASR systems infeasible.This paper presents an iterative polynomial training framework for acoustic modeling, which recursively expands the current acoustic features into their second-order polynomial feature space.In each recursion the dimensionality is reduced by a linear projection, such that increasingly higher order polynomial information is incorporated while keeping the dimensionality of the acoustic models constant.Experimental results obtained for a large-vocabulary continuous speech recognition task show that the proposed method outperforms conventional mixture models. Muhammad Ali Tahir, Heyun Huang, Ralf Schlüter, Hermann Ney, Louis ten Bosch, Bert Cranen, Lou Boves |
INTERSPEECH | 5 |
| 2013 | Language-universal speech audiometry with automated scoringabstractContains fulltext : 119396.pdf (Publisher’s version ) (Open Access) Bart Vaerenberg, Louis ten Bosch, Wojtek Kowalczyk, Martine Coene, Herwig De Smet, Paul Govaerts |
INTERSPEECH | 2 |
| 2013 | Detecting words in speech using linear separability in a bag-of-events vector spaceabstractContains fulltext : 119395.pdf (Publisher’s version ) (Open Access) Maarten Versteegh, Louis ten Bosch |
INTERSPEECH | 2 |
| 2013 | A dynamic programming framework for neural network-based automatic speech segmentation
Van Zyl van Vuuren, Louis ten Bosch, Thomas Niesler |
INTERSPEECH | 2 |
| 2012 | Knowledge-based Quadratic Discriminant Analysis for phonetic classificationabstractModeling the second-order statistics of articulatory trajectories is likely to improve the performance in classifying phone segments compared to using only linear combinations of MFCCs. Nevertheless, the extremely high dimensionality of the feature space spanned by a combination of monomials of degree-1 and degree-2 makes it difficult to effectively exploit the discriminative information in the full covariance matrix. This paper proposes a novel algorithm, dubbed Knowledge-based Quadratic Discriminant Analysis (KnQDA), for reducing the number of dimensions of the space spanned by degree-1 and degree-2 monomials by using phonetic knowledge for selecting the set of degree-2 monomials that are most likely to improve classification. KnQDA seeks a trade-off between overfitting and undertraining, which further improves the learnability. Binary classifications on all pairs of phones in TIMIT show the effectiveness of the proposed method, especially on those phone pairs that overlap strongly in the linear feature space. Heyun Huang, Yang Liu 0007, Louis ten Bosch, Bert Cranen, Lou Boves |
ICASSP | 3 |
| 2012 | Modeling Cue Trading in Human Word RecognitionabstractClassical phonetic studies have shown that acoustic-articulatory cues can be interchanged without affecting the resulting phoneme percept (‘cue trading’). Cue trading has so far mainly been investigated in the context of phoneme identification. In this study, we investigate cue trading during recognition of words, the units of speech through which we communicate. This paper aims to provide a method to quantify cue trading effects by using a computational model of human word recognition. This model takes the acoustic signal as input and represents speech using articulatory feature streams. Importantly, it allows cue trading and underspecification. Its set-up is inspired by the functionality of Fine-Tracker, a recent computational model of human word recognition. This approach makes it possible, for the first time, to quantify cue trading in terms of a trade-off between features and to investigate cue trading in the context of a word recognition task. Index Terms: cue trading, human word recognition, computational modeling, articulatory features. Louis ten Bosch, Odette Scharenborg |
INTERSPEECH | 1 |
| 2012 | Exploring Discriminative Speech Trajectory Structures
Heyun Huang, Louis ten Bosch, Bert Cranen, Lou Boves |
INTERSPEECH | 2 |
| 2012 | Using Sparse Classification Outputs as Feature Observations for Noise-robust ASRabstractContains fulltext : 102118.pdf (author's version ) (Open Access) Contains fulltext : 102118.pdf (Publisher’s version ) (Open Access) Bert Cranen, Jort F. Gemmeke, Louis ten Bosch, Lou Boves, Mathew Magimai-Doss |
INTERSPEECH | 4 |
| 2012 | Combination of Sparse Classification and Multilayer Perceptron for Noise-robust ASRabstractContains fulltext : 101559.pdf (Publisher’s version ) (Open Access) Mathew Magimai-Doss, Jort F. Gemmeke, Bert Cranen, Louis ten Bosch, Lou Boves |
INTERSPEECH | 5 |
| 2012 | Acoustic-phonetic and artificial neural network feature analysis to assess speech quality of stop consonants produced by patients treated for oral or oropharyngeal cancerabstractSpeech impairment often occurs in patients after treatment for head and neck cancer. A specific speech characteristic that influences intelligibility and speech quality is voice-onset-time (VOT) in stop consonants. VOT is one of the functionally most relevant parameters that distinguishes voiced and voiceless stops. The goal of the present study is to investigate the role and validity of acoustic-phonetic and artificial neural network analysis (ANN) of stop consonants in a multidimensional speech assessment protocol. Speech recordings of 51 patients 6 months after treatment for oral or oropharyngeal cancer and of 18 control speakers were evaluated by trained speech pathologists regarding intelligibility and articulation. Acoustic-phonetic analyses and artificial neural network analysis of the phonological feature voicing were performed in voiced /b/, /d/ and voiceless /p/ and /t/. Results revealed that objective acoustic-phonetic analysis and feature analysis for /b, d, p/ distinguish between patients and controls. Within patients, /t, d/ distinguish for tumour location and tumour stage. Measurements of the phonological feature voicing in almost all consonants were significantly correlated with articulation and intelligibility, but not with self-evaluations. Overall, objective acoustic-phonetic and feature analyses of stop consonants are feasible and contribute to further development of a multidimensional speech quality assessment protocol. Marieke de Bruijn, Louis ten Bosch, Joop Kuik, Birgit I. Witte, Johannes A. Langendijk, C. René Leemans, Irma Verdonck-de Leeuw |
Speech Commun. | 2 |
| 2011 | Thresholding Word Activations for Response Scoring - Modelling Psycholinguistic DataabstractIn the present paper we investigate the effect of categorising raw behavioural data or computational model responses. In addition, the effect of averaging over stimuli from potentially different populations is assessed. To this end, we replicate studies on word learning and generalisation abilities using the ACORNS models. Our results show that discrete categories may obscure interesting phenomena in the continuous responses. For example, the finding that learning in the model saturates very early at a uniform high recognition accuracy only holds for categorical representations. Additionally, a large difference in the accuracy for individual words is obscured by averaging over all stimuli. Because different words behaved differently for different speakers, we could not identify a phonetic basis for the differences. Implications and new predictions for infant behaviour are discussed. Christina Bergmann, Louis ten Bosch, Lou Boves |
INTERSPEECH | 2 |
| 2011 | Assessing Acoustic Reduction: Exploiting Local Structure in SpeechabstractThis paper presents a method to quantify the spectral characteristics of reduction in speech.Hämäläinen et al. (2009) proposes a measure of spectral reduction which is able to predict a substantial amount of the variation in duration that linguistically motivated variables do not account for.In this paper, we continue studying acoustic reduction in speech by developing a new acoustic measure of reduction, based on local manifold structure in speech.We show that this measure yields significantly improved statistical models for predicting variation in duration. Louis ten Bosch, Annika Hämäläinen, Mirjam Ernestus |
INTERSPEECH | 1 |
| 2011 | Globality-Locality Consistent Discriminant Analysis for Phone ClassificationabstractContains fulltext : 94442.pdf (author's version ) (Open Access) Heyun Huang, Yang Liu 0007, Jort F. Gemmeke, Louis ten Bosch, Bert Cranen, Lou Boves |
INTERSPEECH | 4 |
| 2011 | Improvements of a Dual-Input DBN for Noise Robust ASRabstractContains fulltext : 94477.pdf (Publisher’s version ) (Open Access) Jort F. Gemmeke, Bert Cranen, Louis ten Bosch, Lou Boves |
INTERSPEECH | 4 |
| 2011 | Modelling Novelty Preference in Word LearningabstractThis paper investigates the effects of novel words on a cogni-tively plausible computational model of word learning. The model is first familiarized with a set of words, achieving high recognition scores and subsequently offered novel words for training. We show that the model is able to recognize the novel words as different from the previously seen words, based on a measure of novelty that we introduce. We then propose a pro-cedure analogous to novelty preference in infants. Results from simulations of word learning show that adding this procedure to our model speeds up training and helps the model attain higher recognition rates. Index Terms: language acquisition, word learning, computa-tional modelling Maarten Versteegh, Louis ten Bosch, Lou Boves |
INTERSPEECH | 2 |
| 2010 | Discovering an optimal set of minimally contrasting acoustic speech units: a point of focus for whole-word pattern matchingabstract\n Contains fulltext :\n 86046.pdf (Publisher’s version ) (Open Access)\n Guillaume Aimetti, Roger K. Moore, Louis ten Bosch |
INTERSPEECH | 3 |
| 2010 | Language acquisition and cross-modal associations: computational simulation of the result of infant studies
Louis ten Bosch, Lou Boves |
INTERSPEECH | 1 |
| 2010 | Using a DBN to integrate sparse classification and GMM-based ASRabstractContains fulltext : 86189.pdf (author's version ) (Open Access) Jort F. Gemmeke, Bert Cranen, Louis ten Bosch, Lou Boves |
INTERSPEECH | 4 |
| 2010 | Active word learning under uncertain input conditionsabstractThis paper presents an analysis of phoneme durations of emotional speech in two languages: Dutch and Korean. The analyzed corpus of emotional speech has been specifically developed for the purpose of cross-linguistic comparison, and is more balanced than any similar corpus available so far: a) it contains expressions by both Dutch and Korean actors and is based on judgments by both Dutch and Korean listeners; b) the same elicitation technique and recording procedure were used for recordings of both languages; and c) the phonetics of the carrier phrase were constructed to be permissible in both languages. The carefully controlled phonetic content of the carrier phrase allows for analysis of the role of specific phonetic features, such as phoneme duration, in emotional expression in Dutch and Korean. In this study the mutual effect of language and emotion on phoneme duration is presented. Maarten Versteegh, Louis ten Bosch, Lou Boves |
INTERSPEECH | 2 |
| 2010 | A Speech Corpus for Modeling Language Acquisition: CAREGIVER
Toomas Altosaar, Louis ten Bosch, Guillaume Aimetti, Christos Koniaris, Kris Demuynck, Henk van den Heuvel |
LREC | 2 |
| 2009 | Discovering keywords from cross-modal input: ecological vs. engineering methods for enhancing acoustic repetitionsabstract\n Contains fulltext :\n 76399.pdf (author's version ) (Open Access)\n Guillaume Aimetti, Roger K. Moore, Louis ten Bosch, Okko Johannes Räsänen, Unto K. Laine |
INTERSPEECH | 3 |
| 2009 | Do multiple caregivers speed up language acquisition?abstractIn this paper we compare three different implementations of language learning to investigate the issue of speaker-dependent initial representations and subsequent generalization. These implementations are used in a comprehensive model of lan-guage acquisition under development in the FP6 FET project ACORNS. All algorithms are embedded in a cognitively and ecologically plausible framework, and perform the task of de-tecting word-like units without any lexical, phonetic, or phono-logical information. The results show that the computational approaches differ with respect to the extent they deal with un-seen speakers, and how generalization depends on the variation observed during training. Index Terms: Language acquisition, Computational modeling 1. Louis ten Bosch, Okko Johannes Räsänen, Joris Driesen, Guillaume Aimetti, Toomas Altosaar, Lou Boves, A. Corns |
INTERSPEECH | 1 |
| 2009 | Adaptive non-negative matrix factorization in a computational model of language acquisitionabstractDuring the early stages of language acquisition, young infants face the task of learning a basic vocabulary without the aid of prior linguistic knowledge. It is believed the long term episodic memory plays an important role in this process. Experiments have shown that infants retain large amounts of very detailed episodic information about the speech they perceive (e.g. [1]). This weakly justifies the fact that some algorithms attempting to model the process of vocabulary acquisition computationally process large amounts of speech data in batch. Non-negative Matrix Factorization (NMF), a technique that is particularly successful in data mining but can also be applied to vocabulary acquisition (e.g. [2]), is such an algorithm. In this paper, we will integrate an adaptive variant of NMF into a computational framework for vocabulary acquisition, foregoing the need for long term storage of speech inputs, and experimentally show its accuracy matches that of the original batch algorithm. Copyright © 2009 ISCA. Joris Driesen, Louis ten Bosch, Hugo Van hamme |
INTERSPEECH | 2 |
| 2009 | Modelling vocabulary growth from birth to young adulthoodabstractThere has been considerable debate over the existence of the ‘vocabulary spurt ’ phenomenon- an apparent acceleration in word learning that is commonly said to occur in children around the age of 18 months. This paper presents an investigation into modelling the phenomenon using data from almost 1800 children. The results indicate that the acquisition of a receptive/productive lexicon can be quite adequately modelled as a single growth function with an ecologically well founded and cognitively plausible interpretation. Hence it is concluded that there is little evidence for the vocabulary spurt phenomenon as a separable aspect of language acquisition. Index Terms: child language acquisition, lexical development, vocabulary size Roger K. Moore, Louis ten Bosch |
INTERSPEECH | 2 |
| 2009 | A Computational Model of Language Acquisition: the Emergence of WordsabstractIn this paper, we discuss a computational model that is able to detect and build word-like representations on the basis of sensory input. The model is designed and tested with a further aim to investigate how infants may learn to communicate by means of spoken language. The computational model makes use of a memory, a perception module, and the concept of 'learning drive'. Learning takes place within a communicative loop between a 'caregiver' and the 'learner'. Experiments carried out on three European languages with different genetic background (Finnish, Swedish, and Dutch) show that a robust word representation can be learned in using less than 100 acoustic tokens (examples) of that word. The model is inspired by the memory structure that is assumed functional for human cognitive processing. Louis ten Bosch, Lou Boves, Hugo Van hamme, Roger K. Moore |
Fundam. Informaticae | 1 |
| 2009 | Modelling pronunciation variation with single-path and multi-path syllable models: Issues to consider
Annika Hämäläinen, Louis ten Bosch, Lou Boves |
Speech Commun. | 2 |
| 2008 | A computational model of language acquisition: focus on word discoveryabstractYoung infants learn words by detecting patterns in the speech signal and by associating these patterns to stimuli presented by non-speech modalities (e.g vision). In this paper, we model this behaviour by designing and testing a computational model of word discovery. The model is able to build word-like representations on the basis of multimodal input data. The discovery of words (and word-like entities) takes place within a communicative loop between two protagonists, a ’carer ’ and the ’learner’. Experiments carried out on three different European languages (Finnish, Swedish, and Dutch) show that a robust word representation can be learned in using about 50 acoustic tokens (examples) of that word. The model is inspired by the memory structure that is assumed functional for human speech processing. Index Terms: language acquisition, unsupervised word detection, computational modelling 1. Louis ten Bosch, Hugo Van hamme, Lou Boves |
INTERSPEECH | 1 |
| 2008 | Phonetic-acoustic and feature analyses by a neural network to assess speech quality in patients treated for head and neck cancerabstractSubjective speech evaluation is the gold standard to assess speech quality of head and neck cancer patients. This study investigates if conventional acoustic-phonetic and novel feature analysis contribute to the development of a multidimensional speech assessment protocol. Speech recordings of 51 patients 6 months post-treatment and of 18 control speakers were subjectively evaluated for intelligibility, nasal resonance and articulation. Self-evaluation of speech problems was assessed by the EORTC QLQ-H&N35 speech subscale. Feature analysis was performed to assess objectively nasality in vowels /a,i,u/. Results revealed that size of the vowel triangle, pressure release of /k/ and nasality in /i/ predict best intelligibility, articulation and nasal resonance and differentiated best between patients and controls. Within patients, /k/ and /x/ differentiated tumour site and tumour classification. Various objective variables were related to speech problems as reported by patients. Marieke de Bruijn, Irma Verdonck-de Leeuw, Louis ten Bosch, Joop Kuik, Hugo Quené, Lou Boves, Johannes A. Langendijk, C. René Leemans |
INTERSPEECH | 3 |
| 2008 | Phonological representations in poor readers
Cecile T. L. Kuijpers, Louis ten Bosch |
INTERSPEECH | 2 |
| 2008 | Evaluating the Relationship between Linguistic and Geographic Distances using a 3D Visualization
Folkert de Vriend, Jan Pieter Kunst, Louis ten Bosch, Charlotte Giesbers, Roeland van Hout |
LREC | 3 |
| 2007 | Modelling Pronunciation Variation using Multi-Path HMMS for SyllablesabstractRecent research suggests that it is more appropriate to model pronunciation variation with syllable-length acoustic models than with triphones. Due to the large number of factors contributing to pronunciation variation at the syllable level, the creation of multi-path model topologies appears necessary. In this paper, we construct multi-path models using phonetic knowledge to initialise the parallel paths, and a data-driven solution for their reestimation. When applied to 94 frequent syllables in a Dutch read speech recognition task, the approach leads to improved recognition performance when compared with a much more complex triphone recogniser. A detailed analysis of the pronunciation variation captured by the parallel paths pinpoints the deficiencies of the approach, and provides insights into how these may be overcome. Annika Hämäläinen, Louis ten Bosch, Lou Boves |
ICASSP (4) | 2 |
| 2007 | A computational model for unsupervised word discoveryabstractWe present an unsupervised algorithm for the discovery of words and word-like fragments from the speech signal, without using an upfront defined lexicon or acoustic phone models.The algorithm is based on a combination of acoustic pattern discovery, clustering, and temporal sequence learning.It exploits the acoustic similarity between multiple acoustic tokens of the same words or word-like fragments.In its current form, the algorithm is able to discover words in speech with low perplexity (connected digits).Although its performance still falls off compared to mainstream ASR approaches, the value of the algorithm is its potential to serve as a computational model in two research directions.First, the algorithm may lead to an approach for speech recognition that is fundamentally liberated from the modelling constraints in conventional ASR.Second, the proposed algorithm can be interpreted as a computational model of language acquisition that takes actual speech as input and is able to find words as 'emergent' properties from raw input. Louis ten Bosch, Bert Cranen |
INTERSPEECH | 1 |
| 2007 | Construction and analysis of multiple paths in syllable modelsabstractIn this paper, we construct multi-path syllable models using phonetic knowledge for initialising the parallel paths, and a data-driven solution for their re-estimation. We hypothesise that the richer topology of multi-path syllable models would be better at accounting for pronunciation variation than context-dependent phone models that can only account for the effects of left and right neighbours. We show that parallel paths that are initialised with phonetic knowledge and then reestimated do indeed result in different trajectories in feature space. Yet, this does not result in better recognition performance. We suggest explanations for this finding, and provide the reader with important insights into the issues playing a role in pronunciation variation modelling with multi-path syllable models. Index Terms: speech recognition, hidden Markov models, multi-path syllable models, Kullback-Leibler distance, Annika Hämäläinen, Louis ten Bosch, Lou Boves |
INTERSPEECH | 2 |
| 2007 | Speech quality after major surgery of the oral cavity and oropharynx with microvascular soft tissue reconstructionabstractSpeech quality of patients with oral or oropharyngeal carcinoma was assessed by perceptual and acoustic-phonetic analyses. Speech recordings of running speech of patients before and 6 and 12 months after treatment for oral or oropharyngeal cancer and of 18 control speakers were evaluated regarding intelligibility, nasality and articulation, which revealed deteriorated speech in 20% of the patients before treatment, and in 75% 6-12 months after treatment. Acoustic analyses comprised formant, duration, perturbation and noise measures of the vowels /i/, /a/, and /u/ and were performed on the speech samples 6 months after treatment and the controls. Patients appeared to have a smaller vowel space compared to controls, which was clearly related to speech intelligibility. Furthermore, voice perturbation appeared to be higher in patients. Although oropharyngeal treatment does not effect the function of the larynx itself, the acoustic coupling between source and filter may effect the smoothness of the voicing characteristics. The presented speech analyses may serve as part of an outcome measurement protocol for assessing efficacy of speech rehabilitation. Irma Verdonck-de Leeuw, Louis ten Bosch, Li Ying Chao, Rico N. P. M. Rinkel, Pepijn A. Borggreven, Lou Boves, C. René Leemans |
INTERSPEECH | 2 |
| 2007 | 'Early recognition' of polysyllabic words in continuous speech
Odette Scharenborg, Louis ten Bosch, Lou Boves |
Comput. Speech Lang. | 2 |
| 2007 | Bridging the gap between human and automatic speech recognition
Louis ten Bosch, Katrin Kirchhoff |
Speech Commun. | 1 |
| 2006 | Acoustic Scores and Symbolic Mismatch Penalties in Phone LatticesabstractThis paper builds on previous work aimed at unraveling the structure of the speech signal using probabilistic representations. The context of this work is a multi-pass speech recognition system in which a phone lattice is created and used as a basis for a lexical decoding pass (search) that allows symbolic mismatches at certain costs. The focus is on the optimization of the costs of the phone insertions, deletions and substitutions that are used in the lexical decoding pass. Two optimization approaches are presented, one related to a multi-pass computational model for human speech recognition, the other based on a decoding that minimizes Bayes' risks. In the final section, the advantages of the two optimization methods are discussed and compared Louis ten Bosch, Annika Hämäläinen, Odette Scharenborg, Lou Boves |
ICASSP (1) | 1 |
| 2006 | On speech variation and word type differentiation by articulatory feature representationsabstractThis paper describes ongoing research aiming at the description of variation in speech as represented by asynchronous articulatory features. We will first illustrate how distances in the articulatory feature space can be used for event detection along speech trajectories in this space. The temporal structure imposed by the cosine distance in articulatory feature space coincides to a large extent with the manual segmentation on phone level. The analysis also indicates that the articulatory feature representation provides better such alignments than the MFCC representation does. Secondly, we will present first results that indicate that articulatory features can be used to probe for acoustic differences in the onsets of Dutch singulars and plurals. Louis ten Bosch, R. Harald Baayen, Mirjam Ernestus |
INTERSPEECH | 1 |
| 2006 | Pronunciation variant-based multi-path HMMs for syllablesabstractRecent research suggests that it is more appropriate to model pronunciation variation with syllable-length acoustic models than with context-dependent phones. Due to the large number of factors contributing to pronunciation variation at the syllable level, the creation of multi-path model topologies appears necessary. In this paper, we propose a novel approach for constructing multi-path models for frequent syllables. The suggested approach uses phonetic knowledge for the initialisation of the parallel paths, and a data-driven solution for their re-estimation. When applied to 94 frequent syllables in a 37-hour corpus of Dutch read speech, it leads to improved recognition performance when compared with a triphone recogniser of similar complexity. Annika Hämäläinen, Louis ten Bosch, Lou Boves |
INTERSPEECH | 2 |
| 2005 | Improving out-of-coverage language modelling in a multimodal dialogue system using small training setsabstractFor automatic speech recognition, the construction of an adequate language model may be difficult when only a limited amount of training text is available. Previous work has shown that in the case of small training sets statistical language models may outperform grammars on out-of-coverage utterances, while showing comparable performance on incoverage input. In this paper, we compare the performance of an automatic speech recognition system using a grammar and a statistical language model including garbage models in the case of very limited in-domain training data. The results show that a bigram language model and a grammar show similar performance, and that the inclusion of garbage models in statistical language models enhances their performance both on in-coverage and out-of-coverage utterances. Louis ten Bosch |
INTERSPEECH | 1 |
| 2005 | ASR decoding in a computational model of human word recognitionabstractRecently, a computational model of human word recognition, called SpeM, has been developed. In contrast to most current models of human word recognition, SpeM is able to process actual acoustic speech input, and decodes the incoming speech stream into lexical and non-lexical items. This model makes the links between HSR and ASR as explicit as possible. In this paper, we focus on unravelling the structure of the complex search space that is used in SpeM and similar decoding strategies. To that end, it discusses a number of properties of phone lattices in relation to canonical phone representations. Furthermore, we elaborate on the close relation between distances in this search space, and distance measures in search spaces that are based on a combination of acoustic and phonetic features. 1. Louis ten Bosch, Odette Scharenborg |
INTERSPEECH | 1 |
| 2005 | On temporal aspects of turn taking in conversational dialogues
Louis ten Bosch, Nelleke Oostdijk, Lou Boves |
Speech Commun. | 1 |
| 2005 | Conversational agent or direct manipulation in human-system interaction
Els den Os, Lou Boves, Stéphane Rossignol, Louis ten Bosch, Louis Vuurpijl |
Speech Commun. | 4 |
| 2004 | Survey of spontaneous speech phenomena in a multimodal dialogue system and some implications for ASRabstractAudio recordings of speakers using speech-driven systems show phenomena that are characteristic for on-line speech responses, such as out-of-task utterances, self-talk and speech disfluencies. This paper focuses on a survey of these phenomena as they were recorded during interactions by subjects using a multimodal system, and reports on experiments concerning the treatment of these phenomena for automatic speech recognition. This study is a starting point for the study of a richer set of on-line phenomena in speech addressed to multimodal systems and the implications for automatic speech recognition. Louis ten Bosch, Lou Boves |
INTERSPEECH | 1 |
| 2003 | Learning rule ranking by dynamic construction of context-free grammars using AND/OR graphs
Anna Corazza, Louis ten Bosch |
INTERSPEECH | 2 |
| 2003 | Recognising 'real-life' speech with spem: a speech-based computational model of human speech recognitionabstractIn this paper, we present a novel computational model of human speech recognition - called SpeM - based on the theory underlying Shortlist. We will show that SpeM, in combination with an automatic phone recogniser (APR), is able to simulate the human speech recognition process from the acoustic signal to the ultimate recognition of words. This joint model takes an acoustic speech file as input and calculates the activation flows of candidate words on the basis of the degree of fit of the candidate words with the input.\n\nExperiments showed that SpeM outperforms Shortlist on the recognition of 'real-life' input. Furthermore, SpeM performs only slightly worse than an off-the-shelf full-blown automatic speech recogniser in which all words are equally probable, while it provides a transparent computationally elegant paradigm for modelling word activations in human word recognition. Odette Scharenborg, Louis ten Bosch, Lou Boves |
INTERSPEECH | 2 |
| 2003 | Modelling human speech recognition using automatic speech recognition paradigms in speMabstractThe following full text is an author's version which may differ from the publisher's version. Odette Scharenborg, James M. McQueen, Louis ten Bosch, Dennis Norris |
INTERSPEECH | 3 |
| 2003 | Emotions, speech and the ASR framework
Louis ten Bosch |
Speech Commun. | 1 |
| 2002 | Probabilistic ranking of constraints
Louis ten Bosch |
INTERSPEECH | 1 |
| 2001 | Pronunciation modeling and lexical adaptation in midsize vocabulary ASR
Louis ten Bosch, Nick Cremelie |
INTERSPEECH | 1 |
| 2001 | Up to what level can acoustical and textual features predict prominenceabstractIn this study acoustical as well as lexical/syntactic correlates of prominence are analyzed and discussed.Prominence is defined at the word level and is based on listener judgments.Spoken sentences from many different speakers, taken from the Dutch Polyphone corpus of telephone speech, are analyzed.A selection of useful acoustical input features is chosen for classification of word prominence, by means of Feed Forward Nets.For an independent test set of 1,000 sentences about 79% of the words are correctly classified as prominent or not.We also developed an algorithm, based on text input, using linguistic/syntactical features derived from text only, to predict prominence.The prediction agrees with the perceived prominence in 81% of the cases for the independent test set.The results of this project show that certain acoustical and linguistic correlates of prominence can be extracted automatically and can be used to accurately predict prominence with a consistency similar to the pominence assignment by naive listeners. Barbertje M. Streefkerk, Louis C. W. Pols, Louis ten Bosch |
INTERSPEECH | 3 |
| 2000 | ASR, dialects, and acoustic/phonological distances
Louis ten Bosch |
INTERSPEECH | 1 |
| 1999 | Acoustical features as predictors for prominence in read aloud dutch sentences used in ANN's
Barbertje M. Streefkerk, Louis C. W. Pols, Louis ten Bosch |
EUROSPEECH | 3 |
| 1998 | Automatic detection of prominence (as defined by listeners' judgements) in read aloud dutch sentencesabstractThis paper describes a first step towards the automatic classification of prominence (as defined by native listeners). As a result of a listening experiment each word in 500 sentences was marked with a rating scale between `0' (non-prominent) and `10' (very prominent). These prominence labels are compared with the following acoustical features: loudness of each vowel, and F0 range and duration of each syllable. A linear relationship between the rating scale of prominence and these acoustical features is found. These acoustical features then are used for a preliminary automatic classification to predict prominence. Barbertje M. Streefkerk, Louis C. W. Pols, Louis ten Bosch |
ICSLP | 3 |
| 1998 | A novel feature transformation for vocal tract length normalization in automatic speech recognitionabstractThis paper proposes a method to transform acoustic models that have been trained with a certain group of speakers for use on different speech in hidden Markov model based (HMM-based) automatic speech recognition. Features are transformed on the basis of assumptions regarding the difference in vocal tract length between the groups of speakers. First, the vocal tract length (VTL) of these groups has been estimated based on the average third formant F/sub 3/. Second, the linear acoustic theory of speech production has been applied to warp the spectral characteristics of the existing models so as to match the incoming speech. The mapping is composed of subsequent nonlinear submappings. By locally linearizing it and comparing results in the output, a linear approximation for the exact mapping was obtained which is accurate as long as the warping is reasonably small. The feature vector, which is computed from a speech frame, consists of the mel scale cepstral coefficients (MFCC) along with delta and delta/sup 2/-cepstra as well as delta and delta/sup 2/ energy. The method has been tested for TI digits data base, containing adult and children speech, consisting of isolated digits and digit strings of different length. The word error rate when trained on adults and tested on children with transformed adult models is decreased by more than a factor of two compared to the nontransformed case. Tom Claes, Ioannis Dologlou, Louis ten Bosch, Dirk Van Compernolle |
IEEE Trans. Speech Audio Process. | 3 |
| 1997 | New transformations of cepstral parameters for automatic vocal tract length normalization in speech recognitionabstractThis paper proposes a method to transform acoustic models (HMM gaussian mixtures) that have been trained on a certain group of speakers for use on speech from a different group of speakers. Cepstral features are transformed on the basis of assumptions regarding the difference in vocal tract length (VTL) between the groups of speakers (VTL normalisation, VTLN). Firstly, the VTL of these groups has been estimated based on the average third formant F . Secondly, the linear acoustic theory of speech production has been applied to warp the spectral characteristics of the existing models so as to match the incoming speech. The mapping is composed of subsequent non-linear submappings. By locally linearizing it, a linear approximation was obtained which is accurate as long as warping is reasonably small. The method has been tested for the TI digits database, containing adult and kids speech, consisting of isolated digits and digit strings of different length. The word error rate when trained on adults and tested on kids with transformed adult models is decreased by more than a factor of 2 compared to the non-transformed case. Tom Claes, Ioannis Dologlou, Louis ten Bosch, Dirk Van Compernolle |
EUROSPEECH | 3 |
| 1996 | On the error criteria in neural networks as a tool for human classification modelling
Louis ten Bosch, Roel Smits |
ICSLP | 1 |
| 1996 | Integration of context-dependent durational knowledge into HMM-based speech recognitionabstractHMMThis paper presents research on integrating context-dependent durational knowledge into HMM-based speech recognition.The first part of the paper presents work on obtaining relations between the parameters of the context-free HMMs and their durational behaviour, in preparation for the context-dependent durational modelling presented in the second part.Duration integration is realised via rescoring in the post-processing step of our N-best monophone recogniser.We use the multi-speaker TIMIT database for our analyses.The single-state dpdf (which is geometrical) is less important than the dpdf of the whole HMM, because in actual practice it is the latter that models a phonetic segment.In this section we firstly derive the closed-form whole-model dpdf for general leftto-right HMM.Left-to-right is by far the most common type of transition topology used for speech recognition.It may include any number of skipping transitions and parallel paths but no feedback loops that contains more than one state.Then an analysis will be given of the general properties of the dpdf, with the help of some examples of useful topologies. Louis ten Bosch, Louis C. W. Pols |
ICSLP | 2 |
| 1996 | Analysis of context-dependent segmental duration for automatic speech recognition
Louis C. W. Pols, Louis ten Bosch |
ICSLP | 3 |
| 1996 | Modelling of phone duration (using the TIMIT database) and its potential benefit for ASR
Louis C. W. Pols, Louis ten Bosch |
Speech Commun. | 3 |
| 1993 | On the automatic classification of pitch movementsabstractIn this paper, we discuss the construction of an algorithm that classifies pitch movements according to the IPO intonation system. We use a pitch stylization technique in order to obtain a continuous pitch contour over time, and a multi-linear classifier for the actual classification. In speaker-independent tests on a corpus of speech read by non-professionals, up to 81 % of the 279 pitch movements in the test corpus were correctly classified. These results are obtained by using information from the sampled speech data files only; a grammar will be used in the second stage of this study. Louis ten Bosch |
EUROSPEECH | 1 |
| 1993 | Impact of dimensionality and correlation of observation vectors in HMM-based speech recognition
Louis ten Bosch, Louis C. W. Pols |
EUROSPEECH | 2 |
| 1989 | From diphones to allophones: from data to rules
Louis ten Bosch, René Collier, Lou Boves |
EUROSPEECH | 1 |