EDBT 2026 Demo / reviewers in the wild / expert
Kris Demuynck
dblp:49/603
· DBLP profile ↗
92ranked-venue papers
19as first author
16since 2021 · last 2025
0000-0001-8525-7160ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 75 · 17 first-author · 13 since 2021Artificial intelligence and machine learning · 55 · 13 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term DetectionabstractQuery-by-example spoken term detection (QbE-STD) is often hindered by reliance on frame-level features and the computationally intensive DTW-based template matching, limiting its practicality. To address these challenges, we propose a novel approach that encodes speech into discrete, speaker-agnostic semantic tokens. This facilitates fast retrieval using text-based search algorithms and effectively handles out-of-vocabulary terms. Our approach focuses on generating consistent token sequences across varying utterances of the same term. We also propose a bidirectional state space modeling within the Mamba encoder, trained in a self-supervised learning framework, to learn contextual frame-level features that are further encoded into discrete tokens. Our analysis shows that our speech tokens exhibit greater speaker invariance than those from existing tokenizers, making them more suitable for QbESTD tasks. Empirical evaluation on LibriSpeech and TIMIT databases indicates that our method outperforms existing baselines while being more efficient. Anup Singh, Kris Demuynck, Vipul Arora 0001 |
ICASSP | 2 |
| 2025 | Weakly Supervised Phonological Features for Pathological Speech AnalysisabstractParalinguistic properties of speech are essential in analyzing and choosing optimal treatment options for patients with speech disorders. However, automatic modeling of these characteristics is difficult due to the lack of labeled speech datasets describing paralinguistic properties, especially at the frame-level. In this paper, we propose a weakly supervised training method which exploits the known acoustic properties of phonemes by training an ASR model with an interpretable frame-level phonological feature bottleneck layer. Subsequently, we assess the viability of these phonological features in speech pathology analysis by developing corresponding models for intelligibility prediction and speech pathology classification. Models using our proposed phonological features perform similar to other state-of-the-art acoustic features on both tasks with a classification accuracy of 75% and a 8.43 RMSE on speech intelligibility prediction. In contrast to others, our phonological features are text-independent and highly interpretable, providing potentially useful insights for speech therapists. Jenthe Thienpondt, Geoffroy Vanderreydt, Abdessalem Hammami, Kris Demuynck |
ICASSP | 4 |
| 2025 | Language-Agnostic Speech Tokenizer for Spoken Term Detection with Efficient RetrievalabstractThe surge in multilingual and code-switched spoken content demands efficient Query-by-Example Spoken Term Detection (STD) systems capable of handling diverse languages. Existing STD systems are monolingual; they typically require large labeled datasets for training and use costly DTW-based matching during inference, limiting their practicality. This paper proposes a novel speech tokenizer that converts speech into language-agnostic tokens. Furthermore, a multi-stage search algorithm enables fast and efficient retrieval from large datasets. In experimental evaluations, the tokens from the proposed tokenizer demonstrate strong speaker invariance, consistent performance across languages, and a capability to generalize effectively to unseen languages, outperforming the baselines significantly. Anup Singh, Kris Demuynck, Vipul Arora 0001 |
INTERSPEECH | 2 |
| 2024 | A novel channel estimate for noise robust speech recognition
Geoffroy Vanderreydt, Kris Demuynck |
Comput. Speech Lang. | 2 |
| 2024 | BioLORD-2023: semantic textual representations fusing large language models and clinical knowledge graph insightsabstractOBJECTIVE: In this study, we investigate the potential of large language models (LLMs) to complement biomedical knowledge graphs in the training of semantic models for the biomedical and clinical domains. MATERIALS AND METHODS: Drawing on the wealth of the Unified Medical Language System knowledge graph and harnessing cutting-edge LLMs, we propose a new state-of-the-art approach for obtaining high-fidelity representations of biomedical concepts and sentences, consisting of 3 steps: an improved contrastive learning phase, a novel self-distillation phase, and a weight averaging phase. RESULTS: Through rigorous evaluations of diverse downstream tasks, we demonstrate consistent and substantial improvements over the previous state of the art for semantic textual similarity (STS), biomedical concept representation (BCR), and clinically named entity linking, across 15+ datasets. Besides our new state-of-the-art biomedical model for English, we also distill and release a multilingual model compatible with 50+ languages and finetuned on 7 European languages. DISCUSSION: Many clinical pipelines can benefit from our latest models. Our new multilingual model enables a range of languages to benefit from our advancements in biomedical semantic representation learning, opening a new avenue for bioinformatics researchers around the world. As a result, we hope to see BioLORD-2023 becoming a precious tool for future biomedical applications. CONCLUSION: In this article, we introduced BioLORD-2023, a state-of-the-art model for STS and BCR designed for the clinical domain. François Remy, Kris Demuynck, Thomas Demeester |
J. Am. Medical Informatics Assoc. | 2 |
| 2024 | FlowHash: Accelerating Audio Search With Balanced Hashing via Normalizing FlowabstractNearest neighbor search on context representation vectors is a formidable task due to challenges posed by high dimensionality, scalability issues, and potential noise within query vectors. Our novel approach leverages normalizing flow within a self-supervised learning framework to effectively tackle these challenges, specifically in the context of audio fingerprinting tasks. Audio fingerprinting systems incorporate two key components: audio encoding and indexing. The existing systems consider these components independently, resulting in suboptimal performance. Our approach optimizes the interplay between these components, facilitating the adaptation of vectors to the indexing structure. Additionally, we distribute vectors in the latent$\mathbb {R}^{K}$space using normalizing flow, resulting in balanced$K$-bit hash codes. This allows indexing vectors using a balanced hash table, where vectors are uniformly distributed across all possible$2^{K}$hash buckets. This significantly accelerates retrieval, achieving speedups of up to 2× and 1.4× compared to the Locality-Sensitive Hashing (LSH) and Product Quantization (PQ), respectively. We empirically demonstrate that our system is scalable, highly effective, and efficient in identifying short audio queries ($\leq$2 s), particularly at high noise and reverberation levels. Anup Singh, Kris Demuynck, Vipul Arora 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | ECAPA2: A Hybrid Neural Network Architecture and Training Strategy for Robust Speaker EmbeddingsabstractIn this paper, we present ECAPA2, a novel hybrid neural network architecture and training strategy to produce robust speaker embeddings. Most speaker verification models are based on either the 1D- or 2D-convolutional operation, often manifested as Time Delay Neural Networks or ResNets, respectively. Hybrid models are relatively unexplored without an intuitive explanation what constitutes best practices in regard to its architectural choices. We motivate the proposed ECAPA2 model in this paper with an analysis of current speaker verification architectures. In addition, we propose a training strategy which makes the speaker embeddings more robust against overlapping speech and short utterance lengths. The presented ECAPA2 architecture and training strategy attains state-of-the-art performance on the VoxCeleb1 test sets with significantly less parameters than current models. Finally, we make a pre-trained model publicly available to promote research on downstream tasks. Jenthe Thienpondt, Kris Demuynck |
ASRU | 2 |
| 2023 | Parameter-Efficient Tuning with Adaptive Bottlenecks for Automatic Speech RecognitionabstractTransfer learning from large multilingual pretrained models, like XLSR, has become the new paradigm for Automatic Speech Recognition (ASR). Considering their ever-increasing size, fine-tuning all the weights has become impractical when the computing budget is limited. Adapters are lightweight trainable modules inserted between layers while the pre-trained part is kept frozen. They form a parameter-efficient fine-tuning method, but they still require a large bottleneck size to match standard fine-tuning performance. In this paper, we propose ABSADAPTER, a method to further reduce the parameter budget for equal task performance. Specifically, ABSADAPTER uses an Adaptive Bottleneck Scheduler to redistribute the adapter’s weights to the layers that need adaptation the most. By training only 8% of the XLSR model, ABSADAPTER achieves close to standard fine-tuning performance on a domain-shifted Air-Traffic Communication (ATC) ASR task. Geoffroy Vanderreydt, Amrutha Prasad, Driss Khalil, Srikanth R. Madikeri, Kris Demuynck, Petr Motlícek |
ASRU | 5 |
| 2023 | Simultaneously Learning Robust Audio Embeddings and Balanced Hash Codes for Query-by-ExampleabstractAudio fingerprinting systems must efficiently and robustly identify query snippets in an extensive database. To this end, state-of-the-art systems use deep learning to generate compact audio fingerprints. These systems deploy indexing methods, which quantize fingerprints to hash codes in an unsupervised manner to expedite the search. However, these methods generate imbalanced hash codes, leading to their suboptimal performance. Therefore, we propose a self-supervised learning framework to compute fingerprints and balanced hash codes in an end-to-end manner to achieve both fast and accurate retrieval performance. We model hash codes as a balanced clustering process, which we regard as an instance of the optimal transport problem. Experimental results indicate that the proposed approach improves retrieval efficiency while preserving high accuracy, particularly at high distortion levels, compared to the competing methods. Moreover, our system is efficient and scalable in computational load and memory storage. Anup Singh, Kris Demuynck, Vipul Arora 0001 |
ICASSP | 2 |
| 2023 | Margin-Mixup: A Method for Robust Speaker Verification In Multi-Speaker AudioabstractThis paper is concerned with the task of speaker verification on audio with multiple overlapping speakers. Most speaker verification systems are designed with the assumption of a single speaker being present in a given audio segment. However, in a real-world setting this assumption does not always hold. In this paper, we demonstrate that current speaker verification systems are not robust against audio with noticeable speaker overlap. To alleviate this issue, we propose margin-mixup, a simple training strategy that can easily be adopted by existing speaker verification pipelines to make the resulting speaker embeddings robust against multi-speaker audio. In contrast to other methods, margin-mixup requires no alterations to regular speaker verification architectures, while attaining better results. On our multi-speaker test set based on VoxCeleb1, the proposed margin-mixup strategy improves the EER on average with 44.4% relative to our state-of-the-art speaker verification baseline systems. Jenthe Thienpondt, Nilesh Madhu, Kris Demuynck |
ICASSP | 3 |
| 2023 | Behavioral Analysis of Pathological Speaker Embeddings of Patients During Oncological Treatment of Oral CancerabstractIn this paper, we analyze the behavior of speaker embeddings of patients during oral cancer treatment. First, we found that pre- and post-treatment speaker embeddings differ significantly, notifying a substantial change in voice characteristics. However, a partial recovery to pre-operative voice traits is observed after 12 months post-operation. Secondly, the same-speaker similarity at distinct treatment stages is similar to healthy speakers, indicating that the embeddings can capture characterizing features of even severely impaired speech. Finally, a speaker verification analysis signifies a stable false positive rate and variable false negative rate when combining speech samples of different treatment stages. This indicates robustness of the embeddings towards other speakers, while still capturing the changing voice characteristics during treatment. To the best of our knowledge, this is the first analysis of speaker embeddings during oral cancer treatment of patients. Jenthe Thienpondt, Caroline M. Speksnijder, Kris Demuynck |
INTERSPEECH | 3 |
| 2022 | Tackling the Score Shift in Cross-Lingual Speaker Verification by Exploiting Language InformationabstractThis paper contains a post-challenge performance analysis on cross-lingual speaker verification of the IDLab submission to the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). We show that current speaker embedding extractors consistently underestimate speaker similarity in within-speaker cross-lingual trials. Consequently, the typical training and scoring protocols do not put enough emphasis on the compensation of intra-speaker language variability. We propose two techniques to increase cross-lingual speaker verification robustness. First, we enhance our previously proposed Large-Margin Fine-Tuning (LM-FT) training stage with a mini-batch sampling strategy which increases the amount of intra-speaker cross-lingual samples within the mini-batch. Second, we incorporate language information in the logistic regression calibration stage. We integrate quality metrics based on soft and hard decisions of a VoxLingua107 language identification model. The proposed techniques result in a 11.7% relative improvement over the baseline model on the VoxSRC-21 test set and contributed to our third place finish in the corresponding challenge. Jenthe Thienpondt, Brecht Desplanques, Kris Demuynck |
ICASSP | 3 |
| 2022 | Transfer Learning for Robust Low-Resource Children's Speech ASR with Transformers and Source-Filter WarpingabstractAutomatic Speech Recognition (ASR) systems are known to exhibit difficulties when transcribing children's speech. This can mainly be attributed to the absence of large children's speech corpora to train robust ASR models and the resulting domain mismatch when decoding children's speech with systems trained on adult data. In this paper, we propose multiple enhancements to alleviate these issues. First, we propose a data augmentation technique based on the source-filter model of speech to close the domain gap between adult and children's speech. This enables us to leverage the data availability of adult speech corpora by making these samples perceptually similar to children's speech. Second, using this augmentation strategy, we apply transfer learning on a Transformer model pre-trained on adult data. This model follows the recently introduced XLS-R architecture, a wav2vec 2.0 model pre-trained on several cross-lingual adult speech corpora to learn general and robust acoustic frame-level representations. Adopting this model for the ASR task using adult data augmented with the proposed source-filter warping strategy and a limited amount of in-domain children's speech significantly outperforms previous state-of-the-art results on the PF-STAR British English Children's Speech corpus with a 4.86% WER on the official test set. Jenthe Thienpondt, Kris Demuynck |
INTERSPEECH | 2 |
| 2022 | Transfer Learning from Multi-Lingual Speech Translation Benefits Low-Resource Speech RecognitionabstractIn this article, we propose a simple yet effective approach to train an end-to-end speech recognition system on languages with limited resources by leveraging a large pre-trained wav2vec2.0 model fine-tuned on a multi-lingual speech translation task. We show that the weights of this model form an excellent initialization for Connectionist Temporal Classification (CTC) speech recognition, a different but closely related task. We explore the benefits of this initialization for various languages, both in-domain and out-of-domain for the speech translation task. Our experiments on the CommonVoice dataset confirm that our approach performs significantly better in-domain, and is often better out-of-domain too. This method is particularly relevant for Automatic Speech Recognition (ASR) with limited data and/or compute budget during training. Geoffroy Vanderreydt, François Remy, Kris Demuynck |
INTERSPEECH | 3 |
| 2021 | The Idlab Voxsrc-20 Submission: Large Margin Fine-Tuning and Quality-Aware Score Calibration in DNN Based Speaker VerificationabstractIn this paper we propose and analyse a large margin fine-tuning strategy and a quality-aware score calibration in text-independent speaker verification. Large margin fine-tuning is a secondary training stage for DNN based speaker verification systems trained with margin-based loss functions. It enables the network to create more robust speaker embeddings by enabling the use of longer training utterances in combination with a more aggressive margin penalty. Score calibration is a common practice in speaker verification systems to map output scores to well-calibrated log-likelihood-ratios, which can be converted to interpretable probabilities. By including quality features in the calibration system, the decision thresholds of the evaluation metrics become quality-dependent and more consistent across varying trial conditions. Applying both enhancements on the ECAPA-TDNN architecture leads to state-of-the-art results on all publicly available VoxCeleb1 test sets and contributed to our winning submissions in the supervised verification tracks of the Vox-Celeb Speaker Recognition Challenge 2020. Jenthe Thienpondt, Brecht Desplanques, Kris Demuynck |
ICASSP | 3 |
| 2021 | Integrating Frequency Translational Invariance in TDNNs and Frequency Positional Information in 2D ResNets to Enhance Speaker VerificationabstractThis paper describes the IDLab submission for the text-independent task of the Short-duration Speaker Verification Challenge 2021 (SdSVC-21). This speaker verification competition focuses on short duration test recordings and cross-lingual trials, along with the constraint of limited availability of in-domain DeepMine Farsi training data. Currently, both Time Delay Neural Networks (TDNNs) and ResNets achieve state-of-the-art results in speaker verification. These architectures are structurally very different and the construction of hybrid networks looks a promising way forward. We introduce a 2D convolutional stem in a strong ECAPA-TDNN baseline to transfer some of the strong characteristics of a ResNet based model to this hybrid CNN-TDNN architecture. Similarly, we incorporate absolute frequency positional encodings in an SE-ResNet34 architecture. These learnable feature map biases along the frequency axis offer this architecture a straightforward way to exploit frequency positional information. We also propose a frequency-wise variant of Squeeze-Excitation (SE) which better preserves frequency-specific information when rescaling the feature maps. Both modified architectures significantly outperform their corresponding baseline on the SdSVC-21 evaluation data and the original VoxCeleb1 test set. A four system fusion containing the two improved architectures achieved a third place in the final SdSVC-21 Task 2 ranking. Jenthe Thienpondt, Brecht Desplanques, Kris Demuynck |
Interspeech | 3 |
| 2020 | ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker VerificationabstractCurrent speaker verification techniques rely on a neural network to extract speaker representations. The successful x-vector architecture is a Time Delay Neural Network (TDNN) that applies statistics pooling to project variable-length utterances into fixed-length speaker characterizing embeddings. In this paper, we propose multiple enhancements to this architecture based on recent trends in the related fields of face verification and computer vision. Firstly, the initial frame layers can be restructured into 1-dimensional Res2Net modules with impactful skip connections. Similarly to SE-ResNet, we introduce Squeeze-and-Excitation blocks in these modules to explicitly model channel interdependencies. The SE block expands the temporal context of the frame layer by rescaling the channels according to global properties of the recording. Secondly, neural networks are known to learn hierarchical features, with each layer operating on a different level of complexity. To leverage this complementary information, we aggregate and propagate features of different hierarchical levels. Finally, we improve the statistics pooling module with channel-dependent frame attention. This enables the network to focus on different subsets of frames during each of the channel's statistics estimation. The proposed ECAPA-TDNN architecture significantly outperforms state-of-the-art TDNN based systems on the VoxCeleb test sets and the 2019 VoxCeleb Speaker Recognition Challenge. Brecht Desplanques, Jenthe Thienpondt, Kris Demuynck |
INTERSPEECH | 3 |
| 2020 | Cross-Lingual Speaker Verification with Domain-Balanced Hard Prototype Mining and Language-Dependent Score NormalizationabstractIn this paper we describe the top-scoring IDLab submission for the text-independent task of the Short-duration Speaker Verification (SdSV) Challenge 2020. The main difficulty of the challenge exists in the large degree of varying phonetic overlap between the potentially cross-lingual trials, along with the limited availability of in-domain DeepMine Farsi training data. We introduce domain-balanced hard prototype mining to fine-tune the state-of-the-art ECAPA-TDNN x-vector based speaker embedding extractor. The sample mining technique efficiently exploits speaker distances between the speaker prototypes of the popular AAM-softmax loss function to construct challenging training batches that are balanced on the domain-level. To enhance the scoring of cross-lingual trials, we propose a language-dependent s-norm score normalization. The imposter cohort only contains data from the Farsi target-domain which simulates the enrollment data always being Farsi. In case a Gaussian-Backend language model detects the test speaker embedding to contain English, a cross-language compensation offset determined on the AAM-softmax speaker prototypes is subtracted from the maximum expected imposter mean score. A fusion of five systems with minor topological tweaks resulted in a final MinDCF and EER of 0.065 and 1.45% respectively on the SdSVC evaluation set. Jenthe Thienpondt, Brecht Desplanques, Kris Demuynck |
INTERSPEECH | 3 |
| 2018 | Explaining Character-Aware Neural Networks for Word-Level Prediction: Do They Discover Linguistic Rules?abstractCharacter-level features are currently used in different neural network-based natural language processing algorithms.However, little is known about the character-level patterns those models learn.Moreover, models are often compared only quantitatively while a qualitative analysis is missing.In this paper, we investigate which character-level patterns neural networks learn and if those patterns coincide with manually-defined word segmentations and annotations.To that end, we extend the contextual decomposition (Murdoch et al., 2018) technique to convolutional neural networks which allows us to compare convolutional neural networks and bidirectional long short-term memory networks.We evaluate and compare these models for the task of morphological tagging on three morphologically different languages and show that these models implicitly discover understandable linguistic rules. Fréderic Godin, Kris Demuynck, Joni Dambre, Wesley De Neve, Thomas Demeester |
EMNLP | 2 |
| 2018 | Cross-lingual Speech Emotion Recognition through Factor AnalysisabstractConventional speech emotion recognition based on the extraction of high level descriptors emerging from low level descriptors seldom delivers promising results in cross-corpus experiments.Therefore it might not perform well in real-life applications.Factor analysis, proven in the fields of language identification and speaker verification, could clear a path towards more robust emotion recognition.This paper proposes an iVector-based approach operating on acoustic MFCC features with a separate modeling of the speaker and emotion variabilities respectively.The speech analysis extracts two fixedlength low-dimensional feature vectors corresponding to the two mentioned sources of variation.To model the speakerrelated nuisance variability speaker factors are extracted using an eigenvoice matrix.After compensating for this speaker variability in the supervector space, the emotion factors (one per targeted emotion) are extracted using an emotion variability matrix.The emotion factors are then fed to a basic emotion classifier.Leave-one-speaker-out cross-validation on the Berlin Database of Emotional Speech EMO-DB (German) and IEMO-CAP (English) datasets lead to results that are competitive with the current state-of-the-art.Cross-lingual experiments demonstrate the excellent robustness of the method: the classification accuracies degrade less than 15% relative when emotion models are trained on one corpus and tested on the other. Brecht Desplanques, Kris Demuynck |
INTERSPEECH | 2 |
| 2018 | Vowel Space as a Tool to Evaluate Articulation ProblemsabstractTreatment for oral tumors can lead to long term changes in the anatomy and physiology of the vocal tract and result in problems with articulation. There are currently no readily available automatic methods to evaluate changes in articulation. We developed a Praat script which plots and measures vowel space coverage. The script reproduces speaker specific vowel space use and speaking-style dependent vowel reduction in normal speech from a Dutch corpus. Speaker identity and speaking style explain more than 60% of the variance in the measured area of the vowel triangle. In recordings of patients treated for oral tumors, vowel space use before and after treatment is still significantly correlated. Articulation before and after treatment is evaluated in a listening experiment and from a maximal articulation speed task. Linear models can explain 50-75% of variance in perceptual ratings and relative articulation rate from values at previous recordings and vowel space measures. R. J. J. H. van Son, Catherine Middag, Kris Demuynck |
INTERSPEECH | 3 |
| 2018 | On the application of reservoir computing networks for noisy image recognition
Azarakhsh Jalalvand, Kris Demuynck, Wesley De Neve, Jean-Pierre Martens |
Neurocomputing | 2 |
| 2017 | Adaptive speaker diarization of broadcast news based on factor analysis
Brecht Desplanques, Kris Demuynck, Jean-Pierre Martens |
Comput. Speech Lang. | 2 |
| 2016 | Language model adaptation for ASR of spoken translations using phrase-based translation models and named entity modelsabstractLanguage model adaptation based on Machine Translation (MT) is a recently proposed approach to improve the Automatic Speech Recognition (ASR) of spoken translations that does not suffer from a common problem in approaches based on rescoring i.e. errors made during recognition cannot be recovered by the MT system. In previous work we presented an efficient implementation for MT-based language model adaptation using a word-based translation model. By omitting renormalization and employing weighted updates, the implementation exhibited virtually no adaptation overhead, enabling its use in a real-time setting. In this paper we investigate whether we can improve recognition accuracy without sacrificing the achieved efficiency. More precisely, we investigate the effect of both state-of-the-art phrase-based translation models and named entity probability estimation. We report relative WER reductions of 6.2% over a word-based LM adaptation technique and 25.3% over an unadapted 3-gram baseline on an English-to-Dutch dataset. Joris Pelemans, Tom Vanallemeersch, Kris Demuynck, Lyan Verwimp, Hugo Van hamme, Patrick Wambacq |
ICASSP | 3 |
| 2016 | STON: Efficient Subtitling in Dutch Using State-of-the-Art Tools
Lyan Verwimp, Brecht Desplanques, Kris Demuynck, Joris Pelemans, Marieke Lycke, Patrick Wambacq |
INTERSPEECH | 3 |
| 2016 | SCALE: A Scalable Language Engineering Toolkit
Joris Pelemans, Lyan Verwimp, Kris Demuynck, Hugo Van hamme, Patrick Wambacq |
LREC | 3 |
| 2015 | Improving n-gram probability estimates by compound-head clusteringabstractCompounding is one of the most productive word formation processes in many languages and is therefore a main source of data sparsity in language modeling. Many solutions have been suggested to model compound words, most of which break the compound into its constituents and train a new model with them. In earlier work, we argued that this approach is suboptimal and we presented a novel technique that clusters new, domain-specific compound words together with their semantic heads. The clusters were then used to build a class-based n-gram model that enabled a reliable estimation of n-gram probabilities, without the need for additional training data. In this paper, we investigate how this “semantic head mapping” can best be made an integral part of the language modeling strategy and find that, with some adaptations, our technique is capable of producing more accurate compound probability estimates than a baseline word-based n-gram language model, which lead to a significant word error rate reduction for Dutch read speech. Joris Pelemans, Kris Demuynck, Hugo Van hamme, Patrick Wambacq |
ICASSP | 2 |
| 2015 | Factor analysis for speaker segmentation and improved speaker diarizationabstractSpeaker diarization includes two steps: speaker segmentation and speaker clustering. Speaker segmentation searches for speaker boundaries, whereas speaker clustering aims at grouping speech segments of the same speaker. In this work, the segmentation is improved by replacing the Bayesian Information Criterion (BIC) with a new iVector-based approach. Unlike BIC-based methods which trigger on any acoustic dissimilarities, the proposed method suppresses phonetic variations and accentuates speaker differences. More specifically our method generates boundaries based on the distance between two speaker factor vectors that are extracted on a frame-by frame basis. The extraction relies on an eigenvoice matrix so that large differences between speaker factor vectors indicate a different speaker. A Mahalanobis-based distance measure, in which the covariance matrix compensates for the remaining and detrimental phonetic variability, is shown to generate accurate boundaries. The detected segments are clustered by a state-of-the-art iVector Probabilistic Linear Discriminant Analysis system. Experiments on the COST278 multilingual broadcast news database show relative reductions of 50% in boundary detection errors. The speaker error rate is reduced by 8% relative. Brecht Desplanques, Kris Demuynck, Jean-Pierre Martens |
INTERSPEECH | 2 |
| 2015 | Efficient language model adaptation for automatic speech recognition of spoken translationsabstractCopyright © 2015 ISCA. Direct integration of translation model (TM) probabilities into a language model (LM) with the purpose of improving automatic speech recognition (ASR) of spoken translations typically requires a number of complex operations for each sentence. Many if not all of the LM probabilities need to be updated, the model needs to be renormalized and the ASR system needs to load a new, updated LM for each sentence. In computer-aided translation environments the time loss induced by these complex operations seriously reduces the potential of ASR as an efficient input method. In this paper we present a novel LM adaptation technique that drastically reduces the complexity of each of these operations. The technique consists of LM probability updates using exponential weights based on TM probabilities for each sentence and does not enforce probability renormalization. Instead of storing each resulting language model in its entirety, we only store the update weights which also reduces disk storage and loading time during ASR. Experiments on Dutch read speech translated from English show that both disk storage and recognition time drop dramatically compared to a baseline system that employs a more conventional way of updating the LM. Joris Pelemans, Tom Vanallemeersch, Kris Demuynck, Hugo Van hamme, Patrick Wambacq |
INTERSPEECH | 3 |
| 2015 | Robust continuous digit recognition using Reservoir Computing
Azarakhsh Jalalvand, Fabian Triefenbach, Kris Demuynck, Jean-Pierre Martens |
Comput. Speech Lang. | 3 |
| 2015 | Integrating meta-information into recurrent neural network language models
Yangyang Shi, Martha A. Larson, Joris Pelemans, Catholijn M. Jonker, Patrick Wambacq, Pascal Wiggers, Kris Demuynck |
Speech Commun. | 7 |
| 2014 | Coping with language data sparsity: Semantic head mapping of compound wordsabstractIn this paper we present a novel clustering technique for compound words. By mapping compounds onto their semantic heads, the technique is able to estimate n-gram probabilities for unseen compounds. We argue that compounds are well represented by their heads which allows the clustering of rare words and reduces the risk of over-generalization. The semantic heads are obtained by a two-step process which consists of constituent generation and best head selection based on corpus statistics. Experiments on Dutch read speech show that our technique is capable of correctly identifying compounds and their semantic heads with a precision of 80.25% and a recall of 85.97%. A class-based language model with compound-head clusters achieves a significant reduction in both perplexity and WER. Joris Pelemans, Kris Demuynck, Hugo Van hamme, Patrick Wambacq |
ICASSP | 2 |
| 2014 | Robust language recognition via adaptive language factor extractionabstractThis paper presents a technique to adapt an acoustically based language classifier to the background conditions and speaker accents. This adaptation improves language classification on a broad spectrum of TV broadcasts. The core of the system consists of an iVector-based setup in which language and channel variabilities are modeled separately. The subsequent language classifier (the backend) operates on the language factors, i.e. those features in the extracted iVectors that explain the observed language variability. The proposed technique adapts the language variability model to the background conditions and to the speaker accents present in the audio. The effect of the adaptation is evaluated on a 28 hours corpus composed of documentaries and monolingual as well as multilingual broadcast news shows. Consistent improvements in the automatic identification of Flemish (Belgian Dutch), English and French are demonstrated for all broadcast types. Brecht Desplanques, Kris Demuynck, Jean-Pierre Martens |
INTERSPEECH | 2 |
| 2014 | Speech Recognition Web Services for Dutch
Joris Pelemans, Kris Demuynck, Hugo Van hamme, Patrick Wambacq |
LREC | 2 |
| 2014 | An improved two-stage mixed language model approach for handling out-of-vocabulary words in large vocabulary continuous speech recognition
Bert Réveil, Kris Demuynck, Jean-Pierre Martens |
Comput. Speech Lang. | 2 |
| 2014 | Large Vocabulary Continuous Speech Recognition With Reservoir-Based Acoustic ModelsabstractThanks to research in neural network based acoustic modeling, progress in Large Vocabulary Continuous Speech Recognition (LVCSR) seems to have gained momentum recently. In search for further progress, the present letter investigates Reservoir Computing (RC) as an alternative new paradigm for acoustic modeling. RC unifies the appealing dynamical modeling capacity of a Recurrent Neural Network (RNN) with the simplicity and robustness of linear regression as a model for training the weights of that network. In previous work, an RC-HMM hybrid yielding very good phone recognition accuracy on TIMIT could be designed, but no proof was offered yet that this success would also transfer to LVCSR. This letter describes the development of an RC-HMM hybrid that provides good recognition on the Wall Street Journal benchmark. For the WSJ0 5k word task, word error rates of 6.2% (bigram language model) and 3.9% (trigram) are obtained on the Nov-92 evaluation set. Given that RC-based acoustic modeling is a fairly new approach, these results open up promising perspectives. Fabian Triefenbach, Kris Demuynck, Jean-Pierre Martens |
IEEE Signal Process. Lett. | 2 |
| 2013 | Porting concepts from DNNs back to GMMsabstractDeep neural networks (DNNs) have been shown to outperform Gaussian Mixture Models (GMM) on a variety of speech recognition benchmarks. In this paper we analyze the differences between the DNN and GMM modeling techniques and port the best ideas from the DNN-based modeling to a GMM-based system. By going both deep (multiple layers) and wide (multiple parallel sub-models) and by sharing model parameters, we are able to close the gap between the two modeling techniques on the TIMIT database. Since the `deep' GMMs retain the maximum-likelihood trained Gaussians as first layer, advanced techniques such as speaker adaptation and model-based noise robustness can be readily incorporated. Regardless of their similarities, the DNNs and the deep GMMs still show a sufficient amount of complementarity to allow effective system combination. Kris Demuynck, Fabian Triefenbach |
ASRU | 1 |
| 2013 | Exemplar-based joint channel and noise compensationabstractIn this paper two models for channel estimation in exemplar-based noise robust speech recognition are proposed. Building on a compositional model that models noisy speech and a combination of noise and speech atoms, the first model iteratively estimates a filter to best compensate the mismatch with the observed noisy speech. The second model estimates separate filters for the noise and speech atoms. We show that both models enable noise-robust ASR even if the channel characteristics of the noisy speech do not match those of the exemplars in the dictionary. Moreover, the second model, which is able to estimate separate filters for speech and noise, is shown to be robust even in the presence of bandwidth-limited sources. Jort F. Gemmeke, Tuomas Virtanen, Kris Demuynck |
ICASSP | 3 |
| 2013 | Context-dependent modeling and speaker normalization applied to reservoir-based phone recognitionabstractReservoir Computing (RC) has recently been introduced as an interesting alternative for acoustic modeling. For phone and continuous digit recognition, the reservoir approach obtained quite promising results. In this work, we further elaborate this concept by porting some well-known techniques used to enhance recognition rates of GMM-based models to Reservoir Computing. In particular, we introduce context-dependent (CD) triphone states to model co-articulation and pronunciation mismatches arising from an imperfect lexicon. We also propose to incorporate two speaker normalization methods in the feature space, namely mean \& variance normalization and vocal tract length normalization. The impact of the investigated techniques is studied in the context of phone recognition on the TIMIT corpus. Our CD-RC-HMM hybrid yields a speaker-independent phone error rate (PER) of 22\% and a speaker-dependent PER of 20.5\%. By combining GMM and RC-based likelihoods at the state level, these scores can be reduced further. Fabian Triefenbach, Azarakhsh Jalalvand, Kris Demuynck, Jean-Pierre Martens |
INTERSPEECH | 3 |
| 2013 | Rapid speaker adaptation in latent speaker space with non-negative matrix factorization
Xueru Zhang, Kris Demuynck, Hugo Van hamme |
Speech Commun. | 2 |
| 2013 | Acoustic Modeling With Hierarchical ReservoirsabstractAccurate acoustic modeling is an essential requirement of a state-of-the-art continuous speech recognizer. The Acoustic Model (AM) describes the relation between the observed speech signal and the non-observable sequence of phonetic units uttered by the speaker. Nowadays, most recognizers use Hidden Markov Models (HMMs) in combination with Gaussian Mixture Models (GMMs) to model the acoustics, but neural-based architectures are on the rise again. In this work, the recently introduced Reservoir Computing (RC) paradigm is used for acoustic modeling. A reservoir is a fixed - and thus non-trained - Recurrent Neural Network (RNN) that is combined with a trained linear model. This approach combines the ability of an RNN to model the recent past of the input sequence with a simple and reliable training procedure. It is shown here that simple reservoir-based AMs achieve reasonable phone recognition and that deep hierarchical and bi-directional reservoir architectures lead to a very competitive Phone Error Rate (PER) of 23.1% on the well-known TIMIT task. Fabian Triefenbach, Azarakhsh Jalalvand, Kris Demuynck, Jean-Pierre Martens |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | A layered approach for dutch large vocabulary continuous speech recognitionabstractIn this paper we investigate whether a layered architecture that has already proven its value for small tasks, works for a system with large lexica (400k words) and language models (5-grams) as well. The architecture was designed to decouple phone and word recognition which allows for the integration of more complex linguistic components, especially at the sub-word level. It was tested on the Dutch language which - with its large variety of accents and rich morphology - is ideally suited to benefit from this integration. The results reveal that the architecture is already competitive to an all-in-one approach in which acoustic models, language models and lexicon are all applied simultaneously. Candidates for further improvement to the system based on a conditional phone confusion model are suggested. Joris Pelemans, Kris Demuynck, Patrick Wambacq |
ICASSP | 2 |
| 2012 | Latent variable speaker adaptation of Gaussian mixture weights and meansabstractWe describe a novel fast speaker adaptation algorithm for large vocabulary speech recognition systems, which adapts both the Gaussian means and the mixture weights. Gaussian means are expressed as a linear combination of eigenvoices estimated with principal component analysis. The non-negative Gaussian mixture weights are expressed as a linear combination of a set of latent vectors estimated with non-negative matrix factorization. Experiments on the Wall Street Journal database show that the combination of weight and mean adaptation consistently improves the performance compared to eigenvoice adaptation only. Improvements up to 5.8% relative word error rate reduction were observed with 40 eigenvoices and 40 latent weight vectors. Furthermore, combining weight and mean adaptation outperformed both weight and mean adaptation on itself, even if the latter uses more latent vectors. Xueru Zhang, Kris Demuynck, Hugo Van hamme |
ICASSP | 2 |
| 2012 | Dutch Automatic Speech Recognition on the Web: Towards a General Purpose SystemabstractIn this paper we present our state-of-the-art automatic speech recognition system for Dutch that we made available on the web. The free, online disclosure of our software aims at allowing non-specialists to adopt ASR technology effortlessly. Access is possible via a standard web browser or as a web service in automated tools. We discuss the way the web application was built and focus on usability criteria - especially interoperability. To overcome user differences we provide input conversion and basic parameter selection. Extensions of the current system and a path to robust, general purpose ASR are suggested. Joris Pelemans, Kris Demuynck, Patrick Wambacq |
INTERSPEECH | 2 |
| 2012 | Improving large vocabulary continuous speech recognition by combining GMM-based and reservoir-based acoustic modelingabstractIn earlier work we have shown that good phoneme recognition is possible with a so-called reservoir, a special type of recurrent neural network. In this paper, different architectures based on Reservoir Computing (RC) for large vocabulary continuous speech recognition are investigated. Besides experiments with HMM hybrids, it is shown that a RC-HMM tandem can achieve the same recognition accuracy as a classical HMM, which is a promising result for such a fairly new paradigm. It is also demonstrated that a state-level combination of the scores of the tandem and the baseline HMM leads to a significant improvement over the baseline. A word error rate reduction of the order of 20% relative is possible. Fabian Triefenbach, Kris Demuynck, Jean-Pierre Martens |
SLT | 2 |
| 2011 | Integrating meta-information into exemplar-based speech recognition with segmental conditional random fieldsabstractExemplar based recognition systems are characterized by the fact that, instead of abstracting large amounts of data into compact models, they store the observed data enriched with some annotations and infer on-the-fly from the data by finding those exemplars that resemble the input speech best. One advantage of exemplar based systems is that next to deriving what the current phone or word is, one can easily derive a wealth of meta-information concerning the chunk of audio under investigation. In this work we harvest meta-information from the set of best matching exemplars, that is thought to be relevant for the recognition such as word boundary predictions and speaker entropy. Integrating this meta-information into the recognition framework using segmental conditional random fields, reduced the WER of the exemplar based system on the WSJ Nov92 20k task from 8.2% to 7.6%. Adding the HMM-score and multiple HMM phone detectors as features further reduced the error rate to 6.6%. Kris Demuynck, Dino Seppi, Dirk Van Compernolle, Patrick Nguyen, Geoffrey Zweig |
ICASSP | 1 |
| 2011 | Progress in example based automatic speech recognitionabstractIn this paper we present a number of improvements that were recently made to the template based speech recognition system developed at ESAT. Combining these improvements resulted in a decrease in word error rate from 9.6% to 8.2% on the Nov92, 20k trigram, Wall Street Journal task. The improvements are along different lines. Apart from the time warping already applied within the DTW, it was found beneficial to apply additional length compensation on the template score. The single best score was replaced by a weighted k-NN average, while maintaining natural successor information as an ensemble cost. The local geometry of the acoustic space is now taken into account by assigning a diagonal covariance matrix to each input frame. Context sensitivity of short templates is increased by taking cross boundary scores into account for sorting the N best templates. Furthermore boundaries on the template segmentations may be relaxed. Finally context dependent word templates are now being used for short words. Several other variants that were not retained in the final system are discussed as well. Kris Demuynck, Dino Seppi, Hugo Van hamme, Dirk Van Compernolle |
ICASSP | 1 |
| 2011 | Rapid speaker adaptation with speaker adaptive training and non-negative matrix factorizationabstractIn this paper, we describe a novel speaker adaptation algorithm based on Gaussian mixture weight adaptation. A small number of latent speaker vectors are estimated with non-negative matrix factorization (NMF). These base vectors encode the correlations between Gaussian activations as learned from the train data. Expressing the speaker dependent Gaussian mixture weights as a linear combination of a small number of base vectors, reduces the number of parameters that must be estimated from the enrollment data. In order to learn meaningful correlations between Gaussian activations from the train data, the NMF-based weight adaptation was combined with vocal tract length normalization (VTLN) and feature-space maximum likelihood linear regression (fMLLR) based speaker adaptive training based. Evaluation on the 5k closed and 20k open vocabulary Wall Street Journal tasks shows a 4% relative word error rate reduction over the speaker independent recognition system which already incorporates VTLN. The proposed fast adaptation algorithm, using a single enrollment sentence only, results in similar performance as fMLLR adapting on 40 enrollment sentences. Xueru Zhang, Kris Demuynck, Hugo Van hamme |
ICASSP | 2 |
| 2011 | Speech recognitionwith segmental conditional random fields: A summary of the JHU CLSP 2010 Summer WorkshopabstractThis paper summarizes the 2010 CLSP Summer Workshop on speech recognition at Johns Hopkins University. The key theme of the workshop was to improve on state-of-the-art speech recognition systems by using Segmental Conditional Random Fields (SCRFs) to integrate multiple types of information. This approach uses a state of-the-art baseline as a springboard from which to add a suite of novel features including ones derived from acoustic templates, deep neural net phoneme detections, duration models, modulation features, and whole word point-process models. The SCRF framework is able to appropriately weight these different information sources to produce significant gains on both die Broadcast News and Wall Street Journal tasks. Geoffrey Zweig, Patrick Nguyen, Dirk Van Compernolle, Kris Demuynck, Les E. Atlas, Pascal Clark, Gregory Sell, Meihong Wang, Fei Sha, Hynek Hermansky, Damianos Karakos, Aren Jansen, Samuel Thomas 0001, Sivaram G. S. V. S., Samuel R. Bowman, Justine T. Kao |
ICASSP | 4 |
| 2011 | Template-Based Automatic Speech Recognition Meets ProsodyabstractIn this paper, we use prosodic information to improve the accuracy of our template-based automatic speech recognizer. Prosodic information is harvested adopting a data-driven approach. A number of prosodic features is extracted, then combined into major groups, and finally studied separately and together. All acoustic evidence, both segmental and suprasegmental, is modelled non-parametrically. The different sources of information are conveniently combined with segmental conditional random fields. Prosody enhances the accuracy of the state-of-the-art baseline by reducing the word error rate by 7% relative on the nov92, 20κ trigram, Wall Street Journal task. Copyright © 2011 ISCA. Dino Seppi, Kris Demuynck, Dirk Van Compernolle |
INTERSPEECH | 2 |
| 2010 | Histogram equalization and noise masking for robust speech recognitionabstractMismatch between training and test conditions deteriorates the performance of speech recognizers. This paper investigates the combination of parametric histogram equalization (pHEQ) and noise masking to compensate for the mismatch caused by additive noise. The proposed front-end maps the distribution of the observed power spectrum vectors to a target distribution. The target distribution matches the distribution of the noise free training data except for an artificially reduced signal-to-noise ratio. Different power spectrum estimation algorithms are used to estimate the noise distribution as used internally by pHEQ more reliably under non-stationary noise conditions. The proposed front-end is evaluated on the Aurora4 database and shows a significant improvement w.r.t. mean-normalized Mel-frequency spectral coefficients. Moreover, the performance could be further improved if better estimates of the instantaneous noise power spectrum were available. Xueru Zhang, Kris Demuynck, Hugo Van hamme |
ICASSP | 2 |
| 2010 | Feature versus model based noise robustnessabstractOver the years, the focus in noise robust speech recognition has shifted from noise robust features to model based techniques such as parallel model combination and uncertainty decoding. In this paper, we contrast prime examples of both approaches in the context of large vocabulary recognition systems such as used for automatic audio indexing and transcription. We look at the approximations the techniques require to keep the computational load reasonable, the resulting computational cost, and the accuracy measured on the Aurora4 benchmark. The results show that a well designed feature based scheme is capable of providing recognition accuracies at least as good as the model based approaches at a substantially lower computational cost. © 2010 ISCA. Kris Demuynck, Xueru Zhang, Dirk Van Compernolle, Hugo Van hamme |
INTERSPEECH | 1 |
| 2010 | A Speech Corpus for Modeling Language Acquisition: CAREGIVER
Toomas Altosaar, Louis ten Bosch, Guillaume Aimetti, Christos Koniaris, Kris Demuynck, Henk van den Heuvel |
LREC | 5 |
| 2009 | The ESAT 2008 system for N-Best Dutch speech recognition benchmarkabstractThis paper describes the ESAT 2008 Broadcast News transcription system for the N-Best 2008 benchmark, developed in part for testing the recent SPRAAK Speech Recognition Toolkit. ESAT system was developed for the Southern Dutch Broadcast News subtask of N-Best using standard methods of modern speech recognition. A combination of improvements were made in commonly overlooked areas such as text normalization, pronunciation modeling, lexicon selection and morphological modeling, virtually solving the out-of-vocabulary (OOV) problem for Dutch by reducing OOV-rate to 0.06% on the N-Best development data and 0.23% on the evaluation data. Recognition experiments were run with several configurations comparing one-pass vs. two-pass decoding, high-order vs. low-order n-gram models, lexicon sizes and different types of morphological modeling. The system achieved 7.23% word error rate (WER) on the broadcast news development data and 20.3% on the much more difficult evaluation data of N-Best. Kris Demuynck, Antti Puurula, Dirk Van Compernolle, Patrick Wambacq |
ASRU | 1 |
| 2009 | Evaluation of phone lattice based speech decodingabstractPreviously, we proposed a flexible two-layered speech recogniser architecture, called FLaVoR. In the first layer an unconstrained, task independent phone recogniser generates a phone lattice. Only in the second layer the task specific lexicon and language model are applied to decode the phone lattice and produce a word level recognition result. In this paper, we present a further evaluation of the FLaVoR architecture. The performance of a classical single-layered architecture and the FLaVoR architecture are compared on two recognition tasks, using the same acoustic, lexical and language models. On the large vocabulary Wall Street Journal 5k and 20k benchmark tasks, the two-layered architecture resulted in slightly but not significantly better word error rates. On a reading error detection task for a reading tutor for children, the FLaVoR architecture clearly outperformed the single-layered architecture. Index Terms: ASR architecture, phone lattice decoding, system assessment 1. Jacques Duchateau, Kris Demuynck, Hugo Van hamme |
INTERSPEECH | 2 |
| 2009 | Developing a reading tutor: Design and evaluation of dedicated speech recognition and synthesis modules
Jacques Duchateau, Yuk On Kong, Leen Cleuren, Lukas Latacz, Jan Roelens, Abdurrahman Samir, Kris Demuynck, Pol Ghesquière, Werner Verhelst, Hugo Van hamme |
Speech Commun. | 7 |
| 2008 | Unsupervised learning of auditory filter banks using non-negative matrix factorisationabstractNon-negative matrix factorisation (NMF) is an unsupervised learning technique that decomposes a non-negative data matrix into a product of two lower rank non-negative matrices. The non-negativity constraint results in a parts-based and often sparse representation of the data. We use NMF to factorise a matrix with spectral slices of continuous speech to automatically find a feature set for speech recognition. The resulting decomposition yields a filter bank design with remarkable similarities to perceptually motivated designs, supporting the hypothesis that human hearing and speech production are well matched to each other. We point out that the divergence cost criterion used by NMF is linearly dependent on energy, which may influence the design. We will however argue that this does not significantly affect the interpretation of our results. Furthermore, we compare our filter bank with several hearing models found in literature. Evaluating the filter bank for speech recognition shows that the same recognition performance is achieved as with classical MEL-based features. Alexander Bertrand, Kris Demuynck, Veronique Stouten, Hugo Van hamme |
ICASSP | 2 |
| 2008 | Fast speaker adaptation using non-negative matrix factorizationabstractThis paper describes a new method for fast speaker adaptation in large vocabulary recognition systems. As in most HMM-based recognizers, the observation densities are modeled as a weighted sum of Gaussian densities. Instead of adapting the means of the Gaussian densities, which is typically done, the weights for the Gaussian densities in the states are adapted. By applying non-negative matrix factorization (NMF) in the proposed method, very fast adaptation was achieved. Experiments on the Wall Street Journal benchmark recognition task show relative improvements between 5% and 15%, while the adaptation converges within 0.2 seconds. Analysis of the latent speakers found by NMF learns that these latent speakers reflect the gender of the speaker most prominently, even when vocal tract length normalization is used, and that they reflect the speaker’s age more clearly than the speaker’s regional influences or dialect. Jacques Duchateau, Tobias Leroy, Kris Demuynck, Hugo Van hamme |
ICASSP | 3 |
| 2008 | SPRAAK: an open source "SPeech recognition and automatic annotation kit"
Kris Demuynck, Jan Roelens, Dirk Van Compernolle, Patrick Wambacq |
INTERSPEECH | 1 |
| 2008 | Discovering Phone Patterns in Spoken Utterances by Non-Negative Matrix FactorizationabstractWe present a technique to automatically discover the (word-sized) phone patterns that are present in speech utterances. These patterns are learnt from a set of phone lattices generated from the utterances. Just like children acquiring language, our system does not have prior information on what the meaningful patterns are. By applying the non-negative matrix factorization algorithm to a fixed-length high-dimensional vector representation of the speech utterances, a decomposition in terms of additive units is obtained. We illustrate that these units correspond to words in case of a small vocabulary task. Our result also raises questions about whether explicit segmentation and clustering are needed in an unsupervised learning context. Veronique Stouten, Kris Demuynck, Hugo Van hamme |
IEEE Signal Process. Lett. | 2 |
| 2007 | Outlier Correction for Local Distance Measures in Example Based Speech RecognitionabstractExample based speech recognition is critically dependent on the quality of the acoustic distance measure between input and reference vectors. In the past, the commonly used Euclidean distance has been refined to take into account the covariance of the different sounds, resulting in a class dependent distance measure. However, using the same measure for the whole class is still too crude: vectors in the tails of the distribution (outliers) are unduly considered equally representative of the class as those in the centre. In this paper, we derive two techniques inspired by non-parametric density estimation that explicitly adjust the distance measure based on the position of the reference vector in its class. Experiments on three low-level acoustic tasks show that "data sharpening" results in a substantial improvement, while "adaptive kernels" have minimal effect. Mathias De Wachter, Kris Demuynck, Dirk Van Compernolle |
ICASSP (4) | 2 |
| 2007 | Automatically learning the units of speech by non-negative matrix factorisationabstractWe present an unsupervised technique to discover the (word-sized) speech units in which a corpus of utterances can be de-composed. First, a fixed-length high-dimensional vector rep-resentation of the utterances is obtained. Then, the resulting matrix is decomposed in terms of additive units by applying the non-negative matrix factorisation algorithm. On a small vocab-ulary task, the obtained basis vectors each represent one of the uttered words. We also investigate the amount of speech data that is needed to obtain a correct set of basis vectors. By de-creasing the number of occurrences of the words in the corpus, an indication of the learning rate of the system is obtained. Index Terms: matrix factorisation, word segmentation, phone lattices, language acquisition. Veronique Stouten, Kris Demuynck, Hugo Van hamme |
INTERSPEECH | 2 |
| 2007 | Evaluating acoustic distance measures for template based recognitionabstractIn this paper we investigate the behaviour of different acoustic distance measures for template based speech recognition in light of the combination of acoustic distances, linguistic knowledge and template concatenation fluency costs. To that end, different acoustic distance measures are compared on tasks with varying levels of fluency/linguistic constraints. We show that the adoption of those constraints invariably results in an acoustically clearly suboptimal template sequence being chosen as the winning hypothesis. There are strong implications for the design of acoustic distance measures: distance measures that are optimal for frame based classification may prove to be suboptimal for full sentence recognition. In particular, we show this is the case when comparing the Euclidean and the recently introduced adaptive kernel local Mahalanobis distance measures. Index Terms: template based, example based, episodic, nonparametric 1. Mathias De Wachter, Kris Demuynck, Patrick Wambacq, Dirk Van Compernolle |
INTERSPEECH | 2 |
| 2007 | Template-Based Continuous Speech RecognitionabstractDespite their known weaknesses, hidden Markov models (HMMs) have been the dominant technique for acoustic modeling in speech recognition for over two decades. Still, the advances in the HMM framework have not solved its key problems: it discards information about time dependencies and is prone to overgeneralization. In this paper, we attempt to overcome these problems by relying on straightforward template matching. The basis for the recognizer is the well-known DTW algorithm. However, classical DTW continuous speech recognition results in an explosion of the search space. The traditional top-down search is therefore complemented with a data-driven selection of candidates for DTW alignment. We also extend the DTW framework with a flexible subword unit mechanism and a class sensitive distance measure—two components suggested by state-of-the-art HMM systems. The added flexibility of the unit selection in the template-based framework leads to new approaches to speaker and environment adaptation. The template matching system reaches a performance somewhat worse than the best published HMM results for the Resource Management benchmark, but thanks to complementarity of errors between the HMM and DTW systems, the combination of both leads to a decrease in word error rate with 17% compared to the HMM results. Mathias De Wachter, Mike Matton, Kris Demuynck, Patrick Wambacq, Ronald Cools, Dirk Van Compernolle |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Robust phone lattice decodingabstractMost ASR systems adopt an all-in-one approach: acoustic model, lexicon and language model are all applied simultaneously, thus forming a single large search space. This way, both lexicon and language model help in constraining the search at an early stage which greatly improves its efficiency. However, such close integration comes at a cost: all resources must be kept simple. Achieving higher accuracy in unconstrained LVCSR tasks will require more complex resources while at the same time the 'unconstrainedness' of the task reduces the effectiveness of the all-in-one approach. Therefore, we propose a modular two-layered architecture. First, a pure acoustic-phonemic search generates a dense phone network. Next a robust decoder finds those words from the lexicon that match well with the phone sequences encoded in the phone network. In this paper we investigate the properties the robust word decoder must have and we propose an efficient search algorithm. Kris Demuynck, Dirk Van Compernolle, Hugo Van hamme |
INTERSPEECH | 1 |
| 2006 | Boosting HMM performance with a memory upgradeabstractThe state-of-the-art in automatic speech recognition is distinctly Markovian. The ubiquitous 'beads-on-a-string' approach, where sentences are explained as a sequence of words, words as a sequence of phones and phones as a sequence of acoustically stable states, is bound to lose a lot of dynamic information. In this paper we show that a combination with example-based recognition can be used to recapture some of that information. A new approach to combine Hidden Markov Model (HMM) and phone-examplebased continuous speech recognition is presented. Experiments show that the combination outperforms the HMM recognizer, and indicate that adding long-span information is especially beneficial. Mathias De Wachter, Kris Demuynck, Dirk Van Compernolle |
INTERSPEECH | 2 |
| 2004 | A locally weighted distance measure for example based speech recognitionabstractState-of-the-art speech recognition relies on a state-dependent distance measure. In HMM systems, the distance measure is trained into state-dependent covariance matrices using a maximum likelihood or discriminative criterion. This "automatic" adjustment of the distance measure is traditionally considered an inherent advantage of HMMs over DTW (dynamic time warping) recognizers, as those typically rely on a uniform Euclidean distance. We show how to incorporate a non-uniform weighted distance measure into an example-based recognition system. By doing so, we manage to combine the superior segmental behaviour of DTW with the near-optimal acoustic distance measure as found in HMMs. The non-uniform distance measure enforces modifications to the k nearest neighbours search, an essential component in our large vocabulary DTW approach. We show that the complexity of our solution remains within bounds. The validity of the full approach is verified by experimental results on the resource management and TIDigits tasks. Mathias De Wachter, Kris Demuynck, Patrick Wambacq, Dirk Van Compernolle |
ICASSP (1) | 2 |
| 2004 | Synthesizing speech from speech recognition parametersabstractThe merits of different signal preprocessing schemes for speech recognizers are usually assessed purely on the basis of the resulting recognition accuracy. Such benchmarks give a good indication as to whether one preprocessing is better than another, but little knowledge is acquired about why it is better or how it could be further improved. In order to gain more insight in the preprocessing, we seek to re-synthesize speech from speech recognition features. This way, we are able to pin-point some deficiencies in our current preprocessing scheme. Additional analysis of successful new preprocessing schemes may allow us one day to identify precisely those properties that are desirable in a feature set. Next to these purely scientific aims, the re-synthesis of speech from recognition features is of interest to thin-client speech applications, and as an alternative to the classical LPC source-filter model for speech manipulation. Kris Demuynck, Oscar Garcia, Dirk Van Compernolle |
INTERSPEECH | 1 |
| 2004 | Automatic Phonemic Labeling and Segmentation of Spoken Dutch
Kris Demuynck, Tom Laureys, Patrick Wambacq, Dirk Van Compernolle |
LREC | 1 |
| 2003 | FLavor: a flexible architecture for LVCSRabstractThis paper describes a new architecture for large vocabulary continuous speech recognition (LVCSR), which will be developed within the project FLaVoR (Flexible Large Vocabulary Recognition). The proposed architecture abandons the standard all-in-one search strategy with integrated acoustic, lexical and language model information. Instead, a modular framework is proposed which allows for the integration of more complex linguistic components. The search process consists of two layers. First, a pure acoustic-phonemic search generates a dense phoneme network enriched with meta-data. Then, the output of the first layer is used by sophisticated language technology components for word decoding in the second layer. Preliminary experiments prove the feasibility of the approach. Kris Demuynck, Tom Laureys, Dirk Van Compernolle, Hugo Van hamme |
INTERSPEECH | 1 |
| 2003 | Robust speech recognition using model-based feature enhancementabstractMaintaining a high level of robustness for Automatic Speech Recognition (ASR) systems is especially challenging when the background noise has a time-varying nature. We have implemented a Model-Based Feature Enhancement (MBFE) technique that not only can easily be embedded in the feature extraction module of a recogniser, but also is intrinsically suited for the removal of non-stationary additive noise. To this end we combine statistical models of the cepstral feature vectors of both clean speech and noise, using a Vector Taylor Series approximation in the power spectral domain. Based on this combined HMM, a global MMSE-estimate of the clean speech is then calculated. Because of the scalability of the applied models, MBFE is flexible and computationally feasible. Recognition experiments with this feature enhancement technique on the Aurora2 connected digit recognition task showed significant improvements on the noise robustness of the HTK recogniser. Veronique Stouten, Hugo Van hamme, Kris Demuynck, Patrick Wambacq |
INTERSPEECH | 3 |
| 2003 | Data driven example based continuous speech recognitionabstractThe dominant acoustic modeling methodology based on Hidden Markov Models is known to have certain weaknesses. Partial solutions to these flaws have been presented, but the fundamental problem remains: compression of the data to a compact HMM discards useful information such as time dependencies and speaker information. In this paper, we look at pure example based recognition as a solution to this problem. By replacing the HMM with the underlying examples, all information in the training data is retained. We show how information about speaker and environment can be used, introducing a new interpretation of adaptation. The basis for the recognizer is the well-known DTW algorithm, which has often been used for small tasks. However, large vocabulary speech recognition introduces new demands, resulting in an explosion of the search space. We show how this problem can be tackled using a data driven approach which selects appropriate speech examples as candidates for DTW-alignment. Mathias De Wachter, Kris Demuynck, Dirk Van Compernolle, Patrick Wambacq |
INTERSPEECH | 2 |
| 2002 | Doing away with the Viterbi approximationabstractIn this paper, we investigate the use of the total likelihood (the weighted sum of the likelihoods of all possible state sequences) instead of the approximation with the Viterbi likelihood (the like-lihood of the best state sequence) normally used in speech recognition. Next to its use in a recognizer, the use of total likelihoods in the context of an automatic word aligrunent task is also addressed shortly. We describe how the search algorithm must be modified and how word lattices based on total likelihoods can be constructed. The total likelihood framework also requires us to make a distinction between upgrading the language model scores or downgrading the acoustic model scores in the recognizer. To help in deciding between these two alternatives, some theoretical foundation is given to the practice of making a weighted combination of language and acoustic scores. Finally, the total likelihood and the Viterbi framework are compared in terms of accuracy and computational effort on the Wall Street Journal recognition task, while the accuracy of word alignments is evaluated on a large Dutch corpus. Kris Demuynck, Dirk Van Compernolle, Patrick Wambacq |
ICASSP | 1 |
| 2002 | Confidence scoring based on backward language modelsabstractIn this paper we introduce the backward N-gram language model (LM) scores as a confidence measure in large vocabulary continuous speech recognition. Jacques Duchateau, Kris Demuynck, Patrick Wambacq |
ICASSP | 2 |
| 2002 | Autoregressive acoustical modelling of free field cough soundabstractIn this paper the performance and assumptions of linear prediction acoustical modelling are assessed on the free field cough sound. Four distinct free field cough classes originating from animal and human species in different health conditions are considered. Firstly based on the prediction signal-to-noise ratio the model order for each cough class was chosen to 14. Secondly for each cough class the vocal tract formants are estimated from the linear prediction parameters. Finally the occurrence of subglottal resonances in the cough-sound is investigated. Annemie Van Hirtum, Daniel Berckmans, Kris Demuynck, Dirk Van Compernolle |
ICASSP | 3 |
| 2002 | Automatic generation of phonetic transcriptions for large speech corporaabstractWe describe a method for the automatic production of phonetic transcriptions in large speech corpora. First, we focus on the application of different techniques for the generation of pronunciation variants. Then, we explain the application of a speech recognition system for selecting the acoustically best matching phonetic transcription. The system is evaluated on different test sets selected from the Spoken Dutch Corpus, ranging from read-aloud text to spontaneous speech, and achieves promising first results. 1. Kris Demuynck, Tom Laureys, Steven Gillis |
INTERSPEECH | 1 |
| 2002 | An Improved Algorithm for the Automatic Segmentation of Speech Corpora
Tom Laureys, Kris Demuynck, Jacques Duchateau, Patrick Wambacq |
LREC | 2 |
| 2002 | Word Segmentation in the Spoken Dutch Corpus
Jean-Pierre Martens, Diana Binnenpoorte, Kris Demuynck, Ruben Van Parys, Tom Laureys, Wim Goedertier, Jacques Duchateau |
LREC | 3 |
| 2001 | Class definition in discriminant feature analysisabstractThe aim of discriminant feature analysis techniques in the signal processing of speech recognition systems is to find a feature vector transformation which maps a high dimensional input vector onto a low dimensional vector while retaining a maximum amount of information in the feature vector to discriminate between predefined classes. This paper points out the significance of the definition of the classes in the discriminant feature analysis technique. Three choices for the definition of the classes are investigated: the phonemes, the states in context independent acoustic models and the tied states in context dependent acoustic models. These choices for the classes were applied to (1) standard LDA (linear discriminant analysis) for reference and to (2) MIDA, an improved, mutual information based discriminant analysis technique. Evaluation of the resulting linear feature transforms on a large vocabulary continuous speech recognition task shows, depending on the technique, the best choice for the classes. Jacques Duchateau, Kris Demuynck, Dirk Van Compernolle, Patrick Wambacq |
INTERSPEECH | 2 |
| 2000 | Discriminative resolution enhancement in acoustic modellingabstractThe accuracy of the acoustic models in large vocabulary recognition systems can be improved by increasing the resolution in the acoustic feature space. This can be obtained by increasing the number of Gaussian densities in the models by splitting of the Gaussians. This paper proposes a novel algorithm for this splitting operation. It is based on the phonetic decision tree used for the state tying in context dependent modelling. The advantage of the method is that it improves the capability of the acoustic models to discriminate between the different tied states. The proposed splitting algorithm was evaluated on the Wall Street Journal recognition task. Comparison with a commonly used splitting algorithm clearly shows that our method can provide smaller (thus faster) acoustic models and results in lower error rates. Jacques Duchateau, Kris Demuynck, Patrick Wambacq |
ICASSP | 2 |
| 2000 | An efficient search space representation for large vocabulary continuous speech recognition
Kris Demuynck, Jacques Duchateau, Dirk Van Compernolle, Patrick Wambacq |
Speech Commun. | 1 |
| 1999 | Optimal feature sub-space selection based on discriminant analysis
Kris Demuynck, Jacques Duchateau, Dirk Van Compernolle |
EUROSPEECH | 1 |
| 1999 | Accuracy versus complexity in context dependent phone modelingabstractThis paper presents two different directions to build HMM models which give enough acoustic resolution and t in limited user resources. They both refer to scaling down the acoustic models which are built with tied gaussian HMMs. The total number of gaussians is reduced by a pairwise merging, and the number of gaussians per state is reduced by selecting them based on the so called occupancy criterion. Experiments carried out on the WSJ recognition task show that after scaling down, no further training is needed when the number of gaussians or the number of gaussians per state is reduced up to a factor three. This is an advantage as retraining can not be executed by the final system user. Jacques Duchateau, Kris Demuynck, Ioannis Dologlou, Patrick Wambacq, Dirk Van Compernolle, Hugo Van hamme |
EUROSPEECH | 3 |
| 1998 | VE platform: A Base System for Distributed Virtual RealityabstractThe VEplatform project aims at the development of a highly scalable system which allows distributed virtual reality applications to be built. The system is decentralised i.e. there is no central component which controls it and which can become a bottleneck. At this moment VEplatform is already a working system allowing many participants to join a virtual world. In this article the software architecture of VEplatform is explained and some hints to the future are given. Kris Demuynck, Jan Broeckhove, Frans Arickx |
Computer Graphics International | 1 |
| 1998 | Improved feature decorrelation for HMM-based speech recognitionabstractIn most HMM-based recognition systems, a mixture of diagonal covariance gaussians is used to model the observation density functions in the states. The use of diagonal covariance gaussians however assumes that the underlying data vectors have uncorrelated vector components: if each gaussian is replaced with its full covariant counterpart, the off-diagonal elements in the covariance matrices should be small. To that end, most recognition systems have some kind of decorrelation matrix near the end of the preprocessing. Examples are the inverse cosine transform used with cepstral coefficients, and principal component analysis (PCA) or linear discriminant analysis (LDA) of the features. However, none of these transforms is optimal if it comes to reducing the mismatch introduced by setting the off-diagonal elements in the covariance matrices to zero. The algorithm described in this paper reduces the local correlations between feature vector components inside the gaussians with a single global linear transform at the end of the preprocessing stage. The algorithm is optimal in the sense that we calculate the linear transformation that minimises the sum of the square of all off-diagonal elements over all gaussians. The algorithm is compared with principal component analysis, linear discriminant analysis and the recently published maximum likelihood modelling for semi-tied covariance matrices. The decorrelation method is also evaluated on two speech recognition tasks. A significant relative improvement was achieved in both cases. Kris Demuynck, Jacques Duchateau, Dirk Van Compernolle, Patrick Wambacq |
ICSLP | 1 |
| 1998 | Improved parameter tying for efficient acoustic model evaluation in large vocabulary continuous speech recognitionabstractIn an HMM based large vocabulary continuous speech recognition system, the evaluation of - context dependent - acoustic models is very time consuming. In Semi-Continuous HMMs, a state is modelled as a mixture of elementary - generally gaussian - probability density functions. Observation probability calculations of these states can be made faster by reducing the size of the mixture of gaussians used to model them. In this paper, we propose different criteria to decide which gaussians should remain in the mixture for a state, and which ones can be removed. The performance of the criteria is compared on context dependent tied state models using the WSJ recognition task. Our novel criterion, which decides to remove a gaussian in a state if it is based on too few acoustic data, outperforms the other described criteria. Jacques Duchateau, Kris Demuynck, Dirk Van Compernolle, Patrick Wambacq |
ICSLP | 2 |
| 1998 | The VEplatform system: A system for distributed virtual reality
Kris Demuynck, Jan Broeckhove, Frans Arickx |
Future Gener. Comput. Syst. | 1 |
| 1998 | Fast and accurate acoustic modelling with semi-continuous HMMs
Jacques Duchateau, Kris Demuynck, Dirk Van Compernolle |
Speech Commun. | 2 |
| 1997 | A static lexicon network representation for cross-word context dependent phonesabstractTo cope with the prohibitive growth of lexical tree based search-graphs when using cross-word context dependent (CD) phone models, an efficient novel search-topology was developed. The lexicon is stored as a compact static network with no language model (LM) information attached to it. The static representation avoids the cost of dynamic tree expansion, facilitates the integration of additional pronunciation information (e.g. assimilation rules) and is easier to integrate in existing search engines. Moreover, the network representation also results in a compact structure when words have alternative pronunciations, and due to its construction, it offers partial LM forwarding at no extra cost. Next, all knowledge sources (pronunciation information, language model and acoustic models) are combined by a slightly modified token-passing algorithm, resulting in a one pass time-synchronous recognition system. Kris Demuynck, Jacques Duchateau, Dirk Van Compernolle |
EUROSPEECH | 1 |
| 1997 | A novel node splitting criterion in decision tree construction for semi-continuous HMMs
Jacques Duchateau, Kris Demuynck, Dirk Van Compernolle |
EUROSPEECH | 2 |
| 1996 | Reduced semi-continuous models for large vocabulary continuous speech recognition in DutchabstractSemi-continuous Density HMM's have -due to the decoupling between the set of gaussians and the other HMM-parameters -more possibilities than Continuous Density HMM's to match the number of parameters in the model to the available train data.The computational load of the SC-HMM's however is huge compared to the load of their continuous counterparts, because of the large mixture weighting vector and because of the fact that for each frame all gaussians have to be evaluated.This paper describes the different steps taken to reduce the computational load of the SC-HMM's, resulting in faster and better models. Kris Demuynck, Jacques Duchateau, Dirk Van Compernolle |
ICSLP | 1 |
| 1994 | Pseudo-segment based speech recognition using neural recurrent whole-word recognizersabstractDescribes a recurrent neural network based, isolated word speech recognizer. The recognizer uses 2 MLPs. A first, static MLP is used for classification of frames in phonemes. Next, a time compression step is applied. The resulting pseudo-segments are then used as inputs for a second, dynamic MLP that integrates the information over time to decide the current word. The authors apply this approach on an isolated digit recognition task and compare the results with hybrid MLP/HMM approach using the same static MLP.> Philippe Le Cerf, Kris Demuynck, Jacques Duchateau, Dirk Van Compernolle |
ICASSP (1) | 2 |