Hugo Van hamme

dblp:77/6519 · DBLP profile ↗
← Back
182ranked-venue papers
11as first author
33since 2021 · last 2025
0000-0003-1331-5186ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 152 · 11 first-author · 31 since 2021Artificial intelligence and machine learning · 106 · 8 first-author · 21 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2025 Graph Connectionist Temporal Classification for Phoneme Recognition
abstract
Automatic Phoneme Recognition (APR) systems are often trained using pseudo phoneme-level annotations generated from text through Grapheme-to-Phoneme (G2P) systems. These G2P systems frequently output multiple possible pronunciations per word, but the standard Connectionist Temporal Classification (CTC) loss cannot account for such ambiguity during training. In this work, we adapt Graph Temporal Classification (GTC) to the APR setting. GTC enables training from a graph of alternative phoneme sequences, allowing the model to consider multiple pronunciations per word as valid supervision. Our experiments on English and Dutch data sets show that incorporating multiple pronunciations per word into the training loss consistently improves phoneme error rates compared to a baseline trained with CTC. These results suggest that integrating pronunciation variation into the loss function is a promising strategy for training APR systems from noisy G2P-based supervision.
Henry Grafé, Hugo Van hamme
ASRU2
2025 SSVD: Structured SVD for Parameter-Efficient Fine-Tuning and Benchmarking under Domain Shift in ASR
abstract
Parameter-efficient fine-tuning (PEFT) has emerged as a scalable solution for adapting large foundation models. While low-rank adaptation (LoRA) is widely used in speech applications, its state-of-the-art variants, e.g., VeRA, DoRA, PiSSA, and SVFT, are developed mainly for language and vision tasks, with limited validation in speech. This work presents the first comprehensive integration and benchmarking of these PEFT methods within ESPnet. We further introduce structured SVDguided (SSVD) fine-tuning, which selectively rotates input-associated right singular vectors while keeping output-associated vectors fixed to preserve semantic mappings. This design enables robust domain adaptation with minimal trainable parameters and improved efficiency. We evaluate all methods on domain-shifted speech recognition tasks, including child speech and dialectal variation, across model scales from 0.1B to 2B. All implementations are released in ESPnet to support reproducibility and future work.
Pu Wang 0007, Shinji Watanabe 0001, Hugo Van hamme
ASRU3
2025 Self-Incremental Training for Personalized Voice Command Recognition in a Wireless Audio Sensor Network
abstract
This paper studies self-incremental training in the context of personalized Deep Neural Networks (DNNs) for voice command recognition tailored for resource-constrained sensor nodes. The learning task runs when new unsupervised data becomes available within a Wireless Audio Sensor Network (WASN). After collecting a new multi-sensor dataset of voice commands, we experimentally investigate network-level policies to assign pseudo-labels to the new data. Our baseline analysis shows an accuracy improvement of up to +15% with respect to models pretrained on a large keyword corpus dataset. The multi-sensor labeling strategy closely approximates the performance achieved in a single-sensor scenario providing a clean signal, while we observe +4.7% compared to other sensors with degraded signal quality.
Manuele Rusci, Hugo Van hamme, Tinne Tuytelaars
ICASSP2
2025 Leveraging Geographic Metadata for Dialect-Aware Speech Recognition
abstract
status: Published
Pouya Mehralian, Hugo Van hamme
INTERSPEECH2
2025 Challenges and practical guidelines for atypical speech data collection, annotation, usage and sharing: A multi-project perspective
abstract
Contains fulltext : 325867.pdf (Publisher’s version ) (Open Access)
Zhengjun Yue, Mara Barberis, Tanvina Patel, Judith Dineley, Willemijn Doedens, Lottie Stipdonk, Elke De Witte, Erfan Loweimi, Hugo Van hamme, Djaina Satoer, Marina B. Ruiter, Laureano Moro-Velázquez, Nicholas Cummins, Odette Scharenborg
INTERSPEECH10
2025 Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation
abstract
End-to-end transformer-based automatic speech recognition (ASR) systems often capture multiple speech traits in their learned representations that are highly entangled, leading to a lack of interpretability. In this study, we propose the explainable Disentangled-Transformer, which disentangles the internal representations into sub-embeddings with explicit content and speaker traits based on varying temporal resolutions. Experimental results show that the proposed Disentangled-Transformer produces a clear speaker identity, separated from the speech content, for speaker diarization while improving ASR performance.
Pu Wang 0007, Hugo Van hamme
IPAS2
2024 Unsupervised Accent Adaptation Through Masked Language Model Correction of Discrete Self-Supervised Speech Units
abstract
Self-supervised pre-trained speech models have strongly improved speech recognition, yet they are still sensitive to domain shifts and accented or atypical speech. Many of these models rely on quantisation or clustering to learn discrete acoustic units. We propose to correct the discovered discrete units for accented speech back to a standard pronunciation in an unsupervised manner. A masked language model is trained on discrete units from a standard accent and iteratively corrects an accented token sequence by masking unexpected cluster sequences and predicting their common variant. Small accent adapter blocks are inserted in the pre-trained model and fine-tuned by predicting the corrected clusters, which leads to an increased robustness of the pre-trained model towards a target accent, and this without supervision. We are able to improve a state-of-the-art HuBERT Large model on a downstream accented speech recognition task by altering the training regime with the proposed method.
Jakob Poncelet, Hugo Van hamme
ICASSP2
2024 Automatic recognition and detection of aphasic natural speech
abstract
sponsorship: Fonds Wetenschappelijk Onderzoek|1SH1Q24N
Mara Barberis, Pieter De Clercq, Bastiaan Tamm, Hugo Van hamme, Maaike Vandermosten
INTERSPEECH4
2024 Unsupervised Online Continual Learning for Automatic Speech Recognition
abstract
sponsorship: Research supported by Research Foundation Flanders (FWO) under grant S004923N of the SBO programme. (Research Foundation Flanders (FWO)|S004923N)
Steven Vander Eeckt, Hugo Van hamme
INTERSPEECH2
2024 Efficient Extraction of Noise-Robust Discrete Units from Self-Supervised Speech Models
abstract
Continuous speech can be converted into a discrete sequence by deriving discrete units from the hidden features of self-supervised learned (SSL) speech models. Although SSL models are becoming larger and trained on more data, they are often sensitive to real-life distortions like additive noise or reverberation, which translates to a shift in discrete units. We propose a parameter-efficient approach to generate noiserobust discrete units from pre-trained SSL models by training a small encoder-decoder model, with or without adapters, to simultaneously denoise and discretise the hidden features of the SSL model. The model learns to generate a clean discrete sequence for a noisy utterance, conditioned on the SSL features. The proposed denoiser outperforms several pre-training methods on the tasks of noisy discretisation and noisy speech recognition, and can be finetuned to the target environment with a few recordings of unlabeled target data.
Jakob Poncelet, Hugo Van hamme
SLT3
2024 Scale-aware dual-branch complex convolutional recurrent network for monaural speech enhancement
Meng Sun 0001, Xiongwei Zhang, Hugo Van hamme
Comput. Speech Lang.4
2023 Whisper-Slu: Extending a Pretrained Speech-to-Text Transformer for Low Resource Spoken Language Understanding
abstract
Human-computer interactions require systems that work out of the box without requiring lots of data to adapt to a new task or user. In this research, we address low resource spoken language understanding tasks such as named entity recognition (NER), intent recognition (IR), and slot filling (SF) to research how a pretrained model can be modified for a new task, then finetuned with few labelled data. We propose extending the Whisper model with task-specific modules for NER, SF, and IR, leveraging a Markov network as output structure. We develop a novel approach to finetuning by removing irrelevant weights and reorganizing the embeddings to drastically improve the performance in a low-resource setting. Our approach outperforms previous models without external language models and demonstrates effective transfer learning, even with very limited training data. The models exhibit a small footprint, making them suitable for applications requiring robustness, few-shot learning, and efficiency.
Quentin Meeus, Marie-Francine Moens, Hugo Van hamme
ASRU3
2023 ICASSP 2023 Auditory EEG Decoding Challenge
abstract
This paper describes the auditory EEG challenge which was organized as one of the Signal Processing Grand Challenges of ICASSP 2023. This challenge consists of two tasks in which the goal is to relate electroencephalogram (EEG) signals to the presented speech stimulus. In the first task, named match-mismatch, the goal is to determine which of the two speech segments matches with a given EEG segment. In the second task, a regression task, the goal is to reconstruct the speech envelope from the EEG.
Lies Bollens, Mohammad Jalilpour-Monesi, Bernd Accou, Jonas Vanthornhout, Hugo Van hamme, Tom Francart
ICASSP5
2023 Weight Averaging: A Simple Yet Effective Method to Overcome Catastrophic Forgetting in Automatic Speech Recognition
abstract
Adapting a trained Automatic Speech Recognition (ASR) model to new tasks results in catastrophic forgetting of old tasks, limiting the model’s ability to learn continually and to be extended to new speakers, dialects, languages, etc. Focusing on End-to-End ASR, in this paper, we propose a simple yet effective method to overcome catastrophic forgetting: weight averaging. By simply taking the average of the previous and the adapted model, our method achieves high performance on both the old and new tasks. It can be further improved by introducing a knowledge distillation loss during the adaptation. We illustrate the effectiveness of our method on both monolingual and multilingual ASR. In both cases, our method strongly outperforms all baselines, even in its simplest form.
Steven Vander Eeckt, Hugo Van hamme
ICASSP2
2023 Using Adapters to Overcome Catastrophic Forgetting in End-to-End Automatic Speech Recognition
abstract
Learning a set of tasks in sequence remains a challenge for artificial neural networks, which, in such scenarios, tend to suffer from Catastrophic Forgetting (CF). The same applies to End-to-End (E2E) Automatic Speech Recognition (ASR) models, even for monolingual tasks. In this paper, we aim to overcome CF for E2E ASR by inserting adapters, small architectures of few parameters which allow a general model to be fine-tuned to a specific task, into our model. We make these adapters task-specific, while regularizing the parameters of the model shared by all tasks, thus stimulating the model to fully exploit the adapters while keeping the shared parameters to work well for all tasks. Our method outperforms all baselines on two monolingual experiments while being more storage efficient and without requiring the storage of data from previous tasks.
Steven Vander Eeckt, Hugo Van hamme
ICASSP2
2023 Cross-Lingual Transfer Learning for Alzheimer's Detection from Spontaneous Speech
abstract
Alzheimer’s disease (AD) is a progressive neurodegenerative disease most often associated with memory deficits and cognitive decline. With the aging population, there has been much interest in automated methods for cognitive impairment detection. One approach that has attracted attention in recent years is AD detection through spontaneous speech. While the results are promising, it is not certain whether the learned speech features can be generalized across languages. To fill this gap, the ADReSS-M challenge was organized. This paper presents our submission to this ICASSP-2023 Signal Processing Grand Challenge (SPGC). The model was trained on 228 English samples of a picture description task and was transferred to Greek using only 8 samples. We obtained an accuracy of 82.6% for AD detection, a root-mean-square error of 4.345 for cognitive score prediction, and ranked 2nd place in the competition out of 24 competitors.
Bastiaan Tamm, Rik Vandenberghe, Hugo Van hamme
ICASSP3
2023 Rehearsal-Free Online Continual Learning for Automatic Speech Recognition
abstract
sponsorship: Research supported by Research Foundation Flanders (FWO) under grant S004923N of the SBO programme. (Research Foundation Flanders (FWO) SBO programme|S004923N)
Steven Vander Eeckt, Hugo Van hamme
INTERSPEECH2
2023 Parameter-efficient Dysarthric Speech Recognition Using Adapter Fusion and Householder Transformation
abstract
sponsorship: The research was supported by KU Leuven Special Research Fund grant C24M/22/025 and the Flemish Government under the "Onderzoeksprogramma Artificiele Intelligentie (AI) Vlaanderen" programme. (KU Leuven Special Research Fund|C24M/22/025, Flemish Government under the "Onderzoeksprogramma Artificiele Intelligentie (AI) Vlaanderen" programme)
Jinzi Qi, Hugo Van hamme
INTERSPEECH2
2022 Learning Subject-Invariant Representations from Speech-Evoked EEG Using Variational Autoencoders
abstract
The electroencephalogram (EEG) is a powerful method to understand how the brain processes speech. Linear models have recently been replaced for this purpose with deep neural networks and yield promising results. In related EEG classification fields, it is shown that explicitly modeling subject-invariant features improves generalization of models across subjects and benefits classification accuracy. In this work, we adapt factorized hierarchical variational autoencoders to exploit parallel EEG recordings of the same stimuli. We model EEG into two disentangled latent spaces. Subject accuracy reaches 98.96% and 1.60% on respectively the subject and content latent space, whereas binary content classification experiments reach an accuracy of 51.51% and 62.91% on respectively the subject and content latent space.
Lies Bollens, Tom Francart, Hugo Van hamme
ICASSP3
2022 Multitask Learning for Low Resource Spoken Language Understanding
abstract
We explore the benefits that multitask learning offer to speech processing as we train models on dual objectives with automatic speech recognition and intent classification or sentiment classification. Our models, although being of modest size, show improvements over models trained end-to-end on intent classification. We compare different settings to find the optimal disposition of each task module compared to one another. Finally, we study the performance of the models in low-resource scenario by training the models with as few as one example per class. We show that multitask learning in these scenarios compete with a baseline model trained on text features and performs considerably better than a pipeline model. On sentiment classification, we match the performance of an end-to-end model with ten times as many parameters. We consider 4 tasks and 4 datasets in Dutch and English.
Quentin Meeus, Marie-Francine Moens, Hugo Van hamme
INTERSPEECH3
2022 Relating the fundamental frequency of speech with EEG using a dilated convolutional network
abstract
sponsorship: The authors thank all the subjects for the recordings as well as Wendy Verheijen, Bernd Accou, Kyara Cloes, Amelie Algoet, Jolien Smeulders, Lore Kerkhofs, Sara Peeters, Merel Dillen, Ilham Gamgami, Amber Verhoeven, Lies Bollens, Vitor Vasconcelos and Amber Aerts for their help with data collection. Funding was provided by the KU Leuven Special Research Fund C24/18/099 (C2 project to Tom Francart and Hugo Van hamme), FWO research project G0D6720N, and an FWO post-doctoral fellowship to Jonas Vanthornhout (1290821N). (KU Leuven Special Research Fund|C24/18/099, FWO research project|G0D6720N, FWO|1290821N)
Corentin Puffay, Jana Van Canneyt, Jonas Vanthornhout, Hugo Van hamme, Tom Francart
INTERSPEECH4
2022 Pre-trained Speech Representations as Feature Extractors for Speech Quality Assessment in Online Conferencing Applications
abstract
Speech quality in online conferencing applications is typically assessed through human judgements in the form of the mean opinion score (MOS) metric. Since such a labor-intensive approach is not feasible for large-scale speech quality assessments in most settings, the focus has shifted towards automated MOS prediction through end-to-end training of deep neural networks (DNN). Instead of training a network from scratch, we propose to leverage the speech representations from the pre-trained wav2vec-based XLS-R model. However, the number of parameters of such a model exceeds task-specific DNNs by several orders of magnitude, which poses a challenge for resulting fine-tuning procedures on smaller datasets. Therefore, we opt to use pre-trained speech representations from XLS-R in a feature extraction rather than a fine-tuning setting, thereby significantly reducing the number of trainable model parameters. We compare our proposed XLS-R-based feature extractor to a Mel-frequency cepstral coefficient (MFCC)-based one, and experiment with various combinations of bidirectional long short term memory (Bi-LSTM) and attention pooling feedforward (AttPoolFF) networks trained on the output of the feature extractors. We demonstrate the increased performance of pre-trained XLS-R embeddings in terms a reduced root mean squared error (RMSE) on the ConferencingSpeech 2022 MOS prediction task.
Bastiaan Tamm, Helena Balabin, Rik Vandenberghe, Hugo Van hamme
INTERSPEECH4
2022 Bottleneck Low-rank Transformers for Low-resource Spoken Language Understanding
abstract
sponsorship: The research was supported by the program of China Scholarship Council No.201906090275 and the Flemish Government under "Onderzoeksprogramma AI Vlaanderen". (China Scholarship Council|201906090275, Flemish Government under "Onderzoeksprogramma AI Vlaanderen")
Pu Wang 0007, Hugo Van hamme
INTERSPEECH2
2022 Learning to Jointly Transcribe and Subtitle for End-To-End Spontaneous Speech Recognition
abstract
TV subtitles are a rich source of transcriptions of many types of speech, ranging from read speech in news reports to conversational and spontaneous speech in talk shows and soaps. However, subtitles are not verbatim (i.e. exact) transcriptions of speech, so they cannot be used directly to improve an Automatic Speech Recognition (ASR) model. We propose a multitask dual-decoder Transformer model that jointly performs ASR and automatic subtitling. The ASR decoder (possibly pre-trained) predicts the verbatim output and the subtitle decoder generates a subtitle, while sharing the encoder. The two decoders can be independent or connected. The model is trained to perform both tasks jointly, and is able to effectively use subtitle data. We show improvements on regular ASR and on spontaneous and conversational ASR by incorporating the additional subtitle decoder. The method does not require preprocessing (aligning, filtering, pseudo-labeling,…) of the subtitles.
Jakob Poncelet, Hugo Van hamme
SLT2
2022 Weak-Supervised Dysarthria-Invariant Features for Spoken Language Understanding Using an Fhvae and Adversarial Training
abstract
The scarcity of training data and the large speaker variation in dysarthric speech lead to poor accuracy and poor speaker generalization of spoken language understanding systems for dysarthric speech. Through work on the speech features, we focus on improving the model generalization ability with limited dysarthric data. Factorized Hierarchical Variational Auto-Encoders (FHVAE) trained unsupervisedly have shown their advantage in disentangling content and speaker representations. Earlier work showed that the dysarthria shows in both feature vectors. Here, we add adversarial training to bridge the gap between the control and dysarthric speech data domains. We extract dysarthric and speaker invariant features using weak supervision. The extracted features are evaluated on a Spoken Language Understanding task and yield a higher accuracy on unseen speakers with more severe dysarthria compared to features from the basic FHVAE model or plain filterbanks.
Jinzi Qi, Hugo Van hamme
SLT2
2021 Comparison of Self-Supervised Speech Pre-Training Methods on Flemish Dutch
abstract
Recent research in speech processing exhibits a growing interest in unsupervised and self-supervised representation learning from unlabelled data to alleviate the need for large amounts of annotated data. We investigate several popular pre-training methods and apply them to Flemish Dutch. We compare off-the-shelf English pre-trained models to models trained on an increasing amount of Flemish data. We find that the most important factors for positive transfer to downstream speech recognition tasks include a substantial amount of data and a matching pre-training domain. Ideally, we also finetune on an annotated subset in the target language. All pre-trained models improve linear phone separability in Flemish, but not all methods improve Automatic Speech Recognition. We experience superior performance with wav2vec 2.0 and we obtain a 30% WER improvement by finetuning the multilingually pre-trained XLSR-53 model on Flemish Dutch, after integration into an HMM-DNN acoustic model.
Jakob Poncelet, Hugo Van hamme
ASRU2
2021 Audiovisual Transfer Learning for Audio Tagging and Sound Event Detection
abstract
We study the merit of transfer learning for two sound recognition problems, i.e., audio tagging and sound event detection. Employing feature fusion, we adapt a baseline system utilizing only spectral acoustic inputs to also make use of pretrained auditory and visual features, extracted from networks built for different tasks and trained with external data. We perform experiments with these modified models on an audiovisual multi-label data set, of which the training partition contains a large number of unlabeled samples and a smaller amount of clips with weak annotations, indicating the clip-level presence of 10 sound categories without specifying the temporal boundaries of the active auditory events. For clip-based audio tagging, this transfer learning method grants marked improvements. Addition of the visual modality on top of audio also proves to be advantageous in this context. When it comes to generating transcriptions of audio recordings, the benefit of pretrained features depends on the requested temporal resolution: for coarse-grained sound event detection, their utility remains notable. But when more fine-grained predictions are required, performance gains are strongly reduced due to a mismatch between the problem at hand and the goals of the models from which the pretrained vectors were obtained.
Wim Boes, Hugo Van hamme
Interspeech2
2021 Extracting Different Levels of Speech Information from EEG Using an LSTM-Based Model
abstract
Decoding the speech signal that a person is listening to from the human brain via electroencephalography (EEG) can help us understand how our auditory system works. Linear models have been used to reconstruct the EEG from speech or vice versa. Recently, Artificial Neural Networks (ANNs) such as Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) based architectures have outperformed linear models in modeling the relation between EEG and speech. Before attempting to use these models in real-world applications such as hearing tests or (second) language comprehension assessment we need to know what level of speech information is being utilized by these models. In this study, we aim to analyze the performance of an LSTM-based model using different levels of speech features. The task of the model is to determine which of two given speech segments is matched with the recorded EEG. We used low- and high-level speech features including: envelope, mel spectrogram, voice activity, phoneme identity, and word embedding. Our results suggest that the model exploits information about silences, intensity, and broad phonetic classes from the EEG. Furthermore, the mel spectrogram, which contains all this information, yields the highest accuracy (84%) among all the features.
Mohammad Jalilpour-Monesi, Bernd Accou, Tom Francart, Hugo Van hamme
Interspeech4
2021 Speech Disorder Classification Using Extended Factorized Hierarchical Variational Auto-Encoders
abstract
sponsorship: The research was supported by KUL grant CELSA/18/027 and the Flemish Government under "Onderzoeksprogramma AI Vlaanderen". (KUL grant|CELSA/18/027, Flemish Government under "Onderzoeksprogramma AI Vlaanderen")
Jinzi Qi, Hugo Van hamme
Interspeech2
2021 A Study into Pre-Training Strategies for Spoken Language Understanding on Dysarthric Speech
abstract
sponsorship: The research was supported by the program of China Scholarship Council No. 201906090275, KUL grant CELSA/18/027 and the Flemish Government under "Onderzoeksprogramma AI Vlaanderen". (China Scholarship Council|201906090275, KUL grant|CELSA/18/027, Flemish Government under "Onderzoeksprogramma AI Vlaanderen")
Pu Wang 0007, Bagher BabaAli, Hugo Van hamme
Interspeech3
2021 A Light Transformer For Speech-To-Intent Applications
abstract
Spoken language understanding (SLU) systems can make life more agreeable, safer (e.g. in a car) or can increase the independence of physically challenged users. However, due to the many sources of variation in speech, a well-trained system is hard to transfer to other conditions like a different language or to speech impaired users. A remedy is to design a user-taught SLU system that can learn fully from scratch from users' demonstrations, which in turn requires that the system's model quickly converges after only a few training samples. In this paper, we propose a light transformer structure by using a simplified relative position encoding with the goal to reduce the model size and improve efficiency. The light transformer works as an alternative speech encoder for an existing user-taught multitask SLU system. Experimental results on three datasets with challenging speech conditions prove our approach outperforms the existed system and other state-of-art models with half of the original model size and training time.
Pu Wang 0007, Hugo Van hamme
SLT2
2021 Low resource end-to-end spoken language understanding with capsule networks
Jakob Poncelet, Vincent Renkens, Hugo Van hamme
Comput. Speech Lang.3
2021 Show me where the action is!
abstract
Abstract Reality TV shows have gained popularity, motivating many production houses to bring new variants for us to watch. Compared to traditional TV shows, reality TV shows have spontaneous unscripted footage. Computer vision techniques could partially replace the manual labour needed to record and process this spontaneity. However, automated real-world video recording and editing is a challenging topic. In this paper, we propose a system that utilises state-of-the-art video and audio processing algorithms to, on the one hand, automatically steer cameras, replacing camera operators and on the other hand, detect all audiovisual action cues in the recorded video, to ease the job of the film editor. This publication has hence two main contributions. The first, automating the steering of multiple Pan-Tilt-Zoom PTZ cameras to take aesthetically pleasing medium shots of all the people present. These shots need to comply with the cinematographic rules and are based on the poses acquired by a pose detector. Secondly, when a huge amount of audio-visual data has been collected, it becomes labour intensive for a human editor retrieve the relevant fragments. As a second contribution, we combine state-of-the-art audio and video processing techniques for sound activity detection, action recognition, face recognition, and pose detection to decrease the required manual labour during and after recording. These techniques used during post-processing produce meta-data allowing for footage filtering, decreasing the search space. We extended our system further by producing timelines uniting generated meta-data, allowing the editor to have a quick overview. We evaluated our system on three in-the-wild reality TV recording sessions of 24 hours (× 8 cameras) each taken in real households.
Timothy Callemein, Tom Roussel, Ali Diba, Floris De Feyter, Wim Boes, Luc Van Eycken, Luc Van Gool, Hugo Van hamme, Tinne Tuytelaars, Toon Goedemé
Multim. Tools Appl.8
2020 On the long-term learning ability of LSTM LMs
Wim Boes, Robbe Van Rompaey, Lyan Verwimp, Joris Pelemans, Hugo Van hamme, Patrick Wambacq
ESANN5
2020 An LSTM Based Architecture to Relate Speech Stimulus to Eeg
abstract
Modeling the relationship between natural speech and a recorded electroencephalogram (EEG) helps us understand how the brain processes speech and has various applications in neuroscience and brain-computer interfaces. In this context, so far mainly linear models have been used. However, the decoding performance of the linear model is limited due to the complex and highly non-linear nature of the auditory processing in the human brain. We present a novel Long Short-Term Memory (LSTM)-based architecture as a nonlinear model for the classification problem of whether a given pair of (EEG, speech envelope) correspond to each other or not. The model maps short segments of the EEG and the envelope to a common embedding space using a CNN in the EEG path and an LSTM in the speech path. The latter also compensates for the brain response delay. In addition, we use transfer learning to fine-tune the model for each subject. The mean classification accuracy of the proposed model reaches 85%, which is significantly higher than that of a state of the art Convolutional Neural Network (CNN)-based model (73%) and the linear model (69%).
Mohammad Jalilpour-Monesi, Bernd Accou, Jair Montoya-Martínez, Tom Francart, Hugo Van hamme
ICASSP5
2020 Multitask Learning with Capsule Networks for Speech-to-Intent Applications
abstract
Voice controlled applications can be a great aid to society, especially for physically challenged people. However this requires robustness to all kinds of variations in speech. A spoken language understanding system that learns from interaction with and demonstrations from the user, allows the use of such a system in different settings and for different types of speech, even for deviant or impaired speech, while also allowing the user to choose a phrasing. The user gives a command and enters its intent through an interface, after which the model learns to map the speech directly to the right action. Since the effort of the user should be as low as possible, capsule networks have drawn interest due to potentially needing little training data compared to deeper neural networks. In this paper, we show how capsules can incorporate multitask learning, which often can improve the performance of a model when the task is difficult. The basic capsule network will be expanded with a regularisation to create more structure in its output: it learns to identify the speaker of the utterance by forcing the required information into the capsule vectors. To this end we move from a speaker dependent to a speaker independent setting.
Jakob Poncelet, Hugo Van hamme
ICASSP2
2020 State gradients for analyzing memory in LSTM language models
Lyan Verwimp, Hugo Van hamme, Patrick Wambacq
Comput. Speech Lang.2
2019 Practical Applicability of Deep Neural Networks for Overlapping Speaker Separation
abstract
status: Published
Pieter Appeltans, Jeroen Zegers, Hugo Van hamme
INTERSPEECH3
2019 CNN-LSTM Models for Multi-Speaker Source Separation Using Bayesian Hyper Parameter Optimization
abstract
sponsorship: This work was funded by the SB PhD grant of the Research Foundation Flanders (FWO) with project number 1S66217N. (SB PhD grant of the Research Foundation Flanders (FWO)|1S66217N)
Jeroen Zegers, Hugo Van hamme
INTERSPEECH2
2019 Audiovisual Transformer Architectures for Large-Scale Classification and Synchronization of Weakly Labeled Audio Events
abstract
We tackle the task of environmental event classification by drawing inspiration from the transformer neural network architecture used in machine translation. We modify this attention-based feedforward structure in such a way that allows the resulting model to use audio as well as video to compute sound event predictions. We perform extensive experiments with these adapted transformers on an audiovisual data set, obtained by appending relevant visual information to an existing large-scale weakly labeled audio collection. The employed multi-label data contains clip-level annotation indicating the presence or absence of 17 classes of environmental sounds, and does not include temporal information. We show that the proposed modified transformers strongly improve upon previously introduced models and in fact achieve state-of-the-art results. We also make a compelling case for devoting more attention to research in multimodal audiovisual classification by proving the usefulness of visual information for the task at hand, namely audio event recognition. In addition, we visualize internal attention patterns of the audiovisual transformers and in doing so demonstrate their potential for performing multimodal synchronization.
Wim Boes, Hugo Van hamme
ACM Multimedia2
2019 Hyperspectral image classification using Non-negative Tensor Factorization and 3D Convolutional Neural Networks
Sayeh Mirzaei, Hugo Van hamme, Shima Khosravani
Signal Process. Image Commun.2
2018 Multi-Scenario Deep Learning for Multi-Speaker Source Separation
abstract
Research in deep learning for multi-speaker source separation has received a boost in the last years. However, most studies are restricted to mixtures of a specific number of speakers, called a specific scenario. While some works included experiments for different scenarios, research towards combining data of different scenarios or creating a single model for multiple scenarios have been very rare. In this work it is shown that data of a specific scenario is relevant for solving another scenario. Furthermore, it is concluded that a single model, trained on different scenarios is capable of matching performance of scenario specific models.
Jeroen Zegers, Hugo Van hamme
ICASSP2
2018 Capsule Networks for Low Resource Spoken Language Understanding
abstract
Designing a spoken language understanding system for command-and-control applications can be challenging because of a wide variety of domains and users or because of a lack of training data. In this paper we discuss a system that learns from scratch from user demonstrations. This method has the advantage that the same system can be used for many domains and users without modifications and that no training data is required prior to deployment. The user is required to train the system, so for a user friendly experience it is crucial to minimize the required amount of data. In this paper we investigate whether a capsule network can make efficient use of the limited amount of available training data. We compare the proposed model to an approach based on Non-negative Matrix Factorisation which is the state-of-the-art in this setting and another deep learning approach that was recently introduced for end-to-end spoken language understanding. We show that the proposed model outperforms the baseline models for three command-and-control applications: controlling a small robot, a vocally guided card game and a home automation task.
Vincent Renkens, Hugo Van hamme
INTERSPEECH2
2018 State Gradients for RNN Memory Analysis
Lyan Verwimp, Hugo Van hamme, Vincent Renkens, Patrick Wambacq
INTERSPEECH2
2018 Memory Time Span in LSTMs for Multi-Speaker Source Separation
abstract
With deep learning approaches becoming state-of-the-art in many speech (as well as non-speech) related machine learning tasks, efforts are being taken to delve into the neural networks which are often considered as a black box. In this paper it is analyzed how recurrent neural network (RNNs) cope with temporal dependencies by determining the relevant memory time span in a long short-term memory (LSTM) cell. This is done by leaking the state variable with a controlled lifetime and evaluating the task performance. This technique can be used for any task to estimate the time span the LSTM exploits in that specific scenario. The focus in this paper is on the task of separating speakers from overlapping speech. We discern two effects: A long term effect, probably due to speaker characterization and a short term effect, probably exploiting phone-size formant tracks.
Jeroen Zegers, Hugo Van hamme
INTERSPEECH2
2018 TF-LM: TensorFlow-based Language Modeling Toolkit
Lyan Verwimp, Hugo Van hamme, Patrick Wambacq
LREC2
2018 The CAMETRON Lecture Recording System: High Quality Video Recording and Editing with Minimal Human Supervision
Dries Hulens, Bram Aerts, Punarjay Chakravarty, Ali Diba, Toon Goedemé, Tom Roussel, Jeroen Zegers, Tinne Tuytelaars, Luc Van Eycken, Luc Van Gool, Hugo Van hamme, Joost Vennekens
MMM (1)11
2018 Information-Weighted Neural Cache Language Models for ASR
abstract
Neural cache language models (LMs) extend the idea of regular cache language models by making the cache probability dependent on the similarity between the current context and the context of the words in the cache. We make an extensive comparison of `regular' cache models with neural cache models, both in terms of perplexity and WER after rescoring first-pass ASR results. Furthermore, we propose two extensions to this neural cache model that make use of the content value/information weight of the word: firstly, combining the cache probability and LM probability with an information-weighted interpolation and secondly, selectively adding only content words to the cache. We obtain a 29.9%/32.1% (validation/test set) relative improvement in perplexity with respect to a baseline LSTM LM on theWikiText-2 dataset, outperforming previous work on neural cache LMs. Additionally, we observe significant WER reductions with respect to the baseline model on the WSJ ASR task.
Lyan Verwimp, Joris Pelemans, Hugo Van hamme, Patrick Wambacq
SLT3
2017 Character-Word LSTM Language Models
abstract
We present a Character-Word Long Short-Term Memory Language Model which both reduces the perplexity with respect to a baseline word-level language model and reduces the number of parameters of the model.Character information can reveal structural (dis)similarities between words and can even be used when a word is out-of-vocabulary, thus improving the modeling of infrequent and unknown words.By concatenating word and character embeddings, we achieve up to 2.77% relative improvement on English compared to a baseline model with a similar amount of parameters and 4.57% on Dutch.Moreover, we also outperform baseline word-level models with a larger number of parameters.
Lyan Verwimp, Joris Pelemans, Hugo Van hamme, Patrick Wambacq
EACL (1)3
2017 Improving Source Separation via Multi-Speaker Representations
abstract
Copyright © 2017 ISCA. Lately there have been novel developments in deep learning towards solving the cocktail party problem. Initial results are very promising and allow for more research in the domain. One technique that has not yet been explored in the neural network approach to this task is speaker adaptation. Intuitively, information on the speakers that we are trying to separate seems fundamentally important for the speaker separation task. However, retrieving this speaker information is challenging since the speaker identities are not known a priori and multiple speakers are simultaneously active. There is thus some sort of chicken and egg problem. To tackle this, source signals and i-vectors are estimated alternately. We show that blind multi-speaker adaptation improves the results of the network and that (in our case) the network is not capable of adequately retrieving this useful speaker information itself.
Jeroen Zegers, Hugo Van hamme
INTERSPEECH2
2017 Automatic relevance determination for nonnegative dictionary learning in the gamma-Poisson model
Vincent Renkens, Hugo Van hamme
Signal Process.2
2017 Joint Denoising and Dereverberation Using Exemplar-Based Sparse Representations and Decaying Norm Constraint
abstract
Exemplar-based nonnegative models, where the noisy speech is decomposed as a sparse nonnegative linear combination of the speech and noise exemplars stored in a dictionary, have been successfully used for speech denoising. This paper extends this technique for the single-channel speech enhancement in noisy reverberant environments using a novel approximation of the noisy reverberant speech in the frequency domain and nonnegative matrix deconvolution. In the proposed model, the room impulse response (RIR) in the magnitude short-time Fourier transform domain is defined such that its decaying structure can also be estimated from the test data itself, whereas the existing models used a suboptimal binwise clamping procedure to impose such a decaying structure that does not hold in a typical RIR. This paper presents multiplicative updates for estimating the RIR, its decay, and the underlying anechoic speech and noise. The proposed model is evaluated on a synthetically created dataset created by convolving TIMIT recordings with RIRs measured from different rooms and varying speaker-and-microphone locations, and adding background noises taken from the CHiME corpus. Simulation results show that the proposed model results in a better RIR estimate over the existing model and improves various instrumental speech quality measures.
Deepak Baby, Hugo Van hamme
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Weakly Supervised Learning of Hidden Markov Models for Spoken Language Acquisition
abstract
In this paper, a spoken command and control interface that acquires spoken language through demonstrations from the user is discussed. The user can train the system by uttering a command and subsequently demonstrating the required action through an alternative interface. From the demonstration, a bag of semantic concepts representation that represents which semantic concepts are present in the demonstration is extracted. In the previous work, we have proposed a method for learning words for these concepts by linking the bag of semantic concepts representation to a bag of features representation of the acoustics. In this method, the order in which the words occur is lost. However, in many cases, the order in which the words occur is important to be able to determine the correct action. In this paper, the vocabulary acquisition based on nonnegative matrix factorization is jointly trained with a hidden Markov model (HMM), making it possible to use the bag of concepts representation as a weak supervision for HMM learning. This model can better utilize the timing information to improve the results and the order in which the words occur is retained making it possible to learn vocabulary and grammar. The proposed system is tested on several command and control tasks and it is shown that for unimpaired speech the resulting system outperforms the system solely based on vocabulary acquisition.
Vincent Renkens, Hugo Van hamme
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Supervised speech dereverberation in noisy environments using exemplar-based sparse representations
abstract
Exemplar-based techniques, where the noisy speech is decomposed as a linear combination of the speech and noise exemplars stored in a dictionary, have been successfully used for speech enhancement in noisy environments. This paper extends this technique to achieve speech dereverberation in noisy environments by means of a nonnegative approximation of the noisy reverberant speech in the frequency domain. A novel approach for estimating the room impulse response (RIR) together with the speech and noise estimates using a non-negative matrix deconvolution (NMD)-based technique is proposed. In addition, we extend an existing technique based on nonnegative matrix factorisation (NMF) that performs speech derever-beration in noise-free environments to noisy scenarios. New estimators for jointly obtaining the RIR and exemplar weights for the NMD and NMF-based formulations are presented. The proposed techniques are evaluated on the noise-free and noisy reverberant speech in the CHiME-2 WSJ0 database and are shown to yield better speech enhancement in terms of signal-to-distortion ratio (SDR), perceptual evaluation of speech quality (PESQ) and cepstral distance (CD) measures.
Deepak Baby, Hugo Van hamme
ICASSP2
2016 Language model adaptation for ASR of spoken translations using phrase-based translation models and named entity models
abstract
Language model adaptation based on Machine Translation (MT) is a recently proposed approach to improve the Automatic Speech Recognition (ASR) of spoken translations that does not suffer from a common problem in approaches based on rescoring i.e. errors made during recognition cannot be recovered by the MT system. In previous work we presented an efficient implementation for MT-based language model adaptation using a word-based translation model. By omitting renormalization and employing weighted updates, the implementation exhibited virtually no adaptation overhead, enabling its use in a real-time setting. In this paper we investigate whether we can improve recognition accuracy without sacrificing the achieved efficiency. More precisely, we investigate the effect of both state-of-the-art phrase-based translation models and named entity probability estimation. We report relative WER reductions of 6.2% over a word-based LM adaptation technique and 25.3% over an unadapted 3-gram baseline on an English-to-Dutch dataset.
Joris Pelemans, Tom Vanallemeersch, Kris Demuynck, Lyan Verwimp, Hugo Van hamme, Patrick Wambacq
ICASSP5
2016 Data selection for noise robust exemplar matching
abstract
Exemplar-based acoustic modeling is based on labeled training segments that are compared with the unseen test utterances with respect to a dissimilarity measure. Using a larger number of accurately labeled exemplars provides better generalization thus improved recognition performance which comes with increased computation and memory requirements. We have recently developed a noise robust exemplar matching-based automatic speech recognition system which uses a large number of undercomplete dictionaries containing speech exemplars of the same length and label to recognize noisy speech. In this work, we investigate several speech exemplar selection techniques proposed for undercomplete speech dictionaries to find a trade-off between the recognition accuracy and the acoustic model size in terms of the amount of speech exemplars used for recognition. The exemplar selection criterion has be to chosen carefully as the amount of redundancy in these dictionaries is very limited compared to overcomplete dictionaries containing plenty of exemplars. The recognition accuracies obtained on the small vocabulary track of the 2nd CHiME Challenge and the AURORA-2 database using the complete and pruned dictionaries are compared to investigate the performance of each selection criterion.
Emre Yilmaz 0001, Jort F. Gemmeke, Hugo Van hamme
ICASSP3
2016 Active speaker detection with audio-visual co-training
abstract
In this work, we show how to co-train a classifier for active speaker detection using audio-visual data. First, audio Voice Activity Detection (VAD) is used to train a personalized video-based active speaker classifier in a weakly supervised fashion. The video classifier is in turn used to train a voice model for each person. The individual voice models are then used to detect active speakers. There is no manual supervision - audio weakly supervises video classification, and the co-training loop is completed by using the trained video classifier to supervise the training of a personalized audio voice classifier.
Punarjay Chakravarty, Jeroen Zegers, Tinne Tuytelaars, Hugo Van hamme
ICMI4
2016 Joint Sound Source Separation and Speaker Recognition
abstract
Non-negative Matrix Factorization (NMF) has already been applied to learn speaker characterizations from single or non-simultaneous speech for speaker recognition applications. It is also known for its good performance in (blind) source separation for simultaneous speech. This paper explains how NMF can be used to jointly solve the two problems in a multichannel speaker recognizer for simultaneous speech. It is shown how state-of-the-art multichannel NMF for blind source separation can be easily extended to incorporate speaker recognition. Experiments on the CHiME corpus show that this method outperforms the sequential approach of first applying source separation, followed by speaker recognition that uses state-of-the-art i-vector techniques.
Jeroen Zegers, Hugo Van hamme
INTERSPEECH2
2016 SCALE: A Scalable Language Engineering Toolkit
Joris Pelemans, Lyan Verwimp, Kris Demuynck, Hugo Van hamme, Patrick Wambacq
LREC4
2016 Incrementally learn the relevance of words in a dictionary for spoken language acquisition
abstract
This paper discusses a spoken language acquisition system for a command-and-control interface. The proposed system learns a set of words through coupled commands and demonstrations. The user can teach the system a new command by demonstrating the uttered command through an alternative interface. With these coupled commands and demonstrations, the system can learn the acoustic representations of the used words coupled with the meaning or semantics. In previous work the focus was mainly on a batch learning scheme to train the model. All the commands and demonstrations had to be stored and the model had to be retrained from scratch every time a new demonstration was given by the user. This work presents a Bayesian learning scheme where the dictionary of learned words can be updated when new data is presented. The dictionary can automatically expand to add new words or shrink to forget old words. The proposed system is tested on a language acquisition task where the user suddenly starts using new words. The results show that the proposed system can learn the new words quicker than a baseline where the size of the dictionary cannot be adjusted.
Vincent Renkens, Vikrant Tomar, Hugo Van hamme
SLT3
2016 Under-determined reverberant audio source separation using Bayesian Non-negative Matrix Factorization
Sayeh Mirzaei, Hugo Van hamme, Yaser Norouzi
Speech Commun.2
2016 Noise robust exemplar matching with alpha-beta divergence
Emre Yilmaz 0001, Jort F. Gemmeke, Hugo Van hamme
Speech Commun.3
2016 Unseen Noise Estimation Using Separable Deep Auto Encoder for Speech Enhancement
abstract
Unseen noise estimation is a key yet challenging step to make a speech enhancement algorithm work in adverse environments. At worst, the only prior knowledge we know about the encountered noise is that it is different from the involved speech. Therefore, by subtracting the components which cannot be adequately represented by a well defined speech model, the noises can be estimated and removed. Given the good performance of deep learning in signal representation, a deep auto encoder (DAE) is employed in this work for accurately modeling the clean speech spectrum. In the subsequent stage of speech enhancement, an extra DAE is introduced to represent the residual part obtained by subtracting the estimated clean speech spectrum (by using the pre-trained DAE) from the noisy speech spectrum. By adjusting the estimated clean speech spectrum and the unknown parameters of the noise DAE, one can reach a stationary point to minimize the total reconstruction error of the noisy speech spectrum. The enhanced speech signal is thus obtained by transforming the estimated clean speech spectrum back into time domain. The above proposed technique is called separable deep auto encoder (SDAE). Given the under-determined nature of the above optimization problem, the clean speech reconstruction is confined in the convex hull spanned by a pre-trained speech dictionary. New learning algorithms are investigated to respect the non-negativity of the parameters in the SDAE. Experimental results on TIMIT with 20 noise types at various noise levels demonstrate the superiority of the proposed method over the conventional baselines.
Meng Sun 0001, Xiongwei Zhang, Hugo Van hamme, Thomas Fang Zheng
IEEE ACM Trans. Audio Speech Lang. Process.3
2015 Exemplar-based speech enhancement for deep neural network based automatic speech recognition
abstract
Deep neural network (DNN) based acoustic modelling has been successfully used for a variety of automatic speech recognition (ASR) tasks, thanks to its ability to learn higher-level information using multiple hidden layers. This paper investigates the recently proposed exemplar-based speech enhancement technique using coupled dictionaries as a pre-processing stage for DNN-based systems. In this setting, the noisy speech is decomposed as a weighted sum of atoms in an input dictionary containing exemplars sampled from a domain of choice, and the resulting weights are applied to a coupled output dictionary containing exemplars sampled in the short-time Fourier transform (STFT) domain to directly obtain the speech and noise estimates for speech enhancement. In this work, settings using input dictionary of exemplars sampled from the STFT, Mel-integrated magnitude STFT and modulation envelope spectra are evaluated. Experiments performed on the AURORA-4 database revealed that these pre-processing stages can improve the performance of the DNN-HMM-based ASR systems with both clean and multi-condition training.
Deepak Baby, Jort F. Gemmeke, Tuomas Virtanen, Hugo Van hamme
ICASSP4
2015 Improving n-gram probability estimates by compound-head clustering
abstract
Compounding is one of the most productive word formation processes in many languages and is therefore a main source of data sparsity in language modeling. Many solutions have been suggested to model compound words, most of which break the compound into its constituents and train a new model with them. In earlier work, we argued that this approach is suboptimal and we presented a novel technique that clusters new, domain-specific compound words together with their semantic heads. The clusters were then used to build a class-based n-gram model that enabled a reliable estimation of n-gram probabilities, without the need for additional training data. In this paper, we investigate how this “semantic head mapping” can best be made an integral part of the language modeling strategy and find that, with some adaptations, our technique is capable of producing more accurate compound probability estimates than a baseline word-based n-gram language model, which lead to a significant word error rate reduction for Dutch read speech.
Joris Pelemans, Kris Demuynck, Hugo Van hamme, Patrick Wambacq
ICASSP3
2015 Who's Speaking?: Audio-Supervised Classification of Active Speakers in Video
abstract
Active speakers have traditionally been identified in video by detecting their moving lips. This paper demonstrates the same using spatio-temporal features that aim to capture other cues: movement of the head, upper body and hands of active speakers. Speaker directional information, obtained using sound source localization from a microphone array is used to supervise the training of these video features.
Punarjay Chakravarty, Sayeh Mirzaei, Tinne Tuytelaars, Hugo Van hamme
ICMI4
2015 Investigating modulation spectrogram features for deep neural network-based automatic speech recognition
abstract
Copyright © 2015 ISCA. Deep neural network (DNN) based acoustic modelling has been shown to yield significant improvements over Gaussian Mixture Models (GMM) for a variety of automatic speech recognition (ASR) tasks. In addition, it is also becoming popular to use rich speech representations, such as full-resolution spectrograms and perceptually motivated features, as input to the DNNs as they are less sensitive to the increase in the input dimensionality. In this work, we evaluate the performance of a DNN trained on the perceptually motivated modulation envelope spectrogram features that model the temporal amplitude modulations within sub-band speech signals. The proposed approach is shown to outperform DNNs trained on a variety of conventional features such as Mel, PLP and STFT features on both TIMIT phone recognition and the AURORA-4 word recognition tasks. It is also shown that the approach outperforms a sophisticated auditory model based on Gabor filter bank features on TIMIT and the channel matched conditions of the AURORA-4 database.
Deepak Baby, Hugo Van hamme
INTERSPEECH2
2015 A multi-channel speech enhancement framework for robust NMF-based speech recognition for speech-impaired users
abstract
In this paper a multi-channel speech enhancement framework for distant speech acquisition in noisy and reverberant environments for Non-negative Matrix Factorization (NMF)-based Automatic Speech Recognition (ASR) is proposed. The system is evaluated for its use in an assistive vocal interface for physically impaired and speech-impaired users. The framework utilises the Spatially Pre-processed Speech Distortion Weighted Multi-channel Wiener Filter (SP-SDW-MWF) in combination with a postfilter to reduce noise and reverberation. Additionally, the estimation uncertainty of the speech enhancement framework is propagated through the Mel-Frequency Cepstrum Coefficients (MFCC) feature extraction to allow for feature compensation in a later stage. Results indicate that a) using a trade-off parameter between noise reduction and speech distortion has a positive effect on the recognition performance with respect to the well-known GSC and MWF and b) the addition of a postfilter and the feature compensation increases performance with respect to several baselines for a non-pathological and pathological speaker.
Gert Dekkers, Toon van Waterschoot, Bart Vanrumste, Bert Van Den Broeck, Jort F. Gemmeke, Hugo Van hamme, Peter Karsmakers
INTERSPEECH6
2015 Efficient language model adaptation for automatic speech recognition of spoken translations
abstract
Copyright © 2015 ISCA. Direct integration of translation model (TM) probabilities into a language model (LM) with the purpose of improving automatic speech recognition (ASR) of spoken translations typically requires a number of complex operations for each sentence. Many if not all of the LM probabilities need to be updated, the model needs to be renormalized and the ASR system needs to load a new, updated LM for each sentence. In computer-aided translation environments the time loss induced by these complex operations seriously reduces the potential of ASR as an efficient input method. In this paper we present a novel LM adaptation technique that drastically reduces the complexity of each of these operations. The technique consists of LM probability updates using exponential weights based on TM probabilities for each sentence and does not enforce probability renormalization. Instead of storing each resulting language model in its entirety, we only store the update weights which also reduces disk storage and loading time during ASR. Experiments on Dutch read speech translated from English show that both disk storage and recognition time drop dramatically compared to a baseline system that employs a more conventional way of updating the LM.
Joris Pelemans, Tom Vanallemeersch, Kris Demuynck, Hugo Van hamme, Patrick Wambacq
INTERSPEECH4
2015 Mutually exclusive grounding for weakly supervised non-negative matrix factorisation
abstract
Copyright © 2015 ISCA. Non-negative Matrix Factorisation (NMF) has been successfully applied for learning the meaning of a small set of vocal commands without any prior knowledge of the language. This kind of learning is useful if flexibility in terms of the acoustic and language model is required, for example in assistive technologies for dysarthric speakers because they do not comply with common models. Vocal commands are grounded through the addition of semantic labels that represent the action corresponding to the command. The Kullback Leibler Divergence (KLD) is used to evaluate the acoustic model. The KLD is optimal for Poisson distributed data making it an appropriate metric for the acoustic features because they are a count of acoustic events. The semantic labels are however activations, so a multinomial likelihood function seems more appropriate because they are mutually exclusive. In this paper a cost function to evaluate the semantic model based on the multinomial likelihood function is proposed that aims to better suit its distribution. To minimise the proposed cost function a new set of update rules and a new normalisation scheme are proposed.
Vincent Renkens, Hugo Van hamme
INTERSPEECH2
2015 Noise robust exemplar matching for speech enhancement: applications to automatic speech recognition
abstract
Copyright © 2015 ISCA. We present a novel automatic speech recognition (ASR) scheme which uses the recently proposed noise robust exemplar matching framework for speech enhancement in the front-end. The proposed system employs a GMM-HMM back-end to recognize the enhanced speech signals unlike the prior work focusing on template matching only. Speech enhancement is achieved using multiple dictionaries containing speech exemplars representing a single speech unit and several noise exemplars of the same length. These combined dictionaries are used to approximate the noisy segments and the speech component is obtained as a linear combination of the speech exemplars in the combined dictionaries yielding the minimum total reconstruction error. The performance of the proposed system is evaluated on the small vocabulary track of the 2nd CHiME Challenge and the AURORA-2 database and the results have shown the effectiveness of the proposed approach in improving the noise robustness of a conventional ASR system.
Emre Yilmaz 0001, Deepak Baby, Hugo Van hamme
INTERSPEECH3
2015 Two-stage blind audio source counting and separation of stereo instantaneous mixtures using Bayesian tensor factorisation
abstract
In this paper, the authors address the tasks of audio source counting and separation for two‐channel instantaneous mixtures. This goal is achieved in two steps. First, a novel scheme is proposed for estimating the number of sources and the corresponding channel intensity difference (CID) values. For this purpose, an angular spectrum is evaluated as a function of the ratio of the magnitude spectrogram of the two channels and the peak locations of that spectrum are obtained. In the second stage, a new approach is developed for extracting the individual source signals exploiting a Bayesian non‐parametric modelling. The mean field variational Bayesian approach is applied for inferring the unknown parameters. Classification is then performed on the inferred active CID values to obtain the individual source magnitude spectrograms. This way, the number of spectral components used for modelling each source is found automatically from the data. The Bayesian approach is compared with the standard Kullback–Leibler non‐negative tensor factorisation method to illustrate the effectiveness of Bayesian modelling. The performance of the source separation is measured by obtaining the existing metrics for multichannel blind source separation evaluation. The experiments are performed on instantaneous mixtures from the dev2 database.
Sayeh Mirzaei, Yaser Norouzi, Hugo Van hamme
IET Signal Process.3
2015 A stable approach for model order selection in nonnegative matrix factorization
Meng Sun 0001, Xiongwei Zhang, Hugo Van hamme
Pattern Recognit. Lett.3
2015 Blind audio source counting and separation of anechoic mixtures using the multichannel complex NMF framework
abstract
In this paper, we address the tasks of audio source counting and separation for a stereo anechoic mixture of audio signals. This will be achieved in two stages. In the first stage, a novel approach is introduced for estimating the number of sources as well as the channel mixing coefficients. For this purpose, a 2-D spectrum is evaluated against both the phase and amplitude differences of the two channels. Hence, obtaining the peak locations of the spectrum yields the number of the sources and the corresponding channel coefficients. In the second stage, an extension of a single channel complex matrix factorization method to multichannel is developed to extract the individual source signals. We find primary estimates of the sources via binary masking and then apply the complex factorization to the complex spectrogram of each source. The obtained factors are then utilized as initial values in the complex multichannel factorization model. We also suggest a method for estimating the number of required components for modeling each source. The separation performance improvement over the conventional methods is investigated by calculating BSS evaluation metrics. The comparison is also carried out in terms of source counting and localization with the recently proposed DeMIX-Anechoic method.
Sayeh Mirzaei, Hugo Van hamme, Yaser Norouzi
Signal Process.2
2015 Coupled Dictionaries for Exemplar-Based Speech Enhancement and Automatic Speech Recognition
abstract
Exemplar-based speech enhancement systems work by decomposing the noisy speech as a weighted sum of speech and noise exemplars stored in a dictionary and use the resulting speech and noise estimates to obtain a time-varying filter in the full-resolution frequency domain to enhance the noisy speech. To obtain the decomposition, exemplars sampled in lower dimensional spaces are preferred over the full-resolution frequency domain for their reduced computational complexity and the ability to better generalize to unseen cases. But the resulting filter may be sub-optimal as the mapping of the obtained speech and noise estimates to the full-resolution frequency domain yields a low-rank approximation. This paper proposes an efficient way to directly compute the full-resolution frequency estimates of speech and noise using coupled dictionaries: an input dictionary containing atoms from the desired exemplar space to obtain the decomposition and a coupled output dictionary containing exemplars from the full-resolution frequency domain. We also introduce modulation spectrogram features for the exemplar-based tasks using this approach. The proposed system was evaluated for various choices of input exemplars and yielded improved speech enhancement performances on the AURORA-2 and AURORA-4 databases. We further show that the proposed approach also results in improved word error rates (WERs) for the speech recognition tasks using HMM-GMM and deep-neural network (DNN) based systems.
Deepak Baby, Tuomas Virtanen, Jort F. Gemmeke, Hugo Van hamme
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Coupled dictionary training for exemplar-based speech enhancement
abstract
In exemplar-based speech enhancement systems, lower dimensional features are preferred over the full-scale DFT features for their reduced computational complexity and the ability to better generalize for the unseen cases. But in order to obtain the Wiener-like filter for noisy DFT enhancement, the speech and noise estimates obtained in the feature space need to be mapped to the DFT space, which yield a low-rank approximation of the estimates resulting in a sub-optimal filter. This paper proposes a novel method using coupled dictionaries where the exemplars for the required feature space and the DFT space are jointly extracted and the estimates are directly obtained in the DFT space following the decomposition in the chosen feature space. Simulation experiments revealed that the proposed approach, where the activations of exemplars calculated using the Mel resolution are directly used to obtain the Wiener filter in the DFT space, results in improved signal-to-distortion ratio (SDR) when compared to the system without coupled dictionaries. To further motivate the use of coupled dictionaries, the paper also investigates the use of modulation envelope features for the exemplar-based speech enhancement.
Deepak Baby, Tuomas Virtanen, Tom Barker, Hugo Van hamme
ICASSP4
2014 Coping with language data sparsity: Semantic head mapping of compound words
abstract
In this paper we present a novel clustering technique for compound words. By mapping compounds onto their semantic heads, the technique is able to estimate n-gram probabilities for unseen compounds. We argue that compounds are well represented by their heads which allows the clustering of rare words and reduces the risk of over-generalization. The semantic heads are obtained by a two-step process which consists of constituent generation and best head selection based on corpus statistics. Experiments on Dutch read speech show that our technique is capable of correctly identifying compounds and their semantic heads with a precision of 80.25% and a recall of 85.97%. A class-based language model with compound-head clusters achieves a significant reduction in both perplexity and WER.
Joris Pelemans, Kris Demuynck, Hugo Van hamme, Patrick Wambacq
ICASSP3
2014 Active-set newton algorithm for non-negative sparse coding of audio
abstract
We propose a new algorithm to efficiently obtain non-negative sparse representations for audio. The spectrum of an audio signal is represented as a sparse linear combination of atoms taken from an overcomplete dictionary. The algorithm is based on minimizing the generalized Kullback-Leibler divergence between an observed magnitude spectrum and a non-negative linear combination of atoms, plus an ℓ1regularization term. The proposed method consists of an active-set method that iteratively updates a set of active atoms that have non-zero weights, using a Newton step where the weights of the active atoms are updated. The proposed method was evaluated using mixtures of two speakers, and it was shown to yield more than 10 times faster convergence in comparison to an established algorithm based on multiplicative update rules. Moreover, the ℓ1regularization was found to decrease the computation time and to improve the source separation performance.
Tuomas Virtanen, Bhiksha Raj, Jort F. Gemmeke, Hugo Van hamme
ICASSP4
2014 Noise-robust speech recognition with exemplar-based sparse representations using Alpha-Beta divergence
abstract
In this paper, we investigate the performance of a noise-robust sparse representations (SR)-based recognizer using the Alpha-Beta (AB)-divergence to compare the noisy speech segments and exemplars. The baseline recognizer, which approximates noisy speech segments as a linear combination of speech and noise exemplars of variable length, uses the generalized Kullback-Leibler divergence to quantify the approximation quality. Incorporating a reconstruction error-based back-end, the recognition performance highly depends on the congruence of the divergence measure and used speech features. Having two tuning parameters, namely α and β, the AB-divergence provides improved robustness against background noise and outliers. These parameters can be adjusted for better performance depending on the distribution of speech and noise exemplars in the high-dimensional feature space. Moreover, various well-known distance/divergence measures such as the Euclidean distance, generalized Kullback-Leibler divergence, Itakura-Saito divergence and Hellinger distance are special cases of the AB-divergence for different (α, β) values. The goal of this work is to investigate the optimal divergence for mel-scaled magnitude spectral features by performing recognition experiments at several SNR levels using different (α, β) pairs. The results demonstrate the effectiveness of the AB-divergence compared to the generalized Kullback-Leibler divergence especially at the lower SNR levels.
Emre Yilmaz 0001, Jort F. Gemmeke, Hugo Van hamme
ICASSP3
2014 Modelling primitive streaming of simple tone sequences through factorisation of modulation pattern tensors
abstract
Copyright © 2014 ISCA. We present a novel method for determining how the perceptual organisation of simple alternating tone sequences is likely to occur in human listeners. By training a tensor model representation using features which incorporate both low-frequency modulation rate and phase, a set of components is learned. Test patterns are modelled using these learned components, and the sum of component activations is used to predict either an 'integrated' or 'segregated' auditory stream percept. We find that for the basic streaming paradigm tested, our proposed model and method is able to correctly predict either segregation or integration in the majority of cases.
Tom Barker, Hugo Van hamme, Tuomas Virtanen
INTERSPEECH2
2014 Blind speech source localization, counting and separation for 2-channel convolutive mixtures in a reverberant environment
abstract
Copyright © 2014 ISCA. In this paper, the tasks of speech source localization, source counting and source separation are addressed for an unknown number of sources in a stereo recording scenario. In the first stage, the angles of arrival of individual source signals are estimated through a peak finding scheme applied to the angular spectrum which has been derived using non-linear GCC-PHAT. Then, based on the known channel mixture coefficients, we propose an approach for separating the sources based on Maximum Likelihood (ML) estimation. The predominant source in each time-frequency bin is identified through ML assuming a diffuse noise model. The separation performance is improved over a binary time-frequency masking method. The performance is measured by obtaining the existing metrics for blind source separation evaluation. The experiments are performed on synthetic speech mixtures in both anechoic and reverberant environments.
Sayeh Mirzaei, Hugo Van hamme, Yaser Norouzi
INTERSPEECH2
2014 An evaluation of unsupervised acoustic model training for a dysarthric speech interface
abstract
Copyright © 2014 ISCA. In this paper, we investigate unsupervised acoustic model training approaches for dysarthric-speech recognition. These models are first, frame-based Gaussian posteriorgrams, obtained from Vector Quantization (VQ), second, so-called Acoustic Unit Descriptors (AUDs), which are hidden Markov models of phone-like units, that are trained in an unsupervised fashion, and, third, posteriorgrams computed on the AUDs. Experiments were carried out on a database collected from a home automation task and containing nine speakers, of which seven are considered to utter dysarthric speech. All unsupervised modeling approaches delivered significantly better recognition rates than a speaker-independent phoneme recognition baseline, showing the suitability of unsupervised acoustic model training for dysarthric speech. While the AUD models led to the most compact representation of an utterance for the subsequent semantic inference stage, posteriorgram-based representations resulted in higher recognition rates, with the Gaussian posteriorgram achieving the highest slot filling F-score of 97.02%.
Oliver Walter, Vladimir Despotovic, Reinhold Häb-Umbach, Jort F. Gemmeke, Bart Ons, Hugo Van hamme
INTERSPEECH6
2014 Automatic assessment of children's reading with the FLaVoR decoding using a phone confusion model
abstract
Copyright © 2014 ISCA. Reading skills of children can be improved with the help of automatic reading tutors (ART), i.e. interactive software with an appealing interface which supports and challenges the child in the reading task, provides instantaneous feedback and automatically assesses its reading skills. For this purpose, ARTs benefit from automatic speech recognition technology for tracking the child's responses and detecting reading miscues (errors). In previous work, a novel speech recognition architecture has been proposed which adopts a two-layered structure: first a phone recognizer uses task-independent acoustic and language models to generate a phone lattice which is then decoded using a lexicon of expected words and task-dependent finite state grammars. This approach has shown significant improvements in reading miscue detection. In this paper, we extend this technique by employing a more flexible decoding scheme that allows substitution, deletion and insertion of phones. Specifically, the phone lattice generated in the first layer is extended based on a phone confusion matrix that models the typical phone confusions in a language. The proposed system has provided improved miscue detection on the CHOREC database compared to a baseline system without a phone confusion model.
Emre Yilmaz 0001, Joris Pelemans, Hugo Van hamme
INTERSPEECH3
2014 Speech Recognition Web Services for Dutch
Joris Pelemans, Kris Demuynck, Hugo Van hamme, Patrick Wambacq
LREC3
2014 Learning Like a Toddler: Watching Television Series to Learn Vocabulary from Images and Audio
abstract
This paper presents the initial findings of our efforts to build an unsupervised multimodal vocabulary learning scheme in a realistic scenario. For this purpose, a new multimodal dataset, called Musti3D, has been created. The Musti3D database contains episodes from an animation series for toddlers. Annotated with audiovisual information, this database is used for the investigation of a non-negative matrix factorization (NMF)-based audiovisual learning technique. The performance of the technique, i.e. correctly matching the audio and visual representations of the objects, has been evaluated by gradually reducing the level of supervision starting from the ground truth transcriptions. Moreover, we have performed experiments using different visual representations and time spans for combining the audiovisual information. The preliminary results show the feasibility of the proposed audiovisual learning framework.
Emre Yilmaz 0001, Konstantinos Rematas, Tinne Tuytelaars, Hugo Van hamme
ACM Multimedia4
2014 Exemplar-based noise robust automatic speech recognition using modulation spectrogram features
abstract
We propose a novel exemplar-based feature enhancement method for automatic speech recognition which uses coupled dictionaries: an input dictionary containing atoms sampled in the modulation (envelope) spectrogram domain and an output dictionary with atoms in the Mel or full-resolution frequency domain. The input modulation representation is chosen for its separation properties of speech and noise and for its relation with human auditory processing. The output representation is one which can be processed by the ASR back-end. The proposed method was investigated on the AURORA-2 and AURORA-4 databases and improved word error rates (WER) were obtained when compared to the system which uses Mel features in the input exemplars. The paper also proposes a hybrid system which combines the baseline and the proposed algorithm on the AURORA-2 database which in turn also yielded improvement over both the algorithms.
Deepak Baby, Tuomas Virtanen, Jort F. Gemmeke, Tom Barker, Hugo Van hamme
SLT5
2014 Dysarthric vocal interfaces with minimal training data
abstract
Over the past decade, several speech-based electronic assistive technologies (EATs) have been developed that target users with dysarthric speech. These EATs include vocal command & control systems, but also voice-input voice-output communication aids (VIVOCAs). In these systems, the vocal interfaces are based on automatic speech recognition systems (ASR), but this approach requires much training data and detailed annotation. In this work we evaluate an alternative approach, which works by mining utterance-based representations of speech for recurrent acoustic patterns, with the goal of achieving usable recognition accuracies with less speaker-specific training data. Comparisons with a conventional ASR system on dysarthric speech databases show that the proposed approach offers a substantial reduction in the amount of training data needed to achieve the same recognition accuracies.
Jort F. Gemmeke, Siddharth Sehgal, Stuart P. Cunningham, Hugo Van hamme
SLT4
2014 Acquisition of ordinal words using weakly supervised NMF
abstract
This paper issues in the design of a vocal interface for a robot that can learn to understand spoken utterances through demonstration. Weakly supervised non-negative matrix factorization (NMF) is used as a machine learning algorithm where acoustic data are augmented with semantic labels representing the meaning of the command. Many parameters that the robot needs in order to execute the commands have an ordinal structure. Constrained subspace NMF (CSNMF) is proposed as an extension to NMF that aims to better deal with ordinal data and thus increase the learning rate of the grounding information with an ordinal structure. Furthermore automatic relevance determination is used to deal with model order selection. The use of CSNMF yields a significant improvement in the learning rate and accuracy when recognising ordinal parameters.
Vincent Renkens, Steven Janssens, Bart Ons, Jort F. Gemmeke, Hugo Van hamme
SLT5
2014 Fast vocabulary acquisition in an NMF-based self-learning vocal user interface
abstract
In command-and-control applications, a vocal user interface (VUI) is useful for handsfree control of various devices, especially for people with a physical disability. The spoken utterances are usually restricted to a predefined list of phrases or to a restricted grammar, and the acoustic models work well for normal speech. While some state-of-the-art methods allow for user adaptation of the predefined acoustic models and lexicons, we pursue a fully adaptive VUI by learning both vocabulary and acoustics directly from interaction examples. A learning curve usually has a steep rise in the beginning and an asymptotic ceiling at the end. To limit tutoring time and to guarantee good performance in the long run, the word learning rate of the VUI should be fast and the learning curve should level off at a high accuracy. In order to deal with these performance indicators, we propose a multi-level VUI architecture and we investigate the effectiveness of alternative processing schemes. In the low-level layer, we explore the use of MIDA features (Mutual Information Discrimination Analysis) against conventional MFCC features. In the mid-level layer, we enhance the acoustic representation by means of phone posteriorgrams and clustering procedures. In the high-level layer, we use the NMF (Non-negative Matrix Factorization) procedure which has been demonstrated to be an effective approach for word learning. We evaluate and discuss the performance and the feasibility of our approach in a realistic experimental setting of the VUI-user learning context.
Bart Ons, Jort F. Gemmeke, Hugo Van hamme
Comput. Speech Lang.3
2014 Speaker age estimation using i-vectors
Mohamad Hasan Bahari, Mitchell McLaren, Hugo Van hamme, David A. van Leeuwen
Eng. Appl. Artif. Intell.3
2014 Non-Negative Factor Analysis of Gaussian Mixture Model Weight Adaptation for Language and Dialect Recognition
abstract
Recent studies show that Gaussian mixture model (GMM) weights carry less, yet complimentary, information to GMM means for language and dialect recognition. However, state-of-the-art language recognition systems usually do not use this information. In this research, a non-negative factor analysis (NFA) approach is developed for GMM weight decomposition and adaptation. This modeling, which is conceptually simple and computationally inexpensive, suggests a new low-dimensional utterance representation method using a factor analysis similar to that of the i-vector framework. The obtained subspace vectors are then applied in conjunction with i-vectors to the language/dialect recognition problem. The suggested approach is evaluated on the NIST 2011 and RATS language recognition evaluation (LRE) corpora and on the QCRI Arabic dialect recognition evaluation (DRE) corpus. The assessment results show that the proposed adaptation method yields more accurate recognition results compared to three conventional weight adaptation approaches, namely maximum likelihood re-estimation, non-negative matrix factorization, and a subspace multinomial model. Experimental results also show that the intermediate-level fusion of i-vectors and NFA subspace vectors improves the performance of the state-of-the-art i-vector framework especially for the case of short utterances.
Mohamad Hasan Bahari, Najim Dehak, Hugo Van hamme, Lukás Burget, Ahmed Ali 0002, James R. Glass
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 Noise Robust Exemplar Matching Using Sparse Representations of Speech
abstract
Performing automatic speech recognition using exemplars (templates) holds the promise to provide a better duration and coarticulation modeling compared to conventional approaches such as hidden Markov models (HMMs). Exemplars are spectrographic representations of speech segments extracted from the training data, each associated with a speech unit, e.g. phones, syllables, half-words or words, and preserve the complete spectro-temporal content of the speech. Conventional exemplar-matching approaches to automatic speech recognition systems, such as those based on dynamic time warping, have typically focused on evaluation in clean conditions. In this paper, we propose a novel noise robust exemplar matching framework for automatic speech recognition. This recognizer approximates noisy speech segments as a weighted sum of speech and noise exemplars and performs recognition by comparing the reconstruction errors of different classes with respect to a divergence measure. We evaluate the system performance in keyword recognition on the small vocabulary track of the 2nd CHiME Challenge and connected digit recognition on the AURORA-2 database. The results show that the proposed system achieves comparable results with state-of-the-art noise robust recognition systems.
Emre Yilmaz 0001, Jort F. Gemmeke, Hugo Van hamme
IEEE ACM Trans. Audio Speech Lang. Process.3
2013 NMF-based keyword learning from scarce data
abstract
This research is situated in a project aimed at the development of a vocal user interface (VUI) that learns to understand its users specifically persons with a speech impairment. The vocal interface adapts to the speech of the user by learning the vocabulary from interaction examples. Word learning is implemented through weakly supervised non-negative matrix factorization (NMF). The goal of this study is to investigate how we can improve word learning when the number of interaction examples is low. We demonstrate two approaches to train NMF models on scarce data: 1) training word models using smoothed training data, and 2) training word models that strictly correspond to the grounding information derived from a few interaction examples. We found that both approaches can substantially improve word learning from scarce training data.
Bart Ons, Jort F. Gemmeke, Hugo Van hamme
ASRU3
2013 Accent recognition using i-vector, Gaussian Mean Supervector and Gaussian posterior probability supervector for spontaneous telephone speech
abstract
In this paper, three utterance modelling approaches, namely Gaussian Mean Supervector (GMS), i-vector and Gaussian Posterior Probability Supervector (GPPS), are applied to the accent recognition problem. For each utterance modeling method, three different classifiers, namely the Support Vector Machine (SVM), the Naive Bayesian Classifier (NBC) and the Sparse Representation Classifier (SRC), are employed to find out suitable matches between the utterance modelling schemes and the classifiers. The evaluation database is formed by using English utterances of speakers whose native languages are Russian, Hindi, American English, Thai, Vietnamese and Cantonese. These utterances are drawn from the National Institute of Standards and Technology (NIST) 2008 Speaker Recognition Evaluation (SRE) database. The study results show that GPPS and i-vector are more effective than GMS in this accent recognition task. It is also concluded that among the employed classifiers, the best matches for i-vector and GPPS are SVM and SRC, respectively.
Mohamad Hasan Bahari, Rahim Saeidi, Hugo Van hamme, David A. van Leeuwen
ICASSP3
2013 Embedding time warping in exemplar-based sparse representations of speech
abstract
This paper describes a new sparse representation model for speech that allows time warping as an extension to a recently proposed sparse representations-based speech recognition system. This recognition system uses exemplars to model the acoustics which are labeled speech occurrences of different length extracted from the training data. Exemplars are organized in multiple dictionaries on the basis of their class and length. Input speech segments are approximated as a sparse linear combination of the exemplars using these dictionaries and a reconstruction error-based decoding is adopted in order to find the best matching class sequence. With the current sparse representation model using a dictionary and a weight vector to approximate an input speech segment, it is not possible to compare input speech segments with exemplars of different lengths. The goal of this work is to introduce a novel sparse representation model which allows time warping using a third matrix which linearly combines consecutive frames in order to shrink or expand the approximation. Preliminary results have shown the feasibility of the proposed sparse representation model.
Emre Yilmaz 0001, Jort F. Gemmeke, Hugo Van hamme
ICASSP3
2013 A diagonalized newton algorithm for non-negative sparse coding
abstract
Signal models where non-negative vector data are represented by a sparse linear combination of non-negative basis vectors have attracted much attention in problems including image classification, document topic modeling, sound source segregation and robust speech recognition. In this paper, an iterative algorithm based on Newton updates to minimize the Kullback-Leibler divergence between data and model is proposed. It finds the sparse activation weights of the basis vectors more efficiently than the expectation-maximization (EM) algorithm. To avoid the computational burden of a matrix inversion, a diagonal approximation is made and therefore the algorithm is called diagonal Newton Algorithm (DNA). It is several times faster than EM, especially for undercomplete problems. But DNA also performs surprisingly well on overcomplete problems.
Hugo Van hamme
ICASSP1
2013 Self-taught assistive vocal interfaces: an overview of the ALADIN project
abstract
This paper gives an overview of research within the ALADIN project, which aims to develop an assistive vocal interface for people with a physical impairment. In contrast to existing ap-proaches, the vocal interface is trained by the end-user himself, which means it can be used with any vocabulary and grammar, and that it is maximally adapted to the — possibly dysarthric — speech of the user. This paper describes the overall learn-ing framework, the user-centred design and evaluation aspects, database collection and approaches taken to combat problems such as noise and erroneous input. Index Terms: vocal user interface, user-centred design, self-taught learning, speech database, dysarthric speech
Jort F. Gemmeke, Bart Ons, Netsanet M. Tessema, Hugo Van hamme, Janneke van de Loo, Guy De Pauw, Walter Daelemans, Jonathan Huyghe, Jan Derboven, Lode Vuegen, Bert Van Den Broeck, Peter Karsmakers, Bart Vanrumste
INTERSPEECH4
2013 Model order estimation using Bayesian NMF for discovering phone patterns in spoken utterances
abstract
In earlier work, we have shown that vocabulary discovery from spoken utterances and subsequent recognition of the acquired vocabulary can be achieved through Non-negative Matrix Factorization (NMF). An open issue for this task is to determine automatically how many different word representations should be included in the model. In this paper, Bayesian NMF is applied to estimate the model order. The per-utterance word activations are given a gamma prior while the word models are assumed deterministic. Two Bayesian approaches are applied for obtaining optimal parameter values. First, the penalized joint log-likelihood of the parameters is considered as the objective function. Then, maximal marginal likelihood estimator (MMLE) is implemented which obtains the word models maximizing the likelihood after integration over the activations. The variational Bayesian algorithm, which maximizes a lower bound of the marginal log-likelihood, is applied to this optimization problem. The number of required latent components or basis vectors (model order) is estimated by evaluating likelihood metrics. The inferred model order is validated by observing error criteria on a test set. Experiments on synthetic data as well as real speech show that MMLE is more effective for the purpose of model order selection. Copyright © 2013 ISCA.
Sayeh Mirzaei, Hugo Van hamme, Yaser Norouzi
INTERSPEECH2
2013 Joint training of non-negative Tucker decomposition and discrete density hidden Markov models
Meng Sun 0001, Hugo Van hamme
Comput. Speech Lang.2
2013 Rapid speaker adaptation in latent speaker space with non-negative matrix factorization
Xueru Zhang, Kris Demuynck, Hugo Van hamme
Speech Commun.3
2012 Weakly supervised keyword learning using sparse representations of speech
abstract
When applied to speech, Non-negative Matrix Factorization is capable of learning a small vocabulary of words, foregoing any prior linguistic knowledge. This makes it adequate for small-scale speech applications where flexibility is of the utmost importance, e.g. assistive technology for the speech impaired. However, its performance depends on the way its inputs are represented. We propose the use of exemplar-based sparse representations of speech, and explore the influence of some of these representation's basic parameters, such as the total number of exemplars considered and the sparseness imposed on them. We show that the resulting learning performance compares favorably with those of previously proposed approaches.
Joris Driesen, Jort F. Gemmeke, Hugo Van hamme
ICASSP3
2012 Fast word acquisition in an NMF-based learning framework
abstract
A speech recognition system that automatically learns word models for a small vocabulary from examples of its usage, without using prior linguistic information, can be of great use in cognitive robotics, human-machine interfaces, and assistive devices. In the latter case, the user's speech capabilities may also be affected. In this paper, we consider a NMF-based learning framework capable of doing this, and experimentally show that its learning rate crucially depends on how the speech data is represented. Higher-level units of speech, which hide some of the complex variability of the acoustics, are found to yield faster learning rates.
Joris Driesen, Hugo Van hamme
ICASSP2
2012 Tri-factorization learning of sub-word units with application to vocabulary acquisition
abstract
In prior work, we proposed a method for vocabulary acquisition based on a co-occurrence model and non-negative matrix factorization. The vocabulary is described in terms of co-occurrence statistics of frame-level acoustic descriptions and suffers from poor scalability to larger vocabularies. Much like whole-word HMM models, there is no reuse of a sub-word units such as phone models. In this paper, we apply the co-occurrence framework to learn a set of sub-word units unsupervisedly using a matrix tri-factorization and propose a method for computing their posteriorgram and finally show vocabulary acquisition from the posteriorgram. The method outperforms our prior work in that it can learn from a smaller set of labeled data and shows a better recognition accuracy.
Meng Sun 0001, Hugo Van hamme
ICASSP2
2012 Latent variable speaker adaptation of Gaussian mixture weights and means
abstract
We describe a novel fast speaker adaptation algorithm for large vocabulary speech recognition systems, which adapts both the Gaussian means and the mixture weights. Gaussian means are expressed as a linear combination of eigenvoices estimated with principal component analysis. The non-negative Gaussian mixture weights are expressed as a linear combination of a set of latent vectors estimated with non-negative matrix factorization. Experiments on the Wall Street Journal database show that the combination of weight and mean adaptation consistently improves the performance compared to eigenvoice adaptation only. Improvements up to 5.8% relative word error rate reduction were observed with 40 eigenvoices and 40 latent weight vectors. Furthermore, combining weight and mean adaptation outperformed both weight and mean adaptation on itself, even if the latter uses more latent vectors.
Xueru Zhang, Kris Demuynck, Hugo Van hamme
ICASSP3
2012 Age Estimation from Telephone Speech using i-vectors
abstract
Motivated by the success of i-vectors in the field of speaker recognition, this paper proposes a new approach for age estimation from telephone speech patterns based on i-vectors. In this method, each utterance is modeled by its corresponding i-vector. Then, Support Vector Regression (SVR) is applied to estimate the age of speakers. The proposed method is trained and tested on telephone conversations of the National Institute for Standard in Technology (NIST) 2010 and 2008 Speaker Recognition Evaluations databases. Evaluation results show that the proposed method outperforms different conventional methods in speaker age estimation.
Mohamad Hasan Bahari, Mitchell McLaren, Hugo Van hamme, David A. van Leeuwen
INTERSPEECH3
2012 A Self-Learning Assistive Vocal Interface Based on Vocabulary Learning and Grammar Induction
abstract
This paper introduces research within the ALADIN project, which aims to develop an assistive vocal interface for people with a physical impairment. In contrast to existing approaches, the vocal interface is self-learning which means it can be used with any language, dialect, vocabulary and grammar. The paper describes the overall learning framework, and the two components that will provide vocabulary learning and grammar induction. In addition, the paper describes encouraging results of early implementations of these vocabulary and grammar learning components, applied to recorded sessions of a vocally guided card game, patience. Index Terms: language acquisition, word finding, grammar induction, non-negative matrix factorization, concept tagging
Jort F. Gemmeke, Janneke van de Loo, Guy De Pauw, Joris Driesen, Hugo Van hamme, Walter Daelemans
INTERSPEECH5
2012 Advances in noise robust digit recognition using hybrid exemplar-based techniques
abstract
Expressing noisy speech spectra as a linear combination of speech and noise exemplars has been shown to be a powerful tool to achieve noise robust ASR. Such a model has been used both to do feature enhancement (FE) and to directly provide noise robust speech state probabilities using a method called sparse classification (SC). The goal of this work is threefold: First, we integrate various SC advances recently proposed in literature, second, we improve upon the results obtained with FE through retraining and multi-condition training of the acoustic models used in the recognizer and finally, we propos the use of a single hybrid SC-FE system. In our experiments onAURORA-2 we obtain an impressive 3% and 5% average WER on matched and on mismatched noise types, respectively.
Jort F. Gemmeke, Hugo Van hamme
INTERSPEECH2
2012 Robust Tracking for Automatic Reading Tutors
abstract
Reading tutor software uses automatic speech recognition technology to support children in developing their reading skills. In many forms of exercise and evaluation, tracking the reading position is a relevant task or even a prerequisite, e.g. to provide assistance on the pronunciation of a word or to advance the screen to the next page. In this paper, we introduce a new robust tracking algorithm, which measures the similarity between the recognized phones and the phonetic transcription of words displayed on a screen using an efficient dynamic programming algorithm. The criteria for accepting a word reading attempt and thus advancing the cursor can hence be expressed phonetically. In addition, the most likely state of the Hidden Markov Model (HMM) used to decode the speech serves as a fallback for cases of phone matching failure. The new tracker's performance is compared with two other trackers which use either the most likely HMM state or phone matching. The evaluation metrics quantify both the frequency of timely movements and loss of tracking synchronicity. The proposed approach performs significantly better than the others achieving a Timing Accuracy of Tracking of 81.03% compared to 50.63% of the phone matching approach and 32.36% of the state-based approach.
Emre Yilmaz 0001, Dirk Van Compernolle, Hugo Van hamme
INTERSPEECH3
2012 Supervised input space scaling for non-negative matrix factorization
Joris Driesen, Hugo Van hamme
Signal Process.2
2011 An hierarchical exemplar-based sparse model of speech, with an application to ASR
abstract
We propose a hierarchical exemplar-based model of speech, as well as a new algorithm, to efficiently find sparse linear combinations of exemplars in dictionaries containing hundreds of thousands exemplars. We use a variant of hierarchical agglomerative clustering to find a hierarchy connecting all exemplars, so that each exemplar is a parent to two child nodes. We use a modified version of a multiplicative-updates based algorithm to find sparse representations starting from a small active set of exemplars from the dictionary. Namely, on each iteration we replace exemplars that have an increasing weight by their child-nodes. We illustrate the properties of the proposed method by investigating computational effort, accuracy of the eventual sparse representation and speech recognition accuracy on a digit recognition task.
Jort F. Gemmeke, Hugo Van hamme
ASRU2
2011 Progress in example based automatic speech recognition
abstract
In this paper we present a number of improvements that were recently made to the template based speech recognition system developed at ESAT. Combining these improvements resulted in a decrease in word error rate from 9.6% to 8.2% on the Nov92, 20k trigram, Wall Street Journal task. The improvements are along different lines. Apart from the time warping already applied within the DTW, it was found beneficial to apply additional length compensation on the template score. The single best score was replaced by a weighted k-NN average, while maintaining natural successor information as an ensemble cost. The local geometry of the acoustic space is now taken into account by assigning a diagonal covariance matrix to each input frame. Context sensitivity of short templates is increased by taking cross boundary scores into account for sorting the N best templates. Furthermore boundaries on the template segmentations may be relaxed. Finally context dependent word templates are now being used for short words. Several other variants that were not retained in the final system are discussed as well.
Kris Demuynck, Dino Seppi, Hugo Van hamme, Dirk Van Compernolle
ICASSP3
2011 Unsupervised vocabulary discovery using non-negative matrix factorization with graph regularization
abstract
In this paper, we present a model for unsupervised pattern discovery using non-negative matrix factorization (NMF) with graph regularization. Though the regularization can be applied to many applications, we illustrate its effectiveness in a task of vocabulary acquisition in which a spoken utterance is represented by its histogram of the acoustic co-occurrences. The regularization expresses that temporally close co-occurrences should tend to end up in the same learned pattern. A novel algorithm that converges to a local optimum of the regularized cost function is proposed. Our experiments show that the graph regularized NMF model always performs better than the primary NMF model on the task of unsupervised acquisition of a small vocabulary.
Meng Sun 0001, Hugo Van hamme
ICASSP2
2011 Rapid speaker adaptation with speaker adaptive training and non-negative matrix factorization
abstract
In this paper, we describe a novel speaker adaptation algorithm based on Gaussian mixture weight adaptation. A small number of latent speaker vectors are estimated with non-negative matrix factorization (NMF). These base vectors encode the correlations between Gaussian activations as learned from the train data. Expressing the speaker dependent Gaussian mixture weights as a linear combination of a small number of base vectors, reduces the number of parameters that must be estimated from the enrollment data. In order to learn meaningful correlations between Gaussian activations from the train data, the NMF-based weight adaptation was combined with vocal tract length normalization (VTLN) and feature-space maximum likelihood linear regression (fMLLR) based speaker adaptive training based. Evaluation on the 5k closed and 20k open vocabulary Wall Street Journal tasks shows a 4% relative word error rate reduction over the speaker independent recognition system which already incorporates VTLN. The proposed fast adaptation algorithm, using a single enrollment sentence only, results in similar performance as fMLLR adapting on 40 enrollment sentences.
Xueru Zhang, Kris Demuynck, Hugo Van hamme
ICASSP3
2011 Image pattern discovery by using the spatial closeness of visual code words
abstract
A graph regularized non-negative matrix factorization (NMF) model is proposed for image pattern discovery. Each image is represented by its histogram of visual words (i.e. bag-of-words) and the image contents are discovered by the NMF model. The graph regularization preserves the spatial closeness of visual code words in the obtained patterns, thus improving the bag-of-words representation against its main shortcoming: the loss of spatial information. Experiments on a subset of the Caltech256 database show the efficacy of the proposed model.
Meng Sun 0001, Hugo Van hamme
ICIP2
2011 Rapid Speaker Adaptation using Maximum Likelihood Neural Regression
abstract
In this paper, a new method called Maximum Likelihood Neural Regression (MLNR) is introduced for Rapid Speaker Adaptation (RSA). MLNR, which is conceptually simple, adapts the Gaussian means of a speaker independent (SI) model to the data of a new speaker by assuming a non-linear mapping from the SI Gaussian means to the adapted Gaussian means. It performs a non linear regression between maximum likelihood (ML) estimates of the means and the speaker independent means using General Regression Neural Networks (GRNN). Evaluation on the Wall Street Journal benchmark shows that the suggested scheme outperforms different conventional approaches.
Mohamad Hasan Bahari, Hugo Van hamme
ICME2
2011 Modelling vocabulary acquisition, adaptation and generalization in infants using adaptive Bayesian PLSA
Joris Driesen, Hugo Van hamme
Neurocomputing2
2011 Sparse conjugate directions pursuit with application to fixed-size kernel models
abstract
This work studies an optimization scheme for computing sparse approximate solutions of over-determined linear systems. Sparse Conjugate Directions Pursuit (SCDP) aims to construct a solution using only a small number of nonzero (i.e. nonsparse) coefficients. Motivations of this work can be found in a setting of machine learning where sparse models typically exhibit better generalization performance, lead to fast evaluations, and might be exploited to define scalable algorithms. The main idea is to build up iteratively a conjugate set of vectors of increasing cardinality, in each iteration solving a small linear subsystem. By exploiting the structure of this conjugate basis, an algorithm is found (i) converging in at most D iterations for D -dimensional systems, (ii) with computational complexity close to the classical conjugate gradient algorithm, and (iii) which is especially efficient when a few iterations suffice to produce a good approximation. As an example, the application of SCDP to Fixed-Size Least Squares Support Vector Machines (FS-LSSVM) is discussed resulting in a scheme which efficiently finds a good model size for the FS-LSSVM setting, and is scalable to large-scale machine learning tasks. The algorithm is empirically verified in a classification context. Further discussion includes algorithmic issues such as component selection criteria, computational analysis, influence of additional hyper-parameters, and determination of a suitable stopping criterion.
Peter Karsmakers, Kristiaan Pelckmans, Kris De Brabanter, Hugo Van hamme, Johan A. K. Suykens
Mach. Learn.4
2011 Advances in Missing Feature Techniques for Robust Large-Vocabulary Continuous Speech Recognition
abstract
Missing feature theory (MFT) has demonstrated great potential for improving the noise robustness in speech recognition. MFT was mostly applied in the log-spectral domain since this is also the representation in which the masks have a simple formulation. However, with diagonally structured covariance matrices in the log-spectral domain, recognition performance can only be maintained at the cost of increasing the number of Gaussians drastically. In this paper, MFT can be applied for static and dynamic features in any feature domain that is a linear transform of log-spectra. A crucial part in MFT-systems is the computation of reliability masks from noisy data. The proposed system operates on either binary masks where hard decisions are made about the reliability of the data or on fuzzy masks which use a soft decision criterion. For real-life deployments, a compensation for convolutional noise is also required. Channel compensation in speech recognition typically involves estimating an additive shift in the log-spectral or cepstral domain. To deal with the fact that some features are considered as unreliable, a maximum-likelihood estimation technique is integrated in the back-end recognition process of the MFT system to estimate the channel. Hence, the resulting MFT-based recognizer can deal with both additive and convolutional noise and shows promising results on the Aurora4 large-vocabulary database.
Maarten Van Segbroeck, Hugo Van hamme
IEEE Trans. Speech Audio Process.2
2010 Histogram equalization and noise masking for robust speech recognition
abstract
Mismatch between training and test conditions deteriorates the performance of speech recognizers. This paper investigates the combination of parametric histogram equalization (pHEQ) and noise masking to compensate for the mismatch caused by additive noise. The proposed front-end maps the distribution of the observed power spectrum vectors to a target distribution. The target distribution matches the distribution of the noise free training data except for an artificially reduced signal-to-noise ratio. Different power spectrum estimation algorithms are used to estimate the noise distribution as used internally by pHEQ more reliably under non-stationary noise conditions. The proposed front-end is evaluated on the Aurora4 database and shows a significant improvement w.r.t. mean-normalized Mel-frequency spectral coefficients. Moreover, the performance could be further improved if better estimates of the instantaneous noise power spectrum were available.
Xueru Zhang, Kris Demuynck, Hugo Van hamme
ICASSP3
2010 Feature versus model based noise robustness
abstract
Over the years, the focus in noise robust speech recognition has shifted from noise robust features to model based techniques such as parallel model combination and uncertainty decoding. In this paper, we contrast prime examples of both approaches in the context of large vocabulary recognition systems such as used for automatic audio indexing and transcription. We look at the approximations the techniques require to keep the computational load reasonable, the resulting computational cost, and the accuracy measured on the Aurora4 benchmark. The results show that a well designed feature based scheme is capable of providing recognition accuracies at least as good as the model based approaches at a substantially lower computational cost. © 2010 ISCA.
Kris Demuynck, Xueru Zhang, Dirk Van Compernolle, Hugo Van hamme
INTERSPEECH4
2010 Learning from images and speech with Non-negative Matrix Factorization enhanced by input space scaling
abstract
Computional learning from multimodal data is often done with matrix factorization techniques such as NMF (Non-negative Matrix Factorization), pLSA (Probabilistic Latent Semantic Analysis) or LDA (Latent Dirichlet Allocation). The different modalities of the input are to this end converted into features that are easily placed in a vectorized format. An inherent weakness of such a data representation is that only a subset of these data features actually aids the learning. In this paper, we first describe a simple NMF-based recognition framework operating on speech and image data. We then propose and demonstrate a novel algorithm that scales the inputs of this framework in order to optimize its recognition performance.
Joris Driesen, Hugo Van hamme, W. Bastiaan Kleijn
SLT2
2009 Adaptive non-negative matrix factorization in a computational model of language acquisition
abstract
During the early stages of language acquisition, young infants face the task of learning a basic vocabulary without the aid of prior linguistic knowledge. It is believed the long term episodic memory plays an important role in this process. Experiments have shown that infants retain large amounts of very detailed episodic information about the speech they perceive (e.g. [1]). This weakly justifies the fact that some algorithms attempting to model the process of vocabulary acquisition computationally process large amounts of speech data in batch. Non-negative Matrix Factorization (NMF), a technique that is particularly successful in data mining but can also be applied to vocabulary acquisition (e.g. [2]), is such an algorithm. In this paper, we will integrate an adaptive variant of NMF into a computational framework for vocabulary acquisition, foregoing the need for long term storage of speech inputs, and experimentally show its accuracy matches that of the original batch algorithm. Copyright © 2009 ISCA.
Joris Driesen, Louis ten Bosch, Hugo Van hamme
INTERSPEECH3
2009 Evaluation of phone lattice based speech decoding
abstract
Previously, we proposed a flexible two-layered speech recogniser architecture, called FLaVoR. In the first layer an unconstrained, task independent phone recogniser generates a phone lattice. Only in the second layer the task specific lexicon and language model are applied to decode the phone lattice and produce a word level recognition result. In this paper, we present a further evaluation of the FLaVoR architecture. The performance of a classical single-layered architecture and the FLaVoR architecture are compared on two recognition tasks, using the same acoustic, lexical and language models. On the large vocabulary Wall Street Journal 5k and 20k benchmark tasks, the two-layered architecture resulted in slightly but not significantly better word error rates. On a reading error detection task for a reading tutor for children, the FLaVoR architecture clearly outperformed the single-layered architecture. Index Terms: ASR architecture, phone lattice decoding, system assessment 1.
Jacques Duchateau, Kris Demuynck, Hugo Van hamme
INTERSPEECH3
2009 Application of noise robust MDT speech recognition on the SPEECON and speechdat-car databases
abstract
Contains fulltext : 79368.pdf (author's version ) (Open Access)
Jort F. Gemmeke, Maarten Van Segbroeck, Bert Cranen, Hugo Van hamme
INTERSPEECH5
2009 Applying non-negative matrix factorization on time-frequency reassignment spectra for missing data mask estimation
abstract
The application of Missing Data Theory (MDT) has shown to improve the robustness of automatic speech recognition (ASR) systems. A crucial part in a MDT-based recognizer is the computation of the reliability masks from noisy data. To estimate accurate masks in environments with unknown, non-stationary noise statistics, we need to rely on a strong model for the speech. In this paper, an unsupervised technique using non-negative matrix factorization (NMF) discovers phone-sized time-frequency patches into which speech can be decomposed. The input matrix for the NMF is constructed using a high resolution and reassigned time-frequency representation. This representation facilitates an accurate detection of the patches that are active in unseen noisy speech. After further denoising of the patch activations, speech and noise can be reconstructed from which missing feature masks are estimated. Recognition experiments on the Aurora2 database demonstrate the effectiveness of this technique. Index Terms: noise robust speech recognition, missing data techniques, mask estimation, speech separation, non-negative matrix factorization 1.
Maarten Van Segbroeck, Hugo Van hamme
INTERSPEECH2
2009 A Computational Model of Language Acquisition: the Emergence of Words
abstract
In this paper, we discuss a computational model that is able to detect and build word-like representations on the basis of sensory input. The model is designed and tested with a further aim to investigate how infants may learn to communicate by means of spoken language. The computational model makes use of a memory, a perception module, and the concept of 'learning drive'. Learning takes place within a communicative loop between a 'caregiver' and the 'learner'. Experiments carried out on three European languages with different genetic background (Finnish, Swedish, and Dutch) show that a robust word representation can be learned in using less than 100 acoustic tokens (examples) of that word. The model is inspired by the memory structure that is assumed functional for human cognitive processing.
Louis ten Bosch, Lou Boves, Hugo Van hamme, Roger K. Moore
Fundam. Informaticae3
2009 Developing a reading tutor: Design and evaluation of dedicated speech recognition and synthesis modules
Jacques Duchateau, Yuk On Kong, Leen Cleuren, Lukas Latacz, Jan Roelens, Abdurrahman Samir, Kris Demuynck, Pol Ghesquière, Werner Verhelst, Hugo Van hamme
Speech Commun.10
2009 Unsupervised learning of time-frequency patches as a noise-robust representation of speech
Maarten Van Segbroeck, Hugo Van hamme
Speech Commun.2
2009 Automatic voice onset time estimation from reassignment spectra
Veronique Stouten, Hugo Van hamme
Speech Commun.2
2008 Unsupervised learning of auditory filter banks using non-negative matrix factorisation
abstract
Non-negative matrix factorisation (NMF) is an unsupervised learning technique that decomposes a non-negative data matrix into a product of two lower rank non-negative matrices. The non-negativity constraint results in a parts-based and often sparse representation of the data. We use NMF to factorise a matrix with spectral slices of continuous speech to automatically find a feature set for speech recognition. The resulting decomposition yields a filter bank design with remarkable similarities to perceptually motivated designs, supporting the hypothesis that human hearing and speech production are well matched to each other. We point out that the divergence cost criterion used by NMF is linearly dependent on energy, which may influence the design. We will however argue that this does not significantly affect the interpretation of our results. Furthermore, we compare our filter bank with several hearing models found in literature. Evaluating the filter bank for speech recognition shows that the same recognition performance is achieved as with classical MEL-based features.
Alexander Bertrand, Kris Demuynck, Veronique Stouten, Hugo Van hamme
ICASSP4
2008 Fast speaker adaptation using non-negative matrix factorization
abstract
This paper describes a new method for fast speaker adaptation in large vocabulary recognition systems. As in most HMM-based recognizers, the observation densities are modeled as a weighted sum of Gaussian densities. Instead of adapting the means of the Gaussian densities, which is typically done, the weights for the Gaussian densities in the states are adapted. By applying non-negative matrix factorization (NMF) in the proposed method, very fast adaptation was achieved. Experiments on the Wall Street Journal benchmark recognition task show relative improvements between 5% and 15%, while the adaptation converges within 0.2 seconds. Analysis of the latent speakers found by NMF learns that these latent speakers reflect the gender of the speaker most prominently, even when vocal tract length normalization is used, and that they reflect the speaker’s age more clearly than the speaker’s regional influences or dialect.
Jacques Duchateau, Tobias Leroy, Kris Demuynck, Hugo Van hamme
ICASSP4
2008 Estimation of the voicing cut-off frequency contour of natural speech based on harmonic and aperiodic energies
abstract
We present a new algorithm for the automatic estimation of the voicing cut-off frequency (VCO), i.e., the frequency that separates the periodic low-frequency part from the aperiodic high-frequency part in voiced segments of natural speech. Starting from the power spectrum of a two pitch period speech frame, we define the VCO to be located at the frequency for which the sum of the periodic and aperiodic energy in the spectral band below and above that frequency respectively, is maximised. By formulating the problem in terms of a score function we are able to apply a dynamic programming based smoothing technique. Remarkably smooth and accurate VCO contours were obtained, despite the simplicity of the proposed algorithm. In a formal evaluation the algorithm compares favourably to two existing VCO estimation techniques.
Kris Hermus, Laurent Girin, Hugo Van hamme, Sufian Irhimeh
ICASSP3
2008 Robust speech recognition using missing data techniques in the prospect domain and fuzzy masks
abstract
Missing data theory (MDT) has been applied to handle the problem of noise-robust speech recognition. Conventional MDT-systems require acoustic models that are expressed in the log-spectral rather than in the cepstral domain, which leads to a loss in accuracy. Therefore, we have already introduced a MDT-technique that can be applied in any feature domain that is a linear transform of log-spectra. This MDT-system requires hard decisions about the reliability of each spectral component. When computed from noisy data, misclassification errors in the mask are hardly unavoidable and the recognition rate will significantly degrade. The risk of misclassifications can be reduced by estimating a probability that the component is reliable, e.g. a fuzzy mask. In this paper, we extend our MDT-system to be applied in the probabilistic decision framework. Experiments on the Aurora2 database demonstrate a further increase in recognition accuracy, especially at low SNRs.
Maarten Van Segbroeck, Hugo Van hamme
ICASSP2
2008 A computational model of language acquisition: focus on word discovery
abstract
Young infants learn words by detecting patterns in the speech signal and by associating these patterns to stimuli presented by non-speech modalities (e.g vision). In this paper, we model this behaviour by designing and testing a computational model of word discovery. The model is able to build word-like representations on the basis of multimodal input data. The discovery of words (and word-like entities) takes place within a communicative loop between two protagonists, a ’carer ’ and the ’learner’. Experiments carried out on three different European languages (Finnish, Swedish, and Dutch) show that a robust word representation can be learned in using about 50 acoustic tokens (examples) of that word. The model is inspired by the memory structure that is assumed functional for human speech processing. Index Terms: language acquisition, unsupervised word detection, computational modelling 1.
Louis ten Bosch, Hugo Van hamme, Lou Boves
INTERSPEECH2
2008 Improving the multigram algorithm by using lattices as input
abstract
The multigram algorithm is a statistical technique that can be used for extracting recurring patterns from a sequential input. When provided with a symbol sequence representing a speech signal, it is able to extract word-like patterns from it, despite the large amount of subsequences that can represent a single word. For this, it uses statistical information derived from the entire input. However, due to the abstraction of speech to symbols, much of the information originally present in the signal is no longer available to the algorithm. In this paper we propose a way of using a richer abstraction of the signal in the form of a lattice. Furthermore, a way of grounding recurring patterns to concepts in other modalities will be presented. Finally, the information learned by the algorithm using both kinds of input is tested in a recognition experiment. This will show that the use of lattices leads to a significant improvement in terms of recognition rate.
Joris Driesen, Hugo Van hamme
INTERSPEECH2
2008 Lip synchronization: from phone lattice to PCA eigen-projections using neural networks
abstract
Lip synchronization is the process of generating natural lip movements from a speech signal. In this work we address the lip-sync problem using an automatic phone recognizer that generates a phone lattice carrying posterior probabilities. The acoustic feature vector contains the posterior probabilities of all the phones over a time window centered at the current time point. Hence this representation characterizes the phone recognition output including the confusion patterns caused by its limited accuracy. A 3D face model with varying texture is computed by analyzing a video recording of the speaker using a 3D morphable model. Training a neural network using 30 000 data vectors from an audiovisual recording in Dutch resulted in a very good simulation of the face on independent data sets of the same or of a different speaker.
Samer Al Moubayed, Michaël De Smet, Hugo Van hamme
INTERSPEECH3
2008 Discriminative model combination and language model selection in a reading tutor for children
abstract
In this paper, we suggest the use of general acoustic and language models to deal with the mismatch between the training and testing data of a reading tutor for children. The testing data consist of isolated real and non-existing (pseudo) words, while the training data consist of continuous readings of Dutch sentences. General acoustic (e.g. context independent) and language models (e.g. bigram phone language models) are proposed as they implicitly better model the hesitant nature of the testing data. Discriminative model combination (DMC) is modified to provide different weights for different phones and was utilized to combine the new models into the baseline system. Combination of general acoustic and language models into the baseline system using DMC significantly lowers the system phone error rate, by 3.5 % relative to the baseline system for the non-existing (pseudo) words.
Abdurrahman Samir, Jacques Duchateau, Hugo Van hamme
INTERSPEECH3
2008 Comparison of variable selection methods and classifiers for native accent identification
abstract
Acoustic differences are so subtle in a native accent identification (AID) task that a brute force frame-based Gaussian Mixture Model (GMM) fails to discover the tiny distinctions [1]. Apart from the frame-based framework, in this paper we propose a vector-based speaker modeling method, to which common support vector machine (SVM) kernels can be applied. The vector-based speaker model is composed of the concatenation of the average acoustic representations of all phonemes. SVM and GMM classifiers are compared on the speaker models. Moreover, based on the observation that accents only differ in a limited number of phonemes, a variable selection framework is indispensable to select accent relevant features. We investigate a forward selection method, Analysis of Variance (ANOVA) , and a backward selection method, SVM- Recursive Feature Elimination (SVM-RFE). We find that the multiclass SVM-RFE achieves comparable performance with the ANOVA on optimally selected variable sets, while it obtains excellent performance with very few features in low dimensions. Results demonstrate the effectiveness of the proposed speaker models together with the SVM classifier both in low dimensions and in high dimensions as well as the necessity of variable selection. Index Terms: variable selection, native accent identification, support vector machines, recursive feature elimination, cross
Tingyao Wu, Peter Karsmakers, Hugo Van hamme, Dirk Van Compernolle
INTERSPEECH3
2008 HAC-models: a novel approach to continuous speech recognition
abstract
In this paper, a bottom-up, activation-based paradigm for continuous speech recognition is described. Speech is described by co-occurrence statistics of acoustic events over an analysis window of variable length, leading to a vectorial representation of high but fixed dimension called "Histogram of Acoustic Co-occurrence" (HAC). During training, recurring acoustic patterns are discovered and associated to words through non-negative matrix factorisation. During testing, word activations are computed from the HAC-representation and their time of occurrence is estimated. Hence, words in a continuous utterance can be detected, ordered and located. Copyright © 2008 ISCA.
Hugo Van hamme
INTERSPEECH1
2008 Children's Oral Reading Corpus (CHOREC): Description and Assessment of Annotator Agreement
Leen Cleuren, Jacques Duchateau, Pol Ghesquière, Hugo Van hamme
LREC4
2008 Recording Speech of Children, Non-Natives and Elderly People for HLT Applications: the JASMIN-CGN Corpus
Catia Cucchiarini, Joris Driesen, Hugo Van hamme, Eric Sanders
LREC3
2008 Discovering Phone Patterns in Spoken Utterances by Non-Negative Matrix Factorization
abstract
We present a technique to automatically discover the (word-sized) phone patterns that are present in speech utterances. These patterns are learnt from a set of phone lattices generated from the utterances. Just like children acquiring language, our system does not have prior information on what the meaningful patterns are. By applying the non-negative matrix factorization algorithm to a fixed-length high-dimensional vector representation of the speech utterances, a decomposition in terms of additive units is obtained. We illustrate that these units correspond to words in case of a small vocabulary task. Our result also raises questions about whether explicit segmentation and clustering are needed in an unsupervised learning context.
Veronique Stouten, Kris Demuynck, Hugo Van hamme
IEEE Signal Process. Lett.3
2007 DCT-Based Amplitude and Frequency Modulated Harmonic-Plus-Noise Modelling for Text-to-Speech Synthesis
abstract
We present a harmonic-plus-noise modelling (HNM) strategy in the context of corpus-based text-to-speech (TTS) synthesis, in which whole speech phonemes are modelled in their integrity, contrary to the traditional frame-based approach. The pitch and amplitude trajectories of each phoneme are modelled with a low-order DCT expansion. The parameter analysis algorithm is to a large extent aided and guided by the pitch contours, and by the phonetic annotation and segmentation information that is available in any TTS system. The major advantages of our model are: few parameter interpolation points during synthesis (one per phoneme), flexible time and pitch modifications, and a reduction in the number of model parameters which is favourable for low bit rate coding in TTS for embedded applications. Listening tests on TTS sentences have shown that very natural speech can be obtained, despite the compactness of the signal representation.
Kris Hermus, Hugo Van hamme, Werner Verhelst, Sufian Irhimeh, Jan De Moortel
ICASSP (4)2
2007 Automatic assessment of children's reading level
abstract
In this paper, an automatic system for the assessment of reading in children is described and evaluated. The assessment is based on a reading test with 40 words, presented one by one to the child by means of a computerized reading tutor. The score that expresses the child’s reading performance is calculated as the total time needed to read the 40 words divided by the number of correctly read words. In each grade, children are classified in 5 groups based on their score as provided by human annotators. We show that when the score for a child is assessed automatically using a speech recognizer, a classification can be obtained with a substantial agreement (Cohen’s Kappa over 0.6) with the human classification. As all children in the experiments were classified either correctly or in an adjoining group, we can conclude that the proposed system can provide large time gains in current manual classification procedures. Index Terms: computer aided language learning, reading assessment, ASR for children.
Jacques Duchateau, Leen Cleuren, Hugo Van hamme, Pol Ghesquière
INTERSPEECH3
2007 Fixed-size kernel logistic regression for phoneme classification
abstract
Kernel logistic regression (KLR) is a popular non-linear classification technique. Unlike an empirical risk minimization approach such as employed by Support Vector Machines (SVMs), KLR yields probabilistic outcomes based on a maximum likelihood argument which are particularly important in speech recognition. Different from other KLR implementations we use a Nyström approximation to solve large scale problems with estimation in the primal space such as done in fixed-size Least Squares Support Vector Machines (LS-SVMs). In the speech experiments it is investigated how a natural KLR extension to multi-class classification compares to binary KLR models coupled via a one-versus-one coding scheme. Moreover, a comparison to SVMs is made. Index Terms: phoneme classification, kernel logistic regression, large-scale, multi-class
Peter Karsmakers, Kristiaan Pelckmans, Johan A. K. Suykens, Hugo Van hamme
INTERSPEECH4
2007 Vector-quantization based mask estimation for missing data automatic speech recognition
abstract
The application of Missing Data Theory (MDT) has shown to improve the robustness of automatic speech recognition (ASR) systems. A crucial part in a MDT-based recognizer is the computation of the reliability masks from noisy data. To estimate accurate masks in environments with unknown, non-stationary noise statistics only weak assumptions can be made about the noise and we need to rely on a strong model for the speech. In this paper, we present a missing data detector that uses harmonicity in the noisy input signal and a vector quantizer (VQ) to confine speech models to a subspace. The resulting system can deal with additive and convolutional noise and shows promising results on the Aurora4 large vocabulary database. Index Terms: speech recognition, noise robustness, missing data mask estimation, speech separation
Maarten Van Segbroeck, Hugo Van hamme
INTERSPEECH2
2007 Automatically learning the units of speech by non-negative matrix factorisation
abstract
We present an unsupervised technique to discover the (word-sized) speech units in which a corpus of utterances can be de-composed. First, a fixed-length high-dimensional vector rep-resentation of the utterances is obtained. Then, the resulting matrix is decomposed in terms of additive units by applying the non-negative matrix factorisation algorithm. On a small vocab-ulary task, the obtained basis vectors each represent one of the uttered words. We also investigate the amount of speech data that is needed to obtain a correct set of basis vectors. By de-creasing the number of occurrences of the words in the corpus, an indication of the learning rate of the system is obtained. Index Terms: matrix factorisation, word segmentation, phone lattices, language acquisition.
Veronique Stouten, Kris Demuynck, Hugo Van hamme
INTERSPEECH3
2007 Estimation of the Voicing Cut-Off Frequency Contour Based on a Cumulative Harmonicity Score
abstract
We present a new algorithm for the estimation of the voicing cut-off frequency (VCO), i.e., the frequency that separates the harmonic low-frequency part from the aperiodic high-frequency part in voiced speech. The VCO is estimated as the frequency for which the sum of the harmonicity scores of all pitch harmonics below that frequency is maximized. The algorithm is combined with a powerful dynamic programming approach to track the VCO estimates over time. Remarkably accurate and smooth VCO contours are obtained, despite the simplicity of the algorithm. Applications include a.o. (sinusoidal) speech modeling, coding, and synthesis, as well as harmonic speech analysis for, e.g., automatic speech recognition.
Kris Hermus, Hugo Van hamme, Sufian Irhimeh
IEEE Signal Process. Lett.2
2006 Application of Minimum Statistics and Minima Controlled Recursive Averaging Methods to Estimate a Cepstral Noise Model for Robust ASR
abstract
Many compensation techniques, both in the model and feature domain, require an estimate of the noise statistics to compensate for the clean speech degradation in adverse environments. We explore how two spectral noise estimation approaches can be applied in the context of model-based feature enhancement. The minimum statistics method and the improved minima controlled recursive averaging method are used to estimate the noise power spectrum based only on the noisy speech. The noise mean and variance estimates are nonlinearly transformed to the cepstral domain and used in the Gaussian noise model of MBFE. We show that the resulting system achieves an accuracy on the Aurora2 task that is comparable to MBFE with prior knowledge on noise. Finally, this performance can be significantly improved when the MS or EMCRA noise mean is reestimated based on a clean speech model
Veronique Stouten, Hugo Van hamme, Patrick Wambacq
ICASSP (1)2
2006 Maximum Likelihood Based Temporal Frame Selection
abstract
In this paper, we propose a maximum likelihood (ML) based frame selection approach. A fixed frame rate adopted in most state-of-the-art speech recognition systems can face some problems, such as accidentally meeting noisy frames, assigning the same importance to each frame, and pitch asynchronous representation. As an attempt to avoid those problems, our approach selects reliable frames from a fine resolution along the time axis. In a phoneme recognition task, we show that significant improvements are achieved with the frame selection approach comparing to a system with a fixed frame rate.
Tingyao Wu, Dirk Van Compernolle, Jacques Duchateau, Hugo Van hamme
ICASSP (1)4
2006 Handling Time-Derivative Features in a Missing Data Framework for Robust Automatic Speech Recognition
abstract
We present a novel approach to handling dynamic (time derivative or delta) features for automatic speech recognition using a HMM/GMM-architecture and based on missing data techniques for noise robustness. The static and the dynamic features are imputed in the observations based on an acoustic model expressed in a domain that is a linear transform of the log-spectra and taking bounds into account. The reliability masks of the dynamic features are ternary. We describe a method for computing oracle masks for dynamic features. We also propose a simple method to derive dynamic masks from the reliability mask of the static features. We find that using bounds in the imputation is advantageous, both for oracle masks and for masks derived from the noisy observations
Hugo Van hamme
ICASSP (1)1
2006 Developing an automatic assessment tool for children²s oral reading
abstract
Automation of oral reading assessment and of feedback in a reading tutor is a very challenging task. This paper describes our research aiming at developing such automated systems. First topic is the recording and annotation of CHOREC, the Flemish database of children’s oral reading we develop in order to characterize oral reading processes statistically. Next, we propose a classification of both oral reading strategies and errors, which provides the basis of the envisaged assessment and feedback. Finally, experimental results show that our two-layered recognition system is able to provide high reading miscue detection rates, while only few correctly read words are erroneously tagged as miscue. Index Terms: reading assessment, database annotation, speech technology, education.
Leen Cleuren, Jacques Duchateau, Alain Sips, Pol Ghesquière, Hugo Van hamme
INTERSPEECH5
2006 Robust phone lattice decoding
abstract
Most ASR systems adopt an all-in-one approach: acoustic model, lexicon and language model are all applied simultaneously, thus forming a single large search space. This way, both lexicon and language model help in constraining the search at an early stage which greatly improves its efficiency. However, such close integration comes at a cost: all resources must be kept simple. Achieving higher accuracy in unconstrained LVCSR tasks will require more complex resources while at the same time the 'unconstrainedness' of the task reduces the effectiveness of the all-in-one approach. Therefore, we propose a modular two-layered architecture. First, a pure acoustic-phonemic search generates a dense phone network. Next a robust decoder finds those words from the lexicon that match well with the phone sequences encoded in the phone network. In this paper we investigate the properties the robust word decoder must have and we propose an efficient search algorithm.
Kris Demuynck, Dirk Van Compernolle, Hugo Van hamme
INTERSPEECH3
2006 Handling convolutional noise in missing data automatic speech recognition
abstract
Missing Data Techniques have already shown their effectiveness in dealing with additive noise in automatic speech recognition systems. For real-life deployments, a compensation for linear filtering distortions is also required. Channel compensation in speech recognition typically involves estimating an additive shift in the log-spectral or cepstral domain. This paper explores a maximum likelihood technique to estimate this model offset while some data are missing. Recognition experiments on the Aurora2 recognition task demonstrate the effectiveness of this technique. In particular, we show that our method is more accurate than previously published methods and can handle narrow-band data. Index Terms: speech recognition, missing data techniques, convolutional distortion, channel estimation.
Maarten Van Segbroeck, Hugo Van hamme
INTERSPEECH2
2006 Single frame selection for phoneme classification
abstract
Our former study [1] has shown that maximum likelihood (ML) based frame selection, which selects reliable frames from a high resolution along the time axis, helps to improve the discrimination between phonemes. In this paper, we present our recent research on single frame selection for a phoneme classification task. A new single selection, which only selects one frame for one state in an Hidden Markov Model (HMM), is proposed. The new technique takes likelihoods of frames and their positions in a phoneme segment into account at the same time, and selects very few frames to represent the spectral evolution of the phoneme. Furthermore, we also show that for a low model complexity, a phoneme model trained by selected frames is more discriminative than a model using all frames.
Tingyao Wu, Dirk Van Compernolle, Jacques Duchateau, Hugo Van hamme
INTERSPEECH4
2006 JASMIN-CGN: Extension of the Spoken Dutch Corpus with Speech of Elderly People, Children and Non-natives in the Human-Machine Interaction Modality
Catia Cucchiarini, Hugo Van hamme, Olga van Herwijnen, Felix Smits
LREC2
2006 Model-based feature enhancement with uncertainty decoding for noise robust ASR
Veronique Stouten, Hugo Van hamme, Patrick Wambacq
Speech Commun.2
2005 Effect of Phase-Sensitive Environment Model and Higher Order VTS on Noisy Speech Feature Enhancement
abstract
Model-based techniques for robust speech recognition often require the statistics of noisy speech. In this paper, we propose two modifications to obtain more accurate versions of the statistics of the combined HMM (starting from a clean speech and a noise model). Usually, the phase difference between speech and noise is neglected in the acoustic environment model. However, we show how a phase-sensitive environment model can be efficiently integrated in the context of multi-stream model-based feature enhancement and gives rise to more accurate covariance matrices for the noisy speech. Also, by expanding the vector Taylor series up to the second order term, an improved noisy speech mean can be obtained. Finally, we explain how the front-end clean speech model itself can be improved by a preprocessing of the training data. Recognition results on the Aurora4 database illustrate the effect on the noise robustness for each of these modifications.
Veronique Stouten, Hugo Van hamme, Patrick Wambacq
ICASSP (1)2
2005 Statistical language models for large vocabulary spontaneous speech recognition in dutch
abstract
In state-of-the-art large vocabulary automatic recognition systems, a large statistical language model is used, typically an N-gram. However in order to estimate this model, a large database of sentences or texts in the same style as the recognition task is needed. For spontaneous speech one doesn't dispose of such database since it should consist of accurate thus expensive orthographic transcriptions of spoken audio. This paper investigates how readily available large news paper corpora can be used to improve language models for spontaneous speech recognition although both language styles differ considerably. A technique is proposed that does a perplexity based automatic selection of appropriate news paper articles and that subsequently uses these texts in the language model estimation. Recognition experiments on spontaneous broadcast speech in Dutch showed significant improvements using this technique.
Jacques Duchateau, Dong Hoon Van Uytsel, Hugo Van hamme, Patrick Wambacq
INTERSPEECH3
2005 PROSPECT features and their application to missing data techniques for vocal tract length normalization
abstract
fmodel Speaker normalization by (piecewise) linear warping of the frequency axis is a popular method because of its simplicity and effectiveness. However, when this so-called vocal tract length normalization is applied to map test speakers with a shorter vocal tract onto acoustic models trained on speakers with a longer vocal tract, there is important information missing in the frequency bins at the high end of the spectrum. Usually, this missing information is reconstructed by ad hoc rules or through extrapolation of the spectrum. In this paper, we present a new method to estimate the content of those bins. The proposed solution is derived from Missing Data Techniques, that are used for noise robust speech recognizers. To alleviate the accuracy loss associated with Missing Data Techniques that are usually expressed in the spectral domain, we apply the PROSPECT feature representation introduced about a year ago. We demonstrate the superiority of our approach on the TIDigits database. fNyq f Nyq α fwarp arctan(α −1) forig fknee piecewise linear fNyq linear fdata 1.
Wim Jansen, Hugo Van hamme
INTERSPEECH2
2005 Kalman and unscented kalman filter feature enhancement for noise robust ASR
abstract
Model-based feature enhancement is an ASR front-end technique to increase the robustness of the recogniser in noisy environments. However, its MMSE-estimates of the clean speech feature vectors are based only on the static components at the current frame. In this paper, we show how the Kalman filter framework can be seen as a natural extension that incorporates both the current and the previous frames in the enhancement process. Because multiple Kalman filters are run in parallel, the global clean speech estimate is given by a weighted linear combination of the individual MMSE-estimates. Also, the unscented transformation is considered to avoid the linearisation of the cepstral domain observation equation. We present experimental results on the Aurora2 database for both the multi-modal Kalman and the unscented Kalman filter feature enhancement. 1.
Veronique Stouten, Hugo Van hamme, Patrick Wambacq
INTERSPEECH2
2004 Robust speech recognition using cepstral domain missing data techniques and noisy masks
abstract
Missing data techniques (MDT) have been shown to be an effective method for curing the performance degradation of HMM-based speech recognition systems operating on noisy signals. However, a major drawback of the approach is that MDT requires that the acoustic model be expressed as a mixture of diagonal Gaussians in the log-spectral domain, whereas a higher accuracy can be obtained with Gaussian mixtures in the cepstral domain. The paper describes a recognizer based on the recently described cepstral-domain MDT approach using missing data masks computed from the noisy signal. It exploits a novel decision criterion that integrates harmonicity with signal-to-noise ratio and which makes minimal assumptions on the noise. The system is shown to exhibit a recognition accuracy that is comparable to the ETSI advanced front-end reference.
Hugo Van hamme
ICASSP (1)1
2004 Joint removal of additive and convolutional noise with model-based feature enhancement
abstract
In this paper we describe how we successfully extended the model-based feature enhancement (MBFE) algorithm to jointly remove additive and convolutional noise from corrupted speech. Although a model of the clean speech can incorporate prior knowledge into the feature enhancement process, this model no longer yields an accurate fit if a different microphone is used. To cure the resulting performance degradation, we merge a new iterative EM algorithm to estimate the channel, and the MBFE-algorithm to remove nonstationary additive noise. In the latter, the parameters of a shifted clean speech HMM and a noise HMM are first combined by a vector Taylor series approximation and then the state-conditional MMSE-estimates of the clean speech are calculated. Recognition experiments confirmed the superior performance on the Aurora4 recognition task. An average relative reduction in WER of 12% and 2.8% on the clean and multi condition training respectively, was obtained compared to the Advanced Front-End standard.
Veronique Stouten, Hugo Van hamme, Patrick Wambacq
ICASSP (1)2
2004 PROSPECT features and their application to missing data techniques for robust speech recognition
abstract
Missing data theory has been applied to the problem of speech recognition in adverse environments. The resulting systems require acoustic models that are expressed in the spectral rather than in the cepstral domain, which leads to loss of accuracy. Cepstral Missing Data Techniques (CMDT) surmount this disadvantage, but require significantly more computation. In this paper, we study alternatives to the cepstral representation that lead to more efficient MDT systems. The proposed solution, PROSPECT features (Projected Spectra), can be interpreted as a novel speech representation, or as an approximation of the inverse covariance (precision) matrix of the Gaussian distributions modeling the log-spectra.
Hugo Van hamme
INTERSPEECH1
2004 Accounting for the uncertainty of speech estimates in the context of model-based feature enhancement
abstract
In this paper we present two techniques to cover the gap between the true and the estimated clean speech features in the context of Model-Based Feature Enhancement (MBFE) for noise robust speech recognition.While in the output of every feature enhancement algorithm some residual uncertainty remains, currently this information is mostly discarded.Firstly, we explain how the generation of not only a global MMSEestimate of clean speech, but also several alternative (stateconditional) estimates are supplied to the back-end for recognition.Secondly, we explore the benefits of calculating the variance of the front-end estimate and incorporating this in the acoustic models of the recogniser.Experiments on the Aurora2 task confirmed the superior performance of the resulting system: an average increase in recognition accuracy from 85.65% to 88.50% was obtained for the clean training condition.
Hugo Van hamme, Patrick Wambacq, Veronique Stouten
INTERSPEECH1
2004 Use and Evaluation of Prosodic Annotations in Dutch
Jacques Duchateau, Tim Ceyssens, Hugo Van hamme
LREC3
2004 Evaluation and Adaptation of the Celex Dutch Morphological Database
Tom Laureys, Guy De Pauw, Hugo Van hamme, Walter Daelemans, Dirk Van Compernolle
LREC3
2003 FLavor: a flexible architecture for LVCSR
abstract
This paper describes a new architecture for large vocabulary continuous speech recognition (LVCSR), which will be developed within the project FLaVoR (Flexible Large Vocabulary Recognition). The proposed architecture abandons the standard all-in-one search strategy with integrated acoustic, lexical and language model information. Instead, a modular framework is proposed which allows for the integration of more complex linguistic components. The search process consists of two layers. First, a pure acoustic-phonemic search generates a dense phoneme network enriched with meta-data. Then, the output of the first layer is used by sophisticated language technology components for word decoding in the second layer. Preliminary experiments prove the feasibility of the approach.
Kris Demuynck, Tom Laureys, Dirk Van Compernolle, Hugo Van hamme
INTERSPEECH4
2003 Assessment of dereverberation algorithms for large vocabulary speech recognition systems
abstract
The performance of large vocabulary recognition systems, for instance in a dictation application, typically deteriorates severely when used in a reverberant environment. This can be partially avoided by adding a dereverberation algorithm as a speech signal preprocessing step. The purpose of this paper is to compare the effect of different speech dereverberation algorithms on the performance of a recognition system. Experiments were conducted on the Wall Street Journal dictation benchmark. Reverberation was added to the clean acoustic data in the benchmark both by simulation and by re-recording the data in a reverberant room. Moreover additive noise was added to investigate its effect on the dereverberation algorithms. We found that dereverberation based on a delay-and-sum beamforming algorithm has the best performance of the investigated algorithms.
Koen Eneman, Jacques Duchateau, Marc Moonen, Dirk Van Compernolle, Hugo Van hamme
INTERSPEECH5
2003 Two correction models for likelihoods in robust speech recognition using missing feature theory
abstract
In Missing Feature Theory (MFT), it is assumed that some of the features that are extracted from an observation are missing or unreliable. Applied to spectral features for noisy speech recognition, the clean feature values are known to be less than the observed noisy features. Based on this inequality constraint, an HMM-state-dependent clean speech value of the missing features can be inferred through maximum likelihood estimation. This paper describes two observed biases of the likelihood evaluated at the estimate. Theoretical and experimental evidence are provided that an upper bound on the accuracy is improved by applying computationally simple corrections for the number of free variables in the likelihood maximization and for the global acoustic space density function.
Hugo Van hamme
INTERSPEECH1
2003 Robust speech recognition using missing feature theory in the cepstral or LDA domain
abstract
When applying Missing Feature Theory to noise robust speech recognition, spectral features are labeled as either reliable or unreliable in the time-frequency plane. The acoustic model evaluation of the unreliable features is modified to express that their clean values are unknown or confined within bounds. Classically, MFT requires an assumption of statistical independence in the spectral domain, which deteriorates the accuracy on clean speech. In this paper, MFT is expressed in any domain that is a linear transform of (log-)spectra, for example for cepstra and their timederivatives. The acoustic model evaluation is recast as a nonnegative least squares problem. Approximate solutions are proposed and the success of the method is shown through experiments on the AURORA-2 database.
Hugo Van hamme
INTERSPEECH1
2003 Robust speech recognition using model-based feature enhancement
abstract
Maintaining a high level of robustness for Automatic Speech Recognition (ASR) systems is especially challenging when the background noise has a time-varying nature. We have implemented a Model-Based Feature Enhancement (MBFE) technique that not only can easily be embedded in the feature extraction module of a recogniser, but also is intrinsically suited for the removal of non-stationary additive noise. To this end we combine statistical models of the cepstral feature vectors of both clean speech and noise, using a Vector Taylor Series approximation in the power spectral domain. Based on this combined HMM, a global MMSE-estimate of the clean speech is then calculated. Because of the scalability of the applied models, MBFE is flexible and computationally feasible. Recognition experiments with this feature enhancement technique on the Aurora2 connected digit recognition task showed significant improvements on the noise robustness of the HTK recogniser.
Veronique Stouten, Hugo Van hamme, Kris Demuynck, Patrick Wambacq
INTERSPEECH2
2003 Evaluation of model-based feature enhancement on the AURORA-4 task
abstract
In this paper we focus on the challenging task of noise robustness for large vocabulary Continuous Speech Recognition (LVCSR) systems in non-stationary noise environments. We have extended our Model-Based Feature Enhancement (MBFE) algorithm – that we earlier successfully applied to small vocabulary CSR in the AURORA-2 framework – to cope with the new demands that are imposed by the large vocabulary size in the AURORA-4 task. To incorporate a priori knowledge of the background noise, we combine scalable Hidden Markov Models (HMMs) of the cepstral feature vectors of both clean speech and noise, using a Vector Taylor Series approximation in the power spectral domain. Then, a global MMSE-estimate of the clean speech is calculated based on this combined HMM. This technique is easily embeddable in the feature extraction module of a recogniser and is intrinsically suited for the removal of non-stationary additive noise. Our approach is validated on the AURORA-4 task, revealing a significant gain in noise robustness over the baseline.
Veronique Stouten, Hugo Van hamme, Jacques Duchateau, Patrick Wambacq
INTERSPEECH2
2002 Investigation of speech recognition over IP channels
abstract
In this paper we investigate the effects of IP channels on speech recognition systems and methods to recover the associated performance degradation. There are three major VoIP (voice over IP) distortion sources: speech encoding-decoding (codecs), packet loss and jitter (time-delay). To speech recognition systems distortions are mainly from packet loss and the speech codecs. Their effects on the recognizer's performance are systematically investigated by using four different ITU-T recommended speech codecs. The results show that the speech codecs introduce bigger degradation than the packet losses (random and burst). To recover the codec degradations we have applied the MLLR adaptation and a data-mixed retraining method. These techniques reduce the degradation by about 50%.
Jim Van Sciver, Jeff Z. Ma, Filiep Vanpoucke, Hugo Van hamme
ICASSP4
2000 Model-based feature enhancement for noisy speech recognition
abstract
In this paper, a new feature enhancement algorithm called model-based feature enhancement (MBFE) is introduced for noise robust speech recognition. In MBFE, statistical models (i.e., Gaussian HMM's) of the clean speech feature vectors and of the perturbing noise feature vectors are used to construct the optimal MMSE estimator of the clean speech feature vectors. The estimated clean speech features are then fed to a recognizer. The performance of MBFE is studied experimentally on a connected-digits recognition task in several additive noise conditions (synthetic white and impulsive noise, car noise, and machine tool noise are considered). The performance of MBFE is also compared to that of a state-of-the-art implementation of nonlinear spectral subtraction.
Christophe Couvreur, Hugo Van hamme
ICASSP2
2000 Evaluation of various confidence-based strategies for isolated word rejection
abstract
Three baseline isolated word rejection strategies are initially proposed and their rejection performance is evaluated on two different types of out-of-vocabulary (OOV) utterances: OOV similar in nature to the in-vocabulary (IV) ones versus OOV consisting of non-speech events such as coughs, clicks, smacks, etc. A general OOV model, referred to as garbage, is added in parallel to the IV models and then a confidence measure, based on the contrast of the first best, hypothesis score with the garbage score, is used as an IV-OOV classifier. The discriminative power of some other confidence measures, utilising the distance between the first two hypotheses in the N-best list, is also investigated. Further, a considerable improvement (up to 34%) of the baseline classification error rate is achieved when the garbage model is supplied with word transition penalties.
Elena Tsiporkova, Filiep Vanpoucke, Hugo Van hamme
ICASSP3
2000 Dialect adaptation for Mandarin Chinese speech recognition
Frédéric Beaugendre, Tom Claes, Hugo Van hamme
INTERSPEECH3
1999 Accuracy versus complexity in context dependent phone modeling
abstract
This paper presents two different directions to build HMM models which give enough acoustic resolution and t in limited user resources. They both refer to scaling down the acoustic models which are built with tied gaussian HMMs. The total number of gaussians is reduced by a pairwise merging, and the number of gaussians per state is reduced by selecting them based on the so called occupancy criterion. Experiments carried out on the WSJ recognition task show that after scaling down, no further training is needed when the number of gaussians or the number of gaussians per state is reduced up to a factor three. This is an advantage as retraining can not be executed by the final system user.
Jacques Duchateau, Kris Demuynck, Ioannis Dologlou, Patrick Wambacq, Dirk Van Compernolle, Hugo Van hamme
EUROSPEECH7
1996 An adaptive-beam pruning technique for continuous speech recognition
abstract
Pruning is an essential paradigm to build HMM-based large vocabulary speech recognisers that use reasonable computing resources.Unlikely sentence, word or subword hypotheses are removed from the search space when their likelihood falls outside a beam relative to the best scoring hypothesis.A method for automatically steering this beam such that the search space attains a predened size is presented.
Hugo Van hamme, Filip Van Aelten
ICSLP1
1994 ARDOSS: autoregressive domain spectral subtraction for robust speech recognition in additive noise
Hugo Van hamme
ICSLP1
1994 Comparison of acoustic features and robustness tests of a real-time recogniser using a hardware telephone line simulator
Hugo Van hamme, Guido Gallopyn, Ludwig Weynants, Bart D'hoore, Hervé Bourlard
ICSLP1
1988 Karhunen-Loeve analysis of dynamic sequences of thermographic images for early breast cancer detection
abstract
The Karhunen-Loeve transform (KLT) is applied to the analysis of dynamic sequences of thermograms describing the temporal evolution of body surface temperature following the application of an external thermal stimulus. The KLT may be evaluated either along the spatial or temporal dimensions of the data; the duality of both representations is emphasized. An example is presented to illustrate that the KLT allows an efficient data reduction and facilities tumor detection by highlighting physiologically important abnormalities in the time behavior of thermal patterns.>
Michael Unser, Hugo Van hamme, Patrick de Muynck, E. Van Denhaute, Jan Cornelis 0001
CVPR2