VLDB 2026 Research / reviewers in the wild / expert
Mikko Kurimo
dblp:k/MikkoKurimo
· DBLP profile ↗
128ranked-venue papers
14as first author
30since 2021 · last 2026
0000-0001-5278-7974ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 97 · 11 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 92 · 10 first-author · 24 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A study on the layer-wise transferability of self-supervised learning features for children's speech processing tasks
Abhijit Sinha, Hemant Kumar Kathania, Mikko Kurimo |
Speech Commun. | 3 |
| 2025 | Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) is crucial for improving human-computer interaction. Despite strides in monolingual SER, extending them to build a multilingual system remains challenging. Our goal is to train a single model capable of multilingual SER by distilling knowledge from multiple teacher models. To address this, we introduce a novel language-aware multi-teacher knowledge distillation method to advance SER in English, Finnish, and French. It leverages Wav2Vec2.0 as the foundation of monolingual teacher models and then distills their knowledge into a single multilingual student model. The student model demonstrates state-of-the-art performance, with a weighted recall of 72.9 on the English dataset and an unweighted recall of 63.4 on the Finnish dataset, surpassing fine-tuning and knowledge distillation baselines. Our method excels in improving recall for sad and neutral emotions, although it still faces challenges in recognizing anger and happiness. Mehedi Hasan Bijoy, Dejan Porjazovski, Tamás Grósz, Mikko Kurimo |
INTERSPEECH | 4 |
| 2025 | Is your model big enough? Training and interpreting large-scale monolingual speech foundation models
Yaroslav Getman, Tamás Grósz, Tommi Lehtonen, Mikko Kurimo |
INTERSPEECH | 4 |
| 2025 | Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland SwedishabstractMispronunciation detection (MD) models are the cornerstones of many language learning applications. Unfortunately, most systems are built for English and other major languages, while low-resourced language varieties, such as Finland Swedish (FS), lack such tools. In this paper, we introduce our MD model for FS, trained on 89 hours of first language (L1) speakers' spontaneous speech and tested on 33 minutes of L2 transcribed read-aloud speech. We trained a multilingual wav2vec 2.0 model with entropy regularization, followed by temperature scaling and top-k normalization after the inference to better adapt it for MD. The main novelty of our method lies in its simplicity, requiring minimal L2 data. The process is also language-independent, making it suitable for other low-resource languages. Our proposed algorithm allows us to balance Recall (43.2%) and Precision (29.8%), compared with the baseline model's Recall (77.5%) and Precision (17.6%). Nhan Phan, Mikko Kuronen, Maria Kautonen, Riikka Ullakonoja, Anna von Zansen, Yaroslav Getman, Ekaterina Voskoboinik, Tamás Grósz, Mikko Kurimo |
INTERSPEECH | 9 |
| 2025 | Beyond Traditional Speech Modifications : Utilizing Self Supervised Features for Enhanced Zero-Shot Children ASR
Abhijit Sinha, Hemant Kumar Kathania, Mikko Kurimo |
INTERSPEECH | 3 |
| 2024 | Collecting Linguistic Resources for Assessing Children's Pronunciation of Nordic LanguagesabstractThis paper reports on the experience collecting a number of corpora of Nordic languages spoken by children. The aim of the data collection is providing annotated data to develop and evaluate computer assisted pronunciation assessment systems both for non-native children learning a Nordic language (L2) and for L1 children with speech sound disorder (SSD). The paper presents the challenges encountered recording and annotating data for Finnish, Swedish and Norwegian, as well as the ethical considerations related with making this data publicly available. We hope that sharing this experience will encourage others to collect similar data for other languages. Of the different data collections, we were able to make the Norwegian corpus publicly available in the hope that it will serve as a reference in pronunciation assessment research. Anne Marte Haug Olstad, Anna-Riikka Smolander, Sofia Strömbergsson, Sari Ylinen, Minna Lehtonen, Mikko Kurimo, Yaroslav Getman, Tamás Grósz, Xinwei Cao, Torbjørn Svendsen, Giampiero Salvi |
LREC/COLING | 6 |
| 2024 | Investigating the Clusters Discovered By Pre-Trained AV-HuBERTabstractSelf-supervised models, such as HuBERT and its audio-visual version AV-HuBERT, have demonstrated excellent performance on various tasks. The main factor for their success is the pre-training procedure, which requires only raw data without human transcription. During the self-supervised pre-training phase, HuBERT is trained to discover latent clusters in the training data, but these clusters are discarded, and only the last hidden layer is used by the conventional finetuning step. We investigate what latent information the AV-HuBERT model managed to uncover via its clusters and can we use them directly for speech recognition. To achieve this, we consider the sequence of cluster ids as a ’language’ developed by the AV-HuBERT and attempt to translate it to English text via small LSTM-based models. These translation models enable us to investigate the relations between the clusters and the English alphabet, shedding light on groups of latent clusters specialized to recognise specific phonetic groups. Our results demonstrate that using the pre-trained system as a quantizer, we are able to compress the video to as low as 275 bit/sec while maintaining acceptable speech recognition accuracy. Furthermore, compared to the conventional finetuning step, our solution has considerably lower computational cost. Anja Virkkunen, Marek Sarvas, Guangpu Huang, Tamás Grósz, Mikko Kurimo |
ICASSP | 5 |
| 2024 | Exploring adaptation techniques of large speech foundation models for low-resource ASR: a case study on Northern SámiabstractPublisher Copyright: © 2024 International Speech Communication Association. All rights reserved. Yaroslav Getman, Tamás Grósz, Katri Hiovain-Asikainen, Mikko Kurimo |
INTERSPEECH | 4 |
| 2024 | What happens in continued pre-training? Analysis of self-supervised speech models with continued pre-training for colloquial Finnish ASRabstractPublisher Copyright: © 2024 International Speech Communication Association. All rights reserved. Yaroslav Getman, Tamás Grósz, Mikko Kurimo |
INTERSPEECH | 3 |
| 2024 | Oversampling, Augmentation and Curriculum Learning for Speaking Assessment with Limited Training DataabstractPublisher Copyright: © 2024 International Speech Communication Association. All rights reserved. Tin Mei Lun, Ekaterina Voskoboinik, Ragheb Al-Ghezi, Tamás Grósz, Mikko Kurimo |
INTERSPEECH | 5 |
| 2024 | CaptainA self-study mobile app for practising speaking: task completion assessment and feedback with generative AI
Nhan Phan, Anna von Zansen, Maria Kautonen, Tamás Grósz, Mikko Kurimo |
INTERSPEECH | 5 |
| 2024 | Automated content assessment and feedback for Finnish L2 learners in a picture description speaking taskabstractWe propose a framework to address several unsolved challenges in second language (L2) automatic speaking assessment (ASA) and feedback. The challenges include: 1. ASA of visual task completion, 2. automated content grading and explanation of spontaneous L2 speech, 3. corrective feedback generation for L2 learners, and 4. all the above for a language that has minimal speech data of L2 learners. The proposed solution combines visual natural language generation (NLG), automatic speech recognition (ASR) and prompting a large language model (LLM) for low-resource L2 learners. We describe the solution and the outcomes of our case study for a picture description task in Finnish. Our results indicate substantial agreement with human experts in grading, explanation and feedback. This framework has the potential for a significant impact in constructing next-generation computer-assisted language learning systems to provide automatic scoring with feedback for learners of low-resource languages. Nhan Phan, Anna von Zansen, Maria Kautonen, Ekaterina Voskoboinik, Tamás Grósz, Raili Hildén, Mikko Kurimo |
INTERSPEECH | 7 |
| 2024 | Out-of-distribution generalisation in spoken language understandingabstractTest data is said to be out-of-distribution (OOD) when it unex- pectedly differs from the training data, a common challenge in real-world use cases of machine learning. Although OOD gen- eralisation has gained interest in recent years, few works have focused on OOD generalisation in spoken language understand- ing (SLU) tasks. To facilitate research on this topic, we intro- duce a modified version of the popular SLU dataset SLURP, featuring data splits for testing OOD generalisation in the SLU task. We call our modified dataset SLURP For OOD gener- alisation, or SLURPFOOD. Utilising our OOD data splits, we find end-to-end SLU models to have limited capacity for gen- eralisation. Furthermore, by employing model interpretability techniques, we shed light on the factors contributing to the gen- eralisation difficulties of the models. To improve the generali- sation, we experiment with two techniques, which improve the results on some, but not all the splits, emphasising the need for new techniques. Dejan Porjazovski, Anssi Moisio, Mikko Kurimo |
INTERSPEECH | 3 |
| 2024 | Spectral warping based data augmentation for low resource children's speaker verificationabstractAbstract In this paper, we present our effort to develop an automatic speaker verification (ASV) system for low resources children’s data. For the children’s speakers, very limited amount of speech data is available in majority of the languages for training the ASV system. Developing an ASV system under low resource conditions is a very challenging problem. To develop the robust baseline system, we merged out of domain adults’ data with children’s data to train the ASV system and tested with children’s speech. This kind of system leads to acoustic mismatches between training and testing data. To overcome this issue, we have proposed spectral warping based data augmentation. We modified adult speech data using spectral warping method (to simulate like children’s speech) and added it to the training data to overcome data scarcity and mismatch between adults’ and children’s speech. The proposed data augmentation gives 20.46% and 52.52% relative improvement (in equal error rate) for Indian Punjabi and British English speech databases, respectively. We compared our proposed method with well known data augmentation methods: SpecAugment, speed perturbation (SP) and vocal tract length perturbation (VTLP), and found that the proposed method performed best. The proposed spectral warping method is publicly available at https://github.com/kathania/Speaker-Verification-spectral-warping . Hemant Kumar Kathania, Virender Kadyan, Sudarsana Reddy Kadiri, Mikko Kurimo |
Multim. Tools Appl. | 4 |
| 2024 | Comparison and analysis of new curriculum criteria for end-to-end ASRabstractTraditionally, teaching a human and a Machine Learning (ML) model is quite different, but organized and structured learning has the ability to enable faster and better understanding of the underlying concepts. For example, when humans learn to speak, they first learn how to utter basic phones and then slowly move towards more complex structures such as words and sentences. Motivated by this observation, researchers have started to adapt this approach for training ML models. Since the main concept, the gradual increase in difficulty, resembles the notion of the curriculum in education, the methodology became known as Curriculum Learning (CL). In this work, we design and test new CL approaches to train Automatic Speech Recognition systems, specifically focusing on the so-called end-to-end models. These models consist of a single, large-scale neural network that performs the recognition task, in contrast to the traditional way of having several specialized components focusing on different subtasks (e.g., acoustic and language modeling). We demonstrate that end-to-end models can achieve better performances if they are provided with an organized training set consisting of examples that exhibit an increasing level of difficulty. To impose structure on the training set and to define the notion of an easy example, we explored multiple solutions that use either external, static scoring methods or incorporate feedback from the model itself. In addition, we examined the effect of pacing functions that control how much data is presented to the network during each training epoch. Our proposed curriculum learning strategies were tested on the task of speech recognition on two data sets, one containing spontaneous Finnish speech where volunteers were asked to speak about a given topic, and one containing planned English speech. Empirical results showed that a good curriculum strategy can yield performance improvements and speed-up convergence. After a given number of epochs, our best strategy achieved a 5.6% and 3.4% decrease in terms of test set word error rate for the Finnish and English data sets, respectively. Georgios Karakasidis, Mikko Kurimo, Peter Bell 0001, Tamás Grósz |
Speech Commun. | 2 |
| 2024 | From Raw Speech to Fixed Representations: A Comprehensive Evaluation of Speech Embedding TechniquesabstractSpeech embeddings, fixed-size representations derived from raw audio data, play a crucial role in diverse machine learning applications. Despite the abundance of speech embedding techniques, selecting the most suitable one remains challenging. Existing studies often focus on intrinsic or extrinsic aspects, seldom exploring both simultaneously. Furthermore, comparing the state-of-the-art pre-trained models with prior speech embedding solutions is notably scarce in the literature. To address these gaps, we undertake a comprehensive evaluation of both small and large-scale speech embedding models, which, in our opinion, needs to incorporate both intrinsic and extrinsic assessments. The intrinsic experiments delve into the models' ability to pick speaker-related characteristics and assess their discriminative capacities, providing insights into their inherent capabilities and internal workings. Concurrently, the extrinsic experiments evaluate whether the models learned semantic cues during pre-training. The findings underscore the superior performance of the large-scale pre-trained models, albeit at an elevated computational cost. The base self-supervised models show comparable results to their large counterparts, making them a better choice for many applications. Furthermore, we show that by selecting the most crucial dimensions, the models' performance often does not suffer drastically and even improves in some cases. This research contributes valuable insights into the nuanced landscape of speech embeddings, aiding researchers and practitioners in making informed choices for various applications. Dejan Porjazovski, Tamás Grósz, Mikko Kurimo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Principled Comparisons for End-to-End Speech Recognition: Attention vs Hybrid at the 1000-Hour ScaleabstractEnd-to-End speech recognition has become the center of attention for speech recognition research, but Hybrid Hidden Markov Model Deep Neural Network (HMM/DNN) -systems remain a competitive approach in terms of performance. End-to-End models may be better at very large data scales, and HMM / DNN-systems may have an advantage in low-resource scenarios, but the thousand-hour scale is particularly interesting for comparisons. At that scale experiments have not been able to conclusively demonstrate which approach is best, or if the heterogeneous approaches yield similar results. In this work, we work towards answering that question for Attention-based Encoder-Decoder models compared with HMM / DNN-systems. We present two simple experimental design principles, and how to build systems adhering to those principles. We demonstrate how those principles remove confounding variables related to both data, and neural architecture and training. We apply the principles in a set of experiments on three diverse thousand-hour-scale tasks. In our experiments, the HMM / DNN-systems yield equal or better results in almost all cases. Aku Rouhe, Tamás Grósz, Mikko Kurimo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Investigating wav2vec2 context representations and the effects of fine-tuning, a case-study of a Finnish modelabstractSelf-supervised speech models, such as the wav2vec2, have become extremely popular in the past few years. Their main appeal is that after their pre-training on a large amount of audio, they require only a small amount of supervised, finetuning data to achieve outstanding results. Despite their immense success, very little is understood about the pre-trained models and how finetuning changes them. In this work, we take the first steps towards a better understanding of wav2vec2 systems using model interpretation tools such as visualization and latent embedding clustering. Through our analysis, we gain new insights into the abilities of the pre-trained networks and the effect that finetuning has on them. We demonstrate that the clusters learned by the pre-trained model are just as important a factor as the supervised training data distribution in determining the accuracy of the finetuned system, which could aid us in selecting the most suitable pre-trained model for the supervised data. Tamás Grósz, Yaroslav Getman, Ragheb Al-Ghezi, Aku Rouhe, Mikko Kurimo |
INTERSPEECH | 5 |
| 2023 | Advancing Audio Emotion and Intent Recognition with Large Pre-Trained Models and Bayesian InferenceabstractLarge pre-trained models are essential in paralinguistic systems, demonstrating effectiveness in tasks like emotion recognition and stuttering detection. In this paper, we employ large pre-trained models for the ACM Multimedia Computational Paralinguistics Challenge, addressing the Requests and Emotion Share tasks. We explore audio-only and hybrid solutions leveraging audio and text modalities. Our empirical results consistently show the superiority of the hybrid approaches over the audio-only models. Moreover, we introduce a Bayesian layer as an alternative to the standard linear output layer. The multimodal fusion approach achieves an 85.4% UAR on HC-Requests and 60.2% on HC-Complaints. The ensemble model for the Emotion Share task yields the best ρ value of .614. The Bayesian wav2vec2 approach, explored in this study, allows us to easily build ensembles, at the cost of fine-tuning only one model. Moreover, we can have usable confidence values instead of the usual overconfident posterior probabilities. Dejan Porjazovski, Yaroslav Getman, Tamás Grósz, Mikko Kurimo |
ACM Multimedia | 4 |
| 2022 | When to Laugh and How Hard? A Multimodal Approach to Detecting Humor and Its IntensityabstractPrerecorded laughter accompanying dialog in comedy TV shows encourages the audience to laugh by clearly marking humorous moments in the show. We present an approach for automatically detecting humor in the Friends TV show using multimodal data. Our model is capable of recognizing whether an utterance is humorous or not and assess the intensity of it. We use the prerecorded laughter in the show as annotation as it marks humor and the length of the audience’s laughter tells us how funny a given joke is. We evaluate the model on episodes the model has not been exposed to during the training phase. Our results show that the model is capable of correctly detecting whether an utterance is humorous 78% of the time and how long the audience’s laughter reaction should last with a mean absolute error of 600 milliseconds. Khalid Al-Najjar, Mika Hämäläinen, Jörg Tiedemann, Jorma Laaksonen, Mikko Kurimo |
COLING | 5 |
| 2022 | wav2vec2-based Speech Rating System for Children with Speech Sound DisorderabstractThe computational resources were provided by Aalto ScienceIT. This work was supported by NordForsk through the funding to Technology-enhanced foreign and second-language learning of Nordic languages, project number 103893. Yaroslav Getman, Ragheb Al-Ghezi, Katja Voskoboinik, Tamás Grósz, Mikko Kurimo, Giampiero Salvi, Torbjørn Svendsen, Sofia Strömbergsson |
INTERSPEECH | 5 |
| 2022 | Comparison and Analysis of New Curriculum Criteria for End-to-End ASRabstractThe computational resources were provided by Aalto ScienceIT. We are grateful for the Academy of Finland project funding number 345790 in ICT 2023 programme's project”Understanding speech and scene with ears and eyes” Georgios Karakasidis, Tamás Grósz, Mikko Kurimo |
INTERSPEECH | 3 |
| 2022 | Low Resource Comparison of Attention-based and Hybrid ASR Exploiting wav2vec 2.0abstractFunding Information: We are grateful for the Academy of Finland project funding, numbers: 337073, 345790. We acknowledge the computational resources provided by the Aalto Science-IT project. Publisher Copyright: Copyright © 2022 ISCA. Aku Rouhe, Anja Virkkunen, Juho Leinonen 0002, Mikko Kurimo |
INTERSPEECH | 4 |
| 2022 | Wav2vec2-based Paralinguistic Systems to Recognise Vocalised Emotions and StutteringabstractWith the rapid advancement in automatic speech recognition and natural language understanding, a complementary field (paralinguistics) emerged, focusing on the non-verbal content of speech. The ACM Multimedia 2022 Computational Paralinguistics Challenge introduced several exciting tasks of this field. In this work, we focus on tackling two Sub-Challenges using modern, pre-trained models called wav2vec2. Our experimental results demonstrated that wav2vec2 is an excellent tool for detecting the emotions behind vocalisations and recognising different types of stutterings. Albeit they achieve outstanding results on their own, our results demonstrated that wav2vec2-based systems could be further improved by ensembling them with other models. Our best systems outperformed the competition baselines by a considerable margin, achieving an unweighted average recall of 44.0 (absolute improvement of 6.6% over baseline) on the Vocalisation Sub-Challenge and 62.1 (absolute improvement of 21.7% over baseline) on the Stuttering Sub-Challenge. Tamás Grósz, Dejan Porjazovski, Yaroslav Getman, Sudarsana Reddy Kadiri, Mikko Kurimo |
ACM Multimedia | 5 |
| 2022 | Automatic Rating of Spontaneous Speech for Low-Resource LanguagesabstractAutomatic spontaneous speaking assessment systems bring numerous advantages to second language (L2) learning and assessment such as promoting self-learning and reducing language teachers' workload. Conventionally, these systems are developed for languages with a large number of learners due to the abundance of training data, yet languages with fewer learners such as Finnish and Swedish remain at a disadvantage due to the scarcity of required training data. Nevertheless, recent advancements in self-supervised deep learning make it possible to develop automatic speech recognition systems with a reasonable amount of training data. In turn, this advancement makes it feasible to develop systems for automatically assessing spoken proficiency of learners of underresourced languages: L2 Finnish and Finland Swedish. Our work evaluates the overall performance of the L2 ASR systems as well as the the rating systems compared to human reference ratings for both languages. Ragheb Al-Ghezi, Yaroslav Getman, Ekaterina Voskoboinik, Mittul Singh, Mikko Kurimo |
SLT | 5 |
| 2022 | A formant modification method for improved ASR of children's speechabstractDifferences in acoustic characteristics between children’s and adults’ speech degrade performance of automatic speech recognition systems when systems trained using adults’ speech are used to recognize children’s speech. This performance degradation is due to the acoustic mismatch between training and testing. One of the main sources of the acoustic mismatch is the difference in vocal tract resonances (formant frequencies) between adult and child speakers. The present study aims to reduce the mismatch in formant frequencies by modifying formants of children’s speech to better correspond to formants of adults’ speech. This is carried out by warping the linear prediction (LP) spectrum computed from children’s speech. The warped LP spectra computed in a frame-based manner from children’s speech are used with the corresponding LP residuals to synthesize speech whose formant structure is closer to that of adults’ speech. When used in testing of an ASR system trained using adults’ speech, the warping reduces the spectral mismatch in speech between training and testing and improves the system performance in recognition of children’s speech. Experiments were conducted using narrowband (8 kHz) and wideband (16 kHz) speech of adult and child speakers from the WSJCAM0 and PF_STAR databases, respectively, and by recognizing children’s speech using acoustic models trained with adults’ speech. The proposed method gave relative improvements of 24% and 11% for the DNN and TDNN acoustic models, respectively, for narrowband speech. For wideband speech, the technique gave relative improvements of 27% and 13% for the DNN and TDNN acoustic models, respectively. The performance of the proposed method was also compared to two speaker adaptation methods: vocal tract length normalization (VTLN) and speaking rate adaptation (SRA). This comparison showed the best recognition performance for the proposed method. We also combined the proposed method with VTLN and SRA, and found that the combined method gave a further reduction in WER. Moreover, our experiments carried out for noisy speech using various types of additive noise and signal-to-noise ratios showed that the proposed method performs well also for degraded speech. Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Paavo Alku, Mikko Kurimo |
Speech Commun. | 4 |
| 2021 | Vowel Non-Vowel Based Spectral Warping and Time Scale Modification for Improvement in Children's ASRabstractAcoustic differences between children’s and adults’ speech causes the degradation in the automatic speech recognition system performance when system trained on adults’ speech and tested on children’s speech. The key acoustic mismatch factors are formant, speaking rate, and pitch. In this paper, we proposed a linear prediction based spectral warping method by using the knowledge of vowel and non-vowel regions in speech signals to mitigate the formant frequencies differences between child and adult speakers. The proposed method gives 31% relative improvement over the baseline system. We have also investigated time scale modification using RTISILA and SOLAFS algorithms and found that our proposed method performs better. Combining the proposed method with RTISILA and SOLAFS results in a further error rate reduction. The final combined system gives 49% relative improvement compared to the baseline system. Hemant Kumar Kathania, Mikko Kurimo |
ICASSP | 3 |
| 2021 | Self-Supervised End-to-End ASR for Low Resource L2 SwedishabstractFunding Information: This work is part of Digitala project which is funded by the Academy of Finland (grant numbers 322619, 322625, 322965). The computational resources were provided by Aalto ScienceIT. Funding Information: This work is part of Digitala project which is funded by the Academy of Finland (grant numbers 322619, 322625, 322965). The computational resources were provided by Aalto Scien-ceIT. Publisher Copyright: Copyright © 2021 ISCA. Ragheb Al-Ghezi, Yaroslav Getman, Aku Rouhe, Raili Hildén, Mikko Kurimo |
Interspeech | 5 |
| 2021 | Advances in subword-based HMM-DNN speech recognition across languagesabstractWe describe a novel way to implement subword language models in speech recognition systems based on weighted finite state transducers, hidden Markov models, and deep neural networks. The acoustic models are built on graphemes in a way that no pronunciation dictionaries are needed, and they can be used together with any type of subword language model, including character models. The advantages of short subword units are good lexical coverage, reduced data sparsity, and avoiding vocabulary mismatches in adaptation. Moreover, constructing neural network language models (NNLMs) is more practical, because the input and output layers are small. We also propose methods for combining the benefits of different types of language model units by reconstructing and combining the recognition lattices. We present an extensive evaluation of various subword units on speech datasets of four languages: Finnish, Swedish, Arabic, and English. The results show that the benefits of short subwords are even more consistent with NNLMs than with traditional n-gram language models. Combination across different acoustic models and language models with various units improve the results further. For all the four datasets we obtain the best results published so far. Our approach performs well even for English, where the phoneme-based acoustic models and word-based language models typically dominate: The phoneme-based baseline performance can be reached and improved by 4% using graphemes only when several grapheme-based models are combined. Furthermore, combining both grapheme and phoneme models yields the state-of-the-art error rate of 15.9% for the MGB 2018 dev17b test. For all four languages we also show that the language models perform reasonably well when only limited training data is available. Peter Smit, Sami Virpioja, Mikko Kurimo |
Comput. Speech Lang. | 3 |
| 2021 | Morphologically motivated word classes for very large vocabulary speech recognition of Finnish and Estonian
Matti Varjokallio, Sami Virpioja, Mikko Kurimo |
Comput. Speech Lang. | 3 |
| 2020 | Study of Formant Modification for Children ASRabstractThe performance of automatic speech recognition systems for children’s speech is known to suffer from the large variation and mismatch in the acoustic and linguistic attributes between children’s and adults’ speech. One of the various identified sources of mismatch is the difference in formant frequencies between adults and children. In this paper, we propose a formant modification method to mitigate differences between adults’ and children’s speech and to improve the performance of ASR for children. The explored technique gives a relative 27% improvement in system performance compared to a hybrid DNN-HMM baseline. We also compare the system performance with related speaker adaptation methods like vocal tract length normalization (VTLN) and speaking rate adaptation (SRA) and find that the proposed method gives improvements over them, as well. Combining the proposed method with VTLN and SRA results in a further reduction of WER. We also found that the proposed method performs well even for noisy speech. Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Paavo Alku, Mikko Kurimo |
ICASSP | 4 |
| 2020 | Speaker-Aware Training of Attention-Based End-to-End Speech Recognition Using Neural Speaker EmbeddingsabstractIn speaker-aware training, a speaker embedding is appended to DNN input features. This allows the DNN to effectively learn representations, which are robust to speaker variability.We apply speaker-aware training to attention-based end-to-end speech recognition. We show that it can improve over a purely end-to-end baseline. We also propose speaker-aware training as a viable method to leverage untranscribed, speaker annotated data.We apply state-of-the-art embedding approaches, both i-vectors and neural embeddings, such as x-vectors. We experiment with embeddings trained in two conditions: on the fixed ASR data, and on a large untranscribed dataset. We run our experiments on the TED-LIUM and Wall Street Journal datasets. No embedding consistently outperforms all others, but in many settings neural embeddings outperform i-vectors. Aku Rouhe, Tuomas Kaseva, Mikko Kurimo |
ICASSP | 3 |
| 2020 | Finnish ASR with Deep Transformer Modelsabstract| openaire: EC/H2020/780069/EU//MeMAD Abhilash Jain, Aku Rouhe, Stig-Arne Grönroos, Mikko Kurimo |
INTERSPEECH | 4 |
| 2020 | Data Augmentation Using Prosody and False Starts to Recognize Non-Native Children's SpeechabstractThis paper describes AaltoASR's speech recognition system for the INTERSPEECH 2020 shared task on Automatic Speech Recognition (ASR) for non-native children's speech. The task is to recognize non-native speech from children of various age groups given a limited amount of speech. Moreover, the speech being spontaneous has false starts transcribed as partial words, which in the test transcriptions leads to unseen partial words. To cope with these two challenges, we investigate a data augmentation-based approach. Firstly, we apply the prosody-based data augmentation to supplement the audio data. Secondly, we simulate false starts by introducing partial-word noise in the language modeling corpora creating new words. Acoustic models trained on prosody-based augmented data outperform the models using the baseline recipe or the SpecAugment-based augmentation. The partial-word noise also helps to improve the baseline language model. Our ASR system, a combination of these schemes, is placed third in the evaluation period and achieves the word error rate of 18.71%. Post-evaluation period, we observe that increasing the amounts of prosody-based augmented data leads to better performance. Furthermore, removing low-confidence-score words from hypotheses can lead to further gains. These two improvements lower the ASR error rate to 17.99%. Hemant Kumar Kathania, Mittul Singh, Tamás Grósz, Mikko Kurimo |
INTERSPEECH | 4 |
| 2020 | FinChat: Corpus and Evaluation Setup for Finnish Chat Conversations on Everyday TopicsabstractCreating open-domain chatbots requires large amounts of conversational data and related benchmark tasks to evaluate them. Standardized evaluation tasks are crucial for creating automatic evaluation metrics for model development; otherwise, comparing the models would require resource-expensive human evaluation. While chatbot challenges have recently managed to provide a plethora of such resources for English, resources in other languages are not yet available. In this work, we provide a starting point for Finnish open-domain chatbot research. We describe our collection efforts to create the Finnish chat conversation corpus FinChat, which is made available publicly. FinChat includes unscripted conversations on seven topics from people of different ages. Using this corpus, we also construct a retrieval-based evaluation task for Finnish chatbot development. We observe that off-the-shelf chatbot models trained on conversational corpora do not perform better than chance at choosing the right answer based on automatic metrics, while humans can do the same task almost perfectly. Similarly, in a human evaluation, responses to questions from the evaluation set generated by the chatbots are predominantly marked as incoherent. Thus, FinChat provides a challenging evaluation set, meant to encourage chatbot development in Finnish. Katri Leino, Juho Leinonen 0002, Mittul Singh, Sami Virpioja, Mikko Kurimo |
INTERSPEECH | 5 |
| 2020 | Releasing a Toolkit and Comparing the Performance of Language Embeddings Across Various Spoken Language Identification DatasetsabstractPeer reviewed Matias Lindgren, Tommi Jauhiainen, Mikko Kurimo |
INTERSPEECH | 3 |
| 2020 | Morfessor EM+Prune: Improved Subword Segmentation with Expectation Maximization and PruningabstractData-driven segmentation of words into subword units has been used in various natural language processing applications such as automatic speech recognition and statistical machine translation for almost 20 years. Recently it has became more widely adopted, as models based on deep neural networks often benefit from subword units even for morphologically simpler languages. In this paper, we discuss and compare training algorithms for a unigram subword model, based on the Expectation Maximization algorithm and lexicon pruning. Using English, Finnish, North Sami, and Turkish data sets, we show that this approach is able to find better solutions to the optimization problem defined by the Morfessor Baseline model than its original recursive training algorithm. The improved optimization also leads to higher morphological segmentation accuracy when compared to a linguistic gold standard. We publish implementations of the new algorithms in the widely-used Morfessor software package. Stig-Arne Grönroos, Sami Virpioja, Mikko Kurimo |
LREC | 3 |
| 2020 | Transfer learning and subword sampling for asymmetric-resource one-to-many neural translation
Stig-Arne Grönroos, Sami Virpioja, Mikko Kurimo |
Mach. Transl. | 3 |
| 2019 | Spherediar: An Effective Speaker Diarization System for Meeting DataabstractIn this paper, we present SphereDiar, a speaker diarization system composed of three novel subsystems: the Sphere-Speaker (SS) neural network, designed for speaker embedding extraction, a segmentation method called Homogeneity Based Segmentation (HBS) and a clustering algorithm called Top Two Silhouettes (Top2S). The system is evaluated on a set of over 200 manually transcribed multiparty meetings. The evaluation reveals that the system can be further simplified by omitting the use of HBS. Furthermore, we illustrate that SphereDiar achieves state-of-the-art results with two different meeting data sets. Tuomas Kaseva, Aku Rouhe, Mikko Kurimo |
ASRU | 3 |
| 2019 | Transparent Pronunciation Scoring Using Articulatorily Weighted Phoneme Edit DistanceabstractFor researching effects of gamification in foreign language learning for children in the "Say It Again, Kid!" project we developed a feedback paradigm that can drive gameplay in pronunciation learning games. We describe our scoring system based on the difference between a reference phone sequence and the output of a multilingual CTC phoneme recogniser. We present a white-box scoring model of mapped weighted Levenshtein edit distance between reference and error with error weights for articulatory differences computed from a training set of scored utterances. The system can produce a human-readable list of each detected mispronunciation's contribution to the utterance score. We compare our scoring method to established black box methods. Reima Karhila, Anna-Riikka Smolander, Sari Ylinen, Mikko Kurimo |
INTERSPEECH | 4 |
| 2019 | Subword RNNLM Approximations for Out-Of-Vocabulary Keyword SearchabstractIn spoken Keyword Search, the query may contain out-of-vocabulary (OOV) words not observed when training the speech recognition system. Using subword language models (LMs) in the first-pass recognition makes it possible to recognize the OOV words, but even the subword n-gram LMs suffer from data sparsity. Recurrent Neural Network (RNN) LMs alleviate the sparsity problems but are not suitable for first-pass recognition as such. One way to solve this is to approximate the RNNLMs by back-off n-gram models. In this paper, we propose to interpolate the conventional n-gram models and the RNNLM approximation for better OOV recognition. Furthermore, we develop a new RNNLM approximation method suitable for subword units: It produces variable-order n-grams to include long-span approximations and considers also n-grams that were not originally observed in the training corpus. To evaluate these models on OOVs, we setup Arabic and Finnish Keyword Search tasks concentrating only on OOV words. On these tasks, interpolating the baseline RNNLM approximation and a conventional LM outperforms the conventional LM in terms of the Maximum Term Weighted Value for single-character subwords. Moreover, replacing the baseline approximation with the proposed method achieves the best performance on both multi- and single-character subwords. Mittul Singh, Sami Virpioja, Peter Smit, Mikko Kurimo |
INTERSPEECH | 4 |
| 2019 | RL-KLM: automating keystroke-level modeling with reinforcement learningabstractThe Keystroke-Level Model (KLM) is a popular model for predicting users' task completion times with graphical user interfaces. KLM predicts task completion times as a linear function of elementary operators. However, the policy, or the assumed sequence of the operators that the user executes, needs to be prespeciffed by the analyst. This paper investigates Reinforcement Learning (RL) as an algorithmic method to obtain the policy automatically. We define the KLM as an Markov Decision Process, and show that when solved with RL methods, this approach yields user-like policies in simple but realistic interaction tasks. RL-KLM offers a quick way to obtain a global upper bound for user performance. It opens up new possibilities to use KLM in computational interaction. However, scalability and validity remain open issues. Katri Leino, Antti Oulasvirta, Mikko Kurimo |
IUI | 3 |
| 2018 | Captaina: Integrated Pronunciation Practice and Data Collection Portal
Aku Rouhe, Reima Karhila, Aija Elg, Minnaleena Toivola, Peter Smit, Anna-Riikka Smolander, Mikko Kurimo |
INTERSPEECH | 7 |
| 2018 | First-Pass Techniques for Very Large Vocabulary Speech Recognition ff Morphologically Rich LanguagesabstractIn speech recognition of morphologically rich languages, very large vocabulary sizes are required to achieve good error rates. Especially traditional n-gram language models trained over word sequences suffer from data sparsity issues. The language modelling can often be improved by segmenting the words to sequences of subword units that are more frequent. Another solution is to cluster the words into classes and apply a class-based language model. We show that linearly interpolating n-gram models trained over words, subwords, and word classes improves the first-pass speech recognition accuracy in very large vocabulary speech recognition tasks for two morphologically rich and agglutinative languages, Finnish and Estonian. To overcome performance issues, we also introduce a novel language model look-ahead method utilizing a class bigram model. The method improves the results over a unigram look-ahead model with the same recognition speed, the difference increasing for small real-time factors. The improved model combination and look-ahead model are useful in cases where real-time recognition is required or when the improved hypotheses help with further recognition passes. For instance, neural network language models are mostly applied by rescoring the generated hypotheses due to higher computational costs. Matti Varjokallio, Sami Virpioja, Mikko Kurimo |
SLT | 3 |
| 2017 | Character-based units for unlimited vocabulary continuous speech recognitionabstractWe study character-based language models in the state-of-the-art speech recognition framework. This approach has advantages over both word-based systems and so-called end-to-end ASR systems that do not have separate acoustic and language models. We describe the necessary modifications needed to build an effective character-based ASR system using the Kaldi toolkit and evaluate the models based on words, statistical morphs, and characters for both Finnish and Arabic. The morph-based models yield the best recognition results for both well-resourced and lower-resourced tasks, but the character-based models are close to their performance in the lower-resource tasks, outperforming the word-based models. Character-based models are especially good at predicting novel word forms that were not seen in the training data. Using character-based neural network language models is both computationally efficient and provides a larger gain compared to the morph and word-based systems. Peter Smit, Siva Reddy Gangireddy, Seppo Enarvi, Sami Virpioja, Mikko Kurimo |
ASRU | 5 |
| 2017 | Aalto system for the 2017 Arabic multi-genre broadcast challengeabstractWe describe the speech recognition systems we have created for MGB-3, the 3rdMulti Genre Broadcast challenge, which this year consisted of a task of building a system for transcribing Egyptian Dialect Arabic speech, using a big audio corpus of primarily Modern Standard Arabic speech and only a small amount (5 hours) of Egyptian adaptation data. Our system, which was a combination of different acoustic models, language models and lexical units, achieved a Multi-Reference Word Error Rate of 29.25%, which was the lowest in the competition. Also on the old MGB-2 task, which was run again to indicate progress, we achieved the lowest error rate: 13.2%. The result is a combination of the application of state-of-the-art speech recognition methods such as simple dialect adaptation for a Time-Delay Neural Network (TDNN) acoustic model (−27% errors compared to the baseline), Recurrent Neural Network Language Model (RNNLM) rescoring (an additional −5%), and system combination with Minimum Bayes Risk (MBR) decoding (yet another −10%). We also explored the use of morph and character language models, which was particularly beneficial in providing a rich pool of systems for the MBR decoding. Peter Smit, Siva Reddy Gangireddy, Seppo Enarvi, Sami Virpioja, Mikko Kurimo |
ASRU | 5 |
| 2017 | ASR in Classroom Today: Automatic Visualization of Conceptual Network in Science Classrooms
Daniela Caballero, Roberto Araya, Hanna Kronholm, Jouni Viiri, André Mansikkaniemi, Sami Lehesvuori, Tuomas Virtanen, Mikko Kurimo |
EC-TEL | 8 |
| 2017 | LDA-based context dependent recurrent neural network language model using document-based topic distribution of wordsabstractAdding context information into recurrent neural network language models (RNNLMs) have been investigated recently to improve the effectiveness of learning RNNLM. Conventionally, a fast approximate topic representation for a block of words was proposed by using corpus-based topic distribution of word incorporating latent Dirichlet allocation (LDA) model. It is then updated for each subsequent word using an exponential decay. However, words could represent different topics in different documents. In this paper, we form document-based distribution over topics for each word using LDA model and apply it in the computation of fast approximate exponentially decaying features. We have shown experimental results on a well known Penn Treebank corpus and found that our approach outperforms the conventional LDA-based context RNNLM approach. Moreover, we carried out speech recognition experiments on Wall Street Journal corpus and achieved word error rate (WER) improvements over the other approach. Md. Akmal Haidar, Mikko Kurimo |
ICASSP | 2 |
| 2017 | SIAK - A Game for Foreign Language Pronunciation Learning
Reima Karhila, Sari Ylinen, Seppo Enarvi, Kalle J. Palomäki, Aleksander Nikulin, Olli Rantula, Vertti Viitanen, Krupakar Dhinakaran, Anna-Riikka Smolander, Heini Kallio, Katja Junttila, Maria Uther, Perttu Hämäläinen, Mikko Kurimo |
INTERSPEECH | 14 |
| 2017 | Automatic Construction of the Finnish Parliament Speech CorpusabstractAutomatic speech recognition (ASR) systems require large amounts of transcribed speech data, for training state-of-the-art deep neural network (DNN) acoustic models. Transcribed speech is a scarce and expensive resource, and ASR systems are prone to underperform in domains where there is not a lot of training data available. In this work, we open up a vast and previously unused resource of transcribed speech for Finnish, by retrieving and aligning all the recordings and meeting transcripts from the web portal of the Parliament of Finland. Short speech-text segment pairs are retrieved from the audio and text material, by using the Levenshtein algorithm to align the first-pass ASR hypotheses with the corresponding meeting transcripts. DNN acoustic models are trained on the automatically constructed corpus, and performance is compared to other models trained on a commercially available speech corpus. Model performance is evaluated on Finnish parliament speech, by dividing the testing set into seen and unseen speakers. Performance is also evaluated on broadcast speech to test the general applicability of the parliament speech corpus. We also study the use of meeting transcripts in language model adaptation, to achieve additional gains in speech recognition accuracy of Finnish parliament speech. André Mansikkaniemi, Peter Smit, Mikko Kurimo |
INTERSPEECH | 3 |
| 2017 | Reading Validation for Pronunciation Evaluation in the Digitala Project
Aku Rouhe, Reima Karhila, Peter Smit, Mikko Kurimo |
INTERSPEECH | 4 |
| 2017 | Improved Subword Modeling for WFST-Based Speech RecognitionabstractBecause in agglutinative languages the number of observed word forms is very high, subword units are often utilized in speech recognition. However, the proper use of subword units requires careful consideration of details such as silence modeling, position-dependent phones, and combination of the units. In this paper, we implement subword modeling in the Kaldi toolkit by creating modified lexicon by finite-state transducers to represent the subword units correctly. We experiment with multiple types of word boundary markers and achieve the best results by adding a marker to the left or right side of a subword unit whenever it is not preceded or followed by a word boundary, respectively. We also compare three different toolkits that provide data-driven subword segmentations. In our experiments on a variety of Finnish and Estonian datasets, the best subword models do outperform word-based models and naive subword implementations. The largest relative reduction in WER is a 23% over word-based models for a Finnish read speech dataset. The results are also better than any previously published ones for the same datasets, and the improvement on all datasets is more than 5%. Peter Smit, Sami Virpioja, Mikko Kurimo |
INTERSPEECH | 3 |
| 2017 | Automatic Speech Recognition With Very Large Conversational Finnish and Estonian VocabulariesabstractToday, the vocabulary size for language models in large vocabulary speech recognition is typically several hundreds of thousands of words. While this is already sufficient in some applications, the out-of-vocabulary words are still limiting the usability in others. In agglutinative languages the vocabulary for conversational speech should include millions of word forms to cover the spelling variations due to colloquial pronunciations, in addition to the word compounding and inflections. Very large vocabularies are also needed, for example, when the recognition of rare proper names is important. Previously, very large vocabularies have been efficiently modeled in conventional n-gram language models either by splitting words into subword units or by clustering words into classes. While vocabulary size is not as critical anymore in modern speech recognition systems, training time and memory consumption become an issue when state-of-the-art neural network language models are used. In this paper, we investigate techniques that address the vocabulary size issue by reducing the effective vocabulary size and by processing large vocabularies more efficiently. The experimental results in conversational Finnish and Estonian speech recognition indicate that properly defined word classes improve recognition accuracy. Subword n-gram models are not better on evaluation data than word n-gram models constructed from a vocabulary that includes all the words in the training corpus. However, when recurrent neural network (RNN) language models are used, their ability to utilize long contexts gives a larger gain to subword-based modeling. Our best results are from RNN language models that are based on statistical morphs. We show that the suitable size for a subword vocabulary depends on the language. Using time delay neural network acoustic models, we were able to achieve new state of the art in Finnish and Estonian conversational speech recognition, 27.1% word error rate in the Finnish task and 21.9% in the Estonian task. Seppo Enarvi, Peter Smit, Sami Virpioja, Mikko Kurimo |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2016 | Experiments on Adaptation Methods to Improve Acoustic Modeling for French Speech RecognitionabstractInternational audience Saeideh Mirzaei, Pierrick Milhorat, Jérôme Boudy, Gérard Chollet, Mikko Kurimo |
ICPRAM | 5 |
| 2016 | TheanoLM - An Extensible Toolkit for Neural Network Language ModelingabstractWe present a new tool for training neural network language models (NNLMs), scoring sentences, and generating text. The tool has been written using Python library Theano, which allows researcher to easily extend it and tune any aspect of the training process. Regardless of the flexibility, Theano is able to generate extremely fast native code that can utilize a GPU or multiple CPU cores in order to parallelize the heavy numerical computations. The tool has been evaluated in difficult Finnish and English conversational speech recognition tasks, and significant improvement was obtained over our best back-off n-gram models. The results that we obtained in the Finnish task were compared to those from existing RNNLM and RWTHLM toolkits, and found to be as good or better, while training times were an order of magnitude shorter. Seppo Enarvi, Mikko Kurimo |
INTERSPEECH | 2 |
| 2016 | Recurrent Neural Network Language Model with Incremental Updated Context Information Generated Using Bag-of-Words Representation
Md. Akmal Haidar, Mikko Kurimo |
INTERSPEECH | 2 |
| 2016 | Digitala: An Augmented Test and Review Process Prototype for High-Stakes Spoken Foreign Language Examination
Reima Karhila, Aku Rouhe, Peter Smit, André Mansikkaniemi, Heini Kallio, Erik Lindroos, Raili Hildén, Martti Vainio, Mikko Kurimo |
INTERSPEECH | 9 |
| 2016 | A Comparative Study of Minimally Supervised Morphological SegmentationabstractThis article presents a comparative study of a subfield of morphology learning referred to as minimally supervised morphological segmentation. In morphological segmentation, word forms are segmented into morphs, the surface forms of morphemes. In the minimally supervised data-driven learning setting, segmentation models are learned from a small number of manually annotated word forms and a large set of unannotated word forms. In addition to providing a literature survey on published methods, we present an in-depth empirical comparison on three diverse model families, including a detailed error analysis. Based on the literature survey, we conclude that the existing methodology contains substantial work on generative morph lexicon-based approaches and methods based on discriminative boundary detection. As for which approach has been more successful, both the previous work and the empirical evaluation presented here strongly imply that the current state of the art is yielded by the discriminative boundary detection methodology. Teemu Ruokolainen, Oskar Kohonen, Kairit Sirts, Stig-Arne Grönroos, Mikko Kurimo, Sami Virpioja |
Comput. Linguistics | 5 |
| 2016 | Comparing human and automatic speech recognition in a perceptual restoration experiment
Ulpu Remes, Ana Ramírez López, Lauri Juvela, Kalle J. Palomäki, Guy J. Brown, Paavo Alku, Mikko Kurimo |
Comput. Speech Lang. | 7 |
| 2015 | Designing multichannel source separation based on single-channel source separationabstractIn this paper, an extension of independent vector analysis (IVA), model-based IVA, is proposed for multichannel source separation. For obtaining better source models, we introduce a single-channel source separation method, and utilize the outputs as source variances in time-frequency-variant Gaussian source model. The demixing matrices are estimated in the same way as a state-of-the-art IVA method, auxiliary-function-based IVA (AuxIVA). Experimental evaluations show that the proposed approach is effective and improves the source separation performance of IVA. In addition, several post-filters aiming to realize multichannel Wiener filter (MWF) are investigated. This setup proves to further increase the performance of IVA. The presented method shows a potential to provide a general way to improve the separation performance from single-channel source separation to multichannel source separation. Ana Ramírez López, Nobutaka Ono, Ulpu Remes, Kalle J. Palomäki, Mikko Kurimo |
ICASSP | 5 |
| 2015 | Adaptation of Morph-Based Speech Recognition for Foreign Names and AcronymsabstractIn this paper, we improve morph-based speech recognition system by focusing adaptation efforts on acronyms (ACRs) and foreign proper names (FPNs). An unsupervised language model (LM) adaptation framework based on two-pass decoding is used. Vocabulary adaptation is applied alongside unsupervised LM adaptation. The aim is to improve both language and pronunciation modeling for FPNs and ACRs. A smart selection algorithm is used to find the most likely topically related foreign words and acronyms from in-domain text. New pronunciation rules are generated for the selected words. Different kinds of morpheme adaptation operations are also evaluated on the ACR and FPN candidate words, to ensure optimal results are gained from pronunciation adaptation. Statistically significant improvements in average word error rate (WER), and term error rate (TER), are achieved using a combination of unsupervised LM adaptation with vocabulary adaptation focused on ACRs and FPNs. André Mansikkaniemi, Mikko Kurimo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Bounded Conditional Mean Imputation with Observation Uncertainties and Acoustic Model AdaptationabstractAutomatic speech recognition systems use noise compensation and acoustic model adaptation to increase robustness towards speaker and environmental variation. The current work focuses on noise compensation with bounded conditional mean imputation (BCMI). BCMI approaches are missing-data methods which operate on the assumption that noise-corrupted observations can be divided into reliable and unreliable components. BCMI methods substitute the unreliable components with a clean speech posterior distribution. The posterior means can be used as clean speech estimates and the posterior variances can be introduced in acoustic model likelihood calculation as observation uncertainties. In addition, we propose in the current work that similar uncertainties are introduced in acoustic model adaptation. Evaluation with speech data recorded in diverse public and car environments indicates that the proposed uncertainties improve adaptation performance. When uncertainties were used in acoustic model likelihood calculation and adaptation, the proposed imputation and adaptation system introduced 15%-84% relative error reductions to an uncompensated baseline system performance. Ulpu Remes, Ana Ramírez López, Kalle J. Palomäki, Mikko Kurimo |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | Morfessor FlatCat: An HMM-Based Method for Unsupervised and Semi-Supervised Learning of Morphology
Stig-Arne Grönroos, Sami Virpioja, Peter Smit, Mikko Kurimo |
COLING | 4 |
| 2014 | Painless Semi-Supervised Morphological Segmentation using Conditional Random FieldsabstractTeemu Ruokolainen, Oskar Kohonen, Sami Virpioja, Mikko Kurimo. Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, volume 2: Short Papers. 2014. Teemu Ruokolainen, Oskar Kohonen, Sami Virpioja, Mikko Kurimo |
EACL | 4 |
| 2014 | Accelerated Estimation of Conditional Random Fields using a Pseudo-Likelihood-inspired Perceptron VariantabstractTeemu Ruokolainen, Miikka Silfverberg, Mikko Kurimo, Krister Linden. Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, volume 2: Short Papers. 2014. Teemu Ruokolainen, Miikka Silfverberg, Mikko Kurimo, Krister Lindén |
EACL | 3 |
| 2014 | Morfessor 2.0: Toolkit for statistical morphological segmentationabstractMorfessor is a family of probabilistic machine learning methods for finding the morphological segmentation from raw text data.Recent developments include the development of semi-supervised methods for utilizing annotated data.Morfessor 2.0 is a rewrite of the original, widely-used Morfessor 1.0 software, with well documented command-line tools and library interface.It includes new features such as semi-supervised learning, online training, and integrated evaluation code. Peter Smit, Sami Virpioja, Stig-Arne Grönroos, Mikko Kurimo |
EACL | 4 |
| 2014 | Unsupervised feature extraction for multimedia event detection and ranking using audio contentabstractIn this paper, we propose a new approach to classify and rank multimedia events based purely on audio content using video data from TRECVID-2013 multimedia event detection (MED) challenge. We perform several layers of nonlinear mappings to extract a set of unsupervised features from an initial set of temporal and spectral features to obtain a superior presentation of the atomic audio units. Additionally, we propose a novel weighted divergence measure for kernel based classifiers. The extensive set of experiments confirms that augmentation of the proposed steps results in an improved accuracy for most of the event classes. Ehsan Amid, Annamaria Mesaros, Kalle J. Palomäki, Jorma Laaksonen, Mikko Kurimo |
ICASSP | 5 |
| 2014 | On the role of missing data imputation and NMF feature enhancement in building synthetic voices using reverberant speech
Dhananjaya Gowda, Heikki Kallasjoki, Reima Karhila, Cristian Contan, Kalle J. Palomäki, Mircea Giurgiu, Mikko Kurimo |
INTERSPEECH | 7 |
| 2014 | Spectral tilt modelling with GMMs for intelligibility enhancement of narrowband telephone speech
Emma Jokinen, Ulpu Remes, Marko Takanen, Kalle J. Palomäki, Mikko Kurimo, Paavo Alku |
INTERSPEECH | 5 |
| 2014 | A Toolkit for Efficient Learning of Lexical Units for Speech Recognition
Matti Varjokallio, Mikko Kurimo |
LREC | 2 |
| 2014 | A word-level token-passing decoder for subword n-gram LVCSRabstractThe decoder is a key component of any modern speech recognizer. Morphologically rich languages pose special challenges for the decoder design, as a very large recognition vocabulary is required to avoid high out-of-vocabulary (OOV) rates. To alleviate these issues, the n-gram models are often trained over subwords instead of words. A subword n-gram model is able to assign probabilities to unseen word forms. We review token-passing decoding and suggest a novel way of creating the decoding graph for subword n-grams on word-level. This approach has the advantage of a better control over the recognition vocabulary, including removal of nonsense words and the possibility to include important OOV-words to the graph. The different decoders are evaluated in a Finnish large vocabulary continuous speech recognition (LVCSR) task. Matti Varjokallio, Mikko Kurimo |
SLT | 2 |
| 2013 | Learning a subword vocabulary based on unigram likelihoodabstractUsing words as vocabulary units for tasks like speech recognition is infeasible for many morphologically rich languages, including Finnish. Thus, subword units are commonly used for language modeling. This work presents a novel algorithm for creating a subword vocabulary, based on the unigram likelihood of a text corpus. The method is evaluated with entropy measure and a Finnish LVCSR task. Unigram entropy of the text corpus is shown to be a good indicator for the quality of higher order n-gram models, also resulting in high speech recognition accuracy. Matti Varjokallio, Mikko Kurimo, Sami Virpioja |
ASRU | 2 |
| 2013 | Supervised Morphological Segmentation in a Low-Resource Learning Setting using Conditional Random Fields
Teemu Ruokolainen, Oskar Kohonen, Sami Virpioja, Mikko Kurimo |
CoNLL | 4 |
| 2013 | HMM-based speech synthesis adaptation using noisy data: Analysis and evaluation methodsabstractThis paper investigates the role of noise in speaker-adaptation of HMM-based text-to-speech (TTS) synthesis and presents a new evaluation procedure. Both a new listening test based on ITU-T recommendation 835 and a perceptually motivated objective measure, frequency-weighted segmental SNR, improve the evaluation of synthetic speech when noise is present. The evaluation of voices adapted with noisy data show that the noise plays a relatively small but noticeable role in the quality of synthetic speech: Naturalness and speaker similarity are not affected in a significant way by the noise, but listeners prefer the voices trained from cleaner data. Noise removal, even when it degrades natural speech quality, improves the synthetic voice. Reima Karhila, Ulpu Remes, Mikko Kurimo |
ICASSP | 3 |
| 2013 | Analysis of breathy, modal and pressed phonation based on low frequency spectral density
Dhananjaya Gowda, Mikko Kurimo |
INTERSPEECH | 2 |
| 2013 | Robust formant detection using group delay function and stabilized weighted linear prediction
Dhananjaya Gowda, Jouni Pohjalainen, Mikko Kurimo, Paavo Alku |
INTERSPEECH | 3 |
| 2013 | Unsupervised topic adaptation for morph-based speech recognition
André Mansikkaniemi, Mikko Kurimo |
INTERSPEECH | 2 |
| 2013 | Lombard modified text-to-speech synthesis for improved intelligibility: submission for the hurricane challenge 2013abstractThis paper describes modification of a TTS system for im-proving the intelligibility of speech in various noise conditions. First, the GlottHMM vocoder is used for training a voice with modal speech data. The vocoder and voice parameters are then modified to mimic the properties of Lombard effect based on a small amount of Lombard speech from the same speaker. More specifically, the durations are increased, fundamental frequency is raised, spectral tilt is decreased, the harmonic-to-noise ratio is increased, and a pressed glottal flow pulses are used in cre-ating excitation. The formants of the speech are also enhanced and finally the speech is compressed in order to increase noise robustness of the voice. The evaluation results of the Hurricane Challenge 2013 indicate that the modified voice is mostly less intelligible than the unmodified natural speech, as expected, but more intelligible than the reference TTS voice, especially in the low SNR conditions. Antti Suni, Reima Karhila, Tuomo Raitio, Mikko Kurimo, Martti Vainio, Paavo Alku |
INTERSPEECH | 4 |
| 2013 | Personalising speech-to-speech translation: Unsupervised cross-lingual speaker adaptation for HMM-based speech synthesis
John Dines, Lakshmi Babu Saheer, Matthew Gibson, William J. Byrne, Keiichiro Oura, Keiichi Tokuda, Junichi Yamagishi, Simon King 0001, Mirjam Wester, Teemu Hirsimäki, Reima Karhila, Mikko Kurimo |
Comput. Speech Lang. | 13 |
| 2012 | Creating synthetic voices for children by adapting adult average voice using stacked transformations and VTLNabstractThis paper describes experiments in creating personalised children's voices for HMM-based synthesis by adapting either an adult or child average voice. The adult average voice is trained from a large adult speech database, whereas the child average voice is trained using a small database of children's speech. Here we present the idea to use stacked transformations for creating synthetic child voices, where the child average voice is first created from the adult average voice through speaker adaptation using all the pooled speech data from multiple children and then adding child specific speaker adaptation on top of it. VTLN is applied to speech synthesis to see whether it helps the speaker adaptation when only a small amount of adaptation data is available. The listening test results show that the stacked transformations significantly improve speaker adaptation for small amounts of data, but the additional benefit provided by VTLN is not yet clear. Reima Karhila, Rama Sanand Doddipatla, Mikko Kurimo, Peter Smit |
ICASSP | 3 |
| 2012 | Improving Discriminative Training for Robust Acoustic Models in Large Vocabulary Continuous Speech Recognition
Janne Pylkkönen, Mikko Kurimo |
INTERSPEECH | 2 |
| 2012 | Optimization-Based Control for the Extended Baum-Welch Algorithm
Janne Pylkkönen, Mikko Kurimo |
INTERSPEECH | 2 |
| 2012 | Bandwidth Extension of Telephone Speech to Low Frequencies Using Sinusoidal Synthesis and a Gaussian Mixture ModelabstractThe quality of narrowband telephone speech is degraded by the limited audio bandwidth. This paper describes a method that extends the bandwidth of telephone speech to the frequency range 0-300 Hz. The method generates the lowest harmonics of voiced speech using sinusoidal synthesis. The energy in the extension band is estimated from spectral features using a Gaussian mixture model. The amplitudes and phases of the synthesized sinusoidal components are adjusted based on the amplitudes and phases of the narrowband input speech, which provides adaptivity to varying input bandwidth characteristics. The proposed method was evaluated with listening tests in combination with another bandwidth extension method for the frequency range 4-8 kHz. While the low-frequency bandwidth extension was not found to improve perceived quality, the method reduced dissimilarity with wideband speech. Hannu Pulakka, Ulpu Remes, Santeri Yrttiaho, Kalle J. Palomäki, Mikko Kurimo, Paavo Alku |
IEEE Trans. Speech Audio Process. | 5 |
| 2012 | Analysis of Extended Baum-Welch and Constrained Optimization for Discriminative Training of HMMsabstractDiscriminative training is an essential part in building a state-of-the-art speech recognition system. The Extended Baum–Welch (EBW) algorithm is the most popular method to carry out this demanding large-scale optimization task. This paper presents a novel analysis of the EBW algorithm which shows that EBW is performing a specific kind of constrained optimization. The constraints show an interesting connection between the improvement of the discriminative criterion and the Kullback–Leibler divergence (KLD). Based on the analysis, a novel method for controlling the EBW algorithm is proposed. The presented analysis uses decomposed formulae for Gaussian mixture KLDs which correspond to the ones used in the Constrained Line Search (CLS) optimization algorithm. The CLS algorithm for discriminative training is therefore also briefly presented and its connections to EBW studied. Large vocabulary speech recognition experiments are used to evaluate the proposed controlling of EBW, which is shown to outperform the common heuristics in model robustness. Comparison of EBW to CLS also shows differences in robustness in favor to EBW. The constraints for Gaussian parameter optimization as well as the special mixture weight estimation method used with EBW are shown to be the key factors for good performance. Janne Pylkkönen, Mikko Kurimo |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Speech bandwidth extension using Gaussian mixture model-based estimation of the highband mel spectrumabstractThe quality and intelligibility of narrowband telephone speech can be enhanced by artificial bandwidth extension. This study combines Gaussian mixture model-based (GMM) mel spectrum extension with a filter bank implementation for generating the missing spectral content in the highband at 4-8 kHz. The narrowband mel spectrum is calculated from input speech and the GMM is used to estimate the mel spectrum in the highband. An excitation signal for the highband is generated as a combination of upsampled linear prediction residual and modulated noise. The excitation is divided into sub-bands that are weighted and summed to realize the estimated mel spectrum. The bandwidth-extended output is obtained as the sum of the artificial highband signal and narrowband speech. Listening tests indicate that this method is preferred over narrowband speech and over a previously presented artificial bandwidth extension method which is implemented in some mobile phone models. Hannu Pulakka, Ulpu Remes, Kalle J. Palomäki, Mikko Kurimo, Paavo Alku |
ICASSP | 4 |
| 2011 | Using stacked transformations for recognizing foreign accented speechabstractA common problem in speech recognition for foreign accented speech is that there is not enough training data for an accent-specific or a speaker-specific recognizer. Speaker adaptation can be used to improve the accuracy of a speaker independent recognizer, but a lot of adaptation data is needed for speakers with a strong foreign accent. In this paper we propose a rather simple and successful technique of stacked transformations where the baseline models trained for native speakers are first adapted by using accent-specific data and then by another transformation using speaker-specific data. Because the accent-specific data can be collected offline, the first transformation can be more detailed and comprehensive, and the second one less detailed and fast. Experimental results are provided for speaker adaptation in English spoken by Finnish speakers. The evaluation results confirm that the stacked transformations are very helpful for fast speaker adaptation. Peter Smit, Mikko Kurimo |
ICASSP | 2 |
| 2011 | Noise Robust Feature Extraction Based on Extended Weighted Linear Prediction in LVCSRabstractThis paper introduces extended weighted linear prediction (XLP) to noise robust short-time spectrum analysis in the feature extraction process of a speech recognition system. XLP is a generalization of standard linear prediction (LP) and temporally weighted linear prediction (WLP) which have already been applied to noise robust speech recognition with good results. With XLP, higher controllability to the temporal weighting of different parts of the noisy speech is gained by taking the lags of the signal into account in prediction. Here, the performance of XLP is put up against WLP and conventional spectrum analysis methods FFT and LP on a large vocabulary continuous speech recognition (LVCSR) scheme using real world noisy data containing additive and convolutive noise. The results show improvements over the reference methods in several cases. Index Terms: linear prediction, temporal weighting, noise robust, speech recognition Sami Keronen, Jouni Pohjalainen, Paavo Alku, Mikko Kurimo |
INTERSPEECH | 4 |
| 2011 | Low-Frequency Bandwidth Extension of Telephone Speech Using Sinusoidal Synthesis and Gaussian Mixture Model
Hannu Pulakka, Ulpu Remes, Santeri Yrttiaho, Kalle J. Palomäki, Mikko Kurimo, Paavo Alku |
INTERSPEECH | 5 |
| 2011 | A Study on Combining VTLN and SAT to Improve the Performance of Automatic Speech Recognition
Rama Sanand Doddipatla, Mikko Kurimo |
INTERSPEECH | 2 |
| 2011 | Missing-Feature Reconstruction With a Bounded Nonlinear State-Space ModelabstractMissing-feature reconstruction can improve speech recognition performance in unknown noisy environments. In this work, we examine using a nonlinear state-space model (NSSM) for missing-feature reconstruction and propose estimation with observed bounds to improve the NSSM performance. Evaluated in large-vocabulary continuous speech recognition task with babble and impulsive noise, using observed bounds in NSSM state estimation significantly improved the method performance. Ulpu Remes, Kalle J. Palomäki, Tapani Raiko, Antti Honkela, Mikko Kurimo |
IEEE Signal Process. Lett. | 5 |
| 2010 | Unsupervised cross-lingual speaker adaptation for HMM-based speech synthesis using two-pass decision tree constructionabstractThis paper demonstrates how unsupervised cross-lingual adaptation of HMM-based speech synthesis models may be performed without explicit knowledge of the adaptation data language. A two-pass decision tree construction technique is deployed for this purpose. Using parallel translated datasets, cross-lingual and intralingual adaptation are compared in a controlled manner. Listener evaluations reveal that the proposed method delivers performance approaching that of unsupervised intralingual adaptation. Matthew Gibson, Teemu Hirsimäki, Reima Karhila, Mikko Kurimo, William J. Byrne |
ICASSP | 4 |
| 2010 | Efficient estimation of maximum entropy language models with n-gram features: an SRILM extensionabstractWe present an extension to the SRILM toolkit for training maximum entropy language models with N -gram features. The extension uses a hierarchical parameter estimation procedure [1] for making the training time and memory consumption feasible for moderately large training data (hundreds of millions of words). Experiments on two speech recognition tasks indicate that the models trained with our implementation perform equally to or better than N -gram models built with interpolated Kneser-Ney discounting. Tanel Alumäe, Mikko Kurimo |
INTERSPEECH | 2 |
| 2010 | Unsupervised cross-lingual speaker adaptation for accented speech recognitionabstractIn this paper we present investigations on how the acoustic models in automatic speech recognition can be adapted across languages in unsupervised fashion to improve recognition of speech with a foreign accent. Recognition systems were trained on large Finnish and English corpora, and tested both on monolingual and bilingual material. Adaptation with bilingual and monolingual recognisers was compared. We found out that recognition of foreign accented English with help of Finnish adaptation training data from the same speaker was not improved significantly. However, the recognition of native Finnish using foreign accented English adaptation data was improved significantly. Reima Karhila, Mikko Kurimo |
SLT | 2 |
| 2010 | Thousands of Voices for HMM-Based Speech Synthesis-Analysis and Application of TTS Systems Built on Various ASR CorporaabstractIn conventional speech synthesis, large amounts of phonetically balanced speech data recorded in highly controlled recording studio environments are typically required to build a voice. Although using such data is a straightforward solution for high quality synthesis, the number of voices available will always be limited, because recording costs are high. On the other hand, our recent experiments with HMM-based speech synthesis systems have demonstrated that speaker-adaptive HMM-based speech synthesis (which uses an “average voice model” plus model adaptation) is robust to non-ideal speech data that are recorded under various conditions and with varying microphones, that are not perfectly clean, and/or that lack phonetic balance. This enables us to consider building high-quality voices on “non-TTS” corpora such as ASR corpora. Since ASR corpora generally include a large number of speakers, this leads to the possibility of producing an enormous number of voices automatically. In this paper, we demonstrate the thousands of voices for HMM-based speech synthesis that we have made from several popular ASR corpora such as the Wall Street Journal (WSJ0, WSJ1, and WSJCAM0), Resource Management, Globalphone, and SPEECON databases. We also present the results of associated analysis based on perceptual evaluation, and discuss remaining issues. Junichi Yamagishi, Bela Usabaev, Simon King 0001, Oliver Watts, John Dines, Jilei Tian, Rile Hu, Keiichiro Oura, Yi-Jian Wu, Keiichi Tokuda, Reima Karhila, Mikko Kurimo |
IEEE Trans. Speech Audio Process. | 13 |
| 2009 | Weighted linear prediction for speech analysis in noisy conditionsabstract1 k p; where E = X n (s n p X k=1 a k s n k ) 2 I WLP is a generalization of LP, introducing temporal weighting of the squared prediction error [2]: E = X n (s n p X k=1 a k s n k ) 2 W n IW n is the weighting function I For constant W n , WLP becomes identical to LP! Stabilized Weighted Linear Prediction (SWLP) Jouni Pohjalainen, Heikki Kallasjoki, Kalle J. Palomäki, Mikko Kurimo, Paavo Alku |
INTERSPEECH | 4 |
| 2009 | Thousands of voices for HMM-based speech synthesisabstractOur recent experiments with HMM-based speech synthesis systems have demonstrated that speaker-adaptive HMM-based speech synthesis (which uses an 'average voice model' plus model adaptation) is robust to non-ideal speech data that are recorded under various conditions and with varying microphones, that are not perfectly clean, and/or that lack of phonetic balance. This enables us consider building high-quality voices on 'non-TTS' corpora such as ASR corpora. Since ASR corpora generally include a large number of speakers, this leads to the possibility of producing an enormous number of voices automatically. In this paper we show thousands of voices for HMM-based speech synthesis that we have made from several popular ASR corpora such as the Wall Street Journal databases (WSJ0/WSJ1/WSJCAM0), Resource Management, Globalphone and Speecon. We report some perceptual evaluation results and outline the outstanding issues. Junichi Yamagishi, Bela Usabaev, Simon King 0001, Oliver Watts, John Dines, Jilei Tian, Rile Hu, Keiichiro Oura, Keiichi Tokuda, Reima Karhila, Mikko Kurimo |
INTERSPEECH | 12 |
| 2009 | Importance of High-Order N-Gram Models in Morph-Based Speech RecognitionabstractSpeech recognition systems trained for morphologically rich languages face the problem of vocabulary growth caused by prefixes, suffixes, inflections, and compound words. Solutions proposed in the literature include increasing the size of the vocabulary and segmenting words into morphs. However, in many cases, the methods have only been experimented with low-order n-gram models or compared to word-based models that do not have very large vocabularies. In this paper, we study the importance of using high-order variable-length n-gram models when the language models are trained over morphs instead of whole words. Language models trained on a very large vocabulary are compared with models based on different morph segmentations. Speech recognition experiments are carried out on two highly inflecting and agglutinative languages, Finnish and Estonian. The results suggest that high-order models can be essential in morph-based speech recognition, even when lattices are generated for two-pass recognition. The analysis of recognition errors reveal that the high-order morph language models improve especially the recognition of previously unseen words. Teemu Hirsimäki, Janne Pylkkönen, Mikko Kurimo |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | Speech to speech machine translation: Biblical chatter from Finnish to English
David Ellis, Mathias Creutz, Timo Honkela, Mikko Kurimo |
IJCNLP | 4 |
| 2007 | Vocabulary Decomposition for Estonian Open Vocabulary Speech Recognition
Antti Puurula, Mikko Kurimo |
ACL | 2 |
| 2007 | Segregation of Speakers for Speaker Adaptation in TV News AudioabstractSpeaker adaptation is commonly used to compensate speaker variation in large vocabulary continuous speech recognition. In a multi-speaker environment where speakers change frequently speaker segregation is needed to divide the input audio stream to speaker turns. Speaker turns define the current speaker at each time and speaker adaptation can thus be done based on speaker turns. The novelty of this paper is that the speaker-specific transformations are estimated incrementally and in tandem with speaker segregation. Therefore we need a transformation that can be reliably estimated based on one speaker turn alone. We propose the constrained maximum likelihood linear regression (CMLLR) for this. In testing with Finnish TV news audio, speaker adaptation reduced the average letter error rate 25% relative to baseline. Ulpu Remes, Janne Pylkkönen, Mikko Kurimo |
ICASSP (4) | 3 |
| 2007 | Morfessor and variKN machine learning tools for speech and language technologyabstractThis paper introduces two recent open source software packages developed for unsupervised natural language modeling. The Morfessor program segments words automatically into morpheme-like units without any rule-based morphological analyzers. The VariKN toolkit trains language models producing a compact set of high-order n-grams utilizing state-of-art Kneser-Ney smoothing. As an example, this paper shows how to construct a language model for speech recognition in multiple languages utilizing only a minimal amount of linguistic resources. Morfessor and VariKN also have other applications in text understanding, information retrieval and machine translation. Unsupervised machine learning techniques are particularly well suited for the development of systems for less-resourced languages, because they do not depend on manually designed morphological or syntactical analyzers or annotated data. Index Terms: variable length language modeling, speech recognition, unsupervised morphology, open source software Vesa Siivola, Mathias Creutz, Mikko Kurimo |
INTERSPEECH | 3 |
| 2007 | Comparison of subspace methods for Gaussian mixture models in speech recognitionabstractSpeech recognizers typically use high-dimensional feature vectors to capture the essential cues for speech recognition purposes. The acoustics are then commonly modeled with a Hidden Markov Model with Gaussian Mixture Models as observation probability density functions. Using unrestricted Gaussian parameters might lead to intolerable model costs both evaluation- and storagewise, which limits their practical use only to some high-end systems. The classical approach to tackle with these problems is to assume independent features and constrain the covariance matrices to being diagonal. This can be thought as constraining the second order parameters to lie in a fixed subspace consisting of rank-1 terms. In this paper we discuss the differences between recently proposed subspace methods for GMMs with emphasis placed on the applicability of the models to a practical LVCSR system. Index Terms: speech recognition, acoustic modeling, Gaussian mixture, multivariate normal distribution, subspace method Matti Varjokallio, Mikko Kurimo |
INTERSPEECH | 2 |
| 2007 | Analysis of Morph-Based Speech Recognition and the Modeling of Out-of-Vocabulary Words Across Languages
Mathias Creutz, Teemu Hirsimäki, Mikko Kurimo, Antti Puurula, Janne Pylkkönen, Vesa Siivola, Matti Varjokallio, Ebru Arisoy, Murat Saraclar, Andreas Stolcke |
HLT-NAACL | 3 |
| 2007 | Indexing confusion networks for morph-based spoken document retrievalabstractIn this paper, we investigate methods for improving the performance of morph-based spoken document retrieval in Finnish by extracting relevant index terms from confusion networks. Our approach uses morpheme-like subword units ("morphs") for recognition and indexing. This alleviates the problem of out-of-vocabulary words, especially with inflectional languages like Finnish. Confusion networks offer a convenient representation of alternative recognition candidates by aligning mutually exclusive terms and by giving the posterior probability of each term. The rank of the competing terms and their posterior probability is used to estimate term frequency for indexing. Comparing against 1-best recognizer transcripts, we show that retrieval effectiveness is significantly improved. Finally, the effect of pruning in recognition is analyzed, showing that when recognition speed is increased, the reduction in retrieval performance due to the increase in the 1-best error rate can be compensated by using confusion networks. Ville T. Turunen, Mikko Kurimo |
SIGIR | 2 |
| 2006 | Unsupervised segmentation of words into morphemes - morpho challenge 2005 application to automatic speech recognitionabstractWithin the EU Network of Excellence PASCAL, a challenge was organized to design a statistical machine learning algorithm that segments words into the smallest meaning-bearing units of language, morphemes. Ideally, these are basic vocabulary units suitable for different tasks, such as speech and text understanding, machine translation, information retrieval, and statistical language modeling. Twelve research groups participated in the challenge and had submitted segmentation results obtained by their algorithms. In this paper, we evaluate the application of these segmentation algorithms to large vocabulary speech recognition using statistical n-gram language models based on the proposed word segments instead of entire words. Experiments were done for two agglutinative and morphologically rich languages: Finnish and Turkish. We also investigate combining various segmentations to improve the performance of the recognizer. Index Terms: speech recognition, language modelling, morphemes, unsupervised learning. Mikko Kurimo, Mathias Creutz, Matti Varjokallio, Ebru Arisoy, Murat Saraclar |
INTERSPEECH | 1 |
| 2006 | Using latent semantic indexing for morph-based spoken document retrievalabstractPreviously, phone-based and word-based approaches have been used for spoken document retrieval.The former suffers from high error rates and the latter from limited vocabulary of the recognizer.Our method relies on unlimited vocabulary continuous speech recognizer that uses morpheme-like units discovered in an unsupervised manner.The morpheme-like units, or "morphs" for short, have been successfully used also as index terms.One problem using morphs as index terms is that the segmentation does not always separate the same stem for different inflected forms of the same word.This resembles the problem of synonyms.In this paper, we apply latent semantic indexing to morph based retrieval.The idea is to project morphs that correspond to the same word, as well as other semantically related terms, to the same dimension.The results show clear improvements in Finnish spoken document retrieval performance. Ville T. Turunen, Mikko Kurimo |
INTERSPEECH | 2 |
| 2006 | Compact n-gram models by incremental growing and clustering of historiesabstractThis work concerns building n-gram language models that are suitable for large vocabulary speech recognition in devices that have a restricted amount of memory and space available. Our target language is Finnish, and in order to evade the problems of its rich morphology, we use sub-word units, morphs, as model units instead of the words. In the proposed model we apply incremental growing and clustering of the morph n-gram histories. By selecting the histories using maximum a posteriori estimation, and clustering them with information radius measure, we obtain a clustered varigram model. We show that for restricted model sizes this model gives better cross-entropy and speech recognition results than the conventional n-gram models, and also better recognition results than non-clustered varigram models built with another recently introduced method. Index Terms: language models, clustering, information radius, speech recognition Sami Virpioja, Mikko Kurimo |
INTERSPEECH | 2 |
| 2006 | Unlimited vocabulary speech recognition for agglutinative languages
Mikko Kurimo, Antti Puurula, Ebru Arisoy, Vesa Siivola, Teemu Hirsimäki, Janne Pylkkönen, Tanel Alumäe, Murat Saraclar |
HLT-NAACL | 1 |
| 2006 | Unlimited vocabulary speech recognition with morph language models applied to Finnish
Teemu Hirsimäki, Mathias Creutz, Vesa Siivola, Mikko Kurimo, Sami Virpioja, Janne Pylkkönen |
Comput. Speech Lang. | 4 |
| 2005 | Methods for combining language models in speech recognitionabstractStatistical language models have a vital part in contemporary speech recognition systems and a lot of language models have been presented in the literature. The best results have been achieved when different language models have been used together. Several combination methods have been presented, but few comparisons of the different methods has been done. In this work, three combination methods that have been used with language models are studied. In addition, a new approach based on likelihood density function estimation using histograms is presented. The methods are evaluated in speech recognition experiments and perplexity calculations. The test data consist of Finnish news articles and four language models work as the component models. In the perplexity experiments, all combining methods produced statistically significant improvement compared to the 4-gram model that worked as a baseline. The best result, 46 % improvement to the 4-gram model, was achieved when combining three language models together by using the new bin estimation method. In the speech recognition experiments, 4 % reduction to the word error and over 7 % reduction to the phoneme error was achieved by unigram rescaling method. 1. Simo Broman, Mikko Kurimo |
INTERSPEECH | 2 |
| 2005 | To recover from speech recognition errors in spoken document retrievalabstractAn important difference between the retrieval of spoken and written documents is that the indexing of the speech data is usually based on automatic speech transcripts that contain recognition errors. However, there are several ways of reducing the effect of incorrect index terms in the retrieval. This paper presents retrieval experiments with unlimited vocabulary speech recognizer that utilizes a lexicon of unsupervised morpheme-like units. Based on this recognizer, three different methods are evaluated for error recovery. First, the recognized words are expanded by adding the recognized morphemes, too. Second, the words are expanded by adding the best rival morpheme candidates that were pruned away by the recognizer. Third, the queries are expanded by the potentially relevant terms found from text documents, which were retrieved from parallel text corpora by the original queries. The best results are obtained by that latter method which significantly improves the precision compared to the original queries and brings the spoken document retrieval precision to the same level as the corresponding text document retrieval. Mikko Kurimo, Ville T. Turunen |
INTERSPEECH | 1 |
| 2005 | SpeechFind: Advances in Spoken Document Retrieval for a National Gallery of the Spoken WordabstractAdvances in formulating spoken document retrieval for a new National Gallery of the Spoken Word (NGSW) are addressed. NGSW is the first large-scale repository of its kind, consisting of speeches, news broadcasts, and recordings from the 20th century. After presenting an overview of the audio stream content of the NGSW, with sample audio files from U.S. Presidents from 1893 to the present, an overall system diagram is proposed with a discussion of critical tasks associated with effective audio information retrieval. These include advanced audio segmentation, speech recognition model adaptation for acoustic background noise and speaker variability, and information retrieval using natural language processing for text query requests that include document and query expansion. For segmentation, a new evaluation criterion entitled fused error score (FES) is proposed, followed by application of the CompSeg segmentation scheme on DARPA Hub4 Broadcast News (30.5% relative improvement in FES) and NGSW data. Transcript generation is demonstrated for a six-decade portion of the NGSW corpus. Novel model adaptation using structure maximum likelihood eigenspace mapping shows a relative 21.7% improvement. Issues regarding copyright assessment and metadata construction are also addressed for the purposes of a sustainable audio collection of this magnitude. Advanced parameter-embedded watermarking is proposed with evaluations showing robustness to correlated noise attacks. Our experimental online system entitled "SpeechFind" is presented, which allows for audio retrieval from a portion of the NGSW corpus. Finally, a number of research challenges such as language modeling and lexicon for changing time periods, speaker trait and identification tracking, as well as new directions, are discussed in order to address the overall task of robust phrase searching in unrestricted audio corpora. John H. L. Hansen, Rongqing Huang, Michael Seadle, John R. Deller Jr., Aparna Gurijala, Mikko Kurimo, Pongtep Angkititrakul |
IEEE Trans. Speech Audio Process. | 7 |
| 2004 | An evaluation of a spoken document retrieval baseline system in finishabstractThis paper presents a baseline spoken document retrieval system in Finnish. Due to its agglutinative structure, Finnish speech can not be adequately transcribed using the standard large vocabulary continuous speech recognition approaches. The definition of a sufficient lexicon and the training of the statistical language models are difficult, because the words appear transformed by many inflections and compounds. In this work we apply a recently developed unlimited vocabulary speech recognition system that allows the use of n-gram language models based on morpheme-like subword units discovered in an unsupervised manner. In addition to word-based indexing, we also propose an indexing based on the subword units provided directly by our speech recognizer. In an initial evaluation of newsreading in Finnish, we obtained a fairly low recognition error rate and average document retrieval precisions close to that from human reference transcripts. Mikko Kurimo, Ville T. Turunen, Inger Ekman |
INTERSPEECH | 1 |
| 2004 | Duration modeling techniques for continuous speech recognitionabstractPhone durations play a significant part in the comprehension of speech. The duration information is still mostly disregarded in automatic speech recognizers due to the use of hidden Markov models (HMMs) which are deficient in modeling phone durations properly. Previous results have shown that using different approaches for explicit duration modeling have improved the isolated word recognition in English. However, a unified comparison between the methods has not been reported In this paper three techniques for explicit duration modeling are compared and evaluated in a large vocabulary continuous speech recognition task. The target language was Finnish, in which phone durations are especially important for proper understanding. The results show that the choice of the duration modeling technique depends on the speed requirements of the recognizer. The best technique required a slightly longer running time than without an explicit duration model, but achieved an 8 % relative improvement to the letter error rate. 1. Janne Pylkkönen, Mikko Kurimo |
INTERSPEECH | 2 |
| 2003 | On lexicon creation for turkish LVCSRabstractIn this paper, we address the lexicon design problem in Turkish large vocabulary speech recognition. Although we focus only on Turkish, the methods described here are general enough that they can be considered for other agglutinative languages like Finnish, Korean etc. In an agglutinative language, several words can be created from a single root word using a rich collection of morphological rules. So, a virtually infinite size lexicon is required to cover the language if words are used as the basic units. The standard approach to this problem is to discover a number of primitive units so that a large set of words can be created by compounding those units. Two broad classes of methods are available for splitting words into their sub-units; morphology-based and data-driven methods. Although the word splitting significantly reduces the out of vocabulary rate, it shrinks the context and increases acoustic confusibility. We have used two methods to address the latter. In one method, we use word counts to avoid splitting of high frequency lexical units, and in the other method, we recompound splits according to a probabilistic measure. We present experimental results that show the methods are very effective to lower the word error rate at the expense of lexicon size. 1. Kadri Hacioglu, Bryan L. Pellom, Tolga Çiloglu, Özlem Öztürk, Mikko Kurimo, Mathias Creutz |
INTERSPEECH | 5 |
| 2003 | Unlimited vocabulary speech recognition based on morphs discovered in an unsupervised mannerabstractWe study continuous speech recognition based on sub-word units found in an unsupervised fashion. For agglutinative languages like Finnish, traditional word-based n-gram language modeling does not work well due to the huge number of different word forms. We use a method based on the Minimum Description Length principle to split words statistically into subword units allowing efficient language modeling and unlimited vocabulary. The perplexity and speech recognition experiments on Finnish speech data show that the resulting model outperforms both word and syllable based trigram models. Compared to the word trigram model, the out-of-vocabulary rate is reduced from 20% to 0% and the word error rate from 56% to 32%. Vesa Siivola, Teemu Hirsimäki, Mathias Creutz, Mikko Kurimo |
INTERSPEECH | 4 |
| 2002 | An Efficiently Focusing Large Vocabulary Language Model
Mikko Kurimo, Krista Lagus |
ICANN | 1 |
| 2002 | Thematic indexing of spoken documents by using self-organizing maps
Mikko Kurimo |
Speech Commun. | 1 |
| 2001 | Large vocabulary statistical language modeling for continuous speech recognition in finnishabstractStatistical language modeling (SLM) is an essential part in any large-vocabulary continuous speech recognition (LVCSR) system. The development of the standard SLM methods has been strongly affected by the goals of LVCSR in English. The structure of Finnish is substantially different from English, so if the standard SLMs are directly applied, the success is by no means granted. In this paper we describe our first attempts of building a LVCSR for Finnish and the new SLMs that we have tried. One of our objective has been the indexing and recognition of broadcast news, so special issues of our interest are topic detection, word stemming and modeling words that are poorly covered in the training data. Our new methods are based on neural computing using the self-organizing map (SOM) which has recently been shown to successfully extract and approximate latent semantic structures from massive text collections. 1. Vesa Siivola, Mikko Kurimo, Krista Lagus |
INTERSPEECH | 2 |
| 2000 | Fast latent semantic indexing of spoken documents by using self-organizing mapsabstractThis paper describes a new latent semantic indexing (LSI) method for spoken audio documents. The framework is indexing broadcast news from radio and TV as a combination of large vocabulary continuous speech recognition (LVCSR), natural language processing (NLP) and information retrieval (IR). For indexing, the documents are presented as vectors of word counts, whose dimensionality is rapidly reduced by random mapping (RM). The obtained vectors are projected into the latent semantic subspace determined by SVD, where the vectors are then smoothed by a self-organizing map (SOM). The smoothing by the closest document clusters is important here, because the documents are often short and have a high word error rate (WER). As the clusters in the semantic subspace reflect the news topics, the SOMs provide an easy way to visualize the index and query results and to explore the database. Test results are reported for TREC's spoken document retrieval databases (www.idiap.ch/kurimo/thisl.html). Mikko Kurimo |
ICASSP | 1 |
| 1998 | Self-organization in mixture densities of HMM based speech recognition
Mikko Kurimo |
ESANN | 1 |
| 1998 | Improving vocabulary independent HMM decoding results by using the dynamically expanding contextabstractA method is presented to correct phoneme strings produced by a vocabulary independent speech recognizer. The method first extracts the N best matching result strings using mixture density hidden Markov models (HMMs) trained by neural networks. Then the strings are corrected by the rules generated automatically by the dynamically expanding context (DEC). Finally, the corrected string candidates and the extra alternatives proposed by the DEC are ranked according to the likelihood score of the best HMM path to generate the obtained string. The experiments show that N need not be very large and the method is able to decrease recognition errors from a test data that even has no common words with the training data of the speech recognizer. Mikko Kurimo |
ICASSP | 1 |
| 1997 | Comparison results for segmental training algorithms for mixture density HMMsabstractThis work presents experiments on four segmental training algorithms for mixture density HMMs. The segmental versions of SOM and LVQ3 suggested by the author are compared against the conventional segmental K-means and the segmental GPD. The recognition task used as a test bench is the speaker dependent, but vocabulary independent automatic speech recognition. The output density function of each state in each model is a mixture of multivariate Gaussian densities. Neural network methods SOM and LVQ are applied to learn the parameters of the density models from the mel-cepstrum features of the training samples. The segmental training improves the segmentation and the model parameters by turns to obtain the best possible result, because the segmentation and the segment classification depend on each other. It suffices to start the training process by dividing the training samples approximatively into phoneme samples. 1. INTRODUCTION The recognition task used as a test bench for the trainin... Mikko Kurimo |
EUROSPEECH | 1 |
| 1997 | Training mixture density HMMs with SOM and LVQ
Mikko Kurimo |
Comput. Speech Lang. | 1 |
| 1996 | Using the self-organizing map to speed up the probability density estimation for speech recognition with mixture density HMMs
Mikko Kurimo, Panu Somervuo |
ICSLP | 1 |
| 1993 | Using LVQ to enhance semi-continuous hidden Markov models for phonemesabstractExperiments are made to enhance the discrimination ability of the SCHMMs by applying Learning Vector Quantization. The SCHMMs are used for the modeling of phonemes in a speaker-dependent speech recognition application to create the phonetic transcriptions of spoken utterances. The probability density functions for the cepstral feature vectors produced in each state of each model are modeled by mixtures of multivariate Gaussian density functions. The mean vectors of the Gaussian densities are chosen by clustering the feature vectors of the training samples by using the Self-Organizing Map (SOM). Then the Gaussians are modified to correspond better to the Bayesian decision surfaces between phonemes by tuning the mean vectors by the LVQ. The experiments indicate that by this careful placement of the mean vectors the recognition error rates for the SCHMMs decrease significantly. LVQ algorithms can also be successfully applied after Baum-Welch or Viterbi training to slightly modify the Gaus... Mikko Kurimo |
EUROSPEECH | 1 |
| 1992 | Application of self-organizing maps and LVQ in training continuous density hidden Markov models for phonemesabstractWe present experiments in using neural network based methods to initialize continuous observation density hidden Markov models (CDHMMs). Proper initialization provides an easy way to avoid excessive amount of iterations, when maximum likelihood algorithms are used to estimate the parameters of CDHMMs. This is important in, for example, phoneme based automatic speech recognition, where the output density functions of the states of HMMs are complex and a lot of training data must be used. In our work CDHMMs are used as phoneme models in the task of transcribing speech into phoneme sequences. The probability density function of the output distribution for a state is approximated by mixture of a large number of multivariate Gaussian density functions (typically 25). We present experiments of initializing the means of mixture Gaussians by Self-Organizing Maps (SOMs) and Learning Vector Quantization (LVQ). The results of the experiments indicate that initialization by SOMs speeds up the conv... Mikko Kurimo, Kari Torkkola |
ICSLP | 1 |
| 1991 | Improving short-time speech frame recognition results by using contextabstractThis paper focuses on comparing three approaches to improve the accuracy of classifying short-time speech frames into phoneme classes by taking into account the classifications of nearby frames, also individually classified. We investigate whether this improvement has an effect to the accuracy of transcribing speech into phoneme sequences using two different decoding schemes, one based on simple durational rules, and the other on hidden Markov models (HMMs). The experiments indicate that recognition accuracies can indeed be improved significantly by taking the local context into account. 1 INTRODUCTION "More is to be gained by discovering suboptimal ways of handling context than by discovering optimal ways of handling the local structure", Haralick stated in [2]. In this paper we explore some options to take advantage of local context in improving classification accuracy. Our framework is a phonemic speech recognizer based on classifying shorttime feature vectors. The viewpoint taken i... Kari Torkkola, Mikko Kokkonen, Mikko Kurimo, Pekka Utela |
EUROSPEECH | 3 |