Okko Johannes Räsänen

dblp:35/9230 · also Okko Räsänen · DBLP profile ↗
← Back
78ranked-venue papers
28as first author
21since 2021 · last 2026
0000-0002-0537-0946ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 57 · 18 first-author · 15 since 2021Artificial intelligence and machine learning · 56 · 23 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 8 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Corrigendum to "FinnAffect: An affective speech corpus for spontaneous Finnish" [Speech Communication 175 (2025) 103327]
Kalle Lahtinen, Liisa Mustanoja, Okko Johannes Räsänen
Speech Commun.3
2025 Investigating Affect Mining Techniques for Annotation Sample Selection in the Creation of Finnish Affective Speech Corpus
abstract
Study of affect in speech requires suitable data, as emotional expression and perception vary across languages. Until now, no corpus has existed for natural expression of affect in spontaneous Finnish, existing data being acted or from a very specific communicative setting. This paper presents the first such corpus, created by annotating 12,000 utterances for emotional arousal and valence, sampled from three large-scale Finnish speech corpora. To ensure diverse affective expression, sample selection was conducted with an affect mining approach combining acoustic, cross-linguistic speech emotion, and text sentiment features. We compare this method to random sampling in terms of annotation diversity, and conduct post-hoc analyses to identify sampling choices that would have maximized the diversity. As an outcome, the work introduces a spontaneous Finnish affective speech corpus and informs sampling strategies for affective speech corpus creation in other languages or domains.
Kalle Lahtinen, Einari Vaaras, Liisa Mustanoja, Okko Johannes Räsänen
INTERSPEECH4
2025 A model of early word acquisition based on realistic-scale audiovisual naming events
abstract
Infants gradually learn to parse continuous speech into words and connect names with objects, yet the mechanisms behind development of early word perception skills remain unknown. We studied the extent to which early words can be acquired through statistical learning from regularities in audiovisual sensory input. We simulated word learning in infants up to 12 months of age in a realistic setting, using a model that solely learns from statistical regularities in unannotated raw speech and pixel-level visual input. Crucially, the quantity of object naming events was carefully designed to match that accessible to infants of comparable ages. Results show that the model effectively learns to recognize words and associate them with corresponding visual objects, with a vocabulary growth rate comparable to that observed in infants. The findings support the viability of general statistical learning for early word perception, demonstrating how learning can operate without assuming any prior linguistic capabilities. • A model of cross-situational audiovisual word acquisition in infants aged 0–12 months. • Learning simulated with realistic-scale speech and audiovisual naming events. • The learner model succeeds in learning words and word meanings from the input. • Feasibility proof for generic statistical learning in early language bootstrapping.
Khazar Khorrami, Okko Johannes Räsänen
Speech Commun.2
2025 FinnAffect: An affective speech corpus for spontaneous Finnish
abstract
Affective expression plays a major role in everyday spoken and written language. In order to study how affect is expressed by Finnish language users in day-to-day life, data consisting of samples from naturalistic and unscripted contexts is required. The present work describes the first spontaneous speech corpus for Finnish with affect-related annotations, containing 12,000 transcribed samples of unscripted speech paired with continuous-valued scores of valence and arousal marked by five native Finnish speakers. We first describe the creation of the corpus, based on combining speech samples from three large-scale Finnish speech corpora, from which we chose samples for annotation using an active learning-based affect mining approach. We then report characteristics of the resulting corpus and annotation consistency, followed by speech emotion recognition (SER) experiments with several classifiers and regression models to test the feasibility of the corpus for SER system development and evaluation. Annotation analyses reveal mean Pearson correlations between annotator scores and the mean of all annotators to be ρ m e a n = 0 . 856 for valence and ρ m e a n = 0 . 898 for arousal. The SER experiments on discretized labels result in an average unweighted average recall (UAR) of 0.458 for ternary valence classification and 0.719 for binary arousal classification using a fine-tuned ExHuBERT model for valence prediction and a support vector machine (SVM) classifier for arousal prediction, reaching comparable levels to those reported earlier for spontaneous speech. For the regression task, concordance correlation coefficients of 0.270 and 0.689 were obtained for valence and arousal, respectively, when using a WavLM-based model trained on MSP-Podcast corpus and fine-tuned on the target data. Overall, the analyses suggest that the corpus provides a feasible basis for later study on affective expression in spontaneous Finnish.
Kalle Lahtinen, Liisa Mustanoja, Okko Johannes Räsänen
Speech Commun.3
2025 Text-Based Audio Retrieval by Learning From Similarities Between Audio Captions
abstract
This letter proposes to use similarities of audio captions for estimating audio-caption relevances to be used for training text-based audio retrieval systems. Current audio-caption datasets (e.g., Clotho) contain audio samples paired with annotated captions, but lack relevance information about audio samples and captions beyond the annotated ones. Besides, mainstream approaches (e.g., CLAP) usually treat the annotated pairs as positives and consider all other audio-caption combinations as negatives, assuming a binary relevance between audio samples and captions. To infer the relevance between audio samples and arbitrary captions, we propose a method that computes non-binary audio-caption relevance scores based on the textual similarities of audio captions. We measure textual similarities of audio captions by calculating the cosine similarity of their Sentence-BERT embeddings and then transform these similarities into audio-caption relevance scores using a logistic function, thereby linking audio samples through their annotated captions to all other captions in the dataset. To integrate the computed relevances into training, we employ a listwise ranking objective, where relevance scores are converted into probabilities of ranking audio samples for a given textual query. We show the effectiveness of the proposed method by demonstrating improvements in text-based audio retrieval compared to methods that use binary audio-caption relevances for training.
Huang Xie, Khazar Khorrami, Okko Johannes Räsänen, Tuomas Virtanen
IEEE Signal Process. Lett.3
2024 Age-Dependent Analysis and Stochastic Generation of Child-Directed Speech
Okko Johannes Räsänen, Daniil Kocharov
CogSci1
2024 The Difficulty and Importance of Estimating the Lower and Upper Bounds of Infant Speech Exposure
Joseph Coffey, Okko Johannes Räsänen, Camila Scaff, Alejandrina Cristià
INTERSPEECH2
2023 Analysing the Impact of Audio Quality on the Use of Naturalistic Long-Form Recordings for Infant-Directed Speech Research
María Andrea Cruz Blandón, Alejandrina Cristià, Okko Johannes Räsänen
CogSci3
2023 Computational Insights to Acquisition of Phonemes, Words, and Word Meanings in Early Language: Sequential or Parallel Acquisition?
Khazar Khorrami, María Andrea Cruz Blandón, Okko Johannes Räsänen
CogSci3
2023 Is Reliability of Cognitive Measures in Children Dependent on Participant Age? A Case Study with Two Large-Scale Datasets
Okko Johannes Räsänen, María Andrea Cruz Blandón, Jukka Leppänen
CogSci1
2023 On Negative Sampling for Contrastive Audio-Text Retrieval
abstract
This paper investigates negative sampling for contrastive learning in the context of audio-text retrieval. The strategy for negative sampling refers to selecting negatives (either audio clips or textual descriptions) from a pool of candidates for a positive audio-text pair. We explore sampling strategies via model-estimated within-modality and cross-modality relevance scores for audio and text samples. With a constant training setting on the retrieval system from [1], we study eight sampling strategies, including hard and semi-hard negative sampling. Experimental results show that retrieval performance varies dramatically among different strategies. Particularly, by selecting semi-hard negatives with cross-modality scores, the retrieval system gains improved performance in both text-to-audio and audio-to-text retrieval. Besides, we show that feature collapse occurs while sampling hard negatives with cross-modality scores.
Huang Xie, Okko Johannes Räsänen, Tuomas Virtanen
ICASSP2
2023 BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language models
abstract
International audience
Marvin Lavechin, Yaya Sy, Hadrien Titeux, María Andrea Cruz Blandón, Okko Johannes Räsänen, Hervé Bredin, Emmanuel Dupoux, Alejandrina Cristià
INTERSPEECH5
2023 Syllable Discovery and Cross-Lingual Generalization in a Visually Grounded, Self-Supervised Speech Model
abstract
In this paper, we show that representations capturing syllabic units emerge when training a self-supervised speech model with a visually-grounded training objective. We demonstrate that a nearly identical model architecture (HuBERT) trained with a masked language modeling loss does not exhibit this same ability, suggesting that the visual grounding objective is responsible for the emergence of this phenomenon. We propose the use of a minimum cut algorithm to automatically predict syllable boundaries in speech, followed by a 2-stage clustering method to group identical syllables together. We show that our model not only outperforms a state-of-the-art syllabic segmentation method on the language it was trained on (English), but also generalizes in a zero-shot fashion to Estonian. Finally, we show that the same model is capable of zero-shot generalization for a word segmentation task on 4 other languages from the Zerospeech Challenge, in some cases beating the previous state-of-the-art.
Puyuan Peng, Shang-Wen Li 0001, Okko Johannes Räsänen, Abdel-rahman Mohamed, David F. Harwath
INTERSPEECH3
2023 Development of a speech emotion recognizer for large-scale child-centered audio recordings from a hospital environment
abstract
In order to study how early emotional experiences shape infant development, one approach is to analyze the emotional content of speech heard by infants, as captured by child-centered daylong recordings, and as analyzed by automatic speech emotion recognition (SER) systems. However, since large-scale daylong audio is initially unannotated and differs from typical speech corpora from controlled environments, there are no existing in-domain SER systems for the task. Based on existing literature, it is also unclear what is the best approach to deploy a SER system for a new domain. Consequently, in this study, we investigated alternative strategies for deploying a SER system for large-scale child-centered audio recordings from a neonatal hospital environment, comparing cross-corpus generalization, active learning (AL), and domain adaptation (DA) methods in the process. We first conducted simulations with existing emotion-labeled speech corpora to find the best strategy for SER system deployment. We then tested how the findings generalize to our new initially unannotated dataset. As a result, we found that the studied AL method provided overall the most consistent results, being less dependent on the specifics of the training corpora or speech features compared to the alternative methods. However, in situations without the possibility to annotate data, unsupervised DA proved to be the best approach. We also observed that deployment of a SER system for real-world daylong child-centered audio recordings achieved a SER performance level comparable to those reported in literature, and that the amount of human effort required for the system deployment was overall relatively modest.
Einari Vaaras, Sari Ahlqvist-Björkroth, Konstantinos Drossos, Liisa Lehtonen, Okko Johannes Räsänen
Speech Commun.5
2023 Automatic Assessment of Parkinson's Disease Using Speech Representations of Phonation and Articulation
abstract
Speech from people with Parkinson's disease (PD) are likely to be degraded on phonation, articulation, and prosody. Motivated to describe articulation deficits comprehensively, we investigated 1) the universal phonological features that model articulation manner and place, also known as speech attributes, and 2) glottal features capturing phonation characteristics. These were further supplemented by, and compared with, prosodic features using a popular compact feature set and standard MFCC. Temporal characteristics of these features were modeled by convolutional neural networks. Besides the features, we were also interested in the speech tasks for collecting data for automatic PD speech assessment, like sustained vowels, text reading, and spontaneous monologue. For this, we utilized a recently collected Finnish PD corpus (PDSTU) as well as a Spanish database (PC-GITA). The experiments were formulated as regression problems against expert ratings of PD-related symptoms, including ratings of speech intelligibility, voice impairment, overall severity of communication disorder on PDSTU, as well as on the Unified Parkinson's Disease Rating Scale (UPDRS) on PC-GITA. The experimental results show: 1) the speech attribute features can well indicate the severity of pathologies in parkinsonian speech; 2) combining phonation features with articulatory features improves the PD assessment performance, but requires high-quality recordings to be applicable; 3) read speech leads to more accurate automatic ratings than the use of sustained vowels, but not if the amount of speech is limited to correspond to the sustained vowels in duration; and 4) jointly using data from several speech tasks can further improve the automatic PD assessment performance.
Yuanyuan Liu 0002, Mittapalle Kiran Reddy, Nelly Penttilä, Tiina Ihalainen, Paavo Alku, Okko Johannes Räsänen
IEEE ACM Trans. Audio Speech Lang. Process.6
2022 Unsupervised Audio-Caption Aligning Learns Correspondences Between Individual Sound Events and Textual Phrases
abstract
We investigate unsupervised learning of correspondences between sound events and textual phrases through aligning audio clips with textual captions describing the content of a whole audio clip. We align originally unaligned and unannotated audio clips and their captions by scoring the similarities between audio frames and words, as encoded by modality-specific encoders and using a ranking-loss criterion to optimize the model. After training, we obtain clip-caption similarity by averaging frame-word similarities and estimate event-phrase correspondences by calculating frame-phrase similarities. We evaluate the method with two cross-modal tasks: audio-caption retrieval, and phrase-based sound event detection (SED). Experimental results show that the proposed method can globally associate audio clips with captions as well as locally learn correspondences between individual sound events and textual phrases in an unsupervised manner.
Huang Xie, Okko Johannes Räsänen, Konstantinos Drossos, Tuomas Virtanen
ICASSP2
2022 Analysis of Self-Supervised Learning and Dimensionality Reduction Methods in Clustering-Based Active Learning for Speech Emotion Recognition
abstract
Peer reviewed
Einari Vaaras, Manu Airaksinen, Okko Johannes Räsänen
INTERSPEECH3
2021 Zero-Shot Audio Classification with Factored Linear and Nonlinear Acoustic-Semantic Projections
abstract
In this paper, we study zero-shot learning in audio classification through factored linear and nonlinear acoustic-semantic projections between audio instances and sound classes. Zero-shot learning in audio classification refers to classification problems that aim at recognizing audio instances of sound classes, which have no available training data but only semantic side information. In this paper, we address zero-shot learning by employing factored linear and nonlinear acoustic-semantic projections. We develop factored linear projections by applying rank decomposition to a bilinear model, and use nonlinear activation functions, such as tanh, to model the non-linearity between acoustic embeddings and semantic embeddings. Compared with the prior bilinear model, experimental results show that the proposed projection methods are effective for improving classification performance of zero-shot learning in audio classification.
Huang Xie, Okko Johannes Räsänen, Tuomas Virtanen
ICASSP2
2021 Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models
abstract
Systems that can find correspondences between multiple modalities, such as between speech and images, have great potential to solve different recognition and data analysis tasks in an unsupervised manner. This work studies multimodal learning in the context of visually grounded speech (VGS) models, and focuses on their recently demonstrated capability to extract spatiotemporal alignments between spoken words and the corresponding visual objects without ever been explicitly trained for object localization or word recognition. As the main contributions, we formalize the alignment problem in terms of an audiovisual alignment tensor that is based on earlier VGS work, introduce systematic metrics for evaluating model performance in aligning visual objects and spoken words, and propose a new VGS model variant for the alignment task utilizing cross-modal attention layer. We test our model and a previously proposed model in the alignment task using SPEECH-COCO captions coupled with MSCOCO images. We compare the alignment performance using our proposed evaluation metrics to the semantic retrieval task commonly used to evaluate VGS models. We show that cross-modal attention layer not only helps the model to achieve higher semantic cross-modal retrieval performance, but also leads to substantial improvements in the alignment performance between image object and spoken words.
Khazar Khorrami, Okko Johannes Räsänen
Interspeech2
2021 Automatic Analysis of the Emotional Content of Speech in Daylong Child-Centered Recordings from a Neonatal Intensive Care Unit
abstract
Researchers have recently started to study how the emotional speech heard by young infants can affect their developmental outcomes. As a part of this research, hundreds of hours of daylong recordings from preterm infants' audio environments were collected from two hospitals in Finland and Estonia in the context of so-called APPLE study. In order to analyze the emotional content of speech in such a massive dataset, an automatic speech emotion recognition (SER) system is required. However, there are no emotion labels or existing indomain SER systems to be used for this purpose. In this paper, we introduce this initially unannotated large-scale real-world audio dataset and describe the development of a functional SER system for the Finnish subset of the data. We explore the effectiveness of alternative state-of-the-art techniques to deploy a SER system to a new domain, comparing cross-corpus generalization, WGAN-based domain adaptation, and active learning in the task. As a result, we show that the best-performing models are able to achieve a classification performance of 73.4% unweighted average recall (UAR) and 73.2% UAR for a binary classification for valence and arousal, respectively. The results also show that active learning achieves the most consistent performance compared to the two alternatives.
Einari Vaaras, Sari Ahlqvist-Björkroth, Konstantinos Drossos, Okko Johannes Räsänen
Interspeech4
2021 Language-Independent Approach for Automatic Computation of Vowel Articulation Features in Dysarthric Speech Assessment
abstract
Imprecise vowel articulation can be observed in people with Parkinson's disease (PD). Acoustic features measuring vowel articulation have been demonstrated to be effective indicators of PD in its assessment. Standard clinical vowel articulation features of vowel working space area (VSA), vowel articulation index (VAI) and formants centralization ratio (FCR), are derived the first two formants of the three corner vowels /a/, /i/ and /u/. Conventionally, manual annotation of the corner vowels from speech data is required before measuring vowel articulation. This process is time-consuming. The present work aims to reduce human effort in clinical analysis of PD speech by proposing an automatic pipeline for vowel articulation assessment. The method is based on automatic corner vowel detection using a language universal phoneme recognizer, followed by statistical analysis of the formant data. The approach removes the restrictions of prior knowledge of speaking content and the language in question. Experimental results on a Finnish PD speech corpus demonstrate the efficacy and reliability of the proposed automatic method in deriving VAI, VSA, FCR and F2i/F2u (the second formant ratio for vowels /i/ and /u/). The automatically computed parameters are shown to be highly correlated with features computed with manual annotations of corner vowels. In addition, automatically and manually computed vowel articulation features have comparable correlations with experts' ratings on speech intelligibility, voice impairment and overall severity of communication disorder. Language-independence of the proposed approach is further validated on a Spanish PD database, PC-GITA, as well as on TORGO corpus of English dysarthric speech.
Yuanyuan Liu 0002, Nelly Penttilä, Tiina Ihalainen, Juulia Lintula, Rachel Convey, Okko Johannes Räsänen
IEEE ACM Trans. Audio Speech Lang. Process.6
2020 Measuring prosodic predictability in children's home language environments
Kyle MacDonald, Marisa Casillas, Okko Johannes Räsänen, Anne S. Warlaumont
CogSci3
2020 Unsupervised Discovery of Recurring Speech Patterns Using Probabilistic Adaptive Metrics
abstract
Unsupervised spoken term discovery (UTD) aims at finding recurring segments of speech from a corpus of acoustic speech data. One potential approach to this problem is to use dynamic time warping (DTW) to find well-aligning patterns from the speech data. However, automatic selection of initial candidate segments for the DTW-alignment and detection of "sufficiently good" alignments among those require some type of pre-defined criteria, often operationalized as threshold parameters for pair-wise distance metrics between signal representations. In the existing UTD systems, the optimal hyperparameters may differ across datasets, limiting their applicability to new corpora and truly low-resource scenarios. In this paper, we propose a novel probabilistic approach to DTW-based UTD named as PDTW. In PDTW, distributional characteristics of the processed corpus are utilized for adaptive evaluation of alignment quality, thereby enabling systematic discovery of pattern pairs that have similarity what would be expected by coincidence. We test PDTW on Zero Resource Speech Challenge 2017 datasets as a part of 2020 implementation of the challenge. The results show that the system performs consistently on all five tested languages using fixed hyperparameters, clearly outperforming the earlier DTW-based system in terms of coverage of the detected patterns.
Okko Johannes Räsänen, María Andrea Cruz Blandón
INTERSPEECH1
2019 Data Augmentation Strategies for Neural Network F0 Estimation
abstract
This study explores various speech data augmentation methods for the task of noise-robust fundamental frequency (F0) estimation with neural networks. The explored augmentation strategies are split into additive noise and channel-based augmentation and into vocoder-based augmentation methods. In vocoder-based augmentation, a glottal vocoder is used to enhance the accuracy of ground truth F0 used for training of the neural network, as well as to expand the training data diversity in terms of F0 patterns and vocal tract lengths of the talkers. Evaluations on the PTDB-TUG corpus indicate that noise and channel augmentation can be used to greatly increase the noise robustness of trained models, and that vocoder-based ground truth enhancement further increases model performance. For smaller datasets, vocoder-based diversity augmentation can also be used to increase performance. The best-performing proposed method greatly outperformed the compared F0 estimation methods in terms of noise robustness.
Manu Airaksinen, Lauri Juvela, Paavo Alku, Okko Johannes Räsänen
ICASSP4
2019 Cycle-consistent Adversarial Networks for Non-parallel Vocal Effort Based Speaking Style Conversion
abstract
Speaking style conversion (SSC) is the technology of converting natural speech signals from one style to another. In this study, we propose the use of cycle-consistent adversarial networks (CycleGANs) for converting styles with varying vocal effort, and focus on conversion between normal and Lombard styles as a case study of this problem. We propose a parametric approach that uses the Pulse Model in Log domain (PML) vocoder to extract speech features. These features are mapped using the CycleGAN from utterances in the source style to the corresponding features of target speech. Finally, the mapped features are converted to a Lombard speech waveform with the PML. The CycleGAN was compared in subjective listening tests with 2 other standard mapping methods used in conversion, and the CycleGAN was found to have the best performance in terms of speech quality and in terms of the magnitude of the perceptual change between the two styles.
Shreyas Seshadri, Lauri Juvela, Junichi Yamagishi, Okko Johannes Räsänen, Paavo Alku
ICASSP4
2019 A Computational Model of Early Language Acquisition from Audiovisual Experiences of Young Infants
abstract
Earlier research has suggested that human infants might use statistical dependencies between speech and non-linguistic multimodal input to bootstrap their language learning before they know how to segment words from running speech. However, feasibility of this hypothesis in terms of real-world infant experiences has remained unclear. This paper presents a step towards a more realistic test of the multimodal bootstrapping hypothesis by describing a neural network model that can learn word segments and their meanings from referentially ambiguous acoustic input. The model is tested on recordings of real infant-caregiver interactions using utterance-level labels for concrete visual objects that were attended by the infant when caregiver spoke an utterance containing the name of the object, and using random visual labels for utterances during absence of attention. The results show that beginnings of lexical knowledge may indeed emerge from individually ambiguous learning scenarios. In addition, the hidden layers of the network show gradually increasing selectivity to phonetic categories as a function of layer depth, resembling models trained for phone recognition in a supervised manner.
Okko Johannes Räsänen, Khazar Khorrami
INTERSPEECH1
2019 Augmented CycleGANs for Continuous Scale Normal-to-Lombard Speaking Style Conversion
abstract
Lombard speech is a speaking style associated with increased vocal effort that is naturally used by humans to improve intelligibility in the presence of noise. It is hence desirable to have a system capable of converting speech from normal to Lombard style. Moreover, it would be useful if one could adjust the degree of Lombardness in the converted speech so that the system is more adaptable to different noise environments. In this study, we propose the use of recently developed Augmented cycle-consistent adversarial networks (Augmented CycleGANs) for conversion between normal and Lombard speaking styles. The proposed system gives a smooth control on the degree of Lombardness of the mapped utterances by traversing through different points in the latent space of the trained model. We utilize a parametric approach that uses the Pulse Model in Log domain (PML) vocoder to extract features from normal speech that are then mapped to Lombard-style features using the Augmented CycleGAN. Finally, the mapped features are converted to Lombard speech with PML. The model is trained on multi-language data recorded in different noise conditions, and we compare its effectiveness to a previously proposed CycleGAN system in experiments for intelligibility and quality of mapped speech.
Shreyas Seshadri, Lauri Juvela, Paavo Alku, Okko Johannes Räsänen
INTERSPEECH4
2019 Automatic word count estimation from daylong child-centered recordings in various language environments using language-independent syllabification of speech
abstract
Automatic word count estimation (WCE) from audio recordings can be used to quantify the amount of verbal communication in a recording environment. One key application of WCE is to measure language input heard by infants and toddlers in their natural environments, as captured by daylong recordings from microphones worn by the infants. Although WCE is nearly trivial for high-quality signals in high-resource languages, daylong recordings are substantially more challenging due to the unconstrained acoustic environments and the presence of near- and far-field speech. Moreover, many use cases of interest involve languages for which reliable ASR systems or even well-defined lexicons are not available. A good WCE system should also perform similarly for low- and high-resource languages in order to enable unbiased comparisons across different cultures and environments. Unfortunately, the current state-of-the-art solution, the LENA system, is based on proprietary software and has only been optimized for American English, limiting its applicability. In this paper, we build on existing work on WCE and present the steps we have taken towards a freely available system for WCE that can be adapted to different languages or dialects with a limited amount of orthographically transcribed speech data. Our system is based on language-independent syllabification of speech, followed by a language-dependent mapping from syllable counts (and a number of other acoustic features) to the corresponding word count estimates. We evaluate our system on samples from daylong infant recordings from six different corpora consisting of several languages and socioeconomic environments, all manually annotated with the same protocol to allow direct comparison. We compare a number of alternative techniques for the two key components in our system: speech activity detection and automatic syllabification of speech. As a result, we show that our system can reach relatively consistent WCE accuracy across multiple corpora and languages (with some limitations). In addition, the system outperforms LENA on three of the four corpora consisting of different varieties of English. We also demonstrate how an automatic neural network-based syllabifier, when trained on multiple languages, generalizes well to novel languages beyond the training data, outperforming two previously proposed unsupervised syllabifiers as a feature extractor for WCE.
Okko Johannes Räsänen, Shreyas Seshadri, Julien Karadayi, Eric Riebling, John P. Bunce, Alejandrina Cristià, Florian Metze, Marisa Casillas, Celia Rosemberg, Elika Bergelson, Melanie Soderstrom
Speech Commun.1
2019 SylNet: An Adaptable End-to-End Syllable Count Estimator for Speech
abstract
Automatic syllable count estimation (SCE) is used in a variety of applications ranging from speaking rate estimation to detecting social activity from wearable microphones or developmental research concerned with quantifying speech heard by language-learning children in different environments. The majority of previously utilized SCE methods have relied on heuristic digital signal processing (DSP) methods, and only a small number of bi-directional long short-term memory (BLSTM) approaches have made use of modern machine learning approaches in the SCE task. This letter presents a novel end-to-end method called SylNet for automatic syllable counting from speech, built on the basis of a recent developments in neural network architectures. We describe how the entire model can be optimized directly to minimize SCE error on the training data without annotations aligned at the syllable level, and how it can be adapted to new languages using limited speech data with known syllable counts. Experiments on several different languages reveal that SylNet generalizes to languages beyond its training data and further improves with adaptation. It also outperforms several previously proposed methods for syllabification, including end-to-end BLSTMs.
Shreyas Seshadri, Okko Johannes Räsänen
IEEE Signal Process. Lett.2
2018 Time-regularized Linear Prediction for Noise-robust Extraction of the Spectral Envelope of Speech
abstract
Feature extraction of speech signals is typically performed in short-time frames by assuming that the signal is stationary within each frame. For the extraction of the spectral envelope of speech, which conveys the formant frequencies produced by the resonances of the slowly varying vocal tract, an often used frame length is within 20-30 ms. However, this kind of conventional frame-based spectral analysis is oblivious of the broader temporal context of the signal and is prone to degradation by, for example, environmental noise. In this paper, we propose a new frame-based linear prediction (LP) analysis method that includes a regularization term that penalizes energy differences in consecutive frames of an all-pole spectral envelope model. This integrates the slowly varying nature of the vocal tract as a part of the analysis. Objective evaluations related to feature distortion and phonetic representational capability were performed by studying the properties of the mel-frequency cepstral coefficient (MFCC) representations computed from different spectral estimation methods under noisy conditions using the TIMIT database. The results show that the proposed time-regularized LP approach exhibits superior MFCC distortion behavior while simultaneously having the greatest average separability of different phoneme categories in comparison to the other methods.
Manu Airaksinen, Lauri Juvela, Okko Johannes Räsänen, Paavo Alku
INTERSPEECH3
2018 Comparison of Syllabification Algorithms and Training Strategies for Robust Word Count Estimation across Different Languages and Recording Conditions
abstract
Word count estimation (WCE) from audio recordings has a number of applications, including quantifying the amount of speech that language-learning infants hear in their natural environments, as captured by daylong recordings made with devices worn by infants. To be applicable in a wide range of scenarios and also low-resource domains, WCE tools should be extremely robust against varying signal conditions and require minimal access to labeled training data in the target domain. For this purpose, earlier work has used automatic syllabification of speech, followed by a least-squares-mapping of syllables to word counts. This paper compares a number of previously proposed syllabifiers in the WCE task, including a supervised bi-directional long short-term memory (BLSTM) network that is trained on a language for which high quality syllable annotations are available (a “high resource language”), and reports how the alternative methods compare on different languages and signal conditions. We also explore additive noise and varying-channel data augmentation strategies for BLSTM training, and show how they improve performance in both matching and mismatching languages. Intriguingly, we also find that even though the BLSTM works on languages beyond its training data, the unsupervised algorithms can still outperform it in challenging signal conditions on novel languages.
Okko Johannes Räsänen, Shreyas Seshadri, Marisa Casillas
INTERSPEECH1
2018 Comparison of spectral tilt measures for sentence prominence in speech - Effects of dimensionality and adverse noise conditions
Sofoklis Kakouros, Okko Johannes Räsänen, Paavo Alku
Speech Commun.2
2017 Connecting stimulus-driven attention to the properties of infant-directed speech - Is exaggerated intonation also more surprising?
Okko Johannes Räsänen, Sofoklis Kakouros, Melanie Soderstrom
CogSci1
2017 Dirichlet process mixture models for clustering i-vector data
abstract
Non-parametric Bayesian methods have recently gained popularity in several research areas dealing with unsupervised learning. These models are capable of simultaneously learning the cluster models as well as their number based on properties of a dataset. The most commonly applied models are using Dirichlet process priors and Gaussian models, called as Dirichlet process Gaussian mixture models (DPGMMs). Recently, von Mises-Fisher mixture models (VMMs) have also been gaining popularity in modelling high-dimensional unit-normalized features such as text documents and gene expression data. VMMs are potentially more efficient in modeling certain speech representations such as i-vector data when compared to the GMM-based models, as they work with unit-normalized features based on cosine distance. The current work investigates the applicability of Dirichlet process VMMs (DPVMMs) for i-vector-based speaker clustering and verification, showing that they indeed show superior performance in comparison to DPGMMs in the tasks. In addition, we introduce an implementation of the DPVMMs with variational inference that is publicly available for use.
Shreyas Seshadri, Ulpu Remes, Okko Johannes Räsänen
ICASSP3
2017 Evaluation of Spectral Tilt Measures for Sentence Prominence Under Different Noise Conditions
abstract
Spectral tilt has been suggested to be a correlate of prominence in speech, although several studies have not replicated this empirically. This may be partially due to the lack of a standard method for tilt estimation from speech, rendering interpretations and comparisons between studies difficult. In addition, little is known about the performance of tilt estimators for prominence detection in the presence of noise. In this work, we investigate and compare several standard tilt measures on quantifying prominence in spoken Dutch and under different levels of additive noise. We also compare these measures with other acoustic correlates of prominence, namely, energy, F0, and duration. Our results provide further empirical support for the finding that tilt is a systematic correlate of prominence, at least in Dutch, even though energy, F0, and duration appear still to be more robust features for the task. In addition, our results show that there are notable differences between different tilt estimators in their ability to discriminate prominent words from non-prominent ones in different levels of noise.
Sofoklis Kakouros, Okko Johannes Räsänen, Paavo Alku
INTERSPEECH2
2017 Speaking Style Conversion from Normal to Lombard Speech Using a Glottal Vocoder and Bayesian GMMs
abstract
Speaking style conversion is the technology of converting natural speech signals from one style to another. In this study, we focus on normal-to-Lombard conversion. This can be used, for example, to enhance the intelligibility of speech in noisy environments. We propose a parametric approach that uses a vocoder to extract speech features. These features are mapped using Bayesian GMMs from utterances spoken in normal style to the corresponding features of Lombard speech. Finally, the mapped features are converted to a Lombard speech waveform with the vocoder. Two vocoders were compared in the proposed normal-to-Lombard conversion: a recently developed glottal vocoder that decomposes speech into glottal flow excitation and vocal tract, and the widely used STRAIGHT vocoder. The conversion quality was evaluated in two subjective listening tests measuring subjective similarity and naturalness. The similarity test results show that the system is able to convert normal speech into Lombard speech for the two vocoders. However, the subjective naturalness of the converted Lombard speech was clearly better using the glottal vocoder in comparison to STRAIGHT.
Ana Ramírez López, Shreyas Seshadri, Lauri Juvela, Okko Johannes Räsänen, Paavo Alku
INTERSPEECH4
2017 Comparison of Non-Parametric Bayesian Mixture Models for Syllable Clustering and Zero-Resource Speech Processing
abstract
Peer reviewed
Shreyas Seshadri, Ulpu Remes, Okko Johannes Räsänen
INTERSPEECH3
2017 An online model for vowel imitation learning
Heikki Rasilo, Okko Johannes Räsänen
Speech Commun.2
2016 Statistical Learning of Prosodic Patterns and Reversal of Perceptual Cues for Sentence Prominence
Sofoklis Kakouros, Okko Johannes Räsänen
CogSci2
2016 A Cognitive Approach to Modeling Sentence Level Prominence Based on Stimulus Unpredictability
Sofoklis Kakouros, Okko Johannes Räsänen
CogSci2
2016 Analyzing distributional learning of phonemic categories in unsupervised deep neural networks
Okko Johannes Räsänen, Tasha Nagamine, Nima Mesgarani
CogSci1
2016 Analyzing the Contribution of Top-Down Lexical and Bottom-Up Acoustic Cues in the Detection of Sentence Prominence
abstract
Copyright © 2016 ISCA. Recent work has suggested that prominence perception could be driven by the predictability of the acoustic prosodic features of speech. On the other hand, lexical predictability and part of speech information are also known to correlate with prominence. In this paper, we investigate how the bottom-up acoustic and top-down lexical cues contribute to sentence prominence by using both types of features in unsupervised and supervised systems for automatic prominence detection. The study is conducted using a corpus of Dutch continuous speech with manually annotated prominence labels. Our results show that unpredictability of speech patterns is a consistent and important cue for prominence at both the lexical and acoustic levels, and also that lexical predictability and part-of-speech information can be used as efficient features in supervised prominence classifiers.
Sofoklis Kakouros, Joris Pelemans, Lyan Verwimp, Patrick Wambacq, Okko Johannes Räsänen
INTERSPEECH5
2016 3PRO - An unsupervised method for the automatic detection of sentence prominence in speech
Sofoklis Kakouros, Okko Johannes Räsänen
Speech Commun.2
2016 Sequence Prediction With Sparse Distributed Hyperdimensional Coding Applied to the Analysis of Mobile Phone Use Patterns
abstract
Modeling and prediction of temporal sequences is central to many signal processing and machine learning applications. Prediction based on sequence history is typically performed using parametric models, such as fixed-order Markov chains ( n -grams), approximations of high-order Markov processes, such as mixed-order Markov models or mixtures of lagged bigram models, or with other machine learning techniques. This paper presents a method for sequence prediction based on sparse hyperdimensional coding of the sequence structure and describes how higher order temporal structures can be utilized in sparse coding in a balanced manner. The method is purely incremental, allowing real-time online learning and prediction with limited computational resources. Experiments with prediction of mobile phone use patterns, including the prediction of the next launched application, the next GPS location of the user, and the next artist played with the phone media player, reveal that the proposed method is able to capture the relevant variable-order structure from the sequences. In comparison with the n -grams and the mixed-order Markov models, the sparse hyperdimensional predictor clearly outperforms its peers in terms of unweighted average recall and achieves an equal level of weighted average recall as the mixed-order Markov chain but without the batch training of the mixed-order model.
Okko Johannes Räsänen, Jukka Saarinen
IEEE Trans. Neural Networks Learn. Syst.1
2015 Analyzing the Predictability of Lexeme-specific Prosodic Features as a Cue to Sentence Prominence
Sofoklis Kakouros, Okko Johannes Räsänen
CogSci2
2015 Generating Hyperdimensional Distributed Representations from Continuous-Valued Multivariate Sensory Input
Okko Johannes Räsänen
CogSci1
2015 Cross-situational cues are relevant for early word segmentation
Okko Johannes Räsänen, Heikki Rasilo
CogSci1
2015 Computational evidence for effects of memory decay, familiarity preference and mutual exclusivity in cross-situational learning
Heikki Rasilo, Okko Johannes Räsänen
CogSci2
2015 Automatic detection of sentence prominence in speech using predictability of word-level acoustic features
Sofoklis Kakouros, Okko Johannes Räsänen
INTERSPEECH2
2015 Unsupervised word discovery from speech using automatic segmentation into syllable-like units
Okko Johannes Räsänen, Gabriel Doyle, Michael C. Frank
INTERSPEECH1
2015 Weakly-supervised word learning is improved by an active online algorithm
Heikki Rasilo, Okko Johannes Räsänen
INTERSPEECH2
2015 Feature selection methods and their combinations in high-dimensional classification of speaker likability, intelligibility and personality traits
Jouni Pohjalainen, Okko Johannes Räsänen, Serdar Kadioglu
Comput. Speech Lang.2
2014 Statistical Unpredictability of F0 Trajectories as a Cue to Sentence Stress
Sofoklis Kakouros, Okko Johannes Räsänen
CogSci2
2014 Basic cuts revisited: Temporal segmentation of speech into phone-like units with statistical learning at a pre-linguistic level
Okko Johannes Räsänen
CogSci1
2014 Perception of sentence stress in English infant directed speech
abstract
Various studies have examined the acoustic features in infant directed speech (IDS) and adult directed speech (ADS). However, there are few speech corpora with prominence annotation from multiple listeners or analysis of the acoustic properties of the stressed versus unstressed words, most studies and corpora focusing on syllabic stress. In order to fill this gap, the current study analyzes the acoustic properties of sentence stress in a corpus of English IDS. More specifically, the work is one of the first analyzing IDS as perceived by adult listeners, providing inter-annotator agreement ratings and an analysis of the acoustic correlates of sentence stress with regard to the most important prosodic features encountered in the literature: fundamental frequency, intensity, word duration, and spectral tilt. The analysis shows that all of the analyzed features correlate with the perception of stress, indicating that the sentential prominence in IDS is conveyed by similar acoustic characteristics that are known to be relevant for stress perception in ADS.
Sofoklis Kakouros, Okko Johannes Räsänen
INTERSPEECH2
2014 Modeling Dependencies in Multiple Parallel Data Streams with Hyperdimensional Computing
abstract
This work presents an approach for modeling statistical dependencies in multivariate discrete sequences by using hyperdimensional random vectors. The system takes any number of parallel sequences as inputs and learns to predict the future states of these streams using the mutual dependencies between the inputs. Performance of the system is tested in an activity recognition task with data from multiple worn sensors. The results show that the approach outperforms the existing baseline results in the task and demonstrate that the system is capable to account for the varying reliability of different input streams.
Okko Johannes Räsänen, Sofoklis Kakouros
IEEE Signal Process. Lett.1
2013 Attention based temporal filtering of sensory signals for data redundancy reduction
abstract
Since modern computational devices are required to store and process increasing amounts of data generated from various sources, efficient algorithms for identification of significant information in the data are becoming essential. Sensory recordings are one example where automatic and continuous storing and processing of large amounts of data is needed. Therefore, algorithms that can alleviate the computational load of the devices and reduce their storage requirements by removing uninformative data are important. In this work we propose a method for data reduction based on theories of human attention. The method detects temporally salient events based on the context in which they occur and retains only those sections of the input signal. The algorithm is tested as a pre-processing stage in a weakly supervised keyword learning experiment where it is shown to significantly improve the quality of the codebooks used in the pattern discovery process.
Sofoklis Kakouros, Okko Johannes Räsänen, Unto K. Laine
ICASSP2
2013 Automatic self-supervised learning of associations between speech and text
abstract
One of the key challenges in artificial cognitive systems is to develop effective algorithms that learn without human supervision to understand qualitatively different realisations of the same abstraction and therefore also acquire an ability to transcribe a sensory data stream to completely different modality. This is also true in the so-called Big Data problem. Through learning of associations between multiple types of data of the same phenomenon, it is possible to capture hidden dynamics that govern processes that yielded the measured data. In this thesis, a methodological framework for automatic discovery of statistical associations between two qualitatively different data streams is proposed. The simulations are run on a noisy, high bit-rate, sensory signal (speech) and temporally discrete categorical data (text). In order to distinguish the approach from traditional automatic speech recognition systems, it does not utilize any phonetic or linguistic knowledge in the recognition. It merely learns statistically sound units of speech and text and their mutual mappings in an unsupervised manner. The experiments on child directed speech with limited vocabulary show that, after a period of learning, the method acquires a promising ability to transcribe continuous speech to its textual representation.
Juha Knuuttila, Okko Johannes Räsänen, Unto K. Laine
INTERSPEECH2
2013 Random subset feature selection in automatic recognition of developmental disorders, affective states, and level of conflict from speech
Okko Johannes Räsänen, Jouni Pohjalainen
INTERSPEECH1
2013 Feedback and imitation by a caregiver guides a virtual infant to learn native phonemes and the skill of speech inversion
Heikki Rasilo, Okko Johannes Räsänen, Unto K. Laine
Speech Commun.2
2012 Acoustic analysis supports the existence of a single distributional learning mechanism in structural rule learning from an artificial language
Okko Johannes Räsänen, Heikki Rasilo
CogSci1
2012 Hierarchical unsupervised discovery of user context from multivariate sensory data
abstract
A system capable for purely unsupervised learning of sensory context models is presented in this work. The system is based on discovery of short-term activity motifs from the sensory data and statistical analysis of these motifs on a larger time scale. Detected context segments are then clustered into high-level context categories and the data corresponding to these categories are used to train on-line classifiers for different contexts. Experiments show that the method is capable of segmenting sensory recordings into epochs of high-level environmental contexts based purely on audio signal, and that the classifiers trained from the obtained segments are selective towards specific contexts.
Okko Johannes Räsänen
ICASSP1
2012 Context induced merging of synonymous word models in computational modeling of early language acquisition
abstract
It has been shown that both infants and machines are able to discover recurring word-like patterns from continuous speech in the absence of supervision. However, these early models for words do not always generalize well across different acoustic variants of the same words. Instead, several parallel models for words or multiple fragments of a word are initially learned. In this work, we study a two-stage computational framework for refining the initially acquired representations of acoustic word patterns. In the first stage, the automatically discovered word patterns are studied in the context of visual word referents, enabling grounding of the word forms to the systematically co-occurring objects and actions in the environment. In the second stage, synonymy of the words is measured in terms of the similarity of their environmental contexts. The word models that share similar external context are merged together, producing a lexicon with a smaller number of parallel models for each word and with a greater generalization capability from each model towards new realizations of the word. The experimental results show that the context-based equivalence and merging of parallel models leads to a more compact and higher quality lexicon than a learning process based purely on acoustic similarities.
Okko Johannes Räsänen
ICASSP1
2012 Feature Selection for Speaker Traits
abstract
This study focuses on handling high-dimensional classification problems by means of feature selection. The data sets used are provided by the organizers of the Interspeech 2012 Speaker Trait Challenge. A combination of two feature selection approaches gives results that approach or exceed the challenge baselines using a knearest-neighbor classifier. One of the feature selection methods is based on covering the data set with correct unsupervised or supervised classifications according to individual features. The other selection method applies a measure of statistical dependence between discretized features and class labels. Index Terms: pattern recognition, feature selection, high-dimensional data, speaker characteristics
Jouni Pohjalainen, Serdar Kadioglu, Okko Johannes Räsänen
INTERSPEECH3
2012 Non-auditory cognitive capabilities in computational modeling of early language acquisition
abstract
Computational models of early language acquisition (LA) play an important role in understanding the acquisition and processing of spoken language. Since language is an extremely complex phenomenon, computational studies typically address only a specific aspect of the LA at a time. This calls for a huge number of assumptions regarding the other cognitive processes of the learning system, and these assumptions can have significant consequences to the ecological plausibility of the simulations. In this paper, we review the developmental status of a number of cognitive processes during the first year of infant’s life that are typically involved in the computational simulations of LA. How these findings are related to the plausibility of different simplifications and assumptions in computational models are also discussed.
Okko Johannes Räsänen
INTERSPEECH1
2012 Average Spectrotemporal Structure of Continuous Speech Matches with the Frequency Resolution of Human Hearing
Okko Johannes Räsänen
INTERSPEECH1
2012 Modeling spoken language acquisition with a generic cognitive architecture for associative learning
abstract
Human neo-cortex can be viewed as a modality invariant system for pattern discovery and associative learning. Similarly, research in the field of distributional learning suggests that much of human language acquisition can be explained by generic statistical learning mechanisms. The current paper argues that pattern processing capabilities of the human brain can be better understood if the process of early language acquisition is modeled using an entire cognitive architecture capable of unsupervised pattern discovery and associative learning. A highlevel motivation and description for generic processing principles in such architecture are given, followed by examples of our current work in the field.
Okko Johannes Räsänen, Heikki Rasilo, Unto K. Laine
INTERSPEECH1
2012 A method for noise-robust context-aware pattern discovery and recognition from categorical sequences
Okko Johannes Räsänen, Unto K. Laine
Pattern Recognit.1
2012 Computational modeling of phonetic and lexical learning in early language acquisition: Existing models and future directions
Okko Johannes Räsänen
Speech Commun.1
2011 Method for Speech Inversion with Large Scale Statistical Evaluation
abstract
Abstract An articulatory model of speech production is created for the purpose of studying the links between speech production and perception. A computationally effective method for speech inversion in proposed, using a two-pole predictor structure in order to maintain better articulatory dynamics when compared to conventional dynamic programming methods. Preliminary tests for the effect of inversion are performed for 2500 Finnish syllables extracted from continuous speech, consisting of 125 different syllable classes. A cluster selectivity test shows that the syllables are more reliably clustered using the automatically obtained parametric representation of articulatory gestures rather than the original formant representation that is used as a starting point for the inversion. Index Terms: Articulatory model, speech inversion, motor theory, vocal tract 1. Introduction to articulatory modeling Speech events are more conveniently described in articulatory than acoustic sense. Individual articulators move rather slowly and smoothly when compared to spectral characteristics of speech signals. Since the relative trajectories of different articulators remain rather similar in the production of speech sounds regardless of the speaker, the modeling of speech perception with articulatory modeling may help to overcome many of the problems that arise from the ambiguity in the purely acoustic domain. In the 19th century research on the area of articulatory modeling boomed when the first electrical and then digital models for speech production could be implemented. Researchers have often referred to articulatory models developed by Coker [1], Mermelstein [2] or Maeda [3], for example when studying the speech inverse problem. Maeda’s model’s seven articulatory parameters were estimated from x-ray tracings using so-called arbitrary factor analysis in order to have the parameters maximally uncorrelated to each other. Mermelstein’s geometrical articulatory model depicts the positions of articulators in the midsagittal plane. Lips, jaw, tongue, velum and hyoid are considered as movable structures. In 1990’s and 2000’s more complex vocal tract and tongue models were developed. E.g. Dang and Honda have created a 3D articulatory model which used physiological constraints typical to human articulation in inverting vowel-to-vowel sequences [4].
Heikki Rasilo, Unto K. Laine, Okko Johannes Räsänen, Toomas Altosaar
INTERSPEECH3
2010 Fully unsupervised word learning from continuous speech using transitional probabilities of atomic acoustic events
abstract
This work presents a learning algorithm based on transitional probabilities of atomic acoustic events (vector quantized spectral features). The algorithm learns models for word-like units in speech without any supervision, and without a priori knowledge of phonemic or linguistic units. The learned models can be used to segment novel utterances into word-like units, supporting the theory that transitional probabilities of acoustic events could work as a bootstrapping mechanism of language learning. The performance of the algorithm is evaluated using a corpus of Finnish infant-directed speech.
Okko Johannes Räsänen
INTERSPEECH1
2010 Estimation studies of vocal tract shape trajectory using a variable length and lossy kelly-lochbaum model
abstract
This work demonstrates the use of a modified KellyLochbaum (KL) vocal tract (VT) model in dynamic mapping from speech signals to articulatory configurations. The sixteen section KL model is equipped with a variable length segment for lip rounding and an accurate model for lip radiation impedance. Profiles for the eight Finnish vowels are used to form so called anchor points in the articulatory and spectral domain. These profiles are modulated by cosine functions to produce clusters of vowel variants around the anchor points, leading to the filling of the vowel triangle with over 189000 variants. The resulting profile and formant frequency data are stored in a codebook that is used in the trajectory estimation task, proposing a number of profile candidates for each speech frame based on the observed formant frequencies. The final trajectory is estimated by minimizing the articulatory distance across all frames. The first trajectory estimation results are promising and in good balance with the present phonetic literature.
Heikki Rasilo, Unto K. Laine, Okko Johannes Räsänen
INTERSPEECH3
2009 Discovering keywords from cross-modal input: ecological vs. engineering methods for enhancing acoustic repetitions
abstract
\n Contains fulltext :\n 76399.pdf (author's version ) (Open Access)\n
Guillaume Aimetti, Roger K. Moore, Louis ten Bosch, Okko Johannes Räsänen, Unto K. Laine
INTERSPEECH4
2009 Do multiple caregivers speed up language acquisition?
abstract
In this paper we compare three different implementations of language learning to investigate the issue of speaker-dependent initial representations and subsequent generalization. These implementations are used in a comprehensive model of lan-guage acquisition under development in the FP6 FET project ACORNS. All algorithms are embedded in a cognitively and ecologically plausible framework, and perform the task of de-tecting word-like units without any lexical, phonetic, or phono-logical information. The results show that the computational approaches differ with respect to the extent they deal with un-seen speakers, and how generalization depends on the variation observed during training. Index Terms: Language acquisition, Computational modeling 1.
Louis ten Bosch, Okko Johannes Räsänen, Joris Driesen, Guillaume Aimetti, Toomas Altosaar, Lou Boves, A. Corns
INTERSPEECH2
2009 Self-learning vector quantization for pattern discovery from speech
abstract
A novel and computationally straightforward clustering algorithm was developed for vector quantization (VQ) of speech signals for a task of unsupervised pattern discovery (PD) from speech. The algorithm works in purely incremental mode, is computationally extremely feasible, and achieves comparable classification quality with the well-known k-means algorithm in the PD task. In addition to presenting the algorithm, general findings regarding the relationship between the amounts of training material, convergence of the clustering algorithm, and the ultimate quality of VQ codebooks are discussed. Index Terms: speech recognition, pattern discovery, time series analysis, vector quantization, data clustering
Okko Johannes Räsänen, Unto K. Laine, Toomas Altosaar
INTERSPEECH1
2009 An improved speech segmentation quality measure: the r-value
abstract
Phone segmentation in ASR is usually performed indirectly by Viterbi decoding of HMM output. Direct approaches also exist, e.g., blind speech segmentation algorithms. In either case, performance of automatic speech segmentation algorithms is often measured using automated evaluation algorithms and used to optimize a segmentation system’s performance. However, evaluation approaches reported in literature were found to be lacking. Also, we have determined that increases in phone boundary location detection rates are often due to increased over-segmentation levels and not to algorithmic improvements, i.e., by simply adding random boundaries a better hit-rate can be achieved when using current quality measures. Since established measures were found to be insensitive to this type of random boundary insertion, a new R-value quality measure is introduced that indicates how close a segmentation algorithm’s performance is to an ideal point of operation. Index terms: blind speech segmentation, segmentation evaluation. 1.
Okko Johannes Räsänen, Unto K. Laine, Toomas Altosaar
INTERSPEECH1
2009 A noise robust method for pattern discovery in quantized time series: the concept matrix approach
abstract
An efficient method for pattern discovery from discrete time series is introduced in this paper. The method utilizes two parallel streams of data, a discrete unit time-series and a set of labeled events, From these inputs it builds associative models between systematically co-occurring structures existing in both streams. The models are based on transitional probabilities of events at several different time scales. Learning and recognition processes are incremental, making the approach suitable for online learning tasks. The capabilities of the algorithm are demonstrated in a continuous speech recognition task operating in varying noise levels.
Okko Johannes Räsänen, Unto K. Laine, Toomas Altosaar
INTERSPEECH1
2008 Computational language acquisition by statistical bottom-up processing
abstract
Statistical learning of patterns from perceptual input is an increasingly central topic in cognitive processing including human language acquisition. We present an unsupervised computational method for statistical word learning by analysis of transitional probabilities of subsequent phone pairs. Results indicate that word differentiation is possible with this type of approach and are in line with previous behavioral findings. Index Terms: computational language acquisition, speech segmentation, speech clustering, statistical learning
Okko Johannes Räsänen, Unto K. Laine, Toomas Altosaar
INTERSPEECH1