EDBT 2026 Demo / reviewers in the wild / expert
Paavo Alku
dblp:99/2726
· DBLP profile ↗
216ranked-venue papers
32as first author
32since 2021 · last 2026
0000-0002-8173-9418ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 179 · 29 first-author · 19 since 2021Artificial intelligence and machine learning · 130 · 21 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Classification of phonation types in singing and speaking voice using self-supervised learning models
Prathamesh Parasharam Patil, Mittapalle Kiran Reddy, Paavo Alku |
Speech Commun. | 3 |
| 2025 | Wavelet Scattering Network Features for Intensity Category Classification and Prediction of SPL from SpeechabstractSpeakers change vocal intensity in daily life to communicate over long distances and to express vocal emotions. Humans produce speech using different intensity categories (e.g. soft, normal and loud voice) and they can regulate intensity across a wide sound pressure level (SPL) range. Knowing the intensity category or the SPL of speech is beneficial in speech-based biomarking of health. Recent studies have explored the vocal intensity category classification and prediction of SPL from speech, which has been recorded without SPL calibration information and is presented on an arbitrary amplitude scale. Using speech signals in such scenario, this study investigates the wavelet scattering network (WSN) features in two tasks: (1) classification of speech into four intensity categories (soft, normal, loud, very loud) (multi-class classification task) and (2) prediction of SPL (regression task). In the former task, the WSN features showed absolute accuracy improvements of 4-14% compared to reference features. For the latter task, the WSN features improved the prediction of SPL by an average of 1-2 dB compared to the reference features. Manila Kodali, Sudarsana Reddy Kadiri, Shri Narayanan, Paavo Alku |
ICASSP | 4 |
| 2025 | Automatic classification of vocal intensity categories from amplitude-normalized speech signals by comparing acoustic features and classifier modelsabstractRegulation of vocal intensity is a fundamental phenomenon in speech communication. Speakers use different intensity categories (e.g., soft, normal, and loud voice) to generate different vocal emotions or to communicate in noisy conditions or over varying distances. Vocal intensity categories have been studied in fundamental research of speech, but much less is known about their automatic classification. This study investigates the classification of vocal intensity categories from speech signals in a scenario, where the original level information of speech is absent and the signal is presented on a normalized amplitude scale. Different acoustic features were studied together with machine learning (ML) and deep learning (DL) classifiers using two different labeling approaches. Speech signals recorded from 50 speakers reciting sentences in four intensity categories (soft, normal, loud, and very loud) were analyzed. Altogether 15 feature sets including different cepstral, spectral and handcrafted (eGeMAPS) features were compared. Three ML classifiers (support vector machine, random forest and AdaBoost), and four DL classifiers (deep neural network, convolutional neural network, recurrent neural network and bidirectional long short-term memory network) were compared. The best classification accuracy of 86.0% was obtained by combining the best performing cepstral and spectral features and using the bidirectional long short-term memory classifier. • Multi-class classification of vocal intensity categories is studied. • The classification is studied using amplitude-normalized speech signals. • Various labelling approaches, features and classifiers are compared. • DL models outperformed ML models, with BiLSTM achieving the best performance. Manila Kodali, Luna Ansari, Sudarsana Reddy Kadiri, Shri Narayanan, Paavo Alku |
Speech Commun. | 5 |
| 2025 | Towards robust heart failure detection in digital telephony environments by utilizing transformer-based codec inversionabstractThis study introduces the Codec Transformer Network (CTN) to enhance the reliability of automatic heart failure (HF) detection from coded telephone speech by addressing codec-related challenges in digital telephony. The study specifically addresses the codec mismatch between training and inference in HF detection. CTN is designed to map the mel-spectrogram representations of encoded speech signals back to their original, non-encoded forms, thereby recovering HF-related discriminative information. The effectiveness of CTN is demonstrated in conjunction with three HF detectors, based on Support Vector Machine, Random Forest, and K-Nearest Neighbors classifiers. The results show that CTN effectively retrieves the discriminative information between patients and controls, and performs comparably to or better than a baseline approach, based on multi-condition training. Saska Tirronen, Farhad Javanmardi, Hilla Pohjalainen, Sudarsana Reddy Kadiri, Mittapalle Kiran Reddy, Pyry Helkkula, Kasimir Kaitue, Mikko Minkkinen, Heli Tolppanen, Tuomo Nieminen, Paavo Alku |
Speech Commun. | 11 |
| 2024 | Fine-tuning of Pre-trained Models for Classification of Vocal Intensity Category from Speech SignalsabstractSpeakers regulate vocal intensity on many occasions for example to be heard over a long distance or to express vocal emotions. Humans can regulate vocal intensity over a wide sound pressure level (SPL) range and therefore speech can be categorized into different vocal intensity categories. Recent machine learning experiments have studied classification of vocal intensity category from speech signals which have been recorded without SPL information and which are represented on arbitrary amplitude scales. By fine-tuning four pre-trained models (wav2vec2-BASE, wav2vec2-LARGE, HuBERT, audio speech transformers), this paper studies classification of speech into four intensity categories (soft, normal, loud, very loud), when speech is presented on such arbitrary amplitude scale. The fine-tuned model embeddings showed absolute improvements of 5% and 10-12% in accuracy compared to baselines for the target intensity category label and the SPL-based intensity category label, respectively. Manila Kodali, Sudarsana Reddy Kadiri, Paavo Alku |
INTERSPEECH | 3 |
| 2024 | A comparison of data augmentation methods in voice pathology detectionabstractTo distinguish pathological voices from healthy voices, automatic voice pathology detection systems can be built using machine learning (ML) and deep learning (DL) techniques. To fully exploit such systems, large quantities of training data are typically required. The amount of training data is, however, small in the area of pathological voice, and therefore data augmentation (DA) becomes a potential technology to artificially increase the quantity of training data. This study presents a systematic comparison between various DA methods in the detection of pathological voice, including three time domain methods (noise addition, pitch shifting and time stretching), one time-frequency domain method (SpecAugment), and two vocoder-based methods (harmonic-to-noise ratio (HNR) modification and glottal pulse length modification). Detection systems were built using four popular spectral feature representations (static mel-frequency cepstral coefficients (MFCCs), dynamic MFCCs, spectrogram and mel-spectrogram). As classifiers, two widely used ML models (support vector machine (SVM) and random forest (RF)) and two DL models (long short-term memory (LSTM) network and convolutional neural network (CNN) with 1-dimensional (1-D) and 2-dimensional (2-D) architectures) were used. These systems were trained using a small number of training samples from two popular databases of pathological voice (HUPA and SVD) to find the best feature/classifier combination for each database. As a result, one ML-based detection system (mel-spectrogram/SVM for HUPA and SVD) and two DL-based detection systems (dynamic MFCCs/2-D CNN for HUPA and mel-spectrogram/2-D CNN for SVD) were selected for the comparison of the DA methods. The results show that by using DA in the system training, detection accuracy increased compared to the baseline systems that were trained without using DA. This improvement in accuracy was, however, clearly larger for the 2D-CNN system than for the SVM system. Furthermore, all six DA methods improved accuracy of the 2-D CNN system compared to the baseline system for both databases. The highest improvements were achieved using the time-frequency domain SpecAugment DA method, which improved accuracy by 1.5% and 3.8% (absolute) for the HUPA and SVD database, respectively. Farhad Javanmardi, Sudarsana Reddy Kadiri, Paavo Alku |
Comput. Speech Lang. | 3 |
| 2024 | Investigation of self-supervised pre-trained models for classification of voice quality from speech and neck surface accelerometer signalsabstractPrior studies in the automatic classification of voice quality have mainly studied the use of the acoustic speech signal as input. Recently, a few studies have been carried out by jointly using both speech and neck surface accelerometer (NSA) signals as inputs, and by extracting mel-frequency cepstral coefficients (MFCCs) and glottal source features. This study examines simultaneously-recorded speech and NSA signals in the classification of voice quality (breathy, modal, and pressed) using features derived from three self-supervised pre-trained models (wav2vec2-BASE, wav2vec2-LARGE, and HuBERT) and using a support vector machine (SVM) as well as convolutional neural networks (CNNs) as classifiers. Furthermore, the effectiveness of the pre-trained models is compared in feature extraction between glottal source waveforms and raw signal waveforms for both speech and NSA inputs. Using two signal processing methods (quasi-closed phase (QCP) glottal inverse filtering and zero frequency filtering (ZFF)), glottal source waveforms are estimated from both speech and NSA signals. The study has three main goals: (1) to study whether features derived from pre-trained models improve classification accuracy compared to conventional features (spectrogram, mel-spectrogram, MFCCs, i-vector, and x-vector), (2) to investigate which of the two modalities (speech vs. NSA) is more effective as input in the classification task with pre-trained model-based features, and (3) to evaluate whether the deep learning-based CNN classifier can enhance the classification accuracy in comparison to the SVM classifier. The results revealed that the use of the NSA input showed better classification performance compared to the speech signal. Between the features, the pre-trained model-based features showed better classification accuracies, both for speech and NSA inputs compared to the conventional features. The two classifiers performed equally well for all the pre-trained model-based features for both speech and NSA signals. It was also found that the HuBERT features performed better than the wav2vec2-BASE and wav2vec2-LARGE features for both speech and NSA inputs. In particular, when compared to the conventional features, the HuBERT features showed an absolute accuracy improvement of 3%–6% for speech and NSA signals in the classification of voice quality. Sudarsana Reddy Kadiri, Farhad Javanmardi, Paavo Alku |
Comput. Speech Lang. | 3 |
| 2024 | Automatic classification of the severity level of Parkinson's disease: A comparison of speaking tasks, features, and classifiersabstractAutomatic speech-based severity level classification of Parkinson’s disease (PD) enables objective assessment and earlier diagnosis. While many studies have been conducted on the binary classification task to distinguish speakers in PD from healthy controls (HCs), clearly fewer studies have addressed multi-class PD severity level classification problems. Furthermore, in studying the three main issues of speech-based classification systems—speaking tasks, features, and classifiers—previous investigations on the severity level classification have yielded inconclusive results due to the use of only a few, and sometimes just one, type of speaking task, feature, or classifier in each study. Hence, a systematic comparison is conducted in this study between different speaking tasks, features, and classifiers. Five speaking tasks (vowel task, sentence task, diadochokinetic (DDK) task, read text task, and monologue task), four features (phonation, articulation, prosody, and their fusion), and four classifier architectures (support vector machine (SVM), random forest (RF), multilayer perceptron (MLP), and AdaBoost) were compared. The classification task studied was a 3-class problem to classify PD severity level as healthy vs. mild vs. severe. Two MDS-UPDRS scales (MDS-UPDRS-III and MDS-UPDRS-S) were used for the ground truth severity level labels. The results showed that the use of the monologue task and the articulation and fusion of features improved classification accuracy significantly compared to the use of the other speaking tasks and features. The best classification systems resulted in a rate of accuracy of 58% (using the monologue task with the articulation features) for the MDS-UPDR-III scale and 56% (using the monologue task with fusion of features) for the MDS-UPDRS-S scale. Manila Kodali, Sudarsana Reddy Kadiri, Paavo Alku |
Comput. Speech Lang. | 3 |
| 2024 | AVID: A speech database for machine learning studies on vocal intensityabstractVocal intensity, which is quantified typically with the sound pressure level (SPL), is a key feature of speech. To measure SPL from speech recordings, a standard calibration tone (with a reference SPL of 94 dB or 114 dB) needs to be recorded together with speech. However, most of the popular databases that are used in areas such as speech and speaker recognition have been recorded without calibration information by expressing speech on arbitrary amplitude scales. Therefore, information about vocal intensity of the recorded speech, including SPL, is lost. In the current study, we introduce a new open and calibrated speech/electroglottography (EGG) database named Aalto Vocal Intensity Database (AVID). AVID includes speech and EGG produced by 50 speakers (25 males, 25 females) who varied their vocal intensity in four categories (soft, normal, loud and very loud). Recordings were conducted using a constant mouth-to-microphone distance and by recording a calibration tone. The speech data was labelled sentence-wise using a total of 19 labels that support the utilisation of the data in machine learing (ML) -based studies of vocal intensity based on supervised learning. In order to demonstrate how the AVID data can be used to study vocal intensity, we investigated one multi-class classification task (classification of speech into soft, normal, loud and very loud intensity classes) and one regression task (prediction of SPL of speech). In both tasks, we deliberately warped the level of the input speech by normalising the signal to have its maximum amplitude equal to 1.0, that is, we simulated a scenario that is prevalent in current speech databases. The results show that using the spectrogram feature with the support vector machine classifier gave an accuracy of 82% in the multi-class classification of the vocal intensity category. In the prediction of SPL, using the spectrogram feature with the support vector regressor gave an mean absolute error of about 2 dB and a coefficient of determination of 92%. We welcome researchers interested in classification and regression problems to utilise AVID in the study of vocal intensity, and we hope that the current results could serve as baselines for future ML studies on the topic. Paavo Alku, Manila Kodali, Laura Laaksonen, Sudarsana Reddy Kadiri |
Speech Commun. | 1 |
| 2024 | Pre-trained models for detection and severity level classification of dysarthria from speechabstractAutomatic detection and severity level classification of dysarthria from speech enables non-invasive and effective diagnosis that helps clinical decisions about medication and therapy of patients. In this work, three pre-trained models (wav2vec2-BASE, wav2vec2-LARGE, and HuBERT) are studied to extract features to build automatic detection and severity level classification systems for dysarthric speech. The experiments were conducted using two publicly available databases (UA-Speech and TORGO). One machine learning-based model (support vector machine, SVM) and one deep learning-based model (convolutional neural network, CNN) was used as the classifier. In order to compare the performance of the wav2vec2-BASE, wav2vec2-LARGE, and HuBERT features, three popular acoustic feature sets, namely, mel-frequency cepstral coefficients (MFCCs), openSMILE and extended Geneva minimalistic acoustic parameter set (eGeMAPS) were considered. Experimental results revealed that the features derived from the pre-trained models outperformed the three baseline features. It was also found that the HuBERT features performed better than the wav2vec2-BASE and wav2vec2-LARGE features. In particular, when compared to the best-performing baseline feature (openSMILE), the HuBERT features showed in the detection problem absolute accuracy improvements that varied between 1.33% (the SVM classifier, the TORGO database) and 2.86% (the SVM classifier, the UA-Speech database). In the severity level classification problem, the HuBERT features showed absolute accuracy improvements that varied between 6.54% (the SVM classifier, the TORGO database) and 10.46% (the SVM classifier, the UA-Speech database) compared to the best-performing baseline feature (eGeMAPS). Farhad Javanmardi, Sudarsana Reddy Kadiri, Paavo Alku |
Speech Commun. | 3 |
| 2024 | Automatic classification of neurological voice disorders using wavelet scattering featuresabstractNeurological voice disorders are caused by problems in the nervous system as it interacts with the larynx. In this paper, we propose to use wavelet scattering transform (WST)-based features in automatic classification of neurological voice disorders. As a part of WST, a speech signal is processed in stages with each stage consisting of three operations–convolution, modulus and averaging–to generate low-variance data representations that preserve discriminability across classes while minimizing differences within a class. The proposed WST-based features were extracted from speech signals of patients suffering from either spasmodic dysphonia (SD) or recurrent laryngeal nerve palsy (RLNP) and from speech signals of healthy speakers of the Saarbruecken voice disorder (SVD) database. Two machine learning algorithms (support vector machine (SVM) and feed forward neural network (NN)) were trained separately using the WST-based features, to perform two binary classification tasks (healthy vs. SD and healthy vs. RLNP) and one multi-class classification task (healthy vs. SD vs. RLNP). The results show that WST-based features outperformed state-of-the-art features in all three tasks. Furthermore, the best overall classification performance was achieved by the NN classifier trained using WST-based features. Yagnavajjula Madhu Keerthana, Mittapalle Kiran Reddy, Paavo Alku, K. Sreenivasa Rao, Pabitra Mitra |
Speech Commun. | 3 |
| 2024 | Exploring the Impact of Fine-Tuning the Wav2vec2 Model in Database-Independent Detection of Dysarthric SpeechabstractMany acoustic features and machine learning models have been studied to build automatic detection systems to distinguish dysarthric speech from healthy speech. These systems can help to improve the reliability of diagnosis. However, speech recorded for diagnosis in real-life clinical conditions can differ from the training data of the detection system in terms of, for example, recording conditions, speaker identity, and language. These mismatches may lead to a reduction in detection performance in practical applications. In this study, we investigate the use of the wav2vec2 model as a feature extractor together with a support vector machine (SVM) classifier to build automatic detection systems for dysarthric speech. The performance of the wav2vec2 features is evaluated in two cross-database scenarios, language-dependent and language-independent, to study their generalizability to unseen speakers, recording conditions, and languages before and after fine-tuning the wav2vec2 model. The results revealed that the fine-tuned wav2vec2 features showed better generalization in both scenarios and gave an absolute accuracy improvement of 1.46%-8.65% compared to the non-fine-tuned wav2vec2 features. Farhad Javanmardi, Sudarsana Reddy Kadiri, Paavo Alku |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Wav2vec-Based Detection and Severity Level Classification of Dysarthria From SpeechabstractAutomatic detection and severity level classification of dysarthria directly from acoustic speech signals can be used as a tool in medical diagnosis. In this work, the pre-trained wav2vec 2.0 model is studied as a feature extractor to build detection and severity level classification systems for dysarthric speech. The experiments were carried out with the popularly used UA-speech database. In the detection experiments, the results revealed that the best performance was obtained using the embeddings from the first layer of the wav2vec model that yielded an absolute improvement of 1.23% in accuracy compared to the best performing baseline feature (spectrogram). In the studied severity level classification task, the results revealed that the embeddings from the final layer gave an absolute improvement of 10.62% in accuracy compared to the best baseline features (mel-frequency cepstral coefficients). Farhad Javanmardi, Saska Tirronen, Manila Kodali, Sudarsana Reddy Kadiri, Paavo Alku |
ICASSP | 5 |
| 2023 | Automatic Classification of Vocal Intensity Category from SpeechabstractRegulation of vocal intensity is a fundamental phenomenon in speech communication. Vocal intensity can be quantified using sound pressure level (SPL), which can be measured easily by recording a standard calibration signal with speech and by comparing the energy of the recorded speech signal with that of the calibration tone. Unfortunately, speech recordings are mostly conducted without the SPL calibration signal, and speech signals are saved to databases using arbitrary amplitude scales. Therefore, neither the SPL nor the intensity category (e.g. soft or loud phonation) of a saved speech signal can be determined afterwards. Even though the original level information of speech is lost when the signal is presented on arbitrary amplitude scales, the speech signal contains other acoustic cues of vocal intensity. In the current study, we study machine learning and deep learning -based methods in automatic classification of vocal intensity category when the input speech is expressed using an arbitrary amplitude scale. A new gender-balanced database consisting of speech produced in four vocal intensity categories (soft, normal, loud, and very loud) was first recorded. Support vector machine and deep neural network (DNN) models were used to develop automatic classification systems using spectrograms, mel-spectrograms, and mel-frequency cepstral coefficients as features. The DNN classifier using the mel-spectrogram showed the best classification accuracy of about 90%. The database is made publicly available at https://bit.ly/3tLPGRx. Manila Kodali, Sudarsana Reddy Kadiri, Laura Laaksonen, Paavo Alku |
ICASSP | 4 |
| 2023 | Utilizing Wav2Vec In Database-Independent Voice Disorder DetectionabstractAutomatic detection of voice disorders from acoustic speech signals can help to improve reliability of medical diagnosis. However, the real-life environment in which speech signals are recorded for diagnosis can be different from the environment in which the detection system’s training data was originally collected. This mismatch between the recording conditions can decrease detection performance in practical scenarios. In this work, we propose to use a pre-trained wav2vec 2.0 model as a feature extractor to build automatic detection systems for voice disorders. The embeddings from the first layers of the context network contain information about phones, and these features are useful in voice disorder detection. We evaluate the performance of the wav2vec features in single-database and crossdatabase scenarios to study their generalizability to unseen speakers and recording conditions. The results indicate that the wav2vec features generalize better than popular spectral and cepstral baseline features. Saska Tirronen, Farhad Javanmardi, Manila Kodali, Sudarsana Reddy Kadiri, Paavo Alku |
ICASSP | 5 |
| 2023 | Severity Classification of Parkinson's Disease from Speech using Single Frequency Filtering-based FeaturesabstractDeveloping objective methods for assessing the severity of Parkinson's disease (PD) is crucial for improving the diagnosis and treatment. This study proposes two sets of novel features derived from the single frequency filtering (SFF) method: (1) SFF cepstral coefficients (SFFCC) and (2) MFCCs from the SFF (MFCC-SFF) for the severity classification of PD. Prior studies have demonstrated that SFF offers greater spectro-temporal resolution compared to the short-time Fourier transform. The study uses the PC-GITA database, which includes speech of PD patients and healthy controls produced in three speaking tasks (vowels, sentences, text reading). Experiments using the SVM classifier revealed that the proposed features outperformed the conventional MFCCs in all three speaking tasks. The proposed SFFCC and MFCC-SFF features gave a relative improvement of 5.8% and 2.3% for the vowel task, 7.0% & 1.8% for the sentence task, and 2.4% and 1.1% for the read text task, in comparison to MFCC features. Sudarsana Reddy Kadiri, Manila Kodali, Paavo Alku |
INTERSPEECH | 3 |
| 2023 | Classification of Vocal Intensity Category from Speech using the Wav2vec2 and Whisper EmbeddingsabstractIn speech communication, talkers regulate vocal intensity resulting in speech signals of different intensity categories (e.g., soft, loud). Intensity category carries important information about the speaker's health and emotions. However, many speech databases lack calibration information, and therefore sound pressure level cannot be measured from the recorded data. Machine learning, however, can be used in intensity category classification even though calibration information is not available. This study investigates pre-trained model embeddings (Wav2vec2 and Whisper) in classification of vocal intensity category (soft, normal, loud, and very loud) from speech signals expressed using arbitrary amplitude scales. We use a new database consisting of two speaking tasks (sentence and paragraph). Support vector machine is used as a classifier. Our results show that the pre-trained model embeddings outperformed three baseline features, providing improvements of up to 7%(absolute) in accuracy. Manila Kodali, Sudarsana Reddy Kadiri, Paavo Alku |
INTERSPEECH | 3 |
| 2023 | Refining a deep learning-based formant tracker using linear prediction methodsabstractIn this study, formant tracking is investigated by refining the formants tracked by an existing data-driven tracker, DeepFormants, using the formants estimated in a model-driven manner by linear prediction (LP)-based methods. As LP-based formant estimation methods, conventional covariance analysis (LP-COV) and the recently proposed quasi-closed phase forward–backward (QCP-FB) analysis are used. In the proposed refinement approach, the contours of the three lowest formants are first predicted by the data-driven DeepFormants tracker, and the predicted formants are replaced frame-wise with local spectral peaks shown by the model-driven LP-based methods. The refinement procedure can be plugged into the DeepFormants tracker with no need for any new data learning. Two refined DeepFormants trackers were compared with the original DeepFormants and with five known traditional trackers using the popular vocal tract resonance (VTR) corpus. The results indicated that the data-driven DeepFormants trackers outperformed the conventional trackers and that the best performance was obtained by refining the formants predicted by DeepFormants using QCP-FB analysis. In addition, by tracking formants using VTR speech that was corrupted by additive noise, the study showed that the refined DeepFormants trackers were more resilient to noise than the reference trackers. In general, these results suggest that LP-based model-driven approaches, which have traditionally been used in formant estimation, can be combined with a modern data-driven tracker easily with no further training to improve the tracker’s performance. Paavo Alku, Sudarsana Reddy Kadiri, Dhananjaya Gowda |
Comput. Speech Lang. | 1 |
| 2023 | Analysis of Instantaneous Frequency Components of Speech Signals for Epoch ExtractionabstractThe major impulse-like excitation in the speech signal is due to abrupt closure of the vocal folds, which takes place at the glottal closure instant (GCI) or epoch in each cycle. GCIs are used in many areas of speech science and technology, such as in prosody modification, voice source analysis, formant extraction and speech synthesis. It is difficult to observe these discontinuities (corresponding to GCIs) in the speech signal because of the superimposed time-varying response of the vocal tract system. This paper examines the phase part of different frequency components of the speech signal to extract epochs. Three analysis methods to decompose the speech signal into different frequency components are considered. These methods are the short-time Fourier transform (STFT), narrow bandpass filtering (NBPF), and single frequency filtering (SFF). The locations of the discontinuities in the speech signal are obtained from the instantaneous frequency (IF) (i.e., the time derivative of the phase) of each of the frequency components. A method for automatic detection of epochs using the amplitude weighted IF is proposed. Performance of the proposed epoch detection method is compared with four state-of-the-art methods in clean and telephone quality speech. The performance of the proposed method is comparable with the performance of the existing epoch detection methods for clean speech but better for telephone quality speech. Sudarsana Reddy Kadiri, Paavo Alku, Bayya Yegnanarayana |
Comput. Speech Lang. | 2 |
| 2023 | Classification of functional dysphonia using the tunable Q wavelet transformabstractFunctional dysphonia (FD) refers to an abnormality in voice quality in the absence of an identifiable lesion. In this paper, we propose an approach based on the tunable Q wavelet transform (TQWT) to automatically classify two types of FD (hyperfunctional dysphonia and hypofunctional dysphonia) from a healthy voice using the acoustic voice signal. Using TQWT, voice signals were decomposed into sub-bands and the entropy values extracted from the sub-bands were utilized as features for the studied 3-class classification problem. In addition, the Mel-frequency cepstral coefficient (MFCC) and glottal features were extracted from the acoustic voice signal and the estimated glottal source signal, respectively. A convolutional neural network (CNN) classifier was trained separately for the TQWT, MFCC and glottal features. Experiments were conducted using voice signals of 57 healthy speakers and 113 FD patients (72 with hyperfunctional dysphonia and 41 with hypofunctional dysphonia) taken from the VOICED database. These experiments revealed that the TQWT features yielded an absolute improvement of 5.5% and 4.5% compared to the baseline MFCC features and glottal features, respectively. Furthermore, the highest classification accuracy (67.91%) was obtained using the combination of the TQWT and glottal features, which indicates the complementary nature of these features. Mittapalle Kiran Reddy, Yagnavajjula Madhu Keerthana, Paavo Alku |
Speech Commun. | 3 |
| 2023 | Automatic Assessment of Parkinson's Disease Using Speech Representations of Phonation and ArticulationabstractSpeech from people with Parkinson's disease (PD) are likely to be degraded on phonation, articulation, and prosody. Motivated to describe articulation deficits comprehensively, we investigated 1) the universal phonological features that model articulation manner and place, also known as speech attributes, and 2) glottal features capturing phonation characteristics. These were further supplemented by, and compared with, prosodic features using a popular compact feature set and standard MFCC. Temporal characteristics of these features were modeled by convolutional neural networks. Besides the features, we were also interested in the speech tasks for collecting data for automatic PD speech assessment, like sustained vowels, text reading, and spontaneous monologue. For this, we utilized a recently collected Finnish PD corpus (PDSTU) as well as a Spanish database (PC-GITA). The experiments were formulated as regression problems against expert ratings of PD-related symptoms, including ratings of speech intelligibility, voice impairment, overall severity of communication disorder on PDSTU, as well as on the Unified Parkinson's Disease Rating Scale (UPDRS) on PC-GITA. The experimental results show: 1) the speech attribute features can well indicate the severity of pathologies in parkinsonian speech; 2) combining phonation features with articulatory features improves the PD assessment performance, but requires high-quality recordings to be applicable; 3) read speech leads to more accurate automatic ratings than the use of sustained vowels, but not if the amount of speech is limited to correspond to the sustained vowels in duration; and 4) jointly using data from several speech tasks can further improve the automatic PD assessment performance. Yuanyuan Liu 0002, Mittapalle Kiran Reddy, Nelly Penttilä, Tiina Ihalainen, Paavo Alku, Okko Johannes Räsänen |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Exemplar-Based Sparse Representations for Detection of Parkinson's Disease From SpeechabstractParkinson's disease (PD) is a progressive neurological disorder which affects the motor system. The automatic detection of PD improves the diagnosis of the disease, and it can be done in a non-invasive manner from speech. In this paper, we investigate the use of an exemplar-based sparse representation (SR) classification approach for detecting PD from speech. Exemplars are speech feature vectors extracted from the training data. The idea is to formulate the detection task as a problem of finding sparse representations of test speech feature vectors with respect to training speech exemplars. The main advantage of using the SR approach instead of conventional machine learning (ML)-based approaches is that the training step–which is time-consuming and sometimes requires unorganized hyper-parameter tuning–is not needed. Furthermore, SRs are more robust to redundancy and noise in the data. In this work, we study SR classification approaches based on two sparse coding models, namely, l1-regularized least squares ($l_{1}$LS) and non-negative least squares (NNLS). We propose a strategy based on class-specific dictionaries for improving performance of the$l_{1}$LS- and NNLS-based SR classification. To investigate the detection performance, the$l_{1}$LS- and NNLS-based approaches are applied and compared with the traditional PD detection approach based on ML classification algorithms using the PC-GITA PD dataset and an openly available dataset consisting of mobile device voice recordings from healthy and PD patients. The results indicate that the proposed NNLS-based SR classification approach performs better than the traditional ML-based methods in discriminating PD patients from healthy subjects. Mittapalle Kiran Reddy, Paavo Alku |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Comparing 1-dimensional and 2-dimensional spectral feature representations in voice pathology detection using machine learning and deep learning classifiersabstractThis work was supported by the Academy of Finland (grant number 313390). The computational resources were provided by Aalto ScienceIT. Farhad Javanmardi, Sudarsana Reddy Kadiri, Manila Kodali, Paavo Alku |
INTERSPEECH | 4 |
| 2022 | Convolutional Neural Networks for Classification of Voice Qualities from Speech and Neck Surface Accelerometer SignalsabstractThis work was supported by the Academy of Finland (grant number 313390). The computational resources were provided by Aalto ScienceIT. Sudarsana Reddy Kadiri, Farhad Javanmardi, Paavo Alku |
INTERSPEECH | 3 |
| 2022 | A formant modification method for improved ASR of children's speechabstractDifferences in acoustic characteristics between children’s and adults’ speech degrade performance of automatic speech recognition systems when systems trained using adults’ speech are used to recognize children’s speech. This performance degradation is due to the acoustic mismatch between training and testing. One of the main sources of the acoustic mismatch is the difference in vocal tract resonances (formant frequencies) between adult and child speakers. The present study aims to reduce the mismatch in formant frequencies by modifying formants of children’s speech to better correspond to formants of adults’ speech. This is carried out by warping the linear prediction (LP) spectrum computed from children’s speech. The warped LP spectra computed in a frame-based manner from children’s speech are used with the corresponding LP residuals to synthesize speech whose formant structure is closer to that of adults’ speech. When used in testing of an ASR system trained using adults’ speech, the warping reduces the spectral mismatch in speech between training and testing and improves the system performance in recognition of children’s speech. Experiments were conducted using narrowband (8 kHz) and wideband (16 kHz) speech of adult and child speakers from the WSJCAM0 and PF_STAR databases, respectively, and by recognizing children’s speech using acoustic models trained with adults’ speech. The proposed method gave relative improvements of 24% and 11% for the DNN and TDNN acoustic models, respectively, for narrowband speech. For wideband speech, the technique gave relative improvements of 27% and 13% for the DNN and TDNN acoustic models, respectively. The performance of the proposed method was also compared to two speaker adaptation methods: vocal tract length normalization (VTLN) and speaking rate adaptation (SRA). This comparison showed the best recognition performance for the proposed method. We also combined the proposed method with VTLN and SRA, and found that the combined method gave a further reduction in WER. Moreover, our experiments carried out for noisy speech using various types of additive noise and signal-to-noise ratios showed that the proposed method performs well also for degraded speech. Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Paavo Alku, Mikko Kurimo |
Speech Commun. | 3 |
| 2022 | Glottal flow characteristics in vowels produced by speakers with heart failureabstractHeart failure (HF) is one of the most life-threatening diseases globally. HF is an under-diagnosed condition, and more screening tools are needed to detect it. A few recent studies have suggested that HF also affects the functioning of the speech production mechanism by causing generation of edema in the vocal folds and by impairing the lung function. It has not yet been studied whether these possible effects of HF on the speech production mechanism are large enough to cause acoustically measurable differences to distinguish speech produced in HF from that produced by healthy speakers. Therefore, the goal of the present study was to compare speech production between HF patients and healthy controls by focusing on the excitation signal generated at the level of the vocal folds, the glottal flow. The glottal flow was computed from speech using the quasi-closed phase glottal inverse filtering method and the estimated flow was parameterized with 12 glottal parameters. The sound pressure level (SPL) was measured from speech as an additional parameter. The statistical analyses conducted on the parameters indicated that most of the glottal parameters and SPL were significantly different between the HF patients and healthy controls. The results showed that the HF patients generally produced a more rounded glottal pulse and a lower SPL level compared to the healthy controls, indicating incomplete glottal closure and inappropriate leakage of air through the glottis. The results observed in this preliminary study indicate that glottal features are capable of distinguishing speakers with HF from healthy controls. Therefore, the study suggests that glottal features constitute a potential feature extraction approach which should be taken into account in future large-scale investigations in studying the automatic detection of HF from speech. Mittapalle Kiran Reddy, Hilla Pohjalainen, Pyry Helkkula, Kasimir Kaitue, Mikko Minkkinen, Heli Tolppanen, Tuomo Nieminen, Paavo Alku |
Speech Commun. | 8 |
| 2022 | End-to-End Pathological Speech Detection Using Wavelet Scattering NetworkabstractIn recent years, developing robust systems for automatic detection of pathological speech has attracted increasing interest among researchers and clinicians. This study proposes an end-to-end approach based on wavelet scattering network (WSN) for detection of pathological speech. In the proposed approach, the WSN (which involves no learning) extracts suitable information from the input raw speech signal and this information is then passed through a multi-layer perceptron (MLP) in order to classify the speech signal as either healthy or pathological. The results show that the proposed approach outperformed a convolutional neural network (CNN) based end-to-end system in distinguishing pathological speech from healthy speech. Furthermore, the proposed system achieved comparable performance with a state-of-the-art traditional system based on hand-crafted features for uncompressed speech, but gave better performance than the traditional system for compressed speech of low bit rates. Mittapalle Kiran Reddy, Yagnavajjula Madhu Keerthana, Paavo Alku |
IEEE Signal Process. Lett. | 3 |
| 2021 | Glottal features for classification of phonation type from speech and neck surface accelerometer signalsabstractGlottal source characteristics vary between phonation types due to the tension of laryngeal muscles with the respiratory effort. Previous studies in the classification of phonation type have mainly used speech signals recorded by microphone. Recently, two studies were published in the classification of phonation type using neck surface accelerometer (NSA) signals. However, there are no previous studies comparing the use of the acoustic speech signal vs. the NSA signal as input in classifying phonation type. Therefore, the current study investigates simultaneously recorded speech and NSA signals in the classification of three phonation types (breathy, modal, pressed). The general goal is to understand which of the two signals (speech vs. NSA) is more effective in the classification task. We hypothesize that by using the same feature set for both signals, classification accuracy is higher for the NSA signal, which is more closely related to the physical vibration of the vocal folds and less affected by the vocal tract compared to the acoustical speech signal. Glottal source waveforms were computed using two signal processing methods, quasi-closed phase (QCP) glottal inverse filtering and zero frequency filtering (ZFF), and a group of time-domain and frequency-domain scalar features were computed from the obtained waveforms. In addition, the study investigated the use of mel-frequency cepstral coefficients (MFCCs) derived from the glottal source waveforms computed by QCP and ZFF. Classification experiments with support vector machine classifiers revealed that the NSA signal showed better discrimination of the phonation types compared to the speech signal when the same feature set was used. Furthermore, it was observed that the glottal features showed complementary information with the conventional MFCC features resulting in the best classification accuracy both for the NSA signal (86.9%) and the speech signal (80.6%). Sudarsana Reddy Kadiri, Paavo Alku |
Comput. Speech Lang. | 2 |
| 2021 | Automatic assessment of intelligibility in speakers with dysarthria from coded telephone speech using glottal features
N. P. Narendra, Paavo Alku |
Comput. Speech Lang. | 2 |
| 2021 | The automatic detection of heart failure using speech signalsabstractHeart failure (HF) is a major global health concern and is increasing in prevalence. It affects the larynx and breathing – thereby the quality of speech. In this article, we propose an approach for the automatic detection of people with HF using the speech signal. The proposed method explores mel-frequency cepstral coefficient (MFCC) features, glottal features, and their combination to distinguish HF from healthy speech. The glottal features were extracted from the voice source signal estimated using glottal inverse filtering. Four machine learning algorithms , namely, support vector machine , Extra Tree , AdaBoost , and feed-forward neural network (FFNN), were trained separately for individual features and their combination. It was observed that the MFCC features yielded higher classification accuracies compared to glottal features. Furthermore, the complementary nature of glottal features was investigated by combining these features with the MFCC features. Our results show that the FFNN classifier trained using a reduced set of glottal + MFCC features achieved the best overall performance in both speaker-dependent and speaker-independent scenarios. Mittapalle Kiran Reddy, Pyry Helkkula, Yagnavajjula Madhu Keerthana, Kasimir Kaitue, Mikko Minkkinen, Heli Tolppanen, Tuomo Nieminen, Paavo Alku |
Comput. Speech Lang. | 8 |
| 2021 | Extraction and Utilization of Excitation Information of Speech: A ReviewabstractSpeech production can be regarded as a process where a time-varying vocal tract system (filter) is excited by a time-varying excitation. In addition to its linguistic message, the speech signal also carries information about, for example, the gender and age of the speaker. Moreover, the speech signal includes acoustical cues about several speaker traits, such as the emotional state and the state of health of the speaker. In order to understand the production of these acoustical cues by the human speech production mechanism and utilize this information in speech technology, it is necessary to extract features describing both the excitation and the filter of the human speech production mechanism. While the methods to estimate and parameterize the vocal tract system are well established, the excitation appears less studied. This article provides a review of signal processing approaches used for the extraction of excitation information from speech. This article highlights the importance of excitation information in the analysis and classification of phonation type and vocal emotions, in the analysis of nonverbal laughter sounds, and in studying pathological voices. Furthermore, recent developments of deep learning techniques in the context of extraction and utilization of the excitation information are discussed. Sudarsana Reddy Kadiri, Paavo Alku, Bayya Yegnanarayana |
Proc. IEEE | 2 |
| 2021 | The Detection of Parkinson's Disease From Speech Using Voice Source InformationabstractDeveloping automatic methods to detect Parkinson's disease (PD) from speech has attracted increasing interest as these techniques can potentially be used in telemonitoring health applications. This article studies the utilization of voice source information in the detection of PD using two classifier architectures: traditional pipeline approach and end-to-end approach. The former consists of feature extraction and classifier stages. In feature extraction, the baseline acoustic features-consisting of articulation, phonation, and prosody features-were computed and voice source information was extracted using glottal features that were estimated by iterative adaptive inverse filtering (IAIF) and quasi-closed phase (QCP) glottal inverse filtering methods. Support vector machine classifiers were developed utilizing the baseline and glottal features extracted from every speech utterance and the corresponding healthy/PD labels. The end-to-end approach uses deep learning models which were trained using both raw speech waveforms and raw voice source waveforms. In the latter, two glottal inverse filtering methods (IAIF and QCP) and zero frequency filtering method were utilized. The deep learning architecture consists of a combination of convolutional layers followed by a multilayer perceptron. Experiments were performed using PC-GITA speech database. From the traditional pipeline systems, the highest classification accuracy (67.93%) was given by combination of baseline and QCP-based glottal features. From the end-to-end-systems, the highest accuracy (68.56%) was given by the system trained using QCP-based glottal flow signals. Even though classification accuracies were modest for all systems, the study is encouraging as the extraction of voice source information was found to be most effective in both approaches. N. P. Narendra, Björn W. Schuller, Paavo Alku |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Comparison of Glottal Closure Instants Detection Algorithms for Emotional SpeechabstractIn production of voiced speech, epochs or glottal closure instants (GCIs) refer to the instants of significant excitation of the vocal tract. Extraction of GCIs is used as a pre-processing stage in many areas of speech technology, such as in prosody modification, speech synthesis and voice source analysis. In the past decades, several GCI detection algorithms have been developed and most of them provide excellent results for speech signals produced using modal (normal) type of phonation. There are, however, no studies comparing multiple state-of-the-art GCI detection methods in emotional speech. In this paper, we compare six GCI detection algorithms using emotional speech and known evaluation metrics. We use the Berlin EMO-DB acted emotional speech database which contains seven emotions and simultaneous electroglottography (EGG) recordings as ground truth. The results show that all six GCI detection algorithms give best performance in processing speech of neutral emotion and that the performance degrade particularly in emotions of high arousal (anger and joy). To improve the performance of GCI detection in emotional speech, the study underlines the importance of local average pitch period estimates. Sudarsana Reddy Kadiri, Paavo Alku, Bayya Yegnanarayana |
ICASSP | 2 |
| 2020 | Study of Formant Modification for Children ASRabstractThe performance of automatic speech recognition systems for children’s speech is known to suffer from the large variation and mismatch in the acoustic and linguistic attributes between children’s and adults’ speech. One of the various identified sources of mismatch is the difference in formant frequencies between adults and children. In this paper, we propose a formant modification method to mitigate differences between adults’ and children’s speech and to improve the performance of ASR for children. The explored technique gives a relative 27% improvement in system performance compared to a hybrid DNN-HMM baseline. We also compare the system performance with related speaker adaptation methods like vocal tract length normalization (VTLN) and speaking rate adaptation (SRA) and find that the proposed method gives improvements over them, as well. Combining the proposed method with VTLN and SRA results in a further reduction of WER. We also found that the proposed method performs well even for noisy speech. Hemant Kumar Kathania, Sudarsana Reddy Kadiri, Paavo Alku, Mikko Kurimo |
ICASSP | 3 |
| 2020 | Parkinson's Disease Detection from Speech Using Single Frequency Filtering Cepstral CoefficientsabstractParkinson's disease (PD) is a progressive deterioration of the human central nervous system. Detection of PD (discriminating patients with PD from healthy subjects) from speech is a useful approach due to its non-invasive nature. This study proposes to use novel cepstral coefficients derived from the single frequency filtering (SFF) method, called as single frequency filtering cepstral coefficients (SFFCCs) for the detection of PD. SFF has been shown to provide higher spectro-temporal resolution compared to the short-time Fourier transform. The current study uses the PC-GITA database, which consists of speech from speakers with PD and healthy controls (50 males, 50 females). Our proposed detection system is based on the i-vectors derived from SFFCCs using SVM as a classifier. In the detection of PD, better performance was achieved when the i-vectors were computed from the proposed SFFCCs compared to the popular conventional MFCCs. Furthermore, we investigated the effect of temporal variations by deriving the shifted delta cepstral (SDC) coefficients using SFFCCs. These experiments revealed that the i-vectors derived from the proposed SFFCCs+SDC features gave an absolute improvement of 9% compared to the i-vectors derived from the baseline MFCCs+SDC features, indicating the importance of temporal variations in the detection of PD. Sudarsana Reddy Kadiri, Rashmi Kethireddy, Paavo Alku |
INTERSPEECH | 3 |
| 2020 | ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
Xin Wang 0037, Junichi Yamagishi, Massimiliano Todisco, Héctor Delgado, Andreas Nautsch, Nicholas W. D. Evans, Md. Sahidullah, Ville Vestman, Tomi Kinnunen, Kong-Aik Lee, Lauri Juvela, Paavo Alku, Yu-Huai Peng, Hsin-Te Hwang, Yu Tsao 0001, Hsin-Min Wang, Sébastien Le Maguer, Zhen-Hua Ling |
Comput. Speech Lang. | 12 |
| 2020 | Duration of the rhotic approximant /ɹ/ in spastic dysarthria of different severity levels
Krishna Gurugubelli, Anil Kumar Vuppala, N. P. Narendra, Paavo Alku |
Speech Commun. | 4 |
| 2020 | Analysis and classification of phonation types in speech and singing voice
Sudarsana Reddy Kadiri, Paavo Alku, Bayya Yegnanarayana |
Speech Commun. | 2 |
| 2020 | Automatic intelligibility assessment of dysarthric speech using glottal parameters
N. P. Narendra, Paavo Alku |
Speech Commun. | 2 |
| 2020 | Time-Varying Quasi-Closed-Phase Analysis for Accurate Formant Tracking in Speech SignalsabstractIn this paper, we propose a new method for the accurate estimation and tracking of formants in speech signals using time-varying quasi-closed-phase (TVQCP) analysis. Conventional formant tracking methods typically adopt a two-stage estimateand-track strategy wherein an initial set of formant candidates are estimated using short-time analysis (e.g., 10-50 ms), followed by a tracking stage based on dynamic programming or a linear state-space model. One of the main disadvantages of these approaches is that the tracking stage, however good it may be, cannot improve upon the formant estimation accuracy of the first stage. The proposed TVQCP method provides a single-stage formant tracking that combines the estimation and tracking stages into one. TVQCP analysis combines three approaches to improve formant estimation and tracking: (1) it uses temporally weighted quasi-closed-phase analysis to derive closed-phase estimates of the vocal tract with reduced interference from the excitation source, (2) it increases the residual sparsity by using the L1 optimization and (3) it uses time-varying linear prediction analysis overlong time windows (e.g., 100-200 ms) to impose a continuity constraint on the vocal tract model and hence on the formant trajectories. Formant tracking experiments with a wide variety of synthetic and natural speech signals show that the proposed TVQCP method performs better than conventional and popular formant tracking tools, such as Wavesurfer and Praat (based on dynamic programming), the KARMA algorithm (based on Kalman filtering), and DeepFormants (based on deep neural networks trained in a supervised manner). Matlab scripts for the proposed method can be found at: Dhananjaya Gowda, Sudarsana Reddy Kadiri, Brad H. Story, Paavo Alku |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Data Augmentation Strategies for Neural Network F0 EstimationabstractThis study explores various speech data augmentation methods for the task of noise-robust fundamental frequency (F0) estimation with neural networks. The explored augmentation strategies are split into additive noise and channel-based augmentation and into vocoder-based augmentation methods. In vocoder-based augmentation, a glottal vocoder is used to enhance the accuracy of ground truth F0 used for training of the neural network, as well as to expand the training data diversity in terms of F0 patterns and vocal tract lengths of the talkers. Evaluations on the PTDB-TUG corpus indicate that noise and channel augmentation can be used to greatly increase the noise robustness of trained models, and that vocoder-based ground truth enhancement further increases model performance. For smaller datasets, vocoder-based diversity augmentation can also be used to increase performance. The best-performing proposed method greatly outperformed the compared F0 estimation methods in terms of noise robustness. Manu Airaksinen, Lauri Juvela, Paavo Alku, Okko Johannes Räsänen |
ICASSP | 3 |
| 2019 | Waveform Generation for Text-to-speech Synthesis Using Pitch-synchronous Multi-scale Generative Adversarial NetworksabstractThe state-of-the-art in text-to-speech (TTS) synthesis has recently improved considerably due to novel neural waveform generation methods, such as WaveNet. However, these methods suffer from their slow sequential inference process, while their parallel versions are difficult to train and even more computationally expensive. Meanwhile, generative adversarial networks (GANs) have achieved impressive results in image generation and are making their way into audio applications; parallel inference is among their lucrative properties. By adopting recent advances in GAN training techniques, this investigation studies waveform generation for TTS in two domains (speech signal and glottal excitation). Listening test results show that while direct waveform generation with GAN is still far behind WaveNet, a GAN-based glottal excitation model can achieve quality and voice similarity on par with a WaveNet vocoder. Lauri Juvela, Bajibabu Bollepalli, Junichi Yamagishi, Paavo Alku |
ICASSP | 4 |
| 2019 | Cycle-consistent Adversarial Networks for Non-parallel Vocal Effort Based Speaking Style ConversionabstractSpeaking style conversion (SSC) is the technology of converting natural speech signals from one style to another. In this study, we propose the use of cycle-consistent adversarial networks (CycleGANs) for converting styles with varying vocal effort, and focus on conversion between normal and Lombard styles as a case study of this problem. We propose a parametric approach that uses the Pulse Model in Log domain (PML) vocoder to extract speech features. These features are mapped using the CycleGAN from utterances in the source style to the corresponding features of target speech. Finally, the mapped features are converted to a Lombard speech waveform with the PML. The CycleGAN was compared in subjective listening tests with 2 other standard mapping methods used in conversion, and the CycleGAN was found to have the best performance in terms of speech quality and in terms of the magnitude of the perceptual change between the two styles. Shreyas Seshadri, Lauri Juvela, Junichi Yamagishi, Okko Johannes Räsänen, Paavo Alku |
ICASSP | 5 |
| 2019 | Lombard Speech Synthesis Using Transfer Learning in a Tacotron Text-to-Speech SystemabstractCurrently, there is increasing interest to use sequence-to-sequence models in text-to-speech (TTS) synthesis with attention like that in Tacotron models. These models are end-to-end, meaning that they learn both co-articulation and duration properties directly from text and speech. Since these models are entirely data-driven, they need large amounts of data to generate synthetic speech of good quality. However, in challenging speaking styles, such as Lombard speech, it is difficult to record sufficiently large speech corpora. Therefore, we propose a transfer learning method to adapt a TTS system of normal speaking style to Lombard style. We also experiment with a WaveNet vocoder along with a traditional vocoder (WORLD) in the synthesis of Lombard speech. The subjective and objective evaluation results indicated that the proposed adaptation system coupled with the WaveNet vocoder clearly outperformed the conventional deep neural network based TTS system in the synthesis of Lombard speech Bajibabu Bollepalli, Lauri Juvela, Paavo Alku |
INTERSPEECH | 3 |
| 2019 | GELP: GAN-Excited Linear Prediction for Speech Synthesis from Mel-SpectrogramabstractRecent advances in neural network -based text-to-speech have reached human level naturalness in synthetic speech. The present sequence-to-sequence models can directly map text to mel-spectrogram acoustic features, which are convenient for modeling, but present additional challenges for vocoding (i.e., waveform generation from the acoustic features). Highquality synthesis can be achieved with neural vocoders, such as WaveNet, but such autoregressive models suffer from slow sequential inference. Meanwhile, their existing parallel inference counterparts are difficult to train and require increasingly large model sizes. In this paper, we propose an alternative training strategy for a parallel neural vocoder utilizing generative adversarial networks, and integrate a linear predictive synthesis filter into the model. Results show that the proposed model achieves significant improvement in inference speed, while outperforming a WaveNet in copy-synthesis quality. Lauri Juvela, Bajibabu Bollepalli, Junichi Yamagishi, Paavo Alku |
INTERSPEECH | 4 |
| 2019 | Mel-Frequency Cepstral Coefficients of Voice Source Waveforms for Classification of Phonation Types in SpeechabstractVoice source characteristics in different phonation types vary due to the tension of laryngeal muscles along with the respiratory effort. This study investigates the use of mel-frequency cepstral coefficients (MFCCs) derived from voice source waveforms for classification of phonation types in speech. The cepstral coefficients are computed using two source waveforms: (1) glottal flow waveforms estimated by the quasi-closed phase (QCP) glottal inverse filtering method and (2) approximate voice source waveforms obtained using the zero frequency filtering (ZFF) method. QCP estimates voice source waveforms based on the source-filter decomposition while ZFF yields source waveforms without explicitly computing the source-filter decomposition. Experiments using MFCCs computed from the two source waveforms show improved accuracy in classification of phonation types compared to the existing voice source features and conventional MFCC features. Further, it is observed that the proposed features have complimentary information to the existing features. Sudarsana Reddy Kadiri, Paavo Alku |
INTERSPEECH | 2 |
| 2019 | Augmented CycleGANs for Continuous Scale Normal-to-Lombard Speaking Style ConversionabstractLombard speech is a speaking style associated with increased vocal effort that is naturally used by humans to improve intelligibility in the presence of noise. It is hence desirable to have a system capable of converting speech from normal to Lombard style. Moreover, it would be useful if one could adjust the degree of Lombardness in the converted speech so that the system is more adaptable to different noise environments. In this study, we propose the use of recently developed Augmented cycle-consistent adversarial networks (Augmented CycleGANs) for conversion between normal and Lombard speaking styles. The proposed system gives a smooth control on the degree of Lombardness of the mapped utterances by traversing through different points in the latent space of the trained model. We utilize a parametric approach that uses the Pulse Model in Log domain (PML) vocoder to extract features from normal speech that are then mapped to Lombard-style features using the Augmented CycleGAN. Finally, the mapped features are converted to Lombard speech with PML. The model is trained on multi-language data recorded in different noise conditions, and we compare its effectiveness to a previously proposed CycleGAN system in experiments for intelligibility and quality of mapped speech. Shreyas Seshadri, Lauri Juvela, Paavo Alku, Okko Johannes Räsänen |
INTERSPEECH | 3 |
| 2019 | Vocal effort compensation for MFCC feature extraction in a shouted versus normal speaker recognition task
Emma Jokinen, Rahim Saeidi, Tomi Kinnunen, Paavo Alku |
Comput. Speech Lang. | 4 |
| 2019 | OPENGLOT - An open environment for the evaluation of glottal inverse filtering
Paavo Alku, Tiina Murtola, Jarmo Malinen, Juha Kuortti, Brad H. Story, Manu Airaksinen, Mika Salmi, Erkki Vilkman, Ahmed Geneid |
Speech Commun. | 1 |
| 2019 | Normal-to-Lombard adaptation of speech synthesis using long short-term memory recurrent neural networks
Bajibabu Bollepalli, Lauri Juvela, Manu Airaksinen, Cassia Valentini-Botinhao, Paavo Alku |
Speech Commun. | 5 |
| 2019 | Analysis of phonation onsets in vowel production, using information from glottal area and flow estimate
Tiina Murtola, Jarmo Malinen, Ahmed Geneid, Paavo Alku |
Speech Commun. | 4 |
| 2019 | Dysarthric speech classification from coded telephone speech using glottal features
N. P. Narendra, Paavo Alku |
Speech Commun. | 2 |
| 2019 | Estimation of the glottal source from coded telephone speech using deep neural networks
N. P. Narendra, Manu Airaksinen, Brad H. Story, Paavo Alku |
Speech Commun. | 4 |
| 2019 | GlotNet - A Raw Waveform Model for the Glottal Excitation in Statistical Parametric Speech SynthesisabstractRecently, generative neural network models which operate directly on raw audio, such as WaveNet, have improved the state of the art in text-to-speech synthesis (TTS). Moreover, there is increasing interest in using these models as statistical vocoders for generating speech waveforms from various acoustic features. However, there is also a need to reduce the model complexity, without compromising the synthesis quality. Previously, glottal pulseforms (i.e., time-domain waveforms corresponding to the source of human voice production mechanism) have been successfully synthesized in TTS by glottal vocoders using straightforward deep feedforward neural networks. Therefore, it is natural to extend the glottal waveform modeling domain to use the more powerful WaveNet-like architecture. Furthermore, due to their inherent simplicity, glottal excitation waveforms permit scaling down the waveform generator architecture. In this study, we present a raw waveform glottal excitation model, called GlotNet, and compare its performance with the corresponding direct speech waveform model, WaveNet, using equivalent architectures. The models are evaluated as part of a statistical parametric TTS system. Listening test results show that both approaches are rated highly in voice similarity to the target speaker, and obtain similar quality ratings with large models. Furthermore, when the model size is reduced, the quality degradation is less severe for GlotNet. Lauri Juvela, Bajibabu Bollepalli, Vassilis Tsiaras, Paavo Alku |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Speech Waveform Synthesis from MFCC Sequences with Generative Adversarial NetworksabstractThis paper proposes a method for generating speech from filterbank mel frequency cepstral coefficients (MFCC), which are widely used in speech applications, such as ASR, but are generally considered unusable for speech synthesis. First, we predict fundamental frequency and voicing information from MFCCs with an autoregressive recurrent neural net. Second, the spectral envelope information contained in MFCCs is converted to all-pole filters, and a pitch-synchronous excitation model matched to these filters is trained. Finally, we introduce a generative adversarial network -based noise model to add a realistic high-frequency stochastic component to the modeled excitation signal. The results show that high quality speech reconstruction can be obtained, given only MFCC information at test time. Lauri Juvela, Bajibabu Bollepalli, Xin Wang 0037, Hirokazu Kameoka, Manu Airaksinen, Junichi Yamagishi, Paavo Alku |
ICASSP | 7 |
| 2018 | Time-regularized Linear Prediction for Noise-robust Extraction of the Spectral Envelope of SpeechabstractFeature extraction of speech signals is typically performed in short-time frames by assuming that the signal is stationary within each frame. For the extraction of the spectral envelope of speech, which conveys the formant frequencies produced by the resonances of the slowly varying vocal tract, an often used frame length is within 20-30 ms. However, this kind of conventional frame-based spectral analysis is oblivious of the broader temporal context of the signal and is prone to degradation by, for example, environmental noise. In this paper, we propose a new frame-based linear prediction (LP) analysis method that includes a regularization term that penalizes energy differences in consecutive frames of an all-pole spectral envelope model. This integrates the slowly varying nature of the vocal tract as a part of the analysis. Objective evaluations related to feature distortion and phonetic representational capability were performed by studying the properties of the mel-frequency cepstral coefficient (MFCC) representations computed from different spectral estimation methods under noisy conditions using the TIMIT database. The results show that the proposed time-regularized LP approach exhibits superior MFCC distortion behavior while simultaneously having the greatest average separability of different phoneme categories in comparison to the other methods. Manu Airaksinen, Lauri Juvela, Okko Johannes Räsänen, Paavo Alku |
INTERSPEECH | 4 |
| 2018 | Speaker-independent Raw Waveform Model for Glottal ExcitationabstractRecent speech technology research has seen a growing interest in using WaveNets as statistical vocoders, i.e., generating speech waveforms from acoustic features. These models have been shown to improve the generated speech quality over classical vocoders in many tasks, such as text-to-speech synthesis and voice conversion. Furthermore, conditioning WaveNets with acoustic features allows sharing the waveform generator model across multiple speakers without additional speaker codes. However, multi-speaker WaveNet models require large amounts of training data and computation to cover the entire acoustic space. This paper proposes leveraging the source-filter model of speech production to more effectively train a speaker-independent waveform generator with limited resources. We present a multi-speaker 'GlotNet' vocoder, which utilizes a WaveNet to generate glottal excitation waveforms, which are then used to excite the corresponding vocal tract filter to produce speech. Listening tests show that the proposed model performs favourably to a direct WaveNet vocoder trained with the same model architecture and data. Lauri Juvela, Vassilis Tsiaras, Bajibabu Bollepalli, Manu Airaksinen, Junichi Yamagishi, Paavo Alku |
INTERSPEECH | 6 |
| 2018 | Dysarthric Speech Classification Using Glottal Features Computed from Non-words, Words and SentencesabstractDysarthria is a neuro-motor disorder resulting from the disruption of normal activity in speech production leading to slow, slurred and imprecise (low intelligible) speech. Automatic classification of dysarthria from speech can be used as a potential clinical tool in medical treatment. This paper examines the effectiveness of glottal source parameters in dysarthric speech classification from three categories of speech signals, namely non-words, words and sentences. In addition to the glottal parameters, two sets of acoustic parameters extracted by the openSMILE toolkit are used as baseline features. A dysarthric speech classification system is proposed by training support vector machines (SVMs) using features extracted from speech utterances and their labels indicating dysarthria/healthy. Classification accuracy results indicate that the glottal parameters contain discriminating information required for the identification of dysarthria. Additionally, the complementary nature of the glottal parameters is demonstrated when these parameters, in combination with the openSMILE-based acoustic features, result in improved classification accuracy. Analysis of classification accuracies of the glottal and openSMILE features for non-words, words and sentences is carried out. Results indicate that in terms of classification accuracy the word level is best suited in identifying the presence of dysarthria. N. P. Narendra, Paavo Alku |
INTERSPEECH | 2 |
| 2018 | Comparison of spectral tilt measures for sentence prominence in speech - Effects of dimensionality and adverse noise conditions
Sofoklis Kakouros, Okko Johannes Räsänen, Paavo Alku |
Speech Commun. | 3 |
| 2018 | Estimation of the glottal flow from speech pressure signals: Evaluation of three variants of iterative adaptive inverse filtering using computational physical modelling of voice productionabstractThe aim of this study is to comparatively review and evaluate three variants of the glottal inverse filtering algorithm based on iterative adaptive inverse filtering (IAIF): the Standard algorithm, and two recently proposed variants that use iterative optimal preemphasis (IOP) and a glottal flow model (GFM), respectively. To enable an objective evaluation, a computational physical model of voice production is used to generate time-domain signals pertaining to both the input glottal flow and the output speech pressure, for a wide range of vowels, fundamental frequencies, and voice qualities (involving co-variation of phonation type and loudness). Furthermore, for a fair comparison, the three key parameters of IAIF are selected by an exhaustive search to minimize the root-mean-square error between the estimated and reference glottal flow derivative in each analyzed frame and performance is assessed with two time-domain and two frequency-domain error measures. A conventional evaluation is also carried out with fixed parameter values determined by cross-validation. Results indicate that IOP tends to yield the lowest errors for nonback vowels (reducing errors by 31% on average compared with Standard), especially for not too high fundamental frequencies and not too pressed voice qualities; GFM becomes competitive for normal phonations when fixed parameter values are used; and in other cases, Standard IAIF is still recommended. In addition, the results suggest that not only the overall spectral tilt (as controlled by IOP and GFM) but also the balance between the levels of different spectral regions, can be important for accurate estimation of the glottal flow. Parham Mokhtari, Brad H. Story, Paavo Alku, Hiroshi Ando |
Speech Commun. | 3 |
| 2018 | Parameterization of a computational physical model for glottal flow using inverse filtering and high-speed videoendoscopy
Tiina Murtola, Paavo Alku, Jarmo Malinen, Ahmed Geneid |
Speech Commun. | 2 |
| 2018 | Speaker recognition from whispered speech: A tutorial survey and an application of time-varying linear prediction
Ville Vestman, Dhananjaya Gowda, Md. Sahidullah, Paavo Alku, Tomi Kinnunen |
Speech Commun. | 4 |
| 2018 | A Comparison Between STRAIGHT, Glottal, and Sinusoidal Vocoding in Statistical Parametric Speech SynthesisabstractA vocoder is used to express a speech waveform with a controllable parametric representation that can be converted back into a speech waveform. Vocoders representing their main categories (mixed excitation, glottal, and sinusoidal vocoders) were compared in this study with formal and crowd-sourced listening tests. The vocoder quality was measured within the context of analysis-synthesis as well as text-to-speech (TTS) synthesis in a modern statistical parametric speech synthesis framework. Furthermore, the TTS experiments were divided into synthesis with vocoder-specific features and synthesis with a shared envelope model, where the waveform generation method of the vocoders is mainly responsible for the quality differences. Finally, all of the tests included four distinct voices as a way to investigate the effect of different speakers on the synthesized speech quality. The obtained results suggest that the choice of the voice has a profound impact on the overall quality of the vocoder-generated speech, and the best vocoder for each voice can vary case by case. The single best-rated TTS system was obtained with the glottal vocoder GlottDNN using a male voice with low expressiveness. However, the results indicate that the sinusoidal vocoder PML (pulse model in log-domain) has the best overall performance across the performed tests. Finally, when controlling for the spectral models of the vocoders, the observed differences are similar to the baseline results. This indicates that the waveform generation method of a vocoder is essential for quality improvements. Manu Airaksinen, Lauri Juvela, Bajibabu Bollepalli, Junichi Yamagishi, Paavo Alku |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | Frequency-warped time-weighted linear prediction for glottal vocodingabstractAuto-regressive modeling is a prevalent source-filter separation method of speech. Conventional linear prediction (LP) and its derivatives such as weighted linear prediction (WeLP) produce parametric spectral models within a linear frequency scale, whereas frequency-warped linear prediction (WaLP) can be used to take into account the frequency sensitivity of the human auditory system. From the perspective of glottal vocoding, the principles behind WeLP have been found to be beneficial for an accurate separation of the glottal source signal and the vocal tract transfer function, but this approach can not utilize the auditory benefits of frequency warping. On the other hand, the WaLP approach suffers from less accurate source-filter separation properties. In this study, a generalized frequency-warped time-weighted linear prediction (WWLP) analysis is proposed. Experiments with WWLP are performed within the context of glottal vocoding. The subjective listening test results show that WWLP-based spectral envelope modeling is able to increase quality over previously developed methods in some of the test cases. Manu Airaksinen, Bajibabu Bollepalli, Jouni Pohjalainen, Paavo Alku |
ICASSP | 4 |
| 2017 | Lombard speech synthesis using long short-term memory recurrent neural networksabstractIn statistical parametric speech synthesis (SPSS), a few studies have investigated the Lombard effect, specifically by using hidden Markov model (HMM)-based systems. Recently, artificial neural networks have demonstrated promising results in SPSS, specifically by using long short-term memory recurrent neural networks (LSTMs). The Lombard effect, however, has not been studied in the LSTM-based speech synthesis systems. In this study, we propose three methods for Lombard speech adaptation in LSTM-based speech synthesis. In particular, (1) we augment Lombard specific information with the linguistic features as input, (2) scale the hidden activations using the learning hidden unit contributions (LHUC) method, and (3) fine-tune the LSTMs trained on normal speech with a small Lombard speech data. To investigate the effectiveness of the proposed methods, we carry out experiments using small (10 utterances) and large (500 utterances) Lombard speech data. Experimental results confirm the adaptability of the LSTMs, and similarity tests show that the LSTMs can achieve significantly better adaptation performance than the HMMs in both small and large data conditions. Bajibabu Bollepalli, Manu Airaksinen, Paavo Alku |
ICASSP | 3 |
| 2017 | Non-parallel voice conversion using i-vector PLDA: towards unifying speaker verification and transformationabstractText-independent speaker verification (recognizing speakers regardless of content) and non-parallel voice conversion (transforming voice identities without requiring content-matched training utterances) are related problems. We adopt i-vector method to voice conversion. An i-vector is a fixed-dimensional representation of a speech utterance that enables treating voice conversion in utterance domain, as opposed to frame domain. The high dimensionality (800) and small number of training utterances (24) necessitates using prior information of speakers. We adopt probabilistic linear discriminant analysis (PLDA) for voice conversion. The proposed approach requires neither parallel utterances, transcriptions nor time alignment procedures at any stage. Tomi Kinnunen, Lauri Juvela, Paavo Alku, Junichi Yamagishi |
ICASSP | 3 |
| 2017 | Normal-to-shouted speech spectral mapping for speaker recognition under vocal effort mismatchabstractSpeaker recognition performance degrades substantially in case of vocal effort mismatch (e.g. shouted vs. normal speech) between test and enrollment utterances. Such a mismatch is often encountered, for example, in forensic speaker recognition. This paper introduces a novel spectral mapping method which, when employed jointly with a statistical mapping technique, converts the Mel-frequency band energies of normal speech towards their counterparts in shouted speech. The aim is to obtain more robust performance in speaker recognition by tackling vocal effort mismatch between enrollment and test utterances. The processing is performed on the speech signal before feature extraction. The proposed approach was evaluated by testing the performance of a state-of-the-art i-vector-based speaker recognition system with and without applying the spectral mapping processing to the enrollment data. The results show that pre-processing with the proposed approach results in considerable improvement in correct identification rates. Ana Ramírez López, Rahim Saeidi, Lauri Juvela, Paavo Alku |
ICASSP | 4 |
| 2017 | Effects of Training Data Variety in Generating Glottal Pulses from Acoustic Features with DNNsabstractGlottal volume velocity waveform, the acoustical excitation of voiced speech, cannot be acquired through direct measurements in normal production of continuous speech. Glottal inverse filtering (GIF), however, can be used to estimate the glottal flow from recorded speech signals. Unfortunately, the usefulness of GIF algorithms is limited since they are sensitive to noise and call for high-quality recordings. Recently, efforts have been taken to expand the use of GIF by training deep neural networks (DNNs) to learn a statistical mapping between frame-level acoustic features and glottal pulses estimated by GIF. This framework has been successfully utilized in statistical speech synthesis in the form of the GlottDNN vocoder which uses a DNN to generate glottal pulses to be used as the synthesizer’s excitation waveform. In this study, we investigate how the DNN-based generation of glottal pulses is affected by training data variety. The evaluation is done using both objective measures as well as subjective listening tests of synthetic speech. The results suggest that the performance of the glottal pulse generation with DNNs is affected particularly by how well the training corpus suits GIF: processing low-pitched male speech and sustained phonations shows better performance than processing high-pitched female voices or continuous speech. Manu Airaksinen, Paavo Alku |
INTERSPEECH | 2 |
| 2017 | Generative Adversarial Network-Based Glottal Waveform Model for Statistical Parametric Speech SynthesisabstractRecent studies have shown that text-to-speech synthesis quality can be improved by using glottal vocoding. This refers to vocoders that parameterize speech into two parts, the glottal excitation and vocal tract, that occur in the human speech production apparatus. Current glottal vocoders generate the glottal excitation waveform by using deep neural networks (DNNs). However, the squared error-based training of the present glottal excitation models is limited to generating conditional average waveforms, which fails to capture the stochastic variation of the waveforms. As a result, shaped noise is added as post-processing. In this study, we propose a new method for predicting glottal waveforms by generative adversarial networks (GANs). GANs are generative models that aim to embed the data distribution in a latent space, enabling generation of new instances very similar to the original by randomly sampling the latent distribution. The glottal pulses generated by GANs show a stochastic component similar to natural glottal pulses. In our experiments, we compare synthetic speech generated using glottal waveforms produced by both DNNs and GANs. The results show that the newly proposed GANs achieve synthesis quality comparable to that of widely-used DNNs, without using an additive noise component. Bajibabu Bollepalli, Lauri Juvela, Paavo Alku |
INTERSPEECH | 3 |
| 2017 | Reducing Mismatch in Training of DNN-Based Glottal Excitation Models in a Statistical Parametric Text-to-Speech SystemabstractNeural network-based models that generate glottal excitation waveforms from acoustic features have been found to give improved quality in statistical parametric speech synthesis. Until now, however, these models have been trained separately from the acoustic model. This creates mismatch between training and synthesis, as the synthesized acoustic features used for the excitation model input differ from the original inputs, with which the model was trained on. Furthermore, due to the errors in predicting the vocal tract filter, the original excitation waveforms do not provide perfect reconstruction of the speech waveform even if predicted without error. To address these issues and to make the excitation model more robust against errors in acoustic modeling, this paper proposes two modifications to the excitation model training scheme. First, the excitation model is trained in a connected manner, with inputs generated by the acoustic model. Second, the target glottal waveforms are re-estimated by performing glottal inverse filtering with the predicted vocal tract filters. The results show that both of these modifications improve performance measured in MSE and MFCC distortion, and slightly improve the subjective quality of the synthetic speech. Lauri Juvela, Bajibabu Bollepalli, Junichi Yamagishi, Paavo Alku |
INTERSPEECH | 4 |
| 2017 | Evaluation of Spectral Tilt Measures for Sentence Prominence Under Different Noise ConditionsabstractSpectral tilt has been suggested to be a correlate of prominence in speech, although several studies have not replicated this empirically. This may be partially due to the lack of a standard method for tilt estimation from speech, rendering interpretations and comparisons between studies difficult. In addition, little is known about the performance of tilt estimators for prominence detection in the presence of noise. In this work, we investigate and compare several standard tilt measures on quantifying prominence in spoken Dutch and under different levels of additive noise. We also compare these measures with other acoustic correlates of prominence, namely, energy, F0, and duration. Our results provide further empirical support for the finding that tilt is a systematic correlate of prominence, at least in Dutch, even though energy, F0, and duration appear still to be more robust features for the task. In addition, our results show that there are notable differences between different tilt estimators in their ability to discriminate prominent words from non-prominent ones in different levels of noise. Sofoklis Kakouros, Okko Johannes Räsänen, Paavo Alku |
INTERSPEECH | 3 |
| 2017 | Speaking Style Conversion from Normal to Lombard Speech Using a Glottal Vocoder and Bayesian GMMsabstractSpeaking style conversion is the technology of converting natural speech signals from one style to another. In this study, we focus on normal-to-Lombard conversion. This can be used, for example, to enhance the intelligibility of speech in noisy environments. We propose a parametric approach that uses a vocoder to extract speech features. These features are mapped using Bayesian GMMs from utterances spoken in normal style to the corresponding features of Lombard speech. Finally, the mapped features are converted to a Lombard speech waveform with the vocoder. Two vocoders were compared in the proposed normal-to-Lombard conversion: a recently developed glottal vocoder that decomposes speech into glottal flow excitation and vocal tract, and the widely used STRAIGHT vocoder. The conversion quality was evaluated in two subjective listening tests measuring subjective similarity and naturalness. The similarity test results show that the system is able to convert normal speech into Lombard speech for the two vocoders. However, the subjective naturalness of the converted Lombard speech was clearly better using the glottal vocoder in comparison to STRAIGHT. Ana Ramírez López, Shreyas Seshadri, Lauri Juvela, Okko Johannes Räsänen, Paavo Alku |
INTERSPEECH | 5 |
| 2017 | Glottal Source Estimation from Coded Telephone Speech Using a Deep Neural NetworkabstractIn speech analysis, the information about the glottal source is obtained from speech by using glottal inverse filtering (GIF). The accuracy of state-of-the-art GIF methods is sufficiently high when the input speech signal is of high-quality (i.e., with little noise or reverberation). However, in realistic conditions, particularly when GIF is computed from coded telephone speech, the accuracy of GIF methods deteriorates severely. To robustly estimate the glottal source under coded condition, a deep neural network (DNN)-based method is proposed. The proposed method utilizes a DNN to map the speech features extracted from the coded speech to the glottal flow waveform estimated from the corresponding clean speech. To generate the coded telephone speech, adaptive multi-rate (AMR) codec is utilized which is a widely used speech compression method. The proposed glottal source estimation method is compared with two existing GIF methods, closed phase covariance analysis (CP) and iterative adaptive inverse filtering (IAIF). The results indicate that the proposed DNN-based method is capable of estimating glottal flow waveforms from coded telephone speech with a considerably better accuracy in comparison to CP and IAIF. N. P. Narendra, Manu Airaksinen, Paavo Alku |
INTERSPEECH | 3 |
| 2017 | Time-Varying Autoregressions for Speaker Verification in Reverberant ConditionsabstractAutomatic speaker verification (ASV) systems are vulnerable to spoofing attacks using speech generated by voice conversion and speech synthesis techniques. Commonly, a countermeasure (CM) system is integrated with an ASV system for improved protection against spoofing attacks. But integration of the two systems is challenging and often leads to increased false rejection rates. Furthermore, the performance of CM severely degrades if in-domain development data are unavailable. In this study, therefore, we propose a solution that uses two separate background models — one from human speech and another from spoofed data. During test, the ASV score for an input utterance is computed as the difference of the log-likelihood against the target model and the combination of the log-likelihoods against two background models. Evaluation experiments are conducted using the joint ASV and CM protocol of ASVspoof 2015 corpus consisting of text-independent ASV tasks with short utterances. Our proposed system reduces error rates in the presence of spoofing attacks by using out-of-domain spoofed data for system development, while maintaining the performance for zero-effort imposter attacks compared to the baseline system. Ville Vestman, Dhananjaya Gowda, Md. Sahidullah, Paavo Alku, Tomi Kinnunen |
INTERSPEECH | 4 |
| 2017 | Glottal Vocoding With Frequency-Warped Time-Weighted Linear PredictionabstractLinear prediction (LP) is a prevalent source-filter separation method of speech production. One of the drawbacks of conventional LP-based approaches is the biasing of estimated formants by harmonic peaks. Methods such as discrete all-pole modeling and weighted LP have been proposed to overcome this problem, but they all use a linear frequency scale. This study proposes a new LP technique, frequency-warped time-weighted linear prediction (WWLP), to provide spectral envelope estimates robust to harmonic peaks that work on a warped frequency scale that approximates the sensitivities of the human auditory system. Experiments are performed within the context of vocoding in statistical parametric speech synthesis. Subjective listening test results show that WWLP-based spectral envelope modeling increases quality over previously developed methods. Manu Airaksinen, Bajibabu Bollepalli, Jouni Pohjalainen, Paavo Alku |
IEEE Signal Process. Lett. | 4 |
| 2017 | Quadratic Programming Approach to Glottal Inverse Filtering by Joint Norm-1 and Norm-2 OptimizationabstractThis study proposes an approach for glottal inverse filtering of acoustic speech signals using quadratic programming (QPR). The method aims to jointly model the effect of vocal tract and lip radiation with a single filter whose coefficients are optimized using QPR. This optimization is based on the principles of closed phase analysis, where the contribution of the glottal source is attenuated in optimizing the inverse model of the vocal tract. By expressing the optimization problem in terms of the output of a filter, we can apply physically motivated optimization such as flatness of the closed phase. The proposed method was objectively evaluated using a synthetic Liljencrants-Fant model based test set of sustained vowels, as well as a real speech test set where the glottal flow estimates' closed phases were compared in terms of their flatness. The results based on synthetic speech indicate that the proposed method is robust to changes in f0, and state-of-the-art quality results were obtained for high-pitched voices, when f0is in the range 330-450 Hz. The results based on real speech indicate that the proposed method produces glottal flow estimates that have flatter closed phases with less formant ripple in comparison to estimates computed with known reference methods. Manu Airaksinen, Tom Bäckström, Paavo Alku |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | The Linear Predictive Modeling of Speech From Higher-Lag Autocorrelation Coefficients Applied to Noise-Robust Speaker RecognitionabstractA linear predictive spectral estimation method based on higher-lag autocorrelation coefficients is proposed for the noise-robust feature extraction from speech. The method, called higher-lag linear prediction, is derived from a signal prediction model that is optimized in the mean square sense using a cost function that has two prediction error terms, the first of which is similar to that of conventional linear prediction and the second of which is a delayed version introducing an integer delay of M samples. This basic form is developed further into the combined higher-lag linear prediction (CHLLP) model by simultaneously taking advantage of the zero-lag and higher-lag predictions. The CHLLP model was used in the computation of mel-frequency cepstral coefficients and compared with several reference feature extraction methods in speaker recognition. The experiments were conducted by using a modern i-vector-based system. Noise-corruption was done using both additive car, babble, and factory noise in different signal-to-noise ratio conditions as well as speech recordings from real noisy conditions. The results indicate that CHLLP outperformed the reference feature extraction methods in almost all the comparisons in the noise-corrupted conditions and the performance of CHLLP was only slightly inferior to the nonparametric FFT-based spectral modeling in the clean condition. Paavo Alku, Rahim Saeidi |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Intelligibility Enhancement of Telephone Speech Using Gaussian Process Regression for Normal-to-Lombard Spectral Tilt ConversionabstractNoise in the environment can decrease the quality and intelligibility of a telephone conversation. This study focuses on the intelligibility enhancement of narrowband telephone speech in a near-end noise scenario using a postprocessing method based on normal-to-Lombard spectral tilt conversion. The proposed technique uses nonparallel, conversational normal, and Lombard speech together with Gaussian process regression in order to mimic the flattening of the spectral tilt that occurs in the production of natural speech in noisy conditions. The performance of the proposed method was evaluated in comparison to two reference methods, a fixed high-pass filter, and a baseline spectral tilt conversion, as well as in comparison to unprocessed speech in terms of intelligibility and listening preference in noisy conditions and in terms of pressedness in silent conditions. The results indicate that while the proposed technique provides a similar benefit in terms of intelligibility as fixed high-pass filtering, it is also able to produce a notable increase in pressedness. This suggests that the developed processing of the spectral tilt can compete with fixed high-pass filtering in intelligibility enhancement, but it is also able to convert speech to become perceptually closer to natural Lombard speech. Emma Jokinen, Ulpu Remes, Paavo Alku |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | A subjective listening test of six different artificial bandwidth extension approaches in English, Chinese, German, and KoreanabstractIn studies on artificial bandwidth extension (ABE), there is a lack of international coordination in subjective tests between multiple methods and languages. Here we present the design of absolute category rating listening tests evaluating 12 ABE variants of six approaches in multiple languages, namely in American English, Chinese, German, and Korean. Since the number of ABE variants caused a higher-than-recommended length of the listening test, ABE variants were distributed into two separate listening tests per language. The paper focuses on the listening test design, which aimed at merging the subjective scores of both tests and thus allows for a joint analysis of all ABE variants under test at once. A language-dependent analysis, evaluating ABE variants in the context of the underlying coded narrowband speech condition showed statistical significant improvement in English, German, and Korean for some ABE solutions. Johannes Abel, Magdalena Kaniewska, Cyril Guillaume, Wouter Tirry, Hannu Pulakka, Ville Myllylä, Jari Sjoberg, Paavo Alku, Itai Katsir, David Malah, Israel Cohen, M. A. Tugtekin Turan, Engin Erzin, Thomas Schlien, Peter Vary, Amr H. Nour-Eldin, Peter Kabal, Tim Fingscheidt |
ICASSP | 8 |
| 2016 | Quasi closed phase analysis of speech signals using time varying weighted linear prediction for accurate formant trackingabstractRecent research on temporally weighted linear prediction shows that quasi closed phase (QCP) analysis of speech signals provides better modeling of the vocal tract and the glottal source. Quasi closed phase analysis gives more weightage on the closed phase of the glottal cycle, at the same time deemphasizing the region around the instant of significant excitation which is often poorly predicted. However, all the traditional analysis techniques including the QCP analysis is performed over short intervals of time. They do not impose any continuity constraints either on the vocal tract system or the glottal source. Such constraints are often imposed at a later stage to either smooth or track the estimated features over time. Time varying linear prediction (TVLP) provides a framework for modeling speech with a long-term continuity constraint imposed on the vocal tract shape. In this paper, we propose a new method for accurate modeling and tracking of the vocal tract resonances by integrating the advantages of a QCP analysis with that of TVLP. Formant tracking experiments show consistent improvement in performance over traditional LP or TVLP methods under a variety of conditions including different voice types and over a wide range of fundamental frequency. Dhananjaya Gowda, Manu Airaksinen, Paavo Alku |
ICASSP | 3 |
| 2016 | High-pitched excitation generation for glottal vocoding in statistical parametric speech synthesis using a deep neural networkabstractAchieving high quality and naturalness in statistical parametric synthesis of female voices remains to be difficult despite recent advances in the study area. Vocoding is one such key element in all statistical speech synthesizers that is known to affect the synthesis quality and naturalness. The present study focuses on a special type of vocoding, glottal vocoders, which aim to parameterize speech based on modelling the real excitation of (voiced) speech, the glottal flow. More specifically, we compare three different glottal vocoders by aiming at improved synthesis naturalness of female voices. Two of the vocoders are previously known, both utilizing an old glottal inverse filtering (GIF) method in estimating the glottal flow. The third on, denoted as Quasi Closed Phase - Deep Neural Net (QCP-DNN), takes advantage of a recently proposed new GIF method that shows improved accuracy in estimating the glottal flow from high-pitched speech. Subjective listening tests conducted on an US English female voice show that the proposed QCP-DNN method gives significant improvement in synthetic naturalness compared to the two previously developed glottal vocoders. Lauri Juvela, Bajibabu Bollepalli, Manu Airaksinen, Paavo Alku |
ICASSP | 4 |
| 2016 | GlottDNN - A Full-Band Glottal Vocoder for Statistical Parametric Speech SynthesisabstractGlottHMM is a previously developed vocoder that has been successfully used in HMM-based synthesis by parameterizing speech into two parts (glottal flow, vocal tract) according to the functioning of the real human voice production mechanism. In this study, a new glottal vocoding method, GlottDNN, is proposed. The GlottDNN vocoder is built on the principles of its predecessor, GlottHMM, but the new vocoder introduces three main improvements: GlottDNN (1) takes advantage of a new, more accurate glottal inverse filtering method, (2) uses a new method of deep neural network (DNN) -based glottal excitation generation, and (3) proposes a new approach of band-wise processing of full-band speech. The proposed GlottDNN vocoder was evaluated as part of a full-band state-of-the-art DNN-based text-to-speech (TTS) synthesis system, and compared against the release version of the original GlottHMM vocoder, and the well-known STRAIGHT vocoder. The results of the subjective listening test indicate that GlottDNN improves the TTS quality over the compared methods. Manu Airaksinen, Bajibabu Bollepalli, Lauri Juvela, Zhizheng Wu 0001, Simon King 0001, Paavo Alku |
INTERSPEECH | 6 |
| 2016 | Automatic Glottal Inverse Filtering with Non-Negative Matrix Factorization
Manu Airaksinen, Lauri Juvela, Tom Bäckström, Paavo Alku |
INTERSPEECH | 4 |
| 2016 | Time-Varying Quasi-Closed-Phase Weighted Linear Prediction Analysis of Speech for Accurate Formant Detection and Tracking
Dhananjaya Gowda, Paavo Alku |
INTERSPEECH | 2 |
| 2016 | Intelligibility Enhancement at the Receiving End of the Speech Transmission System - Effects of Far-End Noise Reduction
Emma Jokinen, Paavo Alku |
INTERSPEECH | 2 |
| 2016 | The Use of Read versus Conversational Lombard Speech in Spectral Tilt Modeling for Intelligibility Enhancement in Near-End Noise Conditions
Emma Jokinen, Ulpu Remes, Paavo Alku |
INTERSPEECH | 3 |
| 2016 | Majorisation-Minimisation Based Optimisation of the Composite Autoregressive System with Application to Glottal Inverse FilteringabstractThe composite autoregressive system can be used to estimate a speech source-filter decomposition in a rigorous manner, thus having potential use in glottal inverse filtering. By introducing a suitable prior, spectral tilt can be introduced into the source component estimation to better correspond to human voice production. However, the current expectation-maximisation based composite autoregressive model optimisation leaves room for improvement in terms of speed. Inspired by majorisation-minimisation techniques used for nonnegative matrix factorisation, this work derives new update rules for the model, resulting in faster convergence compared to the original approach. Additionally, we present a new glottal inverse filtering method based on the composite autoregressive system and compare it with inverse filtering methods currently used in glottal excitation modelling for parametric speech synthesis. These initial results show that the proposed method performs comparatively well, sometimes outperforming the reference methods. Lauri Juvela, Hirokazu Kameoka, Manu Airaksinen, Junichi Yamagishi, Paavo Alku |
INTERSPEECH | 5 |
| 2016 | Using Text and Acoustic Features in Predicting Glottal Excitation Waveforms for Parametric Speech Synthesis with Recurrent Neural NetworksabstractThis work studies the use of deep learning methods to directly model glottal excitation waveforms from context dependent text features in a text-to-speech synthesis system. Glottal vocoding is integrated into a deep neural network-based text-to-speech framework where text and acoustic features can be flexibly used as both network inputs or outputs. Long short-term memory recurrent neural networks are utilised in two stages: first, in mapping text features to acoustic features and second, in predicting glottal waveforms from the text and/or acoustic features. Results show that using the text features directly yields similar quality to the prediction of the excitation from acoustic features, both outperforming a baseline system based on using a fixed glottal pulse for excitation generation. Lauri Juvela, Xin Wang 0037, Shinji Takaki, Manu Airaksinen, Junichi Yamagishi, Paavo Alku |
INTERSPEECH | 6 |
| 2016 | Analysis of Face Mask Effect on Speaker Recognition
Rahim Saeidi, Ilkka Huhtakallio, Paavo Alku |
INTERSPEECH | 3 |
| 2016 | Comparing human and automatic speech recognition in a perceptual restoration experiment
Ulpu Remes, Ana Ramírez López, Lauri Juvela, Kalle J. Palomäki, Guy J. Brown, Paavo Alku, Mikko Kurimo |
Comput. Speech Lang. | 6 |
| 2016 | Phase modification for increasing the intelligibility of telephone speech in near-end noise conditions - evaluation of two methods
Emma Jokinen, Hannu Pulakka, Paavo Alku |
Speech Commun. | 3 |
| 2016 | Phase perception of the glottal excitation and its relevance in statistical parametric speech synthesis
Tuomo Raitio, Lauri Juvela, Antti Suni, Martti Vainio, Paavo Alku |
Speech Commun. | 5 |
| 2016 | Feature Extraction Using Power-Law Adjusted Linear Prediction With Application to Speaker Recognition Under Severe Vocal Effort MismatchabstractLinear prediction is one of the most established techniques in signal estimation, and it is widely utilized in speech signal processing. It has been long understood that the nerve firing rate of human auditory system can be approximated by power law non-linearity, and this has been the motivation behind using perceptual linear prediction in extracting acoustic features in a variety of speech processing applications. In this paper, we revisit the application of power law non-linearity in speech spectrum estimation by compressing/expanding power spectrum in autocorrelation-based linear prediction. The development of so-called LP- α is motivated by a desire to obtain spectral features that present less mismatch than conventionally used spectrum estimation methods when speech of normal loudness is compared to speech under vocal effort. The effectiveness of the proposed approach is demonstrated in a speaker recognition task conducted under severe vocal effort mismatch comparing shouted versus normal speech mode. Rahim Saeidi, Paavo Alku, Tom Bäckström |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Noise robust estimation of the voice source using a deep neural networkabstractIn the analysis of speech production, information about the voice source can be obtained non-invasively with glottal inverse filtering (GIF) methods. Current state-of-the-art GIF methods are capable of producing high-quality estimates in suitable conditions (e.g. low noise and reverberation), but their performance deteriorates in nonideal conditions because they require noise-sensitive parameter estimation. This study proposes a method for noise robust estimation of the voice source by creating a mapping using a deep neural network (DNN) between robust low-level speech features and the desired reference, a time-domain glottal flow computed by a GIF method. The method was evaluated with two GIF methods, of which one (quasi closed phase analysis, QCP) requires additional parameter estimation and the other (iterative adaptive inverse filtering, IAIF) does not. The results show that the proposed method outperforms the QCP method with SNRs less than 50-20 dB, but the simple IAIF method only with very low SNRs. Manu Airaksinen, Tuomo Raitio, Paavo Alku |
ICASSP | 3 |
| 2015 | Glottal inverse filtering based on quadratic programming
Manu Airaksinen, Tom Bäckström, Paavo Alku |
INTERSPEECH | 3 |
| 2015 | AM-FM based filter bank analysis for estimation of spectro-temporal envelopes and its application for speaker recognition in noisy reverberant environments
Dhananjaya Gowda, Rahim Saeidi, Paavo Alku |
INTERSPEECH | 3 |
| 2015 | Comparison of Gaussian process regression and Gaussian mixture models in spectral tilt modelling for intelligibility enhancement of telephone speech
Emma Jokinen, Ulpu Remes, Paavo Alku |
INTERSPEECH | 3 |
| 2015 | Speech quality evaluation of artificial bandwidth extension: comparing subjective judgments and instrumental predictions
Hannu Pulakka, Ville Myllylä, Anssi Rämö, Paavo Alku |
INTERSPEECH | 4 |
| 2015 | Phase perception of the glottal excitation of vocoded speech
Tuomo Raitio, Lauri Juvela, Antti Suni, Martti Vainio, Paavo Alku |
INTERSPEECH | 5 |
| 2015 | Accounting for uncertainty of i-vectors in speaker recognition using uncertainty propagation and modified imputationabstractOne of the biggest challenges in speaker recognition is incom-plete observations in test phase caused by availability of only short duration utterances. The problem with short utterances is that speaker recognition needs to be handled by having in-formation from only limited amount of acoustic classes. By considering limited observations from a test speaker, the re-sulting i-vector as a representative of short utterance will be uncertain; the shorter the duration, the higher the uncertainty. In recent studies, an uncertainty decoding technique has been employed in probabilistic linear discriminant analysis (PLDA) modeling in order to account for uncertain i-vectors. In this paper, we propose to extend uncertainty handling using simpli-fied PLDA scoring and modified imputation. We experiment with a state-of-the-art speaker recognition system focusing on uncertainty caused by controlled utterance duration. The uncer-tainties after i-vector extraction are being propagated through pre-processing steps and both uncertainty decoding and modi-fied imputation are considered. Our experimental results indi-cate improved equal error rate and detection cost attained by us-ing uncertainty-of-observation techniques in dealing with short duration utterances. Index Terms: speaker verification, duration, uncertainty de-coding, modified imputation Rahim Saeidi, Paavo Alku |
INTERSPEECH | 2 |
| 2015 | Speaker recognition for speech under face coverabstractSpeech under face cover constitute a case that is increasingly met by forensic speech experts. Wearing face cover mostly hap-pens when an individual strives to conceal his or her identity. Based on the material of face cover and the level of contact with speech production organs, speech production becomes affected by face mask and a part of speech energy gets absorbed in the mask. There has been little research on how speech acoustics is affected by different face masks and how face covers might affect performance of automatic speaker recognition systems. In the present paper, we have collected speech under face mask with the aim of studying the effects of wearing different masks on state-of-the-art text-independent automatic speaker recogni-tion system. The preliminary speaker recognition rates along with mask identification experiments are presented in this pa-per. Rahim Saeidi, Tuija Niemi, Hanna Karppelin, Jouni Pohjalainen, Tomi Kinnunen, Paavo Alku |
INTERSPEECH | 6 |
| 2014 | Comparison of post-processing methods for intelligibility enhancement of narrowband speech in a mobile phone frameworkabstractPost-processing methods can be used in mobile communications to improve the intelligibility of speech in adverse background noise conditions. This study addresses the improved intelligibility and the speech quality achieved with a well-known approach, dynamic range compression, by comparing it to two other real-time postprocessing methods based on energy reallocation. In addition, the effects of utilizing amplitude normalization instead of energy normalization on the performance of the post-processing methods are investigated. The evaluations were conducted using subjective tests in several background noise conditions. The results indicate that the two energy reallocating approaches outperform dynamic range compression both in intelligibility and quality and that amplitude normalization causes the performance of the tested post-processing methods to degrade in some conditions. Emma Jokinen, Marko Takanen, Paavo Alku |
ICASSP | 3 |
| 2014 | Multi-scale modulation filtering in automatic detection of emotions in telephone speechabstractThis study investigates emotion detection from noise-corrupted telephone speech. A generic modulation filtering approach for audio pattern recognition is proposed that utilizes inherent long-term properties of acoustic features in different classes. When applied to binary classification along the activation and valence dimensions, filtering the baseline short-time timbral features in both the training and detection phase leads to significant improvement especially in noise robustness. Automatic selection of training data based on the filter's prediction residual further improves the results. Jouni Pohjalainen, Paavo Alku |
ICASSP | 2 |
| 2014 | Gaussian mixture linear predictionabstractThis work introduces an approach to linear predictive signal analysis utilizing a Gaussian mixture autoregressive model. By initializing different autoregressive states of the model to approximately correspond to the target signal and the expected type of undesired signal components, such as background noise, the iterative parameter estimation converges towards a focused linear prediction model of the target signal. Differently initialized and trained variants of mixture linear prediction are evaluated using objective spectrum distortion measures as well as in feature extraction for speech detection in the presence of ambient noise. In these evaluations, the novel analysis methods perform better than the Fourier transform and conventional linear prediction. Jouni Pohjalainen, Paavo Alku |
ICASSP | 2 |
| 2014 | Parameterization of the glottal source with the phase plane plot
Manu Airaksinen, Paavo Alku |
INTERSPEECH | 2 |
| 2014 | Automatic estimation of the lip radiation effect in glottal inverse filtering
Manu Airaksinen, Tom Bäckström, Paavo Alku |
INTERSPEECH | 3 |
| 2014 | Spectral tilt modelling with GMMs for intelligibility enhancement of narrowband telephone speech
Emma Jokinen, Ulpu Remes, Marko Takanen, Kalle J. Palomäki, Mikko Kurimo, Paavo Alku |
INTERSPEECH | 6 |
| 2014 | Enhancement of speech intelligibility in near-end noise conditions with phase modification
Emma Jokinen, Marko Takanen, Hannu Pulakka, Paavo Alku |
INTERSPEECH | 4 |
| 2014 | Filtering and subspace selection for spectral features in detecting speech under physical stressabstractThis paper investigates approaches to modeling the time evolution of short-time spectral features in paralinguistic s peech type classification, where we focus on detection of speech in fluenced by physical exertion. The time series model consist s of autoregressive processes of multiple time scales and orders and is trained to describe the long-term dynamics of a given target speech class. The model is applied in two ways in improving long-term modeling in the detection task: 1) to perform predictive filtering of the features and 2) to automatically select instantaneous classification subspaces. The spectrum analysis me thod underlying the short-time features is also varied between t he standard discrete Fourier transform and a time-weighted linear predictive method which yields smooth all-pole spectrum envelope models. Configurations of the proposed methods are eval uated in the Physical Load task of the Interspeech 2014 Computational Paralinguistics Challenge and show improvement over the baseline timbral classifier and the challenge baseline. Also the interrelationships among the methods are discussed. Index Terms: computational paralinguistics, physical load, modulation filtering, spectrum analysis Jouni Pohjalainen, Paavo Alku |
INTERSPEECH | 2 |
| 2014 | Subjective voice quality evaluation of artificial bandwidth extension: comparing different audio bandwidths and speech codecs
Hannu Pulakka, Anssi Rämö, Ville Myllylä, Henri Toukomaa, Paavo Alku |
INTERSPEECH | 5 |
| 2014 | Deep neural network based trainable voice source model for synthesis of speech with varying vocal effortabstractThis paper studies a deep neural network (DNN) based voice source modelling method in the synthesis of speech with varying vocal effort. The new trainable voice source model learns a mapping between the acoustic features and the time-domain pitch-synchronous glottal flow waveform using a DNN. The voice source model is trained with various speech material from breathy, normal, and Lombard speech. In synthesis, a normal voice is first adapted to a desired style, and using the flexible DNN-based voice source model, a style-specific excitation waveform is automatically generated based on the adapted acoustic features. The proposed voice source model is compared to a robust and high-quality excitation modelling method based on manually selected mean glottal flow pulses for each vocal effort level and using a spectral matching filter to correctly match the voice source spectrum to a desired style. Subjective evaluations show that the proposed DNN-based method is rated comparable to the baseline method, but avoids the manual selection of the pulses and is computationally faster than a system using a spectral matching filter. Tuomo Raitio, Antti Suni, Lauri Juvela, Martti Vainio, Paavo Alku |
INTERSPEECH | 5 |
| 2014 | Automatic glottal inverse filtering with the Markov chain Monte Carlo method
Harri Auvinen, Tuomo Raitio, Manu Airaksinen, Samuli Siltanen, Brad H. Story, Paavo Alku |
Comput. Speech Lang. | 6 |
| 2014 | Glottal source processing: From analysis to applications
Thomas Drugman, Paavo Alku, Abeer Alwan, Bayya Yegnanarayana |
Comput. Speech Lang. | 2 |
| 2014 | An adaptive post-filtering method producing an artificial Lombard-like effect for intelligibility enhancement of narrowband telephone speech
Emma Jokinen, Marko Takanen, Martti Vainio, Paavo Alku |
Comput. Speech Lang. | 4 |
| 2014 | Synthesis and perception of breathy, normal, and Lombard speech in the presence of noise
Tuomo Raitio, Antti Suni, Martti Vainio, Paavo Alku |
Comput. Speech Lang. | 4 |
| 2014 | Mixture Linear Prediction in Speaker Verification Under Vocal Effort MismatchabstractThis paper describes an approach to robust signal analysis using iterative parameter re-estimation of a mixture autoregressive (AR) model. The model's focus can be adjusted by initialization of the target and non-target states. The variant examined in this study uses an i.i.d. mixture AR model and is designed to tackle the spectral biasing effect caused by the voice excitation in speech signals with variable fundamental frequency. In our speaker verification experiments, this method performed competitively against standard spectrum analysis techniques in non-mismatch conditions and showed significant improvements in vocal effort mismatch conditions. Jouni Pohjalainen, Cemal Hanilçi, Tomi Kinnunen, Paavo Alku |
IEEE Signal Process. Lett. | 4 |
| 2014 | Quasi Closed Phase Glottal Inverse Filtering Analysis With Weighted Linear PredictionabstractThis study presents a new glottal inverse filtering (GIF) technique based on closed phase analysis over multiple fundamental periods. The proposed quasi closed phase (QCP) analysis method utilizes weighted linear prediction (WLP) with a specific attenuated main excitation (AME) weight function that attenuates the contribution of the glottal source in the linear prediction model optimization. This enables the use of the autocorrelation criterion in linear prediction in contrast to the covariance criterion used in conventional closed phase analysis. The QCP method was compared to previously developed methods by using synthetic vowels produced with the conventional source-filter model as well as with a physical modeling approach. The obtained objective measures show that the QCP method improves the GIF performance in terms of errors in typical glottal source parametrizations for both low- and high-pitched vowels. Additionally, QCP was tested in a physiologically oriented vocoder, where the analysis/synthesis quality was evaluated with a subjective listening test indicating improved perceived quality for normal speaking style. Manu Airaksinen, Tuomo Raitio, Brad H. Story, Paavo Alku |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2013 | Speaker identification from shouted speech: Analysis and compensationabstractText-independent speaker identification is studied using neutral and shouted speech in Finnish to analyze the effect of vocal mode mismatch between training and test utterances. Standard mel-frequency cepstral coefficient (MFCC) features with Gaussian mixture model (GMM) recognizer are used for speaker identification. The results indicate that speaker identification accuracy reduces from perfect (100 %) to 8.71 % under vocal mode mismatch. Because of this dramatic degradation in recognition accuracy, we propose to use a joint density GMM mapping technique for compensating the MFCC features. This mapping is trained on a disjoint emotional speech corpus to create a completely speaker- and speech mode independent emotion-neutralizing mapping. As a result of the compensation, the 8.71 % identification accuracy increases to 32.00 % without degrading the non-mismatched train-test conditions much. Cemal Hanilçi, Tomi Kinnunen, Rahim Saeidi, Jouni Pohjalainen, Paavo Alku, Figen Ertas |
ICASSP | 5 |
| 2013 | Automatic detection of anger in telephone speech with robust autoregressive modulation filteringabstractA new system for automatic detection of angry speech is proposed. Using simulation of far-end-noise-corrupted telephone speech and the widely used Berlin database of emotional speech, autoregressive prediction of features across speech frames is shown to contribute significantly to both the clean speech performance and the robustness of the system. The autoregressive models are learned from the training data in order to capture long-term temporal dynamics of the features. Additionally, linear predictive spectrum analysis outperforms conventional Fourier spectrum analysis in terms of robustness in the computation of mel-frequency cepstral coefficients in the feature extraction stage. Jouni Pohjalainen, Paavo Alku |
ICASSP | 2 |
| 2013 | Comparing glottal-flow-excited statistical parametric speech synthesis methodsabstractThis paper studies the performance of glottal flow signal based excitation methods in statistical parametric speech synthesis. The current state of the art in excitationmodeling is reviewed and three excitation methods are selected for experiments. Two of the methods are based on the principal component analysis (PCA) decomposition of estimated glottal flow pulses. While the first one uses only the mean of the pulses, the second method uses 12 principal components in addition to the mean signal for modeling the glottal flow waveform. The third method utilizes a glottal flow pulse library from which pulses are selected according to target and concatenation costs. Subjective listening tests are carried out to determine the quality and similarity of the synthetic speech of one male and one female speaker. The results show that the PCA-based methods are rated best both in quality and similarity, but adding more components does not yield any improvements. Tuomo Raitio, Antti Suni, Martti Vainio, Paavo Alku |
ICASSP | 4 |
| 2013 | Quasi closed phase analysis for glottal inverse filtering
Manu Airaksinen, Brad H. Story, Paavo Alku |
INTERSPEECH | 3 |
| 2013 | Effect of MPEG audio compression on HMM-based speech synthesisabstractIn this paper, the effect of MPEG audio compression on HMMbased speech synthesis is studied. Speech signals are encoded with various compression rates and analyzed using the GlottHMM vocoder. Objec ... Bajibabu Bollepalli, Tuomo Raitio, Paavo Alku |
INTERSPEECH | 3 |
| 2013 | Robust formant detection using group delay function and stabilized weighted linear prediction
Dhananjaya Gowda, Jouni Pohjalainen, Mikko Kurimo, Paavo Alku |
INTERSPEECH | 4 |
| 2013 | Comparison of spectrum estimators in speaker verification: mismatch conditions induced by vocal effortabstractWe study the problem of vocal effort mismatch in speaker verification. Changes in speaker’s vocal effort induce changes in fundamental frequency (F0) and formant structure which introduce unwanted intra-speaker variations to features. We compare seven alternative spectrum estimators in the context of melfrequency cepstral coefficient (MFCC) extraction for speaker verification. The compared variants include traditional FFT spectrum and six parametric all-pole models. Experimental results on the NIST 2010 speaker recognition evaluation (SRE) corpus utilizing both GMM-UBM and more recent GMM supervector classifier indicate that spectrum estimation has a considerable impact on speaker verification accuracy under mismatched vocal effort conditions. The highest recognition accuracy was achieved using a particular variant of temporally weighted all-pole model, stabilized weighted linear prediction (SWLP). Index Terms: speaker recognition, vocal effort mismatch, spectrum estimation Cemal Hanilçi, Tomi Kinnunen, Padmanabhan Rajan, Jouni Pohjalainen, Paavo Alku, Figen Ertas |
INTERSPEECH | 5 |
| 2013 | Frequency-adaptive post-filtering for intelligibility enhancement of narrowband telephone speech
Emma Jokinen, Marko Takanen, Paavo Alku |
INTERSPEECH | 3 |
| 2013 | Speech quality prediction for artificial bandwidth extension algorithmsabstractDuring the transition period from narrowband to wideband speech transmission services, Artificial Bandwidth Extension (ABE) algorithms are able to reduce the perceptual degradation of narrowband-transmitted speech signals by extending the audio bandwidth. In this paper, we analyze whether the resulting speech quality can be predicted reliably with instrumental models. Estimations from the new ITU standard POLQA, its predecessor WB-PESQ and the diagnostic DIAL model are compared to subjective listener judgments. This comparison reveals that the instrumental measures are not fully able to cope with ABE-processed speech, particularly in predicting ABE rank orders reliably. Reasons for this finding and corresponding diagnoses are discussed. Index Terms: speech quality, artificial bandwidth extension, instrumental quality prediction, speech transmission, diagnosis Sebastian Möller 0001, Emilia Kelaidi, Friedemann Köster, Nicolas Côté, Patrick Bauer, Tim Fingscheidt, Thomas Schlien, Hannu Pulakka, Paavo Alku |
INTERSPEECH | 9 |
| 2013 | Extended weighted linear prediction using the autocorrelation snapshot - a robust speech analysis method and its application to recognition of vocal emotionsabstractTemporally weighted linear predictive methods have recently been successfully used for robust feature extraction in speech and speaker recognition. This paper introduces their general formulation, where various efficient temporal weighting functions can be included in the optimization of the all-pole coefficients of a linear predictive model. Temporal weighting is imposed by multiplying elements of instantaneous autocorrelation “snapshot” matrices computed from speech data. With this novel autocorrelation-snapshot formulation of weighted linear prediction, it is demonstrated that different temporal aspects of speech can be emphasized in order to enhance robustness of feature extraction in speech emotion recognition. Jouni Pohjalainen, Paavo Alku |
INTERSPEECH | 2 |
| 2013 | Analysis and synthesis of shouted speechabstractIn this study, the acoustic properties of shouted speech are analyzed in relation to normal speech, and various synthesis techniques for shouting are investigated. The analysis shows large differences between the two styles, which induces difficulties in synthesis. Analysis-synthesis experiments show that the use of spectral estimation methods that are not biased by the sparse harmonics of shouted speech is beneficial. The synthesis of shouting is performed through adaptation and voice conversion. Subjective evaluation of synthesis reveals that, despite quality degradation, the impression of shouting and use of vocal effort is fairly well preserved. In addition, the use of specific spectral estimation methods is found to be beneficial also in adaptation. Tuomo Raitio, Antti Suni, Jouni Pohjalainen, Manu Airaksinen, Martti Vainio, Paavo Alku |
INTERSPEECH | 6 |
| 2013 | Using group delay functions from all-pole models for speaker recognitionabstractPopular features for speech processing, such as mel-frequency cepstral coefficients (MFCCs), are derived from the short-term magnitude spectrum, whereas the phase spectrum remains un-used. While the common argument to use only the magnitude spectrum is that the human ear is phase-deaf, phase-based fea-tures have remained less explored due to additional signal pro-cessing difficulties they introduce. A useful representation of the phase is the group delay function, but its robust computa-tion remains difficult. This paper advocates the use of group delay functions derived from parametric all-pole models instead of their direct computation from the discrete Fourier transform. Using a subset of the vocal effort data in the NIST 2010 speaker recognition evaluation (SRE) corpus, we show that group delay features derived via parametric all-pole models improve recog-nition accuracy, especially under high vocal effort. Addition-ally, the group delay features provide comparable or improved accuracy over conventional magnitude-based MFCC features. Thus, the use of group delay functions derived from all-pole models provide an effective way to utilize information from the phase spectrum of speech signals. Index Terms: speaker verification, group delay functions, high vocal effort Padmanabhan Rajan, Tomi Kinnunen, Cemal Hanilçi, Jouni Pohjalainen, Paavo Alku |
INTERSPEECH | 5 |
| 2013 | Lombard modified text-to-speech synthesis for improved intelligibility: submission for the hurricane challenge 2013abstractThis paper describes modification of a TTS system for im-proving the intelligibility of speech in various noise conditions. First, the GlottHMM vocoder is used for training a voice with modal speech data. The vocoder and voice parameters are then modified to mimic the properties of Lombard effect based on a small amount of Lombard speech from the same speaker. More specifically, the durations are increased, fundamental frequency is raised, spectral tilt is decreased, the harmonic-to-noise ratio is increased, and a pressed glottal flow pulses are used in cre-ating excitation. The formants of the speech are also enhanced and finally the speech is compressed in order to increase noise robustness of the voice. The evaluation results of the Hurricane Challenge 2013 indicate that the modified voice is mostly less intelligible than the unmodified natural speech, as expected, but more intelligible than the reference TTS voice, especially in the low SNR conditions. Antti Suni, Reima Karhila, Tuomo Raitio, Mikko Kurimo, Martti Vainio, Paavo Alku |
INTERSPEECH | 6 |
| 2012 | Comparing spectrum estimators in speaker verification under additive noise degradationabstractDifferent short-term spectrum estimators for speaker verification under additive noise are considered. Conventionally, mel-frequency cepstral coefficients (MFCCs) are computed from discrete Fourier transform (DFT) spectra of windowed speech frames. Recently, linear prediction (LP) and its temporally weighted variants have been substituted as the spectrum analysis method in speech and speaker recognition. In this paper, 12 different short-term spectrum estimation methods are compared for speaker verification under additive noise contamination. Experimental results conducted on NIST 2002 SRE show that the spectrum estimation method has a large effect on recognition performance and stabilized weighted LP (SWLP) and minimum variance distortionless response (MVDR) methods yield approximately 7 % and 8 % relative improvements over the standard DFT method at −10 dB SNR level of factory and babble noises, respectively in terms of equal error rate (EER). Cemal Hanilçi, Tomi Kinnunen, Rahim Saeidi, Jouni Pohjalainen, Paavo Alku, Figen Ertas, Johan Sandberg, Maria Sandsten |
ICASSP | 5 |
| 2012 | Robust speech analysis by lag-weighted linear predictionabstractThis study introduces an approach for linear predictive spectrum analysis based on emphasizing selected time-domain properties in the analyzed signal in combination with a stabilization operation. A stable weighted linear predictive method based on a novel autocorrelation-based weighting scheme is described and its spectral properties are demonstrated. The robustness of the proposed method is compared with conventional techniques in terms of an Euclidean MFCC distortion measure in different additive noise conditions. In the experimental evaluation, the novel speech analysis technique outperforms the other evaluated methods. Jouni Pohjalainen, Paavo Alku |
ICASSP | 2 |
| 2012 | Conversational evaluation of artificial bandwidth extension of telephone speech using a mobile handsetabstractArtificial bandwidth extension methods have been developed to improve the quality and intelligibility of narrowband telephone speech. Bandwidth extension methods have typically been evaluated with objective measures or subjective listening-only tests, whereas realistic conversational evaluations have been rare. This paper presents a conversational evaluation of two bandwidth extension methods together with narrowband and wideband speech. The evaluation was performed using a mobile handset with a wired earpiece and microphone both in silence and in simulated street noise. The results indicate that one of the evaluated bandwidth extension methods was significantly preferred over narrowband speech in silence. The results also suggest slight preference for this bandwidth extension method over narrowband speech in street noise. True wideband speech was considered superior to bandwidth-extended and narrowband speech especially in silence. Hannu Pulakka, Laura Laaksonen, Ville Myllylä, Santeri Yrttiaho, Paavo Alku |
ICASSP | 5 |
| 2012 | On measuring the intelligibility of synthetic speech in noise - Do we need a realistic noise environment?abstractAssessing the intelligibility of synthetic speech is important in creating synthetic voices to be used in real life applications, especially for the ones involving interfering noise. This raises the question how to measure the intelligibility of synthetic speech to correctly simulate such conditions. Conventionally, this has been done using a simple listening test setup where diotic speech and noise are played to both ears with headphones. This is indeed very different from the real noise environment where speech and noise are spatially distributed. This paper addresses the question whether a realistic noise environment should be used to test the intelligibility of synthetic speech. Three different test conditions, one with multichannel reproduction of noise and speech, and two headphone setups are evaluated. Tests are performed with natural and synthetic speech, including speech especially intended for noisy conditions. The results indicate a general trend in all setups but also some interesting differences. Tuomo Raitio, Marko Takanen, Olli Santala, Antti Suni, Martti Vainio, Paavo Alku |
ICASSP | 6 |
| 2012 | Improved formant frequency estimation from high-pitched vowels by downgrading the contribution of the glottal source with weighted linear predictionabstractSince performance of conventional linear prediction (LP) deteriorates in formant estimation of high-pitched voices, several all-pole modeling methods robust to F0 have been developed. This study compares five such previously known methods and proposes a new technique, Weighted Linear Prediction with Attenuated Main Excitation (WLP-AME). WLP-AME utilizes weighted linear prediction in which the square of the prediction error is multiplied with a weighting function that downgrades the contribution of the glottal source in the model optimization. Consequently, the resulting all-pole model is affected more by the vocal tract characteristics, which leads to more accurate formant estimates. By using synthetic vowels created with a physical modeling approach, the study shows that WLP-AME yields improved formant frequency estimates for high-pitched vowels in comparison to the previously known methods. Paavo Alku, Jouni Pohjalainen, Martti Vainio, Anne-Maria Laukkanen, Brad H. Story |
INTERSPEECH | 1 |
| 2012 | Utilizing Markov Chain Monte Carlo (MCMC) Method for Improved Glottal Inverse FilteringabstractThis paper presents a new glottal inverse filtering (GIF) method that utilizes Markov chain Monte Carlo (MCMC) algorithm. First, initial estimates of the vocal tract and glottal flow are eval-uated by an existing GIF method, the iterative adaptive inverse filtering (IAIF). Simultaneously, the initially estimated glottal flow is synthesized using the Klatt model and filtered with the estimated vocal tract filter. In the MCMC estimation process, the first few poles of the initial vocal tract model and the Klatt parameter are refined in order to minimize the error between the original and the synthetic signals. MCMC converges to the optimal result, and the final estimate of the vocal tract is found by averaging the parameter values of the Markov chain. Ex-periments show that the MCMC-based GIF method gives more accurate results compared to the original IAIF method. Harri Auvinen, Tuomo Raitio, Samuli Siltanen, Paavo Alku |
INTERSPEECH | 4 |
| 2012 | Utilization of the Lombard effect in post-filtering for intelligibility enhancement of telephone speech
Emma Jokinen, Paavo Alku, Martti Vainio |
INTERSPEECH | 2 |
| 2012 | Towards Glottal Source Controllability in Expressive Speech SynthesisabstractIn order to obtain more human like sounding humanmachine interfaces we must first be able to give them expressive capabilities in the way of emotional and stylistic features so as to closely adequate them to the intended task. If we want to replicate those features it is not enough to merely replicate the prosodic information of fundamental frequency and speaking rhythm. The proposed additional layer is the modification of the glottal model, for which we make use of the GlottHMM parameters. This paper analyzes the viability of such an approach by verifying that the expressive nuances are captured by the aforementioned features, obtaining 95% recognition rates on styled speaking and 82% on emotional speech. Then we evaluate the effect of speaker bias and recording environment on the source modeling in order to quantify possible problems when analyzing multi-speaker databases. Finally we propose a speaking styles separation for Spanish based on prosodic features and check its perceptual significance. Jaime Lorenzo-Trueba, Roberto Barra-Chicote, Tuomo Raitio, Nicolas Obin, Paavo Alku, Junichi Yamagishi, Juan Manuel Montero-Martínez |
INTERSPEECH | 5 |
| 2012 | Voice source analysis using biomechanical modeling and glottal inverse filteringabstractThis paper studies the use of glottal inverse filtering together with a biomechanical model of the vocal folds to simulate the glottal flow waveform. The glottal flow waveform is first estimated by inverse filtering the acoustic speech pressure signal of natural speech. The estimated glottal flow is used as a template in an optimization process which searches for a set of parameters for a deterministic vocal fold model such that the model output reproduces the estimated glottal flow. The results indicate that the method can reproduce the main deterministic components of the glottal flow signal with good accuracy. Index Terms: vocal folds, glottal flow, biomechanical simulation, glottal inverse filtering. Alan Pinheiro, Tuomo Raitio, Danyane Gomes, Paavo Alku |
INTERSPEECH | 4 |
| 2012 | Automatic Detection of High Vocal Effort in Telephone SpeechabstractA system is proposed for the automatic detection of high vocal effort in speech. The system is evaluated using both PCMcoded speech and AMR-coded telephone speech. In addition, the effect of far-end noise in the telephone conditions is studied using both matched-condition training and cases with additive noise mismatch. The proposed system is based on Bayesian classification of mel-frequency cepstral feature vectors. Concerning the MFCC feature extraction process, the substitution of a spectrum analysis method emphasizing the fine structure improves the results in the noisy cases. Jouni Pohjalainen, Tuomo Raitio, Hannu Pulakka, Paavo Alku |
INTERSPEECH | 4 |
| 2012 | Wideband Parametric Speech Synthesis Using Warped Linear PredictionabstractThis paper studies the use of warped linear prediction (WLP) for wideband parametric speech synthesis. As the sampling fre-quency is increased from the usual 16 kHz, linear frequency res-olution of conventional linear prediction (LP) cannot efficiently model the speech spectrum. By using frequency warping that weights perceptually the most important formant information, spectral models with better accuracy and lower model orders can be utilized. In this work, WLP is embedded in a paramet-ric speech synthesizer to efficiently create wideband synthetic speech. Experiments show that WLP-based wideband synthetic speech is rated better compared to narrowband speech and wide-band LP-based speech. Index Terms: statistical parametric speech synthesis, wide-band, warped linear prediction, WLP Tuomo Raitio, Antti Suni, Martti Vainio, Paavo Alku |
INTERSPEECH | 4 |
| 2012 | Effect of noise type and level on focus related fundamental frequency changesabstractSpeech in noise, or Lombard speech, is characterized by increased intensity and higher fundamental frequency as well as lengthened segmental durations as speakers try to maintain a beneficial signal-to-noise ratio to fill both communicative and self-monitoring requirements. The phenomenon has been studied with regard to different noise types and different noise levels, as well as with respect to different communicative tasks (e.g., reading out loud vs. speaking to a real listener). However, there are no studies where the effect has been measured with different noises keeping the loudness levels equal. Here we study the Lombard effect with three different noise types at three levels with equal loudness while varying focus structure to elicit different pitch contours. The results show that people adapt their intonation contours depending on both noise level and type even when the noises are similar with respect to their perceived loudness. This points to a special role for pitch in Lombard speech. Martti Vainio, Daniel Aalto, Antti Suni, Anja Arnhold, Tuomo Raitio, Henri Seijo, Juhani Järvikivi, Paavo Alku |
INTERSPEECH | 8 |
| 2012 | Regularized All-Pole Models for Speaker Verification Under Noisy EnvironmentsabstractRegularization of linear prediction based mel-frequency cepstral coefficient (MFCC) extraction in speaker verification is considered. Commonly, MFCCs are extracted from the discrete Fourier transform (DFT) spectrum of speech frames. In this paper, DFT spectrum estimate is replaced with the recently proposed regularized linear prediction (RLP) method. Regularization of temporally weighted variants, weighted LP (WLP) and stabilized WLP (SWLP) which have earlier shown success in speech and speaker recognition, is also introduced. A novel type of double autocorrelation (DAC) lag windowing is also proposed to enhance robustness. Experiments on the NIST 2002 corpus indicate that regularized all-pole methods (RLP, RWLP and RSWLP) yield large improvement on recognition accuracy under additive factory and babble noise conditions in terms of both equal error rate (EER) and minimum detection cost function (MinDCF). Cemal Hanilçi, Tomi Kinnunen, Figen Ertas, Rahim Saeidi, Jouni Pohjalainen, Paavo Alku |
IEEE Signal Process. Lett. | 6 |
| 2012 | Conversational Evaluation of Speech Bandwidth Extension Using a Mobile HandsetabstractThe quality of narrowband telephone speech can be improved by artificial bandwidth extension. So far, bandwidth extension methods have been evaluated with objective measures and subjective listening-only tests, whereas realistic conversational evaluations have been rare. This article presents a conversational evaluation of two bandwidth extension methods in comparison with narrowband and wideband references. The evaluation was carried out using a mobile handset with a wired earpiece and microphone both in silence and in simulated street noise. The results indicate that speech processed with one of the bandwidth extension methods was preferred over narrowband speech. Hannu Pulakka, Laura Laaksonen, Ville Myllylä, Santeri Yrttiaho, Paavo Alku |
IEEE Signal Process. Lett. | 5 |
| 2012 | Bandwidth Extension of Telephone Speech to Low Frequencies Using Sinusoidal Synthesis and a Gaussian Mixture ModelabstractThe quality of narrowband telephone speech is degraded by the limited audio bandwidth. This paper describes a method that extends the bandwidth of telephone speech to the frequency range 0-300 Hz. The method generates the lowest harmonics of voiced speech using sinusoidal synthesis. The energy in the extension band is estimated from spectral features using a Gaussian mixture model. The amplitudes and phases of the synthesized sinusoidal components are adjusted based on the amplitudes and phases of the narrowband input speech, which provides adaptivity to varying input bandwidth characteristics. The proposed method was evaluated with listening tests in combination with another bandwidth extension method for the frequency range 4-8 kHz. While the low-frequency bandwidth extension was not found to improve perceived quality, the method reduced dissimilarity with wideband speech. Hannu Pulakka, Ulpu Remes, Santeri Yrttiaho, Kalle J. Palomäki, Mikko Kurimo, Paavo Alku |
IEEE Trans. Speech Audio Process. | 6 |
| 2011 | Glottal inverse filtering using stabilised weighted linear predictionabstractThis paper presents and evaluates an inverse filtering technique of the speech signal which is based on the Stabilized Weighted Lin ear Prediction (SWLP) of speech. SWLP emphasizes the speech samples that fit the underlying speech production model well, by imposing temporal weighting of the square of the residual signal. The performance of SWLP is compared to the conventional Linear Prediction based inverse filtering techniques, such as the Autocorrelation and Closed Phase Covariance method. All the inverse filtering approaches are evaluated on a database of speech signals generated by a physical model of the voice production system. Results show that the estimated glottal flows using SWLP are closer to the original glottal flow than those estimated by the Autocorrelation approach, while its performance is comparable to the Closed Phase Covariance approach. George P. Kafentzis, Yannis Stylianou, Paavo Alku |
ICASSP | 3 |
| 2011 | Shout detection in noiseabstractFor the task of detecting shouted speech in a noisy environment, this paper introduces a system based on mel frequency cepstral coefficient (MFCC) feature extraction, unsupervised frame dropping and Gaussian mixture model (GMM) classification. The evaluation material consists of phonemically identical speech and shouting as well as environmental noise of varying levels. The performance of the shout detection system is analyzed by varying the MFCC feature extraction with respect to 1) the feature vector length and 2) the spectrum estimation method. As for feature vector length, the best performance is obtained using 30 MFCC coefficients, which is more than what is conventionally used. In spectrum estimation, a scheme that combines a linear prediction spectrum envelope with spectral fine structure outperforms the conventional FFT. Jouni Pohjalainen, Paavo Alku, Tomi Kinnunen |
ICASSP | 2 |
| 2011 | Speech bandwidth extension using Gaussian mixture model-based estimation of the highband mel spectrumabstractThe quality and intelligibility of narrowband telephone speech can be enhanced by artificial bandwidth extension. This study combines Gaussian mixture model-based (GMM) mel spectrum extension with a filter bank implementation for generating the missing spectral content in the highband at 4-8 kHz. The narrowband mel spectrum is calculated from input speech and the GMM is used to estimate the mel spectrum in the highband. An excitation signal for the highband is generated as a combination of upsampled linear prediction residual and modulated noise. The excitation is divided into sub-bands that are weighted and summed to realize the estimated mel spectrum. The bandwidth-extended output is obtained as the sum of the artificial highband signal and narrowband speech. Listening tests indicate that this method is preferred over narrowband speech and over a previously presented artificial bandwidth extension method which is implemented in some mobile phone models. Hannu Pulakka, Ulpu Remes, Kalle J. Palomäki, Mikko Kurimo, Paavo Alku |
ICASSP | 5 |
| 2011 | Utilizing glottal source pulse library for generating improved excitation signal for HMM-based speech synthesisabstractThis paper describes a source modeling method for hidden Markov model (HMM) based speech synthesis for improved naturalness. A speech corpus is first decomposed into the glottal source signal and the model of the vocal tract filter using glottal inverse filtering, and parametrized into excitation and spectral features. Additionally, a library of glottal source pulses is extracted from the estimated voice source signal. In the synthesis stage, the excitation signal is generated by selecting appropriate pulses from the library according to the target cost of the excitation features and a concatenation cost between adjacent glottal source pulses. Finally, speech is synthesized by filtering the excitation signal by the vocal tract filter. Experiments show that the naturalness of the synthetic speech is better or equal, and speaker similarity is better, compared to a system using only single glottal source pulse. Tuomo Raitio, Antti Suni, Hannu Pulakka, Martti Vainio, Paavo Alku |
ICASSP | 5 |
| 2011 | Noise Robust Feature Extraction Based on Extended Weighted Linear Prediction in LVCSRabstractThis paper introduces extended weighted linear prediction (XLP) to noise robust short-time spectrum analysis in the feature extraction process of a speech recognition system. XLP is a generalization of standard linear prediction (LP) and temporally weighted linear prediction (WLP) which have already been applied to noise robust speech recognition with good results. With XLP, higher controllability to the temporal weighting of different parts of the noisy speech is gained by taking the lags of the signal into account in prediction. Here, the performance of XLP is put up against WLP and conventional spectrum analysis methods FFT and LP on a large vocabulary continuous speech recognition (LVCSR) scheme using real world noisy data containing additive and convolutive noise. The results show improvements over the reference methods in several cases. Index Terms: linear prediction, temporal weighting, noise robust, speech recognition Sami Keronen, Jouni Pohjalainen, Paavo Alku, Mikko Kurimo |
INTERSPEECH | 3 |
| 2011 | Detection of Shouted Speech in the Presence of Ambient NoiseabstractThis study focuses on the detection of shouted speech in realistic noisy conditions. An automatic system based on modified mel frequency cepstral coefficient (MFCC) feature extraction and Gaussian mixture model (GMM) classification is developed. The performance of the automatic system is compared against human perception measured by a listening test. At moderate noise levels, the automatic system outperforms humans. In severe conditions, classification by humans is clearly better. Jouni Pohjalainen, Tuomo Raitio, Paavo Alku |
INTERSPEECH | 3 |
| 2011 | Low-Frequency Bandwidth Extension of Telephone Speech Using Sinusoidal Synthesis and Gaussian Mixture Model
Hannu Pulakka, Ulpu Remes, Santeri Yrttiaho, Kalle J. Palomäki, Mikko Kurimo, Paavo Alku |
INTERSPEECH | 6 |
| 2011 | Analysis of HMM-Based Lombard Speech SynthesisabstractHumans modify their voice in interfering noise in order to maintain the intelligibility of their speech – this is called the Lombard effect. This ability, however, has not been extensively modeled in speech synthesis. Here we compare several methods of synthesizing speech in noise using a physiologically based statistical speech synthesis system (GlottHMM). The results show that in a realistic street noise situation the synthetic Lombard speech is judged by listeners both as appropriate for the situation and as intelligible as natural Lombard speech. Of the different types of models, one using adaptation and extrapolation performed the best. Tuomo Raitio, Antti Suni, Martti Vainio, Paavo Alku |
INTERSPEECH | 4 |
| 2011 | Bandwidth Extension of Telephone Speech Using a Neural Network and a Filter Bank Implementation for Highband Mel SpectrumabstractThe limited audio bandwidth used in narrowband telephone systems degrades both the quality and the intelligibility of speech. This paper presents a new method for the bandwidth extension of telephone speech. Frequency components are added to the frequency band 4-8 kHz using only the information in the narrowband speech. A neural network is used to estimate the mel spectrum in the extension band in short time frames based on features calculated from the narrowband speech. A wideband excitation signal is generated by spectral folding from the narrowband linear prediction residual and a filter bank is utilized to divide the excitation into four sub-bands that cover the extension band. These sub-bands are weighted such that the estimated mel spectrum is realized. Bandwidth-extended speech is obtained by summing the weighted sub-bands and the original narrowband signal. Listening tests show that this new method improves speech quality compared with narrowband telephone speech and with a previously published bandwidth extension method. Hannu Pulakka, Paavo Alku |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | HMM-Based Speech Synthesis Utilizing Glottal Inverse FilteringabstractThis paper describes an hidden Markov model (HMM)-based speech synthesizer that utilizes glottal inverse filtering for generating natural sounding synthetic speech. In the proposed method, speech is first decomposed into the glottal source signal and the model of the vocal tract filter through glottal inverse filtering, and thus parametrized into excitation and spectral features. The source and filter features are modeled individually in the framework of HMM and generated in the synthesis stage according to the text input. The glottal excitation is synthesized through interpolating and concatenating natural glottal flow pulses, and the excitation signal is further modified according to the spectrum of the desired voice source characteristics. Speech is synthesized by filtering the reconstructed source signal with the vocal tract filter. Experiments show that the proposed system is capable of generating natural sounding speech, and the quality is clearly better compared to two HMM-based speech synthesis systems based on widely used vocoder techniques. Tuomo Raitio, Antti Suni, Junichi Yamagishi, Hannu Pulakka, Jani Nurminen, Martti Vainio, Paavo Alku |
IEEE Trans. Speech Audio Process. | 7 |
| 2010 | Extended weighted linear prediction (XLP) analysis of speech and its application to speaker verification in adverse conditionsabstractThis paper introduces a generalized formulation of linear prediction (LP), including both conventional and temporally weighted LP analysis methods as special cases. The temporally weighted methods have recently been successfully applied to noise robust spectrum analysis in speech and speaker recognition applications. In comparison to those earlier methods, the new generalized approach allows more versatility in weighting different parts of the data in the LP analysis. Two such weighted methods are evaluated and compared to the conventional spectrum modeling methods FFT and LP, as well as the temporally weighted methods WLP and SWLP, by substituting each of them in turn as the spectrum estimation method of the MFCC feature extraction stage of a GMM-UBM based speaker verification system. The new methods are shown to lead to performance improvement in several cases involving channel distortion and additive noise mismatch between the training and recognition conditions. Index Terms: linear prediction, speaker verification, mel frequency cepstral coefficients Jouni Pohjalainen, Rahim Saeidi, Tomi Kinnunen, Paavo Alku |
INTERSPEECH | 4 |
| 2010 | Laryngeal voice quality in the expression of focus
Martti Vainio, Matti Airas, Juhani Järvikivi, Paavo Alku |
INTERSPEECH | 4 |
| 2010 | Temporally Weighted Linear Prediction Features for Tackling Additive Noise in Speaker VerificationabstractText-independent speaker verification under additive noise corruption is considered. In the popular mel-frequency cepstral coefficient (MFCC) front-end, the conventional Fourier-based spectrum estimation is substituted with weighted linear predictive methods, which have earlier shown success in noise-robust speech recognition. Two temporally weighted variants of linear predictive modeling are introduced to speaker verification and they are compared to FFT, which is normally used in computing MFCCs, and to conventional linear prediction. The effect of speech enhancement (spectral subtraction) on the system performance with each of the four feature representations is also investigated. Experiments by the authors on the NIST 2002 SRE corpus indicate that the accuracy of the conventional and proposed features are close to each other on clean data. For factory noise at 0 dB SNR level, baseline FFT and the better of the proposed features give EERs of 17.4% and 15.6%, respectively. These accuracies improve to 11.6% and 11.2%, respectively, when spectral subtraction is included as a preprocessing method. The new features hold a promise for noise-robust speaker verification. Rahim Saeidi, Jouni Pohjalainen, Tomi Kinnunen, Paavo Alku |
IEEE Signal Process. Lett. | 4 |
| 2009 | On separating glottal source and vocal tract information in telephony speaker verificationabstractThe popular mel-frequency cepstral coefficients (MFCCs) capture a mixture of speaker-related, phonemic and channel information. Speaker-related information could be further broken down according to articulatory criteria. How these underlying components are exactly mixed in the features is not well understood. To this end, in this paper we aim at separating the spectra of glottal source and vocal tract using glottal inverse filtering, with an application to speaker recognition over telephone lines. Our experiments on the 10 sec-10 sec condition of the NIST 2006 SRE corpus suggest that the mel-frequency cepstrum of the voice source is not too useful for recognizing speakers. On the contrary, fusing the vocal tract spectrum with conventional MFCCs improves accuracy, suggesting that vocal tract information should be enhanced. Tomi Kinnunen, Paavo Alku |
ICASSP | 2 |
| 2009 | Weighted linear prediction for speech analysis in noisy conditionsabstract1 k p; where E = X n (s n p X k=1 a k s n k ) 2 I WLP is a generalization of LP, introducing temporal weighting of the squared prediction error [2]: E = X n (s n p X k=1 a k s n k ) 2 W n IW n is the weighting function I For constant W n , WLP becomes identical to LP! Stabilized Weighted Linear Prediction (SWLP) Jouni Pohjalainen, Heikki Kallasjoki, Kalle J. Palomäki, Mikko Kurimo, Paavo Alku |
INTERSPEECH | 5 |
| 2009 | New method for delexicalization and its application to prosodic tagging for text-to-speech synthesisabstractThis paper describes a new flexible delexicalization method based on glottal excited parametric speech synthesis scheme. The system utilizes inverse filtered glottal flow and all-pole modelling of the vocal tract. The method provides a possibility to retain and manipulate all relevant prosodic features of any kind of speech. Most importantly, the features include voice quality, which has not been properly modeled in earlier delexicalization methods. The functionality of the new method was tested in a prosodic tagging experiment aimed at providing word prominence data for a text-to-speech synthesis system. The experiment confirmed the usefulness of the method and further corroborated earlier evidence that linguistic factors influence the perception of prosodic prominence. Martti Vainio, Antti Suni, Tuomo Raitio, Jani Nurminen, Juhani Järvikivi, Paavo Alku |
INTERSPEECH | 6 |
| 2009 | Stabilised weighted linear prediction
Carlo Magi, Jouni Pohjalainen, Tom Bäckström, Paavo Alku |
Speech Commun. | 4 |
| 2008 | DC-constrained linear prediction for glottal inverse filtering
Paavo Alku, Carlo Magi, Tom Bäckström |
INTERSPEECH | 1 |
| 2008 | HMM-based Finnish text-to-speech system utilizing glottal inverse filteringabstractAbstract This paper describes an HMM-based speech synthesis sys-tem that utilizes glottal inverse filtering for generating naturalsounding synthetic speech. In the proposed system, speech isfirst parametrized into spectral and excitation features using aglottal inverse filtering based method. The parameters are fedinto an HMM system for training and then generated from thetrained HMM according to text input. Glottal flow pulses ex-tracted from real speech are used as a voice source, and thevoice source is further modified according to the all-pole modelparameters generated by the HMM. Preliminary experimentsshow that the proposed system is capable of generating naturalsounding speech, and the quality is clearly better compared to asystem utilizing a conventional impulse train excitation model.Index Terms: speech synthesis, glottal inverse filtering, HMM 1. Introduction The ultimate goal of text-to-speech synthesis (TTS) is to enablecreating natural sounding speech from arbitrary text. More-over, the current trend in TTS research calls for systems thatenable producing speech in different speaking styles with dif-ferent speaker characteristics and even emotions. In order tofulfill these stringent general requirements, two major synthe-sis techniques have attracted increasing interest in the speechresearch community during the past decade. These two alter-natives are (1) the unit selection technique and (2) the hiddenMarkov model (HMM) based approach. The former has beenshown to yield synthetic speech of highly natural quality. How-ever, unit selection techniques do not allow for easy adaptationof the TTS system to different speaking styles and speaker char-acteristics. In addition, their implementation requires databasesof extensive sizes, which severely limit the use of this TTS tech-nique, for example, in mobile terminals. HMM-based tech-niques, in turn, benefit from better adaptability and a clearlysmaller memory requirement. However, the current HMM sys-tems often suffer from degraded naturalness in quality. It canbe argued that a potential reason for the reduced naturalness inthe current HMM-based TTS systems can be explained by theuse of signal generation techniques which are oversimplified toproperly mimic natural speech pressure waveforms.A large part of what can be characterized as naturalnessin speech emerges from different voice characteristics as wellas their context dependent changes. Therefore, it is justifiedin speech synthesis to search for methods aiming at accuratemodeling of different voice characteristics as well as prosodicfeatures of speech. Towards these goals, HMM-based synthe-sizers have been developed with special emphasis on voice char-acteristics such as speaker individualities, speaking styles, andemotions [1]. Moreover, some recent studies have introducedimprovements to the parametric HMM systems’ signal genera-tion techniques by utilizing, for example, mixed excitation [2]and residual modeling [3]. These techniques have been shownto improve the quality of synthetic speech compared to systemsutilizing a traditional impulse train excitation model. However,the quality of the systems using these techniques still remainsfar from the quality of natural speech.In the real human voice production mechanism, the excita-tion of (voiced) speech is represented by the glottal volume ve-locity waveform generated by the vibrating vocal folds. This ex-citation signal, the glottal source, has naturally attracted interestin speech synthesis and many techniques have been proposed tomimic the glottal source of natural speech. One such techniqueis the Liljencrants-Fant (LF) model of the differentiated glottalsource that has been used both in traditional rule-based synthe-sis [4, 5] as well as within an HMM-based speech synthesizer[6]. However, the use of artificial glottal flow pulses usuallyresults in a somewhat buzzy quality due to a strong harmonicstructure at higher frequencies. To overcome this problem, theidea of utilizing glottal flow pulses extracted from real speechwith the help of glottal inverse filtering has been proposed [7, 8].However, previous studies based on glottal flow pulses extractedfrom natural speech are limited to special purposes such as thegeneration of isolated vowels, and the benefits from combiningautomatic glottal inverse filtering with an HMM-based speechsynthesizer have not been utilized.In this paper, a novel HMM-based speech synthesis sys-tem that utilizes glottal inverse filtering for generating naturalsounding synthetic speech is presented. The rest of the paper isorganized as follows: Section 2 describes the proposed speechsynthesis system. The results of the experiments with the newsynthesizer are presented in Section 3. Discussion on the pro-posed speech synthesis system and future plans are presented inSection 4, and final conclusions are presented in Section 5. Tuomo Raitio, Antti Suni, Hannu Pulakka, Martti Vainio, Paavo Alku |
INTERSPEECH | 5 |
| 2008 | Simple proofs of root locations of two symmetric linear prediction models
Carlo Magi, Tom Bäckström, Paavo Alku |
Signal Process. | 3 |
| 2008 | Evaluation of an Artificial Speech Bandwidth Extension Method in Three LanguagesabstractQuality and intelligibility of narrowband telephone speech can be improved by artificial bandwidth extension (ABE), which extends the speech bandwidth using only information available in the narrowband speech signal. This paper reports a three-language evaluation of an ABE method that has recently been launched in several of Nokia's mobile telephone models. The method extends the speech bandwidth to frequencies above the telephone band by first utilizing spectral folding and then modifying the magnitude spectrum of the extension band with spline curves. The performance of the method was evaluated by formal listening tests in American English, Russian, and Mandarin Chinese. The results of the listening tests indicate that ABE processing improved the subjective quality of coded narrowband speech in all these languages. Differences between bandwidth-extended American English test sentences and their original wideband counterparts were also evaluated using both an objective distance measure that simulates the characteristics of human hearing and a conventional spectral distortion measure. The average objective error was calculated for different categories of speech sounds. The error was found to be smallest in nasals and semivowels and largest in fricative sounds. Hannu Pulakka, Laura Laaksonen, Martti Vainio, Jouni Pohjalainen, Paavo Alku |
IEEE Trans. Speech Audio Process. | 5 |
| 2007 | Comparison of multiple voice source parameters in different phonation typesabstractA large sample of vowels produced by male and female speakers were inverse filtered and parameterized using 21 different glottal flow parameters. The performance of the different parameters in expression of the phonation type was then tested using objective statistical methods. The comparison of the results revealed marked differences in the parameters ’ performance, and therefore, guidelines for parameter use and comparison were established. Index Terms: voice quality, phonation type, inverse filtering, voice source, parameterization Matti Airas, Paavo Alku |
INTERSPEECH | 2 |
| 2007 | Stabilised weighted linear prediction - a robust all-pole method for speech processing
Carlo Magi, Tom Bäckström, Paavo Alku |
INTERSPEECH | 3 |
| 2007 | The effect of highband harmonic structure in the artificial bandwidth expansion of telephone speech
Hannu Pulakka, Paavo Alku, Laura Laaksonen, Päivi Valve |
INTERSPEECH | 2 |
| 2007 | Minimum Separation of Line Spectral FrequenciesabstractWe provide a theoretical lower limit on the distance of line spectral frequencies for both the line spectrum pair decomposition and the immittance spectrum pair decomposition. The result applies to line spectral frequencies computed from linear predictive polynomials with all roots within a zero-centered circle of radius r<1 Tom Bäckström, Carlo Magi, Paavo Alku |
IEEE Signal Process. Lett. | 3 |
| 2007 | Neural Network-Based Artificial Bandwidth Expansion of SpeechabstractThe limited bandwidth of 0.3-3.4 kHz in current telephone systems reduces both the quality and the intelligibility of speech. Artificial bandwidth expansion is a method that expands the bandwidth of the narrowband speech signal in the receiving end of the transmission link by adding new frequency components to the higher frequencies, i.e., up to 8 kHz. In this paper, a new method for artificial bandwidth expansion, termed Neuroevolution Artificial Bandwidth Expansion (NEABE) is proposed. The method uses spectral folding to create the initial spectral components above the telephone band. The spectral envelope is then shaped in the frequency domain, based on a set of parameters given by a neural network. Subjective listening tests were used to evaluate the performance of the proposed algorithm, and the results showed that NEABE speech was preferred over narrowband speech in about 80% of the test cases Juho Kontio, Laura Laaksonen, Paavo Alku |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Quality improvement of telephone speech by artificial bandwidth expansion - listening tests in three languages
Hannu Pulakka, Laura Laaksonen, Paavo Alku |
INTERSPEECH | 3 |
| 2005 | Objective Quality Measures for Glottal Inverse Filtering of Speech Pressure SignalsabstractGlottal inverse filtering is a process where the effects of the vocal tract are cancelled from the speech signal in order to estimate the voice source. Traditionally, inverse filtering methods have involved a high level of manual tuning of parameters, such as the vocal tract model order. We present objective heuristics for the measurement of the quality of the resulting glottal flow estimate. In addition, we propose an automatic method for determining the order of the vocal tract all-pole model in inverse filtering based on phase-plane analysis and estimation of the glottal flow kurtosis. Tom Bäckström, Matti Airas, Laura Lehto, Paavo Alku |
ICASSP (1) | 4 |
| 2005 | Artificial Bandwidth Expansion Method to Improve Intelligibility and Quality of AMR-Coded Narrowband SpeechabstractSpeech quality suffers from the limited bandwidth of cellular telephone systems, making it sound muffled. In addition, intelligibility is degraded due to missing higher frequency components. The proposed enhancement system is designed to improve both intelligibility and quality of narrowband speech by expanding the bandwidth and creating new spectral components to high frequencies in the receiving end of the transmission link. The algorithm can be used together with conventional narrowband speech codecs and it is designed to be robust in different noise conditions. In addition, the computational load of the algorithm is reasonable. Laura Laaksonen, Juho Kontio, Paavo Alku |
ICASSP (1) | 3 |
| 2005 | A toolkit for voice inverse filtering and parametrisation
Matti Airas, Hannu Pulakka, Tom Bäckström, Paavo Alku |
INTERSPEECH | 4 |
| 2005 | Group delay function as a means to assess quality of glottal inverse filtering
Paavo Alku, Matti Airas, Tom Bäckström, Hannu Pulakka |
INTERSPEECH | 1 |
| 2005 | Subglottal pressure and NAQ variation in voice production of classically trained baritone singersabstractThe subglottal pressure (Ps) and voice source characteristics of five professional baritone singers were analyzed. Glottal adduction was estimated with amplitude quotient (AQ), defined as the ratio between peak-to-peak pulse amplitude and the negative peak of the differentiated flow glottogram, and with normalized amplitude quotient (NAQ), defined as AQ divided by fundamental period length. Previous studies show that NAQ and its variation with Ps represent an effective parameter in the analysis of voice source characteristics. Therefore, the present study aims at increasing our knowledge of these two parameters further by finding out how they vary with pitch and Ps in operatic baritone singers, singing at high and low pitch. Ten equally spaced Ps values were selected from three takes of the syllable [pae], repeated with a continuously decreasing vocal loudness and initiated at maximum vocal loudness. The vowel sounds following the selected Ps peaks were inverse filtered. Data on peak-to-peak pulse amplitude, maximum flow declination rate, AQ and NAQ will be presented. Eva Björkner, Johan Sundberg, Paavo Alku |
INTERSPEECH | 3 |
| 2004 | Evaluation of an inverse filtering technique using physical modeling of voice productionabstractGlottal flows and sound pressure waveforms of four different fundamental frequencies were generated using a computational model of vocal fold vibration and acoustic wave propagation in order to evaluate the performance of an inverse filtering method. Four time-based parameters of the glottal flow were used in order to assess the accuracy of the inverse filtering technique. The results show that for most of the cases analyzed the relative error was less than 5 % when the time-based parameters extracted from the estimated glottal flows were compared to those obtained from the original flow waveforms produced by the physical model of the vocal fold vibration. Paavo Alku, Matti Airas, Brad H. Story |
INTERSPEECH | 1 |
| 2004 | Analysis of the voice source in different phonation types: simultaneous high-sped imaging of the vocal fold vibration and glottal inverse filteringabstractGlottal flow waveforms estimated by inverse filtering acoustic speech pressure signals were compared to glottal area functions obtained by digital high-speed imaging of the vocal fold vibration. Sp ... Hannu Pulakka, Paavo Alku, Svante Granqvist, Stellan Hertegard, Hans Larsson, Anne-Maria Laukkanen, Per-Ake Lindestad, Erkki Vilkman |
INTERSPEECH | 2 |
| 2004 | Linear predictive method for improved spectral modeling of lower frequencies of speech with small prediction ordersabstractAn all-pole modeling technique, Linear Prediction with Low-frequency Emphasis (LPLE), which emphasizes the lower frequency range of the input signal, is presented. The method is based on first interpreting conventional linear predictive (LP) analyses of successive prediction orders with parallel structures using the concept of symmetric linear prediction. In these implementations, symmetric linear prediction is preceded by simple pre-filters, which are of either low or high frequency characteristics. Combining those symmetric linear predictors that are not preceded by high-frequency pre-filters yields the proposed LPLE predictor. It is proved that the all-pole filters computed by LPLE are always stable. The results achieved with vowels show that the proposed method is well-suited for those applications, where low-order all-pole models with improved modeling of the lowest formants, are needed. Paavo Alku, Tom Bäckström |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | A time-domain interpretation for the LSP decompositionabstractThe line spectrum pair (LSP) decomposition is a widely used method in speech coding. In this article, we will show that the LSP polynomials, whose trivial zeros have been removed, are equivalent to two optimal (in the mean square sense) predictors in which a sample is predicted from linear combinations of its previous averaged and differentiated values. Tom Bäckström, Paavo Alku, Tuomas Paatero, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | All-pole modeling of wide-band speech with symmetric linear predictionabstractA new linear predictive technique, all-pole modeling with symmetric linear prediction (ASLP), is presented. The starting point of the method is an implementation of conventional linear prediction (LP) with a parallel structure, where two symmetric linear predictors are combined to prefilters represented by first order FIRs. Modification of these prefilters yields the ASLP predictor, which is always minimum phase. Experiments indicate that the new method models the formant structure of wide-band speech more accurately than conventional LP, when the prediction order is smaller than the one required by the sampling frequency. Paavo Alku, Tom Bäckström |
ICASSP (1) | 1 |
| 2003 | On the stability of constrained linear predictive modelsabstractStability of the all-pole model in conventional, unconstrained linear prediction with the autocorrelation criterion is well known. By exerting constraints to the optimisation problem it is possible to define models of order m + l with m parameters. However, traditionally constraints have led to models whose stability is not guaranteed. In this paper, we discuss constrained linear predictive models where the constraint is one-dimensional (l = 1) and derive stability criteria for these models. Tom Bäckström, Paavo Alku |
ICASSP (6) | 2 |
| 2003 | Linear predictive method with low-frequency emphasis
Paavo Alku, Tom Bäckström |
INTERSPEECH | 1 |
| 2003 | A constrained linear predictive model with the minimum-phase property
Tom Bäckström, Paavo Alku |
Signal Process. | 2 |
| 2003 | All-pole modeling technique based on weighted sum of LSP polynomialsabstractThis study presents a new technique called weighted-sum line spectrum pair (WLSP) where an all-pole filter is defined by using a sum of weighted line spectrum pair polynomials. The WLSP yields a stable all-pole filter of order m, whose autocorrelation function coincides with that of the input signal between indices 0 and m-1. By sacrificing the exact matching at index m, the WLSP models the autocorrelation of the input signal at the indices above m more accurately than conventional linear prediction (LP). Experiments with vowels show that, in comparison to the conventional LP, WLSP yields all-pole spectra that model formants with an increased dynamic range between formant peaks and spectral valleys. Tom Bäckström, Paavo Alku |
IEEE Signal Process. Lett. | 2 |
| 2003 | On line spectral frequenciesabstractThe commonly used line spectral frequencies form the roots of symmetric and antisymmetric polynomials constructed from a linear predictor. We provide a new, simpler proof that the symmetric and antisymmetric polynomials can be regarded as optimal constrained predictors that correspond to predicting from the low-pass and high-pass filtered signal, respectively. W. Bastiaan Kleijn, Tom Bäckström, Paavo Alku |
IEEE Signal Process. Lett. | 3 |
| 2002 | All-pole modeling technique based on the Weighted Sum of the LSP polynomialsabstractWith the order of prediction equal to m, conventional linear prediction (LP) yields an all-pole filter that matches exactly the autocorrelation function of the input signal between indices 0 and m. This study presents a new technique, Weighted-Sum Line Spectrum Pair (WLSP), where an all-pole filter is defined by using a sum of weighted LSP polynomials. WLSP yields a stable all-pole filter of order m, whose autocorrelation function coincides to that of the input signal between indices 0 and m −1. By sacrificing the exact matching of the autocorrelations at index m, WLSP models the autocorrelation of the input signal at the indices above m more accurately than conventional LP. The current paper presents mathematical properties of WLSP together with preliminary results on modeling of vowels. The results indicate that WLSP, in comparison to conventional LP is able to yield all-pole filters that model especially the upper formants with a larger dynamic range between formant peaks and spectral valleys. Paavo Alku, Tom Bäckström |
ICASSP | 1 |
| 2002 | A time domain reformulation of linear prediction equivalent to the LSP decompositionabstractThe Line spectrum pair (LSP) decomposition is a widely used method in speech coding. In this paper, we will present a reformulation of conventional linear prediction which is equivalent to the LSP decomposition. The paper shows that the symmetric and antisymmetric polynomials of the LSP decomposition are equivalent to two filters, that are determined by predicting a signal sample using its averaged and differentiated previous values. Tom Bäckström, Paavo Alku, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2002 | All-pole modeling of wide-band speech using weighted sum of the LSP polynomials
Paavo Alku, Tom Bäckström |
INTERSPEECH | 1 |
| 2002 | Measuring the effect of fundamental frequency raising as a strategy for increasing vocal intensity in soft, normal and loud phonation
Paavo Alku, Juha Vintturi, Erkki Vilkman |
Speech Commun. | 1 |
| 2002 | Time-domain parameterization of the closing phase of glottal airflow waveform from voices over a large intensity rangeabstractThe aim of this paper is to analyze and compare two time-domain parameterization methods of the glottal flow waveform on a large intensity range. The first parameter is the classical closing quotient which indicates the portion of a period where the glottis is closing. The second parameter is the normalized amplitude quotient which is defined using the ratio between the maximum flow amplitude and the negative peak amplitude of the differentiated glottal flow. The parameters are shown to be strongly correlated, and the normalized amplitude quotient to be a more accurate, consistent and robust measure than the closing quotient. The subjects, five female and six male, produced sustained phonations on a large intensity range. On this material, the normalized amplitude quotient is shown to vary systematically with sound pressure level, and it reveals information that for the closing quotient is hidden in the local variance. Tom Bäckström, Paavo Alku, Erkki Vilkman |
IEEE Trans. Speech Audio Process. | 2 |
| 2001 | The use of fundamental frequency raising as a strategy for increasing vocal intensity in soft, normal, and loud phonationabstractA method is presented to estimate the effect of intentional raising of fundamental frequency (F0) on vocal intensity. The method, Energy of the Synthesised Period (ESP), is based on computation of the energy of a hypothetical speech sound synthesised using a single period of the glottal volume velocity waveform and a digital filter that models the vocal tract. Both the glottal flow and the vocal tract filter are estimated by inverse filtering. The results show that, in producing loud voice, speakers use F0 to increase the number of glottal closures per time unit, which increases rapid fluctuations in the speech pressure waveform, which, in turn, raises vocal intensity. The average increase of sound pressure level due to this active use of F0 was approximately 4 dB in loud speech. Paavo Alku, Juha Vintturi, Erkki Vilkman |
INTERSPEECH | 1 |
| 2001 | One-delayed-mass model for efficient synthesis of glottal flowabstractA lumped physical model of the glottal source is presented. Vocal folds are described as single masses, but vertical phase differences between upper and lower margins of the folds are taken into account by appropriately describing the non-linear interaction of the mechanical model with aerodynamics. This results in a modified one-mass model, or a “one-delayed-mass model”. Analysis on numerical simulations shows that the system behaves qualitatively as higher-dimensional models (such as the two-mass model by Ishizaka and Flanagan); in particular, control over flow skewness is guaranteed, allowing for synthesis of realistic glottal flow waveforms. As only one degree of freedom (one mass) is needed in the model, structure and number of parameters are drastically reduced, thus making it suitable for real-time synthesis applications. Federico Avanzini, Paavo Alku, Matti Karjalainen |
INTERSPEECH | 2 |
| 2000 | Analysis of voice production in breathy, normal and pressed phonation by comparing inverse filtering and videokymography
Paavo Alku, Jan G. Svec, Erkki Vilkman, Frantisek Sram |
INTERSPEECH | 1 |
| 2000 | MEG-measurements of brain activity reveal the link between human speech production and perception
Paavo Alku, Hannu Tiitinen, Kalle J. Palomäki, Päivi Sivonen |
INTERSPEECH | 1 |
| 2000 | Neuromagnetic study on localization of speech sounds
Kalle J. Palomäki, Paavo Alku, Ville Mäkinen, Patrick J. C. May, Hannu Tiitinen |
INTERSPEECH | 2 |
| 2000 | A linear predictive method for highly compressed presentation of speech spectraabstractOur study proposes a new linear predictive algorithm, Linear Prediction with Sample Grouping (LPSG), for spectral modelling of speech. This method reformulates computation of linear prediction by grouping and extrapolating samples used in the prediction. In LPSG the number of samples used in the computation of the prediction is larger than the number of parameters to define the optimal predictor. Consequently, the proposed method makes it possible to obtain all-pole models for speech spectra that can be defined with a very compressed set of parameters. Quantisation of the prediction parameters of LPSG was compared in the present study to conventional linear prediction (LP) using a very low order of prediction. It appeared that LPSG yields better spectral matching and smaller residual energies in comparison to LP. Susanna Varho, Paavo Alku |
ISCAS | 2 |
| 1999 | On the linearity of the relationship between the sound pressure level and the negative peak amplitude of the differentiated glottal flow in vowel production
Paavo Alku, Juha Vintturi, Erkki Vilkman |
Speech Commun. | 1 |
| 1998 | A new linear predictive method for compression of speech signals
Paavo Alku, Susanna Varho |
ICSLP | 1 |
| 1998 | Analyzing the effect of secondary excitations of the vocal tract on vocal intensity in different loudness conditions
Paavo Alku, Juha Vintturi, Erkki Vilkman |
ICSLP | 1 |
| 1998 | Estimation of amplitude features of the glottal flow by inverse filtering speech pressure signals
Paavo Alku, Erkki Vilkman, Anne-Maria Laukkanen |
Speech Commun. | 1 |
| 1998 | Separated Linear Prediction - A new all-pole modelling technique for speech analysis
Susanna Varho, Paavo Alku |
Speech Commun. | 2 |
| 1997 | Parabolic spectral parameter - A new method for quantification of the glottal flow
Paavo Alku, Helmer Strik, Erkki Vilkman |
Speech Commun. | 1 |
| 1996 | A frequency domain method for parametrization of the voice source
Paavo Alku, Erkki Vilkman |
ICSLP | 1 |
| 1996 | Amplitude domain quotient for characterization of the glottal volume velocity waveform estimated by inverse filtering
Paavo Alku, Erkki Vilkman |
Speech Commun. | 1 |
| 1994 | Estimation of the glottal pulseform based on discrete all-pole modeling
Paavo Alku, Erkki Vilkman |
ICSLP | 1 |
| 1992 | An automatic method to estimate the time-based parameters of the glottal pulseformabstractA method of automatically extracting the time-based parameters (the fundamental period, the open quotient, the speed quotient, and the closing quotient) of the glottal waveform is presented. The only input that is needed in the analysis is the acoustical speech signal. In the first stage of the method an estimate for the glottal excitation is computed using iterative adaptive inverse filtering (IAIF). The extraction of the time-based parameters is then computed by using both the obtained glottal flow and its derivative. The method was tested with synthetic and natural speech of three different phonation types. The results show that the method gives good estimates for the time-based parameters of the glottal flow except for the analysis of speech that is produced using high fundamental frequency and pressed phonation.> Paavo Alku |
ICASSP | 1 |
| 1992 | Inverse filtering of the glottal waveform using the Itakura-saito distortion measure
Paavo Alku |
ICSLP | 1 |
| 1992 | Glottal wave analysis with Pitch Synchronous Iterative Adaptive Inverse Filtering
Paavo Alku |
Speech Commun. | 1 |
| 1991 | Glottal wave analysis with pitch synchronous iterative adaptive inverse filtering
Paavo Alku |
EUROSPEECH | 1 |
| 1990 | Glottal-LPC based coding of telephone band vowels with simple all-pole excitation
Paavo Alku |
ICSLP | 1 |
| 1990 | A comparison of egg and a new automatic inverse filtering method in phonation change from breathy to normal
Paavo Alku, Erkki Vilkman, Unto K. Laine |
ICSLP | 1 |
| 1989 | A new glottal LPC method of low complexity for speech analysis and coding
Paavo Alku, Unto K. Laine |
EUROSPEECH | 1 |
| 1989 | Speech processing in the object-oriented DSP environment quicksig
Matti Karjalainen, Toomas Altosaar, Paavo Alku, Lauri Lehtinen, Seppo Helle |
EUROSPEECH | 3 |
| 1988 | QuickSig-an object-oriented signal processing environmentabstractAn object-oriented DSP (digital signal-processing) environment called QuickSig is described that is based on recent developments in object-oriented programming (New Flavors on Symbolics Lisp machines). The design philosophy of QuickSig has been to extend the Lisp language by a layer of general DSP constructs, abstracts, and structures like signals, filters, windows, graphical presentations, and related signal-processing operations. QuickSig is a fast prototyping system for algorithmic development. It is extendable to include new ways of modeling signals and signal processing, both numerical and symbolic. The main features of the present system and some features that are under development are reported.> Matti Karjalainen, Toomas Altosaar, Paavo Alku |
ICASSP | 3 |