Jón Guðnason

dblp:14/970 · also Jon Gudnason · DBLP profile ↗
← Back
43ranked-venue papers
8as first author
13since 2021 · last 2027
0000-0001-6560-5543ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 8 first-author · 7 since 2021
YearPublicationVenuePosition
2027 Dual contrastive learning with clinically guided multimodal fusion for speech-based severe sleep apnea risk screening
Peizheng Wang, Arnab Majumdar, Wen-Te Liu, Jiunn-Horng Kang, Jón Guðnason, Ying-Ying Chen, I-Jung Liu, Kang-Yun Lee, Tzu-Tao Chen, Hsin-Chien Lee, Yi-Chih Lin, Yi-Chun Kuan, Yu-Hsiang Chang, Cheng-Yu Tsai
Expert Syst. Appl.5
2026 FPSC: A Sustainable Pipeline for Building a Faroese Parliamentary Speech Corpus
Dávid í Lág, Barbara Scalvini, Carlos Daniel Hernandez Mena, Jón Guðnason
LREC4
2025 Physiologically-Informed Feature Analysis of Acquired Speech Disorders for Stroke Assessment
Giulia Sanguedolce, Jón Guðnason, Dragos-Cristian Gruia, Emilie D'Olne, Fatemeh Geranmayeh, Patrick A. Naylor
INTERSPEECH2
2024 SamróMur MilljóN: An ASR Corpus of One Million Verified Read Prompts in Icelandic
abstract
The platform samromur.is, or “Samrómur” for short, is a crowdsourcing web application built on Mozilla’s Common Voice, designed to accumulate speech data for the advancement of language technologies in Icelandic. Over the years, Samrómur has proven to be remarkably successful in amassing a significant number of high-quality audio clips from thousands of users. However, the challenge of manually verifying the entirety of the collected data has hindered its effective exploitation, especially in the realm of Automatic Speech Recognition (ASR), its original purpose. In this paper, we introduce the “Samrómur Milljón” corpus, an ASR dataset comprising one million audio clips from Samrómur. These clips have been automatically verified using state-of-the-art speech recognition systems such as NeMo, Wav2Vec2, and Whisper. Additionally, we present the ASR results obtained from creating acoustic models based on Samrómur Milljón. These results demonstrate significant promise when compared to other acoustic models trained with a similar volume of Icelandic data from different sources.
Carlos Daniel Hernandez Mena, Þorsteinn Daði Gunnarsson, Jón Guðnason
LREC/COLING3
2023 Relative Dynamic Time Warping Comparison for Pronunciation Errors
abstract
We propose using a dynamic time warping (DTW) difference-to-sum ratio to classify speech as either matching or diverging from a linguistic standard. This measure effectively recognises non-native Norwegian speakers’ mispronunciations in words and phonetic segments. The contributions of the approach include (a) using DTW comparisons from two parallel sources, which represent the linguistic standard (e.g. native speakers) and an error model, to identify pronunciation errors; (b) recognising a heterogeneous standard, in this case the highly variable range of Norwegian dialects, instead of only a specified canonical phoneme sequence; (c) handling unanticipated pronunciation variants, both acceptable and unacceptable, beyond those seen in the standard and error models; and (d) requiring minimal training or pretraining data in the target language, which helps to make pronunciation error detection accessible even in lowresource languages without functional ASR.
Caitlin Richter, Jón Guðnason
ICASSP2
2023 Fine-Grained Emotional Control of Text-to-Speech: Learning to Rank Inter- and Intra-Class Emotion Intensities
abstract
State-of-the-art Text-To-Speech (TTS) models are capable of producing high-quality speech. The generated speech, however, is usually neutral in emotional expression, whereas very often one would want fine-grained emotional control of words or phonemes. Although still challenging, the first TTS models have been recently proposed that are able to control voice by manually assigning emotion intensity. Unfortunately, due to the neglect of intra-class distance, the intensity differences are often unrecognizable. In this paper, we propose a fine-grained controllable emotional TTS, that considers both inter- and intra-class distances and be able to synthesize speech with recognizable intensity difference. Our subjective and objective experiments demonstrate that our model exceeds two state-of-the-art controllable TTS models for controllability, emotion expressiveness and naturalness.
Jón Guðnason, Damian Borth
ICASSP2
2023 Epoch-Based Spectrum Estimation for Speech
Jón Guðnason, Guolin Fang, Mike Brookes
INTERSPEECH1
2023 Orthography-based Pronunciation Scoring for Better CAPT Feedback
Caitlin Richter, Ragnar Pálsson, Luke O'Brien, Kolbrún Friðriksdóttir, Branislav Bédi, Eydís Huld Magnúsdóttir, Jón Guðnason
INTERSPEECH7
2023 Learning Emotional Representations from Imbalanced Speech Data for Speech Emotion Recognition and Emotional Text-to-Speech
abstract
Effective speech emotional representations play a key role in Speech Emotion Recognition (SER) and Emotional Text-To-Speech (TTS) tasks. However, emotional speech samples are more difficult and expensive to acquire compared with Neutral style speech, which causes one issue that most related works unfortunately neglect: imbalanced datasets. Models might overfit to the majority Neutral class and fail to produce robust and effective emotional representations. In this paper, we propose an Emotion Extractor to address this issue. We use augmentation approaches to train the model and enable it to extract effective and generalizable emotional representations from imbalanced datasets. Our empirical results show that (1) for the SER task, the proposed Emotion Extractor surpasses the state-of-the-art baseline on three imbalanced datasets; (2) the produced representations from our Emotion Extractor benefit the TTS model, and enable it to synthesize more expressive speech.
Jón Guðnason, Damian Borth
INTERSPEECH2
2022 National Language Technology Platform (NLTP): overall view
abstract
The work in progress on the CEF Action National Language Technology Platform (NLTP) is presented. The Action aims at combining the most advanced Language Technology (LT) tools and solutions in a new state-of-the-art, Artificial Intelli- gence (AI) driven, National Language Technology Platform (NLTP).
Arturs Vasilevskis, Janis Ziedins, Marko Tadic, Zeljka Motika, Mark Fishel, Hrafn Loftsson, Jón Guðnason, Claudia Borg, Keith Cortis, Judie Attard, Donatienne Spiteri
EAMT7
2022 Generative Data Augmentation Guided by Triplet Loss for Speech Emotion Recognition
abstract
Speech Emotion Recognition (SER) is crucial for humancomputer interaction but still remains a challenging problem because of two major obstacles: data scarcity and imbalance.Many datasets for SER are substantially imbalanced, where data utterances of one class (most often Neutral) are much more frequent than those of other classes.Furthermore, only a few data resources are available for many existing spoken languages.To address these problems, we exploit a GAN-based augmentation model guided by a triplet network, to improve SER performance given imbalanced and insufficient training data.We conduct experiments and demonstrate: 1) With a highly imbalanced dataset, our augmentation strategy significantly improves the SER performance (+8% recall score compared with the baseline).2) Moreover, in a cross-lingual benchmark, where we train a model with enough source language utterances but very few target language utterances (around 50 in our experiments), our augmentation strategy brings benefits for the SER performance of all three target languages.
Hamed Hemati, Jón Guðnason, Damian Borth
INTERSPEECH3
2022 Samrómur Children: An Icelandic Speech Corpus
abstract
Samrómur Children is an Icelandic speech corpus intended for the field of automatic speech recognition. It contains 131 hours of read speech from Icelandic children aged between 4 to 17 years. The test portion was meticulously selected to cover a wide range of ages as possible; we aimed to have exactly the same amount of data per age range. The speech was collected with the crowd-sourcing platform Samrómur.is, which is inspired on the “Mozilla’s Common Voice Project”. The corpus was developed within the framework of the “Language Technology Programme for Icelandic 2019 − 2023”; the goal of the project is to make Icelandic available in language-technology applications. Samrómur Children is the first corpus in Icelandic with children’s voices for public use under a Creative Commons license. Additionally, we present baseline experiments and results using Kaldi.
Carlos Daniel Hernandez Mena, David Erik Mollberg, Michal Borsky, Jón Guðnason
LREC4
2021 Acoustic Measure of Vocal Strain Based on Glottal Airflow Periodicity
abstract
In the clinical practice of dysphonia, the effects of treatment are traditionally monitored by a sequence of auditory-perceptual assessments aimed at measuring vocal quality for the patient. Alternatively, acoustic measurement of vocal quality promises to automate perceptual assessments while keeping the assessments accurate and non-invasive. However, acoustic measures of vocal quality need to be further developed in both functional and technical terms. On the one hand, many of them are susceptible to non-dysphonic perturbations from articulatory movements in continuous speech, while on the other, their accuracy in approximating the generally nonlinear mapping from observation to vocal quality is limited by their use of a linear model. This paper presents an acoustic measure of vocal strain, a specific vocal quality that typically co-occurs with the development of vocal-fold nodules in vocal hyper-function. Vocal strain merits acoustic measurement more than other vocal qualities because its perceptual assessment typically exhibits a lower intra- and inter-rater reliability than the assessment of other vocal qualities. Based on an assumed correlation between vocal strain and the degree of periodicity in vocal-fold vibrations, this paper presents an acoustic measure in which a nonlinear regression model is used to predict the strain from some periodicity features extracted from a glottal airflow estimate. When tested on a set of listener-rated utterances composed mostly of continuous speech, the proposed glottal measure outperformed a direct-analysis measure in producing strain assessments which are consistent with perceptual ratings.
Yu-Ren Chien, Jón Guðnason
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Samrómur: Crowd-sourcing Data Collection for Icelandic Speech Recognition
abstract
This contribution describes an ongoing project of speech data collection, using the web application Samrómur which is built upon Common Voice, Mozilla Foundation’s web platform for open-source voice collection. The goal of the project is to build a large-scale speech corpus for Automatic Speech Recognition (ASR) for Icelandic. Upon completion, Samrómur will be the largest open speech corpus for Icelandic collected from the public domain. We discuss the methods used for the crowd-sourcing effort and show the importance of marketing and good media coverage when launching a crowd-sourcing campaign. Preliminary results exceed our expectations, and in one month we collected data that we had estimated would take three months to obtain. Furthermore, our initial dataset of around 45 thousand utterances has good demographic coverage, is gender-balanced and with proper age distribution. We also report on the task of validating the recordings, which we have not promoted, but have had numerous hours invested by volunteers.
David Erik Mollberg, Ólafur Helgi Jónsson, Sunneva THorsteinsdóttir, Steinþór Steingrímsson, Eydís Huld Magnúsdóttir, Jón Guðnason
LREC6
2020 Language Technology Programme for Icelandic 2019-2023
abstract
In this paper, we describe a new national language technology programme for Icelandic. The programme, which spans a period of five years, aims at making Icelandic usable in communication and interactions in the digital world, by developing accessible, open-source language resources and software. The research and development work within the programme is carried out by a consortium of universities, institutions, and private companies, with a strong emphasis on cooperation between academia and industries. Five core projects will be the main content of the programme: language resources, speech recognition, speech synthesis, machine translation, and spell and grammar checking. We also describe other national language technology programmes and give an overview over the history of language technology in Iceland.
Anna Björk Nikulásdóttir, Jón Guðnason, Anton Karl Ingason, Hrafn Loftsson, Eiríkur Rögnvaldsson, Einar Freyr Sigurðsson, Steinþór Steingrímsson
LREC2
2019 F0 Variability Measures Based on Glottal Closure Instants
Yu-Ren Chien, Michal Borsky, Jón Guðnason
INTERSPEECH3
2019 The Althingi ASR System
Inga Rún Helgadóttir, Anna Björk Nikulásdóttir, Michal Borsky, Judy Y. Fong, Róbert Kjaran, Jón Guðnason
INTERSPEECH6
2019 Bootstrapping a Text Normalization System for an Inflected Language. Numbers as a Test Case
Anna Björk Nikulásdóttir, Jón Guðnason
INTERSPEECH2
2019 Lattice Re-Scoring During Manual Editing for Automatic Error Correction of ASR Transcripts
Anna V. Rúnarsdóttir, Inga Rún Helgadóttir, Jón Guðnason
INTERSPEECH3
2018 Open ASR for Icelandic: Resources and a Baseline System
Anna Björk Nikulásdóttir, Inga Rún Helgadóttir, Matthias Petursson, Jón Guðnason
LREC4
2018 Risamálheild: A Very Large Icelandic Text Corpus
Steinþór Steingrímsson, Sigrún Helgadóttir, Eiríkur Rögnvaldsson, Starkaður Barkarson, Jón Guðnason
LREC5
2018 An Icelandic Pronunciation Dictionary for TTS
abstract
This paper describes an Icelandic pronunciation dictionary for speech applications and its processing for use in a text-to-speech system for Icelandic. Cleaning and correction procedures were implemented to create a consistent training set for grapheme-to-phoneme conversion modeling, needed for the automatic extension of the dictionary. Experiments with the original version of the dictionary and the cleaned version described in this paper as training sets for a joint sequence g2p algorithm show a clear benefit of using clean data for training, both in terms of PER and in terms of categories of errors made by the g2p algorithm. The results of the dictionary processing where also used to create an initial version of an open source database for Icelandic speech applications.
Anna Björk Nikulásdóttir, Jón Guðnason, Eiríkur Rögnvaldsson
SLT2
2017 Classification of voice modes using neck-surface accelerometer data
abstract
This study analyzes signals recorded using a neck-surface accelerometer from subjects producing speech with different voice modes. The purpose is to explore if the recorded waveforms can capture the glottal vibratory patterns which can be related to the movement of the vocal folds and thus voice quality. The accelerometer waveforms do not contain the supraglottal resonances, and these characteristics make the proposed method suitable for real-life voice quality assessment and monitoring as it does not breach patient privacy. The experiments with a Gaussian mexture model classifier demonstrate that different voice qualities produce distinctly different accelerometer waveforms. The system achieved 80.2% and 89.5% for frame- and utterance-level accuracy, respectively, for classifying among modal, breathy, pressed, and rough voice modes using a speaker-dependent classifier. Finally, the article presents characteristic waveforms for each modality and discusses their attributes.
Michal Borsky, Marion Cocude, Daryush D. Mehta, Matías Zanartu, Jón Guðnason
ICASSP5
2017 Objective Severity Assessment from Disordered Voice Using Estimated Glottal Airflow
Yu-Ren Chien, Michal Borsky, Jón Guðnason
INTERSPEECH3
2017 Building ASR Corpora Using Eyra
Jón Guðnason, Matthias Petursson, Róbert Kjaran, Simon Klüpfel, Anna Björk Nikulásdóttir
INTERSPEECH1
2017 Building an ASR Corpus Using Althingi's Parliamentary Speeches
Inga Rún Helgadóttir, Róbert Kjaran, Anna Björk Nikulásdóttir, Jón Guðnason
INTERSPEECH4
2017 Modal and Nonmodal Voice Quality Classification Using Acoustic and Electroglottographic Features
abstract
The goal of this study was to investigate the performance of different feature types for voice quality classification using multiple classifiers. The study compared the COVAREP feature set; which included glottal source features, frequency warped cepstrum and harmonic model features; against the mel-frequency cepstral coefficients (MFCCs) computed from the acoustic voice signal, acoustic-based glottal inverse filtered (GIF) waveform, and electroglottographic (EGG) waveform. Our hypothesis was that MFCCs can capture the perceived voice quality from either of these three voice signals. Experiments were carried out on recordings from 28 participants with normal vocal status who were prompted to sustain vowels with modal and non-modal voice qualities. Recordings were rated by an expert listener using the Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V), and the ratings were transformed into a dichotomous label (presence or absence) for the prompted voice qualities of modal voice, breathiness, strain, and roughness. The classification was done using support vector machines, random forests, deep neural networks and Gaussian mixture model classifiers, which were built as speaker independent using a leave-one-speaker-out strategy. The best classification accuracy of 79.97% was achieved for the full COVAREP set. The harmonic model features were the best performing subset, with 78.47% accuracy, and the static+dynamic MFCCs scored at 74.52%. A closer analysis showed that MFCC and dynamic MFCC features were able to classify modal, breathy, and strained voice quality dimensions from the acoustic and GIF waveforms. Reduced classification performance was exhibited by the EGG waveform.
Michal Borsky, Daryush D. Mehta, Jarrad H. Van Stan, Jón Guðnason
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Evaluation of Glottal Inverse Filtering Algorithms Using a Physiologically Based Articulatory Speech Synthesizer
abstract
Glottal inverse filtering aims to estimate the glottal airflow signal from a speech signal for applications such as speaker recognition and clinical voice assessment. Nonetheless, evaluation of inverse filtering algorithms has been challenging due to the practical difficulties of directly measuring glottal airflow. Apart from this, it is acknowledged that the performance of many methods degrade in voice conditions that are of great interest, such as breathiness, high pitch, soft voice, and running speech. This paper presents a comprehensive, objective, and comparative evaluation of state-of-the-art inverse filtering algorithms that takes advantage of speech and glottal airflow signals generated by a physiological speech synthesizer. The synthesizer provides a physics-based simulation of the voice production process and thus an adequate test bed for revealing the temporal and spectral performance characteristics of each algorithm. Included in the synthetic data are continuous speech utterances and sustained vowels, which are produced with multiple voice qualities (pressed, slightly pressed, modal, slightly breathy, and breathy), fundamental frequencies, and subglottal pressures to simulate the natural variations in real speech. In evaluating the accuracy of a glottal flow estimate, multiple error measures are used, including an error in the estimated signal that measures overall waveform deviation, as well as an error in each of several clinically relevant features extracted from the glottal flow estimate. Waveform errors calculated from glottal flow estimation experiments exhibited mean values around 30% for sustained vowels, and around 40% for continuous speech, of the amplitude of true glottal flow derivative. Closed-phase approaches showed remarkable stability across different voice qualities and subglottal pressures. The algorithms of choice, as suggested by significance tests, are closed-phase covariance analysis for the analysis of sustained vowels, and sparse linear prediction for the analysis of continuous speech. Results of data subset analysis suggest that analysis of close rounded vowels is an additional challenge in glottal flow estimation.
Yu-Ren Chien, Daryush D. Mehta, Jón Guðnason, Matías Zanartu, Thomas F. Quatieri
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Classification of Voice Modality Using Electroglottogram Waveforms
Michal Borsky, Daryush D. Mehta, Julius P. Gudjohnsen, Jón Guðnason
INTERSPEECH4
2015 Evaluation of speech inverse filtering techniques using a physiologically based synthesizer
abstract
Glottal inverse filtering methods are designed to derive a glottal flow waveform from a speech signal. In this paper, we evaluate and compare such methods using a speech synthesizer that simulates voice production in a physiologically-based manner that includes complexities such as nonlinear source-tract coupling. Five inverse filtering techniques are evaluated on 90 synthesized speech waveforms generated by setting six vowel configurations, three glottal models, and five fundamental frequencies. Using normalized mean square error as the primary performance metric of the estimated glottal flow derivative, results show that the accuracy of all methods depends on the configuration of the vocal tract, glottis and the fundamental frequency. Averaged over these conditions, the closed phase covariance and one weighted covariance algorithm yield lower error rates (0.41 ± 0.2) than iterative and adaptive inverse filtering (0.49 ± 0.1) and complex cepstrum decomposition (0.76 ± 0.1).
Jón Guðnason, Daryush D. Mehta, Thomas F. Quatieri
ICASSP1
2015 Detection of cardiovascular reactivity in speech
Laurens van der Werff, Jón Guðnason, Kamilla R. Johannsdottir
INTERSPEECH2
2014 Closed phase estimation for inverse filtering the oral airflow waveform
abstract
Glottal closed phase estimation during speech production is critical to inverse filtering and, although addressed for radiated acoustic pressure analysis, must be better understood for the analysis of the oral airflow volume velocity signal that provides important properties of healthy and disordered voices. This paper compares the estimation of the closed phase from the acoustic speech signal and the oral airflow waveform recorded using a pneumotachograph mask. Results are presented for ten adult speakers with normal voices who sustained a set of vowels at a comfortable pitch and loudness. With electroglottography as reference, the identification rate and accuracy of glottal closure instants for the oral airflow are 96.8 % and 0.28 ms, whereas these metrics are 99.4 % and 0.10 ms for the acoustic signal. We conclude that glottal closure detection is adequate for close phase inverse filtering but that improvements to detection of glottal opening instants on the oral airflow signal are warranted.
Jón Guðnason, Daryush D. Mehta, Thomas F. Quatieri
ICASSP1
2012 Data-driven voice source waveform analysis and synthesis
Jón Guðnason, Mark R. P. Thomas, Daniel P. W. Ellis, Patrick A. Naylor
Speech Commun.1
2012 Detection of Glottal Closure Instants From Speech Signals: A Quantitative Review
abstract
The pseudo-periodicity of voiced speech can be exploited in several speech processing applications. This requires however that the precise locations of the glottal closure instants (GCIs) are available. The focus of this paper is the evaluation of automatic methods for the detection of GCIs directly from the speech waveform. Five state-of-the-art GCI detection algorithms are compared using six different databases with contemporaneous electroglottographic recordings as ground truth, and containing many hours of speech by multiple speakers. The five techniques compared are the Hilbert Envelope-based detection (HE), the Zero Frequency Resonator-based method (ZFR), the Dynamic Programming Phase Slope Algorithm (DYPSA), the Speech Event Detection using the Residual Excitation And a Mean-based Signal (SEDREAMS) and the Yet Another GCI Algorithm (YAGA). The efficacy of these methods is first evaluated on clean speech, both in terms of reliabililty and accuracy. Their robustness to additive noise and to reverberation is also assessed. A further contribution of the paper is the evaluation of their performance on a concrete application of speech processing: the causal-anticausal decomposition of speech. It is shown that for clean speech, SEDREAMS and YAGA are the best performing techniques, both in terms of identification rate and accuracy. ZFR and SEDREAMS also show a superior robustness to additive noise and reverberation.
Thomas Drugman, Mark R. P. Thomas, Jón Guðnason, Patrick A. Naylor, Thierry Dutoit
IEEE Trans. Speech Audio Process.3
2012 Estimation of Glottal Closing and Opening Instants in Voiced Speech Using the YAGA Algorithm
abstract
Accurate estimation of glottal closing instants (GCIs) and opening instants (GOIs) is important for speech processing applications that benefit from glottal-synchronous processing including pitch tracking, prosodic speech modification, speech dereverberation, synthesis and study of pathological voice. We propose the Yet Another GCI/GOI Algorithm (YAGA) to detect GCIs from speech signals by employing multiscale analysis, the group delay function, andN-best dynamic programming. A novel GOI detector based upon the consistency of the candidates' closed quotients relative to the estimated GCIs is also presented. Particular attention is paid to the precise definition of the glottal closed phase, which we define as the analysis interval that produces minimum deviation from an all-pole model of the speech signal with closed-phase linear prediction (LP). A reference algorithm analyzing both electroglottograph (EGG) and speech signals is described for evaluation of the proposed speech-based algorithm. In addition to the development of a GCI/GOI detector, an important outcome of this work is in demonstrating that GOIs derived from the EGG signal are not necessarily well-suited to closed-phase LP analysis. Evaluation of YAGA against the APLAWD and SAM databases show that GCI identification rates of up to 99.3% can be achieved with an accuracy of 0.3 ms and GOI detection can be achieved equally reliably with an accuracy of 0.5 ms.
Mark R. P. Thomas, Jón Guðnason, Patrick A. Naylor
IEEE Trans. Speech Audio Process.2
2010 Voice source estimation for artificial bandwidth extension of telephone speech
abstract
Artificial bandwidth extension (ABWE) of speech signals aims to estimate wideband speech (50 Hz – 7 kHz) from narrowband signals (300 Hz – 3.4 kHz). Applying the source-filter model of speech, many existing algorithms estimate vocal tract filter parameters independently of the source signal. However, many current methods for extending the narrowband voice source signal are limited to straightforward signal processing techniques which are only effective for high-band estimation. This paper presents a method for ABWE that employs novel data-driven modelling and an existing spectral mirroring technique to estimate the wideband source signal in both the high and low extension bands. A state-of-the-art Hidden Markov Model-based estimator evaluates the temporal and spectral envelopes in the missing frequency bands, with which the ABWE speech signal is synthesized. Informal listening tests comparing two existing source estimation techniques and two permutations of the proposed approach show an improvement in the perceived bandwidth of speech signals, in particular towards low frequencies. Subjective tests on the same data show a preference for the proposed techniques over the existing methods under test.
Mark R. P. Thomas, Jón Guðnason, Patrick A. Naylor, Bernd Geiser, Peter Vary
ICASSP2
2009 Data-driven voice soruce waveform modelling
abstract
This paper presents a data-driven approach to the modelling of voice source waveforms. The voice source is a signal that is estimated by inverse-filtering speech signals with an estimate of the vocal tract filter. It is used in speech analysis, synthesis, recognition and coding to decompose a speech signal into its source and vocal tract filter components. Existing approaches parameterize the voice source signal with physically- or mathematically-motivated models. Though the models are well-defined, estimation of their parameters is not well understood and few are capable of reproducing the large variety of voice source waveforms. Here we present a data-driven approach to classify types of voice source waveforms based upon their mel frequency cepstrum coefficients with Gaussian mixture modelling. A set of ldquoprototyperdquo waveform classes is derived from a weighted average of voice source cycles from real data. An unknown speech signal is then decomposed into its prototype components and resynthesized. Results indicate that with sixteen voice source classes, low resynthesis errors can be achieved.
Mark R. P. Thomas, Jón Guðnason, Patrick A. Naylor
ICASSP2
2009 Voice source waveform analysis and synthesis using principal component analysis and Gaussian mixture modelling
abstract
The paper presents a voice source waveform modeling techniques based on principal component analysis (PCA) and Gaussian mixture modeling (GMM). The voice source is obtained by inverse-filtering speech with the estimated vocal tract filter. This decomposition is useful in speech analysis, synthesis, recognition and coding. Here, a data-driven approach is presented for signal decomposition and classification based on the principal components of the voice source. The principal components are analyzed and the 'prototype' voice source signals corresponding to the Gaussian mixture means are examined. We show how an unknown signal can be decomposed into its components and/or prototypes and resynthesized. We show how the techniques are suited for both low bitrate or high quality analysis/synthesis schemes.
Jón Guðnason, Mark R. P. Thomas, Patrick A. Naylor, Daniel P. W. Ellis
INTERSPEECH1
2008 Voice source cepstrum coefficients for speaker identification
abstract
We propose a novel feature set for speaker recognition that is based on the voice source signal. The feature extraction process uses closed-phase LPC analysis to estimate the vocal tract transfer function. The LPC spectrum envelope is converted to cepstrum coefficients which are used to derive the voice source features. Unlike approaches based on inverse-filtering, our procedure is robust to LPC analysis errors and low-frequency phase distortion. We have performed text-independent closed-set speaker identification experiments on the TIMIT and the YOHO databases using a standard Gaussian mixture model technique. Compared to using mel- frequency cepstrum coefficients, the misclassification rate for the TIMIT database reduced from 1.51% to 0.16% when combined with the proposed voice source features. For the YOHO database the mis- classification rate decreased from 13.79% to 10.07%. The new feature vector also compares favourably to other proposed voice source feature sets.
Jón Guðnason, Mike Brookes
ICASSP1
2007 Estimation of Glottal Closure Instants in Voiced Speech Using the DYPSA Algorithm
abstract
We present the Dynamic Programming Projected Phase-Slope Algorithm (DYPSA) for automatic estimation of glottal closure instants (GCIs) in voiced speech. Accurate estimation of GCIs is an important tool that can be applied to a wide range of speech processing tasks including speech analysis, synthesis and coding. DYPSA is automatic and operates using the speech signal alone without the need for an EGG signal. The algorithm employs the phase-slope function and a novel phase-slope projection technique for estimating GCI candidates from the speech signal. The most likely candidates are then selected using a dynamic programming technique to minimize a cost function that we define. We review and evaluate three existing methods of GCI estimation and compare the new DYPSA algorithm to them. Results are presented for the APLAWD and SAM databases for which 95.7% and 93.1% of GCIs are correctly identified
Patrick A. Naylor, Anastasis Kounoudes, Jón Guðnason, Mike Brookes
IEEE Trans. Speech Audio Process.3
2006 A quantitative assessment of group delay methods for identifying glottal closures in voiced speech
abstract
Measures based on the group delay of the LPC residual have been used by a number of authors to identify the time instants of glottal closure in voiced speech. In this paper, we discuss the theoretical properties of three such measures and we also present a new measure having useful properties. We give a quantitative assessment of each measure's ability to detect glottal closure instants evaluated using a speech database that includes a direct measurement of glottal activity from a Laryngograph/EGG signal. We find that when using a fixed-length analysis window, the best measures can detect the instant of glottal closure in 97% of larynx cycles with a standard deviation of 0.6 ms and that in 9% of these cycles an additional excitation instant is found that normally corresponds to glottal opening. We show that some improvement in detection rate may be obtained if the analysis window length is adapted to the speech pitch. If the measures are applied to the preemphasized speech instead of to the LPC residual, we find that the timing accuracy worsens but the detection rate improves slightly. We assess the computational cost of evaluating the measures and we present new recursive algorithms that give a substantial reduction in computation in all cases.
Mike Brookes, Patrick A. Naylor, Jón Guðnason
IEEE Trans. Speech Audio Process.3
2005 Automatic recognition of MSTAR targets using radar shadow and superresolution features
abstract
Automatic target recognition from high range resolution radar profiles remains an important and challenging problem. In this paper, we present a novel feature set for this task that combines a representation of the target's radar shadow with a noise-robust superresolution characterisation of the target scattering centres derived from the MUSIC algorithm. Using an HMM to represent aspect dependence, we demonstrate that the inclusion of the shadow features results in a significant improvement in recognition performance. We evaluate our proposed feature set on a closed-set identification task using targets from the MSTAR database and show that it results in lower recognition error rates than previously published methods using the same data.
Jingjing Cui 0005, Jón Guðnason, Mike Brookes
ICASSP (5)2
2002 Distribution based classification using Gaussian Mixture Models
abstract
A central task in classification is a measure of similarity between a dataset and a class that is characterised by a probability density function. The Bhattacharyya distance and the Kullback-Liebler divergence measure have been successful in comparing two multivariate normal density functions but their use is impracticable when the data is modelled using complex distributions such as Gaussian Mixture Models. The similarity is computed by combining the Bhattacharyya distances between corresponding mixtures in the reference and the test data model. In this paper we compare the performance of the Likelihood Ratio Test to a novel technique that defines a similarity measure between data and reference models having Gaussian Mixture probability density functions. When fitting a Gaussian Mixture Model to the test dataset our procedure ensures a one to one correspondence between the mixtures of the dataset and those of the reference model. This procedure has been tested using experiments, with both synthetic data and a Speaker Verification evaluation database. The performance was assessed using Detection Error Trade-off curves and demonstrates that the new measure performs significantly better than Likelihood Ratio Test.
Jón Guðnason, Mike Brookes
ICASSP1