Eduardo Lleida

dblp:14/4997 · also Eduardo Lleida-Solano · DBLP profile ↗
← Back
110ranked-venue papers
12as first author
8since 2021 · last 2024
0000-0001-9137-4013ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 96 · 10 first-author · 7 since 2021Artificial intelligence and machine learning · 84 · 8 first-author · 6 since 2021
YearPublicationVenuePosition
2024 Predefined Prototypes for Intra-Class Separation and Disentanglement
abstract
International audience
Antonio Almudévar, Théo Mariotte, Alfonso Ortega Giménez, Marie Tahon, Luis Vicente, Antonio Miguel, Eduardo Lleida
INTERSPEECH7
2023 Variational Classifier for Unsupervised Anomalous Sound Detection under Domain Generalization
Antonio Almudévar, Alfonso Ortega Giménez, Luis Vicente, Antonio Miguel, Eduardo Lleida
INTERSPEECH5
2023 On the Use of High Frequency Information for Voice Pathology Classification
Dayana Ribas González, Eduardo Lleida
INTERSPEECH3
2022 aDCF Loss Function for Deep Metric Learning in End-to-End Text-Dependent Speaker Verification Systems
abstract
Metric learning approaches have widely expanded to the training of Speaker Verification (SV) systems based on Deep Neural Networks (DNNs), by using a loss function more consistent with the evaluation process than the traditional identification losses. However, these methods do not consider the performance measure and can involve high computational cost, for example, the need for a careful pair or triplet data selection. This paper proposes the approximated Detection Cost Function (aDCF) loss, which is a loss function based on the measure of the decision errors in SV systems, namely the False Rejection Rate (FRR) and the False Acceptance Rate (FAR). With aDCF loss as the training objective function, the end-to-end system learns how to minimize decision errors. Furthermore, we replace the typical linear layer as the last layer of DNN by a cosine distance layer, which reduces the difference between the metric in the training process and the metric during evaluation. aDCF loss function was evaluated in RSR2015-Part I and RSR2015-Part II datasets for text-dependent speaker verification. The system trained with aDCF loss outperforms all the state-of-the-art functions employed in this paper in both parts of the database.
Victoria Mingote, Antonio Miguel, Dayana Ribas González, Alfonso Ortega Giménez, Eduardo Lleida
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 Memory Layers with Multi-Head Attention Mechanisms for Text-Dependent Speaker Verification
abstract
In this paper, we explore an approach based on memory layers and multi-head attention mechanisms to improve in an efficient way the performance of text-dependent speaker verification (SV) systems. The most extended SV systems based on Deep Neural Networks (DNN) extract the embedding of the utterance from the average pooling of the temporal dimension after processing. Unlike previous works, we can exploit the phonetic knowledge needed for text-dependent SV systems by combining the temporal attention of multiple parallel heads with the phonetic embeddings extracted from a phonetic classification network, which helps to guide to the attention mechanism with the role of the positional embedding. The addition of a memory layer to a text-dependent SV system was tested on the RSR2015-part II and DeepMine-part I databases, where, in both cases outperformed the baseline result and the reference system based on the same transformer network without the memory layer.
Victoria Mingote, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
ICASSP4
2021 Unsupervised Representation Learning for Speech Activity Detection in the Fearless Steps Challenge 2021
abstract
International audience
Pablo Gimeno, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
Interspeech4
2021 Log-Likelihood-Ratio Cost Function as Objective Loss for Speaker Verification Systems
abstract
International audience
Victoria Mingote, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
Interspeech4
2021 Generalizing AUC Optimization to Multiclass Classification for Audio Segmentation With Limited Training Data
abstract
Area under the ROC curve (AUC) optimisation techniques developed for neural networks have recently demonstrated their capabilities in different audio and speech related tasks. However, due to its intrinsic nature, AUC optimisation has focused only on binary tasks so far. In this paper, we introduce an extension to the AUC optimisation framework so that it can be easily applied to an arbitrary number of classes, aiming to overcome the issues derived from training data limitations in deep learning solutions. Building upon the multiclass definitions of the AUC metric found in the literature, we define two new training objectives using a one-versus-one and a one-versus-rest approach. In order to demonstrate its potential, we apply them in an audio segmentation task with limited training data that aims to differentiate 3 classes: foreground music, background music and no music. Experimental results show that our proposal can improve the performance of audio segmentation systems significantly compared to traditional training criteria such as cross entropy.
Pablo Gimeno, Victoria Mingote, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
IEEE Signal Process. Lett.5
2020 Knowledge Distillation and Random Erasing Data Augmentation for Text-Dependent Speaker Verification
abstract
This paper explores the Knowledge Distillation (KD) approach and a data augmentation technique to improve the generalization ability and robustness of text-dependent speaker verification (SV) systems. The KD method consists of two neural networks, known as Teacher and Student, where the student is trained to replicate the predictions from the teacher, so it learns their variability during the training process. To provide robustness to the distillation process, we apply Random Erasing (RE), a data augmentation technique which was created to improve the generalization ability of the neural networks. We have developed two alternatives of the combination of KD and RE, which, produce a more robust system with better performance, since the student network can learn from teacher predictions of data not existing in the original dataset. All alternatives were tested on RSR2015-Part I database, where the proposed variants outperform reference system based on a single network using RE.
Victoria Mingote, Antonio Miguel, Dayana Ribas González, Alfonso Ortega Giménez, Eduardo Lleida
ICASSP5
2020 Partial AUC Optimisation Using Recurrent Neural Networks for Music Detection with Limited Training Data
Pablo Gimeno, Victoria Mingote, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
INTERSPEECH5
2020 Training Speaker Enrollment Models by Network Optimization
Victoria Mingote, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH4
2020 Shouted Speech Compensation for Speaker Verification Robust to Vocal Effort Conditions
abstract
The performance of speaker verification systems degrades when vocal effort conditions between enrollment and test (e.g., shouted vs. normal speech) are different. This is a potential situation in non-cooperative speaker verification tasks. In this paper, we present a study on different methods for linear compensation of embeddings making use of Gaussian mixture models to cluster shouted and normal speech domains. These compensation techniques are borrowed from the area of robustness for automatic speech recognition and, in this work, we apply them to compensate the mismatch between shouted and normal conditions in speaker verification. Before compensation, shouted condition is automatically detected by means of logistic regression. The process is computationally light and it is performed in the back-end of an x-vector system. Experimental results show that applying the proposed approach in the presence of vocal effort mismatch yields up to 13.8% equal error rate relative improvement with respect to a system that applies neither shouted speech detection nor compensation.
Santi Prieto, Alfonso Ortega Giménez, Iván López-Espejo, Eduardo Lleida
INTERSPEECH4
2020 Optimization of the area under the ROC curve using neural network supervectors for text-dependent speaker verification
Victoria Mingote, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
Comput. Speech Lang.4
2019 Speech Enhancement with Wide Residual Networks in Reverberant Environments
abstract
This paper proposes a speech enhancement method which exploits the high potential of residual connections in a Wide Residual Network architecture. This is supported on single dimensional convolutions computed alongside the time domain, which is a powerful approach to process contextually correlated representations through the temporal domain, such as speech feature sequences. We find the residual mechanism extremely useful for the enhancement task since the signal always has a linear shortcut and the non-linear path enhances it in several steps by adding or subtracting corrections. The enhancement capability of the proposal is assessed by objective quality metrics evaluated with simulated and real samples of reverberated speech signals. Results show that the proposal outperforms the state-of-the-art method called WPE, which is known to effectively reduce reverberation and greatly enhance the signal. The proposed model, trained with artificial synthesized reverberation data, was able to generalize to real room impulse responses for a variety of conditions (e.g. different room sizes, $RT_{60}$, near & far field). Furthermore, it achieves accuracy for real speech with reverberation from two different datasets.
Jorge Llombart, Dayana Ribas González, Antonio Miguel, Luis Vicente, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH6
2019 Progressive Speech Enhancement with Residual Connections
abstract
This paper studies the Speech Enhancement based on Deep Neural Networks. The proposed architecture gradually follows the signal transformation during enhancement by means of a visualization probe at each network block. Alongside the process, the enhancement performance is visually inspected and evaluated in terms of regression cost. This progressive scheme is based on Residual Networks. During the process, we investigate a residual connection with a constant number of channels, including internal state between blocks, and adding progressive supervision. The insights provided by the interpretation of the network enhancement process leads us to design an improved architecture for the enhancement purpose. Following this strategy, we are able to obtain speech enhancement results beyond the state-of-the-art, achieving a favorable trade-off between dereverberation and the amount of spectral distortion.
Jorge Llombart, Dayana Ribas González, Antonio Miguel, Luis Vicente, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH6
2019 Language Recognition Using Triplet Neural Networks
Victoria Mingote, Diego Castán, Mitchell McLaren, Mahesh Kumar Nandwana, Alfonso Ortega Giménez, Eduardo Lleida, Antonio Miguel
INTERSPEECH6
2019 Optimization of False Acceptance/Rejection Rates and Decision Threshold for End-to-End Text-Dependent Speaker Verification Systems
Victoria Mingote, Antonio Miguel, Dayana Ribas González, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH5
2019 ViVoLAB Speaker Diarization System for the DIHARD 2019 Challenge
Ignacio Viñals, Pablo Gimeno, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
INTERSPEECH5
2019 Phonetically-Aware Embeddings, Wide Residual Networks with Time-Delay Neural Networks and Self Attention Models for the 2018 NIST Speaker Recognition Evaluation
Ignacio Viñals, Dayana Ribas González, Victoria Mingote, Jorge Llombart, Pablo Gimeno, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH8
2018 Estimation of the Number of Speakers with Variational Bayesian PLDA in the DIHARD Diarization Challenge
Ignacio Viñals, Pablo Gimeno, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
INTERSPEECH5
2018 Speaker and language recognition and characterization: Introduction to the CSL special issue
Eduardo Lleida, Luis Javier Rodríguez-Fuentes
Comput. Speech Lang.1
2017 Tied Hidden Factors in Neural Networks for End-to-End Speaker Recognition
abstract
In this paper we propose a method to model speaker and session variability and able to generate likelihood ratios using neural networks in an end-To-end phrase dependent speaker verification system. As in Joint Factor Analysis, the model uses tied hidden variables to model speaker and session variability and a MAP adaptation of some of the parameters of the model. In the training procedure our method jointly estimates the network parameters and the values of the speaker and channel hidden variables. This is done in a two-step backpropagation algorithm, first the network weights and factor loading matrices are updated and then the hidden variables, whose gradients are calculated by aggregating the corresponding speaker or session frames, since these hidden variables are tied. The last layer of the network is defined as a linear regression probabilistic model whose inputs are the previous layer outputs. This choice has the advantage that it produces likelihoods and additionally it can be adapted during the enrolment using MAP without the need of a gradient optimization. The decisions are made based on the ratio of the output likelihoods of two neural network models, speaker adapted and universal background model. The method was evaluated on the RSR2015 database.
Antonio Miguel, Jorge Llombart, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH4
2017 Domain Adaptation of PLDA Models in Broadcast Diarization by Means of Unsupervised Speaker Clustering
Ignacio Viñals, Alfonso Ortega Giménez, Jesús Villalba 0001, Antonio Miguel, Eduardo Lleida
INTERSPEECH5
2016 Analysis of speech quality measures for the task of estimating the reliability of speaker verification decisions
Jesús Villalba 0001, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
Speech Commun.4
2016 Bayesian Networks to Model the Variability of Speaker Verification Scores in Adverse Environments
abstract
State-of-the-art speaker recognition technology attains great performance in controlled conditions. However, when the speech segments suffer distortions like noise or reverberation performance can severely deteriorate, this fact motivated us to investigate how score distributions diverge from the ideal ones in degraded conditions. We propose a Bayesian network model that assumes that two scores exist: one observed and another one hidden. The observed score or noisy score is the one given by the speaker verification system. Meanwhile, the hidden score or clean score is the ideal score that we would obtain in a trial with high-quality speech. A set of quality measures helps to relate both scores. We applied this network to two tasks. The first one consists in rejecting unreliable trials, i.e., trials that we cannot assure whether they are target or nontarget. We prove that this method outperforms previous approaches, based on another type of Bayesian networks. The second task is to compute an improved likelihood ratio, dependent on the quality measures. This ratio improved calibration in noisy conditions.
Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
IEEE ACM Trans. Audio Speech Lang. Process.4
2015 Variational Bayesian PLDA for speaker diarization in the MGB challenge
abstract
This paper describes the ViVoLab speaker diarization system for the Multi-Genre Broadcast (MGB) Challenge at ASRU2015. The challenge data consisted of BBC TV programmes of different genres. Diarization followed a longitudinal setup, i.e., the speakers of the current episode had to be linked to the speakers in previous episodes of the same show. We propose a system based on the i-vector paradigm. After an initial segmentation step, we compute an i-vector per speech segment. Then, a generative model based on Bayesian PLDA clusters the speakers. In this model, the speaker labels are latent variables that we optimize by variational Bayes iterations. The number of speakers in each episode was decided by maximizing the variational lower bound. The system includes several phases of segment-merging and re-clustering. We re-compute i-vectors after each merging step, which reduces the i-vector uncertainty. This approach attained a DER around 30% in the development set.
Jesús Villalba 0001, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
ASRU4
2015 Spoofing detection with DNN and one-class SVM for the ASVspoof 2015 challenge
abstract
Speaker verification systems have achieved great performance in recent times. However, we usually measure performance on a ideal scenarios with naive impostors that do not modify their voices to impersonate the target speakers. The fact of impersonating a legitimate user is known as spoofing attack. Recent works show the vulnerability of current speaker verification technology to several types of attacks. Most of these works use non-public databases and different performance measures, which makes difficult to compare approaches. The spoofing challenge (ASVspoof 2015) tries to overcome this problem by proposing a common evaluation framework. This paper describes our submission to the challenge. We proposed to use spectral log-filter-bank and relative phase shift features as input to classifiers based on deep neural networks (DNN). The first of our classifiers used DNN posteriors to decide if the trial is spoof or non-spoof. The second used a bottleneck feature from the DNN as input to a one-class SVM. The one-class SVM models the distribution of legitimate speech, not needing spoofing data for training. We fused the score of the different classifiers to produce our final submission. Our system attained very competitive results with EER<0.05% in 9 out of 10 spoofing types.
Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH4
2014 Unscented transform for ivector-based noisy speaker recognition
abstract
Recently, a new version of the iVector modelling has been proposed for noise robust speaker recognition, where the nonlinear function that relates clean and noisy cepstral coefficients is approximated by a first order vector Taylor series (VTS). In this paper, it is proposed to substitute the first order VTS by an unscented transform, where unlike VTS, the nonlinear function is not applied over the clean model parameters directly, but over a set of sampled points. The resulting points in the transformed space are then used to calculate the model parameters. For very low signal-to-noise ratio improvements in equal error rate of about 7% for a clean backend and of 14.50% for a multistyle backend are obtained.
David Martínez González, Lukás Burget, Themos Stafylakis, Yun Lei, Patrick Kenny, Eduardo Lleida
ICASSP6
2014 Unsupervised adaptation of PLDA by using variational Bayes methods
abstract
State-of-the-art speaker recognition relays on models that need a large amount of training data. This models are successful in tasks like NIST SRE because there is sufficient data available. However, in real applications, we usually do not have so much data and, in many cases, the speaker labels are unknown. We present a method to adapt a PLDA model from a domain with a large amount of labeled data to another with unlabeled data. We describe a generative model that produces both sets of data where the unknown labels are modeled like latent variables. We used variational Bayes to estimate the hidden variables. We performed experiments adapting a model trained on Switchboard to NIST SRE without labels. The adapted model is evaluated on NIST SRE10. Compared to the non-adapted model, EER improved by 42% and 49% by adapting with 200 and with all the NIST speakers respectively.
Jesús Villalba 0001, Eduardo Lleida
ICASSP2
2014 Factor analysis with sampling methods for text dependent speaker recognition
Antonio Miguel, Jesús Villalba 0001, Alfonso Ortega Giménez, Eduardo Lleida, Carlos Vaquero
INTERSPEECH4
2014 Low bit rate compression methods of feature vectors for distributed speech recognition
José Enrique García Laínez, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
Speech Commun.4
2013 Segmentation-by-classification system based on factor analysis
abstract
This paper proposes a novel audio segmentation-by-classification system based on Factor Analysis (FA) with a channel compensation matrix for each class and scoring the fixed-length segments as the log-likelihood ratio between class/no-class. The scores are smoothed and the most probable sequence is computed with a Viterbi algorithm. The system described here is designed to segment and classify the audio files coming from broadcast programs into five different classes: speech (SP), speech with noise (SN), speech with music (SM), music (MU) or others (OT). This task was proposed in the Albayzin 2010 evaluation campaign. The system is compared with the winning system of the evaluation achieving lower error rate in SP and SN. These classes represent 3/4 of the total amount of the data. Therefore, the FA segmentation system gets a reduction in the average segmentation error rate.
Diego Castán, Alfonso Ortega Giménez, Jesús Villalba 0001, Antonio Miguel, Eduardo Lleida
ICASSP5
2013 Prosodic features and formant modeling for an ivector-based language recognition system
abstract
The prosody of a language is encoded in syllable length, loudness and pitch. These attributes make humans perceive rhythm, stress and intonation in speech. Depending on the language, these speech properties vary, making language classification possible. On the other hand, formants are the resonance frequencies of the vocal tract, depend heavily on the position adopted by the articulatory organs, and are especially useful to disambiguate vowel sounds. In this paper prosodic and formant information are combined to build a generative language identification system based on Gaussian models fed with iVectors. The system is evaluated on the NIST LRE09 database and the inclusion of formant information gives about 50% relative improvement for the 30 s task over a prosodic system without it. The fusion with a state-of-the-art acoustic system based on shifted delta cepstral coefficients (SDC) shows the complementarity of both approaches.
David Martínez González, Eduardo Lleida, Alfonso Ortega Giménez, Antonio Miguel
ICASSP2
2013 Handling i-vectors from different recording conditions using multi-channel simplified PLDA in speaker recognition
abstract
In this work, we address the problem of having i-vectors that have been produced in different channel conditions. Traditionally, this problem has been handled training the LDA covariance matrices pooling the data of all the conditions or averaging the covariance matrices of each condition in different ways. We present a PLDA variant that we call, multi-channel SPLDA, where the speaker space distribution is common to all i-vectors and the channel space distribution depends on the type of channel where the segment has been recorded. We test our approach on the telephone part of the NIST SRE10 extended condition where we added some additive noises to the test segments. We compare results of a SPLDA model trained only with clean data, SPLDA trained with pooled noisy and clean data and our MCSPLDA model.
Jesús Villalba 0001, Eduardo Lleida
ICASSP2
2013 Suprasegmental information modelling for autism disorder spectrum and specific language impairment classification
David Martínez González, Dayana Ribas González, Eduardo Lleida, Alfonso Ortega Giménez, Antonio Miguel
INTERSPEECH3
2013 A new Bayesian network to assess the reliability of speaker verification decisions
abstract
In some situations the quality of the signals involved in a speaker verification trial is not as good as needed to take a reliable decision. In this work, we present a new method based on Bayesian networks and quality measures to estimate if the trial decision is reliable. We present experiments on the NIST SRE2010 dataset degraded with additive noise. A system well calibrated for clean speech, produces a large actual DCF on the degraded dataset. We use our method to discard the unreliable trials and achieve a dramatic improvement of the cost values. We also prove that our method outperforms previously published approaches.
Jesús Villalba 0001, Eduardo Lleida, Alfonso Ortega Giménez, Antonio Miguel
INTERSPEECH2
2013 The I3a speaker recognition system for NIST SRE12: post-evaluation analysis
abstract
The I3A submission for the recent NIST 2012 speaker recognition evaluation (SRE) was based on the i-vector approach with a multi-channel PLDA classifier. This PLDA is modified so that, for each i-vector, the between-class covariance depends on the type of channel where the segment was recorded (telephone,interviews,clean, noisy, etc). In this paper, we present the description of our submission and a detailed post-evaluation analysis of the results. We analyze several factors affecting performance: enrollment data selection, classifier type, scoring technique, calibration, known and unknown non-targets, target speakers included or not in development, segment duration, noise level and noise type. Some of these factor are new in this evaluation. After post-evaluation, actual costs improve by 15– 43% depending on the common condition.
Jesús Villalba 0001, Eduardo Lleida, Alfonso Ortega Giménez, Antonio Miguel
INTERSPEECH2
2013 Handling recordings acquired simultaneously over multiple channels with PLDA
abstract
In some speaker recognition scenarios we find conversations recorded simultaneously over multiple channels.That is the case of the interviews in the NIST SRE dataset.To take advantage of that, we propose a modification of the PLDA model that considers two different inter-session variability terms.The first term is tied between all the recordings belonging to the same conversation whereas the second is not.Thus, the former mainly intends to capture the variability due to the phonetic content of the conversation while the latter tries to capture the channel variability.We test this approach on the NIST SRE12 core condition using multiple channels per interview to enroll the speakers.The proposed approach improves the minimum DCF by 26-29 % on telephone speech and by 1-8% on interviews compared to the standard PLDA (scored by the book).
Jesús Villalba 0001, Mireia Díez, Amparo Varona, Eduardo Lleida
INTERSPEECH4
2013 Quality Assessment for Speaker Diarization and Its Application in Speaker Characterization
abstract
There are many applications related to speaker characterization, specially in telephone environments, where large datasets are available but not directly useful since there are two speakers involved in every recording. Even with very accurate speaker diarization systems, we can expect to find some recordings with low diarization accuracy. The use of these recordings may reduce the accuracy of any speaker characterization technology. Therefore, it is highly desirable to detect those recordings where the speakers are correctly segmented, in order to discard or process manually the remaining ones before feeding them into the application. In this work we propose a set of confidence measures to assess the quality of a hypothetical diarization output, in order to detect those recordings that are correctly segmented. We show that these confidence measures enable us to retrieve most of the desired recordings from a given dataset, discarding those recordings that degrade the overall accuracy of an application that make use of speaker characterization technologies.
Carlos Vaquero, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
IEEE Trans. Speech Audio Process.4
2012 Beam-Search Formant Tracking Algorithm Based on Trajectory Functions for Continuous Speech
José Enrique García Laínez, Dayana Ribas González, Antonio Miguel, Eduardo Lleida, José Ramón Calvo de Lara
CIARP4
2012 The BLZ Submission to the NIST 2011 LRE: Data Collection, System Development and Performance
abstract
This paper describes the most relevant features of a collaborative multi-site submission to the NIST 2011 Language Recognition Evaluation (LRE), consisting of one primary and three contrastive systems, each fusing different combinations of 13 state-of-the-art (acoustic and phonotactic) language recognition subsystems.The collaboration focused on collecting and sharing training data for those target languages for which few development data were provided by NIST, and on defining a common development dataset to train backend and fusion parameters and select the best fusions.Official and post-key results are presented and compared, revealing that the greedy approach applied to select the best fusions provided suboptimal but very competitive performance.Several factors contributed to the high performance attained by BLZ systems, including the availability of training data for low resource target languages, the reliability of the development dataset (consisting only of data audited by NIST), the diversity of modeling approaches, features and datasets in the systems considered for fusion, and the effectiveness of the search for optimal fusions.
Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel, Alberto Abad, David Martínez González, Jesús Villalba 0001, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH10
2012 A prelingual tool for the education of altered voices
William Ricardo Rodríguez-Dueñas, Oscar Saz-Torralba, Eduardo Lleida
Speech Commun.3
2011 Multi-site heterogeneous system fusions for the Albayzin 2010 Language Recognition Evaluation
abstract
Best language recognition performance is commonly obtained by fusing the scores of several heterogeneous systems. Regardless the fusion approach, it is assumed that different systems may contribute complementary information, either because they are developed on different datasets, or because they use different features or different modeling approaches. Most authors apply fusion as a final resource for improving performance based on an existing set of systems. Though relative performance gains decrease as larger sets of systems are considered, best performance is usually attained by fusing all the available systems, which may lead to high computational costs. In this paper, we aim to discover which technologies combine the best through fusion and to analyse the factors (data, features, modeling methodologies, etc.) that may explain such a good performance. Results are presented and discussed for a number of systems provided by the participating sites and the organizing team of the Albayzin 2010 Language Recognition Evaluation. We hope the conclusions of this work help research groups make better decisions in developing language recognition technology.
Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel, David Martínez González, Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida, Alberto Abad, Oscar Koller, Isabel Trancoso, Paula Lopez-Otero, Laura Docío Fernández, Carmen García-Mateo, Rahim Saeidi, Mehdi Soufifar, Tomi Kinnunen, Torbjørn Svendsen, Pasi Fränti
ASRU10
2011 Intra-session variability compensation and a hypothesis generation and selection strategy for speaker segmentation
abstract
This paper addresses the problem of speaker segmentation in two-speaker telephone conversations, using an eigenvoice based factor analysis approach. We present a set of improvements in the speaker segmentation system. First, we study two methods to compensate for intra-session variability, that is the variability present in a speaker during a single session. Secondly we propose a method to generate segmentation hypotheses that combined with a given confidence measure, enables the selection of correct hypotheses improving the overall segmentation performance. The proposed improvements are evaluated on the NIST Speaker Recognition Evaluation 2008 summed channel test condition, obtaining 28% relative improvement in terms of speaker segmentation error.
Carlos Vaquero, Alfonso Ortega Giménez, Eduardo Lleida
ICASSP3
2011 Hierarchical Audio Segmentation with HMM and Factor Analysis in Broadcast News Domain
Diego Castán, Carlos Vaquero, Alfonso Ortega Giménez, David Martínez González, Jesús Villalba 0001, Eduardo Lleida
INTERSPEECH6
2011 I3A Language Recognition System for Albayzin 2010 LRE
David Martínez González, Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH5
2011 Partitioning of Two-Speaker Conversation Datasets
Carlos Vaquero, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH3
2011 Bayesian Networks for Discrete Observation Distributions in Speech Recognition
abstract
Traditionally, in speech recognition, the hidden Markov model state emission probability distributions are usually associated to continuous random variables, by using Gaussian mixtures. Thus, complex multimodal inter-feature dependencies are not accurately modeled by Gaussian models, since they are unimodal distributions and mixtures of Gaussians are needed in these complex cases, but this is done in a loose and inefficient way. Graphical models provide a precise and simple mechanism to model the dependencies among two or more variables. This paper proposes the use of discrete random variables as observations and graphical models to extract the internal dependence structure in the feature vectors. Therefore, speech features are quantized to a small number of levels, in order to obtain a tractable model. These quantized speech features provide a mechanism to increase the robustness against noise uncertainty. In addition, discrete random variables allow the learning of joint statistics of the observation densities. A method to estimate a graphical model with a constrained number of dependencies is shown in this paper, being a special kind of Bayesian network. Experimental results show that by using this modeling, better performance can be obtained compared to standard baseline systems.
Antonio Miguel, Alfonso Ortega Giménez, Luis Buera, Eduardo Lleida
IEEE Trans. Speech Audio Process.4
2010 Speaker Verification in Noisy Environment Using Missing Feature Approach
Dayana Ribas González, Jesús Villalba 0001, Eduardo Lleida, José Ramón Calvo de Lara
CIARP3
2010 Non-linear predictive vector quantization of feature vectors for distributed speech recognition
José Enrique García Laínez, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
INTERSPEECH4
2010 Confidence measures for speaker segmentation and their relation to speaker verification
Carlos Vaquero, Alfonso Ortega Giménez, Jesús Villalba 0001, Antonio Miguel, Eduardo Lleida
INTERSPEECH5
2010 The Alborada-I3A Corpus of Disordered Speech
Oscar Saz-Torralba, Eduardo Lleida, Carlos Vaquero, William Ricardo Rodríguez-Dueñas
LREC2
2010 Unsupervised Data-Driven Feature Vector Normalization With Acoustic Model Adaptation for Robust Speech Recognition
abstract
In this paper, an unsupervised data-driven robust speech recognition approach is proposed based on a joint feature vector normalization and acoustic model adaptation. Feature vector normalization reduces the acoustic mismatch between training and testing conditions by mapping the feature vectors towards the training space. Model adaptation modifies the parameters of the acoustic models to match the test space. However, since neither is optimal, both approaches use an intermediate space between training and testing spaces to map either the feature vectors or acoustic models. The joint optimization of both approaches provides a common intermediate space with a better match between normalized feature vectors and adapted acoustic models. In this paper, feature vector normalization is based on a minimum mean square error (MMSE) criterion. A class dependent multi-environment model linear normalization (CD-MEMLIN) based on two classes (silence/speech) with a cross probability model (CD-MEMLIN-CPM) is used. CD-MEMLIN-CPM assumes that each class of clean and noisy spaces can be modeled with a Gaussian mixture model (GMM), training a linear transformation for each pair of Gaussians in an unsupervised data-driven training process. This feature vector normalization maps the recognition space feature vector to a normalized space. The acoustic model adaptation maps the training space to the normalized space by defining a set of linear transformations over an expanded HMM-state space, compensating for those degradations that the feature vector normalization is not able to model, like rotations. Experiments have been carried out with the Spanish SpeechDat Car database and Aurora 2 databases using both the standard Mel-frequency cepstral coefficient (MFCC) and advanced ETSI front-ends. Consistent improvements were reached for both corpora and front-ends. Using the standard MFCC front-end, a 92.08% average improvement on WER for Spanish SpeechDat Car and a 69.75% average improvement for clean condition evaluation of Aurora 2 was obtained, improving those results reached with ETSI advanced front-end (83.28% and 67.41%, respectively). Using the ETSI advanced front-end with the proposed solution, a 75.47% average improvement was obtained for the clean condition evaluation of Aurora 2 database.
Luis Buera, Antonio Miguel, Oscar Saz-Torralba, Alfonso Ortega Giménez, Eduardo Lleida
IEEE Trans. Speech Audio Process.5
2009 A study of pronunciation verification in a speech therapy application
abstract
Techniques are presented for detecting phoneme level mispronunciations in utterances obtained from a population of impaired children speakers. The intended application of these approaches is to use the resulting confidence measures to provide feedback to patients concerning the quality of pronunciations in utterances arising within interactive speech therapy sessions. The pronunciation verification scenario involves presenting utterances of known words to a phonetic decoder and generating confusion networks from the resulting phone lattices. Confidence measures are derived from the posterior probabilities obtained from the confusion networks. Phoneme level mispronunciation detection performance was significantly improved with respect to a baseline system by optimizing acoustic models and pronunciation models in the phonetic decoder and applying a nonlinear mapping to the confusion network posteriors.
Shou-Chun Yin, Richard C. Rose, Oscar Saz-Torralba, Eduardo Lleida
ICASSP4
2009 Unsupervised training scheme with non-stereo data for empirical feature vector compensation
Luis Buera, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida, Richard M. Stern
INTERSPEECH4
2009 Differential vector quantization of feature vectors for distributed speech recognition
abstract
Distributed speech recognition arises for solving computational limitations of mobile devices like PDAs or mobile phones. Due to bandwidth restrictions, it is necessary to develop efficient transmission techniques of acoustic features in Automatic Speech Recognition applications. This paper presents a technique for compressing acoustic feature vectors based on Differential Vector Quantization. It is a combination of Vector Quantization and Differential encoding schemes. Recognition experiments have been carried out, showing that the proposed method outperforms the ETSI standard VQ system, and classical VQ schemes for different codebook lengths and situations. With the proposed scheme, bit rates as low as 2.1 kbps can be used without decreasing the performance of the ASR system in terms of WER compared with a system without quantization. Index Terms: speech recognition, distributed systems, vector quantization
José Enrique García Laínez, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
INTERSPEECH4
2009 Local projections and support vector based feature selection in speech recognition
Antonio Miguel, Alfonso Ortega Giménez, Luis Buera, Eduardo Lleida
INTERSPEECH4
2009 Graphical models for discrete hidden Markov models in speech recognition
Antonio Miguel, Alfonso Ortega Giménez, Luis Buera, Eduardo Lleida
INTERSPEECH4
2009 Real-time live broadcast news subtitling system for Spanish
Alfonso Ortega Giménez, José Enrique García Laínez, Antonio Miguel, Eduardo Lleida
INTERSPEECH4
2009 Combination of acoustic and lexical speaker adaptation for disordered speech recognition
abstract
This paper presents an approach to provide of lexical adaptation in Automatic Speech Recognition (ASR) of the disordered speech from a group of young impaired speakers. The outcome of an Acoustic Phonetic Decoder (APD) is used to learn new lexical variants of the 57-word vocabulary and add them to a lexicon personalized to each user. The possibilities of combination of this lexical adaptation with acoustic adaptation achieved through traditional Maximum A Posteriori (MAP) approaches are furtherer explored, and the results show the importance of matching the lexicon in the ASR decoding phase to the lexicon used for the acoustic adaptation.
Oscar Saz-Torralba, Eduardo Lleida, Antonio Miguel
INTERSPEECH2
2009 Tools and Technologies for Computer-Aided Speech and Language Therapy
Oscar Saz-Torralba, Shou-Chun Yin, Eduardo Lleida, Richard C. Rose, Carlos Vaquero, William Ricardo Rodríguez-Dueñas
Speech Commun.3
2008 E-inclusion technologies for the speech handicapped
abstract
This paper addresses the problem that disabled people face when accessing the new systems and technologies that are available nowadays. The use of speech technologies, specially helpful for motor handicapped people, becomes unapproachable when these people also suffer speech impairments, making the gap in the society wider for them. As a way to include speech impaired people in the technological society of today, two lines of work have been carried out. On one hand, a computer-aided speech therapy software has been developed for the speech training of children with different disabilities. This tool, available for free distribution, makes use of different state-of-the-art speech technologies to train different levels of the language. As a result of this work, the software is being used currently in several centers for special education with a very encouraging feedback about the capabilities of the system. On the other hand, research on the use of automatic speech recognition (ASR) systems for the speech impaired has been carried out. This work has focused on current techniques of speaker adaptation to know how these techniques, fruitfully used in other tasks, can deal with this specific kind of speech. The use of Maximum A Posterior (MAP) obtains an improvement of 60.61% compared to the results of a baseline speaker independent model.
Carlos Vaquero, Oscar Saz-Torralba, Eduardo Lleida, William Ricardo Rodríguez-Dueñas
ICASSP3
2008 Feature vector normalization with combined standard and throat microphones for robust ASR
abstract
We propose on-line unsupervised compensation technique for robust speech recognition that combines standard and throat microphone feature vectors. The solution, called Multi-Environment Model-based LInear Normalization with Throat microphone information, MEMLINT, is an extension of MEM-LIN formulation. Hence, standard microphone noisy space and throat microphone space are modelled as GMMs and a set of linear transformations are learnt from data associated to each pair of Gaussians (one for each GMM) using training stereo data. On the other hand, to compensate some kinds of degra-dation which are not considered in MEMLINT, we propose to use jointly an on-line unsupervised acoustic model adaptation method based on rotation transformations over an expanded HMM-state space (augMented stAte space acousTic dEcoder, MATE). Some experiments with an own recorded database were carried out, showing that the proposed approach significantly outperforms the single microphone approach. Index Terms: Throat microphone, robust speech recognition, feature vector normalization.
Luis Buera, Antonio Miguel, Oscar Saz-Torralba, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH5
2008 Verifying pronunciation accuracy from speakers with neuromuscular disorders
abstract
This paper presents a study of confidence measure based techniques for detecting phoneme level mispronunciations in utterances from impaired children with neuromuscular disorders. Several different adaptation scenarios are investigated to determine the effects of mismatched speaker characteristics and mismatched task domain on the ability to verify the phoneme level pronunciations. These techniques are evaluated in the context of a speech corpus where utterances were elicited from children in interactive speech therapy sessions involving a multimodal game-like environment. Results are presented in terms of phone detection charateristics where, for example, equal error rates of as low as 16.2 % were obtained for detecting instances where phonemes were deleted by impaired speakers. 1.
Shou-Chun Yin, Richard C. Rose, Oscar Saz-Torralba, Eduardo Lleida
INTERSPEECH4
2008 Capturing Local Variability for Speaker Normalization in Speech Recognition
abstract
The new model reduces the impact of local spectral and temporal variability by estimating a finite set of spectral and temporal warping factors which are applied to speech at the frame level. Optimum warping factors are obtained while decoding in a locally constrained search. The model involves augmenting the states of a standard hidden Markov model (HMM), providing an additional degree of freedom. It is argued in this paper that this represents an efficient and effective method for compensating local variability in speech which may have potential application to a broader array of speech transformations. The technique is presented in the context of existing methods for frequency warping-based speaker normalization for ASR. The new model is evaluated in clean and noisy task domains using subsets of the Aurora 2, the Spanish Speech-Dat-Car, and the TIDIGITS corpora. In addition, some experiments are performed on a Spanish language corpus collected from a population of speakers with a range of speech disorders. It has been found that, under clean or not severely degraded conditions, the new model provides improvements over the standard HMM baseline. It is argued that the framework of local warping is an effective general approach to providing more flexible models of speaker variability.
Antonio Miguel, Eduardo Lleida, Richard C. Rose, Luis Buera, Oscar Saz-Torralba, Alfonso Ortega Giménez
IEEE Trans. Speech Audio Process.2
2007 Robust speech recognition with on-line unsupervised acoustic feature compensation
abstract
An on-line unsupervised hybrid compensation technique is proposed to reduce the mismatch between training and testing conditions. It combines multi-environment model based linear normalization with cross-probability model based on GMMs (MEMLIN CPM) with a novel acoustic model adaptation method based on rotation transformations. Hence, a set of rotation transformations is estimated with clean and MEMLIN CPM-normalized training data by linear regression in an unsupervised process. Thus, in testing, each MEMLIN CPM normalized frame is decoded using a modified Viterbi algorithm and expanded acoustic models, which are obtained from the reference ones and the set of rotation transformations. To test the proposed solution, some experiments with Spanish SpeechDat Car database were carried out. MEMLIN CPM over standard ETSI front-end parameters reaches 83.89% of average improvement in WER, while the introduced hybrid solution goes up to 92.07%. Also, the proposed hybrid technique was tested with Aurora 2 database, obtaining an average improvement of 68.88% with clean training.
Luis Buera, Antonio Miguel, Eduardo Lleida, Oscar Saz-Torralba, Alfonso Ortega Giménez
ASRU3
2007 On the jointly unsupervised feature vector normalization and acoustic model compensation for robust speech recognition
Luis Buera, Antonio Miguel, Eduardo Lleida, Oscar Saz-Torralba, Alfonso Ortega Giménez
INTERSPEECH3
2007 Evaluation of the combined use of MEMLIN and MLLR on the non-native adaptation task of hiwire project database
abstract
This paper describes the performance of the combination of Multi-Environment Model-based LInear Normalization, MEMLIN, which provides an estimation of the uncorrupted feature vector, with Maximum Likelihood Linear Regression, MLLR, for the collected database under the auspices of the IST-EU STREP project HIWIRE. In this work the results for the nonnative adaptation task (NNA) are presented. The HIWIRE project database consist on command and control aeronautics application utterances pronounced by non-native speakers which are digitally corrupted with airplane cockpit noise. Thus, three noise conditions are defined: low, medium and high noise. In the proposed system, each MEMLIN-normalized feature vector is decoded using the MLLR-adapted acoustic models. The experiments show that an important improvement is reached combining MEMLIN and MLLR methods for all kinds of non-native speakers and noise conditions. Index Terms: Hiwire project, robust speech recognition, nonnative adaptation task.
Luis Buera, Antonio Miguel, Oscar Saz-Torralba, Eduardo Lleida, Alfonso Ortega Giménez
INTERSPEECH4
2007 Cepstral Vector Normalization Based on Stereo Data for Robust Speech Recognition
abstract
In this paper, a set of feature vector normalization methods based on the minimum mean square error (MMSE) criterion and stereo data is presented. They include multi-environment model-based linear normalization (MEMLIN), polynomial MEMLIN (P-MEMLIN), multi-environment model-based histogram normalization (MEMHIN), and phoneme-dependent MEMLIN (PD-MEMLIN). Those methods model clean and noisy feature vector spaces using Gaussian mixture models (GMMs). The objective of the methods is to learn a transformation between clean and noisy feature vectors associated with each pair of clean and noisy model Gaussians. The direct approach to learn the transformation is by using stereo data; that is, noisy feature vectors and the corresponding clean feature vectors. In this paper, however, a nonstereo data based training procedure, is presented. The transformations can be modeled just like a bias vector (MEMLIN), or by using a first-order polynomial (P-MEMLIN) or a nonlinear function based on histogram equalization (MEMHIN). Further improvements are obtained by using phoneme-dependent bias vector transformation (PD-MEMLIN). In PD-MEMLIN, the clean and noisy feature vector spaces are split into several phonemes, and each of them is modeled as a GMM. Those methods achieve significant word error rate improvements over others that are based on similar targets. The experimental results using the SpeechDat Car database show an average improvement in word error rate greater than 68% in all cases compared to the baseline when using the original clean acoustic models, and up to 83% when training acoustic models on the new normalized feature space
Luis Buera, Eduardo Lleida, Antonio Miguel, Alfonso Ortega Giménez, Oscar Saz-Torralba
IEEE Trans. Speech Audio Process.2
2006 Stability Control in a Two-Channel Speech Reinforcement System for Vehicles
abstract
This paper presents a two-channel speech reinforcement system for cars able to improve the communication between the front and the rear passengers. One of the problems of this kind of systems is that they must operate in closed-loop, as acoustic feedback paths appear due to the short distance between loudspeakers and microphones. This feedback paths can make the system become unstable and acoustic echo control is needed in order to ensure stability. The system must perform two plant identifications for each channel. One of them is an open-loop identification and the other one is closed-loop. We propose here the use of echo suppression filters specially designed for closed-loop subsystems along with echo suppression filters for open-loop subsystems based on the optimal filtering theory. Results about the performance of the proposed system are provided
Alfonso Ortega Giménez, Eduardo Lleida, Enrique Masgrau, Luis Buera, Antonio Miguel
ICASSP (5)2
2006 Time-dependent cross-probability model for multi-environment model based LInear normalization
Luis Buera, Eduardo Lleida, Juan A. Nolazco-Flores, Antonio Miguel, Alfonso Ortega Giménez
INTERSPEECH2
2006 Local transformation models for speech recognition
abstract
This paper presents a novel acoustic modeling framework that naturally extends the Hidden Markov Model (HMM) approach. The novel models reduce the errors caused by speaker variability by means of a local spectral mismatch reduction. A more complex and flexible speech production scheme can be assumed, in which the local temporal and frequency elastic deformations of the speech are captured by the model. In the new framework the states of a standard HMM, which are usually associated with temporal transitions, are expanded so that a new degree of freedom for the model is provided and it is then possible to estimate an optimum frequency warping factor at the same time as the decoder finds the best state sequence. In the local spectral warping based models the states become time-frequency related states and the number of parameters of the model is comparable to the standard HMM since they share a certain amount of parameters as it will be shown. The novel models are evaluated in the noise-free TIDIGITS corpus, which includes connected digits uttered by male, female and children. It has been found that, under speaker group (age-gender) mismatch conditions, the local frequency warping reduced Word Error Rate (WER) in mean by a 70%, using the initial models. When matched speaker group conditions were tested the error was reduced in mean in a 9.7% after reestimating the models. Index Terms: speaker variability, local frequency warping.
Antonio Miguel, Eduardo Lleida, Alfons Juan-Císcar, Luis Buera, Alfonso Ortega Giménez, Oscar Saz-Torralba
INTERSPEECH2
2006 Study of time and frequency variability in pathological speech and error reduction methods for automatic speech recognition
abstract
In this work, we study the variations in the time and frequency domains inside a Spanish language corpus of speakers with nonpathological and pathological speech. We show how pathological speech has a greater variability in the duration of the words than non-pathological speech, while in the frequency domain we show that the vowels confusability increases by a 18%. The baseline experiments in Automatic Speech Recognition (ASR) with this corpus demonstrate that this variability causes a loss in the performance of ASR systems. To reduce the impact of time and frequency variability we use a recent Vocal Tract Length Normalization (VTLN) system: MATE (augMented stAte space acousTic modEl), as a way of improving the performance of ASR systems when dealing with speakers who suffer any kind of speech pathology. Experiments with MATE show a 17.04 % and 11.19 % WER reduction by using frequency and time MATE respectively. 1.
Oscar Saz-Torralba, Antonio Miguel, Eduardo Lleida, Alfonso Ortega Giménez, Luis Buera
INTERSPEECH3
2006 Design and acquisition of a telephone spontaneous speech dialogue corpus in Spanish: DIHANA
José-Miguel Benedí, Eduardo Lleida, Amparo Varona, María José Castro Bleda, Isabel Galiano, Raquel Justo, Iñigo López de Letona, Antonio Miguel
LREC2
2005 Lip Reading for Robust Speech Recognition on Embedded Devices
abstract
In this article a complete audio-visual speech recognition system suitable for embedded devices is presented. As visual feature extraction algorithms active shape models (ASM) and discrete cosine transformation (DCT) have been investigated and discussed for an embedded implementation. The audio-visual information integration has also been designed by taking into account device limitations. It is well known that the use of visual cues improves the recognition results especially in scenarios with high level of acoustical noise. We wanted to compare the performance of lip reading and the conventional noise reduction systems in these degraded scenarios, as well as the combination of both kinds of solutions. Important improvements are obtained especially for nonstationary background noise like voice interference, car acceleration or indicator clicks. For this kind of noise lip reading outperforms the results obtained with conventional noise reduction technologies.
Jesus F. Guitarte Perez, Alejandro F. Frangi, Eduardo Lleida, Klaus Lukas
ICASSP (1)3
2005 Robust speech recognition in cars using phoneme dependent multi-environment linear normalization
abstract
In this paper a Phoneme-Dependent Multi-Environment Models based LInear feature Normalization, PD-MEMLIN, is presented. The target of this algorithm is to learn the difference between clean and noisy feature vectors associated to a pair of gaussians of the same phoneme (one for a clean model, and the other one for a noisy model), for each basic defined environment. These differences are estimated in a previous training process with stereo data. In order to compensate some of the problems of the independence assumption of the feature vectors components and the mismatch error between perfect and proposed transformations, two approaches have been proposed too: a multi-environment rotation transformation algorithm, and the use of transformed space acoustic models. Some experiments with SpeechDat Car database were carried out in order to study the behavior of the proposed techniques in a real acoustic environment. The experimental results show an average improvement of more than 77% using PD-MEMLIN, and more than 85% using transformed space acoustic models and multienvironment rotation transformation, concerning the baseline.
Luis Buera, Eduardo Lleida, Antonio Miguel, Alfonso Ortega Giménez
INTERSPEECH2
2005 Augmented state space acoustic decoding for modeling local variability in speech
abstract
This paper presents a decoding method for automatic speech recognition (ASR) that reduces the impact of local spectral and temporal variabilities on ASR performance. The procedure involves augmenting the standard Viterbi search for an optimum state sequence with a locally constrained search for optimum degrees of spectral warping or temporal warping applied to individual analysis frames. It is argued in the paper that this represents an efficient and effective method for compensating for local variability in speech which may have potential application to a broader array of speech transformations. The techniques are presented in the context of existing methods for frequency warping based speaker normalization and existing methods for computation of dynamic features for ASR. The modified decoding algorithms were evaluated in both clean and noisy task domains using
Antonio Miguel, Eduardo Lleida, Richard C. Rose, Luis Buera, Alfonso Ortega Giménez
INTERSPEECH2
2005 Acoustic feedback cancellation in speech reinforcement systems for vehicles
Alfonso Ortega Giménez, Eduardo Lleida, Enrique Masgrau, Luis Buera, Antonio Miguel
INTERSPEECH2
2005 Speech Reinforcement System for Car Cabin Communications
abstract
A speech reinforcement system is presented to improve communication between the front and the rear passengers in large motor vehicles. This type of communication can be difficult due to a number of factors, including distance between speakers, noise and lack of visual contact. The system described makes use of a set of microphones to pick up the speech of each passenger, then it amplifies these signals and plays them back to the cabin through the car audio loudspeaker system. The two main problems are noise amplification and electro-acoustic coupling between loudspeakers and microphones. To overcome these problems the system uses a set of acoustic echo cancellers, echo suppression filters and noise reduction stages. In this paper, the stability of a speech reinforcement system is studied. We propose a solution based on echo cancellers and residual echo suppression filters. The spectral estimation method for the power spectral density of the residual echo existing after the echo canceller is presented along with the derivation of the optimal residual echo suppression filter. Some results about the performance of the proposed system are also provided.
Alfonso Ortega Giménez, Eduardo Lleida, Enrique Masgrau
IEEE Trans. Speech Audio Process.2
2004 Multi-environment models based linear normalization for speech recognition in car conditions
abstract
A multi-environment adaptation technique, based on minimum mean squared error estimation, is proposed. MEMLIN (multi-environment models based linear normalization) consists of a feature adaptation using stereo data and several basic defined environments. The target of this algorithm is to learn the difference between clean and noisy feature vectors associated to a pair of Gaussians (one for a clean model, and the other for a noisy model), for each basic environment. This knowledge, the associated Gaussians, the conditional probability between clean and noisy Gaussians, and the environment are the data used to compensate the mismatch between clean and noisy vectors. This algorithm obtains important improvements regarding other techniques that look for similar targets. The experimental results with the SpeechDat Car database shows an average improvement of more than 68%, concerning the baseline, over 7 different defined environments.
Luis Buera, Eduardo Lleida, Antonio Miguel, Alfonso Ortega Giménez
ICASSP (1)2
2004 AV@CAR: A Spanish Multichannel Multimodal Corpus for In-Vehicle Automatic Audio-Visual Speech Recognition
Alfonso Ortega Giménez, Federico Sukno, Eduardo Lleida, Alejandro F. Frangi, Antonio Miguel, Luis Buera, Ernesto Zacur
LREC3
2003 Residual echo power estimation for speech reinforcement systems in vehicles
abstract
In acoustic echo cancelation systems, some residual echo exists after the acoustic echo canceler (AEC) due to the fact that the adaptive filter does not model exactly the impulse response of the Loudspeaker-Enclosure-Microphone (LEM) path.This is specially important in feedback acoustic environments like speech reinforcement systems for cars where this residual echo can make the system become unstable. In order to suppress this residual echo remaining after the AEC, postfiltering is the most used technique.The optimal filter that ensures stability without attenuating the speech signal depends on the power spectral density (psd) of the residual echo that must be estimated. This paper presents a residual echo psd estimation method needed to obtain the optimal echo suppression filter in speech reinforcement systems for cars.
Alfonso Ortega Giménez, Eduardo Lleida, Enrique Masgrau
INTERSPEECH2
2002 Cabin car communication system to improve communications inside a car
abstract
This paper presents a cabin car communication system (CCCS) to improve the communication among passengers inside a car. Noise, distance between speakers and many other factors make difficult to maintain a conversation inside a car. The CCCS picks up the speech of each passenger, amplifies it, and uses the car loudspeaker system to return it into the cabin. Two problems arise when designing a CCCS; the electro-acoustic coupling between loudspeakers and microphones, and the amplification of the inside car noise. As a result of the first problem, the system may become unstable. To maintain the stability of the system, the CCCS makes use of a robust acoustic echo cancellation scheme based on system identification and a Wiener echo suppressor. Using a noise reduction system based on Wiener filtering reduces second problem, noise amplification. Experimental results showing the performance of the system in terms of acoustic echo and noise reduction and speech reinforce are presented. A system with 2-input/2-output channels has been built on a DSP board for medium size cars and minivan vehicles.
Alfonso Ortega Giménez, Eduardo Lleida, Enrique Masgrau, Fernando Gallego
ICASSP2
2001 Acoustic echo control and noise reduction for cabin car communication
abstract
A Cabin Car Communication System (CCCS) has the goal of improving the communication among passengers inside the car. Wind, road and engine noise, the distance between passengers and other factors make difficult the communication inside vehicles. The driver must often look away from the road and passengers move out of normal seating positions. The CCCS makes use of a set of microphones to pick up the speech and the car-audio loudspeakers to reinforce the sound level. This scenario presents a great challenger for acoustic echo control and noise reduction. Acoustic echo control must prevent the overall system from howling and becoming unstable with the additional problem that the system must always work with double talk. The noise reduction must clean the microphone signal to avoid the reinforce of the noise inside the car. In this paper, we describe a combined acoustic echo control and noise reduction algorithm suitable for cabin car communication systems. We present experimental results in terms of echo return loss enhancement, stability and maximum reinforce without howling. 1.
Eduardo Lleida, Enrique Masgrau, Alfonso Ortega Giménez
INTERSPEECH1
2001 A new method for epoch detection based on the Cohen's class of time frequency representations
abstract
This paper presents a new method for detecting the instants of glottal closure (IGC), or epochs, in noisy environments based on the Cohen's class time-frequency representations (TFR). We define a detection function inspired in a time-frequency formulation for optimum detection and apply a morphologic closing over it to determine the epochs. It compares favorably with other methods (e.g., the SIFT-based and the Frobenius norm [FN] function) in different levels of Gaussian noise. Experiments are carried over a data base composed of ten speakers.
Juan L. Navarro-Mesa, Eduardo Lleida, Asunción Moreno
IEEE Signal Process. Lett.2
2000 Utterance verification in continuous speech recognition: decoding and training procedures
abstract
This paper introduces a set of acoustic modeling and decoding techniques for utterance verification (UV) in hidden Markov model (HMM) based continuous speech recognition (CSR). Utterance verification in this work implies the ability to determine when portions of a hypothesized word string correspond to incorrectly decoded vocabulary words or out-of-vocabulary words that may appear in an utterance. This capability is implemented here as a likelihood ratio (LR) based hypothesis testing procedure for, verifying individual words in a decoded string. There are two UV techniques that are presented here. The first is a procedure for estimating the parameters of UV models during training according to an optimization criterion which is directly related to the LR measure used in UV. The second technique is a speech recognition decoding procedure where the "best" decoded path is defined to be that which optimizes a LR criterion. These techniques were evaluated in terms of their ability to improve UV performance on a speech dialog task over the public switched telephone network. The results of an experimental study presented in the paper shows that LR based parameter estimation results in a significant improvement in UV performance for this task. The study also found that the use of the LR based decoding procedure, when used in conjunction with models trained using the LR criterion, can provide as much as an 11% improvement in UV performance when compared to existing UV procedures. Finally, it was also found that the performance of the LR decoder was highly dependent on the use of the LR criterion in training acoustic models. Several observations are made in the paper concerning the formation of confidence measures for UV and the interaction of these techniques with statistical language models used in ASR.
Eduardo Lleida, Richard C. Rose
IEEE Trans. Speech Audio Process.1
1999 Microphone array design for robust speech acquisition and recognition
Julián Fernández-Navajas, Eduardo Lleida, Enrique Masgrau
EUROSPEECH2
1999 Performance comparison of several adaptive schemes for microphone array beamforming
Enrique Masgrau, Luis Aguilar, Eduardo Lleida
EUROSPEECH3
1999 An improved speech endpoint detection system in noisy environments by means of third-order spectra
abstract
We exploit the properties of the third-order spectra in two proposals to obtain speech detection functions. One is obtained from the principal domain of the bispectrum and the other one from the integrated bispectrum. We have developed a threshold-based system in which the detection functions can be easily integrated. Experiments show the improvement on the detection scores over the energy-based function in noisy environments.
Juan L. Navarro-Mesa, Asunción Moreno, Eduardo Lleida
IEEE Signal Process. Lett.3
1998 Robust continuous speech recognition system based on a microphone array
abstract
A robust speech recognition system for videoconference applications is presented based on a microphone array. By means of a microphone array, the speech recognition system is able to know the position of the users and increase the signal-to-noise ratio (SNR) between the desired speaker signal and the interference from the other users. The user positions are estimated by means of the combination of a direction of arrival (DOA) estimation method with a speaker identification system. The beamforming is performed by using the spatial references of the desired speaker and the interference locations. A minimum variance algorithm with spatial constraints working in the frequency domain is used to design the weights of the broadband microphone array. Results of the speech recognition system are reported in a simulated environment with several users asking questions to a geographic data base.
Eduardo Lleida, Julián Fernández-Navajas, Enrique Masgrau
ICASSP1
1997 Speech recognition using automatically derived acoustic baseforms
abstract
This paper investigates procedures for obtaining user-configurable speech recognition vocabularies. These procedures use example utterances of vocabulary words to perform unsupervised automatic acoustic baseform determination in terms of a set of speaker independent subword acoustic units. Several procedures, differing both in the definition of subword acoustic model context and in the phonotactic constraints used in decoding have been investigated. The tendency of input utterances to contain out-of-vocabulary or non-speech information is accounted for using likelihood ratio based utterance verification procedures. Comparisons of different definitions of the likelihood ratio used for utterance verification and of different criteria for estimating parameters used in the likelihood ratio test have been performed. The performance of these techniques has been evaluated on utterances taken from a trial of a voice label recognition service.
Richard C. Rose, Eduardo Lleida
ICASSP2
1997 Non-quadratic criterion algorithms for speech enhancement
Enrique Masgrau, Eduardo Lleida, Luis Vicente
EUROSPEECH2
1996 Efficient decoding and training procedures for utterance verification in continuous speech recognition
abstract
It is often necessary in speech recognition to include a mechanism for verifying decoded utterances in order to account for incorrectly decoded vocabulary words and utterances corresponding to words or sounds that are not included in a prespecified lexicon. This paper describes an utterance verification procedure for hidden Markov model (HMM) based continuous speech recognition that is based on a likelihood ratio (LR) criterion. There are two important contributions. The first is a search algorithm which directly optimizes a likelihood ratio criterion. This search algorithm is important because it allows decoding to be performed in speech recognition according to the same measure of confidence that is used in hypothesis testing. The second contribution is a corresponding training procedure for estimating model parameters which also directly optimizes the same likelihood ratio criterion. These techniques are applied to spontaneous spoken queries in the context of a "movie locator" dialog system.
Eduardo Lleida, Richard C. Rose
ICASSP1
1996 Pitch detection and voiced/unvoiced decision algorithm based on wavelet transforms
abstract
An improvement o f an existing Pitch Detection Algorithm is presented in this paper.The solution reduces the computational load of its precedent algorithm and introduces a voiced/unvoiced decision step to reduce the number of errors.The eciency of this improved system is tested with a semi-automatically segmented speech data base according to the information delivered by an attached laryngograph signal.The results show its periodicity detection.
Léonard Janer, Juan José Bonet, Eduardo Lleida
ICSLP3
1996 Wavelet transforms for non-uniform speech recogntion systems
Léonard Janer, Josep Martí, Climent Nadeu, Eduardo Lleida
ICSLP4
1996 Likelihood ratio decoding and confidence measures for continuous speech recognition
Eduardo Lleida, Richard C. Rose
ICSLP1
1996 A user-configurable system for voice label recognition
Richard C. Rose, Eduardo Lleida, G. W. Erhart, R. V. Grubbe
ICSLP2
1995 Semantic decoding of speech in constrained domains
Antonio Bonafonte, José B. Mariño, Eduardo Lleida
EUROSPEECH3
1993 TELEMACO - a real time keyword spotting application for voice dialling
abstract
The problem of detecting a given set of words in fluent speech is one of the most interesting topics in speech recognition for practical real time applications. This paper present the TELEMACO system for automatic voice dialling which is based on the use of the keyword spotting technology to detect the dialling commands in fluent speech used by the IBERCOM Spanish telephone system. The user interface is based on a PC computer with a DSP board. The DSP board runs the speech recognition task and the interaction with the telephone line. The keyword vocabulary is composed by commands to dial, answer, hang-up, cancel, recall, store, etc. Each keyword is modeled by means of a discrete Hidden Markov Model. To model the non-keyword speech, syllabic fillers models and background models are used. The keyword spotting algorithm is a null grammar time-synchronous Viterbi search with two search spaces. The first search is over all the models (keywords and fillers) and the second search is only over the filler model. Thus, we can compare the behaviour of the filler model with the candidate keyword for each detection and decide if the keyword has been uttered or not. This process is done frame by frame. When a keyword is detected, the DSP board send the recognition word to the PC to take the corresponding action. The system has been implemented in a Windows environment.
Eduardo Lleida, José B. Mariño, Arturo Moreno
EUROSPEECH1
1993 Out-of-vocabulary word modelling and rejection for keyword spotting
abstract
This paper presents a combination of out-of-vocabulary (OOV) word modeling and rejection techniques in an attempt to accept utterances embedding a keyword and reject utterances with nonkeywords. The goal of this research is to develop a robust, task-independent Spanish keyword spotter and to develop a method for optimizing confidence thresholds for a particular context. To model OOV words, we employed both word and sub-word units as fillers, combined with n-gram language models. We also introduce a methodology for optimizing confidence thresholds to control the tradeoffs between acceptance, confirmation, and rejection of utterances. Our experiments are based on a Mexican Spanish auto-attendant system using the SpeechWorks recognizer release 6.5 Second Edition, in which we achieved a reduction in error of 8.9% as compared to the baseline system. Most of the error reduction is attributed to better keyword detection in utterances that contain both keywords and OOV words.
Eduardo Lleida, José B. Mariño, Josep M. Salavedra, Antonio Bonafonte, Enric Monte-Moreno
EUROSPEECH1
1993 Albayzin speech database: design of the phonetic corpus
abstract
This paper describes the phonetic content of Albayzin, a spoken database for Spanish designed for speech recognition purposes. A statistical study of a large sample of spontaneous speech is presented, and the phonetic and statistical criteria for the final constitution of the database are discussed. Finally, the contents of the phonetic database are analyzed
Asunción Moreno, Dolors Poch-Olivé, Antonio Bonafonte, Eduardo Lleida, Joaquim Llisterri, José B. Mariño, Climent Nadeu
EUROSPEECH4
1992 On the AR modelling of the one-sided autocorrelation sequence for noisy speech recognition
abstract
Speech recognition in noisy environments remains an unsolved problem even in the case of isolated word recognition with small vocabularies. Recently, several techniques have been proposed to alleviate this problem. Concretely, two closely related parameterization techniques based on an AR modelling in the autocorrelation domain called SMC [1] and OSALPC [2] have shown good results using speech contaminated by additive white noise. The aim of this paper is twofold: to compare several techniques based on an AR modelling in the autocorrelation domain, including SMC and OSALPC, and to find the optimum model order and cepstral liftering for noisy conditions.
Javier Hernando, Climent Nadeu, Eduardo Lleida
ICSLP3
1992 Syllabic fillers for Spanish HMM keyword spotting
Eduardo Lleida, José B. Mariño, Josep M. Salavedra, Antonio Bonafonte
ICSLP1
1992 Smoothing hidden Markov models ay means of a self organizing feature map
Enric Monte-Moreno, José B. Mariño, Eduardo Lleida
ICSLP3
1991 Demisyllable-based HMM spotting for continuous speech recognition
abstract
The authors describe the acoustic processor of a Spanish continuous speech recognition system based on demisyllable units. The acoustic processor is based on a spotting algorithm which takes as input the unknown utterance, the HMM (hidden Markov model) of the reference demisyllables, and the lexical knowledge in terms of a finite state network. The spotting algorithm is a modified version of the one-step Viterbi algorithm with multiple hypotheses. The output of the system is a lattice of word hypotheses suitable to be parsed by a linguistic analyzer. The proposed acoustic processor was tested using the integers from 0 to 1000 and telephonic numbers in a speaker-independent approach. The results show the good performance of the demisyllable as a recognition unit for the Spanish language and the efficiency of the spotting algorithm.>
Eduardo Lleida, José B. Mariño, Climent Nadeu, Joan Salavedra
ICASSP1
1991 Two level continuous speech recognition using demisyllable-based HMM word spotting
abstract
Peer Reviewed
Eduardo Lleida, José B. Mariño, Climent Nadeu, Albert Oliveras
EUROSPEECH1
1990 Statistical feature selection for isolated word recognition
abstract
A procedure for feature selection in isolated word recognition is discussed. The feature selection is performed in two steps. The first step takes into account the temporal correlation among feature vectors in order to obtain a transformation matrix which projects the initial template of N feature vectors to a new space where they are uncorrelated. This step gives a new template of M feature vectors, where M>
Eduardo Lleida, Climent Nadeu, Enric Monte-Moreno, José B. Mariño
ICASSP1
1989 Recognition of numbers and strings of numbers by using demisyllables: one speaker experiment
abstract
This communication reports the use of demisyllables for continuous speech recognition in a specific application: the recognition of Spanish numbers. After a brief outline of the recognition system, a description of demisyllable syntactic constraints and one-speaker reference generation is provided. Finally, the recognition performance is assessed by means of two experiments: the recognition of integer numbers from zero to one thousand and telephone numbers uttered in a Spanish way (strings of integers from zero to ninety nine), in both applications the results that the system yielded were excellent.
José B. Mariño, Climent Nadeu, Asunción Moreno, Eduardo Lleida, Enric Monte-Moreno
EUROSPEECH4
1989 New backpropagation algorithm using quadratic potential functions, and an experiment on isolated word recognition
abstract
This paper presents a new algorithm to train multilayered perceptrons, using quadratical potential functions. This new algorithm is compared in an isolated word recognition task, with the back propagation algorithm that uses linear combinations of the inputs. Some pattern recognition techniques are also used to reduce the dimensinality of the input pattern in order to reduce the computational burden of the training and the recognition. The algorithm that uses quadratic potential functions yields better results in the recognition task.
Enric Monte-Moreno, Eduardo Lleida, José B. Mariño
EUROSPEECH2
1989 Modeling of the analytic spectrum for speech recognition
Climent Nadeu, Eduardo Lleida, Javier Hernando
EUROSPEECH2