Antonio Miguel

dblp:94/8307 · also Antonio Miguel Artiaga · DBLP profile ↗
← Back
68ranked-venue papers
8as first author
10since 2021 · last 2024
0000-0001-5803-4316ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 55 · 6 first-author · 7 since 2021
YearPublicationVenuePosition
2024 Predefined Prototypes for Intra-Class Separation and Disentanglement
abstract
International audience
Antonio Almudévar, Théo Mariotte, Alfonso Ortega Giménez, Marie Tahon, Luis Vicente, Antonio Miguel, Eduardo Lleida
INTERSPEECH6
2023 Variational Classifier for Unsupervised Anomalous Sound Detection under Domain Generalization
Antonio Almudévar, Alfonso Ortega Giménez, Luis Vicente, Antonio Miguel, Eduardo Lleida
INTERSPEECH4
2023 Direction of Arrival Estimation of Sound Sources Using Icosahedral CNNs
abstract
In this paper, we present a new model for Direction of Arrival (DOA) estimation of sound sources based on an Icosahedral Convolutional Neural Network (CNN) applied over SRP-PHAT power maps computed from the signals received by a microphone array. This icosahedral CNN is equivariant to the 60 rotational symmetries of the icosahedron, which represent a good approximation of the continuous space of spherical rotations, and can be implemented using standard 2D convolutional layers, having a lower computational cost than most of the spherical CNNs. In addition, instead of using fully connected layers after the icosahedral convolutions, we propose a new soft-argmax function that can be seen as a differentiable version of the argmax function and allows us to solve the DOA estimation as a regression problem interpreting the output of the convolutional layers as a probability distribution. We prove that using models that fit the equivariances of the problem allows us to outperform other state-of-the-art models with a lower computational cost and more robustness, obtaining root mean square localization errors lower than$10^{\circ }$even in scenarios with a reverberation time$\mathbf {T_{60}}$of$1.5 \,\mathrm{s}$.
David Diaz-Guerra, Antonio Miguel, José Ramón Beltrán
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 aDCF Loss Function for Deep Metric Learning in End-to-End Text-Dependent Speaker Verification Systems
abstract
Metric learning approaches have widely expanded to the training of Speaker Verification (SV) systems based on Deep Neural Networks (DNNs), by using a loss function more consistent with the evaluation process than the traditional identification losses. However, these methods do not consider the performance measure and can involve high computational cost, for example, the need for a careful pair or triplet data selection. This paper proposes the approximated Detection Cost Function (aDCF) loss, which is a loss function based on the measure of the decision errors in SV systems, namely the False Rejection Rate (FRR) and the False Acceptance Rate (FAR). With aDCF loss as the training objective function, the end-to-end system learns how to minimize decision errors. Furthermore, we replace the typical linear layer as the last layer of DNN by a cosine distance layer, which reduces the difference between the metric in the training process and the metric during evaluation. aDCF loss function was evaluated in RSR2015-Part I and RSR2015-Part II datasets for text-dependent speaker verification. The system trained with aDCF loss outperforms all the state-of-the-art functions employed in this paper in both parts of the database.
Victoria Mingote, Antonio Miguel, Dayana Ribas González, Alfonso Ortega Giménez, Eduardo Lleida
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Memory Layers with Multi-Head Attention Mechanisms for Text-Dependent Speaker Verification
abstract
In this paper, we explore an approach based on memory layers and multi-head attention mechanisms to improve in an efficient way the performance of text-dependent speaker verification (SV) systems. The most extended SV systems based on Deep Neural Networks (DNN) extract the embedding of the utterance from the average pooling of the temporal dimension after processing. Unlike previous works, we can exploit the phonetic knowledge needed for text-dependent SV systems by combining the temporal attention of multiple parallel heads with the phonetic embeddings extracted from a phonetic classification network, which helps to guide to the attention mechanism with the role of the positional embedding. The addition of a memory layer to a text-dependent SV system was tested on the RSR2015-part II and DeepMine-part I databases, where, in both cases outperformed the baseline result and the reference system based on the same transformer network without the memory layer.
Victoria Mingote, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
ICASSP2
2021 Unsupervised Representation Learning for Speech Activity Detection in the Fearless Steps Challenge 2021
abstract
International audience
Pablo Gimeno, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
Interspeech3
2021 Log-Likelihood-Ratio Cost Function as Objective Loss for Speaker Verification Systems
abstract
International audience
Victoria Mingote, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
Interspeech2
2021 gpuRIR: A python library for room impulse response simulation with GPU acceleration
David Diaz-Guerra, Antonio Miguel, José Ramón Beltrán
Multim. Tools Appl.2
2021 Generalizing AUC Optimization to Multiclass Classification for Audio Segmentation With Limited Training Data
abstract
Area under the ROC curve (AUC) optimisation techniques developed for neural networks have recently demonstrated their capabilities in different audio and speech related tasks. However, due to its intrinsic nature, AUC optimisation has focused only on binary tasks so far. In this paper, we introduce an extension to the AUC optimisation framework so that it can be easily applied to an arbitrary number of classes, aiming to overcome the issues derived from training data limitations in deep learning solutions. Building upon the multiclass definitions of the AUC metric found in the literature, we define two new training objectives using a one-versus-one and a one-versus-rest approach. In order to demonstrate its potential, we apply them in an audio segmentation task with limited training data that aims to differentiate 3 classes: foreground music, background music and no music. Experimental results show that our proposal can improve the performance of audio segmentation systems significantly compared to traditional training criteria such as cross entropy.
Pablo Gimeno, Victoria Mingote, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
IEEE Signal Process. Lett.4
2021 Robust Sound Source Tracking Using SRP-PHAT and 3D Convolutional Neural Networks
abstract
In this article, we present a new single sound source DOA estimation and tracking system based on the well-known SRP-PHAT algorithm and a three-dimensional Convolutional Neural Network. It uses SRP-PHAT power maps as input features of a fully convolutional causal architecture that uses 3D convolutional layers to accurately perform the tracking of a sound source even in highly reverberant scenarios where most of the state of the art techniques fail. Unlike previous methods, since we do not use bidirectional recurrent layers and all our convolutional layers are causal in the time dimension, our system is feasible for real-time applications and it provides a new DOA estimation for each new SRP-PHAT map. To train the model, we introduce a new procedure to simulate random trajectories as they are needed during the training, equivalent to an infinite-size dataset with high flexibility to modify its acoustical conditions such as the reverberation time. We use both acoustical simulations on a large range of reverberation times and the actual recordings of the LOCATA dataset to prove the robustness of our system and its good performance even using low-resolution SRP-PHAT maps.
David Diaz-Guerra, Antonio Miguel, José Ramón Beltrán
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Knowledge Distillation and Random Erasing Data Augmentation for Text-Dependent Speaker Verification
abstract
This paper explores the Knowledge Distillation (KD) approach and a data augmentation technique to improve the generalization ability and robustness of text-dependent speaker verification (SV) systems. The KD method consists of two neural networks, known as Teacher and Student, where the student is trained to replicate the predictions from the teacher, so it learns their variability during the training process. To provide robustness to the distillation process, we apply Random Erasing (RE), a data augmentation technique which was created to improve the generalization ability of the neural networks. We have developed two alternatives of the combination of KD and RE, which, produce a more robust system with better performance, since the student network can learn from teacher predictions of data not existing in the original dataset. All alternatives were tested on RSR2015-Part I database, where the proposed variants outperform reference system based on a single network using RE.
Victoria Mingote, Antonio Miguel, Dayana Ribas González, Alfonso Ortega Giménez, Eduardo Lleida
ICASSP2
2020 Partial AUC Optimisation Using Recurrent Neural Networks for Music Detection with Limited Training Data
Pablo Gimeno, Victoria Mingote, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
INTERSPEECH4
2020 Training Speaker Enrollment Models by Network Optimization
Victoria Mingote, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH2
2020 Optimization of the area under the ROC curve using neural network supervectors for text-dependent speaker verification
Victoria Mingote, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
Comput. Speech Lang.2
2019 Disentangling and Learning Robust Representations with Natural Clustering
abstract
Learning representations that disentangle the underlying factors of variability in data is an intuitive way to achieve generalization in deep models. In this work, we address the scenario where generative factors present a multimodal distribution due to the existence of class distinction in the data. We propose N-VAE, a model which is capable of separating factors of variation which are exclusive to certain classes from factors that are shared among classes. This model implements an explicitly compositional latent variable structure by defining a class-conditioned latent space and a shared latent space. We show its usefulness for detecting and disentangling class-dependent generative factors as well as its capacity to generate artificial samples which contain characteristics unseen in the training data.
Javier Antorán, Antonio Miguel
ICMLA2
2019 Speech Enhancement with Wide Residual Networks in Reverberant Environments
abstract
This paper proposes a speech enhancement method which exploits the high potential of residual connections in a Wide Residual Network architecture. This is supported on single dimensional convolutions computed alongside the time domain, which is a powerful approach to process contextually correlated representations through the temporal domain, such as speech feature sequences. We find the residual mechanism extremely useful for the enhancement task since the signal always has a linear shortcut and the non-linear path enhances it in several steps by adding or subtracting corrections. The enhancement capability of the proposal is assessed by objective quality metrics evaluated with simulated and real samples of reverberated speech signals. Results show that the proposal outperforms the state-of-the-art method called WPE, which is known to effectively reduce reverberation and greatly enhance the signal. The proposed model, trained with artificial synthesized reverberation data, was able to generalize to real room impulse responses for a variety of conditions (e.g. different room sizes, $RT_{60}$, near & far field). Furthermore, it achieves accuracy for real speech with reverberation from two different datasets.
Jorge Llombart, Dayana Ribas González, Antonio Miguel, Luis Vicente, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH3
2019 Progressive Speech Enhancement with Residual Connections
abstract
This paper studies the Speech Enhancement based on Deep Neural Networks. The proposed architecture gradually follows the signal transformation during enhancement by means of a visualization probe at each network block. Alongside the process, the enhancement performance is visually inspected and evaluated in terms of regression cost. This progressive scheme is based on Residual Networks. During the process, we investigate a residual connection with a constant number of channels, including internal state between blocks, and adding progressive supervision. The insights provided by the interpretation of the network enhancement process leads us to design an improved architecture for the enhancement purpose. Following this strategy, we are able to obtain speech enhancement results beyond the state-of-the-art, achieving a favorable trade-off between dereverberation and the amount of spectral distortion.
Jorge Llombart, Dayana Ribas González, Antonio Miguel, Luis Vicente, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH3
2019 Language Recognition Using Triplet Neural Networks
Victoria Mingote, Diego Castán, Mitchell McLaren, Mahesh Kumar Nandwana, Alfonso Ortega Giménez, Eduardo Lleida, Antonio Miguel
INTERSPEECH7
2019 Optimization of False Acceptance/Rejection Rates and Decision Threshold for End-to-End Text-Dependent Speaker Verification Systems
Victoria Mingote, Antonio Miguel, Dayana Ribas González, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH2
2019 ViVoLAB Speaker Diarization System for the DIHARD 2019 Challenge
Ignacio Viñals, Pablo Gimeno, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
INTERSPEECH4
2019 Phonetically-Aware Embeddings, Wide Residual Networks with Time-Delay Neural Networks and Self Attention Models for the 2018 NIST Speaker Recognition Evaluation
Ignacio Viñals, Dayana Ribas González, Victoria Mingote, Jorge Llombart, Pablo Gimeno, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH6
2018 Estimation of the Number of Speakers with Variational Bayesian PLDA in the DIHARD Diarization Challenge
Ignacio Viñals, Pablo Gimeno, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
INTERSPEECH4
2017 Tied Hidden Factors in Neural Networks for End-to-End Speaker Recognition
abstract
In this paper we propose a method to model speaker and session variability and able to generate likelihood ratios using neural networks in an end-To-end phrase dependent speaker verification system. As in Joint Factor Analysis, the model uses tied hidden variables to model speaker and session variability and a MAP adaptation of some of the parameters of the model. In the training procedure our method jointly estimates the network parameters and the values of the speaker and channel hidden variables. This is done in a two-step backpropagation algorithm, first the network weights and factor loading matrices are updated and then the hidden variables, whose gradients are calculated by aggregating the corresponding speaker or session frames, since these hidden variables are tied. The last layer of the network is defined as a linear regression probabilistic model whose inputs are the previous layer outputs. This choice has the advantage that it produces likelihoods and additionally it can be adapted during the enrolment using MAP without the need of a gradient optimization. The decisions are made based on the ratio of the output likelihoods of two neural network models, speaker adapted and universal background model. The method was evaluated on the RSR2015 database.
Antonio Miguel, Jorge Llombart, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH1
2017 Domain Adaptation of PLDA Models in Broadcast Diarization by Means of Unsupervised Speaker Clustering
Ignacio Viñals, Alfonso Ortega Giménez, Jesús Villalba 0001, Antonio Miguel, Eduardo Lleida
INTERSPEECH4
2016 Analysis of speech quality measures for the task of estimating the reliability of speaker verification decisions
Jesús Villalba 0001, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
Speech Commun.3
2016 Bayesian Networks to Model the Variability of Speaker Verification Scores in Adverse Environments
abstract
State-of-the-art speaker recognition technology attains great performance in controlled conditions. However, when the speech segments suffer distortions like noise or reverberation performance can severely deteriorate, this fact motivated us to investigate how score distributions diverge from the ideal ones in degraded conditions. We propose a Bayesian network model that assumes that two scores exist: one observed and another one hidden. The observed score or noisy score is the one given by the speaker verification system. Meanwhile, the hidden score or clean score is the ideal score that we would obtain in a trial with high-quality speech. A set of quality measures helps to relate both scores. We applied this network to two tasks. The first one consists in rejecting unreliable trials, i.e., trials that we cannot assure whether they are target or nontarget. We prove that this method outperforms previous approaches, based on another type of Bayesian networks. The second task is to compute an improved likelihood ratio, dependent on the quality measures. This ratio improved calibration in noisy conditions.
Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Variational Bayesian PLDA for speaker diarization in the MGB challenge
abstract
This paper describes the ViVoLab speaker diarization system for the Multi-Genre Broadcast (MGB) Challenge at ASRU2015. The challenge data consisted of BBC TV programmes of different genres. Diarization followed a longitudinal setup, i.e., the speakers of the current episode had to be linked to the speakers in previous episodes of the same show. We propose a system based on the i-vector paradigm. After an initial segmentation step, we compute an i-vector per speech segment. Then, a generative model based on Bayesian PLDA clusters the speakers. In this model, the speaker labels are latent variables that we optimize by variational Bayes iterations. The number of speakers in each episode was decided by maximizing the variational lower bound. The system includes several phases of segment-merging and re-clustering. We re-compute i-vectors after each merging step, which reduces the i-vector uncertainty. This approach attained a DER around 30% in the development set.
Jesús Villalba 0001, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
ASRU3
2015 Spoofing detection with DNN and one-class SVM for the ASVspoof 2015 challenge
abstract
Speaker verification systems have achieved great performance in recent times. However, we usually measure performance on a ideal scenarios with naive impostors that do not modify their voices to impersonate the target speakers. The fact of impersonating a legitimate user is known as spoofing attack. Recent works show the vulnerability of current speaker verification technology to several types of attacks. Most of these works use non-public databases and different performance measures, which makes difficult to compare approaches. The spoofing challenge (ASVspoof 2015) tries to overcome this problem by proposing a common evaluation framework. This paper describes our submission to the challenge. We proposed to use spectral log-filter-bank and relative phase shift features as input to classifiers based on deep neural networks (DNN). The first of our classifiers used DNN posteriors to decide if the trial is spoof or non-spoof. The second used a bottleneck feature from the DNN as input to a one-class SVM. The one-class SVM models the distribution of legitimate speech, not needing spoofing data for training. We fused the score of the different classifiers to produce our final submission. Our system attained very competitive results with EER<0.05% in 9 out of 10 spoofing types.
Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH2
2014 Factor analysis with sampling methods for text dependent speaker recognition
Antonio Miguel, Jesús Villalba 0001, Alfonso Ortega Giménez, Eduardo Lleida, Carlos Vaquero
INTERSPEECH1
2014 Low bit rate compression methods of feature vectors for distributed speech recognition
José Enrique García Laínez, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
Speech Commun.3
2013 Segmentation-by-classification system based on factor analysis
abstract
This paper proposes a novel audio segmentation-by-classification system based on Factor Analysis (FA) with a channel compensation matrix for each class and scoring the fixed-length segments as the log-likelihood ratio between class/no-class. The scores are smoothed and the most probable sequence is computed with a Viterbi algorithm. The system described here is designed to segment and classify the audio files coming from broadcast programs into five different classes: speech (SP), speech with noise (SN), speech with music (SM), music (MU) or others (OT). This task was proposed in the Albayzin 2010 evaluation campaign. The system is compared with the winning system of the evaluation achieving lower error rate in SP and SN. These classes represent 3/4 of the total amount of the data. Therefore, the FA segmentation system gets a reduction in the average segmentation error rate.
Diego Castán, Alfonso Ortega Giménez, Jesús Villalba 0001, Antonio Miguel, Eduardo Lleida
ICASSP4
2013 Prosodic features and formant modeling for an ivector-based language recognition system
abstract
The prosody of a language is encoded in syllable length, loudness and pitch. These attributes make humans perceive rhythm, stress and intonation in speech. Depending on the language, these speech properties vary, making language classification possible. On the other hand, formants are the resonance frequencies of the vocal tract, depend heavily on the position adopted by the articulatory organs, and are especially useful to disambiguate vowel sounds. In this paper prosodic and formant information are combined to build a generative language identification system based on Gaussian models fed with iVectors. The system is evaluated on the NIST LRE09 database and the inclusion of formant information gives about 50% relative improvement for the 30 s task over a prosodic system without it. The fusion with a state-of-the-art acoustic system based on shifted delta cepstral coefficients (SDC) shows the complementarity of both approaches.
David Martínez González, Eduardo Lleida, Alfonso Ortega Giménez, Antonio Miguel
ICASSP4
2013 Suprasegmental information modelling for autism disorder spectrum and specific language impairment classification
David Martínez González, Dayana Ribas González, Eduardo Lleida, Alfonso Ortega Giménez, Antonio Miguel
INTERSPEECH5
2013 A new Bayesian network to assess the reliability of speaker verification decisions
abstract
In some situations the quality of the signals involved in a speaker verification trial is not as good as needed to take a reliable decision. In this work, we present a new method based on Bayesian networks and quality measures to estimate if the trial decision is reliable. We present experiments on the NIST SRE2010 dataset degraded with additive noise. A system well calibrated for clean speech, produces a large actual DCF on the degraded dataset. We use our method to discard the unreliable trials and achieve a dramatic improvement of the cost values. We also prove that our method outperforms previously published approaches.
Jesús Villalba 0001, Eduardo Lleida, Alfonso Ortega Giménez, Antonio Miguel
INTERSPEECH4
2013 The I3a speaker recognition system for NIST SRE12: post-evaluation analysis
abstract
The I3A submission for the recent NIST 2012 speaker recognition evaluation (SRE) was based on the i-vector approach with a multi-channel PLDA classifier. This PLDA is modified so that, for each i-vector, the between-class covariance depends on the type of channel where the segment was recorded (telephone,interviews,clean, noisy, etc). In this paper, we present the description of our submission and a detailed post-evaluation analysis of the results. We analyze several factors affecting performance: enrollment data selection, classifier type, scoring technique, calibration, known and unknown non-targets, target speakers included or not in development, segment duration, noise level and noise type. Some of these factor are new in this evaluation. After post-evaluation, actual costs improve by 15– 43% depending on the common condition.
Jesús Villalba 0001, Eduardo Lleida, Alfonso Ortega Giménez, Antonio Miguel
INTERSPEECH4
2013 Quality Assessment for Speaker Diarization and Its Application in Speaker Characterization
abstract
There are many applications related to speaker characterization, specially in telephone environments, where large datasets are available but not directly useful since there are two speakers involved in every recording. Even with very accurate speaker diarization systems, we can expect to find some recordings with low diarization accuracy. The use of these recordings may reduce the accuracy of any speaker characterization technology. Therefore, it is highly desirable to detect those recordings where the speakers are correctly segmented, in order to discard or process manually the remaining ones before feeding them into the application. In this work we propose a set of confidence measures to assess the quality of a hypothetical diarization output, in order to detect those recordings that are correctly segmented. We show that these confidence measures enable us to retrieve most of the desired recordings from a given dataset, discarding those recordings that degrade the overall accuracy of an application that make use of speaker characterization technologies.
Carlos Vaquero, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
IEEE Trans. Speech Audio Process.3
2012 Beam-Search Formant Tracking Algorithm Based on Trajectory Functions for Continuous Speech
José Enrique García Laínez, Dayana Ribas González, Antonio Miguel, Eduardo Lleida, José Ramón Calvo de Lara
CIARP3
2011 Multi-site heterogeneous system fusions for the Albayzin 2010 Language Recognition Evaluation
abstract
Best language recognition performance is commonly obtained by fusing the scores of several heterogeneous systems. Regardless the fusion approach, it is assumed that different systems may contribute complementary information, either because they are developed on different datasets, or because they use different features or different modeling approaches. Most authors apply fusion as a final resource for improving performance based on an existing set of systems. Though relative performance gains decrease as larger sets of systems are considered, best performance is usually attained by fusing all the available systems, which may lead to high computational costs. In this paper, we aim to discover which technologies combine the best through fusion and to analyse the factors (data, features, modeling methodologies, etc.) that may explain such a good performance. Results are presented and discussed for a number of systems provided by the participating sites and the organizing team of the Albayzin 2010 Language Recognition Evaluation. We hope the conclusions of this work help research groups make better decisions in developing language recognition technology.
Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel, David Martínez González, Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida, Alberto Abad, Oscar Koller, Isabel Trancoso, Paula Lopez-Otero, Laura Docío Fernández, Carmen García-Mateo, Rahim Saeidi, Mehdi Soufifar, Tomi Kinnunen, Torbjørn Svendsen, Pasi Fränti
ASRU8
2011 I3A Language Recognition System for Albayzin 2010 LRE
David Martínez González, Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH3
2011 Bayesian Networks for Discrete Observation Distributions in Speech Recognition
abstract
Traditionally, in speech recognition, the hidden Markov model state emission probability distributions are usually associated to continuous random variables, by using Gaussian mixtures. Thus, complex multimodal inter-feature dependencies are not accurately modeled by Gaussian models, since they are unimodal distributions and mixtures of Gaussians are needed in these complex cases, but this is done in a loose and inefficient way. Graphical models provide a precise and simple mechanism to model the dependencies among two or more variables. This paper proposes the use of discrete random variables as observations and graphical models to extract the internal dependence structure in the feature vectors. Therefore, speech features are quantized to a small number of levels, in order to obtain a tractable model. These quantized speech features provide a mechanism to increase the robustness against noise uncertainty. In addition, discrete random variables allow the learning of joint statistics of the observation densities. A method to estimate a graphical model with a constrained number of dependencies is shown in this paper, being a special kind of Bayesian network. Experimental results show that by using this modeling, better performance can be obtained compared to standard baseline systems.
Antonio Miguel, Alfonso Ortega Giménez, Luis Buera, Eduardo Lleida
IEEE Trans. Speech Audio Process.1
2010 Non-linear predictive vector quantization of feature vectors for distributed speech recognition
José Enrique García Laínez, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
INTERSPEECH3
2010 Confidence measures for speaker segmentation and their relation to speaker verification
Carlos Vaquero, Alfonso Ortega Giménez, Jesús Villalba 0001, Antonio Miguel, Eduardo Lleida
INTERSPEECH4
2010 Unsupervised Data-Driven Feature Vector Normalization With Acoustic Model Adaptation for Robust Speech Recognition
abstract
In this paper, an unsupervised data-driven robust speech recognition approach is proposed based on a joint feature vector normalization and acoustic model adaptation. Feature vector normalization reduces the acoustic mismatch between training and testing conditions by mapping the feature vectors towards the training space. Model adaptation modifies the parameters of the acoustic models to match the test space. However, since neither is optimal, both approaches use an intermediate space between training and testing spaces to map either the feature vectors or acoustic models. The joint optimization of both approaches provides a common intermediate space with a better match between normalized feature vectors and adapted acoustic models. In this paper, feature vector normalization is based on a minimum mean square error (MMSE) criterion. A class dependent multi-environment model linear normalization (CD-MEMLIN) based on two classes (silence/speech) with a cross probability model (CD-MEMLIN-CPM) is used. CD-MEMLIN-CPM assumes that each class of clean and noisy spaces can be modeled with a Gaussian mixture model (GMM), training a linear transformation for each pair of Gaussians in an unsupervised data-driven training process. This feature vector normalization maps the recognition space feature vector to a normalized space. The acoustic model adaptation maps the training space to the normalized space by defining a set of linear transformations over an expanded HMM-state space, compensating for those degradations that the feature vector normalization is not able to model, like rotations. Experiments have been carried out with the Spanish SpeechDat Car database and Aurora 2 databases using both the standard Mel-frequency cepstral coefficient (MFCC) and advanced ETSI front-ends. Consistent improvements were reached for both corpora and front-ends. Using the standard MFCC front-end, a 92.08% average improvement on WER for Spanish SpeechDat Car and a 69.75% average improvement for clean condition evaluation of Aurora 2 was obtained, improving those results reached with ETSI advanced front-end (83.28% and 67.41%, respectively). Using the ETSI advanced front-end with the proposed solution, a 75.47% average improvement was obtained for the clean condition evaluation of Aurora 2 database.
Luis Buera, Antonio Miguel, Oscar Saz-Torralba, Alfonso Ortega Giménez, Eduardo Lleida
IEEE Trans. Speech Audio Process.2
2009 Unsupervised training scheme with non-stereo data for empirical feature vector compensation
Luis Buera, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida, Richard M. Stern
INTERSPEECH2
2009 Differential vector quantization of feature vectors for distributed speech recognition
abstract
Distributed speech recognition arises for solving computational limitations of mobile devices like PDAs or mobile phones. Due to bandwidth restrictions, it is necessary to develop efficient transmission techniques of acoustic features in Automatic Speech Recognition applications. This paper presents a technique for compressing acoustic feature vectors based on Differential Vector Quantization. It is a combination of Vector Quantization and Differential encoding schemes. Recognition experiments have been carried out, showing that the proposed method outperforms the ETSI standard VQ system, and classical VQ schemes for different codebook lengths and situations. With the proposed scheme, bit rates as low as 2.1 kbps can be used without decreasing the performance of the ASR system in terms of WER compared with a system without quantization. Index Terms: speech recognition, distributed systems, vector quantization
José Enrique García Laínez, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
INTERSPEECH3
2009 Local projections and support vector based feature selection in speech recognition
Antonio Miguel, Alfonso Ortega Giménez, Luis Buera, Eduardo Lleida
INTERSPEECH1
2009 Graphical models for discrete hidden Markov models in speech recognition
Antonio Miguel, Alfonso Ortega Giménez, Luis Buera, Eduardo Lleida
INTERSPEECH1
2009 Real-time live broadcast news subtitling system for Spanish
Alfonso Ortega Giménez, José Enrique García Laínez, Antonio Miguel, Eduardo Lleida
INTERSPEECH3
2009 Combination of acoustic and lexical speaker adaptation for disordered speech recognition
abstract
This paper presents an approach to provide of lexical adaptation in Automatic Speech Recognition (ASR) of the disordered speech from a group of young impaired speakers. The outcome of an Acoustic Phonetic Decoder (APD) is used to learn new lexical variants of the 57-word vocabulary and add them to a lexicon personalized to each user. The possibilities of combination of this lexical adaptation with acoustic adaptation achieved through traditional Maximum A Posteriori (MAP) approaches are furtherer explored, and the results show the importance of matching the lexicon in the ASR decoding phase to the lexicon used for the acoustic adaptation.
Oscar Saz-Torralba, Eduardo Lleida, Antonio Miguel
INTERSPEECH3
2008 Feature vector normalization with combined standard and throat microphones for robust ASR
abstract
We propose on-line unsupervised compensation technique for robust speech recognition that combines standard and throat microphone feature vectors. The solution, called Multi-Environment Model-based LInear Normalization with Throat microphone information, MEMLINT, is an extension of MEM-LIN formulation. Hence, standard microphone noisy space and throat microphone space are modelled as GMMs and a set of linear transformations are learnt from data associated to each pair of Gaussians (one for each GMM) using training stereo data. On the other hand, to compensate some kinds of degra-dation which are not considered in MEMLINT, we propose to use jointly an on-line unsupervised acoustic model adaptation method based on rotation transformations over an expanded HMM-state space (augMented stAte space acousTic dEcoder, MATE). Some experiments with an own recorded database were carried out, showing that the proposed approach significantly outperforms the single microphone approach. Index Terms: Throat microphone, robust speech recognition, feature vector normalization.
Luis Buera, Antonio Miguel, Oscar Saz-Torralba, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH2
2008 Improving Robustness in Frequency Warping-Based Speaker Normalization
abstract
This letter addresses the issue of frequency warping-based speaker normalization in noisy acoustic environments. Techniques are developed for improving the robustness of localized estimates of frequency warping transformations that are applied to individual observation vectors. It is shown that automatic speech recognition (ASR) performance can be improved by using speaker class-dependent distributions characterizing frequency warping transformations associated with individual hidden Markov model states. The effect of these techniques is demonstrated over a range of noise conditions on the Aurora 2 speech corpus.
Richard C. Rose, Antonio Miguel, Alireza Keyvani
IEEE Signal Process. Lett.2
2008 Capturing Local Variability for Speaker Normalization in Speech Recognition
abstract
The new model reduces the impact of local spectral and temporal variability by estimating a finite set of spectral and temporal warping factors which are applied to speech at the frame level. Optimum warping factors are obtained while decoding in a locally constrained search. The model involves augmenting the states of a standard hidden Markov model (HMM), providing an additional degree of freedom. It is argued in this paper that this represents an efficient and effective method for compensating local variability in speech which may have potential application to a broader array of speech transformations. The technique is presented in the context of existing methods for frequency warping-based speaker normalization for ASR. The new model is evaluated in clean and noisy task domains using subsets of the Aurora 2, the Spanish Speech-Dat-Car, and the TIDIGITS corpora. In addition, some experiments are performed on a Spanish language corpus collected from a population of speakers with a range of speech disorders. It has been found that, under clean or not severely degraded conditions, the new model provides improvements over the standard HMM baseline. It is argued that the framework of local warping is an effective general approach to providing more flexible models of speaker variability.
Antonio Miguel, Eduardo Lleida, Richard C. Rose, Luis Buera, Oscar Saz-Torralba, Alfonso Ortega Giménez
IEEE Trans. Speech Audio Process.1
2007 Robust speech recognition with on-line unsupervised acoustic feature compensation
abstract
An on-line unsupervised hybrid compensation technique is proposed to reduce the mismatch between training and testing conditions. It combines multi-environment model based linear normalization with cross-probability model based on GMMs (MEMLIN CPM) with a novel acoustic model adaptation method based on rotation transformations. Hence, a set of rotation transformations is estimated with clean and MEMLIN CPM-normalized training data by linear regression in an unsupervised process. Thus, in testing, each MEMLIN CPM normalized frame is decoded using a modified Viterbi algorithm and expanded acoustic models, which are obtained from the reference ones and the set of rotation transformations. To test the proposed solution, some experiments with Spanish SpeechDat Car database were carried out. MEMLIN CPM over standard ETSI front-end parameters reaches 83.89% of average improvement in WER, while the introduced hybrid solution goes up to 92.07%. Also, the proposed hybrid technique was tested with Aurora 2 database, obtaining an average improvement of 68.88% with clean training.
Luis Buera, Antonio Miguel, Eduardo Lleida, Oscar Saz-Torralba, Alfonso Ortega Giménez
ASRU2
2007 On the jointly unsupervised feature vector normalization and acoustic model compensation for robust speech recognition
Luis Buera, Antonio Miguel, Eduardo Lleida, Oscar Saz-Torralba, Alfonso Ortega Giménez
INTERSPEECH2
2007 Evaluation of the combined use of MEMLIN and MLLR on the non-native adaptation task of hiwire project database
abstract
This paper describes the performance of the combination of Multi-Environment Model-based LInear Normalization, MEMLIN, which provides an estimation of the uncorrupted feature vector, with Maximum Likelihood Linear Regression, MLLR, for the collected database under the auspices of the IST-EU STREP project HIWIRE. In this work the results for the nonnative adaptation task (NNA) are presented. The HIWIRE project database consist on command and control aeronautics application utterances pronounced by non-native speakers which are digitally corrupted with airplane cockpit noise. Thus, three noise conditions are defined: low, medium and high noise. In the proposed system, each MEMLIN-normalized feature vector is decoded using the MLLR-adapted acoustic models. The experiments show that an important improvement is reached combining MEMLIN and MLLR methods for all kinds of non-native speakers and noise conditions. Index Terms: Hiwire project, robust speech recognition, nonnative adaptation task.
Luis Buera, Antonio Miguel, Oscar Saz-Torralba, Eduardo Lleida, Alfonso Ortega Giménez
INTERSPEECH2
2007 Cepstral Vector Normalization Based on Stereo Data for Robust Speech Recognition
abstract
In this paper, a set of feature vector normalization methods based on the minimum mean square error (MMSE) criterion and stereo data is presented. They include multi-environment model-based linear normalization (MEMLIN), polynomial MEMLIN (P-MEMLIN), multi-environment model-based histogram normalization (MEMHIN), and phoneme-dependent MEMLIN (PD-MEMLIN). Those methods model clean and noisy feature vector spaces using Gaussian mixture models (GMMs). The objective of the methods is to learn a transformation between clean and noisy feature vectors associated with each pair of clean and noisy model Gaussians. The direct approach to learn the transformation is by using stereo data; that is, noisy feature vectors and the corresponding clean feature vectors. In this paper, however, a nonstereo data based training procedure, is presented. The transformations can be modeled just like a bias vector (MEMLIN), or by using a first-order polynomial (P-MEMLIN) or a nonlinear function based on histogram equalization (MEMHIN). Further improvements are obtained by using phoneme-dependent bias vector transformation (PD-MEMLIN). In PD-MEMLIN, the clean and noisy feature vector spaces are split into several phonemes, and each of them is modeled as a GMM. Those methods achieve significant word error rate improvements over others that are based on similar targets. The experimental results using the SpeechDat Car database show an average improvement in word error rate greater than 68% in all cases compared to the baseline when using the original clean acoustic models, and up to 83% when training acoustic models on the new normalized feature space
Luis Buera, Eduardo Lleida, Antonio Miguel, Alfonso Ortega Giménez, Oscar Saz-Torralba
IEEE Trans. Speech Audio Process.3
2006 Detecting Replay Attacks in Audiovisual Identity Verification
abstract
We describe an algorithm that detects a lack of correspondence between speech and lip motion by detecting and monitoring the degree of synchrony between live audio and visual signals. It is simple, effective, and computationally inexpensive; providing a useful degree of robustness against basic replay attacks and against speech or image forgeries. The method is based on a cross-correlation analysis between two streams of features, one from the audio signal and the other from the image sequence. We argue that such an algorithm forms an effective first barrier against several kinds of replay attack that would defeat existing verification systems based on standard multimodal fusion techniques. In order to provide an evaluation mechanism for the new technique we have augmented the protocols that accompany the BANCA multimedia corpus by defining new scenarios. We obtain 0% equal-error rate (EER) on the simplest scenario and 35% on a more challenging one
Hervé Bredin, Antonio Miguel, Ian H. Witten, Gérard Chollet
ICASSP (1)2
2006 Stability Control in a Two-Channel Speech Reinforcement System for Vehicles
abstract
This paper presents a two-channel speech reinforcement system for cars able to improve the communication between the front and the rear passengers. One of the problems of this kind of systems is that they must operate in closed-loop, as acoustic feedback paths appear due to the short distance between loudspeakers and microphones. This feedback paths can make the system become unstable and acoustic echo control is needed in order to ensure stability. The system must perform two plant identifications for each channel. One of them is an open-loop identification and the other one is closed-loop. We propose here the use of echo suppression filters specially designed for closed-loop subsystems along with echo suppression filters for open-loop subsystems based on the optimal filtering theory. Results about the performance of the proposed system are provided
Alfonso Ortega Giménez, Eduardo Lleida, Enrique Masgrau, Luis Buera, Antonio Miguel
ICASSP (5)5
2006 On the Interaction Between Speaker Normalization, Environment Compensation, and Discriminant Feature Space Transformations
abstract
This paper presents a study of the interaction between frequency warping based speaker normalization algorithms, environment compensation algorithms, and discriminant feature space transformations (DFT) in providing consistent reductions in ASR word error rate (WER) over a range of acoustic degradations. Performance improvements obtained using speaker normalization algorithms, including vocal tract length normalization (VTLN) and a newly proposed augmented state space acoustic decoder, are shown to improve substantially when applied in a discriminant feature space where acoustic environment compensation has been applied. Furthermore, the effects on ASR performance of the DFT are also shown to be enhanced by reducing within class variability by applying the DFT on a speaker and an environment normalized feature space
Richard C. Rose, Alireza Keyvani, Antonio Miguel
ICASSP (1)3
2006 Time-dependent cross-probability model for multi-environment model based LInear normalization
Luis Buera, Eduardo Lleida, Juan A. Nolazco-Flores, Antonio Miguel, Alfonso Ortega Giménez
INTERSPEECH4
2006 Local transformation models for speech recognition
abstract
This paper presents a novel acoustic modeling framework that naturally extends the Hidden Markov Model (HMM) approach. The novel models reduce the errors caused by speaker variability by means of a local spectral mismatch reduction. A more complex and flexible speech production scheme can be assumed, in which the local temporal and frequency elastic deformations of the speech are captured by the model. In the new framework the states of a standard HMM, which are usually associated with temporal transitions, are expanded so that a new degree of freedom for the model is provided and it is then possible to estimate an optimum frequency warping factor at the same time as the decoder finds the best state sequence. In the local spectral warping based models the states become time-frequency related states and the number of parameters of the model is comparable to the standard HMM since they share a certain amount of parameters as it will be shown. The novel models are evaluated in the noise-free TIDIGITS corpus, which includes connected digits uttered by male, female and children. It has been found that, under speaker group (age-gender) mismatch conditions, the local frequency warping reduced Word Error Rate (WER) in mean by a 70%, using the initial models. When matched speaker group conditions were tested the error was reduced in mean in a 9.7% after reestimating the models. Index Terms: speaker variability, local frequency warping.
Antonio Miguel, Eduardo Lleida, Alfons Juan-Císcar, Luis Buera, Alfonso Ortega Giménez, Oscar Saz-Torralba
INTERSPEECH1
2006 Study of time and frequency variability in pathological speech and error reduction methods for automatic speech recognition
abstract
In this work, we study the variations in the time and frequency domains inside a Spanish language corpus of speakers with nonpathological and pathological speech. We show how pathological speech has a greater variability in the duration of the words than non-pathological speech, while in the frequency domain we show that the vowels confusability increases by a 18%. The baseline experiments in Automatic Speech Recognition (ASR) with this corpus demonstrate that this variability causes a loss in the performance of ASR systems. To reduce the impact of time and frequency variability we use a recent Vocal Tract Length Normalization (VTLN) system: MATE (augMented stAte space acousTic modEl), as a way of improving the performance of ASR systems when dealing with speakers who suffer any kind of speech pathology. Experiments with MATE show a 17.04 % and 11.19 % WER reduction by using frequency and time MATE respectively. 1.
Oscar Saz-Torralba, Antonio Miguel, Eduardo Lleida, Alfonso Ortega Giménez, Luis Buera
INTERSPEECH2
2006 Design and acquisition of a telephone spontaneous speech dialogue corpus in Spanish: DIHANA
José-Miguel Benedí, Eduardo Lleida, Amparo Varona, María José Castro Bleda, Isabel Galiano, Raquel Justo, Iñigo López de Letona, Antonio Miguel
LREC8
2005 Robust speech recognition in cars using phoneme dependent multi-environment linear normalization
abstract
In this paper a Phoneme-Dependent Multi-Environment Models based LInear feature Normalization, PD-MEMLIN, is presented. The target of this algorithm is to learn the difference between clean and noisy feature vectors associated to a pair of gaussians of the same phoneme (one for a clean model, and the other one for a noisy model), for each basic defined environment. These differences are estimated in a previous training process with stereo data. In order to compensate some of the problems of the independence assumption of the feature vectors components and the mismatch error between perfect and proposed transformations, two approaches have been proposed too: a multi-environment rotation transformation algorithm, and the use of transformed space acoustic models. Some experiments with SpeechDat Car database were carried out in order to study the behavior of the proposed techniques in a real acoustic environment. The experimental results show an average improvement of more than 77% using PD-MEMLIN, and more than 85% using transformed space acoustic models and multienvironment rotation transformation, concerning the baseline.
Luis Buera, Eduardo Lleida, Antonio Miguel, Alfonso Ortega Giménez
INTERSPEECH3
2005 Augmented state space acoustic decoding for modeling local variability in speech
abstract
This paper presents a decoding method for automatic speech recognition (ASR) that reduces the impact of local spectral and temporal variabilities on ASR performance. The procedure involves augmenting the standard Viterbi search for an optimum state sequence with a locally constrained search for optimum degrees of spectral warping or temporal warping applied to individual analysis frames. It is argued in the paper that this represents an efficient and effective method for compensating for local variability in speech which may have potential application to a broader array of speech transformations. The techniques are presented in the context of existing methods for frequency warping based speaker normalization and existing methods for computation of dynamic features for ASR. The modified decoding algorithms were evaluated in both clean and noisy task domains using
Antonio Miguel, Eduardo Lleida, Richard C. Rose, Luis Buera, Alfonso Ortega Giménez
INTERSPEECH1
2005 Acoustic feedback cancellation in speech reinforcement systems for vehicles
Alfonso Ortega Giménez, Eduardo Lleida, Enrique Masgrau, Luis Buera, Antonio Miguel
INTERSPEECH5
2004 Multi-environment models based linear normalization for speech recognition in car conditions
abstract
A multi-environment adaptation technique, based on minimum mean squared error estimation, is proposed. MEMLIN (multi-environment models based linear normalization) consists of a feature adaptation using stereo data and several basic defined environments. The target of this algorithm is to learn the difference between clean and noisy feature vectors associated to a pair of Gaussians (one for a clean model, and the other for a noisy model), for each basic environment. This knowledge, the associated Gaussians, the conditional probability between clean and noisy Gaussians, and the environment are the data used to compensate the mismatch between clean and noisy vectors. This algorithm obtains important improvements regarding other techniques that look for similar targets. The experimental results with the SpeechDat Car database shows an average improvement of more than 68%, concerning the baseline, over 7 different defined environments.
Luis Buera, Eduardo Lleida, Antonio Miguel, Alfonso Ortega Giménez
ICASSP (1)3
2004 AV@CAR: A Spanish Multichannel Multimodal Corpus for In-Vehicle Automatic Audio-Visual Speech Recognition
Alfonso Ortega Giménez, Federico Sukno, Eduardo Lleida, Alejandro F. Frangi, Antonio Miguel, Luis Buera, Ernesto Zacur
LREC5