EDBT 2026 Demo / reviewers in the wild / expert
Alfonso Ortega Giménez
dblp:121/1854-1 · also Alfonso Ortega 0001
· DBLP profile ↗
72ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0002-3886-7748ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 60 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 58 · 5 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Aligning Multimodal Representations through an Information BottleneckabstractContrastive losses have been extensively used as a tool for multimodal representation learning. However, it has been empirically observed that their use is not effective to learn an aligned representation space. In this paper, we argue that this phenomenon is caused by the presence of modality-specific information in the representation space. Although some of the most widely used contrastive losses maximize the mutual information between representations of both modalities, they are not designed to remove the modality-specific information. We give a theoretical description of this problem through the lens of the Information Bottleneck Principle. We also empirically analyze how different hyperparameters affect the emergence of this phenomenon in a controlled experimental setup. Finally, we propose a regularization term in the loss function that is derived by means of a variational approximation and aims to increase the representational alignment. We analyze in a set of controlled experiments and real-world applications the advantages of including this regularization term. Antonio Almudévar, José Miguel Hernández-Lobato, Sameer Khurana, Ricard Marxer, Alfonso Ortega Giménez |
ICML | 5 |
| 2024 | Unsupervised multiple domain translation through controlled Disentanglement in variational autoencoderabstractUnsupervised Multiple Domain Translation is the task of transforming data from one domain to other domains without having paired data to train the systems. Typically, methods based on Generative Adversarial Networks (GANs) are used to address this task. However, our proposal exclusively relies on a modified version of a Variational Autoencoder. This modification consists of the use of two latent variables disentangled in a controlled way by design. One of this latent variables is imposed to depend exclusively on the domain, while the other one must depend on the rest of the variability factors of the data. Additionally, the conditions imposed over the domain latent variable allow for better control and understanding of the latent space. We empirically demonstrate that our approach works on different vision datasets improving the performance of other well known methods. Finally, we prove that, indeed, one of the latent variables stores all the information related to the domain and the other one hardly contains any domain information. Antonio Almudévar, Théo Mariotte, Alfonso Ortega Giménez, Marie Tahon |
ICASSP | 3 |
| 2024 | An Explainable Proxy Model for Multilabel Audio SegmentationabstractAudio signal segmentation is a key task for automatic audio indexing. It consists of detecting the boundaries of class-homogeneous segments in the signal. In many applications, explainable AI is a vital process for transparency of decision-making with machine learning. In this paper, we propose an explainable multilabel segmentation model that solves speech activity (SAD), music (MD), noise (ND), and overlapped speech detection (OSD) simultaneously. This proxy uses the non-negative matrix factorization (NMF) to map the embeddings used for the segmentation to the frequency domain. Experiments conducted on two datasets show similar performances as the pre-trained black box model while strong explainable features arise. Specifically, the frequency bins used for the decision can be easily identified at both the segment level (local explanations) and global level (class prototypes). Théo Mariotte, Antonio Almudévar, Marie Tahon, Alfonso Ortega Giménez |
ICASSP | 4 |
| 2024 | Predefined Prototypes for Intra-Class Separation and DisentanglementabstractInternational audience Antonio Almudévar, Théo Mariotte, Alfonso Ortega Giménez, Marie Tahon, Luis Vicente, Antonio Miguel, Eduardo Lleida |
INTERSPEECH | 3 |
| 2024 | Explainable by-design Audio Segmentation through Non-Negative Matrix Factorization and Probing
Martin Lebourdais, Théo Mariotte, Antonio Almudévar, Marie Tahon, Alfonso Ortega Giménez |
INTERSPEECH | 5 |
| 2023 | Variational Classifier for Unsupervised Anomalous Sound Detection under Domain Generalization
Antonio Almudévar, Alfonso Ortega Giménez, Luis Vicente, Antonio Miguel, Eduardo Lleida |
INTERSPEECH | 2 |
| 2022 | aDCF Loss Function for Deep Metric Learning in End-to-End Text-Dependent Speaker Verification SystemsabstractMetric learning approaches have widely expanded to the training of Speaker Verification (SV) systems based on Deep Neural Networks (DNNs), by using a loss function more consistent with the evaluation process than the traditional identification losses. However, these methods do not consider the performance measure and can involve high computational cost, for example, the need for a careful pair or triplet data selection. This paper proposes the approximated Detection Cost Function (aDCF) loss, which is a loss function based on the measure of the decision errors in SV systems, namely the False Rejection Rate (FRR) and the False Acceptance Rate (FAR). With aDCF loss as the training objective function, the end-to-end system learns how to minimize decision errors. Furthermore, we replace the typical linear layer as the last layer of DNN by a cosine distance layer, which reduces the difference between the metric in the training process and the metric during evaluation. aDCF loss function was evaluated in RSR2015-Part I and RSR2015-Part II datasets for text-dependent speaker verification. The system trained with aDCF loss outperforms all the state-of-the-art functions employed in this paper in both parts of the database. Victoria Mingote, Antonio Miguel, Dayana Ribas González, Alfonso Ortega Giménez, Eduardo Lleida |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Memory Layers with Multi-Head Attention Mechanisms for Text-Dependent Speaker VerificationabstractIn this paper, we explore an approach based on memory layers and multi-head attention mechanisms to improve in an efficient way the performance of text-dependent speaker verification (SV) systems. The most extended SV systems based on Deep Neural Networks (DNN) extract the embedding of the utterance from the average pooling of the temporal dimension after processing. Unlike previous works, we can exploit the phonetic knowledge needed for text-dependent SV systems by combining the temporal attention of multiple parallel heads with the phonetic embeddings extracted from a phonetic classification network, which helps to guide to the attention mechanism with the role of the positional embedding. The addition of a memory layer to a text-dependent SV system was tested on the RSR2015-part II and DeepMine-part I databases, where, in both cases outperformed the baseline result and the reference system based on the same transformer network without the memory layer. Victoria Mingote, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida |
ICASSP | 3 |
| 2021 | Unsupervised Representation Learning for Speech Activity Detection in the Fearless Steps Challenge 2021abstractInternational audience Pablo Gimeno, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida |
Interspeech | 2 |
| 2021 | Log-Likelihood-Ratio Cost Function as Objective Loss for Speaker Verification SystemsabstractInternational audience Victoria Mingote, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida |
Interspeech | 3 |
| 2021 | Generalizing AUC Optimization to Multiclass Classification for Audio Segmentation With Limited Training DataabstractArea under the ROC curve (AUC) optimisation techniques developed for neural networks have recently demonstrated their capabilities in different audio and speech related tasks. However, due to its intrinsic nature, AUC optimisation has focused only on binary tasks so far. In this paper, we introduce an extension to the AUC optimisation framework so that it can be easily applied to an arbitrary number of classes, aiming to overcome the issues derived from training data limitations in deep learning solutions. Building upon the multiclass definitions of the AUC metric found in the literature, we define two new training objectives using a one-versus-one and a one-versus-rest approach. In order to demonstrate its potential, we apply them in an audio segmentation task with limited training data that aims to differentiate 3 classes: foreground music, background music and no music. Experimental results show that our proposal can improve the performance of audio segmentation systems significantly compared to traditional training criteria such as cross entropy. Pablo Gimeno, Victoria Mingote, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida |
IEEE Signal Process. Lett. | 3 |
| 2020 | Knowledge Distillation and Random Erasing Data Augmentation for Text-Dependent Speaker VerificationabstractThis paper explores the Knowledge Distillation (KD) approach and a data augmentation technique to improve the generalization ability and robustness of text-dependent speaker verification (SV) systems. The KD method consists of two neural networks, known as Teacher and Student, where the student is trained to replicate the predictions from the teacher, so it learns their variability during the training process. To provide robustness to the distillation process, we apply Random Erasing (RE), a data augmentation technique which was created to improve the generalization ability of the neural networks. We have developed two alternatives of the combination of KD and RE, which, produce a more robust system with better performance, since the student network can learn from teacher predictions of data not existing in the original dataset. All alternatives were tested on RSR2015-Part I database, where the proposed variants outperform reference system based on a single network using RE. Victoria Mingote, Antonio Miguel, Dayana Ribas González, Alfonso Ortega Giménez, Eduardo Lleida |
ICASSP | 4 |
| 2020 | Partial AUC Optimisation Using Recurrent Neural Networks for Music Detection with Limited Training Data
Pablo Gimeno, Victoria Mingote, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida |
INTERSPEECH | 3 |
| 2020 | Training Speaker Enrollment Models by Network Optimization
Victoria Mingote, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida |
INTERSPEECH | 3 |
| 2020 | Shouted Speech Compensation for Speaker Verification Robust to Vocal Effort ConditionsabstractThe performance of speaker verification systems degrades when vocal effort conditions between enrollment and test (e.g., shouted vs. normal speech) are different. This is a potential situation in non-cooperative speaker verification tasks. In this paper, we present a study on different methods for linear compensation of embeddings making use of Gaussian mixture models to cluster shouted and normal speech domains. These compensation techniques are borrowed from the area of robustness for automatic speech recognition and, in this work, we apply them to compensate the mismatch between shouted and normal conditions in speaker verification. Before compensation, shouted condition is automatically detected by means of logistic regression. The process is computationally light and it is performed in the back-end of an x-vector system. Experimental results show that applying the proposed approach in the presence of vocal effort mismatch yields up to 13.8% equal error rate relative improvement with respect to a system that applies neither shouted speech detection nor compensation. Santi Prieto, Alfonso Ortega Giménez, Iván López-Espejo, Eduardo Lleida |
INTERSPEECH | 2 |
| 2020 | Optimization of the area under the ROC curve using neural network supervectors for text-dependent speaker verification
Victoria Mingote, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida |
Comput. Speech Lang. | 3 |
| 2019 | Speech Enhancement with Wide Residual Networks in Reverberant EnvironmentsabstractThis paper proposes a speech enhancement method which exploits the high potential of residual connections in a Wide Residual Network architecture. This is supported on single dimensional convolutions computed alongside the time domain, which is a powerful approach to process contextually correlated representations through the temporal domain, such as speech feature sequences. We find the residual mechanism extremely useful for the enhancement task since the signal always has a linear shortcut and the non-linear path enhances it in several steps by adding or subtracting corrections. The enhancement capability of the proposal is assessed by objective quality metrics evaluated with simulated and real samples of reverberated speech signals. Results show that the proposal outperforms the state-of-the-art method called WPE, which is known to effectively reduce reverberation and greatly enhance the signal. The proposed model, trained with artificial synthesized reverberation data, was able to generalize to real room impulse responses for a variety of conditions (e.g. different room sizes, $RT_{60}$, near & far field). Furthermore, it achieves accuracy for real speech with reverberation from two different datasets. Jorge Llombart, Dayana Ribas González, Antonio Miguel, Luis Vicente, Alfonso Ortega Giménez, Eduardo Lleida |
INTERSPEECH | 5 |
| 2019 | Progressive Speech Enhancement with Residual ConnectionsabstractThis paper studies the Speech Enhancement based on Deep Neural Networks. The proposed architecture gradually follows the signal transformation during enhancement by means of a visualization probe at each network block. Alongside the process, the enhancement performance is visually inspected and evaluated in terms of regression cost. This progressive scheme is based on Residual Networks. During the process, we investigate a residual connection with a constant number of channels, including internal state between blocks, and adding progressive supervision. The insights provided by the interpretation of the network enhancement process leads us to design an improved architecture for the enhancement purpose. Following this strategy, we are able to obtain speech enhancement results beyond the state-of-the-art, achieving a favorable trade-off between dereverberation and the amount of spectral distortion. Jorge Llombart, Dayana Ribas González, Antonio Miguel, Luis Vicente, Alfonso Ortega Giménez, Eduardo Lleida |
INTERSPEECH | 5 |
| 2019 | Language Recognition Using Triplet Neural Networks
Victoria Mingote, Diego Castán, Mitchell McLaren, Mahesh Kumar Nandwana, Alfonso Ortega Giménez, Eduardo Lleida, Antonio Miguel |
INTERSPEECH | 5 |
| 2019 | Optimization of False Acceptance/Rejection Rates and Decision Threshold for End-to-End Text-Dependent Speaker Verification Systems
Victoria Mingote, Antonio Miguel, Dayana Ribas González, Alfonso Ortega Giménez, Eduardo Lleida |
INTERSPEECH | 4 |
| 2019 | ViVoLAB Speaker Diarization System for the DIHARD 2019 Challenge
Ignacio Viñals, Pablo Gimeno, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida |
INTERSPEECH | 3 |
| 2019 | Phonetically-Aware Embeddings, Wide Residual Networks with Time-Delay Neural Networks and Self Attention Models for the 2018 NIST Speaker Recognition Evaluation
Ignacio Viñals, Dayana Ribas González, Victoria Mingote, Jorge Llombart, Pablo Gimeno, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida |
INTERSPEECH | 7 |
| 2018 | Estimation of the Number of Speakers with Variational Bayesian PLDA in the DIHARD Diarization Challenge
Ignacio Viñals, Pablo Gimeno, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida |
INTERSPEECH | 3 |
| 2017 | Tied Hidden Factors in Neural Networks for End-to-End Speaker RecognitionabstractIn this paper we propose a method to model speaker and session variability and able to generate likelihood ratios using neural networks in an end-To-end phrase dependent speaker verification system. As in Joint Factor Analysis, the model uses tied hidden variables to model speaker and session variability and a MAP adaptation of some of the parameters of the model. In the training procedure our method jointly estimates the network parameters and the values of the speaker and channel hidden variables. This is done in a two-step backpropagation algorithm, first the network weights and factor loading matrices are updated and then the hidden variables, whose gradients are calculated by aggregating the corresponding speaker or session frames, since these hidden variables are tied. The last layer of the network is defined as a linear regression probabilistic model whose inputs are the previous layer outputs. This choice has the advantage that it produces likelihoods and additionally it can be adapted during the enrolment using MAP without the need of a gradient optimization. The decisions are made based on the ratio of the output likelihoods of two neural network models, speaker adapted and universal background model. The method was evaluated on the RSR2015 database. Antonio Miguel, Jorge Llombart, Alfonso Ortega Giménez, Eduardo Lleida |
INTERSPEECH | 3 |
| 2017 | Domain Adaptation of PLDA Models in Broadcast Diarization by Means of Unsupervised Speaker Clustering
Ignacio Viñals, Alfonso Ortega Giménez, Jesús Villalba 0001, Antonio Miguel, Eduardo Lleida |
INTERSPEECH | 2 |
| 2016 | Analysis of speech quality measures for the task of estimating the reliability of speaker verification decisions
Jesús Villalba 0001, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida |
Speech Commun. | 2 |
| 2016 | Bayesian Networks to Model the Variability of Speaker Verification Scores in Adverse EnvironmentsabstractState-of-the-art speaker recognition technology attains great performance in controlled conditions. However, when the speech segments suffer distortions like noise or reverberation performance can severely deteriorate, this fact motivated us to investigate how score distributions diverge from the ideal ones in degraded conditions. We propose a Bayesian network model that assumes that two scores exist: one observed and another one hidden. The observed score or noisy score is the one given by the speaker verification system. Meanwhile, the hidden score or clean score is the ideal score that we would obtain in a trial with high-quality speech. A set of quality measures helps to relate both scores. We applied this network to two tasks. The first one consists in rejecting unreliable trials, i.e., trials that we cannot assure whether they are target or nontarget. We prove that this method outperforms previous approaches, based on another type of Bayesian networks. The second task is to compute an improved likelihood ratio, dependent on the quality measures. This ratio improved calibration in noisy conditions. Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Variational Bayesian PLDA for speaker diarization in the MGB challengeabstractThis paper describes the ViVoLab speaker diarization system for the Multi-Genre Broadcast (MGB) Challenge at ASRU2015. The challenge data consisted of BBC TV programmes of different genres. Diarization followed a longitudinal setup, i.e., the speakers of the current episode had to be linked to the speakers in previous episodes of the same show. We propose a system based on the i-vector paradigm. After an initial segmentation step, we compute an i-vector per speech segment. Then, a generative model based on Bayesian PLDA clusters the speakers. In this model, the speaker labels are latent variables that we optimize by variational Bayes iterations. The number of speakers in each episode was decided by maximizing the variational lower bound. The system includes several phases of segment-merging and re-clustering. We re-compute i-vectors after each merging step, which reduces the i-vector uncertainty. This approach attained a DER around 30% in the development set. Jesús Villalba 0001, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida |
ASRU | 2 |
| 2015 | Spoofing detection with DNN and one-class SVM for the ASVspoof 2015 challengeabstractSpeaker verification systems have achieved great performance in recent times. However, we usually measure performance on a ideal scenarios with naive impostors that do not modify their voices to impersonate the target speakers. The fact of impersonating a legitimate user is known as spoofing attack. Recent works show the vulnerability of current speaker verification technology to several types of attacks. Most of these works use non-public databases and different performance measures, which makes difficult to compare approaches. The spoofing challenge (ASVspoof 2015) tries to overcome this problem by proposing a common evaluation framework. This paper describes our submission to the challenge. We proposed to use spectral log-filter-bank and relative phase shift features as input to classifiers based on deep neural networks (DNN). The first of our classifiers used DNN posteriors to decide if the trial is spoof or non-spoof. The second used a bottleneck feature from the DNN as input to a one-class SVM. The one-class SVM models the distribution of legitimate speech, not needing spoofing data for training. We fused the score of the different classifiers to produce our final submission. Our system attained very competitive results with EER<0.05% in 9 out of 10 spoofing types. Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida |
INTERSPEECH | 3 |
| 2014 | Factor analysis with sampling methods for text dependent speaker recognition
Antonio Miguel, Jesús Villalba 0001, Alfonso Ortega Giménez, Eduardo Lleida, Carlos Vaquero |
INTERSPEECH | 3 |
| 2014 | Low bit rate compression methods of feature vectors for distributed speech recognition
José Enrique García Laínez, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida |
Speech Commun. | 2 |
| 2013 | Segmentation-by-classification system based on factor analysisabstractThis paper proposes a novel audio segmentation-by-classification system based on Factor Analysis (FA) with a channel compensation matrix for each class and scoring the fixed-length segments as the log-likelihood ratio between class/no-class. The scores are smoothed and the most probable sequence is computed with a Viterbi algorithm. The system described here is designed to segment and classify the audio files coming from broadcast programs into five different classes: speech (SP), speech with noise (SN), speech with music (SM), music (MU) or others (OT). This task was proposed in the Albayzin 2010 evaluation campaign. The system is compared with the winning system of the evaluation achieving lower error rate in SP and SN. These classes represent 3/4 of the total amount of the data. Therefore, the FA segmentation system gets a reduction in the average segmentation error rate. Diego Castán, Alfonso Ortega Giménez, Jesús Villalba 0001, Antonio Miguel, Eduardo Lleida |
ICASSP | 2 |
| 2013 | Prosodic features and formant modeling for an ivector-based language recognition systemabstractThe prosody of a language is encoded in syllable length, loudness and pitch. These attributes make humans perceive rhythm, stress and intonation in speech. Depending on the language, these speech properties vary, making language classification possible. On the other hand, formants are the resonance frequencies of the vocal tract, depend heavily on the position adopted by the articulatory organs, and are especially useful to disambiguate vowel sounds. In this paper prosodic and formant information are combined to build a generative language identification system based on Gaussian models fed with iVectors. The system is evaluated on the NIST LRE09 database and the inclusion of formant information gives about 50% relative improvement for the 30 s task over a prosodic system without it. The fusion with a state-of-the-art acoustic system based on shifted delta cepstral coefficients (SDC) shows the complementarity of both approaches. David Martínez González, Eduardo Lleida, Alfonso Ortega Giménez, Antonio Miguel |
ICASSP | 3 |
| 2013 | Suprasegmental information modelling for autism disorder spectrum and specific language impairment classification
David Martínez González, Dayana Ribas González, Eduardo Lleida, Alfonso Ortega Giménez, Antonio Miguel |
INTERSPEECH | 4 |
| 2013 | A new Bayesian network to assess the reliability of speaker verification decisionsabstractIn some situations the quality of the signals involved in a speaker verification trial is not as good as needed to take a reliable decision. In this work, we present a new method based on Bayesian networks and quality measures to estimate if the trial decision is reliable. We present experiments on the NIST SRE2010 dataset degraded with additive noise. A system well calibrated for clean speech, produces a large actual DCF on the degraded dataset. We use our method to discard the unreliable trials and achieve a dramatic improvement of the cost values. We also prove that our method outperforms previously published approaches. Jesús Villalba 0001, Eduardo Lleida, Alfonso Ortega Giménez, Antonio Miguel |
INTERSPEECH | 3 |
| 2013 | The I3a speaker recognition system for NIST SRE12: post-evaluation analysisabstractThe I3A submission for the recent NIST 2012 speaker recognition evaluation (SRE) was based on the i-vector approach with a multi-channel PLDA classifier. This PLDA is modified so that, for each i-vector, the between-class covariance depends on the type of channel where the segment was recorded (telephone,interviews,clean, noisy, etc). In this paper, we present the description of our submission and a detailed post-evaluation analysis of the results. We analyze several factors affecting performance: enrollment data selection, classifier type, scoring technique, calibration, known and unknown non-targets, target speakers included or not in development, segment duration, noise level and noise type. Some of these factor are new in this evaluation. After post-evaluation, actual costs improve by 15– 43% depending on the common condition. Jesús Villalba 0001, Eduardo Lleida, Alfonso Ortega Giménez, Antonio Miguel |
INTERSPEECH | 3 |
| 2013 | Quality Assessment for Speaker Diarization and Its Application in Speaker CharacterizationabstractThere are many applications related to speaker characterization, specially in telephone environments, where large datasets are available but not directly useful since there are two speakers involved in every recording. Even with very accurate speaker diarization systems, we can expect to find some recordings with low diarization accuracy. The use of these recordings may reduce the accuracy of any speaker characterization technology. Therefore, it is highly desirable to detect those recordings where the speakers are correctly segmented, in order to discard or process manually the remaining ones before feeding them into the application. In this work we propose a set of confidence measures to assess the quality of a hypothetical diarization output, in order to detect those recordings that are correctly segmented. We show that these confidence measures enable us to retrieve most of the desired recordings from a given dataset, discarding those recordings that degrade the overall accuracy of an application that make use of speaker characterization technologies. Carlos Vaquero, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | The BLZ Submission to the NIST 2011 LRE: Data Collection, System Development and PerformanceabstractThis paper describes the most relevant features of a collaborative multi-site submission to the NIST 2011 Language Recognition Evaluation (LRE), consisting of one primary and three contrastive systems, each fusing different combinations of 13 state-of-the-art (acoustic and phonotactic) language recognition subsystems.The collaboration focused on collecting and sharing training data for those target languages for which few development data were provided by NIST, and on defining a common development dataset to train backend and fusion parameters and select the best fusions.Official and post-key results are presented and compared, revealing that the greedy approach applied to select the best fusions provided suboptimal but very competitive performance.Several factors contributed to the high performance attained by BLZ systems, including the availability of training data for low resource target languages, the reliability of the development dataset (consisting only of data audited by NIST), the diversity of modeling approaches, features and datasets in the systems considered for fusion, and the effectiveness of the search for optimal fusions. Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel, Alberto Abad, David Martínez González, Jesús Villalba 0001, Alfonso Ortega Giménez, Eduardo Lleida |
INTERSPEECH | 9 |
| 2011 | Multi-site heterogeneous system fusions for the Albayzin 2010 Language Recognition EvaluationabstractBest language recognition performance is commonly obtained by fusing the scores of several heterogeneous systems. Regardless the fusion approach, it is assumed that different systems may contribute complementary information, either because they are developed on different datasets, or because they use different features or different modeling approaches. Most authors apply fusion as a final resource for improving performance based on an existing set of systems. Though relative performance gains decrease as larger sets of systems are considered, best performance is usually attained by fusing all the available systems, which may lead to high computational costs. In this paper, we aim to discover which technologies combine the best through fusion and to analyse the factors (data, features, modeling methodologies, etc.) that may explain such a good performance. Results are presented and discussed for a number of systems provided by the participating sites and the organizing team of the Albayzin 2010 Language Recognition Evaluation. We hope the conclusions of this work help research groups make better decisions in developing language recognition technology. Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel, David Martínez González, Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida, Alberto Abad, Oscar Koller, Isabel Trancoso, Paula Lopez-Otero, Laura Docío Fernández, Carmen García-Mateo, Rahim Saeidi, Mehdi Soufifar, Tomi Kinnunen, Torbjørn Svendsen, Pasi Fränti |
ASRU | 9 |
| 2011 | Intra-session variability compensation and a hypothesis generation and selection strategy for speaker segmentationabstractThis paper addresses the problem of speaker segmentation in two-speaker telephone conversations, using an eigenvoice based factor analysis approach. We present a set of improvements in the speaker segmentation system. First, we study two methods to compensate for intra-session variability, that is the variability present in a speaker during a single session. Secondly we propose a method to generate segmentation hypotheses that combined with a given confidence measure, enables the selection of correct hypotheses improving the overall segmentation performance. The proposed improvements are evaluated on the NIST Speaker Recognition Evaluation 2008 summed channel test condition, obtaining 28% relative improvement in terms of speaker segmentation error. Carlos Vaquero, Alfonso Ortega Giménez, Eduardo Lleida |
ICASSP | 2 |
| 2011 | Hierarchical Audio Segmentation with HMM and Factor Analysis in Broadcast News Domain
Diego Castán, Carlos Vaquero, Alfonso Ortega Giménez, David Martínez González, Jesús Villalba 0001, Eduardo Lleida |
INTERSPEECH | 3 |
| 2011 | I3A Language Recognition System for Albayzin 2010 LRE
David Martínez González, Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida |
INTERSPEECH | 4 |
| 2011 | Partitioning of Two-Speaker Conversation Datasets
Carlos Vaquero, Alfonso Ortega Giménez, Eduardo Lleida |
INTERSPEECH | 2 |
| 2011 | Bayesian Networks for Discrete Observation Distributions in Speech RecognitionabstractTraditionally, in speech recognition, the hidden Markov model state emission probability distributions are usually associated to continuous random variables, by using Gaussian mixtures. Thus, complex multimodal inter-feature dependencies are not accurately modeled by Gaussian models, since they are unimodal distributions and mixtures of Gaussians are needed in these complex cases, but this is done in a loose and inefficient way. Graphical models provide a precise and simple mechanism to model the dependencies among two or more variables. This paper proposes the use of discrete random variables as observations and graphical models to extract the internal dependence structure in the feature vectors. Therefore, speech features are quantized to a small number of levels, in order to obtain a tractable model. These quantized speech features provide a mechanism to increase the robustness against noise uncertainty. In addition, discrete random variables allow the learning of joint statistics of the observation densities. A method to estimate a graphical model with a constrained number of dependencies is shown in this paper, being a special kind of Bayesian network. Experimental results show that by using this modeling, better performance can be obtained compared to standard baseline systems. Antonio Miguel, Alfonso Ortega Giménez, Luis Buera, Eduardo Lleida |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Non-linear predictive vector quantization of feature vectors for distributed speech recognition
José Enrique García Laínez, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida |
INTERSPEECH | 2 |
| 2010 | Confidence measures for speaker segmentation and their relation to speaker verification
Carlos Vaquero, Alfonso Ortega Giménez, Jesús Villalba 0001, Antonio Miguel, Eduardo Lleida |
INTERSPEECH | 2 |
| 2010 | Unsupervised Data-Driven Feature Vector Normalization With Acoustic Model Adaptation for Robust Speech RecognitionabstractIn this paper, an unsupervised data-driven robust speech recognition approach is proposed based on a joint feature vector normalization and acoustic model adaptation. Feature vector normalization reduces the acoustic mismatch between training and testing conditions by mapping the feature vectors towards the training space. Model adaptation modifies the parameters of the acoustic models to match the test space. However, since neither is optimal, both approaches use an intermediate space between training and testing spaces to map either the feature vectors or acoustic models. The joint optimization of both approaches provides a common intermediate space with a better match between normalized feature vectors and adapted acoustic models. In this paper, feature vector normalization is based on a minimum mean square error (MMSE) criterion. A class dependent multi-environment model linear normalization (CD-MEMLIN) based on two classes (silence/speech) with a cross probability model (CD-MEMLIN-CPM) is used. CD-MEMLIN-CPM assumes that each class of clean and noisy spaces can be modeled with a Gaussian mixture model (GMM), training a linear transformation for each pair of Gaussians in an unsupervised data-driven training process. This feature vector normalization maps the recognition space feature vector to a normalized space. The acoustic model adaptation maps the training space to the normalized space by defining a set of linear transformations over an expanded HMM-state space, compensating for those degradations that the feature vector normalization is not able to model, like rotations. Experiments have been carried out with the Spanish SpeechDat Car database and Aurora 2 databases using both the standard Mel-frequency cepstral coefficient (MFCC) and advanced ETSI front-ends. Consistent improvements were reached for both corpora and front-ends. Using the standard MFCC front-end, a 92.08% average improvement on WER for Spanish SpeechDat Car and a 69.75% average improvement for clean condition evaluation of Aurora 2 was obtained, improving those results reached with ETSI advanced front-end (83.28% and 67.41%, respectively). Using the ETSI advanced front-end with the proposed solution, a 75.47% average improvement was obtained for the clean condition evaluation of Aurora 2 database. Luis Buera, Antonio Miguel, Oscar Saz-Torralba, Alfonso Ortega Giménez, Eduardo Lleida |
IEEE Trans. Speech Audio Process. | 4 |
| 2009 | Unsupervised training scheme with non-stereo data for empirical feature vector compensation
Luis Buera, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida, Richard M. Stern |
INTERSPEECH | 3 |
| 2009 | Differential vector quantization of feature vectors for distributed speech recognitionabstractDistributed speech recognition arises for solving computational limitations of mobile devices like PDAs or mobile phones. Due to bandwidth restrictions, it is necessary to develop efficient transmission techniques of acoustic features in Automatic Speech Recognition applications. This paper presents a technique for compressing acoustic feature vectors based on Differential Vector Quantization. It is a combination of Vector Quantization and Differential encoding schemes. Recognition experiments have been carried out, showing that the proposed method outperforms the ETSI standard VQ system, and classical VQ schemes for different codebook lengths and situations. With the proposed scheme, bit rates as low as 2.1 kbps can be used without decreasing the performance of the ASR system in terms of WER compared with a system without quantization. Index Terms: speech recognition, distributed systems, vector quantization José Enrique García Laínez, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida |
INTERSPEECH | 2 |
| 2009 | Local projections and support vector based feature selection in speech recognition
Antonio Miguel, Alfonso Ortega Giménez, Luis Buera, Eduardo Lleida |
INTERSPEECH | 2 |
| 2009 | Graphical models for discrete hidden Markov models in speech recognition
Antonio Miguel, Alfonso Ortega Giménez, Luis Buera, Eduardo Lleida |
INTERSPEECH | 2 |
| 2009 | Real-time live broadcast news subtitling system for Spanish
Alfonso Ortega Giménez, José Enrique García Laínez, Antonio Miguel, Eduardo Lleida |
INTERSPEECH | 1 |
| 2008 | Feature vector normalization with combined standard and throat microphones for robust ASRabstractWe propose on-line unsupervised compensation technique for robust speech recognition that combines standard and throat microphone feature vectors. The solution, called Multi-Environment Model-based LInear Normalization with Throat microphone information, MEMLINT, is an extension of MEM-LIN formulation. Hence, standard microphone noisy space and throat microphone space are modelled as GMMs and a set of linear transformations are learnt from data associated to each pair of Gaussians (one for each GMM) using training stereo data. On the other hand, to compensate some kinds of degra-dation which are not considered in MEMLINT, we propose to use jointly an on-line unsupervised acoustic model adaptation method based on rotation transformations over an expanded HMM-state space (augMented stAte space acousTic dEcoder, MATE). Some experiments with an own recorded database were carried out, showing that the proposed approach significantly outperforms the single microphone approach. Index Terms: Throat microphone, robust speech recognition, feature vector normalization. Luis Buera, Antonio Miguel, Oscar Saz-Torralba, Alfonso Ortega Giménez, Eduardo Lleida |
INTERSPEECH | 4 |
| 2008 | Capturing Local Variability for Speaker Normalization in Speech RecognitionabstractThe new model reduces the impact of local spectral and temporal variability by estimating a finite set of spectral and temporal warping factors which are applied to speech at the frame level. Optimum warping factors are obtained while decoding in a locally constrained search. The model involves augmenting the states of a standard hidden Markov model (HMM), providing an additional degree of freedom. It is argued in this paper that this represents an efficient and effective method for compensating local variability in speech which may have potential application to a broader array of speech transformations. The technique is presented in the context of existing methods for frequency warping-based speaker normalization for ASR. The new model is evaluated in clean and noisy task domains using subsets of the Aurora 2, the Spanish Speech-Dat-Car, and the TIDIGITS corpora. In addition, some experiments are performed on a Spanish language corpus collected from a population of speakers with a range of speech disorders. It has been found that, under clean or not severely degraded conditions, the new model provides improvements over the standard HMM baseline. It is argued that the framework of local warping is an effective general approach to providing more flexible models of speaker variability. Antonio Miguel, Eduardo Lleida, Richard C. Rose, Luis Buera, Oscar Saz-Torralba, Alfonso Ortega Giménez |
IEEE Trans. Speech Audio Process. | 6 |
| 2007 | Robust speech recognition with on-line unsupervised acoustic feature compensationabstractAn on-line unsupervised hybrid compensation technique is proposed to reduce the mismatch between training and testing conditions. It combines multi-environment model based linear normalization with cross-probability model based on GMMs (MEMLIN CPM) with a novel acoustic model adaptation method based on rotation transformations. Hence, a set of rotation transformations is estimated with clean and MEMLIN CPM-normalized training data by linear regression in an unsupervised process. Thus, in testing, each MEMLIN CPM normalized frame is decoded using a modified Viterbi algorithm and expanded acoustic models, which are obtained from the reference ones and the set of rotation transformations. To test the proposed solution, some experiments with Spanish SpeechDat Car database were carried out. MEMLIN CPM over standard ETSI front-end parameters reaches 83.89% of average improvement in WER, while the introduced hybrid solution goes up to 92.07%. Also, the proposed hybrid technique was tested with Aurora 2 database, obtaining an average improvement of 68.88% with clean training. Luis Buera, Antonio Miguel, Eduardo Lleida, Oscar Saz-Torralba, Alfonso Ortega Giménez |
ASRU | 5 |
| 2007 | On the jointly unsupervised feature vector normalization and acoustic model compensation for robust speech recognition
Luis Buera, Antonio Miguel, Eduardo Lleida, Oscar Saz-Torralba, Alfonso Ortega Giménez |
INTERSPEECH | 5 |
| 2007 | Evaluation of the combined use of MEMLIN and MLLR on the non-native adaptation task of hiwire project databaseabstractThis paper describes the performance of the combination of Multi-Environment Model-based LInear Normalization, MEMLIN, which provides an estimation of the uncorrupted feature vector, with Maximum Likelihood Linear Regression, MLLR, for the collected database under the auspices of the IST-EU STREP project HIWIRE. In this work the results for the nonnative adaptation task (NNA) are presented. The HIWIRE project database consist on command and control aeronautics application utterances pronounced by non-native speakers which are digitally corrupted with airplane cockpit noise. Thus, three noise conditions are defined: low, medium and high noise. In the proposed system, each MEMLIN-normalized feature vector is decoded using the MLLR-adapted acoustic models. The experiments show that an important improvement is reached combining MEMLIN and MLLR methods for all kinds of non-native speakers and noise conditions. Index Terms: Hiwire project, robust speech recognition, nonnative adaptation task. Luis Buera, Antonio Miguel, Oscar Saz-Torralba, Eduardo Lleida, Alfonso Ortega Giménez |
INTERSPEECH | 5 |
| 2007 | Cepstral Vector Normalization Based on Stereo Data for Robust Speech RecognitionabstractIn this paper, a set of feature vector normalization methods based on the minimum mean square error (MMSE) criterion and stereo data is presented. They include multi-environment model-based linear normalization (MEMLIN), polynomial MEMLIN (P-MEMLIN), multi-environment model-based histogram normalization (MEMHIN), and phoneme-dependent MEMLIN (PD-MEMLIN). Those methods model clean and noisy feature vector spaces using Gaussian mixture models (GMMs). The objective of the methods is to learn a transformation between clean and noisy feature vectors associated with each pair of clean and noisy model Gaussians. The direct approach to learn the transformation is by using stereo data; that is, noisy feature vectors and the corresponding clean feature vectors. In this paper, however, a nonstereo data based training procedure, is presented. The transformations can be modeled just like a bias vector (MEMLIN), or by using a first-order polynomial (P-MEMLIN) or a nonlinear function based on histogram equalization (MEMHIN). Further improvements are obtained by using phoneme-dependent bias vector transformation (PD-MEMLIN). In PD-MEMLIN, the clean and noisy feature vector spaces are split into several phonemes, and each of them is modeled as a GMM. Those methods achieve significant word error rate improvements over others that are based on similar targets. The experimental results using the SpeechDat Car database show an average improvement in word error rate greater than 68% in all cases compared to the baseline when using the original clean acoustic models, and up to 83% when training acoustic models on the new normalized feature space Luis Buera, Eduardo Lleida, Antonio Miguel, Alfonso Ortega Giménez, Oscar Saz-Torralba |
IEEE Trans. Speech Audio Process. | 4 |
| 2006 | Stability Control in a Two-Channel Speech Reinforcement System for VehiclesabstractThis paper presents a two-channel speech reinforcement system for cars able to improve the communication between the front and the rear passengers. One of the problems of this kind of systems is that they must operate in closed-loop, as acoustic feedback paths appear due to the short distance between loudspeakers and microphones. This feedback paths can make the system become unstable and acoustic echo control is needed in order to ensure stability. The system must perform two plant identifications for each channel. One of them is an open-loop identification and the other one is closed-loop. We propose here the use of echo suppression filters specially designed for closed-loop subsystems along with echo suppression filters for open-loop subsystems based on the optimal filtering theory. Results about the performance of the proposed system are provided Alfonso Ortega Giménez, Eduardo Lleida, Enrique Masgrau, Luis Buera, Antonio Miguel |
ICASSP (5) | 1 |
| 2006 | Time-dependent cross-probability model for multi-environment model based LInear normalization
Luis Buera, Eduardo Lleida, Juan A. Nolazco-Flores, Antonio Miguel, Alfonso Ortega Giménez |
INTERSPEECH | 5 |
| 2006 | Local transformation models for speech recognitionabstractThis paper presents a novel acoustic modeling framework that naturally extends the Hidden Markov Model (HMM) approach. The novel models reduce the errors caused by speaker variability by means of a local spectral mismatch reduction. A more complex and flexible speech production scheme can be assumed, in which the local temporal and frequency elastic deformations of the speech are captured by the model. In the new framework the states of a standard HMM, which are usually associated with temporal transitions, are expanded so that a new degree of freedom for the model is provided and it is then possible to estimate an optimum frequency warping factor at the same time as the decoder finds the best state sequence. In the local spectral warping based models the states become time-frequency related states and the number of parameters of the model is comparable to the standard HMM since they share a certain amount of parameters as it will be shown. The novel models are evaluated in the noise-free TIDIGITS corpus, which includes connected digits uttered by male, female and children. It has been found that, under speaker group (age-gender) mismatch conditions, the local frequency warping reduced Word Error Rate (WER) in mean by a 70%, using the initial models. When matched speaker group conditions were tested the error was reduced in mean in a 9.7% after reestimating the models. Index Terms: speaker variability, local frequency warping. Antonio Miguel, Eduardo Lleida, Alfons Juan-Císcar, Luis Buera, Alfonso Ortega Giménez, Oscar Saz-Torralba |
INTERSPEECH | 5 |
| 2006 | Study of time and frequency variability in pathological speech and error reduction methods for automatic speech recognitionabstractIn this work, we study the variations in the time and frequency domains inside a Spanish language corpus of speakers with nonpathological and pathological speech. We show how pathological speech has a greater variability in the duration of the words than non-pathological speech, while in the frequency domain we show that the vowels confusability increases by a 18%. The baseline experiments in Automatic Speech Recognition (ASR) with this corpus demonstrate that this variability causes a loss in the performance of ASR systems. To reduce the impact of time and frequency variability we use a recent Vocal Tract Length Normalization (VTLN) system: MATE (augMented stAte space acousTic modEl), as a way of improving the performance of ASR systems when dealing with speakers who suffer any kind of speech pathology. Experiments with MATE show a 17.04 % and 11.19 % WER reduction by using frequency and time MATE respectively. 1. Oscar Saz-Torralba, Antonio Miguel, Eduardo Lleida, Alfonso Ortega Giménez, Luis Buera |
INTERSPEECH | 4 |
| 2005 | Robust speech recognition in cars using phoneme dependent multi-environment linear normalizationabstractIn this paper a Phoneme-Dependent Multi-Environment Models based LInear feature Normalization, PD-MEMLIN, is presented. The target of this algorithm is to learn the difference between clean and noisy feature vectors associated to a pair of gaussians of the same phoneme (one for a clean model, and the other one for a noisy model), for each basic defined environment. These differences are estimated in a previous training process with stereo data. In order to compensate some of the problems of the independence assumption of the feature vectors components and the mismatch error between perfect and proposed transformations, two approaches have been proposed too: a multi-environment rotation transformation algorithm, and the use of transformed space acoustic models. Some experiments with SpeechDat Car database were carried out in order to study the behavior of the proposed techniques in a real acoustic environment. The experimental results show an average improvement of more than 77% using PD-MEMLIN, and more than 85% using transformed space acoustic models and multienvironment rotation transformation, concerning the baseline. Luis Buera, Eduardo Lleida, Antonio Miguel, Alfonso Ortega Giménez |
INTERSPEECH | 4 |
| 2005 | Augmented state space acoustic decoding for modeling local variability in speechabstractThis paper presents a decoding method for automatic speech recognition (ASR) that reduces the impact of local spectral and temporal variabilities on ASR performance. The procedure involves augmenting the standard Viterbi search for an optimum state sequence with a locally constrained search for optimum degrees of spectral warping or temporal warping applied to individual analysis frames. It is argued in the paper that this represents an efficient and effective method for compensating for local variability in speech which may have potential application to a broader array of speech transformations. The techniques are presented in the context of existing methods for frequency warping based speaker normalization and existing methods for computation of dynamic features for ASR. The modified decoding algorithms were evaluated in both clean and noisy task domains using Antonio Miguel, Eduardo Lleida, Richard C. Rose, Luis Buera, Alfonso Ortega Giménez |
INTERSPEECH | 5 |
| 2005 | Acoustic feedback cancellation in speech reinforcement systems for vehicles
Alfonso Ortega Giménez, Eduardo Lleida, Enrique Masgrau, Luis Buera, Antonio Miguel |
INTERSPEECH | 1 |
| 2005 | Speech Reinforcement System for Car Cabin CommunicationsabstractA speech reinforcement system is presented to improve communication between the front and the rear passengers in large motor vehicles. This type of communication can be difficult due to a number of factors, including distance between speakers, noise and lack of visual contact. The system described makes use of a set of microphones to pick up the speech of each passenger, then it amplifies these signals and plays them back to the cabin through the car audio loudspeaker system. The two main problems are noise amplification and electro-acoustic coupling between loudspeakers and microphones. To overcome these problems the system uses a set of acoustic echo cancellers, echo suppression filters and noise reduction stages. In this paper, the stability of a speech reinforcement system is studied. We propose a solution based on echo cancellers and residual echo suppression filters. The spectral estimation method for the power spectral density of the residual echo existing after the echo canceller is presented along with the derivation of the optimal residual echo suppression filter. Some results about the performance of the proposed system are also provided. Alfonso Ortega Giménez, Eduardo Lleida, Enrique Masgrau |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Multi-environment models based linear normalization for speech recognition in car conditionsabstractA multi-environment adaptation technique, based on minimum mean squared error estimation, is proposed. MEMLIN (multi-environment models based linear normalization) consists of a feature adaptation using stereo data and several basic defined environments. The target of this algorithm is to learn the difference between clean and noisy feature vectors associated to a pair of Gaussians (one for a clean model, and the other for a noisy model), for each basic environment. This knowledge, the associated Gaussians, the conditional probability between clean and noisy Gaussians, and the environment are the data used to compensate the mismatch between clean and noisy vectors. This algorithm obtains important improvements regarding other techniques that look for similar targets. The experimental results with the SpeechDat Car database shows an average improvement of more than 68%, concerning the baseline, over 7 different defined environments. Luis Buera, Eduardo Lleida, Antonio Miguel, Alfonso Ortega Giménez |
ICASSP (1) | 4 |
| 2004 | AV@CAR: A Spanish Multichannel Multimodal Corpus for In-Vehicle Automatic Audio-Visual Speech Recognition
Alfonso Ortega Giménez, Federico Sukno, Eduardo Lleida, Alejandro F. Frangi, Antonio Miguel, Luis Buera, Ernesto Zacur |
LREC | 1 |
| 2004 | Nonlinear distortion cancellation in OFDM systems using an adaptive LINC structureabstractThe linear amplification using nonlinear components (LINC) technique is a well-known power amplifier linearization method to reduce out-of-band interferences in a nonconstant envelope modulation system, such as the OFDM modulation. The major drawback of LINC transmitters is the inherited sensitivity to gain and phase imbalances between the two amplifier branches. In this paper a novel adaptive full-digital base band method, which corrects any gain and phase imbalances in LINC transmitters mainly due to the un-matching between two branches, is described. Its main advantage is the ability to track the input signal variations and adapt to the changes of amplifier nonlinear characteristics. Other effects are included in the analysis such as quadrature modulator and demodulator impairments. A computer simulation has been carried out to verify method functionality. Paloma García-Dúcar, Alfonso Ortega Giménez, Jesús de Mingo, Antonio Valdovinos |
PIMRC | 2 |
| 2003 | Residual echo power estimation for speech reinforcement systems in vehiclesabstractIn acoustic echo cancelation systems, some residual echo exists after the acoustic echo canceler (AEC) due to the fact that the adaptive filter does not model exactly the impulse response of the Loudspeaker-Enclosure-Microphone (LEM) path.This is specially important in feedback acoustic environments like speech reinforcement systems for cars where this residual echo can make the system become unstable. In order to suppress this residual echo remaining after the AEC, postfiltering is the most used technique.The optimal filter that ensures stability without attenuating the speech signal depends on the power spectral density (psd) of the residual echo that must be estimated. This paper presents a residual echo psd estimation method needed to obtain the optimal echo suppression filter in speech reinforcement systems for cars. Alfonso Ortega Giménez, Eduardo Lleida, Enrique Masgrau |
INTERSPEECH | 1 |
| 2002 | Cabin car communication system to improve communications inside a carabstractThis paper presents a cabin car communication system (CCCS) to improve the communication among passengers inside a car. Noise, distance between speakers and many other factors make difficult to maintain a conversation inside a car. The CCCS picks up the speech of each passenger, amplifies it, and uses the car loudspeaker system to return it into the cabin. Two problems arise when designing a CCCS; the electro-acoustic coupling between loudspeakers and microphones, and the amplification of the inside car noise. As a result of the first problem, the system may become unstable. To maintain the stability of the system, the CCCS makes use of a robust acoustic echo cancellation scheme based on system identification and a Wiener echo suppressor. Using a noise reduction system based on Wiener filtering reduces second problem, noise amplification. Experimental results showing the performance of the system in terms of acoustic echo and noise reduction and speech reinforce are presented. A system with 2-input/2-output channels has been built on a DSP board for medium size cars and minivan vehicles. Alfonso Ortega Giménez, Eduardo Lleida, Enrique Masgrau, Fernando Gallego |
ICASSP | 1 |
| 2001 | Acoustic echo control and noise reduction for cabin car communicationabstractA Cabin Car Communication System (CCCS) has the goal of improving the communication among passengers inside the car. Wind, road and engine noise, the distance between passengers and other factors make difficult the communication inside vehicles. The driver must often look away from the road and passengers move out of normal seating positions. The CCCS makes use of a set of microphones to pick up the speech and the car-audio loudspeakers to reinforce the sound level. This scenario presents a great challenger for acoustic echo control and noise reduction. Acoustic echo control must prevent the overall system from howling and becoming unstable with the additional problem that the system must always work with double talk. The noise reduction must clean the microphone signal to avoid the reinforce of the noise inside the car. In this paper, we describe a combined acoustic echo control and noise reduction algorithm suitable for cabin car communication systems. We present experimental results in terms of echo return loss enhancement, stability and maximum reinforce without howling. 1. Eduardo Lleida, Enrique Masgrau, Alfonso Ortega Giménez |
INTERSPEECH | 3 |