VLDB 2026 Research / reviewers in the wild / expert
Ángel M. Gómez
dblp:82/2944 · also Angel M. Gomez, Angel Manuel Gomez Garcia
· DBLP profile ↗
56ranked-venue papers
15as first author
9since 2021 · last 2026
0000-0002-9995-3068ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 12 first-author · 4 since 2021Artificial intelligence and machine learning · 26 · 7 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
6 papers |
Audio and music processing · 100% | |
| Network and information security
1 paper |
Biometric security · 100% | |
| Artificial intelligence
6 papers |
Speech recognition and synthesis · 98% Video understanding and tracking · 2% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 30 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Biometric security
anti-spoofing |
1.0 | 1 | 2026 | RIRplay: Generation of a Replay Stereo Corpus for Voice Biometrics Anti-Spoofing · IEEE Trans. Inf. Forensics Secur. 2026 |
Biometric security › anti-spoofing
replay attack detection |
1.0 | 1 | 2026 | RIRplay: Generation of a Replay Stereo Corpus for Voice Biometrics Anti-Spoofing · IEEE Trans. Inf. Forensics Secur. 2026 |
Biometric security
speaker recognition |
1.0 | 1 | 2026 | RIRplay: Generation of a Replay Stereo Corpus for Voice Biometrics Anti-Spoofing · IEEE Trans. Inf. Forensics Secur. 2026 |
Bioinformatics and computational biology
protein function prediction |
0.5 | 1 | 2021 | Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function · Bioinform. 2021 |
Audio and music processing › speech enhancement
beamforming and postfiltering |
0.4 | 1 | 2020 | Online Multichannel Speech Enhancement Based on Recursive EM and DNN-Based Speech Presence Estimation · IEEE ACM Trans. Audio Speech Lang. Process. 2020 |
Audio and music processing › speech enhancement
multichannel speech enhancement |
0.4 | 1 | 2020 | Online Multichannel Speech Enhancement Based on Recursive EM and DNN-Based Speech Presence Estimation · IEEE ACM Trans. Audio Speech Lang. Process. 2020 |
Audio and music processing
speech enhancement |
0.4 | 1 | 2020 | Online Multichannel Speech Enhancement Based on Recursive EM and DNN-Based Speech Presence Estimation · IEEE ACM Trans. Audio Speech Lang. Process. 2020 |
Audio and music processing › speech enhancement
speech presence probability estimation |
0.4 | 1 | 2020 | Online Multichannel Speech Enhancement Based on Recursive EM and DNN-Based Speech Presence Estimation · IEEE ACM Trans. Audio Speech Lang. Process. 2020 |
Audio and music processing › speaker recognition
speaker verification |
0.4 | 1 | 2019 | A Gated Recurrent Convolutional Neural Network for Robust Spoofing Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2019 |
Audio and music processing › speaker recognition › speaker verification
spoofing detection |
0.4 | 1 | 2019 | A Gated Recurrent Convolutional Neural Network for Robust Spoofing Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2019 |
Audio and music processing
speech coding |
0.4 | 2 | 2016 | An Error Mitigation Technique for Erasure Channels Based on a Wavelet Representation of the Speech Excitation Signal · IEEE Trans. Multim. 2016 One-Pulse FEC Coding for Robust CELP-Coded Speech Transmission Over Erasure Channels · IEEE Trans. Multim. 2011 |
Audio and music processing
acoustic simulation |
0.3 | 1 | 2026 | RIRplay: Generation of a Replay Stereo Corpus for Voice Biometrics Anti-Spoofing · IEEE Trans. Inf. Forensics Secur. 2026 |
Audio and music processing
speech processing |
0.3 | 1 | 2026 | RIRplay: Generation of a Replay Stereo Corpus for Voice Biometrics Anti-Spoofing · IEEE Trans. Inf. Forensics Secur. 2026 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.3 | 2 | 2013 | MMSE-Based Missing-Feature Reconstruction With Temporal Modeling for Robust Speech Recognition · IEEE Trans. Speech Audio Process. 2013 Efficient MMSE Estimation and Uncertainty Processing for Multienvironment Robust Speech Recognition · IEEE Trans. Speech Audio Process. 2011 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › robust speech recognition
noise-robust speech recognition |
0.3 | 2 | 2013 | MMSE-Based Missing-Feature Reconstruction With Temporal Modeling for Robust Speech Recognition · IEEE Trans. Speech Audio Process. 2013 Efficient MMSE Estimation and Uncertainty Processing for Multienvironment Robust Speech Recognition · IEEE Trans. Speech Audio Process. 2011 |
Natural language and speech › Speech recognition and synthesis
speech synthesis |
0.3 | 1 | 2017 | Direct Speech Reconstruction From Articulatory Sensor Data by Machine Learning · IEEE ACM Trans. Audio Speech Lang. Process. 2017 |
Audio and music processing › speech coding
packet loss concealment |
0.2 | 1 | 2016 | An Error Mitigation Technique for Erasure Channels Based on a Wavelet Representation of the Speech Excitation Signal · IEEE Trans. Multim. 2016 |
Natural language and speech › Speech recognition and synthesis › speech coding
packet loss concealment |
0.2 | 2 | 2010 | A Multipulse-Based Forward Error Correction Technique for Robust CELP-Coded Speech Transmission Over Erasure Channels · IEEE Trans. Speech Audio Process. 2010 MMSE-Based Packet Loss Concealment for CELP-Coded Speech Recognition · IEEE Trans. Speech Audio Process. 2010 |
Natural language and speech › Speech recognition and synthesis
speech coding |
0.2 | 2 | 2010 | A Multipulse-Based Forward Error Correction Technique for Robust CELP-Coded Speech Transmission Over Erasure Channels · IEEE Trans. Speech Audio Process. 2010 MMSE-Based Packet Loss Concealment for CELP-Coded Speech Recognition · IEEE Trans. Speech Audio Process. 2010 |
Coding theory › error-correcting codes
forward error correction |
0.2 | 2 | 2011 | One-Pulse FEC Coding for Robust CELP-Coded Speech Transmission Over Erasure Channels · IEEE Trans. Multim. 2011 Combining Media-Specific FEC and Error Concealment for Robust Distributed Speech Recognition Over Loss-Prone Packet Channels · IEEE Trans. Multim. 2006 |
Natural language and speech › Speech recognition and synthesis
missing-feature reconstruction |
0.2 | 1 | 2013 | MMSE-Based Missing-Feature Reconstruction With Temporal Modeling for Robust Speech Recognition · IEEE Trans. Speech Audio Process. 2013 |
Bioinformatics and computational biology › protein analysis › protein bioinformatics
protein representation learning |
0.1 | 1 | 2021 | Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function · Bioinform. 2021 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › robust speech recognition
feature compensation |
0.1 | 1 | 2011 | Efficient MMSE Estimation and Uncertainty Processing for Multienvironment Robust Speech Recognition · IEEE Trans. Speech Audio Process. 2011 |
Natural language and speech › Speech recognition and synthesis › speech coding
code-excited linear prediction |
0.1 | 1 | 2010 | A Multipulse-Based Forward Error Correction Technique for Robust CELP-Coded Speech Transmission Over Erasure Channels · IEEE Trans. Speech Audio Process. 2010 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
robust speech recognition |
0.1 | 1 | 2010 | MMSE-Based Packet Loss Concealment for CELP-Coded Speech Recognition · IEEE Trans. Speech Audio Process. 2010 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
distributed speech recognition |
0.1 | 2 | 2010 | On the Ramsey Class of Interleavers for Robust Speech Recognition in Burst-Like Packet Loss · IEEE Trans. Speech Audio Process. 2007 MMSE-Based Packet Loss Concealment for CELP-Coded Speech Recognition · IEEE Trans. Speech Audio Process. 2010 |
Audio and music processing › speech coding
linear predictive coding |
0.1 | 1 | 2016 | An Error Mitigation Technique for Erasure Channels Based on a Wavelet Representation of the Speech Excitation Signal · IEEE Trans. Multim. 2016 |
Audio and music processing › speech recognition
distributed speech recognition |
0.1 | 1 | 2006 | Combining Media-Specific FEC and Error Concealment for Robust Distributed Speech Recognition Over Loss-Prone Packet Channels · IEEE Trans. Multim. 2006 |
Audio and music processing
speech recognition |
0.1 | 1 | 2006 | Combining Media-Specific FEC and Error Concealment for Robust Distributed Speech Recognition Over Loss-Prone Packet Channels · IEEE Trans. Multim. 2006 |
Computer vision › Video understanding and tracking
temporal modeling |
0.0 | 1 | 2013 | MMSE-Based Missing-Feature Reconstruction With Temporal Modeling for Robust Speech Recognition · IEEE Trans. Speech Audio Process. 2013 |
Methods — techniques the papers use, named apart from their topics
room impulse response simulation · 2.0neural network training · 2.0deep neural network · 1.1minimum mean square error estimation · 0.8gaussian mixture model · 0.7recurrent neural network · 0.6bidirectional RNN · 0.6unsupervised pretraining · 0.5multilayer perceptron · 0.5recursive expectation-maximization · 0.4maximum likelihood estimation · 0.4deep neural network mask estimation · 0.4hidden markov model · 0.4signal-to-noise masks · 0.4gated recurrent convolutional neural network · 0.4haar wavelet transform · 0.2codebook · 0.2soft-data decoding · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RIRplay: Generation of a Replay Stereo Corpus for Voice Biometrics Anti-SpoofingabstractWhile recent efforts in countering spoofing attacks on voice biometric systems have primarily focused on detecting synthetic speech, Physical Access (PA) attacks, such as audio replay, still pose a serious and unresolved challenge. This research gap has been mainly due to the lack of new, realistic speech corpora for training and testing effective and generalizable countermeasure systems. Given the difficulty in collecting actual audio samples from this kind of attack, simulation has been proposed as an alternative to provide audio replay training data. The objective of this work is the generation of a novel simulated database, called RIRplay, that is both realistic, in the sense of reproducing the actual spoofing process, and representative of a wide variety of possible acoustic contexts. Our results show that training with the RIRplay corpus reduces the Equal Error Rate (EER) by nearly 10 percentage points on the challenging ASVspoof 2021 evaluation set, from 36.89% to 28.04%, compared to models trained on the ASVspoof 2019 corpus, demonstrating significant improvements in out-of-domain generalization. Jose C. Sanchez-Valera, Antonio M. Peinado, Juan M. Martín-Doñas, Alejandro Gómez Alanís, Ángel M. Gómez, Massimiliano Todisco |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Exploring Self-supervised Embeddings and Synthetic Data Augmentation for Robust Audio Deepfake DetectionabstractThis work explores the performance of large speech self- supervised models as robust audio deepfake detectors. Despite the current trend of fine-tuning the upstream network, in this paper, we revisit the use of pre-trained models as feature extractors to adapt specialized downstream audio deepfake classifiers. The goal is to keep the general knowledge of the audio foundation model to extract discriminative features to feed up a simplified deepfake classifier. In addition, the generalization capabilities of the system are improved by augmenting the training corpora using additional synthetic data from different vocoder algorithms. This strategy is also complemented by various data augmentations covering challenging acoustic conditions. Our proposal is evaluated under different benchmark datasets for audio deepfake and anti-spoofing tasks, showing state-of-the-art performance. Furthermore, we analyze the relevant parts of the downstream classifier to achieve a robust system. Juan M. Martín-Doñas, Aitor Álvarez 0001, Eros Roselló, Ángel M. Gómez, Antonio M. Peinado |
INTERSPEECH | 4 |
| 2024 | Anti-spoofing Ensembling Model: Dynamic Weight Allocation in Ensemble Models for Improved Voice Biometrics SecurityabstractThis paper proposes an ensembling model as spoofed speech countermeasure, with a particular focus on synthetic voice. Despite the recent advances in speaker verification based on deep neural networks, this technology is still susceptible to various malicious attacks, so that some kind of countermeasures are needed. While an increasing number of anti-spoofing techniques can be found in the literature, the combination of multiple models, or ensemble models, still proves to be one of the best approaches. However, current iterations often rely on fixed weight assignments, potentially neglecting the unique strengths of each individual model. In response, we propose a novel ensembling model, an adaptive neural network-based approach that dynamically adjusts weights based on input utterances. Our experimental findings show that this approach outperforms traditional weighted score averaging techniques, showcasing its ability to adapt to diverse audio characteristics effectively. Eros Roselló, Ángel M. Gómez, Iván López-Espejo, Antonio M. Peinado, Juan M. Martín-Doñas |
INTERSPEECH | 2 |
| 2023 | A conformer-based classifier for variable-length utterance processing in anti-spoofingabstractThe success achieved by conformers in Automatic Speech Recognition (ASR) leads us to their application in other domains, such as spoofing detection for automatic speaker verification (ASV), where the conformer self-attention mechanism might effectively model and detect the artifacts introduced in spoofed speech signals. Also, conformers can naturally handle the variable duration of speech utterances. However, as with transformers, the conformer performance may degrade when trained with limited data. To address this issue, we propose utilizing conformers in conjunction with self-supervised learning, specifically leveraging a pre-trained model called wav2vec 2.0, which is pre-trained using a substantial amount of bonafide data. Our experimental results demonstrate that our proposed method achieves one of the best results in the recent ASVspoof 2021 logical access (LA) and deep fake (DF) databases. Eros Roselló, Alejandro Gómez Alanís, Ángel M. Gómez, Antonio M. Peinado |
INTERSPEECH | 3 |
| 2022 | An analysis of protein language model embeddings for fold predictionabstractThe identification of the protein fold class is a challenging problem in structural biology. Recent computational methods for fold prediction leverage deep learning techniques to extract protein fold-representative embeddings mainly using evolutionary information in the form of multiple sequence alignment (MSA) as input source. In contrast, protein language models (LM) have reshaped the field thanks to their ability to learn efficient protein representations (protein-LM embeddings) from purely sequential information in a self-supervised manner. In this paper, we analyze a framework for protein fold prediction using pre-trained protein-LM embeddings as input to several fine-tuning neural network models, which are supervisedly trained with fold labels. In particular, we compare the performance of six protein-LM embeddings: the long short-term memory-based UniRep and SeqVec, and the transformer-based ESM-1b, ESM-MSA, ProtBERT and ProtT5; as well as three neural networks: Multi-Layer Perceptron, ResCNN-BGRU (RBG) and Light-Attention (LAT). We separately evaluated the pairwise fold recognition (PFR) and direct fold classification (DFC) tasks on well-known benchmark datasets. The results indicate that the combination of transformer-based embeddings, particularly those obtained at amino acid level, with the RBG and LAT fine-tuning models performs remarkably well in both tasks. To further increase prediction accuracy, we propose several ensemble strategies for PFR and DFC, which provide a significant performance boost over the current state-of-the-art results. All this suggests that moving from traditional protein representations to protein-LM embeddings is a very promising approach to protein fold-related tasks. Amelia Villegas-Morcillo, Ángel M. Gómez, Victoria E. Sánchez |
Briefings Bioinform. | 2 |
| 2021 | PANACEA Cough Sound-Based Diagnosis of COVID-19 for the DiCOVA 2021 ChallengeabstractThe COVID-19 pandemic has led to the saturation of public health services worldwide.In this scenario, the early diagnosis of SARS-Cov-2 infections can help to stop or slow the spread of the virus and to manage the demand upon health services.This is especially important when resources are also being stretched by heightened demand linked to other seasonal diseases, such as the flu.In this context, the organisers of the DiCOVA 2021 challenge have collected a database with the aim of diagnosing COVID-19 through the use of coughing audio samples.This work presents the details of the automatic system for COVID-19 detection from cough recordings presented by team PANACEA.This team consists of researchers from two European academic institutions and one company: EURECOM (France), University of Granada (Spain), and Biometric Vox S.L. (Spain).We developed several systems based on established signal processing and machine learning methods.Our best system employs a Teager energy operator cepstral coefficients (TECCs) based frontend and Light gradient boosting machine (LightGBM) backend.The AUC obtained by this system on the test set is 76.31% which corresponds to a 10% improvement over the official baseline. Madhu R. Kamble, José A. González 0001, Teresa Grau, Juan M. Espín, Lorenzo Cascioli, Alejandro Gómez Alanís, Jose Patino 0001, Roberto Font, Antonio M. Peinado, Ángel M. Gómez, Nicholas W. D. Evans, Maria A. Zuluaga, Massimiliano Todisco |
Interspeech | 11 |
| 2021 | Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular functionabstractMOTIVATION: Protein function prediction is a difficult bioinformatics problem. Many recent methods use deep neural networks to learn complex sequence representations and predict function from these. Deep supervised models require a lot of labeled training data which are not available for this task. However, a very large amount of protein sequences without functional labels is available. RESULTS: We applied an existing deep sequence model that had been pretrained in an unsupervised setting on the supervised task of protein molecular function prediction. We found that this complex feature representation is effective for this task, outperforming hand-crafted features such as one-hot encoding of amino acids, k-mer counts, secondary structure and backbone angles. Also, it partly negates the need for complex prediction models, as a two-layer perceptron was enough to achieve competitive performance in the third Critical Assessment of Functional Annotation benchmark. We also show that combining this sequence representation with protein 3D structure information does not lead to performance improvement, hinting that 3D structure is also potentially learned during the unsupervised pretraining. AVAILABILITY AND IMPLEMENTATION: Implementations of all used models can be found at https://github.com/stamakro/GCN-for-Structure-and-Function. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Amelia Villegas-Morcillo, Stavros Makrodimitris, Roeland C. H. J. van Ham, Ángel M. Gómez, Victoria E. Sánchez, Marcel J. T. Reinders |
Bioinform. | 4 |
| 2021 | FoldHSphere: deep hyperspherical embeddings for protein fold recognitionabstractBACKGROUND: Current state-of-the-art deep learning approaches for protein fold recognition learn protein embeddings that improve prediction performance at the fold level. However, there still exists aperformance gap at the fold level and the (relatively easier) family level, suggesting that it might be possible to learn an embedding space that better represents the protein folds. RESULTS: In this paper, we propose the FoldHSphere method to learn a better fold embedding space through a two-stage training procedure. We first obtain prototype vectors for each fold class that are maximally separated in hyperspherical space. We then train a neural network by minimizing the angular large margin cosine loss to learn protein embeddings clustered around the corresponding hyperspherical fold prototypes. Our network architectures, ResCNN-GRU and ResCNN-BGRU, process the input protein sequences by applying several residual-convolutional blocks followed by a gated recurrent unit-based recurrent layer. Evaluation results on the LINDAHL dataset indicate that the use of our hyperspherical embeddings effectively bridges the performance gap at the family and fold levels. Furthermore, our FoldHSpherePro ensemble method yields an accuracy of 81.3% at the fold level, outperforming all the state-of-the-art methods. CONCLUSIONS: Our methodology is efficient in learning discriminative and fold-representative embeddings for the protein domains. The proposed hyperspherical embeddings are effective at identifying the protein fold class by pairwise comparison, even when amino acid sequence similarities are low. Amelia Villegas-Morcillo, Victoria E. Sánchez, Ángel M. Gómez |
BMC Bioinform. | 3 |
| 2021 | Protein Fold Recognition From Sequences Using Convolutional and Recurrent Neural NetworksabstractThe identification of a protein fold type from its amino acid sequence provides important insights about the protein 3D structure. In this paper, we propose a deep learning architecture that can process protein residue-level features to address the protein fold recognition task. Our neural network model combines 1D-convolutional layers with gated recurrent unit (GRU) layers. The GRU cells, as recurrent layers, cope with the processing issues associated to the highly variable protein sequence lengths and so extract a fold-related embedding of fixed size for each protein domain. These embeddings are then used to perform the pairwise fold recognition task, which is based on transferring the fold type of the most similar template structure. We compare our model with several template-based and deep learning-based methods from the state-of-the-art. The evaluation results over the well-known LINDAHL and SCOP_TEST sets, along with a proposed LINDAHL test set updated to SCOP 1.75, show that our embeddings perform significantly better than these methods, specially at the fold level. Supplementary material, which can be found on the Computer Society Digital Library at http://doi.ieeecomputersociety.org/10.1109/TCBB.2020.3012732, source code and trained models are available at http://sigmat.ugr.es/~amelia/CNN-GRU-RF+/. Amelia Villegas-Morcillo, Ángel M. Gómez, Juan Andres Morales-Cordovilla, Victoria E. Sánchez |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | Online Multichannel Speech Enhancement Based on Recursive EM and DNN-Based Speech Presence EstimationabstractThis article presents a recursive expectation-maximization algorithm for online multichannel speech enhancement. A deep neural network mask estimator is used to compute the speech presence probability, which is then improved by means of statistical spatial models of the noisy speech and noise signals. The clean speech signal is estimated using beamforming, single-channel linear postfiltering and speech presence masking. The clean speech statistics and speech presence probabilities are finally used to compute the acoustic parameters for beamforming and postfiltering by means of maximum likelihood estimation. This iterative procedure is carried out on a frame-by-frame basis. The algorithm integrates the different estimates in a common statistical framework suitable for online scenarios. Moreover, our method can successfully exploit spectral, spatial and temporal speech properties. Our proposed algorithm is tested in different noisy environments using the multichannel recordings of the CHiME-4 database. The experimental results show that our method outperforms other related state-of-the-art approaches in noise reduction performance, while allowing low-latency processing for real-time applications. Juan M. Martín-Doñas, Jesper Jensen 0001, Zheng-Hua Tan, Ángel M. Gómez, Antonio M. Peinado |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | A Light Convolutional GRU-RNN Deep Feature Extractor for ASV Spoofing Detection
Alejandro Gómez Alanís, Antonio M. Peinado, José A. González 0001, Ángel M. Gómez |
INTERSPEECH | 4 |
| 2019 | Multi-Channel Block-Online Source Extraction Based on Utterance AdaptationabstractThis paper deals with multi-channel speech recognition in scenarios with multiple speakers. Recently, the spectral characteristics of a target speaker, extracted from an adaptation utterance, have been used to guide a neural network mask estimator to focus on that speaker. In this work we present two variants of speakeraware neural networks, which exploit both spectral and spatial information to allow better discrimination between target and interfering speakers. Thus, we introduce either a spatial preprocessing prior to the mask estimation or a spatial plus spectral speaker characterization block whose output is directly fed into the neural mask estimator. The target speaker’s spectral and spatial signature is extracted from an adaptation utterance recorded at the beginning of a session. We further adapt the architecture for low-latency processing by means of block-online beamforming that recursively updates the signal statistics. Experimental results show that the additional spatial information clearly improves source extraction, in particular in the same-gender case, and that our proposal achieves state-of-the-art performance in terms of distortion reduction and recognition accuracy. Juan M. Martín-Doñas, Jens Heitkaemper, Reinhold Häb-Umbach, Ángel M. Gómez, Antonio M. Peinado |
INTERSPEECH | 4 |
| 2019 | A Gated Recurrent Convolutional Neural Network for Robust Spoofing DetectionabstractAutomatic speaker verification (ASV) systems are exposed to spoofing attacks which may compromise their security. While anti-spoofing techniques have been mainly studied for clean scenarios, it has also been shown that they perform poorly in noisy environments. In this work, we aim at improving the performance of spoofing detection for ASV in clean and noisy scenarios. To achieve this, we first propose the use of Gated Recurrent Convolutional Neural Networks (GRCNNs) as a deep feature extractor to robustly represent speech signals as utterance-level embeddings, which are later used by a back-end recognizer for the final genuine/spoofed classification. Then, to enhance the robustness of the system in noisy conditions, we propose the use of signal-to-noise masks (SNMs) as new input features to inform the anti-spoofing system about the time-frequency regions of the input spectral features that are mostly affected by noise and, hence, should be neglected when computing the embeddings. To evaluate our proposals, experiments were carried out on the clean and noisy versions of the ASVspoof 2015 corpus for detecting logical access attacks, as well as on the ASVspoof 2017 database to detect replay attacks. Additional results are provided for the ASVspoof 2019 corpus, including both logical and physical scenarios. The experimental results show that our proposal clearly outperforms some well-known methods based on classical features and other similar deep feature based systems for both clean and noisy conditions. Alejandro Gómez Alanís, Antonio M. Peinado, José A. González 0001, Ángel M. Gómez |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | End-to-end prediction of protein-protein interaction based on embedding and recurrent neural networks
Francisco Gonzalez-Lopez, Juan Andres Morales-Cordovilla, Amelia Villegas-Morcillo, Ángel M. Gómez, Victoria E. Sánchez |
BIBM | 4 |
| 2018 | A Deep Identity Representation for Noise Robust Spoofing Detection
Alejandro Gómez Alanís, Antonio M. Peinado, José A. González 0001, Ángel M. Gómez |
INTERSPEECH | 4 |
| 2018 | Speech excitation signal recovering based on a novel error mitigation scheme under erasure channel conditions
Domingo López-Oller, Nadir Benamirouche, Ángel M. Gómez, José L. Pérez-Córdoba |
Speech Commun. | 3 |
| 2018 | A Deep Learning Loss Function Based on the Perceptual Evaluation of the Speech QualityabstractThis letter proposes a perceptual metric for speech quality evaluation, which is suitable, as a loss function, for training deep learning methods. This metric, derived from the perceptual evaluation of the speech quality algorithm, is computed in a per-frame basis and from the power spectra of the reference and processed speech signal. Thus, two disturbance terms, which account for distortion once auditory masking and threshold effects are factored in, amend the mean square error (MSE) loss function by introducing perceptual criteria based on human psychoacoustics. The proposed loss function is evaluated for noisy speech enhancement with deep neural networks. Experimental results show that our metric achieves significant gains in speech quality (evaluated using an objective metric and a listening test) when compared to using MSE or other perceptual-based loss functions from the literature. Juan M. Martín-Doñas, Ángel M. Gómez, José A. González 0001, Antonio M. Peinado |
IEEE Signal Process. Lett. | 2 |
| 2017 | Dual-channel DNN-based speech enhancement for smartphonesabstractSpeech communications in real-world scenarios need high performance enhancement algorithms to address the distortions that can degrade the intelligibility and quality of the speech signal. Current portable devices usually integrate multiple microphones that can conveniently be exploited to improve the signal quality. In this paper we present a dual-microphone speech enhancement approach suitable for smartphones with primary (front) and reference (back) microphones. Our proposal is based on the use of deep neural networks which are able to obtain a non-linear mapping function between noisy and clean speech signals. We explore two different architectures: a feedforward deep neural network (DNN) with temporal context and a gated recurrent unit (GRU) recurrent neural network (RNN). The proposed system is evaluated under different acoustic conditions in close- and far-talk device positions. A comparison with other single- and dual-channel approaches shows that our proposal obtains the best performance in terms of perceptual quality. Juan M. Martín-Doñas, Ángel M. Gómez, Iván López-Espejo, Antonio M. Peinado |
MMSP | 2 |
| 2017 | Dual-channel VTS feature compensation for noise-robust speech recognition on mobile devicesabstractOne way to improve automatic speech recognition (ASR) performance on the latest mobile devices, which can be employed on a variety of noisy environments, consists of taking advantage of the small microphone arrays embedded in them. Since the performance of the classic beamforming techniques with small microphone arrays is rather limited, specific techniques are being developed to efficiently exploit this novel feature for noise‐robust ASR purposes. In this study, a novel dual‐channel minimum mean square error‐based feature compensation method relying on a vector Taylor series (VTS) expansion of a dual‐channel speech distortion model is proposed. In contrast to the single‐channel VTS approach (which can be considered as the state‐of‐the‐art for feature compensation), the authors’ technique particularly benefits from the spatial properties of speech and noise. Their proposal is assessed on a dual‐microphone smartphone (a particular case of interest) by means of the AURORA2‐2C synthetic corpus. Word recognition results, also validated with real noisy speech data, demonstrate the higher accuracy of their method by clearly outperforming minimum variance distortionless response beamforming and a single‐channel VTS feature compensation approach, especially at low signal‐to‐noise ratios. Iván López-Espejo, Antonio M. Peinado, Ángel M. Gómez, José A. González 0001 |
IET Signal Process. | 3 |
| 2017 | A statistical analysis of the kernel-based MMSE estimator with application to image reconstruction
Antonio M. Peinado, Ján Koloda, Ángel M. Gómez, Victoria E. Sánchez |
Signal Process. Image Commun. | 3 |
| 2017 | Direct Speech Reconstruction From Articulatory Sensor Data by Machine LearningabstractThis paper describes a technique that generates speech acoustics from articulator movements. Our motivation is to help people who can no longer speak following laryngectomy, a procedure that is carried out tens of thousands of times per year in the Western world. Our method for sensing articulator movement, permanent magnetic articulography, relies on small, unobtrusive magnets attached to the lips and tongue. Changes in magnetic field caused by magnet movements are sensed and form the input to a process that is trained to estimate speech acoustics. In the experiments reported here this “Direct Synthesis” technique is developed for normal speakers, with glued-on magnets, allowing us to train with parallel sensor and acoustic data. We describe three machine learning techniques for this task, based on Gaussian mixture models, deep neural networks, and recurrent neural networks (RNNs). We evaluate our techniques with objective acoustic distortion measures and subjective listening tests over spoken sentences read from novels (the CMU Arctic corpus). Our results show that the best performing technique is a bidirectional RNN (BiRNN), which employs both past and future contexts to predict the acoustics from the sensor data. BiRNNs are not suitable for synthesis in real time but fixed-lag RNNs give similar results and, because they only look a little way into the future, overcome this problem. Listening tests show that the speech produced by this method has a natural quality that preserves the identity of the speaker. Furthermore, we obtain up to 92% intelligibility on the challenging CMU Arctic material. To our knowledge, these are the best results obtained for a silent-speech system without a restricted vocabulary and with an unobtrusive device that delivers audio in close to real time. This work promises to lead to a technology that truly will give people whose larynx has been removed their voices back. José A. González 0001, Lam Aun Cheah, Ángel M. Gómez, Phil D. Green, James M. Gilbert, Stephen R. Ell, Roger K. Moore, Ed Holdsworth |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | System-compatible robustness improvement for new generation dect decoders by G.722 soft-decision decodingabstractThe ITU-T Recommendation G.722 about subband adaptive differential pulse code modulation (SB-ADPCM) is the mandatory wideband speech codec in the new generation digital enhanced cordless telephony (NG-DECT). Although in ADPCM the difference signal instead of the original signal is quantized and adaptive prediction is employed, redundancy is yet observed within the quantized samples. In this paper we apply a soft-decision speech decoding technique which exploits this redundancy in terms of a priori knowledge and the channel reliability information to NG-DECT. In that way, we propose a novel scheme in a standard-compliant fashion which improves the robustness of the decoder. The performance of our proposal is evaluated in terms of speech quality and a noticeable improvement over the standard codec and its own packet loss concealment algorithm is observed. Domingo López-Oller, Sai Han, Ángel M. Gómez, José L. Pérez-Córdoba, Tim Fingscheidt |
ICASSP | 3 |
| 2016 | An Error Mitigation Technique for Erasure Channels Based on a Wavelet Representation of the Speech Excitation SignalabstractThe importance of packet-based speech transmissions has grown since it offers cheaper and efficient communications. However, frame erasures are a common hurdle in these networks and concealment techniques are necessary to ensure a minimum quality of service. In this paper, we propose a mitigation technique focused on the reconstruction of the linear prediction coding (LPC) coefficients and the excitation signal of the lost frame by using a replacement technique. These replacements are obtained by means of a minimum mean square error estimation based on a source model of the speech parameters (LPC coefficients and the excitation signal). As this approach critically relies on the quantization and representation of the excitation signal, we explore the Haar wavelet transform as a novel approach to represent the excitation signal for error mitigation. Thus, this paper describes how optimal codebook and estimates can be computed in a Haar transformed domain. As a result, the excitation signal of a frame can be decomposed in several partitions where each one is independently reconstructed. Objective and subjective tests are conducted in order to assess the quality of the concealed speech signal resulting from our proposal. Both evaluations confirm noticeable improvements over the default mitigation method included in the two tested standard codecs, adaptive multirate, and Internet low bitrate codec. Domingo López-Oller, Ángel M. Gómez, José L. Pérez-Córdoba, Victoria E. Sánchez |
IEEE Trans. Multim. | 2 |
| 2013 | Sparse signal model for ultrasonic nondestructive evaluation of CFRP composite platesabstractSignal processing has been proven to be an useful tool to characterize damaged materials under ultrasonic nondestructive evaluation. In this work, we hypothesize that the transfer function of multilayered materials for a through-transmission configuration can be represented as a classical all-pole model with sparse coefficients. To test this hypothesis, we propose an analysis-by-synthesis scheme which, by assuming an underlying sparse digital signal model of the specimen, infers the order and extent of the model parameters corresponding to a certain impact damage level. Then, we exploit the sparse structure of the obtained digital filter for practical NDE applications, with emphasis on impact damage identification of carbon-fiber reinforced polymer plates. Nicolas Bochud, Ángel M. Gómez, Guillermo Rus-Carlborg, Antonio M. Peinado |
ICASSP | 2 |
| 2013 | Backwards-compatible error propagation recovery for the amr codec over erasure channelsabstractThis paper presents a recovery scheme for the error-propagation distortion which frequently appears after a frame erasure in CELP-based speech coders, in particular the AMR codec. The extensive use of predictive filters and parameter encoding allow a high-quality speech synthesis in these codecs, but makes them more vulnerable to frame erasures. Thus, when a frame is lost, an additional distortion appears in the subsequent frame, although that was correctly received, further degrading the speech quality. This degradation can also propagate over several frames, being even more damaging than the loss itself. This well known fact has motivated the development of techniques which prevent or mitigate the error propagation. Nevertheless, the previously proposed methods in some respect modify the transmission scheme (by including additional frames, FEC codes, etc.) making them incompatible with the original decoder. In this work, we apply a steganographic technique to embed recovery data to assist the decoder after a frame loss. This data mainly consist of resynchronization pulses and correction vectors for the excitation signal and the spectral envelope, respectively. PESQ results confirm that our proposal achieves a higher robustness against error propagation while the full backwards-compatibility with the AMR standard is retained. Ángel M. Gómez, José L. Pérez-Córdoba, Bernd Geiser |
ICASSP | 1 |
| 2013 | Speech Spectral Envelope Enhancement by HMM-Based Analysis/ResynthesisabstractWe propose a speech enhancement-by-resynthesis framework whose strength lies in a common statistical speech model that is shared by the analysis and synthesis stages. First, a spectro-temporal analysis is performed and masked spectro-temporal regions are identified using a noise model. Then, HMM synthesis is used to reconstruct the spectral envelope in masked regions in a manner which is conditioned on the reliable regions, preventing the resynthesis from regressing to the training data mean. As a demonstration we enhance noise-corrupted speech utterances from a small vocabulary corpus for which good statistical models are available. Perceptual evaluation of speech quality and log spectral distances demonstrate considerable performance improvements over baseline approaches that do not exploit strong speech knowledge. The letter is accompanied by audio examples. José L. Carmona, Jon Barker, Ángel M. Gómez, Ning Ma 0002 |
IEEE Signal Process. Lett. | 3 |
| 2013 | MMSE-Based Missing-Feature Reconstruction With Temporal Modeling for Robust Speech RecognitionabstractThis paper addresses the problem of feature compensation in the log-spectral domain by using the missing-data (MD) approach to noise robust speech recognition, that is, the log-spectral features can be either almost unaffected by noise or completely masked by it. First, a general MD framework based on minimum mean square error (MMSE) estimation is introduced which exploits the correlation across frequency bands to reconstruct the missing features. This framework allows the derivation of different MD imputation approaches and, in particular, a novel technique taking advantage of truncated Gaussian distributions is presented. While the proposed technique provides excellent results at high and medium signal-to-noise ratios (SNRs), its performance diminishes at low SNRs where very few reliable features are available. The reconstruction technique is therefore extended to exploit temporal constraints using two different approaches. In the first approach, time-frequency patches of speech containing a number of consecutive frames are modeled using a Gaussian mixture model (GMM). In the second one, the sequential structure of speech is alternatively modeled by a hidden Markov model (HMM). The proposed techniques are evaluated on Aurora-2 and Aurora-4 databases using both oracle and estimated masks. In both cases, the proposed techniques outperform the recognition performance obtained by the baseline system and other related techniques. Also, the introduction of a temporal modeling turns out to be very effective in reconstructing spectra at low SNRs. In particular, HMMs show the highest capability of accounting for time correlations and, therefore, achieve the best results. José A. González 0001, Antonio M. Peinado, Ning Ma 0002, Ángel M. Gómez, Jon Barker |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | Model-based cepstral analysis for ultrasonic non-destructive evaluation of compositesabstractThe use of model-based cepstral features has been shown as an effective characterization of damaged materials tested with ultrasonic non-destructive evaluation (NDE) techniques. In this work, we focus our study on carbon-fiber reinforced polymer plates and show that the use of signal models with physical meaning can provide a cepstral representation with a high discriminative power. First, we introduce a complete digital signal model based on a physical analysis of wave propagation inside the plate. The resulting model has several drawbacks: a high number of parameters to estimate and the difficulty of expressing it as a classical rational transfer function, which does not allow a model parameter estimation through classical least-squares signal modeling techniques. In order to overcome these problems, we propose two simplifications of the physical model also based on a mechanical analysis of the system. We carry out a set of damage recognition experiments showing that cepstra extracted from these models are more discriminative than other previously used methods such as the LPC cepstrum (all-pole model) or a simple FFT cepstrum. Borja Fuentes, José L. Carmona, Nicolas Bochud, Ángel M. Gómez, Antonio M. Peinado |
ICASSP | 4 |
| 2012 | Combining missing-data reconstruction and uncertainty decoding for robust speech recognitionabstractThis paper proposes a novel approach for noise-robust speech recognition which combines a missing-data (MD) derived spectral reconstruction technique and uncertainty decoding based on the weighted Viterbi algorithm (WVA). First, the noisy feature vectors are compensated by using a novel MD imputation technique based on the integration of truncated Gaussian pdfs. Although the proposed MD estimator has both the advantages of MD techniques and the use of cepstral features, it may still be affected by a number of uncertainty sources. In order to deal with these uncertainties, WVA-based uncertainty decoding is proposed. Our experiments on the Aurora-2 and Aurora-4 tasks show that the proposed MD estimator outperforms other MD imputation techniques. Also, we show that the combination of MD imputation with WVA provides better results than the combination with other uncertainty processing techniques such as the use of evidence pdfs for the estimated features. José A. González 0001, Antonio M. Peinado, Ángel M. Gómez, Ning Ma 0002, Jon Barker |
ICASSP | 3 |
| 2012 | Log-spectral feature reconstruction based on an occlusion model for noise robust speech recognition
José A. González 0001, Antonio M. Peinado, Ángel M. Gómez, Ning Ma 0002 |
INTERSPEECH | 3 |
| 2012 | Improving objective intelligibility prediction by combining correlation and coherence based methods with a measure based on the negative distortion ratio
Ángel M. Gómez, Belinda Schwerin, Kuldip K. Paliwal |
Speech Commun. | 1 |
| 2011 | Robust parametrization for non-destructive evaluation of composites using ultrasonic signalsabstractAnticipating and characterizing damages in layered carbon fiber-reinforced polymers is a challenging problem. Non-destructive evaluation using ultrasonic signals is a well-established method to obtain physically relevant parameters to characterize damages in isotropic homogeneous materials. However, ultrasonic signals obtained from composites require special care in signal interpretation due to their structural complexity. In this paper, some enhancements on the interpretation are done by adapting classical parametrization techniques to extract relevant features from the ultrasonic signals. Thus, a cepstral-based feature extractor is firstly designed and optimized by using a classification system based on cepstral distances. Then, this feature extractor is applied in an analysis-by-synthesis scheme which, by using a numerical model of the specimen, infers the values of the damage parameters. Nicolas Bochud, Ángel M. Gómez, Guillermo Rus-Carlborg, José L. Carmona, Antonio M. Peinado |
ICASSP | 2 |
| 2011 | Objective Intelligibility Prediction of Speech by Combining Correlation and Distortion Based TechniquesabstractA number of techniques based on correlation measurements have recently been proposed to provide an objective measure of intelligibility. These techniques are able to detect nonlinear distortions and provide intelligibility scores highly correlated with those given by human listeners. However, the performance of these techniques has not been found satisfactory for measuring the speech intelligibility of speech enhancement algorithms. In this paper we first investigate the different correlation-based methods, in the context of speech enhancement. We then propose to combine these correlation-based techniques with spectral distance based ones. Results presented show that objective intelligibility prediction is significantly improved by this combination. Ángel M. Gómez, Belinda Schwerin, Kuldip K. Paliwal |
INTERSPEECH | 1 |
| 2011 | Efficient MMSE Estimation and Uncertainty Processing for Multienvironment Robust Speech RecognitionabstractThis paper presents a feature compensation framework based on minimum mean square error (MMSE) estimation and stereo training data for robust speech recognition. In our proposal, we model the clean and noisy feature spaces in order to obtain clean feature estimates. However, unlike other well-known MMSE compensation methods such as SPLICE or MEMLIN, which model those spaces with Gaussian mixture models (GMMs), in our case every feature space is characterized by a set of prototype vectors which can be alternatively considered as a vector quantization (VQ) codebook. The discrete nature of this feature space characterization introduces two significative advantages. First, it allows the implementation of a very efficient MMSE estimator in terms of accuracy and computational cost. On the other hand, time correlations can be exploited by means of hidden Markov modeling (HMM). In addition, a novel subregion-based modeling is applied in order to accurately represent the transformation between the clean and noisy domains. In order to deal with unknown environments, a multiple-model approach is also explored. Since this approach has been shown quite sensitive to incorrect environment classification, we adapt two uncertainty processing techniques, soft-data decoding and exponential weighting, to our estimation framework. As a result, environment miss-classifications are concealed, allowing a better performance under unknown environments. The experimental results on noisy digit recognition show a relative improvement of 87.93% in word accuracy regarding the baseline when clean acoustic models are used, while a 4.54% is achieved with multi-style trained models. José A. González 0001, Antonio M. Peinado, Ángel M. Gómez, José L. Carmona |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | One-Pulse FEC Coding for Robust CELP-Coded Speech Transmission Over Erasure ChannelsabstractIn this paper, we present an improved quantization scheme for the redundancy data of a forward error correction (FEC) technique proposed for the transmission of code-excited linear prediction (CELP)-coded speech over erasure channels. The use of a FEC-based error protection scheme is motivated by the well-known fact that, after a frame erasure, the previous excitation is not available and a desynchronization between the encoder and the decoder long-term prediction (LTP) filters appears, causing an additional distortion which is propagated to subsequent frames. LTP synchronization can be recovered by means of a single-pulse representation of the previous excitation. No additional delay is introduced by this technique which only requires a small transmission bandwidth increase. In this paper, we focus on the efficient encoding of this pulse. Thus, an optimization procedure, which takes into account the overall synthesis error, is proposed in order to provide better pulse-position and pulse-amplitude quantization codebooks. Moreover, by extending the previous procedure, an efficient joint position-amplitude quantization can be obtained. Objective quality tests applied to our proposal show that, by means of the proposed codebooks, the number of bits required to represent the resynchronization pulse is effectively reduced. In addition, a discontinuous transmission mechanism is derived from the cost functional used during joint position-amplitude quantization, further reducing the bit-rate. Ángel M. Gómez, José L. Carmona, José A. González 0001, Victoria E. Sánchez |
IEEE Trans. Multim. | 1 |
| 2010 | Efficient VQ-based MMSE estimation for robust speech recognitionabstractThis paper presents a feature compensation technique based on the minimum mean square error (MMSE) estimation for robust speech recognition. Similarly to other MMSE compensation methods based on stereo data, our approach models the differences between clean and noisy feature spaces, and the resulting MMSE estimate of the clean feature vector is obtained as a piece-wise linear transformation of the noisy one. However, unlike other well-known MMSE techniques such as SPLICE or MEMLIN, which model the feature spaces with GMMs, in our proposal each feature space is characterized by a set of cells obtained by means of VQ quantization. This VQ-based approach allows a very efficient implementation of the MMSE estimator. Also, the possible degradation inherent to any VQ process is overcome by a strategy based on considering different subregions inside each cell and a subregion-based mean and variance compensation. The experimental results show that, along with a a very efficient MMSE estimator, our technique achieves even better recognition accuracies than SPLICE and MEMLIN. José A. González 0001, Antonio M. Peinado, Ángel M. Gómez, José L. Carmona, Juan Andres Morales-Cordovilla |
ICASSP | 3 |
| 2010 | A multipulse FEC scheme based on amplitude estimation for CELP codecs over packet networks
José L. Carmona, Ángel M. Gómez, Antonio M. Peinado, José L. Pérez-Córdoba, José A. González 0001 |
INTERSPEECH | 2 |
| 2010 | MMSE-Based Packet Loss Concealment for CELP-Coded Speech RecognitionabstractIn this paper, we analyze the performance of network speech recognition (NSR) over IP networks, adapting and proposing new solutions to the packet loss problem for code excited linear prediction (CELP) codecs. NSR has a client-server architecture which places the recognizer at the server side using a standard speech codec for speech transmission. Its main advantage is that no changes are required for the existing client devices and networks. However, the use of speech codecs degrades its performance, mainly in the presence of packet losses. First, we study the degradations introduced by CELP codecs in lossy packet networks. Later, we propose a reconstruction technique based on minimum mean square error (MMSE) estimation using hidden Markov models. This approach also allows us to obtain reliability measures associated to each estimate. We show how to use this information to improve the recognition performance by means of soft-data decoding and weighted Viterbi algorithm. The experimental results are obtained for two well-known CELP codecs, G.729 and AMR 12.2 kbps, carrying out recognition from decoded speech. Finally, we analyze an efficient and improved implementation of the proposed techniques using an NSR system which extracts speech recognition features directly from the bit-stream parameters. The experimental results show that the different proposed NSR systems achieve a comparable performance to distributed speech recognition (DSR). José L. Carmona, Antonio M. Peinado, José L. Pérez-Córdoba, Ángel M. Gómez |
IEEE Trans. Speech Audio Process. | 4 |
| 2010 | A Multipulse-Based Forward Error Correction Technique for Robust CELP-Coded Speech Transmission Over Erasure ChannelsabstractThe widely used code-excited linear prediction (CELP) paradigm relies on a strong interframe dependency which renders CELP-based codecs vulnerable to packet loss. The use of long-term prediction (LTP) or adaptive codebooks (ACB) is the main source of interframe dependency in these codecs, since they employ the excitation from previous frames. After a frame erasure, previous excitation is unavailable and a desynchronization between the encoder and the decoder appears, causing an additional distortion which is propagated to the subsequent frames. In this paper, we propose a novel media-specific Forward Error Correction (FEC) technique which retrieves LTP-resynchronization with no additional delay at the cost of a very small bit of overhead. In particular, the proposed FEC code contains a multipulse signal which replaces the excitation of the previous frame (i.e., ACB memory) when this has been lost. This multipulse description of the previous excitation is optimized to minimize the perceptual error between the synthesized speech signal and the original one. To this end, we develop a multipulse formulation which includes the additional CELP processing and, in addition, can cope with the presence of advanced LTP filters and the usual subframe segmentation applied in modern codecs. Finally, a quantization scheme is proposed to encode pulse parameters. Objective and subjective quality tests applied to our proposal show that the propagation error due to LTP filter can practically be removed with a very little bandwidth increase. Ángel M. Gómez, José L. Carmona, Antonio M. Peinado, Victoria E. Sánchez |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | A robust scheme for distributed speech recognition over loss-prone packet channels
Ángel M. Gómez, Antonio M. Peinado, Victoria E. Sánchez, José L. Carmona |
Speech Commun. | 1 |
| 2008 | A scalable coding scheme based on interframe dependency limitationabstractWhile VoIP (voice over IP) is gaining importance in comparison with other types of telephony, packet loss remains as the main source of degradation in VoIP systems. Traditional speech codecs, such as those based on the CELP (code excited linear prediction) paradigm, can achieve low bit-rates at the cost of introducing interframe dependencies. As a result, the effect of a packet loss burst is propagated to the frames correctly received after the burst. iLBC (internet low bit-rate codec) alleviates this problem by removing the interframe dependencies at the cost of a higher bit-rate. In this paper we propose a combination of iLBC with an ACELP (algebraic CELP) codec in which a variable number of ACELP-coded frames is inserted between every two iLBC-coded frames. The experimental results show that the combined codec can achieve a performance close to that of iLBC at different loss conditions but with a smaller bit-rate. Also, scalability is achieved by modifying the number of inserted ACELP-coded frames. José L. Carmona, José L. Pérez-Córdoba, Antonio M. Peinado, Ángel M. Gómez, José A. González 0001 |
ICASSP | 4 |
| 2008 | Intelligibility evaluation of Ramsey-derived interleavers for internet voice streaming with the iLBC codecabstractThis paper focuses on the application of a previously proposed interleaving, derived from the Ramsey convolutional class, in a voice streaming context with bursty packet losses. This kind of interleaving has already shown significant improvements in a distributed speech recognition context, in comparison with the widely used minimum latency block interleavers (MLBI). Here, the effectiveness of these interleavers is evaluated with an internet-oriented speech codec, such as iLBC. Since iLBC avoids error propagation due to lost frames and only uses the previously received frame to recover from those, Ramsey inter-leaving turns out especially suitable for this codec. In order to measure the performance of the system, the ITU PESQ algo-rithm is applied along with an intelligibility criterion based on the accuracy obtained through an automatic speech recognizer (ASR). In addition, an informal subjective test is carried out to corroborate the ASR scores. Results show that the proposed Ramsey-derived interleaving provides at least the same quality and better intelligibility than the MLBI ones when it is applied to iLBC codec. 1. Ángel M. Gómez, José L. Carmona, Antonio M. Peinado, Victoria E. Sánchez, José A. González 0001 |
INTERSPEECH | 1 |
| 2008 | Error concealment based on MMSE estimation for multimedia wireless and IP applicationsabstractThis paper presents a framework for the error concealment of multimedia signals transmitted over wireless and IP networks. This framework is based on two elements. First, the replacements for the erroneous channel outputs are obtained from a minimum mean square error (MMSE) estimation which takes into account both the source and all the available information from the transmission channel. Then, hidden Markov models (HMMs) provide the required stochastic modeling which allows the estimation. This modeling makes it possible a natural integration of the two types of information (source and available data from the channel). This framework is developed for both, wireless channels with errors at the bit level and packet channels degraded by packet loss. In the first case, a very high performance can be obtained even without the need of soft decision or any type of channel SNR estimation. In the second case, the lack of channel outputs during packet losses must be compensated either by enhancing the source model (increasing the model order) or by introducing media-specific FEC codes which may substitute the missing packets. This second option has yielded higher performance than the first one. Antonio M. Peinado, Ángel M. Gómez, Victoria E. Sánchez |
PIMRC | 2 |
| 2007 | iLBC-Based Transparametrization: A Real Alternative to DSR for Speech Recognition Over Packet NetworksabstractThis paper proposes a method for the remote recognition of speech coded with the iLBC codec, which is employed by a number of VoIP systems. While the usual way of performing recognition of coded speech is to decode first the speech signal and use it as input to the recognition engine, our system directly converts the iLBC parameters into recognition features. The main advantage of this approach is to avoid any type of decoding post-processing which, although originally conceived to improve the speech perception, can be harmful for a recognition system. Our method ensures the compatibility between the speech spectra provided by the iLBC codec and those employed for cepstrum computation and introduces a robust and suitable packet loss concealment strategy. Our experimental results show that the proposed system achieves a performance better than that obtained from iLBC-decoded speech and similar to that of a distributed speech recognition system over a clean or degraded transmission channel. José L. Carmona, Antonio M. Peinado, José L. Pérez-Córdoba, Ángel M. Gómez, Victoria E. Sánchez |
ICASSP (4) | 4 |
| 2007 | An Integrated Scheme for Robust Distributed Speech Recognition Over Lossy Packet NetworksabstractIn this work we present a complete set of techniques devoted to offer robustness against frame losses in distributed speech recognition over packet-switched networks. The proposed scheme is composed of tree techniques, two of them are applied at the sender and the last one in the recognizer itself. On one hand, a media-specific forward error correction (FEC) technique is used to allow the recovery of information within the bursts. On the other hand, a recognizer-based technique well known by its remarkable ability to reduce the effects of long consecutive frame losses during recognition, the weighted Viterbi algorithm (WVA), is used to handle the additional information introduced by FEC codes. Moreover, a double stream strategy whereby interleaving can be applied along with FEC codes without any delay increase, is also applied. The application of interleaving allows to reduce the perceived burst length at the receiver, further improving the recognition performance. As a result, the proposed scheme can provide an acceptable performance even under extremely adverse channel conditions. Ángel M. Gómez, Antonio M. Peinado, Victoria E. Sánchez, Antonio J. Rubio |
ICASSP (4) | 1 |
| 2007 | On the Ramsey Class of Interleavers for Robust Speech Recognition in Burst-Like Packet LossabstractThis paper focuses on the application of the Ramsey class of interleavers to the problem of achieving robust distributed speech recognition in the presence of burst-like packet loss. Although block interleavers of minimal latency have been already proposed by other authors in order to cope with this problem, the Ramsey class offers a more powerful kind of interleavers. Taking into account the behavior of the error concealment techniques commonly used in these recognition systems, it is possible to design Ramsey interleavers which, in comparison with the block interleavers of minimal latency, can counteract longer bursts with the same latency and thus improve the robustness of the recognition system. The effectiveness of the proposed design criterion is verified by means of simulations and experimental comparisons Ángel M. Gómez, Antonio M. Peinado, Victoria E. Sánchez, Antonio J. Rubio |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Interleaving and MMSE estimation with VQ replicas for distributed speech recognition over lossy packet networks
Ángel M. Gómez, Antonio M. Peinado, Victoria E. Sánchez, José L. Carmona, Antonio J. Rubio |
INTERSPEECH | 1 |
| 2006 | Multi-flow block interleaving applied to distributed speech recognition over IP networksabstractAbstract Interleaving has shown to be a useful technique to provide robustdistributed speech recognition over IP networks. This is due toits ability to disperse consecutive losses. However, this ability isrelated to the delay introduced by the interleaver. In this work,we propose a novel multi-flow block interleaver which exploitsthe presence of several streams and allows to reduce the involveddelay. Experimental results have shown that this interleaver ap-proximates the performance of end-to-end interleavers but with afraction of their delay. As disadvantage, this interleaver must beplaced in a common node where more than one flow are available. Index Terms : distributed speech recognition, IP networks, inter-leaving, active networks. 1. Introduction Since its beginning, Internet has been growing in size, incorporat-ing many new networks, as well as in functionality, adding newservices. As many other features have been integrated into In-ternet, such as mailing, instant messaging, telephony and so on,speech enabled services (SES) are also being incorporated. Theseservices provide ubiquitous speech recognition, allowing multipleusers to remotely access and share high performance recognitionengines.A very attractive approach to speech recognition over IP net-works is the distributed speech recognition (DSR) solution [1]. Asmany other services over Internet, it is based on a client-serverarchitecture. On one hand, a simple and low power client ( Ángel M. Gómez, Juan J. Ramos-Muñoz, Antonio M. Peinado, Victoria E. Sánchez |
INTERSPEECH | 1 |
| 2006 | An integrated solution for error concealment in DSR systems over wireless channels
Antonio M. Peinado, Ángel M. Gómez, Victoria E. Sánchez, José L. Pérez-Córdoba, Antonio J. Rubio |
INTERSPEECH | 2 |
| 2006 | Combining Media-Specific FEC and Error Concealment for Robust Distributed Speech Recognition Over Loss-Prone Packet ChannelsabstractThis paper presents a mixed recovery scheme for robust distributed speech recognition (DSR) implemented over a packet channel which suffers packet losses. The scheme combines media-specific forward error correction (FEC) and error concealment (EC). Media-specific FEC is applied at the client side, where FEC bits representing strongly quantized versions of the speech vectors are introduced. At the server side, the information provided by those FEC bits is used by the EC algorithm to improve the recognition performance. We investigate the adaptation of two different EC techniques, namely minimum mean square error (MMSE) estimation, which operates at the decoding stage, and weighted Viterbi recognition (WVR), where EC is applied at the recognition stage, in order to be used along with FEC. The experimental results show that a significant increase in recognition accuracy can be obtained with very little bandwidth increase, which may be null in practice, and a limited increase in latency, which in any case is not so critical for an application such as DSR Ángel M. Gómez, Antonio M. Peinado, Victoria E. Sánchez, Antonio J. Rubio |
IEEE Trans. Multim. | 1 |
| 2006 | Recognition of coded speech transmitted over wireless channelsabstractNetwork-based speech recognition (NSR) and distributed speech recognition (DSR) have been proposed as solutions to translate speech recognition technologies to mobile environments. NSR is the most straightforward solution since it does not require any modification in the mobile phone, however DSR offers higher robustness against codec compression and transmission channel degradation. This paper explores an alternative approach for remote speech recognition which combines the advantages of NSR and DSR. In this scheme, a standard speech codec is used for speech transmission but the recognition is performed from the received codec parameters. In particular, we focus on the effect of transmission channel errors, which can cause a more severe performance reduction on speech recognition than codec distortion. First, we show that an NSR solution can approach DSR through a reconstruction technique along with an adapted noise reduction technique originally proposed for acoustic noise. Then, these results are improved by working with recognition features directly extracted from the codec bitstream by means of parameter transcoding. Required modifications on current networks in order to access the bitstream are described. The network upgrading with the tandem free operation (TFO) protocol is an attractive solution. This upgrade not only offers an overall improvement on the end-to-end speech quality, but would also allow a recognition performance similar, and even higher in poor channel conditions, to that obtained by DSR when parameter transcoding along with the proposed mitigation techniques are applied Ángel M. Gómez, Antonio M. Peinado, Victoria E. Sánchez, Antonio J. Rubio |
IEEE Trans. Wirel. Commun. | 1 |
| 2005 | Packet Loss Concealment Based on VQ Replicas and MMSE Estimation Applied to Distributed Speech RecognitionabstractThis paper proposes a new packet loss concealment technique based on the inclusion in each packet of a few FEC bits, representing data replicas, combined with a minimum mean square error estimation (MMSE). This technique is developed for an Aurora-2 distributed speech recognition system working over an IP network. In addition to the data representing the transmitted speech frames, each packet includes some FEC bits representing a strongly VQ-quantized version (replicas) of previous and subsequent frames. When a loss burst occurs, the lost frames can be reconstructed from the VQ replicas. In order to mitigate the degradation introduced by the coarse VQ quantization of the replicas, a model-based MMSE estimation is applied. The experimental results show that, under a strongly degraded channel, it is possible to obtain up to 83.31 % of word accuracy with only 4 FEC bits or 88.47 % with 8 FEC bits per packet, when the Aurora mitigation algorithm only obtains 76.98 %. Antonio M. Peinado, Ángel M. Gómez, Victoria E. Sánchez, José L. Pérez-Córdoba, Antonio J. Rubio |
ICASSP (1) | 2 |
| 2005 | Joint source-channel coding of LSP parameters for bursty channels
José L. Pérez-Córdoba, Antonio M. Peinado, Ángel M. Gómez, Antonio J. Rubio |
INTERSPEECH | 3 |
| 2004 | Mitigation of channel errors in EFR-based speech recognitionabstractNetwork-based speech recognition (NSR) using the conventional speech channel with the enhanced full rate (EFR) or the adaptive multi-rate (AMR) codec is a very attractive approach since no change to existing mobile phones is needed. However, NSR reveals a degrading performance due to both transmission channel errors and the speech encoding process in comparison with distributed speech recognition (DSR), where speech features are efficiently coded and transmitted on a data channel. We focus on the degradation of the speech features caused by channel errors in an NSR system and propose methods to improve the quality of these features. Applying these methods, it turns out that the performance of an NSR system based on EFR coding is comparable to that based on DSR. Ángel M. Gómez, Antonio M. Peinado, Victoria E. Sánchez, José L. Pérez-Córdoba, Antonio J. Rubio |
ICASSP (1) | 1 |
| 2003 | A source model mitigation technique for distributed speech recognition over lossy packet channelsabstractIn this paper, we develop a new mitigation technique for a distributed speech recognition system over IP. We have designed and tested several methods to improve the interpolation used in the Aurora DSR ETSI standard without any significant increase of computational cost at the decoder. These methods make use of the information contained in the data-source, because, in IP networks, unlike in cellular networks, no information is received during packet losses. When a packet loss occurs, the lost information can be reconstructed through estimations from the N nearest received packets. Due to the enormous amount of combinations from previous and next received speech vector sequences, we have developed a methodology that drastically reduces the amount of required estimations. Ángel M. Gómez, Antonio M. Peinado, Victoria E. Sánchez, Antonio J. Rubio |
INTERSPEECH | 1 |
| 2003 | Entropy-optimized channel error mitigation with application to speech recognition over wireless
Victoria E. Sánchez, Antonio M. Peinado, Ángel M. Gómez, José L. Pérez-Córdoba |
INTERSPEECH | 3 |