Alejandro Gómez Alanís

dblp:226/1967 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-9797-8974ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 RIRplay: Generation of a Replay Stereo Corpus for Voice Biometrics Anti-Spoofing
abstract
While recent efforts in countering spoofing attacks on voice biometric systems have primarily focused on detecting synthetic speech, Physical Access (PA) attacks, such as audio replay, still pose a serious and unresolved challenge. This research gap has been mainly due to the lack of new, realistic speech corpora for training and testing effective and generalizable countermeasure systems. Given the difficulty in collecting actual audio samples from this kind of attack, simulation has been proposed as an alternative to provide audio replay training data. The objective of this work is the generation of a novel simulated database, called RIRplay, that is both realistic, in the sense of reproducing the actual spoofing process, and representative of a wide variety of possible acoustic contexts. Our results show that training with the RIRplay corpus reduces the Equal Error Rate (EER) by nearly 10 percentage points on the challenging ASVspoof 2021 evaluation set, from 36.89% to 28.04%, compared to models trained on the ASVspoof 2019 corpus, demonstrating significant improvements in out-of-domain generalization.
Jose C. Sanchez-Valera, Antonio M. Peinado, Juan M. Martín-Doñas, Alejandro Gómez Alanís, Ángel M. Gómez, Massimiliano Todisco
IEEE Trans. Inf. Forensics Secur.4
2024 Promptformer: Prompted Conformer Transducer for ASR
abstract
Context cues carry information which can improve multi-turn interactions in automatic speech recognition (ASR) systems. In this paper, we introduce a novel mechanism inspired by hyper-prompting to fuse textual context with acoustic representations in the attention mechanism. Results on a test set with multi-turn interactions show that our method achieves 5.9% relative word error rate reduction (rWERR) over a strong baseline. We show that our method does not degrade in the absence of context and leads to improvements even if the model is trained without context. We further show that leveraging a pre-trained sentence-piece model for context embedding generation can outperform an external BERT model.
Sergio Duarte Torres, Arunasish Sen, Aman Rana, Lukas Drude, Alejandro Gómez Alanís, Andreas Schwarz, Leif Rädel, Volker Leutnant
ICASSP5
2023 A conformer-based classifier for variable-length utterance processing in anti-spoofing
abstract
The success achieved by conformers in Automatic Speech Recognition (ASR) leads us to their application in other domains, such as spoofing detection for automatic speaker verification (ASV), where the conformer self-attention mechanism might effectively model and detect the artifacts introduced in spoofed speech signals. Also, conformers can naturally handle the variable duration of speech utterances. However, as with transformers, the conformer performance may degrade when trained with limited data. To address this issue, we propose utilizing conformers in conjunction with self-supervised learning, specifically leveraging a pre-trained model called wav2vec 2.0, which is pre-trained using a substantial amount of bonafide data. Our experimental results demonstrate that our proposed method achieves one of the best results in the recent ASVspoof 2021 logical access (LA) and deep fake (DF) databases.
Eros Roselló, Alejandro Gómez Alanís, Ángel M. Gómez, Antonio M. Peinado
INTERSPEECH2
2021 PANACEA Cough Sound-Based Diagnosis of COVID-19 for the DiCOVA 2021 Challenge
abstract
The COVID-19 pandemic has led to the saturation of public health services worldwide.In this scenario, the early diagnosis of SARS-Cov-2 infections can help to stop or slow the spread of the virus and to manage the demand upon health services.This is especially important when resources are also being stretched by heightened demand linked to other seasonal diseases, such as the flu.In this context, the organisers of the DiCOVA 2021 challenge have collected a database with the aim of diagnosing COVID-19 through the use of coughing audio samples.This work presents the details of the automatic system for COVID-19 detection from cough recordings presented by team PANACEA.This team consists of researchers from two European academic institutions and one company: EURECOM (France), University of Granada (Spain), and Biometric Vox S.L. (Spain).We developed several systems based on established signal processing and machine learning methods.Our best system employs a Teager energy operator cepstral coefficients (TECCs) based frontend and Light gradient boosting machine (LightGBM) backend.The AUC obtained by this system on the test set is 76.31% which corresponds to a 10% improvement over the official baseline.
Madhu R. Kamble, José A. González 0001, Teresa Grau, Juan M. Espín, Lorenzo Cascioli, Alejandro Gómez Alanís, Jose Patino 0001, Roberto Font, Antonio M. Peinado, Ángel M. Gómez, Nicholas W. D. Evans, Maria A. Zuluaga, Massimiliano Todisco
Interspeech7
2021 On Joint Optimization of Automatic Speaker Verification and Anti-Spoofing in the Embedding Space
abstract
Biometric systems are exposed to spoofing attacks which may compromise their security, and voice biometrics based on automatic speaker verification (ASV), is no exception. To increase the robustness against such attacks, anti-spoofing systems have been proposed for the detection of replay, synthesis and voice conversion-based attacks. However, most proposed anti-spoofing techniques are loosely integrated with the ASV system. In this work, we develop a new integration neural network which jointly processes the embeddings extracted from ASV and anti-spoofing systems in order to detect both zero-effort impostors and spoofing attacks. Moreover, we propose a new loss function based on the minimization of the area under the expected (AUE) performance and spoofability curve (EPSC), which allows us to optimize the integration neural network on the desired operating range in which the biometric system is expected to work. To evaluate our proposals, experiments were carried out on the recent ASVspoof 2019 corpus, including both logical access (LA) and physical access (PA) scenarios. The experimental results show that our proposal clearly outperforms some well-known techniques based on the integration at the score- and embedding-level. Specifically, our proposal achieves up to 23.62% and 22.03% relative equal error rate (EER) improvement over the best performing baseline in the LA and PA scenarios, respectively, as well as relative gains of 27.62% and 29.15% on the AUE metric.
Alejandro Gómez Alanís, José A. González 0001, S. Pavankumar Dubagunta, Antonio M. Peinado, Mathew Magimai-Doss
IEEE Trans. Inf. Forensics Secur.1
2019 A Light Convolutional GRU-RNN Deep Feature Extractor for ASV Spoofing Detection
Alejandro Gómez Alanís, Antonio M. Peinado, José A. González 0001, Ángel M. Gómez
INTERSPEECH1
2019 A Gated Recurrent Convolutional Neural Network for Robust Spoofing Detection
abstract
Automatic speaker verification (ASV) systems are exposed to spoofing attacks which may compromise their security. While anti-spoofing techniques have been mainly studied for clean scenarios, it has also been shown that they perform poorly in noisy environments. In this work, we aim at improving the performance of spoofing detection for ASV in clean and noisy scenarios. To achieve this, we first propose the use of Gated Recurrent Convolutional Neural Networks (GRCNNs) as a deep feature extractor to robustly represent speech signals as utterance-level embeddings, which are later used by a back-end recognizer for the final genuine/spoofed classification. Then, to enhance the robustness of the system in noisy conditions, we propose the use of signal-to-noise masks (SNMs) as new input features to inform the anti-spoofing system about the time-frequency regions of the input spectral features that are mostly affected by noise and, hence, should be neglected when computing the embeddings. To evaluate our proposals, experiments were carried out on the clean and noisy versions of the ASVspoof 2015 corpus for detecting logical access attacks, as well as on the ASVspoof 2017 database to detect replay attacks. Additional results are provided for the ASVspoof 2019 corpus, including both logical and physical scenarios. The experimental results show that our proposal clearly outperforms some well-known methods based on classical features and other similar deep feature based systems for both clean and noisy conditions.
Alejandro Gómez Alanís, Antonio M. Peinado, José A. González 0001, Ángel M. Gómez
IEEE ACM Trans. Audio Speech Lang. Process.1
2018 A Deep Identity Representation for Noise Robust Spoofing Detection
Alejandro Gómez Alanís, Antonio M. Peinado, José A. González 0001, Ángel M. Gómez
INTERSPEECH1