Jesús Villalba 0001

dblp:211/9841 · also Jesús Antonio Villalba López · DBLP profile ↗
← Back
96ranked-venue papers
17as first author
50since 2021 · last 2026
0000-0001-9459-8426ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 84 · 15 first-author · 40 since 2021Artificial intelligence and machine learning · 70 · 14 first-author · 37 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech
abstract
There are not enough established benchmarks for the task fo speech summarization. Creating new benchmarks demands human annotation, as LLMs could embed systemic errors and bias into datasets. We test ten annotation workflows varying input modality (audio, transcript, or both) and the inclusion of editing (self or peer-editing) to investigate potential quality tradeoffs from using human annotators to summarize audio. We compare human audio-based summaries to human transcript-based summaries to track the impact of the different information modalities on summary quality. We also compare the human outputs against four LLM benchmarks (three text, one audio) to examine whether human-written summaries are less informative than highly fluent automated outputs. We find that audio-based summaries are less informative and more compressed than transcript summaries. However, iterative peer-editing with audio mitigates this difference, enabling audio-based summaries to be as informative as their transcript counterparts and LLM summaries. These findings validate iterative peer-editing among human annotators for the creation of benchmarks informed by both lexical and prosodic information. This enables crucial dataset collection even in setting where transcripts are unavailable.
Kaavya Chaparala, Thomas Thebaud, Jesús Villalba 0001, Laureano Moro-Velázquez, Peter Viechnicki, Najim Dehak
LREC3
2026 Towards Fair Speech Recognition: Mitigating Demographic Bias in End-to-End ASR Systems
Maliha Jahan, Thomas Thebaud, Zsuzsanna Fagyal, Jesús Villalba 0001, Mark Hasegawa-Johnson, Laureano Moro-Velázquez, Najim Dehak
LREC4
2025 Multi-Target Backdoor Attacks Against Speaker Recognition
abstract
In this work, we propose a multi-target backdoor attack against speaker identification using position-independent clicking sounds as triggers. Unlike previous single-target approaches, our method targets up to 50 speakers simultaneously, achieving success rates of up to 95.04%. To simulate more realistic attack conditions, we vary the signal-to-noise ratio between speech and trigger, demonstrating a trade-off between stealth and effectiveness. We further extend the attack to the speaker verification task by selecting the most similar training speaker—based on cosine similarity—as a proxy target. The attack is most effective when target and enrolled speaker pairs are highly similar, reaching success rates of up to 90% in such cases.
Alexandrine Fortier, Sonal Joshi, Thomas Thebaud, Jesús Villalba 0001, Najim Dehak, Patrick Cardinal
ASRU4
2025 The JHU-MIT System for NIST SRE24: Post-Evaluation Analysis
abstract
We present the JHU-MIT submission for NIST SRE24, along with post-evaluation analysis and key insights. In the audio fixed condition, our system used Res2Net50 and ResNet100 embeddings; the open condition additionally included an ECAPA-TDNN with a multilingual Wav2Vec2 front-end, which emerged as the best single system. The audio back-ends consisted of either PLDA adapted to SRE24 Dev or a mixture of PLDA models tuned to different subconditions. To avoid overfitting, we optimized back-end hyperparameters via twofold cross-validation. For the visual condition, we leveraged pretrained ResNet100-Subcenter-ArcFace embeddings. Agglomerative clustering was used to diarize speaker and face identities in multi-speaker videos. The primary audio fixed system achieved Act. Cp=0.574, while the open condition reached Cp=0.366 on SRE24 Eval. The visual system yielded Cp=0.169, and audiovisual fusion further improved performance, achieving Cp=0.101 (fixed) and Cp=0.087 (open).
Jesús Villalba 0001, Jonas Borgstrom, Prabhav Singh, L. Paola García-Perera, Pedro A. Torres-Carrasquillo, Najim Dehak
ASRU1
2025 Paired by the Teacher: Turning Unpaired Data into High-Fidelity Pairs for Low-Resource Text Generation
abstract
We present Paired by the Teacher (PbT), a twostage teacher-student pipeline that synthesizes accurate input-output pairs without human labels or parallel data.In many low-resource natural language generation (NLG) scenarios, practitioners may have only raw outputs, like highlights, recaps, or questions, or only raw inputs, such as articles, dialogues, or paragraphs, but seldom both.This mismatch forces small models to learn from very few examples or rely on costly, broad-scope synthetic examples produced by large LLMs.PbT addresses this by asking a teacher LLM to compress each unpaired example into a concise intermediate representation (IR), and training a student to reconstruct inputs from IRs.This enables outputs to be paired with student-generated inputs, yielding high-quality synthetic data.We evaluate PbT on five benchmarks-document summarization (XSum, CNNDM), dialogue summarization (SAMSum, DialogSum), and question generation (SQuAD)-as well as an unpaired setting on SwitchBoard (paired with Dialog-Sum summaries).An 8B student trained only on PbT data outperforms models trained on 70 B teacher-generated corpora and other unsupervised baselines, coming within 1.2 ROUGE-L of human-annotated pairs and closing 82% of the oracle gap at one-third the annotation cost of direct synthesis.Human evaluation on SwitchBoard further confirms that only PbT produces concise, faithful summaries aligned with the target style, highlighting its advantage of generating in-domain sources that avoid the mismatch, limiting direct synthesis.
Yen-Ju Lu, Thomas Thebaud, Laureano Moro-Velázquez, Najim Dehak, Jesús Villalba 0001
EMNLP5
2025 Unveiling Performance Bias in ASR Systems: A Study on Gender, Age, Accent, and More
abstract
With the recent advancements in speech recognition, it is crucial to ensure these systems are free from performance biases against any speaker subgroups. This study examined the performance of twenty variants of seven Automatic Speech Recognition models across four datasets in English language: L2 Arctic, Speech Accent Archive, CORAAL, and SBCSAE. We employed Poisson regression and drop-in-deviance tests to identify which attributes significantly contribute to the Word Error Rate. Our analysis revealed biases related to attributes such as native language, location, occupation, and birthplace. Most systems did not exhibit bias related to factors like gender and age. Additionally, we conducted an experiment to detect bias related to "variant" (accent and dialect) by combining the CORAAL (African American Vernacular English (AAVE)) and SBCSAE (General American English (GAE)) datasets, aiming to identify the sources of any observed bias. We found that both speaker variability and dialectal difference contribute to observed bias for variant.
Maliha Jahan, Priyam Mazumdar, Thomas Thebaud, Mark Hasegawa-Johnson, Jesús Villalba 0001, Najim Dehak, Laureano Moro-Velázquez
ICASSP5
2025 Detecting Neurodegenerative Diseases using Frame-Level Handwriting Embeddings
abstract
In this study, we explored the use of spectrograms to represent handwriting signals for assessing neurodegenerative diseases, including 42 healthy controls (CTL), 35 subjects with Parkinson’s Disease (PD), 21 with Alzheimer’s Disease (AD), and 15 with Parkinson’s Disease Mimics (PDM). We applied CNN and CNN-BLSTM models for binary classification using both multi-channel fixed-size and frame-based spectrograms. Our results showed that handwriting tasks and spectrogram channel combinations significantly impacted classification performance. The highest F1-score (89.8%) was achieved for AD vs. CTL, while PD vs. CTL reached 74.5%, and PD vs. PDM scored 77.97%. CNN consistently outperformed CNN-BLSTM. Different sliding window lengths were tested for constructing frame-based spectrograms. A 1-second window worked best for AD, longer windows improved PD classification, and window length had little effect on PD vs. PDM.
Sarah Laouedj, Jesús Villalba 0001, Thomas Thebaud, Laureano Moro-Velázquez, Najim Dehak
ICASSP3
2025 FaiST: A Benchmark Dataset for Fairness in Speech Technology
Maliha Jahan, Yinglun Sun, Priyam Mazumdar, Zsuzsanna Fagyal, Thomas Thebaud, Jesús Villalba 0001, Mark Hasegawa-Johnson, Najim Dehak, Laureano Moro-Velázquez
INTERSPEECH6
2025 EmoJudge: LLM Based Post-Hoc Refinement for Multimodal Speech Emotion Recognition
Prabhav Singh, Jesús Villalba 0001
INTERSPEECH2
2025 Count Your Speakers! Multitask Learning for Multimodal Speaker Diarization
Prabhav Singh, Jesús Villalba 0001, Najim Dehak
INTERSPEECH2
2025 Multimodal Emotion Diarization: Frame-Wise Integration of Text and Audio Representations
Ziv Tamir, Thomas Thebaud, Jesús Villalba 0001, Najim Dehak, Oren Kurland
INTERSPEECH3
2025 Joint Diarization and Separation Using SepFormer With Non-Autoregressive Attractors
Magdalena Rybicka, Konrad Kowalczyk, Thomas Thebaud, Najim Dehak, Jesús Villalba 0001
IEEE Signal Process. Lett.5
2024 Multimodal Emotion Recognition Harnessing the Complementarity of Speech, Language, and Vision
abstract
In the realm of audiovisual emotion recognition, a significant challenge lies in developing neural network architectures capable of effectively harnessing and integrating multimodal information. This study introduces an advanced methodology for the Empathic Virtual Agent Challenge (EVAC), utilizing state-of-the-art speech, language, and image models. Specifically, we leverage cutting-edge pre-trained models, including multilingual variants fine-tuned in French for each modality, and integrate them using late fusion techniques. Through extensive experimentation and validation, we demonstrate the efficacy of our approach in achieving competitive results on the challenge dataset. Our findings highlight that multimodal approaches outperform unimodal methods across Core Affect Presence and Intensity and Appraisal Dimensions tasks, underscoring the effectiveness of integrating diverse modalities. This underscores the importance of leveraging multiple sources of information to capture nuanced emotional states more accurately and robustly in real-world applications.
Thomas Thebaud, Anna Favaro, Yaohan Guan, Prabhav Singh, Jesús Villalba 0001, Laureano Moro-Velázquez, Najim Dehak
ICMI6
2024 Noise-robust Speech Separation with Fast Generative Correction
Helin Wang, Jesús Villalba 0001, Laureano Moro-Velázquez, Jiarui Hai, Thomas Thebaud, Najim Dehak
INTERSPEECH2
2024 Exploring the Complementary Nature of Speech and Eye Movements for Profiling Neurological Disorders
Anna Favaro, Thomas Thebaud, Jesús Villalba 0001, Najim Dehak, Laureano Moro-Velázquez
INTERSPEECH4
2024 CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
abstract
We introduce Condition-Aware Self-Supervised Learning Representation (CA-SSLR), a generalist conditioning model broadly applicable to various speech-processing tasks. Compared to standard fine-tuning methods that optimize for downstream models, CA-SSLR integrates language and speaker embeddings from earlier layers, making the SSL model aware of the current language and speaker context. This approach reduces the reliance on the input audio features while preserving the integrity of the base SSLR. CA-SSLR improves the model’s capabilities and demonstrates its generality on unseen tasks with minimal task-specific tuning. Our method employs linear modulation to dynamically adjust internal representations, enabling fine-grained adaptability without significantly altering the original model behavior. Experiments show that CA-SSLR reduces the number of trainable parameters, mitigates overfitting, and excels in under-resourced and unseen tasks. Specifically, CA-SSLR achieves a 10\% relative reduction in LID errors, a 37\% improvement in ASR CER on the ML-SUPERB benchmark, and a 27\% decrease in SV EER on VoxCeleb-1, demonstrating its effectiveness.
Yen-Ju Lu, Thomas Thebaud, Laureano Moro-Velázquez, Ariya Rastrow, Najim Dehak, Jesús Villalba 0001
NeurIPS7
2024 Clean Label Attacks Against SLU Systems
abstract
Poisoning backdoor attacks involve an adversary manipulating the training data to induce certain behaviors in the victim model by inserting a trigger in the signal at inference time. We adapted clean label backdoor (CLBD)-data poisoning attacks, which do not modify the training labels, on state-of-the-art speech recognition models that support/perform a Spoken Language Understanding task, achieving 99.8% attack success rate by poisoning 10% of the training data. We analyzed how varying the signal-strength of the poison, percent of samples poisoned, and choice of trigger impact the attack. We also found that CLBD attacks are most successful when applied to training samples that are inherently hard for a proxy model. Using this strategy, we achieved an attack success rate of 99.3% by poisoning a meager 1.5% of the training data. Finally, we applied two previously developed defenses against gradient-based attacks, and found that they attain mixed success against poisoning.
Henry Li Xinyuan, Sonal Joshi, Thomas Thebaud, Jesús Villalba 0001, Najim Dehak, Sanjeev Khudanpur
SLT4
2024 Slowness Regularized Contrastive Predictive Coding for Acoustic Unit Discovery
abstract
Self-supervised methods such as Contrastive predictive Coding (CPC) have greatly improved the quality of the unsupervised representations. These representations significantly reduce the amount of labeled data needed for downstream task performance, such as automatic speech recognition. CPC learns representations by learning to predict future frames given current frames. Based on the observation that the acoustic information, e.g., phones, changes slower than the feature extraction rate in CPC, we propose regularization techniques that impose slowness constraints on the features. Here we propose two regularization techniques: Self-expressing constraint and Left-or-Right regularization. We evaluate the proposed model on ABX and linear phone classification tasks, acoustic unit discovery, and automatic speech recognition. The regularized CPC trained on 100 hours of unlabeled data matches the performance of the baseline CPC trained on 360 hours of unlabeled data. We also show that our regularization techniques are complementary to data augmentation and can further boost the system's performance. In monolingual, cross-lingual, or multilingual settings, with/without data augmentation, regardless of the amount of data used for training, our regularized models outperformed the baseline CPC models on the ABX task.
Saurabhchand Bhati, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Time-Domain Speech Super-Resolution With GAN Based Modeling for Telephony Speaker Verification
abstract
Automatic Speaker Verification(ASV) technology has become commonplace in virtual assistants. However, its performance suffers when there is a mismatch between the train and test domains. Mixed bandwidth training, i.e., pooling training data from both domains, is a preferred choice for developing a universal model that works for both narrowband and wideband domains. We propose complementing this technique by performing neural upsampling of narrowband signals, also known as bandwidth extension. We aim to discover and analyze high-performing time-domain Generative Adversarial Network (GAN) based models to improve our downstream state-of-the-art ASV system. We choose GANs since they 1) are powerful for learning conditional distribution and 2) allow flexibleplug-inusage as a pre-processor during the training of downstream tasks (ASV) with data augmentation. Prior work mainly focused on feature-domain bandwidth extension and limited experimental setups. We address these limitations by 1) using time-domain extension models, 2) reporting results on three real test sets, 3) extending training data, and 4) devising new test-time schemes. We compare supervised (conditional GAN) and unsupervised GANs (CycleGAN) and demonstrate an average relative improvement in the equal error rate of 8.6% and 7.7%, respectively. For further analysis, we study changes in the visual quality of the spectrogram, audio perceptual quality, t-SNE embeddings, and ASV score distributions. We show that our bandwidth extension leads to phenomena such as a shift of telephone (test) embeddings towards wideband (train) signals, a negative correlation of perceptual quality with downstream performance, and condition-independent score calibration.
Saurabh Kataria 0001, Jesús Villalba 0001, Laureano Moro-Velázquez, Piotr Zelasko, Najim Dehak
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 End-to-End Neural Speaker Diarization With Non-Autoregressive Attractors
abstract
Despite many recent developments in speaker diarization, it remains a challenge and an active area of research to make diarization robust and effective in real-life scenarios. Well-established clustering-based methods are showing good performance and qualities. However, such systems are built of several independent, separately optimized modules, which may cause non-optimum performance. End-to-end neural speaker diarization (EEND) systems are considered the next stepping stone in pursuing high-performance diarization. Nevertheless, this approach also suffers limitations, such as dealing with long recordings and scenarios with a large (more than four) or unknown number of speakers in the recording. The appearance of EEND with encoder-decoder-based attractors (EEND-EDA) enabled us to deal with recordings that contain a flexible number of speakers thanks to an LSTM-based EDA module. A competitive alternative over the referenced EEND-EDA baseline is the EEND with non-autoregressive attractor (EEND-NAA) estimation, proposed recently by the authors of this article. NAA back-end incorporates k-means clustering as part of the attractor estimation and an attractor refinement module based on a Transformer decoder. However, in our previous work on EEND-NAA, we assumed a known number of speakers, and the experimental evaluation was limited to 2-speaker recordings only. In this article, we describe in detail our recent EEND-NAA approach and propose further improvements to the EEND-NAA architecture, introducing three novel variants of the NAA back-end, which can handle recordings containing speech of a variable and unknown number of speakers. Conducted experiments include simulated mixtures generated using the Switchboard and NIST SRE datasets and real-life recordings from the CALLHOME and DIHARD II datasets. In experimental evaluation, the proposed systems achieve up to 51% relative improvement for the simulated scenario and up to 15% for real recordings over the baseline EEND-EDA.
Magdalena Rybicka, Jesús Villalba 0001, Thomas Thebaud, Najim Dehak, Konrad Kowalczyk
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Model-Based Fairness Metric for Speaker Verification
abstract
Ensuring that technological advancements benefit all groups of people equally is crucial. The first step towards fairness is identifying existing inequalities. The naive comparison of group error rates may lead to wrong conclusions. We introduce a new method to determine whether a speaker verification system is fair toward several population subgroups. We propose to model miss and false alarm probabilities as a function of multiple factors, including the population group effects, e.g., male and female, and a series of confounding variables, e.g., speaker effects, language, nationality, etc. This model can estimate error rates related to a group effect without the influence of confounding effects. We experiment with a synthetic dataset where we control group and confounding effects. Our metric achieves significantly lower false positive and false negative rates w.r.t. baseline. We also experiment with VoxCeleb and NIST SRE21 datasets on different ASV systems and present our conclusions.
Maliha Jahan, Laureano Moro-Velázquez, Thomas Thebaud, Najim Dehak, Jesús Villalba 0001
ASRU5
2023 Joint Energy-Based Model for Robust Speech Classification System Against Dirty-Label Backdoor Poisoning Attacks
abstract
Our novel technique utilizes a Joint Energy-based Model (JEM) that integrates both discriminative and generative approaches to increase resistance against dirty-label backdoor attacks. Our approach is especially effective when the trigger is short or hardly perceivable. We simulate the attack on the Speech Commands Dataset consisting of 1s audio clips. During training, we use JEM to model a view of the input implemented by a randomly selected 610ms window. During inference, we combine all (40) possible views utilizing a generative part of JEM. The resulting system has slightly decreased accuracy but significantly increased resistance shown in multiple scenarios. Interestingly, replacing JEM with a standard discriminative model (Disc) provides increased resistance with a lesser effect compared to JEM but maintains accuracy. We introduce an extension motivated by semi-supervised training that further improves JEM but not Disc. JEM can also benefit from Gaussian noise during evaluation.
Martin Sustek, Sonal Joshi, Henry Li, Thomas Thebaud, Jesús Villalba 0001, Sanjeev Khudanpur, Najim Dehak
ASRU5
2023 Clustering Unsupervised Representations as Defense Against Poisoning Attacks on Speech Commands Classification System
abstract
Poisoning attacks entail attackers intentionally tampering with training data. In this paper, we consider a dirty-label poisoning attack scenario on a speech commands classification system. The threat model assumes that certain utterances from one of the classes (source class) are poisoned by superimposing a trigger on it, and its label is changed to another class selected by the attacker (target class). We propose a filtering defense against such an attack. First, we use DIstillation with NO labels (DINO) to learn unsupervised representations for all the training examples. Next, we use K-means and LDA to cluster these representations. Finally, we keep the utterances with the most repeated label in their cluster for training and discard the rest. For a 10% poisoned source class, we demonstrate a drop in attack success rate from 99.75% to 0.25%. We test our defense against a variety of threat models, including different target and source classes, as well as trigger variations.
Thomas Thebaud, Sonal Joshi, Henry Li, Martin Sustek, Jesús Villalba 0001, Sanjeev Khudanpur, Najim Dehak
ASRU5
2023 Advances in Language Recognition in Low Resource African Languages: The JHU-MIT Submission for NIST LRE22
Jesús Villalba 0001, Jonas Borgstrom, Maliha Jahan, Saurabh Kataria 0001, L. Paola García-Perera, Pedro A. Torres-Carrasquillo, Najim Dehak
INTERSPEECH1
2023 Segmental SpeechCLIP: Utilizing Pretrained Image-text Models for Audio-Visual Learning
Saurabhchand Bhati, Jesús Villalba 0001, Laureano Moro-Velázquez, Thomas Thebaud, Najim Dehak
INTERSPEECH2
2023 Do Phonatory Features Display Robustness to Characterize Parkinsonian Speech Across Corpora?
Anna Favaro, Tianyu Cao 0003, Thomas Thebaud, Jesús Villalba 0001, Ankur A. Butala, Najim Dehak, Laureano Moro-Velázquez
INTERSPEECH4
2023 Self-FiLM: Conditioning GANs with self-supervised representations for bandwidth extension based speaker recognition
abstract
Speech super-resolution/Bandwidth Extension (BWE) can improve downstream tasks like Automatic Speaker Verification (ASV).We introduce a simple novel technique called Self-FiLM to inject self-supervision into existing BWE models via Feature-wise Linear Modulation.We hypothesize that such information captures domain/environment information, which can give zero-shot generalization.Self-FiLM Conditional GAN (CGAN) gives 18% relative improvement in Equal Error Rate and 8.5% in minimum Decision Cost Function using state-ofthe-art ASV system on SRE21 test.We further by 1) deep feature loss from time-domain models and 2) re-training of data2vec 2.0 models on naturalistic wideband (VoxCeleb) and telephone data (SRE Superset etc.).Lastly, we integrate selfsupervision with CycleGAN to present a completely unsupervised solution that matches the semi-supervised performance.
Saurabh Kataria 0001, Jesús Villalba 0001, Laureano Moro-Velázquez, Thomas Thebaud, Najim Dehak
INTERSPEECH2
2023 DuTa-VC: A Duration-aware Typical-to-atypical Voice Conversion Approach with Diffusion Probabilistic Model
Helin Wang, Thomas Thebaud, Jesús Villalba 0001, Myra Sydnor, Becky Lammers, Najim Dehak, Laureano Moro-Velázquez
INTERSPEECH3
2022 Non-contrastive self-supervised learning of utterance-level speech representations
Raghavendra Pappagari, Piotr Zelasko, Laureano Moro-Velázquez, Jesús Villalba 0001, Najim Dehak
INTERSPEECH5
2022 Defense against Adversarial Attacks on Hybrid Speech Recognition System using Adversarial Fine-tuning with Denoiser
Sonal Joshi, Saurabh Kataria 0001, Yiwen Shao, Piotr Zelasko, Jesús Villalba 0001, Sanjeev Khudanpur, Najim Dehak
INTERSPEECH5
2022 AdvEst: Adversarial Perturbation Estimation to Classify and Detect Adversarial Attacks against Speaker Identification
abstract
Adversarial attacks pose a severe security threat to the state-ofthe-art speaker identification systems, thereby making it vital to propose countermeasures against them.Building on our previous work that used representation learning to classify and detect adversarial attacks, we propose an improvement to it using Ad-vEst, a method to estimate adversarial perturbation.First, we prove our claim that training the representation learning network using adversarial perturbations as opposed to adversarial examples (consisting of the combination of clean signal and adversarial perturbation) is beneficial because it eliminates nuisance information.At inference time, we use a time-domain denoiser to estimate the adversarial perturbations from adversarial examples.Using our improved representation learning approach to obtain attack embeddings (signatures), we evaluate their performance for three applications: known attack classification, attack verification, and unknown attack detection.We show that common attacks in the literature (Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD), Carlini-Wagner (CW) with different Lp threat models) can be classified with an accuracy of ∼ 96%.We also detect unknown attacks with an equal error rate (EER) of ∼9%, which is absolute improvement of ∼12% from our previous work.
Sonal Joshi, Saurabh Kataria 0001, Jesús Villalba 0001, Najim Dehak
INTERSPEECH3
2022 Joint domain adaptation and speech bandwidth extension using time-domain GANs for speaker verification
abstract
Speech systems developed for a particular choice of acoustic domain and sampling frequency do not translate easily to others.The usual practice is to learn domain adaptation and bandwidth extension models independently.Contrary to this, we propose to learn both tasks together.Particularly, we learn to map narrowband conversational telephone speech to wideband microphone speech.We developed parallel and non-parallel learning solutions which utilize both paired and unpaired data.First, we first discuss joint and disjoint training of multiple generative models for our tasks.Then, we propose a two-stage learning solution where we use a pre-trained domain adaptation system for pre-processing in bandwidth extension training.We evaluated our schemes on a Speaker Verification downstream task.We used the JHU-MIT experimental setup for NIST SRE21, which comprises SRE16, SRE-CTS Superset and SRE21.Our results provide the first evidence that learning both tasks is better than learning just one.On SRE16, our best system achieves 22% relative improvement in Equal Error Rate w.r.t. a direct learning baseline and 8% w.r.t. a strong bandwidth expansion system.
Saurabh Kataria 0001, Jesús Villalba 0001, Laureano Moro-Velázquez, Najim Dehak
INTERSPEECH2
2022 End-to-End Neural Speaker Diarization with an Iterative Refinement of Non-Autoregressive Attention-based Attractors
Magdalena Rybicka, Jesús Villalba 0001, Najim Dehak, Konrad Kowalczyk
INTERSPEECH2
2022 Chunking Defense for Adversarial Attacks on ASR
Yiwen Shao, Jesús Villalba 0001, Sonal Joshi, Saurabh Kataria 0001, Sanjeev Khudanpur, Najim Dehak
INTERSPEECH2
2022 Vsameter: Evaluation of a New Open-Source Tool to Measure Vowel Space Area and Related Metrics
abstract
Vowel space area (VSA) is an applicable metric for studying speech production deficits and intelligibility. Previous works suggest that the VSA accounts for almost 50% of the intelligibility variance, being an essential component of global intelligibility estimates. However, almost no study publishes a tool to estimate VSA automatically with publicly available codes. In this paper, we propose an open-source tool called VSAmeter to measure VSA and vowel articulation index (VAI) automatically and validate it with the VSA and VAI obtained from a dataset in which the formants and phone segments have been annotated manually. The results show that VSA and VAI values obtained by our proposed method strongly correlate with those generated by manually extracted F1 and F2 and alignments. Such a method can be utilized in speech applications, e.g., the automatic measurement of VAI for the evaluation of speakers with dysarthria.
Tianyu Cao 0003, Laureano Moro-Velázquez, Piotr Zelasko, Jesús Villalba 0001, Najim Dehak
SLT4
2022 A Multi-Modal Array of Interpretable Features to Evaluate Language and Speech Patterns in Different Neurological Disorders
abstract
Speech-based automatic approaches for evaluating neurological disorders (NDs) depend on feature extraction before the classification pipeline. It is preferable for these features to be interpretable to facilitate their development as diagnostic tools. This study focuses on the analysis of interpretable features obtained from the spoken responses of 88 subjects with NDs and controls (CN). Subjects with NDs have Alzheimer's disease (AD), Parkinson's disease (PD), or Parkinson's disease mimics (PDM). We configured three complementary sets of features related to cognition, speech, and language, and conducted a statistical analysis to examine which features differed between NDs and CN. Results suggested that features capturing response informativeness, reaction times, vocabulary richness, and syntactic complexity provided separability between AD and CN. Similarly, fundamental frequency variability helped differentiate PD from CN, while the number of salient informational units PDM from CN.
Anna Favaro, Chelsie Motley, Tianyu Cao 0003, Miguel Iglesias, Ankur A. Butala, Esther S. Oh, Robert D. Stevens 0002, Jesús Villalba 0001, Najim Dehak, Laureano Moro-Velázquez
SLT8
2022 Unsupervised Speech Segmentation and Variable Rate Representation Learning Using Segmental Contrastive Predictive Coding
abstract
Typically, unsupervised segmentation of speech into the phone- and word-like units are treated as separate tasks and are often done via different methods which do not fully leverage the inter-dependence of the two tasks. Here, we unify them and propose a technique that can jointly perform both, showing that these two tasks indeed benefit from each other. Recent attempts employ self-supervised learning, such as contrastive predictive coding (CPC), where the next frame is predicted given past context. However, CPC only looks at the audio signal’s frame-level structure. We overcome this limitation with a segmental contrastive predictive coding (SCPC) framework to model the signal structure at a higher level, e.g., phone level. A convolutional neural network learns frame-level representation from the raw waveform via noise-contrastive estimation (NCE). A differentiable boundary detector finds variable-length segments, which are then used to optimize a segment encoder via NCE to learn segment representations. The differentiable boundary detector allows us to train frame-level and segment-level encoders jointly. Experiments show that our single model outperforms existing phone and word segmentation methods on TIMIT and Buckeye datasets. We analyze the impact of the threshold on boundary detector performance, and our results suggest that automatically learning the boundary threshold can be as effective as manually tuning that threshold. We discover that phone class impacts the boundary detection performance, and the boundaries between successive vowels or semivowels are the most difficult. Finally, we use SCPC to extract speech features at the segment level rather than at the uniformly spaced frame level (e.g., 10 ms) and produce variable rate representations that change according to the contents of the utterance. We can lower the feature extraction rate from the typical 100 Hz to as low as 14.5 Hz on average while still outperforming the hand-crafted features such as MFCC on the linear phone classification task.
Saurabhchand Bhati, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Beyond Isolated Utterances: Conversational Emotion Recognition
abstract
Speech emotion recognition is the task of recognizing the speaker's emotional state given a recording of their utterance. While most of the current approaches focus on inferring emotion from isolated utterances, we argue that this is not sufficient to achieve conversational emotion recognition (CER) which deals with recognizing emotions in conversations. In this work, we propose several approaches for CER by treating it as a sequence labeling task. We investigated transformer architecture for CER and, compared it with ResNet-34 and BiLSTM architectures in both contextual and contextless scenarios using IEMOCAP corpus. Based on the inner workings of the self-attention mechanism, we proposed DiverseCatAugment (DCA), an augmentation scheme, which improved the transformer model performance by an absolute 3.3% micro-f1 on conversations and 3.6% on isolated utterances. We further enhanced the performance by introducing an interlocutor-aware transformer model where we learn a dictionary of interlocutor index embeddings to exploit diarized conversations.
Raghavendra Pappagari, Piotr Zelasko, Jesús Villalba 0001, Laureano Moro-Velázquez, Najim Dehak
ASRU3
2021 Focus on the Present: A Regularization Method for the ASR Source-Target Attention Layer
abstract
This paper introduces a novel method to diagnose the source-target attention in state-of-the-art end-to-end speech recognition models with joint connectionist temporal classification (CTC) and attention training. Our method is based on the fact that both, CTC and source-target attention, are acting on the same encoder representations. To understand the functionality of the attention, CTC is applied to compute the token posteriors given the attention outputs. We found that the source-target attention heads are able to predict several tokens ahead of the current one. Inspired by the observation, a new regularization method is proposed which leverages CTC to make source-target attention more focused on the frames corresponding to the output token being predicted by the decoder. Experiments reveal stable improvements up to 7% and 13% relatively with the proposed regularization on TED-LIUM 2 and Librispeech.
Nanxin Chen, Piotr Zelasko, Jesús Villalba 0001, Najim Dehak
ICASSP3
2021 Improving Reconstruction Loss Based Speaker Embedding in Unsupervised and Semi-Supervised Scenarios
abstract
Text-to-speech (TTS) models trained to minimize the spectrogram reconstruction loss can learn speaker embeddings without explicit speaker identity supervision, unlike x-vector speaker identification (SID) systems. Leveraging this way of speaker embedding learning can be useful in unsupervised or semi-supervised scenarios where non, or only some, of the training data have speaker labels. Thus, in this paper, we evaluate speaker embeddings learned by training the spectrogram prediction network under unsupervised and semi-supervised scenarios. We experimented with different data sampling strategies. The best one was sampling two different segments from the same utterance, namely A and B, where the spectrogram of B is predicted given the B phone sequence and the speaker embedding extracted from A. This method improved by 3.4% relative in EER, compared to using the same utterance for both A and B without segmenting. In the unsupervised scenario, the best speaker embedding outperformed i-vectors, the state-of-the-art unsupervised speaker embedding, in speaker verification by 12.9% relative in EER. We observed high correlation between reconstruction loss and speaker embedding quality. In the semi-supervised scenario, having more unlabeled data in training led to a better performance in speaker verification. Adding 5314 unlabeled speakers to 800 labeled speakers improved EER by 10.8 % relative.
Piotr Zelasko, Jesús Villalba 0001, Najim Dehak
ICASSP3
2021 Perceptual Loss Based Speech Denoising with an Ensemble of Audio Pattern Recognition and Self-Supervised Models
abstract
Deep learning based speech denoising still suffers from the challenge of improving perceptual quality of enhanced signals. We introduce a generalized framework called Perceptual Ensemble Regularization Loss (PERL) built on the idea of perceptual losses. Perceptual loss discourages distortion to certain speech properties and we analyze it using six large-scale pre-trained models: speaker classification, acoustic model, speaker embedding, emotion classification, and two self-supervised speech encoders (PASE+, wav2vec 2.0). We first build a strong baseline (w/o PERL) using Conformer Transformer Networks on the popular enhancement benchmark called VCTK-DEMAND. Using auxiliary models one at a time, we find acoustic event and self-supervised model PASE+ to be most effective. Our best model (PERL-AE) only uses acoustic event model (utilizing AudioSet) to outperform state-of-the-art methods on major perceptual metrics. To explore if denoising can leverage full framework, we use all networks but find that our seven-loss formulation suffers from the challenges of Multi-Task Learning. Finally, we report a critical observation that state-of-the-art Multi-Task weight learning methods cannot outperform hand tuning, perhaps due to challenges of domain mismatch and weak complementarity of losses.
Saurabh Kataria 0001, Jesús Villalba 0001, Najim Dehak
ICASSP2
2021 CopyPaste: An Augmentation Method for Speech Emotion Recognition
abstract
Data augmentation is a widely used strategy for training robust machine learning models. It partially alleviates the problem of limited data for tasks like speech emotion recognition (SER), where collecting data is expensive and challenging. This study proposes CopyPaste, a perceptually motivated novel augmentation procedure for SER. Assuming that the presence of emotions other than neutral dictates a speaker’s overall perceived emotion in a recording, concatenation of an emotional (emotion E) and a neutral utterance can still be labeled with emotion E. We hypothesize that SER performance can be improved using these concatenated utterances in model training. To verify this, three CopyPaste schemes are tested on two deep learning models: one trained independently and another using transfer learning from an x-vector model, a speaker recognition model. We observed that all three CopyPaste schemes improve SER performance on all the three datasets considered: MSP-Podcast, Crema-D, and IEMOCAP. Additionally, CopyPaste performs better than noise augmentation and, using them together improves the SER performance further. Our experiments on noisy test sets suggested that CopyPaste is effective even in noisy test conditions.
Raghavendra Pappagari, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak
ICASSP2
2021 Segmental Contrastive Predictive Coding for Unsupervised Word Segmentation
abstract
Automatic detection of phoneme or word-like units is one of the core objectives in zero-resource speech processing. Recent attempts employ self-supervised training methods, such as contrastive predictive coding (CPC), where the next frame is predicted given past context. However, CPC only looks at the audio signal's frame-level structure. We overcome this limitation with a segmental contrastive predictive coding (SCPC) framework that can model the signal structure at a higher level e.g. at the phoneme level. In this framework, a convolutional neural network learns frame-level representation from the raw waveform via noise-contrastive estimation (NCE). A differentiable boundary detector finds variable-length segments, which are then used to optimize a segment encoder via NCE to learn segment representations. The differentiable boundary detector allows us to train frame-level and segment-level encoders jointly. Typically, phoneme and word segmentation are treated as separate tasks. We unify them and experimentally show that our single model outperforms existing phoneme and word segmentation methods on TIMIT and Buckeye datasets. We analyze the impact of boundary threshold and when is the right time to include the segmental loss in the learning process.
Saurabhchand Bhati, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak
Interspeech2
2021 Align-Denoise: Single-Pass Non-Autoregressive Speech Recognition
Nanxin Chen, Piotr Zelasko, Laureano Moro-Velázquez, Jesús Villalba 0001, Najim Dehak
Interspeech4
2021 Deep Feature CycleGANs: Speaker Identity Preserving Non-Parallel Microphone-Telephone Domain Adaptation for Speaker Verification
abstract
With the increase in the availability of speech from varied domains, it is imperative to use such out-of-domain data to improve existing speech systems. Domain adaptation is a prominent pre-processing approach for this. We investigate it for adapt microphone speech to the telephone domain. Specifically, we explore CycleGAN-based unpaired translation of microphone data to improve the x-vector/speaker embedding network for Telephony Speaker Verification. We first demonstrate the efficacy of this on real challenging data and then, to improve further, we modify the CycleGAN formulation to make the adaptation task-specific. We modify CycleGAN's identity loss, cycle-consistency loss, and adversarial loss to operate in the deep feature space. Deep features of a signal are extracted from an auxiliary (speaker embedding) network and, hence, preserves speaker identity. Our 3D convolution-based Deep Feature Discriminators (DFD) show relative improvements of 5-10% in terms of equal error rate. To dive deeper, we study a challenging scenario of pooling (adapted) microphone and telephone data with data augmentations and telephone codecs. Finally, we highlight the sensitivity of CycleGAN hyper-parameters and introduce a parameter called probability of adaptation.
Saurabh Kataria 0001, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak
Interspeech2
2021 Automatic Detection and Assessment of Alzheimer Disease Using Speech and Language Technologies in Low-Resource Scenarios
Raghavendra Pappagari, Sonal Joshi, Laureano Moro-Velázquez, Piotr Zelasko, Jesús Villalba 0001, Najim Dehak
Interspeech6
2021 Spine2Net: SpineNet with Res2Net and Time-Squeeze-and-Excitation Blocks for Speaker Recognition
Magdalena Rybicka, Jesús Villalba 0001, Piotr Zelasko, Najim Dehak, Konrad Kowalczyk
Interspeech2
2021 Representation Learning to Classify and Detect Adversarial Attacks Against Speaker and Speech Recognition Systems
abstract
Adversarial attacks have become a major threat for machine learning applications. There is a growing interest in studying these attacks in the audio domain, e.g, speech and speaker recognition; and find defenses against them. In this work, we focus on using representation learning to classify/detect attacks w.r.t. the attack algorithm, threat model or signal-to-adversarial-noise ratio. We found that common attacks in the literature can be classified with accuracies as high as 90%. Also, representations trained to classify attacks against speaker identification can be used also to classify attacks against speaker verification and speech recognition. We also tested an attack verification task, where we need to decide whether two speech utterances contain the same attack. We observed that our models did not generalize well to attack algorithms not included in the attack representation model training. Motivated by this, we evaluated an unknown attack detection task. We were able to detect unknown attacks with equal error rates of about 19%, which is promising.
Jesús Villalba 0001, Sonal Joshi, Piotr Zelasko, Najim Dehak
Interspeech1
2021 Non-Autoregressive Transformer for Speech Recognition
abstract
Very deep transformers outperform conventional bidirectional long short-term memory networks for automatic speech recognition (ASR) by a significant margin. However, being autoregressive models, their computational complexity is still a prohibitive factor in their deployment into production systems. To amend this problem, we study two different non-autoregressive transformer structures for ASR: Audio-Conditional Masked Language Model (A-CMLM) and Audio-Factorized Masked Language Model (A-FMLM). When training these frameworks, the decoder input tokens are randomly replaced by special mask tokens. Then, the network is optimized to predict the masked tokens by taking both the unmasked context tokens and the input speech into consideration. During inference, we start from all masked tokens and the network iteratively predicts missing tokens based on partial results. A new decoding strategy is proposed as an example, which starts from the most confident predictions to the rest. Results on Mandarin (AISHELL), Japanese (CSJ), English (LibriSpeech) benchmarks show promising results to train such a non-autoregressive network for ASR. Especially in AISHELL, the proposed method outperformed the Kaldi ASR system and matched the performance of the state-of-the-art autoregressive transformer with 7× speedup.
Nanxin Chen, Shinji Watanabe 0001, Jesús Villalba 0001, Piotr Zelasko, Najim Dehak
IEEE Signal Process. Lett.3
2021 Study of Pre-Processing Defenses Against Adversarial Attacks on State-of-the-Art Speaker Recognition Systems
abstract
Adversarial examples are designed to fool the speaker recognition (SR) system by adding a carefully crafted human-imperceptible noise to the speech signals. Posing a severe security threat to state-of-the-art SR systems, it becomes vital to deep-dive and study their vulnerabilities. Moreover, it is of greater importance to propose countermeasures that can protect the systems against these attacks. Addressing these concerns, we first investigated how state-of-the-art x-vector based SR systems are affected by white-box adversarial attacks, i.e., when the adversary has full knowledge of the system. x-Vector based SR systems are evaluated against white-box adversarial attacks common in the literature like fast gradient sign method (FGSM), basic iterative method (BIM)–a.k.a. iterative-FGSM–, projected gradient descent (PGD), and Carlini-Wagner (CW) attack. To mitigate against these attacks, we investigated four pre-processing defenses which do not need adversarial examples during training. The four pre-processing defenses–viz. randomized smoothing, DefenseGAN, variational autoencoder (VAE), and Parallel Wave-GAN vocoder (PWG) are compared against the baseline defense of adversarial training. Performing powerful adaptive white-box adversarial attack (i.e., when the adversary has full knowledge of the system, including the defense), our conclusions indicate that SR systems were extremely vulnerable under BIM, PGD, and CW attacks. Among the proposed pre-processing defenses, PWG combined with randomized smoothing offers the most protection against the attacks, with accuracy averaging 93% compared to 52% in the undefended system and an absolute improvement > 90% for BIM attacks with L∞ > 0.001 and CW attack.
Sonal Joshi, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak
IEEE Trans. Inf. Forensics Secur.2
2020 Feature Enhancement with Deep Feature Losses for Speaker Verification
abstract
Speaker Verification still suffers from the challenge of generalization to novel adverse environments. We leverage on the recent advancements made by deep learning based speech enhancement and propose a feature-domain supervised denoising based solution. We propose to use Deep Feature Loss which optimizes the enhancement network in the hidden activation space of a pre-trained auxiliary speaker embedding network. We experimentally verify the approach on simulated and real data. A simulated testing setup is created using various noise types at different SNR levels. For evaluation on real data, we choose BabyTrain corpus which consists of children recordings in uncontrolled environments. We observe consistent gains in every condition over the state-of-the-art augmented Factorized-TDNN x-vector system. On BabyTrain corpus, we observe relative gains of 10.38% and 12.40% in minDCF and EER respectively.
Saurabh Kataria 0001, Phani S. Nidadavolu, Jesús Villalba 0001, Nanxin Chen, L. Paola García-Perera, Najim Dehak
ICASSP3
2020 Using X-Vectors to Automatically Detect Parkinson's Disease from Speech
abstract
The promise of new neuroprotective treatments to stop or slow the advance of Parkinson's Disease (PD) urges for new biomarkers or detection schemes that can deliver a faster diagnosis. Given that speech is affected by PD, the combination of deep neural networks and speech processing can provide automatic detection schemes. Accordingly, in this study we analyze for the first time a new state-of-the-art speaker recognition technique, x-Vectors, in a different scenario: the automatic detection of PD from speech. The proposed approach is compared with another speaker recognition technique, i-Vectors, employed in previous works and used as baseline in this study. A corpus with 43 PD patients and 46 control speakers was used to evaluate the performance of these two techniques at two sampling frequencies: 8 and 16 kHz.The x-Vector approach provided the best results in terms of accuracy and AUC reaching values of 90% and 0.94, respectively. Consequently, results suggest that speaker embeddings obtained using deep neural networks are successful extracting acoustic information relative to patterns in articulation, prosody and/or phonation common in persons with PD.
Laureano Moro-Velázquez, Jesús Villalba 0001, Najim Dehak
ICASSP2
2020 Unsupervised Feature Enhancement for Speaker Verification
abstract
The task of making speaker verification systems robust to adverse scenarios remains a challenging and an active area of research. We developed an unsupervised feature enhancement approach in log-filter bank space with the end goal of improving speaker verification performance. We experimented with using both real speech recorded in adverse environments and degraded speech obtained by simulation to train the enhancement systems. The effectiveness of this approach was shown by testing on several real, simulated noisy, and reverberant test sets. The approach yielded significant improvements on both real and simulated sets when data augmentation was not used in speaker verification pipeline. We also experimented with training the x-vector and PLDA systems with enhanced augmented features instead of augmented features and observed better performance on real test conditions (4.2% relative improvement in minDCF on SRI).
Phani S. Nidadavolu, Saurabh Kataria 0001, Jesús Villalba 0001, L. Paola García-Perera, Najim Dehak
ICASSP3
2020 X-Vectors Meet Emotions: A Study On Dependencies Between Emotion and Speaker Recognition
abstract
In this work, we explore the dependencies between speaker recognition and emotion recognition. We first show that knowledge learned for speaker recognition can be reused for emotion recognition through transfer learning. Then, we show the effect of emotion on speaker recognition. For emotion recognition, we show that using a simple linear model is enough to obtain good performance on the features extracted from pre-trained models such as the x-vector model. Then, we improve emotion recognition performance by finetuning for emotion classification. We evaluated our experiments on three different types of datasets: IEMOCAP, MSP-Podcast, and Crema-D. By fine-tuning, we obtained 30.40%, 7.99%, and 8.61% absolute improvement on IEMOCAP, MSP-Podcast, and Crema-D respectively over baseline model with no pre-training. Finally, we present results on the effect of emotion on speaker verification. We observed that speaker verification performance is prone to changes in test speaker emotions. We found that trials with angry utterances performed worst in all three datasets. We hope our analysis will initiate a new line of research in the speaker recognition community.
Raghavendra Pappagari, Tianzi Wang, Jesús Villalba 0001, Nanxin Chen, Najim Dehak
ICASSP3
2020 Self-Expressing Autoencoders for Unsupervised Spoken Term Discovery
abstract
Unsupervised spoken term discovery consists of two tasks: finding the acoustic segment boundaries and labeling acoustically similar segments with the same labels. We perform segmentation based on the assumption that the frame feature vectors are more similar within a segment than across the segments. Therefore, for strong segmentation performance, it is crucial that the features represent the phonetic properties of a frame more than other factors of variability. We achieve this via a self-expressing autoencoder framework. It consists of a single encoder and two decoders with shared weights. The encoder projects the input features into a latent representation. One of the decoders tries to reconstruct the input from these latent representations and the other from the self-expressed version of them. We use the obtained features to segment and cluster the speech data. We evaluate the performance of the proposed method in the Zero Resource 2020 challenge unit discovery task. The proposed system consistently outperforms the baseline, demonstrating the usefulness of the method in learning representations.
Saurabhchand Bhati, Jesús Villalba 0001, Piotr Zelasko, Najim Dehak
INTERSPEECH2
2020 Learning Speaker Embedding from Text-to-Speech
abstract
Zero-shot multi-speaker Text-to-Speech (TTS) generates target speaker voices given an input text and the corresponding speaker embedding.In this work, we investigate the effectiveness of the TTS reconstruction objective to improve representation learning for speaker verification.We jointly trained endto-end Tacotron 2 TTS and speaker embedding networks in a self-supervised fashion.We hypothesize that the embeddings will contain minimal phonetic information since the TTS decoder will obtain that information from the textual input.TTS reconstruction can also be combined with speaker classification to enhance these embeddings further.Once trained, the speaker encoder computes representations for the speaker verification task, while the rest of the TTS blocks are discarded.We investigated training TTS from either manual or ASR-generated transcripts.The latter allows us to train embeddings on datasets without manual transcripts.We compared ASR transcripts and Kaldi phone alignments as TTS inputs, showing that the latter performed better due to their finer resolution.Unsupervised TTS embeddings improved EER by 2.06% absolute with regard to i-vectors for the LibriTTS dataset.TTS with speaker classification loss improved EER by 0.28% and 0.73% absolutely from a model using only speaker classification loss in LibriTTS and Voxceleb1 respectively.
Piotr Zelasko, Jesús Villalba 0001, Shinji Watanabe 0001, Najim Dehak
INTERSPEECH3
2020 x-Vectors Meet Adversarial Attacks: Benchmarking Adversarial Robustness in Speaker Verification
Jesús Villalba 0001, Yuekai Zhang, Najim Dehak
INTERSPEECH1
2020 Black-Box Attacks on Spoofing Countermeasures Using Transferability of Adversarial Examples
Yuekai Zhang, Ziyan Jiang, Jesús Villalba 0001, Najim Dehak
INTERSPEECH3
2020 State-of-the-art speaker recognition with neural network embeddings in NIST SRE18 and Speakers in the Wild evaluations
Jesús Villalba 0001, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, L. Paola García-Perera, Fred Richardson, Réda Dehak, Pedro A. Torres-Carrasquillo, Najim Dehak
Comput. Speech Lang.1
2019 Low-Resource Domain Adaptation for Speaker Recognition Using Cycle-Gans
abstract
Current speaker recognition technology provides great performance with the x-vector approach. However, performance decreases when the evaluation domain is different from the training domain, an issue usually addressed with domain adaptation approaches. Recently, unsupervised domain adaptation using cycle-consistent Generative Adversarial Networks (CycleGAN) has received a lot of attention. Cycle-GAN learn mappings between features of two domains given non-parallel data. We investigate their effectiveness in low resource scenario i.e. when limited amount of target domain data is available for adaptation, a case unexplored in previous works. We experiment with two adaptation tasks: microphone to telephone and a novel reverberant to clean adaptation with the end goal of improving speaker recognition performance. Number of speakers present in source and target domains are 7000 and 191 respectively. By adding noise to the target domain during CycleGAN training, we were able to achieve better performance compared to the adaptation system whose CycleGAN was trained on a larger target data. On reverberant to clean adaptation task, our models improved EER by 18.3% relative on VOiCES dataset compared to a system trained on clean data. They also slightly improved over the state-of-the-art Weighted Prediction Error (WPE) de-reverberation algorithm.
Phani S. Nidadavolu, Saurabh Kataria 0001, Jesús Villalba 0001, Najim Dehak
ASRU3
2019 Hierarchical Transformers for Long Document Classification
abstract
BERT, which stands for Bidirectional Encoder Representations from Transformers, is a recently introduced language representation model based upon the transfer learning paradigm. We extend its fine-tuning procedure to address one of its major limitations - applicability to inputs longer than a few hundred words, such as transcripts of human call conversations. Our method is conceptually simple. We segment the input into smaller chunks and feed each of them into the base model. Then, we propagate each output through a single recurrent layer, or another transformer, followed by a softmax activation. We obtain the final classification decision after the last segment has been consumed. We show that both BERT extensions are quick to fine-tune and converge after as little as 1 epoch of training on a small, domain-specific data set. We successfully apply them in three different tasks involving customer call satisfaction prediction and topic classification, and obtain a significant improvement over the baseline models in two of them.
Raghavendra Pappagari, Piotr Zelasko, Jesús Villalba 0001, Yishay Carmiel, Najim Dehak
ASRU3
2019 Language Model Integration Based on Memory Control for Sequence to Sequence Speech Recognition
abstract
In this paper, we explore several new schemes to train a seq2seq model to integrate a pre-trained language model (LM). Our proposed fusion methods focus on the memory cell state and the hidden state in the seq2seq decoder long short-term memory (LSTM), and the memory cell state is updated by the LM unlike the prior studies. This means the memory retained by the main seq2seq would be adjusted by the external LM. These fusion methods have several variants depending on the architecture of this memory cell update and the use of memory cell and hidden states which directly affects the final label inference. We performed the experiments to show the effectiveness of the proposed methods in a mono-lingual ASR setup on the Librispeech corpus and in a transfer learning setup from a multilingual ASR (MLASR) base model to a low-resourced language. In Librispeech, our best model improved WER by 3.7%, 2.4% for test clean, test other relatively to the shallow fusion baseline, with multilevel decoding. In transfer learning from an MLASR base model to the IARPA Babel Swahili model, the best scheme improved the transferred model on eval set by 9.9%, 9.8% in CER, WER relatively to the 2-stage transfer baseline.
Shinji Watanabe 0001, Takaaki Hori, Murali Karthick Baskar, Hirofumi Inaguma, Jesús Villalba 0001, Najim Dehak
ICASSP6
2019 Investigation on Neural Bandwidth Extension of Telephone Speech for Improved Speaker Recognition
abstract
We extend our previous work on training mixed-bandwidth (BW) speaker recognition system by predicting missing information in upperband (UB) of upsampled telephone speech. Mixed-BW systems combine speech from narrowband (NB) and wideband (WB) speech corpora by basic upsampling of NB speech with low-pass filter interpolator, resulting in no information loss in the original WB speech. In this work, we explore the usage of a deep residual full-convolutional neural network (CNN) and a bidirectional long short term memory (BLSTM) network along with a previously proposed deep neural network (DNN) for bandwidth extension (BWE) of NB telephone speech. Speaker recognition systems trained with bandwidth extended features improved in performance over mixed-BW and NB baseline systems. In terms of detection cost function (DCF), the CNN-BWE system improved by 10.78% and 15.96% (relative) in the Speakers In The Wild (SITW) eval core and assist-multi-speaker condition respectively w.r.t. the NB baseline; and improved by 3.21% and 4.13% w.r.t. to the mixed-BW baseline.
Phani S. Nidadavolu, Vicente Iglesias, Jesús Villalba 0001, Najim Dehak
ICASSP3
2019 Cycle-GANs for Domain Adaptation of Acoustic Features for Speaker Recognition
abstract
It is well known that domain mismatch between the training and evaluation data hinders the performance of any machine learning system. Various factors contribute to domain mismatch. In speaker recognition systems, it mainly occurs due to the mismatch in recording conditions and language. Most speaker recognition corpora are telephone speech. Meanwhile, a few evaluation data sets like Speakers In The Wild (SITW) are microphone speech. In this work, we explore domain adaptation at acoustic feature level by learning feature mappings between domains using cycle consistent generative adversarial networks (cycle-GANs), without any parallel data between domains. Microphone features mapped to telephone domain are used to evaluate speaker recognition system trained only on telephone data. We achieved 9.37% and 2.82% relative improvement in equal error rate (EER) and detection cost function (DCF) on SITW eval set.
Phani S. Nidadavolu, Jesús Villalba 0001, Najim Dehak
ICASSP2
2019 Tied Mixture of Factor Analyzers Layer to Combine Frame Level Representations in Neural Speaker Embeddings
Nanxin Chen, Jesús Villalba 0001, Najim Dehak
INTERSPEECH2
2019 ASSERT: Anti-Spoofing with Squeeze-Excitation and Residual Networks
abstract
We present JHU's system submission to the ASVspoof 2019 Challenge: Anti-Spoofing with Squeeze-Excitation and Residual neTworks (ASSERT).Anti-spoofing has gathered more and more attention since the inauguration of the ASVspoof Challenges, and ASVspoof 2019 dedicates to address attacks from all three major types: text-to-speech, voice conversion, and replay.Built upon previous research work on Deep Neural Network (DNN), ASSERT is a pipeline for DNN-based approach to anti-spoofing.ASSERT has four components: feature engineering, DNN models, network optimization and system combination, where the DNN models are variants of squeeze-excitation and residual networks.We conducted an ablation study of the effectiveness of each component on the ASVspoof 2019 corpus, and experimental results showed that ASSERT obtained more than 93% and 17% relative improvements over the baseline systems in the two sub-challenges in ASVspooof 2019, ranking ASSERT one of the top performing systems.Code and pretrained models will be made publicly available.
Cheng-I Lai, Nanxin Chen, Jesús Villalba 0001, Najim Dehak
INTERSPEECH3
2019 The JHU Speaker Recognition System for the VOiCES 2019 Challenge
David Snyder, Jesús Villalba 0001, Nanxin Chen, Daniel Povey, Gregory Sell, Najim Dehak, Sanjeev Khudanpur
INTERSPEECH2
2019 State-of-the-Art Speaker Recognition for Telephone and Video Speech: The JHU-MIT Submission for NIST SRE18
Jesús Villalba 0001, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, Fred Richardson, Suwon Shon, François Grondin, Réda Dehak, L. Paola García-Perera, Daniel Povey, Pedro A. Torres-Carrasquillo, Sanjeev Khudanpur, Najim Dehak
INTERSPEECH1
2018 Measuring Uncertainty in Deep Regression Models: The Case of Age Estimation from Speech
abstract
Age estimation from speech recently received a lot of attention. Approaches such as i-vectors and deep learning have been successfully applied to this task achieving great performance. However, one drawback of those methods is that they produce a hard age estimation without any kind of confidence measure about the quality of the prediction. Designing systems with the ability to provide a confidence measure about their output is extremely valuable for several applications where the cost of making bad decisions is worse than making no decision, e.g., forensics. In this paper, we propose a novel framework to jointly predict the age and its estimation uncertainty in a context of neural regression model. This model is trained using probabilistic fashion instead of using the classical minimum mean square error objective used for regression tasks. The probabilistic output corresponds to a Gaussian posterior. The proposed neural network will estimate both the posterior mean which corresponds to the predicted age and the variance which quantifies the uncertainty of the prediction. We evaluated our approach on two different datasets NIST SRE 2008 - 2010 and Switchboard.
Nanxin Chen, Jesús Villalba 0001, Yishay Carmiel, Najim Dehak
ICASSP2
2018 Joint Verification-Identification in end-to-end Multi-Scale CNN Framework for Topic Identification
abstract
We present an end-to-end multi-scale Convolutional Neural Network (CNN) framework for topic identification (topic ID). In this work, we examined multi -scale CNN for classification using raw text input. Topical word embeddings are learnt at multiple scales using parallel convolutional layers. A technique to integrate verification and identification objectives is examined to improve topic ID performance. With this approach, we achieved significant improvement in identification task. We evaluated our framework on two contrasting datasets: 20 newsgroups and Fisher. We obtained 92.93% accuracy on Fisher and 86.12% on 20 newsgroups, which to our know ledge are the best published results on these datasets at the moment.
Raghavendra Pappagari, Jesús Villalba 0001, Najim Dehak
ICASSP2
2018 An Investigation of Non-linear i-vectors for Speaker Verification
Nanxin Chen, Jesús Villalba 0001, Najim Dehak
INTERSPEECH2
2018 Deep Neural Networks for Emotion Recognition Combining Audio and Transcripts
abstract
In this paper, we propose to improve emotion recognition by combining acoustic information and conversation transcripts.On the one hand, a LSTM network was used to detect emotion from acoustic features like f0, shimmer, jitter, MFCC, etc.On the other hand, a multi-resolution CNN was used to detect emotion from word sequences.This CNN consists of several parallel convolutions with different kernel sizes to exploit contextual information at different levels.A temporal pooling layer aggregates the hidden representations of different words into a unique sequence level embedding, from which we computed the emotion posteriors.We optimized a weighted sum of classification and verification losses.The verification loss tries to bring embeddings from same emotions closer while separating embeddings from different emotions.We also compared our CNN with state-of-the-art text-based hand-crafted features (e-vector).We evaluated our approach on the USC-IEMOCAP dataset as well as the dataset consisting of US English telephone speech.In the former, we used human-annotated transcripts while in the latter, we used ASR transcripts.The results showed fusing audio and transcript information improved unweighted accuracy by relative 24% for IEMOCAP and relative 3.4% for the telephone data compared to a single acoustic system.
Raghavendra Pappagari, Purva Kulkarni, Jesús Villalba 0001, Yishay Carmiel, Najim Dehak
INTERSPEECH4
2018 Effectiveness of Single-Channel BLSTM Enhancement for Language Identification
abstract
This paper proposes to apply deep neural network (DNN)-based single-channel speech enhancement (SE) to language identification. The 2017 language recognition evaluation (LRE17) introduced noisy audios from videos, in addition to the telephone conversation from past challenges. Because of that, adapting models from telephone speech to noisy speech from the video domain was required to obtain optimum performance. However, such adaptation requires knowledge of the audio domain and availability of in-domain data. Instead of adaptation, we propose to use a speech enhancement step to clean up the noisy audio as preprocessing for language identification. We used a bi-directional long short-term memory (BLSTM) neural network, which given log-Mel noisy features predicts a spectral mask indicating how clean each time-frequency bin is. The noisy spectrogram is multiplied by this predicted mask to obtain the enhanced magnitude spectrogram, and it is transformed back into the time domain by using the unaltered noisy speech phase. The experiments show significant improvement to language identification of noisy speech, for systems with and without domain adaptation, while preserving the identification performance in the telephone audio domain. In the best adapted state-of-the-art bottleneck i-vector system the relative improvement is 11.3% for noisy speech.
Peter Sibbern Frederiksen, Jesús Villalba 0001, Shinji Watanabe 0001, Zheng-Hua Tan, Najim Dehak
INTERSPEECH2
2018 End-to-end Deep Neural Network Age Estimation
Pegah Ghahremani, Phani S. Nidadavolu, Nanxin Chen, Jesús Villalba 0001, Daniel Povey, Sanjeev Khudanpur, Najim Dehak
INTERSPEECH4
2018 Investigation on Bandwidth Extension for Speaker Recognition
Phani S. Nidadavolu, Cheng-I Lai, Jesús Villalba 0001, Najim Dehak
INTERSPEECH3
2018 Diarization is Hard: Some Experiences and Lessons Learned for the JHU Team in the Inaugural DIHARD Challenge
Gregory Sell, David Snyder, Alan McCree, Daniel Garcia-Romero, Jesús Villalba 0001, Matthew Maciejewski, Vimal Manohar, Najim Dehak, Daniel Povey, Shinji Watanabe 0001, Sanjeev Khudanpur
INTERSPEECH5
2017 Tied Variational Autoencoder Backends for i-Vector Speaker Recognition
Jesús Villalba 0001, Niko Brümmer, Najim Dehak
INTERSPEECH1
2017 Domain Adaptation of PLDA Models in Broadcast Diarization by Means of Unsupervised Speaker Clustering
Ignacio Viñals, Alfonso Ortega Giménez, Jesús Villalba 0001, Antonio Miguel, Eduardo Lleida
INTERSPEECH3
2016 Analysis of speech quality measures for the task of estimating the reliability of speaker verification decisions
Jesús Villalba 0001, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
Speech Commun.1
2016 Bayesian Networks to Model the Variability of Speaker Verification Scores in Adverse Environments
abstract
State-of-the-art speaker recognition technology attains great performance in controlled conditions. However, when the speech segments suffer distortions like noise or reverberation performance can severely deteriorate, this fact motivated us to investigate how score distributions diverge from the ideal ones in degraded conditions. We propose a Bayesian network model that assumes that two scores exist: one observed and another one hidden. The observed score or noisy score is the one given by the speaker verification system. Meanwhile, the hidden score or clean score is the ideal score that we would obtain in a trial with high-quality speech. A set of quality measures helps to relate both scores. We applied this network to two tasks. The first one consists in rejecting unreliable trials, i.e., trials that we cannot assure whether they are target or nontarget. We prove that this method outperforms previous approaches, based on another type of Bayesian networks. The second task is to compute an improved likelihood ratio, dependent on the quality measures. This ratio improved calibration in noisy conditions.
Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Variational Bayesian PLDA for speaker diarization in the MGB challenge
abstract
This paper describes the ViVoLab speaker diarization system for the Multi-Genre Broadcast (MGB) Challenge at ASRU2015. The challenge data consisted of BBC TV programmes of different genres. Diarization followed a longitudinal setup, i.e., the speakers of the current episode had to be linked to the speakers in previous episodes of the same show. We propose a system based on the i-vector paradigm. After an initial segmentation step, we compute an i-vector per speech segment. Then, a generative model based on Bayesian PLDA clusters the speakers. In this model, the speaker labels are latent variables that we optimize by variational Bayes iterations. The number of speakers in each episode was decided by maximizing the variational lower bound. The system includes several phases of segment-merging and re-clustering. We re-compute i-vectors after each merging step, which reduces the i-vector uncertainty. This approach attained a DER around 30% in the development set.
Jesús Villalba 0001, Alfonso Ortega Giménez, Antonio Miguel, Eduardo Lleida
ASRU1
2015 Spoofing detection with DNN and one-class SVM for the ASVspoof 2015 challenge
abstract
Speaker verification systems have achieved great performance in recent times. However, we usually measure performance on a ideal scenarios with naive impostors that do not modify their voices to impersonate the target speakers. The fact of impersonating a legitimate user is known as spoofing attack. Recent works show the vulnerability of current speaker verification technology to several types of attacks. Most of these works use non-public databases and different performance measures, which makes difficult to compare approaches. The spoofing challenge (ASVspoof 2015) tries to overcome this problem by proposing a common evaluation framework. This paper describes our submission to the challenge. We proposed to use spectral log-filter-bank and relative phase shift features as input to classifiers based on deep neural networks (DNN). The first of our classifiers used DNN posteriors to decide if the trial is spoof or non-spoof. The second used a bottleneck feature from the DNN as input to a one-class SVM. The one-class SVM models the distribution of legitimate speech, not needing spoofing data for training. We fused the score of the different classifiers to produce our final submission. Our system attained very competitive results with EER<0.05% in 9 out of 10 spoofing types.
Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH1
2014 Unsupervised adaptation of PLDA by using variational Bayes methods
abstract
State-of-the-art speaker recognition relays on models that need a large amount of training data. This models are successful in tasks like NIST SRE because there is sufficient data available. However, in real applications, we usually do not have so much data and, in many cases, the speaker labels are unknown. We present a method to adapt a PLDA model from a domain with a large amount of labeled data to another with unlabeled data. We describe a generative model that produces both sets of data where the unknown labels are modeled like latent variables. We used variational Bayes to estimate the hidden variables. We performed experiments adapting a model trained on Switchboard to NIST SRE without labels. The adapted model is evaluated on NIST SRE10. Compared to the non-adapted model, EER improved by 42% and 49% by adapting with 200 and with all the NIST speakers respectively.
Jesús Villalba 0001, Eduardo Lleida
ICASSP1
2014 Factor analysis with sampling methods for text dependent speaker recognition
Antonio Miguel, Jesús Villalba 0001, Alfonso Ortega Giménez, Eduardo Lleida, Carlos Vaquero
INTERSPEECH2
2013 Segmentation-by-classification system based on factor analysis
abstract
This paper proposes a novel audio segmentation-by-classification system based on Factor Analysis (FA) with a channel compensation matrix for each class and scoring the fixed-length segments as the log-likelihood ratio between class/no-class. The scores are smoothed and the most probable sequence is computed with a Viterbi algorithm. The system described here is designed to segment and classify the audio files coming from broadcast programs into five different classes: speech (SP), speech with noise (SN), speech with music (SM), music (MU) or others (OT). This task was proposed in the Albayzin 2010 evaluation campaign. The system is compared with the winning system of the evaluation achieving lower error rate in SP and SN. These classes represent 3/4 of the total amount of the data. Therefore, the FA segmentation system gets a reduction in the average segmentation error rate.
Diego Castán, Alfonso Ortega Giménez, Jesús Villalba 0001, Antonio Miguel, Eduardo Lleida
ICASSP3
2013 Handling i-vectors from different recording conditions using multi-channel simplified PLDA in speaker recognition
abstract
In this work, we address the problem of having i-vectors that have been produced in different channel conditions. Traditionally, this problem has been handled training the LDA covariance matrices pooling the data of all the conditions or averaging the covariance matrices of each condition in different ways. We present a PLDA variant that we call, multi-channel SPLDA, where the speaker space distribution is common to all i-vectors and the channel space distribution depends on the type of channel where the segment has been recorded. We test our approach on the telephone part of the NIST SRE10 extended condition where we added some additive noises to the test segments. We compare results of a SPLDA model trained only with clean data, SPLDA trained with pooled noisy and clean data and our MCSPLDA model.
Jesús Villalba 0001, Eduardo Lleida
ICASSP1
2013 A new Bayesian network to assess the reliability of speaker verification decisions
abstract
In some situations the quality of the signals involved in a speaker verification trial is not as good as needed to take a reliable decision. In this work, we present a new method based on Bayesian networks and quality measures to estimate if the trial decision is reliable. We present experiments on the NIST SRE2010 dataset degraded with additive noise. A system well calibrated for clean speech, produces a large actual DCF on the degraded dataset. We use our method to discard the unreliable trials and achieve a dramatic improvement of the cost values. We also prove that our method outperforms previously published approaches.
Jesús Villalba 0001, Eduardo Lleida, Alfonso Ortega Giménez, Antonio Miguel
INTERSPEECH1
2013 The I3a speaker recognition system for NIST SRE12: post-evaluation analysis
abstract
The I3A submission for the recent NIST 2012 speaker recognition evaluation (SRE) was based on the i-vector approach with a multi-channel PLDA classifier. This PLDA is modified so that, for each i-vector, the between-class covariance depends on the type of channel where the segment was recorded (telephone,interviews,clean, noisy, etc). In this paper, we present the description of our submission and a detailed post-evaluation analysis of the results. We analyze several factors affecting performance: enrollment data selection, classifier type, scoring technique, calibration, known and unknown non-targets, target speakers included or not in development, segment duration, noise level and noise type. Some of these factor are new in this evaluation. After post-evaluation, actual costs improve by 15– 43% depending on the common condition.
Jesús Villalba 0001, Eduardo Lleida, Alfonso Ortega Giménez, Antonio Miguel
INTERSPEECH1
2013 Handling recordings acquired simultaneously over multiple channels with PLDA
abstract
In some speaker recognition scenarios we find conversations recorded simultaneously over multiple channels.That is the case of the interviews in the NIST SRE dataset.To take advantage of that, we propose a modification of the PLDA model that considers two different inter-session variability terms.The first term is tied between all the recordings belonging to the same conversation whereas the second is not.Thus, the former mainly intends to capture the variability due to the phonetic content of the conversation while the latter tries to capture the channel variability.We test this approach on the NIST SRE12 core condition using multiple channels per interview to enroll the speakers.The proposed approach improves the minimum DCF by 26-29 % on telephone speech and by 1-8% on interviews compared to the standard PLDA (scored by the book).
Jesús Villalba 0001, Mireia Díez, Amparo Varona, Eduardo Lleida
INTERSPEECH1
2012 The BLZ Submission to the NIST 2011 LRE: Data Collection, System Development and Performance
abstract
This paper describes the most relevant features of a collaborative multi-site submission to the NIST 2011 Language Recognition Evaluation (LRE), consisting of one primary and three contrastive systems, each fusing different combinations of 13 state-of-the-art (acoustic and phonotactic) language recognition subsystems.The collaboration focused on collecting and sharing training data for those target languages for which few development data were provided by NIST, and on defining a common development dataset to train backend and fusion parameters and select the best fusions.Official and post-key results are presented and compared, revealing that the greedy approach applied to select the best fusions provided suboptimal but very competitive performance.Several factors contributed to the high performance attained by BLZ systems, including the availability of training data for low resource target languages, the reliability of the development dataset (consisting only of data audited by NIST), the diversity of modeling approaches, features and datasets in the systems considered for fusion, and the effectiveness of the search for optimal fusions.
Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel, Alberto Abad, David Martínez González, Jesús Villalba 0001, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH8
2011 Multi-site heterogeneous system fusions for the Albayzin 2010 Language Recognition Evaluation
abstract
Best language recognition performance is commonly obtained by fusing the scores of several heterogeneous systems. Regardless the fusion approach, it is assumed that different systems may contribute complementary information, either because they are developed on different datasets, or because they use different features or different modeling approaches. Most authors apply fusion as a final resource for improving performance based on an existing set of systems. Though relative performance gains decrease as larger sets of systems are considered, best performance is usually attained by fusing all the available systems, which may lead to high computational costs. In this paper, we aim to discover which technologies combine the best through fusion and to analyse the factors (data, features, modeling methodologies, etc.) that may explain such a good performance. Results are presented and discussed for a number of systems provided by the participating sites and the organizing team of the Albayzin 2010 Language Recognition Evaluation. We hope the conclusions of this work help research groups make better decisions in developing language recognition technology.
Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel, David Martínez González, Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida, Alberto Abad, Oscar Koller, Isabel Trancoso, Paula Lopez-Otero, Laura Docío Fernández, Carmen García-Mateo, Rahim Saeidi, Mehdi Soufifar, Tomi Kinnunen, Torbjørn Svendsen, Pasi Fränti
ASRU7
2011 Hierarchical Audio Segmentation with HMM and Factor Analysis in Broadcast News Domain
Diego Castán, Carlos Vaquero, Alfonso Ortega Giménez, David Martínez González, Jesús Villalba 0001, Eduardo Lleida
INTERSPEECH5
2011 I3A Language Recognition System for Albayzin 2010 LRE
David Martínez González, Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH2
2011 Towards Fully Bayesian Speaker Recognition: Integrating Out the Between-Speaker Covariance
abstract
We propose a variational Bayes solution to integrate out the model parameters in a generative i-vector speaker recognizer. The existing state-of-the-art in generative i-vector modelling plugs in fixed maximum-likelihood point-estimates of model parameters. This recipe may suffer from over-fitting of especially the between-speaker covariance. We show how to integrate out the between-speaker covariance and demonstrate dramatic improvements on NIST SRE 2010.
Jesús Villalba 0001, Niko Brümmer
INTERSPEECH1
2010 Speaker Verification in Noisy Environment Using Missing Feature Approach
Dayana Ribas González, Jesús Villalba 0001, Eduardo Lleida, José Ramón Calvo de Lara
CIARP2
2010 Confidence measures for speaker segmentation and their relation to speaker verification
Carlos Vaquero, Alfonso Ortega Giménez, Jesús Villalba 0001, Antonio Miguel, Eduardo Lleida
INTERSPEECH3