VLDB 2026 Research / reviewers in the wild / expert
Najim Dehak
dblp:18/2810
· DBLP profile ↗
146ranked-venue papers
7as first author
65since 2021 · last 2026
0000-0002-4489-5753ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 123 · 5 first-author · 51 since 2021Artificial intelligence and machine learning · 103 · 5 first-author · 45 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational SpeechabstractThere are not enough established benchmarks for the task fo speech summarization. Creating new benchmarks demands human annotation, as LLMs could embed systemic errors and bias into datasets. We test ten annotation workflows varying input modality (audio, transcript, or both) and the inclusion of editing (self or peer-editing) to investigate potential quality tradeoffs from using human annotators to summarize audio. We compare human audio-based summaries to human transcript-based summaries to track the impact of the different information modalities on summary quality. We also compare the human outputs against four LLM benchmarks (three text, one audio) to examine whether human-written summaries are less informative than highly fluent automated outputs. We find that audio-based summaries are less informative and more compressed than transcript summaries. However, iterative peer-editing with audio mitigates this difference, enabling audio-based summaries to be as informative as their transcript counterparts and LLM summaries. These findings validate iterative peer-editing among human annotators for the creation of benchmarks informed by both lexical and prosodic information. This enables crucial dataset collection even in setting where transcripts are unavailable. Kaavya Chaparala, Thomas Thebaud, Jesús Villalba 0001, Laureano Moro-Velázquez, Peter Viechnicki, Najim Dehak |
LREC | 6 |
| 2026 | Towards Fair Speech Recognition: Mitigating Demographic Bias in End-to-End ASR Systems
Maliha Jahan, Thomas Thebaud, Zsuzsanna Fagyal, Jesús Villalba 0001, Mark Hasegawa-Johnson, Laureano Moro-Velázquez, Najim Dehak |
LREC | 7 |
| 2026 | Interpretable Features for the Assessment of Neurodegenerative Diseases Through Handwriting AnalysisabstractMotor dysfunction is a common sign of neurodegenerative diseases (NDs) such as Parkinson's disease (PD) and Alzheimer's disease (AD), but may be difficult to detect, especially in the early stages. In this work, we examine the behavior of a wide array of interpretable features extracted from the handwriting signals of 113 subjects performing multiple tasks on a digital tablet, as part of the Neurological Signals dataset. The aim is to measure their effectiveness in characterizing NDs, including AD and PD. To this end, task-agnostic and task-specific features are extracted from 14 distinct tasks. Subsequently, through statistical analysis and a series of classification experiments, we investigate which features provide greater discriminative power between NDs and healthy controls and amongst different NDs. Preliminary results indicate that the tasks at hand can all be effectively leveraged to distinguish between the considered set of NDs, specifically by measuring the stability, the speed of writing, the time spent not writing, and the pressure variations between groups from our handcrafted interpretable features, which shows a statistically significant difference between groups, across multiple tasks. Using various binary classification algorithms on the computed features, we obtain up to 87% accuracy for the discrimination between AD and healthy controls (CTL), and up to 69% for the discrimination between PD and CTL. Thomas Thebaud, Anna Favaro, Casey Chen, Gabriel Chávez, Laureano Moro-Velázquez, Emile Moukhebeir, Ankur A. Butala, Najim Dehak |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | Multi-Target Backdoor Attacks Against Speaker RecognitionabstractIn this work, we propose a multi-target backdoor attack against speaker identification using position-independent clicking sounds as triggers. Unlike previous single-target approaches, our method targets up to 50 speakers simultaneously, achieving success rates of up to 95.04%. To simulate more realistic attack conditions, we vary the signal-to-noise ratio between speech and trigger, demonstrating a trade-off between stealth and effectiveness. We further extend the attack to the speaker verification task by selecting the most similar training speaker—based on cosine similarity—as a proxy target. The attack is most effective when target and enrolled speaker pairs are highly similar, reaching success rates of up to 90% in such cases. Alexandrine Fortier, Sonal Joshi, Thomas Thebaud, Jesús Villalba 0001, Najim Dehak, Patrick Cardinal |
ASRU | 5 |
| 2025 | Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM
Thomas Thebaud, Yen-Ju Lu, Matthew Wiesner, Peter Viechnicki, Najim Dehak |
ASRU | 5 |
| 2025 | The JHU-MIT System for NIST SRE24: Post-Evaluation AnalysisabstractWe present the JHU-MIT submission for NIST SRE24, along with post-evaluation analysis and key insights. In the audio fixed condition, our system used Res2Net50 and ResNet100 embeddings; the open condition additionally included an ECAPA-TDNN with a multilingual Wav2Vec2 front-end, which emerged as the best single system. The audio back-ends consisted of either PLDA adapted to SRE24 Dev or a mixture of PLDA models tuned to different subconditions. To avoid overfitting, we optimized back-end hyperparameters via twofold cross-validation. For the visual condition, we leveraged pretrained ResNet100-Subcenter-ArcFace embeddings. Agglomerative clustering was used to diarize speaker and face identities in multi-speaker videos. The primary audio fixed system achieved Act. Cp=0.574, while the open condition reached Cp=0.366 on SRE24 Eval. The visual system yielded Cp=0.169, and audiovisual fusion further improved performance, achieving Cp=0.101 (fixed) and Cp=0.087 (open). Jesús Villalba 0001, Jonas Borgstrom, Prabhav Singh, L. Paola García-Perera, Pedro A. Torres-Carrasquillo, Najim Dehak |
ASRU | 6 |
| 2025 | Paired by the Teacher: Turning Unpaired Data into High-Fidelity Pairs for Low-Resource Text GenerationabstractWe present Paired by the Teacher (PbT), a twostage teacher-student pipeline that synthesizes accurate input-output pairs without human labels or parallel data.In many low-resource natural language generation (NLG) scenarios, practitioners may have only raw outputs, like highlights, recaps, or questions, or only raw inputs, such as articles, dialogues, or paragraphs, but seldom both.This mismatch forces small models to learn from very few examples or rely on costly, broad-scope synthetic examples produced by large LLMs.PbT addresses this by asking a teacher LLM to compress each unpaired example into a concise intermediate representation (IR), and training a student to reconstruct inputs from IRs.This enables outputs to be paired with student-generated inputs, yielding high-quality synthetic data.We evaluate PbT on five benchmarks-document summarization (XSum, CNNDM), dialogue summarization (SAMSum, DialogSum), and question generation (SQuAD)-as well as an unpaired setting on SwitchBoard (paired with Dialog-Sum summaries).An 8B student trained only on PbT data outperforms models trained on 70 B teacher-generated corpora and other unsupervised baselines, coming within 1.2 ROUGE-L of human-annotated pairs and closing 82% of the oracle gap at one-third the annotation cost of direct synthesis.Human evaluation on SwitchBoard further confirms that only PbT produces concise, faithful summaries aligned with the target style, highlighting its advantage of generating in-domain sources that avoid the mismatch, limiting direct synthesis. Yen-Ju Lu, Thomas Thebaud, Laureano Moro-Velázquez, Najim Dehak, Jesús Villalba 0001 |
EMNLP | 4 |
| 2025 | Unveiling Performance Bias in ASR Systems: A Study on Gender, Age, Accent, and MoreabstractWith the recent advancements in speech recognition, it is crucial to ensure these systems are free from performance biases against any speaker subgroups. This study examined the performance of twenty variants of seven Automatic Speech Recognition models across four datasets in English language: L2 Arctic, Speech Accent Archive, CORAAL, and SBCSAE. We employed Poisson regression and drop-in-deviance tests to identify which attributes significantly contribute to the Word Error Rate. Our analysis revealed biases related to attributes such as native language, location, occupation, and birthplace. Most systems did not exhibit bias related to factors like gender and age. Additionally, we conducted an experiment to detect bias related to "variant" (accent and dialect) by combining the CORAAL (African American Vernacular English (AAVE)) and SBCSAE (General American English (GAE)) datasets, aiming to identify the sources of any observed bias. We found that both speaker variability and dialectal difference contribute to observed bias for variant. Maliha Jahan, Priyam Mazumdar, Thomas Thebaud, Mark Hasegawa-Johnson, Jesús Villalba 0001, Najim Dehak, Laureano Moro-Velázquez |
ICASSP | 6 |
| 2025 | Detecting Neurodegenerative Diseases using Frame-Level Handwriting EmbeddingsabstractIn this study, we explored the use of spectrograms to represent handwriting signals for assessing neurodegenerative diseases, including 42 healthy controls (CTL), 35 subjects with Parkinson’s Disease (PD), 21 with Alzheimer’s Disease (AD), and 15 with Parkinson’s Disease Mimics (PDM). We applied CNN and CNN-BLSTM models for binary classification using both multi-channel fixed-size and frame-based spectrograms. Our results showed that handwriting tasks and spectrogram channel combinations significantly impacted classification performance. The highest F1-score (89.8%) was achieved for AD vs. CTL, while PD vs. CTL reached 74.5%, and PD vs. PDM scored 77.97%. CNN consistently outperformed CNN-BLSTM. Different sliding window lengths were tested for constructing frame-based spectrograms. A 1-second window worked best for AD, longer windows improved PD classification, and window length had little effect on PD vs. PDM. Sarah Laouedj, Jesús Villalba 0001, Thomas Thebaud, Laureano Moro-Velázquez, Najim Dehak |
ICASSP | 6 |
| 2025 | SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion TransformerabstractIn this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer that operates on latent features. SoloAudio supports both audio-oriented and language-oriented TSE by utilizing a CLAP model as the feature extractor for target sounds. Furthermore, SoloAudio leverages synthetic audio generated by state-of-the-art text-to-audio models for training, demonstrating strong generalization to out-of-domain data and unseen sound events. We evaluate this approach on the FSD Kaggle 2018 mixture dataset and real data from AudioSet, where SoloAudio achieves the state-of-the-art results on both in-domain and out-of-domain data, and exhibits impressive zero-shot and few-shot capabilities. Source code1and demos2are released. Helin Wang, Jiarui Hai, Yen-Ju Lu, Karan Thakkar, Mounya Elhilali, Najim Dehak |
ICASSP | 6 |
| 2025 | SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and SynthesisabstractIn this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot text-based speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates classifier-free guidance to enhance the stability of the generation process. A watermark Encodec is proposed to embed frame-level watermarks into the edited regions of the speech so that which parts were edited can be detected. In addition, the waveform reconstruction leverages the original unedited speech segments, providing superior recovery compared to the Encodec model. Our approach achieves state-of-the-art performance in the RealEdit speech editing task and the LibriTTS text-to-speech task, surpassing previous methods. Furthermore, SSR-Speech excels in multi-span speech editing and also demonstrates remarkable robustness to background sounds. The source code1and demos2are released. Helin Wang, Meng Yu 0003, Jiarui Hai, Chen Chen 0075, Rilin Chen, Najim Dehak, Dong Yu 0001 |
ICASSP | 7 |
| 2025 | ADCeleb: A Longitudinal Speech Dataset from Public Figures for Early Detection of Alzheimer's Disease
Kunxiao Gao, Anna Favaro, Najim Dehak, Laureano Moro-Velázquez |
INTERSPEECH | 3 |
| 2025 | FaiST: A Benchmark Dataset for Fairness in Speech Technology
Maliha Jahan, Yinglun Sun, Priyam Mazumdar, Zsuzsanna Fagyal, Thomas Thebaud, Jesús Villalba 0001, Mark Hasegawa-Johnson, Najim Dehak, Laureano Moro-Velázquez |
INTERSPEECH | 8 |
| 2025 | The Interspeech 2025 Challenge on Speech Emotion Recognition in Naturalistic Conditions
Abinay Reddy Naini, Lucas Goncalves, Ali N. Salman, Pravin Mote, Ismail Rasim Ülgen, Thomas Thebaud, Laureano Moro-Velázquez, L. Paola García-Perera, Najim Dehak, Berrak Sisman, Carlos Busso |
INTERSPEECH | 9 |
| 2025 | Count Your Speakers! Multitask Learning for Multimodal Speaker Diarization
Prabhav Singh, Jesús Villalba 0001, Najim Dehak |
INTERSPEECH | 3 |
| 2025 | Multimodal Emotion Diarization: Frame-Wise Integration of Text and Audio Representations
Ziv Tamir, Thomas Thebaud, Jesús Villalba 0001, Najim Dehak, Oren Kurland |
INTERSPEECH | 4 |
| 2025 | Joint Diarization and Separation Using SepFormer With Non-Autoregressive Attractors
Magdalena Rybicka, Konrad Kowalczyk, Thomas Thebaud, Najim Dehak, Jesús Villalba 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Finding Spoken Identifications: Using GPT-4 Annotation for an Efficient and Fast Dataset Creation PipelineabstractThe growing emphasis on fairness in speech-processing tasks requires datasets with speakers from diverse subgroups that allow training and evaluating fair speech technology systems. However, creating such datasets through manual annotation can be costly. To address this challenge, we present a semi-automated dataset creation pipeline that leverages large language models. We use this pipeline to generate a dataset of speakers identifying themself or another speaker as belonging to a particular race, ethnicity, or national origin group. We use OpenaAI’s GPT-4 to perform two complex annotation tasks- separating files relevant to our intended dataset from the irrelevant ones (filtering) and finding and extracting information on identifications within a transcript (tagging). By evaluating GPT-4’s performance using human annotations as ground truths, we show that it can reduce resources required by dataset annotation while barely losing any important information. For the filtering task, GPT-4 had a very low miss rate of 6.93%. GPT-4’s tagging performance showed a trade-off between precision and recall, where the latter got as high as 97%, but precision never exceeded 45%. Our approach reduces the time required for the filtering and tagging tasks by 95% and 80%, respectively. We also present an in-depth error analysis of GPT-4’s performance. Maliha Jahan, Helin Wang, Thomas Thebaud, Yinglun Sun, Giang Ha Le, Zsuzsanna Fagyal, Odette Scharenborg, Mark Hasegawa-Johnson, Laureano Moro-Velázquez, Najim Dehak |
LREC/COLING | 10 |
| 2024 | DPM-TSE: A Diffusion Probabilistic Model for Target Sound ExtractionabstractCommon target sound extraction (TSE) approaches primarily relied on discriminative approaches in order to separate the target sound while minimizing interference from the unwanted sources, with varying success in separating the target from the background. This study introduces DPM-TSE, a generative method based on diffusion probabilistic modeling (DPM) for Target Sound Extraction (TSE), to achieve both cleaner target renderings as well as improved separability from unwanted sounds. The technique also tackles the noise floor of DPM by introducing a correction method for noise schedules and sample steps. This approach is evaluated using both objective and subjective quality metrics on the FSD Kaggle 2018 dataset. The results show that DPM-TSE has a significant improvement in perceived quality in terms of target extraction and purity. Jiarui Hai, Helin Wang, Dongchao Yang, Karan Thakkar, Najim Dehak, Mounya Elhilali |
ICASSP | 5 |
| 2024 | Multimodal Emotion Recognition Harnessing the Complementarity of Speech, Language, and VisionabstractIn the realm of audiovisual emotion recognition, a significant challenge lies in developing neural network architectures capable of effectively harnessing and integrating multimodal information. This study introduces an advanced methodology for the Empathic Virtual Agent Challenge (EVAC), utilizing state-of-the-art speech, language, and image models. Specifically, we leverage cutting-edge pre-trained models, including multilingual variants fine-tuned in French for each modality, and integrate them using late fusion techniques. Through extensive experimentation and validation, we demonstrate the efficacy of our approach in achieving competitive results on the challenge dataset. Our findings highlight that multimodal approaches outperform unimodal methods across Core Affect Presence and Intensity and Appraisal Dimensions tasks, underscoring the effectiveness of integrating diverse modalities. This underscores the importance of leveraging multiple sources of information to capture nuanced emotional states more accurately and robustly in real-world applications. Thomas Thebaud, Anna Favaro, Yaohan Guan, Prabhav Singh, Jesús Villalba 0001, Laureano Moro-Velázquez, Najim Dehak |
ICMI | 8 |
| 2024 | Leveraging Universal Speech Representations for Detecting and Assessing the Severity of Mild Cognitive Impairment Across Languages
Anna Favaro, Tianyu Cao 0003, Najim Dehak, Laureano Moro-Velázquez |
INTERSPEECH | 3 |
| 2024 | Noise-robust Speech Separation with Fast Generative Correction
Helin Wang, Jesús Villalba 0001, Laureano Moro-Velázquez, Jiarui Hai, Thomas Thebaud, Najim Dehak |
INTERSPEECH | 6 |
| 2024 | Exploring the Complementary Nature of Speech and Eye Movements for Profiling Neurological Disorders
Anna Favaro, Thomas Thebaud, Jesús Villalba 0001, Najim Dehak, Laureano Moro-Velázquez |
INTERSPEECH | 5 |
| 2024 | CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech ProcessingabstractWe introduce Condition-Aware Self-Supervised Learning Representation (CA-SSLR), a generalist conditioning model broadly applicable to various speech-processing tasks. Compared to standard fine-tuning methods that optimize for downstream models, CA-SSLR integrates language and speaker embeddings from earlier layers, making the SSL model aware of the current language and speaker context.
This approach reduces the reliance on the input audio features while preserving the integrity of the base SSLR. CA-SSLR improves the model’s capabilities and demonstrates its generality on unseen tasks with minimal task-specific tuning. Our method employs linear modulation to dynamically adjust internal representations, enabling fine-grained adaptability without significantly altering the original model behavior. Experiments show that CA-SSLR reduces the number of trainable parameters, mitigates overfitting, and excels in under-resourced and unseen tasks. Specifically, CA-SSLR achieves a 10\% relative reduction in LID errors, a 37\% improvement in ASR CER on the ML-SUPERB benchmark, and a 27\% decrease in SV EER on VoxCeleb-1, demonstrating its effectiveness. Yen-Ju Lu, Thomas Thebaud, Laureano Moro-Velázquez, Ariya Rastrow, Najim Dehak, Jesús Villalba 0001 |
NeurIPS | 6 |
| 2024 | Clean Label Attacks Against SLU SystemsabstractPoisoning backdoor attacks involve an adversary manipulating the training data to induce certain behaviors in the victim model by inserting a trigger in the signal at inference time. We adapted clean label backdoor (CLBD)-data poisoning attacks, which do not modify the training labels, on state-of-the-art speech recognition models that support/perform a Spoken Language Understanding task, achieving 99.8% attack success rate by poisoning 10% of the training data. We analyzed how varying the signal-strength of the poison, percent of samples poisoned, and choice of trigger impact the attack. We also found that CLBD attacks are most successful when applied to training samples that are inherently hard for a proxy model. Using this strategy, we achieved an attack success rate of 99.3% by poisoning a meager 1.5% of the training data. Finally, we applied two previously developed defenses against gradient-based attacks, and found that they attain mixed success against poisoning. Henry Li Xinyuan, Sonal Joshi, Thomas Thebaud, Jesús Villalba 0001, Najim Dehak, Sanjeev Khudanpur |
SLT | 5 |
| 2024 | Slowness Regularized Contrastive Predictive Coding for Acoustic Unit DiscoveryabstractSelf-supervised methods such as Contrastive predictive Coding (CPC) have greatly improved the quality of the unsupervised representations. These representations significantly reduce the amount of labeled data needed for downstream task performance, such as automatic speech recognition. CPC learns representations by learning to predict future frames given current frames. Based on the observation that the acoustic information, e.g., phones, changes slower than the feature extraction rate in CPC, we propose regularization techniques that impose slowness constraints on the features. Here we propose two regularization techniques: Self-expressing constraint and Left-or-Right regularization. We evaluate the proposed model on ABX and linear phone classification tasks, acoustic unit discovery, and automatic speech recognition. The regularized CPC trained on 100 hours of unlabeled data matches the performance of the baseline CPC trained on 360 hours of unlabeled data. We also show that our regularization techniques are complementary to data augmentation and can further boost the system's performance. In monolingual, cross-lingual, or multilingual settings, with/without data augmentation, regardless of the amount of data used for training, our regularized models outperformed the baseline CPC models on the ABX task. Saurabhchand Bhati, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Time-Domain Speech Super-Resolution With GAN Based Modeling for Telephony Speaker VerificationabstractAutomatic Speaker Verification(ASV) technology has become commonplace in virtual assistants. However, its performance suffers when there is a mismatch between the train and test domains. Mixed bandwidth training, i.e., pooling training data from both domains, is a preferred choice for developing a universal model that works for both narrowband and wideband domains. We propose complementing this technique by performing neural upsampling of narrowband signals, also known as bandwidth extension. We aim to discover and analyze high-performing time-domain Generative Adversarial Network (GAN) based models to improve our downstream state-of-the-art ASV system. We choose GANs since they 1) are powerful for learning conditional distribution and 2) allow flexibleplug-inusage as a pre-processor during the training of downstream tasks (ASV) with data augmentation. Prior work mainly focused on feature-domain bandwidth extension and limited experimental setups. We address these limitations by 1) using time-domain extension models, 2) reporting results on three real test sets, 3) extending training data, and 4) devising new test-time schemes. We compare supervised (conditional GAN) and unsupervised GANs (CycleGAN) and demonstrate an average relative improvement in the equal error rate of 8.6% and 7.7%, respectively. For further analysis, we study changes in the visual quality of the spectrogram, audio perceptual quality, t-SNE embeddings, and ASV score distributions. We show that our bandwidth extension leads to phenomena such as a shift of telephone (test) embeddings towards wideband (train) signals, a negative correlation of perceptual quality with downstream performance, and condition-independent score calibration. Saurabh Kataria 0001, Jesús Villalba 0001, Laureano Moro-Velázquez, Piotr Zelasko, Najim Dehak |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | End-to-End Neural Speaker Diarization With Non-Autoregressive AttractorsabstractDespite many recent developments in speaker diarization, it remains a challenge and an active area of research to make diarization robust and effective in real-life scenarios. Well-established clustering-based methods are showing good performance and qualities. However, such systems are built of several independent, separately optimized modules, which may cause non-optimum performance. End-to-end neural speaker diarization (EEND) systems are considered the next stepping stone in pursuing high-performance diarization. Nevertheless, this approach also suffers limitations, such as dealing with long recordings and scenarios with a large (more than four) or unknown number of speakers in the recording. The appearance of EEND with encoder-decoder-based attractors (EEND-EDA) enabled us to deal with recordings that contain a flexible number of speakers thanks to an LSTM-based EDA module. A competitive alternative over the referenced EEND-EDA baseline is the EEND with non-autoregressive attractor (EEND-NAA) estimation, proposed recently by the authors of this article. NAA back-end incorporates k-means clustering as part of the attractor estimation and an attractor refinement module based on a Transformer decoder. However, in our previous work on EEND-NAA, we assumed a known number of speakers, and the experimental evaluation was limited to 2-speaker recordings only. In this article, we describe in detail our recent EEND-NAA approach and propose further improvements to the EEND-NAA architecture, introducing three novel variants of the NAA back-end, which can handle recordings containing speech of a variable and unknown number of speakers. Conducted experiments include simulated mixtures generated using the Switchboard and NIST SRE datasets and real-life recordings from the CALLHOME and DIHARD II datasets. In experimental evaluation, the proposed systems achieve up to 51% relative improvement for the simulated scenario and up to 15% for real recordings over the baseline EEND-EDA. Magdalena Rybicka, Jesús Villalba 0001, Thomas Thebaud, Najim Dehak, Konrad Kowalczyk |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Model-Based Fairness Metric for Speaker VerificationabstractEnsuring that technological advancements benefit all groups of people equally is crucial. The first step towards fairness is identifying existing inequalities. The naive comparison of group error rates may lead to wrong conclusions. We introduce a new method to determine whether a speaker verification system is fair toward several population subgroups. We propose to model miss and false alarm probabilities as a function of multiple factors, including the population group effects, e.g., male and female, and a series of confounding variables, e.g., speaker effects, language, nationality, etc. This model can estimate error rates related to a group effect without the influence of confounding effects. We experiment with a synthetic dataset where we control group and confounding effects. Our metric achieves significantly lower false positive and false negative rates w.r.t. baseline. We also experiment with VoxCeleb and NIST SRE21 datasets on different ASV systems and present our conclusions. Maliha Jahan, Laureano Moro-Velázquez, Thomas Thebaud, Najim Dehak, Jesús Villalba 0001 |
ASRU | 4 |
| 2023 | Joint Energy-Based Model for Robust Speech Classification System Against Dirty-Label Backdoor Poisoning AttacksabstractOur novel technique utilizes a Joint Energy-based Model (JEM) that integrates both discriminative and generative approaches to increase resistance against dirty-label backdoor attacks. Our approach is especially effective when the trigger is short or hardly perceivable. We simulate the attack on the Speech Commands Dataset consisting of 1s audio clips. During training, we use JEM to model a view of the input implemented by a randomly selected 610ms window. During inference, we combine all (40) possible views utilizing a generative part of JEM. The resulting system has slightly decreased accuracy but significantly increased resistance shown in multiple scenarios. Interestingly, replacing JEM with a standard discriminative model (Disc) provides increased resistance with a lesser effect compared to JEM but maintains accuracy. We introduce an extension motivated by semi-supervised training that further improves JEM but not Disc. JEM can also benefit from Gaussian noise during evaluation. Martin Sustek, Sonal Joshi, Henry Li, Thomas Thebaud, Jesús Villalba 0001, Sanjeev Khudanpur, Najim Dehak |
ASRU | 7 |
| 2023 | Clustering Unsupervised Representations as Defense Against Poisoning Attacks on Speech Commands Classification SystemabstractPoisoning attacks entail attackers intentionally tampering with training data. In this paper, we consider a dirty-label poisoning attack scenario on a speech commands classification system. The threat model assumes that certain utterances from one of the classes (source class) are poisoned by superimposing a trigger on it, and its label is changed to another class selected by the attacker (target class). We propose a filtering defense against such an attack. First, we use DIstillation with NO labels (DINO) to learn unsupervised representations for all the training examples. Next, we use K-means and LDA to cluster these representations. Finally, we keep the utterances with the most repeated label in their cluster for training and discard the rest. For a 10% poisoned source class, we demonstrate a drop in attack success rate from 99.75% to 0.25%. We test our defense against a variety of threat models, including different target and source classes, as well as trigger variations. Thomas Thebaud, Sonal Joshi, Henry Li, Martin Sustek, Jesús Villalba 0001, Sanjeev Khudanpur, Najim Dehak |
ASRU | 7 |
| 2023 | Advances in Language Recognition in Low Resource African Languages: The JHU-MIT Submission for NIST LRE22
Jesús Villalba 0001, Jonas Borgstrom, Maliha Jahan, Saurabh Kataria 0001, L. Paola García-Perera, Pedro A. Torres-Carrasquillo, Najim Dehak |
INTERSPEECH | 7 |
| 2023 | Segmental SpeechCLIP: Utilizing Pretrained Image-text Models for Audio-Visual Learning
Saurabhchand Bhati, Jesús Villalba 0001, Laureano Moro-Velázquez, Thomas Thebaud, Najim Dehak |
INTERSPEECH | 5 |
| 2023 | Do Phonatory Features Display Robustness to Characterize Parkinsonian Speech Across Corpora?
Anna Favaro, Tianyu Cao 0003, Thomas Thebaud, Jesús Villalba 0001, Ankur A. Butala, Najim Dehak, Laureano Moro-Velázquez |
INTERSPEECH | 6 |
| 2023 | Self-FiLM: Conditioning GANs with self-supervised representations for bandwidth extension based speaker recognitionabstractSpeech super-resolution/Bandwidth Extension (BWE) can improve downstream tasks like Automatic Speaker Verification (ASV).We introduce a simple novel technique called Self-FiLM to inject self-supervision into existing BWE models via Feature-wise Linear Modulation.We hypothesize that such information captures domain/environment information, which can give zero-shot generalization.Self-FiLM Conditional GAN (CGAN) gives 18% relative improvement in Equal Error Rate and 8.5% in minimum Decision Cost Function using state-ofthe-art ASV system on SRE21 test.We further by 1) deep feature loss from time-domain models and 2) re-training of data2vec 2.0 models on naturalistic wideband (VoxCeleb) and telephone data (SRE Superset etc.).Lastly, we integrate selfsupervision with CycleGAN to present a completely unsupervised solution that matches the semi-supervised performance. Saurabh Kataria 0001, Jesús Villalba 0001, Laureano Moro-Velázquez, Thomas Thebaud, Najim Dehak |
INTERSPEECH | 5 |
| 2023 | DuTa-VC: A Duration-aware Typical-to-atypical Voice Conversion Approach with Diffusion Probabilistic Model
Helin Wang, Thomas Thebaud, Jesús Villalba 0001, Myra Sydnor, Becky Lammers, Najim Dehak, Laureano Moro-Velázquez |
INTERSPEECH | 6 |
| 2022 | Non-contrastive self-supervised learning of utterance-level speech representations
Raghavendra Pappagari, Piotr Zelasko, Laureano Moro-Velázquez, Jesús Villalba 0001, Najim Dehak |
INTERSPEECH | 6 |
| 2022 | Defense against Adversarial Attacks on Hybrid Speech Recognition System using Adversarial Fine-tuning with Denoiser
Sonal Joshi, Saurabh Kataria 0001, Yiwen Shao, Piotr Zelasko, Jesús Villalba 0001, Sanjeev Khudanpur, Najim Dehak |
INTERSPEECH | 7 |
| 2022 | AdvEst: Adversarial Perturbation Estimation to Classify and Detect Adversarial Attacks against Speaker IdentificationabstractAdversarial attacks pose a severe security threat to the state-ofthe-art speaker identification systems, thereby making it vital to propose countermeasures against them.Building on our previous work that used representation learning to classify and detect adversarial attacks, we propose an improvement to it using Ad-vEst, a method to estimate adversarial perturbation.First, we prove our claim that training the representation learning network using adversarial perturbations as opposed to adversarial examples (consisting of the combination of clean signal and adversarial perturbation) is beneficial because it eliminates nuisance information.At inference time, we use a time-domain denoiser to estimate the adversarial perturbations from adversarial examples.Using our improved representation learning approach to obtain attack embeddings (signatures), we evaluate their performance for three applications: known attack classification, attack verification, and unknown attack detection.We show that common attacks in the literature (Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD), Carlini-Wagner (CW) with different Lp threat models) can be classified with an accuracy of ∼ 96%.We also detect unknown attacks with an equal error rate (EER) of ∼9%, which is absolute improvement of ∼12% from our previous work. Sonal Joshi, Saurabh Kataria 0001, Jesús Villalba 0001, Najim Dehak |
INTERSPEECH | 4 |
| 2022 | Joint domain adaptation and speech bandwidth extension using time-domain GANs for speaker verificationabstractSpeech systems developed for a particular choice of acoustic domain and sampling frequency do not translate easily to others.The usual practice is to learn domain adaptation and bandwidth extension models independently.Contrary to this, we propose to learn both tasks together.Particularly, we learn to map narrowband conversational telephone speech to wideband microphone speech.We developed parallel and non-parallel learning solutions which utilize both paired and unpaired data.First, we first discuss joint and disjoint training of multiple generative models for our tasks.Then, we propose a two-stage learning solution where we use a pre-trained domain adaptation system for pre-processing in bandwidth extension training.We evaluated our schemes on a Speaker Verification downstream task.We used the JHU-MIT experimental setup for NIST SRE21, which comprises SRE16, SRE-CTS Superset and SRE21.Our results provide the first evidence that learning both tasks is better than learning just one.On SRE16, our best system achieves 22% relative improvement in Equal Error Rate w.r.t. a direct learning baseline and 8% w.r.t. a strong bandwidth expansion system. Saurabh Kataria 0001, Jesús Villalba 0001, Laureano Moro-Velázquez, Najim Dehak |
INTERSPEECH | 4 |
| 2022 | End-to-End Neural Speaker Diarization with an Iterative Refinement of Non-Autoregressive Attention-based Attractors
Magdalena Rybicka, Jesús Villalba 0001, Najim Dehak, Konrad Kowalczyk |
INTERSPEECH | 3 |
| 2022 | Chunking Defense for Adversarial Attacks on ASR
Yiwen Shao, Jesús Villalba 0001, Sonal Joshi, Saurabh Kataria 0001, Sanjeev Khudanpur, Najim Dehak |
INTERSPEECH | 6 |
| 2022 | Vsameter: Evaluation of a New Open-Source Tool to Measure Vowel Space Area and Related MetricsabstractVowel space area (VSA) is an applicable metric for studying speech production deficits and intelligibility. Previous works suggest that the VSA accounts for almost 50% of the intelligibility variance, being an essential component of global intelligibility estimates. However, almost no study publishes a tool to estimate VSA automatically with publicly available codes. In this paper, we propose an open-source tool called VSAmeter to measure VSA and vowel articulation index (VAI) automatically and validate it with the VSA and VAI obtained from a dataset in which the formants and phone segments have been annotated manually. The results show that VSA and VAI values obtained by our proposed method strongly correlate with those generated by manually extracted F1 and F2 and alignments. Such a method can be utilized in speech applications, e.g., the automatic measurement of VAI for the evaluation of speakers with dysarthria. Tianyu Cao 0003, Laureano Moro-Velázquez, Piotr Zelasko, Jesús Villalba 0001, Najim Dehak |
SLT | 5 |
| 2022 | A Multi-Modal Array of Interpretable Features to Evaluate Language and Speech Patterns in Different Neurological DisordersabstractSpeech-based automatic approaches for evaluating neurological disorders (NDs) depend on feature extraction before the classification pipeline. It is preferable for these features to be interpretable to facilitate their development as diagnostic tools. This study focuses on the analysis of interpretable features obtained from the spoken responses of 88 subjects with NDs and controls (CN). Subjects with NDs have Alzheimer's disease (AD), Parkinson's disease (PD), or Parkinson's disease mimics (PDM). We configured three complementary sets of features related to cognition, speech, and language, and conducted a statistical analysis to examine which features differed between NDs and CN. Results suggested that features capturing response informativeness, reaction times, vocabulary richness, and syntactic complexity provided separability between AD and CN. Similarly, fundamental frequency variability helped differentiate PD from CN, while the number of salient informational units PDM from CN. Anna Favaro, Chelsie Motley, Tianyu Cao 0003, Miguel Iglesias, Ankur A. Butala, Esther S. Oh, Robert D. Stevens 0002, Jesús Villalba 0001, Najim Dehak, Laureano Moro-Velázquez |
SLT | 9 |
| 2022 | Textual Data Augmentation for Arabic-English Code-Switching Speech RecognitionabstractThe pervasiveness of intra-utterance code-switching (CS) in spoken content requires that speech recognition (ASR) systems handle mixed language. Designing a CS-ASR system has many challenges, mainly due to data scarcity, grammatical structure complexity, and domain mismatch. The most common method for addressing CS is to train an ASR system with the available transcribed CS speech, along with monolingual data. In this work, we propose a zero-shot learning methodology for CS-ASR by augmenting the monolingual data with artificially generating CS text. We based our approach on random lexical replacements and Equivalence Constraint (EC) while exploiting aligned translation pairs to generate random and grammatically valid CS content. Our empirical results show a 65.5% relative reduction in language model perplexity, and 7.7% in ASR WER on two ecologically valid CS test sets. The human evaluation of the generated text using EC suggests that more than 80% is of adequate quality. Amir Hussein, Shammur Absar Chowdhury, Ahmed Abdelali, Najim Dehak, Ahmed Ali 0002, Sanjeev Khudanpur |
SLT | 4 |
| 2022 | Discovering phonetic inventories with crosslingual automatic speech recognition
Piotr Zelasko, Siyuan Feng 0001, Laureano Moro-Velázquez, Ali Abavisani, Saurabhchand Bhati, Odette Scharenborg, Mark Hasegawa-Johnson, Najim Dehak |
Comput. Speech Lang. | 8 |
| 2022 | Unsupervised Speech Segmentation and Variable Rate Representation Learning Using Segmental Contrastive Predictive CodingabstractTypically, unsupervised segmentation of speech into the phone- and word-like units are treated as separate tasks and are often done via different methods which do not fully leverage the inter-dependence of the two tasks. Here, we unify them and propose a technique that can jointly perform both, showing that these two tasks indeed benefit from each other. Recent attempts employ self-supervised learning, such as contrastive predictive coding (CPC), where the next frame is predicted given past context. However, CPC only looks at the audio signal’s frame-level structure. We overcome this limitation with a segmental contrastive predictive coding (SCPC) framework to model the signal structure at a higher level, e.g., phone level. A convolutional neural network learns frame-level representation from the raw waveform via noise-contrastive estimation (NCE). A differentiable boundary detector finds variable-length segments, which are then used to optimize a segment encoder via NCE to learn segment representations. The differentiable boundary detector allows us to train frame-level and segment-level encoders jointly. Experiments show that our single model outperforms existing phone and word segmentation methods on TIMIT and Buckeye datasets. We analyze the impact of the threshold on boundary detector performance, and our results suggest that automatically learning the boundary threshold can be as effective as manually tuning that threshold. We discover that phone class impacts the boundary detection performance, and the boundaries between successive vowels or semivowels are the most difficult. Finally, we use SCPC to extract speech features at the segment level rather than at the uniformly spaced frame level (e.g., 10 ms) and produce variable rate representations that change according to the contents of the utterance. We can lower the feature extraction rate from the typical 100 Hz to as low as 14.5 Hz on average while still outperforming the hand-crafted features such as MFCC on the linear phone classification task. Saurabhchand Bhati, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Joint Prediction of Truecasing and Punctuation for Conversational Speech in Low-Resource ScenariosabstractCapitalization and punctuation are important cues for comprehending written texts and conversational transcripts. Yet, many ASR systems do not produce punctuated and case-formatted speech transcripts. We propose to use a multi-task system that can exploit the relations between casing and punctuation to improve their prediction performance. Whereas text data for predicting punctuation and truecasing is seemingly abundant, we argue that written text resources are inadequate as training data for conversational models. We quantify the mismatch between written and conversational text domains by comparing the joint distributions of punctuation and word cases, and by testing our model cross-domain. Further, we show that by training the model in the written text domain and then transfer learning to conversations, we can achieve reasonable performance with less data. Raghavendra Pappagari, Piotr Zelasko, Agnieszka Mikolajczyk, Piotr Pezik, Najim Dehak |
ASRU | 5 |
| 2021 | Beyond Isolated Utterances: Conversational Emotion RecognitionabstractSpeech emotion recognition is the task of recognizing the speaker's emotional state given a recording of their utterance. While most of the current approaches focus on inferring emotion from isolated utterances, we argue that this is not sufficient to achieve conversational emotion recognition (CER) which deals with recognizing emotions in conversations. In this work, we propose several approaches for CER by treating it as a sequence labeling task. We investigated transformer architecture for CER and, compared it with ResNet-34 and BiLSTM architectures in both contextual and contextless scenarios using IEMOCAP corpus. Based on the inner workings of the self-attention mechanism, we proposed DiverseCatAugment (DCA), an augmentation scheme, which improved the transformer model performance by an absolute 3.3% micro-f1 on conversations and 3.6% on isolated utterances. We further enhanced the performance by introducing an interlocutor-aware transformer model where we learn a dictionary of interlocutor index embeddings to exploit diarized conversations. Raghavendra Pappagari, Piotr Zelasko, Jesús Villalba 0001, Laureano Moro-Velázquez, Najim Dehak |
ASRU | 5 |
| 2021 | Focus on the Present: A Regularization Method for the ASR Source-Target Attention LayerabstractThis paper introduces a novel method to diagnose the source-target attention in state-of-the-art end-to-end speech recognition models with joint connectionist temporal classification (CTC) and attention training. Our method is based on the fact that both, CTC and source-target attention, are acting on the same encoder representations. To understand the functionality of the attention, CTC is applied to compute the token posteriors given the attention outputs. We found that the source-target attention heads are able to predict several tokens ahead of the current one. Inspired by the observation, a new regularization method is proposed which leverages CTC to make source-target attention more focused on the frames corresponding to the output token being predicted by the decoder. Experiments reveal stable improvements up to 7% and 13% relatively with the proposed regularization on TED-LIUM 2 and Librispeech. Nanxin Chen, Piotr Zelasko, Jesús Villalba 0001, Najim Dehak |
ICASSP | 4 |
| 2021 | Improving Reconstruction Loss Based Speaker Embedding in Unsupervised and Semi-Supervised ScenariosabstractText-to-speech (TTS) models trained to minimize the spectrogram reconstruction loss can learn speaker embeddings without explicit speaker identity supervision, unlike x-vector speaker identification (SID) systems. Leveraging this way of speaker embedding learning can be useful in unsupervised or semi-supervised scenarios where non, or only some, of the training data have speaker labels. Thus, in this paper, we evaluate speaker embeddings learned by training the spectrogram prediction network under unsupervised and semi-supervised scenarios. We experimented with different data sampling strategies. The best one was sampling two different segments from the same utterance, namely A and B, where the spectrogram of B is predicted given the B phone sequence and the speaker embedding extracted from A. This method improved by 3.4% relative in EER, compared to using the same utterance for both A and B without segmenting. In the unsupervised scenario, the best speaker embedding outperformed i-vectors, the state-of-the-art unsupervised speaker embedding, in speaker verification by 12.9% relative in EER. We observed high correlation between reconstruction loss and speaker embedding quality. In the semi-supervised scenario, having more unlabeled data in training led to a better performance in speaker verification. Adding 5314 unlabeled speakers to 800 labeled speakers improved EER by 10.8 % relative. Piotr Zelasko, Jesús Villalba 0001, Najim Dehak |
ICASSP | 4 |
| 2021 | How Phonotactics Affect Multilingual and Zero-Shot ASR PerformanceabstractThe idea of combining multiple languages’ recordings to train a single automatic speech recognition (ASR) model brings the promise of the emergence of universal speech representation. Recently, a Transformer encoder-decoder model has been shown to leverage multilingual data well in IPA transcriptions of languages presented during training. However, the representations it learned were not successful in zero-shot transfer to unseen languages. Because that model lacks an explicit factorization of the acoustic model (AM) and language model (LM), it is unclear to what degree the performance suffered from differences in pronunciation or the mismatch in phono-tactics. To gain more insight into the factors limiting zero-shot ASR transfer, we replace the encoder-decoder with a hybrid ASR system consisting of a separate AM and LM. Then, we perform an extensive evaluation of monolingual, multilingual, and crosslingual (zero-shot) acoustic and language models on a set of 13 phonetically diverse languages. We show that the gain from modeling crosslingual phonotactics is limited, and imposing a too strong model can hurt the zero-shot transfer. Furthermore, we find that a multilingual LM hurts a multilingual ASR system’s performance, and retaining only the target language’s phonotactic data in LM training is preferable. Siyuan Feng 0001, Piotr Zelasko, Laureano Moro-Velázquez, Ali Abavisani, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
ICASSP | 7 |
| 2021 | Perceptual Loss Based Speech Denoising with an Ensemble of Audio Pattern Recognition and Self-Supervised ModelsabstractDeep learning based speech denoising still suffers from the challenge of improving perceptual quality of enhanced signals. We introduce a generalized framework called Perceptual Ensemble Regularization Loss (PERL) built on the idea of perceptual losses. Perceptual loss discourages distortion to certain speech properties and we analyze it using six large-scale pre-trained models: speaker classification, acoustic model, speaker embedding, emotion classification, and two self-supervised speech encoders (PASE+, wav2vec 2.0). We first build a strong baseline (w/o PERL) using Conformer Transformer Networks on the popular enhancement benchmark called VCTK-DEMAND. Using auxiliary models one at a time, we find acoustic event and self-supervised model PASE+ to be most effective. Our best model (PERL-AE) only uses acoustic event model (utilizing AudioSet) to outperform state-of-the-art methods on major perceptual metrics. To explore if denoising can leverage full framework, we use all networks but find that our seven-loss formulation suffers from the challenges of Multi-Task Learning. Finally, we report a critical observation that state-of-the-art Multi-Task weight learning methods cannot outperform hand tuning, perhaps due to challenges of domain mismatch and weak complementarity of losses. Saurabh Kataria 0001, Jesús Villalba 0001, Najim Dehak |
ICASSP | 3 |
| 2021 | CopyPaste: An Augmentation Method for Speech Emotion RecognitionabstractData augmentation is a widely used strategy for training robust machine learning models. It partially alleviates the problem of limited data for tasks like speech emotion recognition (SER), where collecting data is expensive and challenging. This study proposes CopyPaste, a perceptually motivated novel augmentation procedure for SER. Assuming that the presence of emotions other than neutral dictates a speaker’s overall perceived emotion in a recording, concatenation of an emotional (emotion E) and a neutral utterance can still be labeled with emotion E. We hypothesize that SER performance can be improved using these concatenated utterances in model training. To verify this, three CopyPaste schemes are tested on two deep learning models: one trained independently and another using transfer learning from an x-vector model, a speaker recognition model. We observed that all three CopyPaste schemes improve SER performance on all the three datasets considered: MSP-Podcast, Crema-D, and IEMOCAP. Additionally, CopyPaste performs better than noise augmentation and, using them together improves the SER performance further. Our experiments on noisy test sets suggested that CopyPaste is effective even in noisy test conditions. Raghavendra Pappagari, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak |
ICASSP | 5 |
| 2021 | Align or attend? Toward More Efficient and Accurate Spoken Word Discovery Using Speech-to-Image RetrievalabstractMultimodal word discovery (MWD) is often treated as a byproduct of the speech-to-image retrieval problem. However, our theoretical analysis shows that some kind of alignment/attention mechanism is crucial for a MWD system to learn meaningful word-level representation. We verify our theory by conducting retrieval and word discovery experiments on MSCOCO and Flickr8k, and empirically demonstrate that both neural MT with self-attention and statistical MT achieve word discovery scores that are superior to those of a state-of-the-art neural retrieval system, outperforming it by 2% and 5% alignment F1 scores respectively. Liming Wang 0003, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
ICASSP | 5 |
| 2021 | Segmental Contrastive Predictive Coding for Unsupervised Word SegmentationabstractAutomatic detection of phoneme or word-like units is one of the core objectives in zero-resource speech processing. Recent attempts employ self-supervised training methods, such as contrastive predictive coding (CPC), where the next frame is predicted given past context. However, CPC only looks at the audio signal's frame-level structure. We overcome this limitation with a segmental contrastive predictive coding (SCPC) framework that can model the signal structure at a higher level e.g. at the phoneme level. In this framework, a convolutional neural network learns frame-level representation from the raw waveform via noise-contrastive estimation (NCE). A differentiable boundary detector finds variable-length segments, which are then used to optimize a segment encoder via NCE to learn segment representations. The differentiable boundary detector allows us to train frame-level and segment-level encoders jointly. Typically, phoneme and word segmentation are treated as separate tasks. We unify them and experimentally show that our single model outperforms existing phoneme and word segmentation methods on TIMIT and Buckeye datasets. We analyze the impact of boundary threshold and when is the right time to include the segmental loss in the learning process. Saurabhchand Bhati, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak |
Interspeech | 5 |
| 2021 | Align-Denoise: Single-Pass Non-Autoregressive Speech Recognition
Nanxin Chen, Piotr Zelasko, Laureano Moro-Velázquez, Jesús Villalba 0001, Najim Dehak |
Interspeech | 5 |
| 2021 | WaveGrad 2: Iterative Refinement for Text-to-Speech SynthesisabstractThis paper introduces WaveGrad 2, a non-autoregressive generative model for text-to-speech synthesis. WaveGrad 2 is trained to estimate the gradient of the log conditional density of the waveform given a phoneme sequence. The model takes an input phoneme sequence, and through an iterative refinement process, generates an audio waveform. This contrasts to the original WaveGrad vocoder which conditions on mel-spectrogram features, generated by a separate model. The iterative refinement process starts from Gaussian noise, and through a series of refinement steps (e.g., 50 steps), progressively recovers the audio sequence. WaveGrad 2 offers a natural way to trade-off between inference speed and sample quality, through adjusting the number of refinement steps. Experiments show that the model can generate high fidelity audio, approaching the performance of a state-of-the-art neural TTS system. We also report various ablation studies over different model configurations. Audio samples are available at this https URL. Nanxin Chen, Yu Zhang 0033, Heiga Zen, Ron J. Weiss, Mohammad Norouzi 0002, Najim Dehak |
Interspeech | 6 |
| 2021 | Deep Feature CycleGANs: Speaker Identity Preserving Non-Parallel Microphone-Telephone Domain Adaptation for Speaker VerificationabstractWith the increase in the availability of speech from varied domains, it is imperative to use such out-of-domain data to improve existing speech systems. Domain adaptation is a prominent pre-processing approach for this. We investigate it for adapt microphone speech to the telephone domain. Specifically, we explore CycleGAN-based unpaired translation of microphone data to improve the x-vector/speaker embedding network for Telephony Speaker Verification. We first demonstrate the efficacy of this on real challenging data and then, to improve further, we modify the CycleGAN formulation to make the adaptation task-specific. We modify CycleGAN's identity loss, cycle-consistency loss, and adversarial loss to operate in the deep feature space. Deep features of a signal are extracted from an auxiliary (speaker embedding) network and, hence, preserves speaker identity. Our 3D convolution-based Deep Feature Discriminators (DFD) show relative improvements of 5-10% in terms of equal error rate. To dive deeper, we study a challenging scenario of pooling (adapted) microphone and telephone data with data augmentations and telephone codecs. Finally, we highlight the sensitivity of CycleGAN hyper-parameters and introduce a parameter called probability of adaptation. Saurabh Kataria 0001, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak |
Interspeech | 5 |
| 2021 | Automatic Detection and Assessment of Alzheimer Disease Using Speech and Language Technologies in Low-Resource Scenarios
Raghavendra Pappagari, Sonal Joshi, Laureano Moro-Velázquez, Piotr Zelasko, Jesús Villalba 0001, Najim Dehak |
Interspeech | 7 |
| 2021 | Spine2Net: SpineNet with Res2Net and Time-Squeeze-and-Excitation Blocks for Speaker Recognition
Magdalena Rybicka, Jesús Villalba 0001, Piotr Zelasko, Najim Dehak, Konrad Kowalczyk |
Interspeech | 4 |
| 2021 | Representation Learning to Classify and Detect Adversarial Attacks Against Speaker and Speech Recognition SystemsabstractAdversarial attacks have become a major threat for machine learning applications. There is a growing interest in studying these attacks in the audio domain, e.g, speech and speaker recognition; and find defenses against them. In this work, we focus on using representation learning to classify/detect attacks w.r.t. the attack algorithm, threat model or signal-to-adversarial-noise ratio. We found that common attacks in the literature can be classified with accuracies as high as 90%. Also, representations trained to classify attacks against speaker identification can be used also to classify attacks against speaker verification and speech recognition. We also tested an attack verification task, where we need to decide whether two speech utterances contain the same attack. We observed that our models did not generalize well to attack algorithms not included in the attack representation model training. Motivated by this, we evaluated an unknown attack detection task. We were able to detect unknown attacks with equal error rates of about 19%, which is promising. Jesús Villalba 0001, Sonal Joshi, Piotr Zelasko, Najim Dehak |
Interspeech | 4 |
| 2021 | Non-Autoregressive Transformer for Speech RecognitionabstractVery deep transformers outperform conventional bidirectional long short-term memory networks for automatic speech recognition (ASR) by a significant margin. However, being autoregressive models, their computational complexity is still a prohibitive factor in their deployment into production systems. To amend this problem, we study two different non-autoregressive transformer structures for ASR: Audio-Conditional Masked Language Model (A-CMLM) and Audio-Factorized Masked Language Model (A-FMLM). When training these frameworks, the decoder input tokens are randomly replaced by special mask tokens. Then, the network is optimized to predict the masked tokens by taking both the unmasked context tokens and the input speech into consideration. During inference, we start from all masked tokens and the network iteratively predicts missing tokens based on partial results. A new decoding strategy is proposed as an example, which starts from the most confident predictions to the rest. Results on Mandarin (AISHELL), Japanese (CSJ), English (LibriSpeech) benchmarks show promising results to train such a non-autoregressive network for ASR. Especially in AISHELL, the proposed method outperformed the Kaldi ASR system and matched the performance of the state-of-the-art autoregressive transformer with 7× speedup. Nanxin Chen, Shinji Watanabe 0001, Jesús Villalba 0001, Piotr Zelasko, Najim Dehak |
IEEE Signal Process. Lett. | 5 |
| 2021 | What Helps Transformers Recognize Conversational Structure? Importance of Context, Punctuation, and Labels in Dialog Act RecognitionabstractAbstract Dialog acts can be interpreted as the atomic units of a conversation, more fine-grained than utterances, characterized by a specific communicative function. The ability to structure a conversational transcript as a sequence of dialog acts—dialog act recognition, including the segmentation—is critical for understanding dialog. We apply two pre-trained transformer models, XLNet and Longformer, to this task in English and achieve strong results on Switchboard Dialog Act and Meeting Recorder Dialog Act corpora with dialog act segmentation error rates (DSER) of 8.4% and 14.2%. To understand the key factors affecting dialog act recognition, we perform a comparative analysis of models trained under different conditions. We find that the inclusion of a broader conversational context helps disambiguate many dialog act classes, especially those infrequent in the training data. The presence of punctuation in the transcripts has a massive effect on the models’ performance, and a detailed analysis reveals specific segmentation patterns observed in its absence. Finally, we find that the label set specificity does not affect dialog act segmentation performance. These findings have significant practical implications for spoken language understanding applications that depend heavily on a good-quality segmentation being available. Piotr Zelasko, Raghavendra Pappagari, Najim Dehak |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | Study of Pre-Processing Defenses Against Adversarial Attacks on State-of-the-Art Speaker Recognition SystemsabstractAdversarial examples are designed to fool the speaker recognition (SR) system by adding a carefully crafted human-imperceptible noise to the speech signals. Posing a severe security threat to state-of-the-art SR systems, it becomes vital to deep-dive and study their vulnerabilities. Moreover, it is of greater importance to propose countermeasures that can protect the systems against these attacks. Addressing these concerns, we first investigated how state-of-the-art x-vector based SR systems are affected by white-box adversarial attacks, i.e., when the adversary has full knowledge of the system. x-Vector based SR systems are evaluated against white-box adversarial attacks common in the literature like fast gradient sign method (FGSM), basic iterative method (BIM)–a.k.a. iterative-FGSM–, projected gradient descent (PGD), and Carlini-Wagner (CW) attack. To mitigate against these attacks, we investigated four pre-processing defenses which do not need adversarial examples during training. The four pre-processing defenses–viz. randomized smoothing, DefenseGAN, variational autoencoder (VAE), and Parallel Wave-GAN vocoder (PWG) are compared against the baseline defense of adversarial training. Performing powerful adaptive white-box adversarial attack (i.e., when the adversary has full knowledge of the system, including the defense), our conclusions indicate that SR systems were extremely vulnerable under BIM, PGD, and CW attacks. Among the proposed pre-processing defenses, PWG combined with randomized smoothing offers the most protection against the attacks, with accuracy averaging 93% compared to 52% in the undefended system and an absolute improvement > 90% for BIM attacks with L∞ > 0.001 and CW attack. Sonal Joshi, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2020 | Feature Enhancement with Deep Feature Losses for Speaker VerificationabstractSpeaker Verification still suffers from the challenge of generalization to novel adverse environments. We leverage on the recent advancements made by deep learning based speech enhancement and propose a feature-domain supervised denoising based solution. We propose to use Deep Feature Loss which optimizes the enhancement network in the hidden activation space of a pre-trained auxiliary speaker embedding network. We experimentally verify the approach on simulated and real data. A simulated testing setup is created using various noise types at different SNR levels. For evaluation on real data, we choose BabyTrain corpus which consists of children recordings in uncontrolled environments. We observe consistent gains in every condition over the state-of-the-art augmented Factorized-TDNN x-vector system. On BabyTrain corpus, we observe relative gains of 10.38% and 12.40% in minDCF and EER respectively. Saurabh Kataria 0001, Phani S. Nidadavolu, Jesús Villalba 0001, Nanxin Chen, L. Paola García-Perera, Najim Dehak |
ICASSP | 6 |
| 2020 | Using X-Vectors to Automatically Detect Parkinson's Disease from SpeechabstractThe promise of new neuroprotective treatments to stop or slow the advance of Parkinson's Disease (PD) urges for new biomarkers or detection schemes that can deliver a faster diagnosis. Given that speech is affected by PD, the combination of deep neural networks and speech processing can provide automatic detection schemes. Accordingly, in this study we analyze for the first time a new state-of-the-art speaker recognition technique, x-Vectors, in a different scenario: the automatic detection of PD from speech. The proposed approach is compared with another speaker recognition technique, i-Vectors, employed in previous works and used as baseline in this study. A corpus with 43 PD patients and 46 control speakers was used to evaluate the performance of these two techniques at two sampling frequencies: 8 and 16 kHz.The x-Vector approach provided the best results in terms of accuracy and AUC reaching values of 90% and 0.94, respectively. Consequently, results suggest that speaker embeddings obtained using deep neural networks are successful extracting acoustic information relative to patterns in articulation, prosody and/or phonation common in persons with PD. Laureano Moro-Velázquez, Jesús Villalba 0001, Najim Dehak |
ICASSP | 3 |
| 2020 | Unsupervised Feature Enhancement for Speaker VerificationabstractThe task of making speaker verification systems robust to adverse scenarios remains a challenging and an active area of research. We developed an unsupervised feature enhancement approach in log-filter bank space with the end goal of improving speaker verification performance. We experimented with using both real speech recorded in adverse environments and degraded speech obtained by simulation to train the enhancement systems. The effectiveness of this approach was shown by testing on several real, simulated noisy, and reverberant test sets. The approach yielded significant improvements on both real and simulated sets when data augmentation was not used in speaker verification pipeline. We also experimented with training the x-vector and PLDA systems with enhanced augmented features instead of augmented features and observed better performance on real test conditions (4.2% relative improvement in minDCF on SRI). Phani S. Nidadavolu, Saurabh Kataria 0001, Jesús Villalba 0001, L. Paola García-Perera, Najim Dehak |
ICASSP | 5 |
| 2020 | X-Vectors Meet Emotions: A Study On Dependencies Between Emotion and Speaker RecognitionabstractIn this work, we explore the dependencies between speaker recognition and emotion recognition. We first show that knowledge learned for speaker recognition can be reused for emotion recognition through transfer learning. Then, we show the effect of emotion on speaker recognition. For emotion recognition, we show that using a simple linear model is enough to obtain good performance on the features extracted from pre-trained models such as the x-vector model. Then, we improve emotion recognition performance by finetuning for emotion classification. We evaluated our experiments on three different types of datasets: IEMOCAP, MSP-Podcast, and Crema-D. By fine-tuning, we obtained 30.40%, 7.99%, and 8.61% absolute improvement on IEMOCAP, MSP-Podcast, and Crema-D respectively over baseline model with no pre-training. Finally, we present results on the effect of emotion on speaker verification. We observed that speaker verification performance is prone to changes in test speaker emotions. We found that trials with angry utterances performed worst in all three datasets. We hope our analysis will initiate a new line of research in the speaker recognition community. Raghavendra Pappagari, Tianzi Wang, Jesús Villalba 0001, Nanxin Chen, Najim Dehak |
ICASSP | 5 |
| 2020 | Punctuation Prediction in Spontaneous Conversations: Can We Mitigate ASR Errors with Retrofitted Word Embeddings?abstractAutomatic Speech Recognition (ASR) systems introduce word errors, which often confuse punctuation prediction models, turning punctuation restoration into a challenging task. These errors usually take the form of homonyms. We show how retrofitting of the word embeddings on the domain-specific data can mitigate ASR errors. Our main contribution is a method for better alignment of homonym embeddings and the validation of the presented method on the punctuation prediction task. We record the absolute improvement in punctuation prediction accuracy between 6.2% (for question marks) to 9% (for periods) when compared with the state-of-the-art model. Lukasz Augustyniak, Piotr Szymanski, Mikolaj Morzy, Piotr Zelasko, Adrian Szymczak, Jan Mizgajski, Yishay Carmiel, Najim Dehak |
INTERSPEECH | 8 |
| 2020 | Self-Expressing Autoencoders for Unsupervised Spoken Term DiscoveryabstractUnsupervised spoken term discovery consists of two tasks: finding the acoustic segment boundaries and labeling acoustically similar segments with the same labels. We perform segmentation based on the assumption that the frame feature vectors are more similar within a segment than across the segments. Therefore, for strong segmentation performance, it is crucial that the features represent the phonetic properties of a frame more than other factors of variability. We achieve this via a self-expressing autoencoder framework. It consists of a single encoder and two decoders with shared weights. The encoder projects the input features into a latent representation. One of the decoders tries to reconstruct the input from these latent representations and the other from the self-expressed version of them. We use the obtained features to segment and cluster the speech data. We evaluate the performance of the proposed method in the Zero Resource 2020 challenge unit discovery task. The proposed system consistently outperforms the baseline, demonstrating the usefulness of the method in learning representations. Saurabhchand Bhati, Jesús Villalba 0001, Piotr Zelasko, Najim Dehak |
INTERSPEECH | 4 |
| 2020 | Learning Speaker Embedding from Text-to-SpeechabstractZero-shot multi-speaker Text-to-Speech (TTS) generates target speaker voices given an input text and the corresponding speaker embedding.In this work, we investigate the effectiveness of the TTS reconstruction objective to improve representation learning for speaker verification.We jointly trained endto-end Tacotron 2 TTS and speaker embedding networks in a self-supervised fashion.We hypothesize that the embeddings will contain minimal phonetic information since the TTS decoder will obtain that information from the textual input.TTS reconstruction can also be combined with speaker classification to enhance these embeddings further.Once trained, the speaker encoder computes representations for the speaker verification task, while the rest of the TTS blocks are discarded.We investigated training TTS from either manual or ASR-generated transcripts.The latter allows us to train embeddings on datasets without manual transcripts.We compared ASR transcripts and Kaldi phone alignments as TTS inputs, showing that the latter performed better due to their finer resolution.Unsupervised TTS embeddings improved EER by 2.06% absolute with regard to i-vectors for the LibriTTS dataset.TTS with speaker classification loss improved EER by 0.28% and 0.73% absolutely from a model using only speaker classification loss in LibriTTS and Voxceleb1 respectively. Piotr Zelasko, Jesús Villalba 0001, Shinji Watanabe 0001, Najim Dehak |
INTERSPEECH | 5 |
| 2020 | Using State of the Art Speaker Recognition and Natural Language Processing Technologies to Detect Alzheimer's Disease and Assess its Severity
Raghavendra Pappagari, Laureano Moro-Velázquez, Najim Dehak |
INTERSPEECH | 4 |
| 2020 | x-Vectors Meet Adversarial Attacks: Benchmarking Adversarial Robustness in Speaker Verification
Jesús Villalba 0001, Yuekai Zhang, Najim Dehak |
INTERSPEECH | 3 |
| 2020 | That Sounds Familiar: An Analysis of Phonetic Representations Transfer Across LanguagesabstractOnly a handful of the world's languages are abundant with the resources that enable practical applications of speech processing technologies. One of the methods to overcome this problem is to use the resources existing in other languages to train a multilingual automatic speech recognition (ASR) model, which, intuitively, should learn some universal phonetic representations. In this work, we focus on gaining a deeper understanding of how general these representations might be, and how individual phones are getting improved in a multilingual setting. To that end, we select a phonetically diverse set of languages, and perform a series of monolingual, multilingual and crosslingual (zero-shot) experiments. The ASR is trained to recognize the International Phonetic Alphabet (IPA) token sequences. We observe significant improvements across all languages in the multilingual setting, and stark degradation in the crosslingual setting, where the model, among other errors, considers Javanese as a tone language. Notably, as little as 10 hours of the target language training data tremendously reduces ASR error rates. Our analysis uncovered that even the phones that are unique to a single language can benefit greatly from adding training data from other languages - an encouraging result for the low-resource speech community. Piotr Zelasko, Laureano Moro-Velázquez, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
INTERSPEECH | 5 |
| 2020 | Black-Box Attacks on Spoofing Countermeasures Using Transferability of Adversarial Examples
Yuekai Zhang, Ziyan Jiang, Jesús Villalba 0001, Najim Dehak |
INTERSPEECH | 4 |
| 2020 | rVAD: An unsupervised segment-based robust voice activity detection method
Zheng-Hua Tan, Achintya Kumar Sarkar, Najim Dehak |
Comput. Speech Lang. | 3 |
| 2020 | State-of-the-art speaker recognition with neural network embeddings in NIST SRE18 and Speakers in the Wild evaluations
Jesús Villalba 0001, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, L. Paola García-Perera, Fred Richardson, Réda Dehak, Pedro A. Torres-Carrasquillo, Najim Dehak |
Comput. Speech Lang. | 12 |
| 2020 | Analysis of the Effects of Supraglottal Tract Surgical Procedures in Automatic Speaker Recognition PerformanceabstractThis article evaluates the impact in the performance of state-of-the-art automatic speaker recognition schemes of three surgical procedures modifying the supraglottal tract structures of speakers. To do so, a new corpus (Cuco) was recorded, containing the speech of 107 speakers before and after surgery. Speakers were divided into four groups depending on the type of surgery: tonsillectomy, functional endoscopy sinus surgery (FESS), septoplasty, and controls. The analyzed speaker recognition schemes were i-vectors, i-vectors with supervised Universal Background Model, i-vectors employing Time-delay Deep Neural Networks and x-vectors. In all cases, probabilistic linear discriminant analysis was employed in the back-end. Results show changes in the speech of patients who underwent tonsillectomy or FESS after surgery in contrast to controls or patients who had a septoplasty, where not significant variations are observed. These changes increase the Equal Error Rate (EER) of the analyzed speaker recognition schemes for the septoplasty and FESS groups when employing enrollment data recorded before the surgery. Moreover, surgery has a similar influence in the speech of female and male speakers with respect to the analyzed schemes. In consequence, results suggest that it is advisable to update the speaker's enrollment speech after three months following supraglottal tract surgery to ensure that the effects of the operation and post-operative recovery period do not influence the performance of the automatic speaker recognition systems. Laureano Moro-Velázquez, Estefanía Hernández-García, Jorge Andrés Gómez García, Juan Ignacio Godino-Llorente, Najim Dehak |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2019 | Low-Resource Domain Adaptation for Speaker Recognition Using Cycle-GansabstractCurrent speaker recognition technology provides great performance with the x-vector approach. However, performance decreases when the evaluation domain is different from the training domain, an issue usually addressed with domain adaptation approaches. Recently, unsupervised domain adaptation using cycle-consistent Generative Adversarial Networks (CycleGAN) has received a lot of attention. Cycle-GAN learn mappings between features of two domains given non-parallel data. We investigate their effectiveness in low resource scenario i.e. when limited amount of target domain data is available for adaptation, a case unexplored in previous works. We experiment with two adaptation tasks: microphone to telephone and a novel reverberant to clean adaptation with the end goal of improving speaker recognition performance. Number of speakers present in source and target domains are 7000 and 191 respectively. By adding noise to the target domain during CycleGAN training, we were able to achieve better performance compared to the adaptation system whose CycleGAN was trained on a larger target data. On reverberant to clean adaptation task, our models improved EER by 18.3% relative on VOiCES dataset compared to a system trained on clean data. They also slightly improved over the state-of-the-art Weighted Prediction Error (WPE) de-reverberation algorithm. Phani S. Nidadavolu, Saurabh Kataria 0001, Jesús Villalba 0001, Najim Dehak |
ASRU | 4 |
| 2019 | Hierarchical Transformers for Long Document ClassificationabstractBERT, which stands for Bidirectional Encoder Representations from Transformers, is a recently introduced language representation model based upon the transfer learning paradigm. We extend its fine-tuning procedure to address one of its major limitations - applicability to inputs longer than a few hundred words, such as transcripts of human call conversations. Our method is conceptually simple. We segment the input into smaller chunks and feed each of them into the base model. Then, we propagate each output through a single recurrent layer, or another transformer, followed by a softmax activation. We obtain the final classification decision after the last segment has been consumed. We show that both BERT extensions are quick to fine-tune and converge after as little as 1 epoch of training on a small, domain-specific data set. We successfully apply them in three different tasks involving customer call satisfaction prediction and topic classification, and obtain a significant improvement over the baseline models in two of them. Raghavendra Pappagari, Piotr Zelasko, Jesús Villalba 0001, Yishay Carmiel, Najim Dehak |
ASRU | 5 |
| 2019 | Language Model Integration Based on Memory Control for Sequence to Sequence Speech RecognitionabstractIn this paper, we explore several new schemes to train a seq2seq model to integrate a pre-trained language model (LM). Our proposed fusion methods focus on the memory cell state and the hidden state in the seq2seq decoder long short-term memory (LSTM), and the memory cell state is updated by the LM unlike the prior studies. This means the memory retained by the main seq2seq would be adjusted by the external LM. These fusion methods have several variants depending on the architecture of this memory cell update and the use of memory cell and hidden states which directly affects the final label inference. We performed the experiments to show the effectiveness of the proposed methods in a mono-lingual ASR setup on the Librispeech corpus and in a transfer learning setup from a multilingual ASR (MLASR) base model to a low-resourced language. In Librispeech, our best model improved WER by 3.7%, 2.4% for test clean, test other relatively to the shallow fusion baseline, with multilevel decoding. In transfer learning from an MLASR base model to the IARPA Babel Swahili model, the best scheme improved the transferred model on eval set by 9.9%, 9.8% in CER, WER relatively to the 2-stage transfer baseline. Shinji Watanabe 0001, Takaaki Hori, Murali Karthick Baskar, Hirofumi Inaguma, Jesús Villalba 0001, Najim Dehak |
ICASSP | 7 |
| 2019 | Attentive Filtering Networks for Audio Replay Attack DetectionabstractAn attacker may use a variety of techniques to fool an automatic speaker verification system into accepting them as a genuine user. Anti-spoofing methods meanwhile aim to make the system robust against such attacks. The ASVspoof 2017 Challenge focused specifically on replay attacks, with the intention of measuring the limits of replay attack detection as well as developing countermeasures against them. In this work, we propose our replay attacks detection system - Attentive Filtering Network, which is composed of an attention-based filtering mechanism that enhances feature representations in both the frequency and time domains, and a ResNet-based classifier. We show that the network enables us to visualize the automatically acquired feature representations that are helpful for spoofing detection. Attentive Filtering Network attains an evaluation EER of 8.99% on the ASVspoof 2017 Version 2.0 dataset. With system fusion, our best system further obtains a 30% relative improvement over the ASVspoof 2017 enhanced baseline system. Cheng-I Lai, Alberto Abad, Korin Richmond, Junichi Yamagishi, Najim Dehak, Simon King 0001 |
ICASSP | 5 |
| 2019 | Investigation on Neural Bandwidth Extension of Telephone Speech for Improved Speaker RecognitionabstractWe extend our previous work on training mixed-bandwidth (BW) speaker recognition system by predicting missing information in upperband (UB) of upsampled telephone speech. Mixed-BW systems combine speech from narrowband (NB) and wideband (WB) speech corpora by basic upsampling of NB speech with low-pass filter interpolator, resulting in no information loss in the original WB speech. In this work, we explore the usage of a deep residual full-convolutional neural network (CNN) and a bidirectional long short term memory (BLSTM) network along with a previously proposed deep neural network (DNN) for bandwidth extension (BWE) of NB telephone speech. Speaker recognition systems trained with bandwidth extended features improved in performance over mixed-BW and NB baseline systems. In terms of detection cost function (DCF), the CNN-BWE system improved by 10.78% and 15.96% (relative) in the Speakers In The Wild (SITW) eval core and assist-multi-speaker condition respectively w.r.t. the NB baseline; and improved by 3.21% and 4.13% w.r.t. to the mixed-BW baseline. Phani S. Nidadavolu, Vicente Iglesias, Jesús Villalba 0001, Najim Dehak |
ICASSP | 4 |
| 2019 | Cycle-GANs for Domain Adaptation of Acoustic Features for Speaker RecognitionabstractIt is well known that domain mismatch between the training and evaluation data hinders the performance of any machine learning system. Various factors contribute to domain mismatch. In speaker recognition systems, it mainly occurs due to the mismatch in recording conditions and language. Most speaker recognition corpora are telephone speech. Meanwhile, a few evaluation data sets like Speakers In The Wild (SITW) are microphone speech. In this work, we explore domain adaptation at acoustic feature level by learning feature mappings between domains using cycle consistent generative adversarial networks (cycle-GANs), without any parallel data between domains. Microphone features mapped to telephone domain are used to evaluate speaker recognition system trained only on telephone data. We achieved 9.37% and 2.82% relative improvement in equal error rate (EER) and detection cost function (DCF) on SITW eval set. Phani S. Nidadavolu, Jesús Villalba 0001, Najim Dehak |
ICASSP | 3 |
| 2019 | Unsupervised Acoustic Segmentation and Clustering Using Siamese Network EmbeddingsabstractUnsupervised discovery of acoustic units from the raw speech signal forms the core objective of zero-resource speech processing. It involves identifying the acoustic segment boundaries and consistently assigning unique labels to acoustically similar segments. In this work, the possible candidates for segment boundaries are identified in an unsupervised manner from the kernel Gram matrix computed from the Mel-frequency cepstral coefficients (MFCC). These segment boundary candidates are used to train a siamese network, that is intended to learn embeddings that minimize intrasegment distances and maximize the intersegment distances. The siamese embeddings capture phonetic information from longer contexts of the speech signal and enhance the intersegment discriminability. These properties make the siamese embeddings better suited for acoustic segmentation and clustering than the raw MFCC features. The Gram matrix computed from the siamese embeddings provides unambiguous evidence for boundary locations. The initial candidate boundaries are refined using this evidence, and siamese embeddings are extracted for the new acoustic segments. A graph growing approach is used to cluster the siamese embeddings, and a unique label is assigned to acoustically similar segments. The performance of the proposed method for acoustic segmentation and clustering is evaluated on Zero Resource 2017 database. Saurabhchand Bhati, Shekhar Nayak, K. Sri Rama Murty, Najim Dehak |
INTERSPEECH | 4 |
| 2019 | Tied Mixture of Factor Analyzers Layer to Combine Frame Level Representations in Neural Speaker Embeddings
Nanxin Chen, Jesús Villalba 0001, Najim Dehak |
INTERSPEECH | 3 |
| 2019 | ASSERT: Anti-Spoofing with Squeeze-Excitation and Residual NetworksabstractWe present JHU's system submission to the ASVspoof 2019 Challenge: Anti-Spoofing with Squeeze-Excitation and Residual neTworks (ASSERT).Anti-spoofing has gathered more and more attention since the inauguration of the ASVspoof Challenges, and ASVspoof 2019 dedicates to address attacks from all three major types: text-to-speech, voice conversion, and replay.Built upon previous research work on Deep Neural Network (DNN), ASSERT is a pipeline for DNN-based approach to anti-spoofing.ASSERT has four components: feature engineering, DNN models, network optimization and system combination, where the DNN models are variants of squeeze-excitation and residual networks.We conducted an ablation study of the effectiveness of each component on the ASVspoof 2019 corpus, and experimental results showed that ASSERT obtained more than 93% and 17% relative improvements over the baseline systems in the two sub-challenges in ASVspooof 2019, ranking ASSERT one of the top performing systems.Code and pretrained models will be made publicly available. Cheng-I Lai, Nanxin Chen, Jesús Villalba 0001, Najim Dehak |
INTERSPEECH | 4 |
| 2019 | Study of the Performance of Automatic Speech Recognition Systems in Speakers with Parkinson's DiseaseabstractParkinson’s Disease (PD) affects motor capabilities of patients, who in some cases need to use human-computer assistive technologies to regain independence. The objective of this work is to study in detail the differences in error patterns from state-of-the-art Automatic Speech Recognition (ASR) systems on speech from people with and without PD. Two different speech recognizers (attention-based end-to-end and Deep Neural Network - Hidden Markov Models hybrid systems) were trained on a Spanish language corpus and subsequently tested on speech from 43 speakers with PD and 46 without PD. The differences related to error rates, substitutions, insertions and deletions of characters and phonetic units between the two groups were analyzed, showing that the word error rate is 27% higher in speakers with PD than in control speakers, with a moderated correlation between that rate and the developmental stage of the disease. The errors were related to all manner classes, and were more pronounced in the vowel /u/. This study is the first to evaluate ASR systems’ responses to speech from patients at different stages of PD in Spanish. The analyses showed general trends but individual speech deficits must be studied in the future when designing new ASR systems for this population. Laureano Moro-Velázquez, Shinji Watanabe 0001, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
INTERSPEECH | 7 |
| 2019 | Improving Emotion Identification Using Phone Posteriors in Raw Speech Waveform Based DNN
Mousmita Sarma, Pegah Ghahremani, Daniel Povey, Nagendra Kumar Goel, Kandarpa Kumar Sarma, Najim Dehak |
INTERSPEECH | 6 |
| 2019 | MCE 2018: The 1st Multi-Target Speaker Detection and Identification Challenge EvaluationabstractThe Multi-target Challenge aims to assess how well current speech technology is able to determine whether or not a recorded utterance was spoken by one of a large number of blacklisted speakers. It is a form of multi-target speaker detection based on real-world telephone conversations. Data recordings are generated from call center customer-agent conversations. The task is to measure how accurately one can detect 1) whether a test recording is spoken by a blacklisted speaker, and 2) which specific blacklisted speaker was talking. This paper outlines the challenge and provides its baselines, results, and discussions. Suwon Shon, Najim Dehak, Douglas A. Reynolds, James R. Glass |
INTERSPEECH | 2 |
| 2019 | The JHU Speaker Recognition System for the VOiCES 2019 Challenge
David Snyder, Jesús Villalba 0001, Nanxin Chen, Daniel Povey, Gregory Sell, Najim Dehak, Sanjeev Khudanpur |
INTERSPEECH | 6 |
| 2019 | State-of-the-Art Speaker Recognition for Telephone and Video Speech: The JHU-MIT Submission for NIST SRE18
Jesús Villalba 0001, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, Fred Richardson, Suwon Shon, François Grondin, Réda Dehak, L. Paola García-Perera, Daniel Povey, Pedro A. Torres-Carrasquillo, Sanjeev Khudanpur, Najim Dehak |
INTERSPEECH | 16 |
| 2019 | Pretraining by Backtranslation for End-to-End ASR in Low-Resource SettingsabstractWe explore training attention-based encoder-decoder ASR in low-resource settings. These models perform poorly when trained on small amounts of transcribed speech, in part because they depend on having sufficient target-side text to train the attention and decoder networks. In this paper we address this shortcoming by pretraining our network parameters using only text-based data and transcribed speech from other languages. We analyze the relative contributions of both sources of data. Across 3 test languages, our text-based approach resulted in a 20% average relative improvement over a text-based augmentation technique without pretraining. Using transcribed speech from nearby languages gives a further 20-30% relative reduction in character error rate. Matthew Wiesner, Adithya Renduchintala, Shinji Watanabe 0001, Chunxi Liu, Najim Dehak, Sanjeev Khudanpur |
INTERSPEECH | 5 |
| 2018 | Measuring Uncertainty in Deep Regression Models: The Case of Age Estimation from SpeechabstractAge estimation from speech recently received a lot of attention. Approaches such as i-vectors and deep learning have been successfully applied to this task achieving great performance. However, one drawback of those methods is that they produce a hard age estimation without any kind of confidence measure about the quality of the prediction. Designing systems with the ability to provide a confidence measure about their output is extremely valuable for several applications where the cost of making bad decisions is worse than making no decision, e.g., forensics. In this paper, we propose a novel framework to jointly predict the age and its estimation uncertainty in a context of neural regression model. This model is trained using probabilistic fashion instead of using the classical minimum mean square error objective used for regression tasks. The probabilistic output corresponds to a Gaussian posterior. The proposed neural network will estimate both the posterior mean which corresponds to the predicted age and the variance which quantifies the uncertainty of the prediction. We evaluated our approach on two different datasets NIST SRE 2008 - 2010 and Switchboard. Nanxin Chen, Jesús Villalba 0001, Yishay Carmiel, Najim Dehak |
ICASSP | 4 |
| 2018 | Characterizing Performance of Speaker Diarization Systems on Far-Field Speech Using Standard MethodsabstractTo date, the bulk of research on speaker diarization has been conducted on telephone or near-field speech. As the need for technologies capable of handling conversational speech increases, it is necessary to establish the performance of state-of-the-art systems in this domain. In this work we evaluate the performance of an ivector/PLDA-based diarization system on the AMI Meeting Corpus, comparing performance on near-field, far-field, and signal-enhanced conditions. Matthew Maciejewski, David Snyder, Vimal Manohar, Najim Dehak, Sanjeev Khudanpur |
ICASSP | 4 |
| 2018 | Joint Verification-Identification in end-to-end Multi-Scale CNN Framework for Topic IdentificationabstractWe present an end-to-end multi-scale Convolutional Neural Network (CNN) framework for topic identification (topic ID). In this work, we examined multi -scale CNN for classification using raw text input. Topical word embeddings are learnt at multiple scales using parallel convolutional layers. A technique to integrate verification and identification objectives is examined to improve topic ID performance. With this approach, we achieved significant improvement in identification task. We evaluated our framework on two contrasting datasets: 20 newsgroups and Fisher. We obtained 92.93% accuracy on Fisher and 86.12% on 20 newsgroups, which to our know ledge are the best published results on these datasets at the moment. Raghavendra Pappagari, Jesús Villalba 0001, Najim Dehak |
ICASSP | 3 |
| 2018 | An Investigation of Non-linear i-vectors for Speaker Verification
Nanxin Chen, Jesús Villalba 0001, Najim Dehak |
INTERSPEECH | 3 |
| 2018 | Deep Neural Networks for Emotion Recognition Combining Audio and TranscriptsabstractIn this paper, we propose to improve emotion recognition by combining acoustic information and conversation transcripts.On the one hand, a LSTM network was used to detect emotion from acoustic features like f0, shimmer, jitter, MFCC, etc.On the other hand, a multi-resolution CNN was used to detect emotion from word sequences.This CNN consists of several parallel convolutions with different kernel sizes to exploit contextual information at different levels.A temporal pooling layer aggregates the hidden representations of different words into a unique sequence level embedding, from which we computed the emotion posteriors.We optimized a weighted sum of classification and verification losses.The verification loss tries to bring embeddings from same emotions closer while separating embeddings from different emotions.We also compared our CNN with state-of-the-art text-based hand-crafted features (e-vector).We evaluated our approach on the USC-IEMOCAP dataset as well as the dataset consisting of US English telephone speech.In the former, we used human-annotated transcripts while in the latter, we used ASR transcripts.The results showed fusing audio and transcript information improved unweighted accuracy by relative 24% for IEMOCAP and relative 3.4% for the telephone data compared to a single acoustic system. Raghavendra Pappagari, Purva Kulkarni, Jesús Villalba 0001, Yishay Carmiel, Najim Dehak |
INTERSPEECH | 6 |
| 2018 | Effectiveness of Single-Channel BLSTM Enhancement for Language IdentificationabstractThis paper proposes to apply deep neural network (DNN)-based single-channel speech enhancement (SE) to language identification. The 2017 language recognition evaluation (LRE17) introduced noisy audios from videos, in addition to the telephone conversation from past challenges. Because of that, adapting models from telephone speech to noisy speech from the video domain was required to obtain optimum performance. However, such adaptation requires knowledge of the audio domain and availability of in-domain data. Instead of adaptation, we propose to use a speech enhancement step to clean up the noisy audio as preprocessing for language identification. We used a bi-directional long short-term memory (BLSTM) neural network, which given log-Mel noisy features predicts a spectral mask indicating how clean each time-frequency bin is. The noisy spectrogram is multiplied by this predicted mask to obtain the enhanced magnitude spectrogram, and it is transformed back into the time domain by using the unaltered noisy speech phase. The experiments show significant improvement to language identification of noisy speech, for systems with and without domain adaptation, while preserving the identification performance in the telephone audio domain. In the best adapted state-of-the-art bottleneck i-vector system the relative improvement is 11.3% for noisy speech. Peter Sibbern Frederiksen, Jesús Villalba 0001, Shinji Watanabe 0001, Zheng-Hua Tan, Najim Dehak |
INTERSPEECH | 5 |
| 2018 | End-to-end Deep Neural Network Age Estimation
Pegah Ghahremani, Phani S. Nidadavolu, Nanxin Chen, Jesús Villalba 0001, Daniel Povey, Sanjeev Khudanpur, Najim Dehak |
INTERSPEECH | 7 |
| 2018 | Investigation on Bandwidth Extension for Speaker Recognition
Phani S. Nidadavolu, Cheng-I Lai, Jesús Villalba 0001, Najim Dehak |
INTERSPEECH | 4 |
| 2018 | Emotion Identification from Raw Speech Signals Using DNNs
Mousmita Sarma, Pegah Ghahremani, Daniel Povey, Nagendra Kumar Goel, Kandarpa Kumar Sarma, Najim Dehak |
INTERSPEECH | 6 |
| 2018 | Visualizing Phoneme Category Adaptation in Deep Neural NetworksabstractBoth human listeners and machines need to adapt their sound categories whenever a new speaker is encountered. This perceptual learning is driven by lexical information. The aim of this paper is two-fold: investigate whether a deep neural network-based (DNN) ASR system can adapt to only a few examples of ambiguous speech as humans have been found to do; investigate a DNN’s ability to serve as a model of human perceptual learning. Crucially, we do so by looking at intermediate levels of phoneme category adaptation rather than at the output level. We visualize the activations in the hidden layers of the DNN during perceptual learning. The results show that, similar to humans, DNN systems learn speaker-adapted phone category boundaries from a few labeled examples. The DNN adapts its category boundaries not only by adapting the weights of the output layer, but also by adapting the implicit feature maps computed by the hidden layers, suggesting the possibility that human perceptual learning might involve a similar nonlinear distortion of a perceptual space that is intermediate between the acoustic input and the phonological categories. Comparisons between DNNs and humans can thus provide valuable insights into the way humans process speech and improve ASR technology. Odette Scharenborg, Sebastian Tiesmeyer, Mark Hasegawa-Johnson, Najim Dehak |
INTERSPEECH | 4 |
| 2018 | Diarization is Hard: Some Experiences and Lessons Learned for the JHU Team in the Inaugural DIHARD Challenge
Gregory Sell, David Snyder, Alan McCree, Daniel Garcia-Romero, Jesús Villalba 0001, Matthew Maciejewski, Vimal Manohar, Najim Dehak, Daniel Povey, Shinji Watanabe 0001, Sanjeev Khudanpur |
INTERSPEECH | 8 |
| 2018 | Automatic Speech Recognition and Topic Identification from Speech for Almost-Zero-Resource Languages
Matthew Wiesner, Chunxi Liu, Lucas Ondel Yang, Craig Harman, Vimal Manohar, Jan Trmal, Zhongqiang Huang, Najim Dehak, Sanjeev Khudanpur |
INTERSPEECH | 8 |
| 2018 | Punctuation Prediction Model for Conversational SpeechabstractAn ASR system usually does not predict any punctuation or capitalization. Lack of punctuation causes problems in result presentation and confuses both the human reader andoff-the-shelf natural language processing algorithms. To overcome these limitations, we train two variants of Deep Neural Network (DNN) sequence labelling models - a Bidirectional Long Short-Term Memory (BLSTM) and a Convolutional Neural Network (CNN), to predict the punctuation. The models are trained on the Fisher corpus which includes punctuation annotation. In our experiments, we combine time-aligned and punctuated Fisher corpus transcripts using a sequence alignment algorithm. The neural networks are trained on Common Web Crawl GloVe embedding of the words in Fisher transcripts aligned with conversation side indicators and word time infomation. The CNNs yield a better precision and BLSTMs tend to have better recall. While BLSTMs make fewer mistakes overall, the punctuation predicted by the CNN is more accurate - especially in the case of question marks. Our results constitute significant evidence that the distribution of words in time, as well as pre-trained embeddings, can be useful in the punctuation prediction task. Piotr Zelasko, Piotr Szymanski, Jan Mizgajski, Adrian Szymczak, Yishay Carmiel, Najim Dehak |
INTERSPEECH | 6 |
| 2018 | Low-Resource Contextual Topic Identification on SpeechabstractIn topic identification (topic ID) on real-world unstructured audio, an audio instance of variable topic shifts is first broken into sequential segments, and each segment is independently classified. We first present a general purpose method for topic ID on spoken segments in low-resource languages, using a cascade of universal acoustic modeling, translation lexicons to English, and English-language topic classification. Next, instead of classifying each segment independently, we demonstrate that exploring the contextual dependencies across sequential segments can provide large improvements. In particular, we propose an attention-based contextual model which is able to leverage the contexts in a selective manner. We test both our contextual and non-contextual models on four LORELEI languages, and on all but one our attention-based contextual model significantly outperforms the context-independent models. Chunxi Liu, Matthew Wiesner, Shinji Watanabe 0001, Craig Harman, Jan Trmal, Najim Dehak, Sanjeev Khudanpur |
SLT | 6 |
| 2017 | Topic identification of spoken documents using unsupervised acoustic unit discoveryabstractThis paper investigates the application of unsupervised acoustic unit discovery for topic identification (topic ID) of spoken audio documents. The acoustic unit discovery method is based on a non-parametric Bayesian phone-loop model that segments a speech utterance into phone-like categories. The discovered phone-like (acoustic) units are further fed into the conventional topic ID framework. Using multilingual bottleneck features for the acoustic unit discovery, we show that the proposed method outperforms other systems that are based on cross-lingual phoneme recognizer. Santosh Kesiraju, Raghavendra Pappagari, Lucas Ondel Yang, Lukás Burget, Najim Dehak, Sanjeev Khudanpur, Jan Cernocký, Suryakanth V. Gangashetty |
ICASSP | 5 |
| 2017 | An empirical evaluation of zero resource acoustic unit discoveryabstractAcoustic unit discovery (AUD) is a process of automatically identifying a categorical acoustic unit inventory from speech and producing corresponding acoustic unit tokenizations. AUD provides an important avenue for unsupervised acoustic model training in a zero resource setting where expert-provided linguistic knowledge and transcribed speech are unavailable. Therefore, to further facilitate zero-resource AUD process, in this paper, we demonstrate acoustic feature representations can be significantly improved by (i) performing linear discriminant analysis (LDA) in an unsupervised self-trained fashion, and (ii) leveraging resources of other languages through building a multilingual bottleneck (BN) feature extractor to give effective cross-lingual generalization. Moreover, we perform comprehensive evaluations of AUD efficacy on multiple downstream speech applications, and their correlated performance suggests that AUD evaluations are feasible using different alternative language resources when only a subset of these evaluation resources can be available in typical zero resource applications. Chunxi Liu, Jinyi Yang, Santosh Kesiraju, Alena Rott, Lucas Ondel Yang, Pegah Ghahremani, Najim Dehak, Lukás Burget, Sanjeev Khudanpur |
ICASSP | 8 |
| 2017 | Multi-view representation learning via gcca for multimodal analysis of Parkinson's diseaseabstractInformation from different bio-signals such as speech, handwriting, and gait have been used to monitor the state of Parkinson's disease (PD) patients, however, all the multimodal bio-signals may not always be available. We propose a method based on multi-view representation learning via generalized canonical correlation analysis (GCCA) for learning a representation of features extracted from handwriting and gait that can be used as a complement to speech-based features. Three different problems are addressed: classification of PD patients vs. healthy controls, prediction of the neurological state of PD patients according to the UPDRS score, and the prediction of a modified version of the Frenchay dysarthria assessment (m-FDA). According to the results, the proposed approach is suitable to improve the results in the addressed problems, specially in the prediction of the UPDRS, and m-FDA scores. Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Raman Arora, Elmar Nöth, Najim Dehak, Heidi Christensen, Frank Rudzicz, Tobias Bocklet, Milos Cernak, Hamid R. Chinaei, Julius Hannink, Phani S. Nidadavolu, Maria Yancheva, Alyssa Vann, Nikolai Vogler |
ICASSP | 5 |
| 2017 | Evaluation of the Neurological State of People with Parkinson's Disease Using i-Vectors
Nicanor García, Juan Rafael Orozco-Arroyave, Luis Fernando D'Haro, Najim Dehak, Elmar Nöth |
INTERSPEECH | 4 |
| 2017 | The MIT-LL, JHU and LRDE NIST 2016 Speaker Recognition Evaluation System
Pedro A. Torres-Carrasquillo, Fred Richardson, Shahan C. Nercessian, Douglas E. Sturim, William M. Campbell, Youngjune Gwon, Swaroop Vattam, Najim Dehak, Sri Harish Reddy Mallidi, Phani S. Nidadavolu, Réda Dehak |
INTERSPEECH | 8 |
| 2017 | Tied Variational Autoencoder Backends for i-Vector Speaker Recognition
Jesús Villalba 0001, Niko Brümmer, Najim Dehak |
INTERSPEECH | 3 |
| 2016 | Automatic Dialect Detection in Arabic Broadcast SpeechabstractWe investigate different approaches for dialect identification in Arabic broadcast speech, using phonetic, lexical features obtained from a speech recognition system, and acoustic features using the i-vector framework. We studied both generative and discriminate classifiers, and we combined these features using a multi-class Support Vector Machine (SVM). We validated our results on an Arabic/English language identification task, with an accuracy of 100%. We used these features in a binary classifier to discriminate between Modern Standard Arabic (MSA) and Dialectal Arabic, with an accuracy of 100%. We further report results using the proposed method to discriminate between the five most widely used dialects of Arabic: namely Egyptian, Gulf, Levantine, North African, and MSA, with an accuracy of 52%. We discuss dialect identification errors in the context of dialect code-switching between Dialectal Arabic and MSA, and compare the error pattern between manually labeled data, and the output from our classifier. We also release the train and test data as standard corpus for dialect identification. Ahmed Ali 0002, Najim Dehak, Patrick Cardinal, Sameer Khurana, Sree Harsha Yella, James R. Glass, Peter Bell 0001, Steve Renals |
INTERSPEECH | 2 |
| 2016 | Exploiting Hidden-Layer Responses of Deep Neural Networks for Language RecognitionabstractAbstract : The most popular way to apply Deep Neural Network (DNN) for Language IDentification (LID) involves the extraction of bottleneck features from a network that was trained on automatic speech recognition task. These features are modeled using a classical I-vector system. Recently, a more direct DNN approach was proposed, it consists of estimating the language posteriors directly from a stacked frames input. The final decision score is based on averaging the scores for all the frames for a given speech segment. In this paper, we extended the direct DNN approach by modeling all hidden-layer activations rather than just averaging the output scores. One super-vector per utterance is formed by concatenating all hidden-layer responses. The dimensionality of this vector is then reduced using a Principal Component Analysis (PCA). The obtained reduce vector summarizes the most discriminative features for language recognition based on the trained DNNs. We evaluated this approach in NIST 2015 language recognition evaluation. The performances achieved by the proposed approach are very competitive to the classical I-vector baseline. Sri Harish Reddy Mallidi, Lukás Burget, Oldrich Plchot, Najim Dehak |
INTERSPEECH | 5 |
| 2016 | Native Language Detection Using the I-Vector Framework
Mohammed Senoussaoui, Patrick Cardinal, Najim Dehak, Alessandro L. Koerich |
INTERSPEECH | 3 |
| 2016 | On the Use of Acoustic Unit Discovery for Language RecognitionabstractIn this paper, we explore the use of large-scale acoustic unit discovery for language recognition. The deep neural network-based approaches that have achieved recent success in this task require transcribed speech and pronunciation dictionaries, which may be limited in availability and expensive to obtain. We aim to replace the need for such supervision via the unsupervised discovery of acoustic units. In this work, we present a parallelized version of a Bayesian nonparametric model from previous work and use it to learn acoustic units from a few hundred hours of multilingual data. These unit (or senone) sequences are then used as targets to train a deep neural network-based i-vector language recognition system. We find that a score-level fusion of our unsupervised system with an acoustic baseline can shrink the gap significantly between the baseline and a supervised benchmark system built using transcribed English. Subsequent experiments also show that an improved acoustic representation of the data can yield substantial performance gains and that language specificity is important for discovering meaningful acoustic units. We validate the generalizability of our proposed approach by presenting state-of-the-art results that exhibit similar trends on the NIST Language Recognition Evaluations from 2011 and 2015. Stephen H. Shum, David F. Harwath, Najim Dehak, James R. Glass |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Speaker adaptation using the i-vector technique for bottleneck featuresabstractDeep Neural Networks (DNN) have been largely used and successfully applied in the context of speaker independent Automatic Speech Recognition (ASR). However, these models are not easily adapted to model a specific speaker characteristic. Recently, one approach was proposed to address this issue, which consists of using the I-vector representation as input to the DNN. The I-vector is playing the role of providing information about the speaker as well as the environmental conditions for a given recording. This approach achieved a significant improvement in the context of a hybrid system of DNN combined with Hidden Markov Model (HMM). In this paper, we study the effect of speaker adaptation based on the I-vector framework in the context of stacked bottleneck features. These features, extracted from a second level of DNNs, are modelled by a classical Gaussian Mixture Model (GMM) ASR system. The proposed approach achieved an absolute WER improvement of 1.2% on an Arabic Broadcast news task. Index Terms: DNN, I-Vector, Bottleneck Features, Speech Recognition Patrick Cardinal, Najim Dehak, Yu Zhang 0033, James R. Glass |
INTERSPEECH | 2 |
| 2015 | A unified deep neural network for speaker and language recognitionabstractLearned feature representations and sub-phoneme posteriors from Deep Neural Networks (DNNs) have been used separately to produce significant performance gains for speaker and language recognition tasks. In this work we show how these gains are possible using a single DNN for both speaker and language recognition. The unified DNN approach is shown to yield substantial performance improvements on the the 2013 Domain Adaptation Challenge speaker recognition task (55% reduction in EER for the out-of-domain condition) and on the NIST 2011 Language Recognition Evaluation (48% reduction in EER for the 30s test condition). Fred Richardson, Douglas A. Reynolds, Najim Dehak |
INTERSPEECH | 3 |
| 2015 | Deep Neural Network Approaches to Speaker and Language RecognitionabstractThe impressive gains in performance obtained using deep neural networks (DNNs) for automatic speech recognition (ASR) have motivated the application of DNNs to other speech technologies such as speaker recognition (SR) and language recognition (LR). Prior work has shown performance gains for separate SR and LR tasks using DNNs for direct classification or for feature extraction. In this work we present the application of single DNN for both SR and LR using the 2013 Domain Adaptation Challenge speaker recognition (DAC13) and the NIST 2011 language recognition evaluation (LRE11) benchmarks. Using a single DNN trained for ASR on Switchboard data we demonstrate large gains on performance in both benchmarks: a 55% reduction in EER for the DAC13 out-of-domain condition and a 48% reduction in Cavg on the LRE11 30 s test condition. It is also shown that further gains are possible using score or feature fusion leading to the possibility of a single i-vector extractor producing state-of-the-art SR and LR performance. Fred Richardson, Douglas A. Reynolds, Najim Dehak |
IEEE Signal Process. Lett. | 3 |
| 2014 | Recent advances in ASR applied to an Arabic transcription system for Al-JazeeraabstractThis paper describes a detailed comparison of several state-of-the-art speech recognition techniques applied to a limited Ara-bic broadcast news dataset. The different approaches were all trained on 50 hours of transcribed audio from the Al-Jazeera news channel. The best results were obtained using i-vector-based speaker adaptation in a training scenario using the Min-imum Phone Error (MPE) criteria combined with sequential Deep Neural Network (DNN) training. We report results for two different types of test data: broadcast news reports, with a best word error rate (WER) of 17.86%, and a broadcast conver-sations with a best WER of 29.85%. The overall WER on this test set is 25.6%. Index Terms: Arabic, ASR system, Kaldi 1. Patrick Cardinal, Ahmed Ali 0002, Najim Dehak, Yu Zhang 0033, Tuka Al Hanai, James R. Glass, Stephan Vogel |
INTERSPEECH | 3 |
| 2014 | Limited labels for unlimited data: active learning for speaker recognitionabstractIn this paper, we attempt to quantify the amount of labeled data necessary to build a state-of-the-art speaker recognition system. We begin by using i-vectors and the cosine similarity metric to represent an unlabeled set of utterances, then obtain labels from a noiseless oracle in the form of pairwise queries. Finally, we use the resulting speaker clusters to train a PLDA scoring function, which is assessed on the 2010 NIST Speaker Recognition Evaluation. After presenting the initial results of an algorithm that sorts queries based on nearest-neighbor pairs, we develop techniques that further minimize the number of queries needed to obtain state-of-the-art performance. We show the generalizability of our methods in anecdotal fashion by applying our methods to two different distributions of utterances-per-speaker and, ultimately, find that the actual number of pairwise labels needed to obtain state-of-the-art results may be a mere fraction of the queries required to fully label the entire set of utterances. Index Terms: speaker recognition, i-vectors, active learning Stephen H. Shum, Najim Dehak, James R. Glass |
INTERSPEECH | 2 |
| 2014 | A complete KALDI recipe for building Arabic speech recognition systemsabstractIn this paper we present a recipe and language resources for training and testing Arabic speech recognition systems using the KALDI toolkit. We built a prototype broadcast news system using 200 hours GALE data that is publicly available through LDC. We describe in detail the decisions made in building the system: using the MADA toolkit for text normalization and vowelization; why we use 36 phonemes; how we generate pronunciations; how we build the language model. We report results using state-of-the-art modeling and decoding techniques. The scripts are released through KALDI and resources are made available on QCRI's language resources web portal. This is the first effort to share reproducible sizable training and testing results on MSA system. Ahmed Ali 0002, Patrick Cardinal, Najim Dehak, Stephan Vogel, James R. Glass |
SLT | 4 |
| 2014 | Non-Negative Factor Analysis of Gaussian Mixture Model Weight Adaptation for Language and Dialect RecognitionabstractRecent studies show that Gaussian mixture model (GMM) weights carry less, yet complimentary, information to GMM means for language and dialect recognition. However, state-of-the-art language recognition systems usually do not use this information. In this research, a non-negative factor analysis (NFA) approach is developed for GMM weight decomposition and adaptation. This modeling, which is conceptually simple and computationally inexpensive, suggests a new low-dimensional utterance representation method using a factor analysis similar to that of the i-vector framework. The obtained subspace vectors are then applied in conjunction with i-vectors to the language/dialect recognition problem. The suggested approach is evaluated on the NIST 2011 and RATS language recognition evaluation (LRE) corpora and on the QCRI Arabic dialect recognition evaluation (DRE) corpus. The assessment results show that the proposed adaptation method yields more accurate recognition results compared to three conventional weight adaptation approaches, namely maximum likelihood re-estimation, non-negative matrix factorization, and a subspace multinomial model. Experimental results also show that the intermediate-level fusion of i-vectors and NFA subspace vectors improves the performance of the state-of-the-art i-vector framework especially for the case of short utterances. Mohamad Hasan Bahari, Najim Dehak, Hugo Van hamme, Lukás Burget, Ahmed Ali 0002, James R. Glass |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Developing a speaker identification system for the DARPA RATS projectabstractThis paper describes the speaker identification (SID) system developed by the Patrol team for the first phase of the DARPA RATS (Robust Automatic Transcription of Speech) program, which seeks to advance state of the art detection capabilities on audio from highly degraded communication channels. We present results using multiple SID systems differing mainly in the algorithm used for voice activity detection (VAD) and feature extraction. We show that (a) unsupervised VAD performs as well supervised methods in terms of downstream SID performance, (b) noise-robust feature extraction methods such as CFCCs out-perform MFCC front-ends on noisy audio, and (c) fusion of multiple systems provides 24% relative improvement in EER compared to the single best system when using a novel SVM-based fusion algorithm that uses side information such as gender, language, and channel id. Oldrich Plchot, Spyridon Matsoukas, Pavel Matejka, Najim Dehak, Jeff Z. Ma, Sandro Cumani, Ondrej Glembek, Hynek Hermansky, Sri Harish Reddy Mallidi, Nima Mesgarani, Richard M. Schwartz, Mehdi Soufifar, Zheng-Hua Tan, Samuel Thomas 0001, Bing Zhang 0004, Xinhui Zhou |
ICASSP | 4 |
| 2013 | Bayesian distance metric learning on i-vector for speaker verificationabstractThesis (S.M.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 2013. Najim Dehak, James R. Glass |
INTERSPEECH | 2 |
| 2013 | New cosine similarity scorings to implement gender-independent speaker verificationabstractThis paper is a natural extension of our previous work on gender-independent speaker verification systems [1]. In a previous paper, we presented a solution to avoid using gender information in the Probabilistic Linear Discriminant Analysis (PLDA) without any loss of accuracy compared with a genderdependent base-line implementation. In this work, we propose two solutions to make a speaker verification system based on Cosine similarity independent of speaker gender. Our choice of the Cosine similarity is motivated by the fact that it is proved itself as a second state-of-the art- in parallel with PLDA- of i-vector based speaker verification systems. As measured by Equal Error Rate and min DCF’s, performance results on the extended telephone list coreext-coreext condition of SRE2010 1 show no performance decrease in gender-independent Cosine similarity system compared to gender-dependent one. Tests were also successful for genderindependent propositions on a cross gender list as done in [1]. Mohammed Senoussaoui, Patrick Kenny, Pierre Dumouchel, Najim Dehak |
INTERSPEECH | 4 |
| 2013 | Unsupervised Methods for Speaker Diarization: An Integrated and Iterative ApproachabstractIn speaker diarization, standard approaches typically perform speaker clustering on some initial segmentation before refining the segment boundaries in a re-segmentation step to obtain a final diarization hypothesis. In this paper, we integrate an improved clustering method with an existing re-segmentation algorithm and, in iterative fashion, optimize both speaker cluster assignments and segmentation boundaries jointly. For clustering, we extend our previous research using factor analysis for speaker modeling. In continuing to take advantage of the effectiveness of factor analysis as a front-end for extracting speaker-specific features (i.e., i-vectors), we develop a probabilistic approach to speaker clustering by applying a Bayesian Gaussian Mixture Model (GMM) to principal component analysis (PCA)-processed i-vectors. We then utilize information at different temporal resolutions to arrive at an iterative optimization scheme that, in alternating between clustering and re-segmentation steps, demonstrates the ability to improve both speaker cluster assignments and segmentation boundaries in an unsupervised manner. Our proposed methods attain results that are comparable to those of a state-of-the-art benchmark set on the multi-speaker CallHome telephone corpus. We further compare our system with a Bayesian nonparametric approach to diarization and attempt to reconcile their differences in both methodology and performance. Stephen H. Shum, Najim Dehak, Réda Dehak, James R. Glass |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Patrol Team Language Identification System for DARPA RATS P1 EvaluationabstractThis paper describes the language identification (LID) system developed by the Patrol team for the first phase of the DARPA RATS (Robust Automatic Transcription of Speech) program, which seeks to advance state of the art detection capabilities on audio from highly degraded communication channels. We show that techniques originally developed for LID on telephone speech (e.g., for the NIST language recognition evaluations) remain effective on the noisy RATS data, provided that careful consideration is applied when designing the training and development sets. In addition, we show significant improvements from the use of Wiener filtering, neural network based and language dependent i-vector modeling, and fusion. Index Terms: language identification, noisy speech. 1. Pavel Matejka, Oldrich Plchot, Mehdi Soufifar, Ondrej Glembek, Luis Fernando D'Haro, Karel Veselý, Frantisek Grézl, Jeff Z. Ma, Spyridon Matsoukas, Najim Dehak |
INTERSPEECH | 10 |
| 2012 | On the Use of Spectral and Iterative Methods for Speaker DiarizationabstractThis paper extends upon our previous work using i-vectors for speaker diarization. We examine the effectiveness of spectral clustering as an alternative to our previous approach using K-means clustering and adapt a previously-used heuristic to es-timate the number of speakers. Additionally, we consider an iterative optimization scheme and experiment with its ability to improve both cluster assignments and segmentation boundaries in an unsupervised manner. Our proposed methods attain re-sults similar to those of a state-of-the-art benchmark set on the multi-speaker CallHome telephone corpus. Stephen H. Shum, Najim Dehak, James R. Glass |
INTERSPEECH | 2 |
| 2011 | A channel-blind system for speaker verificationabstractThe majority of speaker verification systems proposed in the NIST speaker recognition evaluation are conditioned on the type of data to be processed: telephone or microphone. In this paper, we propose a new speaker verification system that can be applied to both types of data. This system, named blind system, is based on an extension of the total variability framework. Recognition results with the pro posed channel-independent system are comparable to state of the art systems that require conditioning on the channel type. Another ad vantage of our proposed system is that it allows for combining data from multiple channels in the same visualization in order to explore the effects of different microphones and collection environments. Najim Dehak, Zahi N. Karam, Douglas A. Reynolds, Réda Dehak, William M. Campbell, James R. Glass |
ICASSP | 1 |
| 2011 | Towards reduced false-alarms using cohortsabstractThe focus of the 2010 NIST Speaker Recognition Evaluation (SRE) [1] was the low false alarm regime of the detection error trade-off (DET) curve. This paper presents several approaches that specifically target this issue. It begins by highlighting the main problem with operating in the low-false alarm regime. Two sets of methods to tackle this issue are presented that require a large and diverse impostor set: the first set penalizes trials whose enrollment and test utterances are not nearest neighbors of each other while the second takes an adaptive score normalization approach similar to TopNorm [2] and ATNorm [3]. Zahi N. Karam, William M. Campbell, Najim Dehak |
ICASSP | 3 |
| 2011 | The MIT LL 2010 speaker recognition evaluation system: Scalable language-independent speaker recognitionabstractResearch in the speaker recognition community has continued to ad dress methods of mitigating variational nuisances. Telephone and auxiliary-microphone recorded speech emphasize the need for a ro bust way of dealing with unwanted variation. The design of recent 2010 NIST-SRE Speaker Recognition Evaluation (SRE) reflects this research emphasis. In this paper, we present the MIT submission applied to the tasks of the 2010 NIST-SRE with two main goals- language-independent scalable modeling and robust nuisance mitigation. For modeling, exclusive use of inner product-based and cepstral systems produced a language-independent computationally scalable system. For robustness, systems that captured spectral and prosodic information, modeled nuisance subspaces using multiple novel methods, and fused scores of multiple systems were implemented. The performance of the system is presented on a subset of the NIST SRE 2010 core tasks. Douglas E. Sturim, William M. Campbell, Najim Dehak, Zahi N. Karam, Alan McCree, Douglas A. Reynolds, Fred Richardson, Pedro A. Torres-Carrasquillo, Stephen H. Shum |
ICASSP | 3 |
| 2011 | Language Recognition via i-vectors and Dimensionality ReductionabstractIn this paper, a new language identification system is presented based on the total variability approach previously developed in the field of speaker identification. Various techniques are em-ployed to extract the most salient features in the lower dimen-sional i-vector space and the system developed results in excel-lent performance on the 2009 LRE evaluation set without the need for any post-processing or backend techniques. Additional performance gains are observed when the system is combined with other acoustic systems. Najim Dehak, Pedro A. Torres-Carrasquillo, Douglas A. Reynolds, Réda Dehak |
INTERSPEECH | 1 |
| 2011 | Exploiting Intra-Conversation Variability for Speaker DiarizationabstractIn this paper, we propose a new approach to speaker diariza-tion based on the Total Variability approach to speaker verifica-tion. Drawing on previous work done in applying factor anal-ysis priors to the diarization problem, we arrive at a simplified approach that exploits intra-conversation variability in the To-tal Variability space through the use of Principal Component Analysis (PCA). Using our proposed methods, we demonstrate the ability to achieve state-of-the-art performance (0.9 % DER) in the diarization of summed-channel telephone data from the NIST 2008 SRE. Stephen H. Shum, Najim Dehak, Ekapol Chuangsuwanich, Douglas A. Reynolds, James R. Glass |
INTERSPEECH | 2 |
| 2011 | Front-End Factor Analysis for Speaker VerificationabstractThis paper presents an extension of our previous work which proposes a new speaker representation for speaker verification. In this modeling, a new low-dimensional speaker- and channel-dependent space is defined using a simple factor analysis. This space is named the total variability space because it models both speaker and channel variabilities. Two speaker verification systems are proposed which use this new representation. The first system is a support vector machine-based system that uses the cosine kernel to estimate the similarity between the input data. The second system directly uses the cosine similarity as the final decision score. We tested three channel compensation techniques in the total variability space, which are within-class covariance normalization (WCCN), linear discriminate analysis (LDA), and nuisance attribute projection (NAP). We found that the best results are obtained when LDA is followed by WCCN. We achieved an equal error rate (EER) of 1.12% and MinDCF of 0.0094 using the cosine distance scoring on the male English trials of the core condition of the NIST 2008 Speaker Recognition Evaluation dataset. We also obtained 4% absolute EER improvement for both-gender trials on the 10 s-10 s condition compared to the classical joint factor analysis scoring. Najim Dehak, Patrick Kenny, Réda Dehak, Pierre Dumouchel, Pierre Ouellet |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Support vector machines and Joint Factor Analysis for speaker verificationabstractThis article presents several techniques to combine between support vector machines (SVM) and joint factor analysis (JFA) model for speaker verification. In this combination, the SVMs are applied to different sources of information produced by the JFA. These informations are the Gaussian mixture model supervectors and speakers and common factors. We found that using SVM in JFA factors gave the best results especially when within class covariance normalization method is applied in order to compensate for the channel effect. The new combination results are comparable to other classical JFA scoring techniques. Najim Dehak, Patrick Kenny, Réda Dehak, Ondrej Glembek, Pierre Dumouchel, Lukás Burget, Valiantsina Hubeika, Fabio Castaldo |
ICASSP | 1 |
| 2009 | Comparison of scoring methods used in speaker recognition with Joint Factor AnalysisabstractThe aim of this paper is to compare different log-likelihood scoring methods, that different sites used in the latest state-of-the-art Joint Factor Analysis (JFA) Speaker Recognition systems. The algorithms use various assumptions and have been derived from various approximations of the objective functions of JFA. We compare the techniques in terms of speed and performance. We show, that approximations of the true log-likelihood ratio (LLR) may lead to significant speedup without any loss in performance. Ondrej Glembek, Lukás Burget, Najim Dehak, Niko Brümmer, Patrick Kenny |
ICASSP | 3 |
| 2009 | Support vector machines versus fast scoring in the low-dimensional total variability space for speaker verificationabstractThis paper presents a new speaker verification system architecture based on Joint Factor Analysis (JFA) as feature extractor. In this modeling, the JFA is used to define a new low-dimensional space named the total variability factor space, instead of both channel and speaker variability spaces for the classical JFA. The main contribution in this approach, is the use of the cosine kernel in the new total factor space to design two different systems: the first system is Support Vector Machines based, and the second one uses directly this kernel as a decision score. This last scoring method makes the process faster and less computation complex compared to others classical methods. We tested several intersession compensation methods in total factors, and we found that the combination of Linear Discriminate Analysis and Within Class Covariance Normalization achieved the best performance. We achieved a remarkable results using fast scoring method based only on cosine kernel especially for male trials, we yield an EER of 1.12% and MinDCF of 0.0094 on the English trials of the NIST 2008 SRE dataset. Index Terms: Total variability space, cosine kernel, fast scoring, support vector machines. Najim Dehak, Réda Dehak, Patrick Kenny, Niko Brümmer, Pierre Ouellet, Pierre Dumouchel |
INTERSPEECH | 1 |
| 2009 | Cepstral and long-term features for emotion recognitionabstractIn this paper, we describe systems that were developed for the Open Performance Sub-Challenge of the INTERSPEECH 2009 Emotion Challenge. We participate in both two-class and fiveclass emotion detection. For the two-class problem, the best performance is obtained by logistic regression fusion of three systems. These systems use short- and long-term speech features. Fusion allowed to an absolute improvement of 2:6% on the unweighted recall value compared with [1]. For the fiveclass problem, we submitted two individual systems: cepstral GMM vs. long-term GMM-UBM. The best result comes from a cepstral GMM and produces an absolute improvement of 3:5% compared to [6]. Pierre Dumouchel, Najim Dehak, Yazid Attabi, Réda Dehak, Narjès Boufaden |
INTERSPEECH | 2 |
| 2008 | Development of the primary CRIM system for the NIST 2008 speaker recognition evaluationabstractWe describe how we modified the CRIM factor analysis speaker verification system to handle the new cross-channel conditi ons encountered in the 2008 NIST speaker recognition evaluation. Using the 2006 evaluation data for development, we obtained results on a broad spectrum of test conditions that are uniformly bette r than the best results that have been published in the literature. Index Terms: speaker verification, factor analysis Patrick Kenny, Najim Dehak, Pierre Ouellet, Vishwa Gupta, Pierre Dumouchel |
INTERSPEECH | 2 |
| 2008 | A Study of Interspeaker Variability in Speaker VerificationabstractWe propose a new approach to the problem of estimating the hyperparameters which define the interspeaker variability model in joint factor analysis. We tested the proposed estimation technique on the NIST 2006 speaker recognition evaluation data and obtained 10%–15% reductions in error rates on the core condition and the extended data condition (as measured both by equal error rates and the NIST detection cost function). We show that when a large joint factor analysis model is trained in this way and tested on the core condition, the extended data condition and the cross-channel condition, it is capable of performing at least as well as fusions of multiple systems of other types. (The comparisons are based on the best results on these tasks that have been reported in the literature.) In the case of the cross-channel condition, a factor analysis model with 300 speaker factors and 200 channel factors can achieve equal error rates of less than 3.0%. This is a substantial improvement over the best results that have previously been reported on this task. Patrick Kenny, Pierre Ouellet, Najim Dehak, Vishwa Gupta, Pierre Dumouchel |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Linear and non linear kernel GMM supervector machines for speaker verificationabstractThis paper presents a comparison between Support Vector Machines (SVM) speaker verification systems based on linear and non linear kernels defined in GMM supervector space. We describe how these kernel functions are related and we show how the nuisance attribute projection (NAP) technique can be used with both of these kernels to deal with the session variability problem. We demonstrate the importance of GMM model normalization (M-Norm) especially for the non linear kernel. All our experiments were performed on the core condition of NIST 2006 speaker recognition evaluation (all trials). Our best results (an equal error rate of 6.3%) were obtained using NAP and GMM model normalization with the non linear kernel. Réda Dehak, Najim Dehak, Patrick Kenny, Pierre Dumouchel |
INTERSPEECH | 2 |
| 2007 | Continuous prosodic features and formant modeling with joint factor analysis for speaker verificationabstractIn this paper, we introduced the use of formants contours with prosodic contours based on pitch and energy for speaker recognition. These contours are modeled on continuous manners by using the Legendre polynomials on basic unit which represents syllables. The parameters extracted from the Legendre polynomials coefficients plus the syllables duration are modeled with Gaussian Mixture Models (GMM). Factor analysis is used to treat the speaker and channel variability. The results obtained on the core condition of NIST 2006 speaker recognition evaluation show that the use of formant with prosodic information gives an absolute improvement of approximately 3% on equal error rate (EER) compared with the results obtained by prosodic informations alone. However when the formants and the prosodic system scores are fused with a state of the art cepstral joint factor analysis system, we obtain equivalent results to the results obtained when we fused system based on prosodic features alone with the same cepstral joint factor analysis system. This fusion gives a relative improvement of 8.0% (all trials) and 12.0% (English only) on EER compared to cepstral system alone. Najim Dehak, Patrick Kenny, Pierre Dumouchel |
INTERSPEECH | 1 |
| 2007 | Modeling Prosodic Features With Joint Factor Analysis for Speaker VerificationabstractIn this paper, we introduce the use of continuous prosodic features for speaker recognition, and we show how they can be modeled using joint factor analysis. Similar features have been successfully used in language identification. These prosodic features are pitch and energy contours spanning a syllable-like unit. They are extracted using a basis consisting of Legendre polynomials. Since the feature vectors are continuous (rather than discrete), they can be modeled using a standard Gaussian mixture model (GMM). Furthermore, speaker and session variability effects can be modeled in the same way as in conventional joint factor analysis. We find that the best results are obtained when we use the information about the pitch, energy, and the duration of the unit all together. Testing on the core condition of NIST 2006 speaker recognition evaluation data gives an equal error rate of 16.6% and 14.6%, with prosodic features alone, for all trials and English-only trials, respectively. When the prosodic system is fused with a state-of-the-art cepstral joint factor analysis system, we obtain a relative improvement of 8% (all trials) and 12% (English only) compared to the cepstral system alone. Najim Dehak, Pierre Dumouchel, Patrick Kenny |
IEEE Trans. Speech Audio Process. | 1 |