VLDB 2026 Research / reviewers in the wild / expert
Jahangir Alam 0001
dblp:76/10401-1 · also Md. Jahangir Alam 0001
· DBLP profile ↗
74ranked-venue papers
19as first author
28since 2021 · last 2025
0000-0003-1081-3665ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 18 first-author · 23 since 2021Artificial intelligence and machine learning · 51 · 15 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 3 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AdaptiveDrop: A Simple Adaptive Label Noise Filtering Scheme for Enhanced Self-supervised Speaker VerificationabstractUsing clustering-driven annotations to train a neural network can be a tricky task because of label noise. In this paper, we propose a dynamic and adaptive label noise cleansing method, called AdaptiveDrop which combines both label noise filtering and correction simultaneously in cascade to combine their advantages. Contrary to other label noise filtering approaches, our method filters noisy samples on the fly from an early stage of training. We also provide a variant that incorporates sub-centers per each class for enhanced robustness to label noise by continuously tracking the dominant sub-centers via a dictionary table. AdaptiveDrop is a simple general-purpose method, performed end-to-end in only one stage of training, can be integrated with any loss function, and does not require training from scratch on the cleansed dataset. We show through extensive ablation studies for the self-supervised speaker verification task that our method is effective, benefits from long epochs of iterative filtering and provides consistent performance gains across various loss functions and real-world pseudo-labels. Abderrahim Fathan, Jahangir Alam 0001 |
ICASSP | 3 |
| 2025 | LAVViT: Latent Audio-Visual Vision Transformers for Speaker VerificationabstractRecently, Vision Transformers (ViTs) have shown remarkable success in various computer vision applications. In this work, we have explored the potential of ViTs, pre-trained on visual data, for audio-visual speaker verification. To cope with the challenges of large-scale training, we introduce the Latent Audio-Visual Vision Transformer (LAVViT) adapters, where we exploit the existing pre-trained models on visual data without fine-tuning their parameters and train only the parameters of LAVViT adapters. The LAVViT adapters are injected into every layer of the ViT architecture to effectively fuse the audio and visual modalities using a small set of latent tokens, forming an attention bottleneck, thereby reducing the quadratic computational cost of cross-attention across the modalities. The proposed approach has been evaluated on the Voxceleb1 dataset and shows promising performance using only a few trainable parameters. Code is available at https://github.com/praveena2j/LAVViT Gnana Praveen Rajasekhar, Jahangir Alam 0001 |
ICASSP | 2 |
| 2025 | Text-dependent Speaker Verification Challenge 2024: Exploring Shared and User-defined PassphrasesabstractIn contrast to text-independent speaker verification, which has received significant attention from researchers and has many competitions dedicated to it, text-dependent speaker verification (TdSV) has been less explored recently. The TdSV Challenge 2024 was organized to analyze and explore novel methods for this type of speaker verification and aims to motivate participants to develop new approaches to TdSV, conduct comprehensive analyses, and investigate advanced techniques such as self-supervised learning. This challenge builds on the achievements of the short-duration speaker verification (SdSV) Challenges held in 2020 and 2021 and focuses specifically on TdSV in two distinct scenarios. The first scenario involves conventional TdSV, while the second focuses on speaker enrollment using user-defined passphrases. This paper provides a detailed description of both tasks, introduces the evaluation rules, and presents a comprehensive analysis of the results obtained from this challenge. Hossein Zeinali, Kong-Aik Lee, Jahangir Alam 0001, Lukás Burget |
ICASSP | 3 |
| 2025 | Text-Independent Speaker Verification Employing A Novel Hybrid Neural Embedding ExtractorabstractReliable and discriminative speaker embedding extraction lies at the heart of modern neural automatic speaker verification (ASV) systems. In this study, we introduce a novel hybrid neural architecture that enhances embedding quality by integrating frequency- and channel-aware Selective Kernel Attention (SKA) into a 2D convolutional neural network (2D-CNN) feature extractor. This design strengthens the joint modeling of frequency and channel characteristics, resulting in more discriminative speaker representations. The feature extractor feeds into a composite frame-level network, structured as a cascade of Time-Delay Neural Network (TDNN)–Long Short-Term Memory (LSTM) hybrid and fully TDNN layers. To summarize speaker traits at the utterance level, we employ Multi-Level Attentive Statistics Pooling (MLASP), which captures diverse statistical cues and exploits the complementary strengths of the hybrid architecture. MLASP further improves robustness by recovering subtle, previously underutilized features. The full system is trained using the additive angular margin softmax (AAMSoftmax) loss, which promotes tighter intra-speaker clustering and broader inter-speaker separation in the embedding space. We also explore the influence of different CNN-driven feature learning modules on ASV performance and resilience. Evaluations on the VoxCeleb and CNCeleb benchmarks confirm that our proposed method consistently surpasses both standard baselines and state-of-the-art ASV models trained under the same conditions. Jahangir Alam 0001, Md Shahidul Alam |
IJCB | 1 |
| 2025 | A Hybrid Neural Approach to Speaker Verification with an Improved Additive Angular Margin LossabstractExtraction of speaker embeddings plays a crucial role in the neural automatic speaker verification system. Here, we propose a novel hybrid neural embedding framework, which employs frequency- and channel-wise Selective Kernel Attention (SKA) into the 2D-CNN - based feature extraction module to aggregate global frequency-channel information to the attention weights to extract speaker discriminant embeddings. The aforementioned feature extraction module is connected with a frame-level network, which is composed of a Time Delay Neural Network (TDNN)-Long Short Term Memory hybrid network and a fully TDNN network in a cascade fashion. Multi-Level Attentive Statistics Pooling, which incorporates local statistics as context, is adopted for aggregating the speaker information within an utterance-level context by capturing the complementarity of different networks. Additionally, the proposed approach utilizes an improved Additive Angular Margin (AAM) Softmax loss function that integrates a dynamic and adaptive label noise cleansing method, termed AdaptiveDrop. This method seamlessly combines label noise filtering and correction in a cascaded manner, leveraging the strengths of both techniques to enhance robustness. Experimental results on the VoxCeleb dataset reveal that the proposed approach outperforms baseline systems in both supervised and self-supervised speaker verification tasks. Jahangir Alam 0001, Abderrahim Fathan, Md Shahidul Alam |
IJCNN | 1 |
| 2025 | Automatic Labeling and Correction of Noisy Labels for Robust Self-Supervised Speaker Verification
Abderrahim Fathan, Jahangir Alam 0001 |
INTERSPEECH | 2 |
| 2025 | An Investigative Study on Recent Sharpness- and Flatness-Based Optimizers for Enhanced Self-Supervised Speaker Verification
Abderrahim Fathan, Jahangir Alam 0001 |
INTERSPEECH | 2 |
| 2024 | Dynamic Cross Attention for Audio-Visual Person VerificationabstractAlthough person or identity verification has been predominantly explored using individual modalities such as face and voice, audio-visual fusion has recently shown immense potential to outperform unimodal approaches. Audio and visual modalities are often expected to pose strong complementary relationships, which plays a crucial role for effective audio-visual fusion. However, they may not always strongly complement each other, they may also exhibit weak complementary relationships, resulting in poor audio-visual feature representations. In this paper, we propose a Dynamic Cross Attention (DCA) model that can dynamically select the cross-attended or unattended features on the fly based on the strong or weak complementary relationships, respectively, across audio and visual modalities. In particular, a conditional gating layer is designed to evaluate the contribution of the cross-attention mechanism and choose cross-attended features only when they exhibit strong complementary relationships, otherwise unattended features. Extensive experiments are conducted on the Voxceleb1 dataset to demonstrate the robustness of the proposed model. Results indicate that the proposed model consistently improves the performance on multiple variants of cross-attention while outperforming the state-of-the-art methods. Code is available at https://github.com/praveena2j/DCAforPersonVerification Gnana Praveen Rajasekhar, Jahangir Alam 0001 |
FG | 2 |
| 2024 | Audio-Visual Person Verification Based on Recursive Fusion of Joint Cross-AttentionabstractPerson or identity verification has been recently gaining a lot of attention using audio-visual fusion as faces and voices share close associations with each other. Conventional approaches based on audio-visual fusion rely on score-level or early feature-level fusion techniques. Though existing approaches showed improvement over unimodal systems, the potential of audio-visual fusion for person verification is not fully exploited. In this paper, we have investigated the prospect of effectively capturing both intra- and inter-modal relationships across audio and visual modalities, which can play a crucial role in significantly improving the fusion performance over unimodal systems. In particular, we introduce a recursive fusion of a joint cross-attentional model, where a joint audio-visual feature representation is employed in the cross-attention framework in a recursive fashion to progressively refine the feature representations that can efficiently capture the intra- and inter-modal relationships. To further enhance the audio-visual feature representations, we have also explored BLSTMs to improve the temporal modeling of audio-visual feature representations. Extensive experiments are conducted on the Voxceleb1 dataset to evaluate the proposed model. Results indicate that the proposed model shows promising improvement in fusion performance by adeptly capturing the intra- and inter-modal relationships across audio and visual modalities. Code is available at https://github.com/praveena2j/RJCAforSpeakerVerification Gnana Praveen Rajasekhar, Jahangir Alam 0001 |
FG | 2 |
| 2024 | Self-Supervised Speaker Verification Employing A Novel Clustering AlgorithmabstractClustering is an unsupervised learning technique, which leverages a large amount of unlabeled data to learn cluster-wise representations from speech. One of the most popular self-supervised techniques to train a speaker verification system is to predict the pseudo-labels using clustering algorithms and then train the speaker embedding net-work using the generated pseudo-labels in a discriminative manner. Therefore, pseudo-labels - driven self-supervised speaker verification systems’ performance relies heavily on the accuracy of the adopted clustering algorithms. In this contribution, we propose a novel clustering technique that not only (i) combines predictions of augmented samples to provide a complementary supervisory signal for clustering and imposes symmetry within the augmentations but also (ii) enforces representation invariance via Self-Augmented Training (SAT) and maximizes the information-theoretic dependency between samples and their predicted pseudo-labels. Experimental results on the Vox-Celeb dataset show that the proposed clustering framework achieves better clustering performance in terms of a variety of clustering metrics. Proposed framework is also able to provide better self-supervised speaker verification performance than the state-of-the-art approaches trained on the same dataset. Abderrahim Fathan, Jahangir Alam 0001 |
ICASSP | 2 |
| 2024 | On the influence of regularization techniques on label noise robustness: Self-supervised speaker verification as a use caseabstractClustering-based Pseudo-Labels (PLs) are widely used to optimize Speaker Embedding networks and train Self-Supervised Speaker Verification (SV) systems. However, this self-supervised training scheme relies on highly accurate PLs. In this paper, we perform a large investigative study of the effect of several regularization techniques (mixup, label smoothing, employing sub-centers) on the label noise robustness of self-supervised speaker verification systems. We study these techniques and apply them to various recent metric learning loss functions for better generalization of self-supervised speaker verification systems. In particular, we investigate the effect of these losses and regularizations on the robustness of the self-supervised SV task against label noise using various clustering models to generate real-world PLs of different noise patterns and levels. We provide a thorough comparative analysis of the generalization performance of these losses and regularization techniques using different numbers of clusters and propose some combination systems that are effective against label noise and lead to considerable improvements in SV performance. Abderrahim Fathan, Jahangir Alam 0001 |
IJCB | 3 |
| 2024 | Cross-Attention is not always needed: Dynamic Cross-Attention for Audio-Visual Dimensional Emotion RecognitionabstractIn video-based emotion recognition, audio and visual modalities are often expected to have a complementary relationship, which is widely explored using cross-attention. However, they may also exhibit weak complementary relationships, resulting in poor representations of audio-visual features, thus degrading the performance of the system. To address this issue, we propose Dynamic Cross-Attention (DCA) that can dynamically select cross-attended or unattended features on the fly based on their strong or weak complementary relationships respectively. Specifically, a simple yet efficient gating layer is designed to evaluate the contribution of the cross-attention mechanism and choose cross-attended features only when they exhibit a strong complementary relationship, otherwise unattended features. We evaluate the performance of the proposed approach on the challenging RECOLA and Aff-Wild2 datasets. We also compare the proposed approach with other variants of cross-attention and show that the proposed model consistently improves the performance on both datasets. Gnana Praveen Rajasekhar, Jahangir Alam 0001 |
ICME | 2 |
| 2024 | On the impact of several regularization techniques on label noise robustness of self-supervised speaker verification systems
Abderrahim Fathan, Jahangir Alam 0001 |
INTERSPEECH | 3 |
| 2024 | An analytic study on clustering driven self-supervised speaker verification
Abderrahim Fathan, Jahangir Alam 0001 |
Pattern Recognit. Lett. | 2 |
| 2023 | CAMSAT: Augmentation Mix and Self-Augmented Training Clustering for Self-Supervised Speaker RecognitionabstractClustering (CL)-based pseudo-labels (PLs) are widely used to optimize speaker embedding (SE) networks and train self-supervised (SS) speaker verification (SV) systems. However, PL-based SS training depends on high-quality PLs. In this paper, we propose a general-purpose CL algorithm called CAMSAT that outperforms all other baselines used to cluster SEs. Moreover, using the generated PLs to train our SE system allows us to further improve SV performance. CAMSAT is based on two principles: (1) mixing predictions of augmented samples to provide a complementary supervisory signal for CL and enforce symmetry within augmentations (2) Self-Augmented Training to enforce representation invariance and maximize the information-theoretic dependency between samples and their predicted PLs. We provide a thorough comparative analysis of the performance of our CL method vs. all baselines using a variety of CL metrics and perform an ablation study to analyze the contribution of each component. Abderrahim Fathan, Jahangir Alam 0001 |
ASRU | 2 |
| 2023 | Hybrid Neural Network with Cross- and Self-Module Attention Pooling for Text-Independent Speaker VerificationabstractExtraction of a speaker embedding vector plays an important role in deep learning-based speaker verification. In this contribution, to extract speaker discriminant utterance level embeddings, we propose a hybrid neural network that employs both cross- and self-module attention pooling mechanisms. More specifically, the proposed system incorporates a 2D-Convolution Neural Network (CNN)-based feature extraction module in cascade with a frame-level network, which is composed of a fully Time Delay Neural Network (TDNN) network and a TDNN-Long Short Term Memory (TDNN-LSTM) hybrid network in a parallel manner. The proposed system also employs a multi-level cross- and self-module attention pooling for aggregating the speaker information within an utterance-level context by capturing the complementarity between two parallelly connected modules. In order to evaluate the proposed system, we conduct a set of experiments on the Voxceleb corpus, and the proposed hybrid network is able to outperform the conventional approaches trained on the same dataset. Jahangir Alam 0001, Woo Hyun Kang, Abderrahim Fathan |
ICASSP | 1 |
| 2023 | On the Use of Cross- and Self-Module Attentive Statistics Pooling Techniques for Text-Independent Speaker VerificationabstractIn neural speaker verification, statistics pooling plays a key role in the learning and extraction of a speaker embedding vector. In this contribution, we perform an investigative study on the use of cross-module and self-module attention statistics pooling mechanisms to extract speaker discriminant utterance level embeddings. More specifically, we propose a novel hybrid neural network that employs a2D-Convolution Neural Network (CNN)-based feature extraction module in cascade with a frame-level network, which is composed of a fully Time Delay Neural Network (TDNN) and a TDNN-Long Short Term Memory (LSTM) hybrid network in a parallel manner. In order to capture the complementarity between two parallelly connected modules, the proposed system also make use of a cross- and self-module attention pooling for aggregating the speaker information within an utterance-level context. We conduct a set of experiments on the Voxceleb and CNCeleb corpora, and the proposed approach is able to provide better performance than the conventional approaches trained on the same dataset. Jahangir Alam 0001 |
IJCB | 1 |
| 2022 | Robust Self-Supervised Speaker Representation Learning Via Instance Mix RegularizationabstractOver the recent years, various self-supervised contrastive embedding learning methods for deep speaker verification were proposed. The performance of the self-supervised contrastive learning framework highly depends on the data augmentation technique, but due to the sensitive nature of speaker information within the speech signal, most speaker embedding training relies on simple augmentations such as additive noise or simulated reverberation. Thus while the conventional self-supervised speaker embedding systems can yield minimum within-utterance variability, the capability to generalize to out-of-set utterance is limited. In order to alleviate this problem, we propose a novel self-supervised learning framework for speaker verification which combines the angular prototypical loss and the instance mix (i-mix) regularization. The proposed method was evaluated on the VoxCeleb1 dataset and showed noticeable improvement over the standard self-supervised embedding method. Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan |
ICASSP | 2 |
| 2022 | Mel-Spectrogram Image-Based End-to-End Audio Deepfake Detection Under Channel-Mismatched ConditionsabstractThis work focuses on the problem of detecting fake audio clips. To improve current audio spoofing detection models, we propose a selection of multiple audio augmentations spe-cially designed to resemble audio spoofing attacks. These augmentations are experimentally found to be very useful and using them achieves a notable performance of 2.8% EER on the ASVspoof 2019 challenge evaluation set. Unlike the widely employed acoustic features, in this paper we explore the use of Mel-spectrogram image features and employ vari-ous audio codecs to achieve robustness to codec and transmission channel variability present in the ASVspoof2021 Evalu-ation set. To better handle spectral information, crucial to de-tect spoofing, we adopt the WaveletCNN and VGG16 archi-tectures which outperform all baselines. Finally, we find that robustness of countermeasure systems degrades dramatically when provided with speech samples degraded through VoIP network transmission or mismatching audio compression. Abderrahim Fathan, Jahangir Alam 0001, Woo Hyun Kang |
ICME | 2 |
| 2022 | MIM-DG: Mutual information minimization-based domain generalization for speaker verification
Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan |
INTERSPEECH | 2 |
| 2022 | Mixup regularization strategies for spoofing countermeasure system
Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan |
INTERSPEECH | 2 |
| 2022 | End-to-end framework for spoof-aware speaker verification
Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan |
INTERSPEECH | 2 |
| 2022 | Deep learning-based end-to-end spoken language identification system for domain-mismatched scenarioabstractDomain mismatch is a critical issue when it comes to spoken language identification. To overcome the domain mismatch problem, we have applied several architectures and deep learning strategies which have shown good results in cross-domain speaker verification tasks to spoken language identification. Our systems were evaluated on the Oriental Language Recognition (OLR) Challenge 2021 Task 1 dataset, which provides a set of cross-domain language identification trials. Among our experimented systems, the best performance was achieved by using the mel frequency cepstral coefficient (MFCC) and pitch features as input and training the ECAPA-TDNN system with a flow-based regularization technique, which resulted in a Cavg of 0.0631 on the OLR 2021 progress set. Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan |
LREC | 2 |
| 2022 | Flow-ER: A Flow-Based Embedding Regularization Strategy for Robust Speech Representation LearningabstractOver the recent years, various deep learning-based embedding methods were proposed. Although the deep learning-based embedding extraction methods have shown good performance in numerous tasks including speaker verification, language identification and anti-spoofing, their performance is limited when it comes to mismatched conditions due to the variability within them unrelated to the main task. In order to alleviate this problem, we propose a novel training strategy that regularizes the embedding network to have minimum information about the nuisance attributes. To achieve this, our proposed method directly incorporates the information bottleneck scheme into the training process, where the mutual information is estimated using an auxiliary normalizing flow network. The performance of the proposed method is evaluated on different speech processing tasks and found to provide improvement over the standard training strategy in all experimentations. Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan |
SLT | 2 |
| 2022 | Multi-level self-attentive TDNN: A general and efficient approach to summarize speech into discriminative utterance-level representations
João Monteiro 0002, Jahangir Alam 0001, Tiago H. Falk |
Speech Commun. | 2 |
| 2022 | A Multimodal Non-Intrusive Stress Monitoring From the Pleasure-Arousal Emotional DimensionsabstractWith the increasing development of advanced unmanned aerial vehicles (UAVs), communication between operators and these intelligent systems is becoming more stressful. For the safety of UAV flights, automatic psychological stress detection is becoming a key research topic for successful missions. Stress can be reliably estimated via some biological markers which are not appropriate in many cases of human-machine-interaction setups. In this article, we propose a non-intrusive deep learning-based stress level estimation approach. The goal is to identify the region where the operator's emotional state projects in the space defined by the latent dimensional emotions of arousal and valence since the stress region is well delimited in this space. The proposed multimodal approach uses sequential temporal CNN and LSTM with an Attention Weighted Average layer in the vision modality. As a second modality, we investigate local and global descriptors such as Mel-frequency cepstral coefficients, i-vector embeddings as well as Fisher-vector encodings. The multimodal-fusion approach uses a strategy referred to as “late-fusion” that involves the combination of unimodal model outputs as inputs of the decision engine. Since we have to deal with more naturalistic behavior in operator-machine interaction contexts, the One minute Gradual Emotion Challenge dataset was used for predictive model validation. Mohamed Dahmane, Jahangir Alam 0001, Pierre-Luc St-Charles, Marc Lalonde, Kevin Heffner, Samuel Foucher |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | Hybrid Network with Multi-Level Global-Local Statistics Pooling for Robust Text-Independent Speaker RecognitionabstractIn this paper, we propose a new hybrid system for extracting a speaker embedding vector. More specifically, the proposed system employs a multi-level global-local statistics pooling method in order to aggregate the speaker information within short time-span and utterance-level context. In order to evaluate the proposed system, a set of experiments on the NIST SRE 2016, Short-duration speaker verification (SdSV) Challenge 2021, and VoxCeleb datasets were conducted, and the proposed hybrid network was able to outperform the conventional approaches trained on the same dataset. Moreover, our experiments showed that the proposed system is able to achieve stable performance even when using a relatively smaller dataset, which highlights the efficiency of the proposed system in extracting the speaker-dependent information. Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan |
ASRU | 2 |
| 2021 | On the use of blind channel response estimation and a residual neural network to detect physical access attacks to speaker verification systems
Anderson R. Avila, Jahangir Alam 0001, Fabiano O. Costa Prado, Douglas D. O'Shaughnessy, Tiago H. Falk |
Comput. Speech Lang. | 2 |
| 2020 | An Ensemble Based Approach for Generalized Detection of Spoofing Attacks to Automatic Speaker RecognizersabstractAs automatic speaker recognizer systems become mainstream, voice spoofing attacks are on the rise. Common attack strategies include replay, the use of text-to-speech synthesis, and voice conversion systems. While previouslyproposed end-to-end detection frameworks have shown to be effective in spotting attacks for one particular spoofing strategy, they have relied on different models, architectures, and speech representations, depending on the spoofing strategy. In practice, however, one does not have a priori information regarding the strategy an attacker might employ to fool a speaker recognizer, thus it is necessary to devise approaches which are able to detect attacks regardless of the strategy employed to generate them. In this work, we introduce an end-to-end ensemble based approach such that two models - previously shown to perform well on each considered attack strategy - are trained jointly, while a third model learns how to mix their outputs yielding a single score. Experimental results with replay and text-to-speech/voice conversion attacks show the proposed ensemble method achieving similar or superior performance when compared to systems specialized on each spoofing strategy separately. João Monteiro 0002, Jahangir Alam 0001, Tiago H. Falk |
ICASSP | 2 |
| 2020 | An end-to-end approach for the verification problem: learning the right distanceabstractIn this contribution, we augment the metric learning setting by introducing a parametric pseudo-distance, trained jointly with the encoder. Several interpretations are thus drawn for the learned distance-like model’s output. We first show it approximates a likelihood ratio which can be used for hypothesis tests, and that it further induces a large divergence across the joint distributions of pairs of examples from the same and from different classes. Evaluation is performed under the verification setting consisting of determining whether sets of examples belong to the same class, even if such classes are novel and were never presented to the model during training. Empirical evaluation shows such method defines an end-to-end approach for the verification problem, able to attain better performance than simple scorers such as those based on cosine similarity and further outperforming widely used downstream classifiers. We further observe training is much simplified under the proposed approach compared to metric learning with actual distances, requiring no complex scheme to harvest pairs of examples. João Monteiro 0002, Isabela Albuquerque, Jahangir Alam 0001, R. Devon Hjelm, Tiago H. Falk |
ICML | 3 |
| 2020 | SdSV Challenge 2020: Large-Scale Evaluation of Short-Duration Speaker Verification
Hossein Zeinali, Kong-Aik Lee, Jahangir Alam 0001, Lukás Burget |
INTERSPEECH | 3 |
| 2020 | On The Performance of Time-Pooling Strategies for End-to-End Spoken Language IdentificationabstractAutomatic speech processing applications often have to deal with the problem of aggregating local descriptors (i.e., representations of input speech data corresponding to specific portions across the time dimension) and turning them into a single fixed-dimension representation, known as global descriptor, on top of which downstream classification tasks can be performed. In this paper, we provide an empirical assessment of different time pooling strategies when used with state-of-the-art representation learning models. In particular, insights are provided as to when it is suitable to use simple statistics of local descriptors or when more sophisticated approaches are needed. Here, language identification is used as a case study and a database containing ten oriental languages under varying test conditions (short-duration test recordings, confusing languages, unseen languages) is used. Experiments are performed with classifiers trained on top of global descriptors to provide insights on open-set evaluation performance and show that appropriate selection of such pooling strategies yield embeddings able to outperform well-known benchmark systems as well as previously results based on attention only. João Monteiro 0002, Jahangir Alam 0001, Tiago H. Falk |
LREC | 2 |
| 2020 | Generalized end-to-end detection of spoofing attacks to automatic speaker recognizers
João Monteiro 0002, Jahangir Alam 0001, Tiago H. Falk |
Comput. Speech Lang. | 2 |
| 2019 | Development of Voice Spoofing Detection Systems for 2019 Edition of Automatic Speaker Verification and Countermeasures ChallengeabstractA robust speaker verification system is expected to provide high recognition accuracy not only in adverse environments but also in the presence of spoofing attacks, which renders voice spoofing detection as crucial to prevent automatic speaker verification systems from a security breach. In this work, we present anti-spoofing systems developed for tackling spoofing attacks introduced for the ASVspoof 2019 challenge. We employ frame-level descriptors such as discrete Fourier transform, as well as constant Q transform-based spectral and cepstral features as countermeasures. These descriptors are both used on their own with a spoofing detection classifier to detect spoofing attacks, or in tandem with deep bottleneck features, i.e. approximate posteriors parametrized by a neural network designed to discriminate between bonafide and spoof signals. Fisher vector encoding and i-vector representations are further learned from the frame-level descriptors of the signals. For modeling, we employ two classification strategies. We finally build an end-to-end anti-spoofing system by making use of modified versions of light convolution neural networks as well as well-known ResNets. Our primary system for the logical access task and a single end-to-end system for the case of physical access we attain significant improvements over two baseline systems. João Monteiro 0002, Jahangir Alam 0001 |
ASRU | 2 |
| 2019 | Adapting End-to-end Neural Speaker Verification to New Languages and Recording Conditions with Adversarial TrainingabstractIn this article we propose a novel approach for adapting speaker embeddings to new domains based on adversarial training of neural networks. We apply our embeddings to the task of text-independent speaker verification, a challenging, real-world problem in biometric security. We further the development of end-to-end speaker embedding models by combing a novel 1-dimensional, self-attentive residual network, an angular margin loss function and adversarial training strategy. Our model is able to learn extremely compact, 64-dimensional speaker embeddings that deliver competitive performance on a number of popular datasets using simple cosine distance scoring. One the NIST-SRE 2016 task we are able to beat a strong i-vector baseline, while on the Speakers in the Wild task our model was able to outperform both i-vector and x-vector baselines, showing an absolute improvement of 2.19% over the latter. Additionally, we show that the integration of adversarial training consistently leads to a significant improvement over an unadapted model. Gautam Bhattacharya, Jahangir Alam 0001, Patrick Kenny |
ICASSP | 2 |
| 2019 | Generative Adversarial Speaker Embedding Networks for Domain Robust End-to-end Speaker VerificationabstractThis article presents a novel approach for learning domain-invariant speaker embeddings using Generative Adversarial Networks. The main idea is to confuse a domain discriminator so that it cannot tell if embeddings are from the source or target domains. We train several GAN variants using our proposed framework and apply them to the speaker verification task. On the challenging NIST-SRE 2016 dataset, we are able to match the performance of a strong baseline x-vector system. In contrast to the the baseline systems which are dependent on dimensionality reduction (LDA) and an external classifier (PLDA), our proposed speaker embeddings can be scored using simple cosine distance. This is achieved by optimizing our models end-to-end, using an angular margin loss function. Furthermore, we are able to significantly boost verification performance by averaging our different GAN models at the score level, achieving a relative improvement of 7.2% over the baseline. Gautam Bhattacharya, João Monteiro 0002, Jahangir Alam 0001, Patrick Kenny |
ICASSP | 3 |
| 2019 | Blind Channel Response Estimation for Replay Attack Detection
Anderson R. Avila, Jahangir Alam 0001, Douglas D. O'Shaughnessy, Tiago H. Falk |
INTERSPEECH | 2 |
| 2019 | Deep Speaker Recognition: Modular or Monolithic?
Gautam Bhattacharya, Jahangir Alam 0001, Patrick Kenny |
INTERSPEECH | 2 |
| 2019 | CRIM's Speech Transcription and Call Sign Detection System for the ATC Airbus Challenge Task
Vishwa Gupta, Lise Rebout, Gilles Boulianne, Pierre André Ménard, Jahangir Alam 0001 |
INTERSPEECH | 5 |
| 2019 | Combining Speaker Recognition and Metric Learning for Speaker-Dependent Representation Learning
João Monteiro 0002, Jahangir Alam 0001, Tiago H. Falk |
INTERSPEECH | 2 |
| 2019 | Intrusive Quality Measurement of Noisy and Enhanced Speech based on i-Vector SimilarityabstractIn this paper, the i-vector framework is investigated as an intrusive quality measure for noisy and enhanced speech. While widely used across numerous speech applications, the potential of using i-vectors to summarize the quality of a speech recording has been overlooked. This paper aims to fill this gap. We show that the i-vector framework is well-suited for assessing speech signal quality, surpassing well-established instrumental measures such as the Perceptual Evaluation of Speech Quality (PESQ) and Perceptual Objective Listening Quality Analysis (POLQA). Three datasets are used in our experiments. First, the TIMIT database is used to train the i-vector extractor on clean speech. To evaluate the proposed method, the noisy speech corpus (NOIZEUS) and the evaluation set of the 2014 IEEE REVERB challenge are used with both containing subjective ratings of perceived quality. Correlations with mean opinion scores (MOS) as high as 0.90 are achieved. Anderson R. Avila, Jahangir Alam 0001, Douglas D. O'Shaughnessy, Tiago H. Falk |
QoMEX | 2 |
| 2019 | Residual convolutional neural network with attentive feature pooling for end-to-end language identification from short-duration speech
João Monteiro 0002, Jahangir Alam 0001, Tiago H. Falk |
Comput. Speech Lang. | 2 |
| 2018 | Investigating Speech Enhancement and Perceptual Quality for Speech Emotion Recognition
Anderson R. Avila, Jahangir Alam 0001, Douglas D. O'Shaughnessy, Tiago H. Falk |
INTERSPEECH | 2 |
| 2018 | Deeply Fused Speaker Embeddings for Text-Independent Speaker Verification
Gautam Bhattacharya, Jahangir Alam 0001, Vishwa Gupta, Patrick Kenny |
INTERSPEECH | 2 |
| 2017 | Speaker Verification Under Adverse Conditions Using i-Vector Adaptation and Neural Networks
Jahangir Alam 0001, Patrick Kenny, Gautam Bhattacharya, Marcel Kockmann |
INTERSPEECH | 1 |
| 2017 | Deep Speaker Embeddings for Short-Duration Speaker Verification
Gautam Bhattacharya, Jahangir Alam 0001, Patrick Kenny |
INTERSPEECH | 2 |
| 2017 | Analysis and Description of ABC Submission to NIST SRE 2016
Oldrich Plchot, Pavel Matejka, Anna Silnova, Ondrej Novotný, Mireia Díez, Johan Rohdin, Ondrej Glembek, Niko Brümmer, Albert Swart, Jesús Jorrín-Prieto, L. Paola García-Perera, Luis Buera, Patrick Kenny, Jahangir Alam 0001, Gautam Bhattacharya |
INTERSPEECH | 14 |
| 2016 | Tandem Features for Text-Dependent Speaker Verification on the RedDots Corpus
Jahangir Alam 0001, Patrick Kenny, Vishwa Gupta |
INTERSPEECH | 1 |
| 2016 | Modelling speaker and channel variability using deep neural networks for robust speaker verificationabstractWe propose to improve the performance of i-vector based speaker verification by processing the i-vectors with a deep neural network before they are fed to a cosine distance or probabilistic linear discriminant analysis (PLDA) classifier. To this end we build on an existing model that we refer to as Non-linear Within Class Normalization (NWCN) and introduce a novel Speaker Classifier Network (SCN). Both models deliver impressive speaker verification performance, showing a 56% and 68% relative improvement over standard i-vectors when combined with a cosine distance backend. The NWCN model also reduces the equal error rate for PLDA from 1.78% to 1.63%. We also test these models under the constraints of domain mismatch, i.e. when no in-domain training data is available. Under these conditions, SCN features in combination with cosine distance performs better than the PLDA baseline, achieving an equal error rate of 2.92% as compared to 3.37%. Gautam Bhattacharya, Jahangir Alam 0001, Patrick Kenny, Vishwa Gupta |
SLT | 2 |
| 2016 | Text-Dependent Speaker Recognition With Random Digit StringsabstractIn this paper, we explore joint factor analysis (JFA) for text-dependent speaker recognition with random digit strings. The core of the proposed method is a JFA model by which we extract features. These features can either represent overall utterances or individual digits, and are fed into a trainable backend to estimate likelihood ratios. Within this framework, several extensions are proposed. First is a logistic regression method for combining log-likelihood ratios that correspond to individual mixture components. Second is the extraction of phonetically aware Baum-Welch statistics, by using forced alignment instead of the typical posterior probabilities that are derived by the universal background model. We also explore a digit-string-dependent way to apply score normalization that exhibits a notable improvement compared to the standard one. By fusing six JFA features, we attained 2.01% and 3.19% equal error rates on male and female, respectively, on the challenging RSR2015 (part III) dataset. Themos Stafylakis, Jahangir Alam 0001, Patrick Kenny |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Speaker and Channel Factors in Text-Dependent Speaker RecognitionabstractWe reformulate joint factor analysis so that it can serve as a feature extractor for text-dependent speaker recognition. The new formulation is based on left-to-right modeling with tied mixture HMMs and it is designed to deal with problems such as the inadequacy of subspace methods in modeling speaker-phrase variability, UBM mismatches that arise as a result of variable phonetic content, and the need to exploit text-independent resources in text-dependent speaker recognition. We pass the features extracted by factor analysis to a trainable backend which plays a role analogous to that of PLDA in the i-vector/PLDA cascade in text-independent speaker recognition. We evaluate these methods on a proprietary dataset consisting of English and Urdu passphrases collected in Pakistan. By using both text-independent data and text-dependent data for training purposes and by fusing results obtained with multiple front ends at the score level, we achieved equal error rates of around 1.3% and 2% on the English and Urdu portions of this task. Themos Stafylakis, Patrick Kenny, Jahangir Alam 0001, Marcel Kockmann |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | JFA modeling with left-to-right structure and a new backend for text-dependent speaker recognitionabstractThis paper introduces a new formulation of Joint Factor Analysis (JFA) for text-dependent speaker recognition based on left-to-right modeling with tied mixture HMMs. It accommodates many different ways of extracting multiple features to characterize speakers (features may or may not be HMM state-dependent, they may be modeled with subspace or factorial priors and these priors maybe imputed from text-dependent or text-independent background data). We feed these features to a new, trainable classifier for text-dependent speaker recognition in a manner which is broadly analogous to the i-vector/PLDA cascade in text-independent speaker recognition. We have evaluated this approach on a challenging proprietary dataset consisting of telephone recordings of short English and Urdu pass-phrases collected in Pakistan. By fusing results obtained with multiple front ends, equal error rate of around 2% are achievable. Patrick Kenny, Themos Stafylakis, Jahangir Alam 0001, Marcel Kockmann |
ICASSP | 3 |
| 2015 | Development of CRIM system for the automatic speaker verification spoofing and countermeasures challenge 2015abstractThe automatic speaker verification spoofing and countermeasures challenge 2015 provides a common framework for the evaluation of spoofing countermeasures or anti-spoofing techniques in the presence of various seen and unseen spoofing attacks. This contribution proposes a system consisting of amplitude, phase, linear prediction residual, and combined amplitude - phase-based countermeasures for the detection of spoofing attacks. In this task we use following features: Mel-frequency cepstral coefficients (MFCC), product spectrum-based cepstral coefficients, modified group delay cepstral coefficients, weighted linear prediction group delay cepstral coefficients, linear prediction residual cepstral coefficients, cosine normalized phase-based cepstral features (CNPCC), and a combination of MFCC-CNPCC. The product spectrum-based features are influenced by both the amplitude and phase spectra. The Gaussian Mixture Model (GMM) classifier is used for the discrimination of the human and spoofed speech signals. Our primary submitted system is a linear fusion of the sub-systems based on the features mentioned above with fusion weights trained on the development dataset. Experimental results on the challenge evaluation data provided an average EER (equal error rate) of 0.041%, 5.347%, and 2.69% on the known, unknown and all (known + unknown) spoofing attacks, respectively. Among all the systems product spectrum-based cepstral coefficients- and conventional MFCC (without any feature normalization)based systems performed the best in terms of EER measure. On the known, unknown and all conditions the EER obtained by the MFCC and product spectrum-based features are 0.78% & 0.65%, 5.39% & 5.37% and 3.09% & 3.01%, respectively. Jahangir Alam 0001, Patrick Kenny, Gautam Bhattacharya, Themos Stafylakis |
INTERSPEECH | 1 |
| 2015 | Combining amplitude and phase-based features for speaker verification with short duration utterancesabstractDue to the increasing use of fusion in speaker recognition systems, one trend of current research activity focuses on new features that capture complementary information to the MFCC (Mel-frequency cepstral coefficients) for improving speaker recognition performance. The goal of this work is to combine (or fuse) amplitude and phase-based features to improve speaker verification performance. Based on the amplitude and phase spectra we investigate some possible variations to the extraction of cepstral coefficients that produce diversity with respect to fused subsystems. Among the amplitude-based features we consider widely used MFCC, Linear frequency cepstral coefficients, and multitaper spectrum estimationbased MFCC (denoted here as MMFCC). To compute phasebased features we choose modified group delayand all-pole group delay-, linear prediction residual phase-based features. We also consider product spectrum-based cepstral coefficients features that are influenced by both the amplitude and phase spectra. For performance evaluation, text-dependent speaker verification experiments are conducted on the a proprietary dataset known as Voice Trust-Pakistan (VT-Pakistan) corpus. Experimental results show that the fused system provide reduced error rate compared to both the amplitude and phasebased features. On the average fused system provided a relative improvement of 37% over the baseline MFCC systems in terms of EER, DCF (detection cost function) of SRE 2008 and DCF of SRE 2010. Jahangir Alam 0001, Patrick Kenny, Themos Stafylakis |
INTERSPEECH | 1 |
| 2015 | An i-vector backend for speaker verificationabstractWe propose a new approach to the problem of uncertainty modeling in text-dependent speaker verification where speaker factors are used as the feature representation. The state-of-the-art backend in this situation consists in using point estimates of speaker factors to model the joint distribution of pairs of enrollment and test feature vectors under the same-speaker hypothesis. We develop a version of this backend that works with Baum-Welch statistics instead of point estimates. The likelihood ratio calculations for speaker verification turn out to be formally equivalent to evidence calculations with i-vector extractors having non-standard normal priors. Experiments show that this i-vector backend performs well on Part III of the RSR2015 dataset. Patrick Kenny, Themos Stafylakis, Jahangir Alam 0001, Marcel Kockmann |
INTERSPEECH | 3 |
| 2015 | The reddots data collection for speaker recognitionabstractde niveau recherche, publiés ou non, émanant des établissements d'enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Kong-Aik Lee, Anthony Larcher, Guangsen Wang, Patrick Kenny, Niko Brümmer, David A. van Leeuwen, Hagai Aronowitz, Marcel Kockmann, Carlos Vaquero, Bin Ma 0001, Haizhou Li 0001, Themos Stafylakis, Jahangir Alam 0001, Albert Swart, Javier Perez |
INTERSPEECH | 13 |
| 2015 | JFA for speaker recognition with random digit stringsabstractIn this paper, we examine the use of Joint Factor Analysis methods on RSR2015 part III (digits), [1]. A tied-mixture HMM is used for segmentation of the utterances into digits, while Joint Factor Analysis and a trainable backend are deployed for feature extraction and LLR calculation, respectively. A novel approach for digit-dependent fusion of UBMcomponent log-likelihood ratios is introduced, yielding the best results so far. The fusion of 5 different JFA features gives an equal-error rate of 3.6%, compared to 6.3% attained by the a baseline GMM-UBM model with score normalization. JFA for feature extraction JFA vs. i-vectors • The text-independent paradigm of i-vector/PLDA has not been successful in text-dependent speakerrecognition. The speaker-phrase variability is hard to be confined into a low-dimensional subspace. • JFA offers the flexibility of confining the channel effects in a subspace while allowing the speaker-phrace factors to lie on the supervector space, [2]. Main JFA equation S = m + Ux + V y + Dz (1) • The hidden variable x varies from one recording to another and is intended to model channel effects. • In text-independent speaker recognition, the term Dz is usually dropped and speakers are characterized by the low-dimensional vector y. Here, we extract either z or y features, [3]. JFA on utterances segmented into digits • JFA can be extended to utterances that are segmented into HMM states (digits). • Features can be global (digit-independent) or local (digit-dependent), supervectors-sized (z-vectors) or subspace (y-vectors). Segmentation and Baum-Welch stats Tied-Mixture HMM • Train a UBMand use its means and covariance matrices as codebook for a Tied-Mixture HMM (TMM) • The TMM has a single Gaussian codebook and digitdependent weights. • Very efficient for training and evaluating (Viterbi algorithm). • We use it also for extracting Baum-Welch stats for local features instead of the UBM. Training and evaluating the system Training the JFA and backend • Train a JFA model using both local and global features, z or y-vectors. (Several combinations are possible.) • Extract z or y-vectors, project them onto the unitsphere). • Train a Joint-Density Backend per feature. Evaluating the model • Apply Viterbi segmentation, extract z or y-vectors and use the JDB to calculate LLRs for each trial. • Apply score normalization and fuse score-normalized LLRs coming from multiple features. Joint-Density Backend An Alternative to PLDA • We model the joint-distribution of pairs of enrollment and test vectors under the same speaker hypothesis, [4]. • We use ”target” trials from the training set t = [ye , y T t ] T . • We estimate mean and covariance matrix (C). Assuming zero mean, C is as follows: Themos Stafylakis, Patrick Kenny, Jahangir Alam 0001, Marcel Kockmann |
INTERSPEECH | 3 |
| 2015 | Regularized minimum variance distortionless response-based cepstral features for robust continuous speech recognition
Jahangir Alam 0001, Patrick Kenny, Douglas D. O'Shaughnessy |
Speech Commun. | 1 |
| 2014 | JFA-based front ends for speaker recognitionabstractWe discuss the limitations of the i-vector representation of speech segments in speaker recognition and explain how Joint Factor Analysis (JFA) can serve as an alternative feature extractor in a variety of ways. Building on the work of Zhao and Dong, we implemented a variational Bayes treatment of JFA which accommodates adaptation of universal background models (UBMs) in a natural way. This allows us to experiment with several types of features for speaker recognition: speaker factors and diagonal factors in addition to i-vectors, extracted with and without UBM adaptation in each case. We found that, in text-independent speaker verification experiments on NIST data, extracting i-vectors with UBM adaptation led to a 10% reduction in equal error rates although performance did not improve consistently over the whole DET curve. We achieved a further 10% reduction (with a similar inconsistency) by using speaker factors extracted with UBM adaptation as features. In text-dependent speaker recognition experiments on RSR2015 data, we were able to achieve very good performance using a JFA model with diagonal factors but no speaker factors as a feature extractor. Contrary to standard practice, this JFA model was configured so as to model speakerphrase combinations (rather than speakers) and it was trained on utterances of very short duration (rather than whole recording sessions). We also present a variant of the length normalization trick inspired by uncertainty propagation which leads to substantial gains in performance over the whole DET curve. Patrick Kenny, Themos Stafylakis, Pierre Ouellet, Jahangir Alam 0001 |
ICASSP | 4 |
| 2014 | Noise spectrum estimation using Gaussian mixture model-based speech presence probability for robust speech recognitionabstractThis work presents a noise spectrum estimator based on the Gaussian mixture model (GMM)-based speech presence probability (SPP) for robust speech recognition. Estimated noise spectrum is then used to compute a subband a posteriori signal-to-noise ratio (SNR). A sigmoid shape weighting rule is formed based on this subband a posteriori SNR to enhance the speech spectrum in the auditory domain, which is used in the Mel-frequency cepstral coefficient (MFCC) framework for robust feature, denoted here as Robust MFCC (RMFCC) extraction. The performance of the GMM-SPP noise spectrum estimator-based RMFCC feature extractor is evaluated in the context of speech recognition on the AURORA-4 continuous speech recognition task. For comparison we incorporate six existing noise estimation methods into this auditory domain spectrum enhancement framework. The ETSI advanced frontend (ETSI-AFE), power normalized cepstral coefficients (PNCC), and robust compressive gammachirp cepstral coefficients (RCGCC) are also considered for comparison purposes. Experimental speech recognition results show that, in terms of word accuracy, RMFCC provides an average relative improvements of 8.1%, 6.9% and 6.6% over RCGCC, ETSI-AFE, and PNCC, respectively. With GMM-SPP -based noise estimation method an average relative improvement of 3.6% is obtained over other six noise estimation methods in terms of word recognition accuracy. Jahangir Alam 0001, Patrick Kenny, Pierre Dumouchel, Douglas D. O'Shaughnessy |
INTERSPEECH | 1 |
| 2014 | In-domain versus out-of-domain training for text-dependent JFAabstractWe propose a simple and effective strategy to cope with dataset shifts in text-dependent speaker recognition based on Joint Factor Analysis (JFA). We have previously shown how to compensate for lexical variation in text-dependent JFA by adapting the Universal Background Model (UBM) to individual passphrases. A similar type of adaptation can be used to port a JFA model trained on out-of-domain data to a given text-dependent task domain. On the RSR2015 test set we found that this type of adaptation gave essentially the same results as in-domain JFA training. To explore this idea more fully, we experimented with several types of JFA model on the CSLU speaker recognition dataset. Taking a suitably configured JFA model trained on NIST data and adapting it in the proposed way results in a 22% reduction in error rates compared with the GMM/UBM benchmark. Error rates are still much higher than those that can be achieved on the RSR2015 test set with the same strategy but cheating experiments suggest that if large amounts of in-domain training data are available, then JFA modelling is capable in principle of achieving very low error rates even on hard tasks such as CSLU. Patrick Kenny, Themos Stafylakis, Jahangir Alam 0001, Pierre Ouellet, Marcel Kockmann |
INTERSPEECH | 3 |
| 2013 | Speech recognition using regularized minimum variance distortionless response spectrum estimation-based cepstral featuresabstractThis paper presents regularized minimum variance distortion-less response (MVDR)-based cepstral features for robust continuous speech recognition. The mel-frequency cepstral coefficient (MFCC) features, widely used in speech recognition tasks, are usually computed from a direct spectrum estimate, that is, the squared magnitude of the discrete Fourier transform (DFT) of speech frames. Direct spectrum estimation methods (also known as nonparametric estimators) perform poorly under noisy and adverse conditions. To reduce this performance drop we propose to increase robustness of the speech recognition system by extracting more robust features based on the regularized MVDR technique. The proposed method, when evaluated on the AURORA-4 speech recognition task, provides an average relative improvement in word accuracy of 11.3%, 6.1%, and 5.2% over the conventional MFCC, PLP, MVDR and PMVDR-based MFCC features, respectively. Jahangir Alam 0001, Patrick Kenny, Douglas D. O'Shaughnessy |
ICASSP | 1 |
| 2013 | Multiple windowed spectral features for emotion recognitionabstractMFCC (Mel Frequency Cepstral Coefficients) and PLP (Perceptual linear prediction coefficients) or RASTA-PLP have demonstrated good results whether when they are used in combination with prosodic features as suprasegmental (long-term) information or when used stand-alone as segmental (short-time) information. MFCC and PLP feature parameterization aims to represent the speech parameters in a way similar to how sound is perceived by humans. However, MFCC and PLP are usually computed from a Hamming-windowed periodogram spectrum estimate that is characterized by large variance. In this paper we study the effect of averaging spectral estimates obtained using a set of orthogonal tapers (windows) on emotion recognition performance. The multitaper MFCC and PLP are examined separately as short-time information vectors modeled using Gaussian mixture models (GMMs). When tested on the FAU AIBO spontaneous emotion corpus, a relative improvement ranging from 2.2% to 3.9% for both MFCC and PLP systems is achieved by multiple windowed spectral features compared to single windowed ones. Yazid Attabi, Jahangir Alam 0001, Pierre Dumouchel, Patrick Kenny, Douglas D. O'Shaughnessy |
ICASSP | 2 |
| 2013 | PLDA for speaker verification with utterances of arbitrary durationabstractThe duration of speech segments has traditionally been controlled in the NIST speaker recognition evaluations so that researchers working in this framework have been relieved of the responsibility of dealing with the duration variability that arises in practical applications. The fixed dimensional i-vector representation of speech utterances is ideal for working under such controlled conditions and ignoring the fact that i-vectors extracted from short utterances are less reliable than those extracted from long utterances leads to a very simple formulation of the speaker recognition problem. However a more realistic approach seems to be needed to handle duration variability properly. In this paper, we show how to quantify the uncertainty associated with the i-vector extraction process and propagate it into a PLDA classifier. We evaluated this approach using test sets derived from the NIST 2010 core and extended core conditions by randomly truncating the utterances in the female, telephone speech trials so that the durations of all enrollment and test utterances lay in the range 3-60 seconds and we found that it led to substantial improvements in accuracy. Although the likelihood ratio computation for speaker verification is more computationally expensive than in the standard i-vector/PLDA classifier, it is still quite modest as it reduces to computing the probability density functions of two full covariance Gaussians (irrespective of the number of the number of utterances used to enroll a speaker). Patrick Kenny, Themos Stafylakis, Pierre Ouellet, Jahangir Alam 0001, Pierre Dumouchel |
ICASSP | 4 |
| 2013 | Amplitude modulation features for emotion recognition from speechabstractThe goal of speech emotion recognition (SER) is to identify the emotional or physical state of a human being from his or her voice. One of the most important things in a SER task is to extract and select relevant speech features with which most emotions could be recognized. In this paper, we present a smoothed nonlinear energy operator (SNEO)-based amplitude modulation cepstral coefficients (AMCC) feature for recognizing emotions from speech signals. SNEO estimates the energy required to produce the AM-FM signal, and then the estimated energy is separated into its amplitude and frequency components using an energy separation algorithm (ESA). AMCC features are obtained by first decomposing a speech signal using a C-channel gammatone filterbank, computing the AM power spectrum, and taking a discrete cosine transform (DCT) of the root compressed AM power spectrum. Conventional MFCC (Mel-frequency cepstral coefficients) and Mel-warped DFT (discrete Fourier transform) spectrum based cepstral coefficients (MWDCC) features are used for comparing the recognition performances of the proposed features. Emotion recognition experiments are conducted on the FAU AIBO spontaneous emotion corpus. It is observed from the experimental results that the AMCC features provide a relative improvement of approximately 3.5% over the baseline MFCC. Jahangir Alam 0001, Yazid Attabi, Pierre Dumouchel, Patrick Kenny, Douglas D. O'Shaughnessy |
INTERSPEECH | 1 |
| 2013 | Regularized MVDR spectrum estimation-based robust feature extractors for speech recognitionabstractIn this paper, we present two robust feature extractors that use a regularized minimum variance distortionless response (RMVDR) spectrum estimator instead of the discrete Fourier transform-based direct spectrum estimator, used in many front-ends including the conventional MFCC, for estimating the speech power spectrum. Direct spectrum estimators, e.g., single tapered periodogram, have high variance and they perform poorly under noisy and adverse conditions. RMVDR spectrum estimator has low spectral variance and are robust to mismatch conditions. Based on RMVDR spectrum estimator two robust feature extractors, robust RMVDR cepstral coefficients (RRMCC) and normalized RMVDR cepstral coefficients (NRMCC), are proposed that incorporate an auditory domain spectrum enhancement (ASE) method and a medium duration power bias subtraction (MDPBS) technique, respectively, for enhancement of the speech spectrum. Experimental speech recognition results are conducted on the AURORA-4 corpus and performances are compared with the MFCC, PLP, MVDR-MFCC, RMVDR-MFCC, PMVDR, ETSI advancement front-end (ETSI-AFE), PNCC, CFCC, and the robust feature extractor (RFE) of [6]. Experimental results demonstrate that the proposed robust feature extractors outperformed the other robust front-ends in terms of percentage word accuracy on the AURORA-4 large vocabulary continuous speech recognition (LVCSR) task under different mismatch conditions. Jahangir Alam 0001, Patrick Kenny, Douglas D. O'Shaughnessy |
INTERSPEECH | 1 |
| 2013 | Frequency warping and robust speaker verification: a comparison of alternative mel-scale representationsabstractAccuracy of speaker verification is high under controlled condi-tions but falls off rapidly in the presence of interfering sounds. This is because spectral features, such as Mel-frequency cep-stral coefficients (MFCCs), are sensitive to additive noise. MFCCs are a particular realization of warped-frequency rep-resentation with low-frequency focus. But there are several alternative, potentially more robust, warped-frequency repre-sentations. We provide an experimental comparison of five warped-frequency features. They use exactly the same fre-quency warping function, the same number of coefficients and postprocessing, but differ in their internal computations. The compared variants are (1) conventional MFCCs from discrete Fourier transform (DFT), followed by Mel-scaled filterbank, (2) MFCCs via direct warping of DFT, followed by linear-scale fil-terbank, (3) warped linear prediction features, (4) perceptual minimum variance distortionless features and (5) recently pro-posed sparse Mel-scale histogram features. Experiments car-ried out on a subset of the SRE 10 corpus using a scaled-down i-vector system indicate that direct DFT warping outperforms conventional MFCCs in most of the cases. Index Terms: speaker recognition, noise, frequency warping 1. Tomi Kinnunen, Jahangir Alam 0001, Pavel Matejka, Patrick Kenny, Jan Cernocký, Douglas D. O'Shaughnessy |
INTERSPEECH | 2 |
| 2013 | Multitaper MFCC and PLP features for speaker verification using i-vectors
Jahangir Alam 0001, Tomi Kinnunen, Patrick Kenny, Pierre Ouellet, Douglas D. O'Shaughnessy |
Speech Commun. | 1 |
| 2012 | Robust Feature Extraction for Speech Recognition by Enhancing Auditory Spectrum
Jahangir Alam 0001, Patrick Kenny, Douglas D. O'Shaughnessy |
INTERSPEECH | 1 |
| 2011 | Multi-taper MFCC features for speaker verification using I-vectorsabstractThis paper studies the low-variance multi-taper mel-frequency cepstral coefficient (MFCC) features in the state-of-the-art speaker verification. The MFCC features are usually computed using a Hamming-windowed DFT spectrum. Windowing reduces the bias of the spectrum but variance remains high. Recently, low-variance multi-taper MFCC features were studied in speaker verification with promising preliminary results on the NIST 2002 SRE data using a simple GMM-UBM recognizer. In this study our goal is to validate those findings using a up-to-date i-vector classifier on the latest NIST 2010 SRE data. Our experiment on the telephone (det5) and microphone speech (det1, det2, det3 and det4) indicate that the multi-taper approaches perform better than the conventional Hamming window technique. Jahangir Alam 0001, Tomi Kinnunen, Patrick Kenny, Pierre Ouellet, Douglas D. O'Shaughnessy |
ASRU | 1 |
| 2011 | Full-covariance UBM and heavy-tailed PLDA in i-vector speaker verificationabstractIn this paper, we describe recent progress in i-vector based speaker verification. The use of universal background models (UBM) with full-covariance matrices is suggested and thoroughly experimentally tested. The i-vectors are scored using a simple cosine distance and advanced techniques such as Probabilistic Linear Discriminant Analysis (PLDA) and heavy-tailed variant of PLDA (PLDA-HT). Finally, we investigate into dimensionality reduction of i-vectors before entering the PLDA-HT modeling. The results are very competitive: on NIST 2010 SRE task, the results of a single full-covariance LDA-PLDA-HT system approach those of complex fused system. Pavel Matejka, Ondrej Glembek, Fabio Castaldo, Jahangir Alam 0001, Oldrich Plchot, Patrick Kenny, Lukás Burget, Jan Cernocký |
ICASSP | 4 |
| 2009 | An improved perceptual speech enhancement technique employing a psychoacoustically motivated weighting factorabstractSuppression of speech components after perceptual speech enhancement (SE) lowers the noise masking threshold (NMT) level of the enhanced signal. This may re-introduce noise components that are initially masked but not processed by the denoising filter, thereby, favoring the emergence of musical noise. This paper presents a modified perceptual speech enhancement algorithm based on a perceptually motivated weighting factor to effectively suppress the background noise without introducing much distortion in the enhanced signal using the perceptual speech enhancement methods. The performance of the proposed enhancement algorithm is evaluated by the Segmental SNR and Perceptual Evaluation of Speech Quality (PESQ) measures under various noisy environments and yields better results compared to the perceptual speech enhancement methods. Jahangir Alam 0001, Sid-Ahmed Selouani, Douglas D. O'Shaughnessy |
ASRU | 1 |
| 2008 | Speech enhancement based on novel two-step a priori SNR estimators
Jahangir Alam 0001, Douglas D. O'Shaughnessy, Sid-Ahmed Selouani |
INTERSPEECH | 1 |
| 2008 | Speech enhancement using a wiener denoising technique and musical noise reduction
Jahangir Alam 0001, Sid-Ahmed Selouani, Douglas D. O'Shaughnessy, Sofia Ben Jebara |
INTERSPEECH | 1 |