VLDB 2026 Research / reviewers in the wild / expert
Abderrahim Fathan
dblp:293/6969
· DBLP profile ↗
18ranked-venue papers
9as first author
18since 2021 · last 2025
0000-0002-5749-1643ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 12 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AdaptiveDrop: A Simple Adaptive Label Noise Filtering Scheme for Enhanced Self-supervised Speaker VerificationabstractUsing clustering-driven annotations to train a neural network can be a tricky task because of label noise. In this paper, we propose a dynamic and adaptive label noise cleansing method, called AdaptiveDrop which combines both label noise filtering and correction simultaneously in cascade to combine their advantages. Contrary to other label noise filtering approaches, our method filters noisy samples on the fly from an early stage of training. We also provide a variant that incorporates sub-centers per each class for enhanced robustness to label noise by continuously tracking the dominant sub-centers via a dictionary table. AdaptiveDrop is a simple general-purpose method, performed end-to-end in only one stage of training, can be integrated with any loss function, and does not require training from scratch on the cleansed dataset. We show through extensive ablation studies for the self-supervised speaker verification task that our method is effective, benefits from long epochs of iterative filtering and provides consistent performance gains across various loss functions and real-world pseudo-labels. Abderrahim Fathan, Jahangir Alam 0001 |
ICASSP | 1 |
| 2025 | A Hybrid Neural Approach to Speaker Verification with an Improved Additive Angular Margin LossabstractExtraction of speaker embeddings plays a crucial role in the neural automatic speaker verification system. Here, we propose a novel hybrid neural embedding framework, which employs frequency- and channel-wise Selective Kernel Attention (SKA) into the 2D-CNN - based feature extraction module to aggregate global frequency-channel information to the attention weights to extract speaker discriminant embeddings. The aforementioned feature extraction module is connected with a frame-level network, which is composed of a Time Delay Neural Network (TDNN)-Long Short Term Memory hybrid network and a fully TDNN network in a cascade fashion. Multi-Level Attentive Statistics Pooling, which incorporates local statistics as context, is adopted for aggregating the speaker information within an utterance-level context by capturing the complementarity of different networks. Additionally, the proposed approach utilizes an improved Additive Angular Margin (AAM) Softmax loss function that integrates a dynamic and adaptive label noise cleansing method, termed AdaptiveDrop. This method seamlessly combines label noise filtering and correction in a cascaded manner, leveraging the strengths of both techniques to enhance robustness. Experimental results on the VoxCeleb dataset reveal that the proposed approach outperforms baseline systems in both supervised and self-supervised speaker verification tasks. Jahangir Alam 0001, Abderrahim Fathan, Md Shahidul Alam |
IJCNN | 2 |
| 2025 | Automatic Labeling and Correction of Noisy Labels for Robust Self-Supervised Speaker Verification
Abderrahim Fathan, Jahangir Alam 0001 |
INTERSPEECH | 1 |
| 2025 | An Investigative Study on Recent Sharpness- and Flatness-Based Optimizers for Enhanced Self-Supervised Speaker Verification
Abderrahim Fathan, Jahangir Alam 0001 |
INTERSPEECH | 1 |
| 2024 | Self-Supervised Speaker Verification Employing A Novel Clustering AlgorithmabstractClustering is an unsupervised learning technique, which leverages a large amount of unlabeled data to learn cluster-wise representations from speech. One of the most popular self-supervised techniques to train a speaker verification system is to predict the pseudo-labels using clustering algorithms and then train the speaker embedding net-work using the generated pseudo-labels in a discriminative manner. Therefore, pseudo-labels - driven self-supervised speaker verification systems’ performance relies heavily on the accuracy of the adopted clustering algorithms. In this contribution, we propose a novel clustering technique that not only (i) combines predictions of augmented samples to provide a complementary supervisory signal for clustering and imposes symmetry within the augmentations but also (ii) enforces representation invariance via Self-Augmented Training (SAT) and maximizes the information-theoretic dependency between samples and their predicted pseudo-labels. Experimental results on the Vox-Celeb dataset show that the proposed clustering framework achieves better clustering performance in terms of a variety of clustering metrics. Proposed framework is also able to provide better self-supervised speaker verification performance than the state-of-the-art approaches trained on the same dataset. Abderrahim Fathan, Jahangir Alam 0001 |
ICASSP | 1 |
| 2024 | On the influence of regularization techniques on label noise robustness: Self-supervised speaker verification as a use caseabstractClustering-based Pseudo-Labels (PLs) are widely used to optimize Speaker Embedding networks and train Self-Supervised Speaker Verification (SV) systems. However, this self-supervised training scheme relies on highly accurate PLs. In this paper, we perform a large investigative study of the effect of several regularization techniques (mixup, label smoothing, employing sub-centers) on the label noise robustness of self-supervised speaker verification systems. We study these techniques and apply them to various recent metric learning loss functions for better generalization of self-supervised speaker verification systems. In particular, we investigate the effect of these losses and regularizations on the robustness of the self-supervised SV task against label noise using various clustering models to generate real-world PLs of different noise patterns and levels. We provide a thorough comparative analysis of the generalization performance of these losses and regularization techniques using different numbers of clusters and propose some combination systems that are effective against label noise and lead to considerable improvements in SV performance. Abderrahim Fathan, Jahangir Alam 0001 |
IJCB | 1 |
| 2024 | On the impact of several regularization techniques on label noise robustness of self-supervised speaker verification systems
Abderrahim Fathan, Jahangir Alam 0001 |
INTERSPEECH | 1 |
| 2024 | An analytic study on clustering driven self-supervised speaker verification
Abderrahim Fathan, Jahangir Alam 0001 |
Pattern Recognit. Lett. | 1 |
| 2023 | CAMSAT: Augmentation Mix and Self-Augmented Training Clustering for Self-Supervised Speaker RecognitionabstractClustering (CL)-based pseudo-labels (PLs) are widely used to optimize speaker embedding (SE) networks and train self-supervised (SS) speaker verification (SV) systems. However, PL-based SS training depends on high-quality PLs. In this paper, we propose a general-purpose CL algorithm called CAMSAT that outperforms all other baselines used to cluster SEs. Moreover, using the generated PLs to train our SE system allows us to further improve SV performance. CAMSAT is based on two principles: (1) mixing predictions of augmented samples to provide a complementary supervisory signal for CL and enforce symmetry within augmentations (2) Self-Augmented Training to enforce representation invariance and maximize the information-theoretic dependency between samples and their predicted PLs. We provide a thorough comparative analysis of the performance of our CL method vs. all baselines using a variety of CL metrics and perform an ablation study to analyze the contribution of each component. Abderrahim Fathan, Jahangir Alam 0001 |
ASRU | 1 |
| 2023 | Hybrid Neural Network with Cross- and Self-Module Attention Pooling for Text-Independent Speaker VerificationabstractExtraction of a speaker embedding vector plays an important role in deep learning-based speaker verification. In this contribution, to extract speaker discriminant utterance level embeddings, we propose a hybrid neural network that employs both cross- and self-module attention pooling mechanisms. More specifically, the proposed system incorporates a 2D-Convolution Neural Network (CNN)-based feature extraction module in cascade with a frame-level network, which is composed of a fully Time Delay Neural Network (TDNN) network and a TDNN-Long Short Term Memory (TDNN-LSTM) hybrid network in a parallel manner. The proposed system also employs a multi-level cross- and self-module attention pooling for aggregating the speaker information within an utterance-level context by capturing the complementarity between two parallelly connected modules. In order to evaluate the proposed system, we conduct a set of experiments on the Voxceleb corpus, and the proposed hybrid network is able to outperform the conventional approaches trained on the same dataset. Jahangir Alam 0001, Woo Hyun Kang, Abderrahim Fathan |
ICASSP | 3 |
| 2022 | Robust Self-Supervised Speaker Representation Learning Via Instance Mix RegularizationabstractOver the recent years, various self-supervised contrastive embedding learning methods for deep speaker verification were proposed. The performance of the self-supervised contrastive learning framework highly depends on the data augmentation technique, but due to the sensitive nature of speaker information within the speech signal, most speaker embedding training relies on simple augmentations such as additive noise or simulated reverberation. Thus while the conventional self-supervised speaker embedding systems can yield minimum within-utterance variability, the capability to generalize to out-of-set utterance is limited. In order to alleviate this problem, we propose a novel self-supervised learning framework for speaker verification which combines the angular prototypical loss and the instance mix (i-mix) regularization. The proposed method was evaluated on the VoxCeleb1 dataset and showed noticeable improvement over the standard self-supervised embedding method. Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan |
ICASSP | 3 |
| 2022 | Mel-Spectrogram Image-Based End-to-End Audio Deepfake Detection Under Channel-Mismatched ConditionsabstractThis work focuses on the problem of detecting fake audio clips. To improve current audio spoofing detection models, we propose a selection of multiple audio augmentations spe-cially designed to resemble audio spoofing attacks. These augmentations are experimentally found to be very useful and using them achieves a notable performance of 2.8% EER on the ASVspoof 2019 challenge evaluation set. Unlike the widely employed acoustic features, in this paper we explore the use of Mel-spectrogram image features and employ vari-ous audio codecs to achieve robustness to codec and transmission channel variability present in the ASVspoof2021 Evalu-ation set. To better handle spectral information, crucial to de-tect spoofing, we adopt the WaveletCNN and VGG16 archi-tectures which outperform all baselines. Finally, we find that robustness of countermeasure systems degrades dramatically when provided with speech samples degraded through VoIP network transmission or mismatching audio compression. Abderrahim Fathan, Jahangir Alam 0001, Woo Hyun Kang |
ICME | 1 |
| 2022 | MIM-DG: Mutual information minimization-based domain generalization for speaker verification
Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan |
INTERSPEECH | 3 |
| 2022 | Mixup regularization strategies for spoofing countermeasure system
Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan |
INTERSPEECH | 3 |
| 2022 | End-to-end framework for spoof-aware speaker verification
Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan |
INTERSPEECH | 3 |
| 2022 | Deep learning-based end-to-end spoken language identification system for domain-mismatched scenarioabstractDomain mismatch is a critical issue when it comes to spoken language identification. To overcome the domain mismatch problem, we have applied several architectures and deep learning strategies which have shown good results in cross-domain speaker verification tasks to spoken language identification. Our systems were evaluated on the Oriental Language Recognition (OLR) Challenge 2021 Task 1 dataset, which provides a set of cross-domain language identification trials. Among our experimented systems, the best performance was achieved by using the mel frequency cepstral coefficient (MFCC) and pitch features as input and training the ECAPA-TDNN system with a flow-based regularization technique, which resulted in a Cavg of 0.0631 on the OLR 2021 progress set. Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan |
LREC | 3 |
| 2022 | Flow-ER: A Flow-Based Embedding Regularization Strategy for Robust Speech Representation LearningabstractOver the recent years, various deep learning-based embedding methods were proposed. Although the deep learning-based embedding extraction methods have shown good performance in numerous tasks including speaker verification, language identification and anti-spoofing, their performance is limited when it comes to mismatched conditions due to the variability within them unrelated to the main task. In order to alleviate this problem, we propose a novel training strategy that regularizes the embedding network to have minimum information about the nuisance attributes. To achieve this, our proposed method directly incorporates the information bottleneck scheme into the training process, where the mutual information is estimated using an auxiliary normalizing flow network. The performance of the proposed method is evaluated on different speech processing tasks and found to provide improvement over the standard training strategy in all experimentations. Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan |
SLT | 3 |
| 2021 | Hybrid Network with Multi-Level Global-Local Statistics Pooling for Robust Text-Independent Speaker RecognitionabstractIn this paper, we propose a new hybrid system for extracting a speaker embedding vector. More specifically, the proposed system employs a multi-level global-local statistics pooling method in order to aggregate the speaker information within short time-span and utterance-level context. In order to evaluate the proposed system, a set of experiments on the NIST SRE 2016, Short-duration speaker verification (SdSV) Challenge 2021, and VoxCeleb datasets were conducted, and the proposed hybrid network was able to outperform the conventional approaches trained on the same dataset. Moreover, our experiments showed that the proposed system is able to achieve stable performance even when using a relatively smaller dataset, which highlights the efficiency of the proposed system in extracting the speaker-dependent information. Woo Hyun Kang, Jahangir Alam 0001, Abderrahim Fathan |
ASRU | 3 |