Amit Meghanani

dblp:232/4082 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0002-0811-274XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Towards a Unified Benchmark for Arabic Pronunciation Assessment: Qur'anic Recitation as Case Study
Yassine El Kheir, Omnia Ibrahim, Amit Meghanani, Nada Almarwani, Hawau Olamide Toyin, Sadeen Alharbi, Modar Alfadly, Lamya Alkanhal, Ibrahim Selim, Shehab Elbatal, Salima Mdhaffar, Thomas Hain, Yasser Hifny, Mostafa Shahin, Ahmed Ali 0002
INTERSPEECH3
2024 Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
abstract
Acoustic word embeddings (AWEs) are vector representations of spoken words.An effective method for obtaining AWEs is the Correspondence Auto-Encoder (CAE).In the past, the CAE method has been associated with traditional MFCC features.Representations obtained from self-supervised learning (SSL)-based speech models such as HuBERT, Wav2vec2, etc., are outperforming MFCC in many downstream tasks.However, they have not been well studied in the context of learning AWEs.This work explores the effectiveness of CAE with SSL-based speech representations to obtain improved AWEs.Additionally, the capabilities of SSL-based speech models are explored in cross-lingual scenarios for obtaining AWEs.Experiments are conducted on five languages: Polish, Portuguese, Spanish, French, and English.HuBERT-based CAE model achieves the best results for word discrimination in all languages, despite Hu-BERT being pre-trained on English only.Also, the HuBERT-based CAE model works well in cross-lingual settings.It outperforms MFCCbased CAE models trained on the target languages when trained on one source language and tested on target languages.
Amit Meghanani, Thomas Hain
EACL (1)1
2024 SCORE: Self-Supervised Correspondence Fine-Tuning for Improved Content Representations
abstract
There is a growing interest in cost-effective self-supervised fine-tuning (SSFT) of self-supervised learning (SSL)-based speech models to obtain task-specific representations. These task-specific representations are used for robust performance on various downstream tasks by fine-tuning on the labelled data. This work presents a cost-effective SSFT method named Self-supervised Correspondence (SCORE) fine-tuning to adapt the SSL speech representations for content-related tasks. The proposed method uses a correspondence training strategy, aiming to learn similar representations from perturbed speech and original speech. Commonly used data augmentation techniques for content-related tasks (ASR) are applied to obtain perturbed speech. SCORE fine-tuned HuBERT outperforms the vanilla HuBERT on SUPERB benchmark with only a few hours of fine-tuning (< 5 hrs) on a single GPU for automatic speech recognition, phoneme recognition, and query-by-example tasks, with relative improvements of 1.09%, 3.58%, and 12.65%, respectively. SCORE provides competitive results with the recently proposed SSFT method SPIN, using only 1/3 of the processed speech compared to SPIN.
Amit Meghanani, Thomas Hain
ICASSP1
2024 LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks
Amit Meghanani, Thomas Hain
INTERSPEECH1
2023 Deriving Translational Acoustic Sub-Word Embeddings
abstract
There is a growing interest in understanding the representational geometry of acoustic word embeddings (AWEs), which are fixed-dimensional representations of spoken words. However, not much research has been conducted on acoustic sub-word embeddings (ASWEs), which can provide a better understanding of the AWE space. This work focuses on decomposing AWEs to obtain ASWEs while retaining the ability to reconstruct AWEs by translating ASWEs in the embedding space, under constrained settings. Initially, high-quality AWEs are obtained with an Average Precision (AP) score of 0.97 on the word discrimination task. Subsequently, ASWEs are derived through the decomposition of AWEs. Three adapted versions of the AP metric, utilized for evaluating the quality of the derived ASWEs and their translational properties, are proposed. The results demonstrate that the derived ASWEs exhibit high quality, and the reconstruction of AWEs from the ASWEs is achievable by translating them in the embedding space.
Amit Meghanani, Thomas Hain
ASRU1
2021 An Exploration of Log-Mel Spectrogram and MFCC Features for Alzheimer's Dementia Recognition from Spontaneous Speech
abstract
In this work, we explore the effectiveness of log-Mel spectrogram and MFCC features for Alzheimer's dementia (AD) recognition on ADReSS challenge dataset. We use three different deep neural networks (DNN) for AD recognition and mini-mental state examination (MMSE) score prediction: (i) convolutional neural network followed by a long-short term memory network (CNN-LSTM), (ii) pre-trained ResNet18 network followed by LSTM (ResNet-LSTM), and (iii) pyramidal bidirectional LSTM followed by a CNN (pBLSTM-CNN). CNN-LSTM achieves an accuracy of 64.58% with MFCC features and ResNet-LSTM achieves an accuracy of 62.5% using log-Mel spectrograms. pBLSTM-CNN and ResNet-LSTM models achieve root mean square errors (RMSE) of 5.9 and 5.98 in the MMSE score prediction, using the log-Mel spectrograms. Our results beat the baseline accuracy (62.5%) and RMSE (6.14) reported for acoustic features on ADReSS challenge dataset. The results suggest that log-Mel spectrograms and MFCCs are effective features for AD recognition problem when used with DNN models.
Amit Meghanani, Chandran Savithri Anoop, A. G. Ramakrishnan
SLT1
2020 Pitch-synchronous Discrete Cosine Transform Features for Speaker Identification and Verification
abstract
We propose a feature called pitch-synchronous discrete cosine transform (PS-DCT), derived from the voiced part of the speech for speaker identification (SID) and verification (SV) tasks. PS-DCT features are derived from the �time-domain, quasi-stationary waveform shape� of the voiced sounds. We test our PS-DCT feature on TIMIT, Mandarin and YOHO datasets. On TIMIT with 168 and Mandarin with 855 speakers, we obtain the SID accuracies of 99.4 and 96.1, respectively, using a Gaussian mixture model-based classifier. In the i-vector-based SV framework, fusing the �PS-DCT based system� with the �MFCC-based system� at the score level reduces the equal error rate (EER) for both YOHO and Mandarin datasets. In the case of limited test data and session variabilities, we obtain a significant reduction in EER, up to 5.8 (for test data of duration < 3 sec). Copyright © 2020 by SCITEPRESS � Science and Technology Publications, Lda. All rights reserved.
Amit Meghanani, A. G. Ramakrishnan
ICPRAM1