Hardik B. Sailor

dblp:136/3700 · DBLP profile ↗
← Back
22ranked-venue papers
11as first author
9since 2021 · last 2025
0000-0001-6872-5153ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 10 first-author · 9 since 2021Artificial intelligence and machine learning · 18 · 10 first-author · 7 since 2021
YearPublicationVenuePosition
2025 MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
abstract
While self-supervised learning (SSL)-based models have boosted audio deepfake detection accuracy, fully finetuning them is computationally expensive. To address this, we propose a parameter-efficient framework that combines Low-Rank Adaptation with a Mixture-of-Experts router, called Mixture of LoRA Experts (MoLEx). It preserves pre-trained knowledge of SSL models while efficiently finetuning only selected experts, reducing training costs while maintaining robust performance. The observed utility of experts during inference shows the router reactivates the same experts for similar attacks but switches to other experts for novel spoofs, confirming MoLEx’s domainaware adaptability. MoLEx additionally offers flexibility for domain adaptation by allowing extra experts to be trained without modifying the entire model. We mainly evaluate our approach on the ASVSpoof 5 dataset and achieve the state-of-the-art (SOTA) equal error rate (EER) of 5.56% on the evaluation set without augmentation.
Zihan Pan, Hardik B. Sailor
ASRU2
2025 Diversity and complementarity of speech encoders across diverse tasks in a multi-modal large language model
abstract
A Large Language Model (LLM) can be extended to understand speech inputs by using a speech encoder to compute embeddings from the speech, which are then used with a text prompt. Diverse information is expressed in speech and a wide variety of tasks can be performed. Different speech encoders may specialise toward different information types and tasks. This complementarity can be leveraged upon by using multiple speech encoders. This paper presents a comprehensive analysis of the diversity and complementarity between open-source speech encoders, when used in a multi-modal LLM framework. Experiments identify the encoders that excel in each type of downstream task, thereby guiding future system design. The diversity between encoders is measured, showing that Whisper tends to behave more differently. Diversity between encoders is compared across tasks, showing that semantic tasks tend to yield more diverse predictions. Early and late fusion show that complementarity can yield improvements.
Jeremy H. M. Wong, Muhammad Huzaifah 0001, Hardik B. Sailor, Kye Min Tan, Bin Wang 0040, Qiongqiong Wang, Xunlong Zou, Nancy F. Chen, AiTi Aw
ASRU3
2025 Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation
Qiongqiong Wang, Hardik B. Sailor, Tianchi Liu 0004, AiTi Aw
INTERSPEECH2
2024 Attentive Merging of Hidden Embeddings from Pre-trained Speech Model for Anti-spoofing Detection
Zihan Pan, Tianchi Liu 0004, Hardik B. Sailor, Qiongqiong Wang
INTERSPEECH3
2024 Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CTRSVDD) Challenge 2024
abstract
This work details our approach to achieving a leading system with a 1.79% pooled equal error rate (EER) on the evaluation set of the Controlled Singing Voice Deepfake Detection (CtrSVDD). The rapid advancement of generative AI models presents significant challenges for detecting AI-generated deepfake singing voices, attracting increased research attention. The Singing Voice Deepfake Detection (SVDD) Challenge 2024 aims to address this complex task. In this work, we explore the ensemble methods, utilizing speech foundation models to develop robust singing voice anti-spoofing systems. We also introduce a novel Squeeze-and-Excitation Aggregation (SEA) method, which efficiently and effectively integrates representation features from the speech foundation models, surpassing the performance of our other individual systems. Evaluation results confirm the efficacy of our approach in detecting deepfake singing voices. The codes can be accessed at https://github.com/Anmol2059/SVDD2024.
Anmol Guragain, Tianchi Liu 0004, Zihan Pan, Hardik B. Sailor, Qiongqiong Wang
SLT4
2024 Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-Spoofing
abstract
The effects of language mismatch impact speech anti-spoofing systems, while investigations and quantification of these effects remain limited. Existing anti-spoofing datasets are mainly in English, and the high cost of acquiring multilingual datasets hinders training language-independent models. We initiate this work by evaluating top-performing speech anti-spoofing systems that are trained on English data but tested on other languages, observing notable performance declines. We propose an innovative approach - Accent-based data expansion via TTS (ACCENT), which introduces diverse linguistic knowledge to monolingual-trained models, improving their cross-lingual capabilities. We conduct experiments on a large-scale dataset consisting of over 3 million samples, including 1.8 million training samples and nearly 1.2 million testing samples across 12 languages. The language mismatch effects are preliminarily quantified and remarkably reduced over 15% by applying the proposed ACCENT. This easily implementable method shows promise for multilingual and low-resource language scenarios.
Tianchi Liu 0004, Ivan Kukanov, Zihan Pan, Qiongqiong Wang, Hardik B. Sailor, Kong-Aik Lee
SLT5
2022 Spectral Modification Based Data Augmentation For Improving End-to-End ASR For Children's Speech
abstract
Training a robust Automatic Speech Recognition (ASR) system for children's speech recognition is a challenging task due to inherent differences in acoustic attributes of adult and child speech and scarcity of publicly available children's speech dataset. In this paper, a novel segmental spectrum warping and perturbations in formant energy are introduced, to generate a children-like speech spectrum from that of an adult's speech spectrum. Then, this modified adult spectrum is used as augmented data to improve end-to-end ASR systems for children's speech recognition. The proposed data augmentation methods give 6.5% and 6.1% relative reduction in WER on children dev and test sets respectively, compared to the vocal tract length perturbation (VTLP) baseline system trained on Librispeech 100 hours adult speech dataset. When children's speech data is added in training with Librispeech set, it gives a 3.7 % and 5.1% relative reduction in WER, compared to the VTLP baseline system.
Vishwanath Pratap Singh, Hardik B. Sailor, Supratik Bhattacharya
INTERSPEECH2
2021 Warped Ensembles: A Novel Technique for Improving CTC Based End-to-End Speech Recognition
abstract
Combining outputs from predictive models trained for similar tasks generally perform better than using a single model. Models trained on different domains could for example help in improving the performance with the help of information from complementary domains. The weights in the ensemble could also be adjusted to best suit the desired target domain. However, in End-to-End (E2E) based speech recognition solutions, this is not a simple task owing to the lack of time alignments in the outputs. This work presents a novel ensembling technique for E2E ASR systems trained using Connectionist Temporal Classification (CTC) loss, which uses a multi-sequence alignment technique for synchronising output time-frames across models, and then combines the outputs using weights obtained using a calibration technique introduced in our earlier work. Several experiments are conducted on multi-accent English data and the Librispeech corpus. Word Error Rates (WER) are evaluated and compared against the performance of ROVER (a system combination technique), and the best model in the system. When compared to the best model, average relative improvements of 6.46% and 8.94% are observed in unadapted in-domain and out-of-domain experiments respectively. Comparison to ROVER shows average relative improvements of 19.29% and 9.21% in in-domain and out-of-domain experiments.
Kiran Praveen, Hardik B. Sailor
ASRU2
2021 SRI-B End-to-End System for Multilingual and Code-Switching ASR Challenges for Low Resource Indian Languages
Hardik B. Sailor, Kiran Praveen, Vikas Agrawal
Interspeech1
2020 Multilingual Speech Recognition Using Language-Specific Phoneme Recognition as Auxiliary Task for Indian Languages
Hardik B. Sailor, Thomas Hain
INTERSPEECH1
2019 Unsupervised Adaptation of Acoustic Models for ASR Using Utterance-Level Embeddings from Squeeze and Excitation Networks
abstract
This paper proposes the adaptation of neural network-based acoustic models using a Squeeze-and-Excitation (SE) network for automatic speech recognition (ASR). In particular, this work explores to use the SE network to learn utterance-level embeddings. The acoustic modelling is performed using Light Gated Recurrent Units (LiGRU). The utterance embed-dings are learned from hidden unit activations jointly with LiGRU and used to scale respective activations of hidden layers in the LiGRU network. The advantage of such approach is that it does not require domain labels, such as speakers and noise to be known in order to perform the adaptation, thereby providing unsupervised adaptation. The global average and attentive pooling are applied on hidden units to extract utterance-level information that represents the speakers and acoustic conditions. ASR experiments were carried out on the TIMIT and Aurora 4 corpora. The proposed model achieves better performance on both the datasets compared to their respective baselines with relative improvements of 5.59% and 5.54% for TIMIT and Aurora 4 database, respectively. These experiments show the potential of using the conditioning information learned via utterance embeddings in the SE network to adapt acoustic models for speakers, noise, and other acoustic conditions.
Hardik B. Sailor, Salil Deena, Md Asif Jalal, Rasa Lileikyte, Thomas Hain
ASRU1
2019 Whether to Pretrain DNN or not?: An Empirical Analysis for Voice Conversion
Nirmesh J. Shah, Hardik B. Sailor, Hemant A. Patil
INTERSPEECH2
2018 DA-IICT/IIITV System for Low Resource Speech Recognition Challenge 2018
Hardik B. Sailor, Maddala Venkata Siva Krishna, Diksha Chhabra, Ankur T. Patil, Madhu R. Kamble, Hemant A. Patil
INTERSPEECH1
2018 Auditory Filterbank Learning for Temporal Modulation Features in Replay Spoof Speech Detection
Hardik B. Sailor, Madhu R. Kamble, Hemant A. Patil
INTERSPEECH1
2018 Auditory Filterbank Learning Using ConvRBM for Infant Cry Classification
Hardik B. Sailor, Hemant A. Patil
INTERSPEECH1
2017 Unsupervised Filterbank Learning Using Convolutional Restricted Boltzmann Machine for Environmental Sound Classification
Hardik B. Sailor, Dharmesh M. Agrawal, Hemant A. Patil
INTERSPEECH1
2017 Unsupervised Representation Learning Using Convolutional Restricted Boltzmann Machine for Spoof Speech Detection
Hardik B. Sailor, Madhu R. Kamble, Hemant A. Patil
INTERSPEECH1
2016 Filterbank learning using Convolutional Restricted Boltzmann Machine for speech recognition
abstract
Convolutional Restricted Boltzmann Machine (ConvRBM) as a model for speech signal is presented in this paper. We have developed ConvRBM with sampling from noisy rectified linear units (NReLUs). ConvRBM is trained in an unsupervised way to model speech signal of arbitrary lengths. Weights of the model can represent an auditory-like filterbank. Our proposed learned filterbank is also nonlinear with respect to center frequencies of subband filters similar to standard filterbanks (such as Mel, Bark, ERB, etc.). We have used our proposed model as a front-end to learn features and applied to speech recognition task. Performance of ConvRBM features is improved compared to MFCC with relative improvement of 5% on TIMIT test set and 7% on WSJ0 database for both Nov'92 test sets using GMM-HMM systems. With DNN-HMM systems, we achieved relative improvement of 3% on TIMIT test set over MFCC and Mel filterbank (FBANK). On WSJ0 Nov'92 test sets, we achieved relative improvement of 4-14% using ConvRBM features over MFCC features and 3.6-5.6% using ConvRBM filterbank over FBANK features.
Hardik B. Sailor, Hemant A. Patil
ICASSP1
2016 Native Language Identification Using Spectral and Source-Based Features
Avni Rajpal, Tanvina B. Patel, Hardik B. Sailor, Maulik C. Madhavi, Hemant A. Patil, Hiroya Fujisaki
INTERSPEECH3
2016 Unsupervised Deep Auditory Model Using Stack of Convolutional RBMs for Speech Recognition
Hardik B. Sailor, Hemant A. Patil
INTERSPEECH1
2016 Novel Unsupervised Auditory Filterbank Learning Using Convolutional RBM for Speech Recognition
abstract
To learn auditory filterbanks, recently, we have proposed an unsupervised learning model based on convolutional restricted Boltzmann machine (RBM) with rectified linear units. In this paper, theory, training algorithm of our proposed model, and detailed analysis of learned filterbank are being presented. Learning of the model with different databases shows that the model is able to learn cochlear-like impulse responses that are localized in frequency-domain. An auditory-like scale obtained from filterbanks learned from clean and noisy datasets resembles the Mel scale, which is known to mimic perceptually relevant aspect of speech. We have experimented with both cepstral (denoted as ConvRBM-CC) as well as filterbank features (denoted as ConvRBM-BANK). On large vocabulary continuous speech recognition task, we achieved relative improvement of 7.21-17.8% in word error rate (WER) compared to Mel frequency cepstral coefficient (MFCC) features and 1.35-6.82% compared to Mel filterbank (FBANK) features. On AURORA 4 multicondition training database, the relative improvement in WER by 4.8-13.65% was achieved using a Hybrid Deep Neural Network-Hidden Markov Model (DNN-HMM) system with ConvRBM-CC features. Using ConvRBM-BANK features, we achieve absolute reduction of 1.25-3.85% in WER on AURORA 4 test sets compared to FBANK features. A context-dependent DNN-HMM system further improves performance with a relative improvement of 3.6-4.6% on an average for bigram 5k and tri-gram 5k language models. Hence, our proposed learned filterbank performs better than traditional MFCC and Mel-filterbank features for both clean and multicondition automatic speech recognition (ASR) tasks. A system combination of ConvRBM-BANK and FBANK features further improve performance in all ASR tasks. Cross-domain experiments where subband filters trained on one database are used for the ASR task of another database show that model learns generalized representations of speech signals.
Hardik B. Sailor, Hemant A. Patil
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 Effectiveness of PLP-based phonetic segmentation for speech synthesis
abstract
In this paper, use of Viterbi-based algorithm and spectral transition measure (STM)-based algorithm for the task of speech data labeling is being attempted. In the STM framework, we propose use of several spectral features such as recently proposed cochlear filter cepstral coefficients (CFCC), perceptual linear prediction cepstral coefficients (PLPCC) and RelAtive SpecTrAl (RASTA)-based PLPCC in addition to Mel frequency cepstral coefficients (MFCC) for phonetic segmentation task. To evaluate effectiveness of these segmentation algorithms, we require manual accurate phoneme-level labeled data which is not available for low resourced languages such as Gujarati (one of the official languages of India). In order to measure effectiveness of various segmentation algorithms, HMM-based speech synthesis system (HTS) for Gujarati has been built. From the subjective and objective evaluations, it is observed that Viterbi-based and STM with PLPCC-based segmentation algorithms work better than other algorithms.
Nirmesh J. Shah, Bhavik B. Vachhani, Hardik B. Sailor, Hemant A. Patil
ICASSP3