Hemant Yadav

dblp:02/7350 · DBLP profile ↗
← Back
10ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0003-3057-1456ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing AI Agent Evaluation through Critical Step Identification
Sambit Kumar Panda, Subodh Kumar Chaturvedi, Mandlem Chakradhar Reddy, Hemant Yadav
COMPSAC4
2025 JOOCI: a Novel Method for Learning Comprehensive Speech Representations
abstract
Information in speech can be categorized into two groups: Content (what is being said, such as linguistics) and Other (how it is expressed such as information about speaker and paralinguistic features). Current self-supervised learning (SSL) methods are shown to divide the model’s representational-depth or layers in two, with earlier layers specializing in Other and later layers in Content related tasks. This layer-wise division is inherently sub-optimal, as neither information type can use all layers to build hierarchical representations. To address this, we propose JOOCI, a novel speech representation learning method that does not compromise on the representational-depth for either information type. JOOCI outperforms WavLM by 26.5% (relative), and other models of similar size ($\mathbf{1 0 0 M}$ parameters), when evaluated on two speaker recognition and two language tasks from the SUPERB benchmark, demonstrating its effectiveness in Jointly Optimizing Other and Content Information (JOOCI).
Hemant Yadav, Sunayana Sitaram, Rajiv Ratn Shah
ASRU1
2024 MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
Hemant Yadav, Sunayana Sitaram, Rajiv Ratn Shah
INTERSPEECH1
2024 NOA-LSTM: An efficient LSTM cell architecture for time series forecasting
Hemant Yadav, Amit Thakkar
Expert Syst. Appl.1
2023 Mask-Net: Learning Context Aware Invariant Features Using Adversarial Forgetting (Student Abstract)
abstract
Training a robust system, e.g., Speech to Text (STT), requires large datasets. Variability present in the dataset, such as unwanted nuances and biases, is the reason for the need for large datasets to learn general representations. In this work, we propose a novel approach to induce invariance using adversarial forgetting (AF). Our initial experiments on learning invariant features such as accent on the STT task achieve better generalizations in terms of word error rate (WER) compared to traditional models. We observe an absolute improvement of 2.2% and 1.3% on out-of-distribution and in-distribution test sets, respectively.
Hemant Yadav, Rajiv Ratn Shah
AAAI1
2023 Partial Rank Similarity Minimization Method for Quality MOS Prediction of Unseen Speech Synthesis Systems in Zero-Shot and Semi-Supervised Setting
abstract
This paper introduces a novel objective function for quality mean opinion score (MOS) prediction of unseen speech synthesis systems. The proposed function measures the similarity of relative positions of predicted MOS values, in a mini-batch, rather than the actual MOS values. That is the partial rank similarity is measured $(\mathcal{P}RS)$ rather than the individual MOS values as with the L1 loss. Our experiments on out-of-domain speech synthesis systems demonstrate that the $\mathcal{P} R S$ outperforms L1 loss in zero-shot and semi-supervised settings, exhibiting stronger correlation with ground truth. These findings highlight the importance of considering rank order, as done by $\mathcal{P}RS$, when training MOS prediction models. We also argue that mean squared error and linear correlation coefficient metrics may be unreliable for evaluating MOS prediction models. In conclusion, $\mathcal{P} R S$-trained models provide a robust framework for evaluating speech quality and offer insights for developing high-quality speech synthesis systems. Code and models are available at github.com/nii-yamagishilab/partial_rank_similarity/
Hemant Yadav, Erica Cooper, Junichi Yamagishi, Sunayana Sitaram, Rajiv Ratn Shah
ASRU1
2023 Analysing the Masked Predictive Coding Training Criterion for Pre-Training a Speech Representation Model
abstract
Recent developments in pre-trained speech representation utilizing self-supervised learning (SSL) have yielded exceptional results on a variety of downstream tasks. One such technique, known as masked predictive coding (MPC), has been employed by some of the most high-performing models. In this study, we investigate the impact of MPC loss on the type of information learnt at various layers in the HuBERT model, using nine probing tasks. Our findings indicate that the amount of content information learned at various layers of the HuBERT model has a positive correlation to the MPC loss. Additionally, it is also observed that any speaker-related information learned at intermediate layers of the model, is an indirect consequence of the learning process, and therefore cannot be controlled using the MPC loss. These findings may serve as inspiration for further research in the speech community, specifically in the development of new pre-training tasks or the exploration of new pre-training criterion’s that directly preserves both speaker and content information at various layers of a learnt model.
Hemant Yadav, Sunayana Sitaram, Rajiv Ratn Shah
ICASSP1
2022 Intent classification using pre-trained language agnostic embeddings for low resource languages
abstract
Building Spoken Language Understanding (SLU) systems that do not rely on language specific Automatic Speech Recognition (ASR) is an important yet less explored problem in language processing.In this paper, we present a comparative study aimed at employing a pre-trained language agnostic acoustic model to perform SLU in low resource scenarios.Specifically, we use three different embedding settings extracted using Allosaurus, a pre-trained universal phone decoder: (1) Phonelabels (2) Panphone, and (3) Allo embeddings (proposed by us).These embeddings are then used in identifying the spoken intent.We perform experiments across three different languages: English, Sinhala, and Tamil each with different data sizes to simulate high, medium, and low resource scenarios.Our system improves on the state-of-the-art (SOTA) intent classification accuracy by absolute 2.11% for Sinhala and 7.00% for Tamil and achieves competitive results in English.Furthermore, we also present a quantitative analysis to show how the performance scales with the number of training examples.
Hemant Yadav, Akshat Gupta, Sai Krishna Rallabandi, Alan W. Black, Rajiv Ratn Shah
INTERSPEECH1
2022 A Survey of Multilingual Models for Automatic Speech Recognition
abstract
Although Automatic Speech Recognition (ASR) systems have achieved human-like performance for a few languages, the majority of the world’s languages do not have usable systems due to the lack of large speech datasets to train these models. Cross-lingual transfer is an attractive solution to this problem, because low-resource languages can potentially benefit from higher-resource languages either through transfer learning, or being jointly trained in the same multilingual model. The problem of cross-lingual transfer has been well studied in ASR, however, recent advances in Self Supervised Learning are opening up avenues for unlabeled speech data to be used in multilingual ASR models, which can pave the way for improved performance on low-resource languages. In this paper, we survey the state of the art in multilingual ASR models that are built with cross-lingual transfer in mind. We present best practices for building multilingual models from research across diverse languages and techniques, discuss open questions and provide recommendations for future work.
Hemant Yadav, Sunayana Sitaram
LREC1
2020 End-to-End Named Entity Recognition from English Speech
abstract
Named entity recognition (NER) from text has been a widely studied problem and usually extracts semantic information from text. Until now, NER from speech is mostly studied in a two-step pipeline process that includes first applying an automatic speech recognition (ASR) system on an audio sample and then passing the predicted transcript to a NER tagger. In such cases, the error does not propagate from one step to another as both the tasks are not optimized in an end-to-end (E2E) fashion. Recent studies confirm that integrated approaches (e.g., E2E ASR) outperform sequential ones (e.g., phoneme based ASR). In this paper, we introduce a first publicly available NER annotated dataset for English speech and present an E2E approach, which jointly optimizes the ASR and NER tagger components. Experimental results show that the proposed E2E approach outperforms the classical two-step approach. We also discuss how NER from speech can be used to handle out of vocabulary (OOV) words in an ASR system.
Hemant Yadav, Sreyan Ghosh, Yi Yu 0001, Rajiv Ratn Shah
INTERSPEECH1