VLDB 2026 Research / reviewers in the wild / expert
Srinivasan Umesh
dblp:36/5441
· DBLP profile ↗
86ranked-venue papers
14as first author
23since 2021 · last 2025
0000-0002-5957-1444ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 81 · 12 first-author · 22 since 2021Artificial intelligence and machine learning · 49 · 6 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MADASR 2.0: Multi-Lingual Multi-Dialect ASR Challenge in 8 Indian LanguagesabstractWe present MADASR 2.0, a challenge at ASRU 2025 aimed at advancing multilingual and multidialectal automatic speech recognition (ASR) in low-resource Indian languages. Building on the 2023 edition, it introduces a subset of the RESPIN corpus, over 1200 hours of read speech across 8 languages and 33 dialects, with test sets including both read and spontaneous speech. The challenge comprises four tracks varying by training data size and external resource usage, and supports auxiliary tasks like language and dialect identification. We detail the dataset, tasks, baselines, and submissions and analyse trends across tracks and speech styles. Results highlight the continued difficulty of spontaneous ASR, the benefits of multitask and transfer learning, and effective strategies for building dialect-aware ASR systems. MADASR 2.0 offers a standardised benchmark to support future research on inclusive and scalable ASR for linguistically diverse populations. Sumit Sharma 0016, Deekshitha G, Abhayjeet Singh, Amartyaveer, Sathvik Udupa, Sandhya Badiger, Sanjeev Khudanpur, Sunayana Sitaram, Srinivasan Umesh, Bhuvana Ramabhadran, Brian Kingsbury, Hema A. Murthy, Srikanth S. Narayanan, Howard Lakougna, Prasanta Kumar Ghosh |
ASRU | 10 |
| 2024 | Stable Distillation: Regularizing Continued Pre-Training for Low-Resource Automatic Speech RecognitionabstractContinued self-supervised (SSL) pre-training for adapting existing SSL models to the target domain has shown to be extremely effective for low-resource Automatic Speech Recognition (ASR). This paper proposes Stable Distillation, a simple and novel approach for SSL-based continued pre-training that boosts ASR performance in the target domain where both labeled and unlabeled data are limited. Stable Distillation employs self-distillation as regularization for continued pre-training, alleviating the over-fitting issue, a common problem continued pre-training faces when the source and target domains differ. Specifically, first, we perform vanilla continued pre-training on an initial SSL pre-trained model on the target domain ASR dataset and call it the teacher. Next, we take the same initial pre-trained model as a student to perform continued pre-training while enforcing its hidden representations to be close to that of the teacher (via MSE loss). This student is then used for downstream ASR fine-tuning on the target dataset. In practice, Stable Distillation outperforms all our baselines by 0.8 - 7 WER when evaluated in various experimental settings1. Ashish Seth, Sreyan Ghosh, Srinivasan Umesh, Dinesh Manocha |
ICASSP | 3 |
| 2024 | FusDom: Combining in-Domain and Out-of-Domain Knowledge for Continuous Self-Supervised LearningabstractContinued pre-training (CP) offers multiple advantages, like target domain adaptation and the potential to exploit the continuous stream of unlabeled data available online. However, continued pre-training on out-of-domain distributions often leads to catastrophic forgetting of previously acquired knowledge, leading to sub-optimal ASR performance. This paper presents FusDom, a simple and novel methodology for SSL-based continued pre-training. FusDom learns speech representations that are robust and adaptive yet not forgetful of concepts seen in the past. Instead of solving the SSL pre-text task on the output representations of a single model, FusDom leverages two identical pre-trained SSL models, a teacher and a student, with a modified pre-training head to solve the CP SSL pre-text task. This head employs a cross-attention mechanism between the representations of both models while only the student receives gradient updates and the teacher does not. Finally, the student is fine-tuned for ASR. In practice, FusDom outperforms all our baselines across settings significantly, with WER improvements in the range of 0.2 WER - 7.3 WER in the target domain, while retaining the performance in the earlier domain1. Ashish Seth, Sreyan Ghosh, Srinivasan Umesh, Dinesh Manocha |
ICASSP | 3 |
| 2024 | All Ears: Building Self-Supervised Learning based ASR models for Indian Languages at scale
Vasista Sai Lodagala, Abhishek Biswas, Shoutrik Das, Jordan Fernandes, Srinivasan Umesh |
INTERSPEECH | 5 |
| 2024 | Lite ASR Transformer: A Light Weight Transformer Architecture For Automatic Speech RecognitionabstractTransformers are popular sequence-to-sequence models but have large number of parameters and high compute requirements. As an initiative to reduce the energy demand by Transformer models and to get better Transformer models for edge devices, we propose a light weight Transformer in this paper. We attempt to reduce the compute and carbon footprint of the original Transformer architecture by incorporating architectural modifications. The proposed modifications reduce the Transformer parameters by 42.7 % relative to the original Transformer of same depth and width. Our automatic speech recognition experiments on LibriSpeech, SPGISpeech and GigaSpeech datasets show that the proposed light weight Transformer has negligible ASR performance degradation. The compute requirements also reduce by 23% relative to the original Transformer of same depth and width. We also show that the proposed Lite ASR Transformer has acceptable convergence and also the latency is 20% lesser relative to the original Transformer of same depth and width. Narla John Metilda Sagaya Mary, Srinivasan Umesh |
SLT | 2 |
| 2023 | Towards Developing State-of-The-Art TTS Synthesisers for 13 Indian Languages with Signal Processing Aided AlignmentsabstractEnd-to-end (E2E) systems synthesise high-quality speech, but this typically requires a large amount of data. As E2E synthesis progressed from Tacotron to FastSpeech2, it became evident that features representing prosody, particularly subword durations, are important for error-free synthesis. Variants of FastSpeech use a teacher model or forced alignments for training. This paper uses signal processing cues in tandem with forced alignment to produce accurate phone boundaries for the training data. As a result of better duration modelling, good-quality synthesisers are developed. Evaluations indicate that systems developed using the proposed signal processing-aided approach are better than systems developed using other alignment approaches, especially in low-resource scenarios. Our systems also outperform the existing best TTS systems available for 13 Indian languages. Anusha Prakash 0001, Srinivasan Umesh, Hema A. Murthy |
ASRU | 2 |
| 2023 | MAST: Multiscale Audio Spectrogram TransformersabstractWe present Multiscale Audio Spectrogram Transformer (MAST) for audio classification, which brings the concept of multiscale feature hierarchies to the Audio Spectrogram Transformer (AST) [1]. Given an input audio spectrogram, we first patchify and project it into an initial temporal resolution and embedding dimension, post which the multiple stages in MAST progressively expand the embedding dimension while reducing the temporal resolution of the input. We use a pyramid structure that allows early layers of MAST operating at a high temporal resolution but low embedding space to model simple low-level acoustic information and deeper temporally coarse layers to model high-level acoustic information with high-dimensional embeddings. We also extend our approach to present a new Self-Supervised Learning (SSL) method called SS-MAST, which calculates a symmetric contrastive loss between latent representations from a student and a teacher encoder, leveraging patch-drop, a novel audio augmentation approach that we introduce. In practice, MAST significantly outperforms AST by an average accuracy of 3.4% across 8 speech and non-speech tasks from the LAPE Benchmark [2], achieving state-of-the-art results on keyword spotting in Speech Commands. Additionally, our proposed SS-MAST achieves an absolute average improvement of 2.6% over the previously proposed SSAST [3]1. Sreyan Ghosh, Ashish Seth, Srinivasan Umesh, Dinesh Manocha |
ICASSP | 3 |
| 2023 | Data2vec-Aqc: Search for the Right Teaching Assistant in the Teacher-Student Training SetupabstractIn this paper, we propose a new Self-Supervised Learning (SSL) algorithm called data2vec-aqc, for speech representation learning from unlabeled speech data. Our goal is to improve SSL for speech in domains where both unlabeled and labeled data are limited. Building on the recently introduced data2vec [1], we introduce additional modules to the data2vec framework that leverage the benefit of data augmentations, quantized representations, and clustering. The interaction between these modules helps solve the cross-contrastive loss as an additional self-supervised objective. data2vec-aqc achieves up to 14.1% and 20.9% relative WER improvement over the existing state-of-the-art data2vec system over the test-clean and test-other sets, respectively of LibriSpeech, without the use of any language model (LM). Our proposed model also achieves up to 17.8% relative WER gains over the baseline data2vec when fine-tuned on a subset of the Switchboard dataset. Code: https://github.com/Speech-Lab-IITM/data2vec-aqc. Vasista Sai Lodagala, Sreyan Ghosh, Srinivasan Umesh |
ICASSP | 3 |
| 2023 | SLICER: Learning Universal Audio Representations Using Low-Resource Self-Supervised Pre-TrainingabstractWe present a new Self-Supervised Learning (SSL) approach to pre-train encoders on unlabeled audio data that reduces the need for large amounts of labeled data for audio and speech classification. Our primary aim is to learn au-dio representations that can generalize across a large vari-ety of speech and non-speech tasks in a low-resource un-labeled audio pre-training setting. Inspired by the recent success of clustering and contrasting learning paradigms for SSL-based speech representation learning, we propose SLICER (Symmetrical Learning of Instance and Cluster-level Efficient Representations) which brings together the best of both clustering and contrasting learning paradigms. We use a symmetric loss between latent representations from student and teacher encoders and simultaneously solve in-stance and cluster-level contrastive learning tasks. We obtain cluster representations online by just projecting the input spectrogram into an output subspace with dimensions equal to the number of clusters. In addition, we propose a novel mel-spectrogram augmentation procedure k-mix, based on mixup [1], which does not require labels and aids unsupervised representation learning for audio. Overall, SLICER achieves state-of-the-art results on the LAPE Benchmark [2], significantly outperforming all other prior approaches, some-times pre-trained on 10× larger unsupervised data than our setting. Code https: https://github.com/Sreyan88/audio-ssl. Ashish Seth, Sreyan Ghosh, Srinivasan Umesh, Dinesh Manocha |
ICASSP | 3 |
| 2023 | Technology Pipeline for Large Scale Cross-Lingual Dubbing of Lecture Videos into Multiple Indian Languages
Anusha Prakash 0001, Arun Kumar A, Ashish Seth, Bhagyashree Mukherjee, Ishika Gupta, Jom Kuriakose, Jordan Fernandes, K. V. Vikram, Mano Ranjith Kumar, Narla John Metilda Sagaya Mary, Mohammad Wajahat, Mohana N, Mudit Batra, Navina K, Nihal John George, Nithya Ravi, Pruthwik Mishra, Sudhanshu Srivastava 0001, Vasista Sai Lodagala, Vandan Mujadia, Kada Sai Venkata Vineeth, Vrunda N. Sukhadia, Dipti Misra Sharma, Hema A. Murthy, Pushpak Bhattacharyya, Srinivasan Umesh, Rajeev Sangal |
INTERSPEECH | 26 |
| 2023 | The Tag-Team Approach: Leveraging CLS and Language Tagging for Enhancing Multilingual ASR
Kaousheik Jayakumar, Vrunda N. Sukhadia, Arun Kumar A, Srinivasan Umesh |
INTERSPEECH | 4 |
| 2023 | SALTTS: Leveraging Self-Supervised Speech Representations for improved Text-to-Speech Synthesis
Ramanan Sivaguru, Vasista Sai Lodagala, Srinivasan Umesh |
INTERSPEECH | 3 |
| 2022 | Investigation of Robustness of Hubert Features from Different Layers to Domain, Accent and Language VariationsabstractIn this paper, we investigate the use of pre-trained HuBERT model to build downstream Automatic Speech Recognition (ASR) models using data that have differences in domain, accent and even language. We use the standard ESPnet recipe with HuBERT as pre-trained models whose output is fed as input features to a downstream Conformer model built from target domain data. We compare the performance of HuBERT pre-trained features with the baseline Conformer model built with Mel-filterbank features. We observe that as the domain, accent and bandwidth (as in the case of Switchboard data) vary, the relative improvements in performance over baseline decrease significantly. Further, with more labelled data in the target domain, the relative improvement narrows down, and both systems become comparable. We also investigate the effect on ASR performance when output from intermediate layers of HuBERT are used as features and show that these are more suitable for data in a different language, since they capture more of the acoustic representation. Finally, we compare the output from Convolutional Neural Network (CNN) Feature encoder used in pre-trained models with the Mel-filterbank features and show that Mel-filterbanks are often better features for modelling data from different domains. Pratik Kumar, Vrunda N. Sukhadia, Srinivasan Umesh |
ICASSP | 3 |
| 2022 | Investigation of Ensemble features of Self-Supervised Pretrained Models for Automatic Speech RecognitionabstractSelf-supervised learning (SSL) based models have been shown to generate powerful representations that can be used to improve the performance of downstream speech tasks. Several state-of-the-art SSL models are available, and each of these models optimizes a different loss which gives rise to the possibility of their features being complementary. This paper proposes using an ensemble of such SSL representations and models, which exploits the complementary nature of the features extracted by the various pretrained models. We hypothesize that this results in a richer feature representation and shows results for the ASR downstream task. To this end, we use three SSL models that have shown excellent results on ASR tasks, namely HuBERT, Wav2vec2.0, and WaveLM. We explore the ensemble of models fine-tuned for the ASR task and the ensemble of features using the embeddings obtained from the pre-trained models for a downstream ASR task. We get improved performance over individual models and pre-trained features using Librispeech(100h) and WSJ dataset for the downstream tasks. A. Arunkumar, Vrunda N. Sukhadia, Srinivasan Umesh |
INTERSPEECH | 3 |
| 2022 | Joint Encoder-Decoder Self-Supervised Pre-training for ASRabstractSelf-supervised learning (SSL) has shown tremendous success in various speech-related downstream tasks, including Automatic Speech Recognition (ASR).The output embeddings of the SSL model are treated as powerful short-time representations of the speech signal.However, in the ASR task, the main objective is to get the correct sequence of acoustic units, characters, or byte-pair encodings (BPEs).Usually, encoderdecoder architecture works exceptionally well for a sequenceto-sequence task like ASR.Therefore, in this paper, we propose a new paradigm that exploits the power of a decoder during self-supervised learning.We use Hidden Unit BERT (Hu-BERT) SSL framework to compute the conventional masked prediction loss for the encoder.In addition, we have introduced a decoder in the SSL framework and proposed a target preparation strategy for the decoder.Finally, we use a multitask SSL setup wherein we jointly optimize both the encoder and decoder losses.We hypothesize that the presence of a decoder in the SSL model helps it learn an acoustic unit-based language model, which might improve the performance of an ASR downstream task.We compare our proposed SSL model with HuBERT and show up to 25% relative improvement in performance on ASR by finetuning on various LibriSpeech subsets. A. Arunkumar, Srinivasan Umesh |
INTERSPEECH | 2 |
| 2022 | Gram Vaani ASR Challenge on spontaneous telephone speech recordings in regional variations of HindiabstractThis paper describes the corpus and baseline systems for the Gram Vaani Automatic Speech Recognition (ASR) challenge in regional variations of Hindi. The corpus for this challenge comprises the spontaneous telephone speech recordings collected by a social technology enterprise, Gram Vaani. The regional variations of Hindi together with spontaneity of speech, natural background and transcriptions with variable accuracy due to crowdsourcing make it a unique corpus for ASR on spontaneous telephonic speech. Around, 1108 hours of real-world spontaneous speech recordings, including 1000 hours of unlabelled training data, 100 hours of labelled training data, 5 hours of development data and 3 hours of evaluation data, have been released as a part of the challenge. The efficacy of both training and test sets are validated on different ASR systems in both traditional time-delay neural network-hidden Markov model (TDNN-HMM) frameworks and fully-neural end-to-end (E2E) setup. The word error rate (WER) and character error rate (CER) on eval set for a TDNN model trained on 100 hours of labelled data are 29.7 and 15.1, respectively. While, in E2E setup, WER and CER on eval set for a conformer model trained on 100 hours of data are 32.9 and 19.0, respectively. Anish Bhanushali, Grant Bridgman, Deekshitha G, Prasanta Kumar Ghosh, Pratik Kumar, Adithya Raj Kolladath, Nithya Ravi, Aaditeshwar Seth, Ashish Seth, Abhayjeet Singh, Vrunda N. Sukhadia, Srinivasan Umesh, Sathvik Udupa, Lodagala Durga Prasad |
INTERSPEECH | 13 |
| 2022 | Span Classification with Structured Information for Disfluency Detection in Spoken UtterancesabstractExisting approaches in disfluency detection focus on solving a token-level classification task for identifying and removing disfluencies in text.Moreover, most works focus on leveraging only contextual information captured by the linear sequences in text, thus ignoring the structured information in text which is efficiently captured by dependency trees.In this paper, building on the span classification paradigm of entity recognition, we propose a novel architecture for detecting disfluencies in transcripts from spoken utterances, incorporating both contextual information through transformers and long-distance structured information captured by dependency trees, through graph convolutional networks (GCNs).Experimental results show that our proposed model achieves state-of-the-art results on the widely used English Switchboard for disfluency detection and outperforms prior-art by a significant margin.We make all our codes publicly available on GitHub 1 . Sreyan Ghosh, Sonal Kumar, Yaman Singla, Rajiv Ratn Shah, Srinivasan Umesh |
INTERSPEECH | 5 |
| 2022 | DeToxy: A Large-Scale Multimodal Dataset for Toxicity Classification in Spoken UtterancesabstractToxic speech, also known as hate speech, is regarded as one of the crucial issues plaguing online social media today.Most recent work on toxic speech detection is constrained to the modality of text and written conversations with very limited work on toxicity detection from spoken utterances or using the modality of speech.In this paper, we introduce a new dataset DeToxy, the first publicly available toxicity annotated dataset for the English language.DeToxy is sourced from various openly available speech databases and consists of over 2 million utterances.We believe that our dataset would act as a benchmark for the relatively new and un-explored Spoken Language Processing task of detecting toxicity from spoken utterances and boost further research in this space.Finally, we also provide strong unimodal baselines for our dataset and compare traditional two-step and E2E approaches.Our experiments show that in the case of spoken utterances, text-based approaches are largely dependent on gold human-annotated transcripts for their performance and also suffer from the problem of keyword bias.However, the presence of speech files in DeToxy helps facilitates the development of E2E speech models which alleviate both the abovestated problems by better capturing speech clues. Sreyan Ghosh, Samden Lepcha, Sakshi Singh, Rajiv Ratn Shah, Srinivasan Umesh |
INTERSPEECH | 5 |
| 2022 | CCC-WAV2VEC 2.0: Clustering AIDED Cross Contrastive Self-Supervised Learning of Speech RepresentationsabstractWhile Self-Supervised Learning has helped reap the benefit of the scale from the available unlabeled data, the learning paradigms are continously being bettered. We present a new pre-training strategy named ccc-wav2vec 2.0, which uses clustering and an augmentation based cross-contrastive loss as its self-supervised objective. Through the clustering module we scale down the influence of those negative examples that are highly similar to the positive. The Cross-Contrastive loss is computed between the encoder output of the original sample and the quantizer output of its augmentation, and vice-versa, bringing robustness to the pre-training strategy. ccc-wav2vec 2.0 achieves upto 15.6% and 12.7% relative WER improvement over the baseline wav2vec 2.0 on the test-clean and test-other sets respectively of LibriSpeech, without the use of any language model. The proposed method also achieves upto 14.9% relative WER improvement over the baseline wav2vec 2.0, when fine-tuned on Switchboard data. Vasista Sai Lodagala, Sreyan Ghosh, Srinivasan Umesh |
SLT | 3 |
| 2022 | PADA: Pruning Assisted Domain Adaptation for Self-Supervised Speech RepresentationsabstractWhile self-supervised speech representation learning (SSL) models serve a variety of downstream tasks, these models have been observed to overfit to the domain from which the unlabeled data originates. To alleviate this issue, we propose PADA (Pruning Assisted Domain Adaptation). Before performing the target-domain ASR fine-tuning, we discover the redundant weights from pre-trained wav2vec 2.0 models through various pruning strategies. We investigate the effect of Task-Agnostic and Task-Aware pruning and propose a new pruning paradigm called, Cross-Domain Task-Aware Pruning (CD-TAW). CD-TAW obtains the initial pruning mask from a well fine-tuned out-of-domain (OOD) model, thereby making use of the readily available fine-tuned models from the web. The proposed CD-TAW method achieves up to 20.6% relative WER improvement over our baseline when fine-tuned on a 2-hour subset of Switchboard data without language model (LM) decoding. Vasista Sai Lodagala, Sreyan Ghosh, Srinivasan Umesh |
SLT | 3 |
| 2022 | Domain Adaptation of Low-Resource Target-Domain Models Using Well-Trained ASR Conformer ModelsabstractIn encoder-decoder framework for Automatic Speech Recognition (ASR) systems, the decoder of the well-trained ASR model is largely tuned towards the source-domain, hurting the performance of target-domain models in vanilla transfer-learning. On the other hand, the encoder layers of the well-trained ASR model mainly capture the acoustic characteristics. In this paper, the embeddings tapped from the encoder layers of a well-trained ASR model are used as features for domain adaptation of a downstream low resource Conformer target-domain model. We do ablation studies on optimal encoder layers for tapping embeddings and the effect of freezing or updating the well-trained ASR model's encoder layers. Lastly, the application of Spectral Augmentation (SpecAug) on the proposed features improves the target-domain performance further. The proposed method reports an average relative improvement of ~40% over baseline with different source-domain model and target-domain Conformer model combinations. Vrunda N. Sukhadia, Srinivasan Umesh |
SLT | 2 |
| 2022 | S-Vectors and TESA: Speaker Embeddings and a Speaker Authenticator Based on Transformer EncoderabstractOne of the most popular speaker embeddings is x-vectors, which are obtained from an architecture that gradually builds a larger temporal context with layers. In this paper, we propose to derive speaker embeddings from Transformer’s encoder trained for speaker classification. Self-attention, on which Transformer’s encoder is built, attends to all the features over the entire utterance and might be more suitable in capturing the speaker characteristics in an utterance. We refer to the speaker embeddings obtained from the proposed speaker classification model as s-vectors to emphasize that they are obtained from an architecture that heavily relies on self-attention. Through experiments, we demonstrate that s-vectors perform better than x-vectors. In addition to the s-vectors, we also propose a new architecture based on Transformer’s encoder for speaker verification as a replacement for speaker verification based on conventional probabilistic linear discriminant analysis (PLDA). This architecture is inspired by the next sentence prediction task of bidirectional encoder representations from Transformers (BERT), and we feed the s-vectors of two utterances to verify whether they belong to the same speaker. We name this architecture the Transformer encoder speaker authenticator (TESA). Our experiments show that the performance of s-vectors with TESA is better than s-vectors with conventional PLDA-based speaker verification. Narla John Metilda Sagaya Mary, Srinivasan Umesh, Sandesh Varadaraju Katta |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Exploring the use of Common Label Set to Improve Speech Recognition of Low Resource Indian LanguagesabstractIn many Indian languages, written characters are organized on sound phonetic principles, and the ordering of characters is the same across many of them. However, while training conventional end-to-end (E2E) Multilingual speech recognition systems, we treat characters or target subword units from different languages as separate entities. Since the visual rendering of these characters is different, in this paper, we explore the benefits of representing such similar target subword units (e.g., Byte Pair Encoded(BPE) units) through a Common Label Set (CLS). The CLS can be very easily created using automatic methods since the ordering of characters is the same in many Indian Languages. E2E models are trained using a transformer-based encoder-decoder architecture. During testing, given the Mel-filterbank features as input, the system outputs a sequence of BPE units in CLS representation. Depending on the language, we then map the recognized CLS units back to the language-specific grapheme representation. Results show that models trained using CLS improve over monolingual baseline and a multilingual framework with separate symbols for each language. Similar experiments on a subset of the Voxforge dataset also confirm the benefits of CLS. An extension of this idea is to decode an unseen language (Zero-resource) using CLS trained model. Vishwas M. Shetty, Srinivasan Umesh |
ICASSP | 2 |
| 2020 | Investigation of Methods to Improve the Recognition Performance of Tamil-English Code-Switched Data in Transformer FrameworkabstractCode-switching (CS) refers to (inter/intra-word) switching between multiple languages in a single conversation. In multilingual countries like India, CS occurs very often in everyday speech, resulting in a new breed of languages in urban regions like Hinglish (Hindi-English), Tanglish (Tamil-English), etc. Research in Indic CS speech recognition is primarily affected by insufficient data. In this paper, we investigate methods to deal with such very low resource scenarios. Recently, Transformers have shown promising results on automatic speech recognition (ASR) tasks. In a Transformer based framework, we investigate two methods for Tamil-English CS speech recognition, namely, (i) well-trained encoders of Monolingual Transformers as feature extractors to provide language discrimination, (ii) language information as tokens at the targets. Our results show that CS is efficiently handled by the second method, while the first method was efficient in discriminating languages. Narla John Metilda Sagaya Mary, Vishwas M. Shetty, Srinivasan Umesh |
ICASSP | 3 |
| 2020 | Improving the Performance of Transformer Based Low Resource Speech Recognition for Indian LanguagesabstractThe recent success of the Transformer based sequence-to-sequence framework for various Natural Language Processing tasks has motivated its application to Automatic Speech Recognition. In this work, we explore the application of Transformers on low resource Indian languages in a multilingual framework. We explore various methods to incorporate language information into a multilingual Transformer, i.e., (i) at the decoder, (ii) at the encoder. These methods include using language identity tokens or providing language information to the acoustic vectors. Language information to the acoustic vectors can be given in the form of one hot vector or by learning a language embedding. From our experiments, we observed that providing language identity always improved performance. The language embedding learned from our proposed approach, when added to the acoustic feature vector, gave the best result. The proposed approach with retraining gave 6% - 11% relative improvements in character error rates over the monolingual baseline. Vishwas M. Shetty, Narla John Metilda Sagaya Mary, Srinivasan Umesh |
ICASSP | 3 |
| 2018 | Correlational Networks for Speaker Normalization in Automatic Speech Recognition
Rini A. Sharon, Sandeep Kothinti, Srinivasan Umesh |
INTERSPEECH | 3 |
| 2018 | Articulatory and Stacked Bottleneck Features for Low Resource Speech Recognition
Vishwas M. Shetty, Rini A. Sharon, Basil Abraham, Tejaswi Seeram, Anusha Prakash 0001, Nithya Ravi, Srinivasan Umesh |
INTERSPEECH | 7 |
| 2018 | Investigating the Effect of Audio Duration on Dementia Detection Using Acoustic FeaturesabstractThis paper presents recent progress toward our goal to enable area-wide pre-screening methods for the early detection of dementia based on automatically processing conversational speech of a representative group of more than 200 subjects. We focus on conversational speech since it is the natural form of communication that can be recorded unobtrusively, without adding stress to subjects, and without the need of controlled clinical settings. We describe our unsupervised process chain consisting of voice activity detection and speaker diarization followed by extraction of features and detection of early signs of dementia. The unsupervised system achieves up to 0.645 unweighted average recall (UAR) and compares favorably to a system that was carefully designed on manually annotated data. To further lower the burden for subjects, we investigate UAR over speech duration, and find that about 12 minutes of interview are sufficient to achieve the best UAR. Jochen Weiner, Miguel Angrick, Srinivasan Umesh, Tanja Schultz |
INTERSPEECH | 3 |
| 2018 | FMLLR Speaker Normalization With i-Vector: In Pseudo-FMLLR and Distillation FrameworkabstractWhen an automatic speech recognition (ASR) system is deployed for real-world applications, it often receives only one utterance at a time for decoding. This single utterance could be of short duration depending on the ASR task. In these cases, robust estimation of speaker normalizing methods like feature-space maximum likelihood linear regression (FMLLR) and i-vectors may not be feasible. In this paper, we propose two unsupervised speaker normalization techniques-one at feature level and other at model level of acoustic modeling-to overcome the drawbacks of FMLLR and i-vectors in real-time scenarios. At feature level, we propose the use of deep neural networks (DNN) to generate pseudo-FMLLR features from time-synchronous pair of filterbank and FMLLR features. These pseudo-FMLLR features can then be used for DNN acoustic model training and decoding. At model level, we propose a generalized distillation framework, where a teacher DNN trained on FMLLR features guides the training and optimization of a student DNN trained on filterbank features. In both the proposed methods, the ambiguity in choosing the speaker-specific FMLLR transform can be reduced by augmenting i-vectors to the input filterbank features. Experiments conducted on 33-h and 110-h subsets of Switchboard corpus show that the proposed methods provide significant gains over DNNs trained on FMLLR, i-vector appended FMLLR, filterbank and i -vector appended filterbank features, in real-time scenario. Neethu Mariam Joy, Sandeep Kothinti, Srinivasan Umesh |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Transfer Learning and Distillation Techniques to Improve the Acoustic Modeling of Low Resource Languages
Basil Abraham, Tejaswi Seeram, Srinivasan Umesh |
INTERSPEECH | 3 |
| 2017 | Joint Estimation of Articulatory Features and Acoustic Models for Low-Resource Languages
Basil Abraham, Srinivasan Umesh, Neethu Mariam Joy |
INTERSPEECH | 2 |
| 2017 | Generalized Distillation Framework for Speaker Normalization
Neethu Mariam Joy, Sandeep Kothinti, Srinivasan Umesh, Basil Abraham |
INTERSPEECH | 3 |
| 2017 | On Improving Acoustic Models for TORGO Dysarthric Speech Database
Neethu Mariam Joy, Srinivasan Umesh, Basil Abraham |
INTERSPEECH | 2 |
| 2017 | An automated technique to generate phone-to-articulatory label mapping
Basil Abraham, Srinivasan Umesh |
Speech Commun. | 2 |
| 2017 | DNNs for unsupervised extraction of pseudo speaker-normalized features without explicit adaptation data
Neethu Mariam Joy, Murali Karthick Baskar, Srinivasan Umesh |
Speech Commun. | 3 |
| 2016 | Articulatory Feature Extraction Using CTC to Build Articulatory Classifiers Without Forced Frame Alignments for Speech Recognition
Basil Abraham, Srinivasan Umesh, Neethu Mariam Joy |
INTERSPEECH | 2 |
| 2016 | Overcoming Data Sparsity in Acoustic Modeling of Low-Resource Language by Borrowing Data and Model Parameters from High-Resource Languages
Basil Abraham, Srinivasan Umesh, Neethu Mariam Joy |
INTERSPEECH | 2 |
| 2016 | DNNs for Unsupervised Extraction of Pseudo FMLLR Features Without Explicit Adaptation Data
Neethu Mariam Joy, Murali Karthick Baskar, Srinivasan Umesh, Basil Abraham |
INTERSPEECH | 3 |
| 2015 | Speaker adaptation of convolutional neural network using speaker specific subspace vectors of SGMM
Murali Karthick Baskar, Prateek Kolhar, Srinivasan Umesh |
INTERSPEECH | 3 |
| 2015 | Sub-band based histogram equalization in cepstral domain for speech recognition
Vikas Joshi, Raghavendra Bilgi, Srinivasan Umesh, Luz García 0001, M. Carmen Benítez |
Speech Commun. | 3 |
| 2014 | Improving deep neural networks using state projection vectors of subspace Gaussian mixture model as featuresabstractRecent advancement in deep neural network (DNN) has surpassed the conventional hidden Markov model-Gaussian mixture model (HMM-GMM) framework due to its efficient training procedure. Providing better phonetic context information in the input gives improved performance for DNN. The state projection vectors (state specific vectors) in subspace Gaussian mixture model (SGMM) captures the phonetic information in low dimensional vector space. In this paper, we propose to use state specific vectors of SGMM as features thereby providing additional phonetic information for the DNN framework. To each observation vector in the train data, the corresponding state specific vectors of SGMM are aligned to form the state specific vector feature set. Linear discriminant analysis (LDA) feature set are formed by applying LDA to the training data. Since bottleneck features are efficient in extracting useful discriminative information for the phonemes, LDA feature set and state specific vector feature set are converted to bottleneck features. These bottleneck features of both feature sets act as input features to train a single DNN framework. Relative improvement of 8.8% for TIMIT database (core test set) and 9.7% for WSJ corpus is obtained by using the state specific vector bottleneck feature set when compared to the DNN trained only with LDA bottleneck feature set. Also training Deep belief network - DNN (DBN-DNN) using the proposed feature set attains a WER of 20.46% on TIMIT core test set proving the effectiveness of our method. The state specific vectors while acting as features, provide additional useful information related to phoneme variation. Thus by combining it with LDA bottleneck features improved performance is obtained using the DNN framework. Murali Karthick Baskar, Srinivasan Umesh |
SLT | 2 |
| 2014 | Acoustic modelling for speech recognition in Indian languages in an agricultural commodities task domain
Aanchan Mohan, Richard C. Rose, Sina Hamidi Ghalehjegh, Srinivasan Umesh |
Speech Commun. | 4 |
| 2013 | Modified splice and its extension to non-stereo data for noise robust speech recognitionabstractIn this paper, a modification to the training process of the popular SPLICE algorithm has been proposed for noise robust speech recognition. The modification is based on feature correlations, and enables this stereo-based algorithm to improve the performance in all noise conditions, especially in unseen cases. Further, the modified framework is extended to work for non-stereo datasets where clean and noisy training utterances, but not stereo counterparts, are required. Finally, an MLLR-based computationally efficient run-time noise adaptation method in SPLICE framework has been proposed. The modified SPLICE shows 8.6% absolute improvement over SPLICE in Test C of Aurora-2 database, and 2.93% overall. Non-stereo method shows 10.37% and 6.93% absolute improvements over Aurora-2 and Aurora-4 baseline models respectively. Run-time adaptation shows 9.89% absolute improvement in modified framework as compared to SPLICE for Test C, and 4.96% overall w.r.t. standard MLLR adaptation on HMMs. D. S. Pavan Kumar, N. Vishnu Prasad, Vikas Joshi, Srinivasan Umesh |
ASRU | 4 |
| 2013 | Acoustic modeling using transform-based phone-cluster adaptive trainingabstractIn this paper, we propose a new acoustic modeling technique called the Phone-Cluster Adaptive Training. In this approach, the parameters of context-dependent states are obtained by the linear interpolation of several monophone cluster models, which are themselves obtained by adaptation using linear transformation of a canonical Gaussian Mixture Model (GMM). This approach is inspired from the Cluster Adaptive Training (CAT) for speaker adaptation and the Subspace Gaussian Mixture Model (SGMM). The parameters of the model are updated in an adaptive training framework. The interpolation vectors implicitly capture the phonetic context information. The proposed approach shows substantial improvement over the Continuous Density Hidden Markov Model (CDHMM) and a similar performance to that of the SGMM, while using significantly fewer parameters than both the CDHMM and the SGMM. Vimal Manohar, Srinivas C. Bhargav, Srinivasan Umesh |
ASRU | 3 |
| 2013 | Improved cepstral mean and variance normalization using Bayesian frameworkabstractCepstral Mean and Variance Normalization (CMVN) is a computationally efficient normalization technique for noise robust speech recognition. The performance of CMVN is known to degrade for short utterances, due to insufficient data for parameter estimation and loss of discriminable information as all utterances are forced to have zero mean and unit variance. In this work, we propose to use posterior estimates of mean and variance in CMVN, instead of the maximum likelihood estimates. This Bayesian approach, in addition to providing a robust estimate of parameters, is also shown to preserve discriminable information without increase in computational cost, making it particularly relevant for Interactive Voice Response (IVR)-based applications. The relative WER reduction of this approach w.r.t. Cepstral Mean Normalization, CMVN and Histogram Equalization are (i) 40.1%, 27% and 4.3% with the Aurora2 database for all utterances, (ii) 25.7%, 38.6% and 30.4% with the Aurora2 database for short utterances, and (iii) 18.7%, 12.6% and 2.5% with the Aurora4 database. N. Vishnu Prasad, Srinivasan Umesh |
ASRU | 2 |
| 2013 | Modified cepstral mean normalization - transforming to utterance specific non-zero mean
Vikas Joshi, N. Vishnu Prasad, Srinivasan Umesh |
INTERSPEECH | 3 |
| 2012 | Robust speech recognition through selection of speaker and environment transformsabstractIn this paper, we address the problem of robustness to both noise and speaker-variability in automatic speech recognition (ASR). We propose the use of pre-computed Noise and Speaker transforms, and an optimal combination of these two transforms are chosen during test using maximum-likelihood (ML) criterion. These pre-computed transforms are obtained during training by using data obtained from different noise conditions that are usually encountered for that particular ASR task. The environment transforms are obtained during training using constrained-MLLR (CMLLR) framework, while for speaker-transforms we use the analytically determined linear-VTLN matrices. Even though the exact noise environment may not be encountered during test, the ML-based choice of the closest Environment transform provides “sufficient” cleaning and this is corroborated by experimental results with performance comparable to histogram equalization or Vector Taylor Series approaches on Aurora-2 task. The proposed method is simple since it involves only the choice of pre-computed environment and speaker transforms and therefore, can be applied with very little test data unlike many other speaker and noise-compensation methods. Raghavendra Bilgi, Vikas Joshi, Srinivasan Umesh, Luz García 0001, M. Carmen Benítez |
ICASSP | 3 |
| 2012 | Noise and speaker compensation in the Log filter bank domainabstractIn this paper, we propose a method to compensate for noise and speaker-variability directly in the Log filter-bank (FB) domain, so that MFCC features are robust to noise and speaker-variations. For noise-compensation, we use Vector Taylor Series (VTS) approach in the Log FB domain, and speaker-normalization is also done in the Log FB domain using Linear Vocal tract length (VTLN) matrices. For VTLN, optimal selection of warp-factor is done in Log FB domain using canonical GMM model, avoiding the two-pass approach needed by a HMM model. Further, this can be efficiently implemented using sufficient statistics obtained from the GMM and the FB-VTLN-matrices. The warp-factor selection using GMM can also be done in cepstral domain by applying DCT matrices without the usual approximations associated with conventional linear-VTLN. The elegance of the proposed approach is that given the speech data, we obtain directly MFCC features that are robust to noise and speaker-variations. The proposed approach, show a significant relative improvement of 31% over baseline on Aurora-4 task. Vikas Joshi, Raghavendra Bilgi, Srinivasan Umesh, Luz García 0001, M. Carmen Benítez |
ICASSP | 3 |
| 2012 | Computationally efficient speaker identification using fast-MLLR based anchor modelingabstractIn this paper, we propose a computationally efficient method to identify a speaker from a large population of speakers. The proposed method is based on our earlier [1] Fast Maximum Likelihood Linear Linear Regression (MLLR) anchor modeling technique which provides performance comparable to the conventional anchor modeling system and yet reduces computation time significantly by computing likelihood efficiently using sufficient statistics of data and anchor specific MLLR matrix. However, both these systems still require a Gaussian Mixture Model-Universal Background Model (GMM-UBM) based back-end system to choose the optimal speaker, which is computationally heavy. In our proposed method, we show that applying Linear-Discriminant Analysis (LDA) and Within-Class-Covariance Normalization (WCCN) on the Speaker characterization Vector (SCV) of our recently proposed Fast-MLLR method, we can combine the computational efficiency and the discriminant capability to have a system that uses simple cosine-distance measure to identify speakers and yet has significantly superior performance compared to both full-blown GMM-UBM system and the anchor-model system. More importantly, there is no need of the “back-end” system. Experimental result on NIST 2004 SRE shows that the proposed method reduces identification error rate by an absolute 2% and takes only 2/3 of the time taken by efficient Fast-MLLR system and only 20% of the time taken by the stand-alone GMM-UBM system. Achintya Kumar Sarkar, Srinivasan Umesh, Jean-François Bonastre |
ICASSP | 2 |
| 2012 | VTLN Using Analytically Determined Linear-Transformation on Conventional MFCCabstractIn this paper, we propose a method to analytically obtain a linear-transformation on the conventional Mel frequency cepstral coefficients (MFCC) features that corresponds to conventional vocal tract length normalization (VTLN)-warped MFCC features, thereby simplifying the VTLN processing. There have been many attempts to obtain such a linear-transformation, but all the previously proposed approaches either modify the signal processing (and therefore not conventional MFCC), or the linear-transformation does not correspond to conventional VTLN-warping, or the matrices being estimated and are data dependent. In short, the conventional VTLN part of an automatic speech recognition (ASR) system cannot be simply replaced with any of the previously proposed methods. Umesh proposed the idea to use band-limited interpolation for performing VTLN-warping on MFCC using plain cepstra. Motivated from this work, Panchapagesan and Alwan proposed a linear-transformation to perform VTLN-warping on conventional MFCC. However, in their approach, VTLN warping is specified in the Mel-frequency domain and is not equivalent to conventional VTLN. In this paper, we present an approach which also draws inspiration from the work of Umesh , and which we believe for the first time performs conventional VTLN as a linear-transformation on conventional MFCC using the ideas of band-limited interpolation. Deriving such a linear-transformation to perform VTLN, would allow us to use the VTLN-matrices in transform-based adaptation framework with its associated advantages and yet would require the estimation of a single parameter. Using four different tasks, we show that our proposed approach has almost identical recognition performance to conventional VTLN on both clean and noisy speech data. Rama Sanand Doddipatla, Srinivasan Umesh |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Use of VTL-wise models in feature-mapping framework to achieve performance of multiple-background models in speaker verificationabstractRecently, Multiple Background Models (M-BMs) [1, 2] have been shown to be useful in speaker verification, where the M-BMs are formed based on different Vocal Tract Lengths (VTLs) among the population. The speaker models are adapted from the particular Background Model (BM) corresponding to their VTL. During test, log likelihood ratio of the test utterance is calculated between claimant model and the corresponding BM. In this paper, instead of using different BM for different speaker, we propose the use of single gender, channel and VTL independent UBM (root-UBM) using the concept of VTL dependent mapping function. The pro posed concept is inspired by Feature Mapping (FM) technique used in speaker verification to overcome channel variability. In our pro posed method, VTL specific gender independent Gaussian Mixture models (GMMs) are derived from the root-UBM using Maximum a posteriori (MAP) adaptation. The mapping relation is then learned between the root-UBM and the VTL-specific GMM. During training and testing phase, feature vectors are mapped into root-UBM using the best VTL specific model. Then speaker models are adapted from the root-UBM using mapped features. During test, the log likelihood ratio is calculated between target model and root-UBM. Therefore, unlike M-BM system, there is no need to switch to different BMs depending on the claimant. Another advantage of the proposed method is that other additional normalization/compensation techniques can be easily applied since it is in a single UBM frame-work. The experiments are performed on NIST 2004 SRE core condition, and we show that the performance of the proposed method is close to the M-BM system with and without score normalization. Achintya Kumar Sarkar, Srinivasan Umesh |
ICASSP | 2 |
| 2011 | Efficient Speaker and Noise Normalization for Robust Speech Recognition
Vikas Joshi, Raghavendra Bilgi, Srinivasan Umesh, M. Carmen Benítez, Luz García 0001 |
INTERSPEECH | 3 |
| 2011 | Sub-Band Level Histogram Equalization for Robust Speech RecognitionabstractThis paper describes a novel modification of Histogram Equalization approach to robust speech recognition. We propose separate equalization of the high frequency and low frequency bands. We study different combinations of the sub-band equalization and obtain best results when we performs a twostage equalization. First, conventional Histogram Equalization (HEQ) is performed on the cepstral features, which does not completely equalize high frequency and low frequency bands, even though the overall histogram equalization is good. In the second stage, an equalization is done separately on the high frequency and the low frequency components of the above equalized cepstra. We refer to this approach as Sub-band Histogram Equalization (S-HEQ). The new set of features has better equalization of the sub-bands as well as the overall cepstral histogram. Recognition results show a relative improvement of 12% and 15% over conventional HEQ on Aurora-2 and Aurora4 databases respectively. Vikas Joshi, Raghavendra Bilgi, Srinivasan Umesh, Luz García 0001, M. Carmen Benítez |
INTERSPEECH | 3 |
| 2011 | Eigen-Voice Based Anchor Modeling System for Speaker Identification Using MLLR Super-VectorabstractIn this paper, we propose an anchor modeling scheme where instead of conventional “anchor” speakers, we use eigenvectors that span the Eigen-voice space. The computational advantage of conventional Anchor-modeling based speaker identification system comes from representing all speakers in a space spanned by a small number of anchor speakers instead of having separate speaker models. The conventional “anchor” speakers are usually chosen using data-driven clustering and the number of such speakers are also empirically determined. The use of proposed eigenvoice based anchors provide a more systematic way of spanning the speaker-space and in determining the optimal number of anchors. In our proposed method, the eigenvector space is built using the Maximum Likelihood Linear Regression (MLLR) super-vectors of non-target speakers. Further, the proposed method does not require calculation of the likelihood with respect to anchor speaker models to create the speakercharacterization vector as done in conventional anchor systems. Instead, speakers are characterized with respect to eigen-space by projecting the speaker’s MLLR-super vector onto the eigenvoice space. This makes the method computationally efficient. Experimental results show that the proposed method consistently performs better than conventional anchor modeling technique for different number of anchor speakers. Achintya Kumar Sarkar, Srinivasan Umesh |
INTERSPEECH | 2 |
| 2010 | Fast computation of speaker characterization vector using MLLR and sufficient statistics in anchor model framework
Achintya Kumar Sarkar, Srinivasan Umesh |
INTERSPEECH | 2 |
| 2009 | Improving the performance of VTLN under mismatched speaker conditions and making it approach that of matched speaker conditionsabstractThe performance of conventional VTLN for mis-matched train and test speaker conditions (e.g. adult-train child-test) does not approach the performance of matched speaker conditions (e.g. child-train child-test). In this paper, we investigate this problem and propose methods to reduce this gap in performance. We use our recently proposed linear transformation approach to VTLN, that also enables us to study the effect of Jacobian unlike conventional VTLN. The main advantage of transform-based VTLN over adaptation based approaches (like CMLLR), is that it does not require any matrix estimation. We argue that the degraded VTLN performance under mismatched speaker conditions is due to the significant frequency warping that is necessary for normalization which leads to a mis-match between the correlation in the feature components of the test data and the covariance structure of the trained/normalized model. We show that the use of a global de-correlating transform (MLLT) leads to improved VTLN performance. We finally show that using both Jacobian and MLLT together improves the VTLN performance for mis-matched cases with the performance approaching that of matched speaker conditions. Rama Sanand Doddipatla, Shakti Prasad Rath, Srinivasan Umesh |
ICASSP | 3 |
| 2009 | Characterizing speaker variability using spectral envelopes of vowel sounds
A. N. Harish, Rama Sanand Doddipatla, Srinivasan Umesh |
INTERSPEECH | 3 |
| 2009 | Acoustic class specific VTLN-warping using regression class treesabstractIn this paper, we study the use of different frequency warp-factors for different acoustic classes in a computationally efficient frame-work of Vocal Tract Length Normalization (VTLN). This is motivated by the fact that all acoustic classes do not exhibit similar spectral variations as a result of physi-ological differences in vocal tract, and therefore, the use of a single frequency-warp for the entire utterance may not be ap-propriate. We have recently proposed a VTLN method that im-plements VTLN-warping through a linear-transformation (LT) of the conventional MFCC features and efficiently estimates the warp-factor using the same sufficient statistics as that are used in CMLLR adaptation. In this paper we have shown that, in this framework of VTLN, and using the idea of regression class tree, we can obtain separate VTLN-warping for different acoustic classes. The use of regression class tree ensures that warp-factor is estimated for each class even when there is very little data available for that class. The acoustic classes, in general, can be any collection of the Gaussian components in the acoustic model. We have built acoustic classes by using data-driven ap-proach and by using phonetic knowledge. Using WSJ database we have shown the recognition performance of the proposed acoustic class specific warp-factor both for the data driven and the phonetic knowledge based regression class tree definitions and compare it with the case of the single warp-factor. Shakti Prasad Rath, Srinivasan Umesh |
INTERSPEECH | 2 |
| 2009 | Using VTLN matrices for rapid and computationally-efficient speaker adaptation with robustness to first-pass transcription errorsabstractIn this paper, we propose to combine the rapid adaptation capability of conventional Vocal Tract Length Normalization (VTLN) with the computational efficiency of transform-based adaptation such as MLLR or CMLLR. VTLN requires the estimation of only one parameter and is, therefore, most suited for the cases where there is little adaptation data (i.e. rapid adaptation). In contrast, transform-based adaptation methods require the estimation of matrices. However, the drawback of conventional VTLN is that it is computationally expensive since it requires multiple spectral-warping to generate VTLN-warped features. We have recently shown that VTLN-warping can be implemented by a linear-transformation (LT) of the conventional MFCC features. These LTs are analytically pre-computed and stored. In this frame-work of LT VTLN, computational complexity of VTLN is similar to transform-based adaptation since warp-factor estimation can be done using the same sufficient statistics as that are used in CMLLR. We show that VTLN provides significant improvement in performance when there is small adaptation data as compared to transform-based adaptation methods. We also show that the use of an additional decorrelating transform, MLLT, along with the VTLN-matrices, gives performance that is better than MLLR and comparable to SAT with MLLT even for large adaptation data. Further we show that in the mismatched train and test case (i.e. poor first-pass transcription), VTLN provides significant improvement over the transform-based adaptation methods. We compare the performances of different methods on the WSJ, the RM and the TIDIGITS databases. Index Terms: VTLN, Rapid Adaptation, MLLT, CAT, Linear Transform Shakti Prasad Rath, Srinivasan Umesh, Achintya Kumar Sarkar |
INTERSPEECH | 2 |
| 2009 | A study on the influence of covariance adaptation on jacobian compensation in vocal tract length normalization
Rama Sanand Doddipatla, Shakti Prasad Rath, Srinivasan Umesh |
INTERSPEECH | 3 |
| 2009 | Text-independent speaker identification using vocal tract length normalization for building universal background modelabstractIn this paper, we propose to use Vocal Tract Length Normalization (VTLN) to build the Universal Background Model (UBM) for a closed set speaker identification system. Vocal Tract Length (VTL) differences among speakers is a major source of variability in the speech signal. Since the UBM model is trained using data from many speakers, it statistically captures this inherent variation in the speech signal, which results in a “coarse” model in the acoustic space. This may cause the adapted speaker models obtained from the UBM model to have significantly high overlap in the acoustic space. We hypothesize that the use of VTLN will help in compacting the UBM model and thus the speaker adapted models obtained from this compact model will have better speaker-separability in the acoustic space. We perform experiments on MIT, TIMIT and NIST 2004 SRE databases and show that using VTLN we can achieve lesser Identification Error Rates as compared to the conventional GMM-UBM based method. Achintya Kumar Sarkar, Srinivasan Umesh, Shakti Prasad Rath |
INTERSPEECH | 2 |
| 2008 | A computationally efficient approach to warp factor estimation in VTLN using EM algorithm and sufficient statistics
P. T. Akhil, Shakti Prasad Rath, Srinivasan Umesh, Rama Sanand Doddipatla |
INTERSPEECH | 3 |
| 2008 | Use of spectral centre of gravity for generating speaker invariant features for automatic speech recognition
Rama Sanand Doddipatla, Rani R. Sandhya, Srinivasan Umesh |
INTERSPEECH | 4 |
| 2008 | Study of jacobian compensation using linear transformation of conventional MFCC for VTLN
Rama Sanand Doddipatla, Srinivasan Umesh |
INTERSPEECH | 2 |
| 2008 | A shift-based approach to speaker normalization using non-linear frequency-scaling model
Rohit Sinha 0003, Srinivasan Umesh |
Speech Commun. | 2 |
| 2007 | Speaker-Invariant Features for Automatic Speech Recognition
Srinivasan Umesh, Rama Sanand Doddipatla, G. Praveen |
IJCAI | 1 |
| 2007 | Linear transformation approach to VTLN using dynamic frequency warping
Rama Sanand Doddipatla, D. Dinesh Kumar, Srinivasan Umesh |
INTERSPEECH | 3 |
| 2007 | A Study of Filter Bank Smoothing in MFCC Features for Recognition of Children's SpeechabstractIn this paper, we study the effect of filter bank smoothing on the recognition performance of children's speech. Filter bank smoothing of spectra is done during the computation of the Mel filter bank cepstral coefficients (MFCCs). We study the effect of smoothing both for the case when there is vocal-tract length normalization (VTLN) as well as for the case when there is no VTLN. The results from our experiments indicate that unlike conventional VTLN implementation, it is better not to scale the bandwidths of the filters during VTLN - only the filter center frequencies need be scaled. Our interpretation of the above result is that while the formant center frequencies may approximately scale between speakers, the formant bandwidths do not change significantly. Therefore, the scaling of filter bandwidths by a warp-factor during conventional VTLN results in differences in spectral smoothing leading to degradation in recognition performance. Similarly, results from our experiments indicate that for telephone-based speech when there is no normalization it is better to use uniform-bandwidth filters instead of the constant- like filters that are used in the computation of conventional MFCC. Our interpretation is that with constant- filters there is excessive spectral smoothing at higher frequencies which leads to degradation in performance for children's speech. However, the use of constant- filters during VTLN does not create any additional performance degradation. As we will show, during VTLN it is only important that the filter bandwidths are not scaled irrespective of whether we use constant- or uniform-bandwidth filters. With our proposed changes in the filter bank implementation we get comparable performance for adults and about 6% improvement for children both for the case of using VTLN as well as the for the case of not using VTLN on a telephone-based digit recognition task. Srinivasan Umesh, Rohit Sinha 0003 |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Study Of Non-Linear Frequency Warping Functions For Speaker NormalizationabstractIn this paper, we study non-linear frequency-warping functions that are commonly used in speaker normalization. This study is motivated by our recently proposed affine transformation model for speaker normalization [1] which has provided improved recognition performance when compared to uniform scaling model [1, 2]. In this work, using formant data from Peterson & Barney and Hillenbrand vowel databases, we analyze the behavior of scale factor as a function of frequency. The empirical observation [3, 4] shows that while uniform scaling assumption may be valid at higher frequencies, there are significant deviations at low frequencies. We show that while our recently proposed model has behavior similar to the empirical result, the behavior of many of the commonly used non-linear models (including that of Eide-Gish, power law and bilinear transformation) differ significantly from the empirical result. This difference in behavior from the empirical observation may explain the limited improvement in recognition performance provided by these non-linear models when compared to conventional uniform-scaling model. We also show that our proposed model does better fitting to the formant data than these non-linear models. We, therefore, conclude that the affine-transformation model may be a more appropriate non-linear model for speaker normalization. S. V. Bharath Kumar, Srinivasan Umesh, Rohit Sinha 0003 |
ICASSP (1) | 2 |
| 2006 | Vtln Warping Factor Estimation Using Accumulation of Sufficient StatisticsabstractIn this paper we present an efficient and flexible approach to VTLN warping factor estimation. Due to the equivalence of frequency warping and linear transformation of cepstral coefficients, warping factors can be efficiently estimated by accumulating the sufficient statistics for linear transformation estimation, and searching the constrained space of transformations given by the explicit mapping between warping factors and linear transformation matrices. We show that the positive effect of using a properly normalized optimization criterion for warping factor estimation, which has been previously demonstrated for a signal analysis front-end without a filterbank, carries over to a MFCC front-end, resulting in a net improvement in word error rate Jonas Lööf, Hermann Ney, Srinivasan Umesh |
ICASSP (1) | 3 |
| 2005 | Implementing frequency-warping and VTLN through linear transformation of conventional MFCCabstractIn this paper, we show that frequency-warping (including VTLN) can be implemented through linear transformation of conventional MFCC.Unlike the Pitz-Ney [1] continuous domain approach, we directly determine the relation between frequency-warping and the linear-transformation in the discrete-domain.The advantage of such an approach is that it can be applied to any frequency-warping and is not limited to cases where an analytical closed-form solution can be found.The proposed method exploits the bandlimited interpolation idea (in the frequency-domain) to do the necessary frequency-warping and yields exact results as long as the cepstral coefficients are que-frency limited.This idea of quefrencylimitedness shows the importance of the filter-bank smoothing of the spectra which has been ignored in [1,2].Furthermore, unlike [1], since we operate in the discrete domain, we can also apply the usual discrete-cosine transform (i.e.DCT-II) on the logarithm of the filter-bank output to get conventional MFCC features.Therefore, using our proposed method, we can linearly transform conventional MFCC cepstra to do VTLN and we do not require any recomputation of the warped-features.We provide experimental results in support of this approach. Srinivasan Umesh, András Zolnay, Hermann Ney |
INTERSPEECH | 1 |
| 2004 | Non-uniform speaker normalization using affine-transformationabstractWe propose a mathematical model to describe the relation between the formant frequencies of speakers and show that with the proposed affine model, speaker differences separate out as translation factors when a "Mel-like" warping is performed. Using speech data, we estimate the parameters of this warping function and show that it is close to the usual Mel-formula. This model is motivated by Rohit Sinha and S. Umesh's shift-based non-uniform speaker-normalization method (see Proc. IEEE ICASSP, 2002), which provides improvement over conventional maximum-likelihood based speaker normalization methods. We therefore provide a unified framework that relates the relationship between formants of speakers and the method of removing speaker differences (which involves Mel-warping) in a neat mathematical framework which is substantiated by our recognition experiments. S. V. Bharath Kumar, Srinivasan Umesh, Rohit Sinha 0003 |
ICASSP (1) | 2 |
| 2004 | An investigation into front-end signal processing for speaker normalizationabstractOur investigation into the front-end signal processing for maximum likelihood based speaker normalization reveals that, in the linear scaling model, it is more appropriate (and evidently more correct) to assume that the spectral envelopes of any two speakers for the same sound are linearly scaled versions of one another, rather than assuming that the magnitudes of the whole spectra (including pitch harmonics) are scaled. The use of the proposed model and its implementation results in about 4% and 7% relative improvement for adults and children respectively on a digit recognition task. Srinivasan Umesh, Rohit Sinha 0003, S. V. Bharath Kumar |
ICASSP (1) | 1 |
| 2004 | Using VTLN for broadcast news transcriptionabstractVocal tract length normalisation (VTLN) is a commonly used speaker normalisation approach. It is attractive compared to many normalisation schemes as it is typically dependent on only a single parameter, allowing the warp factors to be robustly calculated on little data. However, the scheme normally requires explicitly coding the data at multiple warp factors. Furthermore, it is only possible to approximate the Jacobian associated with the VTLN transformation. A new, simple, linear approximation to VTLN is described in this paper. This linear approximation allows the Jacobian to be exactly computed. It can also be highly efficient in terms of warp factor estimation and application of the warp factors. Both the linear and standard CUED VTLN schemes were evaluated in the 2003 BNE evaluation framework and found to yield similar performance. When used in system combination both VTLN schemes yielded slight gains over the baseline system. Do Yeong Kim, Srinivasan Umesh, Mark J. F. Gales, Thomas Hain, Philip C. Woodland |
INTERSPEECH | 2 |
| 2003 | A method for compensation of Jacobian in speaker normalizationabstractIn the conventional maximum likelihood based speaker normalization approach, the optimal frequency warping factors are estimated by maximizing the likelihood of warped features in a grid search. The conventional method of likelihood computation for warped features does not account for the Jacobian of the transformation. This fact is pointed out by some researchers who have also shown that frequency warping is equivalent to the transformation in the cepstral domain. As an approximation, variance normalization of cepstral features is used before likelihood computation to account for the Jacobian. In this paper, we suggest an alternate method to avoid the Jacobian problem. Our preliminary investigation shows that our proposed method provides improvement in normalization performance compared to the conventional method of warping factor estimation for a digit recognition task. Rohit Sinha 0003, Srinivasan Umesh |
ICASSP (1) | 2 |
| 2002 | Non-uniform scaling based speaker normalizationabstractWe present experimental results that show better speaker nonnalization using our previously reported frequency warping function that is derived purely from speech data. In our previous work, we have numerically computed the frequency warping function for non-uniform scaling, which is similar to mel-scale, such that spectral envelopes from different speakers enunciating the same sound are similar except for a possible translation factor. In this paper, we do a maximum likelihood search for these translation parameters and show that this non-uniform normalization scheme provides about 18 % improvement over the normalization method based on the maximum likelihood estimate of uniform scaling parameters and about 30 % improvement over mel filterbank cepstral coefficient based baseline for a telephone based continuous digit recognition task. The other attractive attribute of the proposed method is the simplicity in generating features with different shifts compared to generating features with different warping factors in earlier methods. Rohit Sinha 0003, Srinivasan Umesh |
ICASSP | 2 |
| 2002 | A simple approach to non-uniform vowel normalizationabstractIn this paper, we present results of non-uniform vowel normalization and show that the frequency-warping necessary to do nonuniform vowel nonnalization is similar to the mel-scale. We compare our methods to Fant's non-uniform vowel normalization method and show that with proposed frequency warping approach we can achieve similar performance without any knowledge of the spoken vowel and the fonnant number. The proposed approach is motivated by a desire to perform non-uniform speaker normalization in automatic speech recognition systems. We also present results of a more comprehensive study of our earlier work on non-uniform scaling which again shows that mel-scale is the appropriate warping function. All the results in this paper are based on data from Peterson & Barney and Hillenbrand et al. vowel databases. Srinivasan Umesh, S. V. Bharath Kumar, M. K. Vinay, Rohit Sinha 0003 |
ICASSP | 1 |
| 2002 | Frequency warping and the Mel scaleabstractWe present experimental results that show that the scale-factor relating the formant frequencies of different speakers increases with decreasing values of formant frequency. Based on these results, we experimentally obtain a frequency warping function aimed at separating speaker dependencies from the inherent characterization of the sound. We find that the frequency warping function is similar to the Mel scale, and we believe that this is the first time that a Mel-like scale has been obtained using only speech. Our results and methods may therefore explain, from a speech point of view, the Mel scale, which was obtained historically from hearing based experiments. Srinivasan Umesh, Leon Cohen, Douglas J. Nelson |
IEEE Signal Process. Lett. | 1 |
| 2000 | Exploiting frequency-scaling invariance properties of the scale transform for automatic speech recognitionabstractAn experimental study of the application of scale-transform to improve the performance of speaker independent continuous speech recognition, is presented in this paper. Three major results are described. First, a comparison was made between the scale-transform based magnitude cepstrum coefficients (STCC) and mel-scale filter bank cepstrum coefficients (MFCC) on a telephone based connected digit recognition task. It was shown that the STCC can obtain a performance that is close to that of the MFCC. Second, a simple frequency-normalization procedure was applied to the scale-transform representation that improved performance on the connected digit recognition task with respect to the MFCC. Finally, in a more controlled experimental setting using the TIMIT database, it was shown that the application of phone-specific frequency warpings improved phone classification performance over using a single speaker-specific warping. This last result may have general implications for all frequency warping based speaker normalization procedures. Srinivasan Umesh, Richard C. Rose, Sarangarajan Parthasarathy |
INTERSPEECH | 1 |
| 1999 | Fitting the Mel scaleabstractWe show that there are many qualitatively different equations, each with few parameters, that fit the experimentally obtained Mel scale. We investigate the often made remark that there are two regions to the Mel scale, the first region ( Srinivasan Umesh, Leon Cohen, Douglas J. Nelson |
ICASSP | 1 |
| 1999 | Scale transform in speech analysisabstractIn this paper, we study the scale transform of the spectral-envelope of speech utterances by different speakers. This study is motivated by the hypothesis that the formant frequencies between different speakers are approximately related by a scaling constant for a given vowel. The scale transform has the fundamental property that the magnitude of the scale-transform of a function X(f) and its scaled version /spl radic//spl alpha/X(/spl alpha/f) are same. The methods presented here are useful in reducing variations in acoustic features. We show that the F-ratio tests indicate better separability of vowels by using scale-transform based features than mel-transform based features. The data used in the comparison of the different features consist of 200 utterances of four vowels that are extracted from the TIMIT database. Srinivasan Umesh, Leon Cohen, Nenad Marinovic, Douglas J. Nelson |
IEEE Trans. Speech Audio Process. | 1 |
| 1998 | Improved scale-cepstral analysis in speechabstractWe present improvements over the original scale-cepstrum proposed by Umesh, Cohen, Marinovic and Nelson (see IEEE Trans. on Speech and Audio Processing, 1996). The scale-cepstrum was motivated by a desire to normalize the first-order effects of differences in vocal-tract lengths for a given vowel. Our subsequent work has shown that a more appropriate frequency-warping than the log-warping used in in the original scale-cepstrum is necessary to account for the frequency dependency of the scale-factor. Using this more appropriate frequency-warping and a modified method of computing the scale-cepstrum we have obtained improved features that provide better separability between vowels than before, and are also robust to noise. Srinivasan Umesh, Leon Cohen, Douglas J. Nelson |
ICASSP | 1 |
| 1997 | Frequency-warping and speaker-normalizationabstractWe have proposed the use of scale-cepstral coefficients as features in speech recognition. We have developed a corresponding frequency-warping function, such that, in the warped domain the formant envelopes of different speakers are approximately translated versions of one and another for any given vowel. These methods were motivated by a desire to achieve speaker-normalization. In this paper, we point out very interesting parallels of the various steps in computing the scale-cepstrum, with those observed in computing features based on physiological models of the auditory system or psychoacoustic experiments. It may therefore be useful to have a better understanding of the need for the various signal-processing steps which may result in the development of more robust recognizers. Srinivasan Umesh, Leon Cohen, Douglas J. Nelson |
ICASSP | 1 |
| 1996 | Computationally efficient estimation of sinusoidal frequency at low SNRabstractWe propose a computationally efficient method for estimation of frequency of a single complex sinusoid at low SNR. This method is motivated by the cross-power spectrum method of Nelson (1993) and the weighted phase averager (WPA) methods of Tretter (1985), Kay (1988), and Lovell et al. (1991). We demonstrate that by a simple preprocessing, we can extend the threshold SNR of the WPA significantly. Further, unlike the WPA, the proposed method can be easily extended for estimation of frequencies of multiple sinusoids that are well-separated in frequency. We also derive the variance of the proposed estimator and provide simulation results comparing the proposed method with the WPA. Srinivasan Umesh, Douglas J. Nelson |
ICASSP | 1 |
| 1996 | Frequency-warping in speech
Srinivasan Umesh, Leon Cohen, Nenad Marinovic, Douglas J. Nelson |
ICSLP | 1 |
| 1992 | Resolving the components of transient signals by a multistage procedureabstractThe authors propose a multistage procedure to achieve good local resolution of significantly overlapping components in a transient signal. The received transient signal is assumed to be a superposition of modulated and time-shifted versions of a single one-sided exponential function whose damping coefficient is known. At each stage of the proposed procedure the authors coarsely identify (as compared to the next stage) broad regions in the time-frequency plane that have significant signal energy. These correspond to possible regions where signal components may be present. In each succeeding stage, they focus on these coarsely identified signal regions to further identify subregions in them that have significant energy. The detection of signal components in the subregions is done in the presence of possible interference from neighboring regions. Implementation of this method is based on the use of generalized eigenvectors of signal and interference correlation matrices. The detection statistics is a linear combination of the magnitude squared projections of the received data onto the individual generalized eigenvectors.> Srinivasan Umesh, Donald W. Tufts |
ICASSP | 1 |