Shatrughan Singh

dblp:255/7655 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
7since 2021 · last 2024
0009-0003-8550-4186ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 since 2021Artificial intelligence and machine learning · 9 · 5 since 2021
YearPublicationVenuePosition
2024 Robust Speaker Personalisation Using Generalized Low-Rank Adaptation for Automatic Speech Recognition
abstract
For voice assistant systems, personalizing automated speech recognition (ASR) to a customer is the proverbial holy grail. Careful selection of hyper-parameters will be necessary for fine-tuning a larger ASR model with little speaker data. It is demonstrated that low-rank adaptation (LoRA) is a useful method for optimizing large language models (LLMs). We adapt the ASR model to specific speakers while lowering computational complexity and memory requirements by utilizing low-rank adaptation. In this work, generalized LoRA is used to refine the state-of-the-art cascaded conformer transducer model. To obtain the speaker-specific model, a small number of weights are added to the existing model and finetuned. Improved ASR accuracy across many speakers is observed in experimental assessments, while efficiency is maintained. Using the proposed method, an average relative improvement of 20% in word error rate is obtained across speakers with limited data.
Arun Baby, George Joseph, Shatrughan Singh
ICASSP3
2022 Multi-stage Progressive Compression of Conformer Transducer for On-device Speech Recognition
abstract
The smaller memory bandwidth in smart devices prompts development of smaller Automatic Speech Recognition (ASR) models. To obtain a smaller model, one can employ the model compression techniques. Knowledge distillation (KD) is a popular model compression approach that has shown to achieve smaller model size with relatively lesser degradation in the model performance. In this approach, knowledge is distilled from a trained large size teacher model to a smaller size student model. Also, the transducer based models have recently shown to perform well for on-device streaming ASR task, while the conformer models are efficient in handling long term dependencies. Hence in this work we employ a streaming transducer architecture with conformer as the encoder. We propose a multi-stage progressive approach to compress the conformer transducer model using KD. We progressively update our teacher model with the distilled student model in a multi-stage setup. On standard LibriSpeech dataset, our experimental results have successfully achieved compression rates greater than 60% without significant degradation in the performance compared to the larger teacher model.
Jash Rathod, Nauman Dawalatabad, Shatrughan Singh, Dhananjaya Gowda
INTERSPEECH3
2021 Two-Pass End-to-End ASR Model Compression
abstract
Speech recognition on smart devices is challenging owing to the small memory footprint. Hence small size ASR models are desirable. With the use of popular transducer-based models, it has become possible to practically deploy streaming speech recognition models on small devices [1]. Recently, the two-pass model [2] combining RNN-T and LAS modules has shown exceptional performance for streaming on-device speech recognition. In this work, we propose a simple and effective approach to reduce the size of the two-pass model for memory-constrained devices. We employ a popular knowledge distillation approach in three stages using the Teacher-Student training technique. In the first stage, we use a trained RNN-T model as a teacher model and perform knowledge distillation to train the student RNN-T model. The second stage uses the shared encoder and trains a LAS rescorer for student model using the trained RNN-T+LAS teacher model. Finally, we perform deep-finetuning for the student model with a shared RNN-T encoder, RNN-T decoder, and LAS rescorer. Our experimental results on standard LibriSpeech dataset show that our system can achieve a high compression rate of 55% without significant degradation in the WER compared to the two-pass teacher model.
Nauman Dawalatabad, Tushar Vatsal, Shatrughan Singh, Dhananjaya Gowda, Chanwoo Kim 0001
ASRU5
2021 HiTNet: Byte-to-BPE Hierarchical Transcription Network for End-to-End Speech Recognition
abstract
In this paper, we propose a new byte to byte-pair-encoding (BPE) Hierarchical Transcription Network (HiTNet) architecture for end-to-end (e2e) automatic speech recognition (ASR). The proposed HiTNet architecture simultaneously encodes as well as decodes information hierarchically at different levels of linguistic granularity such as bytes and BPE. In general this idea can be extended to any levels of granularity including phonemes or graphemes or bytes (character to sub-character in some languages), to sub-words or byte-pair encodings (BPE), to words, and so on. Existing hierarchical e2e ASR models primarily encode the acoustic information in an hierarchical manner governed by weaker linguistic constraints at each level. The language information at each level is neither embedded or used explicitly, nor is the information decoded at each level passed on to the next stage. The proposed architecture primarily decodes information in an hierarchical manner utilizing the linguistic information at each level explicitly, while at the same time utilizing the hierarchically encoded acoustic information at each level. Experiments with a two-level byte-to-BPE (b2B) hierarchical transcription show that the proposed architecture significantly reduces the word error rates of both the byte and BPE decoders compared to baseline byte and BPE based attention encoder-decoder models.
Dhananjaya Gowda, Abhinav Garg, Jiyeon Kim, Mehul Kumar, Nauman Dawalatabad, Aman Maghan, Shatrughan Singh, Chanwoo Kim 0001
ASRU10
2021 Voice to Action: Spoken Language Understanding for Memory-Constrained Systems
abstract
Spoken Language Understanding (SLU) is the task of extracting semantic information from speech. Traditional architectures of SLU involve a two-stage pipeline consisting of an Automatic Speech Recognition (ASR) module that converts speech into text hypotheses, followed by a Natural Language Understanding (NLU) module that extracts semantic information from the text hypotheses. The increasing demand for light-weight modules, especially in on-device applications, promulgated end-to-end SLU, where a single module directly converts speech into semantic information. This work pro-poses a low resource, lightweight end-to-end SLU model that treats SLU as a sequence-to-sequence task and is optimized similar to the ASR task. The proposed model utilizes transfer learning from a heavy-resource, general domain ASR model to achieve an 10.8% absolute improvement in the Word Error Rate (WER) over the base model. We then utilize Knowledge Distillation to achieve a further 31.9% reduction in the model size with 22.7% relative WER improvement. In addition, we compress our models by more than 4 times using Low-Rank Approximation (LRA) method with minimum degradation in accuracy.
Aditya Jayasimha, Aman Maghan, Shatrughan Singh, Dhananjaya Gowda, Chanwoo Kim 0001
ASRU4
2021 Comparative Study of Different Tokenization Strategies for Streaming End-to-End ASR
abstract
Most End-to-End Automatic Speech Recognition (ASR) models use character-based vocabularies: characters, sub-words (BPE), or words. While these work well for training a monolingual model, there are certain limitations when ap-plying these to a multilingual model. Rare characters from character-rich languages like Korean can easily result in large vocabulary size, limiting the model's compactness. Repre-senting text at the level of bytes has also been proposed. However, a byte sequence representation of text is often much longer, which increases the decoding time and makes it computationally expensive for on-device use. Byte-based sub-words (BBPE) are proposed in neural machine translation for word representation but are still unexplored in the ASR domain. In this work, we conduct an empirical study comparing the above three tokenization strategies across three metrics: Word Error Rate (WER), model size, and the decoding time, which are critical for an on-device ASR. We did extensive experiments for both monolingual and bilingual, with languages belonging to same (English and Spanish) and different (English and Korean) language families. Our exper-iments show that BBPE and BPE models yield a similar WER for English and Spanish. While for a character-rich language like Korean, we get 26% and 14% relative WER improvement with BBPE monolingual and bilingual models, respectively. In contrast, the byte models trade-off small model size and a fixed vocabulary at the cost of high xRT. Among all three, we found the BBPE strategy to be the most flexible and optimal for most cases.
Aman Maghan, Dhananjaya Gowda, Shatrughan Singh, Chanwoo Kim 0001
ASRU5
2021 Neural Utterance Confidence Measure for RNN-Transducers and Two Pass Models
abstract
In this paper, we propose methods to compute confidence score on the predictions made by an end-to-end speech recognition model in a 2-pass framework. We use RNN-Transducer for a streaming model, and an attention-based decoder for the second pass model. We use neural technique to compute the confidence score, and experiment with various combinations of features from RNN-Transducer and second pass models. The neural confidence score model is trained as a binary classification task to accept or reject a prediction made by speech recognition model. The model is evaluated in a distributed speech recognition environment, and performs significantly better when features from second pass model are used as compared to the features from streaming model.
Dhananjaya Gowda, Kwangyoun Kim, Shatrughan Singh, Chanwoo Kim 0001
ICASSP6
2020 Hierarchical Multi-Stage Word-to-Grapheme Named Entity Corrector for Automatic Speech Recognition
Abhinav Garg, Dhananjaya Gowda, Shatrughan Singh, Chanwoo Kim 0001
INTERSPEECH4
2020 Utterance Invariant Training for Hybrid Two-Pass End-to-End Speech Recognition
Dhananjaya Gowda, Kwangyoun Kim, Hejung Yang, Abhinav Garg, Jiyeon Kim, Mehul Kumar, Sichen Jin, Shatrughan Singh, Chanwoo Kim 0001
INTERSPEECH10
2020 Utterance Confidence Measure for End-to-End Speech Recognition with Applications to Distributed Speech Recognition Scenarios
Dhananjaya Gowda, Abhinav Garg, Shatrughan Singh, Chanwoo Kim 0001
INTERSPEECH5
2019 End-to-End Training of a Large Vocabulary End-to-End Speech Recognition System
abstract
In this paper, we present an end-to-end training framework for building state-of-the-art end-to-end speech recognition systems. Our training system utilizes a cluster of Central Processing Units (CPUs) and Graphics Processing Units (GPUs). The entire data reading, large scale data augmentation, neural network parameter updates are all performed “on-the-fly”. We use vocal tract length perturbation [1] and an acoustic simulator [2] for data augmentation. The processed features and labels are sent to the GPU cluster. The Horovod allreduce approach is employed to train neural network parameters. We evaluated the effectiveness of our system on the standard Librispeech corpus [3] and the 10,000-hr anonymized Bixby English dataset. Our end-to-end speech recognition system built using this training infrastructure showed a 2.44 % WER on test-clean of the LibriSpeech test set after applying shallow fusion with a Transformer language model (LM). For the proprietary English Bixby open domain test set, we obtained a WER of 7.92 % using a Bidirectional Full Attention (BFA) end-to-end model after applying shallow fusion with an RNN-LM. When the monotonic chunckwise attention (MoCha) based approach is employed for streaming speech recognition, we obtained a WER of 9.95 % on the same Bixby open domain test set.
Chanwoo Kim 0001, Minkyoo Shin, Shatrughan Singh, Larry Heck, Dhananjaya Gowda, Kwangyoun Kim, Mehul Kumar, Jiyeon Kim, Kyungmin Lee, Abhinav Garg, Eunhyang Kim
ASRU3