Abhinav Garg

dblp:202/5696 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0001-5082-5500ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2024 Data Driven Grapheme-to-Phoneme Representations for a Lexicon-Free Text-to-Speech
abstract
Grapheme-to-Phoneme (G2P) is an essential first step in any modern, high-quality Text-to-Speech (TTS) system. Most of the current G2P systems rely on carefully hand-crafted lexicons developed by experts. This poses a two-fold problem. Firstly, the lexicons are generated using a fixed phoneme set, usually, ARPABET or IPA, which might not be the most optimal way to represent phonemes for all languages. Secondly, the man-hours required to produce such an expert lexicon are very high. In this paper, we eliminate both of these issues by using recent advances in self-supervised learning to obtain data-driven phoneme representations instead of fixed representations. We compare our lexicon-free approach against strong baselines that utilize a well-crafted lexicon. Furthermore, we show that our data-driven lexicon-free method performs as good or even marginally better than the conventional rule-based or lexicon-based neural G2Ps in terms of Mean Opinion Score (MOS) while using no prior language lexicon or phoneme set, i.e. no linguistic expertise.
Abhinav Garg, Jiyeon Kim, Sushil Khyalia, Chanwoo Kim 0001, Dhananjaya Gowda
ICASSP1
2023 Self-Supervised Accent Learning for Under-Resourced Accents Using Native Language Data
abstract
In this paper, we propose a novel method to improve the accuracy of an English speech recognizer for a target accent using the corresponding native language data. Collecting labeled data for all accents of English to train an end-to-end neural speech recognizer for English is a difficult and expensive task. Also, finding a pool of representative English speakers for any arbitrary accent to collect unlabeled data can be a difficult task. However, collecting unlabeled speech data for any native language is a much simpler task. It is important to note that the accents of most non-native English speakers are heavily biased by the co-articulation of sounds in their own native language. In view of this, we propose to use unlabeled native language data to learn self-supervised representations during the pre-training stage. The pre-trained model is then fine-tuned using limited labeled English data for the target accent. Experiments using native language data to pre-train an English recognizer followed by fine-tuning using target accented English show significant improvements in word error rates on four different accents (Great Britain, Korean, Chinese, Spanish).
Mehul Kumar, Jiyeon Kim, Dhananjaya Gowda, Abhinav Garg, Chanwoo Kim 0001
ICASSP4
2021 HiTNet: Byte-to-BPE Hierarchical Transcription Network for End-to-End Speech Recognition
abstract
In this paper, we propose a new byte to byte-pair-encoding (BPE) Hierarchical Transcription Network (HiTNet) architecture for end-to-end (e2e) automatic speech recognition (ASR). The proposed HiTNet architecture simultaneously encodes as well as decodes information hierarchically at different levels of linguistic granularity such as bytes and BPE. In general this idea can be extended to any levels of granularity including phonemes or graphemes or bytes (character to sub-character in some languages), to sub-words or byte-pair encodings (BPE), to words, and so on. Existing hierarchical e2e ASR models primarily encode the acoustic information in an hierarchical manner governed by weaker linguistic constraints at each level. The language information at each level is neither embedded or used explicitly, nor is the information decoded at each level passed on to the next stage. The proposed architecture primarily decodes information in an hierarchical manner utilizing the linguistic information at each level explicitly, while at the same time utilizing the hierarchically encoded acoustic information at each level. Experiments with a two-level byte-to-BPE (b2B) hierarchical transcription show that the proposed architecture significantly reduces the word error rates of both the byte and BPE decoders compared to baseline byte and BPE based attention encoder-decoder models.
Dhananjaya Gowda, Abhinav Garg, Jiyeon Kim, Mehul Kumar, Nauman Dawalatabad, Aman Maghan, Shatrughan Singh, Chanwoo Kim 0001
ASRU2
2021 Semi-Supervised Transfer Learning for Language Expansion of End-to-End Speech Recognition Models to Low-Resource Languages
Jiyeon Kim, Mehul Kumar, Dhananjaya Gowda, Abhinav Garg, Chanwoo Kim 0001
ASRU4
2021 A Comparison of Streaming Models and Data Augmentation Methods for Robust Speech Recognition
abstract
In this paper, we present a comparative study on the robustness of two different online streaming speech recognition models: Monotonic Chunkwise Attention (MoChA) and Recurrent Neural Network-Transducer (RNN-T). We explore three recently proposed data augmentation techniques, namely, multi-conditioned training using an acoustic simulator, Vocal Tract Length Perturbation (VTLP) for speaker variability, and SpecAugment. Experimental results show that unidirectional models are in general more sensitive to noisy examples in the training set. It is observed that the final performance of the model depends on the proportion of training examples processed by data augmentation techniques. MoChA models generally perform better than RNN-T models. However, we observe that training of MoChA models seems to be more sensitive to various factors such as the characteristics of training sets and the incorporation of additional augmentations techniques. On the other hand, RNN-T models perform better than MoChA models in terms of latency, inference time, and the stability of training. Additionally, RNN-T models are generally more robust against noise and reverberation. All these advantages make RNN-T models a better choice for streaming on-device speech recognition compared to MoChA models.
Jiyeon Kim, Mehul Kumar, Dhananjaya Gowda, Abhinav Garg, Chanwoo Kim 0001
ASRU4
2021 Streaming End-to-End Speech Recognition with Jointly Trained Neural Feature Enhancement
abstract
In this paper, we present a streaming end-to-end speech recognition model based on Monotonic Chunkwise Attention (MoCha) jointly trained with enhancement layers. Even though the MoCha attention enables streaming speech recognition with recognition accuracy comparable to a full attention-based approach, training this model is sensitive to various factors such as the difficulty of training examples, hyper-parameters, and so on. Because of these issues, speech recognition accuracy of a MoCha-based model for clean speech drops significantly when a multi-style training approach is applied. Inspired by Curriculum Learning [1], we introduce two training strategies: Gradual Application of Enhanced Features (GAEF) and Gradual Reduction of Enhanced Loss (GREL). With GAEF, the model is initially trained using clean features. Subsequently, the portion of outputs from the enhancement layers gradually increases. With GREL, the portion of the Mean Squared Error (MSE) loss for the enhanced output gradually reduces as training proceeds. In experimental results on the LibriSpeech corpus and noisy far-field test sets, the proposed model with GAEF-GREL training strategies shows significantly better results than the conventional multi-style training approach.
Chanwoo Kim 0001, Abhinav Garg, Dhananjaya Gowda, Seongkyu Mun
ICASSP2
2020 Hierarchical Multi-Stage Word-to-Grapheme Named Entity Corrector for Automatic Speech Recognition
Abhinav Garg, Dhananjaya Gowda, Shatrughan Singh, Chanwoo Kim 0001
INTERSPEECH1
2020 Streaming On-Device End-to-End ASR System for Privacy-Sensitive Voice-Typing
Abhinav Garg, Gowtham P. Vadisetti, Dhananjaya Gowda, Sichen Jin, Aditya Jayasimha, Youngho Han, Jiyeon Kim, Junmo Park, Kwangyoun Kim, Young-Yoon Lee, Kyungbo Min, Chanwoo Kim 0001
INTERSPEECH1
2020 Utterance Invariant Training for Hybrid Two-Pass End-to-End Speech Recognition
Dhananjaya Gowda, Kwangyoun Kim, Hejung Yang, Abhinav Garg, Jiyeon Kim, Mehul Kumar, Sichen Jin, Shatrughan Singh, Chanwoo Kim 0001
INTERSPEECH5
2020 Utterance Confidence Measure for End-to-End Speech Recognition with Applications to Distributed Speech Recognition Scenarios
Dhananjaya Gowda, Abhinav Garg, Shatrughan Singh, Chanwoo Kim 0001
INTERSPEECH4
2019 Improved Multi-Stage Training of Online Attention-Based Encoder-Decoder Models
abstract
In this paper, we propose a refined multi-stage multi-task training strategy to improve the performance of onlineattention-based encoder-decoder (AED) models. A three-stage training based on three levels of architectural granularity namely, character encoder, byte pair encoding (BPE) based encoder, and attention decoder, is proposed. Also, multi-task learning based on two-levels of linguistic granularity namely, character and BPE, is used. We explore different pre-training strategies for the encoders including transfer learning from a bidirectional encoder. Our encoder-decoder models with online attention show ~35% and ~10% relative improvement over their baselines for smaller and bigger models, respectively. Our models achieve a word error rate (WER) of 5.04% and 4.48% on the Librispeech test-clean data for the smaller and bigger models respectively after fusion with long short-term memory (LSTM) based external language model (LM).
Abhinav Garg, Dhananjaya Gowda, Kwangyoun Kim, Mehul Kumar, Chanwoo Kim 0001
ASRU1
2019 End-to-End Training of a Large Vocabulary End-to-End Speech Recognition System
abstract
In this paper, we present an end-to-end training framework for building state-of-the-art end-to-end speech recognition systems. Our training system utilizes a cluster of Central Processing Units (CPUs) and Graphics Processing Units (GPUs). The entire data reading, large scale data augmentation, neural network parameter updates are all performed “on-the-fly”. We use vocal tract length perturbation [1] and an acoustic simulator [2] for data augmentation. The processed features and labels are sent to the GPU cluster. The Horovod allreduce approach is employed to train neural network parameters. We evaluated the effectiveness of our system on the standard Librispeech corpus [3] and the 10,000-hr anonymized Bixby English dataset. Our end-to-end speech recognition system built using this training infrastructure showed a 2.44 % WER on test-clean of the LibriSpeech test set after applying shallow fusion with a Transformer language model (LM). For the proprietary English Bixby open domain test set, we obtained a WER of 7.92 % using a Bidirectional Full Attention (BFA) end-to-end model after applying shallow fusion with an RNN-LM. When the monotonic chunckwise attention (MoCha) based approach is employed for streaming speech recognition, we obtained a WER of 9.95 % on the same Bixby open domain test set.
Chanwoo Kim 0001, Minkyoo Shin, Shatrughan Singh, Larry Heck, Dhananjaya Gowda, Kwangyoun Kim, Mehul Kumar, Jiyeon Kim, Kyungmin Lee, Abhinav Garg, Eunhyang Kim
ASRU12
2019 ReCall: Crowdsourcing on Basic Phones to Financially Sustain Voice Forums
abstract
Although voice forums are widely used to enable marginalized communities to produce, consume, and share information, their financial sustainability is a key concern among HCI4D researchers and practitioners. We present ReCall, a crowdsourcing marketplace accessible via phone calls where low-income rural residents vocally transcribe audio files to gain free airtime to participate in voice forums as well as to earn money. We conducted a series of experimental and usability evaluations with 28 low-income people in rural India to examine the effect of phone types, channel types, and review modes on speech transcription performance. We then deployed ReCall for two weeks to 24 low-income rural residents who placed 5,879 phone calls, completed 29,000 micro tasks to yield transcriptions with 85% accuracy, and earned INR 20,500. Our mixed-methods analysis indicates that each minute of crowd work on ReCall gives users eight minutes of free airtime on another voice forum, and thus illustrates a way to address the financial sustainability of voice forums.
Aditya Vashistha, Abhinav Garg, Richard J. Anderson 0001
CHI2
2019 Threats, Abuses, Flirting, and Blackmail: Gender Inequity in Social Media Voice Forums
abstract
HCI4D researchers and practitioners have leveraged voice forums to enable people with literacy, socioeconomic, and connectivity barriers to access, report, and share information. Although voice forums have received impassioned usage from low-income, low-literate, rural, tribal, and disabled communities in diverse HCI4D contexts, the participation of women in these services is almost non-existent. In this paper, we investigate the reasons for the low participation of women in social media voice forums by examining the use of Sangeet Swara in India and Baang in Pakistan by marginalized women and men. Our mixed-methods approach spanning content analysis of audio posts, quantitative analysis of interactions between users, and qualitative interviews with users indicate gender inequity due to deep-rooted patriarchal values. We found that women on these forums faced systemic discrimination and encountered abusive content, flirts, threats, and harassment. We discuss design recommendations to create social media voice forums that foster gender equity in use of these services.
Aditya Vashistha, Abhinav Garg, Richard J. Anderson 0001, Agha Ali Raza
CHI2
2019 Multi-Task Multi-Resolution Char-to-BPE Cross-Attention Decoder for End-to-End Speech Recognition
Dhananjaya Gowda, Abhinav Garg, Kwangyoun Kim, Mehul Kumar, Chanwoo Kim 0001
INTERSPEECH2
2019 Improved Vocal Tract Length Perturbation for a State-of-the-Art End-to-End Speech Recognition System
Chanwoo Kim 0001, Minkyu Shin, Abhinav Garg, Dhananjaya Gowda
INTERSPEECH3