VLDB 2026 Research / reviewers in the wild / expert
Preethi Jyothi
dblp:01/9014
· DBLP profile ↗
78ranked-venue papers
10as first author
44since 2021 · last 2026
0009-0000-4173-3348ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 64 · 7 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 10 first-author · 23 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Linguistically informed automatic speech recognition in Sanskrit
Rishabh Kumar, Devaraj Adiga, Rishav Ranjan, Amrith Krishna, Ganesh Ramakrishnan, Pawan Goyal 0002, Preethi Jyothi |
Comput. Speech Lang. | 7 |
| 2025 | LexGen: Domain-aware Multilingual Lexicon GenerationabstractLexicon or dictionary generation across domains has the potential for societal impact, as it can potentially enhance information accessibility for a diverse user base while preserving language identity. Prior work in the field primarily focuses on bilingual lexical induction, which deals with word alignments using mapping-based or corpora-based approaches. However, these approaches do not cater to domain-specific lexicon generation that consists of domain-specific terminology. This task becomes particularly important in specialized medical, engineering, and other technical domains, owing to the highly infrequent usage of the terms and scarcity of data involving domain-specific terms especially for low-resource languages. We propose a new model to generate dictionary words for 6 Indian languages in the multi-domain setting. Our model consists of domain-specific and domain-generic layers that encode information, and these layers are invoked via a learnable routing technique. We also release a new benchmark dataset consisting of >75K translation pairs across 6 Indian languages spanning 8 diverse domains. We conduct both zero-shot and few-shot experiments across multiple domains to show the efficacy of our proposed model in generalizing to unseen domains and unseen languages. Additionally, we also perform a human post-hoc evaluation on unseen languages. The source code and dataset is present at https://github.com/Atulkmrsingh/lexgen. Ayush Maheshwari, Atul Kumar Singh, N. J. Karthika, Krishnakant Bhatt, Preethi Jyothi, Ganesh Ramakrishnan |
ACL (1) | 5 |
| 2025 | CoSTA: Code-Switched Speech Translation using Aligned Speech-Text InterleavingabstractCode-switching is a widely prevalent linguistic phenomenon in multilingual societies like India. Building speech-to-text models for code-switched speech is challenging due to limited availability of datasets. In this work, we focus on the problem of spoken translation (ST) of code-switched speech in Indian languages to English text. We present a new end-to-end model architecture CoSTA that scaffolds on pretrained automatic speech recognition (ASR) and machine translation (MT) modules (that are more widely available for many languages). Speech and ASR text representations are fused using an aligned interleaving scheme and are fed further as input to a pretrained MT module; the whole pipeline is then trained end-to-end for spoken translation using synthetically created ST data. We also release a new evaluation benchmark for code-switched Bengali- English, Hindi-English, Marathi-English and Telugu-English speech to English text. CoSTA significantly outperforms many competitive cascaded and end-to-end multimodal baselines by up to 3.5 BLEU points. Bhavani Shankar, Preethi Jyothi, Pushpak Bhattacharyya |
COLING | 2 |
| 2025 | LASER: An LLM-based ASR Scoring and Evaluation RubricabstractStandard ASR evaluation metrics like Word Error Rate (WER) tend to unfairly penalize morphological and syntactic nuances that do not significantly alter sentence semantics.We introduce an LLM-based scoring rubric LASER that leverages state-of-the-art LLMs' in-context learning abilities to learn from prompts with detailed examples.Hindi LASER scores using Gemini 2.5 Pro achieved a very high correlation score of 94% with human annotations.Hindi examples in the prompt were also effective in analyzing errors in other Indian languages such as Marathi, Kannada and Malayalam.We also demonstrate how a smaller LLM like Llama 3 can be finetuned on word-pair examples derived from reference and ASR predictions to predict penalty types with close to 89% accuracy.Error type Example variations Penalty Numerical Phrases "1300" vs "Terah sau" or "Ek hajar teen sau" No penalty Abbreviations "ATM" vs "Ay Ti Em" vs "Ay tee yum" No penalty Compound Words "bhajan sangraha" vs "bhajansangraha" No penalty Transliterations (Native spellings) "ayskreem" vs "aaiskrim" or "skul" vs "skool" No penalty Actual transliterations "ice cream" vs "ayskrim" or "aaiskrim" No penalty Acceptable alternate spellings "sundar with a bindu" vs "sundar with a half na" No penalty Proper nouns "Priya" vs "Pria" vs "Preeya" vs "Preya" No penalty Slang and Colloquial terms "Yaha" vs "Ye" or "vaha" vs "vo" or "par" vs "pe" No penalty Small (single character) spelling errors "ladki" vs "ladkee" or "bahut" vs "bahoot" Minor penalty Small grammatical errors (gender/tense/number) "hain" vs "hai" or "uska" vs "uski" vs "usko" Minor penalty Spelling errors that alter meaning "kumar" vs "kamar" or "saman" vs "samanya" Major penalty Incorrect word substitutions "sundar" vs "bhadda" or "mota" vs "chhota" Major penalty Significant omissions or additions "-" vs "sundar" or "mota" vs "-" Major penalty Reordering of words that changes meaning "bahut accha khana" vs "bahut khana accha" Major penalty Amruta Parulekar, Preethi Jyothi |
EMNLP | 2 |
| 2025 | Skip-Salsa: Skip Synchronous Fusion of ASR LLM Decoders
Ashish R. Mittal, Darshan Prabhu, Sunita Sarawagi, Preethi Jyothi |
INTERSPEECH | 4 |
| 2025 | BlockDecoder: Boosting ASR Decoders with Context and Merger ModulesabstractAttention-based encoder decoder models remain a popular choice for state-of-the-art automatic speech recognition (ASR). These models combine a powerful audio encoder that extracts rich acoustic features with a decoder that autoregressively produces the ASR output. The decoder handles two critical tasks: (1) building rich text-only context and (2) merging acoustic information from the encoder to ensure the predictions remain faithful to the audio. We observe a systematic pattern across the attention distributions of decoder layers in prior architectures: the initial layers direct most attention towards building textual context, while the later layers largely focus on merging acoustic and textual information for the final predictions. Leveraging this key insight, we propose **BlockDecoder**, a novel decoder architecture comprising two distinct components: a text encoder that is purely text-based, and a **Merger** that combines information from the audio encoder and text encoder to generate output tokens. Unlike traditional decoders, the **Merger** autoregressively predicts a sequence of K tokens within a *block* of size K, while relying on the same precomputed contextual information from both text and audio encoders across the block. This design choice allows for the efficient reuse of encoder representations. The separation of the decoder into the text encoder and the **Merger** promotes modularity and more flexible control of parameters via the number of text encoder and **Merger** layers. As a result, **BlockDecoder** yields a significant speedup ($\sim2$x) compared to traditional decoders, across diverse datasets, languages, and speech tasks, without any degradation in performance. The code is available at https://github.com/csalt-research/blockdecoder. Darshan Prabhu, Preethi Jyothi |
NeurIPS | 2 |
| 2024 | In-context Mixing (ICM): Code-mixed Prompts for Multilingual LLMsabstractWe introduce a simple and effective prompt ing technique called incontext mixing (ICM) for effective incontext learning (ICL) with multilingual large language models (MLLMs).With ICM, we modify the fewshot examples within ICL prompts to be intrasententially codemixed by randomly swapping content words in the target languages with their English translations.We observe that ICM prompts yield superior performance in NLP tasks such as disfluency correction, grammar error cor rection and text simplification that demand a close correspondence between the input and out put sequences.Significant improvements are observed mainly for lowresource languages that are underrepresented during the pretrain ing and finetuning of MLLMs.We present an extensive set of experiments to analyze when ICM is effective and what design choices con tribute towards its effectiveness.ICM works consistently and significantly better than other prompting techniques across models of varying capacity such as mT0XXL, BloomZ and GPT 4. Code, prompts and datasets are available here. Bhavani Shankar, Preethi Jyothi, Pushpak Bhattacharyya |
ACL (1) | 2 |
| 2024 | Emotion Arithmetic: Emotional Speech Synthesis via Weight Space Interpolation
Pavan Kalyan, Preeti Rao, Preethi Jyothi, Pushpak Bhattacharyya |
INTERSPEECH | 3 |
| 2024 | SALSA: Speedy ASR-LLM Synchronous Aggregation
Ashish R. Mittal, Darshan Prabhu, Sunita Sarawagi, Preethi Jyothi |
INTERSPEECH | 4 |
| 2024 | Improving Self-supervised Pre-training using Accent-Specific Codebooks
Darshan Prabhu, Omkar Nitsure, Preethi Jyothi, Sriram Ganapathy |
INTERSPEECH | 4 |
| 2024 | MULTI-CONVFORMER: Extending Conformer with Multiple Convolution Kernels
Darshan Prabhu, Yifan Peng 0003, Preethi Jyothi, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2024 | WikiDO: A New Benchmark Evaluating Cross-Modal Retrieval for Vision-Language ModelsabstractCross-modal (image-to-text and text-to-image) retrieval is an established task used in evaluation benchmarks to test the performance of vision-language models (VLMs). Several state-of-the-art VLMs (e.g. CLIP, BLIP-2) have achieved near-perfect performance on widely-used image-text retrieval benchmarks such as MSCOCO-Test-5K and Flickr30K-Test-1K. As a measure of out-of-distribution (OOD) generalization, prior works rely on zero-shot performance evaluated on one dataset (Flickr) using a VLM finetuned on another one (MSCOCO). We argue that such comparisons are insufficient to assess the OOD generalization capability of models due to high visual and linguistic similarity between the evaluation and finetuning datasets. To address this gap, we introduce WikiDO (drawn from Wikipedia Diversity Observatory), a novel cross-modal retrieval benchmark to assess the OOD generalization capabilities of pretrained VLMs. This consists of newly scraped 380K image-text pairs from Wikipedia with domain labels, a carefully curated, human-verified a)in-distribution (ID) test set (3K) and b) OOD test set (3K). The image-text pairs are very diverse in topics and geographical locations. We evaluate different VLMs of varying capacity on the \wikido benchmark; BLIP-2 achieves zero-shot performance of $R@1\approx66\%$ on the OOD test set, compared to $\approx$ $81\%$ on COCO and $\approx95\%$ on Flickr. When fine-tuned on WikiDO, the $R@1$ improvement is at most $\approx5\%$ on OOD instances compared to $\approx12\%$ on ID instances. We probe the VLMs with varying finetuning objectives and datasets of varying sizes to identify what aids OOD generalization the most. Our results confirm that WikiDO offers a strong cross-modal benchmark for current VLMs in specifically evaluating for OOD generalization. Our benchmark is hosted as a competition at https://kaggle.com/competitions/wikido24 with public access to dataset and code. Tankala Pavan Kalyan, Piyush Singh Pasi, Sahil Dharod, Azeem Motiwala, Preethi Jyothi, Aditi Chaudhary, Krishna Srinivasan |
NeurIPS | 5 |
| 2023 | Improving Pretraining Techniques for Code-Switched NLPabstractPretrained models are a mainstay in modern NLP applications.Pretraining requires access to large volumes of unlabeled text.While monolingual text is readily available for many of the world's languages, access to large quantities of code-switched text (i.e., text with tokens of multiple languages interspersed within a sentence) is much more scarce.Given this resource constraint, the question of how pretraining using limited amounts of code-switched text could be altered to improve performance for code-switched NLP becomes important to tackle.In this paper, we explore different masked language modeling (MLM) pretraining techniques for code-switched text that are cognizant of language boundaries prior to masking.The language identity of the tokens can either come from human annotators, trained language classifiers, or simple relative frequencybased estimates.We also present an MLM variant by introducing a residual connection from an earlier layer in the pretrained model that uniformly boosts performance on downstream tasks.Experiments on two downstream tasks, Question Answering (QA) and Sentiment Analysis (SA), involving four code-switched language pairs (Hindi-English, Spanish-English, Tamil-English, Malayalam-English) yield relative improvements of up to 5.8 and 2.7 F1 scores on QA (Hindi-English) and SA (Tamil-English), respectively, compared to standard pretraining techniques.To understand our task improvements better, we use a series of probes to study what additional information is encoded by our pretraining techniques and also introduce an auxiliary loss function that explicitly models language identification to further aid the residual MLM variants. Richeek Das, Sahasra Ranjan, Shreya Pathak, Preethi Jyothi |
ACL (1) | 4 |
| 2023 | DITTO: Data-efficient and Fair Targeted Subset Selection for ASR Accent AdaptationabstractSuraj Kothawade, Anmol Mekala, D.Chandra Sekhara Hetha Havya, Mayank Kothyari, Rishabh Iyer, Ganesh Ramakrishnan, Preethi Jyothi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Suraj Kothawade, Anmol Reddy Mekala, D. Chandra Sekhara Hetha Havya, Mayank Kothyari, Rishabh Iyer 0001, Ganesh Ramakrishnan, Preethi Jyothi |
ACL (1) | 7 |
| 2023 | Speech-enriched Memory for Inference-time Adaptation of ASR Models to Word DictionariesabstractDespite the impressive performance of ASR models on mainstream benchmarks, their performance on rare words is unsatisfactory.In enterprise settings, often a focused list of entities (such as locations, names, etc) are available which can be used to adapt the model to the terminology of specific domains.In this paper, we present a novel inference algorithm that improves the prediction of state-of-the-art ASR models using nearest-neighbor-based matching on an inference-time word list.We consider both the Transducer architecture that is useful in the streaming setting, and state-of-the-art encoder-decoder models such as Whisper.In our approach, a list of rare entities is indexed in a memory by synthesizing speech for each entry, and then storing the internal acoustic and language model states obtained from the best possible alignment on the ASR model.The memory is organized as a trie which we harness to perform a stateful lookup during inference.A key property of our extension is that we prevent spurious matches by restricting to only word-level matches.In our experiments on publicly available datasets and private benchmarks, we show that our method is effective in significantly improving rare word recognition. Ashish R. Mittal, Sunita Sarawagi, Preethi Jyothi, George Saon, Gakuto Kurata |
EMNLP | 3 |
| 2023 | Accented Speech Recognition With Accent-specific CodebooksabstractSpeech accents pose a significant challenge to state-of-the-art automatic speech recognition (ASR) systems.Degradation in performance across underrepresented accents is a severe deterrent to the inclusive adoption of ASR.In this work, we propose a novel accent adaptation approach for end-to-end ASR systems using cross-attention with a trainable set of codebooks.These learnable codebooks capture accent-specific information and are integrated within the ASR encoder layers.The model is trained on accented English speech, while the test data also contained accents which were not seen during training.On the Mozilla Common Voice multi-accented dataset, we show that our proposed approach yields significant performance gains not only on the seen English accents (up to 37% relative improvement in word error rate) but also on the unseen accents (up to 5% relative improvement in WER).Further, we illustrate benefits for a zero-shot transfer setup on the L2Artic dataset.We also compare the performance with other approaches based on accent adversarial training. Fbank Feats CTC Darshan Prabhu, Preethi Jyothi, Sriram Ganapathy, Vinit Unni |
EMNLP | 2 |
| 2023 | Towards Zero-Shot Code-Switched Speech RecognitionabstractIn this work, we seek to build effective code-switched (CS) automatic speech recognition systems (ASR) under the zero-shot set-ting where no transcribed CS speech data is available for training. Previously proposed frameworks which conditionally factorize the bilingual task into its constituent monolingual parts are a promising starting point for leveraging monolingual data efficiently. However, these methods require the monolingual modules to perform language segmentation. That is, each monolingual module has to simultaneously detect CS points and transcribe speech segments of one language while ignoring those of other languages – not a trivial task. We propose to simplify each monolingual module by allowing them to transcribe all speech segments indiscriminately with a monolingual script (i.e. transliteration). This simple modification passes the responsibility of CS point detection to subsequent bilingual modules which determine the final output by considering multiple monolingual transliterations along with external language model information. We apply this transliteration-based approach in an end-to-end differentiable neural network and demonstrate its efficacy for zero-shot CS ASR on Mandarin-English SEAME test sets. Brian Yan, Matthew Wiesner, Ondrej Klejch, Preethi Jyothi, Shinji Watanabe 0001 |
ICASSP | 4 |
| 2023 | In-Situ Text-Only Adaptation of Speech Models with Low-Overhead Speech Imputations
Ashish R. Mittal, Sunita Sarawagi, Preethi Jyothi |
ICLR | 3 |
| 2023 | Temporally Aligning Long Audio Interviews with Questions: A Case Study in Multimodal Data IntegrationabstractThe problem of audio-to-text alignment has seen significant amount of research using complete supervision during training. However, this is typically not in the context of long audio recordings wherein the text being queried does not appear verbatim within the audio file. This work is a collaboration with a non-governmental organization called CARE India that collects long audio health surveys from young mothers residing in rural parts of Bihar, India. Given a question drawn from a questionnaire that is used to guide these surveys, we aim to locate where the question is asked within a long audio recording. This is of great value to African and Asian organizations that would otherwise have to painstakingly go through long and noisy audio recordings to locate questions (and answers) of interest. Our proposed framework, INDENT, uses a cross-attention-based model and prior information on the temporal ordering of sentences to learn speech embeddings that capture the semantics of the underlying spoken text. These learnt embeddings are used to retrieve the corresponding audio segment based on text queries at inference time. We empirically demonstrate the significant effectiveness (improvement in R-avg of about 3%) of our model over those obtained using text-based heuristics. We also show how noisy ASR, generated using state-of-the-art ASR models for Indian languages, yields better results when used in place of speech. INDENT, trained only on Hindi data is able to cater to all languages supported by the (semantically) shared text space. We illustrate this empirically on 11 Indic languages. Piyush Singh Pasi, Karthikeya Battepati, Preethi Jyothi, Ganesh Ramakrishnan, Tanmay Mahapatra, Manoj Singh |
IJCAI | 3 |
| 2023 | DisfluencyFixer: A tool to enhance Language Learning through Speech To Speech Disfluency Correction
Vineet Bhat, Preethi Jyothi, Pushpak Bhattacharyya |
INTERSPEECH | 2 |
| 2023 | Unsupervised Code-switched Text Generation from Parallel TextabstractSpeech is a fundamental means of communication that can be seen to provide two channels for transmitting information: the lexical channel of which words are said, and the non-lexical channel of how they are spoken. Both channels shape listener expectations of upcoming communication; however, directly quantifying their relative effect on expectations is challenging. Previous attempts require spoken variations of lexically equivalent dialogue turns or conspicuous acoustic manipulations. This paper introduces a generalised paradigm to study the value of non-lexical information in dialogue across unconstrained lexical content. By quantifying the perceptual value of the non-lexical channel with both accuracy and entropy reduction, we show that non-lexical information produces a consistent effect on expectations of upcoming dialogue: even when it leads to poorer discriminative turn judgements than lexical content alone, it yields higher consensus among participants. Jie Chi, Brian Lu, Jason Eisner, Peter Bell 0001, Preethi Jyothi, Ahmed Ali 0002 |
INTERSPEECH | 5 |
| 2023 | Narrator or Character: Voice Modulation in an Expressive Multi-speaker TTS
Tankala Pavan Kalyan, Preeti Rao, Preethi Jyothi, Pushpak Bhattacharyya |
INTERSPEECH | 3 |
| 2023 | Improving RNN-Transducers with Acoustic LookAhead
Vinit Unni, Ashish R. Mittal, Preethi Jyothi, Sunita Sarawagi |
INTERSPEECH | 3 |
| 2022 | Accurate Online Posterior Alignments for Principled Lexically-Constrained DecodingabstractOnline alignment in machine translation refers to the task of aligning a target word to a source word when the target sequence has only been partially decoded.Good online alignments facilitate important applications such as lexically constrained translation where userdefined dictionaries are used to inject lexical constraints into the translation model.We propose a novel posterior alignment technique that is truly online in its execution and superior in terms of alignment error rates compared to existing methods.Our proposed inference technique jointly considers alignment and token probabilities in a principled manner and can be seamlessly integrated within existing constrained beam-search decoding algorithms.On five language pairs, including two distant language pairs, we achieve consistent drop in alignment error rates.When deployed on seven lexically constrained translation tasks, we achieve significant improvements in BLEU specifically around the constrained positions. Soumya Chatterjee 0002, Sunita Sarawagi, Preethi Jyothi |
ACL (1) | 3 |
| 2022 | Aligning Multilingual Embeddings for Improved Code-switched Natural Language UnderstandingabstractMultilingual pretrained models, while effective on monolingual data, need additional training to work well with code-switched text. In this work, we present a novel idea of training multilingual models with alignment objectives using parallel text so as to explicitly align word representations with the same underlying semantics across languages. Such an explicit alignment step has a positive downstream effect and improves performance on multiple code-switched NLP tasks. We explore two alignment strategies and report improvements of up to 7.32%, 0.76% and 1.9% on Hindi-English Sentiment Analysis, Named Entity Recognition and Question Answering tasks compared to a competitive baseline model. Barah Fazili, Preethi Jyothi |
COLING | 2 |
| 2022 | Zero-shot Disfluency Detection for Indian LanguagesabstractDisfluencies that appear in the transcriptions from automatic speech recognition systems tend to impair the performance of downstream NLP tasks. Disfluency correction models can help alleviate this problem. However, the unavailability of labeled data in low-resource languages impairs progress. We propose using a pretrained multilingual model, finetuned only on English disfluencies, for zero-shot disfluency detection in Indian languages. We present a detailed pipeline to synthetically generate disfluent text and create evaluation datasets for four Indian languages: Bengali, Hindi, Malayalam, and Marathi. Even in the zero-shot setting, we obtain F1 scores of 75 and higher on five disfluency types across all four languages. We also show the utility of synthetically generated disfluencies by evaluating on real disfluent text in Bengali, Hindi, and Marathi. Finetuning the multilingual model on additional synthetic Hindi disfluent text nearly doubles the number of exact matches and yields a 20-point boost in F1 scores when evaluated on real Hindi disfluent text, compared to training with only English disfluent text. Rohit Kundu, Preethi Jyothi, Pushpak Bhattacharyya |
COLING | 2 |
| 2022 | CoCoa: An Encoder-Decoder Model for Controllable Code-switched GenerationabstractCode-switching has seen growing interest in recent years as an important multilingual NLP phenomenon.Generating code-switched text for data augmentation has been sufficiently well-explored.However, there is no prior work on generating code-switched text with fine-grained control on the degree of codeswitching and the lexical choices used to convey formality.We present COCOA, an encoder-decoder translation model that converts monolingual Hindi text to Hindi-English code-switched text with both encoder-side and decoder-side interventions to achieve finegrained controllable generation.COCOA can be invoked at test-time to synthesize codeswitched text that is simultaneously faithful to syntactic and lexical attributes relevant to code-switching.COCOA outputs were subjected to rigorous subjective and objective evaluations.Human evaluations establish that our outputs are of superior quality while being faithful to desired attributes.We show significantly improved BLEU scores when compared with human-generated code-switched references.Compared to competitive baselines, we show 10% reduction in perplexity on a language modeling task and also demonstrate clear improvements on a downstream code-switched sentiment analysis task. Sneha Mondal, Ritika, Shreya Pathak, Preethi Jyothi, Aravindan Raghuveer |
EMNLP | 4 |
| 2022 | Adaptive Discounting of Implicit Language Models in RNN-TransducersabstractRNN-Transducer (RNN-T) models have become synonymous with streaming end-to-end ASR systems. While they perform competitively on a number of evaluation categories, rare words pose a serious challenge to RNN-T models. One main reason for the degradation in performance on rare words is that the language model (LM) internal to RNN-Ts can be-come overconfident and lead to hallucinated predictions that are acoustically inconsistent with the underlying speech. To address this issue, we propose a lightweight adaptive LM dis-counting technique ADAPTLMD, that can be used with any RNN-T architecture without requiring any external resources or additional parameters. ADAPTLMD uses a two-pronged approach: 1. Randomly mask the prediction network output to encourage the RNN-T to not be overly reliant on it’s outputs. 2. Dynamically choose when to discount the implicit LM (ILM) based on rarity of recently predicted tokens and divergence between ILM and implicit acoustic model (IAM) scores. Comparing ADAPTLMD to a competitive RNN-T baseline, we obtain up to 4% and 14% relative reductions in overall WER and rare word PER, respectively, on a conversational, code-mixed Hindi-English ASR task. Vinit Unni, Shreya Khare, Ashish R. Mittal, Preethi Jyothi, Sunita Sarawagi, Samarth Bharadwaj |
ICASSP | 4 |
| 2022 | SPLICEOUT: A Simple and Efficient Audio Augmentation MethodabstractTime masking has become a de facto augmentation technique for speech and audio tasks, including automatic speech recognition (ASR) and audio classification, most notably as a part of SpecAugment.In this work, we propose SPLICEOUT, a simple modification to time masking which makes it computationally more efficient.SPLICEOUT performs comparably to (and sometimes outperforms) SpecAugment on a wide variety of speech and audio tasks, including ASR for seven different languages using varying amounts of training data, as well as on speech translation, sound and music classification, thus establishing itself as a broadly applicable audio augmentation method.SPLICEOUT also provides additional gains when used in conjunction with other augmentation techniques.Apart from the fully-supervised setting, we also demonstrate that SPLICEOUT can complement unsupervised representation learning with performance gains in the semi-supervised and self-supervised settings. Arjit Jain, Pranay Reddy Samala, Deepak Mittal, Preethi Jyothi, Maneesh Kumar Singh 0001 |
INTERSPEECH | 4 |
| 2022 | VAgyojaka: An Annotating and Post-Editing Tool for Automatic Speech Recognition
Rishabh Kumar, Devaraj Adiga, Mayank Kothyari, Jatin Dalal, Ganesh Ramakrishnan, Preethi Jyothi |
INTERSPEECH | 6 |
| 2022 | Linguistically Informed Post-processing for ASR Error correction in Sanskrit
Rishabh Kumar, Devaraj Adiga, Rishav Ranjan, Amrith Krishna, Ganesh Ramakrishnan, Pawan Goyal 0002, Preethi Jyothi |
INTERSPEECH | 7 |
| 2021 | From Machine Translation to Code-Switching: Generating High-Quality Code-Switched TextabstractIshan Tarunesh, Syamantak Kumar, Preethi Jyothi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ishan Tarunesh, Syamantak Kumar, Preethi Jyothi |
ACL/IJCNLP (1) | 3 |
| 2021 | Disfluency Correction using Unsupervised and Semi-supervised LearningabstractNikhil Saini, Drumil Trivedi, Shreya Khare, Tejas Dhamecha, Preethi Jyothi, Samarth Bharadwaj, Pushpak Bhattacharyya. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Nikhil Saini, Drumil Trivedi, Shreya Khare, Tejas I. Dhamecha, Preethi Jyothi, Samarth Bharadwaj, Pushpak Bhattacharyya |
EACL | 5 |
| 2021 | Meta-Learning for Effective Multi-task and Multilingual ModellingabstractIshan Tarunesh, Sushil Khyalia, Vishwajeet Kumar, Ganesh Ramakrishnan, Preethi Jyothi. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Ishan Tarunesh, Sushil Khyalia, Vishwajeet Kumar, Ganesh Ramakrishnan, Preethi Jyothi |
EACL | 5 |
| 2021 | Error-Driven Fixed-Budget ASR Personalization for Accented SpeakersabstractWe consider the task of personalizing ASR models while being constrained by a fixed budget on recording speaker specific utterances. Given a speaker and an ASR model, we propose a method of identifying sentences for which the speaker’s utterances are likely to be harder for the given ASR model to recognize. We assume a tiny amount of speaker-specific data to learn phoneme-level error models which help us select such sentences. We show that speaker’s utterances on the sentences selected using our error model indeed have larger error rates when compared to speaker’s utterances on randomly selected sentences. We find that fine-tuning the ASR model on the sentence utterances selected with the help of error models yield higher WER improvements in comparison to fine-tuning on an equal number of randomly selected sentence utterances. Thus, our method provides an efficient way of collecting speaker utterances under budget constraints for personalizing ASR models. Abhijeet Awasthi, Aman Kansal, Sunita Sarawagi, Preethi Jyothi |
ICASSP | 4 |
| 2021 | Collaborative Learning to Generate Audio-Video JointlyabstractThere have been a number of techniques that have demonstrated the generation of multimedia data for one modality at a time using GANs, such as the ability to generate images, videos, and audio. However, so far, the task of multi-modal generation of data, specifically for audio and videos both, has not been sufficiently well-explored. Towards this, we propose a method that demonstrates that we are able to generate naturalistic samples of video and audio data by the joint correlated generation of audio and video modalities. The proposed method uses multiple discriminators to ensure that the audio, video, and the joint output are also indistinguishable from real-world samples. We present a dataset for this task and show that we are able to generate realistic samples. This method is validated using various standard metrics such as Inception Score, Frechet Inception Distance (FID) and through human evaluation. Vinod K. Kurmi, Vipul Bajaj, Badri Narayana Patro, K. S. Venkatesh, Vinay P. Namboodiri, Preethi Jyothi |
ICASSP | 6 |
| 2021 | An Investigation of End-to-End Models for Robust Speech RecognitionabstractEnd-to-end models for robust automatic speech recognition (ASR) have not been sufficiently well-explored in prior work. With end-to-end models, one could choose to preprocess the input speech using speech enhancement techniques and train the model using enhanced speech. Another alternative is to pass the noisy speech as input and modify the model architecture to adapt to noisy speech. A systematic comparison of these two approaches for end-to-end robust ASR has not been attempted before. We address this gap and present a detailed comparison of speech enhancement-based techniques and three different model-based adaptation techniques covering data augmentation, multi-task learning, and adversarial learning for robust ASR. While adversarial learning is the best-performing technique on certain noise types, it comes at the cost of degrading clean speech WER. On other relatively stationary noise types, a new speech enhancement technique outperformed all the model-based adaptation techniques. This suggests that knowledge of the underlying noise type can meaningfully inform the choice of adaptation technique. Archiki Prasad, Preethi Jyothi, Rajbabu Velmurugan |
ICASSP | 2 |
| 2021 | Cross Lingual Video and Text Retrieval: A New Benchmark Dataset and Algorithm
Jayaprakash Akula, Abhishek Sharma 0010, Rishabh Dabral, Preethi Jyothi, Ganesh Ramakrishnan |
ICMI | 4 |
| 2021 | Perturb, Predict & Paraphrase: Semi-Supervised Learning using Noisy Student for Image CaptioningabstractRecent semi-supervised learning (SSL) methods are predominantly focused on multi-class classification tasks. Classification tasks allow for easy mixing of class labels during augmentation which does not trivially extend to structured outputs such as word sequences that appear in tasks like image captioning. Noisy Student Training is a recent SSL paradigm proposed for image classification that is an extension of self-training and teacher-student learning. In this work, we provide an in-depth analysis of the noisy student SSL framework for the task of image captioning and derive state-of-the-art results. The original algorithm relies on computationally expensive data augmentation steps that involve perturbing the raw images and computing features for each perturbed image. We show that, even in the absence of raw image augmentation, the use of simple model and feature perturbations to the input images for the student model are beneficial to SSL training. We also show how a paraphrase generator could be effectively used for label augmentation to improve the quality of pseudo labels and significantly improve performance. Our final results in the limited labeled data setting (1% of the MS-COCO labeled data) outperform previous state-of-the-art approaches by 2.5 on BLEU4 and 11.5 on CIDEr scores. Arjit Jain, Pranay Reddy Samala, Preethi Jyothi, Deepak Mittal, Maneesh Kumar Singh 0001 |
IJCAI | 3 |
| 2021 | Reduce and Reconstruct: ASR for Low-Resource Phonetic LanguagesabstractThis work presents a seemingly simple but effective technique to improve low-resource ASR systems for phonetic languages. By identifying sets of acoustically similar graphemes in these languages, we first reduce the output alphabet of the ASR system using linguistically meaningful reductions and then reconstruct the original alphabet using a standalone module. We demonstrate that this lessens the burden and improves the performance of low-resource end-to-end ASR systems (because only reduced-alphabet predictions are needed) and that it is possible to design a very simple but effective reconstruction module that recovers sequences in the original alphabet from sequences in the reduced alphabet. We present a finite state transducer-based reconstruction module that operates on the 1-best ASR hypothesis in the reduced alphabet. We demonstrate the efficacy of our proposed technique using ASR systems for two Indian languages, Gujarati and Telugu. With access to only 10 hrs of speech data, we obtain relative WER reductions of up to 7% compared to systems that do not use any reduction. Anuj Diwan, Preethi Jyothi |
Interspeech | 2 |
| 2021 | MUCS 2021: Multilingual and Code-Switching ASR Challenges for Low Resource Indian LanguagesabstractRecently, there is increasing interest in multilingual automatic speech recognition (ASR) where a speech recognition system caters to multiple low resource languages by taking advantage of low amounts of labeled corpora in multiple languages. With multilingualism becoming common in today's world, there has been increasing interest in code-switching ASR as well. In code-switching, multiple languages are freely interchanged within a single sentence or between sentences. The success of low-resource multilingual and code-switching ASR often depends on the variety of languages in terms of their acoustics, linguistic characteristics as well as the amount of data available and how these are carefully considered in building the ASR system. In this challenge, we would like to focus on building multilingual and code-switching ASR systems through two different subtasks related to a total of seven Indian languages, namely Hindi, Marathi, Odia, Tamil, Telugu, Gujarati and Bengali. For this purpose, we provide a total of ~600 hours of transcribed speech data, comprising train and test sets, in these languages including two code-switched language pairs, Hindi-English and Bengali-English. We also provide a baseline recipe for both the tasks with a WER of 30.73% and 32.45% on the test sets of multilingual and code-switching subtasks, respectively. Anuj Diwan, Rakesh Vaideeswaran, Sanket Shah, Ankita Singh, Srinivasa Raghavan K. M., Shreya Khare, Vinit Unni, Saurabh Vyas, Akash Rajpuria, Chiranjeevi Yarra, Ashish R. Mittal, Prasanta Kumar Ghosh, Preethi Jyothi, Kalika Bali, Vivek Seshadri, Sunayana Sitaram, Samarth Bharadwaj, Jai Nanavati, Raoul Nanavati, Karthik Sankaranarayanan |
Interspeech | 13 |
| 2021 | Low Resource ASR: The Surprising Effectiveness of High Resource Transliteration
Shreya Khare, Ashish R. Mittal, Anuj Diwan, Sunita Sarawagi, Preethi Jyothi, Samarth Bharadwaj |
Interspeech | 5 |
| 2021 | Cross-Modal Learning for Audio-Visual Video ParsingabstractIn this paper, we present a novel approach to the audio-visual video parsing (AVVP) task that demarcates events from a video separately for audio and visual modalities. The proposed parsing approach simultaneously detects the temporal boundaries in terms of start and end times of such events. We show how AVVP can benefit from the following techniques geared towards effective cross-modal learning: (i) adversarial training and skip connections (ii) global context aware attention and, (iii) self-supervised pretraining using an audio-video grounding objective to obtain cross-modal audio-video representations. We present extensive experimental evaluations on the Look, Listen, and Parse (LLP) dataset and show that we outperform the state-of-the-art Hybrid Attention Network (HAN) on all five metrics proposed for AVVP. We also present several ablations to validate the effect of pretraining, global attention and adversarial training. Jatin Lamba, Jayaprakash Akula, Rishabh Dabral, Preethi Jyothi, Ganesh Ramakrishnan |
Interspeech | 5 |
| 2021 | Select, Substitute, Search: A New Benchmark for Knowledge-Augmented Visual Question AnsweringabstractMultimodal IR, spanning text corpus, knowledge graph and images, called outside knowledge visual question answering (OKVQA), is of much recent interest. However, the popular data set has serious limitations. A surprisingly large fraction of queries do not assess the ability to integrate cross-modal information. Instead, some are independent of the image, some depend on speculation, some require OCR or are otherwise answerable from the image alone. To add to the above limitations, frequency-based guessing is very effective because of (unintended) widespread answer overlaps between the train and test folds. Overall, it is hard to determine when state-of-the-art systems exploit these weaknesses rather than really infer the answers, because they are opaque and their 'reasoning' process is uninterpretable. An equally important limitation is that the dataset is designed for the quantitative assessment only of the end-to-end answer retrieval task, with no provision for assessing the correct(semantic) interpretation of the input query. In response, we identify a key structural idiom in OKVQA ,viz., S3 (select, substitute and search), and build a new data set and challenge around it. Specifically, the questioner identifies an entity in the image and asks a question involving that entity which can be answered only by consulting a knowledge graph or corpus passage mentioning the entity. Our challenge consists of (i)OKVQA_S3, a subset of OKVQA annotated based on the structural idiom and (ii)S3VQA, a new dataset built from scratch. We also present a neural but structurally transparent OKVQA system, S3, that explicitly addresses our challenge dataset, and outperforms recent competitive baselines. We make our code and data available at https://s3vqa.github.io/. Aman Jain, Mayank Kothyari, Vishwajeet Kumar, Preethi Jyothi, Ganesh Ramakrishnan, Soumen Chakrabarti |
SIGIR | 4 |
| 2020 | How Accents Confound: Probing for Accent Information in End-to-End Speech Recognition SystemsabstractIn this work, we present a detailed analysis of how accent information is reflected in the internal representation of speech in an end-to-end automatic speech recognition (ASR) system. We use a state-of-the-art end-to-end ASR system, comprising convolutional and recurrent layers, that is trained on a large amount of US-accented English speech and evaluate the model on speech samples from seven different English accents. We examine the effects of accent on the internal representation using three main probing techniques: a) Gradient-based explanation methods, b) Information-theoretic measures, and c) Outputs of accent and phone classifiers. We find different accents exhibiting similar trends irrespective of the probing technique used. We also find that most accent information is encoded within the first recurrent layer, which is suggestive of how one could adapt such an end-to-end model to learn representations that are invariant to accents. Archiki Prasad, Preethi Jyothi |
ACL | 2 |
| 2020 | Coupled Training of Sequence-to-Sequence Models for Accented Speech RecognitionabstractAccented speech poses significant challenges for state-of-the-art automatic speech recognition (ASR) systems. Accent is a property of speech that lasts throughout an utterance in varying degrees of strength. This makes it hard to isolate the influence of accent on individual speech sounds. We propose coupled training for encoder-decoder ASR models that acts on pairs of utterances corresponding to the same text spoken by speakers with different accents. This training regime introduces an L2 loss between the attention-weighted representations corresponding to pairs of utterances with the same text, thus acting as a regularizer and encouraging representations from the encoder to be more accent-invariant. We focus on recognizing accented English samples from the Mozilla Common Voice corpus. We obtain significant error rate reductions on accented samples from a large set of diverse accents using coupled training. We also show consistent improvements in performance on heavily accented samples (as determined by a standalone accent classifier). Vinit Unni, Nitish Joshi, Preethi Jyothi |
ICASSP | 3 |
| 2020 | Black-Box Adaptation of ASR for Accented SpeechabstractWe introduce the problem of adapting a black-box, cloud-based ASR system to speech from a target accent. While leading online ASR services obtain impressive performance on main-stream accents, they perform poorly on sub-populations - we observed that the word error rate (WER) achieved by Google's ASR API on Indian accents is almost twice the WER on US accents. Existing adaptation methods either require access to model parameters or overlay an error-correcting module on output transcripts. We highlight the need for correlating outputs with the original speech to fix accent errors. Accordingly, we propose a novel coupling of an open-source accent-tuned local model with the black-box service where the output from the service guides frame-level inference in the local model. Our fine-grained merging algorithm is better at fixing accent errors than existing word-level combination strategies. Experiments on Indian and Australian accents with three leading ASR models as service, show that we achieve as much as 28% relative reduction in WER over both the local and service models. Kartik Khandelwal, Preethi Jyothi, Abhijeet Awasthi, Sunita Sarawagi |
INTERSPEECH | 2 |
| 2020 | Caption Alignment for Low Resource Audio-Visual DataabstractUnderstanding videos via captioning has gained a lot of traction recently. While captions are provided alongside videos, the information about where a caption aligns within a video is missing, which could be particularly useful for indexing and retrieval. Existing work on learning to infer alignments has mostly exploited visual features and ignored the audio signal. Video understanding applications often underestimate the importance of the audio modality. We focus on how to make effective use of the audio modality for temporal localization of captions within videos. We release a new audio-visual dataset that has captions time-aligned by (i) carefully listening to the audio and watching the video, and (ii) watching only the video. Our dataset is audio-rich and contains captions in two languages, English and Marathi (a low-resource language). We further propose an attention-driven multimodal model, for effective utilization of both audio and video for temporal localization. We then investigate (i) the effects of audio in both data preparation and model design, and (ii) effective pretraining strategies (Audioset, ASR-bottleneck features, PASE, etc.) handling low-resource setting to help extract rich audio representations. Vighnesh Reddy Konda, Mayur Warialani, Rakesh Prasanth Achari, Varad Bhatnagar, Jayaprakash Akula, Preethi Jyothi, Ganesh Ramakrishnan, Gholamreza Haffari, Pankaj Singh |
INTERSPEECH | 6 |
| 2020 | Improving Low Resource Code-Switched ASR Using Augmented Code-Switched TTSabstractBuilding Automatic Speech Recognition (ASR) systems for code-switched speech has recently gained renewed attention due to the widespread use of speech technologies in multilingual communities worldwide. End-to-end ASR systems are a natural modeling choice due to their ease of use and superior performance in monolingual settings. However, it is well known that end-to-end systems require large amounts of labeled speech. In this work, we investigate improving code-switched ASR in low resource settings via data augmentation using code-switched text-to-speech (TTS) synthesis. We propose two targeted techniques to effectively leverage TTS speech samples: 1) Mixup, an existing technique to create new training samples via linear interpolation of existing samples, applied to TTS and real speech samples, and 2) a new loss function, used in conjunction with TTS samples, to encourage code-switched predictions. We report significant improvements in ASR performance achieving absolute word error rate (WER) reductions of up to 5%, and measurable improvement in code switching using our proposed techniques on a Hindi-English code-switched ASR task. Yash Sharma 0004, Basil Abraham, Karan Taneja, Preethi Jyothi |
INTERSPEECH | 4 |
| 2020 | Crowdsourcing Speech Data for Low-Resource Languages from Low-Income WorkersabstractVoice-based technologies are essential to cater to the hundreds of millions of new smartphone users. However, most of the languages spoken by these new users have little to no labelled speech data. Unfortunately, collecting labelled speech data in any language is an expensive and resource-intensive task. Moreover, existing platforms typically collect speech data only from urban speakers familiar with digital technology whose dialects are often very different from low-income users. In this paper, we explore the possibility of collecting labelled speech data directly from low-income workers. In addition to providing diversity to the speech dataset, we believe this approach can also provide valuable supplemental earning opportunities to these communities. To this end, we conducted a study where we collected labelled speech data in the Marathi language from three different user groups: low-income rural users, low-income urban users, and university students. Overall, we collected 109 hours of data from 36 participants. Our results show that the data collected from low-income participants is of comparable quality to the data collected from university students (who are typically employed to do this work) and that crowdsourcing speech data from low-income rural and urban workers is a viable method of gathering speech data. Basil Abraham, Danish Goel, Divya Siddarth, Kalika Bali, Manu Chopra, Monojit Choudhury, Pratik Joshi, Preethi Jyothi, Sunayana Sitaram, Vivek Seshadri |
LREC | 8 |
| 2019 | Cross-Lingual Training for Automatic Question GenerationabstractAutomatic question generation (QG) is a challenging problem in natural language understanding.QG systems are typically built assuming access to a large number of training instances where each instance is a question and its corresponding answer.For a new language, such training instances are hard to obtain making the QG problem even more challenging.Using this as our motivation, we study the reuse of an available large QG dataset in a secondary language (e.g.English) to learn a QG model for a primary language (e.g.Hindi) of interest.For the primary language, we assume access to a large amount of monolingual text but only a small QG dataset.We propose a cross-lingual QG model which uses the following training regime: (i) Unsupervised pretraining of language models in both primary and secondary languages and (ii) joint supervised training for QG in both languages.We demonstrate the efficacy of our proposed approach using two different primary languages, Hindi and Chinese.We also create and release a new question answering dataset for Hindi consisting of 6555 sentences. Vishwajeet Kumar, Nitish Joshi, Arijit Mukherjee, Ganesh Ramakrishnan, Preethi Jyothi |
ACL (1) | 5 |
| 2019 | Exploiting Monolingual Speech Corpora for Code-Mixed Speech Recognition
Karan Taneja, Satarupa Guha, Preethi Jyothi, Basil Abraham |
INTERSPEECH | 3 |
| 2019 | Improved feature selection and classification for rheumatoid arthritis disease using weighted decision tree approach (REACT)
Siva Shanmugam, Preethi Jyothi |
J. Supercomput. | 2 |
| 2018 | Code-switched Language Models Using Dual RNNs and Same-Source PretrainingabstractThis work focuses on building language models (LMs) for code-switched text.We propose two techniques that significantly improve these LMs: 1) A novel recurrent neural network unit with dual components that focus on each language in the code-switched text separately 2) Pretraining the LM using synthetic text from a generative model estimated using the training data.We demonstrate the effectiveness of our proposed techniques by reporting perplexities on a Mandarin-English task and derive significant reductions in perplexity. Saurabh Garg 0003, Tanmay Parekh, Preethi Jyothi |
EMNLP | 3 |
| 2018 | Revisiting the Importance of Encoding Logic Rules in Sentiment ClassificationabstractWe analyze the performance of different sentiment classification models on syntacticallycomplex inputs like A-but-B sentences.The first contribution of this analysis addresses reproducible research: to meaningfully compare different models, their accuracies must be averaged over far more random seeds than what has traditionally been reported.With proper averaging in place, we notice that the distillation model described in Hu et al. (2016), which incorporates explicit logic rules for sentiment classification, is ineffective.In contrast, using contextualized ELMo embeddings (Peters et al., 2018a) instead of logic rules yields significantly better performance.Additionally, we provide analysis and visualizations that demonstrate ELMo's ability to implicitly learn logic rules.Finally, a crowdsourced analysis reveals how ELMo outperforms baseline models even on sentences with ambiguous sentiment labels. Kalpesh Krishna, Preethi Jyothi, Mohit Iyyer |
EMNLP | 2 |
| 2018 | Generalizing Across Domains via Cross-Gradient Training
Shiv Shankar, Vihari Piratla, Soumen Chakrabarti, Siddhartha Chaudhuri, Preethi Jyothi, Sunita Sarawagi |
ICLR (Poster) | 5 |
| 2018 | Dual Language Models for Code Switched Speech RecognitionabstractIn this work, we present a simple and elegant approach to language modeling for bilingual code-switched text.Since codeswitching is a blend of two or more different languages, a standard bilingual language model can be improved upon by using structures of the monolingual language models.We propose a novel technique called dual language models, which involves building two complementary monolingual language models and combining them using a probabilistic model for switching between the two.We evaluate the efficacy of our approach using a conversational Mandarin-English speech corpus.We prove the robustness of our model by showing significant improvements in perplexity measures over the standard bilingual language model without the use of any external information.Similar consistent improvements are also reflected in automatic speech recognition error rates. Saurabh Garg 0003, Tanmay Parekh, Preethi Jyothi |
INTERSPEECH | 3 |
| 2018 | Improved Accented Speech Recognition Using Accent Embeddings and Multi-task Learning
Minali Upreti, Preethi Jyothi |
INTERSPEECH | 3 |
| 2018 | Time Aggregation Operators for Multi-label Audio Event Detection
Pankaj Joshi, Digvijaysingh Gautam, Ganesh Ramakrishnan, Preethi Jyothi |
INTERSPEECH | 4 |
| 2018 | Synthesizing Audio for Hindi WordNetabstractIn this paper, we describe our work on the creation of a voice model using a speech synthesis system for the Hindi Language.We use preexisting "voices", use publicly available speech corpora to create a "voice" using the Festival Speech Synthesis System (Black, 1997).Our contribution is two-fold: (1) We scrutinize multiple speech synthesis systems and provide an extensive report on the currently available stateof-the-art systems.We also develop voices using the existing implementations of the aforementioned systems, and (2) We use these voices to generate sample audios for randomly chosen words; manually evaluate the audio generated, and produce audio for all WordNet words using the winner voice model.We also produce audios for the Hindi WordNet Glosses and Example sentences.We describe our efforts to use preexisting implementations for WaveNet -a model to generate raw audio using neural nets (Oord et al., 2016) and generate speech for Hindi.Our lexicographers perform a manual evaluation of the audio generated using multiple voices.A qualitative and quantitative analysis reveals that the voice model generated by us performs the best with an accuracy of 0.44. Diptesh Kanojia, Preethi Jyothi, Pushpak Bhattacharyya |
GWC | 2 |
| 2018 | Hindi Wordnet for Language Teaching: Experiences and Lessons LearntabstractHanumant Redkar, Rajita Shukla, Sandhya Singh, Jaya Saraswati, Laxmi Kashyap, Diptesh Kanojia, Preethi Jyothi, Malhar Kulkarni, Pushpak Bhattacharyya. Proceedings of the 9th Global Wordnet Conference. 2018. Hanumant Harichandra Redkar, Rajita Shukla, Sandhya Singh, Jaya Saraswati, Laxmi Kashyap, Diptesh Kanojia, Preethi Jyothi, Malhar Kulkarni, Pushpak Bhattacharyya |
GWC | 7 |
| 2017 | Leveraging native language speech for accent identification using deep Siamese networksabstractThe problem of automatic accent identification is important for several applications like speaker profiling and recognition as well as for improving speech recognition systems. The accented nature of speech can be primarily attributed to the influence of the speaker's native language on the given speech recording. In this paper, we propose a novel accent identification system whose training exploits speech in native languages along with the accented speech. Specifically, we develop a deep Siamese network based model which learns the association between accented speech recordings and the native language speech recordings. The Siamese networks are trained with i-vector features extracted from the speech recordings using either an unsupervised Gaussian mixture model (GMM) or a supervised deep neural network (DNN) model. We perform several accent identification experiments using the CSLU Foreign Accented English (FAE) corpus. In these experiments, our proposed approach using deep Siamese networks yield significant relative performance improvements of 15.4% on a 10-class accent identification task, over a baseline DNN-based classification system that uses GMM i-vectors. Furthermore, we present a detailed error analysis of the proposed accent identification system. Aditya Siddhant, Preethi Jyothi, Sriram Ganapathy |
ASRU | 2 |
| 2017 | Low-resource grapheme-to-phoneme conversion using recurrent neural networksabstractGrapheme-to-phoneme (G2P) conversion is an important problem for many speech and language processing applications. G2P models are particularly useful for low-resource languages that do not have well-developed pronunciation lexicons. Prominent G2P paradigms are based on initial alignments between grapheme and phoneme sequences. In this work, we devise new alignment strategies that work effectively with recurrent neural network based models when only a small number of pronunciations are available to train the models. In a small data setting, we build G2P models for Pashto, Tagalog and Lithuanian that significantly outperform a joint sequence model and a baseline recurrent neural network based model, giving up to 14% and 9% relative reductions in phone and word error rates when trained on a dataset of 250 words. Preethi Jyothi, Mark Hasegawa-Johnson |
ICASSP | 1 |
| 2017 | ASR for Under-Resourced Languages From Probabilistic TranscriptionabstractIn many under-resourced languages it is possible to find text, and it is possible to find speech, but transcribed speech suitable for training automatic speech recognition (ASR) is unavailable. In the absence of native transcripts, this paper proposes the use of a probabilistic transcript: A probability mass function over possible phonetic transcripts of the waveform. Three sources of probabilistic transcripts are demonstrated. First, self-training is a well-established semisupervised learning technique, in which a cross-lingual ASR first labels unlabeled speech, and is then adapted using the same labels. Second, mismatched crowdsourcing is a recent technique in which nonspeakers of the language are asked to write what they hear, and their nonsense transcripts are decoded using noisy channel models of second-language speech perception. Third, EEG distribution coding is a new technique in which nonspeakers of the language listen to it, and their electrocortical response signals are interpreted to indicate probabilities. ASR was trained in four languages without native transcripts. Adaptation using mismatched crowdsourcing significantly outperformed self-training, and both significantly outperformed a cross-lingual baseline. Both EEG distribution coding and text-derived phone language models were shown to improve the quality of probabilistic transcripts derived from mismatched crowdsourcing. Mark Hasegawa-Johnson, Preethi Jyothi, Daniel McCloy, Majid Mirbagheri, Giovanni M. Di Liberto, Amit Das 0007, Bradley Ekin, Chunxi Liu, Vimal Manohar, Hao Tang 0002, Edmund C. Lalor, Nancy F. Chen, Paul Hager, Tyler Kekona, Rose Sloan, Adrian K. C. Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Adapting ASR for under-resourced languages using mismatched transcriptionsabstractMismatched transcriptions of speech in a target language refers to transcriptions provided by people unfamiliar with the language, using English letter sequences. In this work, we demonstrate the value of such transcriptions in building an ASR system for the target language. For different languages, we use less than an hour of mismatched transcriptions to successfully adapt baseline multilingual models built with no access to native transcriptions in the target language. The adapted models provide up to 25% relative improvement in phone error rates on an unseen evaluation set. Chunxi Liu, Preethi Jyothi, Hao Tang 0002, Vimal Manohar, Rose Sloan, Tyler Kekona, Mark Hasegawa-Johnson, Sanjeev Khudanpur |
ICASSP | 2 |
| 2016 | Automatic Speech Recognition Using Probabilistic Transcriptions in Swahili, Amharic, and Dinka
Amit Das 0007, Preethi Jyothi, Mark Hasegawa-Johnson |
INTERSPEECH | 2 |
| 2016 | Articulatory feature-based pronunciation modeling
Karen Livescu, Preethi Jyothi, Eric Fosler-Lussier |
Comput. Speech Lang. | 2 |
| 2015 | Acquiring Speech Transcriptions Using Mismatched CrowdsourcingabstractTranscribed speech is a critical resource for building statistical speech recognition systems. Recent work has looked towards soliciting transcriptions for large speech corpora from native speakers of the language using crowdsourcing techniques. However, native speakers of the target language may not be readily available for crowdsourcing. We examine the following question: can humans unfamiliar with the target language help transcribe? We follow an information-theoretic approach to this problem: (1) We learn the characteristics of a noisy channel that models the transcribers' systematic perception biases. (2) We use an error-correcting code, specifically a repetition code, to encode the inputs to this channel, in conjunction with a maximum-likelihood decoding rule. To demonstrate the feasibility of this approach, we transcribe isolated Hindi words with the help of Mechanical Turk workers unfamiliar with Hindi. We successfully recover Hindi words with an accuracy of over 85% (and 94% in a 4-best list) using a 15-fold repetition code. We also estimate the conditional entropy of the input to this channel (Hindi words) given the channel output (transcripts from crowdsourced workers) to be less than 2 bits; this serves as a theoretical estimate of the average number of bits of auxiliary information required for errorless recovery. Preethi Jyothi, Mark Hasegawa-Johnson |
AAAI | 1 |
| 2015 | Transcribing continuous speech using mismatched crowdsourcingabstractMismatched crowdsourcing was recently proposed as a poten-tial approach to deriving moderately accurate speech transcrip-tions using crowd workers unfamiliar with the language be-ing spoken. In introducing this approach, we demonstrated its promise with the help of an isolated word recovery task for Hindi. However, it remained open whether mismatched crowd sourcing can yield non-trivial accuracy in a continuous speech task. In this work, we focus on this question and demonstrate a word error rate of under 45 % in a large-vocabulary task (again for Hindi). In achieving this, we develop several new tech-niques capable of scaling effectively to continuous speech. We also provide an information theoretic analysis and estimate the amount of information lost in transcription by the mismatched crowd workers to be under 5 bits. Preethi Jyothi, Mark Hasegawa-Johnson |
INTERSPEECH | 1 |
| 2015 | Improved hindi broadcast ASR by adapting the language model and pronunciation model using a priori syntactic and morphophonemic knowledgeabstractIn this work, we present a new large-vocabulary, broadcast news ASR system for Hindi. Since Hindi has a largely phone-mic orthography, the pronunciation model was automatically generated from text. We experiment with several variants of this model and study the effect of incorporating word bound-ary information with these models. We also experiment with knowledge-based adaptations to the language model in Hindi, derived in an unsupervised manner, that lead to small im-provements in word error rate (WER). Our experiments were conducted on a new corpus assembled from publicly-available Hindi news broadcasts. We evaluate our techniques on an open-vocabulary task and obtain competitive WERs on an unseen test set. Index Terms: Hindi LVCSR system, Broadcast news ASR, Grapheme and phoneme-based models, Knowledge-based language-model adaptation 1. Preethi Jyothi, Mark Hasegawa-Johnson |
INTERSPEECH | 1 |
| 2013 | Discriminative training of WFST factors with application to pronunciation modelingabstractOne of the most popular speech recognition architectures consists of multiple components (like the acoustic, pronunciation and language models) that are modeled as weighted finite state transducer (WFST) factors in a cascade. These factor WFSTs are typically trained in isolation and combined efficiently for decoding. Recent work has explored jointly estimating parameters for these models using considerable amounts of training data. We propose an alternative approach to selectively train factor WFSTs in such an architecture, while still leveraging information from the entire cascade. This technique allows us to effectively estimate parameters of a factor WFST using relatively small amounts of data, if the factor is small. Our approach involves an online training paradigm for linear models adapted for discriminatively training one or more WFSTs in a cascade. We apply this method to train a pronunciation model for recognition on conversational speech, resulting in significant improvements in recognition performance over the baseline model. Index Terms: Pronunciation models, weighted finite state transducers, large-margin training. Preethi Jyothi, Eric Fosler-Lussier, Karen Livescu |
INTERSPEECH | 1 |
| 2013 | Conditional Random Fields in Speech, Audio, and Language ProcessingabstractConditional random fields (CRFs) are probabilistic sequence models that have been applied in the last decade to a number of applications in audio, speech, and language processing. In this paper, we provide a tutorial overview of CRF technologies, pointing to other resources for more in-depth discussion; in particular, we describe the common linear-chain model as well as a number of common extensions within the CRF family of models. An overview of the mathematical techniques used in training and evaluating these models is also provided, as well as a discussion of the relationships with other probabilistic models. Finally, we survey recent work in speech, audio, and language processing to show how the same CRF technology can be deployed in different scenarios. Eric Fosler-Lussier, Yanzhang He, Preethi Jyothi, Rohit Prabhavalkar |
Proc. IEEE | 3 |
| 2012 | Distributed discriminative language models for Google voice-searchabstractThis paper considers large-scale linear discriminative language models trained using a distributed perceptron algorithm. The algorithm is implemented efficiently using a MapReduce/SSTable framework. This work also introduces the use of large amounts of unsupervised data (confidence filtered Google voice-search logs) in conjunction with a novel training procedure that regenerates word lattices for the given data with a weaker acoustic model than the one used to generate the unsupervised transcriptions for the logged data. We observe small but statistically significant improvements in recognition performance after reranking N-best lists of a standard Google voice-search data set. Preethi Jyothi, Leif Johnson, Ciprian Chelba, Brian Strope |
ICASSP | 1 |
| 2012 | Discriminatively learning factorized finite state pronunciation models from dynamic Bayesian networksabstractThis paper describes an approach to efficiently construct, and discriminatively train, a weighted finite state transducer (WFST) representation for an articulatory feature-based model of pronunciation. This model is originally implemented as a dynamic Bayesian network (DBN). The work is motivated by a desire to (1) incorporate such a pronunciation model in WFSTbased recognizers, and to (2) learn discriminative models that are more general than the DBNs. The approach is quite general, though here we show how it applies to a specific model. We use the conditional independence assumptions imposed by the DBN to efficiently convert it into a sequence of WFSTs (factor FSTs) which, when composed, yield the same model as the DBN. We then introduce a linear model of the arc weights of the factor FSTs and discriminatively learn its weights using the averaged perceptron algorithm. We demonstrate the approach using a lexical access task in which we recognize a word given its surface realization. Our experimental results using a phonetically transcribed subset of the Switchboard corpus show that the discriminatively learned model performs significantly better than the original DBN. Preethi Jyothi, Eric Fosler-Lussier, Karen Livescu |
INTERSPEECH | 1 |
| 2011 | Lexical access experiments with context-dependent articulatory feature-based modelsabstractWe address the problem of pronunciation variation in conversational speech with a context-dependent articulatory feature-based model. The model is an extension of previous work using dynamic Bayesian networks, which allow for easy factorization of a state into multiple variables representing the articulatory features. We build context-dependent decision trees for the articulatory feature distributions, which are incorporated into the dynamic Bayesian networks, and experiment with different sets of context variables. We evaluate our models on a lexical access task using a phonetically transcribed subset of the Switchboard corpus. We find that our models outperform a context-dependent phonetic baseline. Preethi Jyothi, Karen Livescu, Eric Fosler-Lussier |
ICASSP | 1 |
| 2010 | Discriminative language modeling using simulated ASR errorsabstractIn this paper, we approach the problem of discriminatively training language models using a weighted finite state transducer (WFST) framework that does not require acoustic training data. The phonetic confusions prevalent in the recognizer are modeled using a confusion matrix that takes into account information from the pronunciation model (word-based phone confusion log likelihoods) and information from the acoustic model (distances between the phonetic acoustic models). This confusion matrix, within the WFST framework, is used to generate confusable word graphs that serve as inputs to the averaged perceptron algorithm to train the parameters of the discriminative language model. Experiments on a large vocabulary speech recognition task show significant word error rate reductions when compared to a baseline using a trigram model trained with the maximum likelihood criterion. Preethi Jyothi, Eric Fosler-Lussier |
INTERSPEECH | 1 |
| 2010 | Investigations into the Crandem Approach to Word Recognition
Rohit Prabhavalkar, Preethi Jyothi, William Hartmann, Jeremy Morris, Eric Fosler-Lussier |
HLT-NAACL | 2 |
| 2009 | A comparison of audio-free speech recognition error prediction methodsabstractPredicting possible speech recognition errors can be invaluable for a number of Automatic Speech Recognition (ASR) applications. In this study, we extend a Weighted Finite State Transducer (WFST) framework for error prediction to facilitate a comparison between two approaches of predicting confusable words: examining recognition errors on the training set to learn phone confusions and utilizing distances between the phonetic acoustic models for the prediction task. We also expand the framework to deal with continuous word recognition and we can accurately predict 60% of the misrecognized sentences (with an average words-per-sentence count of 15) and a little over 70% of the total number of errors from the unseen test data where no acoustic information related to the test data is utilized. Index Terms: Finite State Transducer, Automatic Speech Recognition, Error prediction Preethi Jyothi, Eric Fosler-Lussier |
INTERSPEECH | 1 |