VLDB 2026 Research / reviewers in the wild / expert
Deekshitha G
dblp:330/9917
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MADASR 2.0: Multi-Lingual Multi-Dialect ASR Challenge in 8 Indian LanguagesabstractWe present MADASR 2.0, a challenge at ASRU 2025 aimed at advancing multilingual and multidialectal automatic speech recognition (ASR) in low-resource Indian languages. Building on the 2023 edition, it introduces a subset of the RESPIN corpus, over 1200 hours of read speech across 8 languages and 33 dialects, with test sets including both read and spontaneous speech. The challenge comprises four tracks varying by training data size and external resource usage, and supports auxiliary tasks like language and dialect identification. We detail the dataset, tasks, baselines, and submissions and analyse trends across tracks and speech styles. Results highlight the continued difficulty of spontaneous ASR, the benefits of multitask and transfer learning, and effective strategies for building dialect-aware ASR systems. MADASR 2.0 offers a standardised benchmark to support future research on inclusive and scalable ASR for linguistically diverse populations. Sumit Sharma 0016, Deekshitha G, Abhayjeet Singh, Amartyaveer, Sathvik Udupa, Sandhya Badiger, Sanjeev Khudanpur, Sunayana Sitaram, Srinivasan Umesh, Bhuvana Ramabhadran, Brian Kingsbury, Hema A. Murthy, Srikanth S. Narayanan, Howard Lakougna, Prasanta Kumar Ghosh |
ASRU | 3 |
| 2025 | Improving Dialect Identification in Indian Languages Using Multimodal Features from Dialect Informed ASRabstractDialect identification (DID) addresses the challenge of recog-nizing regional variations within a language. The current deep learning approaches focus on audio-only, text-only, or multi-task setups combining automatic speech recognition (ASR) with DID. This work introduces a novel multimodal architecture that leverages speech and text features to enhance DID performance. Our method integrates ASR-generated speech representations with text embeddings derived from ASR hypotheses using a RoBERTa-based encoder. Additionally, we perform a layer-wise analysis of the IndicWav2Vec model to identify the layers most effective for extracting dialectal features. We evaluate our approach on a subset of the RESPIN dataset featuring eight Indian languages and 33 dialects. Experimental results show that our proposed multimodal DID system achieves an average DID accuracy of 79.81%, consistently outperforming baseline methods. This study is the first to analyse comprehensively DID in Indian languages, providing new insights into their dialectal diversity. Amartyaveer, Sumit Sharma 0016, Sathvik Udupa, Sandhya Badiger, Abhayjeet Singh, Deekshitha G, Jesuraja Bandekar, Savitha Murthy, Prasanta Kumar Ghosh |
ICASSP | 7 |
| 2025 | Comparison of Acoustic and Textual Features for Dysarthria Severity Classification in Amyotrophic Lateral Sclerosis
Y. S. Upendra Vishwanath, Tanuka Bhattacharjee, Deekshitha G, Sathvik Udupa, Chowdam Venkata Thirumala Kumar, Madassu Keerthipriya, Darshan Chikktimmegowda, Dipti Baskar, Yamini Belur, Seena Vengalil, Atchayaram Nalini, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2025 | RESPIN-S1.0: A read speech corpus of 10000+ hours in dialects of nine Indian LanguagesabstractWe introduce RESPIN-S1.0, the largest publicly available dialect-rich read-speech corpus for Indian languages, comprising more than 10,000 hours of validated audio across nine major languages: Bengali, Bhojpuri, Chhattisgarhi, Hindi, Kannada, Magahi, Maithili, Marathi, and Telugu. Indian languages exhibit high dialectal variation and are spoken by populations that remain digitally underserved. Existing speech corpora typically represent only standard dialects and lack domain and linguistic diversity. RESPIN-S1.0 addresses this limitation by collecting speech across more than 38 dialects and two high-impact domains: agriculture and finance. Text data were composed by native dialect speakers and validated through a pipeline combining automated and manual checks. Over 200,000 unique sentences were recorded through a crowdsourced mobile platform and categorised into clean, semi-noisy, and noisy subsets based on transcription quality, with the clean portion alone exceeding 10,000 hours. Along with audio and transcriptions, RESPIN provides dialect-aware phonetic lexicons, speaker metadata, and reproducible train, development, and test splits. To benchmark performance, we evaluate multiple ASR models, including TDNN-HMM, E-Branchformer, Whisper, and wav2vec2-based self-supervised models, and find that fine-tuning on RESPIN significantly improves recognition accuracy over pretrained baselines. A subset of RESPIN-S1.0 has already supported community challenges such as the SLT Code Hackathon 2022 and MADASR@ASRU 2023 and 2025, releasing more than 1,200 hours publicly. This resource supports research in dialectal ASR, language identification, and related speech technologies, establishing a comprehensive benchmark for inclusive, dialect-rich ASR in multilingual low-resource settings. Dataset: https://spiredatasets.ee.iisc.ac.in/respincorpus Code: https://github.com/labspire/respin_baselines.git Abhayjeet Singh, Deekshitha G, Amartya Veer, Jesuraja Bandekar, Savitha Murthy, Sumit Sharma 0016, Sandhya Badiger, Sathvik Udupa, Amala Nagireddi, Srinivasa Raghavan K. M., Rohan Saxena, Jai Nanavati, Raoul Nanavati, Janani Sridharan, Arjun Singh Mehta, Ashish Seth, Sai Praneeth Reddy Mora, Prashanthi V, Gauri Date, Karthika P, Prasanta Kumar Ghosh |
NeurIPS | 3 |
| 2024 | Adapter pre-training for improved speech recognition in unseen domains using low resource adapter tuning of self-supervised models
Sathvik Udupa, Jesuraja Bandekar, Deekshitha G, Sandhya Badiger, Abhayjeet Singh Savitha Murthy, Priyanka Pai, Srinivasa Raghavan K. M., Raoul Nanavati, Prasanta Kumar Ghosh |
INTERSPEECH | 4 |
| 2023 | Gated Multi Encoders and Multitask Objectives for Dialectal Speech Recognition in Indian LanguagesabstractIn this work, several methods have been proposed towards improving the performance of dialectal automatic speech recognition (ASR). A novel encoder architecture has been introduced that is suited for multi-dialect ASR training. Further, we propose Multi-Task Self-Supervised learning (SSL) fine-tuning using CTC and dialect identification. Additionally, the use of different language models (LM) to improve the performance of dialectal ASR has been investigated. Around 800 hours of Bengali and Bhojpuri data, released as a part of the MADASR ASRU challenge have been used to train these models. The work shows that the proposed multi-encoder ASR observes a relative reduction of 7.5% and 9% in WER in Bhojpuri and Bengali, respectively. Additionally, we also observe a 1-2% WER reduction in fine-tuning SSL, further improving performance in these languages. Moreover, we observe advantages in using dialect-specific LM decoding based on predicted dialect. Sathvik Udupa, Jesuraja Bandekar, Deekshitha G, Prasanta Kumar Ghosh, Sandhya Badiger, Abhayjeet Singh, Savitha Murthy, Priyanka Pai, Srinivasa Raghavan K. M., Raoul Nanavati |
ASRU | 3 |
| 2023 | Lightweight, Multi-Speaker, Multi-Lingual Indic Text-to-SpeechabstractThe Lightweight, Multi-speaker, Multi-lingual Indic Text-to-Speech (LIMMITS’23) challenge is organized as part of the ICASSP 2023 signal processing grand challenge. LIMMITS’23 aims at the development of a lightweight, multi-speaker, multi-lingual Text to Speech (TTS) model using datasets in Marathi, Hindi, and Telugu. The challenge encourages the advancement of TTS in Indian Languages as well as the development of techniques involved in TTS data selection and model compression. The 3 tracks of LIMMITS’23 have provided an opportunity for various researchers and practitioners around the world to explore the state of the art in TTS research. Abhayjeet Singh, Amala Nagireddi, Deekshitha G, Jesuraja Bandekar, Roopa R., Sandhya Badiger, Sathvik Udupa, Prasanta Kumar Ghosh, Hema A. Murthy, Heiga Zen, Pranaw Kumar, Kamal Kant, Amol Bole, Bira Chandra Singh, Keiichi Tokuda, Mark Hasegawa-Johnson, Philipp Olbrich |
ICASSP | 3 |
| 2022 | Gram Vaani ASR Challenge on spontaneous telephone speech recordings in regional variations of HindiabstractThis paper describes the corpus and baseline systems for the Gram Vaani Automatic Speech Recognition (ASR) challenge in regional variations of Hindi. The corpus for this challenge comprises the spontaneous telephone speech recordings collected by a social technology enterprise, Gram Vaani. The regional variations of Hindi together with spontaneity of speech, natural background and transcriptions with variable accuracy due to crowdsourcing make it a unique corpus for ASR on spontaneous telephonic speech. Around, 1108 hours of real-world spontaneous speech recordings, including 1000 hours of unlabelled training data, 100 hours of labelled training data, 5 hours of development data and 3 hours of evaluation data, have been released as a part of the challenge. The efficacy of both training and test sets are validated on different ASR systems in both traditional time-delay neural network-hidden Markov model (TDNN-HMM) frameworks and fully-neural end-to-end (E2E) setup. The word error rate (WER) and character error rate (CER) on eval set for a TDNN model trained on 100 hours of labelled data are 29.7 and 15.1, respectively. While, in E2E setup, WER and CER on eval set for a conformer model trained on 100 hours of data are 32.9 and 19.0, respectively. Anish Bhanushali, Grant Bridgman, Deekshitha G, Prasanta Kumar Ghosh, Pratik Kumar, Adithya Raj Kolladath, Nithya Ravi, Aaditeshwar Seth, Ashish Seth, Abhayjeet Singh, Vrunda N. Sukhadia, Srinivasan Umesh, Sathvik Udupa, Lodagala Durga Prasad |
INTERSPEECH | 3 |