EDBT 2026 Demo / reviewers in the wild / expert
Sathvik Udupa
dblp:270/4407
· DBLP profile ↗
20ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0002-2225-5464ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 9 first-author · 18 since 2021Artificial intelligence and machine learning · 15 · 7 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General IntelligenceabstractAudio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challenging. To address this gap, we introduce MMAU-Pro, the most comprehensive and rigorously curated benchmark for assessing audio intelligence in AI systems. MMAU-Pro contains 5,305 instances, where each instance has one or more audios paired with human expert-generated question-answer pairs, spanning speech, sound, music, and their combinations. Unlike existing benchmarks, MMAU-Pro evaluates auditory intelligence across 49 unique skills and multiple complex dimensions, including long-form audio comprehension, spatial audio reasoning, multi-audio understanding, among others. All questions are meticulously designed to require deliberate multi-hop reasoning, including both multiple-choice and open-ended response formats. Importantly, audio data is sourced directly ``from the wild" rather than from existing datasets with known distributions. We evaluate 22 leading open-source and proprietary multimodal AI models, revealing significant limitations: even state-of-the-art models such as Gemini 2.5 Flash and Audio Flamingo 3 achieve only 59.2% and 51.7% accuracy, respectively, approaching random performance in multiple categories. Our extensive analysis highlights specific shortcomings and provides novel insights, offering actionable perspectives for the community to enhance future AI systems' progression toward audio general intelligence. The benchmark and code is available at https://sonalkum.github.io/mmau-pro. Sonal Kumar, Simon Sedlácek, Vaibhavi Lokegaonkar, Fernando López, Wenyi Yu, Nishit Anand, Hyeonggon Ryu, Lichang Chen, Maxim Plicka, Miroslav Hlavácek, William Fineas Ellingwood, Sathvik Udupa, Siyuan Hou, Allison Ferner, Sara Barahona, Cecilia Bolaños, Satish Rahi, Laura Herrera-Alarcón, Satvik Dixit, Rupali S. Patil, Soham Deshmukh, Lasha Koroshinadze, L. Paola García-Perera, Eleni Zanou, Themos Stafylakis, Joon Son Chung, David F. Harwath, Dinesh Manocha, Alicia Lozano-Diez, Santosh Kesiraju, Sreyan Ghosh, Ramani Duraiswami |
AAAI | 12 |
| 2025 | MADASR 2.0: Multi-Lingual Multi-Dialect ASR Challenge in 8 Indian LanguagesabstractWe present MADASR 2.0, a challenge at ASRU 2025 aimed at advancing multilingual and multidialectal automatic speech recognition (ASR) in low-resource Indian languages. Building on the 2023 edition, it introduces a subset of the RESPIN corpus, over 1200 hours of read speech across 8 languages and 33 dialects, with test sets including both read and spontaneous speech. The challenge comprises four tracks varying by training data size and external resource usage, and supports auxiliary tasks like language and dialect identification. We detail the dataset, tasks, baselines, and submissions and analyse trends across tracks and speech styles. Results highlight the continued difficulty of spontaneous ASR, the benefits of multitask and transfer learning, and effective strategies for building dialect-aware ASR systems. MADASR 2.0 offers a standardised benchmark to support future research on inclusive and scalable ASR for linguistically diverse populations. Sumit Sharma 0016, Deekshitha G, Abhayjeet Singh, Amartyaveer, Sathvik Udupa, Sandhya Badiger, Sanjeev Khudanpur, Sunayana Sitaram, Srinivasan Umesh, Bhuvana Ramabhadran, Brian Kingsbury, Hema A. Murthy, Srikanth S. Narayanan, Howard Lakougna, Prasanta Kumar Ghosh |
ASRU | 6 |
| 2025 | Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed TrainingabstractAccurate, low-latency endpointing is crucial for effective spoken dialogue systems. While traditional endpointers often rely on spectrum-based audio features, this work proposes real-time speech endpointing for multi-turn dialogues using streaming, low-bitrate Neural Audio Codec (NAC) features, building upon recent advancements in neural audio codecs. To further reduce cutoff errors, we introduce a novel label delay training scheme. At a fixed median latency of 160 ms, our combined NAC and label delay approach achieves significant relative cutoff error reductions: 42.7% for a single-stream endpointer and 37.5% for a two-stream configuration, compared to baseline methods. Finally, we demonstrate efficient integration with a codec-based pretrained speech large language model, improving its median response time by 1200 ms and reducing its cutoff error by 35%. Sathvik Udupa, Shinji Watanabe 0001, Petr Schwarz, Jan Cernocký |
ASRU | 1 |
| 2025 | Improving Dialect Identification in Indian Languages Using Multimodal Features from Dialect Informed ASRabstractDialect identification (DID) addresses the challenge of recog-nizing regional variations within a language. The current deep learning approaches focus on audio-only, text-only, or multi-task setups combining automatic speech recognition (ASR) with DID. This work introduces a novel multimodal architecture that leverages speech and text features to enhance DID performance. Our method integrates ASR-generated speech representations with text embeddings derived from ASR hypotheses using a RoBERTa-based encoder. Additionally, we perform a layer-wise analysis of the IndicWav2Vec model to identify the layers most effective for extracting dialectal features. We evaluate our approach on a subset of the RESPIN dataset featuring eight Indian languages and 33 dialects. Experimental results show that our proposed multimodal DID system achieves an average DID accuracy of 79.81%, consistently outperforming baseline methods. This study is the first to analyse comprehensively DID in Indian languages, providing new insights into their dialectal diversity. Amartyaveer, Sumit Sharma 0016, Sathvik Udupa, Sandhya Badiger, Abhayjeet Singh, Deekshitha G, Jesuraja Bandekar, Savitha Murthy, Prasanta Kumar Ghosh |
ICASSP | 4 |
| 2025 | Comparison of Acoustic and Textual Features for Dysarthria Severity Classification in Amyotrophic Lateral Sclerosis
Y. S. Upendra Vishwanath, Tanuka Bhattacharjee, Deekshitha G, Sathvik Udupa, Chowdam Venkata Thirumala Kumar, Madassu Keerthipriya, Darshan Chikktimmegowda, Dipti Baskar, Yamini Belur, Seena Vengalil, Atchayaram Nalini, Prasanta Kumar Ghosh |
INTERSPEECH | 4 |
| 2025 | RESPIN-S1.0: A read speech corpus of 10000+ hours in dialects of nine Indian LanguagesabstractWe introduce RESPIN-S1.0, the largest publicly available dialect-rich read-speech corpus for Indian languages, comprising more than 10,000 hours of validated audio across nine major languages: Bengali, Bhojpuri, Chhattisgarhi, Hindi, Kannada, Magahi, Maithili, Marathi, and Telugu. Indian languages exhibit high dialectal variation and are spoken by populations that remain digitally underserved. Existing speech corpora typically represent only standard dialects and lack domain and linguistic diversity. RESPIN-S1.0 addresses this limitation by collecting speech across more than 38 dialects and two high-impact domains: agriculture and finance. Text data were composed by native dialect speakers and validated through a pipeline combining automated and manual checks. Over 200,000 unique sentences were recorded through a crowdsourced mobile platform and categorised into clean, semi-noisy, and noisy subsets based on transcription quality, with the clean portion alone exceeding 10,000 hours. Along with audio and transcriptions, RESPIN provides dialect-aware phonetic lexicons, speaker metadata, and reproducible train, development, and test splits. To benchmark performance, we evaluate multiple ASR models, including TDNN-HMM, E-Branchformer, Whisper, and wav2vec2-based self-supervised models, and find that fine-tuning on RESPIN significantly improves recognition accuracy over pretrained baselines. A subset of RESPIN-S1.0 has already supported community challenges such as the SLT Code Hackathon 2022 and MADASR@ASRU 2023 and 2025, releasing more than 1,200 hours publicly. This resource supports research in dialectal ASR, language identification, and related speech technologies, establishing a comprehensive benchmark for inclusive, dialect-rich ASR in multilingual low-resource settings. Dataset: https://spiredatasets.ee.iisc.ac.in/respincorpus Code: https://github.com/labspire/respin_baselines.git Abhayjeet Singh, Deekshitha G, Amartya Veer, Jesuraja Bandekar, Savitha Murthy, Sumit Sharma 0016, Sandhya Badiger, Sathvik Udupa, Amala Nagireddi, Srinivasa Raghavan K. M., Rohan Saxena, Jai Nanavati, Raoul Nanavati, Janani Sridharan, Arjun Singh Mehta, Ashish Seth, Sai Praneeth Reddy Mora, Prashanthi V, Gauri Date, Karthika P, Prasanta Kumar Ghosh |
NeurIPS | 9 |
| 2024 | Articulatory synthesis using representations learnt through phonetic label-aware contrastive loss
Jesuraja Bandekar, Sathvik Udupa, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2024 | Adapter pre-training for improved speech recognition in unseen domains using low resource adapter tuning of self-supervised models
Sathvik Udupa, Jesuraja Bandekar, Deekshitha G, Sandhya Badiger, Abhayjeet Singh Savitha Murthy, Priyanka Pai, Srinivasa Raghavan K. M., Raoul Nanavati, Prasanta Kumar Ghosh |
INTERSPEECH | 1 |
| 2024 | IndicMOS: Multilingual MOS Prediction for 7 Indian languages
Sathvik Udupa, Soumi Maiti, Prasanta Kumar Ghosh |
INTERSPEECH | 1 |
| 2023 | Gated Multi Encoders and Multitask Objectives for Dialectal Speech Recognition in Indian LanguagesabstractIn this work, several methods have been proposed towards improving the performance of dialectal automatic speech recognition (ASR). A novel encoder architecture has been introduced that is suited for multi-dialect ASR training. Further, we propose Multi-Task Self-Supervised learning (SSL) fine-tuning using CTC and dialect identification. Additionally, the use of different language models (LM) to improve the performance of dialectal ASR has been investigated. Around 800 hours of Bengali and Bhojpuri data, released as a part of the MADASR ASRU challenge have been used to train these models. The work shows that the proposed multi-encoder ASR observes a relative reduction of 7.5% and 9% in WER in Bhojpuri and Bengali, respectively. Additionally, we also observe a 1-2% WER reduction in fine-tuning SSL, further improving performance in these languages. Moreover, we observe advantages in using dialect-specific LM decoding based on predicted dialect. Sathvik Udupa, Jesuraja Bandekar, Deekshitha G, Prasanta Kumar Ghosh, Sandhya Badiger, Abhayjeet Singh, Savitha Murthy, Priyanka Pai, Srinivasa Raghavan K. M., Raoul Nanavati |
ASRU | 1 |
| 2023 | Lightweight, Multi-Speaker, Multi-Lingual Indic Text-to-SpeechabstractThe Lightweight, Multi-speaker, Multi-lingual Indic Text-to-Speech (LIMMITS’23) challenge is organized as part of the ICASSP 2023 signal processing grand challenge. LIMMITS’23 aims at the development of a lightweight, multi-speaker, multi-lingual Text to Speech (TTS) model using datasets in Marathi, Hindi, and Telugu. The challenge encourages the advancement of TTS in Indian Languages as well as the development of techniques involved in TTS data selection and model compression. The 3 tracks of LIMMITS’23 have provided an opportunity for various researchers and practitioners around the world to explore the state of the art in TTS research. Abhayjeet Singh, Amala Nagireddi, Deekshitha G, Jesuraja Bandekar, Roopa R., Sandhya Badiger, Sathvik Udupa, Prasanta Kumar Ghosh, Hema A. Murthy, Heiga Zen, Pranaw Kumar, Kamal Kant, Amol Bole, Bira Chandra Singh, Keiichi Tokuda, Mark Hasegawa-Johnson, Philipp Olbrich |
ICASSP | 7 |
| 2023 | Real-Time MRI Video Synthesis from Time Aligned Phonemes with Sequence-to-Sequence NetworksabstractReal-Time Magnetic resonance imaging (rtMRI) of the midsagittal plane of the mouth is of interest for speech production research. In this work, we focus on estimating utterance level rtMRI video from the spoken phoneme sequence. We obtain time-aligned phonemes from forced alignment, to obtain frame-level phoneme sequences which are aligned with rtMRI frames. We propose a sequence-tosequence learning model with a transformer phoneme encoder and convolutional frame decoder. We then modify the learning by using intermediary features obtained from sampling from a pretrained phoneme-conditioned variational autoencoder (CVAE). We train on 8 subjects in a subject-specific manner and demonstrate the performance with a subjective test. We also use an auxiliary task of air tissue boundary (ATB) segmentation to obtain the objective scores on the proposed models. We show that the proposed method is able to generate realistic rtMRI video for unseen utterances, and adding CVAE is beneficial for learning the sequence-to-sequence mapping for subjects where the mapping is hard to learn. Sathvik Udupa, Prasanta Kumar Ghosh |
ICASSP | 1 |
| 2023 | Improved Acoustic-to-Articulatory Inversion Using Representations from Pretrained Self-Supervised Learning ModelsabstractIn this work, we investigate the effectiveness of pretrained Self-Supervised Learning (SSL) features for learning the mapping for acoustic to articulatory inversion (AAI). Signal processing-based acoustic features such as MFCCs have been predominantly used for the AAI task with deep neural networks. With SSL features working well for various other speech tasks such as speech recognition, emotion classification, etc., we experiment with its efficacy for AAI. We train on SSL features with transformer neural networks-based AAI models of 3 different model complexities and compare its performance with MFCCs in subject-specific (SS), pooled and fine-tuned (FT) configurations with data from 10 subjects, and evaluate with correlation coefficient (CC) score on the unseen sentence test set. We find that acoustic feature reconstruction objective-based SSL features such as TERA and DeCoAR work well for AAI, with SS CCs of these SSL features reaching close to the best FT CCs of MFCC. We also find the results consistent across different model sizes. Sathvik Udupa, C. Siddarth, Prasanta Kumar Ghosh |
ICASSP | 1 |
| 2023 | Exploring a classification approach using quantised articulatory movements for acoustic to articulatory inversion
Jesuraja Bandekar, Sathvik Udupa, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2022 | Gram Vaani ASR Challenge on spontaneous telephone speech recordings in regional variations of HindiabstractThis paper describes the corpus and baseline systems for the Gram Vaani Automatic Speech Recognition (ASR) challenge in regional variations of Hindi. The corpus for this challenge comprises the spontaneous telephone speech recordings collected by a social technology enterprise, Gram Vaani. The regional variations of Hindi together with spontaneity of speech, natural background and transcriptions with variable accuracy due to crowdsourcing make it a unique corpus for ASR on spontaneous telephonic speech. Around, 1108 hours of real-world spontaneous speech recordings, including 1000 hours of unlabelled training data, 100 hours of labelled training data, 5 hours of development data and 3 hours of evaluation data, have been released as a part of the challenge. The efficacy of both training and test sets are validated on different ASR systems in both traditional time-delay neural network-hidden Markov model (TDNN-HMM) frameworks and fully-neural end-to-end (E2E) setup. The word error rate (WER) and character error rate (CER) on eval set for a TDNN model trained on 100 hours of labelled data are 29.7 and 15.1, respectively. While, in E2E setup, WER and CER on eval set for a conformer model trained on 100 hours of data are 32.9 and 19.0, respectively. Anish Bhanushali, Grant Bridgman, Deekshitha G, Prasanta Kumar Ghosh, Pratik Kumar, Adithya Raj Kolladath, Nithya Ravi, Aaditeshwar Seth, Ashish Seth, Abhayjeet Singh, Vrunda N. Sukhadia, Srinivasan Umesh, Sathvik Udupa, Lodagala Durga Prasad |
INTERSPEECH | 14 |
| 2022 | Watch Me Speak: 2D Visualization of Human Mouth during Speech
C. Siddarth, Sathvik Udupa, Prasanta Kumar Ghosh |
INTERSPEECH | 2 |
| 2022 | Streaming model for Acoustic to Articulatory Inversion with transformer networksabstractEstimating speech articulatory movements from speech acoustics is known as Acoustic to Articulatory Inversion (AAI). Recently, transformer-based AAI models have been shown to achieve state-of-art performance. However, in transformer networks, the attention is applied over the whole utterance, thereby needing to obtain the full utterance before the inference, which leads to high latency and is impractical for streaming AAI. To enable streaming during inference, evaluation could be performed on non-overlapping chucks instead of a full utterance. However, due to a mismatch of the attention receptive field during training and evaluation, there could be a drop in AAI performance. To overcome this scenario, in this work we perform experiments with different attention masks and use context from previous predictions during training. Experiments results revealed that using the random start mask attention with the context from previous predictions of transformer decoder performs better than the baseline results. Sathvik Udupa, Aravind Illa, Prasanta Kumar Ghosh |
INTERSPEECH | 1 |
| 2021 | Estimating Articulatory Movements in Speech Production with Transformer NetworksabstractWe estimate articulatory movements in speech production from different modalities - acoustics and phonemes. Acoustic-to articulatory inversion (AAI) is a sequence-to-sequence task. On the other hand, phoneme to articulatory (PTA) motion estimation faces a key challenge in reliably aligning the text and the articulatory movements. To address this challenge, we explore the use of a transformer architecture - FastSpeech, with explicit duration modelling to learn hard alignments between the phonemes and articulatory movements. We also train a transformer model on AAI. We use correlation coefficient (CC) and root mean squared error (rMSE) to assess the estimation performance in comparison to existing methods on both tasks. We observe 154%, 11.8% & 4.8% relative improvement in CC with subject-dependent, pooled and fine-tuning strategies, respectively, for PTA estimation. Additionally, on the AAI task, we obtain 1.5%, 3% and 3.1% relative gain in CC on the same setups compared to the state-of-the-art baseline. We further present the computational benefits of having transformer architecture as representation blocks. Sathvik Udupa, Anwesha Roy, Abhayjeet Singh, Aravind Illa, Prasanta Kumar Ghosh |
Interspeech | 1 |
| 2021 | Web Interface for Estimating Articulatory Movements in Speech Production from Acoustics and Text
Sathvik Udupa, Anwesha Roy, Abhayjeet Singh, Aravind Illa, Prasanta Kumar Ghosh |
Interspeech | 1 |
| 2020 | Voice based classification of patients with Amyotrophic Lateral Sclerosis, Parkinson's Disease and Healthy Controls with CNN-LSTM using transfer learningabstractIn this paper, we consider 2-class and 3-class classification problems for classifying patients with Amyotrophic Lateral Sclerosis (ALS), Parkinson's Disease (PD), and Healthy Controls (HC) using a CNNLSTM network. Classification performance is examined for three different tasks, namely, Spontaneous speech (SPON), Diadochokinetic rate (DIDK) and Sustained phoneme production (PHON). Experiments are conducted using speech data recorded from 60 ALS, 60 PD, and 60 HC subjects. Classifications using SVM and DNN are considered as baseline schemes. Classification accuracy of ALS and HC (indicated by ALS/HC) using CNN-LSTM has shown an improvement of 10.40%, 4.22% and 0.08% for PHON, SPON and DIDK tasks, respectively over the best of the baseline schemes. Furthermore, the CNN-LSTM network achieves the highest PD/HC classification accuracy of 88.5% for the SPON task and the highest 3-class (ALS/PD/HC) classification accuracy of 85.24% for the DIDK task. Experiments using transfer learning at low resource training data show that data from ALS benefits PD/HC classification and vice-versa. Experiments with fine-tuning weights of 3-class (ALS/PD/HC) classifier for 2-class classification (PD/HC or ALS/HC) gives an absolute improvement of 2% classification accuracy in SPON task when compared with randomly initialized 2-class classifier. Jhansi Mallela, Aravind Illa, Suhas B. N., Sathvik Udupa, Yamini Belur, Atchayaram Nalini, Pradeep Reddy, Dipanjan Gope, Prasanta Kumar Ghosh |
ICASSP | 4 |