VLDB 2026 Research / reviewers in the wild / expert
Rohit Paturi
dblp:173/6480
· DBLP profile ↗
8ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SEAL: Speaker Error Correction using Acoustic-conditioned Large Language ModelsabstractSpeaker Diarization (SD) is a crucial component of modern end-to-end ASR pipelines. Traditional SD systems, which are typically audio-based and operate independently of ASR, often introduce speaker errors, particularly during speaker transitions and overlapping speech. Recently, language models including fine-tuned large language models (LLMs) have shown to be effective as a second-pass speaker error corrector by leveraging lexical context in the transcribed output. In this work, we introduce a novel acoustic conditioning approach to provide more fine-grained information from the acoustic diarizer to the LLM. We also show that a simpler constrained decoding strategy reduces LLM hallucinations, while avoiding complicated post-processing. Our approach significantly reduces the speaker error rates by 24-43% across Fisher, Callhome, and RT03-CTS datasets, compared to the first-pass Acoustic SD. Rohit Paturi, Amber Afshan, Sundararajan Srinivasan |
ICASSP | 2 |
| 2024 | Speakers Unembedded: Embedding-free Approach to Long-form Neural Diarization
Vivek Govindan, Rohit Paturi, Sundararajan Srinivasan |
INTERSPEECH | 3 |
| 2024 | AG-LSEC: Audio Grounded Lexical Speaker Error Correction
Rohit Paturi, Sundararajan Srinivasan |
INTERSPEECH | 1 |
| 2023 | Generalized Zero-Shot Audio-to-Intent ClassificationabstractSpoken language understanding systems using audio-only data are gaining popularity, yet their ability to handle unseen intents remains limited. In this study, we propose a generalized zero-shot audio-to-intent classification framework with only a few sample text sentences per intent. To achieve this, we first train a supervised audio-to-intent classifier by making use of a self-supervised pre-trained model. We then leverage a neural audio synthesizer to create audio embeddings for sample text utterances and perform generalized zero-shot classification on unseen intents using cosine similarity. We also propose a multimodal training strategy that incorporates lexical information into the audio representation to improve zero-shot performance. Our multimodal training approach improves the accuracy of zero-shot intent classification on unseen intents of SLURP by 2.75% and 18.2% for the SLURP and internal goal-oriented dialog datasets, respectively, compared to audio-only training. Veera Raghavendra Elluru, Devang Kulshreshtha, Rohit Paturi, Sravan Babu Bodapati, Srikanth Ronanki |
ASRU | 3 |
| 2023 | End-to-End Single-Channel Speaker-Turn Aware Conversational Speech TranslationabstractJuan Pablo Zuluaga-Gomez, Zhaocheng Huang, Xing Niu, Rohit Paturi, Sundararajan Srinivasan, Prashant Mathur, Brian Thompson, Marcello Federico. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Juan Zuluaga-Gomez, Zhaocheng Huang, Xing Niu 0001, Rohit Paturi, Sundararajan Srinivasan, Prashant Mathur, Brian Thompson 0001, Marcello Federico |
EMNLP | 4 |
| 2023 | Lexical Speaker Error Correction: Leveraging Language Models for Speaker Diarization Error Correction
Rohit Paturi, Sundararajan Srinivasan |
INTERSPEECH | 1 |
| 2022 | Directed speech separation for automatic speech recognition of long form conversational speechabstractMany of the recent advances in speech separation are primarily aimed at synthetic mixtures of short audio utterances with high degrees of overlap.Most of these approaches need an additional stitching step to stitch the separated speech chunks for long form audio. Since most of the approaches involve Permutation Invariant training (PIT), the order of separated speech chunks is nondeterministic and leads to difficulty in accurately stitching homogenous speaker chunks for downstream tasks like Automatic Speech Recognition (ASR).Also, most of these models are trained with synthetic mixtures and do not generalize to real conversational data.In this paper, we propose a speaker conditioned separator trained on speaker embeddings extracted directly from the mixed signal using an over-clustering based approach.This model naturally regulates the order of the separated chunks without the need for an additional stitching step.We also introduce a data sampling strategy with real and synthetic mixtures which generalizes well to real conversation speech.With this model and data sampling technique, we show significant improvements in speaker-attributed word error rate (SA-WER) on Hub5 data. Rohit Paturi, Sundararajan Srinivasan, Katrin Kirchhoff, Daniel Garcia-Romero |
INTERSPEECH | 1 |
| 2015 | Age-dependent height estimation and speaker normalization for children's speech using the first three subglottal resonancesabstractThis paper proposes an age-dependent scheme for automatic height estimation and speaker normalization of children’s speech, using the first three subglottal resonances (SGRs). Similar to previous work, our analysis indicates that children above the age of 11 years show different acoustic properties from those under 11. Therefore, an age-dependent model is investigated. The estimation algorithms for the first three SGRs are motivated by our previous research for adults. The algorithms for the first two SGRs have been applied to children’s speech before. This paper proposes a similar approach to estimate Sg3 for children. The algorithm is trained and evaluated on 46 children, aged between 6-17 years, using cross-validation. Average RMS errors in estimating Sg1, Sg2 and Sg3 using the age-dependent model are 51, 128 and 168 Hz, respectively. The height estimation algorithm employs a negative correlation between SGRs and height, and the mean absolute height estimation error was found to be less than 3.8cm for the younger children and 4.9cm for the older children. In addition, using TIDIGITS, a linear frequency warping scheme using age-dependent Sg3 gives statisticallysignificant word error rate reductions (up to 26%) relative to conventional VTLN. Jinxi Guo, Rohit Paturi, Gary Yeung, Steven M. Lulich, Harish Arsikere, Abeer Alwan |
INTERSPEECH | 2 |