Rohit Paturi

dblp:173/6480 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2025 SEAL: Speaker Error Correction using Acoustic-conditioned Large Language Models
abstract
Speaker Diarization (SD) is a crucial component of modern end-to-end ASR pipelines. Traditional SD systems, which are typically audio-based and operate independently of ASR, often introduce speaker errors, particularly during speaker transitions and overlapping speech. Recently, language models including fine-tuned large language models (LLMs) have shown to be effective as a second-pass speaker error corrector by leveraging lexical context in the transcribed output. In this work, we introduce a novel acoustic conditioning approach to provide more fine-grained information from the acoustic diarizer to the LLM. We also show that a simpler constrained decoding strategy reduces LLM hallucinations, while avoiding complicated post-processing. Our approach significantly reduces the speaker error rates by 24-43% across Fisher, Callhome, and RT03-CTS datasets, compared to the first-pass Acoustic SD.
Rohit Paturi, Amber Afshan, Sundararajan Srinivasan
ICASSP2
2024 Speakers Unembedded: Embedding-free Approach to Long-form Neural Diarization
Vivek Govindan, Rohit Paturi, Sundararajan Srinivasan
INTERSPEECH3
2024 AG-LSEC: Audio Grounded Lexical Speaker Error Correction
Rohit Paturi, Sundararajan Srinivasan
INTERSPEECH1
2023 Generalized Zero-Shot Audio-to-Intent Classification
abstract
Spoken language understanding systems using audio-only data are gaining popularity, yet their ability to handle unseen intents remains limited. In this study, we propose a generalized zero-shot audio-to-intent classification framework with only a few sample text sentences per intent. To achieve this, we first train a supervised audio-to-intent classifier by making use of a self-supervised pre-trained model. We then leverage a neural audio synthesizer to create audio embeddings for sample text utterances and perform generalized zero-shot classification on unseen intents using cosine similarity. We also propose a multimodal training strategy that incorporates lexical information into the audio representation to improve zero-shot performance. Our multimodal training approach improves the accuracy of zero-shot intent classification on unseen intents of SLURP by 2.75% and 18.2% for the SLURP and internal goal-oriented dialog datasets, respectively, compared to audio-only training.
Veera Raghavendra Elluru, Devang Kulshreshtha, Rohit Paturi, Sravan Babu Bodapati, Srikanth Ronanki
ASRU3
2023 End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation
abstract
Juan Pablo Zuluaga-Gomez, Zhaocheng Huang, Xing Niu, Rohit Paturi, Sundararajan Srinivasan, Prashant Mathur, Brian Thompson, Marcello Federico. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Juan Zuluaga-Gomez, Zhaocheng Huang, Xing Niu 0001, Rohit Paturi, Sundararajan Srinivasan, Prashant Mathur, Brian Thompson 0001, Marcello Federico
EMNLP4
2023 Lexical Speaker Error Correction: Leveraging Language Models for Speaker Diarization Error Correction
Rohit Paturi, Sundararajan Srinivasan
INTERSPEECH1
2022 Directed speech separation for automatic speech recognition of long form conversational speech
abstract
Many of the recent advances in speech separation are primarily aimed at synthetic mixtures of short audio utterances with high degrees of overlap.Most of these approaches need an additional stitching step to stitch the separated speech chunks for long form audio. Since most of the approaches involve Permutation Invariant training (PIT), the order of separated speech chunks is nondeterministic and leads to difficulty in accurately stitching homogenous speaker chunks for downstream tasks like Automatic Speech Recognition (ASR).Also, most of these models are trained with synthetic mixtures and do not generalize to real conversational data.In this paper, we propose a speaker conditioned separator trained on speaker embeddings extracted directly from the mixed signal using an over-clustering based approach.This model naturally regulates the order of the separated chunks without the need for an additional stitching step.We also introduce a data sampling strategy with real and synthetic mixtures which generalizes well to real conversation speech.With this model and data sampling technique, we show significant improvements in speaker-attributed word error rate (SA-WER) on Hub5 data.
Rohit Paturi, Sundararajan Srinivasan, Katrin Kirchhoff, Daniel Garcia-Romero
INTERSPEECH1
2015 Age-dependent height estimation and speaker normalization for children's speech using the first three subglottal resonances
abstract
This paper proposes an age-dependent scheme for automatic height estimation and speaker normalization of children’s speech, using the first three subglottal resonances (SGRs). Similar to previous work, our analysis indicates that children above the age of 11 years show different acoustic properties from those under 11. Therefore, an age-dependent model is investigated. The estimation algorithms for the first three SGRs are motivated by our previous research for adults. The algorithms for the first two SGRs have been applied to children’s speech before. This paper proposes a similar approach to estimate Sg3 for children. The algorithm is trained and evaluated on 46 children, aged between 6-17 years, using cross-validation. Average RMS errors in estimating Sg1, Sg2 and Sg3 using the age-dependent model are 51, 128 and 168 Hz, respectively. The height estimation algorithm employs a negative correlation between SGRs and height, and the mean absolute height estimation error was found to be less than 3.8cm for the younger children and 4.9cm for the older children. In addition, using TIDIGITS, a linear frequency warping scheme using age-dependent Sg3 gives statisticallysignificant word error rate reductions (up to 26%) relative to conventional VTLN.
Jinxi Guo, Rohit Paturi, Gary Yeung, Steven M. Lulich, Harish Arsikere, Abeer Alwan
INTERSPEECH2