Saket Dingliwal

dblp:236/4652 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Accelerated Test-Time Scaling with Model-Free Speculative Sampling
abstract
Woomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh, Jinwoo Shin, Aram Galstyan, Sravan Babu Bodapati. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Woomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh, Jinwoo Shin, Aram Galstyan, Sravan Babu Bodapati
EMNLP2
2025 SeRA: Self-Reviewing and Alignment of LLMs using Implicit Reward Margins
abstract
Direct alignment algorithms (DAAs), such as direct preference optimization (DPO), have become popular alternatives to Reinforcement Learning from Human Feedback (RLHF) due to their simplicity, efficiency, and stability. However, the preferences used by DAAs are usually collected before alignment training begins and remain unchanged (off-policy). This design leads to two problems where the policy model (1) picks up on spurious correlations in the dataset (as opposed to only learning alignment to human preferences), and (2) overfits to feedback on off-policy trajectories that have less likelihood of being generated by the updated policy model. To address these issues, we introduce Self-Reviewing and Alignment (SeRA), a cost-efficient and effective method that can be readily combined with existing DAAs. SeRA comprises of two components: (1) sample selection using implicit reward margin to alleviate over-optimization on such undesired features, and (2) preference bootstrapping using implicit rewards to augment preference data with updated policy models in a cost-efficient manner. Extensive experiments, including on instruction-following tasks, demonstrate the effectiveness and generality of SeRA in training LLMs with diverse offline preference datasets and and DAAs.
Jongwoo Ko, Saket Dingliwal, Bhavana Ganesh, Sailik Sengupta, Sravan Babu Bodapati, Aram Galstyan
ICLR2
2024 Sequential Editing for Lifelong Training of Speech Recognition Models
Devang Kulshreshtha, Nikolaos Pappas 0004, Brady Houston, Saket Dingliwal, Srikanth Ronanki
INTERSPEECH4
2023 Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale
abstract
Hritik Bansal, Karthik Gopalakrishnan, Saket Dingliwal, Sravan Bodapati, Katrin Kirchhoff, Dan Roth. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Hritik Bansal, Karthik Gopalakrishnan 0001, Saket Dingliwal, Sravan Babu Bodapati, Katrin Kirchhoff, Dan Roth 0001
ACL (1)3
2023 Don't Stop Self-Supervision: Accent Adaptation of Speech Representations via Residual Adapters
Anshu Bhatia, Sanchit Sinha, Saket Dingliwal, Karthik Gopalakrishnan 0001, Sravan Babu Bodapati, Katrin Kirchhoff
INTERSPEECH3
2023 Multilingual Contextual Adapters To Improve Custom Word Recognition In Low-resource Languages
Devang Kulshreshtha, Saket Dingliwal, Brady Houston, Sravan Babu Bodapati
INTERSPEECH2
2022 Domain Prompts: Towards memory and compute efficient domain adaptation of ASR systems
abstract
Automatic Speech Recognition (ASR) systems have found their use in numerous industrial applications in very diverse domains creating a need to adapt to new domains with small memory and deployment overhead. In this work, we introduce domain prompts, a methodology that involves training a small number of domain embedding parameters to prime a Transformer-based Language Model (LM) to a particular domain. Using this domain-adapted LM for rescoring ASR hypotheses can achieve 7-13% WER reduction for a new domain with just 1000 unlabeled textual domain-specific sentences and a handful of additional parameters. Our method can match or even beat the performance of models fully fine-tuned towards a particular domain with only 0.02% of the parameters. Given the parameter efficiency and negligible deployment overhead, our experiments showcase that such a method is an ideal choice for on-the-fly adaptation of LMs used in ASR systems to progressively scale it to new domains.
Saket Dingliwal, Ashish Shenoy, Sravan Babu Bodapati, Ankur Gandhe, Ravi Gadde, Katrin Kirchhoff
INTERSPEECH1
2022 Personalization of CTC Speech Recognition Models
abstract
End-to-end speech recognition models trained using joint Connectionist Temporal Classification (CTC)-Attention loss have gained popularity recently. In these models, a non-autoregressive CTC decoder is often used at inference time due to its speed and simplicity. However, such models are hard to personalize because of their conditional independence assumption that prevents output tokens from previous time steps to influence future predictions. To tackle this, we propose a novel two-way approach that first biases the encoder with attention over a predefined list of rare long-tail and out-of-vocabulary (OOV) words and then uses dynamic boosting and phone alignment network during decoding to further bias the subword pre-dictions. We evaluate our approach on open-source VoxPopuli and in-house medical datasets to showcase a 60% improvement in F1 score on domain-specific rare words over a strong CTC baseline.
Saket Dingliwal, Monica Sunkara, Srikanth Ronanki, Jeff Farris, Katrin Kirchhoff, Sravan Babu Bodapati
SLT1
2021 Few Shot Dialogue State Tracking using Meta-learning
abstract
Saket Dingliwal, Shuyang Gao, Sanchit Agarwal, Chien-Wei Lin, Tagyoung Chung, Dilek Hakkani-Tur. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Saket Dingliwal, Shuyang Gao, Sanchit Agarwal, Chien-Wei Lin, Tagyoung Chung, Dilek Hakkani-Tür
EACL1
2020 Robust Handwriting Recognition with Limited and Noisy Data
abstract
Despite the advent of deep learning in computer vision, the general handwriting recognition problem is far from solved. Most existing approaches focus on handwriting datasets that have clearly written text and carefully segmented labels. In this paper, we instead focus on learning handwritten characters from maintenance logs, a constrained setting where data is very limited and noisy. We break the problem into two consecutive stages of word segmentation and word recognition respectively, and utilize data augmentation techniques to train both stages. Extensive comparisons with popular baselines for scene-text detection and word recognition show that our system achieves a lower error rate and is more suited to handle noisy and difficult documents.
Hai Pham, Amrith Setlur, Saket Dingliwal, Tzu-Hsiang Lin, Barnabás Póczos, Jae Lim, Collin McCormack, Tam Vu 0002
ICFHR3