VLDB 2026 Research / reviewers in the wild / expert
Kiran Praveen
dblp:258/8273
· DBLP profile ↗
6ranked-venue papers
5as first author
5since 2021 · last 2023
0000-0002-7489-417XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Prefix Search Decoding for RNN Transducers
Kiran Praveen, Advait Vinay Dhopeshwarkar, Balaji Radhakrishnan |
INTERSPEECH | 1 |
| 2023 | Language Identification Networks for Multilingual Everyday Recordings
Kiran Praveen, Balaji Radhakrishnan, Kamini Sabu, M. Ali Basha Shaik |
INTERSPEECH | 1 |
| 2021 | Warped Ensembles: A Novel Technique for Improving CTC Based End-to-End Speech RecognitionabstractCombining outputs from predictive models trained for similar tasks generally perform better than using a single model. Models trained on different domains could for example help in improving the performance with the help of information from complementary domains. The weights in the ensemble could also be adjusted to best suit the desired target domain. However, in End-to-End (E2E) based speech recognition solutions, this is not a simple task owing to the lack of time alignments in the outputs. This work presents a novel ensembling technique for E2E ASR systems trained using Connectionist Temporal Classification (CTC) loss, which uses a multi-sequence alignment technique for synchronising output time-frames across models, and then combines the outputs using weights obtained using a calibration technique introduced in our earlier work. Several experiments are conducted on multi-accent English data and the Librispeech corpus. Word Error Rates (WER) are evaluated and compared against the performance of ROVER (a system combination technique), and the best model in the system. When compared to the best model, average relative improvements of 6.46% and 8.94% are observed in unadapted in-domain and out-of-domain experiments respectively. Comparison to ROVER shows average relative improvements of 19.29% and 9.21% in in-domain and out-of-domain experiments. Kiran Praveen, Hardik B. Sailor |
ASRU | 1 |
| 2021 | SRI-B End-to-End System for Multilingual and Code-Switching ASR Challenges for Low Resource Indian Languages
Hardik B. Sailor, Kiran Praveen, Vikas Agrawal |
Interspeech | 2 |
| 2021 | Dynamically Weighted Ensemble Models for Automatic Speech RecognitionabstractIn machine learning, training multiple models for the same task, and using the outputs from all the models helps reduce the variance of the combined result. Using an ensemble of models in classification tasks such as Automatic Speech Recognition (ASR) improves the accuracy across different target domains such as multiple accents, environmental conditions, and other scenarios. It is possible to select model weights for the ensemble in numerous ways. A classifier trained to identify target domain, a simple averaging function, or an exhaustive grid search are the common approaches to obtain suitable weights. All these methods suffer either in choosing sub-optimal weights or by being computationally expensive. We propose a novel and practical method for dynamic weight selection in an ensemble, which can approximate a grid search in a time-efficient manner. We show that a combination of weights always performs better than assigning uniform weights for all models. Our algorithm can utilize a validation set if available or find weights dynamically from the input utterance itself. Experiments conducted for various ASR tasks show that the proposed method outperforms the uniformly weighted ensemble in terms of Word Error Rate (WER) in our experiments. Kiran Praveen, Shakti Prasad Rath, Sandip Shriram Bapat |
SLT | 1 |
| 2019 | Second Language Transfer Learning in Humans and Machines Using Image SupervisionabstractIn the task of language learning, humans exhibit remarkable ability to learn new words from a foreign language with very few instances of image supervision. The question therefore is whether such transfer learning efficiency can be simulated in machines. In this paper, we propose a deep semantic model for transfer learning words from a foreign language (Japanese) using image supervision. The proposed model is a deep audio-visual correspondence network that uses a proxy based triplet loss. The model is trained with large dataset of multi-modal speech/image input in the native language (English). Then, a subset of the model parameters of the audio network are transfer learned to the foreign language words using proxy vectors from the image modality. Using the proxy based learning approach, we show that the proposed machine model achieves transfer learning performance for an image retrieval task which is comparable to the human performance. We also present an analysis that contrasts the errors made by humans and machines in this task. Kiran Praveen, Akshara Soman, Sriram Ganapathy |
ASRU | 1 |