Shikhar Bharadwaj

dblp:216/0480 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2026
0009-0003-7202-0502ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2026 PRiSM: Benchmarking Phone Realization in Speech Models
abstract
Shikhar Bharadwaj, Chin-Jou Li, Yoonjae Kim, Kwanghee Choi, Eunjung Yeo, Ryan Soh-Eun Shim, Hanyu Zhou, Brendon Boldt, Karen Rosero, Kalvin Chang, Darsh Agrawal, Keer Xu, Chao-Han Huck Yang, Jian Zhu, Shinji Watanabe, David R. Mortensen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shikhar Bharadwaj, Chin-Jou Li, Yoonjae Kim, Kwanghee Choi, Eunjung Yeo, Ryan Soh-Eun Shim, Hanyu Zhou, Brendon Boldt, Karen Rosero, Kalvin Chang, Darsh Agrawal, Keer Xu, Chao-Han Huck Yang, Shinji Watanabe 0001, David R. Mortensen
ACL (1)1
2026 POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
abstract
Chin-Jou Li, Kalvin Chang, Shikhar Bharadwaj, Eunjung Yeo, Kwanghee Choi, Jian Zhu, David R. Mortensen, Shinji Watanabe. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chin-Jou Li, Kalvin Chang, Shikhar Bharadwaj, Eunjung Yeo, Kwanghee Choi, David R. Mortensen, Shinji Watanabe 0001
ACL (1)3
2025 VERSA-v2: A Modular and Scalable Toolkit for Speech and Audio Evaluation with Expanded Metrics, Visualization, and LLM Integration
abstract
We present VERSA-v2, a major upgrade of the Versatile Evaluation of Speech and Audio (VERSA) toolkit for standardized and scalable evaluation across speech, audio, and music tasks. It features a modular, object-oriented architecture that simplifies metric integration and now supports over 100 metrics, organized into curated task-specific packs. VERSA-v2 also introduces interactive visualizations, per-metric profiling, and prompt-based evaluation using both text- and audio-based large language models (LLMs). These advancements make VERSA-v2 a robust, extensible, and LLM-enabled platform for comprehensive and interpretable speech and audio evaluation.
Jiatong Shi, Bo-Hao Su, Shikhar Bharadwaj, Shih-Heng Wang, Jionghao Hang, Wei Wang 0010, Wenhao Feng, Yuxun Tang, Nezih Topaloglu, Siddhant Arora, Jinchuan Tian, Hye-Jin Shim, Wangyou Zhang, Wen-Chin Huang, Shinji Watanabe 0001
ASRU3
2025 Identifying and Mitigating Mismatched Language Code in Multilingual ASR
abstract
Multilingual speech recognition systems often use an input language code in order to prompt the transcription in the target language. However, the spoken language in the input audio may not always match the language code, as often prevalent in multilingual societies. This language mismatch can significantly reduce ASR quality. We present a technique to identify and mitigate this issue. We combine off-the-shelf language-ID and language verification models to determine the language code input to the ASR model. The language verification model acts as a gate that decides when to trust the provided language code or use the output of the language-ID model. We compare these approaches with baselines that include vanilla language-ID based and language-independent ASR models. Our experiments on YouTube, SPRING-INX and FLEURS datasets shows the efficacy of the proposed model especially in the mismatched language code setting.
Sepand Mavandadi, Kartik Audhkhasi, Shikhar Bharadwaj, Brian Farris, Tongzhou Chen, Bhuvana Ramabhadran, Sriram Ganapathy
ICASSP4
2025 Context-Driven Dynamic Pruning for Large Speech Foundation Models
Masao Someki, Shikhar Bharadwaj, Atharva Anand Joshi, Chyi-Jiunn Lin, Jinchuan Tian, Jee-Weon Jung, Nathan Susanj, Shinji Watanabe 0001
INTERSPEECH2
2025 OpusLM: A Family of Open Unified Speech Language Models
Jinchuan Tian, Yifan Peng 0003, Jiatong Shi, Siddhant Arora, Shikhar Bharadwaj, Takashi Maekaku, Yusuke Shinohara, Keita Goto, Xiang Yue, Chao-Han Huck Yang, Shinji Watanabe 0001
INTERSPEECH6
2025 EmoNews: A Spoken Dialogue System for Expressive News Conversations
abstract
We develop a task-oriented spoken dialogue system (SDS) that regulates emotional speech based on contextual cues to enable more empathetic news conversations. Despite advancements in emotional text-to-speech (TTS) techniques, task-oriented emotional SDSs remain underexplored due to the compartmentalized nature of SDS and emotional TTS research, as well as the lack of standardized evaluation metrics for social goals. We address these challenges by developing an emotional SDS for news conversations that utilizes a large language model (LLM)-based sentiment analyzer to identify appropriate emotions and PromptTTS to synthesize context-appropriate emotional speech. We also propose subjective evaluation scale for emotional SDSs and judge the emotion regulation performance of the proposed and baseline systems. Experiments showed that our emotional SDS outperformed a baseline system in terms of the emotion regulation and engagement. These results suggest the critical role of speech emotion for more engaging conversations. All our source code is open-sourced.
Ryuki Matsuura, Shikhar Bharadwaj, Dhatchi Kunde Govindarajan
SIGDIAL2
2024 IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
abstract
As large language models (LLMs) see increasing adoption across the globe, it is imperative for LLMs to be representative of the linguistic diversity of the world.India is a linguistically diverse country of 1.4 Billion people.To facilitate research on multilingual LLM evaluation, we release INDICGENBENCHthe largest benchmark for evaluating LLMs on user-facing generation tasks across a diverse set 29 of Indic languages covering 13 scripts and 4 language families.INDICGEN-BENCH is composed of diverse generation tasks like cross-lingual summarization, machine translation, and cross-lingual question answering.INDICGENBENCH extends existing benchmarks to many Indic languages through human curation providing multi-way parallel evaluation data for many under-represented Indic languages for the first time.We evaluate a wide range of proprietary and open-source LLMs including GPT-3.5, GPT-4, PaLM-2, mT5, Gemma, BLOOM and LLaMA on IN-DICGENBENCH in a variety of settings.The largest PaLM-2 models performs the best on most tasks, however, there is a significant performance gap in all languages compared to English showing that further research is needed for the development of more inclusive multilingual language models.INDICGENBENCH is available at www.github.com/google-research- datasets/indic-gen-bench 2 INDICGENBENCH INDICGENBENCH is a high-quality, humancurated benchmark to evaluate text generation capabilities of multilingual models on Indic languages.Our benchmark consists of 5 user-facing tasks (viz., summarization, machine translation, and question answering) across 29 Indic languages spanning 13 writing scripts and 4 language families.For certain tasks, INDICGENBENCH provides the first-ever evaluation dataset for up to 18 Indic languages.Table 1 provides summary of INDICGENBENCH and examples of instances across tasks present in it.Languages in INDICGENBENCH are divided into (relatively) Higher, Medium, and Low resource categories based on the availability of web text resources (see appendix §A for details).
Harman Singh, Nitish Gupta, Shikhar Bharadwaj, Dinesh Tewari, Partha Talukdar
ACL (1)3
2024 Multimodal Modeling for Spoken Language Identification
abstract
Spoken language identification refers to the task of automatically predicting the spoken language in a given utterance. Conventionally, it is modeled as a speech-based language identification task. Prior techniques have been constrained to a single modality; however in the case of video data there is a wealth of other metadata that may be beneficial for this task. In this work, we propose MuSeLI, a Multimodal Spoken Language Identification method, which delves into the use of various metadata sources to enhance language identification. Our study reveals that metadata such as video title, description and geographic location provide substantial information to identify the spoken language of the multimedia recording. We conduct experiments using two diverse public datasets of YouTube videos, and obtain state-of-the-art results on the language identification task. We additionally conduct an ablation study that describes the distinct contribution of each modality for language recognition.
Shikhar Bharadwaj, Shikhar Vashishth, Ankur Bapna, Sriram Ganapathy, Vera Axelrod, Siddharth Dalmia, Daan van Esch, Sandy Ritchie, Partha Talukdar, Jason Riesa
ICASSP1
2023 MASR: Multi-Label Aware Speech Representation
abstract
In the recent years, speech representation learning is constructed primarily as a self-supervised learning (SSL) task, using the raw audio signal alone, while ignoring the side-information that is often available for a given speech recording. In this paper, we propose MASR, a Multi-label Aware Speech Representation learning framework, which addresses the aforementioned limitations. MASR enables the inclusion of multiple external knowledge sources to enhance the utilization of meta-data information. The external knowledge sources are incorporated in the form of sample-level pair-wise similarity matrices that are useful in a hard-mining loss. A key advantage of the MASR framework is that it can be combined with any choice of SSL method. Using MASR representations, we perform evaluations on several downstream tasks such as language identification, speech recognition and other non-semantic tasks such as speaker and emotion recognition. In these experiments, we illustrate significant performance improvements for the MASR over other established benchmarks. We perform a detailed analysis on the language identification task to provide insights on how the proposed loss function enables the representations to separate closely related languages.
Anjali Raj, Shikhar Bharadwaj, Sriram Ganapathy, Shikhar Vashishth
ASRU2
2023 Label Aware Speech Representation Learning For Language Identification
Shikhar Vashishth, Shikhar Bharadwaj, Sriram Ganapathy, Ankur Bapna, Wei Han 0002, Vera Axelrod, Partha Talukdar
INTERSPEECH2
2022 Efficient Constituency Tree based Encoding for Natural Language to Bash Translation
abstract
Bash is a Unix command language used for interacting with the Operating System.Recent works on natural language to Bash translation have made significant advances, but none of the previous methods utilize the problem's inherent structure.We identify this structure and propose a Segmented Invocation Transformer (SIT) that utilizes the information from the constituency parse tree of the natural language text.Our method is motivated by the alignment between segments in the natural language text and Bash command components.Incorporating the structure in the modelling improves the performance of the model.Since such systems must be universally accessible, we benchmark the inference times on a CPU rather than a GPU.We observe a 1.8x improvement in the inference time and a 5x reduction in model parameters.Attribution analysis using Integrated Gradients reveals that the proposed method can capture the problem structure.
Shikhar Bharadwaj, Shirish K. Shevade
NAACL-HLT1
2021 Explainable Natural Language to Bash Translation using Abstract Syntax Tree
abstract
Natural language processing for program synthesis has been widely researched.In this work, we focus on generating Bash commands from natural language invocations with explanations.We propose a novel transformer based solution by utilizing Bash Abstract Syntax Trees and manual pages.Our method incorporates tree structure information in the transformer architecture and provides explanations for its predictions via alignment matrices between user invocation and manual page text.Our method performs on par with the state of the art performance on Natural Language Context to Command task and performs better than fine-tuned T5 and Seq2Seq models.
Shikhar Bharadwaj, Shirish K. Shevade
CoNLL1