EDBT 2026 Demo / reviewers in the wild / expert
Sunayana Sitaram
dblp:27/7642
· DBLP profile ↗
42ranked-venue papers
6as first author
28since 2021 · last 2026
0000-0003-4251-9719ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 5 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic LanguagesabstractPranjal A Chitale, Varun Gumma, Sanchit Ahuja, Prashant Kodali, Manan Uppadhyay, Deepthi Sudharsan, Sunayana Sitaram. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Pranjal A. Chitale, Varun Gumma, Sanchit Ahuja, Prashant Kodali, Manan Uppadhyay, Deepthi Sudharsan, Sunayana Sitaram |
ACL (1) | 7 |
| 2026 | Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot SettingsabstractLarge Language Models (LLMs) are typically evaluated through general or domain-specific benchmarks testing capabilities that often lack grounding in the lived realities of end users. Critical domains such as healthcare require evaluations that extend beyond artificial or simulated tasks to reflect the everyday needs, cultural practices, and nuanced contexts of communities. We propose Samiksha, a community-driven evaluation pipeline co-created with civil-society organizations (CSOs) and community members. Our approach enables scalable, automated benchmarking through a culturally aware, community-driven pipeline in which community feedback informs what to evaluate, how the benchmark is built, and how outputs are scored. We demonstrate this approach in the health domain in India. Our analysis highlights how current multilingual LLMs address nuanced community health queries, while also offering a scalable pathway for contextually grounded and inclusive LLM evaluation. Hamna, Gayatri Bhat, Sourabrata Mukherjee, Faisal M. Lalani, Evan Hadfield, Divya Siddarth, Kalika Bali, Sunayana Sitaram |
CHI | 8 |
| 2026 | Introduction to the Special Issue on Evaluations of Large Language Models Part 2
Jindong Wang 0001, Linyi Yang, Sunayana Sitaram, Qiang Yang 0001, Bhiksha Raj |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2025 | MADASR 2.0: Multi-Lingual Multi-Dialect ASR Challenge in 8 Indian LanguagesabstractWe present MADASR 2.0, a challenge at ASRU 2025 aimed at advancing multilingual and multidialectal automatic speech recognition (ASR) in low-resource Indian languages. Building on the 2023 edition, it introduces a subset of the RESPIN corpus, over 1200 hours of read speech across 8 languages and 33 dialects, with test sets including both read and spontaneous speech. The challenge comprises four tracks varying by training data size and external resource usage, and supports auxiliary tasks like language and dialect identification. We detail the dataset, tasks, baselines, and submissions and analyse trends across tracks and speech styles. Results highlight the continued difficulty of spontaneous ASR, the benefits of multitask and transfer learning, and effective strategies for building dialect-aware ASR systems. MADASR 2.0 offers a standardised benchmark to support future research on inclusive and scalable ASR for linguistically diverse populations. Sumit Sharma 0016, Deekshitha G, Abhayjeet Singh, Amartyaveer, Sathvik Udupa, Sandhya Badiger, Sanjeev Khudanpur, Sunayana Sitaram, Srinivasan Umesh, Bhuvana Ramabhadran, Brian Kingsbury, Hema A. Murthy, Srikanth S. Narayanan, Howard Lakougna, Prasanta Kumar Ghosh |
ASRU | 9 |
| 2025 | JOOCI: a Novel Method for Learning Comprehensive Speech RepresentationsabstractInformation in speech can be categorized into two groups: Content (what is being said, such as linguistics) and Other (how it is expressed such as information about speaker and paralinguistic features). Current self-supervised learning (SSL) methods are shown to divide the model’s representational-depth or layers in two, with earlier layers specializing in Other and later layers in Content related tasks. This layer-wise division is inherently sub-optimal, as neither information type can use all layers to build hierarchical representations. To address this, we propose JOOCI, a novel speech representation learning method that does not compromise on the representational-depth for either information type. JOOCI outperforms WavLM by 26.5% (relative), and other models of similar size ($\mathbf{1 0 0 M}$ parameters), when evaluated on two speaker recognition and two language tasks from the SUPERB benchmark, demonstrating its effectiveness in Jointly Optimizing Other and Content Information (JOOCI). Hemant Yadav, Sunayana Sitaram, Rajiv Ratn Shah |
ASRU | 2 |
| 2025 | Bridging the Language Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMsabstractLarge language models (LLMs) have revolutionized various domains but still struggle with non-Latin scripts and low-resource languages. This paper addresses the critical challenge of improving multilingual performance without extensive fine-tuning. We introduce a novel dynamic learning approach that optimizes prompt strategy, embedding model, and LLM per query at runtime. By adapting configurations dynamically, our method achieves significant improvements over static, best and random baselines. It operates efficiently in both offline and online settings, generalizing seamlessly across new languages and datasets. Leveraging Retrieval-Augmented Generation (RAG) with state-of-the-art multilingual embeddings, we achieve superior task performance across diverse linguistic contexts. Through systematic investigation and evaluation across18 diverse languages using popular question-answering (QA) datasets we show our approach results in 10-15% improvements in multilingual performance over pre-trained models and 4x gains compared to fine-tuned, language-specific models. Somnath Kumar, Vaibhav Balloli, Mercy Ranjit, Kabir Ahuja, Sunayana Sitaram, Kalika Bali, Tanuja Ganu, Akshay Uttama Nambi |
COLING | 5 |
| 2025 | Improving Cross Lingual Transfer by Pretraining with Active ForgettingabstractLarge Language Models (LLMs) demonstrate exceptional capabilities in a multitude of NLP tasks. However, the efficacy of such models to languages other than English is often limited. Prior works have shown that encoder-only models such as BERT or XLM-RoBERTa show impressive cross lingual transfer of their capabilities from English to other languages. In this work, we propose a pretraining strategy that uses active forgetting to achieve similar cross lingual transfer in decoder-only LLMs. We show that LLMs pretrained with active forgetting are highly effective when adapting to new and unseen languages. Through extensive experimentation, we find that LLMs pretrained with active forgetting are able to learn better multilingual representations which translates to better performance in many downstream tasks. Divyanshu Aggarwal, Ashutosh Sathe, Sunayana Sitaram |
EMNLP | 3 |
| 2025 | A Multilingual, Culture-First Approach to Addressing Misgendering in LLM ApplicationsabstractMisgendering is the act of referring to someone by a gender that does not match their chosen identity.It marginalizes and undermines a person's sense of self, causing significant harm.English-based approaches have clear-cut approaches to avoiding misgendering, such as the use of the pronoun "they".However, other languages pose unique challenges due to both grammatical and cultural constructs.In this work we develop methodologies to assess and mitigate misgendering across 42 languages and dialects using a participatory-design approach to design effective and appropriate guardrails across all languages.We test these guardrails in a standard LLM-based application (meeting transcript summarization), where both the data generation and the annotation steps followed a human-in-the-loop approach.We find that the proposed guardrails are very effective in reducing misgendering rates across all languages in the summaries generated, and without incurring loss of quality.Our human-in-the-loop approach demonstrates a method to feasibly scale inclusive and responsible AI-based solutions across multiple languages and cultures.We release the guardrails and synthetic dataset encompassing 42 languages, along with human and LLM-judge evaluations, to encourage further research on this subject. Sunayana Sitaram, Adrian de Wynter, Isobel McCrum, Qilong Gu |
EMNLP | 1 |
| 2025 | Introduction to the Special Issue on Evaluations of Large Language Models: Part 1abstractNo abstract available. Jindong Wang 0001, Linyi Yang, Sunayana Sitaram, Qiang Yang 0001, Bhiksha Raj |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2024 | DOSA: A Dataset of Social Artifacts from Different Indian Geographical SubculturesabstractGenerative models are increasingly being used in various applications, such as text generation, commonsense reasoning, and question-answering. To be effective globally, these models must be aware of and account for local socio-cultural contexts, making it necessary to have benchmarks to evaluate the models for their cultural familiarity. Since the training data for LLMs is web-based and the Web is limited in its representation of information, it does not capture knowledge present within communities that are not on the Web. Thus, these models exacerbate the inequities, semantic misalignment, and stereotypes from the Web. There has been a growing call for community-centered participatory research methods in NLP. In this work, we respond to this call by using participatory research methods to introduce DOSA, the first community-generated Dataset of 615 Social Artifacts, by engaging with 260 participants from 19 different Indian geographic subcultures. We use a gamified framework that relies on collective sensemaking to collect the names and descriptions of these artifacts such that the descriptions semantically align with the shared sensibilities of the individuals from those cultures. Next, we benchmark four popular LLMs and find that they show significant variation across regional sub-cultures in their ability to infer the artifacts. Agrima Seth, Sanchit Ahuja, Kalika Bali, Sunayana Sitaram |
LREC/COLING | 4 |
| 2024 | MAFIA: Multi-Adapter Fused Inclusive Language ModelsabstractPrachi Jain, Ashutosh Sathe, Varun Gumma, Kabir Ahuja, Sunayana Sitaram. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ashutosh Sathe, Varun Gumma, Kabir Ahuja, Sunayana Sitaram |
EACL (1) | 5 |
| 2024 | Teaching LLMs to Abstain across Languages via Multilingual FeedbackabstractShangbin Feng, Weijia Shi, Yike Wang, Wenxuan Ding, Orevaoghene Ahia, Shuyue Stella Li, Vidhisha Balachandran, Sunayana Sitaram, Yulia Tsvetkov. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Shangbin Feng, Yike Wang 0002, Wenxuan Ding 0001, Orevaoghene Ahia, Shuyue Stella Li, Vidhisha Balachandran, Sunayana Sitaram, Yulia Tsvetkov |
EMNLP | 8 |
| 2024 | Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic PromptingabstractSagnik Mukherjee, Muhammad Farid Adilazuarda, Sunayana Sitaram, Kalika Bali, Alham Fikri Aji, Monojit Choudhury. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Sagnik Mukherjee, Muhammad Farid Adilazuarda, Sunayana Sitaram, Kalika Bali, Alham Fikri Aji, Monojit Choudhury |
EMNLP | 3 |
| 2024 | PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural DataabstractEvaluation of multilingual Large Language Models (LLMs) is challenging due to a variety of factors -the lack of benchmarks with sufficient linguistic diversity, contamination of popular benchmarks into LLM pre-training data and the lack of local, cultural nuances in translated benchmarks.In this work, we study human and LLM-based evaluation in a multilingual, multi-cultural setting.We evaluate 30 models across 10 Indic languages by conducting 90K human evaluations and 30K LLMbased evaluations and find that models such as GPT-4o and Llama-3 70B consistently perform best for most Indic languages.We build leaderboards for two evaluation settings -pairwise comparison and direct assessment and analyse the agreement between humans and LLMs.We find that humans and LLMs agree fairly well in the pairwise setting but the agreement drops for direct assessment evaluation especially for languages such as Bengali and Odia.We also check for various biases in human and LLMbased evaluation and find evidence of self-bias in the GPT-based evaluator.Our work presents a significant step towards scaling up multilingual evaluation of LLMs. 1 few-shot cross-lingual transfer.In Ishaan Watts, Varun Gumma, Aditya Yadavalli, Vivek Seshadri, S. Manohar 0001, Sunayana Sitaram |
EMNLP | 6 |
| 2024 | MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
Hemant Yadav, Sunayana Sitaram, Rajiv Ratn Shah |
INTERSPEECH | 2 |
| 2024 | MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and TasksabstractSanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts, Ashutosh Sathe, Millicent Ochieng, Rishav Hada, Prachi Jain, Mohamed Ahmed, Kalika Bali, Sunayana Sitaram. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Sanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts, Ashutosh Sathe, Millicent Ochieng, Rishav Hada, Kalika Bali, Sunayana Sitaram |
NAACL-HLT | 11 |
| 2024 | CultureLLM: Incorporating Cultural Differences into Large Language ModelsabstractLarge language models (LLMs) have been observed to exhibit bias towards certain cultures due to the predominance of training data obtained from English corpora. Considering that multilingual cultural data is often expensive to procure, existing methodologies address this challenge through prompt engineering or culture-specific pre-training. However, these strategies may neglect the knowledge deficiency of low-resource cultures and necessitate substantial computing resources. In this paper, we propose CultureLLM, a cost-effective solution to integrate cultural differences into LLMs. CultureLLM employs the World Value Survey (WVS) as seed data and generates semantically equivalent training data through the proposed semantic data augmentation. Utilizing only $50$ seed samples from WVS with augmented data, we fine-tune culture-specific LLMs as well as a unified model (CultureLLM-One) for $9$ cultures, encompassing both rich and low-resource languages. Extensive experiments conducted on $60$ culture-related datasets reveal that CultureLLM significantly surpasses various counterparts such as GPT-3.5 (by $8.1$\%) and Gemini Pro (by $9.5$\%), demonstrating performance comparable to or exceeding that of GPT-4. Our human study indicates that the generated samples maintain semantic equivalence to the original samples, offering an effective solution for LLMs augmentation. Code is released at https://github.com/Scarelette/CultureLLM. Mengzhuo Chen, Jindong Wang 0001, Sunayana Sitaram, Xing Xie 0001 |
NeurIPS | 4 |
| 2023 | A Comparative Study on the Impact of Model Compression Techniques on Fairness in Language ModelsabstractCompression techniques for deep learning have become increasingly popular, particularly in settings where latency and memory constraints are imposed.Several methods, such as pruning, distillation, and quantization, have been adopted for compressing models, each providing distinct advantages.However, existing literature demonstrates that compressing deep learning models could affect their fairness.Our analysis involves a comprehensive evaluation of pruned, distilled, and quantized language models, which we benchmark across a range of intrinsic and extrinsic metrics for measuring bias in text classification.We also investigate the impact of using multilingual models and evaluation measures.Our findings highlight the significance of considering both the pre-trained model and the chosen compression strategy in developing equitable language technologies.The results also indicate that compression strategies can have an adverse effect on fairness measures. Krithika Ramesh, Arnav Chavan, Shrey Pandit, Sunayana Sitaram |
ACL (1) | 4 |
| 2023 | Partial Rank Similarity Minimization Method for Quality MOS Prediction of Unseen Speech Synthesis Systems in Zero-Shot and Semi-Supervised SettingabstractThis paper introduces a novel objective function for quality mean opinion score (MOS) prediction of unseen speech synthesis systems. The proposed function measures the similarity of relative positions of predicted MOS values, in a mini-batch, rather than the actual MOS values. That is the partial rank similarity is measured $(\mathcal{P}RS)$ rather than the individual MOS values as with the L1 loss. Our experiments on out-of-domain speech synthesis systems demonstrate that the $\mathcal{P} R S$ outperforms L1 loss in zero-shot and semi-supervised settings, exhibiting stronger correlation with ground truth. These findings highlight the importance of considering rank order, as done by $\mathcal{P}RS$, when training MOS prediction models. We also argue that mean squared error and linear correlation coefficient metrics may be unreliable for evaluating MOS prediction models. In conclusion, $\mathcal{P} R S$-trained models provide a robust framework for evaluating speech quality and offer insights for developing high-quality speech synthesis systems. Code and models are available at github.com/nii-yamagishilab/partial_rank_similarity/ Hemant Yadav, Erica Cooper, Junichi Yamagishi, Sunayana Sitaram, Rajiv Ratn Shah |
ASRU | 4 |
| 2023 | DiTTO: A Feature Representation Imitation Approach for Improving Cross-Lingual TransferabstractShanu Kumar, Soujanya Abbaraju, Sandipan Dandapat, Sunayana Sitaram, Monojit Choudhury. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Shanu Kumar, Abbaraju Soujanya, Sandipan Dandapat, Sunayana Sitaram, Monojit Choudhury |
EACL | 4 |
| 2023 | MEGA: Multilingual Evaluation of Generative AIabstractKabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Mohamed Ahmed, Kalika Bali, Sunayana Sitaram. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Akshay Uttama Nambi, Tanuja Ganu, Sameer Segal, Kalika Bali, Sunayana Sitaram |
EMNLP | 12 |
| 2023 | Analysing the Masked Predictive Coding Training Criterion for Pre-Training a Speech Representation ModelabstractRecent developments in pre-trained speech representation utilizing self-supervised learning (SSL) have yielded exceptional results on a variety of downstream tasks. One such technique, known as masked predictive coding (MPC), has been employed by some of the most high-performing models. In this study, we investigate the impact of MPC loss on the type of information learnt at various layers in the HuBERT model, using nine probing tasks. Our findings indicate that the amount of content information learned at various layers of the HuBERT model has a positive correlation to the MPC loss. Additionally, it is also observed that any speaker-related information learned at intermediate layers of the model, is an indirect consequence of the learning process, and therefore cannot be controlled using the MPC loss. These findings may serve as inspiration for further research in the speech community, specifically in the development of new pre-training tasks or the exploration of new pre-training criterion’s that directly preserves both speaker and content information at various layers of a learnt model. Hemant Yadav, Sunayana Sitaram, Rajiv Ratn Shah |
ICASSP | 2 |
| 2022 | LITMUS Predictor: An AI Assistant for Building Reliable, High-Performing and Fair Multilingual NLP SystemsabstractPre-trained multilingual language models are gaining popularity due to their cross-lingual zero-shot transfer ability, but these models do not perform equally well in all languages. Evaluating task-specific performance of a model in a large number of languages is often a challenge due to lack of labeled data, as is targeting improvements in low performing languages through few-shot learning. We present a tool - LITMUS Predictor - that can make reliable performance projections for a fine-tuned task-specific model in a set of languages without test and training data, and help strategize data labeling efforts to optimize performance and fairness objectives. Anirudh Srinivasan, Gauri Kholkar, Rahul Kejriwal, Tanuja Ganu, Sandipan Dandapat, Sunayana Sitaram, Balakrishnan Santhanam, Somak Aditya, Kalika Bali, Monojit Choudhury |
AAAI | 6 |
| 2022 | On the Calibration of Massively Multilingual Language ModelsabstractMassively Multilingual Language Models (MMLMs) have recently gained popularity due to their surprising effectiveness in cross-lingual transfer.While there has been much work in evaluating these models for their performance on a variety of tasks and languages, little attention has been paid on how well calibrated these models are with respect to the confidence in their predictions.We first investigate the calibration of MMLMs in the zero-shot setting and observe a clear case of miscalibration in low-resource languages or those which are typologically diverse from English.Next, we empirically show that calibration methods like temperature scaling and label smoothing do reasonably well in improving calibration in the zero-shot scenario.We also find that fewshot examples in the language can further help reduce calibration errors, often substantially.Overall, our work contributes towards building more reliable multilingual models by highlighting the issue of their miscalibration, understanding what language and model-specific factors influence it, and pointing out the strategies to improve the same. Kabir Ahuja, Sunayana Sitaram, Sandipan Dandapat, Monojit Choudhury |
EMNLP | 2 |
| 2022 | A Survey of Multilingual Models for Automatic Speech RecognitionabstractAlthough Automatic Speech Recognition (ASR) systems have achieved human-like performance for a few languages, the majority of the world’s languages do not have usable systems due to the lack of large speech datasets to train these models. Cross-lingual transfer is an attractive solution to this problem, because low-resource languages can potentially benefit from higher-resource languages either through transfer learning, or being jointly trained in the same multilingual model. The problem of cross-lingual transfer has been well studied in ASR, however, recent advances in Self Supervised Learning are opening up avenues for unlabeled speech data to be used in multilingual ASR models, which can pave the way for improved performance on low-resource languages. In this paper, we survey the state of the art in multilingual ASR models that are built with cross-lingual transfer in mind. We present best practices for building multilingual models from research across diverse languages and techniques, discuss open questions and provide recommendations for future work. Hemant Yadav, Sunayana Sitaram |
LREC | 2 |
| 2022 | Benchmarking Evaluation Metrics for Code-Switching Automatic Speech RecognitionabstractCode-switching poses a number of challenges and opportunities for multilingual automatic speech recognition. In this paper, we focus on the question of robust and fair evaluation metrics. To that end, we develop a reference benchmark data set of code-switching speech recognition hypotheses with human judgments. We define clear guidelines for minimal editing of automatic hypotheses. We validate the guidelines using 4-way inter-annotator agreement. We evaluate a large number of metrics in terms of correlation with human judgments. The metrics we consider vary in terms of representation (orthographic, phonological, semantic), directness (intrinsic vs extrinsic), granularity (e.g. word, character), and similarity computation method. The highest correlation to human judgment is achieved using transliteration followed by text normalization. We release the first corpus for human acceptance of code-switching speech recognition results in dialectal Arabic/English conversation speech. Injy Hamed, Amir Hussein, Oumnia Chellah, Shammur Absar Chowdhury, Hamdy Mubarak, Sunayana Sitaram, Nizar Habash, Ahmed Ali 0002 |
SLT | 6 |
| 2021 | A Survey of Code-switching: Linguistic and Social Perspectives for Language TechnologiesabstractA. Seza Doğruöz, Sunayana Sitaram, Barbara E. Bullock, Almeida Jacqueline Toribio. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. A. Seza Dogruöz, Sunayana Sitaram, Barbara Bullock, Almeida Jacqueline Toribio |
ACL/IJCNLP (1) | 2 |
| 2021 | MUCS 2021: Multilingual and Code-Switching ASR Challenges for Low Resource Indian LanguagesabstractRecently, there is increasing interest in multilingual automatic speech recognition (ASR) where a speech recognition system caters to multiple low resource languages by taking advantage of low amounts of labeled corpora in multiple languages. With multilingualism becoming common in today's world, there has been increasing interest in code-switching ASR as well. In code-switching, multiple languages are freely interchanged within a single sentence or between sentences. The success of low-resource multilingual and code-switching ASR often depends on the variety of languages in terms of their acoustics, linguistic characteristics as well as the amount of data available and how these are carefully considered in building the ASR system. In this challenge, we would like to focus on building multilingual and code-switching ASR systems through two different subtasks related to a total of seven Indian languages, namely Hindi, Marathi, Odia, Tamil, Telugu, Gujarati and Bengali. For this purpose, we provide a total of ~600 hours of transcribed speech data, comprising train and test sets, in these languages including two code-switched language pairs, Hindi-English and Bengali-English. We also provide a baseline recipe for both the tasks with a WER of 30.73% and 32.45% on the test sets of multilingual and code-switching subtasks, respectively. Anuj Diwan, Rakesh Vaideeswaran, Sanket Shah, Ankita Singh, Srinivasa Raghavan K. M., Shreya Khare, Vinit Unni, Saurabh Vyas, Akash Rajpuria, Chiranjeevi Yarra, Ashish R. Mittal, Prasanta Kumar Ghosh, Preethi Jyothi, Kalika Bali, Vivek Seshadri, Sunayana Sitaram, Samarth Bharadwaj, Jai Nanavati, Raoul Nanavati, Karthik Sankaranarayanan |
Interspeech | 16 |
| 2020 | GLUECoS: An Evaluation Benchmark for Code-Switched NLPabstractCode-switching is the use of more than one language in the same conversation or utterance.Recently, multilingual contextual embedding models, trained on multiple monolingual corpora, have shown promising results on cross-lingual and multilingual tasks.We present an evaluation benchmark, GLUECoS, for code-switched languages, that spans several NLP tasks in English-Hindi and English-Spanish.Specifically, our evaluation benchmark includes Language Identification from text, POS tagging, Named Entity Recognition, Sentiment Analysis, Question Answering and a new task for code-switching, Natural Language Inference.We present results on all these tasks using cross-lingual word embedding models and multilingual models.In addition, we fine-tune multilingual models on artificially generated code-switched data.Although multilingual models perform significantly better than cross-lingual models, our results show that in most tasks, across both language pairs, multilingual models fine-tuned on code-switched data perform best, showing that multilingual models can be further optimized for code-switching tasks. Simran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram, Monojit Choudhury |
ACL | 4 |
| 2020 | Crowdsourcing Speech Data for Low-Resource Languages from Low-Income WorkersabstractVoice-based technologies are essential to cater to the hundreds of millions of new smartphone users. However, most of the languages spoken by these new users have little to no labelled speech data. Unfortunately, collecting labelled speech data in any language is an expensive and resource-intensive task. Moreover, existing platforms typically collect speech data only from urban speakers familiar with digital technology whose dialects are often very different from low-income users. In this paper, we explore the possibility of collecting labelled speech data directly from low-income workers. In addition to providing diversity to the speech dataset, we believe this approach can also provide valuable supplemental earning opportunities to these communities. To this end, we conducted a study where we collected labelled speech data in the Marathi language from three different user groups: low-income rural users, low-income urban users, and university students. Overall, we collected 109 hours of data from 36 participants. Our results show that the data collected from low-income participants is of comparable quality to the data collected from university students (who are typically employed to do this work) and that crowdsourcing speech data from low-income rural and urban workers is a viable method of gathering speech data. Basil Abraham, Danish Goel, Divya Siddarth, Kalika Bali, Manu Chopra, Monojit Choudhury, Pratik Joshi, Preethi Jyothi, Sunayana Sitaram, Vivek Seshadri |
LREC | 9 |
| 2018 | Language Modeling for Code-Mixing: The Role of Linguistic Theory based Synthetic DataabstractAdithya Pratapa, Gayatri Bhat, Monojit Choudhury, Sunayana Sitaram, Sandipan Dandapat, Kalika Bali. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Adithya Pratapa, Gayatri Bhat, Monojit Choudhury, Sunayana Sitaram, Sandipan Dandapat, Kalika Bali |
ACL (1) | 4 |
| 2018 | Word Embeddings for Code-Mixed Language ProcessingabstractWe compare three existing bilingual word embedding approaches, and a novel approach of training skip-grams on synthetic code-mixed text generated through linguistic models of code-mixing, on two tasks -sentiment analysis and POS tagging for code-mixed text.Our results show that while CVM and CCA based embeddings perform as well as the proposed embedding technique on semantic and syntactic tasks respectively, the proposed approach provides the best performance for both tasks overall.Thus, this study demonstrates that existing bilingual embedding techniques are not ideal for code-mixed text processing and there is a need for learning multilingual word embedding from the code-mixed text. Adithya Pratapa, Monojit Choudhury, Sunayana Sitaram |
EMNLP | 3 |
| 2018 | Effect of TTS Generated Audio on OOV Detection and Word Error Rate in ASR for Low-resource LanguagesabstractOut-of-Vocabulary (OOV) detection and recovery is an important aspect of reducing Word Error Rate (WER) in Automatic Speech Recognition (ASR). In this paper, we evaluate the effect of OOV detection and recovery for a low-resource language on WER. We use a small seed corpus of continuous speech and improve the vocabulary by incorporating the detected OOV words. We use a syllable-model to learn OOV words and augment the word-model with these words leading to improved recognition. Our research investigates the effect on OOV detection and recovery after adding missing syllable sounds in the syllable model using a Text-to-Speech (TTS) system. Our experiments are conducted using 5 hours of continuous speech Kannada corpus. We use an already available Festival TTS for Hindi to generate Kannada speech. Our initial experiments report an improvement in OOV detection due to addition of missing syllable sounds using a cross-lingual TTS system. Savitha Murthy, Dinkar Sitaram, Sunayana Sitaram |
INTERSPEECH | 3 |
| 2018 | Homophone Identification and Merging for Code-switched Speech RecognitionabstractCode-switching or mixing is the use of multiple languages in a single utterance or conversation. Borrowing occurs when a word from a foreign language becomes part of the vocabulary of a language. In multilingual societies, switching/mixing and borrowing are not always clearly distinguishable. Due to this, transcription of code-switched and borrowed words is often not standardized, and leads to the presence of homophones in the training data. In this work, we automatically identify and disambiguate homophones in code-switched data to improve recognition of code-switched speech. We use a WX-based common pronunciation scheme for both languages being mixed and unify the homophones during training, which results in a lower word error rate for systems built using this data. We also extend this framework to propose a metric for code-switched speech recognition that takes into account homophones in both languages while calculating WER, which can help provide a more accurate picture of errors the ASR system makes on code-switched speech. Brij Mohan Lal Srivastava, Sunayana Sitaram |
INTERSPEECH | 2 |
| 2018 | Discovering Canonical Indian English Accents: A Crowdsourcing-based Approach
Sunayana Sitaram, Varun Manjunath, Varun Bharadwaj, Monojit Choudhury, Kalika Bali, Michael Tjalve |
LREC | 1 |
| 2017 | Speech Synthesis for Mixed-Language Navigation InstructionsabstractText-to-Speech (TTS) systems that can read navigation instructions are one of the most widely used speech interfaces today. Text in the navigation domain may contain named entities such as location names that are not in the language that the TTS database is recorded in. Moreover, named entities can be compound words where individual lexical items belong to different languages. These named entities may be transliterated into the script that the TTS system is trained on. This may result in incorrect pronunciation rules being used for such words. We describe experiments to extend our previous work in generating code-mixed speech to synthesize navigation instructions, with a mixed-lingual TTS system. We conduct subjective listening tests with two sets of users, one being students who are native speakers of an Indian language and very proficient in English, and the other being drivers with low English literacy, but familiarity with location names. We find that in both sets of users, there is a significant preference for our proposed system over a baseline system that synthesizes instructions in English. Khyathi Raghavi Chandu, Sai Krishna Rallabandi, Sunayana Sitaram, Alan W. Black |
INTERSPEECH | 3 |
| 2016 | Speech Synthesis of Code-Mixed Text
Sunayana Sitaram, Alan W. Black |
LREC | 1 |
| 2016 | Polyglot Neural Language Models: A Case Study in Cross-Lingual Phonetic Representation LearningabstractYulia Tsvetkov, Sunayana Sitaram, Manaal Faruqui, Guillaume Lample, Patrick Littell, David Mortensen, Alan W Black, Lori Levin, Chris Dyer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Yulia Tsvetkov, Sunayana Sitaram, Manaal Faruqui, Guillaume Lample, Patrick Littell, David R. Mortensen, Alan W. Black, Lori S. Levin, Chris Dyer |
HLT-NAACL | 2 |
| 2015 | Using articulatory features and inferred phonological segments in zero resource speech processing
Pallavi Baljekar, Sunayana Sitaram, Prasanna Kumar Muthukumar, Alan W. Black |
INTERSPEECH | 2 |
| 2015 | Using acoustics to improve pronunciation for synthesis of low resource languages
Sunayana Sitaram, Serena Jeblee, Alan W. Black |
INTERSPEECH | 1 |
| 2015 | Universal grapheme-based speech synthesis
Sunayana Sitaram, Alok Parlikar, Gopala Krishna Anumanchipalli, Alan W. Black |
INTERSPEECH | 1 |
| 2013 | Bootstrapping Text-to-Speech for speech processing in languages without an orthographyabstractSpeech synthesis technology has reached the stage where given a well-designed corpus of audio and accurate transcription an at least understandable synthesizer can be built without necessarily resorting to new innovations. However many languages do not have a well-defined writing system but such languages could still greatly benefit from speech systems. In this paper we consider the case where we have a (potentially large) single speaker database but have no transcriptions and no standardized way to write transcriptions. To address this scenario we propose a method that allows us to bootstrap synthetic voices purely from speech data. We use a novel combination of automatic speech recognition and automatic word segmentation for the bootstrapping. Our experimental results on speech corpora in two languages, English and German, show that synthetic voices that are built using this method are close to understandable. Our method is language-independent and can thus be used to build synthetic voices from a speech corpus in any new language. Sunayana Sitaram, Sukhada Palkar, Yun-Nung Chen, Alok Parlikar, Alan W. Black |
ICASSP | 1 |