VLDB 2026 Research / reviewers in the wild / expert
Monojit Choudhury
dblp:29/5841
· DBLP profile ↗
71ranked-venue papers
5as first author
30since 2021 · last 2025
0000-0001-7473-7839ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 59 · 5 first-author · 28 since 2021Databases, data management, data science and information retrieval · 8 · 1 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | All Languages Matter: Evaluating LMMs on Culturally Diverse 100 LanguagesabstractExisting Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support low-resource languages, all while effectively integrating corresponding visual cues. In pursuit of culturally diverse global multimodal models, our proposed All Languages Matter Benchmark (ALM-bench) represents the largest and most comprehensive effort to date for evaluating LMMs across 100 languages. ALM-bench challenges existing models by testing their ability to understand and reason about culturally diverse images paired with text in various languages, including many low-resource languages traditionally underrepresented in LMM research. The benchmark offers a robust and nuanced evaluation framework featuring various question formats, including true/false, multiple choice, and open-ended questions, which are further divided into short and long-answer categories. ALM-bench design ensures a comprehensive assessment of a model’s ability to handle varied levels of difficulty in visual and linguistic reasoning. To capture the rich tapestry of global cultures, ALM-bench carefully curates content from 13 distinct cultural aspects, ranging from traditions and rituals to famous personalities and celebrations. Through this, ALM-bench not only provides a rigorous testing ground for state-of-the-art open and closed-source LMMs but also highlights the importance of cultural and linguistic inclusivity, encouraging the development of models that can serve diverse global populations effectively. Our benchmark is publicly available at https://mbzuai-oryx.github.io/ALM-Bench/. Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana, Noor Ahsan, Nevasini Sasikumar, Omkar Thawakar, Henok Biadglign Ademtew, Yahya Hmaiti, Amandeep Kumar, Kartik Kuckreja, Mykola Maslych, Wafa Al Ghallabi, Mihail Minkov Mihaylov, Abdelrahman M. Shaker, Mike Zhang, Mahardika Krisna Ihsani, Amiel Esplana, Monil Gokani, Shachar Mirkin, Harsh Singh, Ashay Srivastava, Endre Hamerlik, Fathinah Asma Izzati, Fadillah A. Maani, Sebastian Cavada, Jenny Chim, Rohit Gupta 0012, Sanjay Manjunath, Kamila Zhumakhanova, Feno Heriniaina Rabevohitra, Azril Hafizi Amirudin, Muhammad Ridzuan, Daniya Najiha Abdul Kareem, Ketan More, Pramesh Shakya, Amirpouya Ghasemaghaei, Amirbek Djanibekov, Dilshod Azizov, Branislava Jankovic, Naman Bhatia, Alvaro Cabrera, Johan S. Obando-Ceron, Olympiah Otieno, Fabian Farestam, Muztoba Rabbani, Sanoojan Baliah, Santosh Sanjeev, Abduragim Shtanchaev, Maheen Fatima, Amrin Kareem, Toluwani Aremu, Nathan A. Z. Xavier, Amit Bhatkal, Hawau Olamide Toyin, Aman Chadha, Hisham Cholakkal, Rao Muhammad Anwer, Michael Felsberg, Jorma Laaksonen, Thamar Solorio, Monojit Choudhury, Ivan Laptev, Mubarak Shah, Salman Khan 0001, Fahad Shahbaz Khan |
CVPR | 65 |
| 2025 | An Interdisciplinary Approach to Human-Centered Machine TranslationabstractMarine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Fred Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé Iii, Kevin Duh, Ge Gao, Alvin C Grissom II, Marzena Karpinska, Elaine C Khoong, William D. Lewis, Andre Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Frédéric Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé III, Kevin Duh, Ge Gao 0001, Alvin Grissom II, Marzena Karpinska, Elaine C. Khoong, William D. Lewis, André F. T. Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon |
EMNLP | 7 |
| 2025 | Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and HindiabstractThe obligatory use of third-person honorifics is a distinctive feature of several South Asian languages, encoding nuanced socio-pragmatic cues such as power, age, gender, fame, and social distance.In this work, (i) We present the first large-scale study of third-person honorific pronoun and verb usage across 10,000 Hindi and Bengali Wikipedia articles with annotations linked to key socio-demographic attributes of the subjects, including gender, age group, fame, and cultural origin.(ii) Our analysis uncovers systematic intra-language regularities but notable cross-linguistic differences: honorifics are more prevalent in Bengali than in Hindi, while non-honorifics dominate while referring to infamous, juvenile, and culturally "exotic" entities.Notably, in both languages, and more prominently in Hindi, men are more frequently addressed with honorifics than women.(iii) To examine whether large language models (LLMs) internalize similar socio-pragmatic norms, we probe six LLMs using controlled generation and translation tasks over 1,000 culturally balanced entities.We find that LLMs diverge from Wikipedia usage, exhibiting alternative preferences in honorific selection across tasks, languages, and sociodemographic attributes.These discrepancies highlight gaps in the socio-cultural alignment of LLMs and open new directions for studying how LLMs acquire, adapt, or distort sociallinguistic norms. Sourabrata Mukherjee, Atharva Mehta, Sougata Saha, Akhil Arora 0001, Monojit Choudhury |
EMNLP | 5 |
| 2025 | Viability of Machine Translation for Healthcare in Low-Resourced LanguagesabstractHellina Hailu Nigatu, Nikita Mehandru, Negasi Haile Abadi, Blen Gebremeskel, Ahmed Alaa, Monojit Choudhury. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Hellina Nigatu, Nikita Mehandru, Negasi Haile Abadi, Blen Gebremeskel, Ahmed Alaa 0001, Monojit Choudhury |
EMNLP | 6 |
| 2025 | Exploring Adapter Design Tradeoffs for Low Resource Music GenerationabstractFine-tuning large-scale music audio generation models, such as MusicGen and Mustango, is a computationally expensive process, often requiring updates to billions of parameters and, therefore, significant hardware resources. Parameter-Efficient Fine-Tuning (PEFT) techniques, particularly adapter-based methods, have emerged as a promising alternative, enabling adaptation with minimal trainable parameters while preserving model performance. However, the design choices for adapters, including their architecture, placement, and size, are numerous, and it is unclear which of these combinations would produce optimal adapters and why, for a given case of low-resource music genre. In this paper, we attempt to answer this question by studying various adapter configurations for two AI music models, MusicGen and Mustango, on two genres: Hindustani Classical and Turkish Makam music. Atharva Mehta, Shivam Chauhan, Monojit Choudhury |
ACM Multimedia | 3 |
| 2025 | SMAB: MAB based word Sensitivity Estimation Framework and its Applications in Adversarial Text GenerationabstractSaurabh Kumar Pandey, Sachin Vashistha, Debrup Das, Somak Aditya, Monojit Choudhury. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Saurabh Kumar Pandey, Sachin Vashistha, Debrup Das, Somak Aditya, Monojit Choudhury |
NAACL (Long Papers) | 5 |
| 2025 | Meta-Cultural Competence: Climbing the Right Hill of Cultural AwarenessabstractSougata Saha, Saurabh Kumar Pandey, Monojit Choudhury. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sougata Saha, Saurabh Kumar Pandey, Monojit Choudhury |
NAACL (Long Papers) | 3 |
| 2025 | Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?abstractSougata Saha, Saurabh Kumar Pandey, Harshit Gupta, Monojit Choudhury. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sougata Saha, Saurabh Kumar Pandey, Monojit Choudhury |
NAACL (Long Papers) | 4 |
| 2025 | From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed SentencesabstractCurrent computational approaches for analysing or generating code-mixed sentences do not explicitly model “naturalness” or “acceptability” of code-mixed sentences, but rely on training corpora to reflect distribution of acceptable code-mixed sentences. Modelling human judgement for the acceptability of code-mixed text can help in distinguishing natural code-mixed text and enable quality-controlled generation of code-mixed text. To this end, we construct Cline—a dataset containing human acceptability judgements for English-Hindi (en-hi) code-mixed text. Cline is the largest of its kind with 16,642 sentences, consisting of samples sourced from two sources: synthetically generated code-mixed text and samples collected from online social media. Our analysis establishes that popular code-mixing metrics such as CMI, Number of Switch Points, Burstines, which are used to filter/curate/compare code-mixed corpora have low correlation with human acceptability judgements, underlining the necessity of our dataset. Experiments using Cline demonstrate that simple Multilayer Perceptron (MLP) models when trained solely using code-mixing metrics as features are outperformed by fine-tuned pre-trained Multilingual Large Language Models (MLLMs). Specifically, among Encoder models XLM-Roberta and Bernice outperform IndicBERT across different configurations. Among Encoder-Decoder models, mBART performs better than mT5, however, Encoder-Decoder models are not able to outperform Encoder-only models. Decoder-only models perform the best when compared with all other MLLMS, with Llama 3.2 - 3B models outperforming similarly sized Qwen, Phi models. Comparison with zero and fewshot capabilitites of ChatGPT show that MLLMs fine-tuned on larger data outperform ChatGPT, providing scope for improvement in code-mixed tasks. Zero-shot transfer from English–Hindi to English-Telugu acceptability judgments using our model checkpoints proves superior to random baselines, enabling application to other code-mixed language pairs and providing further avenues of research. We publicly release our human-annotated dataset, trained checkpoints, code-mix corpus, and code for data generation and model training. Prashant Kodali, Anmol Goel, Likhith Asapu, Vamshi Krishna Bonagiri, Anirudh Govil, Monojit Choudhury, Ponnurangam Kumaraguru, Manish Shrivastava 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2024 | Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language We Prompt Them inabstractEthical reasoning is a crucial skill for Large Language Models (LLMs). However, moral values are not universal, but rather influenced by language and culture. This paper explores how three prominent LLMs – GPT-4, ChatGPT, and Llama2Chat-70B – perform ethical reasoning in different languages and if their moral judgement depend on the language in which they are prompted. We extend the study of ethical reasoning of LLMs by (CITATION) to a multilingual setup following their framework of probing LLMs with ethical dilemmas and policies from three branches of normative ethics: deontology, virtue, and consequentialism. We experiment with six languages: English, Spanish, Russian, Chinese, Hindi, and Swahili. We find that GPT-4 is the most consistent and unbiased ethical reasoner across languages, while ChatGPT and Llama2Chat-70B show significant moral value bias when we move to languages other than English. Interestingly, the nature of this bias significantly vary across languages for all LLMs, including GPT-4. Utkarsh Agarwal, Kumar Tanmay, Aditi Khandelwal, Monojit Choudhury |
LREC/COLING | 4 |
| 2024 | INMT-Lite: Accelerating Low-Resource Language Data Collection via Offline Interactive Neural Machine TranslationabstractA steady increase in the performance of Massively Multilingual Models (MMLMs) has contributed to their rapidly increasing use in data collection pipelines. Interactive Neural Machine Translation (INMT) systems are one class of tools that can utilize MMLMs to promote such data collection in several under-resourced languages. However, these tools are often not adapted to the deployment constraints that native language speakers operate in, as bloated, online inference-oriented MMLMs trained for data-rich languages, drive them. INMT-Lite addresses these challenges through its support of (1) three different modes of Internet-independent deployment and (2) a suite of four assistive interfaces suitable for (3) data-sparse languages. We perform an extensive user study for INMT-Lite with an under-resourced language community, Gondi, to find that INMT-Lite improves the data generation experience of community members along multiple axes, such as cognitive load, task productivity, and interface interaction time and effort, without compromising on the quality of the generated translations.INMT-Lite’s code is open-sourced to further research in this domain. Harshita Diddee, Anurag Shukla, Tanuja Ganu, Vivek Seshadri, Sandipan Dandapat, Monojit Choudhury, Kalika Bali |
LREC/COLING | 6 |
| 2024 | Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting JailbreaksabstractRecent explorations with commercial Large Language Models (LLMs) have shown that non-expert users can jailbreak LLMs by simply manipulating their prompts; resulting in degenerate output behavior, privacy and security breaches, offensive outputs, and violations of content regulator policies. Limited studies have been conducted to formalize and analyze these attacks and their mitigations. We bridge this gap by proposing a formalism and a taxonomy of known (and possible) jailbreaks. We survey existing jailbreak methods and their effectiveness on open-source and commercial LLMs (such as GPT-based models, OPT, BLOOM, and FLAN-T5-XXL). We further discuss the challenges of jailbreak detection in terms of their effectiveness against known attacks. For further analysis, we release a dataset of model outputs across 3700 jailbreak prompts over 4 tasks. Abhinav Rao, Atharva Naik, Sachin Vashistha, Somak Aditya, Monojit Choudhury |
LREC/COLING | 5 |
| 2024 | Do Moral Judgment and Reasoning Capability of LLMs Change with Language? A Study using the Multilingual Defining Issues TestabstractAditi Khandelwal, Utkarsh Agarwal, Kumar Tanmay, Monojit Choudhury. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Aditi Khandelwal, Utkarsh Agarwal, Kumar Tanmay, Monojit Choudhury |
EACL (1) | 4 |
| 2024 | Towards Measuring and Modeling "Culture" in LLMs: A SurveyabstractMuhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Shivdutt Singh, Alham Fikri Aji, Jacki O’Neill, Ashutosh Modi, Monojit Choudhury. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh, Alham Fikri Aji, Jacki O'Neill, Ashutosh Modi, Monojit Choudhury |
EMNLP | 8 |
| 2024 | "They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated ConversationsabstractLarge language models (LLMs) have emerged as an integral part of modern societies, powering user-facing applications such as personal assistants and enterprise applications like recruitment tools.Despite their utility, research indicates that LLMs perpetuate systemic biases.Yet, prior works on LLM harms predominantly focus on Western concepts like race and gender, often overlooking cultural concepts from other parts of the world.Additionally, these studies typically investigate "harm" as a singular dimension, ignoring the various and subtle forms in which harms manifest.To address this gap, we introduce the Covert Harms and Social Threats (CHAST), a set of seven metrics grounded in social science literature.We utilize evaluation models aligned with human assessments to examine the presence of covert harms in LLM-generated conversations, particularly in the context of recruitment.Our experiments reveal that seven out of the eight LLMs included in this study generated conversations riddled with CHAST, characterized by malign views expressed in seemingly neutral language unlikely to be detected by existing methods.Notably, these LLMs manifested more extreme views and opinions when dealing with non-Western concepts like caste, compared to Western ones such as race.Warning: This paper has instances of offensive language to serve as examples. Preetam Prabhu Srikar Dammu, Hayoung Jung, Monojit Choudhury, Tanushree Mitra |
EMNLP | 4 |
| 2024 | Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic PromptingabstractSagnik Mukherjee, Muhammad Farid Adilazuarda, Sunayana Sitaram, Kalika Bali, Alham Fikri Aji, Monojit Choudhury. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Sagnik Mukherjee, Muhammad Farid Adilazuarda, Sunayana Sitaram, Kalika Bali, Alham Fikri Aji, Monojit Choudhury |
EMNLP | 6 |
| 2024 | The Zeno's Paradox of 'Low-Resource' LanguagesabstractThe disparity in the languages commonly studied in Natural Language Processing (NLP) is typically reflected by referring to languages as low vs high-resourced.However, there is limited consensus on what exactly qualifies as a 'low-resource language.'To understand how NLP papers define and study 'low resource' languages, we qualitatively analyzed 150 papers from the ACL Anthology and popular speechprocessing conferences that mention the keyword 'low-resource.' Based on our analysis, we show how several interacting axes contribute to 'low-resourcedness' of a language and why that makes it difficult to track progress for each individual language.We hope our work (1) elicits explicit definitions of the terminology when it is used in papers and (2) provides grounding for the different axes to consider when connoting a language as low-resource. Hellina Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio, Monojit Choudhury |
EMNLP | 5 |
| 2023 | DiTTO: A Feature Representation Imitation Approach for Improving Cross-Lingual TransferabstractShanu Kumar, Soujanya Abbaraju, Sandipan Dandapat, Sunayana Sitaram, Monojit Choudhury. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Shanu Kumar, Abbaraju Soujanya, Sandipan Dandapat, Sunayana Sitaram, Monojit Choudhury |
EACL | 5 |
| 2023 | LLM-powered Data Augmentation for Enhanced Cross-lingual PerformanceabstractThis paper explores the potential of leveraging Large Language Models (LLMs) for data augmentation in multilingual commonsense reasoning datasets where the available training data is extremely limited.To achieve this, we utilise several LLMs, namely Dolly-v2, Sta-bleVicuna, ChatGPT, and GPT-4, to augment three datasets: XCOPA, XWinograd, and XS-toryCloze.Subsequently, we evaluate the effectiveness of fine-tuning smaller multilingual models, mBERT and XLMR, using the synthesised data.We compare the performance of training with data generated in English and target languages, as well as translated Englishgenerated data, revealing the overall advantages of incorporating data generated by LLMs, e.g. a notable 13.4 accuracy score improvement for the best case.Furthermore, we conduct a human evaluation by asking native speakers to assess the naturalness and logical coherence of the generated examples across different languages.The results of the evaluation indicate that LLMs such as ChatGPT and GPT-4 excel at producing natural and coherent text in most languages, however, they struggle to generate meaningful text in certain languages like Tamil.We also observe that ChatGPT falls short in generating plausible alternatives compared to the original dataset, whereas examples from GPT-4 exhibit competitive logical consistency.We release the generated data at https://github.com/mbzuai-nlp/Gen-X. Chenxi Whitehouse, Monojit Choudhury, Alham Fikri Aji |
EMNLP | 2 |
| 2023 | Prover: Generating Intermediate Steps for NLI with Commonsense Knowledge Retrieval and Next-Step PredictionabstractDeepanway Ghosal, Somak Aditya, Monojit Choudhury. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Deepanway Ghosal, Somak Aditya, Monojit Choudhury |
IJCNLP (1) | 3 |
| 2022 | LITMUS Predictor: An AI Assistant for Building Reliable, High-Performing and Fair Multilingual NLP SystemsabstractPre-trained multilingual language models are gaining popularity due to their cross-lingual zero-shot transfer ability, but these models do not perform equally well in all languages. Evaluating task-specific performance of a model in a large number of languages is often a challenge due to lack of labeled data, as is targeting improvements in low performing languages through few-shot learning. We present a tool - LITMUS Predictor - that can make reliable performance projections for a fine-tuned task-specific model in a set of languages without test and training data, and help strategize data labeling efforts to optimize performance and fairness objectives. Anirudh Srinivasan, Gauri Kholkar, Rahul Kejriwal, Tanuja Ganu, Sandipan Dandapat, Sunayana Sitaram, Balakrishnan Santhanam, Somak Aditya, Kalika Bali, Monojit Choudhury |
AAAI | 10 |
| 2022 | Multi Task Learning For Zero Shot Performance Prediction of Multilingual ModelsabstractMassively Multilingual Transformer based Language Models have been observed to be surprisingly effective on zero-shot transfer across languages, though the performance varies from language to language depending on the pivot language(s) used for fine-tuning.In this work, we build upon some of the existing techniques for predicting the zero-shot performance on a task, by modeling it as a multi-task learning problem.We jointly train predictive models for different tasks which helps us build more accurate predictors for tasks where we have test data in very few languages to measure the actual performance of the model.Our approach also lends us the ability to perform a much more robust feature selection, and identify a common set of features that influence zero-shot performance across a variety of tasks. Kabir Ahuja, Shanu Kumar, Sandipan Dandapat, Monojit Choudhury |
ACL (1) | 4 |
| 2022 | Global Readiness of Language Technology for Healthcare: What Would It Take to Combat the Next Pandemic?abstractThe COVID-19 pandemic has brought out both the best and worst of language technology (LT). On one hand, conversational agents for information dissemination and basic diagnosis have seen widespread use, and arguably, had an important role in fighting against the pandemic. On the other hand, it has also become clear that such technologies are readily available for a handful of languages, and the vast majority of the global south is completely bereft of these benefits. What is the state of LT, especially conversational agents, for healthcare across the world’s languages? And, what would it take to ensure global readiness of LT before the next pandemic? In this paper, we try to answer these questions through survey of existing literature and resources, as well as through a rapid chatbot building exercise for 15 Asian and African languages with varying amount of resource-availability. The study confirms the pitiful state of LT even for languages with large speaker bases, such as Sinhala and Hausa, and identifies the gaps that could help us prioritize research and investment strategies in LT for healthcare. Ishani Mondal, Kabir Ahuja, Jacki O'Neill, Kalika Bali, Monojit Choudhury |
COLING | 6 |
| 2022 | The Six Conundrums of Building and Deploying Language Technologies for Social GoodabstractDeployment of speech and language technology for social good (LT4SG), especially those targeted at the welfare of marginalized communities and speakers of low-resource and under-served languages, has been a prominent theme of research within NLP, Speech and the AI communities. Many researchers, especially those working in core NLP/Speech domains, rely on a combination of individual expertise, experiences or ad hoc surveys for prioritizing between language technologies that provide social good to the end-users. This has been criticized by several scholars who argue that it is critical to include the target community during the LT’s design and development process. However, prioritization of communities, languages, technologies and design approaches presents a very large set of complex challenges to the technologists, for which there are no simple or off-the-shelf solutions. In this position paper, we distill our experiential insights into six fundamental conundrums that technologists face and must resolve while deciding which LT technology to build for which community, and by using what approach. We discuss that at the root of these conundrums lie certain fundamental ethical problems of a digital-divide that can be overcome only by resolving deeper ethical dilemmas of distributive justice. We urge the community to reflect on these conundrums and leverage shared experiential insights to reconcile the intent of broadly, any Technology for Social Good, with the ground realities of its deployment. Harshita Diddee, Kalika Bali, Monojit Choudhury, Namrata Mukhija |
COMPASS | 3 |
| 2022 | On the Calibration of Massively Multilingual Language ModelsabstractMassively Multilingual Language Models (MMLMs) have recently gained popularity due to their surprising effectiveness in cross-lingual transfer.While there has been much work in evaluating these models for their performance on a variety of tasks and languages, little attention has been paid on how well calibrated these models are with respect to the confidence in their predictions.We first investigate the calibration of MMLMs in the zero-shot setting and observe a clear case of miscalibration in low-resource languages or those which are typologically diverse from English.Next, we empirically show that calibration methods like temperature scaling and label smoothing do reasonably well in improving calibration in the zero-shot scenario.We also find that fewshot examples in the language can further help reduce calibration errors, often substantially.Overall, our work contributes towards building more reliable multilingual models by highlighting the issue of their miscalibration, understanding what language and model-specific factors influence it, and pointing out the strategies to improve the same. Kabir Ahuja, Sunayana Sitaram, Sandipan Dandapat, Monojit Choudhury |
EMNLP | 4 |
| 2022 | Language Patterns and Behaviour of the Peer Supporters in Multilingual Healthcare Conversational ForumsabstractIn this work, we conduct a quantitative linguistic analysis of the language usage patterns of multilingual peer supporters in two health-focused WhatsApp groups in Kenya comprising of youth living with HIV. Even though the language of communication for the group was predominantly English, we observe frequent use of Kiswahili, Sheng and code-mixing among the three languages. We present an analysis of language choice and its accommodation, different functions of code-mixing, and relationship between sentiment and code-mixing. To explore the effectiveness of off-the-shelf Language Technologies (LT) in such situations, we attempt to build a sentiment analyzer for this dataset. Our experiments demonstrate the challenges of developing LT and therefore effective interventions for such forums and languages. We provide recommendations for language resources that should be built to address these challenges. Ishani Mondal, Kalika Bali, Monojit Choudhury, Jacki O'Neill, Millicent Ochieng, Kagonya Awori, Keshet Ronen |
LREC | 4 |
| 2022 | On the Economics of Multilingual Few-shot Learning: Modeling the Cost-Performance Trade-offs of Machine Translated and Manual DataabstractBorrowing ideas from Production functions in micro-economics, in this paper we introduce a framework to systematically evaluate the performance and cost trade-offs between machinetranslated and manually-created labelled data for task-specific fine-tuning of massively multilingual language models.We illustrate the effectiveness of our framework through a casestudy on the TyDIQA-GoldP dataset.One of the interesting conclusions of the study is that if the cost of machine translation is greater than zero, the optimal performance at least cost is always achieved with at least some or only manually-created data.To our knowledge, this is the first attempt towards extending the concept of production functions to study data collection strategies for training multilingual models, and can serve as a valuable tool for other similar cost vs data trade-offs in NLP. Kabir Ahuja, Monojit Choudhury, Sandipan Dandapat |
NAACL-HLT | 2 |
| 2021 | How Linguistically Fair Are Multilingual Pre-Trained Language Models?abstractMassively multilingual pre-trained language models, such as mBERT and XLM-RoBERTa, have received significant attention in the recent NLP literature for their excellent capability towards crosslingual zero-shot transfer of NLP tasks. This is especially promising because a large number of languages have no or very little labeled data for supervised learning. Moreover, a substantially improved performance on low resource languages without any significant degradation of accuracy for high resource languages lead us to believe that these models will help attain a fairer distribution of language technologies despite the prevalent unfair and extremely skewed distribution of resources across the world’s languages. Nevertheless, these models, and the experimental approaches adopted by the researchers to arrive at those, have been criticised by some for lacking a nuanced and thorough comparison of benefits across languages and tasks. A related and important question that has received little attention is how to choose from a set of models, when no single model significantly outperforms the others on all tasks and languages. As we discuss in this paper, this is often the case, and the choices are usually made without a clear articulation of reasons or underlying fairness assumptions. In this work, we scrutinize the choices made in previous work, and propose a few different strategies for fair and efficient model selection based on the principles of fairness in economics and social choice theory. In particular, we emphasize Rawlsian fairness, which provides an appropriate framework for making fair (with respect to languages, or tasks, or both) choices while selecting multilingual pre-trained language models for a practical or scientific set-up. Monojit Choudhury |
AAAI | 1 |
| 2021 | Language Translation as a Socio-Technical System: Case-Studies of Mixed-Initiative InteractionsabstractSeamless access to information in a rapidly globalizing world demands for availability of information across, ideally all but at the least a large number of, languages. Machine translation has been proposed as a technological solution to this complex problem. However, despite seven decades of research, and recently seen rapid progress in the field - thanks to deep learning and availability of large data-sets, perfect machine translation across a large number of the world’s languages still remains elusive. In fact, it is a distant and perhaps even an impossible goal. Erroneous translations, on the other hand, can be detrimental in critical situations such as talking to a law enforcement officer; or, they could potentially perpetuate social biases or stereotypes, for instance, by producing mis-gendered translations. In this work, we argue that language translation is inherently a socio-technical system, which has to be viewed, studied, and optimized for, as such. The need and context of translation, the socio-demographic factors behind the human translators as well as the consumers of the translated content affect the complexity of the translation system, as much as the accuracy of the technology and its interface. Through a series of case studies on mixed-initiative interaction based approach to translation, we bring out the various socio-technical factors and their complex interactions that one has to bear in mind while designing for the ideal human-machine translation systems. Through these observations, we make multiple recommendations which, at the core, suggest that ”solving” translation in the real sense would require more coordinated efforts between the technical (NLP) and social communities (HCI + CSCW + DEV). Sebastin Santy, Kalika Bali, Monojit Choudhury, Sandipan Dandapat, Tanuja Ganu, Anurag Shukla, Jahanvi Shah, Vivek Seshadri |
COMPASS | 3 |
| 2021 | American Politicians Diverge Systematically, Indian Politicians do so Chaotically: Text Embeddings as a Window into Party Polarization
Amar Budhiraja, Rahul Agrawal, Monojit Choudhury, Joyojeet Pal |
ICWSM | 4 |
| 2020 | The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldabstractLanguage technologies contribute to promoting multilingualism and linguistic diversity around the world.However, only a very small number of the over 7000 languages of the world are represented in the rapidly evolving language technologies and applications.In this paper we look at the relation between the types of languages, resources, and their representation in NLP conferences to understand the trajectory that different languages have followed over time.Our quantitative investigation underlines the disparity between languages, especially in terms of their resources, and calls into question the "language agnostic" status of current models and systems.Through this paper, we attempt to convince the ACL community to prioritise the resolution of the predicaments highlighted here, so that no language is left behind. Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, Monojit Choudhury |
ACL | 5 |
| 2020 | GLUECoS: An Evaluation Benchmark for Code-Switched NLPabstractCode-switching is the use of more than one language in the same conversation or utterance.Recently, multilingual contextual embedding models, trained on multiple monolingual corpora, have shown promising results on cross-lingual and multilingual tasks.We present an evaluation benchmark, GLUECoS, for code-switched languages, that spans several NLP tasks in English-Hindi and English-Spanish.Specifically, our evaluation benchmark includes Language Identification from text, POS tagging, Named Entity Recognition, Sentiment Analysis, Question Answering and a new task for code-switching, Natural Language Inference.We present results on all these tasks using cross-lingual word embedding models and multilingual models.In addition, we fine-tune multilingual models on artificially generated code-switched data.Although multilingual models perform significantly better than cross-lingual models, our results show that in most tasks, across both language pairs, multilingual models fine-tuned on code-switched data perform best, showing that multilingual models can be further optimized for code-switching tasks. Simran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram, Monojit Choudhury |
ACL | 5 |
| 2020 | TaxiNLI: Taking a Ride up the NLU HillabstractPre-trained Transformer-based neural architectures have consistently achieved state-of-theart performance in the Natural Language Inference (NLI) task.Since NLI examples encompass a variety of linguistic, logical, and reasoning phenomena, it remains unclear as to which specific concepts are learnt by the trained systems and where they can achieve strong generalization.To investigate this question, we propose a taxonomic hierarchy of categories that are relevant for the NLI task.We introduce TAXINLI, a new dataset, that has 10k examples from the MNLI dataset (Williams et al., 2018) with these taxonomic labels.Through various experiments on TAXINLI, we observe that whereas for certain taxonomic categories SOTA neural models have achieved near perfect accuracies-a large jump over the previous models-some categories still remain difficult.Our work adds to the growing body of literature that shows the gaps in the current NLI systems and datasets through a systematic presentation and analysis of reasoning categories. Pratik Joshi, Somak Aditya, Aalok Sathe, Monojit Choudhury |
CoNLL | 4 |
| 2020 | Engagement Patterns of Peer-to-Peer Interactions on Mental Health Platforms
Ashish Sharma 0004, Monojit Choudhury, Tim Althoff, Amit Sharma 0007 |
ICWSM | 2 |
| 2020 | Crowdsourcing Speech Data for Low-Resource Languages from Low-Income WorkersabstractVoice-based technologies are essential to cater to the hundreds of millions of new smartphone users. However, most of the languages spoken by these new users have little to no labelled speech data. Unfortunately, collecting labelled speech data in any language is an expensive and resource-intensive task. Moreover, existing platforms typically collect speech data only from urban speakers familiar with digital technology whose dialects are often very different from low-income users. In this paper, we explore the possibility of collecting labelled speech data directly from low-income workers. In addition to providing diversity to the speech dataset, we believe this approach can also provide valuable supplemental earning opportunities to these communities. To this end, we conducted a study where we collected labelled speech data in the Marathi language from three different user groups: low-income rural users, low-income urban users, and university students. Overall, we collected 109 hours of data from 36 participants. Our results show that the data collected from low-income participants is of comparable quality to the data collected from university students (who are typically employed to do this work) and that crowdsourcing speech data from low-income rural and urban workers is a viable method of gathering speech data. Basil Abraham, Danish Goel, Divya Siddarth, Kalika Bali, Manu Chopra, Monojit Choudhury, Pratik Joshi, Preethi Jyothi, Sunayana Sitaram, Vivek Seshadri |
LREC | 6 |
| 2020 | Do Multilingual Users Prefer Chat-bots that Code-mix? Let's Nudge and Find Out!abstractDespite their pervasiveness, current text-based conversational agents (chatbots) are predominantly monolingual, while users are often multilingual. It is well-known that multilingual users mix languages while interacting with others, as well as in their interactions with computer systems (such as query formulation in text-/voice-based search interfaces and digital assistants). Linguists refer to this phenomenon as code-mixing or code-switching. Do multilingual users also prefer chatbots that can respond in a code-mixed language over those which cannot? In order to inform the design of chatbots for multilingual users, we conduct a mixed-method user-study (N=91) where we examine how conversational agents, that code-mix and reciprocate the users' mixing choices over multiple conversation turns, are evaluated and perceived by bilingual users. We design a human-in-the-loop chatbot with two different code-mixing policies -- (a) always code-mix irrespective of user behavior, and (b) nudge with subtle code-mixed cues and reciprocate only if the user, in turn, code-mixes. These two are contrasted with a monolingual chatbot that never code-mixed. Users are asked to interact with the bots, and provide ratings on perceived naturalness and personal preference. They are also asked open-ended questions around what they (dis)liked about the bots. Analysis of the chat logs, users' ratings, and qualitative responses reveal that multilingual users strongly prefer chatbots that can code-mix. We find that self-reported language proficiency is the strongest predictor of user preferences. Compared to the Always code-mix policy, Nudging emerges as a low-risk low-gain policy which is equally acceptable to all users. Nudging as a policy is further supported by the observation that users who rate the code-mixing bot higher typically tend to reciprocate the language mixing pattern of the bot. These findings present a first step towards developing conversational systems that are more human-like and engaging by virtue of adapting to the users' linguistic style. Anshul Bawa, Pranav Khadpe, Pratik Joshi, Kalika Bali, Monojit Choudhury |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2020 | Topical Focus of Political Campaigns and its Impact: Findings from Politicians' Hashtag Use during the 2019 Indian ElectionsabstractWe studied the topical preferences of social media campaigns of India's two main political parties by examining the tweets of 7382 politicians during the key phase of campaigning between Jan - May of 2019 in the run up to the 2019 general election. First, we compare the use of self-promotion and opponent attack, and their respective success online by categorizing 1208 most commonly used hashtags accordingly into the two categories. Second, we classify the tweets applying a qualitative typology to hashtags on the subjects of nationalism, corruption, religion and development. We find that the ruling BJP tended to promote itself over attacking the opposition whereas the main challenger INC was more likely to attack than promote itself. Moreover, while the INC gets more retweets on average, the BJP dominates Twitter's trends by flooding the online space with large numbers of tweets. We consider the implications of our findings hold for political communication strategies in democracies across the world. Anmol Panda, Ramaravind Kommiya Mothilal, Monojit Choudhury, Kalika Bali, Joyojeet Pal |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2019 | Identifying and Analyzing Different Aspects of English-Hindi Code-Switching in TwitterabstractCode-switching or the juxtaposition of linguistic units from two or more languages in a single utterance, has, in recent times, become very common in text, thanks to social media and other computer mediated forms of communication. In this exploratory study of English-Hindi code-switching on Twitter, we automatically create a large corpus of code-switched tweets and devise techniques to identify the relationship between successive components in a code-switched tweet. More specifically, we identify pragmatic functions such as narrative-evaluative, negative reinforcement, translation or semantically equivalent statements, and so on characterizing the relation between successive components. We analyze the difference/similarity between switching patterns in code-switched and monolingual multi-component tweets. We observe strong dominance of narrative-evaluative (non-opinion to opinion or vice versa) switching in case of both code-switched and monolingual multi-component tweets in around 40% of cases. Polarity switching appears to be a prevalent switching phenomenon (10%) specifically in code-switched tweets (three to four times higher than monolingual multi-component tweets) where preference of expressing negative sentiment in Hindi is approximately twice compared to English. Positive reinforcement appears to be an important pragmatic function for English multi-component tweets, whereas negative reinforcement plays a key role for Devanagari multi-component tweets. Our results also indicate that the extent and nature of code-switching also strongly depend on the topic (sports, politics, etc.) of discussion. Koustav Rudra, Ashish Sharma 0004, Kalika Bali, Monojit Choudhury, Niloy Ganguly |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2018 | Language Modeling for Code-Mixing: The Role of Linguistic Theory based Synthetic DataabstractAdithya Pratapa, Gayatri Bhat, Monojit Choudhury, Sunayana Sitaram, Sandipan Dandapat, Kalika Bali. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Adithya Pratapa, Gayatri Bhat, Monojit Choudhury, Sunayana Sitaram, Sandipan Dandapat, Kalika Bali |
ACL (1) | 3 |
| 2018 | Word Embeddings for Code-Mixed Language ProcessingabstractWe compare three existing bilingual word embedding approaches, and a novel approach of training skip-grams on synthetic code-mixed text generated through linguistic models of code-mixing, on two tasks -sentiment analysis and POS tagging for code-mixed text.Our results show that while CVM and CCA based embeddings perform as well as the proposed embedding technique on semantic and syntactic tasks respectively, the proposed approach provides the best performance for both tasks overall.Thus, this study demonstrates that existing bilingual embedding techniques are not ideal for code-mixed text processing and there is a need for learning multilingual word embedding from the code-mixed text. Adithya Pratapa, Monojit Choudhury, Sunayana Sitaram |
EMNLP | 2 |
| 2018 | An Integrated Representation of Linguistic and Social Functions of Code-Switching
Silvana Hartmann, Monojit Choudhury, Kalika Bali |
LREC | 2 |
| 2018 | Discovering Canonical Indian English Accents: A Crowdsourcing-based Approach
Sunayana Sitaram, Varun Manjunath, Varun Bharadwaj, Monojit Choudhury, Kalika Bali, Michael Tjalve |
LREC | 4 |
| 2017 | Estimating Code-Switching on Twitter with a Novel Generalized Word-Level Language Detection TechniqueabstractShruti Rijhwani, Royal Sequiera, Monojit Choudhury, Kalika Bali, Chandra Shekhar Maddila. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Shruti Rijhwani, Royal Sequiera, Monojit Choudhury, Kalika Bali, Chandra Shekhar Maddila |
ACL (1) | 3 |
| 2017 | All that is English may be Hindi: Enhancing language identification through automatic ranking of the likeliness of word borrowing in social mediaabstractIn this paper, we present a set of computational methods to identify the likeliness of a word being borrowed, based on the signals from social media. In terms of Spearman correlation coefficient values, our methods perform more than two times better (nearly 0.62) in predicting the borrowing likeliness compared to the best performing baseline (nearly 0.26) reported in literature. Based on this likeliness estimate we asked annotators to re-annotate the language tags of foreign words in predominantly native contexts. In 88 percent of cases the annotators felt that the foreign language tag should be replaced by native language tag, thus indicating a huge scope for improvement of automatic language identification systems. Jasabanta Patro, Bidisha Samanta, Saurabh Singh 0003, Abhipsa Basu, Prithwish Mukherjee, Monojit Choudhury, Animesh Mukherjee 0001 |
EMNLP | 6 |
| 2016 | Improving Document Ranking for Long Queries with Nested Query Segmentation
Rishiraj Saha Roy, Anusha Suresh, Niloy Ganguly, Monojit Choudhury |
ECIR | 4 |
| 2016 | Understanding Language Preference for Expression of Opinion and Sentiment: What do Hindi-English Speakers do on Twitter?abstractLinguistic research on multilingual societies has indicated that there is usually a preferred language for expression of emotion and sentiment (Dewaele, 2010).Paucity of data has limited such studies to participant interviews and speech transcriptions from small groups of speakers.In this paper, we report a study on 430,000 unique tweets from Indian users, specifically Hindi-English bilinguals, to understand the language of preference, if any, for expressing opinion and sentiment.To this end, we develop classifiers for opinion detection in these languages, and further classifying opinionated tweets into positive, negative and neutral sentiments.Our study indicates that Hindi (i.e., the native language) is preferred over English for expression of negative opinion and swearing.As an aside, we explore some common pragmatic functions of codeswitching through sentiment detection.* * This work was done when the author was a Research Fellow at Microsoft Research Lab India.1 Although some linguists differentiate between Codeswitching and Code-mixing, this paper will use the two terms interchangeably. Koustav Rudra, Shruti Rijhwani, Rafiya Begum, Kalika Bali, Monojit Choudhury, Niloy Ganguly |
EMNLP | 5 |
| 2016 | Functions of Code-Switching in Tweets: An Annotation Framework and Some Initial Experiments
Rafiya Begum, Kalika Bali, Monojit Choudhury, Koustav Rudra, Niloy Ganguly |
LREC | 3 |
| 2016 | Syntactic complexity of Web search queries through the lenses of language models, networks and users
Rishiraj Saha Roy, Smith Agarwal, Niloy Ganguly, Monojit Choudhury |
Inf. Process. Manag. | 4 |
| 2015 | Discovering and understanding word level user intent in Web search queries
Rishiraj Saha Roy, Rahul Katare, Niloy Ganguly, Srivatsan Laxman, Monojit Choudhury |
J. Web Semant. | 5 |
| 2014 | Automatic Discovery of Adposition Typology
Rishiraj Saha Roy, Rahul Katare, Niloy Ganguly, Monojit Choudhury |
COLING | 4 |
| 2014 | POS Tagging of English-Hindi Code-Mixed Social Media ContentabstractCode-mixing is frequently observed in user generated content on social media, especially from multilingual users. The linguistic complexity of such content is compounded by presence of spelling vari-ations, transliteration and non-adherance to formal grammar. We describe our initial efforts to create a multi-level an-notated corpus of Hindi-English code-mixed text collated from Facebook fo-rums, and explore language identifica-tion, back-transliteration, normalization and POS tagging of this data. Our re-sults show that language identification and transliteration for Hindi are two major challenges that impact POS tagging accu-racy. 1 Yogarshi Vyas, Spandana Gella, Kalika Bali, Monojit Choudhury |
EMNLP | 5 |
| 2014 | Query expansion for mixed-script information retrievalabstractFor many languages that use non-Roman based indigenous scripts (e.g., Arabic, Greek and Indic languages) one can often find a large amount of user generated transliterated content on the Web in the Roman script. Such content creates a monolingual or multi-lingual space with more than one script which we refer to as the Mixed-Script space. IR in the mixed-script space is challenging because queries written in either the native or the Roman script need to be matched to the documents written in both the scripts. Moreover, transliterated content features extensive spelling variations. In this paper, we formally introduce the concept of Mixed-Script IR, and through analysis of the query logs of Bing search engine, estimate the prevalence and thereby establish the importance of this problem. We also give a principled solution to handle the mixed-script term matching and spelling variation where the terms across the scripts are modelled jointly in a deep-learning architecture and can be compared in a low-dimensional abstract space. We present an extensive empirical analysis of the proposed method along with the evaluation results in an ad-hoc retrieval setting of mixed-script IR where the proposed method achieves significantly better results (12% increase in MRR and 29% increase in MAP) compared to other state-of-the-art baselines. Parth Gupta, Kalika Bali, Rafael E. Banchs, Monojit Choudhury, Paolo Rosso |
SIGIR | 4 |
| 2014 | Improving unsupervised query segmentation using parts-of-speech sequence informationabstractWe present a generic method for augmenting unsupervised query segmentation by incorporating Parts-of-Speech (POS) sequence information to detect meaningful but rare n-grams. Our initial experiments with an existing English POS tagger employing two different POS tagsets and an unsupervised POS induction technique specifically adapted for queries show that POS information can significantly improve query segmentation performance in all these cases. Rishiraj Saha Roy, Yogarshi Vyas, Niloy Ganguly, Monojit Choudhury |
SIGIR | 4 |
| 2013 | Crowd Prefers the Middle Path: A New IAA Metric for Crowdsourcing Reveals Turker Biases in Query Segmentation
Rohan Ramanath, Monojit Choudhury, Kalika Bali, Rishiraj Saha Roy |
ACL (1) | 2 |
| 2012 | Can Modern Statistical Parsers Lead to Better Natural Language Understanding for Education?
Umair Z. Ahmed, Arpit Kumar, Monojit Choudhury, Kalika Bali |
CICLing (1) | 3 |
| 2012 | Mining Hindi-English Transliteration Pairs from Online Hindi Lyrics
Kanika Gupta, Monojit Choudhury, Kalika Bali |
LREC | 2 |
| 2012 | An Empirical Study of the Occurrence and Co-Occurrence of Named Entities in Natural Language Corpora
K. Saravanan 0001, Monojit Choudhury, Raghavendra Udupa, A. Kumaran 0001 |
LREC | 2 |
| 2012 | An IR-based evaluation framework for web search query segmentationabstractThis paper presents the first evaluation framework for Web search query segmentation based directly on IR performance. In the past, segmentation strategies were mainly validated against manual annotations. Our work shows that the goodness of a segmentation algorithm as judged through evaluation against a handful of human annotated segmentations hardly reflects its effectiveness in an IR-based setup. In fact, state-of the-art algorithms are shown to perform as good as, and sometimes even better than human annotations a fact masked by previous validations. The proposed framework also provides us an objective understanding of the gap between the present best and the best possible segmentation algorithm. We draw these conclusions based on an extensive evaluation of six segmentation strategies, including three most recent algorithms, vis-a-vis segmentations from three human annotators. The evaluation framework also gives insights about which segments should be necessarily detected by an algorithm for achieving the best retrieval results. The meticulously constructed dataset used in our experiments has been made public for use by the research community. Rishiraj Saha Roy, Niloy Ganguly, Monojit Choudhury, Srivatsan Laxman |
SIGIR | 3 |
| 2011 | Network based models of cognitive and social dynamics of human languages
Animesh Mukherjee 0001, Monojit Choudhury, Samer Hassan 0002, Smaranda Muresan |
Comput. Speech Lang. | 2 |
| 2010 | Resource Creation for Training and Testing of Transliteration Systems for Indian Languages
Sowmya V. B., Monojit Choudhury, Kalika Bali, Tirthankar Dasgupta, Anupam Basu |
LREC | 2 |
| 2009 | Large-Coverage Root Lexicon Extraction for Hindi
Cohan Sujay Carlos, Monojit Choudhury, Sandipan Dandapat |
EACL | 2 |
| 2009 | Discovering Global Patterns in Linguistic Networks through Spectral Analysis: A Case Study of the Consonant Inventories
Animesh Mukherjee 0001, Monojit Choudhury, Ravi Kannan |
EACL | 2 |
| 2008 | Modeling the Structure and Dynamics of the Consonant Inventories: A Complex Network Approach
Animesh Mukherjee 0001, Monojit Choudhury, Anupam Basu, Niloy Ganguly |
COLING | 2 |
| 2008 | Invited Talk: Breaking the Zipfian Barrier of NLP
Monojit Choudhury |
IJCNLP | 1 |
| 2008 | Social Network Inspired Models of NLP and Language Evolution
Monojit Choudhury, Animesh Mukherjee 0001, Niloy Ganguly |
IJCNLP | 1 |
| 2008 | Unsupervised Parts-of-Speech Induction for Bengali
Joy Deep Nath, Monojit Choudhury, Animesh Mukherjee 0001, Chris Biemann, Niloy Ganguly |
LREC | 2 |
| 2008 | A Common Parts-of-Speech Tagset Framework for Indian Languages
Baskaran Sankaran, Kalika Bali, Monojit Choudhury, Tanmoy Bhattacharya 0003, Pushpak Bhattacharyya, Girish Nath Jha, K. Saravanan 0001, L. Sobha, Karumuri V. Subbarao |
LREC | 3 |
| 2007 | Redundancy Ratio: An Invariant Property of the Consonant Inventories of the World's Languages
Animesh Mukherjee 0001, Monojit Choudhury, Anupam Basu, Niloy Ganguly |
ACL | 2 |
| 2007 | Investigation and modeling of the structure of texting language
Monojit Choudhury, Rahul Saraf, Vijit Jain, Animesh Mukherjee 0001, Sudeshna Sarkar, Anupam Basu |
Int. J. Document Anal. Recognit. | 1 |
| 2006 | Analysis and Synthesis of the Distribution of Consonants over Languages: A Complex Network Approach
Monojit Choudhury, Animesh Mukherjee 0001, Anupam Basu, Niloy Ganguly |
ACL | 1 |
| 2006 | Battery-aware code partitioning for a text to speech systemabstractThe advent of multi-core embedded processors has brought along new challenges for embedded system design. This paper presents an efficient, battery aware, code partitioning technique for a text to speech system, which is executed on a multi-core embedded processor. The system achieves significant performance improvements both in terms of execution time as well as battery lifetimes. The mentioned technique provides a new paradigm for battery aware embedded system design which can be easily extended to other applications Anirban Lahiri, Anupam Basu, Monojit Choudhury, Srobona Mitra |
DATE | 3 |