EDBT 2026 Demo / reviewers in the wild / expert
Anoop Kunchukuttan
dblp:126/8631
· DBLP profile ↗
26ranked-venue papers
8as first author
17since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cross-Lingual Auto Evaluation for Assessing Multilingual LLMsabstractSumanth Doddapaneni, Mohammed Safi Ur Rahman Khan, Dilip Venkatesh, Raj Dabre, Anoop Kunchukuttan, Mitesh M Khapra. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Sumanth Doddapaneni, Mohammed Safi Ur Rahman Khan, Dilip Venkatesh, Raj Dabre, Anoop Kunchukuttan, Mitesh M. Khapra |
ACL (1) | 5 |
| 2025 | Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian LanguagesabstractAshwin Sankar, Sparsh Jain, Nikhil Narasimhan, Devilal Choudhary, Dhairya Suman, Mohammed Safi Ur Rahman Khan, Anoop Kunchukuttan, Mitesh M Khapra, Raj Dabre. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ashwin Sankar, Sparsh Jain, Nikhil Narasimhan, Devilal Choudhary, Dhairya Suman, Mohammed Safi Ur Rahman Khan, Anoop Kunchukuttan, Mitesh M. Khapra, Raj Dabre |
ACL (1) | 7 |
| 2024 | RomanSetu: Efficiently unlocking multilingual capabilities of Large Language Models via RomanizationabstractJaavid J, Raj Dabre, Aswanth M, Jay Gala, Thanmay Jayakumar, Ratish Puduppully, Anoop Kunchukuttan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jaavid Aktar Husain, Raj Dabre, Aswanth M., Jay Gala, Thanmay Jayakumar, Ratish Puduppully, Anoop Kunchukuttan |
ACL (1) | 7 |
| 2024 | IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian LanguagesabstractMohammed Safi Ur Rahman Khan, Priyam Mehta, Ananth Sankar, Umashankar Kumaravelan, Sumanth Doddapaneni, Suriyaprasaad B, Varun G, Sparsh Jain, Anoop Kunchukuttan, Pratyush Kumar, Raj Dabre, Mitesh M. Khapra. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Mohammed Safi Ur Rahman Khan, Priyam Mehta, Ananth Sankar, Umashankar Kumaravelan, Sumanth Doddapaneni, Suriyaprasaad B, Varun Balan G, Sparsh Jain, Anoop Kunchukuttan, Raj Dabre, Mitesh M. Khapra |
ACL (1) | 9 |
| 2024 | Synthetic Data Generation and Joint Learning for Robust Code-Mixed TranslationabstractThe widespread online communication in a modern multilingual world has provided opportunities to blend more than one language (aka code-mixed language) in a single utterance. This has resulted a formidable challenge for the computational models due to the scarcity of annotated data and presence of noise. A potential solution to mitigate the data scarcity problem in low-resource setup is to leverage existing data in resource-rich language through translation. In this paper, we tackle the problem of code-mixed (Hinglish and Bengalish) to English machine translation. First, we synthetically develop HINMIX, a parallel corpus of Hinglish to English, with ~4.2M sentence pairs. Subsequently, we propose RCMT, a robust perturbation based joint-training model that learns to handle noise in the real-world code-mixed text by parameter sharing across clean and noisy words. Further, we show the adaptability of RCMT in a zero-shot setup for Bengalish to English translation. Our evaluation and comprehensive analyses qualitatively and quantitatively demonstrate the superiority of RCMT over state-of-the-art code-mixed and robust translation methods. Kartik, Sanjana Soni, Anoop Kunchukuttan, Tanmoy Chakraborty 0002, Md. Shad Akhtar |
LREC/COLING | 3 |
| 2023 | IndicSUPERB: A Speech Processing Universal Performance Benchmark for Indian LanguagesabstractA cornerstone in AI research has been the creation and adoption of standardized training and test datasets to earmark the progress of state-of-the-art models. A particularly successful example is the GLUE dataset for training and evaluating Natural Language Understanding (NLU) models for English. The large body of research around self-supervised BERT-based language models revolved around performance improvements on NLU tasks in GLUE. To evaluate language models in other languages, several language-specific GLUE datasets were created. The area of speech language understanding (SLU) has followed a similar trajectory. The success of large self-supervised models such as wav2vec2 enable creation of speech models with relatively easy to access unlabelled data. These models can then be evaluated on SLU tasks, such as the SUPERB benchmark. In this work, we extend this to Indic languages by releasing the IndicSUPERB benchmark. Specifically, we make the following three contributions. (i) We collect Kathbath containing 1,684 hours of labelled speech data across 12 Indian languages from 1,218 contributors located in 203 districts in India. (ii) Using Kathbath, we create benchmarks across 6 speech tasks: Automatic Speech Recognition, Speaker Verification, Speaker Identification (mono/multi), Language Identification, Query By Example, and Keyword Spotting for 12 languages. (iii) On the released benchmarks, we train and evaluate different self-supervised models alongside the a commonly used baseline FBANK. We show that language-specific fine-tuned models are more accurate than baseline on most of the tasks, including a large gap of 76% for Language Identification task. However, for speaker identification, self-supervised models trained on large datasets demonstrate an advantage. We hope IndicSUPERB contributes to the progress of developing speech language understanding models for Indian languages. Tahir Javed, Kaushal Santosh Bhogale, Abhigyan Raman, Anoop Kunchukuttan, Mitesh M. Khapra |
AAAI | 5 |
| 2023 | IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian LanguagesabstractAnanya Sai B, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan, Pratyush Kumar, Mitesh M. Khapra, Raj Dabre. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ananya Sai, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan, Mitesh M. Khapra, Raj Dabre |
ACL (1) | 4 |
| 2023 | Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic LanguagesabstractSumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M. Khapra, Anoop Kunchukuttan, Pratyush Kumar. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Sumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M. Khapra, Anoop Kunchukuttan |
ACL (1) | 6 |
| 2023 | Naamapadam: A Large-Scale Named Entity Annotated Data for Indic LanguagesabstractArnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M. Khapra, Pratyush Kumar, Rudra Murthy, Anoop Kunchukuttan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Arnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M. Khapra, V. Rudra Murthy, Anoop Kunchukuttan |
ACL (1) | 7 |
| 2023 | DecoMT: Decomposed Prompting for Machine Translation Between Related Languages using Large Language ModelsabstractThis study investigates machine translation between related languages i.e., languages within the same family that share linguistic characteristics such as word order and lexical similarity.Machine translation through few-shot prompting leverages a small set of translation pair examples to generate translations for test sentences.This procedure requires the model to learn how to generate translations while simultaneously ensuring that token ordering is maintained to produce a fluent and accurate translation.We propose that for related languages, the task of machine translation can be simplified by leveraging the monotonic alignment characteristic of such languages.We introduce DecoMT, a novel approach of fewshot prompting that decomposes the translation process into a sequence of word chunk translations.Through automatic and human evaluation conducted on multiple related language pairs across various language families, we demonstrate that our proposed approach of decomposed prompting surpasses multiple established few-shot baseline approaches.For example, DecoMT outperforms the strong fewshot prompting BLOOM model with an average improvement of 8 chrF++ scores across the examined languages. Ratish Puduppully, Anoop Kunchukuttan, Raj Dabre, AiTi Aw, Nancy Chen |
EMNLP | 2 |
| 2023 | Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource LanguagesabstractCollecting labelled datasets for speech recognition systems for low-resource languages on a diverse set of domains and speakers is expensive. In this work, we demonstrate an inexpensive and effective alternative by "mining" text and audio pairs for Indian languages from public sources, specifically from the public archives of All India Radio. As a key component, we adapt the Needleman-Wunsch algorithm to align sentences with corresponding audio segments given a long audio and a PDF of its transcript, while being robust to large errors due to OCR, extraneous text, and non-transcribed speech. We thus create Shrutilipi, a dataset which contains over 6,400 hours of labelled audio across 12 Indian languages totalling to 3.3M sentences. We establish the quality of Shrutilipi with 21 human evaluators across the 12 languages. We also establish the diversity of Shrutilipi in terms of represented regions, speakers, and mentioned named entities. Significantly, we show that adding Shrutilipi to the training dataset of ASR systems improves accuracy for both Wav2Vec and Conformer model architectures for 7 languages across benchmarks. Kaushal Santosh Bhogale, Abhigyan Raman, Tahir Javed, Sumanth Doddapaneni, Anoop Kunchukuttan, Mitesh M. Khapra |
ICASSP | 5 |
| 2022 | Towards Building ASR Systems for the Next Billion UsersabstractRecent methods in speech and language technology pretrain very large models which are fine-tuned for specific tasks. However, the benefits of such large models are often limited to a few resource rich languages of the world. In this work, we make multiple contributions towards building ASR systems for low resource languages from the Indian subcontinent. First, we curate 17,000 hours of raw speech data for 40 Indian languages from a wide variety of domains including education, news, technology, and finance. Second, using this raw speech data we pretrain several variants of wav2vec style models for 40 Indian languages. Third, we analyze the pretrained models to find key features: codebook vectors of similar sounding phonemes are shared across languages, representations across layers are discriminative of the language family, and attention heads often pay attention within small local windows. Fourth, we fine-tune this model for downstream ASR for 9 languages and obtain state-of-the-art results on 3 public datasets, including on very low-resource languages such as Sinhala and Nepali. Our work establishes that multilingual pretraining is an effective strategy for building ASR systems for the linguistically diverse speakers of the Indian subcontinent. Tahir Javed, Sumanth Doddapaneni, Abhigyan Raman, Kaushal Santosh Bhogale, Gowtham Ramesh, Anoop Kunchukuttan, Mitesh M. Khapra |
AAAI | 6 |
| 2022 | IndicXNLI: Evaluating Multilingual Inference for Indian LanguagesabstractWhile Indic NLP has made rapid advances recently in terms of the availability of corpora and pre-trained models, benchmark datasets on standard NLU tasks are limited.To this end, we introduce INDICXNLI, an NLI dataset for 11 Indic languages.It has been created by high-quality machine translation of the original English XNLI dataset and our analysis attests to the quality of INDICXNLI.By finetuning different pre-trained LMs on this IN-DICXNLI, we analyze various cross-lingual transfer techniques with respect to the impact of the choice of language models, languages, multi-linguality, mix-language input, etc.These experiments provide us with useful insights into the behaviour of pre-trained models for a diverse set of languages. Divyanshu Aggarwal, Vivek Gupta 0001, Anoop Kunchukuttan |
EMNLP | 3 |
| 2022 | IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic LanguagesabstractAman Kumar, Himani Shrotriya, Prachi Sahu, Amogh Mishra, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, Mitesh M. Khapra, Pratyush Kumar. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Himani Shrotriya, Prachi Sahu, Amogh Mishra, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, Mitesh M. Khapra |
EMNLP | 7 |
| 2022 | Bilingual Tabular Inference: A Case Study on Indic LanguagesabstractChaitanya Agarwal, Vivek Gupta, Anoop Kunchukuttan, Manish Shrivastava. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Chaitanya Agarwal, Vivek Gupta 0001, Anoop Kunchukuttan, Manish Shrivastava 0001 |
NAACL-HLT | 3 |
| 2022 | Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic LanguagesabstractAbstract We present Samanantar, the largest publicly available parallel corpora collection for Indic languages. The collection contains a total of 49.7 million sentence pairs between English and 11 Indic languages (from two language families). Specifically, we compile 12.4 million sentence pairs from existing, publicly available parallel corpora, and additionally mine 37.4 million sentence pairs from the Web, resulting in a 4× increase. We mine the parallel sentences from the Web by combining many corpora, tools, and methods: (a) Web-crawled monolingual corpora, (b) document OCR for extracting sentences from scanned documents, (c) multilingual representation models for aligning sentences, and (d) approximate nearest neighbor search for searching in a large collection of sentences. Human evaluation of samples from the newly mined corpora validate the high quality of the parallel sentences across 11 languages. Further, we extract 83.4 million sentence pairs between all 55 Indic language pairs from the English-centric parallel corpus using English as the pivot language. We trained multilingual NMT models spanning all these languages on Samanantar which outperform existing models and baselines on publicly available benchmarks, such as FLORES, establishing the utility of Samanantar. Our data and models are available publicly at Samanantar and we hope they will help advance research in NMT and multilingual NLP for Indic languages. Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalakshmi J, Divyanshu Kakwani, Aswin Pradeep, Srihari Nagaraj, Vivek Raghavan, Anoop Kunchukuttan, Mitesh Shantadevi Khapra |
Trans. Assoc. Comput. Linguistics | 16 |
| 2021 | A Large-scale Evaluation of Neural Machine Transliteration for Indic LanguagesabstractWe take up the task of largescale evaluation of neural machine transliteration between English and Indian languages, with a focus on multilin gual transliteration to utilize orthographic sim ilarity between Indian languages.We create a corpus of 600K word pairs mined from parallel translation corpora and monolingual corpora, which is the largest transliteration corpora for Indian languages mined from public sources.We perform a detailed analysis of multilingual transliteration and propose an improved mul tilingual training pipeline for Indic languages.We analyse various factors affecting transliter ation quality like language family, translitera tion direction and word origin. Anoop Kunchukuttan, Rahul Kejriwal |
EACL | 1 |
| 2019 | Learning Multilingual Word Embeddings in Latent Metric Space: A Geometric ApproachabstractAbstract We propose a novel geometric approach for learning bilingual mappings given monolingual embeddings and a bilingual dictionary. Our approach decouples the source-to-target language transformation into (a) language-specific rotations on the original embeddings to align them in a common, latent space, and (b) a language-independent similarity metric in this common space to better model the similarity between the embeddings. Overall, we pose the bilingual mapping problem as a classification problem on smooth Riemannian manifolds. Empirically, our approach outperforms previous approaches on the bilingual lexicon induction and cross-lingual word similarity tasks. We next generalize our framework to represent multiple languages in a common latent space. Language-specific rotations for all the languages and a common similarity metric in the latent space are learned jointly from bilingual dictionaries for multiple language pairs. We illustrate the effectiveness of joint learning for multiple languages in an indirect word translation setting. Pratik Jawanpuria, Arjun Balgovind, Anoop Kunchukuttan, Bamdev Mishra |
Trans. Assoc. Comput. Linguistics | 3 |
| 2018 | The IIT Bombay English-Hindi Parallel Corpus
Anoop Kunchukuttan, Pratik Mehta, Pushpak Bhattacharyya |
LREC | 1 |
| 2018 | Leveraging Orthographic Similarity for Multilingual Neural TransliterationabstractWe address the task of joint training of transliteration models for multiple language pairs ( multilingual transliteration). This is an instance of multitask learning, where individual tasks (language pairs) benefit from sharing knowledge with related tasks. We focus on transliteration involving related tasks i.e., languages sharing writing systems and phonetic properties ( orthographically similar languages). We propose a modified neural encoder-decoder model that maximizes parameter sharing across language pairs in order to effectively leverage orthographic similarity. We show that multilingual transliteration significantly outperforms bilingual transliteration in different scenarios (average increase of 58% across a variety of languages we experimented with). We also show that multilingual transliteration models can generalize well to languages/language pairs not encountered during training and hence perform well on the zeroshot transliteration task. We show that further improvements can be achieved by using phonetic feature input. Anoop Kunchukuttan, Mitesh M. Khapra, Gurneet Singh, Pushpak Bhattacharyya |
Trans. Assoc. Comput. Linguistics | 1 |
| 2016 | Substring-based unsupervised transliteration with phonetic and contextual knowledge
Anoop Kunchukuttan, Pushpak Bhattacharyya, Mitesh M. Khapra |
CoNLL | 1 |
| 2016 | Orthographic Syllable as basic unit for SMT between Related LanguagesabstractWe explore the use of the orthographic syllable, a variable-length consonant-vowel sequence, as a basic unit of translation between related languages which use abugida or alphabetic scripts.We show that orthographic syllable level translation significantly outperforms models trained over other basic units (word, morpheme and character) when training over small parallel corpora. Anoop Kunchukuttan, Pushpak Bhattacharyya |
EMNLP | 1 |
| 2015 | Brahmi-Net: A transliteration and script conversion system for languages of the Indian subcontinentabstractWe present Brahmi-Net- an online system for transliteration and script conversion for all ma-jor Indian language pairs (306 pairs). The sys-tem covers 13 Indo-Aryan languages, 4 Dra-vidian languages and English. For training the transliteration systems, we mined paral-lel transliteration corpora from parallel trans-lation corpora using an unsupervised method and trained statistical transliteration systems using the mined corpora. Languages which do not have parallel corpora are supported by transliteration through a bridge language. Our script conversion system supports con-version between all Brahmi-derived scripts as well as ITRANS romanization scheme. For this, we leverage co-ordinated Unicode ranges between Indic scripts and use an extended ITRANS encoding for transliterating between English and Indic scripts. The system also pro-vides top-k transliterations and simultaneous transliteration into multiple output languages. We provide a Python as well as REST API to access these services. The API and the mined transliteration corpus are made available for research use under an open source license. 1 Anoop Kunchukuttan, Ratish Puduppully, Pushpak Bhattacharyya |
HLT-NAACL | 1 |
| 2014 | When Transliteration Met Crowdsourcing : An Empirical Study of Transliteration via Crowdsourcing using Efficient, Non-redundant and Fair Quality Control
Mitesh M. Khapra, Ananthakrishnan Ramanathan, Anoop Kunchukuttan, Karthik Visweswariah, Pushpak Bhattacharyya |
LREC | 3 |
| 2014 | Shata-Anuvadak: Tackling Multiway Translation of Indian Languages
Anoop Kunchukuttan, Abhijit Mishra, Rajen Chatterjee, Ritesh M. Shah, Pushpak Bhattacharyya |
LREC | 1 |
| 2012 | Experiences in Resource Generation for Machine Translation through Crowdsourcing
Anoop Kunchukuttan, Shourya Roy, Pratik Patel, Kushal Ladha, Somya Gupta, Mitesh M. Khapra, Pushpak Bhattacharyya |
LREC | 1 |