EDBT 2026 Demo / reviewers in the wild / expert
Mitesh M. Khapra
dblp:90/7967
· DBLP profile ↗
85ranked-venue papers
9as first author
35since 2021 · last 2025
0009-0008-3687-9922ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 77 · 9 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 1 first-author · 18 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cross-Lingual Auto Evaluation for Assessing Multilingual LLMsabstractSumanth Doddapaneni, Mohammed Safi Ur Rahman Khan, Dilip Venkatesh, Raj Dabre, Anoop Kunchukuttan, Mitesh M Khapra. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Sumanth Doddapaneni, Mohammed Safi Ur Rahman Khan, Dilip Venkatesh, Raj Dabre, Anoop Kunchukuttan, Mitesh M. Khapra |
ACL (1) | 6 |
| 2025 | Can Vision-Language Models Evaluate Handwritten Math?abstractRecent advancements in Vision-Language Models (VLMs) have opened new possibilities in automatic grading of handwritten student responses, particularly in mathematics. However, a comprehensive study to test the ability of VLMs to evaluate and reason over handwritten content remains absent. To address this gap, we introduce FERMAT, a benchmark designed to assess VLMs’ ability to detect, localize and correct errors in handwritten mathematical content. FERMAT spans four key error dimensions - computational, conceptual, notational, and presentation - and comprises over 2,200 handwritten math solutions derived from 609 manually curated problems from grades 7-12 with intentionally introduced perturbations. Using FERMAT we benchmark nine VLMs across three tasks: error detection, localization, and correction. Our results reveal significant shortcomings in current VLMs in reasoning over handwritten text, with Gemini-1.5-Pro achieving the highest error correction rate (77%). We also observed that some models struggle with processing handwritten content, as their accuracy improves when handwritten inputs are replaced with printed text or images. These findings highlight the limitations of current VLMs and reveal new avenues for improvement. We will release FERMAT and all the associated resources in the open-source to drive further research. Oikantik Nath, Hanani Bathina, Mohammed Safi Ur Rahman Khan, Mitesh M. Khapra |
ACL (1) | 4 |
| 2025 | FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and StereotypesabstractJanki Atul Nawale, Mohammed Safi Ur Rahman Khan, Janani D, Mansi Gupta, Danish Pruthi, Mitesh M Khapra. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Janki Nawale, Mohammed Safi Ur Rahman Khan, Janani D, Danish Pruthi, Mitesh M. Khapra |
ACL (1) | 6 |
| 2025 | Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian LanguagesabstractAshwin Sankar, Sparsh Jain, Nikhil Narasimhan, Devilal Choudhary, Dhairya Suman, Mohammed Safi Ur Rahman Khan, Anoop Kunchukuttan, Mitesh M Khapra, Raj Dabre. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ashwin Sankar, Sparsh Jain, Nikhil Narasimhan, Devilal Choudhary, Dhairya Suman, Mohammed Safi Ur Rahman Khan, Anoop Kunchukuttan, Mitesh M. Khapra, Raj Dabre |
ACL (1) | 8 |
| 2025 | Towards Bringing Parity in Pretraining Datasets for Low-resource Indian LanguagesabstractLack of large-scale pretraining data for low resource languages from the Indian sub-continent, leads to their underrepresentation in existing massively multilingual models. In this work, we address this gap by proposing a framework to create large raw audio datasets for such under-represented languages by collating publicly accessible audio content. Leveraging this framework, we present MahaDhwani, a corpus comprising 279K hours of raw audio across 22 Indian languages. To test the utility of MahaDhwani, we pretrain a conformer style model, and then further finetune it to build a multilingual ASR model supporting the 22 languages. Using a hybrid multi-softmax decoder, we balance the benefit of shared parameters which enable crosslingual transfer, and the benefit of dedicated capacity for each language. Our evaluations on the IndicVoices benchmark show the benefits of pre-training, particularly in low-resource settings. We will open-source our framework, code and scripts to reproduce the dataset. Kaushal Santosh Bhogale, Deovrat Mehendale, Tahir Javed, Devbrat Anuragi, Sakshi Joshi, Sai Sundaresan, Aparna Ananthanarayanan, Sharmistha Dey, Sathish Kumar Reddy G, Anusha Srinivasan, Abhigyan Raman, Mitesh M. Khapra |
ICASSP | 13 |
| 2025 | IndicDLP: A Foundational Dataset for Multi-lingual and Multi-domain Document Layout Parsing
Oikantik Nath, Sahithi Kukkala, Mitesh M. Khapra, Ravi Kiran Sarvadevabhatla |
ICDAR (1) | 3 |
| 2025 | NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data
Tahir Javed, Kaushal Santosh Bhogale, Mitesh M. Khapra |
INTERSPEECH | 3 |
| 2025 | Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
Sakshi Joshi, Eldho Ittan George, Tahir Javed, Kaushal Santosh Bhogale, Nikhil Narasimhan, Mitesh M. Khapra |
INTERSPEECH | 6 |
| 2025 | Rasmalai : Resources for Adaptive Speech Modeling in IndiAn Languages with Accents and Intonations
Ashwin Sankar, Yoach Lacombe, Sherry Thomas, Praveen Srinivasa Varadhan, Sanchit Gandhi, Mitesh M. Khapra |
INTERSPEECH | 6 |
| 2025 | The State Of TTS: A Case Study with Human Fooling Rates
Praveen Srinivasa Varadhan, Sherry Thomas, Sai Teja M. S., Suvrat Bhooshan, Mitesh M. Khapra |
INTERSPEECH | 5 |
| 2025 | Quality Estimation and Post-Editing Using LLMs For Indic Languages: How Good Is It?abstractRecently, there have been increasing efforts on Quality Estimation (QE) and Post-Editing (PE) using Large Language Models (LLMs) for Machine Translation (MT). However, the focus has mainly been on high resource languages and the approaches either rely on prompting or combining existing QE models with LLMs, instead of single end-to-end systems. In this paper, we investigate the efficacy of end-to-end QE and PE systems for low-resource languages taking 5 Indian languages as a use-case. We augment existing QE data containing multidimentional quality metric (MQM) error annotations with explanations of errors and PEs with the help of proprietary LLMs (GPT-4), following which we fine-tune Gemma-2-9B, an open-source multilingual LLM to perform QE and PE jointly. While our models attain QE capabilities competitive with or surpassing existing models in both referenceful and referenceless settings, we observe that they still struggle with PE. Further investigation reveals that this occurs because our models lack the ability to accurately identify fine-grained errors in the translation, despite being excellent indicators of overall quality. This opens up opportunities for research in end-to-end QE and PE for low-resource languages. Anushka Singh, Aarya Pakhale, Mitesh M. Khapra, Raj Dabre |
MTSummit (1) | 3 |
| 2024 | IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian LanguagesabstractMohammed Safi Ur Rahman Khan, Priyam Mehta, Ananth Sankar, Umashankar Kumaravelan, Sumanth Doddapaneni, Suriyaprasaad B, Varun G, Sparsh Jain, Anoop Kunchukuttan, Pratyush Kumar, Raj Dabre, Mitesh M. Khapra. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Mohammed Safi Ur Rahman Khan, Priyam Mehta, Ananth Sankar, Umashankar Kumaravelan, Sumanth Doddapaneni, Suriyaprasaad B, Varun Balan G, Sparsh Jain, Anoop Kunchukuttan, Raj Dabre, Mitesh M. Khapra |
ACL (1) | 12 |
| 2024 | Finding Blind Spots in Evaluator LLMs with Interpretable ChecklistsabstractLarge Language Models (LLMs) are increasingly relied upon to evaluate text outputs of other LLMs, thereby influencing leaderboards and development decisions.However, concerns persist over the accuracy of these assessments and the potential for misleading conclusions.In this work, we investigate the effectiveness of LLMs as evaluators for text generation tasks.We propose FBI, a novel framework designed to examine the proficiency of Evaluator LLMs in assessing four critical abilities in other LLMs: factual accuracy, instruction following, coherence in long-form writing, and reasoning proficiency.By introducing targeted perturbations in answers generated by LLMs, that clearly impact one of these key capabilities, we test whether an Evaluator LLM can detect these quality drops.By creating a total of 2400 perturbed answers covering 22 perturbation categories, we conduct a comprehensive study using different evaluation strategies on five prominent LLMs commonly used as evaluators in the literature.Our findings reveal significant shortcomings in current Evaluator LLMs, which failed to identify quality drops in over 50% of cases on average.Singleanswer and pairwise evaluations demonstrated notable limitations, whereas reference-based evaluations showed comparatively better performance.These results underscore the unreliable nature of current Evaluator LLMs and advocate for cautious implementation in practical applications.Code and data are available at https://github.com/AI4Bharat/FBI. Sumanth Doddapaneni, Mohammed Safi Ur Rahman Khan, Sshubam Verma, Mitesh M. Khapra |
EMNLP | 4 |
| 2024 | Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
Srija Anand, Praveen Srinivasa Varadhan, Ashwin Sankar, Giri Raju, Mitesh M. Khapra |
INTERSPEECH | 5 |
| 2024 | Empowering Low-Resource Language ASR via Large-Scale Pseudo Labeling
Kaushal Santosh Bhogale, Deovrat Mehendale, Niharika Parasa, Sathish Kumar Reddy G, Tahir Javed, Mitesh M. Khapra |
INTERSPEECH | 7 |
| 2024 | LAHAJA: A Robust Multi-accent Benchmark for Evaluating Hindi ASR Systems
Tahir Javed, Janki Nawale, Sakshi Joshi, Eldho Ittan George, Kaushal Santosh Bhogale, Deovrat Mehendale, Mitesh M. Khapra |
INTERSPEECH | 7 |
| 2024 | Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
Praveen Srinivasa Varadhan, Ashwin Sankar, Giri Raju, Mitesh M. Khapra |
INTERSPEECH | 4 |
| 2024 | IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTSabstractRecent advancements in text-to-speech (TTS) synthesis show that large-scale models trained with extensive web data produce highly natural-sounding output. However, such data is scarce for Indian languages due to the lack of high-quality, manually subtitled data on platforms like LibriVox or YouTube. To address this gap, we enhance existing large-scale ASR datasets containing natural conversations collected in low-quality environments to generate high-quality TTS training data. Our pipeline leverages the cross-lingual generalization of denoising and speech enhancement models trained on English and applied to Indian languages. This results in IndicVoices-R (IV-R), the largest multilingual Indian TTS dataset derived from an ASR dataset, with 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. IV-R matches the quality of gold-standard TTS datasets like LJSpeech, LibriTTS, and IndicTTS. We also introduce the IV-R Benchmark, the first to assess zero-shot, few-shot, and many-shot speaker generalization capabilities of TTS models on Indian voices, ensuring diversity in age, gender, and style. We demonstrate that fine-tuning an English pre-trained model on a combined dataset of high-quality IndicTTS and our IV-R dataset results in better zero-shot speaker generalization compared to fine-tuning on the IndicTTS dataset alone. Further, our evaluation reveals limited zero-shot generalization for Indian voices in TTS models trained on prior datasets, which we improve by fine-tuning the model on our data containing diverse set of speakers across language families. We open-source code and data for all 22 official Indian languages. Ashwin Sankar, Srija Anand, Praveen Srinivasa Varadhan, Sherry Thomas, Mehak Singal, Shridhar Kumar, Deovrat Mehendale, Aditi Krishana, Giri Raju, Mitesh M. Khapra |
NeurIPS | 10 |
| 2023 | IndicSUPERB: A Speech Processing Universal Performance Benchmark for Indian LanguagesabstractA cornerstone in AI research has been the creation and adoption of standardized training and test datasets to earmark the progress of state-of-the-art models. A particularly successful example is the GLUE dataset for training and evaluating Natural Language Understanding (NLU) models for English. The large body of research around self-supervised BERT-based language models revolved around performance improvements on NLU tasks in GLUE. To evaluate language models in other languages, several language-specific GLUE datasets were created. The area of speech language understanding (SLU) has followed a similar trajectory. The success of large self-supervised models such as wav2vec2 enable creation of speech models with relatively easy to access unlabelled data. These models can then be evaluated on SLU tasks, such as the SUPERB benchmark. In this work, we extend this to Indic languages by releasing the IndicSUPERB benchmark. Specifically, we make the following three contributions. (i) We collect Kathbath containing 1,684 hours of labelled speech data across 12 Indian languages from 1,218 contributors located in 203 districts in India. (ii) Using Kathbath, we create benchmarks across 6 speech tasks: Automatic Speech Recognition, Speaker Verification, Speaker Identification (mono/multi), Language Identification, Query By Example, and Keyword Spotting for 12 languages. (iii) On the released benchmarks, we train and evaluate different self-supervised models alongside the a commonly used baseline FBANK. We show that language-specific fine-tuned models are more accurate than baseline on most of the tasks, including a large gap of 76% for Language Identification task. However, for speaker identification, self-supervised models trained on large datasets demonstrate an advantage. We hope IndicSUPERB contributes to the progress of developing speech language understanding models for Indian languages. Tahir Javed, Kaushal Santosh Bhogale, Abhigyan Raman, Anoop Kunchukuttan, Mitesh M. Khapra |
AAAI | 6 |
| 2023 | IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian LanguagesabstractAnanya Sai B, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan, Pratyush Kumar, Mitesh M. Khapra, Raj Dabre. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ananya Sai, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan, Mitesh M. Khapra, Raj Dabre |
ACL (1) | 6 |
| 2023 | Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic LanguagesabstractSumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M. Khapra, Anoop Kunchukuttan, Pratyush Kumar. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Sumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M. Khapra, Anoop Kunchukuttan |
ACL (1) | 5 |
| 2023 | Naamapadam: A Large-Scale Named Entity Annotated Data for Indic LanguagesabstractArnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M. Khapra, Pratyush Kumar, Rudra Murthy, Anoop Kunchukuttan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Arnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M. Khapra, V. Rudra Murthy, Anoop Kunchukuttan |
ACL (1) | 4 |
| 2023 | Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource LanguagesabstractCollecting labelled datasets for speech recognition systems for low-resource languages on a diverse set of domains and speakers is expensive. In this work, we demonstrate an inexpensive and effective alternative by "mining" text and audio pairs for Indian languages from public sources, specifically from the public archives of All India Radio. As a key component, we adapt the Needleman-Wunsch algorithm to align sentences with corresponding audio segments given a long audio and a PDF of its transcript, while being robust to large errors due to OCR, extraneous text, and non-transcribed speech. We thus create Shrutilipi, a dataset which contains over 6,400 hours of labelled audio across 12 Indian languages totalling to 3.3M sentences. We establish the quality of Shrutilipi with 21 human evaluators across the 12 languages. We also establish the diversity of Shrutilipi in terms of represented regions, speakers, and mentioned named entities. Significantly, we show that adding Shrutilipi to the training dataset of ASR systems improves accuracy for both Wav2Vec and Conformer model architectures for 7 languages across benchmarks. Kaushal Santosh Bhogale, Abhigyan Raman, Tahir Javed, Sumanth Doddapaneni, Anoop Kunchukuttan, Mitesh M. Khapra |
ICASSP | 7 |
| 2023 | Towards Building Text-to-Speech Systems for the Next Billion UsersabstractDeep learning based text-to-speech (TTS) systems have been evolving rapidly with advances in model architectures, training methodologies, and generalization across speakers and languages. However, these advances have not been thoroughly investigated for Indian language speech synthesis. Such investigation is computationally expensive given the number and diversity of Indian languages, relatively lower resource availability, and the diverse set of advances in neural TTS that remain untested. In this paper, we evaluate the choice of acoustic models, vocoders, supplementary loss functions, training schedules, and speaker and language diversity for Dravidian and Indo-Aryan languages. Based on this, we identify monolingual models with FastPitch and HiFi-GAN V1, trained jointly on male and female speakers to perform the best. With this setup, we train and evaluate TTS models for 13 languages and find our models to significantly improve upon existing models in all languages as measured by mean opinion scores. We open-source all models on the Bhashini platform. Gokul Karthik Kumar, Praveen S. V, Mitesh M. Khapra, Karthik Nandakumar |
ICASSP | 4 |
| 2023 | Vistaar: Diverse Benchmarks and Training Sets for Indian Language ASR
Kaushal Santosh Bhogale, Sai Sundaresan, Abhigyan Raman, Tahir Javed, Mitesh M. Khapra |
INTERSPEECH | 5 |
| 2023 | Svarah: Evaluating English ASR Systems on Indian Accents
Tahir Javed, Sakshi Joshi, Vignesh Nagarajan, Sai Sundaresan, Janki Nawale, Abhigyan Raman, Kaushal Santosh Bhogale, Mitesh M. Khapra |
INTERSPEECH | 9 |
| 2022 | Towards Building ASR Systems for the Next Billion UsersabstractRecent methods in speech and language technology pretrain very large models which are fine-tuned for specific tasks. However, the benefits of such large models are often limited to a few resource rich languages of the world. In this work, we make multiple contributions towards building ASR systems for low resource languages from the Indian subcontinent. First, we curate 17,000 hours of raw speech data for 40 Indian languages from a wide variety of domains including education, news, technology, and finance. Second, using this raw speech data we pretrain several variants of wav2vec style models for 40 Indian languages. Third, we analyze the pretrained models to find key features: codebook vectors of similar sounding phonemes are shared across languages, representations across layers are discriminative of the language family, and attention heads often pay attention within small local windows. Fourth, we fine-tune this model for downstream ASR for 9 languages and obtain state-of-the-art results on 3 public datasets, including on very low-resource languages such as Sinhala and Nepali. Our work establishes that multilingual pretraining is an effective strategy for building ASR systems for the linguistically diverse speakers of the Indian subcontinent. Tahir Javed, Sumanth Doddapaneni, Abhigyan Raman, Kaushal Santosh Bhogale, Gowtham Ramesh, Anoop Kunchukuttan, Mitesh M. Khapra |
AAAI | 8 |
| 2022 | Active Evaluation: Efficient NLG Evaluation with Few Pairwise ComparisonsabstractRecent studies have shown the advantages of evaluating NLG systems using pairwise comparisons as opposed to direct assessment.Given k systems, a naive approach for identifying the top-ranked system would be to uniformly obtain pairwise comparisons from all k 2 pairs of systems.However, this can be very expensive as the number of human annotations required would grow quadratically with k.In this work, we introduce Active Evaluation, a framework to efficiently identify the top-ranked system by actively choosing system pairs for comparison using dueling bandit algorithms.We perform extensive experiments with 13 dueling bandits algorithms on 13 NLG evaluation datasets spanning 5 tasks and show that the number of human annotations can be reduced by 80%.To further reduce the number of human annotations, we propose model-based dueling bandit algorithms which combine automatic evaluation metrics with human evaluations.Specifically, we eliminate sub-optimal systems even before the human annotation process and perform human evaluations only on test examples where the automatic metric is highly uncertain.This reduces the number of human annotations required further by 89%.In effect, we show that identifying the top-ranked system requires only a few hundred human annotations, which grow linearly with k.Lastly, we provide practical recommendations and best practices to identify the top-ranked system efficiently. Akash Kumar Mohankumar, Mitesh M. Khapra |
ACL (1) | 2 |
| 2022 | OpenHands: Making Sign Language Recognition Accessible with Pose-based Pretrained Models across LanguagesabstractAI technologies for Natural Languages have made tremendous progress recently.However, commensurate progress has not been made on Sign Languages, in particular, in recognizing signs as individual words or as complete sentences.We introduce OpenHands 1 , a library where we take four key ideas from the NLP community for low-resource languages and apply them to sign languages for wordlevel recognition.First, we propose using pose extracted through pretrained models as the standard modality of data in this work to reduce training time and enable efficient inference, and we release standardized pose datasets for different existing sign language datasets.Second, we train and release checkpoints of 4 posebased isolated sign language recognition models across 6 languages (American, Argentinian, Chinese, Greek, Indian, and Turkish), providing baselines and ready checkpoints for deployment.Third, to address the lack of labelled data, we propose self-supervised pretraining on unlabelled data.We curate and release the largest pose-based pretraining dataset on Indian Sign Language (Indian-SL).Fourth, we compare different pretraining strategies and for the first time establish that pretraining is effective for sign language recognition by demonstrating (a) improved fine-tuning performance especially in low-resource settings, and (b) high crosslingual transfer from Indian-SL to few other sign languages.We open-source all models and datasets in OpenHands with a hope that it makes research in sign languages reproducible and more accessible. Prem Selvaraj, Gokul N. C., Mitesh M. Khapra |
ACL (1) | 4 |
| 2022 | IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic LanguagesabstractAman Kumar, Himani Shrotriya, Prachi Sahu, Amogh Mishra, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, Mitesh M. Khapra, Pratyush Kumar. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Himani Shrotriya, Prachi Sahu, Amogh Mishra, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, Mitesh M. Khapra |
EMNLP | 8 |
| 2022 | Addressing Resource Scarcity across Sign Languages with Multilingual Pretraining and Unified-Vocabulary DatasetsabstractThere are over 300 sign languages in the world, many of which have very limited or no labelled sign-to-text datasets. To address low-resource data scenarios, self-supervised pretraining and multilingual finetuning have been shown to be effective in natural language and speech processing. In this work, we apply these ideas to sign language recognition.We make three contributions.- First, we release SignCorpus, a large pretraining dataset on sign languages comprising about 4.6K hours of signing data across 10 sign languages. SignCorpus is curated from sign language videos on the internet, filtered for data quality, and converted into sequences of pose keypoints thereby removing all personal identifiable information (PII).- Second, we release Sign2Vec, a graph-based model with 5.2M parameters that is pretrained on SignCorpus. We envisage Sign2Vec as a multilingual large-scale pretrained model which can be fine-tuned for various sign recognition tasks across languages.- Third, we create MultiSign-ISLR -- a multilingual and label-aligned dataset of sequences of pose keypoints from 11 labelled datasets across 7 sign languages, and MultiSign-FS -- a new finger-spelling training and test set across 7 languages. On these datasets, we fine-tune Sign2Vec to create multilingual isolated sign recognition models. With experiments on multiple benchmarks, we show that pretraining and multilingual transfer are effective giving significant gains over state-of-the-art results.All datasets, models, and code has been made open-source via the OpenHands toolkit. Gokul NC, Manideep Ladi, Sumit Negi, Prem Selvaraj, Mitesh M. Khapra |
NeurIPS | 6 |
| 2021 | A Systematic Evaluation of Object Detection Networks for Scientific PlotsabstractAre existing object detection methods adequate for detecting text and visual elements in scientific plots which are arguably different than the objects found in natural images? To answer this question, we train and compare the accuracy of Fast/Faster R-CNN, SSD, YOLO and RetinaNet on the PlotQA dataset with over 220,000 scientific plots. At the standard IOU setting of 0.5, most networks perform well with mAP scores greater than 80% in detecting the relatively simple objects in plots. However, the performance drops drastically when evaluated at a stricter IOU of 0.9 with the best model giving a mAP of 35.70%. Note that such a stricter evaluation is essential when dealing with scientific plots where even minor localisation errors can lead to large errors in downstream numerical inferences. Given this poor performance, we propose minor modifications to existing models by combining ideas from different object detection networks. While this significantly improves the performance, there are still two main issues: (i) performance on text objects which are essential for reasoning is very poor, and (ii) inference time is unacceptably large considering the simplicity of plots. To solve this open problem, we make a series of contributions: (a) an efficient region proposal method based on Laplacian edge detectors, (b) a feature representation of region proposals that includes neighbouring information, (c) a linking component to join multiple region proposals for detecting longer textual objects, and (d) a custom loss function that combines a smooth L1-loss with an IOU-based loss. Combining these ideas, our final model is very accurate at extreme IOU values achieving a mAP of 93.44%@0.9 IOU. Simultaneously, our model is very efficient with an inference time 16x lesser than the current models, including one-stage detectors. Our model also achieves a high accuracy on an extrinsic plot-to-table conversion task with an F1 score of 0.77. With these contributions, we make a definitive progress in object detection for plots and enable further exploration on automated reasoning of plots. Pritha Ganguly, Nitesh Methani, Mitesh M. Khapra |
AAAI | 3 |
| 2021 | The Heads Hypothesis: A Unifying Statistical Approach Towards Understanding Multi-Headed Attention in BERTabstractMulti-headed attention heads are a mainstay in transformer-based models. Different methods have been proposed to classify the role of each attention head based on the relations between tokens which have high pair-wise attention. These roles include syntactic (tokens with some syntactic relation), local (nearby tokens), block (tokens in the same sentence) and delimiter (the special [CLS], [SEP] tokens). There are two main challenges with existing methods for classification: (a) there are no standard scores across studies or across functional roles, and (b) these scores are often average quantities measured across sentences without capturing statistical significance. In this work, we formalize a simple yet effective score that generalizes to all the roles of attention heads and employs hypothesis testing on this score for robust inference. This provides us the right lens to systematically analyze attention heads and confidently comment on many commonly posed questions on analyzing the BERT model. In particular, we comment on the co-location of multiple functional roles in the same attention head, the distribution of attention heads across layers, and effect of fine-tuning for specific NLP tasks on these functional roles. The code is made publicly available at https://github.com/iitmnlp/heads-hypothesis Madhura Pande, Aakriti Budhraja, Preksha Nema, Mitesh M. Khapra |
AAAI | 5 |
| 2021 | Perturbation CheckLists for Evaluating NLG Evaluation MetricsabstractNatural Language Generation (NLG) evaluation is a multifaceted task requiring assessment of multiple desirable criteria, e.g., fluency, coherency, coverage, relevance, adequacy, overall quality, etc. Across existing datasets for 6 NLG tasks, we observe that the human evaluation scores on these multiple criteria are often not correlated.For example, there is a very low correlation between human scores on fluency and data coverage for the task of structured data to text generation.This suggests that the current recipe of proposing new automatic evaluation metrics for NLG by showing that they correlate well with scores assigned by humans for a single criteria (overall quality) alone is inadequate.Indeed, our extensive study involving 25 automatic evaluation metrics across 6 different tasks and 18 different evaluation criteria shows that there is no single metric which correlates well with human scores on all desirable criteria, for most NLG tasks.Given this situation, we propose CheckLists for better design and evaluation of automatic metrics.We design templates which target a specific criteria (e.g., coverage) and perturb the output such that the quality gets affected only along this specific criteria (e.g., the coverage drops).We show that existing evaluation metrics are not robust against even such simple perturbations and disagree with scores assigned by humans to the perturbed output.The proposed templates thus allow for a fine-grained assessment of automatic evaluation metrics exposing their limitations and will facilitate better design, analysis and evaluation of such metrics. 1 Ananya Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, Mitesh M. Khapra |
EMNLP (1) | 5 |
| 2021 | Unsupervised Deep Video DenoisingabstractDeep convolutional neural networks (CNNs) for video denoising are typically trained with supervision, assuming the availability of clean videos. However, in many applications, such as microscopy, noiseless videos are not available. To address this, we propose an Unsupervised Deep Video Denoiser (UDVD1), a CNN architecture designed to be trained exclusively with noisy data. The performance of UDVD is comparable to the supervised state-of-the-art, even when trained only on a single short noisy video. We demonstrate the promise of our approach in real-world imaging applications by denoising raw video, fluorescence-microscopy and electron-microscopy data. In contrast to many current approaches to video denoising, UDVD does not require explicit motion compensation. This is advantageous because motion compensation is computationally expensive, and can be unreliable when the input data are noisy. A gradient-based analysis reveals that UDVD automatically adapts to local motion in the input noisy videos. Thus, the network learns to perform implicit motion compensation, even though it is only trained for denoising. Dev Yashpal Sheth, Sreyas Mohan, Joshua L. Vincent, Ramon Manzorro, Peter A. Crozier, Mitesh M. Khapra, Eero P. Simoncelli, Carlos Fernandez-Granda |
ICCV | 6 |
| 2020 | Towards Transparent and Explainable Attention ModelsabstractAkash Kumar Mohankumar, Preksha Nema, Sharan Narasimhan, Mitesh M. Khapra, Balaji Vasan Srinivasan, Balaraman Ravindran. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Akash Kumar Mohankumar, Preksha Nema, Sharan Narasimhan, Mitesh M. Khapra, Balaji Vasan Srinivasan, Balaraman Ravindran |
ACL | 4 |
| 2020 | Joint Transformer/RNN Architecture for Gesture Typing in Indic LanguagesabstractGesture typing is a method of typing words on a touch-based keyboard by creating a continuous trace passing through the relevant keys.This work is aimed at developing a keyboard that supports gesture typing in Indic languages.We begin by noting that when dealing with Indic languages, one needs to cater to two different sets of users: (i) users who prefer to type in the native Indic script (Devanagari, Bengali, etc.) and (ii) users who prefer to type in the English script but want the transliterated output in the native script.In both cases, we need a model that takes a trace as input and maps it to the intended word.To enable the development of these models, we create and release two datasets.First, we create a dataset containing keyboard traces for 193,658 words from 7 Indic languages.Second, we curate 104,412 English-Indic transliteration pairs from Wikidata across these languages.Using these datasets we build a model that performs path decoding, transliteration and transliteration correction.Unlike prior approaches, our proposed model does not make co-character independence assumptions during decoding.The overall accuracy of our model across the 7 languages varies from 70-95%. Emil Biju, Anirudh Sriram, Mitesh M. Khapra |
COLING | 3 |
| 2020 | On the weak link between importance and prunability of attention headsabstractGiven the success of Transformer-based models, two directions of study have emerged: interpreting role of individual attention heads and down-sizing the models for efficiency.Our work straddles these two streams: We analyse the importance of basing pruning strategies on the interpreted role of the attention heads.We evaluate this on Transformer and BERT models on multiple NLP tasks.Firstly, we find that a large fraction of the attention heads can be randomly pruned with limited effect on accuracy.Secondly, for Transformers, we find no advantage in pruning attention heads identified to be important based on existing studies that relate importance to the location of a head.On the BERT model too we find no preference for top or bottom layers, though the latter are reported to have higher importance.However, strategies that avoid pruning middle layers and consecutive layers perform better.Finally, during fine-tuning the compensation for pruned attention heads is roughly equally distributed across the un-pruned heads.Our results thus suggest that interpretation of attention heads does not strongly inform pruning. Aakriti Budhraja, Madhura Pande, Preksha Nema, Mitesh M. Khapra |
EMNLP (1) | 5 |
| 2020 | Towards Interpreting BERT for Reading Comprehension Based QAabstractBERT and its variants have achieved stateof-the-art performance in various NLP tasks.Since then, various works have been proposed to analyze the linguistic information being captured in BERT.However, the current works do not provide an insight into how BERT is able to achieve near human-level performance on the task of Reading Comprehension based Question Answering.In this work, we attempt to interpret BERT for RCQA.Since BERT layers do not have predefined roles, we define a layer's role or functionality using Integrated Gradients.Based on the defined roles, we perform a preliminary analysis across all layers.We observed that the initial layers focus on query-passage interaction, whereas later layers focus more on contextual understanding and enhancing the answer prediction.Specifically for quantifier questions (how much/how many), we notice that BERT focuses on confusing words (i.e., on other numerical quantities in the passage) in the later layers, but still manages to predict the answer correctly.The fine-tuning and analysis scripts will be publicly available at https://github.com/ iitmnlp/BERT-Analysis-RCQA. Sahana Ramnath, Preksha Nema, Deep Sahni, Mitesh M. Khapra |
EMNLP (1) | 4 |
| 2020 | INCLUDE: A Large Scale Dataset for Indian Sign Language RecognitionabstractIndian Sign Language (ISL) is a complete language with its own grammar, syntax, vocabulary and several unique linguistic attributes. It is used by over 5 million deaf people in India. Currently, there is no publicly available dataset on ISL to evaluate Sign Language Recognition (SLR) approaches. In this work, we present the Indian Lexicon Sign Language Dataset - INCLUDE - an ISL dataset that contains 0.27 million frames across 4,287 videos over 263 word signs from 15 different word categories. INCLUDE is recorded with the help of experienced signers to provide close resemblance to natural conditions. A subset of 50 word signs is chosen across word categories to define INCLUDE-50 for rapid evaluation of SLR meth- ods with hyperparameter tuning. As the first large scale study of SLR on ISL, we evaluate several deep neural networks combining different methods for augmentation, feature extraction, encoding and decoding. The best performing model achieves an accuracy of 94.5% on the INCLUDE-50 dataset and 85.6% on the INCLUDE dataset. This model uses a pre-trained feature extractor and encoder and only trains a decoder. We further explore generalisation by fine-tuning the decoder for an American Sign Language dataset. On the ASLLVD with 48 classes, our model has an accuracy of 92.1%; improving on existing results and providing an efficient method to support SLR for multiple languages. Advaith Sridhar, Rohith Gandhi Ganesan, Mitesh M. Khapra |
ACM Multimedia | 4 |
| 2020 | PlotQA: Reasoning over Scientific PlotsabstractExisting synthetic datasets (Figure QA, DVQA) for reasoning over plots do not contain variability in data labels, real-valued data, or complex reasoning questions. Consequently, proposed models for these datasets do not fully address the challenge of reasoning over plots. In particular, they assume that the answer comes either from a small fixed size vocabulary or from a bounding box within the image. However, in practice, this is an unrealistic assumption because many questions require reasoning and thus have real-valued answers which appear neither in a small fixed size vocabulary nor in the image. In this work, we aim to bridge this gap between existing datasets and real-world plots. Specifically, we propose Plot QA with 28.9 million question-answer pairs over 224,377 plots on data from real-world sources and questions based on crowd-sourced question templates. Further, 80.76% of the out-of-vocabulary (OOV) questions in PlotQA have answers that are not in a fixed vocabulary. Analysis of existing models on Plot QA reveals that they cannot deal with OOV questions: their overall accuracy on our dataset is in single digits. This is not surprising given that these models were not designed for such questions. As a step towards a more holistic model which can address fixed vocabulary as well as OOV questions, we propose a hybrid approach: Specific questions are answered by choosing the answer from a fixed vocabulary or by extracting it from a predicted bounding box in the plot, while other questions are answered with a table question-answering engine which is fed with a structured table generated by detecting visual elements from the image. On the existing DVQA dataset, our model has an accuracy of 58%, significantly improving on the highest reported accuracy of 46%. On Plot QA, our model has an accuracy of 22.52%, which is significantly better than state of the art models. Nitesh Methani, Pritha Ganguly, Mitesh M. Khapra |
WACV | 3 |
| 2020 | Improving Dialog Evaluation with a Multi-reference Adversarial Dataset and Large Scale PretrainingabstractThere is an increasing focus on model-based dialog evaluation metrics such as ADEM, RUBER, and the more recent BERT-based metrics. These models aim to assign a high score to all relevant responses and a low score to all irrelevant responses. Ideally, such models should be trained using multiple relevant and irrelevant responses for any given context. However, no such data is publicly available, and hence existing models are usually trained using a single relevant response and multiple randomly selected responses from other contexts (random negatives). To allow for better training and robust evaluation of model-based metrics, we introduce the DailyDialog++ dataset, consisting of (i) five relevant responses for each context and (ii) five adversarially crafted irrelevant responses for each context. Using this dataset, we first show that even in the presence of multiple correct references, n-gram based metrics and embedding based metrics do not perform well at separating relevant responses from even random negatives. While model-based metrics perform better than n-gram and embedding based metrics on random negatives, their performance drops substantially when evaluated on adversarial examples. To check if large scale pretraining could help, we propose a new BERT-based evaluation metric called DEB, which is pretrained on 727M Reddit conversations and then finetuned on our dataset. DEB significantly outperforms existing models, showing better correlation with human judgments and better performance on random negatives (88.27% accuracy). However, its performance again drops substantially when evaluated on adversarial responses, thereby highlighting that even large-scale pretrained evaluation models are not robust to the adversarial examples in our dataset. The dataset 1 and code 2 are publicly available. Ananya Sai, Akash Kumar Mohankumar, Siddhartha Arora, Mitesh M. Khapra |
Trans. Assoc. Comput. Linguistics | 4 |
| 2019 | Re-Evaluating ADEM: A Deeper Look at Scoring Dialogue ResponsesabstractAutomatically evaluating the quality of dialogue responses for unstructured domains is a challenging problem. ADEM (Lowe et al. 2017) formulated the automatic evaluation of dialogue systems as a learning problem and showed that such a model was able to predict responses which correlate significantly with human judgements, both at utterance and system level. Their system was shown to have beaten word-overlap metrics such as BLEU with large margins. We start with the question of whether an adversary can game the ADEM model. We design a battery of targeted attacks at the neural network based ADEM evaluation system and show that automatic evaluation of dialogue systems still has a long way to go. ADEM can get confused with a variation as simple as reversing the word order in the text! We report experiments on several such adversarial scenarios that draw out counterintuitive scores on the dialogue responses. We take a systematic look at the scoring function proposed by ADEM and connect it to linear system theory to predict the shortcomings evident in the system. We also devise an attack that can fool such a system to rate a response generation system as favorable. Finally, we allude to future research directions of using the adversarial attacks to design a truly automated dialogue evaluation system. Ananya Sai, Mithun Das Gupta, Mitesh M. Khapra, Mukundhan Srinivasan |
AAAI | 3 |
| 2019 | Efficient Video Classification Using Fewer FramesabstractRecently, there has been a lot of interest in building compact models for video classification which have a small memory footprint (<1 GB). While these models are compact, they typically operate by repeated application of a small weight matrix to all the frames in a video. For example, recurrent neural network based methods compute a hidden state for every frame of the video using a recurrent weight matrix. Similarly, cluster-and-aggregate based methods such as NetVLAD have a learnable clustering matrix which is used to assign soft-clusters to every frame in the video. Since these models look at every frame in the video, the number of floating point operations (FLOPs) is still large even though the memory footprint is small. In this work, we focus on building compute-efficient video classification models which process fewer frames and hence have less number of FLOPs. Similar to memory efficient models, we use the idea of distillation albeit in a different setting. Specifically, in our case, a compute-heavy teacher which looks at all the frames in the video is used to train a compute-efficient student which looks at only a small fraction of frames in the video. This is in contrast to a typical memory efficient Teacher-Student setting, wherein both the teacher and the student look at all the frames in the video but the student has fewer parameters. Our work thus complements the research on memory efficient video classification. We do an extensive evaluation with three types of models for video classification, viz., (i) recurrent models (ii) cluster-and-aggregate models and (iii) memory-efficient cluster-and-aggregate models and show that in each of these cases, a see-it-all teacher can be used to train a compute efficient see-very-little student. Overall, we show that the proposed student network can reduce the inference time by 30% and the number of FLOPs by approximately 90% with a negligent drop in the performance. Shweta Bhardwaj, Mukundhan Srinivasan, Mitesh M. Khapra |
CVPR | 3 |
| 2019 | Let's Ask Again: Refine Network for Automatic Question GenerationabstractPreksha Nema, Akash Kumar Mohankumar, Mitesh M. Khapra, Balaji Vasan Srinivasan, Balaraman Ravindran. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Preksha Nema, Akash Kumar Mohankumar, Mitesh M. Khapra, Balaji Vasan Srinivasan, Balaraman Ravindran |
EMNLP/IJCNLP (1) | 3 |
| 2019 | FigureNet : A Deep Learning model for Question-Answering on Scientific PlotsabstractDeep Learning has managed to push boundaries in a wide variety of tasks. One area of interest is to tackle problems in reasoning and understanding, with an aim to emulate human intelligence. In this work, we describe a deep learning model that addresses the reasoning task of question-answering on categorical plots. We introduce a novel architecture FigureNet, that learns to identify various plot elements, quantify the represented values and determine a relative ordering of these statistical values. We test our model on the FigureQA dataset which provides images and accompanying questions for scientific plots like bar graphs and pie charts, augmented with rich annotations. Our approach outperforms the state-of-the-art Relation Networks baseline by approximately 7% on this dataset, with a training time that is over an order of magnitude lesser. Revanth Reddy, Rahul Ramesh, Ameet Deshpande, Mitesh M. Khapra |
IJCNN | 4 |
| 2019 | Studying the plasticity in deep convolutional neural networks using random pruning
Deepak Mittal, Shweta Bhardwaj, Mitesh M. Khapra, Balaraman Ravindran |
Mach. Vis. Appl. | 3 |
| 2019 | Graph Convolutional Network with Sequential Attention for Goal-Oriented Dialogue SystemsabstractDomain-specific goal-oriented dialogue systems typically require modeling three types of inputs, namely, (i) the knowledge-base associated with the domain, (ii) the history of the conversation, which is a sequence of utterances, and (iii) the current utterance for which the response needs to be generated. While modeling these inputs, current state-of-the-art models such as Mem2Seq typically ignore the rich structure inherent in the knowledge graph and the sentences in the conversation context. Inspired by the recent success of structure-aware Graph Convolutional Networks (GCNs) for various NLP tasks such as machine translation, semantic role labeling, and document dating, we propose a memory-augmented GCN for goal-oriented dialogues. Our model exploits (i) the entity relation graph in a knowledge-base and (ii) the dependency graph associated with an utterance to compute richer representations for words and entities. Further, we take cognizance of the fact that in certain situations, such as when the conversation is in a code-mixed language, dependency parsers may not be available. We show that in such situations we could use the global word co-occurrence graph to enrich the representations of utterances. We experiment with four datasets: (i) the modified DSTC2 dataset, (ii) recently released code-mixed versions of DSTC2 dataset in four languages, (iii) Wizard-of-Oz style CAM676 dataset, and (iv) Wizard-of-Oz style MultiWOZ dataset. On all four datasets our method outperforms existing methods, on a wide range of evaluation metrics. Suman Banerjee 0003, Mitesh M. Khapra |
Trans. Assoc. Comput. Linguistics | 2 |
| 2019 | Improving NER Tagging Performance in Low-Resource Languages via Multilingual LearningabstractExisting supervised solutions for Named Entity Recognition (NER) typically rely on a large annotated corpus. Collecting large amounts of NER annotated corpus is time-consuming and requires considerable human effort. However, collecting small amounts of annotated corpus for any language is feasible, but the performance degrades due to data sparsity. We address the data sparsity by borrowing features from the data of a closely related language. We use hierarchical neural networks to train a supervised NER system. The feature borrowing from a closely related language happens via the shared layers of the network. The neural network is trained on the combined dataset of the low-resource language and a closely related language, also termed Multilingual Learning. Unlike existing systems, we share all layers of the network between the two languages. We apply multilingual learning for NER in Indian languages and empirically show the benefits over a monolingual deep learning system and a traditional machine-learning system with some feature engineering. Using multilingual learning, we show that the low-resource language NER performance increases mainly due to (1) increased named entity vocabulary, (2) cross-lingual subword features, and (3) multilingual learning playing the role of regularization. V. Rudra Murthy, Mitesh M. Khapra, Pushpak Bhattacharyya |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2018 | Towards Building Large Scale Multimodal Domain-Aware Conversation SystemsabstractWhile multimodal conversation agents are gaining importance in several domains such as retail, travel etc., deep learning research in this area has been limited primarily due to the lack of availability of large-scale, open chatlogs. To overcome this bottleneck, in this paper we introduce the task of multimodal, domain-aware conversations, and propose the MMD benchmark dataset. This dataset was gathered by working in close coordination with large number of domain experts in the retail domain. These experts suggested various conversations flows and dialog states which are typically seen in multimodal conversations in the fashion domain. Keeping these flows and states in mind, we created a dataset consisting of over 150K conversation sessions between shoppers and sales agents, with the help of in-house annotators using a semi-automated manually intense iterative process. With this dataset, we propose 5 new sub-tasks for multimodal conversations along with their evaluation methodology. We also propose two multimodal neural models in the encode-attend-decode paradigm and demonstrate their performance on two of the sub-tasks, namely text response generation and best image response selection. These experiments serve to establish baseline performance and open new research directions for each of these sub-tasks. Further, for each of the sub-tasks, we present a 'per-state evaluation' of 9 most significant dialog states, which would enable more focused research into understanding the challenges and complexities involved in each of these states. Amrita Saha, Mitesh M. Khapra, Karthik Sankaranarayanan |
AAAI | 2 |
| 2018 | Complex Sequential Question Answering: Towards Learning to Converse Over Linked Question Answer Pairs with a Knowledge GraphabstractWhile conversing with chatbots, humans typically tend to ask many questions, a significant portion of which can be answered by referring to large-scale knowledge graphs (KG). While Question Answering (QA) and dialog systems have been studied independently, there is a need to study them closely to evaluate such real-world scenarios faced by bots involving both these tasks. Towards this end, we introduce the task of Complex Sequential QA which combines the two tasks of (i) answering factual questions through complex inferencing over a realistic-sized KG of millions of entities, and (ii) learning to converse through a series of coherently linked QA pairs. Through a labor intensive semi-automatic process, involving in-house and crowdsourced workers, we created a dataset containing around 200K dialogs with a total of 1.6M turns. Further, unlike existing large scale QA datasets which contain simple questions that can be answered from a single tuple, the questions in our dialogs require a larger subgraph of the KG. Specifically, our dataset has questions which require logical, quantitative, and comparative reasoning as well as their combinations. This calls for models which can: (i) parse complex natural language questions, (ii) use conversation context to resolve coreferences and ellipsis in utterances, (iii) ask for clarifications for ambiguous queries, and finally (iv) retrieve relevant subgraphs of the KG to answer such questions. However, our experiments with a combination of state of the art dialog and QA models show that they clearly do not achieve the above objectives and are inadequate for dealing with such complex real world settings. We believe that this new dataset coupled with the limitations of existing models as reported in this paper should encourage further research in Complex Sequential QA. Amrita Saha, Vardaan Pahuja, Mitesh M. Khapra, Karthik Sankaranarayanan, Sarath Chandar |
AAAI | 3 |
| 2018 | DuoRC: Towards Complex Language Understanding with Paraphrased Reading ComprehensionabstractWe propose DuoRC, a novel dataset for Reading Comprehension (RC) that motivates several new challenges for neural approaches in language understanding beyond those offered by existing RC datasets.DuoRC contains 186,089 unique questionanswer pairs created from a collection of 7680 pairs of movie plots where each pair in the collection reflects two versions of the same movie -one from Wikipedia and the other from IMDb -written by two different authors.We asked crowdsourced workers to create questions from one version of the plot and a different set of workers to extract or synthesize answers from the other version.This unique characteristic of DuoRC where questions and answers are created from different versions of a document narrating the same underlying story, ensures by design, that there is very little lexical overlap between the questions created from one version and the segments containing the answer in the other version.Further, since the two versions have different levels of plot detail, narration style, vocabulary, etc., answering questions from the second version requires deeper language understanding and incorporating external background knowledge.Additionally, the narrative style of passages arising from movie plots (as opposed to typical descriptive passages in existing datasets) exhibits the need to perform complex reasoning over events across multiple sentences.Indeed, we observe that state-of-the-art neural RC models which have achieved near human performance on the SQuAD dataset (Rajpurkar et al., 2016b), even when coupled with tra-ditional NLP techniques to address the challenges presented in DuoRC exhibit very poor performance (F1 score of 37.42% on DuoRC v/s 86% on SQuAD dataset).This opens up several interesting research avenues wherein DuoRC could complement other RC datasets to explore novel neural approaches for studying language understanding. Amrita Saha, Rahul Aralikatte, Mitesh M. Khapra, Karthik Sankaranarayanan |
ACL (1) | 3 |
| 2018 | A Dataset for Building Code-Mixed Goal Oriented Conversation SystemsabstractThere is an increasing demand for goal-oriented conversation systems which can assist users in various day-to-day activities such as booking tickets, restaurant reservations, shopping, etc. Most of the existing datasets for building such conversation systems focus on monolingual conversations and there is hardly any work on multilingual and/or code-mixed conversations. Such datasets and systems thus do not cater to the multilingual regions of the world, such as India, where it is very common for people to speak more than one language and seamlessly switch between them resulting in code-mixed conversations. For example, a Hindi speaking user looking to book a restaurant would typically ask, “Kya tum is restaurant mein ek table book karne mein meri help karoge?” (“Can you help me in booking a table at this restaurant?”). To facilitate the development of such code-mixed conversation models, we build a goal-oriented dialog dataset containing code-mixed conversations. Specifically, we take the text from the DSTC2 restaurant reservation dataset and create code-mixed versions of it in Hindi-English, Bengali-English, Gujarati-English and Tamil-English. We also establish initial baselines on this dataset using existing state of the art models. This dataset along with our baseline implementations will be made publicly available for research purposes. Suman Banerjee 0003, Nikita Moghe, Siddhartha Arora, Mitesh M. Khapra |
COLING | 4 |
| 2018 | Towards Exploiting Background Knowledge for Building Conversation SystemsabstractExisting dialog datasets contain a sequence of utterances and responses without any explicit background knowledge associated with them.This has resulted in the development of models which treat conversation as a sequenceto-sequence generation task (i.e., given a sequence of utterances generate the response sequence).This is not only an overly simplistic view of conversation but it is also emphatically different from the way humans converse by heavily relying on their background knowledge about the topic (as opposed to simply relying on the previous sequence of utterances).For example, it is common for humans to (involuntarily) produce utterances which are copied or suitably modified from background articles they have read about the topic.To facilitate the development of such natural conversation models which mimic the human process of conversing, we create a new dataset containing movie chats wherein each response is explicitly generated by copying and/or modifying sentences from unstructured background knowledge such as plots, comments and reviews about the movie.We establish baseline results on this dataset (90K utterances from 9K conversations) using three different models: (i) pure generation based models which ignore the background knowledge (ii) generation based models which learn to copy information from the background knowledge when required and (iii) span prediction based models which predict the appropriate response span in the background knowledge. Nikita Moghe, Siddhartha Arora, Suman Banerjee 0003, Mitesh M. Khapra |
EMNLP | 4 |
| 2018 | Towards a Better Metric for Evaluating Question Generation SystemsabstractThere has always been criticism for using ngram based similarity metrics, such as BLEU, NIST, etc, for evaluating the performance of NLG systems.However, these metrics continue to remain popular and are recently being used for evaluating the performance of systems which automatically generate questions from documents, knowledge graphs, images, etc.Given the rising interest in such automatic question generation (AQG) systems, it is important to objectively examine whether these metrics are suitable for this task.In particular, it is important to verify whether such metrics used for evaluating AQG systems focus on answerability of the generated question by preferring questions which contain all relevant information such as question type (Wh-types), entities, relations, etc.In this work, we show that current automatic evaluation metrics based on n-gram similarity do not always correlate well with human judgments about answerability of a question.To alleviate this problem and as a first step towards better evaluation metrics for AQG, we introduce a scoring function to capture answerability and show that when this scoring function is integrated with existing metrics, they correlate significantly better with human judgments.The scripts and data developed as a part of this work are made publicly available.1 Preksha Nema, Mitesh M. Khapra |
EMNLP | 2 |
| 2018 | ElimiNet: A Model for Eliminating Options for Reading Comprehension with Multiple Choice QuestionsabstractThe task of Reading Comprehension with Multiple Choice Questions, requires a human (or machine) to read a given {passage, question} pair and select one of the n given options. The current state of the art model for this task first computes a question-aware representation for the passage and then selects the option which has the maximum similarity with this representation. However, when humans perform this task they do not just focus on option selection but use a combination of elimination and selection. Specifically, a human would first try to eliminate the most irrelevant option and then read the passage again in the light of this new information (and perhaps ignore portions corresponding to the eliminated option). This process could be repeated multiple times till the reader is finally ready to select the correct option. We propose ElimiNet, a neural network-based model which tries to mimic this process. Specifically, it has gates which decide whether an option can be eliminated given the {passage, question} pair and if so it tries to make the passage representation orthogonal to this eliminated option (akin to ignoring portions of the passage corresponding to the eliminated option). The model makes multiple rounds of partial elimination to refine the passage representation and finally uses a selection module to pick the best option. We evaluate our model on the recently released large scale RACE dataset and show that it outperforms the current state of the art model on 7 out of the 13 question types in this dataset. Further, we show that taking an ensemble of our elimination-selection based method with a selection based method gives us an improvement of 3.1% over the best-reported performance on this dataset. Soham Parikh, Ananya Sai, Preksha Nema, Mitesh M. Khapra |
IJCAI | 4 |
| 2018 | Generating Descriptions from Structured Data Using a Bifocal Attention Mechanism and Gated OrthogonalizationabstractPreksha Nema, Shreyas Shetty, Parag Jain, Anirban Laha, Karthik Sankaranarayanan, Mitesh M. Khapra. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Preksha Nema, Shreyas Shetty, Parag Jain, Anirban Laha, Karthik Sankaranarayanan, Mitesh M. Khapra |
NAACL-HLT | 6 |
| 2018 | On Controllable Sparse Alternatives to SoftmaxabstractConverting an n-dimensional vector to a probability distribution over n objects is a commonly used component in many machine learning tasks like multiclass classification, multilabel classification, attention mechanisms etc. For this, several probability mapping functions have been proposed and employed in literature such as softmax, sum-normalization, spherical softmax, and sparsemax, but there is very little understanding in terms how they relate with each other. Further, none of the above formulations offer an explicit control over the degree of sparsity. To address this, we develop a unified framework that encompasses all these formulations as special cases. This framework ensures simple closed-form solutions and existence of sub-gradients suitable for learning via backpropagation. Within this framework, we propose two novel sparse formulations, sparsegen-lin and sparsehourglass, that seek to provide a control over the degree of desired sparsity. We further develop novel convex loss functions that help induce the behavior of aforementioned formulations in the multilabel classification setting, showing improved performance. We also demonstrate empirically that the proposed formulations, when used to compute attention weights, achieve better or comparable performance on standard seq2seq tasks like neural machine translation and abstractive summarization. Anirban Laha, Saneem A. Chemmengath, Priyanka Agrawal, Mitesh M. Khapra, Karthik Sankaranarayanan, Harish G. Ramaswamy |
NeurIPS | 4 |
| 2018 | Recovering from Random Pruning: On the Plasticity of Deep Convolutional Neural NetworksabstractRecently there has been a lot of work on pruning filters from deep convolutional neural networks (CNNs) with the intention of reducing computations. The key idea is to rank the filters based on a certain criterion (say, l1-norm, average percentage of zeros, etc) and retain only the top ranked filters. Once the low scoring filters are pruned away the remainder of the network is fine tuned and is shown to give performance comparable to the original unpruned network. In this work, we report experiments which suggest that the comparable performance of the pruned network is not due to the specific criterion chosen but due to the inherent plasticity of deep neural networks which allows them to recover from the loss of pruned filters once the rest of the filters are fine-tuned. Specifically, we show counter-intuitive results wherein by randomly pruning 25-50% filters from deep CNNs we are able to obtain the same performance as obtained by using state of the art pruning methods. We empirically validate our claims by doing an exhaustive evaluation with VGG-16 and ResNet-50. Further, we also evaluate a real world scenario where a CNN trained on all 1000 ImageNet classes needs to be tested on only a small set of classes at test time (say, only animals). We create a new benchmark dataset from ImageNet to evaluate such class specific pruning and show that even here a random pruning strategy gives close to state of the art performance. Lastly, unlike existing approaches which mainly focus on the task of image classification, in this work we also report results on object detection. We show that using a simple random pruning strategy we can achieve significant speed up in object detection (74% improvement in fps) while retaining the same accuracy as that of the original Faster RCNN model. Deepak Mittal, Shweta Bhardwaj, Mitesh M. Khapra, Balaraman Ravindran |
WACV | 3 |
| 2018 | Learning Disentangled Multimodal Representations for the Fashion DomainabstractIn many visual domains (like fashion, furniture, etc.) the search for products on online platforms requires matching textual queries to image content. For example, the user provides a search query in natural language (e.g.,pink floral top) and the results obtained are of a different modality (e.g., the set of images of pink floral tops). Recent work on multimodal representation learning enables such cross-modal matching by learning a common representation space for text and image. While such representations ensure that the n-dimensional representation of pink floral top is very close to representation of corresponding images, they do not ensure that the first k1(2(<; n) correspond to style and so on. In other words, they learn entangled representations where each dimension does not correspond to a specific attribute. We propose two simple variants which can learn disentangled common representations for the fashion domain wherein each dimension would correspond to a specific attribute (color, style, silhoutte, etc.). Our proposed variants can be integrated with any existing multimodal representation learning method. We use a large fashion dataset of over 700K fashion items crawled from multiple fashion e-commerce portals to evaluate the learned representations on four different applications from the fashion domain, namely, cross-modal image retrieval, visual search, image tagging, and query expansion. Our experimental results show that the proposed variants lead to better performance for each of these applications while learning disentangled representations. Amrita Saha, Megha Nawhal, Mitesh M. Khapra, Vikas C. Raykar |
WACV | 3 |
| 2018 | Leveraging Orthographic Similarity for Multilingual Neural TransliterationabstractWe address the task of joint training of transliteration models for multiple language pairs ( multilingual transliteration). This is an instance of multitask learning, where individual tasks (language pairs) benefit from sharing knowledge with related tasks. We focus on transliteration involving related tasks i.e., languages sharing writing systems and phonetic properties ( orthographically similar languages). We propose a modified neural encoder-decoder model that maximizes parameter sharing across language pairs in order to effectively leverage orthographic similarity. We show that multilingual transliteration significantly outperforms bilingual transliteration in different scenarios (average increase of 58% across a variety of languages we experimented with). We also show that multilingual transliteration models can generalize well to languages/language pairs not encountered during training and hence perform well on the zeroshot transliteration task. We show that further improvements can be achieved by using phonetic feature input. Anoop Kunchukuttan, Mitesh M. Khapra, Gurneet Singh, Pushpak Bhattacharyya |
Trans. Assoc. Comput. Linguistics | 2 |
| 2017 | Diversity driven attention model for query-based abstractive summarizationabstractAbstractive summarization aims to generate a shorter version of the document covering all the salient points in a compact and coherent fashion.On the other hand, query-based summarization highlights those points that are relevant in the context of a given query.The encodeattend-decode paradigm has achieved notable success in machine translation, extractive summarization, dialog systems, etc.But it suffers from the drawback of generation of repeated phrases.In this work we propose a model for the query-based summarization task based on the encode-attend-decode paradigm with two key additions (i) a query attention model (in addition to document attention model) which learns to focus on different portions of the query at different time steps (instead of using a static representation for the query) and (ii) a new diversity based attention model which aims to alleviate the problem of repeating phrases in the summary.In order to enable the testing of this model we introduce a new query-based summarization dataset building on debatepedia.Our experiments show that with these two additions the proposed model clearly outperforms vanilla encode-attend-decode models with a gain of 28% (absolute) in ROUGE-L scores. Preksha Nema, Mitesh M. Khapra, Anirban Laha, Balaraman Ravindran |
ACL (1) | 2 |
| 2017 | Generating Natural Language Question-Answer Pairs from a Knowledge Graph Using a RNN Based Question Generation ModelabstractSathish Reddy, Dinesh Raghu, Mitesh M. Khapra, Sachindra Joshi. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Sathish Reddy, Dinesh Raghu, Mitesh M. Khapra, Sachindra Joshi |
EACL (1) | 3 |
| 2017 | Attend, Adapt and Transfer: Attentive Deep Architecture for Adaptive Transfer from multiple sources in the same domain
Janarthanan Rajendran, Aravind S. Lakshminarayanan, Mitesh M. Khapra, P. Prasanna, Balaraman Ravindran |
ICLR (Poster) | 3 |
| 2016 | A Correlational Encoder Decoder Architecture for Pivot Based Sequence GenerationabstractInterlingua based Machine Translation (MT) aims to encode multiple languages into a common linguistic representation and then decode sentences in multiple target languages from this representation. In this work we explore this idea in the context of neural encoder decoder architectures, albeit on a smaller scale and without MT as the end goal. Specifically, we consider the case of three languages or modalities X, Z and Y wherein we are interested in generating sequences in Y starting from information available in X. However, there is no parallel training data available between X and Y but, training data is available between X & Z and Z & Y (as is often the case in many real world applications). Z thus acts as a pivot/bridge. An obvious solution, which is perhaps less elegant but works very well in practice is to train a two stage model which first converts from X to Z and then from Z to Y. Instead we explore an interlingua inspired solution which jointly learns to do the following (i) encode X and Z to a common representation and (ii) decode Y from this common representation. We evaluate our model on two tasks: (i) bridge transliteration and (ii) bridge captioning. We report promising results in both these applications and believe that this is a right step towards truly interlingua inspired encoder decoder architectures. Amrita Saha, Mitesh M. Khapra, Sarath Chandar, Janarthanan Rajendran, Kyunghyun Cho |
COLING | 2 |
| 2016 | Substring-based unsupervised transliteration with phonetic and contextual knowledge
Anoop Kunchukuttan, Pushpak Bhattacharyya, Mitesh M. Khapra |
CoNLL | 3 |
| 2016 | Bridge Correlational Neural Networks for Multilingual Multimodal Representation LearningabstractJanarthanan Rajendran, Mitesh M. Khapra, Sarath Chandar, Balaraman Ravindran. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Janarthanan Rajendran, Mitesh M. Khapra, Sarath Chandar, Balaraman Ravindran |
HLT-NAACL | 2 |
| 2016 | Correlational Neural NetworksabstractCommon representation learning (CRL), wherein different descriptions (or views) of the data are embedded in a common subspace, has been receiving a lot of attention recently. Two popular paradigms here are canonical correlation analysis (CCA)-based approaches and autoencoder (AE)-based approaches. CCA-based approaches learn a joint representation by maximizing correlation of the views when projected to the common subspace. AE-based methods learn a common representation by minimizing the error of reconstructing the two views. Each of these approaches has its own advantages and disadvantages. For example, while CCA-based approaches outperform AE-based approaches for the task of transfer learning, they are not as scalable as the latter. In this work, we propose an AE-based approach, correlational neural network (CorrNet), that explicitly maximizes correlation among the views when projected to the common subspace. Through a series of experiments, we demonstrate that the proposed CorrNet is better than AE and CCA with respect to its ability to learn correlated common representations. We employ CorrNet for several cross-language tasks and show that the representations learned using it perform better than the ones learned using other state-of-the-art approaches. Sarath Chandar, Mitesh M. Khapra, Hugo Larochelle, Balaraman Ravindran |
Neural Comput. | 2 |
| 2015 | Show Me Your Evidence - an Automatic Method for Context Dependent Evidence DetectionabstractEngaging in a debate with oneself or others to take decisions is an integral part of our day-today life.A debate on a topic (say, use of performance enhancing drugs) typically proceeds by one party making an assertion/claim (say, PEDs are bad for health) and then providing an evidence to support the claim (say, a 2006 study shows that PEDs have psychiatric side effects).In this work, we propose the task of automatically detecting such evidences from unstructured text that support a given claim.This task has many practical applications in decision support and persuasion enhancement in a wide range of domains.We first introduce an extensive benchmark data set tailored for this task, which allows training statistical models and assessing their performance.Then, we suggest a system architecture based on supervised learning to address the evidence detection task.Finally, promising experimental results are reported. Ruty Rinott, Lena Dankin, Carlos Alzate, Mitesh M. Khapra, Ehud Aharoni, Noam Slonim |
EMNLP | 4 |
| 2014 | When Transliteration Met Crowdsourcing : An Empirical Study of Transliteration via Crowdsourcing using Efficient, Non-redundant and Fair Quality Control
Mitesh M. Khapra, Ananthakrishnan Ramanathan, Anoop Kunchukuttan, Karthik Visweswariah, Pushpak Bhattacharyya |
LREC | 1 |
| 2014 | An Autoencoder Approach to Learning Bilingual Word Representations
Sarath Chandar, Stanislas Lauly, Hugo Larochelle, Mitesh M. Khapra, Balaraman Ravindran, Vikas C. Raykar, Amrita Saha |
NIPS | 4 |
| 2013 | Cut the noise: Mutually reinforcing reordering and alignments for improved machine translation
Karthik Visweswariah, Mitesh M. Khapra, Ananthakrishnan Ramanathan |
ACL (1) | 2 |
| 2013 | Lost in Translation: Viability of Machine Translation for Cross Language Sentiment Analysis
A. R. Balamurali, Mitesh M. Khapra, Pushpak Bhattacharyya |
CICLing (2) | 2 |
| 2013 | Improving reordering performance using higher order and structural features
Mitesh M. Khapra, Ananthakrishnan Ramanathan, Karthik Visweswariah |
HLT-NAACL | 1 |
| 2012 | Experiences in Resource Generation for Machine Translation through Crowdsourcing
Anoop Kunchukuttan, Shourya Roy, Pratik Patel, Kushal Ladha, Somya Gupta, Mitesh M. Khapra, Pushpak Bhattacharyya |
LREC | 6 |
| 2011 | Together We Can: Bilingual Bootstrapping for WSD
Mitesh M. Khapra, Salil Joshi 0001, Arindam Chatterjee, Pushpak Bhattacharyya |
ACL | 1 |
| 2011 | It Takes Two to Tango: A Bilingual Unsupervised Approach for Estimating Sense Distributions using Expectation Maximization
Mitesh M. Khapra, Salil Joshi 0001, Pushpak Bhattacharyya |
IJCNLP | 1 |
| 2010 | PR + RQ ALMOST EQUAL TO PQ: Transliteration Mining Using Bridge LanguageabstractWe address the problem of mining name transliterations from comparable corpora in languages P and Q in the following resource-poor scenario:Parallel names in PQ are not available for training. Parallel names in PR and RQ are available for training.We propose a novel solution for the problem by computing a common geometric feature space for P,Q and R where name transliterations are mapped to similar vectors. We employ Canonical Correlation Analysis (CCA) to compute the common geometric feature space using only parallel names in PR and RQ and without requiring parallel names in PQ. We test our algorithm on data sets in several languages and show that it gives results comparable to the state-of-the-art transliteration mining algorithms that use parallel names in PQ for training. Mitesh M. Khapra, Raghavendra Udupa, A. Kumaran 0001, Pushpak Bhattacharyya |
AAAI | 1 |
| 2010 | All Words Domain Adapted WSD: Finding a Middle Ground between Supervision and Unsupervision
Mitesh M. Khapra, Anup Kulkarni, Saurabh Sohoney, Pushpak Bhattacharyya |
ACL | 1 |
| 2010 | Value for Money: Balancing Annotation Effort, Lexicon Building and Accuracy for Multilingual WSD
Mitesh M. Khapra, Saurabh Sohoney, Anup Kulkarni, Pushpak Bhattacharyya |
COLING | 1 |
| 2010 | Transliteration Equivalence Using Canonical Correlation Analysis
Raghavendra Udupa, Mitesh M. Khapra |
ECIR | 2 |
| 2010 | Everybody loves a rich cousin: An empirical study of transliteration through bridge languages
Mitesh M. Khapra, A. Kumaran 0001, Pushpak Bhattacharyya |
HLT-NAACL | 1 |
| 2010 | Improving the Multilingual User Experience of Wikipedia Using Cross-Language Name Search
Raghavendra Udupa, Mitesh M. Khapra |
HLT-NAACL | 2 |
| 2010 | Compositional Machine TransliterationabstractMachine transliteration is an important problem in an increasingly multilingual world, as it plays a critical role in many downstream applications, such as machine translation or crosslingual information retrieval systems. In this article, we propose compositional machine transliteration systems, where multiple transliteration components may be composed either to improve existing transliteration quality, or to enable transliteration functionality between languages even when no direct parallel names corpora exist between them. Specifically, we propose two distinct forms of composition: serial and parallel. Serial compositional system chains individual transliteration components, say, X → Y and Y → Z systems, to provide transliteration functionality, X → Z. In parallel composition evidence from multiple transliteration paths between X → Z are aggregated for improving the quality of a direct system. We demonstrate the functionality and performance benefits of the compositional methodology using a state-of-the-art machine transliteration framework in English and a set of Indian languages, namely, Hindi, Marathi, and Kannada. Finally, we underscore the utility and practicality of our compositional approach by showing that a CLIR system integrated with compositional transliteration systems performs consistently on par with, and sometimes better than, that integrated with a direct transliteration system. A. Kumaran 0001, Mitesh M. Khapra, Pushpak Bhattacharyya |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2009 | Projecting Parameters for Multilingual Word Sense Disambiguation
Mitesh M. Khapra, Sapan Shah, Piyush Kedia, Pushpak Bhattacharyya |
EMNLP | 1 |