Diptesh Kanojia

dblp:127/0183 · DBLP profile ↗
← Back
47ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0001-8814-0080ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 10 first-author · 22 since 2021Databases, data management, data science and information retrieval · 11 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 since 2021
YearPublicationVenuePosition
2026 TRACE: Textual Relevance Augmentation and Contextual Encoding for Multimodal Hate Detection
abstract
Social media memes are a challenging domain for hate detection because they intertwine visual and textual cues into culturally nuanced messages. To tackle these challenges, we introduce TRACE, a hierarchical multimodal framework that leverages visually grounded context augmentation, along with a novel caption-scoring network to emphasize hate-relevant content, and parameter-efficient fine-tuning of CLIP’s text encoder. Our experiments demonstrate that selectively fine-tuning deeper text encoder layers significantly enhances performance compared to simpler projection-layer fine-tuning methods. Specifically, our framework achieves state-of-the-art accuracy (0.807) and F1-score (0.806) on the widely-used Hateful Memes dataset, matching the performance of considerably larger models while maintaining efficiency. Moreover, it achieves superior generalization on the MultiOFF offensive meme dataset (F1-score 0.673), highlighting robustness across meme categories. Additional analyses confirm that robust visual grounding and nuanced text representations significantly reduce errors caused by benign confounders. We publicly release our code to facilitate future research.
Girish A. Koushik, Helen Treharne, Aditya Joshi 0001, Diptesh Kanojia
AAAI4
2026 Bridging Domains for Automatic Post-Editing: A Classifier-Guided Multi-Domain Adaptation Framework
abstract
Automatic Post-Editing (APE) is a widely studied approach for enhancing the output quality of Neural Machine Translation (NMT) systems. While most prior work has focused on general-purpose APE, the potential of domain-specific APE, such as for personalized or specialized content, remains underexplored due to the scarcity of domain-labeled training data. In this work, we investigate domain adaptation for APE using adapter-based methods. Our proposed multitask learning-based domain adaptation framework includes the use of a domain classifier to get a weighted combination of parallel domain-specific adapters at inference time, without requiring prior domain knowledge. This design allows the model to leverage cross-domain similarities, making it especially robust in low-resource domain scenarios. Our experimental results on English–German, English–Marathi, and English–Tamil pairs across different domains for each pair show substantial improvements over their respective general-purpose APE baselines. To facilitate further research, we will release human-annotated domain labels for triplets in WMT22 English–Marathi, and WMT24 English–Tamil APE datasets and the code.
Sourabh Dattatray Deoghare, Diptesh Kanojia, Pushpak Bhattacharyya
EAMT (1)2
2026 Improving Search Suggestions for Alphanumeric Queries
Samarth Agrawal, Jayanth Yetukuri, Diptesh Kanojia, Qunzhi Zhou
ECIR (4)3
2026 MUNIChus: MUltilingual News Image Captioning Benchmark
Yuji Chen, Alistair Plum, Hansi Hettiarachchi, Diptesh Kanojia, Saroj Basnet, Marcos Zampieri, Tharindu Ranasinghe
LREC4
2025 Unsupervised Audio-Visual Segmentation with Modality Alignment
abstract
Audio-Visual Segmentation (AVS) aims to identify, at the pixel level, the object in a visual scene that produces a given sound. Current AVS methods rely on costly fine-grained annotations of mask-audio pairs, making them impractical for scalability. To address this, we propose the Modality Correspondence Alignment (MoCA) framework, which seamlessly integrates off-the-shelf foundation models like DINO, SAM, and ImageBind. Our approach leverages existing knowledge within these models and optimizes their joint usage for multimodal associations. Our approach relies on estimating positive and negative image pairs in the feature space. For pixel-level association, we introduce an audio-visual adapter and a novel {pixel matching aggregation} strategy within the image-level contrastive learning framework. This allows for a flexible connection between object appearance and audio signal at the pixel level, with tolerance to imaging variations such as translation and rotation. Extensive experiments on the AVSBench (single and multi-object splits) and AVSS datasets demonstrate that MoCA outperforms unsupervised baseline approaches and some supervised counterparts, particularly in complex scenarios with multiple auditory objects. In terms of mIoU, MoCA achieves a substantial improvement over baselines in both the AVSBench (S4: +17.24%, MS3: +67.64%) and AVSS (+19.23%) audio-visual segmentation challenges.
Swapnil Bhosale, Haosen Yang 0003, Diptesh Kanojia, Jiankang Deng, Xiatian Zhu
AAAI3
2025 Refer to the Reference: Reference-focused Synthetic Automatic Post-Editing Data Generation
abstract
A prevalent approach to synthetic APE data generation uses source (src) sentences in a parallel corpus to obtain translations (mt) through an MT system and treats corresponding reference (ref) sentences as post-edits (pe). While effective, due to independence between ‘mt’ and ‘pe,’ these translations do not adequately reflect errors to be corrected by a human post-editor. Thus, we introduce a novel and simple yet effective reference-focused synthetic APE data generation technique that uses ‘ref’ instead of src’ sentences to obtain corrupted translations (mt_new). The experimental results across English-German, English-Russian, English-Marathi, English-Hindi, and English-Tamil language pairs demonstrate the superior performance of APE systems trained using the newly generated synthetic data compared to those trained using existing synthetic data. Further, APE models trained using a balanced mix of existing and newly generated synthetic data achieve improvements of 0.37, 0.19, 1.01, 2.42, and 2.60 TER points, respectively. We will release the generated synthetic APE data.
Sourabh Dattatray Deoghare, Diptesh Kanojia, Pushpak Bhattacharyya
COLING2
2025 Prompt-based Explainable Quality Estimation for English-Malayalam
abstract
The aim of this project was to curate data for the English-Malayalam language pair for the tasks of Quality Estimation (QE) and Automatic Post-Editing (APE) of Machine Translation. Whilst the primary aim of the project was to create a dataset for a low-resource language pair, we plan to use this dataset to investigate different zero-shot and few-shot prompting strategies including chain-of-thought, towards a unified explainable QE-APE framework.
Archchana Sindhujan, Diptesh Kanojia, Constantin Orasan
MTSummit (2)2
2024 DiffSED: Sound Event Detection with Denoising Diffusion
abstract
Sound Event Detection (SED) aims to predict the temporal boundaries of all the events of interest and their class labels, given an unconstrained audio sample. Taking either the split-and-classify (i.e., frame-level) strategy or the more principled event-level modeling approach, all existing methods consider the SED problem from the discriminative learning perspective. In this work, we reformulate the SED problem by taking a generative learning perspective. Specifically, we aim to generate sound temporal boundaries from noisy proposals in a denoising diffusion process, conditioned on a target audio sample. During training, our model learns to reverse the noising process by converting noisy latent queries to the ground-truth versions in the elegant Transformer decoder framework. Doing so enables the model generate accurate event boundaries from even noisy queries during inference. Extensive experiments on the Urban-SED and EPIC-Sounds datasets demonstrate that our model significantly outperforms existing alternatives, with 40+% faster convergence in training. Code: https://github.com/Surrey-UPLab/DiffSED
Swapnil Bhosale, Sauradip Nag, Diptesh Kanojia, Jiankang Deng, Xiatian Zhu
AAAI3
2024 Product Retrieval and Ranking for Alphanumeric Queries
abstract
This talk addresses the challenge of improving user experience on e-commerce platforms by enhancing product ranking relevant to user's search queries. Queries such as S2716DG consist of alphanumeric characters where a letter or number can signify important detail for the product/model. Speaker describes recent research where we curate samples from existing datasets at eBay, manually annotated with buyer-centric relevance scores, and centrality scores which reflect how well the product title matches the user's intent. We introduce a User-intent Centrality Optimization (UCO) approach for existing models, which optimizes for the user intent in semantic product search. To that end, we propose a dual-loss based optimization to handle hard negatives, i.e., product titles that are semantically relevant but do not reflect the user's intent. Our contributions include curating a challenging evaluation set and implementing UCO, resulting in significant improvements in product ranking efficiency, observed for different evaluation metrics. Our work aims to ensure that the most buyer-centric titles for a query are ranked higher, thereby, enhancing the user experience on e-commerce platforms.
Hadeel Saadany, Swapnil Bhosale, Samarth Agrawal, Constantin Orasan, Diptesh Kanojia
CIKM6
2024 Character-level Language Models for Abbreviation and Long-form Detection
Leonardo Zilio, Shenbin Qian, Diptesh Kanojia, Constantin Orasan
LREC/COLING3
2024 Evaluating Machine Translation for Emotion-loaded User Generated Content (TransEval4Emo-UGC)
abstract
This paper presents a dataset for evaluating the machine translation of emotion-loaded user generated content. It contains human-annotated quality evaluation data and post-edited reference translations. The dataset is available at our GitHub repository.
Shenbin Qian, Constantin Orasan, Félix do Carmo, Diptesh Kanojia
EAMT (2)4
2024 What do Large Language Models Need for Machine Translation Evaluation?
abstract
Shenbin Qian, Archchana Sindhujan, Minnie Kabra, Diptesh Kanojia, Constantin Orasan, Tharindu Ranasinghe, Fred Blain. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Shenbin Qian, Archchana Sindhujan, Minnie Kabra, Diptesh Kanojia, Constantin Orasan, Tharindu Ranasinghe, Frédéric Blain
EMNLP4
2024 StableTalk: Advancing Audio-to-Talking Face Generation with Stable Diffusion and Vision Transformer
Fatemeh Nazarieh, Josef Kittler, Muhammad Awais 0001, Diptesh Kanojia, Zhenhua Feng 0001
ICPR (6)4
2024 A Survey of Multimodal Sarcasm Detection
Shafkat Farabi, Tharindu Ranasinghe, Diptesh Kanojia, Yu Kong 0001, Marcos Zampieri
IJCAI3
2024 AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic Synthesis
abstract
Novel view acoustic synthesis (NVAS) aims to render binaural audio at any target viewpoint, given a mono audio emitted by a sound source at a 3D scene. Existing methods have proposed NeRF-based implicit models to exploit visual cues as a condition for synthesizing binaural audio. However, in addition to low efficiency originating from heavy NeRF rendering, these methods all have a limited ability of characterizing the entire scene environment such as room geometry, material properties, and the spatial relation between the listener and sound source. To address these issues, we propose a novel Audio-Visual Gaussian Splatting (AV-GS) model. To obtain a material-aware and geometry-aware condition for audio synthesis, we learn an explicit point-based scene representation with audio-guidance parameters on locally initialized Gaussian points, taking into account the space relation from the listener and sound source. To make the visual scene model audio adaptive, we propose a point densification and pruning strategy to optimally distribute the Gaussian points, with the per-point contribution in sound propagation (e.g., more points needed for texture-less wall surfaces as they affect sound path diversion). Extensive experiments validate the superiority of our AV-GS over existing alternatives on the real-world RWAS and simulation-based SoundSpaces datasets. Project page: \url{https://surrey-uplab.github.io/research/avgs/}
Swapnil Bhosale, Haosen Yang 0003, Diptesh Kanojia, Jiankang Deng, Xiatian Zhu
NeurIPS3
2024 CreoleVal: Multilingual Multitask Benchmarks for Creoles
abstract
Abstract Creoles represent an under-explored and marginalized group of languages, with few available resources for NLP research. While the genealogical ties between Creoles and a number of highly resourced languages imply a significant potential for transfer learning, this potential is hampered due to this lack of annotated data. In this work we present CreoleVal, a collection of benchmark datasets spanning 8 different NLP tasks, covering up to 28 Creole languages; it is an aggregate of novel development datasets for reading comprehension relation classification, and machine translation for Creoles, in addition to a practical gateway to a handful of preexisting benchmarks. For each benchmark, we conduct baseline experiments in a zero-shot setting in order to further ascertain the capabilities and limitations of transfer learning for Creoles. Ultimately, we see CreoleVal as an opportunity to empower research on Creoles in NLP and computational linguistics, and in general, a step towards more equitable language technology around the globe.
Heather C. Lent, Kushal Tatariya, Raj Dabre, Yiyi Chen 0002, Marcell Fekete, Esther Ploeger, Li Zhou 0010, Ruth-Ann Armstrong, Abee Eijansantos, Catriona Malau, Hans Erik Heje, Ernests Lavrinovics, Diptesh Kanojia, Paul Belony, Marcel Bollmann, Loïc Grobol, Miryam de Lhoneux, Daniel Hershcovich, Michel DeGraff, Anders Søgaard, Johannes Bjerva
Trans. Assoc. Comput. Linguistics13
2023 Evaluation of Chinese-English Machine Translation of Emotion-Loaded Microblog Texts: A Human Annotated Dataset for the Quality Assessment of Emotion Translation
abstract
In this paper, we focus on how current Machine Translation (MT) engines perform on the translation of emotion-loaded texts by evaluating outputs from Google Translate according to a framework proposed in this paper. We propose this evaluation framework based on the Multidimensional Quality Metrics (MQM) and perform detailed error analyses of the MT outputs. From our analysis, we observe that about 50% of MT outputs are erroneous in preserving emotions. After further analysis of the erroneous examples, we find that emotion carrying words and linguistic phenomena such as polysemous words, negation, abbreviation etc., are common causes for these translation errors.
Shenbin Qian, Constantin Orasan, Félix do Carmo, Qiuliang Li, Diptesh Kanojia
EAMT5
2023 Predict and Use: Harnessing Predicted Gaze to Improve Multimodal Sarcasm Detection
abstract
Sarcasm is a complex linguistic construct with incongruity at its very core.Detecting sarcasm depends on the actual content spoken and tonality, facial expressions, the context of an utterance, and personal traits like language proficiency and cognitive capabilities.In this paper, we propose the utilization of synthetic gaze data to improve the task performance for multimodal sarcasm detection in a conversational setting.We enrich an existing multimodal conversational dataset, i.e., MUStARD++ with gaze features.With the help of human participants, we collect gaze features for < 20% of data instances, and we investigate various methods for gaze feature prediction for the rest of the dataset.We perform extrinsic and intrinsic evaluations to assess the quality of the predicted gaze features.We observe a performance gain of up to 6.6% points by adding a new modality, i.e., collected gaze features.When both collected and predicted data are used, we observe a performance gain of 2.3% points on the complete dataset.Interestingly, with only predicted gaze features, too, we observe a gain in performance (1.9% points).We retain and use the feature prediction model, which maximally correlates with collected gaze features.Our model trained on combining collected and synthetic gaze data achieves SoTA performance on the MUStARD++ dataset.To the best of our knowledge, ours is the first predict-and-use model for sarcasm detection.We publicly release 1 the code, gaze data, and our best models for further research.
Divyank Tiwari, Diptesh Kanojia, Anupama Ray, Apoorva Nunna, Pushpak Bhattacharyya
EMNLP2
2022 Harnessing Abstractive Summarization for Fact-Checked Claim Detection
abstract
Social media platforms have become new battlegrounds for anti-social elements, with misinformation being the weapon of choice. Fact-checking organizations try to debunk as many claims as possible while staying true to their journalistic processes but cannot cope with its rapid dissemination. We believe that the solution lies in partial automation of the fact-checking life cycle, saving human time for tasks which require high cognition. We propose a new workflow for efficiently detecting previously fact-checked claims that uses abstractive summarization to generate crisp queries. These queries can then be executed on a general-purpose retrieval system associated with a collection of previously fact-checked claims. We curate an abstractive text summarization dataset comprising noisy claims from Twitter and their gold summaries. It is shown that retrieval performance improves 2x by using popular out-of-the-box summarization models and 3x by fine-tuning them on the accompanying dataset compared to verbatim querying. Our approach achieves Recall@5 and MRR of 35% and 0.3, compared to baseline values of 10% and 0.1, respectively. Our dataset, code, and models are available publicly: https://github.com/varadhbhatnagar/FC-Claim-Det/.
Varad Bhatnagar, Diptesh Kanojia, Kameswari Chebrolu
COLING2
2022 HiNER: A large Hindi Named Entity Recognition Dataset
abstract
Named Entity Recognition (NER) is a foundational NLP task that aims to provide class labels like Person, Location, Organisation, Time, and Number to words in free text. Named Entities can also be multi-word expressions where the additional I-O-B annotation information helps label them during the NER annotation process. While English and European languages have considerable annotated data for the NER task, Indian languages lack on that front- both in terms of quantity and following annotation standards. This paper releases a significantly sized standard-abiding Hindi NER dataset containing 109,146 sentences and 2,220,856 tokens, annotated with 11 tags. We discuss the dataset statistics in all their essential detail and provide an in-depth analysis of the NER tag-set used with our data. The statistics of tag-set in our dataset shows a healthy per-tag distribution especially for prominent classes like Person, Location and Organisation. Since the proof of resource-effectiveness is in building models with the resource and testing the model on benchmark data and against the leader-board entries in shared tasks, we do the same with the aforesaid data. We use different language models to perform the sequence labelling task for NER and show the efficacy of our data by performing a comparative evaluation with models trained on another dataset available for the Hindi NER task. Our dataset helps achieve a weighted F1 score of 88.78 with all the tags and 92.22 when we collapse the tag-set, as discussed in the paper. To the best of our knowledge, no available dataset meets the standards of volume (amount) and variability (diversity), as far as Hindi NER is concerned. We fill this gap through this work, which we hope will significantly help NLP for Hindi. We release this dataset with our code and models for further research at https://github.com/cfiltnlp/HiNER
V. Rudra Murthy, Pallab Bhattacharjee, Rahul Sharnagat, Jyotsana Khatri, Diptesh Kanojia, Pushpak Bhattacharyya
LREC5
2022 PLOD: An Abbreviation Detection Dataset for Scientific Documents
abstract
The detection and extraction of abbreviations from unstructured texts can help to improve the performance of Natural Language Processing tasks, such as machine translation and information retrieval. However, in terms of publicly available datasets, there is not enough data for training deep-neural-networks-based models to the point of generalising well over data. This paper presents PLOD, a large-scale dataset for abbreviation detection and extraction that contains 160k+ segments automatically annotated with abbreviations and their long forms. We performed manual validation over a set of instances and a complete automatic validation for this dataset. We then used it to generate several baseline models for detecting abbreviations and long forms. The best models achieved an F1-score of 0.92 for abbreviations and 0.89 for detecting their corresponding long forms. We release this dataset along with our code and all the models publicly at https://github.com/surrey-nlp/PLOD-AbbreviationDetection
Leonardo Zilio, Hadeel Saadany, Diptesh Kanojia, Constantin Orasan
LREC4
2021 Cognition-aware Cognate Detection
abstract
Diptesh Kanojia, Prashant Sharma, Sayali Ghodekar, Pushpak Bhattacharyya, Gholamreza Haffari, Malhar Kulkarni. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Diptesh Kanojia, Sayali Ghodekar, Pushpak Bhattacharyya, Gholamreza Haffari, Malhar Kulkarni
EACL1
2021 "So You Think You're Funny?": Rating the Humour Quotient in Standup Comedy
abstract
Computational Humour (CH) has attracted the interest of Natural Language Processing and Computational Linguistics communities.Creating datasets for automatic measurement of humour quotient is difficult due to multiple possible interpretations of the content.In this work, we create a multi-modal humourannotated dataset (∼40 hours) using stand-up comedy clips.We devise a novel scoring mechanism to annotate the training data with a humour quotient score using the audience's laughter.The normalized duration (laughter duration divided by the clip duration) of laughter in each clip is used to compute this humour coefficient score on a five-point scale (0-4).This method of scoring is validated by comparing with manually annotated scores, wherein a quadratic weighted kappa of 0.6 is obtained.We use this dataset to train a model that provides a "funniness" score, on a five-point scale, given the audio and its corresponding text.We compare various neural language models for the task of humour-rating and achieve an accuracy of 0.813 in terms of Quadratic Weighted Kappa (QWK).Our "Open Mic" dataset is released for further research along with the code.
Anirudh Mittal, Pranav Jeevan, Prerak Gandhi, Diptesh Kanojia, Pushpak Bhattacharyya
EMNLP (1)4
2020 Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages
abstract
Cognates are variants of the same lexical form across different languages; for example "fonema" in Spanish and "phoneme" in English are cognates, both of which mean "a unit of sound".The task of automatic detection of cognates among any two languages can help downstream NLP tasks such as Cross-lingual Information Retrieval, Computational Phylogenetics, and Machine Translation.In this paper, we demonstrate the use of cross-lingual word embeddings for detecting cognates among fourteen Indian Languages.Our approach introduces the use of context from a knowledge graph to generate improved feature representations for cognate detection.We then evaluate the impact of our cognate detection mechanism on neural machine translation (NMT), as a downstream task.We evaluate our methods to detect cognates on a challenging dataset of twelve Indian languages, namely, Sanskrit, Hindi, Assamese, Oriya, Kannada, Gujarati, Tamil, Telugu, Punjabi, Bengali, Marathi, and Malayalam.Additionally, we create evaluation datasets for two more Indian languages, Konkani and Nepali 1 .We observe an improvement of up to 18% points, in terms of F-score, for cognate detection.Furthermore, we observe that cognates extracted using our method help improve NMT quality by up to 2.76 BLEU.We also release 2 our code, newly constructed datasets and cross-lingual models publicly.
Diptesh Kanojia, Raj Dabre, Shubham Dewangan, Pushpak Bhattacharyya, Gholamreza Haffari, Malhar Kulkarni
COLING1
2020 A Survey on Using Gaze Behaviour for Natural Language Processing
abstract
Gaze behaviour has been used as a way to gather cognitive information for a number of years. In this paper, we discuss the use of gaze behaviour in solving different tasks in natural language processing (NLP) without having to record it at test time. This is because the collection of gaze behaviour is a costly task, both in terms of time and money. Hence, in this paper, we focus on research done to alleviate the need for recording gaze behaviour at run time. We also mention different eye tracking corpora in multiple languages, which are currently available and can be used in natural language processing. We conclude our paper by discussing applications in a domain - education - and how learning gaze behaviour can help in solving the tasks of complex word identification and automatic essay grading.
Sandeep Mathias, Diptesh Kanojia, Abhijit Mishra, Pushpak Bhattacharyya
IJCAI2
2020 Challenge Dataset of Cognates and False Friend Pairs from Indian Languages
abstract
Cognates are present in multiple variants of the same text across different languages (e.g., “hund” in German and “hound” in the English language mean “dog”). They pose a challenge to various Natural Language Processing (NLP) applications such as Machine Translation, Cross-lingual Sense Disambiguation, Computational Phylogenetics, and Information Retrieval. A possible solution to address this challenge is to identify cognates across language pairs. In this paper, we describe the creation of two cognate datasets for twelve Indian languages namely Sanskrit, Hindi, Assamese, Oriya, Kannada, Gujarati, Tamil, Telugu, Punjabi, Bengali, Marathi, and Malayalam. We digitize the cognate data from an Indian language cognate dictionary and utilize linked Indian language Wordnets to generate cognate sets. Additionally, we use the Wordnet data to create a False Friends’ dataset for eleven language pairs. We also evaluate the efficacy of our dataset using previously available baseline cognate detection approaches. We also perform a manual evaluation with the help of lexicographers and release the curated gold-standard dataset with this paper.
Diptesh Kanojia, Malhar Kulkarni, Pushpak Bhattacharyya, Gholamreza Haffari
LREC1
2020 Recommendation Chart of Domains for Cross-Domain Sentiment Analysis: Findings of A 20 Domain Study
abstract
Cross-domain sentiment analysis (CDSA) helps to address the problem of data scarcity in scenarios where labelled data for a domain (known as the target domain) is unavailable or insufficient. However, the decision to choose a domain (known as the source domain) to leverage from is, at best, intuitive. In this paper, we investigate text similarity metrics to facilitate source domain selection for CDSA. We report results on 20 domains (all possible pairs) using 11 similarity metrics. Specifically, we compare CDSA performance with these metrics for different domain-pairs to enable the selection of a suitable source domain, given a target domain. These metrics include two novel metrics for evaluating domain adaptability to help source domain selection of labelled data and utilize word and sentence-based embeddings as metrics for unlabelled data. The goal of our experiments is a recommendation chart that gives the K best source domains for CDSA for a given target domain. We show that the best K source domains returned by our similarity metrics have a precision of over 50%, for varying values of K.
Akash Sheoran, Diptesh Kanojia, Aditya Joshi 0001, Pushpak Bhattacharyya
LREC2
2019 Utilizing Wordnets for Cognate Detection among Indian Languages
abstract
Automatic Cognate Detection (ACD) is a challenging task which has been utilized to help NLP applications like Machine Translation, Information Retrieval and Computational Phylogenetics.Unidentified cognate pairs can pose a challenge to these applications and result in a degradation of performance.In this paper, we detect cognate word pairs among ten Indian languages with Hindi and use deep learning methodologies to predict whether a word pair is cognate or not.We identify IndoWordnet as a potential resource to detect cognate word pairs based on orthographic similarity-based methods and train neural network models using the data obtained from it.We identify parallel corpora as another potential resource and perform the same experiments for them.We also validate the contribution of Wordnets through further experimentation and report improved performance of up to 26%.We discuss the nuances of cognate detection among closely related Indian languages and release the lists of detected cognates as a dataset.We also observe the behaviour of, to an extent, unrelated Indian language pairs and release the lists of detected cognates among them as well.
Diptesh Kanojia, Kevin Patel, Malhar Kulkarni, Pushpak Bhattacharyya, Gholamreza Haffari
GWC1
2018 Eyes are the Windows to the Soul: Predicting the Rating of Text Quality Using Gaze Behaviour
abstract
Sandeep Mathias, Diptesh Kanojia, Kevin Patel, Samarth Agrawal, Abhijit Mishra, Pushpak Bhattacharyya. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Sandeep Mathias, Diptesh Kanojia, Kevin Patel, Samarth Agrawal, Pushpak Bhattacharyya
ACL (1)2
2018 Indian Language Wordnets and their Linkages with Princeton WordNet
Diptesh Kanojia, Kevin Patel, Pushpak Bhattacharyya
LREC1
2018 Synthesizing Audio for Hindi WordNet
abstract
In this paper, we describe our work on the creation of a voice model using a speech synthesis system for the Hindi Language.We use preexisting "voices", use publicly available speech corpora to create a "voice" using the Festival Speech Synthesis System (Black, 1997).Our contribution is two-fold: (1) We scrutinize multiple speech synthesis systems and provide an extensive report on the currently available stateof-the-art systems.We also develop voices using the existing implementations of the aforementioned systems, and (2) We use these voices to generate sample audios for randomly chosen words; manually evaluate the audio generated, and produce audio for all WordNet words using the winner voice model.We also produce audios for the Hindi WordNet Glosses and Example sentences.We describe our efforts to use preexisting implementations for WaveNet -a model to generate raw audio using neural nets (Oord et al., 2016) and generate speech for Hindi.Our lexicographers perform a manual evaluation of the audio generated using multiple voices.A qualitative and quantitative analysis reveals that the voice model generated by us performs the best with an accuracy of 0.44.
Diptesh Kanojia, Preethi Jyothi, Pushpak Bhattacharyya
GWC1
2018 pyiwn: A Python based API to access Indian Language WordNets
abstract
Indian language WordNets have their individual web-based browsing interfaces along with a common interface for In-doWordNet.These interfaces prove to be useful for language learners and in an educational domain, however, they do not provide the functionality of connecting to them and browsing their data through a lucid application programming interface or an API.In this paper, we present our work on creating such an easy-to-use framework which is bundled with the data for Indian language WordNets and provides NLTK WordNet interface like core functionalities in Python.Additionally, we use a pre-built speech synthesis system for Hindi language and augment Hindi data with audios for words, glosses, and example sentences.We provide a detailed usage of our API and explain the functions for ease of the user.Also, we package the IndoWord-Net data along with the source code and provide it openly for the purpose of research.We aim to provide all our work as an open source framework for further development.
Ritesh Panjwani, Diptesh Kanojia, Pushpak Bhattacharyya
GWC2
2018 Semi-automatic WordNet Linking using Word Embeddings
abstract
Wordnets are rich lexico-semantic resources.Linked wordnets are extensions of wordnets, which link similar concepts in wordnets of different languages.Such resources are extremely useful in many Natural Language Processing (NLP) applications, primarily those based on knowledge-based approaches.In such approaches, these resources are considered as gold standard/oracle.Thus, it is crucial that these resources hold correct information.Thereby, they are created by human experts.However, manual maintenance of such resources is a tedious and costly affair.Thus techniques that can aid the experts are desirable.In this paper, we propose an approach to link wordnets.Given a synset of the source language, the approach returns a ranked list of potential candidate synsets in the target language from which the human expert can choose the correct one(s).Our technique is able to retrieve a winner synset in the top 10 ranked list for 60% of all synsets and 70% of noun synsets.
Kevin Patel, Diptesh Kanojia, Pushpak Bhattacharyya
GWC2
2018 Hindi Wordnet for Language Teaching: Experiences and Lessons Learnt
abstract
Hanumant Redkar, Rajita Shukla, Sandhya Singh, Jaya Saraswati, Laxmi Kashyap, Diptesh Kanojia, Preethi Jyothi, Malhar Kulkarni, Pushpak Bhattacharyya. Proceedings of the 9th Global Wordnet Conference. 2018.
Hanumant Harichandra Redkar, Rajita Shukla, Sandhya Singh, Jaya Saraswati, Laxmi Kashyap, Diptesh Kanojia, Preethi Jyothi, Malhar Kulkarni, Pushpak Bhattacharyya
GWC6
2017 Sarcasm Suite: A Browser-Based Engine for Sarcasm Detection and Generation
abstract
Sarcasm Suite is a browser-based engine that deploys five of our past papers in sarcasm detection and generation. The sarcasm detection modules use four kinds of incongruity: sentiment incongruity, semantic incongruity, historical context incongruity and conversational context incongruity. The sarcasm generation module is a chatbot that responds sarcastically to user input. With a visually appealing interface that indicates predictions using `faces' of our co-authors from our past papers, Sarcasm Suite is our first demonstration of our work in computational sarcasm.
Aditya Joshi 0001, Diptesh Kanojia, Pushpak Bhattacharyya, Mark J. Carman
AAAI2
2017 Scanpath Complexity: Modeling Reading Effort Using Gaze Information
abstract
Measuring reading effort is useful for practical purposes such as designing learning material and personalizing text comprehension environment. We propose a quantification of reading effort by measuring the complexity of eye-movement patterns of readers. We call the measure Scanpath Complexity. Scanpath complexity is modeled as a function of various properties of gaze fixations and saccades- the basic parameters of eye movement behavior. We demonstrate the effectiveness of our scanpath complexity measure by showing that its correlation with different measures of lexical and syntactic complexity as well as standard readability metrics is better than popular baseline measures based on fixation alone.
Abhijit Mishra, Diptesh Kanojia, Seema Nagar, Kuntal Dey, Pushpak Bhattacharyya
AAAI2
2016 Predicting Readers' Sarcasm Understandability by Modeling Gaze Behavior
abstract
Sarcasm understandability or the ability to understand textual sarcasm depends upon readers' language proficiency, social knowledge, mental state and attentiveness. We introduce a novel method to predict the sarcasm understandability of a reader. Presence of incongruity in textual sarcasm often elicits distinctive eye-movement behavior by human readers. By recording and analyzing the eye-gaze data, we show that eye-movement patterns vary when sarcasm is understood vis-à-vis when it is not. Motivated by our observations, we propose a system for sarcasm understandability prediction using supervised machine learning. Our system relies on readers' eye-movement parameters and a few textual features, thence, is able to predict sarcasm understandability with an F-score of 93%, which demonstrates its efficacy. The availability of inexpensive embedded-eye-trackers on mobile devices creates avenues for applying such research which benefits web-content creators, review writers and social media analysts alike.
Abhijit Mishra, Diptesh Kanojia, Pushpak Bhattacharyya
AAAI2
2016 Harnessing Cognitive Features for Sarcasm Detection
abstract
In this paper, we propose a novel mechanism for enriching the feature vector, for the task of sarcasm detection, with cognitive features extracted from eye-movement patterns of human readers.Sarcasm detection has been a challenging research problem, and its importance for NLP applications such as review summarization, dialog systems and sentiment analysis is well recognized.Sarcasm can often be traced to incongruity that becomes apparent as the full sentence unfolds.This presence of incongruity-implicit or explicit-affects the way readers eyes move through the text.We observe the difference in the behaviour of the eye, while reading sarcastic and non sarcastic sentences.Motivated by this observation, we augment traditional linguistic and stylistic features for sarcasm detection with the cognitive features obtained from readers eye movement data.We perform statistical classification using the enhanced feature set so obtained.The augmented cognitive features improve sarcasm detection by 3.7% (in terms of Fscore), over the performance of the best reported system.
Abhijit Mishra, Diptesh Kanojia, Seema Nagar, Kuntal Dey, Pushpak Bhattacharyya
ACL (1)2
2016 Leveraging Cognitive Features for Sentiment Analysis
abstract
Sentiments expressed in user-generated short text and sentences are nuanced by subtleties at lexical, syntactic, semantic and pragmatic levels. To address this, we propose to augment traditional features used for sentiment analysis and sarcasm detection, with cognitive features derived from the eye-movement patterns of readers. Statistical classification using our enhanced feature set improves the performance (F-score) of polarity detection by a maximum of 3.7% and 9.3% on two datasets, over the systems that use only traditional features. We perform feature significance analysis, and experiment on a held-out dataset, showing that cognitive features indeed empower sentiment analyzers to handle complex constructs.
Abhijit Mishra, Diptesh Kanojia, Seema Nagar, Kuntal Dey, Pushpak Bhattacharyya
CoNLL2
2016 SlangNet: A WordNet like resource for English Slang
Shehzaad Dhuliawala, Diptesh Kanojia, Pushpak Bhattacharyya
LREC2
2016 That'll Do Fine!: A Coarse Lexical Resource for English-Hindi MT, Using Polylingual Topic Models
Diptesh Kanojia, Aditya Joshi 0001, Pushpak Bhattacharyya, Mark J. Carman
LREC1
2016 Sophisticated Lexical Databases - Simplified Usage: Mobile Applications and Browser Plugins For Wordnets
abstract
India is a country with 22 officially recognized languages and 17 of these have WordNets, a crucial resource.Web browser based interfaces are available for these WordNets, but are not suited for mobile devices which deters people from effectively using this resource.We present our initial work on developing mobile applications and browser extensions to access WordNets for Indian Languages.Our contribution is two fold: (1) We develop mobile applications for the Android, iOS and Windows Phone OS platforms for Hindi, Marathi and Sanskrit WordNets which allow users to search for words and obtain more information along with their translations in English and other Indian languages.(2) We also develop browser extensions for English, Hindi, Marathi, and Sanskrit WordNets, for both Mozilla Firefox, and Google Chrome.We believe that such applications can be quite helpful in a classroom scenario, where students would be able to access the WordNets as dictionaries as well as lexical knowledge bases.This can help in overcoming the language barrier along with furthering language understanding.
Diptesh Kanojia, Raj Dabre, Pushpak Bhattacharyya
GWC1
2016 A picture is worth a thousand words: Using OpenClipArt library for enriching IndoWordNet
abstract
WordNet has proved to be immensely useful for Word Sense Disambiguation, and thence Machine translation, Information Retrieval and Question Answering.It can also be used as a dictionary for educational purposes.The semantic nature of concepts in a Word-Net motivates one to try to express this meaning in a more visual way.In this paper, we describe our work of enriching IndoWordNet with image acquisitions from the OpenClipArt library.We describe an approach used to enrich WordNets for eighteen Indian languages.Our contribution is three fold: (1) We develop a system, which, given a synset in English, finds an appropriate image for the synset.The system uses the OpenclipArt library (OCAL) to retrieve images and ranks them.(2) After retrieving the images, we map the results along with the linkages between Princeton WordNet and Hindi Word-Net, to link several synsets to corresponding images.We choose and sort top three images based on our ranking heuristic per synset.(3) We develop a tool that allows a lexicographer to manually evaluate these images.The top images are shown to a lexicographer by the evaluation tool for the task of choosing the best image representation.The lexicographer also selects the number of relevant images.Using our system, we obtain an Average Precision (P @ 3) score of 0.30.
Diptesh Kanojia, Shehzaad Dhuliawala, Pushpak Bhattacharyya
GWC1
2016 Mapping it differently: A solution to the linking challenges
abstract
This paper reports the work of creating bilingual mappings in English for certain synsets of Hindi wordnet, the need for doing this, the methods adopted and the tools created for the task.Hindi wordnet, which forms the foundation for other Indian language wordnets, has been linked to the English WordNet.To maximize linkages, an important strategy of using direct and hypernymy linkages has been followed.However, the hypernymy linkages were found to be inadequate in certain cases and posed a challenge due to sense granularity of language.Thus, the idea of creating bilingual mappings was adopted as a solution.A bilingual mapping means a linkage between a concept in two different languages, with the help of translation and/or transliteration.Such mappings retain meaningful representations, while capturing semantic similarity at the same time.This has also proven to be a great enhancement of Hindi wordnet and can be a crucial resource for multilingual applications in natural language processing, including machine translation and cross language information retrieval.
Meghna Singh, Rajita Shukla, Jaya Saraswati, Laxmi Kashyap, Diptesh Kanojia, Pushpak Bhattacharyya
GWC5
2015 World WordNet Database Structure: An Efficient Schema for Storing Information of WordNets of the World
abstract
WordNet is an online lexical resource which expresses unique concepts in a language. English WordNet is the first WordNet which was developed at Princeton University. Over a period of time, many language WordNets were developed by various organizations all over the world. It has always been a challenge to store the WordNet data. Some WordNets are stored using file system and some WordNets are stored using different database models. In this paper, we present the World WordNet Database Structure which can be used to efficiently store the WordNet information of all languages of the World. This design can be adapted by most language WordNets to store information such as synset data, semantic and lexical relations, ontology details, language specific features, linguistic information, etc. An attempt is made to develop Application Programming Interfaces to manipulate the data from these databases. This database structure can help in various Natural Language Processing applications like Multilingual Information Retrieval, Word Sense Disambiguation, Machine Translation, etc.
Hanumant Harichandra Redkar, Sudha Bhingardive, Diptesh Kanojia, Pushpak Bhattacharyya
AAAI3
2014 Do not do processing, when you can look up: Towards a Discrimination Net for WSD
abstract
The task of Word Sense Disambiguation (WSD) incorporates in its definition the role of 'context'.We present our work on the development of a tool which allows for automatic acquisition and ranking of 'context clues' for WSD.These clue words are extracted from the contexts of words appearing in a large monolingual corpus.These mined collection of contextual clues form a discrimination net in the sense that for targeted WSD, navigation of the net leads to the correct sense of a word given its context.Utilizing this resource we intend to develop efficient and light weight WSD based on look up and navigation of memoryresident knowledge base, thereby avoiding heavy computation which often prevents incorporation of any serious WSD in MT and search.The need for large quantities of sense marked data too can be reduced.
Diptesh Kanojia, Pushpak Bhattacharyya, Raj Dabre, Siddhartha Gunti, Manish Shrivastava 0001
GWC1
2013 More than meets the eye: Study of Human Cognition in Sense Annotation
Salil Joshi 0001, Diptesh Kanojia, Pushpak Bhattacharyya
HLT-NAACL2