Pushpak Bhattacharyya

dblp:p/PushpakBhattacharyya · also Pushpak Bhattacharya · DBLP profile ↗
← Back
41ranked-venue papers in the field
1as first author
13since 2021 · last 2024
0000-0001-5319-5508ORCID · verified

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 21 (1 first)Information Retrieval & Web Search · 18Data Mining & Knowledge Discovery · 2
YearPublicationVenuePosition
2024 Yes, This Is What I Was Looking For! Towards Multi-modal Medical Consultation Concern Summary Generation
Abhisek Tiwari, Shreyangshu Bera, Sriparna Saha 0001, Pushpak Bhattacharyya, Samrat Ghosh
ECIR (3)4
2024 Material Microstructure Design Using VAE-Regression with a Multimodal Prior
Avadhut Sardeshmukh, Sreedhar Reddy, Gautham B. P., Pushpak Bhattacharyya
PAKDD (6)4
2023 Multi-step Prompting for Few-shot Emotion-Grounded Conversations
abstract
Conversational systems have shown immense growth in their ability to communicate like humans. With the emergence of large pre-trained language models (PLMs) the ability to provide informative responses have improved significantly. Despite the success of PLMs, the ability to identify and generate engaging and empathetic responses is largely dependent on labelled-data. In this work, we design a prompting approach that identifies the emotion of a given utterance and uses the emotion information for generating the appropriate responses for conversational systems. We propose a two-step prompting method that first recognises the emotion in the dialogue utterance and in the second-step uses the predicted emotion to prompt the PLM to generate the corresponding em- pathetic response in a few-shot setting. Experimental results on three publicly available datasets show that our proposed approach outperforms the state-of-the-art approaches for both automatic and manual evaluation.
Mauajama Firdaus, Gopendra Vikram Singh, Asif Ekbal, Pushpak Bhattacharyya
CIKM4
2023 Experience and Evidence are the eyes of an excellent summarizer! Towards Knowledge Infused Multi-modal Clinical Conversation Summarization
abstract
With the advancement of telemedicine, both researchers and medical practitioners are working hand-in-hand to develop various techniques to automate various medical operations, such as diagnosis report generation. In this paper, we first present a multi-modal clinical conversation summary generation task that takes a clinician-patient interaction (both textual and visual information) and generates a succinct synopsis of the conversation. We propose a knowledge-infused, multi-modal, multi-tasking medical domain identification and clinical conversation summary generation (MM-CliConSummation) framework. It leverages an adapter to infuse knowledge and visual features and unify the fused feature vector using a gated mechanism. Furthermore, we developed a multi-modal, multi-intent clinical conversation summarization corpus annotated with intent, symptom, and summary. The extensive set of experiments, both quantitatively and qualitatively, led to the following findings: (a) critical significance of visuals, (b) more precise and medical entity preserving summary with additional knowledge infusion, and (c) a correlation between medical department identification and clinical synopsis generation. Furthermore, the dataset and source code are available at https://github.com/NLP-RL/MM-CliConSummation
Abhisek Tiwari, Anisha Saha, Sriparna Saha 0001, Pushpak Bhattacharyya, Minakshi Dhar
CIKM4
2023 DeCoDE: Detection of Cognitive Distortion and Emotion Cause Extraction in Clinical Conversations
Gopendra Vikram Singh, Soumitra Ghosh, Asif Ekbal, Pushpak Bhattacharyya
ECIR (2)4
2023 "Explain Thyself Bully": Sentiment Aided Cyberbullying Detection with Explanation
Krishanu Maity, Prince Jha, Raghav Jain, Sriparna Saha 0001, Pushpak Bhattacharyya
ICDAR (3)5
2023 VAD-assisted multitask transformer framework for emotion recognition and intensity prediction on suicide notes
Soumitra Ghosh, Asif Ekbal, Pushpak Bhattacharyya
Inf. Process. Manag.3
2022 Dr. Can See: Towards a Multi-modal Disease Diagnosis Virtual Assistant
abstract
Artificial Intelligence-based clinical decision support is gaining ever-growing popularity and demand in both the research and industry communities. One such manifestation is automatic disease diagnosis, which aims to assist clinicians in conducting symptom investigations and disease diagnoses. When we consult with doctors, we often report and describe our health conditions with visual aids. Moreover, many people are unacquainted with several symptoms and medical terms, such as mouth ulcer and skin growth. Therefore, visual form of symptom reporting is a necessity. Motivated by the efficacy of visual form of symptom reporting, we propose and build a novel end-to-end Multi-modal Disease Diagnosis Virtual Assistant (MDD-VA) using reinforcement learning technique. In conversation, users' responses are heavily influenced by the ongoing dialogue context, and multi-modal responses appear to be of no difference. We also propose and incorporate a Context-aware Symptom Image Identification module that leverages discourse context in addition to the symptom image for identifying symptoms effectively. Furthermore, we first curate a multi-modal conversational medical dialogue corpus in English that is annotated with intent, symptoms, and visual information. The proposed MDD-VA outperforms multiple uni-modal baselines in both automatic and human evaluation, which firmly establishes the critical role of symptom information provided by visuals . The dataset and code are available at https://github.com/NLP-RL/DrCanSee
Abhisek Tiwari, Manisimha Manthena, Sriparna Saha 0001, Pushpak Bhattacharyya, Minakshi Dhar, Sarbajeet Tiwari
CIKM4
2022 CARES: CAuse Recognition for Emotion in Suicide Notes
Soumitra Ghosh, Swarup Roy, Asif Ekbal, Pushpak Bhattacharyya
ECIR (2)4
2022 An Efficient Fusion Mechanism for Multimodal Low-resource Setting
abstract
The effective fusion of multiple modalities (i.e., text, acoustic, and visual) is a non-trivial task, as these modalities often carry specific and diverse information and do not contribute equally. The fusion of different modalities could even be more challenging under the low-resource setting, where we have fewer samples for training. This paper proposes a multi-representative fusion mechanism that generates diverse fusions with multiple modalities and then chooses the best fusion among them. To achieve this, we first apply convolution filters on multimodal inputs to generate different and diverse representations of modalities. We then fuse pairwise modalities with multiple representations to get the multiple fusions. Finally, we propose an attention mechanism that only selects the most appropriate fusion, which eventually helps resolve the noise problem by ignoring the noisy fusions. We evaluate our proposed approach on three low-resource multimodal sentiment analysis datasets, i.e., YouTube, MOUD, and ICT-MMMO. Experimental results show the effectiveness of our proposed approach with the accuracies of 59.3%, 83.0%, and 84.1% for the YouTube, MOUD, and ICT-MMMO datasets, respectively.
Dushyant Singh Chauhan, Asif Ekbal, Pushpak Bhattacharyya
SIGIR3
2022 A Multitask Framework for Sentiment, Emotion and Sarcasm aware Cyberbullying Detection from Multi-modal Code-Mixed Memes
abstract
Detecting cyberbullying from memes is highly challenging, because of the presence of the implicit affective content which is also often sarcastic, and multi-modality (image + text). The current work is the first attempt, to the best of our knowledge, in investigating the role of sentiment, emotion and sarcasm in identifying cyberbullying from multi-modal memes in a code-mixed language setting. As a contribution, we have created a benchmark multi-modal meme dataset called MultiBully annotated with bully, sentiment, emotion and sarcasm labels collected from open-source Twitter and Reddit platforms. Moreover, the severity of the cyberbullying posts is also investigated by adding a harmfulness score to each of the memes. The created dataset consists of two modalities, text and image. Most of the texts in our dataset are in code-mixed form, which captures the seamless transitions between languages for multilingual users. Two different multimodal multitask frameworks (BERT+ResNET-Feedback and CLIP-CentralNet) have been proposed for cyberbullying detection (CD), the three auxiliary tasks being sentiment analysis (SA), emotion recognition (ER) and sarcasm detection (SAR). Experimental results indicate that compared to uni-modal and single-task variants, the proposed frameworks improve the performance of the main task, i.e., CD, by 3.18% and 3.10% in terms of accuracy and F1 score, respectively.
Krishanu Maity, Prince Jha, Sriparna Saha 0001, Pushpak Bhattacharyya
SIGIR4
2021 Let's Summarize Scientific Documents! A Clustering-Based Approach via Citation Context
Santosh Kumar Mishra, Naveen Saini, Sriparna Saha 0001, Pushpak Bhattacharyya
NLDB4
2021 Authorship Attribution Using Capsule-Based Fusion Approach
Chanchal Suman, Sriparna Saha 0001, Pushpak Bhattacharyya
NLDB4
2020 Natural Language Generation Using Transformer Network in an Open-Domain Setting
Deeksha Varshney, Asif Ekbal, Ganesh Prasad Nagaraja, Mrigank Tiwari, Abhijith Athreya Mysore Gopinath, Pushpak Bhattacharyya
NLDB6
2019 A Novel Approach Towards Fake News Detection: Deep Learning Augmented with Textual Entailment Features
Tanik Saikh, Asif Ekbal, Pushpak Bhattacharyya
NLDB4
2019 Utilizing Wordnets for Cognate Detection among Indian Languages
abstract
Automatic Cognate Detection (ACD) is a challenging task which has been utilized to help NLP applications like Machine Translation, Information Retrieval and Computational Phylogenetics.Unidentified cognate pairs can pose a challenge to these applications and result in a degradation of performance.In this paper, we detect cognate word pairs among ten Indian languages with Hindi and use deep learning methodologies to predict whether a word pair is cognate or not.We identify IndoWordnet as a potential resource to detect cognate word pairs based on orthographic similarity-based methods and train neural network models using the data obtained from it.We identify parallel corpora as another potential resource and perform the same experiments for them.We also validate the contribution of Wordnets through further experimentation and report improved performance of up to 26%.We discuss the nuances of cognate detection among closely related Indian languages and release the lists of detected cognates as a dataset.We also observe the behaviour of, to an extent, unrelated Indian language pairs and release the lists of detected cognates among them as well.
Diptesh Kanojia, Kevin Patel, Malhar Kulkarni, Pushpak Bhattacharyya, Gholamreza Haffari
GWC4
2018 Synthesizing Audio for Hindi WordNet
abstract
In this paper, we describe our work on the creation of a voice model using a speech synthesis system for the Hindi Language.We use preexisting "voices", use publicly available speech corpora to create a "voice" using the Festival Speech Synthesis System (Black, 1997).Our contribution is two-fold: (1) We scrutinize multiple speech synthesis systems and provide an extensive report on the currently available stateof-the-art systems.We also develop voices using the existing implementations of the aforementioned systems, and (2) We use these voices to generate sample audios for randomly chosen words; manually evaluate the audio generated, and produce audio for all WordNet words using the winner voice model.We also produce audios for the Hindi WordNet Glosses and Example sentences.We describe our efforts to use preexisting implementations for WaveNet -a model to generate raw audio using neural nets (Oord et al., 2016) and generate speech for Hindi.Our lexicographers perform a manual evaluation of the audio generated using multiple voices.A qualitative and quantitative analysis reveals that the voice model generated by us performs the best with an accuracy of 0.44.
Diptesh Kanojia, Preethi Jyothi, Pushpak Bhattacharyya
GWC3
2018 pyiwn: A Python based API to access Indian Language WordNets
abstract
Indian language WordNets have their individual web-based browsing interfaces along with a common interface for In-doWordNet.These interfaces prove to be useful for language learners and in an educational domain, however, they do not provide the functionality of connecting to them and browsing their data through a lucid application programming interface or an API.In this paper, we present our work on creating such an easy-to-use framework which is bundled with the data for Indian language WordNets and provides NLTK WordNet interface like core functionalities in Python.Additionally, we use a pre-built speech synthesis system for Hindi language and augment Hindi data with audios for words, glosses, and example sentences.We provide a detailed usage of our API and explain the functions for ease of the user.Also, we package the IndoWord-Net data along with the source code and provide it openly for the purpose of research.We aim to provide all our work as an open source framework for further development.
Ritesh Panjwani, Diptesh Kanojia, Pushpak Bhattacharyya
GWC3
2018 An Iterative Approach for Unsupervised Most Frequent Sense Detection using WordNet and Word Embeddings
abstract
Given a word, what is the most frequent sense in which it occurs in a given corpus?Most Frequent Sense (MFS) is a strong baseline for unsupervised word sense disambiguation.If we have large amounts of sense-annotated corpora, MFS can be trivially created.However, senseannotated corpora are a rarity.In this paper, we propose a method which can compute MFS from raw corpora.Our approach iteratively exploits the semantic congruity among related words in corpus.Our method performs better compared to another similar work.
Kevin Patel, Pushpak Bhattacharyya
GWC2
2018 Semi-automatic WordNet Linking using Word Embeddings
abstract
Wordnets are rich lexico-semantic resources.Linked wordnets are extensions of wordnets, which link similar concepts in wordnets of different languages.Such resources are extremely useful in many Natural Language Processing (NLP) applications, primarily those based on knowledge-based approaches.In such approaches, these resources are considered as gold standard/oracle.Thus, it is crucial that these resources hold correct information.Thereby, they are created by human experts.However, manual maintenance of such resources is a tedious and costly affair.Thus techniques that can aid the experts are desirable.In this paper, we propose an approach to link wordnets.Given a synset of the source language, the approach returns a ranked list of potential candidate synsets in the target language from which the human expert can choose the correct one(s).Our technique is able to retrieve a winner synset in the top 10 ranked list for 60% of all synsets and 70% of noun synsets.
Kevin Patel, Diptesh Kanojia, Pushpak Bhattacharyya
GWC3
2018 Hindi Wordnet for Language Teaching: Experiences and Lessons Learnt
abstract
Hanumant Redkar, Rajita Shukla, Sandhya Singh, Jaya Saraswati, Laxmi Kashyap, Diptesh Kanojia, Preethi Jyothi, Malhar Kulkarni, Pushpak Bhattacharyya. Proceedings of the 9th Global Wordnet Conference. 2018.
Hanumant Harichandra Redkar, Rajita Shukla, Sandhya Singh, Jaya Saraswati, Laxmi Kashyap, Diptesh Kanojia, Preethi Jyothi, Malhar Kulkarni, Pushpak Bhattacharyya
GWC9
2016 Detecting Most Frequent Sense using Word Embeddings and BabelNet
abstract
Since the inception of the SENSEVAL evaluation exercises there has been a great deal of recent research into Word Sense Disambiguation (WSD).Over the years, various supervised, unsupervised and knowledge based WSD systems have been proposed.Beating the first sense heuristics is a challenging task for these systems.In this paper, we present our work on Most Frequent Sense (MFS) detection using Word Embeddings and BabelNet features.The semantic features from BabelNet viz., synsets, gloss, relations, etc. are used for generating sense embeddings.We compare word embedding of a word with its sense embeddings to obtain the MFS with the highest similarity.The MFS is detected for six languages viz., English, Spanish, Russian, German, French and Italian.However, this approach can be applied to any language provided that word embeddings are available for that language.
Harpreet Singh Arora, Sudha Bhingardive, Pushpak Bhattacharyya
GWC3
2016 IndoWordNet: : Similarity- Computing Semantic Similarity and Relatedness using IndoWordNet
abstract
Semantic similarity and relatedness measures play an important role in natural language processing applications.In this paper, we present the IndoWordNet::Similarity tool and interface, designed for computing the semantic similarity and relatedness between two words in IndoWordNet.A java based tool and a web interface have been developed to compute this semantic similarity and relatedness.Also, Java API has been developed for this purpose.This tool, web interface and the API are made available for the research purpose.
Sudha Bhingardive, Hanumant Harichandra Redkar, Prateek Sappadla, Dhirendra Singh, Pushpak Bhattacharyya
GWC5
2016 Sophisticated Lexical Databases - Simplified Usage: Mobile Applications and Browser Plugins For Wordnets
abstract
India is a country with 22 officially recognized languages and 17 of these have WordNets, a crucial resource.Web browser based interfaces are available for these WordNets, but are not suited for mobile devices which deters people from effectively using this resource.We present our initial work on developing mobile applications and browser extensions to access WordNets for Indian Languages.Our contribution is two fold: (1) We develop mobile applications for the Android, iOS and Windows Phone OS platforms for Hindi, Marathi and Sanskrit WordNets which allow users to search for words and obtain more information along with their translations in English and other Indian languages.(2) We also develop browser extensions for English, Hindi, Marathi, and Sanskrit WordNets, for both Mozilla Firefox, and Google Chrome.We believe that such applications can be quite helpful in a classroom scenario, where students would be able to access the WordNets as dictionaries as well as lexical knowledge bases.This can help in overcoming the language barrier along with furthering language understanding.
Diptesh Kanojia, Raj Dabre, Pushpak Bhattacharyya
GWC3
2016 A picture is worth a thousand words: Using OpenClipArt library for enriching IndoWordNet
abstract
WordNet has proved to be immensely useful for Word Sense Disambiguation, and thence Machine translation, Information Retrieval and Question Answering.It can also be used as a dictionary for educational purposes.The semantic nature of concepts in a Word-Net motivates one to try to express this meaning in a more visual way.In this paper, we describe our work of enriching IndoWordNet with image acquisitions from the OpenClipArt library.We describe an approach used to enrich WordNets for eighteen Indian languages.Our contribution is three fold: (1) We develop a system, which, given a synset in English, finds an appropriate image for the synset.The system uses the OpenclipArt library (OCAL) to retrieve images and ranks them.(2) After retrieving the images, we map the results along with the linkages between Princeton WordNet and Hindi Word-Net, to link several synsets to corresponding images.We choose and sort top three images based on our ranking heuristic per synset.(3) We develop a tool that allows a lexicographer to manually evaluate these images.The top images are shown to a lexicographer by the evaluation tool for the task of choosing the best image representation.The lexicographer also selects the number of relevant images.Using our system, we obtain an Average Precision (P @ 3) score of 0.30.
Diptesh Kanojia, Shehzaad Dhuliawala, Pushpak Bhattacharyya
GWC3
2016 IndoWordNet Conversion to Web Ontology Language (OWL)
abstract
WordNet plays a significant role in Linked Open Data (LOD) cloud.It has numerous application ranging from ontology annotation to ontology mapping.IndoWord-Net is a linked WordNet connecting 18 Indian language WordNets with Hindi as a source WordNet.The Hindi WordNet was initially developed by linking it to English WordNet.In this paper, we present a data representation of IndoWordNet in Web Ontology Language (OWL).The schema of Princeton WordNet has been enhanced to support the representation of IndoWordNet.This IndoWordNet representation in OWL format is now available to link other web resources.This representation is implemented for eight Indian languages.
Apurva Nagvenkar, Jyoti D. Pawar, Pushpak Bhattacharyya
GWC3
2016 Samāsa-Kartā: An Online Tool for Producing Compound Words using IndoWordNet
abstract
Samāsa or compounds are a regular feature of Indian Languages.They are also found in other languages like German, Italian, French, Russian, Spanish, etc. Compound word is constructed from two or more words to form a single word.The meaning of this word is derived from each of the individual words of the compound.To develop a system to generate, identify and interpret compounds, is an important task in Natural Language Processing.This paper introduces a web based tool -Samāsa-Kartā for producing compound words.Here, the focus is on Sanskrit language due to its richness in usage of compounds; however, this approach can be applied to any Indian language as well as other languages.IndoWordNet is used as a resource for words to be compounded.The motivation behind creating compound words is to create, to improve the vocabulary, to reduce sense ambiguity, etc. in order to enrich the WordNet.The Samāsa-Kartā can be used for various applications viz., compound categorization, sandhi creation, morphological analysis, paraphrasing, synset creation, etc.
Hanumant Harichandra Redkar, Nilesh Joshi, Sandhya Singh, Irawati Kulkarni, Malhar Kulkarni, Pushpak Bhattacharyya
GWC6
2016 High, Medium or Low? Detecting Intensity Variation Among polar synonyms in WordNet
abstract
For fine-grained sentiment analysis, we need to go beyond zero-one polarity and find a way to compare adjectives (synonyms) that share the same sense.Choice of a word from a set of synonyms, provides a way to select the exact polarityintensity.For example, choosing to describe a person as benevolent rather than kind 1 changes the intensity of the expression.In this paper, we present a sense based lexical resource, where synonyms are assigned intensity levels, viz., high, medium and low.We show that the measure P (s|w) (probability of a sense s given the word w) can derive the intensity of a word within the sense.We observe a statistically significant positive correlation between P (s|w) and intensity of synonyms for three languages, viz., English, Marathi and Hindi.The average correlation scores are 0.47 for English, 0.56 for Marathi and 0.58 for Hindi.
Raksha Sharma, Pushpak Bhattacharyya
GWC2
2016 Detection of Compound Nouns and Light Verb Constructions using IndoWordNet
abstract
Detection of MultiWord Expressions (MWEs) is one of the fundamental problems in Natural Language Processing.In this paper, we focus on two categories of MWEs -Compound Nouns and Light Verb Constructions.These two categories can be tackled using knowledge bases, rather than pure statistics.We investigate usability of IndoWordNet for the detection of MWEs.Our IndoWordNet based approach uses semantic and ontological features of words that can be extracted from IndoWordNet.This approach has been tested on Indian languages viz., Assamese, Bengali, Hindi, Konkani, Marathi, Odia and Punjabi.Results show that ontological features are found to be very useful for the detection of light verb constructions, while use of semantic properties for the detection of compound nouns is found to be satisfactory.This approach can be easily adapted by other Indian languages.Detected MWEs can be interpolated into WordNets as they help in representing semantic knowledge.
Dhirendra Singh, Sudha Bhingardive, Pushpak Bhattacharyya
GWC3
2016 Mapping it differently: A solution to the linking challenges
abstract
This paper reports the work of creating bilingual mappings in English for certain synsets of Hindi wordnet, the need for doing this, the methods adopted and the tools created for the task.Hindi wordnet, which forms the foundation for other Indian language wordnets, has been linked to the English WordNet.To maximize linkages, an important strategy of using direct and hypernymy linkages has been followed.However, the hypernymy linkages were found to be inadequate in certain cases and posed a challenge due to sense granularity of language.Thus, the idea of creating bilingual mappings was adopted as a solution.A bilingual mapping means a linkage between a concept in two different languages, with the help of translation and/or transliteration.Such mappings retain meaningful representations, while capturing semantic similarity at the same time.This has also proven to be a great enhancement of Hindi wordnet and can be a crucial resource for multilingual applications in natural language processing, including machine translation and cross language information retrieval.
Meghna Singh, Rajita Shukla, Jaya Saraswati, Laxmi Kashyap, Diptesh Kanojia, Pushpak Bhattacharyya
GWC6
2014 Graph Based Algorithm for Automatic Domain Segmentation of WordNet
abstract
We present a graph based algorithm for automatic domain segmentation of Wordnet.We pose the problem as a Markov Random Field Classification problem and show how existing graph based algorithms for Image Processing can be used to solve the problem.Our approach is unsupervised and can be easily adopted for any language.We conduct our experiments for two domains, health and tourism.We achieve F-Score more than .70 in both domains.This work can be useful for many critical problems like word sense disambiguation, domain specific ontology extraction etc.
Brijesh Bhatt, Subhash Kunnath, Pushpak Bhattacharyya
GWC3
2014 Facilitating Multi-Lingual Sense Annotation: Human Mediated Lemmatizer
abstract
Sense marked corpora is essential for supervised word sense disambiguation (WSD).The marked sense ids come from wordnets.However, words in corpora appear in morphed forms, while wordnets store lemma.This situation calls for accurate lemmatizers.The lemma is the gateway to the wordnet.However, the problem is that for many languages, lemmatizers do not exist, and this problem is not easy to solve, since rule based lemmatizers take time and require highly skilled linguists.Satistical stemmers on the other hand do not return legitimate lemma.
Pushpak Bhattacharyya, Ankit Bahuguna, Lavita Talukdar, Bornali Phukan
GWC1
2014 Semi-Automatic Extension of Sanskrit Wordnet using Bilingual Dictionary
abstract
In this paper, we report our methods and results of using, for the first time, semi-automatic approach to enhance an Indian language Wordnet.We apply our methods to enhancing an already existing Sanskrit Wordnet created from Hindi Wordnet (which is created from Princeton Wordnet) using expansion approach.We base our experiment on an existing bilingual Sanskrit English Dictionary and show how lemma in this dictionary can be mapped to Princeton Wordnet through which corresponding Sanskrit synsets can be populated by Sanskrit lexemes.This our method will also show how absence of resources of a pair of languages need not be an obstacle, if another resource of one of them is available.Sanskrit being historically related to languages of Indo-European family, we believe that this semi-automatic approach will help enhance Wordnets of other Indian languages of the same family.
Sudha Bhingardive, Tanuja Ajotikar, Irawati Kulkarni, Malhar Kulkarni, Pushpak Bhattacharyya
GWC5
2014 IndoWordnet Visualizer: A Graphical User Interface for Browsing and Exploring Wordnets of Indian Languages
abstract
In this paper, we are presenting a graphical user interface to browse and explore the In-doWordnet lexical database for various Indian languages.IndoWordnet visualizer extracts the related concepts for a given word and displays a sub graph containing those concepts.The interface is enhanced with different features in order to provide flexibility to the user.IndoWordnet visualizer is made publically available.Though it was initially constructed for making the wordnet validation process easier, it is proving to be very useful in analyzing various Natural Language Processing tasks, viz.,
Devendra Singh Chaplot, Sudha Bhingardive, Pushpak Bhattacharyya
GWC3
2014 Do not do processing, when you can look up: Towards a Discrimination Net for WSD
abstract
The task of Word Sense Disambiguation (WSD) incorporates in its definition the role of 'context'.We present our work on the development of a tool which allows for automatic acquisition and ranking of 'context clues' for WSD.These clue words are extracted from the contexts of words appearing in a large monolingual corpus.These mined collection of contextual clues form a discrimination net in the sense that for targeted WSD, navigation of the net leads to the correct sense of a word given its context.Utilizing this resource we intend to develop efficient and light weight WSD based on look up and navigation of memoryresident knowledge base, thereby avoiding heavy computation which often prevents incorporation of any serious WSD in MT and search.The need for large quantities of sense marked data too can be reduced.
Diptesh Kanojia, Pushpak Bhattacharyya, Raj Dabre, Siddhartha Gunti, Manish Shrivastava 0001
GWC2
2012 TwiSent: a multistage system for analyzing sentiment in twitter
abstract
In this paper, we present TwiSent, a sentiment analysis system for Twitter. Based on the topic searched, TwiSent collects tweets pertaining to it and categorizes them into the different polarity classes positive, negative and objective. However, analyzing micro-blog posts have many inherent challenges compared to the other text genres. Through TwiSent, we address the problems of 1) Spams pertaining to sentiment analysis in Twitter, 2) Structural anomalies in the text in the form of incorrect spellings, nonstandard abbreviations, slangs etc., 3) Entity specificity in the context of the topic searched and 4) Pragmatics embedded in text. The system performance is evaluated on manually annotated gold standard data and on an automatically annotated tweet set based on hashtags. It is a common practise to show the efficacy of a supervised system on an automatically annotated dataset. However, we show that such a system achieves lesser classification accurcy when tested on generic twitter dataset. We also show that our system performs much better than an existing system.
Subhabrata Mukherjee, Akshat Malu, A. R. Balamurali, Pushpak Bhattacharyya
CIKM4
2012 WikiSent: Weakly Supervised Sentiment Analysis through Extractive Summarization with Wikipedia
abstract
This paper describes a weakly supervised system for sentiment analysis in the movie review domain. The objective is to classify a movie review into a polarity class, positive or negative, based on those sentences bearing opinion on the movie alone, leaving out other irrelevant text. Wikipedia incorporates the world knowledge of movie-specific features in the system which is used to obtain an extractive summary of the review, consisting of the reviewer’s opinions about the specific aspects of the movie. This filters out the concepts which are irrelevant or objective with respect to the given movie. The proposed system, WikiSent, does not require any labeled data for training. It achieves a better or comparable accuracy to the existing semi-supervised and unsupervised systems in the domain, on the same dataset. We also perform a general movie review trend analysis using WikiSent.
Subhabrata Mukherjee, Pushpak Bhattacharyya
ECML/PKDD (1)2
2010 On Improving Pseudo-Relevance Feedback Using Pseudo-Irrelevant Documents
Karthik Raman 0001, Raghavendra Udupa, Pushpak Bhattacharyya, Abhijit Bhole
ECIR3
2010 Multilingual PRF: english lends a helping hand
abstract
In this paper, we present a novel approach to Pseudo-Relevance Feedback (PRF) called Multilingual PRF (MultiPRF). The key idea is to harness multilinguality. Given a query in a language, we take the help of another language to ameliorate the well known problems of PRF, viz. (a) The expansion terms from PRF are primarily based on co-occurrence relationships with query terms, and thus other terms which are lexically and semantically related, such as morphological variants and synonyms, are not explicitly captured, and (b) PRF is quite sensitive to the quality of the initially retrieved top k documents and is thus not robust. In MultiPRF, given a query in language L1, it is translated into language L2 and PRF is performed on a collection in language L2 and the resultant feedback model is translated from L2 back into L1. The final feedback model is obtained by combining the translated model with the original feedback model of the query in L1.
Manoj Kumar Chinnakotla, Karthik Raman 0001, Pushpak Bhattacharyya
SIGIR3
2004 Is question answering an acquired skill?
abstract
We present a question answering (QA) system which learns how to detect and rank answer passages by analyzing questions and their answers (QA pairs) provided as training data. We built our system in only a few person-months using o#- the-shelf components: a part-of-speech tagger, a shallow parser, a lexical network, and a few well-known supervised learning algorithms. In contrast, many of the top TREC QA systems are large group e#orts, using customized ontologies, question classifiers, and highly tuned ranking functions. Our ease of deployment arises from using generic, trainable algorithms that exploit simple feature extractors on QA pairs. With TREC QA data, our system achieves mean reciprocal rank (MRR) that compares favorably with the best scores in recent years, and generalizes from one corpus to another. Our key technique is to recover, from the question, fragments of what might have been posed as a structured query, had a suitable schema been available. One fragment comprises selectors: tokens that are likely to appear (almost) unchanged in an answer passage. The other fragment contains question tokens which give clues about the answer type, and are expected to be replaced in the answer passage by tokens which specialize or instantiate the desired answer type. Selectors are like constants in where-clauses in relational queries, and answer types are like column names. We present new algorithms for locating selectors and answer type clues and using them in scoring passages with respect to a question.
Ganesh Ramakrishnan, Soumen Chakrabarti, Deepa Paranjpe, Pushpak Bhattacharyya
WWW4
2003 Text Representation with WordNet Synsets using Soft Sense Disambiguation
Ganesh Ramakrishnan, Pushpak Bhattacharyya
NLDB2