EDBT 2026 Demo / reviewers in the wild / expert
Kripabandhu Ghosh
dblp:74/10289
· DBLP profile ↗
36ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-8130-1221ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 20 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 18 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLMs in Sarcasm Detection? It's elementary! (Or is it?)abstractWhile Large Language Models (LLMs) are frequently cited for their sophisticated pragmatic reasoning (CITATION), recent progress in sarcasm detection increasingly relies on synthetic benchmarks (CITATION). This study exposes a catastrophic generalization gap in this paradigm: we observe that models achieve near-perfect accuracy on synthetic data but collapse to random guessing on organic human speech. By triangulating hidden state geometry, entropy analysis, and causal interventions, we demonstrate that this disparity stems from shortcut learning (CITATION)—models exploit the low-entropy statistical signatures of generated text while remaining “semantically blind” to the pragmatic cues essential for irony. Our findings indicate that high performance on synthetic leaderboards reflects forensic pattern matching rather than the genuine linguistic intelligence assumed in prior work, creating a statistical mirage of competence. Priyanshu Mahato, Aniket Santosh Mishra, Kripabandhu Ghosh |
ACL (1) | 3 |
| 2026 | ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question AnsweringabstractThe rapid proliferation of Large Language Models (LLMs) has significantly contributed to the development of equitable AI systems capable of factual question-answering (QA). However, no known study tests the LLMs' robustness when presented with obfuscated versions of questions. To systematically evaluate these limitations, we propose a novel technique, ObfusQAte, and leveraging the same, introduce ObfusQA, a comprehensive, first-of-its-kind framework with multi-tiered obfuscation levels designed to examine LLM capabilities across three distinct dimensions: (i) Named-Entity Indirection, (ii) Distractor Indirection, and (iii) Contextual Overload. By capturing these fine-grained distinctions in language, ObfusQA provides a comprehensive benchmark for evaluating LLM robustness and adaptability. Our study observes that LLMs exhibit a tendency to fail or generate hallucinated responses when confronted with these increasingly nuanced variations. To foster research in this direction, we make ObfusQAte publicly available. Shubhra Ghosh, Abhilekh Borah, Aditya Kumar Guru, Kripabandhu Ghosh |
LREC | 4 |
| 2026 | Structured Legal Document Generation in India: A Model-Agnostic Wrapper Approach with VidhikDastaavejabstractAutomating legal document drafting can improve efficiency and reduce the burden of manual legal work. Yet, the structured generation of private legal documents remains underexplored, particularly in the Indian context, due to the scarcity of public datasets and the complexity of adapting models for long-form legal drafting. To address this gap, we introduce VidhikDastaavej, a large-scale, anonymized dataset of private legal documents curated in collaboration with an Indian law firm. Covering 133 diverse categories, this dataset is the first resource of its kind and provides a foundation for research in structured legal text generation and Legal AI more broadly. We further propose a Model-Agnostic Wrapper (MAW), a two-stage generation framework that first plans the section structure of a legal draft and then generates each section with retrieval-based prompts. MAW is independent of any specific LLM, making it adaptable across both open- and closed-source models. Comprehensive evaluation, including lexical, semantic, LLM-based, and expert-driven assessments with inter-annotator agreement, shows that the wrapper substantially improves factual accuracy, coherence, and completeness compared to fine-tuned baselines. This work establishes both a new benchmark dataset and a generalizable generation framework, paving the way for future research in AI-assisted legal drafting. Shubham Kumar Nigam, Balaramamahanthi Deepak Patnaik, Noel Shallum, Kripabandhu Ghosh, Arnab Bhattacharya 0001 |
LREC | 4 |
| 2025 | Justice for the Disadvantaged: A Study of Public Reactions on Indian Supreme Court Judgments
Soumilya De, Soumyajit Datta, Koustav Rudra, Saptarshi Ghosh 0001, Ashiqur KhudaBuksh, Kripabandhu Ghosh |
ASONAM (2) | 6 |
| 2025 | NYAYAANUMANA and INLEGALLLAMA: The Largest Indian Legal Judgment Prediction Dataset and Specialized Language Model for Enhanced Decision AnalysisabstractThe integration of artificial intelligence (AI) in legal judgment prediction (LJP) has the potential to transform the legal landscape, particularly in jurisdictions like India, where a significant backlog of cases burdens the legal system. This paper introduces NyayaAnumana, the largest and most diverse corpus of Indian legal cases compiled for LJP, encompassing a total of 7,02,945 preprocessed cases. NyayaAnumana, which combines the words “Nyaya” and “Anumana” that means “judgment” and “inference” respectively for most major Indian languages, includes a wide range of cases from the Supreme Court, High Courts, Tribunal Courts, District Courts, and Daily Orders and, thus, provides unparalleled diversity and coverage. Our dataset surpasses existing datasets like PredEx and ILDC, offering a comprehensive foundation for advanced AI research in the legal domain. In addition to the dataset, we present INLegalLlama, a domain-specific generative large language model (LLM) tailored to the intricacies of the Indian legal system. It is developed through a two-phase training approach over a base LLaMa model. First, Indian legal documents are injected using continual pretraining. Second, task-specific supervised finetuning is done. This method allows the model to achieve a deeper understanding of legal contexts. Our experiments demonstrate that incorporating diverse court data significantly boosts model accuracy, achieving approximately 90% F1-score in prediction tasks. INLegalLlama not only improves prediction accuracy but also offers comprehensible explanations, addressing the need for explainability in AI-assisted legal decisions. Shubham Kumar Nigam, Balaramamahanthi Deepak Patnaik, Shivam Mishra, Noel Shallum, Kripabandhu Ghosh, Arnab Bhattacharya 0001 |
COLING | 5 |
| 2025 | CryptOpiQA: A new Opinion and Question Answering dataset on CryptocurrencyabstractCryptocurrency has attracted a lot of public attention and opinion worldwide. Users have different kinds of information needs regarding such topics and publicly available information is a good resource to satisfy those information needs. In this paper, we investigate the public opinion on cryptocurrency and bitcoin on two social media – Twitter and Reddit. We have created a multi-level dataset CryptOpiQA and garnered valuable insights. The dataset contains both gold standard (manually annotated) and silver standard (inferred from the gold standard) labels. As a part of this dataset, we have also created a Question Answering sub-corpus. We have used state-of-the-art LLMs and advanced techniques such as retrieval augmented generation (RAG) to improve question-answering (QnA) results. We believe this dataset and the analysis will be useful in studying user opinions and Question-Answering on cryptocurrency in the research community. Sougata Sarkar, Aditya Badwal, Amartya Roy, Koustav Rudra, Kripabandhu Ghosh |
COLING | 5 |
| 2025 | Can LLMs be Literary Companions?: Analysing LLMs on Bengali Figures of Speech IdentificationabstractDespite Bengali being among the most spoken languages bearing cultural importance and richness, the NLP endeavors on it, remain relatively limited.Figures of speech (FoS) not only contribute to the phonetic and semantic nuances of a language, but they also exhibit aesthetics, expression, and creativity in literature.To our knowledge, in this paper, we present the first ever Bengali figures of speech classification dataset, BengFoS, on works of six renowned poets of Bengali literature.We deploy state-of-the-art (SoTA) models to this dataset, improve them, and finally dissect them, revealing novel insights on the intrinsic behavior of two open-source LLMs (Llama and DeepSeek) in FoS detection.Though we focused on Bengali, the experimental framework can be reproduced for English as well as for other low-resource languages.1 Kripabandhu Ghosh |
EMNLP | 2 |
| 2023 | Legal IR and NLP: The History, Challenges, and State-of-the-Art
Debasis Ganguly, Jack G. Conrad, Kripabandhu Ghosh, Saptarshi Ghosh 0001, Pawan Goyal 0002, Paheli Bhattacharya, Shubham Kumar Nigam, Shounak Paul |
ECIR (3) | 3 |
| 2023 | Exemplar-Free Continual Transformer with ConvolutionsabstractContinual Learning (CL) involves training a machine learning model in a sequential manner to learn new information while retaining previously learned tasks without the presence of previous training data. Although there has been significant interest in CL, most recent CL approaches in computer vision have focused on convolutional architectures only. However, with the recent success of vision transformers, there is a need to explore their potential for CL. Although there have been some recent CL approaches for vision transformers, they either store training instances of previous tasks or require a task identifier during test time, which can be limiting. This paper proposes a new exemplar-free approach for class/task incremental learning called ConTraCon, which does not require task-id to be explicitly present during inference and avoids the need for storing previous training instances. The proposed approach leverages the transformer architecture and involves re-weighting the key, query, and value weights of the multi-head self-attention layers of a transformer trained on a similar task. The re-weighting is done using convolution, which enables the approach to maintain low parameter requirements per task. Additionally, an image augmentation-based entropic task identification approach is used to predict tasks without requiring task-ids during inference. Experiments on four benchmark datasets demonstrate that the proposed approach outperforms several competitive approaches while requiring fewer parameters.1 Anurag Roy, Vinay Kumar Verma, Sravan Voonna, Kripabandhu Ghosh, Saptarshi Ghosh 0001, Abir Das |
ICCV | 4 |
| 2023 | LeDA: A System for Legal Data AnnotationabstractThis paper presents LeDA, a system for Legal Data Annotation. The system offers the functionality of annotating and categorising text spans representing legal concepts that capture the topic of a document, and also supports a meta-annotator to adjudicate the ground truth created by different annotators. Notably, our system supports a dynamic update of the ontology by enabling the creation of new legal concepts. Currently employed to annotate key legal concepts, LeDA aims to construct concept-based semantic representations for tasks such as similar case retrieval, and judgment prediction. Subinay Adhikary, Dwaipayan Roy 0001, Debasis Ganguly, Shouvik Kumar Guha, Kripabandhu Ghosh |
JURIX | 5 |
| 2023 | Fairness for both Readers and Authors: Evaluating Summaries of User Generated ContentabstractSummarization of textual content has many applications, ranging from summarizing long documents to recent efforts towards summarizing user generated text (e.g., tweets, Facebook or Reddit posts). Traditionally, the focus of summarization has been to generate summaries which can best satisfy the readers. In this work, we look at summarization of user-generated content as a two-sided problem where satisfaction of both readers and authors is crucial. Through three surveys, we show that for user-generated content, traditional evaluation approach of measuring similarity between reference summaries and algorithmic summaries cannot capture author satisfaction. We propose an author satisfaction-based evaluation metric CROSSEM which, we show empirically, can potentially complement the current evaluation paradigm. We further propose the idea of inequality in satisfaction, to account for individual fairness amongst readers and authors. To our knowledge, this is the first attempt towards developing a fair summary evaluation framework for user generated content, and is likely to spawn lot of future research in this space. Garima Chhikara, Kripabandhu Ghosh, Saptarshi Ghosh 0001, Abhijnan Chakraborty |
SIGIR | 2 |
| 2022 | Legal case document similarity: You need both network and text
Paheli Bhattacharya, Kripabandhu Ghosh, Arindam Pal 0001, Saptarshi Ghosh 0001 |
Inf. Process. Manag. | 2 |
| 2021 | ILDC for CJPE: Indian Legal Documents Corpus for Court Judgment Prediction and ExplanationabstractVijit Malik, Rishabh Sanjay, Shubham Kumar Nigam, Kripabandhu Ghosh, Shouvik Kumar Guha, Arnab Bhattacharya, Ashutosh Modi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Vijit Malik, Rishabh Sanjay, Shubham Kumar Nigam, Kripabandhu Ghosh, Shouvik Kumar Guha, Arnab Bhattacharya 0001, Ashutosh Modi |
ACL/IJCNLP (1) | 4 |
| 2021 | Incorporating domain knowledge for extractive summarization of legal case documentsabstractAutomatic summarization of legal case documents is an important and practical challenge. Apart from many domain-independent text summarization algorithms that can be used for this purpose, several algorithms have been developed specifically for summarizing legal case documents. However, most of the existing algorithms do not systematically incorporate domain knowledge that specifies what information should ideally be present in a legal case document summary. To address this gap, we propose an unsupervised summarization algorithm DELSumm which is designed to systematically incorporate guidelines from legal experts into an optimization setup. We conduct detailed experiments over case documents from the Indian Supreme Court. The experiments show that our proposed unsupervised method outperforms several strong baselines in terms of ROUGE scores, including both general summarization algorithms and legal-specific ones. In fact, though our proposed algorithm is unsupervised, it outperforms several supervised summarization models that are trained over thousands of document-summary pairs. Paheli Bhattacharya, Soham Poddar, Koustav Rudra, Kripabandhu Ghosh, Saptarshi Ghosh 0001 |
ICAIL | 4 |
| 2021 | An Analytical Study of Algorithmic and Expert Summaries of Legal CasesabstractAutomatic summarization of legal case documents is an important and challenging problem, where algorithms attempt to generate summaries that match well with expert-generated summaries. This work takes the first step in analyzing expert-generated summaries and algorithmic summaries of legal case documents. We try to uncover how law experts write summaries for a legal document, how various generic as well as domain-specific extractive algorithms generate summaries, and how the expert summaries vary from the algorithmic summaries. We also analyze which important sentences of a legal case document are missed by most algorithms while generating summaries, in terms of the rhetorical roles of the sentences and the positions of the sentences in the legal document. Aniket Deroy, Paheli Bhattacharya, Kripabandhu Ghosh, Saptarshi Ghosh 0001 |
JURIX | 3 |
| 2020 | Fairness for Whom? Understanding the Reader's Perception of Fairness in Text SummarizationabstractWith the surge in user-generated textual information, there has been a recent increase in the use of summarization algorithms for providing an overview of the extensive content. Traditional metrics for evaluation of these algorithms (e.g. ROUGE scores) rely on matching algorithmic summaries to human-generated ones. However, it has been shown that when the textual contents are heterogeneous, e.g., when they come from different socially salient groups, most existing summarization algorithms represent the social groups very differently compared to their distribution in the original data. To mitigate such adverse impacts, some fairness-preserving summarization algorithms have also been proposed. All of these studies have considered normative notions of fairness from the perspective of writers of the contents, neglecting the readers' perceptions of the underlying fairness notions. To bridge this gap, in this work, we study the interplay between the fairness notions and how readers perceive them in textual summaries. Through our experiments, we show that reader's perception of fairness is often context-sensitive. Moreover, standard ROUGE evaluation metrics are unable to quantify the perceived (un)fairness of the summaries. To this end, we propose a human-in-the-loop metric and an automated graph-based methodology to quantify the perceived bias in textual summaries. We demonstrate their utility by quantifying the (un)fairness of several summaries of heterogeneous socio-political microblog datasets. Anurag Shandilya, Abhisek Dash, Abhijnan Chakraborty, Kripabandhu Ghosh, Saptarshi Ghosh 0001 |
IEEE BigData | 4 |
| 2020 | ZSCRGAN: A GAN-based Expectation Maximization Model for Zero-Shot Retrieval of Images from Textual DescriptionsabstractMost existing algorithms for cross-modal Information Retrieval are based on a supervised train-test setup, where a model learns to align the mode of the query (e.g., text) to the mode of the documents (e.g., images) from a given training set. Such a setup assumes that the training set contains an exhaustive representation of all possible classes of queries. In reality, a retrieval model may need to be deployed on previously unseen classes, which implies a zero-shot IR setup. In this paper, we propose a novel GAN-based model for zero-shot text to image retrieval. When given a textual description as the query, our model can retrieve relevant images in a zero-shot setup. The proposed model is trained using an Expectation-Maximization framework. Experiments on multiple benchmark datasets show that our proposed model comfortably outperforms several state-of-the-art zero-shot text to image retrieval models, as well as zero-shot classification and hashing models suitably used for retrieval. Anurag Roy, Vinay Kumar Verma, Kripabandhu Ghosh, Saptarshi Ghosh 0001 |
CIKM | 3 |
| 2020 | Retrieval of Prior Court Cases Using Witness TestimoniesabstractWitness testimonies are important constituents of a court case description and play a significant role in the final decision. We propose two techniques to identify sentences representing witness testimonies. The first technique employs linguistic rules whereas the second technique applies distant supervision where training set is constructed automatically using the output of the first technique. We then represent the identified witness testimonies in a more meaningful structure – event verb (predicate) along with its arguments corresponding to semantic roles A0 and A1 [1]. We demonstrate effectiveness of such representation in retrieving semantically similar prior relevant cases. To the best of our knowledge, this is the first paper to apply NLP techniques to extract witness information from court judgements and use it for retrieving prior court cases. Kripabandhu Ghosh, Sachin Pawar, Girish Keshav Palshikar, Pushpak Bhattacharyya, Vasudeva Varma |
JURIX | 1 |
| 2020 | Hier-SPCNet: A Legal Statute Hierarchy-based Heterogeneous Network for Computing Legal Case Document SimilarityabstractComputing similarity between two legal case documents is a challenging task, for which text-based and network-based measures have been proposed in literature. All prior network-based similarity methods considered a precedent citation network among case documents only (PCNet). However, this approach misses an important source of legal knowledge - the hierarchy of legal statutes that are applicable in a given legal jurisdiction (e.g., country). We propose to augment the PCNet with the hierarchy of legal statutes, to form a heterogeneous network Hier-SPCNet. Experiments over a set of Indian Supreme Court case documents show that Hier-SPCNet enables significantly better document similarity estimation, as compared to existing approaches using PCNet. We also show that the proposed network-based method can complement text-based measures for better estimation of legal document similarity. Paheli Bhattacharya, Kripabandhu Ghosh, Arindam Pal 0001, Saptarshi Ghosh 0001 |
SIGIR | 2 |
| 2020 | TAQE: Tweet Retrieval-Based Infrastructure Damage Assessment During DisastersabstractTwitter is an active communication channel for the spreading of updated information in emergency situations. Retrieving specific information related to infrastructure damage offers the situational views to the concerned authorities, who can take necessary action to disburse help. However, such usages of Twitter demand significant accuracy of the retrieved information. Previous techniques on IR have not been able to capture the semantic variations satisfactorily in the tweets, due to low content quality and vocabulary gap, and consequently have failed to yield considerable performance. This has left ample scope for further improvement in this area of research. There are two major contributions of our work: 1) developing a relevant tweet retrieval framework that provides information about infrastructure damage and 2) assignment of a relative damage score to the affected regions so that the severity of the damage can be assessed. Our proposed technique involves a novel split-query-based mechanism with topic aligned query expansion (TAQE) to retrieve relevant tweets that are subsequently used for measuring the infrastructure damage across different locations. We report empirical results on multiple-crisis-related data sets to establish the efficacy of our approach to these events at different locations. Empirical validation of our proposed approach on manually annotated ground-truth data reveals considerably better performance metrics in terms of precision, recall, Bpref, and MAP over several state-of-the-art techniques. Shalini Priya, Manish Bhanu, Sourav Kumar Dandapat, Kripabandhu Ghosh, Joydeep Chandra |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2019 | Identifying infrastructure damage during earthquake using deep active learningabstractTwitter provides important information for emergency responders in the rescue process during disasters. However, tweets containing relevant information are sparse and are usually hidden in a vast set of noisy contents. This leads to inherent challenges in generating suitable training data that are required for neural network models. In this paper, we study the problem of retrieving the infrastructure damage information from tweets generated from different location during crisis using the model actively trained on past but similar events. We combine RNN and GRU based model coupled with active learning that gets trained on most uncertain samples and captures the latent features of different data distribution. It reduces the uses of around 90% less training data, thereby significantly reducing the manual annotation efforts. We use the model pre-trained using active learning based approach to retrieve the infrastructure damage tweets originated from different regions. We obtain a minimum of 18% gain on F1-measure and considerably on other metrics over recent state-of-the-art IR techniques. Shalini Priya, Saharsh Singh, Sourav Kumar Dandapat, Kripabandhu Ghosh, Joydeep Chandra |
ASONAM | 4 |
| 2019 | A Comparative Study of Summarization Algorithms Applied to Legal Case Judgments
Paheli Bhattacharya, Kaustubh Hiware, Subham Rajgaria, Nilay Pochhi, Kripabandhu Ghosh, Saptarshi Ghosh 0001 |
ECIR (1) | 5 |
| 2019 | Identification of Rhetorical Roles of Sentences in Indian Legal JudgmentsabstractAutomatically understanding the rhetorical roles of sentences in a legal case judgement is an important problem to solve, since it can help in several downstream tasks like summarization of legal judgments, legal search, and so on. The task is challenging since legal case documents are usually not well-structured, and these rhetorical roles may be subjective (as evident from variation of opinions between legal experts). In this paper, we address this task for judgments from the Supreme Court of India. We label sentences in 50 documents using multiple human annotators, and perform an extensive analysis of the human-assigned labels. We also attempt automatic identification of the rhetorical roles of sentences. While prior approaches towards this task used Conditional Random Fields over manually handcrafted features, we explore the use of deep neural models which do not require hand-crafting of features. Experiments show that neural models perform much better in this task than baseline methods which use handcrafted features. Paheli Bhattacharya, Shounak Paul, Kripabandhu Ghosh, Saptarshi Ghosh 0001, Adam Z. Wyner |
JURIX | 3 |
| 2019 | Utilizing microblogs for assisting post-disaster relief operations via matching resource needs and availabilities
Ritam Dutt, Moumita Basu, Kripabandhu Ghosh, Saptarshi Ghosh 0001 |
Inf. Process. Manag. | 3 |
| 2019 | Summarizing User-generated Textual Content: Motivation and Methods for Fairness in Algorithmic SummariesabstractAs the amount of user-generated textual content grows rapidly, text summarization algorithms are increasingly being used to provide users a quick overview of the information content. Traditionally, summarization algorithms have been evaluated only based on how well they match human-written summaries (e.g. as measured by ROUGE scores). In this work, we propose to evaluate summarization algorithms from a completely new perspective that is important when the user-generated data to be summarized comes from different socially salient user groups, e.g. men or women, Caucasians or African-Americans, or different political groups (Republicans or Democrats). In such cases, we check whether the generated summaries fairly represent these different social groups. Specifically, considering that an extractive summarization algorithm selects a subset of the textual units (e.g. microblogs) in the original data for inclusion in the summary, we investigate whether this selection is fair or not. Our experiments over real-world microblog datasets show that existing summarization algorithms often represent the socially salient user-groups very differently compared to their distributions in the original data. More importantly, some groups are frequently under-represented in the generated summaries, and hence get far less exposure than what they would have obtained in the original data. To reduce such adverse impacts, we propose novel fairness-preserving summarization algorithms which produce high-quality summaries while ensuring fairness among various groups. To our knowledge, this is the first attempt to produce fair text summarization, and is likely to open up an interesting research direction. Abhisek Dash, Anurag Shandilya, Arindam Biswas 0004, Kripabandhu Ghosh, Saptarshi Ghosh 0001, Abhijnan Chakraborty |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2019 | Extracting Resource Needs and Availabilities From Microblogs for Aiding Post-Disaster Relief OperationsabstractMicroblogging sites like Twitter are the important sources of real-time information during disaster/emergency events. During such events, the critical situational information posted is immersed in a lot of conversational content; hence, reliable methodologies are needed for extracting the meaningful information. In this paper, we focus on a particular application that is critical for efficient management of post-disaster relief operations - identifying tweets that inform about resource needs and resource availabilities. Two broad types of methodologies can be practically applied to identify such tweets during an ongoing disaster event: 1) supervised classification approaches, where the classifier models are trained on microblogs posted during prior events and applied on those posted during the ongoing event and 2) unsupervised pattern matching and information retrieval approaches that can be directly applied on the microblogs posted during the ongoing event. In this paper, we experiment with several supervised and unsupervised approaches to address the problem, including several neural network-based classification and retrieval models. We also propose two novel neural retrieval models (unsupervised) for the said application, which effectively combine word-level embeddings and character-level embeddings. We conduct experiments on tweets posted during two disaster events and observe that the two approaches perform well in different scenarios. Specifically, if good quality training data are available from prior events, then classification approaches perform better; however, if such training data are not available, then unsupervised retrieval methods outperform supervised classification approaches. Moumita Basu, Anurag Shandilya, Prannay Khosla, Kripabandhu Ghosh, Saptarshi Ghosh 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2018 | Characterizing Infrastructure Damage After Earthquake: A Split-Query Based IR ApproachabstractRetrieving relevant information from social media based on specific requirements has become a focus area for researchers. In this paper, we propose a framework for online retrieval of tweets providing information about possible infrastructure damages, caused due to earthquakes and use the same to determine a damage score for the possibly affected locations. Identifying such tweets would not only provide a holistic view of the affected areas but would also help in taking necessary relief actions. Existing works on this topic fail to effectively capture the semantic variation in the tweets, possibly due to poor content quality, thereby providing scopes for further improvement in the mechanisms involved. Our proposed technique relies on a novel split-query based mechanism along with a pseudo-relevance feedback approach to identify the relevant tweets. The pseudo-relevance feedback approach expands on an initial set of seed tweets obtained using a semi-automatic query generation mechanism that couples topic based clustering with human annotation. Empirical validation of our proposed method on a manually annotated ground truth data reveals a considerable improvement in precision, recall and mean average precision over several baseline methods. Shalini Priya, Manish Bhanu, Sourav Kumar Dandapat, Kripabandhu Ghosh, Joydeep Chandra |
ASONAM | 4 |
| 2017 | Identifying Post-Disaster Resource Needs and Availabilities from MicroblogsabstractMicroblogging sites like Twitter are increasingly being used for aiding post-disaster relief operations. In such situations, identifying needs and availabilities of various types of resources is critical for effective coordination of the relief operations. We focus on the problem of automatically identifying tweets that inform about needs and availabilities of resources, termed as need-tweets and availability-tweets respectively. Traditionally, pattern matching techniques are adopted to identify such tweets. In this work, we present novel retrieval methodologies, based on word embeddings, for automatically identifying need-tweets and availability-tweets. Experiments over tweets posted during two recent disaster events show that the proposed methodologies outperform prior pattern-matching techniques. Moumita Basu, Kripabandhu Ghosh, Somenath Das, Ratnadeep Dey, Somprakash Bandyopadhyay, Saptarshi Ghosh 0001 |
ASONAM | 2 |
| 2017 | Automatic Catchphrase Identification from Legal Court Case DocumentsabstractAutomatically identifying catchphrases from legal court case documents is an important problem in Legal Information Retrieval, which has not been extensively studied. In this work, we propose an unsupervised approach for extraction and ranking of catchphrases from court case documents, by focusing on noun phrases. Using a dataset of gold standard catchphrases created by legal experts from real-life court documents, we compare the proposed approach with several unsupervised and supervised baselines. We show that the proposed methodology achieves statistically significantly better performance compared to all the baselines. Arpan Mandal, Kripabandhu Ghosh, Arindam Pal 0001, Saptarshi Ghosh 0001 |
CIKM | 2 |
| 2017 | Combining Local and Global Word Embeddings for Microblog StemmingabstractStemming is a vital step employed to improve retrieval performance through efficient unification of morphological variants of a word. We propose an unsupervised, context-specific stemming algorithm for microblogs, based on both local and global word embeddings, which is capable of handling the informal, noisy vocabulary of microblogs. Experiments on two standard microblog data collections (TREC 2016 and FIRE 2016) show that, the proposed stemmer enables significantly better retrieval performance than several state-of-the-art stemming algorithms, for the same queries. Anurag Roy, Trishnendu Ghorai, Kripabandhu Ghosh, Saptarshi Ghosh 0001 |
CIKM | 3 |
| 2017 | A Novel Word Embedding Based Stemming Approach for Microblog Retrieval During Disasters
Moumita Basu, Anurag Roy, Kripabandhu Ghosh, Somprakash Bandyopadhyay, Saptarshi Ghosh 0001 |
ECIR | 3 |
| 2016 | Improving Information Retrieval Performance on OCRed Text in the Absence of Clean Text Ground Truth
Kripabandhu Ghosh, Anirban Chakraborty 0002, Swapan K. Parui, Prasenjit Majumder |
Inf. Process. Manag. | 1 |
| 2015 | Clustered Semi-Supervised Relevance FeedbackabstractIn relevance feedback, first-round search results are used to boost second-round search results. Two forms have been traditionally considered: exhaustively labelled feedback, where all first-round results to depth k are annotated for relevance by the user; and blind feedback, where the top-k results are all assumed to be relevant. In this paper, we consider an intermediate, semi-supervised scheme, in which only a subset of results is selected for annotation, and then their labels are propagated to their nearest neighbours. Specifically, we use clustering to determine the nearest-neighbour groups, and seed selection to choose documents for annotation. We find that the effectiveness of this method is indistinguishable from the exhaustive relevance feedback, and is significantly higher than both blind feedback and the use of the annotated subset alone. We show that this approach works well in environments in which some but limited amounts of human feedback are available, such as early case assessment in e-discovery. Kripabandhu Ghosh, Swapan K. Parui |
CIKM | 1 |
| 2015 | Retrieval from Noisy E-Discovery Corpus in the Absence of Training DataabstractOCR errors hurt retrieval performance to a great extent. Research has been done on modelling and correction of OCR errors. However, most of the existing systems use language dependent resources or training texts for studying the nature of errors. Not much research has been reported on improving retrieval performance from erroneous text when no training data is available. We propose a novel algorithm for detecting OCR errors and improving retrieval performance on an E-Discovery corpus. Our contribution is two-fold : (1) identifying erroneous variants of query terms for improvement in retrieval performance, and (2) presenting a scope for a possible error-modelling in the erroneous corpus where clean ground truth text is not available for comparison. Our algorithm does not use any training data or any language specific resources like thesaurus. It also does not use any knowledge about the language except that the word delimiter is blank space. The proposed approach obtained statistically significant improvements in recall over state-of-the-art baselines. Anirban Chakraborty 0002, Kripabandhu Ghosh, Swapan K. Parui |
SIGIR | 2 |
| 2015 | Learning combination weights in data fusion using Genetic Algorithms
Kripabandhu Ghosh, Swapan K. Parui, Prasenjit Majumder |
Inf. Process. Manag. | 1 |
| 2012 | Improving e-discovery using information retrievalabstractE-discovery is the requirement that the documents and information in electronic form stored in corporate systems be produced as evidence in litigation. It has posed great challenges for legal experts. Legal searchers have always looked to find "any and all" evidence for a given case. Thus, a legal search system would essentially be a recall-oriented system. It has been a common practice among expert searchers to formulate Boolean queries to represent their information need. We want to work on three basic problems: Boolean query formulation - Our primary goal is to study Boolean query formulation in the light of the E-discovery task. This will include automatic Boolean query generation, expansion and learning the effect of proximity operators in Boolean searches. Data fusion - We would also like to explore the effectiveness of data fusion techniques in improving recall. Error modeling - Finally, we will work on error modeling methods for noisy legal documents. Kripabandhu Ghosh |
SIGIR | 1 |