Rajdeep Sarkar

dblp:172/0876 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
5since 2021 · last 2024
0000-0002-3942-2978ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 A Hybrid Approach to Aspect Based Sentiment Analysis Using Transfer Learning
abstract
Aspect-Based Sentiment Analysis ( ABSA) aims to identify terms or multiword expressions (MWEs) on which sentiments are expressed and the sentiment polarities associated with them. The development of supervised models has been at the forefront of research in this area. However, training these models requires the availability of manually annotated datasets which is both expensive and time-consuming. Furthermore, the available annotated datasets are tailored to a specific domain, language, and text type. In this work, we address this notable challenge in current state-of-the-art ABSA research. We propose a hybrid approach for Aspect Based Sentiment Analysis using transfer learning. The approach focuses on generating weakly-supervised annotations by exploiting the strengths of both large language models (LLM) and traditional syntactic dependencies. We utilise syntactic dependency structures of sentences to complement the annotations generated by LLMs, as they may overlook domain-specific aspect terms. Extensive experimentation on multiple datasets is performed to demonstrate the efficacy of our hybrid method for the tasks of aspect term extraction and aspect sentiment classification.
Gaurav Negi, Rajdeep Sarkar, Omnia Zayed, Paul Buitelaar
LREC/COLING2
2024 BRECS: Enhanced Binary Representation of Word Embeddings via Cosine Similarity
abstract
Word representations like GloVe and Word2Vec encapsulate semantic and syntactic attributes and constitute the fundamental building block in diverse Natural Language Processing (NLP) applications. Such vector embeddings are typically stored in float32 format, and for a substantial vocabulary size, they impose considerable memory and computational demands due to the resource-intensive float32 operations. Thus, representing words via binary embeddings has emerged as a promising but challenging solution. In this paper, we introduce BRECS, an autoencoder-based Siamese framework for the generation of enhanced binary word embeddings (from the original embeddings). We propose the use of the novel Binary Cosine Similarity (BCS) regularisation in BRECS, which enables it to learn the semantics and structure of the vector space spanned by the original word embeddings, leading to better binary representation generation. We further show that our framework is tailored with independent parameters within the various components, thereby providing it with better learning capability. Extensive experiments across multiple datasets and tasks demonstrate the effectiveness of BRECS, compared to existing baselines for static and contextual binary word embedding generation. The source code is available at https://github.com/rajbsk/brecs.
Rajdeep Sarkar, Sourav Dutta 0001, John P. McCrae
ECAI1
2023 PICKD: In-Situ Prompt Tuning for Knowledge-Grounded Dialogue Generation
Rajdeep Sarkar, Koustava Goswami, Mihael Arcan, John P. McCrae
PAKDD (4)1
2022 Semantic Aware Answer Sentence Selection Using Self-Learning Based Domain Adaptation
abstract
Selecting an appropriate and relevant context forms an essential component for the efficacy of several information retrieval applications like Question Answering (QA) systems. The problem of Answer Sentence Selection (AS2) refers to the task of selecting sentences, from a larger text, that are relevant and contain the answer to users' queries. While there has been a lot of success in building AS2 systems trained on open-domain data (e.g., SQuAD, NQ), they do not generalize well in closed-domain settings, since domain adaptation can be challenging due to poor availability and annotation expense of domain-specific data. This paper proposes SEDAN, an effective self-learning framework to adapt AS2 models for domain-specific applications. We leverage large pre-trained language models to automatically generate domain-specific QA pairs for domain adaptation. We further fine-tune a pre-trained Sentence-BERT architecture to capture semantic relatedness between questions and answer sentences for AS2. Extensive experiments demonstrate the effectiveness of our proposed approach (over existing state-of-the-art AS2 baselines) on different Question Answering benchmark datasets.
Rajdeep Sarkar, Sourav Dutta 0001, Haytham Assem, Mihael Arcan, John P. McCrae
KDD1
2021 Qasar: Self-Supervised Learning Framework for Extractive Question Answering
abstract
Question Answering (QA) has become a foundational research area in Natural Language Understanding (NLU) with widespread applications in search, personal digital assistance, and conversational systems. Despite the success in open-domain question answering, existing extractive question answering models pre-trained using Wikipedia articles (e.g., SQuAD data) perform rather poorly in closed-domain and industrial scenarios. Further, a major limitation in adapting question answering systems to such contexts is the poor availability and the expensive annotation of domain-specific data. Thus, wide applicability of QA models are severely hampered in enterprise systems.In this paper, we aim to address the above challenges by introducing a novel QA framework, Qasar, using self-supervised learning for efficient domain adaptation. We show, for the first time, the advantage of fine-tuning pre-trained QA models for closed-domains by synthetically generated domain-specific questions and answers (from relevant documents) from large language models like T5. Further, we also propose a novel context retrieval component based on question-context semantic relatedness to further boost the accuracy of the Qasar QA framework. Experimental results show significant performance improvements on both open-and closed-domain QA datasets, while requiring no labelling efforts, which we believe will contribute to the ease of deployment of such systems in enterprise settings. The different modules of our framework (synthetic data generation, context retrieval, and question answering) can be fully reproduced by fine-tuning publicly available language models and QA models on SQuAD dataset as discussed in the paper.
Haytham Assem, Rajdeep Sarkar, Sourav Dutta 0001
IEEE BigData2
2020 Unsupervised Deep Language and Dialect Identification for Short Texts
abstract
Automatic Language Identification (LI) or Dialect Identification (DI) of short texts of closely related languages or dialects, is one of the primary steps in many natural language processing pipelines.Language identification is considered a solved task in many cases; however, in the case of very closely related languages, or in an unsupervised scenario (where the languages are not known in advance), performance is still poor.In this paper, we propose the Unsupervised Deep Language and Dialect Identification (UDLDI) method, which can simultaneously learn sentence embeddings and cluster assignments from short texts.The UDLDI model understands the sentence constructions of languages by applying attention to character relations which helps to optimize the clustering of languages.We have performed our experiments on three shorttext datasets for different language families, each consisting of closely related languages or dialects, with very minimal training sets.Our experimental evaluations on these datasets have shown significant improvement over state-of-the-art unsupervised methods and our model has outperformed state-of-the-art LI and DI systems in supervised settings.
Koustava Goswami, Rajdeep Sarkar, Bharathi Raja Chakravarthi, Theodorus Fransen, John P. McCrae
COLING2
2020 Suggest me a movie for tonight: Leveraging Knowledge Graphs for Conversational Recommendation
abstract
Conversational recommender systems focus on the task of suggesting products to users based on the conversation flow.Recently, the use of external knowledge in the form of knowledge graphs has shown to improve the performance in recommendation and dialogue systems.Information from knowledge graphs aids in enriching those systems by providing additional information such as closely related products and textual descriptions of the items.However, knowledge graphs are incomplete since they do not contain all factual information present on the web.Furthermore, when working on a specific domain, knowledge graphs in its entirety contribute towards extraneous information and noise.In this work, we study several subgraph construction methods and compare their performance across the recommendation task.We incorporate pre-trained embeddings from the subgraphs along with positional embeddings in our models.Extensive experiments show that our method has a relative improvement of at least 5.62% compared to the state-of-the-art on multiple metrics on the recommendation task.
Rajdeep Sarkar, Koustava Goswami, Mihael Arcan, John P. McCrae
COLING1
2019 Automated Early Leaderboard Generation from Comparative Tables
Mayank Singh 0001, Rajdeep Sarkar, Atharva Vyas, Pawan Goyal 0002, Animesh Mukherjee 0001, Soumen Chakrabarti
ECIR (1)2
2018 A supervised approach to taxonomy extraction using word embeddings
Rajdeep Sarkar, John P. McCrae, Paul Buitelaar
LREC1
2017 Relay-Linking Models for Prominence and Obsolescence in Evolving Networks
abstract
The rate at which nodes in evolving social networks acquire links (friends, citations) shows complex temporal dynamics. Preferential attachment and link copying models, while enabling elegant analysis, only capture rich-gets-richer effects, not aging and decline. Recent aging models are complex and heavily parameterized; most involve estimating 1-3 parameters per node. These parameters are intrinsic: they explain decline in terms of events in the past of the same node, and do not explain, using the network, where the linking attention might go instead. We argue that traditional characterization of linking dynamics are insufficient to judge the faithfulness of models. We propose a new temporal sketch of an evolving graph, and introduce several new characterizations of a network's temporal dynamics. Then we propose a new family of frugal aging models with no per-node parameters and only two global parameters. Our model is based on a surprising inversion or undoing of triangle completion, where an old node relays a citation to a younger follower in its immediate vicinity. Despite very few parameters, the new family of models shows remarkably better fit with real data. Before concluding, we analyze temporal signatures for various research communities yielding further insights into their comparative dynamics. To facilitate reproducible research, we shall soon make all the codes and the processed dataset available in the public domain.
Mayank Singh 0001, Rajdeep Sarkar, Pawan Goyal 0002, Animesh Mukherjee 0001, Soumen Chakrabarti
KDD2
2017 Tabloids in the Era of Social Media?: Understanding the Production and Consumption of Clickbaits in Twitter
abstract
With the growing shift towards news consumption primarily through social media sites like Twitter, most of the traditional as well as new-age media houses are promoting their news stories by tweeting about them. The competition for user attention in such mediums has led many media houses to use catchy sensational form of tweets to attract more users - a process known as clickbaiting. In this work, using an extensive dataset collected from Twitter, we analyze the social sharing patterns of clickbait and non-clickbait tweets to determine the organic reach of such tweets. We also attempt to study the sections of Twitter users who actively engage themselves in following clickbait and non-clickbait tweets. Comparing the advent of clickbaits with the rise of tabloidization of news, we bring out several important insights regarding the news consumers as well as the media organizations promoting news stories on Twitter.
Abhijnan Chakraborty, Rajdeep Sarkar, Ayushi Mrigen, Niloy Ganguly
Proc. ACM Hum. Comput. Interact.2