Raghav Jain

dblp:265/2414 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Assessing current capabilities for incorporating lipidomics in multiomics data integration
abstract
Comprehensive analyses of multiple biological components including nucleic acids, proteins, metabolites, and lipids (i.e. "multiomics") provide unique insights into complex biological processes. Combining insights from these components through multiomics data integration enhances the depth and nuance of biological understanding available from these measurements. Among methods that integrate data across different technologies (e.g. mass spectrometry, sequencing), those that link components based on biological prior knowledge-pathway analyses-represent the most direct way of translating molecular-level observations into meaningful biological insights. However, significant barriers exist that prevent full utilization of metabolomics and especially lipidomics data in pathway integration. Challenges include the fast turnover and complex interactions of small molecules compared to biological macromolecules, low metabolite annotation rates, isomerism among lipids, and a lack of lipid representation in existing pathway knowledge bases. While these issues stem from a variety of causes, improvements to multiomics pathway integration including better incorporation of lipids into pathway knowledge bases and increased adoption of artificial intelligence approaches can greatly enhance the utility of small molecule -omics data, particularly for lipidomics. Here, we describe the current landscape of multiomics integration tools with an emphasis on support for metabolomics and lipidomics data, we highlight their capabilities and gaps with tangible examples using real multiomics data, and provide our perspective on how these approaches can be improved to better support generation of useful biological insights from complex multiomics data.
Dylan H. Ross, Raghav Jain, Hyeyoon Kim, Javier E. Flores, Soumaydeep Sarkar, Chaevien S. Clendinen, Jennifer E. Kyle, Tao Liu 0003, Sara J. C. Gosline
Briefings Bioinform.2
2025 Using Shapley interactions to understand how models use structure
abstract
Divyansh Singhvi, Diganta Misra, Andrej Erkelens, Raghav Jain, Isabel Papadimitriou, Naomi Saphra. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Divyansh Singhvi, Diganta Misra, Andrej Erkelens, Raghav Jain, Isabel Papadimitriou, Naomi Saphra
ACL (1)4
2024 CLIPSyntel: CLIP and LLM Synergy for Multimodal Question Summarization in Healthcare
abstract
In the era of modern healthcare, swiftly generating medical question summaries is crucial for informed and timely patient care. Despite the increasing complexity and volume of medical data, existing studies have focused solely on text-based summarization, neglecting the integration of visual information. Recognizing the untapped potential of combining textual queries with visual representations of medical conditions, we introduce the Multimodal Medical Question Summarization (MMQS) Dataset. This dataset, a major contribution of our work, pairs medical queries with visual aids, facilitating a richer and more nuanced understanding of patient needs. We also propose a framework, utilizing the power of Contrastive Language Image Pretraining(CLIP) and Large Language Models(LLMs), consisting of four modules that identify medical disorders, generate relevant context, filter medical concepts, and craft visually aware summaries. Our comprehensive framework harnesses the power of CLIP, a multimodal foundation model, and various general-purpose LLMs, comprising four main modules: the medical disorder identification module, the relevant context generation module, the context filtration module for distilling relevant medical concepts and knowledge, and finally, a general-purpose LLM to generate visually aware medical question summaries. Leveraging our MMQS dataset, we showcase how visual cues from images enhance the generation of medically nuanced summaries. This multimodal approach not only enhances the decision-making process in healthcare but also fosters a more nuanced understanding of patient queries, laying the groundwork for future research in personalized and responsive medical care. Disclaimer: The article features graphic medical imagery, a result of the subject's inherent requirements.
Akash Ghosh, Arkadeep Acharya, Raghav Jain, Sriparna Saha 0001, Aman Chadha, Setu Sinha
AAAI3
2024 MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention
abstract
Prince Jha, Raghav Jain, Konika Mandal, Aman Chadha, Sriparna Saha, Pushpak Bhattacharyya. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Prince Jha, Raghav Jain, Konika Mandal, Aman Chadha, Sriparna Saha 0001, Pushpak Bhattacharyya
ACL (1)2
2024 Meme-ingful Analysis: Enhanced Understanding of Cyberbullying in Memes Through Multimodal Explanations
abstract
Prince Jha, Krishanu Maity, Raghav Jain, Apoorv Verma, Sriparna Saha, Pushpak Bhattacharyya. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Prince Jha, Krishanu Maity, Raghav Jain, Apoorv Verma, Sriparna Saha 0001, Pushpak Bhattacharyya
EACL (1)3
2024 MedSumm: A Multimodal Approach to Summarizing Code-Mixed Hindi-English Clinical Queries
Akash Ghosh, Arkadeep Acharya, Prince Jha, Sriparna Saha 0001, Aniket Gaudgaul, Rajdeep Majumdar, Aman Chadha, Raghav Jain, Setu Sinha, Shivani Agarwal 0005
ECIR (5)8
2024 Timeline Summarization in the Era of LLMs
abstract
Timeline summarization is the task of automatically generating concise overviews of documents that capture the key events and their progression on timelines. While this capability is useful for quickly comprehending event sequences without reading lengthy descriptions, timeline summarization remains a relatively underexplored area in recent years when compared to traditional document summarization task and their evolution. The advent of large language models (LLMs) has led some to presume summarization as a solved problem. However, timeline summarization poses unique challenges for LLMs. Our investigation is centered on evaluating the performance of LLMs, against state-of-the-art models in this field. We employed three different approaches: chunking, knowledge graph-based summarization, and TimeRanker. Each of these methods was systematically tested on three benchmark datasets for timeline summarization to assess their effectiveness in capturing and condensing key events and their evolution within timelines. Our findings reveal that while LLMs show promise, timeline summarization remains a complex task that is not yet fully resolved.
Daivik Sojitra, Raghav Jain, Sriparna Saha 0001, Adam Jatowt, Manish Gupta 0001
SIGIR2
2024 Explainable Cyberbullying Detection in Hinglish: A Generative Approach
abstract
The escalating prevalence of online cyberbullying and trolling across various social media platforms has becomes a pressing concern. Extensive research demonstrates the detrimental impact of cyberbullying on the mental well-being of its victims. Given the sheer volume of online content, manual identification of cyberbullying instances proves unfeasible, necessitating the development of automated cyberbullying detection methods. This challenge has attracted considerable attention within the natural language processing (NLP) community, owing to advancements in machine learning techniques. However, most of the methods fail to provide reasoning for their decisions which warrants the use of interpretable models that can explain the model’s output in real-time. Interpretable models rather than black-box models with high performance are becoming popular adhering the “right to explanations” laws. Motivated by this, we create a cyberbullying corpusBullyExplainin code-mixed language, where a post has been annotated with four labels, i.e., bully, sentiment, target, and rationales (explainability). Current work addresses the task of explainable cyberbully detection and proposes a unified generative framework,BullyGenby redefining this multitask problem as a text-to-text generation task. Our framework is capable of not only detecting whether the text is a cyberbully or not but also provides reasoning by predicting rationale, target group and sentiment of the text. Experimental results illustrate the efficacy of our proposed model by outperforming the state-of-the-art and several baselines by a significant margin and conclude that text-to-text generation model could be a good alternative for multitask classification problems.
Krishanu Maity, Raghav Jain, Prince Jha, Sriparna Saha 0001
IEEE Trans. Comput. Soc. Syst.2
2024 Toward Multimodal Complaint Severity Detection From Social Media
abstract
The prevalence of complaints submitted online and the sheer volume of information made available by social media platforms highlight the need for automated complaint analysis tools. In linguistic studies, complaints have been classified according to how much of personal risk the complainant is willing to take. This is crucial information for understanding the motivations of complainants and how people come up with reasonable means of reparation. Few attempts have been made to use the existing multimodal complaint model, which focuses on improving the textual mode with the help of images, to find specific visual features that help identify complaints. Our aim is to find a solution to this problem. In order to detect complaints and the severity level associated with them in a multitask setting, we propose Multimodal framEwork for Complaint and Severity-level detectIon (MECSI), a novel multimodal framework that uses local and global attributes (in both modalities) and relates them to the textual context. To do this, we add severity-level annotation to the newly released CESAMARD dataset, which is a compilation of reviews and images of products listed on the Amazon website. The experimental findings confirmed the superiority of our proposed model over the state-of-the-art model and other strong rival baselines, proving the efficacy of our proposed framework.
Apoorva Singh, Prince Jha, Souryadip Das, Raghav Jain, Sriparna Saha 0001
IEEE Trans. Comput. Soc. Syst.4
2023 Peeking inside the black box: A Commonsense-aware Generative Framework for Explainable Complaint Detection
abstract
Complaining is an illocutionary act in which the speaker communicates his/her dissatisfaction with a set of circumstances and holds the hearer (the complainee) answerable, directly or indirectly.Considering breakthroughs in machine learning approaches, the complaint detection task has piqued the interest of the natural language processing (NLP) community.Most of the earlier studies failed to justify their findings, necessitating the adoption of interpretable models that can explain the model's output in real-time.We introduce an explainable complaint dataset, X-CI, the first benchmark dataset for explainable complaint detection.Each instance in the X-CI dataset is annotated with five labels: complaint label, emotion label, polarity label, complaint severity level, and rationale (explainability), i.e., the causal span explaining the reason for the complaint/noncomplaint label.We address the task of explainable complaint detection and propose a commonsense-aware unified generative framework by reframing the multitask problem as a text-to-text generation task.Our framework can predict the complaint cause, severity level, emotion, and polarity of the text in addition to detecting whether it is a complaint or not.We further establish the advantages of our proposed model on various evaluation metrics over the state-of-the-art models and other baselines when applied to the X-CI dataset in both full and few-shot settings 1 .
Apoorva Singh, Raghav Jain, Prince Jha, Sriparna Saha 0001
ACL (1)2
2023 Investigating the Impact of Multimodality and External Knowledge in Aspect-level Complaint and Sentiment Analysis
abstract
Automated complaint analysis is vital for generating critical insights, which in turn enhance customer satisfaction, product quality, and overall business performance. Nevertheless, conventional methods frequently fail to capture the nuances of aspect-level complaints and inadequately utilize external knowledge, thus creating a gap in effective complaint detection and analysis. In response to this issue, we proactively explore the role of external knowledge and multimodality in this domain. This leads to the development of MGasD (Multimodal Generative framework for aspect-based complaint and sentiment Detection), a multimodal knowledge-infused unified framework. MGasD diverges from traditional methods by reframing the complaint detection problem as a multimodal text-to-text generation task. Significantly, our research includes the development of a novel aspect-level dataset. Annotated for both complaint and sentiment categories across diverse domains such as books, electronics, edibles, fashion, and miscellaneous, this dataset provides a comprehensive platform for the concurrent study of complaints and sentiment. This resource facilitates a more robust understanding of consumer feedback. Our proposed methodology establishes a benchmark performance in the novel aspect-based complaint and sentiment detection tasks based on extensive evaluation. We also demonstrate that our model consistently outperforms all other baselines and state-of-the-art models in both full and few-shot settings (The dataset and code are available at:https://github.com/appy1608/CIKM2023).
Apoorva Singh, Apoorv Verma, Raghav Jain, Sriparna Saha 0001
CIKM3
2023 T-VAKS: A Tutoring-Based Multimodal Dialog System via Knowledge Selection
abstract
Advancements in Conversational Natural Language Processing (NLP) have the potential to address critical social challenges, particularly in achieving the United Nations’ Sustainable Development Goal of quality education. However, the application of NLP in the educational domain, especially language learning, has been limited due to the inherent complexities of the field and the scarcity of available datasets. In this paper, we introduce T-VAKS (Tutoring Virtual Agent with Knowledge Selection), a novel language tutoring multimodal Virtual Agent (VA) designed to assist students in learning a new language, thereby promoting AI for Social Good. T-VAKS aims to bridge the gap between NLP and the educational domain, enabling more effective language tutoring through intelligent virtual agents. Our approach employs an information theory-based knowledge selection module built on top of a multimodal seq2seq generative model, facilitating the generation of appropriate, informative, and contextually relevant tutor responses. The knowledge selection module in turn consists of two sub-modules: (i) knowledge relevance estimation, and (ii) knowledge focusing framework. We evaluate the performance of our proposed end-to-end dialog system against various baseline models and the most recent state-of-the-art models, using multiple evaluation metrics. The results demonstrate that T-VAKS outperforms competing models, highlighting the potential of our approach in enhancing language learning through the use of conversational NLP and virtual agents, ultimately contributing to addressing social challenges and promoting well-being.
Raghav Jain, Tulika Saha, Sriparna Saha 0001
ECAI1
2023 Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language Models
abstract
Temporal reasoning represents a vital component of human communication and understanding, yet remains an underexplored area within the context of Large Language Models (LLMs).Despite LLMs demonstrating significant proficiency in a range of tasks, a comprehensive, large-scale analysis of their temporal reasoning capabilities is missing.Our paper addresses this gap, presenting the first extensive benchmarking of LLMs on temporal reasoning tasks.We critically evaluate 8 different LLMs across 6 datasets using 3 distinct prompting strategies.Additionally, we broaden the scope of our evaluation by including in our analysis 2 Code Generation LMs.Beyond broad benchmarking of models and prompts, we also conduct a finegrained investigation of performance across different categories of temporal tasks.We further analyze the LLMs on varying temporal aspects, offering insights into their proficiency in understanding and predicting the continuity, sequence, and progression of events over time.Our findings reveal a nuanced depiction of the capabilities and limitations of the models within temporal reasoning, offering a comprehensive reference for future research in this pivotal domain.
Raghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha 0001, Adam Jatowt, Sandipan Dandapat
EMNLP1
2023 GenEx: A Commonsense-aware Unified Generative Framework for Explainable Cyberbullying Detection
abstract
With the rise of social media and online communication, the issue of cyberbullying has gained significant prominence.While extensive research is being conducted to develop more effective models for detecting cyberbullying in monolingual languages, a significant gap exists in understanding code-mixed languages and the need for explainability in this context.To address this gap, we have introduced a novel benchmark dataset named Bully-Explain for explainable cyberbullying detection in code-mixed language.In this dataset, each post is meticulously annotated with four labels: bully, sentiment, target, and rationales, indicating the specific phrases responsible for identifying the post as a bully.Our current research presents an innovative unified generative framework, GenEx, which reimagines the multitask problem as a text-to-text generation task.Our proposed approach demonstrates its superiority across various evaluation metrics when applied to the BullyExplain dataset, surpassing other baseline models and current state-of-the-art approaches. 1
Krishanu Maity, Raghav Jain, Prince Jha, Sriparna Saha 0001, Pushpak Bhattacharyya
EMNLP2
2023 "Explain Thyself Bully": Sentiment Aided Cyberbullying Detection with Explanation
Krishanu Maity, Prince Jha, Raghav Jain, Sriparna Saha 0001, Pushpak Bhattacharyya
ICDAR (3)3
2023 Reimagining Complaint Analysis: Adopting Seq2Path for a Generative Text-to-Text Framework
abstract
Apoorva Singh, Raghav Jain, Sriparna Saha. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Apoorva Singh, Raghav Jain, Sriparna Saha 0001
IJCNLP (1)2
2023 Generative Models vs Discriminative Models: Which Performs Better in Detecting Cyberbullying in Memes?
abstract
Accessibility to the internet has led to a massive increase in the usage of social networking and online communication apps over the past decade. With so much content available online, these platforms suffer from governance and monitoring problems making the users susceptible to online cyberbullying and trolling. Recently, this trolling and hate comes majorly from memes that combine text and image modalities. Many studies show how cyberbullying can harm the mental well-being of the affected individuals. Previous studies also tried to study the role of sentiment, emotions, and sarcasm in identifying hateful memes setting up the meme detection task as a multi-modal, multitask problem. In the past, meme detection models were discriminative, but more recently generative models have been used to solve non-generation tasks such as aspect-based sentiment analysis and span detection. Motivated by this, in this work, we propose a unified Multimodal Generative framework, MGex by reframing the multitasking problem of detecting cyberbullying, sentiment, emotion, and sarcasm as a multimodal text-to-text generation problem. For this purpose, we use MultiBully dataset which provides annotation for all these labels. Here, we evaluate and contrast our proposed generative framework with several multitasking baselines and state-of-the-art models. Disclaimer: The article contains offensive text and profanity. This is owing to the nature of the work and does not reflect any opinion or stand of the authors.
Raghav Jain, Krishanu Maity, Prince Jha, Sriparna Saha 0001
IJCNN1
2023 AbCoRD: Exploiting multimodal generative approach for Aspect-based Complaint and Rationale Detection
abstract
Valuable feedback can be found in customer reviews about a product or service, but it can be difficult to identify specific complaints and their underlying causes from the text. While various methods have been used to detect and analyze complaints, there has been little research on examining complaints at the aspect level and the reasons behind these complaints. To address this, we added rationale annotation for aspect-based complaint classes to a publicly available benchmark multimodal complaint dataset (CESAMARD) covering five domains (books, electronics, edibles, fashion, and miscellaneous). Current methods treat these tasks as classification problems and do not use external knowledge. Our study proposes a knowledge-infused, hierarchical, multimodal generative approach for aspect-based complaint and rationale detection that reframes the multitasking problem as a multimodal text-to-text generation task. Our approach achieved benchmark performance in the aspect-based complaint and rationale detection task through an extensive evaluation. We demonstrated that our model consistently outperformed all other baselines and state-of-the-art models in both full and few-shot settings. Our study contributes to the development of more accurate and efficient methods for extracting valuable insights from customer reviews to improve products and service 1
Raghav Jain, Apoorva Singh, Vivek Kumar Gangwar, Sriparna Saha 0001
ACM Multimedia1
2023 Let the Model Make Financial Senses: A Text2Text Generative Approach for Financial Complaint Identification
Sarmistha Das 0001, Apoorva Singh, Raghav Jain, Sriparna Saha 0001, Alka Maurya
PAKDD (3)3
2022 WIDAR - Weighted Input Document Augmented ROUGE
Raghav Jain, Vaibhav Mavi, Anubhav Jangra, Sriparna Saha 0001
ECIR (1)1
2022 Domain Infused Conversational Response Generation for Tutoring based Virtual Agent
abstract
Recent advances in deep learning typically, with the introduction of transformer based models has shown massive improvement and success in many Natural Language Processing (NLP) tasks. One such area which has leveraged immensely is conversational agents or chatbots in open-ended (chit-chat conversations) and task-specific (such as medical or legal dialogue bots etc.) domains. However, in the era of automation, there is still a dearth of works focused on one of the most relevant use cases, i.e., tutoring dialog systems that can help students learn new subjects or topics of their interest. Most of the previous works in this domain are either rule based systems which require a lot of manual efforts or are based on multiple choice type factual questions. In this paper, we propose EDICA (Educational Domain Infused Conversational Agent), a language tutoring Virtual Agent (VA). EDICA employs two mechanisms in order to converse fluently with a student/user over a question and assist them to learn a language: (i) Student/Tutor Intent Classification (SIC-TIC) framework to identify the intent of the student and decide the action of the VA, respectively, in the on-going conversation and (ii) Tutor Response Generation (TRG) framework to generate domain infused and intent/action conditioned tutor responses at every step of the conversation. The VA is able to provide hints, ask questions and correct student's reply by generating an appropriate, informative and relevant tutor response. We establish the superiority of our proposed approach on various evaluation metrics over other baselines and state of the art models.
Raghav Jain, Tulika Saha, Souhitya Chakraborty, Sriparna Saha 0001
IJCNN1
2020 BIRDSAI: A Dataset for Detection and Tracking in Aerial Thermal Infrared Videos
abstract
Monitoring of protected areas to curb illegal activities like poaching and animal trafficking is a monumental task. To augment existing manual patrolling efforts, unmanned aerial surveillance using visible and thermal infrared (TIR) cameras is increasingly being adopted. Automated data acquisition has become easier with advances in unmanned aerial vehicles (UAVs) and sensors like TIR cameras, which allow surveillance at night when poaching typically occurs. However, it is still a challenge to accurately and quickly process large amounts of the resulting TIR data. In this paper, we present the first large dataset collected using a TIR camera mounted on a fixed-wing UAV in multiple African protected areas. This dataset includes TIR videos of humans and animals with several challenging scenarios like scale variations, background clutter due to thermal reflections, large camera rotations, and motion blur. Additionally, we provide another dataset with videos synthetically generated with the publicly available Microsoft AirSim simulation platform using a 3D model of an African savanna and a TIR camera model. Through our benchmarking experiments on state-of-the-art detectors, we demonstrate that leveraging the synthetic data in a domain adaptive setting can significantly improve detection performance. We also evaluate various recent approaches for single and multi-object tracking. With the increasing popularity of aerial imagery for monitoring and surveillance purposes, we anticipate this unique dataset to be used to develop and evaluate techniques for object detection, tracking, and domain adaptation for aerial, TIR videos.
Elizabeth Bondi-Kelly, Raghav Jain, Palash Aggrawal, Saket Anand, Robert Hannaford, Ashish Kapoor, James Piavis, Shital Shah, Lucas Joppa, Bistra Dilkina, Milind Tambe
WACV2