EDBT 2026 Demo / reviewers in the wild / expert
Raghvendra Kumar 0003
dblp:340/4018
· DBLP profile ↗
13ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0002-9488-3099ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and ResourcesabstractIndia's linguistic landscape, spanning 22 scheduled languages and hundreds of marginalized dialects, has driven rapid growth in NLP datasets, benchmarks, and pretrained models.However, no dedicated survey consolidates resources developed specifically for Indian languages.Existing reviews either focus on a few high-resource languages or subsume Indian languages within broader multilingual settings, limiting coverage of low-resource and culturally diverse varieties.To address this gap, we present the first unified survey of Indian NLP resources, covering 200+ datasets, 50+ benchmarks, and 100+ models, tools, and systems across text, speech, multimodal, and culturally grounded tasks.We organize resources by linguistic phenomena, domains, and modalities; analyze trends in annotation, evaluation, and model design; and identify persistent challenges such as data sparsity, uneven language coverage, script diversity, and limited cultural and domain generalization.This survey offers a consolidated foundation for equitable, culturally grounded, and scalable NLP research in the Indian linguistic ecosystem. Raghvendra Kumar 0003, Devankar Raj, Sriparna Saha 0001 |
ACL (1) | 1 |
| 2026 | Small Models, Big Picture! A Language Model Augmentation for Enhanced Reader-Aware Summarization
Raghvendra Kumar 0003, A. S. Poornash, Sriparna Saha 0001 |
ECIR (1) | 1 |
| 2026 | From Comments to Conclusions: Adaptive Reader-Aware Summary Generation in Low-Resource Languages via Agent Debate
Raghvendra Kumar 0003, S. A. Mohammed Salman, Jaya Verma, Sriparna Saha 0001 |
ECIR (1) | 1 |
| 2026 | Sifting Truth From Spectacle! A Multimodal Hindi Dataset for Misinformation Detection With Emotional Cues and SentimentsabstractMisinformation poses a growing threat across media ecosystems, yet research on Hindi, one of the world's most widely spoken languages, remains limited. We introduce a novel multimodal Hindi dataset of 6,544 article–image pairs to advance misinformation detection. Unlike existing English-centric and predominantly unimodal datasets, ours integrates text, images, and affective signals while being carefully cleaned of veracity cues to avoid artefact-driven inflation. Each sample is annotated with sentiment and emotions, making this the first Hindi resource with multimodal and affective dimensions. Through extensive experiments using IndicBART, IndicBERT, mBERT, and Vision Transformer models, we demonstrate the effectiveness of text–image fusion and affective features across multiple configurations. We also analyze the readability characteristics of genuine and misleading articles, providing insights into the linguistic patterns of Hindi misinformation. This dataset establishes a robust benchmark for multimodal misinformation detection and lays essential groundwork for research in Hindi and other low-resource languages. Raghvendra Kumar 0003, Pulkit Bansal, Raunak Kumar Singh, Sriparna Saha 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | COSMMIC: Comment-Sensitive Multimodal Multilingual Indian Corpus for Summarization and Headline GenerationabstractDespite progress in comment-aware multimodal and multilingual summarization for English and Chinese, research in Indian languages remains limited. This study addresses this gap by introducing COSMMIC, a pioneering comment-sensitive multimodal, multilingual dataset featuring nine major Indian languages. COSMMIC comprises 4,959 article-image pairs and 24,484 reader comments, with ground-truth summaries available in all included languages. Our approach enhances summaries by integrating reader insights and feedback. We explore summarization and headline generation across four configurations: (1) using article text alone, (2) incorporating user comments, (3) utilizing images, and (4) combining text, comments, and images. To assess the dataset’s effectiveness, we employ state-of-the-art language models such as LLama3 and GPT-4. We conduct a comprehensive study to evaluate different component combinations, including identifying supportive comments, filtering out noise using a dedicated comment classifier using IndicBERT, and extracting valuable insights from images with a multilingual CLIP-based classifier. This helps determine the most effective configurations for natural language generation (NLG) tasks. Unlike many existing datasets that are either text-only or lack user comments in multimodal settings, COSMMIC uniquely integrates text, images, and user feedback. This holistic approach bridges gaps in Indian language resources, advancing NLP research and fostering inclusivity. Raghvendra Kumar 0003, Mohammed Salman S. A, Aryan Sahu, Tridib Nandi, Pragathi Y. P., Sriparna Saha 0001, José G. Moreno 0001 |
ACL (1) | 1 |
| 2025 | Poetry in Pixels: Prompt Tuning for Poem Image Generation via Diffusion ModelsabstractThe task of text-to-image generation has encountered significant challenges when applied to literary works, especially poetry. Poems are a distinct form of literature, with meanings that frequently transcend beyond the literal words. To address this shortcoming, we propose a PoemToPixel framework designed to generate images that visually represent the inherent meanings of poems. Our approach incorporates the concept of prompt tuning in our image generation framework to ensure that the resulting images closely align with the poetic content. In addition, we propose the PoeKey algorithm, which extracts three key elements in the form of emotions, visual elements, and themes from poems to form instructions which are subsequently provided to a diffusion model for generating corresponding images. Furthermore, to expand the diversity of the poetry dataset across different genres and ages, we introduce MiniPo, a novel multimodal dataset comprising 1001 children’s poems and images. Leveraging this dataset alongside PoemSum, we conducted both quantitative and qualitative evaluations of image generation using our PoemToPixel framework. This paper demonstrates the effectiveness of our approach and offers a fresh perspective on generating images from literary sources. The code and dataset used in this work are publicly available. Sofia Jamil, Bollampalli Areen Reddy, Raghvendra Kumar 0003, Sriparna Saha 0001, K. J. Joseph, Koustava Goswami |
COLING | 3 |
| 2025 | PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt RefinementabstractRecent advancements in text-to-image diffusion models have achieved remarkable success in generating realistic and diverse visual content. A critical factor in this process is the model’s ability to accurately interpret textual prompts. However, these models often struggle with creative expressions, particularly those involving complex, abstract, or highly descriptive language. In this work, we introduce a novel training-free approach tailored to improve image generation for a unique form of creative language: poetic verse, which frequently features layered, abstract, and dual meanings. Our proposed PoemTale Diffusion approach aims to minimise the information that is lost during poetic text-to-image conversion by integrating a multi stage prompt refinement loop into Language Models to enhance the interpretability of poetic texts. To support this, we adapt existing state-of-the-art diffusion models by modifying their self-attention mechanisms with a consistent self-attention technique to generate multiple consistent images, which are then collectively used to convey the poem’s meaning. Moreover, to encourage research in the field of poetry, we introduce the P4I (PoemForImage) dataset, consisting of 1,111 poems sourced from multiple online and offline resources. We engaged a panel of poetry experts for qualitative assessments. The results from both human and quantitative evaluations validate the efficacy of our method and contribute a novel perspective to poem-to-image generation with enhanced information capture in the generated images. Sofia Jamil, Bollampalli Areen Reddy, Raghvendra Kumar 0003, Sriparna Saha 0001, Koustava Goswami |
ECAI | 3 |
| 2025 | DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian CultureabstractArijit Maji, Raghvendra Kumar, Akash Ghosh, Anushka, Nemil Shah, Abhilekh Borah, Vanshika Shah, Nishant Mishra, Sriparna Saha. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Arijit Maji, Raghvendra Kumar 0003, Akash Ghosh, Anushka, Nemil Shah, Abhilekh Borah, Vanshika Shah, Sriparna Saha 0001 |
EMNLP | 2 |
| 2024 | IndicBART Alongside Visual Element: Multimodal Summarization in Diverse Indian Languages
Raghvendra Kumar 0003, Deepak Prakash, Sriparna Saha 0001 |
ICDAR (6) | 1 |
| 2024 | Extracting the Full Story: A Multimodal Approach and Dataset to Crisis Summarization in TweetsabstractIn our digitally connected world, the influx of microblog data poses a formidable challenge in extracting relevant information amid a continuous stream of updates. This challenge intensifies during crises, where the demand for timely and relevant information is crucial. Current summarization techniques often struggle with the intricacies of microblog data in such situations. To address this, our research explores crisis-related microblogs, recognizing the crucial role of multimedia content, such as images, in offering a comprehensive perspective. In response to these challenges, we introduce a multimodal extractive-abstractive summarization model. Leveraging a fusion of TF-IDF scoring and bigram filtering, coupled with the effectiveness of three distinct models—BIGBIRD, CLIP, and bootstrapping language-image pre-training (BLIP)—we aim to overcome the limitations of traditional extractive and text-only approaches. Our model is designed and evaluated on a newly curated Twitter dataset featuring 12 494 tweets and 3090 images across eight crisis events, each accompanied by gold-standard summaries. The experimental findings showcase the remarkable efficacy of our model, surpassing current benchmarks by a notable margin of 16% and 17%. This confirms our model's strength and its relevance in crisis scenarios with the crucial interplay of text and multimedia. Notably, our research contributes to multimodal, abstractive microblog summarization, addressing a key gap in the literature. It is also a valuable tool for swift information extraction in time-sensitive situations. Raghvendra Kumar 0003, Ritika Sinha, Sriparna Saha 0001, Adam Jatowt |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | Diving into a Sea of Opinions: Multi-modal Abstractive Summarization with Comment SensitivityabstractIn the modern era, the rapid expansion of social media and the proliferation of the internet community has led to a multi-fold increase in the richness and range of views and outlooks expressed by readers and viewers. To obtain valuable insights from this vast sea of opinions, we present an inventive and holistic procedure for multi-modal abstractive summarization with comment sensitivity. Our proposed model utilizes both textual and visual modalities and examines the remarks provided by the readers to produce summaries that apprehend the significant points and opinions made by them. Our model features a transformer-based encoder that seamlessly processes both news articles and comments, merging them before transmitting the amalgamated information to the decoder. Additionally, the core segment of our architecture consists of an attention-based merging technique which is trained adversarially by means of a generator and discriminator to bridge the semantic gap between comments and articles. We have used a Bi-LSTM-based branch for image pointer generation. We assess our model on the reader-aware multi-document summarization (RA-MDS) dataset which contains news articles, their summaries, and related comments. We have extended the dataset by adding images pertaining to news articles in the corpus to increase the richness and diversity of the dataset. Our comprehensive experiments reveal that our model outperforms similar pre-trained models and baselines across two of the four evaluated metrics, showcasing its superior performance. Raghvendra Kumar 0003, Ratul Chakraborty, Sriparna Saha 0001, Naveen Saini |
CIKM | 1 |
| 2023 | Multimodal Rumour Detection: Catching News that Never Transpired!
Raghvendra Kumar 0003, Ritika Sinha, Sriparna Saha 0001, Adam Jatowt |
ICDAR (3) | 1 |
| 2023 | Can Multimodal Pointer Generator Transformers Produce Topically Relevant Summaries?abstractDue to the growth in demand for brief and pertinent multimedia material over the past few years, multimodal summarization has attracted a lot of study interest. Recently Transformers have been widely used for various sequence processing tasks due to their fast parallel processing ability compared to LSTMs. Although Multimodal Summarization (MS) has tractioned much research interest of late, a research gap exists in producing topic-relevant multimodal summaries. Since any summary deals with concise information, it should carry the essence of the topic from which it was derived. Further, due to the lack of alignment information among the images and the inter-modal segments, MS systems also face difficulty choosing appropriate pictorial summaries. To study these research questions, we propose a Multitask learning-based Multimodal Pointer Generator Transformer (MPGT), which utilizes the topic information of the samples to produce multimodal summaries. We also augment the popular MSMO dataset for this study with similar “On-Topic” and “Off-Topic” images. Our results show that inter-modal attention among images helps achieve better alignment in the visual modality and improves image precision scores. Our analysis also provides discussions on how we can further enhance topic-relevant MS systems. Sourajit Mukherjee, Adam Jatowt, Raghvendra Kumar 0003, Anubhav Jangra, Sriparna Saha 0001 |
IJCNN | 3 |