Koustava Goswami

dblp:266/1147 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0002-0428-160XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Crossing Borders: A Multimodal Challenge for Indian Poetry Translation and Image Generation
abstract
Indian poetry, known for its linguistic complexity and deep cultural resonance, has a rich and varied heritage spanning thousands of years. However, its layered meanings, cultural allusions, and sophisticated grammatical constructions often pose challenges for comprehension, especially for non-native speakers or readers unfamiliar with its context and language. Despite its cultural significance, existing works on poetry have largely overlooked Indian language poems. In this paper, we propose the Translation and Image Generation (TAI) framework, leveraging Large Language Models (LLMs) and Latent Diffusion Models through appropriate prompt tuning. Our framework supports the United Nations Sustainable Development Goals of Quality Education (SDG 4) and Reduced Inequalities (SDG 10), by enhancing the accessibility of culturally rich Indian-language poetry to a global audience. It includes (1) a translation module that uses an Odds Ratio Preference Alignment Algorithm to accurately translate morphologically rich poetry into English; (2) an image generation module that employs a semantic graph to capture tokens, dependencies, and semantic relationships between metaphors and their meanings, to create visually meaningful representations of Indian poems. Our comprehensive experimental evaluation, including both human and quantitative assessments, demonstrates the superiority of TAI Diffusion in poem image generation tasks, outperforming strong baselines. To further address the scarcity of resources for Indian-language poetry, we introduce the Morphologically Rich Indian Language Poems MorphoVerse Dataset, comprising 1,570 poems across 21 low-resource Indian languages. By addressing the gap in poetry translation and visual comprehension, this work aims to broaden accessibility and enrich the reader’s experience.
Sofia Jamil, Kotla Sai Charan, Sriparna Saha 0001, Koustava Goswami, K. J. Joseph
AAAI4
2026 Step-by-step Layered Design Generation
Faizan Farooq Khan, K. J. Joseph, Koustava Goswami, Mohamed Elhoseiny 0001, Balaji Vasan Srinivasan
AAAI3
2026 Agentic Design Review System
Sayan Nag, K. J. Joseph, Koustava Goswami, Vlad I. Morariu, Balaji Vasan Srinivasan
AAAI3
2025 Poetry in Pixels: Prompt Tuning for Poem Image Generation via Diffusion Models
abstract
The task of text-to-image generation has encountered significant challenges when applied to literary works, especially poetry. Poems are a distinct form of literature, with meanings that frequently transcend beyond the literal words. To address this shortcoming, we propose a PoemToPixel framework designed to generate images that visually represent the inherent meanings of poems. Our approach incorporates the concept of prompt tuning in our image generation framework to ensure that the resulting images closely align with the poetic content. In addition, we propose the PoeKey algorithm, which extracts three key elements in the form of emotions, visual elements, and themes from poems to form instructions which are subsequently provided to a diffusion model for generating corresponding images. Furthermore, to expand the diversity of the poetry dataset across different genres and ages, we introduce MiniPo, a novel multimodal dataset comprising 1001 children’s poems and images. Leveraging this dataset alongside PoemSum, we conducted both quantitative and qualitative evaluations of image generation using our PoemToPixel framework. This paper demonstrates the effectiveness of our approach and offers a fresh perspective on generating images from literary sources. The code and dataset used in this work are publicly available.
Sofia Jamil, Bollampalli Areen Reddy, Raghvendra Kumar 0003, Sriparna Saha 0001, K. J. Joseph, Koustava Goswami
COLING6
2025 PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement
abstract
Recent advancements in text-to-image diffusion models have achieved remarkable success in generating realistic and diverse visual content. A critical factor in this process is the model’s ability to accurately interpret textual prompts. However, these models often struggle with creative expressions, particularly those involving complex, abstract, or highly descriptive language. In this work, we introduce a novel training-free approach tailored to improve image generation for a unique form of creative language: poetic verse, which frequently features layered, abstract, and dual meanings. Our proposed PoemTale Diffusion approach aims to minimise the information that is lost during poetic text-to-image conversion by integrating a multi stage prompt refinement loop into Language Models to enhance the interpretability of poetic texts. To support this, we adapt existing state-of-the-art diffusion models by modifying their self-attention mechanisms with a consistent self-attention technique to generate multiple consistent images, which are then collectively used to convey the poem’s meaning. Moreover, to encourage research in the field of poetry, we introduce the P4I (PoemForImage) dataset, consisting of 1,111 poems sourced from multiple online and offline resources. We engaged a panel of poetry experts for qualitative assessments. The results from both human and quantitative evaluations validate the efficacy of our method and contribute a novel perspective to poem-to-image generation with enhanced information capture in the generated images.
Sofia Jamil, Bollampalli Areen Reddy, Raghvendra Kumar 0003, Sriparna Saha 0001, Koustava Goswami
ECAI5
2025 Do It Yourself (DIY): Modifying Images for Poems in a Zero-Shot Setting Using Weighted Prompt Manipulation
abstract
Poetry is an expressive form of art that invites multiple interpretations, as readers often bring their own emotions, experiences, and cultural backgrounds into their understanding of a poem.Recognizing this, we aim to generate images for poems and improve these images in a zero-shot setting, enabling audiences to modify images as per their requirements.To achieve this, we introduce a novel Weighted Prompt Manipulation (WPM) technique, which systematically modifies attention weights and text embeddings within diffusion models.By dynamically adjusting the importance of specific words, WPM enhances or suppresses their influence in the final generated image, leading to semantically richer and more contextually accurate visualizations.Our approach exploits diffusion models and large language models (LLMs) such as GPT in conjunction with existing poetry datasets, ensuring a comprehensive and structured methodology for improved image generation in the literary domain.To the best of our knowledge, this is the first attempt at integrating weighted prompt manipulation for enhancing imagery in poetic language.Resources related to data and codes are available here: DIY
Sofia Jamil, Kotla Sai Charan, Sriparna Saha 0001, Koustava Goswami, K. J. Joseph
EMNLP4
2025 Subjective Behaviors and Preferences in LLM: Language of Browsing
abstract
Sai Sundaresan, Harshita Chopra, Atanu R. Sinha, Koustava Goswami, Nagasai Saketh Naidu, Raghav Karan, N Anushka. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Sai Sundaresan, Harshita Chopra, Atanu R. Sinha, Koustava Goswami, Nagasai Saketh Naidu, Raghav Karan, N. Anushka
EMNLP4
2025 Reasoning and Planning for Multimodal Large Language Models: A Multilingual and Cross-Domain Exploration
abstract
Recent advancements in Multimodal Large Language Models (MLLMs), coupled with the progress of reinforcement learning, have substantially enhanced reasoning and decision-making across modalities, including text, vision, audio, and video. This tutorial introduces the fundamental principles, methodologies, and practical applications of MLLM reasoning, with a particular emphasis on strengthening reasoning capabilities in multilingual and cross-domain settings. We further discuss the key challenges and limitations of current multimodal reasoning approaches, as well as future directions for advancing the field. By highlighting how MLLMs support enhanced reasoning and planning in cross-lingual and cross-domain contexts, this session aims to equip researchers and practitioners with the conceptual foundations and practical tools needed to effectively integrate MLLM reasoning into their work.
Sarmistha Das 0001, Akash Ghosh, Sriparna Saha 0001, Koustava Goswami, K. J. Joseph
ACM Multimedia4
2025 Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs
abstract
Anirudh Phukan, Divyansh Divyansh, Harshit Kumar Morj, Vaishnavi Vaishnavi, Apoorv Saxena, Koustava Goswami. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Anirudh Phukan, Divyansh, Harshit Kumar Morj, Vaishnavi, Apoorv Saxena, Koustava Goswami
NAACL (Long Papers)6
2024 CoPL: Contextual Prompt Learning for Vision-Language Understanding
abstract
Recent advances in multimodal learning has resulted in powerful vision-language models, whose representations are generalizable across a variety of downstream tasks. Recently, their generalization ability has been further extended by incorporating trainable prompts, borrowed from the natural language processing literature. While such prompt learning techniques have shown impressive results, we identify that these prompts are trained based on global image features which limits itself in two aspects: First, by using global features, these prompts could be focusing less on the discriminative foreground image, resulting in poor generalization to various out-of-distribution test cases. Second, existing work weights all prompts equally whereas intuitively, prompts should be reweighed according to the semantics of the image. We address these as part of our proposed Contextual Prompt Learning (CoPL) framework, capable of aligning the prompts to the localized features of the image. Our key innovations over earlier works include using local image features as part of the prompt learning process, and more crucially, learning to weight these prompts based on local features that are appropriate for the task at hand. This gives us dynamic prompts that are both aligned to local image features as well as aware of local contextual relationships. Our extensive set of experiments on a variety of standard and few-shot datasets show that our method produces substantially improved performance when compared to the current state of the art methods. We also demonstrate both few-shot and out-of-distribution performance to establish the utility of learning dynamic prompts that are aligned to local image features.
Koustava Goswami, Srikrishna Karanam, Prateksha Udhayanan, K. J. Joseph, Balaji Vasan Srinivasan
AAAI1
2024 SafaRi: Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation
Sayan Nag, Koustava Goswami, Srikrishna Karanam
ECCV (44)2
2024 Enhancing Post-Hoc Attributions in Long Document Comprehension via Coarse Grained Answer Decomposition
abstract
Accurately attributing answer text to its source document is crucial for developing a reliable question-answering system.However, attribution for long documents remains largely unexplored.Post-hoc attribution systems are designed to map answer text back to the source document, yet the granularity of this mapping has not been addressed.Furthermore, a critical question arises: What exactly should be attributed?This involves identifying the specific information units within an answer that require grounding.In this paper, we propose and investigate a novel approach to the factual decomposition of generated answers for attribution, employing template-based in-context learning.To accomplish this, we utilize the question and integrate negative sampling during few-shot in-context learning for decomposition.This approach enhances the semantic understanding of both abstractive and extractive answers.We examine the impact of answer decomposition by providing a thorough examination of various attribution approaches, ranging from retrieval-based techniques to LLM-based attributors.
Pritika Ramu, Koustava Goswami, Apoorv Saxena, Balaji Vasan Srinivasan
EMNLP2
2024 An Audit on the Perspectives and Challenges of Hallucinations in NLP
abstract
Pranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs, Mukund Srinath, Koustava Goswami, Sarah Rajtmajer, Shomir Wilson. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Pranav Venkit, Tatiana Chakravorti, Heidi R. Biggs, Mukund Srinath, Koustava Goswami, Sarah Michele Rajtmajer, Shomir Wilson
EMNLP6
2024 Iterative Multi-granular Image Editing using Diffusion Models
abstract
Recent advances in text-guided image synthesis has dramatically changed how creative professionals generate artistic and aesthetically pleasing visual assets. To fully support such creative endeavors, the process should possess the ability to: 1) iteratively edit the generations and 2) control the spatial reach of desired changes (global, local or anything in between). We formalize this pragmatic problem setting as Iterative Multi-granular Editing. While there has been substantial progress with diffusion-based models for image synthesis and editing, they are all one shot (i.e., no iterative editing capabilities) and do not naturally yield multi-granular control (i.e., covering the full spectrum of local-to-global edits). To overcome these drawbacks, we propose EMILIE: Iterative Multi-granular Image Editor. EMILIE introduces a novel latent iteration strategy, which re-purposes a pre-trained diffusion model to facilitate iterative editing. This is complemented by a gradient control operation for multi-granular control. We introduce a new benchmark dataset to evaluate our newly proposed setting. We conduct exhaustive quantitatively and qualitatively evaluation against recent state-of-the-art approaches adapted to our task, to being out the mettle of EMILIE. We hope our work would attract attention to this newly identified, pragmatic problem setting.
K. J. Joseph, Prateksha Udhayanan, Tripti Shukla, Aishwarya Agarwal, Srikrishna Karanam, Koustava Goswami, Balaji Vasan Srinivasan
WACV6
2023 SwitchPrompt: Learning Domain-Specific Gated Soft Prompts for Classification in Low-Resource Domains
abstract
Prompting pre-trained language models leads to promising results across natural language processing tasks but is less effective when applied in low-resource domains, due to the domain gap between the pre-training data and the downstream task.In this work, we bridge this gap with a novel and lightweight prompting methodology called SwitchPrompt for the adaptation of language models trained on datasets from the general domain to diverse low-resource domains.Using domain-specific keywords with a trainable gated prompt, Switch-Prompt offers domain-oriented prompting, that is, effective guidance on the target domains for general-domain language models.Our fewshot experiments on three text classification benchmarks demonstrate the efficacy of the general-domain pre-trained language models when used with SwitchPrompt.They often even outperform their domain-specific counterparts trained with baseline state-of-the-art prompting methods by up to 10.7% performance increase in accuracy.This result indicates that SwitchPrompt effectively reduces the need for domain-specific language model pre-training.
Koustava Goswami, Lukas Lange, Jun Araki, Heike Adel
EACL1
2023 A-STAR: Test-time Attention Segregation and Retention for Text-to-image Synthesis
abstract
While recent developments in text-to-image generative models have led to a suite of high-performing methods capable of producing creative imagery from free-form text, there are several limitations. By analyzing the cross-attention representations of these models, we notice two key issues. First, for text prompts that contain multiple concepts, there is a significant amount of pixel-space overlap (i.e., same spatial regions) among pairs of different concepts This eventually leads to the model being unable to distinguish between the two concepts and one of them being ignored in the final generation. Next, while these models attempt to capture all such concepts during the beginning of denoising (e.g., first few steps) as evidenced by cross-attention maps, this knowledge is not retained by the end of denoising (e.g., last few steps). Such loss of knowledge eventually leads to inaccurate generation outputs.To address these issues, our key innovations include two test-time attention-based loss functions that substantially improve the performance of pretrained baseline text-to-image diffusion models. First, our attention segregation loss reduces the cross-attention overlap between attention maps of different concepts in the text prompt, thereby reducing the confusion/conflict among various concepts and the eventual capture of all concepts in the generated output. Next, our attention retention loss explicitly forces text-to-image diffusion models to retain cross-attention information for all concepts across all denoising time steps, thereby leading to reduced information loss and the preservation of all concepts in the generated output. We conduct extensive experiments with the proposed loss functions on a variety of text prompts and demonstrate they lead to generated images that are significantly semantically closer to the input text when compared to baseline text-to-image diffusion models.
Aishwarya Agarwal, Srikrishna Karanam, K. J. Joseph, Apoorv Saxena, Koustava Goswami, Balaji Vasan Srinivasan
ICCV5
2023 PICKD: In-Situ Prompt Tuning for Knowledge-Grounded Dialogue Generation
Rajdeep Sarkar, Koustava Goswami, Mihael Arcan, John P. McCrae
PAKDD (4)2
2021 Mufin: Enriching Semantic Understanding of Sentence Embedding using Dual Tune Framework
abstract
With the advancements of Natural Language Understanding (NLU), diverse industrial applications like user intent classification, smart chatbots, sentiment analysis and question answering have be-come a primary paradigm. Transformers-based multi-lingual language models such as XLM have performed significantly well in diverse semantic understanding and classification tasks. However, fine-tuning such large pre-trained architectures is resource and compute intensive, limiting its wide adoption in enterprise environments.We present a novel efficient and light-weight frame-work based on sentence embeddings to obtain enhanced multi-lingual text representations for domain-specific NLU applications. Our framework combines the concepts of up-projection, alignment and meta-embeddings enhancing the textual semantic similarity knowledge of smaller sentence embedding architectures. Extensive experiments on diverse cross-lingual classification tasks showcase the proposed framework to be comparable to state-of-the-art large language models (in mono-lingual and zero-shot settings), even with lesser training and resource requirements.
Koustava Goswami, Sourav Dutta 0001, Haytham Assem
IEEE BigData1
2021 Cross-lingual Sentence Embedding using Multi-Task Learning
abstract
Multilingual sentence embeddings capture rich semantic information not only for measuring similarity between texts but also for catering to a broad range of downstream crosslingual NLP tasks.State-of-the-art multilingual sentence embedding models require large parallel corpora to learn efficiently, which confines the scope of these models.In this paper, we propose a novel sentence embedding framework based on an unsupervised loss function for generating effective multilingual sentence embeddings, eliminating the need for parallel corpora.We capture semantic similarity and relatedness between sentences using a multitask loss function for training a dual encoder model mapping different languages onto the same vector space.We demonstrate the efficacy of an unsupervised as well as a weakly supervised variant of our framework on STS, BUCC and Tatoeba benchmark tasks.The proposed unsupervised sentence embedding framework outperforms even supervised stateof-the-art methods for certain under-resourced languages on the Tatoeba dataset and on a monolingual benchmark.Further, we show enhanced zero-shot learning capabilities for more than 30 languages, with the model being trained on only 13 languages.Our model can be extended to a wide range of languages from any language family, as it overcomes the requirement of parallel corpora for training.
Koustava Goswami, Sourav Dutta 0001, Haytham Assem, Theodorus Fransen, John P. McCrae
EMNLP (1)1
2020 Unsupervised Deep Language and Dialect Identification for Short Texts
abstract
Automatic Language Identification (LI) or Dialect Identification (DI) of short texts of closely related languages or dialects, is one of the primary steps in many natural language processing pipelines.Language identification is considered a solved task in many cases; however, in the case of very closely related languages, or in an unsupervised scenario (where the languages are not known in advance), performance is still poor.In this paper, we propose the Unsupervised Deep Language and Dialect Identification (UDLDI) method, which can simultaneously learn sentence embeddings and cluster assignments from short texts.The UDLDI model understands the sentence constructions of languages by applying attention to character relations which helps to optimize the clustering of languages.We have performed our experiments on three shorttext datasets for different language families, each consisting of closely related languages or dialects, with very minimal training sets.Our experimental evaluations on these datasets have shown significant improvement over state-of-the-art unsupervised methods and our model has outperformed state-of-the-art LI and DI systems in supervised settings.
Koustava Goswami, Rajdeep Sarkar, Bharathi Raja Chakravarthi, Theodorus Fransen, John P. McCrae
COLING1
2020 Suggest me a movie for tonight: Leveraging Knowledge Graphs for Conversational Recommendation
abstract
Conversational recommender systems focus on the task of suggesting products to users based on the conversation flow.Recently, the use of external knowledge in the form of knowledge graphs has shown to improve the performance in recommendation and dialogue systems.Information from knowledge graphs aids in enriching those systems by providing additional information such as closely related products and textual descriptions of the items.However, knowledge graphs are incomplete since they do not contain all factual information present on the web.Furthermore, when working on a specific domain, knowledge graphs in its entirety contribute towards extraneous information and noise.In this work, we study several subgraph construction methods and compare their performance across the recommendation task.We incorporate pre-trained embeddings from the subgraphs along with positional embeddings in our models.Extensive experiments show that our method has a relative improvement of at least 5.62% compared to the state-of-the-art on multiple metrics on the recommendation task.
Rajdeep Sarkar, Koustava Goswami, Mihael Arcan, John P. McCrae
COLING2