EDBT 2026 Demo / reviewers in the wild / expert
Ramaneswaran S.
dblp:331/4189 · also Ramaneswaran Selvakumar
· DBLP profile ↗
14ranked-venue papers
1as first author
14since 2021 · last 2025
0000-0002-8340-403XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal InteractionsabstractRamaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi, Sonal Kumar, Sreyan Ghosh, Dinesh Manocha. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ramaneswaran S., Ashish Seth, Nishit Anand, Utkarsh Tyagi, Sonal Kumar, Sreyan Ghosh, Dinesh Manocha |
EMNLP | 1 |
| 2025 | EGOILLUSION: Benchmarking Hallucinations in Egocentric Video UnderstandingabstractAshish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand, Sonal Kumar, Sreyan Ghosh, Ramani Duraiswami, Chirag Agarwal, Dinesh Manocha. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ashish Seth, Utkarsh Tyagi, Ramaneswaran S., Nishit Anand, Sonal Kumar, Sreyan Ghosh, Ramani Duraiswami, Chirag Agarwal, Dinesh Manocha |
EMNLP | 3 |
| 2025 | MMAU: A Massive Multi-Task Audio Understanding and Reasoning BenchmarkabstractThe ability to comprehend audio—which includes speech, non-speech sounds, and music—is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding models on tasks requiring expert-level knowledge and complex reasoning. MMAU comprises 10k carefully curated audio clips paired with human-annotated natural language questions and answers spanning speech, environmental sounds, and music. It includes information extraction and reasoning questions, requiring models to demonstrate 27 distinct skills across unique and challenging tasks. Unlike existing benchmarks, MMAU emphasizes advanced perception and reasoning with domain-specific knowledge, challenging models to tackle tasks akin to those faced by experts. We assess 18 open-source and proprietary (Large) Audio-Language Models, demonstrating the significant challenges
posed by MMAU. Notably, even the most advanced Gemini 2.0 Flash achieves only 59.93% accuracy, and the state-of-the-art open-source Qwen2-Audio achieves only 52.50%, highlighting considerable room for improvement. We believe MMAU will drive the audio and multimodal research community to develop more advanced audio understanding models capable of solving complex audio tasks. S. Sakshi, Utkarsh Tyagi, Sonal Kumar, Ashish Seth, Ramaneswaran S., Oriol Nieto, Ramani Duraiswami, Sreyan Ghosh, Dinesh Manocha |
ICLR | 5 |
| 2025 | PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio ClassificationabstractAshish Seth, Ramaneswaran Selvakumar, Sonal Kumar, Sreyan Ghosh, Dinesh Manocha. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ashish Seth, Ramaneswaran S., Sonal Kumar, Sreyan Ghosh, Dinesh Manocha |
NAACL (Long Papers) | 2 |
| 2024 | ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract DescriptionsabstractSreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Chandra Kiran Evuru, Ramaneswaran S, S Sakshi, Dinesh Manocha. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Sreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Chandra Kiran Reddy Evuru, Ramaneswaran S., S. Sakshi, Dinesh Manocha |
ACL (1) | 5 |
| 2024 | EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation LearningabstractIn this paper, we present EH-MAM (Easyto-Hard adaptive Masked Acoustic Modeling), a novel self-supervised learning approach for speech representation learning.In contrast to the prior methods that use random masking schemes for Masked Acoustic Modeling (MAM), we introduce a novel selective and adaptive masking strategy.Specifically, during SSL training, we progressively introduce harder regions to the model for reconstruction.Our approach automatically selects hard regions and is built on the observation that the reconstruction loss of individual frames in MAM can provide natural signals to judge the difficulty of solving the MAM pre-text task for that frame.To identify these hard regions, we employ a teacher model that first predicts the frame-wise losses and then decides which frames to mask.By learning to create challenging problems, such as identifying harder frames and solving them simultaneously, the model is able to learn more effective representations and thereby acquire a more comprehensive understanding of the speech.Quantitatively, EH-MAM outperforms several state-ofthe-art baselines across various low-resource speech recognition and SUPERB benchmarks by 5%-10%.Additionally, we conduct a thorough analysis to show that the regions masked by EH-MAM effectively capture useful context across speech frames 1 . * Equal contribution, † Equal Ideation 1 Code: Ashish Seth, Ramaneswaran S., S. Sakshi, Sonal Kumar, Sreyan Ghosh, Dinesh Manocha |
EMNLP | 2 |
| 2024 | CompA: Addressing the Gap in Compositional Reasoning in Audio-Language ModelsabstractA fundamental characteristic of audio is its compositional nature. Audio-language models (ALMs) trained using a contrastive approach (e.g., CLAP) that learns a shared representation between audio and language modalities have improved performance in many downstream applications, including zero-shot audio classification, audio retrieval, etc. However, the ability of these models to effectively perform compositional reasoning remains largely unexplored and necessitates additional research. In this paper, we propose CompA, a collection of two expert-annotated benchmarks with a majority of real-world audio samples, to evaluate compositional reasoning in ALMs. Our proposed CompA-order evaluates how well an ALM understands the order or occurrence of acoustic events in audio, and CompA-attribute evaluates attribute-binding of acoustic events. An instance from either benchmark consists of two audio-caption pairs, where both audios have the same acoustic events but with different compositions. An ALM is evaluated on how well it matches the right audio to the right caption. Using this benchmark, we first show that current ALMs perform only marginally better than random chance, thereby struggling with compositional reasoning. Next, we propose CompA-CLAP, where we fine-tune CLAP using a novel learning method to improve its compositional reasoning abilities. To train CompA-CLAP, we first propose improvements to contrastive training with composition-aware hard negatives, allowing for more focused training. Next, we propose a novel modular contrastive loss that helps the model learn fine-grained compositional understanding and overcomes the acute scarcity of openly available compositional audios. CompA-CLAP significantly improves over all our baseline models on the CompA benchmark, indicating its superior compositional reasoning capabilities. Sreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi, Chandra Kiran Reddy Evuru, Ramaneswaran S., Sakshi Singh, Oriol Nieto, Ramani Duraiswami, Dinesh Manocha |
ICLR | 6 |
| 2024 | A Closer Look at the Limitations of Instruction TuningabstractInstruction Tuning (IT), the process of training large language models (LLMs) using instruction-response pairs, has emerged as the predominant method for transforming base pre-trained LLMs into open-domain conversational agents. While IT has achieved notable success and widespread adoption, its limitations and shortcomings remain underexplored. In this paper, through rigorous experiments and an in-depth analysis of the changes LLMs undergo through IT, we reveal various limitations of IT. In particular, we show that (1) IT fails to enhance knowledge or skills in LLMs. LoRA fine-tuning is limited to learning response initiation and style tokens, and full-parameter fine-tuning leads to knowledge degradation. (2) Copying response patterns from IT datasets derived from knowledgeable sources leads to a decline in response quality. (3) Full-parameter fine-tuning increases hallucination by inaccurately borrowing tokens from conceptually similar instances in the IT dataset for generating responses. (4) Popular methods to improve IT do not lead to performance improvements over a simple LoRA fine-tuned model. Our findings reveal that responses generated solely from pre-trained knowledge consistently outperform responses by models that learn any form of new knowledge from IT on open-source datasets. We hope the insights and challenges revealed in this paper inspire future work in related directions. Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S., Deepali Aneja, Zeyu Jin, Ramani Duraiswami, Dinesh Manocha |
ICML | 4 |
| 2024 | Composite Diffusion: whole >= ΣpartsabstractFor artists or graphic designers, the spatial arrangement of a scene is a critical design choice. However, existing text-to-image diffusion models provide limited support for incorporating spatial information. This paper introduces Composite Diffusion as a means for artists to generate high-quality images by composing from sub-scenes. The artists can specify the arrangement of the sub-scenes through a free-form segment layout and can describe the content of each sub-scene using natural text and additional control inputs. We provide a comprehensive and modular framework for Composite Diffusion that enables alternative ways of generating, composing, and harmonizing sub-scenes.We further argue that existing image quality metrics lack a holistic evaluation of image composites. To address this, we propose novel quality criteria especially relevant to composite generation. We believe that our approach provides an intuitive method of art creation. Through extensive user surveys and quantitative and qualitative analysis, we show how it achieves greater spatial, semantic, and creative control over image generation. In addition, our methods do not need to retrain or modify the architecture of the base diffusion models and can work in a plug-and-play manner with the fine-tuned models. Vikram Jamwal, Ramaneswaran S. |
WACV | 2 |
| 2024 | Emotion-Aware Multimodal Fusion for Meme Emotion DetectionabstractThe ever-evolving social media discourse has witnessed an overwhelming use of memes to express opinions or dissent. Besides being misused for spreading malcontent, they are mined by corporations and political parties to glean the public's opinion. Therefore, memes predominantly offer affect-enriched insights towards ascertaining the societal psyche. However, the current approaches are yet to model the affective dimensions expressed in memes effectively. They rely extensively on large multimodal datasets for pre-training and do not generalize well due to constrained visual-linguistic grounding. In this paper, we introduce MOOD (Meme emOtiOns Dataset), which embodies six basic emotions. We then present ALFRED (emotion-Aware muLtimodal Fusion foR Emotion Detection), a novel multimodal neural framework that (i) explicitly models emotion-enriched visual cues, and (ii) employs an efficient cross-modal fusion via a gating mechanism. Our investigation establishes ALFRED's superiority over existing baselines by 4.94% F1. Additionally, ALFRED competes strongly with previous best approaches on the challenging Memotion task. We then discuss ALFRED's domain-agnostic generalizability by demonstrating its dominance on two recently-released datasets - HarMeme and Dank Memes, over other baselines. Further, we analyze ALFRED's interpretability using attention maps. Finally, we highlight the inherent challenges posed by the complex interplay of disparate modality-specific cues toward meme analysis. Ramaneswaran S., Md. Shad Akhtar, Tanmoy Chakraborty 0002 |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | ACLM: A Selective-Denoising based Generative Data Augmentation Approach for Low-Resource Complex NERabstractSreyan Ghosh, Utkarsh Tyagi, Manan Suri, Sonal Kumar, Ramaneswaran S, Dinesh Manocha. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Sreyan Ghosh, Utkarsh Tyagi, Manan Suri, Sonal Kumar, Ramaneswaran S., Dinesh Manocha |
ACL (1) | 5 |
| 2023 | MEMEX: Detecting Explanatory Evidence for Memes via Knowledge-Enriched ContextualizationabstractMemes are a powerful tool for communication over social media.Their affinity for evolving across politics, history, and sociocultural phenomena makes them an ideal communication vehicle.To comprehend the subtle message conveyed within a meme, one must understand the background that facilitates its holistic assimilation.Besides digital archiving of memes and their metadata by a few websites like knowyourmeme.com,currently, there is no efficient way to deduce a meme's context dynamically.In this work, we propose a novel task, MEMEX -given a meme and a related document, the aim is to mine the context that succinctly explains the background of the meme.At first, we develop MCC (Meme Context Corpus), a novel dataset for MEMEX.Further, to benchmark MCC, we propose MIME (MultImodal Meme Explainer), a multimodal neural framework that uses common sense enriched meme representation and a layered approach to capture the cross-modal semantic dependencies between the meme and the context.MIME surpasses several unimodal and multimodal systems and yields an absolute improvement of ⇡ 4% F1-score over the best baseline.Lastly, we conduct detailed analyses of MIME's performance, highlighting the aspects that could lead to optimal modeling of cross-modal contextual associations. Ramaneswaran S., Udit Arora, Md. Shad Akhtar, Tanmoy Chakraborty 0002 |
ACL (1) | 2 |
| 2023 | DALE: Generative Data Augmentation for Low-Resource Legal NLPabstractSreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, S Ramaneswaran, S Sakshi, Utkarsh Tyagi, Dinesh Manocha. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Sreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S., Sakshi Singh, Utkarsh Tyagi, Dinesh Manocha |
EMNLP | 4 |
| 2023 | From Multilingual Complexity to Emotional Clarity: Leveraging Commonsense to Unveil Emotions in Code-Mixed DialoguesabstractUnderstanding emotions during conversation is a fundamental aspect of human communication, driving NLP research for Emotion Recognition in Conversation (ERC).While considerable research has focused on discerning emotions of individual speakers in monolingual dialogues, understanding the emotional dynamics in codemixed conversations has received relatively less attention.This motivates our undertaking of ERC for code-mixed conversations in this study.Recognizing that emotional intelligence encompasses a comprehension of worldly knowledge, we propose an innovative approach that integrates commonsense information with dialogue context to facilitate a deeper understanding of emotions.To achieve this, we devise an efficient pipeline that extracts relevant commonsense from existing knowledge graphs based on the code-mixed input.Subsequently, we develop an advanced fusion technique that seamlessly combines the acquired commonsense information with the dialogue representation obtained from a dedicated dialogue understanding module.Our comprehensive experimentation showcases the substantial performance improvement obtained through the systematic incorporation of commonsense in ERC.Both quantitative assessments and qualitative analyses further corroborate the validity of our hypothesis, reaffirming the pivotal role of commonsense integration in enhancing ERC. Shivani Kumar, Ramaneswaran S., Md. Shad Akhtar, Tanmoy Chakraborty 0002 |
EMNLP | 2 |