VLDB 2026 Research / reviewers in the wild / expert
Deepanway Ghosal
dblp:203/9407
· DBLP profile ↗
21ranked-venue papers
13as first author
14since 2021 · last 2025
0000-0002-3858-4449ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 12 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial ReasoningabstractTraditional reinforcement learning-based robotic control methods are often task-specific and fail to generalize across diverse environments or unseen objects and instructions. Visual Language Models (VLMs) demonstrate strong scene understanding and planning capabilities but lack the ability to generate actionable policies tailored to specific robotic embodiments. To address this, Visual-Language-Action (VLA) models have emerged, yet they face challenges in long-horizon spatial reasoning and grounded task planning. In this work, we propose the Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning, EMMA-X. EMMA-X leverages our constructed hierarchical embodiment dataset based on BridgeV2, containing 60,000 robot manipulation trajectories auto-annotated with grounded task reasoning and spatial guidance. Additionally, we introduce a trajectory segmentation strategy based on gripper states and motion trajectories, which can help mitigate hallucination in grounding subtask reasoning generation. Experimental results demonstrate that EMMA-X achieves superior performance over competitive baselines, particularly in real-world robotic tasks requiring spatial reasoning. Pengfei Hong, Pala Tej Deep, Vernon Toh Yan Han, U-Xuan Tan, Deepanway Ghosal, Soujanya Poria |
ACL (1) | 6 |
| 2025 | MixEval-X: Any-to-any Evaluations from Real-world Data MixtureabstractPerceiving and generating diverse modalities are crucial for AI models to effectively learn from and engage with real-world signals, necessitating reliable evaluations for their development. We identify two major issues in current evaluations: (1) inconsistent standards, shaped by different communities with varying protocols and maturity levels; and (2) significant query, grading, and generalization biases. To address these, we introduce MixEval-X, the first any-to-any, real-world benchmark designed to optimize and standardize evaluations across diverse input and output modalities. We propose multi-modal benchmark mixture and adaptation-rectification pipelines to reconstruct real-world task distributions, ensuring evaluations generalize effectively to real-world use cases. Extensive meta-evaluations show our approach effectively aligns benchmark samples with real-world task distributions. Meanwhile, MixEval-X's model rankings correlate strongly with that of crowd-sourced real-world evaluations (up to 0.98) while being much more efficient. We provide comprehensive leaderboards to rerank existing models and organizations and offer insights to enhance understanding of multi-modal evaluations and inform future research. Jinjie Ni, Deepanway Ghosal, Bo Li 0080, Junhao Zhang 0001, Xiang Yue, Fuzhao Xue, Yuntian Deng, Zian Zheng 0001, Kaichen Zhang, Mahir Shah, Kabir Jain, Yang You 0001, Michael Shieh |
ICLR | 3 |
| 2025 | AlgoPuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Algorithmic Multimodal PuzzlesabstractDeepanway Ghosal, Vernon Toh, Yew Ken Chia, Soujanya Poria. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Deepanway Ghosal, Vernon Toh Yan Han, Yew Ken Chia, Soujanya Poria |
NAACL (Long Papers) | 1 |
| 2024 | Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference OptimizationabstractPeer Reviewed Navonil Majumder, Chia-Yu Hung, Deepanway Ghosal, Wei-Ning Hsu, Rada Mihalcea, Soujanya Poria |
ACM Multimedia | 3 |
| 2024 | Mustango: Toward Controllable Text-to-Music GenerationabstractJan Melechovsky, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, Soujanya Poria. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jan Melechovský, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, Soujanya Poria |
NAACL-HLT | 3 |
| 2023 | ReTAG: Reasoning Aware Table to Analytic Text GenerationabstractThe task of table summarization involves generating text that both succinctly and accurately represents the table or a specific set of highlighted cells within a table.While significant progress has been made in table to text generation techniques, models still mostly generate descriptive summaries, which reiterates the information contained within the table in sentences.Through analysis of popular table to text benchmarks (ToTTo (Parikh et al., 2020) and InfoTabs (Gupta et al., 2020)) we observe that in order to generate the ideal summary, multiple types of reasoning is needed coupled with access to knowledge beyond the scope of the table.To address this gap, we propose RETAG, a table and reasoning aware model that uses vector-quantization to infuse different types of analytical reasoning into the output.RETAG achieves 2.2%, 2.9% improvement on the PARENT metric in the relevant slice of ToTTo and InfoTabs for the table to text generation task over state of the art baselines.Through human evaluation, we observe that output from RETAG is upto 12% more faithful and analytical compared to a strong table-aware model.To the best of our knowledge, RETAG is the first model that can controllably use multiple reasoning methods within a structure-aware sequence to sequence model to surpass state of the art performance in multiple table to text tasks.We extend (and open source 35.6K analytical, 55.9k descriptive instances) the ToTTo, InfoTabs datasets with the reasoning categories used in each reference sentences. Deepanway Ghosal, Preksha Nema, Aravindan Raghuveer |
EMNLP | 1 |
| 2023 | Prover: Generating Intermediate Steps for NLI with Commonsense Knowledge Retrieval and Next-Step PredictionabstractDeepanway Ghosal, Somak Aditya, Monojit Choudhury. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Deepanway Ghosal, Somak Aditya, Monojit Choudhury |
IJCNLP (1) | 1 |
| 2023 | Text-to-Audio Generation using Instruction Guided Latent Diffusion ModelabstractThe immense scale of the recent large language models (LLM) allows many interesting properties, such as, instruction- and chain-of-thought-based fine-tuning, that has significantly improved zero- and few-shot performance in many natural language processing (NLP) tasks. Inspired by such successes, we adopt such an instruction-tuned LLM Flan-T5 as the text encoder for text-to-audio (TTA) generation-a task where the goal is to generate an audio from its textual description. The prior works on TTA either pre-trained a joint text-audio encoder or used a non-instruction-tuned model, such as, T5. Consequently, our latent diffusion model (LDM)-based approach (Tango) outperforms the state-of-the-art AudioLDM on most metrics and stays comparable on the rest on AudioCaps test set, despite training the LDM on a 63 times smaller dataset and keeping the text encoder frozen. This improvement might also be attributed to the adoption of audio pressure level-based sound mixing for training set augmentation, whereas the prior methods take a random mix. Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, Soujanya Poria |
ACM Multimedia | 1 |
| 2022 | CICERO: A Dataset for Contextualized Commonsense Inference in DialoguesabstractThis paper addresses the problem of dialogue reasoning with contextualized commonsense inference.We curate CICERO, a dataset of dyadic conversations with five types of utterance-level reasoning-based inferences: cause, subsequent event, prerequisite, motivation, and emotional reaction.The dataset contains 53,105 of such inferences from 5,672 dialogues.We use this dataset to solve relevant generative and discriminative tasks: generation of cause and subsequent event; generation of prerequisite, motivation, and listener's emotional reaction; and selection of plausible alternatives.Our results ascertain the value of such dialogue-centric commonsense knowledge datasets.It is our hope that CI-CERO will open new research avenues into commonsense-based dialogue reasoning. Deepanway Ghosal, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
ACL (1) | 1 |
| 2022 | Two is Better than Many? Binary Classification as an Effective Approach to Multi-Choice Question AnsweringabstractWe propose a simple refactoring of multichoice question answering (MCQA) tasks as a series of binary classifications.The MCQA task is generally performed by scoring each (question, answer) pair normalized over all the pairs, and then selecting the answer from the pair that yield the highest score.For n answer choices, this is equivalent to an n-class classification setup where only one class (true answer) is correct.We instead show that classifying (question, true answer) as positive instances and (question, false answer) as negative instances is significantly more effective across various models and datasets.We show the efficacy of our proposed approach in different tasks -abductive reasoning, commonsense question answering, science question answering, and sentence completion.Our DeBERTa binary classification model reaches the top or close to the top performance on public leaderboards for these tasks.The source code of the proposed approach is available at https Deepanway Ghosal, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
EMNLP | 1 |
| 2022 | All-in-One: Emotion, Sentiment and Intensity Prediction Using a Multi-Task Ensemble FrameworkabstractWe propose a multi-task ensemble framework that jointly learns multiple related problems. The ensemble model aims to leverage the learned representations of three deep learning models (i.e., CNN, LSTM and GRU) and a hand-crafted feature representation for the predictions. Through multi-task framework, we address four problems of emotion and sentiment analysis, i.e., “emotionclassification&intensity”, “valence,arousal&dominancefor emotion”, “valence&arousalfor sentiment”, and “3-class categorical&5-class ordinal classificationfor sentiment”. The underlying problems cover two granularity (i.e.,coarse-grainedandfine-grained) and a diverse range of domains (i.e.,tweets,Facebook posts,news headlines,blogs,lettersetc.). Experimental results suggest that the proposed multi-task framework outperforms the single-task frameworks in all experiments. Md. Shad Akhtar, Deepanway Ghosal, Asif Ekbal, Pushpak Bhattacharyya, Sadao Kurohashi |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | STaCK: Sentence Ordering with Temporal Commonsense KnowledgeabstractSentence order prediction is the task of finding the correct order of sentences in a randomly ordered document.Correctly ordering the sentences requires an understanding of coherence with respect to the chronological sequence of events described in the text.Documentlevel contextual understanding and commonsense knowledge centered around these events are often essential in uncovering this coherence and predicting the exact chronological order.In this paper, we introduce STaCK -a framework based on graph neural networks and temporal commonsense knowledge to model global information and predict the relative order of sentences.Our graph network accumulates temporal evidence using knowledge of 'past' and 'future' and formulates sentence ordering as a constrained edge classification problem.We report results on five different datasets, and empirically show that the proposed method is naturally suitable for order prediction, thus demonstrating the role of temporal commonsense knowledge.The implementation of this work is available at: https://github.com/declare-lab/ sentence-ordering.Jennifer has her final exam tomorrow.She got so stressed, she pulled an Deepanway Ghosal, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
EMNLP (1) | 1 |
| 2021 | CIDER: Commonsense Inference for Dialogue Explanation and ReasoningabstractWell, I missed several buses.How on earth can you miss several buses?I, ah ..., I have got late.But there's a bus every ten minutes, and you are over 1 hour late.Have you got it now? Deepanway Ghosal, Pengfei Hong, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
SIGDIAL | 1 |
| 2021 | Persuasive dialogue understanding: The baselines and negative results
Hui Chen 0023, Deepanway Ghosal, Navonil Majumder, Amir Hussain 0001, Soujanya Poria |
Neurocomputing | 2 |
| 2020 | KinGDOM: Knowledge-Guided DOMain Adaptation for Sentiment AnalysisabstractCross-domain sentiment analysis has received significant attention in recent years, prompted by the need to combat the domain gap between different applications that make use of sentiment analysis.In this paper, we take a novel perspective on this task by exploring the role of external commonsense knowledge.We introduce a new framework, KinGDOM, which utilizes the ConceptNet knowledge graph to enrich the semantics of a document by providing both domain-specific and domain-general background concepts.These concepts are learned by training a graph convolutional autoencoder that leverages inter-domain concepts in a domain-invariant manner.Conditioning a popular domain-adversarial baseline method with these learned concepts helps improve its performance over state-of-the-art approaches, demonstrating the efficacy of our proposed framework. Deepanway Ghosal, Devamanyu Hazarika, Abhinaba Roy, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
ACL | 1 |
| 2020 | MIME: MIMicking Emotions for Empathetic Response GenerationabstractNavonil Majumder, Pengfei Hong, Shanshan Peng, Jiankun Lu, Deepanway Ghosal, Alexander Gelbukh, Rada Mihalcea, Soujanya Poria. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Navonil Majumder, Pengfei Hong, Shanshan Peng, Jiankun Lu, Deepanway Ghosal, Alexander F. Gelbukh, Rada Mihalcea, Soujanya Poria |
EMNLP (1) | 5 |
| 2019 | DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in ConversationabstractDeepanway Ghosal, Navonil Majumder, Soujanya Poria, Niyati Chhaya, Alexander Gelbukh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Deepanway Ghosal, Navonil Majumder, Soujanya Poria, Niyati Chhaya, Alexander F. Gelbukh |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Contextual Inter-modal Attention for Multi-modal Sentiment AnalysisabstractMulti-modal sentiment analysis offers various challenges, one being the effective combination of different input modalities, namely text, visual and acoustic.In this paper, we propose a recurrent neural network based multi-modal attention framework that leverages the contextual information for utterance-level sentiment prediction.The proposed approach applies attention on multi-modal multi-utterance representations and tries to learn the contributing features amongst them.We evaluate our proposed approach on two multi-modal sentiment analysis benchmark datasets, viz.CMU Multi-modal Opinion-level Sentiment Intensity (CMU-MOSI) corpus and the recently released CMU Multi-modal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) corpus.Evaluation results show the effectiveness of our proposed approach with the accuracies of 82.31% and 79.80% for the MOSI and MO-SEI datasets, respectively.These are approximately 2 and 1 points performance improvement over the state-of-the-art models for the datasets. Deepanway Ghosal, Md. Shad Akhtar, Dushyant Singh Chauhan, Soujanya Poria, Asif Ekbal, Pushpak Bhattacharyya |
EMNLP | 1 |
| 2018 | Deep Ensemble Model with the Fusion of Character, Word and Lexicon Level Information for Emotion and Sentiment Prediction
Deepanway Ghosal, Md. Shad Akhtar, Asif Ekbal, Pushpak Bhattacharyya |
ICONIP (5) | 1 |
| 2018 | Music Genre Recognition Using Deep Neural Networks and Transfer Learning
Deepanway Ghosal, Maheshkumar H. Kolekar |
INTERSPEECH | 1 |
| 2017 | A Multilayer Perceptron based Ensemble Technique for Fine-grained Financial Sentiment AnalysisabstractIn this paper, we propose a novel method for combining deep learning and classical feature based models using a Multi-Layer Perceptron (MLP) network for financial sentiment analysis.We develop various deep learning models based on Convolutional Neural Network (CNN), Long Short Term Memory (LSTM) and Gated Recurrent Unit (GRU).These are trained on top of pre-trained, autoencoder-based, financial word embeddings and lexicon features.An ensemble is constructed by combining these deep learning models and a classical supervised model based on Support Vector Regression (SVR).We evaluate our proposed technique on a benchmark dataset of SemEval-2017 shared task on financial sentiment analysis.The propose model shows impressive results on two datasets, i.e. microblogs and news headlines datasets.Comparisons show that our proposed model performs better than the existing state-of-the-art systems for the above two datasets by 2.0 and 4.1 cosine points, respectively. Md. Shad Akhtar, Deepanway Ghosal, Asif Ekbal, Pushpak Bhattacharyya |
EMNLP | 3 |