EDBT 2026 Demo / reviewers in the wild / expert
Navonil Majumder
dblp:198/3608 · also Navonil Mazumder
· DBLP profile ↗
31ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0002-1449-617XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 10 Open Challenges Steering the Future of Vision-Language-Action ModelsabstractDue to their ability of follow natural language instructions, vision-language-action (VLA) models are increasingly preva- lent in the embodied AI arena, following the widespread suc- cess of their precursors—LLMs and VLMs. In this paper, we discuss 10 principal milestones in the ongoing develop- ment of VLA models—multimodality, reasoning, data, eval- uation, cross-robkot action generalization, efficiency, whole- body coordination, safety, agents, and coordination with hu- mans. Furthermore, we discuss the emerging trends of us- ing spatial understanding, modeling world dynamics, post training, and data synthesis—all aiming to reach these mile- stones. Through these discussions, we hope to bring attention to the research avenues that may accelerate the development of VLA models into wider acceptability. Soujanya Poria, Navonil Majumder, Chia-Yu Hung, Amir Ali Bagherzadeh, Kenneth Kwok, Ziwei Wang 0010, Cheston Tan, Jiajun Wu 0001, David Hsu |
AAAI | 2 |
| 2025 | Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to RefuseabstractLLMs are an integral component of retrieval-augmented generation (RAG) systems. While many studies focus on evaluating the overall quality of end-to-end RAG systems, there is a gap in understanding the appropriateness of LLMs for the RAG task. To address this, we introduce Trust-Score, a holistic metric that evaluates the trustworthiness of LLMs within the RAG framework. Our results show that various prompting methods, such as in-context learning, fail to effectively adapt LLMs to the RAG task as measured by Trust-Score. Consequently, we propose Trust-Align, a method to align LLMs for improved Trust-Score performance. 26 out of 27 models aligned using Trust-Align substantially outperform competitive baselines on ASQA, QAMPARI, and ELI5. Specifically, in LLaMA-3-8b, Trust-Align outperforms FRONT on ASQA (↑12.56), QAMPARI (↑36.04), and ELI5 (↑17.69). Trust-Align also significantly enhances models’ ability to correctly refuse and provide quality citations. We also demonstrate the effectiveness of Trust-Align across different open-weight models, including the LLaMA series (1b to 8b), Qwen-2.5 series (0.5b to 7b), and Phi3.5 (3.8b). We release our code at https://github.com/declare-lab/trust-align. Maojia Song, Shang Hong Sim, Rishabh Bhardwaj, Hai Leong Chieu, Navonil Majumder, Soujanya Poria |
ICLR | 5 |
| 2025 | Reward-Guided Tree Search for Inference Time Alignment of Large Language ModelsabstractChia-Yu Hung, Navonil Majumder, Ambuj Mehrish, Soujanya Poria. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Chia-Yu Hung, Navonil Majumder, Ambuj Mehrish, Soujanya Poria |
NAACL (Long Papers) | 2 |
| 2024 | Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference OptimizationabstractPeer Reviewed Navonil Majumder, Chia-Yu Hung, Deepanway Ghosal, Wei-Ning Hsu, Rada Mihalcea, Soujanya Poria |
ACM Multimedia | 1 |
| 2024 | Mustango: Toward Controllable Text-to-Music GenerationabstractJan Melechovsky, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, Soujanya Poria. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jan Melechovský, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, Soujanya Poria |
NAACL-HLT | 4 |
| 2023 | Evaluating Parameter-Efficient Transfer Learning Approaches on SURE Benchmark for Speech UnderstandingabstractFine-tuning is widely used as the default algorithm for transfer learning from pre-trained models. Parameter inefficiency can however arise when, during transfer learning, all the parameters of a large pre-trained model need to be updated for individual downstream tasks. As the number of parameters grows, fine-tuning is prone to overfitting and catastrophic forgetting. In addition, full fine-tuning can become prohibitively expensive when the model is used for many tasks. To mitigate this issue, parameter-efficient transfer learning algorithms, such as adapters and prefix tuning, have been proposed as a way to introduce a few trainable parameters that can be plugged into large pre-trained language models such as BERT, HuBERT. In this paper, we introduce the Speech UndeRstanding Evaluation (SURE) benchmark for parameter-efficient learning for various speech processing tasks. Additionally, we introduce a new adapter, ConvAdapter, based on 1D convolution. We show that ConvAdapter outperforms the standard adapters while showing comparable performance against prefix tuning and Low-Rank Adaptation with only 0.94% of trainable parameters. Yingting Li, Ambuj Mehrish, Rishabh Bhardwaj, Navonil Majumder, Bo Cheng 0001, Shuai Zhao 0001, Amir Zadeh 0001, Rada Mihalcea, Soujanya Poria |
ICASSP | 4 |
| 2023 | ADAPTERMIX: Exploring the Efficacy of Mixture of Adapters for Low-Resource TTS AdaptationabstractThere are significant challenges for speaker adaptation in textto-speech for languages that are not widely spoken or for speakers with accents or dialects that are not well-represented in the training data. To address this issue, we propose the use of the ”mixture of adapters” method. This approach involves adding multiple adapters within a backbone-model layer to learn the unique characteristics of different speakers. Our approach outperforms the baseline, with a noticeable improvement of 5% observed in speaker preference tests when using only one minute of data for each new speaker. Moreover, following the adapter paradigm, we fine-tune only the adapter parameters (11% of the total model parameters). This is a significant achievement in parameter-efficient speaker adaptation, and one of the first models of its kind. Overall, our proposed approach offers a promising solution to the speech synthesis techniques, particularly for adapting to speakers from diverse backgrounds. Ambuj Mehrish, Abhinav Ramesh Kashyap, Yingting Li, Navonil Majumder, Soujanya Poria |
INTERSPEECH | 4 |
| 2023 | Sentence Embedder Guided Utterance Encoder (SEGUE) for Spoken Language Understanding
Yi Xuan Tan, Navonil Majumder, Soujanya Poria |
INTERSPEECH | 2 |
| 2023 | Text-to-Audio Generation using Instruction Guided Latent Diffusion ModelabstractThe immense scale of the recent large language models (LLM) allows many interesting properties, such as, instruction- and chain-of-thought-based fine-tuning, that has significantly improved zero- and few-shot performance in many natural language processing (NLP) tasks. Inspired by such successes, we adopt such an instruction-tuned LLM Flan-T5 as the text encoder for text-to-audio (TTA) generation-a task where the goal is to generate an audio from its textual description. The prior works on TTA either pre-trained a joint text-audio encoder or used a non-instruction-tuned model, such as, T5. Consequently, our latent diffusion model (LDM)-based approach (Tango) outperforms the state-of-the-art AudioLDM on most metrics and stays comparable on the rest on AudioCaps test set, despite training the LDM on a 63 times smaller dataset and keeping the text encoder frozen. This improvement might also be attributed to the adoption of audio pressure level-based sound mixing for training set augmentation, whereas the prior methods take a random mix. Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, Soujanya Poria |
ACM Multimedia | 2 |
| 2023 | Beneath the Tip of the Iceberg: Current Challenges and New Directions in Sentiment Analysis ResearchabstractSentiment analysis as a field has come a long way since it was first introduced as a task nearly 20 years ago. It has widespread commercial applications in various domains like marketing, risk management, market research, and politics, to name a few. Given its saturation in specific subtasks — such as sentiment polarity classification — and datasets, there is an underlying perception that this field has reached its maturity. In this article, we discuss this perception by pointing out the shortcomings and under-explored, yet key aspects of this field necessary to attaintruesentiment understanding. We analyze the significant leaps responsible for its current relevance. Further, we attempt to chart a possible course for this field that covers many overlooked and unanswered questions. Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Rada Mihalcea |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | CICERO: A Dataset for Contextualized Commonsense Inference in DialoguesabstractThis paper addresses the problem of dialogue reasoning with contextualized commonsense inference.We curate CICERO, a dataset of dyadic conversations with five types of utterance-level reasoning-based inferences: cause, subsequent event, prerequisite, motivation, and emotional reaction.The dataset contains 53,105 of such inferences from 5,672 dialogues.We use this dataset to solve relevant generative and discriminative tasks: generation of cause and subsequent event; generation of prerequisite, motivation, and listener's emotional reaction; and selection of plausible alternatives.Our results ascertain the value of such dialogue-centric commonsense knowledge datasets.It is our hope that CI-CERO will open new research avenues into commonsense-based dialogue reasoning. Deepanway Ghosal, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
ACL (1) | 3 |
| 2022 | Two is Better than Many? Binary Classification as an Effective Approach to Multi-Choice Question AnsweringabstractWe propose a simple refactoring of multichoice question answering (MCQA) tasks as a series of binary classifications.The MCQA task is generally performed by scoring each (question, answer) pair normalized over all the pairs, and then selecting the answer from the pair that yield the highest score.For n answer choices, this is equivalent to an n-class classification setup where only one class (true answer) is correct.We instead show that classifying (question, true answer) as positive instances and (question, false answer) as negative instances is significantly more effective across various models and datasets.We show the efficacy of our proposed approach in different tasks -abductive reasoning, commonsense question answering, science question answering, and sentence completion.Our DeBERTa binary classification model reaches the top or close to the top performance on public leaderboards for these tasks.The source code of the proposed approach is available at https Deepanway Ghosal, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
EMNLP | 2 |
| 2022 | Improving aspect-level sentiment analysis with aspect extraction
Navonil Majumder, Rishabh Bhardwaj, Soujanya Poria, Alexander F. Gelbukh, Amir Hussain 0001 |
Neural Comput. Appl. | 1 |
| 2021 | More Identifiable yet Equally Performant Transformers for Text ClassificationabstractRishabh Bhardwaj, Navonil Majumder, Soujanya Poria, Eduard Hovy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Rishabh Bhardwaj, Navonil Majumder, Soujanya Poria, Eduard H. Hovy |
ACL/IJCNLP (1) | 2 |
| 2021 | Aspect Sentiment Triplet Extraction Using Reinforcement LearningabstractAspect Sentiment Triplet Extraction (ASTE) is the task of extracting triplets of aspect terms, their associated sentiments, and the opinion terms that provide evidence for the expressed sentiments. Previous approaches to ASTE usually simultaneously extract all three components or first identify the aspect and opinion terms, then pair them up to predict their sentiment polarities. In this work, we present a novel paradigm, ASTE-RL, by regarding the aspect and opinion terms as arguments of the expressed sentiment in a hierarchical reinforcement learning (RL) framework. We first focus on sentiments expressed in a sentence, then identify the target aspect and opinion terms for that sentiment. This takes into account the mutual interactions among the triplet's components while improving exploration and sample efficiency. Furthermore, this hierarchical RL setup enables us to deal with multiple and overlapping triplets. In our experiments, we evaluate our model on existing datasets from laptop and restaurant domains and show that it achieves state-of-the-art performance. The implementation of this work is publicly available at https://github.com/declare-lab/ASTE-RL. Samson Yu Bai Jian, Tapas Nayak, Navonil Majumder, Soujanya Poria |
CIKM | 3 |
| 2021 | STaCK: Sentence Ordering with Temporal Commonsense KnowledgeabstractSentence order prediction is the task of finding the correct order of sentences in a randomly ordered document.Correctly ordering the sentences requires an understanding of coherence with respect to the chronological sequence of events described in the text.Documentlevel contextual understanding and commonsense knowledge centered around these events are often essential in uncovering this coherence and predicting the exact chronological order.In this paper, we introduce STaCK -a framework based on graph neural networks and temporal commonsense knowledge to model global information and predict the relative order of sentences.Our graph network accumulates temporal evidence using knowledge of 'past' and 'future' and formulates sentence ordering as a constrained edge classification problem.We report results on five different datasets, and empirically show that the proposed method is naturally suitable for order prediction, thus demonstrating the role of temporal commonsense knowledge.The implementation of this work is available at: https://github.com/declare-lab/ sentence-ordering.Jennifer has her final exam tomorrow.She got so stressed, she pulled an Deepanway Ghosal, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
EMNLP (1) | 2 |
| 2021 | M2H2: A Multimodal Multiparty Hindi Dataset For Humor Recognition in ConversationsabstractHumor recognition in conversations is a challenging task that has recently gained popularity due to its importance in dialogue understanding, including in multimodal settings (i.e., text, acoustics, and visual). The few existing datasets for humor are mostly in English. However, due to the tremendous growth in multilingual content, there is a great demand to build models and systems that support multilingual information access. To this end, we propose a dataset for Multimodal Multiparty Hindi Humor (M2H2) recognition in conversations containing 6,191 utterances from 13 episodes of a very popular TV series ”Shrimaan Shrimati Phir Se”. Each utterance is annotated with humor/non-humor labels and encompasses acoustic, visual, and textual modalities. We propose several strong multimodal baselines and show the importance of contextual and multimodal information for humor recognition in conversations. The empirical results on M2H2 dataset demonstrate that multimodal information complements unimodal information for humor recognition. The dataset and the baselines are available at http://www.iitp.ac.in/~ai-nlp-ml/resources.html and https://github.com/declare-lab/M2H2-dataset. Dushyant Singh Chauhan, Gopendra Vikram Singh, Navonil Majumder, Amir Zadeh 0001, Asif Ekbal, Pushpak Bhattacharyya, Louis-Philippe Morency, Soujanya Poria |
ICMI | 3 |
| 2021 | CIDER: Commonsense Inference for Dialogue Explanation and ReasoningabstractWell, I missed several buses.How on earth can you miss several buses?I, ah ..., I have got late.But there's a bus every ten minutes, and you are over 1 hour late.Have you got it now? Deepanway Ghosal, Pengfei Hong, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
SIGDIAL | 4 |
| 2021 | Leveraging label hierarchy using transfer and multi-task learning: A case study on patent classification
Segun Taofeek Aroyehun, Jason Angel, Navonil Majumder, Alexander F. Gelbukh, Amir Hussain 0001 |
Neurocomputing | 3 |
| 2021 | Persuasive dialogue understanding: The baselines and negative results
Hui Chen 0023, Deepanway Ghosal, Navonil Majumder, Amir Hussain 0001, Soujanya Poria |
Neurocomputing | 3 |
| 2020 | KinGDOM: Knowledge-Guided DOMain Adaptation for Sentiment AnalysisabstractCross-domain sentiment analysis has received significant attention in recent years, prompted by the need to combat the domain gap between different applications that make use of sentiment analysis.In this paper, we take a novel perspective on this task by exploring the role of external commonsense knowledge.We introduce a new framework, KinGDOM, which utilizes the ConceptNet knowledge graph to enrich the semantics of a document by providing both domain-specific and domain-general background concepts.These concepts are learned by training a graph convolutional autoencoder that leverages inter-domain concepts in a domain-invariant manner.Conditioning a popular domain-adversarial baseline method with these learned concepts helps improve its performance over state-of-the-art approaches, demonstrating the efficacy of our proposed framework. Deepanway Ghosal, Devamanyu Hazarika, Abhinaba Roy, Navonil Majumder, Rada Mihalcea, Soujanya Poria |
ACL | 4 |
| 2020 | MIME: MIMicking Emotions for Empathetic Response GenerationabstractNavonil Majumder, Pengfei Hong, Shanshan Peng, Jiankun Lu, Deepanway Ghosal, Alexander Gelbukh, Rada Mihalcea, Soujanya Poria. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Navonil Majumder, Pengfei Hong, Shanshan Peng, Jiankun Lu, Deepanway Ghosal, Alexander F. Gelbukh, Rada Mihalcea, Soujanya Poria |
EMNLP (1) | 1 |
| 2019 | DialogueRNN: An Attentive RNN for Emotion Detection in ConversationsabstractEmotion detection in conversations is a necessary step for a number of applications, including opinion mining over chat history, social media threads, debates, argumentation mining, understanding consumer feedback in live conversations, and so on. Currently systems do not treat the parties in the conversation individually by adapting to the speaker of each utterance. In this paper, we describe a new method based on recurrent neural networks that keeps track of the individual party states throughout the conversation and uses this information for emotion classification. Our model outperforms the state-of-the-art by a significant margin on two different datasets. Navonil Majumder, Soujanya Poria, Devamanyu Hazarika, Rada Mihalcea, Alexander F. Gelbukh, Erik Cambria |
AAAI | 1 |
| 2019 | MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in ConversationsabstractEmotion recognition in conversations (ERC) is a challenging task that has recently gained popularity due to its potential applications.Until now, however, there has been no largescale multimodal multi-party emotional conversational database containing more than two speakers per dialogue.To address this gap, we propose the Multimodal EmotionLines Dataset (MELD), an extension and enhancement of EmotionLines.MELD contains about 13,000 utterances from 1,433 dialogues from the TV-series Friends.Each utterance is annotated with emotion and sentiment labels, and encompasses audio, visual, and textual modalities.We propose several strong multimodal baselines and show the importance of contextual and multimodal information for emotion recognition in conversations.The full dataset is available for use at http:// affective-meld.github.io. Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, Rada Mihalcea |
ACL (1) | 3 |
| 2019 | DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in ConversationabstractDeepanway Ghosal, Navonil Majumder, Soujanya Poria, Niyati Chhaya, Alexander Gelbukh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Deepanway Ghosal, Navonil Majumder, Soujanya Poria, Niyati Chhaya, Alexander F. Gelbukh |
EMNLP/IJCNLP (1) | 2 |
| 2018 | Memory Fusion Network for Multi-view Sequential LearningabstractMulti-view sequential learning is a fundamental problem in machine learning dealing with multi-view sequences. In a multi-view sequence, there exists two forms of interactions between different views: view-specific interactions and cross-view interactions. In this paper, we present a new neural architecture for multi-view sequential learning called the Memory Fusion Network (MFN) that explicitly accounts for both interactions in a neural architecture and continuously models them through time. The first component of the MFN is called the System of LSTMs, where view-specific interactions are learned in isolation through assigning an LSTM function to each view. The cross-view interactions are then identified using a special attention mechanism called the Delta-memory Attention Network (DMAN) and summarized through time with a Multi-view Gated Memory. Through extensive experimentation, MFN is compared to various proposed approaches for multi-view sequential learning on multiple publicly available benchmark datasets. MFN outperforms all the multi-view approaches. Furthermore, MFN outperforms all current state-of-the-art models, setting new state-of-the-art results for all three multi-view datasets. Amir Zadeh 0001, Paul Pu Liang, Navonil Majumder, Soujanya Poria, Erik Cambria, Louis-Philippe Morency |
AAAI | 3 |
| 2018 | A Deep Learning Approach for Multimodal Deception Detection
Gangeshwar Krishnamurthy, Navonil Majumder, Soujanya Poria, Erik Cambria |
CICLing (1) | 2 |
| 2018 | IARM: Inter-Aspect Relation Modeling with Memory Networks in Aspect-Based Sentiment AnalysisabstractSentiment analysis has immense implications in modern businesses through user-feedback mining.Large product-based enterprises like Samsung and Apple make crucial business decisions based on the large quantity of user reviews and suggestions available in different e-commerce websites and social media platforms like Amazon and Facebook.Sentiment analysis caters to these needs by summarizing user sentiment behind a particular object.In this paper, we present a novel approach of incorporating the neighboring aspects related information into the sentiment classification of the target aspect using memory networks.Our method outperforms the state of the art by 1.6% on average in two distinct domains. Navonil Majumder, Soujanya Poria, Alexander F. Gelbukh, Md. Shad Akhtar, Erik Cambria, Asif Ekbal |
EMNLP | 1 |
| 2018 | Multimodal sentiment analysis using hierarchical fusion with context modeling
Navonil Majumder, Devamanyu Hazarika, Alexander F. Gelbukh, Erik Cambria, Soujanya Poria |
Knowl. Based Syst. | 1 |
| 2017 | Context-Dependent Sentiment Analysis in User-Generated VideosabstractSoujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, Louis-Philippe Morency. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh 0001, Louis-Philippe Morency |
ACL (1) | 4 |
| 2017 | Multi-level Multiple Attentions for Contextual Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis involves identifying sentiment in videos and is a developing field of research. Unlike current works, which model utterances individually, we propose a recurrent model that is able to capture contextual information among utterances. In this paper, we also introduce attentionbased networks for improving both context learning and dynamic feature fusion. Our model shows 6-8% improvement over the state of the art on a benchmark dataset. Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh 0001, Louis-Philippe Morency |
ICDM | 4 |