VLDB 2026 Research / reviewers in the wild / expert
Mauajama Firdaus
dblp:223/8272
· DBLP profile ↗
40ranked-venue papers
19as first author
32since 2021 · last 2025
0000-0001-7485-5974ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 14 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GENTEEL-NEGOTIATOR: LLM-Enhanced Mixture-of-Expert-Based Reinforcement Learning Approach for Polite Negotiation DialogueabstractDeveloping intelligent negotiation dialogue systems that resolve conflicts and promote equitable, inclusive, and sustainable outcomes is at the forefront of advancing automated negotiation technology for social good. Negotiation involves balancing cooperation and competition to maximize value without causing offense. Using polite language fosters mutual understanding and creates a respectful and collaborative environment essential for successful negotiations in various domains. Considering this, in this paper, we propose a polite negotiation dialogue system, GENTEEL-NEGOTIATOR for social good applications to boost the overall quality of negotiation outcomes. We focus on developing a negotiation dialogue system for two key application areas, namely tourism and e-commerce. We begin by curating a unique negotiation dialogue dataset, NEGOCHAT for tourism. We further enrich the NEGOCHAT and Integrative Negotiation Dataset (IND) for e-commerce with various negotiation strategies. These datasets are then used to develop the GENTEEL-NEGOTIATOR, leveraging the Large Language Model (LLM) and mixture-of-expert (MoE)-based reinforcement learning approach. The proposed MoE-based method employs heuristic experts dedicated to negotiation, politeness, and dialogue coherence to facilitate the learning of diverse semantics by analyzing the dialogue context. A novel reward function with negotiation strategy congruence, politeness, dialogue coherence, and engagingness rewards is designed to guide the policy’s learning for generating responses. Automatic and human evaluations on NEGOCHAT and IND datasets validate the effectiveness of GENTEEL-NEGOTIATOR in generating polite responses during negotiation while maintaining conversation goals, including coherence and engagingness. Priyanshu Priya, Rishikant Chigrupaatii, Mauajama Firdaus, Asif Ekbal |
AAAI | 3 |
| 2025 | Unmasking offensive content: a multimodal approach with emotional understanding
Gopendra Vikram Singh, Soumitra Ghosh, Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya |
Multim. Tools Appl. | 3 |
| 2024 | Well, Now We Know! Unveiling Sarcasm: Initiating and Exploring Multimodal Conversations with ReasoningabstractSarcasm is a widespread linguistic phenomenon that poses a considerable challenge to explain due to its subjective nature, absence of contextual cues, and rooted personal perspectives. Even though the identification of sarcasm has been extensively studied in dialogue analysis, merely detecting sarcasm falls short of enabling conversational systems to genuinely comprehend the underlying meaning of a conversation and generate fitting responses. It is imperative to not only detect sarcasm but also pinpoint its origination and the rationale behind the sarcastic expressions to capture its authentic essence. In this paper, we delve into the discourse structure of conversations infused with sarcasm and introduce a novel task - Sarcasm Initiation and Reasoning in Conversations (SIRC). Embedded in a multimodal environment and involving a combination of both English and code-mixed interactions, the objective of the task is to discern the trigger or starting point of sarcasm. Additionally, the task involves producing a natural language explanation that rationalizes the satirical dialogues. To this end, we introduce Sarcasm Initiation and Reasoning Dataset (SIRD) to facilitate our task and provide sarcasm initiation annotations and reasoning. We develop a comprehensive model named Sarcasm Initiation and Reasoning Generation (SIRG), which is designed to encompass textual, audio, and visual representations. To achieve this, we introduce a unique shared fusion method that employs cross-attention mechanisms to seamlessly integrate these diverse modalities. Our experimental outcomes, conducted on the SIRC dataset, demonstrate that our proposed framework establishes a new benchmark for both sarcasm initiation and its reasoning generation in the context of multimodal conversations. The code and dataset can be accessed from https://www.iitp.ac.in/∼ai-nlp-ml resources.html#sarcasm-explain and https://github.com/GussailRaat/SIRG-Sarcasm-Initiation-and-Reasoning-Generation. Gopendra Vikram Singh, Mauajama Firdaus, Dushyant Singh Chauhan, Asif Ekbal, Pushpak Bhattacharyya |
AAAI | 2 |
| 2024 | Claim-Centric and Sentiment Guided Graph Attention Network for Rumour DetectionabstractAutomatic rumour detection has gained attention due to the influence of social media on individuals and its pervasiveness. In this work, we construct a representation that takes into account the claim in the source tweet, considering both the propagation graph and the accompanying text alongside tweet sentiment. This is achieved through the implementation of a hierarchical attention mechanism, which not only captures the embedding of documents from individual word vectors but also combines these document representations as nodes within the propagation graph. Furthermore, to address potential overfitting concerns, we employ generative models to augment the existing datasets. This involves rephrasing the claims initially made in the source tweet, thereby creating a more diverse and robust dataset. In addition, we augment the dataset with sentiment labels to improve the performance of the rumour detection task. This holistic and refined approach yields a significant enhancement in the performance of our model across three distinct datasets designed for rumour detection. Quantitative and qualitative analysis proves the effectiveness of our methodology, surpassing the achievements of prior methodologies. Sajad Ramezani, Mauajama Firdaus, Lili Mou |
LREC/COLING | 2 |
| 2024 | Affective Computing for Social Good Applications: Current Advances, Gaps and Opportunities in Conversational Setting
Priyanshu Priya, Mauajama Firdaus, Gopendra Vikram Singh, Asif Ekbal |
ECIR (5) | 2 |
| 2024 | Deciphering Cognitive Distortions in Patient-Doctor Mental Health Conversations: A Multimodal LLM-Based Detection and Reasoning FrameworkabstractCognitive distortion research holds increasing significance as it sheds light on pervasive errors in thinking patterns, providing crucial insights into mental health challenges and fostering the development of targeted interventions and therapies.This paper delves into the complex domain of cognitive distortions which are prevalent distortions in cognitive processes often associated with mental health issues.Focusing on patient-doctor dialogues, we introduce a pioneering method for detecting and reasoning about cognitive distortions utilizing Large Language Models (LLMs).Operating within a multimodal context encompassing audio, video, and textual data, our approach underscores the critical importance of integrating diverse modalities for a comprehensive understanding of cognitive distortions.By leveraging multimodal information, including audio, video, and textual data, our method offers a nuanced perspective that enhances the accuracy and depth of cognitive distortion detection and reasoning in a zero-shot manner.Our proposed hierarchical framework adeptly tackles both detection and reasoning tasks, showcasing significant performance enhancements compared to current methodologies.Through comprehensive analysis, we elucidate the efficacy of our approach, offering promising insights into the diagnosis and understanding of cognitive distortions in multimodal settings.The code and dataset can be found here: https://www.iitp.ac. in/~ai-nlp-ml/resources.html#ZS-CoDR.CoD Reasoning: The patient's final words, "They are always commenting on everything that I'm doing."could be seen as an example of cognitive distortion.This distortion occurs when someone assumes others are always scrutinizing them, despite lacking evidence.The patient's belief that others are constantly monitoring and critiquing their actions is exaggerated and unsupported, demonstrating a distorted perception of external attention.D: Ok, ok.And can you hear what they are actually saying?P: Yeah, they are talking about me. Gopendra Vikram Singh, Sai Vemulapalli, Mauajama Firdaus, Asif Ekbal |
EMNLP | 3 |
| 2024 | Two in One: A multi-task framework for politeness turn identification and phrase extraction in goal-oriented conversations
Priyanshu Priya, Mauajama Firdaus, Asif Ekbal |
Comput. Speech Lang. | 2 |
| 2024 | Zero-shot multitask intent and emotion prediction from multimodal data: A benchmark study
Gopendra Vikram Singh, Mauajama Firdaus, Dushyant Singh Chauhan, Asif Ekbal, Pushpak Bhattacharyya |
Neurocomputing | 2 |
| 2024 | A Unified Framework for Slot based Response Generation in a Multimodal Dialogue System
Mauajama Firdaus, Avinash Madasu, Asif Ekbal |
Multim. Tools Appl. | 1 |
| 2024 | Please Donate to Save a Life: Inducing Politeness to Handle Resistance in Persuasive Dialogue AgentsabstractIn a persuasive conversation forsocial good, even the most compelling and persuasive argument may fail to persuade a persuadee resisting the persuasion. Whereas use of polite tone, apologetic expressions, or deferential modes of reference such as, ‘thank you’, ‘Please’ etc. can make the conversation more interesting, engaging, and persuading to the persuadee. We propose a resistance handling polite persuasive dialogue system (Re-Po-PDS), harnessing an efficient reward function consisting of persuasiveness, politeness, coherence, and non-repetitiveness rewards in a reinforcement learning framework and a politeness transfer model. Due to the lack of polite annotated persuasive data, we first annotate thePersuaionForGooddataset with different politeness labels and name it as PP4G dataset. Then we train two transformer-based persuasive and politeness classifiers to receive persuasive and politeness feedback for our RL-agent. Further, we train a politeness transfer model which is used at the inference time as per persuadee's resistive strategy encountered to form a more polite response. Our experimental results confirm that our proposed model increases the rate of generating polite persuasive responses as compared to the available state-of-the-art dialogue models while also making the dialogues more engaging and retaining. Kshitij Mishra, Mauajama Firdaus, Asif Ekbal |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | XeroPol: Emotion-Aware Contrastive Learning for Zero-Shot Cross-Lingual Politeness Identification in DialoguesabstractPoliteness is key to successful conversations. It depicts the behavior that is socially valued and is often accompanied by emotions. Previously, researchers have focused on detecting politeness in goal-oriented conversations in high-resource English language. The existing studies do not focus on identifying politeness in a resource-scared Indian languages such as Hindi, primarily due to the lack of labeled data. To overcome this limitation, in this article, we propose a novel emotion-aware contrastive learning (CL) method for zero-shot cross-lingual politeness identification (XeroPol) task in dialogues. We introduceContrastiveAligner, a CL-based alignment method for zero-shot cross-lingual transfer.ContrastiveAligneremploys translated data and pushes the model to generate similar utterance embeddings for different languages. As politeness and emotion are interrelated, hence, as the conversation progresses, the variation in emotions tends to pose challenges in identifying politeness in dialogues. Thus, in this work, we also design an auxiliary emotion-aware CL objective using sentiment information, namely theEmoSenti objective, which is expected to implicitly model the emotion change across utterances and help in the primary task of politeness identification. Experiments on MultiDoGo and EmoWOZ datasets demonstrate that the proposed approach significantly outperforms the baselines. Further analysis such as human evaluation on the EmoInHindi dataset validates the efficacy of the entire approach. Priyanshu Priya, Mauajama Firdaus, Asif Ekbal |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2023 | Multi-step Prompting for Few-shot Emotion-Grounded ConversationsabstractConversational systems have shown immense growth in their ability to communicate like humans. With the emergence of large pre-trained language models (PLMs) the ability to provide informative responses have improved significantly. Despite the success of PLMs, the ability to identify and generate engaging and empathetic responses is largely dependent on labelled-data. In this work, we design a prompting approach that identifies the emotion of a given utterance and uses the emotion information for generating the appropriate responses for conversational systems. We propose a two-step prompting method that first recognises the emotion in the dialogue utterance and in the second-step uses the predicted emotion to prompt the PLM to generate the corresponding em- pathetic response in a few-shot setting. Experimental results on three publicly available datasets show that our proposed approach outperforms the state-of-the-art approaches for both automatic and manual evaluation. Mauajama Firdaus, Gopendra Vikram Singh, Asif Ekbal, Pushpak Bhattacharyya |
CIKM | 1 |
| 2023 | Weakly Supervised Explainable Phrasal Reasoning with Neural Fuzzy Logic
Zijun Wu 0002, Zi Xuan Zhang, Atharva Naik, Zhijian Mei, Mauajama Firdaus, Lili Mou |
ICLR | 5 |
| 2023 | A multi-task learning framework for politeness and emotion detection in dialogues for mental health counselling and legal aid
Priyanshu Priya, Mauajama Firdaus, Asif Ekbal |
Expert Syst. Appl. | 2 |
| 2023 | Affect-GCN: a multimodal graph convolutional network for multi-emotion with intensity recognition and sentiment analysis in dialogues
Mauajama Firdaus, Gopendra Vikram Singh, Asif Ekbal, Pushpak Bhattacharyya |
Multim. Tools Appl. | 1 |
| 2023 | I Enjoy Writing and Playing, Do You?: A Personalized and Emotion Grounded Dialogue Agent Using Generative Adversarial NetworkabstractSocial chatbots have gained immense popularity, and their appeal lies in their capacity to respond to diverse requests, but also in their ability to develop an emotional connection with users. To develop and promote social chatbots, we need to concentrate on increasing user interaction and consider both the intellectual and emotional quotient in conversational agents. In this work, we propose the task of empathetic, personalized dialogue generation giving the machine the capability to respond emotionally and in accordance with the persona of the user. We design a generative adversarial framework, named EP-GAN (Empathy and Persona aware Generative Adversarial Network) to generate responses that are sensitive to the emotion of the user and corresponds to the persona. The persona information is encoded as memory vectors that, along with the dialogue history, are fed to the decoder for generation. An interactive adversarial learning framework is implemented to verify whether the generated responses elicit the emotion and persona in dialogues. Experimental results show that the EP-GAN framework significantly outperforms the existing baselines for both automatic and manual evaluation. We achieve an improvement of 1 point BLEU-4 and 2 points in the emotion accuracy metric compared to the best performing baseline. Mauajama Firdaus, Naveen Thangavelu, Asif Ekbal, Pushpak Bhattacharyya |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | EmoInt-Trans: A Multimodal Transformer for Identifying Emotions and Intents in Social ConversationsabstractIn the natural language processing community, open-domain conversational agents, also known as chatbots, are gaining popularity. One of the difficulties is getting them to communicate in an emotionally intelligent manner. To generate dialogues, current neural response generation methods depend solely on end-to-end learning from large scale conversation data. Therefore, we introduce a large-scale multi Emotion and Intent guided Multimodal Dialogue (EmoInt-MD) dataset labelled with 32 emotions and 15 empathetic intents having 32 k dialogues taken from different movie genres. We propose a novel multi-task multimodal contextual Transformer framework for simultaneously identifying the emotions and intents in a given utterance utilizing audio and visual features in addition to the textual information. Experimental analysis proves that the proposed framework outperforms several unimodal and multimodal baselines on theEmoInt-MDdataset. This dataset along with our baseline and proposed framework implementations will be made publicly available for research purposes. Gopendra Vikram Singh, Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Being Polite: Modeling Politeness Variation in a Personalized Dialog AgentabstractPoliteness enhances interactions by improving relations between the participants. If there is a display of rudeness, even the finest conversation can fall through. In addition, if lathered with kindness, even the most angst-prone circumstance can be expressed with far less suffering. Previously, researchers have focused upon including politeness in conversations. But the existing research does not focus on variations in politeness according to the user profile. Therefore, in this article, we propose a novel task of generating polite personalized dialog responses in accordance with the user profile and consistent with the conversational history. We design a novel Polite Personalized Dialog Generation (PoPe-DG) framework that employs a reinforced deliberation network. We create human-annotated politeness templates according to user profiles to induce politeness variation in the generated responses for the proposed task. Precisely, the personality profile is transformed and normalized into a vector using the fusion attention combined with dialogue utterances to build context. Furthermore, the context modules and the annotated templates are appended to initialize the deliberation decoder. Experimental analysis validates that our proposed approach inculcates politeness in responses in accordance with the user profile and the conversational history. Mauajama Firdaus, Arunav Pratap Shandilya, Asif Ekbal, Pushpak Bhattacharyya |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | Predicting Politeness Variations in Goal-Oriented ConversationsabstractPoliteness is an expression in language that eases the conversation toward a positive undertone. If there is a display of rudeness, even the finest communication can fall through. In addition, if lathered with politeness, even the most angst-prone scenario can be expressed with far less hurt. In this article, we address the task of identifying politeness in goal-oriented dialog systems. In this regard, we create politeness-annotated conversational data (PACD) utilizing Microsoft Dialogue Challenge and DSTC1 datasets. For correctly identifying the politeness, we employ a hierarchical transformer network that effectively captures the contextual information (i.e., previous utterances) and current input for predicting the politeness in a given utterance of a dialog. Empirical results demonstrate that our proposed approach outperforms all the defined baselines. Furthermore, through in- and cross-domain experiments, we show the necessity of a PACD to mitigate acts such as rude requests or insults for both socially interactive and task-oriented dialog systems. Kshitij Mishra, Mauajama Firdaus, Asif Ekbal |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2023 | Unity in Diversity: Multilabel Emoji Identification in TweetsabstractEmojis or emoticons are not just a modern trend but have become an essential part of our day-to-day interactions. Predicting a suitable emoji for a given tweet is a challenging task because a wrong emoji prediction for a tweet can change the meaning of the message or can amplify the emotion of the message. This task is particularly challenging since it requires selecting an appropriate emoji from a huge list of prospective emojis that may or may not be equivalent to one another. Humans use multiple emojis to convey their emotions, thereby making the task a multilabel classification problem. In this article, we propose a multilabel emoji prediction system that predicts the appropriate emoji for a given tweet by using different state-of-the-art baselines. Due to the unavailability of a multi-emoji dataset, we create a large-scale multilabel emoji dataset named Mu-Emoji that comprises of more than 0.6 million tweets having varied emojis belonging to both positive and negative sentiments. For our proposed task, we employ graph attention network along with bidirectional encoder representations from transformer encoder for the accurate prediction of emojis. Qualitative and quantitative analyses show that our multilabel emoji prediction baselines perform well compared with the single-emoji prediction baselines for our proposed Mu-Emoji dataset. Our proposed framework also outperforms all the baselines for both single and multilabel emoji prediction tasks. Gopendra Vikram Singh, Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2022 | PoliSe: Reinforcing Politeness Using User Sentiment for Customer Care Response GenerationabstractThe interaction between a consumer and the customer service representative greatly contributes to the overall customer experience. Therefore, to ensure customers’ comfort and retention, it is important that customer service agents and chatbots connect with users on social, cordial, and empathetic planes. In the current work, we automatically identify the sentiment of the user and transform the neutral responses into polite responses conforming to the sentiment and the conversational history. Our technique is basically a reinforced multi-task network- the primary task being ‘polite response generation’ and the secondary task being ‘sentiment analysis’- that uses a Transformer based encoder-decoder. We use sentiment annotated conversations from Twitter as the training data. The detailed evaluation shows that our proposed approach attains superior performance compared to the baseline models. Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya |
COLING | 1 |
| 2022 | Sentiment Guided Aspect Conditioned Dialogue Generation in a Multimodal System
Mauajama Firdaus, Nidhi Thakur, Asif Ekbal |
ECIR (1) | 1 |
| 2022 | EmoInHindi: A Multi-label Emotion and Intensity Annotated Dataset in Hindi for Emotion Recognition in DialoguesabstractThe long-standing goal of Artificial Intelligence (AI) has been to create human-like conversational systems. Such systems should have the ability to develop an emotional connection with the users, consequently, emotion recognition in dialogues has gained popularity. Emotion detection in dialogues is a challenging task because humans usually convey multiple emotions with varying degrees of intensities in a single utterance. Moreover, emotion in an utterance of a dialogue may be dependent on previous utterances making the task more complex. Recently, emotion recognition in low-resource languages like Hindi has been in great demand. However, most of the existing datasets for multi-label emotion and intensity detection in conversations are in English. To this end, we propose a large conversational dataset in Hindi named EmoInHindi for multi-label emotion and intensity recognition in conversations containing 1,814 dialogues with a total of 44,247 utterances. We prepare our dataset in a Wizard-of-Oz manner for mental health and legal counselling of crime victims. Each utterance of dialogue is annotated with one or more emotion categories from 16 emotion labels including neutral and their corresponding intensity. We further propose strong contextual baselines that can detect the emotion(s) and corresponding emotional intensity of an utterance given the conversational context. Gopendra Vikram Singh, Priyanshu Priya, Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya |
LREC | 3 |
| 2022 | Are Emoji, Sentiment, and Emotion Friends? A Multi-task Learning for Emoji, Sentiment, and Emotion Analysis
Gopendra Vikram Singh, Dushyant Singh Chauhan, Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya |
PACLIC | 3 |
| 2022 | Please be polite: Towards building a politeness adaptive dialogue system for goal-oriented conversations
Kshitij Mishra, Mauajama Firdaus, Asif Ekbal |
Neurocomputing | 2 |
| 2022 | Knowing What to Say: Towards knowledge grounded code-mixed response generation for open-domain conversations
Gopendra Vikram Singh, Mauajama Firdaus, Shambhavi, Asif Ekbal |
Knowl. Based Syst. | 2 |
| 2022 | EmoSen: Generating Sentiment and Emotion Controlled Responses in a Multimodal Dialogue SystemabstractAn essential skill for effective communication is the ability to express specific sentiment and emotion in a conversation. Any robust dialogue system should handle the combined effect of both sentiment and emotion while generating responses. This is expected to provide a better experience and concurrently increase users’ satisfaction. Previously, research on either emotion or sentiment controlled dialogue generation has shown great promise in developing the next generation conversational agents, but the simultaneous effect of both is still unexplored. The existing dialogue systems are majorly based on unimodal sources, predominantly the text, and thereby cannot utilize the information present in the other sources, such as video, audio, image, etc. In this article, we present at first a large scale benchmark Sentiment Emotion aware Multimodal Dialogue (SEMD) dataset for the task of sentiment and emotion controlled dialogue generation. The SEMD dataset consists of 55k conversations from 10 TV shows having text, audio, and video information. To utilize multimodal information, we propose multimodal attention based conditional variational autoencoder (M-CVAE) that outperforms several baselines. Quantitative and qualitative analyses show that multimodality, along with contextual information, plays an essential role in generating coherent and diverse responses for any given emotion and sentiment. Mauajama Firdaus, Hardik Chauhan, Asif Ekbal, Pushpak Bhattacharyya |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | More the Merrier: Towards Multi-Emotion and Intensity Controllable Response GenerationabstractThe focus on conversational systems has recently shifted towards creating engaging agents by inculcating emotions into them. Human emotions are highly complex as humans can express multiple emotions with varying intensity in a single utterance, whereas the conversational agents convey only one emotion in their responses. To infuse human-like behaviour in the agents, we introduce the task of multi-emotion controllable response generation with the ability to express different emotions with varying levels of intensity in an open-domain dialogue system. We introduce a Multiple Emotion Intensity aware Multi-party Dialogue (MEIMD) dataset having 34k conversations taken from 8 different TV Series. We finally propose a Multiple Emotion with Intensity-based Dialogue Generation (MEI-DG) framework. The system employs two novel mechanisms: viz. (i) determining the trade-off between the emotion and generic words, while focusing on the intensity of the desired emotions; and (ii) computing the amount of emotion left to be expressed, thereby regulating the generation accordingly. The detailed evaluation shows that our proposed approach attains superior performance compared to the baseline models. Mauajama Firdaus, Hardik Chauhan, Asif Ekbal, Pushpak Bhattacharyya |
AAAI | 1 |
| 2021 | Attribute Centered Multimodal Response Generation in a Dialogue SystemabstractDialogue system has become a prominent platform for human-machine interactions. The ongoing research in vision and language has opened new frontiers for building multimodal dialogue systems that incorporate information from the various complementary sources such as text, images, audios, and videos. For interactive systems, multimodal knowledge in the form of different modalities need to be presented to the user for effective communication. Recent research has mainly focused upon textual generation in a multimodal setting, thereby providing incomplete information to the users. Hence, in our current work, we present an attribute centered image generation framework in a multimodal system that is capable of generating images, conditioned upon the textual features with the help of the taxonomy-attribute combined tree. The visual features interact with multi-head attended textual features through attention based factorized bilinear pooling approach for fine-grained representation. Further, the multimodal representation is extended to generate images using the adversarial network. The loss functions encourage the network to generate images that are very similar to the natural image. We perform our experiments on the Multimodal Dialog Dataset (MMD) to create contextualized images using the fashion attribute features. Empirical studies show that the generated attribute centered images help in making the dialogue more engaging to the users. Mauajama Firdaus, Arunav Pratap Shandilya, Sthita Pragyan Pujari, Asif Ekbal |
IJCNN | 1 |
| 2021 | Multi-Aspect Controlled Response Generation in a Multimodal Dialogue System using Hierarchical Transformer NetworkabstractMultimodality in dialogues has become crucial for a thorough understanding of the intent of the user to provide better responses to fulfill user's demands. Existing dialogue systems suffer from the issue of inconsistency and dull responses. Many goal-oriented conversational systems lack the different aspect information of the products or services to present an informative and exciting response. Aspects such as the price, color, pattern, rating are essential for deciding whether to purchase/order. To alleviate the issues in the existing systems, we propose the novel task of multi-aspect guided dialogue generation. This task is introduced to focus on making the responses focused and consistent with the different aspects mentioned in the current dialogue. In our present work, we design a hierarchical transformer network to capture the dialogue context for generating the responses. For creating the responses with multiple aspects, we explicitly give the aspect vectors at the time of decoding for generation. The information of the aspects specified to the decoder controls the overall generation process. We evaluate our proposed hierarchical framework on the newly created multi-domain multi-modal dialogue (MDMMD) dataset consisting of both text and images. Experimental results show that the proposed system outperforms all the existing and baseline approaches. The aspect controlled responses are immensely consistent with the ongoing dialogue and highly diverse resolving the issues of the current systems. Mauajama Firdaus, Nidhi Thakur, Asif Ekbal |
IJCNN | 1 |
| 2021 | SEPRG: Sentiment aware Emotion controlled Personalized Response GenerationabstractSocial chatbots have gained immense popularity, and their appeal lies not just in their capacity to respond to the diverse requests from users, but also in the ability to develop an emotional connection with users.To further develop and promote social chatbots, we need to concentrate on increasing user interaction and take into account both the intellectual and emotional quotient in the conversational agents.Therefore, in this work, we propose the task of sentiment aware emotion controlled personalized dialogue generation giving the machine the capability to respond emotionally and in accordance with the persona of the user.As sentiment and emotions are highly co-related, we use the sentiment knowledge of the previous utterance to generate the correct emotional response in accordance with the user persona.We design a Transformer based Dialogue Generation framework, that generates responses that are sensitive to the emotion of the user and corresponds to the persona and sentiment as well.Moreover, the persona information is encoded by a different Transformer encoder, along with the dialogue history, is fed to the decoder for generating responses.We annotate the PersonaChat dataset with sentiment information to improve the response quality.Experimental results on the PersonaChat dataset show that the proposed framework significantly outperforms the existing baselines, thereby generating personalized emotional responses in accordance with the sentiment that provides better emotional connection and user satisfaction as desired in a social chatbot. Mauajama Firdaus, Umang Jain, Asif Ekbal, Pushpak Bhattacharyya |
INLG | 1 |
| 2021 | Aspect-Aware Response Generation for Multimodal Dialogue SystemabstractMultimodality in dialogue systems has opened up new frontiers for the creation of robust conversational agents. Any multimodal system aims at bridging the gap between language and vision by leveraging diverse and often complementary information from image, audio, and video, as well as text. For every task-oriented dialog system, different aspects of the product or service are crucial for satisfying the user’s demands. Based upon the aspect, the user decides upon selecting the product or service. The ability to generate responses with the specified aspects in a goal-oriented dialogue setup facilitates user satisfaction by fulfilling the user’s goals. Therefore, in our current work, we propose the task of aspect controlled response generation in a multimodal task-oriented dialog system. We employ a multimodal hierarchical memory network for generating responses that utilize information from both text and images. As there was no readily available data for building such multimodal systems, we create a Multi-Domain Multi-Modal Dialog (MDMMD++) dataset. The dataset comprises the conversations having both text and images belonging to the four different domains, such as hotels, restaurants, electronics, and furniture. Quantitative and qualitative analysis on the newly created MDMMD++ dataset shows that the proposed methodology outperforms the baseline models for the proposed task of aspect controlled response generation. Mauajama Firdaus, Nidhi Thakur, Asif Ekbal |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2020 | MEISD: A Multimodal Multi-Label Emotion, Intensity and Sentiment Dialogue Dataset for Emotion Recognition and Sentiment Analysis in ConversationsabstractEmotion and sentiment classification in dialogues is a challenging task that has gained popularity in recent times.Humans tend to have multiple emotions with varying intensities while expressing their thoughts and feelings.Emotions in an utterance of dialogue can either be independent or dependent on the previous utterances, making the task complex and interesting.Multi-label emotion detection in conversations is a significant task that provides the ability to the system to understand the various emotions of the users interacting.On the other hand, sentiment analysis in dialogue or conversation helps in understanding the perspective of the user with respect to the ongoing conversation.Besides text, additional information in the form of audio and video assists in identifying the correct emotions with the appropriate intensity and sentiments in an utterance of a dialogue.Lately, quite a few datasets have been made available for emotion and sentiment classification in dialogues.Still, these datasets are imbalanced in representing different emotions and consist of only a single emotion.Hence, we present at first a large-scale balanced Multimodal Multi-label Emotion, Intensity, and Sentiment Dialogue dataset (MEISD) collected from different TV series that has textual, audio, and visual features, and then establish a baseline setup for further research. Mauajama Firdaus, Hardik Chauhan, Asif Ekbal, Pushpak Bhattacharyya |
COLING | 1 |
| 2020 | Persona aware Response Generation with EmotionsabstractConversational systems are the perfect examples of human-machine interactions. The conversational agents while interacting with humans lack the ability to express emotions and behave inconsistently, making the conversations boring and non-interactive. In this work, we propose the task of persona aware emotional response generation in which the system can generate specific and consistent responses in accordance to the provided personality information and the conversational history. To make the responses interactive and interesting we intend to infuse the emotions in the responses that help in making the responses more human-like. We propose a persona aware attention framework employing an encoder-decoder approach. We investigate different ways to include the desired emotions in the responses. Experimental results on the PersonaChat dataset shows that our proposed framework outperforms the baseline models and can generate interactive and emotional responses. Mauajama Firdaus, Naveen Thangavelu, Asif Ekbal, Pushpak Bhattacharyya |
IJCNN | 1 |
| 2020 | Incorporating Politeness across Languages in Customer Care Responses: Towards building a Multi-lingual Empathetic Dialogue AgentabstractCustomer satisfaction is an essential aspect of customer care systems. It is imperative for such systems to be polite while handling customer requests/demands. In this paper, we present a large multi-lingual conversational dataset for English and Hindi. We choose data from Twitter having both generic and courteous responses between customer care agents and aggrieved users. We also propose strong baselines that can induce courteous behaviour in generic customer care response in a multi-lingual scenario. We build a deep learning framework that can simultaneously handle different languages and incorporate polite behaviour in the customer care agent’s responses. Our system is competent in generating responses in different languages (here, English and Hindi) depending on the customer’s preference and also is able to converse with humans in an empathetic manner to ensure customer satisfaction and retention. Experimental results show that our proposed models can converse in both the languages and the information shared between the languages helps in improving the performance of the overall system. Qualitative and quantitative analysis shows that the proposed method can converse in an empathetic manner by incorporating courteousness in the responses and hence increasing customer satisfaction. Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya |
LREC | 1 |
| 2019 | Ordinal and Attribute Aware Response Generation in a Multimodal Dialogue SystemabstractMultimodal dialogue systems have opened new frontiers in the traditional goal-oriented dialogue systems. The state-of-the-art dialogue systems are primarily based on unimodal sources, predominantly the text, and hence cannot capture the information present in the other sources such as videos, audios, images etc. With the availability of large scale multimodal dialogue dataset (MMD) (Saha et al., 2018) on the fashion domain, the visual appearance of the products is essential for understanding the intention of the user. Without capturing the information from both the text and image, the system will be incapable of generating correct and desirable responses. In this paper, we propose a novel position and attribute aware attention mechanism to learn enhanced image representation conditioned on the user utterance. Our evaluation shows that the proposed model can generate appropriate responses while preserving the position and attribute information. Experimental results also prove that our proposed approach attains superior performance compared to the baseline models, and outperforms the state-of-the-art approaches on text similarity based evaluation metrics. Hardik Chauhan, Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya |
ACL (1) | 2 |
| 2019 | Exploring Machine Learning and Deep Learning Frameworks for Task-Oriented Dialogue Act ClassificationabstractDialogue Act (DA) Classification plays a signifi-cant role in the understanding of an utterance in a dialogue. Components of Spoken Dialogue System (SDS) such as Natural Language Understanding (NLU) and Dialogue Management (DM) modules can significantly exploit the output of the DA classification. In this paper, we propose a task-oriented DA classifier based on both traditional supervised Machine Learning (ML) as well as Deep Learning (DL) techniques. The type and nature of dialogues basically depend on the domain and the DA itself. So, in order to make the model task-oriented, a new tag-set has been designed by studying the properties of the target domain. On the benchmark SwitchBoard (SWBD) and TRAINS corpus, our proposed models have performed exceptionally well with the new tag-set. Experimental results indicate that our proposed models have achieved good accuracy on both the datasets and outperformed several state of the art approaches and the new tag-set is well suited for task-oriented applications. Tulika Saha, Mauajama Firdaus, Sriparna Saha 0001, Asif Ekbal, Pushpak Bhattacharyya |
IJCNN | 3 |
| 2019 | A Multi-Task Hierarchical Approach for Intent Detection and Slot Filling
Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya |
Knowl. Based Syst. | 1 |
| 2018 | A Deep Learning Based Multi-task Ensemble Model for Intent Detection and Slot Filling in Spoken Language Understanding
Mauajama Firdaus, Shobhit Bhatnagar, Asif Ekbal, Pushpak Bhattacharyya |
ICONIP (4) | 1 |
| 2018 | Intent Detection for Spoken Language Understanding Using a Deep Ensemble Model
Mauajama Firdaus, Shobhit Bhatnagar, Asif Ekbal, Pushpak Bhattacharyya |
PRICAI (1) | 1 |