Gopendra Vikram Singh

dblp:258/4363 · DBLP profile ↗
← Back
27ranked-venue papers
14as first author
27since 2021 · last 2026
0000-0003-1104-5856ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 10 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A safety-aware deep learning framework for scalable schedulability analysis of variable-length real-time task sets
Lalatendu Behera, Gopendra Vikram Singh
Real Time Syst.2
2025 Just a Scratch: Enhancing LLM Capabilities for Self-harm Detection through Intent Differentiation and Emoji Interpretation
abstract
Soumitra Ghosh, Gopendra Vikram Singh, Shambhavi Shambhavi, Sabarna Choudhury, Asif Ekbal. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Soumitra Ghosh, Gopendra Vikram Singh, Shambhavi, Sabarna Choudhury, Asif Ekbal
ACL (1)2
2025 What is Beneath Misogyny: Misogynous Memes Classification and Explanation
abstract
Memes are popular in the modern world and are distributed primarily for entertainment. However, harmful ideologies such as misogyny can be propagated through innocent-looking memes. The detection and understanding of why a meme is misogynous is a research challenge due to its multimodal nature (image and text) and its nuanced manifestations across different societal contexts. We introduce a novel multimodal approach, namely, MM-Misogyny to detect, categorize, and explain misogynistic content in memes. MM-Misogyny processes text and image modalities separately and unifies them into a multimodal context through a cross-attention mechanism. The resulting multimodal context is then easily processed for labeling, categorization, and explanation via a classifier and Large Language Model (LLM). The evaluation of the proposed model is performed on a newly curated dataset (What’s Beneath Misogynous Stereotyping (WBMS)) created by collecting misogynous memes from cyberspace and categorizing them into four categories, namely, Kitchen, Leadership, Working, and Shopping. The model not only detects and classifies misogyny, but also provides a granular understanding of how misogyny operates in operates in domains of life. The results demonstrate the superiority of our approach compared to existing methods. The code and dataset are available at https://github.com/Misogyny.
Kushal Kanwar, Dushyant Singh Chauhan, Gopendra Vikram Singh, Asif Ekbal
IJCAI3
2025 Unmasking offensive content: a multimodal approach with emotional understanding
Gopendra Vikram Singh, Soumitra Ghosh, Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya
Multim. Tools Appl.1
2025 MultiSEAO-Mix: A Multimodal Multitask Framework for Sentiment, Emotion, Support, and Offensive Analysis in Code-Mixed Setting
abstract
Social media platforms have become an open door for users to share their views, resulting in a growing trend of offensive content being shared on social media. Detecting and addressing offensive content is crucial due to its significant impact on society. Although there has been extensive research on the detection of offensive content in the English language, there is a notable gap in detecting offensive content in multimodal settings involving code-mixed languages. In this article, we propose a large scale multimodal code-mixed dataset for Hinglish (Hindi+English)MultiSEAO-Mixfocusing on women and children. TheMultiSEAO-Mixis annotated with offensiveness, sentiment, emotion, and their respective intensities. Additionally, it is also annotated with author support. A multimodal, multitask framework is proposed that considers offensive detection, intensity prediction, and author support as the primary tasks and improves their performance using sentiment, emotion, and corresponding intensities as the auxiliary tasks. Further, we propose a fusion technique that captures the enhanced multimodal representation to improve the performance of our model. Experimental results demonstrate that the proposed multitask framework improves the model performance by more than 4.5 points compared to multitask system without sentiment and emotion as the auxiliary tasks.
Gopendra Vikram Singh, Mamta, Atul Verma, Asif Ekbal
IEEE Trans. Comput. Soc. Syst.1
2024 Benchmarking Cyber Harassment Dialogue Comprehension through Emotion-Informed Manifestations-Determinants Demarcation
abstract
In the digital age, cybercrimes, particularly cyber harassment, have become pressing issues, targeting vulnerable individuals like children, teenagers, and women. Understanding the experiences and needs of the victims is crucial for effective support and intervention. Online conversations between victims and virtual harassment counselors (chatbots) offer valuable insights into cyber harassment manifestations (CHMs) and determinants (CHDs). However, the distinction between CHMs and CHDs remains unclear. This research is the first to introduce concrete definitions for CHMs and CHDs, investigating their distinction through automated methods to enable efficient cyber-harassment dialogue comprehension. We present a novel dataset, Cyber-MaD that contains Cyber harassment dialogues manually annotated with Manifestations and Determinants. Additionally, we design an Emotion-informed Contextual Dual attention Convolution Transformer (E-ConDuCT) framework to extract CHMs and CHDs from cyber harassment dialogues. The framework primarily: a) utilizes inherent emotion features through adjective-noun pairs modeled by an autoencoder, b) employs a unique Contextual Dual attention Convolution Transformer to learn contextual insights; and c) incorporates a demarcation module leveraging task-specific emotional knowledge and a discriminator loss function to differentiate manifestations and determinants. E-ConDuCT outperforms the state-of-the-art systems on the Cyber-MaD corpus, showcasing its potential in the extraction of CHMs and CHDs. Furthermore, its robustness is demonstrated on the emotion cause extraction task using the CARES_CEASE-v2.0 dataset of suicide notes, confirming its efficacy across diverse cause extraction objectives. Access the code and data at 1. https://www.iitp.ac.in/~ai-nlp-ml/resources.html#E-ConDuCT-on-Cyber-MaD, 2. https://github.com/Soumitra816/Manifestations-Determinants.
Soumitra Ghosh, Gopendra Vikram Singh, Jashn Arora, Asif Ekbal
AAAI2
2024 Well, Now We Know! Unveiling Sarcasm: Initiating and Exploring Multimodal Conversations with Reasoning
abstract
Sarcasm is a widespread linguistic phenomenon that poses a considerable challenge to explain due to its subjective nature, absence of contextual cues, and rooted personal perspectives. Even though the identification of sarcasm has been extensively studied in dialogue analysis, merely detecting sarcasm falls short of enabling conversational systems to genuinely comprehend the underlying meaning of a conversation and generate fitting responses. It is imperative to not only detect sarcasm but also pinpoint its origination and the rationale behind the sarcastic expressions to capture its authentic essence. In this paper, we delve into the discourse structure of conversations infused with sarcasm and introduce a novel task - Sarcasm Initiation and Reasoning in Conversations (SIRC). Embedded in a multimodal environment and involving a combination of both English and code-mixed interactions, the objective of the task is to discern the trigger or starting point of sarcasm. Additionally, the task involves producing a natural language explanation that rationalizes the satirical dialogues. To this end, we introduce Sarcasm Initiation and Reasoning Dataset (SIRD) to facilitate our task and provide sarcasm initiation annotations and reasoning. We develop a comprehensive model named Sarcasm Initiation and Reasoning Generation (SIRG), which is designed to encompass textual, audio, and visual representations. To achieve this, we introduce a unique shared fusion method that employs cross-attention mechanisms to seamlessly integrate these diverse modalities. Our experimental outcomes, conducted on the SIRC dataset, demonstrate that our proposed framework establishes a new benchmark for both sarcasm initiation and its reasoning generation in the context of multimodal conversations. The code and dataset can be accessed from https://www.iitp.ac.in/∼ai-nlp-ml resources.html#sarcasm-explain and https://github.com/GussailRaat/SIRG-Sarcasm-Initiation-and-Reasoning-Generation.
Gopendra Vikram Singh, Mauajama Firdaus, Dushyant Singh Chauhan, Asif Ekbal, Pushpak Bhattacharyya
AAAI1
2024 Affective Computing for Social Good Applications: Current Advances, Gaps and Opportunities in Conversational Setting
Priyanshu Priya, Mauajama Firdaus, Gopendra Vikram Singh, Asif Ekbal
ECIR (5)3
2024 Deciphering Cognitive Distortions in Patient-Doctor Mental Health Conversations: A Multimodal LLM-Based Detection and Reasoning Framework
abstract
Cognitive distortion research holds increasing significance as it sheds light on pervasive errors in thinking patterns, providing crucial insights into mental health challenges and fostering the development of targeted interventions and therapies.This paper delves into the complex domain of cognitive distortions which are prevalent distortions in cognitive processes often associated with mental health issues.Focusing on patient-doctor dialogues, we introduce a pioneering method for detecting and reasoning about cognitive distortions utilizing Large Language Models (LLMs).Operating within a multimodal context encompassing audio, video, and textual data, our approach underscores the critical importance of integrating diverse modalities for a comprehensive understanding of cognitive distortions.By leveraging multimodal information, including audio, video, and textual data, our method offers a nuanced perspective that enhances the accuracy and depth of cognitive distortion detection and reasoning in a zero-shot manner.Our proposed hierarchical framework adeptly tackles both detection and reasoning tasks, showcasing significant performance enhancements compared to current methodologies.Through comprehensive analysis, we elucidate the efficacy of our approach, offering promising insights into the diagnosis and understanding of cognitive distortions in multimodal settings.The code and dataset can be found here: https://www.iitp.ac. in/~ai-nlp-ml/resources.html#ZS-CoDR.CoD Reasoning: The patient's final words, "They are always commenting on everything that I'm doing."could be seen as an example of cognitive distortion.This distortion occurs when someone assumes others are always scrutinizing them, despite lacking evidence.The patient's belief that others are constantly monitoring and critiquing their actions is exaggerated and unsupported, demonstrating a distorted perception of external attention.D: Ok, ok.And can you hear what they are actually saying?P: Yeah, they are talking about me.
Gopendra Vikram Singh, Sai Vemulapalli, Mauajama Firdaus, Asif Ekbal
EMNLP1
2024 From Pink and Blue to a Rainbow Hue! Defying Gender Bias through Gender Neutralizing Text Transformations
Gopendra Vikram Singh, Soumitra Ghosh, Neil Dcruze, Asif Ekbal
IJCAI1
2024 Aspect-Based Multimodal Mining: Unveiling Sentiments, Complaints, and Beyond in User-Generated Content
Mamta, Gopendra Vikram Singh, Deepak Raju Kori, Asif Ekbal
ACM Multimedia2
2024 Zero-shot multitask intent and emotion prediction from multimodal data: A benchmark study
Gopendra Vikram Singh, Mauajama Firdaus, Dushyant Singh Chauhan, Asif Ekbal, Pushpak Bhattacharyya
Neurocomputing1
2023 Multi-step Prompting for Few-shot Emotion-Grounded Conversations
abstract
Conversational systems have shown immense growth in their ability to communicate like humans. With the emergence of large pre-trained language models (PLMs) the ability to provide informative responses have improved significantly. Despite the success of PLMs, the ability to identify and generate engaging and empathetic responses is largely dependent on labelled-data. In this work, we design a prompting approach that identifies the emotion of a given utterance and uses the emotion information for generating the appropriate responses for conversational systems. We propose a two-step prompting method that first recognises the emotion in the dialogue utterance and in the second-step uses the predicted emotion to prompt the PLM to generate the corresponding em- pathetic response in a few-shot setting. Experimental results on three publicly available datasets show that our proposed approach outperforms the state-of-the-art approaches for both automatic and manual evaluation.
Mauajama Firdaus, Gopendra Vikram Singh, Asif Ekbal, Pushpak Bhattacharyya
CIKM2
2023 DeCoDE: Detection of Cognitive Distortion and Emotion Cause Extraction in Clinical Conversations
Gopendra Vikram Singh, Soumitra Ghosh, Asif Ekbal, Pushpak Bhattacharyya
ECIR (2)1
2023 Standardizing Distress Analysis: Emotion-Driven Distress Identification and Cause Extraction (DICE) in Multimodal Online Posts
abstract
Due to its growing impact on public opinion, hate speech on social media has garnered increased attention.While automated methods for identifying hate speech have been presented in the past, they have mostly been limited to analyzing textual content.The interpretability of such models has received very little attention, despite the social and legal consequences of erroneous predictions.In this work, we present a novel problem of Distress Identification and Cause Extraction (DICE) from multimodal online posts.We develop a multi-task deep framework for the simultaneous detection of distress content and identify connected causal phrases from the text using emotional information.The emotional information is incorporated into the training process using a zero-shot strategy, and a novel mechanism is devised to fuse the features from the multimodal inputs.Furthermore, we introduce the first-ofits-kind Distress and Cause annotated Multimodal (DCaM) dataset of 20,764 social media posts.We thoroughly evaluate our proposed method by comparing it to several existing benchmarks.Empirical assessment and comprehensive qualitative analysis demonstrate that our proposed method works well on distress detection and cause extraction tasks, improving F1 and ROS scores by 1.95% and 3%, respectively, relative to the best-performing baseline.The code and the dataset can be accessed from the following link: https://www.iitp. ac.in/~ai-nlp-ml/resources.html#DICE.
Gopendra Vikram Singh, Soumitra Ghosh, Atul Verma, Chetna Painkra, Asif Ekbal
EMNLP1
2023 Promoting Gender Equality through Gender-biased Language Analysis in Social Media
abstract
Gender bias is a pervasive issue that impacts women's and marginalized groups' ability to fully participate in social, economic, and political spheres. This study introduces a novel problem of Gender-biased Language Identification and Extraction (GLIdE) from social media interactions and develops a multi-task deep framework that detects gender-biased content and identifies connected causal phrases from the text using emotional information that is present in the input. The method uses a zero-shot strategy with emotional information and a mechanism to represent gender-stereotyped information as a knowledge graph. In this work, we also introduce the first-of-its-kind Gender-biased Analysis Corpus (GAC) of 12,432 social media posts and improve the best-performing baseline for gender-biased language identification and extraction tasks by margins of 4.88% and 5 ROS points, demonstrating this through empirical evaluation and extensive qualitative analysis. By improving the accuracy of identifying and analyzing gender-biased language, this work can contribute to achieving gender equality and promoting inclusive societies, in line with the United Nations Sustainable Development Goals (UN SDGs) and the Leave No One Behind principle (LNOB). We adhere to the principles of transparency and collaboration in line with the UN SDGs by openly sharing our code and dataset.
Gopendra Vikram Singh, Soumitra Ghosh, Asif Ekbal
IJCAI1
2023 MHaDiG: A Multilingual Humor-aided Multiparty Dialogue Generation in multimodal conversational setting
Dushyant Singh Chauhan, Gopendra Vikram Singh, Asif Ekbal, Pushpak Bhattacharyya
Knowl. Based Syst.2
2023 Affect-GCN: a multimodal graph convolutional network for multi-emotion with intensity recognition and sentiment analysis in dialogues
Mauajama Firdaus, Gopendra Vikram Singh, Asif Ekbal, Pushpak Bhattacharyya
Multim. Tools Appl.2
2023 EmoInt-Trans: A Multimodal Transformer for Identifying Emotions and Intents in Social Conversations
abstract
In the natural language processing community, open-domain conversational agents, also known as chatbots, are gaining popularity. One of the difficulties is getting them to communicate in an emotionally intelligent manner. To generate dialogues, current neural response generation methods depend solely on end-to-end learning from large scale conversation data. Therefore, we introduce a large-scale multi Emotion and Intent guided Multimodal Dialogue (EmoInt-MD) dataset labelled with 32 emotions and 15 empathetic intents having 32 k dialogues taken from different movie genres. We propose a novel multi-task multimodal contextual Transformer framework for simultaneously identifying the emotions and intents in a given utterance utilizing audio and visual features in addition to the textual information. Experimental analysis proves that the proposed framework outperforms several unimodal and multimodal baselines on theEmoInt-MDdataset. This dataset along with our baseline and proposed framework implementations will be made publicly available for research purposes.
Gopendra Vikram Singh, Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Unity in Diversity: Multilabel Emoji Identification in Tweets
abstract
Emojis or emoticons are not just a modern trend but have become an essential part of our day-to-day interactions. Predicting a suitable emoji for a given tweet is a challenging task because a wrong emoji prediction for a tweet can change the meaning of the message or can amplify the emotion of the message. This task is particularly challenging since it requires selecting an appropriate emoji from a huge list of prospective emojis that may or may not be equivalent to one another. Humans use multiple emojis to convey their emotions, thereby making the task a multilabel classification problem. In this article, we propose a multilabel emoji prediction system that predicts the appropriate emoji for a given tweet by using different state-of-the-art baselines. Due to the unavailability of a multi-emoji dataset, we create a large-scale multilabel emoji dataset named Mu-Emoji that comprises of more than 0.6 million tweets having varied emojis belonging to both positive and negative sentiments. For our proposed task, we employ graph attention network along with bidirectional encoder representations from transformer encoder for the accurate prediction of emojis. Qualitative and quantitative analyses show that our multilabel emoji prediction baselines perform well compared with the single-emoji prediction baselines for our proposed Mu-Emoji dataset. Our proposed framework also outperforms all the baselines for both single and multilabel emoji prediction tasks.
Gopendra Vikram Singh, Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya
IEEE Trans. Comput. Soc. Syst.1
2022 A Sentiment and Emotion Aware Multimodal Multiparty Humor Recognition in Multilingual Conversational Setting
abstract
In this paper, we hypothesize that humor is closely related to sentiment and emotions. Also, due to the tremendous growth in multilingual content, there is a great demand for building models and systems that support multilingual information access. To end this, we first extend the recently released Multimodal Multiparty Hindi Humor (M2H2) dataset by adding parallel English utterances corresponding to Hindi utterances and then annotating each utterance with sentiment and emotion classes. We name it Sentiment, Humor, and Emotion aware Multilingual Multimodal Multiparty Dataset (SHEMuD). Therefore, we propose a multitask framework wherein the primary task is humor detection, and the auxiliary tasks are sentiment and emotion identification. We design a multitasking framework wherein we first propose a Context Transformer to capture the deep contextual relationships with the input utterances. We then propose a Sentiment and Emotion aware Embedding (SE-Embedding) to get the overall representation of a particular emotion and sentiment w.r.t. the specific humor situation. Experimental results on the SHEMuD show the efficacy of our approach and shows that multitask learning offers an improvement over the single-task framework for both monolingual (4.86 points in Hindi and 5.9 points in English in F1-score) and multilingual (5.17 points in F1-score) setting.
Dushyant Singh Chauhan, Gopendra Vikram Singh, Aseem Arora, Asif Ekbal, Pushpak Bhattacharyya
COLING2
2022 COMMA-DEER: COmmon-sense Aware Multimodal Multitask Approach for Detection of Emotion and Emotional Reasoning in Conversations
abstract
Mental health is a critical component of the United Nations’ Sustainable Development Goals (SDGs), particularly Goal 3, which aims to provide “good health and well-being”. The present mental health treatment gap is exacerbated by stigma, lack of human resources, and lack of research capability for implementation and policy reform. We present and discuss a novel task of detecting emotional reasoning (ER) and accompanying emotions in conversations. In particular, we create a first-of-its-kind multimodal mental health conversational corpus that is manually annotated at the utterance level with emotional reasoning and related emotion. We develop a multimodal multitask framework with a novel multimodal feature fusion technique and a contextuality learning module to handle the two tasks. Leveraging multimodal sources of information, commonsense reasoning, and through a multitask framework, our proposed model produces strong results. We achieve performance gains of 6% accuracy and 4.62% F1 on the emotion detection task and 3.56% accuracy and 3.31% F1 on the ER detection task, when compared to the existing state-of-the-art model.
Soumitra Ghosh, Gopendra Vikram Singh, Asif Ekbal, Pushpak Bhattacharyya
COLING2
2022 EmoInHindi: A Multi-label Emotion and Intensity Annotated Dataset in Hindi for Emotion Recognition in Dialogues
abstract
The long-standing goal of Artificial Intelligence (AI) has been to create human-like conversational systems. Such systems should have the ability to develop an emotional connection with the users, consequently, emotion recognition in dialogues has gained popularity. Emotion detection in dialogues is a challenging task because humans usually convey multiple emotions with varying degrees of intensities in a single utterance. Moreover, emotion in an utterance of a dialogue may be dependent on previous utterances making the task more complex. Recently, emotion recognition in low-resource languages like Hindi has been in great demand. However, most of the existing datasets for multi-label emotion and intensity detection in conversations are in English. To this end, we propose a large conversational dataset in Hindi named EmoInHindi for multi-label emotion and intensity recognition in conversations containing 1,814 dialogues with a total of 44,247 utterances. We prepare our dataset in a Wizard-of-Oz manner for mental health and legal counselling of crime victims. Each utterance of dialogue is annotated with one or more emotion categories from 16 emotion labels including neutral and their corresponding intensity. We further propose strong contextual baselines that can detect the emotion(s) and corresponding emotional intensity of an utterance given the conversational context.
Gopendra Vikram Singh, Priyanshu Priya, Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya
LREC1
2022 Are Emoji, Sentiment, and Emotion Friends? A Multi-task Learning for Emoji, Sentiment, and Emotion Analysis
Gopendra Vikram Singh, Dushyant Singh Chauhan, Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya
PACLIC1
2022 An emoji-aware multitask framework for multimodal sarcasm detection
Dushyant Singh Chauhan, Gopendra Vikram Singh, Aseem Arora, Asif Ekbal, Pushpak Bhattacharyya
Knowl. Based Syst.2
2022 Knowing What to Say: Towards knowledge grounded code-mixed response generation for open-domain conversations
Gopendra Vikram Singh, Mauajama Firdaus, Shambhavi, Asif Ekbal
Knowl. Based Syst.1
2021 M2H2: A Multimodal Multiparty Hindi Dataset For Humor Recognition in Conversations
abstract
Humor recognition in conversations is a challenging task that has recently gained popularity due to its importance in dialogue understanding, including in multimodal settings (i.e., text, acoustics, and visual). The few existing datasets for humor are mostly in English. However, due to the tremendous growth in multilingual content, there is a great demand to build models and systems that support multilingual information access. To this end, we propose a dataset for Multimodal Multiparty Hindi Humor (M2H2) recognition in conversations containing 6,191 utterances from 13 episodes of a very popular TV series ”Shrimaan Shrimati Phir Se”. Each utterance is annotated with humor/non-humor labels and encompasses acoustic, visual, and textual modalities. We propose several strong multimodal baselines and show the importance of contextual and multimodal information for humor recognition in conversations. The empirical results on M2H2 dataset demonstrate that multimodal information complements unimodal information for humor recognition. The dataset and the baselines are available at http://www.iitp.ac.in/~ai-nlp-ml/resources.html and https://github.com/declare-lab/M2H2-dataset.
Dushyant Singh Chauhan, Gopendra Vikram Singh, Navonil Majumder, Amir Zadeh 0001, Asif Ekbal, Pushpak Bhattacharyya, Louis-Philippe Morency, Soujanya Poria
ICMI2