VLDB 2026 Research / reviewers in the wild / expert
Soumitra Ghosh
dblp:21/5014
· DBLP profile ↗
21ranked-venue papers
13as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 9 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Insight in Sight: Complaint Detection and Aspect-Based Reasoning Through Visually-Grounded Reviews With VLLMs
Apoorva Singh, Soumitra Ghosh, Karanjot Singh, Bruno Lepri |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | Just a Scratch: Enhancing LLM Capabilities for Self-harm Detection through Intent Differentiation and Emoji InterpretationabstractSoumitra Ghosh, Gopendra Vikram Singh, Shambhavi Shambhavi, Sabarna Choudhury, Asif Ekbal. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Soumitra Ghosh, Gopendra Vikram Singh, Shambhavi, Sabarna Choudhury, Asif Ekbal |
ACL (1) | 1 |
| 2025 | Unmasking offensive content: a multimodal approach with emotional understanding
Gopendra Vikram Singh, Soumitra Ghosh, Mauajama Firdaus, Asif Ekbal, Pushpak Bhattacharyya |
Multim. Tools Appl. | 2 |
| 2024 | Benchmarking Cyber Harassment Dialogue Comprehension through Emotion-Informed Manifestations-Determinants DemarcationabstractIn the digital age, cybercrimes, particularly cyber harassment, have become pressing issues, targeting vulnerable individuals like children, teenagers, and women. Understanding the experiences and needs of the victims is crucial for effective support and intervention. Online conversations between victims and virtual harassment counselors (chatbots) offer valuable insights into cyber harassment manifestations (CHMs) and determinants (CHDs). However, the distinction between CHMs and CHDs remains unclear. This research is the first to introduce concrete definitions for CHMs and CHDs, investigating their distinction through automated methods to enable efficient cyber-harassment dialogue comprehension. We present a novel dataset, Cyber-MaD that contains Cyber harassment dialogues manually annotated with Manifestations and Determinants. Additionally, we design an Emotion-informed Contextual Dual attention Convolution Transformer (E-ConDuCT) framework to extract CHMs and CHDs from cyber harassment dialogues. The framework primarily: a) utilizes inherent emotion features through adjective-noun pairs modeled by an autoencoder, b) employs a unique Contextual Dual attention Convolution Transformer to learn contextual insights; and c) incorporates a demarcation module leveraging task-specific emotional knowledge and a discriminator loss function to differentiate manifestations and determinants. E-ConDuCT outperforms the state-of-the-art systems on the Cyber-MaD corpus, showcasing its potential in the extraction of CHMs and CHDs. Furthermore, its robustness is demonstrated on the emotion cause extraction task using the CARES_CEASE-v2.0 dataset of suicide notes, confirming its efficacy across diverse cause extraction objectives. Access the code and data at 1. https://www.iitp.ac.in/~ai-nlp-ml/resources.html#E-ConDuCT-on-Cyber-MaD, 2. https://github.com/Soumitra816/Manifestations-Determinants. Soumitra Ghosh, Gopendra Vikram Singh, Jashn Arora, Asif Ekbal |
AAAI | 1 |
| 2024 | From Pink and Blue to a Rainbow Hue! Defying Gender Bias through Gender Neutralizing Text Transformations
Gopendra Vikram Singh, Soumitra Ghosh, Neil Dcruze, Asif Ekbal |
IJCAI | 2 |
| 2023 | DeCoDE: Detection of Cognitive Distortion and Emotion Cause Extraction in Clinical Conversations
Gopendra Vikram Singh, Soumitra Ghosh, Asif Ekbal, Pushpak Bhattacharyya |
ECIR (2) | 2 |
| 2023 | Standardizing Distress Analysis: Emotion-Driven Distress Identification and Cause Extraction (DICE) in Multimodal Online PostsabstractDue to its growing impact on public opinion, hate speech on social media has garnered increased attention.While automated methods for identifying hate speech have been presented in the past, they have mostly been limited to analyzing textual content.The interpretability of such models has received very little attention, despite the social and legal consequences of erroneous predictions.In this work, we present a novel problem of Distress Identification and Cause Extraction (DICE) from multimodal online posts.We develop a multi-task deep framework for the simultaneous detection of distress content and identify connected causal phrases from the text using emotional information.The emotional information is incorporated into the training process using a zero-shot strategy, and a novel mechanism is devised to fuse the features from the multimodal inputs.Furthermore, we introduce the first-ofits-kind Distress and Cause annotated Multimodal (DCaM) dataset of 20,764 social media posts.We thoroughly evaluate our proposed method by comparing it to several existing benchmarks.Empirical assessment and comprehensive qualitative analysis demonstrate that our proposed method works well on distress detection and cause extraction tasks, improving F1 and ROS scores by 1.95% and 3%, respectively, relative to the best-performing baseline.The code and the dataset can be accessed from the following link: https://www.iitp. ac.in/~ai-nlp-ml/resources.html#DICE. Gopendra Vikram Singh, Soumitra Ghosh, Atul Verma, Chetna Painkra, Asif Ekbal |
EMNLP | 2 |
| 2023 | Promoting Gender Equality through Gender-biased Language Analysis in Social MediaabstractGender bias is a pervasive issue that impacts women's and marginalized groups' ability to fully participate in social, economic, and political spheres. This study introduces a novel problem of Gender-biased Language Identification and Extraction (GLIdE) from social media interactions and develops a multi-task deep framework that detects gender-biased content and identifies connected causal phrases from the text using emotional information that is present in the input. The method uses a zero-shot strategy with emotional information and a mechanism to represent gender-stereotyped information as a knowledge graph. In this work, we also introduce the first-of-its-kind Gender-biased Analysis Corpus (GAC) of 12,432 social media posts and improve the best-performing baseline for gender-biased language identification and extraction tasks by margins of 4.88% and 5 ROS points, demonstrating this through empirical evaluation and extensive qualitative analysis. By improving the accuracy of identifying and analyzing gender-biased language, this work can contribute to achieving gender equality and promoting inclusive societies, in line with the United Nations Sustainable Development Goals (UN SDGs) and the Leave No One Behind principle (LNOB). We adhere to the principles of transparency and collaboration in line with the UN SDGs by openly sharing our code and dataset. Gopendra Vikram Singh, Soumitra Ghosh, Asif Ekbal |
IJCAI | 2 |
| 2023 | VAD-assisted multitask transformer framework for emotion recognition and intensity prediction on suicide notes
Soumitra Ghosh, Asif Ekbal, Pushpak Bhattacharyya |
Inf. Process. Manag. | 1 |
| 2023 | Multitasking of sentiment detection and emotion recognition in code-mixed Hinglish data
Soumitra Ghosh, Amit Priyankar, Asif Ekbal, Pushpak Bhattacharyya |
Knowl. Based Syst. | 1 |
| 2023 | A transformer-based multi-task framework for joint detection of aggression and hate on social media dataabstractAbstract Moderators often face a double challenge regarding reducing offensive and harmful content in social media. Despite the need to prevent the free circulation of such content, strict censorship on social media cannot be implemented due to a tricky dilemma – preserving free speech on the Internet while limiting them and how not to overreact. Existing systems do not essentially exploit the correlatedness of hate-offensive content and aggressive posts; instead, they attend to the tasks individually. As a result, the need for cost-effective, sophisticated multi-task systems to effectively detect aggressive and offensive content on social media is highly critical in recent times. This work presents a novel multifaceted transformer-based framework to identify aggressive and hate posts on social media. Through an end-to-end transformer-based multi-task network, our proposed approach addresses the following array of tasks: (a) aggression identification, (b) misogynistic aggression identification, (c) identifying hate-offensive and non-hate-offensive content, (d) identifying hate, profane, and offensive posts, (e) type of offense. We further investigate the role of emotion in improving the system’s overall performance by learning the task of emotion detection jointly with the other tasks. We evaluate our approach on two popular benchmark datasets of aggression and hate speech, covering four languages, and compare the system performance with various state-of-the-art methods. Results indicate that our multi-task system performs significantly well for all the tasks across multiple languages, outperforming several benchmark methods. Moreover, the secondary task of emotion detection substantially improves the system performance for all the tasks, indicating strong correlatedness among the tasks of aggression, hate, and emotion, thus opening avenues for future research. Soumitra Ghosh, Amit Priyankar, Asif Ekbal, Pushpak Bhattacharyya |
Nat. Lang. Eng. | 1 |
| 2023 | SEHC: A Benchmark Setup to Identify Online Hate Speech in EnglishabstractThanks to the digital age, online speech and information may now be disseminated anonymously without regard for repercussions. Regulators face a unique problem with social media platforms because of the speed and volume of material and the lack of editorial supervision. The existing datasets on hate speech or offensive language identification lack diversity in the dataset’s content. In this article, we create a multi-domain hate speech corpus (MHC) of English tweets that includes hate speech against religion, nationality, ethnicity, and gender in general and cover diverse domains, such as current affairs, politics, terrorism, technology, natural disasters, and human/drugs trafficking. Each instance in our dataset is manually annotated as hate or non-hate. We use the existing state-of-the-art models and present a stacked-ensemble-based hate speech classifier (SEHC) to identify hate speech from Twitter data. Our results indicate that the proposed method may serve as a strong baseline for future studies using this dataset. (The dataset is available athttps://www.iitp.ac.in/~ai-nlp-ml/resources.html#MHC.) Soumitra Ghosh, Asif Ekbal, Pushpak Bhattacharyya, Tista Saha, Alka Kumar, Shikha Srivastava |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2022 | EM-PERSONA: EMotion-assisted Deep Neural Framework for PERSONAlity Subtyping from Suicide NotesabstractThe World Health Organization has emphasised the need of stepping up suicide prevention efforts to meet the United Nation’s Sustainable Development Goal target of 2030 (Goal 3: Good health and well-being). We address the challenging task of personality subtyping from suicide notes. Most research on personality subtyping has relied on statistical analysis and feature engineering. Moreover, state-of-the-art transformer models in the automated personality subtyping problem have received relatively less attention. We develop a novel EMotion-assisted PERSONAlity Detection Framework (EM-PERSONA). We annotate the benchmark CEASE-v2.0 suicide notes dataset with personality traits across four dichotomies: Introversion (I)-Extraversion (E), Intuition (N)-Sensing (S), Thinking (T)-Feeling (F), Judging (J)–Perceiving (P). Our proposed method outperforms all baselines on comprehensive evaluation using multiple state-of-the-art systems. Across the four dichotomies, EM-PERSONA improved accuracy by 2.04%, 3.69%, 4.52%, and 3.42%, respectively, over the highest-performing single-task systems. Soumitra Ghosh, Dhirendra Maurya, Asif Ekbal, Pushpak Bhattacharyya |
COLING | 1 |
| 2022 | COMMA-DEER: COmmon-sense Aware Multimodal Multitask Approach for Detection of Emotion and Emotional Reasoning in ConversationsabstractMental health is a critical component of the United Nations’ Sustainable Development Goals (SDGs), particularly Goal 3, which aims to provide “good health and well-being”. The present mental health treatment gap is exacerbated by stigma, lack of human resources, and lack of research capability for implementation and policy reform. We present and discuss a novel task of detecting emotional reasoning (ER) and accompanying emotions in conversations. In particular, we create a first-of-its-kind multimodal mental health conversational corpus that is manually annotated at the utterance level with emotional reasoning and related emotion. We develop a multimodal multitask framework with a novel multimodal feature fusion technique and a contextuality learning module to handle the two tasks. Leveraging multimodal sources of information, commonsense reasoning, and through a multitask framework, our proposed model produces strong results. We achieve performance gains of 6% accuracy and 4.62% F1 on the emotion detection task and 3.56% accuracy and 3.31% F1 on the ER detection task, when compared to the existing state-of-the-art model. Soumitra Ghosh, Gopendra Vikram Singh, Asif Ekbal, Pushpak Bhattacharyya |
COLING | 1 |
| 2022 | CARES: CAuse Recognition for Emotion in Suicide Notes
Soumitra Ghosh, Swarup Roy, Asif Ekbal, Pushpak Bhattacharyya |
ECIR (2) | 1 |
| 2022 | Am I No Good? Towards Detecting Perceived Burdensomeness and Thwarted Belongingness from Suicide NotesabstractThe World Health Organization (WHO) has emphasized the importance of significantly accelerating suicide prevention efforts to fulfill the United Nations' Sustainable Development Goal (SDG) objective of 2030. In this paper, we present an end-to-end multitask system to address a novel task of detection of two interpersonal risk factors of suicide, Perceived Burdensomeness (PB) and Thwarted Belongingness (TB) from suicide notes. We also introduce a manually translated code-mixed suicide notes corpus, CoMCEASE-v2.0, based on the benchmark CEASE-v2.0 dataset, annotated with temporal orientation, PB and TB labels. We exploit the temporal orientation and emotion information in the suicide notes to boost overall performance. For comprehensive evaluation of our proposed method, we compare it to several state-of-the-art approaches on the existing CEASE-v2.0 dataset and the newly announced CoMCEASE-v2.0 dataset. Empirical evaluation suggests that temporal and emotional information can substantially improve the detection of PB and TB. Soumitra Ghosh, Asif Ekbal, Pushpak Bhattacharyya |
IJCAI | 1 |
| 2022 | What Does Your Bio Say? Inferring Twitter Users' Depression Status From Multimodal Profile Information Using Deep LearningabstractPeople suffering from stress and various mental health problems find it easier to express and share their feelings on online platforms, such as Twitter. However, the imposed character limit (280 characters) by Twitter and infrequent online activities of a section of users poses a serious setback in using computational methods for mental health analysis or emotion research. Twitter provides rich metadata information about its users (such as user’s description, geolocation, and profile image URL), which can provide valuable information regarding the mental state of the users. We hypothesize that Twitter’s rich metadata information about their users can provide some valuable depression cues, which may help in an early low-profile evaluation. In this article, we investigate this hypothesis by developing an end-to-end multimodal multitask (MT) system for depression detection (primary task) and emotion recognition (auxiliary task), where the variation of emotion information based on different user descriptions assists the learning of the primary task. The proposed system attains 70% accuracy on the depression detection task outperforming several single-task (ST) baselines built on the various combination of input features. Our findings indicate that Twitters’s rich metadata information can be leveraged to detect depression among users with significant confidence. Soumitra Ghosh, Asif Ekbal, Pushpak Bhattacharyya |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2021 | Context and Knowledge Enriched Transformer Framework for Emotion Recognition in ConversationsabstractEmotion Recognition in Conversation (ERC) is becoming increasingly popular due to the accessibility of an enormous measure of openly accessible conversational information. Moreover, it has potential applications in opinion mining, social media and the health care domain. In this paper, we propose a novel Context and Knowledge Enriched Transformer Framework (CKETF) in which we interpret the contextual information from the utterances using a pre-trained Bidirectional Encoder Representations from Transformers (BERT) model and leverage additive attention based hierarchical transformer for encoding the knowledge sentences. Experiments on the knowledge-grounded Topical Chat dataset shows that both context and external knowledge are important for conversational emotion recognition. We demonstrate through extensive experiments and analysis that our proposed model significantly outperforms the current state-of-the-art methods. Soumitra Ghosh, Deeksha Varshney, Asif Ekbal, Pushpak Bhattacharyya |
IJCNN | 1 |
| 2020 | CEASE, a Corpus of Emotion Annotated Suicide notes in EnglishabstractA suicide note is usually written shortly before the suicide and it provides a chance to comprehend the self-destructive state of mind of the deceased. From a psychological point of view, suicide notes have been utilized for recognizing the motive behind the suicide. To the best of our knowledge, there is no openly accessible suicide note corpus at present, making it challenging for the researchers and developers to deep dive into the area of mental health assessment and suicide prevention. In this paper, we create a fine-grained emotion annotated corpus (CEASE) of suicide notes in English and develop various deep learning models to perform emotion detection on the curated dataset. The corpus consists of 2393 sentences from around 205 suicide notes collected from various sources. Each sentence is annotated with a particular emotion class from a set of 15 fine-grained emotion labels, namely (forgiveness, happiness_peacefulness, love, pride, hopefulness, thankfulness, blame, anger, fear, abuse, sorrow, hopelessness, guilt, information, instructions). For the evaluation, we develop an ensemble architecture, where the base models correspond to three supervised deep learning models, namely Convolutional Neural Network (CNN), Gated Recurrent Unit (GRU) and Long Short Term Memory (LSTM). We obtain the highest test accuracy of 60.17% and cross-validation accuracy of 60.32% Soumitra Ghosh, Asif Ekbal, Pushpak Bhattacharyya |
LREC | 1 |
| 2006 | Comprehensive quality control utilizing the prehybridization third-dye image leads to accurate gene expression measurements by cDNA microarraysabstractBACKGROUND: Gene expression profiling using microarrays has become an important genetic tool. Spotted arrays prepared in academic labs have the advantage of low cost and high design and content flexibility, but are often limited by their susceptibility to quality control (QC) issues. Previously, we have reported a novel 3-color microarray technology that enabled array fabrication QC. In this report we further investigated its advantage in spot-level data QC. RESULTS: We found that inadequate amount of bound probes available for hybridization led to significant, gene-specific compression in ratio measurements, increased data variability, and printing pin dependent heterogeneities. The impact of such problems can be captured through the definition of quality scores, and efficiently controlled through quality-dependent filtering and normalization. We compared gene expression measurements derived using our data processing pipeline with the known input ratios of spiked in control clones, and with the measurements by quantitative real time RT-PCR. In each case, highly linear relationships (R2 > 0.94) were observed, with modest compression in the microarray measurements (correction factor < 1.17). CONCLUSION: Our microarray analytical and technical advancements enabled a better dissection of the sources of data variability and hence a more efficient QC. With that highly accurate gene expression measurements can be achieved using the cDNA microarray technology. Xujing Wang, Shuang Jia, Lisa Meyer, Bixia Xiang, Li-Yen Chen, Carol Moreno, Howard J. Jacob, Soumitra Ghosh, Martin J. Hessner |
BMC Bioinform. | 9 |
| 2003 | Quantitative Quality Control in Microarray Experiments and the Application in Data Filtering, Normalization and False Positive Rate PredictionabstractData preprocessing including proper normalization and adequate quality control before complex data mining is crucial for studies using the cDNA microarray technology. We have developed a simple procedure that integrates data filtering and normalization with quantitative quality control of microarray experiments. Previously we have shown that data variability in a microarray experiment can be very well captured by a quality score q(com) that is defined for every spot, and the ratio distribution depends on q(com). Utilizing this knowledge, our data-filtering scheme allows the investigator to decide on the filtering stringency according to desired data variability, and our normalization procedure corrects the q(com)-dependent dye biases in terms of both the location and the spread of the ratio distribution. In addition, we propose a statistical model for false positive rate determination based on the design and the quality of a microarray experiment. The model predicts that a lower limit of 0.5 for the replicate concordance rate is needed in order to be certain of true positives. Our work demonstrates the importance and advantages of having a quantitative quality control scheme for microarrays. Xujing Wang, Martin J. Hessner, Nirupma Pati, Soumitra Ghosh |
Bioinform. | 5 |