Soujanya Poria

dblp:116/4904 · DBLP profile ↗
← Back
103ranked-venue papers
16as first author
58since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 90 · 16 first-author · 47 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 13 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 10 Open Challenges Steering the Future of Vision-Language-Action Models
abstract
Due to their ability of follow natural language instructions, vision-language-action (VLA) models are increasingly preva- lent in the embodied AI arena, following the widespread suc- cess of their precursors—LLMs and VLMs. In this paper, we discuss 10 principal milestones in the ongoing develop- ment of VLA models—multimodality, reasoning, data, eval- uation, cross-robkot action generalization, efficiency, whole- body coordination, safety, agents, and coordination with hu- mans. Furthermore, we discuss the emerging trends of us- ing spatial understanding, modeling world dynamics, post training, and data synthesis—all aiming to reach these mile- stones. Through these discussions, we hope to bring attention to the research avenues that may accelerate the development of VLA models into wider acceptability.
Soujanya Poria, Navonil Majumder, Chia-Yu Hung, Amir Ali Bagherzadeh, Kenneth Kwok, Ziwei Wang 0010, Cheston Tan, Jiajun Wu 0001, David Hsu
AAAI1
2026 DialogXpert: Driving Intelligent and Emotion-Aware Conversations Through Online Value-Based Reinforcement Learning with LLM Priors
abstract
Large-language-model (LLM) agents excel at reactive dialogue but struggle with proactive, goal-driven interactions due to myopic decoding and costly planning. We introduce DialogXpert, which leverages a frozen LLM to propose a small, high-quality set of candidate actions per turn and employs a compact Q-network over fixed BERT embeddings trained via temporal-difference learning to select optimal moves within this reduced space. By tracking the user's emotions DialogXpert tailors each decision to advance the task while nurturing a genuine, empathetic connection. Across negotiation, emotional support, and tutoring benchmarks, DialogXpert drives conversations to under 3 turns with success rates exceeding 94% and, with a larger LLM prior, pushes success above 97% while markedly improving negotiation outcomes. This framework delivers real-time, strategic, and emotionally intelligent dialogue planning at scale.
Tazeek Bin Abdur Rakib, Ambuj Mehrish, Lay-Ki Soon, Wern Han Lim, Soujanya Poria
AAAI5
2025 Why AI Is WEIRD and Shouldn't Be This Way: Towards AI for Everyone, with Everyone, by Everyone
abstract
This paper presents a vision for creating AI systems that are inclusive at every stage of development, from data collection to model design and evaluation. We address key limitations in the current AI pipeline and its WEIRD* representation, such as lack of data diversity, biases in model performance, and narrow evaluation metrics. We also focus on the need for diverse representation among the developers of these systems, as well as incentives that are not skewed toward certain groups. We highlight opportunities to develop AI systems that are for everyone (with diverse stakeholders in mind), with everyone (inclusive of diverse data and annotators), and by everyone (designed and developed by a globally diverse workforce). *WEIRD = an acronym coined by Joseph Henrich to highlight the coverage limitations of many psychological studies, referring to populations that are Western, Educated, Industrialized, Rich, and Democratic; while we do not fully adopt this term for AI, as its current scope does not perfectly align with the WEIRD dimensions, we believe that today's AI has a similarly "weird" coverage, particularly in terms of who is involved in its development and who benefits from it.
Rada Mihalcea, Oana Ignat, Longju Bai, Angana Borah, Luis Chiruzzo, Zhijing Jin 0001, Claude Kwizera, Joan Nwatu, Soujanya Poria, Thamar Solorio
AAAI9
2025 Pixel-Level Reasoning Segmentation via Multi-turn Conversations
abstract
Dexian Cai, Xiaocui Yang, YongKang Liu, Daling Wang, Shi Feng, Yifei Zhang, Soujanya Poria. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Dexian Cai, Xiaocui Yang, Yongkang Liu 0002, Daling Wang, Shi Feng 0001, Yifei Zhang 0003, Soujanya Poria
ACL (1)7
2025 DiffPO: Diffusion-styled Preference Optimization for Inference Time Alignment of Large Language Models
abstract
Ruizhe Chen, Wenhao Chai, Zhifei Yang, Xiaotian Zhang, Ziyang Wang, Tony Quek, Joey Tianyi Zhou, Soujanya Poria, Zuozhu Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ruizhe Chen, Wenhao Chai, Zhifei Yang 0004, Tony Q. S. Quek, Joey Tianyi Zhou, Soujanya Poria, Zuozhu Liu
ACL (1)8
2025 Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
abstract
Traditional reinforcement learning-based robotic control methods are often task-specific and fail to generalize across diverse environments or unseen objects and instructions. Visual Language Models (VLMs) demonstrate strong scene understanding and planning capabilities but lack the ability to generate actionable policies tailored to specific robotic embodiments. To address this, Visual-Language-Action (VLA) models have emerged, yet they face challenges in long-horizon spatial reasoning and grounded task planning. In this work, we propose the Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning, EMMA-X. EMMA-X leverages our constructed hierarchical embodiment dataset based on BridgeV2, containing 60,000 robot manipulation trajectories auto-annotated with grounded task reasoning and spatial guidance. Additionally, we introduce a trajectory segmentation strategy based on gripper states and motion trajectories, which can help mitigate hallucination in grounding subtask reasoning generation. Experimental results demonstrate that EMMA-X achieves superior performance over competitive baselines, particularly in real-world robotic tasks requiring spatial reasoning.
Pengfei Hong, Pala Tej Deep, Vernon Toh Yan Han, U-Xuan Tan, Deepanway Ghosal, Soujanya Poria
ACL (1)7
2025 M-LongDoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework
abstract
Yew Ken Chia, Liying Cheng, Hou Pong Chan, Maojia Song, Chaoqun Liu, Mahani Aljunied, Soujanya Poria, Lidong Bing. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yew Ken Chia, Liying Cheng, Hou Pong Chan, Maojia Song, Chaoqun Liu, Mahani Aljunied, Soujanya Poria, Lidong Bing
EMNLP7
2025 MOOSE-Chem: Large Language Models for Rediscovering Unseen Chemistry Scientific Hypotheses
abstract
Scientific discovery contributes largely to the prosperity of human society, and recent progress shows that LLMs could potentially catalyst the process. However, it is still unclear whether LLMs can discover novel and valid hypotheses in chemistry. In this work, we investigate this main research question: whether LLMs can automatically discover novel and valid chemistry research hypotheses, given only a research question? With extensive discussions with chemistry experts, we adopt the assumption that a majority of chemistry hypotheses can be resulted from a research background question and several inspirations. With this key insight, we break the main question into three smaller fundamental questions. In brief, they are: (1) given a background question, whether LLMs can retrieve good inspirations; (2) with background and inspirations, whether LLMs can lead to hypothesis; and (3) whether LLMs can identify good hypotheses to rank them higher. To investigate these questions, we construct a benchmark consisting of 51 chemistry papers published in Nature or a similar level in 2024 (all papers are only available online since 2024). Every paper is divided by chemistry PhD students into three components: background, inspirations, and hypothesis. The goal is to rediscover the hypothesis given only the background and a large chemistry literature corpus consisting the ground truth inspiration papers, with LLMs trained with data up to 2023. We also develop an LLM-based multi-agent framework that leverages the assumption, consisting of three stages reflecting the more smaller questions. The proposed method can rediscover many hypotheses with very high similarity with the ground truth ones, covering the main innovations.
Zonglin Yang 0001, Wanhao Liu, Ben Gao, Tong Xie, Wanli Ouyang, Soujanya Poria, Erik Cambria, Dongzhan Zhou
ICLR7
2025 Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
abstract
LLMs are an integral component of retrieval-augmented generation (RAG) systems. While many studies focus on evaluating the overall quality of end-to-end RAG systems, there is a gap in understanding the appropriateness of LLMs for the RAG task. To address this, we introduce Trust-Score, a holistic metric that evaluates the trustworthiness of LLMs within the RAG framework. Our results show that various prompting methods, such as in-context learning, fail to effectively adapt LLMs to the RAG task as measured by Trust-Score. Consequently, we propose Trust-Align, a method to align LLMs for improved Trust-Score performance. 26 out of 27 models aligned using Trust-Align substantially outperform competitive baselines on ASQA, QAMPARI, and ELI5. Specifically, in LLaMA-3-8b, Trust-Align outperforms FRONT on ASQA (↑12.56), QAMPARI (↑36.04), and ELI5 (↑17.69). Trust-Align also significantly enhances models’ ability to correctly refuse and provide quality citations. We also demonstrate the effectiveness of Trust-Align across different open-weight models, including the LLaMA series (1b to 8b), Qwen-2.5 series (0.5b to 7b), and Phi3.5 (3.8b). We release our code at https://github.com/declare-lab/trust-align.
Maojia Song, Shang Hong Sim, Rishabh Bhardwaj, Hai Leong Chieu, Navonil Majumder, Soujanya Poria
ICLR6
2025 The ACM Multimedia 2025 Grand Challenge of Multimodal Conversational Aspect-based Sentiment Analysis
abstract
Understanding fine-grained sentiment dynamics in human conversations is a central goal for next-generation artificial intelligence, especially in scenarios where interactions are rich in both modalities and context. To advance research in this area, we organize the Multimodal Conversational Aspect-based Sentiment Analysis (MCABSA) challenge to the community of aspect-based sentiment analysis. The MCABSA challenge introduces two novel subtasks: 1) Panoptic Sentiment Sextuple Extraction, panoramically recognizing holder, target, aspect, opinion, sentiment, and rationale from multi-turn, multi-party multimodal dialogue; and 2) Sentiment Flipping Analysis, detecting the dynamic sentiment transformation throughout the conversation along with the causal reasons. To support these tasks, we present the PanoSent dataset, a high-quality, large-scale benchmark featuring multi-turn, multi-party dialogues annotated with both explicit and implicit sentiment elements across text, image, audio, and video modalities. PanoSent offers extensive real-world scenario coverage, providing a comprehensive testbed for multimodal conversational sentiment analysis. The challenge has attracted widespread participation from both academia and industry, with over 30 teams registered and more than 100 successful submissions. In this paper, we introduce the task, dataset, and evaluation settings, summarize the systems of the top teams, and discuss the findings of the participants. Further details of the challenge can be found at https://panosent.github.io/MM25-challenge.
Meng Luo 0010, Hao Fei 0001, Bobo Li 0001, Shengqiong Wu, Qian Liu 0012, Soujanya Poria, Erik Cambria, Mong-Li Lee, Wynne Hsu
ACM Multimedia6
2025 AlgoPuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Algorithmic Multimodal Puzzles
abstract
Deepanway Ghosal, Vernon Toh, Yew Ken Chia, Soujanya Poria. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Deepanway Ghosal, Vernon Toh Yan Han, Yew Ken Chia, Soujanya Poria
NAACL (Long Papers)4
2025 Reward-Guided Tree Search for Inference Time Alignment of Large Language Models
abstract
Chia-Yu Hung, Navonil Majumder, Ambuj Mehrish, Soujanya Poria. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Chia-Yu Hung, Navonil Majumder, Ambuj Mehrish, Soujanya Poria
NAACL (Long Papers)4
2025 Toward Robust Multimodal Sentiment Analysis using multimodal foundational models
Xianbing Zhao, Soujanya Poria, Buzhou Tang
Expert Syst. Appl.2
2024 Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
abstract
Aligned language models face a significant limitation as their fine-tuning often results in compromised safety.To tackle this, we propose a simple method RESTA that performs LLM safety realignment.RESTA stands for REstoring Safety through Task Arithmetic.At its core, it involves a simple arithmetic addition of a safety vector to the weights of the compromised model.We demonstrate the effectiveness of RESTA in both parameter-efficient and full fine-tuning, covering a wide range of downstream tasks, including instruction following in Chinese, English, and Hindi, as well as problem-solving capabilities in Code and Math.We also showcase the generalizability of RESTA on three existing safety evaluation benchmarks and a multilingual benchmark dataset proposed as a part of this work, consisting of 550 harmful questions covering 11 categories, each with 5 sub-categories of harm.Overall, RESTA decreases the harmfulness of the compromised model from 18.6% to 5.1% and from 9.2% to 1.5% in parameter-efficient and full finetuning, respectively, while maintaining most of the model's performance on the task.We release the source code at: https://github. com/declare-lab/resta.
Rishabh Bhardwaj, Soujanya Poria
ACL (1)3
2024 HYPERTTS: Parameter Efficient Adaptation in Text to Speech Using Hypernetworks
abstract
Neural speech synthesis, or text-to-speech (TTS), aims to transform a signal from the text domain to the speech domain. While developing TTS architectures that train and test on the same set of speakers has seen significant improvements, out-of-domain speaker performance still faces enormous limitations. Domain adaptation on a new set of speakers can be achieved by fine-tuning the whole model for each new domain, thus making it parameter-inefficient. This problem can be solved by Adapters that provide a parameter-efficient alternative to domain adaptation. Although famous in NLP, speech synthesis has not seen much improvement from Adapters. In this work, we present HyperTTS, which comprises a small learnable network, “hypernetwork”, that generates parameters of the Adapter blocks, allowing us to condition Adapters on speaker representations and making them dynamic. Extensive evaluations of two domain adaptation settings demonstrate its effectiveness in achieving state-of-the-art performance in the parameter-efficient regime. We also compare different variants of , comparing them with baselines in different studies. Promising results on the dynamic adaptation of adapter parameters using hypernetworks open up new avenues for domain-generic multi-speaker TTS systems. The audio samples and code are available at https://github.com/declare-lab/HyperTTS.
Yingting Li, Rishabh Bhardwaj, Ambuj Mehrish, Bo Cheng 0001, Soujanya Poria
LREC/COLING5
2024 A Comprehensive Survey of Sentence Representations: From the BERT Epoch to the CHATGPT Era and Beyond
abstract
Abhinav Ramesh Kashyap, Thanh-Tung Nguyen, Viktor Schlegel, Stefan Winkler, See-Kiong Ng, Soujanya Poria. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Abhinav Ramesh Kashyap, Thanh-Tung Nguyen, Viktor Schlegel, Stefan Winkler 0001, See-Kiong Ng, Soujanya Poria
EACL (1)6
2024 Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
abstract
Ensuring the safe alignment of large language models (LLMs) with human values is critical as they become integral to applications like translation and question answering.Current alignment methods struggle with dynamic user intentions and complex objectives, making models vulnerable to generating harmful content.We propose SAFETY ARITH-METIC, a training-free framework enhancing LLM safety across different scenarios: Base models, Supervised fine-tuned models (SFT), and Edited models.SAFETY ARITH-METIC involves Harm Direction Removal to avoid harmful content and Safety Alignment to promote safe responses.Additionally, we present NOINTENTEDIT, a dataset highlighting edit instances that could compromise model safety if used unintentionally.Our experiments show that SAFETY ARITHMETIC significantly improves safety measures, reduces over-safety, and maintains model utility, outperforming existing methods in ensuring safe content generation.Source codes and dataset can be accessed at: https://github.com/ declare
Rima Hazra, Sayan Layek, Somnath Banerjee 0002, Soujanya Poria
EMNLP4
2024 Chain-of-Knowledge: Grounding Large Language Models via Dynamic Knowledge Adapting over Heterogeneous Sources
abstract
We present chain-of-knowledge (CoK), a novel framework that augments large language models (LLMs) by dynamically incorporating grounding information from heterogeneous sources. It results in more factual rationales and reduced hallucination in generation. Specifically, CoK consists of three stages: reasoning preparation, dynamic knowledge adapting, and answer consolidation. Given a knowledge-intensive question, CoK first prepares several preliminary rationales and answers while identifying the relevant knowledge domains. If there is no majority consensus among the answers from samples, CoK corrects the rationales step by step by adapting knowledge from the identified domains. These corrected rationales can plausibly serve as a better foundation for the final answer consolidation. Unlike prior studies that primarily use unstructured data, CoK also leverages structured knowledge sources such as Wikidata and tables that provide more reliable factual information. To access both unstructured and structured knowledge sources in the dynamic knowledge adapting stage, we propose an adaptive query generator that allows the generation of queries for various types of query languages, including SPARQL, SQL, and natural sentences. Moreover, to minimize error propagation between rationales, CoK corrects the rationales progressively using preceding corrected rationales to generate and correct subsequent rationales. Extensive experiments show that CoK consistently improves the performance of LLMs on knowledge-intensive tasks across different domains.
Xingxuan Li, Yew Ken Chia, Bosheng Ding, Shafiq R. Joty, Soujanya Poria, Lidong Bing
ICLR6
2024 PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis
abstract
While existing Aspect-based Sentiment Analysis (ABSA) has received extensive effort and advancement, there are still gaps in defining a more holistic research target seamlessly integrating multimodality, conversation context, fine-granularity, and also covering the changing sentiment dynamics as well as cognitive causal rationales. This paper bridges the gaps by introducing a multimodal conversational ABSA, where two novel subtasks are proposed: 1) Panoptic Sentiment Sextuple Extraction, panoramically recognizing holder, target, aspect, opinion, sentiment, rationale from multi-turn multi-party multimodal dialogue. 2) Sentiment Flipping Analysis, detecting the dynamic sentiment transformation throughout the conversation with the causal reasons. To benchmark the tasks, we construct PanoSent, a dataset annotated both manually and automatically, featuring high quality, large scale, multimodality, multilingualism, multi-scenarios, and covering both implicit&explicit sentiment elements. To effectively address the tasks, we devise a novel Chain-of-Sentiment reasoning framework, together with a novel multimodal large language model (namely Sentica) and a paraphrase-based verification mechanism. Extensive evaluations demonstrate the superiority of our methods over strong baselines, validating the efficacy of all our proposed methods. The work is expected to open up a new era for the ABSA community, and thus all our codes and data are open at https://PanoSent.github.io/.
Meng Luo 0010, Hao Fei 0001, Bobo Li 0001, Shengqiong Wu, Qian Liu 0012, Soujanya Poria, Erik Cambria, Mong-Li Lee, Wynne Hsu
ACM Multimedia6
2024 Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
abstract
Peer Reviewed
Navonil Majumder, Chia-Yu Hung, Deepanway Ghosal, Wei-Ning Hsu, Rada Mihalcea, Soujanya Poria
ACM Multimedia6
2024 Mustango: Toward Controllable Text-to-Music Generation
abstract
Jan Melechovsky, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, Soujanya Poria. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jan Melechovský, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, Soujanya Poria
NAACL-HLT6
2024 Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
abstract
Siqi Shen, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Soujanya Poria, Rada Mihalcea. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Soujanya Poria, Rada Mihalcea
NAACL-HLT5
2024 Consistency Guided Knowledge Retrieval and Denoising in LLMs for Zero-shot Document-level Relation Triplet Extraction
abstract
Document-level Relation Triplet Extraction (DocRTE) is a fundamental task in information systems that aims to simultaneously extract entities with semantic relations from a document. Existing methods heavily rely on a substantial amount of fully labeled data. However, collecting and annotating data for newly emerging relations is time-consuming and labor-intensive. Recent advanced Large Language Models (LLMs), such as ChatGPT and LLaMA, exhibit impressive long-text generation capabilities, inspiring us to explore an alternative approach for obtaining auto-labeled documents with new relations. In this paper, we propose a Zero-shot Document-level Relation Triplet Extraction (ZeroDocRTE) framework, which Generates labeled data by Retrieval and Denoising Knowledge from LLMs, called GenRDK. Specifically, we propose a chain-of-retrieval prompt to guide ChatGPT to generate labeled long-text data step by step. To improve the quality of synthetic data, we propose a denoising strategy based on the consistency of cross-document knowledge. Leveraging our denoised synthetic data, we proceed to fine-tune the LLaMA2-13B-Chat for extracting document-level relation triplets. We perform experiments for both zero-shot document-level relation and triplet extraction on two public datasets. The experimental results illustrate that our GenRDK framework outperforms strong baselines.
Xiaocui Yang, Rong Tong, Soujanya Poria
WWW6
2024 Hate speech detection: A comprehensive review of recent works
abstract
Abstract There has been surge in the usage of Internet as well as social media platforms which has led to rise in online hate speech targeted on individual or group. In the recent years, hate speech has resulted in one of the challenging problems that can unfurl at a fast pace on digital platforms leading to various issues such as prejudice, violence and even genocide. Considering the acceptance of Artificial Intelligence (AI) and Natural Language Processing (NLP) techniques in varied application domains, it would be intriguing to consider these techniques for automated hate speech detection. In literature, there have been efforts to recognize and categorize hate speech using varied Machine Learning (ML) and Deep Learning (DL) techniques. Hence, considering the need and provocations for hate speech detection we aim to present a comprehensive review that discusses fundamental taxonomy as well as recent advances in the field of online hate speech identification. There is a significant amount of literature related to the initial phases of hate speech detection. The background section provides a detailed explanation of the previous research. The subsequent section that follows is dedicated to examining the recent literature published from the year 2020 onwards. The paper presents some of the hate speech datasets considered for hate speech detection. Furthermore, the paper discusses different data modalities, namely, textual hate speech detection, multi‐modal hate speech detection and multilingual hate speech detection. Apart from systematic review on hate speech detection, the paper also implement several multi‐label models to compare the performance of hate speech detection by employing classic ML technique namely, Logistic Regression and DL technique namely, Long Short‐Term Memory (LSTM) and a multiclass multi‐label architecture. In the implemented architecture, we have derived two new elements to quantify the hatefulness and intensity of hatred to improve the results for hate speech detection using Indonesian tweet dataset. Empirical Analysis of the model reveals that the implemented approach outperforms and is able to achieve improved results for the underlying dataset.
Ankita Gandhi, Param Ahir, Kinjal Adhvaryu, Pooja Shah, Ritika Lohiya, Erik Cambria, Soujanya Poria, Amir Hussain 0001
Expert Syst. J. Knowl. Eng.7
2024 Video2Music: Suitable music generation from videos using an Affective Multimodal Transformer model
Jaeyong Kang, Soujanya Poria, Dorien Herremans
Expert Syst. Appl.2
2023 Uncertainty Guided Label Denoising for Document-level Distant Relation Extraction
abstract
Document-level relation extraction (DocRE)aims to infer complex semantic relations among entities in a document.Distant supervision (DS) is able to generate massive auto-labeled data, which can improve DocRE performance.Recent works leverage pseudo labels generated by the pre-denoising model to reduce noise in DS data.However, unreliable pseudo labels bring new noise, e.g., adding false pseudo labels and losing correct DS labels.Therefore, how to select effective pseudo labels to denoise DS data is still a challenge in document-level distant relation extraction.To tackle this issue, we introduce uncertainty estimation technology to determine whether pseudo labels can be trusted.In this work, we propose a Documentlevel distant Relation Extraction framework with Uncertainty Guided label denoising, UG-DRE.Specifically, we propose a novel instancelevel uncertainty estimation method, which measures the reliability of the pseudo labels with overlapping relations.By further considering the long-tail problem, we design dynamic uncertainty thresholds for different types of relations to filter high-uncertainty pseudo labels.We conduct experiments on two public datasets.Our framework outperforms strong baselines by 1.91 F 1 and 2.28 Ign F 1 on the RE-DocRED dataset.
Xiaocui Yang, Pengfei Hong, Soujanya Poria
ACL (1)6
2023 UDAPTER - Efficient Domain Adaptation Using Adapters
abstract
We propose two methods to make unsupervised domain adaptation (UDA) more parameter efficient using adapters, small bottleneck layers interspersed with every layer of the largescale pre-trained language model (PLM).The first method deconstructs UDA into a two-step process: first by adding a domain adapter to learn domain-invariant information and then by adding a task adapter that uses domaininvariant information to learn task representations in the source domain.The second method jointly learns a supervised classifier while reducing the divergence measure.Compared to strong baselines, our simple methods perform well in natural language inference (MNLI) and the cross-domain sentiment classification task.We even outperform unsupervised domain adaptation methods such as DANN (Ganin et al., 2016) and DSN (Bousmalis et al., 2016) in sentiment classification, and we are within 0.85% F1 for natural language inference task, by fine-tuning only a fraction of the full model parameters.We release our code at https://github.com/declare-lab/domadapter.
Bhavitvya Malik, Abhinav Ramesh Kashyap, Min-Yen Kan, Soujanya Poria
EACL4
2023 LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models
abstract
The success of large language models (LLMs), like GPT-4 and ChatGPT, has led to the development of numerous cost-effective and accessible alternatives that are created by finetuning open-access LLMs with task-specific data (e.g., ChatDoctor) or instruction data (e.g., Alpaca).Among the various fine-tuning methods, adapter-based parameter-efficient fine-tuning (PEFT) is undoubtedly one of the most attractive topics, as it only requires fine-tuning a few external parameters instead of the entire LLMs while achieving comparable or even better performance.To enable further research on PEFT methods of LLMs, this paper presents LLM-Adapters, an easy-to-use framework that integrates various adapters into LLMs and can execute these adapter-based PEFT methods of LLMs for different tasks.The framework includes state-of-the-art open-access LLMs such as LLaMA, BLOOM, and GPT-J, as well as widely used adapters such as Series adapters, Parallel adapter, Prompt-based learning and Reparametrization-based methods.Moreover, we conduct extensive empirical studies on the impact of adapter types, placement locations, and hyper-parameters to the best design for each adapter-based methods.We evaluate the effectiveness of the adapters on fourteen datasets from two different reasoning tasks, Arithmetic Reasoning and Commonsense Reasoning.The results demonstrate that using adapter-based PEFT in smaller-scale LLMs (7B) with few extra trainable parameters yields comparable, and in some cases superior, performance to powerful LLMs (175B) in zero-shot inference on both reasoning tasks.The code and datasets can be found in https://github. com/AGI-Edgerunners/LLM-Adapters.
Lei Wang 0185, Yihuai Lan, Wanyu Xu, Ee-Peng Lim, Lidong Bing, Xing Xu 0001, Soujanya Poria, Roy Ka-Wei Lee
EMNLP8
2023 Evaluating Parameter-Efficient Transfer Learning Approaches on SURE Benchmark for Speech Understanding
abstract
Fine-tuning is widely used as the default algorithm for transfer learning from pre-trained models. Parameter inefficiency can however arise when, during transfer learning, all the parameters of a large pre-trained model need to be updated for individual downstream tasks. As the number of parameters grows, fine-tuning is prone to overfitting and catastrophic forgetting. In addition, full fine-tuning can become prohibitively expensive when the model is used for many tasks. To mitigate this issue, parameter-efficient transfer learning algorithms, such as adapters and prefix tuning, have been proposed as a way to introduce a few trainable parameters that can be plugged into large pre-trained language models such as BERT, HuBERT. In this paper, we introduce the Speech UndeRstanding Evaluation (SURE) benchmark for parameter-efficient learning for various speech processing tasks. Additionally, we introduce a new adapter, ConvAdapter, based on 1D convolution. We show that ConvAdapter outperforms the standard adapters while showing comparable performance against prefix tuning and Low-Rank Adaptation with only 0.94% of trainable parameters.
Yingting Li, Ambuj Mehrish, Rishabh Bhardwaj, Navonil Majumder, Bo Cheng 0001, Shuai Zhao 0001, Amir Zadeh 0001, Rada Mihalcea, Soujanya Poria
ICASSP9
2023 Multiple Contrastive Learning for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis has received extensive attention with the explosion of multimodal data. For multimodal data, representations should have disparate distributions in the feature space under different labels. The paired multi-modal image-text posts should be closer than unpaired. We propose Multimodal fine-grained interaction with the Multiple Contrastive Learning (M2CL) model for image-text multi-modal sentiment detection. Specifically, we first obtain the reinforced global representation of one modality with the assistance of fine-grained information from another via the Multimodal Interaction Component. Then, we introduce the Multiple Contrastive Learning Component, including Supervised Contrastive Learning (SCL) and Dual Multimodal Contrastive Learning (DMCL). SCL accomplishes pushing the posts with the same sentiment closer and pulling the instances of different sentiments apart within each modality. DMCL pushes the paired image-text features together and pulls the unpaired apart with multiple stages. Extensive experiments conducted on three datasets confirm the effectiveness of our approach.
Xiaocui Yang, Shi Feng 0001, Daling Wang, Pengfei Hong, Soujanya Poria
ICASSP5
2023 ADAPTERMIX: Exploring the Efficacy of Mixture of Adapters for Low-Resource TTS Adaptation
abstract
There are significant challenges for speaker adaptation in textto-speech for languages that are not widely spoken or for speakers with accents or dialects that are not well-represented in the training data. To address this issue, we propose the use of the ”mixture of adapters” method. This approach involves adding multiple adapters within a backbone-model layer to learn the unique characteristics of different speakers. Our approach outperforms the baseline, with a noticeable improvement of 5% observed in speaker preference tests when using only one minute of data for each new speaker. Moreover, following the adapter paradigm, we fine-tune only the adapter parameters (11% of the total model parameters). This is a significant achievement in parameter-efficient speaker adaptation, and one of the first models of its kind. Overall, our proposed approach offers a promising solution to the speech synthesis techniques, particularly for adapting to speakers from diverse backgrounds.
Ambuj Mehrish, Abhinav Ramesh Kashyap, Yingting Li, Navonil Majumder, Soujanya Poria
INTERSPEECH5
2023 Sentence Embedder Guided Utterance Encoder (SEGUE) for Spoken Language Understanding
Yi Xuan Tan, Navonil Majumder, Soujanya Poria
INTERSPEECH3
2023 Text-to-Audio Generation using Instruction Guided Latent Diffusion Model
abstract
The immense scale of the recent large language models (LLM) allows many interesting properties, such as, instruction- and chain-of-thought-based fine-tuning, that has significantly improved zero- and few-shot performance in many natural language processing (NLP) tasks. Inspired by such successes, we adopt such an instruction-tuned LLM Flan-T5 as the text encoder for text-to-audio (TTA) generation-a task where the goal is to generate an audio from its textual description. The prior works on TTA either pre-trained a joint text-audio encoder or used a non-instruction-tuned model, such as, T5. Consequently, our latent diffusion model (LDM)-based approach (Tango) outperforms the state-of-the-art AudioLDM on most metrics and stays comparable on the rest on AudioCaps test set, despite training the LDM on a 63 times smaller dataset and keeping the text encoder frozen. This improvement might also be attributed to the adoption of audio pressure level-based sound mixing for training set augmentation, whereas the prior methods take a random mix.
Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, Soujanya Poria
ACM Multimedia4
2023 Few-shot Multimodal Sentiment Analysis Based on Multimodal Probabilistic Fusion Prompts
abstract
Multimodal sentiment analysis has gained significant attention due to the proliferation of multimodal content on social media. However, existing studies in this area rely heavily on large-scale supervised data, which is time-consuming and labor-intensive to collect. Thus, there is a need to address the challenge of few-shot multimodal sentiment analysis. To tackle this problem, we propose a novel method called Multimodal Probabilistic Fusion Prompts (MultiPoint) that leverages diverse cues from different modalities for multimodal sentiment detection in the few-shot scenario. Specifically, we start by introducing a Consistently Distributed Sampling approach called CDS, which ensures that the few-shot dataset has the same category distribution as the full dataset. Unlike previous approaches primarily using prompts based on the text modality, we design unified multimodal prompts to reduce discrepancies between different modalities and dynamically incorporate multimodal demonstrations into the context of each multimodal instance. To enhance the model's robustness, we introduce a probabilistic fusion method to fuse output predictions from multiple diverse prompts for each input. Our extensive experiments on six datasets demonstrate the effectiveness of our approach. First, our method outperforms strong baselines in the multimodal few-shot setting. Furthermore, under the same amount of data (1% of the full dataset), our CDS-based experimental results significantly outperform those based on previously sampled datasets constructed from the same number of instances of each class.
Xiaocui Yang, Shi Feng 0001, Daling Wang, Yifei Zhang 0003, Soujanya Poria
ACM Multimedia5
2023 Beneath the Tip of the Iceberg: Current Challenges and New Directions in Sentiment Analysis Research
abstract
Sentiment analysis as a field has come a long way since it was first introduced as a task nearly 20 years ago. It has widespread commercial applications in various domains like marketing, risk management, market research, and politics, to name a few. Given its saturation in specific subtasks — such as sentiment polarity classification — and datasets, there is an underlying perception that this field has reached its maturity. In this article, we discuss this perception by pointing out the shortcomings and under-explored, yet key aspects of this field necessary to attaintruesentiment understanding. We analyze the significant leaps responsible for its current relevance. Further, we attempt to chart a possible course for this field that covers many overlooked and unanswered questions.
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Rada Mihalcea
IEEE Trans. Affect. Comput.1
2022 CICERO: A Dataset for Contextualized Commonsense Inference in Dialogues
abstract
This paper addresses the problem of dialogue reasoning with contextualized commonsense inference.We curate CICERO, a dataset of dyadic conversations with five types of utterance-level reasoning-based inferences: cause, subsequent event, prerequisite, motivation, and emotional reaction.The dataset contains 53,105 of such inferences from 5,672 dialogues.We use this dataset to solve relevant generative and discriminative tasks: generation of cause and subsequent event; generation of prerequisite, motivation, and listener's emotional reaction; and selection of plausible alternatives.Our results ascertain the value of such dialogue-centric commonsense knowledge datasets.It is our hope that CI-CERO will open new research avenues into commonsense-based dialogue reasoning.
Deepanway Ghosal, Navonil Majumder, Rada Mihalcea, Soujanya Poria
ACL (1)5
2022 So Different Yet So Alike! Constrained Unsupervised Text Style Transfer
abstract
Abhinav Ramesh Kashyap, Devamanyu Hazarika, Min-Yen Kan, Roger Zimmermann, Soujanya Poria. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Abhinav Ramesh Kashyap, Devamanyu Hazarika, Min-Yen Kan, Roger Zimmermann, Soujanya Poria
ACL (1)5
2022 Knowledge Enhanced Reflection Generation for Counseling Dialogues
abstract
In this paper, we study the effect of commonsense and domain knowledge while generating responses in counseling conversations using retrieval and generative methods for knowledge integration.We propose a pipeline that collects domain knowledge through web mining, and show that retrieval from both domainspecific and commonsense knowledge bases improves the quality of generated responses.We also present a model that incorporates knowledge generated by COMET using soft positional encoding and masked self-attention.We show that both retrieved and COMETgenerated knowledge improve the system's performance as measured by automatic metrics and by human evaluation.Lastly, we present a comparative study on the types of knowledge encoded by our system, showing that causal and intentional relationships benefit the generation task more than other types of commonsense relations.
Verónica Pérez-Rosas, Charles Welch, Soujanya Poria, Rada Mihalcea
ACL (1)4
2022 KNOT: Knowledge Distillation Using Optimal Transport for Solving NLP Tasks
abstract
We propose a new approach, Knowledge Distillation using Optimal Transport (KNOT), to distill the natural language semantic knowledge from multiple teacher networks to a student network. KNOT aims to train a (global) student model by learning to minimize the optimal transport cost of its assigned probability distribution over the labels to the weighted sum of probabilities predicted by the (local) teacher models, under the constraints that the student model does not have access to teacher models’ parameters or training data. To evaluate the quality of knowledge transfer, we introduce a new metric, Semantic Distance (SD), that measures semantic closeness between the predicted and ground truth label distributions. The proposed method shows improvements in the global model’s SD performance over the baseline across three NLP tasks while performing on par with Entropy-based distillation on standard accuracy and F1 metrics. The implementation pertaining to this work is publicly available at https://github.com/declare-lab/KNOT.
Rishabh Bhardwaj, Tushar Vaidya, Soujanya Poria
COLING3
2022 DoubleMix: Simple Interpolation-Based Data Augmentation for Text Classification
abstract
This paper proposes a simple yet effective interpolation-based data augmentation approach termed DoubleMix, to improve the robustness of models in text classification. DoubleMix first leverages a couple of simple augmentation operations to generate several perturbed samples for each training data, and then uses the perturbed data and original data to carry out a two-step interpolation in the hidden space of neural models. Concretely, it first mixes up the perturbed data to a synthetic sample and then mixes up the original data and the synthetic perturbed data. DoubleMix enhances models’ robustness by learning the “shifted” features in hidden space. On six text classification benchmark datasets, our approach outperforms several popular text augmentation methods including token-level, sentence-level, and hidden-level data augmentation techniques. Also, experiments in low-resource settings show our approach consistently improves models’ performance when the training data is scarce. Extensive ablation studies and case studies confirm that each component of our approach contributes to the final performance and show that our approach exhibits superior performance on challenging counterexamples. Additionally, visual analysis shows that text features generated by our approach are highly interpretable.
Hui Chen 0023, Wei Han 0002, Diyi Yang, Soujanya Poria
COLING4
2022 SANCL: Multimodal Review Helpfulness Prediction with Selective Attention and Natural Contrastive Learning
abstract
With the boom of e-commerce, Multimodal Review Helpfulness Prediction (MRHP) that identifies the helpfulness score of multimodal product reviews has become a research hotspot. Previous work on this task focuses on attention-based modality fusion, information integration, and relation modeling, which primarily exposes the following drawbacks: 1) the model may fail to capture the really essential information due to its indiscriminate attention formulation; 2) lack appropriate modeling methods that takes full advantage of correlation among provided data. In this paper, we propose SANCL: Selective Attention and Natural Contrastive Learning for MRHP. SANCL adopts a probe-based strategy to enforce high attention weights on the regions of greater significance. It also constructs a contrastive learning framework based on natural matching properties in the dataset. Experimental results on two benchmark datasets with three categories show that SANCL achieves state-of-the-art baseline performance with lower memory consumption.
Wei Han 0002, Hui Chen 0023, Zhen Hai, Soujanya Poria, Lidong Bing
COLING4
2022 PIP: Physical Interaction Prediction via Mental Simulation with Span Selection
Jiafei Duan, Samson Yu Bai Jian, Soujanya Poria, Bihan Wen, Cheston Tan
ECCV (35)3
2022 Vector-Quantized Input-Contextualized Soft Prompts for Natural Language Understanding
abstract
Prompt Tuning has been largely successful as a parameter-efficient method of conditioning large-scale pre-trained language models to perform downstream tasks.Thus far, soft prompt tuning learns a fixed set of task-specific continuous vectors, i.e., soft tokens that remain static across the task samples.A fixed prompt, however, may not generalize well to the diverse kinds of inputs the task comprises.In order to address this, we propose Vector-quantized Input-contextualized Prompts (VIP) 1 as an extension to the soft prompt tuning framework.VIP particularly focuses on two aspectscontextual prompts that learns input-specific contextualization of the soft prompt tokens through a small-scale sentence encoder and quantized prompts that maps the contextualized prompts to a set of learnable codebook vectors through a Vector quantization network.On various language understanding tasks like SuperGLUE, QA, Relation classification, NER and NLI, VIP outperforms the soft prompt tuning (PT) baseline by an average margin of 1.19%.Further, our generalization studies show that VIP learns more robust prompt representations, surpassing PT by a margin of 0.6% -5.3% on Out-of-domain QA and NLI tasks respectively, and by 0.75% on Multi-Task setup over 4 tasks spanning across 12 domains.
Rishabh Bhardwaj, Amrita Saha, Steven C. H. Hoi, Soujanya Poria
EMNLP4
2022 A Dataset for Hyper-Relational Extraction and a Cube-Filling Approach
abstract
Relation extraction has the potential for largescale knowledge graph construction, but current methods do not consider the qualifier attributes for each relation triplet, such as time, quantity or location.The qualifiers form hyperrelational facts which better capture the rich and complex knowledge graph structure.For example, the relation triplet (Leonard Parker, Educated At, Harvard University) can be factually enriched by including the qualifier (End Time, 1967).Hence, we propose the task of hyper-relational extraction to extract more specific and complete facts from text.To support the task, we construct HyperRED, a large-scale and general-purpose dataset.Existing models cannot perform hyper-relational extraction as it requires a model to consider the interaction between three entities.Hence, we propose Cu-beRE, a cube-filling model inspired by tablefilling approaches and explicitly considers the interaction between relation triplets and qualifiers.To improve model scalability and reduce negative class imbalance, we further propose a cube-pruning method.Our experiments show that CubeRE outperforms strong baselines and reveal possible directions for future research.Our code and data are available at github.com/declare-lab/HyperRED.
Yew Ken Chia, Lidong Bing, Sharifah Mahani Aljunied, Luo Si, Soujanya Poria
EMNLP5
2022 Two is Better than Many? Binary Classification as an Effective Approach to Multi-Choice Question Answering
abstract
We propose a simple refactoring of multichoice question answering (MCQA) tasks as a series of binary classifications.The MCQA task is generally performed by scoring each (question, answer) pair normalized over all the pairs, and then selecting the answer from the pair that yield the highest score.For n answer choices, this is equivalent to an n-class classification setup where only one class (true answer) is correct.We instead show that classifying (question, true answer) as positive instances and (question, false answer) as negative instances is significantly more effective across various models and datasets.We show the efficacy of our proposed approach in different tasks -abductive reasoning, commonsense question answering, science question answering, and sentence completion.Our DeBERTa binary classification model reaches the top or close to the top performance on public leaderboards for these tasks.The source code of the proposed approach is available at https
Deepanway Ghosal, Navonil Majumder, Rada Mihalcea, Soujanya Poria
EMNLP4
2022 MM-Align: Learning Optimal Transport-based Alignment Dynamics for Fast and Accurate Inference on Missing Modality Sequences
abstract
Existing multimodal tasks mostly target at the complete input modality setting, i.e., each modality is either complete or completely missing in both training and test sets.However, the randomly missing situations have still been underexplored.In this paper, we present a novel approach named MM-Align to address the missing-modality inference problem.Concretely, we propose 1) an alignment dynamics learning module based on the theory of optimal transport (OT) for indirect missing data imputation; 2) a denoising training algorithm to simultaneously enhance the imputation results and backbone network performance.Compared with previous methods which devote to reconstructing the missing inputs, MM-Align learns to capture and imitate the alignment dynamics between modality sequences.Results of comprehensive experiments on three datasets covering two multimodal tasks empirically demonstrate that our method can perform more accurate and faster inference and relieve overfitting under various missing conditions.
Wei Han 0002, Hui Chen 0023, Min-Yen Kan, Soujanya Poria
EMNLP4
2022 Analyzing Modality Robustness in Multimodal Sentiment Analysis
abstract
Devamanyu Hazarika, Yingting Li, Bo Cheng, Shuai Zhao, Roger Zimmermann, Soujanya Poria. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Devamanyu Hazarika, Yingting Li, Bo Cheng 0001, Shuai Zhao 0001, Roger Zimmermann, Soujanya Poria
NAACL-HLT6
2022 Improving aspect-level sentiment analysis with aspect extraction
Navonil Majumder, Rishabh Bhardwaj, Soujanya Poria, Alexander F. Gelbukh, Amir Hussain 0001
Neural Comput. Appl.3
2021 More Identifiable yet Equally Performant Transformers for Text Classification
abstract
Rishabh Bhardwaj, Navonil Majumder, Soujanya Poria, Eduard Hovy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Rishabh Bhardwaj, Navonil Majumder, Soujanya Poria, Eduard H. Hovy
ACL/IJCNLP (1)3
2021 Aspect Sentiment Triplet Extraction Using Reinforcement Learning
abstract
Aspect Sentiment Triplet Extraction (ASTE) is the task of extracting triplets of aspect terms, their associated sentiments, and the opinion terms that provide evidence for the expressed sentiments. Previous approaches to ASTE usually simultaneously extract all three components or first identify the aspect and opinion terms, then pair them up to predict their sentiment polarities. In this work, we present a novel paradigm, ASTE-RL, by regarding the aspect and opinion terms as arguments of the expressed sentiment in a hierarchical reinforcement learning (RL) framework. We first focus on sentiments expressed in a sentence, then identify the target aspect and opinion terms for that sentiment. This takes into account the mutual interactions among the triplet's components while improving exploration and sample efficiency. Furthermore, this hierarchical RL setup enables us to deal with multiple and overlapping triplets. In our experiments, we evaluate our model on existing datasets from laptop and restaurant domains and show that it achieves state-of-the-art performance. The implementation of this work is publicly available at https://github.com/declare-lab/ASTE-RL.
Samson Yu Bai Jian, Tapas Nayak, Navonil Majumder, Soujanya Poria
CIKM4
2021 STaCK: Sentence Ordering with Temporal Commonsense Knowledge
abstract
Sentence order prediction is the task of finding the correct order of sentences in a randomly ordered document.Correctly ordering the sentences requires an understanding of coherence with respect to the chronological sequence of events described in the text.Documentlevel contextual understanding and commonsense knowledge centered around these events are often essential in uncovering this coherence and predicting the exact chronological order.In this paper, we introduce STaCK -a framework based on graph neural networks and temporal commonsense knowledge to model global information and predict the relative order of sentences.Our graph network accumulates temporal evidence using knowledge of 'past' and 'future' and formulates sentence ordering as a constrained edge classification problem.We report results on five different datasets, and empirically show that the proposed method is naturally suitable for order prediction, thus demonstrating the role of temporal commonsense knowledge.The implementation of this work is available at: https://github.com/declare-lab/ sentence-ordering.Jennifer has her final exam tomorrow.She got so stressed, she pulled an
Deepanway Ghosal, Navonil Majumder, Rada Mihalcea, Soujanya Poria
EMNLP (1)4
2021 Improving Multimodal Fusion with Hierarchical Mutual Information Maximization for Multimodal Sentiment Analysis
abstract
In multimodal sentiment analysis (MSA), the performance of a model highly depends on the quality of synthesized embeddings.These embeddings are generated from the upstream process called multimodal fusion, which aims to extract and combine the input unimodal raw data to produce a richer multimodal representation.Previous work either back-propagates the task loss or manipulates the geometric property of feature spaces to produce favorable fusion results, which neglects the preservation of critical task-related information that flows from input to the fusion results.In this work, we propose a framework named MultiModal InfoMax (MMIM), which hierarchically maximizes the Mutual Information (MI) in unimodal input pairs (inter-modality) and between multimodal fusion result and unimodal input in order to maintain taskrelated information through multimodal fusion.The framework is jointly trained with the main task (MSA) to improve the performance of the downstream MSA task.To address the intractable issue of MI bounds, we further formulate a set of computationally simple parametric and non-parametric methods to approximate their truth value.Experimental results on the two widely used datasets demonstrate the efficacy of our approach.The implementation of this work is publicly available at https://github.com/ declare-lab/Multimodal-Infomax.
Wei Han 0002, Hui Chen 0023, Soujanya Poria
EMNLP (1)3
2021 M2H2: A Multimodal Multiparty Hindi Dataset For Humor Recognition in Conversations
abstract
Humor recognition in conversations is a challenging task that has recently gained popularity due to its importance in dialogue understanding, including in multimodal settings (i.e., text, acoustics, and visual). The few existing datasets for humor are mostly in English. However, due to the tremendous growth in multilingual content, there is a great demand to build models and systems that support multilingual information access. To this end, we propose a dataset for Multimodal Multiparty Hindi Humor (M2H2) recognition in conversations containing 6,191 utterances from 13 episodes of a very popular TV series ”Shrimaan Shrimati Phir Se”. Each utterance is annotated with humor/non-humor labels and encompasses acoustic, visual, and textual modalities. We propose several strong multimodal baselines and show the importance of contextual and multimodal information for humor recognition in conversations. The empirical results on M2H2 dataset demonstrate that multimodal information complements unimodal information for humor recognition. The dataset and the baselines are available at http://www.iitp.ac.in/~ai-nlp-ml/resources.html and https://github.com/declare-lab/M2H2-dataset.
Dushyant Singh Chauhan, Gopendra Vikram Singh, Navonil Majumder, Amir Zadeh 0001, Asif Ekbal, Pushpak Bhattacharyya, Louis-Philippe Morency, Soujanya Poria
ICMI8
2021 Bi-Bimodal Modality Fusion for Correlation-Controlled Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis aims to extract and integrate semantic information collected from multiple modalities to recognize the expressed emotions and sentiment in multimodal data. This research area’s major concern lies in developing an extraordinary fusion scheme that can extract and integrate key information from various modalities. However, previous work is restricted by the lack of leveraging dynamics of independence and correlation between modalities to reach top performance. To mitigate this, we propose the Bi-Bimodal Fusion Network (BBFN), a novel end-to-end network that performs fusion (relevance increment) and separation (difference increment) on pairwise modality representations. The two parts are trained simultaneously such that the combat between them is simulated. The model takes two bimodal pairs as input due to the known information imbalance among modalities. In addition, we leverage a gated control mechanism in the Transformer architecture to further improve the final output. Experimental results on three datasets (CMU-MOSI, CMU-MOSEI, and UR-FUNNY) verifies that our model significantly outperforms the SOTA. The implementation of this work is available at https://github.com/declare-lab/multimodal-deep-learning and https://github.com/declare-lab/BBFN.
Wei Han 0002, Hui Chen 0023, Alexander F. Gelbukh, Amir Zadeh 0001, Louis-Philippe Morency, Soujanya Poria
ICMI6
2021 MTAG: Modal-Temporal Attention Graph for Unaligned Human Multimodal Language Sequences
abstract
Jianing Yang, Yongxin Wang, Ruitao Yi, Yuying Zhu, Azaan Rehman, Amir Zadeh, Soujanya Poria, Louis-Philippe Morency. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Ruitao Yi, Yuying Zhu 0004, Azaan Rehman, Amir Zadeh 0001, Soujanya Poria, Louis-Philippe Morency
NAACL-HLT7
2021 CIDER: Commonsense Inference for Dialogue Explanation and Reasoning
abstract
Well, I missed several buses.How on earth can you miss several buses?I, ah ..., I have got late.But there's a bus every ten minutes, and you are over 1 hour late.Have you got it now?
Deepanway Ghosal, Pengfei Hong, Navonil Majumder, Rada Mihalcea, Soujanya Poria
SIGDIAL6
2021 DOZEN: Cross-Domain Zero Shot Named Entity Recognition with Knowledge Graph
abstract
With the new developments of natural language processing, increasing attention has been given to the task of Named Entity Recognition (NER). However, the vast majority of work focus on a small number of large-scale annotated datasets with a limited number of entities such as person, location and organization. While other datasets have been introduced with domain-specific entities, the smaller size of these largely limits the applicability of state-of-the-art deep models. Even if there are promising new approaches for performing zero-shot learning (ZSL), they are not designed for a cross-domain settings. We propose Cross Domain Zero Shot Named Entity Recognition with Knowledge Graph (DOZEN), which learns the relations between entities across different domains from an existing ontology of external knowledge and a set of analogies linking entities and domains. Experiments performed on both large scale and domain-specific datasets indicate that DOZEN is the most suitable option to extracts unseen entities in a target dataset from a different domain.
Hoang Van Nguyen, Francesco Gelli, Soujanya Poria
SIGIR3
2021 Persuasive dialogue understanding: The baselines and negative results
Hui Chen 0023, Deepanway Ghosal, Navonil Majumder, Amir Hussain 0001, Soujanya Poria
Neurocomputing5
2020 KinGDOM: Knowledge-Guided DOMain Adaptation for Sentiment Analysis
abstract
Cross-domain sentiment analysis has received significant attention in recent years, prompted by the need to combat the domain gap between different applications that make use of sentiment analysis.In this paper, we take a novel perspective on this task by exploring the role of external commonsense knowledge.We introduce a new framework, KinGDOM, which utilizes the ConceptNet knowledge graph to enrich the semantics of a document by providing both domain-specific and domain-general background concepts.These concepts are learned by training a graph convolutional autoencoder that leverages inter-domain concepts in a domain-invariant manner.Conditioning a popular domain-adversarial baseline method with these learned concepts helps improve its performance over state-of-the-art approaches, demonstrating the efficacy of our proposed framework.
Deepanway Ghosal, Devamanyu Hazarika, Abhinaba Roy, Navonil Majumder, Rada Mihalcea, Soujanya Poria
ACL6
2020 SenticNet 6: Ensemble Application of Symbolic and Subsymbolic AI for Sentiment Analysis
abstract
Deep learning has unlocked new paths towards the emulation of the peculiarly-human capability of learning from examples. While this kind of bottom-up learning works well for tasks such as image classification or object detection, it is not as effective when it comes to natural language processing. Communication is much more than learning a sequence of letters and words: it requires a basic understanding of the world and social norms, cultural awareness, commonsense knowledge, etc.; all things that we mostly learn in a top-down manner. In this work, we integrate top-down and bottom-up learning via an ensemble of symbolic and subsymbolic AI tools, which we apply to the interesting problem of polarity detection from text. In particular, we integrate logical reasoning within deep learning architectures to build a new version of SenticNet, a commonsense knowledge base for sentiment analysis.
Erik Cambria, Yang Li 0055, Frank Z. Xing, Soujanya Poria, Kenneth Kwok
CIKM4
2020 MIME: MIMicking Emotions for Empathetic Response Generation
abstract
Navonil Majumder, Pengfei Hong, Shanshan Peng, Jiankun Lu, Deepanway Ghosal, Alexander Gelbukh, Rada Mihalcea, Soujanya Poria. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Navonil Majumder, Pengfei Hong, Shanshan Peng, Jiankun Lu, Deepanway Ghosal, Alexander F. Gelbukh, Rada Mihalcea, Soujanya Poria
EMNLP (1)8
2020 CMU-MOSEAS: A Multimodal Language Dataset for Spanish, Portuguese, German and French
abstract
Modeling multimodal language is a core research area in natural language processing. While languages such as English have relatively large multimodal language resources, other widely spoken languages across the globe have few or no large-scale datasets in this area. This disproportionately affects native speakers of languages other than English. As a step towards building more equitable and inclusive multimodal systems, we introduce the first large-scale multimodal language dataset for Spanish, Portuguese, German and French. The proposed dataset, called CMU-MOSEAS (CMU Multimodal Opinion Sentiment, Emotions and Attributes), is the largest of its kind with 40, 000 total labelled sentences. It covers a diverse set topics and speakers, and carries supervision of 20 labels including sentiment (and subjectivity), emotions, and attributes. Our evaluations on a state-of-the-art multimodal model demonstrates that CMU-MOSEAS enables further research for multilingual studies in multimodal language.
Amir Zadeh 0001, Yansheng Cao, Smon Hessner, Paul Pu Liang, Soujanya Poria, Louis-Philippe Morency
EMNLP (1)5
2020 MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis
abstract
Multimodal Sentiment Analysis is an active area of research that leverages multimodal signals for affective understanding of user-generated videos. The predominant approach, addressing this task, has been to develop sophisticated fusion techniques. However, the heterogeneous nature of the signals creates distributional modality gaps that pose significant challenges. In this paper, we aim to learn effective modality representations to aid the process of fusion. We propose a novel framework, MISA, which projects each modality to two distinct subspaces. The first subspace is modality-invariant, where the representations across modalities learn their commonalities and reduce the modality gap. The second subspace is modality-specific, which is private to each modality and captures their characteristic features. These representations provide a holistic view of the multimodal data, which is used for fusion that leads to task predictions. Our experiments on popular sentiment analysis benchmarks, MOSI and MOSEI, demonstrate significant gains over state-of-the-art models. We also consider the task of Multimodal Humor Detection and experiment on the recently proposed UR_FUNNY dataset. Here too, our model fares better than strong baselines, establishing MISA as a useful multimodal framework.
Devamanyu Hazarika, Roger Zimmermann, Soujanya Poria
ACM Multimedia3
2020 Dialogue systems with audio context
Tom Young, Vlad Pandelea, Soujanya Poria, Erik Cambria
Neurocomputing3
2020 Social Media Marketing and Financial Forecasting
Frank Z. Xing, Soujanya Poria, Erik Cambria, Roy E. Welsch
Inf. Process. Manag.2
2019 DialogueRNN: An Attentive RNN for Emotion Detection in Conversations
abstract
Emotion detection in conversations is a necessary step for a number of applications, including opinion mining over chat history, social media threads, debates, argumentation mining, understanding consumer feedback in live conversations, and so on. Currently systems do not treat the parties in the conversation individually by adapting to the speaker of each utterance. In this paper, we describe a new method based on recurrent neural networks that keeps track of the individual party states throughout the conversation and uses this information for emotion classification. Our model outperforms the state-of-the-art by a significant margin on two different datasets.
Navonil Majumder, Soujanya Poria, Devamanyu Hazarika, Rada Mihalcea, Alexander F. Gelbukh, Erik Cambria
AAAI2
2019 Towards Multimodal Sarcasm Detection (An _Obviously_ Perfect Paper)
abstract
Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, Soujanya Poria. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, Soujanya Poria
ACL (1)6
2019 MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations
abstract
Emotion recognition in conversations (ERC) is a challenging task that has recently gained popularity due to its potential applications.Until now, however, there has been no largescale multimodal multi-party emotional conversational database containing more than two speakers per dialogue.To address this gap, we propose the Multimodal EmotionLines Dataset (MELD), an extension and enhancement of EmotionLines.MELD contains about 13,000 utterances from 1,433 dialogues from the TV-series Friends.Each utterance is annotated with emotion and sentiment labels, and encompasses audio, visual, and textual modalities.We propose several strong multimodal baselines and show the importance of contextual and multimodal information for emotion recognition in conversations.The full dataset is available for use at http:// affective-meld.github.io.
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, Rada Mihalcea
ACL (1)1
2019 Multi-task Learning for Detecting Stance in Tweets
Devamanyu Hazarika, Gangeshwar Krishnamurthy, Soujanya Poria, Roger Zimmermann
CICLing (2)3
2019 Fusing Phonetic Features and Chinese Character Representation for Sentiment Analysis
Haiyun Peng, Soujanya Poria, Yang Li 0055, Erik Cambria
CICLing (2)2
2019 DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation
abstract
Deepanway Ghosal, Navonil Majumder, Soujanya Poria, Niyati Chhaya, Alexander Gelbukh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Deepanway Ghosal, Navonil Majumder, Soujanya Poria, Niyati Chhaya, Alexander F. Gelbukh
EMNLP/IJCNLP (1)3
2019 An Attention-Based Model for Learning Dynamic Interaction Networks
abstract
In the physical world, complex systems are generally created as the composition of multiple primitive components that interact with each other rather than a single monolithic structure. Recently, spatio-temporal graphs received a reasonable amount of attention from the research community since they emerged as a natural representational tool able to capture the interactive and interrelated structure of a complex problem. To better understand the nature of complex systems, there is the need to define models that can easily explain the learned causal relationship. To this end, we propose an attentive model able to learn and project the relational structure into a fixed-size embedding. Such representation naturally captures the dynamic influence that each neighbors exert over a given vertex providing a valuable description of the problem setting. The proposed architecture has been extensively evaluated against strong baselines on toy as well as real-world tasks, such as prediction of household energy load and traffic congestion.
Sandro Cavallari, Soujanya Poria, Erik Cambria, Vincent Wenchen Zheng, Hongyun Cai 0001
IJCNN2
2018 SenticNet 5: Discovering Conceptual Primitives for Sentiment Analysis by Means of Context Embeddings
abstract
With the recent development of deep learning, research in AI has gained new vigor and prominence. While machine learning has succeeded in revitalizing many research fields, such as computer vision, speech recognition, and medical diagnosis, we are yet to witness impressive progress in natural language understanding. One of the reasons behind this unmatched expectation is that, while a bottom-up approach is feasible for pattern recognition, reasoning and understanding often require a top-down approach. In this work, we couple sub-symbolic and symbolic AI to automatically discover conceptual primitives from text and link them to commonsense concepts and named entities in a new three-level knowledge representation for sentiment analysis. In particular, we employ recurrent neural networks to infer primitives by lexical substitution and use them for grounding common and commonsense knowledge by means of multi-dimensional scaling.
Erik Cambria, Soujanya Poria, Devamanyu Hazarika, Kenneth Kwok
AAAI2
2018 Memory Fusion Network for Multi-view Sequential Learning
abstract
Multi-view sequential learning is a fundamental problem in machine learning dealing with multi-view sequences. In a multi-view sequence, there exists two forms of interactions between different views: view-specific interactions and cross-view interactions. In this paper, we present a new neural architecture for multi-view sequential learning called the Memory Fusion Network (MFN) that explicitly accounts for both interactions in a neural architecture and continuously models them through time. The first component of the MFN is called the System of LSTMs, where view-specific interactions are learned in isolation through assigning an LSTM function to each view. The cross-view interactions are then identified using a special attention mechanism called the Delta-memory Attention Network (DMAN) and summarized through time with a Multi-view Gated Memory. Through extensive experimentation, MFN is compared to various proposed approaches for multi-view sequential learning on multiple publicly available benchmark datasets. MFN outperforms all the multi-view approaches. Furthermore, MFN outperforms all current state-of-the-art models, setting new state-of-the-art results for all three multi-view datasets.
Amir Zadeh 0001, Paul Pu Liang, Navonil Majumder, Soujanya Poria, Erik Cambria, Louis-Philippe Morency
AAAI4
2018 Multi-attention Recurrent Network for Human Communication Comprehension
abstract
Human face-to-face communication is a complex multimodal signal. We use words (language modality), gestures (vision modality) and changes in tone (acoustic modality) to convey our intentions. Humans easily process and understand face-to-face communication, however, comprehending this form of communication remains a significant challenge for Artificial Intelligence (AI). AI must understand each modality and the interactions between them that shape the communication. In this paper, we present a novel neural architecture for understanding human communication called the Multi-attention Recurrent Network (MARN). The main strength of our model comes from discovering interactions between modalities through time using a neural component called the Multi-attention Block (MAB) and storing them in the hybrid memory of a recurrent component called the Long-short Term Hybrid Memory (LSTHM). We perform extensive comparisons on six publicly available datasets for multimodal sentiment analysis, speaker trait recognition and emotion recognition. MARN shows state-of-the-art results performance in all the datasets.
Amir Zadeh 0001, Paul Pu Liang, Soujanya Poria, Prateek Vij, Erik Cambria, Louis-Philippe Morency
AAAI3
2018 Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph
abstract
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, Louis-Philippe Morency. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Amir Zadeh 0001, Paul Pu Liang, Soujanya Poria, Erik Cambria, Louis-Philippe Morency
ACL (1)3
2018 A Deep Learning Approach for Multimodal Deception Detection
Gangeshwar Krishnamurthy, Navonil Majumder, Soujanya Poria, Erik Cambria
CICLing (1)3
2018 CASCADE: Contextual Sarcasm Detection in Online Discussion Forums
abstract
The literature in automated sarcasm detection has mainly focused on lexical-, syntactic- and semantic-level analysis of text. However, a sarcastic sentence can be expressed with contextual presumptions, background and commonsense knowledge. In this paper, we propose a ContextuAl SarCasm DEtector (CASCADE), which adopts a hybrid approach of both content- and context-driven modeling for sarcasm detection in online social media discussions. For the latter, CASCADE aims at extracting contextual information from the discourse of a discussion thread. Also, since the sarcastic nature and form of expression can vary from person to person, CASCADE utilizes user embeddings that encode stylometric and personality features of users. When used along with content-based feature extractors such as convolutional neural networks, we see a significant boost in the classification performance on a large Reddit corpus.
Devamanyu Hazarika, Soujanya Poria, Sruthi Gorantla, Erik Cambria, Roger Zimmermann, Rada Mihalcea
COLING2
2018 Contextual Inter-modal Attention for Multi-modal Sentiment Analysis
abstract
Multi-modal sentiment analysis offers various challenges, one being the effective combination of different input modalities, namely text, visual and acoustic.In this paper, we propose a recurrent neural network based multi-modal attention framework that leverages the contextual information for utterance-level sentiment prediction.The proposed approach applies attention on multi-modal multi-utterance representations and tries to learn the contributing features amongst them.We evaluate our proposed approach on two multi-modal sentiment analysis benchmark datasets, viz.CMU Multi-modal Opinion-level Sentiment Intensity (CMU-MOSI) corpus and the recently released CMU Multi-modal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) corpus.Evaluation results show the effectiveness of our proposed approach with the accuracies of 82.31% and 79.80% for the MOSI and MO-SEI datasets, respectively.These are approximately 2 and 1 points performance improvement over the state-of-the-art models for the datasets.
Deepanway Ghosal, Md. Shad Akhtar, Dushyant Singh Chauhan, Soujanya Poria, Asif Ekbal, Pushpak Bhattacharyya
EMNLP4
2018 ICON: Interactive Conversational Memory Network for Multimodal Emotion Detection
abstract
Emotion recognition in conversations is crucial for building empathetic machines.Current work in this domain do not explicitly consider the inter-personal influences that thrive in the emotional dynamics of dialogues.To this end, we propose Interactive COnversational memory Network (ICON), a multimodal emotion detection framework that extracts multimodal features from conversational videos and hierarchically models the selfand interspeaker emotional influences into global memories.Such memories generate contextual summaries which aid in predicting the emotional orientation of utterance-videos.Our model outperforms state-of-the-art networks on multiple classification and regression tasks in two benchmark datasets.
Devamanyu Hazarika, Soujanya Poria, Rada Mihalcea, Erik Cambria, Roger Zimmermann
EMNLP2
2018 IARM: Inter-Aspect Relation Modeling with Memory Networks in Aspect-Based Sentiment Analysis
abstract
Sentiment analysis has immense implications in modern businesses through user-feedback mining.Large product-based enterprises like Samsung and Apple make crucial business decisions based on the large quantity of user reviews and suggestions available in different e-commerce websites and social media platforms like Amazon and Facebook.Sentiment analysis caters to these needs by summarizing user sentiment behind a particular object.In this paper, we present a novel approach of incorporating the neighboring aspects related information into the sentiment classification of the target aspect using memory networks.Our method outperforms the state of the art by 1.6% on average in two distinct domains.
Navonil Majumder, Soujanya Poria, Alexander F. Gelbukh, Md. Shad Akhtar, Erik Cambria, Asif Ekbal
EMNLP2
2018 Conversational Memory Network for Emotion Recognition in Dyadic Dialogue Videos
abstract
Emotion recognition in conversations is crucial for the development of empathetic machines. Present methods mostly ignore the role of inter-speaker dependency relations while classifying emotions in conversations. In this paper, we address recognizing utterance-level emotions in dyadic conversational videos. We propose a deep neural framework, termed conversational memory network, which leverages contextual information from the conversation history. The framework takes a multimodal approach comprising audio, visual and textual features with gated recurrent units to model past utterances of each speaker into memories. Such memories are then merged using attention-based hops to capture inter-speaker dependencies. Experiments show an accuracy improvement of 3-4% over the state of the art.
Devamanyu Hazarika, Soujanya Poria, Amir Zadeh 0001, Erik Cambria, Louis-Philippe Morency, Roger Zimmermann
NAACL-HLT2
2018 Multimodal sentiment analysis using hierarchical fusion with context modeling
Navonil Majumder, Devamanyu Hazarika, Alexander F. Gelbukh, Erik Cambria, Soujanya Poria
Knowl. Based Syst.5
2017 Context-Dependent Sentiment Analysis in User-Generated Videos
abstract
Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, Louis-Philippe Morency. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017.
Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh 0001, Louis-Philippe Morency
ACL (1)1
2017 Benchmarking Multimodal Sentiment Analysis
Erik Cambria, Devamanyu Hazarika, Soujanya Poria, Amir Hussain 0001, R. B. V. Subramanyam
CICLing (2)3
2017 Tensor Fusion Network for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis is an increasingly popular research area, which extends the conventional language-based definition of sentiment analysis to a multimodal setup where other relevant modalities accompany language.In this paper, we pose the problem of multimodal sentiment analysis as modeling intra-modality and inter-modality dynamics.We introduce a novel model, termed Tensor Fusion Network, which learns both such dynamics end-to-end.The proposed approach is tailored for the volatile nature of spoken language in online videos as well as accompanying gestures and voice.In the experiments, our model outperforms state-ofthe-art approaches for both multimodal and unimodal sentiment analysis.
Amir Zadeh 0001, Minghai Chen, Soujanya Poria, Erik Cambria, Louis-Philippe Morency
EMNLP3
2017 Multi-level Multiple Attentions for Contextual Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis involves identifying sentiment in videos and is a developing field of research. Unlike current works, which model utterances individually, we propose a recurrent model that is able to capture contextual information among utterances. In this paper, we also introduce attentionbased networks for improving both context learning and dynamic feature fusion. Our model shows 6-8% improvement over the state of the art on a benchmark dataset.
Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh 0001, Louis-Philippe Morency
ICDM1
2017 Ensemble application of convolutional neural networks and multiple kernel learning for multimodal sentiment analysis
Soujanya Poria, Haiyun Peng, Amir Hussain 0001, Newton Howard, Erik Cambria
Neurocomputing1
2016 SenticNet 4: A Semantic Resource for Sentiment Analysis Based on Conceptual Primitives
abstract
An important difference between traditional AI systems and human intelligence is the human ability to harness commonsense knowledge gleaned from a lifetime of learning and experience to make informed decisions. This allows humans to adapt easily to novel situations where AI fails catastrophically due to a lack of situation-specific rules and generalization capabilities. Commonsense knowledge also provides background information that enables humans to successfully operate in social situations where such knowledge is typically assumed. Since commonsense consists of information that humans take for granted, gathering it is an extremely difficult task. Previous versions of SenticNet were focused on collecting this kind of knowledge for sentiment analysis but they were heavily limited by their inability to generalize. SenticNet 4 overcomes such limitations by leveraging on conceptual primitives automatically generated by means of hierarchical clustering and dimensionality reduction.
Erik Cambria, Soujanya Poria, Rajiv Bajpai, Björn W. Schuller
COLING2
2016 A Deeper Look into Sarcastic Tweets Using Deep Convolutional Neural Networks
abstract
Sarcasm detection is a key task for many natural language processing tasks. In sentiment analysis, for example, sarcasm can flip the polarity of an “apparently positive” sentence and, hence, negatively affect polarity detection performance. To date, most approaches to sarcasm detection have treated the task primarily as a text categorization problem. Sarcasm, however, can be expressed in very subtle ways and requires a deeper understanding of natural language that standard text categorization techniques cannot grasp. In this work, we develop models based on a pre-trained convolutional neural network for extracting sentiment, emotion and personality features for sarcasm detection. Such features, along with the network’s baseline features, allow the proposed models to outperform the state of the art on benchmark datasets. We also address the often ignored generalizability issue of classifying data that have not been seen by the models at learning phase.
Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Prateek Vij
COLING1
2016 Convolutional MKL Based Multimodal Emotion Recognition and Sentiment Analysis
abstract
Technology has enabled anyone with an Internet connection to easily create and share their ideas, opinions and content with millions of other people around the world. Much of the content being posted and consumed online is multimodal. With billions of phones, tablets and PCs shipping today with built-in cameras and a host of new video-equipped wearables like Google Glass on the horizon, the amount of video on the Internet will only continue to increase. It has become increasingly difficult for researchers to keep up with this deluge of multimodal content, let alone organize or make sense of it. Mining useful knowledge from video is a critical need that will grow exponentially, in pace with the global growth of content. This is particularly important in sentiment analysis, as both service and product reviews are gradually shifting from unimodal to multimodal. We present a novel method to extract features from visual and textual modalities using deep convolutional neural networks. By feeding such features to a multiple kernel learning classifier, we significantly outperform the state of the art of multimodal emotion recognition and sentiment analysis on different datasets.
Soujanya Poria, Iti Chaturvedi, Erik Cambria, Amir Hussain 0001
ICDM1
2016 Sentic LDA: Improving on LDA with semantic similarity for aspect-based sentiment analysis
abstract
The advent of the Social Web has provided netizens with new tools for creating and sharing, in a time- and cost-efficient way, their contents, ideas, and opinions with virtually the millions of people connected to the World Wide Web. This huge amount of information, however, is mainly unstructured as specifically produced for human consumption and, hence, it is not directly machine-processable. In order to enable a more efficient passage from unstructured information to structured data, aspect-based opinion mining models the relations between opinion targets contained in a document and the polarity values associated with these. Because aspects are often implicit, however, spotting them and calculating their respective polarity is an extremely difficult task, which is closer to natural language understanding rather than natural language processing. To this end, Sentic LDA exploits common-sense reasoning to shift LDA clustering from a syntactic to a semantic level. Rather than looking at word co-occurrence frequencies, Sentic LDA leverages on the semantics associated with words and multi-word expressions to improve clustering and, hence, outperform state-of-the-art techniques for aspect extraction.
Soujanya Poria, Iti Chaturvedi, Erik Cambria, Federica Bisio
IJCNN1
2016 Fusing audio, visual and textual clues for sentiment analysis from multimodal content
Soujanya Poria, Erik Cambria, Newton Howard, Guang-Bin Huang, Amir Hussain 0001
Neurocomputing1
2016 Aspect extraction for opinion mining with a deep convolutional neural network
Soujanya Poria, Erik Cambria, Alexander F. Gelbukh
Knowl. Based Syst.1
2015 AffectiveSpace 2: Enabling Affective Intuition for Concept-Level Sentiment Analysis
abstract
Predicting the affective valence of unknown multi-word expressions is key for concept-level sentiment analysis. AffectiveSpace 2 is a vector space model, built by means of random projection, that allows for reasoning by analogy on natural language con- cepts. By reducing the dimensionality of affec- tive common-sense knowledge, the model allows semantic features associated with concepts to be generalized and, hence, allows concepts to be intu- itively clustered according to their semantic and affective relatedness. Such an affective intuition (so called because it does not rely on explicit fea- tures, but rather on implicit analogies) enables the inference of emotions and polarity conveyed by multi-word expressions, thus achieving efficient concept-level sentiment analysis.
Erik Cambria, Jie Fu 0001, Federica Bisio, Soujanya Poria
AAAI4
2015 The CLSA Model: A Novel Framework for Concept-Level Sentiment Analysis
Erik Cambria, Soujanya Poria, Federica Bisio, Rajiv Bajpai, Iti Chaturvedi
CICLing (2)2
2015 Modelling Public Sentiment in Twitter: Using Linguistic Patterns to Enhance Supervised Learning
Prerna Chikersal, Soujanya Poria, Erik Cambria, Alexander F. Gelbukh, Chng Eng Siong
CICLing (2)2
2015 Deep Convolutional Neural Network Textual Features and Multiple Kernel Learning for Utterance-level Multimodal Sentiment Analysis
abstract
We present a novel way of extracting features from short texts, based on the activation values of an inner layer of a deep convolutional neural network.We use the extracted features in multimodal sentiment analysis of short video clips representing one sentence each.We use the combined feature vectors of textual, visual, and audio modalities to train a classifier based on multiple kernel learning, which is known to be good at heterogeneous data.We obtain 14% performance improvement over the state of the art and present a parallelizable decision-level data fusion method, which is much faster, though slightly less accurate.
Soujanya Poria, Erik Cambria, Alexander F. Gelbukh
EMNLP1
2015 Towards an intelligent framework for multimodal affective data analysis
Soujanya Poria, Erik Cambria, Amir Hussain 0001, Guang-Bin Huang
Neural Networks1
2014 Dependency-Based Semantic Parsing for Concept-Level Text Analysis
Soujanya Poria, Basant Agarwal, Alexander F. Gelbukh, Amir Hussain 0001, Newton Howard
CICLing (1)1
2014 Sentic patterns: Dependency-based rules for concept-level sentiment analysis
Soujanya Poria, Erik Cambria, Grégoire Winterstein, Guang-Bin Huang
Knowl. Based Syst.1
2014 EmoSenticSpace: A novel framework for affective common-sense reasoning
Soujanya Poria, Alexander F. Gelbukh, Erik Cambria, Amir Hussain 0001, Guang-Bin Huang
Knowl. Based Syst.1
2012 A Classifier Based Approach to Emotion Lexicon Construction
Dipankar Das 0001, Soujanya Poria, Sivaji Bandyopadhyay
NLDB2