Anutosh Maitra

dblp:78/424 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
11since 2021 · last 2027
0000-0002-1346-8783ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2027 Can language models persuade? Exploring the persuasive efficacy of Large Language and Vision Language models
Rohan Kirti, Atharva Deshmukh, Kiran K. Dugana, Yash Rathore, Shipra Shriparn, Sriparna Saha 0001, Roshni R. Ramnani, Anutosh Maitra
Comput. Speech Lang.8
2024 QDETRv: Query-Guided DETR for One-Shot Object Localization in Videos
abstract
In this work, we study one-shot video object localization problem that aims to localize instances of unseen objects in the target video using a single query image of the object. Toward addressing this challenging problem, we extend a popular and successful object detection method, namely DETR (Detection Transformer), and introduce a novel approach –query-guided detection transformer for videos (QDETRv). A distinctive feature of QDETRv is its capacity to exploit information from the query image and spatio-temporal context of the target video, which significantly aids in precisely pinpointing the desired object in the video. We incorporate cross-attention mechanisms that capture temporal relationships across adjacent frames to handle the dynamic context in videos effectively. Further, to ensure strong initialization for QDETRv, we also introduce a novel unsupervised pretraining technique tailored to videos. This involves training our model on synthetic object trajectories with an analogous objective as the query-guided localization task. During this pretraining phase, we incorporate recurrent object queries and loss functions that encourage accurate patch feature reconstruction. These additions enable better temporal understanding and robust representation learning. Our experiments show that the proposed model significantly outperforms the competitive baselines on two public benchmarks, VidOR and ImageNet-VidVRD, extended for one-shot open-set localization tasks.
Yogesh Kumar 0004, Saswat Mallick, Anand Mishra 0001, Sowmya Rasipuram, Anutosh Maitra, Roshni R. Ramnani
AAAI5
2023 Sentiment Aided Graph Attentive Contextualization for Task Oriented Negotiation Dialogue Generation
abstract
Aritra Raut, Sriparna Saha, Anutosh Maitra, Roshni Ramnani. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Aritra Raut, Sriparna Saha 0001, Anutosh Maitra, Roshni R. Ramnani
IJCNLP (1)3
2023 Towards personalized persuasive dialogue generation for adversarial task oriented dialogue setting
Abhisek Tiwari, Abhijeet Khandwe, Sriparna Saha 0001, Roshni R. Ramnani, Anutosh Maitra, Shubhashis Sengupta
Expert Syst. Appl.5
2022 Introducing Multi-modality in Persuasive Task Oriented Virtual Sales Agent
Aritra Raut, Abhisek Tiwari, Sriparna Saha 0001, Anutosh Maitra, Roshni R. Ramnani, Shubhashis Sengupta
ICONIP (3)5
2022 Towards Sentiment and Emotion aided Intent Detection
abstract
Intent detection is one of the crucial Natural Language Understanding(NLU) tasks studied extensively. Misclassification of intents impacts the overall performance of the conversational systems as natural language understanding is the first mean of interaction between a user and a virtual agent. The traditional approach of intent detection is limited only to textual features of user utterances and overlooks other semantic features. The sentiment and emotional state of the speaker are two such semantic features that have essential impacts on intent detection, as they implicitly express user intention conveyed through the user’s message. Depending on the context, these features can help the dialogue agent respond to the same intent with varying degrees of sentiment and emotion. Thus, investigating the role of sentiment and emotion on intent detection is a matter of great interest. The current work investigates the impact of utilizing sentiment and emotion information on intent detection tasks and proposes emotion and sentiment aided intent detection models. We also investigate the impact of sentiment and emotion using three different multitasking frameworks and present a joint model that utilizes the co-relation information across these tasks to correctly identify all these NLU aspects (intent, sentiment, and emotion). The obtained experimental results by the proposed models outperform several baselines and state-of-the-art intent detection models on multiple datasets by a significant margin of 1.8% - 3%, demonstrating the significant role of sentiment and emotion features in intent detection1.
Ashutosh Kumar Trivedi, Sriparna Saha 0001, Anutosh Maitra, Roshni R. Ramnani, Shubhashis Sengupta
ICPR4
2022 Multimodal Depression Detection Using Task-oriented Transformer-based Embedding
abstract
Automatic depression detection from multimodal signals assumes significance since the features from different modalities possesses contrasting but relevant information. Multimodal signal analysis involving the derivation of a joint representation, identification of correct alignment, and selection of appropriate fusing mechanisms is a challenging task. In this paper, we focus on an approach to detecting depression from multimodal signals using task-specific representation and late fusion. Our method uses a task-oriented embedding generated through a GPT2-medium language model that is fine-tuned on a conversational dataset. We present our experimental results benchmarking on Distress Analysis Interview Corpus - a Wizard of Oz (DAIC-WOZ) corpus. Our proposed task-oriented embedding on a simple BiLSTM network outperforms the previous text-based depression prediction methods. The results are further strengthened when audio and video modalities are fused with the text predictions on a late-fusion network.
Sowmya Rasipuram, Junaid Hamid Bhat, Anutosh Maitra, Bishal Shaw, Sriparna Saha 0001
ISCC3
2022 Hollywood Identity Bias Dataset: A Context Oriented Bias Analysis of Movie Dialogues
abstract
Movies reflect society and also hold power to transform opinions. Social biases and stereotypes present in movies can cause extensive damage due to their reach. These biases are not always found to be the need of storyline but can creep in as the author’s bias. Movie production houses would prefer to ascertain that the bias present in a script is the story’s demand. Today, when deep learning models can give human-level accuracy in multiple tasks, having an AI solution to identify the biases present in the script at the writing stage can help them avoid the inconvenience of stalled release, lawsuits, etc. Since AI solutions are data intensive and there exists no domain specific data to address the problem of biases in scripts, we introduce a new dataset of movie scripts that are annotated for identity bias. The dataset contains dialogue turns annotated for (i) bias labels for seven categories, viz., gender, race/ethnicity, religion, age, occupation, LGBTQ, and other, which contains biases like body shaming, personality bias, etc. (ii) labels for sensitivity, stereotype, sentiment, emotion, emotion intensity, (iii) all labels annotated with context awareness, (iv) target groups and reason for bias labels and (v) expert-driven group-validation process for high quality annotations. We also report various baseline performances for bias identification and category detection on our dataset.
Sandhya Singh, Prapti Roy, Nihar Sahoo, Niteesh Mallela, Pushpak Bhattacharyya, Milind Savagaonkar, Nidhi 0002, Roshni R. Ramnani, Anutosh Maitra, Shubhashis Sengupta
LREC10
2022 A persona aware persuasive dialogue policy for dynamic and co-operative goal setting
Abhisek Tiwari, Tulika Saha, Sriparna Saha 0001, Shubhashis Sengupta, Anutosh Maitra, Roshni R. Ramnani, Pushpak Bhattacharyya
Expert Syst. Appl.5
2021 Unsupervised Approach for Knowledge-Graph Creation from Conversation: The Use of Intent Supervision for Slot Filling
abstract
In this paper, we propose an unsupervised approach for knowledge graph (KG) creation from conversational data. We make use of intent classification and slot-filling, the two important components of any dialogue agent, exploit their interconnectedness, and finally construct a KG. We build a supervised intent classifier to extract the intent classes, and then on top of this we run our occlusion based slot-information extraction algorithm. Our algorithm is able to make use of supervised training of intent classifiers for extracting the relevant slot-information in an unsupervised way. To test the effectiveness of our system, we perform both automatic and manual evaluation of our intent-classifier and slot-filling system on three dialog datasets. Finally, we construct a knowledge graph from the dialogue conversation using an algorithm that makes use of our occlusion based slot-information extraction module. Empirical evaluation shows that our occlusion based method is able to successfully extract slot information from conversations, resulting in a high-quality KG.
Zishan Ahmad, Asif Ekbal, Shubhashis Sengupta, Anutosh Maitra, Roshni R. Ramnani, Pushpak Bhattacharyya
IJCNN4
2021 Multi-Modal Dialogue Policy Learning for Dynamic and Co-operative Goal Setting
abstract
Developing an adequate and human-like virtual agent has been one of the primary applications of artificial intelligence. In the last few years, task-oriented dialogue systems have gained huge popularity because of their upsurging relevance and positive outcomes. In real-world, users may not always have a predefined and rigid task goal beforehand; they upgrade/downgrade/change their goal component dynamically depending upon their utility value and agent's serving capability. However, existing virtual agents fail to incorporate this dynamic behavior, leading to either unsuccessful task completion or an ungratified user experience. The paper presents an end to end multimodal dialogue system for dynamic and co-operative goal setting, which incorporates i) a multi-modal semantic state representation in policy learning to deal with multi-modal inputs, ii) a goal manager module in a traditional dialogue manager for handling dynamic and goal unavailability scenarios effectively, iii) an accumulative reward (task/persona/sentiment) for task success, personalized persuasion and user-adaptive behavior, respectively. The obtained experimental results and the comparisons with baselines firmly establish the need and efficacy of the proposed system.
Abhisek Tiwari, Tulika Saha, Sriparna Saha 0001, Shubhashis Sengupta, Anutosh Maitra, Roshni R. Ramnani, Pushpak Bhattacharyya
IJCNN5
2020 Multi-modal Sequence-to-sequence Model for Continuous Affect Prediction in the Wild Using Deep 3D Features
abstract
Continuous affect prediction in the wild is a very interesting problem and is challenging as continuous prediction involves heavy computation. This paper presents the methodologies and techniques used in our contribution to predict continuous emotion dimensions i.e., valence and arousal in ABAW competition on Aff-Wild2 database. Aff-Wild2 database consists of videos in the wild labelled for valence and arousal at frame level. Our proposed methodology uses fusion of both audio and video features (multi-modal) extracted using state-of-the-art methods. These audio-video features are used to train a sequence-to-sequence model that is based on Gated Recurrent Units (GRU). We show promising results on validation data with simple architecture. The overall valence and arousal of the proposed approach is 0.22 and 0.34, which is better than the competition baseline of 0.14 and 0.24 respectively.
Sowmya Rasipuram, Junaid Hamid Bhat, Anutosh Maitra
FG3
2020 Multi-modal Expression Recognition in the Wild Using Sequence Modeling
abstract
As we exceed upon the procedures for modelling the different aspects of behaviour, expression recognition has become a key field of research in Human Computer Interactions. Expression recognition in the wild is a very interesting problem and is challenging as it involves detailed feature extraction and heavy computation. This paper presents the methodologies and techniques used in our contribution to recognize different expressions i.e., neutral, anger, disgust, fear, happiness, sadness, surprise in ABAW competition on Aff-Wild2 database. Aff-Wild2 database consists of videos in the wild labelled for seven different expressions at frame level. We used a bi-modal approach by fusing audio and visual features and train a sequence-to-sequence model that is based on Gated Recurrent Units (GRU) and Long Short Term Memory (LSTM) network. We show experimental results on validation data. The overall accuracy of the proposed approach is 41.5 %, which is better than the competition baseline of 37%.
Sowmya Rasipuram, Junaid Hamid Bhat, Anutosh Maitra
FG3
2020 Using Deep 3D Features and an LSTM Based Sequence Model for Automatic Pain Detection in the Wild
abstract
Automatic pain detection is an important problem in diagnostic and therapeutic applications. In this paper, we aim to develop a computational framework to automatically detect pain in videos in the wild. The videos in the wild vary with respect to gender, age, ethnicity and even other qualitative attributes like upbringing. Previous systems focused on methodologies confined to one particular dataset that is hard to generalize for the population in the wild, or based on invasive methods that collect data using many physiological sensors and induced stressors. We propose a method to automatically detect pain in videos using state-of-the-art expression recognition system along with deep learning. We curated a dataset of 194 videos in the wild with pain and non-pain. We used a sliding window strategy to obtain a fixed-length input sample for the LSTM (Long Short Term Memory) network. We then carefully concatenate the network output of every segment to generate a video-level output. The proposed end-to-end framework can predict binary classification label (pain/non-pain) at video level. Our method achieves promising results on the dataset we collected.
Sowmya Rasipuram, Bukka Nikhil Sai, Dinesh Babu Jayagopi, Anutosh Maitra
FG4
2020 Enabling Interactive Answering of Procedural Questions
Anutosh Maitra, Shivam Garg 0005, Shubhashis Sengupta
NLDB1
2018 Can Taxonomy Help? Improving Semantic Question Matching using Question Taxonomy
abstract
In this paper, we propose a hybrid technique for semantic question matching. It uses a proposed two-layered taxonomy for English questions by augmenting state-of-the-art deep learning models with question classes obtained from a deep learning based question classifier. Experiments performed on three open-domain datasets demonstrate the effectiveness of our proposed approach. We achieve state-of-the-art results on partial ordering question ranking (POQR) benchmark dataset. Our empirical analysis shows that coupling standard distributional features (provided by the question encoder) with knowledge from taxonomy is more effective than either deep learning or taxonomy-based knowledge alone.
Rajkumar Pujari, Asif Ekbal, Pushpak Bhattacharyya, Anutosh Maitra, Tom Geo Jain, Shubhashis Sengupta
COLING5
2007 Situation Awareness Unified Process
abstract
This paper identifies the process for developing information system for situation awareness application domain and identifies novel process artifacts that need to be introduced along with existing approaches. Appropriately engineered method is an important requirement for successful implementation of any software system targeted at situation awareness. When it comes to information system development (ISD) for dynamic organizations, the method engineering plays even more critical role. The existing approaches for architectural description, method composition and process guidance are derived based on experiences from successful past implementations. Yet, how existing team utilizes the experience is the determining factor for its efficient use. Today's organization, where one information system is the result of continuous efforts of multiple teams forming a virtual organization, calls for extending the approaches to get proper benefit from method engineering.
Vikrambhai S. Sorathia, Anutosh Maitra
ICSEA2