Amitava Das 0001

dblp:75/5002-1 · DBLP profile ↗
← Back
33ranked-venue papers
4as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 3 first-author · 15 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 DETONATE - A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization
abstract
Alignment is crucial for text-to-image (T2I) models to ensure that the generated images faithfully capture user intent while maintaining safety and fairness. Direct Preference Optimization (DPO) has emerged as a key alignment technique for large language models (LLMs), and its influence is now extending to T2I systems. This paper introduces DPO-Kernels for T2I models, a novel extension of DPO that enhances alignment across three key dimensions: (i) Hybrid Loss, which integrates embedding-based objectives with the traditional probability-based loss to improve optimization; (ii) Kernelized Representations, leveraging Radial Basis Function (RBF), Polynomial, and Wavelet kernels to enable richer feature transformations, ensuring better separation between safe and unsafe inputs; and (iii) Divergence Selection, expanding beyond DPO’s default Kullback–Leibler (KL) regularizer by incorporating alternative divergence measures such as Wasserstein and Rényi divergences to enhance stability and robustness in alignment training. We introduce DETONATE, the first large-scale benchmark of its kind, comprising approximately 100K curated image pairs, categorized as chosen and rejected. This benchmark encapsulates three critical axes of social bias and discrimination: Race, Gender, and Disability. The prompts are sourced from the hate speech datasets, while the images are generated using state-of-the-art T2I models, including Stable Diffusion 3.5 Large (SD-3.5), Stable Diffusion XL (SD-XL), and Midjourney. Furthermore, to evaluate alignment beyond surface metrics, we introduce the Alignment Quality Index (AQI) for T2I systems: a novel geometric measure that quantifies latent space separability of safe/unsafe image activations, revealing hidden model vulnerabilities. While alignment techniques often risk overfitting, we empirically demonstrate that DPO-Kernels preserve strong generalization bounds using the theory of Heavy-Tailed Self-Regularization (HT-SR).
Renjith Prasad Kaippilly Mana, Abhilekh Borah, Hasnat Md Abdullah, Chathurangi Shyalika, Ritvik Garimella, Rajarshi Roy 0007, Harshul Raj Surana, Nasrin Imanpour, Suranjana Trivedy, Amit P. Sheth, Amitava Das 0001
AAAI12
2025 KnowledgePrompts: Exploring the Abilities of Large Language Models to Solve Proportional Analogies via Knowledge-Enhanced Prompting
abstract
Making analogies is fundamental to cognition. Proportional analogies, which consist of four terms, are often used to assess linguistic and cognitive abilities. For instance, completing analogies like “Oxygen is to Gas as < blank > is to < blank >" requires identifying the semantic relationship (e.g., “type of”) between the first pair of terms (“Oxygen” and “Gas”) and finding a second pair that shares the same relationship (e.g., “Aluminum” and “Metal”). In this work, we introduce a 15K Multiple-Choice Question Answering (MCQA) dataset for proportional analogy completion and evaluate the performance of contemporary Large Language Models (LLMs) in various knowledge-enhanced prompt settings. Specifically, we augment prompts with three types of knowledge: exemplar, structured, and targeted. Our results show that despite extensive training data, solving proportional analogies remains challenging for current LLMs, with the best model achieving an accuracy of 55%. Notably, we find that providing targeted knowledge can better assist models in completing proportional analogies compared to providing exemplars or collections of structured knowledge. Our code and data are available at: https://github.com/Thiliniiw/KnowledgePrompts/
Thilini Wijesiriwardene, Ruwan Wickramarachchi, Sreeram Vennam, Vinija Jain, Aman Chadha, Amitava Das 0001, Ponnurangam Kumaraguru, Amit P. Sheth
COLING6
2025 Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations
abstract
Abhilekh Borah, Chhavi Sharma, Danush Khanna, Utkarsh Bhatt, Gurpreet Singh, Hasnat Md Abdullah, Raghav Kaushik Ravi, Vinija Jain, Jyoti Patel, Shubham Singh, Vasu Sharma, Arpita Vats, Rahul Raja, Aman Chadha, Amitava Das. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Abhilekh Borah, Chhavi Sharma, Danush Khanna, Utkarsh Bhatt, Hasnat Md Abdullah, Raghav Kaushik Ravi, Vinija Jain, Jyoti Patel, Vasu Sharma, Arpita Vats, Rahul Raja, Aman Chadha, Amitava Das 0001
EMNLP15
2025 A deep dive into automated sexism detection using fine-tuned deep learning and large language models
Advaitha Vetagiri, Partha Pakray, Amitava Das 0001
Eng. Appl. Artif. Intell.3
2024 Are my answers medically accurate? Exploiting medical knowledge graphs for medical question answering
Aizan Zafar, Deeksha Varshney, Sovan Kumar Sahoo, Amitava Das 0001, Asif Ekbal
Appl. Intell.4
2024 Racists spreader is narcissistic; sexists is Machiavellian Influence of Psycho-Sociological Facets in hate-speech diffusion prediction
Srinivas PYKL, Amitava Das 0001, Viswanath Pulabaigari
Expert Syst. Appl.2
2024 KI-MAG: A knowledge-infused abstractive question answering system in medical domain
Aizan Zafar, Sovan Kumar Sahoo, Harsh Bhardawaj, Amitava Das 0001, Asif Ekbal
Neurocomputing4
2024 KIMedQA: towards building knowledge-enhanced medical QA models
Aizan Zafar, Sovan Kumar Sahoo, Deeksha Varshney, Amitava Das 0001, Asif Ekbal
J. Intell. Inf. Syst.4
2023 FACTIFY-5WQA: 5W Aspect-based Fact Verification through Question Answering
abstract
Anku Rani, S.M Towhidul Islam Tonmoy, Dwip Dalal, Shreya Gautam, Megha Chakraborty, Aman Chadha, Amit Sheth, Amitava Das. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Anku Rani, S. M. Towhidul Islam Tonmoy, Dwip Dalal, Shreya Gautam, Megha Chakraborty, Aman Chadha, Amit P. Sheth, Amitava Das 0001
ACL (1)8
2023 FACTIFY3M: A benchmark for multimodal fact verification with explainability through 5W Question-Answering
abstract
Megha Chakraborty, Khushbu Pahwa, Anku Rani, Shreyas Chatterjee, Dwip Dalal, Harshit Dave, Ritvik G, Preethi Gurumurthy, Adarsh Mahor, Samahriti Mukherjee, Aditya Pakala, Ishan Paul, Janvita Reddy, Arghya Sarkar, Kinjal Sensharma, Aman Chadha, Amit Sheth, Amitava Das. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Megha Chakraborty, Khushbu Pahwa, Anku Rani, Shreyas Chatterjee, Dwip Dalal, Harshit Dave, Ritvik Garimella, Preethi Gurumurthy, Adarsh Mahor, Samahriti Mukherjee, Aditya Pakala, Ishan Paul, Janvita Reddy, Arghya Sarkar, Kinjal Sensharma, Aman Chadha, Amit P. Sheth, Amitava Das 0001
EMNLP18
2023 Counter Turing Test (CT2): AI-Generated Text Detection is Not as Easy as You May Think - Introducing AI Detectability Index (ADI)
abstract
Megha Chakraborty, S.M Towhidul Islam Tonmoy, S M Mehedi Zaman, Shreya Gautam, Tanay Kumar, Krish Sharma, Niyar Barman, Chandan Gupta, Vinija Jain, Aman Chadha, Amit Sheth, Amitava Das. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Megha Chakraborty, S. M. Towhidul Islam Tonmoy, S. M. Mehedi Zaman, Shreya Gautam, Tanay Kumar, Krish Sharma, Niyar R. Barman, Chandan Gupta, Vinija Jain, Aman Chadha, Amit P. Sheth, Amitava Das 0001
EMNLP12
2023 The Troubling Emergence of Hallucination in Large Language Models - An Extensive Definition, Quantification, and Prescriptive Remediations
abstract
Vipula Rawte, Swagata Chakraborty, Agnibh Pathak, Anubhav Sarkar, S.M Towhidul Islam Tonmoy, Aman Chadha, Amit Sheth, Amitava Das. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Vipula Rawte, Swagata Chakraborty, Agnibh Pathak, Anubhav Sarkar, S. M. Towhidul Islam Tonmoy, Aman Chadha, Amit P. Sheth, Amitava Das 0001
EMNLP8
2022 PESTO: Switching Point Based Dynamic and Relative Positional Encoding for Code-Mixed Languages (Student Abstract)
abstract
NLP applications for code-mixed (CM) or mix-lingual text have gained a significant momentum recently, the main reason being the prevalence of language mixing in social media communications in multi-lingual societies like India, Mexico, Europe, parts of USA etc. Word embeddings are basic building blocks of any NLP system today, yet, word embedding for CM languages is an unexplored territory. The major bottleneck for CM word embeddings is switching points, where the language switches. These locations lack in contextually and statistical systems fail to model this phenomena due to high variance in the seen examples. In this paper we present our initial observations on applying switching point based positional encoding techniques for CM language, specifically Hinglish (Hindi - English). Results are only marginally better than SOTA, but it is evident that positional encoding could be an effective way to train position sensitive language models for CM text.
Kandukuri Sai Teja, Sumanth Manduru, Parth Patwa, Amitava Das 0001
AAAI5
2022 Memotion Analysis through the Lens of Joint Embedding (Student Abstract)
abstract
Joint embedding (JE) is a way to encode multi-modal data into a vector space where text remains as the grounding key and other modalities like image are to be anchored with such keys. Meme is typically an image with embedded text onto it. Although, memes are commonly used for fun, they could also be used to spread hate and fake information. That along with its growing ubiquity over several social platforms has caused automatic analysis of memes to become a widespread topic of research. In this paper, we report our initial experiments on Memotion Analysis problem through joint embeddings. Results are marginally yielding SOTA.
Nethra Gunti, Sathyanarayanan Ramamoorthy, Parth Patwa, Amitava Das 0001
AAAI4
2022 Half-Day Tutorial on Combating Online Hate Speech: The Role of Content, Networks, Psychology, User Behavior, etc
abstract
While the rise in popularity of social media is seen as a hugely positive development, it is also accompanied by a proliferation of hate speech, which has recently become a major concern. On the one hand, hateful content creates an unsafe environment for certain members of society. On the other hand, manual moderation causes distress to content moderators, and the volume of harmful content is far beyond what human moderators can manually flag and react to. Thus, researchers in machine learning, social computing, and other areas have worked on developing tools to help automate the process. While initially studied as a text classification problem, over time, researchers realized that hate speech is multi-faceted and requires analysis of the role of linguistic expressions, context, and network structure, while using inspiration from psychology and user behavior, among others. With this in mind, we provide a holistic view of what the research community has explored so far, and what we believe are promising future research directions.
Sarah Masud, Pinkesh Pinkesh, Amitava Das 0001, Manish Gupta 0001, Preslav Nakov, Tanmoy Chakraborty 0002
WSDM3
2022 Fake spreader is narcissist; Real spreader is Machiavellian prediction of fake news diffusion using psycho-sociological facets
Srinivas PYKL, Amitava Das 0001, Viswanath Pulabaigari
Expert Syst. Appl.2
2021 Hate is the New Infodemic: A Topic-aware Modeling of Hate Speech Diffusion on Twitter
abstract
Online hate speech, particularly over microblogging platforms like Twitter, has emerged as arguably the most severe issue of the past decade. Several countries have reported a steep rise in hate crimes infuriated by malicious hate campaigns. While the detection of hate speech is one of the emerging research areas, the generation and spread of topic-dependent hate in the information network remain under-explored. In this work, we focus on exploring user behavior, which triggers the genesis of hate speech on Twitter and how it diffuses via retweets. We crawl a large-scale dataset of tweets, retweets, user activity history, and follower networks, comprising over 161 million tweets from more than 41 million unique users. We also collect over 600k contemporary news articles published online. We characterize different signals of information that govern these dynamics. Our analyses differentiate the diffusion dynamics in the presence of hate from usual information diffusion. This motivates us to formulate the modeling problem in a topic-aware setting with real-world knowledge. For predicting the initiation of hate speech for any given hashtag, we propose multiple feature-rich models, with the best performing one achieving a macro F1 score of 0.65. Meanwhile, to predict the retweet dynamics on Twitter, we propose RETINA, a novel neural architecture that incorporates exogenous influence using scaled dot-product attention. RETINA achieves a macro F1-score of 0.85, outperforming multiple state-of-the-art models. Our analysis reveals the superlative power of RETINA to predict the retweet dynamics of hateful content compared to the existing diffusion models.
Sarah Masud, Subhabrata Dutta, Sakshi Makkar, Chhavi Jain, Vikram Goyal, Amitava Das 0001, Tanmoy Chakraborty 0002
ICDE6
2020 Minority Positive Sampling for Switching Points - an Anecdote for the Code-Mixing Language Modeling
abstract
Code-Mixing (CM) or language mixing is a social norm in multilingual societies. CM is quite prevalent in social media conversations in multilingual regions like - India, Europe, Canada and Mexico. In this paper, we explore the problem of Language Modeling (LM) for code-mixed Hinglish text. In recent times, there have been several success stories with neural language modeling like Generative Pre-trained Transformer (GPT) (Radford et al., 2019), Bidirectional Encoder Representations from Transformers (BERT) (Devlin et al., 2018) etc.. Hence, neural language models have become the new holy grail of modern NLP, although LM for CM is an unexplored area altogether. To better understand the problem of LM for CM, we initially experimented with several statistical language modeling techniques and consequently experimented with contemporary neural language models. Analysis shows switching-points are the main challenge for the LMCM performance drop, therefore in this paper we introduce the idea of minority positive sampling to selectively induce more sample to achieve better performance. On the contrary, all neural language models demand a huge corpus to train on for better performance. Finally, we are reporting a perplexity of 139 for Hinglish (Hindi-English language pair) LMCM using statistical bi-directional techniques.
Arindam Chatterjere, Vineeth Guptha, Parul Chopra, Amitava Das 0001
LREC4
2020 Meme vs. Non-meme Classification using Visuo-linguistic Association
Chhavi Sharma, Viswanath Pulabaigari, Amitava Das 0001
WEBIST3
2018 Consonant-Vowel Sequences as Subword Units for Code-Mixed Languages
abstract
In this research work, we develop a state-of-art model for identifying sentiment in Hindi-English code-mixed language. We introduce new phonemic sub-word units for Hindi-English code-mixed text along with a hierarchical deep learning model which uses these sub-word units for predicting sentiment. The results indicate that the model yields a significant increase in accuracy as compared to other models.
Chris Andrew Gadde, Santhoshini Reddy, Amitava Das 0001
AAAI5
2017 Semantic Interpretation of Social Network Communities
abstract
A community in a social network is considered to be a group of nodes densely connected internally and sparsely connected externally.Although previous work intensely studied network topology within a community, its semantic interpretation is hardly understood. In this paper, we attempt to understand whether individuals in a community possess similar Personalities, Values and Ethical background. Finally, we show that Personality and Values models could be used as features to discover more accurate community structure compared to the one obtained from only network information.
Tushar Maheshwari, Aishwarya N. Reganti, Tanmoy Chakraborty 0002, Amitava Das 0001
AAAI5
2017 Understanding Psycho-Sociological Vulnerability of ISIS Patronizers in Twitter
abstract
The Islamic State of Iraq and Syria (ISIS) is a Salafi jihadist militant group that has made extensive use of online social media platforms to promulgate its ideologies and evoke many individuals to support the organization. The psycho-sociological background of an individual plays a crucial role in determining his/her vulnerability of being lured into joining the organisation and indulge in terrorist activities, since his/her behavior largely depends on the society s/he was brought up in. Here, we analyse five sociological aspects -- personality, values & ethics, optimism/pessimism, age and gender to understand the psycho-sociological vulnerability of individuals over Twitter. Experimental results suggest that psycho-sociological aspects indeed act as foundation to discover and differentiate between prominent and unobtrusive users in Twitter.
Aishwarya N. Reganti, Tushar Maheshwari, Amitava Das 0001, Tanmoy Chakraborty 0002, Ponnurangam Kumaraguru
ASONAM3
2017 A Societal Sentiment Analysis: Predicting the Values and Ethics of Individuals by Analysing Social Media Content
abstract
Tushar Maheshwari, Aishwarya N. Reganti, Samiksha Gupta, Anupam Jamatia, Upendra Kumar, Björn Gambäck, Amitava Das. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Tushar Maheshwari, Aishwarya N. Reganti, Samiksha Gupta, Anupam Jamatia, Björn Gambäck, Amitava Das 0001
EACL (1)7
2017 FairScholar: Balancing Relevance and Diversity for Scientific Paper Recommendation
Ankesh Anand, Tanmoy Chakraborty 0002, Amitava Das 0001
ECIR3
2017 Quotology - Reading Between the Lines of Quotations
Dwijen Rudra Pal, Amitava Das 0001, Baby Bhattacharya
NLDB2
2016 Collecting and Annotating Indian Social Media Code-Mixed Corpora
Anupam Jamatia, Björn Gambäck, Amitava Das 0001
CICLing (2)3
2016 Comparing the Level of Code-Switching in Corpora
Björn Gambäck, Amitava Das 0001
LREC2
2014 Poetic Machine: Computational Creativity for Automatic Poetry Generation in Bengali
Amitava Das 0001, Björn Gambäck
ICCC1
2014 A Framework for Health Behavior Change using Companionable Robots
abstract
In this paper, we describe a dialogue system framework for a companionable robot, which aims to guide patients to-wards health behavior changes via natu-ral language analysis and generation. The framework involves three broad stages, rapport building and health topic identifi-cation, assess patient’s opinion of change, and designing plan and closing session. The framework uses concepts from psy-chology, computational linguistics, and machine learning and builds on them. One of the goals of the framework is to ensure that the Companionbot builds and main-tains rapport with patients. 1
Bandita Sarma, Amitava Das 0001, Rodney D. Nielsen
INLG2
2012 The 5W Structure for Sentiment Summarization-Visualization-Tracking
Amitava Das 0001, Sivaji Bandyopadhyay, Björn Gambäck
CICLing (1)1
2010 JU_CSE_GREC10: Named Entity Generation at GREC 2010
Amitava Das 0001, Tanik Saikh, Tapabrata Mondal, Sivaji Bandyopadhyay
INLG1
2010 Towards the Global SentiWordNet
Amitava Das 0001, Sivaji Bandyopadhyay
PACLIC1
2008 Language Independent Named Entity Recognition in Indian Languages
Asif Ekbal, Rejwanul Haque, Amitava Das 0001, Venkateswarlu Poka, Sivaji Bandyopadhyay
IJCNLP3