Sunandan Chakraborty

dblp:61/9376 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-3331-6082ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Mod-Guide: An LLM-based Content Moderation Feedback System to Address Insensitive Speech toward Indigenous Ethnic and Religious Minority Communities
abstract
Language operates as a mechanism of both marginalization and resistance, especially for minority communities navigating insensitive and harmful speech online. As content moderation increasingly depends on large language models (LLMs), concerns arise about whether these systems can recognize culturally insensitive speech–language that disregards or marginalizes the cultural and religious perspectives of historically underrepresented communities, often through implicit erasure, misrepresentation, or normative framing, rather than overt hostility. Focusing on Bangladesh’s Hindu and Chakma communities – the country’s largest religious and Indigenous ethnic minorities, respectively – this paper investigates the epistemic limits of LLM-based moderation systems and explores methods for incorporating minority perspectives. We co-created a culturally grounded corpus of insensitive speech with community members and integrated their narratives into moderation pipelines using retrieval augmented generation (RAG). Our tool, Mod-Guide, improves LLM sensitivity to minority viewpoints by leveraging contextual cues derived from lived experience. Through mixed-method evaluations involving both minority and majority participants, we demonstrate that RAG-enhanced moderation responses are more contextually accurate and perceived differently across ethnic lines. This work advances research in human-computer interaction, AI ethics, and social computing by foregrounding restorative justice and hermeneutical inclusion in the design of content moderation systems.
Dipto Das, Achhiya Sultana, Ankit-Singh Chauhan, Saadia Binte Alam, Mohammad Shidujaman, Shion Guha, Sunandan Chakraborty, Syed Ishtiaque Ahmed
COMPASS7
2025 Domain-Specific Knowledge Graphs in RAG-Enhanced Healthcare LLMS
Sydney Anuyah, Mehedi Mahmud Kaushik, Rakesh Shiradkar, Arjan Durresi, Sunandan Chakraborty
IEEE Big Data6
2025 Benchmarking LLMs for Pairwise Causal Discovery in Biomedical and Multi-Domain Contexts
abstract
The safe deployment of large language models (LLMs) in high-stakes fields like biomedicine, requires them to be able to reason about cause and effect. We investigate this ability by testing 13 open-source LLMs on a fundamental task: pairwise causal discovery (PCD) from text. Our benchmark, using 12 diverse datasets, evaluates two core skills: 1) \textbf{Causal Detection} (identifying if a text contains a causal link) and 2) \textbf{Causal Extraction} (pulling out the exact cause and effect phrases). We tested various prompting methods, from simple instructions (zero-shot) to more complex strategies like Chain-of-Thought (CoT) and Few-shot In-Context Learning (FICL). The results show major deficiencies in current models. The best model for detection, DeepSeek-R1-Distill-Llama-70B, only achieved a mean score of 49.57\% ($C_{detect}$), while the best for extraction, Qwen2.5-Coder-32B-Instruct, reached just 47.12\% ($C_{extract}$). Models performed best on simple, explicit, single-sentence relations. However, performance plummeted for more difficult (and realistic) cases, such as implicit relationships, links spanning multiple sentences, and texts containing multiple causal pairs. We provide a unified evaluation framework, built on a dataset validated with high inter-annotator agreement ($κ\ge 0.758$), and make all our data, code, and prompts publicly available to spur further research. \href{https://github.com/sydneyanuyah/CausalDiscovery}{Code available here: https://github.com/sydneyanuyah/CausalDiscovery}
Sydney Anuyah, Sneha Shajee-Mohan, Ankit-Singh Chauhan, Sunandan Chakraborty
IEEE Big Data4
2025 Descriptive Analysis of Online Wildlife Products Using Vision Language Models
Kinshuk Sharma, Juliana Silva Barbosa, Spencer Roberts, Ulhas Gondhali, Gohar Petrossian, Jennifer Jacquet, Juliana Freire, Sunandan Chakraborty
COMPASS8
2025 Automated Knowledge Graph Construction using Large Language Models and Sentence Complexity Modelling
abstract
Sydney Anuyah, Mehedi Mahmud Kaushik, Sri Rama Krishna Reddy Dwarampudi, Rakesh Shiradkar, Arjan Durresi, Sunandan Chakraborty. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Sydney Anuyah, Mehedi Mahmud Kaushik, Sri Rama Krishna Reddy Dwarampudi, Rakesh Shiradkar, Arjan Durresi, Sunandan Chakraborty
EMNLP6
2025 A Cost-Effective LLM-based Approach to Identify Wildlife Trafficking in Online Marketplaces
abstract
Wildlife trafficking remains a critical global issue, significantly impacting biodiversity, ecological stability, and public health. Despite efforts to combat this illicit trade, the rise of e-commerce platforms has made it easier to sell wildlife products, putting new pressure on wild populations of endangered and threatened species. The use of these platforms also opens a new opportunity: as criminals sell wildlife products online, they leave digital traces of their activity that can provide insights into trafficking activities as well as how they can be disrupted. The challenge lies in finding these traces. Online marketplaces publish ads for a plethora of products, and identifying ads for wildlife-related products is like finding a needle in a haystack. Learning classifiers can automate ad identification, but creating them requires costly, time-consuming data labeling that hinders support for diverse ads and research questions. This paper addresses a critical challenge in the data science pipeline for wildlife trafficking analytics: generating quality labeled data for classifiers that select relevant data. While large language models (LLMs) can directly label advertisements, doing so at scale is prohibitively expensive. We propose a cost-effective strategy that leverages LLMs to generate pseudo labels for a small sample of the data and uses these labels to create specialized classification models. Our novel method automatically gathers diverse and representative samples to be labeled while minimizing the labeling costs. Our experimental evaluation shows that our classifiers achieve up to 95% F1 score, outperforming LLMs at a lower cost. We present real use cases that demonstrate the effectiveness of our approach in enabling analyses of different aspects of wildlife trafficking.
Juliana Silva Barbosa, Ulhas Gondhali, Gohar Petrossian, Kinshuk Sharma, Sunandan Chakraborty, Jennifer Jacquet, Juliana Freire
Proc. ACM Manag. Data5
2024 AAVE Corpus Generation and Low-Resource Dialect Machine Translation
abstract
African American Vernacular English (AAVE) is a dialect of the English language spoken in the United States by members of the Black community. The stark differences between AAVE and Standard American English (SAE), as well as a historically negative stigma towards its use, have contributed to an academic performance gap between Black students and their non-Black counterparts. This research works to generate educational resources similar to what is available in English Second Language (ESL) classrooms. Exposure to these resources has been shown to both improve the negative stigma towards the use of AAVE as well as facilitate code-switching between AAVE and SAE. The resources to be generated in this research are a parallel corpora for AAVE and SAE using both professionally translated text and AI-generated text, and a Neural Machine Translation (NMT) model to translate SAE into AAVE using novel network architectures used language to language translation including LSTM, Bi-LSTM, Attention, and Transformer network components. The parallel corpora will be quantitatively reviewed and validated before using tested dialect translation model methods. Methodology will additionally be focused on low-resource machine translation due to the lack of large corpora containing AAVE. Both professional translators and large language model, ChatGPT, will be used to create parallel corpora containing AAVE and SAE. This short paper details the preliminary results of the assessment of these generated corpora as well as the accuracy of dialect machine translation models trained on them.
Eric Graves 0004, Shreyas Aswar, Rujuta Desai, Srilekha Nampelli, Sunandan Chakraborty, Ted Hall
COMPASS5
2022 Mining Latent Disease Factors from Medical Literature using Causality
abstract
Understanding causality is a longstanding goal across many different domains. Different articles, such as those published in medical journals, publish newly discovered knowledge, often causal. In this paper, we use this intuition to build a model that leverages causal relations to unearth factors related to Sjögren’s syndrome. Sjögren’s syndrome is an autoimmune disease affecting up to 3.1 million Americans. The uncommon nature of the disease, coupled with common symptoms of other autoimmune conditions such as rheumatoid arthritis, it is difficult for clinicians to timely diagnose the disease. This is further worsened by suboptimal communication between dentists, and physicians, including rheumatologists and ophthalmologists, because clinical manifestations of this disease require the patients to visit physicians with different specialties. A centralized information system with easy access to common and uncommon factors related to Sjögren’s syndrome may alleviate the problem. We use automatically extracted causal relationships from text related to Sjögren’s syndrome collected from the medical literature to identify a set of factors, such as “signs and symptoms” and “associated conditions”, related to this disease. We show that our approach is capable of retrieving such factors with high precision and recall values. Comparative experiments show that this approach leads to 25% improvement in retrieval F1-score compared to several state-of-the-art biomedical models, including BioBERT and Gram-CNN.
Pranav Gujarathi, Jack VanSchaik, Venkata Mani Babu Karri, Anushri Singh Rajapuri, Biju Cheriyan, Thankam Thyvalikakath, Sunandan Chakraborty
IEEE Big Data7
2022 Detecting Hotspots of Human-Wildlife Conflicts in India using News Articles and Aerial Images
abstract
Human-wildlife conflict (HWC) is one of the most pressing conservation issues at present, with incidents leading to human injury and death, crop and property damage, and livestock predation. Since acquiring real-time data and performing manual analysis on those incidents are costly, we propose to leverage machine learning techniques to build an automated pipeline to construct an HWC knowledge base from historical news articles. Our unsupervised and active learning methods are not only able to recognize the major causes of HWC such as construction, pollution, and farming, but can also classify an unseen news article into its major cause with 90% accuracy. Moreover, our interactive visualizations of the knowledge base illustrate the spatial and temporal trend of human-wildlife conflicts across India for index by cities and animals. Based on our findings that most conflict zones include areas where human settlements are near forested areas, we extend our study to include satellite imagery to identify such proximity zones. We conduct a case study to use this method to identify human-elephant conflict hotspots in northern and western parts of the Indian state of West Bengal. We expect that our findings can inform the public of HWC hotspots and help in much more informed policymaking.
Gokhan Egri, Xinran Han, Zilin Ma, Priyanka Surapaneni, Sunandan Chakraborty
COMPASS5
2022 Note: Using Causality to Mine Sjögren's Syndrome related Factors from Medical Literature
abstract
Research articles published in medical journals often present findings from causal experiments. In this paper, we use this intuition to build a model that leverages causal relations expressed in text to unearth factors related to Sjögren’s syndrome. Sjögren’s syndrome is an auto-immune disease affecting up to 3.1 million Americans. The uncommon nature of the disease, coupled with common symptoms with other autoimmune conditions make the timely diagnosis of this disease very hard. A centralized information system with easy access to common and uncommon factors related to Sjögren’s syndrome may alleviate the problem. We use automatically extracted causal relationships from text related to Sjögren’s syndrome collected from the medical literature to identify a set of factors, such as “signs and symptoms” and “associated conditions”, related to this disease. We show that our approach is capable of retrieving such factors with a high precision and recall values. Comparative experiments show that this approach leads to 25% improvement in retrieval F1-score compared to several state-of-the-art biomedical models, including BioBERT and Gram-CNN.
Pranav Dhananjay Gujarathi, Sai Krishna Reddy Gopi Reddy, Venkata Mani Babu Karri, Ananth Reddy Bhimireddy, Anushri Singh Rajapuri, Manohar Reddy, Mounika Sabbani, Biju Cheriyan, Jack VanSchaik, Thankam Thyvalikakath, Sunandan Chakraborty
COMPASS11
2021 Quantifying Uncertainty in Patient Count Metrics Derived from Imperfect EHR-based Phenotypes
Jack VanSchaik, Sunandan Chakraborty
AMIA2
2020 Extracting Features from Online Forums to Meet Social Needs of Breast Cancer Patients
abstract
Breast cancer patients go through many ordeals when they undergo treatments. Many of these issues are personal, social, or professional. As many of them are not directly medical in nature, these issues are not discussed with their healthcare providers and hence, not included in their treatment plan. However, these issues are vital for the patients' complete recovery. We present a novel approach that acts as the first step in including such personal and social issues resulting from breast cancer treatment into a patient's treatment plan. There are numerous online forums where patients share their experiences and post questions about their treatments and subsequent side effects. We collected data from one such forum called "Online Breast Cancer Forum". On this forum, users (patients) have created threads across many related topics and shared their experiences and questions. We use these message threads to identify critical issues faced by the patient and how they are related to their treatment. We convert the forum data into a bipartite network and turn the network nodes into a high-dimensional feature space. In this feature space, we perform community detection to unearth latent connections between patients and topics. We claim that these latent connections, along with the known ones, will help to create a new knowledge base that will eventually help physicians to estimate non-medical issues for a prescribed treatment. This new knowledge will help the physicians plan a more adaptive and personalized treatment and be better prepared by anticipating potential problems beforehand. We evaluated our method on two baseline methods and show that our method outperforms the baseline methods by 25% on a manually labeled reference dataset.
Maitreyi Mokashi, Enming Zhang, Josette F. Jones, Sunandan Chakraborty
COMPASS4
2020 Affording Extremes: Incivility, Social Media and Democracy in the Indian Context
abstract
In this mixed-methods study of political discourse, we study the affordances of Twitter in the context of free speech in India. We critically examine specific cases of the legal prosecution of free speech and the use of extreme speech in attacks on people to document the risks to citizens when they engage in antagonistic online discourse, particularly against the state or political institutions. We follow this up with quantitative study of the use of extreme speech through 477 hashtags used by a random sample of thousand political actors on Twitter and find that politicians are rewarded, through higher retweet rates, when they engage in extreme or uncivil messaging. We contextualize these findings to the postcolonial history of India and the laws and institutions that enable differential consequences for engaging in various forms of speech. In conclusion, we propose that the affordances of new ICTs like social media need to be carefully considered for their unintended consequences, and that functional access to free speech may differ dramatically based on one's access to institutions.
Anmol Panda, Sunandan Chakraborty, Noopur Raval, Han Zhang 0037, Mugdha Mohapatra, Syeda Zainab Akbar, Joyojeet Pal
ICTD2
2019 Identifying Predictive Causal Factors from News Streams
abstract
Ananth Balashankar, Sunandan Chakraborty, Samuel Fraiberger, Lakshminarayanan Subramanian. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ananth Balashankar, Sunandan Chakraborty, Samuel P. Fraiberger, Lakshminarayanan Subramanian
EMNLP/IJCNLP (1)2
2019 A Co-Training Model with Label Propagation on a Bipartite Graph to Identify Online Users with Disabilities
Sunandan Chakraborty, Erin Brady
ICWSM2
2018 Political Tweets and Mainstream News Impact in India: A Mixed Methods Investigation into Political Outreach
abstract
Citizens' perception of politicians and political issues is increasingly influenced by social media. However, little is known about the potential of second order effects of social media in parts of the world where the majority of voting citizens are not online. In this paper, we examine whether a politician can move to communicating through social media as their primary means of outreach, and still present their message to the mainstream population through traditional media. By studying of the use of Twitter by Indian Prime Minister Narendra Modi between 2009 and 2015, the second-most followed elected official in the world, we present evidence of the impact of social media on print news. We use computational text mining techniques to automatically identify print news reports that use Modi's tweets as a source, alongside manual qualitative coding of tweets to analyze the role of tweet themes in print news coverage. We conclude that while Modi's social media messaging does get coverage in the print news, it is his more "newsworthy" tweets, such as references to celebrities, other politicians, or major events such as holidays that have a greater likelihood of coverage.
Sunandan Chakraborty, Joyojeet Pal, Priyank Chandra, Daniel M. Romero
COMPASS1
2016 The Effects of the Content of FOMC Communications on US Treasury Rates
abstract
This study measures the effects of Federal Open Market Committee text content on the direction of short-and medium-term interest rate movements.Because the words relevant to short-and medium-term interest rates differ, we apply a supervised approach to learn distinct sets of topics for each dependent variable being examined.We generate predictions with and without controlling for factors relevant to interest rate movements, and our prediction results average across multiple training-test splits.Using data from 1999-2016, we achieve 93% and 64% accuracy in predicting Target and Effective Federal Funds Rate movements and 38%-40% accuracy in predicting longer term Treasury Rate movements.We obtain lower but comparable accuracies after controlling for other macroeconomic and market factors.
Chris Rohlfs, Sunandan Chakraborty, Lakshminarayanan Subramanian
EMNLP2
2016 Predicting Socio-Economic Indicators using News Events
abstract
Many socio-economic indicators are sensitive to real-world events. Proper characterization of the events can help to identify the relevant events that drive fluctuations in these indicators. In this paper, we propose a novel generative model of real-world events and employ it to extract events from a large corpus of news articles. We introduce the notion of an event class, which is an abstract grouping of similarly themed events. These event classes are manifested in news articles in the form of event triggers which are specific words that describe the actions or incidents reported in any article. We use the extracted events to predict fluctuations in different socio-economic indicators. Specifically, we focus on food prices and predict the price of 12 different crops based on real-world events that potentially influence food price volatility, such as transport strikes, festivals etc. Our experiments demonstrate that incorporating event information in the prediction tasks reduces the root mean square error (RMSE) of prediction by 22% compared to the standard ARIMA model. We also predict sudden increases in the food prices (i.e. spikes) using events as features, and achieve an average 5-10% increase in accuracy compared to baseline models, including an LDA topic-model based predictive model.
Sunandan Chakraborty, Ashwin Venkataraman, Srikanth Jagabathula, Lakshminarayanan Subramanian
KDD1
2014 On correlation of absence time and search effectiveness
abstract
Online search evaluation metrics are typically derived based on implicit feedback from the users. For instance, computing the number of page clicks, number of queries, or dwell time on a search result. In a recent paper, Dupret and Lalmas introduced a new metric called absence time, which uses the time interval between successive sessions of users to measure their satisfaction with the system. They evaluated this metric on a version of Yahoo! Answers. In this paper, we investigate the effectiveness of absence time in evaluating new features in a web search engine, such as new ranking algorithm or a new user interface. We measured the variation of absence time to the effects of 21 experiments performed on a search engine. Our findings show that the outcomes of absence time agreed with the judgement of human experts performing a thorough analysis of a wide range of online and offline metrics in 14 out of these 21 cases.
Sunandan Chakraborty, Filip Radlinski, Milad Shokouhi, Paul Baecke
SIGIR1
2012 Empowering authors to diagnose comprehension burden in textbooks
abstract
Good textbooks are organized in a systematically progressive fashion so that students acquire new knowledge and learn new concepts based on known items of information. We provide a diagnostic tool for quantitatively assessing the comprehension burden that a textbook imposes on the reader due to non-sequential presentation of concepts. We present a formal definition of comprehension burden and propose an algorithmic approach for computing it. We apply the tool to a corpus of high school textbooks from India and empirically examine its effectiveness in helping authors identify sections of textbooks that can benefit from reorganizing the material presented.
Rakesh Agrawal 0001, Sunandan Chakraborty, Sreenivas Gollapudi, Anitha Kannan, Krishnaram Kenthapadi
KDD2
2010 Managing microfinance with paper, pen and digital slate
abstract
India's extensive Self-Help Group (SHG) microfinance network brings formal savings and credit services to 86 million poor households. Yet, the inability to maintain high-quality records remains a persistent weakness in SHG functioning. We study this problem and present a financial record management application built on a low-cost digital slate prototype. The solution directly accepts handwritten input on ordinary paper forms and provides immediate electronic feedback. A field trial with 200 SHG members in rural India shows that the use of the digital slate solution results in shorter data recording time, fewer incorrect entries, and more complete records. The paper-pen-slate solution performs as well as, and is strongly preferred over, a purely electronic alternative. The digital slate solution is able to comfortably move between paper and digital worlds, achieving efficiency and quality gains while catering to the preferences and budgets of low-income low-literate clients.
Aishwarya Ratan, Kentaro Toyama, Sunandan Chakraborty, Keng Siang Ooi, Mike Koenig, Pushkar V. Chitnis, Matthew Phiong
ICTD3
2007 Samvidha: A ICT system for personalized offline Internet Access for rural schools
abstract
Internet is a huge repository of quality learning materials and continues to grow in a faster rate. The school students may be benefited immensely as these learning materials may well supplement their curricular requirements. But Access to the Internet is costly, because it is very expensive to maintain a persistent Internet connection. For some schools in the developing countries like India, this cost may not be affordable specifically in rural schools. This makes way to a digital divide between the rural and urban schools which is unwanted. For these rural schools, limiting the amount of bandwidth consumed is of paramount importance. It is necessary that the schools be connected to the Internet for the least time, in order to minimize the access cost. In this paper, we present a system Samvidha that allows the rural school students to access the Internet contents in an offline fashion.
Plaban Kumar Bhowmick, Sudeshna Sarkar, Sunandan Chakraborty, Anupam Basu
ICTD3
2007 Shikshak: An Intelligent Tutoring System Authoring tool for rural education
abstract
Low literacy scenario in India and other developing nation demands an alternative learning environment to deal with the problem. Lack of trained teachers, high dropout rates are some of the major problems that need to be addressed. Intelligent Tutoring System (ITS) or ITS Authoring tools (ITSAT) can be thought of as a possible solution to these problems. In this paper we present Shikshak, an ITSAT developed by us and discuss its deployment in the district of Paschim Medinipur, West Bengal along with its sample effect on primary education.
Sunandan Chakraborty, Tamali Bhattacharya, Plaban Kumar Bhowmick, Anupam Basu, Sudeshna Sarkar
ICTD1