Mohna Chakraborty

dblp:299/8728 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0003-3112-7445ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 How Reasoning Influences Intersectional Biases in Vision-Language Models (Student Abstract)
abstract
Vision-Language Models (VLMs) are increasingly deployed across downstream tasks, yet their training data often encode social biases that surface in outputs. Unlike humans, who interpret images through contextual and social cues, VLMs process them through statistical associations, often leading to reasoning that diverges from human reasoning. By analyzing how a VLM reasons, we can understand how inherent biases are perpetuated and can adversely affect downstream performance. To examine this gap, we systematically analyze social biases in five open-source VLMs for an occupation prediction task, on the FairFace dataset. Across 32 occupations and three different prompting styles, we elicit both predictions and reasoning. Our findings show that the biased reasoning patterns systematically underlie intersectional disparities, highlighting the need to align VLM reasoning with human values before downstream deployment.
Adit Desai, Sudipta Roy 0002, Mohna Chakraborty
AAAI3
2025 Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation Framework
abstract
Large language models (LLMs) are increasingly deployed in domains requiring moral understanding, yet their reasoning often remains shallow, and misaligned with human reasoning (Jiang et al., 2021).Unlike humans, whose moral reasoning integrates contextual trade-offs, value systems, and ethical theories, LLMs often rely on surface patterns, leading to biased decisions in morally and ethically complex scenarios.To address this gap, we present a value-grounded framework for evaluating and distilling structured moral reasoning in LLMs.We benchmark 12 open-source models across four moral datasets using a taxonomy of prompts grounded in value systems, ethical theories, and cognitive reasoning strategies.Our evaluation is guided by four questions: (1) Does reasoning improve LLM decision-making over direct prompting?(2) Which types of value/ethical frameworks most effectively guide LLM reasoning?(3) Which cognitive reasoning strategies lead to better moral performance?(4) Can small-sized LLMs acquire moral competence through distillation?We find that prompting with explicit moral structure consistently improves accuracy and coherence, with first-principles reasoning and Schwartz's + care-ethics scaffolds yielding the strongest gains.Furthermore, our supervised distillation approach transfers moral competence from large to small models without additional inference cost.Together, our results offer a scalable path toward interpretable and value-grounded models.
Mohna Chakraborty, Lu Wang 0008, David Jurgens
EMNLP1
2025 Modeling Data Diversity for Joint Instance and Verbalizer Selection in Cold-Start Scenarios
Mohna Chakraborty, Adithya Kulkarni, Qi Li 0012
PAKDD (1)1
2025 Blue Sky: Reducing Performance Gap between Commercial and Open-Source LLMs
abstract
The performance gap between commercial and open-source large language models (LLMs) poses a critical challenge in achieving equitable access to advanced AI technologies, particularly for underfunded institutions. As commercial entities like OpenAI invest substantial resources into proprietary models, open-source alternatives struggle with limitations such as a lack of access to high-quality datasets and feedback, restricting opportunities for research and innovation. We propose strategies needed to democratize AI technology, emphasizing collaboration and knowledge sharing within the community. By fostering a more inclusive environment, we can develop versatile, user-focused models that empower diverse stakeholders and expand the horizons of AI research across various sectors. This paper calls for a holistic approach to bridging this gap through behavioral modeling, leveraging techniques such as reinforcement learning and scenario-based testing to enhance the capabilities of open-source LLMs.
Adithya Kulkarni, Mohna Chakraborty
SDM2
2025 Budget Allocation Exploiting Label Correlation between Instances
abstract
In this study, we introduce an innovative budget allocation method for graph instance annotation in crowdsourcing environments, where both the labels of instances and their correlations are unknown and need to be estimated simultaneously. We model the budget allocation task as a Markov Decision Process (MDP) and develop an optimization framework that minimizes the uncertainties associated with instance labeling and correlation estimation while adhering to budget constraints. To quantify uncertainty, we employ entropy and derive two strategies: OPTUENT-EXP and OPTUENT-OPT. Our reward function further considers the impact of a worker’s label on the entire graph. We conducted extensive experiments using four real-world graph datasets, simulating worker labeling behavior to showcase the effectiveness of our approach. Experimental results demonstrate that our proposed approach can accurately estimate correlations between adjacent nodes while significantly reducing labeling costs. Moreover, across four real-world datasets, our proposed approach consistently outperforms existing baselines in moderate and high budget scenarios, highlighting its robustness and practical scalability.
Adithya Kulkarni, Mohna Chakraborty, Sihong Xie, Qi Li 0012
UAI2
2025 Weakly Supervised Open-Domain Aspect-Based Sentiment Analysis
abstract
Aspect-Based Sentiment Analysis (ABSA) comprises several subtasks: aspect term extraction (ATE), opinion term extraction (OTE), aspect term sentiment extraction (ATSE), aspect-opinion pair extraction (AOPE), and aspect sentiment triplet extraction (ASTE). Existing unified frameworks for ABSA rely heavily on large-scale annotated data, limiting scalability across domains. We propose UAOS, a double-layer unified span extraction framework that performs all five ABSA subtasks under weak supervision. Our approach first extracts aspect-opinion pairs using universal dependency-based rules from unannotated corpora. Sentiment labels for these pairs are generated via a novel zero-shot, domain-agnostic prompt-based method. The resulting weak labels train a unified span extraction architecture equipped with canonical correlation analysis for early stopping and a self-training mechanism to mitigate noise and bias in supervision. Extensive experiments on four ABSA benchmarks demonstrate that UAOS achieves competitive or superior performance compared to fully supervised baselines. It improves upon the state-of-the-art ODAO by +1.54 F1 for ATE, +0.56 for OTE, and +0.82 for AOPE. In ATSE and ASTE, where no weakly supervised baselines exist, UAOS outperforms several supervised models, setting new benchmarks. To assess domain generalizability, we evaluate UAOS on a psychology/education-domain dataset of student reflections spanning four instructional conditions. Without in-domain fine-tuning, it achieves macro F1 scores of 71.05 (ATE), 74.39 (OTE), 68.24 (AOPE), and 60.56 (ASTE). These results highlight the model’s ability to generalize to out-of-distribution, non-commercial text, underscoring its scalability for low-resource ABSA applications.
Mohna Chakraborty, Adithya Kulkarni, Qi Li 0012
ACM Trans. Knowl. Discov. Data1
2023 Zero-shot Approach to Overcome Perturbation Sensitivity of Prompts
abstract
Recent studies have demonstrated that naturallanguage prompts can help to leverage the knowledge learned by pre-trained language models for the binary sentence-level sentiment classification task.Specifically, these methods utilize few-shot learning settings to finetune the sentiment classification model using manual or automatically generated prompts.However, the performance of these methods is sensitive to the perturbations of the utilized prompts.Furthermore, these methods depend on a few labeled instances for automatic prompt generation and prompt ranking.This study aims to find high-quality prompts for the given task in a zero-shot setting.Given a base prompt, our proposed approach automatically generates multiple prompts similar to the base prompt employing positional, reasoning, and paraphrasing techniques and then ranks the prompts using a novel metric.We empirically demonstrate that the top-ranked prompts are high-quality and significantly outperform the base prompt and the prompts generated using few-shot learning for the binary sentence-level sentiment classification task.
Mohna Chakraborty, Adithya Kulkarni, Qi Li 0012
ACL (1)1
2023 Optimal Budget Allocation for Crowdsourcing Labels for Graphs
abstract
Crowdsourcing is an effective and efficient paradigm for obtaining labels for unlabeled corpus employing crowd workers. This work considers the budget allocation problem for a generalized setting on a graph of instances to be labeled where edges encode instance dependencies. Specifically, given a graph and a labeling budget, we propose an optimal policy to allocate the budget among the instances to maximize the overall labeling accuracy. We formulate the problem as a Bayesian Markov Decision Process (MDP), where we define our task as an optimization problem that maximizes the overall label accuracy under budget constraints. Then, we propose a novel stage-wise reward function that considers the effect of worker labels on the whole graph at each timestamp. This reward function is utilized to find an optimal policy for the optimization problem. Theoretically, we show that our proposed policies are consistent when the budget is infinite. We conduct extensive experiments on five real-world graph datasets and demonstrate the effectiveness of the proposed policies to achieve a higher label accuracy under budget constraints.
Adithya Kulkarni, Mohna Chakraborty, Sihong Xie, Qi Li 0012
UAI2
2022 Open-Domain Aspect-Opinion Co-Mining with Double-Layer Span Extraction
abstract
The aspect-opinion extraction tasks extract aspect terms and opinion terms from reviews. The supervised extraction methods achieve state-of-the-art performance but require large-scale human-annotated training data. Thus, they are restricted for open-domain tasks due to the lack of training data. This work addresses this challenge and simultaneously mines aspect terms, opinion terms, and their correspondence in a joint model. We propose an Open-Domain Aspect-Opinion Co-Mining (ODAO) method with a Double-Layer span extraction framework. Instead of acquiring human annotations, ODAO first generates weak labels for unannotated corpus by employing rules-based on universal dependency parsing. Then, ODAO utilizes this weak supervision to train a double-layer span extraction framework to extract aspect terms (ATE), opinion terms (OTE), and aspect-opinion pairs (AOPE). ODAO applies canonical correlation analysis as an early stopping indicator to avoid the model over-fitting to the noise to tackle the noisy weak supervision. ODAO applies a self-training process to gradually enrich the training data to tackle the weak supervision bias issue. We conduct extensive experiments and demonstrate the power of the proposed ODAO. The results on four benchmark datasets for aspect-opinion co-extraction and pair extraction tasks show that ODAO can achieve competitive or even better performance compared with the state-of-the-art fully supervised methods.
Mohna Chakraborty, Adithya Kulkarni, Qi Li 0012
KDD1
2021 Does reusing pre-trained NLP model propagate bugs?
abstract
In this digital era, the textual content has become a seemingly ubiquitous part of our life. Natural Language Processing (NLP) empowers machines to comprehend the intricacies of textual data and eases human-computer interaction. Advancement in language modeling, continual learning, availability of a large amount of linguistic data, and large-scale computational power have made it feasible to train models for downstream tasks related to text analysis, including safety-critical ones, e.g., medical, airlines, etc. Compared to other deep learning (DL) models, NLP-based models are widely reused for various tasks. However, the reuse of pre-trained models in a new setting is still a complex task due to the limitations of the training dataset, model structure, specification, usage, etc. With this motivation, we study BERT, a vastly used language model (LM), from the direction of reusing in the code. We mined 80 posts from Stack Overflow related to BERT and found 4 types of bugs observed in clients’ code. Our results show that 13.75% are fairness, 28.75% are parameter, 15% are token, and 16.25% are version-related bugs.
Mohna Chakraborty
ESEC/SIGSOFT FSE1