EDBT 2026 Demo / reviewers in the wild / expert
Rima Hazra
dblp:247/1245
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-0535-8322ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language ModelsabstractPresent day LLMs face the challenge of managing affordance-based safety risks—situations where outputs inadvertently facilitate harmful actions due to overlooked logical implications. Traditional safety solutions, such as scalar outcome-based reward models, parameter tuning, or heuristic decoding strategies, lack the granularity and proactive nature needed to reliably detect and intervene during subtle yet crucial reasoning steps. Addressing this fundamental gap, we introduce AURA, an innovative, multi-layered framework centered around Process Reward Models (PRMs), providing comprehensive, step level evaluations across logical coherence and safety-awareness. Our framework seamlessly combines introspective self-critique, fine-grained PRM assessments, and adaptive safety-aware decoding to dynamically and proactively guide models toward safer reasoning trajectories. Empirical evidence clearly demonstrates that this approach significantly surpasses existing methods, significantly improving the logical integrity and affordance-sensitive safety of model outputs. This research represents a pivotal step toward safer, more responsible, and contextually aware AI, setting a new benchmark for alignment-sensitive applications. Sayantan Adak, Pratyush Chatterjee, Somnath Banerjee 0002, Rima Hazra, Somak Aditya, Animesh Mukherjee 0001 |
AAAI | 4 |
| 2025 | SafeInfer: Context Adaptive Decoding Time Safety Alignment for Large Language ModelsabstractLanguage models aligned for safety often exhibit fragile and imbalanced mechanisms, increasing the chances of producing unsafe content. In addition, editing techniques to incorporate new knowledge can further compromise safety. To tackle these issues, we propose SafeInfer, a context-adaptive, decoding-time safety alignment strategy for generating safe responses to user queries. safeInfer involves two phases: the 'safety amplification' phase, which uses safe demonstration examples to adjust the model’s hidden states and increase the likelihood of safer outputs, and the 'safety-guided decoding' phase, which influences token selection based on safety-optimized distributions to ensure the generated content adheres to ethical guidelines. Further, we introduce HarmEval, a novel benchmark for comprehensive safety evaluations, designed to address potential misuse scenarios in line with the policies of leading AI technology companies. Somnath Banerjee 0002, Sayan Layek, Soham Tripathy, Shanu Kumar, Animesh Mukherjee 0001, Rima Hazra |
AAAI | 6 |
| 2025 | Turning Logic Against Itself: Probing Model Defenses Through Contrastive QuestionsabstractLarge language models, despite extensive alignment with human values and ethical principles, remain vulnerable to sophisticated jailbreak attacks that exploit their reasoning abilities.Existing safety measures often detect overt malicious intent but fail to address subtle, reasoning-driven vulnerabilities.In this work, we introduce POATE (Polar Opposite query generation, Adversarial Template construction, and Elaboration), a novel jailbreak technique that harnesses contrastive reasoning to provoke unethical responses.POATE crafts semantically opposing intents and integrates them with adversarial templates, steering models toward harmful outputs with remarkable subtlety.We conduct extensive evaluation across six diverse language model families of varying parameter sizes to demonstrate the robustness of the attack, achieving significantly higher attack success rates (~44%) compared to existing methods.To counter this, we propose Intent-Aware CoT and Reverse Thinking CoT, which decompose queries to detect malicious intent and reason in reverse to evaluate and reject harmful responses.These methods enhance reasoning robustness and strengthen the model's defense against adversarial exploits.Our code is publicly available 1 . Rachneet Sachdeva, Rima Hazra, Iryna Gurevych |
EMNLP | 2 |
| 2025 | How (Un)ethical Are Instruction-Centric Responses of LLMs? Unveiling the Vulnerabilities of Safety Guardrails to Harmful QueriesabstractIn this study, we tackle a growing concern around the safety and ethical use of large language models (LLMs). Despite their potential, these models can be tricked into producing harmful or unethical content through various sophisticated methods, including `jailbreaking' techniques and targeted manipulation. Our work zeroes in on a specific issue: to what extent LLMs can be led astray by asking them to generate responses that are instruction-centric such as a pseudocode, a program or a software snippet as opposed to vanilla text. To investigate this question, we introduce TechHazardQA, a dataset containing complex queries which should be answered in both text and instruction-centric formats (e.g., pseudocodes), aimed at identifying triggers for unethical responses. We query a series of LLMs -- Llama-2-13b, Llama-2-7b, Mistral-V2 and Mistral 8X7B -- and ask them to generate both text and instruction-centric responses. For evaluation we report the harmfulness score metric as well as judgements from GPT-4 and humans. Overall, we observe that asking LLMs to produce instruction-centric responses enhances the unethical response generation by 2-38% across the models. As an additional objective, we investigate the impact of model editing using the ROME technique, which further increases the propensity for generating undesirable content. We observe that the propensity to generate unethical content through instruction-centric responses in comparison to text responses increases significantly with a single edit, rising from an average of 18.9% to 56.7% in zero-shot scenarios, from 31.9% to 56.6% in zero-shot CoT, and from 22.8% to 65.7% in few-shot scenarios. Somnath Banerjee 0002, Sayan Layek, Rima Hazra, Animesh Mukherjee 0001 |
ICWSM | 3 |
| 2025 | Navigating the Cultural Kaleidoscope: A Hitchhiker's Guide to Sensitivity in Large Language ModelsabstractSomnath Banerjee, Sayan Layek, Hari Shrawgi, Rajarshi Mandal, Avik Halder, Shanu Kumar, Sagnik Basu, Parag Agrawal, Rima Hazra, Animesh Mukherjee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Somnath Banerjee 0002, Sayan Layek, Hari Shrawgi, Rajarshi Mandal, Avik Halder, Shanu Kumar, Sagnik Basu, Parag Agrawal, Rima Hazra, Animesh Mukherjee 0001 |
NAACL (Long Papers) | 9 |
| 2024 | Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and ActivationsabstractEnsuring the safe alignment of large language models (LLMs) with human values is critical as they become integral to applications like translation and question answering.Current alignment methods struggle with dynamic user intentions and complex objectives, making models vulnerable to generating harmful content.We propose SAFETY ARITH-METIC, a training-free framework enhancing LLM safety across different scenarios: Base models, Supervised fine-tuned models (SFT), and Edited models.SAFETY ARITH-METIC involves Harm Direction Removal to avoid harmful content and Safety Alignment to promote safe responses.Additionally, we present NOINTENTEDIT, a dataset highlighting edit instances that could compromise model safety if used unintentionally.Our experiments show that SAFETY ARITHMETIC significantly improves safety measures, reduces over-safety, and maintains model utility, outperforming existing methods in ensuring safe content generation.Source codes and dataset can be accessed at: https://github.com/ declare Rima Hazra, Sayan Layek, Somnath Banerjee 0002, Soujanya Poria |
EMNLP | 1 |
| 2024 | DistALANER: Distantly Supervised Active Learning Augmented Named Entity Recognition in the Open Source Software Ecosystem
Somnath Banerjee 0002, Avik Dutta, Aaditya Agrawal, Rima Hazra, Animesh Mukherjee 0001 |
ECML/PKDD (10) | 4 |
| 2023 | Duplicate Question Retrieval and Confirmation Time Prediction in Software CommunitiesabstractCommunity Question Answering (CQA) in different domains is growing at a large scale because of the availability of several platforms and huge shareable information among users. With the rapid growth of such online platforms, a massive amount of archived data makes it difficult for moderators to retrieve possible duplicates for a new question and identify and confirm existing question pairs as duplicates at the right time. This problem is even more critical in CQAs corresponding to large software systems like askubuntu where moderators need to be experts to comprehend something as a duplicate. Note that the prime challenge in such CQA platforms is that the moderators are themselves experts and are therefore usually extremely busy with their time being extraordinarily expensive. To facilitate the task of the moderators, in this work, we have tackled two significant issues for the askubuntu CQA platform: (1) retrieval of duplicate questions given a new question and (2) duplicate question confirmation time prediction. In the first task, we focus on retrieving duplicate questions from a question pool for a particular newly posted question. In the second task, we solve a regression problem to rank a pair of questions that could potentially take a long time to get confirmed as duplicates. For duplicate question retrieval, we propose a Siamese neural network based approach by exploiting both text and network-based features, which outperforms several state-of-the-art baseline techniques. Our method outperforms DupPredictor [33] and DUPE [1] by 5% and 7% respectively. For duplicate confirmation time prediction, we have used both the standard machine learning models and neural network along with the text and graph-based features. We obtain Spearman's rank correlation of 0.20 and 0.213 (statistically significant) for text and graph based features respectively. We shall place all our codes and data in the public domain upon acceptance. Rima Hazra, Debanjan Saha, Amruit Sahoo, Somnath Banerjee 0002, Animesh Mukherjee 0001 |
ASONAM | 1 |
| 2022 | Is This Bug Severe? A Text-Cum-Graph Based Model for Bug Severity Prediction
Rima Hazra, Arpit Dwivedi, Animesh Mukherjee 0001 |
ECML/PKDD (6) | 1 |
| 2021 | Joint Autoregressive and Graph Models for Software and Developer Social Networks
Rima Hazra, Hardik Aggarwal, Pawan Goyal 0002, Animesh Mukherjee 0001, Soumen Chakrabarti |
ECIR (1) | 1 |