VLDB 2026 Research / reviewers in the wild / expert
Ninareh Mehrabi
dblp:230/8151
· DBLP profile ↗
13ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SWAN: Semantic Watermarking with Abstract Meaning RepresentationabstractZiping Ye, Gourab Dey, Christos Christodoulopoulos, Charith Peris, Anil Ramakrishna, Weitong Ruan, Aram Galstyan, Kai-Wei Chang, Rahul Gupta, Ninareh Mehrabi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ziping Ye, Gourab Dey, Christos Christodoulopoulos 0001, Charith Peris, Anil Ramakrishna, Weitong Ruan, Aram Galstyan, Kai-Wei Chang 0001, Rahul Gupta 0001, Ninareh Mehrabi |
ACL (1) | 10 |
| 2026 | Breaking the Benchmark: Revealing LLM Bias via Minimal Contextual Augmentation
Kaveh Eskandari Miandoab, Mahammed Kamruzzaman, Arshia Gharooni, Gene Louis Kim, Vasanth Sarathy, Ninareh Mehrabi |
LREC | 6 |
| 2025 | Diagnosing Memorization in Chain-of-Thought Reasoning, One Token at a TimeabstractLarge Language Models (LLMs) perform well on reasoning benchmarks but often fail when inputs alter slightly, raising concerns about the extent to which their success relies on memorization.This issue is especially acute in Chain-of-Thought (CoT) reasoning, where spurious memorized patterns can trigger intermediate errors that cascade into incorrect final answers.We introduce STIM, a novel framework for Source-aware Token-level Identification of Memorization, which attributes each token in a reasoning chain to one of multiple memorization sources -local, mid-range, or long-rangebased on their statistical co-occurrence with the token in the pretraining corpus.Our token-level analysis across tasks and distributional settings reveals that models rely more on memorization in complex or long-tail cases, and that local memorization is often the dominant driver of errors, leading to up to 67% of wrong tokens.We also show that memorization scores from STIM can be effective in predicting the wrong tokens in the wrong reasoning step.STIM offers a powerful tool for diagnosing and improving model reasoning and can generalize to other structured step-wise generation tasks. 1 * Equal Contribution † Work done in Amazon AGI 1 https://github.com/INK-USC/STIMModel Reasoning Steps: Rick killed 10 + 15 = 25 animals last night.Today Rick killed 3 * 3 = 9 wolves.Today Rick killed 15 -3 = 12 cougars.Today Rick killed 9 + 12 = 21 animals.Total animals killed by Rick is 25 + 21 = 46.So the answer is 46. Examples of Simultaneous Memorization Sources of Different StrengthQuestion: Here is a list: [pitanga, pitanga, yumberry, yumberry, pitanga, yumberry, pitanga, yumberry, pitanga, pitanga].How many times does 'pitanga' appear on it?Model Output: 1.The first element is 'pitanga'.... 10.The tenth element is 'pitanga'.Now, counting the occurrences of 'pitanga': We have 'pitanga' appearing 5 times at positions 1, 2, 5, 7, and 10.So the answer is 5.Analysis: Token 5 is incorrect and has the highest influence from local context, where the n-gram "appearing 5" has a high pretraining frequency.STIM scores are: local = 0.81 , mid = 0.24, long = 0.03, suggesting that the model was influenced by locally frequent patterns rather than actual counting.Question: What is ( 100101100 -100101100 * 0.1) * 111110100 + 100101100 * 1111 / 1100100 * 111110100 * 0.1 + (100101100 -100101100 * 0.1 -100101100 * 1111 / 1100100) * 111110100 equal to? (base-2) Model Output:To solve this, we first handle the multiplication and division in base-2, converting to base-10 as necessary:1. (100101100 × 0.1base-2 = 10010110base-2) (which is 180 in base-10) ... So the answer is 26807.536.Analysis: Token 180 is incorrect and has the highest influence from input tokens, having higher pretraining co-occurrence frequency with "111110100", "*" and "100101100" in the input.STIM scores are local=-0.19,mid=0.09,long=0.156, indicating long-range memorization being the primary influence. Huihan Li 0001, Ninareh Mehrabi, Rahul Gupta 0001, Xiang Ren 0001 |
EMNLP | 5 |
| 2025 | DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM ReasoningabstractZero-shot Event Detection (ED), the task of identifying event mentions in natural language text without any training data, is critical for document understanding in specialized domains.Understanding the complex event ontology, extracting domain-specific triggers from the passage, and structuring them appropriately overloads and limits the utility of Large Language Models (LLMs) for zero-shot ED.To this end, we propose DICORE, a divergent-convergent reasoning framework that decouples the task of ED using Dreamer and Grounder.Dreamer encourages divergent reasoning through openended event discovery, which helps to boost event coverage.Conversely, Grounder introduces convergent reasoning to align the freeform predictions with the task-specific instructions using finite-state machine guided constrained decoding.Additionally, an LLM-Judge verifies the final outputs to ensure high precision.Through extensive experiments on six datasets across five domains and nine LLMs, we demonstrate how DICORE consistently outperforms prior zero-shot, transfer-learning, and reasoning baselines, achieving 4-7% average F1 gains over the best baseline -establishing DICORE as a strong zero-shot ED framework. Tanmay Parekh, Kartik Mehta, Ninareh Mehrabi, Kai-Wei Chang 0001, Nanyun Peng 0001 |
EMNLP | 3 |
| 2024 | Tree-of-Traversals: A Zero-Shot Reasoning Algorithm for Augmenting Black-box Language Models with Knowledge GraphsabstractElan Markowitz, Anil Ramakrishna, Jwala Dhamala, Ninareh Mehrabi, Charith Peris, Rahul Gupta, Kai-Wei Chang, Aram Galstyan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Elan Markowitz, Anil Ramakrishna, Jwala Dhamala, Ninareh Mehrabi, Charith Peris, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan |
ACL (1) | 4 |
| 2024 | FLIRT: Feedback Loop In-context Red TeamingabstractNinareh Mehrabi, Palash Goyal, Christophe Dupuy, Qian Hu, Shalini Ghosh, Richard Zemel, Kai-Wei Chang, Aram Galstyan, Rahul Gupta. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ninareh Mehrabi, Palash Goyal, Christophe Dupuy, Shalini Ghosh, Richard S. Zemel, Kai-Wei Chang 0001, Aram Galstyan, Rahul Gupta 0001 |
EMNLP | 1 |
| 2024 | Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language ModelsabstractData is a crucial element in large language model (LLM) alignment.Recent studies have explored using LLMs for efficient data collection.However, LLM-generated data often suffers from quality issues, with underrepresented or absent aspects and low-quality datapoints.To address these problems, we propose DATA ADVISOR, an enhanced LLMbased method for generating data that takes into account the characteristics of the desired dataset.Starting from a set of pre-defined principles in hand, DATA ADVISOR monitors the status of the generated data, identifies weaknesses in the current dataset, and advises the next iteration of data generation accordingly.DATA ADVISOR can be easily integrated into existing data generation methods to enhance data quality and coverage.Experiments on safety alignment of three representative LLMs (i.e., Mistral, Llama2, and Falcon) demonstrate the effectiveness of DATA ADVISOR in enhancing model safety against various fine-grained safety issues without sacrificing model utility.Warning: this paper contains example data that may be offensive or harmful. Fei Wang 0060, Ninareh Mehrabi, Palash Goyal, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan |
EMNLP | 2 |
| 2024 | The steerability of large language models toward data-driven personasabstractJunyi Li, Charith Peris, Ninareh Mehrabi, Palash Goyal, Kai-Wei Chang, Aram Galstyan, Richard Zemel, Rahul Gupta. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Junyi Li 0002, Charith Peris, Ninareh Mehrabi, Palash Goyal, Kai-Wei Chang 0001, Aram Galstyan, Richard S. Zemel, Rahul Gupta 0001 |
NAACL-HLT | 3 |
| 2023 | Resolving Ambiguities in Text-to-Image Generative ModelsabstractNinareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Varun Kumar, Qian Hu, Kai-Wei Chang, Richard Zemel, Aram Galstyan, Rahul Gupta. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ninareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Kai-Wei Chang 0001, Richard S. Zemel, Aram Galstyan, Rahul Gupta 0001 |
ACL (1) | 1 |
| 2022 | Robust Conversational Agents against Imperceptible Toxicity TriggersabstractNinareh Mehrabi, Ahmad Beirami, Fred Morstatter, Aram Galstyan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ninareh Mehrabi, Ahmad Beirami, Fred Morstatter, Aram Galstyan |
NAACL-HLT | 1 |
| 2021 | Exacerbating Algorithmic Bias through Fairness AttacksabstractAlgorithmic fairness has attracted significant attention in recent years, with many quantitative measures suggested for characterizing the fairness of different machine learning algorithms. Despite this interest, the robustness of those fairness measures with respect to an intentional adversarial attack has not been properly addressed. Indeed, most adversarial machine learning has focused on the impact of malicious attacks on the accuracy of the system, without any regard to the system's fairness. We propose new types of data poisoning attacks where an adversary intentionally targets the fairness of a system. Specifically, we propose two families of attacks that target fairness measures. In the anchoring attack, we skew the decision boundary by placing poisoned points near specific target points to bias the outcome. In the influence attack on fairness, we aim to maximize the covariance between the sensitive attributes and the decision outcome and affect the fairness of the model. We conduct extensive experiments that indicate the effectiveness of our proposed attacks. Ninareh Mehrabi, Fred Morstatter, Aram Galstyan |
AAAI | 1 |
| 2021 | Lawyers are Dishonest? Quantifying Representational Harms in Commonsense Knowledge ResourcesabstractWarning: this paper contains content that may be offensive or upsetting.Commonsense knowledge bases (CSKB) are increasingly used for various natural language processing tasks.Since CSKBs are mostly human-generated and may reflect societal biases, it is important to ensure that such biases are not conflated with the notion of commonsense.Here we focus on two widely used CSKBs, ConceptNet and GenericsKB, and establish the presence of bias in the form of two types of representational harms, overgeneralization of polarized perceptions and representation disparity across different demographic groups in both CSKBs.Next, we find similar representational harms for downstream models that use ConceptNet.Finally, we propose a filtering-based approach for mitigating such harms, and observe that our filtered-based approach can reduce the issues in both resources and models but leads to a performance drop, leaving room for future work to build fairer and stronger commonsense models. Ninareh Mehrabi, Fred Morstatter, Jay Pujara, Xiang Ren 0001, Aram Galstyan |
EMNLP (1) | 1 |
| 2019 | Debiasing community detection: the importance of lowly connected nodesabstractCommunity detection is an important task in social network analysis, allowing us to identify and understand the communities within the social structures provided by the network. However, many community detection approaches either fail to assign low-degree (or lowly connected) users to communities, or assign them to trivially small communities that prevent them from being included in analysis. In this work we investigate how excluding these users can bias analysis results. We then introduce an approach that is more inclusive for lowly connected users by incorporating them into larger groups. Experiments show that our approach outperforms the existing state-of-the-art in terms of F1 and Jaccard similarity scores while reducing the bias towards low-degree users. Ninareh Mehrabi, Fred Morstatter, Nanyun Peng 0001, Aram Galstyan |
ASONAM | 1 |