VLDB 2026 Research / reviewers in the wild / expert
Nazneen Fatema Rajani
dblp:166/1729 · also Nazneen Rajani
· DBLP profile ↗
24ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-6301-1960ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing AgentsabstractDespite rapid progress in building conversational AI agents, robustness is still largely untested.Small shifts in user behavior, such as being more impatient, incoherent, or skeptical, can cause sharp drops in agent performance, revealing how brittle current AI agents are.Today's benchmarks fail to capture this fragility: agents may perform well under standard evaluations but degrade spectacularly in more realistic and varied settings.We address this robustness testing gap by introducing TraitBasis, a lightweight, model-agnostic method for systematically stress testing AI agents.TraitBasis learns directions in activation space corresponding to steerable user traits (e.g., impatience or incoherence), which can be controlled, scaled, composed, and applied at inference time without any fine-tuning or extra data.Using TraitBasis, we extend τ -Bench to τ -Trait, where user behaviors are altered via controlled trait vectors.We observe an average 4%-20% performance degradation on τ -Trait across frontier models, highlighting the lack of robustness of current AI agents to variations in user behavior.Together, these results highlight both the critical role of robustness testing and the promise of TraitBasis as a simple, dataefficient, and compositional tool.By powering simulation-driven stress tests and training loops, TraitBasis opens the door to building AI agents that remain reliable in the unpredictable dynamics of real-world human interactions.We plan to open-source τ -Trait across four domains: airline, retail, telecom, and telehealth, so the community can systematically QA their agents under realistic, behaviorally diverse intents and trait scenarios. Muyu He, Soumyadeep Bakshi, James Zou 0001, Nazneen Fatema Rajani |
ACL (1) | 5 |
| 2023 | What's New? Summarizing Contributions in Scientific LiteratureabstractWith thousands of academic articles shared on a daily basis, it has become increasingly difficult to keep up with the latest scientific findings.To overcome this problem, we introduce a new task of disentangled paper summarization, which seeks to generate separate summaries for the paper contributions and the context of the work, making it easier to identify the key findings shared in articles.For this purpose, we extend the S2ORC corpus of academic articles, which spans a diverse set of domains ranging from economics to psychology, by adding disentangled "contribution" and "context" reference labels.Together with the dataset, we introduce and analyze three baseline approaches: 1) a unified model controlled by input code prefixes, 2) a model with separate generation heads specialized in generating the disentangled outputs, and 3) a training strategy that guides the model using additional supervision coming from inbound and outbound citations.We also propose a comprehensive automatic evaluation protocol which reports the relevance, novelty, and disentanglement of generated outputs.Through a human study involving expert annotators, we show that in 79%, of cases our new task is considered more helpful than traditional scientific paper summarization. Hiroaki Hayashi, Wojciech Kryscinski, Bryan McCann, Nazneen Fatema Rajani, Caiming Xiong |
EACL | 4 |
| 2023 | Generative AI meets Responsible AI: Practical Challenges and OpportunitiesabstractGenerative AI models and applications are being rapidly developed and deployed across a wide spectrum of industries and applications ranging from writing and email assistants to graphic design and art generation to educational assistants to coding to drug discovery. However, there are several ethical and social considerations associated with generative AI models and applications. These concerns include lack of interpretability, bias and discrimination, privacy, lack of model robustness, fake and misleading content, copyright implications, plagiarism, and environmental impact associated with training and inference of generative AI models. Krishnaram Kenthapadi, Himabindu Lakkaraju, Nazneen Fatema Rajani |
KDD | 3 |
| 2022 | Conformal Predictor for Improving Zero-Shot Text Classification EfficiencyabstractPre-trained language models (PLMs) have been shown effective for zero-shot (0shot) text classification.0shot models based on natural language inference (NLI) and next sentence prediction (NSP) employ cross-encoder architecture and infer by making a forward pass through the model for each label-text pair separately.This increases the computational cost to make inferences linearly in the number of labels.In this work, we improve the efficiency of such cross-encoder-based 0shot models by restricting the number of likely labels using another fast base classifier-based conformal predictor (CP) calibrated on samples labeled by the 0shot model.Since a CP generates prediction sets with coverage guarantees, it reduces the number of target labels without excluding the most probable label based on the 0shot model.We experiment with three intent and two topic classification datasets.With a suitable CP for each dataset, we reduce the average inference time for NLI-and NSP-based models by 25.6% and 22.2% respectively, without dropping performance below the predefined error rate of 1%. Prafulla Kumar Choubey, Yu Bai 0017, Chien-Sheng Wu, Wenhao Liu 0003, Nazneen Fatema Rajani |
EMNLP | 5 |
| 2022 | HydraSum: Disentangling Style Features in Text Summarization with Multi-Decoder ModelsabstractSummarization systems make numerous "decisions" about summary properties during inference, e.g.degree of copying, specificity and length of outputs, etc.However, these are implicitly encoded within model parameters and specific styles cannot be enforced.To address this, we introduce HYDRASUM, a new summarization architecture that extends the single decoder framework of current models to a mixture-of-experts version with multiple decoders.We show that HYDRASUM's multiple decoders automatically learn contrasting summary styles when trained under the standard training objective without any extra supervision.Through experiments on three summarization datasets (CNN, NEWSROOM and XSUM), we show that HYDRASUM provides a simple mechanism to obtain stylistically-diverse summaries by sampling from either individual decoders or their mixtures, outperforming baseline models.Finally, we demonstrate that a small modification to the gating strategy during training can enforce an even stricter style partitioning, e.g.high-vs low-abstractiveness or high-vs low-specificity, allowing users to sample from a larger area in the generation space and vary summary styles along multiple dimensions. 1Input Article: Insights into the workings of the human body that Leonardo da Vinci could only obtain by dissecting scores of corpses and recording the results in exquisite drawings will be displayed for the first time beside modern 3D films, CT and MRI scans, which show how close the Renaissance genius got to the truth of what lies under the skin.[…] the Edinburgh show will be the first to compare Leonardo's results with scalpel and pen with the best results of modern technology.[…] The exhibition will show how close Leonardo got in some of his last medical experiments to discovering the role of the beating heart in the circulation of the blood, a century before William Harvey worked it out.[…] Edinburgh show will be first to compare Renaissance genius's results with best results of modern technology.Edinburgh show will be first to compare Renaissance genius's results with the best results of modern technology.Edinburgh show will be first to compare Leonardo's results with the best results of modern technology. Low diversity Baseline BARTEdinburgh show will be first to compare Leonardo's results with best results of modern technology.Modern imaging techniques will be displayed alongside Leonardo da Vinci's anatomical drawings in Edinburgh exhibition. Tanya Goyal, Nazneen Fatema Rajani, Wenhao Liu 0003, Wojciech Kryscinski |
EMNLP | 2 |
| 2022 | CTRLsum: Towards Generic Controllable Text SummarizationabstractCurrent summarization systems yield generic summaries that are disconnected from users' preferences and expectations.To address this limitation, we present CTRLSUM, a generic framework to control generated summaries through a set of keywords.During training keywords are extracted automatically without requiring additional human annotations.At test time CTRLSUM features a control function to map control signal to keywords; through engineering the control function, the same trained model is able to be applied to control summaries on various dimensions, while neither affecting the model training process nor the pretrained models.We additionally explore the combination of keywords and text prompts for more control tasks.Experiments demonstrate the effectiveness of CTRLSUM on three domains of summarization datasets and five control tasks: (1) entity-centric and (2) length-controllable summarization, (3) contribution summarization on scientific papers, (4) invention purpose summarization on patent filings, and (5) question-guided summarization on news articles.Moreover, when used in a standard, unconstrained summarization setting, CTRLSUM is comparable or better than strong pretrained systems. 1 Junxian He, Wojciech Kryscinski, Bryan McCann, Nazneen Fatema Rajani, Caiming Xiong |
EMNLP | 4 |
| 2022 | Are Hard Examples also Harder to Explain? A Study with Human and Model-Generated ExplanationsabstractRecent work on explainable NLP has shown that few-shot prompting can enable large pretrained language models (LLMs) to generate grammatical and factual natural language explanations for data labels.In this work, we study the connection between explainability and sample hardness by investigating the following research question -"Are LLMs and humans equally good at explaining data labels for both easy and hard samples?"We answer this question by first collecting humanwritten explanations in the form of generalizable commonsense rules on the task of Winograd Schema Challenge (Winogrande dataset).We compare these explanations with those generated by GPT-3 while varying the hardness of the test samples as well as the in-context samples.We observe that (1) GPT-3 explanations are as grammatical as human explanations regardless of the hardness of the test samples, (2) for easy examples, GPT-3 generates highly supportive explanations but human explanations are more generalizable, and (3) for hard examples, human explanations are significantly better than GPT-3 explanations both in terms of label-supportiveness and generalizability judgements.We also find that hardness of the in-context examples impacts the quality of GPT-3 explanations.Finally, we show that the supportiveness and generalizability aspects of human explanations are also impacted by sample hardness, although by a much smaller margin than models. 1 Swarnadeep Saha, Peter Hase, Nazneen Fatema Rajani, Mohit Bansal |
EMNLP | 3 |
| 2022 | P-Adapters: Robustly Extracting Factual Information from Language Models with Diverse Prompts
Benjamin Newman, Prafulla Kumar Choubey, Nazneen Fatema Rajani |
ICLR | 3 |
| 2022 | iSEA: An Interactive Pipeline for Semantic Error Analysis of NLP ModelsabstractError analysis in NLP models is essential to successful model development and deployment. One common approach for diagnosing errors is to identify subpopulations in the dataset where the model produces the most errors. However, existing approaches typically define subpopulations based on pre-defined features, which requires users to form hypotheses of errors in advance. To complement these approaches, we propose iSEA, an Interactive Pipeline for Semantic Error Analysis in NLP Models, which automatically discovers semantically-grounded subpopulations with high error rates in the context of a human-in-the-loop interactive system. iSEA enables model developers to learn more about their model errors through discovered subpopulations, validate the sources of errors through interactive analysis on the discovered subpopulations, and test hypotheses about model errors by defining custom subpopulations. The tool supports semantic descriptions of error-prone subpopulations at the token and concept level, as well as pre-defined higher-level features. Through use cases and expert interviews, we demonstrate how iSEA can assist error understanding and analysis. Jesse Vig, Nazneen Fatema Rajani |
IUI | 3 |
| 2021 | FastIF: Scalable Influence Functions for Efficient Model Interpretation and DebuggingabstractInfluence functions approximate the "influences" of training data-points for test predictions and have a wide variety of applications.Despite the popularity, their computational cost does not scale well with model and training data size.We present FASTIF, a set of simple modifications to influence functions that significantly improves their run-time.We use k-Nearest Neighbors (kNN) to narrow the search space down to a subset of good candidate data points, identify the configurations that best balance the speed-quality trade-off in estimating the inverse Hessian-vector product, and introduce a fast parallel variant.Our proposed method achieves about 80X speedup while being highly correlated with the original influence values.With the availability of the fast influence functions, we demonstrate their usefulness in four applications.First, we examine whether influential data-points can "explain" test time behavior using the framework of simulatability.Second, we visualize the influence interactions between training and test data-points.Third, we show that we can correct model errors by additional fine-tuning on certain influential data-points, improving the accuracy of a trained MultiNLI model by 2.5% on the HANS dataset.Finally, we experiment with a similar setup but fine-tuning on datapoints not seen during training, improving the model accuracy by 2.8% and 1.7% on HANS and ANLI datasets respectively.Overall, our fast influence functions can be efficiently applied to large models and datasets, and our experiments demonstrate the potential of influence functions in model interpretation and correcting model errors. 1 Nazneen Fatema Rajani, Peter Hase, Mohit Bansal, Caiming Xiong |
EMNLP (1) | 2 |
| 2021 | CoCo: Controllable Counterfactuals for Evaluating Dialogue State Trackers
Semih Yavuz, Kazuma Hashimoto, Jia Li 0015, Nazneen Fatema Rajani, Xifeng Yan, Yingbo Zhou 0002, Caiming Xiong |
ICLR | 6 |
| 2021 | BERTology Meets Biology: Interpreting Attention in Protein Language Models
Jesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong, Richard Socher, Nazneen Fatema Rajani |
ICLR | 6 |
| 2021 | DART: Open-Domain Structured Data Record to Text GenerationabstractLinyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, Nazneen Fatema Rajani. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Linyong Nan, Dragomir R. Radev, Rui Zhang 0037, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma 0001, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta 0015, Tao Yu 0009, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, Nazneen Fatema Rajani |
NAACL-HLT | 24 |
| 2020 | ERASER: A Benchmark to Evaluate Rationalized NLP ModelsabstractJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, Byron C. Wallace. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Jay DeYoung, Nazneen Fatema Rajani, Eric P. Lehman, Caiming Xiong, Richard Socher, Byron C. Wallace |
ACL | 3 |
| 2020 | ESPRIT: Explaining Solutions to Physical Reasoning TasksabstractNazneen Fatema Rajani, Rui Zhang, Yi Chern Tan, Stephan Zheng, Jeremy Weiss, Aadit Vyas, Abhijit Gupta, Caiming Xiong, Richard Socher, Dragomir Radev. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Nazneen Fatema Rajani, Rui Zhang 0037, Yi Chern Tan, Stephan Zheng, Jeremy Weiss, Aadit Vyas, Abhijit Gupta, Caiming Xiong, Richard Socher, Dragomir R. Radev |
ACL | 1 |
| 2020 | Double-Hard Debias: Tailoring Word Embeddings for Gender Bias MitigationabstractWord embeddings derived from humangenerated corpora inherit strong gender bias which can be further amplified by downstream models.Some commonly adopted debiasing approaches, including the seminal Hard Debias algorithm (Bolukbasi et al., 2016), apply post-processing procedures that project pre-trained word embeddings into a subspace orthogonal to an inferred gender subspace.We discover that semantic-agnostic corpus regularities such as word frequency captured by the word embeddings negatively impact the performance of these algorithms.We propose a simple but effective technique, Double-Hard Debias, which purifies the word embeddings against such corpus regularities prior to inferring and removing the gender subspace.Experiments on three bias mitigation benchmarks show that our approach preserves the distributional semantics of the pre-trained word embeddings while reducing gender bias to a significantly larger degree than prior approaches. Xi Victoria Lin, Nazneen Fatema Rajani, Bryan McCann, Vicente Ordonez, Caiming Xiong |
ACL | 3 |
| 2020 | Universal Natural Language Processing with Limited Annotations: Try Few-shot Textual Entailment as a StartabstractA standard way to address different NLP problems is by first constructing a problem-specific dataset, then building a model to fit this dataset.To build the ultimate artificial intelligence, we desire a single machine that can handle diverse new problems, for which task-specific annotations are limited.We bring up textual entailment as a unified solver for such NLP problems.However, current research of textual entailment has not spilled much ink on the following questions: (i) How well does a pretrained textual entailment system generalize across domains with only a handful of domainspecific examples?and (ii) When is it worth transforming an NLP task into textual entailment?We argue that the transforming is unnecessary if we can obtain rich annotations for this task.Textual entailment really matters particularly when the target NLP task has insufficient annotations.Universal NLP 1 can be probably achieved through different routines.In this work, we introduce Universal Few-shot textual Entailment (UFO-ENTAIL).We demonstrate that this framework enables a pretrained entailment model to work well on new entailment domains in a few-shot setting, and show its effectiveness as a unified solver for several downstream NLP tasks such as question answering and coreference resolution when the end-task annotations are limited. Wenpeng Yin 0001, Nazneen Fatema Rajani, Dragomir R. Radev, Richard Socher, Caiming Xiong |
EMNLP (1) | 2 |
| 2020 | ReviewRobot: Explainable Paper Review Generation based on Knowledge SynthesisabstractTo assist human review process, we build a novel ReviewRobot to automatically assign a review score and write comments for multiple categories such as novelty and meaningful comparison.A good review needs to be knowledgeable, namely that the comments should be constructive and informative to help improve the paper; and explainable by providing detailed evidence.ReviewRobot achieves these goals via three steps: (1) We perform domainspecific Information Extraction to construct a knowledge graph (KG) from the target paper under review, a related work KG from the papers cited by the target paper, and a background KG from a large collection of previous papers in the domain.(2) By comparing these three KGs, we predict a review score and detailed structured knowledge as evidence for each review category.(3) We carefully select and generalize human review sentences into templates, and apply these templates to transform the review scores and evidence into natural language comments.Experimental results show that our review score predictor reaches 71.4%-100% accuracy.Human assessment by domain experts shows that 41.7%-70.5% of the comments generated by ReviewRobot are valid and constructive, and better than humanwritten ones for 20% of the time.Thus, Re-viewRobot can serve as an assistant for paper reviewers, program chairs and authors. 1 Qingyun Wang 0005, Qi Zeng 0001, Lifu Huang, Kevin Knight, Heng Ji 0001, Nazneen Fatema Rajani |
INLG | 6 |
| 2019 | Explain Yourself! Leveraging Language Models for Commonsense ReasoningabstractDeep learning models perform poorly on tasks that require commonsense reasoning, which often necessitates some form of worldknowledge or reasoning over information not immediately present in the input.We collect human explanations for commonsense reasoning in the form of natural language sequences and highlighted annotations in a new dataset called Common Sense Explanations (CoS-E).We use CoS-E to train language models to automatically generate explanations that can be used during training and inference in a novel Commonsense Auto-Generated Explanation (CAGE) framework.CAGE improves the state-of-the-art by 10% on the challenging CommonsenseQA task.We further study commonsense reasoning in DNNs using both human and auto-generated explanations including transfer to out-of-domain tasks.Empirical results indicate that we can effectively leverage language models for commonsense reasoning. Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, Richard Socher |
ACL (1) | 1 |
| 2018 | Stacking with Auxiliary Features for Visual Question AnsweringabstractNazneen Fatema Rajani, Raymond Mooney. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Nazneen Fatema Rajani, Raymond J. Mooney |
NAACL-HLT | 1 |
| 2017 | Stacking With Auxiliary FeaturesabstractEnsembling methods are well known for improving prediction accuracy. However, they are limited in the sense that they cannot effectively discriminate among component models. In this paper, we propose stacking with auxiliary features that learns to fuse additional relevant information from multiple component systems as well as input instances to improve performance. We use two types of auxiliary features -- instance features and provenance features. The instance features enable the stacker to discriminate across input instances and the provenance features enable the stacker to discriminate across component systems. When combined together, our algorithm learns to rely on systems that not just agree on an output but also the provenance of this output in conjunction with the properties of the input instance. We demonstrate the success of our approach on three very different and challenging natural language and vision problems: Slot Filling, Entity Discovery and Linking, and ImageNet Object Detection. We obtain new state-of-the-art results on the first two tasks and significant improvements on the ImageNet task, thus verifying the power and generality of our approach. Nazneen Fatema Rajani, Raymond J. Mooney |
IJCAI | 1 |
| 2016 | Combining Supervised and Unsupervised Enembles for Knowledge Base PopulationabstractWe propose an algorithm that combines supervised and unsupervised methods to ensemble multiple systems for two popular Knowledge Base Population (KBP) tasks, Cold Start Slot Filling (CSSF) and Tri-lingual Entity Discovery and Linking (TEDL).We demonstrate that it outperforms the best system for both tasks in the 2015 competition, several ensembling baselines, as well as a state-of-the-art stacking approach.The success of our technique on two different and challenging problems demonstrates the power and generality of our combined approach to ensembling. Nazneen Fatema Rajani, Raymond J. Mooney |
EMNLP | 1 |
| 2015 | Stacked Ensembles of Information Extractors for Knowledge-Base PopulationabstractVidhoon Viswanathan, Nazneen Fatema Rajani, Yinon Bentor, Raymond Mooney. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Vidhoon Viswanathan, Nazneen Fatema Rajani, Yinon Bentor, Raymond J. Mooney |
ACL (1) | 2 |
| 2014 | Extracting topics based on authors, recipients and content in microblogsabstractMicroblogs such as Twitter are important sources for spreading vital information at high speed. They also reflect the general people's reaction and opinion towards major events or stories. With information traveling so quickly, it is helpful to be able to apply unsupervised learning techniques to discover topics for information extraction and analysis. Although graphical models have been traditionally used for topic discovery in microblogs and text streams, previous work may not be as efficient because of the diverse and noisy nature of microblogs. Nazneen Fatema Rajani, Kate McArdle, Jason Baldridge |
SIGIR | 1 |