Nitay Calderon

dblp:305/0373 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
abstract
The “LLM-as-an-annotator” and “LLM-as-a-judge” paradigms employ Large Language Models (LLMs) as annotators, judges, and evaluators in tasks traditionally performed by humans. LLM annotations are widely used, not only in NLP research but also in fields like medicine, psychology, and social science. Despite their role in shaping study results and insights, there is no standard or rigorous procedure to determine whether LLMs can replace human annotators. In this paper, we propose a novel statistical procedure, the Alternative Annotator Test (alt-test), that requires only a modest subset of annotated examples to justify using LLM annotations. Additionally, we introduce a versatile and interpretable measure for comparing LLM annotators and judges. To demonstrate our procedure, we curated a diverse collection of ten datasets, consisting of language and vision-language tasks, and conducted experiments with six LLMs and four prompting techniques. Our results show that LLMs can sometimes replace humans with closed-source LLMs (such as GPT-4o), outperforming the open-source LLMs we examine, and that prompting techniques yield judges of varying quality. We hope this study encourages more rigorous and reliable practices.
Nitay Calderon, Roi Reichart, Rotem Dror
ACL (1)1
2025 Multi-Domain Explainability of Preferences
abstract
Preference mechanisms, such as human preference, LLM-as-a-Judge (LaaJ), and reward models, are central to aligning and evaluating large language models (LLMs).Yet, the underlying concepts that drive these preferences remain poorly understood.In this work, we propose a fully automated method for generating local and global concept-based explanations of preferences across multiple domains.Our method utilizes an LLM to identify concepts (rubrics) that distinguish between chosen and rejected responses, and to represent them with conceptbased vectors.To model the relationships between concepts and preferences, we propose a white-box Hierarchical Multi-Domain Regression model that captures both domain-general and domain-specific effects.To evaluate our method, we curate a dataset spanning eight diverse domains and explain twelve mechanisms.Our method achieves strong preference prediction performance, outperforming baselines while also being explainable.Additionally, we assess explanations in two application-driven settings.First, guiding LLM outputs with concepts from LaaJ explanations yields responses that those judges consistently prefer.Second, prompting LaaJs with concepts explaining humans improves their preference predictions.Together, our work establishes a new paradigm for explainability in the era of LLMs. 1
Nitay Calderon, Liat Ein-Dor, Roi Reichart
EMNLP1
2025 Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance
abstract
NLP benchmarks rely on standardized datasets for training and evaluating models and are crucial for advancing the field.Traditionally, expert annotations ensure high-quality labels; however, the cost of expert annotation does not scale well with the growing demand for larger datasets required by modern models.While crowd-sourcing provides a more scalable solution, it often comes at the expense of annotation precision and consistency.Recent advancements in large language models (LLMs) offer new opportunities to enhance the annotation process, particularly for detecting label errors in existing datasets.In this work, we consider the recent approach of LLM-as-a-judge, leveraging an ensemble of LLMs to flag potentially mislabeled examples.We conduct a case study on four factual consistency datasets from the TRUE benchmark, spanning diverse NLP tasks, and on SummEval, which uses Likertscale ratings of summary quality across multiple dimensions.We empirically analyze the labeling quality of existing datasets and compare expert, crowd-sourced, and LLM-based annotations in terms of the agreement, label quality, and efficiency, demonstrating the strengths and limitations of each annotation method.Our findings reveal a substantial number of label errors, which, when corrected, induce a significant upward shift in reported model performance.This suggests that many of the LLMs' so-called mistakes are due to label errors rather than genuine model failures.Additionally, we discuss the implications of mislabeled data and propose methods to mitigate them in training to improve performance.
Omer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor, Roi Reichart
EMNLP2
2025 NL-Eye: Abductive NLI For Images
abstract
Will a Visual Language Model (VLM)-based bot warn us about slipping if it detects a wet floor? Recent VLMs have demonstrated impressive capabilities, yet their ability to infer outcomes and causes remains underexplored. To address this, we introduce NL-Eye, a benchmark designed to assess VLMs' visual abductive reasoning skills. NL-Eye adapts the abductive Natural Language Inference (NLI) task to the visual domain, requiring models to evaluate the plausibility of hypothesis images based on a premise image and explain their decisions. NL-Eye consists of 350 carefully curated triplet examples (1,050 images) spanning diverse reasoning categories: physical, functional, logical, emotional, cultural, and social. The data curation process involved two steps—writing textual descriptions and generating images using text-to-image models, both requiring substantial human involvement to ensure high-quality and challenging scenes. Our experiments show that VLMs struggle significantly on NL-Eye, often performing at random baseline levels, while humans excel in both plausibility prediction and explanation quality. This demonstrates a deficiency in the abductive reasoning capabilities of modern VLMs. NL-Eye represents a crucial step toward developing VLMs capable of robust multimodal reasoning for real-world applications, including accident-prevention bots and generated video verification.
Mor Ventura, Michael Toker, Nitay Calderon, Zorik Gekhman, Yonatan Bitton, Roi Reichart
ICLR3
2025 On Behalf of the Stakeholders: Trends in NLP Model Interpretability in the Era of LLMs
abstract
Nitay Calderon, Roi Reichart. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Nitay Calderon, Roi Reichart
NAACL (Long Papers)1
2024 Faithful Explanations of Black-box NLP Models Using LLM-generated Counterfactuals
abstract
Causal explanations of the predictions of NLP systems are essential to ensure safety and establish trust. Yet, existing methods often fall short of explaining model predictions effectively or efficiently and are often model-specific. In this paper, we address model-agnostic explanations, proposing two approaches for counterfactual (CF) approximation. The first approach is CF generation, where a large language model (LLM) is prompted to change a specific text concept while keeping confounding concepts unchanged. While this approach is demonstrated to be very effective, applying LLM at inference-time is costly. We hence present a second approach based on matching, and propose a method that is guided by an LLM at training-time and learns a dedicated embedding space. This space is faithful to a given causal graph and effectively serves to identify matches that approximate CFs. After showing theoretically that approximating CFs is required in order to construct faithful explanations, we benchmark our approaches and explain several models, including LLMs with billions of parameters. Our empirical results demonstrate the excellent performance of CF generation models as model-agnostic explainers. Moreover, our matching approach, which requires far less test-time resources, also provides effective explanations, surpassing many baselines. We also find that Top-K techniques universally improve every tested method. Finally, we showcase the potential of LLMs in constructing new benchmarks for model explanation and subsequently validate our conclusions. Our work illuminates new pathways for efficient and accurate approaches to interpreting NLP systems.
Yair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin, Roi Reichart
ICLR2
2024 The Colorful Future of LLMs: Evaluating and Improving LLMs as Emotional Supporters for Queer Youth
abstract
Shir Lissak, Nitay Calderon, Geva Shenkman, Yaakov Ophir, Eyal Fruchter, Anat Brunstein Klomek, Roi Reichart. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Shir Lissak, Nitay Calderon, Geva Shenkman, Yaakov Ophir, Eyal Fruchter, Anat Brunstein Klomek, Roi Reichart
NAACL-HLT2
2023 A Systematic Study of Knowledge Distillation for Natural Language Generation with Pseudo-Target Training
abstract
Modern Natural Language Generation (NLG) models come with massive computational and storage requirements.In this work, we study the potential of compressing them, which is crucial for real-world applications serving millions of users.We focus on Knowledge Distillation (KD) techniques, in which a small student model learns to imitate a large teacher model, allowing to transfer knowledge from the teacher to the student.In contrast to much of the previous work, our goal is to optimize the model for a specific NLG task and a specific dataset.Typically in real-world applications, in addition to labeled data there is abundant unlabeled task-specific data, which is crucial for attaining high compression rates via KD.In this work, we conduct a systematic study of task-specific KD techniques for various NLG tasks under realistic assumptions.We discuss the special characteristics of NLG distillation and particularly the exposure bias problem.Following, we derive a family of Pseudo-Target (PT) augmentation methods, substantially extending prior work on sequence-level KD.We propose the Joint-Teaching method, which applies wordlevel KD to multiple PTs generated by both the teacher and the student.Finally, we validate our findings in an extreme setup with no labeled examples using GPT-4 as the teacher.Our study provides practical model design observations and demonstrates the effectiveness of PT training for task-specific KD in NLG.
Nitay Calderon, Subhabrata Mukherjee, Roi Reichart, Amir Kantor
ACL (1)1
2022 DoCoGen: Domain Counterfactual Generation for Low Resource Domain Adaptation
abstract
Natural language processing (NLP) algorithms have become very successful, but they still struggle when applied to out-of-distribution examples.In this paper we propose a controllable generation approach in order to deal with this domain adaptation (DA) challenge.Given an input text example, our DoCoGen algorithm generates a domain-counterfactual textual example (D-CON) -that is similar to the original in all aspects, including the task label, but its domain is changed to a desired one.Importantly, DoCoGen is trained using only unlabeled examples from multiple domainsno NLP task labels or parallel pairs of textual examples and their domain-counterfactuals are required.We show that DoCoGen can generate coherent counterfactuals consisting of multiple sentences.We use the D-CONs generated by DoCoGen to augment a sentiment classifier and a multi-label intent classifier in 20 and 78 DA setups, respectively, where source-domain labeled data is scarce.Our model outperforms strong baselines and improves the accuracy of a state-of-the-art unsupervised DA algorithm.1 * Both authors equally contributed to this work.Original, Kitchen: A good knife but Quality Control was poor.The knife is solid and very comfortable in hand, however, when I got it new, the blade is slightly bent.I expect it to be in almost perfect condition, but it's not.DoCoGen, Kitchen → Electronics: A good product but Quality Control was poor.The ipod is very easy to use and very comfortable in hand, however, when I got it new, the ipod is slightly flimsy.I expect it to be in almost perfect shape, but it's not.Original, DVD: The direction of this film is excellent.I love all the characters and the way they interact.The storyline is very important also.It's about religious beliefs and neighbors that interact with each other.It's a well-paced and interesting story that's not like anything else I've ever seen.DoCoGen, DVD → Airline: The service on this flight is excellent.I love the staff and the way they interact.The safety is very important also.It's nice to have staff and neighbors that can help each other.It's a well-groomed and professional crew that's not like anything else I've ever experienced.Original, Electronics: That relay board is only good for switching AC loads of 100V or more.If you have a lower voltage load, it's not going to work.For low voltage loads use transistors, MOSFETs or a ULN2803 driver board.DoCoGen, Electronics → Statistics: That model is only good for data of $n$ or more.If you have a lower $n$, it's not going to work.For lower $n$ regression use a linear, logistic or a t-test.
Nitay Calderon, Eyal Ben-David, Amir Feder, Roi Reichart
ACL (1)1
2022 A Functional Information Perspective on Model Interpretation
abstract
Contemporary predictive models are hard to interpret as their deep nets exploit numerous complex relations between input elements. This work suggests a theoretical framework for model interpretability by measuring the contribution of relevant features to the functional entropy of the network with respect to the input. We rely on the log-Sobolev inequality that bounds the functional entropy by the functional Fisher information with respect to the covariance of the data. This provides a principled way to measure the amount of information contribution of a subset of features to the decision function. Through extensive experiments, we show that our method surpasses existing interpretability sampling-based methods on various data signals such as image, text, and audio.
Itai Gat, Nitay Calderon, Roi Reichart, Tamir Hazan
ICML2
2021 From Limited Annotated Raw Material Data to Quality Production Data: A Case Study in the Milk Industry
abstract
Industry 4.0 offers opportunities to combine multiple sensor data sources using IoT technologies for better utilization of raw material in production lines. A common belief that data is readily available (the big data phenomenon), is oftentimes challenged by the need to effectively acquire quality data under severe constraints. In this paper we propose a design methodology, using active learning to enhance learning capabilities, for building a model of production outcome using a constrained amount of raw material training data. The proposed methodology extends existing active learning methods to effectively solve regression-based learning problems and may serve settings where data acquisition requires excessive resources in the physical world. We further suggest a set of qualitative measures to analyze learners performance. The proposed methodology is demonstrated using an actual application in the milk industry, where milk is gathered from multiple small milk farms and brought to a dairy production plant to be processed into cottage cheese.
Roee Shraga, Gil Katz, Yael Badian, Nitay Calderon, Avigdor Gal
CIKM4