Asaf Yehudai

dblp:301/9362 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-3743-7551ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
abstract
The evaluation of Large Language Models (LLMs) increasingly relies on other LLMs acting as judges. However, current evaluation paradigms typically yield a single score or ranking, answering which model is better but not why. While essential for benchmarking, these top-level scores obscure the specific, actionable reasons behind a model's performance. To bridge this gap, we introduce CLEAR, an interactive, open-source package for LLM-based error analysis. CLEAR first generates per-instance textual feedback, then it creates a set of system-level error issues, and quantifies the prevalence of each identified issue. Our package also provides users with an interactive dashboard that allows for a comprehensive error analysis through aggregate visualizations, applies interactive filters to isolate specific issues or score ranges, and drills down to the individual instances that exemplify a particular behavioral pattern. We demonstrate CLEAR analysis for RAG and Math benchmarks, and showcase its utility through a user case study.
Asaf Yehudai, Lilach Eden, Yotam Perlitz, Roy Bar-Haim, Michal Shmueli-Scheuer
AAAI1
2026 Mediocrity is the key for LLM as a Judge Anchor Selection
abstract
The "LLM-as-a-judge" paradigm has become a standard method for evaluating open-ended generation.To address the quadratic scalability costs of pairwise comparisons, popular benchmarks like Arena-Hard and AlpacaEval compare all models against a single anchor.However, despite its widespread use, the impact of anchor selection on the reliability of the results remains largely unexplored.In this work, we systematically investigate the effect of anchor selection by evaluating 22 different anchors on the Arena-Hard-v2.0 dataset.We find that the choice of anchor is critical: a poor anchor can dramatically reduce correlation with human rankings.We identify that common anchor choices (best-performing and worst-performing models) make poor anchors.Because these extreme anchors are consistently better or worse than all other models, they are seldom indicative of the relative ranking of the models.We further quantify the effect size of anchor selection, showing it is comparable to the selection of a judge model.We conclude with actionable recommendations.First, we conduct a power analysis, and compute sufficient benchmark sizes for anchor-based evaluation, finding that standard benchmark sizes are insufficient for pairwise evaluation and fail to distinguish between competitive models reliably.Second, we provide guidelines for selecting informative anchors to ensure reliable and efficient evaluation practices.
Shachar Don-Yehiya, Asaf Yehudai, Leshem Choshen, Omri Abend
ACL (1)2
2025 JuStRank: Benchmarking LLM Judges for System Ranking
abstract
Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available.The scale and versatility of such evaluations make the use of LLMbased judges a compelling solution for this challenge.Crucially, this approach requires first to validate the quality of the LLM judge itself.Previous work has focused on instance-based assessment of LLM judges, where a judge is evaluated over a set of responses, or response pairs, while being agnostic to their source systems.We argue that this setting overlooks critical factors affecting system-level ranking, such as a judge's positive or negative bias towards certain systems.To address this gap, we conduct the first large-scale study of LLM judges as system rankers.System scores are generated by aggregating judgment scores over multiple system outputs, and the judge's quality is assessed by comparing the resulting system ranking to a human-based ranking.Beyond overall judge assessment, our analysis provides a fine-grained characterization of judge behavior, including their decisiveness and bias.
Ariel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim, Lilach Eden, Asaf Yehudai
ACL (1)6
2024 A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns
Asaf Yehudai, Taelin Karidi, Gabriel Stanovsky, Ariel Goldstein, Omri Abend
CogSci1
2024 Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine Translation
abstract
Most works on gender bias focus on intrinsic bias -removing traces of information about a protected group from the model's internal representation.However, these works are often disconnected from the impact of such debiasing on downstream applications, which is the main motivation for debiasing in the first place.In this work, we systematically test how methods for intrinsic debiasing affect neural machine translation models, by measuring the extrinsic bias of such systems under different design choices.We highlight three challenges and mismatches between the debiasing techniques and their end-goal usage, including the choice of embeddings to debias, the mismatch between words and sub-word tokens debiasing, and the effect of translating from English to different target languages.We find that these considerations have a significant impact on downstream performance and the success of debiasing.1
Bar Iluz, Yanai Elazar, Asaf Yehudai, Gabriel Stanovsky
EMNLP3
2024 Achieving Human Parity in Content-Grounded Datasets Generation
abstract
The lack of high-quality data for content-grounded generation tasks has been identified as a major obstacle to advancing these tasks. To address this gap, we propose Genie, a novel method for automatically generating high-quality content-grounded data. It consists of three stages: (a) Content Preparation, (b) Generation: creating task-specific examples from the content (e.g., question-answer pairs or summaries). (c) Filtering mechanism aiming to ensure the quality and faithfulness of the generated data. We showcase this methodology by generating three large-scale synthetic data, making wishes, for Long-Form Question-Answering (LFQA), summarization, and information extraction. In a human evaluation, our generated data was found to be natural and of high quality. Furthermore, we compare models trained on our data with models trained on human-written data -- ELI5 and ASQA for LFQA and CNN-DailyMail for Summarization. We show that our models are on par with or outperforming models trained on human-generated data and consistently outperforming them in faithfulness. Finally, we applied our method to create LFQA data within the medical domain and compared a model trained on it with models trained on other domains.
Asaf Yehudai, Boaz Carmeli, Yosi Mass, Ofir Arviv, Nathaniel Mills, Eyal Shnarch, Leshem Choshen
ICLR1
2023 Evaluating and Improving the Coreference Capabilities of Machine Translation Models
abstract
Machine translation (MT) requires a wide range of linguistic capabilities, which current end-to-end models are expected to learn implicitly by observing aligned sentences in bilingual corpora.In this work, we ask: How well do MT models learn coreference resolution from implicit signal?To answer this question, we develop an evaluation methodology that derives coreference clusters from MT output and evaluates them without requiring annotations in the target language.We further evaluate several prominent open-source and commercial MT systems, translating from English to six target languages, and compare them to state-of-theart coreference resolvers on three challenging benchmarks.Our results show that the monolingual resolvers greatly outperform MT models.Motivated by this result, we experiment with different methods for incorporating the output of coreference resolution models in MT, showing improvement over strong baselines.1
Asaf Yehudai, Arie Cattan, Omri Abend, Gabriel Stanovsky
EACL1
2023 QAID: Question Answering Inspired Few-shot Intent Detection
Asaf Yehudai, Matan Vetzler, Yosi Mass, Koren Lazar, Doron Cohen 0001, Boaz Carmeli
ICLR1
2022 Reinforcement Learning with Large Action Spaces for Neural Machine Translation
abstract
Applying Reinforcement learning (RL) following maximum likelihood estimation (MLE) pre-training is a versatile method for enhancing neural machine translation (NMT) performance. However, recent work has argued that the gains produced by RL for NMT are mostly due to promoting tokens that have already received a fairly high probability in pre-training. We hypothesize that the large action space is a main obstacle to RL’s effectiveness in MT, and conduct two sets of experiments that lend support to our hypothesis. First, we find that reducing the size of the vocabulary improves RL’s effectiveness. Second, we find that effectively reducing the dimension of the action space without changing the vocabulary also yields notable improvement as evaluated by BLEU, semantic similarity, and human evaluation. Indeed, by initializing the network’s final fully connected layer (that maps the network’s internal dimension to the vocabulary dimension), with a layer that generalizes over similar actions, we obtain a substantial improvement in RL performance: 1.5 BLEU points on average.
Asaf Yehudai, Leshem Choshen, Lior Fox, Omri Abend
COLING1
2021 Filling the Gaps in Ancient Akkadian Texts: A Masked Language Modelling Approach
abstract
We present models which complete missing text given transliterations of ancient Mesopotamian documents, originally written on cuneiform clay tablets (2500 BCE -100 CE).Due to the tablets' deterioration, scholars often rely on contextual cues to manually fill in missing parts in the text in a subjective and time-consuming process.We identify that this challenge can be formulated as a masked language modelling task, used mostly as a pretraining objective for contextualized language models.Following, we develop several architectures focusing on the Akkadian language, the lingua franca of the time.We find that despite data scarcity (1M tokens) we can achieve state of the art performance on missing tokens prediction (89% hit@5) using a greedy decoding scheme and pretraining on data from other languages and different time periods.Finally, we conduct human evaluations showing the applicability of our models in assisting experts to transcribe texts in extinct languages.
Koren Lazar, Benny Saret, Asaf Yehudai, Wayne Horowitz, Nathan Wasserman, Gabriel Stanovsky
EMNLP (1)3