Lilach Eden

dblp:264/4901 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0001-5244-3813ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 57% Information extraction and text analysis · 43%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › argument mining
key point analysis
1.432021
Every Bite Is an Experience: Key Point Analysis of Business Reviews · ACL/IJCNLP (1) 2021
Quantitative argument summarization and beyond: Cross-domain key point analysis · EMNLP (1) 2020
From Arguments to Key Points: Towards Automatic Argument Summarization · ACL 2020
Natural language and speech › Language models and text generation
large language model evaluation
1.012026
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy · AAAI 2026
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge
1.012026
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy · AAAI 2026
Natural language and speech › Language models and text generation › large language model evaluation
LLM judge
0.912025
JuStRank: Benchmarking LLM Judges for System Ranking · ACL (1) 2025
Natural language and speech › Language models and text generation › text summarization
opinion summarization
0.822023
From Key Points to Key Point Hierarchy: Structured and Expressive Opinion Summarization · ACL (1) 2023
Every Bite Is an Experience: Key Point Analysis of Business Reviews · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis
textual entailment
0.712023
From Key Points to Key Point Hierarchy: Structured and Expressive Opinion Summarization · ACL (1) 2023
Natural language and speech › Information extraction and text analysis
argument mining
0.412020
Quantitative argument summarization and beyond: Cross-domain key point analysis · EMNLP (1) 2020
Natural language and speech › Language models and text generation › text summarization
argument summarization
0.412020
From Arguments to Key Points: Towards Automatic Argument Summarization · ACL 2020
Natural language and speech › Language models and text generation › text summarization
multi-document summarization
0.112020
Quantitative argument summarization and beyond: Cross-domain key point analysis · EMNLP (1) 2020
Natural language and speech › Language models and text generation
text summarization
0.112020
From Arguments to Key Points: Towards Automatic Argument Summarization · ACL 2020

Methods — techniques the papers use, named apart from their topics

large language model · 1.0interactive visualization · 1.0system score aggregation · 0.9LLM-based judging · 0.9weak supervision · 0.7directional distributional similarity · 0.7review analysis · 0.5argument mining · 0.5salience scoring · 0.4argument-to-key-point matching · 0.4
YearPublicationVenuePosition
2026 CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
abstract
The evaluation of Large Language Models (LLMs) increasingly relies on other LLMs acting as judges. However, current evaluation paradigms typically yield a single score or ranking, answering which model is better but not why. While essential for benchmarking, these top-level scores obscure the specific, actionable reasons behind a model's performance. To bridge this gap, we introduce CLEAR, an interactive, open-source package for LLM-based error analysis. CLEAR first generates per-instance textual feedback, then it creates a set of system-level error issues, and quantifies the prevalence of each identified issue. Our package also provides users with an interactive dashboard that allows for a comprehensive error analysis through aggregate visualizations, applies interactive filters to isolate specific issues or score ranges, and drills down to the individual instances that exemplify a particular behavioral pattern. We demonstrate CLEAR analysis for RAG and Math benchmarks, and showcase its utility through a user case study.
Asaf Yehudai, Lilach Eden, Yotam Perlitz, Roy Bar-Haim, Michal Shmueli-Scheuer
AAAI2
2025 JuStRank: Benchmarking LLM Judges for System Ranking
abstract
Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available.The scale and versatility of such evaluations make the use of LLMbased judges a compelling solution for this challenge.Crucially, this approach requires first to validate the quality of the LLM judge itself.Previous work has focused on instance-based assessment of LLM judges, where a judge is evaluated over a set of responses, or response pairs, while being agnostic to their source systems.We argue that this setting overlooks critical factors affecting system-level ranking, such as a judge's positive or negative bias towards certain systems.To address this gap, we conduct the first large-scale study of LLM judges as system rankers.System scores are generated by aggregating judgment scores over multiple system outputs, and the judge's quality is assessed by comparing the resulting system ranking to a human-based ranking.Beyond overall judge assessment, our analysis provides a fine-grained characterization of judge behavior, including their decisiveness and bias.
Ariel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim, Lilach Eden, Asaf Yehudai
ACL (1)5
2023 From Key Points to Key Point Hierarchy: Structured and Expressive Opinion Summarization
abstract
Key Point Analysis (KPA) has been recently proposed for deriving fine-grained insights from collections of textual comments.KPA extracts the main points in the data as a list of concise sentences or phrases, termed key points, and quantifies their prevalence.While key points are more expressive than word clouds and key phrases, making sense of a long, flat list of key points, which often express related ideas in varying levels of granularity, may still be challenging.To address this limitation of KPA, we introduce the task of organizing a given set of key points into a hierarchy, according to their specificity.Such hierarchies may be viewed as a novel type of Textual Entailment Graph.We develop THINKP, a high quality benchmark dataset of key point hierarchies for business and product reviews, obtained by consolidating multiple annotations.We compare different methods for predicting pairwise relations between key points, and for inferring a hierarchy from these pairwise predictions.In particular, for the task of computing pairwise key point relations, we achieve significant gains over existing strong baselines by applying directional distributional similarity methods to a novel distributional representation of key points, and further boost performance via weak supervision.
Arie Cattan, Lilach Eden, Yoav Kantor, Roy Bar-Haim
ACL (1)2
2021 AI-Assisted Security Controls Mapping for Clouds Built for Regulated Workloads
abstract
Data privacy, security and compliance concerns prevent many enterprises from migrating their critical applications to public cloud infrastructure. To address this, cloud providers offer specialized clouds for heavily regulated industries, which implement prescribed security standards. A critical step in the migration process is to ensure that the customer's security requirements are fully met by the cloud provider. With a few hundreds of services in a typical cloud provider's infrastructure, this becomes a non-trivial task. Few tens to hundreds of security checks exposed by each applicable service need to be matched with several hundreds to thousands of security controls from the customer. Mapping customer's controls to cloud provider's control set is done manually by experts, a process that often takes months to complete, and needs to be repeated with every new customer. Moreover, these mappings have to be re-evaluated following regulatory or business changes, as well as cloud infrastructure upgrades. We present an AI-assisted system for mapping security controls, which drastically reduces the number of candidates a human expert needs to consider, allowing substantial speed-up of the mapping process. We empirically compare several controls mapping models, and show that hierarchical classification using fine-tuned Transformer networks works best. Overall, our empirical results demonstrate that the system performs well on real-world data.
Vikas Agarwal, Roy Bar-Haim, Lilach Eden, Nisha Gupta, Yoav Kantor, Arun Kumar 0002
CLOUD3
2021 Every Bite Is an Experience: Key Point Analysis of Business Reviews
abstract
Roy Bar-Haim, Lilach Eden, Yoav Kantor, Roni Friedman, Noam Slonim. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Roy Bar-Haim, Lilach Eden, Yoav Kantor, Roni Friedman, Noam Slonim
ACL/IJCNLP (1)2
2020 From Arguments to Key Points: Towards Automatic Argument Summarization
abstract
Generating a concise summary from a large collection of arguments on a given topic is an intriguing yet understudied problem.We propose to represent such summaries as a small set of talking points, termed key points, each scored according to its salience.We show, by analyzing a large dataset of crowd-contributed arguments, that a small number of key points per topic is typically sufficient for covering the vast majority of the arguments.Furthermore, we found that a domain expert can often predict these key points in advance.We study the task of argument-to-key point mapping, and introduce a novel large-scale dataset for this task.We report empirical results for an extensive set of experiments with this dataset, showing promising performance.
Roy Bar-Haim, Lilach Eden, Roni Friedman, Yoav Kantor, Dan Lahav, Noam Slonim
ACL2
2020 Quantitative argument summarization and beyond: Cross-domain key point analysis
abstract
When summarizing a collection of views, arguments or opinions on some topic, it is often desirable not only to extract the most salient points, but also to quantify their prevalence.Work on multi-document summarization has traditionally focused on creating textual summaries, which lack this quantitative aspect.Recent work has proposed to summarize arguments by mapping them to a small set of expert-generated key points, where the salience of each key point corresponds to the number of its matching arguments.The current work advances key point analysis in two important respects: first, we develop a method for automatic extraction of key points, which enables fully automatic analysis, and is shown to achieve performance comparable to a human expert.Second, we demonstrate that the applicability of key point analysis goes well beyond argumentation data.Using models trained on publicly available argumentation datasets, we achieve promising results in two additional domains: municipal surveys and user reviews.An additional contribution is an in-depth evaluation of argument-to-key point matching models, where we substantially outperform previous results.
Roy Bar-Haim, Yoav Kantor, Lilach Eden, Roni Friedman, Dan Lahav, Noam Slonim
EMNLP (1)3