VLDB 2026 Research / reviewers in the wild / expert
Lilach Eden
dblp:264/4901
· DBLP profile ↗
7ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0001-5244-3813ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 57% Information extraction and text analysis · 43% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › argument mining
key point analysis |
1.4 | 3 | 2021 | Every Bite Is an Experience: Key Point Analysis of Business Reviews · ACL/IJCNLP (1) 2021 Quantitative argument summarization and beyond: Cross-domain key point analysis · EMNLP (1) 2020 From Arguments to Key Points: Towards Automatic Argument Summarization · ACL 2020 |
Natural language and speech › Language models and text generation
large language model evaluation |
1.0 | 1 | 2026 | CLEAR: Error Analysis via LLM-as-a-Judge Made Easy · AAAI 2026 |
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge |
1.0 | 1 | 2026 | CLEAR: Error Analysis via LLM-as-a-Judge Made Easy · AAAI 2026 |
Natural language and speech › Language models and text generation › large language model evaluation
LLM judge |
0.9 | 1 | 2025 | JuStRank: Benchmarking LLM Judges for System Ranking · ACL (1) 2025 |
Natural language and speech › Language models and text generation › text summarization
opinion summarization |
0.8 | 2 | 2023 | From Key Points to Key Point Hierarchy: Structured and Expressive Opinion Summarization · ACL (1) 2023 Every Bite Is an Experience: Key Point Analysis of Business Reviews · ACL/IJCNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis
textual entailment |
0.7 | 1 | 2023 | From Key Points to Key Point Hierarchy: Structured and Expressive Opinion Summarization · ACL (1) 2023 |
Natural language and speech › Information extraction and text analysis
argument mining |
0.4 | 1 | 2020 | Quantitative argument summarization and beyond: Cross-domain key point analysis · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › text summarization
argument summarization |
0.4 | 1 | 2020 | From Arguments to Key Points: Towards Automatic Argument Summarization · ACL 2020 |
Natural language and speech › Language models and text generation › text summarization
multi-document summarization |
0.1 | 1 | 2020 | Quantitative argument summarization and beyond: Cross-domain key point analysis · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation
text summarization |
0.1 | 1 | 2020 | From Arguments to Key Points: Towards Automatic Argument Summarization · ACL 2020 |
Methods — techniques the papers use, named apart from their topics
large language model · 1.0interactive visualization · 1.0system score aggregation · 0.9LLM-based judging · 0.9weak supervision · 0.7directional distributional similarity · 0.7review analysis · 0.5argument mining · 0.5salience scoring · 0.4argument-to-key-point matching · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLEAR: Error Analysis via LLM-as-a-Judge Made EasyabstractThe evaluation of Large Language Models (LLMs) increasingly relies on other LLMs acting as judges. However, current evaluation paradigms typically yield a single score or ranking, answering which model is better but not why. While essential for benchmarking, these top-level scores obscure the specific, actionable reasons behind a model's performance. To bridge this gap, we introduce CLEAR, an interactive, open-source package for LLM-based error analysis. CLEAR first generates per-instance textual feedback, then it creates a set of system-level error issues, and quantifies the prevalence of each identified issue. Our package also provides users with an interactive dashboard that allows for a comprehensive error analysis through aggregate visualizations, applies interactive filters to isolate specific issues or score ranges, and drills down to the individual instances that exemplify a particular behavioral pattern. We demonstrate CLEAR analysis for RAG and Math benchmarks, and showcase its utility through a user case study. Asaf Yehudai, Lilach Eden, Yotam Perlitz, Roy Bar-Haim, Michal Shmueli-Scheuer |
AAAI | 2 |
| 2025 | JuStRank: Benchmarking LLM Judges for System RankingabstractGiven the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available.The scale and versatility of such evaluations make the use of LLMbased judges a compelling solution for this challenge.Crucially, this approach requires first to validate the quality of the LLM judge itself.Previous work has focused on instance-based assessment of LLM judges, where a judge is evaluated over a set of responses, or response pairs, while being agnostic to their source systems.We argue that this setting overlooks critical factors affecting system-level ranking, such as a judge's positive or negative bias towards certain systems.To address this gap, we conduct the first large-scale study of LLM judges as system rankers.System scores are generated by aggregating judgment scores over multiple system outputs, and the judge's quality is assessed by comparing the resulting system ranking to a human-based ranking.Beyond overall judge assessment, our analysis provides a fine-grained characterization of judge behavior, including their decisiveness and bias. Ariel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim, Lilach Eden, Asaf Yehudai |
ACL (1) | 5 |
| 2023 | From Key Points to Key Point Hierarchy: Structured and Expressive Opinion SummarizationabstractKey Point Analysis (KPA) has been recently proposed for deriving fine-grained insights from collections of textual comments.KPA extracts the main points in the data as a list of concise sentences or phrases, termed key points, and quantifies their prevalence.While key points are more expressive than word clouds and key phrases, making sense of a long, flat list of key points, which often express related ideas in varying levels of granularity, may still be challenging.To address this limitation of KPA, we introduce the task of organizing a given set of key points into a hierarchy, according to their specificity.Such hierarchies may be viewed as a novel type of Textual Entailment Graph.We develop THINKP, a high quality benchmark dataset of key point hierarchies for business and product reviews, obtained by consolidating multiple annotations.We compare different methods for predicting pairwise relations between key points, and for inferring a hierarchy from these pairwise predictions.In particular, for the task of computing pairwise key point relations, we achieve significant gains over existing strong baselines by applying directional distributional similarity methods to a novel distributional representation of key points, and further boost performance via weak supervision. Arie Cattan, Lilach Eden, Yoav Kantor, Roy Bar-Haim |
ACL (1) | 2 |
| 2021 | AI-Assisted Security Controls Mapping for Clouds Built for Regulated WorkloadsabstractData privacy, security and compliance concerns prevent many enterprises from migrating their critical applications to public cloud infrastructure. To address this, cloud providers offer specialized clouds for heavily regulated industries, which implement prescribed security standards. A critical step in the migration process is to ensure that the customer's security requirements are fully met by the cloud provider. With a few hundreds of services in a typical cloud provider's infrastructure, this becomes a non-trivial task. Few tens to hundreds of security checks exposed by each applicable service need to be matched with several hundreds to thousands of security controls from the customer. Mapping customer's controls to cloud provider's control set is done manually by experts, a process that often takes months to complete, and needs to be repeated with every new customer. Moreover, these mappings have to be re-evaluated following regulatory or business changes, as well as cloud infrastructure upgrades. We present an AI-assisted system for mapping security controls, which drastically reduces the number of candidates a human expert needs to consider, allowing substantial speed-up of the mapping process. We empirically compare several controls mapping models, and show that hierarchical classification using fine-tuned Transformer networks works best. Overall, our empirical results demonstrate that the system performs well on real-world data. Vikas Agarwal, Roy Bar-Haim, Lilach Eden, Nisha Gupta, Yoav Kantor, Arun Kumar 0002 |
CLOUD | 3 |
| 2021 | Every Bite Is an Experience: Key Point Analysis of Business ReviewsabstractRoy Bar-Haim, Lilach Eden, Yoav Kantor, Roni Friedman, Noam Slonim. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Roy Bar-Haim, Lilach Eden, Yoav Kantor, Roni Friedman, Noam Slonim |
ACL/IJCNLP (1) | 2 |
| 2020 | From Arguments to Key Points: Towards Automatic Argument SummarizationabstractGenerating a concise summary from a large collection of arguments on a given topic is an intriguing yet understudied problem.We propose to represent such summaries as a small set of talking points, termed key points, each scored according to its salience.We show, by analyzing a large dataset of crowd-contributed arguments, that a small number of key points per topic is typically sufficient for covering the vast majority of the arguments.Furthermore, we found that a domain expert can often predict these key points in advance.We study the task of argument-to-key point mapping, and introduce a novel large-scale dataset for this task.We report empirical results for an extensive set of experiments with this dataset, showing promising performance. Roy Bar-Haim, Lilach Eden, Roni Friedman, Yoav Kantor, Dan Lahav, Noam Slonim |
ACL | 2 |
| 2020 | Quantitative argument summarization and beyond: Cross-domain key point analysisabstractWhen summarizing a collection of views, arguments or opinions on some topic, it is often desirable not only to extract the most salient points, but also to quantify their prevalence.Work on multi-document summarization has traditionally focused on creating textual summaries, which lack this quantitative aspect.Recent work has proposed to summarize arguments by mapping them to a small set of expert-generated key points, where the salience of each key point corresponds to the number of its matching arguments.The current work advances key point analysis in two important respects: first, we develop a method for automatic extraction of key points, which enables fully automatic analysis, and is shown to achieve performance comparable to a human expert.Second, we demonstrate that the applicability of key point analysis goes well beyond argumentation data.Using models trained on publicly available argumentation datasets, we achieve promising results in two additional domains: municipal surveys and user reviews.An additional contribution is an in-depth evaluation of argument-to-key point matching models, where we substantially outperform previous results. Roy Bar-Haim, Yoav Kantor, Lilach Eden, Roni Friedman, Dan Lahav, Noam Slonim |
EMNLP (1) | 3 |