Roy Bar-Haim

dblp:99/5426 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
6since 2021 · last 2026
0000-0001-7365-1662ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 10 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Language models and text generation · 60% Information extraction and text analysis · 38% Knowledge representation and reasoning · 2%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational finance and economics · 100%

Topics — the 14 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
1.922026
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy · AAAI 2026
Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation · EMNLP 2025
Natural language and speech › Language models and text generation › large language model evaluation
LLM judge
1.722025
Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation · EMNLP 2025
JuStRank: Benchmarking LLM Judges for System Ranking · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › argument mining
key point analysis
1.432021
Every Bite Is an Experience: Key Point Analysis of Business Reviews · ACL/IJCNLP (1) 2021
Quantitative argument summarization and beyond: Cross-domain key point analysis · EMNLP (1) 2020
From Arguments to Key Points: Towards Automatic Argument Summarization · ACL 2020
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge
1.012026
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy · AAAI 2026
Natural language and speech › Language models and text generation › text summarization
opinion summarization
0.822023
From Key Points to Key Point Hierarchy: Structured and Expressive Opinion Summarization · ACL (1) 2023
Every Bite Is an Experience: Key Point Analysis of Business Reviews · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis
textual entailment
0.712023
From Key Points to Key Point Hierarchy: Structured and Expressive Opinion Summarization · ACL (1) 2023
Natural language and speech › Information extraction and text analysis
argument mining
0.522020
Quantitative argument summarization and beyond: Cross-domain key point analysis · EMNLP (1) 2020
From Surrogacy to Adoption; From Bitcoin to Cryptocurrency: Debate Topic Expansion · ACL (1) 2019
Natural language and speech › Language models and text generation › text summarization
argument summarization
0.412020
From Arguments to Key Points: Towards Automatic Argument Summarization · ACL 2020
Natural language and speech › Language models and text generation › text summarization
multi-document summarization
0.112020
Quantitative argument summarization and beyond: Cross-domain key point analysis · EMNLP (1) 2020
Natural language and speech › Language models and text generation
text summarization
0.112020
From Arguments to Key Points: Towards Automatic Argument Summarization · ACL 2020
Natural language and speech › Information extraction and text analysis
social media text analysis
0.112011
Identifying and Following Expert Investors in Stock Microblogs · EMNLP 2011
Knowledge, reasoning and agents › Knowledge representation and reasoning › nonmonotonic reasoning › preference handling
preference reasoning
0.112008
Contextual Preferences · ACL 2008
Knowledge, reasoning and agents › Knowledge representation and reasoning
semantic reasoning
0.112007
Semantic Inference at the Lexical-Syntactic Level · AAAI 2007
Mathematical optimization
scalable inference
0.012009
A Compact Forest for Scalable Inference over Entailment and Paraphrase Rules · EMNLP 2009

Methods — techniques the papers use, named apart from their topics

large language model · 1.0interactive visualization · 1.0system score aggregation · 0.9human annotation · 0.9benchmarking · 0.9LLM-based judging · 0.9weak supervision · 0.7directional distributional similarity · 0.7review analysis · 0.5argument mining · 0.5user ranking · 0.1text classification · 0.1rule matching · 0.1forest-based inference · 0.1
YearPublicationVenuePosition
2026 CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
abstract
The evaluation of Large Language Models (LLMs) increasingly relies on other LLMs acting as judges. However, current evaluation paradigms typically yield a single score or ranking, answering which model is better but not why. While essential for benchmarking, these top-level scores obscure the specific, actionable reasons behind a model's performance. To bridge this gap, we introduce CLEAR, an interactive, open-source package for LLM-based error analysis. CLEAR first generates per-instance textual feedback, then it creates a set of system-level error issues, and quantifies the prevalence of each identified issue. Our package also provides users with an interactive dashboard that allows for a comprehensive error analysis through aggregate visualizations, applies interactive filters to isolate specific issues or score ranges, and drills down to the individual instances that exemplify a particular behavioral pattern. We demonstrate CLEAR analysis for RAG and Math benchmarks, and showcase its utility through a user case study.
Asaf Yehudai, Lilach Eden, Yotam Perlitz, Roy Bar-Haim, Michal Shmueli-Scheuer
AAAI4
2025 JuStRank: Benchmarking LLM Judges for System Ranking
abstract
Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available.The scale and versatility of such evaluations make the use of LLMbased judges a compelling solution for this challenge.Crucially, this approach requires first to validate the quality of the LLM judge itself.Previous work has focused on instance-based assessment of LLM judges, where a judge is evaluated over a set of responses, or response pairs, while being agnostic to their source systems.We argue that this setting overlooks critical factors affecting system-level ranking, such as a judge's positive or negative bias towards certain systems.To address this gap, we conduct the first large-scale study of LLM judges as system rankers.System scores are generated by aggregating judgment scores over multiple system outputs, and the judge's quality is assessed by comparing the resulting system ranking to a human-based ranking.Beyond overall judge assessment, our analysis provides a fine-grained characterization of judge behavior, including their decisiveness and bias.
Ariel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim, Lilach Eden, Asaf Yehudai
ACL (1)4
2025 Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation
abstract
We introduce Debate Speech Evaluation as a novel and challenging benchmark for assessing LLM judges.Evaluating debate speeches requires a deep understanding of the speech at multiple levels, including argument strength and relevance, the coherence and organization of the speech, the appropriateness of its style and tone, and so on.This task involves a unique set of cognitive abilities that previously received limited attention in systematic LLM benchmarking.To explore such skills, we leverage a dataset of over 600 meticulously annotated debate speeches and present the first in-depth analysis of how state-of-the-art LLMs compare to human judges on this task.Our findings reveal a nuanced picture: while larger models can approximate individual human judgments in some respects, they differ substantially in their overall judgment behavior.We also investigate the ability of frontier LLMs to generate persuasive, opinionated speeches, showing that models may perform at a human level on this task.
Noy Sternlicht, Ariel Gera, Roy Bar-Haim, Tom Hope, Noam Slonim
EMNLP3
2023 From Key Points to Key Point Hierarchy: Structured and Expressive Opinion Summarization
abstract
Key Point Analysis (KPA) has been recently proposed for deriving fine-grained insights from collections of textual comments.KPA extracts the main points in the data as a list of concise sentences or phrases, termed key points, and quantifies their prevalence.While key points are more expressive than word clouds and key phrases, making sense of a long, flat list of key points, which often express related ideas in varying levels of granularity, may still be challenging.To address this limitation of KPA, we introduce the task of organizing a given set of key points into a hierarchy, according to their specificity.Such hierarchies may be viewed as a novel type of Textual Entailment Graph.We develop THINKP, a high quality benchmark dataset of key point hierarchies for business and product reviews, obtained by consolidating multiple annotations.We compare different methods for predicting pairwise relations between key points, and for inferring a hierarchy from these pairwise predictions.In particular, for the task of computing pairwise key point relations, we achieve significant gains over existing strong baselines by applying directional distributional similarity methods to a novel distributional representation of key points, and further boost performance via weak supervision.
Arie Cattan, Lilach Eden, Yoav Kantor, Roy Bar-Haim
ACL (1)4
2021 AI-Assisted Security Controls Mapping for Clouds Built for Regulated Workloads
abstract
Data privacy, security and compliance concerns prevent many enterprises from migrating their critical applications to public cloud infrastructure. To address this, cloud providers offer specialized clouds for heavily regulated industries, which implement prescribed security standards. A critical step in the migration process is to ensure that the customer's security requirements are fully met by the cloud provider. With a few hundreds of services in a typical cloud provider's infrastructure, this becomes a non-trivial task. Few tens to hundreds of security checks exposed by each applicable service need to be matched with several hundreds to thousands of security controls from the customer. Mapping customer's controls to cloud provider's control set is done manually by experts, a process that often takes months to complete, and needs to be repeated with every new customer. Moreover, these mappings have to be re-evaluated following regulatory or business changes, as well as cloud infrastructure upgrades. We present an AI-assisted system for mapping security controls, which drastically reduces the number of candidates a human expert needs to consider, allowing substantial speed-up of the mapping process. We empirically compare several controls mapping models, and show that hierarchical classification using fine-tuned Transformer networks works best. Overall, our empirical results demonstrate that the system performs well on real-world data.
Vikas Agarwal, Roy Bar-Haim, Lilach Eden, Nisha Gupta, Yoav Kantor, Arun Kumar 0002
CLOUD2
2021 Every Bite Is an Experience: Key Point Analysis of Business Reviews
abstract
Roy Bar-Haim, Lilach Eden, Yoav Kantor, Roni Friedman, Noam Slonim. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Roy Bar-Haim, Lilach Eden, Yoav Kantor, Roni Friedman, Noam Slonim
ACL/IJCNLP (1)1
2020 From Arguments to Key Points: Towards Automatic Argument Summarization
abstract
Generating a concise summary from a large collection of arguments on a given topic is an intriguing yet understudied problem.We propose to represent such summaries as a small set of talking points, termed key points, each scored according to its salience.We show, by analyzing a large dataset of crowd-contributed arguments, that a small number of key points per topic is typically sufficient for covering the vast majority of the arguments.Furthermore, we found that a domain expert can often predict these key points in advance.We study the task of argument-to-key point mapping, and introduce a novel large-scale dataset for this task.We report empirical results for an extensive set of experiments with this dataset, showing promising performance.
Roy Bar-Haim, Lilach Eden, Roni Friedman, Yoav Kantor, Dan Lahav, Noam Slonim
ACL1
2020 Quantitative argument summarization and beyond: Cross-domain key point analysis
abstract
When summarizing a collection of views, arguments or opinions on some topic, it is often desirable not only to extract the most salient points, but also to quantify their prevalence.Work on multi-document summarization has traditionally focused on creating textual summaries, which lack this quantitative aspect.Recent work has proposed to summarize arguments by mapping them to a small set of expert-generated key points, where the salience of each key point corresponds to the number of its matching arguments.The current work advances key point analysis in two important respects: first, we develop a method for automatic extraction of key points, which enables fully automatic analysis, and is shown to achieve performance comparable to a human expert.Second, we demonstrate that the applicability of key point analysis goes well beyond argumentation data.Using models trained on publicly available argumentation datasets, we achieve promising results in two additional domains: municipal surveys and user reviews.An additional contribution is an in-depth evaluation of argument-to-key point matching models, where we substantially outperform previous results.
Roy Bar-Haim, Yoav Kantor, Lilach Eden, Roni Friedman, Dan Lahav, Noam Slonim
EMNLP (1)1
2019 From Surrogacy to Adoption; From Bitcoin to Cryptocurrency: Debate Topic Expansion
abstract
Roy Bar-Haim, Dalia Krieger, Orith Toledo-Ronen, Lilach Edelstein, Yonatan Bilu, Alon Halfon, Yoav Katz, Amir Menczel, Ranit Aharonov, Noam Slonim. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Roy Bar-Haim, Dalia Krieger, Orith Toledo-Ronen, Lilach Edelstein, Yonatan Bilu, Alon Halfon, Yoav Katz, Amir Menczel, Ranit Aharonov, Noam Slonim
ACL (1)1
2018 Learning Sentiment Composition from Sentiment Lexicons
abstract
Sentiment composition is a fundamental sentiment analysis problem. Previous work relied on manual rules and manually-created lexical resources such as negator lists, or learned a composition function from sentiment-annotated phrases or sentences. We propose a new approach for learning sentiment composition from a large, unlabeled corpus, which only requires a word-level sentiment lexicon for supervision. We automatically generate large sentiment lexicons of bigrams and unigrams, from which we induce a set of lexicons for a variety of sentiment composition processes. The effectiveness of our approach is confirmed through manual annotation, as well as sentiment classification experiments with both phrase-level and sentence-level benchmarks.
Orith Toledo-Ronen, Roy Bar-Haim, Alon Halfon, Charles Jochim, Amir Menczel, Ranit Aharonov, Noam Slonim
COLING2
2018 SLIDE - a Sentiment Lexicon of Common Idioms
Charles Jochim, Francesca Bonin, Roy Bar-Haim, Noam Slonim
LREC3
2017 Stance Classification of Context-Dependent Claims
abstract
Roy Bar-Haim, Indrajit Bhattacharya, Francesco Dinuzzo, Amrita Saha, Noam Slonim. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
Roy Bar-Haim, Indrajit Bhattacharya, Francesco Dinuzzo, Amrita Saha, Noam Slonim
EACL (1)1
2015 Knowledge-Based Textual Inference via Parse-Tree Transformations
abstract
Textual inference is an important component in many applications for understanding natural language. Classical approaches to textual inference rely on logical representations for meaning, which may be regarded as "external" to the natural language itself. However, practical applications usually adopt shallower lexical or lexical-syntactic representations, which correspond closely to language structure. In many cases, such approaches lack a principled meaning representation and inference framework. We describe an inference formalism that operates directly on language-based structures, particularly syntactic parse trees. New trees are generated by applying inference rules, which provide a unified representation for varying types of inferences. We use manual and automatic methods to generate these rules, which cover generic linguistic structures as well as specific lexical-based inferences. We also present a novel packed data-structure and a corresponding inference algorithm that allows efficient implementation of this formalism. We proved the correctness of the new algorithm and established its efficiency analytically and empirically. The utility of our approach was illustrated on two tasks: unsupervised relation extraction from a large corpus, and the Recognizing Textual Entailment (RTE) benchmarks.
Roy Bar-Haim, Ido Dagan, Jonathan Berant
J. Artif. Intell. Res.1
2011 Identifying and Following Expert Investors in Stock Microblogs
Roy Bar-Haim, Elad Dinur, Ronen Feldman, Moshe Fresko, Guy Goldstein
EMNLP1
2011 The Stock Sonar - Sentiment Analysis of Stocks Based on a Hybrid Approach
abstract
The Stock Sonar (TSS) is a stock sentiment analysis application based on a novel hybrid approach. While previous work focused on document level sentiment classification, or extracted only generic sentiment at the phrase level, TSS integrates sentiment dictionaries, phrase-level compositional patterns, and predicate-level semantic events. TSS generates precise in-text sentiment tagging as well as sentiment-oriented event summaries for a given stock, which are also aggregated into sentiment scores. Hence, TSS allows investors to get the essence of thousands of articles every day and may help them to make timely, informed trading decisions. The extracted sentiment is also shown to improve the accu- racy of an existing document-level sentiment classifier.
Ronen Feldman, Benjamin Rozenfeld, Roy Bar-Haim, Moshe Fresko
IAAI3
2009 A Compact Forest for Scalable Inference over Entailment and Paraphrase Rules
Roy Bar-Haim, Jonathan Berant, Ido Dagan
EMNLP1
2008 Contextual Preferences
Idan Szpektor, Ido Dagan, Roy Bar-Haim, Jacob Goldberger
ACL3
2008 Natural Language as the Basis for Meaning Representation and Inference
Ido Dagan, Roy Bar-Haim, Idan Szpektor, Iddo Greental, Eyal Shnarch
CICLing2
2008 Part-of-speech tagging of Modern Hebrew text
abstract
Abstract Words in Semitic texts often consist of a concatenation of word segments , each corresponding to a part-of-speech (POS) category. Semitic words may be ambiguous with regard to their segmentation as well as to the POS tags assigned to each segment. When designing POS taggers for Semitic languages, a major architectural decision concerns the choice of the atomic input tokens (terminal symbols). If the tokenization is at the word level, the output tags must be complex, and represent both the segmentation of the word and the POS tag assigned to each word segment. If the tokenization is at the segment level, the input itself must encode the different alternative segmentations of the words, while the output consists of standard POS tags. Comparing these two alternatives is not trivial, as the choice between them may have global effects on the grammatical model. Moreover, intermediate levels of tokenization between these two extremes are conceivable, and, as we aim to show, beneficial. To the best of our knowledge, the problem of tokenization for POS tagging of Semitic languages has not been addressed before in full generality. In this paper, we study this problem for the purpose of POS tagging of Modern Hebrew texts. After extensive error analysis of the two simple tokenization models, we propose a novel, linguistically motivated, intermediate tokenization model that gives better performance for Hebrew over the two initial architectures. Our study is based on the well-known hidden Markov models (HMMs). We start out from a manually devised morphological analyzer and a very small annotated corpus, and describe how to adapt an HMM-based POS tagger for both tokenization architectures. We present an effective technique for smoothing the lexical probabilities using an untagged corpus, and a novel transformation for casting the segment-level tagger in terms of a standard, word-level HMM implementation. The results obtained using our model are on par with the best published results on Modern Standard Arabic, despite the much smaller annotated corpus available for Modern Hebrew.
Roy Bar-Haim, Khalil Sima'an, Yoad Winter
Nat. Lang. Eng.1
2007 Semantic Inference at the Lexical-Syntactic Level
Roy Bar-Haim, Ido Dagan, Iddo Greental, Eyal Shnarch
AAAI1