VLDB 2026 Research / reviewers in the wild / expert
Kfir Bar
dblp:23/8608
· DBLP profile ↗
13ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-1354-2955ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automatic Segmentation of Classical Tibetan Texts into Autochthonous and Allochthonous Regions
Guy Bilitski, Lev Shechter, Sonam Jamtsho, Nir Marciano, Nicola Bajetta, Rebecca Sundén, Omri Drori, Kai Golan Hashiloni, Orr Meir Zwebner, Asaf Shina, Orna Almogi, Dorji Wangchuk, Kfir Bar |
LREC | 13 |
| 2025 | Beyond Pairwise: Global Zero-shot Temporal Graph GenerationabstractTemporal relation extraction (TRE) is a fundamental task in natural language processing (NLP) that involves identifying the temporal relationships between events in a document.Despite the advances in large language models (LLMs), their application to TRE remains limited.Most existing approaches rely on pairwise classification, where event pairs are classified in isolation, leading to computational inefficiency and a lack of global consistency in the resulting temporal graph.In this work, we propose a novel zero-shot method for TRE that generates a document's complete temporal graph in a single step, followed by temporal constraint optimization to refine predictions and enforce temporal consistency across relations.Additionally, we introduce OmniTemp, a new dataset with complete annotations for all pairs of targeted events within a document.Through experiments and analyses, we demonstrate that our method outperforms existing zero-shot approaches and offers a competitive alternative to supervised TRE models. Alon Eirew, Kfir Bar, Ido Dagan |
EMNLP | 2 |
| 2025 | Easy as PIE? Identifying Multi-Word Expressions with LLMsabstractWe investigate the identification of idiomatic expressions-a semantically noncompositional subclass of multiword expressions (MWEs)-in running text using large language models (LLMs) without any fine-tuning.Instead, we adopt a prompt-based approach and evaluate a range of prompting strategies, including zero-shot, few-shot, and chain-of-thought variants, across multiple languages, datasets, and model types.Our experiments show that, with well-crafted prompts, LLMs can perform competitively with supervised models trained on annotated data.These findings highlight the potential of prompt-based LLMs as a flexible and effective alternative for idiomatic expression identification. Kai Golan Hashiloni, Ofri Hefetz, Kfir Bar |
EMNLP | 3 |
| 2025 | Can LLMs Help Encoder Models Maintain Both High Accuracy and Consistency in Temporal Relation Classification?abstractTemporal relation classification (TRC) demands both accuracy and temporal consistency in event timeline extraction. Encoder-based models achieve high accuracy but introduce inconsistencies because they rely on pairwise classification, while LLMs leverage global context to generate temporal graphs, improving consistency at the cost of accuracy. We assess LLM prompting strategies for TRC and their effectiveness in assisting encoder models with cycle resolution. Results show that while LLMs improve consistency, they struggle with accuracy and do not outperform a simple confidence-based cycle resolution approach. Our code is publicly available at: https://github.com/MatufA/timeline-extraction. Adiel Meir, Kfir Bar |
INLG | 2 |
| 2024 | JRC-Names-Retrieval: A Standardized Benchmark for Name SearchabstractMany systems rely on the ability to effectively search through databases of personal and organization entity names in multiple writing scripts. Despite this, there is a relative lack of research studying this problem in isolation. In this work, we discuss this problem in detail and support future research by publishing what we believe is the first comprehensive dataset designed for this task. Additionally, we present a number of baselines against which future work can be compared; among which, we describe a neural solution based on ByT5 (Xue et al. 2022) which demonstrates up to a 12% performance gain over preexisting baselines, indicating that there remains much room for improvement in this space. Philip Blair, Kfir Bar |
LREC/COLING | 2 |
| 2024 | Motivational Interviewing Transcripts Annotated with Global ScoresabstractMotivational interviewing (MI) is a counseling approach that aims to increase intrinsic motivation and commitment to change. Despite its effectiveness in various disorders such as addiction, weight loss, and smoking cessation, publicly available annotated MI datasets are scarce, limiting the development and evaluation of MI language generation models. We present MI-TAGS, a new annotated dataset of MI therapy sessions written in English collected from video recordings available on public sources. The dataset includes 242 MI demonstration transcripts annotated with the MI Treatment Integrity (MITI) 4.2 therapist behavioral codes and global scores, and Client Language EAsy Rating (CLEAR) 1.0 tags for client speech. In this paper we describe the process of data collection, transcription, and annotation, and provide an analysis of the new dataset. Additionally, we explore the potential use of the dataset for training language models to perform several MITI classification tasks; our results suggest that models may be able to automatically provide utterance-level annotation as well as global scores, with performance comparable to human annotators. Ben Cohen, Moreah Zisquit, Stav Yosef, Doron Friedman, Kfir Bar |
LREC/COLING | 5 |
| 2024 | DiaSet: An Annotated Dataset of Arabic ConversationsabstractWe introduce DiaSet, a novel dataset of dialectical Arabic speech, manually transcribed and annotated for two specific downstream tasks: sentiment analysis and named entity recognition. The dataset encapsulates the Palestine dialect, predominantly spoken in Palestine, Israel, and Jordan. Our dataset incorporates authentic conversations between YouTube influencers and their respective guests. Furthermore, we have enriched the dataset with simulated conversations initiated by inviting participants from various locales within the said regions. The participants were encouraged to engage in dialogues with our interviewer. Overall, DiaSet consists of 644.8K tokens and 23.2K annotated instances. Uniform writing standards were upheld during the transcription process. Additionally, we established baseline models by leveraging some of the pre-existing Arabic BERT language models, showcasing the potential applications and efficiencies of our dataset. We make DiaSet publicly available for further research. Abraham Israeli, Aviv Naaman, Guy Maduel, Rawaa Makhoul, Dana Qaraeen, Amir Ejmail, Dina Lisnanskey, Julian Jubran, Shai Fine, Kfir Bar |
LREC/COLING | 10 |
| 2023 | Automatic Translation of Span-Prediction DatasetsabstractOfri Masad, Kfir Bar, Amir Cohen. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ofri Masad, Kfir Bar, Amir Cohen |
IJCNLP (1) | 2 |
| 2022 | How Much Does Lookahead Matter for Disambiguation? Partial Arabic Diacritization Case StudyabstractAbstract We suggest a model for partial diacritization of deep orthographies. We focus on Arabic, where the optional indication of selected vowels by means of diacritics can resolve ambiguity and improve readability. Our partial diacritizer restores short vowels only when they contribute to the ease of understandability during reading a given running text. The idea is to identify those uncertainties of absent vowels that require the reader to look ahead to disambiguate. To achieve this, two independent neural networks are used for predicting diacritics, one that takes the entire sentence as input and another that considers only the text that has been read thus far. Partial diacritization is then determined by retaining precisely those vowels on which the two networks disagree, preferring the reading based on consideration of the whole sentence over the more naïve reading-order diacritization. For evaluation, we prepared a new dataset of Arabic texts with both full and partial vowelization. In addition to facilitating readability, we find that our partial diacritizer improves translation quality compared either to their total absence or to random selection. Lastly, we study the benefit of knowing the text that follows the word in focus toward the restoration of short vowels during reading, and we measure the degree to which lookahead contributes to resolving ambiguities encountered while reading. L’Herbelot had asserted, that the most ancient Korans, written in the Cufic character, had no vowel points; and that these were first invented by Jahia–ben Jamer, who died in the 127th year of the Hegira. “Toderini’s History of Turkish Literature,” Analytical Review (1789) Saeed Esmail, Kfir Bar, Nachum Dershowitz |
Comput. Linguistics | 2 |
| 2021 | Balancing Speed and Accuracy in Neural-Enhanced Phonetic Name Matching
Philip Blair, Carmel Eliav, Fiona Hasanaj, Kfir Bar |
ECML/PKDD (5) | 4 |
| 2014 | Inferring Paraphrases for a Highly Inflected Language from a Monolingual Corpus
Kfir Bar, Nachum Dershowitz |
CICLing (2) | 1 |
| 2012 | Deriving Paraphrases for Highly Inflected Languages from Comparable Documents
Kfir Bar, Nachum Dershowitz |
COLING | 1 |
| 2009 | Automatically Classifying Documents by Ideological and Organizational AffiliationabstractWe show how an Arabic language religious-political document can be automatically classified according to the ideological stream and organizational affiliation that it represents. Tests show that our methods achieve near-perfect accuracy. Moshe Koppel, Navot Akiva, Eli Alshech, Kfir Bar |
ISI | 4 |