EDBT 2026 Demo / reviewers in the wild / expert
Orgad Keller
dblp:32/3363
· DBLP profile ↗
19ranked-venue papers
8as first author
6since 2021 · last 2026
0000-0001-8827-429XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 1 since 2021Theory of computation · 5 · 3 first-authorDatabases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 31% Reinforcement learning · 21% Question answering and dialogue systems · 14% | |
| Theoretical computer science
2 papers |
Algorithmic game theory and mechanism design · 58% Approximation and online algorithms · 42% |
Topics — the 14 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
text summarization |
0.9 | 2 | 2025 | Factually Consistent Summarization via Reinforcement Learning with Textual Entailment Feedback · ACL (1) 2023 Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance · EMNLP 2025 |
Natural language and speech › Information extraction and text analysis
data annotation |
0.9 | 1 | 2025 | Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance · EMNLP 2025 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
noisy label detection |
0.9 | 1 | 2025 | Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance · EMNLP 2025 |
Machine learning › Optimization for machine learning
mirror descent |
0.8 | 1 | 2024 | Multi-turn Reinforcement Learning with Preference Human Feedback · NeurIPS 2024 |
Machine learning › Reinforcement learning
policy optimization |
0.8 | 1 | 2024 | Multi-turn Reinforcement Learning with Preference Human Feedback · NeurIPS 2024 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.8 | 1 | 2024 | Multi-turn Reinforcement Learning with Preference Human Feedback · NeurIPS 2024 |
Approximation and online algorithms
approximation |
0.5 | 1 | 2021 | Targeted Negative Campaigning: Complexity and Approximations · AAAI 2021 |
Natural language and speech › Question answering and dialogue systems
dialogue management |
0.4 | 1 | 2020 | Dynamic Composition for Conversational Domain Exploration · WWW 2020 |
Approximation and online algorithms
approximation algorithms |
0.3 | 1 | 2018 | Approximating Bribery in Scoring Rules · AAAI 2018 |
Algorithmic game theory and mechanism design › social choice › computational social choice › voting manipulation
coalitional manipulation |
0.3 | 1 | 2018 | Approximating Bribery in Scoring Rules · AAAI 2018 |
Algorithmic game theory and mechanism design › social choice
computational social choice |
0.3 | 1 | 2018 | Approximating Bribery in Scoring Rules · AAAI 2018 |
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability › factuality
factual consistency |
0.3 | 1 | 2025 | Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance · EMNLP 2025 |
Natural language and speech › Language models and text generation
alignment |
0.2 | 1 | 2024 | Multi-turn Reinforcement Learning with Preference Human Feedback · NeurIPS 2024 |
Natural language and speech › Question answering and dialogue systems
conversational agents |
0.1 | 1 | 2020 | Dynamic Composition for Conversational Domain Exploration · WWW 2020 |
Methods — techniques the papers use, named apart from their topics
large language model as a judge · 0.9ensemble of large language models · 0.9mirror-descent-based policy optimization · 0.8deep RL · 0.8textual entailment feedback · 0.7reinforcement learning · 0.7approximation algorithm · 0.5randomized reduction · 0.3birkhoff-von neumann decomposition · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficiently Negative: Complexity and Approximations of Targeted Negative CampaigningabstractGiven the ubiquity of negative campaigning in recent political elections, we find it important to study its properties from a theoretical computational perspective. To this end, we present a model where elections can be manipulated by convincing voters to demote specific non-favored candidates, and study its properties in the classic setting of scoring rules. When the goal is constructive (making a preferred candidate win), we prove that finding such a demotion strategy is easy for Plurality and Veto, while generally hard for t-approval and Borda. We also provide a min(t, m - t)-factor approximation for t-approval for every t ∈ {1,..., m - 1} (where m is the number of candidates), and a 3-factor approximation algorithm for Borda. Interestingly enough---following recent trends in political science that show that the effectiveness of negative campaigning depends on the type of candidate and demographic---when assigning varying prices to different possible demotion operations, we are able to provide inapproximability results. When the goal is destructive (making the leading opponent lose), we show that the problem is easy for a broad class of scoring rules and provide an FPTAS for the general case. Avishai Zagoury, Orgad Keller, Avinatan Hassidim, Noam Hazon |
J. Artif. Intell. Res. | 2 |
| 2025 | Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model PerformanceabstractNLP benchmarks rely on standardized datasets for training and evaluating models and are crucial for advancing the field.Traditionally, expert annotations ensure high-quality labels; however, the cost of expert annotation does not scale well with the growing demand for larger datasets required by modern models.While crowd-sourcing provides a more scalable solution, it often comes at the expense of annotation precision and consistency.Recent advancements in large language models (LLMs) offer new opportunities to enhance the annotation process, particularly for detecting label errors in existing datasets.In this work, we consider the recent approach of LLM-as-a-judge, leveraging an ensemble of LLMs to flag potentially mislabeled examples.We conduct a case study on four factual consistency datasets from the TRUE benchmark, spanning diverse NLP tasks, and on SummEval, which uses Likertscale ratings of summary quality across multiple dimensions.We empirically analyze the labeling quality of existing datasets and compare expert, crowd-sourced, and LLM-based annotations in terms of the agreement, label quality, and efficiency, demonstrating the strengths and limitations of each annotation method.Our findings reveal a substantial number of label errors, which, when corrected, induce a significant upward shift in reported model performance.This suggests that many of the LLMs' so-called mistakes are due to label errors rather than genuine model failures.Additionally, we discuss the implications of mislabeled data and propose methods to mitigate them in training to improve performance. Omer Nahum, Nitay Calderon, Orgad Keller, Idan Szpektor, Roi Reichart |
EMNLP | 3 |
| 2024 | Multi-turn Reinforcement Learning with Preference Human FeedbackabstractReinforcement Learning from Human Feedback (RLHF) has become the standard approach for aligning Large Language Models (LLMs) with human preferences, allowing LLMs to demonstrate remarkable abilities in various tasks. Existing methods work by emulating the human preference at the single decision (turn) level, limiting their capabilities in settings that require planning or multi-turn interactions to achieve a long-term goal. In this paper, we address this issue by developing novel methods for Reinforcement Learning (RL) from preference feedback between two full multi-turn conversations. In the tabular setting, we present a novel mirror-descent-based policy optimization algorithm for the general multi-turn preference-based RL problem, and prove its convergence to Nash equilibrium. To evaluate performance, we create a new environment, Education Dialogue, where a teacher agent guides a student in learning a random topic, and show that a deep RL variant of our algorithm outperforms RLHF baselines. Finally, we show that in an environment with explicit rewards, our algorithm recovers the same performance as a reward-based RL baseline, despite relying solely on a weaker preference signal. Lior Shani, Aviv Rosenberg 0002, Asaf Cassel, Oran Lang, Daniele Calandriello, Avital Zipori, Hila Noga, Orgad Keller, Bilal Piot, Idan Szpektor, Avinatan Hassidim, Yossi Matias, Rémi Munos |
NeurIPS | 8 |
| 2023 | Factually Consistent Summarization via Reinforcement Learning with Textual Entailment FeedbackabstractPaul Roit, Johan Ferret, Lior Shani, Roee Aharoni, Geoffrey Cideron, Robert Dadashi, Matthieu Geist, Sertan Girgin, Leonard Hussenot, Orgad Keller, Nikola Momchev, Sabela Ramos Garea, Piotr Stanczyk, Nino Vieillard, Olivier Bachem, Gal Elidan, Avinatan Hassidim, Olivier Pietquin, Idan Szpektor. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Paul Roit, Johan Ferret, Lior Shani, Roee Aharoni, Geoffrey Cideron, Robert Dadashi, Matthieu Geist, Sertan Girgin, Léonard Hussenot, Orgad Keller, Nikola Momchev, Sabela Ramos, Piotr Stanczyk, Nino Vieillard, Olivier Bachem, Gal Elidan, Avinatan Hassidim, Olivier Pietquin, Idan Szpektor |
ACL (1) | 10 |
| 2023 | On the Robustness of Dialogue History Representation in Conversational Question Answering: A Comprehensive Study and a New Prompt-based MethodabstractAbstract Most work on modeling the conversation history in Conversational Question Answering (CQA) reports a single main result on a common CQA benchmark. While existing models show impressive results on CQA leaderboards, it remains unclear whether they are robust to shifts in setting (sometimes to more realistic ones), training data size (e.g., from large to small sets) and domain. In this work, we design and conduct the first large-scale robustness study of history modeling approaches for CQA. We find that high benchmark scores do not necessarily translate to strong robustness, and that various methods can perform extremely differently under different settings. Equipped with the insights from our study, we design a novel prompt-based history modeling approach and demonstrate its strong robustness across various settings. Our approach is inspired by existing methods that highlight historic answers in the passage. However, instead of highlighting by modifying the passage token embeddings, we add textual prompts directly in the passage text. Our approach is simple, easy to plug into practically any model, and highly effective, thus we recommend it as a starting point for future model developers. We also hope that our study and insights will raise awareness to the importance of robustness-focused evaluation, in addition to obtaining high leaderboard scores, leading to better CQA systems.1 Zorik Gekhman, Nadav Oved, Orgad Keller, Idan Szpektor, Roi Reichart |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | Targeted Negative Campaigning: Complexity and Approximations
Avishai Zagoury, Orgad Keller, Avinatan Hassidim, Noam Hazon |
AAAI | 2 |
| 2020 | Dynamic Composition for Conversational Domain ExplorationabstractWe study conversational domain exploration (CODEX), where the user’s goal is to enrich her knowledge of a given domain by conversing with an informative bot. Such conversations should be well grounded in high-quality domain knowledge as well as engaging and open-ended. A CODEX bot should be proactive and introduce relevant information even if not directly asked for by the user. The bot should also appropriately pivot the conversation to undiscovered regions of the domain. To address these dialogue characteristics, we introduce a novel approach termed dynamic composition that decouples candidate content generation from the flexible composition of bot responses. This allows the bot to control the source, correctness and quality of the offered content, while achieving flexibility via a dialogue manager that selects the most appropriate contents in a compositional manner. We implemented a CODEX bot based on dynamic composition and integrated it into the Google Assistant . As an example domain, the bot conversed about the NBA basketball league in a seamless experience, such that users were not aware whether they were conversing with the vanilla system or the one augmented with our CODEX bot. Results are positive and offer insights into what makes for a good conversation. To the best of our knowledge, this is the first real user experiment of open-ended dialogues as part of a commercial assistant system. Idan Szpektor, Deborah Cohen, Gal Elidan, Avinatan Hassidim, Orgad Keller, Sayali Kulkarni, Eran Ofek, Sagie Pudinsky, Asaf Revach, Shimi Salant, Yossi Matias |
WWW | 6 |
| 2019 | New Approximations for Coalitional Manipulation in Scoring RulesabstractWe study the problem of coalitional manipulation---where k manipulators try to manipulate an election on m candidates---for any scoring rule, with focus on the Borda protocol. We do so in both the weighted and unweighted settings. For these problems, recent approximation approaches have tried to minimize k, the number of manipulators needed to make some preferred candidate p win (thus assuming that the number of manipulators is not limited in advance). In contrast, we focus on minimizing the score margin of p which is the difference between the maximum score of a candidate and the score of p. We provide algorithms that approximate the optimum score margin, which are applicable to any scoring rule. For the specific case of the Borda protocol in the unweighted setting, our algorithm provides a superior approximation factor for lower values of k.Our methods are novel and adapt techniques from multiprocessor scheduling by carefully rounding an exponentially-large configuration linear program that is solved by using the ellipsoid method with an efficient separation oracle. We believe that such methods could be beneficial in other social choice settings as well. Orgad Keller, Avinatan Hassidim, Noam Hazon |
J. Artif. Intell. Res. | 1 |
| 2019 | Approximating Weighted and Priced Bribery in Scoring RulesabstractThe classic Bribery problem is to find a minimal subset of voters who need to change their vote to make some preferred candidate win. Its important generalizations consider voters who are weighted and also have different prices. We provide an approximate solution for these problems for a broad family of scoring rules (which includes Borda and t-approval), in the following sense: for constant weights and prices, if there exists a strategy which costs k, we efficiently find a strategy which costs at most k+\widetilde{O}(sqrt(k)). An extension for non-constant weights and prices is also given. Our algorithm is based on a randomized reduction from these Bribery generalizations to weighted coalitional manipulation (WCM). To solve this WCM instance, we apply the Birkhoff-von Neumann (BvN) decomposition to a fractional manipulation matrix. This allows us to limit the size of the possible ballot search space reducing it from exponential to polynomial, while still obtaining good approximation guarantees. Finding a solution in the truncated search space yields a new algorithm for WCM, which is of independent interest. Orgad Keller, Avinatan Hassidim, Noam Hazon |
J. Artif. Intell. Res. | 1 |
| 2018 | Approximating Bribery in Scoring RulesabstractThe classic bribery problem is to find a minimal subset of voters who need to change their vote to make some preferred candidate win.We find an approximate solution for this problem for a broad family of scoring rules (which includes Borda and t-approval), in the following sense: if there is a strategy which requires bribing k voters, we efficiently find a strategy which requires bribing at most k + Õ(√k) voters. Our algorithm is based on a randomized reduction from bribery to coalitional manipulation (UCM). To solve the UCM problem, we apply the Birkhoff-von Neumann (BvN) decomposition to a fractional manipulation matrix. This allows us to limit the size of the possible ballot search space reducing it from exponential to polynomial, while still obtaining good approximation guarantees. Finding the optimal solution in the truncated search space yields a new algorithm for UCM, which is of independent interest. Orgad Keller, Avinatan Hassidim, Noam Hazon |
AAAI | 1 |
| 2018 | Fully Automatic Speaker Separation System, with Automatic Enrolling of Recurrent Speakers
Raphael Cohen, Orgad Keller, Jason Levy, Russell Levy, Micha Breakstone, Amit Ashkenazi |
INTERSPEECH | 2 |
| 2014 | Generalized substring compression
Orgad Keller, Tsvi Kopelowitz, Shir Landau Feibish, Moshe Lewenstein |
Theor. Comput. Sci. | 1 |
| 2013 | Finding the Minimum-Weight k-Path
Avinatan Hassidim, Orgad Keller, Moshe Lewenstein, Liam Roditty |
WADS | 2 |
| 2011 | Approximate string matching with stuck address bits
Amihood Amir, Estrella Eisenberg, Orgad Keller, Avivit Levy, Ely Porat |
Theor. Comput. Sci. | 3 |
| 2010 | Approximate String Matching with Stuck Address Bits
Amihood Amir, Estrella Eisenberg, Orgad Keller, Avivit Levy, Ely Porat |
SPIRE | 3 |
| 2009 | Generalized Substring Compression
Orgad Keller, Tsvi Kopelowitz, Shir Landau Feibish, Moshe Lewenstein |
CPM | 1 |
| 2009 | On the longest common parameterized subsequence
Orgad Keller, Tsvi Kopelowitz, Moshe Lewenstein |
Theor. Comput. Sci. | 1 |
| 2008 | On the Longest Common Parameterized Subsequence
Orgad Keller, Tsvi Kopelowitz, Moshe Lewenstein |
CPM | 1 |
| 2007 | Range Non-overlapping Indexing and Successive List Indexing
Orgad Keller, Tsvi Kopelowitz, Moshe Lewenstein |
WADS | 1 |