EDBT 2026 Demo / reviewers in the wild / expert
Ori Yoran
dblp:290/1285
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 56% Question answering and dialogue systems · 29% Planning, search and constraint satisfaction · 7% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% | |
| Theoretical computer science
1 paper |
Computational complexity · 100% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
code generation |
0.9 | 1 | 2025 | The KoLMogorov Test: Compression by Code Generation · ICLR 2025 |
Computational complexity
kolmogorov complexity |
0.9 | 1 | 2025 | The KoLMogorov Test: Compression by Code Generation · ICLR 2025 |
Natural language and speech › Language models and text generation
LLM agents |
0.8 | 1 | 2024 | AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks? · EMNLP 2024 |
Natural language and speech › Question answering and dialogue systems
multi-hop reasoning |
0.8 | 1 | 2024 | Making Retrieval-Augmented Language Models Robust to Irrelevant Context · ICLR 2024 |
Natural language and speech › Question answering and dialogue systems
open-domain question answering |
0.8 | 1 | 2024 | Making Retrieval-Augmented Language Models Robust to Irrelevant Context · ICLR 2024 |
Natural language and speech › Language models and text generation
retrieval-augmented language models |
0.8 | 1 | 2024 | Making Retrieval-Augmented Language Models Robust to Irrelevant Context · ICLR 2024 |
Natural language and speech › Language models and text generation › LLM agents
web agents |
0.8 | 1 | 2024 | AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks? · EMNLP 2024 |
Information retrieval
web navigation |
0.8 | 1 | 2024 | AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks? · EMNLP 2024 |
Information retrieval
web search |
0.8 | 1 | 2024 | AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks? · EMNLP 2024 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.7 | 1 | 2023 | Answering Questions by Meta-Reasoning over Multiple Chains of Thought · EMNLP 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
metareasoning |
0.7 | 1 | 2023 | Answering Questions by Meta-Reasoning over Multiple Chains of Thought · EMNLP 2023 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering |
0.7 | 1 | 2023 | Answering Questions by Meta-Reasoning over Multiple Chains of Thought · EMNLP 2023 |
Natural language and speech › Language models and text generation › natural language understanding
long document understanding |
0.6 | 1 | 2022 | SCROLLS: Standardized CompaRison Over Long Language Sequences · EMNLP 2022 |
Information retrieval
evaluation |
0.6 | 1 | 2022 | SCROLLS: Standardized CompaRison Over Long Language Sequences · EMNLP 2022 |
Natural language and speech › Question answering and dialogue systems
multimodal question answering |
0.5 | 1 | 2021 | MultiModalQA: complex question answering over text, tables and images · ICLR 2021 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.3 | 1 | 2025 | The KoLMogorov Test: Compression by Code Generation · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
compression · 1.7code generation · 1.7retrieval-augmented language model · 1.5natural language inference filtering · 0.8fine-tuning · 0.8data generation · 0.8large language model · 0.7chain-of-thought prompting · 0.7pre-training · 0.6error-based sampling · 0.6benchmark construction · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to RepresentationsabstractAbstract Language models (LMs) increasingly drive real-world applications that require world knowledge. However, the internal processes through which models turn data into representations of knowledge and beliefs about the world are poorly understood. To facilitate such studies, we present LMEnt, a suite including (1) a knowledge-rich pretraining corpus, fully annotated with entity mentions based on Wikipedia, (2) an entity-based retrieval method over pretraining data that outperforms existing tools by as much as 80.4%, and (3) 12 pretrained LMs with up to 1B parameters and 4K intermediate checkpoints, with comparable performance to popular open-source models on knowledge tasks. Together, these resources provide a controlled environment for analyzing connections between entity mentions in pretraining data and downstream performance. We show the utility of LMEnt by studying knowledge acquisition over training, finding that entity co-occurrence and mention forms—which are difficult to study with existing tools—affect learning trends. Moreover, as LMs form stronger associations between entities, their facts are harder to edit in-context, whereas inconsistencies in model predictions over training are indicative of editing success. We release LMEnt to support studies of knowledge in LMs, including knowledge representations, plasticity, editing, attribution, hallucinations, and learning dynamics. huggingface.co/LMEnt github.com/LMEnt Daniela Gottesman, Alon Gilae-Dotan, Ido Cohen 0002, Yoav Gur-Arieh, Marius Mosbach, Ori Yoran, Mor Geva |
Trans. Assoc. Comput. Linguistics | 6 |
| 2025 | The KoLMogorov Test: Compression by Code GenerationabstractCompression is at the heart of intelligence. A theoretically optimal way to compress any sequence of data is to find the shortest program that outputs that sequence and then halts. However, such Kolmogorov compression is uncomputable, and code generating LLMs struggle to approximate this theoretical ideal, as it requires reasoning, planning and search capabilities beyond those of current models. In this work, we introduce the *KoLMogorov-Test* (KT), a compression-as-intelligence intelligence test for code generation LLMs. In KT a model is presented with a sequence of data at inference time, and asked to generate the shortest program that produces the sequence. We identify several benefits of KT for both evaluation and training: an essentially infinite number of problem instances of varying difficulty is readily available, strong baselines already exist, the evaluation metric (compression) cannot be gamed, and pretraining data contamination is highly unlikely. To evaluate current models, we use audio, text, and DNA data, as well as sequences produced by random synthetic programs. Current flagship models perform poorly - both GPT4-o and Llama-3.1-405B struggle on our natural and synthetic sequences. On our synthetic distribution, we are able to train code generation models with lower compression rates than previous approaches. Moreover, we show that gains on synthetic data generalize poorly to real data, suggesting that new innovations are necessary for additional gains on KT. Ori Yoran, Kunhao Zheng, Fabian Gloeckle, Jonas Gehring, Gabriel Synnaeve, Taco Cohen |
ICLR | 1 |
| 2024 | AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?abstractLanguage agents, built on top of language models (LMs), are systems that can interact with complex environments, such as the open web.In this work, we examine whether such agents can perform realistic and time-consuming tasks on the web, e.g., monitoring real-estate markets or locating relevant nearby businesses.We introduce ASSISTANTBENCH, a challenging new benchmark consisting of 214 realistic tasks that can be automatically evaluated, covering different scenarios and domains.We find that AS-SISTANTBENCH exposes the limitations of current systems, including language models and retrieval-augmented language models, as no model reaches an accuracy of more than 25 points.While closed-book LMs perform well in terms of accuracy, they exhibit low precision and tend to hallucinate facts.State-of-the-art web agents reach a score of near zero.Additionally, we introduce SEEPLANACT (SPA), a new web agent that significantly outperforms previous agents, and an ensemble of SPA and closed-book models reaches the best overall performance.Moreover, we analyze failures of current systems and highlight that open web navigation remains a major challenge. Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin, Ofir Press, Jonathan Berant |
EMNLP | 1 |
| 2024 | Making Retrieval-Augmented Language Models Robust to Irrelevant ContextabstractRetrieval-augmented language models (RALMs) hold promise to produce language understanding systems that are are factual, efficient, and up-to-date. An important desideratum of RALMs, is that retrieved information helps model performance when it is relevant, and does not harm performance when it is not. This is particularly important in multi-hop reasoning scenarios, where misuse of irrelevant evidence can lead to cascading errors. However, recent work has shown that retrieval augmentation can sometimes have a negative effect on performance. In this work, we present a thorough analysis on five open-domain question answering benchmarks, characterizing cases when retrieval reduces accuracy. We then propose two methods to mitigate this issue. First, a simple baseline that filters out retrieved passages that do not entail question-answer pairs according to a natural language inference (NLI) model. This is effective in preventing performance reduction, but at a cost of also discarding relevant passages. Thus, we propose a method for automatically generating data to fine-tune the language model to properly leverage retrieved passages, using a mix of relevant and irrelevant contexts at training time. We empirically show that even 1,000 examples suffice to train the model to be robust to irrelevant contexts while maintaining high performance on examples with relevant ones. Ori Yoran, Tomer Wolfson, Ori Ram, Jonathan Berant |
ICLR | 1 |
| 2024 | Evaluating the Ripple Effects of Knowledge Editing in Language ModelsabstractAbstract Modern language models capture a large body of factual knowledge. However, some facts can be incorrectly induced or become obsolete over time, resulting in factually incorrect generations. This has led to the development of various editing methods that allow updating facts encoded by the model. Evaluation of these methods has primarily focused on testing whether an individual fact has been successfully injected, and if similar predictions for other subjects have not changed. Here we argue that such evaluation is limited, since injecting one fact (e.g., “Jack Depp is the son of Johnny Depp”) introduces a “ripple effect” in the form of additional facts that the model needs to update (e.g., “Jack Depp is the sibling of Lily-Rose Depp”). To address this, we propose novel evaluation criteria that consider the implications of an edit on related facts. Using these criteria, we then construct RippleEdits, a diagnostic benchmark of 5K factual edits, capturing various types of ripple effects. We evaluate prominent editing methods on RippleEdits, showing that they fail to introduce consistent changes in the model’s knowledge. In addition, we find that a simple in-context editing baseline obtains the best scores on our benchmark, suggesting a promising research direction for model editing.1 Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, Mor Geva |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | Answering Questions by Meta-Reasoning over Multiple Chains of ThoughtabstractModern systems for multi-hop question answering (QA) typically break questions into a sequence of reasoning steps, termed chain-ofthought (CoT), before arriving at a final answer.Often, multiple chains are sampled and aggregated through a voting mechanism over the final answers, but the intermediate steps themselves are discarded.While such approaches improve performance, they do not consider the relations between intermediate steps across chains and do not provide a unified explanation for the predicted answer.We introduce Multi-Chain Reasoning (MCR), an approach which prompts large language models to meta-reason over multiple chains of thought, rather than aggregate their answers.MCR examines different reasoning chains, mixes information between them and selects the most relevant facts in generating an explanation and predicting the answer.MCR outperforms strong baselines on 7 multi-hop QA datasets.Moreover, our analysis reveals that MCR explanations exhibit high quality, enabling humans to verify its answers. Ori Yoran, Tomer Wolfson, Ben Bogin, Uri Katz, Daniel Deutch, Jonathan Berant |
EMNLP | 1 |
| 2022 | Turning Tables: Generating Examples from Semi-structured Tables for Endowing Language Models with Reasoning SkillsabstractModels pre-trained with a language modeling objective possess ample world knowledge and language skills, but are known to struggle in tasks that require reasoning.In this work, we propose to leverage semi-structured tables, and automatically generate at scale questionparagraph pairs, where answering the question requires reasoning over multiple facts in the paragraph.We add a pre-training step over this synthetic data, which includes examples that require 16 different reasoning skills such as number comparison, conjunction, and fact composition.To improve data efficiency, we sample examples from reasoning skills where the model currently errs.We evaluate our approach on three reasoning-focused reading comprehension datasets, and show that our model, PReasM, substantially outperforms T5, a popular pre-trained encoder-decoder model.Moreover, sampling examples based on model errors leads to faster training and higher performance. Ori Yoran, Alon Talmor, Jonathan Berant |
ACL (1) | 1 |
| 2022 | SCROLLS: Standardized CompaRison Over Long Language SequencesabstractUri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, Omer Levy. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Uri Shaham 0002, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta 0001, Wenhan Xiong, Mor Geva, Jonathan Berant, Omer Levy |
EMNLP | 5 |
| 2021 | MultiModalQA: complex question answering over text, tables and images
Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, Jonathan Berant |
ICLR | 2 |