EDBT 2026 Demo / reviewers in the wild / expert
Jonathan Berant
dblp:31/8178
· DBLP profile ↗
82ranked-venue papers
10as first author
42since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 79 · 9 first-author · 41 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Comparing human and language models sentence processing difficulties on complex structuresabstractLarge language models (LLMs) that fluently converse with humans are a reality -but do LLMs experience human-like processing difficulties?We systematically compare human and LLM sentence comprehension across seven challenging linguistic structures.We collect sentence comprehension data from humans and five families of state-of-the-art LLMs, varying in size and training procedure in a unified experimental framework.Our results show LLMs overall struggle on the target structures, but especially on garden path (GP) sentences.Indeed, while the strongest models achieve near perfect accuracy on non-GP structures (93.7% for GPT-5), they struggle on GP structures (46.8% for GPT-5).Additionally, when ranking structures based on average performance, rank correlation between humans and models increases with parameter count.For each target structure, we also collect data for their matched baseline without the difficult structure.Comparing performance on the target vs. baseline sentences, the performance gap observed in humans holds for LLMs, with two exceptions: for models that are too weak performance is uniformly low across both sentence types, and for models that are too strong the performance is uniformly high.Together, these reveal convergence and divergence in human and LLM sentence comprehension, offering new insights into the similarity of humans and LLMs. 1 Model Average Subj/Obj NP/VP Depth charge NP/S Double center Interference Red.relative GP non GP Human 28.3 13.3 18.5 28.0 29.7 32.3 36.9 41.7 25.8 32.4 o3 74.5 49.0 66.0 64.0 56.0 98.0 100.0 95.0 66.5 87.3 GPT-4.1 68.7 40.0 63.0 63.0 57.0 90.0 100.0 76.0 59.0 84.3 GPT-5 65.6 32.0 45.0 83.0 45.0 98.0 100.0 65.0 46.8 93.7 Llama-11B-Ins.60.6 50.0 72.0 33.0 62.0 61.0 69.0 79.0 65.7 54.3 Llama-11B 59.6 47.0 66.0 38.0 56.0 62.0 82.0 72.0 60.3 60.7 Llama-90B 49.4 25.0 40.0 43.0 29.0 88.0 93.0 39.0 33.3 74.7 DeepSeek-7B 48.0 32.0 41.0 78.0 44.0 25.0 68.0 53.0 42.5 57.0 DeepSeek-1.5B 46.9 38.0 39.0 69.0 39.0 59.0 43.0 40.0 39.0 57.0 DeepSeek-14B 38.6 15.0 25.0 74.0 7.0 66.0 65.0 25.0 18.0 68.3 Qwen-0.6B47.7 36.0 65.0 51.0 39.0 52.0 52.0 40.0 45.0 51.7 Qwen-8B 47.5 28.0 42.0 56.0 33.0 66.0 77.0 38.0 35.3 66.3 Qwen-14B 45.1 12.0 22.0 79.0 10.0 79.0 91.0 34.0 19.5 83.0 Gemma-1B 44.4 40.0 43.0 65.0 38.0 44.0 41.0 39.0 40.0 50.0 Gemma-1B-Ins.39.7 40.0 52.0 64.0 34.0 20.0 28.0 37.0 40.7 37.3 Gemma-4B 38.6 35.0 38.0 55.0 37.0 35.0 35.0 34.0 36.0 41.7 Samuel Joseph Amouyal, Aya Meltzer-Asscher, Jonathan Berant |
ACL (1) | 3 |
| 2025 | When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language modelsabstractModern Large Language Models (LLMs) have shown human-like abilities in many language tasks, sparking interest in comparing LLMs’ and humans’ language processing. In this paper, we try to answer two questions: 1. What makes garden-path sentences hard to understand for humans? 2. Do the same reasons make garden-path sentences hard for LLMs as well? Based on psycholinguistic research, we formulate hypotheses on why garden-path sentences are hard, and test these hypotheses on human participants and a large suite of LLMs using comprehension questions. Our findings reveal that both LLMs and humans struggle with specific syntactic complexities, with some models showing high correlation with human comprehension. To complement our findings, we test LLM comprehension of garden-path constructions with paraphrasing and text-to-image generation tasks, and find that the results mirror the sentence comprehension question results, further validating our findings on LLM understanding of these constructions. Samuel Joseph Amouyal, Aya Meltzer-Asscher, Jonathan Berant |
ACL (1) | 3 |
| 2025 | Rewarding Progress: Scaling Automated Process Verifiers for LLM ReasoningabstractA promising approach for improving reasoning in large language models is to use process reward models (PRMs). PRMs provide feedback at each step of a multi-step reasoning trace, improving credit assignment over outcome reward models (ORMs) that only provide feedback at the final step. However, collecting dense, per-step human labels is not scalable, and training PRMs from automatically-labeled data has thus far led to limited gains. With the goal of using PRMs to improve a *base* policy via test-time search and reinforcement learning (RL), we ask: ``How should we design process rewards?'' Our key insight is that, to be effective, the process reward for a step should measure
*progress*: a change in the likelihood of producing a correct response in the future, before and after taking the step, as measured under a *prover* policy distinct from the base policy. Such progress values can {distinguish} good and bad steps generated by the base policy, even though the base policy itself cannot. Theoretically, we show that even weaker provers can improve the base policy, as long as they distinguish steps without being too misaligned with the base policy. Our results show that process rewards defined as progress under such provers improve the efficiency of exploration during test-time search and online RL. We empirically validate our claims by training **process advantage verifiers (PAVs)** to measure progress under such provers and show that compared to ORM, they are >8% more accurate, and 1.5-5x more compute-efficient. Equipped with these insights, our PAVs enable **one of the first results** showing a 6x gain in sample efficiency for a policy trained using online RL with PRMs vs. ORMs. Amrith Setlur, Chirag Nagpal, Adam Fisch, Xinyang Geng, Jacob Eisenstein, Rishabh Agarwal, Alekh Agarwal, Jonathan Berant, Aviral Kumar |
ICLR | 8 |
| 2025 | InfAlign: Inference-aware language model alignmentabstractLanguage model alignment is a critical step
in training modern generative language models.
Alignment targets to improve win rate of a sample
from the aligned model against the base model.
Today, we are increasingly using inference-time
algorithms (e.g., Best-of-$N$ , controlled decoding, tree search) to decode from language models
rather than standard sampling. We show that this
train/test mismatch makes standard RLHF framework sub-optimal in view of such inference-time
methods. To this end, we propose a framework for
inference-aware alignment (InfAlign), which
aims to optimize *inference-time win rate* of the
aligned policy against the base model. We prove
that for any inference-time decoding procedure,
the optimal aligned policy is the solution to the
standard RLHF problem with a *transformation*
of the reward. This motivates us to provide the
calibrate-and-transform RL (InfAlign-CTRL)
algorithm to solve this problem, which involves
a reward calibration step and a KL-regularized
reward maximization step with a transformation
of the calibrated reward. For best-of-$N$ sampling
and best-of-$N$ jailbreaking, we propose specific
transformations offering up to 3-8% improvement
on inference-time win rates. Finally, we also show
that our proposed reward calibration method is a
strong baseline for optimizing standard win rate. Ananth Balashankar, Ziteng Sun, Jonathan Berant, Jacob Eisenstein, Michael Collins 0001, Adrian Hutter, Jong Lee, Chirag Nagpal, Flavien Prost, Aradhana Sinha, Ananda Theertha Suresh, Ahmad Beirami |
ICML | 3 |
| 2025 | Theoretical guarantees on the best-of-n alignment policyabstractA simple and effective method for the inference-time alignment of generative models is the best-of-$n$ policy, where $n$ samples are drawn from a reference policy, ranked based on a reward function, and the highest ranking one is selected. A commonly used analytical expression in the literature claims that the KL divergence between the best-of-$n$ policy and the reference policy is equal to $\log (n) - (n-1)/n.$ We disprove the validity of this claim, and show that it is an upper bound on the actual KL divergence. We also explore the tightness of this upper bound in different regimes, and propose a new estimator for the KL divergence and empirically show that it provides a tight approximation. We also show that the win rate of the best-of-$n$ policy against the reference policy is upper bounded by $n/(n+1)$ and derive bounds on the tightness of this characterization. We conclude with analyzing the tradeoffs between win rate and KL divergence of the best-of-$n$ alignment policy, which demonstrate that very good tradeoffs are achievable with $n < 1000$. Ahmad Beirami, Alekh Agarwal, Jonathan Berant, Alexander D'Amour, Jacob Eisenstein, Chirag Nagpal, Ananda Theertha Suresh |
ICML | 3 |
| 2025 | Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven Priors (Extended Abstract)abstractThis paper is an extended abstract of our ICLR 2024 Outstanding Paper Award work. Modeling long-range dependencies across sequences is a longstanding goal in machine learning. While state space models reportedly outperform Transformers on benchmarks like Long Range Arena, we show that random initialization significantly overestimates architectural differences. Pretraining with standard denoising objectives on downstream task data leads to dramatic gains across architectures and minimal performance gaps between Transformers and state space models (SSMs). We demonstrate that properly pretrained vanilla Transformers match S4 performance on Long Range Arena and improve SSM results on PathX-256 by 20 absolute points. Our analysis shows previously-proposed structured parameterizations for SSMs become largely redundant with pretraining. When evaluating architectures on supervised tasks, incorporating data-driven priors via pretraining is essential for reliable performance estimation. Ido Amos, Jonathan Berant, Ankit Gupta 0001 |
IJCAI | 2 |
| 2025 | In-Context Learning with Long-Context Models: An In-Depth ExplorationabstractAmanda Bertsch, Maor Ivgi, Emily Xiao, Uri Alon, Jonathan Berant, Matthew R. Gormley, Graham Neubig. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Amanda Bertsch, Maor Ivgi, Emily Xiao, Uri Alon 0002, Jonathan Berant, Matthew R. Gormley, Graham Neubig |
NAACL (Long Papers) | 5 |
| 2025 | Dolomites: Domain-Specific Long-Form Methodical TasksabstractAbstract Experts in various fields routinely perform methodical writing tasks to plan, organize, and report their work. From a clinician writing a differential diagnosis for a patient, to a teacher writing a lesson plan for students, these tasks are pervasive, requiring to methodically generate structured long-form output for a given input. We develop a typology of methodical tasks structured in the form of a task objective, procedure, input, and output, and introduce DoLoMiTes, a novel benchmark with specifications for 519 such tasks elicited from hundreds of experts from across 25 fields. Our benchmark further contains specific instantiations of methodical tasks with concrete input and output examples (1,857 in total) which we obtain by collecting expert revisions of up to 10 model-generated examples of each task. We use these examples to evaluate contemporary language models, highlighting that automating methodical tasks is a challenging long-form generation problem, as it requires performing complex inferences, while drawing upon the given context as well as domain knowledge. Our dataset is available at https://dolomites-benchmark.github.io/. Chaitanya Malaviya, Priyanka Agrawal, Kuzman Ganchev, Pranesh Srinivasan, Fantine Huot, Jonathan Berant, Mark Yatskar, Dipanjan Das 0001, Mirella Lapata, Christopher Alberti |
Trans. Assoc. Comput. Linguistics | 6 |
| 2024 | AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?abstractLanguage agents, built on top of language models (LMs), are systems that can interact with complex environments, such as the open web.In this work, we examine whether such agents can perform realistic and time-consuming tasks on the web, e.g., monitoring real-estate markets or locating relevant nearby businesses.We introduce ASSISTANTBENCH, a challenging new benchmark consisting of 214 realistic tasks that can be automatically evaluated, covering different scenarios and domains.We find that AS-SISTANTBENCH exposes the limitations of current systems, including language models and retrieval-augmented language models, as no model reaches an accuracy of more than 25 points.While closed-book LMs perform well in terms of accuracy, they exhibit low precision and tend to hallucinate facts.State-of-the-art web agents reach a score of near zero.Additionally, we introduce SEEPLANACT (SPA), a new web agent that significantly outperforms previous agents, and an ensemble of SPA and closed-book models reaches the best overall performance.Moreover, we analyze failures of current systems and highlight that open web navigation remains a major challenge. Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin, Ofir Press, Jonathan Berant |
EMNLP | 6 |
| 2024 | Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven PriorsabstractModeling long-range dependencies across sequences is a longstanding goal in machine learning and has led to architectures, such as state space models, that dramatically outperform Transformers on long sequences. However, these impressive empirical gains have been by and large demonstrated on benchmarks (e.g. Long Range Arena), where models are randomly initialized and trained to predict a target label from an input sequence. In this work, we show that random initialization leads to gross overestimation of the differences between architectures and that pretraining with standard denoising objectives, *using only the downstream task data*, leads to dramatic gains across multiple architectures and to very small gaps between Transformers and state space models (SSMs). In stark contrast to prior works, we find vanilla Transformers to match the performance of S4 on Long Range Arena when properly pretrained, and we improve the best reported results of SSMs on the PathX-256 task by 20 absolute points. Subsequently, we analyze the utility of previously-proposed structured parameterizations for SSMs and show they become mostly redundant in the presence of data-driven initialization obtained through pretraining. Our work shows that, when evaluating different architectures on supervised tasks, incorporation of data-driven priors via pretraining is essential for reliable performance estimation, and can be done efficiently. Ido Amos, Jonathan Berant, Ankit Gupta 0001 |
ICLR | 2 |
| 2024 | Making Retrieval-Augmented Language Models Robust to Irrelevant ContextabstractRetrieval-augmented language models (RALMs) hold promise to produce language understanding systems that are are factual, efficient, and up-to-date. An important desideratum of RALMs, is that retrieved information helps model performance when it is relevant, and does not harm performance when it is not. This is particularly important in multi-hop reasoning scenarios, where misuse of irrelevant evidence can lead to cascading errors. However, recent work has shown that retrieval augmentation can sometimes have a negative effect on performance. In this work, we present a thorough analysis on five open-domain question answering benchmarks, characterizing cases when retrieval reduces accuracy. We then propose two methods to mitigate this issue. First, a simple baseline that filters out retrieved passages that do not entail question-answer pairs according to a natural language inference (NLI) model. This is effective in preventing performance reduction, but at a cost of also discarding relevant passages. Thus, we propose a method for automatically generating data to fine-tune the language model to properly leverage retrieved passages, using a mix of relevant and irrelevant contexts at training time. We empirically show that even 1,000 examples suffice to train the model to be robust to irrelevant contexts while maintaining high performance on examples with relevant ones. Ori Yoran, Tomer Wolfson, Ori Ram, Jonathan Berant |
ICLR | 4 |
| 2024 | Transforming and Combining Rewards for Aligning Large Language ModelsabstractA common approach for aligning language models to human preferences is to first learn a reward model from preference data, and then use this reward model to update the language model. We study two closely related problems that arise in this approach. First, any monotone transformation of the reward model preserves preference ranking; is there a choice that is "better" than others? Second, we often wish to align language models to multiple properties: how should we combine multiple reward models? Using a probabilistic interpretation of the alignment procedure, we identify a natural choice for transformation for (the common case of) rewards learned from Bradley-Terry preference models. The derived transformation is straightforward: we apply a log-sigmoid function to the centered rewards, a method we term "LSC-transformation" (log-sigmoid-centered transformation). This transformation has two important properties. First, it emphasizes improving poorly-performing outputs, rather than outputs that already score well. This mitigates both underfitting (where some prompts are not improved) and reward hacking (where the model learns to exploit misspecification of the reward model). Second, it enables principled aggregation of rewards by linking summation to logical conjunction: the sum of transformed rewards corresponds to the probability that the output is "good" in all measured properties, in a sense we make precise. Experiments aligning language models to be both helpful and harmless using RLHF show substantial improvements over the baseline (non-transformed) approach. Chirag Nagpal, Jonathan Berant, Jacob Eisenstein, Alexander D'Amour, Oluwasanmi Koyejo, Victor Veitch |
ICML | 3 |
| 2024 | SEMQA: Semi-Extractive Multi-Source Question AnsweringabstractTal Schuster, Adam Lelkes, Haitian Sun, Jai Gupta, Jonathan Berant, William Cohen, Donald Metzler. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Tal Schuster, Ádám Dániel Lelkes, Haitian Sun, Jai Gupta 0001, Jonathan Berant, William W. Cohen, Donald Metzler |
NAACL-HLT | 5 |
| 2024 | Retrieval-Pretrained Transformer: Long-range Language Modeling with Self-retrievalabstractAbstract Retrieval-augmented language models (LMs) have received much attention recently. However, typically the retriever is not trained jointly as a native component of the LM, but added post-hoc to an already-pretrained LM, which limits the ability of the LM and the retriever to adapt to one another. In this work, we propose the Retrieval-Pretrained Transformer (RPT), an architecture and training procedure for jointly training a retrieval-augmented LM from scratch and applying it to the task of modeling long texts. Given a recently generated text chunk in a long document, the LM computes query representations, which are then used to retrieve earlier chunks in the document, located potentially tens of thousands of tokens before. Information from retrieved chunks is fused into the LM representations to predict the next target chunk. We train the retriever component with a semantic objective, where the goal is to retrieve chunks that increase the probability of the next chunk, according to a reference LM. We evaluate RPT on four long-range language modeling tasks, spanning books, code, and mathematical writing, and demonstrate that RPT improves retrieval quality and subsequently perplexity across the board compared to strong baselines. Ohad Rubin, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 2 |
| 2023 | Analyzing Transformers in Embedding SpaceabstractUnderstanding Transformer-based models has attracted significant attention, as they lie at the heart of recent technological advances across machine learning.While most interpretability methods rely on running models over inputs, recent work has shown that an inputindependent approach, where parameters are interpreted directly without a forward/backward pass is feasible for some Transformer parameters, and for two-layer attention networks.In this work, we present a conceptual framework where all parameters of a trained Transformer are interpreted by projecting them into the embedding space, that is, the space of vocabulary items they operate on.Focusing mostly on GPT-2 for this paper, we provide diverse evidence to support our argument.First, an empirical analysis showing that parameters of both pretrained and fine-tuned models can be interpreted in embedding space.Second, we present two applications of our framework: (a) aligning the parameters of different models that share a vocabulary, and (b) constructing a classifier without training by "translating" the parameters of a fine-tuned classifier to parameters of a different model that was only pretrained.Overall, our findings show that at least in part, we can abstract away model specifics and understand Transformers in the embedding space.EA EB Layer 18 Head 1 ('women', ' Marie') (' actresses', ' Marie') ('women', ' Anne') ('Women', ' Anne') ('woman', ' Marie') ('Women', ' Marie') Guy Dar, Mor Geva, Ankit Gupta 0001, Jonathan Berant |
ACL (1) | 4 |
| 2023 | Diverse Demonstrations Improve In-context Compositional GeneralizationabstractIn-context learning has shown great success in i.i.d semantic parsing splits, where the training and test sets are drawn from the same distribution.In this setup, models are typically prompted with demonstrations that are similar to the input utterance.However, in the setup of compositional generalization, where models are tested on outputs with structures that are absent from the training set, selecting similar demonstrations is insufficient, as often no example will be similar enough to the input.In this work, we propose a method to select diverse demonstrations that aims to collectively cover all of the structures required in the output program, in order to encourage the model to generalize to new structures from these demonstrations.We empirically show that combining diverse demonstrations with in-context learning substantially improves performance across three compositional generalization semantic parsing datasets in the pure in-context learning setup and when combined with finetuning. 1 * Equal contribution 1 Our code is available at: https://github.com/itayle/ diverse-demonstrations Question: What is the most populous state through which the mississippi runs?Q: What are the major cities in states through which the mississippi runs?A: major(city(loc_2( state(traverse_1(riverid('mississippi')))) )) Q: What are the cities in states through which the mississippi runs?A: city(loc_2( state(traverse_1(riverid('mississippi'))) )) Q: What is the most populous state through which the mississippi runs?(Output) most_populous( state(traverse_1(riverid('mississippi'))) ) (a) Similarity-Based Prompting Q: What are the major cities in states through which the mississippi runs?A: major(city(loc_2( state(traverse_1(riverid('mississippi')))) )) Q: What rivers flow through the state with the largest population?A: river(traverse_2( largest_one(population_1(state (all))))) Q: What is the most populous state through which the mississippi runs?(Output) largest_one(population_1(state(traverse_1(riverid('mississippi'))) )) (b) Diversity-Based Prompting (Ours) Itay Levy, Ben Bogin, Jonathan Berant |
ACL (1) | 3 |
| 2023 | What Are You Token About? Dense Retrieval as Distributions Over the VocabularyabstractOri Ram, Liat Bezalel, Adi Zicher, Yonatan Belinkov, Jonathan Berant, Amir Globerson. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ori Ram, Liat Bezalel, Adi Zicher, Yonatan Belinkov, Jonathan Berant, Amir Globerson |
ACL (1) | 5 |
| 2023 | Answering Questions by Meta-Reasoning over Multiple Chains of ThoughtabstractModern systems for multi-hop question answering (QA) typically break questions into a sequence of reasoning steps, termed chain-ofthought (CoT), before arriving at a final answer.Often, multiple chains are sampled and aggregated through a voting mechanism over the final answers, but the intermediate steps themselves are discarded.While such approaches improve performance, they do not consider the relations between intermediate steps across chains and do not provide a unified explanation for the predicted answer.We introduce Multi-Chain Reasoning (MCR), an approach which prompts large language models to meta-reason over multiple chains of thought, rather than aggregate their answers.MCR examines different reasoning chains, mixes information between them and selects the most relevant facts in generating an explanation and predicting the answer.MCR outperforms strong baselines on 7 multi-hop QA datasets.Moreover, our analysis reveals that MCR explanations exhibit high quality, enabling humans to verify its answers. Ori Yoran, Tomer Wolfson, Ben Bogin, Uri Katz, Daniel Deutch, Jonathan Berant |
EMNLP | 6 |
| 2023 | From Pixels to UI Actions: Learning to Follow Instructions via Graphical User InterfacesabstractMuch of the previous work towards digital agents for graphical user interfaces (GUIs) has relied on text-based representations (derived from HTML or other structured data sources), which are not always readily available. These input representations have been often coupled with custom, task-specific action spaces. This paper focuses on creating agents that interact with the digital world using the same conceptual interface that humans commonly use — via pixel-based screenshots and a generic action space corresponding to keyboard and mouse actions. Building upon recent progress in pixel-based pretraining, we show, for the first time, that it is possible for such agents to outperform human crowdworkers on the MiniWob++ benchmark of GUI-based instruction following tasks. Peter Shaw 0004, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, Kristina Toutanova |
NeurIPS | 4 |
| 2023 | Efficient Long-Text Understanding with Short-Text ModelsabstractAbstract Transformer-based pretrained language models (LMs) are ubiquitous across natural language understanding, but cannot be applied to long sequences such as stories, scientific articles, and long documents due to their quadratic complexity. While a myriad of efficient transformer variants have been proposed, they are typically based on custom implementations that require expensive pretraining from scratch. In this work, we propose SLED: SLiding-Encoder and Decoder, a simple approach for processing long sequences that re-uses and leverages battle-tested short-text pretrained LMs. Specifically, we partition the input into overlapping chunks, encode each with a short-text LM encoder and use the pretrained decoder to fuse information across chunks (fusion-in-decoder). We illustrate through controlled experiments that SLED offers a viable strategy for long text understanding and evaluate our approach on SCROLLS, a benchmark with seven datasets across a wide range of language understanding tasks. We find that SLED is competitive with specialized models that are up to 50x larger and require a dedicated and expensive pretraining step. Maor Ivgi, Uri Shaham 0002, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | Turning Tables: Generating Examples from Semi-structured Tables for Endowing Language Models with Reasoning SkillsabstractModels pre-trained with a language modeling objective possess ample world knowledge and language skills, but are known to struggle in tasks that require reasoning.In this work, we propose to leverage semi-structured tables, and automatically generate at scale questionparagraph pairs, where answering the question requires reasoning over multiple facts in the paragraph.We add a pre-training step over this synthetic data, which includes examples that require 16 different reasoning skills such as number comparison, conjunction, and fact composition.To improve data efficiency, we sample examples from reasoning skills where the model currently errs.We evaluate our approach on three reasoning-focused reading comprehension datasets, and show that our model, PReasM, substantially outperforms T5, a popular pre-trained encoder-decoder model.Moreover, sampling examples based on model errors leads to faster training and higher performance. Ori Yoran, Alon Talmor, Jonathan Berant |
ACL (1) | 3 |
| 2022 | SCROLLS: Standardized CompaRison Over Long Language SequencesabstractUri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, Omer Levy. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Uri Shaham 0002, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta 0001, Wenhan Xiong, Mor Geva, Jonathan Berant, Omer Levy |
EMNLP | 10 |
| 2022 | Unobserved Local Structures Make Compositional Generalization HardabstractWhile recent work has shown that sequence-tosequence models struggle to generalize to new compositions (termed compositional generalization), little is known on what makes compositional generalization hard on a particular test instance.In this work, we investigate the factors that make generalization to certain test instances challenging.We first substantiate that some examples are more difficult than others by showing that different models consistently fail or succeed on the same test instances.Then, we propose a criterion for the difficulty of an example: a test instance is hard if it contains a local structure that was not observed at training time.We formulate a simple decision rule based on this criterion and empirically show it predicts instance-level generalization well across 5 different semantic parsing datasets, substantially better than alternative decision rules.Last, we show local structures can be leveraged for creating difficult adversarial compositional splits and also to improve compositional generalization under limited training budgets by strategically selecting examples for the training set. Ben Bogin, Shivanshu Gupta, Jonathan Berant |
EMNLP | 3 |
| 2022 | Learning to Retrieve Passages without SupervisionabstractOri Ram, Gal Shachaf, Omer Levy, Jonathan Berant, Amir Globerson. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ori Ram, Gal Shachaf, Omer Levy, Jonathan Berant, Amir Globerson |
NAACL-HLT | 4 |
| 2022 | Learning To Retrieve Prompts for In-Context LearningabstractIn-context learning is a recent paradigm in natural language understanding, where a large pretrained language model (LM) observes a test instance and a few training examples as its input, and directly decodes the output without any update to its parameters.However, performance has been shown to strongly depend on the selected training examples (termed prompts).In this work, we propose an efficient method for retrieving prompts for in-context learning using annotated data and an LM.Given an inputoutput pair, we estimate the probability of the output given the input and a candidate training example as the prompt, and label training examples as positive or negative based on this probability.We then train an efficient dense retriever from this data, which is used to retrieve training examples as prompts at test time.We evaluate our approach on three sequence-tosequence tasks where language utterances are mapped to meaning representations, and find that it substantially outperforms prior work and multiple baselines across the board. Ohad Rubin, Jonathan Herzig, Jonathan Berant |
NAACL-HLT | 3 |
| 2022 | Diagonal State Spaces are as Effective as Structured State SpacesabstractModeling long range dependencies in sequential data is a fundamental step towards attaining human-level performance in many modalities such as text, vision, audio and video. While attention-based models are a popular and effective choice in modeling short-range interactions, their performance on tasks requiring long range reasoning has been largely inadequate. In an exciting result, Gu et al. (ICLR 2022) proposed the $\textit{Structured State Space}$ (S4) architecture delivering large gains over state-of-the-art models on several long-range tasks across various modalities. The core proposition of S4 is the parameterization of state matrices via a diagonal plus low rank structure, allowing efficient computation. In this work, we show that one can match the performance of S4 even without the low rank correction and thus assuming the state matrices to be diagonal. Our $\textit{Diagonal State Space}$ (DSS) model matches the performance of S4 on Long Range Arena tasks, speech classification on Speech Commands dataset, while being conceptually simpler and straightforward to implement. Ankit Gupta 0001, Albert Gu, Jonathan Berant |
NeurIPS | 3 |
| 2022 | Break, Perturb, Build: Automatic Perturbation of Reasoning Paths Through Question DecompositionabstractAbstract Recent efforts to create challenge benchmarks that test the abilities of natural language understanding models have largely depended on human annotations. In this work, we introduce the “Break, Perturb, Build” (BPB) framework for automatic reasoning-oriented perturbation of question-answer pairs. BPB represents a question by decomposing it into the reasoning steps that are required to answer it, symbolically perturbs the decomposition, and then generates new question-answer pairs. We demonstrate the effectiveness of BPB by creating evaluation sets for three reading comprehension (RC) benchmarks, generating thousands of high-quality examples without human intervention. We evaluate a range of RC models on our evaluation sets, which reveals large performance gaps on generated examples compared to the original data. Moreover, symbolic perturbations enable fine-grained analysis of the strengths and limitations of models. Last, augmenting the training data with examples generated by BPB helps close the performance gaps, without any drop on the original data distribution. Mor Geva, Tomer Wolfson, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | Span-based Semantic Parsing for Compositional GeneralizationabstractJonathan Herzig, Jonathan Berant. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jonathan Herzig, Jonathan Berant |
ACL/IJCNLP (1) | 2 |
| 2021 | Few-Shot Question Answering by Pretraining Span SelectionabstractOri Ram, Yuval Kirstain, Jonathan Berant, Amir Globerson, Omer Levy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ori Ram, Yuval Kirstain, Jonathan Berant, Amir Globerson, Omer Levy |
ACL/IJCNLP (1) | 3 |
| 2021 | BERTese: Learning to Speak to BERTabstractLarge pre-trained language models have been shown to encode large amounts of world and commonsense knowledge in their parameters, leading to substantial interest in methods for extracting that knowledge.In past work, knowledge was extracted by taking manuallyauthored queries and gathering paraphrases for them using a separate pipeline.In this work, we propose a method for automatically rewriting queries into "BERTese", a paraphrase query that is directly optimized towards better knowledge extraction.To encourage meaningful rewrites, we add auxiliary loss functions that encourage the query to correspond to actual language tokens.We empirically show our approach outperforms competing baselines, obviating the need for complex pipelines.Moreover, BERTese provides some insight into the type of language that helps language models perform knowledge extraction. Adi Haviv, Jonathan Berant, Amir Globerson |
EACL | 2 |
| 2021 | Evaluating the Evaluation of Diversity in Natural Language GenerationabstractDespite growing interest in natural language generation (NLG) models that produce diverse outputs, there is currently no principled method for evaluating the diversity of an NLG system.In this work, we propose a framework for evaluating diversity metrics.The framework measures the correlation between a proposed diversity metric and a diversity parameter, a single parameter that controls some aspect of diversity in generated text.For example, a diversity parameter might be a binary variable used to instruct crowdsourcing workers to generate text with either low or high content diversity.We demonstrate the utility of our framework by: (a) establishing best practices for eliciting diversity judgments from humans, (b) showing that humans substantially outperform automatic metrics in estimating content diversity, and (c) demonstrating that existing methods for controlling diversity by tuning a "decoding parameter" mostly affect form but not meaning.Our framework can advance the understanding of different diversity metrics, an essential step on the road towards better NLG systems. Guy Tevet, Jonathan Berant |
EACL | 2 |
| 2021 | COVR: A Test-Bed for Visually Grounded Compositional Generalization with Real ImagesabstractWhile interest in models that generalize at test time to new compositions has risen in recent years, benchmarks in the visually-grounded domain have thus far been restricted to synthetic images.In this work, we propose COVR, a new test-bed for visually-grounded compositional generalization with real images.To create COVR, we use real images annotated with scene graphs, and propose an almost fully automatic procedure for generating question-answer pairs along with a set of context images.COVR focuses on questions that require complex reasoning, including higherorder operations such as quantification and aggregation.Due to the automatic generation process, COVR facilitates the creation of compositional splits, where models at test time need to generalize to new concepts and compositions in a zero-or few-shot setting.We construct compositional splits using COVR and demonstrate a myriad of cases where state-ofthe-art pre-trained language-and-vision models struggle to compositionally generalize. Ben Bogin, Shivanshu Gupta, Matt Gardner 0001, Jonathan Berant |
EMNLP (1) | 4 |
| 2021 | What's in Your Head? Emergent Behaviour in Multi-Task Transformer ModelsabstractThe primary paradigm for multi-task training in natural language processing is to represent the input with a shared pre-trained language model, and add a small, thin network (head) per task.Given an input, a target head is the head that is selected for outputting the final prediction.In this work, we examine the behaviour of non-target heads, that is, the output of heads when given input that belongs to a different task than the one they were trained for.We find that non-target heads exhibit emergent behaviour, which may either explain the target task, or generalize beyond their original task.For example, in a numerical reasoning task, a span extraction head extracts from the input the arguments to a computation that results in a number generated by a target generative head.In addition, a summarization head that is trained with a target question answering head, outputs query-based summaries when given a question and a context from which the answer is to be extracted.This emergent behaviour suggests that multi-task training leads to nontrivial extrapolation of skills, which can be harnessed for interpretability and generalization. Mor Geva, Uri Katz, Aviv Ben-Arie, Jonathan Berant |
EMNLP (1) | 4 |
| 2021 | Transformer Feed-Forward Layers Are Key-Value MemoriesabstractFeed-forward layers constitute two-thirds of a transformer model's parameters, yet their role in the network remains under-explored.We show that feed-forward layers in transformerbased language models operate as key-value memories, where each key correlates with textual patterns in the training examples, and each value induces a distribution over the output vocabulary.Our experiments show that the learned patterns are human-interpretable, and that lower layers tend to capture shallow patterns, while upper layers learn more semantic ones.The values complement the keys' input patterns by inducing output distributions that concentrate probability mass on tokens likely to appear immediately after each pattern, particularly in the upper layers.Finally, we demonstrate that the output of a feed-forward layer is a composition of its memories, which is subsequently refined throughout the model's layers via residual connections to produce the final output distribution. Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy |
EMNLP (1) | 3 |
| 2021 | Value-aware Approximate AttentionabstractFollowing the success of dot-product attention in Transformers, numerous approximations have been recently proposed to address its quadratic complexity with respect to the input length.However, all approximations thus far have ignored the contribution of the value vectors to the quality of approximation.In this work, we argue that research efforts should be directed towards approximating the true output of the attention sub-layer, which includes the value vectors.We propose a valueaware objective, and show theoretically and empirically that an optimal approximation of a value-aware objective substantially outperforms an optimal approximation that ignores values, in the context of language modeling.Moreover, we show that the choice of kernel function for computing attention similarity can substantially affect the quality of sparse approximations, where kernel functions that are less skewed are more affected by the value vectors. Ankit Gupta 0001, Jonathan Berant |
EMNLP (1) | 2 |
| 2021 | Achieving Model Robustness through Discrete Adversarial TrainingabstractDiscrete adversarial attacks are symbolic perturbations to a language input that preserve the output label but lead to a prediction error.While such attacks have been extensively explored for the purpose of evaluating model robustness, their utility for improving robustness has been limited to offline augmentation only.Concretely, given a trained model, attacks are used to generate perturbed (adversarial) examples, and the model is re-trained exactly once.In this work, we address this gap and leverage discrete attacks for online augmentation, where adversarial examples are generated at every training step, adapting to the changing nature of the model.We propose (i) a new discrete attack, based on best-first search, and (ii) random sampling attacks that unlike prior work are not based on expensive search-based procedures.Surprisingly, we find that random sampling leads to impressive gains in robustness, outperforming the commonly-used offline augmentation, while leading to a speedup at training time of ∼10x.Furthermore, online augmentation with search-based attacks justifies the higher training cost, significantly improving robustness on three datasets.Last, we show that our new attack substantially improves robustness compared to prior methods. Maor Ivgi, Jonathan Berant |
EMNLP (1) | 2 |
| 2021 | Finding needles in a haystack: Sampling Structurally-diverse Training Sets from Synthetic Data for Compositional GeneralizationabstractModern semantic parsers suffer from two principal limitations.First, training requires expensive collection of utterance-program pairs.Second, semantic parsers fail to generalize at test time to new compositions/structures that have not been observed during training.Recent research has shown that automatic generation of synthetic utterance-program pairs can alleviate the first problem, but its potential for the second has thus far been under-explored.In this work, we investigate automatic generation of synthetic utterance-program pairs for improving compositional generalization in semantic parsing.Given a small training set of annotated examples and an "infinite" pool of synthetic examples, we select a subset of synthetic examples that are structurally-diverse and use them to improve compositional generalization.We evaluate our approach on a new split of the schema2QA dataset, and show that it leads to dramatic improvements in compositional generalization as well as moderate improvements in the traditional i.i.d setup.Moreover, structurally-diverse sampling achieves these improvements with as few as 5K examples, compared to 1M examples when sampling uniformly at random -a 200x improvement in data efficiency. Inbar Oren, Jonathan Herzig, Jonathan Berant |
EMNLP (1) | 3 |
| 2021 | Scene Graph tO Image Generation with Contextualized Object Layout RefinementabstractGenerating images from scene graphs is a challenging task that attracted substantial interest recently. Prior works have approached this task by generating an intermediate layout description of the target image. However, the representation of each object in the layout was generated independently, which resulted in high overlap, low coverage, and an overall blurry layout. We propose a novel method that alleviates these issues by generating the entire layout description gradually to improve inter-object dependency. We empirically show on the COCO-STUFF dataset that our approach improves the quality of both the intermediate layout and the final image. Our approach improves the layout coverage by almost 20 points, and drops object overlap to negligible amounts. Our code is available at github.com/yanivbenny/COLoR. Maor Ivgi, Yaniv Benny, Avichai Ben-David, Jonathan Berant, Lior Wolf |
ICIP | 4 |
| 2021 | MultiModalQA: complex question answering over text, tables and images
Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, Jonathan Berant |
ICLR | 9 |
| 2021 | SmBoP: Semi-autoregressive Bottom-up Semantic ParsingabstractThe de-facto standard decoding method for semantic parsing in recent years has been to autoregressively decode the abstract syntax tree of the target program using a top-down depthfirst traversal.In this work, we propose an alternative approach: a Semi-autoregressive Bottom-up Parser (SMBOP) that constructs at decoding step t the top-K sub-trees of height ≤ t.Our parser enjoys several benefits compared to top-down autoregressive parsing.From an efficiency perspective, bottom-up parsing allows to decode all sub-trees of a certain height in parallel, leading to logarithmic runtime complexity rather than linear.From a modeling perspective, a bottom-up parser learns representations for meaningful semantic sub-programs at each step, rather than for semantically-vacuous partial trees.We apply SMBOP on SPIDER, a challenging zero-shot semantic parsing benchmark, and show that SMBOP leads to a 2.2x speed-up in decoding time and a ∼5x speed-up in training time, compared to a semantic parser that uses autoregressive decoding.SMBOP obtains 71.1 denotation accuracy on SPIDER, establishing a new state-of-the-art, and 69.5 exact match, comparable to the 69.6 exact match of the autoregressive RAT-SQL+GRAPPA. Ohad Rubin, Jonathan Berant |
NAACL-HLT | 2 |
| 2021 | Latent Compositional Representations Improve Systematic Generalization in Grounded Question AnsweringabstractAbstract Answering questions that involve multi-step reasoning requires decomposing them and using the answers of intermediate steps to reach the final answer. However, state-of-the-art models in grounded question answering often do not explicitly perform decomposition, leading to difficulties in generalization to out-of-distribution examples. In this work, we propose a model that computes a representation and denotation for all question spans in a bottom-up, compositional manner using a CKY-style parser. Our model induces latent trees, driven by end-to-end (the answer) supervision only. We show that this inductive bias towards tree structures dramatically improves systematic generalization to out-of- distribution examples, compared to strong baselines on an arithmetic expressions benchmark as well as on C losure, a dataset that focuses on systematic generalization for grounded question answering. On this challenging dataset, our model reaches an accuracy of 96.1%, significantly higher than prior models that almost perfectly solve the task on a random, in-distribution split. Ben Bogin, Sanjay Subramanian, Matt Gardner 0001, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 4 |
| 2021 | Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning StrategiesabstractAbstract A key limitation in current datasets for multi-hop reasoning is that the required steps for answering the question are mentioned in it explicitly. In this work, we introduce StrategyQA, a question answering (QA) benchmark where the required reasoning steps are implicit in the question, and should be inferred using a strategy. A fundamental challenge in this setup is how to elicit such creative questions from crowdsourcing workers, while covering a broad range of potential strategies. We propose a data collection procedure that combines term-based priming to inspire annotators, careful control over the annotator population, and adversarial filtering for eliminating reasoning shortcuts. Moreover, we annotate each question with (1) a decomposition into reasoning steps for answering it, and (2) Wikipedia paragraphs that contain the answers to each step. Overall, StrategyQA includes 2,780 examples, each consisting of a strategy question, its decomposition, and evidence paragraphs. Analysis shows that questions in StrategyQA are short, topic-diverse, and cover a wide range of strategies. Empirically, we show that humans perform well (87%) on this task, while our best baseline reaches an accuracy of ∼ 66%. Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth 0001, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 6 |
| 2020 | Injecting Numerical Reasoning Skills into Language ModelsabstractLarge pre-trained language models (LMs) are known to encode substantial amounts of linguistic information.However, high-level reasoning skills, such as numerical reasoning, are difficult to learn from a language-modeling objective only.Consequently, existing models for numerical reasoning have used specialized architectures with limited flexibility.In this work, we show that numerical reasoning is amenable to automatic data generation, and thus one can inject this skill into pre-trained LMs, by generating large amounts of data, and training in a multi-task setup.We show that pre-training our model, GENBERT, on this data, dramatically improves performance on DROP (49.3 → 72.3 F 1 ), reaching performance that matches state-of-the-art models of comparable size, while using a simple and general-purpose encoder-decoder architecture.Moreover, GENBERT generalizes well to math word problem datasets, while maintaining high performance on standard RC tasks.Our approach provides a general recipe for injecting skills into large pre-trained LMs, whenever the skill is amenable to automatic data augmentation.* These authors contributed equally.(b) fine-tuning pre-trained LM numerical reasoning reading compr. Mor Geva, Ankit Gupta 0001, Jonathan Berant |
ACL | 3 |
| 2020 | Obtaining Faithful Interpretations from Compositional Neural NetworksabstractNeural module networks (NMNs) are a popular approach for modeling compositionality: they achieve high accuracy when applied to problems in language and vision, while reflecting the compositional structure of the problem in the network architecture.However, prior work implicitly assumed that the structure of the network modules, describing the abstract reasoning process, provides a faithful explanation of the model's reasoning; that is, that all modules perform their intended behaviour.In this work, we propose and conduct a systematic evaluation of the intermediate outputs of NMNs on NLVR2 and DROP, two datasets which require composing multiple reasoning steps.We find that the intermediate outputs differ from the expected output, illustrating that the network structure does not provide a faithful explanation of model behaviour.To remedy that, we train the model with auxiliary supervision and propose particular choices for module architecture that yield much better faithfulness, at a minimal cost to accuracy. Sanjay Subramanian, Ben Bogin, Nitish Gupta, Tomer Wolfson, Sameer Singh 0001, Jonathan Berant, Matt Gardner 0001 |
ACL | 6 |
| 2020 | A Simple and Effective Model for Answering Multi-span QuestionsabstractModels for reading comprehension (RC) commonly restrict their output space to the set of all single contiguous spans from the input, in order to alleviate the learning problem and avoid the need for a model that generates text explicitly.However, forcing an answer to be a single span can be restrictive, and some recent datasets also include multi-span questions, i.e., questions whose answer is a set of non-contiguous spans in the text.Naturally, models that return single spans cannot answer these questions.In this work, we propose a simple architecture for answering multi-span questions by casting the task as a sequence tagging problem, namely, predicting for each input token whether it should be part of the output or not.Our model substantially improves performance on span extraction questions from DROP and QUOREF by 9.9 and 5.5 EM points respectively. Elad Segal, Avia Efrat, Mor Shoham, Amir Globerson, Jonathan Berant |
EMNLP (1) | 5 |
| 2020 | Leap-Of-Thought: Teaching Pre-Trained Models to Systematically Reason Over Implicit KnowledgeabstractTo what extent can a neural network systematically reason over symbolic facts? Evidence suggests that large pre-trained language models (LMs) acquire some reasoning capacity, but this ability is difficult to control. Recently, it has been shown that Transformer-based models succeed in consistent reasoning over explicit symbolic facts, under a "closed-world" assumption. However, in an open-domain setup, it is desirable to tap into the vast reservoir of implicit knowledge already encoded in the parameters of pre-trained LMs. In this work, we provide a first demonstration that LMs can be trained to reliably perform systematic reasoning combining both implicit, pre-trained knowledge and explicit natural language statements. To do this, we describe a procedure for automatically generating datasets that teach a model new reasoning skills, and demonstrate that models learn to effectively perform inference which involves implicit taxonomic and world knowledge, chaining and counting. Finally, we show that "teaching" models to reason generalizes beyond the training distribution: they successfully compose the usage of multiple reasoning skills in single examples. Our work paves a path towards open-domain systems that constantly improve by interacting with users who can instantly correct a model by adding simple natural language statements. Alon Talmor, Oyvind Tafjord, Peter Clark, Yoav Goldberg, Jonathan Berant |
NeurIPS | 5 |
| 2020 | Differentiable Scene GraphsabstractReasoning about complex visual scenes involves perception of entities and their relations. Scene Graphs (SGs) provide a natural representation for reasoning tasks, by assigning labels to both entities (nodes) and relations (edges). Reasoning systems based on SGs are typically trained in a two-step procedure: first, a model is trained to predict SGs from images, and next a separate model is trained to reason based on the predicted SGs. However, it would seem preferable to train such systems in an end-to-end manner. The challenge, which we address here is that scene-graph representations are non-differentiable and therefore it isn't clear how to use them as intermediate components. Here we propose Differentiable Scene Graphs (DSGs), an image representation that is amenable to differentiable end-to-end optimization, and requires supervision only from the downstream tasks. DSGs provide a dense representation for all regions and pairs of regions, and do not spend modelling capacity on regions of the image that do not contain objects or relations of interest. We evaluate our model on the challenging task of identifying referring relationships (RR) in three benchmark datasets: Visual Genome, VRD and CLEVR. Using DSGs as an intermediate representation leads to new state-of-the-art performance. The full code is available at https://github.com/shikorab/DSG. Moshiko Raboh, Roei Herzig, Jonathan Berant, Gal Chechik, Amir Globerson |
WACV | 3 |
| 2020 | oLMpics - On what Language Model Pre-training CapturesabstractRecent success of pre-trained language models (LMs) has spurred widespread interest in the language capabilities that they possess. However, efforts to understand whether LM representations are useful for symbolic reasoning tasks have been limited and scattered. In this work, we propose eight reasoning tasks, which conceptually require operations such as comparison, conjunction, and composition. A fundamental challenge is to understand whether the performance of a LM on a task should be attributed to the pre-trained representations or to the process of fine-tuning on the task data. To address this, we propose an evaluation protocol that includes both zero-shot evaluation (no fine-tuning), as well as comparing the learning curve of a fine-tuned LM to the learning curve of multiple controls, which paints a rich picture of the LM capabilities. Our main findings are that: (a) different LMs exhibit qualitatively different reasoning abilities, e.g., RoBERTa succeeds in reasoning tasks where BERT fails completely; (b) LMs do not reason in an abstract manner and are context-dependent, e.g., while RoBERTa can compare ages, it can do so only when the ages are in the typical range of human ages; (c) On half of our reasoning tasks all models fail completely. Our findings and infrastructure can help future work on designing new datasets, models, and objective functions for pre-training. Alon Talmor, Yanai Elazar, Yoav Goldberg, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 4 |
| 2020 | Break It Down: A Question Understanding BenchmarkabstractUnderstanding natural language questions entails the ability to break down a question into the requisite steps for computing its answer. In this work, we introduce a Question Decomposition Meaning Representation (QDMR) for questions. QDMR constitutes the ordered list of steps, expressed through natural language, that are necessary for answering a question. We develop a crowdsourcing pipeline, showing that quality QDMRs can be annotated at scale, and release the Break dataset, containing over 83K pairs of questions and their QDMRs. We demonstrate the utility of QDMR by showing that (a) it can be used to improve open-domain question answering on the HotpotQA dataset, (b) it can be deterministically converted to a pseudo-SQL formal language, which can alleviate annotation in semantic parsing applications. Last, we use Break to train a sequence-to-sequence model with copying that parses questions into QDMR structures, and show that it substantially outperforms several natural baselines. Tomer Wolfson, Mor Geva, Ankit Gupta 0001, Yoav Goldberg, Matt Gardner 0001, Daniel Deutch, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 7 |
| 2019 | Representing Schema Structure with Graph Neural Networks for Text-to-SQL ParsingabstractResearch on parsing language to SQL has largely ignored the structure of the database (DB) schema, either because the DB was very simple, or because it was observed at both training and test time.In SPIDER, a recentlyreleased text-to-SQL dataset, new and complex DBs are given at test time, and so the structure of the DB schema can inform the predicted SQL query.In this paper, we present an encoder-decoder semantic parser, where the structure of the DB schema is encoded with a graph neural network, and this representation is later used at both encoding and decoding time.Evaluation shows that encoding the schema structure improves our parser accuracy from 33.8% to 39.4%, dramatically above the current state of the art, which is at 19.7%. Ben Bogin, Jonathan Berant, Matt Gardner 0001 |
ACL (1) | 2 |
| 2019 | MultiQA: An Empirical Investigation of Generalization and Transfer in Reading ComprehensionabstractA large number of reading comprehension (RC) datasets has been created recently, but little analysis has been done on whether they generalize to one another, and the extent to which existing datasets can be leveraged for improving performance on new ones.In this paper, we conduct such an investigation over ten RC datasets, training on one or more source RC datasets, and evaluating generalization, as well as transfer to a target RC dataset.We analyze the factors that contribute to generalization, and show that training on a source RC dataset and transferring to a target dataset substantially improves performance, even in the presence of powerful contextual representations from BERT (Devlin et al., 2019).We also find that training on multiple source RC datasets leads to robust generalization and transfer, and can reduce the cost of example collection for a new RC dataset.Following our analysis, we propose MULTIQA, a BERTbased model, trained on multiple RC datasets, which leads to state-of-the-art performance on five RC datasets.We share our infrastructure for the benefit of the research community. Alon Talmor, Jonathan Berant |
ACL (1) | 2 |
| 2019 | On the Limits of Learning to Actively Learn Semantic RepresentationsabstractOne of the goals of natural language understanding is to develop models that map sentences into meaning representations.However, training such models requires expensive annotation of complex structures, which hinders their adoption.Learning to actively-learn (LTAL) is a recent paradigm for reducing the amount of labeled data by learning a policy that selects which samples should be labeled.In this work, we examine LTAL for learning semantic representations, such as QA-SRL.We show that even an oracle policy that is allowed to pick examples that maximize performance on the test set (and constitutes an upper bound on the potential of LTAL), does not substantially improve performance compared to a random policy.We investigate factors that could explain this finding and show that a distinguishing characteristic of successful applications of LTAL is the interaction between optimization and the oracle policy selection process.In successful applications of LTAL, the examples selected by the oracle policy do not substantially depend on the optimization procedure, while in our setup the stochastic nature of optimization strongly affects the examples selected by the oracle.We conclude that the current applicability of LTAL for improving data efficiency in learning semantic meaning representations is limited. Omri Koshorek, Gabriel Stanovsky, Yichu Zhou, Vivek Srikumar, Jonathan Berant |
CoNLL | 5 |
| 2019 | Global Reasoning over Database Structures for Text-to-SQL ParsingabstractBen Bogin, Matt Gardner, Jonathan Berant. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ben Bogin, Matt Gardner 0001, Jonathan Berant |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding DatasetsabstractMor Geva, Yoav Goldberg, Jonathan Berant. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mor Geva, Yoav Goldberg, Jonathan Berant |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Don't paraphrase, detect! Rapid and Effective Data Collection for Semantic ParsingabstractJonathan Herzig, Jonathan Berant. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jonathan Herzig, Jonathan Berant |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Explaining Queries Over Web Tables to Non-expertsabstractDesigning a reliable natural language (NL) interface for querying tables has been a longtime goal of researchers in both the data management and natural language processing (NLP) communities. Such an interface receives as input an NL question, translates it into a formal query, executes the query and returns the results. Errors in the translation process are not uncommon, and users typically struggle to understand whether their query has been mapped correctly. We address this problem by explaining the obtained formal queries to non-expert users. We introduce novel query explanations that provide a graphic representation of the query cell-based provenance (in its execution on a given table). Our solution augments a state-of-the-art NL interface over web tables, enhancing it in both its training and deployment phase. Experiments, including a user study conducted on Amazon Mechanical Turk, show our solution to improve both the correctness and reliability of the NL interface. Jonathan Berant, Daniel Deutch, Amir Globerson, Tova Milo, Tomer Wolfson |
ICDE | 1 |
| 2019 | Neural network gradient-based learning of black-box function interfaces
Alon Jacovi, Guy Hadash, Einat Kermany, Boaz Carmeli, Ofer Lavi, George Kour, Jonathan Berant |
ICLR (Poster) | 7 |
| 2018 | Weakly Supervised Semantic Parsing with Abstract ExamplesabstractTraining semantic parsers from weak supervision (denotations) rather than strong supervision (programs) complicates training in two ways.First, a large search space of potential programs needs to be explored at training time to find a correct program.Second, spurious programs that accidentally lead to a correct denotation add noise to training.In this work we propose that in closed worlds with clear semantic types, one can substantially alleviate these problems by utilizing an abstract representation, where tokens in both the language utterance and program are lifted to an abstract form.We show that these abstractions can be defined with a handful of lexical rules and that they result in sharing between different examples that alleviates the difficulties in training.To test our approach, we develop the first semantic parser for CNLVR, a challenging visual reasoning dataset, where the search space is large and overcoming spuriousness is critical, because denotations are either TRUE or FALSE, and thus random programs are likely to lead to a correct denotation.Our method substantially improves performance, and reaches 82.5% accuracy, a 14.7% absolute accuracy improvement compared to the best reported accuracy so far. Omer Goldman, Veronica Latcinnik, Ehud Nave, Amir Globerson, Jonathan Berant |
ACL (1) | 5 |
| 2018 | Learning to Search in Long Documents Using Document StructureabstractReading comprehension models are based on recurrent neural networks that sequentially process the document tokens. As interest turns to answering more complex questions over longer documents, sequential reading of large portions of text becomes a substantial bottleneck. Inspired by how humans use document structure, we propose a novel framework for reading comprehension. We represent documents as trees, and model an agent that learns to interleave quick navigation through the document tree with more expensive answer extraction. To encourage exploration of the document tree, we propose a new algorithm, based on Deep Q-Network (DQN), which strategically samples tree nodes at training time. Empirically we find our algorithm improves question answering performance compared to DQN and a strong information-retrieval (IR) baseline, and that ensembling our model with the IR baseline results in further gains in performance. Mor Geva, Jonathan Berant |
COLING | 2 |
| 2018 | Decoupling Structure and Lexicon for Zero-Shot Semantic ParsingabstractBuilding a semantic parser quickly in a new domain is a fundamental challenge for conversational interfaces, as current semantic parsers require expensive supervision and lack the ability to generalize to new domains.In this paper, we introduce a zero-shot approach to semantic parsing that can parse utterances in unseen domains while only being trained on examples in other source domains.First, we map an utterance to an abstract, domainindependent, logical form that represents the structure of the logical form, but contains slots instead of KB constants.Then, we replace slots with KB constants via lexical alignment scores and global inference.Our model reaches an average accuracy of 53.4% on 7 domains in the OVERNIGHT dataset, substantially better than other zero-shot baselines, and performs as good as a parser trained on over 30% of the target domain examples. Jonathan Herzig, Jonathan Berant |
EMNLP | 2 |
| 2018 | Polyglot Semantic Parsing in APIsabstractKyle Richardson, Jonathan Berant, Jonas Kuhn. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Kyle Richardson 0001, Jonathan Berant, Jonas Kuhn |
NAACL-HLT | 2 |
| 2018 | The Web as a Knowledge-Base for Answering Complex QuestionsabstractAlon Talmor, Jonathan Berant. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Alon Talmor, Jonathan Berant |
NAACL-HLT | 2 |
| 2018 | Mapping Images to Scene Graphs with Permutation-Invariant Structured PredictionabstractMachine understanding of complex images is a key goal of artificial intelligence. One challenge underlying this task is that visual scenes contain multiple inter-related objects, and that global context plays an important role in interpreting the scene. A natural modeling framework for capturing such effects is structured prediction, which optimizes over complex labels, while modeling within-label interactions. However, it is unclear what principles should guide the design of a structured prediction model that utilizes the power of deep learning components. Here we propose a design principle for such architectures that follows from a natural requirement of permutation invariance. We prove a necessary and sufficient characterization for architectures that follow this invariance, and discuss its implication on model design. Finally, we show that the resulting model achieves new state of the art results on the Visual Genome scene graph labeling benchmark, outperforming all recent approaches. Roei Herzig, Moshiko Raboh, Gal Chechik, Jonathan Berant, Amir Globerson |
NeurIPS | 4 |
| 2018 | Memory Augmented Policy Optimization for Program Synthesis and Semantic ParsingabstractWe present Memory Augmented Policy Optimization (MAPO), a simple and novel way to leverage a memory buffer of promising trajectories to reduce the variance of policy gradient estimate. MAPO is applicable to deterministic environments with discrete actions, such as structured prediction and combinatorial optimization tasks. We express the expected return objective as a weighted sum of two terms: an expectation over the high-reward trajectories inside the memory buffer, and a separate expectation over trajectories outside the buffer. To make an efficient algorithm of MAPO, we propose: (1) memory weight clipping to accelerate and stabilize training; (2) systematic exploration to discover high-reward trajectories; (3) distributed sampling from inside and outside of the memory buffer to scale up training. MAPO improves the sample efficiency and robustness of policy gradient, especially on tasks with sparse rewards. We evaluate MAPO on weakly supervised program synthesis from natural language (semantic parsing). On the WikiTableQuestions benchmark, we improve the state-of-the-art by 2.6%, achieving an accuracy of 46.3%. On the WikiSQL benchmark, MAPO achieves an accuracy of 74.9% with only weak supervision, outperforming several strong baselines with full supervision. Our source code is available at https://goo.gl/TXBp4e Mohammad Norouzi 0002, Jonathan Berant, Quoc V. Le, Ni Lao |
NeurIPS | 3 |
| 2017 | Coarse-to-Fine Question Answering for Long DocumentsabstractEunsol Choi, Daniel Hewlett, Jakob Uszkoreit, Illia Polosukhin, Alexandre Lacoste, Jonathan Berant. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Eunsol Choi, Daniel Hewlett, Jakob Uszkoreit, Illia Polosukhin, Alexandre Lacoste, Jonathan Berant |
ACL (1) | 6 |
| 2017 | Neural Symbolic Machines: Learning Semantic Parsers on Freebase with Weak SupervisionabstractHarnessing the statistical power of neural networks to perform language understanding and symbolic reasoning is difficult, when it requires executing efficient discrete operations against a large knowledge-base.In this work, we introduce a Neural Symbolic Machine (NSM), which contains (a) a neural "programmer", i.e., a sequence-to-sequence model that maps language utterances to programs and utilizes a key-variable memory to handle compositionality (b) a symbolic "computer", i.e., a Lisp interpreter that performs program execution, and helps find good programs by pruning the search space.We apply REINFORCE to directly optimize the task reward of this structured prediction problem.To train with weak supervision and improve the stability of REINFORCE we augment it with an iterative maximum-likelihood training process.NSM outperforms the state-of-theart on the WEBQUESTIONSSP dataset when trained from question-answer pairs only, without requiring any feature engineering or domain-specific knowledge. Jonathan Berant, Quoc V. Le, Kenneth D. Forbus, Ni Lao |
ACL (1) | 2 |
| 2015 | Building a Semantic Parser OvernightabstractHow do we build a semantic parser in a new domain starting with zero training ex-amples? We introduce a new methodol-ogy for this setting: First, we use a simple grammar to generate logical forms paired with canonical utterances. The logical forms are meant to cover the desired set of compositional operators, and the canon-ical utterances are meant to capture the meaning of the logical forms (although clumsily). We then use crowdsourcing to paraphrase these canonical utterances into natural utterances. The resulting data is used to train the semantic parser. We fur-ther study the role of compositionality in the resulting paraphrases. Finally, we test our methodology on seven domains and show that we can build an adequate se-mantic parser in just a few hours. 1 Jonathan Berant, Percy Liang |
ACL (1) | 2 |
| 2015 | Efficient Global Learning of Entailment GraphsabstractEntailment rules between predicates are fundamental to many semantic-inference applications. Consequently, learning such rules has been an active field of research in recent years. Methods for learning entailment rules between predicates that take into account dependencies between different rules (e.g., entailment is a transitive relation) have been shown to improve rule quality, but suffer from scalability issues, that is, the number of predicates handled is often quite small. In this article, we present methods for learning transitive graphs that contain tens of thousands of nodes, where nodes represent predicates and edges correspond to entailment rules (termed entailment graphs). Our methods are able to scale to a large number of predicates by exploiting structural properties of entailment graphs such as the fact that they exhibit a “tree-like” property. We apply our methods on two data sets and demonstrate that our methods find high-quality solutions faster than methods proposed in the past, and moreover our methods for the first time scale to large graphs containing 20,000 nodes and more than 100,000 edges. Jonathan Berant, Noga Alon, Ido Dagan, Jacob Goldberger |
Comput. Linguistics | 1 |
| 2015 | Knowledge-Based Textual Inference via Parse-Tree TransformationsabstractTextual inference is an important component in many applications for understanding natural language. Classical approaches to textual inference rely on logical representations for meaning, which may be regarded as "external" to the natural language itself. However, practical applications usually adopt shallower lexical or lexical-syntactic representations, which correspond closely to language structure. In many cases, such approaches lack a principled meaning representation and inference framework. We describe an inference formalism that operates directly on language-based structures, particularly syntactic parse trees. New trees are generated by applying inference rules, which provide a unified representation for varying types of inferences. We use manual and automatic methods to generate these rules, which cover generic linguistic structures as well as specific lexical-based inferences. We also present a novel packed data-structure and a corresponding inference algorithm that allows efficient implementation of this formalism. We proved the correctness of the new algorithm and established its efficiency analytically and empirically. The utility of our approach was illustrated on two tasks: unsupervised relation extraction from a large corpus, and the Recognizing Textual Entailment (RTE) benchmarks. Roy Bar-Haim, Ido Dagan, Jonathan Berant |
J. Artif. Intell. Res. | 3 |
| 2015 | Imitation Learning of Agenda-based Semantic ParsersabstractSemantic parsers conventionally construct logical forms bottom-up in a fixed order, resulting in the generation of many extraneous partial logical forms. In this paper, we combine ideas from imitation learning and agenda-based parsing to train a semantic parser that searches partial logical forms in a more strategic order. Empirically, our parser reduces the number of constructed partial logical forms by an order of magnitude, and obtains a 6x-9x speedup over fixed-order parsing, while maintaining comparable accuracy. Jonathan Berant, Percy Liang |
Trans. Assoc. Comput. Linguistics | 1 |
| 2014 | Semantic Parsing via ParaphrasingabstractA central challenge in semantic parsing is handling the myriad ways in which knowledge base predicates can be expressed.Traditionally, semantic parsers are trained primarily from text paired with knowledge base information.Our goal is to exploit the much larger amounts of raw text not tied to any knowledge base.In this paper, we turn semantic parsing on its head.Given an input utterance, we first use a simple method to deterministically generate a set of candidate logical forms with a canonical realization in natural language for each.Then, we use a paraphrase model to choose the realization that best paraphrases the input, and output the corresponding logical form.We present two simple paraphrase models, an association model and a vector space model, and train them jointly from question-answer pairs.Our system PARASEMPRE improves stateof-the-art accuracies on two recently released question-answering datasets. Jonathan Berant, Percy Liang |
ACL (1) | 1 |
| 2014 | Modeling Biological Processes for Reading ComprehensionabstractJonathan Berant, Vivek Srikumar, Pei-Chun Chen, Abby Vander Linden, Brittany Harding, Brad Huang, Peter Clark, Christopher D. Manning. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014. Jonathan Berant, Vivek Srikumar, Pei-Chun Chen, Abby Vander Linden, Brittany Harding, Brad Huang, Peter Clark, Christopher D. Manning |
EMNLP | 1 |
| 2013 | A Two Level Model for Context Sensitive Inference Rules
Oren Melamud, Jonathan Berant, Ido Dagan, Jacob Goldberger, Idan Szpektor |
ACL (1) | 2 |
| 2013 | Semantic Parsing on Freebase from Question-Answer PairsabstractIn this paper, we train a semantic parser that scales up to Freebase.Instead of relying on annotated logical forms, which is especially expensive to obtain at large scale, we learn from question-answer pairs.The main challenge in this setting is narrowing down the huge number of possible logical predicates for a given question.We tackle this problem in two ways: First, we build a coarse mapping from phrases to predicates using a knowledge base and a large text corpus.Second, we use a bridging operation to generate additional predicates based on neighboring predicates.On the dataset of Cai and Yates (2013), despite not having annotated logical forms, our system outperforms their state-of-the-art parser.Additionally, we collected a more realistic and challenging dataset of question-answer pairs and improves over a natural baseline. Jonathan Berant, Andrew Chou, Roy Frostig, Percy Liang |
EMNLP | 1 |
| 2013 | Learning Biological Processes with Global ConstraintsabstractAju Thalappillil Scaria, Jonathan Berant, Mengqiu Wang, Peter Clark, Justin Lewis, Brittany Harding, Christopher D. Manning. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. 2013. Aju Thalappillil Scaria, Jonathan Berant, Mengqiu Wang, Peter Clark, Justin Lewis, Brittany Harding, Christopher D. Manning |
EMNLP | 2 |
| 2012 | Efficient Tree-based Approximation for Entailment Graph Learning
Jonathan Berant, Ido Dagan, Meni Adler, Jacob Goldberger |
ACL (1) | 1 |
| 2012 | Learning Verb Inference Rules from Linguistically-Motivated Evidence
Hila Weisman, Jonathan Berant, Idan Szpektor, Ido Dagan |
EMNLP-CoNLL | 2 |
| 2012 | Learning Entailment Relations by Global Graph Structure OptimizationabstractIdentifying entailment relations between predicates is an important part of applied semantic inference. In this article we propose a global inference algorithm that learns such entailment rules. First, we define a graph structure over predicates that represents entailment relations as directed edges. Then, we use a global transitivity constraint on the graph to learn the optimal set of edges, formulating the optimization problem as an Integer Linear Program. The algorithm is applied in a setting where, given a target concept, the algorithm learns on the fly all entailment rules between predicates that co-occur with this concept. Results show that our global algorithm improves performance over baseline algorithms by more than 10%. Jonathan Berant, Ido Dagan, Jacob Goldberger |
Comput. Linguistics | 1 |
| 2011 | Global Learning of Typed Entailment Rules
Jonathan Berant, Ido Dagan, Jacob Goldberger |
ACL | 1 |
| 2010 | Global Learning of Focused Entailment Graphs
Jonathan Berant, Ido Dagan, Jacob Goldberger |
ACL | 1 |
| 2010 | Recognising Entailment within Discourse
Shachar Mirkin, Jonathan Berant, Ido Dagan, Eyal Shnarch |
COLING | 2 |
| 2009 | A Compact Forest for Scalable Inference over Entailment and Paraphrase Rules
Roy Bar-Haim, Jonathan Berant, Ido Dagan |
EMNLP | 2 |