Ben Bogin

dblp:202/2034 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
10since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 10 since 2021
YearPublicationVenuePosition
2025 DataDecide: How to Predict Best Pretraining Data with Small Experiments
abstract
Because large language models are expensive to pretrain on different datasets, using smaller-scale experiments to decide on data is crucial for reducing costs. Which benchmarks and methods of making decisions from observed performance at small scale most accurately predict the datasets that yield the best large models? To empower open exploration of this question, we release models, data, and evaluations in DataDecide—the most extensive open suite of models over differences in data and scale. We conduct controlled pretraining experiments across 25 corpora with differing sources, deduplication, and filtering up to 100B tokens, model sizes up to 1B parameters, and 3 random seeds. We find that the ranking of models at a single, small size (e.g., 150M parameters) is a strong baseline for predicting best models at our larger target scale (1B) ($\tilde$ 80% of comparisons correct). No scaling law methods among 8 baselines exceed the compute-decision frontier of single-scale predictions, but DataDecide can measure improvement in future scaling laws. We also identify that using continuous likelihood metrics as proxies in small experiments makes benchmarks including MMLU, ARC, HellaSwag, MBPP, and HumanEval $>$ 80% predictable at the target 1B scale with just 0.01% of the compute.
Ian Magnusson, Nguyen Tai, Ben Bogin, David Heineman, Jena D. Hwang, Luca Soldaini, Akshita Bhagia, Jiacheng Liu 0010, Dirk Groeneveld, Oyvind Tafjord, Noah A. Smith, Pang Wei Koh, Jesse Dodge
ICML3
2024 Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
abstract
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar 0009, Li Lucy, Xinxi Lyu, Nathan Lambert 0001, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson 0001, Shannon Shen 0001, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh 0001, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo
ACL (1)7
2024 SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
abstract
Ben Bogin, Kejuan Yang, Shashank Gupta, Kyle Richardson, Erin Bransom, Peter Clark, Ashish Sabharwal, Tushar Khot. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Ben Bogin, Kejuan Yang, Kyle Richardson 0001, Erin Bransom, Peter Clark, Ashish Sabharwal, Tushar Khot
EMNLP1
2024 AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?
abstract
Language agents, built on top of language models (LMs), are systems that can interact with complex environments, such as the open web.In this work, we examine whether such agents can perform realistic and time-consuming tasks on the web, e.g., monitoring real-estate markets or locating relevant nearby businesses.We introduce ASSISTANTBENCH, a challenging new benchmark consisting of 214 realistic tasks that can be automatically evaluated, covering different scenarios and domains.We find that AS-SISTANTBENCH exposes the limitations of current systems, including language models and retrieval-augmented language models, as no model reaches an accuracy of more than 25 points.While closed-book LMs perform well in terms of accuracy, they exhibit low precision and tend to hallucinate facts.State-of-the-art web agents reach a score of near zero.Additionally, we introduce SEEPLANACT (SPA), a new web agent that significantly outperforms previous agents, and an ensemble of SPA and closed-book models reaches the best overall performance.Moreover, we analyze failures of current systems and highlight that open web navigation remains a major challenge.
Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin, Ofir Press, Jonathan Berant
EMNLP4
2024 Leveraging Code to Improve In-Context Learning for Semantic Parsing
abstract
Ben Bogin, Shivanshu Gupta, Peter Clark, Ashish Sabharwal. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Ben Bogin, Shivanshu Gupta, Peter Clark, Ashish Sabharwal
NAACL-HLT1
2023 Diverse Demonstrations Improve In-context Compositional Generalization
abstract
In-context learning has shown great success in i.i.d semantic parsing splits, where the training and test sets are drawn from the same distribution.In this setup, models are typically prompted with demonstrations that are similar to the input utterance.However, in the setup of compositional generalization, where models are tested on outputs with structures that are absent from the training set, selecting similar demonstrations is insufficient, as often no example will be similar enough to the input.In this work, we propose a method to select diverse demonstrations that aims to collectively cover all of the structures required in the output program, in order to encourage the model to generalize to new structures from these demonstrations.We empirically show that combining diverse demonstrations with in-context learning substantially improves performance across three compositional generalization semantic parsing datasets in the pure in-context learning setup and when combined with finetuning. 1 * Equal contribution 1 Our code is available at: https://github.com/itayle/ diverse-demonstrations Question: What is the most populous state through which the mississippi runs?Q: What are the major cities in states through which the mississippi runs?A: major(city(loc_2( state(traverse_1(riverid('mississippi')))) )) Q: What are the cities in states through which the mississippi runs?A: city(loc_2( state(traverse_1(riverid('mississippi'))) )) Q: What is the most populous state through which the mississippi runs?(Output) most_populous( state(traverse_1(riverid('mississippi'))) ) (a) Similarity-Based Prompting Q: What are the major cities in states through which the mississippi runs?A: major(city(loc_2( state(traverse_1(riverid('mississippi')))) )) Q: What rivers flow through the state with the largest population?A: river(traverse_2( largest_one(population_1(state (all))))) Q: What is the most populous state through which the mississippi runs?(Output) largest_one(population_1(state(traverse_1(riverid('mississippi'))) )) (b) Diversity-Based Prompting (Ours)
Itay Levy, Ben Bogin, Jonathan Berant
ACL (1)2
2023 Answering Questions by Meta-Reasoning over Multiple Chains of Thought
abstract
Modern systems for multi-hop question answering (QA) typically break questions into a sequence of reasoning steps, termed chain-ofthought (CoT), before arriving at a final answer.Often, multiple chains are sampled and aggregated through a voting mechanism over the final answers, but the intermediate steps themselves are discarded.While such approaches improve performance, they do not consider the relations between intermediate steps across chains and do not provide a unified explanation for the predicted answer.We introduce Multi-Chain Reasoning (MCR), an approach which prompts large language models to meta-reason over multiple chains of thought, rather than aggregate their answers.MCR examines different reasoning chains, mixes information between them and selects the most relevant facts in generating an explanation and predicting the answer.MCR outperforms strong baselines on 7 multi-hop QA datasets.Moreover, our analysis reveals that MCR explanations exhibit high quality, enabling humans to verify its answers.
Ori Yoran, Tomer Wolfson, Ben Bogin, Uri Katz, Daniel Deutch, Jonathan Berant
EMNLP3
2022 Unobserved Local Structures Make Compositional Generalization Hard
abstract
While recent work has shown that sequence-tosequence models struggle to generalize to new compositions (termed compositional generalization), little is known on what makes compositional generalization hard on a particular test instance.In this work, we investigate the factors that make generalization to certain test instances challenging.We first substantiate that some examples are more difficult than others by showing that different models consistently fail or succeed on the same test instances.Then, we propose a criterion for the difficulty of an example: a test instance is hard if it contains a local structure that was not observed at training time.We formulate a simple decision rule based on this criterion and empirically show it predicts instance-level generalization well across 5 different semantic parsing datasets, substantially better than alternative decision rules.Last, we show local structures can be leveraged for creating difficult adversarial compositional splits and also to improve compositional generalization under limited training budgets by strategically selecting examples for the training set.
Ben Bogin, Shivanshu Gupta, Jonathan Berant
EMNLP1
2021 COVR: A Test-Bed for Visually Grounded Compositional Generalization with Real Images
abstract
While interest in models that generalize at test time to new compositions has risen in recent years, benchmarks in the visually-grounded domain have thus far been restricted to synthetic images.In this work, we propose COVR, a new test-bed for visually-grounded compositional generalization with real images.To create COVR, we use real images annotated with scene graphs, and propose an almost fully automatic procedure for generating question-answer pairs along with a set of context images.COVR focuses on questions that require complex reasoning, including higherorder operations such as quantification and aggregation.Due to the automatic generation process, COVR facilitates the creation of compositional splits, where models at test time need to generalize to new concepts and compositions in a zero-or few-shot setting.We construct compositional splits using COVR and demonstrate a myriad of cases where state-ofthe-art pre-trained language-and-vision models struggle to compositionally generalize.
Ben Bogin, Shivanshu Gupta, Matt Gardner 0001, Jonathan Berant
EMNLP (1)1
2021 Latent Compositional Representations Improve Systematic Generalization in Grounded Question Answering
abstract
Abstract Answering questions that involve multi-step reasoning requires decomposing them and using the answers of intermediate steps to reach the final answer. However, state-of-the-art models in grounded question answering often do not explicitly perform decomposition, leading to difficulties in generalization to out-of-distribution examples. In this work, we propose a model that computes a representation and denotation for all question spans in a bottom-up, compositional manner using a CKY-style parser. Our model induces latent trees, driven by end-to-end (the answer) supervision only. We show that this inductive bias towards tree structures dramatically improves systematic generalization to out-of- distribution examples, compared to strong baselines on an arithmetic expressions benchmark as well as on C losure, a dataset that focuses on systematic generalization for grounded question answering. On this challenging dataset, our model reaches an accuracy of 96.1%, significantly higher than prior models that almost perfectly solve the task on a random, in-distribution split.
Ben Bogin, Sanjay Subramanian, Matt Gardner 0001, Jonathan Berant
Trans. Assoc. Comput. Linguistics1
2020 Obtaining Faithful Interpretations from Compositional Neural Networks
abstract
Neural module networks (NMNs) are a popular approach for modeling compositionality: they achieve high accuracy when applied to problems in language and vision, while reflecting the compositional structure of the problem in the network architecture.However, prior work implicitly assumed that the structure of the network modules, describing the abstract reasoning process, provides a faithful explanation of the model's reasoning; that is, that all modules perform their intended behaviour.In this work, we propose and conduct a systematic evaluation of the intermediate outputs of NMNs on NLVR2 and DROP, two datasets which require composing multiple reasoning steps.We find that the intermediate outputs differ from the expected output, illustrating that the network structure does not provide a faithful explanation of model behaviour.To remedy that, we train the model with auxiliary supervision and propose particular choices for module architecture that yield much better faithfulness, at a minimal cost to accuracy.
Sanjay Subramanian, Ben Bogin, Nitish Gupta, Tomer Wolfson, Sameer Singh 0001, Jonathan Berant, Matt Gardner 0001
ACL2
2019 Representing Schema Structure with Graph Neural Networks for Text-to-SQL Parsing
abstract
Research on parsing language to SQL has largely ignored the structure of the database (DB) schema, either because the DB was very simple, or because it was observed at both training and test time.In SPIDER, a recentlyreleased text-to-SQL dataset, new and complex DBs are given at test time, and so the structure of the DB schema can inform the predicted SQL query.In this paper, we present an encoder-decoder semantic parser, where the structure of the DB schema is encoded with a graph neural network, and this representation is later used at both encoding and decoding time.Evaluation shows that encoding the schema structure improves our parser accuracy from 33.8% to 39.4%, dramatically above the current state of the art, which is at 19.7%.
Ben Bogin, Jonathan Berant, Matt Gardner 0001
ACL (1)1
2019 Global Reasoning over Database Structures for Text-to-SQL Parsing
abstract
Ben Bogin, Matt Gardner, Jonathan Berant. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ben Bogin, Matt Gardner 0001, Jonathan Berant
EMNLP/IJCNLP (1)1
2018 Towards an argumentative content search engine using weak supervision
abstract
Searching for sentences containing claims in a large text corpus is a key component in developing an argumentative content search engine. Previous works focused on detecting claims in a small set of documents or within documents enriched with argumentative content. However, pinpointing relevant claims in massive unstructured corpora, received little attention. A step in this direction was taken in (Levy et al. 2017), where the authors suggested using a weak signal to develop a relatively strict query for claim–sentence detection. Here, we leverage this work to define weak signals for training DNNs to obtain significantly greater performance. This approach allows to relax the query and increase the potential coverage. Our results clearly indicate that the system is able to successfully generalize from the weak signal, outperforming previously reported results in terms of both precision and coverage. Finally, we adapt our system to solve a recent argument mining task of identifying argumentative sentences in Web texts retrieved from heterogeneous sources, and obtain F1 scores comparable to the supervised baseline.
Ran Levy 0001, Ben Bogin, Shai Gretz, Ranit Aharonov, Noam Slonim
COLING2