Peter Shaw 0004

dblp:217/1471-4 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
9since 2021 · last 2024
0000-0002-3187-8938ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 9 since 2021
YearPublicationVenuePosition
2024 BAGEL: Bootstrapping Agents by Guiding Exploration with Language
abstract
Following natural language instructions by executing actions in digital environments (e.g. web-browsers and REST APIs) is a challenging task for language model (LM) agents. Unfortunately, LM agents often fail to generalize to new environments without human demonstrations. This work presents BAGEL, a method for bootstrapping LM agents without human supervision. BAGEL converts a seed set of randomly explored trajectories to synthetic demonstrations via round-trips between two noisy LM components: an LM labeler which converts a trajectory into a synthetic instruction, and a zero-shot LM agent which maps the synthetic instruction into a refined trajectory. By performing these round-trips iteratively, BAGEL quickly converts the initial distribution of trajectories towards those that are well-described by natural language. We adapt the base LM agent at test time with in-context learning by retrieving relevant BAGEL demonstrations based on the instruction, and find improvements of over 2-13% absolute on ToolQA and MiniWob++, with up to 13x reduction in execution failures.
Shikhar Murty, Christopher D. Manning, Peter Shaw 0004, Mandar Joshi, Kenton Lee
ICML3
2023 QUEST: A Retrieval Dataset of Entity-Seeking Queries with Implicit Set Operations
abstract
Formulating selective information needs results in queries that implicitly specify set operations, such as intersection, union, and difference.For instance, one might search for "shorebirds that are not sandpipers" or "science-fiction films shot in England".To study the ability of retrieval systems to meet such information needs, we construct QUEST, a dataset of 3357 natural language queries with implicit set operations, that map to a set of entities corresponding to Wikipedia documents.The dataset challenges models to match multiple constraints mentioned in queries with corresponding evidence in documents and correctly perform various set operations.The dataset is constructed semi-automatically using Wikipedia category names.Queries are automatically composed from individual categories, then paraphrased and further validated for naturalness and fluency by crowdworkers.Crowdworkers also assess the relevance of entities based on their documents and highlight attribution of query constraints to spans of document text.We analyze several modern retrieval systems, finding that they often struggle on such queries.Queries involving negation and conjunction are particularly challenging and systems are further challenged with combinations of these operations. 1 * Work done during an internship at Google.retrieving an exhaustive document set, instead lim-65 iting annotation to the top few results of a baseline 66 information retrieval system.67 To analyze how well retrieval systems handle 68 such queries, we present QUEST, a dataset with 69 natural language queries from four domains, that 70 are mapped to relatively comprehensive sets of en-71 tities corresponding to Wikipedia pages.We use 72 Wikipedia categories and their mapping to entities 73 in Wikipedia as a building block for our dataset 74 construction approach, but do not allow access to 75 this semi-structured data source at inference time, 76 to simulate text-based retrieval.Wikipedia cate-77 gories represent a broad set of natural language 78 descriptions of entity properties and often corre-79 spond to selective information need queries that 80 could be plausibly issued by a search engine user 81 ([At least 90% of the time based on our filtering?]).82 The correspondence between property names and 83 document text is also often subtle and requires so-84 phisticated reasoning to determine relevance, rep-85 resenting the natural language inference challenge 86 inherent in the task, while the knowledge of cate-87 gory membership allows us to construct relatively 88 comprehensive sets of candidate entities for atomic 89 categories and their combinations.90 Our dataset construction process is outlined in 91 Figure 1.The base queries in our dataset are 92 semi-automatically generated using Wikipedia cat-93 egory names.To construct queries, we sample 94 category names and compose them into complex 95 queries by using pre-defined templates (for exam-96 ple, A \ B \ C).Next, we ask crowdworkers to 97 paraphrase these automatically generated queries, 98 while ensuring that the paraphrased queries are 99 fluent and clearly describe what a user could be 00 looking for.These are then validated for natural-01 ness and fluency by a different set of crowdworkers, 02 and filtered according to those criteria.Finally, for 03 a large subset of our dataset, we collect scalar rel-04 evance labels based on the entity documents, and 05 textual attributions mapping query constraints to 06 spans of document text, to aid the development of 07 systems that can make precise inferences based on 08 trusted sources.09 Performing well on this dataset requires sys-10 tems that can match query constraints with cor-11
Chaitanya Malaviya, Peter Shaw 0004, Ming-Wei Chang, Kenton Lee, Kristina Toutanova
ACL (1)2
2023 Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
abstract
Visually-situated language is ubiquitous---sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms. Perhaps due to this diversity, previous work has typically relied on domain-specific recipes with limited sharing of the underlying data, model architectures, and objectives. We present Pix2Struct, a pretrained image-to-text model for purely visual language understanding, which can be finetuned on tasks containing visually-situated language. Pix2Struct is pretrained by learning to parse masked screenshots of web pages into simplified HTML. The web, with its richness of visual elements cleanly reflected in the HTML structure, provides a large source of pretraining data well suited to the diversity of downstream tasks. Intuitively, this objective subsumes common pretraining signals such as OCR, language modeling, and image captioning. In addition to the novel pretraining strategy, we introduce a variable-resolution input representation and a more flexible integration of language and vision inputs, where language prompts such as questions are rendered directly on top of the input image. For the first time, we show that a single pretrained model can achieve state-of-the-art results in six out of nine tasks across four domains: documents, illustrations, user interfaces, and natural images.
Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu 0001, Julian Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw 0004, Ming-Wei Chang, Kristina Toutanova
ICML8
2023 From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces
abstract
Much of the previous work towards digital agents for graphical user interfaces (GUIs) has relied on text-based representations (derived from HTML or other structured data sources), which are not always readily available. These input representations have been often coupled with custom, task-specific action spaces. This paper focuses on creating agents that interact with the digital world using the same conceptual interface that humans commonly use — via pixel-based screenshots and a generic action space corresponding to keyboard and mouse actions. Building upon recent progress in pixel-based pretraining, we show, for the first time, that it is possible for such agents to outperform human crowdworkers on the MiniWob++ benchmark of GUI-based instruction following tasks.
Peter Shaw 0004, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, Kristina Toutanova
NeurIPS1
2022 Generate-and-Retrieve: Use Your Predictions to Improve Retrieval for Semantic Parsing
abstract
A common recent approach to semantic parsing augments sequence-to-sequence models by retrieving and appending a set of training samples, called exemplars. The effectiveness of this recipe is limited by the ability to retrieve informative exemplars that help produce the correct parse, which is especially challenging in low-resource settings. Existing retrieval is commonly based on similarity of query and exemplar inputs. We propose GandR, a retrieval procedure that retrieves exemplars for which outputs are also similar. GandR first generates a preliminary prediction with input-based retrieval. Then, it retrieves exemplars with outputs similar to the preliminary prediction which are used to generate a final prediction. GandR sets the state of the art on multiple low-resource semantic parsing tasks.
Yury Zemlyanskiy, Michiel de Jong, Joshua Ainslie, Panupong Pasupat, Peter Shaw 0004, Linlu Qiu, Sumit Sanghai, Fei Sha
COLING5
2022 Evaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing
abstract
Linlu Qiu, Peter Shaw, Panupong Pasupat, Tianze Shi, Jonathan Herzig, Emily Pitler, Fei Sha, Kristina Toutanova. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Linlu Qiu, Peter Shaw 0004, Panupong Pasupat, Tianze Shi, Jonathan Herzig, Emily Pitler, Fei Sha, Kristina Toutanova
EMNLP2
2022 Improving Compositional Generalization with Latent Structure and Data Augmentation
abstract
Linlu Qiu, Peter Shaw, Panupong Pasupat, Pawel Nowak, Tal Linzen, Fei Sha, Kristina Toutanova. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Linlu Qiu, Peter Shaw 0004, Panupong Pasupat, Pawel Krzysztof Nowak, Tal Linzen, Fei Sha, Kristina Toutanova
NAACL-HLT2
2021 Compositional Generalization and Natural Language Variation: Can a Semantic Parsing Approach Handle Both?
abstract
Peter Shaw, Ming-Wei Chang, Panupong Pasupat, Kristina Toutanova. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Peter Shaw 0004, Ming-Wei Chang, Panupong Pasupat, Kristina Toutanova
ACL/IJCNLP (1)1
2021 Systematic Generalization on gSCAN: What is Nearly Solved and What is Next?
abstract
We analyze the grounded SCAN (gSCAN) benchmark, which was recently proposed to study systematic generalization for grounded language understanding.First, we study which aspects of the original benchmark can be solved by commonly used methods in multimodal research.We find that a generalpurpose Transformer-based model with crossmodal attention achieves strong performance on a majority of the gSCAN splits, surprisingly outperforming more specialized approaches from prior work.Furthermore, our analysis suggests that many of the remaining errors reveal the same fundamental challenge in systematic generalization of linguistic constructs regardless of visual context.Second, inspired by this finding, we propose challenging new tasks for gSCAN by generating data to incorporate relations between objects in the visual environment.Finally, we find that current models are surprisingly data inefficient given the narrow scope of commands in gSCAN, suggesting another challenge for future work.
Linlu Qiu, Hexiang Hu, Bowen Zhang 0002, Peter Shaw 0004, Fei Sha
EMNLP (1)4
2020 Exploring Unexplored Generalization Challenges for Cross-Database Semantic Parsing
abstract
We study the task of cross-database semantic parsing (XSP), where a system that maps natural language utterances to executable SQL queries is evaluated on databases unseen during training. Recently, several datasets, including Spider, were proposed to support development of XSP systems. We propose a challenging evaluation setup for cross-database semantic parsing, focusing on variation across database schemas and in-domain language use. We re-purpose eight semantic parsing datasets that have been well-studied in the setting where in-domain training data is available, and instead use them as additional evaluation data for XSP systems instead. We build a system that performs well on Spider, and find that it struggles to generalize to our re-purposed set. Our setup uncovers several generalization challenges for cross-database semantic parsing, demonstrating the need to use and develop diverse training and evaluation datasets.
Alane Suhr, Ming-Wei Chang, Peter Shaw 0004, Kenton Lee
ACL3
2019 Generating Logical Forms from Graph Representations of Text and Entities
abstract
Structured information about entities is critical for many semantic parsing tasks.We present an approach that uses a Graph Neural Network (GNN) architecture to incorporate information about relevant entities and their relations during parsing.Combined with a decoder copy mechanism, this approach provides a conceptually simple mechanism to generate logical forms with entities.We demonstrate that this approach is competitive with the stateof-the-art across several tasks without pretraining, and outperforms existing approaches when combined with BERT pre-training.
Peter Shaw 0004, Philip Massey, Angelica Chen, Francesco Piccinno, Yasemin Altun
ACL (1)1
2019 Answering Conversational Questions on Structured Data without Logical Forms
abstract
Thomas Mueller, Francesco Piccinno, Peter Shaw, Massimo Nicosia, Yasemin Altun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Thomas Müller 0009, Francesco Piccinno, Peter Shaw 0004, Massimo Nicosia, Yasemin Altun
EMNLP/IJCNLP (1)3