Matt Gardner 0001

dblp:00/8046 · also Matthew Gardner 0001 · DBLP profile ↗
← Back
48ranked-venue papers
6as first author
16since 2021 · last 2022
0000-0001-8458-1727ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 6 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author
YearPublicationVenuePosition
2022 A Meta-framework for Spatiotemporal Quantity Extraction from Text
abstract
News events are often associated with quantities (e.g., the number of COVID-19 patients or the number of arrests in a protest), and it is often important to extract their type, time, and location from unstructured text in order to analyze these quantity events.This paper thus formulates the NLP problem of spatiotemporal quantity extraction, and proposes the first meta-framework for solving it.This meta-framework contains a formalism that decomposes the problem into several information extraction tasks, a shareable crowdsourcing pipeline, and transformer-based baseline models.We demonstrate the meta-framework in three domains-the COVID-19 pandemic, Black Lives Matter protests, and 2020 California wildfires-to show that the formalism is general and extensible, the crowdsourcing pipeline facilitates fast and high-quality data annotation, and the baseline system can handle spatiotemporal quantity extraction well enough to be practically useful.We release all resources for future research on this topic.1
Qiang Ning, Ben Zhou, Hao Wu 0034, Haoruo Peng, Chuchu Fan, Matt Gardner 0001
ACL (1)6
2022 Tailor: Generating and Perturbing Text with Semantic Controls
abstract
Controlled text perturbation is useful for evaluating and improving model generalizability.However, current techniques rely on training a model for every target perturbation, which is expensive and hard to generalize.We present Tailor, a semantically-controlled text generation system.Tailor builds on a pretrained seq2seq model and produces textual outputs conditioned on control codes derived from semantic representations.We craft a set of operations to modify the control codes, which in turn steer generation towards targeted attributes.These operations can be further composed into higher-level ones, allowing for flexible perturbation strategies.We demonstrate the effectiveness of these perturbations in multiple applications.First, we use Tailor to automatically create high-quality contrast sets for four distinct natural language processing (NLP) tasks.These contrast sets contain fewer spurious artifacts and are complementary to manually annotated ones in their lexical diversity.Second, we show that Tailor perturbations can improve model generalization through data augmentation.Perturbing just ∼2% of training data leads to a 5.8-point gain on an NLI challenge set measuring reliance on syntactic heuristics.
Alexis Ross, Sherry Tongshuang Wu, Hao Peng 0009, Matthew E. Peters, Matt Gardner 0001
ACL (1)5
2022 ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension
abstract
Sanjay Subramanian, William Merrill, Trevor Darrell, Matt Gardner, Sameer Singh, Anna Rohrbach. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Sanjay Subramanian, William Merrill, Trevor Darrell, Matt Gardner 0001, Sameer Singh 0001, Anna Rohrbach
ACL (1)4
2022 Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets
abstract
Natural language processing models often exploit spurious correlations between taskindependent features and labels in datasets to perform well only within the distributions they are trained on, while not generalising to different task distributions.We propose to tackle this problem by generating a debiased version of a dataset, which can then be used to train a debiased, off-the-shelf model, by simply replacing its training data.Our approach consists of 1) a method for training data generators to generate high-quality, label-consistent data samples; and 2) a filtering mechanism for removing data points that contribute to spurious correlations, measured in terms of z-statistics.We generate debiased versions of the SNLI and MNLI datasets, 1 and we evaluate on a large suite of debiased, outof-distribution, and adversarial test sets.Results show that models trained on our debiased datasets generalise better than those trained on the original datasets in all settings.On the majority of the datasets, our method outperforms or performs comparably to previous state-ofthe-art debiasing strategies, and when combined with an orthogonal technique, productof-experts, it improves further and outperforms previous best results of SNLI-hard and MNLI-hard.* Work done while at the Allen Institute for AI. 1 All our code and the generated datasets are available at https://github.com/jimmycode/ gen-debiased-nli.Generator (Section 2 & Section 4.1) sample z-filter (Section 3 & Section 4.2)
Yuxiang Wu, Matt Gardner 0001, Pontus Stenetorp, Pradeep Dasigi
ACL (1)2
2022 Successive Prompting for Decomposing Complex Questions
abstract
Answering complex questions that require making latent decisions is a challenging task, especially when limited supervision is available.Recent works leverage the capabilities of large language models (LMs) to perform complex question answering in a few-shot setting by demonstrating how to output intermediate rationalizations while solving the complex question in a single pass.We introduce "Successive Prompting", where we iteratively break down a complex task into a simple task, solve it, and then repeat the process until we get the final solution.Successive prompting decouples the supervision for decomposing complex questions from the supervision for answering simple questions, allowing us to (1) have multiple opportunities to query in-context examples at each reasoning step (2) learn question decomposition separately from question answering, including using synthetic data, and (3) use bespoke (fine-tuned) components for reasoning steps where a large LM does not perform well.The intermediate supervision is typically manually written, which can be expensive to collect.We introduce a way to generate a synthetic dataset which can be used to bootstrap a model's ability to decompose and answer intermediate questions.Our best model (with successive prompting) achieves an improvement of ∼5% absolute F1 on a few-shot version of the DROP dataset when compared with a stateof-the-art model with the same supervision.
Dheeru Dua, Shivanshu Gupta, Sameer Singh 0001, Matt Gardner 0001
EMNLP4
2022 CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation
abstract
The full power of human language-based communication cannot be realized without negation.All human languages have some form of negation.Despite this, negation remains a challenging phenomenon for current natural language understanding systems.To facilitate the future development of models that can process negation effectively, we present CONDAQA, the first English reading comprehension dataset which requires reasoning about the implications of negated statements in paragraphs.We collect paragraphs with diverse negation cues, then have crowdworkers ask questions about the implications of the negated statement in the passage.We also have workers make three kinds of edits to the passage-paraphrasing the negated statement, changing the scope of the negation, and reversing the negation-resulting in clusters of question-answer pairs that are difficult for models to answer with spurious shortcuts.CONDAQA features 14,182 questionanswer pairs with over 200 unique negation cues and is challenging for current state-ofthe-art models.The best performing model on CONDAQA (UNIFIEDQA-V2-3B) achieves only 42% on our consistency metric, well below human performance which is 81%.We release our dataset, along with fully-finetuned, few-shot, and zero-shot evaluations, to facilitate the development of future NLP methods that work on negated language.
Abhilasha Ravichander, Matt Gardner 0001, Ana Marasovic
EMNLP2
2022 Evidentiality-guided Generation for Knowledge-Intensive NLP Tasks
abstract
Retrieval-augmented generation models have shown state-of-the-art performance across many knowledge-intensive NLP tasks such as open-domain question answering and fact verification.These models are trained to generate a final output given retrieved passages that can be irrelevant to an input query, leading to learning spurious cues or memorization.This work introduces a method to incorporate evidentiality of passages-whether a passage contains correct evidence to support the outputinto training the generator.We introduce a multi-task learning framework to jointly generate the final output and predict the evidentiality of each passage.Furthermore, we introduce a new task-agnostic method for obtaining high-quality silver evidentiality labels, addressing the issues of gold evidentiality labels being unavailable in most domains.Our experiments on five datasets across three knowledgeintensive tasks show that our new evidentialityguided generator significantly outperforms its direct counterpart on all of them, and advances the state of the art on three of them.Our analysis shows that the multi-task learning and silver evidentiality mining play key roles.
Akari Asai, Matt Gardner 0001, Hannaneh Hajishirzi
NAACL-HLT2
2021 Competency Problems: On Finding and Removing Artifacts in Language Data
abstract
Much recent work in NLP has documented dataset artifacts, bias, and spurious correlations between input features and output labels.However, how to tell which features have "spurious" instead of legitimate correlations is typically left unspecified.In this work we argue that for complex language understanding tasks, all simple feature correlations are spurious, and we formalize this notion into a class of problems which we call competency problems.For example, the word "amazing" on its own should not give information about a sentiment label independent of the context in which it appears, which could include negation, metaphor, sarcasm, etc.We theoretically analyze the difficulty of creating data for competency problems when human bias is taken into account, showing that realistic datasets will increasingly deviate from competency problems as dataset size increases.This analysis gives us a simple statistical test for dataset artifacts, which we use to show more subtle biases than were described in prior work, including demonstrating that models are inappropriately affected by these less extreme biases.Our theoretical treatment of this problem also allows us to analyze proposed solutions, such as making local edits to dataset instances, and to give recommendations for future data collection and model design efforts that target competency problems.
Matt Gardner 0001, William Merrill, Jesse Dodge, Matthew E. Peters, Alexis Ross, Sameer Singh 0001, Noah A. Smith
EMNLP (1)1
2021 COVR: A Test-Bed for Visually Grounded Compositional Generalization with Real Images
abstract
While interest in models that generalize at test time to new compositions has risen in recent years, benchmarks in the visually-grounded domain have thus far been restricted to synthetic images.In this work, we propose COVR, a new test-bed for visually-grounded compositional generalization with real images.To create COVR, we use real images annotated with scene graphs, and propose an almost fully automatic procedure for generating question-answer pairs along with a set of context images.COVR focuses on questions that require complex reasoning, including higherorder operations such as quantification and aggregation.Due to the automatic generation process, COVR facilitates the creation of compositional splits, where models at test time need to generalize to new concepts and compositions in a zero-or few-shot setting.We construct compositional splits using COVR and demonstrate a myriad of cases where state-ofthe-art pre-trained language-and-vision models struggle to compositionally generalize.
Ben Bogin, Shivanshu Gupta, Matt Gardner 0001, Jonathan Berant
EMNLP (1)3
2021 Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
abstract
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner 0001
EMNLP (1)8
2021 Learning with Instance Bundles for Reading Comprehension
abstract
When training most modern reading comprehension models, all the questions associated with a context are treated as being independent from each other.However, closely related questions and their corresponding answers are not independent, and leveraging these relationships could provide a strong supervision signal to a model.Drawing on ideas from contrastive estimation, we introduce several new supervision losses that compare question-answer scores across multiple related instances.Specifically, we normalize these scores across various neighborhoods of closely contrasting questions and/or answers, adding a cross entropy loss term in addition to traditional maximum likelihood estimation.Our techniques require bundles of related question-answer pairs, which we either mine from within existing data or create using automated heuristics.We empirically demonstrate the effectiveness of training with instance bundles on two datasets-HotpotQA and ROPES-showing up to 9% absolute gains in accuracy.
Dheeru Dua, Pradeep Dasigi, Sameer Singh 0001, Matt Gardner 0001
EMNLP (1)4
2021 Generative Context Pair Selection for Multi-hop Question Answering
abstract
Dheeru Dua, Cicero Nogueira dos Santos, Patrick Ng, Ben Athiwaratkun, Bing Xiang, Matt Gardner, Sameer Singh. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Dheeru Dua, Cícero Nogueira dos Santos, Patrick Ng, Ben Athiwaratkun, Bing Xiang, Matt Gardner 0001, Sameer Singh 0001
EMNLP (1)6
2021 Paired Examples as Indirect Supervision in Latent Decision Models
abstract
Compositional, structured models are appealing because they explicitly decompose problems and provide interpretable intermediate outputs that give confidence that the model is not simply latching onto data artifacts.Learning these models is challenging, however, because end-task supervision only provides a weak indirect signal on what values the latent decisions should take.This often results in the model failing to learn to perform the intermediate tasks correctly.In this work, we introduce a way to leverage paired examples that provide stronger cues for learning latent decisions.When two related training examples share internal substructure, we add an additional training objective to encourage consistency between their latent decisions.Such an objective does not require external supervision for the values of the latent output, or even the end task, yet provides an additional training signal to that provided by individual training examples themselves.We apply our method to improve compositional question answering using neural module networks on the DROP dataset.We explore three ways to acquire paired questions in DROP: (a) discovering naturally occurring paired examples within the dataset, (b) constructing paired examples using templates, and (c) generating paired examples using a question generation model.We empirically demonstrate that our proposed approach improves both in-and outof-distribution generalization and leads to correct latent decision predictions.
Nitish Gupta, Sameer Singh 0001, Matt Gardner 0001, Dan Roth 0001
EMNLP (1)3
2021 Mitigating False-Negative Contexts in Multi-document Question Answering with Retrieval Marginalization
abstract
Question Answering (QA) tasks requiring information from multiple documents often rely on a retrieval model to identify relevant information for reasoning.The retrieval model is typically trained to maximize the likelihood of the labeled supporting evidence.However, when retrieving from large text corpora such as Wikipedia, the correct answer can often be obtained from multiple evidence candidates.Moreover, not all such candidates are labeled as positive during annotation, rendering the training signal weak and noisy.This problem is exacerbated when the questions are unanswerable or when the answers are Boolean, since the model cannot rely on lexical overlap to make a connection between the answer and supporting evidence.We develop a new parameterization of set-valued retrieval that handles unanswerable queries, and we show that marginalizing over this set during training allows a model to mitigate false negatives in supporting evidence annotations.We test our method on two multi-document QA datasets, IIRC and HotpotQA.On IIRC, we show that joint modeling with marginalization improves model performance by 5.5 F1 points and achieves a new state-of-the-art performance of 50.5 F1.We also show that retrieval marginalization results in 4.1 QA F1 improvement over a non-marginalized baseline on HotpotQA in the fullwiki setting. 1 * Majority of the work done as an intern at AI2. 1 Code available at https://github.com/ niansong1996/retrieval_marginalization.An Example in IIRC: Q: How many other Cardinals participated in the 2005 papal conclave with Policarpo?
Ansong Ni, Matt Gardner 0001, Pradeep Dasigi
EMNLP (1)2
2021 A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers
abstract
Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, Matt Gardner. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, Matt Gardner 0001
NAACL-HLT6
2021 Latent Compositional Representations Improve Systematic Generalization in Grounded Question Answering
abstract
Abstract Answering questions that involve multi-step reasoning requires decomposing them and using the answers of intermediate steps to reach the final answer. However, state-of-the-art models in grounded question answering often do not explicitly perform decomposition, leading to difficulties in generalization to out-of-distribution examples. In this work, we propose a model that computes a representation and denotation for all question spans in a bottom-up, compositional manner using a CKY-style parser. Our model induces latent trees, driven by end-to-end (the answer) supervision only. We show that this inductive bias towards tree structures dramatically improves systematic generalization to out-of- distribution examples, compared to strong baselines on an arithmetic expressions benchmark as well as on C losure, a dataset that focuses on systematic generalization for grounded question answering. On this challenging dataset, our model reaches an accuracy of 96.1%, significantly higher than prior models that almost perfectly solve the task on a random, in-distribution split.
Ben Bogin, Sanjay Subramanian, Matt Gardner 0001, Jonathan Berant
Trans. Assoc. Comput. Linguistics3
2020 Benefits of Intermediate Annotations in Reading Comprehension
abstract
Complex, compositional reading comprehension datasets require performing latent sequential decisions that are learned via supervision from the final answer.A large combinatorial space of possible decision paths that result in the same answer, compounded by the lack of intermediate supervision to help choose the right path, makes the learning particularly hard for this task.In this work, we study the benefits of collecting intermediate reasoning supervision along with the answer during data collection.We find that these intermediate annotations can provide two-fold benefits.First, we observe that for any collection budget, spending a fraction of it on intermediate annotations results in improved model performance, for two complex compositional datasets: DROP and Quoref.Second, these annotations encourage the model to learn the correct latent reasoning steps, helping combat some of the biases introduced during the data collection process.
Dheeru Dua, Sameer Singh 0001, Matt Gardner 0001
ACL3
2020 Dynamic Sampling Strategies for Multi-Task Reading Comprehension
abstract
Building general reading comprehension systems, capable of solving multiple datasets at the same time, is a recent aspirational goal in the research community.Prior work has focused on model architectures or generalization to held out datasets, and largely passed over the particulars of the multi-task learning set up.We show that a simple dynamic sampling strategy, selecting instances for training proportional to the multi-task model's current performance on a dataset relative to its singletask performance, gives substantive gains over prior multi-task sampling strategies, mitigating the catastrophic forgetting that is common in multi-task learning.We also demonstrate that allowing instances of different tasks to be interleaved as much as possible between each epoch and batch has a clear benefit in multitask performance over forcing task homogeneity at the epoch or batch level.Our final model shows greatly increased performance over the best model on ORB, a recently-released multitask reading comprehension benchmark.
Ananth Gottumukkala, Dheeru Dua, Sameer Singh 0001, Matt Gardner 0001
ACL4
2020 On Importance Sampling-Based Evaluation of Latent Language Models
abstract
Language models that use additional latent structures (e.g., syntax trees, coreference chains, and knowledge graph links) provide several advantages over traditional language models.However, likelihood-based evaluation of these models is often intractable as it requires marginalizing over the latent space.Existing methods avoid this issue by using importance sampling.Although this approach has asymptotic guarantees, analysis is rarely conducted on the effect of decisions such as sample size, granularity of sample aggregation, and the proposal distribution on the reported estimates.In this paper, we measure the effect these factors have on perplexity estimates for three different latent language models.In addition, we elucidate subtle differences in how importance sampling is applied, which can have substantial effects on the final estimates, as well as provide theoretical results that reinforce the validity of importance sampling for evaluating latent language models.
Robert L. Logan IV, Matt Gardner 0001, Sameer Singh 0001
ACL2
2020 Obtaining Faithful Interpretations from Compositional Neural Networks
abstract
Neural module networks (NMNs) are a popular approach for modeling compositionality: they achieve high accuracy when applied to problems in language and vision, while reflecting the compositional structure of the problem in the network architecture.However, prior work implicitly assumed that the structure of the network modules, describing the abstract reasoning process, provides a faithful explanation of the model's reasoning; that is, that all modules perform their intended behaviour.In this work, we propose and conduct a systematic evaluation of the intermediate outputs of NMNs on NLVR2 and DROP, two datasets which require composing multiple reasoning steps.We find that the intermediate outputs differ from the expected output, illustrating that the network structure does not provide a faithful explanation of model behaviour.To remedy that, we train the model with auxiliary supervision and propose particular choices for module architecture that yield much better faithfulness, at a minimal cost to accuracy.
Sanjay Subramanian, Ben Bogin, Nitish Gupta, Tomer Wolfson, Sameer Singh 0001, Jonathan Berant, Matt Gardner 0001
ACL7
2020 MOCHA: A Dataset for Training and Evaluating Generative Reading Comprehension Metrics
abstract
Posing reading comprehension as a generation problem provides a great deal of flexibility, allowing for open-ended questions with few restrictions on possible answers.However, progress is impeded by existing generation metrics, which rely on token overlap and are agnostic to the nuances of reading comprehension.To address this, we introduce a benchmark for training and evaluating generative reading comprehension metrics: MOdeling Correctness with Human Annotations.MOCHA contains 40K human judgement scores on model outputs from 6 diverse question answering datasets and an additional set of minimal pairs for evaluation.Using MOCHA, we train a Learned Evaluation metric for Reading Comprehension, LERC, to mimic human judgement scores.LERC outperforms baseline metrics by 10 to 36 absolute Pearson points on held-out annotations.When we evaluate robustness on minimal pairs, LERC achieves 80% accuracy, outperforming baselines by 14 to 26 absolute percentage points while leaving significant room for improvement.MOCHA presents a challenging problem for developing accurate and robust generative reading comprehension metrics. 1
Anthony Chen, Gabriel Stanovsky, Sameer Singh 0001, Matt Gardner 0001
EMNLP (1)4
2020 IIRC: A Dataset of Incomplete Information Reading Comprehension Questions
abstract
Humans often have to read multiple documents to address their information needs.However, most existing reading comprehension (RC) tasks only focus on questions for which the contexts provide all the information required to answer them, thus not evaluating a system's performance at identifying a potential lack of sufficient information and locating sources for that information.To fill this gap, we present a dataset, IIRC, with more than 13K questions over paragraphs from English Wikipedia that provide only partial information to answer them, with the missing information occurring in one or more linked documents.The questions were written by crowd workers who did not have access to any of the linked documents, leading to questions that have little lexical overlap with the contexts where the answers appear.This process also gave many questions without answers, and those that require discrete reasoning, increasing the difficulty of the task.We follow recent modeling work on various reading comprehension datasets to construct a baseline model for this dataset, finding that it achieves 31.1% F1 on this task, while estimated human performance is 88.4%.The dataset, code for the baseline system, and a leaderboard can be found at https://allennlp.org/iirc.
James Ferguson, Matt Gardner 0001, Hannaneh Hajishirzi, Tushar Khot, Pradeep Dasigi
EMNLP (1)2
2020 Multi-Step Inference for Reasoning Over Paragraphs
abstract
Complex reasoning over text requires understanding and chaining together free-form predicates and logical connectives.Prior work has largely tried to do this either symbolically or with black-box transformers.We present a middle ground between these two extremes: a compositional model reminiscent of neural module networks that can perform chained logical reasoning.This model first finds relevant sentences in the context and then chains them together using neural modules.Our model gives significant performance improvements (up to 29% relative error reduction when combined with a reranker) on ROPES, a recentlyintroduced complex reasoning dataset.
Jiangming Liu, Matt Gardner 0001, Shay B. Cohen, Mirella Lapata
EMNLP (1)2
2020 TORQUE: A Reading Comprehension Dataset of Temporal Ordering Questions
abstract
A critical part of reading is being able to understand the temporal relationships between events described in a passage of text, even when those relationships are not explicitly stated.However, current machine reading comprehension benchmarks have practically no questions that test temporal phenomena, so systems trained on these benchmarks have no capacity to answer questions such as "what happened before/after [some event]?"We introduce TORQUE, a new English reading comprehension benchmark built on 3.2k news snippets with 21k human-generated questions querying temporal relationships.Results show that RoBERTa-large achieves an exact-match score of 51% on the test set of TORQUE, about 30% behind human performance.1 1 https://allennlp.org/torque.htmlHeavy snow is causing disruption to transport across the UK, with heavy rainfall bringing flooding to the south-west of England.Rescuers searching for a woman trapped in a landslide at her home in Looe, Cornwall, said they had found a body.Q1: What events have already finished?A: searching trapped landslide said found Q2: What events have begun but has not finished?A: snow causing disruption rainfall bringing flooding Q3: What will happen in the future?A: No answers.Q4: What happened before a woman was trapped?A: landslide Q5: What had started before a woman was trapped?A: snow rainfall landslide Q6: What happened while a woman was trapped?A: searching Q7: What happened after a woman was trapped?A: searching said found Q8: What happened at about the same time as the snow?A: rainfall Q9: What happened after the snow started?A: causing disruption bringing flooding searching trapped landslide said found Q10: What happened before the snow started?A: No answers.warm
Qiang Ning, Hao Wu 0034, Rujun Han, Nanyun Peng 0001, Matt Gardner 0001, Dan Roth 0001
EMNLP (1)5
2020 Learning from Task Descriptions
abstract
Typically, machine learning systems solve new tasks by training on thousands of examples.In contrast, humans can solve new tasks by reading some instructions, with perhaps an example or two.To take a step toward closing this gap, we introduce a framework for developing NLP systems that solve new tasks after reading their descriptions, synthesizing prior work in this area.We instantiate this framework with a new English language dataset, ZEST, structured for task-oriented evaluation on unseen tasks.Formulating task descriptions as questions, we ensure each is general enough to apply to many possible inputs, thus comprehensively evaluating a model's ability to solve each task.Moreover, the dataset's structure tests specific types of systematic generalization.We find that the state-of-the-art T5 model achieves a score of 12% on ZEST, leaving a significant challenge for NLP researchers. 1
Orion Weller, Nicholas Lourie, Matt Gardner 0001, Matthew E. Peters
EMNLP (1)3
2020 Neural Module Networks for Reasoning over Text
Nitish Gupta, Dan Roth 0001, Sameer Singh 0001, Matt Gardner 0001
ICLR5
2020 Break It Down: A Question Understanding Benchmark
abstract
Understanding natural language questions entails the ability to break down a question into the requisite steps for computing its answer. In this work, we introduce a Question Decomposition Meaning Representation (QDMR) for questions. QDMR constitutes the ordered list of steps, expressed through natural language, that are necessary for answering a question. We develop a crowdsourcing pipeline, showing that quality QDMRs can be annotated at scale, and release the Break dataset, containing over 83K pairs of questions and their QDMRs. We demonstrate the utility of QDMR by showing that (a) it can be used to improve open-domain question answering on the HotpotQA dataset, (b) it can be deterministically converted to a pseudo-SQL formal language, which can alleviate annotation in semantic parsing applications. Last, we use Break to train a sequence-to-sequence model with copying that parses questions into QDMR structures, and show that it substantially outperforms several natural baselines.
Tomer Wolfson, Mor Geva, Ankit Gupta 0001, Yoav Goldberg, Matt Gardner 0001, Daniel Deutch, Jonathan Berant
Trans. Assoc. Comput. Linguistics5
2019 QUAREL: A Dataset and Models for Answering Questions about Qualitative Relationships
abstract
Many natural la guage questions require recognizing and reasoning with qualitative relationships (e.g., in science, economics, and medicine), but are challenging to answer with corpus-based methods. Qualitative modeling provides tools that support such reasoning, but the semantic parsing task of mapping questions into those models has formidable challenges. We present QUAREL, a dataset of diverse story questions involving qualitative relationships that characterize these challenges, and techniques that begin to address them. The dataset has 2771 questions relating 19 different types of quantities. For example, “Jenny observes that the robot vacuum cleaner moves slower on the living room carpet than on the bedroom carpet. Which carpet has more friction?” We contribute (1) a simple and flexible conceptual framework for representing these kinds of questions; (2) the QUAREL dataset, including logical forms, exemplifying the parsing challenges; and (3) two novel models for this task, built as extensions of type-constrained semantic parsing. The first of these models (called QUASP+) significantly outperforms off-the-shelf tools on QUAREL. The second (QUASP+ZERO) demonstrates zero-shot capability, i.e., the ability to handle new qualitative relationships without requiring additional training data, something not possible with previous models. This work thus makes inroads into answering complex, qualitative questions that require reasoning, and scaling to new relationships at low cost. The dataset and models are available at http://data.allenai.org/quarel.
Oyvind Tafjord, Peter Clark, Matt Gardner 0001, Scott Yih, Ashish Sabharwal
AAAI3
2019 Representing Schema Structure with Graph Neural Networks for Text-to-SQL Parsing
abstract
Research on parsing language to SQL has largely ignored the structure of the database (DB) schema, either because the DB was very simple, or because it was observed at both training and test time.In SPIDER, a recentlyreleased text-to-SQL dataset, new and complex DBs are given at test time, and so the structure of the DB schema can inform the predicted SQL query.In this paper, we present an encoder-decoder semantic parser, where the structure of the DB schema is encoded with a graph neural network, and this representation is later used at both encoding and decoding time.Evaluation shows that encoding the schema structure improves our parser accuracy from 33.8% to 39.4%, dramatically above the current state of the art, which is at 19.7%.
Ben Bogin, Jonathan Berant, Matt Gardner 0001
ACL (1)3
2019 Barack's Wife Hillary: Using Knowledge Graphs for Fact-Aware Language Modeling
abstract
Modeling human language requires the ability to not only generate fluent text but also encode factual knowledge.However, traditional language models are only capable of remembering facts seen at training time, and often have difficulty recalling them.To address this, we introduce the knowledge graph language model (KGLM), a neural language model with mechanisms for selecting and copying facts from a knowledge graph that are relevant to the context.These mechanisms enable the model to render information it has never seen before, as well as generate out-of-vocabulary tokens.We also introduce the Linked WikiText-2 dataset, 1 a corpus of annotated text aligned to the Wikidata knowledge graph whose contents (roughly) match the popular WikiText-2 benchmark (Merity et al., 2017).In experiments, we demonstrate that the KGLM achieves significantly better performance than a strong baseline language model.We additionally compare different language models' ability to complete sentences requiring factual knowledge, and show that the KGLM outperforms even very large language models in generating facts.
Robert L. Logan IV, Nelson F. Liu, Matthew E. Peters, Matt Gardner 0001, Sameer Singh 0001
ACL (1)4
2019 Compositional Questions Do Not Necessitate Multi-hop Reasoning
abstract
Multi-hop reading comprehension (RC) questions are challenging because they require reading and reasoning over multiple paragraphs.We argue that it can be difficult to construct large multi-hop RC datasets.For example, even highly compositional questions can be answered with a single hop if they target specific entity types, or the facts needed to answer them are redundant.Our analysis is centered on HOTPOTQA, where we show that single-hop reasoning can solve much more of the dataset than previously thought.We introduce a single-hop BERT-based RC model that achieves 67 F1-comparable to state-of-theart multi-hop models.We also design an evaluation setting where humans are not shown all of the necessary paragraphs for the intended multi-hop reasoning but can still answer over 80% of questions.Together with detailed error analysis, these results suggest there should be an increasing focus on the role of evidence in multi-hop reasoning and possibly even a shift towards information retrieval style evaluations with large and diverse evidence collections.
Sewon Min, Eric Wallace, Sameer Singh 0001, Matt Gardner 0001, Hannaneh Hajishirzi, Luke Zettlemoyer
ACL (1)4
2019 Global Reasoning over Database Structures for Text-to-SQL Parsing
abstract
Ben Bogin, Matt Gardner, Jonathan Berant. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ben Bogin, Matt Gardner 0001, Jonathan Berant
EMNLP/IJCNLP (1)2
2019 Quoref: A Reading Comprehension Dataset with Questions Requiring Coreferential Reasoning
abstract
Pradeep Dasigi, Nelson F. Liu, Ana Marasović, Noah A. Smith, Matt Gardner. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Pradeep Dasigi, Nelson F. Liu, Ana Marasovic, Noah A. Smith, Matt Gardner 0001
EMNLP/IJCNLP (1)5
2019 QuaRTz: An Open-Domain Dataset of Qualitative Relationship Questions
abstract
Oyvind Tafjord, Matt Gardner, Kevin Lin, Peter Clark. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Oyvind Tafjord, Matt Gardner 0001, Peter Clark
EMNLP/IJCNLP (1)2
2019 Universal Adversarial Triggers for Attacking and Analyzing NLP
abstract
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, Sameer Singh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Eric Wallace, Shi Feng 0005, Nikhil Kandpal, Matt Gardner 0001, Sameer Singh 0001
EMNLP/IJCNLP (1)4
2019 Do NLP Models Know Numbers? Probing Numeracy in Embeddings
abstract
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh, Matt Gardner. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh 0001, Matt Gardner 0001
EMNLP/IJCNLP (1)5
2018 Simple and Effective Multi-Paragraph Reading Comprehension
abstract
We introduce a method of adapting neural paragraph-level question answering models to the case where entire documents are given as input.Most current question answering models cannot scale to document or multi-document input, and naively applying these models to each paragraph independently often results in them being distracted by irrelevant text.We show that it is possible to significantly improve performance by using a modified training scheme that teaches the model to ignore non-answer containing paragraphs.Our method involves sampling multiple paragraphs from each document, and using an objective function that requires the model to produce globally correct output.We additionally identify and improve upon a number of other design decisions that arise when working with document-level data.Experiments on TriviaQA and SQuAD shows our method advances the state of the art, including a 10 point gain on TriviaQA.
Matt Gardner 0001
ACL (1)2
2018 Structured Alignment Networks for Matching Sentences
abstract
Many tasks in natural language processing involve comparing two sentences to compute some notion of relevance, entailment, or similarity.Typically, this comparison is done either at the word level or at the sentence level, with no attempt to leverage the inherent structure of the sentence.When sentence structure is used for comparison, it is obtained during a non-differentiable pre-processing step, leading to propagation of errors.We introduce a model of structured alignments between sentences, showing how to compare two sentences by matching their latent structures.Using a structured attention mechanism, our model matches candidate spans in the first sentence to candidate spans in the second sentence, simultaneously discovering the tree structure of each sentence.Our model is fully differentiable and trained only on the matching objective.We evaluate this model on two tasks, entailment detection and answer sentence selection, and find that modeling latent tree structures results in superior performance.Analysis of the learned sentence structures shows they can reflect some syntactic phenomena.
Yang Liu 0124, Matt Gardner 0001, Mirella Lapata
EMNLP2
2018 Deep Contextualized Word Representations
abstract
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, Luke Zettlemoyer. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner 0001, Kenton Lee, Luke Zettlemoyer
NAACL-HLT4
2017 Open-Vocabulary Semantic Parsing with both Distributional Statistics and Formal Knowledge
abstract
Traditional semantic parsers map language onto compositional, executable queries in a fixed schema. This mapping allows them to effectively leverage the information contained in large, formal knowledge bases (KBs, e.g., Freebase) to answer questions, but it is also fundamentally limiting---these semantic parsers can only assign meaning to language that falls within the KB's manually-produced schema. Recently proposed methods for open vocabulary semantic parsing overcome this limitation by learning execution models for arbitrary language, essentially using a text corpus as a kind of knowledge base. However, all prior approaches to open vocabulary semantic parsing replace a formal KB with textual information, making no use of the KB in their models. We show how to combine the disparate representations used by these two approaches, presenting for the first time a semantic parser that (1) produces compositional, executable representations of language, (2) can successfully leverage the information contained in both a formal KB and a large corpus, and (3) is not limited to the schema of the underlying KB. We demonstrate significantly improved performance over state-of-the-art baselines on an open-domain natural language question answering task.
Matt Gardner 0001, Jayant Krishnamurthy
AAAI1
2017 Neural Semantic Parsing with Type Constraints for Semi-Structured Tables
abstract
We present a new semantic parsing model for answering compositional questions on semi-structured Wikipedia tables.Our parser is an encoder-decoder neural network with two key technical innovations:(1) a grammar for the decoder that only generates well-typed logical forms; and(2) an entity embedding and linking module that identifies entity mentions while generalizing across tables.We also introduce a novel method for training our neural model with question-answer supervision.On the WIKITABLEQUESTIONS data set, our parser achieves a state-of-theart accuracy of 43.3% for a single model and 45.9% for a 5-model ensemble, improving on the best prior score of 38.7% set by a 15-model ensemble.These results suggest that type constraints and entity linking are valuable components to incorporate in neural semantic parsers.
Jayant Krishnamurthy, Pradeep Dasigi, Matt Gardner 0001
EMNLP3
2015 Never-Ending Learning
abstract
Whereas people learn many different types of knowledge from diverse experiences over many years, most current machine learning systems acquire just a single function or data model from just a single data set. We propose a never-ending learning paradigm for machine learning, to better reflect the more ambitious and encompassing type of learning performed by humans. As a case study, we describe the Never-Ending Language Learner (NELL), which achieves some of the desired properties of a never-ending learner, and we discuss lessons learned. NELL has been learning to read the web 24 hours/day since January 2010, and so far has acquired a knowledge base with over 80 million confidence-weighted beliefs (e.g., servedWith(tea, biscuits)). NELL has also learned millions of features and parameters that enable it to read these beliefs from the web. Additionally, it has learned to reason over these beliefs to infer new beliefs, and is able to extend its ontology by synthesizing new relational predicates. NELL can be tracked online at http://rtw.ml.cmu.edu, and followed on Twitter at @CMUNELL.
Tom M. Mitchell, William W. Cohen, Estevam Hruschka, Partha P. Talukdar, Justin Betteridge, Andrew Carlson, Bhavana Dalvi, Matt Gardner 0001, Bryan Kisiel, Jayant Krishnamurthy, Ni Lao, Kathryn Mazaitis, Thahir Mohamed, Ndapandula Nakashole, Emmanouil A. Platanios, Alan Ritter, Mehdi Samadi, Burr Settles, Richard C. Wang, Derry Wijaya, Abhinav Gupta 0001, Xinlei Chen, Abulhair Saparov, Malcolm Greaves, Joel Welling
AAAI8
2015 Efficient and Expressive Knowledge Base Completion Using Subgraph Feature Extraction
abstract
We explore some of the practicalities of using random walk inference methods, such as the Path Ranking Algorithm (PRA), for the task of knowledge base completion.We show that the random walk probabilities computed (at great expense) by PRA provide no discernible benefit to performance on this task, so they can safely be dropped.This allows us to define a simpler algorithm for generating feature matrices from graphs, which we call subgraph feature extraction (SFE).In addition to being conceptually simpler than PRA, SFE is much more efficient, reducing computation by an order of magnitude, and more expressive, allowing for much richer features than paths between two nodes in a graph.We show experimentally that this technique gives substantially better performance than PRA and its variants, improving mean average precision from .432 to .528 on a knowledge base completion task using the NELL KB.
Matt Gardner 0001, Tom M. Mitchell
EMNLP1
2015 Translation Invariant Word Embeddings
abstract
Kejun Huang, Matt Gardner, Evangelos Papalexakis, Christos Faloutsos, Nikos Sidiropoulos, Tom Mitchell, Partha P. Talukdar, Xiao Fu. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015.
Kejun Huang, Matt Gardner 0001, Evangelos E. Papalexakis, Christos Faloutsos, Nicholas D. Sidiropoulos, Tom M. Mitchell, Partha P. Talukdar, Xiao Fu 0001
EMNLP2
2014 Incorporating Vector Space Similarity in Random Walk Inference over Knowledge Bases
abstract
Much work in recent years has gone into the construction of large knowledge bases (KBs), such as Freebase, DBPedia, NELL, and YAGO.While these KBs are very large, they are still very incomplete, necessitating the use of inference to fill in gaps.Prior work has shown how to make use of a large text corpus to augment random walk inference over KBs.We present two improvements to the use of such large corpora to augment KB inference.First, we present a new technique for combining KB relations and surface text into a single graph representation that is much more compact than graphs used in prior work.Second, we describe how to incorporate vector space similarity into random walk inference over KBs, reducing the feature sparsity inherent in using surface text.This allows us to combine distributional similarity with symbolic logical inference in novel and effective ways.With experiments on many relations from two separate KBs, we show that our methods significantly outperform prior work on KB inference, both in the size of problem our methods can handle and in the quality of predictions made.
Matt Gardner 0001, Partha P. Talukdar, Jayant Krishnamurthy, Tom M. Mitchell
EMNLP1
2013 Improving Learning and Inference in a Large Knowledge-Base using Latent Syntactic Cues
abstract
Automatically constructed Knowledge Bases (KBs) are often incomplete and there is a genuine need to improve their coverage.Path Ranking Algorithm (PRA) is a recently proposed method which aims to improve KB coverage by performing inference directly over the KB graph.For the first time, we demonstrate that addition of edges labeled with latent features mined from a large dependency parsed corpus of 500 million Web documents can significantly outperform previous PRAbased approaches on the KB inference task.We present extensive experimental results validating this finding.The resources presented in this paper are publicly available.
Matt Gardner 0001, Partha P. Talukdar, Bryan Kisiel, Tom M. Mitchell
EMNLP1
2010 Speculative Evaluation in Particle Swarm Optimization
Matt Gardner 0001, Andrew W. McNabb, Kevin D. Seppi
PPSN (2)1
2009 An exploration of topologies and communication in large particle swarms
abstract
Particle Swarm Optimization (PSO) has typically been used with small swarms of about 50 particles. However, PSO is more efficiently parallelized with large swarms. We formally describe existing topologies and identify variations which are better suited to large swarms in both sequential and parallel computing environments. We examine the performance of PSO for benchmark functions with respect to swarm size and topology. We develop and demonstrate a new PSO variant which leverages the unique strengths of large swarms. “Hearsay PSO” allows for information to flow quickly through the swarm, even with very loosely connected topologies. These loosely connected topologies are well suited to large scale parallel computing environments because they require very little communication between particles. We consider the case where function evaluations are expensive with respect to communication as well as the case where function evaluations are relatively inexpensive. We also consider a situation where local communication is inexpensive compared to external communication, such as multicore systems in a cluster.
Andrew W. McNabb, Matt Gardner 0001, Kevin D. Seppi
IEEE Congress on Evolutionary Computation2