Greg Durrett

dblp:69/7968 · DBLP profile ↗
← Back
84ranked-venue papers
9as first author
51since 2021 · last 2026
0000-0002-7061-7298ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 77 · 9 first-author · 46 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 ProofWala: A Framework for Multilingual Proof Data Synthesis and Theorem-Proving
abstract
In proof assistants, the physical proximity between two formal mathematical concepts is a strong predictor of their mutual relevance. Furthermore, lemmas with close proximity regularly exhibit similar proof structures. We show that this locality property can be exploited through online learning techniques to obtain solving agents that far surpass offline learners when asked to prove theorems in an unseen mathematical setting. We extensively benchmark two such online solvers implemented in the Tactician platform for the Coq proof assistant: First, Tactician's online $k$-nearest neighbor solver, which can learn from recent proofs, shows a $1.72\times$ improvement in theorems proved over an offline equivalent. Second, we introduce a graph neural network, Graph2Tac, with a novel approach to build hierarchical representations for new definitions. Graph2Tac's online definition task realizes a $1.5\times$ improvement in theorems solved over an offline baseline. The $k$-NN and Graph2Tac solvers rely on orthogonal online data, making them highly complementary. Their combination improves $1.27\times$ over their individual performances. Both solvers outperform all other general-purpose provers for Coq, including CoqHammer, Proverbot9001, and a transformer baseline by at least $1.48\times$ and are available for practical use by end-users.
Amitayush Thakur, George Tsoukalas, Greg Durrett, Swarat Chaudhuri
ITP3
2025 Causal Graph based Event Reasoning using Semantic Relation Experts
abstract
Mahnaz Koupaee, Xueying Bai, Mudan Chen, Greg Durrett, Nathanael Chambers, Niranjan Balasubramanian. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Mahnaz Koupaee, Xueying Bai, Mudan Chen, Greg Durrett, Nathanael Chambers, Niranjan Balasubramanian
ACL (1)4
2025 Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding
abstract
Determining faithfulness of a claim to a source document is an important problem across many domains.This task is generally treated as a binary judgment of whether the claim is supported or unsupported in relation to the source.In many cases, though, whether a claim is supported can be ambiguous.For instance, it may depend on making inferences from given evidence, and different people can reasonably interpret the claim as either supported or unsupported based on their agreement with those inferences.Forcing binary labels upon such claims lowers the reliability of evaluation.In this work, we reframe the task to manage the subjectivity involved with factuality judgments of ambiguous claims.We introduce LLMgenerated edits of summaries as a method of providing a nuanced evaluation of claims: how much does a summary need to be edited to be unambiguous?Whether a claim gets rewritten and how much it changes can be used as an automatic evaluation metric, the Ambiguity Rewrite Metric (ARM), with a much richer feedback signal than a binary judgment of faithfulness.We focus on the area of narrative summarization as it is particularly rife with ambiguity and subjective interpretation.We show that ARM produces a 21% absolute improvement in annotator agreement on claim faithfulness, indicating that subjectivity is reduced.
Melanie Subbiah, Akankshya Mishra, Grace Kim, Liyan Tang, Greg Durrett, Kathy McKeown
EMNLP5
2025 To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
abstract
Chain-of-thought (CoT) via prompting is the de facto method for eliciting reasoning capabilities from large language models (LLMs). But for what kinds of tasks is this extra "thinking" really helpful? To analyze this, we conducted a quantitative meta-analysis covering over 100 papers using CoT and ran our own evaluations of 20 datasets across 14 models. Our results show that CoT gives strong performance benefits primarily on tasks involving math or logic, with much smaller gains on other types of tasks. On MMLU, directly generating the answer without CoT leads to almost identical accuracy as CoT unless the question or model's response contains an equals sign, indicating symbolic operations and reasoning. Following this finding, we analyze the behavior of CoT on these problems by separating planning and execution and comparing against tool-augmented LLMs. Much of CoT's gain comes from improving symbolic execution, but it underperforms relative to using a symbolic solver. Our results indicate that CoT can be applied selectively, maintaining performance while saving inference costs. Furthermore, they suggest a need to move beyond prompt-based CoT to new paradigms that better leverage intermediate computation across the whole range of LLM applications.
Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xi Ye 0003, Kyle Mahowald, Greg Durrett
ICLR10
2025 Understanding Synthetic Context Extension via Retrieval Heads
abstract
Long-context LLMs are increasingly in demand for applications such as retrieval-augmented generation. To defray the cost of pretraining LLMs over long contexts, recent work takes an approach of synthetic context extension: fine-tuning LLMs with synthetically-generated long-context data. However, it remains unclear how and why this synthetic context extension imparts abilities for downstream long-context tasks. In this paper, we investigate fine-tuning on synthetic data for three long-context tasks that require retrieval and reasoning. We vary the realism of "needle'' concepts to be retrieved and diversity of the surrounding "haystack'' context, from using LLMs to construct synthetic documents to using templated relations and creating symbolic datasets. Although models trained on synthetic data underperform models trained on the real data, the impacts of both training settings can be understood via a shared feature of the attention computation, retrieval heads (Wu et al., 2024). The retrieval heads learned from synthetic data have high overlap with retrieval heads learned on real data. Furthermore, there is a strong correlation between the recall of heads learned and the downstream performance of a model, allowing us to interpret and predict the performance of models trained in different settings. Our results shed light on how to interpret synthetic data fine-tuning performance and how to approach creating better data for learning real-world LLM capabilities over long contexts.
Fangcong Yin, Greg Durrett
ICML3
2025 From Distributional to Overton Pluralism: Investigating Large Language Model Alignment
abstract
Thom Lake, Eunsol Choi, Greg Durrett. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Thom Lake, Eunsol Choi, Greg Durrett
NAACL (Long Papers)3
2025 Sparta Alignment: Collectively Aligning Multiple Language Models through Combat
abstract
We propose Sparta Alignment, an algorithm to collectively align multiple LLMs through competition and combat. To complement a single model's lack of diversity in generation and biases in evaluation, multiple LLMs form a 'sparta tribe' to compete against each other in fulfilling instructions while serving as judges for the competition of others. For each iteration, one instruction and two models are selected for a duel, the other models evaluate the two responses, and their evaluation scores are aggregated through a adapted elo-ranking based reputation system, where winners/losers of combat gain/lose weight in evaluating others. The peer-evaluated combat results then become preference pairs where the winning response is preferred over the losing one, and all models learn from these preferences at the end of each iteration. Sparta Alignment enables the self-evolution of multiple LLMs in an iterative and collective competition process. Extensive experiments demonstrate that Sparta Alignment outperforms initial models and 4 self-alignment baselines across 10 out of 12 tasks and datasets with 7.0\% average improvement. Further analysis reveals that Sparta Alignment generalizes more effectively to unseen tasks and leverages the expertise diversity of participating models to produce more logical, direct and informative outputs.
Yuru Jiang, Wenxuan Ding 0001, Shangbin Feng, Greg Durrett, Yulia Tsvetkov
NeurIPS4
2025 AstroVisBench: A Code Benchmark for Scientific Computing and Visualization in Astronomy
abstract
Large Language Models (LLMs) are being explored for applications in scientific research, including their capabilities to synthesize literature, answer research questions, generate research ideas, and even conduct computational experiments.Ultimately, our goal is for these to help scientists derive novel scientific insights. In many areas of science, such insights often arise from processing and visualizing data to understand its patterns. However, evaluating whether an LLM-mediated scientific workflow produces outputs conveying the correct scientific insights is challenging to evaluate and has not been addressed in past work.We introduce AstroVisBench, the first benchmark for both scientific computing and visualization in the astronomy domain.AstroVisBench judges a language model’s ability to both (1) create astronomy-specific workflows to process and analyze data and (2) visualize the results of these workflows through complex plots.Our evaluation of visualizations uses a novel LLM-as-a-judge workflow, which is validated against annotation by five professional astronomers.Using AstroVisBench we present an evaluation of state-of-the-art language models, showing a significant gap in their ability to engage in astronomy research as useful assistants.This evaluation provides a strong end-to-end evaluation for AI scientists that offers a path forward for the development of visualization-based workflows, which are central to a broad range of domains from physics to biology.
Sebastian Joseph, Syed Murtaza Husain, Stella S. R. Offner, Stéphanie Juneau, Paul Torrey, Adam S. Bolton, Juan P. Farias, Niall Gaffney, Greg Durrett, Junyi Jessy Li
NeurIPS9
2025 ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
abstract
Chart understanding presents a unique challenge for large vision-language models (LVLMs), as it requires the integration of sophisticated textual and visual reasoning capabilities. However, current LVLMs exhibit a notable imbalance between these skills, falling short on visual reasoning that is difficult to perform in text. We conduct a case study using a synthetic dataset solvable only through visual reasoning and show that model performance degrades significantly with increasing visual complexity, while human performance remains robust. We then introduce ChartMuseum, a new Chart Question Answering (QA) benchmark containing 1,162 expert-annotated questions spanning multiple reasoning types, curated from real-world charts across 184 sources, specifically built to evaluate complex visual and textual reasoning. Unlike prior chart understanding benchmarks---where frontier models perform similarly and near saturation---our benchmark exposes a substantial gap between model and human performance, while effectively differentiating model capabilities: although humans achieve 93% accuracy, the best-performing model Gemini-2.5-Pro attains only 63.0%, and the leading open-source LVLM Qwen2.5-VL-72B-Instruct achieves only 38.5%. Moreover, on questions requiring primarily visual reasoning, all models experience a 35%-55% performance drop from text-reasoning-heavy question performance. Lastly, our qualitative error analysis reveals specific categories of visual reasoning that are challenging for current LVLMs. Both ChartMuseum and the evaluation code are available at https://github.com/Liyan06/ChartMuseum.
Liyan Tang, Grace Kim, Thom Lake, Wenxuan Ding 0003, Fangcong Yin, Prasann Singhal, Manya Wadhwa, Zeyu Leo Liu, Zayne Sprague, Ramya Namuduri, Bodun Hu, Juan Diego Rodriguez, Puyuan Peng, Greg Durrett
NeurIPS15
2025 CLEVER: A Curated Benchmark for Formally Verified Code Generation
abstract
We introduce ${\rm C{\small LEVER}}$, a high-quality, manually curated benchmark of 161 problems for end-to-end verified code generation in Lean. Each problem consists of (1) the task of generating a specification that matches a held-out ground-truth specification, and (2) the task of generating a Lean implementation that provably satisfies this specification. Unlike prior benchmarks,${\rm C{\small LEVER}}$ avoids test-case supervision, LLM-generated annotations, and specifications that leak implementation logic or allow vacuous solutions. All outputs are verified post-hoc using Lean's type checker to ensure machine-checkable correctness. We use ${\rm C{\small LEVER}}$ to evaluate several few-shot and agentic approaches based on state-of-the-art language models. These methods all struggle to achieve full verification, establishing it as a challenging frontier benchmark for program synthesis and formal reasoning. Our benchmark can be found on [GitHub](https://github.com/trishullab/clever) as well as [HuggingFace](https://huggingface.co/datasets/amitayusht/clever). All our evaluation code is also available [online](https://github.com/trishullab/clever-prover).
Amitayush Thakur, Jasper Lee, George Tsoukalas, Meghana Aparna Sistla, Matthew Zhao, Stefan Zetzsche, Greg Durrett, Yisong Yue, Swarat Chaudhuri
NeurIPS7
2024 SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation
abstract
It is often desirable to distill the capabilities of large language models (LLMs) into smaller student models due to compute and memory constraints.One way to do this for classification tasks is via dataset synthesis, which can be accomplished by generating examples of each label from the LLM.Prior approaches to synthesis use few-shot prompting, which relies on the LLM's parametric knowledge to generate usable examples.However, this leads to issues of repetition, bias towards popular entities, and stylistic differences from human text.In this work, we propose Synthesize by Retrieval and Refinement (SYNTHESIZRR), which uses retrieval augmentation to introduce variety into the dataset synthesis process: as retrieved passages vary, the LLM is "seeded" with different content to generate its examples.We empirically study the synthesis of six datasets, covering topic classification, sentiment analysis, tone detection, and humor, requiring complex synthesis strategies.We find that SYNTHESIZRR 1 greatly improves lexical and semantic diversity, similarity to human-written text, and distillation performance, when compared to 32-shot prompting and four prior approaches.
Abhishek Divekar, Greg Durrett
EMNLP2
2024 MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
abstract
Recognizing if LLM output can be grounded in evidence is central to many tasks in NLP: retrieval-augmented generation, summarization, document-grounded dialogue, and more.Current approaches to this kind of factchecking are based on verifying each piece of a model generation against potential evidence using an LLM.However, this process can be very computationally expensive, requiring many calls to a model to check a single response.In this work, we show how to build small fact-checking models that have GPT-4level performance but for 400x lower cost.We do this by constructing synthetic training data with GPT-4, which involves creating realistic yet challenging instances of factual errors via a structured generation procedure.Training on this data teaches models to check each fact in the claim and recognize synthesis of information across sentences.For evaluation, we unify datasets from recent work on factchecking and grounding LLM generations into a new benchmark, LLM-AGGREFACT.Our best system MiniCheck-FT5 (770M parameters) outperforms all systems of comparable size and reaches GPT-4 accuracy.We release LLM-AGGREFACT, code for data synthesis, and models. 1
Liyan Tang, Philippe Laban, Greg Durrett
EMNLP3
2024 Which questions should I answer? Salience Prediction of Inquisitive Questions
abstract
Inquisitive questions -open-ended, curiositydriven questions people ask as they read -are an integral part of discourse processing (Van Kuppevelt, 1995;Onea, 2016;Kehler and Rohde, 2017) and comprehension (Prince, 2004).Recent work in NLP has taken advantage of question generation capabilities of LLMs to enhance a wide range of applications.But the space of inquisitive questions is vast: many potential questions can be evoked from a given context.So which of those should be prioritized to find answers?Linguistic theories, unfortunately, have not yet provided an answer.This paper presents QSALIENCE, a salience predictor of inquisitive questions.QSALIENCE is instruction-tuned over our dataset of linguistannotated salience scores of 1,766 (context, question) pairs.A question scores high on salience if answering it would greatly enhance the understanding of the text (Van Rooy, 2003).We show that highly salient questions are empirically more likely to be answered in the same article, bridging potential questions (Onea, 2016) with Questions Under Discussion (Roberts, 2012).We further validate our findings by showing that answering salient questions is an indicator of summarization quality in news.
Yating Wu 0002, Ritika Mangla, Alexandros G. Dimakis, Greg Durrett, Junyi Jessy Li
EMNLP4
2024 MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
abstract
While large language models (LLMs) equipped with techniques like chain-of-thought prompting have demonstrated impressive capabilities, they still fall short in their ability to reason robustly in complex settings. However, evaluating LLM reasoning is challenging because system capabilities continue to grow while benchmark datasets for tasks like logical deduction have remained static. We introduce MuSR, a dataset for evaluating language models on multistep soft reasoning tasks specified in a natural language narrative. This dataset has two crucial features. First, it is created through a novel neurosymbolic synthetic-to-natural generation algorithm, enabling the construction of complex reasoning instances that challenge GPT-4 (e.g., murder mysteries roughly 1000 words in length) and which can be scaled further as more capable LLMs are released. Second, our data instances are free text narratives corresponding to real-world domains of reasoning; this makes it simultaneously much more challenging than other synthetically-crafted benchmarks while remaining realistic and tractable for human annotators to solve with high accuracy. We evaluate a range of LLMs and prompting techniques on this dataset and characterize the gaps that remain for techniques like chain-of-thought to perform robust reasoning.
Zayne Sprague, Xi Ye 0003, Kaj Bostrom, Swarat Chaudhuri, Greg Durrett
ICLR5
2024 Coeditor: Leveraging Repo-level Diffs for Code Auto-editing
abstract
Developers often dedicate significant time to maintaining and refactoring existing code. However, most prior work on generative models for code focuses solely on creating new code, overlooking the distinctive needs of editing existing code. In this work, we explore a multi-round code auto-editing setting, aiming to predict edits to a code region based on recent changes within the same codebase. Our model, Coeditor, is a fine-tuned language model specifically designed for code editing tasks. We represent code changes using a line diff format and employ static analysis to form large customized model contexts, ensuring the availability of appropriate information for prediction. We collect a code editing dataset from the commit histories of 1650 open-source Python projects for training and evaluation. In a simplified single-round, single-edit task, Coeditor significantly outperforms GPT-3.5 and SOTA open-source code completion models (bringing exact-match accuracy from 34.7 up to 60.4), demonstrating the benefits of incorporating editing history for code completion. In a multi-round, multi-edit setting, we observe substantial gains by iteratively conditioning on additional user edits. We have open-sourced our code, data, and model weights to encourage future research and have released a VSCode extension powered by our model for interactive IDE usage.
Jiayi Wei, Greg Durrett, Isil Dillig
ICLR2
2024 Complex Claim Verification with Evidence Retrieved in the Wild
abstract
Jifan Chen, Grace Kim, Aniruddh Sriram, Greg Durrett, Eunsol Choi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jifan Chen, Grace Kim, Aniruddh Sriram, Greg Durrett, Eunsol Choi
NAACL-HLT4
2024 X-PARADE: Cross-Lingual Textual Entailment and Information Divergence across Paragraphs
abstract
Juan Rodriguez, Katrin Erk, Greg Durrett. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Juan Diego Rodriguez, Katrin Erk, Greg Durrett
NAACL-HLT3
2024 LoFiT: Localized Fine-tuning on LLM Representations
abstract
Recent work in interpretability shows that large language models (LLMs) can be adapted for new tasks in a learning-free way: it is possible to intervene on LLM representations to elicit desired behaviors for alignment. For instance, adding certain bias vectors to the outputs of certain attention heads is reported to boost the truthfulness of models. In this work, we show that localized fine-tuning serves as an effective alternative to such representation intervention methods. We introduce a framework called Localized Fine-Tuning on LLM Representations (LoFiT), which identifies a subset of attention heads that are most important for learning a specific task, then trains offset vectors to add to the model's hidden representations at those selected heads. LoFiT localizes to a sparse set of heads (3%-10%) and learns the offset vectors from limited training data, comparable to the settings used for representation intervention. For truthfulness and reasoning tasks, we find that LoFiT's intervention vectors are more effective for LLM adaptation than vectors from representation intervention methods such as Inference-time Intervention. We also find that the localization step is important: selecting a task-specific set of attention heads can lead to higher performance than intervening on heads selected for a different task. Finally, across 7 tasks we study, LoFiT achieves comparable performance to other parameter-efficient fine-tuning methods such as LoRA, despite modifying 20x-200x fewer parameters than these methods.
Fangcong Yin, Xi Ye 0003, Greg Durrett
NeurIPS3
2023 Can LMs Learn New Entities from Descriptions? Challenges in Propagating Injected Knowledge
abstract
Pre-trained language models (LMs) are used for knowledge intensive tasks like question answering, but their knowledge gets continuously outdated as the world changes.Prior work has studied targeted updates to LMs, injecting individual facts and evaluating whether the model learns these facts while not changing predictions on other contexts.We take a step forward and study LMs' abilities to make inferences based on injected facts (or propagate those facts): for example, after learning that something is a TV show, does an LM predict that you can watch it?We study this with two clozestyle tasks: an existing dataset of real-world sentences about novel entities (ECBD) as well as a new controlled benchmark with manually designed templates requiring varying levels of inference about injected knowledge.Surprisingly, we find that existing methods for updating knowledge (gradient-based fine-tuning and modifications of this approach) show little propagation of injected knowledge.These methods improve performance on cloze instances only when there is lexical overlap between injected facts and target inferences.Yet, prepending entity definitions in an LM's context improves performance across all settings, suggesting that there is substantial headroom for parameterupdating approaches for knowledge injection.
Yasumasa Onoe, Michael J. Q. Zhang, Shankar Padmanabhan, Greg Durrett, Eunsol Choi
ACL (1)4
2023 EEL: Efficiently Encoding Lattices for Reranking
abstract
Standard decoding approaches for conditional text generation tasks typically search for an output hypothesis with high model probability, but this may not yield the best hypothesis according to human judgments of quality.Reranking to optimize for downstream metrics can better optimize for quality, but many metrics of interest are computed with pre-trained language models, which are slow to apply to large numbers of hypotheses.We explore an approach for reranking hypotheses by using Transformers to efficiently encode lattices of generated outputs, a method we call EEL.With a single Transformer pass over the entire lattice, we can approximately compute a contextualized representation of each token as if it were only part of a single hypothesis in isolation.We combine this approach with a new class of token-factored rerankers (TFRs) that allow for efficient extraction of high reranker-scoring hypotheses from the lattice.Empirically, our approach incurs minimal degradation error compared to the exponentially slower approach of encoding each hypothesis individually.When applying EEL with TFRs across three text generation tasks, our results show both substantial speedup compared to naive reranking and often better performance on downstream metrics than comparable approaches.1
Prasann Singhal, Jiacheng Xu 0001, Xi Ye 0003, Greg Durrett
ACL (1)4
2023 Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors
abstract
Liyan Tang, Tanya Goyal, Alex Fabbri, Philippe Laban, Jiacheng Xu, Semih Yavuz, Wojciech Kryscinski, Justin Rousseau, Greg Durrett. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Liyan Tang, Tanya Goyal, Alexander R. Fabbri, Philippe Laban, Jiacheng Xu 0001, Semih Yavuz, Wojciech Kryscinski, Justin F. Rousseau, Greg Durrett
ACL (1)9
2023 A Block Metropolis-Hastings Sampler for Controllable Energy-based Text Generation
abstract
Recent work has shown that energy-based language modeling is an effective framework for controllable text generation because it enables flexible integration of arbitrary discriminators.However, because energy-based LMs are globally normalized, approximate techniques like Metropolis-Hastings (MH) are required for inference.Past work has largely explored simple proposal distributions that modify a single token at a time, like in Gibbs sampling.In this paper, we develop a novel MH sampler that, in contrast, proposes re-writes of the entire sequence in each step via iterative prompting of a large language model.Our new sampler (a) allows for more efficient and accurate sampling from a target distribution and (b) allows generation length to be determined through the sampling procedure rather than fixed in advance, as past work has required.We perform experiments on two controlled generation tasks, showing both downstream performance gains and more accurate target distribution sampling in comparison with single-token proposal techniques. Token-level sampling (Prior work)How are you?Utterance-level block sampling (Ours) Iteration i: How are you?Iteration i+1: How are you?Proposal: How art you?BERT single-token proposal accept / reject Iteration i: How are you?Iteration i+1: How art thou?
Jarad Forristal, Niloofar Mireshghallah, Greg Durrett, Taylor Berg-Kirkpatrick
CoNLL3
2023 Shortcomings of Question Answering Based Factuality Frameworks for Error Localization
abstract
Despite recent progress in abstractive summarization, models often generate summaries with factual errors.Numerous approaches to detect these errors have been proposed, the most popular of which are question answering (QA)based factuality metrics.These have been shown to work well at predicting summarylevel factuality and have potential to localize errors within summaries, but this latter capability has not been systematically evaluated in past research.In this paper, we conduct the first such analysis and find that, contrary to our expectations, QA-based frameworks fail to correctly identify error spans in generated summaries and are outperformed by trivial exact match baselines.Our analysis reveals a major reason for such poor localization: questions generated by the QG module often inherit errors from non-factual summaries which are then propagated further into downstream modules.Moreover, even human-in-the-loop question generation cannot easily offset these problems.Our experiments conclusively show that there exist fundamental issues with localization using the QA framework which cannot be fixed solely by stronger QA and QG models.
Ryo Kamoi, Tanya Goyal, Greg Durrett
EACL3
2023 Modeling Complex Event Scenarios via Simple Entity-focused Questions
abstract
Event scenarios are often complex and involve multiple event sequences connected through different entity participants.Exploring such complex scenarios requires an ability to branch through different sequences, something that is difficult to achieve with standard event language modeling.To address this, we propose a question-guided generation framework that models events in complex scenarios as answers to questions about participants.At any step in the generation process, the framework uses the previously generated events as context, but generates the next event as an answer to one of three questions: what else a participant did, what else happened to a participant, or what else happened.The participants and the questions themselves can be sampled or be provided as input from a user, allowing for controllable exploration.Our empirical evaluation shows that this question-guided generation provides better coverage of participants, diverse events within a domain, comparable perplexities for modeling event sequences, and more effective control for interactive schema generation 1 .
Mahnaz Koupaee, Greg Durrett, Nathanael Chambers, Niranjan Balasubramanian
EACL2
2023 Assessing Out-of-Domain Language Model Performance from Few Examples
abstract
While pretrained language models have exhibited impressive generalization capabilities, they still behave unpredictably under certain domain shifts.In particular, a model may learn a reasoning process on in-domain training data that does not hold for out-of-domain test data.We address the task of predicting outof-domain (OOD) performance in a few-shot fashion: given a few target-domain examples and a set of models with similar training performance, can we understand how these models will perform on OOD test data?We start from the baseline of looking at model accuracy on the few-shot examples, then investigate how to incorporate analysis of the models' behavior using feature attributions to improve our understanding of generalization.Specifically, we explore a set of "factors" designed to reveal model agreement with certain pathological heuristics that may indicate worse generalization capabilities.On textual entailment, paraphrase recognition, and a synthetic classification task, we show that attribution-based factors can help rank relative model OOD performance.However, accuracy on a few-shot test set is a surprisingly strong baseline, particularly when the system designer does not have in-depth prior knowledge about the domain shift.
Prasann Singhal, Jarad Forristal, Xi Ye 0003, Greg Durrett
EACL4
2023 WiCE: Real-World Entailment for Claims in Wikipedia
abstract
Textual entailment models are increasingly applied in settings like fact-checking, presupposition verification in question answering, or summary evaluation.However, these represent a significant domain shift from existing entailment datasets, and models underperform as a result.We propose WICE, a new fine-grained textual entailment dataset built on natural claim and evidence pairs extracted from Wikipedia.In addition to standard claim-level entailment, WICE provides entailment judgments over subsentence units of the claim, and a minimal subset of evidence sentences that support each subclaim.To support this, we propose an automatic claim decomposition strategy using GPT-3.5 which we show is also effective at improving entailment models' performance on multiple datasets at test time.Finally, we show that real claims in our dataset involve challenging verification and retrieval problems that existing models fail to address. 1
Ryo Kamoi, Tanya Goyal, Juan Diego Rodriguez, Greg Durrett
EMNLP4
2023 QUDeval: The Evaluation of Questions Under Discussion Discourse Parsing
abstract
Questions Under Discussion (QUD) is a versatile linguistic framework in which discourse progresses as continuously asking questions and answering them.Automatic parsing of a discourse to produce a QUD structure thus entails a complex question generation task: given a document and an answer sentence, generate a question that satisfies linguistic constraints of QUD and can be grounded in an anchor sentence in prior context.These questions are known to be curiosity-driven and open-ended.This work introduces the first framework for the automatic evaluation of QUD parsing, instantiating the theoretical constraints of QUD in a concrete protocol.We present QUDE-VAL, a dataset of fine-grained evaluation of 2,190 QUD questions generated from both finetuned systems and LLMs.Using QUDEVAL, we show that satisfying all constraints of QUD is still challenging for modern LLMs, and that existing evaluation metrics poorly approximate parser quality.Encouragingly, human-authored QUDs are scored highly by our human evaluators, suggesting that there is headroom for further progress on language modeling to improve both QUD parsing and QUD evaluation.
Yating Wu 0002, Ritika Mangla, Greg Durrett, Junyi Jessy Li
EMNLP3
2023 Explanation Selection Using Unlabeled Data for Chain-of-Thought Prompting
abstract
Recent work has shown how to prompt large language models with explanations to obtain strong performance on textual reasoning tasks, i.e., the chain-of-thought paradigm.However, subtly different explanations can yield widely varying downstream task accuracy.Explanations that have not been "tuned" for a task, such as off-the-shelf explanations written by nonexperts, may lead to mediocre performance.This paper tackles the problem of how to optimize explanation-infused prompts in a blackbox fashion.We first generate sets of candidate explanations for each example in the prompt using a leave-one-out scheme, then find an effective combination of these explanations with a two-stage framework.We first evaluate explanations for each in-context example in isolation according to two proxy metrics, log likelihood and accuracy on new examples.Then, we search over combinations of explanations to find one that yields high performance against a silver-labeled development set.Across four textual reasoning tasks spanning question answering, mathematical reasoning, and natural language inference, results show that our proxy metrics correlate with ground truth accuracy and our overall method can effectively improve prompts over crowdworker annotations and naive search strategies. 1
Xi Ye 0003, Greg Durrett
EMNLP2
2023 TypeT5: Seq2seq Type Inference using Static Analysis
Jiayi Wei, Greg Durrett, Isil Dillig
ICLR2
2023 Propagating Knowledge Updates to LMs Through Distillation
abstract
Modern language models have the capacity to store and use immense amounts of knowledge about real-world entities, but it remains unclear how to update such knowledge stored in model parameters. While prior methods for updating knowledge in LMs successfully inject atomic facts, updated LMs fail to make inferences based on injected facts. In this work, we demonstrate that a context distillation-based approach can both impart knowledge about entities \emph{and} propagate that knowledge to enable broader inferences. Our approach consists of two stages: transfer set generation and distillation on the transfer set. We first generate a transfer set by prompting a language model to generate continuations from the entity definition. Then, we update the model parameters so that the distribution of the LM (the 'student') matches the distribution of the LM conditioned on the definition (the 'teacher') on the transfer set. Our experiments demonstrate that this approach is more effective at propagating knowledge updates than fine-tuning and other gradient-based knowledge-editing methods. Moreover, it does not compromise performance in other contexts, even when injecting the definitions of up to 150 entities at once.
Shankar Padmanabhan, Yasumasa Onoe, Michael J. Q. Zhang, Greg Durrett, Eunsol Choi
NeurIPS4
2023 SatLM: Satisfiability-Aided Language Models Using Declarative Prompting
abstract
Prior work has combined chain-of-thought prompting in large language models (LLMs) with programmatic representations to perform effective and transparent reasoning. While such an approach works well for tasks that only require forward reasoning (e.g., straightforward arithmetic), it is less effective for constraint solving problems that require more sophisticated planning and search. In this paper, we propose a new satisfiability-aided language modeling (SatLM) approach for improving the reasoning capabilities of LLMs. We use an LLM to generate a declarative task specification rather than an imperative program and leverage an off-the-shelf automated theorem prover to derive the final answer. This approach has two key advantages. The declarative specification is closer to the problem description than the reasoning steps are, so the LLM can parse it out of the description more accurately. Furthermore, by offloading the actual reasoning task to an automated theorem prover, our approach can guarantee the correctness of the answer with respect to the parsed specification and avoid planning errors in the solving process. We evaluate SATLM on 8 different datasets and show that it consistently outperforms program-aided LMs in the imperative paradigm. In particular, SATLM outperforms program-aided LMs by 23% on a challenging subset of the GSM arithmetic reasoning dataset; SATLM also achieves a new SoTA on LSAT and BoardgameQA, surpassing previous models that are trained on the respective training sets.
Xi Ye 0003, Qiaochu Chen, Isil Dillig, Greg Durrett
NeurIPS4
2023 Data Extraction via Semantic Regular Expression Synthesis
abstract
Many data extraction tasks of practical relevance require not only syntactic pattern matching but also semantic reasoning about the content of the underlying text. While regular expressions are very well suited for tasks that require only syntactic pattern matching, they fall short for data extraction tasks that involve both a syntactic and semantic component. To address this issue, we introduce semantic regexes, a generalization of regular expressions that facilitates combined syntactic and semantic reasoning about textual data. We also propose a novel learning algorithm that can synthesize semantic regexes from a small number of positive and negative examples. Our proposed learning algorithm uses a combination of neural sketch generation and compositional type-directed synthesis for fast and effective generalization from a small number of examples. We have implemented these ideas in a new tool called Smore and evaluated it on representative data extraction tasks involving several textual datasets. Our evaluation shows that semantic regexes can better support complex data extraction tasks than standard regular expressions and that our learning algorithm significantly outperforms existing tools, including state-of-the-art neural networks and program synthesis tools.
Qiaochu Chen, Arko Banerjee, Çagatay Demiralp, Greg Durrett, Isil Dillig
Proc. ACM Program. Lang.4
2022 ASPECTNEWS: Aspect-Oriented Summarization of News Documents
abstract
Generic summaries try to cover an entire document and query-based summaries try to answer document-specific questions.But real users' needs often fall in between these extremes and correspond to aspects, high-level topics discussed among similar types of documents.In this paper, we collect a dataset of realistic aspect-oriented summaries, ASPECT-NEWS, which covers different subtopics about articles in news sub-domains.We annotate data across two domains of articles, earthquakes and fraud investigations, where each article is annotated with two distinct summaries focusing on different aspects for each domain.A system producing a single generic summary cannot concisely satisfy both aspects.Our focus in evaluation is how well existing techniques can generalize to these domains without seeing in-domain training data, so we turn to techniques to construct synthetic training data that have been used in query-focused summarization work.We compare several training schemes that differ in how strongly keywords are used and how oracle summaries are extracted.Our evaluation shows that our final approach yields (a) focused summaries, better than those from a generic summarization system or from keyword matching; (b) a system sensitive to the choice of keywords.1
Ojas Ahuja, Jiacheng Xu 0001, Kevin Horecka, Greg Durrett
ACL (1)5
2022 Can Explanations Be Useful for Calibrating Black Box Models?
abstract
NLP practitioners often want to take existing trained models and apply them to data from new domains.While fine-tuning or few-shot learning can be used to adapt a base model, there is no single recipe for making these techniques work; moreover, one may not have access to the original model weights if it is deployed as a black box.We study how to improve a black box model's performance on a new domain by leveraging explanations of the model's behavior.Our approach first extracts a set of features combining human intuition about the task with model attributions generated by black box interpretation techniques, then uses a simple calibrator, in the form of a classifier, to predict whether the base model was correct or not.We experiment with our method on two tasks, extractive question answering and natural language inference, covering adaptation from several pairs of domains with limited target-domain data.The experimental results across all the domain pairs show that explanations are useful for calibrating these models, boosting accuracy when predictions do not have to be returned on every example.We further show that the calibration model transfers to some extent between tasks. 1
Xi Ye 0003, Greg Durrett
ACL (1)2
2022 Generating Literal and Implied Subquestions to Fact-check Complex Claims
abstract
Verifying political claims is a challenging task, as politicians can use various tactics to subtly misrepresent the facts for their agenda.Existing automatic fact-checking systems fall short here, and their predictions like "half-true" are not very useful in isolation, since it is unclear which parts of a claim are true or false.In this work, we focus on decomposing a complex claim into a comprehensive set of yes-no subquestions whose answers influence the veracity of the claim.We present CLAIMDECOMP, a dataset of decompositions for over 1000 claims.Given a claim and its verification paragraph written by fact-checkers, our trained annotators write subquestions covering both explicit propositions of the original claim and its implicit facets, such as additional political context that changes our view of the claim's veracity.We study whether state-of-the-art pre-trained models can learn to generate such subquestions.Our experiments show that these models generate reasonable questions, but predicting implied subquestions based only on the claim (without consulting other evidence) remains challenging.Nevertheless, we show that predicted subquestions can help identify relevant evidence to fact-check the full claim and derive the veracity through their answers, suggesting that claim decomposition can be a useful piece of a fact-checking pipeline.1 1 We release our code and dataset: https://jifan-chen. github.io/ClaimDecompJoe Biden stated on August 31, 2020 in a speech: "When I was vice president, violent crime fell 15% in this country.... The murder rate now is up 26% across the nation this year under Donald Trump."Claim Decomposi-on: focus of this work Claim Q1: Did the crime rate fall by 15% during Joe Biden's presidency?Q2: Did the murder rate in 2020 increase by 26% from 2019?Q3: Is Biden comparing crime rates from the same time interval in his statement?Q4: Is violent crime rate and murder rate directly comparable?Literal Implied
Jifan Chen, Aniruddh Sriram, Eunsol Choi, Greg Durrett
EMNLP4
2022 SNaC: Coherence Error Detection for Narrative Summarization
abstract
Progress in summarizing long texts is inhibited by the lack of appropriate evaluation frameworks.A long summary that appropriately covers the facets of that text must also present a coherent narrative, but current automatic and human evaluation methods fail to identify gaps in coherence.In this work, we introduce SNAC, a narrative coherence evaluation framework for fine-grained annotations of long summaries.We develop a taxonomy of coherence errors in generated narrative summaries and collect spanlevel annotations for 6.6k sentences across 150 book and movie summaries.Our work provides the first characterization of coherence errors generated by state-of-the-art summarization models and a protocol for eliciting coherence judgments from crowdworkers.Furthermore, we show that the collected annotations allow us to benchmark past work in coherence modeling and train a strong classifier for automatically localizing coherence errors in generated summaries.Finally, our SNAC framework can support future work in long document summarization and coherence evaluation, including improved summarization modeling and posthoc summary correction.
Tanya Goyal, Junyi Jessy Li, Greg Durrett
EMNLP3
2022 Discourse Comprehension: A Question Answering Framework to Represent Sentence Connections
abstract
While there has been substantial progress in text comprehension through simple factoid question answering, more holistic comprehension of a discourse still presents a major challenge (Dunietz et al., 2020).Someone critically reflecting on a text as they read it will pose curiosity-driven, often open-ended questions, which reflect deep understanding of the content and require complex reasoning to answer (Ko et al., 2020;Westera et al., 2020).A key challenge in building and evaluating models for this type of discourse comprehension is the lack of annotated data, especially since collecting answers to such questions requires high cognitive load for annotators.This paper presents a novel paradigm that enables scalable data collection targeting the comprehension of news documents, viewing these questions through the lens of discourse.The resulting corpus, DCQA (Discourse Comprehension by Question Answering), captures both discourse and semantic links between sentences in the form of free-form, open-ended questions.On an evaluation set that we annotated on questions from Ko et al. (2020), we show that DCQA provides valuable supervision for answering openended questions.We additionally design pretraining methods utilizing existing questionanswering resources, and use synthetic data to accommodate unanswerable questions.
Wei-Jen Ko, Cutter Dalton, Mark Simmons, Eliza Fisher, Greg Durrett, Junyi Jessy Li
EMNLP5
2022 Natural Language Deduction with Incomplete Information
abstract
A growing body of work studies how to answer a question or verify a claim by generating a natural language "proof": a chain of deductive inferences yielding the answer based on a set of premises.However, these methods can only make sound deductions when they follow from evidence that is given.We propose a new system that can handle the underspecified setting where not all premises are stated at the outset; that is, additional assumptions need to be materialized to prove a claim.By using a natural language generation model to abductively infer a premise given another premise and a conclusion, we can impute missing pieces of evidence needed for the conclusion to be true.Our system searches over two fringes in a bidirectional fashion, interleaving deductive (forward-chaining) and abductive (backwardchaining) generation steps.We sample multiple possible outputs for each step to achieve coverage of the search space, at the same time ensuring correctness by filtering lowquality generations with a round-trip validation procedure.Results on a modified version of the EntailmentBank dataset and a new dataset called Everyday Norms: Why Not? show that abductive generation with validation can recover premises across in-and out-of-domain settings.1
Zayne Sprague, Kaj Bostrom, Swarat Chaudhuri, Greg Durrett
EMNLP4
2022 Massive-scale Decoding for Text Generation using Lattices
abstract
Jiacheng Xu, Siddhartha Jonnalagadda, Greg Durrett. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Jiacheng Xu 0001, Siddhartha Jonnalagadda, Greg Durrett
NAACL-HLT3
2022 The Unreliability of Explanations in Few-shot Prompting for Textual Reasoning
abstract
Does prompting a large language model (LLM) like GPT-3 with explanations improve in-context learning? We study this question on two NLP tasks that involve reasoning over text, namely question answering and natural language inference. We test the performance of four LLMs on three textual reasoning datasets using prompts that include explanations in multiple different styles. For these tasks, we find that including explanations in the prompts for OPT, GPT-3 (davinci), and InstructGPT (text-davinci-001) only yields small to moderate accuracy improvements over standard few-show learning. However, text-davinci-002 is able to benefit more substantially.We further show that explanations generated by the LLMs may not entail the models’ predictions nor be factually grounded in the input, even on simple tasks with extractive explanations. However, these flawed explanations can still be useful as a way to verify LLMs’ predictions post-hoc. Through analysis in our three settings, we show that explanations judged by humans to be good—logically consistent with the input and the prediction—more likely cooccur with accurate predictions. Following these observations, we train calibrators using automatically extracted scores that assess the reliability of explanations, allowing us to improve performance post-hoc across all of our datasets.
Xi Ye 0003, Greg Durrett
NeurIPS2
2022 Type-directed synthesis of visualizations from natural language queries
abstract
We propose a new technique based on program synthesis for automatically generating visualizations from natural language queries. Our method parses the natural language query into a refinement type specification using the intents-and-slots paradigm and leverages type-directed synthesis to generate a set of visualization programs that are most likely to meet the user's intent. Our refinement type system captures useful hints present in the natural language query and allows the synthesis algorithm to reject visualizations that violate well-established design guidelines for the input data set. We have implemented our ideas in a tool called Graphy and evaluated it on NLVCorpus, which consists of 3 popular datasets and over 700 real-world natural language queries. Our experiments show that Graphy significantly outperforms state-of-the-art natural language based visualization tools, including transformer and rule-based ones.
Qiaochu Chen, Shankara Pailoor, Celeste Barnaby, Abby Criswell, Chenglong Wang 0005, Greg Durrett, Isil Dillig
Proc. ACM Program. Lang.6
2022 Automated transpilation of imperative to functional code using neural-guided program synthesis
abstract
While many mainstream languages such as Java, Python, and C# increasingly incorporate functional APIs to simplify programming and improve parallelization/performance, there are no effective techniques that can be used to automatically translate existing imperative code to functional variants using these APIs. Motivated by this problem, this paper presents a transpilation approach based on inductive program synthesis for modernizing existing code. Our method is based on the observation that the overwhelming majority of source/target programs in this setting satisfy an assumption that we call trace-compatibility: not only do the programs share syntactically identical low-level expressions, but these expressions also take the same values in corresponding execution traces. Our method leverages this observation to design a new neural-guided synthesis algorithm that (1) uses a novel neural architecture called cognate grammar network (CGN) and (2) leverages a form of concolic execution to prune partial programs based on intermediate values that arise during a computation. We have implemented our approach in a tool called NGST2 and use it to translate imperative Java and Python code to functional variants that use the Stream and functools APIs respectively. Our experiments show that NGST2 significantly outperforms several baselines and that our proposed neural architecture and pruning techniques are vital for achieving good results.
Benjamin Mariano, Yanju Chen, Yu Feng 0001, Greg Durrett, Isil Dillig
Proc. ACM Program. Lang.4
2021 Conditional Generation of Temporally-ordered Event Sequences
abstract
Shih-Ting Lin, Nathanael Chambers, Greg Durrett. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Shih-Ting Lin, Nathanael Chambers, Greg Durrett
ACL/IJCNLP (1)3
2021 Modeling Fine-Grained Entity Types with Box Embeddings
abstract
Yasumasa Onoe, Michael Boratko, Andrew McCallum, Greg Durrett. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yasumasa Onoe, Michael Boratko, Andrew McCallum, Greg Durrett
ACL/IJCNLP (1)4
2021 Dissecting Generation Modes for Abstractive Summarization Models via Ablation and Attribution
abstract
Jiacheng Xu, Greg Durrett. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jiacheng Xu 0001, Greg Durrett
ACL/IJCNLP (1)2
2021 Flexible Generation of Natural Language Deductions
abstract
An interpretable system for open-domain reasoning needs to express its reasoning process in a transparent form.Natural language is an attractive representation for this purpose -it is both highly expressive and easy for humans to understand.However, manipulating natural language statements in logically consistent ways is hard: models must cope with variation in how meaning is expressed while remaining precise.In this paper, we describe PARAPATTERN, a method for building models to generate deductive inferences from diverse natural language inputs without direct human supervision.We train BART-based models (Lewis et al., 2020) to generate the result of applying a particular logical operation to one or more premise statements.Crucially, we develop a largely automated pipeline for constructing suitable training examples from Wikipedia.We evaluate our models using out-of-domain sentence compositions from the QASC (Khot et al., 2020) and EntailmentBank (Dalvi et al., 2021) datasets as well as targeted perturbation sets.Our results show that our models are substantially more accurate and flexible than baseline systems.PARAPATTERN achieves 85% validity on examples of the 'substitution' operation from EntailmentBank without the use of any in-domain training data, matching the performance of a model fine-tuned for EntailmentBank.The full source code for our method is publicly available.1 . SubstitutionPremises: Staphylococcus epidermis is a microorganism.Microorganisms colonize the skin surface.Paraphrased: Staphylococcus epidermidis is a microorganism.Microbiological colonization of the skin surface."Staphylococcus Epidermidis is a Microorganism."The skin surface is colonized by micro organisms.Conclusion: Staphylococcus epidermis colonizes the skin surface.Premises: During the undergraduate years, seminarians learn the ancient language courses.
Kaj Bostrom, Swarat Chaudhuri, Greg Durrett
EMNLP (1)4
2021 Connecting Attributions and QA Model Behavior on Realistic Counterfactuals
abstract
When a model attribution technique highlights a particular part of the input, a user might understand this highlight as making a statement about counterfactuals (Miller, 2019): if that part of the input were to change, the model's prediction might change as well.This paper investigates how well different attribution techniques align with this assumption on realistic counterfactuals in the case of reading comprehension (RC).RC is a particularly challenging test case, as token-level attributions that have been extensively studied in other NLP tasks such as sentiment analysis are less suitable to represent the reasoning that RC models perform.We construct counterfactual sets for three different RC settings, and through heuristics that can connect attribution methods' outputs to high-level model behavior, we can evaluate how useful different attribution methods and even different formats are for understanding counterfactuals.We find that pairwise attributions are better suited to RC than tokenlevel attributions across these different RC settings, with our best performance coming from a modification that we propose to an existing pairwise attribution method. 1
Xi Ye 0003, Rohan Nair, Greg Durrett
EMNLP (1)3
2021 Robust Question Answering Through Sub-part Alignment
abstract
Current textual question answering (QA) models achieve strong performance on in-domain test sets, but often do so by fitting surfacelevel patterns, so they fail to generalize to outof-distribution settings.To make a more robust and understandable QA system, we model question answering as an alignment problem.We decompose both the question and context into smaller units based on off-the-shelf semantic representations (here, semantic roles), and align the question to a subgraph of the context in order to find the answer.We formulate our model as a structured SVM, with alignment scores computed via BERT, and we can train end-to-end despite using beam search for approximate inference.Our use of explicit alignments allows us to explore a set of constraints with which we can prohibit certain types of bad model behavior arising in cross-domain settings.Furthermore, by investigating differences in scores across different potential answers, we can seek to understand what particular aspects of the input lead the model to choose the answer without relying on post-hoc explanation techniques.We train our model on SQuAD v1.1 and test it on several adversarial and out-of-domain datasets.The results show that our model is more robust than the standard BERT QA model, and constraints derived from alignment scores allow us to effectively trade off coverage and accuracy.
Jifan Chen, Greg Durrett
NAACL-HLT2
2021 Did they answer? Subjective acts and intents in conversational discourse
abstract
Elisa Ferracane, Greg Durrett, Junyi Jessy Li, Katrin Erk. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Elisa Ferracane, Greg Durrett, Junyi Jessy Li, Katrin Erk
NAACL-HLT2
2021 Annotating and Modeling Fine-grained Factuality in Summarization
abstract
Recent pre-trained abstractive summarization systems have started to achieve credible performance, but a major barrier to their use in practice is their propensity to output summaries that are not faithful to the input and that contain factual errors.While a number of annotated datasets and statistical models for assessing factuality have been explored, there is no clear picture of what errors are most important to target or where current techniques are succeeding and failing.We explore both synthetic and human-labeled data sources for training models to identify factual errors in summarization, and study factuality at the word-, dependency-, and sentence-level.Our observations are threefold.First, exhibited factual errors differ significantly across datasets, and commonly-used training sets of simple synthetic errors do not reflect errors made on abstractive datasets like XSUM.Second, human-labeled data with fine-grained annotations provides a more effective training signal than sentence-level annotations or synthetic data.Finally, we show that our best factuality detection model enables training of more factual XSUM summarization models by allowing us to identify non-factual tokens in the training data. 1Reference Summary: An early-medieval gold pendant created from an imitation of a Byzantine coin that was found in a Norfolk field is a "rare find", a museum expert has said. Source Article Fragment:Discovered on land at North Elmham, near Dereham, the circa 600 AD coin was created by French rulers of the time to increase their available currency.[…] The pendant was declared treasure by the Norfolk coroner on Wednesday.An 18th century coin believed to be worth more than #1m has been discovered.A gold pendant created from a necklace was found in a field Entitycentric (Ent-C)The pendant was declared a treasure by the Norfolk coroner on Wednesday.The pendant was declared a treasure by the Ohio coroner on March.
Tanya Goyal, Greg Durrett
NAACL-HLT2
2021 Web question answering with neurosymbolic program synthesis
abstract
In this paper, we propose a new technique based on program synthesis for extracting information from webpages. Given a natural language query and a few labeled webpages, our method synthesizes a program that can be used to extract similar types of information from other unlabeled webpages. To handle websites with diverse structure, our approach employs a neurosymbolic DSL that incorporates both neural NLP models as well as standard language constructs for tree navigation and string manipulation. We also propose an optimal synthesis algorithm that generates all DSL programs that achieve optimal F1 score on the training examples. Our synthesis technique is compositional, prunes the search space by exploiting a monotonicity property of the DSL, and uses transductive learning to select programs with good generalization power. We have implemented these ideas in a new tool called WebQA and evaluate it on 25 different tasks across multiple domains. Our experiments show that WebQA significantly outperforms existing tools such as state-of-the-art question answering models and wrapper induction systems.
Qiaochu Chen, Aaron Lamoreaux, Xinyu Wang 0006, Greg Durrett, Osbert Bastani, Isil Dillig
PLDI4
2020 Fine-Grained Entity Typing for Domain Independent Entity Linking
abstract
Neural entity linking models are very powerful, but run the risk of overfitting to the domain they are trained in. For this problem, a “domain” is characterized not just by genre of text but even by factors as specific as the particular distribution of entities, as neural models tend to overfit by memorizing properties of frequent entities in a dataset. We tackle the problem of building robust entity linking models that generalize effectively and do not rely on labeled entity linking data with a specific entity distribution. Rather than predicting entities directly, our approach models fine-grained entity properties, which can help disambiguate between even closely related entities. We derive a large inventory of types (tens of thousands) from Wikipedia categories, and use hyperlinked mentions in Wikipedia to distantly label data and train an entity typing model. At test time, we classify a mention with this typing model and use soft type predictions to link the mention to the most similar candidate entity. We evaluate our entity linking system on the CoNLL-YAGO dataset (Hoffart et al. 2011) and show that our approach outperforms prior domain-independent entity linking systems. We also test our approach in a harder setting derived from the WikilinksNED dataset (Eshel et al. 2017) where all the mention-entity pairs are unseen during test time. Results indicate that our approach generalizes better than a state-of-the-art neural model on the dataset.
Yasumasa Onoe, Greg Durrett
AAAI2
2020 Neural Syntactic Preordering for Controlled Paraphrase Generation
abstract
Paraphrasing natural language sentences is a multifaceted process: it might involve replacing individual words or short phrases, local rearrangement of content, or high-level restructuring like topicalization or passivization.Past approaches struggle to cover this space of paraphrase possibilities in an interpretable manner.Our work, inspired by pre-ordering literature in machine translation, uses syntactic transformations to softly "reorder" the source sentence and guide our neural paraphrasing model.First, given an input sentence, we derive a set of feasible syntactic rearrangements using an encoder-decoder model.This model operates over a partially lexical, partially syntactic view of the sentence and can reorder big chunks.Next, we use each proposed rearrangement to produce a sequence of position embeddings, which encourages our final encoder-decoder paraphrase model to attend to the source words in a particular order.Our evaluation, both automatic and human, shows that the proposed system retains the quality of the baseline approaches while giving a substantial increase in the diversity of the generated paraphrases.
Tanya Goyal, Greg Durrett
ACL2
2020 Benchmarking Multimodal Regex Synthesis with Complex Structures
abstract
Existing datasets for regular expression (regex) generation from natural language are limited in complexity; compared to regex tasks that users post on StackOverflow, the regexes in these datasets are simple, and the language used to describe them is not diverse.We introduce STRUCTUREDREGEX, a new regex synthesis dataset differing from prior ones in three aspects.First, to obtain structurally complex and realistic regexes, we generate the regexes using a probabilistic grammar with pre-defined macros observed from real-world StackOverflow posts.Second, to obtain linguistically diverse natural language descriptions, we show crowdworkers abstract depictions of the underlying regex and ask them to describe the pattern they see, rather than having them paraphrase synthetic language.Third, we augment each regex example with a collection of strings that are and are not matched by the ground truth regex, similar to how real users give examples.Our quantitative and qualitative analysis demonstrates the advantages of STRUCTUREDREGEX over prior datasets.Further experimental results using various multimodal synthesis techniques highlight the challenge presented by our dataset, including non-local constraints and multi-modal inputs. 1 CatTemp concat( Comp , Comp , Comp ) reprange( Expr ,1,2) Literal reprange( Expr ,1,2) • • • • • • • • • one or two digits then "." and one or two digits SepTemp concat( Seg , Delimiter , Seg , Delimiter , Seg ) CatTemp IntTemp IntTemp Const Const • • • • • • • • • • • • • • • three delimited values, first will be an integer, with other two being either numeric or a string IntTemp and( Cons , Cons ) startwith( Expr ) endwith( Expr ) • • • • • • starts with "C0" and end with 4 digits Cons start[end]with(Expr) | not(start[end]with(Expr)) # must (not) start/end with contain(Expr) | not(contain(Expr)) # must (not) contain rep( ,k) | repatleast( ,k) |reprange( ,k,k) # length constraints AdvStartwithCons | AdvEndwithCons # adversative macro (e.g., start with capitals except A) CondContainCons # conditional macro.(e.g.letter, if contained, must be after a digit) Comp Literal | or(Literal,Literal,...) # literals like digits, letters, strings, or set of literals.rep(Expr,k) | repatleast(Expr,k) | reprange(Expr,k,k)# e.g, 3 digits, 2 -5 letter, etc. optional(Comp) # components can be optional.
Xi Ye 0003, Qiaochu Chen, Isil Dillig, Greg Durrett
ACL4
2020 Calibration of Pre-trained Transformers
abstract
Pre-trained Transformers are now ubiquitous in natural language processing, but despite their high end-task performance, little is known empirically about whether they are calibrated.Specifically, do these models' posterior probabilities provide an accurate empirical measure of how likely the model is to be correct on a given example?We focus on BERT (Devlin et al., 2019) and RoBERTa (Liu et al., 2019) in this work, and analyze their calibration across three tasks: natural language inference, paraphrase detection, and commonsense reasoning.For each task, we consider in-domain as well as challenging outof-domain settings, where models face more examples they should be uncertain about.We show that: (1) when used out-of-the-box, pretrained models are calibrated in-domain, and compared to baselines, their calibration error out-of-domain can be as much as 3.5× lower;(2) temperature scaling is effective at further reducing calibration error in-domain, and using label smoothing to deliberately increase empirical uncertainty helps calibrate posteriors out-of-domain. 1
Shrey Desai, Greg Durrett
EMNLP (1)2
2020 Compressive Summarization with Plausibility and Salience Modeling
abstract
Compressive summarization systems typically rely on a crafted set of syntactic rules to determine what spans of possible summary sentences can be deleted, then learn a model of what to actually delete by optimizing for content selection (ROUGE).In this work, we propose to relax the rigid syntactic constraints on candidate spans and instead leave compression decisions to two data-driven criteria: plausibility and salience.Deleting a span is plausible if removing it maintains the grammaticality and factuality of a sentence, and spans are salient if they contain important information from the summary.Each of these is judged by a pre-trained Transformer model, and only deletions that are both plausible and not salient can be applied.When integrated into a simple extraction-compression pipeline, our method achieves strong in-domain results on benchmark summarization datasets, and human evaluation shows that the plausibility model generally selects for grammatical and factual deletions.Furthermore, the flexibility of our approach allows it to generalize cross-domain: our system fine-tuned on only 500 samples from a new domain can match or exceed an in-domain extractive model trained on much more data.1
Shrey Desai, Jiacheng Xu 0001, Greg Durrett
EMNLP (1)3
2020 Inquisitive Question Generation for High Level Text Comprehension
abstract
Inquisitive probing questions come naturally to humans in a variety of settings, but is a challenging task for automatic systems.One natural type of question to ask tries to fill a gap in knowledge during text comprehension, like reading a news article: we might ask about background information, deeper reasons behind things occurring, or more.Despite recent progress with data-driven approaches, generating such questions is beyond the range of models trained on existing datasets.We introduce INQUISITIVE, a dataset of ∼19K questions that are elicited while a person is reading through a document.Compared to existing datasets, INQUISITIVE questions target more towards high-level (semantic and discourse) comprehension of text.We show that readers engage in a series of pragmatic strategies to seek information.Finally, we evaluate question generation models based on GPT-2 (Radford et al., 2019) and show that our model is able to generate reasonable questions although the task is challenging, and highlight the importance of context to generate INQUIS-ITIVE questions.
Wei-Jen Ko, Te-Yuan Chen, Yiyan Huang, Greg Durrett, Junyi Jessy Li
EMNLP (1)4
2020 Understanding Neural Abstractive Summarization Models via Uncertainty
abstract
An advantage of seq2seq abstractive summarization models is that they generate text in a free-form manner, but this flexibility makes it difficult to interpret model behavior.In this work, we analyze summarization decoders in both blackbox and whitebox ways by studying on the entropy, or uncertainty, of the model's token-level predictions.For two strong pretrained models, PEGASUS (Zhang et al., 2020) and BART (Lewis et al., 2020) on two summarization datasets, we find a strong correlation between low prediction entropy and where the model copies tokens rather than generating novel text.The decoder's uncertainty also connects to factors like sentence position and syntactic distance between adjacent pairs of tokens, giving a sense of what factors make a context particularly selective for the model's next output token.Finally, we study the relationship of decoder uncertainty and attention behavior to understand how attention gives rise to these observed effects in the model.We show that uncertainty is a useful perspective for analyzing summarization and text generation models more broadly.1
Jiacheng Xu 0001, Shrey Desai, Greg Durrett
EMNLP (1)3
2020 LambdaNet: Probabilistic Type Inference using Graph Neural Networks
Jiayi Wei, Maruth Goyal, Greg Durrett, Isil Dillig
ICLR3
2020 Multi-modal synthesis of regular expressions
abstract
In this paper, we propose a multi-modal synthesis technique for automatically constructing regular expressions (regexes) from a combination of examples and natural language. Using multiple modalities is useful in this context because natural language alone is often highly ambiguous, whereas examples in isolation are often not sufficient for conveying user intent. Our proposed technique first parses the English description into a so-called hierarchical sketch that guides our programming-by-example (PBE) engine. Since the hierarchical sketch captures crucial hints, the PBE engine can leverage this information to both prioritize the search as well as make useful deductions for pruning the search space.
Qiaochu Chen, Xinyu Wang 0006, Xi Ye 0003, Greg Durrett, Isil Dillig
PLDI4
2020 Sketch-Driven Regular Expression Generation from Natural Language and Examples
abstract
Recent systems for converting natural language descriptions into regular expressions (regexes) have achieved some success, but typically deal with short, formulaic text and can only produce simple regexes. Real-world regexes are complex, hard to describe with brief sentences, and sometimes require examples to fully convey the user’s intent. We present a framework for regex synthesis in this setting where both natural language (NL) and examples are available. First, a semantic parser (either grammar-based or neural) maps the natural language description into an intermediate sketch, which is an incomplete regex containing holes to denote missing components. Then a program synthesizer searches over the regex space defined by the sketch and finds a regex that is consistent with the given string examples. Our semantic parser can be trained purely from weak supervision based on correctness of the synthesized regex, or it can leverage heuristically derived sketches. We evaluate on two prior datasets (Kushman and Barzilay 2013 ; Locascio et al. 2016 ) and a real-world dataset from Stack Overflow. Our system achieves state-of-the-art performance on the prior datasets and solves 57% of the real-world dataset, which existing neural systems completely fail on. 1
Xi Ye 0003, Qiaochu Chen, Xinyu Wang 0006, Isil Dillig, Greg Durrett
Trans. Assoc. Comput. Linguistics5
2019 Domain Agnostic Real-Valued Specificity Prediction
abstract
Sentence specificity quantifies the level of detail in a sentence, characterizing the organization of information in discourse. While this information is useful for many downstream applications, specificity prediction systems predict very coarse labels (binary or ternary) and are trained on and tailored toward specific domains (e.g., news). The goal of this work is to generalize specificity prediction to domains where no labeled data is available and output more nuanced realvalued specificity ratings.We present an unsupervised domain adaptation system for sentence specificity prediction, specifically designed to output real-valued estimates from binary training labels. To calibrate the values of these predictions appropriately, we regularize the posterior distribution of the labels towards a reference distribution. We show that our framework generalizes well to three different domains with 50%-68% mean absolute error reduction than the current state-of-the-art system trained for news sentence specificity. We also demonstrate the potential of our work in improving the quality and informativeness of dialogue generation systems.
Wei-Jen Ko, Greg Durrett, Junyi Jessy Li
AAAI2
2019 Evaluating Discourse in Structured Text Representations
abstract
Discourse structure is integral to understanding a text and is helpful in many NLP tasks.Learning latent representations of discourse is an attractive alternative to acquiring expensive labeled discourse data.Liu and Lapata (2018) propose a structured attention mechanism for text classification that derives a tree over a text, akin to an RST discourse tree.We examine this model in detail, and evaluate on additional discourse-relevant tasks and datasets, in order to assess whether the structured attention improves performance on the end task and whether it captures a text's discourse structure.We find the learned latent trees have little to no structure and instead focus on lexical cues; even after obtaining more structured trees with proposed model modifications, the trees are still far from capturing discourse structure when compared to discourse dependency trees from an existing discourse parser.Finally, ablation studies show the structured attention provides little benefit, sometimes even hurting performance.1
Elisa Ferracane, Greg Durrett, Junyi Jessy Li, Katrin Erk
ACL (1)2
2019 Embedding Time Expressions for Deep Temporal Ordering Models
abstract
Data-driven models have demonstrated stateof-the-art performance in inferring the temporal ordering of events in text.However, these models often overlook explicit temporal signals, such as dates and time windows.Rule-based methods can be used to identify the temporal links between these time expressions (timexes), but they fail to capture timexes' interactions with events and are hard to integrate with the distributed representations of neural net models.In this paper, we introduce a framework to infuse temporal awareness into such models by learning a pre-trained model to embed timexes.We generate synthetic data consisting of pairs of timexes, then train a character LSTM to learn embeddings and classify the timexes' temporal relation.We evaluate the utility of these embeddings in the context of a strong neural model for event temporal ordering, and show a small increase in performance on the MATRES dataset and more substantial gains on an automatically collected dataset with more frequent event-timex interactions.1
Tanya Goyal, Greg Durrett
ACL (1)2
2019 Effective Use of Transformer Networks for Entity Tracking
abstract
Aditya Gupta, Greg Durrett. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Aditya Gupta 0001, Greg Durrett
EMNLP/IJCNLP (1)2
2019 Query-focused Scenario Construction
abstract
The news coverage of events often contains not one but multiple incompatible accounts of what happened. We develop a query-based system that extracts compatible sets of events (scenarios) from such data, formulated as one-class clustering. Our system incrementally evaluates each event’s compatibility with already selected events, taking order into account. We use synthetic data consisting of article mixtures for scalable training and evaluate our model on a new human-curated dataset of scenarios about real-world news topics. Stronger neural network models and harder synthetic training settings are both important to achieve high performance, and our final scenario construction system substantially outperforms baselines based on prior work.
Su Wang 0001, Greg Durrett, Katrin Erk
EMNLP/IJCNLP (1)2
2019 Neural Extractive Text Summarization with Syntactic Compression
abstract
Jiacheng Xu, Greg Durrett. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jiacheng Xu 0001, Greg Durrett
EMNLP/IJCNLP (1)2
2018 Effective Use of Context in Noisy Entity Linking
abstract
To disambiguate between closely related concepts, entity linking systems need to effectively distill cues from a mention's textual context.We investigate several techniques for using these cues in the task of noisy entity linking on short texts.Our starting point is a stateof-the-art attention-based model from prior work; while this model's attention typically identifies context that is topically relevant, it fails to identify some of the most indicative context words, especially those exhibiting lexical overlap with the true title.Augmenting the model with convolutional networks over characters still leaves it largely unable to pick up on these cues compared to sparse features that target them directly, indicating that automatically learning how to identify relevant character-level context features is a hard problem.Armed with these sparse features, our final system 1 outperforms past work on the WikilinksNED test set by 2.8% absolute.
David Mueller, Greg Durrett
EMNLP2
2018 Picking Apart Story Salads
abstract
During natural disasters and conflicts, information about what happened is often confusing, messy, and distributed across many sources.We would like to be able to automatically identify relevant information and assemble it into coherent narratives of what happened.To make this task accessible to neural models, we introduce Story Salads, mixtures of multiple documents that can be generated at scale.By exploiting the Wikipedia hierarchy, we can generate salads that exhibit challenging inference problems.Story salads give rise to a novel, challenging clustering task, where the objective is to group sentences from the same narratives.We demonstrate that simple bag-of-words similarity clustering falls short on this task and that it is necessary to take into account global context and coherence.
Su Wang 0001, Eric Holgate, Greg Durrett, Katrin Erk
EMNLP3
2018 Spherical Latent Spaces for Stable Variational Autoencoders
abstract
A hallmark of variational autoencoders (VAEs) for text processing is their combination of powerful encoder-decoder models, such as LSTMs, with simple latent distributions, typically multivariate Gaussians.These models pose a difficult optimization problem: there is an especially bad local optimum where the variational posterior always equals the prior and the model does not use the latent variable at all, a kind of "collapse" which is encouraged by the KL divergence term of the objective.In this work, we experiment with another choice of latent distribution, namely the von Mises-Fisher (vMF) distribution, which places mass on the surface of the unit hypersphere.With this choice of prior and posterior, the KL divergence term now only depends on the variance of the vMF distribution, giving us the ability to treat it as a fixed hyperparameter.We show that doing so not only averts the KL collapse, but consistently gives better likelihoods than Gaussians across a range of modeling conditions, including recurrent language modeling and bag-ofwords document modeling.An analysis of the properties of our vMF representations shows that they learn richer and more nuanced structures in their latent representations than their Gaussian counterparts.1
Jiacheng Xu 0001, Greg Durrett
EMNLP2
2017 Identifying Products in Online Cybercrime Marketplaces: A Dataset for Fine-grained Domain Adaptation
abstract
Greg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Rebecca Portnoff, Sadia Afroz, Damon McCoy, Kirill Levchenko, Vern Paxson. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2017.
Greg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Rebecca S. Portnoff, Sadia Afroz 0001, Damon McCoy, Kirill Levchenko, Vern Paxson
EMNLP1
2017 Tools for Automated Analysis of Cybercriminal Markets
abstract
Underground forums are widely used by criminals to buy and sell a host of stolen items, datasets, resources, and criminal services. These forums contain important resources for understanding cybercrime. However, the number of forums, their size, and the domain expertise required to understand the markets makes manual exploration of these forums unscalable. In this work, we propose an automated, top-down approach for analyzing underground forums. Our approach uses natural language processing and machine learning to automatically generate high-level information about underground forums, first identifying posts related to transactions, and then extracting products and prices. We also demonstrate, via a pair of case studies, how an analyst can use these automated approaches to investigate other categories of products and transactions. We use eight distinct forums to assess our tools: Antichat, Blackhat World, Carders, Darkode, Hack Forums, Hell, L33tCrew and Nulled. Our automated approach is fast and accurate, achieving over 80% accuracy in detecting post category, product, and prices.
Rebecca S. Portnoff, Sadia Afroz 0001, Greg Durrett, Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Damon McCoy, Kirill Levchenko, Vern Paxson
WWW3
2016 Learning-Based Single-Document Summarization with Compression and Anaphoricity Constraints
abstract
We present a discriminative model for single-document summarization that integrally combines compression and anaphoricity constraints.Our model selects textual units to include in the summary based on a rich set of sparse features whose weights are learned on a large corpus.We allow for the deletion of content within a sentence when that deletion is licensed by compression rules; in our framework, these are implemented as dependencies between subsentential units of text.Anaphoricity constraints then improve cross-sentence coherence by guaranteeing that, for each pronoun included in the summary, the pronoun's antecedent is included as well or the pronoun is rewritten as a full mention.When trained end-to-end, our final system 1 outperforms prior work on both ROUGE as well as on human judgments of linguistic quality.
Greg Durrett, Taylor Berg-Kirkpatrick, Daniel Klein 0001
ACL (1)1
2016 Capturing Semantic Similarity for Entity Linking with Convolutional Neural Networks
abstract
A key challenge in entity linking is making effective use of contextual information to disambiguate mentions that might refer to different entities in different contexts.We present a model that uses convolutional neural networks to capture semantic correspondence between a mention's context and a proposed target entity.These convolutional networks operate at multiple granularities to exploit various kinds of topic information, and their rich parameterization gives them the capacity to learn which n-grams characterize different topics.We combine these networks with a sparse linear model to achieve state-of-the-art performance on multiple entity linking datasets, outperforming the prior systems of Durrett and Klein (2014) and Nguyen et al. (2014). 1
Matthew Francis-Landau, Greg Durrett, Daniel Klein 0001
HLT-NAACL2
2015 Neural CRF Parsing
abstract
Greg Durrett, Dan Klein. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Greg Durrett, Daniel Klein 0001
ACL (1)1
2015 Disfluency Detection with a Semi-Markov Model and Prosodic Features
abstract
We present a discriminative model for detecting disfluencies in spoken language transcripts.Structurally, our model is a semi-Markov conditional random field with features targeting characteristics unique to speech repairs.This gives a significant performance improvement over standard chain-structured CRFs that have been employed in past work.We then incorporate prosodic features over silences and relative word duration into our semi-CRF model, resulting in further performance gains; moreover, these features are not easily replaced by discrete prosodic indicators such as ToBI breaks.Our final system, the semi-CRF with prosodic information, achieves an F-score of 85.4, which is 1.3 F 1 better than the best prior reported F-score on this dataset.
James Ferguson, Greg Durrett, Daniel Klein 0001
HLT-NAACL2
2014 Less Grammar, More Features
abstract
We present a parser that relies primar-ily on extracting information directly from surface spans rather than on propagat-ing information through enriched gram-mar structure. For example, instead of cre-ating separate grammar symbols to mark the definiteness of an NP, our parser might instead capture the same information from the first word of the NP. Moving context out of the grammar and onto surface fea-tures can greatly simplify the structural component of the parser: because so many deep syntactic cues have surface reflexes, our system can still parse accurately with context-free backbones as minimal as X-bar grammars. Keeping the structural backbone simple and moving features to the surface also allows easy adaptation to new languages and even to new tasks. On the SPMRL 2013 multilingual con-stituency parsing shared task (Seddah et al., 2013), our system outperforms the top single parser system of Björkelund et al. (2013) on a range of languages. In addi-tion, despite being designed for syntactic analysis, our system also achieves state-of-the-art numbers on the structural senti-ment task of Socher et al. (2013). Finally, we show that, in both syntactic parsing and sentiment analysis, many broad linguistic trends can be captured via surface features. 1
David Hall 0006, Greg Durrett, Daniel Klein 0001
ACL (1)2
2014 A Joint Model for Entity Analysis: Coreference, Typing, and Linking
abstract
We present a joint model of three core tasks in the entity analysis stack: coreference resolution (within-document clustering), named entity recognition (coarse semantic typing), and entity linking (matching to Wikipedia entities). Our model is formally a structured conditional random field. Unary factors encode local features from strong baselines for each task. We then add binary and ternary factors to capture cross-task interactions, such as the constraint that coreferent mentions have the same semantic type. On the ACE 2005 and OntoNotes datasets, we achieve state-of-the-art results for all three tasks. Moreover, joint modeling improves performance on each task over strong independent baselines.
Greg Durrett, Daniel Klein 0001
Trans. Assoc. Comput. Linguistics1
2013 Unsupervised Transcription of Historical Documents
Taylor Berg-Kirkpatrick, Greg Durrett, Daniel Klein 0001
ACL (1)2
2013 Decentralized Entity-Level Modeling for Coreference Resolution
Greg Durrett, David Hall 0006, Daniel Klein 0001
ACL (1)1
2013 Easy Victories and Uphill Battles in Coreference Resolution
abstract
Classical coreference systems encode various syntactic, discourse, and semantic phenomena explicitly, using heterogenous features computed from hand-crafted heuristics.In contrast, we present a state-of-the-art coreference system that captures such phenomena implicitly, with a small number of homogeneous feature templates examining shallow properties of mentions.Surprisingly, our features are actually more effective than the corresponding hand-engineered ones at modeling these key linguistic phenomena, allowing us to win "easy victories" without crafted heuristics.These features are successful on syntax and discourse; however, they do not model semantic compatibility well, nor do we see gains from experiments with shallow semantic features from the literature, suggesting that this approach to semantics is an "uphill battle."Nonetheless, our final system 1 outperforms the Stanford system (Lee et al. (2011), the winner of the CoNLL 2011 shared task) by 3.5% absolute on the CoNLL metric and outperforms the IMS system (Björkelund and Farkas (2012), the best publicly available English coreference system) by 1.9% absolute.
Greg Durrett, Daniel Klein 0001
EMNLP1
2013 Supervised Learning of Complete Morphological Paradigms
Greg Durrett, John DeNero
HLT-NAACL1
2012 Syntactic Transfer Using a Bilingual Lexicon
Greg Durrett, Adam Pauls, Daniel Klein 0001
EMNLP-CoNLL1
2010 A Genetic Algorithm to Minimize Chromatic Entropy
Greg Durrett, Muriel Médard, Una-May O'Reilly
EvoCOP1