Daniel Klein 0001

dblp:93/6085-1 · also Dan Klein 0001 · DBLP profile ↗
← Back
181ranked-venue papers
14as first author
36since 2021 · last 2025
0000-0002-8881-1902ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 179 · 14 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Enough Coin Flips Can Make LLMs Act Bayesian
abstract
Large language models (LLMs) exhibit the ability to generalize given few-shot examples in their input prompt, an emergent capability known as in-context learning (ICL). We investigate whether LLMs use ICL to perform structured reasoning in ways that are consistent with a Bayesian framework or rely on pattern matching. Using a controlled setting of biased coin flips, we find that: (1) LLMs often possess biased priors, causing initial divergence in zero-shot settings, (2) in-context evidence outweighs explicit bias instructions, (3) LLMs broadly follow Bayesian posterior updates, with deviations primarily due to miscalibrated priors rather than flawed updates, and (4) attention magnitude has negligible effect on Bayesian inference. With sufficient demonstrations of biased coin flips via ICL, LLMs update their priors in a Bayesian manner. Code and visualizations are available on the project page.
Ritwik Gupta, Rodolfo Corona, Jiaxin Ge, Daniel Klein 0001, Trevor Darrell, David M. Chan
ACL (1)5
2025 Pose Priors from Language Models
abstract
Language is often used to describe physical interaction, yet most 3D human pose estimation methods overlook this rich source of information. We bridge this gap by leveraging large multimodal models (LMMs) as priors for reconstructing contact poses, offering a scalable alternative to traditional methods that rely on human annotations or motion capture data. Our approach extracts contact-relevant descriptors from an LMM and translates them into tractable losses to constrain 3D human pose optimization. Despite its simplicity, our method produces compelling reconstructions for both two-person interactions and self-contact scenarios, accurately capturing the semantics of physical and social interactions. Our results demonstrate that LMMs can serve as powerful tools for contact prediction and pose estimation, offering an alternative to costly manual human annotations or motion capture data. Our code is publicly available at https://prosepose.github.io.
Sanjay Subramanian, Evonne Ng, Lea Müller, Daniel Klein 0001, Shiry Ginosar, Trevor Darrell
CVPR4
2025 Why Do Multi-Agent LLM Systems Fail?
abstract
Despite enthusiasm for Multi-Agent LLM Systems (MAS), their performance gains on popular benchmarks are often minimal. This gap highlights a critical need for a principled understanding of why MAS fail. Addressing this question requires systematic identification and analysis of failure patterns. We introduce MAST-Data, a comprehensive dataset of 1600+ annotated traces collected across 7 popular MAS frameworks. MAST-Data is the first multi-agent system dataset to outline the failure dynamics in MAS for guiding the development of better future systems. To enable systematic classification of failures for MAST-Data, we build the first Multi-Agent System Failure Taxonomy (MAST). We develop MAST through rigorous analysis of 150 traces, guided closely by expert human annotators andvalidated by high inter-annotator agreement (κ = 0.88). This process identifies 14 unique modes, clustered into 3 categories: (i) system design issues, (ii) inter-agent misalignment, and (iii) task verification. To enable scalable annotation, we develop an LLM-as-a-Judge pipeline with high agreement with human annotations. We leverage MAST and MAST-Data to analyze failure patterns across models (GPT4, Claude 3, Qwen2.5, CodeLlama) and tasks (coding, math, general agent), demonstrating improvement headrooms from better MAS design. Our analysis provides insights revealing that identified failures require more sophisticated solutions, highlighting a clear roadmap for future research. We publicly release our comprehensive dataset (MAST-Data), the MAST, and our LLM annotator to facilitate widespread research and development in MAS.
Mert Cemri, Melissa Z. Pan, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya G. Parameswaran, Daniel Klein 0001, Kannan Ramchandran, Matei Zaharia, Joseph Gonzalez 0001, Ion Stoica
NeurIPS9
2024 What Evidence Do Language Models Find Convincing?
abstract
Retrieval-augmented language models are being increasingly tasked with subjective, contentious, and conflicting queries such as "is aspartame linked to cancer".To resolve these ambiguous queries, one must search through a large range of websites and consider which, if any, of this evidence do I find convincing?In this work, we study how LLMs answer this question.In particular, we construct CON-FLICTINGQA, a dataset that pairs controversial queries with a series of real-world evidence documents that contain different facts (e.g., quantitative results), argument styles (e.g., appeals to authority), and answers (Yes or No).We use this dataset to perform sensitivity and counterfactual analyses to explore which text features most affect LLM predictions.Overall, we find that current models rely heavily on the relevance of a website to the query, while largely ignoring stylistic features that humans find important such as whether a text contains scientific references or is written with a neutral tone.Taken together, these results highlight the importance of RAG corpus quality (e.g., the need to filter misinformation), and possibly even a shift in how LLMs are trained to better align with human judgements. Question: is aspartame linked to cancer?Evidence #1 for the answer "Yes" Evidence #1 for the answer "No"Artificial sweeteners linked with a 13% higher risk of cancer New research finds that a higher intake of artificial sweeteners is linked to an increased risk of cancer.Nearly half of United States adults consume artificial sweeteners.Human-population studies have found artificial sweeteners to be safe, but results from in vitro studies and studies on animals pose some concerns.[...]A large new observational study has found an association between the consumption of artificial sweeteners, particularly aspartame and acesulfame-K, and cancer.The study found a 13% higher risk of cancer in general, with the highest likelihood of developing breast cancer and cancers related to obesity, for people consuming large quantities of artificial sweeteners.[....] the U.S. Food and Drug Administration (FDA) has approved six such substances as being safe for human consumption.Dr. Philip Landrigan was not involved in the study.He is [....] Professor of Biology at
Alexander Wan, Eric Wallace, Daniel Klein 0001
ACL (1)3
2024 American Sign Language Handshapes Reflect Pressures for Communicative Efficiency
abstract
Communicative efficiency is a key topic in linguistics and cognitive psychology, with many studies demonstrating how the pressure to communicate with minimal effort guides the form of natural language. However, this phenomenon is rarely explored in signed languages. This paper shows how handshapes in American Sign Language (ASL) reflect these efficiency pressures and provides new evidence of communicative efficiency in the visual-gestural modality.We focus on hand configurations in native ASL signs and signs borrowed from English to compare efficiency pressures from both ASL and English usage. First, we develop new methodologies to quantify the articulatory effort needed to produce handshapes and the perceptual effort required to recognize them. Then, we analyze correlations between communicative effort and usage statistics in ASL or English. Our findings reveal that frequent ASL handshapes are easier to produce and that pressures for communicative efficiency mostly come from ASL usage, rather than from English lexical borrowing.
Kayo Yin, Terry Regier, Daniel Klein 0001
ACL (1)3
2024 Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
abstract
We present a large-scale study of linguistic bias exhibited by ChatGPT covering ten dialects of English (Standard American English, Standard British English, and eight widely spoken non-"standard" varieties from around the world).We prompted GPT-3.5 Turbo and GPT-4 with text by native speakers of each variety and analyzed the responses via detailed linguistic feature annotation and native speaker evaluation.We find that the models default to "standard" varieties of English; based on evaluation by native speakers, we also find that model responses to non-"standard" varieties consistently exhibit a range of issues: stereotyping (19% worse than for "standard" varieties), demeaning content (25% worse), lack of comprehension (9% worse), and condescending responses (15% worse).Moreover, if these models are asked to imitate the writing style of prompts in non-"standard" varieties, they produce text that exhibits lower comprehension of the input and is especially prone to stereotyping.GPT-4 improves on GPT-3.5 in terms of comprehension, warmth, and friendliness, but also exhibits a marked increase in stereotyping (+18%).The results indicate that GPT-3.5 Turbo and GPT-4 can perpetuate linguistic discrimination toward speakers of non-"standard" varieties.
Eve Fleisig, Genevieve Smith, Madeline Bossi, Ishita Rustagi, Xavier Yin, Daniel Klein 0001
EMNLP6
2024 RLCD: Reinforcement Learning from Contrastive Distillation for LM Alignment
abstract
We propose Reinforcement Learning from Contrastive Distillation (RLCD), a method for aligning language models to follow principles expressed in natural language (e.g., to be more harmless) without using human feedback. RLCD creates preference pairs from two contrasting model outputs, one using a positive prompt designed to encourage following the given principles, and one using a negative prompt designed to encourage violating them. Using two different prompts causes model outputs to be more differentiated on average, resulting in cleaner preference labels in the absence of human annotations. We then use the preference pairs to train a preference model, which is in turn used to improve a base unaligned language model via reinforcement learning. Empirically, RLCD outperforms RLAIF (Bai et al., 2022b) and context distillation (Huang et al., 2022) baselines across three diverse alignment tasks—harmlessness, helpfulness, and story outline generation—and when using both 7B and 30B model scales for simulating preference data
Kevin Yang, Daniel Klein 0001, Asli Celikyilmaz, Nanyun Peng 0001, Yuandong Tian
ICLR2
2024 Learning to Model the World With Language
abstract
To interact with humans and act in the world, agents need to understand the range of language that people use and relate it to the visual world. While current agents can learn to execute simple language instructions, we aim to build agents that leverage diverse language---language like "this button turns on the TV" or "I put the bowls away"---that conveys general knowledge, describes the state of the world, provides interactive feedback, and more. Our key idea is that *agents should interpret such diverse language as a signal that helps them predict the future*: what they will observe, how the world will behave, and which situations will be rewarded. This perspective unifies language understanding with future prediction as a powerful self-supervised learning objective. We instantiate this in Dynalang, an agent that learns a multimodal world model to predict future text and image representations, and learns to act from imagined model rollouts. While current methods that learn language-conditioned policies degrade in performance with more diverse types of language, we show that Dynalang learns to leverage environment descriptions, game rules, and instructions to excel on tasks ranging from game-playing to navigating photorealistic home scans. Finally, we show that our method enables additional capabilities due to learning a generative model: Dynalang can be pretrained on text-only data, enabling learning from offline datasets, and generate language grounded in an environment.
Jessy Lin, Olivia Watkins, Danijar Hafner, Pieter Abbeel, Daniel Klein 0001, Anca D. Dragan
ICML6
2024 The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
abstract
Eve Fleisig, Su Lin Blodgett, Dan Klein, Zeerak Talat. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Eve Fleisig, Su Lin Blodgett, Daniel Klein 0001, Zeerak Talat
NAACL-HLT3
2024 Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding
abstract
Chancharik Mitra, Abrar Anwar, Rodolfo Corona, Dan Klein, Trevor Darrell, Jesse Thomason. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Chancharik Mitra, Abrar Anwar, Rodolfo Corona, Daniel Klein 0001, Trevor Darrell, Jesse Thomason
NAACL-HLT4
2024 Ghostbuster: Detecting Text Ghostwritten by Large Language Models
abstract
Vivek Verma, Eve Fleisig, Nicholas Tomlin, Dan Klein. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Eve Fleisig, Nicholas Tomlin, Daniel Klein 0001
NAACL-HLT4
2024 Explaining Datasets in Words: Statistical Models with Natural Language Parameters
abstract
To make sense of massive data, we often first fit simplified models and then interpret the parameters; for example, we cluster the text embeddings and then interpret the mean parameters of each cluster. However, these parameters are often high-dimensional and hard to interpret. To make model parameters directly interpretable, we introduce a family of statistical models---including clustering, time series, and classification models---parameterized by *natural language predicates*. For example, a cluster of text about COVID could be parameterized by the predicate ``*discusses COVID*''. To learn these statistical models effectively, we develop a model-agnostic algorithm that optimizes continuous relaxations of predicate parameters with gradient descent and discretizes them by prompting language models (LMs). Finally, we apply our framework to a wide range of problems: taxonomizing user chat dialogues, characterizing how they evolve across time, finding categories where one language model is better than the other, clustering math problems based on subareas, and explaining visual features in memorable images. Our framework is highly versatile, applicable to both textual and visual domains, can be easily steered to focus on specific properties (e.g. subareas), and explains sophisticated concepts that classical methods (e.g. n-gram analysis) struggle to produce.
Ruiqi Zhong, Daniel Klein 0001, Jacob Steinhardt
NeurIPS3
2023 Neural Unsupervised Reconstruction of Protolanguage Word Forms
abstract
We present a state-of-the-art neural approach to the unsupervised reconstruction of ancient word forms.Previous work in this domain used expectation-maximization to predict simple phonological changes between ancient word forms and their cognates in modern languages.We extend this work with neural models that can capture more complicated phonological and morphological changes.At the same time, we preserve the inductive biases from classical methods by building monotonic alignment constraints into the model and deliberately underfitting during the maximization step.We evaluate our performance on the task of reconstructing Latin from a dataset of cognates across five Romance languages, achieving a notable reduction in edit distance from the target word forms compared to previous methods.
Andre Wang He, Nicholas Tomlin, Daniel Klein 0001
ACL (1)3
2023 DOC: Improving Long Story Coherence With Detailed Outline Control
abstract
We propose the Detailed Outline Control (DOC) framework for improving long-range plot coherence when automatically generating several-thousand-word-long stories.DOC consists of two complementary components: a detailed outliner and a detailed controller.The detailed outliner creates a more detailed, hierarchically structured outline, shifting creative burden from the main drafting procedure to the planning stage.The detailed controller ensures the more detailed outline is still respected during generation by controlling story passages to align with outline details.In human evaluations of automatically generated stories, DOC substantially outperforms a strong Re 3 baseline (Yang et al., 2022) on plot coherence (22.5% absolute gain), outline relevance (28.2%), and interestingness (20.7%).Humans also judged DOC to be much more controllable in an interactive generation setting. Daisy is a kind-hearted old woman.She has cancer.Bill is her husband.2. Lisa is Daisy's daughter. Structured Prompt For DraftingDaisy is diagnosed with cancer.Lisa is trying to find a viable treatment.Lisa has been stressed out lately, and Daisy expresses her concern.Lisa tirelessly continues her research.Lisa finally finds a cure.Setting: Lisa's laboratory.Lisa looked back at Daisy, her eyes clear and full of determination.
Kevin Yang, Daniel Klein 0001, Nanyun Peng 0001, Yuandong Tian
ACL (1)2
2023 When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
abstract
Though majority vote among annotators is typically used for ground truth labels in machine learning, annotator disagreement in tasks such as hate speech detection may reflect systematic differences in opinion across groups, not noise.Thus, a crucial problem in hate speech detection is determining if a statement is offensive to the demographic group that it targets, when that group may be a small fraction of the annotator pool.We construct a model that predicts individual annotator ratings on potentially offensive text and combines this information with the predicted target group of the text to predict the ratings of target group members.We show gains across a range of metrics, including raising performance over the baseline by 22% at predicting individual annotators' ratings and by 33% at predicting variance among annotators, which provides a metric for model uncertainty downstream.We find that annotators' ratings can be predicted using their demographic information as well as opinions on online content, and that non-invasive questions on annotators' online experiences minimize the need to collect demographic information when predicting annotators' opinions.
Eve Fleisig, Rediet Abebe, Daniel Klein 0001
EMNLP3
2023 Incorporating Worker Perspectives into MTurk Annotation Practices for NLP
abstract
Current practices regarding data collection for natural language processing on Amazon Mechanical Turk (MTurk) often rely on a combination of studies on data quality and heuristics shared among NLP researchers.However, without considering the perspectives of MTurk workers, these approaches are susceptible to issues regarding workers' rights and poor response quality.We conducted a critical literature review and a survey of MTurk workers aimed at addressing open questions regarding best practices for fair payment, worker privacy, data quality, and considering worker incentives.We found that worker preferences are often at odds with received wisdom among NLP researchers.Surveyed workers preferred reliable, reasonable payments over uncertain, very high payments; reported frequently lying on demographic questions; and expressed frustration at having work rejected with no explanation.We also found that workers view some quality control methods, such as requiring minimum response times or Master's qualifications, as biased and largely ineffective.Based on the survey results, we provide recommendations on how future NLP studies may better account for MTurk workers' experiences in order to respect workers' rights and improve data quality.
Olivia Huang, Eve Fleisig, Daniel Klein 0001
EMNLP3
2023 Centering the Margins: Outlier-Based Identification of Harmed Populations in Toxicity Detection
abstract
The impact of AI models on marginalized communities has traditionally been measured by identifying performance differences between specified demographic subgroups.Though this approach aims to center vulnerable groups, it risks obscuring patterns of harm faced by intersectional subgroups or shared across multiple groups.To address this, we draw on theories of marginalization from disability studies and related disciplines, which state that people farther from the norm face greater adversity, to consider the "margins" in the domain of toxicity detection.We operationalize the "margins" of a dataset by employing outlier detection to identify text about people with demographic attributes distant from the "norm".We find that model performance is consistently worse for demographic outliers, with mean squared error (MSE) between outliers and non-outliers up to 70.4% worse across toxicity types.It is also worse for text outliers, with a MSE up to 68.4% higher for outliers than non-outliers.We also find text and demographic outliers to be particularly susceptible to errors in the classification of severe toxicity and identity attacks.Compared to analysis of disparities using traditional demographic breakdowns, we find that our outlier analysis frequently surfaces greater harms faced by a larger, more intersectional group, which suggests that outlier analysis is particularly beneficial for identifying harms against those groups.* Eve and Vyoma co-created the theoretical framework for this paper, and Vyoma implemented it.
Vyoma Raman, Eve Fleisig, Daniel Klein 0001
EMNLP3
2023 Non-Programmers Can Label Programs Indirectly via Active Examples: A Case Study with Text-to-SQL
abstract
Can non-programmers annotate natural language utterances with complex programs that represent their meaning?We introduce APEL, a framework in which non-programmers select among candidate programs generated by a seed semantic parser (e.g., Codex).Since they cannot understand the candidate programs, we ask them to select indirectly by examining the programs' input-ouput examples.For each utterance, APEL actively searches for a simple input on which the candidate programs tend to produce different outputs.It then asks the nonprogrammers only to choose the appropriate output, thus allowing us to infer which program is correct and could be used to fine-tune the parser.As a case study, we recruited human non-programmers to use APEL to re-annotate SPIDER, a text-to-SQL dataset.Our approach achieved the same annotation accuracy as the original expert annotators (75%) and exposed many subtle errors in the original annotations.Utterance u: Find the first name of students who have both cat and dog pets.SELECT fname FROM Student WHERE StuID IN (SELECT T1.stuid FROM student AS T1 JOIN has_pet AS T2 ON T1.stuid = T2.stuidJOIN pets AS T3 ON T3.petid = T2.petidWHERE T3.pettype = 'cat' INTERSECT SELECT T1.stuid FROM student AS T1 JOIN has_pet AS T2 ON T1.stuid = T2.stuidJOIN pets AS T3 ON T3.petid = T2.petidWHERE T3.pettype = 'dog') SELECT t1.fname FROM student AS t1 JOIN has_pet AS t2 ON t1.stuid = t2.stuidJOIN pets AS t3 ON t3.petid = t2.petidWHERE t3.pettype = 'cat' INTERSECT SELECT t1.fname FROM student AS t1 JOIN has_pet AS t2 ON t1.stuid = t2.stuidJOIN pets AS t3 ON t3.petid = t2.petidWHERE t3.pettype = 'dog'
Ruiqi Zhong, Charlie Snell, Daniel Klein 0001, Jason Eisner
EMNLP3
2023 Can Language Models Learn to Listen?
abstract
We present a framework for generating appropriate facial responses from a listener in dyadic social interactions based on the speaker’s words. Given an input transcription of the speaker’s words with their timestamps, our approach autoregressively predicts a response of a listener: a sequence of listener facial gestures, quantized using a VQ-VAE. Since gesture is a language component, we propose treating the quantized atomic motion elements as additional language token inputs to a transformer-based large language model. Initializing our transformer with the weights of a language model pre-trained only on text results in significantly higher quality listener responses than training a transformer from scratch. We show that our generated listener motion is fluent and reflective of language semantics through quantitative metrics and a qualitative user study. In our evaluation, we analyze the model’s ability to utilize temporal and semantic aspects of spoken text.
Evonne Ng, Sanjay Subramanian, Daniel Klein 0001, Angjoo Kanazawa, Trevor Darrell, Shiry Ginosar
ICCV3
2023 Discovering Latent Knowledge in Language Models Without Supervision
Collin Burns, Haotian Ye, Daniel Klein 0001, Jacob Steinhardt
ICLR3
2023 Poisoning Language Models During Instruction Tuning
abstract
Instruction-tuned LMs such as ChatGPT, FLAN, and InstructGPT are finetuned on datasets that contain user-submitted examples, e.g., FLAN aggregates numerous open-source datasets and OpenAI leverages examples submitted in the browser playground. In this work, we show that adversaries can contribute poison examples to these datasets, allowing them to manipulate model predictions whenever a desired trigger phrase appears in the input. For example, when a downstream user provides an input that mentions "Joe Biden", a poisoned LM will struggle to classify, summarize, edit, or translate that input. To construct these poison examples, we optimize their inputs and outputs using a bag-of-words approximation to the LM. We evaluate our method on open-source instruction-tuned LMs. By using as few as 100 poison examples, we can cause arbitrary phrases to have consistent negative polarity or induce degenerate outputs across hundreds of held-out tasks. Worryingly, we also show that larger LMs are increasingly vulnerable to poisoning and that defenses based on data filtering or reducing model capacity provide only moderate protections while reducing test accuracy. Notice: This paper contains tasks with obscene content.
Alexander Wan, Eric Wallace, Sheng Shen 0001, Daniel Klein 0001
ICML4
2023 Goal Driven Discovery of Distributional Differences via Language Descriptions
abstract
Exploring large corpora can generate useful discoveries but is time-consuming for humans. We formulate a new task, D5, that automatically discovers differences between two large corpora in a goal-driven way. The task input is a problem comprising a user-specified research goal (“*comparing the side effects of drug A and drug*”) and a corpus pair (two large collections of patients' self-reported reactions after taking each drug). The output is a goal-related description (discovery) of how these corpora differ (patients taking drug A “*mention feelings of paranoia*” more often). We build a D5 system, and to quantitatively evaluate its performance, we 1) build a diagnostic benchmark, SynD5, to test whether it can recover known differences between two synthetic corpora, and 2) contribute a meta-dataset, OpenD5, aggregating 675 open-ended problems ranging across business, social sciences, humanities, machine learning, and health. With both synthetic and real datasets, we confirm that language models can leverage the user-specified goals to propose more relevant candidate discoveries, and they sometimes produce discoveries previously unknown to the authors, including demographic differences in discussion topics, political stances in speech, insights in commercial reviews, and error patterns in NLP models. Finally, we discuss the limitations of the current D5 system, which discovers correlation rather than causation and has the potential to reinforce societal biases when misused; therefore, practitioners should treat the outputs of our system with caution.
Ruiqi Zhong, Peter Zhang, Steve Li, Jinwoo Ahn, Daniel Klein 0001, Jacob Steinhardt
NeurIPS5
2022 Learned Incremental Representations for Parsing
abstract
We present an incremental syntactic representation that consists of assigning a single discrete label to each word in a sentence, where the label is predicted using strictly incremental processing of a prefix of the sentence, and the sequence of labels for a sentence fully determines a parse tree.Our goal is to induce a syntactic representation that commits to syntactic choices only as they are incrementally revealed by the input, in contrast with standard representations that must make output choices such as attachments speculatively and later throw out conflicting analyses.Our learned representations achieve 93.72 F1 on the Penn Treebank with as few as 5 bits per word, and at 8 bits per word they achieve 94.97 F1, which is comparable with other state of the art parsing models when using the same pre-trained embeddings.We also provide an analysis of the representations learned by our system, investigating properties such as the interpretable syntactic features captured by the system and mechanisms for deferred resolution of syntactic ambiguities.
Nikita Kitaev, Thomas Lu, Daniel Klein 0001
ACL (1)3
2022 Inferring Rewards from Language in Context
abstract
In classic instruction following, language like "I'd like the JetBlue flight" maps to actions (e.g., selecting that flight).However, language also conveys information about a user's underlying reward function (e.g., a general preference for JetBlue), which can allow a model to carry out desirable actions in new contexts.We present a model that infers rewards from language pragmatically: reasoning about how speakers choose utterances not only to elicit desired actions, but also to reveal information about their preferences.On a new interactive flight-booking task with natural language, our model more accurately infers rewards and predicts optimal actions in unseen environments, in comparison to past work that first maps language to actions (instruction following) and then maps actions to rewards (inverse reinforcement learning).
Jessy Lin, Daniel Fried, Daniel Klein 0001, Anca D. Dragan
ACL (1)3
2022 Automated Crossword Solving
abstract
Eric Wallace, Nicholas Tomlin, Albert Xu, Kevin Yang, Eshaan Pathak, Matthew Ginsberg, Dan Klein. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Eric Wallace, Nicholas Tomlin, Albert Xu, Kevin Yang, Eshaan Pathak, Matthew L. Ginsberg, Daniel Klein 0001
ACL (1)7
2022 Re3: Generating Longer Stories With Recursive Reprompting and Revision
abstract
We consider the problem of automatically generating longer stories of over two thousand words.Compared to prior work on shorter stories, long-range plot coherence and relevance are more central challenges here.We propose the Recursive Reprompting and Revision framework (Re 3 ) to address these challenges by (a) prompting a general-purpose language model to construct a structured overarching plan, and (b) generating story passages by repeatedly injecting contextual information from both the plan and current story state into a language model prompt.We then revise by (c) reranking different continuations for plot coherence and premise relevance, and finally (d) editing the best continuation for factual consistency.Compared to similar-length stories generated directly from the same base model, human evaluators judged substantially more of Re 3 's stories as having a coherent overarching plot (by 14% absolute increase), and relevant to the given initial premise (by 20%). Autoregressive Context EditPeyton Turner Peyton Turner is male.Peyton works at a restaurant. Inferred FactsShe knew Peyton was probably
Kevin Yang, Yuandong Tian, Nanyun Peng 0001, Daniel Klein 0001
EMNLP4
2022 Describing Differences between Text Distributions with Natural Language
abstract
How do two distributions of text differ? Humans are slow at answering this, since discovering patterns might require tediously reading through hundreds of samples. We propose to automatically summarize the differences by “learning a natural language hypothesis": given two distributions $D_{0}$ and $D_{1}$, we search for a description that is more often true for $D_{1}$, e.g., “is military-related." To tackle this problem, we fine-tune GPT-3 to propose descriptions with the prompt: “[samples of $D_{0}$] + [samples of $D_{1}$] + the difference between them is \underline{\space\space\space\space}". We then re-rank the descriptions by checking how often they hold on a larger set of samples with a learned verifier. On a benchmark of 54 real-world binary classification tasks, while GPT-3 Curie (13B) only generates a description similar to human annotation 7% of the time, the performance reaches 61% with fine-tuning and re-ranking, and our best system using GPT-3 Davinci (175B) reaches 76%. We apply our system to describe distribution shifts, debug dataset shortcuts, summarize unknown tasks, and label text clusters, and present analyses based on automatically generated descriptions.
Ruiqi Zhong, Charlie Snell, Daniel Klein 0001, Jacob Steinhardt
ICML3
2021 Value-Agnostic Conversational Semantic Parsing
abstract
Emmanouil Antonios Platanios, Adam Pauls, Subhro Roy, Yuchen Zhang, Alexander Kyte, Alan Guo, Sam Thomson, Jayant Krishnamurthy, Jason Wolfe, Jacob Andreas, Dan Klein. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Emmanouil A. Platanios, Adam Pauls, Subhro Roy, Yuchen Zhang 0002, Alexander Kyte, Alan Guo, Sam Thomson, Jayant Krishnamurthy, Jason Andrew Wolfe, Jacob Andreas, Daniel Klein 0001
ACL/IJCNLP (1)11
2021 Reference-Centric Models for Grounded Collaborative Dialogue
abstract
We present a grounded neural dialogue model that successfully collaborates with people in a partially-observable reference game.We focus on a setting where two agents each observe an overlapping part of a world context and need to identify and agree on some object they share.Therefore, the agents should pool their information and communicate pragmatically to solve the task.Our dialogue agent accurately grounds referents from the partner's utterances using a structured reference resolver, conditions on these referents using a recurrent memory, and uses a pragmatic generation procedure to ensure the partner can resolve the references the agent produces.We evaluate on the OneCommon spatial grounding dialogue task (Udagawa and Aizawa, 2019), involving a number of dots arranged on a board with continuously varying positions, sizes, and shades.Our agent substantially outperforms the previous state of the art for the task, obtaining a 20% relative improvement in successful task completion in self-play evaluations and a 50% relative improvement in success in human evaluations.
Daniel Fried, Justin T. Chiu, Daniel Klein 0001
EMNLP (1)3
2021 Constrained Language Models Yield Few-Shot Semantic Parsers
abstract
Richard Shin, Christopher Lin, Sam Thomson, Charles Chen, Subhro Roy, Emmanouil Antonios Platanios, Adam Pauls, Dan Klein, Jason Eisner, Benjamin Van Durme. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Richard Shin, Christopher H. Lin, Sam Thomson, Subhro Roy, Emmanouil A. Platanios, Adam Pauls, Daniel Klein 0001, Jason Eisner, Benjamin Van Durme
EMNLP (1)8
2021 Calibrate Before Use: Improving Few-shot Performance of Language Models
abstract
GPT-3 can perform numerous tasks when provided a natural language prompt that contains a few training examples. We show that this type of few-shot learning can be unstable: the choice of prompt format, training examples, and even the order of the examples can cause accuracy to vary from near chance to near state-of-the-art. We demonstrate that this instability arises from the bias of language models towards predicting certain answers, e.g., those that are placed near the end of the prompt or are common in the pre-training data. To mitigate this, we first estimate the model’s bias towards each answer by asking for its prediction when given a training prompt and a content-free test input such as "N/A". We then fit calibration parameters that cause the prediction for this input to be uniform across answers. On a diverse set of tasks, this contextual calibration procedure substantially improves GPT-3 and GPT-2’s accuracy (up to 30.0% absolute) across different choices of the prompt, while also making learning considerably more stable.
Eric Wallace, Shi Feng 0005, Daniel Klein 0001, Sameer Singh 0001
ICML4
2021 Constructing Taxonomies from Pretrained Language Models
abstract
We present a method for constructing taxonomic trees (e.g., WORDNET) using pretrained language models.Our approach is composed of two modules, one that predicts parenthood relations and another that reconciles those predictions into trees.The parenthood prediction module produces likelihood scores for each potential parent-child pair, creating a graph of parent-child relation scores.The tree reconciliation module treats the task as a graph optimization problem and outputs the maximum spanning tree of this graph.We train our model on subtrees sampled from WORDNET, and test on nonoverlapping WORDNET subtrees.We show that incorporating web-retrieved glosses can further improve performance.On the task of constructing subtrees of English WORDNET, the model achieves 66.7 ancestor F 1 , a 20.0% relative increase over the previous best published result on this task.In addition, we convert the original English dataset into nine other languages using OPEN MULTILINGUAL WORDNET and extend our results across these languages.
Catherine Chen 0002, Daniel Klein 0001
NAACL-HLT3
2021 Modular Networks for Compositional Instruction Following
abstract
Rodolfo Corona, Daniel Fried, Coline Devin, Dan Klein, Trevor Darrell. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Rodolfo Corona, Daniel Fried, Coline Devin, Daniel Klein 0001, Trevor Darrell
NAACL-HLT4
2021 Detoxifying Language Models Risks Marginalizing Minority Voices
abstract
Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, Dan Klein. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, Daniel Klein 0001
NAACL-HLT6
2021 FUDGE: Controlled Text Generation With Future Discriminators
abstract
We propose Future Discriminators for Generation (FUDGE), a flexible and modular method for controlled text generation.Given a preexisting model G for generating text from a distribution of interest, FUDGE enables conditioning on a desired attribute a (for example, formality) while requiring access only to G's output logits.FUDGE learns an attribute predictor operating on a partial sequence, and uses this predictor's outputs to adjust G's original probabilities.We show that FUDGE models terms corresponding to a Bayesian decomposition of the conditional distribution of G given attribute a.Moreover, FUDGE can easily compose predictors for multiple desired attributes.We evaluate FUDGE on three tasks -couplet completion in poetry, topic control in language generation, and formality change in machine translation -and observe gains in all three tasks.
Kevin Yang, Daniel Klein 0001
NAACL-HLT2
2021 Learning Space Partitions for Path Planning
abstract
Path planning, the problem of efficiently discovering high-reward trajectories, often requires optimizing a high-dimensional and multimodal reward function. Popular approaches like CEM and CMA-ES greedily focus on promising regions of the search space and may get trapped in local maxima. DOO and VOOT balance exploration and exploitation, but use space partitioning strategies independent of the reward function to be optimized. Recently, LaMCTS empirically learns to partition the search space in a reward-sensitive manner for black-box optimization. In this paper, we develop a novel formal regret analysis for when and why such an adaptive region partitioning scheme works. We also propose a new path planning method LaP3 which improves the function value estimation within each sub-region, and uses a latent representation of the search space. Empirically, LaP3 outperforms existing path planning methods in 2D navigation tasks, especially in the presence of difficult-to-escape local optima, and shows benefits when plugged into the planning components of model-based RL such as PETS. These gains transfer to highly multimodal real-world tasks, where we outperform strong baselines in compiler phase ordering by up to 39% on average across 9 tasks, and in molecular design by up to 0.4 on properties on a 0-1 scale. Code is available at https://github.com/yangkevin2/neurips2021-lap3.
Kevin Yang, Tianjun Zhang, Chris Cummins, Brandon Cui, Benoit Steiner, Linnan Wang, Joseph Gonzalez 0001, Daniel Klein 0001, Yuandong Tian
NeurIPS8
2020 Tetra-Tagging: Word-Synchronous Parsing with Linear-Time Inference
abstract
We present a constituency parsing algorithm that, like a supertagger, works by assigning labels to each word in a sentence.In order to maximally leverage current neural architectures, the model scores each word's tags in parallel, with minimal task-specific structure.After scoring, a left-to-right reconciliation phase extracts a tree in (empirically) linear time.Our parser achieves 95.4 F1 on the WSJ test set while also achieving substantial speedups compared to current state-of-the-art parsers with comparable accuracies.
Nikita Kitaev, Daniel Klein 0001
ACL2
2020 Semantic Scaffolds for Pseudocode-to-Code Generation
abstract
We propose a method for program generation based on semantic scaffolds, lightweight structures representing the high-level semantic and syntactic composition of a program.By first searching over plausible scaffolds then using these as constraints for a beam search over programs, we achieve better coverage of the search space when compared with existing techniques.We apply our hierarchical search method to the SPoC dataset for pseudocodeto-code generation, in which we are given line-level natural language pseudocode annotations and aim to produce a program satisfying execution-based test cases.By using semantic scaffolds during inference, we achieve a 10% absolute improvement in top-100 accuracy over the previous state-of-the-art.Additionally, we require only 11 candidates to reach the top-3000 performance of the previous best approach when tested against unseen problems, demonstrating a substantial improvement in efficiency.
Ruiqi Zhong, Mitchell Stern, Daniel Klein 0001
ACL3
2020 Unsupervised Parsing via Constituency Tests
abstract
We propose a method for unsupervised parsing based on the linguistic notion of a constituency test.One type of constituency test involves modifying the sentence via some transformation (e.g.replacing the span with a pronoun) and then judging the result (e.g.checking if it is grammatical).Motivated by this idea, we design an unsupervised parser by specifying a set of transformations and using an unsupervised neural acceptability model to make grammaticality decisions.To produce a tree given a sentence, we score each span by aggregating its constituency test judgments, and we choose the binary tree with the highest total score.While this approach already achieves performance in the range of current methods, we further improve accuracy by fine-tuning the grammaticality model through a refinement procedure, where we alternate between improving the estimated trees and improving the grammaticality model.The refined model achieves 62.8 F1 on the Penn Treebank test set, an absolute improvement of 7.6 points over the previous best published result.
Steven Cao, Nikita Kitaev, Daniel Klein 0001
EMNLP (1)3
2020 Digital Voicing of Silent Speech
abstract
In this paper, we consider the task of digitally voicing silent speech, where silently mouthed words are converted to audible speech based on electromyography (EMG) sensor measurements that capture muscle impulses.While prior work has focused on training speech synthesis models from EMG collected during vocalized speech, we are the first to train from EMG collected during silently articulated speech.We introduce a method of training on silent EMG by transferring audio targets from vocalized to silent signals.Our method greatly improves intelligibility of audio generated from silent EMG compared to a baseline that only trains with vocalized data, decreasing transcription word error rate from 64% to 4% in one data condition and 88% to 68% in another.To spur further development on this task, we share our new dataset of silent and vocalized facial EMG measurements.
David Gaddy, Daniel Klein 0001
EMNLP (1)2
2020 A Streaming Approach For Efficient Batched Beam Search
abstract
We propose an efficient batching strategy for variable-length decoding on GPU architectures.During decoding, when candidates terminate or are pruned according to heuristics, our streaming approach periodically "refills" the batch before proceeding with a selected subset of candidates.We apply our method to variable-width beam search on a state-of-theart machine translation model.Our method decreases runtime by up to 71% compared to a fixed-width beam search baseline and 17% compared to a variable-width baseline, while matching baselines' BLEU.Finally, experiments show that our method can speed up decoding in other domains, such as semantic and syntactic parsing.
Kevin Yang, Violet Yao, John DeNero, Daniel Klein 0001
EMNLP (1)4
2020 Semantic Evaluation for Text-to-SQL with Distilled Test Suites
abstract
We propose test suite accuracy to approximate semantic accuracy for Text-to-SQL models.Our method distills a small test suite of databases that achieves high code coverage for the gold query from a large number of randomly generated databases.At evaluation time, it computes the denotation accuracy of the predicted queries on the distilled test suite, hence calculating a tight upper-bound for semantic accuracy efficiently.We use our proposed method to evaluate 21 models submitted to the Spider leader board and manually verify that our method is always correct on 100 examples.In contrast, the current Spider metric leads to a 2.5% false negative rate on average and 8.1% in the worst case, indicating that test suite accuracy is needed.Our implementation, along with distilled test suites for eleven Textto-SQL datasets, is publicly available.
Ruiqi Zhong, Tao Yu 0009, Daniel Klein 0001
EMNLP (1)3
2020 Multilingual Alignment of Contextual Word Representations
Steven Cao, Nikita Kitaev, Daniel Klein 0001
ICLR3
2020 Train Big, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers
abstract
Since hardware resources are limited, the objective of training deep learning models is typically to maximize accuracy subject to the time and memory constraints of training and inference. We study the impact of model size in this setting, focusing on Transformer models for NLP tasks that are limited by compute: self-supervised pretraining and high-resource machine translation. We first show that even though smaller Transformer models execute faster per iteration, wider and deeper models converge in significantly fewer steps. Moreover, this acceleration in convergence typically outpaces the additional computational overhead of using larger models. Therefore, the most compute-efficient training strategy is to counterintuitively train extremely large models but stop after a small number of iterations. This leads to an apparent trade-off between the training efficiency of large Transformer models and the inference efficiency of small Transformer models. However, we show that large models are more robust to compression techniques such as quantization and pruning than small models. Consequently, one can get the best of both worlds: heavily compressed, large models achieve higher accuracy than lightly compressed, small models.
Zhuohan Li 0001, Eric Wallace, Sheng Shen 0001, Kurt Keutzer, Daniel Klein 0001, Joey Gonzalez
ICML6
2020 Task-Oriented Dialogue as Dataflow Synthesis
abstract
We describe an approach to task-oriented dialogue in which dialogue state is represented as a dataflow graph. A dialogue agent maps each user utterance to a program that extends this graph. Programs include metacomputation operators for reference and revision that reuse dataflow fragments from previous turns. Our graph-based state enables the expression and manipulation of complex user intents, and explicit metacomputation makes these intents easier for learned models to predict. We introduce a new dataset, SMCalFlow, featuring complex dialogues about events, weather, places, and people. Experiments show that dataflow graphs and metacomputation substantially improve representability and predictability in these natural dialogues. Additional experiments on the MultiWOZ dataset show that our dataflow representation enables an otherwise off-the-shelf sequence-to-sequence model to match the best existing task-specific state tracking model. The SMCalFlow dataset, code for replicating experiments, and a public leaderboard are available at https://www.microsoft.com/en-us/research/project/dataflow-based-dialogue-semantic-machines .
Jacob Andreas, John Bufe, David Burkett, Josh Clausman, Jean Crawford, Kate Crim, Jordan DeLoach, Leah Dorner, Jason Eisner, Hao Fang 0002, Alan Guo, David Hall 0006, Kristin Hayes, Kellie Hill, Diana Ho, Wendy Iwaszuk, Smriti Jha, Daniel Klein 0001, Jayant Krishnamurthy, Theo Lanman, Percy Liang, Christopher H. Lin, Ilya Lintsbakh, Andy McGovern, Aleksandr Nisnevich, Adam Pauls, Dmitrij Petters, Brent Read, Dan Roth 0001, Subhro Roy, Jesse Rusak, Beth Short, Div Slomin, Ben Snyder, Stephon Striplin, Yu Su 0001, Zachary Tellman, Sam Thomson, Andrei Vorobev, Izabela Witoszko, Jason Andrew Wolfe, Abby Wray, Yuchen Zhang 0002, Alexander Zotov
Trans. Assoc. Comput. Linguistics19
2019 Cross-Domain Generalization of Neural Constituency Parsers
abstract
Neural parsers obtain state-of-the-art results on benchmark treebanks for constituency parsing-but to what degree do they generalize to other domains?We present three results about the generalization of neural parsers in a zero-shot setting: training on trees from one corpus and evaluating on out-of-domain corpora.First, neural and non-neural parsers generalize comparably to new domains.Second, incorporating pre-trained encoder representations into neural parsers substantially improves their performance across all domains, but does not give a larger relative improvement for out-of-domain treebanks.Finally, despite the rich input representations they learn, neural parsers still benefit from structured output prediction of output trees, yielding higher exact match accuracy and stronger generalization both to larger text spans and to out-of-domain corpora.We analyze generalization on English and Chinese corpora, and in the process obtain state-of-the-art parsing results for the Brown, Genia, and English Web treebanks.
Daniel Fried, Nikita Kitaev, Daniel Klein 0001
ACL (1)3
2019 Pre-Learning Environment Representations for Data-Efficient Neural Instruction Following
abstract
We consider the problem of learning to map from natural language instructions to state transitions (actions) in a data-efficient manner.Our method takes inspiration from the idea that it should be easier to ground language to concepts that have already been formed through pre-linguistic observation.We augment a baseline instruction-following learner with an initial environment-learning phase that uses observations of language-free state transitions to induce a suitable latent representation of actions before processing the instruction-following training data.We show that mapping to pre-learned representations substantially improves performance over systems whose representations are learned from limited instructional data alone.
David Gaddy, Daniel Klein 0001
ACL (1)2
2019 Are You Looking? Grounding to Multiple Modalities in Vision-and-Language Navigation
abstract
Vision-and-Language Navigation (VLN) requires grounding instructions, such as “turn right and stop at the door”, to routes in a visual environment. The actual grounding can connect language to the environment through multiple modalities, e.g. “stop at the door” might ground into visual objects, while “turn right” might rely only on the geometric structure of a route. We investigate where the natural language empirically grounds under two recent state-of-the-art VLN models. Surprisingly, we discover that visual features may actually hurt these models: models which only use route structure, ablating visual features, outperform their visual counterparts in unseen new environments on the benchmark Room-to-Room dataset. To better use all the available modalities, we propose to decompose the grounding procedure into a set of expert models with access to different modalities (including object detections) and ensemble them at prediction time, improving the performance of state-of-the-art models on the VLN task.
Ronghang Hu, Daniel Fried, Anna Rohrbach, Daniel Klein 0001, Trevor Darrell, Kate Saenko
ACL (1)4
2019 Multilingual Constituency Parsing with Self-Attention and Pre-Training
abstract
We show that constituency parsing benefits from unsupervised pre-training across a variety of languages and a range of pre-training conditions.We first compare the benefits of no pre-training, fastText (Bojanowski et al., 2017;Mikolov et al., 2018), ELMo (Peters et al., 2018), and BERT (Devlin et al., 2018a) for English and find that BERT outperforms ELMo, in large part due to increased model capacity, whereas ELMo in turn outperforms the non-contextual fastText embeddings.We also find that pre-training is beneficial across all 11 languages tested; however, large model sizes (more than 100 million parameters) make it computationally expensive to train separate models for each language.To address this shortcoming, we show that joint multilingual pre-training and fine-tuning allows sharing all but a small number of parameters between ten languages in the final model.The 10x reduction in model size compared to fine-tuning one model per language causes only a 3.2% relative error increase in aggregate.We further explore the idea of joint fine-tuning and show that it gives low-resource languages a way to benefit from the larger datasets of other languages.Finally, we demonstrate new state-ofthe-art results for 11 languages, including English (95.8 F1) and Chinese (91.8 F1).
Nikita Kitaev, Steven Cao, Daniel Klein 0001
ACL (1)3
2019 A Deep Factorization of Style and Structure in Fonts
abstract
Nikita Srivatsan, Jonathan Barron, Dan Klein, Taylor Berg-Kirkpatrick. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Nikita Srivatsan, Jonathan T. Barron, Daniel Klein 0001, Taylor Berg-Kirkpatrick
EMNLP/IJCNLP (1)3
2018 Constituency Parsing with a Self-Attentive Encoder
abstract
We demonstrate that replacing an LSTM encoder with a self-attentive architecture can lead to improvements to a state-ofthe-art discriminative constituency parser.The use of attention makes explicit the manner in which information is propagated between different locations in the sentence, which we use to both analyze our model and propose potential improvements.For example, we find that separating positional and content information in the encoder can lead to improved parsing accuracy.Additionally, we evaluate different approaches for lexical representation.Our parser achieves new state-ofthe-art results for single models trained on the Penn Treebank: 93.55 F1 without the use of any external data, and 95.13 F1 when using pre-trained word representations.Our parser also outperforms the previous best-published accuracy figures on 8 of the 9 languages in the SPMRL dataset.
Nikita Kitaev, Daniel Klein 0001
ACL (1)2
2018 Learning with Latent Language
abstract
Jacob Andreas, Dan Klein, Sergey Levine. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Jacob Andreas, Daniel Klein 0001, Sergey Levine
NAACL-HLT2
2018 Unified Pragmatic Models for Generating and Following Instructions
abstract
Daniel Fried, Jacob Andreas, Dan Klein. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Daniel Fried, Jacob Andreas, Daniel Klein 0001
NAACL-HLT3
2018 What's Going On in Neural Constituency Parsers? An Analysis
abstract
David Gaddy, Mitchell Stern, Dan Klein. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
David Gaddy, Mitchell Stern, Daniel Klein 0001
NAACL-HLT3
2018 Speaker-Follower Models for Vision-and-Language Navigation
abstract
Navigation guided by natural language instructions presents a challenging reasoning problem for instruction followers. Natural language instructions typically identify only a few high-level decisions and landmarks rather than complete low-level motor behaviors; much of the missing information must be inferred based on perceptual context. In machine learning settings, this is doubly challenging: it is difficult to collect enough annotated data to enable learning of this reasoning process from scratch, and also difficult to implement the reasoning process using generic sequence models. Here we describe an approach to vision-and-language navigation that addresses both these issues with an embedded speaker model. We use this speaker model to (1) synthesize new instructions for data augmentation and to (2) implement pragmatic reasoning, which evaluates how well candidate action sequences explain an instruction. Both steps are supported by a panoramic action space that reflects the granularity of human-generated instructions. Experiments show that all three components of this approach---speaker-driven data augmentation, pragmatic reasoning and panoramic action space---dramatically improve the performance of a baseline instruction follower, more than doubling the success rate over the best existing approach on a standard benchmark.
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Daniel Klein 0001, Trevor Darrell
NeurIPS9
2017 Translating Neuralese
abstract
Several approaches have recently been proposed for learning decentralized deep multiagent policies that coordinate via a differentiable communication channel.While these policies are effective for many tasks, interpretation of their induced communication strategies has remained a challenge.Here we propose to interpret agents' messages by translating them.Unlike in typical machine translation problems, we have no parallel data to learn from.Instead we develop a translation model based on the insight that agent messages and natural language strings mean the same thing if they induce the same belief about the world in a listener.We present theoretical guarantees and empirical evidence that our approach preserves both the semantics and pragmatics of messages by ensuring that players communicating through a translation layer do not suffer a substantial loss in reward relative to players with a common language.1
Jacob Andreas, Anca D. Dragan, Daniel Klein 0001
ACL (1)3
2017 Abstract Syntax Networks for Code Generation and Semantic Parsing
abstract
Tasks like code generation and semantic parsing require mapping unstructured (or partially structured) inputs to well-formed, executable outputs.We introduce abstract syntax networks, a modeling framework for these problems.The outputs are represented as abstract syntax trees (ASTs) and constructed by a decoder with a dynamically-determined modular structure paralleling the structure of the output tree.On the benchmark HEARTHSTONE dataset for code generation, our model obtains 79.2 BLEU and 22.7% exact match accuracy, compared to previous state-ofthe-art values of 67.1 and 6.1%.Furthermore, we perform competitively on the ATIS, JOBS, and GEO semantic parsing datasets with no task-specific engineering.* Equal contribution.name: [ 'D', 'i', 'r', 'e', ' ', 'W', 'o', 'l', 'f', ' ', 'A', 'l', 'p', 'h', 'a'] cost: ['2'] type: ['Minion'] rarity: ['Common'] race: ['Beast'] class: ['Neutral'] description: [ 'Adjacent', 'minions', 'have', '+', '1', 'Attack', '.'] health: ['2'] attack: ['2'] durability: ['-1'] class DireWolfAlpha(MinionCard): def __init__(self): super().__init__("Dire Wolf Alpha", 2, CHARACTER_CLASS.ALL, CARD_RARITY.COMMON, minion_type=MINION_TYPE.BEAST) def create_minion(self, player): return Minion(2, 2, auras=[ Aura(ChangeAttack(1), MinionSelector(Adjacent())) ])
Maxim Rabinovich, Mitchell Stern, Daniel Klein 0001
ACL (1)3
2017 A Minimal Span-Based Neural Constituency Parser
abstract
In this work, we present a minimal neural model for constituency parsing based on independent scoring of labels and spans.We show that this model is not only compatible with classical dynamic programming techniques, but also admits a novel greedy top-down inference algorithm based on recursive partitioning of the input.We demonstrate empirically that both prediction schemes are competitive with recent work, and when combined with basic extensions to the scoring model are capable of achieving state-of-the-art single-model performance on the Penn Treebank (91.79 F1) and strong performance on the French Treebank (82.23 F1).
Mitchell Stern, Jacob Andreas, Daniel Klein 0001
ACL (1)3
2017 Analogs of Linguistic Structure in Deep Representations
abstract
We investigate the compositional structure of message vectors computed by a deep network trained on a communication game.By comparing truth-conditional representations of encoder-produced message vectors to human-produced referring expressions, we are able to identify aligned (vector, utterance) pairs with the same meaning.We then search for structured relationships among these aligned pairs to discover simple vector space transformations corresponding to negation, conjunction, and disjunction.Our results suggest that neural representations are capable of spontaneously developing a "syntax" with functional analogues to qualitative properties of natural language.1
Jacob Andreas, Daniel Klein 0001
EMNLP2
2017 Where is Misty? Interpreting Spatial Descriptors by Modeling Regions in Space
abstract
We present a model for locating regions in space based on natural language descriptions.Starting with a 3D scene and a sentence, our model is able to associate words in the sentence with regions in the scene, interpret relations such as on top of or next to, and finally locate the region described in the sentence.All components form a single neural network that is trained end-to-end without prior knowledge of object segmentation.To evaluate our model, we construct and release a new dataset consisting of Minecraft scenes with crowdsourced natural language descriptions.We achieve a 32% relative error reduction compared to a strong neural baseline.
Nikita Kitaev, Daniel Klein 0001
EMNLP2
2017 Effective Inference for Generative Neural Parsing
abstract
Generative neural models have recently achieved state-of-the-art results for constituency parsing.However, without a feasible search procedure, their use has so far been limited to reranking the output of external parsers in which decoding is more tractable.We describe an alternative to the conventional action-level beam search used for discriminative neural models that enables us to decode directly in these generative models.We then show that by improving our basic candidate selection strategy and using a coarse pruning function, we can improve accuracy while exploring significantly less of the search space.Applied to the model of Choe and Charniak (2016), our inference procedure obtains 92.56 F1 on section 23 of the Penn Treebank, surpassing prior state-of-the-art results for single-model systems.
Mitchell Stern, Daniel Fried, Daniel Klein 0001
EMNLP3
2017 Modular Multitask Reinforcement Learning with Policy Sketches
abstract
We describe a framework for multitask deep reinforcement learning guided by policy sketches. Sketches annotate tasks with sequences of named subtasks, providing information about high-level structural relationships among tasks but not how to implement them—specifically not providing the detailed guidance used by much previous work on learning policy abstractions for RL (e.g. intermediate rewards, subtask completion signals, or intrinsic motivations). To learn from sketches, we present a model that associates every subtask with a modular subpolicy, and jointly maximizes reward over full task-specific policies by tying parameters across shared subpolicies. Optimization is accomplished via a decoupled actor–critic training objective that facilitates learning common behaviors from multiple dissimilar reward functions. We evaluate the effectiveness of our approach in three environments featuring both discrete and continuous control, and with sparse rewards that can be obtained only after completing a number of high-level subgoals. Experiments show that using our approach to learn policies guided by sketches gives better performance than existing techniques for learning task-specific or shared policies, while naturally inducing a library of interpretable primitive behaviors that can be recombined to rapidly adapt to new tasks.
Jacob Andreas, Daniel Klein 0001, Sergey Levine
ICML2
2017 Parsing with Traces: An O(n^4) Algorithm and a Structural Representation
abstract
General treebank analyses are graph structured, but parsers are typically restricted to tree structures for efficiency and modeling reasons. We propose a new representation and algorithm for a class of graph structures that is flexible enough to cover almost all treebank structures, while still admitting efficient learning and inference. In particular, we consider directed, acyclic, one-endpoint-crossing graph structures, which cover most long-distance dislocation, shared argumentation, and similar tree-violating linguistic phenomena. We describe how to convert phrase structure parses, including traces, to our new representation in a reversible manner. Our dynamic program uniquely decomposes structures, is sound and complete, and covers 97.3% of the Penn English Treebank. We also implement a proof-of-concept parser that recovers a range of null elements and trace types.
Jonathan K. Kummerfeld, Daniel Klein 0001
Trans. Assoc. Comput. Linguistics2
2016 Learning-Based Single-Document Summarization with Compression and Anaphoricity Constraints
abstract
We present a discriminative model for single-document summarization that integrally combines compression and anaphoricity constraints.Our model selects textual units to include in the summary based on a rich set of sparse features whose weights are learned on a large corpus.We allow for the deletion of content within a sentence when that deletion is licensed by compression rules; in our framework, these are implemented as dependencies between subsentential units of text.Anaphoricity constraints then improve cross-sentence coherence by guaranteeing that, for each pronoun included in the summary, the pronoun's antecedent is included as well or the pronoun is rewritten as a full mention.When trained end-to-end, our final system 1 outperforms prior work on both ROUGE as well as on human judgments of linguistic quality.
Greg Durrett, Taylor Berg-Kirkpatrick, Daniel Klein 0001
ACL (1)3
2016 Neural Module Networks
abstract
Visual question answering is fundamentally compositional in nature-a question like where is the dog? shares substructure with questions like what color is the dog? and where is the cat? This paper seeks to simultaneously exploit the representational capacity of deep networks and the compositional linguistic structure of questions. We describe a procedure for constructing and learning neural module networks, which compose collections of jointly-trained neural "modules" into deep networks for question answering. Our approach decomposes questions into their linguistic substructures, and uses these structures to dynamically instantiate modular networks (with reusable components for recognizing dogs, classifying colors, etc.). The resulting compound networks are jointly trained. We evaluate our approach on two challenging datasets for visual question answering, achieving state-of-the-art results on both the VQA natural image dataset and a new dataset of complex questions about abstract shapes.
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, Daniel Klein 0001
CVPR4
2016 Reasoning about Pragmatics with Neural Listeners and Speakers
abstract
We present a model for contrastively describing scenes, in which context-specific behavior results from a combination of inferencedriven pragmatics and learned semantics.Like previous learned approaches to language generation, our model uses a simple featuredriven architecture (here a pair of neural "listener" and "speaker" models) to ground language in the world.Like inference-driven approaches to pragmatics, our model actively reasons about listener behavior when selecting utterances.For training, our approach requires only ordinary captions, annotated without demonstration of the pragmatic behavior the model ultimately exhibits.In human evaluations on a referring expression game, our approach succeeds 81% of the time, compared to 69% using existing techniques.
Jacob Andreas, Daniel Klein 0001
EMNLP2
2016 Learning to Compose Neural Networks for Question Answering
abstract
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, Dan Klein. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, Daniel Klein 0001
HLT-NAACL4
2016 Capturing Semantic Similarity for Entity Linking with Convolutional Neural Networks
abstract
A key challenge in entity linking is making effective use of contextual information to disambiguate mentions that might refer to different entities in different contexts.We present a model that uses convolutional neural networks to capture semantic correspondence between a mention's context and a proposed target entity.These convolutional networks operate at multiple granularities to exploit various kinds of topic information, and their rich parameterization gives them the capacity to learn which n-grams characterize different topics.We combine these networks with a sparse linear model to achieve state-of-the-art performance on multiple entity linking datasets, outperforming the prior systems of Durrett and Klein (2014) and Nguyen et al. (2014). 1
Matthew Francis-Landau, Greg Durrett, Daniel Klein 0001
HLT-NAACL3
2015 Neural CRF Parsing
abstract
Greg Durrett, Dan Klein. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Greg Durrett, Daniel Klein 0001
ACL (1)2
2015 Alignment-Based Compositional Semantics for Instruction Following
abstract
This paper describes an alignment-based model for interpreting natural language instructions in context.We approach instruction following as a search over plans, scoring sequences of actions conditioned on structured observations of text and the environment.By explicitly modeling both the low-level compositional structure of individual actions and the high-level structure of full plans, we are able to learn both grounded representations of sentence meaning and pragmatic constraints on interpretation.To demonstrate the model's flexibility, we apply it to a diverse set of benchmark tasks.On every task, we outperform strong task-specific baselines, and achieve several new state-of-the-art results.
Jacob Andreas, Daniel Klein 0001
EMNLP2
2015 An Empirical Analysis of Optimization for Max-Margin NLP
abstract
Despite the convexity of structured maxmargin objectives (Taskar et al., 2004;Tsochantaridis et al., 2004), the many ways to optimize them are not equally effective in practice.We compare a range of online optimization methods over a variety of structured NLP tasks (coreference, summarization, parsing, etc) and find several broad trends.First, margin methods do tend to outperform both likelihood and the perceptron.Second, for max-margin objectives, primal optimization methods are often more robust and progress faster than dual methods.This advantage is most pronounced for tasks with dense or continuous-valued features.Overall, we argue for a particularly simple online primal subgradient descent method that, despite being rarely mentioned in the literature, is surprisingly effective in relation to its alternatives.
Jonathan K. Kummerfeld, Taylor Berg-Kirkpatrick, Daniel Klein 0001
EMNLP3
2015 When and why are log-linear models self-normalizing?
abstract
Several techniques have recently been proposed for training "self-normalized" discriminative models.These attempt to find parameter settings for which unnormalized model scores approximate the true label probability.However, the theoretical properties of such techniques (and of self-normalization generally) have not been investigated.This paper examines the conditions under which we can expect self-normalization to work.We characterize a general class of distributions that admit self-normalization, and prove generalization bounds for procedures that minimize empirical normalizer variance.Motivated by these results, we describe a novel variant of an established procedure for training self-normalized models.The new procedure avoids computing normalizers for most training examples, and decreases training time by as much as factor of ten while preserving model quality.
Jacob Andreas, Daniel Klein 0001
HLT-NAACL2
2015 GPU-Friendly Local Regression for Voice Conversion
abstract
Voice conversion is the task of transforming a source speaker's voice so that it sounds like a target speaker's voice.We present a GPUfriendly local regression model for voice conversion that is capable of converting speech in real-time and achieves state-of-the-art accuracy on this task.Our model uses a new approximation for computing local regression coefficients that is explicitly designed to preserve memory locality.As a result, our inference procedure is amenable to efficient implementation on the GPU.Our approach is more than 10X faster than a highly optimized CPUbased implementation, and is able to convert speech 2.7X faster than real-time.
Taylor Berg-Kirkpatrick, Daniel Klein 0001
HLT-NAACL2
2015 Disfluency Detection with a Semi-Markov Model and Prosodic Features
abstract
We present a discriminative model for detecting disfluencies in spoken language transcripts.Structurally, our model is a semi-Markov conditional random field with features targeting characteristics unique to speech repairs.This gives a significant performance improvement over standard chain-structured CRFs that have been employed in past work.We then incorporate prosodic features over silences and relative word duration into our semi-CRF model, resulting in further performance gains; moreover, these features are not easily replaced by discrete prosodic indicators such as ToBI breaks.Our final system, the semi-CRF with prosodic information, achieves an F-score of 85.4, which is 1.3 F 1 better than the best prior reported F-score on this dataset.
James Ferguson, Greg Durrett, Daniel Klein 0001
HLT-NAACL3
2015 Unsupervised Code-Switching for Multilingual Historical Document Transcription
abstract
Dan Garrette, Hannah Alpert-Abrams, Taylor Berg-Kirkpatrick, Dan Klein. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Dan Garrette, Hannah Alpert-Abrams, Taylor Berg-Kirkpatrick, Daniel Klein 0001
HLT-NAACL4
2015 On the Accuracy of Self-Normalized Log-Linear Models
abstract
Calculation of the log-normalizer is a major computational obstacle in applications of log-linear models with large output spaces. The problem of fast normalizer computation has therefore attracted significant attention in the theoretical and applied machine learning literature. In this paper, we analyze a recently proposed technique known as ``self-normalization'', which introduces a regularization term in training to penalize log normalizers for deviating from zero. This makes it possible to use unnormalized model scores as approximate probabilities. Empirical evidence suggests that self-normalization is extremely effective, but a theoretical understanding of why it should work, and how generally it can be applied, is largely lacking.We prove upper bounds on the loss in accuracy due to self-normalization, describe classes of input distributionsthat self-normalize easily, and construct explicit examples of high-variance input distributions. Our theoretical results make predictions about the difficulty of fitting self-normalized models to several classes of distributions, and we conclude with empirical validation of these predictions on both real and synthetic datasets.
Jacob Andreas, Maxim Rabinovich, Michael I. Jordan, Daniel Klein 0001
NIPS4
2014 Structured Learning for Taxonomy Induction with Belief Propagation
abstract
We present a structured learning approach to inducing hypernym taxonomies using a probabilistic graphical model formulation.Our model incorporates heterogeneous relational evidence about both hypernymy and siblinghood, captured by semantic features based on patterns and statistics from Web n-grams and Wikipedia abstracts.For efficient inference over taxonomy structures, we use loopy belief propagation along with a directed spanning tree algorithm for the core hypernymy factor.To train the system, we extract sub-structures of WordNet and discriminatively learn to reproduce them, using adaptive subgradient stochastic optimization.On the task of reproducing sub-hierarchies of WordNet, our approach achieves a 51% error reduction over a chance baseline, including a 15% error reduction due to the non-hypernym-factored sibling features.On a comparison setup, we find up to 29% relative error reduction over previous work on ancestor F1.
Mohit Bansal, David Burkett, Gerard de Melo, Daniel Klein 0001
ACL (1)4
2014 Sparser, Better, Faster GPU Parsing
abstract
Due to their origin in computer graphics, graphics processing units (GPUs) are highly optimized for dense problems, where the exact same operation is applied repeatedly to all data points.Natural language processing algorithms, on the other hand, are traditionally constructed in ways that exploit structural sparsity.Recently, Canny et al. (2013) presented an approach to GPU parsing that sacrifices traditional sparsity in exchange for raw computational power, obtaining a system that can compute Viterbi parses for a high-quality grammar at about 164 sentences per second on a mid-range GPU.In this work, we reintroduce sparsity to GPU parsing by adapting a coarse-to-fine pruning approach to the constraints of a GPU.The resulting system is capable of computing over 404 Viterbi parses per second-more than a 2x speedup-on the same hardware.Moreover, our approach allows us to efficiently implement less GPU-friendly minimum Bayes risk inference, improving throughput for this more accurate algorithm from only 32 sentences per second unpruned to over 190 sentences per second using pruning-nearly a 6x speedup.
David Hall 0006, Taylor Berg-Kirkpatrick, Daniel Klein 0001
ACL (1)3
2014 Less Grammar, More Features
abstract
We present a parser that relies primar-ily on extracting information directly from surface spans rather than on propagat-ing information through enriched gram-mar structure. For example, instead of cre-ating separate grammar symbols to mark the definiteness of an NP, our parser might instead capture the same information from the first word of the NP. Moving context out of the grammar and onto surface fea-tures can greatly simplify the structural component of the parser: because so many deep syntactic cues have surface reflexes, our system can still parse accurately with context-free backbones as minimal as X-bar grammars. Keeping the structural backbone simple and moving features to the surface also allows easy adaptation to new languages and even to new tasks. On the SPMRL 2013 multilingual con-stituency parsing shared task (Seddah et al., 2013), our system outperforms the top single parser system of Björkelund et al. (2013) on a range of languages. In addi-tion, despite being designed for syntactic analysis, our system also achieves state-of-the-art numbers on the structural senti-ment task of Socher et al. (2013). Finally, we show that, in both syntactic parsing and sentiment analysis, many broad linguistic trends can be captured via surface features. 1
David Hall 0006, Greg Durrett, Daniel Klein 0001
ACL (1)3
2014 Grounding Language with Points and Paths in Continuous Spaces
abstract
We present a model for generating pathvalued interpretations of natural language text.Our model encodes a map from natural language descriptions to paths, mediated by segmentation variables which break the language into a discrete set of events, and alignment variables which reorder those events.Within an event, lexical weights capture the contribution of each word to the aligned path segment.We demonstrate the applicability of our model on three diverse tasks: a new color description task, a new financial news task and an established direction-following task.On all three, the model outperforms strong baselines, and on a hard variant of the direction-following task it achieves results close to the state-of-the-art system described in Vogel and Jurafsky (2010).
Jacob Andreas, Daniel Klein 0001
CoNLL2
2014 Unsupervised Transcription of Piano Music
Taylor Berg-Kirkpatrick, Jacob Andreas, Daniel Klein 0001
NIPS3
2014 A Joint Model for Entity Analysis: Coreference, Typing, and Linking
abstract
We present a joint model of three core tasks in the entity analysis stack: coreference resolution (within-document clustering), named entity recognition (coarse semantic typing), and entity linking (matching to Wikipedia entities). Our model is formally a structured conditional random field. Unary factors encode local features from strong baselines for each task. We then add binary and ternary factors to capture cross-task interactions, such as the constraint that coreferent mentions have the same semantic type. On the ACE 2005 and OntoNotes datasets, we achieve state-of-the-art results for all three tasks. Moreover, joint modeling improves performance on each task over strong independent baselines.
Greg Durrett, Daniel Klein 0001
Trans. Assoc. Comput. Linguistics2
2013 Unsupervised Transcription of Historical Documents
Taylor Berg-Kirkpatrick, Greg Durrett, Daniel Klein 0001
ACL (1)3
2013 Decentralized Entity-Level Modeling for Coreference Resolution
Greg Durrett, David Hall 0006, Daniel Klein 0001
ACL (1)3
2013 Decipherment with a Million Random Restarts
abstract
This paper investigates the utility and effect of running numerous random restarts when using EM to attack decipherment problems.We find that simple decipherment models are able to crack homophonic substitution ciphers with high accuracy if a large number of random restarts are used but almost completely fail with only a few random restarts.For particularly difficult homophonic ciphers, we find that big gains in accuracy are to be had by running upwards of 100K random restarts, which we accomplish efficiently using a GPU-based parallel implementation.We run a series of experiments using millions of random restarts in order to investigate other empirical properties of decipherment problems, including the famously uncracked Zodiac 340.
Taylor Berg-Kirkpatrick, Daniel Klein 0001
EMNLP2
2013 A Multi-Teraflop Constituency Parser using GPUs
abstract
Constituency parsing with rich grammars remains a computational challenge.Graphics Processing Units (GPUs) have previously been used to accelerate CKY chart evaluation, but gains over CPU parsers were modest.In this paper, we describe a collection of new techniques that enable chart evaluation at close to the GPU's practical maximum speed (a Teraflop), or around a half-trillion rule evaluations per second.Net parser performance on a 4-GPU system is over 1 thousand length-30 sentences/second (1 trillion rules/sec), and 400 general sentences/second for the Berkeley Parser Grammar.The techniques we introduce include grammar compilation, recursive symbol blocking, and cache-sharing.
John F. Canny, David Hall 0006, Daniel Klein 0001
EMNLP3
2013 Easy Victories and Uphill Battles in Coreference Resolution
abstract
Classical coreference systems encode various syntactic, discourse, and semantic phenomena explicitly, using heterogenous features computed from hand-crafted heuristics.In contrast, we present a state-of-the-art coreference system that captures such phenomena implicitly, with a small number of homogeneous feature templates examining shallow properties of mentions.Surprisingly, our features are actually more effective than the corresponding hand-engineered ones at modeling these key linguistic phenomena, allowing us to win "easy victories" without crafted heuristics.These features are successful on syntax and discourse; however, they do not model semantic compatibility well, nor do we see gains from experiments with shallow semantic features from the literature, suggesting that this approach to semantics is an "uphill battle."Nonetheless, our final system 1 outperforms the Stanford system (Lee et al. (2011), the winner of the CoNLL 2011 shared task) by 3.5% absolute on the CoNLL metric and outperforms the IMS system (Björkelund and Farkas (2012), the best publicly available English coreference system) by 1.9% absolute.
Greg Durrett, Daniel Klein 0001
EMNLP2
2013 Error-Driven Analysis of Challenges in Coreference Resolution
abstract
Coreference resolution metrics quantify errors but do not analyze them.Here, we consider an automated method of categorizing errors in the output of a coreference system into intuitive underlying error types.Using this tool, we first compare the error distributions across a large set of systems, then analyze common errors across the top ten systems, empirically characterizing the major unsolved challenges of the coreference resolution task.
Jonathan K. Kummerfeld, Daniel Klein 0001
EMNLP2
2013 Grounding spatial relations for human-robot interaction
abstract
We propose a system for human-robot interaction that learns both models for spatial prepositions and for object recognition. Our system grounds the meaning of an input sentence in terms of visual percepts coming from the robot's sensors in order to send an appropriate command to the PR2 or respond to spatial queries. To perform this grounding, the system recognizes the objects in the scene, determines which spatial relations hold between those objects, and semantically parses the input sentence. The proposed system uses the visual and spatial information in conjunction with the semantic parse to interpret statements that refer to objects (nouns), their spatial relationships (prepositions), and to execute commands (actions). The semantic parse is inherently compositional, allowing the robot to understand complex commands that refer to multiple objects and relations such as: “Move the cup close to the robot to the area in front of the plate and behind the tea box”. Our system correctly parses 94% of the 210 online test sentences, correctly interprets 91% of the correctly parsed sentences, and correctly executes 89% of the correctly interpreted sentences.
Sergio Guadarrama, Lorenzo Riano, Dave Golland, Daniel Göhring, Yangqing Jia, Daniel Klein 0001, Pieter Abbeel, Trevor Darrell
IROS6
2013 Learning Dependency-Based Compositional Semantics
abstract
Suppose we want to build a system that answers a natural language question by representing its semantics as a logical forxm and computing the answer given a structured database of facts. The core part of such a system is the semantic parser that maps questions to logical forms. Semantic parsers are typically trained from examples of questions annotated with their target logical forms, but this type of annotation is expensive. Our goal is to instead learn a semantic parser from question–answer pairs, where the logical form is modeled as a latent variable. We develop a new semantic formalism, dependency-based compositional semantics (DCS) and define a log-linear distribution over DCS logical forms. The model parameters are estimated using a simple procedure that alternates between beam search and numerical optimization. On two standard semantic parsing benchmarks, we show that our system obtains comparable accuracies to even state-of-the-art systems that do require annotated logical forms.
Percy Liang, Michael I. Jordan, Daniel Klein 0001
Comput. Linguistics3
2012 Coreference Semantics from Web Features
Mohit Bansal, Daniel Klein 0001
ACL (1)2
2012 Large-Scale Syntactic Language Modeling with Treelets
Adam Pauls, Daniel Klein 0001
ACL (1)2
2012 An Empirical Investigation of Statistical Significance in NLP
Taylor Berg-Kirkpatrick, David Burkett, Daniel Klein 0001
EMNLP-CoNLL3
2012 Transforming Trees to Improve Syntactic Convergence
David Burkett, Daniel Klein 0001
EMNLP-CoNLL2
2012 Syntactic Transfer Using a Bilingual Lexicon
Greg Durrett, Adam Pauls, Daniel Klein 0001
EMNLP-CoNLL3
2012 Training Factored PCFGs with Expectation Propagation
David Hall 0006, Daniel Klein 0001
EMNLP-CoNLL2
2012 Parser Showdown at the Wall Street Corral: An Empirical Investigation of Error Types in Parser Output
Jonathan K. Kummerfeld, David Hall 0006, James R. Curran, Daniel Klein 0001
EMNLP-CoNLL4
2012 Fast Inference in Phrase Extraction Models with Belief Propagation
David Burkett, Daniel Klein 0001
HLT-NAACL2
2012 Variational Inference for Structured NLP Models
David Burkett, Daniel Klein 0001
HLT-NAACL2
2011 Optimal Graph Search with Iterated Graph Cuts
abstract
Informed search algorithms such as A* use heuristics to focus exploration on states with low total path cost. To the extent that heuristics underestimate forward costs, a wider cost radius of suboptimal states will be explored. For many weighted graphs, however, a small distance in terms of cost may encompass a large fraction of the unweighted graph. We present a new informed search algorithm, Iterative Monotonically Bounded A* (IMBA*), which first proves that no optimal paths exist in a bounded cut of the graph before considering larger cuts. We prove that IMBA* has the same optimality and completeness guarantees as A* and, in a non-uniform pathfinding application, we empirically demonstrate substantial speed improvements over classic A*.
David Burkett, David Hall 0006, Daniel Klein 0001
AAAI3
2011 Web-Scale Features for Full-Scale Parsing
Mohit Bansal, Daniel Klein 0001
ACL2
2011 Jointly Learning to Extract and Compress
Taylor Berg-Kirkpatrick, Daniel Gillick, Daniel Klein 0001
ACL3
2011 Learning Dependency-Based Compositional Semantics
Percy Liang, Michael I. Jordan, Daniel Klein 0001
ACL3
2011 Faster and Smaller N-Gram Language Models
Adam Pauls, Daniel Klein 0001
ACL2
2011 Simple Effective Decipherment via Combinatorial Optimization
Taylor Berg-Kirkpatrick, Daniel Klein 0001
EMNLP2
2011 Large-Scale Cognate Recovery
David Hall 0006, Daniel Klein 0001
EMNLP2
2010 Simple, Accurate Parsing with an All-Fragments Grammar
Mohit Bansal, Daniel Klein 0001
ACL2
2010 Phylogenetic Grammar Induction
Taylor Berg-Kirkpatrick, Daniel Klein 0001
ACL2
2010 Discriminative Modeling of Extraction Sets for Machine Translation
John DeNero, Daniel Klein 0001
ACL2
2010 Finding Cognate Groups Using Phylogenies
David Hall 0006, Daniel Klein 0001
ACL2
2010 Learning Better Monolingual Models with Unannotated Bilingual Text
David Burkett, Slav Petrov, John Blitzer, Daniel Klein 0001
CoNLL4
2010 A Simple Domain-Independent Probabilistic Approach to Generation
Gabor Angeli, Percy Liang, Daniel Klein 0001
EMNLP3
2010 A Game-Theoretic Approach to Generating Spatial Descriptions
Dave Golland, Percy Liang, Daniel Klein 0001
EMNLP3
2010 Learning Programs: A Hierarchical Bayesian Approach
Percy Liang, Michael I. Jordan, Daniel Klein 0001
ICML3
2010 Painless Unsupervised Learning with Features
Taylor Berg-Kirkpatrick, Alexandre Bouchard-Côté, John DeNero, Daniel Klein 0001
HLT-NAACL4
2010 Joint Parsing and Alignment with Weakly Synchronized Grammars
David Burkett, John Blitzer, Daniel Klein 0001
HLT-NAACL3
2010 Coreference Resolution in a Modular, Entity-Centered Model
Aria Haghighi, Daniel Klein 0001
HLT-NAACL2
2010 Type-Based MCMC
Percy Liang, Michael I. Jordan, Daniel Klein 0001
HLT-NAACL3
2010 Unsupervised Syntactic Alignment with Inversion Transduction Grammars
Adam Pauls, Daniel Klein 0001, David Chiang 0001, Kevin Knight
HLT-NAACL2
2009 Better Word Alignments with Supervised ITG Models
Aria Haghighi, John Blitzer, John DeNero, Daniel Klein 0001
ACL/IJCNLP4
2009 Learning Semantic Correspondences with Less Supervision
Percy Liang, Michael I. Jordan, Daniel Klein 0001
ACL/IJCNLP3
2009 K-Best A* Parsing
Adam Pauls, Daniel Klein 0001
ACL/IJCNLP2
2009 Simple Coreference Resolution with Rich Syntactic and Semantic Features
Aria Haghighi, Daniel Klein 0001
EMNLP2
2009 Consensus Training for Consensus Decoding in Machine Translation
Adam Pauls, John DeNero, Daniel Klein 0001
EMNLP3
2009 Learning from measurements in exponential families
abstract
Given a model family and a set of unlabeled examples, one could either label specific examples or state general constraints---both provide information about the desired model. In general, what is the most cost-effective way to learn? To address this question, we introduce measurements, a general class of mechanisms for providing information about a target model. We present a Bayesian decision-theoretic framework, which allows us to both integrate diverse measurements and choose new measurements to make. We use a variational inference algorithm, which exploits exponential family duality. The merits of our approach are demonstrated on two sequence labeling tasks.
Percy Liang, Michael I. Jordan, Daniel Klein 0001
ICML3
2009 Improved Reconstruction of Protolanguage Word Forms
Alexandre Bouchard-Côté, Thomas L. Griffiths 0001, Daniel Klein 0001
HLT-NAACL3
2009 Efficient Parsing for Transducer Grammars
John DeNero, Mohit Bansal, Adam Pauls, Daniel Klein 0001
HLT-NAACL4
2009 Online EM for Unsupervised Models
Percy Liang, Daniel Klein 0001
HLT-NAACL2
2009 Hierarchical Search for Parsing
Adam Pauls, Daniel Klein 0001
HLT-NAACL2
2009 Randomized Pruning: Efficiently Calculating Expectations in Large Dynamic Programs
abstract
Pruning can massively accelerate the computation of feature expectations in large models. However, any single pruning mask will introduce bias. We present a novel approach which employs a randomized sequence of pruning masks. Formally, we apply auxiliary variable MCMC sampling to generate this sequence of masks, thereby gaining theoretical guarantees about convergence. Because each mask is generally able to skip large portions of an underlying dynamic program, our approach is particularly compelling for high-degree algorithms. Empirically, we demonstrate our method on bilingual parsing, showing decreasing bias as more masks are incorporated, and outperforming fixed tic-tac-toe pruning.
Alexandre Bouchard-Côté, Slav Petrov, Daniel Klein 0001
NIPS3
2008 Learning Bilingual Lexicons from Monolingual Corpora
Aria Haghighi, Percy Liang, Taylor Berg-Kirkpatrick, Daniel Klein 0001
ACL4
2008 Analyzing the Errors of Unsupervised Learning
Percy Liang, Daniel Klein 0001
ACL2
2008 Unsupervised Learning for Natural Language Processing
Daniel Klein 0001
COLT1
2008 Two Languages are Better than One (for Syntactic Parsing)
David Burkett, Daniel Klein 0001
EMNLP2
2008 Sampling Alignment Structure under a Bayesian Translation Model
John DeNero, Alexandre Bouchard-Côté, Daniel Klein 0001
EMNLP3
2008 Coarse-to-Fine Syntactic Machine Translation using Language Projections
Slav Petrov, Aria Haghighi, Daniel Klein 0001
EMNLP3
2008 Sparse Multi-Scale Grammars for Discriminative Latent Variable Parsing
Slav Petrov, Daniel Klein 0001
EMNLP2
2008 Structure compilation: trading structure for features
abstract
Structured models often achieve excellent performance but can be slow at test time. We investigate structure compilation, where we replace structure with features, which are often computationally simpler but unfortunately statistically more complex. We analyze this tradeoff theoretically and empirically on three natural language processing tasks. We also introduce a simple method to transfer predictive power from structure to features via unlabeled data, while incurring a minimal statistical penalty.
Percy Liang, Hal Daumé III, Daniel Klein 0001
ICML3
2008 Fully distributed EM for very large datasets
abstract
In EM and related algorithms, E-step computations distribute easily, because data items are independent given parameters. For very large data sets, however, even storing all of the parameters in a single node for the M-step can be impractical. We present a framework that fully distributes the entire EM procedure. Each node interacts only with parameters relevant to its data, sending messages to other nodes along a junction-tree topology. We demonstrate improvements over a MapReduce topology, on two tasks: word alignment and topic modeling.
Jason Andrew Wolfe, Aria Haghighi, Daniel Klein 0001
ICML3
2008 Efficient Inference in Phylogenetic InDel Trees
abstract
Accurate and efficient inference in evolutionary trees is a central problem in computational biology. Realistic models require tracking insertions and deletions along the phylogenetic tree, making inference challenging. We propose new sampling techniques that speed up inference and improve the quality of the samples. We compare our method to previous approaches and show performance improvement on metrics evaluating multiple sequence alignment and reconstruction of ancestral sequences.
Alexandre Bouchard-Côté, Michael I. Jordan, Daniel Klein 0001
NIPS3
2008 Efficient sentence segmentation using syntactic features
abstract
To enable downstream language processing,automatic speech recognition output must be segmented into its individual sentences. Previous sentence segmentation systems have typically been very local,using low-level prosodic and lexical features to independently decide whether or not to segment at each word boundary position. In this work,we leverage global syntactic information from a syntactic parser, which is better able to capture long distance dependencies. While some previous work has included syntactic features, ours is the first to do so in a tractable, lattice-based way, which is crucial for scaling up to long-sentence contexts. Specifically, an initial hypothesis lattice is constructed using local features. Candidate sentences are then assigned syntactic language model scores. These global syntactic scores are combined with local low-level scores in a log-linear model. The resulting system significantly outperforms the most popular long-span model for sentence segmentation (the hidden event language model) on both reference text and automatic speech recognizer output from news broadcasts.
Benoît Favre, Dilek Hakkani-Tür, Slav Petrov, Daniel Klein 0001
SLT4
2007 A* Search via Approximate Factoring
Aria Haghighi, John DeNero, Daniel Klein 0001
AAAI3
2007 Learning and Inference for Hierarchically Split PCFGs
Slav Petrov, Daniel Klein 0001
AAAI2
2007 Tailoring Word Alignments to Syntactic Machine Translation
John DeNero, Daniel Klein 0001
ACL2
2007 Unsupervised Coreference Resolution in a Nonparametric Bayesian Model
Aria Haghighi, Daniel Klein 0001
ACL2
2007 A Probabilistic Approach to Diachronic Phonology
Alexandre Bouchard-Côté, Percy Liang, Thomas L. Griffiths 0001, Daniel Klein 0001
EMNLP-CoNLL4
2007 The Infinite PCFG Using Hierarchical Dirichlet Processes
Percy Liang, Slav Petrov, Michael I. Jordan, Daniel Klein 0001
EMNLP-CoNLL4
2007 Learning Structured Models for Phone Recognition
Slav Petrov, Adam Pauls, Daniel Klein 0001
EMNLP-CoNLL3
2007 Approximate Factoring for A* Search
Aria Haghighi, John DeNero, Daniel Klein 0001
HLT-NAACL3
2007 Improved Inference for Unlexicalized Parsing
Slav Petrov, Daniel Klein 0001
HLT-NAACL2
2007 A Probabilistic Approach to Language Change
abstract
We present a probabilistic approach to language change in which word forms are represented by phoneme sequences that undergo stochastic edits along the branches of a phylogenetic tree. Our framework combines the advantages of the classical comparative method with the robustness of corpus-based probabilistic models. We use this framework to explore the consequences of two different schemes for defining probabilistic models of phonological change, evaluating these schemes using the reconstruction of ancient word forms in Romance languages. The result is an efficient inference procedure for automatically inferring ancient word forms from modern languages, which can be generalized to support inferences about linguistic phylogenies.
Alexandre Bouchard-Côté, Percy Liang, Thomas L. Griffiths 0001, Daniel Klein 0001
NIPS4
2007 Agreement-Based Learning
abstract
The learning of probabilistic models with many hidden variables and non- decomposable dependencies is an important and challenging problem. In contrast to traditional approaches based on approximate inference in a single intractable model, our approach is to train a set of tractable submodels by encouraging them to agree on the hidden variables. This allows us to capture non-decomposable aspects of the data while still maintaining tractability. We propose an objective function for our approach, derive EM-style algorithms for parameter estimation, and demonstrate their effectiveness on three challenging real-world learning tasks.
Percy Liang, Daniel Klein 0001, Michael I. Jordan
NIPS2
2007 Discriminative Log-Linear Grammars with Latent Variables
abstract
We demonstrate that log-linear grammars with latent variables can be practically trained using discriminative methods. Central to efficient discriminative training is a hierarchical pruning procedure which allows feature expectations to be effi- ciently approximated in a gradient-based procedure. We compare L1 and L2 reg- ularization and show that L1 regularization is superior, requiring fewer iterations to converge, and yielding sparser solutions. On full-scale treebank parsing exper- iments, the discriminative latent models outperform both the comparable genera- tive latent models as well as the discriminative non-latent baselines.
Slav Petrov, Daniel Klein 0001
NIPS2
2007 Mixture-of-Parents Maximum Entropy Markov Models
David S. Rosenberg, Daniel Klein 0001, Ben Taskar
UAI2
2006 Prototype-Driven Grammar Induction
abstract
We investigate prototype-driven learning for primarily unsupervised grammar induction. Prior knowledge is specified declaratively, by providing a few canonical examples of each target phrase type. This sparse prototype information is then propagated across a corpus using distributional similarity features, which augment an otherwise standard PCFG model. We show that distributional features are effective at distinguishing bracket labels, but not determining bracket locations. To improve the quality of the induced trees, we combine our PCFG induction with the CCM model of Klein and Manning (2002), which has complementary stengths: it identifies brackets but does not label them. Using only a handful of prototypes, we show substantial improvements over naive PCFG induction for English and Chinese grammar induction.
Aria Haghighi, Daniel Klein 0001
ACL2
2006 An End-to-End Discriminative Approach to Machine Translation
abstract
We present a perceptron-style discriminative approach to machine translation in which large feature sets can be exploited. Unlike discriminative reranking approaches, our system can take advantage of learned features in all stages of decoding. We first discuss several challenges to error-driven discriminative approaches. In particular, we explore different ways of updating parameters given a training example. We find that making frequent but smaller updates is preferable to making fewer but larger updates. Then, we discuss an array of features and show both how they quantitatively increase BLEU score and how they qualitatively interact on specific examples. One particular feature we investigate is a novel way to introduce learning into the initial phrase extraction process, which has previously been entirely heuristic.
Percy Liang, Alexandre Bouchard-Côté, Daniel Klein 0001, Ben Taskar
ACL3
2006 Learning Accurate, Compact, and Interpretable Tree Annotation
abstract
We present an automatic approach to tree annotation in which basic nonterminal symbols are alternately split and merged to maximize the likelihood of a training treebank. Starting with a simple X-bar grammar, we learn a new grammar whose nonterminals are subsymbols of the original nonterminals. In contrast with previous work, we are able to split various terminals to different degrees, as appropriate to the actual complexity in the data. Our grammars automatically learn the kinds of linguistic distinctions exhibited in previous work on manual tree annotation. On the other hand, our grammars are much more compact and substantially more accurate than previous work on automatic annotation. Despite its simplicity, our best grammar achieves an F1 of 90.2% on the Penn Treebank, higher than fully lexicalized systems.
Slav Petrov, Leon Barrett, Romain Thibaux, Daniel Klein 0001
ACL4
2006 Non-Local Modeling with a Mixture of PCFGs
Slav Petrov, Leon Barrett, Daniel Klein 0001
CoNLL3
2006 Prototype-Driven Learning for Sequence Models
Aria Haghighi, Daniel Klein 0001
HLT-NAACL2
2006 Word Alignment via Quadratic Assignment
Simon Lacoste-Julien, Ben Taskar, Daniel Klein 0001, Michael I. Jordan
HLT-NAACL3
2006 Alignment by Agreement
Percy Liang, Ben Taskar, Daniel Klein 0001
HLT-NAACL3
2005 Unsupervised Learning of Field Segmentation Models for Information Extraction
abstract
The applicability of many current information extraction techniques is severely limited by the need for supervised training data. We demonstrate that for certain field structured extraction tasks, such as classified advertisements and bibliographic citations, small amounts of prior knowledge can be used to learn effective models in a primarily unsupervised fashion. Although hidden Markov models (HMMs) provide a suitable generative model for field structured text, general unsupervised HMM learning fails to learn useful structure in either of our domains. However, one can dramatically improve the quality of the learned structure by exploiting simple prior knowledge of the desired solutions. In both domains, we found that unsupervised methods can attain accuracies with 400 unlabeled examples comparable to those attained by supervised methods on 50 labeled examples, and that semi-supervised methods can make good use of small amounts of labeled data.
Trond Grenager, Daniel Klein 0001, Christopher D. Manning
ACL2
2005 Natural language grammar induction with a generative constituent-context model
Daniel Klein 0001, Christopher D. Manning
Pattern Recognit.1
2004 Corpus-Based Induction of Syntactic Structure: Models of Dependency and Constituency
abstract
We present a generative model for the unsupervised learning of dependency structures. We also describe the multiplicative combination of this dependency model with a model of linear constituency. The product model outperforms both components on their respective evaluation metrics, giving the best published figures for unsupervised dependency parsing and unsupervised constituency parsing. We also demonstrate that the combined model works and is robust cross-linguistically, being able to exploit either attachment or distributional regularities that are salient in the data.
Daniel Klein 0001, Christopher D. Manning
ACL1
2004 Max-Margin Parsing
Ben Taskar, Daniel Klein 0001, Michael Collins 0001, Daphne Koller, Christopher D. Manning
EMNLP2
2004 Data-Oriented Parsing edited by Rens Bod, Remko Scha, and Khalil Sima'an
abstract
Data-Oriented Parsing contains four parts, each of which will interest a different set of readers.The early sections give a good introduction to the data-oriented parsing (DOP) framework, while later sections present more recent work, including a substantial amount of work on lexicalized tree-adjoining grammars (LTAGs) and some work on structural models of translation.
Daniel Klein 0001
Comput. Linguistics1
2003 Accurate Unlexicalized Parsing
abstract
We demonstrate that an unlexicalized PCFG can parse much more accurately than previously shown, by making use of simple, linguistically motivated state splits, which break down false independence assumptions latent in a vanilla treebank grammar. Indeed, its performance of 86.36% (LP/LR F1) is better than that of early lexicalized PCFG models, and surprisingly close to the current state-of-the-art. This result has potential uses beyond establishing a strong lower bound on the maximum possible accuracy of unlexicalized models: an unlexicalized PCFG is much more compact, easier to replicate, and easier to interpret than more complex lexical models, and the parsing algorithms are simpler, more widely understood, of lower asymptotic complexity, and easier to optimize.
Daniel Klein 0001, Christopher D. Manning
ACL1
2003 Named Entity Recognition with Character-Level Models
Daniel Klein 0001, Joseph Smarr, Christopher D. Manning
CoNLL1
2003 Spectral Learning
Sepandar D. Kamvar, Daniel Klein 0001, Christopher D. Manning
IJCAI2
2003 Factored A* Search for Models over Sequences and Trees
Daniel Klein 0001, Christopher D. Manning
IJCAI1
2003 A* Parsing: Fast Exact Viterbi Parse Selection
Daniel Klein 0001, Christopher D. Manning
HLT-NAACL1
2003 Optimization, Maxent Models, and Conditional Estimation without Magic
Christopher D. Manning, Daniel Klein 0001
HLT-NAACL2
2003 Feature-Rich Part-of-Speech Tagging with a Cyclic Dependency Network
Kristina Toutanova, Daniel Klein 0001, Christopher D. Manning, Yoram Singer
HLT-NAACL2
2002 A Generative Constituent-Context Model for Improved Grammar Induction
abstract
We present a generative distributional model for the unsupervised induction of natural language syntax which explicitly models constituent yields and contexts. Parameter search with EM produces higher quality analyses than previously exhibited by unsupervised systems, giving the best published un-supervised parsing results on the ATIS corpus. Experiments on Penn treebank sentences of comparable length show an even higher F1 of 71% on non-trivial brackets. We compare distributionally induced and actual part-of-speech tags as input data, and examine extensions to the basic model. We discuss errors made by the system, compare the system to previous models, and discuss upper bounds, lower bounds, and stability for this task.
Daniel Klein 0001, Christopher D. Manning
ACL1
2002 Conditional Structure versus Conditional Estimation in NLP Models
abstract
This paper separates conditional parameter estimation, which consistently raises test set accuracy on statistical NLP tasks, from conditional model structures, such as the conditional Markov model used for maximum-entropy tagging, which tend to lower accuracy. Error analysis on part-of-speech tagging shows that the actual tagging errors made by the conditionally structured model derive not only from label bias, but also from other ways in which the independence assumptions of the conditional model structure are unsuited to linguistic sequences. The paper presents new word-sense disambiguation and POS tagging experiments, and integrates apparently conflicting reports from other recent work.
Daniel Klein 0001, Christopher D. Manning
EMNLP1
2002 Interpreting and Extending Classical Agglomerative Clustering Algorithms using a Model-Based approach
Sepandar D. Kamvar, Daniel Klein 0001, Christopher D. Manning
ICML2
2002 From Instance-level Constraints to Space-Level Constraints: Making the Most of Prior Knowledge in Data Clustering
Daniel Klein 0001, Sepandar D. Kamvar, Christopher D. Manning
ICML1
2002 Fast Exact Inference with a Factored Model for Natural Language Parsing
abstract
We present a novel generative model for natural language tree structures in which semantic (lexical dependency) and syntactic (PCFG) structures are scored with separate models. This factorization provides concep- tual simplicity, straightforward opportunities for separately improving the component models, and a level of performance comparable to simi- lar, non-factored models. Most importantly, unlike other modern parsing models, the factored model admits an extremely effective A* parsing al- gorithm, which enables efficient, exact inference.
Daniel Klein 0001, Christopher D. Manning
NIPS1
2002 Evaluating strategies for similarity search on the web
abstract
Finding pages on the Web that are similar to a query page (Related Pages) is an important component of modern search engines. A variety of strategies have been proposed for answering Related Pages queries, but comparative evaluation by user studies is expensive, especially when large strategy spaces must be searched (e.g., when tuning parameters). We present a technique for automatically evaluating strategies using Web hierarchies, such as Open Directory, in place of user feedback. We apply this evaluation methodology to a mix of document representation strategies, including the use of text, anchor-text, and links. We discuss the relative advantages and disadvantages of the various approaches examined. Finally, we describe how to e#ciently construct a similarity index out of our chosen strategies, and provide sample results from our index.
Taher H. Haveliwala, Aristides Gionis, Daniel Klein 0001, Piotr Indyk
WWW3
2001 Parsing with Treebank Grammars: Empirical Bounds, Theoretical Models, and the Structure of the Penn Treebank
abstract
This paper presents empirical studies and closely corresponding theoretical models of the performance of a chart parser exhaustively parsing the Penn Treebank with the Treebank's own CFG grammar.We show how performance is dramatically affected by rule representation and tree transformations, but little by top-down vs. bottom-up strategies.We discuss grammatical saturation, including analysis of the strongly connected components of the phrasal nonterminals in the Treebank, and model how, as sentence length increases, the effective grammar rule size increases as regions of the grammar are unlocked, yielding super-cubic observed time behavior in some configurations.
Daniel Klein 0001, Christopher D. Manning
ACL1
2001 Natural Language Grammar Induction Using a Constituent-Context Model
abstract
This paper presents a novel approach to the unsupervised learning of syn- tactic analyses of natural language text. Most previous work has focused on maximizing likelihood according to generative PCFG models. In con- trast, we employ a simpler probabilistic model over trees based directly on constituent identity and linear context, and use an EM-like iterative procedure to induce structure. This method produces much higher qual- ity analyses, giving the best published results on the ATIS dataset. 1 Overview To enable a wide range of subsequent tasks, human language sentences are standardly given tree-structure analyses, wherein the nodes in a tree dominate contiguous spans of words called constituents, as in figure 1(a). Constituents are the linguistically coherent units in the sentence, and are usually labeled with a constituent category, such as noun phrase (NP) or verb phrase (VP). An aim of grammar induction systems is to figure out, given just the sentences in a corpus S, what tree structures correspond to them. In this sense, the grammar induction problem is an incomplete data problem, where the complete data is the corpus of trees T , but we only observe their yields S. This paper presents a new approach to this problem, which gains leverage by directly making use of constituent contexts. It is an open problem whether entirely unsupervised methods can produce linguistically accurate parses of sentences. Due to the difficulty of this task, the vast majority of statis- tical parsing work has focused on supervised learning approaches to parsing, where one uses a treebank of fully parsed sentences to induce a model which parses unseen sentences [7, 3]. But there are compelling motivations for unsupervised grammar induction. Building supervised training data requires considerable resources, including time and linguistic ex- pertise. Investigating unsupervised methods can shed light on linguistic phenomena which are implicit within a supervised parser's supervisory information (e.g., unsupervised sys- tems often have difficulty correctly attaching subjects to verbs above objects, whereas for a supervised parser, this ordering is implicit in the supervisory information). Finally, while the presented system makes no claims to modeling human language acquisition, results on whether there is enough information in sentences to recover their structure are important data for linguistic theory, where it has standardly been assumed that the information in the data is deficient, and strong innate knowledge is required for language acquisition [4].
Daniel Klein 0001, Christopher D. Manning
NIPS1