Christopher D. Manning

dblp:m/ChristopherDManning · DBLP profile ↗
← Back
250ranked-venue papers
9as first author
57since 2021 · last 2025
0000-0001-6155-649XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 231 · 8 first-author · 56 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models
Tolúlopé Ògúnrèmí, Christopher D. Manning, Daniel Jurafsky, Karen Livescu
ASRU2
2025 Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic Evaluations
abstract
Large Language Models (LLMs) exhibit a robust mastery of syntax when processing and generating text.While this suggests internalized understanding of hierarchical syntax and dependency relations, the precise mechanism by which they represent syntactic structure is an open area within interpretability research.Probing provides one way to identify syntactic mechanisms linearly encoded in activations; however, no comprehensive study has yet established whether a model's probing accuracy reliably predicts its downstream syntactic performance.Adopting a "mechanisms vs. outcomes" framework, we evaluate 32 open-weight transformer models and find that syntactic features extracted via probing fail to predict outcomes of targeted syntax evaluations across English linguistic phenomena.Our results highlight a substantial disconnect between latent syntactic representations found via probing and observable syntactic behaviors in downstream tasks.
Ananth Agarwal, Jasper Jian, Christopher D. Manning, Shikhar Murty
EMNLP3
2025 Stronger Baselines for Retrieval-Augmented Generation with Long-Context Language Models
abstract
With the rise of long-context language models (LMs) capable of processing tens of thousands of tokens in a single context window, do multi-stage retrieval-augmented generation (RAG) pipelines still offer measurable benefits over simpler, single-stage approaches?To assess this question, we conduct a controlled evaluation for QA tasks under systematically scaled token budgets, comparing two recent multi-stage pipelines, ReadAgent and RAPTOR, against three baselines, including DOS RAG (Document's Original Structure RAG), a simple retrieve-then-read method that preserves original passage order.Despite its straightforward design, DOS RAG consistently matches or outperforms more intricate methods on multiple long-context QA benchmarks.We trace this strength to a combination of maintaining source fidelity and document structure, prioritizing recall within effective context windows, and favoring simplicity over added pipeline complexity.We recommend establishing DOS RAG as a simple yet strong baseline for future RAG evaluations, paired with stateof-the-art embedding and language models, and benchmarked under matched token budgets, to ensure that added pipeline complexity is justified by clear performance gains as models continue to improve. 1
Alex Laitenberger, Christopher D. Manning, Nelson F. Liu
EMNLP2
2025 AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
abstract
Video detailed captioning is a key task which aims to generate comprehensive and coherent textual descriptions of video content, benefiting both video understanding and generation. In this paper, we propose AuroraCap, a video captioner based on a large multimodal model. We follow the simplest architecture design without additional parameters for temporal modeling. To address the overhead caused by lengthy video sequences, we implement the token merging strategy, reducing the number of input visual tokens. Surprisingly, we found that this strategy results in little performance loss. AuroraCap shows superior performance on various video and image captioning benchmarks, for example, obtaining a CIDEr of 88.9 on Flickr30k, beating GPT-4V (55.3) and Gemini-1.5 Pro (82.2). However, existing video caption benchmarks only include simple descriptions, consisting of a few dozen words, which limits research in this field. Therefore, we develop VDC, a video detailed captioning benchmark with over one thousand carefully annotated structured captions. In addition, we propose a new LLM-assisted metric VDCscore for bettering evaluation, which adopts a divide-and-conquer strategy to transform long caption evaluation into multiple short question-answer pairs. With the help of human Elo ranking, our experiments show that this benchmark better correlates with human judgments of video detailed captioning quality.
Wenhao Chai, Enxin Song, Yilun Du, Chenlin Meng, Vashisht Madhavan, Omer Bar-Tal, Jenq-Neng Hwang, Saining Xie, Christopher D. Manning
ICLR9
2025 h4rm3l: A Language for Composable Jailbreak Attack Synthesis
abstract
Despite their demonstrated valuable capabilities, state-of-the-art (SOTA) widely deployed large language models (LLMs) still have the potential to cause harm to society due to the ineffectiveness of their safety filters, which can be bypassed by prompt transformations called jailbreak attacks. Current approaches to LLM safety assessment, which employ datasets of templated prompts and benchmarking pipelines, fail to cover sufficiently large and diverse sets of jailbreak attacks, leading to the widespread deployment of unsafe LLMs. Recent research showed that novel jailbreak attacks could be derived by composition; however, a formal composable representation for jailbreak attacks, which, among other benefits, could enable the exploration of a large compositional space of jailbreak attacks through program synthesis methods, has not been previously proposed. We introduce h4rm3l, a novel approach that addresses this gap with a human-readable domain-specific language (DSL). Our framework comprises: (1) The h4rm3l DSL, which formally expresses jailbreak attacks as compositions of parameterized string transformation primitives. (2) A synthesizer with bandit algorithms that efficiently generates jailbreak attacks optimized for a target black box LLM. (3) The h4rm3l red-teaming software toolkit that employs the previous two components and an automated harmful LLM behavior classifier that is strongly aligned with human judgment. We demonstrate h4rm3l's efficacy by synthesizing a dataset of 2656 successful novel jailbreak attacks targeting 6 SOTA open-source and proprietary LLMs (GPT-3.5, GPT-4o, Claude-3-Sonnet, Claude-3-Haiku, Llama-3-8B, and Llama-3-70B), and by benchmarking those models against a subset of these synthesized attacks. Our results show that h4rm3l's synthesized attacks are diverse and more successful than existing jailbreak attacks in literature, with success rates exceeding 90% on SOTA LLMs. Warning: This paper and related research artifacts contain offensive and potentially disturbing prompts and model-generated content.
Moussa Doumbouya, Ananjan Nandi, Gabriel Poesia, Davide Ghilardi, Anna Goldie, Federico Bianchi 0001, Daniel Jurafsky, Christopher D. Manning
ICLR8
2025 MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
abstract
Models that rely on subword tokenization have significant drawbacks, such as sensitivity to character-level noise like spelling errors and inconsistent compression rates across different languages and scripts. While character- or byte-level models like ByT5 attempt to address these concerns, they have not gained widespread adoption—processing raw byte streams without tokenization results in significantly longer sequence lengths, making training and inference inefficient. This work introduces MrT5 (MergeT5), a more efficient variant of ByT5 that integrates a token deletion mechanism in its encoder to dynamically shorten the input sequence length. After processing through a fixed number of encoder layers, a learned delete gate determines which tokens are to be removed and which are to be retained for subsequent layers. MrT5 effectively "merges" critical information from deleted tokens into a more compact sequence, leveraging contextual information from the remaining tokens. In continued pre-training experiments, we find that MrT5 can achieve significant gains in inference runtime with minimal effect on performance, as measured by bits-per-byte. Additionally, with multilingual training, MrT5 adapts to the orthographic characteristics of each language, learning language-specific compression rates. Furthermore, MrT5 shows comparable accuracy to ByT5 on downstream evaluations such as XNLI, TyDi QA, and character-level tasks while reducing sequence lengths by up to 75%. Our approach presents a solution to the practical limitations of existing byte-level models.
Julie Kallini, Shikhar Murty, Christopher D. Manning, Christopher Potts, Róbert Csordás
ICLR3
2025 AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
abstract
Fine-grained steering of language model outputs is essential for safety and reliability. Prompting and finetuning are widely used to achieve these goals, but interpretability researchers have proposed a variety of representation-based techniques as well, including sparse autoencoders (SAEs), linear artificial tomography, supervised steering vectors, linear probes, and representation finetuning. At present, there is no benchmark for making direct comparisons between these proposals. Therefore, we introduce AxBench, a large-scale benchmark for steering and concept detection, and report experiments on Gemma-2-2B and 9B. For steering, we find that prompting outperforms all existing methods, followed by finetuning. For concept detection, representation-based methods such as difference-in-means, perform the best. On both evaluations, SAEs are not competitive. We introduce a novel weakly-supervised representational method (Rank-1 Representation Finetuning; ReFT-r1), which is competitive on both tasks while providing the interpretability advantages that prompting lacks. Along with AxBench, we train and publicly release SAE-scale feature dictionaries for ReFT-r1 and DiffMean.
Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang 0078, Jing Huang 0014, Daniel Jurafsky, Christopher D. Manning, Christopher Potts
ICML7
2025 The Surprising Victory of NLP: From Philosophy to Agentic Language Models
abstract
Language Models have been around for decades but have suddenly taken the world by storm. In a surprising third act for anyone doing Natural Language Processing (NLP) in the 70s, 80s, 90s, or 2000s, in much of the popular media, artificial intelligence is now synonymous with language models. In this talk, I want to take a look backward at where language models came from and why they were so slow to emerge, a look inward to give a few thoughts on meaning, intelligence, and what language models understand and know, and a look forward at some topics of recent research: the possibilities for progress with new versions of Universal Transformers and using Large Language Models (LLMs) to build intelligent language-using agents for the digital world. I will argue that material beyond language is not necessary to having meaning and understanding, but it is very useful in most cases, and that composability, adaptability, and learning are vital to intelligence. Rather than simply continuing to build ever-larger LLMs from huge passive collections of text data, we need to explore developing better neural architectures and agents that can learn through interactions. I will discuss how web agents can learn through interactions and how such an interaction-first learning approach can work very effectively in web environments with relatively small language models.
Christopher D. Manning
KDD (2)1
2025 Sneaking Syntax into Transformer Language Models with Tree Regularization
abstract
Ananjan Nandi, Christopher D Manning, Shikhar Murty. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Ananjan Nandi, Christopher D. Manning, Shikhar Murty
NAACL (Long Papers)2
2025 Do Language Models Use Their Depth Efficiently?
abstract
Modern LLMs are increasingly deep, and depth correlates with performance, albeit with diminishing returns. However, do these models use their depth efficiently? Do they compose more features to create higher-order computations that are impossible in shallow models, or do they merely spread the same kinds of computation out over more layers? To address these questions, we analyze the residual stream of the Llama 3.1, Qwen 3, and OLMo 2 family of models. We find: First, comparing the output of the sublayers to the residual stream reveals that layers in the second half contribute much less than those in the first half, with a clear phase transition between the two halves. Second, skipping layers in the second half has a much smaller effect on future computations and output predictions. Third, for multihop tasks, we are unable to find evidence that models are using increased depth to compose subresults in examples involving many hops. Fourth, we seek to directly address whether deeper models are using their additional layers to perform new kinds of computation. To do this, we train linear maps from the residual stream of a shallow model to a deeper one. We find that layers with the same relative depth map best to each other, suggesting that the larger model simply spreads the same computations out over its many layers. All this evidence suggests that deeper models are not using their depth to learn new kinds of computation, but only using the greater depth to perform more fine-grained adjustments to the residual. This may help explain why increasing scale leads to diminishing returns for stacked Transformer architectures.
Róbert Csordás, Christopher D. Manning, Christopher Potts
NeurIPS2
2025 A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection
abstract
Companies spend large amounts of money on public relations campaigns to project a positive brand image.However, sometimes there is a mismatch between what they say and what they do. Oil & gas companies, for example, are accused of "greenwashing" with imagery of climate-friendly initiatives.Understanding the framing, and changes in framing, at scale can help better understand the goals and nature of public relation campaigns.To address this, we introduce a benchmark dataset of expert-annotated video ads obtained from Facebook and YouTube.The dataset provides annotations for 13 framing types for more than 50 companies or advocacy groups across 20 countries.Our dataset is especially designed for the evaluation of vision-language models (VLMs), distinguishing it from past text-only framing datasets.Baseline experiments show some promising results, while leaving room for improvement for future work: GPT-4.1 can detect environmental messages with 79% F1 score, while our best model only achieves 46% F1 score on identifying framing around green innovation.We also identify challenges that VLMs must address, such as implicit framing, handling videos of various lengths, or implicit cultural backgrounds.Our dataset contributes to research in multimodal analysis of strategic communication in the energy sector.
Gaku Morio, Harri Rowlands, Dominik Stammbach, Christopher D. Manning, Peter Henderson 0002
NeurIPS4
2025 Improved Representation Steering for Language Models
abstract
Steering methods for language models (LMs) seek to provide fine-grained and interpretable control over model generations by variously changing model inputs, weights, or representations to adjust behavior. Recent work has shown that adjusting weights or representations is often less effective than steering by prompting, for instance when wanting to introduce or suppress a particular concept. We demonstrate how to improve representation steering via our new Reference-free Preference Steering (RePS), a bidirectional preference-optimization objective that jointly does concept steering and suppression. We train three parameterizations of RePS and evaluate them on AxBench, a large-scale model steering benchmark. On Gemma models with sizes ranging from 2B to 27B, RePS outperforms all existing steering methods trained with a language modeling objective and substantially narrows the gap with prompting -- while promoting interpretability and minimizing parameter count. In suppression, RePS matches the language-modeling objective on Gemma-2 and outperforms it on the larger Gemma-3 variants while remaining resilient to prompt-based jailbreaking attacks that defeat prompting. Overall, our results suggest that RePS provides an interpretable and robust alternative to prompting for both steering and suppression.
Zhengxuan Wu, Qinan Yu, Aryaman Arora, Christopher D. Manning, Christopher Potts
NeurIPS4
2024 Statistical Uncertainty in Word Embeddings: GloVe-V
abstract
Static word embeddings are ubiquitous in computational social science applications and contribute to practical decision-making in a variety of fields including law and healthcare.However, assessing the statistical uncertainty in downstream conclusions drawn from word embedding statistics has remained challenging.When using only point estimates for embeddings, researchers have no streamlined way of assessing the degree to which their model selection criteria or scientific conclusions are subject to noise due to sparsity in the underlying data used to generate the embeddings.We introduce a method to obtain approximate, easy-to-use, and scalable reconstruction error variance estimates for GloVe (Pennington et al., 2014), one of the most widely used word embedding models, using an analytical approximation to a multivariate normal model.To demonstrate the value of embeddings with variance (GloVe-V), we illustrate how our approach enables principled hypothesis testing in core word embedding tasks, such as comparing the similarity between different word pairs in vector space, assessing the performance of different models, and analyzing the relative degree of ethnic or gender bias in a corpus using different word lists.
Andrea Vallebueno, Cassandra Handan-Nader, Christopher D. Manning, Daniel E. Ho
EMNLP3
2024 An Emulator for Fine-tuning Large Language Models using Small Language Models
abstract
Widely used language models (LMs) are typically built by scaling up a two-stage training pipeline: a pre-training stage that uses a very large, diverse dataset of text and a fine-tuning (sometimes, 'alignment') stage that uses targeted examples or other specifications of desired behaviors. While it has been hypothesized that knowledge and skills come from pre-training, and fine-tuning mostly filters this knowledge and skillset, this intuition has not been extensively tested. To aid in doing so, we introduce a novel technique for decoupling the knowledge and skills gained in these two stages, enabling a direct answer to the question, *What would happen if we combined the knowledge learned by a large model during pre-training with the knowledge learned by a small model during fine-tuning (or vice versa)?* Using an RL-based framework derived from recent developments in learning from human preferences, we introduce *emulated fine-tuning (EFT)*, a principled and practical method for sampling from a distribution that approximates (or 'emulates') the result of pre-training and fine-tuning at different scales. Our experiments with EFT show that scaling up fine-tuning tends to improve helpfulness, while scaling up pre-training tends to improve factuality. Beyond decoupling scale, we show that EFT enables test-time adjustment of competing behavioral traits like helpfulness and harmlessness without additional training. Finally, a special case of emulated fine-tuning, which we call LM *up-scaling*, avoids resource-intensive fine-tuning of large pre-trained models by ensembling them with small fine-tuned models, essentially emulating the result of fine-tuning the large pre-trained model. Up-scaling consistently improves helpfulness and factuality of instruction-following models in the Llama, Llama-2, and Falcon families, without additional hyperparameters or training. For reference implementation, see [https://github.com/eric-mitchell/emulated-fine-tuning](https://github.com/eric-mitchell/emulated-fine-tuning).
Eric Mitchell, Rafael Rafailov, Archit Sharma, Chelsea Finn, Christopher D. Manning
ICLR5
2024 Language Model Detectors Are Easily Optimized Against
abstract
The fluency and general applicability of large language models (LLMs) has motivated significant interest in detecting whether a piece of text was written by a language model. While both academic and commercial detectors have been deployed in some settings, particularly education, other research has highlighted the fragility of these systems. In this paper, we demonstrate a data-efficient attack that fine-tunes language models to confuse existing detectors, leveraging recent developments in reinforcement learning of language models. We use the `human-ness' score (often just a log probability) of various open-source and commercial detectors as a reward function for reinforcement learning, subject to a KL-divergence constraint that the resulting model does not differ significantly from the original. For a 7B parameter Llama-2 model, fine-tuning for under a day reduces the AUROC of the OpenAI RoBERTa-Large detector from 0.84 to 0.63, while perplexity on OpenWebText increases from 8.7 to only 9.0; with a larger perplexity budget, we can drive AUROC to 0.30 (worse than random). Similar to traditional adversarial attacks, we find that this increase in 'detector evasion' generalizes to other detectors not used during training. In light of our empirical results, we advise against continued reliance on LLM-generated text detectors. Models, datasets, and selected experiment code will be released at https://github.com/charlottttee/llm-detector-evasion.
Charlotte Nicks, Eric Mitchell, Rafael Rafailov, Archit Sharma, Christopher D. Manning, Chelsea Finn, Stefano Ermon
ICLR5
2024 RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval
abstract
Retrieval-augmented language models can better adapt to changes in world state and incorporate long-tail knowledge. However, most existing methods retrieve only short contiguous chunks from a retrieval corpus, limiting holistic understanding of the overall document context. We introduce the novel approach of recursively embedding, clustering, and summarizing chunks of text, constructing a tree with differing levels of summarization from the bottom up. At inference time, our RAPTOR model retrieves from this tree, integrating information across lengthy documents at different levels of abstraction. Controlled experiments show that retrieval with recursive summaries offers significant improvements over traditional retrieval-augmented LMs on several tasks. On question-answering tasks that involve complex, multi-step reasoning, we show state-of-the-art results; for example, by coupling RAPTOR retrieval with the use of GPT-4, we can improve the best performance on the QuALITY benchmark by 20\% in absolute accuracy.
Parth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna, Anna Goldie, Christopher D. Manning
ICLR6
2024 Fine-Tuning Language Models for Factuality
abstract
The fluency and creativity of large pre-trained language models (LLMs) have led to their widespread use, sometimes even as a replacement for traditional search engines. Yet language models are prone to making convincing but factually inaccurate claims, often referred to as `hallucinations.' These errors can inadvertently spread misinformation or harmfully perpetuate misconceptions. Further, manual fact-checking of model responses is a time-consuming process, making human factuality labels expensive to acquire. In this work, we fine-tune language models to be more factual, without human labeling and targeting more open-ended generation settings than past work. We leverage two key recent innovations in NLP to do so. First, several recent works have proposed methods for judging the factuality of open-ended text by measuring consistency with an external knowledge base or simply a large model's confidence scores. Second, the Direct Preference Optimization algorithm enables straightforward fine-tuning of language models on objectives other than supervised imitation, using a preference ranking over possible model responses. We show that learning from automatically generated factuality preference rankings, generated either through existing retrieval systems or our novel retrieval-free approach, significantly improves the factuality (percent of generated claims that are correct) of Llama-2 on held-out topics compared with RLHF or decoding strategies targeted at factuality. At 7B scale, compared to Llama-2-Chat, we observe 53% and 50% reduction in factual error rate when generating biographies and answering medical questions, respectively. A reference implementation can be found at https://github.com/kttian/llm_factuality_tuning.
Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D. Manning, Chelsea Finn
ICLR4
2024 BAGEL: Bootstrapping Agents by Guiding Exploration with Language
abstract
Following natural language instructions by executing actions in digital environments (e.g. web-browsers and REST APIs) is a challenging task for language model (LM) agents. Unfortunately, LM agents often fail to generalize to new environments without human demonstrations. This work presents BAGEL, a method for bootstrapping LM agents without human supervision. BAGEL converts a seed set of randomly explored trajectories to synthetic demonstrations via round-trips between two noisy LM components: an LM labeler which converts a trajectory into a synthetic instruction, and a zero-shot LM agent which maps the synthetic instruction into a refined trajectory. By performing these round-trips iteratively, BAGEL quickly converts the initial distribution of trajectories towards those that are well-described by natural language. We adapt the base LM agent at test time with in-context learning by retrieving relevant BAGEL demonstrations based on the instruction, and find improvements of over 2-13% absolute on ToolQA and MiniWob++, with up to 13x reduction in execution failures.
Shikhar Murty, Christopher D. Manning, Peter Shaw 0004, Mandar Joshi, Kenton Lee
ICML2
2024 ReportParse: A Unified NLP Tool for Extracting Document Structure and Semantics of Corporate Sustainability Reporting
Gaku Morio, Soh Young In, Jungah Yoon, Harri Rowlands, Christopher D. Manning
IJCAI5
2024 Handwritten Code Recognition for Pen-and-Paper CS Education
abstract
Teaching Computer Science (CS) by having students write programs by hand on paper has key pedagogical advantages: It allows focused learning and requires careful thinking compared to the use of Integrated Development Environments (IDEs) with intelligent support tools or "just trying things out". The familiar environment of pens and paper also lessens the cognitive load of students with no prior experience with computers, for whom the mere basic usage of computers can be intimidating. Finally, this teaching approach opens learning opportunities to students with limited access to computers. However, a key obstacle is the current lack of teaching methods and support software for working with and running handwritten programs. Optical character recognition (OCR) of handwritten code is challenging: Minor OCR errors, perhaps due to varied handwriting styles, easily make code not run, and recognizing indentation is crucial for languages like Python but is difficult to do due to inconsistent horizontal spacing in handwriting. Our approach integrates two innovative methods. The first combines OCR with an indentation recognition module and a language model designed for post-OCR error correction without introducing hallucinations. This method, to our knowledge, surpasses all existing systems in handwritten code recognition. It reduces error from 30% in the state of the art to 5% with minimal hallucination of logical fixes to student programs. The second method leverages a multimodal language model to recognize handwritten programs in an end-to-end fashion. We hope this contribution can stimulate further pedagogical research and contribute to the goal of making CS education universally accessible. We release a dataset of handwritten programs and code to support future research.
Md Sazzad Islam, Moussa Doumbouya, Christopher D. Manning, Chris Piech
L@S3
2024 MoEUT: Mixture-of-Experts Universal Transformers
abstract
Previous work on Universal Transformers (UTs) has demonstrated the importance of parameter sharing across layers. By allowing recurrence in depth, UTs have advantages over standard Transformers in learning compositional generalizations, but layer-sharing comes with a practical limitation of parameter-compute ratio: it drastically reduces the parameter count compared to the non-shared model with the same dimensionality. Naively scaling up the layer size to compensate for the loss of parameters makes its computational resource requirements prohibitive. In practice, no previous work has succeeded in proposing a shared-layer Transformer design that is competitive in parameter count-dominated tasks such as language modeling. Here we propose MoEUT (pronounced "moot"), an effective mixture-of-experts (MoE)-based shared-layer Transformer architecture, which combines several recent advances in MoEs for both feedforward and attention layers of standard Transformers together with novel layer-normalization and grouping schemes that are specific and crucial to UTs. The resulting UT model, for the first time, slightly outperforms standard Transformers on language modeling tasks such as BLiMP and PIQA, while using significantly less compute and memory.
Róbert Csordás, Kazuki Irie, Jürgen Schmidhuber, Christopher Potts, Christopher D. Manning
NeurIPS5
2024 ReFT: Representation Finetuning for Language Models
abstract
Parameter-efficient finetuning (PEFT) methods seek to adapt large neural models via updates to a small number of *weights*. However, much prior interpretability work has shown that *representations* encode rich semantic information, suggesting that editing representations might be a more powerful alternative. We pursue this hypothesis by developing a family of **Representation Finetuning (ReFT)** methods. ReFT methods operate on a frozen base model and learn task-specific interventions on hidden representations. We define a strong instance of the ReFT family, Low-rank Linear Subspace ReFT (LoReFT), and we identify an ablation of this method that trades some performance for increased efficiency. Both are drop-in replacements for existing PEFTs and learn interventions that are 15x--65x more parameter-efficient than LoRA. We showcase LoReFT on eight commonsense reasoning tasks, four arithmetic reasoning tasks, instruction-tuning, and GLUE. In all these evaluations, our ReFTs deliver the best balance of efficiency and performance, and almost always outperform state-of-the-art PEFTs. Upon publication, we will publicly release our generic ReFT training library.
Zhengxuan Wu, Aryaman Arora, Zheng Wang 0078, Atticus Geiger, Daniel Jurafsky, Christopher D. Manning, Christopher Potts
NeurIPS6
2023 Backpack Language Models
abstract
We present Backpacks: a new neural architecture that marries strong modeling performance with an interface for interpretability and control.Backpacks learn multiple non-contextual sense vectors for each word in a vocabulary, and represent a word in a sequence as a contextdependent, non-negative linear combination of sense vectors in this sequence.We find that, after training, sense vectors specialize, each encoding a different aspect of a word.We can interpret a sense vector by inspecting its (non-contextual, linear) projection onto the output space, and intervene on these interpretable hooks to change the model's behavior in predictable ways.We train a 170M-parameter Backpack language model on OpenWebText, matching the loss of a GPT-2 small (124Mparameter) Transformer.On lexical similarity evaluations, we find that Backpack sense vectors outperform even a 6B-parameter Transformer LM's word embeddings.Finally, we present simple algorithms that intervene on sense vectors to perform controllable text generation and debiasing.For example, we can edit the sense vocabulary to tend more towards a topic, or localize a source of gender bias to a sense vector and globally suppress that sense.
John Hewitt, John Thickstun, Christopher D. Manning, Percy Liang
ACL (1)3
2023 Self-Destructing Models: Increasing the Costs of Harmful Dual Uses of Foundation Models
abstract
A growing ecosystem of large, open-source foundation models has reduced the labeled data and technical expertise necessary to apply machine learning to many new problems. Yet foundation models pose a clear dual-use risk, indiscriminately reducing the costs of building both harmful and beneficial machine learning systems. Policy tools such as restricted model access and export controls are the primary methods currently used to mitigate such dual-use risks. In this work, we review potential safe-release strategies and argue that both policymakers and AI researchers would benefit from fundamentally new technologies enabling more precise control over the downstream usage of open-source foundation models. We propose one such approach: the task blocking paradigm, in which foundation models are trained with an additional mechanism to impede adaptation to harmful tasks without sacrificing performance on desirable tasks. We call the resulting models self-destructing models, inspired by mechanisms that prevent adversaries from using tools for harmful purposes. We present an algorithm for training self-destructing models leveraging techniques from meta-learning and adversarial learning, which we call meta-learned adversarial censoring (MLAC). In a small-scale experiment, we show MLAC can largely prevent a BERT-style model from being re-purposed to perform gender identification without harming the model’s ability to perform profession classification.
Peter Henderson 0002, Eric Mitchell, Christopher D. Manning, Daniel Jurafsky, Chelsea Finn
AIES3
2023 Meta-Learning Online Adaptation of Language Models
abstract
Large language models encode impressively broad world knowledge in their parameters.However, the knowledge in static language models falls out of date, limiting the model's effective "shelf life."While online fine-tuning can reduce this degradation, we find that naively fine-tuning on a stream of documents leads to a low level of information uptake.We hypothesize that online fine-tuning does not sufficiently attend to important information.That is, the gradient signal from important tokens representing factual information is drowned out by the gradient from inherently noisy tokens, suggesting that a dynamic, context-aware learning rate may be beneficial.We therefore propose learning which tokens to upweight.We meta-train a small, autoregressive model to reweight the language modeling loss for each token during online fine-tuning, with the objective of maximizing the out-ofdate base question-answering model's ability to answer questions about a document after a single weighted gradient step.We call this approach Context-aware Meta-learned Loss Scaling (CaMeLS).Across three different distributions of documents, our experiments find that CaMeLS provides substantially improved information uptake on streams of thousands of documents compared with standard fine-tuning and baseline heuristics for reweighting token losses.
Nathan Hu, Eric Mitchell, Christopher D. Manning, Chelsea Finn
EMNLP3
2023 Pushdown Layers: Encoding Recursive Structure in Transformer Language Models
abstract
Recursion is a prominent feature of human language, and fundamentally challenging for self-attention due to the lack of an explicit recursive-state tracking mechanism.Consequently, Transformer language models poorly capture long-tail recursive structure and exhibit sample-inefficient syntactic generalization.This work introduces Pushdown Layers, a new self-attention layer that models recursive state via a stack tape that tracks estimated depths of every token in an incremental parse of the observed prefix.Transformer LMs with Pushdown Layers are syntactic language models that autoregressively and synchronously update this stack tape as they predict new tokens, in turn using the stack tape to softly modulate attention over tokens-for instance, learning to "skip" over closed constituents.When trained on a corpus of strings annotated with silver constituency parses, Transformers equipped with Pushdown Layers achieve dramatically better and 3-5x more sample-efficient syntactic generalization, while maintaining similar perplexities.Pushdown Layers are a drop-in replacement for standard self-attention.We illustrate this by finetuning GPT2-medium with Pushdown Layers on an automatically parsed WikiText-103, leading to improvements on several GLUE text classification tasks.
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, Christopher D. Manning
EMNLP4
2023 Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
abstract
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, Christopher Manning. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, Christopher D. Manning
EMNLP8
2023 MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions
abstract
The information stored in large language models (LLMs) falls out of date quickly, and retraining from scratch is often not an option.This has recently given rise to a range of techniques for injecting new facts through updating model weights.Current evaluation paradigms are extremely limited, mainly validating the recall of edited facts, but changing one fact should cause rippling changes to the model's related beliefs.If we edit the UK Prime Minister to now be Rishi Sunak, then we should get a different answer to Who is married to the British Prime Minister?In this work, we present a benchmark, MQUAKE (Multi-hop Question Answering for Knowledge Editing), comprising multi-hop questions that assess whether edited models correctly answer questions where the answer should change as an entailed consequence of edited facts.While we find that current knowledge-editing approaches can recall edited facts accurately, they fail catastrophically on the constructed multi-hop questions.We thus propose a simple memory-based approach, MeLLo, which stores all edited facts externally while prompting the language model iteratively to generate answers that are consistent with the edited facts.While MQUAKE remains challenging, we show that MeLLo scales well with LLMs (up to 175B) and outperforms previous model editors by a large margin. 1
Zexuan Zhong, Zhengxuan Wu, Christopher D. Manning, Christopher Potts, Danqi Chen 0001
EMNLP3
2023 Characterizing intrinsic compositionality in transformers with Tree Projections
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, Christopher D. Manning
ICLR4
2023 DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature
abstract
The increasing fluency and widespread usage of large language models (LLMs) highlight the desirability of corresponding tools aiding detection of LLM-generated text. In this paper, we identify a property of the structure of an LLM's probability function that is useful for such detection. Specifically, we demonstrate that text sampled from an LLM tends to occupy negative curvature regions of the model's log probability function. Leveraging this observation, we then define a new curvature-based criterion for judging if a passage is generated from a given LLM. This approach, which we call DetectGPT, does not require training a separate classifier, collecting a dataset of real or generated passages, or explicitly watermarking generated text. It uses only log probabilities computed by the model of interest and random perturbations of the passage from another generic pre-trained language model (e.g., T5). We find DetectGPT is more discriminative than existing zero-shot methods for model sample detection, notably improving detection of fake news articles generated by 20B parameter GPT-NeoX from 0.81 AUROC for the strongest zero-shot baseline to 0.95 AUROC for DetectGPT.
Eric Mitchell, Yoonho Lee 0001, Alexander Khazatsky, Christopher D. Manning, Chelsea Finn
ICML4
2023 An NLP Benchmark Dataset for Assessing Corporate Climate Policy Engagement
abstract
As societal awareness of climate change grows, corporate climate policy engagements are attracting attention.We propose a dataset to estimate corporate climate policy engagement from various PDF-formatted documents.Our dataset comes from LobbyMap (a platform operated by global think tank InfluenceMap) that provides engagement categories and stances on the documents.To convert the LobbyMap data into the structured dataset, we developed a pipeline using text extraction and OCR.Our contributions are: (i) Building an NLP dataset including 10K documents on corporate climate policy engagement. (ii) Analyzing the properties and challenges of the dataset. (iii) Providing experiments for the dataset using pre-trained language models.The results show that while Longformer outperforms baselines and other pre-trained models, there is still room for significant improvement.We hope our work begins to bridge research on NLP and climate change.
Gaku Morio, Christopher D. Manning
NeurIPS2
2023 Direct Preference Optimization: Your Language Model is Secretly a Reward Model
abstract
While large-scale unsupervised language models (LMs) learn broad world knowledge and some reasoning skills, achieving precise control of their behavior is difficult due to the completely unsupervised nature of their training. Existing methods for gaining such steerability collect human labels of the relative quality of model generations and fine-tune the unsupervised LM to align with these preferences, often with reinforcement learning from human feedback (RLHF). However, RLHF is a complex and often unstable procedure, first fitting a reward model that reflects the human preferences, and then fine-tuning the large unsupervised LM using reinforcement learning to maximize this estimated reward without drifting too far from the original model. In this paper, we leverage a mapping between reward functions and optimal policies to show that this constrained reward maximization problem can be optimized exactly with a single stage of policy training, essentially solving a classification problem on the human preference data. The resulting algorithm, which we call Direct Preference Optimization (DPO), is stable, performant, and computationally lightweight, eliminating the need for fitting a reward model, sampling from the LM during fine-tuning, or performing significant hyperparameter tuning. Our experiments show that DPO can fine-tune LMs to align with human preferences as well as or better than existing methods. Notably, fine-tuning with DPO exceeds RLHF's ability to control sentiment of generations and improves response quality in summarization and single-turn dialogue while being substantially simpler to implement and train.
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, Chelsea Finn
NeurIPS4
2023 ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation
abstract
Abstract Compositional generalization benchmarks for semantic parsing seek to assess whether models can accurately compute meanings for novel sentences, but operationalize this in terms of logical form (LF) prediction. This raises the concern that semantically irrelevant details of the chosen LFs could shape model performance. We argue that this concern is realized for the COGS benchmark (Kim and Linzen, 2020). COGS poses generalization splits that appear impossible for present-day models, which could be taken as an indictment of those models. However, we show that the negative results trace to incidental features of COGS LFs. Converting these LFs to semantically equivalent ones and factoring out capabilities unrelated to semantic interpretation, we find that even baseline models get traction. A recent variable-free translation of COGS LFs suggests similar conclusions, but we observe this format is not semantically equivalent; it is incapable of accurately representing some COGS meanings. These findings inform our proposal for ReCOGS, a modified version of COGS that comes closer to assessing the target semantic capabilities while remaining very challenging. Overall, our results reaffirm the importance of compositional generalization and careful benchmark task design.
Zhengxuan Wu, Christopher D. Manning, Christopher Potts
Trans. Assoc. Comput. Linguistics2
2022 Synthetic Disinformation Attacks on Automated Fact Verification Systems
abstract
Automated fact-checking is a needed technology to curtail the spread of online misinformation. One current framework for such solutions proposes to verify claims by retrieving supporting or refuting evidence from related textual sources. However, the realistic use cases for fact-checkers will require verifying claims against evidence sources that could be affected by the same misinformation. Furthermore, the development of modern NLP tools that can produce coherent, fabricated content would allow malicious actors to systematically generate adversarial disinformation for fact-checkers. In this work, we explore the sensitivity of automated fact-checkers to synthetic adversarial evidence in two simulated settings: ADVERSARIAL ADDITION, where we fabricate documents and add them to the evidence repository available to the fact-checking system, and ADVERSARIAL MODIFICATION, where existing evidence source documents in the repository are automatically altered. Our study across multiple models on three benchmarks demonstrates that these systems suffer significant performance drops against these attacks. Finally, we discuss the growing threat of modern NLG systems as generators of disinformation in the context of the challenges they pose to automated fact-checkers.
Yibing Du, Antoine Bosselut, Christopher D. Manning
AAAI3
2022 Detecting Label Errors by Using Pre-Trained Language Models
abstract
We show that large pre-trained language models are inherently highly capable of identifying label errors in natural language datasets: simply examining out-of-sample data points in descending order of finetuned task loss significantly outperforms more complex error-detection mechanisms proposed in previous work.To this end, we contribute a novel method for introducing realistic, human-originated label noise into existing crowdsourced datasets such as SNLI and TweetNLP.We show that this noise has similar properties to real, handverified label errors, and is harder to detect than existing synthetic noise, creating challenges for model robustness.We argue that human-originated noise is a better standard for evaluation than synthetic noise.Finally, we use crowdsourced verification to evaluate the detection of real errors on IMDB, Amazon Reviews, and Recon, and confirm that pre-trained models perform at a 9-36% higher absolute Area Under the Precision-Recall Curve than existing models.
Derek Chong, Jenny Hong, Christopher D. Manning
EMNLP3
2022 You Only Need One Model for Open-domain Question Answering
abstract
Recent approaches to Open-domain Question Answering refer to an external knowledge base using a retriever model, optionally rerank passages with a separate reranker model and generate an answer using another reader model.Despite performing related tasks, the models have separate parameters and are weaklycoupled during training.We propose casting the retriever and the reranker as internal passage-wise attention mechanisms applied sequentially within the transformer architecture and feeding computed representations to the reader, with the hidden representations progressively refined at each stage.This allows us to use a single question answering model trained end-to-end, which is a more efficient use of model capacity and also leads to better gradient flow.We present a pre-training method to effectively train this architecture and evaluate our model on the Natural Questions and TriviaQA open datasets.For a fixed parameter budget, our model outperforms the previous state-of-the-art model by 1.0 and 0.7 exact match scores.Reading Layer Reranking Layer Retrieval Layer Encoder Query Passage 1 Passage 2 Passage 3 Passage 4 Passage n Encoder Encoder Encoder Encoder Encoder Encoder … … Passage-wise Disjoint Attention Encoder Encoder Encoder Encoder Passage-wise Joint Attention Concat Query and Passage Answer Concat All Decoder Encoder Encoder Encoder … Decoder
Haejun Lee, Akhil Kedia, Ashwin Paranjape, Christopher D. Manning, Kyoung-Gu Woo
EMNLP5
2022 Enhancing Self-Consistency and Performance of Pre-Trained Language Models through Natural Language Inference
abstract
Eric Mitchell, Joseph Noh, Siyan Li, Will Armstrong, Ananth Agarwal, Patrick Liu, Chelsea Finn, Christopher Manning. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Eric Mitchell, Joseph J. Noh, Siyan Li, William S. Armstrong, Ananth Agarwal, Patrick Liu, Chelsea Finn, Christopher D. Manning
EMNLP8
2022 Fixing Model Bugs with Natural Language Patches
abstract
Current approaches for fixing systematic problems in NLP models (e.g., regex patches, finetuning on more data) are either brittle, or labor-intensive and liable to shortcuts.In contrast, humans often provide corrections to each other through natural language.Taking inspiration from this, we explore natural language patches-declarative statements that allow developers to provide corrective feedback at the right level of abstraction, either overriding the model ("if a review gives 2 stars, the sentiment is negative") or providing additional information the model may lack ("if something is described as the bomb, then it is good").We model the task of determining if a patch applies separately from the task of integrating patch information, and show that with a small amount of synthetic data, we can teach models to effectively use real patches on real data-1 to 7 patches improve accuracy by ~1-4 accuracy points on different slices of a sentiment analysis dataset, and F1 by 7 points on a relation extraction dataset.Finally, we show that finetuning on as many as 100 labeled examples may be needed to match the performance of a small set of language patches.
Shikhar Murty, Christopher D. Manning, Scott M. Lundberg, Marco Túlio Ribeiro
EMNLP2
2022 On Measuring the Intrinsic Few-Shot Hardness of Datasets
abstract
While advances in pre-training have led to dramatic improvements in few-shot learning of NLP tasks, there is limited understanding of what drives successful few-shot adaptation in datasets.In particular, given a new dataset and a pre-trained model, what properties of the dataset make it few-shot learnable and are these properties independent of the specific adaptation techniques used?We consider an extensive set of recent few-shot learning methods, and show that their performance across a large number of datasets is highly correlated, showing that few-shot hardness may be intrinsic to datasets, for a given pre-trained model.To estimate intrinsic few-shot hardness, we then propose a simple and lightweight metric called Spread that captures the intuition that fewshot learning is made possible by exploiting feature-space invariances between training and test samples.Our metric better accounts for few-shot hardness compared to existing notions of hardness, and is ~8-100x faster to compute. ⋆ Equal ContributionMethod D1 D2 LMBFF 45.3 -0.4 NullPrompts 43.0 -5.7 BitFit 46.3 -3.5 AdaPET 44.9 -0.3 P-Tuning 46.3 0.3 Few-shot (Avg) 45.2 -2 Full Fine-tuning 45.3 35
Shikhar Murty, Christopher D. Manning
EMNLP3
2022 GreaseLM: Graph REASoning Enhanced Language Models
Xikun Zhang 0001, Antoine Bosselut, Michihiro Yasunaga, Hongyu Ren, Percy Liang, Christopher D. Manning, Jure Leskovec
ICLR6
2022 Fast Model Editing at Scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, Christopher D. Manning
ICLR5
2022 Hindsight: Posterior-guided training of retrievers for improved open-ended generation
Ashwin Paranjape, Omar Khattab, Christopher Potts, Matei Zaharia, Christopher D. Manning
ICLR5
2022 Memory-Based Model Editing at Scale
abstract
Even the largest neural networks make errors, and once-correct predictions can become invalid as the world changes. Model editors make local updates to the behavior of base (pre-trained) models to inject updated knowledge or correct undesirable behaviors. Existing model editors have shown promise, but also suffer from insufficient expressiveness: they struggle to accurately model an edit’s intended scope (examples affected by the edit), leading to inaccurate predictions for test inputs loosely related to the edit, and they often fail altogether after many edits. As a higher-capacity alternative, we propose Semi-Parametric Editing with a Retrieval-Augmented Counterfactual Model (SERAC), which stores edits in an explicit memory and learns to reason over them to modulate the base model’s predictions as needed. To enable more rigorous evaluation of model editors, we introduce three challenging language model editing problems based on question answering, fact-checking, and dialogue generation. We find that only SERAC achieves high performance on all three problems, consistently outperforming existing approaches to model editing by a significant margin. Code, data, and additional project information will be made available at https://sites.google.com/view/serac-editing.
Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning, Chelsea Finn
ICML4
2022 Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset
abstract
One concern with the rise of large language models lies with their potential for significant harm, particularly from pretraining on biased, obscene, copyrighted, and private information. Emerging ethical approaches have attempted to filter pretraining material, but such approaches have been ad hoc and failed to take context into account. We offer an approach to filtering grounded in law, which has directly addressed the tradeoffs in filtering material. First, we gather and make available the Pile of Law, a ~256GB (and growing) dataset of open-source English-language legal and administrative data, covering court opinions, contracts, administrative rules, and legislative records. Pretraining on the Pile of Law may help with legal tasks that have the promise to improve access to justice. Second, we distill the legal norms that governments have developed to constrain the inclusion of toxic or private content into actionable lessons for researchers and discuss how our dataset reflects these norms. Third, we show how the Pile of Law offers researchers the opportunity to learn such filtering rules directly from the data, providing an exciting new research direction in model-based processing.
Peter Henderson 0002, Mark S. Krass, Lucia Zheng, Neel Guha, Christopher D. Manning, Daniel Jurafsky, Daniel E. Ho
NeurIPS5
2022 Deep Bidirectional Language-Knowledge Graph Pretraining
abstract
Pretraining a language model (LM) on text has been shown to help various downstream NLP tasks. Recent works show that a knowledge graph (KG) can complement text data, offering structured background knowledge that provides a useful scaffold for reasoning. However, these works are not pretrained to learn a deep fusion of the two modalities at scale, limiting the potential to acquire fully joint representations of text and KG. Here we propose DRAGON (Deep Bidirectional Language-Knowledge Graph Pretraining), a self-supervised approach to pretraining a deeply joint language-knowledge foundation model from text and KG at scale. Specifically, our model takes pairs of text segments and relevant KG subgraphs as input and bidirectionally fuses information from both modalities. We pretrain this model by unifying two self-supervised reasoning tasks, masked language modeling and KG link prediction. DRAGON outperforms existing LM and LM+KG models on diverse downstream tasks including question answering across general and biomedical domains, with +5% absolute gain on average. In particular, DRAGON achieves notable performance on complex reasoning about language and knowledge (+10% on questions involving long contexts or multi-step reasoning) and low-resource QA (+8% on OBQA and RiddleSense), and new state-of-the-art results on various BioNLP tasks. Our code and trained models are available at https://github.com/michiyasunaga/dragon.
Michihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang 0001, Christopher D. Manning, Percy Liang, Jure Leskovec
NeurIPS5
2022 Neural Generation Meets Real People: Building a Social, Informative Open-Domain Dialogue Agent
abstract
Ethan A. Chi, Ashwin Paranjape, Abigail See, Caleb Chiam, Trenton Chang, Kathleen Kenealy, Swee Kiat Lim, Amelia Hardy, Chetanya Rastogi, Haojun Li, Alexander Iyabor, Yutong He, Hari Sowrirajan, Peng Qi, Kaushik Ram Sadagopan, Nguyet Minh Phu, Dilara Soylu, Jillian Tang, Avanika Narayan, Giovanni Campagna, Christopher Manning. Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2022.
Ethan A. Chi, Ashwin Paranjape, Abigail See, Caleb Chiam, Trenton Chang, Kathleen Kenealy, Swee Kiat Lim, Amelia F. Hardy, Chetanya Rastogi, Alexander Iyabor, Hari Sowrirajan, Peng Qi 0003, Kaushik Ram Sadagopan, Nguyet Minh Phu, Dilara Soylu, Jillian Tang, Avanika Narayan, Giovanni Campagna, Christopher D. Manning
SIGDIAL21
2022 When can I Speak? Predicting initiation points for spoken dialogue agents
abstract
Current spoken dialogue systems initiate their turns after a long period of silence (700-1000ms), which leads to little real-time feedback, sluggish responses, and an overall stilted conversational flow.Humans typically respond within 200ms and successfully predicting initiation points in advance would allow spoken dialogue agents to do the same.In this work, we predict the lead-time to initiation using prosodic features from a pre-trained speech representation model (wav2vec 1.0) operating on user audio and word features from a pre-trained language model (GPT-2) operating on incremental transcriptions.To evaluate errors, we propose two metrics w.r.t.predicted and true lead times.We train and evaluate the models on the Switchboard Corpus and find that our method outperforms features from prior work on both metrics and vastly outperforms the common approach of waiting for 700ms of silence.
Siyan Li, Ashwin Paranjape, Christopher D. Manning
SIGDIAL3
2021 Mind Your Outliers! Investigating the Negative Impact of Outliers on Active Learning for Visual Question Answering
abstract
Siddharth Karamcheti, Ranjay Krishna, Li Fei-Fei, Christopher Manning. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Siddharth Karamcheti, Ranjay Krishna, Li Fei-Fei 0001, Christopher D. Manning
ACL/IJCNLP (1)4
2021 Answering Open-Domain Questions of Varying Reasoning Steps from Text
abstract
We develop a unified system to answer directly from text open-domain questions that may require a varying number of retrieval steps.We employ a single multi-task transformer model to perform all the necessary subtasks-retrieving supporting facts, reranking them, and predicting the answer from all retrieved documents-in an iterative fashion.We avoid crucial assumptions of previous work that do not transfer well to real-world settings, including exploiting knowledge of the fixed number of retrieval steps required to answer each question or using structured metadata like knowledge bases or web links that have limited availability.Instead, we design a system that can answer open-domain questions on any text collection without prior knowledge of reasoning complexity.To emulate this setting, we construct a new benchmark, called BeerQA, by combining existing one-and twostep datasets with a new collection of 530 questions that require three Wikipedia pages to answer, unifying Wikipedia corpora versions in the process.We show that our model demonstrates competitive performance on both existing benchmarks and this new benchmark.We make the new benchmark available at https: //beerqa.github.io/.
Peng Qi 0003, Haejun Lee, Tg Sido, Christopher D. Manning
EMNLP (1)4
2021 Conditional probing: measuring usable information beyond a baseline
abstract
Probing experiments investigate the extent to which neural representations make properties-like part-of-speech-predictable.One suggests that a representation encodes a property if probing that representation produces higher accuracy than probing a baseline representation like non-contextual word embeddings.Instead of using baselines as a point of comparison, we're interested in measuring information that is contained in the representation but not in the baseline.For example, current methods can detect when a representation is more useful than the word identity (a baseline) for predicting part-ofspeech; however, they cannot detect when the representation is predictive of just the aspects of part-of-speech not explainable by the word identity.In this work, we extend a theory of usable information called V-information and propose conditional probing, which explicitly conditions on the information in the baseline.In a case study, we find that after conditioning on non-contextual word embeddings, properties like part-of-speech are accessible at deeper layers of a network than previously thought.
John Hewitt, Kawin Ethayarajh, Percy Liang, Christopher D. Manning
EMNLP (1)4
2021 DReCa: A General Task Augmentation Strategy for Few-Shot Natural Language Inference
abstract
Shikhar Murty, Tatsunori B. Hashimoto, Christopher Manning. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Shikhar Murty, Tatsunori B. Hashimoto, Christopher D. Manning
NAACL-HLT3
2021 Human-like informative conversations: Better acknowledgements using conditional mutual information
abstract
This work aims to build a dialogue agent that can weave new factual content into conversations as naturally as humans.We draw insights from linguistic principles of conversational analysis and annotate human-human conversations from the Switchboard Dialog Act Corpus to examine humans strategies for acknowledgement, transition, detail selection and presentation.When current chatbots (explicitly provided with new factual content) introduce facts into a conversation, their generated responses do not acknowledge the prior turns.This is because models trained with two contexts -new factual content and conversational history -generate responses that are non-specific w.r.t.one of the contexts, typically the conversational history.We show that specificity w.r.t.conversational history is better captured by pointwise conditional mutual information (pcmi h ) than by the established use of pointwise mutual information (pmi).Our proposed method, Fused-PCMI, trades off pmi for pcmi h and is preferred by humans for overall quality over the Max-PMI baseline 60% of the time.Human evaluators also judge responses with higher pcmi h better at acknowledgement 74% of the time.The results demonstrate that systems mimicking human conversational traits (in this case acknowledgement) improve overall quality and more broadly illustrate the utility of linguistic principles in improving dialogue agents.
Ashwin Paranjape, Christopher D. Manning
NAACL-HLT2
2021 Effective Social Chatbot Strategies for Increasing User Initiative
abstract
Many existing chatbots do not effectively support mixed initiative, forcing their users to either respond passively or lead constantly.We seek to improve this experience by introducing new mechanisms to encourage user initiative in social chatbot conversations.Since user initiative in this setting is distinct from initiative in human-human or task-oriented dialogue, we first propose a new definition that accounts for the unique behaviors users take in this context.Drawing from linguistics, we propose three mechanisms to promote user initiative: back-channeling, personal disclosure, and replacing questions with statements.We show that simple automatic metrics of utterance length, number of noun phrases, and diversity of user responses correlate with human judgement of initiative.Finally, we use these metrics to suggest that these strategies do result in statistically significant increases in user initiative, where frequent, but not excessive, back-channeling is the most effective strategy.
Amelia F. Hardy, Ashwin Paranjape, Christopher D. Manning
SIGDIAL3
2021 Large-Scale Quantitative Evaluation of Dialogue Agents' Response Strategies against Offensive Users
abstract
As voice assistants and dialogue agents grow in popularity, so does the abuse they receive.We conducted a large-scale quantitative evaluation of the effectiveness of 4 response types (avoidance, why, empathetic, and counter), and 2 additional factors (using a redirect or a voluntarily provided name) that have not been tested by prior work.We measured their direct effectiveness on real users in-the-wild by the re-offense ratio, length of conversation after the initial response, and number of turns until the next re-offense.Our experiments confirm prior lab studies in showing that empathetic responses perform better than generic avoidance responses as well as counter responses.We show that dialogue agents should almost always guide offensive users to a new topic through the use of redirects and use the user's name if provided.As compared to a baseline avoidance strategy employed by commercial agents, our best strategy is able to reduce the re-offense ratio from 92% to 43%.
Dilara Soylu, Christopher D. Manning
SIGDIAL3
2021 Understanding and predicting user dissatisfaction in a neural generative chatbot
abstract
Neural generative dialogue agents have shown an increasing ability to hold short chitchat conversations, when evaluated by crowdworkers in controlled settings.However, their performance in real-life deployment -talking to intrinsically-motivated users in noisy environments -is less well-explored.In this paper, we perform a detailed case study of a neural generative model deployed as part of Chirpy Cardinal, an Alexa Prize socialbot.We find that unclear user utterances are a major source of generative errors such as ignoring, hallucination, unclearness and repetition.However, even in unambiguous contexts the model frequently makes reasoning errors.Though users express dissatisfaction in correlation with these errors, certain dissatisfaction types (such as offensiveness and privacy objections) depend on additional factors -such as the user's personal attitudes, and prior unaddressed dissatisfaction in the conversation.Finally, we show that dissatisfied user utterances can be used as a semisupervised learning signal to improve the dialogue system.We train a model to predict nextturn dissatisfaction, and show through human evaluation that as a ranking function, it selects higher-quality neural-generated utterances.
Abigail See, Christopher D. Manning
SIGDIAL2
2021 Universal Dependencies
abstract
Abstract Universal dependencies (UD) is a framework for morphosyntactic annotation of human language, which to date has been used to create treebanks for more than 100 languages. In this article, we outline the linguistic theory of the UD framework, which draws on a long tradition of typologically oriented grammatical theories. Grammatical relations between words are centrally used to explain how predicate–argument structures are encoded morphosyntactically in different languages while morphological features and part-of-speech classes give the properties of words. We argue that this theory is a good basis for crosslinguistically consistent annotation of typologically diverse languages in a way that supports computational natural language understanding as well as broader linguistic studies.
Marie-Catherine de Marneffe, Christopher D. Manning, Joakim Nivre, Daniel Zeman
Comput. Linguistics2
2021 Biomedical and clinical English model packages for the Stanza Python NLP library
abstract
OBJECTIVE: The study sought to develop and evaluate neural natural language processing (NLP) packages for the syntactic analysis and named entity recognition of biomedical and clinical English text. MATERIALS AND METHODS: We implement and train biomedical and clinical English NLP pipelines by extending the widely used Stanza library originally designed for general NLP tasks. Our models are trained with a mix of public datasets such as the CRAFT treebank as well as with a private corpus of radiology reports annotated with 5 radiology-domain entities. The resulting pipelines are fully based on neural networks, and are able to perform tokenization, part-of-speech tagging, lemmatization, dependency parsing, and named entity recognition for both biomedical and clinical text. We compare our systems against popular open-source NLP libraries such as CoreNLP and scispaCy, state-of-the-art models such as the BioBERT models, and winning systems from the BioNLP CRAFT shared task. RESULTS: For syntactic analysis, our systems achieve much better performance compared with the released scispaCy models and CoreNLP models retrained on the same treebanks, and are on par with the winning system from the CRAFT shared task. For NER, our systems substantially outperform scispaCy, and are better or on par with the state-of-the-art performance from BioBERT, while being much more computationally efficient. CONCLUSIONS: We introduce biomedical and clinical NLP packages built for the Stanza library. These packages offer performance that is similar to the state of the art, and are also optimized for ease of use. To facilitate research, we make all our models publicly available. We also provide an online demonstration (http://stanza.run/bio).
Yuhao Zhang 0004, Peng Qi 0003, Christopher D. Manning, Curt Langlotz
J. Am. Medical Informatics Assoc.4
2020 Finding Universal Grammatical Relations in Multilingual BERT
abstract
Recent work has found evidence that Multilingual BERT (mBERT), a transformer-based multilingual masked language model, is capable of zero-shot cross-lingual transfer, suggesting that some aspects of its representations are shared cross-lingually.To better understand this overlap, we extend recent work on finding syntactic trees in neural networks' internal representations to the multilingual setting.We show that subspaces of mBERT representations recover syntactic tree distances in languages other than English, and that these subspaces are approximately shared across languages.Motivated by these results, we present an unsupervised analysis method that provides evidence mBERT learns representations of syntactic dependency labels, in the form of clusters which largely agree with the Universal Dependencies taxonomy.This evidence suggests that even without explicit supervision, multilingual masked language models learn certain linguistic universals.
Ethan A. Chi, John Hewitt, Christopher D. Manning
ACL3
2020 Syn-QG: Syntactic and Shallow Semantic Rules for Question Generation
abstract
Question Generation (QG) is fundamentally a simple syntactic transformation; however, many aspects of semantics influence what questions are good to form.We implement this observation by developing Syn-QG, a set of transparent syntactic rules leveraging universal dependencies, shallow semantic parsing, lexical resources, and custom rules which transform declarative sentences into questionanswer pairs.We utilize PropBank argument descriptions and VerbNet state predicates to incorporate shallow semantic content, which helps generate questions of a descriptive nature and produce inferential and semantically richer questions than existing systems.In order to improve syntactic fluency and eliminate grammatically incorrect questions, we employ back-translation over the output of these syntactic rules.A set of crowd-sourced evaluations shows that our system can generate a larger number of highly grammatical and relevant questions than previous QG systems and that back-translation drastically improves grammaticality at a slight cost of generating irrelevant questions.
Kaustubh D. Dhole, Christopher D. Manning
ACL2
2020 Optimizing the Factual Correctness of a Summary: A Study of Summarizing Radiology Reports
abstract
Neural abstractive summarization models are able to generate summaries which have high overlap with human references.However, existing models are not optimized for factual correctness, a critical metric in real-world applications.In this work, we develop a general framework where we evaluate the factual correctness of a generated summary by factchecking it automatically against its reference using an information extraction module.We further propose a training strategy which optimizes a neural summarization model with a factual correctness reward via reinforcement learning.We apply the proposed method to the summarization of radiology reports, where factual correctness is a key requirement.On two separate datasets collected from hospitals, we show via both automatic and human evaluation that the proposed approach substantially improves the factual correctness and overall quality of outputs over a competitive neural summarization system, producing radiology summaries that approach the quality of humanauthored ones.Background: radiographic examination of the chest.clinical history: 80 years of age, male ... Findings: frontal radiograph of the chest demonstrates repositioning of the right atrial lead possibly into the ivc.... a right apical pneumothorax can be seen from the image.moderate right and small left pleural effusions continue.no pulmonary edema is observed.heart size is upper limits of normal. Human Summary: pneumothorax is seen. bilateral pleural effusions continue.Summary A (ROUGE-L = 0.77): no pneumothorax is observed.bilateral pleural effusions continue.Summary B (ROUGE-L = 0.44): pneumothorax is observed on radiograph.bilateral pleural effusions continue to be seen.
Yuhao Zhang 0004, Derek Merck, Emily Bao Tsai, Christopher D. Manning, Curt Langlotz
ACL4
2020 Pre-Training Transformers as Energy-Based Cloze Models
abstract
We introduce Electric, an energy-based cloze model for representation learning over text.Like BERT, it is a conditional generative model of tokens given their contexts.However, Electric does not use masking or output a full distribution over tokens that could occur in a context.Instead, it assigns a scalar energy score to each input token indicating how likely it is given its context.We train Electric using an algorithm based on noise-contrastive estimation and elucidate how this learning objective is closely related to the recently proposed ELECTRA pre-training method.Electric performs well when transferred to downstream tasks and is particularly effective at producing likelihood scores for text: it reranks speech recognition n-best lists better than language models and much faster than masked language models.Furthermore, it offers a clearer and more principled view of what ELECTRA learns during pre-training.
Kevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. Manning
EMNLP (1)4
2020 RNNs can generate bounded hierarchical languages with optimal memory
abstract
Recurrent neural networks empirically generate natural language with high syntactic fidelity.However, their success is not wellunderstood theoretically.We provide theoretical insight into this success, proving in a finiteprecision setting that RNNs can efficiently generate bounded hierarchical languages that reflect the scaffolding of natural language syntax.We introduce Dyck-(k,m), the language of well-nested brackets (of k types) and mbounded nesting depth, reflecting the bounded memory needs and long-distance dependencies of natural language syntax.The best known results use O(k m 2 ) memory (hidden units) to generate these languages.We prove that an RNN with O(m log k) hidden units suffices, an exponential reduction in memory, by an explicit construction.Finally, we show that no algorithm, even with unbounded computation, can suffice with o(m log k) hidden units.
John Hewitt, Michael Hahn 0001, Surya Ganguli, Percy Liang, Christopher D. Manning
EMNLP (1)5
2020 SLM: Learning a Discourse Language Representation with Sentence Unshuffling
abstract
We introduce Sentence-level Language Modeling, a new pre-training objective for learning a discourse language representation in a fully self-supervised manner.Recent pre-training methods in NLP focus on learning either bottom or top-level language representations: contextualized word representations derived from language model objectives at one extreme and a whole sequence representation learned by order classification of two given textual segments at the other.However, these models are not directly encouraged to capture representations of intermediate-size structures that exist in natural languages such as sentences and the relationships among them.To that end, we propose a new approach to encourage learning of a contextualized sentence-level representation by shuffling the sequence of input sentences and training a hierarchical transformer model to reconstruct the original ordering.Through experiments on downstream tasks such as GLUE, SQuAD, and DiscoEval, we show that this feature of our model improves the performance of the original BERT by large margins.
Haejun Lee, Drew A. Hudson, Kangwook Lee 0004, Christopher D. Manning
EMNLP (1)4
2020 ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. Manning
ICLR4
2020 Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection
abstract
Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework. The annotation consists in a linguistically motivated word segmentation; a morphological layer comprising lemmas, universal part-of-speech tags, and standardized morphological features; and a syntactic layer focusing on syntactic relations between predicates, arguments and modifiers. In this paper, we describe version 2 of the universal guidelines (UD v2), discuss the major changes from UD v1 to UD v2, and give an overview of the currently available treebanks for 90 languages.
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajic 0001, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster 0001, Francis M. Tyers, Daniel Zeman
LREC5
2019 BAM! Born-Again Multi-Task Networks for Natural Language Understanding
abstract
It can be challenging to train multi-task neural networks that outperform or even match their single-task counterparts.To help address this, we propose using knowledge distillation where single-task models teach a multi-task model.We enhance this training with teacher annealing, a novel method that gradually transitions the model from distillation to supervised learning, helping the multi-task model surpass its single-task teachers.We evaluate our approach by multi-task fine-tuning BERT on the GLUE benchmark.Our method consistently improves over standard single-task and multi-task training.
Kevin Clark, Minh-Thang Luong, Urvashi Khandelwal, Christopher D. Manning, Quoc V. Le
ACL (1)4
2019 Do Massively Pretrained Language Models Make Better Storytellers?
abstract
Large neural language models trained on massive amounts of text have emerged as a formidable strategy for Natural Language Understanding tasks.However, the strength of these models as Natural Language Generators is less clear.Though anecdotal evidence suggests that these models generate better quality text, there has been no detailed study characterizing their generation abilities.In this work, we compare the performance of an extensively pretrained model, OpenAI GPT2-117 (Radford et al., 2019), to a state-of-the-art neural story generation model (Fan et al., 2018).By evaluating the generated text across a wide variety of automatic metrics, we characterize the ways in which pretrained models do, and do not, make better storytellers.We find that although GPT2-117 conditions more strongly on context, is more sensitive to ordering of events, and uses more unusual words, it is just as likely to produce repetitive and under-diverse text when using likelihood-maximizing decoding algorithms.
Abigail See, Aneesh Pappu, Rohun Saxena, Akhila Yerukola, Christopher D. Manning
CoNLL5
2019 GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
abstract
We introduce GQA, a new dataset for real-world visual reasoning and compositional question answering, seeking to address key shortcomings of previous VQA datasets. We have developed a strong and robust question engine that leverages Visual Genome scene graph structures to create 22M diverse reasoning questions, which all come with functional programs that represent their semantics. We use the programs to gain tight control over the answer distribution and present a new tunable smoothing technique to mitigate question biases. Accompanying the dataset is a suite of new metrics that evaluate essential qualities such as consistency, grounding and plausibility. A careful analysis is performed for baselines as well as state-of-the-art models, providing fine-grained results for different question types and topologies. Whereas a blind LSTM obtains a mere 42.1%, and strong VQA models achieve 54.1%, human performance tops at 89.3%, offering ample opportunity for new research to explore. We hope GQA will provide an enabling resource for the next generation of models with enhanced robustness, improved consistency, and deeper semantic understanding of vision and language.
Drew A. Hudson, Christopher D. Manning
CVPR2
2019 Answering Complex Open-domain Questions Through Iterative Query Generation
abstract
Peng Qi, Xiaowen Lin, Leo Mehr, Zijian Wang, Christopher D. Manning. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Peng Qi 0003, Xiaowen Lin, Leo Mehr, Zijian Wang 0002, Christopher D. Manning
EMNLP/IJCNLP (1)5
2019 Learning by Abstraction: The Neural State Machine
abstract
We introduce the Neural State Machine, seeking to bridge the gap between the neural and symbolic views of AI and integrate their complementary strengths for the task of visual reasoning. Given an image, we first predict a probabilistic graph that represents its underlying semantics and serves as a structured world model. Then, we perform sequential reasoning over the graph, iteratively traversing its nodes to answer a given question or draw a new inference. In contrast to most neural architectures that are designed to closely interact with the raw sensory data, our model operates instead in an abstract latent space, by transforming both the visual and linguistic modalities into semantic concept-based representations, thereby achieving enhanced transparency and modularity. We evaluate our model on VQA-CP and GQA, two recent VQA datasets that involve compositionality, multi-step inference and diverse reasoning skills, achieving state-of-the-art results in both cases. We provide further experiments that illustrate the model's strong generalization capacity across multiple dimensions, including novel compositions of concepts, changes in the answer distribution, and unseen linguistic structures, demonstrating the qualities and efficacy of our approach.
Drew A. Hudson, Christopher D. Manning
NeurIPS2
2019 CoQA: A Conversational Question Answering Challenge
abstract
Humans gather information through conversations involving a series of interconnected questions and answers. For machines to assist in information gathering, it is therefore essential to enable them to answer conversational questions. We introduce CoQA, a novel dataset for building Conversational Question Answering systems. Our dataset contains 127k questions with answers, obtained from 8k conversations about text passages from seven diverse domains. The questions are conversational, and the answers are free-form text with their corresponding evidence highlighted in the passage. We analyze CoQA in depth and show that conversational questions have challenging phenomena not present in existing reading comprehension datasets (e.g., coreference and pragmatic reasoning). We evaluate strong dialogue and reading comprehension models on CoQA. The best system obtains an F1 score of 65.4%, which is 23.4 points behind human performance (88.8%), indicating that there is ample room for improvement. We present CoQA as a challenge to the community at https://stanfordnlp.github.io/coqa .
Siva Reddy, Danqi Chen 0001, Christopher D. Manning
Trans. Assoc. Comput. Linguistics3
2018 Semi-Supervised Sequence Modeling with Cross-View Training
abstract
Unsupervised representation learning algorithms such as word2vec and ELMo improve the accuracy of many supervised NLP models, mainly because they can take advantage of large amounts of unlabeled text.However, the supervised models only learn from taskspecific labeled data during the main training phase.We therefore propose Cross-View Training (CVT), a semi-supervised learning algorithm that improves the representations of a Bi-LSTM sentence encoder using a mix of labeled and unlabeled data.On labeled examples, standard supervised learning is used.On unlabeled examples, CVT teaches auxiliary prediction modules that see restricted views of the input (e.g., only part of a sentence) to match the predictions of the full model seeing the whole input.Since the auxiliary modules and the full model share intermediate representations, this in turn improves the full model.Moreover, we show that CVT is particularly effective when combined with multitask learning.We evaluate CVT on five sequence tagging tasks, machine translation, and dependency parsing, achieving state-of-the-art results. 1
Kevin Clark, Minh-Thang Luong, Christopher D. Manning, Quoc V. Le
EMNLP3
2018 Textual Analogy Parsing: What's Shared and What's Compared among Analogous Facts
abstract
To understand a sentence like "whereas only 10% of White Americans live at or below the poverty line, 28% of African Americans do" it is important not only to identify individual facts, e.g., poverty rates of distinct demographic groups, but also the higher-order relations between them, e.g., the disparity between them.In this paper, we propose the task of Textual Analogy Parsing (TAP) to model this higher-order meaning.The output of TAP is a frame-style meaning representation which explicitly specifies what is shared (e.g., poverty rates) and what is compared (e.g., White Americans vs. African Americans, 10% vs. 28%) between its component facts.Such a meaning representation can enable new applications that rely on discourse understanding such as automated chart generation from quantitative text.We present a new dataset for TAP, baselines, and a model that successfully uses an ILP to enforce the structural constraints of the problem.
Matthew Lamm, Arun Tejasvi Chaganty, Christopher D. Manning, Daniel Jurafsky, Percy Liang
EMNLP3
2018 HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
abstract
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, Christopher D. Manning. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Zhilin Yang 0001, Peng Qi 0003, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, Christopher D. Manning
EMNLP7
2018 Graph Convolution over Pruned Dependency Trees Improves Relation Extraction
abstract
Dependency trees help relation extraction models capture long-range relations between words.However, existing dependency-based models either neglect crucial information (e.g., negation) by pruning the dependency trees too aggressively, or are computationally inefficient because it is difficult to parallelize over different tree structures.We propose an extension of graph convolutional networks that is tailored for relation extraction, which pools information over arbitrary dependency structures efficiently in parallel.To incorporate relevant information while maximally removing irrelevant content, we further apply a novel pruning strategy to the input trees by keeping words immediately around the shortest path between the two entities among which a relation might hold.The resulting model achieves state-of-the-art performance on the large-scale TACRED dataset, outperforming existing sequence and dependency-based neural models.We also show through detailed analysis that this model has complementary strengths to sequence models, and combining them further improves the state of the art.
Yuhao Zhang 0004, Peng Qi 0003, Christopher D. Manning
EMNLP3
2018 Compositional Attention Networks for Machine Reasoning
Drew A. Hudson, Christopher D. Manning
ICLR (Poster)2
2018 Sentences with Gapping: Parsing and Reconstructing Elided Predicates
abstract
Sebastian Schuster, Joakim Nivre, Christopher D. Manning. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Sebastian Schuster 0001, Joakim Nivre, Christopher D. Manning
NAACL-HLT3
2017 Get To The Point: Summarization with Pointer-Generator Networks
abstract
Neural sequence-to-sequence models have provided a viable new approach for abstractive text summarization (meaning they are not restricted to simply selecting and rearranging passages from the original text).However, these models have two shortcomings: they are liable to reproduce factual details inaccurately, and they tend to repeat themselves.In this work we propose a novel architecture that augments the standard sequence-to-sequence attentional model in two orthogonal ways.First, we use a hybrid pointer-generator network that can copy words from the source text via pointing, which aids accurate reproduction of information, while retaining the ability to produce novel words through the generator.Second, we use coverage to keep track of what has been summarized, which discourages repetition.We apply our model to the CNN / Daily Mail summarization task, outperforming the current abstractive state-of-the-art by at least 2 ROUGE points.
Abigail See, Peter J. Liu, Christopher D. Manning
ACL (1)3
2017 Naturalizing a Programming Language via Interactive Learning
abstract
Our goal is to create a convenient natural language interface for performing wellspecified but complex actions such as analyzing data, manipulating text, and querying databases.However, existing natural language interfaces for such tasks are quite primitive compared to the power one wields with a programming language.To bridge this gap, we start with a core programming language and allow users to "naturalize" the core language incrementally by defining alternative, more natural syntax and increasingly complex concepts in terms of compositions of simpler ones.In a voxel world, we show that a community of users can simultaneously teach a common system a diverse language and use it to build hundreds of complex voxel structures.Over the course of three days, these users went from using only the core language to using the naturalized language in 85.9% of the last 10K utterances.
Sida I. Wang, Samuel Ginn, Percy Liang, Christopher D. Manning
ACL (1)4
2017 Importance sampling for unbiased on-demand evaluation of knowledge base population
abstract
Knowledge base population (KBP) systems take in a large document corpus and extract entities and their relations.Thus far, KBP evaluation has relied on judgements on the pooled predictions of existing systems.We show that this evaluation is problematic: when a new system predicts a previously unseen relation, it is penalized even if it is correct.This leads to significant bias against new systems, which counterproductively discourages innovation in the field.Our first contribution is a new importance-sampling based evaluation which corrects for this bias by annotating a new system's predictions ondemand via crowdsourcing.We show this eliminates bias and reduces variance using data from the 2015 TAC KBP task.Our second contribution is an implementation of our method made publicly available as an online KBP evaluation service.We pilot the service by testing diverse state-ofthe-art systems on the TAC KBP 2016 corpus and obtain accurate scores in a cost effective manner.
Arun Tejasvi Chaganty, Ashwin Paranjape, Percy Liang, Christopher D. Manning
EMNLP4
2017 Position-aware Attention and Supervised Data Improve Slot Filling
abstract
Organized relational knowledge in the form of "knowledge graphs" is important for many applications.However, the ability to populate knowledge bases with facts automatically extracted from documents has improved frustratingly slowly.This paper simultaneously addresses two issues that have held back prior work.We first propose an effective new model, which combines an LSTM sequence model with a form of entity position-aware attention that is better suited to relation extraction.Then we build TACRED, a large (119,474 examples) supervised relation extraction dataset, obtained via crowdsourcing and targeted towards TAC KBP relations.The combination of better supervised data and a more appropriate high-capacity model enables much better relation extraction performance.When the model trained on this new dataset replaces the previous relation extraction component of the best TAC KBP 2015 slot filling system, its F 1 score increases markedly from 22.2% to 26.7%.
Yuhao Zhang 0004, Victor Zhong, Danqi Chen 0001, Gabor Angeli, Christopher D. Manning
EMNLP5
2017 Deep Biaffine Attention for Neural Dependency Parsing
Timothy Dozat, Christopher D. Manning
ICLR (Poster)2
2017 Key-Value Retrieval Networks for Task-Oriented Dialogue
abstract
Neural task-oriented dialogue systems often struggle to smoothly interface with a knowledge base.In this work, we seek to address this problem by proposing a new neural dialogue agent that is able to effectively sustain grounded, multi-domain discourse through a novel key-value retrieval mechanism.The model is end-to-end differentiable and does not need to explicitly model dialogue state or belief trackers.We also release a new dataset of 3,031 dialogues that are grounded through underlying knowledge bases and span three distinct tasks in the in-car personal assistant space: calendar scheduling, weather information retrieval, and point-of-interest navigation.Our architecture is simultaneously trained on data from all domains and significantly outperforms a competitive rulebased system and other existing neural dialogue architectures on the provided domains according to both automatic and human evaluation metrics.
Mihail Eric, Lakshmi Krishnan, François Charette, Christopher D. Manning
SIGDIAL Conference4
2016 Combining Natural Logic and Shallow Reasoning for Question Answering
abstract
Broad domain question answering is often difficult in the absence of structured knowledge bases, and can benefit from shallow lexical methods (broad coverage) and logical reasoning (high precision).We propose an approach for incorporating both of these signals in a unified framework based on natural logic.We extend the breadth of inferences afforded by natural logic to include relational entailment (e.g., buy → own) and meronymy (e.g., a person born in a city is born the city's country).Furthermore, we train an evaluation function -akin to gameplayingto evaluate the expected truth of candidate premises on the fly.We evaluate our approach on answering multiple choice science questions, achieving strong results on the dataset.
Gabor Angeli, Neha Nayak, Christopher D. Manning
ACL (1)3
2016 A Fast Unified Model for Parsing and Sentence Understanding
abstract
Samuel R. Bowman, Jon Gauthier, Abhinav Rastogi, Raghav Gupta, Christopher D. Manning, Christopher Potts. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016.
Samuel R. Bowman, Jon Gauthier, Abhinav Rastogi, Christopher D. Manning, Christopher Potts
ACL (1)5
2016 A Thorough Examination of the CNN/Daily Mail Reading Comprehension Task
abstract
Enabling a computer to understand a document so that it can answer comprehension questions is a central, yet unsolved goal of NLP.A key factor impeding its solution by machine learned systems is the limited availability of human-annotated data.Hermann et al. (2015) seek to solve this problem by creating over a million training examples by pairing CNN and Daily Mail news articles with their summarized bullet points, and show that a neural network can then be trained to give good performance on this task.In this paper, we conduct a thorough examination of this new reading comprehension task.Our primary aim is to understand what depth of language understanding is required to do well on this task.We approach this from one side by doing a careful hand-analysis of a small subset of the problems and from the other by showing that simple, carefully designed systems can obtain accuracies of 72.4% and 75.8% on these two datasets, exceeding current state-of-the-art results by over 5% and approaching what we believe is the ceiling for performance on this task.1
Danqi Chen 0001, Jason Bolton, Christopher D. Manning
ACL (1)3
2016 Improving Coreference Resolution by Learning Entity-Level Distributed Representations
abstract
A long-standing challenge in coreference resolution has been the incorporation of entity-level information - features defined over clusters of mentions instead of mention pairs. We present a neural network based coreference system that produces high-dimensional vector representations for pairs of coreference clusters. Using these representations, our system learns when combining clusters is desirable. We train the system with a learning-to-search algorithm that teaches it which local decisions (cluster merges) will lead to a high-scoring final coreference partition. The system substantially outperforms the current state-of-the-art on the English and Chinese portions of the CoNLL 2012 Shared Task dataset despite using few hand-engineered features.
Kevin Clark, Christopher D. Manning
ACL (1)2
2016 Achieving Open Vocabulary Neural Machine Translation with Hybrid Word-Character Models
abstract
Nearly all previous work on neural machine translation (NMT) has used quite restricted vocabularies, perhaps with a subsequent method to patch in unknown words.This paper presents a novel wordcharacter solution to achieving open vocabulary NMT.We build hybrid systems that translate mostly at the word level and consult the character components for rare words.Our character-level recurrent neural networks compute source word representations and recover unknown target words when needed.The twofold advantage of such a hybrid approach is that it is much faster and easier to train than character-based ones; at the same time, it never produces unknown words as in the case of word-based models.On the WMT'15 English to Czech translation task, this hybrid approach offers an addition boost of +2.1-11.4BLEU points over models that already handle unknown words.Our best system achieves a new state-of-the-art result with 20.7 BLEU score.We demonstrate that our character models can successfully learn to not only generate well-formed words for Czech, a highly-inflected language with a very complex vocabulary, but also build correct representations for English source words.
Minh-Thang Luong, Christopher D. Manning
ACL (1)2
2016 Learning Language Games through Interaction
abstract
We introduce a new language learning setting relevant to building adaptive natural language interfaces.It is inspired by Wittgenstein's language games: a human wishes to accomplish some task (e.g., achieving a certain configuration of blocks), but can only communicate with a computer, who performs the actual actions (e.g., removing all red blocks).The computer initially knows nothing about language and therefore must learn it from scratch through interaction, while the human adapts to the computer's capabilities.We created a game called SHRDLURN in a blocks world and collected interactions from 100 people playing it.First, we analyze the humans' strategies, showing that using compositionality and avoiding synonyms correlates positively with task performance.Second, we compare computer strategies, showing that modeling pragmatics on a semantic parsing model accelerates learning for more strategic players.
Sida I. Wang, Percy Liang, Christopher D. Manning
ACL (1)3
2016 Compression of Neural Machine Translation Models via Pruning
abstract
Neural Machine Translation (NMT), like many other deep learning domains, typically suffers from over-parameterization, resulting in large storage sizes. This paper examines three simple magnitude-based pruning schemes to compress NMT models, namely class-blind, class-uniform, and class-distribution, which differ in terms of how pruning thresholds are computed for the different classes of weights in the NMT architecture. We demonstrate the efficacy of weight pruning as a compression technique for a state-of-the-art NMT system. We show that an NMT model with over 200 million parameters can be pruned by 40% with very little performance loss as measured on the WMT'14 English-German translation task. This sheds light on the distribution of redundancy in the NMT architecture. Our main result is that with retraining, we can recover and even surpass the original performance with an 80%-pruned model.
Abigail See, Minh-Thang Luong, Christopher D. Manning
CoNLL3
2016 Deep Reinforcement Learning for Mention-Ranking Coreference Models
abstract
Coreference resolution systems are typically trained with heuristic loss functions that require careful tuning.In this paper we instead apply reinforcement learning to directly optimize a neural mention-ranking model for coreference evaluation metrics.We experiment with two approaches: the REINFORCE policy gradient algorithm and a rewardrescaled max-margin objective.We find the latter to be more effective, resulting in a significant improvement over the current stateof-the-art on the English and Chinese portions of the CoNLL 2012 Shared Task.
Kevin Clark, Christopher D. Manning
EMNLP2
2016 A comparison of Named-Entity Disambiguation and Word Sense Disambiguation
Angel X. Chang, Valentin I. Spitkovsky, Christopher D. Manning, Eneko Agirre
LREC3
2016 Universal Dependencies v1: A Multilingual Treebank Collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajic 0001, Christopher D. Manning, Ryan T. McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, Daniel Zeman
LREC6
2016 Enhanced English Universal Dependencies: An Improved Representation for Natural Language Understanding Tasks
Sebastian Schuster 0001, Christopher D. Manning
LREC2
2016 Understanding Human Language: Can NLP and Deep Learning Help?
abstract
There is a lot of overlap between the core problems of information retrieval (IR) and natural language processing (NLP). An IR system gains from understanding a user need and from understanding documents, and hence being able to determine whether a document has information that satisfies the user need. Much of NLP is about the same thing: Natural language understanding aims to understand the meaning of questions and documents and meaning relationships. The exciting recent application of deep learning approaches in NLP has brought new tools for effectively understanding language semantics. In principle, there should be a lot of synergy, though in practice the concerns of IR on large systems and macro-scale understanding have tended to contrast with the emphasis in NLP on language structure and micro-scale understanding. My talk will emphasize the two topics of how NLP can contribute to understanding textual relationships and how deep learning approaches substantially aid in this goal. One basic -- and very successful tool -- has been the new generation of distributed word representations: neural word embeddings. However, beyond just word meanings, we need to understand how to compose the meanings of larger pieces of text. Two requirements for that are good ways to understand the structure of human language utterances and ways to compose their meanings. Deep learning methods can help for both tasks. Finally, we need to understand relationships between pieces of text, to be able to do tasks such as Natural Language Inference (or Recognizing Textual Entailment) and Question Answering, and I will look at some of our recent work in these areas, both with and without the help of neural networks
Christopher D. Manning
SIGIR1
2015 Leveraging Linguistic Structure For Open Domain Information Extraction
abstract
Gabor Angeli, Melvin Jose Johnson Premkumar, Christopher D. Manning. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Gabor Angeli, Melvin Jose Johnson Premkumar, Christopher D. Manning
ACL (1)3
2015 Text to 3D Scene Generation with Rich Lexical Grounding
abstract
Angel Chang, Will Monroe, Manolis Savva, Christopher Potts, Christopher D. Manning. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Angel X. Chang, Will Monroe, Manolis Savva, Christopher Potts, Christopher D. Manning
ACL (1)5
2015 Entity-Centric Coreference Resolution with Model Stacking
abstract
Kevin Clark, Christopher D. Manning. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Kevin Clark, Christopher D. Manning
ACL (1)2
2015 Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks
abstract
Kai Sheng Tai, Richard Socher, Christopher D. Manning. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Kai Sheng Tai, Richard Socher, Christopher D. Manning
ACL (1)3
2015 Robust Subgraph Generation Improves Abstract Meaning Representation Parsing
abstract
Keenon Werling, Gabor Angeli, Christopher D. Manning. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Keenon Werling, Gabor Angeli, Christopher D. Manning
ACL (1)3
2015 Deep Neural Language Models for Machine Translation
abstract
Neural language models (NLMs) have been able to improve machine translation (MT) thanks to their ability to generalize well to long contexts.Despite recent successes of deep neural networks in speech and vision, the general practice in MT is to incorporate NLMs with only one or two hidden layers and there have not been clear results on whether having more layers helps.In this paper, we demonstrate that deep NLMs with three or four layers outperform those with fewer layers in terms of both the perplexity and the translation quality.We combine various techniques to successfully train deep NLMs that jointly condition on both the source and target contexts.When reranking nbest lists of a strong web-forum baseline, our deep models yield an average boost of 0.5 TER / 0.5 BLEU points compared to using a shallow NLM.Additionally, we adapt our models to a new sms-chat domain and obtain a similar gain of 1.0 TER / 0.5 BLEU points. 1
Thang Luong, Michael Kayser, Christopher D. Manning
CoNLL3
2015 Forum77: An Analysis of an Online Health Forum Dedicated to Addiction Recovery
abstract
Prescription drug abuse is a pressing public health issue, and people who misuse prescription drugs are turning to online forums for help. Are such forums effective? We analyze the process of opioid withdrawal, recovery and relapse on Forum77, MedHelp.org's online health forum for substance abuse recovery. Applying Prochashka's Transtheoretical Model for behavior change, we develop a taxonomy describing phases of addiction expressed by Forum77 members. We examine activity and linguistic features across the phases USING, WITHDRAWING and RECOVERING. We train statistical classifiers to identify addiction phase, relapse and whether a user was RECOVERING at the time of her last post. Applying our classifiers to 2,848 users, we find that while almost 50% relapse, the prognosis for ending in RECOVERING is favorable. Supplementing our results with users' own accounts of their experiences, we discuss Forum77's efficacy and shortcomings, and implications for future technologies.
Diana L. MacLean, Sonal Gupta, Anna Lembke, Christopher D. Manning, Jeffrey Heer
CSCW4
2015 A large annotated corpus for learning natural language inference
abstract
Understanding entailment and contradiction is fundamental to understanding natural language, and inference about entailment and contradiction is a valuable testing ground for the development of semantic representations.However, machine learning research in this area has been dramatically limited by the lack of large-scale resources.To address this, we introduce the Stanford Natural Language Inference corpus, a new, freely available collection of labeled sentence pairs, written by humans doing a novel grounded task based on image captioning.At 570K pairs, it is two orders of magnitude larger than all other resources of its type.This increase in scale allows lexicalized classifiers to outperform some sophisticated existing entailment models, and it allows a neural network-based model to perform competitively on natural language inference benchmarks for the first time.
Samuel R. Bowman, Gabor Angeli, Christopher Potts, Christopher D. Manning
EMNLP4
2015 Effective Approaches to Attention-based Neural Machine Translation
abstract
An attentional mechanism has lately been used to improve neural machine translation (NMT) by selectively focusing on parts of the source sentence during translation.However, there has been little work exploring useful architectures for attention-based NMT.This paper examines two simple and effective classes of attentional mechanism: a global approach which always attends to all source words and a local one that only looks at a subset of source words at a time.We demonstrate the effectiveness of both approaches on the WMT translation tasks between English and German in both directions.With local attention, we achieve a significant gain of 5.0 BLEU points over non-attentional systems that already incorporate known techniques such as dropout.Our ensemble model using different attention architectures yields a new state-of-the-art result in the WMT'15 English to German translation task with 25.9 BLEU points, an improvement of 1.0 BLEU points over the existing best system backed by NMT and an n-gram reranker.1
Thang Luong, Christopher D. Manning
EMNLP3
2015 Distributed Representations of Words to Guide Bootstrapped Entity Classifiers
abstract
Bootstrapped classifiers iteratively generalize from a few seed examples or prototypes to other examples of target labels.However, sparseness of language and limited supervision make the task difficult.We address this problem by using distributed vector representations of words to aid the generalization.We use the word vectors to expand entity sets used for training classifiers in a bootstrapped pattern-based entity extraction system.Our experiments show that the classifiers trained with the expanded sets perform better on entity extraction from four online forums, with 30% F 1 improvement on one forum.The results suggest that distributed representations can provide good directions for generalization in a bootstrapping system.
Sonal Gupta, Christopher D. Manning
HLT-NAACL2
2015 On-the-Job Learning with Bayesian Decision Theory
abstract
Our goal is to deploy a high-accuracy system starting with zero training examples. We consider an “on-the-job” setting, where as inputs arrive, we use real-time crowdsourcing to resolve uncertainty where needed and output our prediction when confident. As the model improves over time, the reliance on crowdsourcing queries decreases. We cast our setting as a stochastic game based on Bayesian decision theory, which allows us to balance latency, cost, and accuracy objectives in a principled way. Computing the optimal policy is intractable, so we develop an approximation based on Monte Carlo Tree Search. We tested our approach on three datasets-- named-entity recognition, sentiment classification, and image classification. On the NER task we obtained more than an order of magnitude reduction in cost compared to full human annotation, while boosting performance relative to the expert provided labels. We also achieve a 8% F1 improvement over having a single human label the whole set, and a 28% F1 improvement over online learning.
Keenon Werling, Arun Tejasvi Chaganty, Percy Liang, Christopher D. Manning
NIPS4
2015 Computational Linguistics and Deep Learning
abstract
Deep Learning waves have lapped at the shores of computational linguistics for several years now, but 2015 seems like the year when the full force of the tsunami hit the major Natural Language Processing (NLP) conferences. However, some pundits are predicting that the final damage will be even worse. Accompanying ICML 2015 in Lille, France, there was another, almost as big, event: the 2015 Deep Learning Workshop. The workshop ended with a panel discussion, and at it, Neil Lawrence said, “NLP is kind of like a rabbit in the headlights of the Deep Learning machine, waiting to be flattened.” Now that is a remark that the computational linguistics community has to take seriously! Is it the end of the road for us? Where are these predictions of steam-rollering coming from?At the June 2015 opening of the Facebook AI Research Lab in Paris, its director Yann LeCun said: “The next big step for Deep Learning is natural language understanding, which aims to give machines the power to understand not just individual words but entire sentences and paragraphs.”1 In a November 2014 Reddit AMA (Ask Me Anything), Geoff Hinton said, “I think that the most exciting areas over the next five years will be really understanding text and videos. I will be disappointed if in five years' time we do not have something that can watch a YouTube video and tell a story about what happened. In a few years time we will put [Deep Learning] on a chip that fits into someone's ear and have an English-decoding chip that's just like a real Babel fish.”2 And Yoshua Bengio, the third giant of modern Deep Learning, has also increasingly oriented his group's research toward language, including recent exciting new developments in neural machine translation systems. It's not just Deep Learning researchers. When leading machine learning researcher Michael Jordan was asked at a September 2014 AMA, “If you got a billion dollars to spend on a huge research project that you get to lead, what would you like to do?”, he answered: “I'd use the billion dollars to build a NASA-size program focusing on natural language processing, in all of its glory (semantics, pragmatics, etc.).” He went on: “Intellectually I think that NLP is fascinating, allowing us to focus on highly structured inference problems, on issues that go to the core of ‘what is thought’ but remain eminently practical, and on a technology that surely would make the world a better place.” Well, that sounds very nice! So, should computational linguistics researchers be afraid? I'd argue, no. To return to the Hitchhiker's Guide to the Galaxy theme that Geoff Hinton introduced, we need to turn the book over and look at the back cover, which says in large, friendly letters: “Don't panic.”There is no doubt that Deep Learning has ushered in amazing technological advances in the last few years. I won't give an extensive rundown of successes, but here is one example. A recent Google blog post told about Neon, the new transcription system for Google Voice.3 After admitting that in the past Google Voice voicemail transcriptions often weren't fully intelligible, the post explained the development of Neon, an improved voicemail system that delivers more accurate transcriptions, like this: “Using a (deep breath) long short-term memory deep recurrent neural network (whew!), we cut our transcription errors by 49%.” Do we not all dream of developing a new approach to a problem which halves the error rate of the previously state-of-the-art system?Michael Jordan, in his AMA, gave two reasons why he wasn't convinced that Deep Learning would solve NLP: “Although current deep learning research tends to claim to encompass NLP, I'm (1) much less convinced about the strength of the results, compared to the results in, say, vision; (2) much less convinced in the case of NLP than, say, vision, the way to go is to couple huge amounts of data with black-box learning architectures.”4Jordan is certainly right about his first point: So far, problems in higher-level language processing have not seen the dramatic error rate reductions from deep learning that have been seen in speech recognition and in object recognition in vision. Although there have been gains from deep learning approaches, they have been more modest than sudden 25% or 50% error reductions. It could easily turn out that this remains the case. The really dramatic gains may only have been possible on true signal processing tasks. On the other hand, I'm much less convinced by his second argument. However, I do have my own two reasons why NLP need not worry about deep learning: (1) It just has to be wonderful for our field for the smartest and most influential people in machine learning to be saying that NLP is the problem area to focus on; and (2) Our field is the domain science of language technology; it's not about the best method of machine learning—the central issue remains the domain problems. The domain problems will not go away. Joseph Reisinger wrote on his blog: “I get pitched regularly by startups doing ‘generic machine learning’ which is, in all honesty, a pretty ridiculous idea. Machine learning is not undifferentiated heavy lifting, it's not commoditizable like EC2, and closer to design than coding.”5 From this perspective, it is people in linguistics, people in NLP, who are the designers. Recently at ACL conferences, there has been an over-focus on numbers, on beating the state of the art. Call it playing the Kaggle game. More of the field's effort should go into problems, approaches, and architectures. Recently, one thing that I've been devoting a lot of time to—together with many other collaborators—is the development of Universal Dependencies.6 The goal is to develop a common syntactic dependency representation and POS and feature label sets that can be used with reasonable linguistic fidelity and human usability across all human languages. That's just one example; there are many other design efforts underway in our field. One other current example is the idea of Abstract Meaning Representation.7Where has Deep Learning helped NLP? The gains so far have not so much been from true Deep Learning (use of a hierarchy of more abstract representations to promote generalization) as from the use of distributed word representations—through the use of real-valued vector representations of words and concepts. Having a dense, multi-dimensional representation of similarity between all words is incredibly useful in NLP, but not only in NLP. Indeed, the importance of distributed representations evokes the “Parallel Distributed Processing” mantra of the earlier surge of neural network methods, which had a much more cognitive-science directed focus (Rumelhart and McClelland 1986). It can better explain human-like generalization, but also, from an engineering perspective, the use of small dimensionality and dense vectors for words allows us to model large contexts, leading to greatly improved language models. Especially seen from this new perspective, the exponentially greater sparsity that comes from increasing the order of traditional word n-gram models seems conceptually bankrupt.I do believe that the idea of deep models will also prove useful. The sharing that occurs within deep representations can theoretically give an exponential representational advantage, and, in practice, offers improved learning systems. The general approach to building Deep Learning systems is compelling and powerful: The researcher defines a model architecture and a top-level loss function and then both the parameters and the representations of the model self-organize so as to minimize this loss, in an end-to-end learning framework. We are starting to see the power of such deep systems in recent work in neural machine translation (Sutskever, Vinyals, and Le 2014; Luong et al. 2015).Finally, I have been an advocate for focusing more on compositionality in models, for language in particular, and for artificial intelligence in general. Intelligence requires being able to understand bigger things from knowing about smaller parts. In particular for language, understanding novel and complex sentences crucially depends on being able to construct their meaning compositionally from smaller parts—words and multi-word expressions—of which they are constituted. Recently, there have been many, many papers showing how systems can be improved by using distributed word representations from “deep learning” approaches, such as word2vec (Mikolov et al. 2013) or GloVe (Pennington, Socher, and Manning 2014). However, this is not actually building Deep Learning models, and I hope in the future that more people focus on the strongly linguistic question of whether we can build meaning composition functions in Deep Learning systems.I encourage people to not get into the rut of doing no more than using word vectors to make performance go up a couple of percent. Even more strongly, I would like to suggest that we might return instead to some of the interesting linguistic and cognitive issues that motivated noncategorical representations and neural network approaches.One example of noncategorical phenomena in language is the POS of words in the gerund V-ing form, such as driving. This form is classically described as ambiguous between a verbal form and a nominal gerund. In fact, however, the situation is more complex, as V-ing forms can appear in any of the four core categories of Chomsky (1970):What is even more interesting is that there is evidence that there is not just an ambiguity but mixed noun–verb status. For example, a classic linguistic text for being a noun is appearing with a determiner, while a classic linguistic test for being a verb is taking a direct object. However, it is well known that the gerund nominalization can do both of these things at once: (1) The not observing this rule is that which the world has blamed in our satorist. (Dryden, Essay Dramatick Poesy, 1684, page 310)(2) The only mental provision she was making for the evening of life, was the collecting and transcribing all the riddles of every sort that she could meet with. (Jane Austen, Emma, 1816)(3) The difficulty is in the getting the gold into Erewhon. (Sam Butler, Erewhon Revisited, 1902) This is oftentimes analyzed by some sort of category-change operation within the levels of a phrase-structure tree, but there is good evidence that this is in fact a case of noncategorical behavior in language.Indeed, this construction was used early on as an example of a “squish” by Ross (1972). Diachronically, the V-ing form shows a history of increasing verbalization, but in many periods it shows a notably non-discrete status. For example, we find clearly graded judgments in this domain: (4) Tom's winning the election was a big upset.(5) ?This teasing John all the time has got to stop.(6) ?There is no marking exams on Fridays.(7) *The cessation hostilities was unexpected. Various combinations of determiner and verb object do not sound so good, but still much better than trying to put a direct object after a nominalization via a derivational morpheme such as -ation. Houston (1985, page 320) shows that assignment of V-ing forms to a discrete part-of-speech classification is less successful (in a predictive sense) than a continuum in explaining the spoken alternation between -ing vs. -in', suggesting that “grammatical categories exist along a continuum which does not exhibit sharp boundaries between the categories.”A different, interesting example was explored by one of my graduate school classmates, Whitney Tabor. Tabor (1994) looked at the use of kind of and sort of, an example that I then used in the introductory chapter of my 1999 textbook (Manning and Schütze, 1999). The nouns kind or sort can head an NP or be used as a hedging adverbial modifier: (8) [That kind [of knife]] isn't used much.(9) We are [kind of] hungry. The interesting thing is that there is a path of reanalysis through ambiguous forms, such as the following pair, which suggests how one form emerged from the other. (10) [a [kind [of dense rock]]](11) [a [[kind of] dense] rock]Tabor (1994) discusses how Old English has kind but few or no uses of kind of. Beginning in Middle English, ambiguous contexts, which provide a breeding ground for the reanalysis, start to appear (the 1570 example in Example (13)), and then, later, examples that are unambiguously the hedging modifier appear (the 1830 example in Example (14)): (12) A nette sent in to the see, and of alle kind of fishis gedrynge (Wyclif, 1382)(13) Their finest and best, is a kind of course red cloth (True Report, 1570)(14) I was kind of provoked at the way you came up (Mass. Spy, 1830) This is history not synchrony. Presumably kids today learn the softener use of kind/sort of first. Did the reader notice an example of it in the quote in my first paragraph? (15) NLP is kind of like a rabbit in the headlights of the deep learning machine (Neil Lawrence, DL workshop panel, 2015) Whitney Tabor modeled this evolution with a small, but already deep, recurrent neural network—one with two hidden layers. He did that in 1994, taking advantage of the opportunity to work with Dave Rumelhart at Stanford.Just recently, there has started to be some new work harnessing the power of distributed representations for modeling and explaining linguistic variation and change. Sagi, Kaufmann, and Clark (2011)—actually using the more traditional method of Latent Semantic Analysis to generate distributed word representations—show how distributed representations can capture a semantic change: the broadening and narrowing of reference over time. They look at examples such as how in Old English deer was any animal, whereas in Middle and Modern English it applies to one clear animal family. The words dog and hound have swapped: In Middle English, hound was used for any kind of canine, while now it is used for a particular sub-kind, whereas the reverse is true for dog.Kulkarni et al. (2015) use neural word embeddings to model the shift in meaning of words such as gay over the last century (exploiting the online Google Books Ngrams corpus). At a recent ACL workshop, Kim et al. (2014) use a similar approach—using word2vec—to look at recent changes in the meaning of words. For example, in Figure 1, they show how around 2000, the meaning of the word cell changed rapidly from being close in meaning to closet and dungeon to being close in meaning to phone and cordless. The meaning of a word in this context is the average over the meanings of all senses of a word, weighted by their frequency of use.These more scientific uses of distributed representations and Deep Learning for modeling phenomena characterize the previous boom in neural networks. There has been a bit of a kerfuffle online lately about citing and crediting work in Deep Learning, and from that perspective, it seems to me that the two people who scarcely get mentioned any more are Dave Rumelhart and Jay McClelland. Starting from the Parallel Distributed Processing Research Group in San Diego, their research program was aimed at a clearly more scientific and cognitive study of neural networks.Now, there are indeed some good questions about the adequacy of neural network approaches for rule-governed linguistic behavior. Old timers in our community should remember that arguing against the adequacy of neural networks for rule-governed linguistic behavior was the foundation for the rise to fame of Steve Pinker—and the foundation of the career of about six of his graduate students. It would take too much space to go through the issues here, but in the end, I think it was a productive debate. It led to a vast amount of work by Paul Smolensky on how basically categorical systems can emerge and be represented in a neural substrate (Smolensky and Legendre 2006). Indeed, Paul Smolensky arguably went too far down the rabbit hole, devoting a large part of his career to developing a new categorical model of phonology, Optimality Theory (Prince and Smolensky 2004). There is a rich body of earlier scientific work that has been neglected. It would be good to return some emphasis within NLP to cognitive and scientific investigation of language rather than almost exclusively using an engineering model of research.Overall, I think we should feel excited and glad to live in a time when Natural Language Processing is seen as so central to both the further development of machine learning and industry application problems. The future is bright. However, I would encourage everyone to think about problems, architectures, cognitive science, and the details of human language, how it is learned, processed, and how it changes, rather than just chasing state-of-the-art numbers on a benchmark task.This Last Words contribution covers part of my 2015 ACL Presidential Address. Thanks to Paola Merlo for suggesting writing it up for publication.
Christopher D. Manning
Comput. Linguistics1
2014 TransPhoner: automated mnemonic keyword generation
abstract
We present TransPhoner: a system that generates keywords for a variety of scenarios including vocabulary learning, phonetic transliteration, and creative word plays. We select effective keywords by considering phonetic, orthographic and semantic word similarity, and word concept imageability. We show that keywords provided by TransPhoner improve learner performance in an online vocabulary learning study, with the improvement being more pronounced for harder words. Participants rated TransPhoner keywords as more helpful than a random keyword baseline, and almost as helpful as manually selected keywords. Comments also indicated higher engagement in the learning task, and more desire to continue learning. We demonstrate additional applications to tasks such as pure phonetic transliteration, generation of mnemonics for complex vocabulary, and topic-based transformation of song lyrics.
Manolis Savva, Angel X. Chang, Christopher D. Manning, Pat Hanrahan
CHI3
2014 Improved Pattern Learning for Bootstrapped Entity Extraction
abstract
Bootstrapped pattern learning for entity extraction usually starts with seed entities and iteratively learns patterns and entities from unlabeled text. Patterns are scored by their ability to extract more positive en-tities and less negative entities. A prob-lem is that due to the lack of labeled data, unlabeled entities are either assumed to be negative or are ignored by the existing pat-tern scoring measures. In this paper, we improve pattern scoring by predicting the labels of unlabeled entities. We use var-ious unsupervised features based on con-trasting domain-specific and general text, and exploiting distributional similarity and edit distances to learned entities. Our system outperforms existing pattern scor-ing algorithms for extracting drug-and-treatment entities from four medical fo-rums. 1
Sonal Gupta, Christopher D. Manning
CoNLL2
2014 NaturalLI: Natural Logic Inference for Common Sense Reasoning
abstract
Common-sense reasoning is important for AI applications, both in NLP and many vision and robotics tasks.We propose NaturalLI: a Natural Logic inference system for inferring common sense facts -for instance, that cats have tails or tomatoes are round -from a very large database of known facts.In addition to being able to provide strictly valid derivations, the system is also able to produce derivations which are only likely valid, accompanied by an associated confidence.We both show that our system is able to capture strict Natural Logic inferences on the Fra-CaS test suite, and demonstrate its ability to predict common sense facts with 49% recall and 91% precision.
Gabor Angeli, Christopher D. Manning
EMNLP2
2014 Combining Distant and Partial Supervision for Relation Extraction
abstract
Broad-coverage relation extraction either requires expensive supervised training data, or suffers from drawbacks inherent to distant supervision.We present an approach for providing partial supervision to a distantly supervised relation extractor using a small number of carefully selected examples.We compare against established active learning criteria and propose a novel criterion to sample examples which are both uncertain and representative.In this way, we combine the benefits of fine-grained supervision for difficult examples with the coverage of a large distantly supervised corpus.Our approach gives a substantial increase of 3.9% endto-end F 1 on the 2013 KBP Slot Filling evaluation, yielding a net F 1 of 37.7%.
Gabor Angeli, Julie Tibshirani, Jean Wu, Christopher D. Manning
EMNLP4
2014 Modeling Biological Processes for Reading Comprehension
abstract
Jonathan Berant, Vivek Srikumar, Pei-Chun Chen, Abby Vander Linden, Brittany Harding, Brad Huang, Peter Clark, Christopher D. Manning. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014.
Jonathan Berant, Vivek Srikumar, Pei-Chun Chen, Abby Vander Linden, Brittany Harding, Brad Huang, Peter Clark, Christopher D. Manning
EMNLP8
2014 Learning Spatial Knowledge for Text to 3D Scene Generation
abstract
We address the grounding of natural language to concrete spatial constraints, and inference of implicit pragmatics in 3D environments.We apply our approach to the task of text-to-3D scene generation.We present a representation for common sense spatial knowledge and an approach to extract it from 3D scene data.In text-to-3D scene generation, a user provides as input natural language text from which we extract explicit constraints on the objects that should appear in the scene.The main innovation of this work is to show how to augment these explicit constraints with learned spatial knowledge to infer missing objects and likely layouts for the objects in the scene.We demonstrate that spatial knowledge is useful for interpreting natural language and show examples of learned knowledge and generated 3D scenes.
Angel X. Chang, Manolis Savva, Christopher D. Manning
EMNLP3
2014 A Fast and Accurate Dependency Parser using Neural Networks
abstract
Almost all current dependency parsers classify based on millions of sparse indi-cator features. Not only do these features generalize poorly, but the cost of feature computation restricts parsing speed signif-icantly. In this work, we propose a novel way of learning a neural network classifier for use in a greedy, transition-based depen-dency parser. Because this classifier learns and uses just a small number of dense fea-tures, it can work very fast, while achiev-ing an about 2 % improvement in unla-beled and labeled attachment scores on both English and Chinese datasets. Con-cretely, our parser is able to parse more than 1000 sentences per second at 92.2% unlabeled attachment score on the English Penn Treebank. 1
Danqi Chen 0001, Christopher D. Manning
EMNLP2
2014 Human Effort and Machine Learnability in Computer Aided Translation
abstract
Spence Green, Sida I. Wang, Jason Chuang, Jeffrey Heer, Sebastian Schuster, Christopher D. Manning. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014.
Spence Green, Sida I. Wang, Jason Chuang, Jeffrey Heer, Sebastian Schuster 0001, Christopher D. Manning
EMNLP6
2014 Glove: Global Vectors for Word Representation
abstract
Recent methods for learning vector space representations of words have succeeded in capturing fine-grained semantic and syntactic regularities using vector arith-metic, but the origin of these regularities has remained opaque. We analyze and make explicit the model properties needed for such regularities to emerge in word vectors. The result is a new global log-bilinear regression model that combines the advantages of the two major model families in the literature: global matrix factorization and local context window methods. Our model efficiently leverages statistical information by training only on the nonzero elements in a word-word co-occurrence matrix, rather than on the en-tire sparse matrix or on individual context windows in a large corpus. The model pro-duces a vector space with meaningful sub-structure, as evidenced by its performance of 75 % on a recent word analogy task. It also outperforms related models on simi-larity tasks and named entity recognition. 1
Jeffrey Pennington, Richard Socher, Christopher D. Manning
EMNLP3
2014 Universal Stanford dependencies: A cross-linguistic typology
Marie-Catherine de Marneffe, Timothy Dozat, Natalia Silveira, Katri Haverinen, Filip Ginter, Joakim Nivre, Christopher D. Manning
LREC7
2014 Event Extraction Using Distant Supervision
Kevin Reschke, Martin Jankowiak, Mihai Surdeanu, Christopher D. Manning, Daniel Jurafsky
LREC4
2014 A Gold Standard Dependency Corpus for English
Natalia Silveira, Timothy Dozat, Marie-Catherine de Marneffe, Samuel R. Bowman, Miriam Connor, John Bauer, Christopher D. Manning
LREC7
2014 Simple MAP Inference via Low-Rank Relaxations
Roy Frostig, Sida I. Wang, Percy Liang, Christopher D. Manning
NIPS4
2014 Global Belief Recursive Neural Networks
Romain Paulus, Richard Socher, Christopher D. Manning
NIPS3
2014 Learning Distributed Representations for Structured Output Prediction
Vivek Srikumar, Christopher D. Manning
NIPS2
2014 Predictive translation memory: a mixed-initiative system for human language translation
abstract
The standard approach to computer-aided language translation is post-editing: a machine generates a single translation that a human translator corrects. Recent studies have shown this simple technique to be surprisingly effective, yet it underutilizes the complementary strengths of precision-oriented humans and recall-oriented machines. We present Predictive Translation Memory, an interactive, mixed-initiative system for human language translation. Translators build translations incrementally by considering machine suggestions that update according to the user's current partial translation. In a large-scale study, we find that professional translators are slightly slower in the interactive mode yet produce slightly higher quality translations despite significant prior experience with the baseline post-editing condition. Our analysis identifies significant predictors of time and quality, and also characterizes interactive aid usage. Subjects entered over 99% of characters via interactive aids, a significantly higher fraction than that shown in previous work.
Spence Green, Jason Chuang, Jeffrey Heer, Christopher D. Manning
UIST4
2014 Research and applications: Induced lexico-syntactic patterns improve information extraction from online medical forums
abstract
OBJECTIVE: To reliably extract two entity types, symptoms and conditions (SCs), and drugs and treatments (DTs), from patient-authored text (PAT) by learning lexico-syntactic patterns from data annotated with seed dictionaries. BACKGROUND AND SIGNIFICANCE: Despite the increasing quantity of PAT (eg, online discussion threads), tools for identifying medical entities in PAT are limited. When applied to PAT, existing tools either fail to identify specific entity types or perform poorly. Identification of SC and DT terms in PAT would enable exploration of efficacy and side effects for not only pharmaceutical drugs, but also for home remedies and components of daily care. MATERIALS AND METHODS: We use SC and DT term dictionaries compiled from online sources to label several discussion forums from MedHelp (http://www.medhelp.org). We then iteratively induce lexico-syntactic patterns corresponding strongly to each entity type to extract new SC and DT terms. RESULTS: Our system is able to extract symptom descriptions and treatments absent from our original dictionaries, such as 'LADA', 'stabbing pain', and 'cinnamon pills'. Our system extracts DT terms with 58-70% F1 score and SC terms with 66-76% F1 score on two forums from MedHelp. We show improvements over MetaMap, OBA, a conditional random field-based classifier, and a previous pattern learning approach. CONCLUSIONS: Our entity extractor based on lexico-syntactic patterns is a successful and preferable technique for identifying specific entity types in PAT. To the best of our knowledge, this is the first paper to extract SC and DT entities from PAT. We exhibit learning of informal terms often used in PAT but missing from typical dictionaries.
Sonal Gupta, Diana L. MacLean, Jeffrey Heer, Christopher D. Manning
J. Am. Medical Informatics Assoc.4
2014 Grounded Compositional Semantics for Finding and Describing Images with Sentences
abstract
Previous work on Recursive Neural Networks (RNNs) shows that these models can produce compositional feature vectors for accurately representing and classifying sentences or images. However, the sentence vectors of previous models cannot accurately represent visually grounded meaning. We introduce the DT-RNN model which uses dependency trees to embed sentences into a vector space in order to retrieve images that are described by those sentences. Unlike previous RNN-based models which use constituency trees, DT-RNNs naturally focus on the action and agents in a sentence. They are better able to abstract from the details of word order and syntactic expression. DT-RNNs outperform other recursive and recurrent neural networks, kernelized CCA and a bag-of-words baseline on the tasks of finding an image that fits a sentence description and vice versa. They also give more similar representations to sentences that describe the same image.
Richard Socher, Andrej Karpathy, Quoc V. Le, Christopher D. Manning, Andrew Y. Ng
Trans. Assoc. Comput. Linguistics4
2014 Cross-lingual Projected Expectation Regularization for Weakly Supervised Learning
abstract
We consider a multilingual weakly supervised learning scenario where knowledge from annotated corpora in a resource-rich language is transferred via bitext to guide the learning in other languages. Past approaches project labels across bitext and use them as features or gold labels for training. We propose a new method that projects model expectations rather than labels, which facilities transfer of model uncertainty across language boundaries. We encode expectations as constraints and train a discriminative CRF model using Generalized Expectation Criteria (Mann and McCallum, 2010). Evaluated on standard Chinese-English and German-English NER datasets, our method demonstrates F1 scores of 64% and 60% when no labeled data is used. Attaining the same accuracy with supervised CRFs requires 12k and 1.5k labeled sentences. Furthermore, when combined with labeled examples, our method yields significant improvements over state-of-the-art supervised methods, achieving best reported numbers to date on Chinese OntoNotes and German CoNLL-03 datasets.
Mengqiu Wang, Christopher D. Manning
Trans. Assoc. Comput. Linguistics2
2013 Effective Bilingual Constraints for Semi-Supervised Learning of Named Entity Recognizers
abstract
Most semi-supervised methods in Natural Language Processing capitalize on unannotated resources in a single language; however, information can be gained from using parallel resources in more than one language, since translations of the same utterance in different languages can help to disambiguate each other. We demonstrate a method that makes effective use of vast amounts of bilingual text (a.k.a. bitext) to improve monolingual systems. We propose a factored probabilistic sequence model that encourages both crosslanguage and intra-document consistency. A simple Gibbs sampling algorithm is introduced for performing approximate inference. Experiments on English-Chinese Named Entity Recognition (NER) using the OntoNotes dataset demonstrate that our method is significantly more accurate than state-ofthe- art monolingual CRF models in a bilingual test setting. Our model also improves on previous work by Burkett et al. (2010), achieving a relative error reduction of 10.8% and 4.5% in Chinese and English, respectively. Furthermore, by annotating a moderate amount of unlabeled bi-text with our bilingual model, and using the tagged data for uptraining, we achieve a 9.2% error reduction in Chinese over the state-ofthe- art Stanford monolingual NER system.
Mengqiu Wang, Wanxiang Che, Christopher D. Manning
AAAI3
2013 Fast and Adaptive Online Training of Feature-Rich Translation Models
Spence Green, Sida I. Wang, Daniel M. Cer, Christopher D. Manning
ACL (1)4
2013 Parsing with Compositional Vector Grammars
Richard Socher, John Bauer, Christopher D. Manning, Andrew Y. Ng
ACL (1)3
2013 Joint Word Alignment and Bilingual Named Entity Recognition Using Dual Decomposition
Mengqiu Wang, Wanxiang Che, Christopher D. Manning
ACL (1)3
2013 The efficacy of human post-editing for language translation
abstract
Language translation is slow and expensive, so various forms of machine assistance have been devised. Automatic machine translation systems process text quickly and cheaply, but with quality far below that of skilled human translators. To bridge this quality gap, the translation industry has investigated post-editing, or the manual correction of machine output. We present the first rigorous, controlled analysis of post-editing and find that post-editing leads to reduced time and, surprisingly, improved quality for three diverse language pairs (English to Arabic, French, and German). Our statistical models and visualizations of experimental data indicate that some simple predictors (like source text part of speech counts) predict translation time, and that post-editing results in very different interaction patterns. From these results we distill implications for the design of new language translation interfaces.
Spence Green, Jeffrey Heer, Christopher D. Manning
CHI3
2013 Philosophers are Mortal: Inferring the Truth of Unseen Facts
Gabor Angeli, Christopher D. Manning
CoNLL2
2013 Better Word Representations with Recursive Neural Networks for Morphology
Thang Luong, Richard Socher, Christopher D. Manning
CoNLL3
2013 Learning Biological Processes with Global Constraints
abstract
Aju Thalappillil Scaria, Jonathan Berant, Mengqiu Wang, Peter Clark, Justin Lewis, Brittany Harding, Christopher D. Manning. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. 2013.
Aju Thalappillil Scaria, Jonathan Berant, Mengqiu Wang, Peter Clark, Justin Lewis, Brittany Harding, Christopher D. Manning
EMNLP7
2013 Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
abstract
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, Christopher Potts. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. 2013.
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, Christopher Potts
EMNLP5
2013 Feature Noising for Log-Linear Structured Prediction
abstract
NLP models have many and sparse features, and regularization is key for balancing model overfitting versus underfitting.A recently repopularized form of regularization is to generate fake training data by repeatedly adding noise to real data.We reinterpret this noising as an explicit regularizer, and approximate it with a second-order formula that can be used during training without actually generating fake data.We show how to apply this method to structured prediction using multinomial logistic regression and linear-chain CRFs.We tackle the key challenge of developing a dynamic program to compute the gradient of the regularizer efficiently.The regularizer is a sum over inputs, so we can estimate it more accurately via a semi-supervised or transductive extension.Applied to text classification and NER, our method provides a >1% absolute performance gain over use of standard L 2 regularization.
Sida I. Wang, Mengqiu Wang, Stefan Wager, Percy Liang, Christopher D. Manning
EMNLP5
2013 Bilingual Word Embeddings for Phrase-Based Machine Translation
abstract
We introduce bilingual word embeddings: semantic embeddings associated across two languages in the context of neural language models.We propose a method to learn bilingual embeddings from a large unlabeled corpus, while utilizing MT word alignments to constrain translational equivalence.The new embeddings significantly out-perform baselines in word semantic similarity.A single semantic similarity feature induced with bilingual embeddings adds near half a BLEU point to the results of NIST08 Chinese-English machine translation task.
Will Y. Zou, Richard Socher, Daniel M. Cer, Christopher D. Manning
EMNLP4
2013 Topic Model Diagnostics: Assessing Domain Relevance via Topical Alignment
abstract
The use of topic models to analyze domain-specific texts often requires manual validation of the latent topics to ensure they are meaningful. We introduce a framework to support large-scale assessment of topical relevance. We measure the correspondence between a set of latent topics and a set of reference concepts to quantify four types of topical misalignment: junk, fused, missing, and repeated topics. Our analysis compares 10,000 topic model variants to 200 expert-provided domain concepts, and demonstrates how our framework can inform choices of model parameters, inference algorithms, and intrinsic measures of topical quality.
Jason Chuang, Sonal Gupta, Christopher D. Manning, Jeffrey Heer
ICML (3)3
2013 Fast dropout training
abstract
Preventing feature co-adaptation by encouraging independent contributions from different features often improves classification and regression performance. Dropout training (Hinton et al., 2012) does this by randomly dropping out (zeroing) hidden units and input features during training of neural networks. However, repeatedly sampling a random subset of input features makes training much slower. Based on an examination of the implied objective function of dropout training, we show how to do fast dropout training by sampling from or integrating a Gaussian approximation, instead of doing Monte Carlo optimization of this objective. This approximation, justified by the central limit theorem and empirical evidence, gives an order of magnitude speedup and more stability. We show how to do fast dropout training for classification, regression, and multilayer neural networks. Beyond dropout, our technique is extended to integrate out other types of noise and small image transformations.
Sida I. Wang, Christopher D. Manning
ICML (2)2
2013 Learning a Product of Experts with Elitist Lasso
Mengqiu Wang, Christopher D. Manning
IJCNLP2
2013 Effect of Non-linear Deep Architecture in Sequence Labeling
Mengqiu Wang, Christopher D. Manning
IJCNLP2
2013 Named Entity Recognition with Bilingual Constraints
Wanxiang Che, Mengqiu Wang, Christopher D. Manning, Ting Liu 0001
HLT-NAACL3
2013 Deep Learning for NLP (without Magic)
Richard Socher, Christopher D. Manning
HLT-NAACL2
2013 Reasoning With Neural Tensor Networks for Knowledge Base Completion
abstract
A common problem in knowledge representation and related fields is reasoning over a large joint knowledge graph, represented as triples of a relation between two entities. The goal of this paper is to develop a more powerful neural network model suitable for inference over these relationships. Previous models suffer from weak interaction between entities or simple linear projection of the vector space. We address these problems by introducing a neural tensor network (NTN) model which allow the entities and relations to interact multiplicatively. Additionally, we observe that such knowledge base models can be further improved by representing each entity as the average of vectors for the words in the entity name, giving an additional dimension of similarity by which entities can share statistical strength. We assess the model by considering the problem of predicting additional true relations between entities given a partial knowledge base. Our model outperforms previous models and can classify unseen relationships in WordNet and FreeBase with an accuracy of 86.2% and 90.0%, respectively.
Richard Socher, Danqi Chen 0001, Christopher D. Manning, Andrew Y. Ng
NIPS3
2013 Zero-Shot Learning Through Cross-Modal Transfer
abstract
This work introduces a model that can recognize objects in images even if no training data is available for the object class. The only necessary knowledge about unseen categories comes from unsupervised text corpora. Unlike previous zero-shot learning models, which can only differentiate between unseen classes, our model can operate on a mixture of objects, simultaneously obtaining state of the art performance on classes with thousands of training images and reasonable performance on unseen classes. This is achieved by seeing the distributions of words in texts as a semantic space for understanding what objects look like. Our deep learning model does not require any manually defined semantic or visual features for either words or images. Images are mapped to be close to semantic word vectors corresponding to their classes, and the resulting image embeddings can be used to distinguish whether an image is of a seen or unseen class. Then, a separate recognition model can be employed for each type. We demonstrate two strategies, the first gives high accuracy on unseen classes, while the second is conservative in its prediction of novelty and keeps the seen classes' accuracy high.
Richard Socher, Milind Ganjoo, Christopher D. Manning, Andrew Y. Ng
NIPS3
2013 Parsing Models for Identifying Multiword Expressions
abstract
Multiword expressions lie at the syntax/semantics interface and have motivated alternative theories of syntax like Construction Grammar. Until now, however, syntactic analysis and multiword expression identification have been modeled separately in natural language processing. We develop two structured prediction models for joint parsing and multiword expression identification. The first is based on context-free grammars and the second uses tree substitution grammars, a formalism that can store larger syntactic fragments. Our experiments show that both models can identify multiword expressions with much higher accuracy than a state-of-the-art system based on word co-occurrence statistics. We experiment with Arabic and French, which both have pervasive multiword expressions. Relative to English, they also have richer morphology, which induces lexical sparsity in finite corpora. To combat this sparsity, we develop a simple factored lexical representation for the context-free parsing model. Morphological analyses are automatically transformed into rich feature tags that are scored jointly with lexical items. This technique, which we call a factored lexicon, improves both standard parsing and multiword expression identification accuracy.
Spence Green, Marie-Catherine de Marneffe, Christopher D. Manning
Comput. Linguistics3
2012 Improving Word Representations via Global Context and Multiple Word Prototypes
Eric H. Huang, Richard Socher, Christopher D. Manning, Andrew Y. Ng
ACL (1)3
2012 Termite: visualization techniques for assessing textual topic models
abstract
Topic models aid analysis of text corpora by identifying latent topics based on co-occurring words. Real-world deployments of topic models, however, often require intensive expert verification and model refinement. In this paper we present Termite, a visual analysis tool for assessing topic model quality. Termite uses a tabular layout to promote comparison of terms both within and across latent topics. We contribute a novel saliency measure for selecting relevant terms and a seriation algorithm that both reveals clustering structure and promotes the legibility of related terms. In a series of examples, we demonstrate how Termite allows analysts to identify coherent and significant themes.
Jason Chuang, Christopher D. Manning, Jeffrey Heer
AVI2
2012 Interpretation and trust: designing model-driven visualizations for text analysis
abstract
Statistical topic models can help analysts discover patterns in large text corpora by identifying recurring sets of words and enabling exploration by topical concepts. However, understanding and validating the output of these models can itself be a challenging analysis task. In this paper, we offer two design considerations - interpretation and trust - for designing visualizations based on data-driven models. Interpretation refers to the facility with which an analyst makes inferences about the data through the lens of a model abstraction. Trust refers to the actual and perceived accuracy of an analyst's inferences. These considerations derive from our experiences developing the Stanford Dissertation Browser, a tool for exploring over 9,000 Ph.D. theses by topical similarity, and a subsequent review of existing literature. We contribute a novel similarity measure for text collections based on a notion of "word-borrowing" that arose from an iterative design process. Based on our experiences and a literature review, we distill a set of design recommendations and describe how they promote interpretable and trustworthy visual analysis tools.
Jason Chuang, Daniel Ramage, Christopher D. Manning, Jeffrey Heer
CHI3
2012 Learning Constraints for Consistent Timeline Extraction
David McClosky, Christopher D. Manning
EMNLP-CoNLL2
2012 Semantic Compositionality through Recursive Matrix-Vector Spaces
Richard Socher, Brody Huval, Christopher D. Manning, Andrew Y. Ng
EMNLP-CoNLL3
2012 Multi-instance Multi-label Learning for Relation Extraction
Mihai Surdeanu, Julie Tibshirani, Ramesh Nallapati, Christopher D. Manning
EMNLP-CoNLL4
2012 Probabilistic Finite State Machines for Regression-based MT Evaluation
Mengqiu Wang, Christopher D. Manning
EMNLP-CoNLL2
2012 SUTime: A library for recognizing and normalizing time expressions
Angel X. Chang, Christopher D. Manning
LREC2
2012 Parsing Time: Learning to Interpret Time Expressions
Gabor Angeli, Christopher D. Manning, Daniel Jurafsky
HLT-NAACL2
2012 Entity Clustering Across Languages
Spence Green, Nicholas Andrews, Matthew R. Gormley, Mark Dredze, Christopher D. Manning
HLT-NAACL5
2012 Convolutional-Recursive Deep Learning for 3D Object Classification
abstract
Recent advances in 3D sensing technologies make it possible to easily record color and depth images which together can improve object recognition. Most current methods rely on very well-designed features for this new 3D modality. We in- troduce a model based on a combination of convolutional and recursive neural networks (CNN and RNN) for learning features and classifying RGB-D images. The CNN layer learns low-level translationally invariant features which are then given as inputs to multiple, fixed-tree RNNs in order to compose higher order fea- tures. RNNs can be seen as combining convolution and pooling into one efficient, hierarchical operation. Our main result is that even RNNs with random weights compose powerful features. Our model obtains state of the art performance on a standard RGB-D object dataset while being more accurate and faster during train- ing and testing than comparable architectures such as two-layer CNNs.
Richard Socher, Brody Huval, Bharath Putta Bath, Christopher D. Manning, Andrew Y. Ng
NIPS4
2012 Combining joint models for biomedical event extraction
abstract
BACKGROUND: We explore techniques for performing model combination between the UMass and Stanford biomedical event extraction systems. Both sub-components address event extraction as a structured prediction problem, and use dual decomposition (UMass) and parsing algorithms (Stanford) to find the best scoring event structure. Our primary focus is on stacking where the predictions from the Stanford system are used as features in the UMass system. For comparison, we look at simpler model combination techniques such as intersection and union which require only the outputs from each system and combine them directly. RESULTS: First, we find that stacking substantially improves performance while intersection and union provide no significant benefits. Second, we investigate the graph properties of event structures and their impact on the combination of our systems. Finally, we trace the origins of events proposed by the stacked model to determine the role each system plays in different components of the output. We learn that, while stacking can propose novel event structures not seen in either base model, these events have extremely low precision. Removing these novel events improves our already state-of-the-art F1 to 56.6% on the test set of Genia (Task 1). Overall, the combined system formed via stacking ("FAUST") performed well in the BioNLP 2011 shared task. The FAUST system obtained 1st place in three out of four tasks: 1st place in Genia Task 1 (56.0% F1) and Task 2 (53.9%), 2nd place in the Epigenetics and Post-translational Modifications track (35.0%), and 1st place in the Infectious Diseases track (55.6%). CONCLUSION: We present a state-of-the-art event extraction system that relies on the strengths of structured prediction and model combination through stacking. Akin to results on other tasks, stacking outperforms intersection and union and leads to very strong results. The utility of model combination hinges on complementary views of the data, and we show that our sub-systems capture different graph properties of event structures. Finally, by removing low precision novel events, we show that performance from stacking can be further improved.
David McClosky, Sebastian Riedel 0001, Mihai Surdeanu, Andrew McCallum, Christopher D. Manning
BMC Bioinform.5
2012 Did It Happen? The Pragmatic Complexity of Veridicality Assessment
abstract
Natural language understanding depends heavily on assessing veridicality—whether events mentioned in a text are viewed as happening or not—but little consideration is given to this property in current relation and event extraction systems. Furthermore, the work that has been done has generally assumed that veridicality can be captured by lexical semantic properties whereas we show that context and world knowledge play a significant role in shaping veridicality. We extend the FactBank corpus, which contains semantically driven veridicality annotations, with pragmatically informed ones. Our annotations are more complex than the lexical assumption predicts but systematic enough to be included in computational work on textual understanding. They also indicate that veridicality judgments are not always categorical, and should therefore be modeled as distributions. We build a classifier to automatically assign event veridicality distributions based on our new annotations. The classifier relies not only on lexical features like hedges or negations, but also on structural features and approximations of world knowledge, thereby providing a nuanced picture of the diverse factors that shape veridicality. “All I know is what I read in the papers” —Will Rogers
Marie-Catherine de Marneffe, Christopher D. Manning, Christopher Potts
Comput. Linguistics2
2012 "Without the clutter of unimportant words": Descriptive keyphrases for text visualization
abstract
Keyphrases aid the exploration of text collections by communicating salient aspects of documents and are often used to create effective visualizations of text. While prior work in HCI and visualization has proposed a variety of ways of presenting keyphrases, less attention has been paid to selecting the best descriptive terms. In this article, we investigate the statistical and linguistic properties of keyphrases chosen by human judges and determine which features are most predictive of high-quality descriptive phrases. Based on 5,611 responses from 69 graduate students describing a corpus of dissertation abstracts, we analyze characteristics of human-generated keyphrases, including phrase length, commonness, position, and part of speech. Next, we systematically assess the contribution of each feature within statistical models of keyphrase quality. We then introduce a method for grouping similar terms and varying the specificity of displayed phrases so that applications can select phrases dynamically based on the available screen space and current context of interaction. Precision-recall measures find that our technique generates keyphrases that match those selected by human judges. Crowdsourced ratings of tag cloud visualizations rank our approach above other automatic techniques. Finally, we discuss the role of HCI methods in developing new algorithmic techniques suitable for user-facing applications.
Jason Chuang, Christopher D. Manning, Jeffrey Heer
ACM Trans. Comput. Hum. Interact.2
2011 Event Extraction as Dependency Parsing
David McClosky, Mihai Surdeanu, Christopher D. Manning
ACL3
2011 Part-of-Speech Tagging from 97% to 100%: Is It Time for Some Linguistics?
Christopher D. Manning
CICLing (1)1
2011 Multiword Expression Identification with Tree Substitution Grammars: A Parsing tour de force with French
Spence Green, Marie-Catherine de Marneffe, John Bauer, Christopher D. Manning
EMNLP4
2011 Semi-Supervised Recursive Autoencoders for Predicting Sentiment Distributions
Richard Socher, Jeffrey Pennington, Eric H. Huang, Andrew Y. Ng, Christopher D. Manning
EMNLP5
2011 Risk analysis for intellectual property litigation
abstract
We introduce the problem of risk analysis for Intellectual Property (IP) lawsuits. More specifically, we focus on estimating the risk for participating parties using solely prior factors, i. e., historical and concurrent behavior of the entities involved in the case. This work represents a first step towards building a comprehensive legal risk assessment system for parties involved in litigation. This technology will allow parties to optimize their case parameters to minimize their own risk, or to settle disputes out of court and thereby ease the burden on the judicial system. In addition, it will also help U.S. courts detect and fix any inherent biases in the system.
Mihai Surdeanu, Ramesh Nallapati, George Gregory, Joshua Walker, Christopher D. Manning
ICAIL5
2011 Parsing Natural Scenes and Natural Language with Recursive Neural Networks
Richard Socher, Cliff Chiung-Yu Lin, Andrew Y. Ng, Christopher D. Manning
ICML4
2011 Analyzing the Dynamics of Research by Extracting Key Aspects of Scientific Papers
Sonal Gupta, Christopher D. Manning
IJCNLP2
2011 Partially labeled topic models for interpretable text mining
abstract
Abstract Much of the world's electronic text is annotated with human-interpretable labels, such as tags on web pages and subject codes on academic publications. Effective text mining in this setting requires models that can flexibly account for the textual patterns that underlie the observed labels while still discovering unlabeled topics. Neither supervised classification, with its focus on label prediction, nor purely unsupervised learning, which does not model the labels explicitly, is appropriate. In this paper, we present two new partially supervised generative models of labeled text, Partially Labeled Dirichlet Allocation (PLDA) and the Partially Labeled Dirichlet Process (PLDP). These models make use of the unsupervised learning machinery of topic models to discover the hidden topics within each label, as well as unlabeled, corpus-wide latent topics. We explore applications with qualitative case studies of tagged web pages from del.icio.us and PhD dissertation abstracts, demonstrating improved model interpretability over traditional topic models. We use the many tags present in our del.icio.us dataset to quantitatively demonstrate the new models' higher correlation with human relatedness scores over several strong baselines.
Daniel Ramage, Christopher D. Manning, Susan T. Dumais
KDD2
2011 Dynamic Pooling and Unfolding Recursive Autoencoders for Paraphrase Detection
abstract
Paraphrase detection is the task of examining two sentences and determining whether they have the same meaning. In order to obtain high accuracy on this task, thorough syntactic and semantic analysis of the two statements is needed. We introduce a method for paraphrase detection based on recursive autoencoders (RAE). Our unsupervised RAEs are based on a novel unfolding objective and learn feature vectors for phrases in syntactic trees. These features are used to measure the word- and phrase-wise similarity between two sentences. Since sentences may be of arbitrary length, the resulting matrix of similarity measures is of variable size. We introduce a novel dynamic pooling layer which computes a fixed-sized representation from the variable-sized matrices. The pooled representation is then used as input to a classifier. Our method outperforms other state-of-the-art approaches on the challenging MSRP paraphrase corpus.
Richard Socher, Eric H. Huang, Jeffrey Pennington, Andrew Y. Ng, Christopher D. Manning
NIPS5
2010 Hierarchical Joint Learning: Improving Joint Parsing and Named Entity Recognition with Non-Jointly Labeled Data
Jenny Rose Finkel, Christopher D. Manning
ACL2
2010 "Was It Good? It Was Provocative." Learning the Meaning of Scalar Adjectives
Marie-Catherine de Marneffe, Christopher D. Manning, Christopher Potts
ACL2
2010 Better Arabic Parsing: Baselines, Evaluations, and Analysis
Spence Green, Christopher D. Manning
COLING2
2010 Probabilistic Tree-Edit Models with Structured Latent Variables for Textual Entailment and Question Answering
Mengqiu Wang, Christopher D. Manning
COLING2
2010 Viterbi Training Improves Unsupervised Dependency Parsing
Valentin I. Spitkovsky, Hiyan Alshawi, Daniel Jurafsky, Christopher D. Manning
CoNLL4
2010 A Multi-Pass Sieve for Coreference Resolution
Karthik Raghunathan, Heeyoung Lee 0004, Sudarshan Rangarajan, Nathanael Chambers, Mihai Surdeanu, Daniel Jurafsky, Christopher D. Manning
EMNLP7
2010 Parsing to Stanford Dependencies: Trade-offs between Speed and Accuracy
Daniel M. Cer, Marie-Catherine de Marneffe, Daniel Jurafsky, Christopher D. Manning
LREC4
2010 The Best Lexical Metric for Phrase-Based Statistical MT System Optimization
Daniel M. Cer, Christopher D. Manning, Daniel Jurafsky
HLT-NAACL2
2010 Accurate Non-Hierarchical Phrase-Based Translation
Michel Galley, Christopher D. Manning
HLT-NAACL2
2010 Improved Models of Distortion Cost for Statistical Machine Translation
Spence Green, Michel Galley, Christopher D. Manning
HLT-NAACL3
2010 Subword Variation in Text Message Classification
Robert Munro, Christopher D. Manning
HLT-NAACL2
2010 Ensemble Models for Dependency Parsing: Cheap and Good?
Mihai Surdeanu, Christopher D. Manning
HLT-NAACL2
2010 Which words are hard to recognize? Prosodic, lexical, and disfluency factors that increase speech recognition error rates
Sharon Goldwater, Daniel Jurafsky, Christopher D. Manning
Speech Commun.3
2009 Quadratic-Time Dependency Parsing for Machine Translation
Michel Galley, Christopher D. Manning
ACL/IJCNLP2
2009 Robust Machine Translation Evaluation with Entailment Features
Sebastian Padó, Michel Galley, Daniel Jurafsky, Christopher D. Manning
ACL/IJCNLP4
2009 Nested Named Entity Recognition
Jenny Rose Finkel, Christopher D. Manning
EMNLP2
2009 Labeled LDA: A supervised topic model for credit attribution in multi-labeled corpora
Daniel Ramage, David Hall 0006, Ramesh Nallapati, Christopher D. Manning
EMNLP4
2009 Joint Parsing and Named Entity Recognition
Jenny Rose Finkel, Christopher D. Manning
HLT-NAACL2
2009 Hierarchical Bayesian Domain Adaptation
Jenny Rose Finkel, Christopher D. Manning
HLT-NAACL2
2009 Clustering the tagged web
abstract
Automatically clustering web pages into semantic groups promises improved search and browsing on the web. In this paper, we demonstrate how user-generated tags from large-scale social bookmarking websites such as del.icio.us can be used as a complementary data source to page text and anchor text for improving automatic clustering of web pages. This paper explores the use of tags in 1) K-means clustering in an extended vector space model that includes tags as well as page text and 2) a novel generative clustering algorithm based on latent Dirichlet allocation that jointly models text and tags. We evaluate the models by comparing their output to an established web directory. We find that the naive inclusion of tagging data improves cluster quality versus page text alone, but a more principled inclusion can substantially improve the quality of all models with a statistically significant absolute F-score increase of 4%. The generative model outperforms K-means with another 8% F-score increase.
Daniel Ramage, Paul Heymann, Christopher D. Manning, Hector Garcia-Molina
WSDM3
2009 Measuring machine translation quality as semantic equivalence: A metric based on entailment features
Sebastian Padó, Daniel M. Cer, Michel Galley, Daniel Jurafsky, Christopher D. Manning
Mach. Transl.5
2008 Efficient, Feature-based, Conditional Random Field Parsing
Jenny Rose Finkel, Alex Kleeman, Christopher D. Manning
ACL3
2008 Which Words Are Hard to Recognize? Prosodic, Lexical, and Disfluency Factors that Increase ASR Error Rates
Sharon Goldwater, Daniel Jurafsky, Christopher D. Manning
ACL3
2008 Finding Contradictions in Text
Marie-Catherine de Marneffe, Anna N. Rafferty, Christopher D. Manning
ACL3
2008 Modeling Semantic Containment and Exclusion in Natural Language Inference
Bill MacCartney, Christopher D. Manning
COLING2
2008 A Simple and Effective Hierarchical Phrase Reordering Model
Michel Galley, Christopher D. Manning
EMNLP2
2008 Studying the History of Ideas Using Topic Models
David Hall 0006, Daniel Jurafsky, Christopher D. Manning
EMNLP3
2008 A Phrase-Based Alignment Model for Natural Language Inference
Bill MacCartney, Michel Galley, Christopher D. Manning
EMNLP3
2008 Legal Docket Classification: Where Machine Learning Stumbles
Ramesh Nallapati, Christopher D. Manning
EMNLP2
2008 Lexicon Schemas and Related Data Models: when Standards Meet Users
Thorsten Trippel, Michael Maxwell, Greville Corbett, Cambell Prince, Christopher D. Manning, Stephen Grimes, Steven Moran
LREC5
2008 A Global Joint Model for Semantic Role Labeling
abstract
We present a model for semantic role labeling that effectively captures the linguistic intuition that a semantic argument frame is a joint structure, with strong dependencies among the arguments. We show how to incorporate these strong dependencies in a statistical joint model with a rich set of features over multiple argument phrases. The proposed model substantially outperforms a similar state-of-the-art local model that does not include dependencies among different arguments. We evaluate the gains from incorporating this joint information on the Propbank corpus, when using correct syntactic parse trees as input, and when using automatically derived parse trees. The gains amount to 24.1% error reduction on all arguments and 36.8% on core arguments for gold-standard parse trees on Propbank. For automatic parse trees, the error reductions are 8.3% and 10.3% on all and core arguments, respectively. We also present results on the CoNLL 2005 shared task data set. Additionally, we explore considering multiple syntactic analyses to cope with parser noise and uncertainty.
Kristina Toutanova, Aria Haghighi, Christopher D. Manning
Comput. Linguistics3
2007 The Infinite Tree
Jenny Rose Finkel, Trond Grenager, Christopher D. Manning
ACL3
2007 Regularization, adaptation, and non-independent features improve hidden conditional random fields for phone classification
abstract
We show a number of improvements in the use of Hidden Conditional Random Fields (HCRFs) for phone classification on the TIMIT and Switchboard corpora. We first show that the use of regularization effectively prevents overfitting, improving over other methods such as early stopping. We then show that HCRFs are able to make use of non-independent features in phone classification, at least with small numbers of mixture components, while HMMs degrade due to their strong independence assumptions. Finally, we successfully apply Maximum a Posteriori adaptation to HCRFs, decreasing the phone classification error rate in the Switchboard corpus by around 1% – 5% given only small amounts of adaptation data.
Yun-Hsuan Sung, Constantinos Boulis, Christopher D. Manning, Daniel Jurafsky
ASRU3
2006 An Effective Two-Stage Model for Exploiting Non-Local Dependencies in Named Entity Recognition
abstract
This paper shows that a simple two-stage approach to handle non-local dependencies in Named Entity Recognition (NER) can outperform existing approaches that handle non-local dependencies, while being much more computationally efficient. NER systems typically use sequence models for tractable inference, but this makes them unable to capture the long distance structure present in text. We use a Conditional Random Field (CRF) based NER system using local features to make predictions and then train another CRF which uses both local information and features extracted from the output of the first CRF. Using features capturing non-local dependencies from the same document, our approach yields a 12.6% relative error reduction on the F1 score, over state-of-the-art NER systems using local-information alone, when compared to the 9.3% relative error reduction offered by the best systems that exploit non-local information. Our approach also makes it easy to incorporate non-local information from other documents in the test corpus, and this gives us a 13.3% error reduction over NER systems using local-information alone. Additionally, our running time for inference is just the inference time of two sequential CRFs, which is much less than that directly model the dependencies and do approximate inference.
Vijay Krishnan, Christopher D. Manning
ACL2
2006 Solving the Problem of Cascading Errors: Approximate Bayesian Inference for Linguistic Annotation Pipelines
Jenny Rose Finkel, Christopher D. Manning, Andrew Y. Ng
EMNLP2
2006 Unsupervised Discovery of a Statistical Verb Lexicon
Trond Grenager, Christopher D. Manning
EMNLP2
2006 Generating Typed Dependency Parses from Phrase Structure Parses
Marie-Catherine de Marneffe, Bill MacCartney, Christopher D. Manning
LREC3
2006 Learning to recognize features of valid textual entailments
Bill MacCartney, Trond Grenager, Marie-Catherine de Marneffe, Daniel M. Cer, Christopher D. Manning
HLT-NAACL5
2006 Graphical Model Representations of Word Lattices
abstract
We introduce a method for expressing word lattices within a dynamic graphical model. We describe a variety of choices for doing this, including a technique to relax the time information associated with lattice nodes in a way that trades off hypothesis expansion with presumed segmentation boundary accuracy. Our approach uses a set of time-inhomogeneous and algorithmically expressed conditional probability tables to encode the lattice. The approach was implemented as part of the graphical model toolkit, and word error rate improvements on the Switchboard corpus indicate that our technique is a viable means to incorporate large state space speech recognition systems into a graphical model.
Gang Ji, Jeff A. Bilmes, Jeff Michels, Katrin Kirchhoff, Christopher D. Manning
SLT5
2005 Robust Textual Inference Via Learning and Abductive Reasoning
Rajat Raina, Andrew Y. Ng, Christopher D. Manning
AAAI3
2005 Incorporating Non-local Information into Information Extraction Systems by Gibbs Sampling
abstract
Most current statistical natural language processing models use only local features so as to permit dynamic programming in inference, but this makes them unable to fully account for the long distance structure that is prevalent in language use. We show how to solve this dilemma with Gibbs sampling, a simple Monte Carlo method used to perform approximate inference in factored probabilistic models. By using simulated annealing in place of Viterbi decoding in sequence models such as HMMs, CMMs, and CRFs, it is possible to incorporate non-local structure while preserving tractable inference. We use this technique to augment an existing CRF-based information extraction system with long-distance dependency models, enforcing label consistency and extraction template consistency constraints. This technique results in an error reduction of up to 9% over state-of-the-art systems on two established information extraction tasks.
Jenny Rose Finkel, Trond Grenager, Christopher D. Manning
ACL3
2005 Unsupervised Learning of Field Segmentation Models for Information Extraction
abstract
The applicability of many current information extraction techniques is severely limited by the need for supervised training data. We demonstrate that for certain field structured extraction tasks, such as classified advertisements and bibliographic citations, small amounts of prior knowledge can be used to learn effective models in a primarily unsupervised fashion. Although hidden Markov models (HMMs) provide a suitable generative model for field structured text, general unsupervised HMM learning fails to learn useful structure in either of our domains. However, one can dramatically improve the quality of the learned structure by exploiting simple prior knowledge of the desired solutions. In both domains, we found that unsupervised methods can attain accuracies with 400 unlabeled examples comparable to those attained by supervised methods on 50 labeled examples, and that semi-supervised methods can make good use of small amounts of labeled data.
Trond Grenager, Daniel Klein 0001, Christopher D. Manning
ACL3
2005 Joint Learning Improves Semantic Role Labeling
abstract
Despite much recent progress on accurate semantic role labeling, previous work has largely used independent classifiers, possibly combined with separate label sequence models via Viterbi decoding. This stands in stark contrast to the linguistic observation that a core argument frame is a joint structure, with strong dependencies between arguments. We show how to build a joint model of argument frames, incorporating novel features that model these interactions into discriminative log-linear models. This system achieves an error reduction of 22% on all arguments and 32% on core arguments over a state-of-the art independent classifier for gold-standard parse trees on PropBank.
Kristina Toutanova, Aria Haghighi, Christopher D. Manning
ACL3
2005 A Joint Model for Semantic Role Labeling
Aria Haghighi, Kristina Toutanova, Christopher D. Manning
CoNLL3
2005 Exploring the boundaries: gene and protein identification in biomedical text
abstract
BACKGROUND: Good automatic information extraction tools offer hope for automatic processing of the exploding biomedical literature, and successful named entity recognition is a key component for such tools. METHODS: We present a maximum-entropy based system incorporating a diverse set of features for identifying gene and protein names in biomedical abstracts. RESULTS: This system was entered in the BioCreative comparative evaluation and achieved a precision of 0.83 and recall of 0.84 in the "open" evaluation and a precision of 0.78 and recall of 0.85 in the "closed" evaluation. CONCLUSION: Central contributions are rich use of features derived from the training data at multiple levels of granularity, a focus on correctly identifying entity boundaries, and the innovative use of several external knowledge sources including full MEDLINE abstracts and web searches.
Jenny Rose Finkel, Shipra Dingare, Christopher D. Manning, Malvina Nissim, Beatrice Alex, Claire Grover
BMC Bioinform.3
2005 Natural language grammar induction with a generative constituent-context model
Daniel Klein 0001, Christopher D. Manning
Pattern Recognit.2
2004 Corpus-Based Induction of Syntactic Structure: Models of Dependency and Constituency
abstract
We present a generative model for the unsupervised learning of dependency structures. We also describe the multiplicative combination of this dependency model with a model of linear constituency. The product model outperforms both components on their respective evaluation metrics, giving the best published figures for unsupervised dependency parsing and unsupervised constituency parsing. We also demonstrate that the combined model works and is robust cross-linguistically, being able to exploit either attachment or distributional regularities that are salient in the data.
Daniel Klein 0001, Christopher D. Manning
ACL2
2004 Deep Dependencies from Context-Free Statistical Parsers: Correcting the Surface Dependency Approximation
abstract
We present a linguistically-motivated algorithm for reconstructing nonlocal dependency in broad-coverage context-free parse trees derived from treebanks. We use an algorithm based on loglinear classifiers to augment and reshape context-free trees so as to reintroduce underlying nonlocal dependencies lost in the context-free approximation. We find that our algorithm compares favorably with prior work on English using an existing evaluation metric, and also introduce and argue for a new dependency-based evaluation metric. By this new evaluation metric our algorithm achieves 60% error reduction on gold-standard input trees and 5% error reduction on state-of-the-art machine-parsed input trees, when compared with the best previous work. We also present the first results on non-local dependency reconstruction for a language other than English, comparing performance on English and German. Our new evaluation metric quantitatively corroborates the intuition that in a language with freer word order, the surface dependencies in context-free parse trees are a poorer approximation to underlying dependency structure.
Roger Levy, Christopher D. Manning
ACL2
2004 Language Learning: Beyond Thunderdome
Christopher D. Manning
CoNLL1
2004 Using Feature Conjunctions Across Examples for Learning Pairwise Classifiers
Satoshi Oyama, Christopher D. Manning
ECML2
2004 Verb Sense and Subcategorization: Using Joint Inference to Improve Performance on Complementary Task
Galen Andrew, Trond Grenager, Christopher D. Manning
EMNLP3
2004 Max-Margin Parsing
Ben Taskar, Daniel Klein 0001, Michael Collins 0001, Daphne Koller, Christopher D. Manning
EMNLP5
2004 The Leaf Path Projection View of Parse Trees: Exploring String Kernels for HPSG Parse Selection
Kristina Toutanova, Penka Markova, Christopher D. Manning
EMNLP3
2004 Learning random walk models for inducing word dependency distributions
abstract
Many NLP tasks rely on accurately estimating word dependency probabilities P(ω1|ω2), where the words w1 and w2 have a particular relationship (such as verb-object). Because of the sparseness of counts of such dependencies, smoothing and the ability to use multiple sources of knowledge are important challenges. For example, if the probability P(N|V) of noun N being the subject of verb V is high, and V takes similar objects to V', and V' is synonymous to V", then we want to conclude that P(N|V") should also be reasonably high---even when those words did not cooccur in the training data.To capture these higher order relationships, we propose a Markov chain model, whose stationary distribution is used to give word probability estimates. Unlike the manually defined random walks used in some link analysis algorithms, we show how to automatically learn a rich set of parameters for the Markov chain's transition probabilities. We apply this model to the task of prepositional phrase attachment, obtaining an accuracy of 87.54%.
Kristina Toutanova, Christopher D. Manning, Andrew Y. Ng
ICML2
2003 Accurate Unlexicalized Parsing
abstract
We demonstrate that an unlexicalized PCFG can parse much more accurately than previously shown, by making use of simple, linguistically motivated state splits, which break down false independence assumptions latent in a vanilla treebank grammar. Indeed, its performance of 86.36% (LP/LR F1) is better than that of early lexicalized PCFG models, and surprisingly close to the current state-of-the-art. This result has potential uses beyond establishing a strong lower bound on the maximum possible accuracy of unlexicalized models: an unlexicalized PCFG is much more compact, easier to replicate, and easier to interpret than more complex lexical models, and the parsing algorithms are simpler, more widely understood, of lower asymptotic complexity, and easier to optimize.
Daniel Klein 0001, Christopher D. Manning
ACL2
2003 Is it Harder to Parse Chinese, or the Chinese Treebank?
abstract
L¼ ¥ S " " h S 9{ | ¦S t 9 w{ ¥¬ w .
Roger Levy, Christopher D. Manning
ACL2
2003 Named Entity Recognition with Character-Level Models
Daniel Klein 0001, Joseph Smarr, Christopher D. Manning
CoNLL4
2003 A Generative Model for Semantic Role Labeling
Cynthia A. Thompson, Roger Levy, Christopher D. Manning
ECML3
2003 Optimizing Local Probability Models for Statistical Parsing
Kristina Toutanova, Mark Mitchell, Christopher D. Manning
ECML3
2003 Spectral Learning
Sepandar D. Kamvar, Daniel Klein 0001, Christopher D. Manning
IJCAI3
2003 Factored A* Search for Models over Sequences and Trees
Daniel Klein 0001, Christopher D. Manning
IJCAI2
2003 A* Parsing: Fast Exact Viterbi Parse Selection
Daniel Klein 0001, Christopher D. Manning
HLT-NAACL2
2003 Optimization, Maxent Models, and Conditional Estimation without Magic
Christopher D. Manning, Daniel Klein 0001
HLT-NAACL1
2003 Feature-Rich Part-of-Speech Tagging with a Cyclic Dependency Network
Kristina Toutanova, Daniel Klein 0001, Christopher D. Manning, Yoram Singer
HLT-NAACL3
2003 Log-Linear Models for Label Ranking
abstract
Label ranking is the task of inferring a total order over a predefined set of labels for each given instance. We present a general framework for batch learning of label ranking functions from supervised data. We assume that each instance in the training data is associated with a list of preferences over the label-set, however we do not assume that this list is either com- plete or consistent. This enables us to accommodate a variety of ranking problems. In contrast to the general form of the supervision, our goal is to learn a ranking function that induces a total order over the entire set of labels. Special cases of our setting are multilabel categorization and hierarchical classification. We present a general boosting-based learning algorithm for the label ranking problem and prove a lower bound on the progress of each boosting iteration. The applicability of our approach is demonstrated with a set of experiments on a large-scale text corpus.
Ofer Dekel, Christopher D. Manning, Yoram Singer
NIPS2
2003 Extrapolation methods for accelerating PageRank computations
abstract
We present a novel algorithm for the fast computation of PageRank, a hyperlink-based estimate of the ''importance'' of Web pages. The original PageRank algorithm uses the Power Method to compute successive iterates that converge to the principal eigenvector of the Markov matrix representing the Web link graph. The algorithm presented here, called Quadratic Extrapolation, accelerates the convergence of the Power Method by periodically subtracting off estimates of the nonprincipal eigenvectors from the current iterate of the Power Method. In Quadratic Extrapolation, we take advantage of the fact that the first eigenvalue of a Markov matrix is known to be 1 to compute the nonprincipal eigenvectors using successive iterates of the Power Method. Empirically, we show that using Quadratic Extrapolation speeds up PageRank computation by 25-300% on a Web graph of 80 million nodes, with minimal overhead. Our contribution is useful to the PageRank community and the numerical linear algebra community in general, as it is a fast method for determining the dominant eigenvector of a matrix that is too large for standard fast methods to be practical.
Sepandar D. Kamvar, Taher H. Haveliwala, Christopher D. Manning, Gene H. Golub
WWW3
2002 A Generative Constituent-Context Model for Improved Grammar Induction
abstract
We present a generative distributional model for the unsupervised induction of natural language syntax which explicitly models constituent yields and contexts. Parameter search with EM produces higher quality analyses than previously exhibited by unsupervised systems, giving the best published un-supervised parsing results on the ATIS corpus. Experiments on Penn treebank sentences of comparable length show an even higher F1 of 71% on non-trivial brackets. We compare distributionally induced and actual part-of-speech tags as input data, and examine extensions to the basic model. We discuss errors made by the system, compare the system to previous models, and discuss upper bounds, lower bounds, and stability for this task.
Daniel Klein 0001, Christopher D. Manning
ACL2
2002 The LinGO Redwoods Treebank: Motivation and Preliminary Applications
Stephan Oepen, Kristina Toutanova, Stuart M. Shieber, Christopher D. Manning, Dan Flickinger, Thorsten Brants
COLING4
2002 Feature Selection for a Rich HPSG Grammar Using Decision Trees
Kristina Toutanova, Christopher D. Manning
CoNLL2
2002 Conditional Structure versus Conditional Estimation in NLP Models
abstract
This paper separates conditional parameter estimation, which consistently raises test set accuracy on statistical NLP tasks, from conditional model structures, such as the conditional Markov model used for maximum-entropy tagging, which tend to lower accuracy. Error analysis on part-of-speech tagging shows that the actual tagging errors made by the conditionally structured model derive not only from label bias, but also from other ways in which the independence assumptions of the conditional model structure are unsuited to linguistic sequences. The paper presents new word-sense disambiguation and POS tagging experiments, and integrates apparently conflicting reports from other recent work.
Daniel Klein 0001, Christopher D. Manning
EMNLP2
2002 Extentions to HMM-based Statistical Word Alignment Models
abstract
This paper describes improved HMM-based word level alignment models for statistical machine translation. We present a method for using part of speech tag information to improve alignment accuracy, and an approach to modeling fertility and correspondence to the empty word in an HMM alignment model. We present accuracy results from evaluating Viterbi alignments against human-judged alignments on the Canadian Hansards corpus, as compared to a bigram HMM, and IBM model 4. The results show up to 16% alignment error reduction.
Kristina Toutanova, H. Tolga Ilhan, Christopher D. Manning
EMNLP3
2002 Interpreting and Extending Classical Agglomerative Clustering Algorithms using a Model-Based approach
Sepandar D. Kamvar, Daniel Klein 0001, Christopher D. Manning
ICML3
2002 From Instance-level Constraints to Space-Level Constraints: Making the Most of Prior Knowledge in Data Clustering
Daniel Klein 0001, Sepandar D. Kamvar, Christopher D. Manning
ICML3
2002 Fast Exact Inference with a Factored Model for Natural Language Parsing
abstract
We present a novel generative model for natural language tree structures in which semantic (lexical dependency) and syntactic (PCFG) structures are scored with separate models. This factorization provides concep- tual simplicity, straightforward opportunities for separately improving the component models, and a level of performance comparable to simi- lar, non-factored models. Most importantly, unlike other modern parsing models, the factored model admits an extremely effective A* parsing al- gorithm, which enables efficient, exact inference.
Daniel Klein 0001, Christopher D. Manning
NIPS2
2001 Parsing with Treebank Grammars: Empirical Bounds, Theoretical Models, and the Structure of the Penn Treebank
abstract
This paper presents empirical studies and closely corresponding theoretical models of the performance of a chart parser exhaustively parsing the Penn Treebank with the Treebank's own CFG grammar.We show how performance is dramatically affected by rule representation and tree transformations, but little by top-down vs. bottom-up strategies.We discuss grammatical saturation, including analysis of the strongly connected components of the phrasal nonterminals in the Treebank, and model how, as sentence length increases, the effective grammar rule size increases as regions of the grammar are unlocked, yielding super-cubic observed time behavior in some configurations.
Daniel Klein 0001, Christopher D. Manning
ACL2
2001 Natural Language Grammar Induction Using a Constituent-Context Model
abstract
This paper presents a novel approach to the unsupervised learning of syn- tactic analyses of natural language text. Most previous work has focused on maximizing likelihood according to generative PCFG models. In con- trast, we employ a simpler probabilistic model over trees based directly on constituent identity and linear context, and use an EM-like iterative procedure to induce structure. This method produces much higher qual- ity analyses, giving the best published results on the ATIS dataset. 1 Overview To enable a wide range of subsequent tasks, human language sentences are standardly given tree-structure analyses, wherein the nodes in a tree dominate contiguous spans of words called constituents, as in figure 1(a). Constituents are the linguistically coherent units in the sentence, and are usually labeled with a constituent category, such as noun phrase (NP) or verb phrase (VP). An aim of grammar induction systems is to figure out, given just the sentences in a corpus S, what tree structures correspond to them. In this sense, the grammar induction problem is an incomplete data problem, where the complete data is the corpus of trees T , but we only observe their yields S. This paper presents a new approach to this problem, which gains leverage by directly making use of constituent contexts. It is an open problem whether entirely unsupervised methods can produce linguistically accurate parses of sentences. Due to the difficulty of this task, the vast majority of statis- tical parsing work has focused on supervised learning approaches to parsing, where one uses a treebank of fully parsed sentences to induce a model which parses unseen sentences [7, 3]. But there are compelling motivations for unsupervised grammar induction. Building supervised training data requires considerable resources, including time and linguistic ex- pertise. Investigating unsupervised methods can shed light on linguistic phenomena which are implicit within a supervised parser's supervisory information (e.g., unsupervised sys- tems often have difficulty correctly attaching subjects to verbs above objects, whereas for a supervised parser, this ordering is implicit in the supervisory information). Finally, while the presented system makes no claims to modeling human language acquisition, results on whether there is enough information in sentences to recover their structure are important data for linguistic theory, where it has standardly been assumed that the information in the data is deficient, and strong innate knowledge is required for language acquisition [4].
Daniel Klein 0001, Christopher D. Manning
NIPS2
2000 What's related? Generalizing approaches to related articles in medicine
Howard R. Strasberg, Christopher D. Manning, Thomas C. Rindfleisch, Kenneth L. Melmon
AMIA2
2000 Enriching the Knowledge Sources Used in a Maximum Entropy Part-of-Speech Tagger
abstract
This paper presents results for a maximum-entropy-based part of speech tagger, which achieves superior performance principally by enriching the information sources used for tagging. In particular, we get improved results by incorporating these features: (i) more extensive treatment of capitalization for unknown words; (ii) features for the disambiguation of the tense forms of verbs; (iii) features for disambiguating particles from prepositions and adverbs. The best resulting accuracy for the tagger on the Penn Treebank is 96.86% overall, and 86.91% on previously unseen words.
Kristina Toutanvoa, Christopher D. Manning
EMNLP2
1999 Linguistics in an Age of Engineering
Christopher D. Manning
PACLIC1
1998 The segmentation problem in morphology learning
Christopher D. Manning
CoNLL1
1993 Automatic Acquisition of a Large Subcategorization Dictionary from Corpora
abstract
This paper presents a new method for producing a dictionary of subcategorization frames from unlabelled text corpora. It is shown that statistical filtering of the results of a finite state parser running on the output of a stochastic tagger produces high quality results, despite the error rates of the tagger and the parser. Further, it is argued that this method can be used to learn all subcategorization frames, whereas previous methods are not extensible to a general solution to the problem.
Christopher D. Manning
ACL1