EDBT 2026 Demo / reviewers in the wild / expert
Danqi Chen 0001
dblp:87/7949
· DBLP profile ↗
73ranked-venue papers
6as first author
58since 2021 · last 2025
0000-0001-5308-2634ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 69 · 5 first-author · 57 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How to Train Long-Context Language Models (Effectively)abstractWe study continued training and supervised fine-tuning (SFT) of a language model (LM) to make effective use of long-context information.We first establish a reliable evaluation protocol to guide model development-instead of perplexity or simple needle-in-a-haystack (NIAH) tests, we use a broad set of long-context downstream tasks, and we evaluate models after SFT as this better reveals long-context abilities.Supported by our robust evaluations, we run thorough experiments to decide the data mix for continued pre-training, the instruction tuning dataset, and many other design choices such as position extrapolation.We find that (1) code repositories and books are excellent sources of long data, but it is crucial to combine them with high-quality short-context data;(2) training with a sequence length beyond the evaluation length boosts long-context performance; (3) for SFT, using only short instruction datasets yields strong performance on long-context tasks.Our final model, ProLong-8B, which is initialized from Llama-3 and trained on 40B tokens, demonstrates state-ofthe-art long-context performance among similarly sized models at a length of 128K.Pro-Long outperforms Llama-3.1-8B-Instruct on the majority of long-context tasks despite using only 5% as many tokens during long-context training.Additionally, ProLong can effectively process up to 512K tokens, one of the longest context windows of publicly available LMs. Tianyu Gao 0001, Alexander Wettig, Howard Yen, Danqi Chen 0001 |
ACL (1) | 4 |
| 2025 | Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-rankingabstractRecent work has identified retrieval heads (Wu et al., 2025b), a subset of attention heads responsible for retrieving salient information in long-context language models (LMs), as measured by their copy-paste behavior in Needlein-a-Haystack tasks.In this paper, we introduce QRHEAD (Query-Focused Retrieval Head), an improved set of attention heads that enhance retrieval from long context.We identify QRHEAD by aggregating attention scores with respect to the input query, using a handful of examples from real-world tasks (e.g., long-context QA).We further introduce QR-RETRIEVER, an efficient and effective retriever that uses the accumulated attention mass of QRHEAD as retrieval scores.We use QR-RETRIEVER for long-context reasoning by selecting the most relevant parts with the highest retrieval scores.On multi-hop reasoning tasks LongMemEval and CLIPPER, this yields over 10% performance gains over full context and outperforms strong dense retrievers.We also evaluate QRRETRIEVER as a re-ranker on the BEIR benchmark and find that it achieves strong zero-shot performance, outperforming other LLM-based re-rankers such as RankGPT.Further analysis shows that both the querycontext attention scoring and task selection are crucial for identifying QRHEAD with strong downstream utility.Overall, our work contributes a general-purpose retriever and offers interpretability insights into the long-context capabilities of LMs. 1 Wuwei Zhang, Fangcong Yin, Howard Yen, Danqi Chen 0001, Xi Ye 0003 |
EMNLP | 4 |
| 2025 | Fantastic Copyrighted Beasts and How (Not) to Generate ThemabstractRecent studies show that image and video generation models can be prompted to reproduce copyrighted content from their training data, raising serious legal con- cerns about copyright infringement. Copyrighted characters (e.g., Mario, Batman) present a significant challenge: at least one lawsuit has already awarded damages based on the generation of such characters. Consequently, commercial services like DALL·E have started deploying interventions. However, little research has systematically examined these problems: (1) Can users easily prompt models to generate copyrighted characters, even if it is unintentional?; (2) How effective are the existing mitigation strategies? To address these questions, we introduce a novel evaluation framework with metrics that assess both the generated image’s similarity to copyrighted characters and its consistency with user intent, grounded in a set of popular copyrighted characters from diverse studios and regions. We show that state-of-the-art image and video generation models can still generate characters even if characters’ names are not explicitly mentioned, sometimes with only two generic keywords (e.g., prompting with “videogame, plumber” consistently gener- ates Nintendo’s Mario character). We also introduce semi-automatic techniques to identify such keywords or descriptions that trigger character generation. Using this framework, we evaluate mitigation strategies, including prompt rewriting and new approaches we propose. Our findings reveal that common methods, such as DALL·E’s prompt rewriting, are insufficient alone and require supplementary strategies like negative prompting. Our work provides empirical grounding for discussions on copyright mitigation strategies and offers actionable insights for model deployers implementing these safeguards. Luxi He, Yangsibo Huang, Tinghao Xie, Luke Zettlemoyer, Chiyuan Zhang, Danqi Chen 0001, Peter Henderson 0002 |
ICLR | 9 |
| 2025 | Unintentional Unalignment: Likelihood Displacement in Direct Preference OptimizationabstractDirect Preference Optimization (DPO) and its variants are increasingly used for aligning language models with human preferences.
Although these methods are designed to teach a model to generate preferred responses more frequently relative to dispreferred responses, prior work has observed that the likelihood of preferred responses often decreases during training. The current work sheds light on the causes and implications of this counterintuitive phenomenon, which we term *likelihood displacement*. We demonstrate that likelihood displacement can be *catastrophic*, shifting probability mass from preferred responses to responses with an opposite meaning. As a simple example, training a model to prefer $\texttt{No}$ over $\texttt{Never}$ can sharply increase the probability of $\texttt{Yes}$. Moreover, when aligning the model to refuse unsafe prompts, we show that such displacement can *unintentionally lead to unalignment*, by shifting probability mass from preferred refusal responses to harmful responses (e.g., reducing the refusal rate of Llama-3-8B-Instruct from 74.4% to 33.4%). We theoretically characterize that likelihood displacement is driven by preferences that induce similar embeddings, as measured by a *centered hidden embedding similarity (CHES)* score. Empirically, the CHES score enables identifying which training samples contribute most to likelihood displacement in a given dataset. Filtering out these samples effectively mitigated unintentional unalignment in our experiments. More broadly, our results highlight the importance of curating data with sufficiently distinct preferences, for which we believe the CHES score may prove valuable. Noam Razin, Sadhika Malladi, Adithya Bhaskar, Danqi Chen 0001, Sanjeev Arora, Boris Hanin |
ICLR | 4 |
| 2025 | BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive RetrievalabstractExisting retrieval benchmarks primarily consist of information-seeking queries (e.g., aggregated questions from search engines) where keyword or semantic-based retrieval is usually sufficient. However, many complex real-world queries require in-depth reasoning to identify relevant documents that go beyond surface form matching. For example, finding documentation for a coding question requires understanding the logic and syntax of the functions involved. To better benchmark retrieval on such challenging queries, we introduce BRIGHT, the first text retrieval benchmark that requires intensive reasoning to retrieve relevant documents. Our dataset consists of 1,398 real-world queries spanning diverse domains such as economics, psychology, mathematics, coding, and more. These queries are drawn from naturally occurring or carefully curated human data. Extensive evaluation reveals that even state-of-the-art retrieval models perform poorly on BRIGHT. The leading model on the MTEB leaderboard (Muennighoff et al., 2023), which achieves a score of 59.0 nDCG@10,1 produces a score of nDCG@10 of 18.0 on BRIGHT. We show that incorporating explicit reasoning about the query improves retrieval performance by up to 12.2 points. Moreover, incorporating retrieved documents from the top-performing retriever boosts question answering performance by over 6.6 points. We believe that BRIGHT paves the way for future research on retrieval systems in more realistic and challenging settings. Hongjin Su, Howard Yen, Mengzhou Xia, Niklas Muennighoff, Han-yu Wang, Haisu Liu, Zachary S. Siegel, Michael Tang, Ruoxi Sun 0002, Jinsung Yoon, Sercan Ö. Arik, Danqi Chen 0001, Tao Yu 0009 |
ICLR | 14 |
| 2025 | SORRY-Bench: Systematically Evaluating Large Language Model Safety RefusalabstractEvaluating aligned large language models' (LLMs) ability to recognize and reject unsafe user requests is crucial for safe, policy-compliant deployments. Existing evaluation efforts, however, face three limitations that we address with **SORRY-Bench**, our proposed benchmark. **First**, existing methods often use coarse-grained taxonomies of unsafe topics, and are over-representing some fine-grained topics. For example, among the ten existing datasets that we evaluated, tests for refusals of self-harm instructions are over 3x less represented than tests for fraudulent activities. SORRY-Bench improves on this by using a fine-grained taxonomy of 44 potentially unsafe topics, and 440 class-balanced unsafe instructions, compiled through human-in-the-loop methods. **Second**, evaluations often overlook the linguistic formatting of prompts, like different languages, dialects, and more --- which are only implicitly considered in many evaluations. We supplement SORRY-bench with 20 diverse linguistic augmentations to systematically examine these effects. **Third**, existing evaluations rely on large LLMs (e.g., GPT-4) for evaluation, which can be computationally expensive. We investigate design choices for creating a fast, accurate automated safety evaluator. By collecting 7K+ human annotations and conducting a meta-evaluation of diverse LLM-as-a-judge designs, we show that fine-tuned 7B LLMs can achieve accuracy comparable to GPT-4 scale LLMs, with lower computational cost. Putting these together, we evaluate over 50 proprietary and open-weight LLMs on SORRY-Bench, analyzing their distinctive safety refusal behaviors. We hope our effort provides a building block for systematic evaluations of LLMs' safety refusal capabilities, in a balanced, granular, and efficient manner. Benchmark demo, data, code, and models are available through [https://sorry-bench.github.io](https://sorry-bench.github.io). Tinghao Xie, Xiangyu Qi, Yi Zeng 0005, Yangsibo Huang, T. W. U. Madhushani, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng 0007, Ruoxi Jia 0001, Bo Li 0026, Kai Li 0001, Danqi Chen 0001, Peter Henderson 0002, Prateek Mittal |
ICLR | 14 |
| 2025 | HELMET: How to Evaluate Long-context Models Effectively and ThoroughlyabstractMany benchmarks exist for evaluating long-context language models (LCLMs), yet developers often rely on synthetic tasks such as needle-in-a-haystack (NIAH) or an arbitrary subset of tasks. However, it remains unclear whether these benchmarks reflect the diverse downstream applications of LCLMs, and such inconsistencies further complicate model comparison. We investigate the underlying reasons behind these practices and find that existing benchmarks often provide noisy signals due to limited coverage of applications, insufficient context lengths, unreliable metrics, and incompatibility with base models. In this work, we introduce HELMET (How to Evaluate Long-context Models Effectively and Thoroughly), a comprehensive benchmark encompassing seven diverse, application-centric categories. We also address several issues in previous benchmarks by adding controllable lengths up to 128K tokens, model-based evaluation for reliable metrics, and few-shot prompting for robustly evaluating base models. Consequently, we demonstrate that HELMET offers more reliable and consistent rankings of frontier LCLMs. Through a comprehensive study of 59 LCLMs, we find that (1) synthetic tasks like NIAH do not reliably predict downstream performance; (2) the diverse categories in HELMET exhibit distinct trends and low correlations with each other; and (3) while most LCLMs achieve perfect NIAH scores, open-source models significantly lag behind closed ones when tasks require full-context reasoning or following complex instructions---the gap widens as length increases. Finally, we recommend using our RAG tasks for fast model development, as they are easy to run and better predict other downstream performance; ultimately, we advocate for a holistic evaluation across diverse tasks. Howard Yen, Tianyu Gao 0001, Minmin Hou, Daniel Fleischer, Peter Izsak, Moshe Wasserblat, Danqi Chen 0001 |
ICLR | 8 |
| 2025 | Metadata Conditioning Accelerates Language Model Pre-trainingabstractThe vast diversity of styles, domains, and quality levels present in language model pre-training corpora is essential in developing general model capabilities, but efficiently learning and deploying the correct behaviors exemplified in each of these heterogeneous data sources is challenging. To address this, we propose a new method, termed Metadata Conditioning then Cooldown (MeCo), to incorporate additional learning cues during pre-training. MeCo first provides metadata (e.g., URLs like en.wikipedia.org) alongside the text during training and later uses a cooldown phase with only the standard text, thereby enabling the model to function normally even without metadata. MeCo significantly accelerates pre-training across different model scales (600M to 8B parameters) and training sources (C4, RefinedWeb, and DCLM). For instance, a 1.6B language model trained with MeCo matches the downstream task performance of standard pre-training while using 33% less data. Additionally, MeCo enables us to steer language models by conditioning the inference prompt on either real or fabricated metadata that encodes the desired properties of the output: for example, prepending wikipedia.org to reduce harmful generations or factquizmaster.com (fabricated) to improve common knowledge task performance. We also demonstrate that MeCo is compatible with different types of metadata, such as model-generated topics. MeCo is remarkably simple, adds no computational overhead, and demonstrates promise in producing more capable and steerable language models. Tianyu Gao 0001, Alexander Wettig, Luxi He, Yihe Dong, Sadhika Malladi, Danqi Chen 0001 |
ICML | 6 |
| 2025 | Organize the Web: Constructing Domains Enhances Pre-Training Data CurationabstractModern language models are trained on large, unstructured datasets consisting of trillions of tokens and obtained by crawling the web. The unstructured nature makes it difficult to reason about their contents and develop systematic approaches to data curation. In this paper, we unpack monolithic web corpora by developing taxonomies of their contents and organizing them into domains. We introduce WebOrganizer, a framework for organizing web pages in terms of both their topic and format. Using these two complementary notions of domains, we automatically annotate pre-training data by distilling annotations from a large language model into efficient classifiers. This allows us to study how data from different domains should be mixed to improve models on downstream tasks, and we show that we can combine insights about effective topics and formats to further boost performance. We demonstrate that our domain mixing also improves existing methods that select data based on quality. Furthermore, we study and compare how quality-based methods will implicitly change the domain mixture. Overall, our work demonstrates that constructing and mixing domains provides a valuable complement to quality-based data curation methods, opening new avenues for effective and insightful pre-training data curation. Alexander Wettig, Kyle Lo, Sewon Min, Hannaneh Hajishirzi, Danqi Chen 0001, Luca Soldaini |
ICML | 5 |
| 2025 | Representing Rule-based Chatbots with TransformersabstractDan Friedman, Abhishek Panigrahi, Danqi Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dan Friedman, Abhishek Panigrahi, Danqi Chen 0001 |
NAACL (Long Papers) | 3 |
| 2025 | Precise Information Control in Long-Form Text GenerationabstractA central challenge in language models (LMs) is faithfulness hallucination: the generation of information unsubstantiated by input context. To study this problem, we propose Precise Information Control (PIC), a new task formulation that requires models to generate long-form outputs grounded in a provided set of short self-contained statements, without adding any unsupported ones. PIC includes a full setting that tests a model’s ability to include exactly all input claims, and a partial setting that requires the model to selectively incorporate only relevant claims. We present PIC-Bench, a benchmark of eight long-form generation tasks (e.g., summarization, biography generation) adapted to the PIC setting, where LMs are supplied with well-formed, verifiable input claims. Our evaluation of a range of open and proprietary LMs on PIC-Bench reveals that, surprisingly, state-of-the-art LMs still hallucinate against user-provided input in over 70% of generations. To alleviate this lack of faithfulness, we introduce a post-training framework that uses a weakly supervised preference data construction method to train an 8B PIC-LM with stronger PIC ability—improving from 69.1% to 91.0% F1 in the full PIC setting. When integrated into end-to-end factual generation pipelines, PIC-LM improves exact match recall by 17.1% on ambiguous QA with retrieval, and factual precision by 30.5% on a birthplace fact-checking task, underscoring the potential of precisely grounded generation. Jacqueline He, Howard Yen, Margaret Li, Shuyue Stella Li, Yulia Tsvetkov, Danqi Chen 0001, Pang Wei W. Koh, Luke Zettlemoyer |
NeurIPS | 8 |
| 2025 | The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningabstractReinforcement learning with verifiable rewards (RLVR) is a promising approach for training language models (LMs) on reasoning tasks that elicit emergent long chains of thought (CoTs). Unlike supervised learning, it updates the model using both correct and incorrect samples via policy gradients. To better understand its mechanism, we decompose the learning signal into reinforcing correct responses and penalizing incorrect ones, referred to as **P**ositive and **N**egative **S**ample **R**einforcement (**PSR** and **NSR**), respectively. We train `Qwen2.5-Math-7B`, `Qwen3-4B` and `Llama-3.1-8B-Instruct` on a mathematical reasoning dataset and uncover a surprising result: training with only negative samples — without reinforcing correct responses — can be highly effective: it consistently improves performance over the base model across the entire Pass@$k$ spectrum $k$ up to 256), often matching or surpassing PPO and GRPO. In contrast, reinforcing only correct responses improves Pass@1 but degrades performance at higher $k$, due to reduced diversity. These inference-scaling trends highlight that solely penalizing incorrect responses may contribute more to performance than previously recognized. Through gradient analysis, we show that NSR works by suppressing incorrect generations and redistributing probability mass toward other plausible candidates, guided by the model's prior beliefs. It refines the model's existing knowledge rather than introducing entirely new behaviors. Building on this insight, we propose a simple variant of the RL objective that upweights NSR, and show that it consistently improves overall Pass@$k$ performance on MATH, AIME 2025, and AMC23. Our code is available at [`https://github.com/TianHongZXY/RLVR-Decomposed`](https://github.com/TianHongZXY/RLVR-Decomposed). Mengzhou Xia, Zhepei Wei, Danqi Chen 0001, Yu Meng 0001 |
NeurIPS | 5 |
| 2024 | The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language ModelsabstractPrior work has found that pretrained language models (LMs) fine-tuned with different random seeds can achieve similar in-domain performance but generalize differently on tests of syntactic generalization.In this work, we show that, even within a single model, we can find multiple subnetworks that perform similarly indomain, but generalize vastly differently.To better understand these phenomena, we investigate if they can be understood in terms of "competing subnetworks": the model initially represents a variety of distinct algorithms, corresponding to different subnetworks, and generalization occurs when it ultimately converges to one.This explanation has been used to account for generalization in simple algorithmic tasks ("grokking").Instead of finding competing subnetworks, we find that all subnetworkswhether they generalize or not-share a set of attention heads, which we refer to as the heuristic core.Further analysis suggests that these attention heads emerge early in training and compute shallow, non-generalizing features.The model learns to generalize by incorporating additional attention heads, which depend on the outputs of the "heuristic" heads to compute higher-level features.Overall, our results offer a more detailed picture of the mechanisms for syntactic generalization in pretrained LMs. 1 Adithya Bhaskar, Dan Friedman, Danqi Chen 0001 |
ACL (1) | 3 |
| 2024 | Long-Context Language Modeling with Parallel Context EncodingabstractExtending large language models (LLMs) to process longer inputs is crucial for a wide range of applications.However, the substantial computational cost of transformers and limited generalization of positional encoding restrict the size of their context window.We introduce Context Expansion with Parallel Encoding (CEPE ), a framework that can be applied to any existing decoder-only LLMs to extend their context window.CEPE employs a small encoder to process long inputs chunk by chunk, enabling the frozen decoder to utilize additional contexts via cross-attention.CEPE is efficient, generalizable, and versatile: trained with 8K-token documents, it extends the context window of LLAMA-2 to 128K tokens, offering 10× the throughput with only 1/6 of the memory.CEPE yields strong performance on language modeling and in-context learning.CEPE also excels in retrieval-augmented applications, while existing long-context models degenerate with retrieved contexts.We further introduce a CEPE variant that can extend the context window of instruction-tuned models using only unlabeled data, and showcase its effectiveness on LLAMA-2-CHAT, leading to a strong instruction-following model that can leverage very long contexts on downstream tasks. 1Chapter 01: Dune ... Howard Yen, Tianyu Gao 0001, Danqi Chen 0001 |
ACL (1) | 3 |
| 2024 | LitSearch: A Retrieval Benchmark for Scientific Literature SearchabstractLiterature search questions, such as "Where can I find research on the evaluation of consistency in generated summaries?"pose significant challenges for modern search engines and retrieval systems.These questions often require a deep understanding of research concepts and the ability to reason across entire articles.In this work, we introduce LitSearch, a retrieval benchmark comprising 597 realistic literature search queries about recent ML and NLP papers.Lit-Search is constructed using a combination of (1) questions generated by GPT-4 based on paragraphs containing inline citations from research papers and (2) questions manually written by authors about their recently published papers.All LitSearch questions were manually examined or edited by experts to ensure high quality.We extensively benchmark state-ofthe-art retrieval models and also evaluate two LLM-based reranking pipelines.We find a significant performance gap between BM25 and state-of-the-art dense retrievers, with a 24.8% absolute difference in [email protected] LLMbased reranking strategies further improve the best-performing dense retriever by 4.4%.Additionally, commercial search engines and research tools like Google Search perform poorly on LitSearch, lagging behind the best dense retriever by up to 32 recall points.Taken together, these results show that LitSearch is an informative new testbed for retrieval systems while catering to a real-world use case. Anirudh Ajith, Mengzhou Xia, Alexis Chevalier, Tanya Goyal, Danqi Chen 0001, Tianyu Gao 0001 |
EMNLP | 5 |
| 2024 | Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationabstractThe rapid progress in open-source large language models (LLMs) is significantly advancing AI development. Extensive efforts have been made before model release to align their behavior with human values, with the primary goal of ensuring their helpfulness and harmlessness. However, even carefully aligned models can be manipulated maliciously, leading to unintended behaviors, known as ``jailbreaks". These jailbreaks are typically triggered by specific text inputs, often referred to as adversarial prompts. In this work, we propose the generation exploitation attack, an extremely simple approach that disrupts model alignment by only manipulating variations of decoding methods. By exploiting different generation strategies, including varying decoding hyper-parameters and sampling methods, we increase the attack success rate from $0\%$ to more than $95\%$ across 11 language models including LLaMA2, Vicuna, Falcon, and MPT families, outperforming state-of-the-art attacks with $30\times$ lower computational cost. Finally, we propose an effective alignment method that explores diverse generation strategies, which can reasonably reduce the attack success rate under our attack. Altogether, our study underscores a major failure in current safety evaluation and alignment procedures for open-source LLMs, strongly advocating for more comprehensive red teaming and better alignment before releasing such models. Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li 0001, Danqi Chen 0001 |
ICLR | 5 |
| 2024 | Detecting Pretraining Data from Large Language ModelsabstractAlthough large language models (LLMs) are widely deployed, the data used to train them is rarely disclosed. Given the incredible scale of this data, up to trillions of tokens, it is all but certain that it includes potentially problematic text such as copyrighted materials, personally identifiable information, and test data for widely reported reference benchmarks. However, we currently have no way to know which data of these types is included or in what proportions. In this paper, we study the pretraining data detection problem: given a piece of text and black-box access to an LLM without knowing the pretraining data, can we determine if the model was trained on the provided text? To facilitate this study, we introduce a dynamic benchmark WIKIMIA that uses data created before and after model training to support gold truth detection. We also introduce a new detection method MIN-K PROB based on a simple hypothesis: an unseen example is likely to contain a few outlier words with low probabilities under the LLM, while a seen example is less likely to have words with such low probabilities. MIN-K PROB can be applied without any knowledge about the pretrainig corpus or any additional training, departing from previous detection methods that require training a reference model on data that is similar to the pretraining data. Moreover, our experiments demonstrate that MIN-K PROB achieves a 7.4% improvement on WIKIMIA over these previous methods. We apply MIN-K PROB to two real-world scenarios, copyrighted book detection and contaminated downstream example detection, and find that it to be a consistently effective solution. Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen 0001, Luke Zettlemoyer |
ICLR | 7 |
| 2024 | Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningabstractThe popularity of LLaMA (Touvron et al., 2023a;b) and other recently emerged moderate-sized large language models (LLMs) highlights the potential of building smaller yet powerful LLMs. Regardless, the cost of training such models from scratch on trillions of tokens remains high. In this work, we study structured pruning as an effective means to develop smaller LLMs from pre-trained, larger models. Our approach employs two key techniques: (1) targeted structured pruning, which prunes a larger model to a specified target shape by removing layers, heads, and intermediate and hidden dimensions in an end-to-end manner, and (2) dynamic batch loading, which dynamically updates the composition of sampled data in each training batch based on varying losses across different domains. We demonstrate the efficacy of our approach by presenting the Sheared-LLaMA series, pruning the LLaMA2-7B model down to 1.3B and 2.7B parameters. Sheared-LLaMA models outperform state-of-the-art open-source models of equivalent sizes, such as Pythia, INCITE, OpenLLaMA and the concurrent TinyLlama models, on a wide range of downstream and instruction tuning evaluations, while requiring only 3% of compute compared to training such models from scratch. This work provides compelling evidence that leveraging existing LLMs with structured pruning is a far more cost-effective approach for building competitive small-scale LLMs Mengzhou Xia, Tianyu Gao 0001, Danqi Chen 0001 |
ICLR | 4 |
| 2024 | Evaluating Large Language Models at Evaluating Instruction FollowingabstractAs research in large language models (LLMs) continues to accelerate, LLM-based evaluation has emerged as a scalable and cost-effective alternative to human evaluations for comparing the ever increasing list of models. This paper investigates the efficacy of these “LLM evaluators”, particularly in using them to assess instruction following, a metric that gauges how closely generated text adheres to the given instruction. We introduce a challenging meta-evaluation benchmark, LLMBar, designed to test the ability of an LLM evaluator in discerning instruction-following outputs. The authors manually curated 419 pairs of outputs, one adhering to instructions while the other diverging, yet may possess deceptive qualities that mislead an LLM evaluator, e.g., a more engaging tone. Contrary to existing meta-evaluation, we discover that different evaluators (i.e., combinations of LLMs and prompts) exhibit distinct performance on LLMBar and even the highest-scoring ones have substantial room for improvement. We also present a novel suite of prompting strategies that further close the gap between LLM and human evaluators. With LLMBar, we hope to offer more insight into LLM evaluators and foster future research in developing better instruction-following models. Jiatong Yu, Tianyu Gao 0001, Yu Meng 0001, Tanya Goyal, Danqi Chen 0001 |
ICLR | 6 |
| 2024 | Language Models as Science TutorsabstractNLP has recently made exciting progress toward training language models (LMs) with strong scientific problem-solving skills. However, model development has not focused on real-life use-cases of LMs for science, including applications in education that require processing long scientific documents. To address this, we introduce TutorEval and TutorChat. TutorEval is a diverse question-answering benchmark consisting of questions about long chapters from STEM textbooks, written by experts. TutorEval helps measure real-life usability of LMs as scientific assistants, and it is the first benchmark combining long contexts, free-form generation, and multi-disciplinary scientific knowledge. Moreover, we show that fine-tuning base models with existing dialogue datasets leads to poor performance on TutorEval. Therefore, we create TutorChat, a dataset of 80,000 long synthetic dialogues about textbooks. We use TutorChat to fine-tune Llemma models with 7B and 34B parameters. These LM tutors specialized in math have a 32K-token context window, and they excel at TutorEval while performing strongly on GSM8K and MATH. Our datasets build on open-source materials, and we release our models, data, and evaluations publicly. Alexis Chevalier, Jiayi Geng, Alexander Wettig, Howard Chen 0003, Sebastian Mizera, Toni Annala, Max Jameson Aragon, Arturo Rodríguez Fanlo, Simon Frieder, Simon Machado, Akshara Prabhakar, Ellie Thieu, Jiachen T. Wang, Xindi Wu, Mengzhou Xia, Wenhan Xia, Jiatong Yu, Zhiyong Jason Ren, Sanjeev Arora, Danqi Chen 0001 |
ICML | 22 |
| 2024 | Interpretability Illusions in the Generalization of Simplified ModelsabstractA common method to study deep learning systems is to use simplified model representations—for example, using singular value decomposition to visualize the model’s hidden states in a lower dimensional space. This approach assumes that the results of these simplifications are faithful to the original model. Here, we illustrate an important caveat to this assumption: even if the simplified representations can accurately approximate the full model on the training set, they may fail to accurately capture the model’s behavior out of distribution. We illustrate this by training Transformer models on controlled datasets with systematic generalization splits, including the Dyck balanced-parenthesis languages and a code completion task. We simplify these models using tools like dimensionality reduction and clustering, and then explicitly test how these simplified proxies match the behavior of the original model. We find consistent generalization gaps: cases in which the simplified proxies are more faithful to the original model on the in-distribution evaluations and less faithful on various tests of systematic generalization. This includes cases where the original model generalizes systematically but the simplified proxies fail, and cases where the simplified proxies generalize better. Together, our results raise questions about the extent to which mechanistic interpretations derived using tools like SVD can reliably predict what a model will do in novel situations. Dan Friedman, Andrew K. Lampinen, Lucas Dixon, Danqi Chen 0001, Asma Ghandeharioun |
ICML | 4 |
| 2024 | QuRating: Selecting High-Quality Data for Training Language ModelsabstractSelecting high-quality pre-training data is important for creating capable language models, but existing methods rely on simple heuristics. We introduce QuRating, a method for selecting pre-training data that can capture human intuitions about data quality. In this paper, we investigate four qualities - writing style, required expertise, facts & trivia, and educational value - and find that LLMs are able to discern these qualities, especially when making pairwise judgments of texts. We train a QuRater model to learn scalar ratings from pairwise judgments, and use it to annotate a 260B training corpus with quality ratings for each of the four criteria. In our experiments, we select 30B tokens according to the different quality ratings and train 1.3B-parameter language models on the selected data. We find that it is important to balance quality and diversity. When we sample using quality ratings as logits over documents, our models obtain lower perplexity and stronger in-context learning performance than baselines. Our best model is based on educational value and performs similarly to a model trained with uniform sampling for 50% more steps. Beyond data selection, we use the quality ratings to construct a training curriculum which improves performance without changing the training dataset. We extensively analyze the quality ratings and discuss their characteristics, biases, and wider implications. Alexander Wettig, Aatmik Gupta, Saumya Malik, Danqi Chen 0001 |
ICML | 4 |
| 2024 | LESS: Selecting Influential Data for Targeted Instruction TuningabstractInstruction tuning has unlocked powerful capabilities in large language models (LLMs), using combined datasets to develop general-purpose chatbots. However, real-world applications often require a specialized suite of skills (e.g., reasoning). The challenge lies in identifying the most relevant data from these extensive datasets to effectively develop specific capabilities, a setting we frame as targeted instruction tuning. We propose LESS, an optimizer-aware and practically efficient algorithm to estimate data influences and perform Low-rank gradiEnt Similarity Search for instruction data selection. Crucially, LESS adapts existing influence formulations to work with the Adam optimizer and variable-length instruction data. LESS first constructs a highly reusable and transferable gradient datastore with low-dimensional gradient features and then selects examples based on their similarity to few-shot examples embodying a specific capability. Experiments show that training on a LESS-selected 5% of the data can often outperform training on the full dataset across diverse downstream tasks. Furthermore, the selected data is highly transferable: smaller models can be leveraged to select useful data for larger models and models from different families. Our qualitative analysis shows that our method goes beyond surface form cues to identify data that exemplifies the necessary reasoning skills for the intended downstream application. To facilitate future work, we release code and data at princeton-nlp/LESS. Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, Danqi Chen 0001 |
ICML | 5 |
| 2024 | SimPO: Simple Preference Optimization with a Reference-Free RewardabstractDirect Preference Optimization (DPO) is a widely used offline preference optimization algorithm that reparameterizes reward functions in reinforcement learning from human feedback (RLHF) to enhance simplicity and training stability. In this work, we propose SimPO, a simpler yet more effective approach. The effectiveness of SimPO is attributed to a key design: using the _average_ log probability of a sequence as the implicit reward. This reward formulation better aligns with model generation and eliminates the need for a reference model, making it more compute and memory efficient. Additionally, we introduce a target reward margin to the Bradley-Terry objective to encourage a larger margin between the winning and losing responses, further improving the algorithm's performance. We compare SimPO to DPO and its latest variants across various state-of-the-art training setups, including both base and instruction-tuned models such as Mistral, Llama 3, and Gemma 2. We evaluate on extensive chat-based evaluation benchmarks, including AlpacaEval 2, MT-Bench, and Arena-Hard. Our results demonstrate that SimPO consistently and significantly outperforms existing approaches without substantially increasing response length. Specifically, SimPO outperforms DPO by up to 6.4 points on AlpacaEval 2 and by up to 7.5 points on Arena-Hard. Our top-performing model, built on Gemma-2-9B-it, achieves a 72.4\% length-controlled win rate on AlpacaEval 2, a 59.1\% win rate on Arena-Hard, and ranks 1st on Chatbot Arena among $<$10B models with real user votes. Yu Meng 0001, Mengzhou Xia, Danqi Chen 0001 |
NeurIPS | 3 |
| 2024 | Finding Transformer Circuits With Edge PruningabstractThe path to interpreting a language model often proceeds via analysis of circuits---sparse computational subgraphs of the model that capture specific aspects of its behavior. Recent work has automated the task of discovering circuits. Yet, these methods have practical limitations, as they either rely on inefficient search algorithms or inaccurate approximations. In this paper, we frame circuit discovery as an optimization problem and propose _Edge Pruning_ as an effective and scalable solution. Edge Pruning leverages gradient-based pruning techniques, but instead of removing neurons or components, prunes the _edges_ between components. Our method finds circuits in GPT-2 that use less than half the number of edges than circuits found by previous methods while being equally faithful to the full model predictions on standard circuit-finding tasks. Edge Pruning is efficient on tasks involving up to 100,000 examples, outperforming previous methods in speed and producing substantially better circuits. It also perfectly recovers the ground-truth circuits in two models compiled with Tracr. Thanks to its efficiency, we scale Edge Pruning to CodeLlama-13B, a model over 100x the size of GPT-2.
We use this setting for a case study, where we compare the mechanisms behind instruction prompting and in-context learning.
We find two circuits with more than 99.96% sparsity that match the performance of the full model. Further analysis reveals that the mechanisms in the two settings overlap substantially. This shows that Edge Pruning is a practical and scalable tool for interpretability,
which can shed light on behaviors that only emerge in large models. Adithya Bhaskar, Alexander Wettig, Dan Friedman, Danqi Chen 0001 |
NeurIPS | 4 |
| 2024 | CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMsabstractChart understanding plays a pivotal role when applying Multimodal Large Language Models (MLLMs) to real-world tasks such as analyzing scientific papers or financial reports. However, existing datasets often focus on oversimplified and homogeneous charts with template-based questions, leading to an overly optimistic measure of progress. We demonstrate that although open-source models can appear to outperform strong proprietary models on these benchmarks, a simple stress test with slightly different charts or questions deteriorates performance by up to 34.5%. In this work, we propose CharXiv, a comprehensive evaluation suite involving 2,323 natural, challenging, and diverse charts from scientific papers. CharXiv includes two types of questions: 1) descriptive questions about examining basic chart elements and 2) reasoning questions that require synthesizing information across complex visual elements in the chart. To ensure quality, all charts and questions are handpicked, curated, and verified by human experts. Our results reveal a substantial, previously underestimated gap between the reasoning skills of the strongest proprietary model (i.e., GPT-4o), which achieves 47.1% accuracy, and the strongest open-source model (i.e., InternVL Chat V1.5), which achieves 29.2%. All models lag far behind human performance of 80.5%, underscoring weaknesses in the chart understanding capabilities of existing MLLMs. We hope that CharXiv facilitates future research on MLLM chart understanding by providing a more realistic and faithful measure of progress. Project website: https://charxiv.github.io/ Mengzhou Xia, Luxi He, Howard Chen 0003, Yitao Liu, Richard Zhu, Kaiqu Liang, Xindi Wu, Sadhika Malladi, Alexis Chevalier, Sanjeev Arora, Danqi Chen 0001 |
NeurIPS | 13 |
| 2023 | Measuring Inductive Biases of In-Context Learning with Underspecified DemonstrationsabstractIn-context learning (ICL) is an important paradigm for adapting large language models (LLMs) to new tasks, but the generalization behavior of ICL remains poorly understood.We investigate the inductive biases of ICL from the perspective of feature bias: which feature ICL is more likely to use given a set of underspecified demonstrations in which two features are equally predictive of the labels.First, we characterize the feature biases of GPT-3 models by constructing underspecified demonstrations from a range of NLP datasets and feature combinations.We find that LLMs exhibit clear feature biases-for example, demonstrating a strong bias to predict labels according to sentiment rather than shallow lexical features, like punctuation.Second, we evaluate the effect of different interventions that are designed to impose an inductive bias in favor of a particular feature, such as adding a natural language instruction or using semantically relevant label words.We find that, while many interventions can influence the learner to prefer a particular feature, it can be difficult to overcome strong prior biases.Overall, our results provide a broader picture of the types of features that ICL may be more likely to exploit and how to impose inductive biases that are better aligned with the intended task. 1 Chenglei Si, Dan Friedman, Nitish Joshi, Shi Feng 0005, Danqi Chen 0001, He He 0001 |
ACL (1) | 5 |
| 2023 | Training Trajectories of Language Models Across ScalesabstractMengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin, Ramakanth Pasunuru, Danqi Chen, Luke Zettlemoyer, Veselin Stoyanov. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Mengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin, Ramakanth Pasunuru, Danqi Chen 0001, Luke Zettlemoyer, Veselin Stoyanov |
ACL (1) | 6 |
| 2023 | Should You Mask 15% in Masked Language Modeling?abstractMasked language models (MLMs) conventionally mask 15% of tokens due to the belief that more masking would leave insufficient context to learn good representations; this masking rate has been widely used, regardless of model sizes or masking strategies.In this work, we revisit this important choice of MLM pre-training.We first establish that 15% is not universally optimal, and larger models should adopt a higher masking rate.Specifically, we find that masking 40% outperforms 15% for BERT-large size models on GLUE and SQuAD.Interestingly, an extremely high masking rate of 80% can still preserve 95% fine-tuning performance and most of the accuracy in linguistic probing, challenging the conventional wisdom about the role of the masking rate.We then examine the interplay between masking rates and masking strategies and find that uniform masking requires a higher masking rate compared to sophisticated masking strategies such as span or PMI masking.Finally, we argue that increasing the masking rate has two distinct effects: it leads to more corruption, which makes the prediction task harder; it also enables more predictions, which benefits optimization.Using this framework, we revisit BERT's 80-10-10 corruption strategy.Together, our results contribute to a better understanding of MLM pre-training. 1 Alexander Wettig, Tianyu Gao 0001, Zexuan Zhong, Danqi Chen 0001 |
EACL | 4 |
| 2023 | Adapting Language Models to Compress ContextsabstractTransformer-based language models (LMs) are powerful and widely-applicable tools, but their usefulness is constrained by a finite context window and the expensive computational cost of processing long text documents.We propose to adapt pre-trained LMs into AutoCompressors.These language models are capable of compressing long contexts into compact summary vectors, which are then accessible to the model as soft prompts.Summary vectors are trained with an unsupervised objective, whereby long documents are processed in segments, and summary vectors from all previous segments are used in language modeling.We fine-tune OPT and Llama-2 models on sequences of up to 30,720 tokens and show that AutoCompressors can utilize long contexts to improve perplexity.We evaluate AutoCompressors on in-context learning by compressing task demonstrations and find that summary vectors are good substitutes for plain-text demonstrations, increasing accuracy while reducing inference costs.Finally, we explore the benefits of pre-computing summary vectors for large corpora by applying summary vectors to retrievalaugmented language modeling and a passage re-ranking task.Overall, AutoCompressors emerge as a simple and inexpensive solution to extend the context window of LMs while speeding up inference over long contexts.1 * AC and AW contributed equally.This work was done when AC was at the Institute for Advanced Study and visited the Princeton NLP group.1 Our code and models are publicly available at https://github. Alexis Chevalier, Alexander Wettig, Anirudh Ajith, Danqi Chen 0001 |
EMNLP | 4 |
| 2023 | C-STS: Conditional Semantic Textual SimilarityabstractAmeet Deshpande, Carlos Jimenez, Howard Chen, Vishvak Murahari, Victoria Graf, Tanmay Rajpurohit, Ashwin Kalyan, Danqi Chen, Karthik Narasimhan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Ameet Deshpande, Carlos E. Jimenez, Howard Chen 0003, Vishvak Murahari, Victoria Graf, Tanmay Rajpurohit, Ashwin Kalyan, Danqi Chen 0001, Karthik Narasimhan |
EMNLP | 8 |
| 2023 | Enabling Large Language Models to Generate Text with CitationsabstractLarge language models (LLMs) have emerged as a widely-used tool for information seeking, but their generated outputs are prone to hallucination.In this work, our aim is to allow LLMs to generate text with citations, improving their factual correctness and verifiability.Existing work mainly relies on commercial search engines and human evaluation, making it challenging to reproduce and compare different modeling approaches.We propose ALCE, the first benchmark for Automatic LLMs' Citation Evaluation.ALCE collects a diverse set of questions and retrieval corpora and requires building end-to-end systems to retrieve supporting evidence and generate answers with citations.We develop automatic metrics along three dimensions-fluency, correctness, and citation quality-and demonstrate their strong correlation with human judgements.Our experiments with state-of-the-art LLMs and novel prompting strategies show that current systems have considerable room for improvement-For example, on the ELI5 dataset, even the best models lack complete citation support 50% of the time.Our analyses further highlight promising future directions, including developing better retrievers, advancing long-context LLMs, and improving the ability to synthesize information from multiple sources. 1 When did the US break away from England? Question Short answers (from the dataset) Tianyu Gao 0001, Howard Yen, Jiatong Yu, Danqi Chen 0001 |
EMNLP | 4 |
| 2023 | Privacy Implications of Retrieval-Based Language ModelsabstractRetrieval-based language models (LMs) have demonstrated improved interpretability, factuality, and adaptability compared to their parametric counterparts by incorporating retrieved text from external datastores.While it is well known that parametric models are prone to leaking private data, it remains unclear how the addition of a retrieval datastore impacts model privacy.In this work, we present the first study of privacy risks in retrieval-based LMs, particularly kNN-LMs.Our goal is to explore the optimal design and training procedure in domains where privacy is of concern, aiming to strike a balance between utility and privacy.Crucially, we find that kNN-LMs are more susceptible to leaking private information from their private datastore than parametric models.We further explore mitigations of privacy risks: When privacy information is targeted and readily detected in the text, we find that a simple sanitization step would eliminate the risks while decoupling query and key encoders achieves an even better utility-privacy trade-off.Otherwise, we consider strategies of mixing public and private data in both datastore and encoder training.While these methods offer modest improvements, they leave considerable room for future work.Together, our findings provide insights for practitioners to better understand and mitigate privacy risks in retrieval-based LMs 1 . Yangsibo Huang, Samyak Gupta, Zexuan Zhong, Kai Li 0001, Danqi Chen 0001 |
EMNLP | 5 |
| 2023 | Poisoning Retrieval Corpora by Injecting Adversarial PassagesabstractDense retrievers have achieved state-of-the-art performance in various information retrieval tasks, but to what extent can they be safely deployed in real-world applications?In this work, we propose a novel attack for dense retrieval systems in which a malicious user generates a small number of adversarial passages by perturbing discrete tokens to maximize similarity with a provided set of training queries.When these adversarial passages are inserted into a large retrieval corpus, we show that this attack is highly effective in fooling these systems to retrieve them for queries that were not seen by the attacker.More surprisingly, these adversarial passages can directly generalize to out-ofdomain queries and corpora with a high success attack rate-for instance, we find that 50 generated passages optimized on Natural Questions can mislead >94% of questions posed in financial documents or online forums.We also benchmark and compare a range of state-ofthe-art dense retrievers, both unsupervised and supervised.Although different systems exhibit varying levels of vulnerability, we show they can all be successfully attacked by injecting up to 500 passages, a small fraction compared to a retrieval corpus of millions of passages.1 Zexuan Zhong, Ziqing Huang, Alexander Wettig, Danqi Chen 0001 |
EMNLP | 4 |
| 2023 | MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop QuestionsabstractThe information stored in large language models (LLMs) falls out of date quickly, and retraining from scratch is often not an option.This has recently given rise to a range of techniques for injecting new facts through updating model weights.Current evaluation paradigms are extremely limited, mainly validating the recall of edited facts, but changing one fact should cause rippling changes to the model's related beliefs.If we edit the UK Prime Minister to now be Rishi Sunak, then we should get a different answer to Who is married to the British Prime Minister?In this work, we present a benchmark, MQUAKE (Multi-hop Question Answering for Knowledge Editing), comprising multi-hop questions that assess whether edited models correctly answer questions where the answer should change as an entailed consequence of edited facts.While we find that current knowledge-editing approaches can recall edited facts accurately, they fail catastrophically on the constructed multi-hop questions.We thus propose a simple memory-based approach, MeLLo, which stores all edited facts externally while prompting the language model iteratively to generate answers that are consistent with the edited facts.While MQUAKE remains challenging, we show that MeLLo scales well with LLMs (up to 175B) and outperforms previous model editors by a large margin. 1 Zexuan Zhong, Zhengxuan Wu, Christopher D. Manning, Christopher Potts, Danqi Chen 0001 |
EMNLP | 5 |
| 2023 | A Kernel-Based View of Language Model Fine-TuningabstractIt has become standard to solve NLP tasks by fine-tuning pre-trained language models (LMs), especially in low-data settings. There is minimal theoretical understanding of empirical success, e.g., why fine-tuning a model with $10^8$ or more parameters on a couple dozen training points does not result in overfitting. We investigate whether the Neural Tangent Kernel (NTK)---which originated as a model to study the gradient descent dynamics of infinitely wide networks with suitable random initialization---describes fine-tuning of pre-trained LMs. This study was inspired by the decent performance of NTK for computer vision tasks (Wei et al., 2022). We extend the NTK formalism to Adam and use Tensor Programs (Yang, 2020) to characterize conditions under which the NTK lens may describe fine-tuning updates to pre-trained language models. Extensive experiments on 14 NLP tasks validate our theory and show that formulating the downstream task as a masked word prediction problem through prompting often induces kernel-based dynamics during fine-tuning. Finally, we use this kernel view to propose an explanation for the success of parameter-efficient subspace-based fine-tuning methods. Sadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen 0001, Sanjeev Arora |
ICML | 4 |
| 2023 | Learning Transformer ProgramsabstractRecent research in mechanistic interpretability has attempted to reverse-engineer Transformer models by carefully inspecting network weights and activations. However, these approaches require considerable manual effort and still fall short of providing complete, faithful descriptions of the underlying algorithms. In this work, we introduce a procedure for training Transformers that are mechanistically interpretable by design. We build on RASP [Weiss et al., 2021], a programming language that can be compiled into Transformer weights. Instead of compiling human-written programs into Transformers, we design a modified Transformer that can be trained using gradient-based optimization and then automatically converted into a discrete, human-readable program. We refer to these models as Transformer Programs. To validate our approach, we learn Transformer Programs for a variety of problems, including an in-context learning task, a suite of algorithmic problems (e.g. sorting, recognizing Dyck languages), and NLP tasks including named entity recognition and text classification. The Transformer Programs can automatically find reasonable solutions, performing on par with standard Transformers of comparable size; and, more importantly, they are easy to interpret. To demonstrate these advantages, we convert Transformers into Python programs and use off-the-shelf code analysis tools to debug model errors and identify the “circuits” used to solve different sub-problems. We hope that Transformer Programs open a new path toward the goal of intrinsically interpretable machine learning. Dan Friedman, Alexander Wettig, Danqi Chen 0001 |
NeurIPS | 3 |
| 2023 | Fine-Tuning Language Models with Just Forward PassesabstractFine-tuning language models (LMs) has yielded success on diverse downstream tasks, but as LMs grow in size, backpropagation requires a prohibitively large amount of memory. Zeroth-order (ZO) methods can in principle estimate gradients using only two forward passes but are theorized to be catastrophically slow for optimizing large models. In this work, we propose a memory-efficient zerothorder optimizer (MeZO), adapting the classical ZO-SGD method to operate in-place, thereby fine-tuning LMs with the same memory footprint as inference. For example, with a single A100 80GB GPU, MeZO can train a 30-billion parameter model, whereas fine-tuning with backpropagation can train only a 2.7B LM with the same budget. We conduct comprehensive experiments across model types (masked and autoregressive LMs), model scales (up to 66B), and downstream tasks (classification, multiple-choice, and generation). Our results demonstrate that (1) MeZO significantly outperforms in-context learning and linear probing; (2) MeZO achieves comparable performance to fine-tuning with backpropagation across multiple tasks, with up to 12× memory reduction and up to 2× GPU-hour reduction in our implementation; (3) MeZO is compatible with both full-parameter and parameter-efficient tuning techniques such as LoRA and prefix tuning; (4) MeZO can effectively optimize non-differentiable objectives (e.g., maximizing accuracy or F1). We support our empirical findings with theoretical insights, highlighting how adequate pre-training and task prompts enable MeZO to fine-tune huge models, despite classical ZO analyses suggesting otherwise. Sadhika Malladi, Tianyu Gao 0001, Eshaan Nichani, Alex Damian, Jason D. Lee, Danqi Chen 0001, Sanjeev Arora |
NeurIPS | 6 |
| 2023 | SIGIR 2023 Workshop on Retrieval Enhanced Machine Learning (REML @ SIGIR 2023)abstractMost machine learning models are designed to be self-contained and encode both "knowledge" and "reasoning" in their parameters. However, such models cannot perform effectively for tasks that require knowledge grounding and tasks that deal with non-stationary data, such as news and social media. Besides, these models often require huge number of parameters to encode all the required knowledge. These issues can be addressed via augmentation with a retrieval model. This category of machine learning models, which is called Retrieval-enhanced machine learning (REML), has recently attracted considerable attention in multiple research communities. For instance, REML models have been studied in the context of open-domain question answering, fact verification, and dialogue systems and also in the context of generalization through memorization in language models and memory networks. We believe that the information retrieval community can significantly contribute to this growing research area by designing, implementing, analyzing, and evaluating various aspects of retrieval models with applications to REML tasks. The goal of this full-day hybrid workshop is to bring together researchers from industry and academia to discuss various aspects of retrieval-enhanced machine learning, including effectiveness, efficiency, and robustness of these models in addition to their impact on real-world applications. Michael Bendersky, Danqi Chen 0001, Fernando Diaz 0001, Hamed Zamani |
SIGIR | 2 |
| 2022 | Ditch the Gold Standard: Re-evaluating Conversational Question AnsweringabstractConversational question answering aims to provide natural-language answers to users in information-seeking conversations. Existing conversational QA benchmarks compare models with pre-collected human-human conversations, using ground-truth answers provided in conversational history. It remains unclear whether we can rely on this static evaluation for model development and whether current systems can well generalize to real-world human-machine conversations. In this work, we conduct the first large-scale human evaluation of state-of-the-art conversational QA systems, where human evaluators converse with models and judge the correctness of their answers. We find that the distribution of human machine conversations differs drastically from that of human-human conversations, and there is a disagreement between human and gold-history evaluation in terms of model ranking. We further investigate how to improve automatic evaluations, and propose a question rewriting mechanism based on predicted history, which better correlates with human judgments. Finally, we analyze the impact of various modeling strategies and discuss future directions towards building better conversational question answering systems. Huihan Li 0001, Tianyu Gao 0001, Manan Goenka, Danqi Chen 0001 |
ACL (1) | 4 |
| 2022 | Structured Pruning Learns Compact and Accurate ModelsabstractThe growing size of neural language models has led to increased attention in model compression.The two predominant approaches are pruning, which gradually removes weights from a pre-trained model, and distillation, which trains a smaller compact model to match a larger one.Pruning methods can significantly reduce the model size but hardly achieve large speedups as distillation.However, distillation methods require large amounts of unlabeled data and are expensive to train.In this work, we propose a task-specific structured pruning method CoFi 1 (Coarse-and Fine-grained Pruning), which delivers highly parallelizable subnetworks and matches the distillation methods in both accuracy and latency, without resorting to any unlabeled data.Our key insight is to jointly prune coarse-grained (e.g., layers) and fine-grained (e.g., heads and hidden units) modules, which controls the pruning decision of each parameter with masks of different granularity.We also devise a layerwise distillation strategy to transfer knowledge from unpruned to pruned models during optimization.Our experiments on GLUE and SQuAD datasets show that CoFi yields models with over 10× speedups with a small accuracy drop, showing its effectiveness and efficiency compared to previous pruning and distillation approaches. 2 Mengzhou Xia, Zexuan Zhong, Danqi Chen 0001 |
ACL (1) | 3 |
| 2022 | Finding Dataset Shortcuts with Grammar InductionabstractMany NLP datasets have been found to contain shortcuts: simple decision rules that achieve surprisingly high accuracy.However, it is difficult to discover shortcuts automatically.Prior work on automatic shortcut detection has focused on enumerating features like unigrams or bigrams, which can find only low-level shortcuts, or relied on post-hoc model interpretability methods like saliency maps, which reveal qualitative patterns without a clear statistical interpretation.In this work, we propose to use probabilistic grammars to characterize and discover shortcuts in NLP datasets.Specifically, we use a contextfree grammar to model patterns in sentence classification datasets and use a synchronous context-free grammar to model datasets involving sentence pairs.The resulting grammars reveal interesting shortcut features in a number of datasets, including both simple and high-level features, and automatically identify groups of test examples on which conventional classifiers fail.Finally, we show that the features we discover can be used to generate diagnostic contrast examples and incorporated into standard robust optimization methods to improve worst-group accuracy.1 Dan Friedman, Alexander Wettig, Danqi Chen 0001 |
EMNLP | 3 |
| 2022 | MABEL: Attenuating Gender Bias using Textual Entailment DataabstractPre-trained language models encode undesirable social biases, which are further exacerbated in downstream use.To this end, we propose MABEL (a Method for Attenuating Gender Bias using Entailment Labels), an intermediate pre-training approach for mitigating gender bias in contextualized representations.Key to our approach is the use of a contrastive learning objective on counterfactually augmented, gender-balanced entailment pairs from natural language inference (NLI) datasets.We also introduce an alignment regularizer that pulls identical entailment pairs along opposite gender directions closer.We extensively evaluate our approach on intrinsic and extrinsic metrics, and show that MABEL outperforms previous task-agnostic debiasing approaches in terms of fairness.It also preserves task performance after fine-tuning on downstream tasks.Together, these findings demonstrate the suitability of NLI data as an effective means of bias mitigation, as opposed to only using unlabeled sentences in the literature.Finally, we identify that existing approaches often use evaluation settings that are insufficient or inconsistent.We make an effort to reproduce and compare previous methods, and call for unifying the evaluation settings across gender debiasing methods for better future comparison. 1 Jacqueline He, Mengzhou Xia, Christiane Fellbaum, Danqi Chen 0001 |
EMNLP | 4 |
| 2022 | Don't Prompt, Search! Mining-based Zero-Shot Learning with Language ModelsabstractMasked language models like BERT can perform text classification in a zero-shot fashion by reformulating downstream tasks as text infilling.However, this approach is highly sensitive to the template used to prompt the model, yet practitioners are blind when designing them in strict zero-shot settings.In this paper, we propose an alternative mining-based approach for zero-shot learning.Instead of prompting language models, we use regular expressions to mine labeled examples 1 from unlabeled corpora, which can optionally be filtered through prompting, and used to finetune a pretrained model.Our method is more flexible and interpretable than prompting, and outperforms it on a wide range of tasks when using comparable templates.Our results suggest that the success of prompting can partly be explained by the model being exposed to similar examples during pretraining, which can be directly retrieved through regular expressions. Task Lbl VerbalizersSent.Pos.good good good Mozes van de Kar, Mengzhou Xia, Danqi Chen 0001, Mikel Artetxe |
EMNLP | 3 |
| 2022 | Prompting ELECTRA: Few-Shot Learning with Discriminative Pre-Trained ModelsabstractPre-trained masked language models successfully perform few-shot learning by formulating downstream tasks as text infilling.However, as a strong alternative in full-shot settings, discriminative pre-trained models like ELECTRA do not fit into the paradigm.In this work, we adapt prompt-based few-shot learning to ELECTRA and show that it outperforms masked language models in a wide range of tasks.ELECTRA is pre-trained to distinguish if a token is generated or original.We naturally extend that to prompt-based few-shot learning by training to score the originality of the target options without introducing new parameters.Our method can be easily adapted to tasks involving multi-token predictions without extra computation overhead.Analysis shows that ELECTRA learns distributions that align better with downstream tasks. 1 Mengzhou Xia, Mikel Artetxe, Jingfei Du, Danqi Chen 0001, Veselin Stoyanov |
EMNLP | 4 |
| 2022 | Generating Natural Language Proofs with Verifier-Guided SearchabstractReasoning over natural language is a challenging problem in NLP.In this work, we focus on proof generation: Given a hypothesis and a set of supporting facts, the model generates a proof tree indicating how to derive the hypothesis from supporting facts.Compared to generating the entire proof in one shot, stepwise generation can better exploit the compositionality and generalize to longer proofs but has achieved limited success on real-world data.Existing stepwise methods struggle to generate proof steps that are both logically valid and relevant to the hypothesis.Instead, they tend to hallucinate invalid steps given the hypothesis.In this paper, we present a novel stepwise method, NLProofS (Natural Language Proof Search), which learns to generate relevant steps conditioning on the hypothesis.At the core of our approach, we train an independent verifier to check the validity of the proof steps to prevent hallucination.Instead of generating steps greedily, we search for proofs maximizing a global proof score judged by the verifier.NL-ProofS achieves state-of-the-art performance on EntailmentBank and RuleTaker.Specifically, it improves the correctness of predicted proofs from 27.7% to 33.3% in the distractor setting of EntailmentBank, demonstrating the effectiveness of NLProofS in generating challenging human-authored proofs. 1 Kaiyu Yang, Jia Deng 0001, Danqi Chen 0001 |
EMNLP | 3 |
| 2022 | Training Language Models with Memory AugmentationabstractRecent work has improved language models (LMs) remarkably by equipping them with a non-parametric memory component.However, most existing approaches only introduce memories at testing time or represent them using a separately trained encoder, resulting in suboptimal training of the language model.In this work, we present TRIME, a novel yet simple training approach designed for training LMs with memory augmentation.Our approach uses a training objective that directly takes inbatch examples as accessible memory.We also present new methods for memory construction and data batching, which are used for adapting to different sets of memories-local, longterm, and external memory-at testing time.We evaluate TRIME on multiple language modeling and machine translation benchmarks and show that it is able to achieve significant improvements across all the settings.Concretely, TRIME reduces the perplexity from 18.70 to 15.37 on WIKITEXT-103, by effectively leveraging a large memory set from the training corpus.Compared to standard LM training, TRIME adds negligible computational overhead and is compatible with different neural architectures, making it a versatile solution for training memory-augmented LMs. 1 Zexuan Zhong, Danqi Chen 0001 |
EMNLP | 3 |
| 2022 | Can Rationalization Improve Robustness?abstractHoward Chen, Jacqueline He, Karthik Narasimhan, Danqi Chen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Howard Chen 0003, Jacqueline He, Karthik Narasimhan, Danqi Chen 0001 |
NAACL-HLT | 4 |
| 2022 | Recovering Private Text in Federated Learning of Language ModelsabstractFederated learning allows distributed users to collaboratively train a model while keeping each user’s data private. Recently, a growing body of work has demonstrated that an eavesdropping attacker can effectively recover image data from gradients transmitted during federated learning. However, little progress has been made in recovering text data. In this paper, we present a novel attack method FILM for federated learning of language models (LMs). For the first time, we show the feasibility of recovering text from large batch sizes of up to 128 sentences. Unlike image-recovery methods that are optimized to match gradients, we take a distinct approach that first identifies a set of words from gradients and then directly reconstructs sentences based on beam search and a prior-based reordering strategy. We conduct the FILM attack on several large-scale datasets and show that it can successfully reconstruct single sentences with high fidelity for large batch sizes and even multiple sentences if applied iteratively.We evaluate three defense methods: gradient pruning, DPSGD, and a simple approach to freeze word embeddings that we propose. We show that both gradient pruning and DPSGD lead to a significant drop in utility. However, if we fine-tune a public pre-trained LM on private text without updating word embeddings, it can effectively defend the attack with minimal data utility loss. Together, we hope that our results can encourage the community to rethink the privacy concerns of LM training and its standard practices in the future. Our code is publicly available at https://github.com/Princeton-SysML/FILM . Samyak Gupta, Yangsibo Huang, Zexuan Zhong, Tianyu Gao 0001, Kai Li 0001, Danqi Chen 0001 |
NeurIPS | 6 |
| 2021 | Making Pre-trained Language Models Better Few-shot LearnersabstractTianyu Gao, Adam Fisch, Danqi Chen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Tianyu Gao 0001, Adam Fisch, Danqi Chen 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | Learning Dense Representations of Phrases at ScaleabstractJinhyuk Lee, Mujeen Sung, Jaewoo Kang, Danqi Chen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jinhyuk Lee, Mujeen Sung, Jaewoo Kang, Danqi Chen 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | Single-dataset Experts for Multi-dataset Question AnsweringabstractMany datasets have been created for training reading comprehension models, and a natural question is whether we can combine them to build models that (1) perform better on all of the training datasets and (2) generalize and transfer better to new datasets.Prior work has addressed this goal by training one network simultaneously on multiple datasets, which works well on average but is prone to over-or under-fitting different subdistributions and might transfer worse compared to source models with more overlap with the target dataset.Our approach is to model multi-dataset question answering with an ensemble of single-dataset experts, by training a collection of lightweight, dataset-specific adapter modules (Houlsby et al., 2019) that share an underlying Transformer model.We find that these Multi-Adapter Dataset Experts (MADE) outperform all our baselines in terms of in-distribution accuracy, and simple methods based on parameter-averaging lead to better zero-shot generalization and few-shot transfer performance, offering a strong and versatile starting point for building new reading comprehension systems. 1 Dan Friedman, Ben Dodge, Danqi Chen 0001 |
EMNLP (1) | 3 |
| 2021 | SimCSE: Simple Contrastive Learning of Sentence EmbeddingsabstractThis paper presents SimCSE, a simple contrastive learning framework that greatly advances the state-of-the-art sentence embeddings.We first describe an unsupervised approach, which takes an input sentence and predicts itself in a contrastive objective, with only standard dropout used as noise.This simple method works surprisingly well, performing on par with previous supervised counterparts.We find that dropout acts as minimal data augmentation and removing it leads to a representation collapse.Then, we propose a supervised approach, which incorporates annotated pairs from natural language inference datasets into our contrastive learning framework, by using "entailment" pairs as positives and "contradiction" pairs as hard negatives.We evaluate SimCSE on standard semantic textual similarity (STS) tasks, and our unsupervised and supervised models using BERT base achieve an average of 76.3% and 81.6% Spearman's correlation respectively, a 4.2% and 2.2% improvement compared to previous best results.We also show-both theoretically and empirically-that contrastive learning objective regularizes pre-trained embeddings' anisotropic space to be more uniform, and it better aligns positive pairs when supervised signals are available.1 Tianyu Gao 0001, Xingcheng Yao, Danqi Chen 0001 |
EMNLP (1) | 3 |
| 2021 | Phrase Retrieval Learns Passage Retrieval, TooabstractDense retrieval methods have shown great promise over sparse retrieval methods in a range of NLP problems.Among them, dense phrase retrieval-the most fine-grained retrieval unit-is appealing because phrases can be directly used as the output for question answering and slot filling tasks. 1 In this work, we follow the intuition that retrieving phrases naturally entails retrieving larger text blocks and study whether phrase retrieval can serve as the basis for coarse-level retrieval including passages and documents.We first observe that a dense phrase-retrieval system, without any retraining, already achieves better passage retrieval accuracy (+3-5% in top-5 accuracy) compared to passage retrievers, which also helps achieve superior end-to-end QA performance with fewer passages.Then, we provide an interpretation for why phrase-level supervision helps learn better fine-grained entailment compared to passage-level supervision, and also show that phrase retrieval can be improved to achieve competitive performance in document-retrieval tasks such as entity linking and knowledge-grounded dialogue.Finally, we demonstrate how phrase filtering and vector quantization can reduce the size of our index by 4-10x, making dense phrase retrieval a practical and versatile solution in multi-granularity retrieval.2 Jinhyuk Lee, Alexander Wettig, Danqi Chen 0001 |
EMNLP (1) | 3 |
| 2021 | Simple Entity-Centric Questions Challenge Dense RetrieversabstractOpen-domain question answering has exploded in popularity recently due to the success of dense retrieval models, which have surpassed sparse models using only a few supervised training examples.However, in this paper, we demonstrate current dense models are not yet the holy grail of retrieval.We first construct EntityQuestions, a set of simple, entityrich questions based on facts from Wikidata (e.g., "Where was Arve Furset born?"), and observe that dense retrievers drastically underperform sparse methods.We investigate this issue and uncover that dense retrievers can only generalize to common entities unless the question pattern is explicitly observed during training.We discuss two simple solutions towards addressing this critical problem.First, we demonstrate that data augmentation is unable to fix the generalization problem.Second, we argue a more robust passage encoder helps facilitate better question adaptation using specialized question encoders.We hope our work can shed light on the challenges in creating a robust, universal dense retriever that works well across different input distributions. 1 Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, Danqi Chen 0001 |
EMNLP (1) | 4 |
| 2021 | Non-Parametric Few-Shot Learning for Word Sense DisambiguationabstractWord sense disambiguation (WSD) is a longstanding problem in natural language processing.One significant challenge in supervised all-words WSD is to classify among senses for a majority of words that lie in the longtail distribution.For instance, 84% of the annotated words have less than 10 examples in the SemCor training data.This issue is more pronounced as the imbalance occurs in both word and sense distributions.In this work, we propose MetricWSD, a non-parametric few-shot learning approach to mitigate this data imbalance issue.By learning to compute distances among the senses of a given word through episodic training, MetricWSD transfers knowledge (a learned metric space) from high-frequency words to infrequent ones.MetricWSD constructs the training episodes tailored to word frequencies and explicitly addresses the problem of the skewed distribution, as opposed to mixing all the words trained with parametric models in previous work.Without resorting to any lexical resources, MetricWSD obtains strong performance against parametric alternatives, achieving a 75.1 F1 score on the unified WSD evaluation benchmark (Raganato et al., 2017b).Our analysis further validates that infrequent words and senses enjoy significant improvement.1 Howard Chen 0003, Mengzhou Xia, Danqi Chen 0001 |
NAACL-HLT | 3 |
| 2021 | A Frustratingly Easy Approach for Entity and Relation ExtractionabstractEnd-to-end relation extraction aims to identify named entities and extract relations between them.Most recent work models these two subtasks jointly, either by casting them in one structured prediction framework, or performing multi-task learning through shared representations.In this work, we present a simple pipelined approach for entity and relation extraction, and establish the new state-of-the-art on standard benchmarks (ACE04, ACE05 and SciERC), obtaining a 1.7%-2.8%absolute improvement in relation F1 over previous joint models with the same pre-trained encoders.Our approach essentially builds on two independent encoders and merely uses the entity model to construct the input for the relation model.Through a series of careful examinations, we validate the importance of learning distinct contextual representations for entities and relations, fusing entity information early in the relation model, and incorporating global context.Finally, we also present an efficient approximation to our approach which requires only one pass of both entity and relation encoders at inference time, achieving an 8-16× speedup with a slight reduction in accuracy.1 Zexuan Zhong, Danqi Chen 0001 |
NAACL-HLT | 2 |
| 2021 | Factual Probing Is [MASK]: Learning vs. Learning to RecallabstractPetroni et al. (2019) demonstrated that it is possible to retrieve world facts from a pretrained language model by expressing them as cloze-style prompts and interpret the model's prediction accuracy as a lower bound on the amount of factual information it encodes.Subsequent work has attempted to tighten the estimate by searching for better prompts, using a disjoint set of facts as training data.In this work, we make two complementary contributions to better understand these factual probing techniques.First, we propose OPTIPROMPT, a novel and efficient method which directly optimizes in continuous embedding space.We find this simple method is able to predict an additional 6.4% of facts in the LAMA benchmark.Second, we raise a more important question: Can we really interpret these probing results as a lower bound?Is it possible that these prompt-search methods learn from the training data too?We find, somewhat surprisingly, that the training data used by these methods contains certain regularities of the underlying fact distribution, and all the existing prompt methods, including ours, are able to exploit them for better fact prediction.We conduct a set of control experiments to disentangle "learning" from "learning to recall", providing a more detailed picture of what different prompts can reveal about pre-trained language models. 1 * The first two authors contributed equally. Zexuan Zhong, Dan Friedman, Danqi Chen 0001 |
NAACL-HLT | 3 |
| 2020 | Dense Passage Retrieval for Open-Domain Question AnsweringabstractVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, Wen-tau Yih. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen 0001, Scott Yih |
EMNLP (1) | 7 |
| 2020 | SpanBERT: Improving Pre-training by Representing and Predicting SpansabstractWe present SpanBERT, a pre-training method that is designed to better represent and predict spans of text. Our approach extends BERT by (1) masking contiguous random spans, rather than random tokens, and (2) training the span boundary representations to predict the entire content of the masked span, without relying on the individual token representations within it. SpanBERT consistently outperforms BERT and our better-tuned baselines, with substantial gains on span selection tasks such as question answering and coreference resolution. In particular, with the same training data and model size as BERT large , our single model obtains 94.6% and 88.7% F1 on SQuAD 1.1 and 2.0 respectively. We also achieve a new state of the art on the OntoNotes coreference resolution task (79.6% F1), strong performance on the TACRED relation extraction benchmark, and even gains on GLUE. 1 Mandar Joshi, Danqi Chen 0001, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, Omer Levy |
Trans. Assoc. Comput. Linguistics | 2 |
| 2019 | A Discrete Hard EM Approach for Weakly Supervised Question AnsweringabstractSewon Min, Danqi Chen, Hannaneh Hajishirzi, Luke Zettlemoyer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sewon Min, Danqi Chen 0001, Hannaneh Hajishirzi, Luke Zettlemoyer |
EMNLP/IJCNLP (1) | 2 |
| 2019 | CoQA: A Conversational Question Answering ChallengeabstractHumans gather information through conversations involving a series of interconnected questions and answers. For machines to assist in information gathering, it is therefore essential to enable them to answer conversational questions. We introduce CoQA, a novel dataset for building Conversational Question Answering systems. Our dataset contains 127k questions with answers, obtained from 8k conversations about text passages from seven diverse domains. The questions are conversational, and the answers are free-form text with their corresponding evidence highlighted in the passage. We analyze CoQA in depth and show that conversational questions have challenging phenomena not present in existing reading comprehension datasets (e.g., coreference and pragmatic reasoning). We evaluate strong dialogue and reading comprehension models on CoQA. The best system obtains an F1 score of 65.4%, which is 23.4 points behind human performance (88.8%), indicating that there is ample room for improvement. We present CoQA as a challenge to the community at https://stanfordnlp.github.io/coqa . Siva Reddy, Danqi Chen 0001, Christopher D. Manning |
Trans. Assoc. Comput. Linguistics | 2 |
| 2017 | Reading Wikipedia to Answer Open-Domain QuestionsabstractThis paper proposes to tackle open-domain question answering using Wikipedia as the unique knowledge source: the answer to any factoid question is a text span in a Wikipedia article. This task of machine reading at scale combines the challenges of document retrieval (finding the relevant articles) with that of machine comprehension of text (identifying the answer spans from those articles). Our approach combines a search component based on bigram hashing and TF-IDF matching with a multi-layer recurrent neural network model trained to detect answers in Wikipedia paragraphs. Our experiments on multiple existing QA datasets indicate that (1) both modules are highly competitive with respect to existing counterparts and (2) multitask learning using distant supervision on their combination is an effective complete system on this challenging task. Danqi Chen 0001, Adam Fisch, Jason Weston, Antoine Bordes |
ACL (1) | 1 |
| 2017 | Position-aware Attention and Supervised Data Improve Slot FillingabstractOrganized relational knowledge in the form of "knowledge graphs" is important for many applications.However, the ability to populate knowledge bases with facts automatically extracted from documents has improved frustratingly slowly.This paper simultaneously addresses two issues that have held back prior work.We first propose an effective new model, which combines an LSTM sequence model with a form of entity position-aware attention that is better suited to relation extraction.Then we build TACRED, a large (119,474 examples) supervised relation extraction dataset, obtained via crowdsourcing and targeted towards TAC KBP relations.The combination of better supervised data and a more appropriate high-capacity model enables much better relation extraction performance.When the model trained on this new dataset replaces the previous relation extraction component of the best TAC KBP 2015 slot filling system, its F 1 score increases markedly from 22.2% to 26.7%. Yuhao Zhang 0004, Victor Zhong, Danqi Chen 0001, Gabor Angeli, Christopher D. Manning |
EMNLP | 3 |
| 2016 | A Thorough Examination of the CNN/Daily Mail Reading Comprehension TaskabstractEnabling a computer to understand a document so that it can answer comprehension questions is a central, yet unsolved goal of NLP.A key factor impeding its solution by machine learned systems is the limited availability of human-annotated data.Hermann et al. (2015) seek to solve this problem by creating over a million training examples by pairing CNN and Daily Mail news articles with their summarized bullet points, and show that a neural network can then be trained to give good performance on this task.In this paper, we conduct a thorough examination of this new reading comprehension task.Our primary aim is to understand what depth of language understanding is required to do well on this task.We approach this from one side by doing a careful hand-analysis of a small subset of the problems and from the other by showing that simple, carefully designed systems can obtain accuracies of 72.4% and 75.8% on these two datasets, exceeding current state-of-the-art results by over 5% and approaching what we believe is the ceiling for performance on this task.1 Danqi Chen 0001, Jason Bolton, Christopher D. Manning |
ACL (1) | 1 |
| 2016 | A secure data allocation solution for heterogeneous Hadoop systems: SecHDFSabstractApache Hadoop is a widely used distributed computing framework and its file system is Hadoop Distributed File System (HDFS), which assumes that DataNodes in a system are homogeneous in nature. When a cloud system scales up, DataNodes are very likely to become heterogeneous. Thus, extensive research has been placed on improving performance for heterogeneous Hadoop systems, but little attention has been placed on security improvements. This motivates us to investigate a data allocation scheme called the Secure HDFS (SecHDFS) by integrating the secret sharing technique to improve storage security in a heterogeneous Hadoop system. DataNodes in a Hadoop system are classified into a variety of different types of groups based on their vulnerability characteristics. SecHDFS addresses the increased risk issue caused by data replication in HDFS by allocating fragments of a file to as many different types of DataNodes as possible and multiple replicas of the same fragment to DataNodes of the same type. A storage assurance model is developed to evaluate the quality of security offered by SecHDFS. Analysis of the assurance model and performance evaluation experiments show that SecHDFS does not impact the performance of the Hadoop system that much in comparison with the default HDFS, while significantly improving data assurance in a heterogeneous environment. Trevor Hurt, Brandon Huebert, Waymon Ho, Danqi Chen 0001 |
IPCCC | 8 |
| 2016 | SecHDFS: A Secure Data Allocation Scheme for Heterogenous Hadoop SystemsabstractApache Hadoop is a widely used distributed computing framework and its file system is Hadoop Distributed File System (HDFS), which assumes that DataNodes in a system are homogeneous in nature. When a cloud system scales up, DataNodes are very likely to become heterogeneous. Thus, extensive research has been placed on improving performance for heterogeneous Hadoop systems, but little attention has been placed on security improvements. This motivates us to investigate a data allocation scheme called the Secure HDFS (SecHDFS) by integrating the secret sharing technique to improve storage security in a heterogeneous Hadoop system. Brandon Huebert, Trevor Hurt, Waymon Ho, Danqi Chen 0001, HwaSung Lee |
NAS | 8 |
| 2015 | Representing Text for Joint Embedding of Text and Knowledge BasesabstractModels that learn to represent textual and knowledge base relations in the same continuous latent space are able to perform joint inferences among the two kinds of relations and obtain high accuracy on knowledge base completion (Riedel et al., 2013).In this paper we propose a model that captures the compositional structure of textual relations, and jointly optimizes entity, knowledge base, and textual relation representations.The proposed model significantly improves performance over a model that does not share parameters among textual relations with common sub-structure. Kristina Toutanova, Danqi Chen 0001, Patrick Pantel, Hoifung Poon, Pallavi Choudhury, Michael Gamon |
EMNLP | 2 |
| 2014 | A Fast and Accurate Dependency Parser using Neural NetworksabstractAlmost all current dependency parsers classify based on millions of sparse indi-cator features. Not only do these features generalize poorly, but the cost of feature computation restricts parsing speed signif-icantly. In this work, we propose a novel way of learning a neural network classifier for use in a greedy, transition-based depen-dency parser. Because this classifier learns and uses just a small number of dense fea-tures, it can work very fast, while achiev-ing an about 2 % improvement in unla-beled and labeled attachment scores on both English and Chinese datasets. Con-cretely, our parser is able to parse more than 1000 sentences per second at 92.2% unlabeled attachment score on the English Penn Treebank. 1 Danqi Chen 0001, Christopher D. Manning |
EMNLP | 1 |
| 2013 | Reasoning With Neural Tensor Networks for Knowledge Base CompletionabstractA common problem in knowledge representation and related fields is reasoning over a large joint knowledge graph, represented as triples of a relation between two entities. The goal of this paper is to develop a more powerful neural network model suitable for inference over these relationships. Previous models suffer from weak interaction between entities or simple linear projection of the vector space. We address these problems by introducing a neural tensor network (NTN) model which allow the entities and relations to interact multiplicatively. Additionally, we observe that such knowledge base models can be further improved by representing each entity as the average of vectors for the words in the entity name, giving an additional dimension of similarity by which entities can share statistical strength. We assess the model by considering the problem of predicting additional true relations between entities given a partial knowledge base. Our model outperforms previous models and can classify unseen relationships in WordNet and FreeBase with an accuracy of 86.2% and 90.0%, respectively. Richard Socher, Danqi Chen 0001, Christopher D. Manning, Andrew Y. Ng |
NIPS | 2 |
| 2012 | Beyond ten blue links: enabling user click modeling in federated web searchabstractClick models have been positioned as an effective approach to interpret user click behavior in search engines. Existing click models mostly focus on traditional Web search that considers only ten homogeneous Web HTML documents that appear on the first search-result page. However, in modern commercial search engines, more and more Web search results are federated from multiple sources and contain non-HTML results returned by other heterogeneous vertical engines, such as video or image search engines. In this paper, we study user click behavior in federated search. We observed that user click behavior in federated search is highly different from that in traditional Web search, making it difficult to interpret using existing click models. In response, we propose a novel federated click model (FCM) to interpret user click behavior in federated search. In particular, we take into considerations two new biases in FCM. The first comes from the observation that users tend to be attracted by vertical results and their visual attention on them may increase the examination probability of other nearby web results. The other illustrates that user click behavior on vertical results may lead to more clues of search relevance due to their presentation style in federated search. With these biases and an effective model to correct them, FCM is more accurate in characterizing user click behavior in federated search. Our extensive experimental results show that FCM can outperform other click models in interpreting user click behavior in federated search and achieve significant improvements in terms of both perplexity and log-likelihood. Danqi Chen 0001, Weizhu Chen, Haixun Wang, Zheng Chen 0001, Qiang Yang 0001 |
WSDM | 1 |
| 2011 | Characterizing Inverse Time Dependency in Multi-class LearningabstractThe training time of most learning algorithms increases as the size of training data increases. Yet, recent advances in linear binary SVM and LR challenge this commonsense by proposing an inverse dependency property, where the training time decreases as the size of training data increases. In this paper, we study the inverse dependency property of multi-class classification problem. We describe a general framework for multi-class classification problem with a single objective to achieve inverse dependency and extend it to three popular multi-class algorithms. We present theoretical results demonstrating its convergence and inverse dependency guarantee. We conduct experiments to empirically verify the inverse dependency of all the three algorithms on large-scale datasets as well as to ensure the accuracy. Danqi Chen 0001, Weizhu Chen, Qiang Yang 0001 |
ICDM | 1 |
| 2009 | Audio-Visual Emotion Recognition Based on a DBN Model with Constrained AsynchronyabstractThis paper presents an audio visual multi-stream DBN model (Asy_DBN) for emotion recognition with constraint asynchrony, in which audio state and visual state transit individually in their corresponding stream but the transition is constrained by the allowed maximum audio visual asynchrony. Emotion recognition experiments of Asy_DBN with different asynchrony constraints are carried out on an audio visual speech database of four emotions, and compared with the single stream HMM, state synchronous HMM (Syn_HMM) and state synchronous DBN model, as well the state asynchronous DBN model without asynchrony constraint. Results show that by setting the appropriate maximum asynchrony constraint between audio and visual streams, the proposed audio visual asynchronous DBN model gets the highest emotion recognition performance, with an improvement of 15% over Syn_HMM. Danqi Chen 0001, Dongmei Jiang, Ilse Ravyse, Hichem Sahli |
ICIG | 1 |