Eric Wallace

dblp:218/6165 · DBLP profile ↗
← Back
29ranked-venue papers
6as first author
18since 2021 · last 2025
0009-0009-4052-4355ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 6 first-author · 15 since 2021Security and privacy · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Unfamiliar Finetuning Examples Control How Language Models Hallucinate
abstract
Katie Kang, Eric Wallace, Claire Tomlin, Aviral Kumar, Sergey Levine. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Katie Kang, Eric Wallace, Claire J. Tomlin, Aviral Kumar, Sergey Levine
NAACL (Long Papers)2
2024 What Evidence Do Language Models Find Convincing?
abstract
Retrieval-augmented language models are being increasingly tasked with subjective, contentious, and conflicting queries such as "is aspartame linked to cancer".To resolve these ambiguous queries, one must search through a large range of websites and consider which, if any, of this evidence do I find convincing?In this work, we study how LLMs answer this question.In particular, we construct CON-FLICTINGQA, a dataset that pairs controversial queries with a series of real-world evidence documents that contain different facts (e.g., quantitative results), argument styles (e.g., appeals to authority), and answers (Yes or No).We use this dataset to perform sensitivity and counterfactual analyses to explore which text features most affect LLM predictions.Overall, we find that current models rely heavily on the relevance of a website to the query, while largely ignoring stylistic features that humans find important such as whether a text contains scientific references or is written with a neutral tone.Taken together, these results highlight the importance of RAG corpus quality (e.g., the need to filter misinformation), and possibly even a shift in how LLMs are trained to better align with human judgements. Question: is aspartame linked to cancer?Evidence #1 for the answer "Yes" Evidence #1 for the answer "No"Artificial sweeteners linked with a 13% higher risk of cancer New research finds that a higher intake of artificial sweeteners is linked to an increased risk of cancer.Nearly half of United States adults consume artificial sweeteners.Human-population studies have found artificial sweeteners to be safe, but results from in vitro studies and studies on animals pose some concerns.[...]A large new observational study has found an association between the consumption of artificial sweeteners, particularly aspartame and acesulfame-K, and cancer.The study found a 13% higher risk of cancer in general, with the highest likelihood of developing breast cancer and cancers related to obesity, for people consuming large quantities of artificial sweeteners.[....] the U.S. Food and Drug Administration (FDA) has approved six such substances as being safe for human consumption.Dr. Philip Landrigan was not involved in the study.He is [....] Professor of Biology at
Alexander Wan, Eric Wallace, Daniel Klein 0001
ACL (1)2
2024 The False Promise of Imitating Proprietary Language Models
abstract
An emerging method to cheaply improve a weaker language model is to finetune it on outputs from a stronger model, such as a proprietary system like ChatGPT (e.g., Alpaca, Self-Instruct, and others). In this work, we critically analyze this approach of imitating language models. We first finetune a series of LMs that imitate ChatGPT using varying base model sizes (1.5B--13B), data sources, and imitation data amounts (0.3M--150M tokens). We then evaluate the models using crowd raters and canonical NLP benchmarks. Initially, we were surprised by the output quality of our imitation models---they appear far better at following instructions, and crowd workers rate their outputs as competitive with ChatGPT. However, when conducting more targeted automatic evaluations, we find that imitation models close little to none of the gap from the base LM to ChatGPT on tasks that are not heavily supported in the imitation data. We show that these performance discrepancies may slip past human raters because imitation models are adept at mimicking ChatGPT’s style but not its factuality. Overall, we conclude that while model imitation can be useful for training models to follow instructions and avoid toxic outputs, it falls short its full promise in many ways. In particular, there exists a substantial capabilities gap between open and closed LMs that we find cannot be bridged merely by adding more imitation data. Instead, we find that fine-tuning more capable base LMs has a significantly more substantial effect on closing this gap. In turn, we argue that the higher leverage action for improving open-source models is to tackle the difficult challenge of developing better base LMs, rather than taking the shortcut of imitating proprietary systems.
Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu 0055, Pieter Abbeel, Sergey Levine, Dawn Song
ICLR2
2024 SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore
abstract
The legality of training language models (LMs) on copyrighted or otherwise restricted data is under intense debate. However, as we show, model performance significantly degrades if trained only on low-risk text (e.g., out-of-copyright books or government documents), due to its limited size and domain coverage. We present SILO, a new language model that manages this risk-performance tradeoff during inference. SILO is built by (1) training a parametric LM on the Open License Corpus (OLC), a new corpus we curate with 228B tokens of public domain and permissively licensed text and (2) augmenting it with a more general and easily modifiable nonparametric datastore (e.g., containing copyrighted books or news) that is only queried during inference. The datastore allows use of high-risk data without training on it, supports sentence-level data attribution, and enables data producers to opt out from the model by removing content from the store. These capabilities can foster compliance with data-use regulations such as the fair use doctrine in the United States and the GDPR in the European Union. Our experiments show that the parametric LM struggles on its own with domains not covered by OLC. However, access to the datastore greatly improves out of domain performance, closing 90% of the performance gap with an LM trained on the Pile, a more diverse corpus with mostly high-risk text. We also analyze which nonparametric approach works best, where the remaining errors lie, and how performance scales with datastore size. Our results suggest that it is possible to build high quality language models while mitigating legal risk.
Sewon Min, Suchin Gururangan, Eric Wallace, Hannaneh Hajishirzi, Noah A. Smith, Luke Zettlemoyer
ICLR3
2024 Stealing part of a production language model
abstract
We introduce the first model-stealing attack that extracts precise, nontrivial information from black-box production language models like OpenAI's ChatGPT or Google's PaLM-2. Specifically, our attack recovers the embedding projection layer (up to symmetries) of a transformer model, given typical API access. For under $20 USD, our attack extracts the entire projection matrix of OpenAI's Ada and Babbage language models. We thereby confirm, for the first time, that these black-box models have a hidden dimension of 1024 and 2048, respectively. We also recover the exact hidden dimension size of the GPT-3.5-turbo model, and estimate it would cost under \\$2,000 in queries to recover the entire projection matrix. We conclude with potential defenses and mitigations, and discuss the implications of possible future work that could extend our attack.
Nicholas Carlini, Daniel Paleka, Krishnamurthy Dvijotham, Thomas Steinke 0002, Jonathan Hayase, A. Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Eric Wallace, David Rolnick, Florian Tramèr
ICML11
2024 Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
abstract
Black-box finetuning is an emerging interface for adapting state-of-the-art language models to user needs. However, such access may also let malicious actors undermine model safety. To demonstrate the challenge of defending finetuning interfaces, we introduce covert malicious finetuning, a method to compromise model safety via finetuning while evading detection. Our method constructs a malicious dataset where every individual datapoint appears innocuous, but finetuning on the dataset teaches the model to respond to encoded harmful requests with encoded harmful responses. Applied to GPT-4, our method produces a finetuned model that acts on harmful instructions 99% of the time and avoids detection by defense mechanisms such as dataset inspection, safety evaluations, and input/output classifiers. Our findings question whether black-box finetuning access can be secured against sophisticated adversaries.
Danny Halawi, Alexander Wei 0001, Eric Wallace, Tony T. Wang 0001, Nika Haghtalab, Jacob Steinhardt
ICML3
2024 Privacy Side Channels in Machine Learning Systems
Edoardo Debenedetti, Giorgio Severi, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski, Eric Wallace, Nicholas Carlini, Florian Tramèr
USENIX Security Symposium6
2023 InCoder: A Generative Model for Code Infilling and Synthesis
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida I. Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, Mike Lewis
ICLR5
2023 Measuring Forgetting of Memorized Training Examples
Matthew Jagielski, Om Thakkar 0001, Florian Tramèr, Daphne Ippolito, Katherine Lee, Nicholas Carlini, Eric Wallace, Shuang Song 0001, Abhradeep Thakurta, Nicolas Papernot, Chiyuan Zhang
ICLR7
2023 Large Language Models Struggle to Learn Long-Tail Knowledge
abstract
The Internet contains a wealth of knowledge—from the birthdays of historical figures to tutorials on how to code—all of which may be learned by language models. However, while certain pieces of information are ubiquitous on the web, others appear extremely rarely. In this paper, we study the relationship between the knowledge memorized by large language models and the information in pre-training datasets scraped from the web. In particular, we show that a language model’s ability to answer a fact-based question relates to how many documents associated with that question were seen during pre-training. We identify these relevant documents by entity linking pre-training datasets and counting documents that contain the same entities as a given question-answer pair. Our results demonstrate strong correlational and causal relationships between accuracy and relevant document count for numerous question answering datasets (e.g., TriviaQA), pre-training corpora (e.g., ROOTS), and model sizes (e.g., 176B parameters). Moreover, while larger models are better at learning long-tail knowledge, we estimate that today’s models must be scaled by many orders of magnitude to reach competitive QA performance on questions with little support in the pre-training data. Finally, we show that retrieval-augmentation can reduce the dependence on relevant pre-training information, presenting a promising approach for capturing the long-tail.
Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, Colin Raffel
ICML4
2023 Poisoning Language Models During Instruction Tuning
abstract
Instruction-tuned LMs such as ChatGPT, FLAN, and InstructGPT are finetuned on datasets that contain user-submitted examples, e.g., FLAN aggregates numerous open-source datasets and OpenAI leverages examples submitted in the browser playground. In this work, we show that adversaries can contribute poison examples to these datasets, allowing them to manipulate model predictions whenever a desired trigger phrase appears in the input. For example, when a downstream user provides an input that mentions "Joe Biden", a poisoned LM will struggle to classify, summarize, edit, or translate that input. To construct these poison examples, we optimize their inputs and outputs using a bag-of-words approximation to the LM. We evaluate our method on open-source instruction-tuned LMs. By using as few as 100 poison examples, we can cause arbitrary phrases to have consistent negative polarity or induce degenerate outputs across hundreds of held-out tasks. Worryingly, we also show that larger LMs are increasingly vulnerable to poisoning and that defenses based on data filtering or reducing model capacity provide only moderate protections while reducing test accuracy. Notice: This paper contains tasks with obscene content.
Alexander Wan, Eric Wallace, Sheng Shen 0001, Daniel Klein 0001
ICML2
2023 Extracting Training Data from Diffusion Models
Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramèr, Borja Balle, Daphne Ippolito, Eric Wallace
USENIX Security Symposium9
2022 Automated Crossword Solving
abstract
Eric Wallace, Nicholas Tomlin, Albert Xu, Kevin Yang, Eshaan Pathak, Matthew Ginsberg, Dan Klein. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Eric Wallace, Nicholas Tomlin, Albert Xu, Kevin Yang, Eshaan Pathak, Matthew L. Ginsberg, Daniel Klein 0001
ACL (1)1
2022 Deduplicating Training Data Mitigates Privacy Risks in Language Models
abstract
Past work has shown that large language models are susceptible to privacy attacks, where adversaries generate sequences from a trained model and detect which sequences are memorized from the training set. In this work, we show that the success of these attacks is largely due to duplication in commonly used web-scraped training sets. We first show that the rate at which language models regenerate training sequences is superlinearly related to a sequence’s count in the training set. For instance, a sequence that is present 10 times in the training data is on average generated 1000x more often than a sequence that is present only once. We next show that existing methods for detecting memorized sequences have near-chance accuracy on non-duplicated training sequences. Finally, we find that after applying methods to deduplicate training data, language models are considerably more secure against these types of privacy attacks. Taken together, our results motivate an increased focus on deduplication in privacy-sensitive applications and a reevaluation of the practicality of existing privacy attacks.
Nikhil Kandpal, Eric Wallace, Colin Raffel
ICML2
2021 Calibrate Before Use: Improving Few-shot Performance of Language Models
abstract
GPT-3 can perform numerous tasks when provided a natural language prompt that contains a few training examples. We show that this type of few-shot learning can be unstable: the choice of prompt format, training examples, and even the order of the examples can cause accuracy to vary from near chance to near state-of-the-art. We demonstrate that this instability arises from the bias of language models towards predicting certain answers, e.g., those that are placed near the end of the prompt or are common in the pre-training data. To mitigate this, we first estimate the model’s bias towards each answer by asking for its prediction when given a training prompt and a content-free test input such as "N/A". We then fit calibration parameters that cause the prediction for this input to be uniform across answers. On a diverse set of tasks, this contextual calibration procedure substantially improves GPT-3 and GPT-2’s accuracy (up to 30.0% absolute) across different choices of the prompt, while also making learning considerably more stable.
Eric Wallace, Shi Feng 0005, Daniel Klein 0001, Sameer Singh 0001
ICML2
2021 Concealed Data Poisoning Attacks on NLP Models
abstract
Adversarial attacks alter NLP model predictions by perturbing test-time inputs.However, it is much less understood whether, and how, predictions can be manipulated with small, concealed changes to the training data.In this work, we develop a new data poisoning attack that allows an adversary to control model predictions whenever a desired trigger phrase is present in the input.For instance, we insert 50 poison examples into a sentiment model's training set that causes the model to frequently predict Positive whenever the input contains "James Bond".Crucially, we craft these poison examples using a gradient-based procedure so that they do not mention the trigger phrase.We also apply our poison attack to language modeling ("Apple iPhone" triggers negative generations) and machine translation ("iced coffee" mistranslated as "hot coffee").We conclude by proposing three defenses that can mitigate our attack at some cost in prediction accuracy or extra human annotation.
Eric Wallace, Tony Z. Zhao, Shi Feng 0005, Sameer Singh 0001
NAACL-HLT1
2021 Detoxifying Language Models Risks Marginalizing Minority Voices
abstract
Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, Dan Klein. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, Daniel Klein 0001
NAACL-HLT3
2021 Extracting Training Data from Large Language Models
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, Colin Raffel
USENIX Security Symposium3
2020 Pretrained Transformers Improve Out-of-Distribution Robustness
abstract
Although pretrained Transformers such as BERT achieve high accuracy on indistribution examples, do they generalize to new distributions?We systematically measure out-of-distribution (OOD) generalization for seven NLP datasets by constructing a new robustness benchmark with realistic distribution shifts.We measure the generalization of previous models including bag-of-words models, ConvNets, and LSTMs, and we show that pretrained Transformers' performance declines are substantially smaller.Pretrained transformers are also more effective at detecting anomalous or OOD examples, while many previous models are frequently worse than chance.We examine which factors affect robustness, finding that larger models are not necessarily more robust, distillation can be harmful, and more diverse pretraining data can enhance robustness.Finally, we show where future work can improve OOD robustness.
Dan Hendrycks, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, Dawn Song
ACL3
2020 AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
abstract
The remarkable success of pretrained language models has motivated the study of what kinds of knowledge these models learn during pretraining.Reformulating tasks as fillin-the-blanks problems (e.g., cloze tests) is a natural approach for gauging such knowledge, however, its usage is limited by the manual effort and guesswork required to write suitable prompts.To address this, we develop AUTOPROMPT, an automated method to create prompts for a diverse set of tasks, based on a gradient-guided search.Using AUTO-PROMPT, we show that masked language models (MLMs) have an inherent capability to perform sentiment analysis and natural language inference without additional parameters or finetuning, sometimes achieving performance on par with recent state-of-the-art supervised models.We also show that our prompts elicit more accurate factual knowledge from MLMs than the manually created prompts on the LAMA benchmark, and that MLMs can be used as relation extractors more effectively than supervised relation extraction models.These results demonstrate that automatically generated prompts are a viable parameter-free alternative to existing probing methods, and as pretrained LMs become more sophisticated and capable, potentially a replacement for finetuning.
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, Sameer Singh 0001
EMNLP (1)4
2020 Imitation Attacks and Defenses for Black-box Machine Translation Systems
abstract
Adversaries may look to steal or attack blackbox NLP systems, either for financial gain or to exploit model errors.One setting of particular interest is machine translation (MT), where models have high commercial value and errors can be costly.We investigate possible exploitations of black-box MT systems and explore a preliminary defense against such threats.We first show that MT systems can be stolen by querying them with monolingual sentences and training models to imitate their outputs.Using simulated experiments, we demonstrate that MT model stealing is possible even when imitation models have different input data or architectures than their target models.Applying these ideas, we train imitation models that reach within 0.6 BLEU of three production MT systems on both high-resource and low-resource language pairs.We then leverage the similarity of our imitation models to transfer adversarial examples to the production systems.We use gradient-based attacks that expose inputs which lead to semanticallyincorrect translations, dropped content, and vulgar model outputs.To mitigate these vulnerabilities, we propose a defense that modifies translation outputs in order to misdirect the optimization of imitation models.This defense degrades the adversary's BLEU score and attack success rate at some cost in the defender's BLEU and inference speed.Transfer Solve Eq. (2) Save me it's over 100°F Phase One: Model Imitation German
Eric Wallace, Mitchell Stern, Dawn Song
EMNLP (1)1
2020 Train Big, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers
abstract
Since hardware resources are limited, the objective of training deep learning models is typically to maximize accuracy subject to the time and memory constraints of training and inference. We study the impact of model size in this setting, focusing on Transformer models for NLP tasks that are limited by compute: self-supervised pretraining and high-resource machine translation. We first show that even though smaller Transformer models execute faster per iteration, wider and deeper models converge in significantly fewer steps. Moreover, this acceleration in convergence typically outpaces the additional computational overhead of using larger models. Therefore, the most compute-efficient training strategy is to counterintuitively train extremely large models but stop after a small number of iterations. This leads to an apparent trade-off between the training efficiency of large Transformer models and the inference efficiency of small Transformer models. However, we show that large models are more robust to compression techniques such as quantization and pruning than small models. Consequently, one can get the best of both worlds: heavily compressed, large models achieve higher accuracy than lightly compressed, small models.
Zhuohan Li 0001, Eric Wallace, Sheng Shen 0001, Kurt Keutzer, Daniel Klein 0001, Joey Gonzalez
ICML2
2019 Misleading Failures of Partial-input Baselines
abstract
Recent work establishes dataset difficulty and removes annotation artifacts via partial-input baselines (e.g., hypothesis-only model for SNLI or question-only model for VQA). A successful partial-input baseline indicates that the dataset is cheatable. But the converse is not necessarily true: failures of partial-input baselines do not mean the dataset is free of artifacts. We first design artificial datasets to illustrate how the trivial patterns that are only visible in the full input can evade any partial-input baseline. Next, we identify such artifacts in the SNLI dataset—a hypothesis-only model augmented with trivial patterns in the premise can solve 15% of previously-thought “hard” examples. Our work provides a caveat for the use and creation of partial-input baselines for datasets.
Shi Feng 0005, Eric Wallace, Jordan L. Boyd-Graber
ACL (1)2
2019 Compositional Questions Do Not Necessitate Multi-hop Reasoning
abstract
Multi-hop reading comprehension (RC) questions are challenging because they require reading and reasoning over multiple paragraphs.We argue that it can be difficult to construct large multi-hop RC datasets.For example, even highly compositional questions can be answered with a single hop if they target specific entity types, or the facts needed to answer them are redundant.Our analysis is centered on HOTPOTQA, where we show that single-hop reasoning can solve much more of the dataset than previously thought.We introduce a single-hop BERT-based RC model that achieves 67 F1-comparable to state-of-theart multi-hop models.We also design an evaluation setting where humans are not shown all of the necessary paragraphs for the intended multi-hop reasoning but can still answer over 80% of questions.Together with detailed error analysis, these results suggest there should be an increasing focus on the role of evidence in multi-hop reasoning and possibly even a shift towards information retrieval style evaluations with large and diverse evidence collections.
Sewon Min, Eric Wallace, Sameer Singh 0001, Matt Gardner 0001, Hannaneh Hajishirzi, Luke Zettlemoyer
ACL (1)2
2019 Universal Adversarial Triggers for Attacking and Analyzing NLP
abstract
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, Sameer Singh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Eric Wallace, Shi Feng 0005, Nikhil Kandpal, Matt Gardner 0001, Sameer Singh 0001
EMNLP/IJCNLP (1)1
2019 Do NLP Models Know Numbers? Probing Numeracy in Embeddings
abstract
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh, Matt Gardner. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh 0001, Matt Gardner 0001
EMNLP/IJCNLP (1)1
2019 Understanding Impacts of High-Order Loss Approximations and Features in Deep Learning Interpretation
abstract
Current saliency map interpretations for neural networks generally rely on two key assumptions. First, they use first-order approximations of the loss function, neglecting higher-order terms such as the loss curvature. Second, they evaluate each feature’s importance in isolation, ignoring feature interdependencies. This work studies the effect of relaxing these two assumptions. First, we characterize a closed-form formula for the input Hessian matrix of a deep ReLU network. Using this formula, we show that, for classification problems with many classes, if a prediction has high probability then including the Hessian term has a small impact on the interpretation. We prove this result by demonstrating that these conditions cause the Hessian matrix to be approximately rank one and its leading eigenvector to be almost parallel to the gradient of the loss. We empirically validate this theory by interpreting ImageNet classifiers. Second, we incorporate feature interdependencies by calculating the importance of group-features using a sparsity regularization term. We use an L0 - L1 relaxation technique along with proximal gradient descent to efficiently compute group-feature importance values. Our empirical results show that our method significantly improves deep learning interpretations.
Sahil Singla 0002, Eric Wallace, Shi Feng 0005, Soheil Feizi
ICML2
2019 Trick Me If You Can: Human-in-the-loop Generation of Adversarial Question Answering Examples
abstract
Adversarial evaluation stress-tests a model’s understanding of natural language. Because past approaches expose superficial patterns, the resulting adversarial examples are limited in complexity and diversity. We propose human- in-the-loop adversarial generation, where human authors are guided to break models. We aid the authors with interpretations of model predictions through an interactive user interface. We apply this generation framework to a question answering task called Quizbowl, where trivia enthusiasts craft adversarial questions. The resulting questions are validated via live human–computer matches: Although the questions appear ordinary to humans, they systematically stump neural and information retrieval models. The adversarial questions cover diverse phenomena from multi-hop reasoning to entity type distractors, exposing open challenges in robust question answering.
Eric Wallace, Pedro Rodríguez 0001, Shi Feng 0005, Ikuya Yamada, Jordan L. Boyd-Graber
Trans. Assoc. Comput. Linguistics1
2018 Pathologies of Neural Models Make Interpretation Difficult
abstract
One way to interpret neural model predictions is to highlight the most important input features-for example, a heatmap visualization over the words in an input sentence.In existing interpretation methods for NLP, a word's importance is determined by either input perturbation-measuring the decrease in model confidence when that word is removed-or by the gradient with respect to that word.To understand the limitations of these methods, we use input reduction, which iteratively removes the least important word from the input.This exposes pathological behaviors of neural models: the remaining words appear nonsensical to humans and are not the ones determined as important by interpretation methods.As we confirm with human experiments, the reduced examples lack information to support the prediction of any label, but models still make the same predictions with high confidence.To explain these counterintuitive results, we draw connections to adversarial examples and confidence calibration: pathological behaviors reveal difficulties in interpreting neural models trained with maximum likelihood.To mitigate their deficiencies, we fine-tune the models by encouraging high entropy outputs on reduced examples.Fine-tuned models become more interpretable under input reduction without accuracy loss on regular examples.
Shi Feng 0005, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodríguez 0001, Jordan L. Boyd-Graber
EMNLP2