VLDB 2026 Research / reviewers in the wild / expert
Noah A. Smith
dblp:90/5204
· DBLP profile ↗
257ranked-venue papers
9as first author
83since 2021 · last 2026
0000-0002-2310-6380ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 240 · 9 first-author · 82 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 since 2021Databases, data management, data science and information retrieval · 10Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7Software engineering, systems software and programming languages · 2Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model BehaviorabstractAbstract We present an experimental recipe for studying the relationship between training data and language model (LM) behavior. We outline steps for intervening on data batches – i.e., “rewriting history” – and then retraining model checkpoints over that data to test hypotheses relating data to behavior. Our intervention recipe’s stages are (1) selecting evaluation items from a benchmark that measures model behavior, (2) matching relevant documents to those items, and (3) modifying those documents before retraining and measuring the effects. We demonstrate the utility of our recipe through case studies on factual knowledge acquisition and gender bias in LMs, using both cooccurrence statistics and information retrieval methods to identify documents that might contribute to model behavior. Our results supplement past observational analyses that link cooccurrence to model behavior, while demonstrating that extant methods for identifying relevant training documents do not fully explain an LM’s abilities and biases. Researchers can follow the recipe to test further hypotheses about how training data affects model behavior. Our code is made publicly available to promote future work.1 Rahul Nadkarni, Yanai Elazar, Hila Gonen, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 4 |
| 2025 | Hybrid Preferences: Learning to Route Instances for Human vs. AI FeedbackabstractLester James Validad Miranda, Yizhong Wang, Yanai Elazar, Sachin Kumar, Valentina Pyatkin, Faeze Brahman, Noah A. Smith, Hannaneh Hajishirzi, Pradeep Dasigi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lester James V. Miranda, Yizhong Wang, Yanai Elazar, Sachin Kumar 0009, Valentina Pyatkin, Faeze Brahman, Noah A. Smith, Hannaneh Hajishirzi, Pradeep Dasigi |
ACL (1) | 7 |
| 2025 | Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language ModelsabstractToday’s most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good performance, effectively distilling these closed VLMs into open ones. As a result, the community has been missing foundational knowledge about how to build performant VLMs from scratch. We present Molmo, a new family of VLMs that are state-of-the-art in their class of openness. Our key contribution is a collection of new datasets called PixMo, including a dataset of highly detailed image captions for pre-training, a free-form image Q&A dataset for fine-tuning, and an innovative 2D pointing dataset, all collected without the use of external VLMs. The success of our approach relies on careful modeling choices, a well- tuned training pipeline, and, most critically, the quality of our newly collected datasets. Our best-in-class 72B model not only outperforms others in the class of open weight and data models, but also outperforms larger proprietary models including Claude 3.5 Sonnet, and Gemini 1.5 Pro and Flash, second only to GPT-4o based on both academic benchmarks and on a large human evaluation. Our model weights, new datasets, and source code are available at https://molmo.allenai.org/blog. Matt Deitke, Sangho Lee 0008, Rohun Tripathi, Yue Yang 0006, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira Anderson, Erin Bransom, Kiana Ehsani, Huong Ngo, Yen-Sung Chen, Ajay Patel, Mark Yatskar, Chris Callison-Burch, Andrew Head, Rose Hendrix, Favyen Bastani, Eli VanderBilt, Nathan Lambert 0001, Yvonne Chou, Arnavi Chheda, Jenna Sparks, Sam Skjonsberg, Michael Schmitz 0002, Aaron Sarnat, Byron Bischoff, Pete Walsh 0001, Chris Newell, Piper Wolters, Tanmay Gupta, Kuo-Hao Zeng, Jon Borchardt, Dirk Groeneveld, Crystal Nam, Sophie Lebrecht, Caitlin Wittlif, Carissa Schoenick, Oscar Michel, Ranjay Krishna, Luca Weihs, Noah A. Smith, Hannaneh Hajishirzi, Ross B. Girshick, Ali Farhadi, Aniruddha Kembhavi |
CVPR | 46 |
| 2025 | Eval3D: Interpretable and Fine-grained Evaluation for 3D GenerationabstractDespite the unprecedented progress in the field of 3D generation, current systems still often fail to produce high-quality 3D assets that are visually appealing and geometrically and semantically consistent across multiple viewpoints. To effectively assess the quality of the generated 3D data, there is a need for a reliable 3D evaluation tool. Unfortunately, existing 3D evaluation metrics often overlook the geometric quality of generated assets or merely rely on black-box multimodal large language models for coarse assessment. In this paper, we introduce Eval3D, a fine-grained, interpretable evaluation tool that can faithfully evaluate the quality of generated 3D assets based on various distinct yet complementary criteria. Our key observation is that many desired properties of 3D generation, such as semantic and geometric consistency, can be effectively captured by measuring the consistency among various foundation models and tools. We thus leverage a diverse set of models and tools as probes to evaluate the inconsistency of generated 3D assets across different aspects. Compared to prior work, Eval3D provides pixel-wise measurement, enables accurate 3D spatial feedback, and aligns more closely with human judgments. We comprehensively evaluate existing 3D generation models using Eval3D and highlight the limitations and challenges of current models. Project page: http://eval3d.github.io. Shivam Duggal, Yushi Hu, Oscar Michel, Aniruddha Kembhavi, William T. Freeman, Noah A. Smith, Ranjay Krishna, Antonio Torralba 0001, Ali Farhadi, Wei-Chiu Ma |
CVPR | 6 |
| 2025 | Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-IndexabstractLanguage models are trained mainly on massive text data from the Internet, and it becomes increasingly important to understand this data source. Exact-match search engines enable searching in large text corpora – counting string appearances and retrieving the enclosing documents – yet the high storage overhead hinders their application on Internet-scale data. We present Infini-gram mini, an efficient and scalable system that can make petabyte-level text corpora searchable. Based on the FM-index data structure (Ferragina and Manzini, 2000), which simultaneously indexes and compresses text, our system creates indexes with size only 44% of the corpus. Infini-gram mini greatly improves upon the best existing implementation of FM-index in terms of indexing speed (18\times) and memory use during both indexing (3.2\times reduction) and querying (down to a negligible amount). We index 83TB of Internet text in 99 days with a single 128-core CPU node (or 19 hours if using 137 such nodes). We show one important use case of Infini-gram mini in a large-scale analysis of benchmark contamination. We find several core LM evaluation benchmarks to be heavily contaminated in Internet crawls (up to 74.2% in GSM8K), which could lead to overestimating the capabilities of language models if trained on such data. We host a benchmark contamination bulletin to share the contamination rate of many core and community-contributed benchmarks. We also release a web interface and an API endpoint to serve general search queries on Infini-gram mini indexes. Jiacheng Liu 0010, Yejin Choi 0001, Noah A. Smith, Hannaneh Hajishirzi |
EMNLP | 4 |
| 2025 | On Linear Representations and Pretraining Data Frequency in Language ModelsabstractPretraining data has a direct impact on the behaviors and quality of language models (LMs), but we only understand the most basic principles of this relationship. While most work focuses on pretraining data's effect on downstream task behavior, we investigate its relationship to LM representations. Previous work has discovered that, in language models, some concepts are encoded "linearly" in the representations, but what factors cause these representations to form (or not)? We study the connection between pretraining data frequency and models' linear representations of factual relations (e.g., mapping France to Paris in a capital prediction task). We find evidence that the formation of linear representations is strongly connected to pretraining term frequencies; specifically for subject-relation-object fact triplets, both subject-object co-occurrence frequency and in-context learning accuracy for the relation are highly correlated with linear representations. This is the case across all phases of pretraining, i.e., it is not affected by the model's underlying capability. In OLMo-7B and GPT-J (6B), we discover that a linear representation consistently (but not exclusively) forms when the subjects and objects within a relation co-occur at least 1k and 2k times, respectively, regardless of when these occurrences happen during pretraining (and around 4k times for OLMo-1B). Finally, we train a regression model on measurements of linear representation quality in fully-trained LMs that can predict how often a term was seen in pretraining. Our model achieves low error even on inputs from a different model with a different pretraining dataset, providing a new method for estimating properties of the otherwise-unknown training data of closed-data models. We conclude that the strength of linear representations in LMs contains signal about the models' pretraining corpora that may provide new avenues for controlling and improving model behavior: particularly, manipulating the models' training data to meet specific frequency thresholds. We release our code to support future work. Jack Merullo, Noah A. Smith, Sarah Wiegreffe, Yanai Elazar |
ICLR | 2 |
| 2025 | MUSE: Machine Unlearning Six-Way Evaluation for Language ModelsabstractLanguage models (LMs) are trained on vast amounts of text data, which may include private and copyrighted content. Data owners may request the removal of their data from a trained model due to privacy or copyright concerns. However, exactly unlearning only these datapoints (i.e., retraining with the data removed) is intractable in modern-day models. This has led to the development of many approximate unlearning algorithms. The evaluation of the efficacy of these algorithms has traditionally been narrow in scope, failing to precisely quantify the success and practicality of the algorithm from the perspectives of both the model deployers and the data owners. We address this issue by proposing MUSE, a comprehensive machine unlearning evaluation benchmark that enumerates six diverse desirable properties for unlearned models: (1) no verbatim memorization, (2) no knowledge memorization, (3) no privacy leakage, (4) utility preservation on data not intended for removal, (5) scalability with respect to the size of removal requests, and (6) sustainability over sequential unlearning requests. Using these criteria, we benchmark how effectively eight popular unlearning algorithms on 7B-parameter LMs can unlearn Harry Potter books and news articles. Our results demonstrate that most algorithms can prevent verbatim memorization and knowledge memorization to varying degrees, but only one algorithm does not lead to severe privacy leakage. Furthermore, existing algorithms fail to meet deployer's expectations because they often degrade general model utility and also cannot sustainably accommodate successive unlearning requests or large-scale content removal. Our findings identify key issues with the practicality of existing unlearning algorithms on language models. Jaechan Lee, Yangsibo Huang, Sadhika Malladi, Jieyu Zhao 0001, Ari Holtzman, Daogao Liu, Luke Zettlemoyer, Noah A. Smith, Chiyuan Zhang |
ICLR | 9 |
| 2025 | DataDecide: How to Predict Best Pretraining Data with Small ExperimentsabstractBecause large language models are expensive to pretrain on different datasets, using smaller-scale experiments to decide on data is crucial for reducing costs. Which benchmarks and methods of making decisions from observed performance at small scale most accurately predict the datasets that yield the best large models? To empower open exploration of this question, we release models, data, and evaluations in DataDecide—the most extensive open suite of models over differences in data and scale. We conduct controlled pretraining experiments across 25 corpora with differing sources, deduplication, and filtering up to 100B tokens, model sizes up to 1B parameters, and 3 random seeds. We find that the ranking of models at a single, small size (e.g., 150M parameters) is a strong baseline for predicting best models at our larger target scale (1B) ($\tilde$ 80% of comparisons correct). No scaling law methods among 8 baselines exceed the compute-decision frontier of single-scale predictions, but DataDecide can measure improvement in future scaling laws. We also identify that using continuous likelihood metrics as proxies in small experiments makes benchmarks including MMLU, ARC, HellaSwag, MBPP, and HumanEval $>$ 80% predictable at the target 1B scale with just 0.01% of the compute. Ian Magnusson, Nguyen Tai, Ben Bogin, David Heineman, Jena D. Hwang, Luca Soldaini, Akshita Bhagia, Jiacheng Liu 0010, Dirk Groeneveld, Oyvind Tafjord, Noah A. Smith, Pang Wei Koh, Jesse Dodge |
ICML | 11 |
| 2025 | Does Liking Yellow Imply Driving a School Bus? Semantic Leakage in Language ModelsabstractHila Gonen, Terra Blevins, Alisa Liu, Luke Zettlemoyer, Noah A. Smith. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hila Gonen, Terra Blevins, Alisa Liu, Luke Zettlemoyer, Noah A. Smith |
NAACL (Long Papers) | 5 |
| 2025 | ComPO: Community Preferences for Language Model PersonalizationabstractSachin Kumar, Chan Young Park, Yulia Tsvetkov, Noah A. Smith, Hannaneh Hajishirzi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sachin Kumar 0009, Chan Young Park, Yulia Tsvetkov, Noah A. Smith, Hannaneh Hajishirzi |
NAACL (Long Papers) | 4 |
| 2025 | Signal and Noise: A Framework for Reducing Uncertainty in Language Model EvaluationabstractDeveloping large language models is expensive and often involves making decisions with small experiments, typically by evaluating on large, multi-task evaluation suites. In this work, we analyze specific properties which make a benchmark more reliable and useful for such decisions, and interventions to design higher-quality evaluation benchmarks. We introduce two key metrics that show differences in current benchmarks: signal, a benchmark’s ability to separate better models from worse models, and noise, a benchmark’s sensitivity to random variability between training steps. We demonstrate that benchmarks with a better signal-to-noise ratio are more reliable when making decisions at small scale, and those with less noise have lower scaling law prediction error. These results suggest that improving signal or noise will lead to more useful benchmarks, so we introduce four interventions designed to directly affect signal or noise. For example, we propose that switching to a metric that has better signal and noise (e.g., perplexity rather than accuracy) leads to better reliability and scaling law error. We also find that filtering noisy benchmarks such that they have better signal-to-noise ratio leads to more reliable evaluations. We also find that averaging the output of a model's checkpoints to reduce noise leads to consistent improvements. We conclude by recommending that those creating new benchmarks, or selecting which existing benchmarks to use, aim for high signal and low noise. We use 30 benchmarks for these experiments, and 465 open-weight language models from 60M to 32B parameters, resulting in a new, publicly available dataset of 50K evaluation benchmark results, totaling 200M instances. David Heineman, Valentin Hofmann, Ian Magnusson, Yuling Gu, Noah A. Smith, Hannaneh Hajishirzi, Kyle Lo, Jesse Dodge |
NeurIPS | 5 |
| 2025 | FlexOLMo: Open Language Models for Flexible Data UseabstractWe introduce FlexOLMo, a new class of language models (LMs) that supports (1) distributed training without data sharing, where different model parameters are independently trained on private datasets, and (2) data-flexible inference, where these parameters along with their associated data can be easily included or excluded from model inferences with no further training. FlexOLMo employs a mixture-of-experts (MoE) architecture where each expert is trained independently on private datasets and later integrated through a new nonparametric routing without any joint training across datasets. FlexOLMo is trained on FLEXMIX, a corpus we curate comprising seven restricted sets, either real or realistic approximations, alongside publicly available datasets. We evaluate models with up to 37 billion parameters (20 billion active) on 31 diverse downstream tasks. We show that a general expert trained on public data can be effectively combined with independently trained experts from other data owners significantly benefiting from these restricted sets (an average 41% relative improvement) while allowing flexible opt-out at inference time (e.g., for users without appropriate licenses or permissions). Our approach also outperforms prior model merging methods by 10.1% on average and surpasses the standard MoE trained without data restrictions using the same training FLOPs. Altogether, FlexOLMo enables training on restricted data while keeping data local and supports fine-grained control of data access at inference. Akshita Bhagia, Kevin Farhat, Niklas Muennighoff, Jacob Morrison, Pete Walsh 0001, Dustin Schwenk, Shayne Longpre, Jake Poznanski, Allyson Ettinger, Daogao Liu, Margaret Li, Mike Lewis, Scott Yih, Dirk Groeneveld, Luca Soldaini, Kyle Lo, Noah A. Smith, Luke Zettlemoyer, Pang Wei W. Koh, Hannaneh Hajishirzi, Ali Farhadi, Sewon Min |
NeurIPS | 18 |
| 2025 | The Leaderboard IllusionabstractMeasuring progress is fundamental to the advancement of any scientific field. As benchmarks play an increasingly central role, they also grow more susceptible to distortion.Chatbot Arena has emerged as the go-to leaderboard for ranking the most capable AI systems. Yet, in this work we identify systematic issues that have resulted in a distorted playing field. We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired. We establish that the ability of these providers to choose the best score leads to biased Arena scores due to selective disclosure of performance results. At an extreme, we found one provider testing 27 private variants before making one model public at the second position on the leaderboard. We also establish that proprietary closed models are sampled at higher rates (number of battles) and have fewer models removed from the arena than open-weight and open-source alternatives. Both these policies lead to large data access asymmetries over time. The top two providers have individually received an estimated 19.2% and 20.4% of all data on the arena. In contrast, a combined 83 open-weight models have only received an estimated 29.7% of the total data. With conservative estimates, we show that access to Chatbot Arena data yields substantial benefits; even limited additional data can result in relative performance gains of up to 112% on ArenaHard, a test set from the arena distribution.Together, these dynamics result in overfitting to Arena-specific dynamics rather than general model quality. The Arena builds on the substantial efforts of both the organizers and an open community that maintains this valuable evaluation platform. We offer actionable recommendations to reform the Chatbot Arena's evaluation framework and promote fairer, more transparent benchmarking for the field. Shivalika Singh, Yiyang Nan, Daniel D'souza, Sayash Kapoor, Ahmet Üstün, Oluwasanmi Koyejo, Yuntian Deng, Shayne Longpre, Noah A. Smith, Beyza Ermis, Marzieh Fadaee, Sara Hooker |
NeurIPS | 10 |
| 2025 | Broken Tokens? Your Language Model can Secretly Handle Non-Canonical TokenizationsabstractModern tokenizers employ deterministic algorithms to map text into a single ``canonical" token sequence, yet the same string can be encoded as many non-canonical tokenizations using the language model vocabulary, including tokenizing by character. In this paper, we investigate the robustness of LMs to input encoded with non-canonical tokenizations entirely unseen during training. Surprisingly, when evaluated across 20 benchmarks, we find that instruction-tuned models retain up to 93.4\% of their original performance when given a randomly sampled tokenization, and 90.8\% with character-level tokenization. We find that overall stronger models tend to be more robust, and that robustness diminishes as the tokenization departs farther from the canonical form. Motivated by these results, we identify settings where non-canonical tokenization schemes can \textit{improve} performance, finding that character‑level segmentation improves string manipulation and code understanding tasks by up to 15\%, and right‑aligned digit grouping enhances large‑number arithmetic by over 33\%. Finally, we investigate the source of this robustness, finding that it arises in the instruction-tuning phase. We provide evidence that both base and post-trained models grasp the semantics of non-canonical tokenizations (perceiving them as containing misspellings). However, base models try to mimic the imagined mistakes and degenerate into nonsensical output, while post-trained models are committed to fluent responses. Overall, our findings suggest that models are less committed to their tokenizer than previously believed, and highlight the promise of intervening on tokenization at inference time to boost language model performance. Brian Siyuan Zheng, Alisa Liu, Orevaoghene Ahia, Jonathan Hayase, Yejin Choi 0001, Noah A. Smith |
NeurIPS | 6 |
| 2025 | Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in TransformersabstractAbstract Transformers trained on natural language data have been shown to exhibit hierarchical generalization without explicitly encoding any structural bias. In this work, we investigate sources of inductive bias in transformer models and their training that could cause such preference for hierarchical generalization. We extensively experiment with transformers trained on five synthetic, controlled datasets using several training objectives and show that, while objectives such as sequence-to-sequence modeling, classification, etc., often fail to lead to hierarchical generalization, the language modeling objective consistently leads to transformers generalizing hierarchically. We then study how different generalization behaviors emerge during the training by conducting pruning experiments that reveal the joint existence of subnetworks within the model implementing different generalizations. Finally, we take a Bayesian perspective to understand transformers’ preference for hierarchical generalization: We establish a correlation between whether transformers generalize hierarchically on a dataset and if the simplest explanation of that dataset is provided by a hierarchical grammar compared to regular grammars exhibiting linear generalization. Overall, our work presents new insights on the origins of hierarchical generalization in transformers and provides a theoretical framework for studying generalization in language models. Kabir Ahuja, Vidhisha Balachandran, Madhur Panwar, Tianxing He, Noah A. Smith, Navin Goyal, Yulia Tsvetkov |
Trans. Assoc. Comput. Linguistics | 5 |
| 2024 | OLMo: Accelerating the Science of Language ModelsabstractDirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, William Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah Smith, Hannaneh Hajishirzi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Dirk Groeneveld, Iz Beltagy, Pete Walsh 0001, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert 0001, Kyle Richardson 0001, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi |
ACL (1) | 42 |
| 2024 | Time is Encoded in the Weights of Finetuned Language ModelsabstractWe present time vectors, a simple tool to customize language models to new time periods.Time vectors are created by finetuning a language model on data from a single time (e.g., a year or month), and then subtracting the weights of the original pretrained model.This vector specifies a direction in weight space that, as our experiments show, improves performance on text from that time period.Time vectors specialized to adjacent time periods appear to be positioned closer together in a manifold.Using this structure, we interpolate between time vectors to induce new models that perform better on intervening and future time periods, without any additional training.We demonstrate the consistency of our findings across different tasks, domains, model sizes, and time scales.Our results suggest that time is encoded in the weight space of finetuned models. Kai Nylund, Suchin Gururangan, Noah A. Smith |
ACL (1) | 3 |
| 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining ResearchabstractLuca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar 0009, Li Lucy, Xinxi Lyu, Nathan Lambert 0001, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson 0001, Shannon Shen 0001, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh 0001, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo |
ACL (1) | 31 |
| 2024 | Know Your Audience: The benefits and pitfalls of generating plain language summaries beyond the "general" audienceabstractLanguage models (LMs) show promise as tools for communicating science to the general public by simplifying and summarizing complex language. Because models can be prompted to generate text for a specific audience (e.g., college-educated adults), LMs might be used to create multiple versions of plain language summaries for people with different familiarities of scientific topics. However, it is not clear what the benefits and pitfalls of adaptive plain language are. When is simplifying necessary, what are the costs in doing so, and do these costs differ for readers with different background knowledge? Through three within-subjects studies in which we surface summaries for different envisioned audiences to participants of different backgrounds, we found that while simpler text led to the best reading experience for readers with little to no familiarity in a topic, high familiarity readers tended to ignore certain details in overly plain summaries (e.g., study limitations). Our work provides methods and guidance on ways of adapting plain language summaries beyond the single “general” audience. Tal August, Kyle Lo, Noah A. Smith, Katharina Reinecke |
CHI | 3 |
| 2024 | A Call for Clarity in Beam Search: How It Works and When It StopsabstractText generation with beam search has proven successful in a wide range of applications. We point out that, though largely overlooked in the literature, the commonly-used implementation of beam decoding (e.g., Hugging Face Transformers and fairseq) uses a first come, first served heuristic: it keeps a set of already completed sequences over time steps and stops when the size of this set reaches the beam size. Based on this finding, we introduce a patience factor, a simple modification to this beam decoding implementation, that generalizes the stopping criterion and provides flexibility to the depth of search. Empirical results demonstrate that adjusting this patience factor improves decoding performance of strong pretrained models on news text summarization and machine translation over diverse language pairs, with a negligible inference slowdown. Our approach only modifies one line of code and can be thus readily incorporated in any implementation. Further, we find that different versions of beam decoding result in large performance differences in summarization, demonstrating the need for clarity in specifying the beam search implementation in research work. Our code will be available upon publication. Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras 0001, Dragomir R. Radev, Yejin Choi 0001, Noah A. Smith |
LREC/COLING | 6 |
| 2024 | BLINK: Multimodal Large Language Models Can See but Not Perceive
Yushi Hu, Bangzheng Li, Yu Feng 0013, Haoyu Wang 0005, Xudong Lin 0003, Dan Roth 0001, Noah A. Smith, Wei-Chiu Ma, Ranjay Krishna |
ECCV (23) | 8 |
| 2024 | Voices Unheard: NLP Resources and Models for Yorùbá Regional DialectsabstractOrevaoghene Ahia, Anuoluwapo Aremu, Diana Abagyan, Hila Gonen, David Ifeoluwa Adelani, Daud Abolade, Noah A. Smith, Yulia Tsvetkov. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Orevaoghene Ahia, Aremu Anuoluwapo, Diana Abagyan, Hila Gonen, David Ifeoluwa Adelani, Daud Abolade, Noah A. Smith, Yulia Tsvetkov |
EMNLP | 7 |
| 2024 | Breaking the Curse of Multilinguality with Cross-lingual Expert Language ModelsabstractTerra Blevins, Tomasz Limisiewicz, Suchin Gururangan, Margaret Li, Hila Gonen, Noah A. Smith, Luke Zettlemoyer. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Terra Blevins, Tomasz Limisiewicz, Suchin Gururangan, Margaret Li, Hila Gonen, Noah A. Smith, Luke Zettlemoyer |
EMNLP | 6 |
| 2024 | Evaluating n-Gram Novelty of Language Models Using Rusty-DAWGabstractHow novel are texts generated by language models (LMs) relative to their training corpora?In this work, we investigate the extent to which modern LMs generate n-grams from their training data, evaluating both (i) the probability LMs assign to complete training n-grams and (ii) n-novelty, the proportion of n-grams generated by an LM that did not appear in the training data (for arbitrarily large n).To enable arbitrary-length n-gram search over a corpus in constant time w.r.t.corpus size, we develop RUSTY-DAWG, a novel search tool inspired by indexing of genomic data.We compare the novelty of LM-generated text to humanwritten text and explore factors that affect generation novelty, focusing on the Pythia models.We find that, for n > 4, LM-generated text is less novel than human-written text, though it is more novel for smaller n.Larger LMs and more constrained decoding strategies both decrease novelty.Finally, we show that LMs complete n-grams with lower loss if they are more frequent in the training data.Overall, our results reveal factors influencing the novelty of LMgenerated text, and we release RUSTY-DAWG to facilitate further pretraining data research.1 William Merrill, Noah A. Smith, Yanai Elazar |
EMNLP | 2 |
| 2024 | What's In My Big Data?abstractLarge text corpora are the backbone of language models.
However, we have a limited understanding of the content of these corpora, including general statistics, quality, social factors, and inclusion of evaluation data (contamination).
In this work, we propose What's In My Big Data? (WIMBD), a platform and a set of sixteen analyses that allow us to reveal and compare the contents of large text corpora. WIMBD builds on two basic capabilities---count and search---*at scale*, which allows us to analyze more than 35 terabytes on a standard compute node.
We apply WIMBD to ten different corpora used to train popular language models, including *C4*, *The Pile*, and *RedPajama*.
Our analysis uncovers several surprising and previously undocumented findings about these corpora, including the high prevalence of duplicate, synthetic, and low-quality content, personally identifiable information, toxic language, and benchmark contamination.
For instance, we find that about 50% of the documents in *RedPajama* and *LAION-2B-en* are duplicates. In addition, several datasets used for benchmarking models trained on such corpora are contaminated with respect to important benchmarks, including the Winograd Schema Challenge and parts of GLUE and SuperGLUE.
We open-source WIMBD's code and artifacts to provide a standard set of evaluations for new text-based corpora and to encourage more analyses and transparency around them. Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Pete Walsh 0001, Dirk Groeneveld, Luca Soldaini, Sameer Singh 0001, Hannaneh Hajishirzi, Noah A. Smith, Jesse Dodge |
ICLR | 12 |
| 2024 | SILO Language Models: Isolating Legal Risk In a Nonparametric DatastoreabstractThe legality of training language models (LMs) on copyrighted or otherwise restricted data is under intense debate. However, as we show, model performance significantly degrades if trained only on low-risk text (e.g., out-of-copyright books or government documents), due to its limited size and domain coverage. We present SILO, a new language model that manages this risk-performance tradeoff during inference. SILO is built by (1) training a parametric LM on the Open License Corpus (OLC), a new corpus we curate with 228B tokens of public domain and permissively licensed text and (2) augmenting it with a more general and easily modifiable nonparametric datastore (e.g., containing copyrighted books or news) that is only queried during inference. The datastore allows use of high-risk data without training on it, supports sentence-level data attribution, and enables data producers to opt out from the model by removing content from the store. These capabilities can foster compliance with data-use regulations such as the fair use doctrine in the United States and the GDPR in the European Union. Our experiments show that the parametric LM struggles on its own with domains not covered by OLC. However, access to the datastore greatly improves out of domain performance, closing 90% of the performance gap with an LM trained on the Pile, a more diverse corpus with mostly high-risk text. We also analyze which nonparametric approach works best, where the remaining errors lie, and how performance scales with datastore size. Our results suggest that it is possible to build high quality language models while mitigating legal risk. Sewon Min, Suchin Gururangan, Eric Wallace, Hannaneh Hajishirzi, Noah A. Smith, Luke Zettlemoyer |
ICLR | 6 |
| 2024 | In-Context Pretraining: Language Modeling Beyond Document BoundariesabstractLanguage models are currently trained to predict tokens given document prefixes, enabling them to zero shot long form generation and prompting-style tasks which can be reduced to document completion. We instead present IN-CONTEXT PRETRAINING, a new approach where language models are trained on a sequence of related documents, thereby explicitly encouraging them to read and reason across document boundaries. Our approach builds on the fact that current pipelines train by concatenating random sets of shorter documents to create longer context windows; this improves efficiency even though the prior documents provide no signal for predicting the next document. Given this fact, we can do IN-CONTEXT PRETRAINING by simply changing the document ordering so that each context contains related documents, and directly applying existing pretraining pipelines. However, this document sorting problem is challenging. There are billions of documents and we would like the sort to maximize contextual similarity for every document without repeating any data. To do this, we introduce approximate algorithms for finding related documents with efficient nearest neighbor search and constructing coherent batches with a graph cover algorithm. Our experiments show IN-CONTEXT PRETRAINING offers a scalable and simple approach to significantly enhance LM performance: we see notable improvements in tasks that require more complex contextual reasoning, including in-context learning (+8%), reading comprehension (+15%), faithfulness to previous contexts (+16%), long-context reasoning (+5%), and retrieval augmentation (+9%). Sewon Min, Maria Lomeli, Chunting Zhou, Margaret Li, Xi Victoria Lin, Noah A. Smith, Luke Zettlemoyer, Scott Yih, Mike Lewis |
ICLR | 7 |
| 2024 | How Language Model Hallucinations Can SnowballabstractA major risk of using language models in practical applications is their tendency to hallucinate incorrect statements. Hallucinations are often attributed to knowledge gaps in LMs, but we show that LMs sometimes produce hallucinations that they can separately recognize as incorrect. To do this, we construct three question-answering datasets where LMs often state an incorrect answer which is followed by an explanation with at least one incorrect claim. Crucially, we find that GPT-3.5, GPT-4, and LLaMA2-70B-chat can identify 67%, 87%, and 94% of these incorrect claims, respectively. We show that this phenomenon doesn’t disappear under higher temperatures sampling, beam search, and zero-shot chain-of-thought prompting. These findings reveal that LM hallucinations can snowball: early mistakes by an LM can lead to more mistakes that otherwise would not be made. Muru Zhang, Ofir Press, William Merrill, Alisa Liu, Noah A. Smith |
ICML | 5 |
| 2024 | MAGNET: Improving the Multilingual Fairness of Language Models with Adaptive Gradient-Based TokenizationabstractIn multilingual settings, non-Latin scripts and low-resource languages are usually disadvantaged in terms of language models’ utility, efficiency, and cost. Specifically, previous studies have reported multiple modeling biases that the current tokenization algorithms introduce to non-Latin script languages, the main one being over-segmentation. In this work, we propose MAGNET— multilingual adaptive gradient-based tokenization—to reduce over-segmentation via adaptive gradient-based subword tokenization. MAGNET learns to predict segment boundaries between byte tokens in a sequence via sub-modules within the model, which act as internal boundary predictors (tokenizers). Previous gradient-based tokenization methods aimed for uniform compression across sequences by integrating a single boundary predictor during training and optimizing it end-to-end through stochastic reparameterization alongside the next token prediction objective. However, this approach still results in over-segmentation for non-Latin script languages in multilingual settings. In contrast, MAGNET offers a customizable architecture where byte-level sequences are routed through language-script-specific predictors, each optimized for its respective language script. This modularity enforces equitable segmentation granularity across different language scripts compared to previous methods. Through extensive experiments, we demonstrate that in addition to reducing segmentation disparities, MAGNET also enables faster language modeling and improves downstream utility. Orevaoghene Ahia, Sachin Kumar 0009, Hila Gonen, Valentin Hofmann, Tomasz Limisiewicz, Yulia Tsvetkov, Noah A. Smith |
NeurIPS | 7 |
| 2024 | The Art of Saying No: Contextual Noncompliance in Language ModelsabstractChat-based language models are designed to be helpful, yet they should not comply with every user request. While most existing work primarily focuses on refusal of ``unsafe'' queries, we posit that the scope of noncompliance should be broadened. We introduce a comprehensive taxonomy of contextual noncompliance describing when and how models should not comply with user requests. Our taxonomy spans a wide range of categories including incomplete, unsupported, indeterminate, and humanizing requests (in addition to unsafe requests). To test noncompliance capabilities of language models, we use this taxonomy to develop a new evaluation suite of 1000 noncompliance prompts. We find that most existing models show significantly high compliance rates in certain previously understudied categories with models like GPT-4 incorrectly complying with as many as 30\% of requests.To address these gaps, we explore different training strategies using a synthetically-generated training set of requests and expected noncompliant responses. Our experiments demonstrate that while direct finetuning of instruction-tuned models can lead to both over-refusal and a decline in general capabilities, using parameter efficient methods like low rank adapters helps to strike a good balance between appropriate noncompliance and other capabilities. Faeze Brahman, Sachin Kumar 0009, Vidhisha Balachandran, Pradeep Dasigi, Valentina Pyatkin, Abhilasha Ravichander, Sarah Wiegreffe, Nouha Dziri, Khyathi Raghavi Chandu, Jack Hessel, Yulia Tsvetkov, Noah A. Smith, Yejin Choi 0001, Hannaneh Hajishirzi |
NeurIPS | 12 |
| 2024 | Data Mixture Inference Attack: BPE Tokenizers Reveal Training Data CompositionsabstractThe pretraining data of today's strongest language models remains opaque, even when their parameters are open-sourced.
In particular, little is known about the proportions of different domains, languages, or code represented in the data. While a long line of membership inference attacks aim to identify training examples on an instance level, they do not extend easily to *global* statistics about the corpus. In this work, we tackle a task which we call *data mixture inference*, which aims to uncover the distributional make-up of the pretraining data. We introduce a novel attack based on a previously overlooked source of information — byte-pair encoding (BPE) tokenizers, used by the vast majority of modern language models. Our key insight is that the ordered vocabulary learned by a BPE tokenizer naturally reveals information about the token frequencies in its training data: the first token is the most common byte pair, the second is the most common pair after merging the first token, and so on. Given a tokenizer's merge list along with data samples for each category of interest (e.g., different natural languages), we formulate a linear program that solves for the relative proportion of each category in the tokenizer's training set. Importantly, to the extent to which tokenizer training data is representative of the pretraining data, we indirectly learn about the pretraining data. In controlled experiments, we show that our attack can recover mixture ratios with high precision for tokenizers trained on known mixtures of natural languages, programming languages, and data sources. We then apply our approach to off-the-shelf tokenizers released alongside recent LMs. We confirm much publicly disclosed information about these models, and also make several new inferences: GPT-4o is much more multilingual than its predecessors, training on 10x more non-English data than GPT-3.5, Llama 3 and Claude are trained on predominantly code, and many recent models are trained on 7-16% books. We hope our work sheds light on current design practices for pretraining data, and inspires continued research into data mixture inference for LMs. Jonathan Hayase, Alisa Liu, Yejin Choi 0001, Sewoong Oh, Noah A. Smith |
NeurIPS | 5 |
| 2024 | Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsabstractHumans draw to facilitate reasoning: we draw auxiliary lines when solving geometry problems; we mark and circle when reasoning on maps; we use sketches to amplify our ideas and relieve our limited-capacity working memory. However, such actions are missing in current multimodal language models (LMs). Current chain-of-thought and tool-use paradigms only use text as intermediate reasoning steps. In this work, we introduce Sketchpad, a framework that gives multimodal LMs a visual sketchpad and tools to draw on the sketchpad. The LM conducts planning and reasoning according to the visual artifacts it has drawn. Different from prior work, which uses text-to-image models to enable LMs to draw, Sketchpad enables LMs to draw with lines, boxes, marks, etc., which is closer to human sketching and better facilitates reasoning. \name can also use specialist vision models during the sketching process (e.g., draw bounding boxes with object detection models, draw masks with segmentation models), to further enhance visual perception and reasoning. We experiment on a wide range of math tasks (including geometry, functions, graph, chess) and complex visual reasoning tasks. Sketchpad substantially improves performance on all tasks over strong base models with no sketching, yielding an average gain of 12.7% on math tasks, and 8.6% on vision tasks. GPT-4o with Sketchpad sets a new state of the art on all tasks, including V*Bench (80.3%), BLINK spatial reasoning (83.9%), and visual correspondence (80.8%). We will release all code and data. Yushi Hu, Dan Roth 0001, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, Ranjay Krishna |
NeurIPS | 7 |
| 2024 | Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference FeedbackabstractLearning from preference feedback has emerged as an essential step for improving the generation quality and performance of modern language models (LMs). Despite its widespread use, the way preference-based learning is applied varies wildly, with differing data, learning algorithms, and evaluations used, making disentangling the impact of each aspect difficult. In this work, we identify four core aspects of preference-based learning: preference data, learning algorithm, reward model, and policy training prompts, systematically investigate the impact of these components on downstream model performance, and suggest a recipe for strong learning for preference feedback. Our findings indicate that all aspects are important for performance, with better preference data leading to the largest improvements, followed by the choice of learning algorithm, the use of improved reward models, and finally the use of additional unlabeled prompts for policy training. Notably, PPO outperforms DPO by up to 2.5% in math and 1.2% in general domains. High-quality preference data leads to improvements of up to 8% in instruction following and truthfulness. Despite significant gains of up to 5% in mathematical evaluation when scaling up reward models, we surprisingly observe marginal improvements in other categories. Hamish Ivison, Yizhong Wang, Jiacheng Liu 0010, Zeqiu Wu, Valentina Pyatkin, Nathan Lambert 0001, Noah A. Smith, Yejin Choi 0001, Hannaneh Hajishirzi |
NeurIPS | 7 |
| 2024 | Paloma: A Benchmark for Evaluating Language Model FitabstractEvaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains—varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. We include two new datasets of the top 100 subreddits (e.g., r/depression on Reddit) and programming languages (e.g., Java on GitHub), both sources common in contemporary LMs. With our benchmark, we release 6 baseline 1B LMs carefully controlled to provide fair comparisons about which pretraining corpus is best and code for others to apply those controls to their own experiments. Our case studies demonstrate how the fine-grained results from Paloma surface findings such as that models pretrained without data beyond Common Crawl exhibit anomalous gaps in LM fit to many domains or that loss is dominated by the most frequently occurring strings in the vocabulary. Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Pete Walsh 0001, Yanai Elazar, Kyle Lo, Dirk Groeneveld, Iz Beltagy, Hannaneh Hajishirzi, Noah A. Smith, Kyle Richardson 0001, Jesse Dodge |
NeurIPS | 14 |
| 2024 | Decoding-Time Language Model Alignment with Multiple ObjectivesabstractAligning language models (LMs) to human preferences has emerged as a critical pursuit, enabling these models to better serve diverse user needs. Existing methods primarily focus on optimizing LMs for a single reward function, limiting their adaptability to varied objectives.
Here, we propose $\textbf{multi-objective decoding~(MOD)}$, a decoding-time algorithm that outputs the next token from a linear combination of predictions of all base models, for any given weighting over different objectives.
We exploit a common form among a family of $f$-divergence regularized alignment approaches (such as PPO, DPO, and their variants) to identify a closed-form solution by Legendre transform, and derive an efficient decoding strategy.
Theoretically, we show why existing approaches can be sub-optimal even in natural settings and obtain optimality guarantees for our method.
Empirical results demonstrate the effectiveness of the algorithm. For example, compared to a parameter-merging baseline, MOD achieves 12.8\% overall reward improvement when equally optimizing towards $3$ objectives. Moreover, we experiment with MOD on combining three fully-finetuned
LMs of different model sizes, each aimed at different objectives such as safety, coding, and general user preference. Unlike traditional methods that require careful curation of a mixture of datasets to achieve comprehensive improvement, we can quickly experiment with preference weightings using MOD to find the best combination of models. Our best combination reduces toxicity on Toxigen to nearly 0\% and achieves 7.9--33.3\% improvement across three other metrics ($\textit{i.e.}$, Codex@1, GSM-COT, BBH-COT). Ruizhe Shi, Yifang Chen 0001, Yushi Hu, Alisa Liu, Hannaneh Hajishirzi, Noah A. Smith, Simon S. Du |
NeurIPS | 6 |
| 2024 | Evaluating Copyright Takedown Methods for Language ModelsabstractLanguage models (LMs) derive their capabilities from extensive training on diverse data, including copyrighted material. These models can memorize and generate content similar to their training data, potentially risking legal issues like copyright infringement.Therefore, model creators are motivated to develop mitigation methods that prevent generating particular copyrighted content, an ability we refer to as copyright takedowns. This paper introduces the first evaluation of the feasibility and side effects of copyright takedowns for LMs. We propose CoTaEval, an evaluation framework to assess the effectiveness of copyright takedown methods,the impact on the model's ability to retain uncopyrightable factual knowledge from the copyrighted content, and how well the model maintains its general utility and efficiency.We examine several strategies, including adding system prompts, decoding-time filtering interventions, and unlearning approaches. Our findings indicate that no method excels across all metrics, showing significant room for research in this unique problem setting and indicating potential unresolved challenges for live policy proposals. Boyi Wei, Yangsibo Huang, Noah A. Smith, Chiyuan Zhang, Luke Zettlemoyer, Kai Li 0001, Peter Henderson 0002 |
NeurIPS | 4 |
| 2024 | Morphosyntactic probing of multilingual BERT modelsabstractAbstract We introduce an extensive dataset for multilingual probing of morphological information in language models (247 tasks across 42 languages from 10 families), each consisting of a sentence with a target word and a morphological tag as the desired label, derived from the Universal Dependencies treebanks. We find that pre-trained Transformer models (mBERT and XLM-RoBERTa) learn features that attain strong performance across these tasks. We then apply two methods to locate, for each probing task, where the disambiguating information resides in the input. The first is a new perturbation method that “masks” various parts of context; the second is the classical method of Shapley values. The most intriguing finding that emerges is a strong tendency for the preceding context to hold more information relevant to the prediction than the following context. Judit Ács, Endre Hamerlik, Roy Schwartz 0001, Noah A. Smith, András Kornai |
Nat. Lang. Eng. | 4 |
| 2023 | Self-Instruct: Aligning Language Models with Self-Generated InstructionsabstractYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, Hannaneh Hajishirzi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, Hannaneh Hajishirzi |
ACL (1) | 5 |
| 2023 | Elaboration-Generating Commonsense Question Answering at ScaleabstractIn question answering requiring common sense, language models (e.g., GPT-3) have been used to generate text expressing background knowledge that helps improve performance.Yet the cost of working with such models is very high; in this work, we finetune smaller language models to generate useful intermediate context, referred to here as elaborations.Our framework alternates between updating two language models-an elaboration generator and an answer predictor-allowing each to influence the other.Using less than 0.5% of the parameters of GPT-3, our model outperforms alternatives with similar sizes and closes the gap with GPT-3 on four commonsense question answering benchmarks.Human evaluations show that the quality of the generated elaborations is high. 1 Wenya Wang 0001, Vivek Srikumar, Hannaneh Hajishirzi, Noah A. Smith |
ACL (1) | 4 |
| 2023 | Vera: A General-Purpose Plausibility Estimation Model for Commonsense StatementsabstractToday's language models can be remarkably intelligent yet still produce text that contains trivial commonsense errors.Therefore, we seek a retrospective verification approach that can reflect on the commonsense plausibility of the machine text, and introduce VERA, a general-purpose model that learns to estimate the commonsense plausibility of declarative statements.To support diverse commonsense domains, VERA is trained on ∼7M commonsense statements that are automatically converted from 19 QA datasets and two commonsense knowledge bases, and using a combination of three training objectives.When applied to solving commonsense problems in the verification format, VERA substantially outperforms existing models that can be repurposed for commonsense verification, even including GPT-3.5/ChatGPT/GPT-4, and it further exhibits generalization capabilities to unseen tasks and provides well-calibrated outputs.We find that VERA excels at filtering machinegenerated commonsense knowledge and is useful in detecting erroneous commonsense statements generated by models like ChatGPT in real-world settings. Jiacheng Liu 0010, Wenya Wang 0001, Dianzhuo Wang, Noah A. Smith, Yejin Choi 0001, Hannaneh Hajishirzi |
EMNLP | 4 |
| 2023 | Do All Languages Cost the Same? Tokenization in the Era of Commercial Language ModelsabstractLanguage models have graduated from being research prototypes to commercialized products offered as web APIs, and recent works have highlighted the multilingual capabilities of these products.The API vendors charge their users based on usage, more specifically on the number of "tokens" processed or generated by the underlying language models.What constitutes a token, however, is training data and model dependent with a large variance in the number of tokens required to convey the same information in different languages.In this work, we analyze the effect of this nonuniformity on the fairness of an API's pricing policy across languages.We conduct a systematic analysis of the cost and utility of OpenAI's language model API on multilingual benchmarks in 22 typologically diverse languages.We show evidence that speakers of a large number of the supported languages are overcharged while obtaining poorer results.These speakers tend to also come from regions where the APIs are less affordable to begin with.Through these analyses, we aim to increase transparency around language model APIs' pricing policies and encourage the vendors to make them more equitable. Orevaoghene Ahia, Sachin Kumar 0009, Hila Gonen, Jungo Kasai, David R. Mortensen, Noah A. Smith, Yulia Tsvetkov |
EMNLP | 6 |
| 2023 | We're Afraid Language Models Aren't Modeling AmbiguityabstractAlisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr, Peter West, Alexander Koller, Swabha Swayamdipta, Noah Smith, Yejin Choi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Alisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr, Peter West, Alexander Koller, Swabha Swayamdipta, Noah A. Smith, Yejin Choi 0001 |
EMNLP | 8 |
| 2023 | PromptCap: Prompt-Guided Image Captioning for VQA with GPT-3abstractKnowledge-based visual question answering (VQA) involves questions that require world knowledge beyond the image to yield the correct answer. Large language models (LMs) like GPT-3 are particularly helpful for this task because of their strong knowledge retrieval and reasoning capabilities. To enable LM to understand images, prior work uses a captioning model to convert images into text. However, when summarizing an image in a single caption sentence, which visual entities to describe are often underspecified. Generic image captions often miss visual details essential for the LM to answer visual questions correctly. To address this challenge, we propose PromptCap (Prompt-guided image Captioning), a captioning model designed to serve as a better connector between images and black-box LMs. Different from generic captions, PromptCap takes a natural-language prompt to control the visual entities to describe in the generated caption. The prompt contains a question that the caption should aid in answering. To avoid extra annotation, PromptCap is trained by examples synthesized with GPT-3 and existing datasets. We demonstrate Prompt-Cap’s effectiveness on an existing pipeline in which GPT-3 is prompted with image captions to carry out VQA. Prompt-Cap outperforms generic captions by a large margin and achieves state-of-the-art accuracy on knowledge-based VQA tasks (60.4% on OK-VQA and 59.6% on A-OKVQA). Zero-shot results on WebQA show that PromptCap generalizes well to unseen domains.1 Yushi Hu, Hang Hua, Zhengyuan Yang, Noah A. Smith, Jiebo Luo 0001 |
ICCV | 5 |
| 2023 | TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringabstractDespite thousands of researchers, engineers, and artists actively working on improving text-to-image generation models, systems often fail to produce images that accurately align with the text inputs. We introduce TIFA (Text-to-Image Faithfulness evaluation with question Answering), an automatic evaluation metric that measures the faithfulness of a generated image to its text input via visual question answering (VQA). Specifically, given a text input, we automatically generate several question-answer pairs using a language model. We calculate image faithfulness by checking whether existing VQA models can answer these questions using the generated image. TIFA is a reference-free metric that allows for fine-grained and interpretable evaluations of generated images. TIFA also has better correlations with human judgments than existing metrics. Based on this approach, we introduce TIFA v1.0, a benchmark consisting of 4K diverse text inputs and 25K questions across 12 categories (object, counting, etc.). We present a comprehensive evaluation of existing text-to-image models using TIFA v1.0 and highlight the limitations and challenges of current models. For instance, we find that current text-to-image models, despite doing well on color and material, still struggle in counting, spatial relations, and composing multiple objects. We hope our benchmark will help carefully measure the research progress in text-to-image synthesis and provide valuable insights for further research.1 Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, Noah A. Smith |
ICCV | 7 |
| 2023 | Binding Language Models in Symbolic Languages
Zhoujun Cheng, Tianbao Xie, Peng Shi 0010, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir R. Radev, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, Tao Yu 0009 |
ICLR | 11 |
| 2023 | Selective Annotation Makes Language Models Better Few-Shot Learners
Hongjin Su, Jungo Kasai, Chen Henry Wu, Jiayi Xin, Rui Zhang 0037, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, Tao Yu 0009 |
ICLR | 10 |
| 2023 | RealTime QA: What's the Answer Right Now?abstractWe introduce RealTime QA, a dynamic question answering (QA) platform that announces questions and evaluates systems on a regular basis (weekly in this version). RealTime QA inquires about the current world, and QA systems need to answer questions about novel events or information. It therefore challenges static, conventional assumptions in open-domain QA datasets and pursues instantaneous applications. We build strong baseline models upon large pretrained language models, including GPT-3 and T5. Our benchmark is an ongoing effort, and this paper presents real-time evaluation results over the past year. Our experimental results show that GPT-3 can often properly update its generation results, based on newly-retrieved documents, highlighting the importance of up-to-date information retrieval. Nonetheless, we find that GPT-3 tends to return outdated answers when retrieved documents do not provide sufficient information to find an answer. This suggests an important avenue for future research: can an open-domain QA system identify such unanswerable cases and communicate with the user or even the retrieval module to modify the retrieval results? We hope that RealTime QA will spur progress in instantaneous applications of question answering and beyond. Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Ronan Le Bras 0001, Akari Asai, Xinyan Yu 0001, Dragomir R. Radev, Noah A. Smith, Yejin Choi 0001, Kentaro Inui |
NeurIPS | 8 |
| 2023 | How Far Can Camels Go? Exploring the State of Instruction Tuning on Open ResourcesabstractIn this work we explore recent advances in instruction-tuning language models on a range of open instruction-following datasets. Despite recent claims that open models can be on par with state-of-the-art proprietary models, these claims are often accompanied by limited evaluation, making it difficult to compare models across the board and determine the utility of various resources. We provide a large set of instruction-tuned models from 6.7B to 65B parameters in size, trained on 12 instruction datasets ranging from manually curated (e.g., OpenAssistant) to synthetic and distilled (e.g., Alpaca) and systematically evaluate them on their factual knowledge, reasoning, multilinguality, coding, safety, and open-ended instruction following abilities through a collection of automatic, model-based, and human-based metrics. We further introduce Tülu, our best performing instruction-tuned model suite finetuned on a combination of high-quality open resources.Our experiments show that different instruction-tuning datasets can uncover or enhance specific skills, while no single dataset (or combination) provides the best performance across all evaluations. Interestingly, we find that model and human preference-based evaluations fail to reflect differences in model capabilities exposed by benchmark-based evaluations, suggesting the need for the type of systemic evaluation performed in this work. Our evaluations show that the best model in any given evaluation reaches on average 87% of ChatGPT performance, and 73% of GPT-4 performance, suggesting that further investment in building better base models and instruction-tuning data is required to close the gap. We release our instruction-tuned models, including a fully finetuned 65B Tülu, along with our code, data, and evaluation framework to facilitate future research. Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, Dave Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, Hannaneh Hajishirzi |
NeurIPS | 9 |
| 2023 | Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingabstractLanguage models (LMs) often exhibit undesirable text generation behaviors, including generating false, toxic, or irrelevant outputs.
Reinforcement learning from human feedback (RLHF)---where human preference judgments on LM outputs are transformed into a learning signal---has recently shown promise in addressing these issues. However, such holistic feedback conveys limited information on long text outputs; it does not indicate which aspects of the outputs influenced user preference; e.g., which parts contain what type(s) of errors. In this paper, we use fine-grained human feedback (e.g., which sentence is false, which sub-sentence is irrelevant) as an explicit training signal. We introduce Fine-Grained RLHF, a framework that enables training and learning from reward functions that are fine-grained in two respects: (1) density, providing a reward after every segment (e.g., a sentence) is generated; and (2) incorporating multiple reward models associated with different feedback types (e.g., factual incorrectness, irrelevance, and information incompleteness). We conduct experiments on detoxification and long-form question answering to illustrate how learning with this reward function leads to improved performance, supported by both automatic and human evaluation. Additionally, we show that LM behaviors can be customized using different combinations of fine-grained reward models. We release all data, collected human feedback, and codes at https://FineGrainedRLHF.github.io. Zeqiu Wu, Yushi Hu, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, Hannaneh Hajishirzi |
NeurIPS | 7 |
| 2023 | Transparency Helps Reveal When Language Models Learn MeaningabstractAbstract Many current NLP systems are built from language models trained to optimize unsupervised objectives on large amounts of raw text. Under what conditions might such a procedure acquire meaning? Our systematic experiments with synthetic data reveal that, with languages where all expressions have context-independent denotations (i.e., languages with strong transparency), both autoregressive and masked language models successfully learn to emulate semantic relations between expressions. However, when denotations are changed to be context-dependent with the language otherwise unmodified, this ability degrades. Turning to natural language, our experiments with a specific phenomenon—referential opacity—add to the growing body of evidence that current language models do not represent natural language semantics well. We show this failure relates to the context-dependent nature of natural language form-meaning mappings. Zhaofeng Wu, William Merrill, Hao Peng 0009, Iz Beltagy, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 5 |
| 2022 | Generating Scientific Definitions with Controllable ComplexityabstractUnfamiliar terminology and complex language can present barriers to understanding science. Natural language processing stands to help address these issues by automatically defining unfamiliar terms. We introduce a new task and dataset for defining scientific terms and controlling the complexity of generated definitions as a way of adapting to a specific reader's background knowledge. We test four definition generation methods for this new task, finding that a sequence-to-sequence approach is most successful. We then explore the version of the task in which definitions are generated at a target complexity level. We introduce a novel reranking approach and find in human evaluations that it offers superior fluency while also controlling complexity, compared to several controllable generation baselines. Tal August, Katharina Reinecke, Noah A. Smith |
ACL (1) | 3 |
| 2022 | Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine TextabstractModern neural language models can produce remarkably fluent and grammatical text.So much, in fact, that recent work by Clark et al. (2021) has reported that conventional crowdsourcing can no longer reliably distinguish between machine-authored (GPT-3) and humanauthored writing.As errors in machine generations become ever subtler and harder to spot, it poses a new challenge to the research community for robust machine text evaluation.We propose a new framework called SCARE-CROW for scrutinizing machine text via crowd annotation.To support the broad range of real machine errors that can be identified by laypeople, the ten error categories of SCARECROWsuch as redundancy , commonsense errors , and incoherence -are identified through several rounds of crowd annotation experiments without a predefined ontology.We then use SCARECROW to collect over 41k error spans in human-written and machinegenerated paragraphs of English language news text.We isolate factors for detailed analysis, including parameter count, training data, and various decoding-time configurations.Our approach successfully quantifies measurable gaps between human authored text and generations from models of several sizes, including fourteen configurations of GPT-3.In addition, our analysis unveils new insights, with detailed rationales provided by laypeople, e.g., that the commonsense capabilities have been improving with larger models while math capabilities have not, and that the choices of simple decoding hyperparameters can make remarkable differences on the perceived quality of machine text.We release our training material, annotation toolkit and dataset at Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A. Smith, Yejin Choi 0001 |
ACL (1) | 4 |
| 2022 | ABC: Attention with Bounded-memory ControlabstractHao Peng, Jungo Kasai, Nikolaos Pappas, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz, Noah Smith. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Hao Peng 0009, Jungo Kasai, Nikolaos Pappas 0002, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz 0001, Noah A. Smith |
ACL (1) | 8 |
| 2022 | Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data SelectionabstractSuchin Gururangan, Dallas Card, Sarah Dreier, Emily Gade, Leroy Wang, Zeyu Wang, Luke Zettlemoyer, Noah A. Smith. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Suchin Gururangan, Dallas Card, Sarah K. Dreier, Emily K. Gade, Leroy Z. Wang, Luke Zettlemoyer, Noah A. Smith |
EMNLP | 8 |
| 2022 | Twist Decoding: Diverse Generators Guide Each OtherabstractJungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Hao Peng, Ximing Lu, Dragomir Radev, Yejin Choi, Noah A. Smith. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras 0001, Hao Peng 0009, Ximing Lu, Dragomir R. Radev, Yejin Choi 0001, Noah A. Smith |
EMNLP | 8 |
| 2022 | GENIE: Toward Reproducible and Standardized Human Evaluation for Text GenerationabstractDaniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A. Smith, Daniel Weld. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi 0001, Noah A. Smith, Daniel S. Weld |
EMNLP | 7 |
| 2022 | UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language ModelsabstractTianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, Tao Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Tianbao Xie, Chen Henry Wu, Peng Shi 0010, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong 0005, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao 0002, Dragomir R. Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang 0037, Noah A. Smith, Luke Zettlemoyer, Tao Yu 0009 |
EMNLP | 21 |
| 2022 | Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Ofir Press, Noah A. Smith, Mike Lewis |
ICLR | 2 |
| 2022 | The Engage Corpus: A Social Media Dataset for Text-Based Recommender SystemsabstractSocial media platforms play an increasingly important role as forums for public discourse. Many platforms use recommendation algorithms that funnel users to online groups with the goal of maximizing user engagement, which many commentators have pointed to as a source of polarization and misinformation. Understanding the role of NLP in recommender systems is an interesting research area, given the role that social media has played in world events. However, there are few standardized resources which researchers can use to build models that predict engagement with online groups on social media; each research group constructs datasets from scratch without releasing their version for reuse. In this work, we present a dataset drawn from posts and comments on the online message board Reddit. We develop baseline models for recommending subreddits to users, given the user’s post and comment history. We also study the behavior of our recommender models on subreddits that were banned in June 2020 as part of Reddit’s efforts to stop the dissemination of hate speech. Daniel Cheng, Kyle Yan, Phillip Keung, Noah A. Smith |
LREC | 4 |
| 2022 | Domain Mismatch Doesn't Always Prevent Cross-lingual Transfer LearningabstractCross-lingual transfer learning without labeled target language data or parallel text has been surprisingly effective in zero-shot cross-lingual classification, question answering, unsupervised machine translation, etc. However, some recent publications have claimed that domain mismatch prevents cross-lingual transfer, and their results show that unsupervised bilingual lexicon induction (UBLI) and unsupervised neural machine translation (UNMT) do not work well when the underlying monolingual corpora come from different domains (e.g., French text from Wikipedia but English text from UN proceedings). In this work, we show how a simple initialization regimen can overcome much of the effect of domain mismatch in cross-lingual transfer. We pre-train word and contextual embeddings on the concatenated domain-mismatched corpora, and use these as initializations for three tasks: MUSE UBLI, UN Parallel UNMT, and the SemEval 2017 cross-lingual word similarity task. In all cases, our results challenge the conclusions of prior work by showing that proper initialization can recover a large portion of the losses incurred by domain mismatch. Daniel Edmiston, Phillip Keung, Noah A. Smith |
LREC | 3 |
| 2022 | DEMix Layers: Disentangling Domains for Modular Language ModelingabstractSuchin Gururangan, Mike Lewis, Ari Holtzman, Noah Smith, Luke Zettlemoyer. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A. Smith, Luke Zettlemoyer |
NAACL-HLT | 4 |
| 2022 | Bidimensional Leaderboards: Generate and Evaluate Language Hand in HandabstractJungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander Fabbri, Yejin Choi, Noah Smith. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras 0001, Lavinia Dunagan, Jacob Morrison, Alexander R. Fabbri, Yejin Choi 0001, Noah A. Smith |
NAACL-HLT | 8 |
| 2022 | Transparent Human Evaluation for Image CaptioningabstractJungo Kasai, Keisuke Sakaguchi, Lavinia Dunagan, Jacob Morrison, Ronan Le Bras, Yejin Choi, Noah Smith. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jungo Kasai, Keisuke Sakaguchi, Lavinia Dunagan, Jacob Morrison, Ronan Le Bras 0001, Yejin Choi 0001, Noah A. Smith |
NAACL-HLT | 7 |
| 2022 | NeuroLogic A*esque Decoding: Constrained Text Generation with Lookahead HeuristicsabstractXiming Lu, Sean Welleck, Peter West, Liwei Jiang, Jungo Kasai, Daniel Khashabi, Ronan Le Bras, Lianhui Qin, Youngjae Yu, Rowan Zellers, Noah Smith, Yejin Choi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ximing Lu, Sean Welleck, Peter West, Jungo Kasai, Daniel Khashabi, Ronan Le Bras 0001, Lianhui Qin, Youngjae Yu, Rowan Zellers, Noah A. Smith, Yejin Choi 0001 |
NAACL-HLT | 11 |
| 2022 | Time Waits for No One! Analysis and Challenges of Temporal MisalignmentabstractKelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam, Noah Smith. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Kelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam, Noah A. Smith |
NAACL-HLT | 5 |
| 2022 | Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language DetectionabstractMaarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, Noah Smith. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Maarten Sap, Swabha Swayamdipta, Laura Vianna, Yejin Choi 0001, Noah A. Smith |
NAACL-HLT | 6 |
| 2022 | Saturated Transformers are Constant-Depth Threshold CircuitsabstractAbstract Transformers have become a standard neural network architecture for many NLP problems, motivating theoretical analysis of their power in terms of formal languages. Recent work has shown that transformers with hard attention are quite limited in power (Hahn, 2020), as they can be simulated by constant-depth AND/OR circuits (Hao et al., 2022). However, hard attention is a strong assumption, which may complicate the relevance of these results in practice. In this work, we analyze the circuit complexity of transformers with saturated attention: a generalization of hard attention that more closely captures the attention patterns learnable in practical transformers. We first show that saturated transformers transcend the known limitations of hard-attention transformers. We then prove saturated transformers with floating-point values can be simulated by constant-depth threshold circuits, giving the class TC0 as an upper bound on the formal languages they recognize. William Merrill, Ashish Sabharwal, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated TextabstractElizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, Noah A. Smith. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, Noah A. Smith |
ACL/IJCNLP (1) | 6 |
| 2021 | DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-ExpertsabstractAlisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, Yejin Choi. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, Yejin Choi 0001 |
ACL/IJCNLP (1) | 6 |
| 2021 | Explaining Relationships Between Scientific DocumentsabstractKelvin Luu, Xinyi Wu, Rik Koncel-Kedziorski, Kyle Lo, Isabel Cachola, Noah A. Smith. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kelvin Luu, Rik Koncel-Kedziorski, Kyle Lo, Isabel Cachola, Noah A. Smith |
ACL/IJCNLP (1) | 6 |
| 2021 | Shortformer: Better Language Modeling using Shorter InputsabstractOfir Press, Noah A. Smith, Mike Lewis. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ofir Press, Noah A. Smith, Mike Lewis |
ACL/IJCNLP (1) | 2 |
| 2021 | Challenges in Automated Debiasing for Toxic Language DetectionabstractWarning: this paper contains content that may be offensive or upsetting.Biased associations have been a challenge in the development of classifiers for detecting toxic language, hindering both fairness and accuracy.As potential solutions, we investigate recently introduced debiasing methods for text classification datasets and models, as applied to toxic language detection.Our focus is on lexical (e.g., swear words, slurs, identity mentions) and dialectal markers (specifically African American English).Our comprehensive experiments establish that existing methods are limited in their ability to prevent biased behavior in current toxicity detectors.We then propose an automatic, dialect-aware data correction method, as a proof-of-concept study.Despite the use of synthetic labels, this method reduces dialectal associations with toxicity.Overall, our findings show that debiasing a model trained on biased toxic language data is not as effective as simply relabeling the data to remove existing biases. Maarten Sap, Swabha Swayamdipta, Yejin Choi 0001, Noah A. Smith |
EACL | 5 |
| 2021 | Competency Problems: On Finding and Removing Artifacts in Language DataabstractMuch recent work in NLP has documented dataset artifacts, bias, and spurious correlations between input features and output labels.However, how to tell which features have "spurious" instead of legitimate correlations is typically left unspecified.In this work we argue that for complex language understanding tasks, all simple feature correlations are spurious, and we formalize this notion into a class of problems which we call competency problems.For example, the word "amazing" on its own should not give information about a sentiment label independent of the context in which it appears, which could include negation, metaphor, sarcasm, etc.We theoretically analyze the difficulty of creating data for competency problems when human bias is taken into account, showing that realistic datasets will increasingly deviate from competency problems as dataset size increases.This analysis gives us a simple statistical test for dataset artifacts, which we use to show more subtle biases than were described in prior work, including demonstrating that models are inappropriately affected by these less extreme biases.Our theoretical treatment of this problem also allows us to analyze proposed solutions, such as making local edits to dataset instances, and to give recommendations for future data collection and model design efforts that target competency problems. Matt Gardner 0001, William Merrill, Jesse Dodge, Matthew E. Peters, Alexis Ross, Sameer Singh 0001, Noah A. Smith |
EMNLP (1) | 7 |
| 2021 | Finetuning Pretrained Transformers into RNNsabstractJungo Kasai, Hao Peng, Yizhe Zhang, Dani Yogatama, Gabriel Ilharco, Nikolaos Pappas, Yi Mao, Weizhu Chen, Noah A. Smith. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Jungo Kasai, Hao Peng 0009, Yizhe Zhang 0002, Dani Yogatama, Gabriel Ilharco, Nikolaos Pappas 0002, Weizhu Chen, Noah A. Smith |
EMNLP (1) | 9 |
| 2021 | Effects of Parameter Norm Growth During Transformer Training: Inductive Bias from Gradient DescentabstractThe capacity of neural networks like the widely adopted transformer is known to be very high.Evidence is emerging that they learn successfully due to inductive bias in the training routine, typically a variant of gradient descent (GD).To better understand this bias, we study the tendency for transformer parameters to grow in magnitude (ℓ 2 norm) during training, and its implications for the emergent representations within self attention layers.Empirically, we document norm growth in the training of transformer language models, including T5 during its pretraining.As the parameters grow in magnitude, we prove that the network approximates a discretized network with saturated activation functions.Such "saturated" networks are known to have a reduced capacity compared to the full network family that can be described in terms of formal languages and automata.Our results suggest saturation is a new characterization of an inductive bias implicit in GD of particular interest for NLP.We leverage the emergent discrete structure in a saturated transformer to analyze the role of different attention heads, finding that some focus locally on a small number of positions, while other heads compute global averages, allowing counting.We believe understanding the interplay between these two capabilities may shed further light on the structure of computation within large transformers. William Merrill, Vivek Ramanujan, Yoav Goldberg, Roy Schwartz 0001, Noah A. Smith |
EMNLP (1) | 5 |
| 2021 | Sentence Bottleneck Autoencoders from Transformer Language ModelsabstractRepresentation learning for text via pretraining a language model on a large corpus has become a standard starting point for building NLP systems.This approach stands in contrast to autoencoders, also trained on raw text, but with the objective of learning to encode each input as a vector that allows full reconstruction.Autoencoders are attractive because of their latent space structure and generative properties.We therefore explore the construction of a sentence-level autoencoder from a pretrained, frozen transformer language model.We adapt the masked language modeling objective as a generative, denoising one, while only training a sentence bottleneck and a single-layer modified transformer decoder.We demonstrate that the sentence representations discovered by our model achieve better quality than previous methods that extract representations from pretrained transformers on text similarity tasks, style transfer (an example of controlled generation), and single-sentence classification tasks in the GLUE benchmark, while using fewer parameters than large pretrained models. 1 Ivan Montero, Nikolaos Pappas 0002, Noah A. Smith |
EMNLP (1) | 3 |
| 2021 | Measuring Association Between Labels and Free-Text RationalesabstractIn interpretable NLP, we require faithful rationales that reflect the model's decision-making process for an explained instance.While prior work focuses on extractive rationales (a subset of the input words), we investigate their lessstudied counterpart: free-text natural language rationales.We demonstrate that pipelines, models for faithful rationalization on informationextraction style tasks, do not work as well on "reasoning" tasks requiring free-text rationales.We turn to models that jointly predict and rationalize, a class of widely used high-performance models for free-text rationalization.We investigate the extent to which the labels and rationales predicted by these models are associated, a necessary property of faithful explanation.Via two tests, robustness equivalence and feature importance agreement, we find that stateof-the-art T5-based joint models exhibit desirable properties for explaining commonsense question-answering and natural language inference, indicating their potential for producing faithful free-text rationales. 1 Sarah Wiegreffe, Ana Marasovic, Noah A. Smith |
EMNLP (1) | 3 |
| 2021 | Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine Translation
Jungo Kasai, Nikolaos Pappas 0002, Hao Peng 0009, James Cross 0003, Noah A. Smith |
ICLR | 5 |
| 2021 | Random Feature Attention
Hao Peng 0009, Nikolaos Pappas 0002, Dani Yogatama, Roy Schwartz 0001, Noah A. Smith, Lingpeng Kong |
ICLR | 5 |
| 2021 | Choose Your Own Adventure: Paired Suggestions in Collaborative Writing for Evaluating Story Generation ModelsabstractStory generation is an open-ended and subjective task, which poses a challenge for evaluating story generation models.We present CHOOSE YOUR OWN ADVENTURE, a collaborative writing setup for pairwise model evaluation.Two models generate suggestions to people as they write a short story; we ask writers to choose one of the two suggestions, and we observe which model's suggestions they prefer.The setup also allows further analysis based on the revisions people make to the suggestions.We show that these measures, combined with automatic metrics, provide an informative picture of the models' performance, both in cases where the differences in generation methods are small (nucleus vs. top-k sampling) and large (GPT2 vs. Fusion models). Elizabeth Clark, Noah A. Smith |
NAACL-HLT | 2 |
| 2021 | A Dataset of Information-Seeking Questions and Answers Anchored in Research PapersabstractPradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, Matt Gardner. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, Matt Gardner 0001 |
NAACL-HLT | 5 |
| 2021 | Provable Limitations of Acquiring Meaning from Ungrounded Form: What Will Future Language Models Understand?abstractAbstract Language models trained on billions of tokens have recently led to unprecedented results on many NLP tasks. This success raises the question of whether, in principle, a system can ever “understand” raw text without access to some form of grounding. We formally investigate the abilities of ungrounded systems to acquire meaning. Our analysis focuses on the role of “assertions”: textual contexts that provide indirect clues about the underlying semantics. We study whether assertions enable a system to emulate representations preserving semantic relations like equivalence. We find that assertions enable semantic emulation of languages that satisfy a strong notion of semantic transparency. However, for classes of languages where the same expression can take different values in different contexts, we show that emulation can become uncomputable. Finally, we discuss differences between our formal model and natural language, exploring how our results generalize to a modal setting and other semantic relations. Together, our results suggest that assertions in code or language do not provide sufficient signal to fully emulate semantic representations. We formalize ways in which ungrounded language models appear to be fundamentally limited in their ability to “understand”. William Merrill, Yoav Goldberg, Roy Schwartz 0001, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 4 |
| 2021 | Infusing Finetuning with Semantic DependenciesabstractAbstract For natural language processing systems, two kinds of evidence support the use of text representations from neural language models “pretrained” on large unannotated corpora: performance on application-inspired benchmarks (Peters et al., 2018, inter alia), and the emergence of syntactic abstractions in those representations (Tenney et al., 2019, inter alia). On the other hand, the lack of grounded supervision calls into question how well these representations can ever capture meaning (Bender and Koller, 2020). We apply novel probes to recent language models— specifically focusing on predicate-argument structure as operationalized by semantic dependencies (Ivanova et al., 2012)—and find that, unlike syntax, semantics is not brought to the surface by today’s pretrained models. We then use convolutional graph encoders to explicitly incorporate semantic parses into task-specific finetuning, yielding benefits to natural language understanding (NLU) tasks in the GLUE benchmark. This approach demonstrates the potential for general-purpose (rather than task-specific) linguistic supervision, above and beyond conventional pretraining and finetuning. Several diagnostics help to localize the benefits of our approach.1 Zhaofeng Wu, Hao Peng 0009, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | Don't Stop Pretraining: Adapt Language Models to Domains and TasksabstractLanguage models pretrained on text from a wide variety of sources form the foundation of today's NLP. In light of the success of these broad-coverage models, we investigate whether it is still helpful to tailor a pretrained model to the domain of a target task. We present a study across four domains (biomedical and computer science publications, news, and reviews) and eight classification tasks, showing that a second phase of pretraining in-domain (domain-adaptive pretraining) leads to performance gains, under both high- and low-resource settings. Moreover, adapting to the task's unlabeled data (task-adaptive pretraining) improves performance even after domain-adaptive pretraining. Finally, we show that adapting to a task corpus augmented using simple data selection strategies is an effective alternative, especially when resources for domain-adaptive pretraining might be unavailable. Overall, we consistently find that multi-phase adaptive pretraining offers large gains in task performance. Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, Noah A. Smith |
ACL | 7 |
| 2020 | A Formal Hierarchy of RNN ArchitecturesabstractWe develop a formal hierarchy of the expressive capacity of RNN architectures.The hierarchy is based on two formal properties: space complexity, which measures the RNN's memory, and rational recurrence, defined as whether the recurrent update can be described by a weighted finite-state machine.We place several RNN variants within this hierarchy.For example, we prove the LSTM is not rational, which formally separates it from the related QRNN (Bradbury et al., 2016).We also show how these models' expressive capacity is expanded by stacking multiple layers or composing them with different pooling functions.Our results build on the theory of "saturated" RNNs (Merrill, 2019).While formally extending these findings to unsaturated RNNs is left to future work, we hypothesize that the practical learnable capacity of unsaturated RNNs obeys a similar hierarchy.Experimental findings from training unsaturated networks on formal languages support this conjecture.We report updated experiments in Appendix H. William Merrill, Gail Weiss, Yoav Goldberg, Roy Schwartz 0001, Noah A. Smith, Eran Yahav |
ACL | 5 |
| 2020 | A Mixture of h - 1 Heads is Better than h HeadsabstractMulti-head attentive neural architectures have achieved state-of-the-art results on a variety of natural language processing tasks.Evidence has shown that they are overparameterized; attention heads can be pruned without significant performance loss.In this work, we instead "reallocate" them-the model learns to activate different heads on different inputs.Drawing connections between multi-head attention and mixture of experts, we propose the mixture of attentive experts model (MAE).MAE is trained using a block coordinate descent algorithm that alternates between updating (1) the responsibilities of the experts and (2) their parameters.Experiments on machine translation and language modeling show that MAE outperforms strong baselines on both tasks.Particularly, on the WMT14 English to German translation dataset, MAE improves over "transformer-base" by 0.8 BLEU, with a comparable number of parameters.Our analysis shows that our model learns to specialize different experts to different inputs. 1 Hao Peng 0009, Roy Schwartz 0001, Dianqi Li, Noah A. Smith |
ACL | 4 |
| 2020 | Improving Transformer Models by Reordering their SublayersabstractMultilayer transformer networks consist of interleaved self-attention and feedforward sublayers.Could ordering the sublayers in a different pattern lead to better performance?We generate randomly ordered transformers and train them with the language modeling objective.We observe that some of these models are able to achieve better performance than the interleaved baseline, and that those successful variants tend to have more self-attention at the bottom and more feedforward sublayers at the top.We propose a new transformer pattern that adheres to this property, the sandwich transformer, and show that it improves perplexity on multiple word-level and character-level language modeling benchmarks, at no cost in parameters, memory, or training time.However, the sandwich reordering pattern does not guarantee performance gains across every task, as we demonstrate on machine translation models.Instead, we suggest that further exploration of task-specific sublayer reorderings is needed in order to unlock additional gains. 1 Ofir Press, Noah A. Smith, Omer Levy |
ACL | 2 |
| 2020 | Social Bias Frames: Reasoning about Social and Power Implications of Languageabstractcontains content that may be offensive or upsetting. Maarten Sap, Saadia Gabriel, Lianhui Qin, Daniel Jurafsky, Noah A. Smith, Yejin Choi 0001 |
ACL | 5 |
| 2020 | Recollection versus Imagination: Exploring Human Memory and Cognition via Neural Language ModelsabstractWe investigate the use of NLP as a measure of the cognitive processes involved in storytelling, contrasting imagination and recollection of events. To facilitate this, we collect and release Hippocorpus, a dataset of 7,000 stories about imagined and recalled events. We introduce a measure of narrative flow and use this to examine the narratives for imagined and recalled events. Additionally, we measure the differential recruitment of knowledge attributed to semantic memory versus episodic memory (Tulving, 1972) for imagined and recalled storytelling by comparing the frequency of descriptions of general commonsense events with more specific realis events. Our analyses show that imagined stories have a substantially more linear narrative flow, compared to recalled stories in which adjacent sentences are more disconnected. In addition, while recalled stories rely more on autobiographical events based on episodic memory, imagined stories express more commonsense knowledge based on semantic memory. Finally, our measures reveal the effect of narrativization of memories in stories (e.g., stories about frequently recalled memories flow more linearly; Bartlett, 1932). Our findings highlight the potential of using NLP tools to study the traces of human cognition in language. Maarten Sap, Eric Horvitz, Yejin Choi 0001, Noah A. Smith, James W. Pennebaker |
ACL | 4 |
| 2020 | The Right Tool for the Job: Matching Model and Instance ComplexitiesabstractAs NLP models become larger, executing a trained model requires significant computational resources incurring monetary and environmental costs.To better respect a given inference budget, we propose a modification to contextual representation fine-tuning which, during inference, allows for an early (and fast) "exit" from neural network calculations for simple instances, and late (and accurate) exit for hard instances.To achieve this, we add classifiers to different layers of BERT and use their calibrated confidence scores to make early exit decisions.We test our proposed modification on five different datasets in two tasks: three text classification datasets and two natural language inference benchmarks.Our method presents a favorable speed/accuracy tradeoff in almost all cases, producing models which are up to five times faster than the state of the art, while preserving their accuracy.Our method also requires almost no additional training resources (in either time or parameters) compared to the baseline BERT model.Finally, our method alleviates the need for costly retraining of multiple models at different levels of efficiency; we allow users to control the inference speed/accuracy tradeoff using a single trained model, by setting a single variable at inference time.We publicly release our code.1 * Research completed during an internship at AI2. 1 github.com/allenai/sledgehammerLayer 0 Layer i Layer k Layer n Input Layer l Layer j Is confident?Yes Roy Schwartz 0001, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, Noah A. Smith |
ACL | 5 |
| 2020 | Explain like I am a Scientist: The Linguistic Barriers of Entry to r/scienceabstractAs an online community for discussing research findings, r/science has the potential to contribute to science outreach and communication with a broad audience. Yet previous work suggests that most of the active contributors on r/science are science-educated people rather than a lay general public. One potential reason is that r/science contributors might use a different, more specialized language than used in other subreddits. To investigate this possibility, we analyzed the language used in more than 68 million posts and comments from 12 subreddits from 2018. We show that r/science uses a specialized language that is distinct from other subreddits. Transient (newer) authors of posts and comments on r/science use less specialized language than more frequent authors, and those that leave the community use less specialized language than those that stay, even when comparing their first comments. These findings suggest that the specialized language used in r/science has a gatekeeping effect, preventing participation by people whose language does not align with that used in r/science. By characterizing r/science's specialized language, we contribute guidelines and tools for increasing the number of contributors in r/science. Tal August, Dallas Card, Gary Hsieh, Noah A. Smith, Katharina Reinecke |
CHI | 4 |
| 2020 | Writing Strategies for Science Communication: Data and Computational AnalysisabstractCommunicating complex scientific ideas without misleading or overwhelming the public is challenging.While science communication guides exist, they rarely offer empirical evidence for how their strategies are used in practice.Writing strategies that can be automatically recognized could greatly support science communication efforts by enabling tools to detect and suggest strategies for writers.We compile a set of writing strategies drawn from a wide range of prescriptive sources and develop an annotation scheme allowing humans to recognize them.We collect a corpus of 128K science writing documents in English and annotate a subset of this corpus.1 We use the annotations to train transformer-based classifiers and measure the strategies' use in the larger corpus.We find that the use of strategies, such as storytelling and emphasizing the most important findings, varies significantly across publications with different reader audiences. Tal August, Lauren Kim, Katharina Reinecke, Noah A. Smith |
EMNLP (1) | 4 |
| 2020 | The Multilingual Amazon Reviews CorpusabstractWe present the Multilingual Amazon Reviews Corpus (MARC), a large-scale collection of Amazon reviews for multilingual text classification.The corpus contains reviews in English, Japanese, German, French, Spanish, and Chinese, which were collected between 2015 and 2019.Each record in the dataset contains the review text, the review title, the star rating, an anonymized reviewer ID, an anonymized product ID, and the coarse-grained product category (e.g., 'books', 'appliances', etc.)The corpus is balanced across the 5 possible star ratings, so each rating constitutes 20% of the reviews in each language.For each language, there are 200,000, 5,000, and 5,000 reviews in the training, development, and test sets, respectively.We report baseline results for supervised text classification and zero-shot crosslingual transfer learning by fine-tuning a multilingual BERT model on reviews data.We propose the use of mean absolute error (MAE) instead of classification accuracy for this task, since MAE accounts for the ordinal nature of the ratings. Phillip Keung, Yichao Lu, György Szarvas, Noah A. Smith |
EMNLP (1) | 4 |
| 2020 | Plug and Play Autoencoders for Conditional Text GenerationabstractText autoencoders are commonly used for conditional generation tasks such as style transfer.We propose methods which are plug and play, where any pretrained autoencoder can be used, and only require learning a mapping within the autoencoder's embedding space, training embedding-to-embedding (Emb2Emb).This reduces the need for labeled training data for the task and makes the training procedure more efficient.Crucial to the success of this method is a loss term for keeping the mapped embedding on the manifold of the autoencoder and a mapping which is trained to navigate the manifold by learning offset vectors.Evaluations on style transfer tasks both with and without sequence-to-sequence supervision show that our method performs better than or comparable to strong baselines while being up to four times faster. Florian Mai, Nikolaos Pappas 0002, Ivan Montero, Noah A. Smith, James Henderson 0001 |
EMNLP (1) | 4 |
| 2020 | Grounded Compositional Outputs for Adaptive Language ModelingabstractLanguage models have emerged as a central component across NLP, and a great deal of progress depends on the ability to cheaply adapt them (e.g., through finetuning) to new domains and tasks.A language model's vocabulary-typically selected before training and permanently fixed later-affects its size and is part of what makes it resistant to such adaptation.Prior work has used compositional input embeddings based on surface forms to ameliorate this issue.In this work, we go one step beyond and propose a fully compositional output embedding layer for language models, which is further grounded in information from a structured lexicon (WordNet), namely semantically related words and free-text definitions.To our knowledge, the result is the first word-level language model with a size that does not depend on the training vocabulary.We evaluate the model on conventional language modeling as well as challenging cross-domain settings with an open vocabulary, finding that it matches or outperforms previous state-of-theart output embedding methods and adaptation approaches.Our analysis attributes the improvements to sample efficiency: our model is more accurate for low-frequency words. Nikolaos Pappas 0002, Phoebe Mulcaire, Noah A. Smith |
EMNLP (1) | 3 |
| 2020 | Dataset Cartography: Mapping and Diagnosing Datasets with Training DynamicsabstractSwabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, Yejin Choi. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Swabha Swayamdipta, Roy Schwartz 0001, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, Yejin Choi 0001 |
EMNLP (1) | 6 |
| 2020 | Multilevel Text Alignment with Cross-Document AttentionabstractText alignment finds application in tasks such as citation recommendation and plagiarism detection. Existing alignment methods operate at a single, predefined level and cannot learn to align texts at, for example, sentence and document levels. We propose a new learning approach that equips previously established hierarchical attention encoders for representing documents with a cross-document attention component, enabling structural comparisons across different levels (document-to-document and sentence-to-document). Our component is weakly supervised from document pairs and can align at multiple levels. Our evaluation on predicting document-to-document relationships and sentence-to-document relationships on the tasks of citation recommendation and plagiarism detection shows that our approach outperforms previously established hierarchical, attention encoders based on recurrent and transformer contextualization that are unaware of structural correspondence between documents. Nikolaos Pappas 0002, Noah A. Smith |
EMNLP (1) | 3 |
| 2020 | Multilingual and Interlingual Semantic Representations for Natural Language Processing: A Brief IntroductionabstractWe introduce the Computational Linguistics special issue on Multilingual and Interlingual Semantic Representations for Natural Language Processing. We situate the special issue’s five articles in the context of our fast-changing field, explaining our motivation for this project. We offer a brief summary of the work in the issue, which includes developments on lexical and sentential semantic representations, from symbolic and neural perspectives. Marta R. Costa-jussà, Cristina España-Bonet, Pascale Fung, Noah A. Smith |
Comput. Linguistics | 4 |
| 2020 | Unsupervised Bitext Mining and Translation via Self-trained Contextual EmbeddingsabstractWe describe an unsupervised method to create pseudo-parallel corpora for machine translation (MT) from unaligned text. We use multilingual BERT to create source and target sentence embeddings for nearest-neighbor search and adapt the model via self-training. We validate our technique by extracting parallel sentence pairs on the BUCC 2017 bitext mining task and observe up to a 24.5 point increase (absolute) in F1 scores over previous unsupervised methods. We then improve an XLM-based unsupervised neural MT system pre-trained on Wikipedia by supplementing it with pseudo-parallel text mined from the same corpus, boosting unsupervised translation performance by up to 3.5 BLEU on the WMT’14 French-English and WMT’16 German-English tasks and outperforming the previous state-of-the-art. Finally, we enrich the IWSLT’15 English-Vietnamese corpus with pseudo-parallel Wikipedia sentence pairs, yielding a 1.2 BLEU improvement on the low-resource MT task. We demonstrate that unsupervised bitext mining is an effective way of augmenting MT datasets and complements existing techniques like initializing with pre-trained contextual embeddings. Phillip Keung, Julian Salazar, Yichao Lu, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 4 |
| 2019 | ATOMIC: An Atlas of Machine Commonsense for If-Then ReasoningabstractWe present ATOMIC, an atlas of everyday commonsense reasoning, organized through 877k textual descriptions of inferential knowledge. Compared to existing resources that center around taxonomic knowledge, ATOMIC focuses on inferential knowledge organized as typed if-then relations with variables (e.g., “if X pays Y a compliment, then Y will likely return the compliment”). We propose nine if-then relation types to distinguish causes vs. effects, agents vs. themes, voluntary vs. involuntary events, and actions vs. mental states. By generatively training on the rich inferential knowledge described in ATOMIC, we show that neural models can acquire simple commonsense capabilities and reason about previously unseen events. Experimental results demonstrate that multitask models that incorporate the hierarchical structure of if-then relation types lead to more accurate inference compared to models trained in isolation, as measured by both automatic and human evaluation. Maarten Sap, Ronan Le Bras 0001, Emily Allaway, Chandra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A. Smith, Yejin Choi 0001 |
AAAI | 8 |
| 2019 | Sentence Mover's Similarity: Automatic Evaluation for Multi-Sentence TextsabstractFor evaluating machine-generated texts, automatic methods hold the promise of avoiding collection of human judgments, which can be expensive and time-consuming.The most common automatic metrics, like BLEU and ROUGE, depend on exact word matching, an inflexible approach for measuring semantic similarity.We introduce methods based on sentence mover's similarity; our automatic metrics evaluate text in a continuous space using word and sentence embeddings.We find that sentence-based metrics correlate with human judgments significantly better than ROUGE, both on machine-generated summaries (average length of 3.4 sentences) and human-authored essays (average length of 7.5).We also show that sentence mover's similarity can be used as a reward when learning a generation model via reinforcement learning; we present both automatic and human evaluations of summaries learned in this way, finding that our approach outperforms ROUGE. Elizabeth Clark, Asli Celikyilmaz, Noah A. Smith |
ACL (1) | 3 |
| 2019 | Variational Pretraining for Semi-supervised Text ClassificationabstractWe introduce VAMPIRE, 1 a lightweight pretraining framework for effective text classification when data and computing resources are limited.We pretrain a unigram document model as a variational autoencoder on in-domain, unlabeled data and use its internal states as features in a downstream classifier.Empirically, we show the relative strength of VAMPIRE against computationally expensive contextual embeddings and other popular semi-supervised baselines under low resource settings.We also find that fine-tuning to indomain data is crucial to achieving decent performance from contextual embeddings when working with limited supervision.We accompany this paper with code to pretrain and use VAMPIRE embeddings in downstream tasks. Suchin Gururangan, Tam Dang, Dallas Card, Noah A. Smith |
ACL (1) | 4 |
| 2019 | The Risk of Racial Bias in Hate Speech DetectionabstractWe investigate how annotators' insensitivity to differences in dialect can lead to racial bias in automatic hate speech detection models, potentially amplifying harm against minority populations.We first uncover unexpected correlations between surface markers of African American English (AAE) and ratings of toxicity in several widely-used hate speech datasets.Then, we show that models trained on these corpora acquire and propagate these biases, such that AAE tweets and tweets by self-identified African Americans are up to two times more likely to be labelled as offensive compared to others.Finally, we propose dialect and race priming as ways to reduce the racial bias in annotation, showing that when annotators are made explicitly aware of an AAE tweet's dialect they are significantly less likely to label the tweet as offensive. Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi 0001, Noah A. Smith |
ACL (1) | 5 |
| 2019 | Is Attention Interpretable?abstractAttention mechanisms have recently boosted performance on a range of NLP tasks.Because attention layers explicitly weight input components' representations, it is also often assumed that attention can be used to identify information that models found important (e.g., specific contextualized word tokens).We test whether that assumption holds by manipulating attention weights in already-trained text classification models and analyzing the resulting differences in their predictions.While we observe some ways in which higher attention weights correlate with greater impact on model predictions, we also find many ways in which this does not hold, i.e., where gradient-based rankings of attention weights better predict their effects than their magnitudes.We conclude that while attention noisily predicts input components' overall importance to a model, it is by no means a fail-safe indicator.1 Sofia Serrano, Noah A. Smith |
ACL (1) | 2 |
| 2019 | Evaluating Gender Bias in Machine TranslationabstractWe present the first challenge set and evaluation protocol for the analysis of gender bias in machine translation (MT).Our approach uses two recent coreference resolution datasets composed of English sentences which cast participants into non-stereotypical gender roles (e.g., "The doctor asked the nurse to help her in the operation").We devise an automatic gender bias evaluation method for eight target languages with grammatical gender, based on morphological analysis (e.g., the use of female inflection for the word "doctor").Our analyses show that four popular industrial MT systems and two recent state-of-the-art academic MT models are significantly prone to gender-biased translation errors for all tested target languages. Gabriel Stanovsky, Noah A. Smith, Luke Zettlemoyer |
ACL (1) | 2 |
| 2019 | Low-Resource Parsing with Crosslingual Contextualized RepresentationsabstractDespite advances in dependency parsing, languages with small treebanks still present challenges.We assess recent approaches to multilingual contextual word representations (CWRs), and compare them for crosslingual transfer from a language with a large treebank to a language with a small or nonexistent treebank, by sharing parameters between languages in the parser itself.We experiment with a diverse selection of languages in both simulated and truly low-resource scenarios, and show that multilingual CWRs greatly facilitate low-resource dependency parsing even without crosslingual supervision such as dictionaries or parallel text.Furthermore, we examine the non-contextual part of the learned language models (which we call a "decontextual probe") to demonstrate that polyglot language models better encode crosslingual lexical correspondence compared to aligned monolingual language models.This analysis provides further evidence that polyglot training is an effective approach to crosslingual transfer. Phoebe Mulcaire, Jungo Kasai, Noah A. Smith |
CoNLL | 3 |
| 2019 | Quoref: A Reading Comprehension Dataset with Questions Requiring Coreferential ReasoningabstractPradeep Dasigi, Nelson F. Liu, Ana Marasović, Noah A. Smith, Matt Gardner. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Pradeep Dasigi, Nelson F. Liu, Ana Marasovic, Noah A. Smith, Matt Gardner 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Show Your Work: Improved Reporting of Experimental ResultsabstractJesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz 0001, Noah A. Smith |
EMNLP/IJCNLP (1) | 5 |
| 2019 | RNN Architecture Learning with Sparse RegularizationabstractJesse Dodge, Roy Schwartz, Hao Peng, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jesse Dodge, Roy Schwartz 0001, Hao Peng 0009, Noah A. Smith |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Topics to Avoid: Demoting Latent Confounds in Text ClassificationabstractSachin Kumar, Shuly Wintner, Noah A. Smith, Yulia Tsvetkov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sachin Kumar 0009, Shuly Wintner, Noah A. Smith, Yulia Tsvetkov |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Robust Navigation with Language Pretraining and Stochastic SamplingabstractXiujun Li, Chunyuan Li, Qiaolin Xia, Yonatan Bisk, Asli Celikyilmaz, Jianfeng Gao, Noah A. Smith, Yejin Choi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiujun Li, Chunyuan Li, Qiaolin Xia, Yonatan Bisk, Asli Celikyilmaz, Jianfeng Gao 0001, Noah A. Smith, Yejin Choi 0001 |
EMNLP/IJCNLP (1) | 7 |
| 2019 | PaLM: A Hybrid Parser and Language ModelabstractHao Peng, Roy Schwartz, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Hao Peng 0009, Roy Schwartz 0001, Noah A. Smith |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Knowledge Enhanced Contextual Word RepresentationsabstractMatthew E. Peters, Mark Neumann, Robert Logan, Roy Schwartz, Vidur Joshi, Sameer Singh, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Matthew E. Peters, Mark Neumann, Robert L. Logan IV, Roy Schwartz 0001, Vidur Joshi, Sameer Singh 0001, Noah A. Smith |
EMNLP/IJCNLP (1) | 7 |
| 2019 | Measuring Online Debaters' Persuasive Skill from Text over TimeabstractOnline debates allow people to express their persuasive abilities and provide exciting opportunities for understanding persuasion. Prior studies have focused on studying persuasion in debate content, but without accounting for each debater’s history or exploring the progression of a debater’s persuasive ability. We study debater skill by modeling how participants progress over time in a collection of debates from Debate.org . We build on a widely used model of skill in two-player games and augment it with linguistic features of a debater’s content. We show that online debaters’ skill levels do tend to improve over time. Incorporating linguistic profiles leads to more robust skill estimation than winning records alone. Notably, we find that an interaction feature combining uncertainty cues (hedging) with terms strongly associated with either side of a particular debate (fightin’ words) is more predictive than either feature on its own, indicating the importance of fine- grained linguistic features. Kelvin Luu, Chenhao Tan, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 3 |
| 2019 | Analyzing Privacy Policies at Scale: From Crowdsourcing to Automated AnnotationsabstractWebsite privacy policies are often long and difficult to understand. While research shows that Internet users care about their privacy, they do not have the time to understand the policies of every website they visit, and most users hardly ever read privacy policies. Some recent efforts have aimed to use a combination of crowdsourcing, machine learning, and natural language processing to interpret privacy policies at scale, thus producing annotations for use in interfaces that inform Internet users of salient policy details. However, little attention has been devoted to studying the accuracy of crowdsourced privacy policy annotations, how crowdworker productivity can be enhanced for such a task, and the levels of granularity that are feasible for automatic analysis of privacy policies. In this article, we present a trajectory of work addressing each of these topics. We include analyses of crowdworker performance, evaluation of a method to make a privacy-policy oriented task easier for crowdworkers, a coarse-grained approach to labeling segments of policy text with descriptive themes, and a fine-grained approach to identifying user choices described in policy text. Together, the results from these efforts show the effectiveness of using automated and semi-automated methods for extracting from privacy policies the data practice details that are salient to Internet users’ interests. Shomir Wilson, Florian Schaub, Frederick Liu, Kanthashree Mysore Sathyendra, Daniel Smullen, Sebastian Zimmeck, Rohan Ramanath, Peter Story, Fei Liu 0004, Norman M. Sadeh, Noah A. Smith |
ACM Trans. Web | 11 |
| 2018 | Event2Mind: Commonsense Inference on Events, Intents, and ReactionsabstractWe investigate a new commonsense inference task: given an event described in a short free-form text ("X drinks coffee in the morning"), a system reasons about the likely intents ("X wants to stay awake") and reactions ("X feels alert") of the event's participants.To support this study, we construct a new crowdsourced corpus of 25,000 event phrases covering a diverse range of everyday events and situations.We report baseline performance on this task, demonstrating that neural encoder-decoder models can successfully compose embedding representations of previously unseen events and reason about the likely intents and reactions of the event participants.In addition, we demonstrate how commonsense inference on people's intents and reactions can help unveil the implicit gender inequality prevalent in modern movie scripts. 1 https://tinyurl.com/event2mind Hannah Rashkin, Maarten Sap, Emily Allaway, Noah A. Smith, Yejin Choi 0001 |
ACL (1) | 4 |
| 2018 | Neural Models for Documents with MetadataabstractMost real-world document collections involve various types of metadata, such as author, source, and date, and yet the most commonly-used approaches to modeling text corpora ignore this information.While specialized models have been developed for particular applications, few are widely used in practice, as customization typically requires derivation of a custom inference algorithm.In this paper, we build on recent advances in variational inference methods and propose a general neural framework, based on topic models, to enable flexible incorporation of metadata and allow for rapid exploration of alternative models.Our approach achieves strong performance, with a manageable tradeoff between perplexity, coherence, and sparsity.Finally, we demonstrate the potential of our framework through an exploration of a corpus of articles about US immigration. Dallas Card, Chenhao Tan, Noah A. Smith |
ACL (1) | 3 |
| 2018 | Backpropagating through Structured Argmax using a SPIGOTabstractWe introduce the structured projection of intermediate gradients optimization technique (SPIGOT), a new method for backpropagating through neural networks that include hard-decision structured predictions (e.g., parsing) in intermediate layers.SPIGOT requires no marginal inference, unlike structured attention networks (Kim et al., 2017) and some reinforcement learning-inspired solutions (Yogatama et al., 2017).Like socalled straight-through estimators (Hinton, 2012), SPIGOT defines gradient-like quantities associated with intermediate nondifferentiable operations, allowing backpropagation before and after them; SPIGOT's proxy aims to ensure that, after a parameter update, the intermediate structure will remain well-formed.We experiment on two structured NLP pipelines: syntactic-then-semantic dependency parsing, and semantic parsing followed by sentiment classification.We show that training with SPIGOT leads to a larger improvement on the downstream task than a modularly-trained pipeline, the straight-through estimator, and structured attention, reaching a new state of the art on semantic dependency parsing. Hao Peng 0009, Sam Thomson, Noah A. Smith |
ACL (1) | 3 |
| 2018 | Bridging CNNs, RNNs, and Weighted Finite-State MachinesabstractRecurrent and convolutional neural networks comprise two distinct families of models that have proven to be useful for encoding natural language utterances.In this paper we present SoPa, a new model that aims to bridge these two approaches.SoPa combines neural representation learning with weighted finite-state automata (WFSAs) to learn a soft version of traditional surface patterns.We show that SoPa is an extension of a one-layer CNN, and that such CNNs are equivalent to a restricted version of SoPa, and accordingly, to a restricted form of WFSA.Empirically, on three text classification tasks, SoPa is comparable or better than both a BiLSTM (RNN) baseline and a CNN baseline, and is particularly useful in small data settings. Roy Schwartz 0001, Sam Thomson, Noah A. Smith |
ACL (1) | 3 |
| 2018 | Rational RecurrencesabstractDespite the tremendous empirical success of neural models in natural language processing, many of them lack the strong intuitions that accompany classical machine learning approaches.Recently, connections have been shown between convolutional neural networks (CNNs) and weighted finite state automata (WFSAs), leading to new interpretations and insights.In this work, we show that some recurrent neural networks also share this connection to WFSAs.We characterize this connection formally, defining rational recurrences to be recurrent hidden state update functions that can be written as the Forward calculation of a finite set of WFSAs.We show that several recent neural models use rational recurrences.Our analysis provides a fresh view of these models and facilitates devising new neural architectures that draw inspiration from WFSAs.We present one such model, which performs better than two recent baselines on language modeling and text classification.Our results demonstrate that transferring intuitions from classical models like WFSAs can be an effective approach to designing and understanding neural models. Hao Peng 0009, Roy Schwartz 0001, Sam Thomson, Noah A. Smith |
EMNLP | 4 |
| 2018 | Syntactic Scaffolds for Semantic StructuresabstractWe introduce the syntactic scaffold, an approach to incorporating syntactic information into semantic tasks.Syntactic scaffolds avoid expensive syntactic processing at runtime, only making use of a treebank during training, through a multitask objective.We improve over strong baselines on PropBank semantics, frame semantics, and coreference resolution, achieving competitive performance on all three tasks. Swabha Swayamdipta, Sam Thomson, Kenton Lee, Luke Zettlemoyer, Chris Dyer, Noah A. Smith |
EMNLP | 6 |
| 2018 | Neural Cross-lingual Named Entity Recognition with Minimal ResourcesabstractFor languages with no annotated resources, unsupervised transfer of natural language processing models such as named-entity recognition (NER) from resource-rich languages would be an appealing capability.However, differences in words and word order across languages make it a challenging problem.To improve mapping of lexical items across languages, we propose a method that finds translations based on bilingual word embeddings.To improve robustness to word order differences, we propose to use self-attention, which allows for a degree of flexibility with respect to word order.We demonstrate that these methods achieve state-of-the-art or competitive NER performance on commonly tested languages under a cross-lingual setting, with much lower resource requirements than past approaches.We also evaluate the challenges of applying these methods to Uyghur, a lowresource language.1 Jiateng Xie, Zhilin Yang 0001, Graham Neubig, Noah A. Smith, Jaime G. Carbonell |
EMNLP | 4 |
| 2018 | Creative Writing with a Machine in the Loop: Case Studies on Slogans and StoriesabstractAs the quality of natural language generated by artificial intelligence systems improves, writing interfaces can support interventions beyond grammar-checking and spell-checking, such as suggesting content to spark new ideas. To explore the possibility of machine-in-the-loop creative writing, we performed two case studies using two system prototypes, one for short story writing and one for slogan writing. Participants in our studies were asked to write with a machine in the loop or alone (control condition). They assessed their writing and experience through surveys and an open-ended interview. We collected additional assessments of the writing from Amazon Mechanical Turk crowdworkers. Our findings indicate that participants found the process fun and helpful and could envision use cases for future systems. At the same time, machine suggestions do not necessarily lead to better written artifacts. We therefore suggest novel natural language models and design choices that may better support creative writing. Elizabeth Clark, Anne Spencer Ross, Chenhao Tan, Yangfeng Ji, Noah A. Smith |
IUI | 5 |
| 2018 | The Importance of Calibration for Estimating Proportions from AnnotationsabstractDallas Card, Noah A. Smith. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Dallas Card, Noah A. Smith |
NAACL-HLT | 2 |
| 2018 | Neural Text Generation in Stories Using Entity Representations as ContextabstractElizabeth Clark, Yangfeng Ji, Noah A. Smith. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Elizabeth Clark, Yangfeng Ji, Noah A. Smith |
NAACL-HLT | 3 |
| 2018 | Parsing Tweets into Universal DependenciesabstractYijia Liu, Yi Zhu, Wanxiang Che, Bing Qin, Nathan Schneider, Noah A. Smith. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Wanxiang Che, Bing Qin 0001, Nathan Schneider 0001, Noah A. Smith |
NAACL-HLT | 6 |
| 2018 | Learning Joint Semantic Parsers from Disjoint DataabstractHao Peng, Sam Thomson, Swabha Swayamdipta, Noah A. Smith. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Hao Peng 0009, Sam Thomson, Swabha Swayamdipta, Noah A. Smith |
NAACL-HLT | 4 |
| 2018 | "You are no Jack Kennedy": On Media Selection of Highlights from Presidential DebatesabstractPolitical speeches and debates play an important role in shaping the images of politicians, and the public often relies on media outlets to select bits of political communication from a large pool of utterances. It is an important research question to understand what factors impact this selection process. To quantitatively explore the selection process, we build a three- decade dataset of presidential debate transcripts and post-debate coverage. We first examine the effect of wording and propose a binary classification framework that controls for both the speaker and the debate situation. We find that crowdworkers can only achieve an accuracy of 60% in this task, indicating that media choices are not entirely obvious. Our classifiers outperform crowdworkers on average, mainly in primary debates. We also compare important factors from crowdworkers» free-form explanations with those from data-driven methods and find interesting differences. Few crowdworkers mentioned that "context matters", whereas our data show that well-quoted sentences are more distinct from the previous utterance by the same speaker than less-quoted sentences. Finally, we examine the aggregate effect of media preferences towards different wordings to understand the extent of fragmentation among media outlets. By analyzing a bipartite graph built from quoting behavior in our data, we observe a decreasing trend in bipartisan coverage. Chenhao Tan, Hao Peng 0009, Noah A. Smith |
WWW | 3 |
| 2018 | Framing Effects: Choice of Slogans Used to Advertise Online Experiments Can Boost Recruitment and Lead to Sample BiasesabstractOnline experimentation with volunteers relies on participants' non-financial motivations to complete a study, such as to altruistically support science or to compare oneself to others. Researchers rely on these motivations to attract study participants and often use incentives, like performance comparisons, to encourage participation. Often, these study incentives are advertised using a slogan (e.g., "What is your thinking style?''). Research on framing effects suggests that advertisement slogans attract people with varying demographics and motivations. Could the slogan advertisements for studies risk attracting only specific users? To investigate the existence of potential sample biases, we measured how different slogan frames affected which participants self-selected into studies. We found that slogan frames impact recruitment significantly; changing the slogan frame from a 'supporting science' frame to a 'comparing oneself to others' frame lead to a 9% increase in recruitment for some studies. Additionally, slogans framed as learning more about oneself attract participants significantly more motivated by boredom compared to other slogan frames. We discuss design implications for using frames to improve recruitment and mitigate sources of sample bias in online research with volunteers. Tal August, Nigini Oliveira, Chenhao Tan, Noah A. Smith, Katharina Reinecke |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2017 | Neural Discourse Structure for Text CategorizationabstractWe show that discourse structure, as defined by Rhetorical Structure Theory and provided by an existing discourse parser, benefits text categorization.Our approach uses a recursive neural network and a newly proposed attention mechanism to compute a representation of the text that focuses on salient content, from the perspective of both RST and the task.Experiments consider variants of the approach and illustrate its strengths and weaknesses. Yangfeng Ji, Noah A. Smith |
ACL (1) | 2 |
| 2017 | Deep Multitask Learning for Semantic Dependency ParsingabstractWe present a deep neural architecture that parses sentences into three semantic dependency graph formalisms.By using efficient, nearly arc-factored inference and a bidirectional-LSTM composed with a multi-layer perceptron, our base system is able to significantly improve the state of the art for semantic dependency parsing, without using hand-engineered features or syntax.We then explore two multitask learning approaches-one that shares parameters across formalisms, and one that uses higher-order structures to predict the graphs jointly.We find that both approaches improve performance across formalisms on average, achieving a new state of the art.Our code is open-source and available at https://github.com/Noahs-ARK/NeurboParser. Hao Peng 0009, Sam Thomson, Noah A. Smith |
ACL (1) | 3 |
| 2017 | Friendships, Rivalries, and Trysts: Characterizing Relations between Ideas in TextsabstractUnderstanding how ideas relate to each other is a fundamental question in many domains, ranging from intellectual history to public communication.Because ideas are naturally embedded in texts, we propose the first framework to systematically characterize the relations between ideas based on their occurrence in a corpus of documents, independent of how these ideas are represented.Combining two statistics-cooccurrence within documents and prevalence correlation over time-our approach reveals a number of different ways in which ideas can cooperate and compete.For instance, two ideas can closely track each other's prevalence over time, and yet rarely cooccur, almost like a "cold war" scenario.We observe that pairwise cooccurrence and prevalence correlation exhibit different distributions.We further demonstrate that our approach is able to uncover intriguing relations between ideas through in-depth case studies on news articles and research papers. Chenhao Tan, Dallas Card, Noah A. Smith |
ACL (1) | 3 |
| 2017 | The Effect of Different Writing Tasks on Linguistic Style: A Case Study of the ROC Story Cloze TaskabstractA writer's style depends not just on personal traits but also on her intent and mental state.In this paper, we show how variants of the same writing task can lead to measurable differences in writing style.We present a case study based on the story cloze task (Mostafazadeh et al., 2016a), where annotators were assigned similar writing tasks with different constraints: (1) writing an entire story, (2) adding a story ending for a given story context, and (3) adding an incoherent ending to a story.We show that a simple linear classifier informed by stylistic features is able to successfully distinguish among the three cases, without even looking at the story context.In addition, combining our stylistic features with language model predictions reaches state of the art performance on the story cloze challenge.Our results demonstrate that different task framings can dramatically affect the way people write. 1 1 This paper extends our LSDSem 2017 shared task submission (Schwartz et al., 2017). Roy Schwartz 0001, Maarten Sap, Ioannis Konstas, Leila Zilles, Yejin Choi 0001, Noah A. Smith |
CoNLL | 6 |
| 2017 | What Do Recurrent Neural Network Grammars Learn About Syntax?abstractAdhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, Noah A. Smith. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, Noah A. Smith |
EACL (1) | 6 |
| 2017 | Dynamic Entity Representations in Neural Language ModelsabstractUnderstanding a long document requires tracking how entities are introduced and evolve over time.We present a new type of language model, ENTITYNLM, that can explicitly model entities, dynamically update their representations, and contextually generate their mentions.Our model is generative and flexible; it can model an arbitrary number of entities in context while generating each entity mention at an arbitrary length.In addition, it can be used for several different tasks such as language modeling, coreference resolution, and entity prediction.Experimental results with all these tasks demonstrate that our model consistently outperforms strong baselines and prior work. Yangfeng Ji, Chenhao Tan, Sebastian Martschat, Yejin Choi 0001, Noah A. Smith |
EMNLP | 5 |
| 2017 | Multitask Learning with CTC and Segmental CRF for Speech RecognitionabstractSegmental conditional random fields (SCRFs) and connectionist temporal classification (CTC) are two sequence labeling methods used for end-to-end training of speech recognition models. Both models define a transcription probability by marginalizing decisions about latent segmentation alternatives to derive a sequence probability: the former uses a globally normalized joint model of segment labels and durations, and the latter classifies each frame as either an output symbol or a "continuation" of the previous label. In this paper, we train a recognition model by optimizing an interpolation between the SCRF and CTC losses, where the same recurrent neural network (RNN) encoder is used for feature extraction for both outputs. We find that this multitask objective improves recognition accuracy when decoding with either the SCRF or CTC models. Additionally, we show that CTC can also be used to pretrain the RNN encoder, which improves the convergence rate when learning the joint model. Liang Lu 0001, Lingpeng Kong, Chris Dyer, Noah A. Smith |
INTERSPEECH | 4 |
| 2017 | Greedy Transition-Based Dependency Parsing with Stack LSTMsabstractWe introduce a greedy transition-based parser that learns to represent parser states using recurrent neural networks. Our primary innovation that enables us to do this efficiently is a new control structure for sequential neural networks—the stack long short-term memory unit (LSTM). Like the conventional stack data structures used in transition-based parsers, elements can be pushed to or popped from the top of the stack in constant time, but, in addition, an LSTM maintains a continuous space embedding of the stack contents. Our model captures three facets of the parser's state: (i) unbounded look-ahead into the buffer of incoming words, (ii) the complete history of transition actions taken by the parser, and (iii) the complete contents of the stack of partially built tree fragments, including their internal structures. In addition, we compare two different word representations: (i) standard word vectors based on look-up tables and (ii) character-based models of words. Although standard word embedding models work well in all languages, the character-based models improve the handling of out-of-vocabulary words, particularly in morphologically rich languages. Finally, we discuss the use of dynamic oracles in training the parser. During training, dynamic oracles alternate between sampling parser states from the training data and from the model as it is being learned, making the model more robust to the kinds of errors that will be made at test time. Training our model with dynamic oracles yields a linear-time greedy parser with very competitive performance. Miguel Ballesteros, Chris Dyer, Yoav Goldberg, Noah A. Smith |
Comput. Linguistics | 4 |
| 2016 | Greedy, Joint Syntactic-Semantic Parsing with Stack LSTMsabstractWe present a transition-based parser that jointly produces syntactic and semantic dependencies. It learns a representation of the entire algorithm state, using stack long short-term memories. Our greedy inference algorithm has linear time, including feature extraction. On the CoNLL 2008--9 English shared tasks, we obtain the best published parsing performance among models that jointly learn syntax and semantics. Swabha Swayamdipta, Miguel Ballesteros, Chris Dyer, Noah A. Smith |
CoNLL | 4 |
| 2016 | Training with Exploration Improves a Greedy Stack LSTM ParserabstractWe adapt the greedy Stack-LSTM dependency parser of Dyer et al. (2015) to support a training-with-exploration procedure using dynamic oracles(Goldberg and Nivre, 2013) instead of cross-entropy minimization. This form of training, which accounts for model predictions at training time rather than assuming an error-free action history, improves parsing accuracies for both English and Chinese, obtaining very strong results for both languages. We discuss some modifications needed in order to get training with exploration to work well for a probabilistic neural-network. Miguel Ballesteros, Yoav Goldberg, Chris Dyer, Noah A. Smith |
EMNLP | 4 |
| 2016 | Analyzing Framing through the Casts of Characters in the NewsabstractWe present an unsupervised model for the discovery and clustering of latent "personas" (characterizations of entities).Our model simultaneously clusters documents featuring similar collections of personas.We evaluate this model on a collection of news articles about immigration, showing that personas help predict the coarse-grained framing annotations in the Media Frames Corpus.We also introduce automated model selection as a fair and robust form of feature evaluation. Dallas Card, Justin H. Gross, Amber E. Boydstun, Noah A. Smith |
EMNLP | 4 |
| 2016 | Character Sequence Models for Colorful WordsabstractWe present a neural network architecture to predict a point in space from the sequence of characters in the color's name. Using large scale color--name pairs obtained from an online design forum, we evaluate our model on a color Turing test and find that, given a name, the colors predicted by our model are preferred by annotators to names created by humans. Our datasets and demo system are available online at this http URL. Kazuya Kawakami, Chris Dyer, Bryan R. Routledge, Noah A. Smith |
EMNLP | 4 |
| 2016 | Distilling an Ensemble of Greedy Dependency Parsers into One MST ParserabstractWe introduce two first-order graph-based dependency parsers achieving a new state of the art.The first is a consensus parser built from an ensemble of independently trained greedy LSTM transition-based parsers with different random initializations.We cast this approach as minimum Bayes risk decoding (under the Hamming cost) and argue that weaker consensus within the ensemble is a useful signal of difficulty or ambiguity.The second parser is a "distillation" of the ensemble into a single model.We train the distillation parser using a structured hinge loss objective with a novel cost that incorporates ensemble uncertainty estimates for each possible attachment, thereby avoiding the intractable crossentropy computations required by applying standard distillation objectives to problems with structured outputs.The first-order distillation parser matches or surpasses the state of the art on English, Chinese, and German. Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Noah A. Smith |
EMNLP | 5 |
| 2016 | Semi-Supervised Learning of Sequence Models with Method of MomentsabstractWe propose a fast and scalable method for semi-supervised learning of sequence models, based on anchor words and moment matching. Our method can handle hidden Markov models with feature-based log-linear emissions. Unlike other semi-supervised methods, no decoding passes are necessary on the unlabeled data and no graph needs to be constructed— only one pass is necessary to collect moment statistics. The model parameters are estimated by solving a small quadratic program for each feature. Experiments on part-of-speech (POS) tagging for Twitter and for a low-resource language (Malagasy) show that our method can learn from very few annotated sentences. Zita Marinho, André F. T. Martins, Shay B. Cohen, Noah A. Smith |
EMNLP | 4 |
| 2016 | Friends with Motives: Using Text to Infer Influence on SCOTUSabstractWe present a probabilistic model of the influence of language on the behavior of the U.S. Supreme Court, specifically influence of amicus briefs on Court decisions and opinions.The approach assumes that amici are rational, utility-maximizing agents who try to win votes or affect the language of court opinions.Our model leads to improved predictions of justices' votes and perplexity of opinion language.It is amenable to inspection, allowing us to explore inferences about the persuasiveness of different amici and influenceability of different justices; these are consistent with earlier findings."Language is the central tool of our trade." Yanchuan Sim, Bryan R. Routledge, Noah A. Smith |
EMNLP | 3 |
| 2016 | Segmental Recurrent Neural Networks for End-to-End Speech RecognitionabstractWe study the segmental recurrent neural network for end-to-end acoustic modelling. This model connects the segmental conditional random field (CRF) with a recurrent neural network (RNN) used for feature extraction. Compared to most previous CRF-based acoustic models, it does not rely on an external system to provide features or segmentation boundaries. Instead, this model marginalises out all the possible segmentations, and features are extracted from the RNN trained together with the segmental CRF. In essence, this model is self-contained and can be trained end-to-end. In this paper, we discuss practical training and decoding issues as well as the method to speed up the training in the context of speech recognition. We performed experiments on the TIMIT dataset. We achieved 17.3 phone error rate (PER) from the first-pass decoding --- the best reported result using CRFs, despite the fact that we only used a zeroth-order CRF and without using any language model. Liang Lu 0001, Lingpeng Kong, Chris Dyer, Noah A. Smith, Steve Renals |
INTERSPEECH | 4 |
| 2016 | Recurrent Neural Network GrammarsabstractChris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, Noah A. Smith. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, Noah A. Smith |
HLT-NAACL | 4 |
| 2016 | Generation from Abstract Meaning Representation using Tree TransducersabstractJeffrey Flanigan, Chris Dyer, Noah A. Smith, Jaime Carbonell. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Jeffrey Flanigan, Chris Dyer, Noah A. Smith, Jaime G. Carbonell |
HLT-NAACL | 3 |
| 2016 | Crowdsourcing Annotations for Websites' Privacy Policies: Can It Really Work?abstractWebsite privacy policies are often long and difficult to understand. While research shows that Internet users care about their privacy, they do not have time to understand the policies of every website they visit, and most users hardly ever read privacy policies. Several recent efforts aim to crowdsource the interpretation of privacy policies and use the resulting annotations to build more effective user interfaces that provide users with salient policy summaries. However, very little attention has been devoted to studying the accuracy and scalability of crowdsourced privacy policy annotations, the types of questions crowdworkers can effectively answer, and the ways in which their productivity can be enhanced. Prior research indicates that most Internet users often have great difficulty understanding privacy policies, suggesting limits to the effectiveness of crowdsourcing approaches. In this paper, we assess the viability of crowdsourcing privacy policy annotations. Our results suggest that, if carefully deployed, crowdsourcing can indeed result in the generation of non-trivial annotations and can also help identify elements of ambiguity in policies. We further introduce and evaluate a method to improve the annotation process by predicting and highlighting paragraphs relevant to specific data practices. Shomir Wilson, Florian Schaub, Rohan Ramanath, Norman M. Sadeh, Fei Liu 0004, Noah A. Smith, Frederick Liu |
WWW | 6 |
| 2016 | Many Languages, One ParserabstractWe train one multilingual model for dependency parsing and use it to parse sentences in several languages. The parsing model uses (i) multilingual word clusters and embeddings; (ii) token-level language information; and (iii) language-specific features (fine-grained POS tags). This input representation enables the parser not only to parse effectively in multiple languages, but also to generalize across languages based on linguistic universals and typological similarities, making it more effective to learn from limited annotations. Our parser’s performance compares favorably to strong baselines in a range of data scenarios, including when the target language has a large treebank, a small treebank, or no treebank for training. Waleed Ammar, George Mulcaire, Miguel Ballesteros, Chris Dyer, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 5 |
| 2015 | Weakly-Supervised Grammar-Informed Bayesian CCG Parser LearningabstractCombinatory Categorial Grammar (CCG) is a lexicalized grammar formalism in which words are associated with categories that, in combination with a small universal set of rules, specify the syntactic configurations in which they may occur. Categories are selected from a large, recursively-defined set; this leads to high word-to-category ambiguity, which is one of the primary factors that make learning CCG parsers difficult, especially in the face of little data. Previous work has shown that learning sequence models for CCG tagging can be improved by using linguistically-motivated prior probability distributions over potential categories. We extend this approach to the task of learning a CCG parser from weak supervision. We present a Bayesian formulation for CCG parser induction that assumes only supervision in the form of an incomplete tag dictionary mapping some word types to sets of potential categories. Our approach outperforms a baseline model trained with uniform priors by exploiting universal, intrinsic properties of the CCG formalism to bias the model toward simpler, more cross-linguistically common categories. Dan Garrette, Chris Dyer, Jason Baldridge, Noah A. Smith |
AAAI | 4 |
| 2015 | The Utility of Text: The Case of Amicus Briefs and the Supreme CourtabstractWe explore the idea that authoring a piece of text is an act of maximizing one's expected utility.To make this idea concrete, we consider the societally important decisions of the Supreme Court of the United States.Extensive past work in quantitative political science provides a framework for empirically modeling the decisions of justices and how they relate to text.We incorporate into such a model texts authored by amici curiae (``friends of the court'' separate from the litigants) who seek to weigh in on the decision, then explicitly model their goals in a random utility model.We demonstrate the benefits of this approach in improved vote prediction and the ability to perform counterfactual analysis. Yanchuan Sim, Bryan R. Routledge, Noah A. Smith |
AAAI | 3 |
| 2015 | Transition-Based Dependency Parsing with Stack Long Short-Term MemoryabstractChris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, Noah A. Smith. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, Noah A. Smith |
ACL (1) | 5 |
| 2015 | Sparse Overcomplete Word Vector RepresentationsabstractManaal Faruqui, Yulia Tsvetkov, Dani Yogatama, Chris Dyer, Noah A. Smith. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Manaal Faruqui, Yulia Tsvetkov, Dani Yogatama, Chris Dyer, Noah A. Smith |
ACL (1) | 5 |
| 2015 | A Supertag-Context Model for Weakly-Supervised CCG Parser LearningabstractCombinatory Categorial Grammar (CCG) is a lexicalized grammar formalism in which words are associated with categories that specify the syntactic configurations in which they may occur.We present a novel parsing model with the capacity to capture the associative adjacent-category relationships intrinsic to CCG by parameterizing the relationships between each constituent label and the preterminal categories directly to its left and right, biasing the model toward constituent categories that can combine with their contexts.This builds on the intuitions of Klein and Manning's (2002) "constituentcontext" model, which demonstrated the value of modeling context, but has the advantage of being able to exploit the properties of CCG.Our experiments show that our model outperforms a baseline in which this context information is not captured. Dan Garrette, Chris Dyer, Jason Baldridge, Noah A. Smith |
CoNLL | 4 |
| 2015 | Improved Transition-based Parsing by Modeling Characters instead of Words with LSTMsabstractWe present extensions to a continuousstate dependency parsing method that makes it applicable to morphologically rich languages. Starting with a highperformance transition-based parser that uses long short-term memory (LSTM) recurrent neural networks to learn representations of the parser state, we replace lookup-based word representations with representations constructed from the orthographic representations of the words, also using LSTMs. This allows statistical sharing across word forms that are similar on the surface. Experiments for morphologically rich languages show that the parsing model benefits from incorporating the character-based encodings of words. Miguel Ballesteros, Chris Dyer, Noah A. Smith |
EMNLP | 3 |
| 2015 | Open Extraction of Fine-Grained Political StatementsabstractText data has recently been used as evidence in estimating the political ideologies of individuals, including political elites and social media users.While inferences about people are often the intrinsic quantity of interest, we draw inspiration from open information extraction to identify a new task: inferring the political import of propositions like OBAMA IS A SOCIAL-IST.We present several models that exploit the structure that exists between people and the assertions they make to learn latent positions of people and propositions at the same time, and we evaluate them on a novel dataset of propositions judged on a political spectrum. David Bamman, Noah A. Smith |
EMNLP | 2 |
| 2015 | A Utility Model of Authors in the Scientific CommunityabstractAuthoring a scientific paper is a complex process involving many decisions.We introduce a probabilistic model of some of the important aspects of that process: that authors have individual preferences, that writing a paper requires trading off among the preferences of authors as well as extrinsic rewards in the form of community response to their papers, that preferences (of individuals and the community) and tradeoffs vary over time.Variants of our model lead to improved predictive accuracy of citations given texts and texts given authors.Further, our model's posterior suggests an interesting relationship between seniority and author choices. Yanchuan Sim, Bryan R. Routledge, Noah A. Smith |
EMNLP | 3 |
| 2015 | Extractive Summarization by Maximizing Semantic VolumeabstractThe most successful approaches to extrac-tive text summarization seek to maximize bigram coverage subject to a budget con-straint. In this work, we propose instead to maximize semantic volume. We em-bed each sentence in a semantic space and construct a summary by choosing a sub-set of sentences whose convex hull max-imizes volume in that space. We provide a greedy algorithm based on the Gram-Schmidt process to efficiently perform volume maximization. Our method out-performs the state-of-the-art summariza-tion approaches on benchmark datasets. 1 Dani Yogatama, Fei Liu 0004, Noah A. Smith |
EMNLP | 3 |
| 2015 | Bayesian Optimization of Text RepresentationsabstractWhen applying machine learning to problems in NLP, there are many choices to make about how to represent input texts.They can have a big effect on performance, but they are often uninteresting to researchers or practitioners who simply need a module that performs well.We apply sequential model-based optimization over this space of choices and show that it makes standard linear models competitive with more sophisticated, expensive state-ofthe-art methods based on latent variables or neural networks on various topic classification and sentiment analysis problems.Our approach is a first step towards black-box NLP systems that work with raw text and do not require manual tuning. Dani Yogatama, Lingpeng Kong, Noah A. Smith |
EMNLP | 3 |
| 2015 | Learning Word Representations with Hierarchical Sparse CodingabstractWe propose a new method for learning word representations using hierarchical regularization in sparse coding inspired by the linguistic study of word meanings. We show an efficient learning algorithm based on stochastic proximal methods that is significantly faster than previous approaches, making it possible to perform hierarchical sparse coding on a corpus of billions of word tokens. Experiments on various benchmark tasks—word similarity ranking, syntactic and semantic analogies, sentence completion, and sentiment analysis—demonstrate that the method outperforms or is competitive with state-of-the-art methods. Dani Yogatama, Manaal Faruqui, Chris Dyer, Noah A. Smith |
ICML | 4 |
| 2015 | Contextualized Sarcasm Detection on Twitter
David Bamman, Noah A. Smith |
ICWSM | 2 |
| 2015 | Toward Abstractive Summarization Using Semantic RepresentationsabstractFei Liu, Jeffrey Flanigan, Sam Thomson, Norman Sadeh, Noah A. Smith. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Fei Liu 0004, Jeffrey Flanigan, Sam Thomson, Norman M. Sadeh, Noah A. Smith |
HLT-NAACL | 5 |
| 2015 | Retrofitting Word Vectors to Semantic LexiconsabstractManaal Faruqui, Jesse Dodge, Sujay Kumar Jauhar, Chris Dyer, Eduard Hovy, Noah A. Smith. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Manaal Faruqui, Jesse Dodge, Sujay Kumar Jauhar, Chris Dyer, Eduard H. Hovy, Noah A. Smith |
HLT-NAACL | 6 |
| 2015 | Transforming Dependencies into Phrase StructuresabstractWe present a new algorithm for transforming dependency parse trees into phrase-structure parse trees.We cast the problem as structured prediction and learn a statistical model.Our algorithm is faster than traditional phrasestructure parsing and achieves 90.4% English parsing accuracy and 82.4% Chinese parsing accuracy, near to the state of the art on both benchmarks. Lingpeng Kong, Alexander M. Rush, Noah A. Smith |
HLT-NAACL | 3 |
| 2015 | A Corpus and Model Integrating Multiword Expressions and SupersensesabstractThis paper introduces a task of identifying and semantically classifying lexical expressions in running text.We investigate the online reviews genre, adding semantic supersense annotations to a 55,000 word English corpus that was previously annotated for multiword expressions.The noun and verb supersenses apply to full lexical expressions, whether single-or multiword.We then present a sequence tagging model that jointly infers lexical expressions and their supersenses.Results show that even with our relatively small training corpus in a noisy domain, the joint task can be performed to attain 70% class labeling F 1 . Nathan Schneider 0001, Noah A. Smith |
HLT-NAACL | 2 |
| 2015 | Modeling User Arguments, Interactions, and Attributes for Stance Prediction in Online Debate ForumsabstractOnline debate forums are important social media for people to voice their opinions and debate with each other. Mining user stances or viewpoints from these forums has been a popular research topic. However, most current work does not address an important problem: for a specific issue, there may not be many users participating and expressing their opinions. Despite the sparsity of user stances, users may provide rich side information; for example, users may write arguments to back up their stances, interact with each other, and provide biographical information. In this work, we propose an integrated model to leverage side information. Our proposed method is a regression-based latent factor model which jointly models user arguments, interactions, and attributes. Our method can perform stance prediction for both warm-start and cold-start users. We demonstrate in experiments that our method has promising results on both micro-level and macro-level stance prediction. Minghui Qiu, Yanchuan Sim, Noah A. Smith, Jing Jiang 0001 |
SDM | 3 |
| 2015 | AD3: alternating directions dual decomposition for MAP inference in graphical models
André F. T. Martins, Mário A. T. Figueiredo, Pedro M. Q. Aguiar, Noah A. Smith, Eric P. Xing |
J. Mach. Learn. Res. | 4 |
| 2014 | A Bayesian Mixed Effects Model of Literary CharacterabstractWe consider the problem of automatically inferring latent character types in a collection of 15,099 English novels published between 1700 and 1899.Unlike prior work in which character types are assumed responsible for probabilistically generating all text associated with a character, we introduce a model that employs multiple effects to account for the influence of extra-linguistic information (such as author).In an empirical evaluation, we find that this method leads to improved agreement with the preregistered judgments of a literary scholar, complementing the results of alternative models. David Bamman, Ted Underwood, Noah A. Smith |
ACL (1) | 3 |
| 2014 | A Discriminative Graph-Based Parser for the Abstract Meaning RepresentationabstractMeaning Representation (AMR) is a semantic formalism for which a growing set of annotated examples is available.We introduce the first approach to parse sentences into this representation, providing a strong baseline for future improvement.The method is based on a novel algorithm for finding a maximum spanning, connected subgraph, embedded within a Lagrangian relaxation of an optimization problem that imposes linguistically inspired constraints.Our approach is described in the general framework of structured prediction, allowing future incorporation of additional features and constraints, and may extend to other formalisms as well.Our open-source system, JAMR, is available at: Jeffrey Flanigan, Sam Thomson, Jaime G. Carbonell, Chris Dyer, Noah A. Smith |
ACL (1) | 5 |
| 2014 | Linguistic Structured Sparsity in Text CategorizationabstractWe introduce three linguistically motivated structured regularizers based on parse trees, topics, and hierarchical word clusters for text categorization.These regularizers impose linguistic bias in feature weights, enabling us to incorporate prior knowledge into conventional bagof-words models.We show that our structured regularizers consistently improve classification accuracies compared to standard regularizers that penalize features in isolation (such as lasso, ridge, and elastic net regularizers) on a range of datasets for various text prediction problems: topic classification, sentiment analysis, and forecasting. Dani Yogatama, Noah A. Smith |
ACL (1) | 2 |
| 2014 | A Step Towards Usable Privacy Policy: Automatic Alignment of Privacy Statements
Fei Liu 0004, Rohan Ramanath, Norman M. Sadeh, Noah A. Smith |
COLING | 4 |
| 2014 | Weakly-Supervised Bayesian Learning of a CCG SupertaggerabstractWe present a Bayesian formulation for weakly-supervised learning of a Combinatory Categorial Grammar (CCG) supertagger with an HMM.We assume supervision in the form of a tag dictionary, and our prior encourages the use of crosslinguistically common category structures as well as transitions between tags that can combine locally according to CCG's combinators.Our prior is theoretically appealing since it is motivated by languageindependent, universal properties of the CCG formalism.Empirically, we show that it yields substantial improvements over previous work that used similar biases to initialize an EM-based learner.Additional gains are obtained by further shaping the prior with corpus-specific information that is extracted automatically from raw text and a tag dictionary. Dan Garrette, Chris Dyer, Jason Baldridge, Noah A. Smith |
CoNLL | 4 |
| 2014 | A Dependency Parser for TweetsabstractWe describe a new dependency parser for English tweets, TWEEBOPARSER. The parser builds on several contributions: new syntactic annotations for a corpus of tweets (TWEEBANK), with conventions informed by the domain; adaptations to a statistical parsing algorithm; and a new approach to exploiting out-of-domain Penn Treebank data. Our experiments show that the parser achieves over 80% unlabeled attachment accuracy on our new, high-quality test set and measure the benefit of our contributions. Our dataset and parser can be found at http://www.ark.cs.cmu.edu/TweetNLP. Lingpeng Kong, Nathan Schneider 0001, Swabha Swayamdipta, Archna Bhatia, Chris Dyer, Noah A. Smith |
EMNLP | 6 |
| 2014 | Identifying Relevant Text Fragments to Help Crowdsource Privacy Policy AnnotationsabstractIn today's age of big data, websites are collecting an increasingly wide variety of information about their users. The texts of websites' privacy policies, which serve as legal agreements between service providers and users, are often long and difficult to understand. Automated analysis of those texts has the potential to help users better understand the implications of agreeing to such policies. In this work, we present a technique that combines machine learning and crowdsourcing to semi-automatically extract key aspects of website privacy policies that is scalable, fast, and cost-effective. Rohan Ramanath, Florian Schaub, Shomir Wilson, Fei Liu 0004, Norman M. Sadeh, Noah A. Smith |
HCOMP | 6 |
| 2014 | Making the Most of Bag of Words: Sentence Regularization with Alternating Direction Method of MultipliersabstractIn many high-dimensional learning problems, only some parts of an observation are important to the prediction task; for example, the cues to correctly categorizing a document may lie in a handful of its sentences. We introduce a learning algorithm that exploits this intuition by encoding it in a regularizer. Specifically, we apply the sparse overlapping group lasso with one group for every bundle of features occurring together in a training-data sentence, leading to thousands to millions of overlapping groups. We show how to efficiently solve the resulting optimization challenge using the alternating directions method of multipliers. We find that the resulting method significantly outperforms competitive baselines (standard ridge, lasso, and elastic net regularizers) on a suite of real-world text categorization problems. Dani Yogatama, Noah A. Smith |
ICML | 2 |
| 2014 | Comprehensive Annotation of Multiword Expressions in a Social Web Corpus
Nathan Schneider 0001, Spencer Onuffer, Nora Kazour, Emily Danchik, Michael T. Mordowanec, Henrietta Conrad, Noah A. Smith |
LREC | 7 |
| 2014 | Conditional Random Field Autoencoders for Unsupervised Structured Prediction
Waleed Ammar, Chris Dyer, Noah A. Smith |
NIPS | 3 |
| 2014 | Frame-Semantic ParsingabstractFrame semantics is a linguistic theory that has been instantiated for English in the FrameNet lexicon. We solve the problem of frame-semantic parsing using a two-stage statistical model that takes lexical targets (i.e., content words and phrases) in their sentential contexts and predicts frame-semantic structures. Given a target in context, the first stage disambiguates it to a semantic frame. This model uses latent variables and semi-supervised learning to improve frame disambiguation for targets unseen at training time. The second stage finds the target's locally expressed semantic arguments. At inference time, a fast exact dual decomposition algorithm collectively predicts all the arguments of a frame at once in order to respect declaratively stated linguistic constraints, resulting in qualitatively better structures than naïve local predictors. Both components are feature-based and discriminatively trained on a small set of annotated frame-semantic parses. On the SemEval 2007 benchmark data set, the approach, along with a heuristic identifier of frame-evoking targets, outperforms the prior state of the art by significant margins. Additionally, we present experiments on the much larger FrameNet 1.5 data set. We have released our frame-semantic parser as open-source software. Dipanjan Das 0001, Desai Chen, André F. T. Martins, Nathan Schneider 0001, Noah A. Smith |
Comput. Linguistics | 5 |
| 2014 | Phrase Dependency Machine Translation with Quasi-Synchronous Tree-to-Tree FeaturesabstractRecent research has shown clear improvement in translation quality by exploiting linguistic syntax for either the source or target language. However, when using syntax for both languages (“tree-to-tree” translation), there is evidence that syntactic divergence can hamper the extraction of useful rules (Ding and Palmer 2005 ). Smith and Eisner ( 2006 ) introduced quasi-synchronous grammar, a formalism that treats non-isomorphic structure softly using features rather than hard constraints. Although a natural fit for translation modeling, its flexibility has proved challenging for building real-world systems. In this article, we present a tree-to-tree machine translation system inspired by quasi-synchronous grammar. The core of our approach is a new model that combines phrases and dependency syntax, integrating the advantages of phrase-based and syntax-based translation. We report statistically significant improvements over a phrase-based baseline on five of seven test sets across four language pairs. We also present encouraging preliminary results on the use of unsupervised dependency parsing for syntax-based machine translation. Kevin Gimpel, Noah A. Smith |
Comput. Linguistics | 2 |
| 2014 | Unsupervised Discovery of Biographical Structure from TextabstractWe present a method for discovering abstract event classes in biographies, based on a probabilistic latent-variable model. Taking as input timestamped text, we exploit latent correlations among events to learn a set of event classes (such as Born, Graduates High School, and Becomes Citizen), along with the typical times in a person’s life when those events occur. In a quantitative evaluation at the task of predicting a person’s age for a given event, we find that our generative model outperforms a strong linear regression baseline, along with simpler variants of the model that ablate some features. The abstract event classes that we learn allow us to perform a large-scale analysis of 242,970 Wikipedia biographies. Though it is known that women are greatly underrepresented on Wikipedia—not only as editors (Wikipedia, 2011) but also as subjects of articles (Reagle and Rhue, 2011)—we find that there is a bias in their characterization as well, with biographies of women containing significantly more emphasis on events of marriage and divorce than biographies of men. David Bamman, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 2 |
| 2014 | Discriminative Lexical Semantic Segmentation with Gaps: Running the MWE GamutabstractWe present a novel representation, evaluation measure, and supervised models for the task of identifying the multiword expressions (MWEs) in a sentence, resulting in a lexical semantic segmentation. Our approach generalizes a standard chunking representation to encode MWEs containing gaps, thereby enabling efficient sequence tagging algorithms for feature-rich discriminative models. Experiments on a new dataset of English web text offer the first linguistically-driven evaluation of MWE identification with truly heterogeneous expression types. Our statistical sequence model greatly outperforms a lookup-based segmentation procedure, achieving nearly 60% F1 for MWE identification. Nathan Schneider 0001, Emily Danchik, Chris Dyer, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 4 |
| 2014 | Dynamic Language Models for Streaming TextabstractWe present a probabilistic language model that captures temporal dynamics and conditions on arbitrary non-linguistic context features. These context features serve as important indicators of language changes that are otherwise difficult to capture using text data by itself. We learn our model in an efficient online fashion that is scalable for large, streaming data. With five streaming datasets from two different genres—economics news articles and social media—we evaluate our model on the task of sequential language modeling. Our model consistently outperforms competing models. Dani Yogatama, Chong Wang 0002, Bryan R. Routledge, Noah A. Smith, Eric P. Xing |
Trans. Assoc. Comput. Linguistics | 4 |
| 2013 | Learning Latent Personas of Film Characters
David Bamman, Brendan T. O'Connor 0001, Noah A. Smith |
ACL (1) | 3 |
| 2013 | Learning to Extract International Relations from Political Context
Brendan T. O'Connor 0001, Brandon M. Stewart, Noah A. Smith |
ACL (1) | 3 |
| 2013 | Translating into Morphologically Rich Languages with Synthetic PhrasesabstractTranslation into morphologically rich languages is an important but recalcitrant problem in MT.We present a simple and effective approach that deals with the problem in two phases.First, a discriminative model is learned to predict inflections of target words from rich source-side annotations.Then, this model is used to create additional sentencespecific word-and phrase-level translations that are added to a standard translation model as "synthetic" phrases.Our approach relies on morphological analysis of the target language, but we show that an unsupervised Bayesian model of morphology can successfully be used in place of a supervised analyzer.We report significant improvements in translation quality when translating from English to Russian, Hebrew and Swahili. Victor Chahuneau, Eva Schlinger, Noah A. Smith, Chris Dyer |
EMNLP | 3 |
| 2013 | Learning Topics and Positions from DebatepediaabstractWe explore Debatepedia, a communityauthored encyclopedia of sociopolitical debates, as evidence for inferring a lowdimensional, human-interpretable representation in the domain of issues and positions.We introduce a generative model positing latent topics and cross-cutting positions that gives special treatment to person mentions and opinion words.We evaluate the resulting representation's usefulness in attaching opinionated documents to arguments and its consistency with human judgments about positions.0.1 0.2 0.3 0.4 0.5 comment, minimum, wage, poverty, capitalism nuclear, weapons, iran, states, threat party, vote, republican, political, voters energy, gas, power, fuel, wind Swapna Gottipati, Minghui Qiu, Yanchuan Sim, Jing Jiang 0001, Noah A. Smith |
EMNLP | 5 |
| 2013 | Measuring Ideological Proportions in Political SpeechesabstractWe seek to measure political candidates' ideological positioning from their speeches.To accomplish this, we infer ideological cues from a corpus of political writings annotated with known ideologies.We then represent the speeches of U.S. Presidential candidates as sequences of cues and lags (filler distinguished only by its length in words).We apply a domain-informed Bayesian HMM to infer the proportions of ideologies each candidate uses in each campaign.The results are validated against a set of preregistered, domain expertauthored hypotheses. Yanchuan Sim, Brice D. L. Acree, Justin H. Gross, Noah A. Smith |
EMNLP | 4 |
| 2013 | A Penny for Your Tweets: Campaign Contributions and Capitol Hill Microblogs
Tae Yano, Dani Yogatama, Noah A. Smith |
ICWSM | 3 |
| 2013 | Knowledge-Rich Morphological Priors for Bayesian Language Models
Victor Chahuneau, Noah A. Smith, Chris Dyer |
HLT-NAACL | 2 |
| 2013 | A Simple, Fast, and Effective Reparameterization of IBM Model 2
Chris Dyer, Victor Chahuneau, Noah A. Smith |
HLT-NAACL | 3 |
| 2013 | Improved Part-of-Speech Tagging for Online Conversational Text with Word Clusters
Olutobi Owoputi, Brendan T. O'Connor 0001, Chris Dyer, Kevin Gimpel, Nathan Schneider 0001, Noah A. Smith |
HLT-NAACL | 6 |
| 2013 | Supersense Tagging for Arabic: the MT-in-the-Middle Attack
Nathan Schneider 0001, Behrang Mohit, Chris Dyer, Kemal Oflazer, Noah A. Smith |
HLT-NAACL | 5 |
| 2012 | A Probabilistic Model for Canonicalizing Named Entity Mentions
Dani Yogatama, Yanchuan Sim, Noah A. Smith |
ACL (1) | 3 |
| 2012 | Recall-Oriented Learning of Named Entities in Arabic Wikipedia
Behrang Mohit, Nathan Schneider 0001, Rishav Bhowmick, Kemal Oflazer, Noah A. Smith |
EACL | 5 |
| 2012 | Word Salad: Relating Food Prices and Descriptions
Victor Chahuneau, Kevin Gimpel, Bryan R. Routledge, Lily Scherlis, Noah A. Smith |
EMNLP-CoNLL | 5 |
| 2012 | Graph-Based Lexicon Expansion with Sparsity-Inducing Penalties
Dipanjan Das 0001, Noah A. Smith |
HLT-NAACL | 2 |
| 2012 | Structured Ramp Loss Minimization for Machine Translation
Kevin Gimpel, Noah A. Smith |
HLT-NAACL | 2 |
| 2012 | Concavity and Initialization for Unsupervised Dependency Parsing
Kevin Gimpel, Noah A. Smith |
HLT-NAACL | 2 |
| 2012 | Structured Sparsity in Natural Language Processing: Models, Algorithms and Applications
André F. T. Martins, Mário A. T. Figueiredo, Noah A. Smith |
HLT-NAACL | 3 |
| 2012 | Textual Predictors of Bill Survival in Congressional Committees
Tae Yano, Noah A. Smith, John D. Wilkerson |
HLT-NAACL | 2 |
| 2012 | Empirical Risk Minimization for Probabilistic Grammars: Sample Complexity and Hardness of LearningabstractProbabilistic grammars are generative statistical models that are useful for compositional and sequential structures. They are used ubiquitously in computational linguistics. We present a framework, reminiscent of structural risk minimization, for empirical risk minimization of probabilistic grammars using the log-loss. We derive sample complexity bounds in this framework that apply both to the supervised setting and the unsupervised setting. By making assumptions about the underlying distribution that are appropriate for natural language scenarios, we are able to derive distribution-dependent sample complexity bounds for probabilistic grammars. We also give simple algorithms for carrying out empirical risk minimization using this framework in both the supervised and unsupervised settings. In the unsupervised case, we show that the problem of minimizing empirical risk is NP-hard. We therefore suggest an approximate algorithm, similar to expectation-maximization, to minimize the empirical risk. Shay B. Cohen, Noah A. Smith |
Comput. Linguistics | 2 |
| 2011 | Semi-Supervised Frame-Semantic Parsing for Unknown Predicates
Dipanjan Das 0001, Noah A. Smith |
ACL | 2 |
| 2011 | Unsupervised Word Alignment with Arbitrary Features
Chris Dyer, Jonathan H. Clark, Alon Lavie, Noah A. Smith |
ACL | 4 |
| 2011 | Discovering Sociolinguistic Associations with Structured Sparsity
Jacob Eisenstein, Noah A. Smith, Eric P. Xing |
ACL | 2 |
| 2011 | Unsupervised Structure Prediction with Non-Parallel Multilingual Guidance
Shay B. Cohen, Dipanjan Das 0001, Noah A. Smith |
EMNLP | 3 |
| 2011 | Quasi-Synchronous Phrase Dependency Grammars for Machine Translation
Kevin Gimpel, Noah A. Smith |
EMNLP | 2 |
| 2011 | Dual Decomposition with Many Overlapping Components
André F. T. Martins, Noah A. Smith, Mário A. T. Figueiredo, Pedro M. Q. Aguiar |
EMNLP | 2 |
| 2011 | Structured Sparsity in Structured Prediction
André F. T. Martins, Noah A. Smith, Mário A. T. Figueiredo, Pedro M. Q. Aguiar |
EMNLP | 2 |
| 2011 | Predicting a Scientific Community's Response to an Article
Dani Yogatama, Michael Heilman, Brendan T. O'Connor 0001, Chris Dyer, Bryan R. Routledge, Noah A. Smith |
EMNLP | 6 |
| 2011 | An Augmented Lagrangian Approach to Constrained MAP Inference
André F. T. Martins, Mário A. T. Figueiredo, Pedro M. Q. Aguiar, Noah A. Smith, Eric P. Xing |
ICML | 4 |
| 2011 | Products of weighted logic programsabstractAbstract Weighted logic programming, a generalization of bottom-up logic programming, is a well-suited framework for specifying dynamic programming algorithms. In this setting, proofs correspond to the algorithm's output space, such as a path through a graph or a grammatical derivation, and are given a real-valued score (often interpreted as a probability) that depends on the real weights of the base axioms used in the proof. The desired output is a function over all possible proofs, such as a sum of scores or an optimal score. We describe theproducttransformation, which can merge two weighted logic programs into a new one. The resulting program optimizes a product of proof scores from the original programs, constituting a scoring function known in machine learning as a “product of experts.” Through the addition of intuitive constraining side conditions, we show that several important dynamic programming algorithms can be derived by applyingproductto weighted logic programs corresponding tosimplerweighted logic programs. In addition, we show how the computation of Kullback–Leibler divergence, an information-theoretic measure, can be interpreted usingproduct. Shay B. Cohen, Robert J. Simmons, Noah A. Smith |
Theory Pract. Log. Program. | 3 |
| 2010 | Viterbi Training for PCFGs: Hardness Results and Competitiveness of Uniform Initialization
Shay B. Cohen, Noah A. Smith |
ACL | 2 |
| 2010 | Nonparametric Word Segmentation for Machine Translation
ThuyLinh Nguyen, Stephan Vogel, Noah A. Smith |
COLING | 3 |
| 2010 | Distributed Asynchronous Online Learning for Natural Language Processing
Kevin Gimpel, Dipanjan Das 0001, Noah A. Smith |
CoNLL | 3 |
| 2010 | A Latent Variable Model for Geographic Lexical Variation
Jacob Eisenstein, Brendan T. O'Connor 0001, Noah A. Smith, Eric P. Xing |
EMNLP | 3 |
| 2010 | Turbo Parsers: Dependency Parsing by Approximate Variational Inference
André F. T. Martins, Noah A. Smith, Eric P. Xing, Pedro M. Q. Aguiar, Mário A. T. Figueiredo |
EMNLP | 2 |
| 2010 | From Tweets to Polls: Linking Text Sentiment to Public Opinion Time Series
Brendan T. O'Connor 0001, Ramnath Balasubramanyan, Bryan R. Routledge, Noah A. Smith |
ICWSM | 4 |
| 2010 | What's Worthy of Comment? Content and Comment Volume in Political Blogs
Tae Yano, Noah A. Smith |
ICWSM | 2 |
| 2010 | Variational Inference for Adaptor Grammars
Shay B. Cohen, David M. Blei, Noah A. Smith |
HLT-NAACL | 3 |
| 2010 | Probabilistic Frame-Semantic Parsing
Dipanjan Das 0001, Nathan Schneider 0001, Desai Chen, Noah A. Smith |
HLT-NAACL | 4 |
| 2010 | Softmax-Margin CRFs: Training Log-Linear Models with Cost Functions
Kevin Gimpel, Noah A. Smith |
HLT-NAACL | 2 |
| 2010 | Good Question! Statistical Ranking for Question Generation
Michael Heilman, Noah A. Smith |
HLT-NAACL | 2 |
| 2010 | Tree Edit Models for Recognizing Textual Entailments, Paraphrases, and Answers to Questions
Michael Heilman, Noah A. Smith |
HLT-NAACL | 2 |
| 2010 | Movie Reviews and Revenues: An Experiment in Text Regression
Mahesh Joshi, Dipanjan Das 0001, Kevin Gimpel, Noah A. Smith |
HLT-NAACL | 4 |
| 2010 | Empirical Risk Minimization with Approximations of Probabilistic GrammarsabstractProbabilistic grammars are generative statistical models that are useful for compositional and sequential structures. We present a framework, reminiscent of structural risk minimization, for empirical risk minimization of the parameters of a fixed probabilistic grammar using the log-loss. We derive sample complexity bounds in this framework that apply both to the supervised setting and the unsupervised setting. Shay B. Cohen, Noah A. Smith |
NIPS | 2 |
| 2010 | Covariance in Unsupervised Learning of Probabilistic Grammars
Shay B. Cohen, Noah A. Smith |
J. Mach. Learn. Res. | 2 |
| 2009 | Paraphrase Identification as Probabilistic Quasi-Synchronous Recognition
Dipanjan Das 0001, Noah A. Smith |
ACL/IJCNLP | 2 |
| 2009 | Concise Integer Linear Programming Formulations for Dependency Parsing
André F. T. Martins, Noah A. Smith, Eric P. Xing |
ACL/IJCNLP | 2 |
| 2009 | Cube Summing, Approximate Inference with Non-Local Features, and Dynamic Programming without Semirings
Kevin Gimpel, Noah A. Smith |
EACL | 2 |
| 2009 | Feature-Rich Translation by Quasi-Synchronous Lattice Parsing
Kevin Gimpel, Noah A. Smith |
EMNLP | 2 |
| 2009 | Polyhedral outer approximations with application to natural language parsingabstractRecent approaches to learning structured predictors often require approximate inference for tractability; yet its effects on the learned model are unclear. Meanwhile, most learning algorithms act as if computational cost was constant within the model class. This paper sheds some light on the first issue by establishing risk bounds for max-margin learning with LP relaxed inference and addresses the second issue by proposing a new paradigm that attempts to penalize “timeconsuming” hypotheses. Our analysis relies on a geometric characterization of the outer polyhedra associated with the LP relaxation. We then apply these techniques to the problem of dependency parsing, for which a concise LP formulation is provided that handles non-local output features. A significant improvement is shown over arc-factored models. 1. André F. T. Martins, Noah A. Smith, Eric P. Xing |
ICML | 2 |
| 2009 | Tutorial summary: Structured prediction for natural language processingabstractNo abstract available. Noah A. Smith |
ICML | 1 |
| 2009 | From Episodes to Sagas: Understanding the News by Identifying Temporally Related Story Sequences
Ramnath Balasubramanyan, Frank Lin, William W. Cohen, Matthew Hurst, Noah A. Smith |
ICWSM | 5 |
| 2009 | Shared Logistic Normal Distributions for Soft Parameter Tying in Unsupervised Grammar Induction
Shay B. Cohen, Noah A. Smith |
HLT-NAACL | 2 |
| 2009 | Predicting Risk from Financial Reports with Regression
Shimon Kogan, Dimitry Levin, Bryan R. Routledge, Jacob S. Sagi, Noah A. Smith |
HLT-NAACL | 5 |
| 2009 | Preference Grammars: Softening Syntactic Constraints to Improve Statistical Machine Translation
Ashish Venugopal, Andreas Zollmann, Noah A. Smith, Stephan Vogel |
HLT-NAACL | 3 |
| 2009 | Predicting Response to Political Blog Posts with Topic Models
Tae Yano, William W. Cohen, Noah A. Smith |
HLT-NAACL | 3 |
| 2009 | Nonextensive Information Theoretic Kernels on Measures
André F. T. Martins, Noah A. Smith, Eric P. Xing, Pedro M. Q. Aguiar, Mário A. T. Figueiredo |
J. Mach. Learn. Res. | 2 |
| 2008 | Stacking Dependency Parsers
André F. T. Martins, Dipanjan Das 0001, Noah A. Smith, Eric P. Xing |
EMNLP | 3 |
| 2008 | Dynamic Programming Algorithms as Products of Weighted Logic Programs
Shay B. Cohen, Robert J. Simmons, Noah A. Smith |
ICLP | 3 |
| 2008 | Nonextensive entropic kernelsabstractPositive definite kernels on probability measures have been recently applied in structured data classification problems. Some of these kernels are related to classic information theoretic quantities, such as mutual information and the Jensen-Shannon divergence. Meanwhile, driven by recent advances in Tsallis statistics, nonextensive generalizations of Shannon's information theory have been proposed. This paper bridges these two trends. We introduce the Jensen-Tsallis q-difference, a generalization of the Jensen-Shannon divergence. We then define a new family of nonextensive mutual information kernels, which allow weights to be assigned to their arguments, and which includes the Boolean, Jensen-Shannon, and linear kernels as particular cases. We illustrate the performance of these kernels on text categorization tasks. André F. T. Martins, Mário A. T. Figueiredo, Pedro M. Q. Aguiar, Noah A. Smith, Eric P. Xing |
ICML | 4 |
| 2008 | Relative keyboard input systemabstractThis paper describes a "relative keyboard," where keystrokes are treated as inputs in a continuous space relative to each other, instead of a discrete, unambiguous sequence. A user with the ability to touch-type may type anywhere on the sensing surface without the need for a visual keyboard. An implementation of such a system is explored and evaluated on simulated data and real user data. Daniel R. Rashid, Noah A. Smith |
IUI | 2 |
| 2008 | Logistic Normal Priors for Unsupervised Probabilistic Grammar InductionabstractWe explore a new Bayesian model for probabilistic grammars, a family of distributions over discrete structures that includes hidden Markov models and probabilistic context-free grammars. Our model extends the correlated topic model framework to probabilistic grammars, exploiting the logistic normal distribution as a prior over the grammar parameters. We derive a variational EM algorithm for that model, and then experiment with the task of unsupervised grammar induction for natural language dependency parsing. We show that our model achieves superior results over previous models that use different priors. Shay B. Cohen, Kevin Gimpel, Noah A. Smith |
NIPS | 3 |
| 2008 | Computational Approaches to Morphology and Syntax Brian Roark and Richard Sproat (Oregon Health and Science University and University of Illinois at Urbana-Champaign) Oxford: Oxford University Press (Oxford surveys in syntax and morphology, edited by Robert D. Van Valin Jr, volume 4), 2007, xx+316 pp; hardbound, ISBN 978-0-19-927477-2
Noah A. Smith |
Comput. Linguistics | 1 |
| 2007 | Computationally Efficient M-Estimation of Log-Linear Structure Models
Noah A. Smith, Douglas L. Vail, John D. Lafferty |
ACL | 1 |
| 2007 | Joint Morphological and Syntactic Disambiguation
Shay B. Cohen, Noah A. Smith |
EMNLP-CoNLL | 2 |
| 2007 | Probabilistic Models of Nonprojective Dependency Trees
David A. Smith, Noah A. Smith |
EMNLP-CoNLL | 2 |
| 2007 | What is the Jeopardy Model? A Quasi-Synchronous Grammar for QA
Mengqiu Wang, Noah A. Smith, Teruko Mitamura |
EMNLP-CoNLL | 2 |
| 2007 | Weighted and Probabilistic Context-Free Grammars Are Equally ExpressiveabstractThis article studies the relationship between weighted context-free grammars (WCFGs), where each production is associated with a positive real-valued weight, and probabilistic context-free grammars (PCFGs), where the weights of the productions associated with a nonterminal are constrained to sum to one. Because the class of WCFGs properly includes the PCFGs, one might expect that WCFGs can describe distributions that PCFGs cannot. However, Z. Chi (1999, Computational Linguistics, 25(1):131–160) and S. P. Abney, D. A. McAllester, and P. Pereira (1999, In Proceedings of the 37th Annual Meeting of the Association for Computational Linguistics, pages 542–549, College Park, MD) proved that every WCFG distribution is equivalent to some PCFG distribution. We extend their results to conditional distributions, and show that every WCFG conditional distribution of parses given strings is also the conditional distribution defined by some PCFG, even when the WCFG's partition function diverges. This shows that any parsing or labeling accuracy improvement from conditional estimation of WCFGs or conditional random fields (CRFs) over joint estimation of PCFGs or hidden Markov models (HMMs) is due to the estimation procedure rather than the change in model class, because PCFGs and HMMs are exactly as expressive as WCFGs and chain-structured CRFs, respectively. Noah A. Smith, Mark Johnson 0001 |
Comput. Linguistics | 1 |
| 2006 | Annealing Structural Bias in Multilingual Weighted Grammar InductionabstractWe first show how a structural locality bias can improve the accuracy of state-of-the-art dependency grammar induction models trained by EM from unannotated examples (Klein and Manning, 2004). Next, by annealing the free parameter that controls this bias, we achieve further improvements. We then describe an alternative kind of structural bias, toward "broken" hypotheses consisting of partial structures over segmented sentences, and show a similar pattern of improvement. We relate this approach to contrastive estimation (Smith and Eisner, 2005a), apply the latter to grammar induction in six languages, and show that our new approach improves accuracy by 1-17% (absolute) over CE (and 8-30% over EM), achieving to our knowledge the best results on this task to date. Our method, structural annealing, is a general technique with broad applicability to hidden-structure discovery problems. Noah A. Smith, Jason Eisner |
ACL | 1 |
| 2006 | Vine Parsing and Minimum Risk Reranking for Speed and Precision
Markus Dreyer, David A. Smith, Noah A. Smith |
CoNLL | 3 |
| 2005 | Contrastive Estimation: Training Log-Linear Models on Unlabeled DataabstractConditional random fields (Lafferty et al., 2001) are quite effective at sequence labeling tasks like shallow parsing (Sha and Pereira, 2003) and named-entity extraction (McCallum and Li, 2003). CRFs are log-linear, allowing the incorporation of arbitrary features into the model. To train on unlabeled data, we require unsupervised estimation methods for log-linear models; few exist. We describe a novel approach, contrastive estimation. We show that the new technique can be intuitively understood as exploiting implicit negative evidence and is computationally efficient. Applied to a sequence labeling problem---POS tagging given a tagging dictionary and unlabeled text---contrastive estimation outperforms EM (with the same feature set), is more robust to degradations of the dictionary, and can largely recover by modeling additional features. Noah A. Smith, Jason Eisner |
ACL | 1 |
| 2004 | Annealing Techniques For Unsupervised Statistical Language LearningabstractExploiting unannotated natural language data is hard largely because unsupervised parameter estimation is hard. We describe deterministic annealing (Rose et al., 1990) as an appealing alternative to the Expectation-Maximization algorithm (Dempster et al., 1977). Seeking to avoid search error, DA begins by globally maximizing an easy concave function and maintains a local maximum as it gradually morphs the function into the desired non-concave likelihood function. Applying DA to parsing and tagging models is shown to be straightforward; significant improvements over EM are shown on a part-of-speech tagging task. We describe a variant, skewed DA, which can incorporate a good initializer when it is available, and show significant improvements over EM on a grammar induction task. Noah A. Smith, Jason Eisner |
ACL | 1 |
| 2004 | Bilingual Parsing with Factored Estimation: Using English to Parse Korean
David A. Smith, Noah A. Smith |
EMNLP | 2 |
| 2003 | The Web as a Parallel CorpusabstractParallel corpora have become an essential resource for work in multilingual natural language processing. In this article, we report on our work using the STRAND system for mining parallel text on the World Wide Web, first reviewing the original algorithm and results and then presenting a set of significant enhancements. These enhancements include the use of supervised learning based on structural features of documents to improve classification performance, a new content-based measure of translational equivalence, and adaptation of the system to take advantage of the Internet Archive for mining parallel text from the Web on a large scale. Finally, the value of these techniques is demonstrated in the construction of a significant parallel corpus for a low-density language pair. Philip Resnik, Noah A. Smith |
Comput. Linguistics | 2 |
| 2002 | From Words to Corpora: Recognizing TranslationabstractThis paper presents a technique for discovering translationally equivalent texts. It is comprised of the application of a matching algorithm at two different levels of analysis and a well-founded similarity score. This approach can be applied to any multilingual corpus using any kind of translation lexicon; it is therefore adaptable to varying levels of multilingual resource availability. Experimental results are shown on two tasks: a search for matching thirty-word segments in a corpus where some segments are mutual translations, and classification of candidate pairs of web pages that may or may not be translations of each other. The latter results compare competitively with previous, document-structure-based approaches to the same problem. Noah A. Smith |
EMNLP | 1 |
| 2000 | Cairo: An Alignment Visualization Tool
Noah A. Smith, Michael E. Jahr |
LREC | 1 |