Alexander M. Rush

dblp:67/9012 · DBLP profile ↗
← Back
98ranked-venue papers
7as first author
42since 2021 · last 2025
0000-0002-9900-1606ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 84 · 7 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 since 2021Systems, architecture and hardware · 4 · 2 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Contextual Document Embeddings
abstract
Dense document embeddings are central to neural retrieval. The dominant paradigm is to train and construct embeddings by running encoders directly on individual documents. In this work, we argue that these embeddings, while effective, are implicitly out-of-context for targeted use cases of retrieval, and that a contextualized document embedding should take into account both the document and neighboring documents in context - analogous to contextualized word embeddings. We propose two complementary methods for contextualized document embeddings: first, an alternative contrastive learning objective that explicitly incorporates the document neighbors into the intra-batch contextual loss; second, a new contextual architecture that explicitly encodes neighbor document information into the encoded representation. Results show that both methods achieve better performance than biencoders in several settings, with differences especially pronounced out-of-domain. We achieve state-of-the-art results on the MTEB benchmark with no hard negative mining, score distillation, dataset-specific instructions, intra-GPU example-sharing, or extremely large batch sizes. Our method can be applied to improve performance on any contrastive learning dataset and any biencoder.
John X. Morris, Alexander M. Rush
ICLR2
2025 Simple Guidance Mechanisms for Discrete Diffusion Models
abstract
Diffusion models for continuous data gained widespread adoption owing to their high quality generation and control mechanisms. However, controllable diffusion on discrete data faces challenges given that continuous guidance methods do not directly apply to discrete diffusion. Here, we provide a straightforward derivation of classifier-free and classifier-based guidance for discrete diffusion, as well as a new class of diffusion models that leverage uniform noise and that are more guidable because they can continuously edit their outputs. We improve the quality of these models with a novel continuous-time variational lower bound that yields state-of-the-art performance, especially in settings involving guidance or fast generation. Empirically, we demonstrate that our guidance mechanisms combined with uniform noise diffusion improve controllable generation relative to autoregressive and diffusion baselines on several discrete data domains, including genomic sequences, small molecule design, and discretized image generation.
Yair Schiff, Subham Sekhar Sahoo, Hao Phung, Guanghan Wang, Sam Boshar, Hugo Dalla-torre, Bernardo P. de Almeida, Alexander M. Rush, Thomas Pierrot, Volodymyr Kuleshov
ICLR8
2025 Compute-Constrained Data Selection
abstract
Data selection can reduce the amount of training data needed to finetune LLMs; however, the efficacy of data selection scales directly with its compute. Motivated by the practical challenge of compute-constrained finetuning, we consider the setting in which both the cost of selecting data and training are budgeted for. We first formalize the problem of data selection with a cost-aware utility function, and model the data selection problem as trading off initial-selection cost for training gain. We run a comprehensive sweep of experiments across multiple tasks, varying compute budget by scaling finetuning tokens, model sizes, and data selection compute. Interestingly we find that many powerful data selection methods are almost never compute-optimal, and that cheaper data selection alternatives dominate both from a theoretical and empirical perspective. For compute-optimal training, we find that perplexity and gradient data selection require training-to-selection model size ratios of 5x and 10x, respectively.
Junjie Oscar Yin, Alexander M. Rush
ICLR2
2025 Commit0: Library Generation from Scratch
abstract
With the goal of benchmarking generative systems beyond expert software development ability, we introduce Commit0, a benchmark that challenges AI agents to write libraries from scratch. Agents are provided with a specification document outlining the library’s API as well as a suite of interactive unit tests, with the goal of producing an implementation of this API accordingly. The implementation is validated through running these unit tests. As a benchmark, Commit0 is designed to move beyond static one-shot code generation towards agents that must process long-form natural language specifications, adapt to multi-stage feedback, and generate code with complex dependencies. Commit0 also offers an interactive environment where models receive static analysis and execution feedback on the code they generate. Our experiments demonstrate that while current agents can pass some unit tests, none can yet fully reproduce full libraries. Results also show that interactive feedback is quite useful for models to generate code that passes more unit tests, validating the benchmarks that facilitate its use. We publicly release the benchmark, the interactive environment, and the leaderboard.
Celine Lee, Justin T. Chiu, Claire Cardie, Matthias Gallé, Alexander M. Rush
ICLR7
2025 Multi-Turn Code Generation Through Single-Step Rewards
abstract
We address the problem of code generation from multi-turn execution feedback. Existing methods either generate code without feedback or use complex, hierarchical reinforcement learning to optimize multi-turn rewards. We propose a simple yet scalable approach, $\mu$CODE, that solves multi-turn code generation using only single-step rewards. Our key insight is that code generation is a one-step recoverable MDP, where the correct code can be recovered from any intermediate code state in a single turn. $\mu$CODE iteratively trains both a generator to provide code solutions conditioned on multi-turn execution feedback and a verifier to score the newly generated code. Experimental evaluations show that our approach achieves significant improvements over state-of-the-art baselines. We provide analysis of the design choices of the reward models and policy, and show the efficacy of $\mu$CODE at utilizing the execution feedback.
Arnav Kumar Jain, Gonzalo Gonzalez-Pumariega, Wayne Chen, Alexander M. Rush, Sanjiban Choudhury
ICML4
2025 Triton-Viz: Visualizing GPU Programming in AI Courses
abstract
GPU programming is a critical component in AI system courses, which is notoriously difficult to learn and teach, given its unique features such as massive parallelism and data movement across memory hierarchies. This paper presents Triton-Viz, an innovative visualization toolkit that helps students learn GPU programming using Triton, one of the most widely used programming languages to develop AI applications. Triton-Viz offers an intuitive interface with interactive visualizations of GPU operations from multiple perspectives, including parallelism, memory access, and performance metrics. By integrating into educational materials, Triton-Viz enhances hands-on learning and improves comprehension of GPU programming concepts and AI algorithms in real applications, as demonstrated in a study conducted with Computer Science students. The positive feedback from this study highlights the utility of Triton-Viz in educational settings and its potential to bridge the gap between the theoretical algorithm and practical implementations in AI courses.
Tejas Ramesh, Alexander M. Rush, Xu Liu 0001, Binqian Yin, Keren Zhou 0001, Shuyin Jiao
SIGCSE (1)2
2025 Scaling Data-Constrained Language Models
abstract
The current trend of scaling language models involves increasing both parameter count and training data set size. Extrapolating this trend suggests that training data set size may soon be limited by the amount of text data available on the internet. Motivated by this limit, we investigate scaling language models in data-constrained regimes. Specifically, we run a large set of experiments varying the extent of data repetition and compute budget, ranging up to 900 billion training tokens and 9 billion parameter models. We find that with constrained data for a fixed compute budget, training with up to 4 epochs of repeated data yields negligible changes to loss compared to having unique data. However, with more repetition, the value of adding compute eventually decays to zero. We propose and empirically validate a scaling law for compute optimality that accounts for the decreasing value of repeated tokens and excess parameters. Finally, we experiment with approaches mitigating data scarcity, including augmenting the training data set with code data or removing commonly used filters. Models and data sets from our 400 training runs are freely available at https://github.com/huggingface/datablations.
Niklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf 0008, Colin Raffel
J. Mach. Learn. Res.2
2024 Predicting Text Preference Via Structured Comparative Reasoning
abstract
Jing Nathan Yan, Tianqi Liu, Justin Chiu, Jiaming Shen, Zhen Qin, Yue Yu, Charumathi Lakshmanan, Yair Kurzion, Alexander Rush, Jialu Liu, Michael Bendersky. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Jing Nathan Yan, Tianqi Liu 0002, Justin T. Chiu, Zhen Qin 0001, Yue Yu 0001, Charumathi Lakshmanan, Yair Kurzion, Alexander M. Rush, Michael Bendersky
ACL (1)9
2024 Diffusion Models Without Attention
abstract
In recent advancements in high-fidelity image generation, Denoising Diffusion Probabilistic Models (DDPMs) have emerged as a key player. However, their application at high resolutions presents significant computational challenges. Current methods, such as patchifying, expedite processes in UNet and Transformer architectures but at the expense of rep-resentational capacity. Addressing this, we introduce the Dif-fusion State Space Model (DIFFUSSM), an architecture that supplants attention mechanisms with a more scalable state space model backbone. This approach effectively handles higher resolutions without resorting to global compression, thus preserving detailed image representation throughout the diffusion process. Our focus on FLOP-efficient architectures in diffusion training marks a significant step forward. Comprehensive evaluations on both ImageNet and LSUN datasets at two resolutions demonstrate that DiffuSSMs are on par or even outperform existing diffusion models with attention modules in FID and Inception Score metrics while significantly reducing total FLOP usage.
Jing Nathan Yan, Jiatao Gu, Alexander M. Rush
CVPR3
2024 ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
abstract
Yash Akhauri, Ahmed F. AbouElhamayed, Jordan Dotzel, Zhiru Zhang, Alexander M. Rush, Safeen Huda, Mohamed S. Abdelfattah. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yash Akhauri, Ahmed F. AbouElhamayed, Jordan Dotzel, Zhiru Zhang, Alexander M. Rush, Safeen Huda, Mohamed S. Abdelfattah
EMNLP5
2024 I Could've Asked That: Reformulating Unanswerable Questions
abstract
When seeking information from unfamiliar documents, users frequently pose questions that cannot be answered by the documents.While existing large language models (LLMs) identify these unanswerable questions, they do not assist users in reformulating their questions, thereby reducing their overall utility.We curate COULDASK, an evaluation benchmark composed of existing and new datasets for document-grounded question answering, specifically designed to study reformulating unanswerable questions.We evaluate stateof-the-art open-source and proprietary LLMs on COULDASK.The results demonstrate the limited capabilities of these models in reformulating questions.Specifically, GPT-4 and Llama2-7B successfully reformulate questions only 26% and 12% of the time, respectively.Error analysis shows that 62% of the unsuccessful reformulations stem from the models merely rephrasing the questions or even generating identical questions.We publicly release the benchmark 1 and the code to reproduce the experiments 2 .
Claire Cardie, Alexander M. Rush
EMNLP4
2024 Guess & Sketch: Language Model Guided Transpilation
abstract
Maintaining legacy software requires many software and systems engineering hours. Assembly code programs, which demand low-level control over the computer machine state and have no variable names, are particularly difficult for humans to analyze. Existing conventional program translators guarantee correctness, but are hand-engineered for the source and target programming languages in question. Learned transpilation, i.e. automatic translation of code, offers an alternative to manual re-writing and engineering efforts. Automated symbolic program translation approaches guarantee correctness but struggle to scale to longer programs due to the exponentially large search space. Their rigid rule-based systems also limit their expressivity, so they can only reason about a reduced space of programs. Probabilistic neural language models (LMs) produce plausible outputs for every input, but do so at the cost of guaranteed correctness. In this work, we leverage the strengths of LMs and symbolic solvers in a neurosymbolic approach to learned transpilation for assembly code. Assembly code is an appropriate setting for a neurosymbolic approach, since assembly code can be divided into shorter non-branching basic blocks amenable to the use of symbolic methods. Guess & Sketch extracts alignment and confidence information from features of the LM then passes it to a symbolic solver to resolve semantic equivalence of the transpilation input and output. We test Guess & Sketch on three different test sets of assembly transpilation tasks, varying in difficulty, and show that it successfully transpiles 57.6% more examples than GPT-4 and 39.6% more examples than an engineered transpiler. We also share a training and evaluation dataset for this task.
Celine Lee, Abdulrahman Mahmoud, Michal Kurek, Simone Campanoni, David Brooks 0001, Stephen Chong, Gu-Yeon Wei, Alexander M. Rush
ICLR8
2024 Language Model Inversion
abstract
Given a prompt, language models produce a distribution over all possible next tokens; when the prompt is unknown, can we use this distributional information to recover the prompt? We consider the problem of anguage model inversion and show that next-token probabilities contain a surprising amount of information about the preceding text. Often we can recover the text in cases where it is hidden from the user, motivating a method for recovering unknown prompts given only the model's current distribution output. We consider a variety of model access scenarios, and show how even without predictions for every token in the vocabulary we can recover the probability vector through search and reconstruction of the input. On LLAMA-7B, our inversion method reconstructs prompts with a BLEU of $59$ and token-level F1 of $77$ and recovers $23\%$ of prompts exactly
John X. Morris, Justin T. Chiu, Vitaly Shmatikov, Alexander M. Rush
ICLR5
2024 Entity Disambiguation via Fusion Entity Decoding
abstract
Junxiong Wang, Ali Mousavi, Omar Attia, Ronak Pradeep, Saloni Potdar, Alexander Rush, Umar Farooq Minhas, Yunyao Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Junxiong Wang, Ali Mousavi 0003, Omar Attia, Ronak Pradeep, Saloni Potdar, Alexander M. Rush, Umar Farooq Minhas, Yunyao Li 0001
NAACL-HLT6
2024 The Mamba in the Llama: Distilling and Accelerating Hybrid Models
abstract
Linear RNN architectures, like Mamba, can be competitive with Transformer models in language modeling while having advantageous deployment characteristics. Given the focus on training large-scale Transformer models, we consider the challenge of converting these pretrained models for deployment. We demonstrate that it is feasible to distill large Transformers into linear RNNs by reusing the linear projection weights from attention layers with academic GPU resources. The resulting hybrid model, which incorporates a quarter of the attention layers, achieves performance comparable to the original Transformer in chat benchmarks and outperforms open-source hybrid Mamba models trained from scratch with trillions of tokens in both chat benchmarks and general benchmarks. Moreover, we introduce a hardware-aware speculative decoding algorithm that accelerates the inference speed of Mamba and hybrid models. Overall we show how, with limited computation resources, we can remove many of the original attention layers and generate from the resulting model more efficiently. Our top-performing model, distilled from Llama3-8B-Instruct, achieves a 29.61 length-controlled win rate on AlpacaEval 2 against GPT-4 and 7.35 on MT-Bench, surpassing the best 8B scale instruction-tuned linear RNN model. We also find that the distilled model has natural length extrapolation, showing almost perfect accuracy in the needle-in-a-haystack test at 20x the distillation length. Code and pre-trained checkpoints are open-sourced at [MambaInLlama](https://github.com/jxiw/MambaInLlama) for distillation and [SpeculativeMamba](https://github.com/itsdaniele/speculative\_mamba) for speculative decoding.
Junxiong Wang, Daniele Paliotta, Avner May, Alexander M. Rush, Tri Dao
NeurIPS4
2023 Abductive Commonsense Reasoning Exploiting Mutually Exclusive Explanations
abstract
Abductive reasoning aims to find plausible explanations for an event.This style of reasoning is critical for commonsense tasks where there are often multiple plausible explanations.Existing approaches for abductive reasoning in natural language processing (NLP) often rely on manually generated annotations for supervision; however, such annotations can be subjective and biased.Instead of using direct supervision, this work proposes an approach for abductive commonsense reasoning that exploits the fact that only a subset of explanations is correct for a given context.The method uses posterior regularization to enforce a mutual exclusion constraint, encouraging the model to learn the distinction between fluent explanations and plausible ones.We evaluate our approach on a diverse set of abductive reasoning datasets; experimental results show that our approach outperforms or is comparable to directly applying pretrained language models in a zeroshot manner and other knowledge-augmented zero-shot methods.
Justin T. Chiu, Claire Cardie, Alexander M. Rush
ACL (1)4
2023 Symbolic Planning and Code Generation for Grounded Dialogue
abstract
Large language models (LLMs) excel at processing and generating both text and code.However, LLMs have had limited applicability in grounded task-oriented dialogue as they are difficult to steer toward task objectives and fail to handle novel grounding.We present a modular and interpretable grounded dialogue system that addresses these shortcomings by composing LLMs with a symbolic planner and grounded code execution.Our system consists of a reader and planner: the reader leverages an LLM to convert partner utterances into executable code, calling functions that perform grounding.The translated code's output is stored to track dialogue state, while a symbolic planner determines the next appropriate response.We evaluate our system's performance on the demanding ONECOMMON dialogue task, involving collaborative reference resolution on abstract images of scattered dots.Our system substantially outperforms the previous state-of-the-art, including improving task success in human evaluations from 56% to 69% in the most challenging setting.
Justin T. Chiu, Derek Chen, Saujas Vaduguru, Alexander M. Rush, Daniel Fried
EMNLP5
2023 Text Embeddings Reveal (Almost) As Much As Text
abstract
How much private information do text embeddings reveal about the original text?We investigate the problem of embedding inversion, reconstructing the full text represented in dense text embeddings.We frame the problem as controlled generation: generating text that, when reembedded, is close to a fixed point in latent space.We find that although a naïve model conditioned on the embedding performs poorly, a multi-step method that iteratively corrects and re-embeds text is able to recover 92% of 32-token text inputs exactly.We train our model to decode text embeddings from two state-of-the-art embedding models, and also show that our model can recover important personal information (full names) from a dataset of clinical notes. 1
John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. Rush
EMNLP4
2023 Tree Prompting: Efficient Task Adaptation without Fine-Tuning
abstract
Prompting language models (LMs) is the main interface for applying them to new tasks.However, for smaller LMs, prompting provides low accuracy compared to gradient-based finetuning.Tree Prompting is an approach to prompting which builds a decision tree of prompts, linking multiple LM calls together to solve a task.At inference time, each call to the LM is determined by efficiently routing the outcome of the previous call using the tree.Experiments on classification datasets show that Tree Prompting improves accuracy over competing methods and is competitive with fine-tuning.We also show that variants of Tree Prompting allow inspection of a model's decision-making process. 1
Chandan Singh, John X. Morris, Alexander M. Rush, Jianfeng Gao 0001, Yuntian Deng
EMNLP3
2023 Hop, Union, Generate: Explainable Multi-hop Reasoning without Rationale Supervision
abstract
Explainable multi-hop question answering (QA) not only predicts answers but also identifies rationales, i. e. subsets of input sentences used to derive the answers.This problem has been extensively studied under the supervised setting, where both answer and rationale annotations are given.Because rationale annotations are expensive to collect and not always available, recent efforts have been devoted to developing methods that do not rely on supervision for rationales.However, such methods have limited capacities in modeling interactions between sentences, let alone reasoning across multiple documents.This work proposes a principled, probabilistic approach for training explainable multi-hop QA systems without rationale supervision.Our approach performs multi-hop reasoning by explicitly modeling rationales as sets, enabling the model to capture interactions between documents and sentences within a document.Experimental results show that our approach is more accurate at selecting rationales than the previous methods, while maintaining similar accuracy in predicting answers.
Justin T. Chiu, Claire Cardie, Alexander M. Rush
EMNLP4
2023 Markup-to-Image Diffusion Models with Scheduled Sampling
Yuntian Deng, Noriyuki Kojima, Alexander M. Rush
ICLR3
2023 OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
abstract
Large multimodal models trained on natural documents, which interleave images and text, outperform models trained on image-text pairs on various multimodal benchmarks. However, the datasets used to train these models have not been released, and the collection process has not been fully specified. We introduce the OBELICS dataset, an open web-scale filtered dataset of interleaved image-text documents comprising 141 million web pages extracted from Common Crawl, 353 million associated images, and 115 billion text tokens. We describe the dataset creation process, present comprehensive filtering rules, and provide an analysis of the dataset's content. To show the viability of OBELICS, we train on the dataset vision and language models of 9 and 80 billion parameters, IDEFICS-9B and IDEFICS, and obtain competitive performance on different multimodal benchmarks. We release our dataset, models and code.
Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander M. Rush, Douwe Kiela, Matthieu Cord, Victor Sanh
NeurIPS9
2023 Scaling Data-Constrained Language Models
abstract
The current trend of scaling language models involves increasing both parameter count and training dataset size. Extrapolating this trend suggests that training dataset size may soon be limited by the amount of text data available on the internet. Motivated by this limit, we investigate scaling language models in data-constrained regimes. Specifically, we run a large set of experiments varying the extent of data repetition and compute budget, ranging up to 900 billion training tokens and 9 billion parameter models. We find that with constrained data for a fixed compute budget, training with up to 4 epochs of repeated data yields negligible changes to loss compared to having unique data. However, with more repetition, the value of adding compute eventually decays to zero. We propose and empirically validate a scaling law for compute optimality that accounts for the decreasing value of repeated tokens and excess parameters. Finally, we experiment with approaches mitigating data scarcity, including augmenting the training dataset with code data or removing commonly used filters. Models and datasets from our 400 training runs are freely available at https://github.com/huggingface/datablations.
Niklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf 0008, Colin Raffel
NeurIPS2
2023 Teal: Learning-Accelerated Optimization of WAN Traffic Engineering
abstract
The rapid expansion of global cloud wide-area networks (WANs) has posed a challenge for commercial optimization engines to efficiently solve network traffic engineering (TE) problems at scale. Existing acceleration strategies decompose TE optimization into concurrent subproblems but realize limited parallelism due to an inherent tradeoff between run time and allocation performance.
Zhiying Xu, Francis Y. Yan, Rachee Singh, Justin T. Chiu, Alexander M. Rush, Minlan Yu
SIGCOMM5
2023 End-to-end learning of multiple sequence alignments with differentiable Smith-Waterman
abstract
MOTIVATION: Multiple sequence alignments (MSAs) of homologous sequences contain information on structural and functional constraints and their evolutionary histories. Despite their importance for many downstream tasks, such as structure prediction, MSA generation is often treated as a separate pre-processing step, without any guidance from the application it will be used for. RESULTS: Here, we implement a smooth and differentiable version of the Smith-Waterman pairwise alignment algorithm that enables jointly learning an MSA and a downstream machine learning system in an end-to-end fashion. To demonstrate its utility, we introduce SMURF (Smooth Markov Unaligned Random Field), a new method that jointly learns an alignment and the parameters of a Markov Random Field for unsupervised contact prediction. We find that SMURF learns MSAs that mildly improve contact prediction on a diverse set of protein and RNA families. As a proof of concept, we demonstrate that by connecting our differentiable alignment module to AlphaFold2 and maximizing predicted confidence, we can learn MSAs that improve structure predictions over the initial MSAs. Interestingly, the alignments that improve AlphaFold predictions are self-inconsistent and can be viewed as adversarial. This work highlights the potential of differentiable dynamic programming to improve neural network pipelines that rely on an alignment and the potential dangers of optimizing predictions of protein sequences with methods that are not fully understood. AVAILABILITY AND IMPLEMENTATION: Our code and examples are available at: https://github.com/spetti/SMURF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Samantha Petti, Nicholas Bhattacharya, Roshan Rao, Justas Dauparas, Juannan Zhou, Alexander M. Rush, Peter K. Koo, Sergey Ovchinnikov 0001
Bioinform.7
2023 Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language Models
abstract
State-of-the-art neural language models can now be used to solve ad-hoc language tasks through zero-shot prompting without the need for supervised training. This approach has gained popularity in recent years, and researchers have demonstrated prompts that achieve strong accuracy on specific NLP tasks. However, finding a prompt for new tasks requires experimentation. Different prompt templates with different wording choices lead to significant accuracy differences. PromptIDE allows users to experiment with prompt variations, visualize prompt performance, and iteratively optimize prompts. We developed a workflow that allows users to first focus on model feedback using small data before moving on to a large data regime that allows empirical grounding of promising prompts using quantitative measures of the task. The tool then allows easy deployment of the newly created ad-hoc models. We demonstrate the utility of PromptIDE (demo: http://prompt.vizhub.ai) and our workflow using several real-world use cases.
Hendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover, Johanna Beyer, Hanspeter Pfister, Alexander M. Rush
IEEE Trans. Vis. Comput. Graph.7
2022 Xatu: boosting existing DDoS detection systems using auxiliary signals
abstract
Traditional DDoS attack detection monitors volumetric traffic features to detect attack onset. To reduce false positives, such detection is often conservative---raising an alert only after a sustained period of observed anomalous behavior. However, contemporary attacks tend to be short, which combined with a long detection delay means that most of the attack still reaches and impacts the victim. We propose Xatu, a system that utilizes auxiliary signals to improve the accuracy and timeliness of existing DDoS detection systems. We explore two types of auxiliary signals, attack preparation signals and the history of prior attacks. These signals can be easily mined from existing traffic monitoring systems in many ISP networks. To leverage these auxiliary signals for attack detection, we propose a multi-timescale LSTM model, which derives both long-term and short-term patterns from diverse auxiliary signals. We then leverage survival analysis to quickly detect attacks when they occur while minimizing false positives and thus scrubbing costs. We evaluate Xatu on traffic from a large ISP, using commercial defense alert data to label prevalent attack events. Xatu would help the commercial defense scrub up to 44.1% additional anomalous traffic and would reduce its median detection delay by 9.5 minutes.1
Zhiying Xu, Sivaramakrishnan Ramanathan, Alexander M. Rush, Jelena Mirkovic, Minlan Yu
CoNEXT3
2022 Model Criticism for Long-Form Text Generation
abstract
Language models have demonstrated the ability to generate highly fluent text; however, it remains unclear whether their output retains coherent high-level structure (e.g., story progression).Here, we propose to apply a statistical tool, model criticism in latent space, to evaluate the high-level structure of the generated text.Model criticism compares the distributions between real and generated data in a latent space obtained according to an assumptive generative process.Different generative processes identify specific failure modes of the underlying model.We perform experiments on three representative aspects of highlevel discourse-coherence, coreference, and topicality-and find that transformer-based language models are able to capture topical structures but have a harder time maintaining structural coherence or modeling coreference.
Yuntian Deng, Volodymyr Kuleshov, Alexander M. Rush
EMNLP3
2022 Multitask Prompted Training Enables Zero-Shot Task Generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim 0002, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Mike Tian-Jian Jiang, Matteo Manica, Sheng Shen 0001, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Févry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf 0008, Alexander M. Rush
ICLR40
2022 GenNI: Human-AI Collaboration for Data-Backed Text Generation
abstract
Table2Text systems generate textual output based on structured data utilizing machine learning. These systems are essential for fluent natural language interfaces in tools such as virtual assistants; however, left to generate freely these ML systems often produce misleading or unexpected outputs. GenNI (Generation Negotiation Interface) is an interactive visual system for high-level human-AI collaboration in producing descriptive text. The tool utilizes a deep learning model designed with explicit control states. These controls allow users to globally constrain model generations, without sacrificing the representation power of the deep learning models. The visual interface makes it possible for users to interact with AI systems following a Refine-Forecast paradigm to ensure that the generation system acts in a manner human users find suitable. We report multiple use cases on two experiments that improve over uncontrolled generation approaches, while at the same time providing fine-grained control. A demo and source code are available at https://genni.vizhub.ai.
Hendrik Strobelt, Jambay Kinley, Robert Krüger, Johanna Beyer, Hanspeter Pfister, Alexander M. Rush
IEEE Trans. Vis. Comput. Graph.6
2021 Parameter-Efficient Transfer Learning with Diff Pruning
abstract
Demi Guo, Alexander Rush, Yoon Kim. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Demi Guo, Alexander M. Rush
ACL/IJCNLP (1)2
2021 GRIT: Generative Role-filler Transformers for Document-level Event Entity Extraction
abstract
We revisit the classic problem of documentlevel role-filler entity extraction (REE) for template filling.We argue that sentence-level approaches are ill-suited to the task and introduce a generative transformer-based encoderdecoder framework (GRIT) that is designed to model context at the document level: it can make extraction decisions across sentence boundaries; is implicitly aware of noun phrase coreference structure, and has the capacity to respect cross-role dependencies in the template structure.We evaluate our approach on the MUC-4 dataset, and show that our model performs substantially better than prior work.We also show that our modeling choices contribute to model performance, e.g., by implicitly capturing linguistic knowledge such as recognizing coreferent entity mentions.
Xinya Du, Alexander M. Rush, Claire Cardie
EACL2
2021 Block Pruning For Faster Transformers
abstract
Pre-training has improved model accuracy for both classification and generation tasks at the cost of introducing much larger and slower models.Pruning methods have proven to be an effective way of reducing model size, whereas distillation methods are proven for speeding up inference.We introduce a block pruning approach targeting both small and fast models.Our approach extends structured methods by considering blocks of any size and integrates this structure into the movement pruning paradigm for fine-tuning.We find that this approach learns to prune out full components of the underlying model, such as attention heads.Experiments consider classification and generation tasks, yielding among other results a pruned model that is a 2.4x faster, 74% smaller BERT on SQuAD v1, with a 1% drop on F1, competitive both with distilled models in speed and pruned models in size.
François Lagunas, Ella Charlaix, Victor Sanh, Alexander M. Rush
EMNLP (1)4
2021 Rationales for Sequential Predictions
abstract
Sequence models are a critical component of modern NLP systems, but their predictions are difficult to explain.We consider model explanations though rationales, subsets of context that can explain individual model predictions.We find sequential rationales by solving a combinatorial optimization: the best rationale is the smallest subset of input tokens that would predict the same output as the full sequence.Enumerating all subsets is intractable, so we propose an efficient greedy algorithm to approximate this objective.The algorithm, which is called greedy rationalization, applies to any model.For this approach to be effective, the model should form compatible conditional distributions when making predictions on incomplete subsets of the context.This condition can be enforced with a short finetuning step.We study greedy rationalization on language modeling and machine translation.Compared to existing baselines, greedy rationalization is best at optimizing the sequential objective and provides the most faithful rationales.On a new dataset of annotated sequential rationales, greedy rationales are most similar to human rationales.
Keyon Vafa, Yuntian Deng, David M. Blei, Alexander M. Rush
EMNLP (1)4
2021 SM6: A 16nm System-on-Chip for Accurate and Noise-Robust Attention-Based NLP Applications : The 33rd Hot Chips Symposium - August 22-24, 2021
abstract
In this work, we present SM6, an SoC architecture for real-time denoised speech and NLP pipelines, featuring (1) MSSE: an unsupervised probabilistic sound source separation accelerator, (2) FlexNLP: a programmable inference accelerator for attention-based seq2seq DNNs using adaptive floating-point datatypes for wide dynamic range computations, (3) a dual-core Arm Cortex A53 CPU cluster, which provides on-demand SIMD FFT processing, and operating system support. In adverse acoustic conditions, MSSE allows FlexNLP to store up to 6x smaller ASR models obviating the very inefficient strategy of scaling up the DNN model to achieve noise robustness. MSSE and FlexNLP produce efficiency ranges of 4.33-17.6 Gsamples/s/W and 2.6-7.8TFLOPs/W, respectively, with per-frame end-to-end latencies of 15-45ms.
Thierry Tambe, En-Yu Yang, Glenn G. Ko, Yuji Chai, Coleman Hooper, Marco Donato, Paul N. Whatmough, Alexander M. Rush, David Brooks 0001, Gu-Yeon Wei
HCS8
2021 Learning from others' mistakes: Avoiding dataset biases without modeling them
Victor Sanh, Thomas Wolf 0008, Yonatan Belinkov, Alexander M. Rush
ICLR4
2021 Developmental Stage Classification of Embryos Using Two-Stream Neural Network with Linear-Chain Conditional Random Field
Stanislav Lukyanenko, Won-Dong Jang, Donglai Wei 0001, Robbert Struyven, Brian D. Leahy, Helen Y. Yang, Alexander M. Rush, Dalit Ben-Yosef, Daniel Needleman, Hanspeter Pfister
MICCAI (8)8
2021 EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference
abstract
Transformer-based language models such as BERT provide significant accuracy improvement to a multitude of natural language processing (NLP) tasks. However, their hefty computational and memory demands make them challenging to deploy to resource-constrained edge platforms with strict latency requirements.
Thierry Tambe, Coleman Hooper, Lillian Pentecost, En-Yu Yang, Marco Donato, Victor Sanh, Paul N. Whatmough, Alexander M. Rush, David Brooks 0001, Gu-Yeon Wei
MICRO9
2021 Template Filling with Generative Transformers
abstract
Template filling is generally tackled by a pipeline of two separate supervised systemsone for role-filler extraction and another for template/event recognition.Since pipelines consider events in isolation, they can suffer from error propagation.We introduce a framework based on end-to-end generative transformers for this task (i.e., GTT).It naturally models the dependence between entities both within a single event and across the multiple events described in a document.Experiments demonstrate that this framework substantially outperforms pipeline-based approaches, and other neural end-to-end baselines that do not model between-event dependencies.We further show that our framework specifically improves performance on documents containing multiple events.
Xinya Du, Alexander M. Rush, Claire Cardie
NAACL-HLT2
2021 Low-Complexity Probing via Finding Subnetworks
abstract
The dominant approach in probing neural networks for linguistic properties is to train a new shallow multi-layer perceptron (MLP) on top of the model's internal representations.This approach can detect properties encoded in the model, but at the cost of adding new parameters that may learn the task directly.We instead propose a subtractive pruning-based probe, where we find an existing subnetwork that performs the linguistic task of interest.Compared to an MLP, the subnetwork probe achieves both higher accuracy on pre-trained models and lower accuracy on random models, so it is both better at finding properties of interest and worse at learning on its own.Next, by varying the complexity of each probe, we show that subnetwork probing Pareto-dominates MLP probing in that it achieves higher accuracy given any budget of probe complexity.Finally, we analyze the resulting subnetworks across various tasks to locate where each task is encoded, and we find that lower-level tasks are captured in lower layers, reproducing similar findings in past work.
Victor Sanh, Alexander M. Rush
NAACL-HLT2
2021 How many data points is a prompt worth?
abstract
When fine-tuning pretrained models for classification, researchers either use a generic model head or a task-specific prompt for prediction.Proponents of prompting have argued that prompts provide a method for injecting taskspecific guidance, which is beneficial in lowdata regimes.We aim to quantify this benefit through rigorous testing of prompts in a fair setting: comparing prompted and head-based fine-tuning in equal conditions across many tasks and data sizes.By controlling for many sources of advantage, we find that prompting does indeed provide a benefit, and that this benefit can be quantified per task.Results show that prompting is often worth 100s of data points on average across classification tasks.
Teven Le Scao, Alexander M. Rush
NAACL-HLT2
2021 Low-Rank Constraints for Fast Inference in Structured Models
abstract
Structured distributions, i.e. distributions over combinatorial spaces, are commonly used to learn latent probabilistic representations from observed data. However, scaling these models is bottlenecked by the high computational and memory complexity with respect to the size of the latent representations. Common models such as Hidden Markov Models (HMMs) and Probabilistic Context-Free Grammars (PCFGs) require time and space quadratic and cubic in the number of hidden states respectively. This work demonstrates a simple approach to reduce the computational and memory complexity of a large class of structured models. We show that by viewing the central inference step as a matrix-vector product and using a low-rank constraint, we can trade off model expressivity and speed via the rank. Experiments with neural parameterized structured models for language modeling, polyphonic music modeling, unsupervised grammar induction, and video modeling show that our approach matches the accuracy of standard models at large state spaces while providing practical speedups.
Justin T. Chiu, Yuntian Deng, Alexander M. Rush
NeurIPS3
2020 What is Learned in Visually Grounded Neural Syntax Acquisition
abstract
Visual features are a promising signal for learning bootstrap textual models.However, blackbox learning models make it difficult to isolate the specific contribution of visual components.In this analysis, we consider the case study of the Visually Grounded Neural Syntax Learner (Shi et al., 2019), a recent approach for learning syntax from a visual training signal.By constructing simplified versions of the model, we isolate the core factors that yield the model's strong performance.Contrary to what the model might be capable of learning, we find significantly less expressive versions produce similar predictions and perform just as well, or even better.We also find that a simple lexical signal of noun concreteness plays the main role in the model's predictions as opposed to more complex syntactic reasoning.
Noriyuki Kojima, Hadar Averbuch-Elor, Alexander M. Rush, Yoav Artzi
ACL3
2020 Posterior Control of Blackbox Generation
abstract
Text generation often requires high-precision output that obeys task-specific rules.This fine-grained control is difficult to enforce with off-the-shelf deep learning models.In this work, we consider augmenting neural generation models with discrete control states learned through a structured latent-variable approach.Under this formulation, task-specific knowledge can be encoded through a range of rich, posterior constraints that are effectively trained into the model.This approach allows users to ground internal model decisions based on prior knowledge, without sacrificing the representational power of neural generative models.Experiments consider applications of this approach for text generation.We find that this method improves over standard benchmarks, while also providing fine-grained control.
Xiang Li 0063, Alexander M. Rush
ACL2
2020 Algorithm-Hardware Co-Design of Adaptive Floating-Point Encodings for Resilient Deep Learning Inference
abstract
Conventional hardware-friendly quantization methods, such as fixed-point or integer, tend to perform poorly at very low precision as their shrunken dynamic ranges cannot adequately capture the wide data distributions commonly seen in sequence transduction models. We present an algorithm-hardware co-design centered around a novel floating-point inspired number format, AdaptivFloat, that dynamically maximizes and optimally clips its available dynamic range, at a layer granularity, in order to create faithful encodings of neural network parameters. AdaptivFloat consistently produces higher inference accuracies compared to block floating-point, uniform, IEEE-like float or posit encodings at low bit precision (≤8-bit) across a diverse set of state-of-the-art neural networks, exhibiting narrow to wide weight distribution. Notably, at 4-bit weight precision, only a 2.1 degradation in BLEU score is observed on the AdaptivFloat-quantized Transformer network compared to total accuracy loss when encoded in the above-mentioned prominent datatypes. Furthermore, experimental results on a deep neural network (DNN) processing element (PE), exploiting AdaptivFloat logic in its computational datapath, demonstrate per-operation energy and area that is 0.9× and 1.14×, width, respectively that of an equivalent bit NVDLA-like integer-based PE.
Thierry Tambe, En-Yu Yang, Zishen Wan, Yuntian Deng, Vijay Janapa Reddi, Alexander M. Rush, David Brooks 0001, Gu-Yeon Wei
DAC6
2020 Scaling Hidden Markov Language Models
abstract
The hidden Markov model (HMM) is a fundamental tool for sequence modeling that cleanly separates the hidden state from the emission structure.However, this separation makes it difficult to fit HMMs to large datasets in modern NLP, and they have fallen out of use due to very poor performance compared to fully observed models.This work revisits the challenge of scaling HMMs to language modeling datasets, taking ideas from recent approaches to neural modeling.We propose methods for scaling HMMs to massive state spaces while maintaining efficient exact inference, a compact parameterization, and effective regularization.Experiments show that this approach leads to models that are more accurate than previous HMM and n-gram-based methods, making progress towards the performance of state-of-the-art neural models.
Justin T. Chiu, Alexander M. Rush
EMNLP (1)2
2020 Sequence-Level Mixed Sample Data Augmentation
abstract
Despite their empirical success, neural networks still have difficulty capturing compositional aspects of natural language.This work proposes a simple data augmentation approach to encourage compositional behavior in neural models for sequence-to-sequence problems.Our approach, SeqMix, creates new synthetic examples by softly combining input/output sequences from the training set.We connect this approach to existing techniques such as SwitchOut (Wang et al., 2018) and word dropout (Sennrich et al., 2016), and show that these techniques are all approximating variants of a single objective.SeqMix consistently yields approximately 1.0 BLEU improvement on five different translation datasets over strong Transformer baselines.On tasks that require strong compositional generalization such as SCAN and semantic parsing, Se-qMix also offers further improvements.
Demi Guo, Alexander M. Rush
EMNLP (1)3
2020 Adversarial Semantic Collisions
abstract
We study semantic collisions: texts that are semantically unrelated but judged as similar by NLP models.We develop gradient-based approaches for generating semantic collisions and demonstrate that state-of-the-art models for many tasks which rely on analyzing the meaning and similarity of texts-including paraphrase identification, document retrieval, response suggestion, and extractive summarization-are vulnerable to semantic collisions.For example, given a target query, inserting a crafted collision into an irrelevant document can shift its retrieval rank from 1000 to top 3.We show how to generate semantic collisions that evade perplexity-based filtering and discuss other potential mitigations.Our code is available at https://github.com/ csong27/collision-bert.
Congzheng Song, Alexander M. Rush, Vitaly Shmatikov
EMNLP (1)2
2020 Cascaded Text Generation with Markov Transformers
abstract
The two dominant approaches to neural text generation are fully autoregressive models, using serial beam search decoding, and non-autoregressive models, using parallel decoding with no output dependencies. This work proposes an autoregressive model with sub-linear parallel time generation. Noting that conditional random fields with bounded context can be decoded in parallel, we propose an efficient cascaded decoding approach for generating high-quality output. To parameterize this cascade, we introduce a Markov transformer, a variant of the popular fully autoregressive model that allows us to simultaneously decode with specific autoregressive context cutoffs. This approach requires only a small modification from standard autoregressive training, while showing competitive accuracy/speed tradeoff compared to existing methods on five machine translation datasets.
Yuntian Deng, Alexander M. Rush
NeurIPS2
2020 Latent Template Induction with Gumbel-CRFs
abstract
Learning to control the structure of sentences is a challenging problem in text generation. Existing work either relies on simple deterministic approaches or RL-based hard structures. We explore the use of structured variational autoencoders to infer latent templates for sentence generation using a soft, continuous relaxation in order to utilize reparameterization for training. Specifically, we propose a Gumbel-CRF, a continuous relaxation of the CRF sampling algorithm using a relaxed Forward-Filtering Backward-Sampling (FFBS) approach. As a reparameterized gradient estimator, the Gumbel-CRF gives more stable gradients than score-function based estimators. As a structured inference network, we show that it learns interpretable templates during training, which allows us to control the decoder during testing. We demonstrate the effectiveness of our methods with experiments on data-to-text generation and unsupervised paraphrase generation.
Chuanqi Tan, Bin Bi, Mosha Chen, Yansong Feng 0002, Alexander M. Rush
NeurIPS6
2020 Movement Pruning: Adaptive Sparsity by Fine-Tuning
abstract
Magnitude pruning is a widely used strategy for reducing model size in pure supervised learning; however, it is less effective in the transfer learning regime that has become standard for state-of-the-art natural language processing applications. We propose the use of movement pruning, a simple, deterministic first-order weight pruning method that is more adaptive to pretrained model fine-tuning. We give mathematical foundations to the method and compare it to existing zeroth- and first-order pruning methods. Experiments show that when pruning large pretrained language models, movement pruning shows significant improvements in high-sparsity regimes. When combined with distillation, the approach achieves minimal accuracy loss with down to only 3% of the model parameters.
Victor Sanh, Thomas Wolf 0008, Alexander M. Rush
NeurIPS3
2020 Visual Interaction with Deep Learning Models through Collaborative Semantic Inference
abstract
Automation of tasks can have critical consequences when humans lose agency over decision processes. Deep learning models are particularly susceptible since current black-box approaches lack explainable reasoning. We argue that both the visual interface and model structure of deep learning systems need to take into account interaction design. We propose a framework of collaborative semantic inference (CSI) for the co-design of interactions and models to enable visual collaboration between humans and algorithms. The approach exposes the intermediate reasoning process of models which allows semantic interactions with the visual metaphors of a problem, which means that a user can both understand and control parts of the model reasoning process. We demonstrate the feasibility of CSI with a co-designed case study of a document summarization system.
Sebastian Gehrmann, Hendrik Strobelt, Robert Krüger, Hanspeter Pfister, Alexander M. Rush
IEEE Trans. Vis. Comput. Graph.5
2019 MASR: A Modular Accelerator for Sparse RNNs
abstract
Recurrent neural networks (RNNs) are becoming the de-facto solution for speech recognition. RNNs exploit long-term temporal relationships in data by applying repeated, learned transformations. Unlike fully-connected (FC) layers with single vector matrix operations, RNN layers consist of hundreds of such operations chained over time. This poses challenges unique to RNNs that are not found in convolutional neural networks(CNNs) or FC models, namely large dynamic activation. In this paper we present MASR, a principled and modular architecture that accelerates bidirectional RNNs for on-chip ASR. MASR is designed to exploit sparsity in both dynamic activations and static weights. The architecture is enhanced by a series of dynamic activation optimizations that enable compact storage, ensure no energy is wasted computing null operations, and maintain high MAC utilization for highly parallel accelerator designs. In comparison to current state-of-the-art sparse neural network accelerators (e.g., EIE), MASR provides 2×area 3×energy, and 1.6×performance benefits. The modular nature of MASR enables designs that efficiently scale from resource-constrained low-power IoT applications to large-scale, highly parallel datacenter deployments.
Udit Gupta 0001, Brandon Reagen, Lillian Pentecost, Marco Donato, Thierry Tambe, Alexander M. Rush, Gu-Yeon Wei, David Brooks 0001
PACT6
2019 Don't Take the Premise for Granted: Mitigating Artifacts in Natural Language Inference
abstract
Natural Language Inference (NLI) datasets often contain hypothesis-only biases-artifacts that allow models to achieve non-trivial performance without learning whether a premise entails a hypothesis.We propose two probabilistic methods to build models that are more robust to such biases and better transfer across datasets.In contrast to standard approaches to NLI, our methods predict the probability of a premise given a hypothesis and NLI label, discouraging models from ignoring the premise.We evaluate our methods on synthetic and existing NLI datasets by training on datasets containing biases and testing on datasets containing no (or different) hypothesis-only biases.Our results indicate that these methods can make NLI models more robust to dataset-specific artifacts, transferring better than a baseline architecture in 9 out of 12 NLI datasets.Additionally, we provide an extensive analysis of the interplay of our methods with known biases in NLI datasets, as well as the effects of encouraging models to ignore biases and fine-tuning on target datasets.1 * * Equal contribution 1 Our code is available at https://github.com/ azpoliak/robust-nli.2 This hypothesis contradicts the premise and would likely not be inferred.
Yonatan Belinkov, Adam Poliak, Stuart M. Shieber, Benjamin Van Durme, Alexander M. Rush
ACL (1)5
2019 Compound Probabilistic Context-Free Grammars for Grammar Induction
abstract
We study a formalization of the grammar induction problem that models sentences as being generated by a compound probabilistic context free grammar.In contrast to traditional formulations which learn a single stochastic grammar, our context-free rule probabilities are modulated by a per-sentence continuous latent variable, which induces marginal dependencies beyond the traditional context-free assumptions.Inference in this grammar is performed by collapsed variational inference, in which an amortized variational posterior is placed on the continuous variable, and the latent trees are marginalized with dynamic programming.Experiments on English and Chinese show the effectiveness of our approach compared to recent state-of-the-art methods for grammar induction from words with neural language models.
Chris Dyer, Alexander M. Rush
ACL (1)3
2019 Simple Unsupervised Summarization by Contextual Matching
abstract
We propose an unsupervised method for sentence summarization using only language modeling.The approach employs two language models, one that is generic (i.e.pretrained), and the other that is specific to the target domain.We show that by using a productof-experts criteria these are enough for maintaining continuous contextual matching while maintaining output fluency.Experiments on both abstractive and extractive sentence summarization data sets show promising results of our method without being exposed to any paired data.
Jiawei Zhou 0012, Alexander M. Rush
ACL (1)2
2019 Avoiding Latent Variable Collapse with Generative Skip Models
abstract
Variational autoencoders (VAEs) learn distributions of high-dimensional data. They model data with a deep latent-variable model and then fit the model by maximizing a lower bound of the log marginal likelihood. VAEs can capture complex distributions, but they can also suffer from an issue known as "latent variable collapse," especially if the likelihood model is powerful. Specifically, the lower bound involves an approximate posterior of the latent variables; this posterior "collapses" when it is set equal to the prior, i.e., when the approximate posterior is independent of the data. While VAEs learn good generative models, latent variable collapse prevents them from learning useful representations. In this paper, we propose a simple new way to avoid latent variable collapse by including skip connections in our generative model; these connections enforce strong links between the latent variables and the likelihood function. We study generative skip models both theoretically and empirically. Theoretically, we prove that skip models increase the mutual information between the observations and the inferred latent variables. Empirically, we study images (MNIST and Omniglot) and text (Yahoo). Compared to existing VAE architectures, we show that generative skip models maintain similar predictive performance but lead to less collapse and provide more meaningful representations of the data.
Adji B. Dieng, Alexander M. Rush, David M. Blei
AISTATS3
2019 Commonsense Knowledge Mining from Pretrained Models
abstract
Joe Davison, Joshua Feldman, Alexander Rush. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Joe Davison, Joshua Feldman, Alexander M. Rush
EMNLP/IJCNLP (1)3
2019 Neural Linguistic Steganography
abstract
Zachary Ziegler, Yuntian Deng, Alexander Rush. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zachary M. Ziegler, Yuntian Deng, Alexander M. Rush
EMNLP/IJCNLP (1)3
2019 Tensor Variable Elimination for Plated Factor Graphs
abstract
A wide class of machine learning algorithms can be reduced to variable elimination on factor graphs. While factor graphs provide a unifying notation for these algorithms, they do not provide a compact way to express repeated structure when compared to plate diagrams for directed graphical models. To exploit efficient tensor algebra in graphs with plates of variables, we generalize undirected factor graphs to plated factor graphs and variable elimination to a tensor variable elimination algorithm that operates directly on plated factor graphs. Moreover, we generalize complexity bounds based on treewidth and characterize the class of plated factor graphs for which inference is tractable. As an application, we integrate tensor variable elimination into the Pyro probabilistic programming language to enable exact inference in discrete latent variable models with repeated structure. We validate our methods with experiments on both directed and undirected graphical models, including applications to polyphonic music modeling, animal movement modeling, and latent sentiment analysis.
Fritz Obermeyer, Eli Bingham, Martin Jankowiak, Neeraj Pradhan, Justin T. Chiu, Alexander M. Rush, Noah D. Goodman
ICML6
2019 Latent Normalizing Flows for Discrete Sequences
abstract
Normalizing flows are a powerful class of generative models for continuous random variables, showing both strong model flexibility and the potential for non-autoregressive generation. These benefits are also desired when modeling discrete random variables such as text, but directly applying normalizing flows to discrete sequences poses significant additional challenges. We propose a VAE-based generative model which jointly learns a normalizing flow-based distribution in the latent space and a stochastic mapping to an observed discrete space. In this setting, we find that it is crucial for the flow-based distribution to be highly multimodal. To capture this property, we propose several normalizing flow architectures to maximize model flexibility. Experiments consider common discrete sequence tasks of character-level language modeling and polyphonic music generation. Our results indicate that an autoregressive flow-based model can match the performance of a comparable autoregressive baseline, and a non-autoregressive flow-based model can improve generation speed with a penalty to performance.
Zachary M. Ziegler, Alexander M. Rush
ICML2
2019 Generating Abstractive Summaries with Finetuned Language Models
abstract
Neural abstractive document summarization is commonly approached by models that exhibit a mostly extractive behavior.This behavior is facilitated by a copy-attention which allows models to copy words from a source document.While models in the mostly extractive news summarization domain benefit from this inductive bias, they commonly fail to paraphrase or compress information from the source document.Recent advances in transferlearning from large pretrained language models give rise to alternative approaches that do not rely on copy-attention and instead learn to generate concise and abstractive summaries.In this paper, as part of the TL;DR challenge, we compare the abstractiveness of summaries from different summarization approaches and show that transfer-learning can be efficiently utilized without any changes to the model architecture.We demonstrate that the approach leads to a higher level of abstraction for a similar performance on the TL;DR challenge tasks, enabling true natural language compression.
Sebastian Gehrmann, Zachary M. Ziegler, Alexander M. Rush
INLG3
2019 Seq2seq-Vis: A Visual Debugging Tool for Sequence-to-Sequence Models
abstract
Neural sequence-to-sequence models have proven to be accurate and robust for many sequence prediction tasks, and have become the standard approach for automatic translation of text. The models work with a five-stage blackbox pipeline that begins with encoding a source sequence to a vector space and then decoding out to a new target sequence. This process is now standard, but like many deep learning methods remains quite difficult to understand or debug. In this work, we present a visual analysis tool that allows interaction and "what if"-style exploration of trained sequence-to-sequence models through each stage of the translation process. The aim is to identify which patterns have been learned, to detect model errors, and to probe the model with counterfactual scenario. We demonstrate the utility of our tool through several real-world sequence-to-sequence use cases on large-scale models.
Hendrik Strobelt, Sebastian Gehrmann, Michael Behrisch 0001, Adam Perer, Hanspeter Pfister, Alexander M. Rush
IEEE Trans. Vis. Comput. Graph.6
2018 Bottom-Up Abstractive Summarization
abstract
Neural network-based methods for abstractive summarization produce outputs that are more fluent than other techniques, but which can be poor at content selection.This work proposes a simple technique for addressing this issue: use a data-efficient content selector to overdetermine phrases in a source document that should be part of the summary.We use this selector as a bottom-up attention step to constrain the model to likely phrases.We show that this approach improves the ability to compress text, while still generating fluent summaries.This two-step process is both simpler and higher performing than other end-to-end content selection models, leading to significant improvements on ROUGE for both the CNN-DM and NYT corpus.Furthermore, the content selector can be trained with as little as 1,000 sentences, making it easy to transfer a trained summarizer to a new domain.
Sebastian Gehrmann, Yuntian Deng, Alexander M. Rush
EMNLP3
2018 Entity Tracking Improves Cloze-style Reading Comprehension
abstract
Reading comprehension tasks test the ability of models to process long-term context and remember salient information.Recent work has shown that relatively simple neural methods such as the Attention Sum-Reader can perform well on these tasks; however, these systems still significantly trail human performance.Analysis suggests that many of the remaining hard instances are related to the inability to track entity-references throughout documents.This work focuses on these hard entity tracking cases with two extensions: (1) additional entity features, and (2) training with a multi-task tracking objective.We show that these simple modifications improve performance both independently and in combination, and we outperform the previous state of the art on the LAMBADA dataset, particularly on difficult entity examples.
Luong Hoang, Sam Wiseman, Alexander M. Rush
EMNLP3
2018 Training for Diversity in Image Paragraph Captioning
abstract
Image paragraph captioning models aim to produce detailed descriptions of a source image.These models use similar techniques as standard image captioning models, but they have encountered issues in text generation, notably a lack of diversity between sentences, that have limited their effectiveness.In this work, we consider applying sequence-level training for this task.We find that standard self-critical training produces poor results, but when combined with an integrated penalty on trigram repetition produces much more diverse paragraphs.This simple training approach improves on the best result on the Visual Genome paragraph captioning dataset from 16.9 to 30.6 CIDEr, with gains on METEOR and BLEU as well, without requiring any architectural changes.
Luke Melas-Kyriazi, Alexander M. Rush, George Han
EMNLP2
2018 Learning Neural Templates for Text Generation
abstract
While neural, encoder-decoder models have had significant empirical success in text generation, there remain several unaddressed problems with this style of generation.Encoderdecoder models are largely (a) uninterpretable, and (b) difficult to control in terms of their phrasing or content.This work proposes a neural generation system using a hidden semimarkov model (HSMM) decoder, which learns latent, discrete templates jointly with learning to generate.We show that this model learns useful templates, and that these templates make generation both more interpretable and controllable.Furthermore, we show that this approach scales to real data sets and achieves strong performance nearing that of encoderdecoder text generation models.
Sam Wiseman, Stuart M. Shieber, Alexander M. Rush
EMNLP3
2018 Semi-Amortized Variational Autoencoders
abstract
Amortized variational inference (AVI) replaces instance-specific local inference with a global inference network. While AVI has enabled efficient training of deep generative models such as variational autoencoders (VAE), recent empirical work suggests that inference networks can produce suboptimal variational parameters. We propose a hybrid approach, to use AVI to initialize the variational parameters and run stochastic variational inference (SVI) to refine them. Crucially, the local SVI procedure is itself differentiable, so the inference network and generative model can be trained end-to-end with gradient-based optimization. This semi-amortized approach enables the use of rich generative models without experiencing the posterior-collapse phenomenon common in training VAEs for problems like text generation. Experiments show this approach outperforms strong autoregressive and variational baselines on standard text and image datasets.
Sam Wiseman, Andrew C. Miller, David A. Sontag, Alexander M. Rush
ICML5
2018 Weightless: Lossy weight encoding for deep neural network compression
abstract
The large memory requirements of deep neural networks limit their deployment and adoption on many devices. Model compression methods effectively reduce the memory requirements of these models, usually through applying transformations such as weight pruning or quantization. In this paper, we present a novel scheme for lossy weight encoding co-designed with weight simplification techniques. The encoding is based on the Bloomier filter, a probabilistic data structure that can save space at the cost of introducing random errors. Leveraging the ability of neural networks to tolerate these imperfections and by re-training around the errors, the proposed technique, named Weightless, can compress weights by up to 496x without loss of model accuracy. This results in up to a 1.51x improvement over the state-of-the-art.
Brandon Reagen, Udit Gupta 0001, Robert Adolf, Michael Mitzenmacher, Alexander M. Rush, Gu-Yeon Wei, David Brooks 0001
ICML5
2018 Adversarially Regularized Autoencoders
abstract
Deep latent variable models, trained using variational autoencoders or generative adversarial networks, are now a key technique for representation learning of continuous structures. However, applying similar methods to discrete structures, such as text sequences or discretized images, has proven to be more challenging. In this work, we propose a more flexible method for training deep latent variable models of discrete structures. Our approach is based on the recently proposed Wasserstein Autoencoder (WAE) which formalizes adversarial autoencoders as an optimal transport problem. We first extend this framework to model discrete sequences, and then further explore different learned priors targeting a controllable representation. Unlike many other latent variable generative models for text, this adversarially regularized autoencoder (ARAE) allows us to generate fluent textual outputs as well as perform manipulations in the latent space to induce change in the output space. Finally we show that the latent representation can be trained to perform unaligned textual style transfer, giving improvements both in automatic measures and human evaluation.
Junbo Jake Zhao, Kelly Zhang, Alexander M. Rush, Yann LeCun
ICML4
2018 End-to-End Content and Plan Selection for Data-to-Text Generation
abstract
Learning to generate fluent natural language from structured data with neural networks has become an common approach for NLG.This problem can be challenging when the form of the structured data varies between examples.This paper presents a survey of several extensions to sequence-to-sequence models to account for the latent content selection process, particularly variants of copy attention and coverage decoding.We further propose a training method based on diverse ensembling to encourage models to learn distinct sentence templates during training.An empirical evaluation of these techniques shows an increase in the quality of generated text across five automated metrics, as well as human evaluation.
Sebastian Gehrmann, Falcon Z. Dai, Henry Elder, Alexander M. Rush
INLG4
2018 Latent Alignment and Variational Attention
abstract
Neural attention has become central to many state-of-the-art models in natural language processing and related domains. Attention networks are an easy-to-train and effective method for softly simulating alignment; however, the approach does not marginalize over latent alignments in a probabilistic sense. This property makes it difficult to compare attention to other alignment approaches, to compose it with probabilistic models, and to perform posterior inference conditioned on observed data. A related latent approach, hard attention, fixes these issues, but is generally harder to train and less accurate. This work considers variational attention networks, alternatives to soft and hard attention for learning latent variable alignment models, with tighter approximation bounds based on amortized variational inference. We further propose methods for reducing the variance of gradients to make these approaches computationally feasible. Experiments show that for machine translation and visual question answering, inefficient exact latent variable models outperform standard neural attention, but these gains go away when using hard attention based training. On the other hand, variational attention retains most of the performance gain but with training speed comparable to neural attention.
Yuntian Deng, Justin T. Chiu, Demi Guo, Alexander M. Rush
NeurIPS5
2018 LSTMVis: A Tool for Visual Analysis of Hidden State Dynamics in Recurrent Neural Networks
abstract
Recurrent neural networks, and in particular long short-term memory (LSTM) networks, are a remarkably effective tool for sequence modeling that learn a dense black-box hidden representation of their sequential input. Researchers interested in better understanding these models have studied the changes in hidden state representations over time and noticed some interpretable patterns but also significant noise. In this work, we present LSTMVis, a visual analysis tool for recurrent neural networks with a focus on understanding these hidden state dynamics. The tool allows users to select a hypothesis input range to focus on local state changes, to match these states changes to similar patterns in a large data set, and to align these results with structural annotations from their domain. We show several use cases of the tool for analyzing specific hidden state properties on dataset containing nesting, phrase structure, and chord progressions, and demonstrate how the tool can be used to isolate patterns for further statistical analysis. We characterize the domain, the different stakeholders, and their goals and tasks. Long-term usage data after putting the tool online revealed great interest in the machine learning community.
Hendrik Strobelt, Sebastian Gehrmann, Hanspeter Pfister, Alexander M. Rush
IEEE Trans. Vis. Comput. Graph.4
2017 Adapting Sequence Models for Sentence Correction
abstract
In a controlled experiment of sequence-tosequence approaches for the task of sentence correction, we find that characterbased models are generally more effective than word-based models and models that encode subword information via convolutions, and that modeling the output data as a series of diffs improves effectiveness over standard approaches.Our strongest sequence-to-sequence model improves over our strongest phrase-based statistical machine translation model, with access to the same data, by 6 M 2 (0.5 GLEU) points.Additionally, in the data environment of the standard CoNLL-2014 setup, we demonstrate that modeling (and tuning against) diffs yields similar or better M 2 scores with simpler models and/or significantly less data than previous sequence-to-sequence approaches.
Allen Schmaltz, Alexander M. Rush, Stuart M. Shieber
EMNLP3
2017 Challenges in Data-to-Document Generation
abstract
Recent neural models have shown significant progress on the problem of generating short descriptive texts conditioned on a small number of database records.In this work, we suggest a slightly more difficult data-to-text generation task, and investigate how effective current approaches are on this task.In particular, we introduce a new, large-scale corpus of data records paired with descriptive documents, propose a series of extractive evaluation methods for analyzing performance, and obtain baseline results using current neural generation methods.Experiments show that these models produce fluent text, but fail to convincingly approximate humangenerated documents.Moreover, even templated baselines exceed the performance of these neural models on some metrics, though copy-and reconstructionbased extensions lead to noticeable improvements.
Sam Wiseman, Stuart M. Shieber, Alexander M. Rush
EMNLP3
2017 Structured Attention Networks
Carl Denton, Luong Hoang, Alexander M. Rush
ICLR (Poster)4
2017 Lie-Access Neural Turing Machines
Greg Yang, Alexander M. Rush
ICLR (Poster)2
2017 Image-to-Markup Generation with Coarse-to-Fine Attention
abstract
We present a neural encoder-decoder model to convert images into presentational markup based on a scalable coarse-to-fine attention mechanism. Our method is evaluated in the context of image-to-LaTeX generation, and we introduce a new dataset of real-world rendered mathematical expressions paired with LaTeX markup. We show that unlike neural OCR techniques using CTC-based models, attention-based approaches can tackle this non-standard OCR task. Our approach outperforms classical mathematical OCR systems by a large margin on in-domain rendered data, and, with pretraining, also performs well on out-of-domain handwritten data. To reduce the inference complexity associated with the attention-based approaches, we introduce a new coarse-to-fine attention layer that selects a support region before applying attention.
Yuntian Deng, Anssi Kanervisto, Jeffrey Ling, Alexander M. Rush
ICML4
2016 Character-Aware Neural Language Models
abstract
We describe a simple neural language model that relies only on character-level inputs. Predictions are still made at the word-level. Our model employs a convolutional neural network (CNN) and a highway net work over characters, whose output is given to a long short-term memory (LSTM) recurrent neural network language model (RNN-LM). On the English Penn Treebank the model is on par with the existing state-of-the-art despite having 60% fewer parameters. On languages with rich morphology (Arabic, Czech, French, German, Spanish, Russian), the model outperforms word-level/morpheme-level LSTM baselines, again with fewer parameters. The results suggest that on many languages, character inputs are sufficient for language modeling. Analysis of word representations obtained from the character composition part of the model reveals that the model is able to encode, from characters only, both semantic and orthographic information.
Yacine Jernite, David A. Sontag, Alexander M. Rush
AAAI4
2016 Sequence-Level Knowledge Distillation
abstract
Neural machine translation (NMT) offers a novel alternative formulation of translation that is potentially simpler than statistical approaches.However to reach competitive performance, NMT models need to be exceedingly large.In this paper we consider applying knowledge distillation approaches (Bucila et al., 2006;Hinton et al., 2015) that have proven successful for reducing the size of neural models in other domains to the problem of NMT.We demonstrate that standard knowledge distillation applied to word-level prediction can be effective for NMT, and also introduce two novel sequence-level versions of knowledge distillation that further improve performance, and somewhat surprisingly, seem to eliminate the need for beam search (even when applied on the original teacher model).Our best student model runs 10 times faster than its state-of-the-art teacher with little loss in performance.It is also significantly better than a baseline model trained without knowledge distillation: by 4.2/1.7 BLEU with greedy decoding/beam search.Applying weight pruning on top of knowledge distillation results in a student model that has 13× fewer parameters than the original teacher model, with a decrease of 0.4 BLEU.
Alexander M. Rush
EMNLP2
2016 An Embedding Model for Predicting Roll-Call Votes
abstract
We develop a novel embedding-based model for predicting legislative roll-call votes from bill text.The model introduces multidimensional ideal vectors for legislators as an alternative to single dimensional ideal point models for quantitatively analyzing roll-call data.These vectors are learned to correspond with pre-trained word embeddings which allows us to analyze which features in a bill text are most predictive of political support.Our model is quite simple, while at the same time allowing us to successfully predict legislator votes on specific bills with higher accuracy than past methods.
Peter Kraft, Hirsh Jain, Alexander M. Rush
EMNLP3
2016 Word Ordering Without Syntax
abstract
Recent work on word ordering has argued that syntactic structure is important, or even required, for effectively recovering the order of a sentence. We find that, in fact, an n-gram language model with a simple heuristic gives strong results on this task. Furthermore, we show that a long short-term memory (LSTM) language model is even more effective at recovering order, with our basic model outperforming a state-of-the-art syntactic model by 11.5 BLEU points. Additional data and larger beams yield further gains, at the expense of training and search time.
Allen Schmaltz, Alexander M. Rush, Stuart M. Shieber
EMNLP2
2016 Sequence-to-Sequence Learning as Beam-Search Optimization
abstract
Sequence-to-Sequence (seq2seq) modeling has rapidly become an important generalpurpose NLP tool that has proven effective for many text-generation and sequence-labeling tasks.Seq2seq builds on deep neural language modeling and inherits its remarkable accuracy in estimating local, next-word distributions.In this work, we introduce a model and beamsearch training scheme, based on the work of Daumé III and Marcu (2005), that extends seq2seq to learn global sequence scores.This structured approach avoids classical biases associated with local training and unifies the training loss with the test-time usage, while preserving the proven model architecture of seq2seq and its efficient training approach.We show that our system outperforms a highlyoptimized attention-based seq2seq system and other baselines on three different sequence to sequence tasks: word ordering, parsing, and machine translation.
Sam Wiseman, Alexander M. Rush
EMNLP2
2016 Abstractive Sentence Summarization with Attentive Recurrent Neural Networks
abstract
Abstractive Sentence Summarization generates a shorter version of a given sentence while attempting to preserve its meaning.We introduce a conditional recurrent neural network (RNN) which generates a summary of an input sentence.The conditioning is provided by a novel convolutional attention-based encoder which ensures that the decoder focuses on the appropriate input words at each step of generation.Our model relies only on learned features and is easy to train in an end-to-end fashion on large data sets.Our experiments show that the model significantly outperforms the recently proposed state-of-the-art method on the Gigaword corpus while performing competitively on the DUC-2004 shared task.
Sumit Chopra, Michael Auli, Alexander M. Rush
HLT-NAACL3
2016 Learning Global Features for Coreference Resolution
abstract
There is compelling evidence that coreference prediction would benefit from modeling global information about entity-clusters.Yet, state-of-the-art performance can be achieved with systems treating each mention prediction independently, which we attribute to the inherent difficulty of crafting informative clusterlevel features.We instead propose to use recurrent neural networks (RNNs) to learn latent, global representations of entity clusters directly from their mentions.We show that such representations are especially useful for the prediction of pronominal mentions, and can be incorporated into an end-to-end coreference system that outperforms the state of the art without requiring any additional search.
Sam Wiseman, Alexander M. Rush, Stuart M. Shieber
HLT-NAACL2
2015 Learning Anaphoricity and Antecedent Ranking Features for Coreference Resolution
abstract
Sam Wiseman, Alexander M. Rush, Stuart Shieber, Jason Weston. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Sam Wiseman, Alexander M. Rush, Stuart M. Shieber, Jason Weston
ACL (1)2
2015 A Neural Attention Model for Abstractive Sentence Summarization
abstract
Summarization based on text extraction is inherently limited, but generation-style abstractive methods have proven challenging to build.In this work, we propose a fully data-driven approach to abstractive sentence summarization.Our method utilizes a local attention-based model that generates each word of the summary conditioned on the input sentence.While the model is structurally simple, it can easily be trained end-to-end and scales to a large amount of training data.The model shows significant performance gains on the DUC-2004 shared task compared with several strong baselines.
Alexander M. Rush, Sumit Chopra, Jason Weston
EMNLP1
2015 A Fast Variational Approach for Learning Markov Random Field Language Models
abstract
Language modelling is a fundamental building block of natural language processing. However, in practice the size of the vocabulary limits the distributions applicable for this task: specifically, one has to either resort to local optimization methods, such as those used in neural language models, or work with heavily constrained distributions. In this work, we take a step towards overcoming these difficulties. We present a method for global-likelihood optimization of a Markov random field language model exploiting long-range contexts in time independent of the corpus size. We take a variational approach to optimizing the likelihood and exploit underlying symmetries to greatly simplify learning. We demonstrate the efficiency of this method both for language modelling and for part-of-speech tagging.
Yacine Jernite, Alexander M. Rush, David A. Sontag
ICML2
2015 Transforming Dependencies into Phrase Structures
abstract
We present a new algorithm for transforming dependency parse trees into phrase-structure parse trees.We cast the problem as structured prediction and learn a statistical model.Our algorithm is faster than traditional phrasestructure parsing and achieves 90.4% English parsing accuracy and 82.4% Chinese parsing accuracy, near to the state of the art on both benchmarks.
Lingpeng Kong, Alexander M. Rush, Noah A. Smith
HLT-NAACL2
2014 A Constrained Viterbi Relaxation for Bidirectional Word Alignment
abstract
Bidirectional models of word alignment are an appealing alternative to post-hoc combinations of directional word aligners.Unfortunately, most bidirectional formulations are NP-Hard to solve, and a previous attempt to use a relaxationbased decoder yielded few exact solutions (6%).We present a novel relaxation for decoding the bidirectional model of DeNero and Macherey (2011).The relaxation can be solved with a modified version of the Viterbi algorithm.To find optimal solutions on difficult instances, we alternate between incrementally adding constraints and applying optimality-preserving coarse-to-fine pruning.The algorithm finds provably exact solutions on 86% of sentence pairs and shows improvements over directional models.
Yin-Wen Chang, Alexander M. Rush, John DeNero, Michael Collins 0001
ACL (1)2
2013 Spectral Learning of Refinement HMMs
Karl Stratos, Alexander M. Rush, Shay B. Cohen, Michael Collins 0001
CoNLL2
2013 Optimal Beam Search for Machine Translation
abstract
Beam search is a fast and empirically effective method for translation decoding, but it lacks formal guarantees about search error.We develop a new decoding algorithm that combines the speed of beam search with the optimal certificate property of Lagrangian relaxation, and apply it to phrase-and syntax-based translation decoding.The new method is efficient, utilizes standard MT algorithms, and returns an exact solution on the majority of translation examples in our test data.The algorithm is 3.5 times faster than an optimized incremental constraint-based decoder for phrase-based translation and 4 times faster for syntax-based translation.
Alexander M. Rush, Yin-Wen Chang, Michael Collins 0001
EMNLP1
2012 Improved Parsing and POS Tagging Using Inter-Sentence Consistency Constraints
Alexander M. Rush, Roi Reichart, Michael Collins 0001, Amir Globerson
EMNLP-CoNLL1
2012 Vine Pruning for Efficient Multi-Pass Dependency Parsing
Alexander M. Rush, Slav Petrov
HLT-NAACL1
2012 A Tutorial on Dual Decomposition and Lagrangian Relaxation for Inference in Natural Language Processing
abstract
Dual decomposition, and more generally Lagrangian relaxation, is a classical method for combinatorial optimization; it has recently been applied to several inference problems in natural language processing (NLP). This tutorial gives an overview of the technique. We describe example algorithms, describe formal guarantees for the method, and describe practical issues in implementing the algorithms. While our examples are predominantly drawn from the NLP literature, the material should be of general relevance to inference problems in machine learning. A central theme of this tutorial is that Lagrangian relaxation is naturally applied in conjunction with a broad class of combinatorial algorithms, allowing inference in models that go significantly beyond previous work on Lagrangian relaxation for inference in graphical models.
Alexander M. Rush, Michael Collins 0001
J. Artif. Intell. Res.1
2011 Exact Decoding of Syntactic Translation Models through Lagrangian Relaxation
Alexander M. Rush, Michael Collins 0001
ACL1
2010 Dual Decomposition for Parsing with Non-Projective Head Automata
Terry Koo, Alexander M. Rush, Michael Collins 0001, Tommi S. Jaakkola, David A. Sontag
EMNLP2
2010 On Dual Decomposition and Linear Programming Relaxations for Natural Language Processing
Alexander M. Rush, David A. Sontag, Michael Collins 0001, Tommi S. Jaakkola
EMNLP1