EDBT 2026 Demo / reviewers in the wild / expert
Mike Lewis
dblp:19/6214
· DBLP profile ↗
66ranked-venue papers
11as first author
36since 2021 · last 2025
0000-0003-0679-6612ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 66 · 11 first-author · 36 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Byte Latent Transformer: Patches Scale Better Than TokensabstractArtidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason E Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srini Iyer. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez 0001, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srinivasan Iyer 0001 |
ACL (1) | 12 |
| 2025 | BTS: Harmonizing Specialized Experts into a Generalist LLMabstractQizhen Zhang, Prajjwal Bhargava, Chloe Bi, Chris X. Cai, Jakob Nicolaus Foerster, Jeremy Fu, Punit Singh Koura, Ruan Silva, Sheng Shen, Emily Dinan, Suchin Gururangan, Mike Lewis. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Qizhen Zhang 0002, Prajjwal Bhargava, Chloe Bi, Chris X. Cai, Jakob N. Foerster, Jeremy Fu, Punit Singh Koura, Ruan Silva, Sheng Shen 0016, Emily Dinan, Suchin Gururangan, Mike Lewis |
EMNLP | 12 |
| 2025 | Law of the Weakest Link: Cross Capabilities of Large Language ModelsabstractThe development and evaluation of Large Language Models (LLMs) have largely focused on individual capabilities. However, this overlooks the intersection of multiple abilities across different types of expertise that are often required for real-world tasks, which we term **cross capabilities**. To systematically explore this concept, we first define seven core individual capabilities and then pair them to form seven common cross capabilities, each supported by a manually constructed taxonomy. Building on these definitions, we introduce *CrossEval*, a benchmark comprising 1,400 human-annotated prompts, with 100 prompts for each individual and cross capability. To ensure reliable evaluation, we involve expert annotators to assess 4,200 model responses, gathering 8,400 human ratings with detailed explanations to serve as reference examples. Our findings reveal that current LLMs consistently exhibit the ``Law of the Weakest Link,'' where cross-capability performance is significantly constrained by the weakest component. Across 58 cross-capability scores from 17 models, 38 scores are lower than all individual capabilities, while 20 fall between strong and weak, but closer to the weaker ability. These results highlight LLMs' underperformance in cross-capability tasks, emphasizing the need to identify and improve their weakest capabilities as a key research priority. The code, benchmarks, and evaluations are available on our [project website](https://www.llm-cross-capabilities.org). Ming Zhong 0005, Aston Zhang, Wenhan Xiong, Chenguang Zhu 0001, Zhengxing Chen, Chloe Bi, Mike Lewis, Sravya Popuri, Sharan Narang, Melanie Kambadur, Dhruv Mahajan 0001, Sergey Edunov, Jiawei Han 0001, Laurens van der Maaten |
ICLR | 10 |
| 2025 | FlexOLMo: Open Language Models for Flexible Data UseabstractWe introduce FlexOLMo, a new class of language models (LMs) that supports (1) distributed training without data sharing, where different model parameters are independently trained on private datasets, and (2) data-flexible inference, where these parameters along with their associated data can be easily included or excluded from model inferences with no further training. FlexOLMo employs a mixture-of-experts (MoE) architecture where each expert is trained independently on private datasets and later integrated through a new nonparametric routing without any joint training across datasets. FlexOLMo is trained on FLEXMIX, a corpus we curate comprising seven restricted sets, either real or realistic approximations, alongside publicly available datasets. We evaluate models with up to 37 billion parameters (20 billion active) on 31 diverse downstream tasks. We show that a general expert trained on public data can be effectively combined with independently trained experts from other data owners significantly benefiting from these restricted sets (an average 41% relative improvement) while allowing flexible opt-out at inference time (e.g., for users without appropriate licenses or permissions). Our approach also outperforms prior model merging methods by 10.1% on average and surpasses the standard MoE trained without data restrictions using the same training FLOPs. Altogether, FlexOLMo enables training on restricted data while keeping data local and supports fine-grained control of data access at inference. Akshita Bhagia, Kevin Farhat, Niklas Muennighoff, Jacob Morrison, Pete Walsh 0001, Dustin Schwenk, Shayne Longpre, Jake Poznanski, Allyson Ettinger, Daogao Liu, Margaret Li, Mike Lewis, Scott Yih, Dirk Groeneveld, Luca Soldaini, Kyle Lo, Noah A. Smith, Luke Zettlemoyer, Pang Wei W. Koh, Hannaneh Hajishirzi, Ali Farhadi, Sewon Min |
NeurIPS | 13 |
| 2024 | Self-Alignment with Instruction BacktranslationabstractWe present a scalable method to build a high quality instruction following language model by automatically labelling human-written text with corresponding instructions. Our approach, named instruction backtranslation, starts with a language model finetuned on a small amount of seed data, and a given web corpus. The seed model is used to construct training examples by generating instruction prompts for web documents (self-augmentation), and then selecting high quality examples from among these candidates (self-curation). This data is then used to finetune a stronger model. Finetuning LLaMa on two iterations of our approach yields a model that outperforms all other LLaMa-based models on the Alpaca leaderboard not relying on distillation data, demonstrating highly effective self-alignment. Xian Li 0003, Chunting Zhou, Timo Schick, Omer Levy, Luke Zettlemoyer, Jason Weston, Mike Lewis |
ICLR | 8 |
| 2024 | RA-DIT: Retrieval-Augmented Dual Instruction TuningabstractRetrieval-augmented language models (RALMs) improve performance by accessing long-tail and up-to-date knowledge from external data stores, but are challenging to build. Existing approaches require either expensive retrieval-specific modifications to LM pre-training or use post-hoc integration of the data store that leads to suboptimal performance. We introduce Retrieval-Augmented Dual Instruction Tuning (RA-DIT), a lightweight fine-tuning methodology that provides a third option by retrofitting any LLM with retrieval capabilities. Our approach operates in two distinct fine-tuning steps: (1) one updates a pre-trained LM to better use retrieved information, while (2) the other updates the retriever to return more relevant results, as preferred by the LM. By fine-tuning over tasks that require both knowledge utilization and contextual awareness, we demonstrate that each stage yields significant performance improvements, and using both leads to additional gains. Our best model, RA-DIT 65B, achieves state-of-the-art performance across a range of knowledge-intensive zero- and few-shot learning benchmarks, significantly outperforming existing in-context RALM approaches by up to +8.9% in 0-shot setting and +1.4% in 5-shot setting on average. Xi Victoria Lin, Xilun Chen 0002, Mingda Chen, Maria Lomeli, Richard James 0001, Pedro Rodríguez 0001, Jacob Kahn, Gergely Szilvasy, Mike Lewis, Luke Zettlemoyer, Scott Yih |
ICLR | 10 |
| 2024 | In-Context Pretraining: Language Modeling Beyond Document BoundariesabstractLanguage models are currently trained to predict tokens given document prefixes, enabling them to zero shot long form generation and prompting-style tasks which can be reduced to document completion. We instead present IN-CONTEXT PRETRAINING, a new approach where language models are trained on a sequence of related documents, thereby explicitly encouraging them to read and reason across document boundaries. Our approach builds on the fact that current pipelines train by concatenating random sets of shorter documents to create longer context windows; this improves efficiency even though the prior documents provide no signal for predicting the next document. Given this fact, we can do IN-CONTEXT PRETRAINING by simply changing the document ordering so that each context contains related documents, and directly applying existing pretraining pipelines. However, this document sorting problem is challenging. There are billions of documents and we would like the sort to maximize contextual similarity for every document without repeating any data. To do this, we introduce approximate algorithms for finding related documents with efficient nearest neighbor search and constructing coherent batches with a graph cover algorithm. Our experiments show IN-CONTEXT PRETRAINING offers a scalable and simple approach to significantly enhance LM performance: we see notable improvements in tasks that require more complex contextual reasoning, including in-context learning (+8%), reading comprehension (+15%), faithfulness to previous contexts (+16%), long-context reasoning (+5%), and retrieval augmentation (+9%). Sewon Min, Maria Lomeli, Chunting Zhou, Margaret Li, Xi Victoria Lin, Noah A. Smith, Luke Zettlemoyer, Scott Yih, Mike Lewis |
ICLR | 10 |
| 2024 | Efficient Streaming Language Models with Attention SinksabstractDeploying Large Language Models (LLMs) in streaming applications such as multi-round dialogue, where long interactions are expected, is urgently needed but poses two major challenges.
Firstly, during the decoding stage, caching previous tokens' Key and Value states (KV) consumes extensive memory.
Secondly, popular LLMs cannot generalize to longer texts than the training sequence length.
Window attention, where only the most recent KVs are cached, is a natural approach --- but we show that it fails when the text length surpasses the cache size.
We observe an interesting phenomenon, namely attention sink, that keeping the KV of initial tokens will largely recover the performance of window attention. In this paper, we first demonstrate that the emergence of attention sink is due to the strong attention scores towards initial tokens as a ``sink'' even if they are not semantically important.
Based on the above analysis, we introduce StreamingLLM, an efficient framework that enables LLMs trained with a finite length attention window to generalize to infinite sequence length without any fine-tuning.
We show that StreamingLLM can enable Llama-2, MPT, Falcon, and Pythia to perform stable and efficient language modeling with up to 4 million tokens and more.
In addition, we discover that adding a placeholder token as a dedicated attention sink during pre-training can further improve streaming deployment. In streaming settings, StreamingLLM outperforms the sliding window recomputation baseline by up to 22.2$\times$ speedup.
Code and datasets are provided at https://github.com/mit-han-lab/streaming-llm. Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 0003, Mike Lewis |
ICLR | 5 |
| 2024 | REPLUG: Retrieval-Augmented Black-Box Language ModelsabstractWeijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Richard James, Mike Lewis, Luke Zettlemoyer, Wen-tau Yih. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Sewon Min, Michihiro Yasunaga, Minjoon Seo, Richard James 0001, Mike Lewis, Luke Zettlemoyer, Scott Yih |
NAACL-HLT | 6 |
| 2024 | Effective Long-Context Scaling of Foundation ModelsabstractWenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, Hao Ma. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Wenhan Xiong, Igor Molybog, Prajjwal Bhargava, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, Hao Ma 0001 |
NAACL-HLT | 19 |
| 2023 | Contrastive Decoding: Open-ended Text Generation as OptimizationabstractXiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, Mike Lewis. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Xiang Li 0063, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori B. Hashimoto, Luke Zettlemoyer, Mike Lewis |
ACL (1) | 8 |
| 2023 | FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationabstractSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Scott Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi |
EMNLP | 4 |
| 2023 | InCoder: A Generative Model for Code Infilling and Synthesis
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida I. Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, Mike Lewis |
ICLR | 10 |
| 2023 | Progressive Prompts: Continual Learning for Language Models
Anastasia Razdaibiedina, Yuning Mao, Madian Khabsa, Mike Lewis, Amjad Almahairi |
ICLR | 5 |
| 2023 | Retrieval-Augmented Multimodal Language ModelingabstractRecent multimodal models such as DALL-E and CM3 have achieved remarkable progress in text-to-image and image-to-text generation. However, these models store all their knowledge (e.g., the appearance of the Eiffel Tower) in the model parameters, requiring increasingly larger models and training data to capture more knowledge. To integrate knowledge in a more scalable and modular way, we propose a retrieval-augmented multimodal model, which enables a base multimodal model (generator) to refer to relevant text and images fetched by a retriever from external memory (e.g., documents on the web). Specifically, for the retriever, we use a pretrained CLIP, and for the generator, we train a CM3 Transformer on the LAION dataset. Our resulting model, named Retrieval-Augmented CM3 (RA-CM3), is the first multimodal model that can retrieve and generate both text and images. We show that RA-CM3 significantly outperforms baseline multimodal models such as DALL-E and CM3 on both image and caption generation tasks (12 FID and 17 CIDEr improvements on MS-COCO), while requiring much less compute for training (<30% of DALL-E). Moreover, we show that RA-CM3 exhibits novel capabilities such as faithful image generation and multimodal in-context learning (e.g., image generation from demonstrations). Michihiro Yasunaga, Armen Aghajanyan, Richard James 0001, Jure Leskovec, Percy Liang, Mike Lewis, Luke Zettlemoyer, Scott Yih |
ICML | 7 |
| 2023 | Coder Reviewer Reranking for Code GenerationabstractSampling diverse programs from a code language model and reranking with model likelihood is a popular method for code generation but it is prone to preferring degenerate solutions. Inspired by collaborative programming, we propose Coder-Reviewer reranking. We augment Coder language models from past work, which generate programs given language instructions, with Reviewer models, which evaluate the likelihood of the instruction given the generated programs. We perform an extensive study across six datasets with eight models from three model families. Experimental results show that Coder-Reviewer reranking leads to consistent and significant improvement (up to 17% absolute accuracy gain) over reranking with the Coder model only. When combined with executability filtering, Coder-Reviewer reranking can often outperform the minimum Bayes risk method. Coder-Reviewer reranking is easy to implement by prompting, can generalize to different programming languages, and works well with off-the-shelf hyperparameters. Tao Yu 0009, Tatsunori B. Hashimoto, Mike Lewis, Scott Yih, Daniel Fried, Sida I. Wang |
ICML | 4 |
| 2023 | MEGABYTE: Predicting Million-byte Sequences with Multiscale TransformersabstractAutoregressive transformers are spectacular models for short sequences but scale poorly to long sequences such as high-resolution images, podcasts, code, or books. We proposed Megabyte, a multi-scale decoder architecture that enables end-to-end differentiable modeling of sequences of over one million bytes. Megabyte segments sequences into patches and uses a local submodel within patches and a global model between patches. This enables sub-quadratic self-attention, much larger feedforward layers for the same compute, and improved parallelism during decoding---unlocking better performance at reduced cost for both training and generation. Extensive experiments show that Megabyte allows byte-level models to perform competitively with subword models on long context language modeling, achieve state-of-the-art density estimation on ImageNet, and model audio from raw files. Together, these results establish the viability of tokenization-free autoregressive sequence modeling at scale. Lili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, Mike Lewis |
NeurIPS | 6 |
| 2023 | LIMA: Less Is More for AlignmentabstractLarge language models are trained in two stages: (1) unsupervised pretraining from raw text, to learn general-purpose representations, and (2) large scale instruction tuning and reinforcement learning, to better align to end tasks and user preferences.
We measure the relative importance of these two stages by training LIMA, a 65B parameter LLaMa language model fine-tuned with the standard supervised loss on only 1,000 carefully curated prompts and responses, without any reinforcement learning or human preference modeling.
LIMA demonstrates remarkably strong performance, learning to follow specific response formats from only a handful of examples in the training data, including complex queries that range from planning trip itineraries to speculating about alternate history.
Moreover, the model tends to generalize well to unseen tasks that did not appear in the training data.
In a controlled human study, responses from LIMA are either equivalent or strictly preferred to GPT-4 in 43\% of cases; this statistic is as high as 58\% when compared to Bard and 65\% versus DaVinci003, which was trained with human feedback.
Taken together, these results strongly suggest that almost all knowledge in large language models is learned during pretraining, and only limited instruction tuning data is necessary to teach models to produce high quality output. Chunting Zhou, Puxin Xu, Srinivasan Iyer 0001, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Lili Yu, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, Omer Levy |
NeurIPS | 13 |
| 2023 | Questions Are All You Need to Train a Dense Passage RetrieverabstractAbstract We introduce ART, a new corpus-level autoencoding approach for training dense retrieval models that does not require any labeled training data. Dense retrieval is a central challenge for open-domain tasks, such as Open QA, where state-of-the-art methods typically require large supervised datasets with custom hard-negative mining and denoising of positive examples. ART, in contrast, only requires access to unpaired inputs and outputs (e.g., questions and potential answer passages). It uses a new passage-retrieval autoencoding scheme, where (1) an input question is used to retrieve a set of evidence passages, and (2) the passages are then used to compute the probability of reconstructing the original question. Training for retrieval based on question reconstruction enables effective unsupervised learning of both passage and question encoders, which can be later incorporated into complete Open QA systems without any further finetuning. Extensive experiments demonstrate that ART obtains state-of-the-art results on multiple QA retrieval benchmarks with only generic initialization from a pre-trained language model, removing the need for labeled data and task-specific losses.1 Our code and model checkpoints are available at: https://github.com/DevSinghSachan/art. Devendra Singh Sachan, Mike Lewis, Dani Yogatama, Luke Zettlemoyer, Joelle Pineau, Manzil Zaheer |
Trans. Assoc. Comput. Linguistics | 2 |
| 2023 | LegoNN: Building Modular Encoder-Decoder ModelsabstractState-of-the-art encoder-decoder models (e.g. for machine translation (MT) or automatic speech recognition (ASR)) are constructed and trained end-to-end as an atomic unit. No component of the model can be (re-)used without the others, making it impossible to share parts, e.g. a high resourced decoder, across tasks. We describe LegoNN, a procedure for building encoder-decoder architectures in a way so that its parts can be applied to other tasks without the need for any fine-tuning. To achieve this reusability, the interface between encoder and decoder modules is grounded to a sequence of marginal distributions over a pre-defined discrete vocabulary. We present two approaches for ingesting these marginals; one is differentiable, allowing the flow of gradients across the entire network, and the other is gradient-isolating. To enable the portability of decoder modules between MT tasks for different source languages and across other tasks like ASR, we introduce a modality agnostic encoder which consists of a length control mechanism to dynamically adapt encoders' output lengths in order to match the expected input length range of pre-trained decoders. We present several experiments to demonstrate the effectiveness of LegoNN models: a trained language generation LegoNN decoder module from German-English (De-En) MT task can be reused without any fine-tuning for the Europarl English ASR and the Romanian-English (Ro-En) MT tasks, matching or beating the performance of baseline. After fine-tuning, LegoNN models improve the Ro-En MT task by 1.5 BLEU points and achieve 12.5% relative WER reduction on the Europarl ASR task. To show how the approach generalizes, we compose a LegoNN ASR model from three modules – each has been learned within different end-to-end trained models on three different datasets – achieving an overall WER reduction of 19.5%. Siddharth Dalmia, Dmytro Okhonko, Mike Lewis, Sergey Edunov, Shinji Watanabe 0001, Florian Metze, Luke Zettlemoyer, Abdel-rahman Mohamed |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Noisy Channel Language Model Prompting for Few-Shot Text ClassificationabstractWe introduce a noisy channel approach for language model prompting in few-shot text classification.Instead of computing the likelihood of the label given the input (referred as direct models), channel models compute the conditional probability of the input given the label, and are thereby required to explain every word in the input.We use channel models for recently proposed few-shot learning methods with no or very limited updates to the language model parameters, via either in-context demonstration or prompt tuning.Our experiments show that, for both methods, channel models significantly outperform their direct counterparts, which we attribute to their stability, i.e., lower variance and higher worstcase accuracy.We also present extensive ablations that provide recommendations for when to use channel prompt tuning instead of other competitive methods (e.g., direct head tuning): channel prompt tuning is preferred when the number of training examples is small, labels in the training data are imbalanced, or generalization to unseen labels is required. Sewon Min, Mike Lewis, Hannaneh Hajishirzi, Luke Zettlemoyer |
ACL (1) | 2 |
| 2022 | Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?abstractLarge language models (LMs) are able to incontext learn-perform a new task via inference alone by conditioning on a few input-label pairs (demonstrations) and making predictions for new inputs.However, there has been little understanding of how the model learns and which aspects of the demonstrations contribute to end task performance.In this paper, we show that ground truth demonstrations are in fact not required-randomly replacing labels in the demonstrations barely hurts performance on a range of classification and multi-choce tasks, consistently over 12 different models including GPT-3.Instead, we find that other aspects of the demonstrations are the key drivers of end task performance, including the fact that they provide a few examples of (1) the label space, (2) the distribution of the input text, and (3) the overall format of the sequence.Together, our analysis provides a new way of understanding how and why in-context learning works, while opening up new questions about how much can be learned from large language models through inference alone. Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, Luke Zettlemoyer |
EMNLP | 5 |
| 2022 | Improving Passage Retrieval with Zero-Shot Question GenerationabstractWe propose a simple and effective re-ranking method for improving passage retrieval in open question answering.The re-ranker re-scores retrieved passages with a zero-shot question generation model, which uses a pre-trained language model to compute the probability of the input question conditioned on a retrieved passage.This approach can be applied on top of any retrieval method (e.g.neural or keywordbased), does not require any domain-or taskspecific training (and therefore is expected to generalize better to data distribution shifts), and provides rich cross-attention between query and passage (i.e. it must explain every token in the question).When evaluated on a number of open-domain retrieval datasets, our re-ranker improves strong unsupervised retrieval models by 6%-18% absolute and strong supervised models by up to 12% in terms of top-20 passage retrieval accuracy.We also obtain new stateof-the-art results on full open-domain question answering by simply adding the new re-ranker to existing models with no further changes.1 Devendra Singh Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan, Scott Yih, Joelle Pineau, Luke Zettlemoyer |
EMNLP | 2 |
| 2022 | HTLM: Hyper-Text Pre-Training and Prompting of Language Models
Armen Aghajanyan, Dmytro Okhonko, Mike Lewis, Mandar Joshi, Hu Xu 0001, Gargi Ghosh, Luke Zettlemoyer |
ICLR | 3 |
| 2022 | 8-bit Optimizers via Block-wise Quantization
Tim Dettmers, Mike Lewis, Sam Shleifer, Luke Zettlemoyer |
ICLR | 2 |
| 2022 | Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Ofir Press, Noah A. Smith, Mike Lewis |
ICLR | 3 |
| 2022 | Tricks for Training Sparse Translation ModelsabstractDheeru Dua, Shruti Bhosale, Vedanuj Goswami, James Cross, Mike Lewis, Angela Fan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Dheeru Dua, Shruti Bhosale, Vedanuj Goswami, James Cross 0003, Mike Lewis, Angela Fan |
NAACL-HLT | 5 |
| 2022 | DEMix Layers: Disentangling Domains for Modular Language ModelingabstractSuchin Gururangan, Mike Lewis, Ari Holtzman, Noah Smith, Luke Zettlemoyer. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A. Smith, Luke Zettlemoyer |
NAACL-HLT | 2 |
| 2022 | MetaICL: Learning to Learn In ContextabstractSewon Min, Mike Lewis, Luke Zettlemoyer, Hannaneh Hajishirzi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Sewon Min, Mike Lewis, Luke Zettlemoyer, Hannaneh Hajishirzi |
NAACL-HLT | 2 |
| 2022 | Sparse Distillation: Speeding Up Text Classification by Using Bigger Student ModelsabstractQinyuan Ye, Madian Khabsa, Mike Lewis, Sinong Wang, Xiang Ren, Aaron Jaech. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Qinyuan Ye, Madian Khabsa, Mike Lewis, Sinong Wang, Xiang Ren 0001, Aaron Jaech |
NAACL-HLT | 3 |
| 2022 | GPT3.int8(): 8-bit Matrix Multiplication for Transformers at ScaleabstractLarge language models have been widely adopted but require significant GPU memory for inference. We develop a procedure for Int8 matrix multiplication for feed-forward and attention projection layers in transformers, which cut the memory needed for inference by half while retaining full precision performance. With our method, a 175B parameter 16/32-bit checkpoint can be loaded, converted to Int8, and used immediately without performance degradation. This is made possible by understanding and working around properties of highly systematic emergent features in transformer language models that dominate attention and transformer predictive performance. To cope with these features, we develop a two-part quantization procedure, {\bf LLM.int8()}. We first use vector-wise quantization with separate normalization constants for each inner product in the matrix multiplication, to quantize most of the features. However, for the emergent outliers, we also include a new mixed-precision decomposition scheme, which isolates the outlier feature dimensions into a 16-bit matrix multiplication while still more than 99.9\% of values are multiplied in 8-bit. Using LLM.int8(), we show empirically it is possible to perform inference in LLMs with up to 175B parameters without any performance degradation. This result makes such models much more accessible, for example making it possible to use OPT-175B/BLOOM on a single server with consumer GPUs. We open source our software. Tim Dettmers, Mike Lewis, Younes Belkada, Luke Zettlemoyer |
NeurIPS | 2 |
| 2021 | Shortformer: Better Language Modeling using Shorter InputsabstractOfir Press, Noah A. Smith, Mike Lewis. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ofir Press, Noah A. Smith, Mike Lewis |
ACL/IJCNLP (1) | 3 |
| 2021 | Joint Verification and Reranking for Open Fact Checking Over TablesabstractMichael Sejr Schlichtkrull, Vladimir Karpukhin, Barlas Oguz, Mike Lewis, Wen-tau Yih, Sebastian Riedel. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Michael Sejr Schlichtkrull, Vladimir Karpukhin, Barlas Oguz, Mike Lewis, Scott Yih, Sebastian Riedel 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | Nearest Neighbor Machine Translation
Urvashi Khandelwal, Angela Fan, Daniel Jurafsky, Luke Zettlemoyer, Mike Lewis |
ICLR | 5 |
| 2021 | BASE Layers: Simplifying Training of Large, Sparse ModelsabstractWe introduce a new balanced assignment of experts (BASE) layer for large language models that greatly simplifies existing high capacity sparse layers. Sparse layers can dramatically improve the efficiency of training and inference by routing each token to specialized expert modules that contain only a small fraction of the model parameters. However, it can be difficult to learn balanced routing functions that make full use of the available experts; existing approaches typically use routing heuristics or auxiliary expert-balancing loss functions. In contrast, we formulate token-to-expert allocation as a linear assignment problem, allowing an optimal assignment in which each expert receives an equal number of tokens. This optimal assignment scheme improves efficiency by guaranteeing balanced compute loads, and also simplifies training by not requiring any new hyperparameters or auxiliary losses. Code is publicly released. Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal 0001, Luke Zettlemoyer |
ICML | 1 |
| 2021 | Multitasking Inhibits Semantic DriftabstractWhen intelligent agents communicate to accomplish shared goals, how do these goals shape the agents' language?We study the dynamics of learning in latent language policies (LLPs), in which instructor agents generate natural-language subgoal descriptions and executor agents map these descriptions to lowlevel actions.LLPs can solve challenging long-horizon reinforcement learning problems and provide a rich model for studying taskoriented language use.But previous work has found that LLP training is prone to semantic drift (use of messages in ways inconsistent with their original natural language meanings).Here, we demonstrate theoretically and empirically that multitask training is an effective counter to this problem: we prove that multitask training eliminates semantic drift in a well-studied family of signaling games, and show that multitask training of neural LLPs in a complex strategy game reduces drift and while improving sample efficiency. Athul Paul Jacob, Mike Lewis, Jacob Andreas |
NAACL-HLT | 2 |
| 2020 | BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionabstractMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, Luke Zettlemoyer. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Mike Lewis, Yinhan Liu, Naman Goyal 0001, Marjan Ghazvininejad, Abdel-rahman Mohamed, Omer Levy, Veselin Stoyanov, Luke Zettlemoyer |
ACL | 1 |
| 2020 | Asking and Answering Questions to Evaluate the Factual Consistency of SummariesabstractPractical applications of abstractive summarization models are limited by frequent factual inconsistencies with respect to their input.Existing automatic evaluation metrics for summarization are largely insensitive to such errors.We propose QAGS, 1 an automatic evaluation protocol that is designed to identify factual inconsistencies in a generated summary.QAGS is based on the intuition that if we ask questions about a summary and its source, we will receive similar answers if the summary is factually consistent with the source.To evaluate QAGS, we collect human judgments of factual consistency on model-generated summaries for the CNN/DailyMail (Hermann et al., 2015) and XSUM (Narayan et al., 2018) summarization datasets.QAGS has substantially higher correlations with these judgments than other automatic evaluation metrics.Also, QAGS offers a natural form of interpretability: The answers and questions generated while computing QAGS indicate which tokens of a summary are inconsistent and why.We believe QAGS is a promising tool in automatically generating usable and factually consistent text.Code for QAGS will be available at https://github. com/W4ngatang/qags.Article: On Friday, 28-year-old Usman Khan stabbed reportedly several people at Fishmongers' Hall in London with a large knife, then fled up London Bridge.Members of the public confronted him; one man sprayed Khan with a fire extinguisher, others struck him with their fists and took his knife, and another, a Polish chef named ukasz, harried him with a five-foot narwhal tusk.[. . .] Summary : On Friday afternoon , a man named Faisal Khan entered a Cambridge University building and started attacking people with a knife and a fire extinguisher .Question 1: What did the attacker have ?Article answer: a large knife Summary answer: a knife and a fire extinguisher Question 2: When did the attack take place ? Kyunghyun Cho, Mike Lewis |
ACL | 3 |
| 2020 | Conversational Semantic ParsingabstractArmen Aghajanyan, Jean Maillard, Akshat Shrivastava, Keith Diedrick, Michael Haeger, Haoran Li, Yashar Mehdad, Veselin Stoyanov, Anuj Kumar, Mike Lewis, Sonal Gupta. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Armen Aghajanyan, Jean Maillard, Akshat Shrivastava, Keith Diedrick, Michael Haeger, Haoran Li 0007, Yashar Mehdad, Veselin Stoyanov, Mike Lewis, Sonal Gupta |
EMNLP (1) | 10 |
| 2020 | Grounded Adaptation for Zero-shot Executable Semantic ParsingabstractWe propose Grounded Adaptation for Zeroshot Executable Semantic Parsing (GAZP) to adapt an existing semantic parser to new environments (e.g.new database schemas).GAZP combines a forward semantic parser with a backward utterance generator to synthesize data (e.g.utterances and SQL queries) in the new environment, then selects cycleconsistent examples to adapt the parser.Unlike data-augmentation, which typically synthesizes unverified examples in the training environment, GAZP synthesizes examples in the new environment whose inputoutput consistency are verified.On the Spider, Sparc, and CoSQL zero-shot semantic parsing tasks, GAZP improves logical form and execution accuracy of the baseline parser.Our analyses show that GAZP outperforms dataaugmentation in the training environment, performance increases with the amount of GAZPsynthesized data, and cycle-consistency is central to successful adaptation. Victor Zhong, Mike Lewis, Sida I. Wang, Luke Zettlemoyer |
EMNLP (1) | 2 |
| 2020 | Generalization through Memorization: Nearest Neighbor Language Models
Urvashi Khandelwal, Omer Levy, Daniel Jurafsky, Luke Zettlemoyer, Mike Lewis |
ICLR | 5 |
| 2020 | Pre-training via ParaphrasingabstractWe introduce MARGE, a pre-trained sequence-to-sequence model learned with an unsupervised multi-lingual multi-document paraphrasing objective. MARGE provides an alternative to the dominant masked language modeling paradigm, where we self-supervise the \emph{reconstruction} of target text by \emph{retrieving} a set of related texts (in many languages) and conditioning on them to maximize the likelihood of generating the original. We show it is possible to jointly learn to do retrieval and reconstruction, given only a random initialization. The objective noisily captures aspects of paraphrase, translation, multi-document summarization, and information retrieval, allowing for strong zero-shot performance on several tasks. For example, with no additional task-specific training we achieve BLEU scores of up to 35.8 for document translation. We further show that fine-tuning gives strong performance on a range of discriminative and generative tasks in many languages, making MARGE the most generally applicable pre-training method to date. Mike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan, Sida I. Wang, Luke Zettlemoyer |
NeurIPS | 1 |
| 2020 | Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksabstractLarge pre-trained language models have been shown to store factual knowledge in their parameters, and achieve state-of-the-art results when fine-tuned on downstream NLP tasks. However, their ability to access and precisely manipulate knowledge is still limited, and hence on knowledge-intensive tasks, their performance lags behind task-specific architectures. Additionally, providing provenance for their decisions and updating their world knowledge remain open research problems. Pre-trained models with a differentiable access mechanism to explicit non-parametric memory can overcome this issue, but have so far been only investigated for extractive downstream tasks. We explore a general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation. We introduce RAG models where the parametric memory is a pre-trained seq2seq model and the non-parametric memory is a dense vector index of Wikipedia, accessed with a pre-trained neural retriever. We compare two RAG formulations, one which conditions on the same retrieved passages across the whole generated sequence, the other can use different passages per token. We fine-tune and evaluate our models on a wide range of knowledge-intensive NLP tasks and set the state-of-the-art on three open domain QA tasks, outperforming parametric seq2seq models and task-specific retrieve-and-extract architectures. For language generation tasks, we find that RAG models generate more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline. Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal 0001, Heinrich Küttler, Mike Lewis, Scott Yih, Tim Rocktäschel, Sebastian Riedel 0001, Douwe Kiela |
NeurIPS | 8 |
| 2020 | Multilingual Denoising Pre-training for Neural Machine TranslationabstractThis paper demonstrates that multilingual denoising pre-training produces significant performance gains across a wide variety of machine translation (MT) tasks. We present mBART—a sequence-to-sequence denoising auto-encoder pre-trained on large-scale monolingual corpora in many languages using the BART objective (Lewis et al., 2019 ). mBART is the first method for pre-training a complete sequence-to-sequence model by denoising full texts in multiple languages, whereas previous approaches have focused only on the encoder, decoder, or reconstructing parts of the text. Pre-training a complete model allows it to be directly fine-tuned for supervised (both sentence-level and document-level) and unsupervised machine translation, with no task- specific modifications. We demonstrate that adding mBART initialization produces performance gains in all but the highest-resource settings, including up to 12 BLEU points for low resource MT and over 5 BLEU points for many document-level and unsupervised models. We also show that it enables transfer to language pairs with no bi-text or that were not in the pre-training corpus, and present extensive analysis of which factors contribute the most to effective pre-training. 1 Yinhan Liu, Jiatao Gu, Naman Goyal 0001, Xian Li 0003, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, Luke Zettlemoyer |
Trans. Assoc. Comput. Linguistics | 7 |
| 2019 | Strategies for Structuring Story GenerationabstractWriters often rely on plans or sketches to write long stories, but most current language models generate word by word from left to right. We explore coarse-to-fine models for creating narrative texts of several hundred words, and introduce new models which decompose stories by abstracting over actions and entities. The model first generates the predicate-argument structure of the text, where different mentions of the same entity are marked with placeholder tokens. It then generates a surface realization of the predicate-argument structure, and finally replaces the entity placeholders with context-sensitive names and references. Human judges prefer the stories from our models to a wide range of previous approaches to hierarchical text generation. Extensive analysis shows that our methods can help improve the diversity and coherence of events and entities in generated stories. Angela Fan, Mike Lewis, Yann N. Dauphin |
ACL (1) | 2 |
| 2019 | Span-based Hierarchical Semantic Parsing for Task-Oriented DialogabstractPanupong Pasupat, Sonal Gupta, Karishma Mandyam, Rushin Shah, Mike Lewis, Luke Zettlemoyer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Panupong Pasupat, Sonal Gupta, Karishma Mandyam, Rushin Shah, Mike Lewis, Luke Zettlemoyer |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Generative Question Answering: Learning to Answer the Whole Question
Mike Lewis, Angela Fan |
ICLR (Poster) | 1 |
| 2019 | Hierarchical Decision Making by Generating and Following Natural Language InstructionsabstractWe explore using latent natural language instructions as an expressive and compositional representation of complex actions for hierarchical decision making. Rather than directly selecting micro-actions, our agent first generates a latent plan in natural language, which is then executed by a separate model. We introduce a challenging real-time strategy game environment in which the actions of a large number of units must be coordinated across long time scales. We gather a dataset of 76 thousand pairs of instructions and executions from human play, and train instructor and executor models. Experiments show that models using natural language as a latent variable significantly outperform models that directly imitate human actions. The compositional structure of language proves crucial to its effectiveness for action representation. We also release our code, models and data. Hengyuan Hu, Denis Yarats, Qucheng Gong, Yuandong Tian, Mike Lewis |
NeurIPS | 5 |
| 2018 | Hierarchical Neural Story GenerationabstractWe explore story generation: creative systems that can build coherent and fluent passages of text about a topic. We collect a large dataset of 300K human-written stories paired with writing prompts from an online forum. Our dataset enables hierarchical story generation, where the model first generates a premise, and then transforms it into a passage of text. We gain further improvements with a novel form of model fusion that improves the relevance of the story to the prompt, and adding a new gated multi-scale self-attention mechanism to model long-range context. Experiments show large improvements over strong baselines on both automated and human evaluations. Human judges prefer stories generated by our approach to those from a strong non-hierarchical model by a factor of two to one. Angela Fan, Mike Lewis, Yann N. Dauphin |
ACL (1) | 2 |
| 2018 | A Dataset for Telling the Stories of Social Media VideosabstractVideo content on social media platforms constitutes a major part of the communication between people, as it allows everyone to share their stories.However, if someone is unable to consume video, either due to a disability or network bandwidth, this severely limits their participation and communication.Automatically telling the stories using multi-sentence descriptions of videos would allow bridging this gap.To learn and evaluate such models, we introduce VideoStory, a new large-scale dataset for video description as a new challenge for multisentence video description.Our VideoStory captions dataset is complementary to prior work and contains 20k videos posted publicly on a social media platform amounting to 396 hours of video with 123k sentences, temporally aligned to the video.* *Work done while SG was intern at Facebook AI Research. Spandana Gella, Mike Lewis, Marcus Rohrbach |
EMNLP | 2 |
| 2018 | Neural Compositional Denotational Semantics for Question AnsweringabstractAnswering compositional questions requiring multi-step reasoning is challenging.We introduce an end-to-end differentiable model for interpreting questions about a knowledge graph (KG), which is inspired by formal approaches to semantics.Each span of text is represented by a denotation in a KG and a vector that captures ungrounded aspects of meaning.Learned composition modules recursively combine constituent spans, culminating in a grounding for the complete sentence which answers the question.For example, to interpret "not green", the model represents "green" as a set of KG entities and "not" as a trainable ungrounded vector-and then uses this vector to parameterize a composition function that performs a complement operation.For each sentence, we build a parse chart subsuming all possible parses, allowing the model to jointly learn both the composition operators and output structure by gradient descent from endtask supervision.The model learns a variety of challenging semantic operators, such as quantifiers, disjunctions and composed relations, and infers latent syntactic structure.It also generalizes well to longer questions than seen in its training data, in contrast to RNN, its treebased variants, and semantic parsing baselines. Nitish Gupta, Mike Lewis |
EMNLP | 2 |
| 2018 | Semantic Parsing for Task Oriented Dialog using Hierarchical RepresentationsabstractTask oriented dialog systems typically first parse user utterances to semantic frames comprised of intents and slots.Previous work on task oriented intent and slot-filling work has been restricted to one intent per query and one slot label per token, and thus cannot model complex compositional requests.Alternative semantic parsing systems have represented queries as logical forms, but these are challenging to annotate and parse.We propose a hierarchical annotation scheme for semantic parsing that allows the representation of compositional queries, and can be efficiently and accurately parsed by standard constituency parsing models.We release a dataset of 44k annotated queries 1 , and show that parsing models outperform sequence-to-sequence approaches on this dataset. Sonal Gupta, Rushin Shah, Mrinal Mohit, Mike Lewis |
EMNLP | 5 |
| 2018 | Hierarchical Text Generation and Planning for Strategic DialogueabstractEnd-to-end models for goal-orientated dialogue are challenging to train, because linguistic and strategic aspects are entangled in latent state vectors. We introduce an approach to learning representations of messages in dialogues by maximizing the likelihood of subsequent sentences and actions, which decouples the semantics of the dialogue utterance from its linguistic realization. We then use these latent sentence representations for hierarchical language generation, planning and reinforcement learning. Experiments show that our approach increases the end-task reward achieved by the model, improves the effectiveness of long-term planning using rollouts, and allows self-play reinforcement learning to improve decision making without diverging from human language. Our hierarchical latent-variable model outperforms previous work both linguistically and strategically. Denis Yarats, Mike Lewis |
ICML | 2 |
| 2017 | Deep Semantic Role Labeling: What Works and What's NextabstractWe introduce a new deep learning model for semantic role labeling (SRL) that significantly improves the state of the art, along with detailed analyses to reveal its strengths and limitations.We use a deep highway BiLSTM architecture with constrained decoding, while observing a number of recent best practices for initialization and regularization.Our 8-layer ensemble model achieves 83.2 F1 on the CoNLL 2005 test set and 83.4 F1 on CoNLL 2012, roughly a 10% relative error reduction over the previous state of the art.Extensive empirical analysis of these gains show that (1) deep models excel at recovering long-distance dependencies but can still make surprisingly obvious errors, and (2) that there is still room for syntactic parsers to improve these results. Luheng He, Kenton Lee, Mike Lewis, Luke Zettlemoyer |
ACL (1) | 3 |
| 2017 | End-to-end Neural Coreference ResolutionabstractWe introduce the first end-to-end coreference resolution model and show that it significantly outperforms all previous work without using a syntactic parser or handengineered mention detector.The key idea is to directly consider all spans in a document as potential mentions and learn distributions over possible antecedents for each.The model computes span embeddings that combine context-dependent boundary representations with a headfinding attention mechanism.It is trained to maximize the marginal likelihood of gold antecedent spans from coreference clusters and is factored to enable aggressive pruning of potential mentions.Experiments demonstrate state-of-the-art performance, with a gain of 1.5 F1 on the OntoNotes benchmark and by 3.1 F1 using a 5-model ensemble, despite the fact that this is the first approach to be successfully trained with no external resources. Kenton Lee, Luheng He, Mike Lewis, Luke Zettlemoyer |
EMNLP | 3 |
| 2017 | Deal or No Deal? End-to-End Learning of Negotiation DialoguesabstractMuch of human dialogue occurs in semicooperative settings, where agents with different goals attempt to agree on common decisions.Negotiations require complex communication and reasoning skills, but success is easy to measure, making this an interesting task for AI.We gather a large dataset of human-human negotiations on a multi-issue bargaining task, where agents who cannot observe each other's reward functions must reach an agreement (or a deal) via natural language dialogue.For the first time, we show it is possible to train end-to-end models for negotiation, which must learn both linguistic and reasoning skills with no annotated dialogue states.We also introduce dialogue rollouts, in which the model plans ahead by simulating possible complete continuations of the conversation, and find that this technique dramatically improves performance.Our code and dataset are publicly available.1 Mike Lewis, Denis Yarats, Yann N. Dauphin, Devi Parikh, Dhruv Batra |
EMNLP | 1 |
| 2016 | Human-in-the-Loop ParsingabstractThis paper demonstrates that it is possible for a parser to improve its performance with a human in the loop, by posing simple questions to non-experts.For example, given the first sentence of this abstract, if the parser is uncertain about the subject of the verb "pose," it could generate the question What would pose something?with candidate answers this paper and a parser.Any fluent speaker can answer this question, and the correct answer resolves the original uncertainty.We apply the approach to a CCG parser, converting uncertain attachment decisions into natural language questions about the arguments of verbs.Experiments show that crowd workers can answer these questions quickly, accurately and cheaply.Our human-in-the-loop parser improves on the state of the art with less than 2 questions per sentence on average, with a gain of 1.7 F1 on the 10% of sentences whose parses are changed. Luheng He, Julian Michael, Mike Lewis, Luke Zettlemoyer |
EMNLP | 3 |
| 2016 | Global Neural CCG Parsing with Optimality GuaranteesabstractWe introduce the first global recursive neural parsing model with optimality guarantees during decoding. To support global features, we give up dynamic programs and instead search directly in the space of all possible subtrees. Although this space is exponentially large in the sentence length, we show it is possible to learn an efficient A* parser. We augment existing parsing models, which have informative bounds on the outside score, with a global model that has loose bounds but only needs to model non-local phenomena. The global model is trained with a new objective that encourages the parser to explore a tiny fraction of the search space. The approach is applied to CCG parsing, improving state-of-the-art accuracy by 0.4 F1. The parser finds the optimal parse for 99.9% of held-out sentences, exploring on average only 190 subtrees. Kenton Lee, Mike Lewis, Luke Zettlemoyer |
EMNLP | 2 |
| 2016 | LSTM CCG ParsingabstractWe demonstrate that a state-of-the-art parser can be built using only a lexical tagging model and a deterministic grammar, with no explicit model of bi-lexical dependencies.Instead, all dependencies are implicitly encoded in an LSTM supertagger that assigns CCG lexical categories.The parser significantly outperforms all previously published CCG results, supports efficient and optimal A * decoding, and benefits substantially from semisupervised tri-training.We give a detailed analysis, demonstrating that the parser can recover long-range dependencies with high accuracy and that the semi-supervised learning enables significant accuracy gains.By running the LSTM on a GPU, we are able to parse over 2600 sentences per second while improving state-of-the-art accuracy by 1.1 F1 in domain and up to 4.5 F1 out of domain. Mike Lewis, Kenton Lee, Luke Zettlemoyer |
HLT-NAACL | 1 |
| 2015 | Question-Answer Driven Semantic Role Labeling: Using Natural Language to Annotate Natural LanguageabstractThis paper introduces the task of questionanswer driven semantic role labeling (QA-SRL), where question-answer pairs are used to represent predicate-argument structure.For example, the verb "introduce" in the previous sentence would be labeled with the questions "What is introduced?", and "What introduces something?", each paired with the phrase from the sentence that gives the correct answer.Posing the problem this way allows the questions themselves to define the set of possible roles, without the need for predefined frame or thematic role ontologies.It also allows for scalable data collection by annotators with very little training and no linguistic expertise.We gather data in two domains, newswire text and Wikipedia articles, and introduce simple classifierbased models for predicting which questions to ask and what their answers should be.Our results show that non-expert annotators can produce high quality QA-SRL data, and also establish baseline performance levels for future work on this task. Luheng He, Mike Lewis, Luke Zettlemoyer |
EMNLP | 2 |
| 2015 | Joint A* CCG Parsing and Semantic Role LabellingabstractJoint models of syntactic and semantic parsing have the potential to improve performance on both tasks-but to date, the best results have been achieved with pipelines.We introduce a joint model using CCG, which is motivated by the close link between CCG syntax and semantics.Semantic roles are recovered by labelling the deep dependency structures produced by the grammar.Furthermore, because CCG is lexicalized, we show it is possible to factor the parsing model over words and introduce a new A * parsing algorithmwhich we demonstrate is faster and more accurate than adaptive supertagging.Our joint model is the first to substantially improve both syntactic and semantic accuracy over a comparable pipeline, and also achieves state-of-the-art results for a nonensemble semantic role labelling model. Mike Lewis, Luheng He, Luke Zettlemoyer |
EMNLP | 1 |
| 2014 | A* CCG Parsing with a Supertag-factored ModelabstractWe introduce a new CCG parsing model which is factored on lexical category assignments.Parsing is then simply a deterministic search for the most probable category sequence that supports a CCG derivation.The parser is extremely simple, with a tiny feature set, no POS tagger, and no statistical model of the derivation or dependencies.Formulating the model in this way allows a highly effective heuristic for A * parsing, which makes parsing extremely fast.Compared to the standard C&C CCG parser, our model is more accurate out-of-domain, is four times faster, has higher coverage, and is greatly simplified.We also show that using our parser improves the performance of a state-ofthe-art question answering system. 1 Mike Lewis, Mark Steedman |
EMNLP | 1 |
| 2014 | Extracting common sense knowledge from text for robot planningabstractAutonomous robots often require domain knowledge to act intelligently in their environment. This is particularly true for robots that use automated planning techniques, which require symbolic representations of the operating environment and the robot's capabilities. However, the task of specifying domain knowledge by hand is tedious and prone to error. As a result, we aim to automate the process of acquiring general common sense knowledge of objects, relations, and actions, by extracting such information from large amounts of natural language text, written by humans for human readers. We present two methods for knowledge acquisition, requiring only limited human input, which focus on the inference of spatial relations from text. Although our approach is applicable to a range of domains and information, we only consider one type of knowledge here, namely object locations in a kitchen environment. As a proof of concept, we test our approach using an automated planner and show how the addition of common sense knowledge can improve the quality of the generated plans. Peter Kaiser 0001, Mike Lewis, Ronald P. A. Petrick, Tamim Asfour, Mark Steedman |
ICRA | 2 |
| 2014 | Improved CCG Parsing with Semi-supervised SupertaggingabstractCurrent supervised parsers are limited by the size of their labelled training data, making improving them with unlabelled data an important goal. We show how a state-of-the-art CCG parser can be enhanced, by predicting lexical categories using unsupervised vector-space embeddings of words. The use of word embeddings enables our model to better generalize from the labelled data, and allows us to accurately assign lexical categories without depending on a POS-tagger. Our approach leads to substantial improvements in dependency parsing results over the standard supervised CCG parser when evaluated on Wall Street Journal (0.8%), Wikipedia (1.8%) and biomedical (3.4%) text. We compare the performance of two recently proposed approaches for classification using a wide variety of word embeddings. We also give a detailed error analysis demonstrating where using embeddings outperforms traditional feature sets, and showing how including POS features can decrease accuracy. Mike Lewis, Mark Steedman |
Trans. Assoc. Comput. Linguistics | 1 |
| 2013 | Unsupervised Induction of Cross-Lingual Semantic RelationsabstractCreating a language-independent meaning representation would benefit many crosslingual NLP tasks.We introduce the first unsupervised approach to this problem, learning clusters of semantically equivalent English and French relations between referring expressions, based on their named-entity arguments in large monolingual corpora.The clusters can be used as language-independent semantic relations, by mapping clustered expressions in different languages onto the same relation.Our approach needs no parallel text for training, but outperforms a baseline that uses machine translation on a cross-lingual question answering task.We also show how to use the semantics to improve the accuracy of machine translation, by using it in a simple reranker. Mike Lewis, Mark Steedman |
EMNLP | 1 |
| 2013 | Combined Distributional and Logical SemanticsabstractWe introduce a new approach to semantics which combines the benefits of distributional and formal logical semantics. Distributional models have been successful in modelling the meanings of content words, but logical semantics is necessary to adequately represent many function words. We follow formal semantics in mapping language to logical representations, but differ in that the relational constants used are induced by offline distributional clustering at the level of predicate-argument structure. Our clustering algorithm is highly scalable, allowing us to run on corpora the size of Gigaword. Different senses of a word are disambiguated based on their induced types. We outperform a variety of existing approaches on a wide-coverage question answering task, and demonstrate the ability to make complex multi-sentence inferences involving quantifiers on the FraCaS suite. Mike Lewis, Mark Steedman |
Trans. Assoc. Comput. Linguistics | 1 |