EDBT 2026 Demo / reviewers in the wild / expert
Omer Levy
dblp:117/4866
· DBLP profile ↗
50ranked-venue papers
11as first author
22since 2021 · last 2025
0000-0001-7300-8191ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 9 first-author · 21 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal ModelabstractWe introduce Transfusion, a recipe for training a multi-modal model over discrete and continuous data.
Transfusion combines the language modeling loss function (next token prediction) with diffusion to train a single transformer over mixed-modality sequences.
We pretrain multiple Transfusion models up to 7B parameters from scratch on a mixture of text and image data, establishing scaling laws with respect to a variety of uni- and cross-modal benchmarks.
Our experiments show that Transfusion scales significantly better than quantizing images and training a language model over discrete image tokens.
By introducing modality-specific encoding and decoding layers, we can further improve the performance of Transfusion models, and even compress each image to just 16 patches.
We further demonstrate that scaling our Transfusion recipe to 7B parameters and 2T multi-modal tokens produces a model that can generate images and text on a par with similar scale diffusion models and language models, reaping the benefits of both worlds. Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, Omer Levy |
ICLR | 10 |
| 2024 | Altogether: Image Captioning via Re-aligning Alt-textabstractHu Xu, Po-Yao Huang, Xiaoqing Tan, Ching-Feng Yeh, Jacob Kahn, Christine Jou, Gargi Ghosh, Omer Levy, Luke Zettlemoyer, Wen-tau Yih, Shang-Wen Li, Saining Xie, Christoph Feichtenhofer. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Hu Xu 0001, Po-Yao Huang 0001, Xiaoqing Ellen Tan, Ching-Feng Yeh, Jacob Kahn, Christine Jou, Gargi Ghosh, Omer Levy, Luke Zettlemoyer, Scott Yih, Shang-Wen Li 0001, Saining Xie, Christoph Feichtenhofer |
EMNLP | 8 |
| 2024 | Self-Alignment with Instruction BacktranslationabstractWe present a scalable method to build a high quality instruction following language model by automatically labelling human-written text with corresponding instructions. Our approach, named instruction backtranslation, starts with a language model finetuned on a small amount of seed data, and a given web corpus. The seed model is used to construct training examples by generating instruction prompts for web documents (self-augmentation), and then selecting high quality examples from among these candidates (self-curation). This data is then used to finetune a stronger model. Finetuning LLaMa on two iterations of our approach yields a model that outperforms all other LLaMa-based models on the Alpaca leaderboard not relying on distillation data, demonstrating highly effective self-alignment. Xian Li 0003, Chunting Zhou, Timo Schick, Omer Levy, Luke Zettlemoyer, Jason Weston, Mike Lewis |
ICLR | 5 |
| 2024 | Branch-Solve-Merge Improves Large Language Model Evaluation and GenerationabstractSwarnadeep Saha, Omer Levy, Asli Celikyilmaz, Mohit Bansal, Jason Weston, Xian Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Swarnadeep Saha, Omer Levy, Asli Celikyilmaz, Mohit Bansal, Jason Weston, Xian Li 0003 |
NAACL-HLT | 2 |
| 2024 | Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context LengthabstractThe quadratic complexity and weak length extrapolation of Transformers limits their ability to scale to long sequences, and while sub-quadratic solutions like linear attention and state space models exist, they empirically underperform Transformers in pretraining efficiency and downstream task accuracy. We introduce MEGALODON, an neural architecture for efficient sequence modeling with unlimited context length. MEGALODON inherits the architecture of MEGA (exponential moving average with gated attention), and further introduces multiple technical components to improve its capability and stability, including complex exponential moving average (CEMA), timestep normalization layer, normalized attention mechanism and pre-norm with two-hop residual configuration. In a controlled head-to-head comparison with LLAMA2, MEGALODON achieves better efficiency than Transformer in the scale of 7 billion parameters and 2 trillion training tokens. MEGALODON reaches a training loss of 1.70, landing mid-way between LLAMA2-7B (1.75) and LLAMA2-13B (1.67). This result is robust throughout a wide range of benchmarks, where MEGALODON consistently outperforms Transformers across different tasks, domains, and modalities. Xuezhe Ma, Wenhan Xiong, Beidi Chen, Lili Yu, Hao Zhang 0025, Jonathan May, Luke Zettlemoyer, Omer Levy, Chunting Zhou |
NeurIPS | 9 |
| 2023 | Causes and Cures for Interference in Multilingual TranslationabstractMultilingual machine translation models can benefit from synergy between different language pairs, but also suffer from interference.While there is a growing number of sophisticated methods that aim to eliminate interference, our understanding of interference as a phenomenon is still limited.This work identifies the main factors that contribute to interference in multilingual machine translation.Through systematic experimentation, we find that interference (or synergy) are primarily determined by model size, data size, and the proportion of each language pair within the total dataset.We observe that substantial interference occurs mainly when the model is very small with respect to the available training data, and that using standard transformer configurations with less than one billion parameters largely alleviates interference and promotes synergy.Moreover, we show that tuning the sampling temperature to control the proportion of each language pair in the data is key to balancing the amount of interference between low and high resource language pairs effectively, and can lead to superior performance overall. Uri Shaham 0002, Maha Elbayad, Vedanuj Goswami, Omer Levy, Shruti Bhosale |
ACL (1) | 4 |
| 2023 | Instruction Induction: From Few Examples to Natural Language Task DescriptionsabstractLarge language models are able to perform a task by conditioning on a few input-output demonstrations -a paradigm known as incontext learning.We show that language models can explicitly infer an underlying task from a few demonstrations by prompting them to generate a natural language instruction that fits the examples.To explore this ability, we introduce the instruction induction challenge, compile a dataset consisting of 24 tasks, and define a novel evaluation metric based on executing the generated instruction.We discover that, to a large extent, the ability to generate instructions does indeed emerge when using a model that is both large enough and aligned to follow instructions; InstructGPT achieves 65.7% of human performance in our execution-based metric, while the original GPT-3 model reaches only 9.8% of human performance.This surprising result suggests that instruction induction might be a viable learning paradigm in and of itself, where instead of fitting a set of latent continuous parameters to the data, one searches for the best description in the natural language hypothesis space. 1 Or Honovich, Uri Shaham 0002, Samuel R. Bowman, Omer Levy |
ACL (1) | 4 |
| 2023 | Unnatural Instructions: Tuning Language Models with (Almost) No Human LaborabstractInstruction tuning enables pretrained language models to perform new tasks from inferencetime natural language descriptions.These approaches rely on vast amounts of human supervision in the form of crowdsourced datasets or user interactions.In this work, we introduce Unnatural Instructions: a large dataset of creative and diverse instructions, collected with virtually no human labor.We collect 64,000 examples by prompting a language model with three seed examples of instructions and eliciting a fourth.This set is then expanded by prompting the model to rephrase each instruction, creating a total of approximately 240,000 examples of instructions, inputs, and outputs.Experiments show that despite containing a fair amount of noise, training on Unnatural Instructions rivals the effectiveness of training on open-source manually-curated datasets, surpassing the performance of models such as T0++ and Tk-Instruct across various benchmarks.These results demonstrate the potential of model-generated data as a cost-effective alternative to crowdsourcing for dataset expansion and diversification. Example 1Instruction: You are given a science question (easy-level) and four answer options (associated with "A", "B", "C", "D").Your task is to find the correct answer based on scientific facts, knowledge, and reasoning.Do not generate anything else apart from one of the following characters: 'A', 'B, 'C', 'D'.There is only one correct answer for each question. Input: Which part of a bicycle BEST moves in a circle? (A) Seat (B) Frame (C) Foot pedal (D) KickstandConstraints: The output should be one of the following characters: 'A', 'B, 'C', 'D'. Example 2Instruction: You are given a negative review and your task is to convert it to a positive review by one or more making minimal changes.Avoid changing the context of the review.Input: we stood there in shock, because we never expected this. Constraints: None.Example 3 Instruction: In this task, you are given two sentences taken from a conversation, and your job is to classify whether these given sentences are sequential or not.We will mark the given sentence pair as 'True' if it's sequential, otherwise 'False'.The two sentences are spoken by two different people. Or Honovich, Thomas Scialom, Omer Levy, Timo Schick |
ACL (1) | 3 |
| 2023 | Scaling Laws for Generative Mixed-Modal Language ModelsabstractGenerative language models define distributions over sequences of tokens that can represent essentially any combination of data modalities (e.g., any permutation of image tokens from VQ-VAEs, speech tokens from HuBERT, BPE tokens for language or code, and so on). To better understand the scaling properties of such mixed-modal models, we conducted over 250 experiments using seven different modalities and model sizes ranging from 8 million to 30 billion, trained on 5-100 billion tokens. We report new mixed-modal scaling laws that unify the contributions of individual modalities and the interactions between them. Specifically, we explicitly model the optimal synergy and competition due to data and model size as an additive term to previous uni-modal scaling laws. We also find four empirical phenomena observed during the training, such as emergent coordinate-ascent style training that naturally alternates between modalities, guidelines for selecting critical hyper-parameters, and connections between mixed-modal competition and training stability. Finally, we test our scaling law by training a 30B speech-text model, which significantly outperforms the corresponding unimodal models. Overall, our research provides valuable insights into the design and training of mixed-modal generative models, an important new class of unified models that have unique distributional properties. Armen Aghajanyan, Lili Yu, Alexis Conneau, Wei-Ning Hsu, Karen Hambardzumyan, Stephen Roller, Naman Goyal 0001, Omer Levy, Luke Zettlemoyer |
ICML | 9 |
| 2023 | Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image GenerationabstractThe ability to collect a large dataset of human preferences from text-to-image users is usually limited to companies, making such datasets inaccessible to the public. To address this issue, we create a web app that enables text-to-image users to generate images and specify their preferences. Using this web app we build Pick-a-Pic, a large, open dataset of text-to-image prompts and real users’ preferences over generated images. We leverage this dataset to train a CLIP-based scoring function, PickScore, which exhibits superhuman performance on the task of predicting human preferences. Then, we test PickScore’s ability to perform model evaluation and observe that it correlates better with human rankings than other automatic evaluation metrics. Therefore, we recommend using PickScore for evaluating future text-to-image generation models, and using Pick-a-Pic prompts as a more relevant dataset than MS-COCO. Finally, we demonstrate how PickScore can enhance existing text-to-image models via ranking. Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, Omer Levy |
NeurIPS | 6 |
| 2023 | LIMA: Less Is More for AlignmentabstractLarge language models are trained in two stages: (1) unsupervised pretraining from raw text, to learn general-purpose representations, and (2) large scale instruction tuning and reinforcement learning, to better align to end tasks and user preferences.
We measure the relative importance of these two stages by training LIMA, a 65B parameter LLaMa language model fine-tuned with the standard supervised loss on only 1,000 carefully curated prompts and responses, without any reinforcement learning or human preference modeling.
LIMA demonstrates remarkably strong performance, learning to follow specific response formats from only a handful of examples in the training data, including complex queries that range from planning trip itineraries to speculating about alternate history.
Moreover, the model tends to generalize well to unseen tasks that did not appear in the training data.
In a controlled human study, responses from LIMA are either equivalent or strictly preferred to GPT-4 in 43\% of cases; this statistic is as high as 58\% when compared to Bard and 65\% versus DaVinci003, which was trained with human feedback.
Taken together, these results strongly suggest that almost all knowledge in large language models is learned during pretraining, and only limited instruction tuning data is necessary to teach models to produce high quality output. Chunting Zhou, Puxin Xu, Srinivasan Iyer 0001, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Lili Yu, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, Omer Levy |
NeurIPS | 15 |
| 2022 | SCROLLS: Standardized CompaRison Over Long Language SequencesabstractUri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, Omer Levy. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Uri Shaham 0002, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta 0001, Wenhan Xiong, Mor Geva, Jonathan Berant, Omer Levy |
EMNLP | 11 |
| 2022 | Models In a Spelling Bee: Language Models Implicitly Learn the Character Composition of TokensabstractStandard pretrained language models operate on sequences of subword tokens without direct access to the characters that compose each token's string representation.We probe the embedding layer of pretrained language models and show that models learn the internal character composition of whole word and subword tokens to a surprising extent, without ever seeing the characters coupled with the tokens.Our results show that the embedding layers of RoBERTa and GPT2 each hold enough information to accurately spell up to a third of the vocabulary and reach high character ngram overlap across all token types.We further test whether enriching subword models with character information can improve language modeling, and observe that this method has a near-identical learning curve as training without spelling-based enrichment.Overall, our results suggest that language modeling objectives incentivize the model to implicitly learn some notion of spelling, and that explicitly teaching the model how to spell does not appear to enhance its performance on such tasks.1 Itay Itzhak, Omer Levy |
NAACL-HLT | 2 |
| 2022 | Learning to Retrieve Passages without SupervisionabstractOri Ram, Gal Shachaf, Omer Levy, Jonathan Berant, Amir Globerson. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ori Ram, Gal Shachaf, Omer Levy, Jonathan Berant, Amir Globerson |
NAACL-HLT | 3 |
| 2022 | Simple Local Attentions Remain Competitive for Long-Context TasksabstractWenhan Xiong, Barlas Oguz, Anchit Gupta, Xilun Chen, Diana Liskovich, Omer Levy, Scott Yih, Yashar Mehdad. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Wenhan Xiong, Barlas Oguz, Anchit Gupta, Xilun Chen 0002, Diana Liskovich, Omer Levy, Scott Yih, Yashar Mehdad |
NAACL-HLT | 6 |
| 2021 | Few-Shot Question Answering by Pretraining Span SelectionabstractOri Ram, Yuval Kirstain, Jonathan Berant, Amir Globerson, Omer Levy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ori Ram, Yuval Kirstain, Jonathan Berant, Amir Globerson, Omer Levy |
ACL/IJCNLP (1) | 5 |
| 2021 | Cryptonite: A Cryptic Crossword Benchmark for Extreme Ambiguity in LanguageabstractCurrent NLP datasets targeting ambiguity can be solved by a native speaker with relative ease.We present Cryptonite, a large-scale dataset based on cryptic crosswords, which is both linguistically complex and naturally sourced.Each example in Cryptonite is a cryptic clue, a short phrase or sentence with a misleading surface reading, whose solving requires disambiguating semantic, syntactic, and phonetic wordplays, as well as world knowledge.Cryptic clues pose a challenge even for experienced solvers, though top-tier experts can solve them with almost 100% accuracy.Cryptonite is a challenging task for current models; fine-tuning T5-Large on 470k cryptic clues achieves only 7.6% accuracy, on par with the accuracy of a rule-based clue solver (8.6%). Avia Efrat, Uri Shaham 0002, Dan Kilman, Omer Levy |
EMNLP (1) | 4 |
| 2021 | Transformer Feed-Forward Layers Are Key-Value MemoriesabstractFeed-forward layers constitute two-thirds of a transformer model's parameters, yet their role in the network remains under-explored.We show that feed-forward layers in transformerbased language models operate as key-value memories, where each key correlates with textual patterns in the training examples, and each value induces a distribution over the output vocabulary.Our experiments show that the learned patterns are human-interpretable, and that lower layers tend to capture shallow patterns, while upper layers learn more semantic ones.The values complement the keys' input patterns by inducing output distributions that concentrate probability mass on tokens likely to appear immediately after each pattern, particularly in the upper layers.Finally, we demonstrate that the output of a feed-forward layer is a composition of its memories, which is subsequently refined throughout the model's layers via residual connections to produce the final output distribution. Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy |
EMNLP (1) | 4 |
| 2021 | How to Train BERT with an Academic BudgetabstractWhile large language models à la BERT are used ubiquitously in NLP, pretraining them is considered a luxury that only a few wellfunded industry labs can afford.How can one train such models with a more modest budget?We present a recipe for pretraining a masked language model in 24 hours using a single lowend deep learning server.We demonstrate that through a combination of software optimizations, design choices, and hyperparameter tuning, it is possible to produce models that are competitive with BERT BASE on GLUE tasks at a fraction of the original pretraining cost. 1 Peter Izsak, Moshe Berchansky, Omer Levy |
EMNLP (1) | 3 |
| 2021 | Can Latent Alignments Improve Autoregressive Machine Translation?abstractLatent alignment objectives such as CTC and AXE significantly improve non-autoregressive machine translation models.Can they improve autoregressive models as well?We explore the possibility of training autoregressive machine translation models with latent alignment objectives, and observe that, in practice, this approach results in degenerate models.We provide a theoretical explanation for these empirical results, and prove that latent alignment objectives are incompatible with teacher forcing. Adi Haviv, Lior Vassertail, Omer Levy |
NAACL-HLT | 3 |
| 2021 | Neural Machine Translation without EmbeddingsabstractMany NLP models operate over sequences of subword tokens produced by hand-crafted tokenization rules and heuristic subword induction algorithms.A simple universal alternative is to represent every computerized text as a sequence of bytes via UTF-8, obviating the need for an embedding layer since there are fewer token types (256) than dimensions.Surprisingly, replacing the ubiquitous embedding layer with one-hot representations of each byte does not hurt performance; experiments on byte-to-byte machine translation from English to 10 different languages show a consistent improvement in BLEU, rivaling character-level and even standard subwordlevel models.A deeper investigation reveals that the combination of embeddingless models with decoder-input dropout amounts to token dropout, which benefits byte-to-byte models in particular. 1 Uri Shaham 0002, Omer Levy |
NAACL-HLT | 2 |
| 2021 | Understanding large-scale software systems - structure and flows
Omer Levy, Dror G. Feitelson |
Empir. Softw. Eng. | 1 |
| 2020 | BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionabstractMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, Luke Zettlemoyer. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Mike Lewis, Yinhan Liu, Naman Goyal 0001, Marjan Ghazvininejad, Abdel-rahman Mohamed, Omer Levy, Veselin Stoyanov, Luke Zettlemoyer |
ACL | 6 |
| 2020 | Improving Transformer Models by Reordering their SublayersabstractMultilayer transformer networks consist of interleaved self-attention and feedforward sublayers.Could ordering the sublayers in a different pattern lead to better performance?We generate randomly ordered transformers and train them with the language modeling objective.We observe that some of these models are able to achieve better performance than the interleaved baseline, and that those successful variants tend to have more self-attention at the bottom and more feedforward sublayers at the top.We propose a new transformer pattern that adheres to this property, the sandwich transformer, and show that it improves perplexity on multiple word-level and character-level language modeling benchmarks, at no cost in parameters, memory, or training time.However, the sandwich reordering pattern does not guarantee performance gains across every task, as we demonstrate on machine translation models.Instead, we suggest that further exploration of task-specific sublayer reorderings is needed in order to unlock additional gains. 1 Ofir Press, Noah A. Smith, Omer Levy |
ACL | 3 |
| 2020 | Generalization through Memorization: Nearest Neighbor Language Models
Urvashi Khandelwal, Omer Levy, Daniel Jurafsky, Luke Zettlemoyer, Mike Lewis |
ICLR | 2 |
| 2020 | Structural Language Models of CodeabstractWe address the problem of any-code completion - generating a missing piece of source code in a given program without any restriction on the vocabulary or structure. We introduce a new approach to any-code completion that leverages the strict syntax of programming languages to model a code snippet as a tree - structural language modeling (SLM). SLM estimates the probability of the program’s abstract syntax tree (AST) by decomposing it into a product of conditional probabilities over its nodes. We present a neural model that computes these conditional probabilities by considering all AST paths leading to a target node. Unlike previous techniques that have severely restricted the kinds of expressions that can be generated in this task, our approach can generate arbitrary code in any programming language. Our model significantly outperforms both seq2seq and a variety of structured approaches in generating Java and C# code. Our code, data, and trained models are available at http://github.com/tech-srl/slm-code-generation/. An online demo is available at http://AnyCodeGen.org. Uri Alon 0002, Roy Sadaka, Omer Levy, Eran Yahav |
ICML | 3 |
| 2020 | Aligned Cross Entropy for Non-Autoregressive Machine TranslationabstractNon-autoregressive machine translation models significantly speed up decoding by allowing for parallel prediction of the entire target sequence. However, modeling word order is more challenging due to the lack of autoregressive factors in the model. This difficultly is compounded during training with cross entropy loss, which can highly penalize small shifts in word order. In this paper, we propose aligned cross entropy (AXE) as an alternative loss function for training of non-autoregressive models. AXE uses a differentiable dynamic program to assign loss based on the best possible monotonic alignment between target tokens and model predictions. AXE-based training of conditional masked language models (CMLMs) substantially improves performance on major WMT benchmarks, while setting a new state of the art for non-autoregressive models. Marjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, Omer Levy |
ICML | 4 |
| 2020 | SpanBERT: Improving Pre-training by Representing and Predicting SpansabstractWe present SpanBERT, a pre-training method that is designed to better represent and predict spans of text. Our approach extends BERT by (1) masking contiguous random spans, rather than random tokens, and (2) training the span boundary representations to predict the entire content of the masked span, without relying on the individual token representations within it. SpanBERT consistently outperforms BERT and our better-tuned baselines, with substantial gains on span selection tasks such as question answering and coreference resolution. In particular, with the same training data and model size as BERT large , our single model obtains 94.6% and 88.7% F1 on SQuAD 1.1 and 2.0 respectively. We also achieve a new state of the art on the OntoNotes coreference resolution task (79.6% F1), strong performance on the TACRED relation extraction benchmark, and even gains on GLUE. 1 Mandar Joshi, Danqi Chen 0001, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, Omer Levy |
Trans. Assoc. Comput. Linguistics | 6 |
| 2019 | Mask-Predict: Parallel Decoding of Conditional Masked Language ModelsabstractMarjan Ghazvininejad, Omer Levy, Yinhan Liu, Luke Zettlemoyer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Marjan Ghazvininejad, Omer Levy, Yinhan Liu, Luke Zettlemoyer |
EMNLP/IJCNLP (1) | 2 |
| 2019 | BERT for Coreference Resolution: Baselines and AnalysisabstractMandar Joshi, Omer Levy, Luke Zettlemoyer, Daniel Weld. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mandar Joshi, Omer Levy, Luke Zettlemoyer, Daniel S. Weld |
EMNLP/IJCNLP (1) | 2 |
| 2019 | code2seq: Generating Sequences from Structured Representations of Code
Uri Alon 0002, Shaked Brody, Omer Levy, Eran Yahav |
ICLR (Poster) | 3 |
| 2019 | GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, Samuel R. Bowman |
ICLR (Poster) | 5 |
| 2019 | Understanding large-scale software: a hierarchical viewabstractProgram comprehension accounts for a large portion of software development costs and effort. The academic literature contains research on program comprehension of short code snippets, but comprehension at the system level is no less important. We claim that comprehending a software system is a distinct activity that differs from code comprehension. We interview experienced developers, architects, and managers in the software industry and open-source community, to uncover the meaning of program comprehension at the system level. The interviews demonstrate, among other things, that system comprehension is detached from code and programming language, and includes scope that is not captured in the code. It focuses on the structure of the system and less on the code itself. This is a continuous, iterative process, which mixes white-box and black-box approaches at different layers of the system, and combines both bottom-up and top-down comprehension strategies. Omer Levy, Dror G. Feitelson |
ICPC | 1 |
| 2019 | Are Sixteen Heads Really Better than One?abstractMulti-headed attention is a driving force behind recent state-of-the-art NLP models. By applying multiple attention mechanisms in parallel, it can express sophisticated functions beyond the simple weighted average. However we observe that, in practice, a large proportion of attention heads can be removed at test time without significantly impacting performance, and that some layers can even be reduced to a single head. Further analysis on machine translation models reveals that the self-attention layers can be significantly pruned, while the encoder-decoder layers are more dependent on multi-headedness. Paul Michel, Omer Levy, Graham Neubig |
NeurIPS | 2 |
| 2019 | SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding SystemsabstractIn the last year, new models and methods for pretraining and transfer learning have driven striking performance improvements across a range of language understanding tasks. The GLUE benchmark, introduced a little over one year ago, offers a single-number metric that summarizes progress on a diverse set of such tasks, but performance on the benchmark has recently surpassed the level of non-expert humans, suggesting limited headroom for further research. In this paper we present SuperGLUE, a new benchmark styled after GLUE with a new set of more difficult language understanding tasks, a software toolkit, and a public leaderboard. SuperGLUE is available at https://super.gluebenchmark.com. Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, Samuel R. Bowman |
NeurIPS | 7 |
| 2019 | code2vec: learning distributed representations of codeabstractWe present a neural model for representing snippets of code as continuous distributed vectors (``code embeddings''). The main idea is to represent a code snippet as a single fixed-length code vector, which can be used to predict semantic properties of the snippet. To this end, code is first decomposed to a collection of paths in its abstract syntax tree. Then, the network learns the atomic representation of each path while simultaneously learning how to aggregate a set of them. We demonstrate the effectiveness of our approach by using it to predict a method's name from the vector representation of its body. We evaluate our approach by training a model on a dataset of 12M methods. We show that code vectors trained on this dataset can predict method names from files that were unobserved during training. Furthermore, we show that our model learns useful method name vectors that capture semantic similarities, combinations, and analogies. A comparison of our approach to previous techniques over the same dataset shows an improvement of more than 75%, making it the first to successfully predict method names based on a large, cross-project corpus. Our trained model, visualizations and vector similarities are available as an interactive online demo at http://code2vec.org. The code, data and trained models are available at https://github.com/tech-srl/code2vec. Uri Alon 0002, Meital Zilberstein, Omer Levy, Eran Yahav |
Proc. ACM Program. Lang. | 3 |
| 2018 | Ultra-Fine Entity TypingabstractWe introduce a new entity typing task: given a sentence with an entity mention, the goal is to predict a set of free-form phrases (e.g.skyscraper, songwriter, or criminal) that describe appropriate types for the target entity.This formulation allows us to use a new type of distant supervision at large scale: head words, which indicate the type of the noun phrases they appear in.We show that these ultra-fine types can be crowd-sourced, and introduce new evaluation sets that are much more diverse and fine-grained than existing benchmarks.We present a model that can predict open types, and is trained using a multitask objective that pools our new head-word supervision with prior supervision from entity linking.Experimental results demonstrate that our model is effective in predicting entity types at varying granularity; it achieves state of the art performance on an existing fine-grained entity typing benchmark, and sets baselines for our newly-introduced datasets.1 Eunsol Choi, Omer Levy, Yejin Choi 0001, Luke Zettlemoyer |
ACL (1) | 2 |
| 2018 | Simulating Action Dynamics with Neural Process Networks
Antoine Bosselut, Omer Levy, Ari Holtzman, Corin Ennis, Dieter Fox, Yejin Choi 0001 |
ICLR (Poster) | 2 |
| 2018 | A general path-based representation for predicting program propertiesabstractPredicting program properties such as names or expression types has a wide range of applications. It can ease the task of programming, and increase programmer productivity. A major challenge when learning from programs is how to represent programs in a way that facilitates effective learning. Uri Alon 0002, Meital Zilberstein, Omer Levy, Eran Yahav |
PLDI | 3 |
| 2017 | Named Entity Disambiguation for Noisy TextabstractWe address the task of Named Entity Disambiguation (NED) for noisy text.We present WikilinksNED, a large-scale NED dataset of text fragments from the web, which is significantly noisier and more challenging than existing newsbased datasets.To capture the limited and noisy local context surrounding each mention, we design a neural model and train it with a novel method for sampling informative negative examples.We also describe a new way of initializing word and entity embeddings that significantly improves performance.Our model significantly outperforms existing state-ofthe-art methods on WikilinksNED while achieving comparable performance on a smaller newswire dataset. Yotam Eshel, Noam Cohen, Kira Radinsky, Shaul Markovitch, Ikuya Yamada, Omer Levy |
CoNLL | 6 |
| 2017 | Zero-Shot Relation Extraction via Reading ComprehensionabstractWe show that relation extraction can be reduced to answering simple reading comprehension questions, by associating one or more natural-language questions with each relation slot.This reduction has several advantages: we can (1) learn relationextraction models by extending recent neural reading-comprehension techniques, (2) build very large training sets for those models by combining relation-specific crowd-sourced questions with distant supervision, and even (3) do zero-shot learning by extracting new relation types that are only specified at test-time, for which we have no labeled training examples.Experiments on a Wikipedia slot-filling task demonstrate that the approach can generalize to new questions for known relation types with high accuracy, and that zero-shot generalization to unseen relation types is possible, at lower accuracy levels, setting the bar for future work on this task. Omer Levy, Minjoon Seo, Eunsol Choi, Luke Zettlemoyer |
CoNLL | 1 |
| 2017 | A Strong Baseline for Learning Cross-Lingual Word Embeddings from Sentence AlignmentsabstractWhile cross-lingual word embeddings have been studied extensively in recent years, the qualitative differences between the different algorithms remain vague.We observe that whether or not an algorithm uses a particular feature set (sentence IDs) accounts for a significant performance gap among these algorithms.This feature set is also used by traditional alignment algorithms, such as IBM Model-1, which demonstrate similar performance to stateof-the-art embedding algorithms on a variety of benchmarks.Overall, we observe that different algorithmic approaches for utilizing the sentence ID feature space result in similar performance.This paper draws both empirical and theoretical parallels between the embedding and alignment literature, and suggests that adding additional sources of information, which go beyond the traditional signal of bilingual sentence-aligned corpora, may substantially improve cross-lingual word embeddings, and that future baselines should at least take such features into account. Omer Levy, Anders Søgaard, Yoav Goldberg |
EACL (1) | 1 |
| 2016 | Modeling Extractive Sentence Intersection via Subtree EntailmentabstractSentence intersection captures the semantic overlap of two texts, generalizing over paradigms such as textual entailment and semantic text similarity. Despite its modeling power, it has received little attention because it is difficult for non-experts to annotate. We analyze 200 pairs of similar sentences and identify several underlying properties of sentence intersection. We leverage these insights to design an algorithm that decomposes the sentence intersection task into several simpler annotation tasks, facilitating the construction of a high quality dataset via crowdsourcing. We implement this approach and provide an annotated dataset of 1,764 sentence intersections. Omer Levy, Ido Dagan, Gabriel Stanovsky, Judith Eckle-Kohler, Iryna Gurevych |
COLING | 1 |
| 2015 | Learning to Exploit Structured Resources for Lexical InferenceabstractMassive knowledge resources, such as Wikidata, can provide valuable informa-tion for lexical inference, especially for proper-names. Prior resource-based ap-proaches typically select the subset of each resource’s relations which are relevant for a particular given task. The selection process is done manually, limiting these approaches to smaller resources such as WordNet, which lacks coverage of proper-names and recent terminology. This paper presents a supervised framework for auto-matically selecting an optimized subset of resource relations for a given target infer-ence task. Our approach enables the use of large-scale knowledge resources, thus providing a rich source of high-precision inferences over proper-names.1 1 Vered Shwartz, Omer Levy, Ido Dagan, Jacob Goldberger |
CoNLL | 2 |
| 2015 | Do Supervised Distributional Methods Really Learn Lexical Inference Relations?abstractOmer Levy, Steffen Remus, Chris Biemann, Ido Dagan. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Omer Levy, Steffen Remus, Chris Biemann, Ido Dagan |
HLT-NAACL | 1 |
| 2015 | Improving Distributional Similarity with Lessons Learned from Word EmbeddingsabstractRecent trends suggest that neural-network-inspired word embedding models outperform traditional count-based distributional models on word similarity and analogy detection tasks. We reveal that much of the performance gains of word embeddings are due to certain system design choices and hyperparameter optimizations, rather than the embedding algorithms themselves. Furthermore, we show that these modifications can be transferred to traditional distributional models, yielding similar gains. In contrast to prior reports, we observe mostly local or insignificant performance differences between the methods, with no global advantage to any single approach over the others. Omer Levy, Yoav Goldberg, Ido Dagan |
Trans. Assoc. Comput. Linguistics | 1 |
| 2014 | Focused Entailment Graphs for Open IE PropositionsabstractOpen IE methods extract structured propositions from text.However, these propositions are neither consolidated nor generalized, and querying them may lead to insufficient or redundant information.This work suggests an approach to organize open IE propositions using entailment graphs.The entailment relation unifies equivalent propositions and induces a specific-to-general structure.We create a large dataset of gold-standard proposition entailment graphs, and provide a novel algorithm for automatically constructing them.Our analysis shows that predicate entailment is extremely context-sensitive, and that current lexical-semantic resources do not capture many of the lexical inferences induced by proposition entailment. Omer Levy, Ido Dagan, Jacob Goldberger |
CoNLL | 1 |
| 2014 | Linguistic Regularities in Sparse and Explicit Word RepresentationsabstractRecent work has shown that neuralembedded word representations capture many relational similarities, which can be recovered by means of vector arithmetic in the embedded space.We show that Mikolov et al.'s method of first adding and subtracting word vectors, and then searching for a word similar to the result, is equivalent to searching for a word that maximizes a linear combination of three pairwise word similarities.Based on this observation, we suggest an improved method of recovering relational similarities, improving the state-of-the-art results on two recent word-analogy datasets.Moreover, we demonstrate that analogy recovery is not restricted to neural word embeddings, and that a similar amount of relational similarities can be recovered from traditional distributional word representations. Omer Levy, Yoav Goldberg |
CoNLL | 1 |
| 2014 | Neural Word Embedding as Implicit Matrix Factorization
Omer Levy, Yoav Goldberg |
NIPS | 1 |
| 2012 | Teaching Machines to Learn by MetaphorsabstractHumans have an uncanny ability to learn new concepts with very few examples. Cognitive theories have suggested that this is done by utilizing prior experience of related tasks. We propose to emulate this process in machines, by transforming new problems into old ones. These transformations are called metaphors. Obviously, the learner is not given a metaphor, but must acquire one through a learning process. We show that learning metaphors yield better results than existing transfer learning methods. Moreover, we argue that metaphors give a qualitative assessment of task relatedness. Omer Levy, Shaul Markovitch |
AAAI | 1 |