Roy Schwartz 0001

dblp:19/376-1 · DBLP profile ↗
← Back
38ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0002-5351-6209ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 9 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models
abstract
Text-to-image generation models suffer from alignment problems, where generated images fail to accurately capture the objects and relations in the text prompt. Prior work has focused on improving alignment by refining the diffusion process, ignoring the role of the text encoder, which guides the diffusion. In this work, we investigate how semantic information is distributed across token representations in text-to-image prompts, analyzing it at two levels: (1) in-item representation—whether individual tokens represent their lexical item (i.e., a word or expression conveying a single concept), and (2) cross-item interaction—whether information flows between tokens of different lexical items. We use patching techniques to uncover encoding patterns, and find that information is usually concentrated in only one or two of the item’s tokens; for example, in the item “San Francisco’s Golden Gate Bridge”, the token “Gate” sufficiently captures the entire expression while the other tokens could effectively be discarded. Lexical items also tend to remain isolated; for instance, in the prompt “a green dog”, the token “dog” encodes no visual information about “green”. However, in some cases, items do influence each other’s representation, often leading to misinterpretations—e.g., in the prompt “a pool by a table”, the token “pool” represents a “pool table” after contextualization. Our findings highlight the critical role of token-level encoding in image generation, and demonstrate that simple interventions at the encoding stage can substantially improve alignment and generation quality.
Guy Kaplan, Michael Toker, Yuval Reif, Yonatan Belinkov, Roy Schwartz 0001
ACL (1)5
2025 On Pruning State-Space LLMs
abstract
Recent work proposed state-space models (SSMs) as an efficient alternative to transformer-based LLMs.Can these models be pruned to further reduce their computation costs?We adapt several pruning methods to the SSM structure, and apply them to four SSMbased LLMs across multiple tasks.We find that such models are quite robust to some pruning methods (e.g., WANDA), while using other methods lead to fast performance degradation. 1
Tamer Ghattas, Michael Hassid, Roy Schwartz 0001
EMNLP3
2025 From Tokens to Words: On the Inner Lexicon of LLMs
abstract
Natural language is composed of words, but modern large language models (LLMs) process sub-words as input. A natural question raised by this discrepancy is whether LLMs encode words internally, and if so how. We present evidence that LLMs engage in an intrinsic detokenization process, where subword sequences are combined into coherent whole-word representations at their last token. Our experiments show that this process primarily takes place within the early and middle layers of the model. We further demonstrate its robustness to arbitrary splits (e.g., “cats” to “ca” and “ts”), typos, and importantly—to out-of-vocabulary words: when feeding the last token internal representations of such words to the model as input, it can “understand” them as the complete word despite never seeing such representations as input during training. Our findings suggest that LLMs maintain a latent vocabulary beyond the tokenizer’s scope. These insights provide a practical, finetuning-free application for expanding the vocabulary of pre-trained models. By enabling the addition of new vocabulary words, we reduce input length and inference iterations, which reduces both space and model latency, with little to no loss in model accuracy.
Guy Kaplan, Matanel Oren, Yuval Reif, Roy Schwartz 0001
ICLR4
2024 Transformers are Multi-State RNNs
abstract
Transformers are considered conceptually different from the previous generation of stateof-the-art NLP models-recurrent neural networks (RNNs).In this work, we demonstrate that decoder-only transformers can in fact be conceptualized as unbounded multistate RNNs-an RNN variant with unlimited hidden state size.We further show that transformers can be converted into bounded multistate RNNs by fixing the size of their hidden state, effectively compressing their keyvalue cache.We introduce a novel, trainingfree compression policy-Token Omission Via Attention (TOVA). 1 Our experiments with four long range tasks and several LLMs show that TOVA outperforms several baseline compression policies.Particularly, our results are nearly on par with the full model, using in some cases only 1 /8 of the original cache size, which translates to 4.8X higher throughput.Our results shed light on the connection between transformers and RNNs, and help mitigate one of LLMs' most painful computational bottlenecks-the size of their key-value cache. 2 * Equal contribuation 1 Literally "good" in Hebrew.
Matanel Oren, Michael Hassid, Yarden Nir-Buchbinder, Yossi Adi, Roy Schwartz 0001
EMNLP5
2024 Beyond Performance: Quantifying and Mitigating Label Bias in LLMs
abstract
Large language models (LLMs) have shown remarkable adaptability to diverse tasks, by leveraging context prompts containing instructions, or minimal input-output examples.However, recent work revealed they also exhibit label bias-an undesirable preference toward predicting certain answers over others.Still, detecting and measuring this bias reliably and at scale has remained relatively unexplored.In this study, we evaluate different approaches to quantifying label bias in a model's predictions, conducting a comprehensive investigation across 279 classification tasks and ten LLMs.Our investigation reveals substantial label bias in models both before and after debiasing attempts, as well as highlights the importance of outcomes-based evaluation metrics, which were not previously used in this regard.We further propose a novel label bias calibration method tailored for few-shot prompting, which outperforms recent calibration approaches for both improving performance and mitigating label bias.Our results emphasize that label bias in the predictions of LLMs remains a barrier to their reliability.1
Yuval Reif, Roy Schwartz 0001
NAACL-HLT2
2024 Morphosyntactic probing of multilingual BERT models
abstract
Abstract We introduce an extensive dataset for multilingual probing of morphological information in language models (247 tasks across 42 languages from 10 families), each consisting of a sentence with a target word and a morphological tag as the desired label, derived from the Universal Dependencies treebanks. We find that pre-trained Transformer models (mBERT and XLM-RoBERTa) learn features that attain strong performance across these tasks. We then apply two methods to locate, for each probing task, where the disambiguating information resides in the input. The first is a new perturbation method that “masks” various parts of context; the second is the classical method of Shapley values. The most intriguing finding that emerges is a strong tendency for the preceding context to hold more information relevant to the prediction than the following context.
Judit Ács, Endre Hamerlik, Roy Schwartz 0001, Noah A. Smith, András Kornai
Nat. Lang. Eng.3
2023 VASR: Visual Analogies of Situation Recognition
abstract
A core process in human cognition is analogical mapping: the ability to identify a similar relational structure between different situations. We introduce a novel task, Visual Analogies of Situation Recognition, adapting the classical word-analogy task into the visual domain. Given a triplet of images, the task is to select an image candidate B' that completes the analogy (A to A' is like B to what?). Unlike previous work on visual analogy that focused on simple image transformations, we tackle complex analogies requiring understanding of scenes. We leverage situation recognition annotations and the CLIP model to generate a large set of 500k candidate analogies. Crowdsourced annotations for a sample of the data indicate that humans agree with the dataset label ~80% of the time (chance level 25%). Furthermore, we use human annotations to create a gold-standard dataset of 3,820 validated analogies. Our experiments demonstrate that state-of-the-art models do well when distractors are chosen randomly (~86%), but struggle with carefully chosen distractors (~53%, compared to 90% human accuracy). We hope our dataset will encourage the development of new analogy-making models. Website: https://vasr-dataset.github.io/
Yonatan Bitton, Ron Yosef, Eli Strugo, Dafna Shahaf, Roy Schwartz 0001, Gabriel Stanovsky
AAAI5
2023 Finding the SWEET Spot: Analysis and Improvement of Adaptive Inference in Low Resource Settings
abstract
Adaptive inference is a simple method for reducing inference costs.The method works by maintaining multiple classifiers of different capacities, and allocating resources to each test instance according to its difficulty.In this work, we compare the two main approaches for adaptive inference, Early-Exit and Multi-Model, when training data is limited.First, we observe that for models with the same architecture and size, individual Multi-Model classifiers outperform their Early-Exit counterparts by an average of 2.3%.We show that this gap is caused by Early-Exit classifiers sharing model parameters during training, resulting in conflicting gradient updates of model weights.We find that despite this gap, Early-Exit still provides a better speed-accuracy trade-off due to the overhead of the Multi-Model approach.To address these issues, we propose SWEET, 1 an Early-Exit fine-tuning method that assigns each classifier its own set of unique model weights, not updated by other classifiers.We compare SWEET's speed-accuracy curve to standard Early-Exit and Multi-Model baselines and find that it outperforms both methods at fast speeds while maintaining comparable scores to Early-Exit at slow speeds.Moreover, SWEET individual classifiers outperform Early-Exit ones by 1.1% on average.SWEET enjoys the benefits of both methods, paving the way for further reduction of inference costs in NLP.We publicly release our code.2
Daniel Rotem, Michael Hassid, Jonathan Mamou, Roy Schwartz 0001
ACL (1)4
2023 Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional Images
abstract
Weird, unusual, and uncanny images pique the curiosity of observers because they challenge commonsense. For example, an image released during the 2022 world cup depicts the famous soccer stars Lionel Messi and Cristiano Ronaldo playing chess, which playfully violates our expectation that their competition should occur on the football field.1Humans can easily recognize and interpret these unconventional images, but can AI models do the same? We introduce WHOOPS!, a new dataset and benchmark for visual commonsense. The dataset is comprised of purposefully commonsense-defying images created by designers using publicly-available image generation tools like Midjourney. We consider several tasks posed over the dataset. In addition to image captioning, cross-modal matching, and visual question answering, we introduce a difficult explanation generation task, where models must identify and explain why a given image is unusual. Our results show that state-of-the-art models such as GPT3 and BLIP2 still lag behind human performance on WHOOPS!. We hope our dataset will inspire the development of AI models with stronger visual commonsense reasoning abilities.2
Nitzan Guetta, Yonatan Bitton, Jack Hessel, Ludwig Schmidt, Yuval Elovici, Gabriel Stanovsky, Roy Schwartz 0001
ICCV7
2023 Textually Pretrained Speech Language Models
abstract
Speech language models (SpeechLMs) process and generate acoustic data only, without textual supervision. In this work, we propose TWIST, a method for training SpeechLMs using a warm-start from a pretrained textual language models. We show using both automatic and human evaluations that TWIST outperforms a cold-start SpeechLM across the board. We empirically analyze the effect of different model design choices such as the speech tokenizer, the pretrained textual model, and the dataset size. We find that model and dataset scale both play an important role in constructing better-performing SpeechLMs. Based on our observations, we present the largest (to the best of our knowledge) SpeechLM both in terms of number of parameters and training data. We additionally introduce two spoken versions of the StoryCloze textual benchmark to further improve model evaluation and advance future research in the field. We make speech samples, code and models publicly available.
Michael Hassid, Tal Remez, Tu Anh Nguyen, Itai Gat, Alexis Conneau, Felix Kreuk, Jade Copet, Alexandre Défossez, Gabriel Synnaeve, Emmanuel Dupoux, Roy Schwartz 0001, Yossi Adi
NeurIPS11
2023 Efficient Methods for Natural Language Processing: A Survey
abstract
Abstract Recent work in natural language processing (NLP) has yielded appealing results from scaling model parameters and training data; however, using only scale to improve performance means that resource consumption also grows. Such resources include data, time, storage, or energy, all of which are naturally limited and unevenly distributed. This motivates research into efficient methods that require fewer resources to achieve similar results. This survey synthesizes and relates current methods and findings in efficient NLP. We aim to provide both guidance for conducting NLP under limited resources, and point towards promising research directions for developing more efficient methods.
Marcos V. Treviso, Ji-Ung Lee, Tianchu Ji, Betty van Aken, Manuel R. Ciosici, Michael Hassid, Kenneth Heafield, Sara Hooker, Colin Raffel, Pedro Henrique Martins, André F. T. Martins, Jessica Zosa Forde, Peter A. Milder, Edwin Simpson, Noam Slonim, Jesse Dodge, Emma Strubell, Niranjan Balasubramanian, Leon Derczynski, Iryna Gurevych, Roy Schwartz 0001
Trans. Assoc. Comput. Linguistics22
2022 ABC: Attention with Bounded-memory Control
abstract
Hao Peng, Jungo Kasai, Nikolaos Pappas, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz, Noah Smith. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Hao Peng 0009, Jungo Kasai, Nikolaos Pappas 0002, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz 0001, Noah A. Smith
ACL (1)7
2022 WinoGAViL: Gamified Association Benchmark to Challenge Vision-and-Language Models
abstract
While vision-and-language models perform well on tasks such as visual question answering, they struggle when it comes to basic human commonsense reasoning skills. In this work, we introduce WinoGAViL: an online game of vision-and-language associations (e.g., between werewolves and a full moon), used as a dynamic evaluation benchmark. Inspired by the popular card game Codenames, a spymaster gives a textual cue related to several visual candidates, and another player tries to identify them. Human players are rewarded for creating associations that are challenging for a rival AI model but still solvable by other human players. We use the game to collect 3.5K instances, finding that they are intuitive for humans (>90% Jaccard index) but challenging for state-of-the-art AI models, where the best model (ViLT) achieves a score of 52%, succeeding mostly where the cue is visually salient. Our analysis as well as the feedback we collect from players indicate that the collected associations require diverse reasoning skills, including general knowledge, common sense, abstraction, and more. We release the dataset, the code and the interactive game, allowing future data collection that can be used to develop models with better association abilities.
Yonatan Bitton, Nitzan Guetta, Ron Yosef, Yuval Elovici, Mohit Bansal, Gabriel Stanovsky, Roy Schwartz 0001
NeurIPS7
2021 Effects of Parameter Norm Growth During Transformer Training: Inductive Bias from Gradient Descent
abstract
The capacity of neural networks like the widely adopted transformer is known to be very high.Evidence is emerging that they learn successfully due to inductive bias in the training routine, typically a variant of gradient descent (GD).To better understand this bias, we study the tendency for transformer parameters to grow in magnitude (ℓ 2 norm) during training, and its implications for the emergent representations within self attention layers.Empirically, we document norm growth in the training of transformer language models, including T5 during its pretraining.As the parameters grow in magnitude, we prove that the network approximates a discretized network with saturated activation functions.Such "saturated" networks are known to have a reduced capacity compared to the full network family that can be described in terms of formal languages and automata.Our results suggest saturation is a new characterization of an inductive bias implicit in GD of particular interest for NLP.We leverage the emergent discrete structure in a saturated transformer to analyze the role of different attention heads, finding that some focus locally on a small number of positions, while other heads compute global averages, allowing counting.We believe understanding the interplay between these two capabilities may shed further light on the structure of computation within large transformers.
William Merrill, Vivek Ramanujan, Yoav Goldberg, Roy Schwartz 0001, Noah A. Smith
EMNLP (1)4
2021 Random Feature Attention
Hao Peng 0009, Nikolaos Pappas 0002, Dani Yogatama, Roy Schwartz 0001, Noah A. Smith, Lingpeng Kong
ICLR4
2021 Automatic Generation of Contrast Sets from Scene Graphs: Probing the Compositional Consistency of GQA
abstract
Yonatan Bitton, Gabriel Stanovsky, Roy Schwartz, Michael Elhadad. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Yonatan Bitton, Gabriel Stanovsky, Roy Schwartz 0001, Michael Elhadad
NAACL-HLT3
2021 Extracting a Knowledge Base of Mechanisms from COVID-19 Papers
abstract
Tom Hope, Aida Amini, David Wadden, Madeleine van Zuylen, Sravanthi Parasa, Eric Horvitz, Daniel Weld, Roy Schwartz, Hannaneh Hajishirzi. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Tom Hope, Aida Amini, Dave Wadden, Madeleine van Zuylen, Sravanthi Parasa, Eric Horvitz, Daniel S. Weld, Roy Schwartz 0001, Hannaneh Hajishirzi
NAACL-HLT8
2021 Provable Limitations of Acquiring Meaning from Ungrounded Form: What Will Future Language Models Understand?
abstract
Abstract Language models trained on billions of tokens have recently led to unprecedented results on many NLP tasks. This success raises the question of whether, in principle, a system can ever “understand” raw text without access to some form of grounding. We formally investigate the abilities of ungrounded systems to acquire meaning. Our analysis focuses on the role of “assertions”: textual contexts that provide indirect clues about the underlying semantics. We study whether assertions enable a system to emulate representations preserving semantic relations like equivalence. We find that assertions enable semantic emulation of languages that satisfy a strong notion of semantic transparency. However, for classes of languages where the same expression can take different values in different contexts, we show that emulation can become uncomputable. Finally, we discuss differences between our formal model and natural language, exploring how our results generalize to a modal setting and other semantic relations. Together, our results suggest that assertions in code or language do not provide sufficient signal to fully emulate semantic representations. We formalize ways in which ungrounded language models appear to be fundamentally limited in their ability to “understand”.
William Merrill, Yoav Goldberg, Roy Schwartz 0001, Noah A. Smith
Trans. Assoc. Comput. Linguistics3
2020 A Formal Hierarchy of RNN Architectures
abstract
We develop a formal hierarchy of the expressive capacity of RNN architectures.The hierarchy is based on two formal properties: space complexity, which measures the RNN's memory, and rational recurrence, defined as whether the recurrent update can be described by a weighted finite-state machine.We place several RNN variants within this hierarchy.For example, we prove the LSTM is not rational, which formally separates it from the related QRNN (Bradbury et al., 2016).We also show how these models' expressive capacity is expanded by stacking multiple layers or composing them with different pooling functions.Our results build on the theory of "saturated" RNNs (Merrill, 2019).While formally extending these findings to unsaturated RNNs is left to future work, we hypothesize that the practical learnable capacity of unsaturated RNNs obeys a similar hierarchy.Experimental findings from training unsaturated networks on formal languages support this conjecture.We report updated experiments in Appendix H.
William Merrill, Gail Weiss, Yoav Goldberg, Roy Schwartz 0001, Noah A. Smith, Eran Yahav
ACL4
2020 A Mixture of h - 1 Heads is Better than h Heads
abstract
Multi-head attentive neural architectures have achieved state-of-the-art results on a variety of natural language processing tasks.Evidence has shown that they are overparameterized; attention heads can be pruned without significant performance loss.In this work, we instead "reallocate" them-the model learns to activate different heads on different inputs.Drawing connections between multi-head attention and mixture of experts, we propose the mixture of attentive experts model (MAE).MAE is trained using a block coordinate descent algorithm that alternates between updating (1) the responsibilities of the experts and (2) their parameters.Experiments on machine translation and language modeling show that MAE outperforms strong baselines on both tasks.Particularly, on the WMT14 English to German translation dataset, MAE improves over "transformer-base" by 0.8 BLEU, with a comparable number of parameters.Our analysis shows that our model learns to specialize different experts to different inputs. 1
Hao Peng 0009, Roy Schwartz 0001, Dianqi Li, Noah A. Smith
ACL2
2020 The Right Tool for the Job: Matching Model and Instance Complexities
abstract
As NLP models become larger, executing a trained model requires significant computational resources incurring monetary and environmental costs.To better respect a given inference budget, we propose a modification to contextual representation fine-tuning which, during inference, allows for an early (and fast) "exit" from neural network calculations for simple instances, and late (and accurate) exit for hard instances.To achieve this, we add classifiers to different layers of BERT and use their calibrated confidence scores to make early exit decisions.We test our proposed modification on five different datasets in two tasks: three text classification datasets and two natural language inference benchmarks.Our method presents a favorable speed/accuracy tradeoff in almost all cases, producing models which are up to five times faster than the state of the art, while preserving their accuracy.Our method also requires almost no additional training resources (in either time or parameters) compared to the baseline BERT model.Finally, our method alleviates the need for costly retraining of multiple models at different levels of efficiency; we allow users to control the inference speed/accuracy tradeoff using a single trained model, by setting a single variable at inference time.We publicly release our code.1 * Research completed during an internship at AI2. 1 github.com/allenai/sledgehammerLayer 0 Layer i Layer k Layer n Input Layer l Layer j Is confident?Yes
Roy Schwartz 0001, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, Noah A. Smith
ACL1
2020 Dataset Cartography: Mapping and Diagnosing Datasets with Training Dynamics
abstract
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, Yejin Choi. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Swabha Swayamdipta, Roy Schwartz 0001, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, Yejin Choi 0001
EMNLP (1)2
2019 Show Your Work: Improved Reporting of Experimental Results
abstract
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz 0001, Noah A. Smith
EMNLP/IJCNLP (1)4
2019 RNN Architecture Learning with Sparse Regularization
abstract
Jesse Dodge, Roy Schwartz, Hao Peng, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jesse Dodge, Roy Schwartz 0001, Hao Peng 0009, Noah A. Smith
EMNLP/IJCNLP (1)2
2019 PaLM: A Hybrid Parser and Language Model
abstract
Hao Peng, Roy Schwartz, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Hao Peng 0009, Roy Schwartz 0001, Noah A. Smith
EMNLP/IJCNLP (1)2
2019 Knowledge Enhanced Contextual Word Representations
abstract
Matthew E. Peters, Mark Neumann, Robert Logan, Roy Schwartz, Vidur Joshi, Sameer Singh, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Matthew E. Peters, Mark Neumann, Robert L. Logan IV, Roy Schwartz 0001, Vidur Joshi, Sameer Singh 0001, Noah A. Smith
EMNLP/IJCNLP (1)4
2018 Bridging CNNs, RNNs, and Weighted Finite-State Machines
abstract
Recurrent and convolutional neural networks comprise two distinct families of models that have proven to be useful for encoding natural language utterances.In this paper we present SoPa, a new model that aims to bridge these two approaches.SoPa combines neural representation learning with weighted finite-state automata (WFSAs) to learn a soft version of traditional surface patterns.We show that SoPa is an extension of a one-layer CNN, and that such CNNs are equivalent to a restricted version of SoPa, and accordingly, to a restricted form of WFSA.Empirically, on three text classification tasks, SoPa is comparable or better than both a BiLSTM (RNN) baseline and a CNN baseline, and is particularly useful in small data settings.
Roy Schwartz 0001, Sam Thomson, Noah A. Smith
ACL (1)1
2018 Rational Recurrences
abstract
Despite the tremendous empirical success of neural models in natural language processing, many of them lack the strong intuitions that accompany classical machine learning approaches.Recently, connections have been shown between convolutional neural networks (CNNs) and weighted finite state automata (WFSAs), leading to new interpretations and insights.In this work, we show that some recurrent neural networks also share this connection to WFSAs.We characterize this connection formally, defining rational recurrences to be recurrent hidden state update functions that can be written as the Forward calculation of a finite set of WFSAs.We show that several recent neural models use rational recurrences.Our analysis provides a fresh view of these models and facilitates devising new neural architectures that draw inspiration from WFSAs.We present one such model, which performs better than two recent baselines on language modeling and text classification.Our results demonstrate that transferring intuitions from classical models like WFSAs can be an effective approach to designing and understanding neural models.
Hao Peng 0009, Roy Schwartz 0001, Sam Thomson, Noah A. Smith
EMNLP2
2018 SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference
abstract
Given a partial description like "she opened the hood of the car," humans can reason about the situation and anticipate what might come next ("then, she examined the engine").In this paper, we introduce the task of grounded commonsense inference, unifying natural language inference and commonsense reasoning.We present Swag, a new dataset with 113k multiple choice questions about a rich spectrum of grounded situations.To address the recurring challenges of the annotation artifacts and human biases found in many existing datasets, we propose Adversarial Filtering (AF), a novel procedure that constructs a de-biased dataset by iteratively training an ensemble of stylistic classifiers, and using them to filter the data.To account for the aggressive adversarial filtering, we use state-of-theart language models to massively oversample a diverse set of potential counterfactuals.Empirical results demonstrate that while humans can solve the resulting inference problems with high accuracy (88%), various competitive models struggle on our task.We provide comprehensive analysis that indicates significant opportunities for future research.
Rowan Zellers, Yonatan Bisk, Roy Schwartz 0001, Yejin Choi 0001
EMNLP3
2018 A Dataset of Peer Reviews (PeerRead): Collection, Insights and NLP Applications
abstract
Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine van Zuylen, Sebastian Kohlmeier, Eduard Hovy, Roy Schwartz. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine van Zuylen, Sebastian Kohlmeier, Eduard H. Hovy, Roy Schwartz 0001
NAACL-HLT7
2017 The Effect of Different Writing Tasks on Linguistic Style: A Case Study of the ROC Story Cloze Task
abstract
A writer's style depends not just on personal traits but also on her intent and mental state.In this paper, we show how variants of the same writing task can lead to measurable differences in writing style.We present a case study based on the story cloze task (Mostafazadeh et al., 2016a), where annotators were assigned similar writing tasks with different constraints: (1) writing an entire story, (2) adding a story ending for a given story context, and (3) adding an incoherent ending to a story.We show that a simple linear classifier informed by stylistic features is able to successfully distinguish among the three cases, without even looking at the story context.In addition, combining our stylistic features with language model predictions reaches state of the art performance on the story cloze challenge.Our results demonstrate that different task framings can dramatically affect the way people write. 1 1 This paper extends our LSDSem 2017 shared task submission (Schwartz et al., 2017).
Roy Schwartz 0001, Maarten Sap, Ioannis Konstas, Leila Zilles, Yejin Choi 0001, Noah A. Smith
CoNLL1
2017 Automatic Selection of Context Configurations for Improved Class-Specific Word Representations
abstract
This paper is concerned with identifying contexts useful for training word representation models for different word classes such as adjectives (A), verbs (V), and nouns (N).We introduce a simple yet effective framework for an automatic selection of class-specific context configurations.We construct a context configuration space based on universal dependency relations between words, and efficiently search this space with an adapted beam search algorithm.In word similarity tasks for each word class, we show that our framework is both effective and efficient.Particularly, it improves the Spearman's ρ correlation with human scores on SimLex-999 over the best previously proposed class-specific contexts by 6 (A), 6 (V) and 5 (N) ρ points.With our selected context configurations, we train on only 14% (A), 26.2% (V), and 33.6% (N) of all dependency-based contexts, resulting in a reduced training time.Our results generalise: we show that the configurations our algorithm learns for one English training setup outperform previously proposed context types in another training setup for English.Moreover, basing the configuration space on universal dependencies, it is possible to transfer the learned configurations to German and Italian.We also demonstrate improved per-class results over other context types in these two languages.
Ivan Vulic, Roy Schwartz 0001, Ari Rappoport, Roi Reichart, Anna Korhonen
CoNLL2
2016 Symmetric Patterns and Coordinations: Fast and Enhanced Representations of Verbs and Adjectives
abstract
State-of-the-art word embeddings, which are often trained on bag-of-words (BOW) contexts, provide a high quality representation of aspects of the semantics of nouns.However, their quality decreases substantially for the task of verb similarity prediction.In this paper we show that using symmetric pattern contexts (SPs, e.g., "X and Y") improves word2vec verb similarity performance by up to 15% and is also instrumental in adjective similarity prediction.The unsupervised SP contexts are even superior to a variety of dependency contexts extracted using a supervised dependency parser.Moreover, we observe that SPs and dependency coordination contexts (Coor) capture a similar type of information, and demonstrate that Coor contexts are superior to other dependency contexts including the set of all dependency contexts, although they are still inferior to SPs.Finally, there are substantially fewer SP contexts compared to alternative representations, leading to a massive reduction in training time.On an 8G words corpus and a 32 core machine, the SP model trains in 11 minutes, compared to 5 and 11 hours with BOW and all dependency contexts, respectively.
Roy Schwartz 0001, Roi Reichart, Ari Rappoport
HLT-NAACL1
2015 Symmetric Pattern Based Word Embeddings for Improved Word Similarity Prediction
abstract
We present a novel word level vector representation based on symmetric patterns (SPs).For this aim we automatically acquire SPs (e.g., "X and Y") from a large corpus of plain text, and generate vectors where each coordinate represents the cooccurrence in SPs of the represented word with another word of the vocabulary.Our representation has three advantages over existing alternatives: First, being based on symmetric word relationships, it is highly suitable for word similarity prediction.Particularly, on the SimLex999 word similarity dataset, our model achieves a Spearman's ρ score of 0.517, compared to 0.462 of the state-of-the-art word2vec model.Interestingly, our model performs exceptionally well on verbs, outperforming stateof-the-art baselines by 20.2-41.5%.Second, pattern features can be adapted to the needs of a target NLP application.For example, we show that we can easily control whether the embeddings derived from SPs deem antonym pairs (e.g.(big,small)) as similar or dissimilar, an important distinction for tasks such as word classification and sentiment analysis.Finally, we show that a simple combination of the word similarity scores generated by our method and by word2vec results in a superior predictive power over that of each individual model, scoring as high as 0.563 in Spearman's ρ on SimLex999.This emphasizes the differences between the signals captured by each of the models.
Roy Schwartz 0001, Roi Reichart, Ari Rappoport
CoNLL1
2014 Minimally Supervised Classification to Semantic Categories using Automatically Acquired Symmetric Patterns
Roy Schwartz 0001, Roi Reichart, Ari Rappoport
COLING1
2013 Authorship Attribution of Micro-Messages
abstract
Work on authorship attribution has traditionally focused on long texts.In this work, we tackle the question of whether the author of a very short text can be successfully identified.We use Twitter as an experimental testbed.We introduce the concept of an author's unique "signature", and show that such signatures are typical of many authors when writing very short texts.We also present a new authorship attribution feature ("flexible patterns") and demonstrate a significant improvement over our baselines.Our results show that the author of a single tweet can be identified with good accuracy in an array of flavors of the authorship attribution task.
Roy Schwartz 0001, Oren Tsur, Ari Rappoport, Moshe Koppel
EMNLP1
2012 Learnability-Based Syntactic Annotation Design
Roy Schwartz 0001, Omri Abend, Ari Rappoport
COLING1
2011 Neutralizing Linguistically Problematic Annotations in Unsupervised Dependency Parsing Evaluation
Roy Schwartz 0001, Omri Abend, Roi Reichart, Ari Rappoport
ACL1