Vered Shwartz

dblp:166/2038 · DBLP profile ↗
← Back
30ranked-venue papers
8as first author
18since 2021 · last 2025
0000-0002-1151-4379ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
YearPublicationVenuePosition
2025 Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities
abstract
Shivam Chandhok, Wan-Cyuan Fan, Vered Shwartz, Vineeth N. Balasubramanian, Leonid Sigal. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shivam Chandhok, Wan-Cyuan Fan, Vered Shwartz, Vineeth N. Balasubramanian, Leonid Sigal
ACL (1)3
2025 CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs' Cultural Knowledge Through Human-AI Red-Teaming
abstract
Yu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yu Ying Chiu, Bill Y. Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi 0001
ACL (1)10
2025 Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events
abstract
The commonsense reasoning capabilities of vision-language models (VLMs), especially in abductive reasoning and defeasible reasoning, remain poorly understood. Most benchmarks focus on typical visual scenarios [1], [23], [42], making it difficult to discern whether model performance stems from keen perception and reasoning skills, or reliance on pure statistical recall. We argue that by focusing on atypical events in videos, clearer insights can be gained on the core capabilities of VLMs. Explaining and understanding such out-of-distribution events requires models to extend beyond basic pattern recognition and regurgitation of their prior knowledge. To this end, we introduce Black-SwanSuite, a benchmark for evaluating VLMs’ ability to reason about unexpected events through abductive and defeasible tasks. Our tasks artificially limit the amount of visual information provided to models while questioning them about hidden unexpected events, or provide new visual information that could change an existing hypothesis about the event. We curate a comprehensive benchmark suite comprising over 3,800 MCQ, 4,900 generative and 6,700 yes/no questions, spanning 1,655 videos. After extensively evaluating various state-of-the-art VLMs, including GPT-4o and Gemini 1.5 Pro, as well as open-source VLMs such as LLaVA-Video, we find significant performance gaps of up to 32% from humans on these tasks. Our findings reveal key limitations in current VLMs, emphasizing the need for enhanced model architectures and training strategies. Our data and leaderboard is available at https://blackswan.cs.ubc.ca.
Aditya Chinchure, Sahithya Ravi, Raymond T. Ng, Vered Shwartz, Boyang Li 0001, Leonid Sigal
CVPR4
2025 A Comparative Approach for Auditing Multilingual Phonetic Transcript Archives
abstract
Abstract Curating datasets that span multiple languages is challenging. To make the collection more scalable, researchers often incorporate one or more imperfect classifiers in the process, like language identification models. These models, however, are prone to failure, resulting in some language partitions being unreliable for downstream tasks. We introduce a statistical test, the Preference Proportion Test, for identifying such unreliable partitions. By annotating only 20 samples for a language partition, we are able to identify systematic transcription errors for 10 language partitions in a recent large multilingual transcribed audio archive, X-IPAPack (Zhu et al., 2024). We find that filtering these low-quality partitions out when training models for the downstream task of phonetic transcription brings substantial benefits, most notably a 25.7% relative improvement on transcribing recordings in out-of-distribution languages. Our work contributes an effective method for auditing multilingual audio archives.1
Farhan Samir, Emily P. Ahn, Shreya Prakash, Márton Sóskuthy, Vered Shwartz
Trans. Assoc. Comput. Linguistics5
2024 Small But Funny: A Feedback-Driven Approach to Humor Distillation
abstract
Sahithya Ravi, Patrick Huber, Akshat Shrivastava, Vered Shwartz, Arash Einolghozati. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Sahithya Ravi, Patrick Huber, Akshat Shrivastava, Vered Shwartz, Arash Einolghozati
ACL (1)4
2024 Stance Reasoner: Zero-Shot Stance Detection on Social Media with Explicit Reasoning
abstract
Social media platforms are rich sources of opinionated content. Stance detection allows the automatic extraction of users’ opinions on various topics from such content. We focus on zero-shot stance detection, where the model’s success relies on (a) having knowledge about the target topic; and (b) learning general reasoning strategies that can be employed for new topics. We present Stance Reasoner, an approach to zero-shot stance detection on social media that leverages explicit reasoning over background knowledge to guide the model’s inference about the document’s stance on a target. Specifically, our method uses a pre-trained language model as a source of world knowledge, with the chain-of-thought in-context learning approach to generate intermediate reasoning steps. Stance Reasoner outperforms the current state-of-the-art models on 3 Twitter datasets, including fully supervised models. It can better generalize across targets, while at the same time providing explicit and interpretable explanations for its predictions.
Maksym Taranukhin, Vered Shwartz, Evangelos E. Milios
LREC/COLING2
2024 Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models
abstract
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, Vered Shwartz. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Yejin Choi 0001, Yoav Goldberg, Maarten Sap, Vered Shwartz
EACL (1)8
2024 From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models
abstract
Despite recent advancements in visionlanguage models, their performance remains suboptimal on images from non-western cultures, due to underrepresentation in training datasets.Various benchmarks have been proposed to test models' cultural inclusivity, but they have limited coverage of cultures and do not adequately assess cultural diversity across universal as well as culture-specific local concepts.To address these limitations, we introduce the GLOBALRG benchmark, comprising two challenging tasks: retrieval across universals and cultural visual grounding.The former task entails retrieving culturally-diverse images for universal concepts from 50 countries, while the latter aims at grounding culture-specific concepts within images from 15 countries.Our evaluation across a wide range of models reveals that the performance varies significantly across cultures -underscoring the necessity for enhancing multicultural understanding in vision-language models.Our
Mehar Bhatia, Sahithya Ravi, Aditya Chinchure, Eunjeong Hwang, Vered Shwartz
EMNLP5
2024 Locating Information Gaps and Narrative Inconsistencies Across Languages: A Case Study of LGBT People Portrayals on Wikipedia
abstract
To explain social phenomena and identify systematic biases, much research in computational social science focuses on comparative text analyses.These studies often rely on coarse corpuslevel statistics or local word-level analyses, mainly in English.We introduce the INFOGAP method-an efficient and reliable approach to locating information gaps and inconsistencies in articles at the fact level, across languages.We evaluate INFOGAP by analyzing LGBT people's portrayals, across 2.7K biography pages on English, Russian, and French Wikipedias.We find large discrepancies in factual coverage across the languages.Moreover, our analysis reveals that biographical facts carrying negative connotations are more likely to be highlighted in Russian Wikipedia.Crucially, INFOGAP both facilitates large scale analyses, and pinpoints local document-and fact-level information gaps, laying a new foundation for targeted and nuanced comparative language analysis at scale. 1
Farhan Samir, Chan Young Park, Anjalie Field, Vered Shwartz, Yulia Tsvetkov
EMNLP4
2023 What happens before and after: Multi-Event Commonsense in Event Coreference Resolution
abstract
Event coreference models cluster event mentions pertaining to the same real-world event.Recent models rely on contextualized representations to recognize coreference among lexically or contextually similar mentions.However, models typically fail to leverage commonsense inferences, which is particularly limiting for resolving lexically-divergent mentions.We propose a model that extends event mentions with temporal commonsense inferences.Given a complex sentence with multiple events, e.g., "The man killed his wife and got arrested", with the target event "arrested", our model generates plausible events that happen before the target event -such as "the police arrived", and after it, such as "he was sentenced".We show that incorporating such inferences into an existing event coreference model improves its performance, and we analyze the coreferences in which such temporal knowledge is required. ① Mention Detection② Baseline Pairwise Scorer (spent, recovering): 0.14 (spent, gunshots): 0.05 (spent, shot): 0.1 (gunshots, shot): 0.82 … gunshots, shot, … recovering hospitalized spent Document 2The third coworker, Bryant Dalton, 39, spent two weeks in the hospital and still is recovering from gunshots to the neck and shoulder, prosecutors said. Document 1Bryant Dalton, 39, was shot in the neck and is hospitalized in good condition.… Document N … (spent, recovering): 0.14 (spent, gunshots): 0.05 (spent, shot): 0.1 (gunshots, shot): 0.82 … ③ Agglomerative Clustering gunshots, shot, … recovering spent, hospitalized spent, recovering, gunshots, shot, hospitalized ② Our Pairwise Scorer (spent, hospitalized): 0.1 (spent, hospitalized): 0.75 f(ctx spent , ctx hospitalized ) f(ctx spent , ctx hospitalized , cs spent , cs hospitalized ,)
Sahithya Ravi, Chris Tanner, Raymond T. Ng, Vered Shwartz
EACL4
2023 GD-COMET: A Geo-Diverse Commonsense Inference Model
abstract
With the increasing integration of AI into everyday life, it's becoming crucial to design AI systems that serve users from diverse backgrounds by making them culturally aware.In this paper, we present GD-COMET, a geo-diverse version of the COMET commonsense inference model.GD-COMET goes beyond Western commonsense knowledge and is capable of generating inferences pertaining to a broad range of cultures.We demonstrate the effectiveness of GD-COMET through a comprehensive human evaluation across 5 diverse cultures, as well as extrinsic evaluation on a geo-diverse task.The evaluation shows that GD-COMET captures and generates culturally nuanced commonsense knowledge, demonstrating its potential to benefit NLP applications across the board and contribute to making NLP more inclusive.PersonX eats a dutch baby PersonX wants PersonX feels Others feel sad, disgusting, happy, angry COMET PersonX is seen as hungry, greedy, disgusting, mean, starving As a result, a full stomach, a full belly, to satisfy hunger, to be full, to have a snack PersonX wants satiated, full, happy PersonX feels Others feel empty, craving, happy, keen PersonX is seen as hungry, ravenous, starving, greedy 🌎 GD- COMET satisfied, full, hungry, happy, guilty to throw up, drink water, wash the baby
Mehar Bhatia, Vered Shwartz
EMNLP2
2023 MemeCap: A Dataset for Captioning and Interpreting Memes
abstract
Memes are a widely popular tool for web users to express their thoughts using visual metaphors.Understanding memes requires recognizing and interpreting visual metaphors with respect to the text inside or around the meme, often while employing background knowledge and reasoning abilities.We present the task of meme captioning and release a new dataset, MEMECAP.Our dataset contains 6.3K memes along with the title of the post containing the meme, the meme captions, the literal image caption, and the visual metaphors.Despite the recent success of vision and language (VL) models on tasks such as image captioning and visual question answering, our extensive experiments using state-of-the-art VL models show that they still struggle with visual metaphors, and perform substantially worse than humans.
Eunjeong Hwang, Vered Shwartz
EMNLP2
2023 Knowledge Graph Compression Enhances Diverse Commonsense Generation
abstract
Generating commonsense explanations requires reasoning about commonsense knowledge beyond what is explicitly mentioned in the context.Existing models use commonsense knowledge graphs such as ConceptNet to extract a subgraph of relevant knowledge pertaining to concepts in the input.However, due to the large coverage and, consequently, vast scale of ConceptNet, the extracted subgraphs may contain loosely related, redundant and irrelevant information, which can introduce noise into the model.We propose to address this by applying a differentiable graph compression algorithm that focuses on more salient and relevant knowledge for the task.The compressed subgraphs yield considerably more diverse outputs when incorporated into models for the tasks of generating commonsense and abductive explanations.Moreover, our model achieves better quality-diversity tradeoff than a large language model with 100 times the number of parameters.Our generic approach can be applied to additional NLP tasks that can benefit from incorporating external knowledge.1
Eunjeong Hwang, Veronika Thost, Vered Shwartz, Tengfei Ma 0001
EMNLP3
2023 VLC-BERT: Visual Question Answering with Contextualized Commonsense Knowledge
abstract
There has been a growing interest in solving Visual Question Answering (VQA) tasks that require the model to reason beyond the content present in the image. In this work, we focus on questions that require commonsense reasoning. In contrast to previous methods which inject knowledge from static knowledge bases, we investigate the incorporation of contextualized knowledge using Commonsense Transformer (COMET), an existing knowledge model trained on human-curated knowledge bases. We propose a method to generate, select, and encode external commonsense knowledge alongside visual and textual cues in a new pre-trained Vision-Language-Commonsense transformer model, VLC-BERT. Through our evaluation on the knowledge-intensive OK-VQA and A-OKVQA datasets, we show that VLC-BERT is capable of outperforming existing models that utilize static knowledge bases. Furthermore, through a detailed analysis, we explain which questions benefit, and which don’t, from contextualized commonsense knowledge from COMET. Code: https://github.com/aditya10/VLC-BERT
Sahithya Ravi, Aditya Chinchure, Leonid Sigal, Renjie Liao 0001, Vered Shwartz
WACV5
2022 It's not Rocket Science: Interpreting Figurative Language in Narratives
abstract
Abstract Figurative language is ubiquitous in English. Yet, the vast majority of NLP research focuses on literal language. Existing text representations by design rely on compositionality, while figurative language is often non- compositional. In this paper, we study the interpretation of two non-compositional figurative languages (idioms and similes). We collected datasets of fictional narratives containing a figurative expression along with crowd-sourced plausible and implausible continuations relying on the correct interpretation of the expression. We then trained models to choose or generate the plausible continuation. Our experiments show that models based solely on pre-trained language models perform substantially worse than humans on these tasks. We additionally propose knowledge-enhanced models, adopting human strategies for interpreting figurative language types: inferring meaning from the context and relying on the constituent words’ literal meanings. The knowledge-enhanced models improve the performance on both the discriminative and generative tasks, further bridging the gap from human performance.
Tuhin Chakrabarty, Yejin Choi 0001, Vered Shwartz
Trans. Assoc. Comput. Linguistics3
2021 Learning to Rationalize for Nonmonotonic Reasoning with Distant Supervision
abstract
The black-box nature of neural models has motivated a line of research that aims to generate natural language rationales to explain why a model made certain predictions. Such rationale generation models, to date, have been trained on dataset-specific crowdsourced rationales, but this approach is costly and is not generalizable to new tasks and domains. In this paper, we investigate the extent to which neural models can reason about natural language rationales that explain model predictions, relying only on distant supervision with no additional annotation cost for human-written rationales. We investigate multiple ways to automatically generate rationales using pre-trained language models, neural knowledge models, and distant supervision from related tasks, and train generative models capable of composing explanatory rationales for unseen instances. We demonstrate our approach on the defeasible inference task, a nonmonotonic reasoning task in which an inference may be strengthened or weakened when new information (an update) is introduced. Our model shows promises at generating post-hoc rationales explaining why an inference is more or less likely given the additional information, however, it mostly generates trivial rationales reflecting the fundamental limitations of neural language models. Conversely, the more realistic setup of jointly predicting the update or its type and generating rationale is more challenging, suggesting an important future direction.
Faeze Brahman, Vered Shwartz, Rachel Rudinger, Yejin Choi 0001
AAAI2
2021 Paragraph-level Commonsense Transformers with Recurrent Memory
abstract
Human understanding of narrative texts requires making commonsense inferences beyond what is stated in the text explicitly. A recent model, COMET, can generate such inferences along several dimensions such as pre- and post-conditions, motivations, and mental states of the participants. However, COMET was trained on short phrases, and is therefore discourse-agnostic. When presented with each sentence of a multi-sentence narrative, it might generate inferences that are inconsistent with the rest of the narrative. We present the task of discourse-aware commonsense inference. Given a sentence within a narrative, the goal is to generate commonsense inferences along predefined dimensions, while maintaining coherence with the rest of the narrative. Such large-scale paragraph-level annotation is hard to get and costly, so we use available sentence-level annotations to efficiently and automatically construct a distantly supervised corpus. Using this corpus, we train PARA-COMET, a discourse-aware model that incorporates paragraph-level information to generate coherent commonsense inferences from narratives. PARA-COMET captures both semantic knowledge pertaining to prior world knowledge, and episodic knowledge involving how current events relate to prior and future events in a narrative. Our results confirm that PARA-COMET outperforms the sentence-level baselines, particularly in generating inferences that are both coherent and novel.
Saadia Gabriel, Chandra Bhagavatula, Vered Shwartz, Ronan Le Bras 0001, Maxwell Forbes, Yejin Choi 0001
AAAI3
2021 Surface Form Competition: Why the Highest Probability Answer Isn't Always Right
abstract
Large language models have shown promising results in zero-shot settings (Brown et al., 2020;Radford et al., 2019).For example, they can perform multiple choice tasks simply by conditioning on a question and selecting the answer with the highest probability.We introduce Domain Conditional Pointwise Mutual Information, an alternative scoring function that directly compensates for surface form competition by simply reweighing each option according to its a priori likelihood within the context of a specific task.It achieves consistent gains in zero-shot performance over both calibrated (Zhao et al., 2021) and uncalibrated scoring functions on all GPT-2 and GPT-3 models on a variety of multiple choice datasets.
Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi 0001, Luke Zettlemoyer
EMNLP (1)3
2020 Do Neural Language Models Overcome Reporting Bias?
abstract
Mining commonsense knowledge from corpora suffers from reporting bias, over-representing the rare at the expense of the trivial (Gordon and Van Durme, 2013).We study to what extent pre-trained language models overcome this issue.We find that while their generalization capacity allows them to better estimate the plausibility of frequent but unspoken of actions, outcomes, and properties, they also tend to overestimate that of the very rare, amplifying the bias that already exists in their training corpus.
Vered Shwartz, Yejin Choi 0001
COLING1
2020 Social Chemistry 101: Learning to Reason about Social and Moral Norms
abstract
Social norms-the unspoken commonsense rules about acceptable social behavior-are crucial in understanding the underlying causes and intents of people's actions in narratives.For example, underlying an action such as "wanting to call cops on my neighbor" are social norms that inform our conduct, such as "It is expected that you report crimes."We present SOCIAL CHEMISTRY, a new conceptual formalism to study people's everyday social norms and moral judgments over a rich spectrum of real life situations described in natural language.We introduce SOCIAL-CHEM-101, a large-scale corpus that catalogs 292k rules-of-thumb such as "It is rude to run a blender at 5am" as the basic conceptual units.Each rule-of-thumb is further broken down with 12 different dimensions of people's judgments, including social judgments of good and bad, moral foundations, expected cultural pressure, and assumed legality, which together amount to over 4.5 million annotations of categorical labels and free-text descriptions.Comprehensive empirical results based on state-of-the-art neural models demonstrate that computational modeling of social norms is a promising research direction.Our model framework, NEURAL NORM TRANSFORMER, learns and generalizes SOCIAL-CHEM-101 to successfully reason about previously unseen situations, generating relevant (and potentially novel) attribute-aware social rules-of-thumb.Punching a friend who stole from me.RoT 1: It is unacceptable to injure a person.RoT 2: People should not steal from others.RoT 3: It is bad to betray a friend.RoT 4: It is OK to want to take revenge.
Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, Yejin Choi 0001
EMNLP (1)3
2020 Back to the Future: Unsupervised Backprop-based Decoding for Counterfactual and Abductive Commonsense Reasoning
abstract
Lianhui Qin, Vered Shwartz, Peter West, Chandra Bhagavatula, Jena D. Hwang, Ronan Le Bras, Antoine Bosselut, Yejin Choi. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Lianhui Qin, Vered Shwartz, Peter West, Chandra Bhagavatula, Jena D. Hwang, Ronan Le Bras 0001, Antoine Bosselut, Yejin Choi 0001
EMNLP (1)2
2020 "You are grounded!": Latent Name Artifacts in Pre-trained Language Models
abstract
Pre-trained language models (LMs) may perpetuate biases originating in their training corpus to downstream models.We focus on artifacts associated with the representation of given names (e.g., Donald), which, depending on the corpus, may be associated with specific entities, as indicated by next token prediction (e.g., Trump).While helpful in some contexts, grounding happens also in underspecified or inappropriate contexts.For example, endings generated for 'Donald is a' substantially differ from those of other names, and often have more-than-average negative sentiment.We demonstrate the potential effect on downstream tasks with reading comprehension probes where name perturbation changes the model answers.As a silver lining, our experiments suggest that additional pre-training on different corpora may mitigate this bias. ModelMain Corpus Type Gen. Cls.Named Entities from News Named Entities from History Model Minimal News History Infrml Avg Minimal News History Infrml Avg GPT 0.0 7.0 12.7 1.4 5.3 0.0 21.9 39.1 7.8 17.2 GPT2-small 22.5 63.4 50.7 15.5 38.0 12.5 29.7 56.2 12.5 27.7 GPT2-medium 33.8 64.8 49.3 12.7 40.2 21.9 32.8 62.5 4.7 30.5 GPT2-large 43.7 66.2 47.9 16.9 43.7 29.7 29.7 56.2 12.5 32.0 GPT2-XL 50.7 62.0 45.1 21.1 44.7 28.1 31.2 60.9 14.1 33.6 TransformerXL 14.1 18.3 15.5 12.7 15.2 35.9 43.8 51.6 37.5 42.2 XLNet-base 4.2 33.8 12.7 4.2 13.7 0.0 34.4 23.4 3.1 15.2 XLNet-large 11.3 40.8 23.9 9.9 21.5 6.2 29.7 31.2 7.8 18.7 Average 22.5 44.5 32.2 11.8 27.7 16.8 31.7 47.6 12.5 27.1
Vered Shwartz, Rachel Rudinger, Oyvind Tafjord
EMNLP (1)1
2020 Unsupervised Commonsense Question Answering with Self-Talk
abstract
Natural language understanding involves reading between the lines with implicit background knowledge.Current systems either rely on pretrained language models as the sole implicit source of world knowledge, or resort to external knowledge bases (KBs) to incorporate additional relevant knowledge.We propose an unsupervised framework based on self-talk as a novel alternative to multiple-choice commonsense tasks.Inspired by inquiry-based discovery learning (Bruner, 1961), our approach inquires language models with a number of information seeking questions such as "what is the definition of ..." to discover additional background knowledge.Empirical results demonstrate that the self-talk procedure substantially improves the performance of zeroshot language model baselines on four out of six commonsense benchmarks, and competes with models that obtain knowledge from external KBs.While our approach improves performance on several benchmarks, the selftalk induced knowledge even when leading to correct answers is not always seen as helpful by human judges, raising interesting questions about the inner-workings of pre-trained language models for commonsense reasoning.
Vered Shwartz, Peter West, Ronan Le Bras 0001, Chandra Bhagavatula, Yejin Choi 0001
EMNLP (1)1
2019 Revisiting Joint Modeling of Cross-document Entity and Event Coreference Resolution
abstract
Recognizing coreferring events and entities across multiple texts is crucial for many NLP applications.Despite the task's importance, research focus was given mostly to withindocument entity coreference, with rather little attention to the other variants.We propose a neural architecture for cross-document coreference resolution.Inspired by Lee et al. (2012), we jointly model entity and event coreference.We represent an event (entity) mention using its lexical span, surrounding context, and relation to entity (event) mentions via predicate-arguments structures.Our model outperforms the previous state-of-the-art event coreference model on ECB+, while providing the first entity coreference results on this corpus.Our analysis confirms that all our representation elements, including the mention span itself, its context, and the relation to other mentions contribute to the model's success.
Shany Barhom, Vered Shwartz, Alon Eirew, Michael Bugert, Nils Reimers 0001, Ido Dagan
ACL (1)2
2019 Diversify Your Datasets: Analyzing Generalization via Controlled Variance in Adversarial Datasets
abstract
Phenomenon-specific "adversarial" datasets have been recently designed to perform targeted stress-tests for particular inference types.Recent work (Liu et al., 2019a) proposed that such datasets can be utilized for training NLI and other types of models, often allowing to learn the phenomenon in focus and improve on the challenge dataset, indicating a "blind spot" in the original training data.Yet, although a model can improve in such a training process, it might still be vulnerable to other challenge datasets targeting the same phenomenon but drawn from a different distribution, such as having a different syntactic complexity level.In this work, we extend this method to drive conclusions about a model's ability to learn and generalize a target phenomenon rather than to "learn" a dataset, by controlling additional aspects in the adversarial datasets.We demonstrate our approach on two inference phenomena -dative alternation and numerical reasoning, elaborating, and in some cases contradicting, the results of Liu et al.. Our methodology enables building better challenge datasets for creating more robust models, and may yield better model understanding and subsequent overarching improvements.
Ohad Rozen, Vered Shwartz, Roee Aharoni, Ido Dagan
CoNLL2
2019 Still a Pain in the Neck: Evaluating Text Representations on Lexical Composition
abstract
Building meaningful phrase representations is challenging because phrase meanings are not simply the sum of their constituent meanings. Lexical composition can shift the meanings of the constituent words and introduce implicit information. We tested a broad range of textual representations for their capacity to address these issues. We found that, as expected, contextualized word representations perform better than static word embeddings, more so on detecting meaning shift than in recovering implicit information, in which their performance is still far from that of humans. Our evaluation suite, consisting of six tasks related to lexical composition effects, can serve future research aiming to improve representations.
Vered Shwartz, Ido Dagan
Trans. Assoc. Comput. Linguistics1
2018 Paraphrase to Explicate: Revealing Implicit Noun-Compound Relations
abstract
Revealing the implicit semantic relation between the constituents of a nouncompound is important for many NLP applications.It has been addressed in the literature either as a classification task to a set of pre-defined relations or by producing free text paraphrases explicating the relations.Most existing paraphrasing methods lack the ability to generalize, and have a hard time interpreting infrequent or new noun-compounds.We propose a neural model that generalizes better by representing paraphrases in a continuous space, generalizing for both unseen noun-compounds and rare paraphrases.Our model helps improving performance on both the noun-compound paraphrasing and classification tasks.
Vered Shwartz, Ido Dagan
ACL (1)1
2017 Hypernyms under Siege: Linguistically-motivated Artillery for Hypernymy Detection
abstract
The fundamental role of hypernymy in NLP has motivated the development of many methods for the automatic identification of this relation, most of which rely on word distribution.We investigate an extensive number of such unsupervised measures, using several distributional semantic models that differ by context type and feature weighting.We analyze the performance of the different methods based on their linguistic motivation.Comparison to the state-of-the-art supervised methods shows that while supervised methods generally outperform the unsupervised ones, the former are sensitive to the distribution of training instances, hurting their reliability.Being based on general linguistic hypotheses and independent from training data, unsupervised measures are more robust, and therefore are still useful artillery for hypernymy detection.
Vered Shwartz, Enrico Santus, Dominik Schlechtweg
EACL (1)1
2016 Improving Hypernymy Detection with an Integrated Path-based and Distributional Method
abstract
Detecting hypernymy relations is a key task in NLP, which is addressed in the literature using two complementary approaches.Distributional methods, whose supervised variants are the current best performers, and path-based methods, which received less research attention.We suggest an improved path-based algorithm, in which the dependency paths are encoded using a recurrent neural network, that achieves results comparable to distributional methods.We then extend the approach to integrate both pathbased and distributional signals, significantly improving upon the state-of-the-art on this task.
Vered Shwartz, Yoav Goldberg, Ido Dagan
ACL (1)1
2015 Learning to Exploit Structured Resources for Lexical Inference
abstract
Massive knowledge resources, such as Wikidata, can provide valuable informa-tion for lexical inference, especially for proper-names. Prior resource-based ap-proaches typically select the subset of each resource’s relations which are relevant for a particular given task. The selection process is done manually, limiting these approaches to smaller resources such as WordNet, which lacks coverage of proper-names and recent terminology. This paper presents a supervised framework for auto-matically selecting an optimized subset of resource relations for a given target infer-ence task. Our approach enables the use of large-scale knowledge resources, thus providing a rich source of high-precision inferences over proper-names.1 1
Vered Shwartz, Omer Levy, Ido Dagan, Jacob Goldberger
CoNLL1