Yoav Goldberg

dblp:68/5296 · DBLP profile ↗
← Back
110ranked-venue papers
15as first author
41since 2021 · last 2026
0000-0002-6497-829XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 104 · 15 first-author · 39 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Factual Retrieval in LLMs Is a Redundant, Distributed and Non-Contiguous Process
abstract
Large language models (LLMs) store and recall factual knowledge, yet the precise mechanism of how entity representations are transformed to enable specific attribute retrieval remains underexplored.In this work, we investigate this mechanism through the lens of an "attribute-computation path"-a sequence of computational steps over the entity representation required to elicit a target attribute.We then propose an iterative patching protocol to identify a minimal subset of layers necessary for this computation.Applying our method to LLaMA 3.1 8B and Qwen3 8B, we find that these paths are non-contiguous, often skipping layers, and that models possess multiple, functionally-equivalent paths for the same entity and fact, highlighting a high degree of redundancy in attribute computation.This implies that knowledge computation is highly distributed, potentially explaining the localizationediting mismatch and suggesting that knowledge storage and retrieval in LLMs is far from being well understood.
Hail Hochman, Natalie Shapira, Yoav Goldberg
ACL (1)3
2026 MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
abstract
Abstract Automated agents, powered by large language models (LLMs), are emerging as the go-to tool for querying information. However, evaluation benchmarks for LLM agents rarely feature natural questions that are both information-seeking and genuinely time-consuming for humans. To address this gap we introduce MoNaCo, a benchmark of 1,315 natural and time-consuming questions that require dozens, and at times hundreds, of intermediate steps to solve— far more than any existing QA benchmark. To build MoNaCo, we developed a decomposed annotation pipeline to elicit and manually answer real-world time-consuming questions at scale. Frontier LLMs evaluated on MoNaCo achieve at most 61.2% F1, hampered by low recall and hallucinations. Our results underscore the limitations of LLM-powered agents in handling the complexity and sheer breadth of real-world information-seeking tasks—with MoNaCo providing an effective resource for tracking such progress. The MoNaCo benchmark, codebase, prompts, and models predictions are all publicly available at: https://tomerwolgithub.github.io/monaco.
Tomer Wolfson, Harsh Trivedi, Mor Geva, Yoav Goldberg, Dan Roth 0001, Tushar Khot, Ashish Sabharwal, Reut Tsarfaty
Trans. Assoc. Comput. Linguistics4
2025 The Power of Reframing: Using LLMs in Synthesizing RV Monitors
Itay Cohen 0001, Klaus Havelund, Doron A. Peled, Yoav Goldberg
RV4
2024 Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
abstract
This paper explores the impact of extending input lengths on the capabilities of Large Language Models (LLMs).Despite LLMs advancements in recent times, their performance consistency across different input lengths is not well understood.We investigate this aspect by introducing a novel QA reasoning framework, specifically designed to assess the impact of input length.We isolate the effect of input length using multiple versions of the same sample, each being extended with padding of different lengths, types and locations.Our findings show a notable degradation in LLMs' reasoning performance at much shorter input lengths than their technical maximum.We show that the degradation trend appears in every version of our dataset, although at different intensities.Additionally, our study reveals that the traditional metric of next word prediction correlates negatively with performance of LLMs' on our reasoning dataset.We analyse our results and identify failure modes that can serve as useful guides for future research, potentially informing strategies to address the limitations observed in LLMs.
Mosh Levy, Alon Jacoby, Yoav Goldberg
ACL (1)3
2024 Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models
abstract
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, Vered Shwartz. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Yejin Choi 0001, Yoav Goldberg, Maarten Sap, Vered Shwartz
EACL (1)6
2024 Evaluating D-MERIT of Partial-annotation on Information Retrieval
abstract
Royi Rassin, Yaron Fairstein, Oren Kalinsky, Guy Kushilevitz, Nachshon Cohen, Alexander Libov, Yoav Goldberg. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Royi Rassin, Yaron Fairstein, Oren Kalinsky, Guy Kushilevitz, Nachshon Cohen, Alexander Libov, Yoav Goldberg
EMNLP7
2024 Extracting automata from recurrent neural networks using queries and counterexamples (extended version)
Gail Weiss, Yoav Goldberg, Eran Yahav
Mach. Learn.2
2024 Correction to: Extracting automata from recurrent neural networks using queries and counterexamples (extended version)
Gail Weiss, Yoav Goldberg, Eran Yahav
Mach. Learn.2
2024 NoviCode: Generating Programs from Natural Language Utterances by Novices
abstract
Abstract Current Text-to-Code models demonstrate impressive capabilities in generating executable code from natural language snippets. However, current studies focus on technical instructions and programmer-oriented language, and it is an open question whether these models can effectively translate natural language descriptions given by non-technical users and express complex goals, to an executable program that contains an intricate flow—composed of API access and control structures as loops, conditions, and sequences. To unlock the challenge of generating a complete program from a plain non-technical description we present NoviCode, a novel NL Programming task, which takes as input an API and a natural language description by a novice non-programmer, and provides an executable program as output. To assess the efficacy of models on this task, we provide a novel benchmark accompanied by test suites wherein the generated program code is assessed not according to their form, but according to their functional execution. Our experiments show that, first, NoviCode is indeed a challenging task in the code synthesis domain, and that generating complex code from non-technical instructions goes beyond the current Text-to-Code paradigm. Second, we show that a novel approach wherein we align the NL utterances with the compositional hierarchical structure of the code, greatly enhances the performance of LLMs on this task, compared with the end-to-end Text-to-Code counterparts.
Asaf Achi Mordechai, Yoav Goldberg, Reut Tsarfaty
Trans. Assoc. Comput. Linguistics2
2023 Conjunct Resolution in the Face of Verbal Omissions
abstract
Verbal omissions are complex syntactic phenomena in VP coordination structures.They occur when verbs and (some of) their arguments are omitted from subsequent clauses after being explicitly stated in an initial clause.Recovering these omitted elements is necessary for accurate interpretation of the sentence, and while humans easily and intuitively fill in the missing information, state-of-the-art models continue to struggle with this task.Previous work is limited to small-scale datasets, synthetic data creation methods, and to resolution methods in the dependency-graph level.In this work we propose a conjunct resolution task that operates directly on the text and makes use of a split-andrephrase paradigm in order to recover the missing elements in the coordination structure.To this end, we first formulate a pragmatic framework of verbal omissions which describes the different types of omissions, and develop an automatic scalable collection method.Based on this method, we curate a large dataset, containing over 10K examples of naturally-occurring verbal omissions with crowd-sourced annotations of the resolved conjuncts.We train various neural baselines for this task, and show that while our best method obtains decent performance, it leaves ample space for improvement.We propose our dataset, metrics and models as a starting point for future research on this topic.
Royi Rassin, Yoav Goldberg, Reut Tsarfaty
ACL (1)2
2023 Linear Guardedness and its Implications
abstract
Methods for erasing human-interpretable concepts from neural representations that assume linearity have been found to be tractable and useful.However, the impact of this removal on the behavior of downstream classifiers trained on the modified representations is not fully understood.In this work, we formally define the notion of log-linear guardedness as the inability of an adversary to predict the concept directly from the representation, and study its implications.We show that, in the binary case, under certain assumptions, a downstream log-linear model cannot recover the erased concept.However, we demonstrate that a multiclass log-linear model can be constructed that indirectly recovers the concept in some cases, pointing to the inherent limitations of log-linear guardedness as a downstream bias mitigation technique.These findings shed light on the theoretical limitations of linear erasure methods and highlight the need for further research on the connections between intrinsic and extrinsic bias in neural models.
Shauli Ravfogel, Yoav Goldberg, Ryan Cotterell
ACL (1)2
2023 Understanding Transformer Memorization Recall Through Idioms
abstract
To produce accurate predictions, language models (LMs) must balance between generalization and memorization.Yet, little is known about the mechanism by which transformer LMs employ their memorization capacity.When does a model decide to output a memorized phrase, and how is this phrase then retrieved from memory?In this work, we offer the first methodological framework for probing and characterizing recall of memorized sequences in transformer LMs.First, we lay out criteria for detecting model inputs that trigger memory recall, and propose idioms as inputs that typically fulfill these criteria.Next, we construct a dataset of English idioms and use it to compare model behavior on memorized vs. non-memorized inputs.Specifically, we analyze the internal prediction construction process by interpreting the model's hidden representations as a gradual refinement of the output probability distribution.We find that across different model sizes and architectures, memorized predictions are a two-step process: early layers promote the predicted token to the top of the output distribution, and upper layers increase model confidence.This suggests that memorized information is stored and retrieved in the early layers of the network.Last, we demonstrate the utility of our methodology beyond idioms in memorized factual statements.Overall, our work makes a first step towards understanding memory recall, and provides a methodological basis for future studies of transformer memorization.1
Adi Haviv, Ido Cohen 0002, Jacob Gidron, Roei Schuster, Yoav Goldberg, Mor Geva
EACL5
2023 LingMess: Linguistically Informed Multi Expert Scorers for Coreference Resolution
abstract
Current state-of-the-art coreference systems are based on a single pairwise scoring component, which assigns to each pair of mention spans a score reflecting their tendency to corefer to each other.We observe that different kinds of mention pairs require different information sources to assess their score.We present LINGMESS, a linguistically motivated categorization of mention-pairs into 6 types of coreference decisions and learn a dedicated trainable scoring function for each category.This significantly improves the accuracy of the pairwise scorer as well as of the overall coreference performance on the English Ontonotes coreference corpus and 5 additional datasets.1
Shon Otmazgin, Arie Cattan, Yoav Goldberg
EACL3
2023 Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks
abstract
Data contamination has become prevalent and challenging with the rise of models pretrained on large automatically-crawled corpora.For closed models, the training data becomes a trade secret, and even for open models, it is not trivial to detect contamination.Strategies such as leaderboards with hidden answers, or using test data which is guaranteed to be unseen, are expensive and become fragile with time.Assuming that all relevant actors value clean test data and will cooperate to mitigate data contamination, what can be done?We propose three strategies that can make a difference: (1) Test data made public should be encrypted with a public key and licensed to disallow derivative distribution; (2) demand training exclusion controls from closed API holders, and protect your test data by refusing to evaluate without them; (3) avoid data which appears with its solution on the internet, and release the web-page context of internet-derived data along with the data.These strategies are practical and can be effective in preventing data contamination.
Alon Jacovi, Avi Caciularu, Omer Goldman, Yoav Goldberg
EMNLP4
2023 Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
abstract
Text-conditioned image generation models often generate incorrect associations between entities and their visual attributes. This reflects an impaired mapping between linguistic binding of entities and modifiers in the prompt and visual binding of the corresponding elements in the generated image. As one example, a query like ``a pink sunflower and a yellow flamingo'' may incorrectly produce an image of a yellow sunflower and a pink flamingo. To remedy this issue, we propose SynGen, an approach which first syntactically analyses the prompt to identify entities and their modifiers, and then uses a novel loss function that encourages the cross-attention maps to agree with the linguistic binding reflected by the syntax. Specifically, we encourage large overlap between attention maps of entities and their modifiers, and small overlap with other entities and modifier words. The loss is optimized during inference, without retraining or fine-tuning the model. Human evaluation on three datasets, including one new and challenging set, demonstrate significant improvements of SynGen compared with current state of the art methods. This work highlights how making use of sentence structure during inference can efficiently and substantially improve the faithfulness of text-to-image generation.
Royi Rassin, Eran Hirsch, Daniel Glickman, Shauli Ravfogel, Yoav Goldberg, Gal Chechik
NeurIPS5
2023 Extending the boundaries of cancer therapeutic complexity with literature text mining
Danna Niezni, Hillel Taub-Tabib, Yuval Harris, Hagit Sason, Yakir Amrusi, Dana Meron Azagury, Maytal Avrashami, Shaked Launer-Wachs, Jonathan Borchardt, Milo Kusold, Aryeh Tiktinsky, Tom Hope, Yoav Goldberg, Yosi Shamay
Artif. Intell. Medicine13
2023 Diagnosing AI Explanation Methods with Folk Concepts of Behavior
Alon Jacovi, Jasmijn Bastings, Sebastian Gehrmann, Yoav Goldberg, Katja Filippova
J. Artif. Intell. Res.4
2023 From centralized to ad-hoc knowledge base construction for hypotheses generation
Shaked Launer-Wachs, Hillel Taub-Tabib, Jennie Tokarev Madem, Orr Bar-Natan, Yoav Goldberg, Yosi Shamay
J. Biomed. Informatics5
2022 Large Scale Substitution-based Word Sense Induction
abstract
We present a word-sense induction method based on pre-trained masked language models (MLMs), which can cheaply scale to large vocabularies and large corpora.The result is a corpus which is sense-tagged according to a corpus-derived sense inventory and where each sense is associated with indicative words.Evaluation on English Wikipedia that was sense-tagged using our method shows that both the induced senses, and the per-instance sense assignment, are of high quality even compared to WSD methods, such as Babelfy.Furthermore, by training a static word embeddings algorithm on the sense-tagged corpus, we obtain high-quality static senseful embeddings.These outperform existing senseful embeddings methods on the WiC dataset and on a new outlier detection dataset we developed.The data driven nature of the algorithm allows to induce corpora-specific senses, which may not appear in standard sense inventories, as we demonstrate using a case study on the scientific domain.
Matan Eyal, Shoval Sadde, Hillel Taub-Tabib, Yoav Goldberg
ACL (1)4
2022 Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
abstract
Transformer-based language models (LMs) are at the core of modern NLP, but their internal prediction construction process is opaque and largely not understood.In this work, we make a substantial step towards unveiling this underlying prediction process, by reverseengineering the operation of the feed-forward network (FFN) layers, one of the building blocks of transformer models.We view the token representation as a changing distribution over the vocabulary, and the output from each FFN layer as an additive update to that distribution.Then, we analyze the FFN updates in the vocabulary space, showing that each update can be decomposed to sub-updates corresponding to single FFN parameter vectors, each promoting concepts that are often human-interpretable.We then leverage these findings for controlling LM predictions, where we reduce the toxicity of GPT2 by almost 50%, and for improving computation efficiency with a simple early exit rule, saving 20% of computation on average. 1
Mor Geva, Avi Caciularu, Kevin Ro Wang, Yoav Goldberg
EMNLP4
2022 Adversarial Concept Erasure in Kernel Space
abstract
The representation space of neural models for textual data emerges in an unsupervised manner during training.Understanding how those representations encode human-interpretable concepts is a fundamental problem.One prominent approach for the identification of concepts in neural representations is searching for a linear subspace whose erasure prevents the prediction of the concept from the representations.However, while many linear erasure algorithms are tractable and interpretable, neural networks do not necessarily represent concepts in a linear manner.To identify non-linearly encoded concepts, we propose a kernelization of a linear minimax game for concept erasure.We demonstrate that it is possible to prevent specific nonlinear adversaries from predicting the concept.However, the protection does not transfer to different nonlinear adversaries.Therefore, exhaustively erasing a non-linearly encoded concept remains an open problem.
Shauli Ravfogel, Francisco Vargas 0001, Yoav Goldberg, Ryan Cotterell
EMNLP3
2022 Linear Adversarial Concept Erasure
abstract
Modern neural models trained on textual data rely on pre-trained representations that emerge without direct supervision. As these representations are increasingly being used in real-world applications, the inability to control their content becomes an increasingly important problem. In this work, we formulate the problem of identifying a linear subspace that corresponds to a given concept, and removing it from the representation. We formulate this problem as a constrained, linear minimax game, and show that existing solutions are generally not optimal for this task. We derive a closed-form solution for certain objectives, and propose a convex relaxation that works well for others. When evaluated in the context of binary gender removal, the method recovers a low-dimensional subspace whose removal mitigates bias by intrinsic and extrinsic evaluation. Surprisingly, we show that the method—despite being linear—is highly expressive, effectively mitigating bias in the output layers of deep, nonlinear classifiers while maintaining tractability and interpretability.
Shauli Ravfogel, Michael Twiton, Yoav Goldberg, Ryan Cotterell
ICML3
2022 A Dataset for N-ary Relation Extraction of Drug Combinations
abstract
Aryeh Tiktinsky, Vijay Viswanathan, Danna Niezni, Dana Meron Azagury, Yosi Shamay, Hillel Taub-Tabib, Tom Hope, Yoav Goldberg. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Aryeh Tiktinsky, Vijay Viswanathan 0002, Danna Niezni, Dana Meron Azagury, Yosi Shamay, Hillel Taub-Tabib, Tom Hope, Yoav Goldberg
NAACL-HLT8
2022 Text-based NP Enrichment
abstract
Abstract Understanding the relations between entities denoted by NPs in a text is a critical part of human-like natural language understanding. However, only a fraction of such relations is covered by standard NLP tasks and benchmarks nowadays. In this work, we propose a novel task termed text-based NP enrichment (TNE), in which we aim to enrich each NP in a text with all the preposition-mediated relations—either explicit or implicit—that hold between it and other NPs in the text. The relations are represented as triplets, each denoted by two NPs related via a preposition. Humans recover such relations seamlessly, while current state-of-the-art models struggle with them due to the implicit nature of the problem. We build the first large-scale dataset for the problem, provide the formal framing and scope of annotation, analyze the data, and report the results of fine-tuned language models on the task, demonstrating the challenge it poses to current technology. A webpage with a data-exploration UI, a demo, and links to the code, models, and leaderboard, to foster further research into this challenging problem can be found at: yanaiela.github.io/TNE/.
Yanai Elazar, Victoria Basmova, Yoav Goldberg, Reut Tsarfaty
Trans. Assoc. Comput. Linguistics3
2021 Including Signed Languages in Natural Language Processing
abstract
Kayo Yin, Amit Moryossef, Julie Hochgesang, Yoav Goldberg, Malihe Alikhani. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Kayo Yin, Amit Moryossef, Julie Hochgesang, Yoav Goldberg, Malihe Alikhani
ACL/IJCNLP (1)4
2021 Counterfactual Interventions Reveal the Causal Effect of Relative Clause Representations on Agreement Prediction
abstract
When language models process syntactically complex sentences, do they use their representations of syntax in a manner that is consistent with the grammar of the language?We propose AlterRep, an intervention-based method to address this question.For any linguistic feature of a given sentence, AlterRep generates counterfactual representations by altering how the feature is encoded, while leaving intact all other aspects of the original representation.By measuring the change in a model's word prediction behavior when these counterfactual representations are substituted for the original ones, we can draw conclusions about the causal effect of the linguistic feature in question on the model's behavior.We apply this method to study how BERT models of different sizes process relative clauses (RCs).We find that BERT variants use RC boundary information during word prediction in a manner that is consistent with the rules of English grammar; this RC boundary information generalizes to a considerable extent across different RC types, suggesting that BERT represents RCs as an abstract linguistic category.
Shauli Ravfogel, Grusha Prasad, Tal Linzen, Yoav Goldberg
CoNLL4
2021 Bootstrapping Relation Extractors using Syntactic Search by Examples
abstract
The advent of neural-networks in NLP brought with it substantial improvements in supervised relation extraction.However, obtaining a sufficient quantity of training data remains a key challenge.In this work we propose a process for bootstrapping training datasets which can be performed quickly by non-NLP-experts.We take advantage of search engines over syntactic-graphs (Such as Shlain et al. ( 2020)) which expose a friendly by-example syntax.We use these to obtain positive examples by searching for sentences that are syntactically similar to user input examples.We apply this technique to relations from TACRED and Do-cRED and show that the resulting models are competitive with models trained on manually annotated data and on data obtained from distant supervision.The models also outperform models trained using NLG data augmentation techniques.Extending the search-based approach with the NLG method further improves the results.
Matan Eyal, Asaf Amrami, Hillel Taub-Tabib, Yoav Goldberg
EACL4
2021 Scalable Evaluation and Improvement of Document Set Expansion via Neural Positive-Unlabeled Learning
abstract
We consider the situation in which a user has collected a small set of documents on a cohesive topic, and they want to retrieve additional documents on this topic from a large collection.Information Retrieval (IR) solutions treat the document set as a query, and look for similar documents in the collection.We propose to extend the IR approach by treating the problem as an instance of positive-unlabeled (PU) learning-i.e., learning binary classifiers from only positive (the query documents) and unlabeled (the results of the IR engine) data.Utilizing PU learning for text with big neural networks is a largely unexplored field.We discuss various challenges in applying PU learning to the setting, showing that the standard implementations of state-of-the-art PU solutions fail.We propose solutions for each of the challenges and empirically validate them with ablation tests.We demonstrate the effectiveness of the new method using a series of experiments of retrieving PubMed abstracts adhering to fine-grained topics, showing improvements over the common IR solution and other baselines.
Alon Jacovi, Gang Niu 0001, Yoav Goldberg, Masashi Sugiyama
EACL3
2021 Back to Square One: Artifact Detection, Training and Commonsense Disentanglement in the Winograd Schema
abstract
The Winograd Schema (WS) has been proposed as a test for measuring commonsense capabilities of models.Recently, pre-trained language model-based approaches have boosted performance on some WS benchmarks but the source of improvement is still not clear.This paper suggests that the apparent progress on WS may not necessarily reflect progress in commonsense reasoning.To support this claim, we first show that the current evaluation method of WS is sub-optimal and propose a modification that uses twin sentences for evaluation.We also propose two new baselines that indicate the existence of artifacts in WS benchmarks.We then develop a method for evaluating WS-like sentences in a zero-shot setting to account for the commonsense reasoning abilities acquired during the pretraining and observe that popular language models perform randomly in this setting when using our more strict evaluation.We conclude that the observed progress is mostly due to the use of supervision in training WS models, which is not likely to successfully support all the required commonsense reasoning skills and knowledge.1
Yanai Elazar, Hongming Zhang 0009, Yoav Goldberg, Dan Roth 0001
EMNLP (1)3
2021 Contrastive Explanations for Model Interpretability
abstract
Contrastive explanations clarify why an event occurred in contrast to another.They are inherently intuitive to humans to both produce and comprehend.We propose a method to produce contrastive explanations in the latent space, via a projection of the input representation, such that only the features that differentiate two potential decisions are captured.Our modification allows model behavior to consider only contrastive reasoning, and uncover which aspects of the input are useful for and against particular decisions.Additionally, for a given input feature, our contrastive explanations can answer for which label, and against which alternative label, is the feature useful.We produce contrastive explanations via both highlevel abstract concept attribution and low-level input token/span attribution for two NLP classification benchmarks.Our findings demonstrate the ability of label-contrastive explanations to provide fine-grained interpretability of model decisions.1
Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi 0001, Yoav Goldberg
EMNLP (1)6
2021 Effects of Parameter Norm Growth During Transformer Training: Inductive Bias from Gradient Descent
abstract
The capacity of neural networks like the widely adopted transformer is known to be very high.Evidence is emerging that they learn successfully due to inductive bias in the training routine, typically a variant of gradient descent (GD).To better understand this bias, we study the tendency for transformer parameters to grow in magnitude (ℓ 2 norm) during training, and its implications for the emergent representations within self attention layers.Empirically, we document norm growth in the training of transformer language models, including T5 during its pretraining.As the parameters grow in magnitude, we prove that the network approximates a discretized network with saturated activation functions.Such "saturated" networks are known to have a reduced capacity compared to the full network family that can be described in terms of formal languages and automata.Our results suggest saturation is a new characterization of an inductive bias implicit in GD of particular interest for NLP.We leverage the emergent discrete structure in a saturated transformer to analyze the role of different attention heads, finding that some focus locally on a small number of positions, while other heads compute global averages, allowing counting.We believe understanding the interplay between these two capabilities may shed further light on the structure of computation within large transformers.
William Merrill, Vivek Ramanujan, Yoav Goldberg, Roy Schwartz 0001, Noah A. Smith
EMNLP (1)3
2021 Asking It All: Generating Contextualized Questions for any Semantic Role
abstract
Asking questions about a situation is an inherent step towards understanding it.To this end, we introduce the task of role question generation, which, given a predicate mention and a passage, requires producing a set of questions asking about all possible semantic roles of the predicate.We develop a two-stage model for this task, which first produces a contextindependent question prototype for each role and then revises it to be contextually appropriate for the passage.Unlike most existing approaches to question generation, our approach does not require conditioning on existing answers in the text.Instead, we condition on the type of information to inquire about, regardless of whether the answer appears explicitly in the text, could be inferred from it, or should be sought elsewhere.Our evaluation demonstrates that we generate diverse and well-formed questions for a large, broadcoverage ontology of predicates and roles.
Valentina Pyatkin, Paul Roit, Julian Michael, Yoav Goldberg, Reut Tsarfaty, Ido Dagan
EMNLP (1)4
2021 Thinking Like Transformers
abstract
What is the computational model behind a Transformer? Where recurrent neural networks have direct parallels in finite state machines, allowing clear discussion and thought around architecture variants or trained models, Transformers have no such familiar parallel. In this paper we aim to change that, proposing a computational model for the transformer-encoder in the form of a programming language. We map the basic components of a transformer-encoder—attention and feed-forward computation—into simple primitives, around which we form a programming language: the Restricted Access Sequence Processing Language (RASP). We show how RASP can be used to program solutions to tasks that could conceivably be learned by a Transformer, and how a Transformer can be trained to mimic a RASP solution. In particular, we provide RASP programs for histograms, sorting, and Dyck-languages. We further use our model to relate their difficulty in terms of the number of required layers and attention heads: analyzing a RASP program implies a maximum number of heads and layers necessary to encode a task in a transformer. Finally, we see how insights gained from our abstraction might be used to explain phenomena seen in recent works.
Gail Weiss, Yoav Goldberg, Eran Yahav
ICML2
2021 Does BERT Pretrained on Clinical Notes Reveal Sensitive Data?
abstract
Eric Lehman, Sarthak Jain, Karl Pichotta, Yoav Goldberg, Byron Wallace. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Eric P. Lehman, Karl Pichotta, Yoav Goldberg, Byron C. Wallace
NAACL-HLT4
2021 Ab Antiquo: Neural Proto-language Reconstruction
abstract
Historical linguists have identified regularities in the process of historic sound change.The comparative method utilizes those regularities to reconstruct proto-words based on observed forms in daughter languages.Can this process be efficiently automated?We address the task of proto-word reconstruction, in which the model is exposed to cognates in contemporary daughter languages, and has to predict the proto word in the ancestor language.We provide a novel dataset for this task, encompassing over 8,000 comparative entries, and show that neural sequence models outperform conventional methods applied to this task so far.Error analysis reveals a variability in the ability of neural model to capture different phonological changes, correlating with the complexity of the changes.Analysis of learned embeddings reveals the models learn phonologically meaningful generalizations, corresponding to well-attested phonological shifts documented by historical linguistics.
Carlo Meloni, Shauli Ravfogel, Yoav Goldberg
NAACL-HLT3
2021 Measuring and Improving Consistency in Pretrained Language Models
abstract
Abstract Consistency of a model—that is, the invariance of its behavior under meaning-preserving alternations in its input—is a highly desirable property in natural language processing. In this paper we study the question: Are Pretrained Language Models (PLMs) consistent with respect to factual knowledge? To this end, we create ParaRel🤘, a high-quality resource of cloze-style query English paraphrases. It contains a total of 328 paraphrases for 38 relations. Using ParaRel🤘, we show that the consistency of all PLMs we experiment with is poor— though with high variance between relations. Our analysis of the representational spaces of PLMs suggests that they have a poor structure and are currently not suitable for representing knowledge robustly. Finally, we propose a method for improving model consistency and experimentally demonstrate its effectiveness.1
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schütze, Yoav Goldberg
Trans. Assoc. Comput. Linguistics7
2021 Erratum: Measuring and Improving Consistency in Pretrained Language Models
abstract
Abstract During production of this paper, an error was introduced to the formula on the bottom of the right column of page 1020. In the last two terms of the formula, the n and m subscripts were swapped. The correct formula is:Lc=∑n=1k∑m=n+1kDKL(Qnri∥Qmri)+DKL(Qmri∥Qnri)The paper has been updated.
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schütze, Yoav Goldberg
Trans. Assoc. Comput. Linguistics7
2021 Amnesic Probing: Behavioral Explanation With Amnesic Counterfactuals
abstract
Abstract A growing body of work makes use of probing in order to investigate the working of neural models, often considered black boxes. Recently, an ongoing debate emerged surrounding the limitations of the probing paradigm. In this work, we point out the inability to infer behavioral conclusions from probing results, and offer an alternative method that focuses on how the information is being used, rather than on what information is encoded. Our method, Amnesic Probing, follows the intuition that the utility of a property for a given task can be assessed by measuring the influence of a causal intervention that removes it from the representation. Equipped with this new analysis tool, we can ask questions that were not possible before, for example, is part-of-speech information important for word prediction? We perform a series of analyses on BERT to answer these types of questions. Our findings demonstrate that conventional probing performance is not correlated to task importance, and we call for increased scrutiny of claims that draw behavioral or causal conclusions from probing results.1
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, Yoav Goldberg
Trans. Assoc. Comput. Linguistics4
2021 Aligning Faithful Interpretations with their Social Attribution
abstract
Abstract We find that the requirement of model interpretations to be faithful is vague and incomplete. With interpretation by textual highlights as a case study, we present several failure cases. Borrowing concepts from social science, we identify that the problem is a misalignment between the causal chain of decisions (causal attribution) and the attribution of human behavior to the interpretation (social attribution). We reformulate faithfulness as an accurate attribution of causality to the model, and introduce the concept of aligned faithfulness: faithful causal chains that are aligned with their expected social behavior. The two steps of causal attribution and social attribution together complete the process of explaining behavior. With this formalization, we characterize various failures of misaligned faithful highlight interpretations, and propose an alternative causal chain to remedy the issues. Finally, we implement highlight explanations of the proposed causal format using contrastive explanations.
Alon Jacovi, Yoav Goldberg
Trans. Assoc. Comput. Linguistics2
2021 Provable Limitations of Acquiring Meaning from Ungrounded Form: What Will Future Language Models Understand?
abstract
Abstract Language models trained on billions of tokens have recently led to unprecedented results on many NLP tasks. This success raises the question of whether, in principle, a system can ever “understand” raw text without access to some form of grounding. We formally investigate the abilities of ungrounded systems to acquire meaning. Our analysis focuses on the role of “assertions”: textual contexts that provide indirect clues about the underlying semantics. We study whether assertions enable a system to emulate representations preserving semantic relations like equivalence. We find that assertions enable semantic emulation of languages that satisfy a strong notion of semantic transparency. However, for classes of languages where the same expression can take different values in different contexts, we show that emulation can become uncomputable. Finally, we discuss differences between our formal model and natural language, exploring how our results generalize to a modal setting and other semantic relations. Together, our results suggest that assertions in code or language do not provide sufficient signal to fully emulate semantic representations. We formalize ways in which ungrounded language models appear to be fundamentally limited in their ability to “understand”.
William Merrill, Yoav Goldberg, Roy Schwartz 0001, Noah A. Smith
Trans. Assoc. Comput. Linguistics2
2021 Revisiting Few-shot Relation Classification: Evaluation Data and Classification Schemes
abstract
We explore few-shot learning (FSL) for relation classification (RC). Focusing on the realistic scenario of FSL, in which a test instance might not belong to any of the target categories (none-of-the-above, [NOTA]), we first revisit the recent popular dataset structure for FSL, pointing out its unrealistic data distribution. To remedy this, we propose a novel methodology for deriving more realistic few-shot test data from available datasets for supervised RC, and apply it to the TACRED dataset. This yields a new challenging benchmark for FSL-RC, on which state of the art models show poor performance. Next, we analyze classification schemes within the popular embedding-based nearest-neighbor approach for FSL, with respect to constraints they impose on the embedding space. Triggered by this analysis, we propose a novel classification scheme in which the NOTA category is represented as learned vectors, shown empirically to be an appealing option for FSL.
Ofer Sabo, Yanai Elazar, Yoav Goldberg, Ido Dagan
Trans. Assoc. Comput. Linguistics3
2020 Unsupervised Domain Clusters in Pretrained Language Models
abstract
The notion of "in-domain data" in NLP is often over-simplistic and vague, as textual data varies in many nuanced linguistic aspects such as topic, style or level of formality.In addition, domain labels are many times unavailable, making it challenging to build domainspecific systems.We show that massive pretrained language models implicitly learn sentence representations that cluster by domains without supervision -suggesting a simple datadriven definition of domains in textual data.We harness this property and propose domain data selection methods based on such models, which require only a small set of in-domain monolingual data.We evaluate our data selection methods for neural machine translation across five diverse domains, where they outperform an established approach as measured by both BLEU and by precision and recall of sentence selection with respect to an oracle.
Roee Aharoni, Yoav Goldberg
ACL2
2020 Simple, Interpretable and Stable Method for Detecting Words with Usage Change across Corpora
abstract
The problem of comparing two bodies of text and searching for words that differ in their usage between them arises often in digital humanities and computational social science.This is commonly approached by training word embeddings on each corpus, aligning the vector spaces, and looking for words whose cosine distance in the aligned space is large.However, these methods often require extensive filtering of the vocabulary to perform well, and-as we show in this work-result in unstable, and hence less reliable, results.We propose an alternative approach that does not use vector space alignment, and instead considers the neighbors of each word.The method is simple, interpretable and stable.We demonstrate its effectiveness in 9 different setups, considering different corpus splitting criteria (age, gender and profession of tweet authors, time of tweet) and different languages (English, French and Hebrew).
Hila Gonen, Ganesh Jawahar, Djamé Seddah, Yoav Goldberg
ACL4
2020 Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness?
abstract
With the growing popularity of deep-learning based NLP models, comes a need for interpretable systems.But what is interpretability, and what constitutes a high-quality interpretation?In this opinion piece we reflect on the current state of interpretability evaluation research.We call for more clearly differentiating between different desired criteria an interpretation should satisfy, and focus on the faithfulness criteria.We survey the literature with respect to faithfulness evaluation, and arrange the current approaches around three assumptions, providing an explicit form to how faithfulness is "defined" by the community.We provide concrete guidelines on how evaluation of interpretation methods should and should not be conducted.Finally, we claim that the current binary definition for faithfulness sets a potentially unrealistic bar for being considered faithful.We call for discarding the binary notion of faithfulness in favor of a more graded one, which we believe will be of greater practical utility.
Alon Jacovi, Yoav Goldberg
ACL2
2020 A Two-Stage Masked LM Method for Term Set Expansion
abstract
We tackle the task of Term Set Expansion (TSE): given a small seed set of example terms from a semantic class, finding more members of that class.The task is of great practical utility, and also of theoretical utility as it requires generalization from few examples.Previous approaches to the TSE task can be characterized as either distributional or pattern-based.We harness the power of neural masked language models (MLM) and propose a novel TSE algorithm, which combines the pattern-based and distributional approaches.Due to the small size of the seed set, fine-tuning methods are not effective, calling for more creative use of the MLM.The gist of the idea is to use the MLM to first mine for informative patterns with respect to the seed set, and then to obtain more members of the seed class by generalizing these patterns.Our method outperforms stateof-the-art TSE algorithms.
Guy Kushilevitz, Shaul Markovitch, Yoav Goldberg
ACL3
2020 A Formal Hierarchy of RNN Architectures
abstract
We develop a formal hierarchy of the expressive capacity of RNN architectures.The hierarchy is based on two formal properties: space complexity, which measures the RNN's memory, and rational recurrence, defined as whether the recurrent update can be described by a weighted finite-state machine.We place several RNN variants within this hierarchy.For example, we prove the LSTM is not rational, which formally separates it from the related QRNN (Bradbury et al., 2016).We also show how these models' expressive capacity is expanded by stacking multiple layers or composing them with different pooling functions.Our results build on the theory of "saturated" RNNs (Merrill, 2019).While formally extending these findings to unsaturated RNNs is left to future work, we hypothesize that the practical learnable capacity of unsaturated RNNs obeys a similar hierarchy.Experimental findings from training unsaturated networks on formal languages support this conjecture.We report updated experiments in Appendix H.
William Merrill, Gail Weiss, Yoav Goldberg, Roy Schwartz 0001, Noah A. Smith, Eran Yahav
ACL3
2020 Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection
abstract
The ability to control for the kinds of information encoded in neural representation has a variety of use cases, especially in light of the challenge of interpreting these models.We present Iterative Null-space Projection (INLP), a novel method for removing information from neural representations.Our method is based on repeated training of linear classifiers that predict a certain property we aim to remove, followed by projection of the representations on their null-space.By doing so, the classifiers become oblivious to that target property, making it hard to linearly separate the data according to it.While applicable for multiple uses, we evaluate our method on bias and fairness use-cases, and show that our method is able to mitigate bias in word embeddings, as well as to increase fairness in a setting of multi-class classification.
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, Yoav Goldberg
ACL5
2020 Facts2Story: Controlling Text Generation by Key Facts
abstract
Recent advancements in self-attention neural network architectures have raised the bar for openended text generation.Yet, while current methods are capable of producing a coherent text which is several hundred words long, attaining control over the content that is being generated-as well as evaluating it-are still open questions.We propose a controlled generation task which is based on expanding a sequence of facts, expressed in natural language, into a longer narrative.We introduce human-based evaluation metrics for this task, as well as a method for deriving a large training dataset.We evaluate three methods on this task, based on fine-tuning pre-trained models.We show that while auto-regressive, unidirectional Language Models such as GPT2 produce better fluency, they struggle to adhere to the requested facts.We propose a plan-andcloze model (using fine-tuned XLNet) which produces competitive fluency while adhering to the requested content.
Eyal Orbach, Yoav Goldberg
COLING2
2020 Exposing Shallow Heuristics of Relation Extraction Models with Challenge Data
abstract
The process of collecting and annotating training data may introduce distribution artifacts which may limit the ability of models to learn correct generalization behavior.We identify failure modes of SOTA relation extraction (RE) models trained on TACRED, which we attribute to limitations in the data annotation process.We collect and annotate a challengeset we call Challenging RE (CRE), based on naturally occurring corpus examples, to benchmark this behavior.Our experiments with four state-of-the-art RE models show that they have indeed adopted shallow heuristics that do not generalize to the challenge-set data.Further, we find that alternative question answering modeling performs significantly better than the SOTA models on the challenge-set, despite worse overall TACRED performance.By adding some of the challenge data as training examples, the performance of the model improves.Finally, we provide concrete suggestion on how to improve RE data collection to alleviate this behavior.
Shachar Rosenman, Alon Jacovi, Yoav Goldberg
EMNLP (1)3
2020 Synthesizing Control for a System with Black Box Environment, Based on Deep Learning
Simon Iosti, Doron A. Peled, Khen Aharon, Saddek Bensalem, Yoav Goldberg
ISoLA (2)5
2020 Leap-Of-Thought: Teaching Pre-Trained Models to Systematically Reason Over Implicit Knowledge
abstract
To what extent can a neural network systematically reason over symbolic facts? Evidence suggests that large pre-trained language models (LMs) acquire some reasoning capacity, but this ability is difficult to control. Recently, it has been shown that Transformer-based models succeed in consistent reasoning over explicit symbolic facts, under a "closed-world" assumption. However, in an open-domain setup, it is desirable to tap into the vast reservoir of implicit knowledge already encoded in the parameters of pre-trained LMs. In this work, we provide a first demonstration that LMs can be trained to reliably perform systematic reasoning combining both implicit, pre-trained knowledge and explicit natural language statements. To do this, we describe a procedure for automatically generating datasets that teach a model new reasoning skills, and demonstrate that models learn to effectively perform inference which involves implicit taxonomic and world knowledge, chaining and counting. Finally, we show that "teaching" models to reason generalizes beyond the training distribution: they successfully compose the usage of multiple reasoning skills in single examples. Our work paves a path towards open-domain systems that constantly improve by interacting with users who can instantly correct a model by adding simple natural language statements.
Alon Talmor, Oyvind Tafjord, Peter Clark, Yoav Goldberg, Jonathan Berant
NeurIPS4
2020 oLMpics - On what Language Model Pre-training Captures
abstract
Recent success of pre-trained language models (LMs) has spurred widespread interest in the language capabilities that they possess. However, efforts to understand whether LM representations are useful for symbolic reasoning tasks have been limited and scattered. In this work, we propose eight reasoning tasks, which conceptually require operations such as comparison, conjunction, and composition. A fundamental challenge is to understand whether the performance of a LM on a task should be attributed to the pre-trained representations or to the process of fine-tuning on the task data. To address this, we propose an evaluation protocol that includes both zero-shot evaluation (no fine-tuning), as well as comparing the learning curve of a fine-tuned LM to the learning curve of multiple controls, which paints a rich picture of the LM capabilities. Our main findings are that: (a) different LMs exhibit qualitatively different reasoning abilities, e.g., RoBERTa succeeds in reasoning tasks where BERT fails completely; (b) LMs do not reason in an abstract manner and are context-dependent, e.g., while RoBERTa can compare ages, it can do so only when the ages are in the typical range of human ages; (c) On half of our reasoning tasks all models fail completely. Our findings and infrastructure can help future work on designing new datasets, models, and objective functions for pre-training.
Alon Talmor, Yanai Elazar, Yoav Goldberg, Jonathan Berant
Trans. Assoc. Comput. Linguistics3
2020 Break It Down: A Question Understanding Benchmark
abstract
Understanding natural language questions entails the ability to break down a question into the requisite steps for computing its answer. In this work, we introduce a Question Decomposition Meaning Representation (QDMR) for questions. QDMR constitutes the ordered list of steps, expressed through natural language, that are necessary for answering a question. We develop a crowdsourcing pipeline, showing that quality QDMRs can be annotated at scale, and release the Break dataset, containing over 83K pairs of questions and their QDMRs. We demonstrate the utility of QDMR by showing that (a) it can be used to improve open-domain question answering on the HotpotQA dataset, (b) it can be deterministically converted to a pseudo-SQL formal language, which can alleviate annotation in semantic parsing applications. Last, we use Break to train a sequence-to-sequence model with copying that parses questions into QDMR structures, and show that it substantially outperforms several natural baselines.
Tomer Wolfson, Mor Geva, Ankit Gupta 0001, Yoav Goldberg, Matt Gardner 0001, Daniel Deutch, Jonathan Berant
Trans. Assoc. Comput. Linguistics4
2019 How Does Grammatical Gender Affect Noun Representations in Gender-Marking Languages?
abstract
Many natural languages assign grammatical gender also to inanimate nouns in the language.In such languages, words that relate to the gender-marked nouns are inflected to agree with the noun's gender.We show that this affects the word representations of inanimate nouns, resulting in nouns with the same gender being closer to each other than nouns with different gender.While "embedding debiasing" methods fail to remove the effect, we demonstrate that a careful application of methods that neutralize grammatical gender signals from the words' context when training word embeddings is effective in removing it.Fixing the grammatical gender bias yields a positive effect on the quality of the resulting word embeddings, both in monolingual and crosslingual settings.We note that successfully removing gender signals, while achievable, is not trivial to do and that a language-specific morphological analyzer, together with careful usage of it, are essential for achieving good results.
Hila Gonen, Yova Kementchedjhieva, Yoav Goldberg
CoNLL3
2019 Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets
abstract
Mor Geva, Yoav Goldberg, Jonathan Berant. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mor Geva, Yoav Goldberg, Jonathan Berant
EMNLP/IJCNLP (1)2
2019 Language Modeling for Code-Switching: Evaluation, Integration of Monolingual Data, and Discriminative Training
abstract
Hila Gonen, Yoav Goldberg. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Hila Gonen, Yoav Goldberg
EMNLP/IJCNLP (1)2
2019 Transfer Learning Between Related Tasks Using Expected Label Proportions
abstract
Matan Ben Noach, Yoav Goldberg. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Matan Ben Noach, Yoav Goldberg
EMNLP/IJCNLP (1)2
2019 Transfer Learning for Related Reinforcement Learning Tasks via Image-to-Image Translation
abstract
Despite the remarkable success of Deep RL in learning control policies from raw pixels, the resulting models do not generalize. We demonstrate that a trained agent fails completely when facing small visual changes, and that fine-tuning—the common transfer learning paradigm—fails to adapt to these changes, to the extent that it is faster to re-train the model from scratch. We show that by separating the visual transfer task from the control policy we achieve substantially better sample efficiency and transfer behavior, allowing an agent trained on the source task to transfer well to the target tasks. The visual mapping from the target to the source domain is performed using unaligned GANs, resulting in a control policy that can be further improved using imitation learning from imperfect demonstrations. We demonstrate the approach on synthetic visual variants of the Breakout game, as well as on transfer between subsequent levels of Road Fighter, a Nintendo car-driving game. A visualization of our approach can be seen in \url{https://youtu.be/4mnkzYyXMn4} and \url{https://youtu.be/KCGTrQi6Ogo}.
Shani Gamrian, Yoav Goldberg
ICML2
2019 Improving Quality and Efficiency in Plan-based Neural Data-to-text Generation
abstract
We follow the step-by-step approach to neural data-to-text generation we proposed in Moryossef et al. (2019), in which the generation process is divided into a text-planning stage followed by a plan-realization stage.We suggest four extensions to that framework: (1) we introduce a trainable neural planning component that can generate effective plans several orders of magnitude faster than the original planner; (2) we incorporate typing hints that improve the model's ability to deal with unseen relations and entities; (3) we introduce a verification-by-reranking stage that substantially improves the faithfulness of the resulting texts; (4) we incorporate a simple but effective referring expression generation module.These extensions result in a generation process that is faster, more fluent, and more accurate.
Amit Moryossef, Yoav Goldberg, Ido Dagan
INLG2
2019 A Little Is Enough: Circumventing Defenses For Distributed Learning
abstract
Distributed learning is central for large-scale training of deep-learning models. However, it is exposed to a security threat in which Byzantine participants can interrupt or control the learning process. Previous attack models assume that the rogue participants (a) are omniscient (know the data of all other participants), and (b) introduce large changes to the parameters. Accordingly, most defense mechanisms make a similar assumption and attempt to use statistically robust methods to identify and discard values whose reported gradients are far from the population mean. We observe that if the empirical variance between the gradients of workers is high enough, an attacker could take advantage of this and launch a non-omniscient attack that operates within the population variance. We show that the variance is indeed high enough even for simple datasets such as MNIST, allowing an attack that is not only undetected by existing defenses, but also uses their power against them, causing those defense mechanisms to consistently select the byzantine workers while discarding legitimate ones. We demonstrate our attack method works not only for preventing convergence but also for repurposing of the model behavior (``backdooring''). We show that less than 25\% of colluding workers are sufficient to degrade the accuracy of models trained on MNIST, CIFAR10 and CIFAR100 by 50\%, as well as to introduce backdoors without hurting the accuracy for MNIST and CIFAR10 datasets, but with a degradation for CIFAR100.
Gilad Baruch, Moran Baruch, Yoav Goldberg
NeurIPS3
2019 Learning Deterministic Weighted Automata with Queries and Counterexamples
abstract
We present an algorithm for reconstruction of a probabilistic deterministic finite automaton (PDFA) from a given black-box language model, such as a recurrent neural network (RNN). The algorithm is a variant of the exact-learning algorithm L*, adapted to work in a probabilistic setting under noise. The key insight of the adaptation is the use of conditional probabilities when making observations on the model, and the introduction of a variation tolerance when comparing observations. When applied to RNNs, our algorithm returns models with better or equal word error rate (WER) and normalised distributed cumulative gain (NDCG) than achieved by n-gram or weighted finite automata (WFA) approximations of the same networks. The PDFAs capture a richer class of languages than n-grams, and are guaranteed to be stochastic and deterministic -- unlike the WFAs.
Gail Weiss, Yoav Goldberg, Eran Yahav
NeurIPS2
2019 Mining fall-related information in clinical notes: Comparison of rule-based and novel word embedding-based machine learning approaches
abstract
BACKGROUND: Natural language processing (NLP) of health-related data is still an expertise demanding, and resource expensive process. We created a novel, open source rapid clinical text mining system called NimbleMiner. NimbleMiner combines several machine learning techniques (word embedding models and positive only labels learning) to facilitate the process in which a human rapidly performs text mining of clinical narratives, while being aided by the machine learning components. OBJECTIVE: This manuscript describes the general system architecture and user Interface and presents results of a case study aimed at classifying fall-related information (including fall history, fall prevention interventions, and fall risk) in homecare visit notes. METHODS: We extracted a corpus of homecare visit notes (n = 1,149,586) for 89,459 patients from a large US-based homecare agency. We used a gold standard testing dataset of 750 notes annotated by two human reviewers to compare the NimbleMiner's ability to classify documents regarding whether they contain fall-related information with a previously developed rule-based NLP system. RESULTS: NimbleMiner outperformed the rule-based system in almost all domains. The overall F- score was 85.8% compared to 81% by the rule based-system with the best performance for identifying general fall history (F = 89% vs. F = 85.1% rule-based), followed by fall risk (F = 87% vs. F = 78.7% rule-based), fall prevention interventions (F = 88.1% vs. F = 78.2% rule-based) and fall within 2 days of the note date (F = 83.1% vs. F = 80.6% rule-based). The rule-based system achieved slightly better performance for fall within 2 weeks of the note date (F = 81.9% vs. F = 84% rule-based). DISCUSSION & CONCLUSIONS: NimbleMiner outperformed other systems aimed at fall information classification, including our previously developed rule-based approach. These promising results indicate that clinical text mining can be implemented without the need for large labeled datasets necessary for other types of machine learning. This is critical for domains with little NLP developments, like nursing or allied health professions.
Maxim Topaz, Ludmila Murga, Katherine M. Gaddis, Margaret V. McDonald, Ofrit Bar-Bachar, Yoav Goldberg, Kathryn H. Bowles
J. Biomed. Informatics6
2019 Where's My Head? Definition, Dataset and Models for Numeric Fused-Heads Identification and Resolution
abstract
We provide the first computational treatment of fused-heads constructions (FHs), focusing on the numeric fused-heads (NFHs). FHs constructions are noun phrases in which the head noun is missing and is said to be “fused” with its dependent modifier. This missing information is implicit and is important for sentence understanding. The missing references are easily filled in by humans but pose a challenge for computational models. We formulate the handling of FHs as a two stages process: Identification of the FH construction and resolution of the missing head. We explore the NFH phenomena in large corpora of English text and create (1) a data set and a highly accurate method for NFH identification; (2) a 10k examples (1 M tokens) crowd-sourced data set of NFH resolution; and (3) a neural baseline for the NFH resolution task. We release our code and data set, to foster further research into this challenging problem.
Yanai Elazar, Yoav Goldberg
Trans. Assoc. Comput. Linguistics2
2018 Word Sense Induction with Neural biLM and Symmetric Patterns
abstract
An established method for Word Sense Induction (WSI) uses a language model to predict probable substitutes for target words, and induces senses by clustering these resulting substitute vectors.We replace the ngram-based language model (LM) with a recurrent one.Beyond being more accurate, the use of the recurrent LM allows us to effectively query it in a creative way, using what we call dynamic symmetric patterns.The combination of the RNN-LM and the dynamic symmetric patterns results in strong substitute vectors for WSI, allowing to surpass the current state-of-the-art on the SemEval 2013 WSI shared task by a large margin.
Asaf Amrami, Yoav Goldberg
EMNLP2
2018 Adversarial Removal of Demographic Attributes from Text Data
abstract
Recent advances in Representation Learning and Adversarial Training seem to succeed in removing unwanted features from the learned representation.We show that demographic information of authors is encoded in-and can be recovered from-the intermediate representations learned by text-based neural classifiers.The implication is that decisions of classifiers trained on textual data are not agnostic to-and likely condition on-demographic attributes.When attempting to remove such demographic information using adversarial training, we find that while the adversarial component achieves chance-level development-set accuracy during training, a post-hoc classifier, trained on the encoded sentences from the first part, still manages to reach substantially higher classification accuracies on the same data.This behavior is consistent across several tasks, demographic properties and datasets.We explore several techniques to improve the effectiveness of the adversarial component.Our main conclusion is a cautionary one: do not rely on the adversarial training to achieve invariant representation to sensitive features.
Yanai Elazar, Yoav Goldberg
EMNLP2
2018 LaVAN: Localized and Visible Adversarial Noise
abstract
Most works on adversarial examples for deep-learning based image classifiers use noise that, while small, covers the entire image. We explore the case where the noise is allowed to be visible but confined to a small, localized patch of the image, without covering any of the main object(s) in the image. We show that it is possible to generate localized adversarial noises that cover only 2% of the pixels in the image, none of them over the main object, and that are transferable across images and locations, and successfully fool a state-of-the-art Inception v3 model with very high success rates.
Danny Karmon, Daniel Zoran, Yoav Goldberg
ICML3
2018 Extracting Automata from Recurrent Neural Networks Using Queries and Counterexamples
abstract
We present a novel algorithm that uses exact learning and abstraction to extract a deterministic finite automaton describing the state dynamics of a given trained RNN. We do this using Angluin’s \lstar algorithm as a learner and the trained RNN as an oracle. Our technique efficiently extracts accurate automata from trained RNNs, even when the state vectors are large and require fine differentiation.
Gail Weiss, Yoav Goldberg, Eran Yahav
ICML2
2018 Neural Network Methods for Natural Language Processing Yoav Goldberg (Bar Ilan University)Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, volume 37), 2017, xxii+287 pp; paperback, ISBN 9781627052986, $74.95; ebook, ISBN 9781627052955, $59.96; doi: 10.2200/S00762ED1V01Y201703HLT037
abstract
Deep learning has attracted dramatic attention in recent years, both in academia and industry. The popular term deep learning generally refers to neural network methods. Indeed, many core ideas and methods were born years ago in the era of “shallow” neural networks. However, recent development of computation resources and accumulation of data, and of course new algorithmic techniques, has enabled this branch of machine learning to dominate many areas of artificial intelligence, first for perception tasks like speech recognition and computer vision, and gradually for natural language processing (NLP) since around 2013.Natural language is an intricate object for computers to handle. Philosophical debates aside, the field of NLP has witnessed a paradigm shift from rule-based methods to statistical approaches, which have been dominant since the 1990s. Following this background, deep learning goes further down the statistical route, and gradually becomes the de facto technique of the mainstream statistical landscape.This book covers the two exciting topics of neural networks and natural language processing. More specifically, it focuses on how neural network methods are applied on natural language data. With this guideline, the structure of the book appears smoother from a neural network entry: It first lays the background of neural network methods, and then discusses the traits of natural language data, including challenges to address and sources of information that we can exploit, so that specialized neural network models introduced later are designed in ways that accommodate natural language data. On the other hand, some fundamentals in natural language processing are not covered in the book, for example, linguistic theories and backgrounds of the natural language processing tasks, and proper preparation of corpus data. Based on this structure, the book is intended for practitioners from both deep learning and natural language processing to have a common ground and a shared understanding of what has been achieved at the intersection of these two fields. NLP practitioners can become well armed with the neural network tools to work on their natural language data, whereas neural network practitioners may feel that the content of the book is a bit light, although sufficient and effective enough for an entry into working with natural language data.After the first, introductory chapter, the book is divided into four parts that roughly follow the structure of the book mentioned above.This part introduces the basic machinery of neural networks, and contains four chapters. Chapter 2 provides the background of supervised machine learning, including concepts like parameterized functions, train, test, and validation sets, training as optimization, and, in particular, the use of gradient-based methods for optimization. Readers familiar with machine learning may safely skip this chapter. The models presented in this chapter are linear and log-linear models. Their limitations are discussed in Chapter 3, which motivates the need for nonlinear models, and sets the backdrop for the introduction of feed-forward neural networks presented in Chapter 4. Finally, Chapter 5 discusses the training of neural networks. Unlike most presentations from other sources, this chapter comes with a more algorithmic than mathematical flavor, by presenting computation graph abstraction as well as related software. It also provides a handy subsection that discusses practical choices for training neural networks.As mentioned earlier, neural network practitioners may feel that the neural network content of the book is a bit light, and this part can be almost entirely skipped by these readers. However, for people coming from more traditional branches of statistical learning, Chapter 5 is still well worth reading.This part discusses the traits of natural language data, the object to which we would like to apply neural networks. There are seven chapters in this part. Chapter 6 presents a categorization of natural language classification problems and discusses the information sources that we can exploit in natural language data. Chapter 7 provides concrete examples of natural language features for solving various NLP tasks. These two chapters are probably quite dense for people coming from machine learning, and they serve to prepare them with the familiarity needed to work with natural language data. Chapter 8 is where neural networks come in; the chapter discusses how to represent textual features as inputs for neural network models. Chapter 9 describes the language modeling task and discusses the feed-forward neural language model. The neural language model also produces the byproduct of word representations, which form the subject of Chapters 10 and 11. In particular, Chapter 10 presents approaches to learning word representations, and Chapter 11 discusses the usage of word representations outside the context of neural networks, like word similarity and word analogies. Chapter 12 is an independent chapter that describes a specific feed-forward neural network architecture for the task of natural language inference.This part of the book, especially Chapter 8, which connects neural networks with natural language data, is the core of the content that distinguishes this book from other materials that cover either neural networks or natural language processing.This part is composed of five chapters that introduce the specialized architectures of convolutional neural networks (CNNs) (Chapter 13) and recurrent neural networks (RNNs) (Chapters 14–17). Chapter 13 mainly introduces 1D CNNs, which are specialized at learning ngram patterns. Chapter 14 describes the modeling of sequences and stacks with recurrent neural networks in an abstract way. This is to be made concrete by the succeeding two chapters. In Chapter 15, concrete instantiations of RNNs like the Long Short-Term Memory (LSTM) and the Gated Recurrent Unit (GRU) are described, and in Chapter 16, concrete applications of modeling with the RNN abstraction to NLP tasks are presented, including sentiment classification, grammaticality detection, part-of-speech tagging, document classification, and dependency parsing. Chapter 17 also includes concrete applications of RNNs, but these tasks involve generating natural language, which are usually modeled with a conditioned RNN language model. The most typical example of these tasks is probably machine translation.As the distribution of the chapters suggests, recurrent neural networks clearly receive more emphases. Indeed, RNNs alleviate the reliance on the Markov assumption and have the potential to model very long sequences. Their capabilities have led to breakthroughs in various sequence processing tasks, making them the celebrated models in research frontiers with proven performance.This part contains four chapters that are relatively independent. Chapter 18 presents recursive neural networks for modeling trees. The capability of modeling trees is important for natural language because of its hierarchical structure. Chapter 19 is devoted to structured prediction, because certain NLP tasks like named entity recognition can be cast in this framework. Chapter 20 discusses multi-task learning and semi-supervised learning. These approaches have not yet grown into a full-fledged stage, but are still important topics for research and offer helpful techniques for many tasks.The final chapter, Chapter 21, briefly reviews the content presented in the book, and discusses challenges that are yet to be addressed.The application of neural networks to natural language processing has revolutionized this long-standing research field, pushing forward the state of the art of many tasks. Nonetheless, the goal of equipping computers with human language capability is still far from solved, and the field continues to develop at a fast pace. This book provides valuable materials for newcomers into this exciting arena of cross-disciplinary research, by preparing relevant information of both neural networks and natural language processing. The book mainly presents mature neural network approaches to natural language processing, because it is hardly possible for a book to keep up to date with such fast development—although at 287 pages, the book is already quite long compared with other books in the synthesis lectures series, which are usually monographs of 50 to 150 pages.
Yoav Goldberg, Graeme Hirst, Yang Liu 0005, Meng Zhang 0019
Comput. Linguistics1
2017 Morphological Inflection Generation with Hard Monotonic Attention
abstract
We present a neural model for morphological inflection generation which employs a hard attention mechanism, inspired by the nearly-monotonic alignment commonly found between the characters in a word and the characters in its inflection.We evaluate the model on three previously studied morphological inflection generation datasets and show that it provides state of the art results in various setups compared to previous neural and nonneural approaches.Finally we present an analysis of the continuous representations learned by both the hard and soft attention (Bahdanau et al., 2015) models for the task, shedding some light on the features such models extract.
Roee Aharoni, Yoav Goldberg
ACL (1)2
2017 Exploring the Syntactic Abilities of RNNs with Multi-task Learning
abstract
Recent work has explored the syntactic abilities of RNNs using the subject-verb agreement task, which diagnoses sensitivity to sentence structure.RNNs performed this task well in common cases, but faltered in complex sentences (Linzen et al., 2016).We test whether these errors are due to inherent limitations of the architecture or to the relatively indirect supervision provided by most agreement dependencies in a corpus.We trained a single RNN to perform both the agreement task and an additional task, either CCG supertagging or language modeling.Multitask training led to significantly lower error rates, in particular on complex sentences, suggesting that RNNs have the ability to evolve more sophisticated syntactic representations than shown before.We also show that easily available agreement training data can improve performance on other syntactic tasks, in particular when only a limited amount of training data is available for those tasks.The multi-task paradigm can also be leveraged to inject grammatical knowledge into language models.
Émile Enguehard, Yoav Goldberg, Tal Linzen
CoNLL2
2017 A Strong Baseline for Learning Cross-Lingual Word Embeddings from Sentence Alignments
abstract
While cross-lingual word embeddings have been studied extensively in recent years, the qualitative differences between the different algorithms remain vague.We observe that whether or not an algorithm uses a particular feature set (sentence IDs) accounts for a significant performance gap among these algorithms.This feature set is also used by traditional alignment algorithms, such as IBM Model-1, which demonstrate similar performance to stateof-the-art embedding algorithms on a variety of benchmarks.Overall, we observe that different algorithmic approaches for utilizing the sentence ID feature space result in similar performance.This paper draws both empirical and theoretical parallels between the embedding and alignment literature, and suggests that adding additional sources of information, which go beyond the traditional signal of bilingual sentence-aligned corpora, may substantially improve cross-lingual word embeddings, and that future baselines should at least take such features into account.
Omer Levy, Anders Søgaard, Yoav Goldberg
EACL (1)3
2017 Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, Yoav Goldberg
ICLR (Poster)5
2017 On-the-fly Operation Batching in Dynamic Computation Graphs
abstract
Dynamic neural networks toolkits such as PyTorch, DyNet, and Chainer offer more flexibility for implementing models that cope with data of varying dimensions and structure, relative to toolkits that operate on statically declared computations (e.g., TensorFlow, CNTK, and Theano). However, existing toolkits - both static and dynamic - require that the developer organize the computations into the batches necessary for exploiting high-performance data-parallel algorithms and hardware. This batching task is generally difficult, but it becomes a major hurdle as architectures become complex. In this paper, we present an algorithm, and its implementation in the DyNet toolkit, for automatically batching operations. Developers simply write minibatch computations as aggregations of single instance computations, and the batching algorithm seamlessly executes them, on the fly, in computationally efficient batches. On a variety of tasks, we obtain throughput similar to manual batches, as well as comparable speedups over single-instance learning on architectures that are impractical to batch manually.
Graham Neubig, Yoav Goldberg, Chris Dyer
NIPS2
2017 Greedy Transition-Based Dependency Parsing with Stack LSTMs
abstract
We introduce a greedy transition-based parser that learns to represent parser states using recurrent neural networks. Our primary innovation that enables us to do this efficiently is a new control structure for sequential neural networks—the stack long short-term memory unit (LSTM). Like the conventional stack data structures used in transition-based parsers, elements can be pushed to or popped from the top of the stack in constant time, but, in addition, an LSTM maintains a continuous space embedding of the stack contents. Our model captures three facets of the parser's state: (i) unbounded look-ahead into the buffer of incoming words, (ii) the complete history of transition actions taken by the parser, and (iii) the complete contents of the stack of partially built tree fragments, including their internal structures. In addition, we compare two different word representations: (i) standard word vectors based on look-up tables and (ii) character-based models of words. Although standard word embedding models work well in all languages, the character-based models improve the handling of out-of-vocabulary words, particularly in morphologically rich languages. Finally, we discuss the use of dynamic oracles in training the parser. During training, dynamic oracles alternate between sampling parser states from the training data and from the model as it is being learned, making the model more robust to the kinds of errors that will be made at test time. Training our model with dynamic oracles yields a linear-time greedy parser with very competitive performance.
Miguel Ballesteros, Chris Dyer, Yoav Goldberg, Noah A. Smith
Comput. Linguistics3
2016 Coordination Annotation Extension in the Penn Tree Bank
abstract
Coordination is an important and common syntactic construction which is not handled well by state of the art parsers.Coordinations in the Penn Treebank are missing internal structure in many cases, do not include explicit marking of the conjuncts and contain various errors and inconsistencies.In this work, we initiated manual annotation process for solving these issues.We identify the different elements in a coordination phrase and label each element with its function.We add phrase boundaries when these are missing, unify inconsistencies, and fix errors.The outcome is an extension of the PTB that includes consistent and detailed structures for coordinations.We make the coordination annotation publicly available, in hope that they will facilitate further research into coordination disambiguation. 1
Jessica Ficler, Yoav Goldberg
ACL (1)2
2016 Improving Hypernymy Detection with an Integrated Path-based and Distributional Method
abstract
Detecting hypernymy relations is a key task in NLP, which is addressed in the literature using two complementary approaches.Distributional methods, whose supervised variants are the current best performers, and path-based methods, which received less research attention.We suggest an improved path-based algorithm, in which the dependency paths are encoded using a recurrent neural network, that achieves results comparable to distributional methods.We then extend the approach to integrate both pathbased and distributional signals, significantly improving upon the state-of-the-art on this task.
Vered Shwartz, Yoav Goldberg, Ido Dagan
ACL (1)2
2016 Semi Supervised Preposition-Sense Disambiguation using Multilingual Data
abstract
Prepositions are very common and very ambiguous, and understanding their sense is critical for understanding the meaning of the sentence. Supervised corpora for the preposition-sense disambiguation task are small, suggesting a semi-supervised approach to the task. We show that signals from unannotated multilingual data can be used to improve supervised preposition-sense disambiguation. Our approach pre-trains an LSTM encoder for predicting the translation of a preposition, and then incorporates the pre-trained encoder as a component in a supervised classification system, and fine-tunes it for the task. The multilingual signals consistently improve results on two preposition-sense datasets.
Hila Gonen, Yoav Goldberg
COLING2
2016 Training with Exploration Improves a Greedy Stack LSTM Parser
abstract
We adapt the greedy Stack-LSTM dependency parser of Dyer et al. (2015) to support a training-with-exploration procedure using dynamic oracles(Goldberg and Nivre, 2013) instead of cross-entropy minimization. This form of training, which accounts for model predictions at training time rather than assuming an error-free action history, improves parsing accuracies for both English and Chinese, obtaining very strong results for both languages. We discuss some modifications needed in order to get training with exploration to work well for a probabilistic neural-network.
Miguel Ballesteros, Yoav Goldberg, Chris Dyer, Noah A. Smith
EMNLP2
2016 A Neural Network for Coordination Boundary Prediction
abstract
We propose a neural-network based model for coordination boundary prediction. The network is designed to incorporate two signals: the similarity between conjuncts and the observation that replacing the whole coordination phrase with a conjunct tends to produce a coherent sentences. The modeling makes use of several LSTM networks. The model is trained solely on conjunction annotations in a Treebank, without using external resources. We show improvements on predicting coordination boundaries on the PTB compared to two state-of-the-art parsers; as well as improvement over previous coordination boundary prediction systems on the Genia corpus.
Jessica Ficler, Yoav Goldberg
EMNLP2
2016 Universal Dependencies v1: A Multilingual Treebank Collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajic 0001, Christopher D. Manning, Ryan T. McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, Daniel Zeman
LREC4
2016 Improving sentence compression by learning to predict gaze
abstract
We show how eye-tracking corpora can be used to improve sentence compression models, presenting a novel multi-task learning algorithm based on multi-layer LSTMs.We obtain performance competitive with or better than state-of-the-art approaches.
Sigrid Klerke, Yoav Goldberg, Anders Søgaard
HLT-NAACL2
2016 A Primer on Neural Network Models for Natural Language Processing
abstract
Over the past few years, neural networks have re-emerged as powerful machine-learning models, yielding state-of-the-art results in fields such as image recognition and speech processing. More recently, neural network models started to be applied also to textual natural language signals, again with very promising results. This tutorial surveys neural network models from the perspective of natural language processing research, in an attempt to bring natural-language researchers up to speed with the neural techniques. The tutorial covers input encoding for natural language tasks, feed-forward networks, convolutional networks, recurrent networks and recursive networks, as well as the computation graph abstraction for automatic gradient computation.
Yoav Goldberg
J. Artif. Intell. Res.1
2016 Simple and Accurate Dependency Parsing Using Bidirectional LSTM Feature Representations
abstract
We present a simple and effective scheme for dependency parsing which is based on bidirectional-LSTMs (BiLSTMs). Each sentence token is associated with a BiLSTM vector representing the token in its sentential context, and feature vectors are constructed by concatenating a few BiLSTM vectors. The BiLSTM is trained jointly with the parser objective, resulting in very effective feature extractors for parsing. We demonstrate the effectiveness of the approach by applying it to a greedy transition-based parser as well as to a globally optimized graph-based parser. The resulting parsers have very simple architectures, and match or surpass the state-of-the-art accuracies on English and Chinese.
Eliyahu Kiperwasser, Yoav Goldberg
Trans. Assoc. Comput. Linguistics2
2016 Easy-First Dependency Parsing with Hierarchical Tree LSTMs
abstract
We suggest a compositional vector representation of parse trees that relies on a recursive combination of recurrent-neural network encoders. To demonstrate its effectiveness, we use the representation as the backbone of a greedy, bottom-up dependency parser, achieving very strong accuracies for English and Chinese, without relying on external word embeddings. The parser’s implementation is available for download at the first author’s webpage.
Eliyahu Kiperwasser, Yoav Goldberg
Trans. Assoc. Comput. Linguistics2
2016 Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies
abstract
The success of long short-term memory (LSTM) neural networks in language processing is typically attributed to their ability to capture long-distance statistical regularities. Linguistic regularities are often sensitive to syntactic structure; can such dependencies be captured by LSTMs, which do not have explicit structural representations? We begin addressing this question using number agreement in English subject-verb dependencies. We probe the architecture’s grammatical competence both using training objectives with an explicit grammatical target (number prediction, grammaticality judgments) and using language models. In the strongly supervised settings, the LSTM achieved very high overall accuracy (less than 1% errors), but errors increased when sequential and structural information conflicted. The frequency of such errors rose sharply in the language-modeling setting. We conclude that LSTMs can capture a non-trivial amount of grammatical structure given targeted supervision, but stronger architectures may be required to further reduce errors; furthermore, the language modeling signal is insufficient for capturing syntax-sensitive dependencies, and should be supplemented with more direct supervision if such dependencies need to be captured.
Tal Linzen, Emmanuel Dupoux, Yoav Goldberg
Trans. Assoc. Comput. Linguistics3
2015 Semi-supervised Dependency Parsing using Bilexical Contextual Features from Auto-Parsed Data
abstract
We present a semi-supervised approach to improve dependency parsing accuracy by using bilexical statistics derived from auto-parsed data.The method is based on estimating the attachment potential of head-modifier words, by taking into account not only the head and modifier words themselves, but also the words surrounding the head and the modifier.When integrating the learned statistics as features in a graph-based parsing model, we observe nice improvements in accuracy when parsing various English datasets.
Eliyahu Kiperwasser, Yoav Goldberg
EMNLP2
2015 Template Kernels for Dependency Parsing
abstract
A common approach to dependency parsing is scoring a parse via a linear function of a set of indicator features.These features are typically manually constructed from templates that are applied to parts of the parse tree.The templates define which properties of a part should combine to create features.Existing approaches consider only a small subset of the possible combinations, due to statistical and computational efficiency considerations.In this work we present a novel kernel which facilitates efficient parsing with feature representations corresponding to a much larger set of combinations.We integrate the kernel into a parse reranking system and demonstrate its effectiveness on four languages from the CoNLL-X shared task. 1
Hillel Taub-Tabib, Yoav Goldberg, Amir Globerson
HLT-NAACL2
2015 Improving Distributional Similarity with Lessons Learned from Word Embeddings
abstract
Recent trends suggest that neural-network-inspired word embedding models outperform traditional count-based distributional models on word similarity and analogy detection tasks. We reveal that much of the performance gains of word embeddings are due to certain system design choices and hyperparameter optimizations, rather than the embedding algorithms themselves. Furthermore, we show that these modifications can be transferred to traditional distributional models, yielding similar gains. In contrast to prior reports, we observe mostly local or insignificant performance differences between the methods, with no global advantage to any single approach over the others.
Omer Levy, Yoav Goldberg, Ido Dagan
Trans. Assoc. Comput. Linguistics2
2014 Linguistic Regularities in Sparse and Explicit Word Representations
abstract
Recent work has shown that neuralembedded word representations capture many relational similarities, which can be recovered by means of vector arithmetic in the embedded space.We show that Mikolov et al.'s method of first adding and subtracting word vectors, and then searching for a word similar to the result, is equivalent to searching for a word that maximizes a linear combination of three pairwise word similarities.Based on this observation, we suggest an improved method of recovering relational similarities, improving the state-of-the-art results on two recent word-analogy datasets.Moreover, we demonstrate that analogy recovery is not restricted to neural word embeddings, and that a similar amount of relational similarities can be recovered from traditional distributional word representations.
Omer Levy, Yoav Goldberg
CoNLL2
2014 Neural Word Embedding as Implicit Matrix Factorization
Omer Levy, Yoav Goldberg
NIPS2
2014 Constrained Arc-Eager Dependency Parsing
abstract
Arc-eager dependency parsers process sentences in a single left-to-right pass over the input and have linear time complexity with greedy decoding or beam search. We show how such parsers can be constrained to respect two different types of conditions on the output dependency graph: span constraints, which require certain spans to correspond to subtrees of the graph, and arc constraints, which require certain arcs to be present in the graph. The constraints are incorporated into the arc-eager transition system as a set of preconditions for each transition and preserve the linear time complexity of the parser.
Joakim Nivre, Yoav Goldberg, Ryan T. McDonald
Comput. Linguistics2
2014 A Tabular Method for Dynamic Oracles in Transition-Based Parsing
abstract
We develop parsing oracles for two transition-based dependency parsers, including the arc-standard parser, solving a problem that was left open in (Goldberg and Nivre, 2013). We experimentally show that using these oracles during training yields superior parsing accuracies on many languages.
Yoav Goldberg, Francesco Sartorio, Giorgio Satta
Trans. Assoc. Comput. Linguistics1
2013 A Non-Monotonic Arc-Eager Transition System for Dependency Parsing
Matthew Honnibal, Yoav Goldberg
CoNLL2
2013 Word Segmentation, Unknown-word Resolution, and Morphological Agreement in a Hebrew Parsing System
abstract
We present a constituency parsing system for Modern Hebrew. The system is based on the PCFG-LA parsing method of Petrov et al. 2006 , which is extended in various ways in order to accommodate the specificities of Hebrew as a morphologically rich language with a small treebank. We show that parsing performance can be enhanced by utilizing a language resource external to the treebank, specifically, a lexicon-based morphological analyzer. We present a computational model of interfacing the external lexicon and a treebank-based parser, also in the common case where the lexicon and the treebank follow different annotation schemes. We show that Hebrew word-segmentation and constituency-parsing can be performed jointly using CKY lattice parsing. Performing the tasks jointly is effective, and substantially outperforms a pipeline-based model. We suggest modeling grammatical agreement in a constituency-based parser as a filter mechanism that is orthogonal to the grammar, and present a concrete implementation of the method. Although the constituency parser does not make many agreement mistakes to begin with, the filter mechanism is effective in fixing the agreement mistakes that the parser does make. These contributions extend outside of the scope of Hebrew processing, and are of general applicability to the NLP community. Hebrew is a specific case of a morphologically rich language, and ideas presented in this work are useful also for processing other languages, including English. The lattice-based parsing methodology is useful in any case where the input is uncertain. Extending the lexical coverage of a treebank-derived parser using an external lexicon is relevant for any language with a small treebank.
Yoav Goldberg, Michael Elhadad
Comput. Linguistics1
2013 Training Deterministic Parsers with Non-Deterministic Oracles
abstract
Greedy transition-based parsers are very fast but tend to suffer from error propagation. This problem is aggravated by the fact that they are normally trained using oracles that are deterministic and incomplete in the sense that they assume a unique canonical path through the transition system and are only valid as long as the parser does not stray from this path. In this paper, we give a general characterization of oracles that are nondeterministic and complete, present a method for deriving such oracles for transition systems that satisfy a property we call arc decomposition, and instantiate this method for three well-known transition systems from the literature. We say that these oracles are dynamic, because they allow us to dynamically explore alternative and nonoptimal paths during training — in contrast to oracles that statically assume a unique optimal path. Experimental evaluation on a wide range of data sets clearly shows that using dynamic oracles to train greedy parsers gives substantial improvements in accuracy. Moreover, this improvement comes at no cost in terms of efficiency, unlike other techniques like beam search.
Yoav Goldberg, Joakim Nivre
Trans. Assoc. Comput. Linguistics1
2012 A Dynamic Oracle for Arc-Eager Dependency Parsing
Yoav Goldberg, Joakim Nivre
COLING1
2011 Rich Parameterization Improves RNA Structure Prediction
Shay Zakov, Yoav Goldberg, Michael Elhadad, Michal Ziv-Ukelson
RECOMB2
2010 Inspecting the Structural Biases of Dependency Parsing Algorithms
Yoav Goldberg, Michael Elhadad
CoNLL1
2010 An Efficient Algorithm for Easy-First Non-Directional Dependency Parsing
Yoav Goldberg, Michael Elhadad
HLT-NAACL1
2009 Enhancing Unlexicalized Parsing Performance Using a Wide Coverage Lexicon, Fuzzy Tag-Set Mapping, and EM-HMM-Based Lexical Probabilities
Yoav Goldberg, Reut Tsarfaty, Meni Adler, Michael Elhadad
EACL1
2009 On the Role of Lexical Features in Sequence Labeling
Yoav Goldberg, Michael Elhadad
EMNLP1
2008 Unsupervised Lexicon-Based Resolution of Unknown Words for Full Morphological Analysis
Meni Adler, Yoav Goldberg, David Gabay, Michael Elhadad
ACL2
2008 EM Can Find Pretty Good HMM POS-Taggers (When Given a Good Start)
Yoav Goldberg, Meni Adler, Michael Elhadad
ACL1
2008 A Single Generative Model for Joint Morphological Segmentation and Syntactic Parsing
Yoav Goldberg, Reut Tsarfaty
ACL1
2008 Identification of Transliterated Foreign Words in Hebrew Script
Yoav Goldberg, Michael Elhadad
CICLing1
2008 Tagging a Hebrew Corpus: the Case of Participles
Meni Adler, Yael Dahan Netzer, Yoav Goldberg, David Gabay, Michael Elhadad
LREC3
2008 Word-Based or Morpheme-Based? Annotation Strategies for Modern Hebrew Clitics
Reut Tsarfaty, Yoav Goldberg
LREC2
2007 SVM Model Tampering and Anchored Learning: A Case Study in Hebrew NP Chunking
Yoav Goldberg, Michael Elhadad
ACL1
2006 Noun Phrase Chunking in Hebrew: Influence of Lexical and Morphological Features
abstract
We present a method for Noun Phrase chunking in Hebrew. We show that the traditional definition of base-NPs as non-recursive noun phrases does not apply in Hebrew, and propose an alternative definition of Simple NPs. We review syntactic properties of Hebrew related to noun phrases, which indicate that the task of Hebrew SimpleNP chunking is harder than base-NP chunking in English. As a confirmation, we apply methods known to work well for English to Hebrew data. These methods give low results (F from 76 to 86) in Hebrew. We then discuss our method, which applies SVM induction over lexical and morphological features. Morphological features improve the average precision by ~0.5%, recall by ~1%, and F-measure by ~0.75, resulting in a system with average performance of 93% precision, 93.4% recall and 93.2 F-measure.
Yoav Goldberg, Meni Adler, Michael Elhadad
ACL1
2006 Strategic Intelligence Analysis: From Information Processing to Meaning-Making
Yair Neuman, Liran Elihay, Meni Adler, Yoav Goldberg, Amir Winer
ISI4