EDBT 2026 Demo / reviewers in the wild / expert
Yanai Elazar
dblp:223/4533
· DBLP profile ↗
28ranked-venue papers
9as first author
21since 2021 · last 2026
0009-0000-8138-4533ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 9 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model BehaviorabstractAbstract We present an experimental recipe for studying the relationship between training data and language model (LM) behavior. We outline steps for intervening on data batches – i.e., “rewriting history” – and then retraining model checkpoints over that data to test hypotheses relating data to behavior. Our intervention recipe’s stages are (1) selecting evaluation items from a benchmark that measures model behavior, (2) matching relevant documents to those items, and (3) modifying those documents before retraining and measuring the effects. We demonstrate the utility of our recipe through case studies on factual knowledge acquisition and gender bias in LMs, using both cooccurrence statistics and information retrieval methods to identify documents that might contribute to model behavior. Our results supplement past observational analyses that link cooccurrence to model behavior, while demonstrating that extant methods for identifying relevant training documents do not fully explain an LM’s abilities and biases. Researchers can follow the recipe to test further hypotheses about how training data affects model behavior. Our code is made publicly available to promote future work.1 Rahul Nadkarni, Yanai Elazar, Hila Gonen, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 2 |
| 2025 | Calibrating Large Language Models with Sample ConsistencyabstractAccurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and elude conventional calibration techniques due to their proprietary nature and massive scale. In this work, we derive model confidence from the distribution of multiple randomly sampled generations, using three measures of consistency. We extensively evaluate eleven open and closed-source models on nine reasoning datasets. Results show that consistency-based calibration methods outperform existing post-hoc approaches in terms of calibration error. Meanwhile, we find that factors such as intermediate explanations, model scaling, and larger sample sizes enhance calibration, while instruction-tuning makes calibration more difficult. Moreover, confidence scores obtained from consistency can potentially enhance model performance. Finally, we offer guidance on choosing suitable consistency metrics for calibration, tailored to model characteristics such as the exposure to instruction-tuning and RLHF. Qing Lyu 0001, Kumar Shridhar, Chaitanya Malaviya, Li Zhang 0039, Yanai Elazar, Niket Tandon, Marianna Apidianaki, Mrinmaya Sachan, Chris Callison-Burch |
AAAI | 5 |
| 2025 | Hybrid Preferences: Learning to Route Instances for Human vs. AI FeedbackabstractLester James Validad Miranda, Yizhong Wang, Yanai Elazar, Sachin Kumar, Valentina Pyatkin, Faeze Brahman, Noah A. Smith, Hannaneh Hajishirzi, Pradeep Dasigi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lester James V. Miranda, Yizhong Wang, Yanai Elazar, Sachin Kumar 0009, Valentina Pyatkin, Faeze Brahman, Noah A. Smith, Hannaneh Hajishirzi, Pradeep Dasigi |
ACL (1) | 3 |
| 2025 | Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining DataabstractThe impressive capabilities of large language models (LLMs) have sparked debate over whether these models genuinely generalize to unseen tasks or predominantly rely on memorizing vast amounts of pretraining data. To explore this issue, we introduce an extended concept of memorization, distributional memorization, which measures the correlation between the LLM output probabilities and the pretraining data frequency. To effectively capture task-specific pretraining data frequency, we propose a novel task-gram language model, which is built by counting the co-occurrence of semantically related $n$-gram pairs from task inputs and outputs in the pretraining corpus. Using the Pythia models trained on the Pile dataset, we evaluate four distinct tasks: machine translation, factual question answering, world knowledge understanding, and math reasoning. Our findings reveal varying levels of memorization, with the strongest effect observed in factual question answering. Furthermore, while model performance improves across all tasks as LLM size increases, only factual question answering shows an increase in memorization, whereas machine translation and reasoning tasks exhibit greater generalization, producing more novel outputs. This study demonstrates that memorization plays a larger role in simpler, knowledge-intensive tasks, while generalization is the key for harder, reasoning-based tasks, providing a scalable method for analyzing large pretraining corpora in greater depth. Xinyi Wang 0003, Antonis Antoniades, Yanai Elazar, Alfonso Amayuelas, Alon Albalak, Kexun Zhang, William Yang Wang |
ICLR | 3 |
| 2025 | On Linear Representations and Pretraining Data Frequency in Language ModelsabstractPretraining data has a direct impact on the behaviors and quality of language models (LMs), but we only understand the most basic principles of this relationship. While most work focuses on pretraining data's effect on downstream task behavior, we investigate its relationship to LM representations. Previous work has discovered that, in language models, some concepts are encoded "linearly" in the representations, but what factors cause these representations to form (or not)? We study the connection between pretraining data frequency and models' linear representations of factual relations (e.g., mapping France to Paris in a capital prediction task). We find evidence that the formation of linear representations is strongly connected to pretraining term frequencies; specifically for subject-relation-object fact triplets, both subject-object co-occurrence frequency and in-context learning accuracy for the relation are highly correlated with linear representations. This is the case across all phases of pretraining, i.e., it is not affected by the model's underlying capability. In OLMo-7B and GPT-J (6B), we discover that a linear representation consistently (but not exclusively) forms when the subjects and objects within a relation co-occur at least 1k and 2k times, respectively, regardless of when these occurrences happen during pretraining (and around 4k times for OLMo-1B). Finally, we train a regression model on measurements of linear representation quality in fully-trained LMs that can predict how often a term was seen in pretraining. Our model achieves low error even on inputs from a different model with a different pretraining dataset, providing a new method for estimating properties of the otherwise-unknown training data of closed-data models. We conclude that the strength of linear representations in LMs contains signal about the models' pretraining corpora that may provide new avenues for controlling and improving model behavior: particularly, manipulating the models' training data to meet specific frequency thresholds. We release our code to support future work. Jack Merullo, Noah A. Smith, Sarah Wiegreffe, Yanai Elazar |
ICLR | 4 |
| 2024 | OLMo: Accelerating the Science of Language ModelsabstractDirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, William Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah Smith, Hannaneh Hajishirzi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Dirk Groeneveld, Iz Beltagy, Pete Walsh 0001, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert 0001, Kyle Richardson 0001, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, Hannaneh Hajishirzi |
ACL (1) | 17 |
| 2024 | Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining ResearchabstractLuca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar 0009, Li Lucy, Xinxi Lyu, Nathan Lambert 0001, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson 0001, Shannon Shen 0001, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh 0001, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, Kyle Lo |
ACL (1) | 10 |
| 2024 | Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine TranslationabstractMost works on gender bias focus on intrinsic bias -removing traces of information about a protected group from the model's internal representation.However, these works are often disconnected from the impact of such debiasing on downstream applications, which is the main motivation for debiasing in the first place.In this work, we systematically test how methods for intrinsic debiasing affect neural machine translation models, by measuring the extrinsic bias of such systems under different design choices.We highlight three challenges and mismatches between the debiasing techniques and their end-goal usage, including the choice of embeddings to debias, the mismatch between words and sub-word tokens debiasing, and the effect of translating from English to different target languages.We find that these considerations have a significant impact on downstream performance and the success of debiasing.1 Bar Iluz, Yanai Elazar, Asaf Yehudai, Gabriel Stanovsky |
EMNLP | 2 |
| 2024 | Evaluating n-Gram Novelty of Language Models Using Rusty-DAWGabstractHow novel are texts generated by language models (LMs) relative to their training corpora?In this work, we investigate the extent to which modern LMs generate n-grams from their training data, evaluating both (i) the probability LMs assign to complete training n-grams and (ii) n-novelty, the proportion of n-grams generated by an LM that did not appear in the training data (for arbitrarily large n).To enable arbitrary-length n-gram search over a corpus in constant time w.r.t.corpus size, we develop RUSTY-DAWG, a novel search tool inspired by indexing of genomic data.We compare the novelty of LM-generated text to humanwritten text and explore factors that affect generation novelty, focusing on the Pythia models.We find that, for n > 4, LM-generated text is less novel than human-written text, though it is more novel for smaller n.Larger LMs and more constrained decoding strategies both decrease novelty.Finally, we show that LMs complete n-grams with lower loss if they are more frequent in the training data.Overall, our results reveal factors influencing the novelty of LMgenerated text, and we release RUSTY-DAWG to facilitate further pretraining data research.1 William Merrill, Noah A. Smith, Yanai Elazar |
EMNLP | 3 |
| 2024 | Detection and Measurement of Syntactic Templates in Generated TextabstractThe diversity of text can be measured beyond word-level features, however existing diversity evaluation focuses primarily on word-level features.Here we propose a method for evaluating diversity over syntactic features to characterize general repetition in models, beyond frequent n-grams.Specifically, we define syntactic templates (e.g., strings comprising parts-of-speech) and show that models tend to produce templated text in downstream tasks at a higher rate than what is found in human-reference texts We find that most (76%) templates in modelgenerated text can be found in pre-training data (compared to only 35% of human-authored text), and are not overwritten during fine-tuning or alignment processes such as RLHF.The connection between templates in generated text and the pre-training data allows us to analyze syntactic templates in models where we do not have the pre-training data.We also find that templates as features are able to differentiate between models, tasks, and domains, and are useful for qualitatively evaluating common model constructions.Finally, we demonstrate the use of templates as a useful tool for analyzing style memorization of training data in LLMs 1 . Chantal Shaib, Yanai Elazar, Junyi Jessy Li, Byron C. Wallace |
EMNLP | 2 |
| 2024 | What's In My Big Data?abstractLarge text corpora are the backbone of language models.
However, we have a limited understanding of the content of these corpora, including general statistics, quality, social factors, and inclusion of evaluation data (contamination).
In this work, we propose What's In My Big Data? (WIMBD), a platform and a set of sixteen analyses that allow us to reveal and compare the contents of large text corpora. WIMBD builds on two basic capabilities---count and search---*at scale*, which allows us to analyze more than 35 terabytes on a standard compute node.
We apply WIMBD to ten different corpora used to train popular language models, including *C4*, *The Pile*, and *RedPajama*.
Our analysis uncovers several surprising and previously undocumented findings about these corpora, including the high prevalence of duplicate, synthetic, and low-quality content, personally identifiable information, toxic language, and benchmark contamination.
For instance, we find that about 50% of the documents in *RedPajama* and *LAION-2B-en* are duplicates. In addition, several datasets used for benchmarking models trained on such corpora are contaminated with respect to important benchmarks, including the Winograd Schema Challenge and parts of GLUE and SuperGLUE.
We open-source WIMBD's code and artifacts to provide a standard set of evaluations for new text-based corpora and to encourage more analyses and transparency around them. Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Pete Walsh 0001, Dirk Groeneveld, Luca Soldaini, Sameer Singh 0001, Hannaneh Hajishirzi, Noah A. Smith, Jesse Dodge |
ICLR | 1 |
| 2024 | The Bias Amplification Paradox in Text-to-Image GenerationabstractPreethi Seshadri, Sameer Singh, Yanai Elazar. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Preethi Seshadri, Sameer Singh 0001, Yanai Elazar |
NAACL-HLT | 3 |
| 2024 | Paloma: A Benchmark for Evaluating Language Model FitabstractEvaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains—varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. We include two new datasets of the top 100 subreddits (e.g., r/depression on Reddit) and programming languages (e.g., Java on GitHub), both sources common in contemporary LMs. With our benchmark, we release 6 baseline 1B LMs carefully controlled to provide fair comparisons about which pretraining corpus is best and code for others to apply those controls to their own experiments. Our case studies demonstrate how the fine-grained results from Paloma surface findings such as that models pretrained without data beyond Common Crawl exhibit anomalous gaps in LM fit to many domains or that loss is dominated by the most frequently occurring strings in the vocabulary. Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Pete Walsh 0001, Yanai Elazar, Kyle Lo, Dirk Groeneveld, Iz Beltagy, Hannaneh Hajishirzi, Noah A. Smith, Kyle Richardson 0001, Jesse Dodge |
NeurIPS | 9 |
| 2022 | Text-based NP EnrichmentabstractAbstract Understanding the relations between entities denoted by NPs in a text is a critical part of human-like natural language understanding. However, only a fraction of such relations is covered by standard NLP tasks and benchmarks nowadays. In this work, we propose a novel task termed text-based NP enrichment (TNE), in which we aim to enrich each NP in a text with all the preposition-mediated relations—either explicit or implicit—that hold between it and other NPs in the text. The relations are represented as triplets, each denoted by two NPs related via a preposition. Humans recover such relations seamlessly, while current state-of-the-art models struggle with them due to the implicit nature of the problem. We build the first large-scale dataset for the problem, provide the formal framing and scope of annotation, analyze the data, and report the results of fine-tuned language models on the task, demonstrating the challenge it poses to current technology. A webpage with a data-exploration UI, a demo, and links to the code, models, and leaderboard, to foster further research into this challenging problem can be found at: yanaiela.github.io/TNE/. Yanai Elazar, Victoria Basmova, Yoav Goldberg, Reut Tsarfaty |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | First Align, then Predict: Understanding the Cross-Lingual Ability of Multilingual BERTabstractMultilingual pretrained language models have demonstrated remarkable zero-shot crosslingual transfer capabilities.Such transfer emerges by fine-tuning on a task of interest in one language and evaluating on a distinct language, not seen during the fine-tuning.Despite promising results, we still lack a proper understanding of the source of this transfer.Using a novel layer ablation technique and analyses of the model's internal representations, we show that multilingual BERT, a popular multilingual language model, can be viewed as the stacking of two sub-networks: a multilingual encoder followed by a taskspecific language-agnostic predictor.While the encoder is crucial for cross-lingual transfer and remains mostly unchanged during finetuning, the task predictor has little importance on the transfer and can be reinitialized during fine-tuning.We present extensive experiments with three distinct tasks, seventeen typologically diverse languages and multiple domains to support our hypothesis.RANDOM-INIT of layers SRC-TRG REF ∆1-2 ∆3-4 ∆5-6 ∆7-8 ∆9-10 ∆11-12 Parsing EN -EN 88.98 -0.96 -0.66 -0.93 -0.55 0.04 -0.09RU -RU 85.15 -0.82 -1.38 -1.51 -0.86 -0.29 0.18 AR -AR 59.54 -0.78 -2.14 -1.20 -0.67 -0.27 0.08 EN -X 53.23 -15.77 -6.51 -3.39 -1.47 0.29 1.00 RU -X 55.41 -7.69 -3.71 -3.13 -1.70 0.92 0.94 AR -X 27.97 -4.91 -3.17 -1.48 -1.68 -0.36 -0.14 POS EN -EN 96.51 -0.30 -0.25 -0.40 -0.00 0.05 0.02 RU -RU 96.90 -0.52 -0.55 -0.40 -0.07 0.02 -0.03 AR -AR 79.28 -0.35 -0.49-0.36 -0.19 -0.05 -0.00 EN -X 79.37 -8.94 -2.49-1.66 -0.88 0.20 -0.14 RU -X 79.25 -10.08 -2.83 -1.65 -2.74 0.01 -0.45 AR -X 64.81 -6.73 -3.50 -1.63 -1.56 -0.73 -1.29 NER EN -EN 83.30 -2.66 -2.14 -1.43 -0.63 -0.23 -0.12 RU -RU 88.20 -2.08 -2.13 -1.52 -0.64 -0.33 -0.13 AR -AR 87.97 -2.37 -2.11 -0.96 -0.39 -0.15 0.21 EN -X 64.17 -8.28 -5.09 -3.07 -0.79 -0.47 -0.13 RU -X 62.13 -15.85 -9.36 -5.50 -2.44 -1.16 -0.06 AR -X 65.59 -16.10 -8.42 -3. Benjamin Muller, Yanai Elazar, Benoît Sagot, Djamé Seddah |
EACL | 2 |
| 2021 | Back to Square One: Artifact Detection, Training and Commonsense Disentanglement in the Winograd SchemaabstractThe Winograd Schema (WS) has been proposed as a test for measuring commonsense capabilities of models.Recently, pre-trained language model-based approaches have boosted performance on some WS benchmarks but the source of improvement is still not clear.This paper suggests that the apparent progress on WS may not necessarily reflect progress in commonsense reasoning.To support this claim, we first show that the current evaluation method of WS is sub-optimal and propose a modification that uses twin sentences for evaluation.We also propose two new baselines that indicate the existence of artifacts in WS benchmarks.We then develop a method for evaluating WS-like sentences in a zero-shot setting to account for the commonsense reasoning abilities acquired during the pretraining and observe that popular language models perform randomly in this setting when using our more strict evaluation.We conclude that the observed progress is mostly due to the use of supervision in training WS models, which is not likely to successfully support all the required commonsense reasoning skills and knowledge.1 Yanai Elazar, Hongming Zhang 0009, Yoav Goldberg, Dan Roth 0001 |
EMNLP (1) | 1 |
| 2021 | Contrastive Explanations for Model InterpretabilityabstractContrastive explanations clarify why an event occurred in contrast to another.They are inherently intuitive to humans to both produce and comprehend.We propose a method to produce contrastive explanations in the latent space, via a projection of the input representation, such that only the features that differentiate two potential decisions are captured.Our modification allows model behavior to consider only contrastive reasoning, and uncover which aspects of the input are useful for and against particular decisions.Additionally, for a given input feature, our contrastive explanations can answer for which label, and against which alternative label, is the feature useful.We produce contrastive explanations via both highlevel abstract concept attribution and low-level input token/span attribution for two NLP classification benchmarks.Our findings demonstrate the ability of label-contrastive explanations to provide fine-grained interpretability of model decisions.1 Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi 0001, Yoav Goldberg |
EMNLP (1) | 4 |
| 2021 | Measuring and Improving Consistency in Pretrained Language ModelsabstractAbstract Consistency of a model—that is, the invariance of its behavior under meaning-preserving alternations in its input—is a highly desirable property in natural language processing. In this paper we study the question: Are Pretrained Language Models (PLMs) consistent with respect to factual knowledge? To this end, we create ParaRel🤘, a high-quality resource of cloze-style query English paraphrases. It contains a total of 328 paraphrases for 38 relations. Using ParaRel🤘, we show that the consistency of all PLMs we experiment with is poor— though with high variance between relations. Our analysis of the representational spaces of PLMs suggests that they have a poor structure and are currently not suitable for representing knowledge robustly. Finally, we propose a method for improving model consistency and experimentally demonstrate its effectiveness.1 Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schütze, Yoav Goldberg |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | Erratum: Measuring and Improving Consistency in Pretrained Language ModelsabstractAbstract During production of this paper, an error was introduced to the formula on the bottom of the right column of page 1020. In the last two terms of the formula, the n and m subscripts were swapped. The correct formula is:Lc=∑n=1k∑m=n+1kDKL(Qnri∥Qmri)+DKL(Qmri∥Qnri)The paper has been updated. Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schütze, Yoav Goldberg |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | Amnesic Probing: Behavioral Explanation With Amnesic CounterfactualsabstractAbstract A growing body of work makes use of probing in order to investigate the working of neural models, often considered black boxes. Recently, an ongoing debate emerged surrounding the limitations of the probing paradigm. In this work, we point out the inability to infer behavioral conclusions from probing results, and offer an alternative method that focuses on how the information is being used, rather than on what information is encoded. Our method, Amnesic Probing, follows the intuition that the utility of a property for a given task can be assessed by measuring the influence of a causal intervention that removes it from the representation. Equipped with this new analysis tool, we can ask questions that were not possible before, for example, is part-of-speech information important for word prediction? We perform a series of analyses on BERT to answer these types of questions. Our findings demonstrate that conventional probing performance is not correlated to task importance, and we call for increased scrutiny of claims that draw behavioral or causal conclusions from probing results.1 Yanai Elazar, Shauli Ravfogel, Alon Jacovi, Yoav Goldberg |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | Revisiting Few-shot Relation Classification: Evaluation Data and Classification SchemesabstractWe explore few-shot learning (FSL) for relation classification (RC). Focusing on the realistic scenario of FSL, in which a test instance might not belong to any of the target categories (none-of-the-above, [NOTA]), we first revisit the recent popular dataset structure for FSL, pointing out its unrealistic data distribution. To remedy this, we propose a novel methodology for deriving more realistic few-shot test data from available datasets for supervised RC, and apply it to the TACRED dataset. This yields a new challenging benchmark for FSL-RC, on which state of the art models show poor performance. Next, we analyze classification schemes within the popular embedding-based nearest-neighbor approach for FSL, with respect to constraints they impose on the embedding space. Triggered by this analysis, we propose a novel classification scheme in which the NOTA category is represented as learned vectors, shown empirically to be an appealing option for FSL. Ofer Sabo, Yanai Elazar, Yoav Goldberg, Ido Dagan |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | Null It Out: Guarding Protected Attributes by Iterative Nullspace ProjectionabstractThe ability to control for the kinds of information encoded in neural representation has a variety of use cases, especially in light of the challenge of interpreting these models.We present Iterative Null-space Projection (INLP), a novel method for removing information from neural representations.Our method is based on repeated training of linear classifiers that predict a certain property we aim to remove, followed by projection of the representations on their null-space.By doing so, the classifiers become oblivious to that target property, making it hard to linearly separate the data according to it.While applicable for multiple uses, we evaluate our method on bias and fairness use-cases, and show that our method is able to mitigate bias in word embeddings, as well as to increase fairness in a setting of multi-class classification. Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, Yoav Goldberg |
ACL | 2 |
| 2020 | oLMpics - On what Language Model Pre-training CapturesabstractRecent success of pre-trained language models (LMs) has spurred widespread interest in the language capabilities that they possess. However, efforts to understand whether LM representations are useful for symbolic reasoning tasks have been limited and scattered. In this work, we propose eight reasoning tasks, which conceptually require operations such as comparison, conjunction, and composition. A fundamental challenge is to understand whether the performance of a LM on a task should be attributed to the pre-trained representations or to the process of fine-tuning on the task data. To address this, we propose an evaluation protocol that includes both zero-shot evaluation (no fine-tuning), as well as comparing the learning curve of a fine-tuned LM to the learning curve of multiple controls, which paints a rich picture of the LM capabilities. Our main findings are that: (a) different LMs exhibit qualitatively different reasoning abilities, e.g., RoBERTa succeeds in reasoning tasks where BERT fails completely; (b) LMs do not reason in an abstract manner and are context-dependent, e.g., while RoBERTa can compare ages, it can do so only when the ages are in the typical range of human ages; (c) On half of our reasoning tasks all models fail completely. Our findings and infrastructure can help future work on designing new datasets, models, and objective functions for pre-training. Alon Talmor, Yanai Elazar, Yoav Goldberg, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 2 |
| 2019 | How Large Are Lions? Inducing Distributions over Quantitative AttributesabstractMost current NLP systems have little knowledge about quantitative attributes of objects and events.We propose an unsupervised method for collecting quantitative information from large amounts of web data, and use it to create a new, very large resource consisting of distributions over physical quantities associated with objects, adjectives, and verbs which we call Distribution over Quantities (DOQ) 1 .This contrasts with recent work in this area which has focused on making only relative comparisons such as "Is a lion bigger than a wolf?".Our evaluation shows that DOQ compares favorably with state of the art results on existing datasets for relative comparisons of nouns and adjectives, and on a new dataset we introduce.* Work carried out during an internship at Google.† Work carried out during employment at Google. 1 The resource is available at https:// github.com/google-research-datasets/ distribution-over-quantities Yanai Elazar, Abhijit Mahabal, Deepak Ramachandran, Tania Bedrax-Weiss, Dan Roth 0001 |
ACL (1) | 1 |
| 2019 | Adversarial Removal of Demographic Attributes RevisitedabstractMaria Barrett, Yova Kementchedjhieva, Yanai Elazar, Desmond Elliott, Anders Søgaard. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Maria Barrett, Yova Kementchedjhieva, Yanai Elazar, Desmond Elliott, Anders Søgaard |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Privacy and Fairness in Recommender Systems via Adversarial Training of User RepresentationsabstractLatent factor models for recommender systems represent users and items as low dimensional vectors. Privacy risks of such systems have previously been studied mostly in the context of recovery of personal information in the form of usage records from the training data. However, the user representations themselves may be used together with external data to recover private user information such as gender and age. In this paper we show that user vectors calculated by a common recommender system can be exploited in this way. We propose the privacy-adversarial framework to eliminate such leakage of private information, and study the trade-off between recommender performance and leakage both theoretically and empirically using a benchmark dataset. An advantage of the proposed method is that it also helps guarantee fairness of results, since all implicit knowledge of a set of attributes is scrubbed from the representations used by the model, and thus can't enter into the decision making. We discuss further applications of this method towards the generation of deeper and more insightful recommendations. Yehezkel S. Resheff, Yanai Elazar, Shimon Shahar, Oren Sar Shalom |
ICPRAM | 2 |
| 2019 | Where's My Head? Definition, Dataset and Models for Numeric Fused-Heads Identification and ResolutionabstractWe provide the first computational treatment of fused-heads constructions (FHs), focusing on the numeric fused-heads (NFHs). FHs constructions are noun phrases in which the head noun is missing and is said to be “fused” with its dependent modifier. This missing information is implicit and is important for sentence understanding. The missing references are easily filled in by humans but pose a challenge for computational models. We formulate the handling of FHs as a two stages process: Identification of the FH construction and resolution of the missing head. We explore the NFH phenomena in large corpora of English text and create (1) a data set and a highly accurate method for NFH identification; (2) a 10k examples (1 M tokens) crowd-sourced data set of NFH resolution; and (3) a neural baseline for the NFH resolution task. We release our code and data set, to foster further research into this challenging problem. Yanai Elazar, Yoav Goldberg |
Trans. Assoc. Comput. Linguistics | 1 |
| 2018 | Adversarial Removal of Demographic Attributes from Text DataabstractRecent advances in Representation Learning and Adversarial Training seem to succeed in removing unwanted features from the learned representation.We show that demographic information of authors is encoded in-and can be recovered from-the intermediate representations learned by text-based neural classifiers.The implication is that decisions of classifiers trained on textual data are not agnostic to-and likely condition on-demographic attributes.When attempting to remove such demographic information using adversarial training, we find that while the adversarial component achieves chance-level development-set accuracy during training, a post-hoc classifier, trained on the encoded sentences from the first part, still manages to reach substantially higher classification accuracies on the same data.This behavior is consistent across several tasks, demographic properties and datasets.We explore several techniques to improve the effectiveness of the adversarial component.Our main conclusion is a cautionary one: do not rely on the adversarial training to achieve invariant representation to sensitive features. Yanai Elazar, Yoav Goldberg |
EMNLP | 1 |