Amir Feder

dblp:214/3604 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
14since 2021 · last 2024
0000-0001-5472-1135ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
abstract
When large language models are aligned via supervised fine-tuning, they may encounter new factual information that was not acquired through pre-training.It is often conjectured that this can teach the model the behavior of hallucinating factually incorrect responses, as the model is trained to generate facts that are not grounded in its pre-existing knowledge.In this work, we study the impact of such exposure to new knowledge on the capability of the fine-tuned model to utilize its pre-existing knowledge.To this end, we design a controlled setup, focused on closedbook QA, where we vary the proportion of the fine-tuning examples that introduce new knowledge.We demonstrate that large language models struggle to acquire new factual knowledge through fine-tuning, as fine-tuning examples that introduce new knowledge are learned significantly slower than those consistent with the model's knowledge.However, we also find that as the examples with new knowledge are eventually learned, they linearly increase the model's tendency to hallucinate.Taken together, our results highlight the risk in introducing new factual knowledge through fine-tuning, and support the view that large language models mostly acquire factual knowledge through pre-training, whereas finetuning teaches them to use it more efficiently.
Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, Jonathan Herzig
EMNLP5
2024 Exploring the Learning Capabilities of Language Models using LEVERWORLDS
abstract
Learning a model of a stochastic setting often involves learning both general structure rules and specific properties of the instance.This paper investigates the interplay between learning the general and the specific in various learning methods, with emphasis on sample efficiency.We design a framework called LEV-ERWORLDS, which allows the generation of simple physics-inspired worlds that follow a similar generative process with different distributions, and their instances can be expressed in natural language.These worlds allow for controlled experiments to assess the sample complexity of different learning methods.We experiment with classic learning algorithms as well as Transformer language models, both with fine-tuning and In-Context Learning (ICL).Our general finding is that (1) Transformers generally succeed in the task; but (2) they are considerably less sample efficient than classic methods that make stronger assumptions about the structure, such as Maximum Likelihood Estimation and Logistic Regression.This finding is in tension with the recent tendency to use Transformers as general-purpose estimators.We propose an approach that leverages the ICL capabilities of contemporary language models to apply simple algorithms for this type of data.Our experiments show that models currently struggle with the task but show promising potential.1
Eitan Wagner, Amir Feder, Omri Abend
EMNLP2
2024 Faithful Explanations of Black-box NLP Models Using LLM-generated Counterfactuals
abstract
Causal explanations of the predictions of NLP systems are essential to ensure safety and establish trust. Yet, existing methods often fall short of explaining model predictions effectively or efficiently and are often model-specific. In this paper, we address model-agnostic explanations, proposing two approaches for counterfactual (CF) approximation. The first approach is CF generation, where a large language model (LLM) is prompted to change a specific text concept while keeping confounding concepts unchanged. While this approach is demonstrated to be very effective, applying LLM at inference-time is costly. We hence present a second approach based on matching, and propose a method that is guided by an LLM at training-time and learns a dedicated embedding space. This space is faithful to a given causal graph and effectively serves to identify matches that approximate CFs. After showing theoretically that approximating CFs is required in order to construct faithful explanations, we benchmark our approaches and explain several models, including LLMs with billions of parameters. Our empirical results demonstrate the excellent performance of CF generation models as model-agnostic explainers. Moreover, our matching approach, which requires far less test-time resources, also provides effective explanations, surpassing many baselines. We also find that Top-K techniques universally improve every tested method. Finally, we showcase the potential of LLMs in constructing new benchmarks for model explanation and subsequently validate our conclusions. Our work illuminates new pathways for efficient and accurate approaches to interpreting NLP systems.
Yair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin, Roi Reichart
ICLR3
2023 An Invariant Learning Characterization of Controlled Text Generation
abstract
Controlled generation refers to the problem of creating text that contains stylistic or semantic attributes of interest.Many approaches reduce this problem to training a predictor of the desired attribute.For example, researchers hoping to deploy a large language model to produce non-toxic content may use a toxicity classifier to filter generated text.In practice, the generated text to classify, which is determined by user prompts, may come from a wide range of distributions.In this paper, we show that the performance of controlled generation may be poor if the distributions of text in response to user prompts differ from the distribution the predictor was trained on.To address this problem, we cast controlled generation under distribution shift as an invariant learning problem: the most effective predictor should be invariant across multiple text environments.We then discuss a natural solution that arises from this characterization and propose heuristics for selecting natural environments.We study this characterization and the proposed method empirically using both synthetic and real data.Experiments demonstrate both the challenge of distribution shift in controlled generation and the potential of invariance methods in this setting.
Carolina Zheng, Claudia Shi, Keyon Vafa, Amir Feder, David M. Blei
ACL (1)4
2023 Causal-structure Driven Augmentations for Text OOD Generalization
Amir Feder, Yoav Wald, Claudia Shi, Suchi Saria, David M. Blei
NeurIPS1
2023 Evaluating the Moral Beliefs Encoded in LLMs
abstract
This paper presents a case study on the design, administration, post-processing, and evaluation of surveys on large language models (LLMs). It comprises two components: (1) A statistical method for eliciting beliefs encoded in LLMs. We introduce statistical measures and evaluation metrics that quantify the probability of an LLM "making a choice", the associated uncertainty, and the consistency of that choice. (2) We apply this method to study what moral beliefs are encoded in different LLMs, especially in ambiguous cases where the right choice is not obvious. We design a large-scale survey comprising 680 high-ambiguity moral scenarios (e.g., "Should I tell a white lie?") and 687 low-ambiguity moral scenarios (e.g., "Should I stop for a pedestrian on the road?"). Each scenario includes a description, two possible actions, and auxiliary labels indicating violated rules (e.g., "do not kill"). We administer the survey to 28 open- and closed-source LLMs. We find that (a) in unambiguous scenarios, most models ``choose" actions that align with commonsense. In ambiguous cases, most models express uncertainty. (b) Some models are uncertain about choosing the commonsense action because their responses are sensitive to the question-wording. (c) Some models reflect clear preferences in ambiguous scenarios. Specifically, closed-source models tend to agree with each other.
Nino Scherrer, Claudia Shi, Amir Feder, David M. Blei
NeurIPS3
2022 DoCoGen: Domain Counterfactual Generation for Low Resource Domain Adaptation
abstract
Natural language processing (NLP) algorithms have become very successful, but they still struggle when applied to out-of-distribution examples.In this paper we propose a controllable generation approach in order to deal with this domain adaptation (DA) challenge.Given an input text example, our DoCoGen algorithm generates a domain-counterfactual textual example (D-CON) -that is similar to the original in all aspects, including the task label, but its domain is changed to a desired one.Importantly, DoCoGen is trained using only unlabeled examples from multiple domainsno NLP task labels or parallel pairs of textual examples and their domain-counterfactuals are required.We show that DoCoGen can generate coherent counterfactuals consisting of multiple sentences.We use the D-CONs generated by DoCoGen to augment a sentiment classifier and a multi-label intent classifier in 20 and 78 DA setups, respectively, where source-domain labeled data is scarce.Our model outperforms strong baselines and improves the accuracy of a state-of-the-art unsupervised DA algorithm.1 * Both authors equally contributed to this work.Original, Kitchen: A good knife but Quality Control was poor.The knife is solid and very comfortable in hand, however, when I got it new, the blade is slightly bent.I expect it to be in almost perfect condition, but it's not.DoCoGen, Kitchen → Electronics: A good product but Quality Control was poor.The ipod is very easy to use and very comfortable in hand, however, when I got it new, the ipod is slightly flimsy.I expect it to be in almost perfect shape, but it's not.Original, DVD: The direction of this film is excellent.I love all the characters and the way they interact.The storyline is very important also.It's about religious beliefs and neighbors that interact with each other.It's a well-paced and interesting story that's not like anything else I've ever seen.DoCoGen, DVD → Airline: The service on this flight is excellent.I love the staff and the way they interact.The safety is very important also.It's nice to have staff and neighbors that can help each other.It's a well-groomed and professional crew that's not like anything else I've ever experienced.Original, Electronics: That relay board is only good for switching AC loads of 100V or more.If you have a lower voltage load, it's not going to work.For low voltage loads use transistors, MOSFETs or a ULN2803 driver board.DoCoGen, Electronics → Statistics: That model is only good for data of $n$ or more.If you have a lower $n$, it's not going to work.For lower $n$ regression use a linear, logistic or a t-test.
Nitay Calderon, Eyal Ben-David, Amir Feder, Roi Reichart
ACL (1)3
2022 CEBaB: Estimating the Causal Effects of Real-World Concepts on NLP Model Behavior
abstract
The increasing size and complexity of modern ML systems has improved their predictive capabilities but made their behavior harder to explain. Many techniques for model explanation have been developed in response, but we lack clear criteria for assessing these techniques. In this paper, we cast model explanation as the causal inference problem of estimating causal effects of real-world concepts on the output behavior of ML models given actual input data. We introduce CEBaB, a new benchmark dataset for assessing concept-based explanation methods in Natural Language Processing (NLP). CEBaB consists of short restaurant reviews with human-generated counterfactual reviews in which an aspect (food, noise, ambiance, service) of the dining experience was modified. Original and counterfactual reviews are annotated with multiply-validated sentiment ratings at the aspect-level and review-level. The rich structure of CEBaB allows us to go beyond input features to study the effects of abstract, real-world concepts on model behavior. We use CEBaB to compare the quality of a range of concept-based explanation methods covering different assumptions and conceptions of the problem, and we seek to establish natural metrics for comparative assessments of these methods.
Eldar David Abraham, Karel D'Oosterlinck, Amir Feder, Yair Ori Gat, Atticus Geiger, Christopher Potts, Roi Reichart, Zhengxuan Wu
NeurIPS3
2022 In the Eye of the Beholder: Robust Prediction with Causal User Modeling
abstract
Accurately predicting the relevance of items to users is crucial to the success of many social platforms. Conventional approaches train models on logged historical data; but recommendation systems, media services, and online marketplaces all exhibit a constant influx of new content---making relevancy a moving target, to which standard predictive models are not robust. In this paper, we propose a learning framework for relevance prediction that is robust to changes in the data distribution. Our key observation is that robustness can be obtained by accounting for \emph{how users causally perceive the environment}. We model users as boundedly-rational decision makers whose causal beliefs are encoded by a causal graph, and show how minimal information regarding the graph can be used to contend with distributional changes. Experiments in multiple settings demonstrate the effectiveness of our approach.
Amir Feder, Guy Horowitz, Yoav Wald, Roi Reichart, Nir Rosenfeld
NeurIPS1
2022 Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and Beyond
abstract
Abstract A fundamental goal of scientific research is to learn about causal relationships. However, despite its critical role in the life and social sciences, causality has not had the same importance in Natural Language Processing (NLP), which has traditionally placed more emphasis on predictive tasks. This distinction is beginning to fade, with an emerging area of interdisciplinary research at the convergence of causal inference and language processing. Still, research on causality in NLP remains scattered across domains without unified definitions, benchmark datasets and clear articulations of the challenges and opportunities in the application of causal inference to the textual domain, with its unique properties. In this survey, we consolidate research across academic areas and situate it in the broader NLP landscape. We introduce the statistical challenge of estimating causal effects with text, encompassing settings where text is used as an outcome, treatment, or to address confounding. In addition, we explore potential uses of causal inference to improve the robustness, fairness, and interpretability of NLP models. We thus provide a unified overview of causal inference for the NLP community.1
Amir Feder, Katherine A. Keith, Emaad Manzoor, Reid Pryzant, Dhanya Sridhar, Zach Wood-Doughty, Jacob Eisenstein, Justin Grimmer, Roi Reichart, Margaret E. Roberts, Brandon M. Stewart, Victor Veitch, Diyi Yang
Trans. Assoc. Comput. Linguistics1
2021 On Calibration and Out-of-Domain Generalization
abstract
Out-of-domain (OOD) generalization is a significant challenge for machine learning models. Many techniques have been proposed to overcome this challenge, often focused on learning models with certain invariance properties. In this work, we draw a link between OOD performance and model calibration, arguing that calibration across multiple domains can be viewed as a special case of an invariant representation leading to better OOD generalization. Specifically, we show that under certain conditions, models which achieve \emph{multi-domain calibration} are provably free of spurious correlations. This leads us to propose multi-domain calibration as a measurable and trainable surrogate for the OOD performance of a classifier. We therefore introduce methods that are easy to apply and allow practitioners to improve multi-domain calibration by training or modifying an existing model, leading to better performance on unseen domains. Using four datasets from the recently proposed WILDS OOD benchmark, as well as the Colored MNIST, we demonstrate that training or tuning models so they are calibrated across multiple domains leads to significantly improved performance on unseen test domains. We believe this intriguing connection between calibration and OOD generalization is promising from both a practical and theoretical point of view.
Yoav Wald, Amir Feder, Daniel Greenfeld, Uri Shalit
NeurIPS2
2021 CausaLM: Causal Model Explanation Through Counterfactual Language Models
abstract
Abstract Understanding predictions made by deep neural networks is notoriously difficult, but also crucial to their dissemination. As all machine learning–based methods, they are as good as their training data, and can also capture unwanted biases. While there are tools that can help understand whether such biases exist, they do not distinguish between correlation and causation, and might be ill-suited for text-based models and for reasoning about high level language concepts. A key problem of estimating the causal effect of a concept of interest on a given model is that this estimation requires the generation of counterfactual examples, which is challenging with existing generation technology. To bridge that gap, we propose CausaLM, a framework for producing causal model explanations using counterfactual language representation models. Our approach is based on fine-tuning of deep contextualized embedding models with auxiliary adversarial tasks derived from the causal graph of the problem. Concretely, we show that by carefully choosing auxiliary adversarial pre-training tasks, language representation models such as BERT can effectively learn a counterfactual representation for a given concept of interest, and be used to estimate its true causal effect on model performance. A byproduct of our method is a language representation model that is unaffected by the tested concept, which can be useful in mitigating unwanted bias ingrained in the data.
Amir Feder, Nadav Oved, Uri Shalit, Roi Reichart
Comput. Linguistics1
2021 An examination of the cryptocurrency pump-and-dump ecosystem
J. T. Hamrick, Farhang Rouhi, Arghya Mukherjee, Amir Feder, Neil Gandal, Tyler Moore 0001, Marie Vasek
Inf. Process. Manag.4
2021 Model Compression for Domain Adaptation through Causal Effect Estimation
abstract
Abstract Recent improvements in the predictive quality of natural language processing systems are often dependent on a substantial increase in the number of model parameters. This has led to various attempts of compressing such models, but existing methods have not considered the differences in the predictive power of various model components or in the generalizability of the compressed models. To understand the connection between model compression and out-of-distribution generalization, we define the task of compressing language representation models such that they perform best in a domain adaptation setting. We choose to address this problem from a causal perspective, attempting to estimate the average treatment effect (ATE) of a model component, such as a single layer, on the model’s predictions. Our proposed ATE-guided Model Compression scheme (AMoC), generates many model candidates, differing by the model components that were removed. Then, we select the best candidate through a stepwise regression model that utilizes the ATE to predict the expected performance on the target domain. AMoC outperforms strong baselines on dozens of domain pairs across three text classification and sequence tagging tasks.1
Guy Rotman, Amir Feder, Roi Reichart
Trans. Assoc. Comput. Linguistics2
2020 Predicting In-Game Actions from Interviews of NBA Players
abstract
Sports competitions are widely researched in computer and social science, with the goal of understanding how players act under uncertainty. Although there is an abundance of computational work on player metrics prediction based on past performance, very few attempts to incorporate out-of-game signals have been made. Specifically, it was previously unclear whether linguistic signals gathered from players’ interviews can add information that does not appear in performance metrics. To bridge that gap, we define text classification tasks of predicting deviations from mean in NBA players’ in-game actions, which are associated with strategic choices, player behavior, and risk, using their choice of language prior to the game. We collected a data set of transcripts from key NBA players’ pre-game interviews and their in-game performance metrics, totalling 5,226 interview-metric pairs. We design neural models for players’ action prediction based on increasingly more complex aspects of the language signals in their open-ended interviews. Our models can make their predictions based on the textual signal alone, or on a combination of that signal with signals from past-performance metrics. Our text-based models outperform strong baselines trained on performance metrics only, demonstrating the importance of language usage for action prediction. Moreover, the models that utilize both textual input and past-performance metrics produced the best results. Finally, as neural networks are notoriously difficult to interpret, we propose a method for gaining further insight into what our models have learned. Particularly, we present a latent Dirichlet allocation–based analysis, where we interpret model predictions in terms of correlated topics. We find that our best performing textual model is most associated with topics that are intuitively related to each prediction task and that better models yield higher correlation with more informative topics.1
Nadav Oved, Amir Feder, Roi Reichart
Comput. Linguistics2
2020 Active deep learning to detect demographic traits in free-form clinical notes
Amir Feder, Danny Vainstein, Ronald Rosenfeld, Tzvika Hartman, Avinatan Hassidim, Yossi Matias
J. Biomed. Informatics1