Alon Jacovi

dblp:218/5900 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0002-7263-2061ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 8 first-author · 10 since 2021
YearPublicationVenuePosition
2025 ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated Simulatability
abstract
Concept-based explanations work by mapping complex model computations to humanunderstandable concepts.Evaluating such explanations is very difficult, as it includes not only the quality of the induced space of possible concepts but also how effectively the chosen concepts are communicated to users.Existing evaluation metrics often focus solely on the former, neglecting the latter.
Antonin Poché, Alon Jacovi, Agustin M. Picard, Victor Boutin, Fanny Jourdan
ACL (1)2
2024 A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
abstract
Alon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig, Or Honovich, Michael Tseng, Michael Collins, Roee Aharoni, Mor Geva. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Alon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig, Or Honovich, Michael Tseng, Michael Collins 0001, Roee Aharoni, Mor Geva
ACL (1)1
2024 Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP
abstract
Improvements in language models' capabilities have pushed their applications towards longer contexts, making long-context evaluation and development an active research area.However, many disparate use cases are grouped together under the umbrella term of "long-context", defined simply by the total length of the model's input, including -for example -Needle-in-a-Haystack tasks, book summarization, and information aggregation.Given their varied difficulty, in this position paper we argue that conflating different tasks by their context length is unproductive.As a community, we require a more precise vocabulary to understand what makes long-context tasks similar or different.We propose to unpack the taxonomy of longcontext based on the properties that make them more difficult with longer contexts.We propose two orthogonal axes of difficulty: (I) Dispersion: How hard is it to find the necessary information in the context?(II) Scope: How much necessary information is there to find?We survey the literature on long context, provide justification for this taxonomy as an informative descriptor, and situate the literature with respect to it.We conclude that the most difficult and interesting settings, whose necessary information is very long and highly dispersed within the input, is severely under-explored.By using a descriptive vocabulary and discussing the relevant properties of difficulty in long context, we can implement more informed research in this area.We call for a careful design of tasks and benchmarks with distinctly long context, taking into account the characteristics that make it qualitatively different from shorter context.
Omer Goldman, Alon Jacovi, Aviv Slobodkin, Aviya Maimon, Ido Dagan, Reut Tsarfaty
EMNLP2
2024 TACT: Advancing Complex Aggregative Reasoning with Information Extraction Tools
abstract
Large Language Models (LLMs) often do not perform well on queries that require the aggregation of information across texts. To better evaluate this setting and facilitate modeling efforts, we introduce TACT - Text And Calculations through Tables, a dataset crafted to evaluate LLMs' reasoning and computational abilities using complex instructions. TACT contains challenging instructions that demand stitching information scattered across one or more texts, and performing complex integration on this information to generate the answer. We construct this dataset by leveraging an existing dataset of texts and their associated tables. For each such tables, we formulate new queries, and gather their respective answers. We demonstrate that all contemporary LLMs perform poorly on this dataset, achieving an accuracy below 38%. To pinpoint the difficulties and thoroughly dissect the problem, we analyze model performance across three components: table-generation, Pandas command-generation, and execution. Unexpectedly, we discover that each component presents substantial challenges for current LLMs. These insights lead us to propose a focused modeling framework, which we refer to as IE as a tool. Specifically, we propose to add "tools" for each of the above steps, and implement each such tool with few-shot prompting. This approach shows an improvement over existing prompting techniques, offering a promising direction for enhancing model capabilities in these tasks.
Avi Caciularu, Alon Jacovi, Eyal Ben-David, Sasha Goldshtein, Tal Schuster, Jonathan Herzig, Gal Elidan, Amir Globerson
NeurIPS2
2023 Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks
abstract
Data contamination has become prevalent and challenging with the rise of models pretrained on large automatically-crawled corpora.For closed models, the training data becomes a trade secret, and even for open models, it is not trivial to detect contamination.Strategies such as leaderboards with hidden answers, or using test data which is guaranteed to be unseen, are expensive and become fragile with time.Assuming that all relevant actors value clean test data and will cooperate to mitigate data contamination, what can be done?We propose three strategies that can make a difference: (1) Test data made public should be encrypted with a public key and licensed to disallow derivative distribution; (2) demand training exclusion controls from closed API holders, and protect your test data by refusing to evaluate without them; (3) avoid data which appears with its solution on the internet, and release the web-page context of internet-derived data along with the data.These strategies are practical and can be effective in preventing data contamination.
Alon Jacovi, Avi Caciularu, Omer Goldman, Yoav Goldberg
EMNLP1
2023 Diagnosing AI Explanation Methods with Folk Concepts of Behavior
Alon Jacovi, Jasmijn Bastings, Sebastian Gehrmann, Yoav Goldberg, Katja Filippova
J. Artif. Intell. Res.1
2021 Scalable Evaluation and Improvement of Document Set Expansion via Neural Positive-Unlabeled Learning
abstract
We consider the situation in which a user has collected a small set of documents on a cohesive topic, and they want to retrieve additional documents on this topic from a large collection.Information Retrieval (IR) solutions treat the document set as a query, and look for similar documents in the collection.We propose to extend the IR approach by treating the problem as an instance of positive-unlabeled (PU) learning-i.e., learning binary classifiers from only positive (the query documents) and unlabeled (the results of the IR engine) data.Utilizing PU learning for text with big neural networks is a largely unexplored field.We discuss various challenges in applying PU learning to the setting, showing that the standard implementations of state-of-the-art PU solutions fail.We propose solutions for each of the challenges and empirically validate them with ablation tests.We demonstrate the effectiveness of the new method using a series of experiments of retrieving PubMed abstracts adhering to fine-grained topics, showing improvements over the common IR solution and other baselines.
Alon Jacovi, Gang Niu 0001, Yoav Goldberg, Masashi Sugiyama
EACL1
2021 Contrastive Explanations for Model Interpretability
abstract
Contrastive explanations clarify why an event occurred in contrast to another.They are inherently intuitive to humans to both produce and comprehend.We propose a method to produce contrastive explanations in the latent space, via a projection of the input representation, such that only the features that differentiate two potential decisions are captured.Our modification allows model behavior to consider only contrastive reasoning, and uncover which aspects of the input are useful for and against particular decisions.Additionally, for a given input feature, our contrastive explanations can answer for which label, and against which alternative label, is the feature useful.We produce contrastive explanations via both highlevel abstract concept attribution and low-level input token/span attribution for two NLP classification benchmarks.Our findings demonstrate the ability of label-contrastive explanations to provide fine-grained interpretability of model decisions.1
Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi 0001, Yoav Goldberg
EMNLP (1)1
2021 Amnesic Probing: Behavioral Explanation With Amnesic Counterfactuals
abstract
Abstract A growing body of work makes use of probing in order to investigate the working of neural models, often considered black boxes. Recently, an ongoing debate emerged surrounding the limitations of the probing paradigm. In this work, we point out the inability to infer behavioral conclusions from probing results, and offer an alternative method that focuses on how the information is being used, rather than on what information is encoded. Our method, Amnesic Probing, follows the intuition that the utility of a property for a given task can be assessed by measuring the influence of a causal intervention that removes it from the representation. Equipped with this new analysis tool, we can ask questions that were not possible before, for example, is part-of-speech information important for word prediction? We perform a series of analyses on BERT to answer these types of questions. Our findings demonstrate that conventional probing performance is not correlated to task importance, and we call for increased scrutiny of claims that draw behavioral or causal conclusions from probing results.1
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, Yoav Goldberg
Trans. Assoc. Comput. Linguistics3
2021 Aligning Faithful Interpretations with their Social Attribution
abstract
Abstract We find that the requirement of model interpretations to be faithful is vague and incomplete. With interpretation by textual highlights as a case study, we present several failure cases. Borrowing concepts from social science, we identify that the problem is a misalignment between the causal chain of decisions (causal attribution) and the attribution of human behavior to the interpretation (social attribution). We reformulate faithfulness as an accurate attribution of causality to the model, and introduce the concept of aligned faithfulness: faithful causal chains that are aligned with their expected social behavior. The two steps of causal attribution and social attribution together complete the process of explaining behavior. With this formalization, we characterize various failures of misaligned faithful highlight interpretations, and propose an alternative causal chain to remedy the issues. Finally, we implement highlight explanations of the proposed causal format using contrastive explanations.
Alon Jacovi, Yoav Goldberg
Trans. Assoc. Comput. Linguistics1
2020 Towards Faithfully Interpretable NLP Systems: How Should We Define and Evaluate Faithfulness?
abstract
With the growing popularity of deep-learning based NLP models, comes a need for interpretable systems.But what is interpretability, and what constitutes a high-quality interpretation?In this opinion piece we reflect on the current state of interpretability evaluation research.We call for more clearly differentiating between different desired criteria an interpretation should satisfy, and focus on the faithfulness criteria.We survey the literature with respect to faithfulness evaluation, and arrange the current approaches around three assumptions, providing an explicit form to how faithfulness is "defined" by the community.We provide concrete guidelines on how evaluation of interpretation methods should and should not be conducted.Finally, we claim that the current binary definition for faithfulness sets a potentially unrealistic bar for being considered faithful.We call for discarding the binary notion of faithfulness in favor of a more graded one, which we believe will be of greater practical utility.
Alon Jacovi, Yoav Goldberg
ACL1
2020 Exposing Shallow Heuristics of Relation Extraction Models with Challenge Data
abstract
The process of collecting and annotating training data may introduce distribution artifacts which may limit the ability of models to learn correct generalization behavior.We identify failure modes of SOTA relation extraction (RE) models trained on TACRED, which we attribute to limitations in the data annotation process.We collect and annotate a challengeset we call Challenging RE (CRE), based on naturally occurring corpus examples, to benchmark this behavior.Our experiments with four state-of-the-art RE models show that they have indeed adopted shallow heuristics that do not generalize to the challenge-set data.Further, we find that alternative question answering modeling performs significantly better than the SOTA models on the challenge-set, despite worse overall TACRED performance.By adding some of the challenge data as training examples, the performance of the model improves.Finally, we provide concrete suggestion on how to improve RE data collection to alleviate this behavior.
Shachar Rosenman, Alon Jacovi, Yoav Goldberg
EMNLP (1)2
2019 Neural network gradient-based learning of black-box function interfaces
Alon Jacovi, Guy Hadash, Einat Kermany, Boaz Carmeli, Ofer Lavi, George Kour, Jonathan Berant
ICLR (Poster)1