VLDB 2026 Research / reviewers in the wild / expert
Mor Geva
dblp:203/9159
· DBLP profile ↗
41ranked-venue papers
9as first author
37since 2021 · last 2026
0000-0001-9529-6315ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 9 first-author · 37 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Constructing Interpretable Features from Compositional Neuron GroupsabstractA central goal for mechanistic interpretability has been to identify the right units of analysis in large language models (LLMs) that causally explain their outputs.While early work focused on individual neurons, evidence that neurons often encode multiple concepts has motivated a shift toward analyzing directions in activation space.A key question is how to find directions that capture interpretable features in an unsupervised manner.Current methods rely on dictionary learning with sparse autoencoders (SAEs), commonly trained over residual stream activations to learn directions from scratch.However, SAEs often struggle in causal evaluations and lack intrinsic interpretability, as their learning is not explicitly tied to the computations of the model.Here, we tackle these limitations by directly decomposing MLP activations with seminonnegative matrix factorization (SNMF), such that the learned features are (a) sparse linear combinations of co-activated neurons, and (b) mapped to their activating inputs, making them directly interpretable.Experiments on Llama 3.1, Gemma 2 and GPT-2 show that SNMF derived features outperform SAEs and a strong supervised baseline (difference-in-means) on causal steering, while aligning with humaninterpretable concepts.Further analysis reveals that specific neuron combinations are reused across semantically-related features, exposing a hierarchical structure in the MLP's activation space.Together, these results position SNMF as a simple and effective tool for identifying interpretable features and dissecting concept representations in LLMs. Or Shafran, Atticus Geiger, Mor Geva |
ACL (1) | 3 |
| 2026 | Universal Jailbreak Suffixes Are Strong Attention HijackersabstractAbstract We study suffix-based jailbreaks—a powerful family of attacks against large language models (LLMs) that optimize adversarial suffixes to circumvent safety alignment. Focusing on the widely used foun-dational GCG attack (Zou et al., 2023b), we observe that suffixes vary in efficacy: some are markedly more universal—generalizing to many unseen harmful instructions—than others. We first show that a shallow, critical mechanism drives GCG’s effectiveness. This mechanism builds on the information flow from the adversarial suffix to the final chat template tokens before generation. Quantifying the dominance of this mechanism during generation, we find GCG irregularly and aggressively hijacks the contex-tualization process. Crucially, we tie hijacking to the universality phenomenon, with more universal suffixes being stronger hijackers. Subsequently, we show that these insights have practical implications: GCG’s universality can be efficiently enhanced (up to ×5 in some cases) at no additional computational cost, and can also be surgically mitigated, at least halving the attack’s success with minimal utility loss.1 Matan Ben-Tov, Mor Geva, Mahmood Sharif |
Trans. Assoc. Comput. Linguistics | 2 |
| 2026 | LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to RepresentationsabstractAbstract Language models (LMs) increasingly drive real-world applications that require world knowledge. However, the internal processes through which models turn data into representations of knowledge and beliefs about the world are poorly understood. To facilitate such studies, we present LMEnt, a suite including (1) a knowledge-rich pretraining corpus, fully annotated with entity mentions based on Wikipedia, (2) an entity-based retrieval method over pretraining data that outperforms existing tools by as much as 80.4%, and (3) 12 pretrained LMs with up to 1B parameters and 4K intermediate checkpoints, with comparable performance to popular open-source models on knowledge tasks. Together, these resources provide a controlled environment for analyzing connections between entity mentions in pretraining data and downstream performance. We show the utility of LMEnt by studying knowledge acquisition over training, finding that entity co-occurrence and mention forms—which are difficult to study with existing tools—affect learning trends. Moreover, as LMs form stronger associations between entities, their facts are harder to edit in-context, whereas inconsistencies in model predictions over training are indicative of editing success. We release LMEnt to support studies of knowledge in LMs, including knowledge representations, plasticity, editing, attribution, hallucinations, and learning dynamics. huggingface.co/LMEnt github.com/LMEnt Daniela Gottesman, Alon Gilae-Dotan, Ido Cohen 0002, Yoav Gur-Arieh, Marius Mosbach, Ori Yoran, Mor Geva |
Trans. Assoc. Comput. Linguistics | 7 |
| 2026 | MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of DocumentsabstractAbstract Automated agents, powered by large language models (LLMs), are emerging as the go-to tool for querying information. However, evaluation benchmarks for LLM agents rarely feature natural questions that are both information-seeking and genuinely time-consuming for humans. To address this gap we introduce MoNaCo, a benchmark of 1,315 natural and time-consuming questions that require dozens, and at times hundreds, of intermediate steps to solve— far more than any existing QA benchmark. To build MoNaCo, we developed a decomposed annotation pipeline to elicit and manually answer real-world time-consuming questions at scale. Frontier LLMs evaluated on MoNaCo achieve at most 61.2% F1, hampered by low recall and hallucinations. Our results underscore the limitations of LLM-powered agents in handling the complexity and sheer breadth of real-world information-seeking tasks—with MoNaCo providing an effective resource for tracking such progress. The MoNaCo benchmark, codebase, prompts, and models predictions are all publicly available at: https://tomerwolgithub.github.io/monaco. Tomer Wolfson, Harsh Trivedi, Mor Geva, Yoav Goldberg, Dan Roth 0001, Tushar Khot, Ashish Sabharwal, Reut Tsarfaty |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language ModelsabstractVision-language models (VLMs) excel at extracting and reasoning about information from images.Yet, their capacity to leverage internal knowledge about specific entities remains underexplored.This work investigates the disparity in model performance when answering factual questions about an entity described in text versus depicted in an image.Our results reveal a significant accuracy drop -reaching 18% for some models -when the entity is presented visually instead of textually.To study this gap we present POPVQA, a dataset which allows separating entity recognition and question answering, and use it to benchmark several models.We hypothesize that this decline arises from limitations in how information flows from image tokens to query tokens.Thus, we use mechanistic interpretability tools to reveal that, although image tokens are preprocessed by the vision encoder, meaningful information flow from these tokens occurs only in the much deeper layers.Furthermore, critical image processing happens in the language model's middle layers, allowing few layers for consecutive reasoning, highlighting a potential inefficiency in how the model utilizes its layers for reasoning.These insights shed light on the internal mechanics of VLMs and offer pathways for enhancing their reasoning capabilities.POPVQA can be found at this link. Ido Cohen 0002, Daniela Gottesman, Mor Geva, Raja Giryes |
ACL (1) | 3 |
| 2025 | Inferring Functionality of Attention Heads from their ParametersabstractAttention heads are one of the building blocks of large language models (LLMs).Prior work on investigating their operation mostly focused on analyzing their behavior during inference for specific circuits or tasks.In this work, we seek a comprehensive mapping of the operations they implement in a model.We propose MAPS (Mapping Attention head Param-eterS), an efficient framework that infers the functionality of attention heads from their parameters, without any model training or inference.We showcase the utility of MAPS for answering two types of questions: (a) given a predefined operation, mapping how strongly heads across the model implement it, and (b) given an attention head, inferring its salient functionality.Evaluating MAPS on 20 operations across 6 popular LLMs shows its estimations correlate with the head's outputs during inference and are causally linked to the model's predictions.Moreover, its mappings reveal attention heads of certain operations that were overlooked in previous studies, and valuable insights on function universality and architecture biases in LLMs.Next, we present an automatic pipeline and analysis that leverage MAPS to characterize the salient operations of a given head.Our pipeline produces plausible operation descriptions for most heads, as assessed by human judgment, while revealing diverse operations.We release our code and mappings Amit Elhelo, Mor Geva |
ACL (1) | 2 |
| 2025 | Enhancing Automated Interpretability with Output-Centric Feature DescriptionsabstractAutomated interpretability pipelines generate natural language descriptions for the concepts represented by features in large language models (LLMs), such as “plants” or “the first word in a sentence”. These descriptions are derived using inputs that activate the feature, which may be a dimension or a direction in the model’s representation space. However, identifying activating inputs is costly, and the mechanistic role of a feature in model behavior is determined both by how inputs cause a feature to activate and by how feature activation affects outputs. Using steering evaluations, we reveal that current pipelines provide descriptions that fail to capture the causal effect of the feature on outputs. To fix this, we propose efficient, output-centric methods for automatically generating feature descriptions. These methods use the tokens weighted higher after feature stimulation or the highest weight tokens after applying the vocabulary “unembedding” head directly to the feature. Our output-centric descriptions better capture the causal effect of a feature on model outputs than input-centric descriptions, but combining the two leads to the best performance on both input and output evaluations. Lastly, we show that output-centric descriptions can be used to find inputs that activate features previously thought to be “dead”. Yoav Gur-Arieh, Roy Mayan, Chen Agassy, Atticus Geiger, Mor Geva |
ACL (1) | 5 |
| 2025 | Precise In-Parameter Concept Erasure in Large Language ModelsabstractLarge language models (LLMs) often acquire knowledge during pretraining that is undesirable in downstream deployments, e.g., sensitive information or copyrighted content.Existing approaches for removing such knowledge rely on fine-tuning, training low-rank adapters or fact-level editing, but these are either too coarse, too shallow, or ineffective.In this work, we propose PISCES (Precise Inparameter Suppression for Concept EraSure), a novel framework for precisely erasing entire concepts from model parameters by directly editing directions that encode them in parameter space.PISCES uses a disentangler model to decompose MLP vectors into interpretable features, identifies those associated with a target concept using automated interpretability techniques, and removes them from model parameters.Experiments on Gemma 2 and Llama 3.1 over various concepts show that PISCES achieves modest gains in efficacy over leading erasure methods, reducing accuracy on the target concept to as low as 7.7%, while dramatically improving erasure specificity (by up to 31%) and robustness (by up to 38%).Overall, these results demonstrate that feature-based in-parameter editing enables a more precise and reliable approach for removing conceptual knowledge in language models. Yoav Gur-Arieh, Clara Suslik, Yihuai Hong, Fazl Barez, Mor Geva |
EMNLP | 5 |
| 2025 | Intrinsic Test of Unlearning Using Parametric Knowledge TracesabstractThe task of "unlearning" certain concepts in large language models (LLMs) has gained attention for its role in mitigating harmful, private, or incorrect outputs.Current evaluations mostly rely on behavioral tests, without monitoring residual knowledge in model parameters, which can be adversarially exploited to recover erased information.We argue that unlearning should also be assessed internally by tracking changes in the parametric traces of unlearned concepts.To this end, we propose a general evaluation methodology that uses vocabulary projections to inspect concepts encoded in model parameters.We apply this approach to localize "concept vectors" -parameter vectors encoding concrete conceptsand construct CONCEPTVECTORS, a benchmark of hundreds of such concepts and their parametric traces in two open-source LLMs.Evaluation on CONCEPTVECTORS shows that existing methods minimally alter concept vectors, mostly suppressing them at inference time, while direct ablation of these vectors removes the associated knowledge and reduces adversarial susceptibility.Our findings reveal limitations of behavior-only evaluations and advocate for parameter-based assessments.We release our code and benchmark at https://github. com/yihuaihong/ConceptVectors.* Bias term is omitted for brevity. Yihuai Hong, Haiqin Yang, Shauli Ravfogel, Mor Geva |
EMNLP | 5 |
| 2025 | Towards Interpreting Visual Information Processing in Vision-Language ModelsabstractVision-Language Models (VLMs) are powerful tools for processing and understanding text and images. We study the processing of visual tokens in the language model component of LLaVA, a prominent VLM. Our approach focuses on analyzing the localization of object information, the evolution of visual token representations across layers, and the mechanism of integrating visual information for predictions. Through ablation studies, we demonstrated that object identification accuracy drops by over 70\% when object-specific tokens are removed. We observed that visual token representations become increasingly interpretable in the vocabulary space across layers, suggesting an alignment with textual tokens corresponding to image content. Finally, we found that the model extracts object information from these refined representations at the last token position for prediction, mirroring the process in text-only language models for factual association tasks. These findings provide crucial insights into how VLMs process and integrate visual information, bridging the gap between our understanding of language and vision models, and paving the way for more interpretable and controllable multimodal systems. Clement Neo, C.-H. Luke Ong, Philip Torr 0001, Mor Geva, David Krueger 0001, Fazl Barez |
ICLR | 4 |
| 2025 | Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus AreasabstractLarge Vision Language Models (VLMs) have long struggled with spatial reasoning tasks. Surprisingly, even simple spatial reasoning tasks, such as recognizing “under” or “behind” relationships between only two objects, pose significant challenges for current VLMs. We believe it is crucial to use the lens of mechanism interpretability, opening up the model and diving into model’s internal states to examine the interactions between image and text tokens during spatial reasoning. Our analysis of attention behaviors reveals significant differences in how VLMs allocate attention to image versus text. By tracing the areas of images that receive the highest attention scores throughout intermediate layers, we observe a notable pattern: errors often coincide with attention being misdirected towards irrelevant objects within the image. Moreover, such attention patterns exhibit substantial differences between familiar (e.g., “on the left side of ”) and unfamiliar (e.g.,“in front of ”) spatial relationships. Motivated by these findings, we propose ADAPTVIS based on inference-time confidence scores to sharpen the attention on highly relevant regions when the model exhibits high confidence, while smoothing and broadening the attention window to consider a wider context when confidence is lower. This training-free decoding method shows significant improvement (e.g., up to a 50 absolute point improvement) on spatial reasoning benchmarks such as WhatsUp and VSR with negligible additional cost. Shiqi Chen 0002, Tongyao Zhu, Ruochen Zhou, Jinghan Zhang 0006, Siyang Gao, Juan Carlos Niebles, Mor Geva, Junxian He, Jiajun Wu 0001, Manling Li |
ICML | 7 |
| 2024 | The Hidden Space of Transformer Language AdaptersabstractWe analyze the operation of transformer language adapters, which are small modules trained on top of a frozen language model to adapt its predictions to new target languages.We show that adapted predictions mostly evolve in the source language the model was trained on, while the target language becomes pronounced only in the very last layers of the model.Moreover, the adaptation process is gradual and distributed across layers, where it is possible to skip small groups of adapters without decreasing adaptation performance.Last, we show that adapters operate on top of the model's frozen representation space while largely preserving its structure, rather than on an "isolated" subspace.Our findings provide a deeper view into the adaptation process of language models to new languages, showcasing the constraints imposed on it by the underlying model and introduces practical implications to enhance its efficiency. 1 Jesujoba O. Alabi, Marius Mosbach, Matan Eyal, Dietrich Klakow, Mor Geva |
ACL (1) | 5 |
| 2024 | RAVEL: Evaluating Interpretability Methods on Disentangling Language Model RepresentationsabstractIndividual neurons participate in the representation of multiple high-level concepts.To what extent can different interpretability methods successfully disentangle these roles?To help address this question, we introduce RAVEL (Resolving Attribute-Value Entanglements in Language Models), a dataset that enables tightly controlled, quantitative comparisons between a variety of existing interpretability methods.We use the resulting conceptual framework to define the new method of Multi-task Distributed Alignment Search (MDAS), which allows us to find distributed representations satisfying multiple causal criteria.With Llama2-7B as the target language model, MDAS achieves state-of-the-art results on RAVEL, demonstrating the importance of going beyond neuron-level analyses to identify features distributed across activations.We release our benchmark at https://github.com/ explanare/ravel. Jing Huang 0014, Zhengxuan Wu, Christopher Potts, Mor Geva, Atticus Geiger |
ACL (1) | 4 |
| 2024 | A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning ChainsabstractAlon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig, Or Honovich, Michael Tseng, Michael Collins, Roee Aharoni, Mor Geva. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Alon Jacovi, Yonatan Bitton, Bernd Bohnet, Jonathan Herzig, Or Honovich, Michael Tseng, Michael Collins 0001, Roee Aharoni, Mor Geva |
ACL (1) | 9 |
| 2024 | Do Large Language Models Latently Perform Multi-Hop Reasoning?abstractWe study whether Large Language Models (LLMs) latently perform multi-hop reasoning with complex prompts such as "The mother of the singer of 'Superstition' is".We look for evidence of a latent reasoning pathway where an LLM (1) latently identifies "the singer of 'Superstition"' as Stevie Wonder, the bridge entity, and (2) uses its knowledge of Stevie Wonder's mother to complete the prompt.We analyze these two hops individually and consider their co-occurrence as indicative of latent multi-hop reasoning.For the first hop, we test if changing the prompt to indirectly mention the bridge entity instead of any other entity increases the LLM's internal recall of the bridge entity.For the second hop, we test if increasing this recall causes the LLM to better utilize what it knows about the bridge entity.We find strong evidence of latent multi-hop reasoning for the prompts of certain relation types, with the reasoning pathway used in more than 80% of the prompts.However, the utilization is highly contextual, varying across different types of prompts.Also, on average, the evidence for the second hop and the full multi-hop traversal is rather moderate and only substantial for the first hop.Moreover, we find a clear scaling trend with increasing model size for the first hop of reasoning but not for the second hop.Our experimental findings suggest potential challenges and opportunities for future development and applications of LLMs. Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, Sebastian Riedel 0001 |
ACL (1) | 4 |
| 2024 | Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity AnswersabstractFactual questions can typically be answered correctly at different levels of granularity.For example, both "August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?".Standard question answering (QA) evaluation protocols, however, do not take this into account explicitly and instead compare a predicted answer against reference answers of a single granularity level.In this work, we propose GRANOLA QA, a novel evaluation setting where a predicted answer is evaluated in terms of accuracy and informativeness against a set of multigranularity answers.We present a simple methodology for enriching existing datasets with multi-granularity answers, and create GRANOLA-EQ, a multi-granularity version of the ENTITYQUESTIONS dataset.We evaluate models using a range of decoding methods on GRANOLA-EQ, including a new algorithm called Decoding with Response Aggregation (DRAG), that is geared towards aligning the answer granularity with the model's uncertainty.Our experiments show that large language models with standard decoding methods tend to generate specific answers, which are often incorrect.In contrast, when evaluated on multi-granularity answers, DRAG yields a nearly 20 point increase in accuracy on average, which further increases for rare entities, revealing that standard evaluation and decoding schemes may underestimate the knowledge encapsulated in language models. Gal Yona, Roee Aharoni, Mor Geva |
ACL (1) | 3 |
| 2024 | Jump to Conclusions: Short-Cutting Transformers with Linear TransformationsabstractTransformer-based language models create hidden representations of their inputs at every layer, but only use final-layer representations for prediction. This obscures the internal decision-making process of the model and the utility of its intermediate representations. One way to elucidate this is to cast the hidden representations as final representations, bypassing the transformer computation in-between. In this work, we suggest a simple method for such casting, using linear transformations. This approximation far exceeds the prevailing practice of inspecting hidden representations from all layers, in the space of the final layer. Moreover, in the context of language modeling, our method produces more accurate predictions from hidden layers, across various model scales, architectures, and data distributions. This allows “peeking” into intermediate representations, showing that GPT-2 and BERT often predict the final output already in early layers. We then demonstrate the practicality of our method to recent early exit strategies, showing that when aiming, for example, at retention of 95% accuracy, our approach saves additional 7.9% layers for GPT-2 and 5.4% layers for BERT. Last, we extend our method to linearly approximate sub-modules, finding that attention is most tolerant to this change. Our code and learned mappings are publicly available at https://github.com/sashayd/mat. Alexander Yom Din, Taelin Karidi, Leshem Choshen, Mor Geva |
LREC/COLING | 4 |
| 2024 | Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop QueriesabstractLarge language models (LLMs) can solve complex multi-step problems, but little is known about how these computations are implemented internally.Motivated by this, we study how LLMs answer multi-hop queries such as "The spouse of the performer of Imagine is".These queries require two information extraction steps: a latent one for resolving the first hop ("the performer of Imagine") into the bridge entity (John Lennon), and another for resolving the second hop ("the spouse of John Lennon") into the target entity (Yoko Ono).Understanding how the latent step is computed internally is key to understanding the overall computation.By carefully analyzing the internal computations of transformer-based LLMs, we discover that the bridge entity is resolved in the early layers of the model.Then, only after this resolution, the two-hop query is solved in the later layers.Because the second hop commences in later layers, there could be cases where these layers no longer encode the necessary knowledge for correctly predicting the answer.Motivated by this, we propose a novel "back-patching" analysis method whereby a hidden representation from a later layer is patched back to an earlier layer.We find that in up to 66% of previously incorrect cases there exists a back-patch that results in the correct generation of the answer, showing that the later layers indeed sometimes lack the needed functionality.Overall, our methods and findings open further opportunities for understanding and improving latent reasoning in transformer-based LLMs. Eden Biran, Daniela Gottesman, Sohee Yang, Mor Geva, Amir Globerson |
EMNLP | 4 |
| 2024 | Estimating Knowledge in Large Language Models Without Generating a Single TokenabstractTo evaluate knowledge in large language models (LLMs), current methods query the model and then evaluate its generated responses.In this work, we ask whether evaluation can be done before the model has generated any text.Concretely, is it possible to estimate how knowledgeable a model is about a certain entity, only from its internal computation?We study this question with two tasks: given a subject entity, the goal is to predict (a) the ability of the model to answer common questions about the entity, and (b) the factuality of open-ended responses generated by the model about the entity.Experiments with a variety of LLMs show that KEEN, a simple probe trained over internal subject representations, succeeds at both tasks -correlating with both the QA accuracy of the model per-subject and FActScore, a recent factuality metric in open-ended generation.Moreover, KEEN naturally aligns with the model's hedging behavior and faithfully reflects changes in the model's knowledge after fine-tuning.Lastly, we show a more interpretable yet equally performant variant of KEEN, which highlights a small set of tokens indicative of clusters and gaps in the model's knowledge.Being simple and lightweight, KEEN can be leveraged to guide decisions such as when it is appropriate to apply further training or augment queries with retrieval. Daniela Gottesman, Mor Geva |
EMNLP | 2 |
| 2024 | Backward Lens: Projecting Language Model Gradients into the Vocabulary SpaceabstractUnderstanding how Transformer-based Language Models (LMs) learn and recall information is a key goal of the deep learning community.Recent interpretability methods project weights and hidden states obtained from the forward pass to the models' vocabularies, helping to uncover how information flows within LMs.In this work, we extend this methodology to LMs' backward pass and gradients.We first prove that a gradient matrix can be cast as a low-rank linear combination of its forward and backward passes' inputs.We then develop methods to project these gradients into vocabulary items and explore the mechanics of how new information is stored in the LMs' neurons.Our code is available Shahar Katz, Yonatan Belinkov, Mor Geva, Lior Wolf |
EMNLP | 3 |
| 2024 | From Insights to Actions: The Impact of Interpretability and Analysis Research on NLPabstractInterpretability and analysis (IA) research is a growing subfield within NLP with the goal of developing a deeper understanding of the behavior or inner workings of NLP systems and methods.Despite growing interest in the subfield, a criticism of this work is that it lacks actionable insights and therefore has little impact on NLP.In this paper, we seek to quantify the impact of IA research on the broader field of NLP.We approach this with a mixed-methods analysis 1 of: (1) a citation graph of 185K+ papers built from all papers published at ACL and EMNLP conferences from 2018 to 2023, and their references and citations, and (2) a survey of 138 members of the NLP community.Our quantitative results show that IA work is well-cited outside of IA, and central in the NLP citation graph.Through qualitative analysis of survey responses and manual annotation of 556 papers, we find that NLP researchers build on findings from IA work and perceive it as important for progress in NLP, multiple subfields, and rely on its findings and terminology for their own work.Many novel methods are proposed based on IA findings and highly influenced by them, but highly influential non-IA work cites IA findings without being driven by them.We end by summarizing what is missing in IA work today and provide a call to action, to pave the way for a more impactful future of IA research. Marius Mosbach, Vagrant Gautam, Tomás Vergara Browne, Dietrich Klakow, Mor Geva |
EMNLP | 5 |
| 2024 | Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?abstractWe posit that large language models (LLMs) should be capable of expressing their intrinsic uncertainty in natural language.For example, if the LLM is equally likely to output two contradicting answers to the same question, then its generated response should reflect this uncertainty by hedging its answer (e.g., "I'm not sure, but I think...").We formalize faithful response uncertainty based on the gap between the model's intrinsic confidence in the assertions it makes and the decisiveness by which they are conveyed.This example-level metric reliably indicates whether the model reflects its uncertainty, as it penalizes both excessive and insufficient hedging.We evaluate a variety of aligned LLMs at faithfully communicating uncertainty on several knowledgeintensive question answering tasks.Our results provide strong evidence that modern LLMs are poor at faithfully conveying their uncertainty, and that better alignment is necessary to improve their trustworthiness. Gal Yona, Roee Aharoni, Mor Geva |
EMNLP | 3 |
| 2024 | The Hidden Language of Diffusion ModelsabstractText-to-image diffusion models have demonstrated an unparalleled ability to generate high-quality, diverse images from a textual prompt. However, the internal representations learned by these models remain an enigma. In this work, we present Conceptor, a novel method to interpret the internal representation of a textual concept by a diffusion model. This interpretation is obtained by decomposing the concept into a small set of human-interpretable textual elements. Applied over the state-of-the-art Stable Diffusion model, Conceptor reveals non-trivial structures in the representations of concepts. For example, we find surprising visual connections between concepts, that transcend their textual semantics. We additionally discover concepts that rely on mixtures of exemplars, biases, renowned artistic styles, or a simultaneous fusion of multiple meanings of the concept.
Through a large battery of experiments, we demonstrate Conceptor's ability to provide meaningful, robust, and faithful decompositions for a wide variety of abstract, concrete, and complex textual concepts, while allowing to naturally connect each decomposition element to its corresponding visual impact on the generated images. Hila Chefer, Oran Lang, Mor Geva, Volodymyr Polosukhin, Assaf Shocher, Michal Irani, Inbar Mosseri, Lior Wolf |
ICLR | 3 |
| 2024 | Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsabstractUnderstanding the internal representations of large language models (LLMs) can help explain models' behavior and verify their alignment with human values. Given the capabilities of LLMs in generating human-understandable text, we propose leveraging the model itself to explain its internal representations in natural language. We introduce a framework called Patchscopes and show how it can be used to answer a wide range of questions about an LLM's computation. We show that many prior interpretability methods based on projecting representations into the vocabulary space and intervening on the LLM computation can be viewed as instances of this framework. Moreover, several of their shortcomings such as failure in inspecting early layers or lack of expressivity can be mitigated by Patchscopes. Beyond unifying prior inspection techniques, Patchscopes also opens up *new* possibilities such as using a more capable model to explain the representations of a smaller model, and multihop reasoning error correction. Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, Mor Geva |
ICML | 5 |
| 2024 | Evaluating the Ripple Effects of Knowledge Editing in Language ModelsabstractAbstract Modern language models capture a large body of factual knowledge. However, some facts can be incorrectly induced or become obsolete over time, resulting in factually incorrect generations. This has led to the development of various editing methods that allow updating facts encoded by the model. Evaluation of these methods has primarily focused on testing whether an individual fact has been successfully injected, and if similar predictions for other subjects have not changed. Here we argue that such evaluation is limited, since injecting one fact (e.g., “Jack Depp is the son of Johnny Depp”) introduces a “ripple effect” in the form of additional facts that the model needs to update (e.g., “Jack Depp is the sibling of Lily-Rose Depp”). To address this, we propose novel evaluation criteria that consider the implications of an edit on related facts. Using these criteria, we then construct RippleEdits, a diagnostic benchmark of 5K factual edits, capturing various types of ripple effects. We evaluate prominent editing methods on RippleEdits, showing that they fail to introduce consistent changes in the model’s knowledge. In addition, we find that a simple in-context editing baseline obtains the best scores on our benchmark, suggesting a promising research direction for model editing.1 Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, Mor Geva |
Trans. Assoc. Comput. Linguistics | 5 |
| 2023 | Analyzing Transformers in Embedding SpaceabstractUnderstanding Transformer-based models has attracted significant attention, as they lie at the heart of recent technological advances across machine learning.While most interpretability methods rely on running models over inputs, recent work has shown that an inputindependent approach, where parameters are interpreted directly without a forward/backward pass is feasible for some Transformer parameters, and for two-layer attention networks.In this work, we present a conceptual framework where all parameters of a trained Transformer are interpreted by projecting them into the embedding space, that is, the space of vocabulary items they operate on.Focusing mostly on GPT-2 for this paper, we provide diverse evidence to support our argument.First, an empirical analysis showing that parameters of both pretrained and fine-tuned models can be interpreted in embedding space.Second, we present two applications of our framework: (a) aligning the parameters of different models that share a vocabulary, and (b) constructing a classifier without training by "translating" the parameters of a fine-tuned classifier to parameters of a different model that was only pretrained.Overall, our findings show that at least in part, we can abstract away model specifics and understand Transformers in the embedding space.EA EB Layer 18 Head 1 ('women', ' Marie') (' actresses', ' Marie') ('women', ' Anne') ('Women', ' Anne') ('woman', ' Marie') ('Women', ' Marie') Guy Dar, Mor Geva, Ankit Gupta 0001, Jonathan Berant |
ACL (1) | 2 |
| 2023 | Understanding Transformer Memorization Recall Through IdiomsabstractTo produce accurate predictions, language models (LMs) must balance between generalization and memorization.Yet, little is known about the mechanism by which transformer LMs employ their memorization capacity.When does a model decide to output a memorized phrase, and how is this phrase then retrieved from memory?In this work, we offer the first methodological framework for probing and characterizing recall of memorized sequences in transformer LMs.First, we lay out criteria for detecting model inputs that trigger memory recall, and propose idioms as inputs that typically fulfill these criteria.Next, we construct a dataset of English idioms and use it to compare model behavior on memorized vs. non-memorized inputs.Specifically, we analyze the internal prediction construction process by interpreting the model's hidden representations as a gradual refinement of the output probability distribution.We find that across different model sizes and architectures, memorized predictions are a two-step process: early layers promote the predicted token to the top of the output distribution, and upper layers increase model confidence.This suggests that memorized information is stored and retrieved in the early layers of the network.Last, we demonstrate the utility of our methodology beyond idioms in memorized factual statements.Overall, our work makes a first step towards understanding memory recall, and provides a methodological basis for future studies of transformer memorization.1 Adi Haviv, Ido Cohen 0002, Jacob Gidron, Roei Schuster, Yoav Goldberg, Mor Geva |
EACL | 6 |
| 2023 | Don't Blame the Annotator: Bias Already Starts in the Annotation InstructionsabstractIn recent years, progress in NLU has been driven by benchmarks.These benchmarks are typically collected by crowdsourcing, where annotators write examples based on annotation instructions crafted by dataset creators.In this work, we hypothesize that annotators pick up on patterns in the crowdsourcing instructions, which bias them to write many similar examples that are then over-represented in the collected data.We study this form of bias, termed instruction bias, in 14 recent NLU benchmarks, showing that instruction examples often exhibit concrete patterns, which are propagated by crowdworkers to the collected data.This extends previous work (Geva et al., 2019) and raises a new concern of whether we are modeling the dataset creator's instructions, rather than the task.Through a series of experiments, we show that, indeed, instruction bias can lead to overestimation of model performance, and that models struggle to generalize beyond biases originating in the crowdsourcing instructions.We further analyze the influence of instruction bias in terms of pattern frequency and model size, and derive concrete recommendations for creating future NLU benchmarks. 1 Mihir Parmar, Swaroop Mishra, Mor Geva, Chitta Baral |
EACL | 3 |
| 2023 | LM vs LM: Detecting Factual Errors via Cross ExaminationabstractA prominent weakness of modern language models (LMs) is their tendency to generate factually incorrect text, which hinders their usability.A natural question is whether such factual errors can be detected automatically.Inspired by truth-seeking mechanisms in law, we propose a factuality evaluation framework for LMs that is based on cross-examination.Our key idea is that an incorrect claim is likely to result in inconsistency with other claims that the model generates.To discover such inconsistencies, we facilitate a multi-turn interaction between the LM that generated the claim and another LM (acting as an examiner) which introduces questions to discover inconsistencies.We empirically evaluate our method on factual claims made by multiple recent LMs on four benchmarks, finding that it outperforms existing methods and baselines, often by a large gap.Our results demonstrate the potential of using interacting LMs to capture factual errors. Roi Cohen, May Hamri, Mor Geva, Amir Globerson |
EMNLP | 3 |
| 2023 | Dissecting Recall of Factual Associations in Auto-Regressive Language ModelsabstractTransformer-based language models (LMs) are known to capture factual knowledge in their parameters.While previous work looked into where factual associations are stored, only little is known about how they are retrieved internally during inference.We investigate this question through the lens of information flow.Given a subject-relation query, we study how the model aggregates information about the subject and relation to predict the correct attribute.With interventions on attention edges, we first identify two critical points where information propagates to the prediction: one from the relation positions followed by another from the subject positions.Next, by analyzing the information at these points, we unveil a three-step internal mechanism for attribute extraction.First, the representation at the lastsubject position goes through an enrichment process, driven by the early MLP sublayers, to encode many subject-related attributes.Second, information from the relation propagates to the prediction.Third, the prediction representation "queries" the enriched subject to extract the attribute.Perhaps surprisingly, this extraction is typically done via attention heads, which often encode subject-attribute mappings in their parameters.Overall, our findings introduce a comprehensive view of how factual associations are stored and extracted internally in LMs, facilitating future research on knowledge localization and editing. 1 Mor Geva, Jasmijn Bastings, Katja Filippova, Amir Globerson |
EMNLP | 1 |
| 2023 | CRoW: Benchmarking Commonsense Reasoning in Real-World TasksabstractRecent efforts in natural language processing (NLP) commonsense reasoning research have yielded a considerable number of new datasets and benchmarks.However, most of these datasets formulate commonsense reasoning challenges in artificial scenarios that are not reflective of the tasks which real-world NLP systems are designed to solve.In this work, we present CROW, a manually-curated, multitask benchmark that evaluates the ability of models to apply commonsense reasoning in the context of six real-world NLP tasks.CROW is constructed using a multi-stage data collection pipeline that rewrites examples from existing datasets using commonsense-violating perturbations.We use CROWto study how NLP systems perform across different dimensions of commonsense knowledge, such as physical, temporal, and social reasoning.We find a significant performance gap when NLP systems are evaluated on CROWcompared to humans, showcasing that commonsense reasoning is far from being solved in real-world task settings.We make our dataset and leaderboard available to the research community.1 * Equal contribution 1 https://github.com/mismayil/crowDialogue Agent: Hi, would you like some free candies?Human: Sure.What are you handing these out for?Agent: Well, we're trying to gather some Mete Ismayilzada, Debjit Paul, Syrielle Montariol, Mor Geva, Antoine Bosselut |
EMNLP | 4 |
| 2022 | SCROLLS: Standardized CompaRison Over Long Language SequencesabstractUri Shaham, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta, Wenhan Xiong, Mor Geva, Jonathan Berant, Omer Levy. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Uri Shaham 0002, Elad Segal, Maor Ivgi, Avia Efrat, Ori Yoran, Adi Haviv, Ankit Gupta 0001, Wenhan Xiong, Mor Geva, Jonathan Berant, Omer Levy |
EMNLP | 9 |
| 2022 | Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceabstractTransformer-based language models (LMs) are at the core of modern NLP, but their internal prediction construction process is opaque and largely not understood.In this work, we make a substantial step towards unveiling this underlying prediction process, by reverseengineering the operation of the feed-forward network (FFN) layers, one of the building blocks of transformer models.We view the token representation as a changing distribution over the vocabulary, and the output from each FFN layer as an additive update to that distribution.Then, we analyze the FFN updates in the vocabulary space, showing that each update can be decomposed to sub-updates corresponding to single FFN parameter vectors, each promoting concepts that are often human-interpretable.We then leverage these findings for controlling LM predictions, where we reduce the toxicity of GPT2 by almost 50%, and for improving computation efficiency with a simple early exit rule, saving 20% of computation on average. 1 Mor Geva, Avi Caciularu, Kevin Ro Wang, Yoav Goldberg |
EMNLP | 1 |
| 2022 | Break, Perturb, Build: Automatic Perturbation of Reasoning Paths Through Question DecompositionabstractAbstract Recent efforts to create challenge benchmarks that test the abilities of natural language understanding models have largely depended on human annotations. In this work, we introduce the “Break, Perturb, Build” (BPB) framework for automatic reasoning-oriented perturbation of question-answer pairs. BPB represents a question by decomposing it into the reasoning steps that are required to answer it, symbolically perturbs the decomposition, and then generates new question-answer pairs. We demonstrate the effectiveness of BPB by creating evaluation sets for three reading comprehension (RC) benchmarks, generating thousands of high-quality examples without human intervention. We evaluate a range of RC models on our evaluation sets, which reveals large performance gaps on generated examples compared to the original data. Moreover, symbolic perturbations enable fine-grained analysis of the strengths and limitations of models. Last, augmenting the training data with examples generated by BPB helps close the performance gaps, without any drop on the original data distribution. Mor Geva, Tomer Wolfson, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | What's in Your Head? Emergent Behaviour in Multi-Task Transformer ModelsabstractThe primary paradigm for multi-task training in natural language processing is to represent the input with a shared pre-trained language model, and add a small, thin network (head) per task.Given an input, a target head is the head that is selected for outputting the final prediction.In this work, we examine the behaviour of non-target heads, that is, the output of heads when given input that belongs to a different task than the one they were trained for.We find that non-target heads exhibit emergent behaviour, which may either explain the target task, or generalize beyond their original task.For example, in a numerical reasoning task, a span extraction head extracts from the input the arguments to a computation that results in a number generated by a target generative head.In addition, a summarization head that is trained with a target question answering head, outputs query-based summaries when given a question and a context from which the answer is to be extracted.This emergent behaviour suggests that multi-task training leads to nontrivial extrapolation of skills, which can be harnessed for interpretability and generalization. Mor Geva, Uri Katz, Aviv Ben-Arie, Jonathan Berant |
EMNLP (1) | 1 |
| 2021 | Transformer Feed-Forward Layers Are Key-Value MemoriesabstractFeed-forward layers constitute two-thirds of a transformer model's parameters, yet their role in the network remains under-explored.We show that feed-forward layers in transformerbased language models operate as key-value memories, where each key correlates with textual patterns in the training examples, and each value induces a distribution over the output vocabulary.Our experiments show that the learned patterns are human-interpretable, and that lower layers tend to capture shallow patterns, while upper layers learn more semantic ones.The values complement the keys' input patterns by inducing output distributions that concentrate probability mass on tokens likely to appear immediately after each pattern, particularly in the upper layers.Finally, we demonstrate that the output of a feed-forward layer is a composition of its memories, which is subsequently refined throughout the model's layers via residual connections to produce the final output distribution. Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy |
EMNLP (1) | 1 |
| 2021 | Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning StrategiesabstractAbstract A key limitation in current datasets for multi-hop reasoning is that the required steps for answering the question are mentioned in it explicitly. In this work, we introduce StrategyQA, a question answering (QA) benchmark where the required reasoning steps are implicit in the question, and should be inferred using a strategy. A fundamental challenge in this setup is how to elicit such creative questions from crowdsourcing workers, while covering a broad range of potential strategies. We propose a data collection procedure that combines term-based priming to inspire annotators, careful control over the annotator population, and adversarial filtering for eliminating reasoning shortcuts. Moreover, we annotate each question with (1) a decomposition into reasoning steps for answering it, and (2) Wikipedia paragraphs that contain the answers to each step. Overall, StrategyQA includes 2,780 examples, each consisting of a strategy question, its decomposition, and evidence paragraphs. Analysis shows that questions in StrategyQA are short, topic-diverse, and cover a wide range of strategies. Empirically, we show that humans perform well (87%) on this task, while our best baseline reaches an accuracy of ∼ 66%. Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth 0001, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Injecting Numerical Reasoning Skills into Language ModelsabstractLarge pre-trained language models (LMs) are known to encode substantial amounts of linguistic information.However, high-level reasoning skills, such as numerical reasoning, are difficult to learn from a language-modeling objective only.Consequently, existing models for numerical reasoning have used specialized architectures with limited flexibility.In this work, we show that numerical reasoning is amenable to automatic data generation, and thus one can inject this skill into pre-trained LMs, by generating large amounts of data, and training in a multi-task setup.We show that pre-training our model, GENBERT, on this data, dramatically improves performance on DROP (49.3 → 72.3 F 1 ), reaching performance that matches state-of-the-art models of comparable size, while using a simple and general-purpose encoder-decoder architecture.Moreover, GENBERT generalizes well to math word problem datasets, while maintaining high performance on standard RC tasks.Our approach provides a general recipe for injecting skills into large pre-trained LMs, whenever the skill is amenable to automatic data augmentation.* These authors contributed equally.(b) fine-tuning pre-trained LM numerical reasoning reading compr. Mor Geva, Ankit Gupta 0001, Jonathan Berant |
ACL | 1 |
| 2020 | Break It Down: A Question Understanding BenchmarkabstractUnderstanding natural language questions entails the ability to break down a question into the requisite steps for computing its answer. In this work, we introduce a Question Decomposition Meaning Representation (QDMR) for questions. QDMR constitutes the ordered list of steps, expressed through natural language, that are necessary for answering a question. We develop a crowdsourcing pipeline, showing that quality QDMRs can be annotated at scale, and release the Break dataset, containing over 83K pairs of questions and their QDMRs. We demonstrate the utility of QDMR by showing that (a) it can be used to improve open-domain question answering on the HotpotQA dataset, (b) it can be deterministically converted to a pseudo-SQL formal language, which can alleviate annotation in semantic parsing applications. Last, we use Break to train a sequence-to-sequence model with copying that parses questions into QDMR structures, and show that it substantially outperforms several natural baselines. Tomer Wolfson, Mor Geva, Ankit Gupta 0001, Yoav Goldberg, Matt Gardner 0001, Daniel Deutch, Jonathan Berant |
Trans. Assoc. Comput. Linguistics | 2 |
| 2019 | Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding DatasetsabstractMor Geva, Yoav Goldberg, Jonathan Berant. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mor Geva, Yoav Goldberg, Jonathan Berant |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Learning to Search in Long Documents Using Document StructureabstractReading comprehension models are based on recurrent neural networks that sequentially process the document tokens. As interest turns to answering more complex questions over longer documents, sequential reading of large portions of text becomes a substantial bottleneck. Inspired by how humans use document structure, we propose a novel framework for reading comprehension. We represent documents as trees, and model an agent that learns to interleave quick navigation through the document tree with more expensive answer extraction. To encourage exploration of the document tree, we propose a new algorithm, based on Deep Q-Network (DQN), which strategically samples tree nodes at training time. Empirically we find our algorithm improves question answering performance compared to DQN and a strong information-retrieval (IR) baseline, and that ensembling our model with the IR baseline results in further gains in performance. Mor Geva, Jonathan Berant |
COLING | 1 |