Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Dana Arad

dblp:331/3635 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Trustworthy machine learning · 56% Vision and language · 18% Language models and text generation · 14%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
2.632025
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs · NeurIPS 2025
MIB: A Mechanistic Interpretability Benchmark · ICML 2025
SAEs Are Good for Steering - If You Select the Right Features · EMNLP 2025
Machine learning › Trustworthy machine learning › machine unlearning
concept unlearning
1.012026
CRISP: Persistent Concept Unlearning via Sparse Autoencoders · ACL (1) 2026
Machine learning › Trustworthy machine learning
hallucination
1.012026
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models · ACL (1) 2026
Machine learning › Trustworthy machine learning
machine unlearning
1.012026
CRISP: Persistent Concept Unlearning via Sparse Autoencoders · ACL (1) 2026
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering
0.912025
SAEs Are Good for Steering - If You Select the Right Features · EMNLP 2025
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
circuit analysis
0.912025
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.912025
MIB: A Mechanistic Interpretability Benchmark · ICML 2025
Natural language and speech › Language models and text generation
model steering
0.912025
SAEs Are Good for Steering - If You Select the Right Features · EMNLP 2025
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder
0.912025
SAEs Are Good for Steering - If You Select the Right Features · EMNLP 2025
Computer vision › Vision and language
vision-language model
0.912025
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.812024
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines · ACL (1) 2024
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
0.812024
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines · ACL (1) 2024
Natural language and speech › Language models and text generation
large language model
0.312026
CRISP: Persistent Concept Unlearning via Sparse Autoencoders · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

sparse autoencoder · 2.7prompt engineering · 1.0activation suppression · 1.0representation patching · 0.9feature scoring · 0.9distributed alignment search · 0.9circuit analysis · 0.9attribution patching · 0.9intermediate representation decoding · 0.8
YearPublicationVenuePosition
2026 CRISP: Persistent Concept Unlearning via Sparse Autoencoders
abstract
As large language models (LLMs) are increasingly deployed in real-world applications, the need to selectively remove unwanted knowledge while preserving model utility has become paramount. Recent work has explored sparse autoencoders (SAEs) to perform precise interventions on monosemantic features. However, most SAE-based methods operate at inference time, which does not create persistent changes in the model's parameters. Such interventions can be bypassed or reversed by malicious actors with parameter access. We introduce CRISP, a parameter-efficient method for persistent concept unlearning using SAEs. CRISP automatically identifies salient SAE features across multiple layers and suppresses their activations. We experiment with two LLMs and show that our method outperforms prior approaches on safety-critical unlearning tasks from the WMDP benchmark, successfully removing harmful knowledge while preserving general and in-domain capabilities. Feature-level analysis reveals that CRISP achieves semantically coherent separation between target and benign concepts, allowing precise suppression of the target features.
Tomer Ashuach, Dana Arad, Aaron Mueller, Martin Tutek, Yonatan Belinkov
ACL (1)2
2026 Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
abstract
William Rudman, Michal Golovanevsky, Dana Arad, Yonatan Belinkov, Carsten Eickhoff, Ritambhara Singh, Kyle Mahowald. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
William Rudman, Michal Golovanevsky, Dana Arad, Yonatan Belinkov, Carsten Eickhoff, Ritambhara Singh, Kyle Mahowald
ACL (1)3
2025 SAEs Are Good for Steering - If You Select the Right Features
abstract
Sparse Autoencoders (SAEs) have been proposed as an unsupervised approach to learn a decomposition of a model's latent space.This enables useful applications such as steeringinfluencing the output of a model towards a desired concept-without requiring labeled data.Current methods identify SAE features to steer by analyzing the input tokens that activate them.However, recent work has highlighted that activations alone do not fully describe the effect of a feature on the model's output.In this work, we draw a distinction between two types of features: input features, which mainly capture patterns in the model's input, and output features, which have a human-understandable effect on the model's output.We propose input and output scores to characterize and locate these types of features, and show that high values for both scores rarely co-occur in the same features.These findings have practical implications: after filtering out features with low output scores, we obtain 2-3x improvements when steering with SAEs, making them competitive with supervised methods. 1
Dana Arad, Aaron Mueller, Yonatan Belinkov
EMNLP1
2025 MIB: A Mechanistic Interpretability Benchmark
abstract
How can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretability Benchmark, with two tracks spanning four tasks and five models. MIB favors methods that precisely and concisely recover relevant causal pathways or causal variables in neural language models. The circuit localization track compares methods that locate the model components---and connections between them---most important for performing a task (e.g., attribution patching or information flow routes). The causal variable track compares methods that featurize a hidden vector, e.g., sparse autoencoders (SAE) or distributed alignment search (DAS), and align those features to a task-relevant causal variable. Using MIB, we find that attribution and mask optimization methods perform best on circuit localization. For causal variable localization, we find that the supervised DAS method performs best, while SAEs features are not better than neurons, i.e., non-featurized hidden vectors. These findings illustrate that MIB enables meaningful comparisons, and increases our confidence that there has been real progress in the field.
Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, Iván Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fiotto-Kaufman, Tal Haklay, Michael Hanna 0001, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, Martin Tutek, Amir Zur, David Bau, Yonatan Belinkov
ICML4
2025 Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
abstract
Vision-Language models (VLMs) show impressive abilities to answer questions on visual inputs (e.g., counting objects in an image), yet demonstrate higher accuracies when performing an analogous task on text (e.g., counting words in a text). We investigate this accuracy gap by identifying and comparing the circuits---the task-specific computational sub-graphs---in different modalities. We show that while circuits are largely disjoint between modalities, they implement relatively similar functionalities: the differences lie primarily in processing modality-specific data positions (an image or a text sequence). Zooming in on the image data representations, we observe they become aligned with the higher-performing analogous textual representations only towards later layers, too late in processing to effectively influence subsequent positions. To overcome this, we patch the representations of visual data tokens from later layers back into earlier layers. In experiments with multiple tasks and models, this simple intervention closes a third of the performance gap between the modalities, on average. Our analysis sheds light on the multi-modal performance gap in VLMs and suggests a training-free approach for reducing it.
Yaniv Nikankin, Dana Arad, Yossi Gandelsman, Yonatan Belinkov
NeurIPS2
2024 Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
abstract
Text-to-image diffusion models (T2I) use a latent representation of a text prompt to guide the image generation process. However, the process by which the encoder produces the text representation is unknown. We propose the Diffusion Lens, a method for analyzing the text encoder of T2I models by generating images from its intermediate representations. Using the Diffusion Lens, we perform an extensive analysis of two recent T2I models. Exploring compound prompts, we find that complex scenes describing multiple objects are composed progressively and more slowly compared to simple scenes; Exploring knowledge retrieval, we find that representation of uncommon concepts require further computation compared to common concepts, and that knowledge retrieval is gradual across layers. Overall, our findings provide valuable insights into the text encoder component in T2I pipelines.
Michael Toker, Hadas Orgad, Mor Ventura, Dana Arad, Yonatan Belinkov
ACL (1)4
2024 Predicting Fact Contributions from Query Logs with Machine Learning
Dana Arad, Daniel Deutch, Nave Frost
EDBT1
2024 ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
abstract
Dana Arad, Hadas Orgad, Yonatan Belinkov. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Dana Arad, Hadas Orgad, Yonatan Belinkov
NAACL-HLT1
2022 LearnShapley: Learning to Predict Rankings of Facts Contribution Based on Query Logs
abstract
To explain query results, a recent line of work has proposed to leverage the game-theoretic notion of Shapley values to quantify the contribution of each input fact to each result. Despite significant recent breakthroughs improving the complexity of computing Shapley values in query answering, the computation remains quite costly. To this end, we propose an approach that aims at ranking input facts based on their (hidden) Shapley values. Our method utilizes a repository of queries over the same database for which we do store exact Shapley values. Intuitively, some queries bear similarity in the ways they transform data, and consequently in the contribution of database facts to their outputs. In this manner, given a new query and a query result, we can learn and predict the ranking of contributing facts. Our contributions are three-fold. First, we introduce DBShap, a curated dataset of queries and query results, along with the contributing facts and respective Shapley values. Second, we define the task of predicting the ranking of facts contribution w.r.t a query and query result. Finally, we propose a solution for the prediction task based on BERT.
Dana Arad, Daniel Deutch, Nave Frost
CIKM1