VLDB 2026 Research / reviewers in the wild / expert
Tal Haklay
dblp:355/0317
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Trustworthy machine learning · 86% Transfer learning and domain adaptation · 11% Information extraction and text analysis · 3% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
2.5 | 3 | 2025 | MIB: A Mechanistic Interpretability Benchmark · ICML 2025 Position-aware Automatic Circuit Discovery · ACL (1) 2025 Linearity of Relation Decoding in Transformer Language Models · ICLR 2024 |
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability |
2.5 | 3 | 2025 | MIB: A Mechanistic Interpretability Benchmark · ICML 2025 Position-aware Automatic Circuit Discovery · ACL (1) 2025 Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking · ICLR 2024 |
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
circuit analysis |
0.9 | 1 | 2025 | Position-aware Automatic Circuit Discovery · ACL (1) 2025 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.8 | 1 | 2024 | Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking · ICLR 2024 |
Natural language and speech › Information extraction and text analysis
entity tracking |
0.2 | 1 | 2024 | Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
sparse autoencoder · 0.9large language model · 0.9edge attribution patching · 0.9distributed alignment search · 0.9attribution patching · 0.9probing · 0.8linear transformation · 0.8circuit analysis · 0.8activation patching · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Position-aware Automatic Circuit DiscoveryabstractA widely used strategy to discover and understand language model mechanisms is circuit analysis.A circuit is a minimal subgraph of a model's computation graph that executes a specific task.We identify a gap in existing circuit discovery methods: they assume circuits are position-invariant, treating model components as equally relevant across input positions.This limits their ability to capture cross-positional interactions or mechanisms that vary across positions.To address this gap, we propose two improvements to incorporate positionality into circuits, even on tasks containing variablelength examples.First, we extend edge attribution patching, a gradient-based method for circuit discovery, to differentiate between token positions.Second, we introduce the concept of a dataset schema, which defines token spans with similar semantics across examples, enabling position-aware circuit discovery in datasets with variable length examples.We additionally develop an automated pipeline for schema generation and application using large language models.Our approach enables fully automated discovery of position-sensitive circuits, yielding better trade-offs between circuit size and faithfulness compared to prior work. 1 Belinkov.2021.Causal analysis of syntactic agreement mechanisms in neural language models. Tal Haklay, Hadas Orgad, David Bau, Aaron Mueller, Yonatan Belinkov |
ACL (1) | 1 |
| 2025 | MIB: A Mechanistic Interpretability BenchmarkabstractHow can we know whether new mechanistic interpretability methods achieve real improvements?
In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretability Benchmark, with two tracks spanning four tasks and five models. MIB favors methods that precisely and concisely recover relevant causal pathways or causal variables in neural language models. The circuit localization track compares methods that locate the model components---and connections between them---most important for performing a task (e.g., attribution patching or information flow routes). The causal variable track compares methods that featurize a hidden vector, e.g., sparse autoencoders (SAE) or distributed alignment search (DAS), and align those features to a task-relevant causal variable. Using MIB, we find that attribution and mask optimization methods perform best on circuit localization. For causal variable localization, we find that the supervised DAS method performs best, while SAEs features are not better than neurons, i.e., non-featurized hidden vectors. These findings illustrate that MIB enables meaningful comparisons, and increases our confidence that there has been real progress in the field. Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, Iván Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fiotto-Kaufman, Tal Haklay, Michael Hanna 0001, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, Martin Tutek, Amir Zur, David Bau, Yonatan Belinkov |
ICML | 9 |
| 2024 | Linearity of Relation Decoding in Transformer Language ModelsabstractMuch of the knowledge encoded in transformer language models (LMs) may be expressed in terms of relations: relations between words and their synonyms, entities and their attributes, etc. We show that, for a subset of relations, this computation is well-approximated by a single linear transformation on the subject representation. Linear relation representations may be obtained by constructing a first-order approximation to the LM from a single prompt, and they exist for a variety of factual, commonsense, and linguistic relations. However, we also identify many cases in which LM predictions capture relational knowledge accurately, but this knowledge is not linearly encoded in their representations. Our results thus reveal a simple, interpretable, but heterogeneously deployed knowledge representation strategy in transformer LMs. Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, David Bau |
ICLR | 3 |
| 2024 | Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity TrackingabstractFine-tuning on generalized tasks such as instruction following, code generation, and mathematics has been shown to enhance language models' performance on a range of tasks. Nevertheless, explanations of how such fine-tuning influences the internal computations in these models remain elusive. We study how fine-tuning affects the internal mechanisms implemented in language models. As a case study, we explore the property of entity tracking, a crucial facet of language comprehension, where models fine-tuned on mathematics have substantial performance gains. We identify a mechanism that enables entity tracking and show that (i) both the original model and its fine-tuned version implement entity tracking with the same circuit. In fact, the entity tracking circuit of the fine-tuned version performs better than the full original model. (ii) The circuits of all the models implement roughly the same functionality, that is entity tracking is performed by tracking the position of the correct entity in both the original model and its fine-tuned version. (iii) Performance boost in the fine-tuned model is primarily attributed to its improved ability to handle positional information. To uncover these findings, we employ two methods: DCM, which automatically detects model components responsible for specific semantics, and CMAP, a new approach for patching activations across models to reveal improved mechanisms. Our findings suggest that fine-tuning enhances, rather than fundamentally alters, the mechanistic operation of the model. Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, David Bau |
ICLR | 3 |