Jack Merullo

dblp:248/8361 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
10since 2021 · last 2025
0009-0005-9673-6809ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Trustworthy machine learning · 38% Representation and self-supervised learning · 24% Language models and text generation · 17%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 19 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
3.352025
On Linear Representations and Pretraining Data Frequency in Language Models · ICLR 2025
Talking Heads: Understanding Inter-Layer Communication in Transformer Language Models · NeurIPS 2024
Circuit Component Reuse Across Tasks in Transformer Language Models · ICLR 2024
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
2.232024
Talking Heads: Understanding Inter-Layer Communication in Transformer Language Models · NeurIPS 2024
Circuit Component Reuse Across Tasks in Transformer Language Models · ICLR 2024
Characterizing Mechanisms for Factual Recall in Language Models · EMNLP 2023
Machine learning › Transfer learning and domain adaptation › feature-based transfer learning
feature transfer
0.912025
Transferring Linear Features Across Language Models With Model Stitching · NeurIPS 2025
Natural language and speech › Language models and text generation
in-context learning
0.912025
Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight Forgetting · ICLR 2025
Machine learning › Representation and self-supervised learning
in-weights learning
0.912025
Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight Forgetting · ICLR 2025
Machine learning › Representation and self-supervised learning
linear representation
0.912025
On Linear Representations and Pretraining Data Frequency in Language Models · ICLR 2025
Machine learning › Representation and self-supervised learning
model stitching
0.912025
Transferring Linear Features Across Language Models With Model Stitching · NeurIPS 2025
Machine learning › Representation and self-supervised learning
pre-training
0.912025
Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight Forgetting · ICLR 2025
Machine learning › Representation and self-supervised learning › pre-training
pretraining data
0.912025
On Linear Representations and Pretraining Data Frequency in Language Models · ICLR 2025
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder
0.912025
Transferring Linear Features Across Language Models With Model Stitching · NeurIPS 2025
Natural language and speech › Language models and text generation
text representation
0.912025
Transferring Linear Features Across Language Models With Model Stitching · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
circuit analysis
0.812024
Circuit Component Reuse Across Tasks in Transformer Language Models · ICLR 2024
Information retrieval › retrieval models
axiomatic retrieval model
0.812024
Axiomatic Causal Interventions for Reverse Engineering Relevance Computation in Neural Retrieval Models · SIGIR 2024
Information retrieval › retrieval models › neural retrieval
neural ranking model
0.812024
Axiomatic Causal Interventions for Reverse Engineering Relevance Computation in Neural Retrieval Models · SIGIR 2024
Information retrieval
retrieval models
0.812024
Axiomatic Causal Interventions for Reverse Engineering Relevance Computation in Neural Retrieval Models · SIGIR 2024
Computer vision › Vision and language
cross-modal alignment
0.712023
Linearly Mapping from Image to Text Space · ICLR 2023
Natural language and speech › Language models and text generation › large language model › knowledge in language models
factual recall
0.712023
Characterizing Mechanisms for Factual Recall in Language Models · EMNLP 2023
Computer vision › Vision and language
multimodal representation
0.712023
Linearly Mapping from Image to Text Space · ICLR 2023
Natural language and speech › Information extraction and text analysis › corpus linguistics
corpus analysis
0.112019
Investigating Sports Commentator Bias within a Large Corpus of American Football Broadcasts · EMNLP/IJCNLP (1) 2019

Methods — techniques the papers use, named apart from their topics

steering vector transfer · 0.9regression model · 0.9pretraining modulation · 0.9model stitching · 0.9affine mapping · 0.9active forgetting · 0.9singular value decomposition · 0.8mechanistic interpretability · 0.8circuit analysis · 0.8causal intervention · 0.8attention head intervention · 0.8attention head analysis · 0.8activation analysis · 0.8
YearPublicationVenuePosition
2025 Dual Process Learning: Controlling Use of In-Context vs. In-Weights Strategies with Weight Forgetting
abstract
Language models have the ability to perform in-context learning (ICL), allowing them to flexibly adapt their behavior based on context. This contrasts with in-weights learning (IWL), where memorized information is encoded in model parameters after iterated observations of data. An ideal model should be able to flexibly deploy both of these abilities. Despite their apparent ability to learn in-context, language models are known to struggle when faced with unseen or rarely seen tokens (Land & Bartolo, 2024). Hence, we study $\textbf{structural in-context learning}$, which we define as the ability of a model to execute in-context learning on arbitrary novel tokens -- so called because the model must generalize on the basis of e.g. sentence structure or task structure, rather than content encoded in token embeddings. We study structural in-context algorithms on both synthetic and naturalistic tasks using toy models, masked language models, and autoregressive language models. We find that structural ICL appears before quickly disappearing early in LM pretraining. While it has been shown that ICL can diminish during training (Singh et al., 2023), we find that prior work does not account for structural ICL. Building on Chen et al. (2024) 's active forgetting method, we introduce pretraining and finetuning methods that can modulate the preference for structural ICL and IWL. Importantly, this allows us to induce a $\textit{dual process strategy}$ where in-context and in-weights solutions coexist within a single model.
Suraj Anand, Michael A. Lepori, Jack Merullo, Ellie Pavlick
ICLR3
2025 On Linear Representations and Pretraining Data Frequency in Language Models
abstract
Pretraining data has a direct impact on the behaviors and quality of language models (LMs), but we only understand the most basic principles of this relationship. While most work focuses on pretraining data's effect on downstream task behavior, we investigate its relationship to LM representations. Previous work has discovered that, in language models, some concepts are encoded "linearly" in the representations, but what factors cause these representations to form (or not)? We study the connection between pretraining data frequency and models' linear representations of factual relations (e.g., mapping France to Paris in a capital prediction task). We find evidence that the formation of linear representations is strongly connected to pretraining term frequencies; specifically for subject-relation-object fact triplets, both subject-object co-occurrence frequency and in-context learning accuracy for the relation are highly correlated with linear representations. This is the case across all phases of pretraining, i.e., it is not affected by the model's underlying capability. In OLMo-7B and GPT-J (6B), we discover that a linear representation consistently (but not exclusively) forms when the subjects and objects within a relation co-occur at least 1k and 2k times, respectively, regardless of when these occurrences happen during pretraining (and around 4k times for OLMo-1B). Finally, we train a regression model on measurements of linear representation quality in fully-trained LMs that can predict how often a term was seen in pretraining. Our model achieves low error even on inputs from a different model with a different pretraining dataset, providing a new method for estimating properties of the otherwise-unknown training data of closed-data models. We conclude that the strength of linear representations in LMs contains signal about the models' pretraining corpora that may provide new avenues for controlling and improving model behavior: particularly, manipulating the models' training data to meet specific frequency thresholds. We release our code to support future work.
Jack Merullo, Noah A. Smith, Sarah Wiegreffe, Yanai Elazar
ICLR1
2025 Transferring Linear Features Across Language Models With Model Stitching
abstract
In this work, we demonstrate that affine mappings between residual streams of language models is a cheap way to effectively transfer represented features between models. We apply this technique to transfer the \textit{weights} of Sparse Autoencoders (SAEs) between models of different sizes to compare their representations. We find that small and large models learn highly similar representation spaces, which motivates training expensive components like SAEs on a smaller model and transferring to a larger model at a FLOPs savings. For example, using a small-to-large transferred SAE as initialization can lead to 50% cheaper training runs when training SAEs on larger models. Next, we show that transferred probes and steering vectors can effectively recover ground truth performance. Finally, we dive deeper into feature-level transferability, finding that semantic and structural features transfer noticeably differently while specific classes of functional features have their roles faithfully mapped. Overall, our findings illustrate similarities and differences in the linear representation spaces of small and large models and demonstrate a method for improving the training efficiency of SAEs.
Alan Chen 0008, Jack Merullo, Alessandro Stolfo, Ellie Pavlick
NeurIPS2
2024 Transformer Mechanisms Mimic Frontostriatal Gating Operations When Trained on Human Working Memory Tasks
Aaron Traylor, Jack Merullo, Michael J. Frank, Ellie Pavlick
CogSci2
2024 Circuit Component Reuse Across Tasks in Transformer Language Models
abstract
Recent work in mechanistic interpretability has shown that behaviors in language models can be successfully reverse-engineered through circuit analysis. A common criticism, however, is that each circuit is task-specific, and thus such analysis cannot contribute to understanding the models at a higher level. In this work, we present evidence that insights (both low-level findings about specific heads and higher-level findings about general algorithms) can indeed generalize across tasks. Specifically, we study the circuit discovered in (Wang, 2022) for the Indirect Object Identification (IOI) task and 1.) show that it reproduces on a larger GPT2 model, and 2.) that it is mostly reused to solve a seemingly different task: Colored Objects (Ippolito & Callison-Burch, 2023). We provide evidence that the process underlying both tasks is functionally very similar, and contains about a 78% overlap in in-circuit attention heads. We further present a proof-of-concept intervention experiment, in which we adjust four attention heads in middle layers in order to ‘repair’ the Colored Objects circuit and make it behave like the IOI circuit. In doing so, we boost accuracy from 49.6% to 93.7% on the Colored Objects task and explain most sources of error. The intervention affects downstream attention heads in specific ways predicted by their interactions in the IOI circuit, indicating that this subcircuit behavior is invariant to the different task inputs. Overall, our results provide evidence that it may yet be possible to explain large language models' behavior in terms of a relatively small number of interpretable task-general algorithmic building blocks and computational components.
Jack Merullo, Carsten Eickhoff, Ellie Pavlick
ICLR1
2024 Language Models Implement Simple Word2Vec-style Vector Arithmetic
abstract
Jack Merullo, Carsten Eickhoff, Ellie Pavlick. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jack Merullo, Carsten Eickhoff, Ellie Pavlick
NAACL-HLT1
2024 Talking Heads: Understanding Inter-Layer Communication in Transformer Language Models
abstract
Although it is known that transformer language models (LMs) pass features from early layers to later layers, it is not well understood how this information is represented and routed by the model. We analyze a mechanism used in two LMs to selectively inhibit items in a context in one task, and find that it underlies a commonly used abstraction across many context-retrieval behaviors. Specifically, we find that models write into low-rank subspaces of the residual stream to represent features which are then read out by later layers, forming low-rank *communication channels* (Elhage et al., 2021) between layers. A particular 3D subspace in model activations in GPT-2 can be traversed to positionally index items in lists, and we show that this mechanism can explain an otherwise arbitrary-seeming sensitivity of the model to the order of items in the prompt. That is, the model has trouble copying the correct information from context when many items ``crowd" this limited space. By decomposing attention heads with the Singular Value Decomposition (SVD), we find that previously described interactions between heads separated by one or more layers can be predicted via analysis of their weight matrices alone. We show that it is possible to manipulate the internal model representations as well as edit model weights based on the mechanism we discover in order to significantly improve performance on our synthetic Laundry List task, which requires recall from a list, often improving task accuracy by over 20\%. Our analysis reveals a surprisingly intricate interpretable structure learned from language model pretraining, and helps us understand why sophisticated LMs sometimes fail in simple domains, facilitating future analysis of more complex behaviors.
Jack Merullo, Carsten Eickhoff, Ellie Pavlick
NeurIPS1
2024 Axiomatic Causal Interventions for Reverse Engineering Relevance Computation in Neural Retrieval Models
abstract
Neural models have demonstrated remarkable performance across diverse ranking tasks. However, the processes and internal mechanisms along which they determine relevance are still largely unknown. Existing approaches for analyzing neural ranker behavior with respect to IR properties rely either on assessing overall model behavior or employing probing methods that may offer an incomplete understanding of causal mechanisms. To provide a more granular understanding of internal model decision-making processes, we propose the use of causal interventions to reverse engineer neural rankers, and demonstrate how mechanistic interpretability methods can be used to isolate components satisfying term-frequency axioms within a ranking model. We identify a group of attention heads that detect duplicate tokens in earlier layers of the model, then communicate with downstream heads to compute overall document relevance. More generally, we propose that this style of mechanistic analysis opens up avenues for reverse engineering the processes neural retrieval models use to compute relevance. This work aims to initiate granular interpretability efforts that will not only benefit retrieval model development and training, but ultimately ensure safer deployment of these models.
Catherine Chen 0001, Jack Merullo, Carsten Eickhoff
SIGIR2
2023 Characterizing Mechanisms for Factual Recall in Language Models
abstract
Language Models (LMs) often must integrate facts they memorized in pretraining with new information that appears in a given context.These two sources can disagree, causing competition within the model, and it is unclear how an LM will resolve the conflict.On a dataset that queries for knowledge of world capitals, we investigate both distributional and mechanistic determinants of LM behavior in such situations.Specifically, we measure the proportion of the time an LM will use a counterfactual prefix (e.g., "The capital of Poland is London") to overwrite what it learned in pretraining ("Warsaw").On Pythia and GPT2, the training frequency of both the query country ("Poland") and the in-context city ("London") highly affect the models' likelihood of using the counterfactual.We then use head attribution to identify individual attention heads that either promote the memorized answer or the in-context answer in the logits.By scaling up or down the value vector of these heads, we can control the likelihood of using the in-context answer on new data.This method can increase the rate of generating the in-context answer to 88% of the time simply by scaling a single head at runtime.Our work contributes to a body of evidence showing that we can often localize model behaviors to specific components and provides a proof of concept for how future methods might control model behavior dynamically at runtime.
Qinan Yu, Jack Merullo, Ellie Pavlick
EMNLP2
2023 Linearly Mapping from Image to Text Space
Jack Merullo, Louis Castricato, Carsten Eickhoff, Ellie Pavlick
ICLR1
2019 Investigating Sports Commentator Bias within a Large Corpus of American Football Broadcasts
abstract
Jack Merullo, Luke Yeh, Abram Handler, Alvin Grissom II, Brendan O’Connor, Mohit Iyyer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jack Merullo, Luke Yeh, Abram Handler, Alvin Grissom II, Brendan T. O'Connor 0001, Mohit Iyyer
EMNLP/IJCNLP (1)1