Lee Sharkey

dblp:294/0023 · also Lee D. Sharkey · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 76% Representation and self-supervised learning · 12% Deep learning architectures and training · 7%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
3.342025
Bilinear MLPs enable weight-based mechanistic interpretability · ICLR 2025
Sparse Autoencoders Do Not Find Canonical Units of Analysis · ICLR 2025
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning · NeurIPS 2024
Machine learning › Trustworthy machine learning
interpretability
2.432025
Sparse Autoencoders Do Not Find Canonical Units of Analysis · ICLR 2025
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning · NeurIPS 2024
Sparse Autoencoders Find Highly Interpretable Features in Language Models · ICLR 2024
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder
2.432025
Sparse Autoencoders Do Not Find Canonical Units of Analysis · ICLR 2025
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning · NeurIPS 2024
Sparse Autoencoders Find Highly Interpretable Features in Language Models · ICLR 2024
Machine learning › Deep learning architectures and training
activation function
0.912025
Bilinear MLPs enable weight-based mechanistic interpretability · ICLR 2025
Machine learning › Representation and self-supervised learning › feature transformation
feature decomposition
0.912025
Sparse Autoencoders Do Not Find Canonical Units of Analysis · ICLR 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding › dictionary learning
sparse dictionary learning
0.812024
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning · NeurIPS 2024
Machine learning › Trustworthy machine learning › neural network interpretability
superposition
0.812024
Sparse Autoencoders Find Highly Interpretable Features in Language Models · ICLR 2024
Machine learning › Trustworthy machine learning
robustness
0.612022
Goal Misgeneralization in Deep Reinforcement Learning · ICML 2022
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift
0.612022
Goal Misgeneralization in Deep Reinforcement Learning · ICML 2022

Methods — techniques the papers use, named apart from their topics

tensor decomposition · 0.9meta-SAE · 0.9eigendecomposition · 0.9SAE stitching · 0.9BatchTopK SAE · 0.9sparse autoencoder · 0.8KL divergence minimization · 0.8
YearPublicationVenuePosition
2025 Sparse Autoencoders Do Not Find Canonical Units of Analysis
abstract
A common goal of mechanistic interpretability is to decompose the activations of neural networks into features: interpretable properties of the input computed by the model. Sparse autoencoders (SAEs) are a popular method for finding these features in LLMs, and it has been postulated that they can be used to find a canonical set of units: a unique and complete list of atomic features. We cast doubt on this belief using two novel techniques: SAE stitching to show they are incomplete, and meta-SAEs to show they are not atomic. SAE stitching involves inserting or swapping latents from a larger SAE into a smaller one. Latents from the larger SAE can be divided into two categories: novel latents, which improve performance when added to the smaller SAE, indicating they capture novel information, and reconstruction latents, which can replace corresponding latents in the smaller SAE that have similar behavior. The existence of novel features indicates incompleteness of smaller SAEs. Using meta-SAEs - SAEs trained on the decoder matrix of another SAE - we find that latents in SAEs often decompose into combinations of latents from a smaller SAE, showing that larger SAE latents are not atomic. The resulting decompositions are often interpretable; e.g. a latent representing "Einstein" decomposes into "scientist", "Germany", and "famous person". To train meta-SAEs we introduce BatchTopK SAEs, an improved variant of the popular TopK SAE method, that only enforces a fixed average sparsity. Even if SAEs do not find canonical units of analysis, they may still be useful tools. We suggest that future research should either pursue different approaches for identifying such units, or pragmatically choose the SAE size suited to their task. We provide an interactive dashboard to explore meta-SAEs: https://metasaes.streamlit.app/
Patrick Leask, Bart Bussmann, Michael T. Pearce, Joseph Isaac Bloom, Curt Tigges, Noura Al Moubayed, Lee Sharkey, Neel Nanda
ICLR7
2025 Bilinear MLPs enable weight-based mechanistic interpretability
abstract
A mechanistic understanding of how MLPs do computation in deep neural net- works remains elusive. Current interpretability work can extract features from hidden activations over an input dataset but generally cannot explain how MLP weights construct features. One challenge is that element-wise nonlinearities introduce higher-order interactions and make it difficult to trace computations through the MLP layer. In this paper, we analyze bilinear MLPs, a type of Gated Linear Unit (GLU) without any element-wise nonlinearity that neverthe- less achieves competitive performance. Bilinear MLPs can be fully expressed in terms of linear operations using a third-order tensor, allowing flexible analysis of the weights. Analyzing the spectra of bilinear MLP weights using eigendecom- position reveals interpretable low-rank structure across toy tasks, image classifi- cation, and language modeling. We use this understanding to craft adversarial examples, uncover overfitting, and identify small language model circuits directly from the weights alone. Our results demonstrate that bilinear layers serve as an interpretable drop-in replacement for current activation functions and that weight- based interpretability is viable for understanding deep-learning models.
Michael T. Pearce, Thomas Dooms, Alice Rigg, José Oramas M., Lee Sharkey
ICLR5
2024 Sparse Autoencoders Find Highly Interpretable Features in Language Models
abstract
One of the roadblocks to a better understanding of neural networks' internals is \textit{polysemanticity}, where neurons appear to activate in multiple, semantically distinct contexts. Polysemanticity prevents us from identifying concise, human-understandable explanations for what neural networks are doing internally. One hypothesised cause of polysemanticity is \textit{superposition}, where neural networks represent more features than they have neurons by assigning features to an overcomplete set of directions in activation space, rather than to individual neurons. Here, we attempt to identify those directions, using sparse autoencoders to reconstruct the internal activations of a language model. These autoencoders learn sets of sparsely activating features that are more interpretable and monosemantic than directions identified by alternative approaches, where interpretability is measured by automated methods. Moreover, we show that with our learned set of features, we can pinpoint the features that are causally responsible for counterfactual behaviour on the indirect object identification task \citep{wang2022interpretability} to a finer degree than previous decompositions. This work indicates that it is possible to resolve superposition in language models using a scalable, unsupervised method. Our method may serve as a foundation for future mechanistic interpretability work, which we hope will enable greater model transparency and steerability.
Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, Lee Sharkey
ICLR5
2024 Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning
abstract
Identifying the features learned by neural networks is a core challenge in mechanistic interpretability. Sparse autoencoders (SAEs), which learn a sparse, overcomplete dictionary that reconstructs a network's internal activations, have been used to identify these features. However, SAEs may learn more about the structure of the datatset than the computational structure of the network. There is therefore only indirect reason to believe that the directions found in these dictionaries are functionally important to the network. We propose end-to-end (e2e) sparse dictionary learning, a method for training SAEs that ensures the features learned are functionally important by minimizing the KL divergence between the output distributions of the original model and the model with SAE activations inserted. Compared to standard SAEs, e2e SAEs offer a Pareto improvement: They explain more network performance, require fewer total features, and require fewer simultaneously active features per datapoint, all with no cost to interpretability. We explore geometric and qualitative differences between e2e SAE features and standard SAE features. E2e dictionary learning brings us closer to methods that can explain network behavior concisely and accurately. We release our library for training e2e SAEs and reproducing our analysis at https://github.com/ApolloResearch/e2e_sae.
Dan Braun, Jordan Taylor, Nicholas Goldowsky-Dill, Lee Sharkey
NeurIPS4
2022 Goal Misgeneralization in Deep Reinforcement Learning
abstract
We study goal misgeneralization, a type of out-of-distribution robustness failure in reinforcement learning (RL). Goal misgeneralization occurs when an RL agent retains its capabilities out-of-distribution yet pursues the wrong goal. For instance, an agent might continue to competently avoid obstacles, but navigate to the wrong place. In contrast, previous works have typically focused on capability generalization failures, where an agent fails to do anything sensible at test time.We provide the first explicit empirical demonstrations of goal misgeneralization and present a partial characterization of its causes.
Lauro Langosco di Langosco, Jack Koch, Lee Sharkey, Jacob Pfau, David Krueger 0001
ICML3