Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Arthur Conmy

dblp:304/2706 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Trustworthy machine learning · 79% Language models and text generation · 10% Transfer learning and domain adaptation · 6%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
4.662025
Scaling Sparse Feature Circuits For Studying In-Context Learning · ICML 2025
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability · ICML 2025
Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
2.842024
Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders · NeurIPS 2024
Successor Heads: Recurring, Interpretable Attention Heads In The Wild · ICLR 2024
Towards Automated Circuit Discovery for Mechanistic Interpretability · NeurIPS 2023
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder
1.932025
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability · ICML 2025
Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders · NeurIPS 2024
Scaling Sparse Feature Circuits For Studying In-Context Learning · ICML 2025
Natural language and speech › Language models and text generation
in-context learning
0.912025
Scaling Sparse Feature Circuits For Studying In-Context Learning · ICML 2025
Machine learning › Trustworthy machine learning
language model interpretability
0.912025
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability · ICML 2025
Machine learning › Transfer learning and domain adaptation
task vector
0.912025
Scaling Sparse Feature Circuits For Studying In-Context Learning · ICML 2025
Machine learning › Trustworthy machine learning › language model interpretability
attention head analysis
0.812024
Successor Heads: Recurring, Interpretable Attention Heads In The Wild · ICLR 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
sparse feature learning
0.812024
Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders · NeurIPS 2024
Security and privacy of machine learning
model stealing
0.812024
Stealing part of a production language model · ICML 2024
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
circuit analysis
0.712023
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small · ICLR 2023
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
circuit discovery
0.712023
Towards Automated Circuit Discovery for Mechanistic Interpretability · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

sparse autoencoder · 1.7black-box query attacks · 1.5feature disentanglement metrics · 0.9circuit finding · 0.9activation analysis · 0.9vector arithmetic · 0.8mechanistic interpretability · 0.8l1 penalty · 0.8gated sparse autoencoder · 0.8circuit analysis · 0.7
YearPublicationVenuePosition
2025 SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
abstract
Sparse autoencoders (SAEs) are a popular technique for interpreting language model activations, and there is extensive recent work on improving SAE effectiveness. However, most prior work evaluates progress using unsupervised proxy metrics with unclear practical relevance. We introduce SAEBench, a comprehensive evaluation suite that measures SAE performance across eight diverse metrics, spanning interpretability, feature disentanglement and practical applications like unlearning. To enable systematic comparison, we open-source a suite of over 200 SAEs across seven recently proposed SAE architectures and training algorithms. Our evaluation reveals that gains on proxy metrics do not reliably translate to better practical performance. For instance, while Matryoshka SAEs slightly underperform on existing proxy metrics, they substantially outperform other architectures on feature disentanglement metrics; moreover, this advantage grows with SAE scale. By providing a standardized framework for measuring progress in SAE development, SAEBench enables researchers to study scaling trends and make nuanced comparisons between different SAE architectures and training methodologies. Our interactive interface enables researchers to flexibly visualize relationships between metrics across hundreds of open-source SAEs at www.neuronpedia.org/sae-bench
Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Isaac Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Callum McDougall, Kola Ayonrinde, Demian Till, Matthew Wearden, Arthur Conmy, Samuel Marks, Neel Nanda
ICML13
2025 Scaling Sparse Feature Circuits For Studying In-Context Learning
abstract
Sparse autoencoders (SAEs) are a popular tool for interpreting large language model activations, but their utility in addressing open questions in interpretability remains unclear. In this work, we demonstrate their effectiveness by using SAEs to deepen our understanding of the mechanism behind in-context learning (ICL). We identify abstract SAE features that (i) encode the model’s knowledge of which task to execute and (ii) whose latent vectors causally induce the task zero-shot. This aligns with prior work showing that ICL is mediated by task vectors. We further demonstrate that these task vectors are well approximated by a sparse sum of SAE latents, including these task-execution features. To explore the ICL mechanism, we scale the sparse feature circuits methodology of Marks et al. (2024) to the Gemma 1 2B model for the more complex task of ICL. Through circuit finding, we discover task-detecting features with corresponding SAE latents that activate earlier in the prompt, that detect when tasks have been performed. They are causally linked with task-execution features through the attention and MLP sublayers.
Dmitrii Kharlapenko, Stepan Shabalin, Arthur Conmy, Neel Nanda
ICML3
2024 Successor Heads: Recurring, Interpretable Attention Heads In The Wild
abstract
In this work we describe successor heads: attention heads that increment tokens with a natural ordering, such as numbers, months, and days. For example, successor heads increment 'Monday' into 'Tuesday'. We explain the successor head behavior with an approach rooted in mechanistic interpretability, the field that aims to explain how models complete tasks in human-understandable terms. Existing research in this area has struggled to find recurring, mechanistically interpretable large language model (LLM) components beyond small toy models. Further, existing results have led to very little insight to explain the internals of the larger models that are used in practice. In this paper, we analyze the behavior of successor heads in LLMs and find that they implement abstract representations that are common to different architectures. Successor heads form in LLMs with as few as 31 million parameters, and at least as many as 12 billion parameters, such as GPT-2, Pythia, and Llama-2. We find a set of 'mod 10' features that underlie how successor heads increment in LLMs across different architectures and sizes. We perform vector arithmetic with these features to edit head behavior and provide insights into numeric representations within LLMs. Additionally, we study the behavior of successor heads on natural language data, where we find that successor heads are important for achieving a low loss on examples involving succession, and also identify interpretable polysemanticity in a Pythia successor head.
Rhys Gould, Euan Ong, George Ogden, Arthur Conmy
ICLR4
2024 Stealing part of a production language model
abstract
We introduce the first model-stealing attack that extracts precise, nontrivial information from black-box production language models like OpenAI's ChatGPT or Google's PaLM-2. Specifically, our attack recovers the embedding projection layer (up to symmetries) of a transformer model, given typical API access. For under $20 USD, our attack extracts the entire projection matrix of OpenAI's Ada and Babbage language models. We thereby confirm, for the first time, that these black-box models have a hidden dimension of 1024 and 2048, respectively. We also recover the exact hidden dimension size of the GPT-3.5-turbo model, and estimate it would cost under \\$2,000 in queries to recover the entire projection matrix. We conclude with potential defenses and mitigations, and discuss the implications of possible future work that could extend our attack.
Nicholas Carlini, Daniel Paleka, Krishnamurthy Dvijotham, Thomas Steinke 0002, Jonathan Hayase, A. Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Eric Wallace, David Rolnick, Florian Tramèr
ICML10
2024 Improving Sparse Decomposition of Language Model Activations with Gated Sparse Autoencoders
abstract
Recent work has found that sparse autoencoders (SAEs) are an effective technique for unsupervised discovery of interpretable features in language models' (LMs) activations, by finding sparse, linear reconstructions of those activations. We introduce the Gated Sparse Autoencoder (Gated SAE), which achieves a Pareto improvement over training with prevailing methods. In SAEs, the L1 penalty used to encourage sparsity introduces many undesirable biases, such as shrinkage -- systematic underestimation of feature activations. The key insight of Gated SAEs is to separate the functionality of (a) determining which directions to use and (b) estimating the magnitudes of those directions: this enables us to apply the L1 penalty only to the former, limiting the scope of undesirable side effects. Through training SAEs on LMs of up to 7B parameters we find that, in typical hyper-parameter ranges, Gated SAEs solve shrinkage, are similarly interpretable, and require half as many firing features to achieve comparable reconstruction fidelity.
Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, János Kramár, Rohin Shah, Neel Nanda
NeurIPS2
2023 Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, Jacob Steinhardt
ICLR3
2023 Towards Automated Circuit Discovery for Mechanistic Interpretability
abstract
Through considerable effort and intuition, several recent works have reverse-engineered nontrivial behaviors of transformer models. This paper systematizes the mechanistic interpretability process they followed. First, researchers choose a metric and dataset that elicit the desired model behavior. Then, they apply activation patching to find which abstract neural network units are involved in the behavior. By varying the dataset, metric, and units under investigation, researchers can understand the functionality of each component. We automate one of the process' steps: finding the connections between the abstract neural network units that form a circuit. We propose several algorithms and reproduce previous interpretability results to validate them. For example, the ACDC algorithm rediscovered 5/5 of the component types in a circuit in GPT-2 Small that computes the Greater-Than operation. ACDC selected 68 of the 32,000 edges in GPT-2 Small, all of which were manually found by previous work. Our code is available at https://github.com/ArthurConmy/Automatic-Circuit-Discovery
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, Adrià Garriga-Alonso
NeurIPS1
2022 Stylegan-Induced Data-Driven Regularization for Inverse Problems
abstract
Recent advances in generative adversarial networks (GANs) have opened up the possibility of generating high-resolution photo-realistic images that were impossible to produce previously. The ability of GANs to sample from high-dimensional distributions has naturally motivated researchers to leverage their power for modeling the image prior in inverse problems. We extend this line of research by developing a Bayesian image reconstruction framework that utilizes the full potential of a pre-trained StyleGAN2 generator, which is the currently dominant GAN architecture, for constructing the prior distribution on the underlying image. Our proposed approach, which we refer to as learned Bayesian reconstruction with generative models (L-BRGM), entails joint optimization over the style-code and the input latent code, and enhances the expressive power of a pre-trained StyleGAN2 generator by allowing the style-codes to be different for different generator layers. Considering the inverse problems of image inpainting and super-resolution, we demonstrate that the proposed approach is competitive with, and sometimes superior to, state-of-the-art GAN-based image reconstruction methods.
Arthur Conmy, Subhadip Mukherjee, Carola-Bibiane Schönlieb
ICASSP1