Adam Karvonen

dblp:372/4888 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 73% Representation and self-supervised learning · 24% Language models and text generation · 3%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
2.532025
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability · ICML 2025
Learning Multi-Level Features with Matryoshka Sparse Autoencoders · ICML 2025
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder
2.532025
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability · ICML 2025
Learning Multi-Level Features with Matryoshka Sparse Autoencoders · ICML 2025
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning
1.022025
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models · NeurIPS 2024
Learning Multi-Level Features with Matryoshka Sparse Autoencoders · ICML 2025
Machine learning › Representation and self-supervised learning › hierarchical representation › hierarchical representation learning
hierarchical feature representation
0.912025
Learning Multi-Level Features with Matryoshka Sparse Autoencoders · ICML 2025
Machine learning › Trustworthy machine learning
language model interpretability
0.912025
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability · ICML 2025

Methods — techniques the papers use, named apart from their topics

sparse autoencoder · 2.5nested dictionaries · 0.9feature disentanglement metrics · 0.9p-annealing · 0.8
YearPublicationVenuePosition
2025 Learning Multi-Level Features with Matryoshka Sparse Autoencoders
abstract
Sparse autoencoders (SAEs) have emerged as a powerful tool for interpreting neural networks by extracting the concepts represented in their activations. However, choosing the size of the SAE dictionary (i.e. number of learned concepts) creates a tension: as dictionary size increases to capture more relevant concepts, sparsity incentivizes features to be split or absorbed into more specific features, leaving high-level features missing or warped. We introduce Matryoshka SAEs, a novel variant that addresses these issues by simultaneously training multiple nested dictionaries of increasing size, forcing the smaller dictionaries to independently reconstruct the inputs without using the larger dictionaries. This organizes features hierarchically - the smaller dictionaries learn general concepts, while the larger dictionaries learn more specific concepts, without incentive to absorb the high-level features. We train Matryoshka SAEs on Gemma-2-2B and TinyStories and find superior performance on sparse probing and targeted concept erasure tasks, more disentangled concept representations, and reduced feature absorption. While there is a minor tradeoff with reconstruction performance, we believe Matryoshka SAEs are a superior alternative for practical tasks, as they enable training arbitrarily large SAEs while retaining interpretable features at different levels of abstraction.
Bart Bussmann, Noa Nabeshima, Adam Karvonen, Neel Nanda
ICML3
2025 SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
abstract
Sparse autoencoders (SAEs) are a popular technique for interpreting language model activations, and there is extensive recent work on improving SAE effectiveness. However, most prior work evaluates progress using unsupervised proxy metrics with unclear practical relevance. We introduce SAEBench, a comprehensive evaluation suite that measures SAE performance across eight diverse metrics, spanning interpretability, feature disentanglement and practical applications like unlearning. To enable systematic comparison, we open-source a suite of over 200 SAEs across seven recently proposed SAE architectures and training algorithms. Our evaluation reveals that gains on proxy metrics do not reliably translate to better practical performance. For instance, while Matryoshka SAEs slightly underperform on existing proxy metrics, they substantially outperform other architectures on feature disentanglement metrics; moreover, this advantage grows with SAE scale. By providing a standardized framework for measuring progress in SAE development, SAEBench enables researchers to study scaling trends and make nuanced comparisons between different SAE architectures and training methodologies. Our interactive interface enables researchers to flexibly visualize relationships between metrics across hundreds of open-source SAEs at www.neuronpedia.org/sae-bench
Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Isaac Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Callum McDougall, Kola Ayonrinde, Demian Till, Matthew Wearden, Arthur Conmy, Samuel Marks, Neel Nanda
ICML1
2024 Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
abstract
What latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representations has shown significant promise. However, evaluating the quality of these SAEs is difficult because we lack a ground-truth collection of interpretable features which we expect good SAEs to identify. We thus propose to measure progress in interpretable dictionary learning by working in the setting of LMs trained on Chess and Othello transcripts. These settings carry natural collections of interpretable features—for example, “there is a knight on F3”—which we leverage into metrics for SAE quality. To guide progress in interpretable dictionary learning, we introduce a new SAE training technique, $p$-annealing, which demonstrates improved performance on our metric.
Adam Karvonen, Benjamin Wright, Can Rager, Rico Angell, Jannik Brinkmann, Logan Smith, Claudio Mayrink Verdun, David Bau, Samuel Marks
NeurIPS1