VLDB 2026 Research / reviewers in the wild / expert
Adam Karvonen
dblp:372/4888
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Trustworthy machine learning · 73% Representation and self-supervised learning · 24% Language models and text generation · 3% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
2.5 | 3 | 2025 | SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability · ICML 2025 Learning Multi-Level Features with Matryoshka Sparse Autoencoders · ICML 2025 Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder |
2.5 | 3 | 2025 | SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability · ICML 2025 Learning Multi-Level Features with Matryoshka Sparse Autoencoders · ICML 2025 Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning |
1.0 | 2 | 2025 | Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models · NeurIPS 2024 Learning Multi-Level Features with Matryoshka Sparse Autoencoders · ICML 2025 |
Machine learning › Representation and self-supervised learning › hierarchical representation › hierarchical representation learning
hierarchical feature representation |
0.9 | 1 | 2025 | Learning Multi-Level Features with Matryoshka Sparse Autoencoders · ICML 2025 |
Machine learning › Trustworthy machine learning
language model interpretability |
0.9 | 1 | 2025 | SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
sparse autoencoder · 2.5nested dictionaries · 0.9feature disentanglement metrics · 0.9p-annealing · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Multi-Level Features with Matryoshka Sparse AutoencodersabstractSparse autoencoders (SAEs) have emerged as a powerful tool for interpreting neural networks by extracting the concepts represented in their activations. However, choosing the size of the SAE dictionary (i.e. number of learned concepts) creates a tension: as dictionary size increases to capture more relevant concepts, sparsity incentivizes features to be split or absorbed into more specific features, leaving high-level features missing or warped. We introduce Matryoshka SAEs, a novel variant that addresses these issues by simultaneously training multiple nested dictionaries of increasing size, forcing the smaller dictionaries to independently reconstruct the inputs without using the larger dictionaries. This organizes features hierarchically - the smaller dictionaries learn general concepts, while the larger dictionaries learn more specific concepts, without incentive to absorb the high-level features. We train Matryoshka SAEs on Gemma-2-2B and TinyStories and find superior performance on sparse probing and targeted concept erasure tasks, more disentangled concept representations, and reduced feature absorption. While there is a minor tradeoff with reconstruction performance, we believe Matryoshka SAEs are a superior alternative for practical tasks, as they enable training arbitrarily large SAEs while retaining interpretable features at different levels of abstraction. Bart Bussmann, Noa Nabeshima, Adam Karvonen, Neel Nanda |
ICML | 3 |
| 2025 | SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model InterpretabilityabstractSparse autoencoders (SAEs) are a popular technique for interpreting language model activations, and there is extensive recent work on improving SAE effectiveness. However, most prior work evaluates progress using unsupervised proxy metrics with unclear practical relevance. We introduce SAEBench, a comprehensive evaluation suite that measures SAE performance across eight diverse metrics, spanning interpretability, feature disentanglement and practical applications like unlearning. To enable systematic comparison, we open-source a suite of over 200 SAEs across seven recently proposed SAE architectures and training algorithms. Our evaluation reveals that gains on proxy metrics do not reliably translate to better practical performance. For instance, while Matryoshka SAEs slightly underperform on existing proxy metrics, they substantially outperform other architectures on feature disentanglement metrics; moreover, this advantage grows with SAE scale. By providing a standardized framework for measuring progress in SAE development, SAEBench enables researchers to study scaling trends and make nuanced comparisons between different SAE architectures and training methodologies. Our interactive interface enables researchers to flexibly visualize relationships between metrics across hundreds of open-source SAEs at www.neuronpedia.org/sae-bench Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Isaac Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Callum McDougall, Kola Ayonrinde, Demian Till, Matthew Wearden, Arthur Conmy, Samuel Marks, Neel Nanda |
ICML | 1 |
| 2024 | Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game ModelsabstractWhat latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representations has shown significant promise. However, evaluating the quality of these SAEs is difficult because we lack a ground-truth collection of interpretable features which we expect good SAEs to identify. We thus propose to measure progress in interpretable dictionary learning by working in the setting of LMs trained on Chess and Othello transcripts. These settings carry natural collections of interpretable features—for example, “there is a knight on F3”—which we leverage into metrics for SAE quality. To guide progress in interpretable dictionary learning, we introduce a new SAE training technique, $p$-annealing, which demonstrates improved performance on our metric. Adam Karvonen, Benjamin Wright, Can Rager, Rico Angell, Jannik Brinkmann, Logan Smith, Claudio Mayrink Verdun, David Bau, Samuel Marks |
NeurIPS | 1 |