VLDB 2026 Research / reviewers in the wild / expert
Nicholas Goldowsky-Dill
dblp:344/6035
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Trustworthy machine learning · 71% Information extraction and text analysis · 15% Representation and self-supervised learning · 13% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
1.6 | 2 | 2025 | Detecting Strategic Deception with Linear Probes · ICML 2025 Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning · NeurIPS 2024 |
Natural language and speech › Information extraction and text analysis › text classification
deception detection |
0.9 | 1 | 2025 | Detecting Strategic Deception with Linear Probes · ICML 2025 |
Machine learning › Trustworthy machine learning › interpretability › representation probing
linear probing |
0.9 | 1 | 2025 | Detecting Strategic Deception with Linear Probes · ICML 2025 |
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability |
0.8 | 1 | 2024 | Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder |
0.8 | 1 | 2024 | Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding › dictionary learning
sparse dictionary learning |
0.8 | 1 | 2024 | Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
linear probe · 0.9KL divergence minimization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Detecting Strategic Deception with Linear ProbesabstractAI models might use deceptive strategies as part of scheming or misaligned behaviour. Monitoring outputs alone is insufficient, since the AI might produce seemingly benign outputs while its internal reasoning is misaligned. We thus evaluate if linear probes can robustly detect deception by monitoring model activations. We test two probe-training datasets, one with contrasting instructions to be honest or deceptive (following Zou et al. (2023)) and one of responses to simple roleplaying scenarios. We test whether these probes generalize to realistic settings where Llama-3.3-70B-Instruct behaves deceptively, such as concealing insider trading Scheurer et al. (2023) and purposely underperforming on safety evaluations Benton et al. (2024). We find that our probe distinguishes honest and deceptive responses with AUROCs between 0.96 and 0.999 on our evaluation datasets. If we set the decision threshold to have a 1% false positive rate on chat data not related to deception, our probe catches 95-99% of the deceptive responses. Overall we think white-box probes are promising for future monitoring systems, but current performance is insufficient as a robust defence against deception. Our probes’ outputs can be viewed at https://data.apolloresearch.ai/dd/ and our code at https://github.com/ApolloResearch/deception-detection. Nicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim, Marius Hobbhahn |
ICML | 1 |
| 2024 | Identifying Functionally Important Features with End-to-End Sparse Dictionary LearningabstractIdentifying the features learned by neural networks is a core challenge in mechanistic interpretability. Sparse autoencoders (SAEs), which learn a sparse, overcomplete dictionary that reconstructs a network's internal activations, have been used to identify these features. However, SAEs may learn more about the structure of the datatset than the computational structure of the network. There is therefore only indirect reason to believe that the directions found in these dictionaries are functionally important to the network. We propose end-to-end (e2e) sparse dictionary learning, a method for training SAEs that ensures the features learned are functionally important by minimizing the KL divergence between the output distributions of the original model and the model with SAE activations inserted. Compared to standard SAEs, e2e SAEs offer a Pareto improvement: They explain more network performance, require fewer total features, and require fewer simultaneously active features per datapoint, all with no cost to interpretability. We explore geometric and qualitative differences between e2e SAE features and standard SAE features. E2e dictionary learning brings us closer to methods that can explain network behavior concisely and accurately. We release our library for training e2e SAEs and reproducing our analysis at
https://github.com/ApolloResearch/e2e_sae. Dan Braun, Jordan Taylor, Nicholas Goldowsky-Dill, Lee Sharkey |
NeurIPS | 3 |