VLDB 2026 Research / reviewers in the wild / expert
Wes Gurnee
dblp:288/2237
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Theory of computation · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 58% Language models and text generation · 21% Reinforcement learning · 6% | |
| Theoretical computer science
1 paper |
Algorithmic game theory and mechanism design · 100% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability |
2.5 | 3 | 2025 | Remarkable Robustness of LLMs: Stages of Inference? · NeurIPS 2025 Not All Language Model Features Are One-Dimensionally Linear · ICLR 2025 Confidence Regulation Neurons in Language Models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
interpretability |
2.4 | 3 | 2025 | Remarkable Robustness of LLMs: Stages of Inference? · NeurIPS 2025 Confidence Regulation Neurons in Language Models · NeurIPS 2024 Refusal in Language Models Is Mediated by a Single Direction · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model |
2.4 | 3 | 2025 | Remarkable Robustness of LLMs: Stages of Inference? · NeurIPS 2025 Confidence Regulation Neurons in Language Models · NeurIPS 2024 Language Models Represent Space and Time · ICLR 2024 |
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
sparse autoencoder |
0.9 | 1 | 2025 | Not All Language Model Features Are One-Dimensionally Linear · ICLR 2025 |
Machine learning › Generative modeling › autoregressive model
next-token prediction |
0.8 | 1 | 2024 | Confidence Regulation Neurons in Language Models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
0.8 | 1 | 2024 | Refusal in Language Models Is Mediated by a Single Direction · NeurIPS 2024 |
Computer vision › Video understanding and tracking › spatio-temporal modeling
spatiotemporal representation |
0.8 | 1 | 2024 | Language Models Represent Space and Time · ICLR 2024 |
Machine learning › Trustworthy machine learning
uncertainty modeling |
0.8 | 1 | 2024 | Confidence Regulation Neurons in Language Models · NeurIPS 2024 |
Machine learning › Reinforcement learning › model-based reinforcement learning
world model |
0.8 | 1 | 2024 | Language Models Represent Space and Time · ICLR 2024 |
Algorithmic game theory and mechanism design › social choice › computational social choice
multiwinner voting |
0.6 | 1 | 2022 | Combatting Gerrymandering with Social Choice: The Design of Multi-member Districts · EC 2022 |
Algorithmic game theory and mechanism design › social choice › voting
ranked choice voting |
0.6 | 1 | 2022 | Combatting Gerrymandering with Social Choice: The Design of Multi-member Districts · EC 2022 |
Algorithmic game theory and mechanism design
social choice |
0.6 | 1 | 2022 | Combatting Gerrymandering with Social Choice: The Design of Multi-member Districts · EC 2022 |
Machine learning › Deep learning architectures and training
transformer |
0.3 | 1 | 2025 | Remarkable Robustness of LLMs: Stages of Inference? · NeurIPS 2025 |
Security and privacy of machine learning › adversarial attack
jailbreak attack |
0.2 | 1 | 2024 | Refusal in Language Models Is Mediated by a Single Direction · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
rank-one weight edit · 1.5mechanistic analysis · 1.5activation steering · 1.5sparse autoencoder · 0.9layer swapping · 0.9layer deletion · 0.9intervention experiment · 0.9behavioral probing · 0.9representation analysis · 0.8probing · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Not All Language Model Features Are One-Dimensionally LinearabstractRecent work has proposed that language models perform computation by manipulating one-dimensional representations of concepts ("features") in activation space. In contrast, we explore whether some language model representations may be inherently multi-dimensional. We begin by developing a rigorous definition of irreducible multi-dimensional features based on whether they can be decomposed into either independent or non-co-occurring lower-dimensional features. Motivated by these definitions, we design a scalable method that uses sparse autoencoders to automatically find multi-dimensional features in GPT-2 and Mistral 7B. These auto-discovered features include strikingly interpretable examples, e.g. $\textit{circular}$ features representing days of the week and months of the year. We identify tasks where these exact circles are used to solve computational problems involving modular arithmetic in days of the week and months of the year. Next, we provide evidence that these circular features are indeed the fundamental unit of computation in these tasks with intervention experiments on Mistral 7B and Llama 3 8B, and we examine the continuity of the days of the week feature in Mistral 7B. Overall, our work argues that understanding multi-dimensional features is necessary to mechanistically decompose some model behaviors. Joshua Engels, Eric J. Michaud, Isaac Liao, Wes Gurnee, Max Tegmark |
ICLR | 4 |
| 2025 | Remarkable Robustness of LLMs: Stages of Inference?abstractWe investigate the robustness of Large Language Models (LLMs) to structural interventions by deleting and swapping adjacent layers during inference. Surprisingly, models retain 72–95\% of their original top-1 prediction accuracy without any fine-tuning. We find that performance degradation is not uniform across layers: interventions to the early and final layers cause the most degradation, while the model is remarkably robust to dropping middle layers. This pattern of localized sensitivity motivates our hypothesis of four stages of inference, observed across diverse model families and sizes: (1) detokenization, where local context is integrated to lift raw token embeddings into higher-level representations; (2) feature engineering, where task- and entity-specific features are iteratively refined; (3) prediction ensembling, where hidden states are aggregated into plausible next-token predictions; and (4) residual calibration, where irrelevant features are suppressed to finalize the output distribution. Synthesizing behavioral and mechanistic evidence, we provide a hypothesis for interpreting depth-dependent computations in LLMs. Vedang Lad, Jin Hwa Lee, Wes Gurnee, Max Tegmark |
NeurIPS | 3 |
| 2024 | Language Models Represent Space and TimeabstractThe capabilities of large language models (LLMs) have sparked debate over whether such systems just learn an enormous collection of superficial statistics or a set of more coherent and grounded representations that reflect the real world. We find evidence for the latter by analyzing the learned representations of three spatial datasets (world, US, NYC places) and three temporal datasets (historical figures, artworks, news headlines) in the Llama-2 family of models. We discover that LLMs learn linear representations of space and time across multiple scales. These representations are robust to prompting variations and unified across different entity types (e.g. cities and landmarks). In addition, we identify individual "space neurons" and "time neurons" that reliably encode spatial and temporal coordinates. While further investigation is needed, our results suggest modern LLMs learn rich spatiotemporal representations of the real world and possess basic ingredients of a world model. Wes Gurnee, Max Tegmark |
ICLR | 1 |
| 2024 | Refusal in Language Models Is Mediated by a Single DirectionabstractConversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this refusal behavior is widespread across chat models, its underlying mechanisms remain poorly understood. In this work, we show that refusal is mediated by a one-dimensional subspace, across 13 popular open-source chat models up to 72B parameters in size. Specifically, for each model, we find a single direction such that erasing this direction from the model's residual stream activations prevents it from refusing harmful instructions, while adding this direction elicits refusal on even harmless instructions. Leveraging this insight, we propose a novel white-box jailbreak method that surgically disables a model's ability to refuse, with minimal effect on other capabilities. This interpretable rank-one weight edit results in an effective jailbreak technique that is simpler and more efficient than fine-tuning. Finally, we mechanistically analyze how adversarial suffixes suppress propagation of the refusal-mediating direction. Our findings underscore the brittleness of current safety fine-tuning methods. More broadly, our work showcases how an understanding of model internals can be leveraged to develop practical methods for controlling model behavior. Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, Neel Nanda |
NeurIPS | 6 |
| 2024 | Confidence Regulation Neurons in Language ModelsabstractDespite their widespread use, the mechanisms by which large language models (LLMs) represent and regulate uncertainty in next-token predictions remain largely unexplored. This study investigates two critical components believed to influence this uncertainty: the recently discovered entropy neurons and a new set of components that we term token frequency neurons. Entropy neurons are characterized by an unusually high weight norm and influence the final layer normalization (LayerNorm) scale to effectively scale down the logits. Our work shows that entropy neurons operate by writing onto an \textit{unembedding null space}, allowing them to impact the residual stream norm with minimal direct effect on the logits themselves. We observe the presence of entropy neurons across a range of models, up to 7 billion parameters. On the other hand, token frequency neurons, which we discover and describe here for the first time, boost or suppress each token’s logit proportionally to its log frequency, thereby shifting the output distribution towards or away from the unigram distribution. Finally, we present a detailed case study where entropy neurons actively manage confidence: the setting of induction, i.e. detecting and continuing repeated subsequences. Alessandro Stolfo, Ben Wu 0001, Wes Gurnee, Yonatan Belinkov, Xingyi Song, Mrinmaya Sachan, Neel Nanda |
NeurIPS | 3 |
| 2022 | Combatting Gerrymandering with Social Choice: The Design of Multi-member DistrictsabstractThe Fair Representation Act, first introduced in 2017 and reintroduced in 2019 and 2021, would mandate the use of multi-member districts (MMDs) to elect members to the United States House of Representatives, i.e., having fewer, larger districts each with multiple representatives. The bill is supported by good governance organizations such as FairVote; the American Academy of Arts and Sciences in 2020 released a report advocating states to use multi-member districts - however, "on the condition that they adopt a non-winner-take-all election model." Despite the popular focus on single-member district (SMD) elections, such MMDs have a long history in the United States, especially at the state and local level. In 1962, 41 state legislatures had MMDs, often with winner-take-all models; even today, 10 state legislatures elect representatives for at least one chamber in such a manner. City councils, state parties, and other organizations often adopt more sophisticated techniques, using variations on Ranked Choice Voting (RCV) to elect multiple winners from each of several districts. Nikhil Garg 0001, Wes Gurnee, David Rothschild, David B. Shmoys |
EC | 2 |