Alexander Matt Turner

dblp:236/6253 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 52% Language models and text generation · 21% Reinforcement learning · 16%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
AI safety
1.122022
Parametrically Retargetable Decision-Makers Tend To Seek Power · NeurIPS 2022
Optimal Policies Tend To Seek Power · NeurIPS 2021
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.912025
Distillation Robustifies Unlearning · NeurIPS 2025
Machine learning › Trustworthy machine learning
machine unlearning
0.912025
Distillation Robustifies Unlearning · NeurIPS 2025
Machine learning › Trustworthy machine learning › machine unlearning
robust unlearning
0.912025
Distillation Robustifies Unlearning · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability › representation engineering
activation editing
0.812024
Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering
0.812024
Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024
Natural language and speech › Language models and text generation › model steering
language model steering
0.812024
Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024
Machine learning › Reinforcement learning
safe reinforcement learning
0.412020
Avoiding Side Effects in Complex Environments · NeurIPS 2020
Machine learning › Reinforcement learning › safe reinforcement learning
side effect avoidance
0.412020
Avoiding Side Effects in Complex Environments · NeurIPS 2020
Natural language and speech › Language models and text generation › alignment
LLM behavior control
0.212024
Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024

Methods — techniques the papers use, named apart from their topics

markov decision process · 1.1knowledge distillation · 0.9fine-tuning · 0.9contrastive activation addition · 0.8functional analysis · 0.6formal theory · 0.5attainable utility preservation · 0.4
YearPublicationVenuePosition
2025 Distillation Robustifies Unlearning
abstract
Current LLM unlearning methods are not robust. A few steps of finetuning can revert their effects. We begin by showing that this is true even for an idealized form of unlearning: training to imitate a model that was never trained on unwanted information. This shows that training a model can drastically modify its input-output behavior while leaving its underlying capabilities intact. In light of this dynamic, we show our main result. Training a randomly initialized student on the outputs of an unlearned model transfers behaviors while leaving latent capabilities behind. In short, distillation robustifies unlearning. Based on this result, we propose Unlearn-Noise-Distill-on-Outputs (UNDO), a scalable method that distills an unlearned model into a noised copy of itself. UNDO introduces a tunable tradeoff between compute cost and robustness, establishing a new Pareto frontier on synthetic language and arithmetic tasks. At its strongest setting, UNDO matches the robustness of a model retrained from scratch with perfect data filtering while using only 60-80% of the compute and requiring only 0.01% of the pretraining data to be labeled. We also show that UNDO robustifies unlearning on the more realistic Weapons of Mass Destruction Proxy (WMDP) benchmark. Since distillation is widely used in practice, incorporating an unlearning step beforehand offers a convenient path to robust capability removal.
Bruce W. Lee, Addie Foote, Alex Infanger, Leni Shor, Harish Kamath, Jacob Goldman-Wetzler, Bryce Woodworth, Alex Cloud, Alexander Matt Turner
NeurIPS9
2024 Steering Llama 2 via Contrastive Activation Addition
abstract
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, Alexander Turner. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, Alexander Matt Turner
ACL (1)6
2022 Parametrically Retargetable Decision-Makers Tend To Seek Power
abstract
If capable AI agents are generally incentivized to seek power in service of the objectives we specify for them, then these systems will pose enormous risks, in addition to enormous benefits. In fully observable environments, most reward functions have an optimal policy which seeks power by keeping options open and staying alive. However, the real world is neither fully observable, nor must trained agents be even approximately reward-optimal. We consider a range of models of AI decision-making, from optimal, to random, to choices informed by learning and interacting with an environment. We discover that many decision-making functions are retargetable, and that retargetability is sufficient to cause power-seeking tendencies. Our functional criterion is simple and broad. We show that a range of qualitatively dissimilar decision-making procedures incentivize agents to seek power. We demonstrate the flexibility of our results by reasoning about learned policy incentives in Montezuma's Revenge. These results suggest a safety risk: Eventually, retargetable training procedures may train real-world agents which seek power over humans.
Alexander Matt Turner, Prasad Tadepalli
NeurIPS1
2021 Optimal Policies Tend To Seek Power
abstract
Some researchers speculate that intelligent reinforcement learning (RL) agents would be incentivized to seek resources and power in pursuit of the objectives we specify for them. Other researchers point out that RL agents need not have human-like power-seeking instincts. To clarify this discussion, we develop the first formal theory of the statistical tendencies of optimal policies. In the context of Markov decision processes, we prove that certain environmental symmetries are sufficient for optimal policies to tend to seek power over the environment. These symmetries exist in many environments in which the agent can be shut down or destroyed. We prove that in these environments, most reward functions make it optimal to seek power by keeping a range of options available and, when maximizing average reward, by navigating towards larger sets of potential terminal states.
Alexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch, Prasad Tadepalli
NeurIPS1
2020 Conservative Agency via Attainable Utility Preservation
abstract
Reward functions are easy to misspecify; although designers can make corrections after observing mistakes, an agent pursuing a misspecified reward function can irreversibly change the state of its environment. If that change precludes optimization of the correctly specified reward function, then correction is futile. For example, a robotic factory assistant could break expensive equipment due to a reward misspecification; even if the designers immediately correct the reward function, the damage is done. To mitigate this risk, we introduce an approach that balances optimization of the primary reward function with preservation of the ability to optimize auxiliary reward functions. Surprisingly, even when the auxiliary reward functions are randomly generated and therefore uninformative about the correctly specified reward function, this approach induces conservative, effective behavior.
Alexander Matt Turner, Dylan Hadfield-Menell, Prasad Tadepalli
AIES1
2020 Avoiding Side Effects in Complex Environments
abstract
Reward function specification can be difficult. Rewarding the agent for making a widget may be easy, but penalizing the multitude of possible negative side effects is hard. In toy environments, Attainable Utility Preservation (AUP) avoided side effects by penalizing shifts in the ability to achieve randomly generated goals. We scale this approach to large, randomly generated environments based on Conway's Game of Life. By preserving optimal value for a single randomly generated reward function, AUP incurs modest overhead while leading the agent to complete the specified task and avoid many side effects. Videos and code are available at https://avoiding-side-effects.github.io/.
Alexander Matt Turner, Neale Ratzlaff, Prasad Tadepalli
NeurIPS1