EDBT 2026 Demo / reviewers in the wild / expert
Alexander Matt Turner
dblp:236/6253
· DBLP profile ↗
6ranked-venue papers
4as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 52% Language models and text generation · 21% Reinforcement learning · 16% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
AI safety |
1.1 | 2 | 2022 | Parametrically Retargetable Decision-Makers Tend To Seek Power · NeurIPS 2022 Optimal Policies Tend To Seek Power · NeurIPS 2021 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.9 | 1 | 2025 | Distillation Robustifies Unlearning · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
machine unlearning |
0.9 | 1 | 2025 | Distillation Robustifies Unlearning · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › machine unlearning
robust unlearning |
0.9 | 1 | 2025 | Distillation Robustifies Unlearning · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › interpretability › representation engineering
activation editing |
0.8 | 1 | 2024 | Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024 |
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering |
0.8 | 1 | 2024 | Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024 |
Natural language and speech › Language models and text generation › model steering
language model steering |
0.8 | 1 | 2024 | Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024 |
Machine learning › Reinforcement learning
safe reinforcement learning |
0.4 | 1 | 2020 | Avoiding Side Effects in Complex Environments · NeurIPS 2020 |
Machine learning › Reinforcement learning › safe reinforcement learning
side effect avoidance |
0.4 | 1 | 2020 | Avoiding Side Effects in Complex Environments · NeurIPS 2020 |
Natural language and speech › Language models and text generation › alignment
LLM behavior control |
0.2 | 1 | 2024 | Steering Llama 2 via Contrastive Activation Addition · ACL (1) 2024 |
Methods — techniques the papers use, named apart from their topics
markov decision process · 1.1knowledge distillation · 0.9fine-tuning · 0.9contrastive activation addition · 0.8functional analysis · 0.6formal theory · 0.5attainable utility preservation · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Distillation Robustifies UnlearningabstractCurrent LLM unlearning methods are not robust. A few steps of finetuning can revert their effects. We begin by showing that this is true even for an idealized form of unlearning: training to imitate a model that was never trained on unwanted information. This shows that training a model can drastically modify its input-output behavior while leaving its underlying capabilities intact. In light of this dynamic, we show our main result. Training a randomly initialized student on the outputs of an unlearned model transfers behaviors while leaving latent capabilities behind. In short, distillation robustifies unlearning. Based on this result, we propose Unlearn-Noise-Distill-on-Outputs (UNDO), a scalable method that distills an unlearned model into a noised copy of itself. UNDO introduces a tunable tradeoff between compute cost and robustness, establishing a new Pareto frontier on synthetic language and arithmetic tasks. At its strongest setting, UNDO matches the robustness of a model retrained from scratch with perfect data filtering while using only 60-80% of the compute and requiring only 0.01% of the pretraining data to be labeled. We also show that UNDO robustifies unlearning on the more realistic Weapons of Mass Destruction Proxy (WMDP) benchmark. Since distillation is widely used in practice, incorporating an unlearning step beforehand offers a convenient path to robust capability removal. Bruce W. Lee, Addie Foote, Alex Infanger, Leni Shor, Harish Kamath, Jacob Goldman-Wetzler, Bryce Woodworth, Alex Cloud, Alexander Matt Turner |
NeurIPS | 9 |
| 2024 | Steering Llama 2 via Contrastive Activation AdditionabstractNina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, Alexander Turner. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, Alexander Matt Turner |
ACL (1) | 6 |
| 2022 | Parametrically Retargetable Decision-Makers Tend To Seek PowerabstractIf capable AI agents are generally incentivized to seek power in service of the objectives we specify for them, then these systems will pose enormous risks, in addition to enormous benefits. In fully observable environments, most reward functions have an optimal policy which seeks power by keeping options open and staying alive. However, the real world is neither fully observable, nor must trained agents be even approximately reward-optimal. We consider a range of models of AI decision-making, from optimal, to random, to choices informed by learning and interacting with an environment. We discover that many decision-making functions are retargetable, and that retargetability is sufficient to cause power-seeking tendencies. Our functional criterion is simple and broad. We show that a range of qualitatively dissimilar decision-making procedures incentivize agents to seek power. We demonstrate the flexibility of our results by reasoning about learned policy incentives in Montezuma's Revenge. These results suggest a safety risk: Eventually, retargetable training procedures may train real-world agents which seek power over humans. Alexander Matt Turner, Prasad Tadepalli |
NeurIPS | 1 |
| 2021 | Optimal Policies Tend To Seek PowerabstractSome researchers speculate that intelligent reinforcement learning (RL) agents would be incentivized to seek resources and power in pursuit of the objectives we specify for them. Other researchers point out that RL agents need not have human-like power-seeking instincts. To clarify this discussion, we develop the first formal theory of the statistical tendencies of optimal policies. In the context of Markov decision processes, we prove that certain environmental symmetries are sufficient for optimal policies to tend to seek power over the environment. These symmetries exist in many environments in which the agent can be shut down or destroyed. We prove that in these environments, most reward functions make it optimal to seek power by keeping a range of options available and, when maximizing average reward, by navigating towards larger sets of potential terminal states. Alexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch, Prasad Tadepalli |
NeurIPS | 1 |
| 2020 | Conservative Agency via Attainable Utility PreservationabstractReward functions are easy to misspecify; although designers can make corrections after observing mistakes, an agent pursuing a misspecified reward function can irreversibly change the state of its environment. If that change precludes optimization of the correctly specified reward function, then correction is futile. For example, a robotic factory assistant could break expensive equipment due to a reward misspecification; even if the designers immediately correct the reward function, the damage is done. To mitigate this risk, we introduce an approach that balances optimization of the primary reward function with preservation of the ability to optimize auxiliary reward functions. Surprisingly, even when the auxiliary reward functions are randomly generated and therefore uninformative about the correctly specified reward function, this approach induces conservative, effective behavior. Alexander Matt Turner, Dylan Hadfield-Menell, Prasad Tadepalli |
AIES | 1 |
| 2020 | Avoiding Side Effects in Complex EnvironmentsabstractReward function specification can be difficult. Rewarding the agent for making a widget may be easy, but penalizing the multitude of possible negative side effects is hard. In toy environments, Attainable Utility Preservation (AUP) avoided side effects by penalizing shifts in the ability to achieve randomly generated goals. We scale this approach to large, randomly generated environments based on Conway's Game of Life. By preserving optimal value for a single randomly generated reward function, AUP incurs modest overhead while leading the agent to complete the specified task and avoid many side effects. Videos and code are available at https://avoiding-side-effects.github.io/. Alexander Matt Turner, Neale Ratzlaff, Prasad Tadepalli |
NeurIPS | 1 |