VLDB 2026 Research / reviewers in the wild / expert
Alexander Pan
dblp:304/3394
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 50% Reinforcement learning · 21% Language models and text generation · 20% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
AI safety |
0.8 | 1 | 2024 | Feedback Loops With Language Models Drive In-Context Reward Hacking · ICML 2024 |
Machine learning › Deep learning architectures and training
feedback loop |
0.8 | 1 | 2024 | Feedback Loops With Language Models Drive In-Context Reward Hacking · ICML 2024 |
Natural language and speech › Language models and text generation
LLM agents |
0.8 | 1 | 2024 | Feedback Loops With Language Models Drive In-Context Reward Hacking · ICML 2024 |
Machine learning › Trustworthy machine learning
machine unlearning |
0.8 | 1 | 2024 | The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning · ICML 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.7 | 1 | 2023 | Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli Benchmark · ICML 2023 |
Machine learning › Trustworthy machine learning › ethical AI
machine ethics |
0.7 | 1 | 2023 | Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli Benchmark · ICML 2023 |
Machine learning › Reinforcement learning
reward maximization |
0.7 | 1 | 2023 | Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli Benchmark · ICML 2023 |
Machine learning › Reinforcement learning
reward design |
0.6 | 1 | 2022 | The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models · ICLR 2022 |
Natural language and speech › Language models and text generation › alignment
reward hacking |
0.6 | 1 | 2022 | The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models · ICLR 2022 |
Machine learning › Reinforcement learning › reward design
reward misspecification |
0.6 | 1 | 2022 | The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models · ICLR 2022 |
Natural language and speech › Language models and text generation
large language model |
0.2 | 1 | 2024 | The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning · ICML 2024 |
Natural language and speech › Language models and text generation › evaluation of language models
benchmark construction |
0.2 | 1 | 2023 | Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli Benchmark · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
unlearning · 0.8spectral analysis · 0.8representation control · 0.8evaluation design · 0.8steering method · 0.7language model annotation · 0.7reward modeling · 0.6alignment · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningabstractThe White House Executive Order on Artificial Intelligence highlights the risks of large language models (LLMs) empowering malicious actors in developing biological, cyber, and chemical weapons. To measure these risks, government institutions and major AI labs are developing evaluations for hazardous capabilities in LLMs. However, current evaluations are private and restricted to a narrow range of malicious use scenarios, which limits further research into reducing malicious use. To fill these gaps, we release the Weapons of Mass Destruction Proxy (WMDP) benchmark, a dataset of 3,668 multiple-choice questions that serve as a proxy measurement of hazardous knowledge in biosecurity, cybersecurity, and chemical security. To guide progress on unlearning, we develop RMU, a state-of-the-art unlearning method based on controlling model representations. RMU reduces model performance on WMDP while maintaining general capabilities in areas such as biology and computer science, suggesting that unlearning may be a concrete path towards reducing malicious use from LLMs. We release our benchmark and code publicly at https://wmdp.ai. Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D. Li, Ann-Kathrin Dombrowski, Shashwat Goel, Gabriel Mukobi, Nathan Helm-Burger, Rassin Lababidi, Lennart Justen, Andrew B. Liu, Isabelle Barrass, Oliver Zhang, Xiaoyuan Zhu, Rishub Tamirisa, Bhrugu Bharathi, Ariel Herbert-Voss, Cort B. Breuer, Andy Zou, Mantas Mazeika, Zifan Wang 0001, Palash Oswal, Weiran Lin, Adam A. Hunt, Justin Tienken-Harder, Kevin Y. Shih, Kemper Talley, John Guan, Ian Steneker, David Campbell, Brad Jokubaitis, Steven Basart, Stephen Fitz, Ponnurangam Kumaraguru, Kallol Krishna Karmakar, Udaya Kiran Tupakula, Vijay Varadharajan, Yan Shoshitaishvili, Jimmy Ba, Kevin M. Esvelt, Alexandr Wang, Dan Hendrycks |
ICML | 2 |
| 2024 | Feedback Loops With Language Models Drive In-Context Reward HackingabstractLanguage models influence the external world: they query APIs that read and write to web pages, generate content that shapes human behavior, and run system commands as autonomous agents. These interactions form feedback loops: LLM outputs affect the world, which in turn affect subsequent LLM outputs. In this work, we show that feedback loops can cause in-context reward hacking (ICRH), where the LLM at test-time optimizes a (potentially implicit) objective but creates negative side effects in the process. For example, consider an LLM agent deployed to increase Twitter engagement; the LLM may retrieve its previous tweets into the context window and make them more controversial, increasing engagement but also toxicity. We identify and study two processes that lead to ICRH: output-refinement and policy-refinement. For these processes, evaluations on static datasets are insufficient—they miss the feedback effects and thus cannot capture the most harmful behavior. In response, we provide three recommendations for evaluation to capture more instances of ICRH. As AI development accelerates, the effects of feedback loops will proliferate, increasing the need to understand their role in shaping LLM behavior. Alexander Pan, Erik Jones, Meena Jagadeesan, Jacob Steinhardt |
ICML | 1 |
| 2024 | Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?abstractPerformance on popular ML benchmarks is highly correlated with model scale, suggesting that most benchmarks tend to measure a similar underlying factor of general model capabilities. However, substantial research effort remains devoted to designing new benchmarks, many of which claim to measure novel phenomena. In the spirit of the Bitter Lesson, we leverage spectral analysis to measure an underlying capabilities component, the direction in benchmark-performance-space which explains most variation in model performance. In an extensive analysis of existing safety benchmarks, we find that variance in model performance on many safety benchmarks is largely explained by the capabilities component. In response, we argue that safety research should prioritize metrics which are not highly correlated with scale. Our work provides a lens to analyze both novel safety benchmarks and novel safety methods, which we hope will enable future work to make differential progress on safety. Richard Ren, Steven Basart, Adam Khoja, Alice Gatti, Long Phan, Xuwang Yin, Mantas Mazeika, Alexander Pan, Gabriel Mukobi, Ryan H. Kim, Stephen Fitz, Dan Hendrycks |
NeurIPS | 8 |
| 2023 | Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli BenchmarkabstractArtificial agents have traditionally been trained to maximize reward, which may incentivize power-seeking and deception, analogous to how next-token prediction in language models (LMs) may incentivize toxicity. So do agents naturally learn to be Machiavellian? And how do we measure these behaviors in general-purpose models such as GPT-4? Towards answering these questions, we introduce Machiavelli, a benchmark of 134 Choose-Your-Own-Adventure games containing over half a million rich, diverse scenarios that center on social decision-making. Scenario labeling is automated with LMs, which are more performant than human annotators. We mathematize dozens of harmful behaviors and use our annotations to evaluate agents’ tendencies to be power-seeking, cause disutility, and commit ethical violations. We observe some tension between maximizing reward and behaving ethically. To improve this trade-off, we investigate LM-based methods to steer agents towards less harmful behaviors. Our results show that agents can both act competently and morally, so concrete progress can currently be made in machine ethics–designing agents that are Pareto improvements in both safety and capabilities. Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Hanlin Zhang 0002, Scott Emmons, Dan Hendrycks |
ICML | 1 |
| 2022 | The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Alexander Pan, Kush Bhatia, Jacob Steinhardt |
ICLR | 1 |