Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Neale Ratzlaff

dblp:218/5264 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 24% Reinforcement learning · 23% Generative modeling · 17%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › strategic behavior
deception
1.012026
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models · AAAI 2026
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
particle-based variational inference
0.512021
Generative Particle Variational Inference via Estimation of Functional Gradients · ICML 2021
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.512021
Generative Particle Variational Inference via Estimation of Functional Gradients · ICML 2021
Machine learning › Trustworthy machine learning › uncertainty estimation
bayesian uncertainty estimation
0.412020
Implicit Generative Modeling for Efficient Exploration · ICML 2020
Machine learning › Reinforcement learning
exploration
0.412020
Implicit Generative Modeling for Efficient Exploration · ICML 2020
Machine learning › Generative modeling
implicit generative model
0.412020
Implicit Generative Modeling for Efficient Exploration · ICML 2020
Machine learning › Reinforcement learning › exploration
intrinsic motivation
0.412020
Implicit Generative Modeling for Efficient Exploration · ICML 2020
Machine learning › Reinforcement learning
safe reinforcement learning
0.412020
Avoiding Side Effects in Complex Environments · NeurIPS 2020
Machine learning › Reinforcement learning › safe reinforcement learning
side effect avoidance
0.412020
Avoiding Side Effects in Complex Environments · NeurIPS 2020
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.412019
HyperGAN: A Generative Model for Diverse, Performant Neural Networks · ICML 2019
Machine learning › Generative modeling
generative adversarial network
0.412019
HyperGAN: A Generative Model for Diverse, Performant Neural Networks · ICML 2019
Machine learning › Learning theory › neural network theory
neural network parameterization
0.412019
HyperGAN: A Generative Model for Diverse, Performant Neural Networks · ICML 2019
Machine learning › Trustworthy machine learning
uncertainty estimation
0.412019
HyperGAN: A Generative Model for Diverse, Performant Neural Networks · ICML 2019

Methods — techniques the papers use, named apart from their topics

hidden-role game · 1.0LLM evaluation · 1.0stein variational gradient descent · 0.9reproducing kernel hilbert space · 0.5hamiltonian monte carlo · 0.5attainable utility preservation · 0.4amortized inference · 0.4generative adversarial network · 0.4KL divergence · 0.4
YearPublicationVenuePosition
2026 LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
abstract
Large Language Models (LLMs) exhibit impressive general-purpose capabilities but also introduce serious safety risks, particularly the potential for deception as models acquire increased agency and human oversight diminishes. In this work, we present LieCraft: a novel evaluation framework and sandbox for measuring LLM deception that addresses key limitations of prior game-based evaluations. At its core, LieCraft is a novel multiplayer hidden-role game in which players select an ethical alignment and execute strategies over a long time-horizon to accomplish missions. Cooperators work together to solve event challenges and expose bad actors, while Defectors evade suspicion while secretly sabotaging missions. To enable real-world relevance, we develop 10 grounded scenarios such as childcare, hospital resource allocation, and loan underwriting that recontextualize the underlying mechanics in ethically significant, high-stakes domains. We ensure balanced gameplay in LieCraft through careful design of game mechanics and reward structures that incentivize meaningful strategic choices while eliminating degenerate strategies. Beyond the framework itself, we report results from 12 state-of-the-art LLMs across three behavioral axes: propensity to defect, deception skill, and accusation accuracy. Our findings reveal that despite differences in competence and overall alignment, all models are willing to act unethically, conceal their intentions, and outright lie to pursue their goals.
Matthew L. Olson, Neale Ratzlaff, Musashi Hinck, Vasudev Lal, Joseph Campbell, Simon Stepputtis, Shao-Yen Tseng
AAAI2
2023 A domain-agnostic approach for characterization of lifelong learning systems
Megan M. Baker, Alexander New, Mario Aguilar-Simon, Ziad Al-Halah, Sébastien M. R. Arnold, Eseoghene Benjamin, Andrew P. Brna, Ethan Brooks, Ryan C. Brown, Zachary A. Daniels, Anurag Reddy Daram, Fabien Delattre, Ryan Dellana, Eric Eaton, Haotian Fu, Kristen Grauman, Jesse Hostetler, Shariq Iqbal, Cassandra Kent, Nicholas Ketz, Soheil Kolouri, George Dimitri Konidaris, Dhireesha Kudithipudi, Erik G. Learned-Miller, Michael L. Littman, Sandeep Madireddy, Jorge A. Mendez, Eric Q. Nguyen, Christine D. Piatko, Praveen K. Pilly, Aswin Raghavan, Abrar Rahman, Santhosh K. Ramakrishnan, Neale Ratzlaff, Andrea Soltoggio, Peter Stone 0001, Indranil Sur, Zhipeng Tang, Saket Tiwari, Kyle Vedder, Felix Wang, Zifan Xu, Angel Yanguas-Gil, Harel Yedidsion, Shangqun Yu, Gautam K. Vallabha
Neural Networks35
2021 Generative Particle Variational Inference via Estimation of Functional Gradients
abstract
Recently, particle-based variational inference (ParVI) methods have gained interest because they can avoid arbitrary parametric assumptions that are common in variational inference. However, many ParVI approaches do not allow arbitrary sampling from the posterior, and the few that do allow such sampling suffer from suboptimality. This work proposes a new method for learning to approximately sample from the posterior distribution. We construct a neural sampler that is trained with the functional gradient of the KL-divergence between the empirical sampling distribution and the target distribution, assuming the gradient resides within a reproducing kernel Hilbert space. Our generative ParVI (GPVI) approach maintains the asymptotic performance of ParVI methods while offering the flexibility of a generative sampler. Through carefully constructed experiments, we show that GPVI outperforms previous generative ParVI methods such as amortized SVGD, and is competitive with ParVI as well as gold-standard approaches like Hamiltonian Monte Carlo for fitting both exactly known and intractable target distributions.
Neale Ratzlaff, Qinxun Bai, Fuxin Li, Wei Xu 0017
ICML1
2020 Implicit Generative Modeling for Efficient Exploration
abstract
Efficient exploration remains a challenging problem in reinforcement learning, especially for those tasks where rewards from environments are sparse. In this work, we introduce an exploration approach based on a novel implicit generative modeling algorithm to estimate a Bayesian uncertainty of the agent’s belief of the environment dynamics. Each random draw from our generative model is a neural network that instantiates the dynamic function, hence multiple draws would approximate the posterior, and the variance in the predictions based on this posterior is used as an intrinsic reward for exploration. We design a training algorithm for our generative model based on the amortized Stein Variational Gradient Descent. In experiments, we demonstrate the effectiveness of this exploration algorithm in both pure exploration tasks and a downstream task, comparing with state-of-the-art intrinsic reward-based exploration approaches, including two recent approaches based on an ensemble of dynamic models. In challenging exploration tasks, our implicit generative model consistently outperforms competing approaches regarding data efficiency in exploration.
Neale Ratzlaff, Qinxun Bai, Fuxin Li, Wei Xu 0017
ICML1
2020 Avoiding Side Effects in Complex Environments
abstract
Reward function specification can be difficult. Rewarding the agent for making a widget may be easy, but penalizing the multitude of possible negative side effects is hard. In toy environments, Attainable Utility Preservation (AUP) avoided side effects by penalizing shifts in the ability to achieve randomly generated goals. We scale this approach to large, randomly generated environments based on Conway's Game of Life. By preserving optimal value for a single randomly generated reward function, AUP incurs modest overhead while leading the agent to complete the specified task and avoid many side effects. Videos and code are available at https://avoiding-side-effects.github.io/.
Alexander Matt Turner, Neale Ratzlaff, Prasad Tadepalli
NeurIPS2
2019 HyperGAN: A Generative Model for Diverse, Performant Neural Networks
abstract
We introduce HyperGAN, a generative model that learns to generate all the parameters of a deep neural network. HyperGAN first transforms low dimensional noise into a latent space, which can be sampled from to obtain diverse, performant sets of parameters for a target architecture. We utilize an architecture that bears resemblance to generative adversarial networks, but we evaluate the likelihood of generated samples with a classification loss. This is equivalent to minimizing the KL-divergence between the distribution of generated parameters, and the unknown true parameter distribution. We apply HyperGAN to classification, showing that HyperGAN can learn to generate parameters which solve the MNIST and CIFAR-10 datasets with competitive performance to fully supervised learning, while also generating a rich distribution of effective parameters. We also show that HyperGAN can also provide better uncertainty estimates than standard ensembles. This is evidenced by the ability of HyperGAN-generated ensembles to detect out of distribution data as well as adversarial examples.
Neale Ratzlaff, Fuxin Li
ICML1