VLDB 2026 Research / reviewers in the wild / expert
Sam Earle
dblp:257/3172
· DBLP profile ↗
14ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0003-3783-9486ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language ModelsabstractWe are in the midst of large-scale industrial and academic efforts to automate the processes of scientific, technological and creative production through AI-driven assistants. Historically, a fundamental property of these processes in their human form has been their open-endedness: their capacity for generating a seemingly endless supply of novel and meaningful new forms. Do artificial agents have any capacity for such fruitful unguided discovery? To answer this question, we turn to Picbreeder, the canonical exemplar of human-driven open-ended search, in which users collaboratively generated a diverse library of images through interactive evolution of small neural networks. We replicate Picbreeder, replacing human users with frontier Vision Language Models (VLMs). We observe clear qualitative differences between the output of our system and the historical human baseline, and attempt to characterize them using metrics of phylogenetic complexity and visual and semantic salience and novelty. In an effort to identify some of the causal factors contributing these differences, we study the addition of exploratory noise to the agents' selection process, of behavioral diversity between agents, and of narrative momentum in the form of memory of past actions. We make our code available at https://github.com/smearle/picbreeder-vlm. Sam Earle, Kai Arulkumaran, Akarsh Kumar, Andrew Dai 0001, Julian Togelius, Sebastian Risi |
GECCO | 1 |
| 2025 | DreamGarden: A Designer Assistant for Growing Games from a Single Prompt
Sam Earle, Samyak Parajuli, Andrzej Banburski-Fahey |
CHI | 1 |
| 2025 | ScriptDoctor: Automatic Generation of PuzzleScript Games via Large Language Models and Tree SearchabstractThere is much interest in using large pre-trained models in Automatic Game Design (AGD), whether via the generation of code, assets, or more abstract conceptualization of design ideas. But so far, this interest largely stems from the ad hoc use of such generative models under persistent human supervision. Much work remains to show how these tools can be integrated into longer-time-horizon AGD pipelines, in which systems interface with game engines to test generated content autonomously. To this end, we introduce ScriptDoctor, a Large Language Model (LLM)-driven system for automatically generating and testing games in PuzzleScript, an expressive but highly constrained description language for turn-based puzzle games over 2D gridworlds. ScriptDoctor generates and tests game design ideas in an iterative loop, where human-authored examples are used to ground the system's output, compilation errors from the PuzzleScript engine are used to elicit functional code, and search-based agents play-test generated games. ScriptDoctor serves as a concrete example of the potential of automated, openended LLM-based workflows in generating novel game content. Sam Earle, Ahmed Khalifa 0001, Muhammad Umair Nasir, Zehua Jiang, Graham Todd, Andrzej Banburski-Fahey, Julian Togelius |
CoG | 1 |
| 2024 | Amorphous Fortress: Exploring Emergent Behavior and Complexity in Multi-Agent 0-Player GamesabstractWe introduce the Amorphous Fortress-an abstract, open-ended artificial life simulation framework. In this system, entities are represented as finite-state machines (FSMs) which allow for multi-agent interaction within a constrained space. These agents are created by randomly generating and evolving the FSMs; sampling from pre-defined states and transitions. This environment was designed to explore the emergent AI behaviors found implicitly in simulation games such as Dwarf Fortress or The Sims. We apply two evolutionary algorithms to this environment, hill-climber and MAP-Elites, to explore the various levels of depth and interaction from the generated FSMs and to generate diverse sets of environments that exhibit dynamics estimated to be complex by analyses of agents' FSM architecture and activation. This paper combines the work of two previous non-archival workshop papers. M Charity, Sam Earle, Dipika Rajesh, Mayu Wilson, Julian Togelius |
CEC | 2 |
| 2024 | Scaling, Control and Generalization in Reinforcement Learning Level GeneratorsabstractProcedural Content Generation via Reinforcement Learning (PCGRL) has been introduced as a means by which controllable designer agents can be trained based only on a set of computable metrics acting as a proxy for the level’s quality and key characteristics. While PCGRL offers a unique set of affordances for game designers, it is constrained by the compute-intensive process of training RL agents, and has so far been limited to generating relatively small levels. To address this issue of scale, we implement several PCGRL environments in Jax so that all aspects of learning and simulation happen in parallel on the GPU, resulting in faster environment simulation; removing the CPU-GPU transfer of information bottleneck during RL training; and ultimately resulting in significantly improved training speed. We replicate several key results from prior works in this new framework, letting models train for much longer than previously studied, and evaluating their behavior after 1 billion timesteps. Aiming for greater control for human designers, we introduce randomized level sizes and frozen “pinpoints” of pivotal game tiles as further ways of countering overfitting. To test the generalization ability of learned generators, we evaluate models on large, out-of-distribution map sizes, and find that partial observation sizes learn more robust design strategies. Sam Earle, Zehua Jiang, Julian Togelius |
CoG | 1 |
| 2024 | Missed Connections: Lateral Thinking Puzzles for Large Language ModelsabstractThe Connections puzzle published each day by the New York Times tasks players with dividing a bank of sixteen words into four groups of four words that each relate to a common theme. Solving the puzzle requires both common linguistic knowledge (i.e. definitions and typical usage) as well as, in many cases, lateral or abstract thinking. This is because the four categories ascend in complexity, with the most challenging category often requiring thinking about words in uncommon ways or as parts of larger phrases. We investigate the capacity for automated AI systems to play Connections and explore the game’s potential as an automated benchmark for abstract reasoning and a way to measure the semantic information encoded by data-driven linguistic systems. In particular, we study both a sentence-embedding baseline and modern large language models (LLMs). We report their accuracy on the task, measure the impacts of chain-of-thought prompting, and discuss their failure modes. Overall, we find that the Connections task is challenging yet feasible, and a strong test-bed for future work. Graham Todd, Timothy Merino, Sam Earle, Julian Togelius |
CoG | 3 |
| 2024 | DreamCraft: Text-Guided Generation of Functional 3D Environments in MinecraftabstractProcedural Content Generation (PCG) algorithms enable the automatic generation of complex and diverse artifacts. However, they don’t provide high-level control over the generated content and typically require domain expertise. In contrast, text-to-3D methods allow users to specify desired characteristics in natural language, offering a high amount of flexibility and expressivity. But unlike PCG, such approaches cannot guarantee functionality, which is crucial for certain applications like game design. In this paper, we present a method for generating functional 3D artifacts from free-form text prompts in the open-world game Minecraft. Our method, DreamCraft, trains quantized Neural Radiance Fields (NeRFs) to represent artifacts that, when viewed in-game, match given text descriptions. We find that DreamCraft produces more aligned in-game artifacts than a baseline that post-processes the output of an unconstrained NeRF. Thanks to the quantized representation of the environment, functional constraints can be integrated using specialized loss terms. We show how this can be leveraged to generate 3D structures that match a target distribution or obey certain adjacency rules over the block types. DreamCraft inherits a high degree of expressivity and controllability from the NeRF, while still being able to incorporate functional constraints through domain-specific objectives. Sam Earle, Filippos Kokkinos, Yuhe Nie, Julian Togelius, Roberta Raileanu |
FDG | 1 |
| 2024 | LLMatic: Neural Architecture Search Via Large Language Models And Quality Diversity OptimizationabstractLarge language models (LLMs) have emerged as powerful tools capable of accomplishing a broad spectrum of tasks. Their abilities span numerous areas, and one area where they have made a significant impact is in the domain of code generation. Here, we propose using the coding abilities of LLMs to introduce meaningful variations to code defining neural networks. Meanwhile, Quality-Diversity (QD) algorithms are known to discover diverse and robust solutions. By merging the code-generating abilities of LLMs with the diversity and robustness of QD solutions, we introduce LLMatic, a Neural Architecture Search (NAS) algorithm. While LLMs struggle to conduct NAS directly through prompts, LLMatic uses a procedural approach, leveraging QD for prompts and network architecture to create diverse and high-performing networks. We test LLMatic on the CIFAR-10 and NAS-bench-201 benchmarks, demonstrating that it can produce competitive networks while evaluating just 2, 000 candidates, even without prior knowledge of the benchmark domain or exposure to any previous top-performing models for the benchmark. The open-sourced code is available at https://github.com/umair-nasir14/LLMatic. Muhammad Umair Nasir, Sam Earle, Julian Togelius, Steven James 0001, Christopher W. Cleghorn |
GECCO | 2 |
| 2023 | Controllable Path of DestructionabstractPath of Destruction (PoD) is a self-supervised method for learning iterative generators. The core idea is to produce a training set by destroying a set of artifacts, and for each destructive step create a training instance based on the corresponding repair action. A generator trained on this dataset can then generate new artifacts by "repairing" from arbitrary states. The PoD method is very data-efficient in terms of original training examples and well-suited to functional artifacts composed of categorical data, such as game levels and discrete 3D structures. In this paper, we extend the Path of Destruction method to allow designer control over aspects of the generated artifacts. Controllability is introduced by adding conditional inputs to the state-action pairs that make up the repair trajectories. We test the controllable PoD method in a 2D dungeon setting, as well as in the domain of small 3D Lego cars. Matthew Siper, Sam Earle, Zehua Jiang, Ahmed Khalifa 0001, Julian Togelius |
CoG | 2 |
| 2023 | Level Generation Through Large Language ModelsabstractLarge Language Models (LLMs) are powerful tools, capable of leveraging their training on natural language to write stories, generate code, and answer questions. But can they generate functional video game levels? Game levels, with their complex functional constraints and spatial relationships in more than one dimension, are very different from the kinds of data an LLM typically sees during training. Datasets of game levels are also hard to come by, potentially taxing the abilities of these data-hungry models. We investigate the use of LLMs to generate levels for the game Sokoban, finding that LLMs are indeed capable of doing so, and that their performance scales dramatically with dataset size. We also perform preliminary experiments on controlling LLM level generators and discuss promising areas for future work. Graham Todd, Sam Earle, Muhammad Umair Nasir, Michael Cerny Green, Julian Togelius |
FDG | 2 |
| 2022 | Learning Controllable 3D Level GeneratorsabstractProcedural Content Generation via Reinforcement Learning (PCGRL) foregoes the need for large human-authored data-sets and allows agents to train explicitly on functional constraints, using computable, user-defined measures of quality instead of target output. We explore the application of PCGRL to 3D domains, in which content-generation tasks naturally have greater complexity and potential pertinence to real-world applications. Here, we introduce several PCGRL tasks for the 3D domain, Minecraft. These tasks will challenge RL-based generators using affordances often found in 3D environments, such as jumping, multiple dimensional movement, and gravity. We train agents to optimize each of these tasks to explore the capabilities of existing in PCGRL. The agents are able to generate relatively complex and diverse levels, and generalize to random initial states and control targets. Controllability tests in the presented tasks demonstrate their utility to analyze success and failure for 3D generators. We argue that these generators could serve both as co-creative tools for game designers, and as pre-trained environment generators in curriculum learning for player agents. Zehua Jiang, Sam Earle, Michael Cerny Green, Julian Togelius |
FDG | 2 |
| 2022 | Illuminating diverse neural cellular automata for level generationabstractWe present a method of generating diverse collections of neural cellular automata (NCA) to design video game levels. While NCAs have so far only been trained via supervised learning, we present a quality diversity (QD) approach to generating a collection of NCA level generators. By framing the problem as a QD problem, our approach can train diverse level generators, whose output levels vary based on aesthetic or functional criteria. To efficiently generate NCAs, we train generators via Covariance Matrix Adaptation MAP-Elites (CMA-ME), a quality diversity algorithm which specializes in continuous search spaces. We apply our new method to generate level generators for several 2D tile-based games: a maze game, Sokoban, and Zelda. Our results show that CMA-ME can generate small NCAs that are diverse yet capable, often satisfying complex solvability criteria for deterministic agents. We compare against a Compositional Pattern-Producing Network (CPPN) baseline trained to produce diverse collections of generators and show that the NCA representation yields a better exploration of level-space. Sam Earle, Justin Snider, Matthew C. Fontaine, Stefanos Nikolaidis, Julian Togelius |
GECCO | 1 |
| 2021 | Learning Controllable Content GeneratorsabstractIt has recently been shown that reinforcement learning can be used to train generators capable of producing high-quality game levels, with quality defined in terms of some user-specified heuristic. To ensure that these generators' output is sufficiently diverse (that is, not amounting to the reproduction of a single optimal level configuration), the generation process is constrained such that the initial seed results in some variance in the generator's output. However, this results in a loss of control over the generated content for the human user. We propose to train generators capable of producing controllably diverse output, by making them “goal-aware.” To this end, we add conditional inputs representing how close a generator is to some heuristic, and also modify the reward mechanism to incorporate that value. Testing on multiple domains, we show that the resulting level generators are capable of exploring the space of possible levels in a targeted, controllable manner, producing levels of comparable quality as their goal-unaware counterparts, that are diverse along designer-specified dimensions. Sam Earle, Maria Edwards, Ahmed Khalifa 0001, Philip Bontrager, Julian Togelius |
CoG | 1 |
| 2021 | Video Games as a Testbed for Open-Ended PhenomenaabstractUnderstanding and engineering open-endedness, or the indefinite generation of novelty and complexity at arbitrary scales, has long been studied by implementing nature-inspired simulations specifically designed for artificial life studies. This paper argues that video games serve as a complementary domain for research on open-endedness. In support of this claim, experiments in this paper evaluate the effects of age-based and spatial destructive events in two game domains: an interactive Game of Life and the city-building game SimCity. These games are played by a neural-network-controlled gameplay agent trying to maximize reward. Results indicate that experiments with SimCity are more likely to identify statistically significant differences in complexity as a result of applied destructive events, highlighting the utility of this game domain for studying artificial life phenomena. Sam Earle, Julian Togelius, Lisa B. Soros |
CoG | 1 |