Brennan Shacklett

dblp:197/7158 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-2894-1158ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Robot navigation and mapping · 27% Autonomous driving · 23% Reinforcement learning · 21%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
GPUs and heterogeneous computing · 64% Parallel and multicore computing · 30% Cloud and datacenter computing · 6%
Computer graphics and multimedia
4 papers
Rendering · 88% Computer animation and physical simulation · 12%

Topics — the 17 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing › GPU-accelerated scientific computing
GPU-accelerated simulation
1.522025
GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPS · ICLR 2025
An Extensible, Data-Oriented Architecture for High-Performance, Many-World Simulation · ACM Trans. Graph. 2023
Rendering
ray tracing
0.922024
High-Throughput Batch Rendering for Embodied AI · SIGGRAPH Asia 2024
R2E2: low-latency path tracing of terabyte-scale scenes using thousands of cloud CPUs · ACM Trans. Graph. 2022
Robotics › Autonomous driving › simulation
closed-loop simulation
0.912025
GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPS · ICLR 2025
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.912025
GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPS · ICLR 2025
Computer vision › 3D vision
3d scene understanding
0.812024
Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation · CVPR 2024
Robotics › Robot navigation and mapping
embodied navigation
0.812024
Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation · CVPR 2024
Robotics › Robot navigation and mapping
object goal navigation
0.812024
Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation · CVPR 2024
Rendering › ray tracing
path tracing
0.612022
R2E2: low-latency path tracing of terabyte-scale scenes using thousands of cloud CPUs · ACM Trans. Graph. 2022
Machine learning › Reinforcement learning
deep reinforcement learning
0.512021
Large Batch Simulation for Deep Reinforcement Learning · ICLR 2021
Machine learning › Efficient and distributed learning
distributed training
0.512021
Large Batch Simulation for Deep Reinforcement Learning · ICLR 2021
Robotics › Robot navigation and mapping
embodied AI simulation
0.512021
Megaverse: Simulating Embodied Agents at One Million Experiences per Second · ICML 2021
Machine learning › Transfer learning and domain adaptation › domain generalization
synthetic-to-real generalization
0.212024
Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation · CVPR 2024
GPUs and heterogeneous computing
GPU rendering
0.212024
High-Throughput Batch Rendering for Embodied AI · SIGGRAPH Asia 2024
Computer animation and physical simulation
rigid body simulation
0.212023
An Extensible, Data-Oriented Architecture for High-Performance, Many-World Simulation · ACM Trans. Graph. 2023
Cloud and datacenter computing
serverless computing
0.212022
R2E2: low-latency path tracing of terabyte-scale scenes using thousands of cloud CPUs · ACM Trans. Graph. 2022
Parallel and multicore computing
parallel programming models
0.112017
Encoding, Fast and Slow: Low-Latency Video Processing Using Thousands of Tiny Threads · NSDI 2017
Parallel and multicore computing › parallel programming models
task parallelism
0.112017
Encoding, Fast and Slow: Low-Latency Video Processing Using Thousands of Tiny Threads · NSDI 2017

Methods — techniques the papers use, named apart from their topics

rasterization · 2.3GPU ray tracing · 2.3ray tracing · 2.0entity-component-system · 2.0GPU acceleration · 2.0reinforcement learning · 1.7CUDA · 1.7serverless cloud computing · 1.1load balancing · 1.1BVH · 1.1embodied agent training · 0.8dataset construction · 0.8service-oriented design · 0.6batch reinforcement learning · 0.5
YearPublicationVenuePosition
2025 GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPS
abstract
Multi-agent learning algorithms have been successful at generating superhuman planning in various games but have had limited impact on the design of deployed multi-agent planners. A key bottleneck in applying these techniques to multi-agent planning is that they require billions of steps of experience. To enable the study of multi-agent planning at scale, we present GPUDrive, a GPU-accelerated, multi-agent simulator built on top of the Madrona Game Engine capable of generating over a million simulation steps per second. Observation, reward, and dynamics functions are written directly in C++, allowing users to define complex, heterogeneous agent behaviors that are lowered to high-performance CUDA. Despite these low-level optimizations, GPUDrive is fully accessible through Python, offering a seamless and efficient workflow for multi-agent, closed-loop simulation. Using GPUDrive, we train reinforcement learning agents on the Waymo Open Motion Dataset, achieving efficient goal-reaching in minutes and scaling to thousands of scenarios in hours. We open-source the code and pre-trained agents at \url{www.github.com/Emerge-Lab/gpudrive}.
Saman Kazemkhani, Aarav Pandya, Daphne Cornelisse, Brennan Shacklett, Eugene Vinitsky
ICLR4
2024 Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation
abstract
We contribute the Habitat Synthetic Scenes Dataset (HSSD-200), a dataset of 211 high-quality 3D scenes, and use it to test navigation agent generalization to realistic 3D environments. Our dataset represents real interiors and contains a diverse set of 18,656 models of real-world objects. We investigate the impact of synthetic 3D scene dataset scale and realism on the task of training embodied agents to find and navigate to objects (ObjectGoal navigation). By comparing to synthetic 3D scene datasets from prior work, we find that scale helps in generalization, but the benefits quickly saturate, making visual fidelity and correlation to real-world scenes more important. Our experiments show that agents trained on our smaller-scale dataset can outperform agents trained on much larger datasets. Surprisingly, we observe that agents trained on just 122 scenes from our dataset outperform agents trained on 10,000 scenes from the ProcTHOR-10K dataset in terms of zero-shot generalization in real-world scanned environments.
Mukul Khanna, Yongsen Mao, Hanxiao Jiang 0001, Sanjay Haresh, Brennan Shacklett, Dhruv Batra, Alexander Clegg, Eric Undersander, Angel X. Chang, Manolis Savva
CVPR5
2024 High-Throughput Batch Rendering for Embodied AI
abstract
Fig. 1.A batch of renders from the HSSD dataset, a popular embodied AI dataset with complex apartment-scale scenes (7.4 million triangles per scene on average).When rendering a batch of 1024 views at 128×128 pixels each, our ray tracer generates frames at an aggregate throughput of 25K frames/sec on a H100 GPU.We perform a systematic study of batch rendering performance under a range of embodied AI workloads and find that ray tracing based solutions, rather than rasterization based solutions, are a preferred rendering solution when executing on widely used datacenter-class GPUs.In this paper we study the problem of efficiently rendering images for embodied AI training workloads, where agent training involves rendering millions to billions of independent, low-resolution frames, often with simple lighting and shading, that serve as the agent's observations of the world.To enable high-throughput training from images, we design a flexible, batchmode rendering interface that allows state-of-the-art GPU-accelerated batch world simulators to efficiently communicate with high-performance rendering backends.Using this interface we architect and compare two highperformance renderers: one based on the GPU hardware-accelerated graphics pipeline and a second based on a GPU software implementation of ray tracing.To evaluate these renderers and encourage further research by the graphics community in this area, we build a rendering benchmark for this under-explored regime.We find that the ray tracing renderer outperforms the rasterization-based solution across the benchmark on a datacenter-class GPU, while also performing competitively in geometrically complex environments on a high-end consumer GPU.When tasked to render large batches of independent 128×128 images, the ray tracer can exceed 100,000 frames per second per GPU for simple scenes, and exceed 10,000 frames per second per GPU on geometrically complex scenes from the HSSD dataset.
Luc Guy Rosenzweig, Brennan Shacklett, Warren Xia, Kayvon Fatahalian
SIGGRAPH Asia2
2024 Learning to Move Like Professional Counter-Strike Players
abstract
Abstract In multiplayer, first‐person shooter games like Counter‐Strike: Global Offensive (CS:GO), coordinated movement is a critical component of high‐level strategic play. However, the complexity of team coordination and the variety of conditions present in popular game maps make it impractical to author hand‐crafted movement policies for every scenario. We show that it is possible to take a data‐driven approach to creating human‐like movement controllers for CS:GO. We curate a team movement dataset comprising 123 hours of professional game play traces, and use this dataset to train a transformer‐based movement model that generates human‐like team movement for all players in a “Retakes” round of the game. Importantly, the movement prediction model is efficient. Performing inference for all players takes less than 0.5 ms per game step (amortized cost) on a single CPU core, making it plausible for use in commercial games today. Human evaluators assess that our model behaves more like humans than both commercially‐available bots and procedural movement controllers scripted by experts (16% to 59% higher by TrueSkill rating of “human‐like”). Using experiments involving in‐game bot vs. bot self‐play, we demonstrate that our model performs simple forms of teamwork, makes fewer common movement mistakes, and yields movement distributions, player lifetimes, and kill locations similar to those observed in professional CS:GO match play.
David Durst, Feng Xie 0008, Vishnu Sarukkai, Brennan Shacklett, Iuri Frosio, Chen Tessler, Joohwan Kim, Carly Taylor, Gilbert Louis Bernstein, Sanjiban Choudhury, Pat Hanrahan, Kayvon Fatahalian
Comput. Graph. Forum4
2023 An Extensible, Data-Oriented Architecture for High-Performance, Many-World Simulation
abstract
Training AI agents to perform complex tasks in simulated worlds requires millions to billions of steps of experience. To achieve high performance, today's fastest simulators for training AI agents adopt the idea of batch simulation: using a single simulation engine to simultaneously step many environments in parallel. We introduce a framework for productively authoring novel training environments (including custom logic for environment generation, environment time stepping, and generating agent observations and rewards) that execute as high-performance, GPU-accelerated batched simulators. Our key observation is that the entity-component-system (ECS) design pattern, popular for expressing CPU-side game logic today, is also well-suited for providing the structure needed for high-performance batched simulators. We contribute the first fully-GPU accelerated ECS implementation that natively supports batch environment simulation. We demonstrate how ECS abstractions impose structure on a training environment's logic and state that allows the system to efficiently manage state, amortize work, and identify GPU-friendly coherent parallel computations within and across different environments. We implement several learning environments in this framework, and demonstrate GPU speedups of two to three orders of magnitude over open source CPU baselines and 5-33× over strong baselines running on a 32-thread CPU. An implementation of the OpenAI hide and seek 3D environment written in our framework, which performs rigid body physics and ray tracing in each simulator step, achieves over 1.9 million environment steps per second on a single GPU.
Brennan Shacklett, Luc Guy Rosenzweig, Bidipta Sarkar, Andrew Szot, Erik Wijmans, Vladlen Koltun, Dhruv Batra, Kayvon Fatahalian
ACM Trans. Graph.1
2022 R2E2: low-latency path tracing of terabyte-scale scenes using thousands of cloud CPUs
abstract
In this paper we explore the viability of path tracing massive scenes using a "supercomputer" constructed on-the-fly from thousands of small, serverless cloud computing nodes. We present R2E2 (Really Elastic Ray Engine) a scene decomposition-based parallel renderer that rapidly acquires thousands of cloud CPU cores, loads scene geometry from a pre-built scene BVH into the aggregate memory of these nodes in parallel, and performs full path traced global illumination using an inter-node messaging service designed for communicating ray data. To balance ray tracing work across many nodes, R2E2 adopts a service-oriented design that statically replicates geometry and texture data from frequently traversed scene regions onto multiple nodes based on estimates of load, and dynamically assigns ray tracing work to lightly loaded nodes holding the required data. We port pbrt's ray-scene intersection components to the R2E2 architecture, and demonstrate that scenes with up to a terabyte of geometry and texture data (where as little as 1/250th of the scene can fit on any one node) can be path traced at 4K resolution, in tens of seconds using thousands of tiny serverless nodes on the AWS Lambda platform.
Sadjad Fouladi, Brennan Shacklett, Fait Poms, Arjun Arora, Alex Ozdemir, Deepti Raghavan, Pat Hanrahan, Kayvon Fatahalian, Keith Winstein
ACM Trans. Graph.2
2021 Large Batch Simulation for Deep Reinforcement Learning
Brennan Shacklett, Erik Wijmans, Aleksei Petrenko, Manolis Savva, Dhruv Batra, Vladlen Koltun, Kayvon Fatahalian
ICLR1
2021 Megaverse: Simulating Embodied Agents at One Million Experiences per Second
abstract
We present Megaverse, a new 3D simulation platform for reinforcement learning and embodied AI research. The efficient design of our engine enables physics-based simulation with high-dimensional egocentric observations at more than 1,000,000 actions per second on a single 8-GPU node. Megaverse is up to 70x faster than DeepMind Lab in fully-shaded 3D scenes with interactive objects. We achieve this high simulation performance by leveraging batched simulation, thereby taking full advantage of the massive parallelism of modern GPUs. We use Megaverse to build a new benchmark that consists of several single-agent and multi-agent tasks covering a variety of cognitive challenges. We evaluate model-free RL on this benchmark to provide baselines and facilitate future research.
Aleksei Petrenko, Erik Wijmans, Brennan Shacklett, Vladlen Koltun
ICML3
2017 Encoding, Fast and Slow: Low-Latency Video Processing Using Thousands of Tiny Threads
Sadjad Fouladi, Riad S. Wahby, Brennan Shacklett, Karthikeyan Balasubramaniam, William Zeng, Rahul Bhalerao, Anirudh Sivaraman, George Porter, Keith Winstein
NSDI3