VLDB 2026 Research / reviewers in the wild / expert
Glen Berseth
dblp:147/5478
· DBLP profile ↗
47ranked-venue papers
11as first author
29since 2021 · last 2025
0000-0001-7351-8028ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 8 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 2 since 2021Systems, architecture and hardware · 8 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Non-Adversarial Inverse Reinforcement Learning via Successor Feature MatchingabstractIn inverse reinforcement learning (IRL), an agent seeks to replicate expert demonstrations through interactions with the environment.
Traditionally, IRL is treated as an adversarial game, where an adversary searches over reward models, and a learner optimizes the reward through repeated RL procedures.
This game-solving approach is both computationally expensive and difficult to stabilize.
In this work, we propose a novel approach to IRL by _direct policy search_:
by exploiting a linear factorization of the return as the inner product of successor features and a reward vector, we design an IRL algorithm by policy gradient descent on the gap between the learner and expert features.
Our non-adversarial method does not require learning an explicit reward function and can be solved seamlessly with existing RL algorithms.
Remarkably, our approach works in state-only settings without expert action labels, a setting which behavior cloning (BC) cannot solve.
Empirical results demonstrate that our method learns from as few as a single expert demonstration and achieves improved performance on various control tasks. Arnav Kumar Jain, Harley Wiltzer, Jesse Farebrother, Irina Rish, Glen Berseth, Sanjiban Choudhury |
ICLR | 5 |
| 2025 | Towards Improving Exploration through Sibling Augmented GFlowNetsabstractExploration is a key factor for the success of an active learning agent, especially when dealing with sparse extrinsic terminal rewards and long trajectories. We introduce Sibling Augmented Generative Flow Networks (SA-GFN), a novel framework designed to enhance exploration and training efficiency of Generative Flow Networks (GFlowNets). SA-GFN uses a decoupled dual network architecture, comprising of a main Behavior Network and an exploratory Sibling Network, to enable a diverse exploration of the underlying distribution using intrinsic rewards. Inspired by the ideas on exploration from reinforcement learning, SA-GFN provides a general-purpose exploration and learning paradigm that integrates with multiple GFlowNet training objectives and is especially helpful for exploration over a wide range of sparse or low reward distributions and task structures. An extensive set of experiments across a diverse range of tasks, reward structures and trajectory lengths, along with a thorough set of ablations, demonstrate the superior performance of SA-GFN in terms of exploration efficacy and convergence speed as compared to the existing methods. In addition, SA-GFN's versatility and compatibility with different GFlowNet training objectives and intrinsic reward methods underscores its broad applicability in various problem domains. Kanika Madan, Alex Lamb, Emmanuel Bengio, Glen Berseth, Yoshua Bengio |
ICLR | 4 |
| 2025 | Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous InferenceabstractRealtime environments change even as agents perform action inference and learning, thus requiring high interaction frequencies to effectively minimize regret. However, recent advances in machine learning involve larger neural networks with longer inference times, raising questions about their applicability in realtime systems where reaction time is crucial. We present an analysis of lower bounds on regret in realtime reinforcement learning (RL) environments to show that minimizing long-term regret is generally impossible within the typical sequential interaction and learning paradigm, but often becomes possible when sufficient asynchronous compute is available. We propose novel algorithms for staggering asynchronous inference processes to ensure that actions are taken at consistent time intervals, and demonstrate that use of models with high action inference times is only constrained by the environment's effective stochasticity over the inference horizon, and not by action frequency. Our analysis shows that the number of inference processes needed scales linearly with increasing inference times while enabling use of models that are multiple orders of magnitude larger than existing approaches when learning from a realtime simulation of Game Boy games such as Pokemon and Tetris. Matthew Riemer, Gopeshh Subbaraj, Glen Berseth, Irina Rish |
ICLR | 3 |
| 2025 | Mitigating Plasticity Loss in Continual Reinforcement Learning by Reducing ChurnabstractPlasticity, or the ability of an agent to adapt to new tasks, environments, or distributions, is crucial for continual learning. In this paper, we study the loss of plasticity in deep continual RL from the lens of churn: network output variability induced by the data in each training batch. We demonstrate that (1) the loss of plasticity is accompanied by the exacerbation of churn due to the gradual rank decrease of the Neural Tangent Kernel (NTK) matrix; (2) reducing churn helps prevent rank collapse and adjusts the step size of regular RL gradients adaptively. Moreover, we introduce Continual Churn Approximated Reduction (C-CHAIN) and demonstrate it improves learning performance and outperforms baselines in a diverse range of continual learning environments on OpenAI Gym Control, ProcGen, DeepMind Control Suite, and MinAtar benchmarks. Hongyao Tang, Johan S. Obando-Ceron, Pablo Samuel Castro, Aaron C. Courville, Glen Berseth |
ICML | 5 |
| 2025 | Outsourced Diffusion Sampling: Efficient Posterior Inference in Latent Spaces of Generative ModelsabstractAny well-behaved generative model over a variable $\mathbf{x}$ can be expressed as a deterministic transformation of an exogenous (outsourced’) Gaussian noise variable $\mathbf{z}$: $\mathbf{x}=f_\theta(\mathbf{z})$. In such a model (eg, a VAE, GAN, or continuous-time flow-based model), sampling of the target variable $\mathbf{x} \sim p_\theta(\mathbf{x})$ is straightforward, but sampling from a posterior distribution of the form $p(\mathbf{x}\mid\mathbf{y}) \propto p_\theta(\mathbf{x})r(\mathbf{x},\mathbf{y})$, where $r$ is a constraint function depending on an auxiliary variable $\mathbf{y}$, is generally intractable. We propose to amortize the cost of sampling from such posterior distributions with diffusion models that sample a distribution in the noise space ($\mathbf{z}$). These diffusion samplers are trained by reinforcement learning algorithms to enforce that the transformed samples $f_\theta(\mathbf{z})$ are distributed according to the posterior in the data space ($\mathbf{x}$). For many models and constraints, the posterior in noise space is smoother than in data space, making it more suitable for amortized inference. Our method enables conditional sampling under unconditional GAN, (H)VAE, and flow-based priors, comparing favorably with other inference methods. We demonstrate the proposed outsourced diffusion sampling in several experiments with large pretrained prior models: conditional image generation, reinforcement learning with human feedback, and protein structure generation. Siddarth Venkatraman, Mohsin Hasan, Minsu Kim 0004, Luca Scimeca, Marcin Sendera, Yoshua Bengio, Glen Berseth, Nikolay Malkin |
ICML | 7 |
| 2025 | Stable Gradients for Stable Learning at Scale in Deep Reinforcement LearningabstractScaling deep reinforcement learning networks is challenging and often results in degraded performance, yet the root causes of this failure mode remain poorly understood. Several recent works have proposed mechanisms to address this, but they are often complex and fail to highlight the causes underlying this difficulty. In this work, we conduct a series of empirical analyses which suggest that the combination of non-stationarity with gradient pathologies, due to suboptimal architectural choices, underlie the challenges of scale. We propose a series of direct interventions that stabilize gradient flow, enabling robust performance across a range of network depths and widths. Our interventions are simple to implement and compatible with well-established algorithms, and result in an effective mechanism that enables strong performance even at large scales. We validate our findings on a variety of agents and suites of environments. Roger Creus Castanyer, Johan S. Obando-Ceron, Pierre-Luc Bacon, Glen Berseth, Aaron C. Courville, Pablo Samuel Castro |
NeurIPS | 5 |
| 2024 | Improving Intrinsic Exploration by Creating Stationary ObjectivesabstractExploration bonuses in reinforcement learning guide long-horizon exploration by defining custom intrinsic objectives. Count-based methods use the frequency of state visits to derive an exploration bonus. In this paper, we identify that any intrinsic reward function derived from count-based methods is non-stationary and hence induces a difficult objective to optimize for the agent. The key contribution of our work lies in transforming the original non-stationary rewards into stationary rewards through an augmented state representation. For this purpose, we introduce the Stationary Objectives For Exploration (SOFE) framework. SOFE requires *identifying* sufficient statistics for different exploration bonuses and finding an *efficient* encoding of these statistics to use as input to a deep network. SOFE is based on proposing state augmentations that expand the state space but hold the promise of simplifying the optimization of the agent's objective. Our experiments show that SOFE improves the agents' performance in challenging exploration problems, including sparse-reward tasks, pixel-based observations, 3D navigation, and procedurally generated environments. Roger Creus Castanyer, Joshua Romoff, Glen Berseth |
ICLR | 3 |
| 2024 | Closing the Gap between TD Learning and Supervised Learning - A Generalisation Point of ViewabstractSome reinforcement learning (RL) algorithms have the capability of recombining together pieces of previously seen experience to solve a task never seen before during training. This oft-sought property is one of the few ways in which dynamic programming based RL algorithms are considered different from supervised learning (SL) based RL algorithms. Yet, recent RL methods based on off-the-shelf SL algorithms achieve excellent results without an explicit mechanism for stitching; it remains unclear whether those methods forgo this important stitching property. This paper studies this question in the setting of goal-reaching problems. We show that the desirable stitching property corresponds to a form of generalization: after training on a distribution of (state, goal) pairs, one would like to evaluate on (state, goal) pairs not seen together in the training data. Our analysis shows that this sort of generalization is different from i.i.d. generalization. This connection between stitching and generalization reveals why we should not expect existing RL methods based on SL to perform stitching, even in the limit of large datasets and models. We experimentally validate this result on carefully constructed datasets.
This connection suggests a simple remedy, the same remedy for improving generalization in supervised learning: data augmentation. We propose a naive temporal data augmentation approach and demonstrate that adding it to RL methods based on SL enables them to successfully stitch together experience, so that they succeed in navigating between states and goals unseen together during training. Raj Ghugare, Matthieu Geist, Glen Berseth, Benjamin Eysenbach |
ICLR | 3 |
| 2024 | Searching for High-Value Molecules Using Reinforcement Learning and TransformersabstractReinforcement learning (RL) over text representations can be effective for finding high-value policies that can search over graphs. However, RL requires careful structuring of the search space and algorithm design to be effective in this challenge. Through extensive experiments, we explore how different design choices for text grammar and algorithmic choices for training can affect an RL policy's ability to generate molecules with desired properties. We arrive at a new RL-based molecular design algorithm (ChemRLformer) and perform a thorough analysis using 25 molecule design tasks, including computationally complex protein docking simulations. From this analysis, we discover unique insights in this problem space and show that ChemRLformer achieves state-of-the-art performance while being more straightforward than prior work by demystifying which design choices are actually helpful for text-based molecule design. Raj Ghugare, Santiago Miret, Adriana Hugessen, Mariano Phielipp, Glen Berseth |
ICLR | 5 |
| 2024 | Intelligent Switching for Reset-Free RLabstractIn the real world, the strong episode resetting mechanisms that are needed to train
agents in simulation are unavailable. The resetting assumption limits the potential
of reinforcement learning in the real world, as providing resets to an agent usually
requires the creation of additional handcrafted mechanisms or human interventions.
Recent work aims to train agents (forward) with learned resets by constructing
a second (backward) agent that returns the forward agent to the initial state. We
find that the termination and timing of the transitions between these two agents
are crucial for algorithm success. With this in mind, we create a new algorithm,
Reset Free RL with Intelligently Switching Controller (RISC) which intelligently
switches between the two agents based on the agent’s confidence in achieving its
current goal. Our new method achieves state-of-the-art performance on several
challenging environments for reset-free RL. Darshan Patil, Janarthanan Rajendran, Glen Berseth, Sarath Chandar |
ICLR | 3 |
| 2024 | Reasoning with Latent Diffusion in Offline Reinforcement LearningabstractOffline reinforcement learning (RL) holds promise as a means to learn high-reward policies from a static dataset, without the need for further environment interactions. However, a key challenge in offline RL lies in effectively stitching portions of suboptimal trajectories from the static dataset while avoiding extrapolation errors arising due to a lack of support in the dataset. Existing approaches use conservative methods that are tricky to tune and struggle with multi-modal data or rely on noisy Monte Carlo return-to-go samples for reward conditioning. In this work, we propose a novel approach that leverages the expressiveness of latent diffusion to model in-support trajectory sequences as compressed latent skills. This facilitates learning a Q-function while avoiding extrapolation error via batch-constraining. The latent space is also expressive and gracefully copes with multi-modal data. We show that the learned temporally-abstract latent space encodes richer task-specific information for offline RL tasks as compared to raw state-actions. This improves credit assignment and facilitates faster reward propagation during Q-learning. Our method demonstrates state-of-the-art performance on the D4RL benchmarks, particularly excelling in long-horizon, sparse-reward tasks. Siddarth Venkatraman, Shivesh Khaitan, Ravi Tej Akella, John M. Dolan, Jeff G. Schneider, Glen Berseth |
ICLR | 6 |
| 2024 | Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment CollaborationabstractLarge, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train "generalist" X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. The project website is robotics-transformer-x.github.io. Abigail O'Neill, Abhiram Maddukuri, Abhishek Gupta 0004, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew E. Wang, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess-Limerick, Bernhard Schölkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu 0003, Charlotte Le, Chelsea Finn, Chen Wang 0053, Chenfeng Xu, Cheng Chi 0001, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Drieß, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Paul Foster, Fangchen Liu, Federico Ceola, Fei Xia 0002, Feiyu Zhao, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guanzhi Wang, Hao Su 0001, Haoshu Fang, Henghui Bao, Heni Ben Amor, Henrik I. Christensen, Hiroki Furuta, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim 0001, Jaimyn Drake, Jan Peters 0001, Jan Schneider 0007, Jasmine Hsu, Jeannette Bohg, Jeffrey T. Bingham, Jensen Gao, Jiaheng Hu, Jiajun Wu 0001, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan 0001, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jitendra Malik, João Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Kenneth Y. Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Zhang 0002, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Yunliang Chen 0001, Lerrel Pinto, Li Fei-Fei 0001, Liam Tan, Linxi Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang 0003, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma 0001, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen 0001, Nicolas Heess, Nikhil J. Joshi, Niko Sünderhauf, Norman Di Palo, Nur Muhammad Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R. Sanketi, Patrick Tree Miller, Patrick Yin, Paul Wohlhart, Peng Xu 0010, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Rafael Rafailov, Ria Doshi, Roberto Martin Martin, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante-Gomez, Sean Kirmani, Sergey Levine, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham D. Sonawani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair 0003, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vincent Vanhoucke, Wolfram Burgard, Xiaolong Wang 0004, Xinghao Zhu, Xinyang Geng, Liangwei Xu, Yecheng Jason Ma 0001, Yejin Kim 0003, Yevgen Chebotar, Yilin Wu 0003, Yonatan Bisk, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang 0001, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zichen Jeff Cui, Zichen Zhang 0016, Zipeng Lin |
ICRA | 71 |
| 2024 | Simplifying Constraint Inference with Inverse Reinforcement LearningabstractLearning safe policies has presented a longstanding challenge for the reinforcement learning (RL) community. Various formulations of safe RL have been proposed; However, fundamentally, tabula rasa RL must learn safety constraints through experience, which is problematic for real-world applications. Imitation learning is often preferred in real-world settings because the experts' safety preferences are embedded in the data the agent imitates. However, imitation learning is limited in its extensibility to new tasks, which can only be learned by providing the agent with expert trajectories. For safety-critical applications with sub-optimal or inexact expert data, it would be preferable to learn only the safety aspects of the policy through imitation, while still allowing for task learning with RL. The field of inverse constrained RL, which seeks to infer constraints from expert data, is a promising step in this direction. However, prior work in this area has relied on complex tri-level optimizations in order to infer safe behavior (constraints). This challenging optimization landscape leads to sub-optimal performance on several benchmark tasks. In this work, we present a simplified version of constraint inference that performs as well or better than prior work across a collection of continuous-control benchmarks. Moreover, besides improving performance, this simplified framework is easier to implement, tune, and more readily lends itself to various extensions, such as offline constraint inference. Adriana Hugessen, Harley Wiltzer, Glen Berseth |
NeurIPS | 3 |
| 2024 | Improving Deep Reinforcement Learning by Reducing the Chain Effect of Value and Policy ChurnabstractDeep neural networks provide Reinforcement Learning (RL) powerful function approximators to address large-scale decision-making problems. However, these approximators introduce challenges due to the non-stationary nature of RL training. One source of the challenges in RL is that output predictions can churn, leading to uncontrolled changes after each batch update for states not included in the batch. Although such a churn phenomenon exists in each step of network training, it remains under-explored on how churn occurs and impacts RL. In this work, we start by characterizing churn in a view of Generalized Policy Iteration with function approximation, and we discover a chain effect of churn that leads to a cycle where the churns in value estimation and policy improvement compound and bias the learning dynamics throughout the iteration. Further, we concretize the study and focus on the learning issues caused by the chain effect in different settings, including greedy action deviation in value-based methods, trust region violation in proximal policy optimization, and dual bias of policy value in actor-critic methods. We then propose a method to reduce the chain effect across different settings, called Churn Approximated ReductIoN (CHAIN), which can be easily plugged into most existing DRL algorithms. Our experiments demonstrate the effectiveness of our method in both reducing churn and improving learning performance across online and offline, value-based and policy-based RL settings. Hongyao Tang, Glen Berseth |
NeurIPS | 2 |
| 2024 | Amortizing intractable inference in diffusion models for vision, language, and controlabstractDiffusion models have emerged as effective distribution estimators in vision, language, and reinforcement learning, but their use as priors in downstream tasks poses an intractable posterior inference problem. This paper studies *amortized* sampling of the posterior over data, $\mathbf{x}\sim p^{\rm post}(\mathbf{x})\propto p(\mathbf{x})r(\mathbf{x})$, in a model that consists of a diffusion generative model prior $p(\mathbf{x})$ and a black-box constraint or likelihood function $r(\mathbf{x})$. We state and prove the asymptotic correctness of a data-free learning objective, *relative trajectory balance*, for training a diffusion model that samples from this posterior, a problem that existing methods solve only approximately or in restricted cases. Relative trajectory balance arises from the generative flow network perspective on diffusion models, which allows the use of deep reinforcement learning techniques to improve mode coverage. Experiments illustrate the broad potential of unbiased inference of arbitrary posteriors under diffusion priors: in vision (classifier guidance), language (infilling under a discrete diffusion LLM), and multimodal data (text-to-image generation). Beyond generative modeling, we apply relative trajectory balance to the problem of continuous control with a score-based behavior prior, achieving state-of-the-art results on benchmarks in offline reinforcement learning. Code is available at [this link](https://github.com/GFNOrg/diffusion-finetuning). Siddarth Venkatraman, Moksh Jain, Luca Scimeca, Minsu Kim 0004, Marcin Sendera, Mohsin Hasan, Luke Rowe, Sarthak Mittal, Pablo Lemos, Emmanuel Bengio, Alexandre Adam, Jarrid Rector-Brooks, Yoshua Bengio, Glen Berseth, Nikolay Malkin |
NeurIPS | 14 |
| 2023 | Bootstrapping Adaptive Human-Machine Interfaces with Offline Reinforcement LearningabstractAdaptive interfaces can help users perform sequential decision-making tasks like robotic teleoperation given noisy, high-dimensional command signals (e.g., from a brain-computer interface). Recent advances in human-in-the-loop machine learning enable such systems to improve by interacting with users, but tend to be limited by the amount of data that they can collect from individual users in practice. In this paper, we propose a reinforcement learning algorithm to address this by training an interface to map raw command signals to actions using a combination of offline pre-training and online fine-tuning. To address the challenges posed by noisy command signals and sparse rewards, we develop a novel method for representing and inferring the user's long-term intent for a given trajectory. We primarily evaluate our method's ability to assist users who can only communicate through noisy, high-dimensional input channels through a user study in which 12 participants performed a simulated navigation task by using their eye gaze to modulate a 128-dimensional command signal from their webcam. The results show that our method enables successful goal navigation more often than a baseline directional interface, by learning to denoise user commands signals and provide shared autonomy assistance. We further evaluate on a simulated Sawyer pushing task with eye gaze control, and the Lunar Lander game with simulated user commands, and find that our method improves over baseline interfaces in these domains as well. Extensive ablation experiments with simulated user commands empirically motivate each component of our method. Jensen Gao, Siddharth Reddy, Glen Berseth, Anca D. Dragan, Sergey Levine |
IROS | 3 |
| 2023 | Maximum State Entropy Exploration using Predecessor and Successor RepresentationsabstractAnimals have a developed ability to explore that aids them in important tasks such as locating food, exploring for shelter, and finding misplaced items. These exploration skills necessarily track where they have been so that they can plan for finding items with relative efficiency. Contemporary exploration algorithms often learn a less efficient exploration strategy because they either condition only on the current state or simply rely on making random open-loop exploratory moves. In this work, we propose $\eta\psi$-Learning, a method to learn efficient exploratory policies by conditioning on past episodic experience to make the next exploratory move. Specifically, $\eta\psi$-Learning learns an exploration policy that maximizes the entropy of the state visitation distribution of a single trajectory. Furthermore, we demonstrate how variants of the predecessor representation and successor representations can be combined to predict the state visitation entropy. Our experiments demonstrate the efficacy of $\eta\psi$-Learning to strategically explore the environment and maximize the state coverage with limited samples. Arnav Kumar Jain, Lucas Lehnert, Irina Rish, Glen Berseth |
NeurIPS | 4 |
| 2023 | Towards Learning to Imitate from a Single Video DemonstrationabstractAgents that can learn to imitate behaviours observed in video -- without having direct access to internal state or action information of the observed agent -- are more suitable for learning in the natural world. However, formulating a reinforcement learning (RL) agent that facilitates this goal remains a significant challenge. We approach this challenge using contrastive training to learn a reward function by comparing an agent's behaviour with a single demonstration. We use a Siamese recurrent neural network architecture to learn rewards in space and time between motion clips while training an RL policy to minimize this distance. Through experimentation, we also find that the inclusion of multi-task data and additional image encoding losses improve the temporal consistency of the learned rewards and, as a result, significantly improve policy learning. We demonstrate our approach on simulated humanoid, dog, and raptor agents in 2D and quadruped and humanoid agents in 3D. We show that our method outperforms current state-of-the-art techniques and can learn to imitate behaviours from a single video demonstration. Glen Berseth, Florian Golemo, Christopher Joseph Pal |
J. Mach. Learn. Res. | 1 |
| 2023 | Heterogeneous Crowd Simulation Using Parametric Reinforcement LearningabstractAgent-based synthetic crowd simulation affords the cost-effective large-scale simulation and animation of interacting digital humans. Model-based approaches have successfully generated a plethora of simulators with a variety of foundations. However, prior approaches have been based on statically defined models predicated on simplifying assumptions, limited video-based datasets, or homogeneous policies. Recent works have applied reinforcement learning to learn policies for navigation. However, these approaches may learn static homogeneous rules, are typically limited in their generalization to trained scenarios, and limited in their usability in synthetic crowd domains. In this article, we present a multi-agent reinforcement learning-based approach that learns a parametric predictive collision avoidance and steering policy. We show that training over a parameter space produces a flexible model across crowd configurations. That is, our goal-conditioned approach learns a parametric policy that affords heterogeneous synthetic crowds. We propose a model-free approach without centralization of internal agent information, control signals, or agent communication. The model is extensively evaluated. The results show policy generalization across unseen scenarios, agent parameters, and out-of-distribution parameterizations. The learned model has comparable computational performance to traditional methods. Qualitatively the model produces both expected (laminar flow, shuffling, bottleneck) and unexpected (side-stepping) emergent qualitative behaviours, and quantitatively the approach is performant across measures of movement quality. Kaidong Hu, M. Brandon Haworth, Glen Berseth, Vladimir Pavlovic 0001, Petros Faloutsos, Mubbasir Kapadia |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | CoMPS: Continual Meta Policy Search
Glen Berseth, Grace Zhang, Chelsea Finn, Sergey Levine |
ICLR | 1 |
| 2022 | AnyMorph: Learning Transferable Polices By Inferring Agent MorphologyabstractThe prototypical approach to reinforcement learning involves training policies tailored to a particular agent from scratch for every new morphology. Recent work aims to eliminate the re-training of policies by investigating whether a morphology-agnostic policy, trained on a diverse set of agents with similar task objectives, can be transferred to new agents with unseen morphologies without re-training. This is a challenging problem that required previous approaches to use hand-designed descriptions of the new agent’s morphology. Instead of hand-designing this description, we propose a data-driven method that learns a representation of morphology directly from the reinforcement learning objective. Ours is the first reinforcement learning algorithm that can train a policy to generalize to new agent morphologies without requiring a description of the agent’s morphology in advance. We evaluate our approach on the standard benchmark for agent-agnostic control, and improve over the current state of the art in zero-shot generalization to new agents. Importantly, our method attains good performance without an explicit description of morphology. Brandon Trabucco, Mariano Phielipp, Glen Berseth |
ICML | 3 |
| 2022 | ASHA: Assistive Teleoperation via Human-in-the-Loop Reinforcement LearningabstractBuilding assistive interfaces for controlling robots through arbitrary, high-dimensional, noisy inputs (e.g., webcam images of eye gaze) can be challenging, especially when it involves inferring the user's desired action in the absence of a natural ‘default’ interface. Reinforcement learning from online user feedback on the system's performance presents a natural solution to this problem, and enables the interface to adapt to individual users. However, this approach tends to require a large amount of human-in-the-loop training data, especially when feedback is sparse. We propose a hierarchical solution that learns efficiently from sparse user feedback: we use offline pre-training to acquire a latent embedding space of useful, high-level robot behaviors, which, in turn, enables the system to focus on using online user feedback to learn a mapping from user inputs to desired high-level behaviors. The key insight is that access to a pre-trained policy enables the system to learn more from sparse rewards than a naïve RL algorithm: using the pre-trained policy, the system can make use of successful task executions to relabel, in hindsight, what the user actually meant to do during unsuccessful executions. We evaluate our method primarily through a user study with 12 participants who perform tasks in three simulated robotic manipulation domains using a webcam and their eye gaze: flipping light switches, opening a shelf door to reach objects inside, and rotating a valve. The results show that our method successfully learns to map 128-dimensional gaze features to 7-dimensional joint torques from sparse rewards in under 10 minutes of online training, and seamlessly helps users who employ different gaze strategies, while adapting to distributional shift in webcam inputs, tasks, and environments Sean Chen, Jensen Gao, Siddharth Reddy, Glen Berseth, Anca D. Dragan, Sergey Levine |
ICRA | 4 |
| 2022 | Hierarchical Reinforcement Learning for Precise Soccer Shooting Skills using a Quadrupedal RobotabstractWe address the problem of enabling quadrupedal robots to perform precise shooting skills in the real world using reinforcement learning. Developing algorithms to enable a legged robot to shoot a soccer ball to a given target is a challenging problem that combines robot motion control and planning into one task. To solve this problem, we need to consider the dynamics limitation and motion stability during the control of a dynamic legged robot. Moreover, we need to consider motion planning to shoot the hard-to-model deformable ball rolling on the ground with uncertain friction to a desired location. In this paper, we propose a hierarchical framework that leverages deep reinforcement learning to train (a) a robust motion control policy that can track arbitrary motions and (b) a planning policy to decide the desired kicking motion to shoot a soccer ball to a target. We deploy the proposed framework on an A1 quadrupedal robot and enable it to accurately shoot the ball to random targets in the real world. Yandong Ji, Zhongyu Li 0003, Xue Bin Peng, Sergey Levine, Glen Berseth, Koushil Sreenath |
IROS | 6 |
| 2021 | SMiRL: Surprise Minimizing Reinforcement Learning in Unstable Environments
Glen Berseth, Daniel Geng, Coline Devin, Nicholas Rhinehart, Chelsea Finn, Dinesh Jayaraman, Sergey Levine |
ICLR | 1 |
| 2021 | X2T: Training an X-to-Text Typing Interface with Online Learning from User Feedback
Jensen Gao, Siddharth Reddy, Glen Berseth, Nicholas Hardy, Nikhilesh Natraj, Karunesh Ganguly, Anca D. Dragan, Sergey Levine |
ICLR | 3 |
| 2021 | Reinforcement Learning for Robust Parameterized Locomotion Control of Bipedal RobotsabstractDeveloping robust walking controllers for bipedal robots is a challenging endeavor. Traditional model-based locomotion controllers require simplifying assumptions and careful modelling; any small errors can result in unstable control. To address these challenges for bipedal locomotion, we present a model-free reinforcement learning framework for training robust locomotion policies in simulation, which can then be transferred to a real bipedal Cassie robot. To facilitate sim-to-real transfer, domain randomization is used to encourage the policies to learn behaviors that are robust across variations in system dynamics. The learned policies enable Cassie to perform a set of diverse and dynamic behaviors, while also being more robust than traditional controllers and prior learning-based methods that use residual control. We demonstrate this on versatile walking behaviors such as tracking a target walking velocity, walking height, and turning yaw. (Video1) Zhongyu Li 0003, Xuxin Cheng, Xue Bin Peng, Pieter Abbeel, Sergey Levine, Glen Berseth, Koushil Sreenath |
ICRA | 6 |
| 2021 | DisCo RL: Distribution-Conditioned Reinforcement Learning for General-Purpose PoliciesabstractCan we use reinforcement learning to learn general-purpose policies that can perform a wide range of different tasks, resulting in flexible and reusable skills? Contextual policies provide this capability in principle, but the representation of the context determines the degree of generalization and expressivity. Categorical contexts preclude generalization to entirely new tasks. Goal-conditioned policies may enable some generalization, but cannot capture all tasks that might be desired. In this paper, we propose goal distributions as a general and broadly applicable task representation suitable for contextual policies. Goal distributions are general in the sense that they can represent any state-based reward function when equipped with an appropriate distribution class, while the particular choice of distribution class allows us to trade off expressivity and learnability. We develop an off-policy algorithm called distribution-conditioned reinforcement learning (DisCo RL) to efficiently learn these policies. We evaluate DisCo RL on a variety of robot manipulation tasks and find that it significantly outperforms prior methods on tasks that require generalization to new goal distributions. Soroush Nasiriany, Vitchyr Pong, Ashvin Nair, Alexander Khazatsky, Glen Berseth, Sergey Levine |
ICRA | 5 |
| 2021 | Information is Power: Intrinsic Control via Information CaptureabstractHumans and animals explore their environment and acquire useful skills even in the absence of clear goals, exhibiting intrinsic motivation. The study of intrinsic motivation in artificial agents is concerned with the following question: what is a good general-purpose objective for an agent? We study this question in dynamic partially-observed environments, and argue that a compact and general learning objective is to minimize the entropy of the agent's state visitation estimated using a latent state-space model. This objective induces an agent to both gather information about its environment, corresponding to reducing uncertainty, and to gain control over its environment, corresponding to reducing the unpredictability of future world states. We instantiate this approach as a deep reinforcement learning agent equipped with a deep variational Bayes filter. We find that our agent learns to discover, represent, and exercise control of dynamic objects in a variety of partially-observed environments sensed with visual observations without extrinsic reward. Nicholas Rhinehart, Jenny Wang, Glen Berseth, John D. Co-Reyes, Danijar Hafner, Chelsea Finn, Sergey Levine |
NeurIPS | 3 |
| 2021 | Interactive Architectural Design with Diverse Solution ExplorationabstractIn architectural design, architects explore a vast amount of design options to maximize various performance criteria, while adhering to specific constraints. In an effort to assist architects in such a complex endeavour, we propose IDOME, an interactive system for computer-aided design optimization. Our approach balances automation and control by efficiently exploring, analyzing, and filtering space layouts to inform architects' decision-making better. At each design iteration, IDOME provides a set of alternative building layouts which satisfy user-defined constraints and optimality criteria concerning a user-defined space parametrization. When the user selects a design generated by IDOME, the system performs a similar optimization process with the same (or different) parameters and objectives. A user may iterate this exploration process as many times as needed. In this work, we focus on optimizing built environments using architectural metrics by improving the degree of visibility, accessibility, and information gaining for navigating a proposed space. This approach, however, can be extended to support other kinds of analysis as well. We demonstrate the capabilities of IDOME through a series of examples, performance analysis, user studies, and a usability test. The results indicate that IDOME successfully optimizes the proposed designs concerning the chosen metrics and offers a satisfactory experience for users with minimal training. Glen Berseth, M. Brandon Haworth, Muhammad Usman 0010, Davide Schaumann, Mahyar Khayatkhoei, Mubbasir Kapadia, Petros Faloutsos |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Deep Integration of Physical Humanoid Control and Crowd NavigationabstractMany multi-agent navigation approaches make use of simplified representations such as a disk. These simplifications allow for fast simulation of thousands of agents but limit the simulation accuracy and fidelity. In this paper, we propose a fully integrated physical character control and multi-agent navigation method. In place of sample complex online planning methods, we extend the use of recent deep reinforcement learning techniques. This extension improves on multi-agent navigation models and simulated humanoids by combining Multi-Agent and Hierarchical Reinforcement Learning. We train a single short term goal-conditioned low-level policy to provide directed walking behaviour. This task-agnostic controller can be shared by higher-level policies that perform longer-term planning. The proposed approach produces reciprocal collision avoidance, robust navigation, and emergent crowd behaviours. Furthermore, it offers several key affordances not previously possible in multi-agent navigation including tunable character morphology and physically accurate interactions with agents and the environment. Our results show that the proposed method outperforms prior methods across environments and tasks, as well as, performing well in terms of zero-shot generalization over different numbers of agents and computation time. M. Brandon Haworth, Glen Berseth, Seonghyeon Moon, Petros Faloutsos, Mubbasir Kapadia |
MIG | 2 |
| 2018 | Progressive Reinforcement Learning with Distillation for Multi-Skilled Motion Control
Glen Berseth, Paul Cernek, Michiel van de Panne |
ICLR (Poster) | 1 |
| 2018 | Model-Based Action Exploration for Learning Dynamic Motion SkillsabstractDeep reinforcement learning has achieved great strides in solving challenging motion control tasks. Recently, there has been significant work on methods for exploiting the data gathered during training, but there has been less work on how to best generate the data to learn from. For continuous action domains, the most common method for generating exploratory actions involves sampling from a Gaussian distribution centred around the mean action output by a policy. Although these methods can be quite capable, they do not scale well with the dimensionality of the action space, and can be dangerous to apply on hardware. We consider learning a forward dynamics model to predict the result, (xt+1), of taking a particular action, (u), given a specific observation of the state, (xt). With this model we perform internal lookahead predictions of outcomes and seek actions we believe have a reasonable chance of success. This method alters the exploratory action space, thereby increasing learning speed and enables higher quality solutions to difficult problems, such as robotic locomotion and juggling. Glen Berseth, Alex Kyriazis, Ivan Zinin, William Choi, Michiel van de Panne |
IROS | 1 |
| 2018 | Feedback Control For Cassie With Deep Reinforcement LearningabstractBipedal locomotion skills are challenging to develop. Control strategies often use local linearization of the dynamics in conjunction with reduced-order abstractions to yield tractable solutions. In these model-based control strategies, the controller is often not fully aware of many details, including torque limits, joint limits, and other non-linearities that are necessarily excluded from the control computations for simplicity. Deep reinforcement learning (DRL) offers a promising model-free approach for controlling bipedal locomotion which can more fully exploit the dynamics. However, current results in the machine learning literature are often based on ad-hoc simulation models that are not based on corresponding hardware. Thus it remains unclear how well DRL will succeed on realizable bipedal robots. In this paper, we demonstrate the effectiveness of DRL using a realistic model of Cassie, a bipedal robot. By formulating a feedback control problem as finding the optimal policy for a Markov Decision Process, we are able to learn robust walking controllers that imitate a reference motion with DRL. Controllers for different walking speeds are learned by imitating simple time-scaled versions of the original reference motion. Controller robustness is demonstrated through several challenging tests, including sensory delay, walking blindly on irregular terrain and unexpected pushes at the pelvis. We also show we can interpolate between individual policies and that robustness can be improved with an interpolated policy. Zhaoming Xie, Glen Berseth, Patrick Clary, Jonathan W. Hurst, Michiel van de Panne |
IROS | 2 |
| 2018 | Interactive spatial analytics for human-aware building designabstractWe present a computational spatial analytics tool for designing environments that better support human-related factors. Our system performs both static and dynamic analyses: the first relates to the building geometry and organization, while the second additionally considers the crowd movement in the space. The results are presented to the designers in the form of numerical values, traces and heat maps displayed on top of the floor plan. We demonstrate our approach with a user study whereby novice architects have tested the proposed approach to iteratively improve a building accessibility in real-time with respect to a selected number of static and dynamic metrics. The results indicate that the users were able to successfully improve their design solutions and thus generate more human-aware environments. The usability and effectiveness of the tool where also measured, yielding positive scores. The modular and flexible nature of the tool enables further extension to incorporate additional static and dynamic spatial metrics. Muhammad Usman 0010, Davide Schaumann, M. Brandon Haworth, Glen Berseth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 4 |
| 2017 | Perceptual evaluation of space in virtual environmentsabstractFloor plan designs and their spatial analysis are typically constrained to blueprints and 2D projections of 3D models. Computing appropriate spatial measures from such representations provides a standard way of quantifying important aspects of the design. We wish to investigate whether a person's perceptual exploration of a space would agree with such spatial measures, that is, whether a person can roughly infer such measures by exploring a space. We perform two studies, one involving novices and the other experts. First, we conduct a perceptual study to discover whether a novice user's perception of spatial measures depends on the mode used to explore the space. Our analysis considers three spatial measures, grounded in Space-Syntax, that characterize key aspects of a design such as visibility, accessibility, and organization. We compare three modes of exploration: 2D blueprints, first-person view in a 3D simulation, and a 3D virtual reality simulation with teleportation. A correlation analysis between the users' perceptual ratings and the spatial measures, indicates that virtual reality is the most effective of the three methods, while 2D blueprints and 3D first-person exploration often fail entirely to convey the spatial measures. In the second study, experts are asked to evaluate and rank the design blueprints for each measure. The expert observations are in strong agreement with the spatial measures for accessibility and organization, but not for visibility in some cases. This indicates that even experts have difficulty understanding spatial aspects of an architecture design from 2D blueprints alone. Muhammad Usman 0010, M. Brandon Haworth, Glen Berseth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 3 |
| 2017 | Crowd sourced co-design of floor plans using simulation guided gamesabstractCrowd-aware environment design is a complex combinatorial decision process, where small changes in a design may affect crowd flow patterns in unexpected and potentially unintuitive ways. Existing solutions rely on expert intuition, best practices, or automation. To address the dimensionality and complexity of the design process, we propose leveraging automation and human creativity at a large scale akin to crowd sourcing, within a gamified collaborative design framework. Using our system, "players" (novice users or experts) can rapidly iterate on their designs while soliciting feedback from computer simulations of crowd movement and the designs of other players. Our approach affords a new way of thinking of the solution space in that it inherently supports competitive collaboration, co-design, and crowd sourced solutions. We evaluate our framework through a preliminary user study. Nilay Chakraborty, Glen Berseth, M. Brandon Haworth, Petros Faloutsos, Muhammad Usman 0010, Mubbasir Kapadia |
MIG | 2 |
| 2017 | On density-flow relationships during crowd evacuationabstractAbstract Traffic and pedestrian dynamics communities often use a standard qualitative classification, namely, level of service (LoS), to describe the relationship between the crowd flow and crowd density in an environment. However, this classification has not yet been rigorously studied in the application of synthetic crowds, which are derived using a variety of approaches and may model certain behaviors better than others. Although synthetic crowds can be simulated to extrapolate crowd flow for rigorous quantitative analysis, these may be at odds with the qualitative LoS. In order to successfully use computer‐assisted design, it is important to have sound quantitative metrics as the basis for analysis and optimization. In this paper, we present a systematic empirical analysis of LoS for synthetic crowds. Using established crowd simulation techniques, we quantify the relation between crowd density and crowd flow for evacuation scenarios across different simulators to explore conformity to qualitative LoS classifications. Following this study, we perform environment optimization experiments under various LoS conditions. Finally, we test the generality of optimizing under these LoS conditions. Our results motivate the need for further study, using real and synthetic crowd datasets across representative environment benchmarks. M. Brandon Haworth, Muhammad Usman 0010, Glen Berseth, Mubbasir Kapadia, Petros Faloutsos |
Comput. Animat. Virtual Worlds | 3 |
| 2017 | CODE: Crowd-optimized design of environmentsabstractAbstract We present crowd‐optimized design of environments (CODE): a “crowd‐aware” computational tool for designing environments (e.g., building floor plans). Our system analyses the impact of newly added environment elements (e.g., pillars or doorways) on the resulting crowd flow, using current‐generation crowd simulators. The results of the simulation are used to provide feedback to the designer in terms of aggregate statistics and heat maps. Additionally, our system is able to “automatically” optimize the placement of environment elements to maximize crowd flow in egress scenarios, while satisfying constraints that are imposed by the designer. Using CODE, architects and environment designers can iteratively refine upon their original design to quickly accommodate the dynamic properties of crowd simulations in an interactive fashion. CODE is modular and flexible so that designers may build environments, select from different crowd simulators, and specify varying crowd configurations. M. Brandon Haworth, Muhammad Usman 0010, Glen Berseth, Mahyar Khayatkhoei, Mubbasir Kapadia, Petros Faloutsos |
Comput. Animat. Virtual Worlds | 3 |
| 2017 | DeepLoco: dynamic locomotion skills using hierarchical deep reinforcement learningabstractLearning physics-based locomotion skills is a difficult problem, leading to solutions that typically exploit prior knowledge of various forms. In this paper we aim to learn a variety of environment-aware locomotion skills with a limited amount of prior knowledge. We adopt a two-level hierarchical control framework. First, low-level controllers are learned that operate at a fine timescale and which achieve robust walking gaits that satisfy stepping-target and style objectives. Second, high-level controllers are then learned which plan at the timescale of steps by invoking desired step targets for the low-level controller. The high-level controller makes decisions directly based on high-dimensional inputs, including terrain maps or other suitable representations of the surroundings. Both levels of the control policy are trained using deep reinforcement learning. Results are demonstrated on a simulated 3D biped. Low-level controllers are learned for a variety of motion styles and demonstrate robustness with respect to force-based disturbances, terrain variations, and style interpolation. High-level controllers are demonstrated that are capable of following trails through terrains, dribbling a soccer ball towards a target location, and navigating through static or dynamic obstacles. Xue Bin Peng, Glen Berseth, KangKang Yin, Michiel van de Panne |
ACM Trans. Graph. | 2 |
| 2016 | ACCLMesh: curvature-based navigation mesh generationabstractAbstract The proposed method computes a navigation mesh for arbitrary and dynamic 3D environments based on curvature and is robust and efficient. This method addresses a number of known limitations in state‐of‐the‐art techniques to produce navigation meshes that are tightly coupled to the original geometry, incorporate geometric details that are crucial for movement decisions, can robustly handle complex surfaces and can efficiently repair the navigation mesh to accommodate dynamically changing environments. The method is integrated into a standard navigation and collision avoidance system to simulate thousands of agents on complex 3D surfaces in real time. Copyright © 2016 John Wiley & Sons, Ltd. Glen Berseth, Mubbasir Kapadia, Petros Faloutsos |
Comput. Animat. Virtual Worlds | 1 |
| 2016 | Terrain-adaptive locomotion skills using deep reinforcement learningabstractReinforcement learning offers a promising methodology for developing skills for simulated characters, but typically requires working with sparse hand-crafted features. Building on recent progress in deep reinforcement learning (DeepRL), we introduce a mixture of actor-critic experts (MACE) approach that learns terrain-adaptive dynamic locomotion skills using high-dimensional state and terrain descriptions as input, and parameterized leaps or steps as output actions. MACE learns more quickly than a single actor-critic approach and results in actor-critic experts that exhibit specialization. Additional elements of our solution that contribute towards efficient learning include Boltzmann exploration and the use of initial actor biases to encourage specialization. Results are demonstrated for multiple planar characters and terrain classes. Xue Bin Peng, Glen Berseth, Michiel van de Panne |
ACM Trans. Graph. | 2 |
| 2015 | ACCLMesh: curvature-based navigation mesh generationabstractWe propose a method to robustly and efficiently compute a navigation mesh for arbitrary and dynamic 3D environments based on curvature. This method addresses a number of known limitations in state-of-the-art techniques to produce navigation meshes that are tightly coupled to the original geometry, incorporate geometric details that are crucial for movement decisions and robustly handle complex surfaces. We integrate the method into a standard navigation and collision-avoidance system to simulate thousands of agents on complex 3D surfaces in real-time. Glen Berseth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 1 |
| 2015 | Evaluating and optimizing level of service for crowd evacuationsabstractLevel of service (LoS) is a standard indicator, widely used in crowd management and urban design, for characterizing the service afforded by environments to crowds of specific densities. However, current LoS indicators are qualitative and rely on expert analysis. Computational approaches for crowd analysis and environment design require robust measures for characterizing the relationship between environments and crowd flow. M. Brandon Haworth, Muhammad Usman 0010, Glen Berseth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 3 |
| 2015 | Environment optimization for crowd evacuationabstractAbstract The layout of a building, real or virtual, affects the flow patterns of its intended users. It is well established, for example, that the placement of pillars at proper locations can often facilitate pedestrian flow during the evacuation of a building. Such considerations are therefore important for architects, game level developers, and others whose domains involve agents navigating through buildings. In this paper, we take the first steps towards developing a simulation framework that can be used to study the optimal placement of architectural elements, such as pillars or doors, for the purposes of facilitating dense pedestrian flow during the evacuation of a building. In particular, we show that the steering algorithms used to model the local navigation abilities of the agents significantly affect the results, which motivates the need for a statistically valid approach and further study. Copyright © 2015 John Wiley & Sons, Ltd. Glen Berseth, Muhammad Usman 0010, M. Brandon Haworth, Mubbasir Kapadia, Petros Faloutsos |
Comput. Animat. Virtual Worlds | 1 |
| 2015 | Dynamic terrain traversal skills using reinforcement learningabstractThe locomotion skills developed for physics-based characters most often target flat terrain. However, much of their potential lies with the creation of dynamic, momentum-based motions across more complex terrains. In this paper, we learn controllers that allow simulated characters to traverse terrains with gaps, steps, and walls using highly dynamic gaits. This is achieved using reinforcement learning, with careful attention given to the action representation, non-parametric approximation of both the value function and the policy; epsilon-greedy exploration; and the learning of a good state distance metric. The methods enable a 21-link planar dog and a 7-link planar biped to navigate challenging sequences of terrain using bounding and running gaits. We evaluate the impact of the key features of our skill learning pipeline on the resulting performance. Xue Bin Peng, Glen Berseth, Michiel van de Panne |
ACM Trans. Graph. | 2 |
| 2014 | Characterizing and optimizing game level difficultyabstractBalancing the interactions between game level design and intended player experience is a difficult and time consuming process. Automating aspects of this process with respect to user-defined constraints has beneficial implications for game designers. A change in level layout may affect the available routes and subsequent player interactions for a number of agents within the level. Small changes in the placement of game elements may lead to significant changes in terms of the challenge experienced by the player on the path to their goal. Estimating the effect of this change requires that the designer take into account new paths of all interacting agents and how these may affect the player. As the number of these agents grow to crowd size, estimating the effect of these changes becomes grows difficult. We present a user-in-the-loop framework for tackling this task by optimizing enemy agent settings and the placement of game elements that affect the flow of agents within the level, with respect to estimated difficulty. Using static path analysis we estimate difficulty based on agent interactions with the player. To exemplify the usefulness of the framework, we show that small changes in level layout lead to significant changes in game difficulty, and optimizations with respect to the characterization of difficulty can be used to attain desired difficulty levels. Glen Berseth, M. Brandon Haworth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 1 |
| 2013 | SteerPlex: Estimating Scenario Complexity for Simulated CrowdsabstractThe complexity of interactive virtual worlds has increased dramatically in recent years, with a rise in mature solutions for designing large-scale environments and populating them with hundreds and thousands of autonomous characters. An interesting problem that arises in this context, and that has received little attention to date, is whether we can predict the complexity of a steering scenario by analyzing the configuration of the environment and the agents involved. We statically analyze an input scenario and compute a set of novel salient features which characterize the expected interactions between agents and obstacles during simulation. Using a statistical approach, we automatically derive the relative influence of each feature on the complexity of a scenario in order to derive a single numerical quantity of expected scenario complexity. We validate our proposed metric by demonstrating a strong negative correlation between the statically computed expected complexity and the dynamic performance of three published crowd simulation techniques. Glen Berseth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 1 |