EDBT 2026 Demo / reviewers in the wild / expert
Joshua Romoff
dblp:192/1508
· DBLP profile ↗
9ranked-venue papers
1as first author
4since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Reinforcement learning · 96% Efficient and distributed learning · 4% | |
| Human-computer interaction and pervasive computing
1 paper |
Games and playful interaction · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › exploration › novelty-based exploration
count-based exploration |
0.8 | 1 | 2024 | Improving Intrinsic Exploration by Creating Stationary Objectives · ICLR 2024 |
Machine learning › Reinforcement learning
exploration |
0.8 | 1 | 2024 | Improving Intrinsic Exploration by Creating Stationary Objectives · ICLR 2024 |
Machine learning › Reinforcement learning › exploration
intrinsic motivation |
0.8 | 1 | 2024 | Improving Intrinsic Exploration by Creating Stationary Objectives · ICLR 2024 |
Machine learning › Reinforcement learning › reward design
reward shaping |
0.8 | 1 | 2024 | Improving Intrinsic Exploration by Creating Stationary Objectives · ICLR 2024 |
Machine learning › Reinforcement learning › markov decision process
constrained markov decision process |
0.6 | 1 | 2022 | Direct Behavior Specification via Constrained Reinforcement Learning · ICML 2022 |
Machine learning › Reinforcement learning
constrained reinforcement learning |
0.6 | 1 | 2022 | Direct Behavior Specification via Constrained Reinforcement Learning · ICML 2022 |
Machine learning › Reinforcement learning
safe reinforcement learning |
0.6 | 1 | 2022 | Direct Behavior Specification via Constrained Reinforcement Learning · ICML 2022 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.5 | 1 | 2021 | Deep Reinforcement Learning for Navigation in AAA Video Games · IJCAI 2021 |
Machine learning › Reinforcement learning › large-scale reinforcement learning › distributed reinforcement learning
actor-learner architecture |
0.4 | 1 | 2019 | Gossip-based Actor-Learner Architectures for Deep Reinforcement Learning · NeurIPS 2019 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative multi-agent reinforcement learning |
0.4 | 1 | 2019 | TarMAC: Targeted Multi-Agent Communication · ICML 2019 |
Machine learning › Reinforcement learning › large-scale reinforcement learning
distributed reinforcement learning |
0.4 | 1 | 2019 | Gossip-based Actor-Learner Architectures for Deep Reinforcement Learning · NeurIPS 2019 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.4 | 1 | 2019 | TarMAC: Targeted Multi-Agent Communication · ICML 2019 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
temporal abstraction |
0.4 | 1 | 2019 | Separable value functions across time-scales · ICML 2019 |
Machine learning › Reinforcement learning
value function |
0.4 | 1 | 2019 | Separable value functions across time-scales · ICML 2019 |
Machine learning › Reinforcement learning
value function approximation |
0.3 | 1 | 2017 | Hybrid Reward Architecture for Reinforcement Learning · NIPS 2017 |
Machine learning › Reinforcement learning
partially observable reinforcement learning |
0.1 | 1 | 2019 | TarMAC: Targeted Multi-Agent Communication · ICML 2019 |
Machine learning › Reinforcement learning › reward design
reward decomposition |
0.1 | 1 | 2017 | Hybrid Reward Architecture for Reinforcement Learning · NIPS 2017 |
Methods — techniques the papers use, named apart from their topics
navigation mesh · 1.0deep reinforcement learning · 1.0state augmentation · 0.8deep network · 0.8lagrangian method · 0.6CMDP · 0.6targeted messaging · 0.4multi-round communication · 0.4gossip protocol · 0.4a2c · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Improving Intrinsic Exploration by Creating Stationary ObjectivesabstractExploration bonuses in reinforcement learning guide long-horizon exploration by defining custom intrinsic objectives. Count-based methods use the frequency of state visits to derive an exploration bonus. In this paper, we identify that any intrinsic reward function derived from count-based methods is non-stationary and hence induces a difficult objective to optimize for the agent. The key contribution of our work lies in transforming the original non-stationary rewards into stationary rewards through an augmented state representation. For this purpose, we introduce the Stationary Objectives For Exploration (SOFE) framework. SOFE requires *identifying* sufficient statistics for different exploration bonuses and finding an *efficient* encoding of these statistics to use as input to a deep network. SOFE is based on proposing state augmentations that expand the state space but hold the promise of simplifying the optimization of the agent's objective. Our experiments show that SOFE improves the agents' performance in challenging exploration problems, including sparse-reward tasks, pixel-based observations, 3D navigation, and procedurally generated environments. Roger Creus Castanyer, Joshua Romoff, Glen Berseth |
ICLR | 2 |
| 2023 | Learning Computational Efficient Bots with Costly FeaturesabstractDeep reinforcement learning (DRL) techniques have become increasingly used in various fields for decision-making processes. However, a challenge that often arises is the trade-off between both the computational efficiency of the decision-making process and the ability of the learned agent to solve a particular task. This is particularly critical in real-time settings such as video games where the agent needs to take relevant decisions at a very high frequency, with a very limited inference timeIn this work, we propose a generic offline learning approach where the computation cost of the input features is taken into account. We derive the Budgeted Decision Transformer as an extension of the Decision Transformer that incorporates cost constraints to limit its cost at inference. As a result, the model can dynamically choose the best input features at each timestep. We demonstrate the effectiveness of our method on several tasks, including D4RL benchmarks and complex 3D environments similar to those found in video games, and show that it can achieve similar performance while using significantly fewer computational resources compared to classical approaches. Anthony Kobanda, Valliappan C. A., Joshua Romoff, Ludovic Denoyer |
CoG | 3 |
| 2022 | Direct Behavior Specification via Constrained Reinforcement LearningabstractThe standard formulation of Reinforcement Learning lacks a practical way of specifying what are admissible and forbidden behaviors. Most often, practitioners go about the task of behavior specification by manually engineering the reward function, a counter-intuitive process that requires several iterations and is prone to reward hacking by the agent. In this work, we argue that constrained RL, which has almost exclusively been used for safe RL, also has the potential to significantly reduce the amount of work spent for reward specification in applied RL projects. To this end, we propose to specify behavioral preferences in the CMDP framework and to use Lagrangian methods to automatically weigh each of these behavioral constraints. Specifically, we investigate how CMDPs can be adapted to solve goal-based tasks while adhering to several constraints simultaneously. We evaluate this framework on a set of continuous control tasks relevant to the application of Reinforcement Learning for NPC design in video games. Julien Roy, Roger Girgis, Joshua Romoff, Pierre-Luc Bacon, Christopher Joseph Pal |
ICML | 3 |
| 2021 | Deep Reinforcement Learning for Navigation in AAA Video GamesabstractIn video games, \non-player characters (NPCs) are used to enhance the players' experience in a variety of ways, e.g., as enemies, allies, or innocent bystanders. A crucial component of NPCs is navigation, which allows them to move from one point to another on the map. The most popular approach for NPC navigation in the video game industry is to use a navigation mesh (NavMesh), which is a graph representation of the map, with nodes and edges indicating traversable areas. Unfortunately, complex navigation abilities that extend the character's capacity for movement, e.g., grappling hooks, jetpacks, teleportation, or double-jumps, increase the complexity of the NavMesh, making it intractable in many practical scenarios. Game designers are thus constrained to only add abilities that can be handled by a NavMesh. As an alternative to the NavMesh, we propose to use Deep Reinforcement Learning (Deep RL) to learn how to navigate 3D maps in video games using any navigation ability. We test our approach on complex 3D environments that are notably an order of magnitude larger than maps typically used in the Deep RL literature. One of these environments is from a recently released AAA video game called Hyper Scape. We find that our approach performs surprisingly well, achieving at least 90% success rate in a variety of scenarios using complex navigation abilities. Eloi Alonso, Maxim Peter, David Goumard, Joshua Romoff |
IJCAI | 4 |
| 2019 | TarMAC: Targeted Multi-Agent CommunicationabstractWe propose a targeted communication architecture for multi-agent reinforcement learning, where agents learn both what messages to send and whom to address them to while performing cooperative tasks in partially-observable environments. This targeting behavior is learnt solely from downstream task-specific reward without any communication supervision. We additionally augment this with a multi-round communication approach where agents coordinate via multiple rounds of communication before taking actions in the environment. We evaluate our approach on a diverse set of cooperative multi-agent tasks, of varying difficulties, with varying number of agents, in a variety of environments ranging from 2D grid layouts of shapes and simulated traffic junctions to 3D indoor environments, and demonstrate the benefits of targeted and multi-round communication. Moreover, we show that the targeted communication strategies learned by agents are interpretable and intuitive. Finally, we show that our architecture can be easily extended to mixed and competitive environments, leading to improved performance and sample complexity over recent state-of-the-art approaches. Théophile Gervet, Joshua Romoff, Dhruv Batra, Devi Parikh, Michael G. Rabbat, Joelle Pineau |
ICML | 3 |
| 2019 | Separable value functions across time-scales
Joshua Romoff, Peter Henderson 0002, Ahmed Touati, Yann Ollivier, Joelle Pineau, Emma Brunskill |
ICML | 1 |
| 2019 | Gossip-based Actor-Learner Architectures for Deep Reinforcement LearningabstractMulti-simulator training has contributed to the recent success of Deep Reinforcement Learning (Deep RL) by stabilizing learning and allowing for higher training throughputs. In this work, we propose Gossip-based Actor-Learner Architectures (GALA) where several actor-learners (such as A2C agents) are organized in a peer-to-peer communication topology, and exchange information through asynchronous gossip in order to take advantage of a large number of distributed simulators. We prove that GALA agents remain within an epsilon-ball of one-another during training when using loosely coupled asynchronous communication. By reducing the amount of synchronization between agents, GALA is more computationally efficient and scalable compared to A2C, its fully-synchronous counterpart. GALA also outperforms A2C, being more robust and sample efficient. We show that we can run several loosely coupled GALA agents in parallel on a single GPU and achieve significantly higher hardware utilization and frame-rates than vanilla A2C at comparable power draws. Mido Assran, Joshua Romoff, Nicolas Ballas, Joelle Pineau, Michael G. Rabbat |
NeurIPS | 2 |
| 2019 | Randomized Value Functions via Multiplicative Normalizing Flows
Ahmed Touati, Harsh Satija, Joshua Romoff, Joelle Pineau, Pascal Vincent |
UAI | 3 |
| 2017 | Hybrid Reward Architecture for Reinforcement LearningabstractOne of the main challenges in reinforcement learning (RL) is generalisation. In typical deep RL methods this is achieved by approximating the optimal value function with a low-dimensional representation using a deep network. While this approach works well in many domains, in domains where the optimal value function cannot easily be reduced to a low-dimensional representation, learning can be very slow and unstable. This paper contributes towards tackling such challenging domains, by proposing a new method, called Hybrid Reward Architecture (HRA). HRA takes as input a decomposed reward function and learns a separate value function for each component reward function. Because each component typically only depends on a subset of all features, the corresponding value function can be approximated more easily by a low-dimensional representation, enabling more effective learning. We demonstrate HRA on a toy-problem and the Atari game Ms. Pac-Man, where HRA achieves above-human performance. Harm van Seijen, Mehdi Fatemi, Romain Laroche, Joshua Romoff, Tavian Barnes, Jeffrey Tsang |
NIPS | 4 |