Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sébastien Racanière

dblp:02/9996 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
0since 2021 · last 2020
0000-0003-2285-8633ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-authorSystems, architecture and hardware · 2Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Reinforcement learning · 30% Generative modeling · 17% Optimization for machine learning · 11%
Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 100%

Topics — the 24 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
model-based reinforcement learning
1.032019
Woulda, Coulda, Shoulda: Counterfactually-Guided Policy Search · ICLR (Poster) 2019
Imagination-Augmented Agents for Deep Reinforcement Learning · NIPS 2017
Recurrent Environment Simulators · ICLR (Poster) 2017
Machine learning › Optimization for machine learning
differentiable games
0.722019
Differentiable Game Mechanics · J. Mach. Learn. Res. 2019
The Mechanics of n-Player Differentiable Games · ICML 2018
Machine learning › Reinforcement learning
model-free reinforcement learning
0.722019
An Investigation of Model-Free Planning · ICML 2019
Imagination-Augmented Agents for Deep Reinforcement Learning · NIPS 2017
Machine learning › Learning paradigms › curriculum learning
automatic curriculum generation
0.412020
Automated curriculum generation through setter-solver interactions · ICLR 2020
Machine learning › Learning paradigms
curriculum learning
0.412020
Automated curriculum generation through setter-solver interactions · ICLR 2020
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.412020
Disentangling by Subspace Diffusion · NeurIPS 2020
Machine learning › Time series and sequential data
dynamical system learning
0.412020
Hamiltonian Generative Networks · ICLR 2020
Machine learning › Generative modeling
energy-based model
0.412020
Hamiltonian Generative Networks · ICLR 2020
Machine learning › Probabilistic and Bayesian machine learning › dynamical system
hamiltonian dynamics
0.412020
Hamiltonian Generative Networks · ICLR 2020
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
0.412020
Disentangling by Subspace Diffusion · NeurIPS 2020
Machine learning › Generative modeling
normalizing flow
0.412020
Normalizing Flows on Tori and Spheres · ICML 2020
Machine learning › Generative modeling › normalizing flow
normalizing flows on manifolds
0.412020
Normalizing Flows on Tori and Spheres · ICML 2020
Machine learning › Generative modeling › diffusion model › efficient diffusion model
subspace diffusion model
0.412020
Disentangling by Subspace Diffusion · NeurIPS 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
counterfactual reasoning
0.412019
Woulda, Coulda, Shoulda: Counterfactually-Guided Policy Search · ICLR (Poster) 2019
Machine learning › Optimization for machine learning
gradient-based optimization
0.412019
Differentiable Game Mechanics · J. Mach. Learn. Res. 2019
Machine learning › Reinforcement learning
policy search
0.412019
Woulda, Coulda, Shoulda: Counterfactually-Guided Policy Search · ICLR (Poster) 2019
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
search-based planning
0.412019
An Investigation of Model-Free Planning · ICML 2019
Knowledge, reasoning and agents › Multi-agent systems › game theory
multi-player games
0.312018
The Mechanics of n-Player Differentiable Games · ICML 2018
Algorithmic game theory and mechanism design › solution concepts in games › equilibrium concepts
nash equilibrium
0.312018
The Mechanics of n-Player Differentiable Games · ICML 2018
Machine learning › Reinforcement learning
policy learning
0.312017
Imagination-Augmented Agents for Deep Reinforcement Learning · NIPS 2017
Machine learning › Reinforcement learning › reinforcement learning environment
simulation environment
0.312017
Recurrent Environment Simulators · ICLR (Poster) 2017
Algorithmic game theory and mechanism design › non-cooperative game
potential game
0.112019
Differentiable Game Mechanics · J. Mach. Learn. Res. 2019
Machine learning › Reinforcement learning › model-based reinforcement learning › model-based planning
planning with learned models
0.112017
Imagination-Augmented Agents for Deep Reinforcement Learning · NIPS 2017
Robotics › Motion planning and robot control
robot control
0.112017
Recurrent Environment Simulators · ICLR (Poster) 2017

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 0.4normalizing flow · 0.4neural network · 0.4lie group theory · 0.4hamiltonian mechanics · 0.4diffusion · 0.4de rham decomposition · 0.4gradient descent · 0.4game jacobian decomposition · 0.4end-to-end training · 0.4convolutional network · 0.4adversarial training · 0.4LSTM · 0.4symplectic gradient adjustment · 0.3potential game · 0.3hamiltonian games · 0.3
YearPublicationVenuePosition
2020 Automated curriculum generation through setter-solver interactions
Sébastien Racanière, Andrew K. Lampinen, Adam Santoro, David P. Reichert, Vlad Firoiu, Timothy P. Lillicrap
ICLR1
2020 Hamiltonian Generative Networks
Peter Toth, Danilo Jimenez Rezende, Andrew Jaegle, Sébastien Racanière, Aleksandar Botev, Irina Higgins
ICLR4
2020 Normalizing Flows on Tori and Spheres
abstract
Normalizing flows are a powerful tool for building expressive distributions in high dimensions. So far, most of the literature has concentrated on learning flows on Euclidean spaces. Some problems however, such as those involving angles, are defined on spaces with more complex geometries, such as tori or spheres. In this paper, we propose and compare expressive and numerically stable flows on such spaces. Our flows are built recursively on the dimension of the space, starting from flows on circles, closed intervals or spheres.
Danilo Jimenez Rezende, George Papamakarios, Sébastien Racanière, Michael S. Albergo, Gurtej Kanwar, Phiala E. Shanahan, Kyle Cranmer
ICML3
2020 Disentangling by Subspace Diffusion
abstract
We present a novel nonparametric algorithm for symmetry-based disentangling of data manifolds, the Geometric Manifold Component Estimator (GEOMANCER). GEOMANCER provides a partial answer to the question posed by Higgins et al.(2018): is it possible to learn how to factorize a Lie group solely from observations of the orbit of an object it acts on? We show that fully unsupervised factorization of a data manifold is possible if the true metric of the manifold is known and each factor manifold has nontrivial holonomy – for example, rotation in 3D. Our algorithm works by estimating the subspaces that are invariant under random walk diffusion, giving an approximation to the de Rham decomposition from differential geometry. We demonstrate the efficacy of GEOMANCER on several complex synthetic manifolds. Our work reduces the question of whether unsupervised disentangling is possible to the question of whether unsupervised metric learning is possible, providing a unifying insight into the geometric nature of representation learning.
David Pfau, Irina Higgins, Aleksandar Botev, Sébastien Racanière
NeurIPS4
2019 Woulda, Coulda, Shoulda: Counterfactually-Guided Policy Search
Lars Buesing, Theophane Weber, Yori Zwols, Nicolas Heess, Sébastien Racanière, Arthur Guez, Jean-Baptiste Lespiau
ICLR (Poster)5
2019 An Investigation of Model-Free Planning
abstract
The field of reinforcement learning (RL) is facing increasingly challenging domains with combinatorial complexity. For an RL agent to address these challenges, it is essential that it can plan effectively. Prior work has typically utilized an explicit model of the environment, combined with a specific planning algorithm (such as tree search). More recently, a new family of methods have been proposed that learn how to plan, by providing the structure for planning via an inductive bias in the function approximator (such as a tree structured neural network), trained end-to-end by a model-free RL algorithm. In this paper, we go even further, and demonstrate empirically that an entirely model-free approach, without special structure beyond standard neural network components such as convolutional networks and LSTMs, can learn to exhibit many of the characteristics typically associated with a model-based planner. We measure our agent’s effectiveness at planning in terms of its ability to generalize across a combinatorial and irreversible state space, its data efficiency, and its ability to utilize additional thinking time. We find that our agent has many of the characteristics that one might expect to find in a planning algorithm. Furthermore, it exceeds the state-of-the-art in challenging combinatorial domains such as Sokoban and outperforms other model-free approaches that utilize strong inductive biases toward planning.
Arthur Guez, Mehdi Mirza, Karol Gregor, Rishabh Kabra, Sébastien Racanière, Theophane Weber, David Raposo, Adam Santoro, Laurent Orseau, Tom Eccles, Greg Wayne, David Silver 0001, Timothy P. Lillicrap
ICML5
2019 Differentiable Game Mechanics
abstract
Deep learning is built on the foundational guarantee that gradient descent on an objective function converges to local minima. Unfortunately, this guarantee fails in settings, such as generative adversarial nets, that exhibit multiple interacting losses. The behavior of gradient-based methods in games is not well understood -- and is becoming increasingly important as adversarial and multi-objective architectures proliferate. In this paper, we develop new tools to understand and control the dynamics in $n$-player differentiable games. The key result is to decompose the game Jacobian into two components. The first, symmetric component, is related to potential games, which reduce to gradient descent on an implicit function. The second, antisymmetric component, relates to Hamiltonian games, a new class of games that obey a conservation law akin to conservation laws in classical mechanical systems. The decomposition motivates Symplectic Gradient Adjustment (SGA), a new algorithm for finding stable fixed points in differentiable games. Basic experiments show SGA is competitive with recently proposed algorithms for finding stable fixed points in GANs -- while at the same time being applicable to, and having guarantees in, much more general cases.
Alistair Letcher, David Balduzzi, Sébastien Racanière, James Martens, Jakob N. Foerster, Karl Tuyls, Thore Graepel
J. Mach. Learn. Res.3
2018 The Mechanics of n-Player Differentiable Games
abstract
The cornerstone underpinning deep learning is the guarantee that gradient descent on an objective converges to local minima. Unfortunately, this guarantee fails in settings, such as generative adversarial nets, where there are multiple interacting losses. The behavior of gradient-based methods in games is not well understood – and is becoming increasingly important as adversarial and multi-objective architectures proliferate. In this paper, we develop new techniques to understand and control the dynamics in general games. The key result is to decompose the second-order dynamics into two components. The first is related to potential games, which reduce to gradient descent on an implicit function; the second relates to Hamiltonian games, a new class of games that obey a conservation law, akin to conservation laws in classical mechanical systems. The decomposition motivates Symplectic Gradient Adjustment (SGA), a new algorithm for finding stable fixed points in general games. Basic experiments show SGA is competitive with recently proposed algorithms for finding local Nash equilibria in GANs – whilst at the same time being applicable to – and having guarantees in – much more general games.
David Balduzzi, Sébastien Racanière, James Martens, Jakob N. Foerster, Karl Tuyls, Thore Graepel
ICML2
2017 Recurrent Environment Simulators
Silvia Chiappa, Sébastien Racanière, Daan Wierstra, Shakir Mohamed
ICLR (Poster)2
2017 Imagination-Augmented Agents for Deep Reinforcement Learning
abstract
We introduce Imagination-Augmented Agents (I2As), a novel architecture for deep reinforcement learning combining model-free and model-based aspects. In contrast to most existing model-based reinforcement learning and planning methods, which prescribe how a model should be used to arrive at a policy, I2As learn to interpret predictions from a trained environment model to construct implicit plans in arbitrary ways, by using the predictions as additional context in deep policy networks. I2As show improved data efficiency, performance, and robustness to model misspecification compared to several strong baselines.
Sébastien Racanière, Theophane Weber, David P. Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adrià Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li 0001, Razvan Pascanu, Peter W. Battaglia, Demis Hassabis, David Silver 0001, Daan Wierstra
NIPS1
2013 An FPGA-Based Data Flow Engine for Gaussian Copula Model
abstract
The Gaussian Copula Model (GCM) plays an important role in the state-of-the-art financial analysis field for modeling the dependence of financial assets. However, the existing implementations of GCM are all computationallydemanding and time-consuming. In this paper, we propose a Dataflow Engine (DFE) design to accelerate the GCM computation. Specifically, a commonly used CPU-friendly GCM algorithm is converted into a fully-pipelined dataflow graph through four steps of optimization: recomposing the algorithm to be pipeline-friendly, removing unnecessary computation, sharing common computing results, and reducing the computing precision while maintaining the same level of accuracy for the computation results. The performance of the proposed DFE design is compared with three CPU-based implementations that are well-optimized. Experimental results show that our DFE solution not only generates fairly accurate result, but also achieves a maximum of 467x speedup over a single-thread CPU-based solution, 120x speedup over a multi-thread CPUbased solution, and 47x speedup over an MPI-based solution.
Huabin Ruan, Xiaomeng Huang, Haohuan Fu, Guangwen Yang 0002, Wayne Luk, Sébastien Racanière, Oliver Pell, Wenjing Han
FCCM6
2012 Rapid computation of value and risk for derivatives portfolios
abstract
SUMMARY We report new results from an on‐going project to accelerate derivatives computations. Our earlier work was focused on accelerating the valuation of credit derivatives. In this paper, we extend our work in two ways: by applying the same techniques, first, to accelerate the computation of portfolio level risk for credit derivatives and, second, to different asset classes using a different type of mathematical model, which together present challenges that are quite different to those dealt with in our earlier work. Specifically, we report acceleration over 270 times faster than a single Intel Core for a multi‐asset Monte Carlo model. We also explore the implications for risk. Copyright © 2011 John Wiley & Sons, Ltd.
Stephen Weston, James Spooner, Sébastien Racanière, Oskar Mencer
Concurr. Comput. Pract. Exp.3
2011 Accelerating Large-Scale HPC Applications Using FPGAs
abstract
Field Programmable Gate Arrays (FPGAs) are conventionally considered as 'glue-logic'. However, modern FPGAs are extremely competitive compared to state-of-the-art CPUs for commercial HPC workloads, such as those found in Oil and Gas and Finance. For example, an FPGA accelerated system can be 31-37 times faster than an equivalently sized conventional machine, and consume 1/39 of the power. The key to achieving the best performance in FPGA accelerators, while maintaining correctness, is optimization of arithmetic units and data types to suit the range/precision at each point in the computation. The flexibility of the FPGA to implement non-standard arithmetic, combined with a data-flow programming model that instantiates a separate unit for each arithmetic operator in the code provides a wide design space. As such, FPGA computing offers significant opportunity for arithmetic research into 'large scale' HPC applications, where there is an opportunity to move away from standard IEEE formats, either to improve precision compared to the CPU version or to increase speed.
Robert G. Dimond, Sébastien Racanière, Oliver Pell
IEEE Symposium on Computer Arithmetic2