Philemon Brakel

dblp:82/10570 · also Philémon Brakel · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 67% Deep learning architectures and training · 12% Learning theory · 6%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
actor-critic methods
1.022024
Offline Actor-Critic Reinforcement Learning Scales to Large Models · ICML 2024
An Actor-Critic Algorithm for Sequence Prediction · ICLR (Poster) 2017
Machine learning › Reinforcement learning
multi-task reinforcement learning
0.812024
Offline Actor-Critic Reinforcement Learning Scales to Large Models · ICML 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Offline Actor-Critic Reinforcement Learning Scales to Large Models · ICML 2024
Machine learning › Reinforcement learning
exploration
0.412019
Recall Traces: Backtracking Models for Efficient Reinforcement Learning · ICLR (Poster) 2019
Machine learning › Learning theory › online learning
sequence prediction
0.312017
An Actor-Critic Algorithm for Sequence Prediction · ICLR (Poster) 2017
Machine learning › Deep learning architectures and training › feedforward neural network
ladder networks
0.212016
Deconstructing the Ladder Network Architecture · ICML 2016
Machine learning › Learning paradigms
semi-supervised learning
0.212016
Deconstructing the Ladder Network Architecture · ICML 2016
Machine learning › Reinforcement learning › deep reinforcement learning
scaling laws for reinforcement learning
0.212024
Offline Actor-Critic Reinforcement Learning Scales to Large Models · ICML 2024
Machine learning › Generative modeling
energy-based model
0.212013
Training energy-based models for time-series imputation · J. Mach. Learn. Res. 2013
Machine learning › Deep learning architectures and training
sequence modeling
0.212013
Training energy-based models for time-series imputation · J. Mach. Learn. Res. 2013
Machine learning › Time series and sequential data › time series analysis
time series imputation
0.212013
Training energy-based models for time-series imputation · J. Mach. Learn. Res. 2013
Machine learning › Deep learning architectures and training
modular neural network
0.112012
Oger: modular learning architectures for large-scale sequential processing · J. Mach. Learn. Res. 2012
Machine learning › Efficient and distributed learning
scalable learning
0.112012
Oger: modular learning architectures for large-scale sequential processing · J. Mach. Learn. Res. 2012

Methods — techniques the papers use, named apart from their topics

transformer · 0.8perceiver · 0.8behavioral cloning · 0.8reinforcement learning · 0.4backtracking models · 0.4policy gradient · 0.3actor-critic · 0.3noise injection · 0.2ablation study · 0.2energy-based model training · 0.2
YearPublicationVenuePosition
2025 Exploiting Policy Idling for Dexterous Manipulation
abstract
Learning based methods for dexterous manipulation have made notable progress in recent years, and they can now produce solutions to complex tasks. However, learned policies often still lack reliability and exhibit limited robustness to important factors of variation. One failure pattern that can be observed across many settings is that policies idle, i.e. they cease to move beyond a small region of states, often indefinitely, when they reach certain states. This policy idling is often a reflection of the training data. For instance, it can occur when the data contains small actions in areas where the robot needs to perform high-precision motions, e.g., when preparing to grasp an object or object insertion. Prior works have tried to mitigate this phenomenon e.g. by filtering the training data or modifying the control frequency. However, these approaches can negatively impact policy performance in other ways. As an alternative, we investigate how to leverage the detectability of idling behavior to inform exploration and policy improvement. Our approach, Pause-Induced Perturbations (PIP), applies perturbations at detected idling states, thus helping it to escape problematic basins of attraction. On a range of challenging simulated dual-arm tasks, we find that this simple approach can already noticeably improve test-time performance, with no additional supervision or training. Furthermore, since the robot tends to idle at critical points in a movement, we also find that learning from the resulting episodes leads to better iterative policy improvement compared to prior approaches. Our perturbation strategy also leads to a 15-35% improvement in absolute success rate on a real-world insertion task that requires complex multi-finger manipulation.
Annie S. Chen, Philemon Brakel, Antonia Bronars, Annie Xie, Sandy H. Huang, Oliver Groth, Maria Bauzá 0001, Markus Wulfmeier, Nicolas Heess, Dushyant Rao
IROS2
2024 Offline Actor-Critic Reinforcement Learning Scales to Large Models
abstract
We show that offline actor-critic reinforcement learning can scale to large models - such as transformers - and follows similar scaling laws as supervised learning. We find that offline actor-critic algorithms can outperform strong, supervised, behavioral cloning baselines for multi-task training on a large dataset; containing both sub-optimal and expert behavior on 132 continuous control tasks. We introduce a Perceiver-based actor-critic model and elucidate the key features needed to make offline RL work with self- and cross-attention modules. Overall, we find that: i) simple offline actor critic algorithms are a natural choice for gradually moving away from the currently predominant paradigm of behavioral cloning, and ii) via offline RL it is possible to learn multi-task policies that master many domains simultaneously, including real robotics tasks, from sub-optimal demonstrations or self-generated data.
Jost Tobias Springenberg, Abbas Abdolmaleki, Jingwei Zhang 0001, Oliver Groth, Michael Bloesch, Thomas Lampe, Philemon Brakel, Sarah Bechtle, Steven Kapturowski, Roland Hafner, Nicolas Heess, Martin A. Riedmiller
ICML7
2022 Learning Coordinated Terrain-Adaptive Locomotion by Imitating a Centroidal Dynamics Planner
abstract
We propose a simple imitation learning procedure for learning locomotion controllers that can walk over very challenging terrains. We use trajectory optimization (TO) to produce a large dataset of trajectories over procedurally generated terrains and use Reinforcement Learning (RL) to imitate these trajectories. We demonstrate with a realistic model of the ANYmal robot that the learned controllers transfer to unseen terrains and provide an effective initialization for fine-tuning on challenging terrains that require exteroception and precise foot placements. Our setup combines TO and RL in a simple fashion that overcomes the computational limitations and need for a robust tracking controller of the former and the exploration and reward-tuning difficulties of the latter.
Philemon Brakel, Steven Bohez, Leonard Hasenclever, Nicolas Heess, Konstantinos Bousmalis
IROS1
2019 Recall Traces: Backtracking Models for Efficient Reinforcement Learning
Anirudh Goyal, Philemon Brakel, William Fedus, Soumye Singhal, Timothy P. Lillicrap, Sergey Levine, Hugo Larochelle, Yoshua Bengio
ICLR (Poster)2
2017 A network of deep neural networks for Distant Speech Recognition
abstract
Despite the remarkable progress recently made in distant speech recognition, state-of-the-art technology still suffers from a lack of robustness, especially when adverse acoustic conditions characterized by non-stationary noises and reverberation are met. A prominent limitation of current systems lies in the lack of matching and communication between the various technologies involved in the distant speech recognition process. The speech enhancement and speech recognition modules are, for instance, often trained independently. Moreover, the speech enhancement normally helps the speech recognizer, but the output of the latter is not commonly used, in turn, to improve the speech enhancement. To address both concerns, we propose a novel architecture based on a network of deep neural networks, where all the components are jointly trained and better cooperate with each other thanks to a full communication scheme between them. Experiments, conducted using different datasets, tasks and acoustic conditions, revealed that the proposed framework can overtake other competitive solutions, including recent joint training approaches.
Mirco Ravanelli, Philemon Brakel, Maurizio Omologo, Yoshua Bengio
ICASSP2
2017 An Actor-Critic Algorithm for Sequence Prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron C. Courville, Yoshua Bengio
ICLR (Poster)2
2017 Improving Speech Recognition by Revising Gated Recurrent Units
abstract
Speech recognition is largely taking advantage of deep learning, showing that substantial benefits can be obtained by modern Recurrent Neural Networks (RNNs). The most popular RNNs are Long Short-Term Memory (LSTMs), which typically reach state-of-the-art performance in many tasks thanks to their ability to learn long-term dependencies and robustness to vanishing gradients. Nevertheless, LSTMs have a rather complex design with three multiplicative gates, that might impair their efficient implementation. An attempt to simplify LSTMs has recently led to Gated Recurrent Units (GRUs), which are based on just two multiplicative gates. This paper builds on these efforts by further revising GRUs and proposing a simplified architecture potentially more suitable for speech recognition. The contribution of this work is two-fold. First, we suggest to remove the reset gate in the GRU design, resulting in a more efficient single-gate architecture. Second, we propose to replace tanh with ReLU activations in the state update equations. Results show that, in our implementation, the revised architecture reduces the per-epoch training time with more than 30% and consistently improves recognition performance across different tasks, input features, and noisy conditions when compared to a standard GRU.
Mirco Ravanelli, Philemon Brakel, Maurizio Omologo, Yoshua Bengio
INTERSPEECH2
2016 End-to-end attention-based large vocabulary speech recognition
abstract
Many state-of-the-art Large Vocabulary Continuous Speech Recognition (LVCSR) Systems are hybrids of neural networks and Hidden Markov Models (HMMs). Recently, more direct end-to-end methods have been investigated, in which neural architectures were trained to model sequences of characters [1,2]. To our knowledge, all these approaches relied on Connectionist Temporal Classification [3] modules. We investigate an alternative method for sequence modelling based on an attention mechanism that allows a Recurrent Neural Network (RNN) to learn alignments between sequences of input frames and output labels. We show how this setup can be applied to LVCSR by integrating the decoding RNN with an n-gram language model and by speeding up its operation by constraining selections made by the attention mechanism and by reducing the source sequence lengths by pooling information over time. Recognition accuracies similar to other HMM-free RNN-based approaches are reported for the Wall Street Journal corpus.
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, Yoshua Bengio
ICASSP4
2016 Batch normalized recurrent neural networks
abstract
Recurrent Neural Networks (RNNs) are powerful models for sequential data that have the potential to learn long-term dependencies. However, they are computationally expensive to train and difficult to parallelize. Recent work has shown that normalizing intermediate representations of neural networks can significantly improve convergence rates in feed-forward neural networks [1]. In particular, batch normalization, which uses mini-batch statistics to standardize features, was shown to significantly reduce training time. In this paper, we investigate how batch normalization can be applied to RNNs. We show for both a speech recognition task and language modeling that the way we apply batch normalization leads to a faster convergence of the training criterion but doesn't seem to improve the generalization performance.
César Laurent, Gabriel Pereyra, Philemon Brakel, Ying Zhang 0033, Yoshua Bengio
ICASSP3
2016 Deconstructing the Ladder Network Architecture
abstract
The Ladder Network is a recent new approach to semi-supervised learning that turned out to be very successful. While showing impressive performance, the Ladder Network has many components intertwined, whose contributions are not obvious in such a complex architecture. This paper presents an extensive experimental investigation of variants of the Ladder Network in which we replaced or removed individual components to learn about their relative importance. For semi-supervised tasks, we conclude that the most important contribution is made by the lateral connections, followed by the application of noise, and the choice of what we refer to as the ‘combinator function’. As the number of labeled training examples increases, the lateral connections and the reconstruction criterion become less important, with most of the generalization improvement coming from the injection of noise in each layer. Finally, we introduce a combinator function that reduces test error rates on Permutation-Invariant MNIST to 0.57% for the supervised setting, and to 0.97% and 1.0% for semi-supervised settings with 1000 and 100 labeled examples, respectively.
Mohammad Pezeshki, Linxi Fan, Philemon Brakel, Aaron C. Courville, Yoshua Bengio
ICML3
2016 Towards End-to-End Speech Recognition with Deep Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) are effective models for reducing spectral variations and modeling spectral correlations in acoustic features for automatic speech recognition (ASR).Hybrid speech recognition systems incorporating CNNs with Hidden Markov Models/Gaussian Mixture Models (HMMs/GMMs) have achieved the state-of-the-art in various benchmarks.Meanwhile, Connectionist Temporal Classification (CTC) with Recurrent Neural Networks (RNNs), which is proposed for labeling unsegmented sequences, makes it feasible to train an 'end-to-end' speech recognition system instead of hybrid settings.However, RNNs are computationally expensive and sometimes difficult to train.In this paper, inspired by the advantages of both CNNs and the CTC approach, we propose an end-to-end speech framework for sequence labeling, by combining hierarchical CNNs with CTC directly without recurrent connections.By evaluating the approach on the TIMIT phoneme recognition task, we show that the proposed model is not only computationally efficient, but also competitive with the existing baseline systems.Moreover, we argue that CNNs have the capability to model temporal correlations with appropriate context information.
Ying Zhang 0033, Mohammad Pezeshki, Philemon Brakel, Saizheng Zhang, César Laurent, Yoshua Bengio, Aaron C. Courville
INTERSPEECH3
2016 Batch-normalized joint training for DNN-based distant speech recognition
abstract
Improving distant speech recognition is a crucial step towards flexible human-machine interfaces. Current technology, however, still exhibits a lack of robustness, especially when adverse acoustic conditions are met. Despite the significant progress made in the last years on both speech enhancement and speech recognition, one potential limitation of state-of-the-art technology lies in composing modules that are not well matched because they are not trained jointly.
Mirco Ravanelli, Philemon Brakel, Maurizio Omologo, Yoshua Bengio
SLT2
2013 Bidirectional truncated recurrent neural networks for efficient speech denoising
abstract
We propose a bidirectional truncated recurrent neural network architecture for speech denoising. Recent work showed that deep recurrent neural networks perform well at speech denoising tasks and outperform feed forward architectures [1]. However, recurrent neural networks are difficult to train and their simulation does not allow for much parallelization. Given the increasing availability of parallel computing architectures like GPUs this is disadvantageous. The architecture we propose aims to retain the positive properties of recurrent neural networks and deep learning while remaining highly parallelizable. Unlike a standard recurrent neural network, it processes information from both past and future time steps. We evaluate two variants of this architecture on the Aurora2 task for robust ASR where they show promising results. The models outperform the ETSI2 advanced front end and the SPLICE algorithm under matching noise conditions.
Philemon Brakel, Dirk Stroobandt, Benjamin Schrauwen
INTERSPEECH1
2013 Training energy-based models for time-series imputation
Philemon Brakel, Dirk Stroobandt, Benjamin Schrauwen
J. Mach. Learn. Res.1
2012 Training Restricted Boltzmann Machines with Multi-tempering: Harnessing Parallelization
Philemon Brakel, Sander Dieleman, Benjamin Schrauwen
ICANN (2)1
2012 Energy-Based Temporal Neural Networks for Imputing Missing Values
Philemon Brakel, Benjamin Schrauwen
ICONIP (2)1
2012 Oger: modular learning architectures for large-scale sequential processing
David Verstraeten, Benjamin Schrauwen, Sander Dieleman, Philemon Brakel, Pieter Buteneers, Dejan Pecevski
J. Mach. Learn. Res.4