VLDB 2026 Research / reviewers in the wild / expert
Damien Ernst
dblp:46/1769
· DBLP profile ↗
40ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-3035-8260ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Computer networks · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Theoretical Justification for Asymmetric Actor-Critic AlgorithmsabstractIn reinforcement learning for partially observable environments, many successful algorithms have been developed within the asymmetric learning paradigm. This paradigm leverages additional state information available at training time for faster learning. Although the proposed learning objectives are usually theoretically sound, these methods still lack a precise theoretical justification for their potential benefits. We propose such a justification for asymmetric actor-critic algorithms with linear function approximators by adapting a finite-time convergence analysis to this setting. The resulting finite-time bound reveals that the asymmetric critic eliminates error terms arising from aliasing in the agent state. Gaspard Lambrechts, Damien Ernst, Aditya Mahajan |
ICML | 2 |
| 2025 | Automated Class Imbalance Learning via Few-shot Bayesian Optimization with Meta-learned Deep Kernel SurrogatesabstractThe class imbalance problem is a critical challenge in real-world applications, such as fault diagnosis, intrusion detection, and fraud detection, where the data exhibit highly skewed class distributions. Traditional methods to address class imbalance, such as resampling approaches, require careful model selection and hyperparameter tuning, which are complex and time-consuming. Automated Class Imbalance Learning (AutoCIL) has recently emerged as a promising paradigm, leveraging Combined Algorithm Selection and Hyperparameter Optimization (CASH) to automate this process. However, existing methods often suffer from inefficiencies and ineffectiveness, especially under resource constraints. In this paper, we propose a novel method called AutoCILFBO – Automated Class Imbalance Learning via Few-shot Bayesian Optimization with Meta-learned Deep Kernel Surrogates. Our approach introduces few-shot Bayesian optimization with deep kernel Gaussian processes tailored for class imbalance domains. Specifically, we meta-learn a shared probabilistic deep kernel surrogate model from a collection of pre-evaluated class imbalance optimization tasks, enabling rapid adaptation to target tasks. Experimental results demonstrate that our method outperforms existing approaches across 16 tasks with statistically significant improvements in terms of efficiency and effectiveness. Shuo Wang 0005, Damien Ernst |
IJCNN | 3 |
| 2025 | Exploration of Rationale-Extraction Methods for Closed-Domain Question Answering with a New Sentence-Level Rationale Dataset
Lize Pirenne, Samy Mokeddem, Damien Ernst, Gilles Louppe |
NLDB (2) | 3 |
| 2023 | IMP-MARL: a Suite of Environments for Large-scale Infrastructure Management Planning via MARLabstractWe introduce IMP-MARL, an open-source suite of multi-agent reinforcement learning (MARL) environments for large-scale Infrastructure Management Planning (IMP), offering a platform for benchmarking the scalability of cooperative MARL methods in real-world engineering applications.In IMP, a multi-component engineering system is subject to a risk of failure due to its components' damage condition.Specifically, each agent plans inspections and repairs for a specific system component, aiming to minimise maintenance costs while cooperating to minimise system failure risk.With IMP-MARL, we release several environments including one related to offshore wind structural systems, in an effort to meet today's needs to improve management strategies to support sustainable and reliable energy systems.Supported by IMP practical engineering environments featuring up to 100 agents, we conduct a benchmark campaign, where the scalability and performance of state-of-the-art cooperative MARL methods are compared against expert-based heuristic policies. The results reveal that centralised training with decentralised execution methods scale better with the number of agents than fully centralised or decentralised RL approaches, while also outperforming expert-based heuristic policies in most IMP environments.Based on our findings, we additionally outline remaining cooperation and scalability challenges that future MARL methods should still address.Through IMP-MARL, we encourage the implementation of new environments and the further development of MARL methods. Pascal Leroy, Pablo G. Morato, Jonathan Pisane, Athanasios Kolios, Damien Ernst |
NeurIPS | 5 |
| 2023 | Distributional reinforcement learning with unconstrained monotonic neural networks
Thibaut Théate, Antoine Wehenkel, Adrien Bolland, Gilles Louppe, Damien Ernst |
Neurocomputing | 5 |
| 2023 | Warming up recurrent neural networks to maximise reachable multistability greatly improves learning
Gaspard Lambrechts, Florent De Geeter, Nicolas Vecoven, Damien Ernst, Guillaume Drion |
Neural Networks | 4 |
| 2022 | Jointly Learning Environments and Control Policies with Projected Stochastic Gradient AscentabstractWe consider the joint design and control of discrete-time stochastic dynamical systems over a finite time horizon. We formulate the problem as a multi-step optimization problem under uncertainty seeking to identify a system design and a control policy that jointly maximize the expected sum of rewards collected over the time horizon considered. The transition function, the reward function and the policy are all parametrized, assumed known and differentiable with respect to their parameters. We then introduce a deep reinforcement learning algorithm combining policy gradient methods with model-based optimization techniques to solve this problem. In essence, our algorithm iteratively approximates the gradient of the expected return via Monte-Carlo sampling and automatic differentiation and takes projected gradient ascent steps in the space of environment and policy parameters. This algorithm is referred to as Direct Environment and Policy Search (DEPS). We assess the performance of our algorithm in three environments concerned with the design and control of a mass-spring-damper system, a small-scale off-grid power system and a drone, respectively. In addition, our algorithm is benchmarked against a state-of-the-art deep reinforcement learning algorithm used to tackle joint design and control problems. We show that DEPS performs at least as well or better in all three environments, consistently yielding solutions with higher returns in fewer iterations. Finally, solutions produced by our algorithm are also compared with solutions produced by an algorithm that does not jointly optimize environment and policy parameters, highlighting the fact that higher returns can be achieved when joint optimization is performed. Adrien Bolland, Ioannis Boukas, Mathias Berger, Damien Ernst |
J. Artif. Intell. Res. | 4 |
| 2021 | An application of deep reinforcement learning to algorithmic trading
Thibaut Théate, Damien Ernst |
Expert Syst. Appl. | 2 |
| 2021 | A deep reinforcement learning framework for continuous intraday market bidding
Ioannis Boukas, Damien Ernst, Thibaut Théate, Adrien Bolland, Alexandre Huynen, Martin Buchwald, Christelle Wynants, Bertrand Cornélusse |
Mach. Learn. | 2 |
| 2020 | On Overfitting and Asymptotic Bias in Batch Reinforcement Learning with Partial Observability (Extended Abstract)abstractWhen an agent has limited information on its environment, the suboptimality of an RL algorithm can be decomposed into the sum of two terms: a term related to an asymptotic bias (suboptimality with unlimited data) and a term due to overfitting (additional suboptimality due to limited data). In the context of reinforcement learning with partial observability, this paper provides an analysis of the tradeoff between these two error sources. In particular, our theoretical analysis formally characterizes how a smaller state representation increases the asymptotic bias while decreasing the risk of overfitting. Vincent François-Lavet, Guillaume Rabusseau, Joelle Pineau, Damien Ernst, Raphaël Fonteneau |
IJCAI | 4 |
| 2019 | On Overfitting and Asymptotic Bias in Batch Reinforcement Learning with Partial ObservabilityabstractThis paper provides an analysis of the tradeoff between asymptotic bias (suboptimality with unlimited data) and overfitting (additional suboptimality due to limited data) in the context of reinforcement learning with partial observability. Our theoretical analysis formally characterizes that while potentially increasing the asymptotic bias, a smaller state representation decreases the risk of overfitting. This analysis relies on expressing the quality of a state representation by bounding $L_1$ error terms of the associated belief states. Theoretical results are empirically illustrated when the state representation is a truncated history of observations, both on synthetic POMDPs and on a large-scale POMDP in the context of smartgrids, with real-world data. Finally, similarly to known results in the fully observable setting, we also briefly discuss and empirically illustrate how using function approximators and adapting the discount factor may enhance the tradeoff between asymptotic bias and overfitting in the partially observable context. Vincent François-Lavet, Guillaume Rabusseau, Joelle Pineau, Damien Ernst, Raphaël Fonteneau |
J. Artif. Intell. Res. | 4 |
| 2017 | Approximate Bayes Optimal Policy Search using Neural Networksabstractpeer reviewed Michael Castronovo, Vincent François-Lavet, Raphaël Fonteneau, Damien Ernst, Adrien Couëtoux |
ICAART (2) | 4 |
| 2017 | An App-based Algorithmic Approach for Harvesting Local and Renewable Energy using Electric Vehiclesabstractpeer reviewed Antoine Dubois, Antoine Wehenkel, Raphaël Fonteneau, Frédéric Olivier, Damien Ernst |
ICAART (1) | 5 |
| 2016 | Decision Making from Confidence Measurement on the Reward Growth using Supervised Learning - A Study Intended for Large-scale Video Gamesabstractpeer reviewed D. Taralla, Z. Qiu, Antonio Sutera, Raphaël Fonteneau, Damien Ernst |
ICAART (2) | 5 |
| 2014 | Using approximate dynamic programming for estimating the revenues of a hydrogen-based high-capacity storage deviceabstractThis paper proposes a methodology to estimate the maximum revenue that can be generated by a company that operates a high-capacity storage device to buy or sell electricity on the day-ahead electricity market. The methodology exploits the Dynamic Programming (DP) principle and is specified for hydrogen-based storage devices that use electrolysis to produce hydrogen and fuel cells to generate electricity from hydrogen. Experimental results are generated using historical data of energy prices on the Belgian market. They show how the storage capacity and other parameters of the storage device influence the optimal revenue. The main conclusion drawn from the experiments is that it may be advisable to invest in large storage tanks to exploit the inter-seasonal price fluctuations of electricity. Vincent François-Lavet, Raphaël Fonteneau, Damien Ernst |
ADPRL | 3 |
| 2013 | Optimized look-ahead trees: Extensions to large and continuous action spacesabstractThis paper studies look-ahead tree based control policies from the viewpoint of online decision making with constraints on the computational budget allowed per decision (expressed as number of calls to the generative model). We consider optimized look-ahead tree (OLT) policies, a recently introduced family of hybrid techniques, which combine the advantages of look-ahead trees (high precision) with the advantages of direct policy search (low online cost) and which are specifically designed for limited online budgets. We present two extensions of the basic OLT algorithm that on the one side allow tackling deterministic optimal control problems with large and continuous action spaces and that on the other side can also help to further reduce the online complexity. Damien Ernst, Francis Maes |
ADPRL | 2 |
| 2013 | Outbound SPIT filter with optimal performance guarantees
Sylvain Martin, Mohamed Nassar 0001, Damien Ernst, Guy Leduc |
Comput. Networks | 4 |
| 2013 | Scenario Trees and Policy Selection for Multistage Stochastic Programming Using Machine LearningabstractIn the context of multistage stochastic optimization problems, we propose a hybrid strategy for generalizing to nonlinear decision rules, using machine learning, a finite data set of constrained vector-valued recourse decisions optimized using scenario-tree techniques from multistage stochastic programming. The decision rules are based on a statistical model inferred from a given scenario-tree solution and are selected by out-of-sample simulation given the true problem. Because the learned rules depend on the given scenario tree, we repeat the procedure for a large number of randomly generated scenario trees and then select the best solution (policy) found for the true problem. The scheme leads to an ex post selection of the scenario tree itself. Numerical tests evaluate the dependence of the approach on the machine learning aspects and show cases where one can obtain near-optimal solutions, starting with a “weak” scenario-tree generator that randomizes the branching structure of the trees. Boris Defourny, Damien Ernst, Louis Wehenkel |
INFORMS J. Comput. | 2 |
| 2013 | Optimal discovery with probabilistic expert advice: finite time analysis and macroscopic optimality
Sébastien Bubeck, Damien Ernst, Aurélien Garivier |
J. Mach. Learn. Res. | 2 |
| 2013 | Monte Carlo Search Algorithm Discovery for Single-Player GamesabstractMuch current research in AI and games is being devoted to Monte Carlo search (MCS) algorithms. While the quest for a single unified MCS algorithm that would perform well on all problems is of major interest for AI, practitioners often know in advance the problem they want to solve, and spend plenty of time exploiting this knowledge to customize their MCS algorithm in a problem-driven way. We propose an MCS algorithm discovery scheme to perform this in an automatic and reproducible way. First, we introduce a grammar over MCS algorithms that enables inducing a rich space of candidate algorithms. Afterwards, we search in this space for the algorithm that performs best on average for a given distribution of training problems. We rely on multiarmed bandits to approximately solve this optimization problem. The experiments, generated on three different domains, show that our approach enables discovering algorithms that outperform several well-known MCS algorithms such as upper confidence bounds applied to trees and nested Monte Carlo search. We also show that the discovered algorithms are generally quite robust with respect to changes in the distribution over the training problems. Francis Maes, David Lupien St-Pierre, Damien Ernst |
IEEE Trans. Comput. Intell. AI Games | 3 |
| 2012 | Policy Search in a Space of Simple Closed-form Formulas: Towards Interpretability of Reinforcement Learning
Francis Maes, Raphaël Fonteneau, Louis Wehenkel, Damien Ernst |
Discovery Science | 4 |
| 2012 | Learning to Play K-armed Bandit Problems
Francis Maes, Louis Wehenkel, Damien Ernst |
ICAART (1) | 3 |
| 2012 | Contextual multi-armed bandits for web server defenseabstractIn this paper we argue that contextual multi-armed bandit algorithms could open avenues for designing self-learning security modules for computer networks and related tasks. The paper has two contributions: a conceptual and an algorithmical one. The conceptual contribution is to formulate the real-world problem of preventing HTTP-based attacks on web servers as a one-shot sequential learning problem, namely as a contextual multi-armed bandit. Our second contribution is to present CMABFAS, a new and computationally very cheap algorithm for general contextual multi-armed bandit learning that specifically targets domains with finite actions. We illustrate how CMABFAS could be used to design a fully self-learning meta filter for web servers that does not rely on feedback from the end-user (i.e., does not require labeled data) and report first convincing simulation results. Sylvain Martin, Damien Ernst, Guy Leduc |
IJCNN | 3 |
| 2011 | Approximate reinforcement learning: An overviewabstractReinforcement learning (RL) allows agents to learn how to optimally interact with complex environments. Fueled by recent advances in approximation-based algorithms, RL has obtained impressive successes in robotics, artificial intelligence, control, operations research, etc. However, the scarcity of survey papers about approximate RL makes it difficult for newcomers to grasp this intricate field. With the present overview, we take a step toward alleviating this situation. We review methods for approximate RL, starting from their dynamic programming roots and organizing them into three major classes: approximate value iteration, policy iteration, and policy search. Each class is subdivided into representative categories, highlighting among others offline and online algorithms, policy gradient methods, and simulation-based techniques. We also compare the different categories of methods, and outline possible ways to enhance the reviewed algorithms. Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska |
ADPRL | 2 |
| 2011 | Active exploration by searching for experiments that falsify the computed control policyabstractWe propose a strategy for experiment selection - in the context of reinforcement learning - based on the idea that the most interesting experiments to carry out at some stage are those that are the most liable to falsify the current hypothesis about the optimal control policy. We cast this idea in a context where a policy learning algorithm and a model identification method are given a priori. Experiments are selected if, using the learnt environment model, they are predicted to yield a revision of the learnt control policy. Algorithms and simulation results are provided for a deterministic system with discrete action space. They show that the proposed approach is promising. Raphaël Fonteneau, Susan A. Murphy, Louis Wehenkel, Damien Ernst |
ADPRL | 4 |
| 2011 | Optimal Sample Selection for Batch-mode Reinforcement Learning
Emmanuel Rachelson, François Schnitzler, Louis Wehenkel, Damien Ernst |
ICAART (1) | 4 |
| 2011 | Cross-Entropy Optimization of Control Policies With Adaptive Basis FunctionsabstractThis paper introduces an algorithm for direct search of control policies in continuous-state discrete-action Markov decision processes. The algorithm looks for the best closed-loop policy that can be represented using a given number of basis functions (BFs), where a discrete action is assigned to each BF. The type of the BFs and their number are specified in advance and determine the complexity of the representation. Considerable flexibility is achieved by optimizing the locations and shapes of the BFs, together with the action assignments. The optimization is carried out with the cross-entropy method and evaluates the policies by their empirical return from a representative set of initial states. The return for each representative state is estimated using Monte Carlo simulations. The resulting algorithm for cross-entropy policy search with adaptive BFs is extensively evaluated in problems with two to six state variables, for which it reliably obtains good policies with only a small number of BFs. In these experiments, cross-entropy policy search requires vastly fewer BFs than value-function techniques with equidistant BFs, and outperforms policy search with a competing optimization algorithm called DIRECT. Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2010 | A Cautious Approach to Generalization in Reinforcement Learning
Raphaël Fonteneau, Susan A. Murphy, Louis Wehenkel, Damien Ernst |
ICAART (1) | 4 |
| 2010 | Upper Confidence Bound Based Decision Making Strategies and Dynamic Spectrum AccessabstractIn this paper, we consider the problem of exploiting spectrum resources for a secondary user (SU) of a wireless communication network. We suggest that Upper Confidence Bound (UCB) algorithms could be useful to design decision making strategies for SUs to exploit intelligently the spectrum resources based on their past observations. The algorithms use an index that provides an optimistic estimation of the availability of the resources to the SU. The suggestion is supported by some experimental results carried out on a specific dynamic spectrum access (DSA) framework. Wassim Jouini, Damien Ernst, Christophe Moy, Jacques Palicot |
ICC | 2 |
| 2009 | Policy search with cross-entropy optimization of basis functionsabstractThis paper introduces a novel algorithm for approximate policy search in continuous-state, discrete-action Markov decision processes (MDPs). Previous policy search approaches have typically used ad-hoc parameterizations developed for specific MDPs. In contrast, the novel algorithm employs a flexible policy parameterization, suitable for solving general discrete-action MDPs. The algorithm looks for the best closed-loop policy that can be represented using a given number of basis functions, where a discrete action is assigned to each basis function. The locations and shapes of the basis functions are optimized, together with the action assignments. This allows a large class of policies to be represented. The optimization is carried out with the cross-entropy method and evaluates the policies by their empirical return from a representative set of initial states. We report simulation experiments in which the algorithm reliably obtains good policies with only a small number of basis functions, albeit at sizable computational costs. Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska |
ADPRL | 2 |
| 2009 | Planning under uncertainty, ensembles of disturbance trees and kernelized discrete action spacesabstractOptimizing decisions on an ensemble of incomplete disturbance trees and aggregating their first stage decisions has been shown as a promising approach to (model-based) planning under uncertainty in large continuous action spaces and in small discrete ones. The present paper extends this approach and deals with large but highly structured action spaces, through a kernel-based aggregation scheme. The technique is applied to a test problem with a discrete action space of 6561 elements adapted from the NIPS 2005 SensorNetwork benchmark. Boris Defourny, Damien Ernst, Louis Wehenkel |
ADPRL | 2 |
| 2009 | Inferring bounds on the performance of a control policy from a sample of trajectoriesabstractWe propose an approach for inferring bounds on the finite-horizon return of a control policy from an off-policy sample of trajectories collecting state transitions, rewards, and control actions. In this paper, the dynamics, control policy, and reward function are supposed to be deterministic and Lipschitz continuous. Under these assumptions, a polynomial algorithm, in terms of the sample size and length of the optimization horizon, is derived to compute these bounds, and their tightness is characterized in terms of the sample density. Raphaël Fonteneau, Susan A. Murphy, Louis Wehenkel, Damien Ernst |
ADPRL | 4 |
| 2009 | Reinforcement Learning Versus Model Predictive Control: A Comparison on a Power System ProblemabstractThis paper compares reinforcement learning (RL) with model predictive control (MPC) in a unified framework and reports experimental results of their application to the synthesis of a controller for a nonlinear and deterministic electrical power oscillations damping problem. Both families of methods are based on the formulation of the control problem as a discrete-time optimal control problem. The considered MPC approach exploits an analytical model of the system dynamics and cost function and computes open-loop policies by applying an interior-point solver to a minimization problem in which the system dynamics are represented by equality constraints. The considered RL approach infers in a model-free way closed-loop policies from a set of system trajectories and instantaneous cost values by solving a sequence of batch-mode supervised learning problems. The results obtained provide insight into the pros and cons of the two approaches and show that RL may certainly be competitive with MPC even in contexts where a good deterministic system model is available. Damien Ernst, Mevludin Glavic, Florin Capitanescu, Louis Wehenkel |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2008 | Consistency of fuzzy model-based reinforcement learningabstractReinforcement learning (RL) is a widely used paradigm for learning control. Computing exact RL solutions is generally only possible when process states and control actions take values in a small discrete set. In practice, approximate algorithms are necessary. In this paper, we propose an approximate, model-based Q-iteration algorithm that relies on a fuzzy partition of the state space, and on a discretization of the action space. Using assumptions on the continuity of the dynamics and of the reward function, we show that the resulting algorithm is consistent, i.e., that the optimal solution is obtained asymptotically as the approximation accuracy increases. An experimental study indicates that a continuous reward function is also important for a predictable improvement in performance as the approximation accuracy increases. Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska |
FUZZ-IEEE | 2 |
| 2007 | Fuzzy Approximation for Convergent Model-Based Reinforcement LearningabstractReinforcement learning (RL) is a learning control paradigm that provides well-understood algorithms with good convergence and consistency properties. Unfortunately, these algorithms require that process states and control actions take only discrete values. Approximate solutions using fuzzy representations have been proposed in the literature for the case when the states and possibly the actions are continuous. However, the link between these mainly heuristic solutions and the larger body of work on approximate RL, including convergence results, has not been made explicit. In this paper, we propose a fuzzy approximation structure for the Q-value iteration algorithm, and show that the resulting algorithm is convergent. The proof is based on an extension of previous results in approximate RL. We then propose a modified, serial version of the algorithm that is guaranteed to converge at least as fast as the original algorithm. An illustrative simulation example is also provided. Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska |
FUZZ-IEEE | 2 |
| 2007 | Estimation of rotor angles of synchronous machines using artificial neural networks and local PMU-based quantities
Alberto Del Angel, Pierre Geurts, Damien Ernst, Mevludin Glavic, Louis Wehenkel |
Neurocomputing | 3 |
| 2006 | Extremely randomized trees
Pierre Geurts, Damien Ernst, Louis Wehenkel |
Mach. Learn. | 2 |
| 2005 | Tree-Based Batch Mode Reinforcement LearningabstractReinforcement learning aims to determine an optimal control policy from interaction with a system or from observations gathered from a system. In batch mode, it can be achieved by approximating the so-called Q-function based on a set of four-tuples (xt, ut , rt, xt+1) where xt denotes the system state at time t, ut the control action taken, rt the instantaneous reward obtained and xt+1 the successor state of the system, and by determining the control policy from this Q-function. The Q-function approximation may be obtained from the limit of a sequence of (batch mode) supervised learning problems. Within this framework we describe the use of several classical tree-based supervised learning methods (CART, Kd-tree, tree bagging) and two newly proposed ensemble algorithms, namely extremely and totally randomized trees. We study their performances on several examples and find that the ensemble methods based on regression trees perform well in extracting relevant information about the optimal control policy from sets of four-tuples. In particular, the totally randomized trees give good results while ensuring the convergence of the sequence, whereas by relaxing the convergence constraint even better accuracy results are provided by the extremely randomized trees. Damien Ernst, Pierre Geurts, Louis Wehenkel |
J. Mach. Learn. Res. | 1 |
| 2003 | Iteratively Extending Time Horizon Reinforcement Learning
Damien Ernst, Pierre Geurts, Louis Wehenkel |
ECML | 1 |
| 2000 | Application of Reinforcement Learning to Electrical Power System Closed-Loop Emergency Control
Christophe Druet, Damien Ernst, Louis Wehenkel |
PKDD | 2 |