EDBT 2026 Demo / reviewers in the wild / expert
Lucian Busoniu
dblp:65/1615
· DBLP profile ↗
23ranked-venue papers
12as first author
3since 2021 · last 2026
0000-0001-8017-1296ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 10 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-authorSystems, architecture and hardware · 3Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 50% Deep learning architectures and training · 50% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
deep reinforcement learning |
0.5 | 1 | 2021 | Spectral Normalisation for Deep Reinforcement Learning: An Optimisation Perspective · ICML 2021 |
Machine learning › Deep learning architectures and training › normalization
spectral normalization |
0.5 | 1 | 2021 | Spectral Normalisation for Deep Reinforcement Learning: An Optimisation Perspective · ICML 2021 |
Methods — techniques the papers use, named apart from their topics
spectral normalization · 0.5lipschitz constraint · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The SeaClear system: An intelligent multi-robot solution for autonomous cleanup of marine debris on the seabedabstractMarine debris poses an alarming threat to ocean environments. Conventional methods of sea and ocean cleaning rely heavily on manual collection, a process that has repeatedly demonstrated its inefficiency and extensive demand for resources. This paper presents the SeaClear system, a novel multi-robot platform designed to autonomously detect and collect marine debris, thereby offering a more efficient solution to this environmental challenge. An overview of the system is presented, followed by a detailed description of each robot’s capabilities. Leveraging artificial intelligence, the system employs the deep-learning-based computer vision algorithm You Only Look Once (YOLO) for the detection of underwater litter, addressing the challenges of poor visibility and hydrodynamic disturbances of underwater environments. Additionally, the paper explores the implemented navigation and control methodologies, which are an essential part of the workflow of the system. The performance of the designed system is validated via field tests conducted in a real-world underwater environment. Finally, directions for future work are proposed. • The infrastructure of a multi-robot system for autonomous underwater debris detection, mapping, and collection is presented. • The suitability of YOLO-based deep learning for real-time debris detection in shallow waters is validated. • The sensing and control scheme that enables the operation of the multi-robotic platform is presented. • The practical performance of the system is confirmed with field experiments to validate the feasibility of AI-driven underwater operations. Athina Ilioudi, Stefan Sosnowski, Elisabeth Banken, Petar Bevanda, Jan Brüdigam, Lucian Busoniu, Yves Chardard, Cosmin Delea, Bart De Schutter, Antun Duras, Claudia Hertel-ten Eikelder, Shahab Heshmati-Alamdari, Vicu-Mihalis Maer, Ivana Palunko, Iva Pozniak, Vicko Prkacin, Domagoj Tolic |
Eng. Appl. Artif. Intell. | 6 |
| 2022 | TS Fuzzy Observer-Based Controller Design for a Class of Discrete-Time Nonlinear SystemsabstractThis article presents an observer-based control design approach for a class of nonlinear discrete-time systems. The model nonlinearities are handled in two ways: 1) A Takagi–Sugeno fuzzy representation is used for nonlinearities that depend on measured states, and 2) nonlinearities that depend on unmeasured states are kept in their original form and handled using a slope-bound condition. The observer-based controller design conditions are given as linear matrix inequalities. The approach we propose significantly improves results in the literature by providing less restrictive design conditions. These improvements are illustrated in a detailed analytical and numerical comparison on a synthetic example; while a pendulum-on-a-cart example shows that the approach works both in simulation and in real-time experiments. Zoltán Nagy 0006, Zsófia Lendek, Lucian Busoniu |
IEEE Trans. Fuzzy Syst. | 3 |
| 2021 | Spectral Normalisation for Deep Reinforcement Learning: An Optimisation PerspectiveabstractMost of the recent deep reinforcement learning advances take an RL-centric perspective and focus on refinements of the training objective. We diverge from this view and show we can recover the performance of these developments not by changing the objective, but by regularising the value-function estimator. Constraining the Lipschitz constant of a single layer using spectral normalisation is sufficient to elevate the performance of a Categorical-DQN agent to that of a more elaborated agent on the challenging Atari domain. We conduct ablation studies to disentangle the various effects normalisation has on the learning dynamics and show that is sufficient to modulate the parameter updates to recover most of the performance of spectral normalisation. These findings hint towards the need to also focus on the neural component and its learning dynamics to tackle the peculiarities of Deep Reinforcement Learning. Florin Gogianu, Tudor Berariu, Mihaela Rosca, Claudia Clopath, Lucian Busoniu, Razvan Pascanu |
ICML | 5 |
| 2019 | Hardware and control design of a ball balancing robotabstractThis paper presents the construction of a new ball balancing robot (ballbot), together with the design of a controller to balance it vertically around a given position in the plane. Requirements on physical size and agility lead to the choice of ball, motors, gears, omnidirectional wheels, and body frame. The electronic hardware architecture is presented in detail, together with timing results showing that real-time control can be achieved. Finally, we design a linear quadratic regulator for balancing, starting from a 2D model of the robot. Experimental balancing results are satisfactory, maintaining the robot in a disc 0.3 m in diameter. Ioana Lal, Marius Nicoara, Alexandru Codrean, Lucian Busoniu |
DDECS | 4 |
| 2018 | Optimistic planning with an adaptive number of action switches for near-optimal nonlinear control
Koppány Máthé, Lucian Busoniu, Rémi Munos, Bart De Schutter |
Eng. Appl. Artif. Intell. | 2 |
| 2016 | Analysis and a home assistance application of online AEMS2 planningabstractWe consider an online planning algorithm for partially observable Markov decision processes (POMDPs), called Anytime Error Minimization Search 2 (AEMS2). Despite the considerable success it has enjoyed in robotics and other problems, no quantitative analysis exists of the relationship between its near-optimality and the computation invested. Exploiting ideas from fully-observable MDP planning, we provide here such an analysis, in which the relationship is modulated via a measure of problem complexity called near-optimality exponent. We illustrate the exponent for some interesting POMDP structures, and examine the role of the informative heuristics used by AEMS2 in the guarantees. In the second part of the paper, we introduce a domestic assistance problem in which a robot monitors partially observable switches and turns them off if needed. AEMS2 successfully solves this task in real experiments, and also works better than several state of the art planners in simulation comparisons. Elod Páll, Levente Tamas, Lucian Busoniu |
IROS | 3 |
| 2016 | Online learning for optimistic planning
Lucian Busoniu, Alexander Daniels, Robert Babuska |
Eng. Appl. Artif. Intell. | 1 |
| 2014 | An analysis of optimistic, best-first search for minimax sequential decision makingabstractWe consider problems in which a maximizer and a minimizer agent take actions in turn, such as games or optimal control with uncertainty modeled as an opponent. We extend the ideas of optimistic optimization to this setting, obtaining a search algorithm that has been previously considered as the best-first search variant of the B* method. We provide a novel analysis of the algorithm relying on a certain structure for the values of action sequences, under which earlier actions are more important than later ones. An asymptotic branching factor is defined as a measure of problem complexity, and it is used to characterize the relationship between computation invested and near-optimality. In particular, when action importance decreases exponentially, convergence rates are obtained. Throughout, examples illustrate analytical concepts such as the branching factor. In an empirical study, we compare the optimistic best-first algorithm with two classical game tree search methods, and apply it to a challenging HIV infection control problem. Lucian Busoniu, Rémi Munos, Elod Páll |
ADPRL | 1 |
| 2013 | Optimistic planning for continuous-action deterministic systemsabstractWe consider the class of online planning algorithms for optimal control, which compared to dynamic programming are relatively unaffected by large state dimensionality. We introduce a novel planning algorithm called SOOP that works for deterministic systems with continuous states and actions. SOOP is the first method to explore the true solution space, consisting of infinite sequences of continuous actions, without requiring knowledge about the smoothness of the system. SOOP can be used parameter-free at the cost of more model calls, but we also propose a more practical variant tuned by a parameter α, which balances finer discretization with longer planning horizons. Experiments on three problems show SOOP reliably ranks among the best algorithms, fully dominating competing methods when the problem requires both long horizons and fine discretization. Lucian Busoniu, Alexander Daniels, Rémi Munos, Robert Babuska |
ADPRL | 1 |
| 2013 | Optimistic planning for belief-augmented Markov Decision ProcessesabstractThis paper presents the Bayesian Optimistic Planning (BOP) algorithm, a novel model-based Bayesian reinforcement learning approach. BOP extends the planning approach of the Optimistic Planning for Markov Decision Processes (OP-MDP) algorithm [10], [9] to contexts where the transition model of the MDP is initially unknown and progressively learned through interactions within the environment. The knowledge about the unknown MDP is represented with a probability distribution over all possible transition models using Dirichlet distributions, and the BOP algorithm plans in the belief-augmented state space constructed by concatenating the original state vector with the current posterior distribution over transition models. We show that BOP becomes Bayesian optimal when the budget parameter increases to infinity. Preliminary empirical validations show promising performance. Raphaël Fonteneau, Lucian Busoniu, Rémi Munos |
ADPRL | 2 |
| 2012 | Experience Replay for Real-Time Reinforcement Learning ControlabstractReinforcement-learning (RL) algorithms can automatically learn optimal control strategies for nonlinear, possibly stochastic systems. A promising approach for RL control is experience replay (ER), which learns quickly from a limited amount of data, by repeatedly presenting these data to an underlying RL algorithm. Despite its benefits, ER RL has been studied only sporadically in the literature, and its applications have largely been confined to simulated systems. Therefore, in this paper, we evaluate ER RL on real-time control experiments that involve a pendulum swing-up problem and the vision-based control of a goalkeeper robot. These real-time experiments are complemented by simulation studies and comparisons with traditional RL. As a preliminary, we develop a general ER framework that can be combined with essentially any incremental RL technique, and instantiate this framework for the approximate Q-learning and SARSA algorithms. The successful real-time learning results that are presented here are highly encouraging for the applicability of ER RL in practice. Sander Adam, Lucian Busoniu, Robert Babuska |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2012 | A Survey of Actor-Critic Reinforcement Learning: Standard and Natural Policy GradientsabstractPolicy-gradient-based actor-critic algorithms are amongst the most popular algorithms in the reinforcement learning framework. Their advantage of being able to search for optimal policies using low-variance gradient estimates has made them useful in several real-life applications, such as robotics, power control, and finance. Although general surveys on reinforcement learning techniques already exist, no survey is specifically dedicated to actor-critic algorithms in particular. This paper, therefore, describes the state of the art of actor-critic algorithms, with a focus on methods that can work in an online setting and use function approximation in order to deal with continuous state and action spaces. After starting with a discussion on the concepts of reinforcement learning and the origins of actor-critic algorithms, this paper describes the workings of the natural gradient, which has made its way into many actor-critic algorithms over the past few years. A review of several standard and natural actor-critic algorithms is given, and the paper concludes with an overview of application areas and a discussion on open issues. Ivo Grondman, Lucian Busoniu, Gabriel A. D. Lopes, Robert Babuska |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2012 | Efficient Model Learning Methods for Actor-Critic ControlabstractWe propose two new actor-critic algorithms for reinforcement learning. Both algorithms use local linear regression (LLR) to learn approximations of the functions involved. A crucial feature of the algorithms is that they also learn a process model, and this, in combination with LLR, provides an efficient policy update for faster learning. The first algorithm uses a novel model-based update rule for the actor parameters. The second algorithm does not use an explicit actor but learns a reference model which represents a desired behavior, from which desired control actions can be calculated using the inverse of the learned process model. The two novel methods and a standard actor-critic algorithm are applied to the pendulum swing-up problem, in which the novel methods achieve faster learning than the standard algorithm. Ivo Grondman, Maarten Vaandrager, Lucian Busoniu, Robert Babuska, Erik Schuitema |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2011 | Approximate reinforcement learning: An overviewabstractReinforcement learning (RL) allows agents to learn how to optimally interact with complex environments. Fueled by recent advances in approximation-based algorithms, RL has obtained impressive successes in robotics, artificial intelligence, control, operations research, etc. However, the scarcity of survey papers about approximate RL makes it difficult for newcomers to grasp this intricate field. With the present overview, we take a step toward alleviating this situation. We review methods for approximate RL, starting from their dynamic programming roots and organizing them into three major classes: approximate value iteration, policy iteration, and policy search. Each class is subdivided into representative categories, highlighting among others offline and online algorithms, policy gradient methods, and simulation-based techniques. We also compare the different categories of methods, and outline possible ways to enhance the reviewed algorithms. Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska |
ADPRL | 1 |
| 2011 | Optimistic planning for sparsely stochastic systemsabstractWe propose an online planning algorithm for finite-action, sparsely stochastic Markov decision processes, in which the random state transitions can only end up in a small number of possible next states. The algorithm builds a planning tree by iteratively expanding states, where each expansion exploits sparsity to add all possible successor states. Each state to expand is actively chosen to improve the knowledge about action quality, and this allows the algorithm to return a good action after a strictly limited number of expansions. More specifically, the active selection method is optimistic in that it chooses the most promising states first, so the novel algorithm is called optimistic planning for sparsely stochastic systems. We note that the new algorithm can also be seen as model-predictive (receding-horizon) control. The algorithm obtains promising numerical results, including the successful online control of a simulated HIV infection with stochastic drug effectiveness. Lucian Busoniu, Rémi Munos, Bart De Schutter, Robert Babuska |
ADPRL | 1 |
| 2011 | Cross-Entropy Optimization of Control Policies With Adaptive Basis FunctionsabstractThis paper introduces an algorithm for direct search of control policies in continuous-state discrete-action Markov decision processes. The algorithm looks for the best closed-loop policy that can be represented using a given number of basis functions (BFs), where a discrete action is assigned to each BF. The type of the BFs and their number are specified in advance and determine the complexity of the representation. Considerable flexibility is achieved by optimizing the locations and shapes of the BFs, together with the action assignments. The optimization is carried out with the cross-entropy method and evaluates the policies by their empirical return from a representative set of initial states. The return for each representative state is estimated using Monte Carlo simulations. The resulting algorithm for cross-entropy policy search with adaptive BFs is extensively evaluated in problems with two to six state variables, for which it reliably obtains good policies with only a small number of BFs. In these experiments, cross-entropy policy search requires vastly fewer BFs than value-function techniques with equidistant BFs, and outperforms policy search with a competing optimization algorithm called DIRECT. Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2010 | Control delay in Reinforcement Learning for real-time dynamic systems: A memoryless approachabstractRobots controlled by Reinforcement Learning (RL) are still rare. A core challenge to the application of RL to robotic systems is to learn despite the existence of control delay - the delay between measuring a system's state and acting upon it. Control delay is always present in real systems. In this work, we present two novel temporal difference (TD) learning algorithms for problems with control delay. These algorithms improve learning performance by taking the control delay into account. We test our algorithms in a gridworld, where the delay is an integer multiple of the time step, as well as in the simulation of a robotic system, where the delay can have any value. In both tests, our proposed algorithms outperform classical TD learning algorithms, while maintaining low computational complexity. Erik Schuitema, Lucian Busoniu, Robert Babuska, Pieter P. Jonker |
IROS | 2 |
| 2009 | Policy search with cross-entropy optimization of basis functionsabstractThis paper introduces a novel algorithm for approximate policy search in continuous-state, discrete-action Markov decision processes (MDPs). Previous policy search approaches have typically used ad-hoc parameterizations developed for specific MDPs. In contrast, the novel algorithm employs a flexible policy parameterization, suitable for solving general discrete-action MDPs. The algorithm looks for the best closed-loop policy that can be represented using a given number of basis functions, where a discrete action is assigned to each basis function. The locations and shapes of the basis functions are optimized, together with the action assignments. This allows a large class of policies to be represented. The optimization is carried out with the cross-entropy method and evaluates the policies by their empirical return from a representative set of initial states. We report simulation experiments in which the algorithm reliably obtains good policies with only a small number of basis functions, albeit at sizable computational costs. Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska |
ADPRL | 1 |
| 2008 | Consistency of fuzzy model-based reinforcement learningabstractReinforcement learning (RL) is a widely used paradigm for learning control. Computing exact RL solutions is generally only possible when process states and control actions take values in a small discrete set. In practice, approximate algorithms are necessary. In this paper, we propose an approximate, model-based Q-iteration algorithm that relies on a fuzzy partition of the state space, and on a discretization of the action space. Using assumptions on the continuity of the dynamics and of the reward function, we show that the resulting algorithm is consistent, i.e., that the optimal solution is obtained asymptotically as the approximation accuracy increases. An experimental study indicates that a continuous reward function is also important for a predictable improvement in performance as the approximation accuracy increases. Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska |
FUZZ-IEEE | 1 |
| 2008 | A Comprehensive Survey of Multiagent Reinforcement LearningabstractMultiagent systems are rapidly finding applications in a variety of domains, including robotics, distributed control, telecommunications, and economics. The complexity of many tasks arising in these domains makes them difficult to solve with preprogrammed agent behaviors. The agents must, instead, discover a solution on their own, using learning. A significant part of the research on multiagent learning concerns reinforcement learning techniques. This paper provides a comprehensive survey of multiagent reinforcement learning (MARL). A central issue in the field is the formal statement of the multiagent learning goal. Different viewpoints on this issue have led to the proposal of many different goals, among which two focal points can be distinguished: stability of the agents' learning dynamics, and adaptation to the changing behavior of the other agents. The MARL algorithms described in the literature aim---either explicitly or implicitly---at one of these two goals or at a combination of both, in a fully cooperative, fully competitive, or more general setting. A representative selection of these algorithms is discussed in detail in this paper, together with the specific issues that arise in each category. Additionally, the benefits and challenges of MARL are described along with some of the problem domains where the MARL techniques have been applied. Finally, an outlook for the field is provided. Lucian Busoniu, Robert Babuska, Bart De Schutter |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2007 | Fuzzy Approximation for Convergent Model-Based Reinforcement LearningabstractReinforcement learning (RL) is a learning control paradigm that provides well-understood algorithms with good convergence and consistency properties. Unfortunately, these algorithms require that process states and control actions take only discrete values. Approximate solutions using fuzzy representations have been proposed in the literature for the case when the states and possibly the actions are continuous. However, the link between these mainly heuristic solutions and the larger body of work on approximate RL, including convergence results, has not been made explicit. In this paper, we propose a fuzzy approximation structure for the Q-value iteration algorithm, and show that the resulting algorithm is convergent. The proof is based on an extension of previous results in approximate RL. We then propose a modified, serial version of the algorithm that is guaranteed to converge at least as fast as the original algorithm. An illustrative simulation example is also provided. Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska |
FUZZ-IEEE | 1 |
| 2006 | Multi-Agent Reinforcement Learning: A SurveyabstractMulti-agent systems are rapidly finding applications in a variety of domains, including robotics, distributed control, telecommunications, economics. Many tasks arising in these domains require that the agents learn behaviors online. A significant part of the research on multi-agent learning concerns reinforcement learning techniques. However, due to different viewpoints on central issues, such as the formal statement of the learning goal, a large number of different methods and approaches have been introduced. In this paper we aim to present an integrated survey of the field. First, the issue of the multi-agent learning goal is discussed, after which a representative selection of algorithms is reviewed. Finally, open issues are identified and future research directions are outlined Lucian Busoniu, Robert Babuska, Bart De Schutter |
ICARCV | 1 |
| 2006 | Decentralized Reinforcement Learning Control of a Robotic ManipulatorabstractMulti-agent systems are rapidly finding applications in a variety of domains, including robotics, distributed control, telecommunications, etc. Learning approaches to multi-agent control, many of them based on reinforcement learning (RL), are investigated in complex domains such as teams of mobile robots. However, the application of decentralized RL to low-level control tasks is not as intensively studied. In this paper, we investigate centralized and decentralized RL, emphasizing the challenges and potential advantages of the latter. These are then illustrated on an example: learning to control a two-link rigid manipulator. Some open issues and future research directions in decentralized RL are outlined Lucian Busoniu, Bart De Schutter, Robert Babuska |
ICARCV | 1 |