Robert Babuska

dblp:65/1151 · DBLP profile ↗
← Back
119ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0001-9578-8598ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 98 · 6 first-author · 11 since 2021Systems, architecture and hardware · 21 · 6 since 2021Human-computer interaction and ubiquitous computing · 16Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021
YearPublicationVenuePosition
2026 Asynchronous Neuro-Evolutionary Symbolic Regression with Maturity-Based Replacement
abstract
We consider a neuro-evolutionary symbolic regression, an approach in which mathematical formulas are internally represented by feedforward neural networks. Gradient-free global exploration of the network topology space is combined with gradient-based parameter learning. From the perspective of evolutionary algorithms, this approach has a potential weakness: newly generated networks produced by genetic operators usually require intensive parameter tuning. If this optimization has not yet converged, such networks may be prematurely discarded due to poor performance, thereby losing the opportunity to contribute their true potential to the evolutionary process. Such a waste of innovative, potentially promising topologies can reduce the algorithm's efficiency. To address this limitation, we propose a neuro-evolutionary algorithm that relies on asynchronous refinement of network parameters and a maturity-based replacement strategy. Thus, newly created networks are allowed to coexist and develop alongside their parental networks, while being compared only with peers of similar maturity. The proposed method was evaluated on six benchmarks and compared to the original evolutionary approach.
Jirí Kubalík, Nada Fucelová, Robert Babuska
GECCO3
2025 Neuro-Evolutionary Approach to Physics-Aware Symbolic Regression
abstract
Symbolic regression is a technique that can automatically derive analytic models from data. Traditionally, symbolic regression has been implemented primarily through genetic programming that evolves populations of candidate solutions sampled by genetic operators, crossover and mutation. More recently, neural networks have been employed to learn the entire analytical model, i.e., its structure and coefficients, using regularized gradient-based optimization. Although this approach tunes the model's coefficients better, it is prone to premature convergence to suboptimal model structures. Here, we propose a neuro-evolutionary symbolic regression method that combines the strengths of evolutionary-based search for optimal neural network (NN) topologies with gradient-based tuning of the network's parameters. Due to the inherent high computational demand of evolutionary algorithms, it is not feasible to learn the parameters of every candidate NN topology to the full convergence. Thus, our method employs a memory-based strategy and population perturbations to enhance exploitation and reduce the risk of being trapped in suboptimal NNs. In this way, each NN topology can be trained using only a short sequence of back-propagation iterations. The proposed method was experimentally evaluated on three real-world test problems and has been shown to outperform other NN-based approaches regarding the quality of the models obtained.
Jirí Kubalík, Robert Babuska
GECCO2
2025 Towards Transparent, Physically Consistent Machine Learning Models
Robert Babuska
ICINCO1
2025 Embedded Hierarchical MPC for Autonomous Navigation
abstract
To efficiently deploy robotic systems in society, mobile robots must move autonomously and safely through complex environments. Nonlinear model predictive control (MPC) methods provide a natural way to find a dynamically feasible trajectory through the environment without colliding with nearby obstacles. However, the limited computation power available on typical embedded robotic systems, such as quadrotors, poses a challenge to running MPC in real time, including its most expensive tasks: constraints generation and optimization. To address this problem, we propose a novel hierarchical MPC scheme that consists of a planning and a tracking layer. The planner constructs a trajectory with a long prediction horizon at a slow rate, while the tracker ensures trajectory tracking at a relatively fast rate. We prove that the proposed framework avoids collisions and is recursively feasible. Furthermore, we demonstrate its effectiveness in simulations and lab experiments with a quadrotor that needs to reach a goal position in a complex static environment. The code is efficiently implemented on the quadrotor's embedded computer to ensure real-time feasibility. Compared to a state-of-the-art single-layer MPC formulation, this allows us to increase the planning horizon by a factor of 5, which results in significantly better performance.
Dennis Benders, Johannes Köhler 0001, Thijs Niesten, Robert Babuska, Javier Alonso-Mora, Laura Ferranti
IEEE Trans. Robotics4
2024 Robotic Grasping of Harvested Tomato Trusses Using Vision and Online Learning
abstract
Currently, truss tomato weighing and packaging require significant manual work. The main obstacle to automation lies in the difficulty of developing a reliable robotic grasping system for already harvested trusses. We propose a method to grasp trusses that are stacked in a crate with considerable clutter, which is how they are commonly stored and transported after harvest. The method consists of a deep learning-based vision system to first identify the individual trusses in the crate and then determine a suitable grasping location on the stem. To this end, we have introduced a grasp pose ranking algorithm with online learning capabilities. After selecting the most promising grasp pose, the robot executes a pinch grasp without needing touch sensors or geometric models. Lab experiments with a robotic manipulator equipped with an eye-in-hand RGB-D camera showed a 100% clearance rate when tasked to pick all trusses from a pile. 93% of the trusses were successfully grasped on the first try, while the remaining 7% required more attempts.
Luuk van den Bent, Tomás Coleman, Robert Babuska
ICRA3
2022 ViewFormer: NeRF-Free Neural Rendering from Few Images Using Transformers
Jonas Kulhanek, Erik Derner, Torsten Sattler, Robert Babuska
ECCV (15)4
2022 Where to Look Next: Learning Viewpoint Recommendations for Informative Trajectory Planning
abstract
Search missions require motion planning and navigation methods for information gathering that continuously replan based on new observations of the robot's surroundings. Current methods for information gathering, such as Monte Carlo Tree Search, are capable of reasoning over long horizons, but they are computationally expensive. An alternative for fast online execution is to train, offline, an information gathering policy, which indirectly reasons about the information value of new observations. However, these policies lack safety guarantees and do not account for the robot dynamics. To overcome these limitations we train an information-aware policy via deep reinforcement learning, that guides a receding-horizon trajectory optimization planner. In particular, the policy continuously recommends a reference viewpoint to the local planner, such that the resulting dynamically feasible and collision-free trajectories lead to observations that maximize the information gain and reduce the uncertainty about the environment. In simulation tests in previously unseen environments, our method consistently outperforms greedy next-best-view policies and achieves competitive performance compared to Monte Carlo Tree Search, in terms of information gains and coverage time, with a reduction in execution time by three orders of magnitude.
Max Lodel, Bruno Brito, Álvaro Serra-Gómez, Laura Ferranti, Robert Babuska, Javier Alonso-Mora
ICRA5
2022 OpenDR: An Open Toolkit for Enabling High Performance, Low Footprint Deep Learning for Robotics
abstract
Existing Deep Learning (DL) frameworks typically do not provide ready-to-use solutions for robotics, where very specific learning, reasoning, and embodiment problems exist. Their relatively steep learning curve and the different methodologies employed by DL compared to traditional approaches, along with the high complexity of DL models, which often leads to the need of employing specialized hardware accelerators, further increase the effort and cost needed to employ DL models in robotics. Also, most of the existing DL methods follow a static inference paradigm, as inherited by the traditional computer vision pipelines, ignoring active perception, which can be employed to actively interact with the environment in order to increase perception accuracy. In this paper, we present the Open Deep Learning Toolkit for Robotics (OpenDR). OpenDR aims at developing an open, non-proprietary, efficient, and modular toolkit that can be easily used by robotics companies and research institutions to efficiently develop and deploy AI and cognition technologies to robotics applications, providing a solid step towards addressing the aforementioned challenges. We also detail the design choices, along with an abstract interface that was created to overcome these challenges. This interface can describe various robotic tasks, spanning beyond traditional DL cognition and inference, as known by existing frameworks, incorporating openness, homogeneity and robotics-oriented perception e.g., through active perception, as its core design principles.
Nikolaos Passalis, S. Pedrazzi, Robert Babuska, Wolfram Burgard, D. Dias, F. Ferro, Moncef Gabbouj, Ole Green, Alexandros Iosifidis, Erdal Kayacan, Jens Kober, O. Michel, Nikos Nikolaidis 0001, Paraskevi Nousi, Roel Pieters, Maria Tzelepi, Abhinav Valada, Anastasios Tefas
IROS3
2021 Guiding Robot Model Construction with Prior Features
abstract
Virtually all robot control methods benefit from the availability of an accurate mathematical model of the robot. However, obtaining a sufficient amount of informative data for constructing dynamic models can be difficult, especially when the models are to be learned during robot deployment. Under such circumstances, standard data-driven model learning techniques often yield models that do not comply with the physics of the robot. We extend a symbolic regression algorithm based on Single Node Genetic Programming by including the prior model information into the model construction process. In this way, symbolic regression automatically builds models that compensate for theoretical or empirical model deficiencies. We experimentally demonstrate the approach on two real-world systems: the TurtleBot 2 mobile robot and the Parrot Bebop 2 drone. The results show that the proposed model-learning algorithm produces realistic models that fit well the training data even when using small training sets. Passing the prior model information to the algorithm significantly improves the model accuracy while speeding up the search.
Erik Derner, Jirí Kubalík, Robert Babuska
IROS3
2021 DeepKoCo: Efficient latent planning with a task-relevant Koopman representation
abstract
This paper presents DeepKoCo, a novel modelbased agent that learns a latent Koopman representation from images. This representation allows DeepKoCo to plan efficiently using linear control methods, such as linear model predictive control. Compared to traditional agents, DeepKoCo learns taskrelevant dynamics, thanks to the use of a tailored lossy autoencoder network that allows DeepKoCo to learn latent dynamics that reconstruct and predict only observed costs, rather than all observed dynamics. As our results show, DeepKoCo achieves a similar final performance as traditional model-free methods on complex control tasks, while being considerably more robust to distractor dynamics, making the proposed agent more amenable for real-life applications.
Bas van der Heijden, Laura Ferranti, Jens Kober, Robert Babuska
IROS4
2021 Inclined Quadrotor Landing using Deep Reinforcement Learning
abstract
Landing a quadrotor on an inclined surface is a challenging maneuver. The final state of any inclined landing trajectory is not an equilibrium, which precludes the use of most conventional control methods. We propose a deep reinforcement learning approach to design an autonomous landing controller for inclined surfaces. Using the proximal policy optimization (PPO) algorithm with sparse rewards and a tailored curriculum learning approach, an inclined landing policy can be trained in simulation in less than 90 minutes on a standard laptop. The policy then directly runs on a real Crazyflie 2.1 quadrotor and successfully performs real inclined landings in a flying arena. A single policy evaluation takes approximately 2.5 ms, which makes it suitable for a future embedded implementation on the quadrotor.
Jacob E. Kooi, Robert Babuska
IROS2
2021 Multi-objective symbolic regression for physics-aware dynamic modeling
Jirí Kubalík, Erik Derner, Robert Babuska
Expert Syst. Appl.3
2020 Symbolic regression driven by training data and prior knowledge
abstract
In symbolic regression, the search for analytic models is typically driven purely by the prediction error observed on the training data samples. However, when the data samples do not sufficiently cover the input space, the prediction error does not provide sufficient guidance toward desired models. Standard symbolic regression techniques then yield models that are partially incorrect, for instance, in terms of their steady-state characteristics or local behavior. If these properties were considered already during the search process, more accurate and relevant models could be produced. We propose a multi-objective symbolic regression approach that is driven by both the training data and the prior knowledge of the properties the desired model should manifest. The properties given in the form of formal constraints are internally represented by a set of discrete data samples on which candidate models are exactly checked. The proposed approach was experimentally evaluated on three test problems with results clearly demonstrating its capability to evolve realistic models that fit the training data well while complying with the prior knowledge of the desired model characteristics at the same time. It outperforms standard symbolic regression by several orders of magnitude in terms of the mean squared deviation from a reference model.
Jirí Kubalík, Erik Derner, Robert Babuska
GECCO3
2020 Simultaneous task allocation and motion scheduling for complex tasks executed by multiple robots
abstract
The coordination of multiple robots operating simultaneously in the same workspace requires the integration of task allocation and motion scheduling. We focus on tasks in which the robot's actions are not confined to small volumes, but can also occupy a large time-varying portion of the workspace, such as in welding along a line. The optimization of such tasks presents a considerable challenge mainly due to the fact that different variants of task execution exist, for instance, there can be multiple starting points of lines or closed curves, differentfilling patterns of areas, etc. We propose a generic and computationally efficient optimization method which is based on constraint programming. It takes into account the kinematics of the robots and guarantees that the motions of the robots are collision-free while minimizing the overall makespan. We evaluate our approach on several use-cases of varying complexity: cutting, additive manufacturing, spot welding, inserting and tightening bolts, performed by a dual-arm robot. In terms of the makespan, the result is superior to task execution by one robot arm as well as by two arms not working simultaneously.
Jan Kristof Behrens, Karla Stépánová, Robert Babuska
ICRA3
2020 Efficient Object Search Through Probability-Based Viewpoint Selection
abstract
The ability to search for objects is a precondition for various robotic tasks. In this paper, we address the problem of finding objects in partially known indoor environments. Using the knowledge of the floor plan and the mapped objects, we consider object-object and object-room co-occurrences as prior information for identifying promising locations where an unmapped object can be present. We propose an efficient search strategy that determines the best pose of the robot based on the analysis of the candidate locations. We optimize the probability of finding the target object and the distance travelled through a cost function.To evaluate our method, several experiments in simulated and real-world environments were performed. The results show that the robot successfully finds the target object in the environment while covering only a small portion of the search space. The real-world experiments with the TurtleBot 2 mobile robot validate the proposed approach and demonstrate that the method performs well also in real environments.
Alejandra C. Hernández, Erik Derner, Clara Gómez, Ramón Barber, Robert Babuska
IROS5
2019 Genetic programming methods for reinforcement learning
abstract
Reinforcement Learning (RL) algorithms can be used to optimally solve dynamic decision-making and control problems. With continuous-valued state and input variables, RL algorithms must rely on function approximators to represent the value function and policy mappings. Commonly used numerical approximators, such as neural networks or basis function expansions, have two main drawbacks: they are black-box models offering no insight in the mappings learnt, and they require significant trial and error tuning of their meta-parameters. In addition, results obtained with deep neural networks suffer from the lack of reproducibility. In this talk, we discuss a family of new approaches to constructing smooth approximators for RL by means of genetic programming and more specifically by symbolic regression. We show how to construct process models and value functions represented by parsimonious analytic expressions using state-of-the-art algorithms, such as Single Node Genetic Programming and Multi-Gene Genetic Programming. We will include examples of nonlinear control problems that can be successfully solved by reinforcement learning with symbolic regression and illustrate some of the challenges this exciting field of research is currently facing.
Robert Babuska
GECCO1
2019 Reinforcement learning based compensation methods for robot manipulators
Yudha P. Pane, Subramanya Nageshrao, Jens Kober, Robert Babuska
Eng. Appl. Artif. Intell.4
2018 Data-driven Construction of Symbolic Process Models for Reinforcement Learning
abstract
Reinforcement learning (RL) is a suitable approach for controlling systems with unknown or time-varying dynamics. RL in principle does not require a model of the system, but before it learns an acceptable policy, it needs many unsuccessful trials, which real robots usually cannot withstand. It is well known that RL can be sped up and made safer by using models learned online. In this paper, we propose to use symbolic regression to construct compact, parsimonious models described by analytic equations, which are suitable for realtime robot control. Single node genetic programming (SNGP) is employed as a tool to automatically search for equations fitting the available data. We demonstrate the approach on two benchmark examples: a simulated mobile robot and the pendulum swing-up problem; the latter both in simulations and real-time experiments. The results show that through this approach we can find accurate models even for small batches of training data. Based on the symbolic model found, RL can control the system well.
Erik Derner, Jirí Kubalík, Robert Babuska
ICRA3
2018 Reinforcement Learning with Symbolic Input-Output Models
abstract
It is well known that reinforcement learning (RL) can benefit from the use of a dynamic prediction model which is learned on data samples collected online from the process to be controlled. Most RL algorithms are formulated in the state-space domain and use state-space models. However, learning state-space models is difficult, mainly because in the vast majority of problems the full state cannot be measured on the system or reconstructed from the measurements. To circumvent this limitation, we propose to use input-output models of the NARX (nonlinear autoregressive with exogenous input) type. Symbolic regression is employed to construct parsimonious models and the corresponding value functions. Thanks to this approach, we can learn accurate models and compute optimal policies even from small amounts of training data. We demonstrate the approach on two simulated examples, a hopping robot and a 1-DOF robot arm, and on a real inverted pendulum system. Results show that our proposed method can reliably determine a good control policy based on a symbolic input-output process model and value function.
Erik Derner, Jirí Kubalík, Robert Babuska
IROS3
2018 Decentralized Reinforcement Learning of Robot Behaviors
David Leonardo Leottau, Javier Ruiz-del-Solar, Robert Babuska
Artif. Intell.3
2018 Policy derivation methods for critic-only reinforcement learning in continuous spaces
Eduard Alibekov, Jirí Kubalík, Robert Babuska
Eng. Appl. Artif. Intell.3
2018 Experience Selection in Deep Reinforcement Learning for Control
abstract
Experience replay is a technique that allows off-policy reinforcement-learning methods to reuse past experiences. The stability and speed of convergence of reinforcement learning, as well as the eventual performance of the learned policy, are strongly dependent on the experiences being replayed. Which experiences are replayed depends on two important choices. The first is which and how many experiences to retain in the experience replay buffer. The second choice is how to sample the experiences that are to be replayed from that buffer. We propose new methods for the combined problem of experience retention and experience sampling. We refer to the combination as experience selection. We focus our investigation specifically on the control of physical systems, such as robots, where exploration is costly. To determine which experiences to keep and which to replay, we investigate different proxies for their immediate and long-term utility. These proxies include age, temporal difference error and the strength of the applied exploration noise. Since no currently available method works in all situations, we propose guidelines for using prior knowledge about the characteristics of the control problem at hand to choose the appropriate experience replay strategy.
Tim de Bruin, Jens Kober, Karl Tuyls, Robert Babuska
J. Mach. Learn. Res.4
2018 A Multiple-Model Reliability Prediction Approach for Condition-Based Maintenance
abstract
Numerous prognostic methods have been developed, aiming at predicting future system reliability with the highest possible accuracy. It is striking that the relation with the subsequent maintenance optimization process is generally overlooked, while it is important in practice. Additionally, almost all existing methods are based on a single degradation measure, and focus on systems with only one degradation and failure mode. In practice, however, multiple degradation measures are often available and needed to adequately predict future system degradation. Moreover, systems may suffer from various kinds of faults, all resulting in different degradation behaviors. To accommodate these properties, we establish a link between failure prognosis and maintenance optimization, and accordingly propose a multivariate multiple-model approach to system reliability prediction. We conclude that in the presence of multiple degradation modes and provided they are correctly identified, a multiple-model approach outperforms a single-model approach with respect to the prediction accuracy. Moreover, in the presence of multiple degradation and failure modes, overall predictions of the remaining useful life as generated by common prognostic approaches are not directly suited for maintenance decision making, as different kinds of system failures and maintenance activities are associated with different costs. In contrast, our approach yields conditional predictions of future system reliability, which much better suit the maintenance optimization process.
Kim Verbert, Bart De Schutter, Robert Babuska
IEEE Trans. Reliab.3
2017 Automated tuning and configuration of path planning algorithms
abstract
A large number of novel path planning methods for a wide range of problems have been described in literature over the past few decades. These algorithms can often be configured using a set of parameters that greatly influence their performance. In a typical use case, these parameters are only very slightly tuned or even left untouched. Systematic approaches to tune parameters of path planning algorithms have been largely unexplored. At the same time, there is a rising interest in the planning and robotics communities regarding the real world application of these theoretically developed and simulation-tested planning algorithms. In this work, we propose the use of Sequential Model-based Algorithm Configuration (SMAC) tools to address these concerns. We show that it is possible to improve the performance of a planning algorithm for a specific problem without the need of in-depth knowledge of the algorithm itself. We compare five planners that see a lot of practical usage on three typical industrial pick-and-place tasks to demonstrate the effectiveness of the method.
Ruben Burger, Mukunda Bharatheesha, Marc van Bert, Robert Babuska
ICRA4
2017 Enhanced Symbolic Regression Through Local Variable Transformations
Jirí Kubalík, Erik Derner, Robert Babuska
IJCCI3
2017 Combining knowledge and historical data for system-level fault diagnosis of HVAC systems
Kim Verbert, Robert Babuska, Bart De Schutter
Eng. Appl. Artif. Intell.2
2017 Bayesian and Dempster-Shafer reasoning for knowledge-based fault diagnosis-A comparative study
Kim Verbert, Robert Babuska, Bart De Schutter
Eng. Appl. Artif. Intell.2
2017 Railway Track Circuit Fault Diagnosis Using Recurrent Neural Networks
abstract
Timely detection and identification of faults in railway track circuits are crucial for the safety and availability of railway networks. In this paper, the use of the long-short-term memory (LSTM) recurrent neural network is proposed to accomplish these tasks based on the commonly available measurement signals. By considering the signals from multiple track circuits in a geographic area, faults are diagnosed from their spatial and temporal dependences. A generative model is used to show that the LSTM network can learn these dependences directly from the data. The network correctly classifies 99.7% of the test input sequences, with no false positive fault detections. In addition, the t-Distributed Stochastic Neighbor Embedding (t-SNE) method is used to examine the resulting network, further showing that it has learned the relevant dependences in the data. Finally, we compare our LSTM network with a convolutional network trained on the same task. From this comparison, we conclude that the LSTM network architecture is better suited for the railway track circuit fault detection and identification tasks than the convolutional network.
Tim de Bruin, Kim Verbert, Robert Babuska
IEEE Trans. Neural Networks Learn. Syst.3
2016 Deep convolutional neural networks for detection of rail surface defects
abstract
In this paper, we propose a deep convolutional neural network solution to the analysis of image data for the detection of rail surface defects. The images are obtained from many hours of automated video recordings. This huge amount of data makes it impossible to manually inspect the images and detect rail surface defects. Therefore, automated detection of rail defects can help to save time and costs, and to ensure rail transportation safety. However, one major challenge is that the extraction of suitable features for detection of rail surface defects is a non-trivial and difficult task. Therefore, we propose to use convolutional neural networks as a viable technique for feature learning. Deep convolutional neural networks have recently been applied to a number of similar domains with success. We compare the results of different network architectures characterized by different sizes and activation functions. In this way, we explore the efficiency of the proposed deep convolutional neural network for detection and classification. The experimental results are promising and demonstrate the capability of the proposed approach.
Shahrzad Faghih-Roohi, Siamak Hajizadeh, Alfredo Núñez, Robert Babuska, Bart De Schutter
IJCNN4
2016 Improved deep reinforcement learning for robotics through distribution-based experience retention
abstract
Recent years have seen a growing interest in the use of deep neural networks as function approximators in reinforcement learning. In this paper, an experience replay method is proposed that ensures that the distribution of the experiences used for training is between that of the policy and a uniform distribution. Through experiments on a magnetic manipulation task it is shown that the method reduces the need for sustained exhaustive exploration during learning. This makes it attractive in scenarios where sustained exploration is in-feasible or undesirable, such as for physical systems like robots and for life long learning. The method is also shown to improve the generalization performance of the trained policy, which can make it attractive for transfer learning. Finally, for small experience databases the method performs favorably when compared to the recently proposed alternative of using the temporal difference error to determine the experience sample distribution, which makes it an attractive option for robots with limited memory capacity.
Tim de Bruin, Jens Kober, Karl Tuyls, Robert Babuska
IROS4
2016 Decentralized Reinforcement Learning Applied to Mobile Robots
David Leonardo Leottau, Aashish Vatsyayan, Javier Ruiz-del-Solar, Robert Babuska
RoboCup4
2016 Nested compliant admittance control for robotic mechanical assembly of misaligned and tightly toleranced parts
abstract
In this paper, we propose a closed-loop force sensor based nested admittance/impedance control strategy to actively estimate and minimize the effects of geometric misalignment that naturally occur during assembly tasks with compliant robots. The method allows the robot to be used with a stiff impedance control setting, which is beneficial for free air motion performance, yet allows to adjust for large misalignment errors between parts that need be assembled. First, the stability bounds on the control parameters of this method are established through numerical simulation, after which they are compared with the experimentally determined parameters. Trial peg-in-hole insertion experiments are performed with a 6-DOF KUKA LWR-4+ robot under various degrees of rotational misalignment, where metal pegs are being inserted into metal holes under tight tolerances. The proposed method allows successful peg insertions even under large rotational misalignments of up to 20° (13 fold increase compared to 1.5° we achieved with traditional impedance control alone), without the need to adjust the position trajectories with complex models on the fly. Moreover, it provided a 5 fold reduction of the average forces exerted on the environment compared with using impedance control alone.
Nicky Mol, Jan Smisek, Robert Babuska, André Schiele
SMC3
2016 Online learning for optimistic planning
Lucian Busoniu, Alexander Daniels, Robert Babuska
Eng. Appl. Artif. Intell.3
2016 Fault diagnosis using spatial and temporal information with application to railway track circuits
Kim Verbert, Bart De Schutter, Robert Babuska
Eng. Appl. Artif. Intell.3
2016 Learning Sequential Composition Control
abstract
Sequential composition is an effective supervisory control method for addressing control problems in nonlinear dynamical systems. It executes a set of controllers sequentially to achieve a control specification that cannot be realized by a single controller. As these controllers are designed offline, sequential composition cannot address unmodeled situations that might occur during runtime. This paper proposes a learning approach to augment the standard sequential composition framework by using online learning to handle unforeseen situations. New controllers are acquired via learning and added to the existing supervisory control structure. In the proposed setting, learning experiments are restricted to take place within the domain of attraction (DOA) of the existing controllers. This guarantees that the learning process is safe (i.e., the closed loop system is always stable). In addition, the DOA of the new learned controller is approximated after each learning trial. This keeps the learning process short as learning is terminated as soon as the DOA of the learned controller is sufficiently large. The proposed approach has been implemented on two nonlinear systems: 1) a nonlinear mass-damper system and 2) an inverted pendulum. The results show that in both cases a new controller can be rapidly learned and added to the supervisory control structure.
Esmaeil Najafi, Robert Babuska, Gabriel A. D. Lopes
IEEE Trans. Cybern.2
2016 Unified Modeling and Control of Walking and Running on the Spring-Loaded Inverted Pendulum
abstract
This paper addresses the control of steady state and transition behaviors for the bipedal spring-loaded inverted pendulum (SLIP) model. We present an event-driven control approach that enables the realization of active running, walking, and walk-run transitions in a unified framework. The synthesis of the controlled behaviors is illustrated by the notion of hybrid automaton in which different gaits are generated as the sequential composition of SLIP's primary phases of motion. We also propose a novel analytical approximate solution to the otherwise nonintegrable double-stance dynamics of the SLIP model. The analytical simplicity of the solution is utilized in the design and analysis of dynamic walking gaits suitable for online implementation. The accuracy of the approximate solution and its influence on the stability properties of the controlled system are carefully analyzed. Finally, we present two simulation examples. The first demonstrates the practicality of the proposed control strategy in creating human-like gaits and gait transitions. In the second example, we use the controlled SLIP as a planner for the control of a multibody bipedal robot model, and embed SLIP-like behaviors into a physics-based robot simulation model. The results corroborate both the practical utility and effectiveness of the proposed approach.
Mohammad Shahbazi, Robert Babuska, Gabriel A. D. Lopes
IEEE Trans. Robotics2
2015 Analytical approximation for the double-stance phase of a walking robot
abstract
This paper introduces an approximate analytical solution to the otherwise non-integrable double-stance dynamics of the bipedal spring-loaded inverted pendulum (SLIP). Despite the apparent structural simplicity of the SLIP, the exact analytical solution to its stance dynamics cannot be found. Approximate maps have been proposed for the monoped SLIP runner (encompassing a single-stance phase). Still, even in an approximate form, a solution to the double-stance dynamics of the bipedal SLIP walker remained an open problem. We propose a double-stance map that can be readily utilized especially in the design of control systems for active dynamic walking. The accuracy of the derived map over a feasible range of locomotion properties is analyzed numerically, and a control application based on this solution is presented. Simulations for an arbitrary chosen energy level reveals that the devised controller enlarges the stable walking domain of the standard SLIP considerably.
Mohammad Shahbazi, Robert Babuska, Gabriel A. D. Lopes
ICRA2
2015 Attitude and altitude estimation and control on board a Flapping Wing Micro Air Vehicle
abstract
The autonomous capabilities of light-weight Flapping Wing Micro Air Vehicles (FWMAVs) have much to gain from onboard state estimation and attitude control. In this article, we present the first FWMAV with robust onboard state estimation and attitude control. The tailed FWMAV DelFly II was used, with the main goal to achieve active stabilization in the (passively unstable) hover condition. The attitude is estimated using an Inertial Measurement Unit with a gyroscope, accelerometer and magnetometer and the altitude is estimated using a barometer. A major challenge lies in the disturbance of the accelerometer measurements by the flapping motion of the wings. We propose a mechanical damping mechanism and flap-cycle based filtering to resolve this issue. The pitch estimates have a mean error of 1.5° with respect to the ground-truth measurement from a motion capture system. Using the onboard pitch estimate we can control the attitude of the FWMAV in the forward flight regime with a 30% lower standard deviation than in a trimmed flight. With a different set of gains, the FWMAV is able to perform a hovering flight - showing that a tailed FWMAV has enough control authority for this task. In a fully autonomous hover experiment, the DelFly II stays within a sphere of 0.75 m radius.
J. L. Verboom, Sjoerd Tijmons, Christophe De Wagter, B. D. W. Remes, Robert Babuska, Guido de Croon
ICRA5
2015 Reinforcement Learning for Port-Hamiltonian Systems
abstract
Passivity-based control (PBC) for port-Hamiltonian systems provides an intuitive way of achieving stabilization by rendering a system passive with respect to a desired storage function. However, in most instances the control law is obtained without any performance considerations and it has to be calculated by solving a complex partial differential equation (PDE). In order to address these issues we introduce a reinforcement learning (RL) approach into the energy-balancing passivity-based control (EB-PBC) method, which is a form of PBC in which the closed-loop energy is equal to the difference between the stored and supplied energies. We propose a technique to parameterize EB-PBC that preserves the systems's PDE matching conditions, does not require the specification of a global desired Hamiltonian, includes performance criteria, and is robust. The parameters of the control law are found by using actor-critic (AC) RL, enabling the search for near-optimal control policies satisfying a desired closed-loop energy landscape. The advantage is that the solutions learned can be interpreted in terms of energy shaping and damping injection, which makes it possible to numerically assess stability using passivity theory. From the RL perspective, our proposal allows for the class of port-Hamiltonian systems to be incorporated in the AC framework, speeding up the learning thanks to the resulting parameterization of the policy. The method has been successfully applied to the pendulum swing-up problem in simulations and real-life experiments.
Olivier Sprangers, Robert Babuska, Subramanya Nageshrao, Gabriel A. D. Lopes
IEEE Trans. Cybern.2
2014 Modeling and Control of Legged Locomotion via Switching Max-Plus Models
abstract
We present a gait generation framework for multi-legged robots based on max-plus algebra that is endowed with intrinsically safe gait transitions. The time schedule of each foot liftoff and touchdown is modeled by sets of max-plus linear equations. The resulting discrete-event system is translated to continuous time via piecewise constant leg phase velocities; thus, it is compatible with traditional central pattern generator approaches. Different gaits and gait parameters are interleaved by utilizing different max-plus system matrices. We present various gait transition schemes and show that optimal transitions, in the sense of minimizing the stance time variation, allow for constant acceleration and deceleration on legged platforms. The framework presented in this paper relies on a compact representation of the gait space, provides guarantees regarding the transient and steady-state behavior, and results in simple implementations on legged robotic platforms.
Gabriel A. D. Lopes, Bart Kersbergen, Ton J. J. van den Boom, Bart De Schutter, Robert Babuska
IEEE Trans. Robotics5
2013 Optimistic planning for continuous-action deterministic systems
abstract
We consider the class of online planning algorithms for optimal control, which compared to dynamic programming are relatively unaffected by large state dimensionality. We introduce a novel planning algorithm called SOOP that works for deterministic systems with continuous states and actions. SOOP is the first method to explore the true solution space, consisting of infinite sequences of continuous actions, without requiring knowledge about the smoothness of the system. SOOP can be used parameter-free at the cost of more model calls, but we also propose a more practical variant tuned by a parameter α, which balances finer discretization with longer planning horizons. Experiments on three problems show SOOP reliably ranks among the best algorithms, fully dominating competing methods when the problem requires both long horizons and fine discretization.
Lucian Busoniu, Alexander Daniels, Rémi Munos, Robert Babuska
ADPRL4
2013 On the convergence of Ant Colony Optimization with stench pheromone
abstract
Ant Colony Optimization (ACO) has proved to be a powerful metaheuristic for combinatorial optimization problems. From a theoretical point of view, the convergence of the ACO algorithm is an important issue. In this paper, we analyze the convergence properties of a recently introduced ACO algorithm, called ACO with stench pheromone (ACO-SP), which can be used to solve dynamic traffic routing problems through finding the minimum cost routes in a traffic network. This new algorithm has two different types of pheromone: the regular pheromone that is used to attract artificial ants to the arc in the network with the lowest cost, and the stench pheromone that is used to push ants away when too many ants converge to that arc. As a first step of a convergence proof for ACO-SP, we consider a network with two arcs. We show that the process of pheromone update will transit among different modes, and finally stay in a stable mode, thus proving convergence for this given case.
Zhe Cong, Bart De Schutter, Robert Babuska
IEEE Congress on Evolutionary Computation3
2013 Evolutionary co-optimization of control and system parameters for a resonating robot arm
abstract
In this paper we simultaneously optimize the parameters describing the morphology of a robot arm and the parameters of its nonlinear controller. A novel concept of a pick-and-place robot arm is considered, which is called the resonating arm (RA). It uses a nonlinear spring mechanism to generate pick-and-place motions without the need for powerful actuators. This improves energy efficiency, cost and weight of the robot arm. Because of the complex interactions of the spring mechanism and the controller, we use evolutionary co-optimization to optimize the RA system as a whole. The results reveal that evolutionary co-optimization yields near optimal solutions for a 1 degree of freedom (1-DOF) RA, which require 43% less torque than the solution found through a separate optimization of the system and the control parameters. In case of a 2-DOF RA, evolutionary co-optimization resulted in credible solutions as well, but with less consistency.
Jurren Pen, Wouter Caarls, Martijn Wisse, Robert Babuska
ICRA4
2013 Solutions to finite horizon cost problems using actor-critic reinforcement learning
abstract
Actor-critic reinforcement learning algorithms have shown to be a successful tool in learning the optimal control for a range of (repetitive) tasks on systems with (partially) unknown dynamics, which may or may not be nonlinear. Most of the reinforcement learning literature published up to this point only deals with modeling the task at hand as a Markov decision process with an infinite horizon cost function. In practice, however, it is sometimes desired to have a solution for the case where the cost function is defined over a finite horizon, which means that the optimal control problem will be time-varying and thus harder to solve. This paper adapts two previously introduced actor-critic algorithms from the infinite horizon setting to the finite horizon setting and applies them to learning a task on a nonlinear system, without needing any assumptions or knowledge about the system dynamics, using radial basis function networks. Simulations on a typical nonlinear motion control problem are carried out, showing that actor-critic algorithms are capable of solving the difficult problem of time-varying optimal control. Moreover, the benefit of using a model learning technique is shown.
Ivo Grondman, Hao Xu 0002, Sarangapani Jagannathan, Robert Babuska
IJCNN4
2013 Robot learning and use of affordances in goal-directed tasks
abstract
An affordance is a relation between an object, an action, and the effect of that action in a given environmental context. One key benefit of the concept of affordance is that it provides information about the consequence of an action which can be stored and reused in a range of tasks that a robot needs to learn and perform. In this paper, we address the challenge of the on-line learning and use of affordances simultaneously while performing goal-directed tasks. This requires efficient online performance to ensure the robot is able to achieve its goal fast. By providing conceptual knowledge of action possibilities and desired effects, we show that a humanoid robot NAO can learn and use affordances in two different task settings. We demonstrate the effectiveness of this approach by integrating affordances into an Extended Classifier System for learning general rules in a reinforcement learning framework. Our experimental results show significant speedups in learning how a robot solves a given task.
Chang Wang 0005, Koen V. Hindriks, Robert Babuska
IROS3
2013 Parametric Bayesian Filters for Nonlinear Stochastic Dynamical Systems: A Survey
abstract
Nonlinear stochastic dynamical systems are commonly used to model physical processes. For linear and Gaussian systems, the Kalman filter is optimal in minimum mean squared error sense. However, for nonlinear or non-Gaussian systems, the estimation of states or parameters is a challenging problem. Furthermore, it is often required to process data online. Therefore, apart from being accurate, the feasible estimation algorithm also needs to be fast. In this paper, we review Bayesian filters that possess the aforementioned properties. Each filter is presented in an easy way to implement algorithmic form. We focus on parametric methods, among which we distinguish three types of filters: filters based on analytical approximations (extended Kalman filter, iterated extended Kalman filter), filters based on statistical approximations (unscented Kalman filter, central difference filter, Gauss-Hermite filter), and filters based on the Gaussian sum approximation (Gaussian sum filter). We discuss each of these filters, and compare them with illustrative examples.
Pawel Stano, Zsófia Lendek, Jelmer Braaksma, Robert Babuska, Cees de Keizer, Arnold J. den Dekker
IEEE Trans. Cybern.4
2012 Adaptive fuzzy observer and robust controller for a 2-DOF robot arm
abstract
Recently, adaptive fuzzy observers have been introduced that are capable of estimating uncertainties along with the states of a nonlinear system represented by an uncertain Takagi-Sugeno (TS) model. In this paper, we use such an adaptive observer to estimate the uncertainties in the state matrices of a two-degrees-of-freedom robot arm model. The TS model of the robot arm is constructed using the sector nonlinearity approach. The estimates are used in updating the model, and the updated model is used to design a controller for the robot arm. We analyze the improvement in the achievable controller performance when using the adaptive observer.
S. Bindiganavile Nagesh, Zsófia Lendek, Amol A. Khalate, Robert Babuska
FUZZ-IEEE4
2012 Experience Replay for Real-Time Reinforcement Learning Control
abstract
Reinforcement-learning (RL) algorithms can automatically learn optimal control strategies for nonlinear, possibly stochastic systems. A promising approach for RL control is experience replay (ER), which learns quickly from a limited amount of data, by repeatedly presenting these data to an underlying RL algorithm. Despite its benefits, ER RL has been studied only sporadically in the literature, and its applications have largely been confined to simulated systems. Therefore, in this paper, we evaluate ER RL on real-time control experiments that involve a pendulum swing-up problem and the vision-based control of a goalkeeper robot. These real-time experiments are complemented by simulation studies and comparisons with traditional RL. As a preliminary, we develop a general ER framework that can be combined with essentially any incremental RL technique, and instantiate this framework for the approximate Q-learning and SARSA algorithms. The successful real-time learning results that are presented here are highly encouraging for the applicability of ER RL in practice.
Sander Adam, Lucian Busoniu, Robert Babuska
IEEE Trans. Syst. Man Cybern. Part C3
2012 A Survey of Actor-Critic Reinforcement Learning: Standard and Natural Policy Gradients
abstract
Policy-gradient-based actor-critic algorithms are amongst the most popular algorithms in the reinforcement learning framework. Their advantage of being able to search for optimal policies using low-variance gradient estimates has made them useful in several real-life applications, such as robotics, power control, and finance. Although general surveys on reinforcement learning techniques already exist, no survey is specifically dedicated to actor-critic algorithms in particular. This paper, therefore, describes the state of the art of actor-critic algorithms, with a focus on methods that can work in an online setting and use function approximation in order to deal with continuous state and action spaces. After starting with a discussion on the concepts of reinforcement learning and the origins of actor-critic algorithms, this paper describes the workings of the natural gradient, which has made its way into many actor-critic algorithms over the past few years. A review of several standard and natural actor-critic algorithms is given, and the paper concludes with an overview of application areas and a discussion on open issues.
Ivo Grondman, Lucian Busoniu, Gabriel A. D. Lopes, Robert Babuska
IEEE Trans. Syst. Man Cybern. Part C4
2012 Efficient Model Learning Methods for Actor-Critic Control
abstract
We propose two new actor-critic algorithms for reinforcement learning. Both algorithms use local linear regression (LLR) to learn approximations of the functions involved. A crucial feature of the algorithms is that they also learn a process model, and this, in combination with LLR, provides an efficient policy update for faster learning. The first algorithm uses a novel model-based update rule for the actor parameters. The second algorithm does not use an explicit actor but learns a reference model which represents a desired behavior, from which desired control actions can be calculated using the inverse of the learned process model. The two novel methods and a standard actor-critic algorithm are applied to the pendulum swing-up problem, in which the novel methods achieve faster learning than the standard algorithm.
Ivo Grondman, Maarten Vaandrager, Lucian Busoniu, Robert Babuska, Erik Schuitema
IEEE Trans. Syst. Man Cybern. Part B4
2012 Machine Learning Algorithms in Bipedal Robot Control
abstract
Over the past decades, machine learning techniques, such as supervised learning, reinforcement learning, and unsupervised learning, have been increasingly used in the control engineering community. Various learning algorithms have been developed to achieve autonomous operation and intelligent decision making for many complex and challenging control problems. One of such problems is bipedal walking robot control. Although still in their early stages, learning techniques have demonstrated promising potential to build adaptive control systems for bipedal robots. This paper gives a review of recent advances on the state-of-the-art learning algorithms and their applications to bipedal robot control. The effects and limitations of different learning techniques are discussed through a representative selection of examples from the literature. Guidelines for future research on learning control of bipedal robots are provided in the end.
W. Art Chaovalitwongse, Robert Babuska
IEEE Trans. Syst. Man Cybern. Part C3
2011 Approximate reinforcement learning: An overview
abstract
Reinforcement learning (RL) allows agents to learn how to optimally interact with complex environments. Fueled by recent advances in approximation-based algorithms, RL has obtained impressive successes in robotics, artificial intelligence, control, operations research, etc. However, the scarcity of survey papers about approximate RL makes it difficult for newcomers to grasp this intricate field. With the present overview, we take a step toward alleviating this situation. We review methods for approximate RL, starting from their dynamic programming roots and organizing them into three major classes: approximate value iteration, policy iteration, and policy search. Each class is subdivided into representative categories, highlighting among others offline and online algorithms, policy gradient methods, and simulation-based techniques. We also compare the different categories of methods, and outline possible ways to enhance the reviewed algorithms.
Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska
ADPRL4
2011 Optimistic planning for sparsely stochastic systems
abstract
We propose an online planning algorithm for finite-action, sparsely stochastic Markov decision processes, in which the random state transitions can only end up in a small number of possible next states. The algorithm builds a planning tree by iteratively expanding states, where each expansion exploits sparsity to add all possible successor states. Each state to expand is actively chosen to improve the knowledge about action quality, and this allows the algorithm to return a good action after a strictly limited number of expansions. More specifically, the active selection method is optimistic in that it chooses the most promising states first, so the novel algorithm is called optimistic planning for sparsely stochastic systems. We note that the new algorithm can also be seen as model-predictive (receding-horizon) control. The algorithm obtains promising numerical results, including the successful online control of a simulated HIV infection with stochastic drug effectiveness.
Lucian Busoniu, Rémi Munos, Bart De Schutter, Robert Babuska
ADPRL4
2011 Decentralized Kalman filter comparison for distributed-parameter systems: A case study for a 1D heat conduction process
abstract
In this paper we compare four methods for decentralized Kalman filtering for distributed-parameter systems, which after spatial and temporal discretization, result in large-scale linear discrete-time systems. These methods are: parallel information filter, distributed information filter, distributed Kalman filter with consensus filter, and distributed Kalman filter with weighted averaging. These filters are suitable for sensor networks, where the sensor nodes perform not only sensing and computations, but also communicate estimates among each other. We consider an application of sensor networks to a heat conduction process. The performance of the decentralized filters is evaluated and compared to the centralized Kalman filter.
Zulkifli Hidayat, Robert Babuska, Bart De Schutter, Alfredo Núñez
ETFA2
2011 Optimal gait switching for legged locomotion
abstract
Switching gaits in many-legged robots can present challenges due to the combinatorial nature of the gait space. In this paper we present an intrinsically safe gait switching generator that minimizes the velocity variance of all the legs in stance, allowing for smooth acceleration in legged robots. The gait switching generator is modeled as a max-plus linear discrete event system which is translated to continuous time via a reference trajectory generator.
Bart Kersbergen, Gabriel A. D. Lopes, Ton J. J. van den Boom, Bart De Schutter, Robert Babuska
IROS5
2011 Sequential stability analysis and observer design for distributed TS fuzzy systems
Zsófia Lendek, Robert Babuska, Bart De Schutter
Fuzzy Sets Syst.2
2011 Erratum to "Adaptive observers for TS fuzzy systems with unknown polynomial inputs" [Fuzzy Sets and Systems 161 (2010) 2043-2065]
Zsófia Lendek, Jimmy Lauber, Thierry-Marie Guerra, Robert Babuska, Bart De Schutter
Fuzzy Sets Syst.4
2011 Cross-Entropy Optimization of Control Policies With Adaptive Basis Functions
abstract
This paper introduces an algorithm for direct search of control policies in continuous-state discrete-action Markov decision processes. The algorithm looks for the best closed-loop policy that can be represented using a given number of basis functions (BFs), where a discrete action is assigned to each BF. The type of the BFs and their number are specified in advance and determine the complexity of the representation. Considerable flexibility is achieved by optimizing the locations and shapes of the BFs, together with the action assignments. The optimization is carried out with the cross-entropy method and evaluates the policies by their empirical return from a representative set of initial states. The return for each representative state is estimated using Monte Carlo simulations. The resulting algorithm for cross-entropy policy search with adaptive BFs is extensively evaluated in problems with two to six state variables, for which it reliably obtains good policies with only a small number of BFs. In these experiments, cross-entropy policy search requires vastly fewer BFs than value-function techniques with equidistant BFs, and outperforms policy search with a competing optimization algorithm called DIRECT.
Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska
IEEE Trans. Syst. Man Cybern. Part B4
2010 Generalized pheromone update for Ant Colony Learning in continuous state spaces
abstract
In this paper, we discuss the Ant Colony Learning (ACL) paradigm for non-linear systems with continuous state spaces. ACL is a novel control policy learning methodology, based on Ant Colony Optimization. In ACL, a collection of agents, called ants, jointly interact with the system at hand in order to find the optimal mapping between states and actions. Through the stigmergic interaction by pheromones, the ants are guided by each others experience towards better control policies. In order to deal with continuous state spaces, we generalize the concept of pheromones and the local and global pheromone update rules. As a result of this generalization, we can integrate both crisp and fuzzy partitioning of the state space into the ACL framework. We compare the performance of ACL with these two partitioning methods by applying it to the control problem of swinging-up and stabilizing an under-actuated pendulum.
Jelmer van Ast, Robert Babuska, Bart De Schutter
IEEE Congress on Evolutionary Computation2
2010 Online self-organizing adaptive fuzzy controller: Application to a nonlinear servo system
abstract
This paper presents a self-organizing adaptive fuzzy controller that works online. No prior knowledge about the differential equations governing the plant nor offline training is needed. Starting from very simple topologies, the algorithm uses the data obtained online during the normal operation of the system to modify the structure of the fuzzy controller. This is achieved in two phases: first, the consequents of the current fuzzy rules are adapted; in the second phase, new membership functions are added online. To show its capabilities, a real experiment with a nonlinear servo system has been carried out with satisfactory results.
Ana Belén Cara, Zsófia Lendek, Robert Babuska, Héctor Pomares, Ignacio Rojas
FUZZ-IEEE3
2010 On non-PDC local observers for TS fuzzy systems
abstract
In this paper we propose a method to design local observers for Takagi-Sugeno fuzzy models obtained from nonlinear systems by the sector nonlinearity approach. When a global observer cannot be designed, using our method it is still possible to design observers that are valid in a well-defined region of the state-space. The design is based on a nonquadratic Lyapunov function. Depending on whether or not the scheduling vector is a function of the states to be estimated, the conditions are formulated as an LMI or a BMI problem, respectively. The results are illustrated on simulation examples, for which classical observer design conditions are unfeasible.
Zsófia Lendek, Thierry-Marie Guerra, Robert Babuska
FUZZ-IEEE3
2010 Distributed nonlinear estimation for robot localization using weighted consensus
abstract
Distributed linear estimation theory has received increased attention in recent years due to several promising industrial applications. Distributed nonlinear estimation, however is still a relatively unexplored field despite the need in numerous practical situations for techniques that can handle nonlinearities. This paper presents a unified way of describing distributed implementations of three commonly used nonlinear estimators: the Extended Kalman Filter, the Unscented Kalman Filter and the Particle Filter. Leveraging on the presented framework, we propose new distributed versions of these methods, in which the nonlinearities are locally managed by the various sensors whereas the different estimates are merged based on a weighted average consensus process. The proposed versions are shown to outperform the few published ones in two robot localization test cases.
Andrea Simonetto, Tamás Keviczky, Robert Babuska
ICRA3
2010 Motion estimation based on predator/prey vision
abstract
We present an unscented Kalman filter based state estimator for a fast moving rigid body (such as a mobile robot) endowed with two video cameras. We focus on forward velocity estimation towards the computation of standard energy cost functions for legged locomotion. Points are chosen as image features and the model of each camera is based on the traditional pinhole projection. The resulting filter's state is composed of the rigid body pose and velocities, together with a measure of depth for each tracked point. By taking inspiration from nature's large predatory and grazing mammals eye configuration, we suggest, via simulation results, a solution for the question of finding the best orientation of two cameras, between side and frontal facing, for velocity estimation in a forward moving robot.
David van der Lijn, Gabriel A. D. Lopes, Robert Babuska
IROS3
2010 Control delay in Reinforcement Learning for real-time dynamic systems: A memoryless approach
abstract
Robots controlled by Reinforcement Learning (RL) are still rare. A core challenge to the application of RL to robotic systems is to learn despite the existence of control delay - the delay between measuring a system's state and acting upon it. Control delay is always present in real systems. In this work, we present two novel temporal difference (TD) learning algorithms for problems with control delay. These algorithms improve learning performance by taking the control delay into account. We test our algorithms in a gridworld, where the delay is an integer multiple of the time step, as well as in the simulation of a robotic system, where the delay can have any value. In both tests, our proposed algorithms outperform classical TD learning algorithms, while maintaining low computational complexity.
Erik Schuitema, Lucian Busoniu, Robert Babuska, Pieter P. Jonker
IROS3
2010 Adaptive observers for TS fuzzy systems with unknown polynomial inputs
Zsófia Lendek, Jimmy Lauber, Thierry-Marie Guerra, Robert Babuska, Bart De Schutter
Fuzzy Sets Syst.4
2009 Policy search with cross-entropy optimization of basis functions
abstract
This paper introduces a novel algorithm for approximate policy search in continuous-state, discrete-action Markov decision processes (MDPs). Previous policy search approaches have typically used ad-hoc parameterizations developed for specific MDPs. In contrast, the novel algorithm employs a flexible policy parameterization, suitable for solving general discrete-action MDPs. The algorithm looks for the best closed-loop policy that can be represented using a given number of basis functions, where a discrete action is assigned to each basis function. The locations and shapes of the basis functions are optimized, together with the action assignments. This allows a large class of policies to be represented. The optimization is carried out with the cross-entropy method and evaluates the policies by their empirical return from a representative set of initial states. We report simulation experiments in which the algorithm reliably obtains good policies with only a small number of basis functions, albeit at sizable computational costs.
Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska
ADPRL4
2009 Modeling of Hot Rolling Industrial Process Using Fuzzy Logic
Alaa F. Sheta, Ertan Öznergiz, M. A. Abdelrahman, Robert Babuska
CAINE4
2009 Stability of Cascaded Fuzzy Systems and Observers
abstract
A large class of nonlinear systems can be well approximated by Takagi-Sugeno (TS) fuzzy models with linear or affine consequents. It is well known that the stability of these consequent models does not ensure the stability of the overall fuzzy system. Therefore, several stability conditions have been developed for TS fuzzy systems. We study a special class of nonlinear dynamic systems that can be decomposed into cascaded subsystems, which are represented as TS fuzzy models. We analyze the stability of the overall TS system based on the stability of the subsystems and prove that the stability of the subsystems implies the stability of the overall system. The main benefit of this approach is that it relaxes the conditions imposed when the system is globally analyzed, thereby solving some of the feasibility problems. Another benefit is that by using this approach, the dimension of the associated linear matrix inequality (LMI) problem can be reduced. For naturally distributed applications, such as multiagent systems, the construction and tuning of a centralized observer may not be feasible. Therefore, we also extend the cascaded approach to the observer design and use fuzzy observers to individually estimate the states of these subsystems. A theoretical proof of stability and simulation examples are presented. The results show that the distributed observer achieves the same performance as the centralized one, while leading to increased modularity, reduced complexity, lower computational costs, and easier tuning. Applications of such cascaded systems include multiagent systems, distributed process control, and hierarchical large-scale systems.
Zsófia Lendek, Robert Babuska, Bart De Schutter
IEEE Trans. Fuzzy Syst.2
2008 Ant Colony Optimization for optimal control
abstract
Ant Colony Optimization (ACO) has proven to be a very powerful optimization heuristic for Combinatorial Optimization Problems (COPs). It has been demonstrated to work well when applied to various NP-complete problems, such as the traveling salesman problem. In this paper, an ACO approach to optimal control is proposed. This approach requires that a continuous-time, continuous-state model of the system, together with a finite action set, is formulated as a discrete, non-deterministic automaton. The control problem is then translated into a stochastic COP. This method is applied to the time-optimal swing-up and stabilization of a pendulum.
Jelmer van Ast, Robert Babuska, Bart De Schutter
IEEE Congress on Evolutionary Computation2
2008 A general modeling framework for swarms
abstract
Swarms are characterized by the ability to generate complex behavior from the coupling of simple individuals. While the swarm approach to distributed systems of moving agents is gradually finding a way to engineering applications, a true successful demonstration of an engineered swarm is still missing. One of the reasons for this is the gap between the complexity of the swarms studied in fundamental research and the complexity needed for the application to interesting control problems. In the majority of the research on swarm intelligent systems, the moving agents in the swarm are modeled as simple reactive agents. This model comprises too little intelligence to fully exploit the potential of swarms. In this paper, a general comprehensive swarm framework is introduced and related to the established state of the art. Such a framework is novel and it is a first and important step in the development and analysis of more complex and intelligent swarms.
Jelmer van Ast, Robert Babuska, Bart De Schutter
IEEE Congress on Evolutionary Computation2
2008 Consistency of fuzzy model-based reinforcement learning
abstract
Reinforcement learning (RL) is a widely used paradigm for learning control. Computing exact RL solutions is generally only possible when process states and control actions take values in a small discrete set. In practice, approximate algorithms are necessary. In this paper, we propose an approximate, model-based Q-iteration algorithm that relies on a fuzzy partition of the state space, and on a discretization of the action space. Using assumptions on the continuity of the dynamics and of the reward function, we show that the resulting algorithm is consistent, i.e., that the optimal solution is obtained asymptotically as the approximation accuracy increases. An experimental study indicates that a continuous reward function is also important for a predictable improvement in performance as the approximation accuracy increases.
Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska
FUZZ-IEEE4
2008 Stability analysis and observer design for decentralized TS fuzzy systems
abstract
A large class of nonlinear systems can be well approximated by Takagi-Sugeno (TS) fuzzy models, with linear or affine consequents. It is well-known that the stability of these consequent models does not ensure the stability of the overall fuzzy system. Stability conditions developed for TS fuzzy systems in general rely on the feasibility of an associated system of linear matrix inequalities, whose complexity may grow exponentially with the number of rules. We study distributed systems, where the subsystems are represented as TS fuzzy models. For such systems, a centralized analysis is often unfeasible. We analyze the stability of the overall TS system based on the stability of the subsystems and the strength of the interconnection terms. For naturally distributed applications, such as multi-agent systems, when adding new subsystems “on-line”, the construction and tuning of a centralized observer is often intractable. Therefore, we also propose a decentralized approach to observer design. Applications of such systems include distributed process control, traffic networks, and economic systems.
Zsófia Lendek, Robert Babuska, Bart De Schutter
FUZZ-IEEE2
2008 Adaptive fuzzy control of a non-linear servo-drive: Theory and experimental results
Domenico Bellomo, David Naso, Robert Babuska
Eng. Appl. Artif. Intell.3
2008 Distributed Kalman filtering for cascaded systems
Zsófia Lendek, Robert Babuska, Bart De Schutter
Eng. Appl. Artif. Intell.2
2008 A Comprehensive Survey of Multiagent Reinforcement Learning
abstract
Multiagent systems are rapidly finding applications in a variety of domains, including robotics, distributed control, telecommunications, and economics. The complexity of many tasks arising in these domains makes them difficult to solve with preprogrammed agent behaviors. The agents must, instead, discover a solution on their own, using learning. A significant part of the research on multiagent learning concerns reinforcement learning techniques. This paper provides a comprehensive survey of multiagent reinforcement learning (MARL). A central issue in the field is the formal statement of the multiagent learning goal. Different viewpoints on this issue have led to the proposal of many different goals, among which two focal points can be distinguished: stability of the agents' learning dynamics, and adaptation to the changing behavior of the other agents. The MARL algorithms described in the literature aim---either explicitly or implicitly---at one of these two goals or at a combination of both, in a fully cooperative, fully competitive, or more general setting. A representative selection of these algorithms is discussed in detail in this paper, together with the specific issues that arise in each category. Additionally, the benefits and challenges of MARL are described along with some of the problem domains where the MARL techniques have been applied. Finally, an outlook for the field is provided.
Lucian Busoniu, Robert Babuska, Bart De Schutter
IEEE Trans. Syst. Man Cybern. Part C2
2007 Fuzzy Approximation for Convergent Model-Based Reinforcement Learning
abstract
Reinforcement learning (RL) is a learning control paradigm that provides well-understood algorithms with good convergence and consistency properties. Unfortunately, these algorithms require that process states and control actions take only discrete values. Approximate solutions using fuzzy representations have been proposed in the literature for the case when the states and possibly the actions are continuous. However, the link between these mainly heuristic solutions and the larger body of work on approximate RL, including convergence results, has not been made explicit. In this paper, we propose a fuzzy approximation structure for the Q-value iteration algorithm, and show that the resulting algorithm is convergent. The proof is based on an extension of previous results in approximate RL. We then propose a modified, serial version of the algorithm that is guaranteed to converge at least as fast as the original algorithm. An illustrative simulation example is also provided.
Lucian Busoniu, Damien Ernst, Bart De Schutter, Robert Babuska
FUZZ-IEEE4
2007 Stability of Cascaded Takagi-Sugeno Fuzzy Systems
abstract
A large class of nonlinear systems can be well approximated by Takagi-Sugeno (TS) fuzzy models, with local models often chosen linear or affine. It is well-known that the stability of these local models does not ensure the stability of the overall fuzzy system. Therefore, several stability conditions have been developed for TS fuzzy systems. We study a special class of nonlinear dynamic systems, that can be decomposed into cascaded subsystems. These subsystems are represented as TS fuzzy models. We analyze the stability of the overall TS system based on the stability of the subsystems. For a general nonlinear, cascaded system, global asymptotic stability of the individual subsystems is not sufficient for the stability of the cascade. However, for the case of TS fuzzy systems, we prove that the stability of the subsystems implies the stability of the overall system. The main benefit of this approach is that it relaxes the conditions imposed when the system is globally analyzed, therefore solving some of the feasibility problems. Another benefit is, that by using this approach, the dimension of the associated linear matrix inequality (LMI) problem can be reduced. Applications of such cascaded systems include multi-agent systems, distributed process control and hierarchical large-scale systems.
Zsófia Lendek, Robert Babuska, Bart De Schutter
FUZZ-IEEE2
2006 A Fuzzy-logic System for Detecting Oscillations in Control Loops
abstract
The automatic detection of oscillations in control loops is essential for effective performance monitoring. However, the methods known from the literature are often sensitive to normal system responses such as step changes in the reference or the rejection of disturbances. Therefore, a novel criterion is proposed in this paper. It uses on-line spectral analysis over a moving window and subsequent fuzzy decision making based on the magnitude and duration of oscillations as criteria. An offline adaptation mechanism is available to tune the system with the help of data and expert knowledge. The usefulness of this criterion has been demonstrated by using real-time data from dissolved oxygen and pH control loops in a fermentation process.
Robert Babuska, Jelmer van Ast, Samir Mesic
FUZZ-IEEE1
2006 Virtual Sensor for the Angle-of-Attack Signal in Small Commercial Aircraft
abstract
An aircraft carries on board many sensors which measure a wide variety of variables. Due to the relations between the measured signals, a certain level redundancy is available. This redundancy can be used to estimate a particular variable based on signals that represent other variables. Such an estimator can be used as a virtual sensor. This paper describes the design of a virtual sensor for the Angle-of-Attack signal in a small commercial aircraft. In order to effectively use all available knowledge and data, and to comply with the stringent design requirements, the virtual sensor combines a number of technologies: a white-box linear time-varying model, a gray-box nonlinear Takagi-Sugeno (TS) fuzzy model and a black-box neural network compensator, whose purpose is to reduce the estimation error of the linear parameter varying model. The TS model and the neural network are trained by using data from nonlinear aircraft simulations. The inputs of the neural network are selected by a genetic search algorithm with a backward elimination procedure. Extensive evaluation has shown that the design requirements are amply met and that the proposed design methodology has a good potential for future applications in aircraft and other high-performance systems.
Marcel Oosterom, Robert Babuska
FUZZ-IEEE2
2006 Multi-Agent Reinforcement Learning: A Survey
abstract
Multi-agent systems are rapidly finding applications in a variety of domains, including robotics, distributed control, telecommunications, economics. Many tasks arising in these domains require that the agents learn behaviors online. A significant part of the research on multi-agent learning concerns reinforcement learning techniques. However, due to different viewpoints on central issues, such as the formal statement of the learning goal, a large number of different methods and approaches have been introduced. In this paper we aim to present an integrated survey of the field. First, the issue of the multi-agent learning goal is discussed, after which a representative selection of algorithms is reviewed. Finally, open issues are identified and future research directions are outlined
Lucian Busoniu, Robert Babuska, Bart De Schutter
ICARCV2
2006 Decentralized Reinforcement Learning Control of a Robotic Manipulator
abstract
Multi-agent systems are rapidly finding applications in a variety of domains, including robotics, distributed control, telecommunications, etc. Learning approaches to multi-agent control, many of them based on reinforcement learning (RL), are investigated in complex domains such as teams of mobile robots. However, the application of decentralized RL to low-level control tasks is not as intensively studied. In this paper, we investigate centralized and decentralized RL, emphasizing the challenges and potential advantages of the latter. These are then illustrated on an example: learning to control a two-link rigid manipulator. Some open issues and future research directions in decentralized RL are outlined
Lucian Busoniu, Bart De Schutter, Robert Babuska
ICARCV3
2006 Dynamic Exploration in Q(lambda)-learning
abstract
Reinforcement learning has proved its value in solving complex optimization tasks. However, the learning time for even simple problems is typically very long. Efficient exploration of the state-action space is therefore crucial for effective learning. This paper introduces a new type of exploration, called dynamic exploration. It differs from the existing exploration methods (both directed and undirected) in that it makes exploration a function of the action selected in the previous time step. In our approach, states can either belong to long-path states, where the optimal action is the same as the optimal action in the previous state, or to switch states, where the action is different. In realistic learning problems, the number of long-path states exceeds the number of switch states. Given this information, the exploration method can explore the state-space more efficiently. Experiments on different gridworld optimization tasks demonstrate the reduction of learning time with dynamic exploration.
Jelmer van Ast, Robert Babuska
IJCNN2
2006 Reinforcement Learning Control for Biped Robot Walking on Uneven Surfaces
abstract
Biped robots based on the concept of (passive) dynamic walking are far simpler than the traditional fullyI controlled walking robots, while achieving a more natural gait and consuming less energy. However, lightly actuated dynamic walking robots, which rely on the natural limit cycle of their mechanical structure, are very sensitive to ground disturbances. Already a very small step down can cause the robot to lose stability. In this paper, we investigate the use of reinforcement learning to make a dynamic walking robot more robust against ground disturbances. The learning controller is applied to a simulated two-link biped which is an abstraction of a mechanical prototype developed at the Delft Biorobotics Laboratory. The learning controller has been designed such that it can be applied as a straightforward extension of the proportionalI-derivative (PD) controller currently used to drive the robot's pneumatic actuators. The learning controller is therefore suitable for the future implementation in the robot hardware. Simulation results demonstrate that the biped quickly learns to overcome step-down disturbances on the floor up to 10% of the leg length, without compromising the natural walking style provided by the PD controller, which was optimized for walking on an even surface.
Jelmer Braaksma, Robert Babuska, Daan G. E. Hobbelen
IJCNN3
2006 Genetic polynomial regression as input selection algorithm for non-linear identification
Koen Maertens, Josse De Baerdemaeker, Robert Babuska
Soft Comput.3
2005 Evaluation of adaptive fuzzy controllers: a real-world experiment
abstract
Several stable adaptive fuzzy control schemes based on feedback linearization and Lyapunov synthesis were proposed in the literature. However, most of such controllers have been only tested on relatively simple simulations examples, in which the effects of noise, uncertainties, computational times and other fundamental problems related to hardware implementation are neglected. In this paper, we study and compare direct and indirect adaptive control schemes by means of an experimental benchmark (two coupled DC machines). For the indirect schemes, we consider both the standard adaptive laws based on tracking error and composite adaptive laws based on tracking and prediction error. The goal of this work is to gain insight in the benefits and drawbacks of the different variants of adaptive fuzzy controllers and to evaluate their potential for practical applications
Domenico Bellomo, David Naso, Robert Babuska
FUZZ-IEEE3
2005 Perspectives of fuzzy systems and control
Antonio Sala 0001, Thierry-Marie Guerra, Robert Babuska
Fuzzy Sets Syst.3
2004 FUZZSAM - visualization of fuzzy clustering results by modified Sammon mapping
abstract
Since in practical data mining problems high-dimensional data are clustered, the resulting clusters are high-dimensional geometrical objects, which are difficult to analyze and interpret. Cluster validity measures try to solve this problem by providing a single numerical value. As a low dimensional graphical representation of the clusters could be much more informative than such a single value, this paper proposes a new tool for the visualization of fuzzy clustering results. By using the basic properties of fuzzy clustering algorithms, this new tool maps the cluster centers and the data such that the distances between the clusters and the data-points are preserved. During the iterative mapping process, the algorithm uses the membership values of the data and minimizes an objective function similar to the original clustering algorithm. Comparing to the original Sammon mapping not only reliable cluster shapes are obtained but the numerical complexity of the algorithm is also drastically reduced. The algorithm has been applied to several data sets and the numerical results show performance superior to principal component analysis and the classical Sammon mapping based projection. The examples demonstrate that proposed FUZZSAMM algorithm is a useful tool in user-guided clustering.
János Abonyi, Robert Babuska
FUZZ-IEEE2
2004 Fuzzy clustering for selecting structure of nonlinear models with mixed discrete and continuous inputs
abstract
A method for selecting regressors in nonlinear models with mixed discrete (categorical) and continuous inputs is proposed. Given a set of input-output data and an initial superset of potential inputs, the relevant inputs are selected by a model-free search algorithm. Fuzzy clustering is used to quantize continuous data into subsets that can be handled in a similar way as discrete data. Two simulation examples and one real-world data set are included to illustrate the performance of the proposed method and compare it with the performance of regression trees. For small to medium size problems (up to 15 candidate inputs), the proposed method works effectively. For larger problems, the computational load becomes too high.
Daniela Girimonte, Robert Babuska, János Abonyi
FUZZ-IEEE2
2004 Parameter Convergence in Adaptive Fuzzy Control
Domenico Bellomo, David Naso, Robert Babuska
ICINCO (1)3
2004 Hybrid Control Design for a Robot Manipulator in a Shield Tunneling Machine
Jelmer Braaksma, J. Ben Klaassens, Robert Babuska, Cees de Keizer
ICINCO (2)3
2004 Effective optimization for fuzzy model predictive control
abstract
This paper addresses the optimization in fuzzy model predictive control. When the prediction model is a nonlinear fuzzy model, nonconvex, time-consuming optimization is necessary, with no guarantee of finding an optimal solution. A possible way around this problem is to linearize the fuzzy model at the current operating point and use linear predictive control (i.e., quadratic programming). For long-range predictive control, however, the influence of the linearization error may significantly deteriorate the performance. In our approach, this is remedied by linearizing the fuzzy model along the predicted input and output trajectories. One can further improve the model prediction by iteratively applying the optimized control sequence to the fuzzy model and linearizing along the so obtained simulated trajectories. Four different methods for the construction of the optimization problem are proposed, making difference between the cases when a single linear model or a set of linear models are used. By choosing an appropriate method, the user can achieve a desired tradeoff between the control performance and the computational load. The proposed techniques have been tested and evaluated using two simulated industrial benchmarks: pH control in a continuous stirred tank reactor and a high-purity distillation column.
Stanimir Mollov, Robert Babuska, János Abonyi, Henk B. Verbruggen
IEEE Trans. Fuzzy Syst.2
2004 Input selection for nonlinear regression models
abstract
A simple and effective method for the selection of significant inputs in nonlinear regression models is proposed. Given a set of input-output data and an initial superset of potential inputs, the relevant inputs are selected by checking whether after deleting a particular input, the data set is still consistent with the basic property of a function. In order to be able to handle real-valued and noisy data in a sensible manner, fuzzy clustering is first applied. The obtained clusters are compared by using a similarity measure in order to find inconsistencies within the data. Several examples using simulated and real-world data sets are presented to demonstrate the effectiveness of the algorithm.
Radek Sindelár, Robert Babuska
IEEE Trans. Fuzzy Syst.2
2003 Design of optimal membership functions for fuzzy gain-scheduled control
abstract
The design of optimal membership functions for model-based fuzzy gain-scheduled control is addressed. The antecedent membership functions in the controller are computed such that the closed-loop behavior complies with the specifications over the entire operating range. It is shown that better performance is obtained than with the standard Parallel Distributed Control (PDC) approach, which is based on using the model membership functions in the controller. A real-world application example of aircraft gain-scheduled control is presented.
Robert Babuska, Marcel Oosterom
FUZZ-IEEE1
2003 Multiobjective identification of Takagi-Sugeno fuzzy models
abstract
The problem of identifying the parameters of the constituent local linear models of Takagi-Sugeno fuzzy models is considered. In order to address the tradeoff between global model accuracy and interpretability of the local models as linearizations of a nonlinear system, two multiobjective identification algorithms are studied. Particular attention is paid to the analysis of conflicts between objectives, and we show that such information can be easily computed from the solution of the multiobjective optimization. This information is useful to diagnose the model and tune the weighting/priorities of the multiobjective optimization. Moreover, the result of the conflict analysis can be used as a constructive tool to modify the fuzzy model structure (including membership functions) in order to meet the multiple objectives. Simple illustrative examples as well as experimental results show the usefulness of the method.
Tor Arne Johansen, Robert Babuska
IEEE Trans. Fuzzy Syst.2
2003 Fuzzy gain scheduling: controller and observer design based on Lyapunov method and convex optimization
abstract
Addresses model-based fuzzy control. A constructive and automated method for the design of a gain-scheduling controller is presented. Based on a given Takagi-Sugeno fuzzy model of the plant, the controller is designed such that stability and prescribed performance of the closed loop are guaranteed. These properties are valid in a wide working range around an equilibrium without restrictions to slowly varying trajectories. The synthesis is based on linear matrix inequalities and convex optimization techniques. If required, a fuzzy state estimator and an extended controller can be included, providing a zero steady-state error in the presence of disturbances and modeling errors. The proposed method has been applied to a control of a laboratory liquid-level process. Hence, the performance has been evaluated in simulations as well as in real-time control.
Petr Korba, Robert Babuska, Henk B. Verbruggen, Paul Martin Frank
IEEE Trans. Fuzzy Syst.2
2003 Comments on the benchmarks in "A proposal for improving the accuracy of Linguistic Modeling" and related articles
abstract
In the above paper by Cordon and Herrara (IEEE Trans. Fuzzy Syst., vol. 8, p. 335-44, 2000), the so-called accurate linguistic modeling (ALM) method was proposed to improve the accuracy of linguistic fuzzy models. A number of examples are given to demonstrate the benefits of the approach. We show that: 1) these examples are not suitable as benchmarks or demonstrators of nonlinear modeling techniques and 2) better results can be obtained by using both standard regression tools as well as other fuzzy modeling techniques. We argue that benchmark examples that are used in articles to demonstrate the effectiveness of fuzzy modeling techniques should be selected with great care. Critical analysis of the results should be made and linear models should be regarded as a lower bound on the acceptable performance.
Johannes A. Roubos, Robert Babuska
IEEE Trans. Fuzzy Syst.2
2002 Improved covariance estimation for Gustafson-Kessel clustering
abstract
This article presents two techniques to improve the calculation of the fuzzy covariance matrix in the Gustafson-Kessel (GK) clustering algorithm. The first one overcomes problems that occur in the standard GK clustering when the number of data samples is small or when the data within a cluster are linearly correlated. The improvement is achieved by fixing the ratio between the maximal and minimal eigenvalue of the covariance matrix. The second technique is useful when the GK algorithm is employed in the extraction of Takagi-Sugeno fuzzy model from data. It reduces the risk of overfitting when the number of training samples is low in comparison to the number of clusters. This is achieved by adding a scaled unity matrix to the calculated covariance matrix. Numerical examples are presented to demonstrate the benefits of the proposed techniques
Robert Babuska, Peter J. van der Veen, Uzay Kaymak
FUZZ-IEEE1
2002 Fault-tolerant model-based predictive control using multiple Takagi-Sugeno fuzzy models
abstract
We address fault-tolerant control for nonlinear plants. We perform soft fault detection and isolation for partial fault detection. A method is proposed that combines model predictive controllers with fuzzy Takagi-Sugeno models. An example is presented to illustrate the functionality of the proposed approach.
Alexandar Ichtev, J. Hellendoom, Robert Babuska, Stanimir Mollov
FUZZ-IEEE3
2002 Robust stability constraints for fuzzy model predictive control
abstract
This paper addresses the synthesis of a predictive controller for a nonlinear process based on a fuzzy model of the Takagi-Sugeno (T-S) type, resulting in a stable closed-loop control system. Conditions are given that guarantee closed-loop robust asymptotic stability for open-loop bounded-input-bounded-output (BIBO) stable processes with an additive l/sub 1/-norm bounded model uncertainty. The idea is closely related to (small-gain-based) l/sub 1/-control theory, but due to the time-varying approach, the resulting robust stability constraints are less conservative. Therefore the fuzzy model is viewed as a linear time-varying system rather than a nonlinear one. The goal is to obtain constraints on the control signal and its increment that guarantee robust stability. Robust global asymptotic stability and offset-free reference tracking are guaranteed for asymptotically constant reference trajectories and disturbances.
Stanimir Mollov, Ton J. J. van den Boom, Federico Cuesta, Aníbal Ollero, Robert Babuska
IEEE Trans. Fuzzy Syst.5
2002 Modified Gath-Geva fuzzy clustering for identification of Takagi-Sugeno fuzzy models
abstract
The construction of interpretable Takagi-Sugeno (TS) fuzzy models by means of clustering is addressed. First, it is shown how the antecedent fuzzy sets and the corresponding consequent parameters of the TS model can be derived from clusters obtained by the Gath-Geva (GG) algorithm. To preserve the partitioning of the antecedent space, linearly transformed input variables can be used in the model. This may, however, complicate the interpretation of the rules. To form an easily interpretable model that does not use the transformed input variables, a new clustering algorithm is proposed, based on the expectation-maximization (EM) identification of Gaussian mixture models. This new technique is applied to two well-known benchmark problems: the MPG (miles per gallon) prediction and a simulated second-order nonlinear process. The obtained results are compared with results from the literature.
János Abonyi, Robert Babuska, Ferenc Szeifert
IEEE Trans. Syst. Man Cybern. Part B2
2002 Soft computing applications in aircraft sensor management and flight control law reconfiguration
abstract
A sensor management system based on soft computing techniques has been developed and implemented in the flight control system of a small commercial aircraft. Unlike in the conventional sensor management system, the signals from sensors are assigned weights based on fuzzy membership functions and the consolidated signal is computed as a weighted average. This approach improves the quality of the consolidated signal and reduces transients due to sensor failures. This soft voting is extended to soft flight control law reconfiguration. In addition, a virtual sensor has been introduced as an arbitrator which enables the isolation of the failed sensor in the duplex operation and the detection of a sensor failure in the simplex operation. The effectiveness of the proposed methods is demonstrated by using an extensive simulation model of a small commercial aircraft, developed by airframe and control system manufacturers on the basis of an existing business jet. Furthermore, the system has been successfully evaluated and compared to standard techniques by means of pilot-in-the-loop simulations on the Research Flight Simulator of the National Aerospace Laboratory in The Netherlands. This application, developed within a Brite/EuRam research project, is characterized by the effective combination of novel soft computing techniques with standard, well proven methods of the aircraft industry. The properties of the conventional sensor management system have been retained, with the additional advantage that the quality of the consolidated signal is improved, the failure-induced transients are reduced, and the consolidated signal remains available up to the last valid sensor.
Marcel Oosterom, Robert Babuska, Henk B. Verbruggen
IEEE Trans. Syst. Man Cybern. Part C2
2001 Accurate, Transparent, and Compact Fuzzy Models for Function Approximation and Dynamic Modeling through Multi-objective Evolutionary Optimization
Fernando Jiménez, Antonio F. Skarmeta, Johannes A. Roubos, Robert Babuska
EMO4
2001 Fault-Tolerant Output-Feedback Control Via Fuzzy State Blending
abstract
A fault-tolerant scheme for sensor faults in output-feedback control is developed by using standard linear design methods combined with try decision logic and 'soft' reconfiguration based on fuzzy state blending. The main advantages of the presented method are the use of standard deterministic control-oriented observers without the need to design special diagnostic observers for the different faults, and the simplicity and transparency of the decision logic. The control scheme is tolerant with regard to total and partial sensor faults occurring separately or simultaneously (under certain specified conditions). The decision logic gives information both on the localization of the fault and its severity. Experimental real-time results are presented for a laboratory system.
Robert Babuska
FUZZ-IEEE1
2001 Fault Detection and Isolation Using Multiple Takagi-Sugeno Fuzzy Models
abstract
In this paper, soft fault detection and isolation (FDI) for nonlinear plants is addressed. A new method has been developed that combines fuzzy Takagi-Sugeno (TS) models, used for residual generation, with quadratic programming for residual evaluation and fault isolation. This method can isolate and identify total and partial failures, both single and multiple. The information from the FDI module can be used for control reconfiguration. One simulation and one real-time example are given to illustrate the functionality of the proposed approach.
Alexandar Ichtev, Hans Hellendoorn, Robert Babuska
FUZZ-IEEE3
2001 Analysis of Interactions in MIMO Takagi-Sugeno Fuzzy Models
abstract
Input-output interactions in the inherently nonlinear Takagi-Sugeno (TS) MIMO fuzzy models can not be analyzed by using standard methods such as the relative gain array (RGA). Two methods are proposed that exploit the specific TS structure. The first one is based on a number of RGAs which can indicate sufficiently well the interactions in the model. Another tool to analyze input-output interactions is the output sensitivity, computed as a partial derivative of the output with respect to the considered input. The use of these techniques is illustrated with the help of a TS fuzzy model of a high-purity distillation column.
Stanimir Mollov, Robert Babuska, Henk B. Verbruggen
FUZZ-IEEE2
2001 Fuzzy Gain Scheduling for Flight Control Laws
abstract
The state-of-the-art methodology for the design of digital flight-by-wire flight control laws is based on dimming and linearization of a nonlinear aircraft model at selected operating points and subsequent tuning of linear control laws. Despite recent advances in the development of computer-aided control design toots, selection of the operating points and the design of the gain schedule still has to be done manually, in a heuristic manner. In order to reduce the design effort, an automated procedure has been developed. The number of operating points and their locations are determined automatically on the basis of the changes in the aerodynamics over the flight envelope, by using fuzzy clustering. This approach also directly provides the interpolation mechanism (membership functions) for the local control law parameters. The procedure has been developed in close cooperation with airframe and control system manufacturers and was evaluated in pilot-in-the-loop flight simulator tests.
Marcel Oosterom, Robert Babuska
FUZZ-IEEE2
2001 Efficient Training Algorithm for Takagi-Sugeno Type Neuro-Fuzzy Network
abstract
The paper describes an algorithm that can be used to train the Takagi-Sugeno (TS) type neuro-fuzzy network very efficiently. The training algorithm is very efficient in the sense that it can bring the performance index of the network, such as the sum squared error (SSE), down to the desired error goal much faster than that the classical backpropagation algorithm (BPA). The proposed training algorithm is based on a slight modification of the Levenberg-Marquardt training algorithm (LMA) which takes into account the modified error index extension of sum squared error as the new performance index of the network. The Levenberg-Marquardt algorithm uses the Jacobian matrix in order to approximate the Hessian matrix and that is the most important and difficult step in implementing this LMA. Therefore, a simple technique has been described to compute first the transpose of Jacobian matrix by comparing two equations and thereafter by further transposing the former one the actual Jacobian matrix is computed that is found to be robust against the modified error index extension. Furthermore, care has been taken to suppress or control the oscillation magnitude during the training of neuro-fuzzy network. Finally, the above training algorithm is tested on neuro-fuzzy modeling and prediction applications of time series and nonlinear plant.
Ajoy Kumar Palit, Robert Babuska
FUZZ-IEEE2
2001 Fuzzy modeling with multivariate membership functions: gray-box identification and control design
abstract
A novel framework for fuzzy modeling and model-based control design is described. The fuzzy model is of the Takagi-Sugeno (TS) type with constant consequents. It uses multivariate antecedent membership functions obtained by Delaunay triangulation of their characteristic points. The number and position of these points are determined by an iterative insertion algorithm. Constrained optimization is used to estimate the consequent parameters, where the constraints are based on control-relevant a priori knowledge about the modeled process. Finally, methods for control design through linearization and inversion of this model are developed. The proposed techniques are demonstrated by means of two benchmark examples: identification of the well-known Box-Jenkins gas furnace and inverse model-based control of a pH process. The obtained results are compared with results from the literature.
János Abonyi, Robert Babuska, Ferenc Szeifert
IEEE Trans. Syst. Man Cybern. Part B2
2001 Rule base reduction: some comments on the use of orthogonal transforms
abstract
Comments on recent publications about the use of orthogonal transforms to order and select rules in a fuzzy rule base. The techniques are well-known from linear algebra, and we comment on their usefulness in fuzzy modeling. The application of rank-revealing methods based on singular value decomposition (SVD) to rule reduction gives rather conservative results. They are essentially subset selection methods, and we show that such methods do not produce an "importance ordering", contrary to what has been stated in the literature. The orthogonal least-squares (OLS) method, which evaluates the contribution of the rules to the output, is more attractive for systems modeling. However, it has been shown to sometimes assign high importance to rules that are correlated in the premise. This hampers the generalization capabilities of the resulting model. We discuss the performance of rank-revealing reduction methods and advocate the use of a less complex method based on the pivoted QR decomposition. Further, we show how detection of redundant rules can be introduced in OLS by a simple extension of the algorithm. The methods are applied to a problem known from the literature and compared to results reported by other researchers.
Magne Setnes, Robert Babuska
IEEE Trans. Syst. Man Cybern. Syst.2
2000 Local and global identification and interpretation of parameters in Takagi-Sugeno fuzzy models
abstract
Addresses the interpretation of parameters in Takagi-Sugeno (TS) fuzzy models. The analysis is presented for the dynamic gain and steady-state representation, but it holds for parameters related to the dynamics as well. The TS model interpolates between local linear models. The overall gain obtained by interpolating the gains of the local models can be interpreted as the local dynamic gain of the entire fuzzy model. This locally interpreted gain is not identical to the dynamic gain obtained by linearization of the fuzzy model at the considered equilibrium. We analyze the origin of this difference with regard to the applied identification method. In order to keep the analysis simple and transparent, a fuzzy model of a Hammerstein system is studied. The results show that fuzzy models obtained by local identification (weighted least squares for each rule) typically yields a poor steady-state representation and the model can only be locally interpreted. On the contrary, a fuzzy model obtained by global identification (one least-square solution for the entire model) can result in a qualitatively bad local interpretation of the gain even though approximates the real process well. Therefore, this model can only be used for prediction or local linearization through Taylor expansion. It is shown that the difference between the globally and locally interpreted gain can be reduced by using a priori knowledge in global identification. The steady-state representation of fuzzy models obtained by local identification can be improved by using inference based on the smoothed maximum operator (instead of the weighted mean).
János Abonyi, Robert Babuska
FUZZ-IEEE2
2000 Decoupling MBPC using fuzzy models
abstract
A decoupling strategy for affine discrete-time nonlinear systems is proposed. The idea is to invert a nonlinear model and to use the inverted model to compensate for the coupling. Linear model based predictive controller is employed to provide the reference to be followed. The control system comprises: 1) a feedback decoupling scheme, 2) a linear model-based predictive controller (MBPC), and 3) a constraints mapping that transforms the actual input constraints into constraints on the linear system. The nonlinear system to be decoupled is expressed by a Takagi-Sugeno type fuzzy model.
Stanimir Mollov, Robert Babuska, Henk B. Verbruggen
FUZZ-IEEE2
2000 Performance evaluation of a fuzzy rule base for control purpose
abstract
This paper introduces the concept of rule base consistency in a Takagi-Sugeno fuzzy model. A consistency measure is defined based on the agreement between the local behaviour of the output and the local behaviour obtained by aggregating the consequent parameters. This consistency measure is useful, for instance, when designing a controller for a fuzzy model obtained by identification methods. The proposed method is illustrated with the help of two examples.
D. Passaquay, M. Bross, Robert Babuska, Serge Boverie, André Titli
FUZZ-IEEE3
1999 Fuzzy model-based predictive control using Takagi-Sugeno models
Johannes A. Roubos, Stanimir Mollov, Robert Babuska, Henk B. Verbruggen
Int. J. Approx. Reason.3
1999 Fuzzy relational classifier trained by fuzzy clustering
abstract
A novel approach to nonlinear classification is presented, in the training phase of the classifier, the training data is first clustered in an unsupervised way by fuzzy c-means or a similar algorithm. The class labels are not used in this step. Then, a fuzzy relation between the clusters and the class identifiers is computed. This approach allows the number of prototypes to be independent of the number of actual classes. For the classification of unseen patterns, the membership degrees of the feature vector in the clusters are first computed by using the distance measure of the clustering algorithm. Then, the output fuzzy set is obtained by relational composition. This fuzzy set contains the membership degrees of the pattern in the given classes. A crisp decision is obtained by defuzzification, which gives either a single class or a "reject" decision, when a unique class cannot be selected based on the available information. The principle of the proposed method is demonstrated on an artificial data set and the applicability of the method is shown on the identification of live-stock from recorded sound sequences. The obtained results are compared with two other classifiers.
Magne Setnes, Robert Babuska
IEEE Trans. Syst. Man Cybern. Part B2
1998 Transparent Fuzzy Modelling
Magne Setnes, Robert Babuska, Henk B. Verbruggen
Int. J. Hum. Comput. Stud.2
1998 Adaptive fuzzy control of satellite attitude by reinforcement learning
abstract
The attitude control of a satellite is often characterized by a limit cycle, caused by measurement inaccuracies and noise in the sensor output. In order to reduce the limit cycle, a nonlinear fuzzy controller was applied. The controller was tuned by means of reinforcement learning without using any model of the sensors or the satellite. The reinforcement signal is computed as a fuzzy performance measure using a noncompensatory aggregation of two control subgoals. Convergence of the reinforcement learning scheme is improved by computing the temporal difference error over several time steps and adapting the critic and the controller at a lower sampling rate. The results show that an adaptive fuzzy controller can better cope with the sensor noise and nonlinearities than a standard linear controller.
Walter M. van Buijtenen, Gerard Schram, Robert Babuska, Henk B. Verbruggen
IEEE Trans. Fuzzy Syst.3
1998 Similarity measures in fuzzy rule base simplification
abstract
In fuzzy rule-based models acquired from numerical data, redundancy may be present in the form of similar fuzzy sets that represent compatible concepts. This results in an unnecessarily complex and less transparent linguistic description of the system. By using a measure of similarity, a rule base simplification method is proposed that reduces the number of fuzzy sets in the model. Similar fuzzy sets are merged to create a common fuzzy set to replace them in the rule base. If the redundancy in the model is high, merging similar fuzzy sets might result in equal rules that also can be merged, thereby reducing the number of rules as well. The simplified rule base is computationally more efficient and linguistically more tractable. The approach has been successfully applied to fuzzy models of real world systems.
Magne Setnes, Robert Babuska, Uzay Kaymak, H. R. van Nauta Lemke
IEEE Trans. Syst. Man Cybern. Part B2
1998 Rule-based modeling: precision and transparency
abstract
This article is a reaction to recent publications on rule-based modeling using fuzzy set theory and fuzzy logic. The interest in fuzzy systems has recently shifted from the seminal ideas about complexity reduction toward data-driven construction of fuzzy systems. Many algorithms have been introduced that aim at numerical approximation of functions by rules, but pay little attention to the interpretability of the resulting rule base. We show that fuzzy rule-based models acquired from measurements can be both accurate and transparent by using a low number of rules. The rules are generated by product-space clustering and describe the system in terms of the characteristic local behavior of the system in regions identified by the clustering algorithm. The fuzzy transition between rules makes it possible to achieve precision along with a good qualitative description in linguistic terms. The latter is useful for expert evaluation, rule-base maintenance, operator training, control systems design, user interfacing, etc. We demonstrate the approach on a modeling problem from a recently published article.
Magne Setnes, Robert Babuska, Henk B. Verbruggen
IEEE Trans. Syst. Man Cybern. Part C2
1997 Knowledge-based fuzzy model for performance prediction of a rock-cutting trencher
M. H. den Hartog, Robert Babuska, H. J. R. Deketh, M. Alvarez Grima, P. N. W. Verhoef, Henk B. Verbruggen
Int. J. Approx. Reason.2