VLDB 2026 Research / reviewers in the wild / expert
Freek Stulp
dblp:73/478
· DBLP profile ↗
50ranked-venue papers
15as first author
13since 2021 · last 2026
0000-0001-9555-9517ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 14 first-author · 12 since 2021Systems, architecture and hardware · 27 · 9 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Ergodic Approach to Robotic Surface Finishing With Learned Motion PreferencesabstractSurface finishing is a time-consuming, dangerous task, difficult to automate despite its necessity in many manufacturing processes. Its automation, particularly through robotics, increases productivity and relieves workers from health-critical tasks. However, challenges remain, as automated offline planning tools can result in certain areas being either neglected or overly processed. Ergodic control offers the possibility to cover target probability distributions in an online manner, by taking into account the observed coverage history. However, existing ergodic control approaches provide little flexibility in designing and adapting coverage strategies. Moreover, they come with simplifying assumptions, such as point-based dynamics, which are no longer valid for tasks where the robot is in contact with strongly varying curvatures on non-trivial surface geometries. In this work, we introduce a closed-form ergodic control framework that includes the tool imprint in the system modeling while simultaneously permitting the intuitive transfer of finish strategies, namely preferred motion directions. We build on the Spectral Multiscale Coverage (SMC) approach, augmenting it with a tool imprint model, as well as both target distributions and state-dependent movement directions extracted from human demonstrations. Through evaluations in a surface finishing task using a torque-controlled, 7-DoF, robot arm we show that our approach optimally covers surfaces according to the tool contact area, with robust error convergence. Stefan Schneyer, Korbinian Nottensteiner, Alin Albu-Schäffer, Freek Stulp, João Silvério |
IEEE Trans. Robotics | 4 |
| 2025 | RACCOON: Grounding Embodied Question-Answering with State Summaries from Existing Robot ModulesabstractExplainability is vital for establishing user trust, also in robotics. Recently, foundation models (e.g. vision-language models, VLMs) fostered a wave of embodied agents that answer arbitrary queries about their environment and their interactions with it. However, naively prompting VLMs to answer queries based on camera images does not take into account existing robot architectures which represent the robot's tasks, skills, and beliefs about the state of the world. To overcome this limitation, we propose RACCOON, a framework that combines foundation models' responses with a robot's internal knowledge. Inspired by Retrieval-Augmented Generation (RAG), RACCOON selects relevant context, retrieves information from the robot's state, and utilizes it to refine prompts for an LLM to answer questions accurately. This bridges the gap between the model's adaptability and the robot's domain expertise. Samuel Bustamante-Gomez, Markus Knauer, Jeremias Thun, Stefan Schneyer, Alin Albu-Schäffer, Bernhard M. Weber, Freek Stulp |
ICRA | 7 |
| 2024 | Unknown Object Grasping for Assistive RoboticsabstractWe propose a novel pipeline for unknown object grasping in shared robotic autonomy scenarios. State-of-the-art methods for fully autonomous scenarios are typically learning-based approaches optimised for a specific end-effector, that generate grasp poses directly from sensor input. In the domain of assistive robotics, we seek instead to utilise the user’s cognitive abilities for enhanced satisfaction, grasping performance, and alignment with their high level task-specific goals. Given a pair of stereo images, we perform unknown object instance segmentation and generate a 3D reconstruction of the object of interest. In shared control, the user then guides the robot end-effector across a virtual hemisphere centered around the object to their desired approach direction. A physics-based grasp planner finds the most stable local grasp on the reconstruction, and finally the user is guided by shared control to this grasp. In experiments on the DLR EDAN platform, we report a grasp success rate of 87% for 10 unknown objects, and demonstrate the method’s capability to grasp objects in structured clutter and from shelves. Elle Miller, Maximilian Durner, Matthias Humt, Gabriel Quere, Wout Boerdijk, Ashok M. Sundaram, Freek Stulp, Jörn Vogel |
ICRA | 7 |
| 2024 | Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment CollaborationabstractLarge, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train "generalist" X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. The project website is robotics-transformer-x.github.io. Abigail O'Neill, Abhiram Maddukuri, Abhishek Gupta 0004, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew E. Wang, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess-Limerick, Bernhard Schölkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu 0003, Charlotte Le, Chelsea Finn, Chen Wang 0053, Chenfeng Xu, Cheng Chi 0001, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Drieß, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Paul Foster, Fangchen Liu, Federico Ceola, Fei Xia 0002, Feiyu Zhao, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guanzhi Wang, Hao Su 0001, Haoshu Fang, Henghui Bao, Heni Ben Amor, Henrik I. Christensen, Hiroki Furuta, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim 0001, Jaimyn Drake, Jan Peters 0001, Jan Schneider 0007, Jasmine Hsu, Jeannette Bohg, Jeffrey T. Bingham, Jensen Gao, Jiaheng Hu, Jiajun Wu 0001, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan 0001, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jitendra Malik, João Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Kenneth Y. Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Zhang 0002, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Yunliang Chen 0001, Lerrel Pinto, Li Fei-Fei 0001, Liam Tan, Linxi Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang 0003, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma 0001, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen 0001, Nicolas Heess, Nikhil J. Joshi, Niko Sünderhauf, Norman Di Palo, Nur Muhammad Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R. Sanketi, Patrick Tree Miller, Patrick Yin, Paul Wohlhart, Peng Xu 0010, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Rafael Rafailov, Ria Doshi, Roberto Martin Martin, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante-Gomez, Sean Kirmani, Sergey Levine, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham D. Sonawani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair 0003, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vincent Vanhoucke, Wolfram Burgard, Xiaolong Wang 0004, Xinghao Zhu, Xinyang Geng, Liangwei Xu, Yecheng Jason Ma 0001, Yejin Kim 0003, Yevgen Chebotar, Yilin Wu 0003, Yonatan Bisk, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang 0001, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zichen Jeff Cui, Zichen Zhang 0016, Zipeng Lin |
ICRA | 64 |
| 2024 | A probabilistic approach for learning and adapting shared control skills with the human in the loopabstractAssistive robots promise to be of great help to wheelchair users with motor impairments, for example for activities of daily living. Using shared control to provide task-specific assistance – for instance with the Shared Control Templates (SCT) framework – facilitates user control, even with low-dimensional input signals. However, designing SCTs is a laborious task requiring robotic expertise. To facilitate their design, we propose a method to learn one of their core components – active constraints – from demonstrated end-effector trajectories. We use a probabilistic model, Kernelized Movement Primitives, which additionally allows adaptation from user commands to improve the shared control skills, during both design and execution. We demonstrate that the SCTs so acquired can be successfully used to pick up an object, as well as adjusted for new environmental constraints, with our assistive robot EDAN. Gabriel Quere, Freek Stulp, David Filliat, João Silvério |
ICRA | 2 |
| 2024 | Fitting Parameters of Linear Dynamical Systems to Regularize Forcing Terms in Dynamical Movement PrimitivesabstractDue to their flexibility and ease of use, Dynamical Movement Primitives (DMPs) are widely used in robotics applications and research. DMPs combine linear dynamical systems to achieve robustness to perturbations and adaptation to moving targets with non-linear function approximators to fit a wide range of demonstrated trajectories.We propose a novel DMP formulation with a generalized logistic function as a delayed goal system. This formulation inherently has low initial jerk, and generates the bell-shaped velocity profiles that are typical of human movement. As the novel formulation is more expressive, it is able to fit a wide range of human demonstrations well, also without a non-linear forcing term. We exploit this increased expressiveness by automating the fitting of the dynamical system parameters through opti-mization. Our experimental evaluation demonstrates that this optimization regularizes the forcing term, and improves the interpolation accuracy of parametric DMPs. Freek Stulp, Adria Colome, Carme Torras |
ICRA | 1 |
| 2023 | Guiding Reinforcement Learning with Shared Control TemplatesabstractPurposeful interaction with objects usually requires certain constraints to be respected, e.g. keeping a bottle upright to avoid spilling. In reinforcement learning, such constraints are typically encoded in the reward function. As a consequence, constraints can only be learned by violating them. This often precludes learning on the physical robot, as it may take many trials to learn the constraints, and the necessity to violate them during the trial-and-error learning may be unsafe. We have serendipitously discovered that constraint representations for shared control - in particular Shared Control Templates (SCTs) - are ideally suited for safely guiding RL. Representing constraints explicitly, rather than implicitly in the reward function, also simplifies the design of the reward function. The main advantage of the approach is safer, faster learning without constraint violations (even with sparse reward functions). We demonstrate this in a pouring task in simulation and on a real robot, where learning the task requires only 65 episodes in 16 minutes. Abhishek Padalkar, Gabriel Quere, Franz Steinmetz, Antonin Raffin, Matthias Nieuwenhuisen, João Silvério, Freek Stulp |
ICRA | 7 |
| 2023 | ROSMC: A High-Level Mission Operation Framework for Heterogeneous Robotic TeamsabstractHeterogeneous teams of multiple mobile robots will be important for future scientific explorations of extraterrestrial surfaces or hazardous areas. Mission operation in such harsh, unknown environments poses diverse challenges. Robots need to cooperate autonomously due to the large network latency to the ground station while operators need to adapt the ongoing mission flexibly based on new discoveries obtained during execution. Furthermore, shared situational awareness between operators and roboticists is highly required to deal with execution failures promptly. To overcome these challenges, this paper proposes the high-level mission operation framework ROSMC. The concept of mission synchronization to robots enables continuous mission adaptations and future planning by operators while robots execute the mission autonomously. The ROS-based GUIs enable operators to intuitively create and monitor the mission for robots as well as to communicate with roboticists smoothly. The proposed framework was evaluated by a pilot study with a simulator and demonstrated at a Moon-analogue field on Mt. Etna in Sicily, Italy, involving 3 robots and around 70 researchers for 4 weeks. Ryo Sakagami, Sebastian G. Brunner, Andreas Dömel, Armin Wedler, Freek Stulp |
ICRA | 5 |
| 2022 | CATs: Task Planning for Shared Control of Assistive Robots with Variable AutonomyabstractFrom robotic space assistance to healthcare robotics, there is increasing interest in robots that offer adaptable levels of autonomy. In this paper, we propose an action representation and planning framework that is able to generate plans that can be executed with both shared control and supervised autonomy, even switching between them during task execution. The action representation - Constraint Action Templates (CATs) - combine the advantages of Action Templates [1] and Shared Control Templates [2]. We demonstrate that CATs enable our planning framework to generate goal-directed plans for variations of a typical task of daily living, and that users can execute them on the wheelchair-robot EDAN in shared control or in autonomous mode. Samuel Bustamante-Gomez, Gabriel Quere, Daniel Leidner, Jörn Vogel, Freek Stulp |
ICRA | 5 |
| 2022 | Multi-Phase Multi-Modal Haptic TeleoperationabstractVirtual Fixtures facilitate teleoperation, for in-stance by guiding the human operator. Developing these Virtual Fixtures in tasks with tight tolerances remains challenging. Fixtures with a high stiffness allow for more precise guidance, whereas a lower stiffness is required to allow for corrections. We observed that many assembly operations can be split into different phases - approaching, positioning, in-contact manipulation - each with different accuracy requirements. Therefore, we propose to use multi-modal fixtures, satisfying the different requirements of these phases: i.e. a position-based Trajectory Fixture for approaching and a more accurate Visual Servoing Fixture for the positioning phase. A state estimation and arbitration component ensures smooth transitions between the fixtures to provide optimal support for the operator and to achieve global availability paired with local precision at the same time. It also allows a high stiffness to be used throughout, thus achieving good guidance for all phases. The approach is validated in an application from a space scenario, consisting of the assembly of a CubeSat subsystem. The empirical results from a pilot study on this task show that our approach is faster and requires less interaction force from the operator than the baseline method. Maximilian Mühlbauer 0001, Franz Steinmetz, Freek Stulp, Thomas Hulin, Alin Albu-Schäffer |
IROS | 3 |
| 2021 | Friction Estimation for Tendon-Driven Robotic HandsabstractIn tendon-driven robotic hands, tendons are usually routed along several pulleys. The resulting friction is often substantial, and must therefore be modelled and estimated, for instance for accurate control and contact detection. Common approaches for friction estimation consider special dedicated setups, where the parameters of a static or dynamic friction model at a single contact point are determined. In this paper, we rather combine such individual friction models into an overall friction model for the entire finger. Furthermore, we propose a method for estimating the parameters of this overall model in situ, i.e. from trajectories executed on the assembled hand, avoiding the need for dedicated setups. An important component of the proposed model is a varying bias for treating friction at low velocities, allowing a simpler static friction model to be used. We demonstrate that our approach enables contacts to be detected more accurately on the DLR David hand, without additional sensors. Friedrich Lange, Martin Pfanne, Franz Steinmetz, Sebastian Wolf 0001, Freek Stulp |
ICRA | 5 |
| 2021 | Learning and Interactive Design of Shared Control TemplatesabstractControlling a robotic arm to achieve manipulation tasks is challenging for humans. Especially if only low-dimensional input signals can be provided, as is often the case for users with motor impairments. Using shared control to provide task-specific guidance and constraints facilitates control – for instance with the Shared Control Templates (SCT) framework – and enables even complex activities of daily living to be performed successfully. However, designing SCTs is a laborious task requiring robotic expertise. To make such design easier and faster, we propose a method for semi-automatically designing SCTs on the basis of demonstrations. Furthermore, we propose two similarity metrics, and demonstrate how these can be used to transfer knowledge from one SCT to another. We demonstrate that the SCTs so acquired can be successfully used in shared control for everyday tasks such as opening a drawer or a cupboard on our assistive robot EDAN. Gabriel Quere, Samuel Bustamante-Gomez, Annette Hagengruber, Jörn Vogel, Franz Steinmetz, Freek Stulp |
IROS | 6 |
| 2021 | Flexible Robotic Assembly Based on Ontological Representation of Tasks, Skills, and ResourcesabstractTechnology has sufficiently matured to enable, in principle, flexible and autonomous robotic assembly systems. However, in practice, it requires making all the relevant (implicit) knowledge that system engineers and workers have – about products to be assembled, tasks to be performed, as well as robots and their skills – available to the system explicitly. Only then can the planning and execution components of a robotic assembly pipeline communicate with each other in the same language and solve tasks autonomously without human intervention. This is why we have developed the Factory of the Future (FoF) ontology. At its core, this ontology models the tasks that are necessary to assemble a product and the robotic skills that can be employed to complete said tasks. The FoF ontology is based on existing standards. We started with theoretical considerations and iteratively adapted it based on practical experience gained from incorporating more and more components required for automated planning and assembly. Furthermore, we propose tools to extend the ontology for specific scenarios with knowledge about parts, robots, tools, and skills from various sources. The resulting scenario ontology serves us as world model for the robotic systems and other components of the assembly process. A central runtime interface to this world model provides fast and easy access to the knowledge during execution. In this work, we also show the integration of a graphical user front-end, an assembly planner, a workspace reconfigurator, and more components of the assembly pipeline that all communicate with the help of the FoF ontology. Overall, our integration of the FoF ontology with the other components of a robotic assembly pipeline shows that using an ontology is a practical method to establish a common language and understanding between the involved components. Philipp Matthias Schäfer, Franz Steinmetz, Stefan Schneyer, Timo Bachmann, Thomas Eiband, Florian Samuel Lay, Abhishek Padalkar, Christoph Sürig, Freek Stulp, Korbinian Nottensteiner |
KR | 9 |
| 2020 | Probabilistic Effect Prediction through Semantic Augmentation and Physical SimulationabstractNowadays, robots are mechanically able to perform highly demanding tasks, where AI-based planning methods are used to schedule a sequence of actions that result in the desired effect. However, it is not always possible to know the exact outcome of an action in advance, as failure situations may occur at any time. To enhance failure tolerance, we propose to predict the effects of robot actions by augmenting collected experience with semantic knowledge and leveraging realistic physics simulations. That is, we consider semantic similarity of actions in order to predict outcome probabilities for previously unknown tasks. Furthermore, physical simulation is used to gather simulated experience that makes the approach robust even in extreme cases. We show how this concept is used to predict action success probabilities and how this information can be exploited throughout future planning trials. The concept is evaluated in a series of real world experiments conducted with the humanoid robot Rollin' Justin. Adrian Simon Bauer, Peter Schmaus, Freek Stulp, Daniel Leidner |
ICRA | 3 |
| 2020 | Robust, Locally Guided Peg-in-Hole using Impedance-Controlled RobotsabstractWe present an approach for the autonomous, robust execution of peg-in-hole assembly tasks. We build on a sampling-based state estimation framework, in which samples are weighted according to their consistency with the position and joint torque measurements. The key idea is to reuse these samples in a motion generation step, where they are assigned a second task-specific weight. The algorithm thereby guides the peg towards the goal along the configuration space. An advantage of the approach is that the user only needs to provide: the geometry of the objects as mesh data, as well as a rough estimate of the object poses in the workspace, and a desired goal state. Another advantage is that the local, online nature of our algorithm leads to robust behavior under uncertainty. The approach is validated in the case of our robotic setup and under varying uncertainties for the classical peg-in-hole problem subject to two different geometries. Korbinian Nottensteiner, Freek Stulp, Alin Albu-Schäffer |
ICRA | 2 |
| 2020 | Shared Control Templates for Assistive RoboticsabstractLight-weight robotic manipulators can be used to restore the manipulation capability of people with a motor disability. However, manipulating the environment poses a complex task, especially when the control interface is of low bandwidth, as may be the case for users with impairments. Therefore, we propose a constraint-based shared control scheme to define skills which provide support during task execution. This is achieved by representing a skill as a sequence of states, with specific user command mappings and different sets of constraints being applied in each state. New skills are defined by combining different types of constraints and conditions for state transitions, in a human-readable format. We demonstrate its versatility in a pilot experiment with three activities of daily living. Results show that even complex, high-dimensional tasks can be performed with a low-dimensional interface using our shared control approach. Gabriel Quere, Annette Hagengruber, Maged Iskandar, Samuel Bustamante-Gomez, Daniel Leidner, Freek Stulp, Jörn Vogel |
ICRA | 6 |
| 2020 | Continuous, Real-Time Emotion Annotation: A Novel Joystick-Based Analysis FrameworkabstractEmotion labels are usually obtained via either manual annotation, which is tedious and time-consuming, or questionnaires, which neglect the time-varying nature of emotions and depend on human's unreliable introspection. To overcome these limitations, we developed a continuous, real-time, joystick-based emotion annotation framework. To assess the same, 30 subjects each watched 8 emotion-inducing videos. They were asked to indicate their instantaneous emotional state in a valence-arousal (V-A) space, using a joystick. Subsequently, five analyses were undertaken: (i) a System Usability Scale (SUS) questionnaire unveiled the framework's excellent usability; (ii) MANOVA analysis of the mean V-A ratings and (iii) trajectory similarity analyses of the annotations confirmed the successful elicitation of emotions; (iv) Change point analysis of the annotations, revealed a direct mapping between emotional events and annotations, thereby enabling automatic detection of emotionally salient points in the videos; and (v) Support Vector Machines (SVM) were trained on classification of 5 second chunks of annotations as well as their change-points. The classification results confirmed that ratings patterns were cohesive across the participants. These analyses confirm the value, validity, and usability of our annotation framework. They also showcase novel tools for gaining greater insights into the emotional experience of the participants. Karan Sharma, Claudio Castellini, Freek Stulp, Egon L. van den Broek |
IEEE Trans. Affect. Comput. | 3 |
| 2020 | A Survey on Policy Search Algorithms for Learning Robot Controllers in a Handful of TrialsabstractMost policy search (PS) algorithms require thousands of training episodes to find an effective policy, which is often infeasible with a physical robot. This survey article focuses on the extreme other end of the spectrum: how can a robot adapt with only a handful of trials (a dozen) and a few minutes? By analogy with the word “big-data,” we refer to this challenge as “micro-data reinforcement learning.” In this article, we show that a first strategy is to leverage prior knowledge on the policy structure (e.g., dynamic movement primitives), on the policy parameters (e.g., demonstrations), or on the dynamics (e.g., simulators). A second strategy is to create data-driven surrogate models of the expected reward (e.g., Bayesian optimization) or the dynamical model (e.g., model-based PS), so that the policy optimizer queries the model instead of the real system. Overall, all successful micro-data algorithms combine these two strategies by varying the kind of model and prior knowledge. The current scientific challenges essentially revolve around scaling up to complex robots, designing generic priors, and optimizing the computing time. Konstantinos Chatzilygeroudis, Vassilis Vassiliades, Freek Stulp, Sylvain Calinon, Jean-Baptiste Mouret |
IEEE Trans. Robotics | 3 |
| 2019 | Teleoperating Robots from the International Space Station: Microgravity Effects on Performance with Force FeedbackabstractSending humans to Mars' surface to build habitats is, for now, prohibitively dangerous and costly. An alternative is to have humans in orbiters, teleoperating robots on Mars to construct habitats, deferring human arrival until these habitats are finished. This paper describes the Kontur-2 experiments, in which the feasibility of this scenario was tested with the International Space Station as an orbiter, a cosmonaut operating a force-feedback joystick as an input device for teleoperation, and Earth as the planet where the teleoperated robot is located. In particular, we focus on human teleoperation performance, which is known to deteriorate under conditions of spaceflight. We investigate whether the provision of force feedback at the joystick is as beneficial as under terrestrial conditions. Our results show that, to support humans operating in weightlessness, haptic assistance needs to be adjusted to the altered environmental condition. Bernhard M. Weber, Ribin Balachandran, Cornelia Riecke, Freek Stulp, Martin Stelzer |
IROS | 4 |
| 2019 | Policy search in continuous action domains: An overview
Olivier Sigaud, Freek Stulp |
Neural Networks | 2 |
| 2019 | A functional data analysis approach for continuous 2-D emotion annotationsabstractThe standard paradigm in Affective Computing involves acquiring one/several markers (e.g., physiological signals) of emotions and training models on these to predict emotions. However, due to the internal nature of emotions, labelling/annotation of emotional experience is done manually by humans us ing specially developed annotation tools. To effectively exploit the resulting subjective annotations for developing affective systems, their quality needs to be assessed. This entails, (i) evaluating the variations in annotations, across different subjects and emotional stimuli, to detect spurious/unexpected patterns; and (ii) developing strategies to effectively combine these subjective annotations into a ground truth annotation. This article builds on our previous work by presenting a novel Functional Data Analysis based approach to assess the quality of annotations. Specifically, the bivariate annotation time-series are transformed into functions, such that each resulting functional annotation then becomes a sample element for analysis like Multivariate Functional Principal Component Analysis (MFPCA) that evaluate variation across all annotations. The resulting scores from MFPCA provide interesting insights into annotation patterns and facilitate the use of multivariate statistical techniques to address both (i) and (ii). Given the presented efficacy of these methods, we believe they offer an exciting new approach to assessing the quality of annotations. Karan Sharma, Marius Wagner, Claudio Castellini, Egon L. van den Broek, Freek Stulp, Friedhelm Schwenker |
Web Intell. | 5 |
| 2018 | Optimizing Contextual Ergonomics Models in Human-Robot InteractionabstractCurrent ergonomic assessment procedures require observation and manual annotation of postures by an expert, after which ergonomic scores are inferred from these annotations. Our aim is to automate this procedure and to enable robots to optimize their behavior with respect to such scores. A particular challenge is that ergonomic scoring requires accurate biomechanical simulations which are computationally too expensive to use in robot control loops or optimization. To address this, we learn Contextual Ergonomics Models, which are Gaussian Process Latent Variable Models that have been trained with full musculoskeletal simulations for specific tasks contexts. Contextual Ergonomics Models enable search in a low-dimensional latent space, whilst the cost function can be defined in terms of the full high-dimensional musculoskeletal model, which can be quickly reconstructed from the latent space. We demonstrate how optimizing Contextual Ergonomics Models leads to significantly reduced muscle activation in an experiment with eight subjects performing a drilling task. Antonio Gonzales Marin, Mohammad S. Shourijeh, Pavel E. Galibarov, Michael Damsgaard, Lars Fritzsch, Freek Stulp |
IROS | 6 |
| 2017 | Tensor Based Knowledge Transfer Across Skill Categories for Robot ControlabstractAdvances in hardware and learning for control are enabling robots to perform increasingly dextrous and dynamic control tasks. These skills typically require a prohibitive amount of exploration for reinforcement learning, and so are commonly achieved by imitation learning from manual demonstration. The costly non-scalable nature of manual demonstration has motivated work into skill generalisation, e.g., through contextual policies and options. Despite good results, existing work along these lines is limited to generalising across variants of one skill such as throwing an object to different locations. In this paper we go significantly further and investigate generalisation across qualitatively different classes of control skills. In particular, we introduce a class of neural network controllers that can realise four distinct skill classes: reaching, object throwing, casting, and ball-in-cup. By factorising the weights of the neural network, we are able to extract transferrable latent skills, that enable dramatic acceleration of learning in cross-task transfer. With a suitable curriculum, this allows us to learn challenging dextrous control tasks like ball-in-cup from scratch with pure reinforcement learning. Chenyang Zhao 0007, Timothy M. Hospedales, Freek Stulp, Olivier Sigaud |
IJCAI | 3 |
| 2015 | Co-manipulation with multiple probabilistic virtual guidesabstractIn co-manipulation, humans and robots solve manipulation tasks together. Virtual guides are important tools for co-manipulation, as they constrain the movement of the robot to avoid undesirable effects, such as collisions with the environment. Defining virtual guides is often a laborious task requiring expert knowledge. This restricts the usefulness of virtual guides in environments where new tasks may need to be solved, or where multiple tasks need to be solved sequentially, but in an unknown order. To this end, we propose a framework for multiple probabilistic virtual guides, and demonstrate a concrete implementation of such guides using kinesthetic teaching and Gaussian mixture models. Our approach enables non-expert users to design virtual guides through demonstration. Also, they may demonstrate novel guides, even if already known guides are active. Finally, users are able to intuitively select the appropriate guide from a set of guides through physical interaction with the robot. We evaluate our approach in a pick-and-place task, where users are to place objects at one of several positions in a cupboard. Gennaro Raiola, Xavier Lamy, Freek Stulp |
IROS | 3 |
| 2015 | Facilitating intention prediction for humans by optimizing robot motionsabstractMembers of a team are able to coordinate their actions by anticipating the intentions of others. Achieving such implicit coordination between humans and robots requires humans to be able to quickly and robustly predict the robot's intentions, i.e. the robot should demonstrate a behavior that is legible. Whereas previous work has sought to explicitly optimize the legibility of behavior, we investigate legibility as a property that arises automatically from general requirements on the efficiency and robustness of joint human-robot task completion. We do so by optimizing fast and successful completion of joint human-robot tasks through policy improvement with stochastic optimization. Two experiments with human subjects show that robots are able to adapt their behavior so that humans become better at predicting the robot's intentions early on, which leads to faster and more robust overall task completion. Freek Stulp, Jonathan Grizou, Baptiste Busch, Manuel Lopes 0001 |
IROS | 1 |
| 2015 | Many regression algorithms, one unified model: A review
Freek Stulp, Olivier Sigaud |
Neural Networks | 1 |
| 2014 | Simultaneous on-line Discovery and Improvement of Robotic Skill optionsabstractThe regularity of everyday tasks enables us to reuse existing solutions for task variations. For instance, most door-handles require the same basic skill (reach, grasp, turn, pull), but small adaptations of the basic skill are required to adapt to the variations that exist (e.g. levers vs. knobs). We introduce the algorithm “Simultaneous On-line Discovery and Improvement of Robotic Skills” (SODIRS) that is able to autonomously discover and optimize skill options for such task variations. We formalize the problem in a reinforcement learning context, and use the PIBB algorithm [2] to continually optimize skills with respect to a cost function. SODIRS discovers new subskills, or “skill options”, by clustering the costs of trials, and determining whether perceptual features are able to predict which cluster a trial will belong to. This enables SODIRS to build a decision tree, in which the leaves contain skill options for task variations. We demonstrate SODIRS' performance in simulation, as well as on a Meka humanoid robot performing the ball-in-cup task. Freek Stulp, Laura Herlant, Antoine Hoarau, Gennaro Raiola |
IROS | 1 |
| 2012 | Path Integral Policy Improvement with Covariance Matrix Adaptation
Freek Stulp, Olivier Sigaud |
ICML | 1 |
| 2012 | Comparing motion generation and motion recall for everyday mobile manipulation tasksabstractWhen first posed with the problem 15 × 15, we may generate the answer by applying a set of rules, e.g breaking the problem down into (10 + 5) × 15 and solving the subcomponents of these simpler multiplications first [2]. But after having solved this problem several times, we simply recall that the answer to 15 × 15 is 225. This distinction between generation and recall can also be applied to motor planning [2], as described in the next two sections. Carmen Lopera, Hilario Tome, Adolfo Rodríguez Tsouroukdissian, Freek Stulp |
IROS | 4 |
| 2012 | Adaptive exploration for continual reinforcement learningabstractMost experiments on policy search for robotics focus on isolated tasks, where the experiment is split into two distinct phases: (1) the learning phase, where the robot learns the task through exploration; (2) the exploitation phase, where exploration is turned off, and the robot demonstrates its performance on the task it has learned. In this paper, we present an algorithm that enables robots to continually and autonomously alternate between these phases. We do so by combining the `Policy Improvement with Path Integrals' direct reinforcement learning algorithm with the covariance matrix adaptation rule from the `Cross-Entropy Method' optimization algorithm. This integration is possible because both algorithms iteratively update parameters with probability-weighted averaging. A practical advantage of the novel algorithm, called PI2-CMA, is that it alleviates the user from having to manually tune the degree of exploration. We evaluate PI2-CMA's ability to continually and autonomously tune exploration on two tasks. Freek Stulp |
IROS | 1 |
| 2012 | Learning and Reasoning with Action-Related Places for Robust Mobile ManipulationabstractWe propose the concept of Action-Related Place (ARPlace) as a powerful and flexible representation of task-related place in the context of mobile manipulation. ARPlace represents robot base locations not as a single position, but rather as a collection of positions, each with an associated probability that the manipulation action will succeed when located there. ARPlaces are generated using a predictive model that is acquired through experience-based learning, and take into account the uncertainty the robot has about its own location and the location of the object to be manipulated. When executing the task, rather than choosing one specific goal position based only on the initial knowledge about the task context, the robot instantiates an ARPlace, and bases its decisions on this ARPlace, which is updated as new information about the task becomes available. To show the advantages of this least-commitment approach, we present a transformational planner that reasons about ARPlaces in order to optimize symbolic plans. Our empirical evaluation demonstrates that using ARPlaces leads to more robust and efficient mobile manipulation in the face of state estimation uncertainty on our simulated robot. Freek Stulp, Andreas Fedrizzi, Lorenz Mösenlechner, Michael Beetz |
J. Artif. Intell. Res. | 1 |
| 2012 | Reinforcement Learning With Sequences of Motion Primitives for Robust ManipulationabstractPhysical contact events often allow a natural decomposition of manipulation tasks into action phases and subgoals. Within the motion primitive paradigm, each action phase corresponds to a motion primitive, and the subgoals correspond to the goal parameters of these primitives. Current state-of-the-art reinforcement learning algorithms are able to efficiently and robustly optimize the parameters of motion primitives in very high-dimensional problems. These algorithms often consider only shape parameters, which determine the trajectory between the start- and end-point of the movement. In manipulation, however, it is also crucial to optimize the goal parameters, which represent the subgoals between the motion primitives. We therefore extend the policy improvement with path integrals (PI2) algorithm to simultaneously optimize shape and goal parameters. Applying simultaneous shape and goal learning to sequences of motion primitives leads to the novel algorithm PI2Seq. We use our methods to address a fundamental challenge in manipulation: improving the robustness of everyday pick-and-place tasks. Freek Stulp, Evangelos A. Theodorou, Stefan Schaal |
IEEE Trans. Robotics | 1 |
| 2011 | Learning to grasp under uncertaintyabstractWe present an approach that enables robots to learn motion primitives that are robust towards state estimation uncertainties. During reaching and preshaping, the robot learns to use line manipulation strategies to maneuver the object into a pose at which closing the hand to perform the grasp is more likely to succeed. In contrast, common assumptions in grasp planning and motion planning for reaching are that these tasks can be performed independently, and that the robot has perfect knowledge of the pose of the objects in the environment. We implement our approach using Dynamic Movement Primitives and the probabilistic model-free reinforcement learning algorithm Policy Improvement with Path Integrals (PI2). The cost function that PI2optimizes is a simple boolean that penalizes failed grasps. The key to acquiring robust motion primitives is to sample the actual pose of the object from a distribution that represents the state estimation uncertainty. During learning, the robot will thus optimize the chance of grasping an object from this distribution, rather than at one specific pose. In our empirical evaluation, we demonstrate how the motion primitives become more robust when grasping simple cylindrical objects, as well as more complex, non-convex objects. We also investigate how well the learned motion primitives generalize towards new object positions and other state estimation uncertainty distributions. Freek Stulp, Evangelos A. Theodorou, Jonas Buchli, Stefan Schaal |
ICRA | 1 |
| 2011 | Movement segmentation using a primitive libraryabstractSegmenting complex movements into a sequence of primitives remains a difficult problem with many applications in the robotics and vision communities. In this work, we show how the movement segmentation problem can be reduced to a sequential movement recognition problem. To this end, we reformulate the original Dynamic Movement Primitive (DMP) formulation as a linear dynamical system with control inputs. Based on this new formulation, we develop an Expectation-Maximization algorithm to estimate the duration and goal position of a partially observed trajectory. With the help of this algorithm and the assumption that a library of movement primitives is present, we present a movement segmentation framework. We illustrate the usefulness of the new DMP formulation on the two applications of online movement recognition and movement segmentation. Franziska Meier, Evangelos A. Theodorou, Freek Stulp, Stefan Schaal |
IROS | 3 |
| 2011 | Learning motion primitive goals for robust manipulationabstractApplying model-free reinforcement learning to manipulation remains challenging for several reasons. First, manipulation involves physical contact, which causes discontinuous cost functions. Second, in manipulation, the end-point of the movement must be chosen carefully, as it represents a grasp which must be adapted to the pose and shape of the object. Finally, there is uncertainty in the object pose, and even the most carefully planned movement may fail if the object is not at the expected position. To address these challenges we 1) present a simplified, computationally more efficient version of our model-free reinforcement learning algorithm PI2; 2) extend PI2so that it simultaneously learns shape parameters and goal parameters of motion primitives; 3) use shape and goal learning to acquire motion primitives that are robust to object pose uncertainty. We evaluate these contributions on a manipulation platform consisting of a 7-DOF arm with a 4-DOF hand. Freek Stulp, Evangelos A. Theodorou, Mrinal Kalakrishnan, Peter Pastor, Ludovic Righetti, Stefan Schaal |
IROS | 1 |
| 2009 | Action-related place-based mobile manipulationabstractIn mobile manipulation, the position to which the robot navigates has a large influence on the ease with which a subsequent manipulation action can be performed. Whether a manipulation action succeeds depends on many factors, such as the robot's hardware configuration, the controllers the robot uses to achieve navigation and manipulation, the task context, and uncertainties in state estimation. In this paper, we present `ARPLACE', an action-related place which takes these factors, and the context in which the actions are performed into account. Through experience-based learning, the robot first learns a so-called generalized success model, which discerns between positions from which manipulation succeeds or fails. On-line, this model is used to compute a ARPLACE, a probability distribution that maps positions to a predicted probability of successful manipulation, and takes the uncertainty in the robot and object's position into account. In an empirical evaluation, we demonstrate that using ARPLACEs for least-commitment navigation improves the success rate of subsequent manipulation tasks substantially. Freek Stulp, Andreas Fedrizzi, Michael Beetz |
IROS | 1 |
| 2008 | A real time system for model-based interpretation of the dynamics of facial expressionsabstractOur system runs at 10 fps on a 2.0 GHz processor and an image resolution of 640times480 pixels. High quality objective functions that are learned from annotated example images ensure both an accurate and fast computation of the model parameters. Our demonstrator for facial expression estimation has been presented at several events with political audience and on TV. However, the approach of robust face models fitting, forms the basis of various more applications such as gaze detection or gender estimation. The drawback of our approach is that the data base from which the objective function is learned needs to cover all aspects of face properties. If, for instance, the database did not contain images of bearded men the objective function will fail when confronted with such an image. Furthermore, the data base has to be manually annotated. Although no expert knowledge is required, this task requires a considerable amount of time. An online fitting demonstration is available. Christoph Mayer 0001, Matthias Wimmer, Freek Stulp, Zahid Riaz, Anton Roth, Martin Eggers, Bernd Radig |
FG | 3 |
| 2008 | An ASM fitting method based on machine learning that provides a robust parameter initialization for AAM fittingabstractDue to their use of information contained in texture, active appearance models (AAM) generally outperform active shape models (ASM) in terms of fitting accuracy. Although many extensions and improvements over the original AAM have been proposed, on of the main drawbacks of AAMs remains its dependence on good initial model parameters to achieve accurate fitting results. In this paper, we determine the initial model parameters for AAM fitting with ASM fitting, and use machine learning techniques to improve the scope and accuracy of ASM fitting. Combining the precision of AAM fitting with the large radius of convergence of learned ASM fitting improves the results by an order of magnitude, as our empirical evaluation on a database of publicly available benchmark images demonstrates. Matthias Wimmer, Shinya Fujie, Freek Stulp, Tetsunori Kobayashi, Bernd Radig |
FG | 3 |
| 2008 | The Assistive Kitchen - A demonstration scenario for cognitive technical systemsabstractThis paper introduces the assistive kitchen as a comprehensive demonstration and challenge scenario for technical cognitive systems. We describe its hardware and software infrastructure. Within the assistive kitchen application, we select particular domain activities as research subjects and identify the cognitive capabilities needed for perceiving, interpreting, analyzing, and executing these activities as research foci. We conclude by outlining open research issues that need to be solved to realize the scenarios successfully. Michael Beetz, Freek Stulp, Bernd Radig, Jan Bandouch, Nico Blodow, Mihai Emanuel Dolha, Andreas Fedrizzi, Dominik Jain, Ulrich Klank, Ingo Kresse, Alexis Maldonado, Zoltan-Csaba Marton, Lorenz Mösenlechner, Federico Ruiz-Ugalde, Radu Bogdan Rusu, Moritz Tenorth |
RO-MAN | 2 |
| 2008 | Refining the Execution of Abstract Actions with Learned Action ModelsabstractRobots reason about abstract actions, such as "go to position `l'", in order to decide what to do or to generate plans for their intended course of action. The use of abstract actions enables robots to employ small action libraries, which reduces the search space for decision making. When executing the actions, however, the robot must tailor the abstract actions to the specific task and situation context at hand. In this article we propose a novel robot action execution system that learns success and performance models for possible specializations of abstract actions. At execution time, the robot uses these models to optimize the execution of abstract actions to the respective task contexts. The robot can so use abstract actions for efficient reasoning, without compromising the performance of action execution. We show the impact of our action execution model in three robotic domains and on two kinds of action execution problems: (1) the instantiation of free action parameters to optimize the expected performance of action sequences; (2) the automatic introduction of additional subgoals to make action sequences more reliable. Freek Stulp, Michael Beetz |
J. Artif. Intell. Res. | 1 |
| 2008 | Learning Local Objective Functions for Robust Face Model FittingabstractModel-based techniques have proven to be successful in interpreting the large amount of information contained in images. Associated fitting algorithms search for the global optimum of an objective function, which should correspond to the best model fit in a given image. Although fitting algorithms have been the subject of intensive research and evaluation, the objective function is usually designed ad hoc, based on implicit and domain-dependent knowledge. In this article, we address the root of the problem by learning more robust objective functions. First, we formulate a set of desirable properties for objective functions and give a concrete example function that has these properties. Then, we propose a novel approach that learns an objective function from training data generated by manual image annotations and this ideal objective function. In this approach, critical decisions such as feature selection are automated, and the remaining manual steps hardly require domain-dependent knowledge. Furthermore, an extensive empirical evaluation demonstrates that the obtained objective functions yield more robustness. Learned objective functions enable fitting algorithms to determine the best model fit more accurately than with designed objective functions. Matthias Wimmer, Freek Stulp, Sylvia Pietzsch, Bernd Radig |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Enabling Users to Guide the Design of Robust Model Fitting AlgorithmsabstractModel-based image interpretation extracts high-level information from images using a priori knowledge about the object of interest. The computational challenge in model fitting is to determine the model parameters that best match a given image, which corresponds to finding the global optimum of the objective function. When it comes to the robustness and accuracy of fitting models to specific images, humans still outperform state- of-the-art model fitting systems. Therefore, we propose a method in which non-experts can guide the process of designing model fitting algorithms. In particular, this paper demonstrates how to obtain robust objective functions for face model fitting applications, by learning their calculation rules from example images annotated by humans. We evaluate the obtained function using a publicly available image database and compare it to a recent state-of-the-art approach in terms of accuracy. Matthias Wimmer, Freek Stulp, Bernd Radig |
ICCV | 2 |
| 2007 | Seamless Execution of Action SequencesabstractOne of the most notable and recognizable features of robot motion is the abrupt transitions between actions in action sequences. In contrast, humans and animals perform sequences of actions efficiently, and with seamless transitions between subsequent actions. This smoothness is not a goal in itself, but a side-effect of the evolutionary optimization of other performance measures. In this paper, we argue that such jagged motion is an inevitable consequence of the way human designers and planners reason about abstract actions. We then present subgoal refinement, a procedure that optimizes action sequences. Sub-goal refinement determines action parameters that are not relevant to why the action was selected, and optimizes these parameters with respect to expected execution performance. This performance is computed using action models, which are learned from observed experience. We integrate subgoal refinement in an existing planning system, and demonstrate how requiring optimal performance causes smooth motion in three robotic domains. Freek Stulp, Wolfram Koska, Alexis Maldonado, Michael Beetz |
ICRA | 1 |
| 2006 | Learning Robust Objective Functions for Model Fitting in Image Understanding ApplicationsabstractModel-based methods in computer vision have proven to be a good approach for compressing the large amount of information in images. Fitting algorithms search for those parameters of the model that optimise the objective function given a certain image. Although fitting algorithms have been the subject of intensive research and evaluation, the objective function is usually designed ad hoc and heuristically with much implicit domain-dependent knowledge. This paper formulates a set of requirements that robust objective functions should satisfy. Furthermore, we propose a novel approach that learns the objective function from training images that have been annotated with the preferred model parameters. The requirements are automatically enforced during the learning phase, which yields generally applicable objective functions. We compare the performance of our approach to other approaches. For this purpose, we propose a set of indicators that evaluate how well an objective function meets the stated requirements. 1 Matthias Wimmer, Freek Stulp, Stephan Tschechne, Bernd Radig |
BMVC | 2 |
| 2006 | Implicit Coordination in Robotic Teams using Learned Prediction ModelsabstractMany application tasks require the cooperation of two or more robots. Humans are good at cooperation in shared workspaces, because they anticipate and adapt to the intentions and actions of others. In contrast, multi-agent and multi-robot systems rely on communication to exchange their intentions. This causes problems in domains where perfect communication is not guaranteed, such as rescue robotics, autonomous vehicles participating in traffic, or robotic soccer. In this paper, we introduce a computational model for implicit coordination, and apply it to a typical coordination task from robotic soccer: regaining ball possession. The computational model specifies that performance prediction models are necessary for coordination, so we learn them off-line from observed experience. By taking the perspective of the team mates, these models are then used to predict utilities of others, and optimize a shared performance model for joint actions. In several experiments conducted with our robotic soccer team, we evaluate the performance of implicit coordination Freek Stulp, Michael Isik, Michael Beetz |
ICRA | 1 |
| 2006 | Coordination Without Negotiation in Teams of Heterogeneous Robots
Michael Isik, Freek Stulp, Gerd Mayer, Hans Utz |
RoboCup | 2 |
| 2005 | Optimized Execution of Action Chains Using Learned Performance Models of Abstract Actions
Freek Stulp, Michael Beetz |
IJCAI | 1 |
| 2004 | Sharing Belief in Teams of Heterogeneous Robots
Hans Utz, Freek Stulp, Arndt Mühlenfeld |
RoboCup | 2 |
| 2003 | Autonomous Robot Controllers Capable of Acquiring Repertoires of Complex Skills
Michael Beetz, Freek Stulp, Alexandra Kirsch, Armin Müller, Sebastian Buck 0001 |
RoboCup | 2 |
| 2002 | Machine control using radial basis value functions and inverse state projectionabstractTypical real world machine control tasks have some characteristics which makes them difficult to solve: Their state spaces are high-dimensional and continuous, and it may be impossible to reach a satisfying target state by exploration or human control. To overcome these problems, in this paper, we propose (1) to use radial basis functions for value function approximation in continuous space reinforcement learning and (2) the use of learned inverse projection functions for state space exploration. We apply our approach to path planning in dynamic environments and to an aircraft autolanding simulation, and evaluate its performance. Sebastian Buck 0001, Freek Stulp, Michael Beetz, Thorsten Schmitt |
ICARCV | 2 |