EDBT 2026 Demo / reviewers in the wild / expert
Joschka Boedecker
dblp:84/5457 · also Joschka Bödecker
· DBLP profile ↗
36ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0002-3486-7345ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 2 first-author · 12 since 2021Systems, architecture and hardware · 12 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Gaussian process-based motor hotspot hunting with concurrent optimization of TMS coil location and orientationabstractTranscranial magnetic stimulation (TMS) is a widely used non-invasive brain stimulation technique in neuroscience research and clinical applications. TMS-based motor hotspot hunting describes the process of identifying the optimal scalp location to elicit robust and reliable motor responses. It is critical to ensure reproducibility of TMS parameters, as well as to determine safe and precise stimulation intensities in both healthy participants and patients. Typically, this process targets motor responses in contralateral short hand muscles. However, hotspot hunting remains challenging due to the vast parameter space and time constraints. To address this, we present an approach that concurrently optimizes both spatial and angular TMS parameters for hotspot hunting using Gaussian processes and Bayesian optimization. We systematically evaluated five state-of-the-art acquisition functions on electromyographic TMS data from eight healthy individuals enhanced by simulated data from generative models. Our results consistently demonstrate that optimizing spatial and angular TMS parameters simultaneously enhances the efficacy and spatial precision of hotspot hunting. Furthermore, we provide mechanistic insights into the acquisition function behavior and the impact of coil rotation constraints, revealing critical limitations in current hotspot-hunting strategies. Specifically, we show that arbitrary constraints on coil rotation angle are suboptimal, as they reduce flexibility and fail to account for individual variability. We further demonstrate that acquisition functions differ in sampling strategies and performance. Functions overly emphasizing exploitation tend to converge prematurely to local optima, whereas those balancing exploration and exploitation-particularly Thompson sampling-achieve superior performance. These findings highlight the importance of acquisition function selection and the necessity of removing restrictive coil rotation constraints for effective hotspot hunting. Our work advances TMS-based hotspot identification, potentially reducing participant burden and improving safety in both research and clinical applications beyond the motor cortex. David Luis Schultheiss, Zsolt Turi, Joschka Boedecker, Andreas Vlachos 0002 |
PLoS Comput. Biol. | 3 |
| 2026 | The Unreasonable Effectiveness of Discrete-Time Gaussian Process Mixtures for Robot Policy LearningabstractWe present Mixture of Discrete-time Gaussian Processes (MiDiGaP), a novel approach for flexible policy representation and imitation learning in robot manipulation. MiDiGaP enables learning from as few as five demonstrations using only camera observations and generalizes across a wide range of challenging tasks. It excels at long-horizon behaviors such as making coffee, highly constrained motions such as opening doors, dynamic actions such as scooping with a spatula, and multimodal tasks such as hanging a mug. MiDiGaP learns these tasks on a CPU in less than a minute and scales linearly to large datasets. We also develop a rich suite of tools for inferencetime steering using evidence such as collision signals and robot kinematic constraints. This steering enables novel generalization capabilities, including obstacle avoidance and cross-embodiment policy transfer. MiDiGaP achieves state-of-the-art performance on diverse few-shot manipulation benchmarks. On constrained RLBench tasks, it improves policy success by 76 percentage points and reduces trajectory cost by 67%. On multimodal tasks, it improves policy success by 48 percentage points while improving sample efficiency 7-fold. In cross-embodiment transfer, it more than doubles policy success. We make the code publicly available. Jan Ole von Hartz, Adrian Röfer, Joschka Boedecker, Abhinav Valada |
IEEE Trans. Robotics | 3 |
| 2025 | Salvage: Shapley-distribution Approximation Learning Via Attribution Guided Exploration for Explainable Image ClassificationabstractThe integration of deep learning into critical vision application areas has given rise to a necessity for techniques that can explain the rationale behind predictions. In this paper, we address this need by introducing Salvage, a novel removal-based explainability method for image classification. Our approach involves training an explainer model that learns the prediction distribution of the classifier on masked images. We first introduce the concept of Shapley-distributions, which offers a more accurate approximation of classification probability distributions than existing methods. Furthermore, we address the issue of unbalanced important and unimportant features. In such settings, naive uniform sampling of feature subsets often results in a highly unbalanced ratio of samples with high and low prediction likelihoods, which can hinder effective learning. To mitigate this, we propose an informed sampling strategy that leverages approximated feature importance scores, thereby reducing imbalance and facilitating the estimation of underrepresented features. After incorporating these two principles into our method, we conducted an extensive analysis on the ImageNette, MURA, WBC, and Pet datasets. The results show that Salvage outperforms various baseline explainability methods, including attention-, gradient-, and removal-based approaches, both qualitatively and quantitatively. Furthermore, we demonstrate that our explainer model can serve as a fully explainable classifier without a major decrease in classification performance, paving the way for fully explainable image classification. Mehdi Naouar, Hanne Raum, Jens Rahnfeld, Yannick Vogt, Joschka Boedecker, Gabriel Kalweit, Maria Kalweit |
ICLR | 5 |
| 2025 | One For All: A Unified Approach to Classification and Self-explanation
Mehdi Naouar, Yannick Vogt, Joschka Boedecker, Gabriel Kalweit, Maria Kalweit |
MICCAI (14) | 3 |
| 2024 | UDUC: An Uncertainty-Driven Approach for Learning-Based Robust ControlabstractLearning-based techniques have become popular in both model predictive control (MPC) and reinforcement learning (RL). Probabilistic ensemble (PE) models offer a promising approach for modelling system dynamics, showcasing the ability to capture uncertainty and scalability in high-dimensional control scenarios. However, PE models are susceptible to mode collapse, resulting in non-robust control when faced with environments slightly different from the training set. In this paper, we introduce the uncertainty-driven robust control (UDUC) loss as an alternative objective for training PE models, drawing inspiration from contrastive learning. We analyze the robustness of the UDUC loss through the lens of robust optimization and evaluate its performance on the challenging real-world reinforcement learning (RWRL) benchmark, which involves significant environmental mismatches between the training and testing environments. Yuan Zhang 0027, Jasper Hoffmann, Joschka Boedecker |
ECAI | 3 |
| 2024 | Improving the Efficiency and Efficacy of Multi-Agent Reinforcement Learning on Complex Railway Networks with a Local-Critic ApproachabstractThe complex railway network is a challenging real-world multi-agent system usually involving thousands of agents. Current planning methods heavily depend on expert knowledge to formulate solutions for specific cases and are therefore hardly generalized to new scenarios, on which multi-agent reinforcement learning (MARL) draws significant attention. Despite some successful applications in multi-agent decision-making tasks, MARL is hard to scale to a large number of agents. This paper rethinks the curse of agents in the centralized-training-decentralized-execution (CTDE) paradigm and proposes a local-critic approach to address the issue. By combining the local critic with the PPO algorithm, we design a deep MARL algorithm denoted as local-critic PPO (LCPPO). In experiments, we evaluate the effectiveness of LCPPO on a complex railway network benchmark, Flatland, with various numbers of agents. Noticeably, LCPPO shows prominent generalizability and robustness under the changes of environments. Yuan Zhang 0027, Umashankar Deekshith, Joschka Boedecker |
ICAPS | 4 |
| 2024 | Learning Continuous Control with Geometric Regularity from Robot Intrinsic SymmetryabstractGeometric regularity, which leverages data symmetry, has been successfully incorporated into deep learning architectures such as CNNs, RNNs, GNNs, and Transformers. While this concept has been widely applied in robotics to address the curse of dimensionality when learning from high-dimensional data, the inherent reflectional and rotational symmetry of robot structures has not been adequately explored. Drawing inspiration from cooperative multi-agent reinforcement learning, we introduce novel network structures for single-agent control learning that explicitly capture these symmetries. Moreover, we investigate the relationship between the geometric prior and the concept of Parameter Sharing in multi-agent reinforcement learning. Last but not the least, we implement the proposed framework in online and offline learning methods to demonstrate its ease of use. Through experiments conducted on various challenging continuous control tasks on simulators and real robots, we highlight the significant potential of the proposed geometric regularity in enhancing robot learning capabilities. Shengchao Yan, Baohe Zhang, Yuan Zhang 0027, Joschka Boedecker, Wolfram Burgard |
ICRA | 4 |
| 2024 | Safe Imitation Learning of Nonlinear Model Predictive Control for Flexible RobotsabstractFlexible robots may overcome some of the industry’s major challenges, such as enabling intrinsically safe human-robot collaboration and achieving a higher payload-to-mass ratio. However, controlling flexible robots is complicated due to their complex dynamics, which include oscillatory behavior and a high-dimensional state space. Nonlinear model predictive control (NMPC) offers an effective means to control such robots, but its significant computational demand often limits its application in real-time scenarios. To enable fast control of flexible robots, we propose a framework for a safe approximation of NMPC using imitation learning and a predictive safety filter. Our framework significantly reduces computation time while incurring a slight loss in performance. Compared to NMPC, our framework shows more than an eightfold improvement in computation time when controlling a three-dimensional flexible robot arm in simulation, all while guaranteeing safety constraints. Notably, our approach out-performs state-of-the-art reinforcement learning methods. The development of fast and safe approximate NMPC holds the potential to accelerate the adoption of flexible robots in industry. The project code is available at: tinyurl.com/anmpc4fr Shamil Mamedov, Rudolf Reiter, Seyed Mahdi B. Azad, Ruan Viljoen, Joschka Boedecker, Moritz Diehl, Jan Swevers |
IROS | 5 |
| 2024 | The Surprising Ineffectiveness of Pre-Trained Visual Representations for Model-Based Reinforcement LearningabstractVisual Reinforcement Learning (RL) methods often require extensive amounts of data. As opposed to model-free RL, model-based RL (MBRL) offers a potential solution with efficient data utilization through planning. Additionally, RL lacks generalization capabilities for real-world tasks. Prior work has shown that incorporating pre-trained visual representations (PVRs) enhances sample efficiency and generalization. While PVRs have been extensively studied in the context of model-free RL, their potential in MBRL remains largely unexplored. In this paper, we benchmark a set of PVRs on challenging control tasks in a model-based RL setting. We investigate the data efficiency, generalization capabilities, and the impact of different properties of PVRs on the performance of model-based agents. Our results, perhaps surprisingly, reveal that for MBRL current PVRs are not more sample efficient than learning representations from scratch, and that they do not generalize better to out-of-distribution (OOD) settings. To explain this, we analyze the quality of the trained dynamics model. Furthermore, we show that data diversity and network architecture are the most important contributors to OOD generalization performance. Robert Krug 0002, Narunas Vaskevicius, Luigi Palmieri, Joschka Boedecker |
NeurIPS | 5 |
| 2023 | Survey on LiDAR Perception in Adverse Weather ConditionsabstractAutonomous vehicles rely on a variety of sensors to gather information about their surrounding. The vehicle’s behavior is planned based on the environment perception, making its reliability crucial for safety reasons. The active LiDAR sensor is able to create an accurate 3D representation of a scene, making it a valuable addition for environment perception for autonomous vehicles. Due to light scattering and occlusion, the LiDAR’s performance change under adverse weather conditions like fog, snow or rain. This limitation recently fostered a large body of research on approaches to alleviate the decrease in perception performance. In this survey, we gathered, analyzed, and discussed different aspects on dealing with adverse weather conditions in LiDAR-based environment perception. We address topics such as the availability of appropriate data, raw point cloud processing and denoising, robust perception algorithms and sensor fusion to mitigate adverse weather induced shortcomings. We furthermore identify the most pressing gaps in the current literature and pinpoint promising research directions. Mariella Dreissig, Dominik Scheuble, Florian Piewak, Joschka Boedecker |
IV | 4 |
| 2023 | Patient groups in Rheumatoid arthritis identified by deep learning respond differently to biologic or targeted synthetic DMARDsabstractCycling of biologic or targeted synthetic disease modifying antirheumatic drugs (b/tsDMARDs) in rheumatoid arthritis (RA) patients due to non-response is a problem preventing and delaying disease control. We aimed to assess and validate treatment response of b/tsDMARDs among clusters of RA patients identified by deep learning. We clustered RA patients clusters at first-time b/tsDMARD (cohort entry) in the Swiss Clinical Quality Management in Rheumatic Diseases registry (SCQM) [1999-2018]. We performed comparative effectiveness analyses of b/tsDMARDs (ref. adalimumab) using Cox proportional hazard regression. Within 15 months, we assessed b/tsDMARD stop due to non-response, and separately a ≥20% reduction in DAS28-esr as a response proxy. We validated results through stratified analyses according to most distinctive patient characteristics of clusters. Clusters comprised between 362 and 1481 patients (3516 unique patients). Stratified (validation) analyses confirmed comparative effectiveness results among clusters: Patients with ≥2 conventional synthetic DMARDs and prednisone at b/tsDMARD initiation, male patients, as well as patients with a lower disease burden responded better to tocilizumab than to adalimumab (hazard ratio [HR] 5.46, 95% confidence interval [CI] [1.76-16.94], and HR 8.44 [3.43-20.74], and HR 3.64 [2.04-6.49], respectively). Furthermore, seronegative women without use of prednisone at b/tsDMARD initiation as well as seropositive women with a higher disease burden and longer disease duration had a higher risk of non-response with golimumab (HR 2.36 [1.03-5.40] and HR 5.27 [2.10-13.21], respectively) than with adalimumab. Our results suggest that RA patient clusters identified by deep learning may have different responses to first-line b/tsDMARD. Thus, it may suggest optimal first-line b/tsDMARD for certain RA patients, which is a step forward towards personalizing treatment. However, further research in other cohorts is needed to verify our results. Maria Kalweit, Andrea M. Burden, Joschka Boedecker, Thomas Hügle, Theresa Burkard |
PLoS Comput. Biol. | 3 |
| 2022 | Affordance Learning from Play for Sample-Efficient Policy LearningabstractRobots operating in human-centered environments should have the ability to understand how objects function: what can be done with each object, where this interaction may occur, and how the object is used to achieve a goal. To this end, we propose a novel approach that extracts a self-supervised visual affordance model from human teleoperated play data and leverages it to enable efficient policy learning and motion planning. We combine model-based planning with model-free deep reinforcement learning (RL) to learn policies that favor the same object regions favored by people, while requiring minimal robot interactions with the environment. We evaluate our algorithm, Visual Affordance-guided Policy Optimization (VAPO), with both diverse simulation manipulation tasks and real world robot tidy-up experiments to demonstrate the effectiveness of our affordance-guided policies. We find that our policies train 4 × faster than the baselines and generalize better to novel objects because our visual affordance model can anticipate their affordance regions. Jessica Borja-Diaz, Oier Mees, Gabriel Kalweit, Lukás Hermann, Joschka Boedecker, Wolfram Burgard |
ICRA | 5 |
| 2022 | Deep Surrogate Q-Learning for Autonomous DrivingabstractOpen challenges for deep reinforcement learning systems are their adaptivity to changing environments and their efficiency w.r.t. computational resources and data. In the application of learning lane-change behavior for autonomous driving, the number of required transitions imposes a bottleneck, since test drivers cannot perform an arbitrary amount of lane changes in the real world. In the off-policy setting, additional information on solving the task can be gained by observing actions from others. While in the classical RL setup this knowledge remains unused, we use other drivers as surrogates to learn the agent's value function more efficiently. We propose Surrogate Q-learning that deals with the aforementioned problems and reduces the required driving time drastically. We further propose an efficient implementation based on a permutation equivariant deep neural network architecture of the Q-function to estimate action-values for a variable number of vehicles in sensor range. We evaluate our method in the open traffic simulator SUMO and learn well performing driving policies on the real highD dataset. Maria Kalweit, Gabriel Kalweit, Moritz Werling, Joschka Boedecker |
ICRA | 4 |
| 2021 | Amortized Q-learning with Model-based Action Proposals for Autonomous Driving on HighwaysabstractWell-established optimization-based methods can guarantee an optimal trajectory for a short optimization horizon, typically no longer than a few seconds. As a result, choosing the optimal trajectory for this short horizon may still result in a sub-optimal long-term solution. At the same time, the resulting short-term trajectories allow for effective, comfortable and provable safe maneuvers in a dynamic traffic environment. In this work, we address the question of how to ensure an optimal long-term driving strategy, while keeping the benefits of classical trajectory planning. We introduce a Reinforcement Learning based approach that coupled with a trajectory planner, learns an optimal long-term decision-making strategy for driving on highways. By online generating locally optimal maneuvers as actions, we balance between the infinite low-level continuous action space, and the limited flexibility of a fixed number of predefined standard lane-change actions. We evaluated our method on realistic scenarios in the open-source traffic simulator SUMO and were able to achieve better performance than the 4 benchmark approaches we compared against, including a random action selecting agent, greedy agent, high-level, discrete actions agent and an IDM-based SUMO-controlled agent. Branka Mirchevska, Maria Hügle, Gabriel Kalweit, Moritz Werling, Joschka Boedecker |
ICRA | 5 |
| 2021 | Q-learning with Long-term Action-space Shaping to Model Complex Behavior for Autonomous Lane ChangesabstractIn autonomous driving applications, reinforcement learning agents often have to perform complex behavior, which can translate into optimizing multiple objectives while following certain rules. Encoding traffic rules and desires such as safety and comfort via classical methods based on reward shaping (i.e. a weighted combination of different objectives in the reward signal) or Lagrangian methods (including auxiliary losses in the optimization) can be very hard and cumbersome. In this work, we propose to instead shape the action-space at the maximization step of Q-learning. We further introduce a formulation for fixed-horizon estimation of auxiliary costs under the current target-policy based on truncated value- functions to encode the desire of comfortable driving ensuring interpretable behavior. We compare our algorithm to reward shaping and Lagrangian methods in the application of high- level decision making in autonomous driving, considering rules for safety, keeping right and comfort. We train and evaluate our agent in the open-source simulator SUMO on a variety of scenarios with different driver types and traffic situations. Additionally, we apply our method on the real HighD data set, showing the real-world applicability and simplicity of Q- learning with Action-space Shaping. Gabriel Kalweit, Maria Hügle, Moritz Werling, Joschka Boedecker |
IROS | 4 |
| 2021 | Residual Feedback Learning for Contact-Rich Manipulation Tasks with UncertaintyabstractWhile classic control theory offers state of the art solutions in many problem scenarios, it is often desired to improve beyond the structure of such solutions and surpass their limitations. To this end, residual policy learning (RPL) offers a formulation to improve existing controllers with reinforcement learning (RL) by learning an additive "residual" to the output of a given controller. However, the applicability of such an approach highly depends on the structure of the controller. Often, internal feedback signals of the controller limit an RL algorithm to adequately change the policy and, hence, learn the task. We propose a new formulation that addresses these limitations by also modifying the feedback signals to the controller with an RL policy and show superior performance of our approach on a contact-rich peg-insertion task under position and orientation uncertainty. In addition, we use a recent Cartesian impedance control architecture as the control framework which can be available to us as a black-box while assuming no knowledge about its input/output structure, and show the difficulties of standard RPL. Furthermore, we introduce an adaptive curriculum for the given task to gradually increase the task difficulty in terms of position and orientation uncertainty. A video showing the results can be found at https://youtu.be/SAZm_Krze7U. Alireza Ranjbar, Ngo Anh Vien, Hanna Carolin Ziesche, Joschka Boedecker, Gerhard Neumann |
IROS | 4 |
| 2020 | Dynamic Interaction-Aware Scene Understanding for Reinforcement Learning in Autonomous DrivingabstractThe common pipeline in autonomous driving systems is highly modular and includes a perception component which extracts lists of surrounding objects and passes these lists to a high-level decision component. In this case, leveraging the benefits of deep reinforcement learning for high-level decision making requires special architectures to deal with multiple variable-length sequences of different object types, such as vehicles, lanes or traffic signs. At the same time, the architecture has to be able to cover interactions between traffic participants in order to find the optimal action to be taken. In this work, we propose the novel Deep Scenes architecture, that can learn complex interaction-aware scene representations based on extensions of either 1) Deep Sets or 2) Graph Convolutional Networks. We present the Graph-Q and DeepScene-Q off-policy reinforcement learning algorithms, both outperforming state-ofthe-art methods in evaluations with the publicly available traffic simulator SUMO. Maria Hügle, Gabriel Kalweit, Moritz Werling, Joschka Boedecker |
ICRA | 4 |
| 2020 | Learning Human-Aware Robot Navigation from Physical Interaction via Inverse Reinforcement LearningabstractAutonomous systems, such as delivery robots, are increasingly employed in indoor spaces to carry out activities alongside humans. This development poses the question of how robots can carry out their tasks while, at the same time, behaving in a socially compliant manner. Further, humans need to be able to communicate their preferences in a simple and intuitive way, and robots should adapt their behavior accordingly. This paper investigates force control as a natural means to interact with a mobile robot by pushing it along the desired trajectory. We employ inverse reinforcement learning (IRL) to learn from human interaction and adapt the robot behavior to its users' preferences, thereby eliminating the need to program the desired behavior manually. We evaluate our approach in a real-world experiment where test subjects interact with an autonomously navigating robot in close proximity. The results suggest that force control presents an intuitive means to interact with a mobile robot and show that our robot can quickly adapt to the test subjects' personal preferences. Marina Kollmitz, Torsten Koller, Joschka Boedecker, Wolfram Burgard |
IROS | 3 |
| 2020 | Deep Inverse Q-learning with ConstraintsabstractPopular Maximum Entropy Inverse Reinforcement Learning approaches require the computation of expected state visitation frequencies for the optimal policy under an estimate of the reward function. This usually requires intermediate value estimation in the inner loop of the algorithm, slowing down convergence considerably. In this work, we introduce a novel class of algorithms that only needs to solve the MDP underlying the demonstrated behavior once to recover the expert policy. This is possible through a formulation that exploits a probabilistic behavior assumption for the demonstrations within the structure of Q-learning. We propose Inverse Action-value Iteration which is able to fully recover an underlying reward of an external agent in closed-form analytically. We further provide an accompanying class of sampling-based variants which do not depend on a model of the environment. We show how to extend this class of algorithms to continuous state-spaces via function approximation and how to estimate a corresponding action-value function, leading to a policy as close as possible to the policy of the external agent, while optionally satisfying a list of predefined hard constraints. We evaluate the resulting algorithms called Inverse Action-value Iteration, Inverse Q-learning and Deep Inverse Q-learning on the Objectworld benchmark, showing a speedup of up to several orders of magnitude compared to (Deep) Max-Entropy algorithms. We further apply Deep Constrained Inverse Q-learning on the task of learning autonomous lane-changes in the open-source simulator SUMO achieving competent driving after training on data corresponding to 30 minutes of demonstrations. Gabriel Kalweit, Maria Hügle, Moritz Werling, Joschka Boedecker |
NeurIPS | 4 |
| 2019 | Multimodal Spatio-Temporal Information in End-to-End Networks for Automotive Steering PredictionabstractWe study the end-to-end steering problem using visual input data from an onboard vehicle camera. An empirical comparison between spatial, spatio-temporal and multimodal models is performed assessing each concept's performance from two points of evaluation. First, how close the model is in predicting and imitating a real-life driver's behavior, second, the smoothness of the predicted steering command. The latter is a newly proposed metric. Building on our results, we propose a new recurrent multimodal model. The suggested model has been tested on a custom dataset recorded by BMW, as well as the public dataset provided by Udacity. Results show that it outperforms previously released scores. Further, a steering correction concept from off-lane driving through the inclusion of correction frames is presented. We show that our suggestion leads to promising results empirically. Mohamed Abou-Hussein, Stefan H. Müller-Weinfurtner, Joschka Boedecker |
ICRA | 3 |
| 2019 | Dynamic Input for Deep Reinforcement Learning in Autonomous DrivingabstractIn many real-world decision making problems, reaching an optimal decision requires taking into account a variable number of objects around the agent. Autonomous driving is a domain in which this is especially relevant, since the number of cars surrounding the agent varies considerably over time and affects the optimal action to be taken. Classical methods that process object lists can deal with this requirement. However, to take advantage of recent high-performing methods based on deep reinforcement learning in modular pipelines, special architectures are necessary. For these, a number of options exist, but a thorough comparison of the different possibilities is missing. In this paper, we elaborate limitations of fully-connected neural networks and other established approaches like convolutional and recurrent neural networks in the context of reinforcement learning problems that have to deal with variable sized inputs. We employ the structure of Deep Sets in off-policy reinforcement learning for high-level decision making, highlight their capabilities to alleviate these limitations, and show that Deep Sets not only yield the best overall performance but also offer better generalization to unseen situations than the other approaches. Maria Hügle, Gabriel Kalweit, Branka Mirchevska, Moritz Werling, Joschka Boedecker |
IROS | 5 |
| 2019 | Adaptive long-term control of biological neural networks with Deep Reinforcement Learning
Jan Wülfing, Sreedhar S. Kumar, Joschka Boedecker, Martin A. Riedmiller, Ulrich Egert |
Neurocomputing | 3 |
| 2018 | Controlling biological neural networks with deep reinforcement learning
Jan Wülfing, Sreedhar S. Kumar, Joschka Boedecker, Martin A. Riedmiller, Ulrich Egert |
ESANN | 3 |
| 2018 | Prediction of electrocardiography features points using seismocardiography data: a machine learning approachabstractSeismocardiography (SCG) is a cardiac diagnostic method which evaluates vibrations of the chest via accelerometers. Current prototypical wearable SCG sensor patches are combined with electrocardiography (ECG) sensors which enables the estimation of cardiac timing intervals such as pre-ejection period (PEP). We envision a system with only SCG sensors to increase the battery run time and wearing comfort. Therefore, we present a method to predict the timing of ECG R-peak based on SCG data. The developed machine learning (ML) algorithms were evaluated with respect to algorithm complexity, body postures, and personalization schemes, on a dataset consisting of 10 subjects. Our findings show that R-peaks can be predicted with a mean error of 4:16 ms for supine, 9:43 ms for sitting, and 14:3 ms for standing posture. While personalization results only in minor improvements, using more complex ML classifier is beneficial for standing posture, which is more affected by motion artifacts. Robert Dürichen, Keshav Deep Verma, Seow Yuen Yee, Thomas Rocznik, Philip Schmidt 0001, Joschka Boedecker, Christian Peters |
UbiComp | 6 |
| 2018 | Early Seizure Detection with an Energy-Efficient Convolutional Neural Network on an Implantable MicrocontrollerabstractImplantable, closed-loop devices for automated early detection and stimulation of epileptic seizures are promising treatment options for patients with severe epilepsy that cannot be treated with traditional means. Most approaches for early seizure detection in the literature are, however, not optimized for implementation on ultra-low power microcontrollers required for long-term implantation. In this paper we present a convolutional neural network for the early detection of seizures from in- tracranial EEG signals, designed specifically for this purpose. In addition, we investigate approximations to comply with hardware limits while preserving accuracy. We compare our approach to three previously proposed convolutional neural networks and a feature-based SVM classifier with respect to detection accuracy, latency and computational needs. Evaluation is based on a comprehensive database with long-term EEG recordings. The proposed method outperforms the other detectors with a median sensitivity of 0.96, false detection rate of 10.1 per hour and median detection delay of 3.7 seconds, while being the only approach suited to be realized on a low power microcontroller due to its parsimonious use of computational and memory resources. Maria Hügle, Simon Heller, Manuel Watter, Manuel Blum 0002, Farrokh Manzouri, Matthias Dümpelmann, Andreas Schulze-Bonhage, Peter Woias, Joschka Boedecker |
IJCNN | 9 |
| 2017 | Predicting Time Series with Space-Time Convolutional and Recurrent Neural Networks
Wolfgang Groß, Sascha Lange, Joschka Boedecker, Manuel Blum 0002 |
ESANN | 3 |
| 2017 | Deep reinforcement learning with successor features for navigation across similar environmentsabstractIn this paper we consider the problem of robot navigation in simple maze-like environments where the robot has to rely on its onboard sensors to perform the navigation task. In particular, we are interested in solutions to this problem that do not require localization, mapping or planning. Additionally, we require that our solution can quickly adapt to new situations (e.g., changing navigation goals and environments). To meet these criteria we frame this problem as a sequence of related reinforcement learning tasks. We propose a successor-feature-based deep reinforcement learning algorithm that can learn to transfer knowledge from previously mastered navigation tasks to new problem instances. Our algorithm substantially decreases the required learning time after the first task instance has been solved, which makes it easily adaptable to changing environments. We validate our method in both simulated and real robot experiments with a Robotino and compare it to a set of baseline methods including classical planning-based navigation. Jingwei Zhang 0001, Jost Tobias Springenberg, Joschka Boedecker, Wolfram Burgard |
IROS | 3 |
| 2016 | Autonomous Optimization of Targeted Stimulation of Neuronal NetworksabstractDriven by clinical needs and progress in neurotechnology, targeted interaction with neuronal networks is of increasing importance. Yet, the dynamics of interaction between intrinsic ongoing activity in neuronal networks and their response to stimulation is unknown. Nonetheless, electrical stimulation of the brain is increasingly explored as a therapeutic strategy and as a means to artificially inject information into neural circuits. Strategies using regular or event-triggered fixed stimuli discount the influence of ongoing neuronal activity on the stimulation outcome and are therefore not optimal to induce specific responses reliably. Yet, without suitable mechanistic models, it is hardly possible to optimize such interactions, in particular when desired response features are network-dependent and are initially unknown. In this proof-of-principle study, we present an experimental paradigm using reinforcement-learning (RL) to optimize stimulus settings autonomously and evaluate the learned control strategy using phenomenological models. We asked how to (1) capture the interaction of ongoing network activity, electrical stimulation and evoked responses in a quantifiable 'state' to formulate a well-posed control problem, (2) find the optimal state for stimulation, and (3) evaluate the quality of the solution found. Electrical stimulation of generic neuronal networks grown from rat cortical tissue in vitro evoked bursts of action potentials (responses). We show that the dynamic interplay of their magnitudes and the probability to be intercepted by spontaneous events defines a trade-off scenario with a network-specific unique optimal latency maximizing stimulus efficacy. An RL controller was set to find this optimum autonomously. Across networks, stimulation efficacy increased in 90% of the sessions after learning and learned latencies strongly agreed with those predicted from open-loop experiments. Our results show that autonomous techniques can exploit quantitative relationships underlying activity-response interaction in biological neuronal networks to choose optimal actions. Simple phenomenological models can be useful to validate the quality of the resulting controllers. Sreedhar S. Kumar, Jan Wülfing, Samora Okujeni, Joschka Boedecker, Martin A. Riedmiller, Ulrich Egert |
PLoS Comput. Biol. | 4 |
| 2015 | Embed to Control: A Locally Linear Latent Dynamics Model for Control from Raw ImagesabstractWe introduce Embed to Control (E2C), a method for model learning and control of non-linear dynamical systems from raw pixel images. E2C consists of a deep generative model, belonging to the family of variational autoencoders, that learns to generate image trajectories from a latent space in which the dynamics is constrained to be locally linear. Our model is derived directly from an optimal control formulation in latent space, supports long-term prediction of image sequences and exhibits strong performance on a variety of complex control problems. Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker, Martin A. Riedmiller |
NIPS | 3 |
| 2014 | Approximate real-time optimal control based on sparse Gaussian process modelsabstractIn this paper we present a fully automated approach to (approximate) optimal control of non-linear systems. Our algorithm jointly learns a non-parametric model of the system dynamics - based on Gaussian Process Regression (GPR) - and performs receding horizon control using an adapted iterative LQR formulation. This results in an extremely data-efficient learning algorithm that can operate under real-time constraints. When combined with an exploration strategy based on GPR variance, our algorithm successfully learns to control two benchmark problems in simulation (two-link manipulator, cart-pole) as well as to swing-up and balance a real cart-pole system. For all considered problems learning from scratch, that is without prior knowledge provided by an expert, succeeds in less than 10 episodes of interaction with the system. Joschka Boedecker, Jost Tobias Springenberg, Jan Wülfing, Martin A. Riedmiller |
ADPRL | 1 |
| 2010 | Improving Recurrent Neural Network Performance Using Transfer Entropy
Oliver Obst, Joschka Boedecker, Minoru Asada |
ICONIP (2) | 2 |
| 2009 | Studies on reservoir initialization and dynamics shaping in echo state networks
Joschka Boedecker, Oliver Obst, Norbert Michael Mayer, Minoru Asada |
ESANN | 1 |
| 2007 | Introducing Physical Visualization Sub-league
Rodrigo da Silva Guerra, Joschka Boedecker, Norbert Michael Mayer, Shinzo Yanagimachi, Yasuji Hirosawa, Kazuhiko Yoshikawa, Masaaki Namekawa, Minoru Asada |
RoboCup | 2 |
| 2007 | HMDP: A New Protocol for Motion Pattern Generation Towards Behavior Abstraction
Norbert Michael Mayer, Joschka Boedecker, Kazuhiro Masui, Masaki Ogino, Minoru Asada |
RoboCup | 2 |
| 2006 | 3D2Real: Simulation League Finals in Real Robots
Norbert Michael Mayer, Joschka Boedecker, Rodrigo da Silva Guerra, Oliver Obst, Minoru Asada |
RoboCup | 2 |
| 2005 | Flexible Coordination of Multiagent Team Behavior Using HTN Planning
Oliver Obst, Joschka Boedecker |
RoboCup | 2 |