Hany Abdulsamad

dblp:173/6249 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0001-8683-8784ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Probabilistic and Bayesian machine learning · 46% Reinforcement learning · 25% Planning, search and constraint satisfaction · 13%

Topics — the 19 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
policy optimization
2.242025
Sequential Monte Carlo for Policy Optimization in Continuous POMDPs · NeurIPS 2025
Nesting Particle Filters for Experimental Design in Dynamical Systems · ICML 2024
Model-Free Trajectory-based Policy Optimization with Monotonic Improvement · J. Mach. Learn. Res. 2018
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
sequential monte carlo
1.622025
Sequential Monte Carlo for Policy Optimization in Continuous POMDPs · NeurIPS 2025
Nesting Particle Filters for Experimental Design in Dynamical Systems · ICML 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty › partially observable markov decision process
continuous POMDP
0.912025
Sequential Monte Carlo for Policy Optimization in Continuous POMDPs · NeurIPS 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
partially observable markov decision process
0.912025
Sequential Monte Carlo for Policy Optimization in Continuous POMDPs · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning
probabilistic inference
0.912025
Sequential Monte Carlo for Policy Optimization in Continuous POMDPs · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › experimental design
bayesian experimental design
0.812024
Nesting Particle Filters for Experimental Design in Dynamical Systems · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.812024
Variational Hierarchical Mixtures for Probabilistic Learning of Inverse Dynamics · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
hierarchical mixture model
0.812024
Variational Hierarchical Mixtures for Probabilistic Learning of Inverse Dynamics · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Robotics › Motion planning and robot control › robot control › learning control
inverse dynamics learning
0.812024
Variational Hierarchical Mixtures for Probabilistic Learning of Inverse Dynamics · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.812024
Variational Hierarchical Mixtures for Probabilistic Learning of Inverse Dynamics · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
bayesian mixture model
0.512021
A Variational Infinite Mixture for Probabilistic Inverse Dynamics Learning · ICRA 2021
Machine learning › Learning paradigms
curriculum learning
0.512021
A Probabilistic Interpretation of Self-Paced Learning with Applications to Reinforcement Learning · J. Mach. Learn. Res. 2021
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff
0.512021
A Probabilistic Interpretation of Self-Paced Learning with Applications to Reinforcement Learning · J. Mach. Learn. Res. 2021
Machine learning › Learning paradigms › curriculum learning
self-paced learning
0.512021
A Probabilistic Interpretation of Self-Paced Learning with Applications to Reinforcement Learning · J. Mach. Learn. Res. 2021
Machine learning › Reinforcement learning › policy optimization
monotonic improvement
0.312018
Model-Free Trajectory-based Policy Optimization with Monotonic Improvement · J. Mach. Learn. Res. 2018
Machine learning › Reinforcement learning › policy optimization
proximal policy optimization
0.312018
Model-Free Trajectory-based Policy Optimization with Monotonic Improvement · J. Mach. Learn. Res. 2018
Robotics › Motion planning and robot control
trajectory optimization
0.212016
Model-Free Trajectory Optimization for Reinforcement Learning · ICML 2016
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
particle markov chain monte carlo
0.212024
Nesting Particle Filters for Experimental Design in Dynamical Systems · ICML 2024
Robotics › Motion planning and robot control
robot control
0.212024
Variational Hierarchical Mixtures for Probabilistic Learning of Inverse Dynamics · IEEE Trans. Pattern Anal. Mach. Intell. 2024

Methods — techniques the papers use, named apart from their topics

nested sequential monte carlo · 1.6policy gradient · 0.9feynman-kac model · 0.9variational inference · 0.8particle MCMC · 0.8local regression · 0.8gradient-based policy amortization · 0.8bayesian nonparametrics · 0.8local polynomial models · 0.5bayesian nonparametric mixture · 0.5
YearPublicationVenuePosition
2025 Sequential Monte Carlo for Policy Optimization in Continuous POMDPs
abstract
Optimal decision-making under partial observability requires agents to balance reducing uncertainty (exploration) against pursuing immediate objectives (exploitation). In this paper, we introduce a novel policy optimization framework for continuous partially observable Markov decision processes (POMDPs) that explicitly addresses this challenge. Our method casts policy learning as probabilistic inference in a non-Markovian Feynman--Kac model that inherently captures the value of information gathering by anticipating future observations, without requiring suboptimal approximations or handcrafted heuristics. To optimize policies under this model, we develop a nested sequential Monte Carlo (SMC) algorithm that efficiently estimates a history-dependent policy gradient under samples from the optimal trajectory distribution induced by the POMDP. We demonstrate the effectiveness of our algorithm across standard continuous POMDP benchmarks, where existing methods struggle to act under uncertainty.
Hany Abdulsamad, Sahel Mohammad Iqbal, Simo Särkkä
NeurIPS1
2024 Nesting Particle Filters for Experimental Design in Dynamical Systems
abstract
In this paper, we propose a novel approach to Bayesian experimental design for non-exchangeable data that formulates it as risk-sensitive policy optimization. We develop the Inside-Out SMC$^2$ algorithm, a nested sequential Monte Carlo technique to infer optimal designs, and embed it into a particle Markov chain Monte Carlo framework to perform gradient-based policy amortization. Our approach is distinct from other amortized experimental design techniques, as it does not rely on contrastive estimators. Numerical validation on a set of dynamical systems showcases the efficacy of our method in comparison to other state-of-the-art strategies.
Sahel Iqbal, Adrien Corenflos, Simo Särkkä, Hany Abdulsamad
ICML4
2024 Variational Hierarchical Mixtures for Probabilistic Learning of Inverse Dynamics
abstract
Well-calibrated probabilistic regression models are a crucial learning component in robotics applications as datasets grow rapidly and tasks become more complex. Unfortunately, classical regression models are usually either probabilistic kernel machines with a flexible structure that does not scale gracefully with data or deterministic and vastly scalable automata, albeit with a restrictive parametric form and poor regularization. In this paper, we consider a probabilistic hierarchical modeling paradigm that combines the benefits of both worlds to deliver computationally efficient representations with inherent complexity regularization. The presented approaches are probabilistic interpretations of local regression techniques that approximate nonlinear functions through a set of local linear or polynomial units. Importantly, we rely on principles from Bayesian nonparametrics to formulate flexible models that adapt their complexity to the data and can potentially encompass an infinite number of components. We derive two efficient variational inference techniques to learn these representations and highlight the advantages of hierarchical infinite local regression models, such as dealing with non-smooth functions, mitigating catastrophic forgetting, and enabling parameter sharing and fast predictions. Finally, we validate this approach on large inverse dynamics datasets and test the learned models in real-world control scenarios.
Hany Abdulsamad, Peter Nickl, Pascal Klink, Jan Peters 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 A Variational Infinite Mixture for Probabilistic Inverse Dynamics Learning
abstract
Probabilistic regression techniques in control and robotics applications have to fulfill different criteria of data-driven adaptability, computational efficiency, scalability to high dimensions, and the capacity to deal with different modalities in the data. Classical regressors usually fulfill only a subset of these properties. In this work, we extend seminal work on Bayesian nonparametric mixtures and derive an efficient variational Bayes inference technique for infinite mixtures of probabilistic local polynomial models with well-calibrated certainty quantification. We highlight the model’s power in combining data-driven complexity adaptation, fast prediction, and the ability to deal with discontinuous functions and heteroscedastic noise. We benchmark this technique on a range of large real-world inverse dynamics datasets, showing that the infinite mixture formulation is competitive with classical Local Learning methods and regularizes model complexity by adapting the number of components based on data and without relying on heuristics. Moreover, to showcase the practicality of the approach, we use the learned models for online inverse dynamics control of a Barrett-WAM manipulator, significantly improving the trajectory tracking performance.
Hany Abdulsamad, Peter Nickl, Pascal Klink, Jan Peters 0001
ICRA1
2021 A Probabilistic Interpretation of Self-Paced Learning with Applications to Reinforcement Learning
abstract
Across machine learning, the use of curricula has shown strong empirical potential to improve learning from data by avoiding local optima of training objectives. For reinforcement learning (RL), curricula are especially interesting, as the underlying optimization has a strong tendency to get stuck in local optima due to the exploration-exploitation trade-off. Recently, a number of approaches for an automatic generation of curricula for RL have been shown to increase performance while requiring less expert knowledge compared to manually designed curricula. However, these approaches are seldomly investigated from a theoretical perspective, preventing a deeper understanding of their mechanics. In this paper, we present an approach for automated curriculum generation in RL with a clear theoretical underpinning. More precisely, we formalize the well-known self-paced learning paradigm as inducing a distribution over training tasks, which trades off between task complexity and the objective to match a desired task distribution. Experiments show that training on this induced distribution helps to avoid poor local optima across RL algorithms in different tasks with uninformative rewards and challenging exploration requirements.
Pascal Klink, Hany Abdulsamad, Boris Belousov, Carlo D'Eramo, Jan Peters 0001, Joni Pajarinen
J. Mach. Learn. Res.2
2020 A Nonparametric Off-Policy Policy Gradient
abstract
Reinforcement learning (RL) algorithms still suffer from high sample complexity despite outstanding recent successes. The need for intensive interactions with the environment is especially observed in many widely popular policy gradient algorithms that perform updates using on-policy samples. The priceof such inefficiency becomes evident in real world scenarios such as interaction-driven robot learning, where the success of RL has been rather limited. We address this issue by building on the general sample efficiency of off-policy algorithms. With nonparametric regression and density estimation methods we construct a nonparametric Bellman equation in a principled manner, which allows us to obtain closed-form estimates of the value function, and to analytically express the full policy gradient. We provide a theoretical analysis of our estimate to show that it is consistent under mild smoothness assumptions and empirically show that our approach has better sample efficiency than state-of-the-art policy gradient methods.
Samuele Tosatto, Hany Abdulsamad, Jan Peters 0001
AISTATS3
2019 Chance-Constrained Trajectory Optimization for Non-linear Systems with Unknown Stochastic Dynamics
abstract
Iterative trajectory optimization techniques for non-linear dynamical systems are among the most powerful and sample-efficient methods of model-based reinforcement learning and approximate optimal control. By leveraging time-variant local linear-quadratic approximations of system dynamics and reward, such methods can find both a target-optimal trajectory and time-variant optimal feedback controllers. However, the local linear-quadratic assumptions are a major source of optimization bias that leads to catastrophic greedy updates, raising the issue of proper regularization. Moreover, the approximate models' disregard for any physical state-action limits of the system causes further aggravation of the problem, as the optimization moves towards unreachable areas of the state-action space. In this paper, we address the issue of constrained systems in the scenario of online-fitted stochastic linear dynamics. We propose modeling state and action physical limits as probabilistic chance constraints linear in both state and action and introduce a new trajectory optimization technique that integrates these probabilistic constraints by optimizing a relaxed quadratic program. Our empirical evaluations show a significant improvement in learning robustness, which enables our approach to perform more effective updates and avoid premature convergence observed in state-of-the-art algorithms.
Onur Celik, Hany Abdulsamad, Jan Peters 0001
IROS2
2018 Model-Free Trajectory-based Policy Optimization with Monotonic Improvement
abstract
Many of the recent trajectory optimization algorithms alternate between linear approximation of the system dynamics around the mean trajectory and conservative policy update. One way of constraining the policy change is by bounding the Kullback-Leibler (KL) divergence between successive policies. These approaches already demonstrated great experimental success in challenging problems such as end-to-end control of physical systems. However, the linear approximation of the system dynamics can introduce a bias in the policy update and prevent convergence to the optimal policy. In this article, we propose a new model-free trajectory-based policy optimization algorithm with guaranteed monotonic improvement. The algorithm backpropagates a local, quadratic and time-dependent \qfunc learned from trajectory data instead of a model of the system dynamics. Our policy update ensures exact KL-constraint satisfaction without simplifying assumptions on the system dynamics. We experimentally demonstrate on highly non-linear control tasks the improvement in performance of our algorithm in comparison to approaches linearizing the system dynamics. In order to show the monotonic improvement of our algorithm, we additionally conduct a theoretical analysis of our policy update scheme to derive a lower bound of the change in policy return between successive iterations.
Riad Akrour, Abbas Abdolmaleki, Hany Abdulsamad, Jan Peters 0001, Gerhard Neumann
J. Mach. Learn. Res.3
2016 Model-Free Trajectory Optimization for Reinforcement Learning
abstract
Many of the recent Trajectory Optimization algorithms alternate between local approximation of the dynamics and conservative policy update. However, linearly approximating the dynamics in order to derive the new policy can bias the update and prevent convergence to the optimal policy. In this article, we propose a new model-free algorithm that backpropagates a local quadratic time-dependent Q-Function, allowing the derivation of the policy update in closed form. Our policy update ensures exact KL-constraint satisfaction without simplifying assumptions on the system dynamics demonstrating improved performance in comparison to related Trajectory Optimization algorithms linearizing the dynamics.
Riad Akrour, Gerhard Neumann, Hany Abdulsamad, Abbas Abdolmaleki
ICML3
2016 Optimal control and inverse optimal control by distribution matching
abstract
Optimal control is a powerful approach to achieve optimal behavior. However, it typically requires a manual specification of a cost function which often contains several objectives, such as reaching goal positions at different time steps or energy efficiency. Manually trading-off these objectives is often difficult and requires a high engineering effort. In this paper, we present a new approach to specify optimal behavior. We directly specify the desired behavior by a distribution over future states or features of the states. For example, the experimenter could choose to reach certain mean positions with given accuracy/variance at specified time steps. Our approach also unifies optimal control and inverse optimal control in one framework. Given a desired state distribution, we estimate a cost function such that the optimal controller matches the desired distribution. If the desired distribution is estimated from expert demonstrations, our approach performs inverse optimal control. We evaluate our approach on several optimal and inverse optimal control tasks on non-linear systems using incremental linearizations similar to differential dynamic programming approaches.
Oleg Arenz, Hany Abdulsamad, Gerhard Neumann
IROS2
2015 Reinforcement learning vs human programming in tetherball robot games
abstract
Reinforcement learning of motor skills is an important challenge in order to endow robots with the ability to learn a wide range of skills and solve complex tasks. However, comparing reinforcement learning against human programming is not straightforward. In this paper, we create a motor learning framework consisting of state-of-the-art components in motor skill learning and compare it to a manually designed program on the task of robot tetherball. We use dynamical motor primitives for representing the robot's trajectories and relative entropy policy search to train the motor framework and improve its behavior by trial and error. These algorithmic components allow for high-quality skill learning while the experimental setup enables an accurate evaluation of our framework as robot players can compete against each other. In the complex game of robot tetherball, we show that our learning approach outperforms and wins a match against a high quality hand-crafted system.
Simone Parisi, Hany Abdulsamad, Alexandros Paraschos, Christian Daniel, Jan Peters 0001
IROS2