Paul Gesel

dblp:246/7806 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 5 first-author · 4 since 2021Systems, architecture and hardware · 7 · 5 first-author · 4 since 2021
YearPublicationVenuePosition
2023 Learning Stable Dynamics via Iterative Quadratic Programming
abstract
This paper proposes a novel autonomous dynamic system (ADS) based controller for trajectory learning from demonstration (LfD). We call our method Learning Stable Dynamics via Iterative Quadratic Programming (LSD-IQP). LSD-IQP learns an energy function and an ADS from demonstrations via semi-infinite quadratic programming. Energy function constraints are imposed on the learned ADS to ensure convergence to a single goal position. Unlike other energy-based methods, LSD-IQP allows the energy function to have both local maximums and saddle points. This flexibility enables LSD-IQP to learn a broader class of motions compared to other ADS-based controllers. We demonstrate the capabilities of LSD-IQP via several experiments, including: 1) learning handwritten symbols and comparing the swept error area to several other ADS methods 2) learning a pick-and-place task with novel goal positions for a robot, and 3) learning a point to point motion in the presence of a non-convex obstacle for a robot.
Paul Gesel, Momotaz Begum
ICRA1
2023 Self-Supervised Visual Motor Skills via Neural Radiance Fields
abstract
In this paper, we propose a novel network architecture for visual imitation learning that exploits neural radiance fields (NeRFs) and key-point correspondence for self-supervised visual motor policy learning. The proposed network architecture incorporates a dynamic system output layer for policy learning. Combining the stability and goal adaption properties of dynamic systems with the robustness of keypoint-based correspondence yields a policy that is invariant to significant clutter, occlusions, lighting conditions changes, and spatial variations in goal configurations. Experiments on multiple manipulation tasks show that our method outperforms comparable visual motor policy learning methods on both in-distribution and out-of-distribution scenarios when using a small number of training samples.
Paul Gesel, Noushad Sojib, Momotaz Begum
IROS1
2021 Learning to Optimize Control Policies and Evaluate Reproduction Performance from Human Demonstrations
abstract
We are interested in learning from demonstration (LfD) that can both learn and execute a trajectory and evaluate the quality of a previously unseen trajectory in the domain of assistive robotics. To this end, we propose a novel continuous inverse optimal control (IOC) formulation that simultaneously learns an optimal time-invariant controller and an evaluation metric from human demonstrations. We assume that the expert’s objective function is a weighted combination of physically meaningful basis objective functions. The evaluation metric is derived from the learned expert’s objective function. The benefit of this approach is twofold: 1) the controller can be optimized with respect to the learned evaluation metric and subject to the robot’s dynamic limitations and 2) the evaluation metric can evaluate the quality of a demonstrated trajectory. We validate our approach with two experiments in a robot guided therapy setting: 1) evaluating demonstrated exercises with the learned metric and 2) reproducing both unconstrained trajectories and trajectories subject to the robot’s dynamic constraints.
Paul Gesel, Dain La Roche, Sajay Arthanat, Momotaz Begum
IROS1
2021 Robust Behavior Cloning with Adversarial Demonstration Detection
abstract
Imitation learning (IL) frameworks in robotics typically assume that a domain expert's demonstration always contains a correct way of doing the task. Despite its theoretical convenience, this assumption has limited practical values for an IL-powered robot in real world. There are many reasons for an expert in the real world to provide demonstrations that may contain incorrect or potentially unsafe way of doing a task. In order for IL-powered robots to work in the real world, IL frameworks need to detect such adversarial demonstrations and not learn from them. This paper proposes an IL framework that can autonomously detect and remove adversarial demonstrations, if they exist in the demonstration set, as it directly learns a task policy from the expert. The proposed framework that we term Robust Maximum Entropy behavior cloning (R-MaxEnt) learns a stochastic model that maps states to actions. In doing so, R-MaxEnt solves a minmax problem that leverages the entropy of the model to assign weights to different demonstrations while assigning poor weights to adversarial samples. Our empirical results show that R-MaxEnt outperforms the existing IL approaches in both real and simulated robotics tasks.
Mostafa Hussein, Brendan Crowe, Madison Clark-Turner, Paul Gesel, Marek Petrik, Momotaz Begum
IROS4
2020 Learning Optimized Human Motion via Phase Space Analysis
abstract
This paper proposes a dynamic system based learning from demonstration approach to teach a robot activities of daily living. The approach takes inspiration from human movement literature to formulate trajectory learning as an optimal control problem. We assume a weighted combination of basis objective functions is the true objective function for a demonstrated motion. We derive basis objective functions analogous to those in human movement literature to optimize the robot's motion. This method aims to naturally adapt the learned motion in different situations. To validate our approach, we learn motions from two categories: 1) commonly prescribed therapeutic exercises and 2) tea making. We show the reproduction accuracy of our method and compare torque requirements to the dynamic motion primitive for each motion, with and without an added load.
Paul Gesel, Francesco Mikulis-Borsoi, Dain La Roche, Sajay Arthanat, Momotaz Begum
IROS1
2019 Leveraging Temporal Reasoning for Policy Selection in Learning from Demonstration
abstract
High-level human activities often have rich temporal structures that determine the order in which atomic actions are executed. We propose the Temporal Context Graph (TCG), a temporal reasoning model that integrates probabilistic inference with Allen's interval algebra, to capture these temporal structures. TCGs are capable of modeling tasks with cyclical atomic actions and consisting of sequential and parallel temporal relations. We present Learning from Demonstration as the application domain where the use of TCGs can improve policy selection and address the problem of perceptual aliasing. Experiments validating the model are presented for learning two tasks from demonstration that involve structured human-robot interactions. The source code for this implementation is available at https://github.com/AssistiveRoboticsUNH/TCG.
Estuardo Carpio, Madison Clark-Turner, Paul Gesel, Momotaz Begum
ICRA3
2019 Learning Motion Trajectories from Phase Space Analysis of the Demonstration
abstract
A major goal of learning from demonstration is task generalization via observation of a teacher. In this paper, we propose a novel framework for learning motion from a single demonstration. Our approach reconstructs the demonstrated trajectory's phase space curve via a linear piece-wise regression method. We approximate dynamics of trajectory segments with linear time invariant equations, each yielding closed form solutions. We show convergence to desired phase space states via an energy-based analysis. The robustness of the model is evaluated on a robot for a sequential trajectory task. Additionally, we show the advantages that the phase space model has over the dynamic motion primitive for a kinematic based task.
Paul Gesel, Momotaz Begum, Dain La Roche
ICRA1