Charles Khazoom

dblp:245/5695 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2023
0000-0001-7224-1688ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2023 Benchmarking Potential Based Rewards for Learning Humanoid Locomotion
abstract
The main challenge in developing effective reinforcement learning (RL) pipelines is often the design and tuning the reward functions. Well-designed shaping reward can lead to significantly faster learning. Naively formulated rewards, however, can conflict with the desired behavior and result in overfitting or even erratic performance if not properly tuned. In theory, the broad class of potential based reward shaping (PBRS) can help guide the learning process without affecting the optimal policy. Although several studies have explored the use of potential based reward shaping to accelerate learning convergence, most have been limited to grid-worlds and low-dimensional systems, and RL in robotics has predominantly relied on standard forms of reward shaping. In this paper, we benchmark standard forms of shaping with PBRS for a humanoid robot. We find that in this high-dimensional system, PBRS has only marginal benefits in convergence speed. However, the PBRS reward terms are significantly more robust to scaling than typical reward shaping approaches, and thus easier to tune.
Se Hwan Jeon, Steve Heim, Charles Khazoom, Sangbae Kim
ICRA3
2023 Optimal Scheduling of Models and Horizons for Model Hierarchy Predictive Control
abstract
Model predictive control (MPC) is a powerful tool to control systems with non-linear dynamics and constraints, but its computational demands impose limitations on the dynamics model used for planning. Instead of using a single complex model along the MPC horizon, model hierarchy predictive control (MHPC) reduces solve times by planning over a sequence of models of varying complexity within a single horizon. Choosing this model sequence can become intractable when considering all possible combinations of reduced order models and prediction horizons. We propose a framework to systematically optimize a model schedule for MHPC. We leverage trajectory optimization (TO) to approximate the accumulated cost of the closed-loop controller. We trade off performance and solve times by minimizing the number of decision variables of the MHPC problem along the horizon while keeping the approximate closed-loop cost near optimal. The framework is validated in simulation with a planar humanoid robot as a proof of concept. We find that the approximated closed-loop cost matches the simulated one for most of the model schedules, and show that the proposed approach finds optimal model schedules that transfer directly to simulation, and with total horizons that vary between 1.1 and 1.6 walking steps.
Charles Khazoom, Steve Heim, Daniel González-Díaz, Sangbae Kim
ICRA1
2022 Humanoid Arm Motion Planning for Improved Disturbance Recovery Using Model Hierarchy Predictive Control
abstract
Humans noticeably swing their arms for balancing and locomotion. Although the underlying biomechanical mechanisms have been studied, it is unclear how robots can fully take advantage of these appendages. Most controllers that exploit arms for balance and locomotion rely on feedback and cannot anticipate incoming disturbances and future states. Model predictive controllers readily address these drawbacks but are computationally expensive. Here, we leverage recent work on model hierarchy predictive control (MHPC). We develop an MHPC formulation that plans arm motions in reaction to expected or unexpected disturbances. We tested multiple model compositions using simulated balance experiments with the MIT Humanoid undergoing various disturbances. We found that an MHPC formulation that plans over a full-body kino-dynamic model for a 0.3 s horizon followed by a single rigid body model for 0.5 s horizon runs at 40 Hz and increases the set of disturbances that the robot can withstand. Arms allow the robot to dissipate momentum quickly and move the center of mass independently from the lower body. This kinematic advantage helps generate ground wrenches while avoiding kinematic singularities and keeping the center of mass and center pressure within the support polygon. We note similar advantages when allowing the MHPC to anticipate incoming disturbances.
Charles Khazoom, Sangbae Kim
ICRA1