Miguel Zamora

dblp:178/7796 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
4since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2024 Deep Compliant Control for Legged Robots
abstract
Control policies trained using deep reinforcement learning often generate stiff, high-frequency motions in response to unexpected disturbances. To promote more natural and compliant balance recovery strategies, we propose a simple modification to the typical reinforcement learning training process. Our key insight is that stiff responses to perturbations are due to an agent’s incentive to maximize task rewards at all times, even as perturbations are being applied. As an alternative, we introduce an explicit recovery stage where tracking rewards are given irrespective of the motions generated by the control policy. This allows agents a chance to gradually recover from disturbances before attempting to carry out their main tasks. Through an in-depth analysis, we highlight both the compliant nature of the resulting control policies, as well as the benefits that compliance brings to legged locomotion. In our simulation and hardware experiments, the compliant policy achieves more robust, energy-efficient, and safe interactions with the environment.
Adrian Hartmann, Dongho Kang, Fatemeh Zargarbashi, Miguel Zamora, Stelian Coros
ICRA4
2024 TRTM: Template-based Reconstruction and Target-oriented Manipulation of Crumpled Cloths
abstract
Precise reconstruction and manipulation of the crumpled cloths is challenging due to the high dimensionality of cloth models, as well as the limited observation at self-occluded regions. We leverage the recent progress in the field of single-view reconstruction to template-based reconstruct the crumpled cloths from their top-view depth observations only, with our proposed sim-real registration protocols. In contrast to previous implicit cloth representations, our reconstruction mesh explicitly describes the positions and visibilities of the entire cloth mesh vertices, enabling more efficient dual-arm and single-arm target-oriented manipulations. Experiments demonstrate that our TRTM system can be applied to daily cloths that have similar topologies as our template mesh, but with different shapes, sizes, patterns, and physical properties. Videos, datasets, pre-trained models, and code can be downloaded from our project website: https://wenbwa.github.io/TRTM/.
Gen Li 0010, Miguel Zamora, Stelian Coros
ICRA3
2023 Gradient-Based Trajectory Optimization With Learned Dynamics
abstract
Trajectory optimization methods have achieved an exceptional level of performance on real-world robots in recent years. These methods heavily rely on accurate analytical models of the dynamics, yet some aspects of the physical world can only be captured to a limited extent. An alternative approach is to leverage machine learning techniques to learn a differentiable dynamics model of the system from data. In this work, we use trajectory optimization and model learning for performing highly dynamic and complex tasks with robotic systems in absence of accurate analytical models of the dynamics. We show that a neural network can model highly nonlinear behaviors accurately for large time horizons, from data collected in only 25 minutes of interactions on two distinct robots: (i) the Boston Dynamics Spot and an (ii) RC car. Furthermore, we use the gradients of the neural network to perform gradient-based trajectory optimization. In our hardware experiments, we demonstrate that our learned model can represent complex dynamics for both the Spot and Radio-controlled (RC) car, and gives good performance in combination with trajectory optimization methods.
Bhavya Sukhija, Nathanael Köhler, Miguel Zamora, Simon Zimmermann, Sebastian Curi, Andreas Krause 0001, Stelian Coros
ICRA3
2021 PODS: Policy Optimization via Differentiable Simulation
abstract
Current reinforcement learning (RL) methods use simulation models as simple black-box oracles. In this paper, with the goal of improving the performance exhibited by RL algorithms, we explore a systematic way of leveraging the additional information provided by an emerging class of differentiable simulators. Building on concepts established by Deterministic Policy Gradients (DPG) methods, the neural network policies learned with our approach represent deterministic actions. In a departure from standard methodologies, however, learning these policies does not hinge on approximations of the value function that must be learned concurrently in an actor-critic fashion. Instead, we exploit differentiable simulators to directly compute the analytic gradient of a policy’s value function with respect to the actions it outputs. This, in turn, allows us to efficiently perform locally optimal policy improvement iterations. Compared against other state-of-the-art RL methods, we show that with minimal hyper-parameter tuning our approach consistently leads to better asymptotic behavior across a set of payload manipulation tasks that demand a high degree of accuracy and precision.
Miguel Zamora, Momchil Peychev, Sehoon Ha, Martin T. Vechev, Stelian Coros
ICML1
2016 Real-time and decision taking selection of single-particles during automated cryo-EM sessions based on neuro-fuzzy method
David Gil-Carton, Miguel Zamora, James D. Sutherland, Rosa Barrio, Izaskun Garrido Hernandez, Mikel Valle, Aitor J. Garrido
Expert Syst. Appl.2
2000 Evolutionary computation techniques for behaviour fusion in autonomous mobile robots
abstract
In this paper, we present evolutionary techniques to solve the problem of the conflicts between different behaviours in the context of an autonomous mobile robot. We also describe the working environment, based on a custom programming language (named BG after its inventors, Barber/spl acute/a and Go/spl acute/mez, 1996) and an agent architecture, where we test a series of behaviours that were developed using fuzzy logic. Finally, some results related to a simple navigational task in an unknown environment are presented.
Humberto Martínez, Antonio F. Skarmeta, Fernando Jiménez, Miguel Zamora
CEC4