EDBT 2026 Demo / reviewers in the wild / expert
Diego Romeres
dblp:125/5256
· DBLP profile ↗
31ranked-venue papers
1as first author
28since 2021 · last 2025
0000-0002-8603-2438ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 1 first-author · 23 since 2021Systems, architecture and hardware · 19 · 1 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | User Preference Meets Pareto-Optimality in Multi-Objective Bayesian OptimizationabstractIncorporating user preferences into multi-objective Bayesian optimization (MOBO) allows for personalization of the op- timization procedure. Preferences are often abstracted in the form of an unknown utility function, estimated through pair- wise comparisons of potential outcomes. However, utility-driven MOBO methods can yield solutions that are dominated by nearby solutions, as non-dominance is not enforced. Additionally, classical MOBO commonly relies on estimating the entire Pareto front to identify the Pareto-optimal solutions, which can be expensive and ignore user preferences. Here, we present a new method, termed preference-utility-balanced MOBO (PUB-MOBO), that allows users to disambiguate between near-Pareto candidate solutions. PUB-MOBO combines utility-based MOBO with local multi-gradient descent to refine user-preferred solutions to be near-Pareto-optimal. To this end, we propose a novel preference-dominated utility function that concurrently preserves user-preferences and dominance amongst candidate solutions. A key advantage of PUB-MOBO is that the local search is restricted to a (small) region of the Pareto front directed by user preferences, alleviating the need to estimate the entire Pareto-front. PUB-MOBO is tested on three synthetic benchmark problems: DTLZ1, DTLZ2 and DH1, as well as on three real-world problems: Vehicle Safety, Conceptual Marine Design, and Car Side Impact. PUB-MOBO consistently outperforms state-of-the-art competitors in terms of proximity to the Pareto-front and utility regret across all the problems. Joshua Hang Sai Ip, Ankush Chakrabarty, Ali Mesbah 0002, Diego Romeres |
AAAI | 4 |
| 2025 | Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLMabstractHuman-robot collaboration towards a shared goal requires robots to understand human action and interaction with the surrounding environment. This paper focuses on humanrobot interaction (HRI) based on human-robot dialogue that relies on the robot action confirmation and action step generation using multimodal scene understanding. The state-of-theart approach uses multimodal transformers to generate robot action steps aligned with robot action confirmation from a single clip showing a task composed of multiple micro steps. Although actions towards a long-horizon task depend on each other throughout an entire video, the current approaches mainly focus on clip-level processing and do not leverage long-context information. This paper proposes a long-context Q-former incorporating left and right context dependency in full videos. Furthermore, this paper proposes a text-conditioning approach to feed text embeddings directly into the LLM decoder to mitigate the high abstraction of the information in text by Q-former. Experiments with the YouCook2 corpus show that the accuracy of confirmation generation is a major factor in the performance of action planning. Furthermore, we demonstrate that the longcontext Q-former improves the confirmation and action planning by integrating VideoLLaMA3. Chiori Hori, Yoshiki Masuyama, Siddarth Jain, Radu Corcodel, Devesh K. Jha, Diego Romeres, Jonathan Le Roux |
ASRU | 6 |
| 2025 | Interactive Robot Action Replanning using Multimodal LLM Trained from Human Demonstration VideosabstractUnderstanding human actions could allow robots to perform a large spectrum of complex manipulation tasks and make collaboration with humans easier. Recently, multimodal scene understanding using audio-visual Transformers has been used to generate robot action sequences from videos of human demonstrations. However, automatic action sequence generation is not always perfect due to the distribution gap between the training and test environments. To bridge this gap, human intervention could be very effective, such as telling the robot agent what should be done. Motivated by this, we propose an error-correction-based action replanning approach that regenerates better action sequences using (1) automatically generated actions from a pretrained action generator and (2) human error-correction in natural language. We collected singlearm robot action sequences aligned to human action instruction for the cooking video dataset YouCook2. We trained the proposed errorcorrection-based action replanning model using a pre-trained multimodal LLM model (AVBLIP-2), generating a pair of (a) single-arm robot micro-step action sequences and (b) action descriptions in natural language simultaneously. To assess the performance of error correction, we collected human feedback on correcting errors in the automatically generated robot actions. Experiments show that our proposed interactive replanning model trained in a multitask manner using action sequence and description outperformed the baseline model in all types of scores. Chiori Hori, Motonari Kambara, Komei Sugiura, Kei Ota, Sameer Khurana, Siddarth Jain, Radu Corcodel, Devesh K. Jha, Diego Romeres, Jonathan Le Roux |
ICASSP | 9 |
| 2025 | FDPP: Fine-Tune Diffusion Policy with Human PreferenceabstractImitation learning from human demonstrations enables robots to perform complex manipulation tasks and has recently witnessed huge success. However, these techniques often struggle to adapt behavior to new preferences or changes in the environment. To address these limitations, we propose Fine-tuning Diffusion Policy with Human Preference (FDPP). FDPP learns a reward function through preference-based learning. This reward is then used to fine-tune the pre-trained policy with reinforcement learning (RL), resulting in alignment of pre-trained policy with new human preferences while still solving the original task. Our experiments across various robotic tasks and preferences demonstrate that FDPP effectively customizes policy behavior without compromising performance. Additionally, we show that incorporating Kullback-Leibler (KL) regularization during fine-tuning prevents over-fitting and helps maintain the competencies of the initial policy. Devesh K. Jha, Masayoshi Tomizuka, Diego Romeres |
ICRA | 4 |
| 2025 | PACE: Proactive Assistance in Human-Robot Collaboration Through Action-Completion EstimationabstractThis paper introduces the Proactive Assistance through action-Completion Estimation (PACE) framework, designed to enhance human-robot collaboration through real-time monitoring of human progress. PACE incorporates a novel method that combines Dynamic Time Warping (DTW) with correlation analysis to track human task progression from hand movements. PACE trains a reinforcement learning policy from limited demonstrations to generate a proactive assistance policy that synchronizes robotic actions with human activities, minimizing idle time and enhancing collaboration efficiency. We validate the framework through user studies involving 12 participants, showing significant improvements in interaction fluency, reduced waiting times, and positive user feedback compared to traditional methods. Davide De Lazzari, Matteo Terreran, Giulio Giacomuzzo, Siddarth Jain, Pietro Falco, Ruggero Carli, Diego Romeres |
ICRA | 7 |
| 2025 | RecoveryChaining: Learning Local Recovery Policies for Robust ManipulationabstractModel-based planners and controllers are commonly used to solve complex manipulation problems as they can efficiently optimize diverse objectives and generalize to long horizon tasks. However, they often fail during deployment due to noisy actuation, partial observability and imperfect models. To enable a robot to recover from such failures, we propose to use hierarchical reinforcement learning to learn a recovery policy. The recovery policy is triggered when a failure is detected based on sensory observations and seeks to take the robot to a state from which it can complete the task using the nominal model-based controllers. Our approach, called RecoveryChaining, uses a hybrid action space, where the model-based controllers are provided as additional nominal options which allows the recovery policy to decide how to recover, when to switch to a nominal controller and which controller to switch to even with sparse rewards. We evaluate our approach in three multi-step manipulation tasks with sparse rewards, where it learns significantly more robust recovery policies than those learned by baselines. We successfully transfer recovery policies learned in simulation to a physical robot to demonstrate the feasibility of sim-to-real transfer with our method. Shivam Vats, Devesh K. Jha, Maxim Likhachev, Oliver Kroemer, Diego Romeres |
IROS | 5 |
| 2025 | PyRoboCOP: Python-Based Robotic Control and Optimization Package for Manipulation and Collision AvoidanceabstractContacts are central to most manipulation tasks as they provide additional dexterity to robots to perform challenging tasks. However, frictional contacts leads to complex complementarity constraints. Planning in the presence of contacts requires robust handling of these constraints to find feasible solutions. This paper presents PyRoboCOP which is a lightweight Python-based package for control and optimization of robotic systems described by nonlinear Differential Algebraic Equations (DAEs). In particular, the proposed optimization package can handle systems with contacts that are described by complementarity constraints. We also present a general framework for specifying obstacle avoidance constraints using complementarity constraints. The package performs direct transcription of the DAEs into a set of nonlinear equations by performing orthogonal collocation on finite elements. The resulting optimization problem belongs to the class of Mathematical Programs with Complementarity Constraints (MPCCs). MPCCs fail to satisfy commonly assumed constraint qualifications and require special handling of the complementarity constraints in order for NonLinear Program (NLP) solvers to solve them effectively. PyRoboCOP provides automatic reformulation of the complementarity constraints that enables NLP solvers to perform optimization of robotic systems. The package is interfaced with for obtaining sparse derivatives by automatic differentiation and for performing optimization. We provide extensive numerical examples for various different robotic systems with collision avoidance as well as contact constraints represented using complementarity constraints. We provide comparisons with other open source optimization packages like and . The code is open sourced and available at https://github.com/merlresearch/PyRoboCOP.Note to Practitioners—PyRoboCOP is intended to be an easy-to-use software package written in Python which can be used for optimization, estimation and control for a large class of robotic systems. Including, in particular, contact-rich applications to deal with complex scenarios that arise when making and breaking contacts during a task. Typical problems that can be solved with our work are trajectory and control sequence optimization, parameter estimation. To make the proposed software package easier for practitioners, the paper provides access to the package and a large number of example problems. Furthermore, the package also provides a guide describing the details of all the methods a user might have to implement for their own system. Compared to some of the other packages, PyRoboCOP works with NumPy object arrays which is the native computing package in Python. We believe that this will make it much easier to learn and use compared to some of the other optimal control packages. Arvind U. Raghunathan, Devesh K. Jha, Diego Romeres |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | Interactive Planning Using Large Language Models for Partially Observable Robotic TasksabstractDesigning robotic agents to perform open vocabulary tasks has been the long-standing goal in robotics and AI. Recently, Large Language Models (LLMs) have achieved impressive results in creating robotic agents for performing open vocabulary tasks. However, planning for these tasks in the presence of uncertainties is challenging as it requires "chain-of-thought" reasoning, aggregating information from the environment, updating state estimates, and generating actions based on the updated state estimates. In this paper, we present an interactive planning technique for partially observable tasks using LLMs. In the proposed method, an LLM is used to collect missing information from the environment using a robot, and infer the state of the underlying problem from collected observations while guiding the robot to perform the required actions. We also use a fine-tuned Llama 2 model via self-instruct and compare its performance against a pre-trained LLM like GPT-4. Results are demonstrated on several tasks in simulation as well as real-world environments. Lingfeng Sun, Devesh K. Jha, Chiori Hori, Siddarth Jain, Radu Corcodel, Xinghao Zhu, Masayoshi Tomizuka, Diego Romeres |
ICRA | 8 |
| 2024 | Multi-level Reasoning for Robotic Assembly: From Sequence Inference to Contact SelectionabstractAutomating the assembly of objects from their parts is a complex problem with innumerable applications in manufacturing, maintenance, and recycling. Unlike existing research, which is limited to target segmentation, pose regression, or using fixed target blueprints, our work presents a holistic multi-level framework for part assembly planning consisting of part assembly sequence inference, part motion planning, and robot contact optimization. We present the Part Assembly Sequence Transformer (PAST) – a sequence-to-sequence neural network – to infer assembly sequences recursively from a target blueprint. We then use a motion planner and optimization to generate part movements and contacts. To train PAST, we introduce D4PAS: a large-scale Dataset for Part Assembly Sequences consisting of physically valid sequences for industrial objects. Experimental results show that our approach generalizes better than prior methods while needing significantly less computational time for inference. Further details on our experiments and results are available in the video. Xinghao Zhu, Devesh K. Jha, Diego Romeres, Lingfeng Sun, Masayoshi Tomizuka, Anoop Cherian |
ICRA | 3 |
| 2024 | Reinforcement Learning for Athletic Intelligence: Lessons from the 1st "AI Olympics with RealAIGym" Competition
Felix Wiebe, Niccolò Turcato, Alberto Dalla Libera, Théo Vincent, Shubham Vyas, Giulio Giacomuzzo, Ruggero Carli, Diego Romeres, Akhil Sathuluri, Markus Zimmermann, Boris Belousov, Jan Peters 0001, Frank Kirchner, Shivesh Kumar |
IJCAI | 9 |
| 2024 | DECAF: a Discrete-Event based Collaborative Human-Robot Framework for Furniture AssemblyabstractThis paper proposes a task planning framework for collaborative Human-Robot scenarios, specifically focused on assembling complex systems such as furniture. The human is characterized as an uncontrollable agent, implying for example that the agent is not bound by a pre-established sequence of actions and instead acts according to its own preferences. Meanwhile, the task planner computes reactively the optimal actions for the collaborative robot to efficiently complete the entire assembly task in the least time possible.We formalize the problem as a Discrete Event Markov Decision Problem (DE-MDP), a comprehensive framework that incorporates a variety of asynchronous behaviors, human change of mind, and failure recovery as stochastic events. Although the problem could theoretically be addressed by constructing a graph of all possible actions, such an approach would be constrained by computational limitations. The proposed formulation offers an alternative solution utilizing Reinforcement Learning to derive an optimal policy for the robot. Experiments were conducted both in simulation and on a real system with human subjects assembling a chair in collaboration with a 7-DoF manipulator. Giulio Giacomuzzo, Matteo Terreran, Siddarth Jain, Diego Romeres |
IROS | 4 |
| 2024 | Autonomous Robotic Assembly: From Part Singulation to Precise AssemblyabstractImagine a robot that can assemble a functional product from the individual parts presented in any configuration to the robot. Designing such a robotic system is a complex problem which presents several open challenges. To bypass these challenges, the current generation of assembly systems is built with a lot of system integration effort to provide the structure and precision necessary for assembly. These systems are mostly responsible for part singulation, part kitting, and part detection, which is accomplished by intelligent system design. In this paper, we present autonomous assembly of a gear box with minimum requirements on structure. The assembly parts are randomly placed in a two-dimensional work environment for the robot. The proposed system makes use of several different manipulation skills such as sliding for grasping, in-hand manipulation, and insertion to assemble the gear box. All these tasks are run in a closed-loop fashion using vision, tactile, and Force-Torque (F/T) sensors. We perform extensive hardware experiments to show the robustness of the proposed methods as well as the overall system. See supplementary video at https://www.youtube.com/watch?v=cZ9M1DQ23OI. Kei Ota, Devesh K. Jha, Siddarth Jain, William Yerazunis, Radu Corcodel, Yash Shukla, Antonia Bronars, Diego Romeres |
IROS | 8 |
| 2024 | Open Human-Robot Collaboration using Decentralized Inverse Reinforcement LearningabstractThe growing interest in human-robot collaboration (HRC), where humans and robots cooperate towards shared goals, has seen significant advancements over the past decade. While previous research has addressed various challenges, several key issues remain unresolved. Many domains within HRC involve activities that do not necessarily require human presence throughout the entire task. Existing literature typically models HRC as a closed system, where all agents are present for the entire duration of the task. In contrast, an open model offers flexibility by allowing an agent to enter and exit the collaboration as needed, enabling them to concurrently manage other tasks. In this paper, we introduce a novel multiagent framework called oDec-MDP, designed specifically to model open HRC scenarios where agents can join or leave tasks flexibly during execution. We generalize a recent multiagent inverse reinforcement learning method - Dec-AIRL to learn from open systems modeled using the oDec-MDP. Our method is validated through experiments conducted in both a simplified toy firefighting domain and a realistic dyadic human-robot collaborative assembly. Results show that our framework and learning method improves upon its closed system counterpart. Prasanth Sengadu Suresh, Siddarth Jain, Prashant Doshi, Diego Romeres |
IROS | 4 |
| 2024 | A Black-Box Physics-Informed Estimator Based on Gaussian Process Regression for Robot Inverse Dynamics IdentificationabstractLearning the inverse dynamics of robots directly from data, adopting a black-box approach, is interesting for several real-world scenarios where limited knowledge about the system is available. In this article, we propose a black-box model based on Gaussian process (GP) regression for the identification of the inverse dynamics of robotic manipulators. The proposed model relies on a novel multidimensional kernel, calledLagrangian Inspired Polynomial(LIP) kernel. The LIP kernel is based on two main ideas. First, instead of directly modeling the inverse dynamics components, we model as GPs the kinetic and potential energy of the system. The GP prior on the inverse dynamics components is derived from those on the energies by applying the properties of GPs under linear operators. Second, as regards the energy prior definition, we prove a polynomial structure of the kinetic and potential energy, and we derive a polynomial kernel that encodes this property. As a consequence, the proposed model allows also to estimate the kinetic and potential energy without requiring any label on these quantities. Results on simulation and on two real robotic manipulators, namely a 7 DOF Franka Emika Panda, and a 6 DOF MELFA RV4FL, show that the proposed model outperforms state-of-the-art black-box estimators based both on Gaussian processes and neural networks in terms of accuracy, generality, and data efficiency. The experiments on the MELFA robot also demonstrate that our approach achieves performance comparable to fine-tuned model-based estimators, despite requiring less prior information. The code of the proposed model is publicly available. Giulio Giacomuzzo, Ruggero Carli, Diego Romeres, Alberto Dalla Libera |
IEEE Trans. Robotics | 3 |
| 2023 | Simultaneous Tactile Estimation and Control of Extrinsic ContactabstractWe propose a method that simultaneously estimates and controls extrinsic contact with tactile feedback. The method enables challenging manipulation tasks that require controlling light forces and accurate motions in contact, such as balancing an unknown object on a thin rod standing upright. A factor graph-based framework fuses a sequence of tactile and kinematic measurements to estimate and control the interaction between gripper-object-environment, including the location and wrench at the extrinsic contact between the grasped object and the environment and the grasp wrench transferred from the gripper to the object. The same framework simultaneously plans the gripper motions that make it possible to estimate the state while satisfying regularizing control objectives to prevent slip, such as minimizing the grasp wrench and minimizing frictional force at the extrinsic contact. We show results with sub-millimeter contact localization error and good slip prevention even on slippery environments, for multiple contact formations (point, line, patch contact) and transitions between them. See supplementary video and results at https://sites.google.com/view/sim-tact. Sangwoon Kim, Devesh K. Jha, Diego Romeres, Parag Patre, Alberto Rodriguez 0003 |
ICRA | 3 |
| 2023 | Learning Generalizable Pivoting SkillsabstractThe skill of pivoting an object with a robotic system is challenging for the external forces that act on the system, mainly given by contact interaction. The complexity increases when the same skills are required to generalize across different objects. This paper proposes a framework for learning robust and generalizable pivoting skills, which consists of three steps. First, we learn a pivoting policy on an “unitary” object using Reinforcement Learning (RL). Then, we obtain the object's feature space by supervised learning to encode the kinematic properties of arbitrary objects. Finally, to adapt the unitary policy to multiple objects, we learn data-driven projections based on the object features to adjust the state and action space of the new pivoting task. The proposed approach is entirely trained in simulation. It requires only one depth image of the object and can zero-shot transfer to real-world objects. We demonstrate robustness to sim-to-real transfer and generalization to multiple objects. Xiang Zhang 0020, Siddarth Jain, Baichuan Huang, Masayoshi Tomizuka, Diego Romeres |
ICRA | 5 |
| 2023 | Style-transfer based Speech and Audio-visual Scene understanding for Robot Action Sequence Acquisition from Videos
Chiori Hori, Puyuan Peng, David F. Harwath, Kei Ota, Siddarth Jain, Radu Corcodel, Devesh K. Jha, Diego Romeres, Jonathan Le Roux |
INTERSPEECH | 9 |
| 2023 | Constrained Dynamic Movement Primitives for Collision Avoidance in Novel EnvironmentsabstractDynamic movement primitives are widely used for learning skills that can be demonstrated to a robot by a skilled human or controller. While their generalization capabilities and simple formulation make them very appealing to use, they possess no strong guarantees to satisfy operational safety constraints for a task. We present constrained dynamic movement primitives (CDMPs), which can allow for positional constraint satisfaction in the robot workspace. Our method solves a non-linear optimization to perturb an existing DMP's forcing weights to admit a Zeroing Barrier Function (ZBF), which certifies positional workspace constraint satisfaction. We demonstrate our approach under different positional constraints on the end-effector movement on multiple physical robots, such as obstacle avoidance and workspace limitations. Seiji Shaw, Devesh K. Jha, Arvind U. Raghunathan, Radu Corcodel, Diego Romeres, George Dimitri Konidaris, Daniel Nikovski |
IROS | 5 |
| 2023 | Introduction to the special issue on Intelligent Control and Optimisation
Seán F. McLoone, Kevin Guelton, Thierry-Marie Guerra, Gian Antonio Susto, Jus Kocijan, Diego Romeres |
Eng. Appl. Artif. Intell. | 6 |
| 2022 | Transfer Learning for Bayesian Optimization with Principal Component AnalysisabstractBayesian Optimization has been widely used for black-box optimization. Especially in the field of machine learning, BO has obtained remarkable results in hyperparameters optimization. However, the best hyperparameters depend on the specific task and traditionally the BO algorithm needs to be repeated for each task. On the other hand, the relationship between hyperparameters and objectives has similar tendency among tasks. Therefore, transfer learning is an important technology to accelerate the optimization of novel task by leveraging the knowledge acquired in prior tasks. In this work, we propose a new transfer learning strategy for BO. We use information geometry based principal component analysis (PCA) to extract a low-dimension manifold from a set of Gaussian process (GP) posteriors that models the objective functions of the prior tasks. Then, the low dimensional parameters of this manifold can be optimized to adapt to a new task and set a prior distribution for the objective function of the novel task. Experiments on hyperparameters optimization benchmarks show that our proposed algorithm, called BO-PCA, accelerates the learning of an unseen task (less data are required) while having low computational cost. Hideyuki Masui, Diego Romeres, Daniel Nikovski |
ICMLA | 2 |
| 2022 | PyROBOCOP: Python-based Robotic Control & Optimization Package for ManipulationabstractPyROBOCOP is a Python-based package for control, optimization and estimation of robotic systems described by nonlinear Differential Algebraic Equations (DAEs). In particular, the package can handle systems with contacts that are described by complementarity constraints and provides a general framework for specifying obstacle avoidance constraints. The package performs direct transcription of the DAEs into a set of nonlinear equations by performing orthogonal collocation on finite elements. PyROBOCOP provides automatic reformulation of the complementarity constraints that are tractable to NLP solvers to perform optimization of robotic systems. The package is interfaced with ADOL-C [1] for obtaining sparse derivatives by automatic differentiation and IPOPT [2] for performing optimization. We evaluate PyROBOCOP on several manipulation problems for control and estimation. Arvind U. Raghunathan, Devesh K. Jha, Diego Romeres |
ICRA | 3 |
| 2022 | Robust Pivoting: Exploiting Frictional Stability Using Bilevel OptimizationabstractGeneralizable manipulation requires that robots be able to interact with novel objects and environment. This requirement makes manipulation extremely challenging as a robot has to reason about complex frictional interaction with uncertainty in physical properties of the object. In this paper, we study robust optimization for control of pivoting manipulation in the presence of uncertainties. We present insights about how friction can be exploited to compensate for the inaccuracies in the estimates of the physical properties during manipulation. In particular, we derive analytical expressions for stability margin provided by friction during pivoting manipulation. This margin is then used in a bilevel trajectory optimization algorithm to design a controller that maximizes this stability margin to provide robustness against uncertainty in physical properties of the object. We demonstrate our proposed method using a 6 DoF manipulator for manipulating several different objects. Yuki Shirai, Devesh K. Jha, Arvind U. Raghunathan, Diego Romeres |
ICRA | 4 |
| 2022 | Active Exploration for Robotic ManipulationabstractRobotic manipulation stands as a largely unsolved problem despite significant advances in robotics and machine learning in recent years. One of the key challenges in manipulation is the exploration of the dynamics of the environment when there is continuous contact between the objects being manipulated. This paper proposes a model-based active exploration approach that enables efficient learning in sparse-reward robotic manipulation tasks. The proposed method estimates an information gain objective using an ensemble of probabilistic models and deploys model predictive control (MPC) to plan actions online that maximize the expected reward while also performing directed exploration. We evaluate our proposed algorithm in simulation and on a real robot, trained from scratch with our method, on a challenging ball pushing task on tilted tables, where the target ball position is not known to the agent a-priori. Our real-world robot experiment serves as a fundamental application of active exploration in model-based reinforcement learning of complex robotic manipulation tasks. Project page https://sites.google.com/view/aerm. Boris Belousov, Georgia Chalvatzaki, Diego Romeres, Devesh K. Jha, Jan Peters 0001 |
IROS | 4 |
| 2022 | Model-Based Policy Search Using Monte Carlo Gradient Estimation With Real Systems ApplicationabstractIn this article, we present a model-based reinforcement learning (MBRL) algorithm namedMonte Carlo probabilistic inference for learning control(MC-PILCO). This algorithm relies on Gaussian processes (GPs) to model the system dynamics and on a Monte Carlo approach to estimate the policy gradient. This defines a framework in which we ablate the choice of the components, which are the selection of the cost function, the optimization of policies using dropout, and an improved data efficiency through the use of structured kernels in the GP models. The combination of the aforementioned aspects affects dramatically the performance of MC-PILCO. Numerical comparisons in a simulated cart–pole environment show that MC-PILCO exhibits better data efficiency and control performance w.r.t. state-of-the-art GP-based MBRL algorithms. Finally, we apply MC-PILCO to real systems, considering, in particular, systems with partially measurable states. We discuss the importance of modeling both the measurement system and the state estimators during policy optimization. The effectiveness of the proposed solutions has been tested in simulation and on two real systems, which are the Furuta pendulum and the ball-and-plate rig. Fabio Amadio, Alberto Dalla Libera, Riccardo Antonello, Daniel Nikovski, Ruggero Carli, Diego Romeres |
IEEE Trans. Robotics | 6 |
| 2021 | Personalizing Individual Comfort in the Group SettingabstractMaintaining individual thermal comfort in indoor spaces shared by multiple occupants is difficult because it requires both intuition about the thermal properties of the room, as well as an understanding of the thermal comfort preferences of each individual. We explore an approach to optimizing individual thermal comfort within a group through temperature set-point optimization of HVAC equipment. We propose a weakly-supervised algorithm to learn the individual thermal comfort preferences and an autoencoding framework to learn static approximations of room thermodynamics. We further propose two approaches to learn a control law that sets the HVAC set-points subject to the preferred user temperatures. The proposed method is tested on a real data-set obtained from workers in an open office. The results show that, on average, the temperature in the room at each user's location can be regulated to within 0.5C of the user's desired temperature. Emil Laftchiev, Diego Romeres, Daniel Nikovski |
AAAI | 2 |
| 2021 | Tactile-RL for Insertion: Generalization to Objects of Unknown GeometryabstractObject insertion is a classic contact-rich manipulation task. The task remains challenging, especially when considering general objects of unknown geometry, which significantly limits the ability to understand the contact configuration between the object and the environment. We study the problem of aligning the object and environment with a tactile-based feedback insertion policy. The insertion process is modeled as an episodic policy that iterates between insertion attempts followed by pose corrections. We explore different mechanisms to learn such a policy based on Reinforcement Learning. The key contribution of this paper is to demonstrate that it is possible to learn a tactile insertion policy that generalizes across different object geometries, and an ablation study of the key design choices for the learning agent: 1) the type of learning scheme: supervised vs. reinforcement learning; 2) the type of learning schedule: unguided vs. curriculum learning ; 3) the type of sensing modality: force/torque vs. tactile; and 4) the type of tactile representation: tactile RGB vs. tactile flow. We show that the optimal configuration of the learning agent (RL + curriculum + tactile flow) exposed to 4 training objects yields an closed-loop insertion policy that inserts 4 novel objects with over 85.0% success rate and within 3~4 consecutive attempts. Comparisons between F/T and tactile sensing, shows that while an F/T-based policy learns more efficiently, a tactile-based policy provides better generalization. See supplementary video and results at https://sites.google.com/view/tactileinsertion. Siyuan Dong, Devesh K. Jha, Diego Romeres, Sangwoon Kim, Daniel Nikovski, Alberto Rodriguez 0003 |
ICRA | 3 |
| 2021 | Trajectory Optimization for Manipulation of Deformable Objects: Assembly of Belt Drive UnitsabstractThis paper presents a novel trajectory optimization formulation to solve the robotic assembly of the belt drive unit. Robotic manipulations involving contacts and deformable objects are challenging in both dynamic modeling and trajectory planning. For modeling, variations in the belt tension and contact forces between the belt and the pulley could dramatically change the system dynamics. For trajectory planning, it is computationally expensive to plan trajectories for such hybrid dynamical systems as it usually requires planning for discrete modes separately. In this work, we formulate the belt drive unit assembly task as a trajectory optimization problem with complementarity constraints to avoid explicitly imposing contact mode sequences. The problem is solved as a mathematical program with complementarity constraints (MPCC) to obtain feasible and efficient assembly trajectories. We validate the proposed method both in simulations with a physics engine and in real-world experiments with a robotic manipulator. Shiyu Jin, Diego Romeres, Arvind Ragunathan, Devesh K. Jha, Masayoshi Tomizuka |
ICRA | 2 |
| 2021 | Learning Disagreement Regions with Deep Neural Networks to Reduce Practical Complexity of Mixed-Integer MPCabstractEfficiently computing solutions to mixed-integer optimization-based control problems, such as in model predictive control (MPC) of hybrid systems, is extremely challenging due to the exponential worst-case complexity. The practical time-complexity of computing good control actions can be reduced by using a combination of two solvers: a strong solver that generates optimal or near-optimal closed-loop solutions with a large number of iterations, and a weak solver that converges quickly to suboptimal closed-loop solutions. In this paper, we propose the use of deep neural networks to learn sub-regions of the admissible state-space where replacing the strong solver with the weak solver maintains constraint satisfaction properties and does not result in a significant deterioration of performance. We illustrate the practical time-complexity reduction of the proposed solver selection mechanism on a station-keeping problem for a satellite. Ankush Chakrabarty, Rien Quirynen, Diego Romeres, Stefano Di Cairano |
SMC | 3 |
| 2020 | Local Policy Optimization for Trajectory-Centric Reinforcement LearningabstractThe goal of this paper is to present a method for simultaneous trajectory and local stabilizing policy optimization to generate local policies for trajectory-centric model-based reinforcement learning (MBRL). This is motivated by the fact that global policy optimization for non-linear systems could be a very challenging problem both algorithmically and numerically. However, a lot of robotic manipulation tasks are trajectory-centric, and thus do not require a global model or policy. Due to inaccuracies in the learned model estimates, an open-loop trajectory optimization process mostly results in very poor performance when used on the real system. Motivated by these problems, we try to formulate the problem of trajectory optimization and local policy synthesis as a single optimization problem. It is then solved simultaneously as an instance of nonlinear programming. We provide some results for analysis as well as achieved performance of the proposed technique under some simplifying assumptions. Patrik Kolaric, Devesh K. Jha, Arvind U. Raghunathan, Frank L. Lewis, Mouhacine Benosman, Diego Romeres, Daniel Nikovski |
ICRA | 6 |
| 2019 | Sim-to-Real Transfer Learning using Robustified Controllers in Robotic Tasks involving Complex DynamicsabstractLearning robot tasks or controllers using deep reinforcement learning has been proven effective in simulations. Learning in simulation has several advantages. For example, one can fully control the simulated environment, including halting motions while performing computations. Another advantage when robots are involved, is that the amount of time a robot is occupied learning a task-rather than being productive-can be reduced by transferring the learned task to the real robot. Transfer learning requires some amount of fine-tuning on the real robot. For tasks which involve complex (non-linear) dynamics, the fine-tuning itself may take a substantial amount of time. In order to reduce the amount of fine-tuning we propose to learn robustified controllers in simulation. Robustified controllers are learned by exploiting the ability to change simulation parameters (both appearance and dynamics) for successive training episodes. An additional benefit for this approach is that it alleviates the precise determination of physics parameters for the simulator, which is a non-trivial task. We demonstrate our proposed approach on a real setup in which a robot aims to solve a maze game, which involves complex dynamics due to static friction and potentially large accelerations. We show that the amount of fine-tuning in transfer learning for a robustified controller is substantially reduced compared to a non-robustified controller. Jeroen van Baar, Alan Sullivan, Radu Cordorel, Devesh K. Jha, Diego Romeres, Daniel Nikovski |
ICRA | 5 |
| 2019 | Semiparametrical Gaussian Processes Learning of Forward Dynamical Models for Navigating in a Circular MazeabstractThis paper presents a problem of model learning for the purpose of learning how to navigate a ball to a goal state in a circular maze environment with two degrees of freedom. The motion of the ball in the maze environment is influenced by several non-linear effects such as dry friction and contacts, which are difficult to model physically. We propose a semiparametric model to estimate the motion dynamics of the ball based on Gaussian Process Regression equipped with basis functions obtained from physics first principles. The accuracy of this semiparametric model is shown not only in estimation but also in prediction at n-steps ahead and its compared with standard algorithms for model learning. The learned model is then used in a trajectory optimization algorithm to compute ball trajectories. We propose the system presented in the paper as a benchmark problem for reinforcement and robot learning, for its interesting and challenging dynamics and its relative ease of reproducibility. Diego Romeres, Devesh K. Jha, Alberto Dalla Libera, William Yerazunis, Daniel Nikovski |
ICRA | 1 |