Eduardo Sebastián

dblp:59/480 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0001-9671-4056ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Systems, architecture and hardware · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 LATMOS: Latent Automaton Task Model from Observation Sequences
abstract
Robot task planning from high-level instructions is an important step towards deploying fully autonomous robot systems in the service sector. Three key aspects of robot task planning present challenges yet to be resolved simultaneously, namely, (i) factorization of complex tasks specifications into simpler executable subtasks, (ii) understanding of the current task state from raw observations, and (iii) planning and verification of task executions. To address these challenges, we propose LATMOS, an automata-theory-inspired task model that, given observations from correct task executions, is able to factorize the task, while supporting verification and planning operations. LATMOS combines an observation encoder to extract features from potentially high-dimensional observations with a sequence model that encapsulates an automaton with symbols in the latent feature space. We conduct evaluations in three task model learning setups: (i) abstract tasks described by logical formulas, (ii) real-world human tasks described by videos and natural language prompts and (iii) a robot task described by image and state observations. The results show improved plan generation and verification capabilities of LATMOS across different observation modalities and tasks.
Weixiao Zhan, Qiyue Dong, Eduardo Sebastián, Nikolay Atanasov 0001
IROS3
2025 AVOCADO: Adaptive Optimal Collision Avoidance Driven by Opinion
abstract
We present AdaptiVe Optimal Collision Avoidance Driven by Opinion (AVOCADO), a novel navigation approach to address holonomic robot collision avoidance when the robot does not know how cooperative the other agents in the environment are. AVOCADO departs from a velocity obstacle's (VO) formulation akin to the optimal reciprocal collision avoidance method. However, instead of assuming reciprocity, it poses an adaptive control problem to adapt to the cooperation level of other robots and agents in real time. This is achieved through a novel nonlinear opinion dynamics design that relies solely on sensor observations. As a by-product, we leverage tools from the opinion dynamics formulation to naturally avoid the deadlocks in geometrically symmetric scenarios that typically suffer VO-based planners. Extensive numerical simulations show that AVOCADO surpasses existing motion planners in mixed cooperative/noncooperative navigation environments in terms of success rate, time to goal and computational time. In addition, we conduct multiple real experiments that verify that AVOCADO is able to avoid collisions in environments crowded with other robots and humans.
Diego Martinez-Baselga, Eduardo Sebastián, Eduardo Montijano, Luis Riazuelo, Carlos Sagüés, Luis Montano
IEEE Trans. Robotics2
2025 Physics-Informed Multiagent Reinforcement Learning for Distributed Multirobot Problems
abstract
The networked nature of multi-robot systems presents challenges in the context of multi-agent reinforcement learning. Centralized control policies do not scale with increasing numbers of robots, whereas independent control policies do not exploit the information provided by other robots, exhibiting poor performance in cooperative-competitive tasks. In this work we propose a physics-informed reinforcement learning approach able to learn distributed multi-robot control policies that are both scalable and make use of all the available information to each robot. Our approach has three key characteristics. First, it imposes a port-Hamiltonian structure on the policy representation, respecting energy conservation properties of physical robot systems and the networked nature of robot team interactions. Second, it uses self-attention to ensure a sparse policy representation able to handle time-varying information at each robot from the interaction graph. Third, we present a soft actor-critic reinforcement learning algorithm parameterized by our self-attention port-Hamiltonian control policy, which accounts for the correlation among robots during training while overcoming the need of value function factorization. Extensive simulations in different multi-robot scenarios demonstrate the success of the proposed approach, surpassing previous multi-robot reinforcement learning solutions in scalability, while achieving similar or superior performance (with averaged cumulative reward up to$\times 2$greater than the state-of-the-art with robot teams$\times 6$larger than the number of robots at training time). We also validate our approach on multiple real robots in the Georgia Tech Robotarium under imperfect communication, demonstrating zero-shot sim-to-real transfer and scalability across number of robots.
Eduardo Sebastián, Thai Duong 0001, Nikolay Atanasov 0001, Eduardo Montijano, Carlos Sagüés
IEEE Trans. Robotics1
2023 LEMURS: Learning Distributed Multi-Robot Interactions
abstract
This paper presents LEMURS, an algorithm for learning scalable multi-robot control policies from cooperative task demonstrations. We propose a port-Hamiltonian description of the multi-robot system to exploit universal physical constraints in interconnected systems and achieve closed-loop stability. We represent a multi-robot control policy using an architecture that combines self-attention mechanisms and neural ordinary differential equations. The former handles time-varying communication in the robot team, while the latter respects the continuous-time robot dynamics. Our representation is distributed by construction, enabling the learned control policies to be deployed in robot teams of different sizes. We demonstrate that LEMURS can learn interactions and cooperative behaviors from demonstrations of multi-agent navigation and flocking tasks.
Eduardo Sebastián, Thai Duong 0001, Nikolay Atanasov 0001, Eduardo Montijano, Carlos Sagüés
ICRA1
2022 Adaptive Multirobot Implicit Control of Heterogeneous Herds
abstract
This article presents a novel control strategy to herd groups of noncooperative evaders by means of a team of robotic herders. In herding problems, the motion of the evaders is typically determined bystrongly nonlinearandheterogeneous reactivedynamics, which makes the development of flexible control solutions a challenging problem. In this context, we propose Implicit Control, an approach that leverages numerical analysis theory to find suitable herding inputs even when the nonlinearities in the evaders’ dynamics yieldimplicit equations. The intuition behind this methodology consists in driving the input, rather than computing it, toward theunknownvalue that achieves the desired dynamic behavior of the herd. The same idea is exploited to develop an adaptation law, with stability guarantees, that copes with uncertainties in the herd’s models. Moreover, our solution is completed with a novel caging technique based on uncertainty models and control barrier functions, together with a distributed estimator to overcome the need of complete perfect measurements. Different simulations and experiments validate the generality and flexibility of the proposal.
Eduardo Sebastián, Eduardo Montijano, Carlos Sagüés
IEEE Trans. Robotics1
2021 Multi-robot Implicit Control of Herds
abstract
This paper presents a novel control strategy to herd a group of non-cooperative evaders by means of a team of robotic herders. In herding problems, the motion of the evaders is typically determined by strong nonlinear reactive dynamics, escaping from the herders. Many applications demand the herding of numerous and/or heterogeneous entities, making the development of flexible control solutions challenging. In this context, our main contribution is a control approach that finds suitable herding actions even when the nonlinearities in the evaders’ dynamics yield to implicit equations. We resort to numerical analysis theory to characterise the existence conditions of such actions and propose two design methods to compute them, one transforming the continuous time implicit system into an expanded explicit system, and the other applying a numerical method to find the action in discrete time. Simulations and real experiments validate the proposal in different scenarios.
Eduardo Sebastián, Eduardo Montijano
ICRA1
2021 Online voltage prediction using gaussian process regression for fault-tolerant photovoltaic standalone applications
abstract
Abstract This paper presents a fault detection system for photovoltaic standalone applications based on Gaussian Process Regression (GPR). The installation is a communication repeater from the Confederación Hidrográfica del Ebro (CHE), public institution which manages the hydrographic system of Aragón, Spain. Therefore, fault-tolerance is a mandatory requirement, complex to fulfill since it depends on the meteorology, the state of the batteries and the power demand. To solve it, we propose an online voltage prediction solution where GPR is applied in a real and large dataset of two years to predict the behavior of the installation up to 48 hour. The dataset captures electrical and thermal measures of the lead-acid batteries which sustain the installation. In particular, the crucial aspect to avoid failures is to determine the voltage at the end of the night, so different GPR methods are studied. Firstly, the photovoltaic standalone installation is described, along with the dataset. Then, there is an overview of GPR, emphasizing in the key aspects to deal with real and large datasets. Besides, three online recursive multistep GPR model alternatives are tailored, justifying the selection of the hyperparameters: Regular GPR, Sparse GPR and Multiple Experts (ME) GPR. An exhaustive assessment is performed, validating the results with those obtained by Long Short-Term Memory (LSTM) and Nonlinear Autoregressive Exogenous Model (NARX) networks. A maximum error of 127 mV and 308 mV at the end of the night with Sparse and ME, respectively, corroborates GPR as a promising tool.
José Miguel Sanz-Alcaine, Eduardo Sebastián, Iván Sanz-Gorrachategui, Carlos Bernal 0001, Antonio Bono-Nuez, Milutin Pajovic, Philip V. Orlik
Neural Comput. Appl.2
2019 A Multi-robot Cooperative Control Strategy for Non-linear Entrapment Problems
abstract
This paper presents a multi-robot entrapment problem where a group of non-cooperative preys are driven to a set of desired positions by controlling a team of robotic hunters. The main source of complexity is the non-linear behavior of the preys, which is tightly coupled with the position of the hunters. In the paper we analyze the use of a linearized control solution, providing a general framework to solve the entrapment problem for a non-specific number of preys and hunters, and any behavioral model. In order to apply a linear controller, we discuss the search of the operating point for each hunter and its particular control law. Finally, we evaluate the algorithm with different examples of prey dynamics via simulations, analyzing their convergence and the number of hunters needed to control them depending on the size of the group of preys.
Eduardo Sebastián, Eduardo Montijano
ETFA1
2005 Adaptive fuzzy sliding mode controller for the snorkel underwater vehicle
Eduardo Sebastián, Miguel Ángel Sotelo
ICINCO1