Nestor Arana-Arexolaleiba

dblp:220/8316 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-3305-8108ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 88% Multi-agent systems · 12%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › hierarchical reinforcement learning
modular reinforcement learning
0.712023
skrl: Modular and Flexible Library for Reinforcement Learning · J. Mach. Learn. Res. 2023
Machine learning › Reinforcement learning
reinforcement learning library
0.712023
skrl: Modular and Flexible Library for Reinforcement Learning · J. Mach. Learn. Res. 2023
Machine learning › Reinforcement learning › large-scale reinforcement learning
distributed reinforcement learning
0.212023
skrl: Modular and Flexible Library for Reinforcement Learning · J. Mach. Learn. Res. 2023
Knowledge, reasoning and agents › Multi-agent systems › multi-agent learning
multi-agent training
0.212023
skrl: Modular and Flexible Library for Reinforcement Learning · J. Mach. Learn. Res. 2023
YearPublicationVenuePosition
2026 Novel framework for automated testing of ill-defined human-robot interaction environments
abstract
As automated systems advance in complexity, comprehensive testing becomes crucial, particularly for human–robot interaction (HRI) environments where human unpredictability creates ill-defined testing domains that challenge conventional software testing approaches. In interactive robotics, evaluation criteria extend beyond performance to include critical safety considerations. This paper introduces a novel automated testing framework combining runtime monitoring with constraint-based techniques for HRI environments. The framework employs a three-level cognitive oracle architecture – observation, interpretation, and diagnosis – that automatically evaluates the correctness of human and robot actions without requiring expert human intervention. The approach uses constraint-based modeling to handle the non-deterministic nature of HRI scenarios while ensuring safety compliance. Validation through five test cases in a refrigerator disassembly simulation demonstrates the framework’s effectiveness in detecting safety violations and procedural errors under environmental uncertainties.
Aitor Agirre, Íñigo Elguea-Aguinaco, Nestor Arana-Arexolaleiba, Leire Etxeberria Elorza, Joseba A. Agirre-Bastegieta
J. Syst. Softw.3
2025 Towards robust shielded reinforcement learning through adaptive constraints and exploration: The fear field framework
abstract
Machine Learning (ML) techniques, including Reinforcement Learning (RL), demonstrate potential as decisionmaking controllers.However, enhancing the robustness required for real-world deployment remains imperative.Within the realm of Safe RL, Shielded RL emerges as a solution, employing shields to block actions leading to unsafe states and offering safe alternatives through known policies.Yet, many Shielded RL methods rely on dynamic environment models, which may inaccurately predict future states, compromising controller robustness.We introduce the Fear Field framework to mitigate this issue for discrete Markov Decision Processbased (MDP) shields with strictly connected unsafe state spaces and fully observable states, which adjusts safe operation constraints based on disparities between model predictions and actual environmental dynamics.We employ parallel learning and Curriculum Learning (CL) strategies to mitigate lengthy training times in high state-space size environments.Additionally, an adaptive exploration algorithm enhances convergence rates amidst significant environmental dynamic shifts.In our case study, integrating CL and the adaptive exploration algorithm with the Fear Field framework reduces unsafe state occurrences by two orders of magnitude while enhancing convergence time following sudden environmental changes.The Fear Field framework significantly reduces unsafe states in the Frozen Lake Gridworld environment at low computational expense when model predictions deviate from reality, with negligible costs otherwise.
Haritz Odriozola-Olalde, Maider Zamalloa, Nestor Arana-Arexolaleiba, Jon Pérez 0001
Eng. Appl. Artif. Intell.3
2025 Novel automated interactive reinforcement learning framework with a constraint-based supervisor for procedural tasks
abstract
Learning to perform procedural motion or manipulation tasks in unstructured or uncertain environments poses significant challenges for intelligent agents. Although reinforcement learning algorithms have demonstrated positive results on simple tasks, the hard-to-engineer reward functions and the impractical amount of trial-and-error iterations these agents require in long-experience streams still present challenges for deployment in industrially relevant environments. In this regard, interactive reinforcement learning has emerged as a promising approach to mitigate these limitations, whereby a human supervisor provides evaluative or corrective feedback to the learning agent during training. However, the requirement of a human-in-the-loop approach throughout the learning process can be impractical for tasks that span several hours. This study aims to overcome this limitation by automating the learning process and substituting human feedback with an artificial supervisor grounded in constraint-based modeling techniques. In contrast to the logical constraints commonly used for conventional reinforcement learning, constraint-based modeling techniques offer enhanced adaptability in terms of conceptualizing and modeling the human knowledge of a task. This modeling capability allows an automated supervisor to acquire a closer approximation to human reasoning by dividing complex tasks into more manageable components and identifying the associated subtask and contextual cues in which the agent is involved. The supervisor then adjusts the evaluative and corrective feedback to suit the specific subtask under consideration. The framework was assessed using three actor-critic agents in a human–robot interaction environment, demonstrating a sample efficiency improvement of 50% and success rates of ≥ 95% in simulation and 90% in real-world implementation.
Íñigo Elguea-Aguinaco, Aitor Agirre, Unai Izagirre-Aizpitarte, Ibai Inziarte-Hidalgo, Simon Bøgh, Nestor Arana-Arexolaleiba
Knowl. Based Syst.6
2024 A Review on Reinforcement Learning for Motion Planning of Robotic Manipulators
abstract
Effective motion planning is an indispensable prerequisite for the optimal performance of robotic manipulators in any task. In this regard, the research and application of reinforcement learning in robotic manipulators for motion planning have gained great relevance in recent years. The ability of reinforcement learning agents to adapt to variable environments, especially those featuring dynamic obstacles, has propelled their increasing application in this domain. Notwithstanding, a clear need remains for a resource that critically examines the progress, challenges, and future directions of this machine learning control technique in motion planning. This article undertakes a comprehensive review of the landscape of reinforcement learning, offering a retrospective analysis of its application in motion planning from 2018 to the present. The exploration extends to the trends associated with reinforcement learning in the context of serial manipulators and motion planning, as well as the various technological challenges currently presented by this machine learning control technique. The overarching objective of this review is to serve as a valuable resource for the robotics community, facilitating the ongoing development of systems controlled by reinforcement learning. By delving into the primary challenges intrinsic to this technology, the review seeks to enhance the understanding of reinforcement learning’s role in motion planning and provides insights that may suggest future research directions in this domain.
Íñigo Elguea-Aguinaco, Ibai Inziarte-Hidalgo, Simon Bøgh, Nestor Arana-Arexolaleiba
Int. J. Intell. Syst.4
2023 skrl: Modular and Flexible Library for Reinforcement Learning
abstract
skrl is an open-source modular library for reinforcement learning written in Python and designed with a focus on readability, simplicity, and transparency of algorithm implementations. In addition to supporting environments that use the traditional interfaces from OpenAI Gym/Farama Gymnasium, DeepMind and others, it provides the facility to load, configure, and operate NVIDIA Isaac Gym, Isaac Orbit, and Omniverse Isaac Gym environments. Furthermore, it enables the simultaneous training of several agents with customizable scopes (subsets of environments among all available ones), which may or may not share resources, in the same run. The library's documentation can be found at https://skrl.readthedocs.io and its source code is available on GitHub at https://github.com/Toni-SM/skrl.
Antonio Serrano-Muñoz, Dimitrios Chrysostomou, Simon Bøgh, Nestor Arana-Arexolaleiba
J. Mach. Learn. Res.4
2019 Semi-automatic quality inspection of solar cell based on Convolutional Neural Networks
abstract
Quality control of solar cells is a very important part of the production process. A little crack or joint failure can cause bad performance of the cell in the future, partly because the defective areas can be electrically disconnected from the active zones. Nowadays, one of the techniques to carry out this control is electroluminescence (EL), which allows obtaining high-resolution images of the cells where a visual and non-invasive inspection of defects can be done. This inspection is mostly performed by trained human operators. However, as the eyes become tired after a working day and the subjectivity of the operators, the accuracy with which the defect detection is done may be compromised. In order to solve this problem, a method to assist the operator in the inspection of polycrystalline silicon solar cells surface from EL images based on Convolutional Neural Networks is proposed. The method would classify the cells as defective and non-defective, and suggest those cells that are defective for re-inspection. Also, it would propose a segmentation map of the defects in the cell. To compensate for the lack of image samples in the dataset, each cell image is divided into regions by a sliding window. Then, each region is classified as defective or non-defective. And finally, all classifications related to the cell are resembled obtaining a segmented image of defective areas in the cell.
Julen Balzategui, Luka Eciolaza Echeverría, Nestor Arana-Arexolaleiba, Jon Altube, Jean-Philippe Aguerre, Iñaki Legarda-Ereño, Aitor Apraiz
ETFA3