Alberto Ballesteros

dblp:129/4820 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
4since 2021 · last 2024
0000-0002-7194-213XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 6 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author
YearPublicationVenuePosition
2024 Towards a Node Active Replication Schema for Highly Reliable Distributed Control Systems Based on TSN
abstract
Given their nature, many control applications that arise from the integration of Operation Technologies (OT) and Information Technologies (IT) are built on top of highly reliable real-time (RT) Distributed Control Systems (DCSs). Since a DCS is made up of several computing nodes that exchange information through a communication subsystem, to achieve high reliability it is necessary that both this subsystem and the service from the nodes are very reliable. To provide RT highly reliable communications while benefiting from Ethernet's advantages, Industry and Academia are pushing the Time-Sensitive Networking Ethernet standards (TSN). On the other hand, one of the most used strategies to ensure a highly reliable service from the nodes is to use fault tolerance in the form of active replication. Our general goal is to develop a complete fault-tolerant architecture (addressing faults both in the communication subsystem and in the nodes) for highly reliable real-time DCSs based on TSN. In this paper we show our ongoing work towards an active replication schema for the nodes of this architecture.
Joan Evangelisti, Manuel Barranco, Julián Proenza, Alberto Ballesteros, Mateu Jover
ETFA4
2024 Mapping IEC 61850 GOOSE Messages into Time-Sensitive Networking
abstract
Modern electrical Substation Automation Systems (SAS) are designed following the guidelines defined in the IEC 61850 standard. This standard specifies the necessary information models and communication services for SAS in such a way that they are independent of the implementation. This allows system designers to choose the specific communication technology that best fits their needs. Time-Sensitive Networking (TSN) is arising as one of the most appealing technologies for this purpose. However, it is necessary to map the communication services to TSN, ensuring that the real-time and fault-tolerance requirements of the messages are met. In particular, efficiently mapping the messages of the Generic Object Oriented Substation Events (GOOSE) service is challenging. This is because it exhibits a transmission pattern that does not align with the types of traffic defined in TSN. In this paper we analyze this transmission pattern, identify and characterize its subpatterns, propose the most suitable mapping for each of them and discuss the efficiency gain with respect the typical approaches used to do this mapping.
Mateu Jover, Alberto Ballesteros, Manuel Barranco, Julián Proenza
ETFA2
2024 Characterizing the Tradeoff between Fault Tolerance and Cost of Redundant TSN Networks
abstract
New emerging Distributed Control Systems (DCSs), like Substation Automation Systems (SASs) of Smart Power Grids, raise new requirements on their underlying control networks. To meet these new requirements, both Industry and Academia are promoting the Time-Sensitive Networking (TSN) Ethernet standards. In particular, TSN includes mechanisms to exchange information simultaneously through several paths of practically any spatially redundant network topology. This topological flexibility can offer a better balance between fault tol-erance (FT) and redundancy cost (extra number of components) than classical Industrial Ethernets. However, the mentioned TSN mechanisms may also increase the cost in terms of extra latency and jitter, which could jeopardize real-time communications. In this paper we show our ongoing work to experimentally assess this extra latency and jitter and, thus, characterize the benefits of TSN in terms of balance between FT and cost.
Mateu Jover, Manuel Barranco, Josep Naranjo, Julián Proenza, Alberto Ballesteros
ETFA5
2021 Exploring the use of Deep Reinforcement Learning to allocate tasks in Critical Adaptive Distributed Embedded Systems
abstract
Critical Adaptive Distributed Embedded Systems (CADES) must carry out a set of funcionalities while fulfilling their associated real-time and dependability requirements. Moreover, they must be able to reconfigure themselves in a bounded time as the operational context changes. Finding a proper configuration can be non-trivial and time-consuming. Several studies have proposed Deep Reinforcement Learning (DRL) approaches to solve combinatorial optimization problems. In this paper, we explore the application of such approaches to CADES by solving a simple tasks allocation problem using DRL and comparing the results with three popular heuristics. The results show that DRL beats two of them and gets very close to the third, while requiring significantly less time to generate a solution.
Ramón Rotaeche, Alberto Ballesteros, Julián Proenza
ETFA2
2019 Temporal Replication of Messages for Adaptive Systems using a Holistic Approach
abstract
Critical Adaptive Distributed Embedded Systems (ADES) must meet high real-time and dependability requirements, while autonomously rearranging themselves to operate in dynamic operational contexts. The DFT4FTT project proposes a self-reconfigurable complete infrastructure, whose different architectural levels provide a set of real-time (RT), fault-tolerance (FT) and flexibility mechanisms that collaborate to adequately support critical ADESs. To efficiently tolerate transient faults in the network of an ADES, this paper describes our ongoing work on providing a dynamic temporal replication of messages that takes into account all the DFT4FTT fault-tolerance mechanisms from a holistic point of view.
Alberto Ballesteros, Manuel Barranco, Sergi Arguimbau, Marc Costa, Julián Proenza
ETFA1
2019 Fault Tolerance in Highly Reliable Ethernet-Based Industrial Systems
abstract
Many industrial systems have specific requirements derived from the applications they execute. Specifically, the interaction of a distributed embedded control system (DECS) with the real-world imposes strict real-time (RT) and reliability requirements. For a system to be RT, it has to produce a proper result in a bounded time. On top of that, for a system to be reliable, it has to operate continuously during its mission time, and in cases in which very high reliability is needed, fault tolerance (FT) techniques are used. Moreover, these systems are often deployed in dynamic environments where the operational conditions may change in an unpredictable manner. Therefore, there is an increasing interest in creating DECSs that are capable of modifying their behavior autonomously and dynamically in response to unexpectedly changing requirements or conditions. In recent years, there is a growing trend toward using Ethernet as the network technology for DECSs. Unfortunately, the original specification of this technology lacks appropriate services to fulfill the most demanding requirements of industrial systems. In this regard, many Ethernet-based protocols and standards have been proposed along the past years to deal with these limitations. In this paper, we survey solutions that have been proposed to achieve FT in Ethernet-based DECSs, considering faults both in their nodes and communication subsystem. In addition, we discuss adaptive FT techniques that can be used to increase the flexibility of adaptive DECS. Finally, we identify future trends and open challenges to build highly reliable DECS in the future.
Ines Alvarez, Alberto Ballesteros, Manuel Barranco, David Gessner, Sinisa Derasevic, Julián Proenza
Proc. IEEE2
2019 A Fault-Tolerant Ethernet for Hard Real-Time Adaptive Systems
abstract
Distributed embedded systems (DESs) that perform critical tasks in unpredictable environments must be reliable, hard real-time, and adaptive. Since a DES comprises nodes that rely on a network, the network must provide adequate support: it must be reliable, convey messages on time, and meet new real-time requirements as the nodes adapt. Ethernet is ill-suited for such hard real-time adaptive systems, but it can be made suitable. The flexible time-triggered (FTT) paradigm already supports hard real-time message exchanges and the necessary flexibility to meet evolving hard real-time requirements, but its Ethernet implementations had reliability limitations. To address these, we designed FTT replicated star for Ethernet (FTTRS), a communication subsystem that tolerates permanent and transient faults, even if they occur simultaneously, while keeping the paradigm's key features: support for both the timely exchange of periodic and sporadic real-time messages, and support for updating the real-time parameters of these messages at runtime. In this paper, we present FTTRS, the first Ethernet-based communication subsystem specifically designed for highly reliable hard real-time adaptive DESs.
David Gessner, Julián Proenza, Manuel Barranco, Alberto Ballesteros
IEEE Trans. Ind. Informatics4
2017 Towards a dynamic task allocation scheme for highly-reliable adaptive distributed embedded systems
abstract
An adaptive distributed embedded system is able to automatize some processes at the same time it modifies its behaviour autonomously and dynamically in response to changing operating conditions. To support adaptivity it is necessary that the underlying Distributed Embedded System (DES) is able to dynamically change the assignment of the processing and network resources. In this regard, the DFT4FTT project aims at providing a complete DES that can support applications with real-time, reliability and adaptivity requirements. This paper describes the first steps towards the design of the task allocation scheme used in the DFT4FTT architecture, responsible for dynamically distributing the workload among the nodes of the DES, taking into account the changes in the environment and in the system itself. This allocation scheme not only provides flexibility from a functional point of view, but also from a fault tolerance point of view. Moreover, its modular design makes it possible to tune the desired level of autonomy in the adaptivity, from a simple support for application reconfiguration to a complete automatic reconfiguration assisted with machine learning algorithms.
Alberto Ballesteros, Julián Proenza, Pere Palmer
ETFA1
2016 First implementation and test of reintegration mechanisms for node replicas in the FT4FTT Architecture
abstract
Distributed Embedded Control Systems (DECSs) used for critical applications must usually abide by strict real-time and dependability requirements. Correspondingly, the FT4FTT project proposes a complete fault-tolerant (FT) architecture for RT DECSs. The Flexible Time-Triggered Ethernet (FTT-Ethernet) communication protocol fulfills the RT requirements, while the FT mechanisms added on top of it, which are based on channel duplication and active replication of nodes, provide the FT behaviour. Temporary faults affecting the channel or the nodes, which are the most probable type of faults in DESs, can manifest in such a way that a node replica loses its coordination with the others and, thereby, it also loses its communication and/or computation capability from then on, leading to attrition of the redundancy initially provided by the active replication of nodes. This paper describes the implementation and test of specific mechanisms that are devised to determine which replicas are temporarily faulty and to promptly reintegrate them.
Alberto Ballesteros, Sinisa Derasevic, Manuel Barranco, Julián Proenza
ETFA1
2016 Improving maintenance of FT4FTT: Extending it to monitor and log its available redundancy via internet
abstract
The FT4FTT project aims at proposing a complete Fault-Tolerant (FT) architecture for Real-Time (RT) critical adaptative Distributed Embedded Control Systems (DECSs) based on Ethernet. FT4FTT tolerates permanent faults in the channel and nodes by using a duplicated Flexible Time-Triggered (FTT) Switched Ethernet star and active replication of the nodes. It also includes mechanisms for node replicas to diagnose and reintegrate after temporary faults affecting the channel or their internal circuitry. However, FT4FTT has no mechanism to deal with channel and node redundancy attrition provoked by permanent faults. This paper presents our ongoing work to extend FT4FTT to both monitor/log its available redundancy, and to remotely access this information via Internet. This will allow to carry out proper maintenance actions, for instance, to timely restore the adequate redundancy level, forecast repairs, and assess the flexibility of the FT mechanisms of adaptative systems.
Manuel Barranco, Adel Zendouh, Alberto Ballesteros, Julián Proenza
ETFA3
2016 First implementation and test of a node replication scheme on top of the flexible time-triggered replicated star for ethernet
abstract
Distributed embedded systems typically have real-time and dependability requirements. Moreover, they must also be flexible to changing conditions when they are deployed in dynamic environments. The FT4FTT project aims at providing a switched Ethernet architecture that can support distributed control applications that are predictable, highly-reliable and adaptive. FT4FTT relies on the Flexible Time-Triggered Replicated Star for Ethernet (FTTRS) to tolerate channel faults. Moreover, nodes' hardware faults are tolerated by means of active node replication with majority voting. In order to coordinately trigger the execution of the tasks in the replicas, we designed the CD4NR mechanism, in which the network assists in deciding what to execute and when. This paper presents the first implementation of the CD4NR mechanism on a real prototype of FTTRS and the first testing of the complete system. For this we developed an experimental setup, based on the hardware-in-the-loop technique, running a real-time control application.
Alberto Ballesteros, Sinisa Derasevic, David Gessner, Francisca Font, Ines Alvarez, Manuel Barranco, Julián Proenza
WFCS1
2015 First experimental evaluation of the consistent replicated voting in the hard real-time ethernet switching architecture
abstract
Distributed Embedded Systems (DESs) typically have dependability and real-time requirements. Moreover, when they are deployed in dynamic environments, they must be flexible enough to adapt to changes in the operation requirements. The Fault Tolerance for Flexible Time-Triggered Ethernet (FT4FTT) project aims at providing a Switched-Ethernet architecture, based on the Flexible Time-Triggered communication paradigm (FTT), that is flexible and highly reliable. In particular, FT4FTT provides node fault-tolerance by means of active replication with majority voting. In this sense, FT4FTT includes the Consistent Replicated Voting (CRV) protocol to enforce replica determinism, even in presence of faults, while maximizing the reliability that can be achieved thanks to the node redundancy and the communication subsystem itself. This papers presents a first implementation of this protocol in a real prototype, and shows the on-going experimental evaluation been carried out to asses its correctness.
Sinisa Derasevic, Maties Melia, Alberto Ballesteros, Manuel Barranco, Julián Proenza
ETFA3
2015 Experimental evaluation of network component crashes and trigger message omissions in the Flexible Time-Triggered Replicated Star for Ethernet
abstract
A distributed embedded system (DES) is made up of a set of computing nodes interconnected by a network. If we want the DES to continue to operate even if a subset of its network elements fail, the network must be fault-tolerant. In particular, this requires that the architecture of the network provides redundant paths between nodes and that any elements critical for the operation of the network are replicated. In the context of DES that must not only be highly reliable, but also provide sufficient flexibility to adapt to unpredictable requirement changes, the Flexible Time-Triggered Replicated Star for Ethernet (FTTRS) has been proposed. One of the core features of FTTRS is precisely its fault-tolerant network architecture. In this paper we present a proof-of-concept prototype of FTTRS and a series of fault-injection experiments. These experiments show that FTTRS can tolerate the crash of any single network element, as well as the crash of various combinations of multiple network elements. A variety of omission failures affecting the most critical FTTRS message (called the trigger message) are also tolerated.
David Gessner, Alberto Ballesteros, Andreu Adrover, Julián Proenza
WFCS2
2014 Achieving elementary cycle synchronization between masters in the flexible time-triggered replicated star for ethernet
abstract
For a distributed embedded system (DES) to operate continuously in a dynamic environment, it must be flexible and highly reliable. This applies in particular to its communication subsystem. The Flexible Time-Triggered Replicated Star for Ethernet (FTTRS) aims at providing such a subsystem by means of a highly-reliable switched-Ethernet architecture based on the Flexible Time-Triggered paradigm (FTT), a master/slave communication paradigm where the master periodically polls the slaves using so-called trigger messages (TMs). In particular, FTTRS interconnects nodes by redundant communication paths provided by two switches, each embedding an FTT master that manages the communication. This allows FTTRS to tolerate the failure of one switch without interrupting the communication as long as the masters are replica determinate, i.e., provide identical service to the slaves. The master replica determinism entails the masters broadcasting their TMs in a lockstep fashion: when one master broadcasts a TM, the other should do the same quasi-simultaneously. In this paper we present a solution inspired by the Precision Time Protocol (PTP) for achieving this lockstep transmission and preliminary results showing the precision with which we can synchronize the masters on a software prototype.
Alberto Ballesteros, Julián Proenza, David Gessner, Guillermo Rodríguez-Navas, Thilo Sauter
ETFA1
2014 Towards an experimental assessment of the slave elementary cycle synchronization in the Flexible Time-Triggered Replicated Star for Ethernet
abstract
The communication subsystem of distributed embedded systems (DES) that must operate continuously and satisfy unpredictable requirement changes must be reliable and flexible. Recently the Flexible Time-Triggered Replicated Star for Ethernet (FTTRS) has been proposed as a communication subsystem that satisfies these two attributes. It is based on the master/multi-slave Flexible-Time Triggered (FTT) communication paradigm and relies on two custom switches, each with its own embedded FTT master. Both masters are active simultaneously and provide the same service. Specifically, they simultaneously and periodically broadcast so-called trigger messages (TMs) in a redundant manner to make them robust to transient channel faults. One of the functions of these TMs is to divide the communication time into rounds called elementary cycles (ECs). For the correct operation of FTTRS, it is important that all slaves agree when each EC starts and ends. A mechanism to achieve this has been recently proposed. This paper presents a first implementation of this mechanism and a series of experimental tests that constitute a first step towards building a prototype of an FTTRS network.
David Gessner, Ines Alvarez, Alberto Ballesteros, Manuel Barranco, Julián Proenza
ETFA3
2014 Towards extending the OMNeT++ INET framework for simulating fault injection in ethernet-based Flexible Time-Triggered systems
abstract
Traditional distributed embedded systems are configured using static environment information and thus do not support dynamic behavior of the system. The necessary flexibility in the system may be provided by the Flexible Time-Triggered (FTT) communication paradigm. If, in addition, it is required that the system operates continuously, the suitable fault tolerance mechanisms that provide high reliability have to be developed and deployed in the system. To be able to successfully assess those mechanisms, it is reasonable to develop simulation models that support injection of various types of faults. This paper describes an OMNeT++ simulation model for distributed systems that are based on the Hard Real-Time Ethernet Switching (HaRTES) implementation of the FTT paradigm. The contribution of the paper is twofold. First, we provide a library of components that are required for modeling FTT networks with arbitrary number of FTT slaves connected to a HaRTES switch, and second, we used the developed components to build an FTT system that is suitable for assessing some of recently proposed mechanisms for tolerating certain transient faults in the communication channel.
Mladen Knezic, Alberto Ballesteros, Julián Proenza
ETFA2
2013 Towards preventing error propagation in a real-time Ethernet switch
abstract
Flexible Time-Triggered communication (FTT) allows a distributed embedded system (DES) to adapt to changing real-time requirements at runtime. This facilitates the continuous operation of the DES under dynamic environments that change over time. However, for continuous operation, high reliability in the nodes of the DES is also crucial. This can be achieved using node replication, as long as failure independence between replicas is ensured, which calls for preventing the propagation of errors. Our goal is to prevent the propagation of Byzantine node behaviours and to ensure that local errors in the channel cannot disturb the global communication. For this, we construct the HaRTES/PG switch, a new switch based on the HaRTES implementation of FTT for Ethernet. This paper presents as a first step a study of the possible errors that may lead to Byzantine node behaviours and a global communication disturbance in HaRTES, as well as some ideas on how to prevent the propagation of these errors in HaRTES/PG.
Alberto Ballesteros, David Gessner, Julián Proenza, Manuel Barranco, Paulo Pedreiras
ETFA1
2011 Designing sfiCAN: A star-based physical fault injector for CAN
abstract
This paper presents the design and a preliminary implementation of sfiCAN: a physical fault injector for the CAN field-bus that allows the creation of a great variety of complex fault scenarios. The fault injector replaces the CAN bus topology with a star topology, whose central element is a hub with fault injection mechanisms. The fault injector is easily configured, with great flexibility, from a PC connected to a dedicated port of the hub. For this it uses a fault-injection specification, which is translated to a configuration protocol on top of CAN. This protocol is only used in-between fault injection tests and therefore does not interfere with the execution of any test. The purpose of the fault injector is to test the behavior of the nodes of a CAN network in the presence of channel errors, in particular, of the nodes' CAN controllers and the software executing on them, for which the star topology is transparent.
David Gessner, Manuel Barranco, Alberto Ballesteros, Julián Proenza
ETFA3