EDBT 2026 Demo / reviewers in the wild / expert
Julián Proenza
dblp:02/4615
· DBLP profile ↗
71ranked-venue papers
4as first author
12since 2021 · last 2024
0000-0001-7238-0557ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 58 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 since 2021Software engineering, systems software and programming languages · 2Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards a Node Active Replication Schema for Highly Reliable Distributed Control Systems Based on TSNabstractGiven their nature, many control applications that arise from the integration of Operation Technologies (OT) and Information Technologies (IT) are built on top of highly reliable real-time (RT) Distributed Control Systems (DCSs). Since a DCS is made up of several computing nodes that exchange information through a communication subsystem, to achieve high reliability it is necessary that both this subsystem and the service from the nodes are very reliable. To provide RT highly reliable communications while benefiting from Ethernet's advantages, Industry and Academia are pushing the Time-Sensitive Networking Ethernet standards (TSN). On the other hand, one of the most used strategies to ensure a highly reliable service from the nodes is to use fault tolerance in the form of active replication. Our general goal is to develop a complete fault-tolerant architecture (addressing faults both in the communication subsystem and in the nodes) for highly reliable real-time DCSs based on TSN. In this paper we show our ongoing work towards an active replication schema for the nodes of this architecture. Joan Evangelisti, Manuel Barranco, Julián Proenza, Alberto Ballesteros, Mateu Jover |
ETFA | 3 |
| 2024 | Mapping IEC 61850 GOOSE Messages into Time-Sensitive NetworkingabstractModern electrical Substation Automation Systems (SAS) are designed following the guidelines defined in the IEC 61850 standard. This standard specifies the necessary information models and communication services for SAS in such a way that they are independent of the implementation. This allows system designers to choose the specific communication technology that best fits their needs. Time-Sensitive Networking (TSN) is arising as one of the most appealing technologies for this purpose. However, it is necessary to map the communication services to TSN, ensuring that the real-time and fault-tolerance requirements of the messages are met. In particular, efficiently mapping the messages of the Generic Object Oriented Substation Events (GOOSE) service is challenging. This is because it exhibits a transmission pattern that does not align with the types of traffic defined in TSN. In this paper we analyze this transmission pattern, identify and characterize its subpatterns, propose the most suitable mapping for each of them and discuss the efficiency gain with respect the typical approaches used to do this mapping. Mateu Jover, Alberto Ballesteros, Manuel Barranco, Julián Proenza |
ETFA | 4 |
| 2024 | Characterizing the Tradeoff between Fault Tolerance and Cost of Redundant TSN NetworksabstractNew emerging Distributed Control Systems (DCSs), like Substation Automation Systems (SASs) of Smart Power Grids, raise new requirements on their underlying control networks. To meet these new requirements, both Industry and Academia are promoting the Time-Sensitive Networking (TSN) Ethernet standards. In particular, TSN includes mechanisms to exchange information simultaneously through several paths of practically any spatially redundant network topology. This topological flexibility can offer a better balance between fault tol-erance (FT) and redundancy cost (extra number of components) than classical Industrial Ethernets. However, the mentioned TSN mechanisms may also increase the cost in terms of extra latency and jitter, which could jeopardize real-time communications. In this paper we show our ongoing work to experimentally assess this extra latency and jitter and, thus, characterize the benefits of TSN in terms of balance between FT and cost. Mateu Jover, Manuel Barranco, Josep Naranjo, Julián Proenza, Alberto Ballesteros |
ETFA | 4 |
| 2024 | An Improved Worst-Case Response Time Analysis for AVB Traffic in Time-Sensitive NetworksabstractTime-Sensitive Networking (TSN) has become one of the most relevant communication networks in many application areas. Among several traffic classes supported by TSN networks, Audio-Video Bridging (AVB) traffic requires a Worst-Case Response Time Analysis (WCRTA) to ensure that AVB frames meet their time requirements. In this paper, we evaluate the existing WCRTAs that cover various features of TSN, including Scheduled Traffic (ST) interference and preemption. When considering the effect of the ST interference, we detect optimism problems in two of the existing WCRTAs, namely (i) the analysis based on the busy period calculation and (ii) the analysis based on the eligible interval. Therefore, we propose a new analysis including a new ST interference calculation that can extend the analysis based on the eligible interval approach. The new analysis covers the effect of the ST interference, the preemption by the ST traffic, and the multi-hop architecture. The resulting WCRTA, while safe, shows a significant improvement in terms of pessimism level compared to the existing analysis approaches relying on either the concept of busy period or the Network Calculus model. Daniel Bujosa, Julián Proenza, Alessandro Vittorio Papadopoulos, Thomas Nolte, Mohammad Ashjaei |
RTSS | 2 |
| 2023 | Introducing Guard Frames to Ensure Schedulability of All TSN Traffic ClassesabstractOffline scheduling of Scheduled Traffic (ST) in Time-Sensitive Networks (TSN) without taking into account the quality of service of non-ST traffic, e.g., time-sensitive traffic such as Audio-Video Bridging (AVB) traffic, can potentially cause deadline misses for non-ST traffic. In this paper, we report our ongoing work to propose a solution that, regardless of the ST scheduling algorithm being used, can ensure meeting timing requirements for non-ST traffic. To do this, we define a frame called Guard Frame (GF) that will be scheduled together with all ST frames. We show that a proper design for the GFs will leave necessary porosity in the ST schedules to ensure that all non-ST traffic will meet their timing requirements. Daniel Bujosa, Julián Proenza, Alessandro Vittorio Papadopoulos, Thomas Nolte, Mohammad Ashjaei |
ETFA | 2 |
| 2023 | Opportunities and Specific Plans for Migrating from PRP to TSN in Substation Automation SystemsabstractElectrical substations are vital for the power grid, and Substation Automation Systems (SASs) have been employed to enhance substation functionality and safety. As the energy landscape evolves, substations face new challenges such as accommodating an increasing number of prosumers. Thus, SASs require a reliable substation communication network (SCN) capable of supporting real-time control and diverse applications. While Ethernet-based SCN technologies have emerged, they often fall short in meeting all requirements, including TCP/IP support, cost-effective fault tolerance, and managing traffic with different real-time demands. Time-Sensitive Networking (TSN) standards have shown promise in addressing these limitations by providing novel mechanisms. In this paper we compare TSN with the Parallel Redundancy Protocol (PRP) demonstrating that TSN offers better functionality and efficiency. In the direction of designing a comprehensive TSN-based architecture for SASs’ Distributed Control Systems (DCSs) we start here by proposing a roadmap for the fault tolerance aspects. Mateu Jover, Manuel Barranco, Julián Proenza |
ETFA | 3 |
| 2022 | Implementing a First CNC for Scheduling and Configuring TSN NetworksabstractNovel industrial applications are leading to important changes in industrial systems. One of the most important changes is the need for systems that are capable to adapt to changes in the environment or the system itself. Because of their nature many of these applications are distributed, and their network infrastructure is key to guarantee the correct operation of the overall system. Furthermore, in order for a distributed system to be able to adapt, its network must be flexible enough to support changes in the traffic during runtime. The Time-Sensitive Networking (TSN) Task Group has proposed a series of standards that aim at providing deterministic real-time communications over Ethernet. TSN also provides centralised online configuration and control architectures which enable the online configuration of the network. A key part in TSN’s centralised architectures is the Centralised Network Configuration element (CNC). In this work we present a first implementation of a CNC capable of scheduling time-triggered traffic and deploying such configuration in the network using the Network Configuration (NETCONF) protocol. We also assess the correctness of our implementation using an industrial use case provided by Volvo Construction Equipment. Ines Alvarez, Andreu Servera, Julián Proenza, Mohammad Ashjaei, Saad Mubeen |
ETFA | 3 |
| 2022 | The Effects of Clock Synchronization in TSN Networks with Legacy End-StationsabstractIn this paper, we present our ongoing work on proposing solutions to integrate legacy end-stations into Time-Sensitive Network (TSN) communication systems where the legacy end-stations are synchronized via their legacy clock synchronization protocol. To this end, we experimentally identify the effects of lacking synchronization or partial synchronization in TSN networks. In the experiments we show the effects of clock synchronization in different scenarios on jitter and clock drifts. Based on the experiments, we propose preliminary solutions to overcome the identified effects. Daniel Bujosa, Andreas Johansson, Mohammad Ashjaei, Alessandro Vittorio Papadopoulos, Julián Proenza, Thomas Nolte |
ETFA | 5 |
| 2022 | Migrating Legacy Ethernet-Based Traffic with Spatial Redundancy to TSN networksabstractDistributed Control Systems (DCSs) for emerging industrial control applications impose new communication requirements that cannot be satisfied by current Industrial Ethernet protocols. As a result, industry is pushing the Time-Sensitive Networking (TSN) standards as the de-facto Ethernet-based linklayer to fulfill these requirements. Adequate roadmaps are needed to support a smooth transition from Industrial-Ethernet-based legacy systems to TSN-based ones. In this context some works propose mechanisms to migrate, i.e. map, route and schedule, legacy traffic to TSN. However none of them considers traffic including streams with spatial redundancy requirements and, thus, they cannot be used to migrate legacy highly-reliable DCSs. The present work extends a previous toolchain to migrate, for the first time, legacy critical traffic that includes spatially redundant streams. Particularly, since redundancy is costly, this work proposes and compares two routing methods that consider one redundant stream per traffic. Mateu Jover, Manuel Barranco, Ines Alvarez, Julián Proenza |
ETFA | 4 |
| 2021 | LETRA: Mapping Legacy Ethernet-Based Traffic into TSN Traffic ClassesabstractThis paper proposes a method to efficiently map the legacy Ethernet-based traffic into Time Sensitive Networking (TSN) traffic classes considering different traffic characteristics. Traffic mapping is one of the essential steps for industries to gradually move towards TSN, which in turn significantly mitigates the management complexity of industrial communication systems. In this paper, we first identify the legacy Ethernet traffic characteristics and properties. Based on the legacy traffic characteristics we presented a mapping methodology to map them into different TSN traffic classes. We implemented the mapping method as a tool, named Legacy Ethernet-based Traffic Mapping Tool or LETRA, together with a TSN traffic scheduling and performed a set of evaluations on different synthetic networks. The results show that the proposed mapping method obtains up to 90% improvement in the schedulability ratio of the traffic compared to an intuitive mapping method on a multi-switch network architecture. Daniel Bujosa, Mohammad Ashjaei, Alessandro Vittorio Papadopoulos, Julián Proenza, Thomas Nolte |
ETFA | 4 |
| 2021 | Exploring the use of Deep Reinforcement Learning to allocate tasks in Critical Adaptive Distributed Embedded SystemsabstractCritical Adaptive Distributed Embedded Systems (CADES) must carry out a set of funcionalities while fulfilling their associated real-time and dependability requirements. Moreover, they must be able to reconfigure themselves in a bounded time as the operational context changes. Finding a proper configuration can be non-trivial and time-consuming. Several studies have proposed Deep Reinforcement Learning (DRL) approaches to solve combinatorial optimization problems. In this paper, we explore the application of such approaches to CADES by solving a simple tasks allocation problem using DRL and comparing the results with three popular heuristics. The results show that DRL beats two of them and gets very close to the third, while requiring significantly less time to generate a solution. Ramón Rotaeche, Alberto Ballesteros, Julián Proenza |
ETFA | 3 |
| 2021 | CSRP: An Enhanced Protocol for Consistent Reservation of Resources in AVB/TSNabstractThe IEEE Audio Video Bridging (AVB) Task Group (TG) was created to provide Ethernet with soft real-time guarantees. Later on, the TG was renamed to Time-Sensitive Networking (TSN) and its scope broadened to support hard real-time and critical applications. The Stream Reservation Protocol (SRP) is a key work of the TGs as it allows reserving resources in the network, guaranteeing the required quality of service. AVB's SRP is based on a distributed architecture, whereas TSN's is based on centralized ones. The distributed version of SRP is supported and used in TSN. Nevertheless, it was not designed to provide properties that are important for critical applications. In this article, we model SRP using UPPAAL and we study the termination and consistency. We verify that SRP does not provide such properties. Furthermore, we propose an improved protocol called Consistent Stream Reservation Protocol (CSRP) and we formally verify its correctness using UPPAAL. Daniel Bujosa, Ines Alvarez, Julián Proenza |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Clock Synchronization in Integrated TSN-EtherCAT NetworksabstractMoving towards new technologies, such as Time Sensitive Networking (TSN), in industries should be gradual with a proper integration process instead of replacing the existing ones to make it beneficial in terms of cost and performance. Within this context, this paper identifies the challenges of integrating a legacy EtherCAT network, as a commonly used technology in the automation domain, into a TSN network. We show that clock synchronization plays an essential role when it comes to EtherCAT-TSN network integration with important requirements. We propose a clock synchronization mechanism based on the TSN standards to obtain a precise synchronization among EtherCAT nodes, resulting to an efficient data transmission. Based on a formal verification framework using UPPAAL tool we show that the integrated EtherCAT-TSN network with the proposed clock synchronization mechanism achieves at least 3 times higher synchronization precision compared to not using any synchronization. Daniel Bujosa, Daniel Hallmans, Mohammad Ashjaei, Alessandro Vittorio Papadopoulos, Julián Proenza, Thomas Nolte |
ETFA | 5 |
| 2019 | Simulation of the Proactive Transmission of Replicated Frames Mechanism over TSNabstractThe Time-Sensitive Networking (TSN) Task Group (TG) is providing Ethernet with timing guarantees, reconfiguration services and fault tolerance mechanisms. Some of TSN's targeted applications are real-time critical applications, which must provide a correct service continuously. To support these applications the TSN TG standardised a spatial redundancy mechanism. Even though spatial redundancy can tolerate permanent and temporary faults, it is not cost-effective. Instead, temporary faults can be tolerated using time redundancy. We proposed the Proactive Transmission of Replicated Frames (PTRF) mechanism to tolerate temporary faults in the links. In this work we present a new PTRF approach, a PTRF simulation model and a comparison of the approaches using exhaustive fault injection. Ines Alvarez, Drago Cavka, Julián Proenza, Manuel Barranco |
ETFA | 3 |
| 2019 | Temporal Replication of Messages for Adaptive Systems using a Holistic ApproachabstractCritical Adaptive Distributed Embedded Systems (ADES) must meet high real-time and dependability requirements, while autonomously rearranging themselves to operate in dynamic operational contexts. The DFT4FTT project proposes a self-reconfigurable complete infrastructure, whose different architectural levels provide a set of real-time (RT), fault-tolerance (FT) and flexibility mechanisms that collaborate to adequately support critical ADESs. To efficiently tolerate transient faults in the network of an ADES, this paper describes our ongoing work on providing a dynamic temporal replication of messages that takes into account all the DFT4FTT fault-tolerance mechanisms from a holistic point of view. Alberto Ballesteros, Manuel Barranco, Sergi Arguimbau, Marc Costa, Julián Proenza |
ETFA | 5 |
| 2019 | Formal Verification of the FTTRS Mechanisms for the Consistent Update of the Traffic ScheduleabstractCritical Adaptive Distributed Embedded Systems (ADESs) are nowadays the focus of many researchers. ADESs are envisioned to dynamically modify their behavior to support changes of their real-time and dependability requirements at runtime as the conditions of the environment in which they operate vary. To provide ADESs with an adequate communication infrastructure, our research group proposed the Flexible-Time-Triggered Replicated Star (FTTRS). FTTRS provides highly reliable communication services on top of Ethernet, while keeping the adaptivity benefits that the Flexible-Time-Triggered (FTT) communication paradigm offers from a real-time perspective. This paper formally verifies, by means of model checking, the correctness of the mechanisms FTTRS includes to enforce consistent changes of the communication scheduling at runtime. Daniel Bujosa, Sergi Arguimbau, Patricia Arguimbau, Julián Proenza, Manuel Barranco |
ETFA | 4 |
| 2019 | Analysing Termination and Consistency in the AVB's Stream Reservation ProtocolabstractThe Audio Video Bridging Task Group (AVB TG) from the IEEE proposed a series of standards to provide Ethernet with soft real-time guarantees. Later on, the group was renamed to Time-Sensitive Networking and its scope was broadened to provide new services to support critical applications. The Stream Reservation Protocol (SRP) stands out among the projects developed by the groups. Nonetheless, SRP was originally designed for audio/video applications and does not take into account properties that are important for critical systems; such as termination and consistency. In this work we study the termination and consistency of SRP at different levels, using a model we developed of this protocol in Uppaal. We see that SRP does not provide termination nor consistency, we discuss how this can impact critical applications and we propose solutions for all the issues detected. Daniel Bujosa, Ines Alvarez, Drago Cavka, Julián Proenza |
ETFA | 4 |
| 2019 | First exploration of the potential of diverse training and voting for increasing the accuracy of CNNsabstractMachine learning techniques are attracting a huge amount of interest from both industry and academia. For instance, Convolutional deep Neural Networks (CNNs) have recently enjoyed a notable success in image understanding. The automotive industry is already using image classifiers for Advanced Driver-Assistance Systems and in the development of the upcoming autonomous cars, which will have to guarantee high levels of reliability. The certification of systems based on machine learning is an open issue but it is clear that any improvement in the performance of image classifiers is to be welcomed. CNNs need to be trained to act as image classifiers. This training leads to slightly different classification capacity depending on some training parameters. In this paper we present a first exploration on the use of schemes based on voting on the results of several CNNs trained differently, as a means to increase the final classification performance, and thus the reliability, of this type of systems. Julián Proenza, Yolanda González Cid, Patricia Arguimbau |
ETFA | 1 |
| 2019 | Fault Tolerance in Highly Reliable Ethernet-Based Industrial SystemsabstractMany industrial systems have specific requirements derived from the applications they execute. Specifically, the interaction of a distributed embedded control system (DECS) with the real-world imposes strict real-time (RT) and reliability requirements. For a system to be RT, it has to produce a proper result in a bounded time. On top of that, for a system to be reliable, it has to operate continuously during its mission time, and in cases in which very high reliability is needed, fault tolerance (FT) techniques are used. Moreover, these systems are often deployed in dynamic environments where the operational conditions may change in an unpredictable manner. Therefore, there is an increasing interest in creating DECSs that are capable of modifying their behavior autonomously and dynamically in response to unexpectedly changing requirements or conditions. In recent years, there is a growing trend toward using Ethernet as the network technology for DECSs. Unfortunately, the original specification of this technology lacks appropriate services to fulfill the most demanding requirements of industrial systems. In this regard, many Ethernet-based protocols and standards have been proposed along the past years to deal with these limitations. In this paper, we survey solutions that have been proposed to achieve FT in Ethernet-based DECSs, considering faults both in their nodes and communication subsystem. In addition, we discuss adaptive FT techniques that can be used to increase the flexibility of adaptive DECS. Finally, we identify future trends and open challenges to build highly reliable DECS in the future. Ines Alvarez, Alberto Ballesteros, Manuel Barranco, David Gessner, Sinisa Derasevic, Julián Proenza |
Proc. IEEE | 6 |
| 2019 | A Fault-Tolerant Ethernet for Hard Real-Time Adaptive SystemsabstractDistributed embedded systems (DESs) that perform critical tasks in unpredictable environments must be reliable, hard real-time, and adaptive. Since a DES comprises nodes that rely on a network, the network must provide adequate support: it must be reliable, convey messages on time, and meet new real-time requirements as the nodes adapt. Ethernet is ill-suited for such hard real-time adaptive systems, but it can be made suitable. The flexible time-triggered (FTT) paradigm already supports hard real-time message exchanges and the necessary flexibility to meet evolving hard real-time requirements, but its Ethernet implementations had reliability limitations. To address these, we designed FTT replicated star for Ethernet (FTTRS), a communication subsystem that tolerates permanent and transient faults, even if they occur simultaneously, while keeping the paradigm's key features: support for both the timely exchange of periodic and sporadic real-time messages, and support for updating the real-time parameters of these messages at runtime. In this paper, we present FTTRS, the first Ethernet-based communication subsystem specifically designed for highly reliable hard real-time adaptive DESs. David Gessner, Julián Proenza, Manuel Barranco, Alberto Ballesteros |
IEEE Trans. Ind. Informatics | 2 |
| 2018 | Towards a Fault-Tolerant Architecture Based on Time Sensitive NetworkingabstractThe Time Sensitive Networking (TSN) Task Group has been working on describing a set of standards that will provide enhanced capabilities to standard Ethernet. Specifically, they work to provide Ethernet with real-time, reliability and reconfiguration capacities. Nevertheless, this set of standards (commonly referred to as TSN) does not cover some reliability aspects that are relevant for the correct operation of critical distributed control systems. Thus, in this work we present a first proposal of a highly reliable architecture and a set of mechanisms based on TSN to support the real-time and reliability requirements of these critical systems. Ines Alvarez, Manuel Barranco, Julián Proenza |
ETFA | 3 |
| 2017 | Towards a time redundancy mechanism for critical frames in time-sensitive networkingabstractTime-Sensitive Networking (TSN) is a set of technical standards that is being developed to provide Ethernet with hard real-time, reliability and flexibility services. In the last years, there has been a growing interest in increasing the connectivity of all kind of devices. This trend has reached industrial environments, where the demanding timing and reliability constraints imposed the use of specialised networks with specific features to support these requirements. Moreover, the industry has shown interest in using Ethernet as the network technology in industrial environments, due to its low cost, high bandwidth and extensive use. The ability of TSN to support both, data-oriented and traditional control traffic over the same network makes it an appealing technology to implement the next generation of industrial networks with high connectivity. Nevertheless, TSN does not cover some reliability aspects important for its deployment in critical systems. In this work we propose the implementation of time redundancy of frames in order to tolerate temporary faults in the channel and, therefore, increase the reliability of the network. Ines Alvarez, Julián Proenza, Manuel Barranco, Mladen Knezic |
ETFA | 2 |
| 2017 | Towards a dynamic task allocation scheme for highly-reliable adaptive distributed embedded systemsabstractAn adaptive distributed embedded system is able to automatize some processes at the same time it modifies its behaviour autonomously and dynamically in response to changing operating conditions. To support adaptivity it is necessary that the underlying Distributed Embedded System (DES) is able to dynamically change the assignment of the processing and network resources. In this regard, the DFT4FTT project aims at providing a complete DES that can support applications with real-time, reliability and adaptivity requirements. This paper describes the first steps towards the design of the task allocation scheme used in the DFT4FTT architecture, responsible for dynamically distributing the workload among the nodes of the DES, taking into account the changes in the environment and in the system itself. This allocation scheme not only provides flexibility from a functional point of view, but also from a fault tolerance point of view. Moreover, its modular design makes it possible to tune the desired level of autonomy in the adaptivity, from a simple support for application reconfiguration to a complete automatic reconfiguration assisted with machine learning algorithms. Alberto Ballesteros, Julián Proenza, Pere Palmer |
ETFA | 2 |
| 2016 | A first performance analysis of the Admission Control in the HaRTES Ethernet switchabstractThere is a growing interest in developing embedded systems capable of being deployed in dynamic environments that may change in unpredictable manners. When such systems are Distributed Embedded Systems (DESs) they must exhibit flexibility at all levels of their architecture, including the network. On the other hand, there is a clear trend in industry towards using Ethernet-based protocols at the network level of DESs. Nevertheless, Ethernet lacks appropriate support for real-time (RT) communications, mixing different RT traffic and on-line management of the Quality of Service (QoS). Several implementations of the Flexible Time-Tiggered (FTT) protocol over Ethernet were proposed to cope with these drawbacks. FTT is a master/multi-slave protocol that is able to simultaneously convey real and non-real-time traffic and provides mechanisms for dynamically changing the QoS of the network, including Admission Control (AC). The AC is a fundamental component for on-line network management, since it guarantees that each participant gets the required QoS. This paper presents the implementation in OMNeT++ of a simulation model of the AC in the FTT HaRTES switch as well as a preliminary performance study using that model. Ines Alvarez, Mladen Knezic, Luís Almeida 0001, Julián Proenza |
ETFA | 4 |
| 2016 | First implementation and test of reintegration mechanisms for node replicas in the FT4FTT ArchitectureabstractDistributed Embedded Control Systems (DECSs) used for critical applications must usually abide by strict real-time and dependability requirements. Correspondingly, the FT4FTT project proposes a complete fault-tolerant (FT) architecture for RT DECSs. The Flexible Time-Triggered Ethernet (FTT-Ethernet) communication protocol fulfills the RT requirements, while the FT mechanisms added on top of it, which are based on channel duplication and active replication of nodes, provide the FT behaviour. Temporary faults affecting the channel or the nodes, which are the most probable type of faults in DESs, can manifest in such a way that a node replica loses its coordination with the others and, thereby, it also loses its communication and/or computation capability from then on, leading to attrition of the redundancy initially provided by the active replication of nodes. This paper describes the implementation and test of specific mechanisms that are devised to determine which replicas are temporarily faulty and to promptly reintegrate them. Alberto Ballesteros, Sinisa Derasevic, Manuel Barranco, Julián Proenza |
ETFA | 4 |
| 2016 | Improving maintenance of FT4FTT: Extending it to monitor and log its available redundancy via internetabstractThe FT4FTT project aims at proposing a complete Fault-Tolerant (FT) architecture for Real-Time (RT) critical adaptative Distributed Embedded Control Systems (DECSs) based on Ethernet. FT4FTT tolerates permanent faults in the channel and nodes by using a duplicated Flexible Time-Triggered (FTT) Switched Ethernet star and active replication of the nodes. It also includes mechanisms for node replicas to diagnose and reintegrate after temporary faults affecting the channel or their internal circuitry. However, FT4FTT has no mechanism to deal with channel and node redundancy attrition provoked by permanent faults. This paper presents our ongoing work to extend FT4FTT to both monitor/log its available redundancy, and to remotely access this information via Internet. This will allow to carry out proper maintenance actions, for instance, to timely restore the adequate redundancy level, forecast repairs, and assess the flexibility of the FT mechanisms of adaptative systems. Manuel Barranco, Adel Zendouh, Alberto Ballesteros, Julián Proenza |
ETFA | 4 |
| 2016 | A first qualitative comparison of the admission control in FTT-SE, HaRTES and AVBabstractEthernet is gaining importance in fields such as automation, avionics and automotive. In these fields novel multimedia-based applications must coexist with traditional control systems, which leads to high diversity in size, intensity and timing requirements of the traffic traversing the channel. Multimedia traffic is characterised by having large size, low intensity and soft real-time requirements, while control traffic usually conveys small amounts of information with a high intensity and hard real-time requirements. Moreover, many modern applications must support on-line connection and disconnection of participants. Since Ethernet was designed as a general purpose data network protocol it lacks appropriate support for real-time communications and dynamic quality of service management. Several protocols were proposed to cope with these drawbacks, including Flexible Time-Triggered Switched Ethernet and, more recently, Audio Video Bridging. In this paper we discuss the importance of the admission control and make a comparison of the implementations carried out in the aforementioned protocols. Ines Alvarez, Luís Almeida 0001, Julián Proenza |
WFCS | 3 |
| 2016 | First implementation and test of a node replication scheme on top of the flexible time-triggered replicated star for ethernetabstractDistributed embedded systems typically have real-time and dependability requirements. Moreover, they must also be flexible to changing conditions when they are deployed in dynamic environments. The FT4FTT project aims at providing a switched Ethernet architecture that can support distributed control applications that are predictable, highly-reliable and adaptive. FT4FTT relies on the Flexible Time-Triggered Replicated Star for Ethernet (FTTRS) to tolerate channel faults. Moreover, nodes' hardware faults are tolerated by means of active node replication with majority voting. In order to coordinately trigger the execution of the tasks in the replicas, we designed the CD4NR mechanism, in which the network assists in deciding what to execute and when. This paper presents the first implementation of the CD4NR mechanism on a real prototype of FTTRS and the first testing of the complete system. For this we developed an experimental setup, based on the hardware-in-the-loop technique, running a real-time control application. Alberto Ballesteros, Sinisa Derasevic, David Gessner, Francisca Font, Ines Alvarez, Manuel Barranco, Julián Proenza |
WFCS | 7 |
| 2016 | Designing fault-diagnosis and reintegration to prevent node redundancy attrition in highly reliable control systems based on FTT-EthernetabstractDistributed Embedded Control Systems (DECSs) used for Real-Time (RT) critical applications must satisfy stringent time requirements and attain high reliability. FTT-Ethernet provides nodes of DECSs with real-time communication capabilities, but does not include Fault Tolerance (FT) mechanisms. The FT4FTT project aims at proposing a complete FT architecture for RT critical DECSs. It uses a duplicated switched FTT-Ethernet star and active node replication with consistent distributed majority voting to respectively tolerate channel and node faults. However, FT4FTT, in its current state, still lacks mechanisms to prevent node redundancy attrition due to temporary faults affecting the nodes and channel, which are the most likely types of faults in DESs. This paper presents our ongoing work to complete the FT4FTT architecture with appropriate fault-diagnosis and reintegration mechanisms that overcome this limitation. Sinisa Derasevic, Manuel Barranco, Julián Proenza |
WFCS | 3 |
| 2016 | Guest Editorial Special Section on Communication in AutomationabstractAutomation systems are composed of tightly integrated mechanical, electronic, and computer equipment being designed to reduce or even eliminate the need for human intervention in the realization of many different tasks. Also, the automation of systems and processes is often characterized by some interesting features, such as reduced exploration costs and intrinsically higher efficiency, safety, and quality. The papers in this special focus on the use of automation in communication systems. Stefano Vitturi, Paulo Pedreiras, Julián Proenza, Thilo Sauter |
IEEE Trans. Ind. Informatics | 3 |
| 2015 | An OMNET++ model to asses node fault-tolerance mechanisms for FTT-Ethernet DESsabstractDistributed embedded systems (DESs) that operate in dynamic environments require emerging flexibility and adaptivity communication requirements. When those DESs are deployed for critical applications, they must also employ appropriate fault-tolerance (FT) mechanisms to attain a high level of reliability. The FTT-Ethernet communication protocol supports the flexibility needed in dynamic environments, but does not provide adequate fault tolerance. In order to overcome this limitation the ongoing FT4FTT project proposes a communication architecture that includes fault-tolerance capabilities at different levels of DESs relying on FTT-Ethernet. In particular, it provides communication and execution mechanisms to tolerate node failures by means of active node replication with majority voting. This paper builds upon a previous OMNET++ model of an FTT-Ethernet-based DES in order to add, simulate and assess those mechanisms. Specifically, it models the communication mechanisms envisaged to enforce replica determinism in the voting procedure, as well as to trigger and coordinate the tasks executed in the replicas. Sinisa Derasevic, Manuel Barranco, Julián Proenza |
ETFA | 3 |
| 2015 | First experimental evaluation of the consistent replicated voting in the hard real-time ethernet switching architectureabstractDistributed Embedded Systems (DESs) typically have dependability and real-time requirements. Moreover, when they are deployed in dynamic environments, they must be flexible enough to adapt to changes in the operation requirements. The Fault Tolerance for Flexible Time-Triggered Ethernet (FT4FTT) project aims at providing a Switched-Ethernet architecture, based on the Flexible Time-Triggered communication paradigm (FTT), that is flexible and highly reliable. In particular, FT4FTT provides node fault-tolerance by means of active replication with majority voting. In this sense, FT4FTT includes the Consistent Replicated Voting (CRV) protocol to enforce replica determinism, even in presence of faults, while maximizing the reliability that can be achieved thanks to the node redundancy and the communication subsystem itself. This papers presents a first implementation of this protocol in a real prototype, and shows the on-going experimental evaluation been carried out to asses its correctness. Sinisa Derasevic, Maties Melia, Alberto Ballesteros, Manuel Barranco, Julián Proenza |
ETFA | 5 |
| 2015 | Towards a layered architecture for the Flexible Time-Triggered Replicated Star for EthernetabstractDistributed embedded systems (DES) have traditionally been designed assuming that the requirements they need to satisfy are known in advance. If this is not the case, and a DES should operate autonomously without interruption, it needs to be adaptive. For this, flexible approaches are necessary and this applies in particular to the network of the DES. However, if the probability of faults occurring is non-negligible, then flexibility alone is not enough and fault tolerance is also necessary. The Flexible Time-Triggered Replicated Star for Ethernet (FTTRS) is a set of protocols and mechanisms together with a specific network topology that builds on a switched-Ethernet implementation of the Flexible Time-Triggered (FTT) communication paradigm by enhancing it to not only provide flexibility, but also fault tolerance. This paper describes our efforts towards a layered architecture for FTTRS to benefit from the well-known advantages of these architectures, such as making the complexity manageable and easier to communicate, and making the design more future proof by allowing changes in one layer without affecting other layers. David Gessner, Ignasi Furió, Julián Proenza |
ETFA | 3 |
| 2015 | Experimental evaluation of network component crashes and trigger message omissions in the Flexible Time-Triggered Replicated Star for EthernetabstractA distributed embedded system (DES) is made up of a set of computing nodes interconnected by a network. If we want the DES to continue to operate even if a subset of its network elements fail, the network must be fault-tolerant. In particular, this requires that the architecture of the network provides redundant paths between nodes and that any elements critical for the operation of the network are replicated. In the context of DES that must not only be highly reliable, but also provide sufficient flexibility to adapt to unpredictable requirement changes, the Flexible Time-Triggered Replicated Star for Ethernet (FTTRS) has been proposed. One of the core features of FTTRS is precisely its fault-tolerant network architecture. In this paper we present a proof-of-concept prototype of FTTRS and a series of fault-injection experiments. These experiments show that FTTRS can tolerate the crash of any single network element, as well as the crash of various combinations of multiple network elements. A variety of omission failures affecting the most critical FTTRS message (called the trigger message) are also tolerated. David Gessner, Alberto Ballesteros, Andreu Adrover, Julián Proenza |
WFCS | 4 |
| 2014 | Achieving elementary cycle synchronization between masters in the flexible time-triggered replicated star for ethernetabstractFor a distributed embedded system (DES) to operate continuously in a dynamic environment, it must be flexible and highly reliable. This applies in particular to its communication subsystem. The Flexible Time-Triggered Replicated Star for Ethernet (FTTRS) aims at providing such a subsystem by means of a highly-reliable switched-Ethernet architecture based on the Flexible Time-Triggered paradigm (FTT), a master/slave communication paradigm where the master periodically polls the slaves using so-called trigger messages (TMs). In particular, FTTRS interconnects nodes by redundant communication paths provided by two switches, each embedding an FTT master that manages the communication. This allows FTTRS to tolerate the failure of one switch without interrupting the communication as long as the masters are replica determinate, i.e., provide identical service to the slaves. The master replica determinism entails the masters broadcasting their TMs in a lockstep fashion: when one master broadcasts a TM, the other should do the same quasi-simultaneously. In this paper we present a solution inspired by the Precision Time Protocol (PTP) for achieving this lockstep transmission and preliminary results showing the precision with which we can synchronize the masters on a software prototype. Alberto Ballesteros, Julián Proenza, David Gessner, Guillermo Rodríguez-Navas, Thilo Sauter |
ETFA | 2 |
| 2014 | A model for quantifying the reliability of highly-reliable distributed systems based on fieldbus replicated busesabstractDespite the efforts devoted to increase the dependability of highly-reliable distributed fieldbus systems by means of simplex stars and replicated stars/buses, literature lacks of appropriate analyses that quantify the system reliability these topologies yield. In previous work, we proposed models to adequately quantify the system reliability benefits of simplex buses and simplex/replicated stars. However, a model for replicated buses is an open issue that needs to be addressed, as they normally include less components than stars and, thus, can be more reliable and cost-effective. To fill this gap, this paper presents a model that makes it possible to appropriately quantify the reliability that a highly-reliable distributed system can achieve when using a replicated bus. Manuel Barranco, Francisco Pozo, Julián Proenza |
ETFA | 3 |
| 2014 | Appropriate consistent replicated voting for increased reliability in a node replication scheme over FTTabstractIn the context of critical applications there is an increasing interest in having Distributed Embedded Systems (DESs) that are able to operate in dynamic environments, while at the same time reaching a high reliability. The Flexible Time-Triggered communication paradigm (FTT) is designed to support the QoS and real-time requirements of the traffic of these systems. However, FTT does not provide fault tolerance. This paper explains our on-going work towards designing a consistent and highly-reliable voting protocol which supports node replication on DESs that use FTT switched Ethernet. In particular, we propose a protocol for the node replicas to vote consistently on messages exchanged through an FTT Ethernet network that uses time redundancy, while trying to maximize the reliability that can be achieved thanks to the redundancy of the nodes and the communication subsystem itself. Sinisa Derasevic, Manuel Barranco, Julián Proenza |
ETFA | 3 |
| 2014 | Using FTT-ethernet for the coordinated dispatching of tasks and messages for node replicationabstractThe Flexible Time Triggered (FTT) paradigm provides online flexible scheduling for distributed embedded systems but it does not present adequate fault tolerance mechanisms so as to reach a very high reliability. Adding the adequate fault tolerance mechanisms to FTT-based architectures would open room for adaptive yet highly dependable systems. In this work we present a fault-tolerant system architecture for control applications that adds a node replication scheme with voting on top of an FTT-based system. Using a previously proposed network-centric approach we show how to coordinate the execution of the different phases for a typical control application in our system architecture, i.e. we show how to trigger the execution of tasks in node replicas and the transmission of messages in the communication channel, using the underlying FTT protocol. At the end, we demonstrate how to apply this idea of coordinated dispatching to one concrete control application, ball-on-plate. Sinisa Derasevic, Julián Proenza, Manuel Barranco |
ETFA | 2 |
| 2014 | Towards an experimental assessment of the slave elementary cycle synchronization in the Flexible Time-Triggered Replicated Star for EthernetabstractThe communication subsystem of distributed embedded systems (DES) that must operate continuously and satisfy unpredictable requirement changes must be reliable and flexible. Recently the Flexible Time-Triggered Replicated Star for Ethernet (FTTRS) has been proposed as a communication subsystem that satisfies these two attributes. It is based on the master/multi-slave Flexible-Time Triggered (FTT) communication paradigm and relies on two custom switches, each with its own embedded FTT master. Both masters are active simultaneously and provide the same service. Specifically, they simultaneously and periodically broadcast so-called trigger messages (TMs) in a redundant manner to make them robust to transient channel faults. One of the functions of these TMs is to divide the communication time into rounds called elementary cycles (ECs). For the correct operation of FTTRS, it is important that all slaves agree when each EC starts and ends. A mechanism to achieve this has been recently proposed. This paper presents a first implementation of this mechanism and a series of experimental tests that constitute a first step towards building a prototype of an FTTRS network. David Gessner, Ines Alvarez, Alberto Ballesteros, Manuel Barranco, Julián Proenza |
ETFA | 5 |
| 2014 | Towards a reliability analysis of the design space for the communication subsystem of FT4FTTabstractFault Tolerance for Flexible Time-Triggered Ethernet-based systems (FT4FTT) is a project to devise an architecture for distributed embedded systems that provides both flexibility to changing real-time requirements and high reliability through fault tolerance. One of the key parts of such an architecture is the communication subsystem. When designing such a subsystem many decisions have to be made. To understand how such decisions impact the reliability of the final design, in this paper we present a framework to evaluate the reliability of a large number of potential designs. The approach is based on storing a finite subset of the design space for the communication subsystem of FT4FTT in an undirected graph and then generating a continuous-time Markov chain from the graph to evaluate the reliability of each design belonging to the subset. David Gessner, Julián Proenza, Manuel Barranco, Paulo Portugal |
ETFA | 2 |
| 2014 | Towards extending the OMNeT++ INET framework for simulating fault injection in ethernet-based Flexible Time-Triggered systemsabstractTraditional distributed embedded systems are configured using static environment information and thus do not support dynamic behavior of the system. The necessary flexibility in the system may be provided by the Flexible Time-Triggered (FTT) communication paradigm. If, in addition, it is required that the system operates continuously, the suitable fault tolerance mechanisms that provide high reliability have to be developed and deployed in the system. To be able to successfully assess those mechanisms, it is reasonable to develop simulation models that support injection of various types of faults. This paper describes an OMNeT++ simulation model for distributed systems that are based on the Hard Real-Time Ethernet Switching (HaRTES) implementation of the FTT paradigm. The contribution of the paper is twofold. First, we provide a library of components that are required for modeling FTT networks with arbitrary number of FTT slaves connected to a HaRTES switch, and second, we used the developed components to build an FTT system that is suitable for assessing some of recently proposed mechanisms for tolerating certain transient faults in the communication channel. Mladen Knezic, Alberto Ballesteros, Julián Proenza |
ETFA | 3 |
| 2013 | Towards Efficient Probabilistic Scheduling Guarantees for Real-Time Systems Subject to Random Errors and Random Bursts of ErrorsabstractReal-time computing and communication systems are often required to operate with prespecified levels of reliability in harsh environments, which may lead to the exposure of the system to random errors and random bursts of errors. The classical fault-tolerant schedulability analysis in such cases assumes a pseudo-periodic arrival of errors, and does not effectively capture any underlying randomness or burst characteristics. More modern approaches employ much richer stochastic error models to capture these behaviors, but this is at the expense of greatly increased complexity. In this paper, we develop a quantile-based approach to probabilistic schedulability analysis in a bid to improve efficiency whilst still retaining a rich stochastic error model capturing random errors and random bursts of errors. Our principal contribution is the derivation of a simple closed-form expression that tightly bounds the number of errors that a system must be able to tolerate at any time subsequent to its critical instant in order to achieve a specified level of reliability. We apply this technique to develop an efficient 'one-shot' schedulability analysis for a simple fault-tolerant EDF scheduler. The paper concludes that the proposed method is capable of giving efficient probabilistic scheduling guarantees, and may easily be coupled with more representative higher-level job failure models, giving rise to efficient analysis procedures for safety-critical fault-tolerant real-time systems. Michael Short 0001, Julián Proenza |
ECRTS | 2 |
| 2013 | Towards preventing error propagation in a real-time Ethernet switchabstractFlexible Time-Triggered communication (FTT) allows a distributed embedded system (DES) to adapt to changing real-time requirements at runtime. This facilitates the continuous operation of the DES under dynamic environments that change over time. However, for continuous operation, high reliability in the nodes of the DES is also crucial. This can be achieved using node replication, as long as failure independence between replicas is ensured, which calls for preventing the propagation of errors. Our goal is to prevent the propagation of Byzantine node behaviours and to ensure that local errors in the channel cannot disturb the global communication. For this, we construct the HaRTES/PG switch, a new switch based on the HaRTES implementation of FTT for Ethernet. This paper presents as a first step a study of the possible errors that may lead to Byzantine node behaviours and a global communication disturbance in HaRTES, as well as some ideas on how to prevent the propagation of these errors in HaRTES/PG. Alberto Ballesteros, David Gessner, Julián Proenza, Manuel Barranco, Paulo Pedreiras |
ETFA | 3 |
| 2013 | Towards dynamic fault tolerance on FTT-based distributed embedded systemsabstractDistributed embedded systems (DES) have been traditionally designed to operate in static environments that do not change over time. Flexible designs are increasingly being introduced to achieve continuous and correct operation under dynamic environments. Some designs, such as the Flexible Time-Triggered communication paradigm (FTT), are focused on being able to modify the real-time operation upon changing requirements imposed by the environment. The on-going project Fault Tolerance for FTT (FT4FTT) purports to increase the reliability of a DES based on the FTT protocol by introducing static fault tolerance. In this paper we give some hints on how to go one step beyond by adding dynamic fault tolerance to the DES. This would result in new systems that would combine the qualities of flexible real-time operation and flexible and adaptive fault tolerance, much enlarging their sphere of applicability. Sinisa Derasevic, Julián Proenza, David Gessner |
ETFA | 2 |
| 2013 | Towards a Flexible Time-Triggered replicated star for ethernetabstractDistributed embedded systems have traditionally been designed using static approaches, i.e., assuming a static environment. Such approaches, however, cannot guarantee continuous operation under dynamic environments that impose new requirements upon a system as time passes. As a solution, flexible approaches have been proposed. One such approach that allows a system to adapt to changing real-time requirements is the Flexible Time-Triggered (FTT) communication paradigm. Nevertheless, if continuous operation under dynamic environments is desired, then flexibility is not enough. Indeed, it is also crucial for the system to be sufficiently reliable. In this paper we therefore explore some design ideas to make FTT highly reliable through fault tolerance by using replication. As a starting point we will use the switch of the Hard Real-Time Ethernet Switching (HaRTES) implementation of FTT. David Gessner, Julián Proenza, Manuel Barranco, Luís Almeida 0001 |
ETFA | 2 |
| 2013 | A proposal for flexible, real-time and consistent multicast in FTT/HaRTES Switched EthernetabstractHard Real-Time Ethernet Switching (HaRTES) is an implementation of the Flexible Time Triggered (FTT) communication paradigm over Switched Ethernet, which intends to provide hard real-time communication in a flexible manner. This paper presents a first proposal for enhancing HaRTES with a service of total order multicast for synchronous messages. This service uses the centralized online scheduling service of FTT in order to reduce complexity and bandwidth utilization. Guillermo Rodríguez-Navas, Julián Proenza |
ETFA | 2 |
| 2013 | Design and Verification of a Media Redundancy Management Driver for a CAN Star TopologyabstractSome of the severe dependability limitations of Controller Area Network (CAN) can be overcome by replacing its bus topology with a star topology. Thus, a replicated star topology with advanced error-containment and fault-tolerance mechanisms for CAN, called ReCANcentrate, has been proposed. Its two hubs are coupled with each other and create a single logical broadcast domain. This allows each node to easily manage the replicated star by means of a software driver, called reCANdrv, that abstracts away the details of this replication. The goal of reCANdrv is to manage the star's media redundancy transparently for a CAN application, allowing it to exchange information through the star while tolerating faults. This paper describes the design of reCANdrv, the specification as properties of reCANdrv's correct redundancy management, and the verification of these properties by means of model checking. David Gessner, Manuel Barranco, Julián Proenza |
IEEE Trans. Ind. Informatics | 3 |
| 2013 | Using Timed Automata for Modeling Distributed Systems with Clocks: Challenges and SolutionsabstractThe application of model checking for the formal verification of distributed embedded systems requires the adoption of techniques for realistically modeling the temporal behavior of such systems. This paper discusses how to model with timed automata the different types of relationships that may be found among the computer clocks of a distributed system, namely, ideal clocks, drifting clocks, and synchronized clocks. For each kind of relationship, a suitable modeling pattern is thoroughly described and formally verified. Guillermo Rodríguez-Navas, Julián Proenza |
IEEE Trans. Software Eng. | 2 |
| 2012 | Probabilistic scheduling guarantees in distributed real-time systems under error burstsabstractNetworked embedded systems used in many real-time (RT) applications rely on dependable communication. Controller Area Network (CAN) has gained wider acceptance as a standard in a large number of applications, mostly due to its cost effectiveness, predictable performance, and its fault-tolerance capability. Research so far has focused on rather simplistic error models which assume only singleton errors separated by a minimum inter-arrival time. However, these systems are often subject to faults that manifest as error bursts of various lengths which have an adverse effect on the message response times that needs to be accounted for. Furthermore, an important factor to be considered in this context is the random nature of occurrences of faults and errors, which, if addressed in the traditional schedulability analysis by assuming a rigid worst case occurrence scenario, may lead to inaccurate results. In this paper we first present a stochastic fault and error model which has the capability of modeling error bursts in lieu of the commonly used simplistic error assumptions. We then present a methodology which enables the provision of appropriate probabilistic RT guarantees in distributed RT systems for the particular case of message scheduling on CAN under the assumed error assumptions. Hüseyin Aysan, Radu Dobrin, Sasikumar Punnekkat, Julián Proenza |
ETFA | 4 |
| 2012 | Developing TOBE-CAN: Total order broadcast enforcement in CANabstractOne of the drawbacks of the Controller Area Network (CAN) that must be overcome to make it suitable for critical applications is its incapacity for providing a Total Order Broadcast (TOB) communication service. A number of mechanisms were proposed to solve this problem, but each one of them only addresses a specific TOB limitation. Thus, this paper introduces TOBE-CAN, the first solution that comprehensively overcomes all the TOB flaws these previous mechanisms deal with. TOBE-CAN takes advantage of some of these mechanisms and further provides TOB while tolerating faults that are beyond their capabilities. Manuel Barranco, Julián Proenza |
ETFA | 2 |
| 2012 | A first qualitative evaluation of star replication schemes for FTT-CANabstractHighly dependable distributed embedded systems (DES) have traditionally been developed using static approaches, i.e., assuming a mostly constant environment. However, the little flexibility of such approaches does not allow continuous operation under dynamic environments. The Flexible Time-Triggered (FTT) communication paradigm is a promising approach to introduce the required flexibility. However, for continuous operation reliability is also crucial. Replicated star topologies are particularly well-suited to provide an increased reliability. Nevertheless, for FTT-CAN, the implementation of FTT for CAN, no replicated star topology that takes advantage of FTT-CAN's features to increase reliability and error containment exists. This paper discusses important design questions that need to be solved to create such a novel solution. David Gessner, Manuel Barranco, Julián Proenza, Michael Short 0001 |
ETFA | 3 |
| 2012 | Using FTT and stars to simplify node replication in CAN-based systemsabstractNodes, among the components of distributed embedded systems, exhibit the greatest permanent failure rate. Thus, providing tolerance to nodes faults is mandatory whenever high-reliability is required, being node replication the most common technique for that purpose. This paper proposes a novel technique suitable for CAN-based systems that simplifies existing approaches taking advantage of a star topology and the FTT protocol. Julián Proenza, Manuel Barranco, Joan Llodra, Luís Almeida 0001 |
ETFA | 1 |
| 2012 | The design of the CANbids architectureabstractDespite the significant advantages of the Controller Area Network (CAN) there is an extended belief that CAN is not suitable for critical applications, mainly because of several dependability limitations. During the CANbids project each one of these limitations has been addressed and a complete architecture for CAN-based fault-tolerant systems has been devised. This architecture allows building highly-reliable systems. This paper describes the design of such an architecture and the prototyping of its fundamental parts. Julián Proenza, Manuel Barranco, Guillermo Rodríguez-Navas, David Gessner, Fernando Guardiola, Luís Almeida 0001 |
ETFA | 1 |
| 2011 | Towards understanding the sensitivity of the reliability achievable by simplex and replicated star topologies in CANabstractStar-based field buses are gaining importance in the context of highly-dependable systems. However, although the error-containment and fault-tolerance capabilities of different stars have been evaluated, no one had appropriately quantified the system dependability benefits stars actually yield. Thus, in previous work, we quantitatively demonstrated, for the case of CAN, that a simplex and a replicated star called CANcentrate and ReCANcentrate can improve the system reliability when compared with a bus. However, we characterized all the dependability-related aspects of the system and the network to favor whenever possible the bus; except in one case, in which we studied the benefits of the simplex star over the bus depending on the error-containment capabilities of the nodes. Thus, to completely understand the full potential of stars, it is still necessary to assess how variations in each one of those aspects affect the reliability achievable with them when compared with the bus. This paper presents two of the set of analyses we are carrying out in this direction. Manuel Barranco, Julián Proenza |
ETFA | 2 |
| 2011 | Towards the integration of flexible-time-triggered communication and replicated star topologies in CANabstractThere is a growing interest in making the CAN field-bus more suitable for dependable applications. In the past years, several dependability limitations of CAN have already been addressed and a significant number of solutions are available. Nevertheless, the integration of these solutions into a single communication infrastructure is still an open issue. In this paper we discuss the integration of two specific solutions: FTT-CAN and ReCANcentrate. FTT-CAN is a higher-layer protocol that guarantees flexible real-time scheduling of CAN messages; whereas ReCANcentrate is a duplicated star topology for CAN that includes several enhanced mechanisms for media fault tolerance. We show how they are integrated into a single architecture that preserves the properties of each solution. Manuel Barranco, Guillermo Rodríguez-Navas, David Gessner, Julián Proenza |
ETFA | 4 |
| 2011 | Designing sfiCAN: A star-based physical fault injector for CANabstractThis paper presents the design and a preliminary implementation of sfiCAN: a physical fault injector for the CAN field-bus that allows the creation of a great variety of complex fault scenarios. The fault injector replaces the CAN bus topology with a star topology, whose central element is a hub with fault injection mechanisms. The fault injector is easily configured, with great flexibility, from a PC connected to a dedicated port of the hub. For this it uses a fault-injection specification, which is translated to a configuration protocol on top of CAN. This protocol is only used in-between fault injection tests and therefore does not interfere with the execution of any test. The purpose of the fault injector is to test the behavior of the nodes of a CAN network in the presence of channel errors, in particular, of the nodes' CAN controllers and the software executing on them, for which the star topology is transparent. David Gessner, Manuel Barranco, Alberto Ballesteros, Julián Proenza |
ETFA | 4 |
| 2011 | Injection of aggregated error flags as a means to guarantee consistent error detection in CANabstractAlthough the specification of CAN states that this protocol provides data consistency, it is well know that said property does not hold for certain specific error scenarios affecting the last bits of a CAN frame, and a number of solutions have been already suggested. Morever, for a long time it has been thought that the errors affecting the initial or intermediate bits of a CAN frame cannot cause any inconsistency. In this paper we show that this assumption is false, and that such kinds of message inconsistencies are also possible for certain combinations of multiple channel errors. After describing these unreported scenarios of inconsistency, we present a mechanism that guarantees the consistent detection of said scenarios and elimininates the possibility of suffering this kind of inconsistencies. This mechanism is therefore useful for the design of highly-dependable applications over CAN. Guillermo Rodríguez-Navas, Christian Winter 0004, Julián Proenza |
ETFA | 3 |
| 2010 | First prototype and experimental assessment of media management in ReCANcentrateabstractAlthough the use of star topologies to improve dependability in field-buses is gaining in importance, as in TTP/C and FlexRay, a mature technology such as the Controller Area Network (CAN) remained essentially a bus-only network. Thus, we proposed a CAN-compliant replicated star topology called ReCANcentrate, which has advanced error-containment and fault-tolerance mechanisms. Its two hubs are coupled with each other and create a single logical broadcast domain that allowed us to propose, in a previous work, a strategy for each node to easily manage the replicated star by means of a software driver that abstracts away the details of the replication. This paper describes the main functionalities of this driver, as well as the first tests we have conducted, on a real ReCANcentrate prototype, to verify the correctness and the performance of the driver in the absence and in the presence of faults. Manuel Barranco, David Gessner, Julián Proenza, Luís Almeida 0001 |
ETFA | 3 |
| 2009 | Demonstrating the Feasibility of Media Management in ReCANcentrateabstractStar topologies are rising the interest of newer field-bus communication technologies like TTP/C and FlexRay, given the dependability advantages stars can offer. However, it is also possible to take advantage of a mature technology such as Controller Area Network (CAN), while benefiting from stars. For that, we developed a CAN-compliant replicated star called ReCANcentrate. It includes two hubs that are coupled with each other, thereby forcing a single broadcast domain that allowed us to define, in a previous work, a strategy for each node to easily manage the replicated star. To demonstrate the feasibility of this management, this paper presents its on-going implementation as a driver to be executed at each node. Manuel Barranco, David Gessner, Julián Proenza, Luís Almeida 0001 |
ETFA | 3 |
| 2009 | First Quantitative Results of the Dependability Improvement Achieved by ReCANcentrateabstractThere is a growing interest in using star topologies instead of buses as the communication infrastructure for highly-reliable distributed control systems, given the better dependability stars are supposed to provide. For the controller area network (CAN), we developed a simplex and a replicated star called CANcentrate and ReCANcentrate respectively. In a previous work we modelled the dependability of the CAN bus and CANcentrate using stochastic activity networks (SANs). There we presented the first quantitative analysis of the error-containment benefits of a simplex star when considering permanent hardware faults. This paper quantitatively analyzes, for the first time, how a replicated star such as ReCANcentrate can improve both error-containment and reliability, also considering permanent hardware faults. We explain our modelling strategy using SANs and show some first and novel results. Manuel Barranco, Julián Proenza, Luís Almeida 0001 |
ETFA | 2 |
| 2009 | A First Design for CANsistant: A Mechanism to Prevent Inconsistent Omissions in CAN in the Presence of Multiple ErrorsabstractDespite the significant advantages of the controller area network (CAN) there is an extended belief that CAN is not suitable for critical applications, mainly because of several dependability limitations. One of them is its limited data consistency. Several solutions to this problem have been previously proposed but they are not able to efficiently ensure consistent broadcasts in the presence of multiple channel errors. This paper introduces a circuit called CANsistant, that detects all scenarios potentially leading to the inconsistent omission of a frame in the presence of up to 4 channel errors and, if necessary, retransmits the affected frame. Julián Proenza, Ernesto Sigg |
ETFA | 1 |
| 2008 | Designing and verifying media management in ReCANcentrateabstractTo overcome some dependability limitations of CAN that arise from its non-redundant bus topology, we have proposed a CAN-compliant replicated star topology, ReCANcentrate, whose hubs incorporate the necessary fault-treatment and fault-tolerance mechanisms. This paper presents ongoing work regarding the design and formal verification of the strategy each node of ReCANcentrate uses to manage the transmissions and the receptions on the replicated star, as well as to tolerate faults. Manuel Barranco, Julián Proenza, Luís Almeida 0001 |
ETFA | 2 |
| 2008 | Analytical Assessment of the Precision Degradation Caused by Faults in a Fault-Tolerant Master/Slave Clock Synchronization Service for CANabstractThe main goal of a clock synchronization service is to keep a consistent perception of time among the nodes of the system. In this context, consistency means that at any instant, the values of all the clocks in the system do not differ more than a given amount, which is called the precision. Moreover, clock synchronization is said to be fault tolerant if the intended precision is guaranteed despite the occurrence of the faults included in the fault model. In this paper, we consider a specific fault-tolerant master/slave clock synchronization service for the Controller Area Network (CAN) field bus, and analyze its precision under different fault assumptions. The equations obtained in our analysis show that, when using master redundancy, inconsistent channel faults may have a negative impact on the guaranteed precision. Guillermo Rodríguez-Navas, Julián Proenza |
SRDS | 2 |
| 2008 | Orthogonal, Fault-Tolerant, and High-Precision Clock Synchronization for the Controller Area NetworkabstractThe controller area network (CAN) is facing a great opportunity. The maturity of this technology makes many researchers believe that CAN may be adopted in more critical systems. However, the suitability of CAN for these challenging applications strongly depends on our capacity to integrate all the solutions already available into a single, comprehensive architecture. We claim that clock synchronization plays a fundamental role in such architecture. Therefore, the means to achieve a solution fulfilling the expected requirements on reliability, cost, and precision must be deeply investigated. This paper discusses the relevance of clock synchronization in the future of CAN systems and describes a novel solution to supply this service. This solution exhibits several advantages: it provides very high precision, causes very low communication and computation overhead, and includes mechanisms to provide fault tolerance. Moreover, and in contrast to previous proposals, it is designed to be orthogonal to the rest of the system. Thus, it can be directly incorporated to any CAN system, without having to replace any of the components, which reduces the cost increment caused by the new service. Guillermo Rodríguez-Navas, Sebastià Roca, Julián Proenza |
IEEE Trans. Ind. Informatics | 3 |
| 2007 | Modelling MajorCAN with UPPAALabstractThe controller area network (CAN) protocol produces data inconsistencies in some scenarios. A previous work proposed a new protocol called MajorCAN which is a small modification to CAN. MajorCAN does not present the reported error scenarios thus ensuring data consistency. Although MajorCAN has been thoroughly simulated, no formal verification has been performed so far. In this paper we describe how we have modelled MajorCAN using a network of timed automata in Uppaal. This is the first step of its formal verification by means of model checking. Matias Bonet, Gabriel Donaire, Julián Proenza |
ETFA | 3 |
| 2007 | Modeling and Verification of Master/Slave Clock Synchronization Using Hybrid Automata and Model-Checking
Guillermo Rodríguez-Navas, Julián Proenza, Hans A. Hansson |
ICFEM | 2 |
| 2006 | An active star topology for improving fault confinement in CAN networksabstractThe controller area network (CAN) is a field bus that is nowadays widespread in distributed embedded systems due to its electrical robustness, low price, and deterministic access delay. However, its use in safety-critical applications has been controversial due to dependability limitations, such as those arising from its bus topology. In particular, in a CAN bus, there are multiple components such that if any of them is faulty, a general failure of the communication system may happen. In this paper, we propose a design for an active star topology called CANcentrate. Our design solves the limitations indicated above by means of an active hub, which prevents error propagation from any of its ports to the others. Due to the specific characteristics of this hub, CANcentrate is fully compatible with existing CAN controllers. This paper compares bus and star topologies, analyzes related work, describes the CANcentrate basics, paying special attention to the mechanisms used for detecting faulty ports, and finally describes the implementation and test of a CANcentrate prototype. Manuel Barranco, Julián Proenza, Guillermo Rodríguez-Navas, Luís Almeida 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2006 | Combining operational flexibility and dependability in FTT-CANabstractThe traditional approaches to the design of distributed safety-critical systems, due to fault-tolerance reasons, have mostly considered static cyclic table-based traffic scheduling. However, there is a growing demand for operational flexibility and integration, mainly to improve efficiency in the use of system resources, with the network playing a central role to support such properties. This calls for dynamic online traffic scheduling techniques so that dynamic communication requirements are adequately supported. Nevertheless, using dynamic traffic management mechanisms raises additional problems, in terms of fault-tolerance, related with the weaker knowledge of the future system state caused by the higher level of operational flexibility. Such problems have been recently addressed in the scope of using flexible time-triggered CAN (FTT-CAN) in safety-critical applications in order to benefit from the high operational flexibility of this protocol. This paper gathers and reviews the main mechanisms that were developed to provide dependability to the protocol, namely, master replication and fail-silence enforcement. Joaquim Ferreira 0001, Luís Almeida 0001, José Alberto Fonseca, Paulo Pedreiras, Ernesto Martins, Guillermo Rodríguez-Navas, Joan Rigo, Julián Proenza |
IEEE Trans. Ind. Informatics | 8 |
| 2005 | Towards analyzing the fault-tolerant operation of server-CANabstractThis work-in-progress (WIP) paper presents server-CAN and highlights its operation and possible vulnerabilities from a fault tolerance point of view. The paper extends earlier work on server-CAN by investigating the behaviour of server-CAN in faulty conditions. Different types of faults are described, and their impact on sever-CAN is discussed, which is the subject of on-going research Thomas Nolte, Guillermo Rodríguez-Navas, Julián Proenza, Sasikumar Punnekkat, Hans A. Hansson |
ETFA | 3 |
| 2003 | COTS-based hardware support to timeliness in CAN networksabstractAdvances in programmable hardware have simplified integration of communication facilities in low-cost hardware components. This has proved to be beneficial in the design of distributed embedded systems as it allows the communication subsystem to provide important properties at a low level. The present work follows this approach in order to achieve timeliness in CAN networks. This paper describes the implementation of the LST-CAN protocol in programmable hardware. This protocol is an extension to CAN which ensures timely communication regardless of environmental interferences. Guillermo Rodríguez-Navas, Manuel Barranco, Julián Proenza, Ian Broster |
ETFA (1) | 3 |
| 2003 | An architecture for physical injection of complex fault scenarios in CAN networksabstractIt has been reported that some particular fault scenarios may cause malfunction of the controller area network protocol. Although such scenarios are very unlikely, they become relevant when attempting to use the CAN protocol for critical applications. The fault injector described in this paper induces these fault scenarios at the physical layer of the CAN protocol by means of a software tool and a set of specifically designed circuits. Therefore, and in contrast to previous solutions, this fault injector is suitable to evaluate most of the dependability mechanisms that have been proposed for CAN networks. Guillermo Rodríguez-Navas, Jesús Jiménez, Julián Proenza |
ETFA (2) | 3 |