Rolf Ernst

dblp:e/RolfErnst · DBLP profile ↗
← Back
228ranked-venue papers
13as first author
31since 2021 · last 2026
0000-0003-2414-9566ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 163 · 11 first-author · 18 since 2021Software engineering, systems software and programming languages · 70 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 2 first-author · 10 since 2021Security and privacy · 6Artificial intelligence and machine learning · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Low Latency Communication of Large Data Objects by Subscriber-centric Selective Data Transfer
abstract
Autonomous cyber-physical systems depend on high-resolution sensors that generate multi-gigabyte-per-second data streams for accurate environmental perception and system safety. However, existing publish-subscribe middleware frameworks, such as ROS 2 with DDS, are optimized for small data objects and face significant latency and overhead challenges when handling large sensor data between distributed application nodes. Recognizing the need for more efficient data management, we propose a new companion middleware that prioritizes application-specific data relevance, enabling selective communication of critical information while reducing the burden on communication resources and maintaining interoperability with state-of-the-art publish-subscribe middleware. This companion middleware enables timely and effective sharing of sensor data by focusing on regions of interest relevant to specific tasks, such as traffic light detection in driving scenarios. Experimental evaluations of our open source implementation of this companion middleware on a Linux platform demonstrate that our protocol integrates efficiently with ROS 2, significantly enhancing data management and communication efficiency.
Nora Sperling, Rolf Ernst
ACM Trans. Cyber Phys. Syst.2
2025 A Novel Timing Model for Neural Networks in Safety-critical Systems
abstract
Deep Neural Networks (DNNs) have a largely fixed program path, so average and worst-case timing depend mostly on data values and physical platform properties. Related work on recent high-performance platforms already observed a dominant role of circuit parameters and operating conditions leading to a normal distribution of DNN timing variations. In extensive experiments with different DNNs on very different types of such platforms used for real-time applications, we could confirm that a Gaussian distribution closely approximates between 99% and 99.9% of all DNN executions. Remaining rare outliers are often neglected in literature, usually assuming that they will not affect applications. However, in experiments with long test series we could show that outliers are frequent enough to be relevant for critical system design and even form dense outlier sequences that can seriously challenge latency critical applications, such as vehicle perception. We propose and evaluate a dual timing model capturing both typical response time and outlier distributions and provide design methods to control and mitigate their effects in design and operation. While developed for DNNs, the approach could be relevant for other critical real-time applications on high-performance platforms.
Robin Hapka, Anika Christmann, Rolf Ernst
COMPSAC3
2025 Teleoperation as a Step Towards Fully Autonomous Systems
abstract
In the foreseeable future, highly automated mobile systems, such as vehicles, robots, UAVs, or trains, will be confronted with difficult situations that require external support. The availability of such external support corresponds to level 4 driving automation and is an essential feature in current robotaxis and automated public transportation. While the first generation of level 4 prototypes relied on safety driver support, commercial systems are gradually moving towards support by teleoperation. Designing teleoperation support for level 4 systems is an end-to-end problem involving two main research and practical challenges, the teleoperation function defining the remote human interface with its scene representation and available control functions, and the real-time communication channel involving wired and wireless segments, which must provide reliable end-to-end data transport. Both challenges are tightly linked, and combined solutions are needed to reach the required safe teleoperation. Solutions can make use of the rich sensing and control system of a level 4 vehicle, which however can only be exploited if the communication channel provides adequate real-time access.
Alex Bendrick, Daniel Tappe, Nora Sperling, Rolf Ernst, Andrea Nota, Selma Saidi, Frank Diermeyer
DATE4
2025 Enabling Continuous Low Latency Streaming in Industrial Roaming Scenarios
abstract
Cooperation among vehicles and mobile robotic systems in industrial environments relies on wireless communication. Future roadmaps envision the integration of high-data-rate sensor streams to enhance the performance of such cooperative systems. However, the state-of-the-art in both 802.11 and cellular communication suffers from prolonged and non-deterministic connection losses, lasting up to several seconds, during roaming. As a result, the reliable transmission of large and latency-critical data in such scenarios becomes impractical. For small control data, this issue has been mitigated using redundant data streams. However, this approach is not viable for perception data streams due to excessive resource demands. This work presents an implementation of mechanisms designed to address this challenge, with a focus on seamless deployment in industrial environments using 802.11 and Ethernet TSN backbones. Our physical model truck demonstrator confirmed the feasibility of reliable sample exchange during roaming for data rates up to 25 Mbit/s without active redundancy.
Daniel Tappe, Alex Bendrick, Rolf Ernst
IECON3
2025 Continuous Streaming in Roaming Scenarios: A Model Truck-Based Demonstration
abstract
Industrial automation demands continuous, lowlatency streaming for mobile robots, yet mobility often causes handover delays and connectivity issues. We present a physical testbed that validates a streaming solution using multiconnectivity, fast loss detection, and application-level backward error correction with standard 802.11 hardware and Linux-based systems. Our approach achieves seamless handovers in under 10 ms without redundant transmissions, confirming its feasibility for industrial applications.
Daniel Tappe, Alex Bendrick, Rolf Ernst
WFCS3
2025 Towards Efficient Multi-Frame Clustering in Response Time Analysis for Large Object Communication
abstract
In autonomous systems, growing sizes of application data, primarily related to perception tasks, have to be transmitted over communication infrastructures that provide higher data rates. Knowledge and exploitation of the clustered structure of multi-frame application data have proven to reduce the interference between different real-time critical communication streams, as recently shown for synchronous systems. However, the accompanying increase in frame numbers will foreseeably make the corresponding analysis impractical due to frame-dependent processing times. As a solution, based on the synchronous example mentioned, we demonstrate how to compose multi-frame transmissions to larger frame clusters in the analysis and decouple the analysis complexity from the application data sizes. Based on an automotive and industrial TSN use case, our results show comparable analytical response times and very low computation times at arbitrary data rates and sizes. It thus enables the deployment of efficient configuration methods that arbitrate large data samples transmitted through the network as one whole frame cluster. In addition, we made the developed analysis openly accessible.
Jonas Peeck, Rolf Ernst, Selma Saidi
ACM Trans. Embed. Comput. Syst.2
2024 Conservative Design with High-Performance COTS Architectures - Beyond Traditional Approaches
abstract
Safety-critical industrial systems are subject to many stringent requirements, including non-functional design constraints that are essential for robustness, reliability, and safety. System timing is often treated as an afterthought, even though data age, race conditions, or missed deadlines can affect functional safety. This challenge has been addressed in real-time system design, but the situation has worsened in recent system designs. Many applications require high performance to achieve behavioral safety goals, i.e., safety of intended functionality, using complex heterogeneous architectures with multi-level memory, interconnected chiplets, and accelerators based on the latest technologies. Machine learning applications in factory automation, robotics, and autonomous systems are important examples. The timing of such architectures is heavily influenced by the physical effects of temperature control and chip variations, including aging, leading to even more complex chip-specific software execution timing. The paper shows that identical chips differ in their timing behavior, and that most of the timing variation of a single chip can be effectively modeled using additive white Gaussian noise. We explain how such timing behavior can be exploited in high-assurance design for safety and reliability, following established reliability analysis, test, and redundancy techniques.
Robin Hapka, Rolf Ernst
ETFA2
2024 Fast Vehicular TSN Network Reconfiguration with Application Aware Network Synchronization
abstract
In-Vehicle Networks (IVN) are under intense pressure to be both resource efficient at low cost and flexible enough for the future. The deployment of a large number of new sensors for autonomous driving and self-awareness significantly increases the demand for network resources. In addition, statically designed networks lack flexibility or have large over-provisioning of resources due to the dynamics of network traffic. Application-aware resource management addresses both of these issues. It allows the network to be reconfigured while controlling application requirements, increasing resource utilization, and providing flexibility for future demands. Our timestamp-based reconfiguration protocol ensures fast, lossless reconfiguration of high bandwidth data connections. Based on global time synchronization, it enables precise reconfiguration without increasing the risk of data loss and reduces the risk of inconsistent network configurations by exploiting the object slack created by large object transfers consisting of hundreds of individual packets. Using software-based Linux switches as the network, we demonstrate the use of the reconfiguration protocol under real-world conditions.
Dominik Stöhrmann, Rolf Ernst
ETFA2
2024 Ultra Reliable Hard Real-Time V2X Streaming with Shared Slack Budgeting
abstract
While autonomous driving roadmaps envision large data to be exchanged in V2X scenarios, current V2X communication standards only focus on reliable exchange of small objects. The Wireless Reliable Real-Time Protocol (W2RP) addresses this challenge, however, cannot properly handle scenarios where multiple applications share the channel especially as there are predetermined periodic and dynamic retransmissions phases for each data object. If transient error spikes occur, applications will compete for dynamic channel resources, potentially creating critical overload situations that can lead to safety violations. Furthermore, the channel is shared by multiple applications with different criticality, properties and constraints. We propose a two-level hierarchical approach that assigns a fixed budget to each node, with periodic and dynamic transmission parts of a node’s applications being scheduled by the node itself. Prioritization is used to differentiate applications based on their criticality, thereby enabling reliable data exchange for safety-critical applications. The concept was evaluated using an OMNeT++-based simulation of an Automated Valet Parking use case and proved highly effective in enabling safe sample exchange even under varying channel conditions, while offering significantly more efficient resource reservation compared to a static configuration.
Alex Bendrick, Daniel Tappe, Rolf Ernst
IV3
2024 Continuous multi-access communication for high-resolution low-latency V2X sensor streaming
abstract
Future automated mobility is expected to rely on Cooperative Perception (CP) applications to achieve high degrees of reliable autonomy. CP applications are characterized by the exchange of large sensor data objects, such as camera frames or LIDAR point clouds. The high mobility of nodes in combination with the safety-critical nature of CP data, demands architectures enabling continuous connectivity for reliable transmission of large data objects. This work presents a user-centric architecture combining a cell-free Radio-Access-Network (RAN) with a centralized backbone management. By monitoring multiple available connections through an ultra-lightweight heartbeat-based protocol, the handover procedure can be reduced to a low-latency backbone reconfiguration. Thus, the need for redundant data transmission to achieve loss-free and continuous data transmissions during handover is avoided.
Daniel Tappe, Alex Bendrick, Rolf Ernst
IV3
2024 Reducing Communication Cost and Latency in Autonomous Vehicles with Subscriber-centric Selective Data Distribution
abstract
Driving automation has become a major cost factor in automotive design. High computation demand for machine learning (ML) applications and growing sensor resolution and data rate require expensive hardware technology and networking. Newer network technologies and topologies can at best compensate the growing communication demand, but the network still accounts for a complex wiring harness with many dedicated sensor cables. In this paper, we exploit the context specific sensor data access of ML perception applications to minimize the sensor data traffic. For that purpose, we extend the popular Data Distribution Service (DDS) publish-subscribe middleware by a subscriber-centric software caching feature, which is then supported by an appropriate network scheduling. Using realistic data sets and ML applications from the popular Autoware benchmark and a zonal architecture according to the P802.1DG automotive Time-Sensitive Networking (TSN) network profile, we demonstrate that both cabling structure and network latency can be significantly reduced for both 1 Gbps and 10 Gbps TSN technologies. This result enables faster perception and/or lower cable cost, at no loss in data quality or reliability.
Nora Sperling, Rolf Ernst
VTC Spring2
2024 Large Data Transfer Optimization for Improved Robustness in Real-Time V2X-Communication
abstract
Vehicle-to-everything (V2X) roadmaps envision future applications that require the reliable exchange of large sensor data over a wireless network in real time. Applications include sensor fusion for cooperative perception or remote vehicle control that are subject to stringent real-time and safety constraints. Real-time requirements result from end-to-end latency constraints, while reliability refers to the quest for loss-free sensor data transfer to reach maximum application quality. In wireless networks, both requirements are in conflict, because of the need for error correction. Notably, the established video coding standards are not suitable for this task, as demonstrated in experiments. This article shows that middleware-based backward error correction (BEC) in combination with application controlled selective data transmission is far more effective for this purpose. The mechanisms proposed in this article use application and context knowledge to dynamically adapt the data object volume at high error rates at sustained application resilience. We evaluate popular camera datasets and perception pipelines from the automotive domain and apply two complementary strategies. The results and comparisons show that this approach has great benefits, far beyond the state of the art. It also shows that there is no single strategy that outperforms the other in all use cases.
Alex Bendrick, Nora Sperling, Rolf Ernst
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 An Error Protection Protocol for the Multicast Transmission of Data Samples in V2X Applications
abstract
There is a trend towards communication of larger data objects in wireless vehicle communication. In many cases, communication uses publish-subscribe protocols. Data rate requirements of such protocols are best addressed by wireless multicast protocols, but the existing protocols lack an error protection that is suitable for real-time and safety-critical applications. We present an application-aware protocol that supports the popular DDS (Data Distribution Service) middleware. By exploiting data object deadlines and slack for retransmissions and employing an adaptable, multicast-aware prioritization mechanism, the reliable exchange of large data objects is enabled. The protocol is sufficiently general to be used on top of different communication standards such as 802.11- and cellular-based V2X (Vehicle-to-Everything) technologies. The protocol was implemented in an OMNeT++ simulation model and evaluated against recent state-of-the-art alternatives using parameters and constraints taken from a motivational truck platooning example. Furthermore, the protocol was implemented using an open-source DDS implementation as the basis and tested on a physical wireless demonstrator setup. The evaluation shows that the presented multicast protocol substantially outperforms the alternatives keeping streaming applications operational even under high frame error rates.
Alex Bendrick, Jonas Peeck, Rolf Ernst
ACM Trans. Cyber Phys. Syst.3
2023 Efficient hard real-time implementation of CNNs on multi-core architectures
abstract
Autonomous driving applications rely on processing large amounts of data in order to ensure sufficient perception performance. In this context, Convolutional Neuronal Networks (CNNs), which are used for object detection in camera images, are an integral part of sensor data processing. However, safety-related hard deadlines are counteracted by the required high data rates that stress the platform w.r.t. memory interference. Providing high camera frame rates under hard latency requirements is still an open issue. In a case study focusing on high-performance multi-core architectures, we evaluate in detail the advantages of a private L2 and shared L3 cache architecture for CNN processing and show how to remove malicious data synchronization effects. In this context we deploy MobileNet as well as YOLO on an Intel i5 processor and demonstrate their applicability to a worst-case design under high CNN frame rates. Last, we provide a heuristic optimization scheme that is able to efficiently find feasible high frame rate configurations. Contrary to common assumptions, our results show that high-performance multi-core COTS platforms are suitable for the application of CNNs even under hard deadline constraints and, hence, offer substantial gains in cost and productivity due to their greater ease of programming.
Jonas Peeck, Robin Hapka, Rolf Ernst
COMPSAC3
2023 Invited: Caching in Automated Data Centric Vehicles for Edge Computing Scenarios
abstract
With the current trend towards large data volumes, state-of-the-art communication and data management solutions based on the publish/subscribe pattern reach the network limits, hampering the deployment of advanced edge and cloud computing services. A main reason is the established publisher-centric distribution mechanism that over-utilizes network resources. We suggest a network and subscriber-centric adaptive data caching communication scheme supported by dynamic network management and configuration, which has great potential to solve the data distribution challenges. We show that state-of-the-art publish/subscribe middlewares like Data Distribution Service (DDS) can transparently be combined with such software caching.
Nora Sperling, Alex Bendrick, Dominik Stöhrmann, Rolf Ernst
DAC4
2023 Formal Analysis of Timing Diversity for Autonomous Systems
abstract
The design of autonomous systems, such as for automated driving and avionics, is challenging due to high performance requirements combined with high criticality. Complex applications demand the full performance of commercial high performance multi-core systems of-the-shelf (COTS), with or without accelerators. While these systems are optimized for performance, hard real-time requirements and deterministic timing behavior are major constraints for safety-critical systems. Unfortunately, infrequent timing outliers caused by interleaved hardware-software effects of COTS systems complicate traditional worst-case design. This conflict often prohibits deploying COTS hardware and consequently prevents sophisticated applications, too. Recently, an approach called Timing Diversity was introduced, which proposes to exploit existing dual modular redundant hardware platforms to mask deadline violations. This paper puts Timing Diversity on a theoretical foundation and provides specification for different implementations. It demonstrates that Timing Diversity needs fast recovery to be effective, proposes a recovery strategy and provides a mathematical model for the reliability of the resulting system. Using experimental data in a Linux based system, it shows that fast recovery is useful, making Timing Diversity a realistic option for compute demanding hard real-time applications.
Anika Christmann, Robin Hapka, Rolf Ernst
DATE3
2023 Information Processing Factory 2.0 - Self-awareness for Autonomous Collaborative Systems
abstract
This paper summarizes the talks of a special session on the IPF 2.0 project, a collaborative German-US research project that leverages self-awareness principles for the self-management of distributed systems of autonomous multiprocessor systems-on-chip (MPSoCs).
Nora Sperling, Alex Bendrick, Dominik Stöhrmann, Rolf Ernst, Bryan Donyanavard, Florian Maurer 0003, Oliver Lenke, Anmol Surhonne, Andreas Herkersdorf, Walaa Amer, Caio Batista de Melo, Ping-Xiang Chen, Quang Anh Hoang, Rachid Karami, Biswadip Maity, Paul Nikolian, Mariam Rakka, Dongjoo Seo, Saehanseul Yi, Minjun Seo, Nikil Dutt, Fadi J. Kurdahi
DATE4
2023 Hard Real-Time Streaming of Large Data Objects with Overlapping Backward Error Correction
abstract
Despite expectations in roadmaps, the reliable V2X exchange of large data objects under age constraints remains an open challenge. Earlier work proposed a protocol that improved reliability by an application-level Backward Error Correction (BEC) mechanism that operates on data objects rather than packets. In this paper, we extend the protocol to overlapping data object transmission and investigate how far we can increase reliability by relaxing the object transmission deadline to realistic latency requirements as found in the literature. The extended protocol can be used for reliable coupling of nodes and wired network segments in the automotive and industrial domain that exchange large data. As an example, a camera stream transmission from truck platooning is used that aims at improving effectiveness of logistics as part of a large-scale industrial process. We present an analytical model for the extended error correction protocol and evaluate the reliability improvement with an OMNeT++ simulation model that incorporates a state-of-the-art burst error model. For realistic application deadlines, the results show a significant reliability improvement especially for burst errors, at constant protocol overhead.
Alex Bendrick, Rolf Ernst
IECON2
2023 Application-centric Network Management - Addressing Safety and Real-time in V2X Applications
abstract
The current roadmaps and surveys for future wireless networking typically focus on communication and networking technologies and use representative applications to derive future network requirements. Such a benchmarking approach, however, does not cover the application integration challenge that arises from the many distributed applications sharing a network infrastructure, each with their individual topology and data structure. The paper addresses V2X networks as an important example. Crucial end-to-end application constraints including real-time and safety encourage a closer look at application interference and systematic integration. This perspective paper proposes a two-layer resource management that divides the problem into an application integration and a network management task. Valet parking with high-resolution infrastructure camera support is elaborated as a use case that overarches vehicle network and wireless network management. Experiments demonstrate the benefits of complementing the current network-centric management by an application-centric integration.
Rolf Ernst, Dominik Stöhrmann, Alex Bendrick, Adam Kostrzewa
ACM Trans. Embed. Comput. Syst.1
2023 Robust Cause-Effect Chains with Bounded Execution Time and System-Level Logical Execution Time
abstract
In automotive and industrial real-time software systems, the primary timing constraints relate to cause-effect chains. A cause-effect chain is a sequence of linked tasks and it typically implements the process of reading sensor data, computing algorithms, and driving actuators. The classic timing analysis computes the maximum end-to-end latency of a given cause-effect chain to verify that its end-to-end deadline can be satisfied in all cases. This information is useful but not sufficient in practice: Software is usually evolving and updates may always alter the maximum end-to-end latency. It would be desirable to judge the quality of a software design a priori by quantifying how robust the timing of a given cause-effect chain will be in the presence of software updates. In this article, we derive robustness margins which guarantee that if software extensions stay within certain bounds, then the end-to-end deadline of a cause-effect chain can still be satisfied. Robustness margins are also useful to know if the system model has uncertain parameters. A robust system design can tolerate bounded deviations from the nominal system model without violating timing constraints. The results are applicable to both the bounded execution time programming model and the (system-level) logical execution time programming model. In this article, we study both an industrial use case from the automotive industry and analyze synthetically generated experiments with our open-source tool TORO.
Leonie Köhler, Phil Hertha, Matthias Beckert, Alex Bendrick, Rolf Ernst
ACM Trans. Embed. Comput. Syst.5
2023 Improving Worst-case TSN Communication Times of Large Sensor Data Samples by Exploiting Synchronization
abstract
Higher levels of automated driving also require a more sophisticated environmental perception. Therefore, an increasing number of sensors transmit their data samples as frame bursts to other applications for further processing. As a vehicle has to react to its environment in time, such data is subject to safety-critical latency constraints. To keep up with the resulting data rates, there is an ongoing transition to a Time-Sensitive Networking (TSN)-based communication backbone. However, the use of TSN-related industry standards does not match the automotive requirements of large timely sensor data transmission, nor it offers benefits on time-critical transmissions of single control data packets. By using the full data rate of prioritized IEEE 802.1Q Ethernet, giving time guarantees on large data samples is possible, but with strongly degraded results due to data collision. Resolving such collisions with time-aware shaping comes with significant overhead. Hence, rather than optimizing the parameters of the existing protocol, we propose a system design that synchronizes the transmission times of sensor data samples. This limits network protocol complexity and hardware requirements by avoiding tight time synchronization and time-aware shaping. We demonstrate that individual sensor data samples are transmitted without significant interference, exclusively at full Ethernet data rate. We provide a synchronous event model together with a straightforward response time analysis for synchronous multi-frame sample transmissions. The results show that worst-case latencies of such sample communication, in contrast to non-synchronized approaches, are close to their theoretical minimum as well as to simulative results while keeping the overall network utilization high.
Jonas Peeck, Rolf Ernst
ACM Trans. Embed. Comput. Syst.2
2022 Industry-track: System-Level Logical Execution Time for Automotive Software Development
abstract
The way how automotive software is developed has rapidly evolved with the introduction of heterogeneous hardware/software architectures. Nevertheless, the requirement for deterministic behavior of safety-critical cause-effect chains persists unchanged. As a side effect of the shared platform, complex dependencies between critical and non-critical functions arise, demanding a model-based approach to handle time determinism throughout the design process. Limited to the scope of a single component, the Logical Execution Time (LET) paradigm provides such an abstraction of the runtime behavior. It has been successfully introduced in AUTOSAR to mitigate the design complexity, ensure a deterministic timing behavior and facilitate a lock-free communication. This paper discusses how the scope of LET can be extended to the system level, enabling an efficient design of distributed AUTOSAR software, where robustness towards platform changes plays a key role. System-Level Logical Execution Time (SL-LET) is currently in the process of AUTOSAR standardization, supported by a joint group of industry and academic partners.
Kai-Björn Gemlau, Hermann von Hasseln, Rolf Ernst
EMSOFT3
2022 Efficient Timing Isolation for Mixed-Criticality Communication Stacks in Performance Architectures
abstract
The high complexity of a network stack as found in Linux leads to unpredictable timing behavior and interference that is incompatible to safety standards. Inspired by the microkernel approach, a filter stack separates critical from non-critical traffic using a heterogeneous architecture. Such separation likely leads to performance inefficiencies or limitations in functionality. In this paper we improve a filter architecture by hardware supported temporal and spatial isolation. We show that this approach provides near optimum performance even when applied to a commercial Ethernet interface. Hardware support simplifies the stack architecture, guaranteeing high predictability for critical traffic.
Kai-Björn Gemlau, Nora Sperling, Rolf Ernst
ETFA3
2022 Controlling High-Performance Platform Uncertainties with Timing Diversity
abstract
Autonomous mobile systems combine high performance requirements with safety criticality. High performance hardware/software architectures, however, expose a far more complex runtime behavior than traditional microcontroller architectures. Such high-performance architectures challenge traditional worst-case design that assumes a formally analyzable or at least deterministic worst-case response time (WCRT) that can be reasonably bounded. However, such architectures expose rare but substantial worst-case outliers, which are not only caused by the application itself, but also by the many dynamic influences of software architecture and platform control. Probabilistic methods can capture such outliers, but are only effective, if the outlier probability is sufficiently low and if the methods cover dynamic platform timing. As a main contribution, this paper exploits platform induced timing variety rather than trying to mitigate it. Assuming the typical redundant dual modular redundancy (DMR) implementation that is deployed in safety-critical systems, it introduces the concept of Timing Diversity, where rare outliers in one of the two channels are masked by the other channel with a sufficiently high probability. The paper uses a convolutional neural network (CNN) example in different parameter settings running on Linux operated multi-core platform with typical dynamic control to investigate the proposed concept. The experiments demonstrate the potential of Timing Diversity in leading to substantially higher reliability. Alternatively, the approach permits a reduction of the system WCRT at the same reliability level.
Robin Hapka, Anika Christmann, Rolf Ernst
RTCSA3
2021 Efficient Run-Time Environments for System-Level LET Programming
abstract
Growing requirements of large industrial and automotive software systems have initiated an ongoing move from monolithic and tightly integrated run-time environments (RTE) to virtual platforms implemented on fewer domain computers with heterogeneous physical architectures. This trend has given rise to new programming paradigms to enable specification, implementation and supervision of software systems that are predictable and robust under interference and change. One of those paradigms, the Logical Execution Time (LET), is now part of the automotive software standard, AUTOSAR. While originally applied to single shared-memory multicore processors, System-level LET (SL LET) extends this approach to virtual and distributed platforms providing a powerful paradigm for CPS in future industrial systems. This contribution explains and demonstrates the resulting challenges to the RTE and the opportunities to improve its efficiency, in particular the communication stack.
Kai-Björn Gemlau, Leonie Köhler, Rolf Ernst
DATE3
2021 Online latency monitoring of time-sensitive event chains in safety-critical applications
abstract
Highly-automated driving involves chains of perception, decision, and control functions. These functions involve data-intensive algorithms that motivate the use of a data-centric middleware and a service-oriented architecture. As an example we use the open-source project Autoware.Auto. The function chains define a safety-critical automated control task with weakly-hard real-time constraints. However, providing the required assurance by formal analysis is challenged by the complex hardware/software structure of these systems and their dynamics. We propose an approach that combines measurement, suitable distribution of deadline segments, and application-level online monitoring that serves to supervise the execution of service-oriented software systems with multiple function chains and weakly-hard real-time constraints. We use DDS as middleware and apply it to an Autoware.Auto use case.
Jonas Peeck, Johannes Schlatow, Rolf Ernst
DATE3
2021 Timing diversity as a protective mechanism: work-in-progress
abstract
Dual modular redundancy (DMR) is not only an established solution for systems with high reliability demands, it is even required in aviation certification standards such as DO-254 [5, Clause 2.3.1]. A safety critical avionic application such as the flight control system is designed with up to 6-fold redundancy and the Avionics Full-Duplex Ethernet (AFDX) communication network is also based on the DMR. Even in the automotive domain, DMR is a well known solution. ISO26262 [3, Part 6, Clause 7.4.13] also suggests heterogeneous or diverse redundancy for safety-critical applications including software which must be redundantly executed on independent hardware components to avoid failure due to hardware errors. We exploit this mandatory software redundancy to master timing errors of critical software with minimum additional overhead.
Mischa Möstl, Robin Hapka, Anika Christmann, Rolf Ernst
EMSOFT4
2021 A Middleware Protocol for Time-Critical Wireless Communication of Large Data Samples
abstract
We present a middleware-based protocol that reliably synchronizes large samples consisting of multiple frames efficiently and within application level QoS requirements over a lossy wireless channel. The protocol uses a custom retransmission scheme, exploiting the latency requirements on sample level for frame level scheduling. It can be integrated into the popular DDS middleware. We investigate some technical limits of such a protocol and compare it to existing error protocols in the software stack and in the wireless protocol and combinations thereof. The comparison is based on an Omnet++ simulation using an established wireless channel error model. For evaluation, we take a use case from automated valet parking where infrastructure data provided via a wireless link augments in-vehicle sensor data. The use case respects the related safety requirements. Results show that the application awareness of the presented protocol, significantly improves service availability by transmitting data efficiently in time even under higher frame error rates.
Jonas Peeck, Mischa Möstl, Tasuku Ishigooka, Rolf Ernst
RTSS4
2021 Achieving safety and performance with reconfiguration protocol for ethernet TSN in automotive systems
Adam Kostrzewa, Rolf Ernst
J. Syst. Archit.2
2021 A Platform Programming Paradigm for Heterogeneous Systems Integration
abstract
To cope with growing computing performance requirements, cyber-physical systems architectures are moving toward heterogeneous high-performance computer architectures and networks. Such architectures, however, incur intricate side effects that challenge traditional software design and integration. The programming paradigm can take a key role in mastering software design, as experience in automotive design demonstrates. To cope with the integration challenge, this industry has started introducing a programming paradigm that efficiently preserves application data flow under platform integration and changes with minimum performance loss. This article will revisit this paradigm that is currently used for lock-free multicore programming and explain its extension to the system level. It will then explore its application to two important developments in industrial design. This article will conclude with an evaluation of its properties, its overhead, and its application toward a robust design process.
Kai-Björn Gemlau, Leonie Köhler, Rolf Ernst
Proc. IEEE3
2021 System-level Logical Execution Time: Augmenting the Logical Execution Time Paradigm for Distributed Real-time Automotive Software
abstract
Logical Execution Time (LET) is a timed programming abstraction, which features predictable and composable timing. It has recently gained considerable attention in the automotive industry, where it was successfully applied to master the distribution of software applications on multi-core electronic control units. However, the LET abstraction in its conventional form is only valid within the scope of a single component. With the recent introduction of System-level Logical Execution Time (SL LET), the concept could be transferred to a system-wide scope. This article improves over a first paper on SL LET, by providing matured definitions and an extensive discussion of the concept. It also features a comprehensive evaluation exploring the impacts of SL LET with regard to design, verification, performance, and implementability. The evaluation goes far beyond the contexts in which LET was originally applied. Indeed, SL LET allows us to address many open challenges in the design and verification of complex embedded hardware/software systems addressing predictability, synchronization, composability, and extensibility. Furthermore, we investigate performance trade-offs, and we quantify implementation costs by providing an analysis of the additionally required buffers.
Kai-Björn Gemlau, Leonie Köhler, Rolf Ernst, Sophie Quinton
ACM Trans. Cyber Phys. Syst.3
2020 EDA for Autonomous Behavior Assurance
abstract
Autonomous systems are self-governed and self-adaptive systems that must additionally comply with high assurance correctness and safety criteria. Such autonomous systems cannot be tested and verified in the traditional design process. While all systems hardware and software components can be implemented as usual, test and verification only cover the autonomous system functionality, but do not include the goal-driven autonomous behavior in all possible circumstances. This autonomous behavior is a primary design target. Thus, autonomous systems pose a number of emerging challenges and opportunities to the field of electronic design automation (EDA). Examples include specification of (evolving) requirements involving components and their interaction, defining different assurance levels for bounded operational environments, synthesis of mechanisms to guideline diagnosis and rigorously monitor systems integration after deployment.
Selma Saidi, Dirk Ziegenbein, Jyotirmoy V. Deshmukh, Rolf Ernst
ICCAD4
2020 Fast Failover in Ethernet-Based Automotive Networks
abstract
Application of Ethernet-based networks in the autonomous and automated vehicles requires fail-operational behavior. The network components must not only detect failures, but also reduce their effects on the system’s safety, as it may not be possible to bring the driver back into the control loop. In this work we present a mechanism which, through protocol-based synchronization and fine-grained re-configuration of network components, allows achieving traffic isolation, fault recovery and controlled degradation of network performance. This enables the maintenance of (degraded) network operation in case of a malfunction, a safe stopping of a vehicle or its safe operation until the driver takes control. We offer a detailed evaluation of our approach in a demonstrator setup as well as a simulation and comparison with state-of-the-art approaches.
Adam Kostrzewa, Rolf Ernst
ISORC2
2020 Safe Online Reconfiguration of Mixed-Criticality Real-Time Systems
abstract
In mixed-criticality real-time systems, system resources must be sufficiently separated to guarantee non-interference of critical functions. This leads to fixed boundaries between critical and non-critical parts of a system impeding change, repair, or proactive means to improve reliability. With conventional methods, moving the boundaries at runtime to manage the resource usage without respecting essential isolation requirements in mixed-criticality systems increases the risk of failure and must, therefore, be avoided. In this paper, we present and evaluate a safe online reconfiguration approach for many-core platforms. It satisfies safety standards isolation requirements and supports continuous real-time system operation as required in many application domains. It is particularly suited to provide proactive actions in the face of foreseen system failures improving availability without increased risk of failure during reconfiguration. Experimental results on a realistic usecase indicate that the proposed approach allows predictable online reconfigurations with tolerable latency overhead.
Thawra Kadeed, Borislav Nikolic, Rolf Ernst
PRDC3
2020 Weakly-hard Real-time Guarantees for Earliest Deadline First Scheduling of Independent Tasks
abstract
The current trend in modeling and analyzing real-time systems is toward tighter yet safe timing constraints. Many practical real-time systems can de facto sustain a bounded number of deadline-misses, i.e., they have Weakly-Hard Real-Time (WHRT) constraints rather than hard real-time constraints. Therefore, we strive to provide tight Deadline Miss Models (DMMs) in complement to tight response time bounds for such systems. In this work, we bound the distribution of deadline-misses for task sets running on uniprocessors using the Earliest Deadline First (EDF) scheduling policy. We assume tasks miss their deadlines due to transient overload resulting from sporadic jobs, e.g., interrupt service routines. We use Typical Worst-Case Analysis (TWCA) to tackle the problem in this context. Also, we address the sources of pessimism in computing DMMs, and we discuss the limitations of the proposed analysis. This work is motivated by and validated on a realistic case study inspired by industrial practice (satellite on-board software) and on a set of synthetic test cases. The synthetic experiment is dedicated to extensively study the impact of EDF on DMMs by presenting a comparison between DMMs computed under EDF and Rate Monotonic (RM). The results show the usefulness of this approach for temporarily overloaded systems when EDF scheduling is considered. They also show that EDF is especially useful for WHRT tasks.
Zain Alabedin Haj Hammadeh, Sophie Quinton, Rolf Ernst
ACM Trans. Embed. Comput. Syst.3
2019 Increasing Accuracy of Timing Models: From CPA to CPA+
abstract
Formal analysis methods of embedded systems provide safe, but unfortunately often pessimistic bounds on response times. An important source of pessimism is the common approach to characterize service request either by the amount of data or the number of events to be processed. Several works, e.g. [1]-[4], have demonstrated that a dual model - which includes information on both data and events - is more accurate, especially for more complex scheduling problems. In this paper, we enrich Compositional Performance Analysis (CPA) by a new component interface which, as we show, is consistent with the generic dual model proposed in [3]. Furthermore, we discuss how composition of components should be realized and how the new information should be integrated into the analysis technique. The improved CPA is called CPA+, and we identify different types of scenarios where CPA+ is particularly beneficial.
Leonie Köhler, Borislav Nikolic, Rolf Ernst, Marc Boyer
DATE3
2019 Slot-Based Transmission Protocol for Real-Time NoCs - SBT-NoC
abstract
To facilitate programming, most multi-core processors feature automated mechanisms maintaining coherence between each core's cache. These mechanisms introduce interference, that is, delays caused by concurrent access to a shared resource. This type of interference is hard to predict, leading to the mechanisms being shunned by real-time system designers, at the cost of potential benefits in both running time and system complexity. We believe that formal methods can provide the means to ensure that the effects of this interference are properly exposed and mitigated. Consequently, this paper proposes a nascent framework relying on timed automata to model and analyze the interference caused by cache coherence.
Borislav Nikolic, Robin Hofmann, Rolf Ernst
ECRTS3
2019 A new design for data-centric Ethernet communication with tight synchronization requirements for automated vehicles
abstract
Deterministic communication is a key challenge for modern embedded real-time systems. This is especially true for the automotive domain, where safety-critical functions are distributed over multiple electronic control units (ECUs) across the vehicle. To handle the increased complexity together with the higher data volume, Ethernet is considered as a universal media. Consequently, a great effort has been spent on making Ethernet predictable to satisfy timing and safety requirements. However, the communication stacks (COM-stacks) that build the bridge between applications and the network did not receive the same attention. Until now it has been dealt with as legacy software and overloaded with additional features to support the new multi-level protocols. Moreover, with the shift to automated vehicles, communication requirements will change fundamentally, transforming the car to a data centric system. In this paper we highlight the communication requirements of future autonomous cars and show why today's COM-stack designs are not suited to meet them. We present a novel design approach that ensures predictable timing and resource sharing while focusing on a minimalist and modular design. We also show how it fits to a state-of-the-art middleware, while keeping the complexity as small as possible.
Kai-Björn Gemlau, Jonas Peeck, Nora Sperling, Phil Hertha, Rolf Ernst
IECON5
2019 Improving a Compositional Timing Analysis Framework for Weakly-Hard Real-Time Systems
abstract
Weakly-hard real-time requirements for software tasks or network tasks are an adequate choice, if the system functionality is robust enough to tolerate a few deadline misses in transient overload phases. Weakly-hard real-time requirements can be formulated as (m, k)-constraints, i.e., no more than m deadline misses are allowed in a consecutive sequence of k jobs of a task. Existing formal analysis methods for weakly-hard real-time systems are so far limited to the scope of a single service-providing resource, only the recently published compositional analysis framework TypicalCPA addresses systems with multiple resources and partitioned scheduling. This paper substantially extends and improves TypicalCPA by addressing major open issues. Firstly, our improved version of TypicalCPA considers not only extra activation events of tasks as potential causes of transient overload but also long execution times of tasks which may exceptionally occur. Secondly, extra events and long execution times are now automatically and no longer manually identified as such. Finally, event propagation between components is also improved. We evaluate our proposed method in the context of an industrial case study.
Leonie Köhler, Rolf Ernst
RTAS2
2019 Self-Aware Scheduling for Mixed-Criticality Component-Based Systems
abstract
A basic mixed-criticality requirement in real-time systems is temporal isolation, which ensures that applications receive a guaranteed (CPU) service and impose a bounded interference on other applications. Providing operating system support for temporal isolation is often inefficient, in terms of utilisation and achieved latencies, or complex and hard to implement or model correctly. Correct models are, however, a prerequisite when response times are bounded by formal analyses. We provide a novel approach to this challenge by applying self-aware computing methodologies that involve run-time monitoring to detect (and correct) model deviations of a budget-based scheduler.
Johannes Schlatow, Mischa Möstl, Rolf Ernst
RTAS3
2019 Slack-based Traffic Shaping for Real-time Ethernet Networks
abstract
Ethernet has been identified as the most promising technology to be used as the backbone for the communication infrastructure in future automotive networks. Several traffic shapers have been proposed to meet the requirements of low latency and high throughput in automotive applications. Safety-critical traffic typically has strict timing requirements, however, meeting those requirements well ahead of time often brings no benefits. These spare capacities, called slack, could be exploited for the benefit of less critical but still time-sensitive traffic. Existing traffic shaping policies do not offer efficient possibilities to reduce and/or exploit this slack. In this paper, we propose a novel, scalable approach for slack-based traffic shaping, which can significantly improve the performance of time-sensitive traffic, at the expense of the slack of safety-critical traffic, which still fulfils all timing requirements. The promising results of the experimental evaluation suggest that the proposed approach represents an efficient means for meeting multi-dimensional and versatile requirements of current- and next-generation automotive applications.
Robin Hofmann, Borislav Nikolic, Rolf Ernst
RTCSA3
2019 Integrated Energy Control for Hard Real-Time Networks-on-Chip
abstract
While Networks-on-Chip (NoCs) are the prevalent solution to provide a scalable interconnect for the complex multiprocessing architectures, their associated energy consumptions have immensely increased. Specifically, hard real-time Networks-on-chip must manifest limited energy consumption as reliability issues in such a shared resource jeopardize the whole system safety. In this paper, we propose a safe and efficient approach that allows global and online energy management under temporal guarantees, i.e., all deadlines of critical functions are met. The approach introduces a control-layer to save energy on the NoC data layer through multiple Power-Aware Network Controllers (PANCs). We explore through PANCs the potential efficiency of integrating multiple energy-savings schemes in the face of the diversity of energy dissipation sources. To safely apply the PANCs in hard real-time systems while meeting the deadlines, a formal worst-case timing analysis of the additional latency induced by the control layer is provided. Experimental results demonstrate the diversity of NoC energy-savings under different combinations of energy-savings schemes. Also, the scalability of the approach is provided, inducing small area overhead.
Thawra Kadeed, Sebastian Tobuschat, Rolf Ernst
RTSS3
2019 Safe and efficient power management of hard real-time networks-on-chip
Thawra Kadeed, Sebastian Tobuschat, Adam Kostrzewa, Rolf Ernst
Integr.4
2019 Selective congestion control for mixed-critical networks-on-chip
Sebastian Tobuschat, Adam Kostrzewa, Rolf Ernst
Integr.3
2019 Real-time analysis of priority-preemptive NoCs with arbitrary buffer sizes and router delays
Borislav Nikolic, Sebastian Tobuschat, Leandro Soares Indrusiak, Rolf Ernst, Alan Burns 0001
Real Time Syst.4
2019 Providing Integrity in Real-Time Networks-on-Chip
abstract
Mixed-critical real-time systems must meet strict integrity, resilience, and timing constraints, as specified by safety standards. Due to the increasing threat of random hardware faults, efficiently achieving high reliability and dependability calls for cross-layer fault-tolerance solutions. This paper introduces the Advanced Integrity Q-service (AIQ), a mechanism to ensure the integrity and predictability of on-chip communication under random hardware faults. Devised for cross-layer and hierarchical fault-tolerance solutions, AIQ realizes low-overhead error detection in hardware and delegates error handling to arbitrary strategies in software. Experimental evaluation featuring benchmark applications and an industrial avionics use case shows that AIQ operates with high reliability and availability and low hardware and performance overheads. In a many-core mixed-critical platform under expected real-time scenarios, AIQ performs with execution time overhead between 1.4% and 7.1%.
Eberle A. Rambo, Yunsheng Shang, Rolf Ernst
IEEE Trans. Very Large Scale Integr. Syst.3
2018 Cross-layer dependency analysis with timing dependence graphs
abstract
We present Non-interference Analysis as a model-based method to automatically reveal, track and analyze end-to-end timing dependencies as part of a cross-layer dependency analysis in complex systems. Based on revealed timing dependencies of functional cause-effect chains, this method enables an automated FMEA inspection of timing behavior of individual functions. In consequence, this method can support safety-critical design processes w.r.t. the technical safety concept as mandated by safety standards such as ISO 26262. Our case-study from a state-of-the-art automated research vehicle and synthetic experiments confirm the applicability and scaleability of the proposed method.
Mischa Möstl, Rolf Ernst
DAC2
2018 Design methodologies for enabling self-awareness in autonomous systems
abstract
This paper deals with challenges and possible solutions for incorporating self-awareness principles in EDA design flows for autonomous systems. We present a holistic approach that enables self-awareness across the software/hardware stack, from systems-on-chip to systems-of-systems (autonomous car) contexts. We use the Information Processing Factory (IPF) metaphor as an exemplar to show how self-awareness can be achieved across multiple abstraction levels, and discuss new research challenges. The IPF approach represents a paradigm shift in platform design by envisioning the move towards a consequent platform-centric design in which the combination of self-organizing learning and formal reactive methods guarantee the applicability of such cyber-physical systems in safety-critical and high-availability applications.
Armin Sadighi, Bryan Donyanavard, Thawra Kadeed, Kasra Moazzemi, Tiago Rogério Mück, Ahmed Nassar 0001, Amir-Mohammad Rahmani, Thomas Wild, Nikil Dutt, Rolf Ernst, Andreas Herkersdorf, Fadi J. Kurdahi
DATE10
2018 Verifying Weakly-Hard Real-Time Properties of Traffic Streams in Switched Networks
abstract
In this paper, we introduce the first verification method which is able to provide weakly-hard real-time guarantees for tasks and task chains in systems with multiple resources under partitioned scheduling with fixed priorities. Existing weakly-hard real-time verification techniques are restricted today to systems with a single resource. A weakly-hard real-time guarantee specifies an upper bound on the maximum number m of deadline misses of a task in a sequence of k consecutive executions. Such a guarantee is useful if a task can experience a bounded number of deadline misses without impacting the system mission. We present our verification method in the context of switched networks with traffic streams between nodes, and demonstrate its practical applicability in an automotive case study.
Leonie Köhler, Sophie Quinton, Thomas Boroske, Rolf Ernst
ECRTS4
2018 Weakly-Hard Real-Time Guarantees for Weighted Round-Robin Scheduling of Real-Time Messages
abstract
Communication resources often exist in distributed real-time systems, therefore, providing guarantees on a predefined end-to-end deadline requires a timing analysis of the communication resource. Worst-case response time analysis techniques for guaranteeing the system's schedulability are not expressive enough for weakly-hard real-time systems. In weakly-hard real-time systems, the timing analysis ought to ensure that the distribution of the system deadline misses/mets is precisely bounded. In this paper, we compute weakly-hard real-time guarantees in the form of a deadline miss model using Typical Worst-Case Analysis for real-time messages with weighted round-robin scheduling. We evaluate the proposed analysis's scalability and the tightness of the computed deadline miss models. We illustrate also the applicability of our analysis in an industrial case study.
Zain Alabedin Haj Hammadeh, Rolf Ernst
ETFA2
2018 System Level LET: Mastering Cause-Effect Chains in Distributed Systems
abstract
Cause-effect chains are omnipresent in automotive and other distributed software, where sensor data is first collected and then processed in order to control actuators. It is challenging to find a correct implementation which respects all latency and precedence constraints of the numerous cause-effect chains contained in the software and which is at the same time robust towards software updates and changes. The Logical Execution Time (LET) programming model has proven to be a very promising solution featuring determinism in time and data flow as well as composability, but it is inherently limited to the scope of an electronic control unit (ECU) due to strict requirements w.r.t. clock synchronization and zero-time communication. In this paper, we present and evaluate a new concept, called System Level LET, which generalizes LET and extends it to the scope of the entire, distributed system while preserving determinism in time and data flow as well as composability. System Level LET has relaxed synchronization requirements, is independent of scheduling policies and can be combined with different communication semantics.
Rolf Ernst, Leonie Köhler, Kai-Björn Gemlau
IECON1
2018 Supporting Dynamic Voltage and Frequency Scaling in Networks-On-Chip for Hard Real-Time Systems
abstract
Networks-On-Chip (NoCs) designed to host advanced embedded applications require simultaneously high performance and efficient real-time guarantees under tight power limitations. Static power management is no longer sufficient to fulfill these goals in the context of a highly dynamic environment, e.g. autonomous driving, leading to mode-dependent system's behavior and substantial changes in traffic patterns. For this purpose, we propose a global and self-aware power management for NoCs. Dynamic voltage and frequency adjustments are performed at runtime depending on the number of currently active applications using protocol-based synchronization. The safety of a solution is assured by formal analysis and extensive experimental evaluation using an vehicle assistance functions, as well as synthetic benchmarks. The proposed scheme provides the desired performance, and when compared with the state-of-the-art approach, it significantly decreases power footprint of the NoC - up to 59% in the considered realistic use-case.
Adam Kostrzewa, Thawra Kadeed, Borislav Nikolic, Rolf Ernst
RTCSA4
2018 Platform-Centric Self-Awareness as a Key Enabler for Controlling Changes in CPS
abstract
Future cyber-physical systems will host a large number of coexisting distributed applications on hardware platforms with thousands to millions of networked components communicating over open networks. These applications and networks are subject to continuous change. The current separation of design process and operation in the field will be superseded by a life-long design process of adaptation, infield integration, and update. Continuous change and evolution, application interference, environment dynamics and uncertainty lead to complex effects which must be controlled to serve a growing set of platform and application needs. Self-adaptation based on self-awareness and self-configuration has been proposed as a basis for such a continuous in-field process. Research is needed to develop automated in-field design methods and tools with the required safety, availability, and security guarantees. The paper shows two complementary use cases of self-awareness in architectures, methods, and tools for cyber-physical systems. The first use case focuses on safety and availability guarantees in self-aware vehicle platforms. It combines contracting mechanisms, tool based self-analysis and self-configuration. A software architecture and a runtime environment executing these tools and mechanisms autonomously are presented including aspects of self-protection against failures and security threats. The second use case addresses variability and long term evolution in networked MPSoC integrating hardware and software mechanisms of surveillance, monitoring, and continuous adaptation. The approach resembles the logistics and operation principles of manufacturing plants which gave rise to the metaphoric term of an Information Processing Factory that relies on incremental changes and feedback control. Both use cases are investigated by larger research groups. Despite their different approaches, both use cases face similar design and design automation challenges which will be summarized in the end. We will argue that seemingly unrelated research challenges, such as in machine learning and security, could also profit from the methods and superior modeling capabilities of self-aware systems.
Mischa Möstl, Johannes Schlatow, Rolf Ernst, Nikil Dutt, Ahmed Nassar 0001, Amir-Mohammad Rahmani, Fadi J. Kurdahi, Thomas Wild, Armin Sadighi, Andreas Herkersdorf
Proc. IEEE3
2018 Synthesis of Monitors for Networked Systems With Heterogeneous Safety Requirements
abstract
Complex embedded systems such as automobiles and IoT-systems feature a wide range of applications with varying degrees of safety relevance. As many applications on these devices interlace more and more, new ways to guarantee sufficient isolation between safety levels become necessary. Previous work only regarded monitoring of timing properties for individual tasks on single resources and neglected functional cause-effect chains. This makes their application costly because nearly every task would require monitoring due to unknown system-level timing dependencies. We present cross-layer dependency analysis of timing properties and system-level synthesis strategies for monitors utilizing a reachability formulation of the problem. Our approach regards distributed cause-effect chains on networked platforms for which the accuracy of timing parameters can vary and heterogeneous safety requirements can be specified.
Mischa Möstl, Johannes Schlatow, Rolf Ernst
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 Adaptive load distribution in mixed-critical Networks-on-Chip
abstract
Modern Networks-on-Chip (NoCs) must accommodate a diversity of temporal requirements e.g. provide guarantees for real-time senders with the minimum impact on performance sensitive best-effort (BE) traffic. In this work, we propose a protocol-based adaptive load distribution which by selectively detouring BE traffic i.e. load balancing, allows to significantly improve NoC's performance without costly hardware extensions. The introduced method offers, during runtime, safe and efficient integration of mixed-critical workloads through the coupling of the flow control with the path selection based on the global NoC state. The requested real-time reliability of the interconnect is achieved through predictable synchronization with control messages supported by a formal analysis and an experimental evaluation.
Adam Kostrzewa, Sebastian Tobuschat, Leonardo Ecco, Rolf Ernst
ASP-DAC4
2017 Exploiting sporadic servers to provide budget scheduling for ARINC653 based real-time virtualization environments
abstract
Virtualization techniques for embedded real-time systems typically employ TDMA scheduling to achieve temporal isolation among different virtualized partitions. Due to the fixed TDMA schedule, worst case response times for IRQs and tasks are significantly increased. Recent publications introduced slack based IRQ shaping to mitigate this problem. While providing better response times for IRQs, those mechanisms neither improve task timings nor provide a work conserving scheduling. In order to provide such capabilities while still providing temporal isolation, we introduce a method based on the well known sporadic server model. In combination with a proposed budget scheduler the system is able to schedule a TDMA based configuration while providing better response times and the same amount of temporal isolation. We show correctness of the approach and evaluate it in a hypervisor implementation.
Matthias Beckert, Kai-Björn Gemlau, Rolf Ernst
DATE3
2017 Architecting high-speed command schedulers for open-row real-time SDRAM controllers
abstract
As SDRAM modules get faster and their data buses wider, researchers proposed the use of the open-row policy in command schedulers for real-time SDRAM controllers. While the real-time properties of such schedulers have been thoroughly investigated, their hardware implementation was not. Hence, in this paper, we propose a highly-parallel and multi-stage architecture that implements a state-of-the open-row real-time command scheduler. Moreover, we evaluate such architecture from the hardware overhead and performance perspectives.
Leonardo Ecco, Rolf Ernst
DATE2
2017 Bounding deadline misses in weakly-hard real-time systems with task dependencies
abstract
Real-time systems with functional dependencies between tasks often require end-to-end (as opposed to task-level) guarantees. For many of these systems, it is even possible to accept the possibility of longer end-to-end delays if one can bound their frequency. Such systems are called weakly-hard. In this paper we provide end-to-end deadline miss models for systems with task chains using Typical Worst-Case Analysis (TWCA). This bounds the number of potential deadline misses in a given sequence of activations of a task chain. To achieve this we exploit task chain properties which arise from the priority assignment of tasks in static-priority preemptive systems. This work is motivated by and validated on a realistic case study inspired by industrial practice and derived synthetic test cases.
Zain Alabedin Haj Hammadeh, Rolf Ernst, Sophie Quinton, Rafik Henia, Laurent Rioux
DATE2
2017 Self-awareness in autonomous automotive systems
abstract
Self-awareness has been used in many research fields in order to add autonomy to computing systems. In automotive systems, we face several system layers that must be enriched with self-awareness to build truly autonomous vehicles. This includes functional aspects like autonomous driving itself, its integration on the hardware/software platform, and among others dependability, real-time, and security aspects. However, self-awareness mechanisms of all layers must be considered in combination in order to build a coherent vehicle self-awareness that does not cause conflicting decisions or even catastrophic effects. In this paper, we summarize current approaches for establishing self-awareness on those layers and elaborate why self-awareness needs to be addressed as a cross-layer problem, which we illustrate by practical examples.
Johannes Schlatow, Mischa Möstl, Rolf Ernst, Marcus Nolte, Inga Jatzkowski, Markus Maurer, Christian Herber, Andreas Herkersdorf
DATE3
2017 Contract-based integration of automotive control software
abstract
The functionalities of automotive control are distributed over a large number of independently developed components that are interconnected by complex data dependencies. During integration it is critical to ensure the functional correctness of each component, due to the safety-critical nature of the automotive system. Thus existing integration processes ensure that interfaces are syntactically correct. Still in many cases communicated signals are semantically incompatible. This results in complicated errors that are hard to detect and fix. Moreover, existing component languages do not provide applicable means for the description and control of correspondent requirements. In this paper we present a novel methodology for an automated identification of integration errors in automotive control software. The key aspect of our approach are contracts, which are used to disclose domain level requirements. These contracts are then checked during integration supported by existing tools. A case study involving an existing engine control software shows the applicability of our approach by detecting a significant number of formerly unknown integration errors.
Tobias Sehnke, Matthias Schultalbers, Rolf Ernst
DATE3
2017 Real-time communication analysis for Networks-on-Chip with backpressure
abstract
Networks-on-Chip (NoCs) for safety-critical domains require formal guarantees for the worst-case behavior of all real-time senders. The majority of existing analysis approaches is capable of providing such guarantees only under the assumption that the queues in the routers never overflow, i.e., that no backpressure occurs. This leads to overly pessimistic guarantees or unfulfilled design requirements in many setups using commercially available NoCs where buffer space is limited. Therefore, we propose an alternative analysis methodology providing formal timing guarantees for packet latencies also in a NoC where backpressure occurs. The analysis allows exploiting the behavior of individual traffic streams to determine safe upper bounds on the latency of individual packets. The correctness of the analysis is evaluated experimentally through comparison with simulation results.
Sebastian Tobuschat, Rolf Ernst
DATE2
2017 Budgeting Under-Specified Tasks for Weakly-Hard Real-Time Systems
abstract
In this paper, we present an extension of slack analysis for budgeting in the design of weakly-hard real-time systems. During design, it often happens that some parts of a task set are fully specified while other parameters, e.g. regarding recovery or monitoring tasks, will be available only much later. In such cases, slack analysis can help anticipate how these missing parameters can influence the behavior of the whole system so that a resource budget can be allocated to them. It is, however, sufficient in many application contexts to budget these tasks in order to preserve weakly-hard rather than hard guarantees. We thus present an extension of slack analysis for deriving task budgets for systems with hard and weakly-hard requirements. This work is motivated by and validated on a realistic case study inspired by industrial practice.
Zain Alabedin Haj Hammadeh, Sophie Quinton, Marco Panunzio, Rafik Henia, Laurent Rioux, Rolf Ernst
ECRTS6
2017 Replica-Aware Co-Scheduling for Mixed-Criticality
abstract
Cross-layer fault-tolerance solutions are the key to effectively and efficiently increase the reliability in future safety-critical real-time systems. Replicated software execution with hardware support for error detection is a cross-layer approach that exploits future many-core platforms to increase reliability without resorting to redundancy in hardware. The performance of such systems, however, strongly depends on the scheduler. Standard schedulers, such as Partitioned~Strict Priority Preemptive (SPP) and Time-Division Multiplexing (TDM)-based ones, although widely employed, provide poor performance in face of replicated execution. In this paper, we propose the replica-aware co-scheduling for mixed-critical systems. Experimental results show schedulability improvements of more than 1.5x when compared to TDM and 6.9x when compared to SPP.
Eberle A. Rambo, Rolf Ernst
ECRTS2
2017 Towards model-based integration of component-based automotive software systems
abstract
The increasing complexity of automotive software systems and the desire for more frequent software and even feature updates require new approaches to the design, integration and testing of these systems. Ideally, those approaches enable an in-field updatability of automotive software systems that provides the same degree of safety guarantees as the traditionally lab-based deployment. In this paper, we present a layered modelling approach that formalises the integration procedure of automotive software systems using graph-based models and formal analyses.
Johannes Schlatow, Mischa Möstl, Rolf Ernst, Marcus Nolte, Inga Jatzkowski, Markus Maurer
IECON3
2017 Designing Networks-on-Chip for High Assurance Real-Time Systems
abstract
Conventional fault-tolerance approaches for Networks-on-Chip (NoCs) cannot be applied to high assurance real-time systems due to their different goals and constraints. These systems impose strict integrity, resilience and real-time requirements. All possible effects of hardware errors must be taken into account and the resulting system must be predictable, even in the presence of errors. In this paper, we present a wormhole-switched NoC with virtual channels for high assurance real-time systems hardened against soft errors. All possible duration and impacts of soft errors are taken into account and the resulting NoC operates with formal guarantees. Experimental evaluation shows that the network is able to provide a predictable behavior even in aggressive environments with very high error rates.
Eberle A. Rambo, Christoph Seitz, Selma Saidi, Rolf Ernst
PRDC4
2017 Demo Abstract: Bounding Deadline Misses for Weakly-Hard Real-Time Systems Designed in CAPELLA
abstract
Real-time systems with functional dependencies between tasks often require guarantees on end-to-end delays. For many of these systems, end-to-end deadline misses are accepted if one can limit their frequency. Such systems are called weaklyhard. Recent work has shown that typical worst-case analysis (TWCA) can compute an upper bound on the number of potential deadline misses in a sequence of activations of a task chain. In a joint collaboration between Thales and TU Braunschweig, the use of TWCA to limit the number of deadline misses in an aerial video tracking (AVT) system was evaluated. The AVT case-study, the complete automated model-based tool chain from the design environment to the timing verification using TWCA, as well as the results of the evaluation will be presented in the demonstration. The tool chain involves four tools: the design modeling tool CAPELLA extended by a performance viewpoint which allows annotating the design model with timing properties needed to perform TWCA, the pivot model TEMPO which handles mismatches between the semantics of the design model and the semantics of the model used in TWCA, the scheduling analysis tool pyCPA that performs TWCA and finally the graphical tool TimingGraphics used to visualize the TWCA results. To show the pertinence of the use of TWCA, we will also compare in the demonstration the obtained results with those obtained using worst-case analysis and simulation.
Rafik Henia, Lisa Roux, Nicolas Sordon, Zain Alabedin Haj Hammadeh, Rolf Ernst, Sophie Quinton
RTAS5
2017 Efficient Latency Guarantees for Mixed-Criticality Networks-on-Chip
abstract
Networks-on-Chip (NoCs) for future mixed-criticality systems must handle a growing variety of traffic requirements, ranging from safety-critical real-time traffic to bursty latency-sensitive best-effort traffic. Additionally, safety standards (e.g. ISO 26262) require sufficient independence among different criticality levels or a full system certification according to the highest applicable safety level. Hence, a NoC must provide performance isolation for safety-critical traffic, while sustaining low latency for best-effort traffic. This paper presents a run-time configurable NoC design enabling latency guarantees for safety-critical traffic with reduced adverse impact on the performance of best-effort traffic. In contrast to existing approaches, we prioritize best-effort over safety-critical traffic and only switch priorities when required. Doing this, we exploit the latency slack of safety-critical applications, while providing sufficient independence among different criticality levels w.r.t. timing properties. We present a formal analysis and an experimental evaluation, showing that the approach provides performance isolation for safety-critical applications, while reducing the adverse effects through strict prioritization on best-effort applications.
Sebastian Tobuschat, Rolf Ernst
RTAS2
2017 Ensuring safety and efficiency in networks-on-chip
Adam Kostrzewa, Selma Saidi, Leonardo Ecco, Rolf Ernst
Integr.4
2017 Tackling the Bus Turnaround Overhead in Real-Time SDRAM Controllers
abstract
Synchronous dynamic random access memories (SDRAMs) are widely employed in multiand many-core platforms due to their high-density and low-cost. Nevertheless, their benefits come at the price of a complex two-stage access protocol, which reflects their bank-based structure and an internal level of explicitly managed caching. In scenarios in which requestors demand real-time guarantees, these features pose a predictability challenge and, in order to tackle it, several SDRAM controllers have been proposed. In this context, recent research shows that a combination of bank privatization and open-rowpolicy (exploiting the caching over the boundary of a single request) represents an effective way to tackle the problem. However, such approach uncovered a new challenge: the data bus turnaround overhead. In SDRAMs, a single data bus is shared by read and write operations. Alternating read and write operations is, consequently, highly undesirable, as the data bus must remain idle during a turnaround. Therefore, in this article, we propose a SDRAM controller that reorders read and write commands, which minimizes data bus turnarounds. Moreover, we compare our approach analytically and experimentally with existing real-time SDRAM controllers both from the worst-case latency and power consumption perspectives.
Leonardo Ecco, Rolf Ernst
IEEE Trans. Computers2
2017 Response Time Analysis for Sporadic Server Based Budget Scheduling in Real Time Virtualization Environments
abstract
Virtualization techniques for embedded real-time systems typically employ TDMA scheduling to achieve temporal isolation among different virtualized applications. Recent work already introduced sporadic server based solutions relying on budgets instead of a fixed TDMA schedule. While providing better average-case response times for IRQs and tasks, a formal response time analysis for the worst-case is still missing. In order to confirm the advantage of a sporadic server based budget scheduling, this paper provides a worst-case response time analysis. To improve the sporadic server based budget scheduling even more, we provide a background scheduling implementation which will also be covered by the formal analysis. We show correctness of the analysis approach and compare it against TDMA based systems. In addition to that, we provide response time measurements from a working hypervisor implementation on an ARM based development board.
Matthias Beckert, Rolf Ernst
ACM Trans. Embed. Comput. Syst.2
2017 Response-Time Analysis for Task Chains with Complex Precedence and Blocking Relations
abstract
For the development of complex software systems, we often resort to component-based approaches that separate the different concerns, enhance verifiability and reusability, and for which microkernel-based implementations are a good fit to enforce these concepts. Composing such a system of several interacting software components will, however, lead to complex precedence and blocking relations, which must be taken into account when performing latency analysis. When modelling these systems by classical task graphs, some of these effects are obfuscated and tend to render such an analysis either overly pessimistic or even optimistic. We therefore firstly present a novel task (meta-)model that is more expressive and accurate w.r.t. these (functional) precedence and mutual blocking relations. Secondly, we apply the busy-window approach and formulate a modular response-time analysis on task-chain level suitable but not restricted to static-priority scheduled systems. We show that the conjunction of both concepts allows the calculation of reasonably tight latency bounds for scenarios not adequately covered by related work.
Johannes Schlatow, Rolf Ernst
ACM Trans. Embed. Comput. Syst.2
2016 Dynamic admission control for real-time networks-on-chips
abstract
Networks-on-Chip (NoCs) for real-time systems require solutions for safe and predictable sharing of network resources between transmissions with different quality-of service requirementrs. In this work, we present a mechanism for a global and dynamic admission control in NoCs designed for realtime systems. It introduces an overlay network to synchronize transmissions using arbitration units called Resource Managers (RMs), which allows a global and work-conserving scheduling. We present a formal worst-case timing analysis for the proposed mechanism and demonstrate that this solution not only exposes higher performance in simulation but, even more importantly, consistently reaches smaller formally guaranteed worst-case latencies than TDM for realistic levels of system's utilization. Our mechanism does not require modification of routers and therefore can be used together with any architecture utilizing non-blocking routers.
Adam Kostrzewa, Selma Saidi, Leonardo Ecco, Rolf Ernst
ASP-DAC4
2016 Invited - Towards fail-operational ethernet based in-vehicle networks
abstract
In the future, vehicles are expected to act more and more autonomously. The transition towards highly automated and autonomous driving will push the safety requirements for in-vehicle networks. Such networks must support isolation between mixed-critical traffic (e.g. critical control and non-critical infotainment) and must be fail-operational. This paper will present new concepts and mechanisms to achieve these goals in Ethernet-based networks. It will cover advanced topics such as software defined networking (SDN) to implement isolation, fault recovery, and controlled degradation, e.g. to maintain (degraded) operation until the driver takes over or to reach a safe stop.
Mischa Möstl, Daniel Thiele, Rolf Ernst
DAC3
2016 Guarantees for runnable entities with heterogeneous real-time requirements
Leonie Köhler, Zain Alabedin Haj Hammadeh, Rolf Ernst
DATE3
2016 Slack-based resource arbitration for real-time Networks-on-Chip
Adam Kostrzewa, Selma Saidi, Rolf Ernst
DATE3
2016 Handling complex dependencies in system design
Mischa Möstl, Rolf Ernst
DATE2
2016 Providing formal latency guarantees for ARQ-based protocols in Networks-on-Chip
Eberle A. Rambo, Selma Saidi, Rolf Ernst
DATE3
2016 Formal analysis based evaluation of software defined networking for time-sensitive Ethernet
Daniel Thiele, Rolf Ernst
DATE2
2016 Formal worst-case timing analysis of Ethernet TSN's burst-limiting shaper
Daniel Thiele, Rolf Ernst
DATE2
2016 The EMC2 Project on Embedded Microcontrollers: Technical Progress after Two Years
abstract
Since April 2014 the Artemis/ECSEL project EMC2 is running and provides significant results. EMC2 stands for "Embedded Multi-Core Systems for Mixed Criticality Applications in Dynamic and Changeable Real-Time Environments". In this paper we report recent progress on technical work in the different workpackages and use cases. We highlight progress in the research on system architecture, design methodology, platform and operating systems, and in qualification and certification. Application cases in the fields of automotive, avionics, health care, and industry are presented exploiting the technical results achieved.
Werner Weber, Alfred Hoess, Jan van Deventer, Frank Oppenheimer, Rolf Ernst, Adam Kostrzewa, Philippe Dore, Thierry Goubier, Haris Isakovic, Norbert Druml, Egon Wuchner, Daniel Schneider 0001, Erwin Schoitsch, Eric Armengaud, Thomas Soderqvist, Massimo Traversone, Sascha Uhrig, Juan-Carlos Perez-Cortes, Sergio Sáez, Juha Kuusela, Mark van Helvoort, Xing Cai, Bjørn Nordmoen, Geir Yngve Paulsen, Hans Petter Dahle, Michael Geissel, Jürgen Salecker, Peter Tummeltshammer
DSD5
2016 Minimizing DRAM Rank Switching Overhead for Improved Timing Bounds and Performance
abstract
Multi-rank DRAM modules have been identified as a flexible option for accommodating large mixed critical workloads. However, because all ranks in a module share the same multi-drop data bus, a penalty in the form of idle cycles is necessary when alternating data transfers between different ranks. Moreover, as the data bus clock frequency of DRAM modules becomes higher, such penalty increases significantly and can no longer be neglected. Therefore, in this paper, we propose a mixed critical real-time controller for multi-rank DRAM modules that minimizes rank switches. Our controller works by scheduling batches of data transfers for each rank and performing rank switches only in the end of each batch. We provide a detailed timing analysis of our approach and a comparison with a state-of-the-art counterpart. For a dual-rank scenario, our approach increases DRAM utilisation, thus reducing the latency bounds of hard real-time applications by on average 14% and decreasing the average request latency of soft real-time applications by on average 51%.
Leonardo Ecco, Adam Kostrzewa, Rolf Ernst
ECRTS3
2016 Zero-time communication for automotive multi-core systems under SPP scheduling
abstract
Multi-core CPUs are quickly gaining importance in automotive ECUs. While using multi-core architectures for application integration is meanwhile reasonably well understood, parallelization of existing task sets and partitioning of future computation intensive tasks still shows performance limitations and challenges portability and flexibility. The logical execution time (LET) paradigm has been proposed to control core-to-core communication timing which is one of the bottlenecks for automotive system parallelization. We show how to improve the predictability and performance of core-to-core communication by applying only minor modifications to the currently used partitioned static scheduling strategy. The basic idea is to guarantee better response times for low-priority tasks by boosting its priority, thereby using higher priority system slack. It can be selectively applied to individual tasks, and can implement the LET paradigm on all or on a subset of the tasks of a system. It is applicable to lower granularities down to individual runnables which is highly important in automotive applications with large container tasks. We show correctness of the approach and evaluate its performance in a microkernel implementation where it exhibits high performance.
Matthias Beckert, Mischa Möstl, Rolf Ernst
ETFA3
2016 Formal worst-case performance analysis of time-sensitive Ethernet with frame preemption
abstract
One of the key challenges in future Ethernet-based automotive and industrial networks is the low-latency transport of time-critical data. To date, Ethernet frames are sent non-preemptively. This introduces a major source of delay, as, in the worst-case, a latency-critical frame might be blocked by a frame of lower priority, which started transmission just before the latency-critical frame. The upcoming IEEE 802.3br standard will introduce Ethernet frame preemption to address this problem. While high-priority traffic benefits from preemption, lower-priority (yet still latency-sensitive) traffic experiences a certain overhead, impacting its timing behavior. In this paper, we present a formal timing analysis for Ethernet to derive worst-case latency bounds under preemption. We use a realistic automotive Ethernet setup to analyze the worst-case performance of standard Ethernet and Ethernet TSN under preemption and also compare our results to non-preemptive implementations of these standards.
Daniel Thiele, Rolf Ernst
ETFA2
2016 Safe and dynamic traffic rate control for networks-on-chips
abstract
Networks-on-Chip (NoCs) for real-time systems require solutions for a safe and predictable sharing of resources between transmissions with different quality-of service (QoS) requirements. In this work, we present a mechanism which allows to apply existing wormhole-switched and performance optimized NoCs in safety critical domains, without requiring complex hardware modifications. For this purpose, we introduce a global and dynamic admission control mechanism implemented in the form of an access layer, controlling the rates at which running applications can access the NoC. The mechanism allows to enforce behavioral models for different data streams as well as to dynamically adapt the rates values to the number of currently active applications. We prove this important feature using formal timing analysis. Our approach results in a higher performance and tighter guarantees while simultaneously decreasing hardware (up to 60%) and temporal overhead (up to 80%) when compared with existing solutions.
Adam Kostrzewa, Sebastian Tobuschat, Rolf Ernst, Selma Saidi
NOCS3
2016 Response-Time Analysis for Task Chains in Communicating Threads
abstract
When modelling software components for timing analysis, we typically encounter functional chains of tasks that lead to precedence relations. As these task chains represent a functionally-dependent sequence of operations, in real-time systems, there is usually a requirement for their end-to-end latency. When mapped to software components, functional chains often result in communicating threads. Since threads are scheduled rather than tasks, specific task chain properties arise that can be exploited for response-time analysis. As a core contribution, this paper presents an extension of the busy-window analysis suitable for such task chains in static-priority preemptive systems. We evaluated the extended busy-window analysis in a compositional performance analysis using synthetic test cases and a realistic automotive use case showing far tighter response-time bounds than current approaches.
Johannes Schlatow, Rolf Ernst
RTAS2
2016 Demo Abstract: Response-Time Analysis for Task Chains in Communicating Threads with pyCPA
abstract
Summary form only given. When modelling software components for timing analysis, we typically encounter functional chains of tasks that lead to precedence relations. As these task chains represent a functionally-dependent sequence of operations, in real-time systems, there is usually a requirement for their end-to-end latency. When mapped to software components, functional chains often result in communicating threads. Since threads are scheduled rather than tasks, specific task chain properties arise that can be exploited for response-time analysis by extending the busy-window analysis for such task chains in static-priority preemptive systems. We implemented this analysis by means of an analysis extension for pyCPA, a research-grade implementation of compositional performance analysis (CPA). The major scope of this demo is to show how CPA can be reasonably performed for realistic component-based systems. It also demonstrates how research on and with CPA is conducted using the pyCPA analysis framework. In the course of this demo, we show two approaches for the extraction of an appropriate timing model: 1) the derivation from a contract-based specification of the software components and 2) a tracing-based approach suitable for black-box components. We also demonstrate how this timing model is fed into the analysis extension in order to obtain response-time results for the task chains of interest. Finally, we present how the developed analysis extension speeds up the CPA and therefore enables an automated design-space exploration and optimisation of the threads' priority assignments in order to satisfy the pre-defined latency requirements.
Johannes Schlatow, Jonas Peeck, Rolf Ernst
RTAS3
2016 Special issue of the Euromicro Conference on Real-Time Systems (ECRTS)
Rolf Ernst
Real Time Syst.1
2016 Formal timing analysis of CAN-to-Ethernet gateway strategies in automotive networks
abstract
Due to increased bandwidth and scalability demands, Ethernet technology is finding its way into recent in-vehicle networks. Tomorrow’s heterogeneous networks will feature legacy buses [e.g. controller area network (CAN) or FlexRay] as well as high-speed Ethernet devices, connected by switches and gateways. As Ethernet offers significantly larger frame sizes than CAN, the efficient transmission of CAN data over an Ethernet backbone depends heavily on the way this data is multiplexed into Ethernet frames. This article focuses on the timing impact introduced by various CAN/Ethernet multiplexing strategies at the gateways. We present a formal analysis method to derive upper bounds on end-to-end latencies for complex multiplexing strategies, which is key for the design of safety-critical real-time systems. We capture complex inter-domain signal paths spanning multiple buses, gateways, and switches and show the applicability in a realistic automotive setup.
Daniel Thiele, Johannes Schlatow, Philip Axer, Rolf Ernst
Real Time Syst.4
2016 Guest Editorial for Special Issue of ESWEEK 2015
abstract
No abstract available.
Petru Eles, Rolf Ernst
ACM Trans. Embed. Comput. Syst.2
2015 Designing time partitions for real-time hypervisor with sufficient temporal independence
abstract
Virtualization techniques for embedded real-time systems, as known from the Integrated Modular Avionics (IMA) architecture of the ARINC653 standard, typically employ a TDMA scheduling to achieve temporal isolation among different virtualized partitions. Due to the fixed TDMA schedule, the worst case interrupt response times are significantly increased. An already proposed technique to mitigate this problem is to allow interrupts within an TDMA schedule, in order to achieve better interrupt response times while maintaining a sufficient degree of temporal independence via monitoring. In this paper we propose a novel approach that optimizes the TDMA schedule based on the partitions internal timing behavior and tasks parameters. The developed optimization algorithm generates a maximum amount of slack within the TDMA cycle. This slack is later used to interpose interrupts, while maintaining the interference with a monitor. We show correctness of the approach and evaluate it in a hypervisor implementation.
Matthias Beckert, Rolf Ernst
DAC2
2015 Improving formal timing analysis of switched ethernet by exploiting FIFO scheduling
abstract
Ethernet is an emerging technology in the automotive domain and is capable to overcome the bandwidth and scalability limits of traditional buses like CAN or FlexRay. Formal performance analysis methods are required to verify the timing, e.g. by providing upper bounds on end-to-end latencies, in safety-critical real-time systems, such as automotive control and advanced driver assistance systems. In many real-time capable Ethernet implementations such as IEEE 802.1Q or AVB, frames can be prioritized and frames of equal priority are scheduled in FIFO order at the switch ouput ports. In this paper, we show how to exploit Ethernet's FIFO scheduling in a compositional formal performance analysis to derive tighter timing guarantees. In an automotive Ethernet setup, our proposed analysis leads to a significant reduction in end-to-end latency guarantees.
Daniel Thiele, Philip Axer, Rolf Ernst
DAC3
2015 Worst-case communication time analysis of networks-on-chip with shared virtual channels
Eberle A. Rambo, Rolf Ernst
DATE2
2015 Improved Deadline Miss Models for Real-Time Systems Using Typical Worst-Case Analysis
abstract
We focus on the problem of computing tight deadline miss models for real-time systems, which bound the number of potential deadline misses in a given sequence of activations of a task. In practical applications, such guarantees are often sufficient because many systems are in fact not hard real-time. Our major contribution is a general formulation of that problem in the context of systems where some tasks occasionally experience sporadic overload. Based on this new formulation, we present an algorithm that can take into account fine-grained effects of overload at the input of different tasks when computing deadline miss bounds. Finally, we show in experiments with synthetic as well as industrial data that our algorithm produces bounds that are much tighter than in previous work, in sufficiently short time.
Wenbo Xu 0002, Zain Alabedin Haj Hammadeh, Alexander Kröller, Rolf Ernst, Sophie Quinton
ECRTS4
2015 Parallel feature extraction and heterogeneous object-detection for multi-camera driver assistance systems
abstract
We present a flexible architecture for image-based feature detection and object classification on an FPGA. This architecture is tailored to the requirements of future driver assistance systems, which will make it necessary to detect a wide range of different object types in multi-camera systems requiring highly efficient hardware. In contrast to other designs, which typically address a specific object type or only accelerate early processing steps, the proposed pipeline offers different operation modes to switch resources for either detection or classification speed. In addition, the architecture can incorporate heterogeneous processors for different feature types. The design is tailored to support any object detection system using weak features and cascaded classifiers. For evaluation, a classic Viola Jones Detector is implemented being fully compatible with OpenCV.
Stefan Wonneberger, Peter Mühlfellner, Pedro Ceriotti, Thorsten Graf 0001, Rolf Ernst
FPL5
2015 An approach for physical topology exploration in wired bus networks
abstract
Fieldbus networks are widely used in the automation area. Normally the wired fieldbus networks can detect the existence of nodes and links, but lack the capability of measuring their physical positions. However, the information about nodes physical location is useful for network maintenance and control. Moreover, there is a trend to develop location-based and context-based services in the future. Particularly in some implementations, high bandwidth and critical time performance are not the most essential demands, people may prefer to distribute the nodes in a flexible way and keep tracking their positions. This paper describes a wired bus network, which can support large scale networks with flexible node distribution. Based on this network, we propose a method to measure wire length and explore the bus physical topology. This approach is able to identify the distance between nodes and detect the position of branches. The experiment demonstrates that the approach achieves high precision in the exploration of network physical topology.
Yidi Zeng, Harald Schrom, Rolf Ernst
ISCAS3
2015 Improved DRAM Timing Bounds for Real-Time DRAM Controllers with Read/Write Bundling
abstract
As DRAMs become faster, the penalty to reverse the direction of their data buses increases. Yet, existing real-time memory controllers do not reorder read and write commands. Hence, timing bounds are computed by assuming an alternating pattern of reads and writes, thus accounting for several data bus direction reversals, consequently leading to suboptimal results. Therefore, in this paper, we propose a memory controller that reorders read and write commands, which minimizes reversals. Moreover, we prove through a detailed timing analysis that the effect of the reordering is bounded. Finally, we compare our approach analytically with a state-of-the-art real-time memory controller and show that our timing bounds are up to 27% better.
Leonardo Ecco, Rolf Ernst
RTSS2
2015 Dynamic Control for Mixed-Critical Networks-on-Chip
abstract
Networks-on-Chip (NoCs) for future real-time systems must provide service guarantees for applications with different levels of criticality. In this work, we propose an efficient mechanism for supporting mixed-criticality which combines the global, work-conserving scheduling for the end to end guarantees with the local arbitration in routers. We introduce a dynamic control layer with a central Resource Manager (RM) synchronizing transmissions with a dedicated protocol. The proposed mechanism allows to improve over existing solutions through reducing hardware overhead compared to non-blocking routers with rate control as well as temporal overhead compared to Time-Division Multiplexing (TDM). By using formal analysis, we show that RMs provide efficient service guarantees to all synchronized applications. We validate experimentally, using benchmarks, these guarantees along with the performance of the mechanism and induced overhead.
Adam Kostrzewa, Selma Saidi, Rolf Ernst
RTSS3
2014 Exploiting Shaper Context to Improve Performance Bounds of Ethernet AVB Networks
abstract
New hard real-time Advanced Driver Assistance Systems such as the Collision-Avoidance System push the bandwidth requirements of the communication infrastructure to a new level. Controller Area Network (CAN) and FlexRay are reaching their limits. Ethernet-based automotive networks such as Ethernet AVB are capable of addressing these requirements. However, designing predictable Ethernet networks is more complex than the design of a traditional CAN bus. Formal real-time performance characteristics are key to a successful Ethernet integration. In this paper we present an improved Ethernet AVB performance analysis which exploits traffic-stream correlations. The results are significantly tighter compared to related work.
Philip Axer, Daniel Thiele, Rolf Ernst, Jonas Diemer
DAC3
2014 Sufficient Temporal Independence and Improved Interrupt Latencies in a Real-Time Hypervisor
abstract
Virtualization techniques for hard real-time systems typically employ TDMA scheduling to achieve temporal isolation among partitions. The processing of user-level interrupt handlers is only performed within appropriate time slots, thus significantly increasing interrupt latencies.
Matthias Beckert, Moritz Neukirchner, Rolf Ernst, Stefan M. Petters
DAC3
2014 Typical Worst Case Response-Time Analysis and its Use in Automotive Network Design
abstract
For some automotive applications, worst case performance guarantees are too expensive, but a minimum level of performance must be formally guaranteed. For such applications, we have developed an approach called Typical Worst Case Analysis (TWCA) which can formally bound the number of violations of the computed response-time guarantee in a given time window. In this paper, we demonstrate how it can be used to analyze a real CAN bus with complex load patterns. We investigate the effects of these load patterns and show how the necessary parameters can be derived and verified from traces and specifications. We compare the results to the commonly used base load approximation --- like a 50%-limit for cyclic load --- showing superior accuracy and expressiveness.
Sophie Quinton, Torsten T. Bone, Julien Hennig, Moritz Neukirchner, Mircea Negrean, Rolf Ernst
DAC6
2014 Failure analysis of a network-on-chip for real-time mixed-critical systems
abstract
Multi- and many-core architectures using Networks-on-Chip (NoC) are being explored for use in real-time safety-critical applications for their performance and efficiency. Such systems must provide isolation between tasks that may present distinct criticality levels. The NoC is critical to maintain the isolation property as it is a heavily used shared resource. To meet safety-standard requirements, such architectures require a systematic evaluation of the effects of all possible failures such as in the form of a Failure Mode and Effects Analysis (FMEA). We present the results of a detailed system-level analysis of a typical real-time mixed-critical network-on-chip architecture. This comprises an FMEA and error effects classification regarding duration and isolation violation.
Eberle A. Rambo, Alexander Tschiene, Jonas Diemer, Leonie Köhler, Rolf Ernst
DATE5
2014 Extending typical worst-case analysis using response-time dependencies to bound deadline misses
abstract
Weakly-hard time constraints have been proposed for applications where occasional deadline misses are permitted. Recently, a new approach called Typical Worst-Case Analysis (TWCA) has been introduced which exploits similar constraints to bound response times of systems with sporadic overload. In this paper, we extend that approach for static priority preemptive and non-preemptive scheduling to determine the maximum number of deadline misses for a given deadline. The approach is based on an optimization problem which trades off higher priority interference versus miss count. We formally derive a lattice structure for the possible combinations that lays the ground for an integer linear programming (ILP) formulation. The ILP solution is evaluated showing effectiveness of the approach and far better results than previous TWCA.
Zain Alabedin Haj Hammadeh, Sophie Quinton, Rolf Ernst
EMSOFT3
2014 Efficient 3D triangulation in hardware for dense structure-from-motion in low-speed automotive scenarios
abstract
With the introduction of surround view cameras in modern vehicles and the possibility of calculating dense motion fields in real-time from a moving camera a detailed 3D reconstruction of the static environment is possible (structure-from-motion). Beside the necessity of a motion field between two image frames the task of triangulating those individual 2D point matches to 3D points in the world becomes non real-time on modern CPUs when to be repeated for all image points. In this work we evaluate different approaches to the 3D triangulation optimization problem in a typical structure-from-motion processing chain for an efficient implementation in hardware. An architecture for solving this problem using linear triangulation with an inhomogeneous solution to the equation system is proposed. We evaluate our implementation using FPGAs against a software-implementation with synthetic datasets and from low-speed parking area scenes for numerical accuracy and real-time capabilities. In addition the proposed fixed-point arithmetic implementation is compared against an implementation using floating-point units.
Stefan Wonneberger, Max Kohler, Wojciech Derendarz, Thorsten Graf 0001, Rolf Ernst
FPL5
2014 FMEA-based analysis of a Network-on-Chip for mixed-critical systems
abstract
Network-on-Chip-based multi- and many-core architectures show high potential for use in safety-critical real-time applications, such as Flight Management Systems, considering their superior efficiency. For such use however, safety standards require proof that the architecture meets the specified security goals. This usually involves a Failure Mode and Effects Analysis (FMEA) to reveal the effects of all potential failures. Moreover, the Network-on-Chip (NoC) is a shared component and plays central role in a mixed-critical system, which must guarantee the isolation between tasks, that may have distinct criticality levels. We present an FMEA-based system-level analysis for NoCs designed for mixed-critical systems. It comprises FMEA, error effects classification regarding duration and isolation violation, and technology independent probability assessment. The analysis gives effective insight into fault-related weaknesses of the NoC, and enables considerable improvements to the NoC's resilience with minimal overhead. We apply it to a typical packet-switched NoC and present the results. Although developed for safety critical applications, the approach can be applied to improve the robustness of general systems.
Eberle A. Rambo, Alexander Tschiene, Jonas Diemer, Leonie Köhler, Rolf Ernst
NOCS5
2014 A mixed critical memory controller using bank privatization and fixed priority scheduling
abstract
Mixed critical platforms are those in which applications that have different criticalities, i.e. different levels of importance for system safety, coexist and share resources. Such platforms require a memory controller capable of providing sufficient timing independence for critical applications. Existing real-time memory controllers, however, either do not support mixed criticality or still allow a certain degree of interference between applications. The former issue leads to overly constrained, and hence more expensive, systems. The latter issue forces designers to assume the worst case latency for every individual memory transaction, which can be very conservative when applied to determine the worst-case execution time (WCET) of a task that performs many memory requests. In this paper, we address both issues. The main contributions are: (1) A memory controller that allows a predetermined number of critical and non-critical applications to coexist, while providing an interference-free memory for the former. To achieve that, we treat the memory as a set of independent virtual devices (VDs). Therefore, we also provide (2) a partitioning strategy to properly map mixed critical workloads to VDs. We present experiments that show that our controller allows DRAM sharing with no interference on critical applications and minimal performance overhead on non-critical ones (they perform on average only 15% slower in the shared environment).
Leonardo Ecco, Sebastian Tobuschat, Selma Saidi, Rolf Ernst
RTCSA4
2014 Building timing predictable embedded systems
abstract
A large class of embedded systems is distinguished from general-purpose computing systems by the need to satisfy strict requirements on timing, often under constraints on available resources. Predictable system design is concerned with the challenge of building systems for which timing requirements can be guaranteed a priori . Perhaps paradoxically, this problem has become more difficult by the introduction of performance-enhancing architectural elements, such as caches, pipelines, and multithreading, which introduce a large degree of uncertainty and make guarantees harder to provide. The intention of this article is to summarize the current state of the art in research concerning how to build predictable yet performant systems. We suggest precise definitions for the concept of “predictability”, and present predictability concerns at different abstraction levels in embedded system design. First, we consider timing predictability of processor instruction sets. Thereafter, we consider how programming languages can be equipped with predictable timing semantics, covering both a language-based approach using the synchronous programming paradigm, as well as an environment that provides timing semantics for a mainstream programming language (in this case C). We present techniques for achieving timing predictability on multicores. Finally, we discuss how to handle predictability at the level of networked embedded systems where randomly occurring errors must be considered.
Philip Axer, Rolf Ernst, Heiko Falk, Alain Girault, Daniel Grund, Nan Guan, Bengt Jonsson 0001, Peter Marwedel, Jan Reineke 0001, Christine Rochange, Maurice Sebastian, Reinhard von Hanxleden, Reinhard Wilhelm, Wang Yi 0001
ACM Trans. Embed. Comput. Syst.2
2013 Stochastic response-time guarantee for non-preemptive, fixed-priority scheduling under errors
abstract
Error recovery mechanisms, such as automatic repeat request (ARQ) for e.g. the CAN protocol, are a crucial part of safety critical embedded systems. These can have a strong impact on the timing behavior of the system and an unpropitious combination of error events may cause a real-time application to miss deadlines with potentially hazardous consequences. Therefore, formal analysis of the worst-case timing including errors is indispensable for certification. We present a new convolution-based stochastic analysis in which we model errors as additional execution time to bound the probability for an activation to exceed a response-time value in the worst-case.
Philip Axer, Rolf Ernst
DAC2
2013 Timing analysis of multi-mode applications on AUTOSAR conform multi-core systems
abstract
Many real-time embedded systems execute multi-mode applications, i.e. applications that can change their functionality over time. With the advent of multi-core embedded architectures, the system design process requires appropriate support for accommodating multi-mode applications on multiple cores which share common resources. Various mode change and resource arbitration protocols, and corresponding timing analysis solutions were proposed for either multi-mode or multi-core real-time applications. However, no attention was given to multi-mode applications that share resources when executing on multi-core systems. In this paper, we address this subject in the context of automotive multi-core processors using AUTOSAR. We present an approach for safely handling shared resources across mode changes and provide a corresponding timing analysis method.
Mircea Negrean, Sebastian Klawitter, Rolf Ernst
DATE3
2013 Sensitivity analysis for arbitrary activation patterns in real-time systems
abstract
Response time analysis, which determines whether timing guarantees are satisfied for a given system, has matured to industrial practice and is able to consider even complex activation patterns modelled through arrival curves or minimum distance functions. On the other side, sensitivity analysis, which determines bounds on parameter variations under which constraints are still satisfied, is largely restricted to variation of single-valued parameters as e.g. task periods. In this paper we provide a sensitivity analysis to determine the bounds on the admissible activation pattern of a task, modelled through a minimum distance function. In an evaluation on a set of synthetic testcases we show, that the proposed algorithm provides significantly tighter bounds, than previous exact analyses, that determine allowable parametrizations of activation patterns.
Moritz Neukirchner, Sophie Quinton, Tobias Michaels, Philip Axer, Rolf Ernst
DATE5
2013 Formal analysis of sporadic bursts in real-time systems
abstract
In this paper we propose a new method for the analysis of response times in uni-processor real-time systems where task activation patterns may contain sporadic bursts. We use a burst model to calculate how often response times may exceed the worst-case response time bound obtained while ignoring bursts. This work is of particular interest to deal with dual-cyclic frames in the analysis of CAN buses. Our approach can handle arbitrary activation patterns and the static priority preemptive as well as non-preemptive scheduling policies. Experiments show the applicability and the benefits of the proposed method.
Sophie Quinton, Mircea Negrean, Rolf Ernst
DATE3
2013 Response-Time Analysis of Parallel Fork-Join Workloads with Real-Time Constraints
abstract
The advent of multi- and many-core processors comes with new challenges and opportunities for the designer of embedded real-time applications. By using parallel programming techniques (e.g. OpenMP) software engineers can leverage from the available hardware parallelism and speed up the algorithms. The inherent redundancy of multi-core architectures can also be used to implement fault-tolerance by executing code redundantly on multiple cores in parallel. Parallel programming and redundant execution are typical examples for fork-join tasks in which the program is partially parallelized. However, complex synchronization of parallel segments across multiple cores can cause unanticipated effects. This is especially problematic in hard real-time applications where data must be available in bounded time (e.g. stereo vision for pedestrian detection). The contribution of this work is a novel worst-case response time analysis which accounts for synchronization of fork-join tasks with arbitrary deadlines. We apply the analysis to the Romain framework which extends the L4 micro kernel by redundant multithreading targeted towards fault-tolerant embedded systems. By using formal analysis, we show that parallelizing workloads can lead to drastic performance impairments compared to traditional sequential execution if not done carefully.
Philip Axer, Sophie Quinton, Moritz Neukirchner, Rolf Ernst, Björn Döbel, Hermann Härtig
ECRTS4
2013 Message from the program co-chairs
abstract
Welcome to EMSOFT 2013, the 13th International Conference on Embedded Software, held in Montreal, Quebec, Canada, on September 29 – October 4, 2013.
Rolf Ernst, Oleg Sokolsky
EMSOFT1
2013 Exploration of FPGA-based dense block matching for motion estimation and stereo vision on a single chip
abstract
Camera-based systems in series vehicles have gained in importance in the past several years, which is documented, for example, by the introduction of front-view cameras and applications such as traffic sign or lane detection by all major car manufacturers. Besides a pure or enhanced visualization of the vehicle's environment, camera systems have also been extensively used for the design and implementation of complex driver assistance functions in diverse research scenarios, as they offer the possibility to extract both depth and motion information of static and moving objects. However, the evolution of existing computation-intensive vision applications from research vehicles toward series integration is currently a challenging task, which is due to the absence of highperformance computer architectures that adhere to the existing strict power and cost constraints. This paper addresses this challenge and explores FPGA-based dense block matching, which enables the calculation of depth information and motion estimation on shared hardware resources, regarding its applicability in intelligent vehicles. This includes the introduction of design scalability in time and space, thereby supporting customized application implementations and multiple camera setups. The presented modular concept also enables enhancements with pre- and post-processing features, which can be utilized to refine the obtained matching results. Its usability has been evaluated in diverse application scenarios and reaches high-performance image processing results of up to 740 GOPS at an acceptable energy level of 11 Watts, rendering it a suitable candidate for future series vehicles.
Henning Sahlbach, Rolf Ernst, Stefan Wonneberger, Thorsten Graf 0001
Intelligent Vehicles Symposium2
2013 Towards a Certifiable Integration of SRAM-Based FPGAs in Safety-Critical Automotive Systems
abstract
Advanced interconnected electronic systems play crucial roles in recent vehicle generations and have resulted in a significant increase of mileage and the introduction of several novel automotive features. For complex driver assistance applications, FPGAs have started to replace established embedded or signal processors, providing high-performance processing capabilities at modest energy consumption. However, their certification in safety-critical applications is a challenging task, which is due to their internal configuration memory-based computer architecture, requiring adapted analysis and error mitigation approaches. Using a recommended automotive safety analysis technique, this paper evaluates a generic in-vehicle FPGA-based computer platform regarding its certification limitations in automotive context. A suitable configuration memory safety concept for applications with highest safety integrity levels is then developed by combining established error mitigation mechanisms, which are also evaluated experimentally on an automotive prototyping platform. The obtained concept supports the execution of safety-critical applications on reconfigurable logic and proposes a viable certification path for automotive FPGAs considering recent safety standards.
Henning Sahlbach, Rolf Ernst
PRDC2
2013 IDAMC: A NoC for mixed criticality systems
abstract
Increasing demand for performance and further integration promotes the use of multi- and many-core systems - also in safety-critical embedded systems. In this domain, hardware platforms obviously have to support real-time, predictability constrained applications such as an anti-lock braking system. However, the on-going trend to integrate multiple functions with different criticalities (mixed critical) on a single platform calls for a paradigm shift. Mixed-critical systems require special attention with respect to functional (access protection) and non-functional (performance) isolation. An additional layer of protection and guaranteed service on the underlying infrastructure enables the efficient adoption of such architectures in safety-critical domains. In this paper, we present the IDAMC, a many-core platform which provides mechanisms to integrate applications of different criticalities on a single platform.
Sebastian Tobuschat, Philip Axer, Rolf Ernst, Jonas Diemer
RTCSA3
2013 Monitoring of Workload Arrival Functions for Mixed-Criticality Systems
abstract
Integrating applications with different safety requirements on a common platform requires either certification of all applications to the highest safety level or "sufficient independence" among them. As the former typically is too costly, isolation mechanisms, such as monitoring, are key in the design of mixed-criticality systems. We regard monitoring of activation patterns of real-time applications in mixed-criticality systems. Existing solutions monitor single tasks in isolation. We present a monitoring scheme which allows to monitor groups of tasks jointly. It allows to express correlations between activations and provides improved resource utilization as no isolation between tasks in a group is enforced.
Moritz Neukirchner, Philip Axer, Tobias Michaels, Rolf Ernst
RTSS4
2013 Compositional performance analysis with improved analysis techniques for obtaining viable end-to-end latencies in distributed embedded systems
Jonas Rox, Rolf Ernst
Int. J. Softw. Tools Technol. Transf.2
2013 MORPHEUS: A heterogeneous dynamically reconfigurable platform for designing highly complex embedded systems
abstract
Recently, system designers are facing the challenge of developing systems that have diverse features, are more complex and more powerful, with less power consumption and reduced time to market. These contradictory constraints have forced technology providers to pursue design solutions that will allow design teams to meet the above design targets. In that respect, this paper introduces an innovative technology platform, called MORPHEUS, which intents to provide complete design framework for dealing with the aforementioned challenges. MORPHEUS consists of a state of the art architecture that encompasses heterogeneous reconfigurable accelerators for implementing on the same hardware architecture applications with varying characteristics and a tool chain that, through a software oriented approach, eases the implementation of highly complex applications with heterogeneous characteristics. The proposed approach has been tested and evaluated through state of the art cases studies borrowed from complementary application domains.
Nikos S. Voros, Michael Hübner 0001, Jürgen Becker 0001, Matthias Kühnle, Florian Thoma, Arnaud Grasset, Paul Brelet, Philippe Bonnot 0001, Fabio Campi, Eberhard Schüler, Henning Sahlbach, Sean Whitty, Rolf Ernst, Enrico Billich, Claudia Tischendorf, Ulrich Heinkel, Frank Ieromnimon, Dimitrios Kritharidis, Axel Schneider, Joachim Knäblein, Wolfram Putzke-Röming
ACM Trans. Embed. Comput. Syst.13
2013 Application Space Exploration of a Heterogeneous Run-Time Configurable Digital Signal Processor
abstract
This paper describes the application space exploration of a heterogeneous digital signal processor with dynamic reconfiguration capabilities. The device is built around three reconfigurable engines featuring different flavours and computation granularities that make it suitable for a wide range of signal processing application domains such as video coding, image processing, telecommunications, and cryptography. Performance of signal processing applications is evaluated from measurements performed on a CMOS 90 nm prototype. In order to characterize the application space of the processor, performance is compared with state-of-the-art devices, taking programmability, computational capabilities, and energy efficiency as the main metrics. The device exploits performance and energy efficiency significantly more than general purpose processors, while still maintaining a user-friendly programming approach that mainly relies on software-oriented languages. The device is able to achieve 1.2 to 15 GOPS with an energy efficiency from 2 to 50 GOPS/W when running the selected applications.
Davide Rossi 0001, Claudio Mucci, Fabio Campi, Simone Spolzino, Luca Vanzolini, Henning Sahlbach, Sean Whitty, Rolf Ernst, Wolfram Putzke-Röming, Roberto Guerrieri
IEEE Trans. Very Large Scale Integr. Syst.8
2012 Probabilistic response time bound for CAN messages with arbitrary deadlines
abstract
The controller area network (CAN) is widely used in industrial and the automotive domain and in this context often for hard real-time applications. Formal methods guide the designer to give worst-case guarantees on timing. However, due to bit errors on the communication channel response times can be delayed due to retransmissions. Some methods exist to cover these effects, but are limited e.g. (support only periodic real-time traffic). In this paper we generalize existing methods to support arbitrary deadlines, and derive a probabilistic response time bound which is especially useful with the emergence of the new automotive safety standard ISO 26262.
Philip Axer, Maurice Sebastian, Rolf Ernst
DATE3
2012 Challenges and new trends in probabilistic timing analysis
abstract
Modeling and analysis of timing information are essential to the design of real-time systems. In this domain, research related to probabilistic analysis is motivated by the desire to refine results obtained using worst-case analysis for systems in which the worst-case scenario is not the only relevant one, such as soft real-time systems. This paper presents an overview of the existing solutions for probabilistic timing analysis, focusing on challenges they have to face. We discuss in particular two new trends toward Probabilistic Real-Time Calculus and Typical-Case Analysis which rise to some of these challenges.
Sophie Quinton, Rolf Ernst, Dominique Bertrand, Patrick Meumeu Yomsi
DATE2
2012 Formal analysis of sporadic overload in real-time systems
abstract
This paper presents a new compositional approach providing safe quantitative information about real-time systems. Our method is based on a new model to describe sporadic overload at the input of a system. We show how to derive from such a model safe quantitative information about the response time of each task. Experiments demonstrate the efficiency of this approach on a real-life example. In addition we improve the state of the art in compositional performance analysis by introducing execution time models which take into account several consecutive executions and by using tighter bounds for computing output event models.
Sophie Quinton, Matthias Hanke, Rolf Ernst
DATE3
2012 Using timing analysis for the design of future switched based Ethernet automotive networks
abstract
In this paper, we focus on modeling and analyzing multi-cast and broadcast traffic latencies on switch-level within an Ethernet- based communication network for automotive applications. The analysis is performed adapting existing worst/best case schedulability analysis concepts, techniques, and methods. Under our modeling assumptions, we obtain safe bounds for both the minimum (lower bound) and maximum (upper bound) latencies. The formal analysis results are validated via simulation to determine the probability distribution of the latencies (including the worst/best case ones). We also show that the bounds can be tightened under some assumptions and we sketch opportunities for future work in this area. Finally, we show how formal analysis can be used to quickly explore tradeoffs in the system configuration which delivers the required performance. All results in this work are obtained on a moderately complex yet meaningful automotive example.
Jonas Rox, Rolf Ernst, Paolo Giusto
DATE2
2012 A high-performance dense block matching solution for automotive 6D-vision
abstract
Camera-based driver assistance systems have attracted the attention of all major automotive manufacturers in the past several years and are increasingly utilized to differentiate a vendor's vehicles from its competitors. The calculation of depth information and Motion Estimation can be considered as two fundamental image processing applications in these systems, which have already been evaluated in diverse research scenarios. However, in order to push these computation-intensive features towards series integration, future in-vehicle implementations must adhere to the automotive industry's strict power consumption and cost constraints. As an answer to this challenge, this paper presents a high-performance FPGA-based dense block matching solution, which enables the calculation of both object motion and the extraction of depth information on shared hardware resources. This novel single-design approach significantly reduces the amount of logic resources required, resulting in valuable cost and power savings. The acquired sensor information can be fusioned into 3D positions with an associated 3D motion vector, which enables a robust perception of the vehicle's environment. The modular implementation offers enhanced configuration features at design and execution time and achieves up to 418 GOPS at a moderate energy consumption of 10 Watts, providing a flexible solution for a future series integration.
Henning Sahlbach, Sean Whitty, Rolf Ernst
DATE3
2012 Optimizing performance analysis for synchronous dataflow graphs with shared resources
abstract
Contemporary embedded systems, which process streaming data such as signal, audio, or video data, are an increasingly important part of our lives. Shared resources (e.g. memories) help to reduce the chip area and power consumption of these systems, saving costs in high volume consumer products. Resource sharing, however, introduces new timing interdependencies between system components, which must be analyzed to verify that the initial timing requirements of the application domain are still met. Graphs with synchronous dataflow (SDF) semantics are frequently used to model these systems. In this paper, we present a method to integrate resource sharing into SDF graphs. Using these graphs and a throughput constraint, we will derive deadlines for resource accesses and the amount of memory required for an implementation. Then we derive the resource load directly from the SDF description, and perform a formal schedulability analysis to check if the original timing constraints are still met. Finally, we perform an evaluation of our approach using an image processing application and present our results.
Daniel Thiele, Rolf Ernst
DATE2
2012 Deriving Monitoring Bounds for Distributed Real-Time Systems
abstract
Runtime controllers can be used in distributed embedded systems to throttle or stop software components and thus to limit the timing effects that applications have on each other through scheduling dependencies. Such runtime controllers require bounds on the worst-case admissible resource utilization per task to estimate and to control the worst-case interference between applications. Multi-dimensional sensitivity analysis can be used to derive efficient local controller bounds from global system constraints. In this paper we present a novel distributed algorithm to determine a multi-dimensional sensitivity bound on activation jitter which serves that purpose. Distribution makes it suitable for in-field application in modular designs, a main requirement in many industrial applications. Its properties are formally derived. Extensive experiments evaluate the solution quality and computation time.
Moritz Neukirchner, Steffen Stein, Rolf Ernst
ECRTS3
2012 Mixed critical system design and analysis
abstract
With increasing use of embedded systems in safety critical systems, architectures and design processes for safety have become a primary objective in systems design. Most such systems are also time critical leading to safety and time critical systems. Safety standards impose strong requirements on such systems challenging system performance and cost. Very often, however, only part of the functions is safety and time critical calling for a design approach that both meets the safety requirements and provides efficiency and flexibility for less critical functions. These conflicting requirements have given rise to the new research area of mixed critical system design with enormous practical relevance. The tutorial addresses key aspects of mixed critical system design.
Rolf Ernst, Alan Burns 0001, Lothar Thiele, Jimmy Le Rhun
EMSOFT1
2012 Exploring the worst-case timing of Ethernet AVB for industrial applications
abstract
Predictable and low-latency communication timing is one of the major challenges for employing Ethernet-based networks in industrial automation. The evolving Ethernet AVB standard appears to be a promising architecture, as it provides mechanisms for predictable timing with standard Ethernet hardware. However, the worst-case timing of Ethernet AVB still has to be evaluated. In this paper, we analyze the timing of Ethernet AVB using both simulation and a formal worst-case analysis based on Compositional Performance Analysis known from embedded computing systems. We investigate two industrial scenarios, a typical line topology and a more complex two-level network, and compare the results from analysis and simulation. This allows us get a good indication of the applicability of the current Ethernet AVB with respect to predictable low-latency timing in industrial automation networks. We also gain an understanding of the benefits and limitations of formal Compositional Performance Analysis compared to simulation in this context.
Jonas Diemer, Jonas Rox, Rolf Ernst, Feng Chen 0012, Karl-Theo Kremer, Kai Richter 0001
IECON3
2012 Generalized Weakly-Hard Constraints
Sophie Quinton, Rolf Ernst
ISoLA (2)2
2012 Monitoring Arbitrary Activation Patterns in Real-Time Systems
abstract
Model-based verification of timing properties has become industrial practice in design processes of safety-critical hard real-time systems. To validate the correctness of the used verification model, systems are additionally monitored during regular operation. With a growing variety of activation patterns considered in verification, some of them with infinite range capturing arbitrary activation patterns, the known approaches to monitoring, which assume periodic streams, have become inapplicable or they suffer from large overhead due to piecewise continuous time monitoring. In this paper we present a light-weight monitoring approach for arbitrary activation patterns. It profits from the discrete time property of a minimum distance event representation which is used instead of the continuous time representation used in earlier approaches. The method has a configurable constant runtime overhead in terms of memory and computation and allows conservative monitoring of a given arbitrary minimum distance function. Furthermore, we provide conditions under which the monitoring function is exact.
Moritz Neukirchner, Tobias Michaels, Philip Axer, Sophie Quinton, Rolf Ernst
RTSS5
2011 Real-time communication analysis for networks with two-stage arbitration
abstract
Current on-chip and macro networks use multi-stage arbitration schemes which independently assign different resources such as crossbar inputs and outputs to individual traffic streams. To use these networks in real-time systems, their worst-case behavior must be proved analytically in order to ensure the required timing guarantees. Current analysis approaches, however, do not capture the multi-stage arbitration accurately. In this paper, we propose an analysis that maps the multi-stage arbitration to a schedulability analysis of multiprocessors with shared resources. This allows the exploitation of knowledge about the worst-case behavior of the individual traffic streams, which is required to provide nonsymmetric guarantees. Using this scheduling analysis approach, a detailed analysis solution for a common multi-stage arbitration scheme (iSLIP) is presented. Finally, we evaluate the proposed approach experimentally and compare it to previous work.
Jonas Diemer, Jonas Rox, Mircea Negrean, Steffen Stein, Rolf Ernst
EMSOFT5
2011 Bounding mode change transition latencies for multi-mode real-time distributed applications
abstract
Predicting timing behaviour is essential for the design of embedded real-time systems that can switch between different operational modes at runtime. The settling time of a mode change, called mode change transition latency, is an important system parameter. Known approaches that address the problem of timing analysis for multi-mode real-time systems are restricted to applications without communicating tasks. Also, these assume that transitions are initiated only during a steady state, however, without indicating when a system executes in a steady state. In this paper, we present an analysis algorithm which gives a maximum bound on each mode change transition latency of multi-mode distributed applications thereby overcoming limitations of previous work. We explain the algorithm, prove its correctness, illustrate the steps and provide experimental data that show its usefulness.
Mircea Negrean, Moritz Neukirchner, Steffen Stein, Simon Schliecker, Rolf Ernst
ETFA5
2011 Utilizing Hidden Markov Models for Formal Reliability Analysis of Real-Time Communication Systems with Errors
abstract
In the near future embedded systems will be faced with the phenomena of increasing error rates, caused by a variety of error sources that have to be considered during the design process. In this paper we propose a method to derive the reliability of a real-time capable CAN bus system with errors. Individual errors on the CAN bus might be correlated in arbitrary way, the proposed algorithm will cover this. It is based on a previous work on reliability analysis that has been restricted to uncorrelated bit errors. To extend this approach we first introduce a suitable error model to describe arbitrary correlations between bit errors. As a key novelty we present an extended analysis procedure that takes this error model into account. This new approach will be utilized to determine the effects of burst errors and to demonstrate the necessity of appropriate error models for reliability analysis.
Maurice Sebastian, Philip Axer, Rolf Ernst
PRDC3
2010 Efficient throughput-guarantees for latency-sensitive networks-on-chip
abstract
Networks-on-chip (NoC) for future multi- and many-core processor platforms face an increasing diversity of traffic requirements, ranging from streaming traffic with real-time requirements to bursty best-effort. The best-effort traffic usually results from applications running on general-purpose processors with caches and is very sensitive to latency. Hence, the NoC must provide guaranteed services to some traffic streams, while maintaining low latency and high throughput of best-effort traffic. In this paper, we propose a run-time configurable NoC that enables bandwidth guarantees with minimum impact on latency for best-effort traffic. This is achieved by prioritization and distributed traffic shaping of best-effort traffic. The analysis and evaluation of our quality-of-service scheme show that it can provide tight bandwidth guarantees for streaming traffic. At the same time, the average latencies of best-effort traffic improved by up to 47% compared to a standard prioritization scheme.
Jonas Diemer, Rolf Ernst, Michael Kauschke
ASP-DAC2
2010 A software update service with self-protection capabilities
abstract
Integration of system components is a crucial challenge in the design of embedded real-time systems, as complex non-functional interdependencies may exist. We propose a software update service with self-protection capabilities against unverified system updates - thus solving the integration problem in-system. As modern embedded systems may evolve through software updates, component replacement or even self-optimization, possible system configurations are hard to predict. Thus the designer of system updates does not know the exact system configuration. This turns the proof of system feasibility into a critical challenge. This paper presents the architecture of a framework and associated protocols enabling updates in embedded systems while ensuring safe operation w.r.t. non-functional properties. The proposed process employs contract based principles at the interfaces towards applications to perform an in-system verification. Practical feasibility of our approach is demonstrated by an implementation of the update process, which is analyzed w.r.t. the memory consumption overhead and execution time.
Moritz Neukirchner, Steffen Stein, Harald Schrom, Rolf Ernst
DATE4
2010 Exploiting inter-event stream correlations between output event streams of non-preemptively scheduled tasks
abstract
In this paper we present a new technique which exploits timing-correlation between tasks for scheduling analysis in multiprocessor and distributed systems with non-preemptive scheduled resources. Previously developed techniques also allow capturing and exploiting timing-correlation in distributed systems. However, they focus on timing correlations resulting from data dependencies between tasks. The new technique presented in this paper is orthogonal to the existing ones and allows capturing timing-correlations between the output event streams of tasks resulting from the use of a non-preemptive scheduling policy on a resource. We also show how these timing-correlations can be exploited to calculate tighter bounds for the worst-case response time analysis for tasks activated by such correlated event streams.
Jonas Rox, Rolf Ernst
DATE2
2010 Bounding the shared resource load for the performance analysis of multiprocessor systems
abstract
Predicting timing behavior is key to reliable real-time system design and verification, but becomes increasingly difficult for current multiprocessor systems on chip. The integration of formerly separate functionality into a single multicore system introduces new inter-core timing dependencies, resulting from the common use of the now shared resources. In order to conservatively bound the delay due to the shared resource accesses, upper bounds on the potential amount of conflicting requests from other processors are required. This paper proposes a method that captures the request distances of multiple shared resource accesses by single tasks and also by multiple tasks that are dynamically scheduled on the same processor. Unlike previous work, we acknowledge the fact that on a single processor, tasks will not actually execute in parallel, but in alternation. This consideration leads to a more accurate load model. In a final step, the approach is extended to allow addressing also dynamic cache misses that do not occur at predefined times but surface dynamically during the execution of the tasks.
Simon Schliecker, Mircea Negrean, Rolf Ernst
DATE3
2010 Application-specific memory performance of a heterogeneous reconfigurable architecture
abstract
Heterogeneous reconfigurable processing architectures are often limited by the speed at which they can access data in external memory. Such architectures are designed for flexibility to support a broad range of target applications, including advanced algorithms with significant processing and data requirements. Clearly, strong performance of applications in this category is an extremely relevant metric for demonstrating the full performance potential of heterogeneous computing platforms. One such example, a film grain noise reduction application for high-definition video, which is composed of multiple image processing tasks, requires enormous data rates due to its large input image size and real-time processing constraints. This application is especially representative of highly parallel, heterogeneous, data-intensive programs that can properly exploit the advantages offered by computing platforms with multiple heterogeneous reconfigurable processing elements. To accomplish this task and meet the above requirements, a bandwidth-optimized external memory controller has been designed for use with a heterogeneous reconfigurable architecture and its NoC interconnect. With the help of the application described above, this paper evaluates the proposed architecture in two forms: (1) with a basic memory controller IP and (2) with the advanced memory controller design. The results illustrate the full potential of the computing platform as well as the power of heterogeneous reconfigurable computing combined with high-speed access to large external memories.
Sean Whitty, Henning Sahlbach, Brady Hurlburt, Rolf Ernst, Wolfram Putzke-Röming
DATE4
2010 A Polynomial-Time Algorithm for Computing Response Time Bounds in Static Priority Scheduling Employing Multi-linear Workload Bounds
abstract
Despite accuracy, analysis speed is sometimes a concern for the performance analysis of real-time systems, e.g. if to performed at runtime for online admission tests. As of today, several algorithms to compute an upper bound to the worst-case response time of a task scheduled under static priority preemptive scheduling with polynomial run-time have been proposed. Most approaches assume periodic activation of all tasks, some allow activation jitter. We generalize the approach to support convex activation patterns, by using multi-linear workload approximations and introduce the possibility to model processor availability to the task set under analysis.
Steffen Stein, Matthias Ivers, Jonas Diemer, Rolf Ernst
ECRTS4
2010 A Scalable, High-Performance Motion Estimation Application for a Weakly-Programmable FPGA Architecture
abstract
Computer architectures for advanced driver assistance systems have become increasingly important in the automotive industry. They target safety-critical applications, which process large amounts of incoming sensor data. This is especially the case for image processing applications, which must handle several uncompressed image streams from multiple cameras. As one possible target architecture, FPGAs provide sufficient processing power for complex applications such as lane or object detection. A particularly challenging application is the reliable detection of moving objects, which is the basis for several future driver assistance applications, such as a digital 3D reconstruction of a vehicle's surroundings. This paper presents an advanced Motion Estimation application, which achieves high performance processing of up to 449 FPS for an image resolution of 512 × 384 pixels. The implementation concept relies on weakly-programmable processing elements and a reconfigurable data path, which allows an efficient exploitation of the FPGA's chip area and clock frequencies, leading to a flexible solution that fits future application requirements.
Henning Sahlbach, Sean Whitty, Oliver Bende, Rolf Ernst
FPL4
2010 Back Suction: Service Guarantees for Latency-Sensitive On-chip Networks
abstract
Networks-on-chip for future many-core processor platforms face an increasing diversity of traffic requirements, ranging from streaming traffic with real-time requirements to bursty latency-sensitive best-effort traffic from general-purpose processors with caches. In this paper, we propose Back Suction, a novel flow-control scheme to implement quality-of-service. Traffic with service guarantees is selectively prioritized upon low buffer occupancy of downstream routers. As a result, best-effort traffic is preferred for an improved latency as long as guaranteed service traffic makes sufficient progress. We present a formal analysis and an experimental evaluation of the Back Suction scheme showing improved latency of best effort traffic when compared to current approaches even under formal service guarantees for streaming traffic.
Jonas Diemer, Rolf Ernst
NOCS2
2010 Editorial: Model-driven embedded-system design
abstract
editorial Free Access Share on Editorial: Model-driven embedded-system design Editors: Twan Basten Embedded Systems Institute, Netherlands, Eindhoven University of Technology, Netherlands Embedded Systems Institute, Netherlands, Eindhoven University of Technology, NetherlandsView Profile , Rolf Ernst Technische Universität Braunschweig, Germany Technische Universität Braunschweig, GermanyView Profile Authors Info & Claims ACM Transactions on Embedded Computing SystemsVolume 10Issue 2Article No.: 15pp 1–4https://doi.org/10.1145/1880050.1880051Published:07 January 2011Publication History 0citation572DownloadsMetricsTotal Citations0Total Downloads572Last 12 Months10Last 6 weeks4 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Twan Basten, Rolf Ernst
ACM Trans. Embed. Comput. Syst.2
2010 Real-time performance analysis of multiprocessor systems with shared memory
abstract
Predicting timing behavior is key to reliable real-time system design and verification, but becomes increasingly difficult for current multiprocessor systems on chip. The integration of formerly separate functionality into a single multicore system introduces new intercore timing dependencies resulting from the common use of the now shared resources. This feedback of system timing on local timing makes traditional performance analysis approaches inappropriate. This article presents a general methodology to model the shared resource traffic and consider its effect on the local task execution. The aggregate busy time captures the timing of multiple accesses to a shared memory far better than the traditional models that focus on the timing of individual events. An iterative approach is proposed to tackle the analysis dependencies that exist in systems with event-driven task activation and dynamic resource arbitration.
Simon Schliecker, Rolf Ernst
ACM Trans. Embed. Comput. Syst.2
2009 A link arbitration scheme for quality of service in a latency-optimized network-on-chip
abstract
Networks-on-chip (NoC) for general-purpose multiprocessors require quality of service mechanisms to allow realtime streaming applications to be executed along with latency-sensitive general purpose processing tasks. In this paper, we propose a NoC link arbitration technique that supports bandwidth guarantees along with best effort latency optimizations. In contrast to many existing quality of service mechanisms, our technique prioritizes best effort over guaranteed bandwidth traffic for optimal latency. Distributed traffic shaping is used to offer bandwidth guarantees over previously reserved connections, which are established dynamically using control messages. Initial simulation results show that our arbitration scheme can provide tight bandwidth guarantees for streaming traffic under network overload conditions. At the same time, the average latency of best effort traffic is improved compared to a simple prioritization of streaming traffic.
Jonas Diemer, Rolf Ernst
DATE2
2009 Panel session - Multicore, will Startups drive innovation?
Ahmed Amine Jerraya, Rolf Ernst
DATE2
2009 Response-time analysis of arbitrarily activated tasks in multiprocessor systems with shared resources
abstract
As multiprocessor systems are increasingly used in real-time environments, scheduling and synchronization analysis of these platforms receive growing attention. However, most known schedulability tests lack a general applicability. Common constraints are a periodic or sporadic task activation pattern, with deadlines no larger than the period, or no support for shared resource arbitration, which is frequently required for embedded systems. In this paper, we address these constraints and present a general analysis which allows the calculation of response times for fixed priority task sets with arbitrary activations and deadlines in a partitioned multiprocessor system with shared resources. Furthermore, we derive an improved bound on the blocking time in this setup for the case where the shared resources are protected according to the Multiprocessor Priority Ceiling Protocol (MPCP).
Mircea Negrean, Simon Schliecker, Rolf Ernst
DATE3
2009 Learning early-stage platform dimensioning from late-stage timing verification
abstract
Today's innovations in the automotive sector are, to a great extent, based on electronics. The increasing integration complexity and stringent cost reduction goals turn E/E platform design into a challenging task. Timing/performance is becoming a key aspect of architecture design, because the platform must be dimensioned to provide just the right amount of computing power and network bandwidth, including reserves for future extensions, in order to be cost efficient. In other words, it must be as powerful as needed but as cheap as possible. Finding this sweet spot is a key challenge. Therefore, OEMs and Tier-1 are in search of new methods, processes, and timing analysis techniques that assist in early platform design stages. In this paper, we demonstrate how some selected techniques that are established for verification (in late design stages) can also be used to guide the design (in early stages). We present examples in the areas ECU (OSEK), buses (CAN, FlexRay) and gated networks. Flow and applicability aspects are highlighted. As a key result, we show that and how we can learn from late-stage verification for early-stage design. Finally, we also outline future challenges in the area of multi-core ECUs.
Kai Richter 0001, Marek Jersak, Rolf Ernst
DATE3
2009 Mapping of a film grain removal algorithm to a heterogeneous reconfigurable architecture
abstract
Despite recent advances in FPGA, GPU, and general purpose processor technologies, the challenges posed by real-time digital image processing at high resolutions cannot be fully overcome due to insufficient processing capability, inadequate data transport and control mechanisms, and often prohibitively high costs. To address these issues, we proposed a two-phase solution for a real-time film grain noise reduction application. The first phase is based on a state-of-the-art FPGA platform used as a reference design. The second phase is based on a novel heterogeneous reconfigurable computing platform that offers flexibility not available from other computing paradigms. This paper introduces the heterogeneous platform and briefly reviews our previous work with the application in question, as well as its implementation on the FPGA demonstration board during the first phase. Then we present a decomposition of the application, which allows an efficient mapping to the new heterogeneous computing platform through the use of its diverse reconfigurable computing units and run-time reconfiguration.
Sean Whitty, Henning Sahlbach, Rolf Ernst, Wolfram Putzke-Röming
DATE3
2009 Reliability Analysis of Single Bus Communication with Real-Time Requirements
abstract
Due to continuous technology downscaling modern embedded real-time systems become more and more susceptible to the occurrence of errors. The usage of appropriate countermeasures is necessary to prevent a system failure. In this paper we present a new reliability estimation technique for such systems. As a key novelty a formal analysis method will be introduced that approximates the probability of failure of a priority driven bus during a period of time, enabling fast and accurate reliability calculation. It removes the major drawbacks of existing approaches, e.g. random-based Monte-Carlo simulation that requires long runtimes. However Monte-Carlo simulation serves as reference method to demonstrate the accuracy of our approach by comparing analysis and simulation results. Finally we consider the design of mixed-criticality systems which combine different safety requirements on a single component.
Maurice Sebastian, Rolf Ernst
PRDC2
2009 System Level Performance Analysis for Real-Time Automotive Multicore and Network Architectures
abstract
Software timing aspects have only recently received broad attention in the automotive industry. New design trends and the ongoing work in the AUTOSAR (Automotive Open System Architecture) partnership have significantly increased the industry's awareness to these issues. Now, timing is recognized as a major challenge and has been put explicitly on the agenda of AUTOSAR and other industry-driven research projects. The goals include complementing the existing standard by a timing view and adding methodological steps, if necessary. Clearly, establishing such timing models requires knowing well the implications of modern architectures and topologies. In this paper, we survey existing performance analysis approaches from real-time systems research and compare them to the established layered software architectures of automotive system design. We highlight key challenges for the application of performance analysis in this domain and identify structural as well as behavioral ldquomodeling gapsrdquo. While structural gaps can be overcome by model transformations, behavioral gaps require real extensions to known analyses. We discuss two such extensions in detail, namely, the use of hierarchical event models and the specialties of timing analysis for multicore platforms. This paper concludes with an overview over qualitative comparisons of analysis techniques, both technically and concerning their industrial applicability.
Simon Schliecker, Jonas Rox, Mircea Negrean, Kai Richter 0001, Marek Jersak, Rolf Ernst
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2009 Application development with the FlexWAFE real-time stream processing architecture for FPGAs
abstract
The challenges posed by complex real-time digital image processing at high resolutions cannot be met by current state-of-the-art general-purpose or DSP processors, due to the lack of processing power. On the other hand, large arrays of FPGA-based accelerators are too inefficient to cover the needs of cost sensitive professional markets. We present a new architecture composed of a network of configurable flexible weakly programmable processing elements, Flexible Weakly programmable Advanced Film Engine (FlexWAFE). This architecture delivers both programmability and high efficiency when implemented on an FPGA basis. We demonstrate these claims using a professional next-generation noise reducer with more than 170G image operations/s at 80% FPGA area utilization on four Virtex II-Pro FPGAs. This article will focus on the FlexWAFE architecture principle and implementation on a PCI-Express board.
Amilcar do Carmo Lucas, Henning Sahlbach, Sean Whitty, Sven Heithecker, Rolf Ernst
ACM Trans. Embed. Comput. Syst.5
2009 Response Time Analysis in Multicore ECUs with Shared Resources
abstract
As multiprocessor systems are increasingly used in automotive real-time environments, scheduling and synchronization analysis of these platforms receive growing attention. Upcoming multicore ECUs allow the integration of previously separated functionality for body electronics or sensor fusion onto a single unit, and allow the parallelization of complex computations over multiple cores. The application of multiple CPUs turns an ECU into a highly integrated ldquonetworked systemrdquo microcosm, in which complex interdependencies can be observed due to the use of shared resources even in partitioned scheduling. To deliver predictable performance, resource arbitration protocols are required and have been proposed in literature. This paper presents an novel analytical approach to provide the worst-case response time for real-time tasks in multiprocessor systems with shared resources. The method supports realistic, event- or time-driven task activation schemes and allows to calculate tight bounds on the estimated system performance.
Simon Schliecker, Mircea Negrean, Rolf Ernst
IEEE Trans. Ind. Informatics3
2008 Distributed Performance Control in Organic Embedded Systems
Steffen Stein, Rolf Ernst
ATC2
2008 Formal Methods in System and MpSoC Performance Analysis and Optimisation
Rolf Ernst, Marek Jersak, Hans Sarnowski, Marco Bekooij, Samarjit Chakraborty
DATE1
2008 Methods, Tools and Standards for the Analysis, Evaluation and Design of Modern Automotive Architectures
abstract
Automotive systems are increasingly distributed and complex. Reduced time-to-market, cost and safety concerns require advance validation of the integrated systems and its components, from the functional, timing, and reliability standpoints. In particular, function correctness and performance may depend on communication and computation delays imposed by the selected architecture platform. Hence, the need for methods and tools capable of predicting the system-level timing behaviour (latencies and jitter), resulting from the HW platform selection, the synchronization between tasks and messages, and also from the synchronization and queuing policies of the middleware and RTOS levels. In this paper, we review methods and tools for the evaluation of the function performance and its timing correctness by simulation or by worst case static analysis.
E. Frank, Reinhard Wilhelm, Rolf Ernst, Alberto L. Sangiovanni-Vincentelli, Marco Di Natale
DATE3
2008 Modeling Event Stream Hierarchies with Hierarchical Event Models
abstract
Compositional scheduling analysis couples local scheduling analysis via event streams. While local analysis has successfully been extended to include hierarchical scheduling strategies, event streams are still flat. In this paper, we generalize the concept of a stream hierarchy to embed different types of streams in a higher level structure. We explain why this extension is a natural match to model streams generated by communication stacks that are ubiquitous in networked embedded systems. We formally define the hierarchical event model and give operations to encode, combine, and extract stream properties that can be used in flat or hierarchical local scheduling analysis. Finally, we give an example and demonstrate that the proposed model enables superior analysis results.
Jonas Rox, Rolf Ernst
DATE2
2008 Construction and Deconstruction of Hierarchical Event Streams with Multiple Hierarchical Layers
abstract
Compositional scheduling analysis couples local scheduling analysis via event streams. While local analysis has successfully been extended to include hierarchical scheduling strategies, event streams are still flat. In this paper, we formally define hierarchical event streams, which cannot only be constructed from flat event streams, but also from hierarchical streams allowing event streams with multiple hierarchical layers. We define an hierarchical event model and the operations to construct and deconstruct hierarchical events streams. Finally, we demonstrate how the model can be integrated in an existing analysis approach for distributed systems, enabling superior analysis results.
Jonas Rox, Rolf Ernst
ECRTS2
2008 Modelling and designing reliable on-chip-communication devices in MPSoCs with real-time requirements
abstract
Due to continuous technology downscaling modern MPSoCs become more and more susceptible to the occurrence of internal errors in computational cores as well as in the on-chip-communication infrastructure. The usage of appropriate techniques is necessary to counteract these errors and thus preventing them from originating a system failure. In this paper we will explore the impact of fault tolerance mechanisms for on-chip communication components in real-time systems. Therefore we will introduce a behavioural model of on-chip communication including a simple simulation framework that can easily be adapted to existing system-on-chip bus architectures. Based on that model several simulations will be performed to determine the reliability of an exemplary on-chip-bus. Our experimental results show that design decisions concerning fault tolerance strongly rely on platform and application characteristics like transmission speed or communication amount.
Maurice Sebastian, Rolf Ernst
ETFA2
2008 A bandwidth optimized SDRAM controller for the MORPHEUS reconfigurable architecture
abstract
High-end applications designed for the MORPHEUS computing platform require a massive amount of memory and memory throughput to fully demonstrate MORPHEUS's potential as a high-performance reconfigurable architecture. For example, a proposed film grain noise reduction application for high definition video, which is composed of multiple image processing tasks, requires enormous data rates due to its large input image size and real-time processing constraints. To meet these requirements and to eliminate external memory bottlenecks, a bandwidth- optimized DDR-SDRAM memory controller has been designed for use with the MORPHEUS platform and its Network On Chip interconnect. This paper describes the controller's design requirements and architecture, including the interface to the Network On Chip and the two-stage memory access scheduler, and presents relevant experiments and performance figures.
Sean Whitty, Rolf Ernst
IPDPS2
2008 Sensitivity analysis of complex embedded real-time systems
Razvan Racu, Arne Hamann 0001, Rolf Ernst
Real Time Syst.3
2007 FlexWAFE - A High-end Real-Time Stream Processing Library for FPGAs
abstract
Digital film processing is characterized by a resolution of at least 2K (2048x1536 pixels per frame at 30 bit/pixel and 24 pictures/s, data rate of 2.2 Gbit/s); higher resolutions of 4K (8.8 Gbit/s) and even 8K (35.2 Gbit/s) are on their way. Real-time processing at this data rate is beyond the scope of today's standard and DSP processors, and ASICs are not economically viable due to the small market volume.
Amilcar do Carmo Lucas, Sven Heithecker, Rolf Ernst
DAC3
2007 Automotive Software Integration
abstract
A growing number of networked applications is implemented on increasingly complex automotive platforms with several bus standards and gateways. Together, they challenge the automotive design process. Recent automotive software standards, in particular AUTOSAR that defines a network runtime environment on top of the existing automotive standards, are intended to improve portability and interoperability. AUTOSAR shall replace or extend earlier proprietary software architecture solutions, but it does not yet sufficiently address time and platform modeling and specification. The presentation will give some examples for open issues with respect to performance, timing and interoperability. It will show how recent results in compositional performance analysis can be exploited to analyze such networked systems, and how to apply design space exploration in a complex automotive supply chain. The resulting tools and methods can even be used to optimize the robustness of an architecture which is important to handle updates and extend the lifetime of an architecture
Razvan Racu, Arne Hamann 0001, Rolf Ernst, Kai Richter 0001
DAC3
2007 Performance analysis of complex systems by integration of dataflow graphs and compositional performance analysis
abstract
In this paper we integrate two established approaches to formal multiprocessor performance analysis, namely synchronous dataflow graphs and compositional performance analysis. Both make different trade-offs between precision and applicability. We show how the strengths of both can be combined to achieve a very precise and adaptive model. We couple these models of completely different paradigms by relying on load descriptions of event streams. The results show a superior performance analysis quality
Simon Schliecker, Steffen Stein, Rolf Ernst
DATE3
2007 Methods for multi-dimensional robustness optimization in complex embedded systems
abstract
Design space exploration of embedded systems typically focuses on classical design goals such as cost, timing, buffer sizes, and power consumption. Robustness criteria, i.e. sensitivity of the system to variations of properties like execution and transmission delays, input data rates, CPU clock rates, etc., has found less attention despite its practical relevance.
Arne Hamann 0001, Razvan Racu, Rolf Ernst
EMSOFT3
2007 Influence of different system abstractions on the performance analysis of distributed real-time systems
abstract
System level performance analysis plays a fundamental role in the design process of real-time embedded systems. Several different approaches have been presented so far to address the problem of accurate performance analysis of distributed embedded systems in early design stages. The existing formal analysis methods are based on essentially different concepts of abstraction. However, the influence of these different models on the accuracy of the system analysis is widely unknown, as a direct comparison of performance analysis methods has not been considered so far. We define a set of benchmarks aimed at the evaluation of performance analysis techniques for distributed systems. We apply different analysis methods to the benchmarks and compare the results obtained in terms of accuracy and analysis times, highlighting the specific effects of the various abstractions. We also point out several pitfalls for the analysis accuracy of single approaches and investigate the reasons for pessimistic performance predictions.
Simon Perathoner, Ernesto Wandeler, Lothar Thiele, Arne Hamann 0001, Simon Schliecker, Rafik Henia, Razvan Racu, Rolf Ernst, Michael González Harbour
EMSOFT8
2007 Efficient priority optimization in complex distributed embedded systems through search space adaptation
abstract
In this paper we present a framework for dynamic search space adaptation during evolutionary design space exploration. Compared to previous approaches our framework is capable of adapting the search space dynamically during exploration leading to better search space exploitation in the same exploration time. The application of our framework to priority optimization in complex distributed embedded systems shows that dynamic search space adaptation can significantly increase exploration efficiency, both in terms of exploration time and quality of achieved results.
Arne Hamann 0001, Rolf Ernst
GECCO2
2007 Improved Output Jitter Calculation for Compositional Performance Analysis of Distributed Systems
abstract
Compositional performance analysis iteratively alternates local scheduling analysis techniques and output event model propagation between system components to enable performance analysis of heterogeneous distributed systems. In spite of its high scalability and adaptability, the compositional approach may suffer from overestimated results compared with other system performance verification techniques. The main reason is an incomplete consideration of event sequence correlations. In this paper we present a new technique that improves the output jitter calculation by correlating jitter and response times and offers significantly tighter analysis bounds.
Rafik Henia, Razvan Racu, Rolf Ernst
IPDPS3
2007 Multi-dimensional Robustness Optimization in Heterogeneous Distributed Embedded Systems
abstract
Embedded system optimization typically considers objectives such as cost, timing, buffer sizes, and power consumption. Robustness criteria, i.e. sensitivity of the system to property variations like execution and transmission delays, input data rates, CPU clock rates, etc., has found less attention despite its practical relevance. In this paper we present an approach for optimizing multidimensional robustness criteria in complex distributed embedded systems. The key novelty of our approach is a scalable stochastic multi-dimensional sensitivity analysis technique approximating the sought-after sensitivity front from two sides, i.e. coming from the space of working and from the space of non-working system property combinations. We utilize the proposed stochastic sensitivity analysis to derive multi-dimensional robustness metrics, which are capable of bounding the robustness of given system configurations with little computational effort. The proposed metrics can significantly speed up multidimensional robustness optimization by quickly identifying promising system configurations, whose in-depth robustness evaluation can be performed subsequently to the optimization process
Arne Hamann 0001, Razvan Racu, Rolf Ernst
IEEE Real-Time and Embedded Technology and Applications Symposium3
2007 Scenario Aware Analysis for Complex Event Models and Distributed Systems
abstract
The set of executing tasks in modern hard real time systems may change during system execution. This change, called scenario change, may lead to a transient overload situation due to the interference of different task set executions, thus necessitating timing requirement verification. Previously developed approaches analyzing response times across scenario changes are limited to strict periodic task event models and restricted to uniprocessor systems, while existing methods adapted for the analysis of distributed systems are not suitable for the analysis across scenario changes. In this paper, we eliminate the restrictions concerning task event models and present a scheduling analysis methodology allowing response time calculation across a scenario change for multi-scenario distributed systems.
Rafik Henia, Rolf Ernst
RTSS2
2007 Scalable precision cache analysis for real-time software
abstract
Caches are needed to increase the processor performance, but the temporal behavior is difficult to predict, especially in embedded systems with preemptive scheduling. Current approaches use simplified assumptions or propose complex analysis algorithms to bound the cache-related preemption delay. In this paper, a scalable preemption delay analysis for associative instruction caches to control the analysis precision and the time-complexity is proposed. An accurate preemption delay calculation is integrated into a cache-aware schedulability analysis. The framework is evaluated in several experiments.
Jan Staschulat, Rolf Ernst
ACM Trans. Embed. Comput. Syst.2
2006 Methods for power optimization in distributed embedded systems with real-time requirements
abstract
Dynamic voltagescaling and sleep state control have been shown to be extremely effective in reducing energy consumption in CMOS circuits. Though plenty of research papers have studied the application of these techniques in real-time embedded system design through intelligent task and/or voltage scheduling, most of these results are limited to relatively simple real-time application models. In this paper, a comprehensive real-time application model including periodic, sporadic and bursty tasks as well as distributed real-time constraints such as end-to-end delays is considered. Two methods are presented for reducing energy consumption while satisfying complex real-time constraints for this model. Experimental results show that the methods achieve significant energy savings without violating any deadlines.
Razvan Racu, Arne Hamann 0001, Rolf Ernst, Bren Mochocki, Xiaobo Sharon Hu
CASES3
2006 Improved offset-analysis using multiple timing-references
abstract
In this paper, we present an extension to existing approaches that capture and exploit timing-correlation between tasks for scheduling analysis in distributed systems. Previous approaches consider a unique timing-reference for each set of time-correlated tasks and thus, do not capture the complete timing-correlation between task activations. Our approach is to consider multiple timing-references which allow us to capture more information about the timing-correlation between tasks. We also present an algorithm that exploits the captured information to calculate tighter bounds for the worst-case response time analysis under a static priority preemptive scheduler.
Rafik Henia, Rolf Ernst
DATE2
2006 A reconfigurable HW/SW platform for computation intensive high-resolution real-time digital film applications
abstract
This paper presents a multi-board, multi-FPGA hardware/software architecture, for computation intensive, high resolution (2048times2048pixels), real-time (24 frames per second) digital film processing. It is based on Xilinx Virtex-II Pro FPGAs, large SDRAM memories for multiple frame storage and a PCI express communication network. The architecture reaches record performance running a complex noise reduction algorithm including a 2.5 dimensions DWT and a full 16times16 motion estimation at 24 fps requiring a total of 203 Gops/s net computing performance and a total of 28 Gbit/s DDR-SDRAM frame memory bandwidth. To increase design productivity and yet achieve high clock rates (125MHz), the architecture combines macro component configuration and macro level floorplanning with weak programmability using distributed microcoding. As an example, the core of the bidirectional motion estimation using 2772 CLBs reaching 155 Gop/s (1538 op/pixel) requiring 7 Gbit/s external memory bandwidth was developed in two men-months
Amilcar do Carmo Lucas, Sven Heithecker, Peter Rüffer, Rolf Ernst, Holger Rückert, Gerhard Wischermann, Karin Gebel, Reinhard Fach, Wolfgang Huther, Stefan Eichner, Gunter Scheller
DATE4
2006 A Formal Approach to Multi-Dimensional Sensitivity Analysis of Embedded Real-Time Systems
abstract
System robustness is a major concern in the design of efficient and reliable state-of-the-art heterogenous embedded real-time systems. Due to complex component interactions, resource sharing and functional dependencies, one-dimensional sensitivity analysis cannot cover all effects that modifications of one system property may have on system performance. One reason is that the variation of one property can also affect the values of other system properties requiring new approaches to keep track of simultaneous parameter changes. In this paper we present a heuristic and a stochastic approach suited for the multi-dimensional sensitivity analysis of large heterogenous embedded systems with complex timing constraints. 1
Razvan Racu, Arne Hamann 0001, Rolf Ernst
ECRTS3
2006 Worst case timing analysis of input dependent data cache behavior
abstract
Data caches significantly reduce the average memory access time and are necessary for an efficient design. Due to its direct dependency on input data is difficult to predict the worst case timing behavior, which is crucial for a reliable system. While simulation is too time-consuming, current worst case execution time approaches focus on instruction caches only. Current approaches to data cache analysis restrict cache behavior to predictable data accesses or classify input dependent memory accesses as non-cache able. In this paper we propose a worst case timing analysis for direct mapped data caches that classifies memory accesses as predictable or unpredictable. For unpredictable memory accesses, a novel analysis framework is proposed that tightly bounds the impact on the existing cache contents as well as cache behavior of unpredictable memory accesses themselves. For predictable memory accesses, we use a local cache simulation and dataflow techniques. Furthermore, we describe an implementation of the analysis framework. Several experiments demonstrate its applicability. The approach targets real-time software verification but is also useful for design space exploration.
Jan Staschulat, Rolf Ernst
ECRTS2
2006 Cost-Efficient Worst-Case Execution Time Analysis in Industrial Practice
abstract
To guarantee real-time behavior of an embedded application, a schedulability analysis can be used. Such ananalysis requires the worst case execution time (WCET) of the application. While several academic approaches to conservatively bound the WCET have been proposed in the last decade, common practice in industry remains simulation and software tests. One reason is that industrial requirements are not sufficiently addressed by academic approaches. In this paper we identify important industrial requirements for WCET-analysis tools. Then, we describe the methodolgy of a previously developed WCET-analysis approach and revise important aspects of its methodology and its implementation to address key industrial requirements. In a large-scale case study the WCET-analysis tool is applied to a safety-critical automotive control application to evaluate the applicability of the tool. Furthermore, the faced challenges and the re-targenting costs for a new processor are discussed.
Jan Staschulat, Jörn-Christian Braam, Rolf Ernst, Thomas Rambow, Rainer Schlör, Rainer Busch
ISoLA3
2006 Real-Time Property Verification in Organic Computing Systems
abstract
Integrating new functionality into complex embedded hard real-time systems requires considerable engineering effort. Emerging formal analysis methodologies and tools from real-time research assist system engineers solving this integration problem. For future organic computer systems, however, it is desirable to integrate these approaches into running systems, enabling them to autonomously perform e.g. online acceptance tests and self-optimization in case of system or environmental changes. This results in high system robustness and extensibility without explicit engineering effort. In this paper, we present an approach adapting formal compositional analysis techniques to realize self-awareness and self-adaptation in embedded systems with respect to real-time properties such as latency constraints, buffer sizes, etc. We introduce a framework for distributed online performance analysis running on embedded real-time systems. Based on this framework we implement an acceptance test for the integration of new functionality into an existing embedded real-time system. Furthermore, we present an online optimization algorithm based on the same framework. In a case study, we demonstrate the applicability of the approach and show that online optimization can increase the acceptance rate with reasonable computational effort.
Steffen Stein, Arne Hamann 0001, Rolf Ernst
ISoLA3
2006 A framework for modular analysis and exploration of heterogeneous embedded systems
Arne Hamann 0001, Marek Jersak, Kai Richter 0001, Rolf Ernst
Real Time Syst.4
2005 An Image Processor for Digital Film
abstract
This paper presents an FPGA based hardware architecture, named FlexWAFE, for high resolution, high troughput real-time digital film processing. Complex algorithms require several hundred arithmetic operations per pixel which is beyond the scope of current DSP processors. To simplify programming and yet achieve high clock rates, the architecture combines component configuration with weak programmability. It alleviates the memory bottleneck by an efficient use of internal memory blocks and a multi-stream SDRAM memory scheduler with tightly bounded latency. Several examples of a discrete wavelet transform and a complex noise reducer demonstrate the architecture efficiency.
Amilcar do Carmo Lucas, Rolf Ernst
ASAP2
2005 Traffic shaping for an FPGA based SDRAM controller with complex QoS requirements
abstract
Today high-end video and multimedia processing applications require huge amounts of memory. For cost reasons, the usage of conventional dynamic RAM (SDRAM) is preferred. However, SDRAM access optimization is a complex task, especially if multistream access with different QoS requirements is involved. In [8], a multi-stream DDR-SDRAM controller IP covering combinations of low latency requirements for processor cache access, hard realtime constraints for periodic video signals and hard real-time bursty accesses for video coprocessors was described. To handle these contradictory QoS requirements at high system performance, a combination of a 2-stage scheduling algorithm and static priorities were used. This paper describes an additional flow control which enhances the overall performance. Experiments with an FPGA based high-end video platform demonstrate the superiority of this architecture.
Sven Heithecker, Rolf Ernst
DAC2
2005 TDMA Time Slot and Turn Optimization with Evolutionary Search Techniques
abstract
In this paper we present arithmetic real-coded variation operators tailored for time slot and turn optimization on TDMA-scheduled resources with evolutionary algorithms. Our operators implement an heuristic strategy to converge towards the solution space and are able to escape local minima. Furthermore, we explicitly separate the variation of the admitted loads and the turn-length in order to give the designer increased control over the optimization process. Experimental results show that our variation operators have advantages over string-coded binary variation operators which are frequently used to solve continuous optimization problems.
Arne Hamann 0001, Rolf Ernst
DATE2
2005 Context-Aware Scheduling Analysis of Distributed Systems with Tree-Shaped Task-Dependencies
abstract
In this paper, we present a new technique which exploits timing-correlation between tasks for scheduling analysis in multiprocessor and distributed systems with tree-shaped task-dependencies. Previously developed techniques also allow capturing and exploiting timing-correlation in distributed systems. However they are only suitable for linear systems, where tasks cannot trigger more than one succeeding task. The new technique presented in this paper allows capturing timing-correlation between tasks in parallel paths in a more accurate way, enabling its exploitation to calculate tighter bounds for the worst-case response time analysis for tasks scheduled under a static priority preemptive scheduler.
Rafik Henia, Rolf Ernst
DATE2
2005 Introducing Flexible Quantity Contracts into Distributed SoC and Embedded System Design Processes
abstract
Increasing design complexity eventually leads to a design process that is distributed over several companies. This is already found in the automotive industry but SoC design appears to move in the same direction. Design processes for complex systems are iterative, but iteration hardly reaches beyond company borders. Iterations require availability of preliminary design data and estimations, but due to cost and liability issues suppliers often hesitate to provide such preliminary data. Moreover, companies are rarely able to judge the accuracy and precision of externally estimated data. So, the systems integrator experiences increased design risk. Particular mechanisms are needed to ensure, that the integrated system will meet the overall requirements even if part of the early estimations are wrong or imprecise. Based on work in supply chain management, we propose an inter-company design process that is based on formal techniques from real-time systems engineering and so called flexible quantity contracts. In this process, formal techniques control design risk and flexible contracts regulate cooperation and cost distribution. The process effectively delays the design freeze point beyond the contract conclusion to enable design iterations. We explain the process and give an example.
Judita Kruse, Clive Thomsen, Rolf Ernst, Thomas Volling, Thomas Spengler
DATE3
2005 Context Sensitive Performance Analysis of Automotive Applications
abstract
Accurate timing analysis is key to efficient embedded system synthesis and integration. While industrial control software systems are developed using graphical models, such as Matlab/Simulink or ASCET/SD, exhaustive simulation is not suitable for verifying functional and timing behavior. Formal performance analysis is an alternative, but can lead to wide timing intervals because of input data dependency and complex target architectures. Hence, a designer might want to restrict the formal performance analysis to parts of the software system, called context or process modes. We describe how to define and characterize such context information from graphical models. Further, we extend the formal performance analysis to consider contexts. Results front an automotive application demonstrate the applicability of our approach.
Jan Staschulat, Rolf Ernst, Andreas Schulze, Fabian Wolf
DATE2
2005 Scheduling Analysis of Real-Time Systems with Precise Modeling of Cache Related Preemption Delay
abstract
Accurate timing analysis is key to efficient embedded system synthesis and integration. Caches are needed to increase the processor performance but they are hard to use because of their complex behaviour especially in preemptive scheduling. Current approaches use simplified assumptions or propose exponentially complex scheduling analysis algorithms to bound the cache related preemption delay at a context switch. We present a conservative polynomial algorithm that extends real-time scheduling analysis to consider cache effects due to the preempted and the preempting task for the preemption delay. Dataflow analysis on task level is combined with real-time scheduling analysis to determine the response time including cache related preemption delay for each task accurately. The experiments show significant improvement in analysis precision over previous polynomial approaches for typical embedded benchmarks.
Jan Staschulat, Simon Schliecker, Rolf Ernst
ECRTS3
2005 An FPGA based SDRAM controller with complex QoS scheduling and traffic shaping (abstract only)
abstract
Today high-end video and multimedia processing applications require huge amounts of memory. For cost reasons, the usage of conventional dynamic RAM (SDRAM) is preferred. However, SDRAM access optimization is a complex task, especially if multi-stream access with different QoS (Quality of Service) requirements is involved. At SIPS 2003 conference, we presented a multi-stream DDR-SDRAM controller IP covering combinations of low latency requirements for processor cache access, hard real-time constraints for periodic video signals and hard real-time bursty accesses for video coprocessors. To handle these contradictory QoS requirements at high system performance, a combination of an 2-stage scheduling algorithm and static priorities was used. This poster describes an additional flow control which greatly enhances the overall performance and controlability. The efficient but simple controller design makes the controller well suited for FPGA based designs. Experiments with our FPGA based high-end video platform demonstrate the superiority of this architecture.
Sven Heithecker, Rolf Ernst
FPGA2
2005 Dynamic voltage scaling for the schedulability of jitter-constrained real-time embedded systems
abstract
Jitter is a critical problem for the design of both distributed embedded systems and real-time control systems. This work considers meeting the completion jitter constraints of a set of independent, periodic, hard real-time tasks scheduled according to a preemptive fixed-priority scheme. Control over completion jitter is achieved by judiciously applying dynamic voltage scaling (DVS). Through simulation, the proposed method is shown to be an effective tool to meet jitter constraints on a variety of systems.
Bren Mochocki, Razvan Racu, Rolf Ernst
ICCAD3
2005 Scalable precision cache analysis for preemptive scheduling
abstract
Accurate timing analysis is key to efficient embedded system synthesis and integration. Caches are needed to increase the processor performance but they are hard to use because of their complex behavior especially in preemptive scheduling. Current approaches use simplified assumptions or propose exponentially complex analysis algorithms to bound the cache related preemption delay at a context switch. Existing approaches consider only direct mapped caches or propose non conservative approximation for set associative caches.In this paper we propose a novel cache related preemption delay analysis for set-associative instruction caches where the designer can adjust the analysis precision by scaling the problem complexity. Furthermore, this precise preemption delay analysis is integrated into a scheduling analysis to determine the response time of tasks accurately. In experiments we evaluate this tradeoff between analysis precision and analysis time. The results show an improvement of 22%-71% in analysis precision of cache related preemption delay and 5%-21% in response time analysis compared to previous conservative approaches.
Jan Staschulat, Rolf Ernst
LCTES2
2005 Applying Sensitivity Analysis in Real-Time Distributed Systems
abstract
During real-world design of embedded real-time systems, it cannot be expected that all performance data required for scheduling analysis is fully available up front. In such situations, sensitivity analysis is a promising approach to deal with uncertainties that result from incomplete specifications, early performance estimates, late feature requests, and so on. Sensitivity analysis allows the system designer to keep track of the flexibility of the system, and thus to quickly assess the impact of changes of individual hardware and software components on system performance. In this paper we integrate sensitivity analysis into our system-level performance analysis framework SymTA/S and show its benefits during the design of complex, networked multiprocessor embedded real-time systems.
Razvan Racu, Marek Jersak, Rolf Ernst
IEEE Real-Time and Embedded Technology and Applications Symposium3
2004 Context-Aware Performance Analysis for Efficient Embedded System Design
abstract
Performance analysis has many advantages in theory compared to simulation for the validation of complex embedded systems, but is rarely used in practice. To make analysis more attractive, it is critical to calculate tight analysis bounds. This paper shows that advanced performance analysis techniques taking correlations between successive computation or communication requests as well a correlated load distribution into account can yield much tighter analysis bounds. Cases where such correlations have a large impact on system timing are especially difficult to simulate and, hence, are an ideal target for formal performance analysis.
Marek Jersak, Rafik Henia, Rolf Ernst
DATE3
2004 Multiple process execution in cache related preemption delay analysis
abstract
Cache prediction for preemptive scheduling is an open issue despite its practical importance. First analysis approaches use simplified models for cache behavior or they assume simplified preemption and execution scenarios that seriously impact analysis precision. We present an analysis approach which considers multiple executions of processes and preemption scenarios for static priority periodic scheduling. The results of our experiments show that caches introduce a strong and complex timing dependency between process executions that are not appropriately captured in the simplified models.
Jan Staschulat, Rolf Ernst
EMSOFT2
2004 Design Space Exploration and System Optimization with SymTA/S-Symbolic Timing Analysis for Systems
abstract
The increasing complexity of heterogeneous SoC and distributed systems confronts the system designer with problems how to determine reasonable design alternatives leading to well functioning systems. Ideally, a designer would try all possible system configuration and choose the best one regarding specific system requirements. Unfortunately, such an approach is not possible because the high number of design parameters in complex systems leads to a very large design-space, prohibiting an exhaustive search. Consequently, good search techniques are needed to find optimal, or at least good, design alternatives. In this paper, we present a design space exploration framework for system optimization using SymTA/S, a software tool for formal performance analysis. In contrast to many previous approaches, our approach takes the hierarchical structure of the design space of heterogeneous SoC and distributed systems into account, allowing the designer to control the exploration process. A main technique in our approach is systematic system optimization using traffic shaping.
Arne Hamann 0001, Marek Jersak, Kai Richter 0001, Rolf Ernst
RTSS4
2003 Enabling scheduling analysis of heterogeneous systems with multi-rate data dependencies and rate intervals
abstract
Formal methods are growing in importance for performance analysis of real-time systems, but embedded system heterogeneity limits the application of these methods to subsystems or special cases. One of the problems is the rich variety of interactions between embedded system processes, which cannot be directly expressed with the typical event models used in real-time analysis.This paper shows how to transform complex interaction patterns into the integral representation of minimum and maximum arrival curves, and then to conservatively approximate these arrival curves using standard event models. This approach paves the way to apply the formal approaches known from real-time analysis to heterogeneous embedded systems.
Marek Jersak, Rolf Ernst
DAC2
2003 Formal Methods for Integration of Automotive Software
abstract
Novel functionality, configurability and higher efficiency in automotive systems require sophisticated embedded software as well as distributed software development between manufacturers and control unit suppliers. However, at least for engine control units (ECU), there exists today no well-defined software integration process that satisfies all key requirements of automotive manufacturers. We propose a methodology for safe integration of automotive software functions where required performance information is exchanged while each partner's IP is protected. We claim that, in principle, performance requirements and constraints (timing, memory consumption) for each software component and for the complete ECU can be formally validated, and believe that ultimately such formal analysis will be required for legal certification of an ECU.
Marek Jersak, Kai Richter 0001, Rolf Ernst, Jörn-Christian Braam, Zheng-Yu Jiang, Fabian Wolf
DATE3
2003 Safe Automotive Software Development
Ken Tindell, Hermann Kopetz, Fabian Wolf, Rolf Ernst
DATE4
2003 Scheduling Analysis Integration for Heterogeneous Multiprocessor SoC
abstract
Today, only very few techniques out of the host of work on formal performance and timing analysis have been adopted in MpSoC (multiprocessor system-on-chip) design. One of the key reasons is a mismatch between the scheduling models assumed in most formal approaches and the heterogeneous world of MpSoC scheduling techniques and communication patterns. This heterogeneity results from IP reuse and a plug-and-play design style, required to effectively reach the necessary design productivity. A second problem is the model complexity. While complex, specialized models can find their way into industry niches, their broad acceptance is extremely doubtful. In this paper, we review the existing scheduling analysis techniques with respect to these key requirements and derive a good compromise between model simplicity on the one hand, and applicability to MpSoC design on the other hand. The approach represents system-level scheduling analysis as a flow-analysis problem for event streams that can be configured to reuse the existing local scheduling analysis techniques. We define transformations between few key event stream models to meet the interfacing requirements of the compositional design style. An example demonstrates the application of the approach, as well as the worthiness of the results.
Kai Richter 0001, Razvan Racu, Rolf Ernst
RTSS3
2002 Model composition for scheduling analysis in platform design
abstract
We present a compositional approach to analyze timing behavior of complex platforms with different scheduling strategies. The approach uses event interfacing in order to couple previously incompatible analysis techniques which provide subsystem and component behavior. Based on these interfaces, event propagation using abstract models is used to derive global system timing properties.
Kai Richter 0001, Dirk Ziegenbein, Marek Jersak, Rolf Ernst
DAC4
2002 Associative caches in formal software timing analysis
abstract
Precise cache analysis is crucial to formally determine program running time. As cache simulation is unsafe with respect to the conservative running time bounds for real-time systems, current cache analysis techniques combine basic block level cache modeling with explicit or implicit program path analysis. We present an approach that extends instruction and data cache modeling from the granularity of basic blocks to program segments thereby increasing the overall running time analysis precision. Data flow analysis and local simulation of program segments are combined to safely predict cache line contents for associative caches in software running time analysis. The experiments show significant improvements in analysis precision over previous approaches on a typical embedded processor.
Fabian Wolf, Jan Staschulat, Rolf Ernst
DAC3
2002 System Design for Flexibility
abstract
With the term flexibility, we introduce a new design dimension of an embedded system that quantitatively characterizes its feasibility in implementing not only one, but possibly several alternative behaviors. This is important when designing systems that may adapt their behavior during operation, e.g., due to new environmental conditions, or when dimensioning a platform-based system that must implement a set of different behaviors. A hierarchical graph model is introduced that allows us to model flexibility and cost of a system formally. Based on this model, an efficient exploration algorithm to find the optimal flexibility/cost-tradeoff-curve of a system using the example of the design of a family of set-top boxes is proposed.
Christian Haubelt, Jürgen Teich, Kai Richter 0001, Rolf Ernst
DATE4
2002 Event Model Interfaces for Heterogeneous System Analysis
abstract
Complex embedded systems consist of hardware and software components from different domains, such as control and signal processing, many of them supplied by different IP vendors. The embedded system designer faces the challenge to integrate, optimize and verify the resulting heterogeneous systems. While format verification is available for some subproblems, the analysis of the whole system is currently limited to simulation or emulation. In this paper we tackle the analysis of global resource sharing, scheduling, and buffer sizing in heterogeneous embedded systems. For many practically used preemptive and non-preemptive hardware and software scheduling algorithms of processors and busses, semi-formal analysis techniques are known. However they cannot be used in system level analysis due to incompatibilities of their underlying event models. This paper presents a technique to couple the analysis of local scheduling strategies via an event interface model. We derive transformation rules between the most important event models and provide proofs where necessary. We use expressive examples to illustrate their application.
Kai Richter 0001, Rolf Ernst
DATE2
2002 SPI - a system model for heterogeneously specified embedded systems
abstract
Embedded systems typically include reactive and transformative functions, often described in different languages and semantics which are well established in their respective application domains. Additionally, a large part of the system functionality and components is reused from previous designs including legacy code. There is little hope that a single language will replace this heterogeneous set of languages. A design process must be able to bridge the semantic differences for verification and synthesis and should account for limited knowledge of system properties. This paper presents the system property intervals (SPI) model, which employs behavioral intervals and process modes to allow the common representation of different languages and semantics. This model is the basis of a workbench which is targeted at the design of heterogeneously specified embedded systems.
Dirk Ziegenbein, Kai Richter 0001, Rolf Ernst, Lothar Thiele, Jürgen Teich
IEEE Trans. Very Large Scale Integr. Syst.3
2001 Combining Languages in Embedded System Design
abstract
Often, several languages with different underlying models of computation are used in the design of an individual embedded system. The languages are selected because of their particular suitability for certain applications and optimizations, or because they have become generally accepted as a standard within an application field. The lack of coherency of the computational semantics, methods and tools is a significant obstacle on the way to higher design productivity and design quality. A similar problem occurs when reused components shall be integrated, possibly described in another language and incompletely documented. Examples are reused components or “legacy code.” The talk will start with a short overview of important models of computation. Then, different techniques to consistently combine model semantics are presented. We explain how to use such models for system analysis and scheduling. The embedded tutorial will conclude that unified languages are no necessity in system design and that a single language will face similar problems in system optimization as a combination of current system design languages.
Rolf Ernst
DSD1
2001 Execution cost interval refinement in static software analysis
Fabian Wolf, Rolf Ernst
J. Syst. Archit.2
2001 An approach to automated hardware/software partitioning using a flexible granularity that is driven by high-level estimation techniques
abstract
Hardware/software partitioning is a key issue in the design of embedded systems when performance constraints have to be met and chip area and/or power dissipation are critical. For that reason, diverse approaches to automatic hardware/software partitioning have been proposed since the early 1990s. In all approaches so far, the granularity during partitioning is fixed, i.e., either small system parts (e.g., base blocks) or large system parts (e.g., whole functions/processes) can be swapped at once during partitioning in order to find the best hardware/software tradeoff. Since the deployment of a fixed granularity is likely to result in suboptimum solutions, we present the first approach that features a flexible granularity during hardware/software partitioning. Our approach is comprehensive in so far that the estimation techniques, our multigranularity performance estimation technique described here in detail, that control partitioning, are adapted to the flexible partitioning granularity. In addition, our multilevel objective function is described. It allows us to tradeoff various design constraints/goals (performance/hardware area) against each other. As a result, our approach is applicable to a wider range of applications than approaches with a fixed granularity. We also show that our approach is fast and that the obtained hardware/software partitions are much more efficient (in terms of hardware effort, for example) than in cases where a fixed granularity is deployed.
Jörg Henkel, Rolf Ernst
IEEE Trans. Very Large Scale Integr. Syst.2
2001 FunState-an internal design representation for codesign
abstract
In this paper, an internal design model called FunState (functions driven by state machines) is presented that enables the representation of different types of system components and scheduling mechanisms using a mixture of functional programming and state machines. It is shown how properties relevant for scheduling and verification of specification models such as Boolean dataflow, cyclostatic dataflow, synchronous dataflow, marked graphs, and communicating state machines as well as Petri nets can be represented in the FunState model of computation. Examples of methods suited for FunState are described, such as scheduling and verification. They are based on the representation of the model's state transitions in the form of a periodic graph. The feasibility of the novel approach is shown with an asynchronous transfer mode switch example.
Karsten Strehl, Lothar Thiele, Matthias Gries, Dirk Ziegenbein, Rolf Ernst, Jürgen Teich
IEEE Trans. Very Large Scale Integr. Syst.5
2001 Path clustering in software timing analysis
abstract
Verification of program running time is essential in system design with real-time constraints. Simulation with incomplete test patterns or simple instruction counting are not appropriate for complex architectures. Software running times of embedded systems are process state and input data dependent. Formal analysis of such dependencies leads to software running time intervals rather than single values. These intervals depend on program properties, execution paths, and states of processes, as well as on the target architecture. An approach to analysis of process behavior using running time intervals is presented. It improves our previous work by exploiting program segments with single paths and by taking the execution context into account. The example of an asynchronous transfer mode (ATM) cell handler demonstrates significant improvements in analysis precision. Experimental results show the superiority of the presented approach over well-established approaches.
Fabian Wolf, Rolf Ernst, Wei Ye 0002
IEEE Trans. Very Large Scale Integr. Syst.2
2000 embedded system design with multiple languages: embedded tutorial
abstract
The use of several languages in the design of embedded systems is very convenient for application development and optimization but it can become an obstacle on the way to higher design productivity. This paper explains solutions and future trends.
Rolf Ernst, Ahmed Amine Jerraya
ASP-DAC1
2000 The Future of Flexible HW Platform Architectures Panel Discussion
Rolf Ernst, Grant Martin, Oz Levia, Pierre G. Paulin, Stamatis Vassiliadis, Kees A. Vissers
DATE1
2000 A Reconfigurable Hardware Platform for Digital Real-Time Signal Processing in Television Studios
abstract
By the continuous increase of speed and capacity of field programmable gate arrays (FPGAs) it becomes possible not only to process digital video data by FPGAs in real-time but also to implement the logic of complex video algorithms in one FPGA. For television (TV) studios the adherence to the real-time constraint is a mandatory requirement. Besides the reconfigurability of FPGAs allows the execution of different algorithms on the same hardware. With the reconfigurable hardware platform introduced in this contribution different video standards including high definition television (HDTV) can be processed in real-time. To realize a system with many complex video applications the platforms can be cascaded.
K. Henriss, Peter Rüffer, Rolf Ernst, Sieghard Hasenzahl
FCCM3
1999 Representation of Function Variants for Embedded System Optimization and Synthesis
abstract
Many embedded systems are implemented with a set of alternative function variants to adapt the system to different applications or environments. This paper proposes a novel approach for the coherent representation and selection of function variants in the different phases of the design process. In this context, the modeling of reconfiguration of system parts is supported in a natural way. Using a real example from the video processing domain, the approach is explained and validated. 1 Introduction Many embedded systems are implemented with a fixed core function and a set of alternative function variants to adapt the system to different applications or environments. Examples are TV sets which can be adapted to different standards or automotive control systems to be used in countries with different emission laws. Function variants are mutually exclusive, i. e. only one variant of a set of alternative functions is selected a time. There may be several of those variant sets in one embedde...
Kai Richter 0001, Dirk Ziegenbein, Rolf Ernst, Lothar Thiele, Jürgen Teich
DAC3
1999 Multi-Language System Design
abstract
The design of large systems, like a mobile telecommunication terminal or the electronic parts of an airplane or a car, may require the participation of several groups belonging to different companies and using different design methods, languages and tools. The concept of multi-language specification aims at coordinating different cultures through the unification of the languages, formalism and notations. This hot topic discusses the main issues and approaches to multi-language design. Two research directions are currently being explored by the EDA community. The first is based on the computation models underlying the languages while the second deals with the specification languages themselves.
Ahmed Amine Jerraya, Rolf Ernst
DATE2
1999 System level design and debug of high-performance embedded media systems (tutorial)
Rolf Ernst, Kees A. Vissers, Pieter van der Wolf, Gert-Jan van Rootselaar
ICCAD1
1999 Improved interconnect sharing by identity operation insertion
abstract
The paper presents an approach to reduce interconnect cost by insertion of identity operations in a control and data flow graph (CDFG). Other than previous approaches, it is based on systematic pattern analysis and automated transformation selection. The cost function controlling transformation selection is derived with statistical experiments and is optimized using practical benchmarks. The results show significantly reduced interconnect cost for most register architectures and application examples.
Dirk Herrmann, Rolf Ernst
ICCAD2
1999 FunState - an internal design representation for codesign
abstract
In this paper, an internal design model called FunState (functions driven by state machines) is presented that enables the representation of different types of system components and scheduling mechanisms using a mixture of functional programming and state machines. It is shown how properties relevant for scheduling and verification of specification models like boolean dataflow, cyclostatic dataflow, synchronous dataflow, marked graphs, and communicating state machines as well as Petri nets may be represented in the FunState model. Examples of methods suited for FunState are described, such as scheduling and verification. They are based on the representation of the model's state transitions in form of a periodic graph.
Lothar Thiele, Karsten Strehl, Dirk Ziegenbein, Rolf Ernst, Jürgen Teich
ICCAD4
1998 High-Level Estimation Techniques for Usage in Hardware/Software Co-Design
abstract
High-level estimation techniques are of paramount importance for design decisions like hardware/software partitioning or design space explorations. In both cases an appropriate compromise between accuracy and computation time determines about the feasibility of those estimation techniques. In this paper we present high-level estimation techniques for hardware effort and hardware/software communication time. Our techniques deliver fast results at sufficient accuracy. Furthermore, it is shown in which way these techniques are applied in order to cope with contradictory design goals like performance constraints and hardware effort constraints. As a solution, we present a cost function for the purpose of hardware/software partitioning that offers a dynamic weighting of its components. The conducted experiments show that the usage of our estimation techniques in conjunction with their efficient combination leads to reasonable hardware/software implementations as opposed to approaches that consider single constraints only.
Jörg Henkel, Rolf Ernst
ASP-DAC2
1998 Representation of process mode correlation for scheduling
abstract
The specijcation of embedded systems veq often contains a mi.x~ureof diferent models of computation.In particular the data $oti~and control $oiv associated to the transformative and reactii'e domains, respectively, are tightly coupled.The paper considers classes of applications that feature communicating processes ~vhoseflmctions depend on a]nite set of computation modes.The change behveen these modes is synchronized by data communication.An approach is presented to model the correlation of process modes and to fidly utilize this information for schedlding.A modeling ~ample sho}vs the optimization potential of the n~v approach.
Dirk Ziegenbein, Kai Richter 0001, Rolf Ernst, Jürgen Teich, Lothar Thiele
ICCAD3
1997 A Hardware/Software Partitioner Using a Dynamically Determined Granularity
abstract
Computer aided hardware/software partitioning is one of the keychallenges in hardware/software co-design. While previous approacheshave used a fixed granularity, i.e. the size of the partitioningobjects was fixed, we present a partitioning approach thatdynamically determines the partitioning granularity to adapt optimizationsteps to application properties and to intermediate optimizationresults. Experiments with simulated annealing optimizationshow a faster convergence and far better adaptability to costfunction variations than in previous experiments with fixed granularity.
Jörg Henkel, Rolf Ernst
DAC2
1997 A processor-coprocessor architecture for high end video applications
abstract
High end video applications are still implemented in hardware consisting of many components. Integration of these components on one IC is difficult as they are typically low volume products and often customization is also required, e.g. in studio applications. This is easier on the board level than on an integrated system. Using hardware parameters for customization can partly overcome the flexibility problem with additional hardware costs. Low cost can be obtained by a change in the architecture paradigm to a processor-coprocessor system. This, however, requires careful design space exploration since the performance target is beyond current DSP processors while at the same time flexibility is required. The paper presents the application of high level synthesis (D.D. Gajski et al., 1992) and novel hardware-software cosynthesis tools to design space exploration. It is shown that completely different algorithms can be mapped to the same target system at a much lower cost than the current approaches.
Elmar Maas, Dirk Herrmann, Rolf Ernst, Peter Rüffer, Sieghard Hasenzahl, Martin Seitz
ICASSP3
1997 Embedded program timing analysis based on path clustering and architecture classification
abstract
Formal program running time verification is an important issue in system design required for performance optimization under "first-time-right" design constraints and for real time system verification. Simulation based approaches or simple instruction counting are not appropriate and risky for more complex architectures in particular with data dependent execution paths. Formal analysis techniques have suffered from loose timing bounds leading to significant performance penalties when strictly adhered to. We present an approach which combines simulation and formal techniques in a safe way to improve analysis precision and tighten the timing bounds. Using a set of processor parameters, it is adaptable to arbitrary processor architectures. The results show an unprecedented analysis precision allowing us to reduce performance overhead for provably correct system or interface timing.
Rolf Ernst, Wei Ye 0002
ICCAD1
1997 An Adaptive Window Management System
Stefan Stille, Shailey Minocha, Rolf Ernst
INTERACT3
1995 A prototyping system for verification and evaluation in hardware-software cosynthesis
abstract
We present a system emulator for rapid prototyping of small embedded HW/SW-systems with hard timing constraints generated by a HW/SW cosynthesis system. It consists of a standard core processor and an application specific coprocessor, which is emulated by XILINX FPGAs. A byte slice architecture allows to emulate rather complex coprocessors. The system emulator supports the prototyping, debugging and time measurement in a comfortable way.
Thomas Benner, Rolf Ernst, Ingo Könenkamp, P. Schüler, H.-C. Schaub
RSP2
1995 An Approach to Automatic Display Layout Using Combinatorial Optimization Algorithms
abstract
Abstract The introduction of automatic display layout (ADL), i.e. the automatic placing and sizing of windows in a window‐oriented graphical user interface, is a major contribution towards an improved user interface. Our approach to ADL is to treat this problem as a combinatorial optimization problem. In this article we describe the concepts we used for implementing an experimental system which controls the computer screen contents and its layout. We give two examples of different standard applications into which we included ADL successfully, namely hypertext for a window layout problem and graph‐browser for a hierarchical graph layout problem within a particular window. The results show that automatic (and tool independent) display layout will be possible in the near future even in an interactive environment.
Peter Lüders, Rolf Ernst, Stefan Stille
Softw. Pract. Exp.2
1994 Adaptation of partitioning and high-level synthesis in hardware/software co-synthesis
abstract
Previously, we had presented the system COSYMA for hardware/software co-synthesis of small embedded con-trollers [ErHeBe93]. Target system of COSYMA is a core processor with application specific co–processors. The system speedup for standard programs compared to a sin-gle 33MHz RISC processor solution with fast, single cy-cle access RAM was typically less than 2 due to restric-tions in high-level co–processor synthesis, and incorrectly estimated back end tool performance, such as hardware synthesis, compiler optimization and communication opti-mization. Meanwhile, a high-level synthesis tool for high-performance co–processors in co-synthesis has been devel-oped. This paper explains the requirements and the main features of the high-level synthesis system and its integra-
Jörg Henkel, Rolf Ernst, Ulrich Holtmann, Thomas Benner
ICCAD2
1993 Speculative Computation for Coprocessor Synthesis
abstract
Time critical parts of a hardware-software system suitable for coprocessor implementation often contain nested loops. When loop pipelining is employed for high performance, control dependencies (conditional branches) in any part of the loop can become a dominant limitation to pipeline utilization. Speculative computation based on multiple branch prediction is a systematic approach that overcomes the problem of control dependencies and enables wide parallelism. We present the concept and some practical examples showing a speedup of up to three, with little hardware overhead.>
Ulrich Holtmann, Rolf Ernst
ICCD2
1993 Fast Timing Analysis for Hardware-Software Co-Synthesis
abstract
At the current time, an iterative approach seems to be best suited for hardware/software partitioning in hardware/software co-synthesis with time constraints. To check the timing constraints, the iteration loop contains a timing analysis. Only computation time-intensive RT-level simulation provides sufficient timing precision for complex processor architectures. We present a hardware/software timing analysis, which comes close to the precision of an RT-level simulation in a fraction of the computation time and, thus, removes a bottleneck from iterative hardware/software co-synthesis. We present some results for our co-synthesis system COSYMA.>
Wei Ye 0002, Rolf Ernst, Thomas Benner, Jörg Henkel
ICCD2
1993 Experiments with low-level speculative computation based on multiple branch prediction
abstract
Coprocessor design is one application of high-level synthesis. We want to focus on high-performance coprocessors to speed up time critical parts in hardware-software codesign of embedded controllers. Time critical software parts often contain nested loops, often with data-dependent branches and data-dependent number of iterations. When (loop) pipelining is employed for high performance, the control dependencies become a dominant limitation to pipeline utilization. Branch prediction is a possible approach, but is usually restricted to few instructions and to one branch because of hardware and control overhead. Multiple branch prediction and speculative computation take a more global view on the program flow. We give practical examples of how speculative computation with multiple branch prediction increases performance far beyond a usual ASAP scheduling based on a CDFG. For scheduling, speculative computation requires a modification of the CDFG and, for the allocation phase, the insertion of register sets to save the processor status. The controller needs slight modification. We conclude that manual application of our approach will in general be too difficult, such that it can only be used in connection with synthesis.>
Ulrich Holtmann, Rolf Ernst
IEEE Trans. Very Large Scale Integr. Syst.2
1991 Fault Tolerant VLSI Design with Functional Block Redundancy
abstract
Functional block redundancy is a dynamic redundancy technique for fault tolerance of VLSI circuits with nonregular logic structure, such as gate array designs. It exploits functional similarity of subcircuits, such as repeatedly used counter and shift register functions, to reduce the overhead of standby modules. The example of a manually optimized industrial gate array shows an extremely low overhead factor of 1.8 for complete single fault tolerance, which previously could not be reached for this type of circuit.>
Rolf Ernst, P. Nowottnick
ICCD1
1989 TSG: A Test System Generator for Debugging and Regression Test of High-Level Behavioral Synthesis Tools
abstract
The authors present a simulation-based system for testing high-level behavioral synthesis tools. Applications are tool debugging and automatic regression test. A key feature is a transformation of sequential circuits for application of random test patterns.>
Rolf Ernst, S. Sutarwala, J.-Y. Jou
ITC1