Robin Hapka

dblp:305/9009 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-5201-3212ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 A Novel Timing Model for Neural Networks in Safety-critical Systems
abstract
Deep Neural Networks (DNNs) have a largely fixed program path, so average and worst-case timing depend mostly on data values and physical platform properties. Related work on recent high-performance platforms already observed a dominant role of circuit parameters and operating conditions leading to a normal distribution of DNN timing variations. In extensive experiments with different DNNs on very different types of such platforms used for real-time applications, we could confirm that a Gaussian distribution closely approximates between 99% and 99.9% of all DNN executions. Remaining rare outliers are often neglected in literature, usually assuming that they will not affect applications. However, in experiments with long test series we could show that outliers are frequent enough to be relevant for critical system design and even form dense outlier sequences that can seriously challenge latency critical applications, such as vehicle perception. We propose and evaluate a dual timing model capturing both typical response time and outlier distributions and provide design methods to control and mitigate their effects in design and operation. While developed for DNNs, the approach could be relevant for other critical real-time applications on high-performance platforms.
Robin Hapka, Anika Christmann, Rolf Ernst
COMPSAC1
2024 Conservative Design with High-Performance COTS Architectures - Beyond Traditional Approaches
abstract
Safety-critical industrial systems are subject to many stringent requirements, including non-functional design constraints that are essential for robustness, reliability, and safety. System timing is often treated as an afterthought, even though data age, race conditions, or missed deadlines can affect functional safety. This challenge has been addressed in real-time system design, but the situation has worsened in recent system designs. Many applications require high performance to achieve behavioral safety goals, i.e., safety of intended functionality, using complex heterogeneous architectures with multi-level memory, interconnected chiplets, and accelerators based on the latest technologies. Machine learning applications in factory automation, robotics, and autonomous systems are important examples. The timing of such architectures is heavily influenced by the physical effects of temperature control and chip variations, including aging, leading to even more complex chip-specific software execution timing. The paper shows that identical chips differ in their timing behavior, and that most of the timing variation of a single chip can be effectively modeled using additive white Gaussian noise. We explain how such timing behavior can be exploited in high-assurance design for safety and reliability, following established reliability analysis, test, and redundancy techniques.
Robin Hapka, Rolf Ernst
ETFA1
2023 Efficient hard real-time implementation of CNNs on multi-core architectures
abstract
Autonomous driving applications rely on processing large amounts of data in order to ensure sufficient perception performance. In this context, Convolutional Neuronal Networks (CNNs), which are used for object detection in camera images, are an integral part of sensor data processing. However, safety-related hard deadlines are counteracted by the required high data rates that stress the platform w.r.t. memory interference. Providing high camera frame rates under hard latency requirements is still an open issue. In a case study focusing on high-performance multi-core architectures, we evaluate in detail the advantages of a private L2 and shared L3 cache architecture for CNN processing and show how to remove malicious data synchronization effects. In this context we deploy MobileNet as well as YOLO on an Intel i5 processor and demonstrate their applicability to a worst-case design under high CNN frame rates. Last, we provide a heuristic optimization scheme that is able to efficiently find feasible high frame rate configurations. Contrary to common assumptions, our results show that high-performance multi-core COTS platforms are suitable for the application of CNNs even under hard deadline constraints and, hence, offer substantial gains in cost and productivity due to their greater ease of programming.
Jonas Peeck, Robin Hapka, Rolf Ernst
COMPSAC2
2023 Formal Analysis of Timing Diversity for Autonomous Systems
abstract
The design of autonomous systems, such as for automated driving and avionics, is challenging due to high performance requirements combined with high criticality. Complex applications demand the full performance of commercial high performance multi-core systems of-the-shelf (COTS), with or without accelerators. While these systems are optimized for performance, hard real-time requirements and deterministic timing behavior are major constraints for safety-critical systems. Unfortunately, infrequent timing outliers caused by interleaved hardware-software effects of COTS systems complicate traditional worst-case design. This conflict often prohibits deploying COTS hardware and consequently prevents sophisticated applications, too. Recently, an approach called Timing Diversity was introduced, which proposes to exploit existing dual modular redundant hardware platforms to mask deadline violations. This paper puts Timing Diversity on a theoretical foundation and provides specification for different implementations. It demonstrates that Timing Diversity needs fast recovery to be effective, proposes a recovery strategy and provides a mathematical model for the reliability of the resulting system. Using experimental data in a Linux based system, it shows that fast recovery is useful, making Timing Diversity a realistic option for compute demanding hard real-time applications.
Anika Christmann, Robin Hapka, Rolf Ernst
DATE2
2022 Controlling High-Performance Platform Uncertainties with Timing Diversity
abstract
Autonomous mobile systems combine high performance requirements with safety criticality. High performance hardware/software architectures, however, expose a far more complex runtime behavior than traditional microcontroller architectures. Such high-performance architectures challenge traditional worst-case design that assumes a formally analyzable or at least deterministic worst-case response time (WCRT) that can be reasonably bounded. However, such architectures expose rare but substantial worst-case outliers, which are not only caused by the application itself, but also by the many dynamic influences of software architecture and platform control. Probabilistic methods can capture such outliers, but are only effective, if the outlier probability is sufficiently low and if the methods cover dynamic platform timing. As a main contribution, this paper exploits platform induced timing variety rather than trying to mitigate it. Assuming the typical redundant dual modular redundancy (DMR) implementation that is deployed in safety-critical systems, it introduces the concept of Timing Diversity, where rare outliers in one of the two channels are masked by the other channel with a sufficiently high probability. The paper uses a convolutional neural network (CNN) example in different parameter settings running on Linux operated multi-core platform with typical dynamic control to investigate the proposed concept. The experiments demonstrate the potential of Timing Diversity in leading to substantially higher reliability. Alternatively, the approach permits a reduction of the system WCRT at the same reliability level.
Robin Hapka, Anika Christmann, Rolf Ernst
RTCSA1
2021 Timing diversity as a protective mechanism: work-in-progress
abstract
Dual modular redundancy (DMR) is not only an established solution for systems with high reliability demands, it is even required in aviation certification standards such as DO-254 [5, Clause 2.3.1]. A safety critical avionic application such as the flight control system is designed with up to 6-fold redundancy and the Avionics Full-Duplex Ethernet (AFDX) communication network is also based on the DMR. Even in the automotive domain, DMR is a well known solution. ISO26262 [3, Part 6, Clause 7.4.13] also suggests heterogeneous or diverse redundancy for safety-critical applications including software which must be redundantly executed on independent hardware components to avoid failure due to hardware errors. We exploit this mandatory software redundancy to master timing errors of critical software with minimum additional overhead.
Mischa Möstl, Robin Hapka, Anika Christmann, Rolf Ernst
EMSOFT2