Matthias Becker 0004

dblp:69/3369-4 · DBLP profile ↗
← Back
46ranked-venue papers
16as first author
23since 2021 · last 2026
0000-0002-1276-3609ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 13 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021
YearPublicationVenuePosition
2026 Preemption Threshold Assignment to Improve Schedulability under Memory Constraints
Thilanka Thilakasiri, Matthias Becker 0004
DATE2
2026 From Timing Budgets to WCETs: Robust SIL- and BSW-Aware Clustering and Allocation for Iterative Automotive Software Development
abstract
Automotive ECUs integrate thousands of AUTOSAR runnables, substantial Basic Software (BSW), and heterogeneous multicore hardware. In iterative software-defined vehicle development, engineers must repeatedly revisit designs while maintaining stable runnable clustering and core allocations, which are expensive structural decisions. Beyond timing, Safety Integrity Levels (SILs), BSW overheads, and per-core memory strongly constrain these decisions, yet are rarely modeled jointly. This paper addresses these challenges through a chain-based analysis model that treats SIL constraints and BSW costs as first-class citizens, as well as an integrated toolchain that constructs job-level data-age constraints, forms SIL-compliant clusters, synthesizes multirate tasks, and maps application and BSW tasks to heterogeneous multicore platforms while checking timing and memory feasibility. A case study based on a real-world motion/drive controller from our industrial partner is described, which serves as the basis for our evaluation. The evaluations are conducted using synthetic systems that reflect the characteristics of the case study. Across 13,825 synthesized systems, SIL/BSW-aware clustering substantially reduces pessimism in analysis. In the industrial configuration, our approach yields a 7% decrease in utilization, demonstrating its practical value. A refinement study, which progressively replaces early budget assumptions with WCET samples, indicates that SIL/BSW-aware clustering preserves structural decisions better than less-informed variants under the same resampling setup.
Tobias Denzinger, Matthias Becker 0004, Peter Ulbrich
ECRTS2
2026 Shape-Aware Analysis of End-to-End Latency Under LET
Mario Günzel, Matthias Becker 0004, Daniel Casini
RTAS2
2025 Special Session - Predictable Timing Behavior in Distributed Cyber-Physical Systems
abstract
Ensuring predictable and deterministic behavior in distributed cyber-physical systems (CPS) is essential for guaranteeing safety, reliability, and real-time behavior. However, achieving this predictability is challenging due to network uncertainties, asynchronous execution, and complex timing interactions.
Jian-Jia Chen, Mario Günzel, Dakshina Dasari, Matthias Becker 0004, Edward A. Lee, Timothy Bourke
EMSOFT4
2025 Optimal Task Phasing for End-To-End Latency in Harmonic and Semi-Harmonic Automotive Systems
abstract
In the context of automotive systems, the end-toend latency of a sequence of tasks (a so-called cause-effect chain) is a common metric to ensure correct timing behavior. To control the end-to-end latency, proper task configuration is crucial. While the literature considers the configuration of task periods, optimization of task phases to minimize the end-to-end latency is only sparsely discussed. In this work, we examine the configuration of task phases to optimize the end-to-end latency of a cause-effect chain that communicates under the Logical Execution Time (LET) paradigm. To that end, we develop a strategy for cause-effect chains with harmonic or semi-harmonic periods, which are very common in industrial applications. We prove that our strategy is optimal in the sense that it minimizes the end-to-end latency. Furthermore, our evaluation based on a real-world use-case and on synthetic automotive benchmarks shows that optimizing task phases can reduce end-to-end latencies significantly. Our approach takes at most$49 \mu ~\mathrm{s}$to find the optimal phasing and compute the end-toend latency for cause-effect chains with 50 tasks, reducing the end-to-end latency by 28 % in median.
Mario Günzel, Matthias Becker 0004
RTAS2
2025 Managing real-time constraints through monitoring and analysis-driven edge orchestration
abstract
Emerging real-time applications are increasingly moving to distributed heterogeneous platforms , under the promise of more powerful and flexible resource capabilities. This shift inevitably brings new challenges. The design space to deploy chains of threads is more complex, and sound estimates of worst-case execution times are harder to obtain. Additionally, the environment is more dynamic, requiring additional runtime flexibility on the part of the application itself. In this paper, we present an optimization-based approach to this problem. First, we present a model and real-time analysis for modern distributed edge applications. Second, we propose a design-time optimization problem to show how to set the main parameters characterizing such applications from a time-predictability perspective. Then, we present an orchestration and runtime decision-making mechanism that monitors execution times and allows for runtime reconfigurations , spanning from graceful degradation policies to re-distributions of workload. A prototypical implementation of the proposed approach based on the QNX RTOS and its evaluation on a realistic case study based on an edge-based valet parking application conclude the paper.
Daniel Casini, Paolo Pazzaglia, Matthias Becker 0004
J. Syst. Archit.3
2025 Real-time probabilistic programming
abstract
Complex cyber–physical systems interact in real time and must consider both timing and uncertainty. Developing software for such systems is expensive and difficult, especially when modeling, inference, and real-time behavior must be developed from scratch. In the last decade, a popular general probabilistic modeling paradigm has emerged—called probabilistic programming languages (PPLs)—that simplifies modeling and inference by separating the concerns between probabilistic modeling and inference algorithm implementation. However, these languages have primarily been designed for offline problems, not online real-time systems. In this paper, we combine PPLs and real-time programming primitives by introducing the concept of real-time probabilistic programming languages (RTPPL). We develop an RTPPL called ProbTime and a new approach for fairness-guided optimization of inference accuracy of a ProbTime system under schedulability constraints. Moreover, we illustrate the applicability of ProbTime on an automotive testbed performing indoor positioning and braking.
Lars Hummelgren, Matthias Becker 0004, David Broman
J. Syst. Archit.2
2024 Meeting Job-Level Dependencies by Task Merging
abstract
Industrial applications are often time critical and subject to end-to-end latency constraints. Job-level dependencies can be leveraged to specify a partial ordering on tasks’ jobs already at early design phases, agnostic of the hardware platform or scheduling algorithm, and guarantee that end-to-end latency constraints of task chains are met as long as the job-level dependencies are respected. However, their realization at runtime can introduce overheads and complicates the scheduling and timing analysis. This work presents an approach that merges multi-periodic tasks that are connected by job-level dependencies to a single task. A Constraint Programming formulation is presented that optimally merges such task clusters while all job-level dependencies are respected. Such an approach removes the need to consider job-level dependencies at runtime without being bound to a specific scheduling algorithm. Evaluations highlight the applicability of the approach by system-level experiments and showcase the scalability of the approach using synthetic task clusters.
Matthias Becker 0004
ASPDAC1
2024 Towards Request Arbitration in Edge-Assisted Smart Intersections Under Timing Constraints
abstract
Infrastructure-assisted autonomous driving has the potential to improve the decision-making process of autonomous driving. The car's perception can be augmented with information collected by infrastructure, such as smart intersections, that otherwise cannot be perceived by the car itself. This additional information enables more informed decisions, improving road safety. At the roadside unit, functionality is realized by a chain of tasks. Requests sent by autonomous cars must be handled within a specific time, while it is additionally crucial that the data age of returned data is within a defined bound. In this work, we discuss data age and response time as two competing metrics in data requests from an autonomous car to a smart intersection. A novel policy to arbitrate pending requests of a server is described that utilizes properties of the task chains that follow the Logical Execution Time paradigm to meet both response time and data age constraints of requests. Evaluations demonstrate the importance of the proposed method by improving the acceptance ratio of requests by up to 36 % compared to a server that does not consider data age constraints.
Matthias Becker 0004, Fredrik Asplund
ETFA1
2024 Multi-objective preference-free exact design space exploration of static DSP on multicore platforms
abstract
A challenge in designing resource-constrained embedded systems for digital signal processing (DSP) is their complexity due to their vast design spaces, where only a fraction of implementations are feasible or optimal. A crucial tool to aid in this challenge is automated design space exploration (DSE). However, no exact, multi-objective, and preference-free DSE approach exists for DSP applications on resource-constrained embedded platforms.We propose a novel DSE solution with these ideal characteristics to perform DSE of analyzable DSP applications for tile-based multiprocessing embedded platforms. Our proposal harmonizes the exactness of constraint programming (CP) and the exploration efficiency of genetic algorithms (GA). Through this synergy, no single-objective reduction strategy or a priori objective preferences is required.We evaluate the proposal through state-of-the-art single-objective case studies and multi-objective case studies inspired by these. The evaluations show that our proposal improves the single-objective state-of-the-art and finds high-quality approximate Pareto-frontiers for the multi-objective case study. Therefore, our proposal is a more performant single-objective DSE solution than the state-of-the-art, and it is the first exact, multi-objective, and preference-free DSE approach for the problem addressed.
Rodolfo Jordão, Fahimeh Bahrami, Yu Yang 0020, Matthias Becker 0004, Ingo Sander, Kathrin Rosvall
FDL4
2024 Work-in-Progress: Exploring Limited Preemption Approaches for the Phased Execution Model
abstract
Phased execution models separate computation from access to shared resources to make task execution predictable. These task models minimize interference between tasks, making them suitable for modern complex multi-core platforms. In the phased execution model, tasks perform computations only using the local memory to avoid accessing the shared memory during task execution. All instructions and data, including the intermediate results, are stored in the local memory during execution. Thus, the local memory size becomes a crucial factor in contrast to conventional execution. In the literature, non-preemptive and fully preemptive execution of phased execution models are studied. While the non-preemptive approaches utilize the local memory well, schedulability is reduced due to blocking. On the other hand, fully preemptive execution phases allow for better schedulability but require significantly more local memory capacity to implement preemptions at runtime without violating the model’s execution semantics. Thus, this work evaluates different approaches to limited preemptive scheduling of the phased execution model under partitioned fixed-priority scheduling. We demonstrate that preemption thresholds and non-preemptive regions can successfully be used to satisfy both timing and memory constraints of phased tasks.
Thilanka Thilakasiri, Matthias Becker 0004
RTSS2
2024 The MATERIAL framework: Modeling and AuTomatic code Generation of Edge Real-TIme AppLications under the QNX RTOS
abstract
Modern edge real-time automotive applications are becoming more complex, dynamic, and distributed, moving away from conventional static operating environments to support advanced driving assistance and autonomous driving functionalities. This shift necessitates formulating more complex task models to represent the evolving nature of these applications aptly. Modeling of real-time automotive systems is typically performed leveraging Architectural Languages (ALs) such as Amalthea, which are commonly used by the industry to describe the characteristics of processing platforms, operating systems, and tasks. However, these architectural languages are originally derived for classical automotive applications and need to evolve to meet the needs of next-generation applications. This paper proposes an automatic framework for the modeling and automatic code generation of dynamic automotive applications under the QNX RTOS. To this end, we extend Amalthea to describe chains of communicating tasks with multiple operating modes and to consider the QNX’s reservation-based scheduler, called APS, which allows providing temporal isolation between applications co-located on the same hardware platform. Finally, an evaluation is presented to compare different implementation alternatives under QNX that are automatically generated by our code generation framework.
Matthias Becker 0004, Daniel Casini
J. Syst. Archit.1
2024 Introduction to the Special Issue on Real-Time Computing in the IoT-to-Edge-to-Cloud Continuum
abstract
Special Issue Part 1 (Issue 3) and Part 2 (Issue 4) of AIEDAM are based on a workshop on Learning and Creativity held at the 2002 conference on Artificial Intelligence in Design, AID '02 (www.cad.strath.ac.uk/AID02_workshop/Workshop_webpage.html; Gero, ...
Daniel Casini, Dakshina Dasari, Matthias Becker 0004, Giorgio C. Buttazzo
ACM Trans. Embed. Comput. Syst.3
2024 IDeSyDe: Systematic Design Space Exploration via Design Space Identification
abstract
Design space exploration (DSE) is a key activity in embedded design processes, where a mapping between applications and platforms that meets the process design requirements must be found. Finding such mappings is very challenging due to the complexity of modern embedded platforms and applications. DSE tools aid in this challenge by potentially covering sections of the design space that could be unintuitive to designers, leading to more optimised designs. Despite this potential benefit, DSE tools remain relatively niche in the embedded industry. A significant obstacle hindering their wider adoption is integrating such tools into embedded design processes. We present two contributions that address this integration issue. First, we present the design space identification (DSI) approach for systematically constructing DSE solutions that are modular and tuneable. Modularity means that DSE solutions can be reused to construct other DSE solutions, while tuneability means that the most specific DSE solution is chosen for the target DSE problem. Moreover, DSI enables transparent cooperation between exploration algorithms. Second, we present IDeSyDe, an extensible DSE framework for DSE solutions based on DSI. IDeSyDe allows extensions to be developed in different programming languages in a manner compliant with the DSI approach. We showcase the relevance of these contributions through five different case studies. The case study evaluations showed that non-exploration DSI procedures create overheads, which are marginal compared to the exploration algorithms. Empirically, most evaluations average 2% of the total DSE request. More importantly, the case studies have shown that IDeSyDe indeed provides a modular and incremental framework for constructing DSE solutions. In particular, the last case study required minimal extensions over the previous case studies so that support for a new application type was added to IDeSyDe.
Rodolfo Jordão, Matthias Becker 0004, Ingo Sander
ACM Trans. Design Autom. Electr. Syst.2
2023 An Exact Schedulability Analysis for Global Fixed-Priority Scheduling of the AER Task Model
abstract
Commercial off-the-shelf (COTS) multi-core platforms offer high performance and large availability of processing resources. Increased contention when accessing shared resources is a result of the high parallelism and one of the main challenges when realtime applications are deployed to these platforms. As a result, several execution models have been proposed to avoid contention by separating access to shared resources from execution.
Thilanka Thilakasiri, Matthias Becker 0004
ASP-DAC2
2023 On the QNX IPC: Assessing Predictability for Local and Distributed Real-Time Systems
abstract
With the advent of massively distributed applications such as those required by the IoT-to-Edge-to-Cloud compute continuum (i.e., automotive, smart agriculture, smart manufacturing, and more), real-time communication mechanisms allowing physically distributed nodes to seamlessly communicate as if they were running on the same host acquired noteworthy importance. To this end, the synchronous inter-process communication (IPC) mechanism provided by the QNX operating system (OS) is a promising candidate, as it allows using the application programming interface for communicating both on a single- and multi-node setting. Furthermore, it provides priority and partition inheritance mechanisms to improve predictability when working with the Adaptive Partitioning Scheduler (APS), a reservationbased scheduler provided by the QNX OS. This paper explores the behavior of the QNX synchronous message-passing (SyncMP) IPC with an extensive set of experiments, using them to formalize its behavior and model it from a real-time perspective. Then, it provides a response-time analysis for client-server applications based on the QNX SyncMP building upon self-suspending task theory. Finally, we evaluate the analysis on an application based on the WATERS 2019 Challenge by Bosch.
Matthias Becker 0004, Dakshina Dasari, Daniel Casini
RTAS1
2023 Methods to Realize Preemption in Phased Execution Models
abstract
Phased execution models are a good solution to tame the increased complexity and contention of commercial off-the-shelf (COTS) multi-core platforms, e.g., Acquisition-Execution-Restitution (AER) model, PRedictable Execution Model (PREM). Such models separate execution from access to shared resources on the platform to minimize contention. All data and instructions needed during an execution phase are copied into the local memory of the core before starting to execute. Phased execution models are generally used with non-preemptive scheduling to increase predictability. However, the blocking time in non-preemptive systems can reduce schedulability. Therefore, an investigation of preemption methods for phased execution models is warranted. Although, preemption for phased execution models must be carefully designed to retain its execution semantics, i.e., the handling of local memory during preemption becomes non-trivial. This paper investigates different methods to realize preemption in phased execution models while preserving their semantics. To the best of our knowledge, this is the first paper to explore different approaches to implement preemption in phased execution models from the perspective of data management. We introduce two strategies to realize preemption of execution phases based on different methods of handling local data of the preempted task. Heuristics are used to create time-triggered schedules for task sets that follow the proposed preemption methods. Additionally, a schedulability-aware preemption heuristic is proposed to reduce the number of preemptions by allowing preemption only when it is beneficial in terms of schedulability. Evaluations on a large number of synthetic task sets are performed to compare the proposed preemption models against each other and against a non-preemptive version. Furthermore, our schedulability-aware preemption heuristic has higher schedulability with a clear margin in all our experiments compared to the non-preemptive and fully-preemptive versions.
Thilanka Thilakasiri, Matthias Becker 0004
ACM Trans. Embed. Comput. Syst.2
2022 End-to-End Analysis of Event Chains under the QNX Adaptive Partitioning Scheduler
abstract
Modern autonomous cars run classic AUTOSAR applications alongside advanced driving assistance systems on a single-vehicle computer. Ensuring safety and predictability in such a complex system is challenging and requires temporal isolation between the various components. A promising solution is the POSIX-compliant QNX operating system: it meets the automotive standards for functional safety at the highest level (ISO 26262 ASIL-D) and provides temporal isolation through the Adaptive Partitioning Scheduler (APS), a resource reservation algorithm that guarantees processor bandwidth to groups of threads. These guarantees make it an ideal platform for composing diverse and complex applications on centralized vehicle computers. However, so far, there is no precise description or analysis of the APS reservation mechanism in real-time literature. In this paper, we provide the first description of the behavior of the APS from a real-time point of view and validate the results by running experiments on a real QNX platform. Based on the derived scheduler rules, we develop a response-time analysis to bound the end-to-end latency of event chains under APS. Finally, we evaluate different design strategies on a case study based on a real autonomous construction vehicle.
Dakshina Dasari, Matthias Becker 0004, Daniel Casini, Tobias Stark
RTAS2
2022 Enabling automated integration of architectural languages: An experience report from the automotive domain
abstract
Modern automotive software systems consist of hundreds of heterogeneous software applications, belonging to separated function domains and often developed within distributed automotive ecosystems consisting of original equipment manufactures, tier-1 and tier-2 companies. Hence, the development of modern automotive software systems is a formidable challenge. A well-known instrument for coping with the tremendous heterogeneity and complexity of modern automotive software systems is the use of architectural languages as a way of enabling different and specific views over these systems. However, the use of different architectural languages might come with the cost of reduced interoperability and automation as different languages might have weak to no integration. In this article, we tackle the challenge of integrating two architectural languages heavily used in the automotive domain for the design and timing analysis of automotive software systems: AMALTHEA and Rubus Component Model. The main contributions of this paper are (i) a mapping scheme for the translation of an AMALTHEA architecture into a Rubus Component Model architecture where high-precision timing analysis can be run, and the back annotation of the analysis results on the starting AMALTHEA architecture; (ii) the implementation of the proposed scheme, which uses the concept of model transformations for enabling a full-fledged automated integration; (iii) the application of such automation on three industrial automotive systems being the brake-by-wire, the full blown engine management system and the engine management system. We discuss and evaluate the proposed contributions using an online, experts survey and the above-mentioned use cases. Based on the evaluation results, we conclude that the proposed automation mechanism is correct and applicable in industrial contexts. Besides, we observe that the performance of the automation mechanism does not degrade when translating large models with several thousands of elements. Eventually, we conclude that experts in this field find the proposed contribution industrially relevant.
Alessio Bucaioni, Matthias Becker 0004
J. Syst. Softw.2
2021 Optimizing Inter-Core Data-Propagation Delays in Industrial Embedded Systems under Partitioned Scheduling
abstract
This paper addresses the scheduling of industrial time-critical applications on multi-core embedded systems. A novel scheduling technique under partitioned scheduling is proposed that minimizes inter-core data-propagation delays between tasks that are activated with different periods. The proposed technique is based on the read-execute-write model for the execution of tasks to guarantee temporal isolation when accessing the shared resources. A Constraint Programming formulation is presented to find the schedule for each core. Evaluations are preformed to assess the scalability as well as the resulting schedulability ratio, which is still 18% for two cores that are both utilized 90%. Furthermore, an automotive industrial case study is performed to demonstrate the applicability of the proposed technique to industrial systems. The case study also presents a comparative evaluation of the schedules generated by (i) the proposed technique and (ii) the Rubus-ICE industrial tool suite with respect to jitter, inter-core data-propagation delays and their impact on data age of task chains that span multiple cores.
Lamija Hasanagic, Tin Vidovic, Saad Mubeen, Mohammad Ashjaei, Matthias Becker 0004
ASP-DAC5
2021 Formulation of Design Space Exploration Problems by Composable Design Space Identification
abstract
Design space exploration (DSE) is a key activity in embedded system design methodologies and can be supported by well-defined models of computation (MoCs) and predictable platform architectures. The original design model, covering the application models, platform models and design constraints needs to be converted into a form analyzable by computer-aided decision procedures such as mathematical programming or genetic algorithms. This conversion is the process of design space identification (DSI), which becomes very challenging if the design domain comprises several MoCs and platforms. For a systematic solution to this problem, separation of concerns between the design domain and decision domain is of key importance. We propose in this paper a systematic DSI scheme that is (a) composable, as it enables the stepwise and simultaneous extension of both design and decision domain, and (b) tuneable, because it also enables different DSE solving techniques given the same design model. We exemplify this DSI scheme by an illustrative example that demonstrates the mechanisms for composition and tuning. Additionally, we show how different compositions can lead to the same decision model as an important property of this DSI scheme.
Rodolfo Jordão, Ingo Sander, Matthias Becker 0004
DATE3
2021 From the Synchronous Data Flow Model of Computation to an Automotive Component Model
abstract
The size and complexity of automotive software systems are steadily increasing. Software functions are subject to different requirements and belong to different functional domains of the car. Meanwhile, streaming applications have become increasingly relevant in emerging application areas such as Advanced Driving Assistance Systems. Among models for streaming applications, the Synchronous Data Flow model is well-known for its analysable properties. This work presents transformation rules that allow transforming applications described by the Synchronous Data Flow model to an automotive component model. The proposed transformation rules are implemented in form of a software plugin for an automotive tool suite that allows for timing analysis, code synthesis and deployment to a Real-Time Operating System. To demonstrate the applicability of the proposed approach, a case study of a Kalman filter that is part of a simplified cruise control application is presented. An abstract Synchronous Data Flow model of the filter is transformed into a component that is deployed on an Electronic Control Unit with hard timing guarantees.
Mehmet Onur Aybek, Rodolfo Jordão, John Lundbäck, Kurt-Lennart Lundbäck, Matthias Becker 0004
ETFA5
2021 Constrained Data-Age with Job-Level Dependencies: How to Reconcile Tight Bounds and Overheads
abstract
Many industrial real-time systems rely on the implicit register communication paradigm to minimize overheads and ease distributed development. Here, tasks follow a simple input-processing-output scheme, and data is passed without synchronization by the last-is-best semantics. In these systems, the age of data is the primary real-time objective, which is defined by data-flow chains that span from the system's inputs to outputs. Consequently, a real-time analysis aims to provide guarantees on worst-case data age. In general, there are two main approaches: (1) Task-level scheduling such that inter-task communication is arranged at the beginning and end of a task's execution interval, which guarantees a deterministic yet highly pessimistic data age. (2) Job-level dependencies (JLD) that are added at critical points in the schedule to link specific job instances of tasks of a multi-rate data-flow chain, which provides tighter upper bounds on data ages. However, the drawback is that JLDs induce substantial synchronization overheads, impact the overall schedulability, and are much more challenging to implement. In this paper, we address the trade-off between tight data-age guarantees, synchronization overheads, and schedulability in multi-core settings. Our proposed solution is to combine the potential of job-level optimization with the determinism and low overheads of static, task-level approaches. Therefore, we present a novel execution model to efficiently map data-age constrained tasksets with job-level dependencies on event-triggered systems by automated system analysis and transformation. Experimental results of an extensive real-world case study substantiate that our approach can further tighten data-age bounds, reduce overheads, and ease schedulability.
Tobias Klaus, Matthias Becker 0004, Wolfgang Schröder-Preikschat, Peter Ulbrich
RTAS2
2020 From AMALTHEA to RCM and Back: a Practical Architectural Mapping Scheme
abstract
This paper focuses on the mapping between two industrial architectural languages: AMALTHEA and Rubus Component Model. Both languages are heavily used within the automotive domain for the design and timing analysis of automotive software, respectively. The main contribution of this paper is a mapping scheme between the two architectural languages enabling i) the translation of an AMALTHEA architecture into a Rubus Component Model architecture where high-precision timing analysis can be performed ii) and the back-propagation of the analysis results on the AMALTHEA architecture. We validate the applicability of the proposed mapping scheme using an industrial use case from the automotive domain: the brake-by-wire system. We discuss the industrial relevance and lessons learnt of this work using expert interviews.
Alessio Bucaioni, Matthias Becker 0004, John Lundbäck, Harald Mackamul
SEAA2
2019 Static Allocation of Parallel Tasks to Improve Schedulability in CPU-GPU Heterogeneous Real-Time Systems
abstract
Autonomous driving is one of the main challenges of modern cars. Computer visions and intelligent on-board decision making are crucial in autonomous driving and require heterogeneous processors with high computing capability under low power consumption constraints. The progress of parallel computing using heterogeneous processing units is further supported by software frameworks like OpenCL, OpenMP, CUDA, and C++AMP. These frameworks allow the allocation of parallel computation on different compute resources. This, however, creates a difficulty in allocating the right computation segments to the right processing units in such a way that the complete system meets all its timing requirements. In this paper, we consider pre-runtime static allocations of parallel tasks to perform their execution either sequentially on CPU or in parallel using a GPU. This allows for improving any unbalanced use of GPU accelerators in a heterogeneous environment. By performing several heuristic algorithms, we show that the overuse of accelerators results in a bottle-neck of the entire system execution. The experimental results show that our allocation schemes that target a balanced use of GPU improves the system schedulability up to 90%.
Nandinbaatar Tsog, Matthias Becker 0004, Fredrik Bruhn, Moris Behnam, Mikael Sjödin
IECON2
2019 An Adaptive Resource Provisioning Scheme for Industrial SDN Networks
abstract
Many industrial domains face the challenge of ever growing networks, driven for example by Internet-of-Things and Industry 4.0. This typically comes together with increased network configuration and management efforts. In addition to the increasing network size, these domains typically are subject to adaptive load situations that pose an additional challenge on the network infrastructure.Software defined networking (SDN) is a promising networking paradigm that reduces configuration complexity and management effort in Ethernet networks. In this work, we investigate SDN in context of adaptive scenarios with QoS constraints. Our approach applies monitoring of several thresholds which automatically trigger redistribution of resources via the central SDN controller. This setup leads to an agile system that can dynamically react to load changes while the infrastructure is not overprovisioned. The approach is implemented in a low-level simulation environment where we demonstrate the benefits of the approach using a case study.
Matthias Becker 0004, Zhonghai Lu, Dejiu Chen
INDIN1
2018 Scheduling multi-rate real-time applications on clustered many-core architectures with memory constraints
abstract
Access to shared memory is one of the main challenges for many-core processors. One group of scheduling strategies for such platforms focuses on the division of tasks' access to shared memory and code execution. This allows to orchestrate the access to shared local and off-chip memory in a way such that access contention between different compute cores is avoided by design. In this work, an execution framework is introduced that leverages local memory by statically allocating a subset of tasks to cores. This reduces the access times to shared memory, as off-chip memory access is avoided, and in turn improves the schedulability of such systems. A Constraint Programming (CP) formulation is presented to select the statically allocated tasks and to generate the complete system schedule. Evaluations show that the proposed approach yields an up to 19% higher schedulability ratio than related work, and a case study demonstrates its applicability to industrial problems.
Matthias Becker 0004, Saad Mubeen, Dakshina Dasari, Moris Behnam, Thomas Nolte
ASP-DAC1
2018 Towards QoS-Aware Service-Oriented Communication in E/E Automotive Architectures
abstract
With the raise of increasingly advanced driving assistance systems in modern cars, execution platforms that build on the principle of service-oriented architectures are being proposed. Alongside, service oriented communication is used to provide the required adaptive communication infrastructure on top of automotive Ethernet networks. A middleware is proposed that enables QoS aware service-oriented communication between software components, where the prescribed behavior of each software component is defined by Assume/Guarantee (A-G) contracts. To enable the use of COTS components, that are often not sufficiently verified for the use in automotive systems, the middleware monitors the communication behavior of components and verifies it against the components A/G contract. A violation of the allowed communication behavior then triggers adaption processes in the system while the impact on other communication is minimized. The applicability of the approach is demonstrated by a case study that utilizes a prototype implementation of the proposed approach.
Matthias Becker 0004, Zhonghai Lu, Dejiu Chen
IECON1
2018 Timing Analysis Driven Design-Space Exploration of Cause-Effect Chains in Automotive Systems
abstract
Model-based development and component-based software engineering have emerged as a promising approach to deal with enormous software complexity in automotive systems. This approach supports the development of software architectures by interconnecting (and reusing) software components (SWCs) at various abstraction levels. Automotive software architectures are often modeled with chains of SWCs, also called cause-effect chains that are constrained by timing requirements. Based on the variations in activation patterns of SWCs, a single model of a cause-effect chain at a higher abstraction level can conform to several valid refined models of the chain at a lower abstraction level, which is closer to the system implementation. As a consequence, the total number of valid implementation-level models generated by the existing techniques increases exponentially, thereby significantly increasing the runtime of the timing analysis engines and liming the scalability of the existing techniques. This paper computes an upper bound on the activation pattern combinations that may result from a system of cause-effect chains in a given high-level model of the software architecture. An efficient algorithm is presented that traverses only a reduced number of possible combinations of the cause-effect chains, resulting in the timing analysis of a significantly lower number of implementation-level models of the software architecture. A proof of concept is provided by conducting a case study that shows significant reduction in the runtime of timing analysis engines, i.e., the timing behavior of the considered system is verified by performing the timing analysis of only 27% of all possible combinations of the cause-effect chains.
Matthias Becker 0004, Saad Mubeen
IECON1
2017 A tighter recursive calculus to compute the worst case traversal time of real-time traffic over NoCs
abstract
Network-on-Chip (NoC) is a communication subsystem which has been widely utilized in many-core processors and system-on-chips in general. In this paper, we focus on a Round-Robin Arbitration (RRA) based wormhole-switched NoC which is a common architecture used in most of the existing implementations. In order to execute real-time applications on such a NoC based platform, a number of given real-time requirements need to be fulfilled. One of the most typical requirements is schedulability which refers to determining if real-time packets can be delivered within the given time durations. Timing analysis is a common tool to verify the schedulability of a real-time system. Unfortunately, the existing timing analyses of RRA-based NoCs either provide too pessimistic estimates which results in overly allocated resources or require a large amount of processing which limits the applicability in reality. Therefore, in this paper, we present an improved timing analysis, aiming to provide more accurate estimates along with acceptable computation time. From the evaluation results, we can clearly observe the improvement achieved by the proposed timing analysis.
Meng Liu 0001, Matthias Becker 0004, Moris Behnam, Thomas Nolte
ASP-DAC2
2017 Using segmentation to improve schedulability of RRA-based NoCs with mixed traffic
abstract
Network-on-Chip (NoC) is the interconnect of choice for many-core processors and system-on-chips in general. Most of the existing NoC designs focus on the performance with respect to average throughput, which makes them less applicable for real-time applications especially when applications have hard timing requirements on the worst-case scenarios. In this paper, we focus on a Round-Robin Arbitration (RRA) based wormhole-switched NoC which is a common architecture used in most of the existing implementations. We propose a novel segmentation algorithm targeting RRA-based NoCs in order to improve the schedulability of real-time traffic without modifying the hardware architecture. Additionally, we also address the problem of transmitting both real-time traffic and best-effort traffic in the same NoC. The proposed solutions aim to provide timing guarantees to real-time traffic and achieve low latency for best-effort traffic. According to the evaluation results, the proposed segmentation solution can significantly improve the schedulability of the whole network.
Meng Liu 0001, Matthias Becker 0004, Moris Behnam, Thomas Nolte
ASP-DAC2
2017 Buffer-Aware Analysis for Worst-Case Traversal Time of Real-Time Traffic over RRA-based NoCs
abstract
Network-on-Chip (NoC) is a communication subsystem which has been widely utilized in many-core processors and system-on-chips in general. In order to execute time-critical applications on a NoC-based platform, the timing behavior of the network needs to be predicted during system design. One of the most important timing requirements is regarding schedulability, which refers to determining if a real-time packet can be delivered within a specific time duration. To verify the fulfillment of such timing requirement, a proper timing analysis is mandatory. Our work focuses on a Round-Robin Arbitration (RRA) based wormhole-switched NoC, which is a common architecture used in many of the existing implementations. Recursive Calculus (RC) is one of the existing analysis approaches for RRA-based NoCs which has been utilized in many research works. However, RC does not take buffer-effects into account. As a result, while performing RC on most of the existing RRA-based NoC designs, it can produce unsafe estimates which is not acceptable for time-critical systems. In this paper, we identify the optimistic problem of RC, and we propose a Revised Recursive Calculus (RRC) which extends RC by considering buffer-effects as well as supporting packetization.
Meng Liu 0001, Matthias Becker 0004, Moris Behnam, Thomas Nolte
PDP2
2017 Partitioning and Analysis of the Network-on-Chip on a COTS Many-Core Platform
abstract
Many-core processors can provide the computational power required by future complex embedded systems. However, their adoption is not trivial, since several sources of interference on COTS many-core platforms have adverse effects on the resulting performance. One main source of performance degradation is the contention on the Network-on-Chip (NoC), which is used for communication among the compute cores via the off-chip memory. Available analysis techniques for the traversal time of messages on the NoC do not consider many of the architectural features found on COTS platforms. In this work, we target a state-of-the-art many-core processor, the Kalray MPPAR®. A novel partitioning strategy for reducing the contention on the NoC is proposed. Further, we present an analysis technique dedicated to the proposed partitioning strategy, which considers all architectural features of the COTS NoC. Additionally, it is shown how to configure the parameters for flow-regulation on the NoC, such that the Worst-Case Traversal Time (WCTT) is minimal and buffers never overflow. The benefits of our approach are evaluated based on extensive experiments that show that contention is significantly reduced compared to the unconstrained case, while the proposed analysis outperforms a state-of-the-art analysis for the same platform. An industrial case study shows the tightness of the proposed analysis.
Matthias Becker 0004, Borislav Nikolic, Dakshina Dasari, Benny Akesson, Vincent Nélis, Moris Behnam, Thomas Nolte
RTAS1
2017 A generic framework facilitating early analysis of data propagation delays in multi-rate systems (Invited paper)
abstract
A majority of multi-rate real-time systems are constrained by a multitude of timing requirements, in addition to the traditional deadlines on well-studied response times. This means, the timing predictability of these systems not only depends on the schedulability of certain task sets but also on the timely propagation of data through the chains of tasks from sensors to actuators. In the automotive industry, four different timing constraints corresponding to various data propagation delays are commonly specified on the systems. This paper identifies and addresses the source of pessimism as well as optimism in the calculations for one such delay, namely the reaction delay, in the state-of-the-art analysis that is already implemented in several industrial tools. Furthermore, a generic framework is proposed to compute all the four end-to-end data propagation delays, complying with the established delay semantics, in a scheduler and hardware-agnostic manner. This allows analysis of the system models already at early development phases, where limited system information is present. The paper further introduces mechanisms to generate job-level dependencies, a partial ordering of jobs, which need to be satisfied by any execution platform in order to meet the data propagation timing requirements. The job-level dependencies are first added to all task chains of the system and then reduced to its minimum required set such that the job order is not affected. Moreover, a necessary schedulability test is provided, allowing for varying the number of CPUs. The experimental evaluations demonstrate the tightness in the reaction delay with the proposed framework as compared to the existing state-of-the-art and practice solutions.
Matthias Becker 0004, Saad Mubeen, Dakshina Dasari, Moris Behnam, Thomas Nolte
RTCSA1
2017 End-to-end timing analysis of cause-effect chains in automotive embedded systems
abstract
Automotive embedded systems are subjected to stringent timing requirements that need to be verified. One of the most complex timing requirement in these systems is the data age constraint. This constraint is specified on cause-effect chains and restricts the maximum time for the propagation of data through the chain. Tasks in a cause-effect chain can have different activation patterns and different periods, that introduce over- and under-sampling effects, which additionally aggravate the end-to-end timing analysis of the chain. Furthermore, the level of timing information available at various development stages (from modeling of the software architecture to the software implementation) varies a lot, the complete timing information is available only at the implementation stage. This uncertainty and limited timing information can restrict the end-to-end timing analysis of these chains. In this paper, we present methods to compute end-to-end delays based on different levels of system information. The characteristics of different communication semantics are further taken into account, thereby enabling timing analysis throughout the development process of such heterogeneous software systems. The presented methods are evaluated with extensive experiments. As a proof of concept, an industrial case study demonstrates the applicability of the proposed methods following a state-of-the-practice development process.
Matthias Becker 0004, Dakshina Dasari, Saad Mubeen, Moris Behnam, Thomas Nolte
J. Syst. Archit.1
2017 Using non-preemptive regions and path modification to improve schedulability of real-time traffic over priority-based NoCs
abstract
Network-on-Chip (NoC) is a preferred communication medium for massively parallel platforms. Fixed-priority based scheduling using virtual-channels is one of the promising solutions to support real-time traffic in on-chip networks. Most of the existing works regarding priority-based NoCs use a flit-level preemptive scheduling. Under such a mechanism, preemptions can only happen between the transmissions of successive flits but not during the transmission of a single flit. In this paper, we present a modified framework where the non-preemptive region of each NoC packet increases from a single flit. Using the proposed approach, the response times of certain traffic flows can be reduced, which can thus improve the schedulability of the whole network. As a result, the utilization of NoCs can be improved by admitting more real-time traffic. Schedulability tests regarding the proposed framework are presented along with the proof of the correctness. Additionally, we also propose a path modification approach on top of the non-preemptive region based method to further improve schedulability. A number of experiments have been performed to evaluate the proposed solutions, where we can observe significant improvement on schedulability compared to the original flit-level preemptive NoCs.
Meng Liu 0001, Matthias Becker 0004, Moris Behnam, Thomas Nolte
Real Time Syst.2
2016 Contention-Free Execution of Automotive Applications on a Clustered Many-Core Platform
abstract
Next generations of compute-intensive real-time applications in automotive systems will require more powerful computing platforms. One promising power-efficient solution for such applications is to use clustered many-core architectures. However, ensuring that real-time requirements are satisfied in the presence of contention in shared resources, such as memories, remains an open issue. This work presents a novel contention-free execution framework to execute automotive applications on such platforms. Privatization of memory banks together with defined access phases to shared memory resources is the backbone of the framework. An Integer Linear Programming (ILP) formulation is presented to find the optimal time-triggered schedule for the on-core execution as well as for the access to shared memory. Additionally a heuristic solution is presented that generates the schedule in a fraction of the time required by the ILP. Extensive evaluations show that the proposed heuristic performs only 0.5% away from the optimal solution while it outperforms a baseline heuristic by 67%. The applicability of the approach to industrially sized problems is demonstrated in a case study of a software for Engine Management Systems.
Matthias Becker 0004, Dakshina Dasari, Borislav Nikolic, Benny Akesson, Vincent Nélis, Thomas Nolte
ECRTS1
2016 Tighter time analysis for real-time traffic in on-chip networks with shared priorities
abstract
The Network-on-Chip (NoC) is the preferred interconnection medium for massively parallel platforms. Targeting real-time applications, fixed-priority based NoCs with virtual channels have been proposed as a promising solution. In order to verify if specific time requirements can be satisfied, schedulability tests are typically used. Several analysis approaches have been proposed targeting priority-based NoCs. However, due to the approximation considered in the analyses, the results may involve a large amount of pessimism. The applicability of the analyses is thus limited in practice. In this paper, we identify a number of properties of NoCs with shared priorities. An improved time analysis is proposed where pessimism can be significantly reduced for many cases. In order to evaluate the proposed analysis, a number of experiments have been generated along with a case study based on an automotive application. The improvement can be clearly observed from the evaluation results.
Meng Liu 0001, Matthias Becker 0004, Moris Behnam, Thomas Nolte
NOCS2
2016 Synthesizing Job-Level Dependencies for Automotive Multi-rate Effect Chains
abstract
Today's automotive embedded systems comprise a multitude of functionalities, many with complex timing requirements. Besides task specific timing requirements, such applications often have timing requirements for the propagation of data through a chain of tasks. An important metric for control applications is the data age, which is addressed in this paper. The analysis of such systems is non-trivial because tasks involved in the data propagation may execute at different periods, which leads to over and undersampling within one chain. This paper presents a novel method to compute worst-and best-case end-to-end latencies for such systems. A second contribution synthesizes job-level dependencies for such task sets in a way that data paths which exceed the age constraint are eliminated. An extensive evaluation is performed on synthetic task sets and the applicability to industrial applications is demonstrated in a case study.
Matthias Becker 0004, Dakshina Dasari, Saad Mubeen, Moris Behnam, Thomas Nolte
RTCSA1
2016 Scheduling Real-Time Packets with Non-preemptive Regions on Priority-Based NoCs
abstract
Network-on-Chip (NoC) is a preferred communication medium for massively parallel platforms. Fixed-priority based scheduling using virtual-channels is one of the promising solutions to support real-time traffic in on-chip networks. Most of the existing NoC implementations which can support fixed-priority based scheduling use a flit-level preemptive scheduling. Under such a mechanism, preemptions can happen between the transmissions of successive flits. In this paper, we present a modified framework where the non-preemptive region of each NoC packet increases from a single flit. Using the proposed approach, the response times of certain packet flows can be reduced, which can thus improve the schedulability of the whole network. As a result, the utilization of NoCs can be improved by admitting more real-time traffic. Schedulability tests regarding the proposed framework are presented along with the proof of the correctness. Moreover, a number of experiments as well as a case study based on an automotive application have been generated, where we can clearly observe the improvement of our solution compared to the original flit-level preemptive NoC.
Meng Liu 0001, Matthias Becker 0004, Moris Behnam, Thomas Nolte
RTCSA2
2016 Real-Time Capabilities of HSA Compliant COTS Platforms
abstract
During recent years, the interest in using heterogeneous computing architecture in industrial applications has increased dramatically. These architectures provide the computational power that makes them attractive for many industrial applications. However, most of these existing heterogeneous architectures suffer from the following limitations: difficulties of heterogeneous parallel programming and high communication cost between the computing units. To overcome these disadvantages, several leading hardware manufacturers have formed the HSA Foundation to develop a new hardware architecture: Heterogeneous System Architecture (HSA). In this paper, we investigate the suitability of using HSA for real-time embedded systems. A preliminary experimental study has been conducted to measure massive computing power and timing predictability of HSA.
Nandinbaatar Tsog, Matthias Becker 0004, Marcus Larsson, Fredrik Bruhn, Moris Behnam, Mikael Sjödin
RTSS2
2016 A dependency-graph based priority assignment algorithm for real-time traffic over NoCs with shared virtual-channels
abstract
The Network-on-Chip (NoC) is the on-chip interconnection medium of choice for modern massively parallel processors and System-on-Chip (SoC) in general. Fixed-priority based preemptive scheduling using virtual-channels is a solution to support real-time communications in on-chip networks. Targeting the priority assignment problem in the context of NoCs, heuristic based priority assignment algorithms are more practical, due to the exponentially increased search space as the number of flows goes up. In our previous work, we have proposed a graph-based heuristic priority assignment algorithm (called GHSA) for NoC communications, where we show that taking the dependencies between flows into account can significantly reduce the search space. However, GHSA only works for NoCs with distinct priorities. Routers in such type of platforms may have a large amount of buffer cost when the number of flows is high. The applicability can thus be limited in reality. One solution to reduce the buffer cost is to allow priority sharing of different flows. In this paper, we propose a dependency-graph based priority assignment algorithm (called eGHSA) targeting NoCs with shared virtual-channels. A number of experiments as well as a case study based on an automotive application are generated, which clearly show that eGHSA improves the efficiency compared to the existing solution in the literature.
Meng Liu 0001, Matthias Becker 0004, Moris Behnam, Thomas Nolte
WFCS2
2016 Towards automated deployment of IEC 61131-3 applications on multi-core systems
abstract
The IEC 61131-3 standard, a widely used standard in the automation industry, defines various programming languages for programmable logic controllers. Today, the open source tools that comply with this standard do not support deployment of the applications on multi-core platforms. In this paper, we introduce a novel multi-step approach that aims to support automatic deployment of the automation control applications, developed using the IEC 61131-3 standard, to multi-core platforms. In the first step, the generated sequential code is partitioned. In the second step, the partitioned code is allocated to tasks while the tasks are mapped to various cores, without violating the dependencies, synchronization and communication constraints in the application. In order to provide a proof of concept, we develop a prototype by extending an existing tool that complies with the standard. We also perform a case study and a preliminary evaluation of the prototype.
Saad Mubeen, Matthias Becker 0004, Xiaosha Zhao, Lingjian Gan, Moris Behnam, Thomas Nolte
WFCS2
2015 Investigation on AUTOSAR-Compliant Solutions for Many-Core Architectures
abstract
As of today, AUTOSAR is the de facto standard in the automotive industry, providing a common software architecture and development process for automotive applications. While this standard is originally written for singlecore operated Electronic Control Units (ECU), new guidelines and recommendations have been added recently to provide support for multicore architectures. This update came as a response to the steady increase of the number and complexity of the software functions embedded in modern vehicles, which call for the computing power of multicore execution environments. In this paper, we enumerate and analyze the design options and the challenges of porting AUTOSAR-based automotive applications onto multicore platforms. In particular, we investigate those options when considering the emerging many-core architectures that provide a more "scalable" environment than the traditional multicore systems. Such platforms are suitable to enable massive parallel execution, and their design is more suitable for partitioning and isolating the software components.
Matthias Becker 0004, Dakshina Dasari, Vincent Nélis, Moris Behnam, Luís Miguel Pinho, Thomas Nolte
DSD1
2015 A many-core based execution framework for IEC 61131-3
abstract
Programmable logic controllers are widely used for the control of automation systems. The standard IEC 61131-3 defines the execution model as well as the programming languages for such systems. Nowadays, actuators and sensors connect to the programmable logic controller via automation buses. While such buses, as well as the sensors and actuators, become more and more powerful, a shift away from the current distributed operation of automation systems, close to the field level, becomes possible. Instead, execution of complex control functions can be relocated to more powerful hardware, and technologies. This paper presents an execution framework for IEC 61131-3, based on a many-core processors. The presented execution model exploits the characteristics of the IEC 61131-3 applications as well as the characteristics of the many-core processor, yielding a predictable execution. We present the platform architecture and an algorithm to allocate a number of IEC 61131-3 conform applications. Experimental as well as simulation based evaluation is provided.
Matthias Becker 0004, Kristian Sandström, Moris Behnam, Thomas Nolte
IECON1
2014 Limiting temperature gradients on many-cores by adaptive reallocation of real-time workloads
abstract
The advent of many-core processors came with the increase in computational power needed for future applications. However new challenges arrived at the same time, especially for the real-time community. Each core on such a processor is a heat source and uneven usage can lead to hot spots on the processor, affecting its lifetime and reliability. For real-time systems, it is therefore of paramount importance to keep the temperature differences between the individual cores below critical values, in order to prevent premature failure of the system. We argue that this problem can not be solved by traditional approaches, since the growing number of cores makes them intractable. We rather argue to split the problem in the spacial domain and control the temperature on core level. The cores control their temperature by rearranging the load in a predictable manner during runtime. To achieve this, a feedback controller is implemented on each core. We conclude our work with a simulation based evaluation of the proposed approach comparing its performance against a previously presented algorithm.
Matthias Becker 0004, Kristian Sandström, Moris Behnam, Thomas Nolte
ETFA1