EDBT 2026 Demo / reviewers in the wild / expert
Martin Schoeberl
dblp:29/3006
· DBLP profile ↗
102ranked-venue papers
32as first author
29since 2021 · last 2026
0000-0003-2366-382XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 53 · 22 first-author · 16 since 2021Software engineering, systems software and programming languages · 9 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deployment of Quantized Deep Noise Suppression on Real-Time Edge Platforms
Alessandro Cerioli, Tórur Biskopstø Strøm, Clement Laroche, Tobias Piechowiak, Luca Pezzarossa, Martin Schoeberl |
ISORC | 6 |
| 2026 | Rigorous Design of Time-Predictable Embedded Systems
Ehsan Khodadad, Luca Pezzarossa, Martin Schoeberl |
ISORC | 3 |
| 2026 | SlimFlit: A simple Network-on-Chip for Real-Time Systems
Tjark Petersen, Voica Gavrilut, Luca Pezzarossa, Martin Schoeberl |
ISORC | 4 |
| 2026 | Shared caches for mixed-criticality 5G radio base stationabstractThe advancement of telecommunication technology is necessary to support the rapid development of societies. The latest standard in telecommunications, 5G, offers unprecedented communication speeds, reach, and quality. The 5G standard supports both critical and non-critical telecommunications. The baseband units—handling wireless signal transmission and reception—divide hardware resources to ensure that high- and low-criticality communication tasks do not interfere with each other. This results in underutilized hardware platforms with high redundancy. In this paper, we propose two level 2 cache architectures that allow a single system to handle high- and low-criticality tasks and improve resource utilization while prioritizing critical tasks. The contention-tracking cache tracks contention among tasks of differing criticality for shared resources and locks those resources for critical tasks once a contention limit is exceeded. The criticality timeout cache sets timers for cache lines accessed by critical tasks and decrements them as long as the lines remain unused. Non-critical tasks are not allowed to evict these lines until the timer has finished. We evaluate the behavior of the proposed cache designs using a simulation framework and a real-world workload from a 5G baseband unit. Additionally, we implement the proposed caches in hardware, along with a baseline PLRU cache, and evaluate their performance and resource usage using the T-CREST platform. The simulation and performance results confirm that the proposed cache architectures can prioritize critical memory access, while allowing non-critical tasks to utilize the cache’s free space. Emad Jacob Maroun, Arijus Grotuzas, Martin Schoeberl |
J. Syst. Archit. | 3 |
| 2025 | A Structured Approach to Verification of Digital Hardware in ScalaabstractFunctional verification accounts for a significant portion of the design effort in modern digital hardware development. As projects grow in complexity, maintaining and extending verification code becomes increasingly difficult, particularly in collaborative environments. This calls for a methodology that defines a clear structure and promotes reuse through modular, composable testbench components. In this paper, we present a Scala-based verification framework that adopts a structured approach to building modular and reusable testbenches, inspired by the Universal Verification Methodology (UVM). We analyze the core mechanisms through which UVM achieves modularity and reusability, and identify a minimal subset that provides equivalent functionality with reduced complexity. The result is a lightweight verification framework in Scala 3 using Verilator as a backend, which allows for simple unit-test-style testing as well as complex UVM-style testbench environments. Tjark Petersen, Luca Pezzarossa, Martin Schoeberl |
DSD | 3 |
| 2025 | Time-Predictable Deep Noise Suppression on an Edge DeviceabstractHearing aids and remote conference systems benefit from noise reduction. Current noise reduction approaches include machine-learning models that run on edge devices like hearing aids, AirPods, or headsets. Although not a safety-critical application, audio processing is a real-time application. We present a real-time enabled solution of speech enhancement with generation of$\mathbf{C}$code for embedded devices, executing on a real-time processor, and analyzing the worst-case execution time for that application. Using the Patmos processor and the Platin WCET analysis tool, we can guarantee that we process noise canceling within the given deadline. Alessandro Cerioli, Tórur Biskopstø Strøm, Clement Laroche, Tobias Piechowiak, Luca Pezzarossa, Martin Schoeberl |
ISORC | 6 |
| 2025 | Simulating Contention and Timeout Caches for a Mixed-Criticality 5G Radio Base StationabstractTelecommunication is a critical driver of economic and social development. 5G technologies are state-of-the-art in telecommunication, setting strong and open-ended requirements for implementing systems. Current systems for implementing baseband technologies in 5 G depend on hardware separation to ensure high-criticality tasks and low-criticality tasks do not interfere in such a way as to violate guarantees. To allow for the merging of high- and low-criticality systems into one, this paper presents two level-2 cache architectures: The contention tracking cache tracks contention events between highand low-criticality tasks, blocking any further contention if a specified limit is reached. The criticality timeout cache associates a timer with each cache line for high-criticality tasks that counts down as long as the line is not reused. During the countdown, low-criticality tasks are prohibited from evicting high-criticality cache lines. When the timer runs out, this prohibition is lifted. A simulation framework is employed to accurately model the behavior of cores accessing a memory hierarchy with various cache types. Simulation statistics demonstrate that the proposed cache architectures effectively prioritize memory accesses for critical tasks while allowing non-critical tasks to utilize any available cache space. Emad Jacob Maroun, Martin Schoeberl |
ISORC | 2 |
| 2025 | Optimized Constant Execution Time CodeabstractSingle-path code aims to make WCET analysis easier by eliminating data-dependent control flow. To completely negate the need for WCET analysis, single-path code must also eliminate execution-time variability from memory accesses. To be practically useful, single-path code must be optimized to be competitive with traditional WCET-analyzed code. This paper summarizes the work in Emad Jacob Maroun's PhD dissertation titled”Compiling for Time-Predictability and Performance“. Memory access compensation ensures that singlepath code exhibits constant execution time. The generated code is optimized using an improved transformation that uses generic allocators for general-purpose and predicate registers. The repetition dominance relation is used to reduce unnecessary code execution. Lastly, a heuristic list scheduler enables single-path code to utilize the second issue slot of a dual-issue processor. In addition to achieving constant execution times on a timepredictable processor, the results show varying but significant improvements of up to 145 % in performance and a reduced code size of up to 28 %. Compared to WCET-analyzed traditional code, single-path code is mostly competitive while outright superior in several cases. However, pathological cases of poor performance are still observed. Emad Jacob Maroun, Martin Schoeberl, Peter P. Puschner |
ISORC | 2 |
| 2025 | Quasi-Static Scheduling for Deterministic Timed Concurrent Models on Multi-Core HardwareabstractTo design performant, expressive, and reliable cyber-physical systems (CPSs), researchers extensively perform quasi-static scheduling for concurrent models of computation (MoCs) on multi-core hardware. However, these quasi-static scheduling approaches are developed independently for their corresponding MoCs, despite commonality in the approaches. To help generalize the use of quasi-static scheduling to new and emerging MoCs, this article proposes a unified approach for a class of deterministic timed concurrent models (DTCMs), including prominent models such as synchronous dataflow (SDF), Boolean-controlled dataflow (BDF), scenario-aware dataflow (SADF), and Logical Execution Time (LET). In contrast to scheduling techniques tailored exclusively to specific MoCs, our unified approach leverages a common intermediate formalism called state space finite automata (SSFA), bridging the gap between high-level MoCs and executable schedules. Once identified as DTCMs, new MoCs can directly adopt SSFA-based scheduling, significantly easing adoption. We show that quasi-static schedules facilitated by SSFA are provably free from timing anomalies and enable straightforward worst-case makespan analysis. We demonstrate the approach using the reactor model—an emerging discrete-event MoC—programmed using the Lingua Franca ( LF ) language. Experiments show that quasi-statically scheduled LF programs exhibit lower runtime overhead compared to the dynamically scheduled LF programs, and that the analyzable worst-case makespans enable compile-time deadline checking. Shaokai Lin, Erling Rennemo Jellum, Mirco Theile, Tassilo Tanneberger, Binqi Sun, Chadlia Jerad, Yimo Xu, Guangyu Feng, Magnus Mæhlum, Jian-Jia Chen, Martin Schoeberl, Linh T. X. Phan, Jerónimo Castrillón, Sanjit A. Seshia, Edward A. Lee |
ACM Trans. Embed. Comput. Syst. | 11 |
| 2024 | Hardware Generators with ChiselabstractMost digital hardware is described in hardware description languages, such as VHDL and (System)Verilog. These languages provide limited programming models for hardware construction despite receiving regular updates and extensions. Chisel defines itself as a hardware construction language, which means it shall permit more than the mere description of digital circuits. However, programmatic hardware generation is not new. Scripting languages like Perl generate VHDL or Verilog code from sources like Excel spreadsheets. Chisel, embedded in the general-purpose language Scala, lends itself to writing hardware generators in that language. We consider this Chisel-Scala ecosystem an ideal starting point for programming hardware generators and illustrate this point with examples using various programming models. We are confident that proven technologies from the software development world can be leveraged in the hardware design domain to improve hardware designers' productivity to build the next billion transistor chips. Martin Schoeberl, Hans Jakob Damsgaard, Luca Pezzarossa, Oliver Keszöcze, Erling Rennemo Jellum |
DSD | 1 |
| 2024 | Towards Lingua Franca on the Patmos ProcessorabstractReal-time embedded systems demand higher reliability than any other computer systems. These systems require special modeling paradigms to satisfy time constraints. This paper introduced a design method by combining T-CREST, a time-predictable multi-core hardware, with Lingua Franca, a coordination framework that generates deterministic time-predictable code. We executed a Lingua Franca piece of software on T-CREST platform and performed preliminary experiments demonstrating its correct functionality. Ehsan Khodadad, Luca Pezzarossa, Martin Schoeberl |
ISORC | 3 |
| 2024 | Two-Step Register Allocation for Implementing Single-Path CodeabstractRegister allocation is a crucial step in the compilation pipeline that decides what program values occupy which physical registers. Single-path code’s use of predicated instructions instead of branching control-flow means register allocation must also allocate predicate registers. In this paper, we improve the original single-path transformation to allow generic register allocators to allocate predicate registers. Our improved transformation splits register allocation into two. First, the general-purpose registers are allocated as usual using a generic register allocator. Then, the main steps of the single-path transformation are performed while still using virtual predicate registers. Lastly, register allocation is rerun using the generic allocator to allocate the predicate registers. Our results show the improved single-path transformation increasing performance by up to 80 % and reducing code size by up to 43 % compared to the original transformation that uses a custom predicate allocator. Emad Jacob Maroun, Martin Schoeberl, Peter P. Puschner |
ISORC | 2 |
| 2024 | Exploration of Network Interface Architectures for a Real-Time Network-on-ChipabstractNetwork interfaces play a central role in multicore architectures that use a network-on-chip for communication. Network interface designs have not received much attention in the research community despite this central role.This paper explores different network interface configurations for a real-time network-on-chip architecture. We evaluate the effects of different FIFO queue organizations on the bandwidth and maximum latency of messages in a time-division multiplexing network-on-chip. Martin Schoeberl |
ISORC | 1 |
| 2024 | Predictable and optimized single-path code for predicated processorsabstractSingle-path code is a code generation technique for real-time systems that reduces execution time variability. However, doing so can incur significant execution-time overhead and does not guarantee constant execution times. In this paper, we address the performance challenges of single-path code and solve the variability issue. We present the repetition dominance relation to identify and optimize code blocks that are always executed a fixed number of times. We show that single-path code’s instructions are uniquely easy to schedule, and we explore an extension to the Patmos architecture that allows additional instruction types in the second issue slot. Lastly, we present two techniques for ensuring that functions always perform the same number of accesses to memory, resulting in programs with constant execution time. We compare the performance of single-path code to that of statically analyzed traditional code. Our results show that single-path code’s performance is mostly competitive while outright superior in several cases. However, pathological cases of poor performance are still observed. Emad Jacob Maroun, Martin Schoeberl, Peter P. Puschner |
J. Syst. Archit. | 2 |
| 2024 | Codesign of Reactor-Oriented Hardware and Software for Cyber-Physical SystemsabstractModern cyber-physical systems often make use of heterogeneous systems-on-chip with reconfigurable logic to provide adequate computing power and flexible I/O. However, modeling, verifying, and implementing the computations spanning CPUs and reconfigurable logic are still challenging. The hardware and software components are often designed by different teams and at different levels of abstraction, making it hard to reason about the resulting computation. We propose to lift both hardware and software design to the same level of abstraction by using the Lingua Franca coordination language. Lingua Franca is based on a sparse synchronous model that allows modeling concurrency and timing while keeping a sequential model for the actual computation. We define hardware reactors as a subset of the reactor model of computation underlying Lingua Franca. We also present and evaluate reactor-chisel, a hardware runtime implementing the semantics of hardware reactors, and an extension to the Lingua Franca compiler enabling reactor-oriented hardware–software codesign. Erling Rennemo Jellum, Martin Schoeberl, Edward A. Lee, Milica Orlandic |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2023 | On the Feasibility of using FPGA's for Efficient Topology OptimizationabstractTopology Optimization is a class of structural optimization problems, where the classical goal is to find the best layout of a structure ensuring that it can withstand a set of prescribed forces, while being as light as possible. Since this class of optimization problems are computationally expensive, a vast amount of research is directed at how to increase the speed at which these problems can be solved. Previous work on accelerating topology optimization problems includes designing more efficient algorithmic approaches and better utilizing hardware resources. However, to the best of our knowledge, no previous attempts have been made to accelerate topology optimization using a hardware accelerator implemented on an FPGA. This paper presents a hardware accelerator for topology optimization, designed to solve compliance minimization problems in three dimensions. The accelerator is implemented as an application-specific instruction set processor, using a custom instruction set architecture designed specifically for the minimum compliance problem. The developed accelerator is able to solve problems 4.8–8.2 times faster than a modern computer, while operating at a fraction of the clock frequency. Although the current accelerator has only been proven to work on coarse meshes, it is reasonable to assume that the speedup will also carry over to larger optimization problems. This indicates that FPGA-based acceleration of topology optimization problems is viable. Kasper Juul Hesse Rasmussen, Martin Schoeberl, Niels Aage, Erik Träff |
DSD | 2 |
| 2023 | FPGA-tidbits: Rapid Prototyping of FPGA Accelerators in ChiselabstractWith increasingly complex workloads and the end of Dennard scaling, the need for heterogeneous computing is becoming apparent. SoC FPGAs (System-on-chip Field-Programmable Gate Arrays) are a promising solution to this need. They combine the versatility of CPU s and the reconfigurability and high performance of FPGAs. SoC FPGAs have received a great deal of attention in recent years from both academia and chipmakers. However, the task of hardware-software codesign, posed by these platforms, remains challenging. This is partly due to the lack of vendor-neutral abstractions for building and evaluating designs. Chisel is a promising hardware construction language based on the idea of writing hardware generators. In this paper, we present FPGA-tidbi ts, an open-source, vendor-neutral Chisel library for rapid prototyping of accelerators for SoC FPG As. Erling Rennemo Jellum, Yaman Umuruglu, Milica Orlandic, Martin Schoeberl |
DSD | 4 |
| 2023 | Compiler-Directed Constant Execution Time on Flat Memory SystemsabstractTime predictability is a central requirement for real-time systems. The correct behavior of such a system can only be achieved if the results of programs are ready in time to affect the environment. Execution times of modern systems can vary for many reasons, meaning complex analyses must be performed to ensure that the execution time is bounded and that a task always finishes before its deadline. Care must also be taken to ensure that nefarious actors do not exploit the varying execution time to compromise the system’s integrity. Avoiding variable execution times can greatly simplify systems, is inherently more secure, and eliminates the need for complex analyses. In this paper, we first argue for the value of having programs with constant execution times. We then show how the memory system around a processing core can affect execution times even on systems without intermediate storage like caches or scratch-pads. We present automatic compiler techniques for generating constant execution time programs and evaluate their implementation on the Patmos architecture. We show that combining our two compensation techniques is generally superior to either on their own. We compare the performance of our implementation to the estimates produced by the Platin worst-case execution time analyzer. While our implementation significantly impacts performance, it is generally manageable and has the potential for comparable execution times. Emad Jacob Maroun, Martin Schoeberl, Peter P. Puschner |
ISORC | 2 |
| 2022 | Keynote SpeakersabstractReal-time systems need time-predictable computers to be able to guarantee that computation can be performed within a given deadline. For worst-case execution time analysis, we need detailed knowledge of the processor and memory architecture. Providing the design of a processor in open-source enables the development of worst-cease execution time analysis tools without the unsafe reverse engineering of processor architectures. Open-source software is currently the basis of many Internet services, e.g., an Apache web server running on top of Linux with a web application written in Java. Furthermore, for most programming languages in use today, there are an opensource compilers available. However, hardware designs are seldom published in open-source. Furthermore, many artifacts developed in research, especially hardware designs, are not published in open-source. The two main arguments formulated against publishing research in open source are: Martin Schoeberl |
DSD | 1 |
| 2022 | Open-Source Research on Time-predictable Computer ArchitectureabstractReal-time systems need time-predictable computers to guarantee that computation can be performed within a given deadline. For worst-case execution time analysis we need detailed knowledge of the processor and memory architecture. Providing the design of a processor in open source enables the development of worst-case execution time analysis tools without the unsafe reverse engineering of processor architectures. As an example project, we will present T-CREST, an open-source, time-predictable multicore platform for real-time systems. The project started within an EU-funded project. Most artifacts have been put into open-source, greatly simplifying the coop-eration and also the further adaption of T-CREST in further research work. Furthermore, open-source enables reproducibility and therefore increases the confidence Martin Schoeberl |
DSD | 1 |
| 2022 | Enabling Coverage-Based Verification in ChiselabstractEver-increasing performance demands are pushing hardware designers towards designing domain-specific accelerators. This has created a demand for improving the overall efficiency of the hardware design and verification cycles. The design efficiency was improved with the introduction of Chisel. However, verification efficiency has yet to be tackled. One method that can increase verification efficiency is the use of various types of coverage measures. In this paper, we present our open-source, coverage-related verification tools targeting digital designs described in Chisel. Specifically, we have created a new method allowing for statement coverage at an intermediate representation of Chisel, and several methods for gathering functional coverage directly on a Chisel description. Andrew Dobis, Hans Jakob Damsgaard, Enrico Tolotto, Kasper Juul Hesse Rasmussen, Tjark Petersen, Martin Schoeberl |
ETS | 6 |
| 2022 | Timing Analysis of TSN-Enabled OPC UA PubSubabstractIndustrial automation is changing towards a flat and highly interconnected architecture, with requirements for end-to-end real-time enabled machine-to-machine communications. Technologies such as time-sensitive networking (TSN) and OPC Unified Architecture (OPC UA) publish-subscribe provide the necessary features. While TSN is well explored, OPC UA's execution time behavior remains unknown. This article presents findings made while extending the open62541 OPC UA Pub Sub stack with the 802.1q VLAN tag to enable IEEE 802.1Qbv time-aware scheduling. The results include end-to-end timing measures and worst-case execution time analyses considering various payloads. Time-predictable T-CREST platforms host the publisher and subscriber, and a TSN network handles message transmission. The paper concludes by outlining further research focusing on dynamic memory access, buffer management, and the inclusion of non-priority access to the Ethernet port. Patrick Denzler, Thomas Frühwirth, Daniel Scheuchenstuhl, Martin Schoeberl, Wolfgang Kastner |
WFCS | 4 |
| 2022 | Comparing timed-division multiplexing and best-effort networks-on-chipabstractBest-effort (BE) networks-on-chips (NOCs) are usually preferred over time-division multiplexed (TDM) NOCs in multi-core platforms because they are work-conserving and have lower (zero-load) latency. On the other hand, BE NOCs are significantly more expensive to implement than TDM NOCs because of their virtual channel buffers, allocators/arbiters, and (credit-based) flow control; functionality that a TDM NOC avoids altogether. The objective of this paper is to compare the performance of BE and TDM NOCs, taking hardware cost into consideration. The networks are compared using graphs showing average latency as a function of offered load. For the BE NOCs, we use the BookSim simulator, and for the TDM NOCs, we derive a queuing theory model and an associated TDM NOC simulator. Through experiments with both router architectures, packet length, link width, and different traffic patterns, we show that for the same hardware cost, a TDM NOC can provide higher bandwidth and comparable latency. We also show that the packet length is the most important factor affecting the TDM period, which again is the primary factor affecting latency. The best TDM NOC design for BE traffic uses single flit packets, wide links/flits, and a router with two pipeline stages: link and router traversal. Jens Sparsø, Hans Jakob Damsgaard, Dimitrios Katsamanis, Martin Schoeberl |
J. Syst. Archit. | 4 |
| 2021 | Evaluating a Time-Triggered Runtime System by Distributing a Flight ControllerabstractWith the recent advancements in the Industrial Internet of Things and Industry 4.0, cyber-physical systems have become increasingly inter-connected. It is becoming a challenge to maintain the same quality-of-control and time-predictability of computation and communication required by safety-critical hard real-time systems as previously achieved through non-distributed architectures. This paper examines the problem of implementing and distributing a closed-loop command-control system over an Ethernet network with guaranteed timing bounds. To achieve bounded communication and computation time, we use an open-source software framework running on the T-CREST platform combined with a TTEthernet network star topology. We evaluate its quality-of-control performance in our experimental setup and compare the results against single-core and multi-core implementations. The proposed distributed time-triggered runtime system executes with jitter below 10µs and can perform a stable flight scenario as verified by the benchmark implementation. Eleftherios Kyriakakis, Jens Sparsø, Martin Schoeberl |
ETFA | 3 |
| 2021 | Experiences from Adjusting Industrial Software for Worst-Case Execution Time AnalysisabstractWorst-case execution time (WCET) analysis is a prevalent way to ensure the timely execution of programs in time-critical systems. With the advent of new technologies such as fog computing and time-sensitive networking (TSN), the interest in timing analysis has increased in industrial communication. This paper highlights experiences made while adjusting the publisher of the open62541 OPC UA stack to enable WCET analysis, following a simple process combined with the open-source platform T-CREST. The main challenges are the required knowledge about the code and the specific communication software characteristics like variable message sizes. Other findings indicate the need for other types of annotation for indirect recursion or callback functions. The paper provides the foundation for further research on adjusting the implementation of existing industrial communication protocols for WCET analysis. Patrick Denzler, Thomas Frühwirth, Andreas Kirchberger, Martin Schoeberl, Wolfgang Kastner |
ISORC | 4 |
| 2021 | Synchronizing Real-Time Tasks in Time-Triggered NetworksabstractIn order to guarantee end-to-end latency and minimal jitter in distributed real-time systems, it is necessary to provide tight synchronization between computation and communication. This requires time-predictable execution of tasks across all processing nodes, and the use of a network protocol that can provide a global time base and bounded communication latency. TTEthernet is one such industrial communication protocol. This paper investigates the synchronization of the task execution schedule with the underlying communication schedule, and we propose an open-source software framework for time-triggered end-systems. We present the implementation of a static cyclic task schedule, on a time-predictable platform that is integrated within a TTEthernet network and synchronized with the communication schedule. We evaluate the presented framework by developing a simple one-sensor, one-actuator industrial control example, distributed over three nodes that communicate over a single TTEthernet switch. The presented real-time system can exchange messages with minimal jitter as the distributed tasks are synchronized over the TTEthernet network with about 1.6 us precision. Due to the tight time synchronization, the system can operate stably with zero missed frames, using a single receiver and a single transmitter buffer. Eleftherios Kyriakakis, Jens Sparsø, Peter P. Puschner, Martin Schoeberl |
ISORC | 4 |
| 2021 | Fault-tolerant Clock Synchronization using Precise Time Protocol Multi-Domain AggregationabstractDistributed real-time systems often rely on time-triggered communication and task execution to guarantee end-to-end latency and time-predictable computation. Such systems require a reliable synchronized network time to be shared among end-systems. The IEEE 1588 Precision Time Protocol (PTP) enables such clock synchronization throughout an Ethernet-based network. While security was not addressed in previous versions of the IEEE 1588 standard, in its most recent iteration (IEEE 1588-2019), several security mechanisms and recommendations were included describing different measures that can be taken to improve system security and safety. One proposal to improve security and reliability is to add redundancy to the network through modifications in the topology. However, this recommendation omits implementation details and leaves the question open of how it affects synchronization quality. This work investigates the quality impact and security properties of redundant PTP deployment and proposes an observation window-based multi-domain, PTP end-system, design to increase fault-tolerance and security. We implement the proposed design inside a discrete-event network simulator and evaluate its clock synchronization quality using two test-case network topologies with simulated faults. Eleftherios Kyriakakis, Koen Tange, Niklas Reusch, Eder Ollora Zaballa, Xenofon Fafoutis, Martin Schoeberl, Nicola Dragoni |
ISORC | 6 |
| 2021 | Static Timing Analysis of OPC UA PubSubabstractIndustrial automation is changing towards higher integration and seamless communication. A stepping stone is end-to-end real-time machine-to-machine communication, now becoming feasible with technologies such as time-sensitive networking (TSN) and OPC Unified Architecture (OPC UA) publish-subscribe. While TSN takes care of communication, the OPC UA stack's execution time behavior remains unknown. This paper highlights experiences made while adjusting the OPC UA subscriber of the open62541 stack for worst-case execution time (WCET) analysis. Two directly connected time-predictable T-CREST platforms hosting the publisher and subscriber delivered end-to-end timing measures validating the WCET estimates. The paper concludes by outlining further research with several time-predictable publishers and subscribers. Patrick Denzler, Thomas Frühwirth, Andreas Kirchberger, Martin Schoeberl, Wolfgang Kastner |
WFCS | 4 |
| 2021 | Compiling for time-predictability with dual-issue single-path codeabstractDesigned for real-time systems, the Patmos instruction-set architecture's features ensure a high degree of predictability.One such feature is its dual-issue pipeline, which can issue and execute bundles of up to two instructions at a time.Executing instructions in the second issue slot is a predictable way to increase the throughput of a processor, but without dedicated support from the compiler, this benefit cannot be unlocked.A compiler generates highly predictable programs by generating single-path code.This technique produces code that always follows the same trace of instructions.While Patmos' compiler can already produce single-path code, it does not assign any instructions to the second issue-slot.This limitation is unfortunate, as single-path code inherently possesses a high degree of instruction-level parallelism.In this paper, we present a singlepath code generation technique with support for dual-issue pipelines.It can also support different bundling algorithms, which allows changing algorithms without having to edit other parts of the compiler.We present a simple bundling algorithm plugged into the single-path code generator.It looks for branches and bundles the basic blocks on each path of the branch.While this specific bundling algorithm is too simple to provide a real-world benefit, it highlights the potential that further work on bundling algorithms can unlock. Emad Jacob Maroun, Martin Schoeberl, Peter P. Puschner |
J. Syst. Archit. | 2 |
| 2020 | Formal Semantics of Predictable Pipelines: a Comparative StudyabstractComputer architectures used in safety-critical domains are subjected to worst-case execution time analysis. The presence of performance-driven microarchitectures may trigger undesired timing phenomena, called timing anomalies, and complicate the timing analysis. This paper investigates pipelines specifically designed to simplify the worst-case execution time analysis (also called predictable pipelines). We propose formal and executable models of four research-oriented pipelines and one industrial pipeline to validate some of their claims related to their timing behavior. We indeed validate, via bounded model checking, the absence of a type of timing anomalies called amplification timing anomalies, or its potential presence by identifying prerequisite to situations where they can occur. Mathieu Jan, Mihail Asavoae, Martin Schoeberl, Edward A. Lee |
ASP-DAC | 3 |
| 2020 | Synchronizing Real-Time Tasks in Time-Aware Networks: Work-in-ProgressabstractDistributed safety-critical systems require both time-predictable task execution and communication. On the processor, the execution of the tasks is dictated by a scheduling policy, while on the network, different industrial communication protocols can be deployed to guarantee bounded message latency. In this paper, we investigate the synchronization of the task execution with the underlying communication schedule, and we propose an open-source software framework. We implement a cyclic executive task scheduling policy on a time-predictable platform and synchronize the task execution with the underlying TTEth-ernet communication schedule. We evaluate our framework by developing a simple one-sensor, one-actuator industrial control example, distributed over three nodes. The presented real-time system can exchange messages with minimal jitter, and the distributed tasks synchronize to a precision of ≈ 1.6μs. Eleftherios Kyriakakis, Jens Sparsø, Peter P. Puschner, Martin Schoeberl |
EMSOFT | 4 |
| 2020 | Towards Dual-Issue Single-Path CodeabstractThe Patmos instruction-set architecture is designed for real-time systems. As such, it has features that increase the predictability of code running on it. One important feature is its dual-issue pipeline: instructions may be organized in bundles of two that are issued and executed in parallel. This increases the throughput of the processor in a predictable manner, but only if the compiler makes use of it.Single-path code is a code-generation technique that produces predictable executions by always following the same trace of instructions. The Patmos compiler can already produce single-path code, but it does not use the second issue slot available in the processor. This is less than ideal because the single-path transformation results in code that has a high degree of instruction-level parallelism.In this paper, we present a single-path code generator that can produce bundled instructions. It includes generic support for bundling algorithms, such that implementing them is simple and does not require changing other parts of the compiler.We also present one such bundling algorithm plugged into the single-path code generator. With it, we show that we can produce dual-issue instructions to improve performance. Emad Jacob Maroun, Martin Schoeberl, Peter P. Puschner |
ISORC | 2 |
| 2020 | A time-predictable open-source TTEthernet end-systemabstractCyber-physical systems deployed in areas like automotive, avionics, or industrial control are often distributed systems. The operation of such systems requires coordinated execution of the individual tasks with bounded communication network latency to guarantee quality-of-control. Both the time for computing and communication needs to be bounded and statically analyzable. To provide deterministic communication between end-systems, real-time networks can use a variety of industrial Ethernet standards typically based on time-division scheduling and enforced by real-time enabled network switches. For the computation, end-systems need time-predictable processors where the worst-case execution time of the application tasks can be analyzed statically. This paper presents a time-predictable end-system with support for deterministic communication using the open-source processor Patmos. The proposed architecture is deployed in a TTEthernet network, and the protocol software stack is implemented, and the worst-case execution time is statically analyzed. The developed end-system is evaluated in an experimental network setup composed of six TTEthernet nodes that exchange periodic frames over a TTEthernet switch. Eleftherios Kyriakakis, Maja Lund, Luca Pezzarossa, Jens Sparsø, Martin Schoeberl |
J. Syst. Archit. | 5 |
| 2019 | Actors Revisited for Time-Critical SystemsabstractProgramming time-critical systems is notoriously difficult. In this paper we propose an actor-oriented programming model with a semantic notion of time and a deterministic coordination semantics based on discrete events to exercise precise control over both the computational and timing aspects of the system behavior. Marten Lohstroh, Martin Schoeberl, Andres Goens, Armin Wasicek, Christopher D. Gill, Marjan Sirjani, Edward A. Lee |
DAC | 2 |
| 2019 | Scratchpad Memories with OwnershipabstractA multicore processor for real-time systems needs a time-predictable way to communicate data between different threads running on different cores. Standard multicore processors support data sharing with shared main memory backed up by caches and cache coherence protocol. This sharing solution is hardly time predictable nor does it scale to more than a few cores.This paper presents a shared scratchpad memory (SPM) for time-predictable communication between cores. The base architecture uses time-division multiplexing for the arbitration of the access to the shared SPM. This allows the timing of programs executing on different cores to be completely independent of each other. We extend this architecture by the notion of ownership. A core can own the SPM. Having exclusive access to the SPM reduces the access time to a single clock cycle. The ownership of the SPM can then be transferred to a different core, implementing low latency communication of bulk data. As an extension, we propose to organize this memory as a pool of SPMs that can be owned by different cores and transferred as needed. We evaluate the proposed architecture within the T-CREST multicore architecture. Martin Schoeberl, Tórur Biskopstø Strøm, Oktay Baris, Jens Sparsø |
DATE | 1 |
| 2019 | Demonstration of a Time-predictable Flight Controller on a Multicore ProcessorabstractUnmanned aerial vehicles, or drones, have drawn extensive attention during the last decade together with the maturity of the technology. Often, real-time requirements are needed when they are deployed for critical missions. The increasing computational demand for drones leads to a shift towards multicore architectures. However, the timing analysis of multicore systems is a challenging task due to the timing interference between processor cores. In this demonstration, we deploy a parallelized flight controller system on a multicore architecture. We provide timing analysis of the system, and test it with a processor-in-the-loop setup, which includes a flight simulator connected to the flight controller running on a time-predictable multicore platform. Oktay Baris, Shibarchi Majumder, Tórur Biskopstø Strøm, Anders la Cour-Harbo, Jens Sparsø, Thomas Bak, Martin Schoeberl |
ISORC | 7 |
| 2019 | A Time-predictable TTEthenet NodeabstractDistributed real-time systems need time-predictable computation and communication to facilitate static analysis of timing requirements and deadlines. This paper presents the implementation of a deterministic network protocol, TTEthernet, on the time-predictable Patmos processor. The implementation uses the existing Ethernet controller on the processor and we tested it with a TTEthernet system provided by TTTech Inc. Further testing showed that the controller could send time-triggered messages with bounded latency and a small jitter of approximately 4.5 us. We also provide worst-case execution time analysis of the network code, which demonstrates a time-predictable end-to-end solution. This work enables Patmos to communicate with other nodes in a deterministic way. Thus, extending the possible uses of Patmos. Maja Lund, Luca Pezzarossa, Jens Sparsø, Martin Schoeberl |
ISORC | 4 |
| 2019 | Hardlock: Real-time multicore locking
Tórur Biskopstø Strøm, Jens Sparsø, Martin Schoeberl |
J. Syst. Archit. | 3 |
| 2018 | One-way shared memoryabstractStandard multicore processors use the shared main memory via the on-chip caches for communication between cores. However, this form of communication has two limitations: (1) it is hardly time-predictable and therefore not a good solution for real-time systems and (2) this single shared memory is a bottleneck in the system. This paper presents a communication architecture for time-predictable multicore systems where core-local memories are distributed on the chip. A network-on-chip constantly copies data from a sender core-local memory to a receiver core-local memory. As this copying is performed in one direction we call this architecture a one-way shared memory. With the use of time-division multiplexing for the memory accesses and the network-on-chip routers we achieve a time-predictable solution where the communication latency and bandwidth can be bounded. An example architecture for a 3×3 core processor and 32-bit wide links and memory ports provides a cumulative bandwidth of 29 bytes per clock cycle. Furthermore, the evaluation shows that this architecture, due to its simplicity, is small compared to other network-on-chip solutions. Martin Schoeberl |
DATE | 1 |
| 2018 | Design of a time-predictable multicore processor: The T-CREST projectabstractReal-time systems need to deliver results in time and often this timely production of a result needs to be guaranteed. Static timing analysis can be used to bound the worst-case execution time of tasks. However, this timing analysis is only possible if the processor architecture is analysis friendly. This paper presents the T-CREST processor, a real-time multicore processor developed to be time-predictable and an easy target for static worst-case execution time analysis. We present how to achieve time-predictability at all levels of the architecture, from the processor pipeline, via a network-on-chip, up to the memory controller. The main architectural feature to provide time predictability is to use static arbitration of shared resources in a time-division multiplexing way. Martin Schoeberl |
DATE | 1 |
| 2018 | Faster Function Blocks for Precision Timed Industrial AutomationabstractIn industrial automation, safety-critical control systems need robust timing guarantees in addition to functional correctness. Unfortunately, devices that are typically used in this domain, such as Programmable Logic Controllers, often feature architectures that are not amenable to static timing analysis, for instance relying on general purpose microprocessors or embedded operating systems. As a result, designers often rely on timing values gained from simple measurement of running applications, an approach that only provides very weak guarantees at best. The synchronous approach for IEC 61499 Function Blocks, in contrast, has been demonstrated to be time predictable when run on appropriate hardware, such as simple microprocessors. However, simple microprocessors are often not fast or powerful enough for modern automation requirements. In this paper, we examine how the performance of synchronous IEC 61499 can be improved through the usage of the multi-core T-CREST architecture, data scratchpads, and an optimised compiler. Overall, our improvements resulted in 60% shorter worst-case execution times. Hammond A. Pearce, Partha S. Roop, Morteza Biglari-Abhari, Martin Schoeberl |
ISORC | 4 |
| 2018 | tpIP: A Time-Predictable TCP/IP Stack for Cyber-Physical SystemsabstractCyber-physical systems are networks of computers connected to the physical world. Often the interaction with the physical world is time critical. In that case computation and communication must be performed in real time. However, a standard implementation of a network stack is hardly time predictable. This paper addresses the challenge of real-time communication for time-critical cyber-physical systems with a time-predictable network stack. We present tpIP, a real-time implementation of the TCP/IP stack. We achieve time predictability by two properties: (1) the application interface is based on polling functions, instead of blocking sockets, that fits for periodic real-time tasks; (2) the implementation is carefully crafted to enable static worst-case execution time analysis of all functions. Martin Schoeberl, Rasmus Ulslev Pedersen |
ISORC | 1 |
| 2018 | Hardlock: A Concurrent Real-Time Multicore Locking UnitabstractTo use multicore processors, an application needs to split computation into several threads that execute on different processing cores. As those threads work together towards a common goal, they need to exchange data in a controlled way. A common communication paradigm between cooperating threads is using shared data structures protected by locks. Implementing a lock on top of shared memory can easily result in a bottleneck on a multicore processor due to the congestion on the shared memory. However, the number of locks in use is usually low and using the large external memory to support locks is over-provisioning a resource. This paper presents an efficient implementation of locking by providing dedicated hardware support for locking on-chip. This locking unit supports a restricted number of locks without the need to get off-chip. The unit can process lock acquisitions in 2 clock cycles and releases in 1 clock cycle. Tórur Biskopstø Strøm, Martin Schoeberl |
ISORC | 2 |
| 2018 | Patmos: a time-predictable microprocessor
Martin Schoeberl, Wolfgang Puffitsch, Stefan Hepp, Benedikt Huber, Daniel Wiltsche-Prokesch |
Real Time Syst. | 1 |
| 2017 | Improving Performance of Single-Path Code through a Time-Predictable Memory HierarchyabstractDeriving the Worst-Case Execution Time (WCET) of a task is a challenging process, especially for processor architectures that use caches, out-of-order pipelines, and speculative execution. Despite existing contributions to WCET analysis for these complex architectures, there are open problems. The single-path code generation overcomes these problems by generating time-predictable code that has a single execution trace. However, the simplicity of this approach comes at the cost of longer execution times. This paper addresses performance improvements for single-path code. We propose a time-predictable memory hierarchy with a prefetcher that exploits the predictability of execution traces in single-path code to speed up code execution. The new memory hierarchy reduces both the cache-miss penalty time and the cache-miss rate on the instruction cache. The benefit of the approach is demonstrated through benchmarks that are executed on an FPGA implementation. Bekim Cilku, Wolfgang Puffitsch, Daniel Wiltsche-Prokesch, Martin Schoeberl, Peter P. Puschner |
ISORC | 4 |
| 2017 | A Controller for Dynamic Partial Reconfiguration in FPGA-Based Real-Time SystemsabstractIn real-time systems, the use of hardware accelerators can lead to a worst-case execution-time speed-up, to a simplification of its analysis, and to a reduction of its pessimism. When using FPGA technology, dynamic partial reconfiguration (DPR) can be used to minimize the area, by only loading those accelerators that are needed at any given point in time. The DPR controllers provided by the FPGA vendors satisfy a wide range of requirements and rely on software to manage the reconfiguration. This approach may lead to slow reconfiguration and unpredictable timing. This paper presents an open-source DPR controller specially developed for hard real-time systems and prototyped in connection with the open-source multi-core platform for real-time applications T-CREST. The controller enables a processor to perform reconfiguration in a time-predictable manner and supports different operating modes. The paper also presents a software tool for bitstream conversion, compression, and for reconfiguration time analysis. The DPR controller is evaluated in terms of hardware cost, operating frequency, speed, and bitstream compression ratio vs. reconfiguration time trade-off. A simple application example is also presented with the scope of showing the reconfiguration features of the controller. Luca Pezzarossa, Martin Schoeberl, Jens Sparsø |
ISORC | 2 |
| 2017 | Safety-critical Java for embedded systemsabstractSummary This paper presents the motivation for and outcomes of an engineering research project on certifiable Java for embedded systems. The project supports the upcoming standard for safety‐critical Java, which defines a subset of Java and libraries aiming for development of high criticality systems. The outcome of this project include prototype safety‐critical Java implementations, a time‐predictable Java processor, analysis tools for memory safety, and example applications to explore the usability of safety‐critical Java for this application area. The text summarizes developments and key contributions and concludes with the lessons learned. Copyright © 2016 John Wiley & Sons, Ltd. Martin Schoeberl, Andreas Engelbredt Dalsgaard, René Rydhof Hansen, Stephan Korsholm, Anders P. Ravn, Juan Ricardo Rios, Tórur Biskopstø Strøm, Hans Søndergaard, Andy J. Wellings, Shuai Zhao 0004 |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | Hardware locks for a real-time Java chip multiprocessorabstractSummary A software locking mechanism commonly protects shared resources for multithreaded applications. This mechanism can, especially in chip‐multiprocessor systems, result in a large synchronization overhead. For real‐time systems in particular, this overhead increases the worst‐case execution time and may void a task set's schedulability. This paper presents 2 hardware locking mechanisms to reduce the worst‐case time required to acquire and release synchronization locks. These solutions are implemented for the chip‐multiprocessor version of the Java Optimized Processor. The 2 hardware locking mechanisms are compared with a software locking solution as well as the original locking system of the processor. The hardware cost and performance are evaluated for all presented locking mechanisms. The performance of the better‐performing hardware locks is comparable with that of the original single global lock when contending for the same lock. When several noncontending locks are used, the hardware locks enable true concurrency for critical sections. Benchmarks show that using the hardware locks yields performance ranging from no worse than the original locks to more than twice their best performance. This improvement can allow a larger number of real‐time tasks to be reliably scheduled on a multiprocessor real‐time platform. Copyright © 2016 John Wiley & Sons, Ltd. Tórur Biskopstø Strøm, Wolfgang Puffitsch, Martin Schoeberl |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | A resource-efficient network interface supporting low latency reconfiguration of virtual circuits in time-division multiplexing networks-on-chip
Rasmus Bo Sørensen, Luca Pezzarossa, Martin Schoeberl, Jens Sparsø |
J. Syst. Archit. | 3 |
| 2016 | Lessons learned from the EU project T-CREST
Martin Schoeberl |
DATE | 1 |
| 2016 | Time-Predictable Virtual MemoryabstractVirtual memory is an important feature of modern computer architectures. For hard real-time systems, memory protection is a particularly interesting feature of virtual memory. However, current memory management units are not designed for time-predictability and therefore cannot be used in such systems. This paper investigates the requirements on virtual memory from the perspective of hard real-time systems and presents the design of a time-predictable memory management unit. Our evaluation shows that the proposed design can be implemented efficiently. The design allows address translation and address range checking in constant time of two clock cycles on a cache miss. This constant time is in strong contrast to the possible cost of a miss in a translation look-aside buffer in traditional virtual memory organizations. Compared to a platform without a memory management unit, these two additional clock cycles per cache miss introduce only a small performance overhead. Wolfgang Puffitsch, Martin Schoeberl |
ISORC | 2 |
| 2016 | A Stack Cache for Real-Time SystemsabstractReal-time systems need time-predictable computing platforms to allow for static analysis of the worst-case execution time. Caches are important for good performance, but data caches are hard to analyze for the worst-case execution time. Stack allocated data has different properties related to locality, lifetime, and static analyzability of access addresses compared to static or heap allocated data. Therefore, caching of stack allocated data benefits from having its own cache. In this paper we present a cache architecture optimized for stack allocated data. This cache is additional to the normal data cache. As stack allocated data has a high locality, even a small stack cache gives a high hit rate. A stack cache added to a write-through data cache considerably improves the performance, while a stack cache compared to the harder to analyze write-back cache has about the sameaverage case performance. Martin Schoeberl, Carsten Nielsen |
ISORC | 1 |
| 2016 | Avionics Applications on a Time-Predictable Chip-MultiprocessorabstractAvionics applications need to be certified for the highest criticality standard. This certification includes schedulability analysis and worst-case execution time (WCET) analysis. WCET analysis is only possible when the software is written to be WCET analyzable and when the platform is time-predictable. In this paper we present prototype avionics applications that have been ported to the time-predictable T-CREST platform. The applications are WCET analyzable, and T-CREST is supported by the aiT WCET analyzer. This combination allows us to provide WCET bounds of avionic tasks, even when executing on a multicore processor. André Rocha, Cláudio Silva 0002, Rasmus Bo Sørensen, Jens Sparsø, Martin Schoeberl |
PDP | 5 |
| 2016 | Argo: A Real-Time Network-on-Chip Architecture With an Efficient GALS ImplementationabstractIn this paper, we present an area-efficient, globally asynchronous, locally synchronous network-on-chip (NoC) architecture for a hard real-time multiprocessor platform. The NoC implements message-passing communication between processor cores. It uses statically scheduled time-division multiplexing (TDM) to control the communication over a structure of routers, links, and network interfaces (NIs) to offer real-time guarantees. The area-efficient design is a result of two contributions: 1) asynchronous routers combined with TDM scheduling and 2) a novel NI microarchitecture. Together they result in a design in which data are transferred in a pipelined fashion, from the local memory of the sending core to the local memory of the receiving core, without any dynamic arbitration, buffering, and clock synchronization. The routers use two-phase bundled-data handshake latches based on the Mousetrap latch controller and are extended with a clock gating mechanism to reduce the energy consumption. The NIs integrate the direct memory access functionality and the TDM schedule, and use dual-ported local memories to avoid buffering, flow-control, and synchronization. To verify the design, we have implemented a 4 × 4 bitorus NoC in 65-nm CMOS technology and we present results on area, speed, and energy consumption for the router, NI, NoC, and post layout. Evangelia Kasapaki, Martin Schoeberl, Rasmus Bo Sørensen, Christoph Thomas Muller, Kees Goossens, Jens Sparsø |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Message Passing on a Time-predictable Multicore ProcessorabstractReal-time systems need time-predictable computing platforms. For a multicore processor to be time-predictable, communication between processor cores needs to be time-predictable as well. This paper presents a time-predictable message-passing library for such a platform. We show how to build up abstraction layers from a simple, time-division multiplexed hardware push channel. We develop these time-predictable abstractions and implement them in software. To prove the time-predictability of these functions we analyze their worst-case execution time (WCET) with the aiT WCET analysis tool. We combine these WCET numbers with the calculation of the network latency of a message and then provide a statically computed end-to-end latency for this core-to-core message. Rasmus Bo Sørensen, Wolfgang Puffitsch, Martin Schoeberl, Jens Sparsø |
ISORC | 3 |
| 2015 | Hardware Locks with Priority Ceiling Emulation for a Java Chip-MultiprocessorabstractAccording to the safety-critical Java specification, priority ceiling emulation is a requirement for implementations, as it has preferable properties, such as avoiding priority inversion and being deadlock free on uni-core systems. In this paper we explore our hardware supported implementation of priority ceiling emulation on the multicore Java optimized processor, and compare it to the existing hardware locks on the Java optimized processor. We find that the additional overhead for priority ceiling emulation on a multicore processor is several times higher than simpler, non-premptive locks, mainly due to slow access to shared memory. We also find that PCE is mostly viable with large critical sections. Tórur Biskopstø Strøm, Martin Schoeberl |
ISORC | 2 |
| 2015 | T-CREST: Time-predictable multi-core architecture for embedded systemsabstractReal-time systems need time-predictable platforms to allow static analysis of the worst-case execution time (WCET). Standard multi-core processors are optimized for the average case and are hardly analyzable. Within the T-CREST project we propose novel solutions for time-predictable multi-core architectures that are optimized for the WCET instead of the average-case execution time. The resulting time-predictable resources (processors, interconnect, memory arbiter, and memory controller) and tools (compiler, WCET analysis) are designed to ease WCET analysis and to optimize WCET performance. Compared to other processors the WCET performance is outstanding. The T-CREST platform is evaluated with two industrial use cases. An application from the avionic domain demonstrates that tasks executing on different cores do not interfere with respect to their WCET. A signal processing application from the railway domain shows that the WCET can be reduced for computation-intensive tasks when distributing the tasks on several cores and using the network-on-chip for communication. With three cores the WCET is improved by a factor of 1.8 and with 15 cores by a factor of 5.7. The T-CREST project is the result of a collaborative research and development project executed by eight partners from academia and industry. The European Commission funded T-CREST. Martin Schoeberl, Sahar Abbaspour, Benny Akesson, Neil C. Audsley, Raffaele Capasso, Jamie Garside, Kees Goossens, Sven Goossens, Scott Hansen, Reinhold Heckmann, Stefan Hepp, Benedikt Huber, Alexander Jordan, Evangelia Kasapaki, Jens Knoop, Yonghui Li 0002, Daniel Wiltsche-Prokesch, Wolfgang Puffitsch, Peter P. Puschner, André Rocha, Cláudio Silva 0002, Jens Sparsø, Alessandro Tocchi |
J. Syst. Archit. | 1 |
| 2014 | A Method Cache for PatmosabstractFor real-time systems we need time-predictable processors. This paper presents a method cache as a time-predictable solution for instruction caching. The method cache caches whole methods (or functions) and simplifies worst-case execution time analysis. We have integrated the method cache in the time-predictable processor Patmos. We evaluate the method cache with a large set of embedded benchmarks. Most benchmarks show a good hit rate for a method cache size in the range between 4 and 16 KB. Philipp Degasperi, Stefan Hepp, Wolfgang Puffitsch, Martin Schoeberl |
ISORC | 4 |
| 2014 | Reusable Libraries for Safety-Critical JavaabstractThe large collection of Java class libraries is a main factor of the success of Java. However, these libraries assume that a garbage-collected heap is used. Safety-critical Java uses scope-based memory areas instead of a garbage-collected heap. Therefore, the Java class libraries are problematic to use in safety-critical Java. We have identified common programming patterns in the Java class libraries that make them unsuitable for safety-critical Java. We propose ways to improve the libraries to avoid the impact of the identified problematic patterns. We illustrate these changes by implementing a total of five scope-safe classes from commonly used libraries. Juan Ricardo Rios, Martin Schoeberl |
ISORC | 2 |
| 2014 | An Evaluation of Safety-Critical Java on a Java ProcessorabstractAbstract—The safety-critical Java (SCJ) specification provides a restricted set of the Java language intended for applications that require certification. In order to test the specification, implementations are emerging and the need to evaluate those implementations in a systematic way is becoming important. In this paper we evaluate our SCJ implementation which is based on the Java Optimized Processor JOP and we measure different performance and timeliness criteria relevant to hard real-time systems. Our implementation targets Level 0 and Level 1 of the specification and to test it we use a series of micro benchmarks, an application-based benchmark, and a reduced set of a SCJ technology compatibility kit. We evaluate the accuracy of periods, linear-time memory allocation, aperiodic event handling, dispatch latency for interrupts, context switch preemption latency, and synchronization. Juan Ricardo Rios, Martin Schoeberl |
ISORC | 2 |
| 2014 | WCET-Based Comparison of an Instruction Scratchpad and a Method CacheabstractThis paper compares two proposed alternatives to conventional instruction caches: a scratchpad memory (SPM) and a method cache. The comparison considers the true worst-case execution time (WCET) and the estimated WCET bound of programs using either an SPM or a method cache, using large numbers of randomly generated programs. For these programs, we find that a method cache is preferable to an SPM if the true WCET is used, because it leads to execution times that are no greater than those for SPM, and are often lower. However, we also find that analytical pessimism is a significant problem for a method cache. If WCET bounds are derived by analysis, the WCET bounds for an instruction SPM are often lower than the bounds for a method cache. This means that an SPM may be preferable in practical systems. Jack Whitham, Martin Schoeberl |
ISORC | 2 |
| 2013 | An area-efficient network interface for a TDM-based network-on-chipabstractNetwork interfaces (NIs) are used in multi-core systems where they connect processors, memories, and other IP-cores to a packet switched Network-on-Chip (NOC). The functionality of a NI is to bridge between the read/write transaction interfaces used by the cores and the packet-streaming interface used by the routers and links in the NOC. The paper addresses the design of a NI for a NOC that uses time division multiplexing (TDM). By keeping the essence of TDM in mind, we have developed a new area-efficient NI micro-architecture. The new design completely eliminates the need for FIFO buffers and credit based flow control - resources which are reported to account for 50–85% of the area in existing NI designs. The paper discusses the design considerations, presents the new NI micro-architecture, and reports area figures for a range of implementations. Jens Sparsø, Evangelia Kasapaki, Martin Schoeberl |
DATE | 3 |
| 2013 | A time-predictable stack cacheabstractReal-time systems need time-predictable architectures to support static worst-case execution time (WCET) analysis. One architectural feature, the data cache, is hard to analyze when different data areas (e.g., heap allocated and stack allocated data) share the same cache. This sharing leads to less precise results of the cache analysis part of the WCET analysis. Splitting the data cache for different data areas enables composable data cache analysis. The WCET analysis tool can analyze the accesses to these different data areas independently. In this paper we present the design and implementation of a cache for stack allocated data. Our port of the LLVM C++ compiler supports the management of the stack cache. The combination of stack cache instructions and the hardware implementation of the stack cache is a further step towards time-predictable architectures. Sahar Abbaspour, Florian Brandner, Martin Schoeberl |
ISORC | 3 |
| 2013 | An SDRAM controller for real-time systemsabstractFor real-time systems we need to statically determine worst-case execution times (WCET) of tasks to proof the schedulability of the system. To enable static WCET analysis, the platform needs to be time-predictable. The platform includes the processor, the caches, the memory system, the operating system, and the application software itself. All those components need to be timing analyzable. Current computers use DRAM as a cost effective main memory. However, these DRAM chips have timing requirements that depend on former accesses and also need to be refreshed to retain their content. Standard memory controllers for DRAM memories are optimized to provide maximum bandwidth or throughput at the cost of variable latency for individual memory accesses. In this paper we present an SDRAM controller for realtime systems. The controller is optimized for the worst case and constant latency to provide a base of the memory hierarchy for time-predictable systems. Edgar Lakis, Martin Schoeberl |
ISORC | 2 |
| 2013 | Micro-transactions for concurrent data structuresabstractSUMMARY Transactional memory is a promising technique for enforcing disciplined access to shared data in a multiprocessor system. Transactional memory simplifies the implementation of a variety of concurrent data structures. In this paper, we study the benefits of a modest, real‐time aware, hardware implementation of transactional memory that we callmicro‐transactions. In particular, we argue that hardware support for micro‐transactions allows us to efficiently implement certain data structures. Those data structures are difficult to realize with the atomic operations provided by stock hardware and provide real‐time guarantees for those operations. Our main implementation platform is the Java Optimized Processor system, a field‐programmable gate array (FPGA) implementation of the Java virtual machine, optimized for real‐time Java. We report on the performance of data structures implemented with locks, atomic instructions, and micro‐transactions. Our results suggest that transactional memory is an interesting alternative to traditional concurrency control mechanisms. Copyright © 2012 John Wiley & Sons, Ltd. Fadi Meawad, Karthik Iyer, Martin Schoeberl, Jan Vitek |
Concurr. Comput. Pract. Exp. | 3 |
| 2013 | Data cache organization for accurate timing analysis
Martin Schoeberl, Benedikt Huber, Wolfgang Puffitsch |
Real Time Syst. | 1 |
| 2012 | Worst-Case Execution Time Based Optimization of Real-Time Java ProgramsabstractStandard compilers optimize execution time for the average case. However, in hard real-time systems the worst-case execution time (WCET) is of primary importance. Therefore, a compiler for real-time systems shall include optimizations that aim to minimize the WCET. One effective compiler optimization is method in lining. It is especially important for languages, like Java, where small setter and getter methods are considered good programming style. In this paper we present and explore WCET driven in lining of Java methods. We use the WCET analysis tool for the Java processor JOP to guide to optimization along the worst-case path. The tool JCopter is integrated with the WCET analysis tool and is used to explore different in lining strategies. On real-time benchmarks the optimization results in a reduction of the WCET by a few percent up to a factor of about 2. Stefan Hepp, Martin Schoeberl |
ISORC | 2 |
| 2012 | Hardware Support for Safety-Critical Java Scope ChecksabstractMemory management in Safety-Critical Java (SCJ) is based on time bounded, non garbage collected scoped memory regions used to store temporary objects. Scoped memory regions may have different life times during the execution of a program and hence, to avoid leaving dangling pointers, it is necessary to check that reference assignments are performed only from objects in shorter lived scopes to objects in longer lived scopes (or between objects in the same scoped memory area). SCJ offers, compared to the RTSJ, a simplified memory model where only the immortal and mission memory scoped areas are shared between threads and any other scoped region is thread private. In this paper we present how, due to this simplified model, a single scope nesting level can be used to check the legality of every reference assignment. We also show that with simple hardware extensions a processor can see some improvement in terms of execution time for applications where cross-scope references are frequent. Our proposal was implemented and tested on the Java Optimized Processor (JOP). Juan Ricardo Rios, Martin Schoeberl |
ISORC | 2 |
| 2012 | A Statically Scheduled Time-Division-Multiplexed Network-on-Chip for Real-Time SystemsabstractThis paper explores the design of a circuit-switched network-on-chip (NoC) based on time-division-multiplexing (TDM) for use in hard real-time systems. Previous work has primarily considered application-specific systems. The work presented here targets general-purpose hardware platforms. We consider a system with IP-cores, where the TDM-NoC must provide directed virtual circuits -- all with the same bandwidth -- between all nodes. This may not be a frequent scenario, but a general platform should provide this capability, and it is an interesting point in the design space to study. The paper presents an FPGA-friendly hardware design, which is simple, fast, and consumes minimal resources. Furthermore, an algorithm to find minimum-period schedules for all-to-all virtual circuits on top of typical physical NoC topologies like 2D-mesh, torus, bidirectional torus, tree, and fat-tree is presented. The static schedule makes the NoC time-predictable and enables worst-case execution time analysis of communicating real-time tasks. Martin Schoeberl, Florian Brandner, Jens Sparsø, Evangelia Kasapaki |
NOCS | 1 |
| 2012 | Worst-case execution time analysis-driven object cache designabstractSUMMARY Hard real‐time systems need a time‐predictable computing platform to enable static worst‐case execution time (WCET) analysis. All performance‐enhancing features need to be WCET analyzable. However, standard data caches containing heap‐allocated data are very hard to analyze statically. In this paper we explore a new object cache design, which is driven by the capabilities of static WCET analysis. Simulations of standard benchmarks estimating the expected average case performance usually drive computer architecture design. The design decisions derived from this methodology do not necessarily result in a WCET analysis‐friendly design. Aiming for a time‐predictable design, we therefore propose to employ WCET analysis techniques for the design space exploration of processor architectures. We evaluated different object cache configurations using static analysis techniques. The number of field accesses that can be statically classified as hits is considerable. The analyzed number of cache miss cycles is 3–46% of the access cycles needed without a cache, which agrees with trends obtained using simulations. Standard data caches perform comparably well in the average case, but accesses to heap data result in overly pessimistic WCET estimations. We therefore believe that an early architecture exploration by means of static timing analysis techniques helps to identify configurations suitable for hard real‐time systems. Copyright © 2011 John Wiley & Sons, Ltd. Benedikt Huber, Wolfgang Puffitsch, Martin Schoeberl |
Concurr. Comput. Pract. Exp. | 3 |
| 2012 | Safety-critical Java with cyclic executives on chip-multiprocessorsabstractSUMMARY Chip‐multiprocessors offer increased processing power at a low cost. However, in order to use them for real‐time systems, tasks have to be scheduled efficiently and predictably. It is well known that finding optimal schedules is a computationally hard problem. In this paper we present a solution that uses model checking to find a static schedule, if one exists at all, which gives an implementation of a table driven multiprocessor scheduler. Mutual exclusion to access shared resources is guaranteed by including access constraints in the schedule generation. To evaluate the proposed cyclic executive for multiprocessors, we have implemented it in the context of safety‐critical Java on a Java processor. Copyright © 2011 John Wiley & Sons, Ltd. Anders P. Ravn, Martin Schoeberl |
Concurr. Comput. Pract. Exp. | 2 |
| 2012 | Fast, Interactive Worst-Case Execution Time Analysis With Back-AnnotationabstractFor hard real-time systems, static code analysis is needed to derive a safe bound on the worst-case execution time (WCET). Virtually all prior work has focused on the accuracy of WCET analysis without regard to the speed of analysis. The resulting algorithms are often too slow to be integrated into the development cycle, requiring WCET analysis to be postponed until a final verification phase. In this paper, we propose interactive WCET analysis as a new method to provide near-instantaneous WCET feedback to the developer during software programming. We show that interactive WCET analysis is feasible using tree-based WCET calculation. The feedback is realized with a plugin for the Java editor jEdit, where the WCET values are back-annotated to the Java source at the statement level. Comparison of this tree-based approach with the implicit path enumeration technique (IPET) shows that tree-based analysis scales better with respect to program size and gives similar WCET values. Trevor Harmon, Martin Schoeberl, Raimund Kirner, Raymond Klefstad, K. H. (Kane) Kim, Michael R. Lowry |
IEEE Trans. Ind. Informatics | 2 |
| 2011 | Leros: A Tiny Microcontroller for FPGAsabstractLeros is a tiny microcontroller that is optimized for current low-cost FPGAs. Leros is designed with a balanced logic to on-chip memory relation. The design goal is a microcontroller that can be clocked in about half of the speed a pipelined on-chip memory and consuming less than 300 logic cells. The architecture, which follows from the design goals, is a pipelined 16-bit accumulator processor. An implementation of Leros needs at least one on-chip memory block and a few hundred logic cells. The application areas of Leros are twofold: First, it can be used as an intelligent peripheral device for auxiliary functions in an FPGA based system-on-chip design. Second, the very small size of Leros makes it an attractive soft core for many-core research with low-cost FPGAs. Martin Schoeberl |
FPL | 1 |
| 2011 | Hardware synchronization for embedded multi-core processorsabstractMulti-core processors are about to conquer embedded systems - it is not the question of whether they are coming but how the architectures of the microcontrollers should look with respect to the strict requirements in the field. We present the step from one to multiple cores in this paper, establishing coherence and consistency for different types of shared memory by hardware means. Also support for point-to-point synchronization between the processor cores is realized implementing different hardware barriers. The practical examinations focus on the logical first step from single- to dual-core systems, using an FPGA-development board with two hard PowerPC processor cores. Best and worst-case results, together with intensive bench- marking of all synchronization primitives implemented, show the expected superiority of the hardware solutions. It is also shown that dual-ported memory outperforms single-ported memory if the multiple cores use inherent parallelism by locking shared memory more intelligently using an address-sensitive method. Christian Stoif, Martin Schoeberl, Benito Liccardi, Jan Haase 0001 |
ISCAS | 2 |
| 2011 | A Time-Predictable Object CacheabstractStatic cache analysis for data allocated on the heap is practically impossible for standard data caches. We propose a distinct object cache for heap allocated data. The cache is highly associative to track symbolic object addresses in the static analysis. Cache lines are organized to hold single objects and individual fields are loaded on a miss. This cache organization is statically analyzable and improves the performance. In this paper we present the design and implementation of the object cache in a uniprocessor and chip-multiprocessor version of the Java processor JOP. Martin Schoeberl |
ISORC | 1 |
| 2011 | Design Space Exploration of Object Caches with Cross-ProfilingabstractTo avoid data cache trashing between heap-allocated data and other data areas, a distinct object cache has been proposed for embedded real-time Java processors. This object cache uses high associativity in order to statically track different object pointers for worst-case execution-time analysis. However, before implementing such an object cache, an empirical analysis of different organization forms is needed. We use a cross-profiling technique based on aspect-oriented programming in order to evaluate different object cache organizations with standard Java benchmarks. From the evaluation we conclude that field access exhibits some temporal locality, but almost no spatial locality. Therefore, filling long cache lines on a miss just introduces a high miss penalty without increasing the hit rate enough to make up for the increased miss penalty. For an object cache, it is more efficient to fill individual words within the cache line on a miss. Martin Schoeberl, Walter Binder, Alex Villazón |
ISORC | 1 |
| 2011 | Introduction to the Special Issue: JTRES 2009abstractJava is a very successful language for several application domains: be it desktop applications, large enterprise server applications, web services, or applets for mobile phones. With the Real-Time Specification for Java (RTSJ) and the upcoming definition of a Safety-Critical Java profile, Java enters the domain of soft and hard real-time systems. The Workshop on Java Technologies for Real-Time and Embedded Systems (JTRES) is dedicated to research on Java for embedded real-time systems. JTRES was started in 2003 as a workshop attached to conference and since 2006 JTRES has been a successful standalone workshop. JTRES 2009, the 7th workshop in the JTRES series, was organized in Madrid, Spain by the Universidad Complutense de Madrid, Facultad de Informatica. This special issue of Concurrency and Computation: Practice and Experience contains invited papers from the JTRES 2009 workshop that have been expanded and carefully peer reviewed. The RTSJ is the main standard for real-time systems implemented in Java. Several commercial and research implementations of the RTSJ are available today. However, the optimal implementation of some features of the RTSJ is still an open research questions. The primary goal for asynchronous event handlers (AEH) in the RTSJ is to have a lightweight concurrency mechanism. In the article, Applying Fixed-Priority Preemptive Scheduling with Preemption Threshold to Asynchronous Event Handling in the RTSJ 1, Kim and Wellings propose a scheduling scheme for AEH that minimizes the number of server threads needed. In the article they first define the worst-case scenario that demands the least upper bound of servers for self-suspending and non-self-suspending handlers. Based on the worst-case scenarios, it is proved that the number of servers required to execute a given number of non-self-suspending handlers depends on the number of priority levels, not on the number of handlers. It is also shown that each self-suspending handler is required to have its own server to prevent unbounded priority inversion. The RTSJ supports two kinds of real-time threads: one that can operate on the garbage collected heap and another one that avoids influence by garbage collection by using a region-based memory management. Basanta-Val et al. present in their article, Extending the concurrency model of the Real-Time Specification for Java 2, a more flexible thread model, where the thread can switch between a heap and a non-heap mode. The authors propose a simple extension to the current threading model named RealtimeThread++, in an attempt to introduce more flexibility in the RTSJ concurrency model. The article describes the extension from several points of view: (i) the programmer, identifying scenarios that may benefit from it significantly; (ii) the real-time Java technology perspective, identifying changes required in the current real-time virtual machine to support it; and (iii) the accumulated experience, relating empirical results obtained from a software prototype that supports the extension. The current practice for soft real-time systems in Java is to avoid the cumbersome programming model of scoped memories and, instead, use a real-time garbage collector. Kalibera presents in his article, Replicating Real-Time Garbage Collector 3, an incremental garbage collector. The real-time collector has to relocate objects in the heap to avoid fragmentation. This is usually achieved via an indirection that has to be followed on every read and write to the heap. Kalibera presents an alternative solution, based on object replication, which does not need any special handling for memory reads, but writes are more expensive: every value is written twice. As writes are less frequent than reads, the total overhead is reduced. The presented technique targets uni-processor systems with green-threading embedded, allowing to design simpler and more predictable mutator barriers. The scoped memory of the RTSJ enables region-based dynamic memory management. Using scopes forces the programmer to reason in terms of locality and to correctly size the scopes. Gabervetsky et al. present in their article Quantitative dynamic-memory analysis for Java 4 tool support for RTSJ scopes. The tool synthesizes a scoped-based memory organization where regions are associated with methods and it infers their sizes in parametric forms in terms of relevant program variables. Furthermore, it exhibits a parametric upper bound on the total amount of memory required to run a method. Benchmarks are important to evaluate systems and also to guide enhancements of specific features. While for general-purpose computing many benchmark suits exist, there are only few benchmarks for embedded and real-time Java available. Kalibera et al. adapted a collision detection algorithm to provide a benchmark that can be used in various real-time and non-real-time Java settings, as described in their article, A Family of Real-time Java Benchmarks 5. The benchmark CDx is open-source to support researchers on real-time Java. CDx can be run on standard Java virtual machines, on RTSJ and Safety Critical Java virtual machines, and a C version is provided to compare with the native performance. We thank the authors contributing to this special issue, and we also thank all the reviewers whose dedication ensured a good selection of articles and made this special issue possible. This work was supported by Publishing Arts Research Council under grant number: 98-1846389. Martin Schoeberl, M. Teresa Higuera-Toledano |
Concurr. Comput. Pract. Exp. | 1 |
| 2011 | A Hardware Abstraction Layer in JavaabstractEmbedded systems use specialized hardware devices to interact with their environment, and since they have to be dependable, it is attractive to use a modern, type-safe programming language like Java to develop programs for them. Standard Java, as a platform-independent language, delegates access to devices, direct memory access, and interrupt handling to some underlying operating system or kernel, but in the embedded systems domain resources are scarce and a Java Virtual Machine (JVM) without an underlying middleware is an attractive architecture. The contribution of this article is a proposal for Java packages with hardware objects and interrupt handlers that interface to such a JVM. We provide implementations of the proposal directly in hardware, as extensions of standard interpreters, and finally with an operating system middleware. The latter solution is mainly seen as a migration path allowing Java programs to coexist with legacy system components. An important aspect of the proposal is that it is compatible with the Real-Time Specification for Java (RTSJ). Martin Schoeberl, Stephan Korsholm, Tomas Kalibera, Anders P. Ravn |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2010 | Design and Implementation of Real-Time Transactional MemoryabstractTransactional memory is a promising, optimistic synchronization mechanism for chip-multiprocessor systems. The simplicity of atomic sections, instead of using explicit locks, is also appealing for real-time systems. In this paper an implementation of real-time transactional memory (RTTM) in the context of a real-time Java chip-multiprocessor (CMP) is presented. To provide a predictable and analyzable solution of transactional memory, the transaction buffer is organized fully associative. Evaluation in an FPGA shows that an associativity of up to 64-way is possible without degrading the overall system performance. The paper presents synthesis results for different RTTM configurations and different number of processor cores in the CMP system. A CMP system with up to 8 processor cores with RTTM support is feasible in an Altera Cyclone-II FPGA. Martin Schoeberl, Peter Hilber |
FPL | 1 |
| 2010 | Worst-Case Analysis of Heap Allocations
Wolfgang Puffitsch, Benedikt Huber, Martin Schoeberl |
ISoLA (2) | 3 |
| 2010 | Scheduling of hard real-time garbage collection
Martin Schoeberl |
Real Time Syst. | 1 |
| 2010 | Worst-case execution time analysis for a Java processorabstractAbstract In this paper, we propose a solution for a worst‐case execution time (WCET) analyzable Java system: a combination of a time‐predictable Java processor and a tool that performs WCET analysis at Java bytecode level. We present a Java processor, called JOP, designed for time‐predictable execution of real‐time tasks. The execution time of bytecodes, the instructions of the Java virtual machine, is known to cycle accuracy for JOP. Therefore, JOP simplifies the low‐level WCET analysis. A method cache, which fills whole Java methods into the cache, simplifies cache analysis. The WCET analysis tool is based on integer linear programming. The tool performs the low‐level analysis at the bytecode level and integrates the method cache analysis. An integrated data‐flow analysis performs receiver‐type analysis for dynamic method dispatches and loop‐bound analysis. Furthermore, a model checking approach to WCET analysis is presented where the method cache can be exactly simulated. The combination of the time‐predictable Java processor and the WCET analysis tool is evaluated with standard WCET benchmarks and three real‐time applications. The WCET friendly architecture of JOP and the integrated method cache analysis yield tight WCET bounds. Comparing the exact, but expensive, model checking‐based analysis of the method cache with the static approach demonstrates that the static approximation of the method cache is sufficiently tight for practical purposes. Copyright © 2010 John Wiley & Sons, Ltd. Martin Schoeberl, Wolfgang Puffitsch, Rasmus Ulslev Pedersen, Benedikt Huber |
Softw. Pract. Exp. | 1 |
| 2010 | A real-time Java chip-multiprocessorabstractChip-multiprocessors are an emerging trend for embedded systems. In this article, we introduce a real-time Java multiprocessor called JopCMP. It is a symmetric shared-memory multiprocessor, and consists of up to eight Java Optimized Processor (JOP) cores, an arbitration control device, and a shared memory. All components are interconnected via a system on chip bus. The arbiter synchronizes the access of multiple CPUs to the shared main memory. In this article, three different arbitration policies are presented, evaluated, and compared with respect to their real-time and average-case performance: a fixed priority, a fair-based, and a time-sliced arbiter. Tasks running on different CPUs of a chip-multiprocessor (CMP) influence each others' execution times when accessing a shared memory. Therefore, the system needs an arbiter that is able to limit the worst-case execution time of a task running on a CPU, even though tasks executing simultaneously on other CPUs access the main memory. Our research shows that timing analysis is in fact possible for homogeneous multiprocessor systems with a shared memory. The timing analysis of tasks, executing on the CMP using time-sliced memory arbitration, leads to viable worst-case execution time bounds. The time-sliced arbiter divides the memory access time into equal time slots, one time slot for each CPU. This memory arbitration scheme allows for a calculation of upper bounds of Java application worst-case execution times, depending on the number of CPUs, the time slot size, and the memory access time. Examples of worst-case execution time calculation are presented, and the analyzed results of a real-world application task are compared to measured execution time results. Finally, we evaluate the tradeoffs when using a time-predictable solution compared to using average-case optimized chip-multiprocessors, applying three different benchmarks. These experiments are carried out by executing the programs on the CMP prototype. Christof Pitter, Martin Schoeberl |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2010 | Nonblocking real-time garbage collectionabstractA real-time garbage collector has to fulfill two basic properties: ensure that programs with bounded allocation rates do not run out of memory and provide short blocking times. Even for incremental garbage collectors, two major sources of blocking exist, namely, root scanning and heap compaction. Finding root nodes of an object graph is an integral part of tracing garbage collectors and cannot be circumvented. Heap compaction is necessary to avoid probably unbounded heap fragmentation, which in turn would lead to unacceptably high memory consumption. In this article, we propose solutions to both issues. Thread stacks are local to a thread, and root scanning, therefore, only needs to be atomic with respect to the thread whose stack is scanned. This fact can be utilized by either blocking only the thread whose stack is scanned, or by delegating the responsibility for root scanning to the application threads. The latter solution eliminates blocking due to root scanning completely. The impact of this solution on the execution time of a garbage collector is shown for two different variants of such a root scanning algorithm. During heap compaction, objects are copied. Copying is usually performed atomically to avoid interference with application threads, which could render the state of an object inconsistent. Copying of large objects and especially large arrays introduces long blocking times that are unacceptable for real-time systems. In this article, an interruptible copy unit is presented that implements nonblocking object copy. The unit can be interrupted after a single word move. We evaluate a real-time garbage collector that uses the proposed techniques on a Java processor. With this garbage collector, it is possible to run high-priority hard real-time tasks at 10 kHz parallel to the garbage collection task on a 100 MHz system. Martin Schoeberl, Wolfgang Puffitsch |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2009 | A disruptive computer design idea: Architectures with repeatable timingabstractThis paper argues that repeatable timing is more important and more achievable than predictable timing. It describes microarchitecture approaches to pipelining and memory hierarchy that deliver repeatable timing and promise comparable or better performance compared to established techniques. Specifically, threads are interleaved in a pipeline to eliminate pipeline hazards, and a hierarchical memory architecture is outlined that hides memory latencies. Stephen A. Edwards, Edward A. Lee, Isaac Liu, Hiren D. Patel, Martin Schoeberl |
ICCD | 6 |
| 2009 | Embedded JIT Compilation with CACAO on YARIabstractJava is one of the most popular programming languages for thedevelopment of portable workstation and server applications availabletoday. Because of its clean design and typesafety, it is alsobecoming attractive in the domain of embedded systems. Unfortunately, the dynamic features of the language and its rich class library causeconsiderable overhead in terms of runtime and memory consumption. Efficient techniques to implement Java virtual machines that aresuitable for use in resource constrained environments are thusneeded. In this work we present a solution for very restrictedenvironments based on CACAO. CACAO is a just-in-time compilingvirtual machine implementation, combining high speed and small size. We have modified the original version of CACAO to run without anunderlying operating system within only 1 MB of memory. In additionwe present a new technique to selectively compile methods during theinitialization phase of real-time Java applications to preventunwanted interaction between dynamic compilation and critical tasks. Furthermore we present the YARI soft-core as the execution platformof CACAO within an field-programmable gate array. We compare ourimplementation with two well known Java processors, JOP and Sun'spicoJava-II, on the same technology. Although JOP achieves a higherclock frequency and picoJava-II occupies nearly 4 times the resourceof YARI, our solution is capable to outperform both of them by afactor of up to 2.8 and 2.2 respectively. Florian Brandner, Tommy Thorn, Martin Schoeberl |
ISORC | 3 |
| 2009 | Thread-Local Scope Caching for Real-time JavaabstractThere is increasing convergence between the fields of parallel and embedded computing. The demand for more functionality in embedded devices means that complex multicore architectures will be used. In order to promote scalability and obtain predictability, on-chip processor-local private memory subsystems will be used. Whilst at the hardware level this is technical feasible, the more pressing problem is how such memory is presented to the programmer and how its local access is policed.In this paper we illustrate how Java augmented by the Real-time Specification for Java can be used to present the abstraction of a thread-local scoped memory area. We show how to enforce access to the memory area to a single real-time thread. We implement the model on the JOP multiprocessor system and report on our experiences. Andy J. Wellings, Martin Schoeberl |
ISORC | 2 |
| 2009 | Cross-profiling for Java processorsabstractAbstract Performance evaluation of embedded software is essential in an early development phase so as to ensure that the software will run on the embedded device's limited computing resources. The prevailing approaches either require the deployment of the software on the embedded target, which can be tedious and may be impossible in an early development phase, or rely on simulation, which can be very slow. In this article, we introduce a customizable cross‐profiling framework for embedded Java processors, including processors featuring a method cache. The developer profiles the embedded software in the host environment, completely decoupled from the target system, on any standard Java virtual machine, but the generated profiles represent the execution time metric of the target system. Our cross‐profiling framework is based on bytecode instrumentation. We identify several pointcuts in the execution of bytecode that need to be instrumented in order to estimate the CPU cycle consumption on the target system. An evaluation using the JOP embedded Java processor as target confirms that our approach reconciles high profile accuracy with moderate overhead. Our cross‐profiling framework also enables the performance evaluation of new processor architectures before they are implemented. As a case study, we explore the performance impact of various processor design choices and optimizations, such as different cache sizes or pipeline organizations, and come up with an improved processor design that yields speedups of up to 40% on standard Java benchmarks. Copyright © 2009 John Wiley & Sons, Ltd. Walter Binder, Martin Schoeberl, Philippe Moret, Alex Villazón |
Softw. Pract. Exp. | 2 |
| 2008 | Cache-aware cross-profiling for java processorsabstractPerformance evaluation of embedded software is essential in an early development phase so as to ensure that the software will run on the embedded device's limited computing resources. Prevailing approaches either require the deployment of the software on the embedded target, which can be tedious and may be impossible in an early development phase, or rely on simulation, which can be very slow. In this paper, we introduce a customizable cross-profiling framework for embedded Java processors, including processors featuring a method cache. The developer profiles the embedded software in the host environment, completely decoupled from the target system, on any standard Java Virtual Machine, but the generated profiles represent the execution time metric of the target system. Our cross-profiling framework is based on bytecode instrumentation. We identify several pointcuts in the execution of bytecode that need to be instrumented in order to estimate the CPU cycle consumption on the target system. An evaluation using the JOP embedded Java processor as target confirms that our approach reconciles high profile accuracy with moderate overhead. Our cross-profiling framework also enables the rapid evaluation of the performance impact of possible optimizations, such as different caching strategies. Walter Binder, Alex Villazón, Martin Schoeberl, Philippe Moret |
CASES | 3 |
| 2008 | Toward Libraries for Real-Time JavaabstractReusable libraries are problematic for real-time software in Java. Using Java's standard class library, for example, demands meticulous coding and testing to avoid response time spikes and garbage collection. We propose two design requirements for reusable libraries in real-time systems: worst-case execution time (WCET) bounds and worst- case memory consumption bounds. Furthermore, WCET cannot be known if blocking method calls are used. We have applied these requirements to the design of three Java-based prototypes: a set of collection classes, a networking stack, and trigonometric functions. Our prototypes show that reusable libraries can meet these requirements and thus be viable for real-time systems. Trevor Harmon, Martin Schoeberl, Raimund Kirner, Raymond Klefstad |
ISORC | 2 |
| 2008 | Interrupt Handlers in JavaabstractAn important part of implementing device drivers is to control the interrupt facilities of the hardware platform and to program interrupt handlers. Current methods for handling interrupts in Java use a server thread waiting for the VM to signal an interrupt occurrence. It means that the interrupt is handled at a later time, which has some disadvantages. We present constructs that allow interrupts to be handled directly and not at a later point decided by a scheduler. A desirable feature of our approach is that we do not require a native middelware layer but can handle interrupts entirely with Java code. We have implemented our approach using an interpreter and a Java processor, and give an example demonstrating its use. Stephan Korsholm, Martin Schoeberl, Anders P. Ravn |
ISORC | 2 |
| 2008 | Hardware Objects for JavaabstractJava, as a safe and platform independent language, avoids access to low-level I/O devices or direct memory access. In standard Java, low-level I/O is not a concern; it is handled by the operating system. However, in the embedded domain resources are scarce and a Java virtual machine (JVM) without an underlying middleware is an attractive architecture. When running the JVM on bare metal, we need access to I/O devices from Java; therefore we investigate a safe and efficient mechanism to represent I/O devices as first class Java objects, where device registers are represented by object fields. Access to those registers is safe as Java's type system regulates it. The access is also fast as it is directly performed by the bytecodesgetfield and putfield. Hardware objects thus provide an object-oriented abstraction of low-level hardware devices. As a proof of concept, we have implemented hardware objects in three quite different JVMs: in the Java processor JOP, the JIT compiler CACAO, and in the interpreting embedded JVM SimpleRTJ. Martin Schoeberl, Christian Thalinger, Stephan Korsholm, Anders P. Ravn |
ISORC | 1 |
| 2008 | A Modular Worst-case Execution Time Analysis Tool for Java ProcessorsabstractRecent technologies such as the real-time specification for Java promise to bring Java's advantages to real-time systems. While these technologies have made Java more predictable, they lack a crucial element: support for determining the worst-case execution time (WCET). Without knowledge of WCET, the correct temporal behavior of a Java program cannot be guaranteed. Although considerable research has been applied to the theory of WCET analysis, implementations are much less common, particularly for Java. Recognizing this deficiency, we have created an open-source, extensible tool that supports WCET analysis of Java programs. Designed for flexibility, it is built around a plug- in model that allows features to be incorporated as needed. Users can plug in various processor models, loop bound detectors, and WCET analysis algorithms without having to understand or alter the tool's internals. Trevor Harmon, Martin Schoeberl, Raimund Kirner, Raymond Klefstad |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2008 | A Java processor architecture for embedded real-time systems
Martin Schoeberl |
J. Syst. Archit. | 1 |
| 2007 | Modeling the Function Cache for Worst-Case Execution Time AnalysisabstractStatic worst-case execution time (WCET) analysis is done by modeling the hardware behavior. In this paper we describe a WCET analysis technique to analyze systems with function caches, a special kind of instruction cache that caches whole functions only. This cache was designed with the aim to be more predictable for the worst-case than existing instruction caches. Within this paper we developed a cache analysis technique for the function cache. One of the new concepts of this analysis technique is the local persistence analysis, which allows to precisely model the function cache. Raimund Kirner, Martin Schoeberl |
DAC | 2 |
| 2007 | Time Predictable CPU and DMA Shared Memory AccessabstractIn this paper, we propose a first step towards a time predictable computer architecture for single-chip multiprocessing (CMP). CMP is the actual trend in server and desktop systems. CMP is even considered for embedded realtime systems, where worst-case execution time (WCET) estimates are of primary importance. We attack the problem of WCET analysis for several processing units accesing a shared resource (the main memory) by support from the hardware. In this paper, we combine a time predictable Java processor and a direct memory access (DMA) unit with a regular access pattern (VGA controller). We analyze and evaluate different arbitration schemes with respect to schedulability analysis and WCET analysis. We also implement the various combinations in an FPGA. An FPGA is the ideal platform to verify the different concepts and evaluate the results by running applications with industrial background in real hardware. Christof Pitter, Martin Schoeberl |
FPL | 2 |
| 2007 | A Time-Triggered Network-on-ChipabstractIn this paper we propose a time-triggered network-on-chip (NoC) for on-chip real-time systems. The NoC provides time predictable on-and off-chip communication, a mandatory feature for dependable real-time systems. A regular structured NoC with a pseudo-static communication schedule allows for a high bandwidth. In this paper we argue for a simple, time-triggered NoC structure to achieve maximum bandwidth. We have implemented the proposed TT-NoC in a low-cost FPGA. The base bandwidth is 29 Gbit/s and the peak bandwidth 230 Gbit/s for eight nodes. The idea is in line with current on-chip multiprocessor designs, such as the Cell processor. The simple design of the network and the network interface easies certification of the proposed NoC for safety critical applications. Martin Schoeberl |
FPL | 1 |
| 2007 | A Profile for Safety Critical JavaabstractWe propose a new, minimal specification for real-time Java for safety critical applications. The intention is to provide a profile that supports programming of applications that can be validated against safety critical standards such as DO-178B (1992). The proposed profile is in line with the Java specification request JSR-302: Safety Critical Java Technology, which is still under discussion. In contrast to the current direction of the expert group for the JSR-302 we do not subset the rather complex Real-Time Specification for Java (RTSJ). Nevertheless, our profile can be implemented on top of an RTSJ compliant JVM Martin Schoeberl, Hans Søndergaard, Bent Thomsen, Anders P. Ravn |
ISORC | 1 |
| 2006 | A time predictable Java processorabstractThis paper presents a Java processor, called JOP, designed for time-predictable execution of real-time tasks. JOP is the implementation of the Java virtual machine in hardware. We propose a processor architecture that favors low worst-case execution time (WCET) over average case performance. The resulting processor is an easy target for the low-level WCET analysis Martin Schoeberl |
DATE | 1 |
| 2006 | Real-Time Garbage Collection for JavaabstractAutomatic memory management or garbage collection greatly simplifies the development of large systems. However, garbage collection is usually not used in real-time systems due to the unpredictable temporal behavior of current implementations of a garbage collector. In this paper we propose a concurrent collector that is scheduled periodically in the same way as ordinary application threads. We provide an upper bound for the collector period so that the application threads never run out of memory Martin Schoeberl |
ISORC | 1 |
| 2004 | Java Technology in an FPGA
Martin Schoeberl |
FPL | 1 |
| 2004 | Restrictions of Java for Embedded Real-Time SystemsabstractJava, with its pragmatic approach to object orientation and enhancements over C, got very popular for desktop and server application development. The productivity increment of up to 40% compared with C++ [E. Quinn et al., (1998)] attracts also embedded systems programmers. However, standard Java is not practical on these usually small devices. This paper presents the status of restricted Java environments for embedded and real-time systems. For missing definitions, additional profiles are proposed. Results of the implementation on a Java processor show that it is possible to develop applications in pure Java on resource constraint devices Martin Schoeberl |
ISORC | 1 |