EDBT 2026 Demo / reviewers in the wild / expert
Emad Jacob Maroun
dblp:216/1942
· DBLP profile ↗
9ranked-venue papers
9as first author
8since 2021 · last 2026
0000-0002-3675-3376ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Shared caches for mixed-criticality 5G radio base stationabstractThe advancement of telecommunication technology is necessary to support the rapid development of societies. The latest standard in telecommunications, 5G, offers unprecedented communication speeds, reach, and quality. The 5G standard supports both critical and non-critical telecommunications. The baseband units—handling wireless signal transmission and reception—divide hardware resources to ensure that high- and low-criticality communication tasks do not interfere with each other. This results in underutilized hardware platforms with high redundancy. In this paper, we propose two level 2 cache architectures that allow a single system to handle high- and low-criticality tasks and improve resource utilization while prioritizing critical tasks. The contention-tracking cache tracks contention among tasks of differing criticality for shared resources and locks those resources for critical tasks once a contention limit is exceeded. The criticality timeout cache sets timers for cache lines accessed by critical tasks and decrements them as long as the lines remain unused. Non-critical tasks are not allowed to evict these lines until the timer has finished. We evaluate the behavior of the proposed cache designs using a simulation framework and a real-world workload from a 5G baseband unit. Additionally, we implement the proposed caches in hardware, along with a baseline PLRU cache, and evaluate their performance and resource usage using the T-CREST platform. The simulation and performance results confirm that the proposed cache architectures can prioritize critical memory access, while allowing non-critical tasks to utilize the cache’s free space. Emad Jacob Maroun, Arijus Grotuzas, Martin Schoeberl |
J. Syst. Archit. | 1 |
| 2025 | Simulating Contention and Timeout Caches for a Mixed-Criticality 5G Radio Base StationabstractTelecommunication is a critical driver of economic and social development. 5G technologies are state-of-the-art in telecommunication, setting strong and open-ended requirements for implementing systems. Current systems for implementing baseband technologies in 5 G depend on hardware separation to ensure high-criticality tasks and low-criticality tasks do not interfere in such a way as to violate guarantees. To allow for the merging of high- and low-criticality systems into one, this paper presents two level-2 cache architectures: The contention tracking cache tracks contention events between highand low-criticality tasks, blocking any further contention if a specified limit is reached. The criticality timeout cache associates a timer with each cache line for high-criticality tasks that counts down as long as the line is not reused. During the countdown, low-criticality tasks are prohibited from evicting high-criticality cache lines. When the timer runs out, this prohibition is lifted. A simulation framework is employed to accurately model the behavior of cores accessing a memory hierarchy with various cache types. Simulation statistics demonstrate that the proposed cache architectures effectively prioritize memory accesses for critical tasks while allowing non-critical tasks to utilize any available cache space. Emad Jacob Maroun, Martin Schoeberl |
ISORC | 1 |
| 2025 | Optimizing Stack Cache Usage Through Local Variable PromotionabstractReal-time systems require high confidence in execution times. Static analysis techniques provide this confidence by accounting for the specific execution platform and code being executed to provide an upper bound on the worst-case execution time. To make analysis easier and more accurate, it is imperative to design the platform to be predictable. One such design is the Patmos processor's split caches: The traditional data cache is split into a stack cache and a data cache. The stack cache stores data from the program stack. This data is highly predictable and easy to reason about for a static analyzer. The current implementation of the Patmos compiler does not use the stack cache optimally; caching only spilled registers as part of register allocation. This paper presents the local variable promotion optimization, which enables the Patmos compiler to store function-local variables on the stack cache. Our results show that this optimization is critical for any platform using a stack cache. While not all benchmark programs see a significant performance benefit, those that did saw increases mostly between 50% and 60%. Emad Jacob Maroun, Daniel Seifert, Johannes Gernedl |
ISORC | 1 |
| 2025 | Optimized Constant Execution Time CodeabstractSingle-path code aims to make WCET analysis easier by eliminating data-dependent control flow. To completely negate the need for WCET analysis, single-path code must also eliminate execution-time variability from memory accesses. To be practically useful, single-path code must be optimized to be competitive with traditional WCET-analyzed code. This paper summarizes the work in Emad Jacob Maroun's PhD dissertation titled”Compiling for Time-Predictability and Performance“. Memory access compensation ensures that singlepath code exhibits constant execution time. The generated code is optimized using an improved transformation that uses generic allocators for general-purpose and predicate registers. The repetition dominance relation is used to reduce unnecessary code execution. Lastly, a heuristic list scheduler enables single-path code to utilize the second issue slot of a dual-issue processor. In addition to achieving constant execution times on a timepredictable processor, the results show varying but significant improvements of up to 145 % in performance and a reduced code size of up to 28 %. Compared to WCET-analyzed traditional code, single-path code is mostly competitive while outright superior in several cases. However, pathological cases of poor performance are still observed. Emad Jacob Maroun, Martin Schoeberl, Peter P. Puschner |
ISORC | 1 |
| 2024 | Two-Step Register Allocation for Implementing Single-Path CodeabstractRegister allocation is a crucial step in the compilation pipeline that decides what program values occupy which physical registers. Single-path code’s use of predicated instructions instead of branching control-flow means register allocation must also allocate predicate registers. In this paper, we improve the original single-path transformation to allow generic register allocators to allocate predicate registers. Our improved transformation splits register allocation into two. First, the general-purpose registers are allocated as usual using a generic register allocator. Then, the main steps of the single-path transformation are performed while still using virtual predicate registers. Lastly, register allocation is rerun using the generic allocator to allocate the predicate registers. Our results show the improved single-path transformation increasing performance by up to 80 % and reducing code size by up to 43 % compared to the original transformation that uses a custom predicate allocator. Emad Jacob Maroun, Martin Schoeberl, Peter P. Puschner |
ISORC | 1 |
| 2024 | Predictable and optimized single-path code for predicated processorsabstractSingle-path code is a code generation technique for real-time systems that reduces execution time variability. However, doing so can incur significant execution-time overhead and does not guarantee constant execution times. In this paper, we address the performance challenges of single-path code and solve the variability issue. We present the repetition dominance relation to identify and optimize code blocks that are always executed a fixed number of times. We show that single-path code’s instructions are uniquely easy to schedule, and we explore an extension to the Patmos architecture that allows additional instruction types in the second issue slot. Lastly, we present two techniques for ensuring that functions always perform the same number of accesses to memory, resulting in programs with constant execution time. We compare the performance of single-path code to that of statically analyzed traditional code. Our results show that single-path code’s performance is mostly competitive while outright superior in several cases. However, pathological cases of poor performance are still observed. Emad Jacob Maroun, Martin Schoeberl, Peter P. Puschner |
J. Syst. Archit. | 1 |
| 2023 | Compiler-Directed Constant Execution Time on Flat Memory SystemsabstractTime predictability is a central requirement for real-time systems. The correct behavior of such a system can only be achieved if the results of programs are ready in time to affect the environment. Execution times of modern systems can vary for many reasons, meaning complex analyses must be performed to ensure that the execution time is bounded and that a task always finishes before its deadline. Care must also be taken to ensure that nefarious actors do not exploit the varying execution time to compromise the system’s integrity. Avoiding variable execution times can greatly simplify systems, is inherently more secure, and eliminates the need for complex analyses. In this paper, we first argue for the value of having programs with constant execution times. We then show how the memory system around a processing core can affect execution times even on systems without intermediate storage like caches or scratch-pads. We present automatic compiler techniques for generating constant execution time programs and evaluate their implementation on the Patmos architecture. We show that combining our two compensation techniques is generally superior to either on their own. We compare the performance of our implementation to the estimates produced by the Platin worst-case execution time analyzer. While our implementation significantly impacts performance, it is generally manageable and has the potential for comparable execution times. Emad Jacob Maroun, Martin Schoeberl, Peter P. Puschner |
ISORC | 1 |
| 2021 | Compiling for time-predictability with dual-issue single-path codeabstractDesigned for real-time systems, the Patmos instruction-set architecture's features ensure a high degree of predictability.One such feature is its dual-issue pipeline, which can issue and execute bundles of up to two instructions at a time.Executing instructions in the second issue slot is a predictable way to increase the throughput of a processor, but without dedicated support from the compiler, this benefit cannot be unlocked.A compiler generates highly predictable programs by generating single-path code.This technique produces code that always follows the same trace of instructions.While Patmos' compiler can already produce single-path code, it does not assign any instructions to the second issue-slot.This limitation is unfortunate, as single-path code inherently possesses a high degree of instruction-level parallelism.In this paper, we present a singlepath code generation technique with support for dual-issue pipelines.It can also support different bundling algorithms, which allows changing algorithms without having to edit other parts of the compiler.We present a simple bundling algorithm plugged into the single-path code generator.It looks for branches and bundles the basic blocks on each path of the branch.While this specific bundling algorithm is too simple to provide a real-world benefit, it highlights the potential that further work on bundling algorithms can unlock. Emad Jacob Maroun, Martin Schoeberl, Peter P. Puschner |
J. Syst. Archit. | 1 |
| 2020 | Towards Dual-Issue Single-Path CodeabstractThe Patmos instruction-set architecture is designed for real-time systems. As such, it has features that increase the predictability of code running on it. One important feature is its dual-issue pipeline: instructions may be organized in bundles of two that are issued and executed in parallel. This increases the throughput of the processor in a predictable manner, but only if the compiler makes use of it.Single-path code is a code-generation technique that produces predictable executions by always following the same trace of instructions. The Patmos compiler can already produce single-path code, but it does not use the second issue slot available in the processor. This is less than ideal because the single-path transformation results in code that has a high degree of instruction-level parallelism.In this paper, we present a single-path code generator that can produce bundled instructions. It includes generic support for bundling algorithms, such that implementing them is simple and does not require changing other parts of the compiler.We also present one such bundling algorithm plugged into the single-path code generator. With it, we show that we can produce dual-issue instructions to improve performance. Emad Jacob Maroun, Martin Schoeberl, Peter P. Puschner |
ISORC | 1 |