EDBT 2026 Demo / reviewers in the wild / expert
Alexander Züpke
dblp:162/7703 · also Alexander Zuepke
· DBLP profile ↗
15ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0003-0134-318XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ETM2: Empowering Traditional Memory Bandwidth Regulation using ETM
Alexander Züpke, Ashutosh Pradhan, Daniele Ottaviano, Andrea Bastoni, Marco Caccamo |
RTAS | 1 |
| 2025 | Enabling Security on the Edge: A CHERI Compartmentalized Network StackabstractThe widespread deployment of embedded systems in critical infrastructures, interconnected edge devices like autonomous drones, and smart industrial systems requires robust security measures. Compromised systems increase the risks of operational failures, data breaches, and-in safety-critical environments-potential physical harm to people. Despite these risks, current security measures are often insufficient to fully address the attack surfaces of embedded devices. CHERI provides strong security from the hardware level by enabling fine-grained compartmentalization and memory protection, which can reduce the attack surface and improve the reliability of such devices. In this work, we explore the potential of CHERI to compartmentalize one of the most critical and targeted components of interconnected systems: their network stack. Our case study examines the tradeoffs of isolating applications, TCP/IP libraries, and network drivers on a CheriBSD system deployed on the Arm Morello platform. Our results suggest that CHERI has the potential to enhance security while maintaining performance in embedded-like environments. Donato Ferraro, Andrea Bastoni, Alexander Züpke, Andrea Marongiu |
DATE | 3 |
| 2025 | Arm Dynamiq Shared Unit and Real-Time: An Empirical EvaluationabstractThe increasing complexity of embedded hardware platforms poses significant challenges for real-time workloads. Architectural features such as Intel RDT, Arm QoS, and Arm MPAM are either unavailable on commercial embedded platforms or designed primarily for server environments optimized for average-case performance and might fail to deliver the expected real-time guarantees. Arm DynamIQ Shared Unit (DSU) includes isolation features-among others, hardware per-way cache partitioning-that can improve the real-time guarantees of complex embedded multicore systems and facilitate real-time analysis. However, the DSU also targets average cases, and its real-time capabilities have not yet been evaluated. This paper presents the first comprehensive analysis of three real-world deployments of the Arm DSU on Rockchip RK3568, Rockchip RK3588, and NVIDIA Orin platforms. We integrate support for the DSU at the operating system and hypervisor level and conduct a large-scale evaluation using both synthetic and real-world benchmarks with varying types and intensities of interference. Our results make extensive use of performance counters and indicate that, although effective, the quality of partitioning and isolation provided by the DSU depends on the type and the intensity of the interfering workloads. In addition, we uncover and analyze in detail the correlation between benchmarks and different types and intensities of interference. Ashutosh Pradhan, Daniele Ottaviano, Haozheng Huang, Alexander Züpke, Andrea Bastoni, Marco Caccamo |
RTAS | 5 |
| 2025 | Predictable Memory Bandwidth Regulation for DynamIQ Arm Systems
Ashutosh Pradhan, Daniele Ottaviano, Haozheng Huang, Alexander Züpke, Andrea Bastoni, Marco Caccamo |
RTCSA | 6 |
| 2025 | Work-in-Progress: Toward Real-Time Cross-ISA Execution on the AMD Embedded+ ArchitectureabstractEmerging embedded platforms increasingly rely on heterogeneous processing units to address diverse performance and energy requirements. The recently introduced AMD Embedded+ architecture reflects this trend by interconnecting via PCIe on the same motherboard one AMD x86 host processor with one Arm AArch64+FPGA complex. This implementation is another step forward towards a more compact heterogeneousISA platform designed with embedded applications in mind. While cross-ISA execution has been explored in the past with a focus on performance, programmability, and energy efficiency, its potential for embedded and predictable real-time workloads remains largely unexplored. In this paper, we start exploring such potential by investigating the real-time capabilities of the first commercial platform based on the AMD Embedded+ architecture: the Sapphire Edge+. We (1) outline key research challenges and real-time use-cases, (2) discuss suitable software architectures for the use-cases and highlight associated trade-offs, and (3) report an initial assessment of the potential of such architectures and use-cases via an experimental evaluation of latency and bandwidth on the real hardware. Lukas Neef, Daniele Ottaviano, Denis Hoornaert, Alexander Züpke, Marco Caccamo, Andrea Bastoni |
RTSS | 4 |
| 2025 | Work-in-Progress: A First Practical Look at Arm's MPAM for Real-Time SystemsabstractArm's Memory Partitioning and Monitoring (MPAM) extension introduces standardized mechanisms for partitioning cache and memory bandwidth. From a real-time systems perspective, this can aid in improving predictability in heterogeneous MPSoCs. In this paper, we present the first practical evaluation of MPAM on a COTS platform—the Radxa Orion O6 with the CIX CD8180 SoC. We characterize the SoC's MPAM capabilities and experimentally assess cache portion partitioning and proportional stride memory bandwidth partitioning under controlled interference. Our results show that enabling MPAM features can reduce interference, but their behavior often diverges from expectations based on the specification, with anomalous effects observed across workloads and cores. These findings highlight both the promise of predictability from MPAM for real-time systems and the current challenges arising from optionality, heterogeneity, and limited documentation. We conclude that broader evaluation across future MPAM-enabled SoCs, aided by detailed performance counter analysis, is essential to establish MPAM's practical value for real-time practitioners. Ashutosh Pradhan, Daniele Ottaviano, Alexander Züpke, Andrea Bastoni, Marco Caccamo |
RTSS | 3 |
| 2024 | A Containerized Microservice Architecture for a ROS 2 Autonomous Driving Software: An End-to-End Latency EvaluationabstractThe automotive industry is transitioning from traditional ECU-based systems to software-defined vehicles. A central role of this revolution is played by containers, lightweight virtualization technologies that enable the flexible consolidation of complex software applications on a common hardware platform. Despite their widespread adoption, the impact of containerization on fundamental real-time metrics such as end-to-end latency, communication jitter, as well as memory and CPU utilization has remained virtually unexplored. This paper presents a microservice architecture for a real-world autonomous driving application where containers isolate each service. Our comprehensive evaluation shows the benefits in terms of end-to-end latency of such a solution even over standard bare-Linux deployments. Specifically, in the case of the presented microservice architecture, the mean end-to-end latency can be improved by 5–8%. Also, the maximum latencies were significantly reduced using container deployment. Tobias Betz, Long Wen 0003, Fengjunjie Pan, Gemb Kaljavesi, Alexander Züpke, Andrea Bastoni, Marco Caccamo, Alois C. Knoll, Johannes Betz |
RTCSA | 5 |
| 2024 | Coherence-Aided Memory Bandwidth RegulationabstractWith the increasing adoption of PS-PL (Processor System-Programmable Logic) platforms, also known as CPU+FPGA systems, there arises a need for efficient resource management strategies. This work explores memory bandwidth regulation in such systems, leveraging the capabilities of tightly coupled FPGAs to offer elegant, low-overhead solutions with highly flexible regulation policies. We introduce MemCoRe, a novel approach that exploits the FPGA’s interaction with cache coherence interfaces and cross-trigger signals to achieve finegrained spatiotemporal awareness of processor activity and software-free control. By comparing MemCoRe with state-of-theart software-based approaches, namely MemGuard and MemPol, we demonstrate significant improvements in regulation precision and overhead reduction. Key contributions include nanosecondscale memory bandwidth regulation, off-core memory bandwidth accounting, address-aware regulation, low-overhead token-bucket regulation, and asymmetric on-off core throttling. Our evaluation on a Xilinx Zynq UltraScale+ ZCU102 CPU+FPGA platform showcases MemCoRe’s capability to regulate memory bandwidth with nanosecond-scale precision. Overall, MemCoRe presents a promising avenue for efficient memory bandwidth regulation in PS-PL platforms, with strong applicability to real-time systems. Ivan Izhbirdeev, Denis Hoornaert, Weifan Chen 0003, Alexander Züpke, Youssef Hammad, Marco Caccamo, Renato Mancuso 0001 |
RTSS | 4 |
| 2024 | MemPol: polling-based microsecond-scale per-core memory bandwidth regulationabstractAbstract In today’s multiprocessor systems-on-a-chip, the shared memory subsystem is a known source of temporal interference. The problem causes logically independent cores to affect each other’s performance, leading to pessimistic worst-case execution time analysis. Memory regulation via throttling is one of the most practical techniques to mitigate interference. Traditional regulation schemes rely on a combination of timer and performance counter interrupts to be delivered and processed on the same cores running real-time workload. Unfortunately, to prevent excessive overhead, regulation can only be enforced at a millisecond-scale granularity. In this work, we present a novel regulation mechanism from outside the cores that monitors performance counters for the application core’s activity in main memory at a microsecond scale. The approach is fully transparent to the applications on the cores, and can be implemented using widely available on-chip debug facilities. The presented mechanism also allows more complex composition of metrics to enact load-aware regulation. For instance, it allows redistributing unused bandwidth between cores while keeping the overall memory bandwidth of all cores below a given threshold. We implement our approach on a host of embedded platforms and conduct an in-depth evaluation on the Xilinx Zynq UltraScale+ ZCU102, NXP i.MX8M and NXP S32G2 platforms using the San Diego Vision Benchmark Suite. Alexander Züpke, Andrea Bastoni, Weifan Chen 0003, Marco Caccamo, Renato Mancuso 0001 |
Real Time Syst. | 1 |
| 2023 | MemPol: Policing Core Memory Bandwidth from Outside of the CoresabstractIn today’s multiprocessor systems-on-a-chip (MP- SoC), the shared memory subsystem is a known source of temporal interference. The problem causes logically independent cores to affect each other’s performance, leading to pessimistic worstcase execution time (WCET) analysis. One of the most practical techniques to mitigate interference is memory regulation via throttling. Traditional regulation schemes rely on a combination of timer and performance counter interrupts to be delivered and processed on the same cores running real-time workload. Unfortunately, to prevent excessive overhead, regulation can only be enforced at a millisecond-scale granularity. In this work, we present a novel regulation mechanism from outside the cores that monitors performance counters for the application core’s activity in main memory at a microsecond scale. The approach is fully transparent to the applications on the cores, and can be implemented using widely available onchip debug facilities. The presented mechanism also allows more complex composition of metrics to enact load-aware regulation. For instance, it allows redistributing unused bandwidth between cores while keeping the overall memory bandwidth of all cores below a given threshold. We implement our approach on a host of embedded platforms and carry out an in-depth evaluation on the Xilinx Zynq UltraScale+ZCUl02 platform using the SD-VBS. Alexander Züpke, Andrea Bastoni, Weifan Chen 0003, Marco Caccamo, Renato Mancuso 0001 |
RTAS | 1 |
| 2023 | Compositional verification of embedded real-time systemsabstractIn an embedded real-time system (ERTS), real-time tasks (software) are typically executed on a multicore shared-memory platform (hardware). The number of cores is usually small, contrasted with a larger number of complex tasks that share data to collaborate. Since most ERTSs are safety-critical, it is crucial to rigorously verify their software against various real-time requirements under the actual hardware constraints (concurrent access to data, number of cores). Both the real-time systems and the formal methods communities provide elegant techniques to realize such verification, which nevertheless face major challenges. For instance, model checking (formal methods) suffers from the state-space explosion problem, whereas schedulability analysis (real-time systems) is pessimistic and restricted to simple task models and schedulability properties. In this paper, we propose a scalable and generic approach to formally verify ERTSs. The core contribution is enabling, through joining the forces of both communities, compositional verification to tame the state-space size. To that end, we formalize a realistic ERTS model where tasks are complex with an arbitrary number of jobs and job segments, then show that compositional verification of such model is possible, using a hybrid approach (from both communities), under the state-of-the-art partitioned fixed-priority (P-FP) with limited preemption scheduling algorithm. The approach consists of the following steps, given the above ERTS model and scheduling algorithm. First, we compute fine-grained data sharing overheads for each job segment that reads or writes some data from the shared memory. Second, we generalize an algorithm that, aware of the data sharing overheads, computes an affinity (task-core allocation) guaranteeing the schedulability of hard-real-time (HRT) tasks. Third, we devise a timed automata (TA) model of the ERTS, that takes into account the affinity, the data sharing overheads and the scheduling algorithm, on which we demonstrate that various properties can be verified compositionally, i.e., on a subset of cores instead of the whole ERTS, therefore reducing the state-space size. In particular, we enable the scalable computation of tight worst-case response times (WCRTs) and other tight bounds separating events on different cores, thus overcoming the pessimism of schedulability analysis techniques. We fully automate our approach and show its benefits on three real-world complex ERTSs, namely two autonomous robots and an automotive case study from the WATERS 2017 industrial challenge. Mohammed Foughali, Pierre-Emmanuel Hladik, Alexander Züpke |
J. Syst. Archit. | 3 |
| 2021 | A Real-Time Virtio-Based Framework for Predictable Inter-VM CommunicationabstractEnsuring real-time properties on current heterogeneous multiprocessor systems on a chip is a challenging task. Furthermore, online artificial intelligent applications –which are routinely deployed on such chips– pose increasing pressure on the memory subsystem that becomes a source of unpredictability. Although techniques have been proposed to restore independent access to memory for concurrently executing virtual machines (VM), providing predictable inter-VM communication remains challenging. In this work, we tackle the problem of predictably transferring data between virtual machines and virtualized hardware resources on multiprocessor systems on chips under consideration of memory interference. We design a "broker-based" real-time communication framework for otherwise isolated virtual machines, provide a virtio-based reference implementation on top of the Jailhouse hypervisor, assess its overheads for FreeRTOS virtual machines, and formally analyze its communication flow schedulability under consideration of the implementation overheads. Furthermore, we define a methodology to assess the maximum DRAM memory saturation empirically, evaluate the framework’s performance and compare it with the theoretical schedulability. Gero Schwäricke, Rohan Tabish, Rodolfo Pellizzoni, Renato Mancuso 0001, Andrea Bastoni, Alexander Züpke, Marco Caccamo |
RTSS | 6 |
| 2020 | Turning Futexes Inside-Out: Efficient and Deterministic User Space Synchronization Primitives for Real-Time Systems with IPCPabstractIn Linux and other operating systems, futexes (fast user space mutexes) are the underlying synchronization primitives to implement POSIX synchronization mechanisms, such as blocking mutexes, condition variables, and semaphores. Futexes allow one to implement mutexes with excellent performance by avoiding system calls in the fast path. However, futexes are fundamentally limited to synchronization mechanisms that are expressible as atomic operations on 32-bit variables. At operating system kernel level, futex implementations require complex mechanisms to look up internal wait queues making them susceptible to determinism issues. In this paper, we present an alternative design for futexes by completely moving the complexity of wait queue management from the operating system kernel into user space, i. e. we turn futexes "inside out". The enabling mechanisms for "inside-out futexes" are an efficient implementation of the immediate priority ceiling protocol (IPCP) to achieve non-preemptive critical sections in user space, spinlocks for mutual exclusion, and interwoven services to suspend or wake up threads. The design allows us to implement common thread synchronization mechanisms in user space and to move determinism concerns out of the kernel while keeping the performance properties of futexes. The presented approach is suitable for multi-processor real-time systems with partitioned fixed-priority (P-FP) scheduling on each processor. We evaluate the approach with an implementation for mutexes and condition variables in a real-time operating system (RTOS). Experimental results on 32-bit ARM platforms show that the approach is feasible, and overheads are driven by low-level synchronization primitives. Alexander Züpke |
ECRTS | 1 |
| 2019 | Deterministic Futexes: Addressing WCET and Bounded Interference ConcernsabstractFast User Space Mutexes (Futexes) in Linux are a lightweight mechanism to implement thread synchronization objects like mutexes and condition variables. Since they handle the uncontended case in user space and call the operating system kernel only for suspension and wake-up on contention, futexes are particularly suited for best-effort workloads. Nonetheless, they are widely used to support real-time workloads and are integrated in the Linux real-time patch. This paper studies the suitability of the current futex implementation in Linux for certification according to avionics safety standards. Specifically, the paper details the worst-case execution time (WCET) behavior of futexes and the interference patterns that independent applications may encounter when using futexes. Based on our analysis, we identify some weaknesses in the current futex design in Linux, and we propose and evaluate a futex implementation in the context of PikeOS that is suitable for hard real-time and safety-critical systems which target certification. The proposed solution provides a subset of the functionality of Linux, but reduces WCET complexity and addresses the Freedom of Interference-principle for independent applications requested by safety standards. Alexander Züpke, Robert Kaiser |
RTAS | 1 |
| 2015 | AUTOBEST: a united AUTOSAR-OS and ARINC 653 kernelabstractThis paper presents AUTOBEST, a united AUTOSAR-OS and ARINC 653 RTOS kernel that addresses the requirements of both automotive and avionics domains. We show that their domain-specific requirements have a common basis and can be implemented with a small partitioning microkernel-based design on embedded microcontrollers with memory protection (MPU) support. While both, AUTOSAR and ARINC 653, use a unified task model in the kernel, we address their differences in dedicated user space libraries. Based on the kernel abstractions of futexes and lazy priority switching, these libraries provide domain specific synchronization mechanisms. Our results show that thereby it is possible to get the best of both worlds: AUTOBEST combines avionics safety with the resource-efficiency known from automotive systems. Alexander Züpke, Marc Bommert, Daniel Lohmann |
RTAS | 1 |