Andrea Bastoni

dblp:77/9720 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0001-8256-6160ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ETM2: Empowering Traditional Memory Bandwidth Regulation using ETM
Alexander Züpke, Ashutosh Pradhan, Daniele Ottaviano, Andrea Bastoni, Marco Caccamo
RTAS4
2025 Enabling Security on the Edge: A CHERI Compartmentalized Network Stack
abstract
The widespread deployment of embedded systems in critical infrastructures, interconnected edge devices like autonomous drones, and smart industrial systems requires robust security measures. Compromised systems increase the risks of operational failures, data breaches, and-in safety-critical environments-potential physical harm to people. Despite these risks, current security measures are often insufficient to fully address the attack surfaces of embedded devices. CHERI provides strong security from the hardware level by enabling fine-grained compartmentalization and memory protection, which can reduce the attack surface and improve the reliability of such devices. In this work, we explore the potential of CHERI to compartmentalize one of the most critical and targeted components of interconnected systems: their network stack. Our case study examines the tradeoffs of isolating applications, TCP/IP libraries, and network drivers on a CheriBSD system deployed on the Arm Morello platform. Our results suggest that CHERI has the potential to enhance security while maintaining performance in embedded-like environments.
Donato Ferraro, Andrea Bastoni, Alexander Züpke, Andrea Marongiu
DATE2
2025 Multi-Objective Memory Bandwidth Regulation and Cache Partitioning for Multicore Real-Time Systems
Binqi Sun, Zhihang Wei, Andrea Bastoni, Debayan Roy, Mirco Theile, Tomasz Kloda, Rodolfo Pellizzoni, Marco Caccamo
ECRTS3
2025 Arm Dynamiq Shared Unit and Real-Time: An Empirical Evaluation
abstract
The increasing complexity of embedded hardware platforms poses significant challenges for real-time workloads. Architectural features such as Intel RDT, Arm QoS, and Arm MPAM are either unavailable on commercial embedded platforms or designed primarily for server environments optimized for average-case performance and might fail to deliver the expected real-time guarantees. Arm DynamIQ Shared Unit (DSU) includes isolation features-among others, hardware per-way cache partitioning-that can improve the real-time guarantees of complex embedded multicore systems and facilitate real-time analysis. However, the DSU also targets average cases, and its real-time capabilities have not yet been evaluated. This paper presents the first comprehensive analysis of three real-world deployments of the Arm DSU on Rockchip RK3568, Rockchip RK3588, and NVIDIA Orin platforms. We integrate support for the DSU at the operating system and hypervisor level and conduct a large-scale evaluation using both synthetic and real-world benchmarks with varying types and intensities of interference. Our results make extensive use of performance counters and indicate that, although effective, the quality of partitioning and isolation provided by the DSU depends on the type and the intensity of the interfering workloads. In addition, we uncover and analyze in detail the correlation between benchmarks and different types and intensities of interference.
Ashutosh Pradhan, Daniele Ottaviano, Haozheng Huang, Alexander Züpke, Andrea Bastoni, Marco Caccamo
RTAS6
2025 Predictable Memory Bandwidth Regulation for DynamIQ Arm Systems
Ashutosh Pradhan, Daniele Ottaviano, Haozheng Huang, Alexander Züpke, Andrea Bastoni, Marco Caccamo
RTCSA7
2025 Work-in-Progress: Toward Real-Time Cross-ISA Execution on the AMD Embedded+ Architecture
abstract
Emerging embedded platforms increasingly rely on heterogeneous processing units to address diverse performance and energy requirements. The recently introduced AMD Embedded+ architecture reflects this trend by interconnecting via PCIe on the same motherboard one AMD x86 host processor with one Arm AArch64+FPGA complex. This implementation is another step forward towards a more compact heterogeneousISA platform designed with embedded applications in mind. While cross-ISA execution has been explored in the past with a focus on performance, programmability, and energy efficiency, its potential for embedded and predictable real-time workloads remains largely unexplored. In this paper, we start exploring such potential by investigating the real-time capabilities of the first commercial platform based on the AMD Embedded+ architecture: the Sapphire Edge+. We (1) outline key research challenges and real-time use-cases, (2) discuss suitable software architectures for the use-cases and highlight associated trade-offs, and (3) report an initial assessment of the potential of such architectures and use-cases via an experimental evaluation of latency and bandwidth on the real hardware.
Lukas Neef, Daniele Ottaviano, Denis Hoornaert, Alexander Züpke, Marco Caccamo, Andrea Bastoni
RTSS6
2025 Work-in-Progress: A First Practical Look at Arm's MPAM for Real-Time Systems
abstract
Arm's Memory Partitioning and Monitoring (MPAM) extension introduces standardized mechanisms for partitioning cache and memory bandwidth. From a real-time systems perspective, this can aid in improving predictability in heterogeneous MPSoCs. In this paper, we present the first practical evaluation of MPAM on a COTS platform—the Radxa Orion O6 with the CIX CD8180 SoC. We characterize the SoC's MPAM capabilities and experimentally assess cache portion partitioning and proportional stride memory bandwidth partitioning under controlled interference. Our results show that enabling MPAM features can reduce interference, but their behavior often diverges from expectations based on the specification, with anomalous effects observed across workloads and cores. These findings highlight both the promise of predictability from MPAM for real-time systems and the current challenges arising from optionality, heterogeneity, and limited documentation. We conclude that broader evaluation across future MPAM-enabled SoCs, aided by detailed performance counter analysis, is essential to establish MPAM's practical value for real-time practitioners.
Ashutosh Pradhan, Daniele Ottaviano, Alexander Züpke, Andrea Bastoni, Marco Caccamo
RTSS4
2025 Work-in-Progress: Learning to Refine Priority Assignment in Fixed-Priority Real-Time Scheduling
abstract
We address the problem of priority assignment for global fixed-priority scheduling on multicore real-time systems, where identifying a feasible priority ordering is a combinatorial challenge. We propose a learning-based framework that trains a lightweight policy network via reinforcement learning to refine existing priority assignments toward schedulable solutions. Based on the policy network, we propose an inference-time policy refinement mechanism that improves schedulability without additional training. It combines breadth sampling—generating candidate orderings via stochastic perturbations—with depth refinement, which iteratively enhances promising candidates. A continuous reward function based on a schedulability hazard metric enables effective training. Preliminary experiments show that the proposed method performs better than classical heuristics such as Deadline Monotonic and DkC, demonstrating its potential as an effective learning-assisted approach to real-time scheduling.
Binqi Sun, Linghan Fang, Andrea Bastoni, Marco Caccamo
RTSS3
2024 A Containerized Microservice Architecture for a ROS 2 Autonomous Driving Software: An End-to-End Latency Evaluation
abstract
The automotive industry is transitioning from traditional ECU-based systems to software-defined vehicles. A central role of this revolution is played by containers, lightweight virtualization technologies that enable the flexible consolidation of complex software applications on a common hardware platform. Despite their widespread adoption, the impact of containerization on fundamental real-time metrics such as end-to-end latency, communication jitter, as well as memory and CPU utilization has remained virtually unexplored. This paper presents a microservice architecture for a real-world autonomous driving application where containers isolate each service. Our comprehensive evaluation shows the benefits in terms of end-to-end latency of such a solution even over standard bare-Linux deployments. Specifically, in the case of the presented microservice architecture, the mean end-to-end latency can be improved by 5–8%. Also, the maximum latencies were significantly reduced using container deployment.
Tobias Betz, Long Wen 0003, Fengjunjie Pan, Gemb Kaljavesi, Alexander Züpke, Andrea Bastoni, Marco Caccamo, Alois C. Knoll, Johannes Betz
RTCSA6
2024 Mcti: mixed-criticality task-based isolation
abstract
Abstract The ever-increasing demand for high performance in the time-critical, low-power embedded domain drives the adoption of powerful but unpredictable, heterogeneous Systems-on-Chip. On these platforms, the main source of unpredictability—the shared memory subsystem—has been widely studied, and several approaches to mitigate undesired effects have been proposed over the years. Among them, performance-counter-based regulation methods have proved particularly successful. Unfortunately, such regulation methods require precise knowledge of each task’s memory consumption and cannot be extended to isolate mixed-criticality tasks running on the same core as the regulation budget is shared. Moreover, the desirable combination of these methodologies with well-known time-isolation techniques—such as server-based reservations—is still an uncharted territory and lacks a precise characterization of possible benefits and limitations. Recognizing the importance of such consolidation for designing predictable real-time systems, we introduce MCTI (Mixed-Criticality Task-based Isolation) as a first initial step in this direction. MCTI is a hardware/software co-design architecture that aims to improve both CPU and memory isolations among tasks with different criticalities even when they share the same CPU. In order to ascertain the correct behavior and distill the benefits of MCTI, we implemented and tested the proposed prototype architecture on a widely available off-the-shelf platform. The evaluation of our prototype shows that (1) MCTI helps shield critical tasks from concurrent non-critical tasks sharing the same memory budget, with only a limited increase in response time being observed, and (2) critical tasks running under memory stress exhibit an average response time close to that achieved when running without memory stress.
Denis Hoornaert, Golsana Ghaemi, Andrea Bastoni, Renato Mancuso 0001, Marco Caccamo, Giulio Corradi
Real Time Syst.3
2024 MemPol: polling-based microsecond-scale per-core memory bandwidth regulation
abstract
Abstract In today’s multiprocessor systems-on-a-chip, the shared memory subsystem is a known source of temporal interference. The problem causes logically independent cores to affect each other’s performance, leading to pessimistic worst-case execution time analysis. Memory regulation via throttling is one of the most practical techniques to mitigate interference. Traditional regulation schemes rely on a combination of timer and performance counter interrupts to be delivered and processed on the same cores running real-time workload. Unfortunately, to prevent excessive overhead, regulation can only be enforced at a millisecond-scale granularity. In this work, we present a novel regulation mechanism from outside the cores that monitors performance counters for the application core’s activity in main memory at a microsecond scale. The approach is fully transparent to the applications on the cores, and can be implemented using widely available on-chip debug facilities. The presented mechanism also allows more complex composition of metrics to enact load-aware regulation. For instance, it allows redistributing unused bandwidth between cores while keeping the overall memory bandwidth of all cores below a given threshold. We implement our approach on a host of embedded platforms and conduct an in-depth evaluation on the Xilinx Zynq UltraScale+ ZCU102, NXP i.MX8M and NXP S32G2 platforms using the San Diego Vision Benchmark Suite.
Alexander Züpke, Andrea Bastoni, Weifan Chen 0003, Marco Caccamo, Renato Mancuso 0001
Real Time Syst.2
2024 Edge Generation Scheduling for DAG Tasks Using Deep Reinforcement Learning
abstract
Directed acyclic graph (DAG) tasks are currently adopted in the real-time domain to model complex applications from the automotive, avionics, and industrial domains that implement their functionalities through chains of intercommunicating tasks. This paper studies the problem of scheduling real-time DAG tasks by presenting a novel schedulability test based on the concept oftrivial schedulability. Using this schedulability test, we propose a new DAG scheduling framework (edge generation scheduling—EGS) that attempts to minimize the DAG width by iteratively generating edges while guaranteeing the deadline constraint. We study how to efficiently solve the problem of generating edges by developing a deep reinforcement learning algorithm combined with a graph representation neural network to learn an efficient edge generation policy for EGS. We evaluate the effectiveness of the proposed algorithm by comparing it with state-of-the-art DAG scheduling heuristics and an optimal mixed-integer linear programming baseline. Experimental results show that the proposed algorithm outperforms the state-of-the-art by requiring fewer processors to schedule the same DAG tasks.https://github.com/binqi-sun/egs
Binqi Sun, Mirco Theile, Ziyuan Qin 0002, Daniele Bernardini 0002, Debayan Roy, Andrea Bastoni, Marco Caccamo
IEEE Trans. Computers6
2023 MemPol: Policing Core Memory Bandwidth from Outside of the Cores
abstract
In today’s multiprocessor systems-on-a-chip (MP- SoC), the shared memory subsystem is a known source of temporal interference. The problem causes logically independent cores to affect each other’s performance, leading to pessimistic worstcase execution time (WCET) analysis. One of the most practical techniques to mitigate interference is memory regulation via throttling. Traditional regulation schemes rely on a combination of timer and performance counter interrupts to be delivered and processed on the same cores running real-time workload. Unfortunately, to prevent excessive overhead, regulation can only be enforced at a millisecond-scale granularity. In this work, we present a novel regulation mechanism from outside the cores that monitors performance counters for the application core’s activity in main memory at a microsecond scale. The approach is fully transparent to the applications on the cores, and can be implemented using widely available onchip debug facilities. The presented mechanism also allows more complex composition of metrics to enact load-aware regulation. For instance, it allows redistributing unused bandwidth between cores while keeping the overall memory bandwidth of all cores below a given threshold. We implement our approach on a host of embedded platforms and carry out an in-depth evaluation on the Xilinx Zynq UltraScale+ZCUl02 platform using the SD-VBS.
Alexander Züpke, Andrea Bastoni, Weifan Chen 0003, Marco Caccamo, Renato Mancuso 0001
RTAS2
2023 RDMA-Based Deterministic Communication Architecture for Autonomous Driving
abstract
Autonomous driving is a big challenge for next-generation vehicles and requires multiple computationally-intensive deep neural networks (DNNs) to be implemented on distributed automotive platforms. Distributed software-enabling autonomous functionalities-has strict timing requirements, e.g., low and deterministic end-to-end latency. Such timings rely on the communication technologies used in the automotive platform, as much on the computation performance of CPUs, GPUs, TPUs, and FPGAs. Hence, we advocate the use of Remote Direct Memory Access (RDMA) technology-typically used in data centers-in automotive platforms. As shown by our experiments with real hardware, Soft-RoCE (software implementation of RDMA) offers low latency communication because of minimal CPU involvement and reduced memory copies. Simultaneously, we show that the native implementation of RDMA does not support determinism, i.e., there is a high variation in communication delays in the presence of interfering data packets. To mitigate this issue, we propose a multi-layer communication stack comprising a deterministic scheduler on top of the Soft-RoCE layer. Further, we have developed a C++ library that offers easy-to-use communication interfaces for distributed applications while implementing the proposed architecture. Experiments show that our library (i) reduces the end-to-end latency of distributed object detection by nearly 9% while having an implementation overhead of less than 1.5% and (ii) minimizes the effects of other data traffic on the delay in high-priority communication.
Hazem Abaza, Abhinaba Habishyashi, Debayan Roy, Andrea Bastoni, Zain Alabedin Haj Hammadeh, Shiqing Fan, Selma Saidi, Sergey Tverdyshev
RTCSA4
2023 Co-Optimizing Cache Partitioning and Multi-Core Task Scheduling: Exploit Cache Sensitivity or Not?
abstract
Cache partitioning techniques have been successfully adopted to mitigate interference among concurrently executing real-time tasks on multi-core processors. Considering that the execution time of a cache-sensitive task strongly depends on the cache available for it to use, co-optimizing cache partitioning and task allocation improves the system's schedulability. In this paper, we propose a hybrid multi-layer design space exploration technique to solve this multi-resource management problem. We explore the interplay between cache partitioning and schedulability by systematically interleaving three optimization layers, viz., (i) in the outer layer, we perform a breadth-first search combined with proactive pruning for cache partitioning; (ii) in the middle layer, we exploit a first-fit heuristic for allocating tasks to cores; and (iii) in the inner layer, we use the well-known recurrence relation for the schedulability analysis of non-preemptive fixed-priority (NP-FP) tasks in a uniprocessor setting. Although our focus is on NP-FP scheduling, we evaluate the flexibility of our framework in supporting different scheduling policies (NP-EDF, P-EDF) by plugging in appropriate analysis methods in the inner layer. Experiments show that, compared to the state-of-the-art techniques, the proposed framework can improve the real-time schedulability of NP-FP task sets by an average of 15.2% with a maximum improvement of 233.6% (when tasks are highly cache-sensitive) and a minimum of 1.6% (when cache sensitivity is low). For such task sets, we found that clustering similar- period (or mutually compatible) tasks often leads to higher schedulability (on average 7.6 %) than clustering by cache sensitivity. In our evaluation, the framework also achieves good results for preemptive and dynamic-priority scheduling policies.
Binqi Sun, Debayan Roy, Tomasz Kloda, Andrea Bastoni, Rodolfo Pellizzoni, Marco Caccamo
RTSS4
2021 A Real-Time Virtio-Based Framework for Predictable Inter-VM Communication
abstract
Ensuring real-time properties on current heterogeneous multiprocessor systems on a chip is a challenging task. Furthermore, online artificial intelligent applications –which are routinely deployed on such chips– pose increasing pressure on the memory subsystem that becomes a source of unpredictability. Although techniques have been proposed to restore independent access to memory for concurrently executing virtual machines (VM), providing predictable inter-VM communication remains challenging. In this work, we tackle the problem of predictably transferring data between virtual machines and virtualized hardware resources on multiprocessor systems on chips under consideration of memory interference. We design a "broker-based" real-time communication framework for otherwise isolated virtual machines, provide a virtio-based reference implementation on top of the Jailhouse hypervisor, assess its overheads for FreeRTOS virtual machines, and formally analyze its communication flow schedulability under consideration of the implementation overheads. Furthermore, we define a methodology to assess the maximum DRAM memory saturation empirically, evaluate the framework’s performance and compare it with the theoretical schedulability.
Gero Schwäricke, Rohan Tabish, Rodolfo Pellizzoni, Renato Mancuso 0001, Andrea Bastoni, Alexander Züpke, Marco Caccamo
RTSS5
2011 Is Semi-Partitioned Scheduling Practical?
abstract
Semi-partitioned schedulers are -- in theory -- a particularly promising category of multiprocessor real-time scheduling algorithms. Unfortunately, issues pertaining to their implementation have not been investigated in detail, so their practical viability remains unclear. In this paper, the practical merit of three EDF-based semi-partitioned algorithms is assessed via an experimental comparison based on real-time schedulability under consideration of real, measured overheads. The presented results indicate that semi-partitioning is indeed a sound and practical idea. However, several problematic design choices are identified as well. These shortcomings and other implementation concerns are discussed in detail.
Andrea Bastoni, Björn B. Brandenburg, James H. Anderson
ECRTS1
2010 An Empirical Comparison of Global, Partitioned, and Clustered Multiprocessor EDF Schedulers
abstract
As multicore platforms become ever larger, overhead-related factors play a greater role in determining which real-time scheduling algorithms are preferable. In this paper, such factors are investigated through an empirical comparison of global, partitioned, and clustered EDF scheduling algorithms on a 24-core Intel system. On this platform, global EDF proved to be a non-viable choice for hard real time systems, while clusters of size six practically approximated global approaches. For soft real-time systems, clustered EDF scheduling algorithms proved to be particularly effective. This study suggests that future global scheduling research should focus on small-to-medium multicore platforms rather than large platforms.
Andrea Bastoni, Björn B. Brandenburg, James H. Anderson
RTSS1