Giovani Gracioli

dblp:40/3248 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0001-9747-2386ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 DRAM Bank-Aware Memory Allocation for Embedded Real-Time Virtualization
Eduardo Bischoff Grasel, Giovani Gracioli, Daniel Casini
ISORC2
2025 Memguard-RW: Improved Real-Time Memory Bandwidth Regulation Within a Hypervisor
abstract
Memory bandwidth is a critical factor in the performance of DRAM-based computing architectures, particularly in memory-intensive computations. Modern multi-core processors share critical resources, such as main memory and cache, which impact the predictability of real-time systems due to resource contention. Techniques like memory access regulation, cache partitioning, and static hypervisors aim to mitigate this contention. This paper presents an improvement of a memory control mechanism based on MemGuard, named MemGuard-RW, designed and implemented within a hypervisor. MemGuard-RW uses two performance counters for measuring the memory accesses, one for memory readings and another one for writings, thus decreasing the pessimism on the original memory access budget from MemGuard. Additionally, we extend the MemGuard schedulability analysis considering the two-counter approach. To evaluate the effectiveness of the implementation and analysis, we deployed FreeRTOS as a guest alongside three stress-generating guests, measuring the interference experienced by the FreeRTOS instance and comparing the analysis with two counters with the original analysis with one counter, using modern benchmarks. The results demonstrate that the proposed mechanism successfully regulates memory accesses, showing its potential for enhancing the predictability and performance of real-time systems in multi-core environments. Our proposed analysis with two counters reduces the upper bound of around 20 % for tasks having medium and high memory usage.
Everaldo P. Gomes, Adrien Jakubiak, Giovani Gracioli, Tomasz Kloda
ISORC3
2024 Minimizing cache usage with fixed-priority and earliest deadline first scheduling
abstract
Abstract Cache partitioning is a technique to reduce interference among tasks running on the processors with shared caches. To make this technique effective, cache segments should be allocated to tasks that will benefit the most from having their data and instructions stored in the cache. The requests for cached data and instructions can be retrieved faster from the cache memory instead of fetching them from the main memory, thereby reducing overall execution time. The existing partitioning schemes for real-time systems divide the available cache among the tasks to guarantee their schedulability as the sole and primary optimization criterion. However, it is also preferable, particularly in systems with power constraints or mixed criticalities where low- and high-criticality workloads are executing alongside, to reduce the total cache usage for real-time tasks. Cache minimization as part of design space exploration can also help in achieving optimal system performance and resource utilization in embedded systems. In this paper, we develop optimization algorithms for cache partitioning that, besides ensuring schedulability, also minimize cache usage. We consider both preemptive and non-preemptive scheduling policies on single-processor systems with fixed- and dynamic-priority scheduling algorithms ( Rate Monotonic ( RM ) and Earliest Deadline First ( EDF ), respectively). For preemptive scheduling, we formulate the problem as an integer quadratically constrained program and propose an efficient heuristic achieving near-optimal solutions. For non-preemptive scheduling, we combine linear and binary search techniques with different fixed-priority schedulability tests and Quick Processor-demand Analysis (QPA) for EDF. Our experiments based on synthetic task sets with parameters from real-world embedded applications show that the proposed heuristic: (i) achieves an average optimality gap of 0.79% within 0.1× run time of a mathematical programming solver and (ii) reduces average cache usage by 39.15% compared to existing cache partitioning approaches. Besides, we find that for large task sets with high utilization, non-preemptive scheduling can use less cache than preemptive to guarantee schedulability.
Binqi Sun, Tomasz Kloda, Sergio Arribas García, Giovani Gracioli, Marco Caccamo
Real Time Syst.4
2023 X-Stream: Accelerating streaming segments on MPSoCs for real-time applications
Rohan Tabish, Rodolfo Pellizzoni, Renato Mancuso 0001, Giovani Gracioli, Reza Mirosanlou, Marco Caccamo
J. Syst. Archit.4
2023 Lazy Load Scheduling for Mixed-criticality Applications in Heterogeneous MPSoCs
abstract
Newly emerging multiprocessor system-on-a-chip (MPSoC) platforms provide hard processing cores with programmable logic (PL) for high-performance computing applications. In this article, we take a deep look into these commercially available heterogeneous platforms and show how to design mixed-criticality applications such that different processing components can be isolated to avoid contention on the shared resources such as last-level cache and main memory. Our approach involves software/hardware co-design to achieve isolation between the different criticality domains. At the hardware level, we use a scratchpad memory (SPM) with dedicated interfaces inside the PL to avoid conflicts in the main memory. At the software level, we employ a hypervisor to support cache-coloring such that conflicts at the shared L2 cache can be avoided. In order to move the tasks in/out of the SPM memory, we rely on a DMA engine and propose a new CPU-DMA co-scheduling policy, called Lazy Load , for which we also derive the response time analysis. The results of a case study on image processing demonstrate that the contention on the shared memory subsystem can be avoided when running with our proposed architecture. Moreover, comprehensive schedulability evaluations show that the newly proposed Lazy Load policy outperforms the existing CPU-DMA scheduling approaches and is effective in mitigating the main memory interference in our proposed architecture.
Tomasz Kloda, Giovani Gracioli, Rohan Tabish, Reza Mirosanlou, Renato Mancuso 0001, Rodolfo Pellizzoni, Marco Caccamo
ACM Trans. Embed. Comput. Syst.2
2021 Flexible Cache Partitioning for Multi-Mode Real-Time Systems
abstract
Cache partitioning is a well-studied technique that mitigates the inter-processor cache interference in multiprocessor systems. The resulting optimization problem involves allocating portions of the cache to individual processors. In multi-mode applications (e.g., flight control system that runs in take-off, cruise, or landing mode), the cache memory requirement can change over time, making runtime cache repartitioning necessary. This paper presents a cache partition allocation framework enabling flexible cache partitioning for multi-mode real-time systems. The main objective is to guarantee timing predictability in the steady states and during mode changes. We evaluate the effectiveness of our approach for multiple embedded benchmarks with different ranges of cache size sensitivity. The results show increased schedulability compared to static partitioning approaches.
Ohchul Kwon, Gero Schwäricke, Tomasz Kloda, Denis Hoornaert, Giovani Gracioli, Marco Caccamo
DATE5
2021 Palisade: A framework for anomaly detection in embedded systems
Sean Kauffman, Murray Dunne, Giovani Gracioli, Waleed Khan 0003, Nirmal Benann, Sebastian Fischmeister
J. Syst. Archit.3
2020 Fixed-Priority Memory-Centric Scheduler for COTS-Based Multiprocessors
abstract
Memory-centric scheduling attempts to guarantee temporal predictability on commercial-off-the-shelf (COTS) multiprocessor systems to exploit their high performance for real-time applications. Several solutions proposed in the real-time literature have hardware requirements that are not easily satisfied by modern COTS platforms, like hardware support for strict memory partitioning or the presence of scratchpads. However, even without said hardware support, it is possible to design an efficient memory-centric scheduler. In this article, we design, implement, and analyze a memory-centric scheduler for deterministic memory management on COTS multiprocessor platforms without any hardware support. Our approach uses fixed-priority scheduling and proposes a global "memory preemption" scheme to boost real-time schedulability. The proposed scheduling protocol is implemented in the Jailhouse hypervisor and Erika real-time kernel. Measurements of the scheduler overhead demonstrate the applicability of the proposed approach, and schedulability experiments show a 20% gain in terms of schedulability when compared to contention-based and static fair-share approaches.
Gero Schwäricke, Tomasz Kloda, Giovani Gracioli, Marko Bertogna, Marco Caccamo
ECRTS3
2019 Designing Mixed Criticality Applications on Modern Heterogeneous MPSoC Platforms
abstract
Multiprocessor Systems-on-Chip (MPSoC) integrating hard processing cores with programmable logic (PL) are becoming increasingly common. While these platforms have been originally designed for high performance computing applications, their rich feature set can be exploited to efficiently implement mixed criticality domains serving both critical hard real-time tasks, as well as soft real-time tasks. In this paper, we take a deep look at commercially available heterogeneous MPSoCs that incorporate PL and a multicore processor. We show how one can tailor these processors to support a mixed criticality system, where cores are strictly isolated to avoid contention on shared resources such as Last-Level Cache (LLC) and main memory. In order to avoid conflicts in last-level cache, we propose the use of cache coloring, implemented in the Jailhouse hypervisor. In addition, we employ ScratchPad Memory (SPM) inside the PL to support a multi-phase execution model for real-time tasks that avoids conflicts in shared memory. We provide a full-stack, working implementation on a latest-generation MPSoC platform, and show results based on both a set of data intensive tasks, as well as a case study based on an image processing benchmark application.
Giovani Gracioli, Rohan Tabish, Renato Mancuso 0001, Reza Mirosanlou, Rodolfo Pellizzoni, Marco Caccamo
ECRTS1
2019 Segment Streaming for the Three-Phase Execution Model: Design and Implementation
abstract
Scheduling tasks using the three-phase execution model (load-execute-unload) can effectively reduce the contention on shared resources in real-time systems. Due to system and program constraints, a task is generally segmented and executed over multiple intervals. Several works showed that co-scheduling memory (unload-load) and computation phases can improve the system schedulability by hiding the memory transfer time. However, this is limited to segments of different tasks and hence executing segments of the same task back-to-back is not allowed. In this paper, we propose a new streaming model to allow overlapping the memory and execution phases of segments of the same task. This is accomplished by a segmentation framework implemented within an LLVM-based compiler-level tool along with a Real-Time Operating System (RTOS) API to handle load/unload requests. Memory phases are processed by a DMA engine that loads/unloads the task content into ScratchPad Memory (SPM). We provide a schedulability analysis of the proposed model under fixed priority partitioned scheme and an RTOS implementation of the API on a latest-generation Multiprocessor System-on-Chip (MPSoC).
Muhammad Refaat Soliman, Giovani Gracioli, Rohan Tabish, Rodolfo Pellizzoni, Marco Caccamo
RTSS2
2018 Preface for the special issue of the 6th Brazilian Symposium on Computing System Engineering (SBESC 2016)
Giovani Gracioli, Rivalino Matias
Sci. Comput. Program.1
2014 CAP: Color-aware task partitioning for multicore real-time applications
abstract
Modern multicore platforms feature multiple levels of cache memory placed between the processor and main memory to hide the latency of ordinary memory systems. The primary goal of this cache hierarchy is to improve average execution time (at the cost of predictability). The uncontrolled use of the cache hierarchy by real-time tasks may impact the estimation of their worst-case execution times (WCET). Software cache partitioning through page coloring has been considered a promising approach to isolate task workloads and thus improve WCET estimation. However, when real-time tasks share cache partitions due to false or true sharing, the inter-core delay caused by the cache coherence protocol may cause deadline losses. In this paper, we propose a Color-Aware task Partitioning (CAP) algorithm that assigns tasks to cores respecting their usage of cache partitions (i.e., colors). Tasks that share one or more colors are grouped together and the whole group is assigned to the same processor. Thus, it is possible to avoid inter-core interference. We compared the deadline miss ratio of several generated task sets partitioned by the CAP algorithm and by the worst-fit decreasing heuristic. We executed the partitioned task sets in a modern 8-core processor with shared L3-cache using a real-time operating system. Our results indicate that a color-aware task partitioning algorithm can avoid deadline misses in a multicore processor with shared cache.
Giovani Gracioli, Antônio Augusto Fröhlich
ETFA1
2013 An experimental evaluation of the cache partitioning impact on multicore real-time schedulers
abstract
Shared cache partitioning is a well-known technique used in multicore real-time systems to isolate task workloads and improve system predictability. Presently, the state-of-the-art studies that evaluate shared cache partitioning on multicore processors lack two key issues. First, the cache partitioning mechanism is typically implemented either in a simulation environment or in a general-purpose OS, and so the impact of kernel activities, such as interrupt handlers and context switching, on the task partitions tend to be overlooked. Second, the evaluation is typically restricted to either a global or partitioned scheduler, thereby by falling to compare the performance of cache partitioning when tasks are scheduled by different schedulers. In this work, we design and implement a shared cache partitioning mechanism in a multicore component-based RTOS capable of assigning partitions to internal OS data structures, including task and system stacks and interrupt handlers data. We evaluate our shared cache partitioning mechanism running task sets under global (G-EDF) and partitioned (P-EDF) multicore real-time scheduling algorithms. Our results indicate that a lightweight RTOS does not impact real-time tasks, and shared cache partitioning has different behavior depending on the scheduler and the task's working set size.
Giovani Gracioli, Antônio Augusto Fröhlich
RTCSA1
2013 Implementation and evaluation of global and partitioned scheduling in a real-time OS
Giovani Gracioli, Antônio Augusto Fröhlich, Rodolfo Pellizzoni, Sebastian Fischmeister
Real Time Syst.1
2012 An operating system runtime reprogramming infrastructure for WSN
abstract
WSNs, despite of its limited resources, are expected to operate without human interventions for a long period of time. Nevertheless, the environment might develop unpredicted characteristics or some network functionality might need some changes. Thus, it is necessary a mechanism which allows software reprogramming of the network nodes after its deployment. This paper presents the integration of a data dissemination protocol and ELUS, an OS support environment. The dissemination protocol is responsible for spreading the data across the network, while ELUS isolates the system components in memory position independent units, allowing their updating at execution time. We have evaluated our infrastructure, using real sensor nodes, in terms of memory consumption, dissemination, and reprogramming time.
Rodrigo Vieira Steiner, Giovani Gracioli, Rita de Cassia Cazu Soldi, Antônio Augusto Fröhlich
ISCC2
2012 Tracing and recording interrupts in embedded software
Giovani Gracioli, Sebastian Fischmeister
J. Syst. Archit.1
2009 Tracing interrupts in embedded software
abstract
During the system development, developers often must correct wrong behavior in the software---an activity colloquially called program debugging. Debugging is a complex activity, especially in real-time embedded systems because such systems interact with the physical world and make heavy use of interrupts for timing and driving I/O devices.
Giovani Gracioli, Sebastian Fischmeister
LCTES1