EDBT 2026 Demo / reviewers in the wild / expert
Marko Bertogna
dblp:31/4402
· DBLP profile ↗
78ranked-venue papers
13as first author
14since 2021 · last 2026
0000-0003-2115-4853ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Vision Foundation Models with Drop-in Depthwise Convolution
Carmelo Scribano, Mohammad Mahdi, Nedyalko Prisadnikov, Yuqian Fu, Giorgia Franchini, Danda Pani Paudel, Marko Bertogna, Luc Van Gool |
ICPR (13) | 7 |
| 2025 | From Physical to Digital: Exploring Digital Twins within the Modena Automotive Smart AreaabstractThe Modena Automotive Smart Area (MASA) is a cutting-edge testing environment featuring a variety of dynamic physical assets, including smart cameras, roadside units, and connected vehicles. These assets support numerous digital applications, ranging from real-time safety systems to mobility intelligence and 3D visualization of the MASA area. However, the complexity of the physical environment and the diverse needs of these digital applications necessitate a decoupling strategy to ensure efficient operation. This paper presents the design of the MASA Digital Twin, detailing its hierarchical structure, the associated design challenges, and the technological approaches used in its implementation. The MASA Digital Twin serves as a crucial tool for managing the interplay between physical and digital elements, enabling a more structured and adaptable approach to connected mobility and smart city applications. Marco Picone 0001, Antonello Barbone, Riccardo Morandi, Enrico Rossini, Alessio Masola, Marcello Pietri, Roberto Cavicchioli, Carlo Augusto Grazia, Marco Mamei, Marko Bertogna |
CCNC | 10 |
| 2025 | BETTY Dataset: A Multi-Modal Dataset for Full-Stack AutonomyabstractWe present the BETTY dataset, a large-scale, multi-modal dataset collected on several autonomous racing vehicles, targeting supervised and self-supervised state estimation, dynamics modeling, motion forecasting, perception, and more. Existing large-scale datasets, especially autonomous vehicle datasets, focus primarily on supervised perception, planning, and motion forecasting tasks. Our work enables multi-modal, data-driven methods by including all sensor inputs and the outputs from the software stack, along with semantic metadata and ground truth information. The dataset encompasses 4 years of data, currently comprising over 13 hours and 32 TB, collected on autonomous racing vehicle platforms. This data spans 6 diverse racing environments, including high-speed oval courses, for single and multi-agent algorithm evaluation in feature-sparse scenarios, as well as high-speed road courses with high longitudinal and lateral accelerations and tight, GPSdenied environments. It captures highly dynamic states, such as$63 \mathrm{m} / \mathrm{s}$crashes, loss of tire traction, and operation at the limit of stability. By offering a large breadth of cross-modal and dynamic data, the BETTY dataset enables the training and testing of full autonomy stack pipelines, pushing the performance of all algorithms to the limits. The current dataset is available at https://pitt-mit-iac.github.io/betty-dataset/. Micah Nye, Ayoub Raji, Andrew Saba, Eidan Erlich, Robert Exley, Aragya Goyal, Alexander Matros, Ritesh Misra, Matthew Sivaprakasam, Marko Bertogna, Deva Ramanan, Sebastian A. Scherer |
ICRA | 10 |
| 2025 | Modular Decision-Making and Drivable Areas for Multi-Agent Autonomous RacingabstractThis paper presents an interaction-aware, modular framework for local trajectory planning in autonomous driving, particularly suited for multi-agent racing scenarios. Our framework first identifies viable drivable areas (tunnels), taking into account predictions of other agents’ behaviors, and subsequently utilizes a high-level decision-making module to select the optimal corridor considering both static and moving vehicles. This decision-making module also strategically determines when to follow an opponent or initiate an overtaking maneuver, while ensuring compliance with racing regulations. A Model Predictive Control (MPC) module is then employed to compute an optimal, collision-free trajectory within the chosen corridor. The proposed modular architecture simplifies the computational complexity typically associated with MPC optimization and facilitates independent component testing. Simulations and real-world tests on various racing tracks demonstrate the efficacy of our approach, even in highly dynamic interactive scenarios with multiple simultaneous opponents; videos of these and additional experiments are available at https://atoschi.github.io/tunnels-framework/. Alessandro Toschi, Francesco Prignoli, Marko Bertogna |
IROS | 3 |
| 2025 | Timing guarantees for inference of AI models in embedded systems
Seunghoon Lee 0002, Woosung Kang 0002, Marko Bertogna, Hoon Sung Chwa, Jinkyu Lee 0001 |
Real Time Syst. | 3 |
| 2024 | Integrating IoT and Simulation for Efficient Livestock Waste Spread in the Po ValleyabstractThe livestock industry in the Po Valley is economically significant but generates challenging by-products, notably pig manure. Disposal typically involves spreading manure on agricultural land, requiring daily transportation by hundreds of tankers and leading to soil, water, and air pollution. This method incurs high costs and poses monitoring challenges for government agencies. This research develops a simulated environment that mimics waste dissemination activities in the Po Valley to support the development of an Internet of Things (IoT) architecture for cost-effective environmental monitoring. A custom-built emulator was designed to map driver behaviors, identify limitations, and foresee challenges before real-world implementation, enabling scalable completely renovated and distributed monitoring tools. Giovanni Triboli, Marco Picone 0001, Marko Bertogna |
DS-RT | 3 |
| 2024 | Intelligent Livestock Waste Sensing & Management: Architecture and Challenges in the Po ValleyabstractIntensive livestock production, notably in the Po Valley (Italy), poses significant environmental challenges due to the high concentration of nitrate-rich effluents. These effluents, generated by millions of animals, are directly disposed of onto fertile soil as a common irrigation technique. Unfortunately, this leads to widespread contamination, compromising essential resources like water, soil, and air. This paper focuses on analyzing and modeling the main phases associated with livestock waste management and designing and prototyping an Internet of Things (IoT) enabled architecture for mission management and facilitating real-time monitoring of operations. The possibility to build a comprehensive representation of involved processes and data allows for filling a crucial gap for scientific, commercial, and policy purposes by enabling precise reporting to control agencies and supporting sustainable interventions to mitigate environmental impact. Giovanni Triboli, Marco Picone 0001, Marko Bertogna |
ISCC | 3 |
| 2024 | Guess the Drift with LOP-UKF: LiDAR Odometry and Pacejka Model for Real-Time Racecar Sideslip EstimationabstractThe sideslip angle, crucial for vehicle safety and stability, is determined using both longitudinal and lateral velocities. However, measuring the lateral component often necessitates costly sensors, leading to its common estimation, a topic thoroughly explored in existing literature. This paper introduces LOP-UKF, a novel method for estimating vehicle lateral velocity by integrating Lidar Odometry with the Pacejka tire model predictions, resulting in a robust estimation via an Unscendent Kalman Filter (UKF). This combination represents a distinct alternative to more traditional methodologies, resulting in a reliable solution also in edge cases. We present experimental results obtained using the Dallara AV-21 across diverse circuits and track conditions, demonstrating the effectiveness of our method. Alessandro Toschi, Nicola Musiu, Francesco Gatti, Ayoub Raji, Francesco Amerotti, Micaela Verucchi, Marko Bertogna |
IV | 7 |
| 2024 | A Simulation Benchmark for Autonomous Racing with Large-Scale Human DataabstractDespite the availability of international prize-money competitions, scaled vehicles, and simulation environments, research on autonomous racing and the control of sports cars operating close to the limit of handling has been limited by the high costs of vehicle acquisition and management, as well as the limited physics accuracy of open-source simulators. In this paper, we propose a racing simulation platform based on the simulator Assetto Corsa to test, validate, and benchmark autonomous driving algorithms, including reinforcement learning (RL) and classical Model Predictive Control (MPC), in realistic and challenging scenarios. Our contributions include the development of this simulation platform, several state-of-the-art algorithms tailored to the racing environment, and a comprehensive dataset collected from human drivers. Additionally, we evaluate algorithms in the offline RL setting. All the necessary code (including environment and benchmarks), working examples, and datasets are publicly released and can be found at: https://github.com/dasGringuen/assettocorsagym. Adrian Remonda, Nicklas Hansen 0001, Ayoub Raji, Nicola Musiu, Marko Bertogna, Eduardo E. Veas, Xiaolong Wang 0004 |
NeurIPS | 5 |
| 2024 | Human-Machine Interfaces in Safety-Related Cooperative Driving Automation SystemsabstractThe main objective of connected vehicles is to signif-icantly improve safety for travel participants and nearby people. Cooperative Driving Automation denotes the set of autonomous applications requiring cooperation among connected vehicles, including those with driving automation features engaged. As the operational purpose of the cooperation depends on the application class and the level of driving automation, it follows that the communication strategies and, therefore, the design principles of the human-machine interfaces of participating vehicles must change as well. However, the literature is missing a standardized and coherent strategy for the design and implementation of these interfaces, which may have undesirable impacts on the risk reduction profile of cooperative applications. This paper wants to be a first step in the definition of suitable design principles for such interfaces. We contribute an overview of related solutions providing a selection of relevant works, and we synthesize a set of fundamental human-machine interface design principles for safety - related cooperative applications, hypothesizing an inverse relationship between the level of driving automation and the need for active approaches to maximise the safety of connected vehicles and nearby road users. Andrea Castellano, Martin Klapez, Carlo Augusto Grazia, Elisa Landini, Maurizio Casoni, Marko Bertogna |
WiMob | 6 |
| 2023 | Binary Classification of Agricultural Crops Using Sentinel Satellite Data and Machine Learning TechniquesabstractThe automated process of determining the crop type carried on plots of land, leveraging data provided by earth observation satellites, represents a highly valuable ability that can serve as a foundation for subsequent analyses or as input for calibrating models, such as Decision Support Systems.This paper presents a study on the task of crop classification starting from indices derived from imagery data provided by ESA Satellites Sentinel 1 and 2. We create a valuable tool to verify farmers' claims, especially in relation to state subsidies for specific crops of interest.To this purpose, we focus on perfecting a binary classification for each of five crops of interest (Tomatoes, Soy, Sugar Beet, Rice, and Wheat), aimed to accurately discern the target crop against any other possible crop.The paper investigates various preprocessing techniques to create a dataset suitable for traditional machine learning methods, which presumes that each land plot to classify is represented by a fixed set of features.To deal with inevitable missing observations caused by clouds or other environmental factors, we investigate different imputation strategies (linear interpolation and constant value filling).Complementary, we study the impact of imbalanced classification labels and evaluate the effectiveness of standard balancing techniques.The findings offer practical implications for monitoring and optimizing agricultural practices in the context of precision farming and sustainable agriculture. Paolo Bertellini, Gianluca D'Addese, Giorgia Franchini, Simone Parisi, Carmelo Scribano, Daniele Zanirato, Marko Bertogna |
FedCSIS | 7 |
| 2023 | A survey on real-time DAG scheduling, revisiting the Global-Partitioned Infinity War
Micaela Verucchi, Ignacio Sanudo Olmedo, Marko Bertogna |
Real Time Syst. | 3 |
| 2021 | The HPC-DAG Task Model for Heterogeneous Real-Time SystemsabstractRecent commercial hardware platforms for embedded real-time systems feature heterogeneous processing units and computing accelerators on the same System-on-Chip. When designing complex real-time applications for such architectures, the designer is exposed to a number of difficult choices, like deciding on which compute engine to execute a certain task, or what degree of parallelism to adopt for a given function. To help the designer exploring the wide space of design choices and tune the scheduling parameters, we propose a novel real-time application model, called HPC-DAG (Heterogeneous Parallel Condition Directed Acyclic Graph Model), specifically conceived for heterogeneous platforms. An HPC-DAG allows the system designer to specify alternative implementations of a software component for different processing engines, as well as conditional branches to modelif-then-elsestatements. We also propose a schedulability analysis for the HPC-DAG model and a set of heuristic allocation algorithms aimed at improving schedulability for latency sensitive applications. Our analysis takes into account the cost of preempting a task, which can be non-negligible on certain processors. We show the use of our approach on a realistic case study, and we demonstrate its effectiveness by comparing it with state-of-the-art algorithms previously proposed in literature. Houssam-Eddine Zahaf, Nicola Capodieci, Roberto Cavicchioli, Giuseppe Lipari, Marko Bertogna |
IEEE Trans. Computers | 5 |
| 2021 | The Predictable Execution Model in Practice: Compiling Real Applications for COTS HardwareabstractAdoption of multi- and many-core processors in real-time systems has so far been slowed down, if not totally barred, due do the difficulty in providing analytical real-time guarantees on worst-case execution times. The Predictable Execution Model (PREM) has been proposed to solve this problem, but its practical support requires significant code refactoring, a task better suited for a compilation tool chain than human programmers. Implementing a PREM compiler presents significant challenges to conform to PREM requirements, such as guaranteed upper bounds on memory footprint and the generation of efficient schedulable non-preemptive regions. This article presents a comprehensive description on how a PREM compiler can be implemented, based on several years of experience from the community. We provide accumulated insights on how to best balance conformance to real-time requirements and performance and present novel techniques that extend the applicability from simple benchmark suites to real-world applications. We show that code transformed by the PREM compiler enables timing predictable execution on modern commercial off-the-shelf hardware, providing novel insights on how PREM can protect 99.4% of memory accesses on random replacement policy caches at only 16% performance loss on benchmarks from the PolyBench benchmark suite. Finally, we show that the requirements imposed on the programming model are well-aligned with current coding guidelines for timing critical software, promoting easy adoption. Björn Forsberg, Marco Solieri, Marko Bertogna, Luca Benini, Andrea Marongiu |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | Fixed-Priority Memory-Centric Scheduler for COTS-Based MultiprocessorsabstractMemory-centric scheduling attempts to guarantee temporal predictability on commercial-off-the-shelf (COTS) multiprocessor systems to exploit their high performance for real-time applications. Several solutions proposed in the real-time literature have hardware requirements that are not easily satisfied by modern COTS platforms, like hardware support for strict memory partitioning or the presence of scratchpads. However, even without said hardware support, it is possible to design an efficient memory-centric scheduler. In this article, we design, implement, and analyze a memory-centric scheduler for deterministic memory management on COTS multiprocessor platforms without any hardware support. Our approach uses fixed-priority scheduling and proposes a global "memory preemption" scheme to boost real-time schedulability. The proposed scheduling protocol is implemented in the Jailhouse hypervisor and Erika real-time kernel. Measurements of the scheduler overhead demonstrate the applicability of the proposed approach, and schedulability experiments show a 20% gain in terms of schedulability when compared to contention-based and static fair-share approaches. Gero Schwäricke, Tomasz Kloda, Giovani Gracioli, Marko Bertogna, Marco Caccamo |
ECRTS | 4 |
| 2020 | Real-Time Virtualization For Industrial AutomationabstractThe industry has recently shown a growing interest in running non-critical activities (e.g., design tools) together with real-time control tasks on open-source COTS platforms.In this paper, we present a modular platform for industrial automation based on open-source components. Thanks to the isolation provided by a hypervisor, the same platform can run both the real-time control code and the design tools, reducing the overall hardware costs.Besides illustrating the software architecture and describing the various open-source components, we illustrate extensions needed for convenient applicability, and we address the problem of interference on hardware resources shared by multiple operating systems, which undermines the performance of the real-time control. A set of experimental results show the effectiveness and the performance of the proposed solution. Claudio Scordino, Ida Maria Savino, Luca Cuomo, Luca Miccio, Andrea Tagliavini, Marko Bertogna, Marco Solieri |
ETFA | 6 |
| 2020 | A Systematic Assessment of Embedded Neural Networks for Object DetectionabstractObject detection is arguably one of the most important and complex tasks to enable the advent of next-generation autonomous systems. Recent advancements in deep learning techniques allowed a significant improvement in detection accuracy and latency of modern neural networks, allowing their adoption in automotive, avionics and industrial embedded systems, where performances are required to meet size, weight and power constraints.Multiple benchmarks and surveys exist to compare state-of-the-art detection networks, profiling important metrics, like precision, latency and power efficiency on Commercial-off-the-Shelf (COTS) embedded platforms. However, we observed a fundamental lack of fairness in the existing comparisons, with a number of implicit assumptions that may significantly bias the metrics of interest. This includes using heterogeneous settings for the input size, training dataset, threshold confidences, and, most importantly, platform-specific optimizations, that are especially important when assessing latency and energy-related values. The lack of uniform comparisons is mainly due to the significant effort required to re-implement network models, whenever openly available, on the specific platforms, to properly configure the available acceleration engines for optimizing performance, and to re-train the model using a homogeneous dataset.This paper aims at filling this gap, providing a comprehensive and fair comparison of the best-in-class Convolution Neural Networks (CNNs) for real-time embedded systems, detailing the effort made to achieve an unbiased characterization on cutting-edge system-on-chips. Multi-dimensional trade-offs are explored for achieving a proper configuration of the available programmable accelerators for neural inference, adopting the best available software libraries. To stimulate the adoption of fair benchmarking assessments, the framework is released to the public in an open source repository. Micaela Verucchi, Gianluca Brilli, Davide Sapienza, Mattia Verasani, Marco Arena, Francesco Gatti, Alessandro Capotondi, Roberto Cavicchioli, Marko Bertogna, Marco Solieri |
ETFA | 9 |
| 2020 | Evaluating Controlled Memory Request Injection to Counter PREM Memory Underutilization
Roberto Cavicchioli, Nicola Capodieci, Marco Solieri, Marko Bertogna, Paolo Valente, Andrea Marongiu |
JSSPP | 4 |
| 2020 | Dissecting the CUDA scheduling hierarchy: a Performance and Predictability PerspectiveabstractOver the last few years, the ever-increasing use of Graphic Processing Units (GPUs) in safety-related domains has opened up many research problems in the real-time community. The closed and proprietary nature of the scheduling mechanisms deployed in NVIDIA GPUs, for instance, represents a major obstacle in deriving a proper schedulability analysis for latency-sensitive applications. Existing literature addresses these issues by either (i) providing simplified models for heterogeneous CPUGPU systems and their associated scheduling policies, or (ii) providing insights about these arbitration mechanisms obtained through reverse engineering. In this paper, we take one step further by correcting and consolidating previously published assumptions about the hierarchical scheduling policies of NVIDIA GPUs and their proprietary CUDA application programming interface. We also discuss how such mechanisms evolved with recently released GPU micro-architectures, and how such changes influence the scheduling models to be exploited by real-time system engineers. Ignacio Sanudo Olmedo, Nicola Capodieci, Jorge Martinez 0003, Andrea Marongiu, Marko Bertogna |
RTAS | 5 |
| 2020 | Latency-Aware Generation of Single-Rate DAGs from Multi-Rate Task SetsabstractModern automotive and avionics embedded systems integrate several functionalities that are subject to complex timing requirements. A typical application in these fields is composed of sensing, computation, and actuation. The ever increasing complexity of heterogeneous sensors implies the adoption of multi-rate task models scheduled onto parallel platforms. Aspects like freshness of data or first reaction to an event are crucial for the performance of the system. The Directed Acyclic Graph (DAG) is a suitable model to express the complexity and the parallelism of these tasks. However, deriving age and reaction timing bounds is not trivial when DAG tasks have multiple rates. In this paper, a method is proposed to convert a multi-rate DAG task-set with timing constraints into a single-rate DAG that optimizes schedulability, age and reaction latency, by inserting suitable synchronization constructs. An experimental evaluation is presented for an autonomous driving benchmark, validating the proposed approach against state-of-the-art solutions. Micaela Verucchi, Mirco Theile, Marco Caccamo, Marko Bertogna |
RTAS | 4 |
| 2020 | Contending memory in heterogeneous SoCs: Evolution in NVIDIA Tegra embedded platformsabstractModern embedded platforms are known to be constrained by size, weight and power (SWaP) requirements. In such contexts, achieving the desired performance-per-watt target calls for increasing the number of processors rather than ramping up their voltage and frequency. Hence, generation after generation, modern heterogeneous System on Chips (SoC) present a higher number of cores within their CPU complexes as well as a wider variety of accelerators that leverages massively parallel compute architectures. Previous literature demonstrated that while increasing parallelism is theoretically optimal for improving on average performance, shared memory hierarchies (i.e. caches and system DRAM) act as a bottleneck by exposing the platform processors to severe contention on memory accesses, hence dramatically impacting performance and timing predictability. In this work we characterize how subsequent generations of embedded platforms from the NVIDIA Tegra family balanced the increasing parallelism of each platform's processors with the consequent higher potential on memory interference. We also present an open-source software for generating test scenarios aimed at measuring memory contention in highly heterogeneous SoCs. Nicola Capodieci, Roberto Cavicchioli, Ignacio Sanudo Olmedo, Marco Solieri, Marko Bertogna |
RTCSA | 5 |
| 2020 | Exact response time analysis of fixed priority systems based on sporadic servers
Jorge Martinez 0003, Dakshina Dasari, Arne Hamann 0001, Ignacio Sanudo Olmedo, Marko Bertogna |
J. Syst. Archit. | 5 |
| 2020 | End-to-end latency characterization of task communication models for automotive systems
Jorge Martinez 0003, Ignacio Sanudo Olmedo, Marko Bertogna |
Real Time Syst. | 3 |
| 2019 | System Performance Modelling of Heterogeneous HW Platforms: An Automated Driving Case StudyabstractThe push towards automated and connected driving functionalities mandates the use of heterogeneous HW platforms in order to provide the required computational resources. For these platforms, the established methods for performance modelling in industry are no longer effective. In this paper, we propose an initial modelling concept for heterogeneous platforms which can then be fed into appropriate tools to derive effective performance predictions. The approach is demonstrated for a prototypical automated driving application on the Nvidia Tegra X2 platform. Falk Wurst, Dakshina Dasari, Arne Hamann 0001, Dirk Ziegenbein, Ignacio Sanudo Olmedo, Nicola Capodieci, Marko Bertogna, Paolo Burgio |
DSD | 7 |
| 2019 | Novel Methodologies for Predictable CPU-To-GPU Command OffloadingabstractContainerisation is becoming a cornerstone of modern distributed systems, thanks to their lightweight virtualisation, high portability, and seamless integration with orchestration tools such as Kubernetes. The usage of containers has also gained traction in real-time cyber-physical systems, such as software-defined vehicles, which are characterised by strict timing requirements to ensure safety and performance. Nevertheless, ensuring real-time execution of co-located containers is challenging because of mutual interference due to the sharing of the same processing hardware. Existing parallel computing frameworks such as Ray and its Kubernetes-enabled variant, KubeRay, excel in distributed computation but lack support for scheduling policies that allow guaranteeing real-time timing constraints and CPU resource isolation between containers, such as the SCHED_DEADLINE policy of Linux. To fill this gap, this paper extends Ray to support real-time containers that leverage SCHED_DEADLINE. To this end, we propose KubeDeadline, a novel, modular Kubernetes extension to support SCHED_DEADLINE. We evaluate our approach through extensive experiments, using synthetic workloads and a case study based on the MobileNet and EfficientNet deep neural networks. Our evaluation shows that KubeDeadline ensures deadline compliance in all synthetic workloads, adds minimal deployment overhead (in the order of milliseconds), and achieves lower worst-case response times, up to 4 times lower, than vanilla Kubernetes under background interference. Roberto Cavicchioli, Nicola Capodieci, Marco Solieri, Marko Bertogna |
ECRTS | 4 |
| 2019 | Deterministic Memory Hierarchy and Virtualization for Modern Multi-Core Embedded SystemsabstractOne of the main predictability bottlenecks of modern multi-core embedded systems is contention for access to shared memory resources. Partitioning and software-driven allocation of memory resources is an effective strategy to mitigate contention in the memory hierarchy. Unfortunately, however, many of the strategies adopted so far can have unforeseen side-effects when practically implemented latest-generation, high-performance embedded platforms. Predictability is further jeopardized by cache eviction policies based on random replacement, targeting average performance instead of timing determinism. In this paper, we present a framework of software-based techniques to restore memory access determinism in high-performance embedded systems. Our approach leverages OS-transparent and DMA-friendly cache coloring, in combination with an invalidation-driven allocation (IDA) technique. The proposed method allows protecting important cache blocks from (i) external eviction by tasks concurrently executing on different cores, and (ii) internal eviction by tasks running on the same core. A working implementation obtained by extending the Jailhouse partitioning hypervisor is presented and evaluated with a combination of synthetic and real benchmarks. Tomasz Kloda, Marco Solieri, Renato Mancuso 0001, Nicola Capodieci, Paolo Valente, Marko Bertogna |
RTAS | 6 |
| 2019 | The Parallel Multi-Mode Digraph Task Model for Energy-Aware Real-Time Heterogeneous Multi-Core SystemsabstractMany task models have been proposed to express and analyze the behavior of real-time applications at different levels of precision. Most of them target sequential applications with no support for parallelism. The digraph task model is one of the most general ones, as it allows modeling arbitrary directed graphs (digraphs) for sequential job releases. In this paper, we extend the digraph task model to support intra-task parallelism. For the proposed parallel multi-mode digraph model, we derive sufficient schedulability tests and a dichotomic search to improve the test pessimism for a set of n tasks onto a heterogeneous single-ISA multi-core platform. To reduce the computational complexity of the schedulability test, we also propose heuristics for (i) partitioning parallel digraph tasks onto the heterogeneous cores, and (ii) assigning core operating frequencies to reduce the overall energy consumption, while meeting real-time constraints. The effectiveness of the proposed approach is validated with an exhaustive set of simulations. Houssam-Eddine Zahaf, Giuseppe Lipari, Marko Bertogna, Pierre Boulet |
IEEE Trans. Computers | 3 |
| 2018 | NVIDIA GPU scheduling details in virtualized environments: work-in-progressabstractModern automotive grade embedded platforms feature high performance Graphics Processing Units (GPUs) to support the massively parallel processing power needed for next-generation autonomous driving applications. Hence, a GPU scheduling approach with strong Real-Time guarantees is needed. While previous research efforts focused on reverse engineering the GPU ecosystem in order to understand and control GPU scheduling on NVIDIA platforms, we provide an in depth explanation of the NVIDIA standard approach to GPU application scheduling on a Drive PX platform. Then, we discuss how a privileged scheduling server can be used to enforce arbitrary scheduling policies in a virtualized environment. Nicola Capodieci, Roberto Cavicchioli, Marko Bertogna |
EMSOFT | 3 |
| 2018 | Deadline-Based Scheduling for GPU with Preemption SupportabstractModern automotive-grade embedded computing platforms feature high-performance Graphics Processing Units (GPUs) to support the massively parallel processing power needed for next-generation autonomous driving applications (e.g., Deep Neural Network (DNN) inference, sensor fusion, path planning, etc). As these workload-intensive activities are pushed to higher criticality levels, there is a stronger need for more predictable scheduling algorithms that are able to guarantee predictability without overly sacrificing GPU utilization. Unfortunately, the real-rime literature on GPU scheduling mostly considered limited (or null) preemption capabilities, while previous efforts in broader domains were often based on programming models and APIs that were not designed to support the real-rime requirements of recurring workloads. In this paper, we present the design of a prototype real-time scheduler for GPU activities on an embedded System on a Chip (SoC) featuring a cutting edge GPU architecture by NVIDIA adopted in the autonomous driving domain. The scheduler runs as a software partition on top of the NVIDIA hypervisor, and it leverages latest generation architectural features, such as pixel-level preemption and threadlevel preemption. Such a design allowed us to implement and test a preemptive Earliest Deadline First (EDF) scheduler for GPU tasks providing bandwidth isolations by means of a Constant Bandwidth Server (CBS). Our work involved investigating alternative programming models for compute APIs, allowing us to characterize CPU-to-GPU command submission with more detailed scheduling information. A detailed experimental characterization is presented to show the significant schedulability improvement of recurring real-time GPU tasks. Nicola Capodieci, Roberto Cavicchioli, Marko Bertogna, Aingara Paramakuru |
RTSS | 3 |
| 2018 | Big Data Analytics for Smart Cities: The H2020 CLASS ProjectabstractNo abstract available. Eduardo Quiñones, Marko Bertogna, Erez Hadad, Ana Juan Ferrer, Luca Chiantore, Alfredo Reboa |
SYSTOR | 2 |
| 2018 | A design flow for supporting component-based software development in multiprocessor real-time systems
Alessandro Biondi 0001, Giorgio C. Buttazzo, Marko Bertogna |
Real Time Syst. | 3 |
| 2018 | Analytical Characterization of End-to-End Communication Delays With Logical Execution TimeabstractModern automotive embedded systems are composed of multiple real-time tasks communicating by means of shared variables. The effect of an initial event is typically propagated to an actuation signal through sequences of tasks writing/reading shared variables, creating an effect chain (EC). The responsiveness, performance and stability of the control algorithms of an automotive application typically depend on the propagation delays of selected ECs. Indeed, task jitter can have a negative impact on the system potentially leading to instability. The logical execution time (LET) model has been recently adopted by the automotive industry as a way of reducing jitter and improving the determinism of the system. In this paper, we provide a formal analysis of the LET model for real-time systems composed of periodic tasks with harmonic and nonharmonic periods, analytically characterizing the control performance of LET ECs. We also show that by introducing tasks offsets, the real-time performance of nonharmonic tasks may improve, getting closer to the constant end-to-end latency experienced in the harmonic case. Further, we present a heuristic algorithm to obtain a set of offsets that might reduce end-to-end latencies, improving LET communication determinism. Finally, we apply this technique to an industrial case study consisting of an automotive engine control system. Jorge Martinez 0003, Ignacio Sanudo Olmedo, Marko Bertogna |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | A static scheduling approach to enable safety-critical OpenMP applicationsabstractParallel computation is fundamental to satisfy the performance requirements of advanced safety-critical systems. OpenMP is a good candidate to exploit the performance opportunities of parallel platforms. However, safety-critical systems are often based on static allocation strategies, whereas current OpenMP implementations are based on dynamic schedulers. This paper proposes two OpenMP-compliant static allocation approaches: an optimal but costly approach based on an ILP formulation, and a sub-optimal but tractable approach that computes a worst-case makespan bound close to the optimal one. Alessandra Melani, Maria A. Serrano, Marko Bertogna, Isabella Cerutti, Eduardo Quiñones, Giorgio C. Buttazzo |
ASP-DAC | 3 |
| 2017 | Memory interference characterization between CPU cores and integrated GPUs in mixed-criticality platformsabstractMost of today's mixed criticality platforms feature Systems on Chip (SoC) where a multi-core CPU complex (the host) competes with an integrated Graphic Processor Unit (iGPU, the device) for accessing central memory. The multi-core host and the iGPU share the same memory controller, which has to arbitrate data access to both clients through often undisclosed or non-priority driven mechanisms. Such aspect becomes critical when the iGPU is a high performance massively parallel computing complex potentially able to saturate the available DRAM bandwidth of the considered SoC. The contribution of this paper is to qualitatively analyze and characterize the conflicts due to parallel accesses to main memory by both CPU cores and iGPU, so to motivate the need of novel paradigms for memory centric scheduling mechanisms. We analyzed different well known and commercially available platforms in order to estimate variations in throughput and latencies within various memory access patterns, both at host and device side. Roberto Cavicchioli, Nicola Capodieci, Marko Bertogna |
ETFA | 3 |
| 2017 | An Analysis of Lazy and Eager Limited Preemption Approaches under DAG-Based Global Fixed Priority SchedulingabstractDAG-based scheduling models have been shown to effectively express the parallel execution of current many-core heterogeneous architectures. However, their applicability to real-time settings is limited by the difficulties to find tight estimations of the worst-case timing parameters of tasks that may arbitrarily be preempted/migrated at any instruction. An efficient approach to increase the system predictability is to limit task preemptions to a set of pre-defined points. This limited preemption model supports two different preemption approaches, eager and lazy, which have been analyzed only for sequential task-sets. This paper proposes a new response time analysis that computes an upper bound on the lower priority blocking that each task may incur with eager and lazy preemptions. We evaluate our analysis with both, synthetic DAG-based task-sets and a real case-study from the automotive domain. Results from the analysis demonstrate that, despite the eager approach generates a higher number of priority inversions, the blocking impact is generally smaller than in the lazy approach, leading to a better schedulability performance. Maria A. Serrano, Alessandra Melani, Sebastian Kehr, Marko Bertogna, Eduardo Quiñones |
ISORC | 4 |
| 2017 | Adaptive Coordination in Autonomous Driving: Motivations and PerspectivesabstractAs autonomous cars are entering mainstream, new research directions are opening involving several domains, from hardware design to control systems, from energy efficiency to computer vision. An exciting direction of research is represented by the coordination of the different vehicles, moving the focus from the single one to a collective system. In this paper we propose some challenging examples thatshow the motivations for a coordination approach in autonomous driving. Moreover, we present some techniques borrowed from distributed artificial intelligence that can be exploited to tackle the previously mentioned challenges. Marko Bertogna, Paolo Burgio, Giacomo Cabri, Nicola Capodieci |
WETICE | 1 |
| 2017 | Schedulability Analysis of Conditional Parallel Task Graphs in Multicore SystemsabstractSeveral task models have been introduced in the literature to describe the intrinsic parallelism of real-time activities, including fork/join, synchronous parallel, DAG-based, etc. Although schedulability tests and resource augmentation bounds have been derived for these task models in the context of multicore systems, they are still too pessimistic to describe the execution flow of parallel tasks characterized by multiple (and nested) conditional statements, where it is hard to decide which execution path to select for modeling the worst-case scenario. To overcome this problem, this paper proposes a task model that integrates control flow information by considering conditional parallel tasks (cp-tasks) represented by DAGs with both precedence and conditional edges. For this task model, a set of meaningful parameters are identified and computed by efficient algorithms and a response-time analysis is presented for different scheduling policies. Experimental results are finally reported to evaluate the efficiency of the proposed schedulability tests and their performance with respect to classic tests based on both conditional and non-conditional existing approaches. Alessandra Melani, Marko Bertogna, Vincenzo Bonifaci, Alberto Marchetti-Spaccamela, Giorgio C. Buttazzo |
IEEE Trans. Computers | 2 |
| 2017 | Exact Response Time Analysis for Fixed Priority Memory-Processor Co-SchedulingabstractRecent technological advances have led to an increasing gap between memory and processor performance, since memory bandwidth is progressing at a much slower pace than processor bandwidth. Pre-fetching techniques are traditionally used to bridge this gap and achieve high processor utilization while tolerating high memory latencies. Following this trend, new computational models have been proposed to split task execution in two consecutive phases: a memory phase in which the required instructions and data are pre-fetched to local memory (M-phase), and an execution phase in which the task is executed with no memory contention (C-phase). Decoupling memory and execution phases not only simplifies the timing analysis, but also allows a more efficient (and predictable) pipelining of memory and execution phases through proper co-scheduling algorithms. This paper takes a further step towards the design of smart co-scheduling algorithms for sporadic real-time tasks complying with the memory-computation (M/C) model, by proposing a theoretical framework aimed at tightly characterizing the schedulability improvement obtainable with the adopted M/C task model on single-core systems. In particular, a critical instant is identified for M/C tasks scheduled with fixed priority and an exact response time analysis with pseudo-polynomial complexity is provided. Then, we investigate the problem of priority assignment for M/C tasks, showing that a necessary condition to achieve optimality is to allow different priorities for the two phases. Our experiments show that the proposed techniques provide a significant schedulability improvement with respect to classic execution models, placing an important building block towards the design of more efficient partitioned multi-core systems. Alessandra Melani, Marko Bertogna, Robert I. Davis 0001, Vincenzo Bonifaci, Alberto Marchetti-Spaccamela, Giorgio C. Buttazzo |
IEEE Trans. Computers | 2 |
| 2016 | Response-time analysis of DAG tasks under fixed priority scheduling with limited preemptions
Maria A. Serrano, Alessandra Melani, Marko Bertogna, Eduardo Quiñones |
DATE | 3 |
| 2016 | A Software Stack for Next-Generation Automotive Systems on Many-Core Heterogeneous PlatformsabstractThe advent of commercial-of-the-shelf (COTS) heterogeneous many-core platforms is opening up a series of opportunities in the embedded computing market. Integrating multiple computing elements running at lower frequencies allows obtaining impressive performance capabilities at a reduced power consumption. These platforms can be successfully adopted to build the next-generation of self-driving vehicles, where Advanced Driver Assistance Systems (ADAS) need to process unprecedently higher computing workloads at low power budgets. Unfortunately, the current methodologies for providing real-time guarantees are uneffective when applied to the complex architectures of modern many-cores. Having impressive average performances with no guaranteed bounds on the response times of the critical computing activities is of little if no use to these applications. Project HERCULES will provide the required technological infrastructure to obtain an order-of-magnitude improvement in the cost and power consumption of next generation automotive systems. This paper presents the integrated software framework of the project, which allows achieving predictable performance on top of cutting-edge heterogeneous COTS platforms. The proposed software stack will let both real-time and non real-time application coexist on next-generation, power-efficient embedded platform, with preserved timing guarantees. Paolo Burgio, Marko Bertogna, Ignacio Sanudo Olmedo, Paolo Gai, Andrea Marongiu, Michal Sojka |
DSD | 2 |
| 2016 | A review of priority assignment in real-time systems
Robert I. Davis 0001, Liliana Cucu-Grosjean, Marko Bertogna, Alan Burns 0001 |
J. Syst. Archit. | 3 |
| 2016 | Guest editorial: special issue on multicore systems
Marco Caccamo, Marko Bertogna |
Real Time Syst. | 2 |
| 2016 | On the compatibility of exact schedulability tests for global fixed priority pre-emptive scheduling with Audsley's optimal priority assignment algorithm
Robert I. Davis 0001, Marko Bertogna, Vincenzo Bonifaci |
Real Time Syst. | 2 |
| 2016 | Schedulability Analysis of Hierarchical Real-Time Systems under Shared ResourcesabstractSharing resources in hierarchical real-time systems implemented with reservation servers requires the adoption of special budget management protocols that preserve the bandwidth allocated to a specific component. In addition, blocking times must be accurately estimated to guarantee both the global feasibility of all the servers and the local schedulability of applications running on each component. This paper presents two new local schedulability tests to verify the schedulability of real-time applications running on reservation servers under fixed priority and EDF local schedulers. Reservation servers are implemented with the BROE algorithm. A simple extension to the SRP protocol is also proposed to reduce the blocking time of the server when accessing global resources shared among components. The performance of the new schedulability tests are compared with other solutions proposed in the literature, showing the effectiveness of the proposed improvements. Finally, an implementation of the main protocols on a lightweight RTOS is described, highlighting the main practical issues that have been encountered. Alessandro Biondi 0001, Giorgio C. Buttazzo, Marko Bertogna |
IEEE Trans. Computers | 3 |
| 2015 | Timing characterization of OpenMP4 tasking modelabstractOpenMP is increasingly being supported by the newest high-end embedded many-core processors. Despite the lack of any notion of real-time execution, the latest specification of OpenMP (v4.0) introduces a tasking model that resembles the way real-time embedded applications are modeled and designed, i.e., as a set of periodic task graphs. This makes OpenMP4 a convenient candidate to be adopted in future real-time systems. However, OpenMP4 incorporates as well features to guarantee backward compatibility with previous versions that limit its practical usability in real-time systems. The most notable example is the distinction between tied and untied tasks. Tied tasks force all parts of a task to be executed on the same thread that started the execution, whereas a suspended untied task is allowed to resume execution on a different thread. Moreover, tied tasks are forbidden to be scheduled in threads in which other non-descendant tied tasks are suspended. As a result, the execution model of tied tasks, which is the default model in OpenMP to simplify the coexistence with legacy constructs, clearly restricts the performance and has serious implications on the response time analysis of OpenMP4 applications, making difficult to adopt it in real-time environments. In this paper, we revisit OpenMP design choices, introducing timing predictability as a new and key metric of interest. Our first results confirm that even if tied tasks can be timing analyzed, the quality of the analysis is much worse than with untied tasks. We thus reason about the benefits of using untied tasks, deriving a response time analysis for this model, and so allowing OpenMP4 untied model to be applied to real-time systems. Maria A. Serrano, Alessandra Melani, Roberto Vargas, Andrea Marongiu, Marko Bertogna, Eduardo Quiñones |
CASES | 5 |
| 2015 | Supporting Component-Based Development in Partitioned Multiprocessor Real-Time SystemsabstractThe fast evolution of multicore systems, combined with the need of sharing the same platform for independently developed software, demands for new methodologies and algorithms that allow resource partitioning, while guaranteeing the isolation of concurrent applications. Unfortunately, a major problem that can break the isolation property of concurrent partitions is resource sharing. Although a number of resource access protocols exist for hierarchical uniprocessor systems, no protocols are available today for managing hierarchical partitions implemented on top a multiporcessor platform under partitioned scheduling. This paper presents a framework to support component based design on a multiprocessor platform and proposes a novel reservation server mechanism, called M-BROE, to handle shared resources in multiprocessor systems in the presence of resource reservation scheduling mechanisms. Alessandro Biondi 0001, Giorgio C. Buttazzo, Marko Bertogna |
ECRTS | 3 |
| 2015 | Response-Time Analysis of Conditional DAG Tasks in Multiprocessor SystemsabstractDifferent task models have been proposed to represent the parallel structure of real-time tasks executing on manycore platforms: fork/join, synchronous parallel, DAG-based, etc. Despite different schedulability tests and resource augmentation bounds are available for these task systems, we experience difficulties in applying such results to real application scenarios, where the execution flow of parallel tasks is characterized by multiple (and nested) conditional structures. When a conditional branch drives the number and size of sub-jobs to spawn, it is hard to decide which execution path to select for modeling the worst-case scenario. To circumvent this problem, we integrate control flow information in the task model, considering conditional parallel tasks (cp-tasks) represented by DAGs composed of both precedence and conditional edges. For this task model, we identify meaningful parameters that characterize the schedulability of the system, and derive efficient algorithms to compute them. A response time analysis based on these parameters is then presented for different scheduling policies. A set of simulations shows that the proposed approach allows efficiently checking the schedulability of the addressed systems, and that it significantly tightens the schedulability analysis of non-conditional (e.g., Classic DAG) tasks over existing approaches. Alessandra Melani, Marko Bertogna, Vincenzo Bonifaci, Alberto Marchetti-Spaccamela, Giorgio C. Buttazzo |
ECRTS | 2 |
| 2015 | Global and Partitioned Multiprocessor Fixed Priority Scheduling with Deferred PreemptionabstractThis article introduces schedulability analysis for Global Fixed Priority Scheduling with Deferred Preemption (gFPDS) for homogeneous multiprocessor systems. gFPDS is a superset of Global Fixed Priority Preemptive Scheduling (gFPPS) and Global Fixed Priority Nonpreemptive Scheduling (gFPNS). We show how schedulability can be improved using gFPDS via appropriate choice of priority assignment and final nonpreemptive region lengths, and provide algorithms that optimize schedulability in this way. Via an experimental evaluation we compare the performance of multiprocessor scheduling using global approaches: gFPDS, gFPPS, and gFPNS, and also partitioned approaches employing FPDS, FPPS, and FPNS on each processor. Robert I. Davis 0001, Alan Burns 0001, Vincent Nélis, Stefan M. Petters, Marko Bertogna |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2014 | P-SOCRATES: A Parallel Software Framework for Time-Critical Many-Core SystemsabstractThe advent of next-generation many-core embedded platforms has the chance of intercepting a converging need for predictable high-performance coming from both the High-Performance Computing (HPC) and Embedded Computing (EC) domains. On one side, new kinds of HPC applications are being required by markets needing huge amounts of information to be processed within a bounded amount of time. On the other side, EC systems are increasingly concerned with providing higher performance in real-time, challenging the performance capabilities of current architectures. This converging demand, however, raises the problem about how to guarantee timing requirements in presence of parallel execution. This paper presents the approach of project P-SOCRATES for the design of an integrated framework for the execution of workload-intensive applications with real-time requirements on top of next-generation commercial-off-the-shelf (COTS) platforms based on many-core accelerated architectures. The time-criticality and parallelisation challenges are addressed by merging techniques coming from both HPC and EC domains, identifying the main sources of indeterminism and proposing efficient mapping and scheduling algorithms, along with the associated timing and schedulability analysis, to guarantee the real-time and performance requirements of the applications. Luís Miguel Pinho, Eduardo Quiñones, Marko Bertogna, Andrea Marongiu, Jorge Pereira Carlos, Claudio Scordino, Michele Ramponi |
DSD | 3 |
| 2014 | Optimal Design for Reservation Servers under Shared ResourcesabstractModularity and hierarchical-based design are crucial features that need to be supported in complex embedded systems characterized by multiple applications with timing requirements.Resource reservation is a powerful scheduling mechanism for achieving such goals and providing temporal isolation among different real-time applications. When different applications share mutually exclusive resources, a precise feasibility analysis can still be performed in isolation, using specific resource access protocols, taking into account only the application features and the reservation parameters. This paper presents a methodology for selecting the parameters of each reservation in order to guarantee the feasibility of the served applications and minimize the required bandwidth. Alessandro Biondi 0001, Alessandra Melani, Marko Bertogna, Giorgio C. Buttazzo |
ECRTS | 3 |
| 2014 | Explicit Preemption Placement for Real-Time Conditional CodeabstractIn the limited-preemption scheduling model, tasks cooperate to offer suitable preemption points for reducing the overall preemption overhead. In the fixed preemption-point model, tasks are allowed to preempt only at statically defined preemption points, reducing the variability of the preemption delay and making the system more predictable. Different works have been proposed to determine the optimal selection of preemption points for minimizing the preemption overhead without affecting the system schedulability due to increased non-preemptivity. However, all works are based on very restrictive task models, without being able to deal with common coding structures like branches, conditional statements and loops. In this work, we overcome this limitation, by proposing a pseudo-polynomial-time algorithm that is capable of determining the optimal set of preemption points to minimize the worst-case execution time of jobs represented by control flow graphs with arbitrarily-nested conditional structures, while preserving system schedulability. Exhaustive experiments are included to show that the proposed approach is able to significantly improve the bounds on the worst-case execution times of limited preemptive tasks. Nathan Fisher, Marko Bertogna |
ECRTS | 3 |
| 2014 | On the effectiveness of energy-aware real-time scheduling algorithms on single-core platformsabstractEnergy-aware scheduling is a challenging problem that has been studied for decades, investigating the trade-off between performance and energy consumption. In early CMOS circuits, Dynamic Voltage and Frequency Scaling (DVFS) techniques allowed drastically reducing the power consumption. Recent technological advancements have decreased the portion of dissipation which is affected by speed scaling, making Dynamic Power Management (DPM) algorithms more effective. However, the adoption of simplistic power models often biased the decision on which technique to adopt, decreasing the effectiveness of the selected implementation. This paper discusses the factors to consider when deciding which technique to implement on a given single-core architecture, highlight the limitations of the current mainstream. Mario Bambagini, Marko Bertogna, Giorgio C. Buttazzo |
ETFA | 2 |
| 2013 | Global fixed priority scheduling with deferred pre-emptionabstractThis paper introduces schedulability analysis for global fixed priority scheduling with deferred pre-emption (gFPDS) for homogeneous multiprocessor systems. gFPDS is a superset of global fixed priority pre-emptive scheduling (gFPPS) and global fixed priority non-pre-emptive scheduling (gFPNS). We show how schedulability can be improved via appropriate choice of priority assignment and final non-pre-emptive region lengths, and we provide algorithms which optimize schedulability in this way. An experimental evaluation shows that gFPDS significantly outperforms both gFPPS and gFPNS. Robert I. Davis 0001, Alan Burns 0001, Vincent Nélis, Stefan M. Petters, Marko Bertogna |
RTCSA | 6 |
| 2013 | Limited Pre-emptive Global Fixed Task PriorityabstractIn this paper a limited pre-emptive global fixed task priority scheduling policy for multiprocessors is presented. This scheduling policy is a generalization of global fully pre-emptive and non-pre-emptive fixed task priority policies for platforms with at least two homogeneous processors. The scheduling protocol devised is such that a job can only be blocked at most once by a body of lower priority non-pre-emptive workload. The presented policy dominates both fully pre-emptive and fully non-pre-emptive with respect to schedulability. A sufficient schedulability test is presented for this policy. Several approaches to estimate the blocking generated by lower priority non-pre-emptive regions are presented. As a last contribution it is experimentally shown that, on the average case, the number of pre-emptions observed in a schedule are drastically reduced in comparison to global fully pre-emptive scheduling. Vincent Nélis, Stefan M. Petters, Marko Bertogna, Robert I. Davis 0001 |
RTSS | 4 |
| 2013 | Limited Preemptive Scheduling for Real-Time Systems. A SurveyabstractThe question whether preemptive algorithms are better than nonpreemptive ones for scheduling a set of real-time tasks has been debated for a long time in the research community. In fact, especially under fixed priority systems, each approach has advantages and disadvantages, and no one dominates the other when both predictability and efficiency have to be taken into account in the system design. Recently, limited preemption models have been proposed as a viable alternative between the two extreme cases of fully preemptive and nonpreemptive scheduling. This paper presents a survey of the existing approaches for reducing preemptions and compares them under different metrics, providing both qualitative and quantitative performance evaluations. Giorgio C. Buttazzo, Marko Bertogna |
IEEE Trans. Ind. Informatics | 2 |
| 2012 | Optimal Fixed Priority Scheduling with Deferred Pre-emptionabstractThe schedulability of systems using fixed priority pre-emptive scheduling can be improved by the use of non-pre-emptive regions at the end of each task's execution, an approach referred to as deferred pre-emption. Choosing the appropriate length for the final non-pre-emptive region of each task is a trade-off between improving the worst-case response time of the task itself and increasing the amount of blocking imposed on higher priority tasks. In this paper we present an optimal algorithm for determining both the priority ordering of tasks and the lengths of their final non-pre-emptive regions. This algorithm is optimal for fixed priority scheduling with deferred pre-emption, in the sense that it is guaranteed to find a schedulable combination of priority ordering and final non-pre-emptive region lengths if such a schedulable combination exists. Robert I. Davis 0001, Marko Bertogna |
RTSS | 2 |
| 2011 | Optimal Selection of Preemption Points to Minimize Preemption OverheadabstractA central issue for verifying the schedulability of hard real-time systems is the correct evaluation of task execution times. These values are significantly influenced by the preemption overhead, which mainly includes the cache related delays and the context switch times introduced by each preemption. Since such an overhead significantly depends on the particular point in the code where preemption takes place, this paper proposes a method for placing suitable preemption points in each task in order to maximize the chances of finding a schedulable solution. In a previous work, we presented a method for the optimal selection of preemption points under the restrictive assumption of a fixed preemption cost, identical for each preemption point. In this paper, we remove such an assumption, exploring a more realistic and complex scenario where the preemption cost varies throughout the task code. Instead of modeling the problem with an integer programming formulation, with exponential worst-case complexity, we derive an optimal algorithm that has a linear time and space complexity. This somewhat surprising result allows selecting the best preemption points even in complex scenarios with a large number of potential preemption locations. Experimental results are also presented to show the effectiveness of the proposed approach in increasing the system schedulability. Marko Bertogna, Orges Xhani, Mauro Marinoni, Francesco Esposito, Giorgio C. Buttazzo |
ECRTS | 1 |
| 2011 | Improving Feasibility of Fixed Priority Tasks Using Non-Preemptive RegionsabstractPreemptive schedulers have been widely adopted in single processor real-time systems to avoid the blocking associated with the non-preemptive execution of lower priority tasks and achieve a high processor utilization. However, under fixed priority assignments, there are cases in which limiting preemptions can improve schedulability with respect to a fully preemptive solution. This is true even neglecting preemption overhead, as it will be shown in the paper. In previous works, limited-preemption schedulers have been mainly considered to reduce the preemption overhead, and make the estimation of worst-case execution times more predictable. In this work, we instead show how to improve the feasibility of fixed-priority task systems by executing the last portion of each task in a non-preemptive fashion. A proper dimensioning of such a region of code allows increasing the number of task sets that are schedulable with a fixed priority algorithm. Simulation experiments are also presented to validate the effectiveness of the proposed approach. Marko Bertogna, Giorgio C. Buttazzo |
RTSS | 1 |
| 2011 | Tests for global EDF schedulability analysis
Marko Bertogna, Sanjoy Baruah |
J. Syst. Archit. | 1 |
| 2011 | Feasibility analysis under fixed priority scheduling with limited preemptions
Giorgio C. Buttazzo, Marko Bertogna |
Real Time Syst. | 3 |
| 2010 | Preemption Points Placement for Sporadic Task SetsabstractLimited preemption scheduling has been introduced as a viable alternative to non-preemptive and fully preemptive scheduling when reduced blocking times need to coexist with an acceptable context switch overhead. To achieve this goal, preemptions are allowed only at selected points of the code of each task, decreasing the preemption overhead and simplifying the estimation of worst-case execution parameters. Unfortunately, the problem of how to place these preemption points is rather complex and has not been solved. In this paper, a method is presented for the optimal placement of preemption points under simplifying conditions, namely, a fixed preemption overhead at each point. We will prove that if our method is not able to produce a feasible schedule, then no other possible preemption point placement (including non-preemptive and fully preemptive scheduling) can find a schedulable solution. The presented method is general enough to be applicable to both EDF and Fixed Priority scheduling, with limited modifications. Marko Bertogna, Giorgio C. Buttazzo, Mauro Marinoni, Francesco Esposito, Marco Caccamo |
ECRTS | 1 |
| 2010 | Comparative evaluation of limited preemptive methodsabstractSchedulability analysis of real-time systems requires the knowledge of the worst-case execution time (WCET) of each computational activity. A precise estimation of such a task parameter is quite difficult to achieve, because execution times depend on many factors, including the task structure, the system architecture details, operating system features and so on. While some of these features are not under our control, selecting a proper scheduling algorithm can reduce the runtime overhead and make the WCETs smaller and more predictable. In particular, since task execution times can be significantly affected by preemptions, a number of scheduling methods have been proposed in the real-time literature to limit preemption during task execution. In this paper, we provide a comprehensive overview of the possible scheduling approaches that can be used to contain preemptions and present a comparative study aimed at evaluating their impact on task execution times. Giorgio C. Buttazzo, Marko Bertogna |
ETFA | 3 |
| 2010 | Feasibility Analysis under Fixed Priority Scheduling with Fixed Preemption PointsabstractLimited preemption models have been proposed as a viable alternative between the two extreme cases of fully preemptive and non-preemptive scheduling. In particular, allowing preemption to occur only at predefined preemption points reduces context switch costs, simplifies the access to shared resources, and allows more predictable estimations of worst-case execution times. Current results related to such a model, however, exhibit two major deficiencies: (i) The exact response time analysis has a high computational complexity; (ii) The maximum lengths of then on-preemptive regions was not completely investigated in all possible scenarios. In this paper, we address the problem of scheduling a set of real-time tasks having fixed priorities and fixed preemption points. In particular, under specific but not restrictive assumptions we simplified the feasibility analysis and proposed an efficient feasibility test. Finally, an algorithm for computing the maximum length of fixed non-preemptive regions for each task is described, and some simulation experiments are presented to validate the proposed approach. Giorgio C. Buttazzo, Marko Bertogna |
RTCSA | 3 |
| 2010 | Limited preemption EDF scheduling of sporadic task systemsabstractThe optimality of the Earliest Deadline First scheduler for uniprocessor systems is one of the main reasons behind the popularity of this algorithm among real-time systems. The ability of fully utilizing the computational power of a processing unit however requires the possibility of preempting a task before its completion. When preemptions are disabled, the schedulability overhead could be significant, leading to deadline misses even at system utilizations close to zero. On the other hand, each preemption causes an increase in the runtime overhead due to the operations executed during a context switch and the negative cache effects resulting from interleaving tasks' executions. These factors have been often neglected in previous theoretical works, ignoring the cost of preemption in real applications. A hybrid limited-preemption real-time scheduling algorithm is derived here, that aims to have low runtime overhead while scheduling all systems that can be scheduled by fully preemptive algorithms. This hybrid algorithm permits preemption where necessary for maintaining feasibility, but attempts to avoid unnecessary preemptions during runtime. The positive effects of this approach are not limited to a reduced runtime overhead, but will be extended as well to a simplified handling of shared resources. Marko Bertogna, Sanjoy Baruah |
IEEE Trans. Ind. Informatics | 1 |
| 2009 | Improving Task Responsiveness with Limited PreemptionsabstractThe optimality of preemptive EDF scheduling with relation to the achievable system utilization is a clear advantage of this scheduling policy for single processor real-time systems. However, recent works suggested that the run-time behavior of EDF might be improved by limiting the preemption support only to particular time instants, dividing each task into a sequence of non-preemptive chunks of execution, without affecting the schedulability of the system. In this paper, we will take a closer look to limited preemption EDF scheduling (LP-EDF), evaluating the potential advantages offered by this policy in terms of response-time reduction and improved control performances. In particular, we will show how to increase the responsiveness of a control application by placing non-preemptive regions of maximal length at the end of the code of selected tasks. The effectiveness of the proposed method will be proved both analytically and by extensive simulations. Marko Bertogna |
ETFA | 2 |
| 2009 | The Multi Supply Function Abstraction for MultiprocessorsabstractMulti-core platforms are becoming the dominant computing architecture for next generation embedded systems. Nevertheless, designing, programming, and analyzing such systems is not easy and a solid methodology is still missing. In this paper, we propose two powerful abstractions to model the computing power of a parallel machine, which provide a general interface for developing and analyzing real-time applications in isolation, independently of the physical platform. The proposed abstractions can be applied on top of different types of service mechanisms, such as periodic servers, static partitions, and P-fair time partitions. In addition, we developed the schedulability analysis of a set of real-time tasks on top of a parallel machine that is compliant with the proposed abstractions. Enrico Bini, Giorgio C. Buttazzo, Marko Bertogna |
RTCSA | 3 |
| 2009 | Bounding the Maximum Length of Non-preemptive Regions under Fixed Priority SchedulingabstractThe question whether preemptive systems are better than non-preemptive systems has been debated for a long time, but only partial answers have been provided in the real-time literature and still some issues remain open. In fact, each approach has advantages and disadvantages, and no one dominates the other when both predictability and efficiency have to be taken into account in the system design. In particular, limiting preemptions allows increasing program locality, making timing analysis more predictable with respect to the fully preemptive case. In this paper, we integrate the features of both preemptive and non-preemptive scheduling by considering that each task can switch to non-preemptive mode, at any time, for a bounded interval. Three methods (with different complexity and performance) are presented to calculate the longest non-preemptive interval that can be executed by each task, under fixed priorities, without degrading the schedulability of the task set, with respect to the fully preemptive case. The methods are also compared by simulations to evaluate their effectiveness in reducing the number of preemptions. Giorgio C. Buttazzo, Marko Bertogna |
RTCSA | 3 |
| 2009 | Virtual Multiprocessor Platforms: Specification and UseabstractA new abstraction — the Parallel Supply Function (PSF) — is proposed for representing the computing capabilities offered by virtual platforms implemented atop identical multiprocessors. It is shown that this abstraction is strictly more powerful than previously-proposed ones, from the perspective of more accurately representing the inherent parallelism of the provided computing capabilities. Sufficient tests are derived for determining whether a given real-time task system, represented as a collection of sporadic tasks, is guaranteed to always meet all deadlines when scheduled upon a specified virtual platform using the global EDF scheduling algorithm. Enrico Bini, Marko Bertogna, Sanjoy Baruah |
RTSS | 2 |
| 2009 | Resource holding times: computation and optimization
Marko Bertogna, Nathan Fisher, Sanjoy Baruah |
Real Time Syst. | 1 |
| 2009 | Resource-sharing servers for Open EnvironmentsabstractWe study the problem of executing a collection of independently designed and validated task systems upon a common platform composed of a preemptive processor and additional shared resources. We present an abstract formulation of the problem and identify the major issues that must be addressed in order to solve this problem. We present and prove the correctness of algorithms that address these issues, and thereby obtain a design for an open real-time environment. Marko Bertogna, Nathan Fisher, Sanjoy Baruah |
IEEE Trans. Ind. Informatics | 1 |
| 2009 | Schedulability Analysis of Global Scheduling Algorithms on Multiprocessor PlatformsabstractThis paper addresses the schedulability problem of periodic and sporadic real-time task sets with constrained deadlines preemptively scheduled on a multiprocessor platform composed by identical processors. We assume that a global work-conserving scheduler is used and migration from one processor to another is allowed during a task lifetime. First, a general method to derive schedulability conditions for multiprocessor real-time systems will be presented. The analysis will be applied to two typical scheduling algorithms: earliest deadline first (EDF) and fixed priority (FP). Then, the derived schedulability conditions will be tightened, refining the analysis with a simple and effective technique that significantly improves the percentage of accepted task sets. The effectiveness of the proposed test is shown through an extensive set of synthetic experiments. Marko Bertogna, Michele Cirinei, Giuseppe Lipari |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2008 | EDZL scheduling analysis
Theodore P. Baker, Michele Cirinei, Marko Bertogna |
Real Time Syst. | 3 |
| 2007 | Static-Priority Scheduling and Resource Hold TimesabstractThe duration of time for which each application locks each shared resource is critically important in composing multiple independently-developed applications upon a shared "open" platform. In a companion paper, we formally defined and studied the concept of resource hold time (RHT) - the largest length of time that may elapse between the instant that an application system locks a resource and the instant that it subsequently releases the resource. We extend the discussion and results from to systems scheduled using static-priority scheduling algorithms, with resource access arbitrated using stack resource policy (SRP), or priority ceiling protocol (PCP). We present a method to compute resource hold times for every resource, and an algorithm to decrease them without changing the semantics of the application or compromising application feasibility. Marko Bertogna, Nathan Fisher, Sanjoy Baruah |
IPDPS | 1 |
| 2007 | Resource-Locking Durations in EDF-Scheduled SystemsabstractThe duration of time for which each application locks each shared resource is critically important in composing multiple independently-developed applications upon a shared "open" platform. The concept of resource hold time (RHT) - the largest length of time that may elapse between the instant that an application system locks a resource and the instant that it subsequently releases the resource - is formally defined and studied in this paper. An algorithm is presented for computing resource hold times for every resource in an application that is scheduled using earliest deadline first scheduling, with resource access arbitrated using the stack resource policy. An algorithm is presented for decreasing these RHT's without changing the semantics of the application or compromising application feasibility Nathan Fisher, Marko Bertogna, Sanjoy Baruah |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2007 | Response-Time Analysis for Globally Scheduled Symmetric Multiprocessor PlatformsabstractIn the last years, a progressive migration from single processor chips to multi-core computing devices has taken place in the general-purpose and embedded system market. The development of multi-processor systems is already a core activity for the most important hardware companies. A lot of different solutions have been proposed to overcome the physical limits of single core devices and to address the increasing computational demand of modern multimedia applications. The real-time community followed this trend with an increasing number of results adapting the classical scheduling analysis to parallel computing systems. This paper will contribute to refine the schedulability analysis for symmetric multi-processor (SMP) real-time systems composed by a set of periodic and sporadic tasks. We will focus on both fixed and dynamic priority global scheduling algorithms, where tasks can migrate from one processor to another during execution. By increasing the complexity of the analysis, we will show that an improvement is possible over existing schedulability tests, significantly increasing the number of schedulable task sets detected. The added computational effort is comparable to the cost of techniques widely used in the uniprocessor case. We believe this is a reasonable cost to pay, given the intrinsically higher complexity of multi-processor devices. Marko Bertogna, Michele Cirinei |
RTSS | 1 |
| 2007 | The Design of an EDF-Scheduled Resource-Sharing Open EnvironmentabstractWe study the problem of executing a collection of independently designed and validated task systems upon a common platform comprised of a preemptive processor and additional shared resources. We present an abstract formulation of the problem and identify the major issues that must be addressed in order to solve this problem. We present (and prove the correctness of) algorithms that address these issues, and thereby obtain a design for an open real-time environment in the presence of shared global resources. Nathan Fisher, Marko Bertogna, Sanjoy Baruah |
RTSS | 2 |
| 2005 | Improved Schedulability Analysis of EDF on Multiprocessor PlatformsabstractMultiprocessor hardware platforms are now being considered for embedded systems, due to their high computational power and little additional cost when compared to single processor systems. When scheduling real-time applications on multiprocessor platforms, a possibility is to use global scheduling, where a scheduling algorithm dynamically assign tasks to processors, and tasks can migrate from one processor to another during their execution. In this paper, we tackle the problem of schedulability analysis of sporadic tasks in global scheduling systems, where the scheduler is the earliest deadline first (EDF) algorithm. We provide two main contributions. First, we show that two recently proposed tests perform poorly when the task set contains heavy tasks (i.e. tasks with high utilization). We also show that neither test dominates the other. As a second contribution, we introduce a new schedulability test that improves significantly the percentage of accepted task sets, especially when considering task sets containing heavy tasks. We show the effectiveness of the proposed test through an extensive set of experiments. Marko Bertogna, Michele Cirinei, Giuseppe Lipari |
ECRTS | 1 |
| 2005 | New Schedulability Tests for Real-Time Task Sets Scheduled by Deadline Monotonic on Multiprocessors
Marko Bertogna, Michele Cirinei, Giuseppe Lipari |
OPODIS | 1 |