EDBT 2026 Demo / reviewers in the wild / expert
Rodolfo Pellizzoni
dblp:25/3790
· DBLP profile ↗
101ranked-venue papers
16as first author
19since 2021 · last 2025
0000-0002-7331-804XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 64 · 10 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 2 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DAMA: A Dual Arbitration Mechanism for Mixed-Criticality Applications
Wafic Lawand, Rodolfo Pellizzoni |
ECRTS | 2 |
| 2025 | Multi-Objective Memory Bandwidth Regulation and Cache Partitioning for Multicore Real-Time Systems
Binqi Sun, Zhihang Wei, Andrea Bastoni, Debayan Roy, Mirco Theile, Tomasz Kloda, Rodolfo Pellizzoni, Marco Caccamo |
ECRTS | 7 |
| 2025 | ParRP: Enabling Space Isolation in Caches with Shared DataabstractThis work presents an approach to isolate cache space while supporting shared data. Enabling shared data caching is challenging because it causes interference making it unsuitable for real-time multicores. Unlike prior works, we aim to introduce isolation, but our approach enables caching of shared data, and promotes isolated cache analysis for individual cores. The crux behind our approach is that shared data isolation can be achieved by partitioning the replacement information instead of the cache's data storage. Consequently, this work introduces ParRP, a novel hardware cache partitioning scheme for realtime cache-coherent multicores. Our evaluation using the gem5 simulator shows that, by providing isolation for shared data, the worst-case execution time of multi-threaded tasks can be lower by 2.4× at the cost of a 16.5% decrease in average-case performance. Rodolfo Pellizzoni, Hiren Patel |
RTAS | 2 |
| 2025 | OpenDRAM: A Modular, High-performance Soft Memory Controller for DDR4 DRAMabstractWe propose OpenDRAM , a synthesizable high-performance DDR4 DRAM soft Memory Controller (MC) for FPGAs. Since DRAMs usually operate at a higher frequency compared to MCs (usually \(4\times\) ), to fully utilize DRAM’s bandwidth, the hardened DDR4 physical interface expects the controller to issue four DRAM commands in a single clock cycle. OpenDRAM is a modular, extensible MC, implementing high-performance bank-parallel schedulers. We detail the design of OpenDRAM ’s logic blocks in RTL and their integration with existing AMD’s Memory Interface Generator (MIG) modules for initialization, maintenance, and interfacing. The integrated project was comprehensively validated on an AMD Virtex UltraScale+ FPGA. We evaluate and compare the performance of OpenDRAM with AMD’s MIG controller and another open source controller, OPRECOMP, using synthetic and accelerator kernels. Results show that OpenDRAM surpasses both commercial and open source counterparts, offering performance improvements of up to 157% over AMD’s MIG and 267% over OPRECOMP, primarily owing to its reordering and scheduling mechanisms. To demonstrate its research use case, we prototype five distinct command schedulers, exploring tradeoffs between scheduling aggressiveness and maximum frequency, and show how FPGA-aware design can enhance timing closure. Finally, we release OpenDRAM as the first high-performance, extensible, open source MC for researchers to utilize, extend, and build upon. Danesh Germchi, Amin Katani, Mohamed Hassan 0002, Rodolfo Pellizzoni |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2024 | Exclusive Hierarchies for Predictable Sharing in Last-Level CacheabstractThis work presents an approach to use a last-level cache (LLC) in a memory hierarchy for cache-coherent real-time multicores that delivers a low worst-case latency (WCL) and higher performance than all of its counterparts. Our approach relies on the key observation that an exclusive memory hierarchy, by definition, eliminates back invalidations, which are one of the largest contributors to the WCL when using inclusive memory hierarchies. However, to the best of our knowledge, there are no prior efforts that ensure the predictability of exclusive hierarchies for cache-coherent multicores. Consequently, in this work, we propose PECC, a predictable exclusive cache coherence mechanism, that achieves a lower average data access latency while providing a low WCL bound that scales linearly in the number of cores. Our evaluation shows that PECC reduces the bound by 6% and improves the average performance by 2.33× over the predictable solution with an inclusive LLC. Zhuanhao Wu, Rodolfo Pellizzoni, Hiren D. Patel |
RTAS | 3 |
| 2024 | Duration-based Instruction Cache LockingabstractCache locking is a commonly used mechanism to improve both performance and predictability for embedded programs. Dynamic cache locking methods proposed in the literature, where the locked content is modified during execution, require inserting locking and unlocking instructions in the program’s code. In this paper, we introduce a novel hardware mechanism that leverages the LRU age bits to perform duration-based locking. Our proposed mechanism dynamically locks and unlocks cache lines for different durations at run-time, without the need to modify the program’s code. We further devise a heuristic that analyzes a program’s loop structure and selects the set of addresses to be locked in a L1 instruction cache alongside their locking durations. Evaluation results show that our duration-based locking mechanism achieves comparable results to the dynamic approach while substantially reducing the initialization overhead and avoiding program code modifications. Wafic Lawand, Rodolfo Pellizzoni |
RTCSA | 2 |
| 2023 | A Tight Holistic Memory Latency Bound Through Coordinated Management of Memory Resources
Shorouk Abdelhalim, Danesh Germchi, Mohamed Hossam, Rodolfo Pellizzoni, Mohamed Hassan 0002 |
ECRTS | 4 |
| 2023 | Co-Optimizing Cache Partitioning and Multi-Core Task Scheduling: Exploit Cache Sensitivity or Not?abstractCache partitioning techniques have been successfully adopted to mitigate interference among concurrently executing real-time tasks on multi-core processors. Considering that the execution time of a cache-sensitive task strongly depends on the cache available for it to use, co-optimizing cache partitioning and task allocation improves the system's schedulability. In this paper, we propose a hybrid multi-layer design space exploration technique to solve this multi-resource management problem. We explore the interplay between cache partitioning and schedulability by systematically interleaving three optimization layers, viz., (i) in the outer layer, we perform a breadth-first search combined with proactive pruning for cache partitioning; (ii) in the middle layer, we exploit a first-fit heuristic for allocating tasks to cores; and (iii) in the inner layer, we use the well-known recurrence relation for the schedulability analysis of non-preemptive fixed-priority (NP-FP) tasks in a uniprocessor setting. Although our focus is on NP-FP scheduling, we evaluate the flexibility of our framework in supporting different scheduling policies (NP-EDF, P-EDF) by plugging in appropriate analysis methods in the inner layer. Experiments show that, compared to the state-of-the-art techniques, the proposed framework can improve the real-time schedulability of NP-FP task sets by an average of 15.2% with a maximum improvement of 233.6% (when tasks are highly cache-sensitive) and a minimum of 1.6% (when cache sensitivity is low). For such task sets, we found that clustering similar- period (or mutually compatible) tasks often leads to higher schedulability (on average 7.6 %) than clustering by cache sensitivity. In our evaluation, the framework also achieves good results for preemptive and dynamic-priority scheduling policies. Binqi Sun, Debayan Roy, Tomasz Kloda, Andrea Bastoni, Rodolfo Pellizzoni, Marco Caccamo |
RTSS | 5 |
| 2023 | X-Stream: Accelerating streaming segments on MPSoCs for real-time applications
Rohan Tabish, Rodolfo Pellizzoni, Renato Mancuso 0001, Giovani Gracioli, Reza Mirosanlou, Marco Caccamo |
J. Syst. Archit. | 2 |
| 2023 | Lazy Load Scheduling for Mixed-criticality Applications in Heterogeneous MPSoCsabstractNewly emerging multiprocessor system-on-a-chip (MPSoC) platforms provide hard processing cores with programmable logic (PL) for high-performance computing applications. In this article, we take a deep look into these commercially available heterogeneous platforms and show how to design mixed-criticality applications such that different processing components can be isolated to avoid contention on the shared resources such as last-level cache and main memory. Our approach involves software/hardware co-design to achieve isolation between the different criticality domains. At the hardware level, we use a scratchpad memory (SPM) with dedicated interfaces inside the PL to avoid conflicts in the main memory. At the software level, we employ a hypervisor to support cache-coloring such that conflicts at the shared L2 cache can be avoided. In order to move the tasks in/out of the SPM memory, we rely on a DMA engine and propose a new CPU-DMA co-scheduling policy, called Lazy Load , for which we also derive the response time analysis. The results of a case study on image processing demonstrate that the contention on the shared memory subsystem can be avoided when running with our proposed architecture. Moreover, comprehensive schedulability evaluations show that the newly proposed Lazy Load policy outperforms the existing CPU-DMA scheduling approaches and is effective in mitigating the main memory interference in our proposed architecture. Tomasz Kloda, Giovani Gracioli, Rohan Tabish, Reza Mirosanlou, Renato Mancuso 0001, Rodolfo Pellizzoni, Marco Caccamo |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2022 | Optimizing parallel PREM compilation over nested loop structuresabstractWe consider automatic parallelization of a computational kernel executed according to the PRedictable Execution Model (PREM), where each thread is divided into execution and memory phases. We target a scratchpad-based architecture, where memory phases are executed by a dedicated DMA component. We employ data analysis and loop tiling to split the kernel execution into segments, and schedule them based on a DAG representation of data and execution dependencies. Our main observation is that properly selecting tile sizes is key to optimize the makespan of the kernel. We thus propose a heuristic that efficiently searches for optimized tile size and core assignments over deeply nested loops, and demonstrate its applicability and performance compared to the state-of-the-art in PREM compilation using the PolyBench-NN benchmark suite. Zhao Gu, Rodolfo Pellizzoni |
DAC | 2 |
| 2022 | Parallelism-Aware High-Performance Cache Coherence with Tight Latency Bounds
Reza Mirosanlou, Mohamed Hassan 0002, Rodolfo Pellizzoni |
ECRTS | 3 |
| 2022 | Beyond Just Safety: Delay-aware Security Monitoring for Real-time Control SystemsabstractModern embedded real-time systems (RTS) are increasingly facing more security threats than the past. A simplistic straightforward integration of security mechanisms might not be able to guarantee thesafetyand predictability of such systems. In this article, we focus on integrating security mechanisms into RTS (especiallylegacyRTS). We introduceContego-C, an analytical model to integrate security tasks into RTS that will allow system designers to improve the security posture without affecting temporal and control constraints of the existing real-time control tasks. We also define ametric(named tightness of periodic monitoring) to measure the effectiveness of such integration. We demonstrate our ideas using a proof-of-concept implementation on an ARM-based rover platform and show that Contego-C can improve security without degrading control performance. Monowar Hasan, Sibin Mohan, Rakesh Bobba, Rodolfo Pellizzoni |
ACM Trans. Cyber Phys. Syst. | 4 |
| 2022 | HopliteML: Evolving Application Customized FPGA NoCs with Adaptable Routers and RegulatorsabstractWe can overcome the pessimism in worst-case routing latency analysis of timing-predictable Network-on-Chip (NoC) workloads by single-digit factors through the use of a hybrid field-programmable gate array (FPGA)–optimized NoC and workload-adapted regulation. Timing-predictable FPGA-optimized NoCs such as HopliteBuf integrate stall-free FIFOs that are sized using offline static analysis of a user-supplied flow pattern and rates. For certain bursty traffic and flow configurations, static analysis delivers very large, sometimes infeasible, FIFO size bounds and large worst-case latency bounds. Alternatively, backpressure-based NoCs such as HopliteBP can operate with lower latencies for certain bursty flows. However, they suffer from severe pessimism in the analysis due to the effect of pipelining of packets and interleaving of flows at switch ports. As we show in this article, a hybrid FPGA NoC that seamlessly composes both design styles on a per-switch basis delivers the best of both worlds, with improved feasibility (bounded operation) and tighter latency bounds. We select the NoC switch configuration through a novel evolutionary algorithm based on Maximum Likelihood Estimation (MLE). For synthetic ( RANDOM , LOCAL ) and real-world ( SpMV , Graph ) workloads, we demonstrate ≈2–3× improvements in feasibility and ≈1–6.8× in worst-case latency while requiring an LUT cost only ≈1–1.5× larger than the cheapest HopliteBuf solution. We also deploy and verify our NoC (PL) and MLE framework (PS) on a Pynq-Z1 to adapt and reconfigure NoC switches dynamically. We can further improve a workload’s routability by learning to surgically tune regulation rates for each traffic trace to maximize available routing bandwidth. We capture critical dependency between traces by modelling the regulation space as a multivariate Gaussian distribution and learn the distribution’s parameters using Covariance Matrix Adaptation Evolution Strategy (CMA-ES). We also propose nested learning, which learns switch configurations and regulation rates in tandem. Compared with stand-alone switch learning, this symbiotic nested learning helps achieve ≈ 1.5× lower cost constrained latency, ≈ 3.1× faster individual rates, and ≈ 1.4× faster mean rates. We also evaluate improvements to vanilla NoCs’ routing using only stand-alone rate learning (no switch learning), with ≈ 1.6× lower latency across synthetic and real-world benchmarks. Gurshaant Singh Malik, Ian Elmor Lang, Rodolfo Pellizzoni, Nachiket Kapre |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2021 | Virtual Gang Scheduling of Parallel Real-Time TasksabstractWe consider the problem of executing parallel realtime tasks according to gang scheduling on a multicore system in the presence of shared resource interference. Specifically, we consider sets of gang-tasks with precedence constraints in the form of a DAG. We introduce the novel concept of a virtual gang: a group of parallel tasks that are scheduled together as a single entity. Employing virtual gangs allows us to tightly bound the effect of shared resource interference. It also transforms the original, complex scheduling problem into a form that can be easily implemented and is amenable to exact schedulability analysis, further reducing pessimism. We present and evaluate both optimal and heuristic methods for forming virtual gangs based on a known interference model and while respecting all precedence constraints among tasks. When precedence constraints are not considered, we also compare our approach against existing response-time analysis for globally scheduled gang-tasks, as well as general parallel tasks. The results show that our approach significantly outperforms state-of-the-art multicore schedulability analyses when shared-resource interference is considered. Even in the absence of interference, it performs better than the state-of-the-art for highly parallel tasksets. Rodolfo Pellizzoni, Heechul Yun |
DATE | 2 |
| 2021 | Duetto: Latency Guarantees at Minimal Performance CostabstractThe management of shared hardware resources in multi-core platforms has been characterized by a fundamental trade-off: high-performance arbiters typically employed in COTS systems offer no worst-case guarantees, while dedicated real-time controllers provide timing guarantees at the cost of significantly degrading system performance. In this paper, we overcome this trade-off by introducing Duetto, a novel hardware resource management paradigm. Duetto pairs a real-time arbiter with a high-performance arbiter and a latency estimator module. Based on the observation that the resource is rarely overloaded, Duetto executes the high-performance arbiter most of the time, switching to the real-time arbiter only in the rare cases when the latency estimator deems that timing guarantees risk being violated. We demonstrate our approach on the case study of a multi-bank memory. Our evaluation based on cycle-accurate simulations shows that Duetto can provide the same latency guarantees as the real-time arbiter with limited loss of performance compared to the high-performance arbiter. Reza Mirosanlou, Mohamed Hassan 0002, Rodolfo Pellizzoni |
DATE | 3 |
| 2021 | Worst-case latency analysis for the versal NoC network packet switchabstractThe recent line of Versal FPGA devices from Xilinx Inc. includes a hard Network-On-Chip (NoC) embedded in the programmable logic, designed to be a high-performance system-level interconnect. While the target markets for Versal devices include applications with real-time constraints, such as automotive driver assist, the associated development tools only provide figures for "structural latencies" of data packets, which assume that the network is otherwise idle. In a realistic setting, this information is not enough to ensure deadlines are met, as different packets can contend for NoC switch outputs, which causes packet contents to be buffered while in transit, increasing their latency. In this work, we present a formal description of the NPS switches that compose the Versal NoC from a flit (or packet) scheduling perspective, based on the available cycle-accurate switch simulation code. We then analyze a scenario where network clients transfer data periodically over a single switch, and propose a method for calculating worst-case communication times in this scenario. Ian Elmor Lang, Nachiket Kapre, Rodolfo Pellizzoni |
NOCS | 3 |
| 2021 | A Real-Time Virtio-Based Framework for Predictable Inter-VM CommunicationabstractEnsuring real-time properties on current heterogeneous multiprocessor systems on a chip is a challenging task. Furthermore, online artificial intelligent applications –which are routinely deployed on such chips– pose increasing pressure on the memory subsystem that becomes a source of unpredictability. Although techniques have been proposed to restore independent access to memory for concurrently executing virtual machines (VM), providing predictable inter-VM communication remains challenging. In this work, we tackle the problem of predictably transferring data between virtual machines and virtualized hardware resources on multiprocessor systems on chips under consideration of memory interference. We design a "broker-based" real-time communication framework for otherwise isolated virtual machines, provide a virtio-based reference implementation on top of the Jailhouse hypervisor, assess its overheads for FreeRTOS virtual machines, and formally analyze its communication flow schedulability under consideration of the implementation overheads. Furthermore, we define a methodology to assess the maximum DRAM memory saturation empirically, evaluate the framework’s performance and compare it with the theoretical schedulability. Gero Schwäricke, Rohan Tabish, Rodolfo Pellizzoni, Renato Mancuso 0001, Andrea Bastoni, Alexander Züpke, Marco Caccamo |
RTSS | 3 |
| 2021 | An Analyzable Inter-core Communication Framework for High-Performance Multicore Embedded Systems
Rohan Tabish, Jen-Yang Wen, Rodolfo Pellizzoni, Renato Mancuso 0001, Heechul Yun, Marco Caccamo, Lui Sha |
J. Syst. Archit. | 3 |
| 2020 | Period Adaptation for Continuous Security Monitoring in Multicore Real-Time SystemsabstractWe propose HYDRA-C, a design-time evaluation framework for integrating monitoring mechanisms in multicore real-time systems (RTS). Our goal is to ensure that security (or other monitoring) mechanisms execute in a "continuous" manner - i.e., as often as possible, across cores. This is to ensure that any such mechanisms run with few interruptions, if any. HYDRA-C is intended to allow designers of RTS to integrate monitoring mechanisms without perturbing existing timing properties or execution orders. We demonstrate the framework using a proofof-concept implementation with intrusion detection mechanisms as security tasks. We develop and use both, (a) a custom intrusion detection system (IDS) as well as (b) Tripwire - an open source data integrity checking tool. We compare the performance of HYDRA-C with a state-of-the-art multicore RT security integration approach and find that our method does not impact the schedulability and, on average, can detect intrusions 19.05% faster without impacting the performance of RT tasks. Monowar Hasan, Sibin Mohan, Rodolfo Pellizzoni, Rakesh Bobba |
DATE | 3 |
| 2020 | Analysis of Memory-Contention in Heterogeneous COTS MPSoCsabstractMultiple-Processors Systems-on-Chip (MPSoCs) provide an appealing platform to execute Mixed Criticality Systems (MCS) with both time-sensitive critical tasks and performance-oriented non-critical tasks. Their heterogeneity with a variety of processing elements can address the conflicting requirements of those tasks. Nonetheless, the complex (and hence hard-to-analyze) architecture of Commercial-Off-The-Shelf (COTS) MPSoCs presents a challenge encumbering their adoption for MCS. In this paper, we propose a framework to analyze the memory contention in COTS MPSoCs and provide safe and tight bounds to the delays suffered by any critical task due to this contention. Unlike existing analyses, our solution is based on two main novel approaches. 1) It conducts a hybrid analysis that blends both request-level and task-level analyses into the same framework. 2) It leverages available knowledge about the types of memory requests of the task under analysis as well as contending tasks; specifically, we consider information that is already obtainable by applying existing static analysis tools to each task in isolation. Thanks to these novel techniques, our comparisons with the state-of-the art approaches show that the proposed analysis provides the tightest bounds across all evaluated access scenarios. Mohamed Hassan 0002, Rodolfo Pellizzoni |
ECRTS | 2 |
| 2020 | Learn the Switches: Evolving FPGA NoCs with Stall-Free and Backpressure Based RoutersabstractWe can overcome the pessimism in worst-case routing latency analysis of timing-predictable Network-on-Chip (NoC) workloads by single digit factors through the use of a hybrid FPGA-optimized NoC. Timing-predictable FPGA-optimized NoCs such as HopliteBuf integrate stall-free FIFOs that are sized using offline, static analysis of a user-supplied flow pattern and rates. For certain bursty traffic and flow configurations, the static analysis delivers very large, sometimes infeasible, FIFO size bounds and large worst-case latency bounds. Alternatively, backpressure-based NoCs such as HopliteBP can operate with lower latencies for certain bursty flows. However, they suffer from severe pessimism in the analysis due to the effect of pipelining of packets and interleaving of flows at switch ports. As we show in this paper, a hybrid FPGA NoC that seamlessly composes both design styles on a per-switch basis, delivers the best of both worlds with improved feasibility (bounded operation), and tighter latency bounds. We select the NoC switch configuration though a novel evolutionary algorithm based on Maximum Likelihood Estimation (MLE). For synthetic (RANDOM, LOCAL) and real world (SpMV, Graph) workloads, we demonstrate ≈2– 3× improvements in feasibility, ≈1–6.8× in worst-case latency while only requiring LUT cost ≈1–1.5× larger than the cheapest HopliteBuf solution. We also deploy and verify our NoC (PL) and MLE framework (PS) on a Pynq-Z1 to adapt and reconfigure NoC switches dynamically. Gurshaant Singh Malik, Ian Elmor Lang, Rodolfo Pellizzoni, Nachiket Kapre |
FPL | 3 |
| 2020 | DRAMbulism: Balancing Performance and Predictability through Dynamic PipeliningabstractWorst-case execution bounds for real-time programs are profoundly impacted by the latency of accessing hardware shared resources, such as off-chip DRAM. While many different memory controller designs have been proposed in the literature, there is a trade-off between average-case performance and predictable worst-case bounds, as techniques targeted at improving the former can harm the latter and vice-versa. We find that taking advantage of pipelining between different commands can improve both, but incorporating pipelining effects in worst-case analysis is challenging. In this work, we introduce a novel DRAM controller that successfully balances performance and predictability by employing a dynamic pipelining scheme. We show that the schedule of DRAM commands is akin to a two-stage two-mode pipeline, and hence, design an easily-implementable admission rule that allows us to dynamically add requests to the pipeline without hurting worst-case bounds. Reza Mirosanlou, Mohamed Hassan 0002, Rodolfo Pellizzoni |
RTAS | 3 |
| 2020 | Dynamic Memory Bandwidth Allocation for Real-Time GPU-Based SoC PlatformsabstractHeterogeneous SoC platforms, comprising both general purpose CPUs and accelerators, such as a GPU, are becoming increasingly attractive for real-time and mixed-criticality systems to cope with the computational demand of data parallel applications. However, contention for access to shared main memory can lead to significant performance degradation on both CPU and GPU. Existing work has shown that memory bandwidth throttling is effective in protecting real-time applications from memory-intensive, best-effort (BE) ones; however, due to the inherent pessimism involved in worst-case execution time (WCET) estimation, such approaches can unduly restrict the bandwidth available to BE applications. In this article, we propose a novel memory bandwidth allocation scheme where we dynamically monitor the progress of a real-time application and increase the bandwidth share of BE ones whenever it is safe to do so. Specifically, we demonstrate our approach by protecting a real-time GPU kernel from BE CPU tasks. Based on profiling information, we first build a WCET estimation model for the GPU kernel. Using such model, we then show how to dynamically recompute on-line the maximum memory budget that can be allocated to BE tasks without exceeding the kernel's assigned execution budget. We implement our proposed technique on NVIDIA embedded SoC and demonstrate its effectiveness on a variety of GPU and CPU benchmarks. Homa Aghilinasab, Heechul Yun, Rodolfo Pellizzoni |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | HopliteBuf: Network Calculus-Based Design of FPGA NoCs with Provably Stall-Free FIFOsabstractHopliteBuf is a deflection-free, low-cost, and high-speed FPGA overlay Network-on-chip (NoC) with stall-free buffers. It is an FPGA-friendly 2D unidirectional torus topology built on top of HopliteRT overlay NoC. The stall-free buffers in HopliteBuf are supported by static analysis tools based on network calculus that help determine worst-case FIFO occupancy bounds for a prescribed workload. We implement these FIFOs using cheap LUT SRAMs (Xilinx SRL32s and Intel MLABs) to reduce cost. HopliteBuf is a hybrid microarchitecture that combines the performance benefits of conventional buffered NoCs by using stall-free buffers with the cost advantages of deflection-routed NoCs by retaining the lightweight unidirectional torus topology structure. We present two design variants of the HopliteBuf NoC: (1) single corner-turn FIFO ( W → S ) and (2) dual corner-turn FIFO ( W → S + N ). The single corner-turn ( W → S ) design is simpler and only introduces a buffering requirement for packets changing dimension from the X ring to the downhill Y ring (or West to South). The dual corner-turn variant requires two FIFOs for turning packets going downhill ( W → S ) as well as uphill ( W → N ). The dual corner-turn design overcomes the mathematical analysis challenges associated with single corner-turn designs for communication workloads with cyclic dependencies between flow traversal paths at the expense of a small increase in resource cost. Our static analysis delivers bounds that are not only better (in latency) than HopliteRT but also tighter by 2−3×. Across 100 randomly generated flowsets mapped to a 5×5 system size, HopliteBuf is able to route a larger fraction of these flowsets with <128-deep FIFOs, boost worst-case routing latency by ≈ 2× for mutually feasible flowsets, and support a 10% higher injection rate than HopliteRT. At 20% injection rates, HopliteRT is only able to route 1--2% of the flowsets, while HopliteBuf can deliver 40--50% sustainability. When compared to the W → S bkp backpressure-based router, we observe that our HopliteBuf solution offers 25--30% better feasibility at 30--40% lower LUT cost. Tushar Garg, Saud Wasly, Rodolfo Pellizzoni, Nachiket Kapre |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2019 | Designing Mixed Criticality Applications on Modern Heterogeneous MPSoC PlatformsabstractMultiprocessor Systems-on-Chip (MPSoC) integrating hard processing cores with programmable logic (PL) are becoming increasingly common. While these platforms have been originally designed for high performance computing applications, their rich feature set can be exploited to efficiently implement mixed criticality domains serving both critical hard real-time tasks, as well as soft real-time tasks. In this paper, we take a deep look at commercially available heterogeneous MPSoCs that incorporate PL and a multicore processor. We show how one can tailor these processors to support a mixed criticality system, where cores are strictly isolated to avoid contention on shared resources such as Last-Level Cache (LLC) and main memory. In order to avoid conflicts in last-level cache, we propose the use of cache coloring, implemented in the Jailhouse hypervisor. In addition, we employ ScratchPad Memory (SPM) inside the PL to support a multi-phase execution model for real-time tasks that avoids conflicts in shared memory. We provide a full-stack, working implementation on a latest-generation MPSoC platform, and show results based on both a set of data intensive tasks, as well as a case study based on an image processing benchmark application. Giovani Gracioli, Rohan Tabish, Renato Mancuso 0001, Reza Mirosanlou, Rodolfo Pellizzoni, Marco Caccamo |
ECRTS | 5 |
| 2019 | PREM-Based Optimal Task Segmentation Under Fixed Priority SchedulingabstractRecently, a large number of works have discussed scheduling tasks consisting of a sequence of memory phases, where code and data are moved between main memory and local memory, and computation phases, where the task executes based on the content of local memory only; the key idea is to prevent main memory contention by scheduling the memory phase of one task in parallel with computation phases of tasks running on other cores. This paper provides two main contributions: (1) we present a compiler-level tool, based on the LLVM intermediate representation, that automatically converts a program into a conditional sequence of segments comprising memory and computation phases; (2) we propose an algorithm to find optimal segmentation decisions for a task set scheduled according to a fixed-priority partitioned scheme. Our evaluation shows that the proposed framework can be feasibly applied to realistic programs, and vastly overperforms a baseline greedy approach. Muhammad Refaat Soliman, Rodolfo Pellizzoni |
ECRTS | 2 |
| 2019 | HopliteBuf: FPGA NoCs with Provably Stall-Free FIFOsabstractDeflection-routed NoCs like Hoplite and HopliteRT take advantage of FPGA-specific features to deliver low-cost, high-frequency, FPGA-friendly communication networks. However, they suffer from long packet deflection penalties, low sustained throughputs, and feature limitations such as out-of-order delivery of packets. In this paper, we introduce the HopliteBuf NoC, and an associated static analysis tool, that eliminates deflections entirely while simultaneously adding in-order delivery feature using (1) small, stall-free FIFOs with provable occupancy bounds, and (2) linearization of vertical rings of the torus Hoplite topology to improve provable link utilization. We implement these FIFOs using cheap LUT SRAMs (Xilinx SRL32s, and Intel MLABs) to absorb packet contention. We evaluate conditions for stall-free behavior using static analysis that compute upper bounds on FIFO occupancy based on the communication pattern. Our static analysis deliver bounds that are not only better (in latency) than HopliteRT but also tighter by 2--3×. Across 100 randomly-generated flowsets mapped to a 5×5 system size, HopliteBuf is able to route a larger fraction of these flowsets with \textless 128-deep FIFOs, boost worst-case routing latency by $\approx$2× for mutually feasible flowsets. At 20% injection rates, HopliteRT is only able to route 1--2% of the flowsets while HopliteBuf can deliver 40--50% sustainability. Tushar Garg, Saud Wasly, Rodolfo Pellizzoni, Nachiket Kapre |
FPGA | 3 |
| 2019 | A Novel Side-Channel in Real-Time SchedulersabstractWe demonstrate the presence of a novel scheduler side-channel in preemptive, fixed-priority real-time systems (RTS); examples of such systems can be found in automotive systems, avionic systems, power plants and industrial control systems among others. This side-channel can leak important timing information such as the future arrival times of real-time tasks. This information can then be used to launch devastating attacks, two of which are demonstrated here (on real hardware platforms). Note that it is not easy to capture this timing information due to runtime variations in the schedules, the presence of multiple other tasks in the system and the typical constraints (e.g., deadlines) in the design of RTS. Our ScheduLeak algorithms demonstrate how to effectively exploit this side-channel. A complete implementation is presented on real operating systems (in Real-time Linux and FreeRTOS). Timing information leaked by ScheduLeak can significantly aid other, more advanced, attacks in better accomplishing their goals. Chien-Ying Chen, Sibin Mohan, Rodolfo Pellizzoni, Rakesh Bobba, Negar Kiyavash |
RTAS | 3 |
| 2019 | Bundled Scheduling of Parallel Real-Time TasksabstractDue to increasing demand for computational power, parallel processing is becoming important for real-time systems. The community has proposed and analyzed several task models and scheduling schemes for parallel applications. Gang scheduling, where all threads of the same task are required to run concurrently, can lead to significant execution time improvements for tightly-synchronized applications. However, existing work on real-time gang scheduling for the rigid model assumes that the number of threads of a parallel task remains constant during its execution, which is not the case for many applications. In this paper, we introduce a novel task model where each task is divided into a sequence of segments called bundles, each requiring a different number of cores, and the threads within each bundle are gang scheduled on an identical multiprocessor. We then derive a sufficient schedulability analysis for sporadic bundled tasks under fixed-priority scheduling. Our evaluation reveals that the proposed model can significantly improve schedulability compared to the rigid model while being competitive with non-gang global thread scheduling and federated scheduling. Saud Wasly, Rodolfo Pellizzoni |
RTAS | 2 |
| 2019 | Segment Streaming for the Three-Phase Execution Model: Design and ImplementationabstractScheduling tasks using the three-phase execution model (load-execute-unload) can effectively reduce the contention on shared resources in real-time systems. Due to system and program constraints, a task is generally segmented and executed over multiple intervals. Several works showed that co-scheduling memory (unload-load) and computation phases can improve the system schedulability by hiding the memory transfer time. However, this is limited to segments of different tasks and hence executing segments of the same task back-to-back is not allowed. In this paper, we propose a new streaming model to allow overlapping the memory and execution phases of segments of the same task. This is accomplished by a segmentation framework implemented within an LLVM-based compiler-level tool along with a Real-Time Operating System (RTOS) API to handle load/unload requests. Memory phases are processed by a DMA engine that loads/unloads the task content into ScratchPad Memory (SPM). We provide a schedulability analysis of the proposed model under fixed priority partitioned scheme and an RTOS implementation of the API on a latest-generation Multiprocessor System-on-Chip (MPSoC). Muhammad Refaat Soliman, Giovani Gracioli, Rohan Tabish, Rodolfo Pellizzoni, Marco Caccamo |
RTSS | 4 |
| 2019 | A real-time scratchpad-centric OS with predictable inter/intra-core communication for multi-core embedded systems
Rohan Tabish, Renato Mancuso 0001, Saud Wasly, Rodolfo Pellizzoni, Marco Caccamo |
Real Time Syst. | 4 |
| 2018 | A design-space exploration for allocating security tasks in multicore real-time systemsabstractThe increased capabilities of modern real-time systems (RTS) expose them to various security threats. Recently, frameworks that integrate security tasks without perturbing the real-time tasks have been proposed, but they only target single core systems. However, modern RTS are migrating towards multicore platforms. This makes the problem of integrating security mechanisms more complex, as designers now have multiple choices for where to allocate the security tasks. In this paper we propose HYDRA, a design space exploration algorithm that finds an allocation of security tasks for multicore RTS using the concept of opportunistic execution. HYDRA allows security tasks to operate with existing real-time tasks without perturbing system parameters or normal execution patterns, while still meeting the desired monitoring frequency for intrusion detection. Our evaluation uses a representative real-time control system (along with synthetic task sets for a broader exploration) to illustrate the efficacy of HYDRA. Monowar Hasan, Sibin Mohan, Rodolfo Pellizzoni, Rakesh Bobba |
DATE | 3 |
| 2018 | Analysis of Dynamic Memory Bandwidth Regulation in Multi-core Real-Time SystemsabstractOne of the primary sources of unpredictability in modern multi-core embedded systems is contention over shared memory resources, such as caches, interconnects, and DRAM. Despite significant achievements in the design and analysis of multi-core systems, there is a need for a theoretical framework that can be used to reason on the worst-case behavior of real-time workload when both processors and memory resources are subject to scheduling decisions. In this paper, we focus our attention on dynamic allocation of main memory bandwidth. In particular, we study how to determine the worst-case response time of tasks spanning through a sequence of time intervals, each with a different bandwidth-to-core assignment. We show that the response time computation can be reduced to a maximization problem over assignment of memory requests to different time intervals, and we provide an efficient way to solve such problem. As a case study, we then demonstrate how our proposed analysis can be used to improve the schedulability of Integrated Modular Avionics systems in the presence of memory-intensive workload. Ankit Agrawal 0005, Renato Mancuso 0001, Rodolfo Pellizzoni, Gerhard Fohler |
RTSS | 3 |
| 2018 | Bounding DRAM Interference in COTS Heterogeneous MPSoCs for Mixed Criticality SystemsabstractCommercial off-the-shelf (COTS) heterogeneous multiple processors systems-on-chip (MPSoCs) are appealing platforms for emerging mixed criticality systems (MCSs). To satisfy MCS requirements, the platform must guarantee predictable timing bounds for critical applications, without degrading average performance for noncritical applications. In particular, this paper studies the main memory subsystem, which in modern MPSoCs is typically based on double data rate synchronous dynamic access memory. While there exists previous work on worst-case DRAM latency analysis, such work only covers a small subset of possible COTS configurations, which are not targeted at MCS. Therefore, we derive a generalized interference delay analysis for DRAM main memory that accounts for a breadth of features deployed in COTS platforms. We then explore the design space by studying the effects of each feature on both the worst-case delay for critical applications, and the bandwidth for noncritical applications. Mohamed Hassan 0002, Rodolfo Pellizzoni |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | A Comparative Study of Predictable DRAM ControllersabstractRecently, the research community has introduced several predictable dynamic random-access memory (DRAM) controller designs that provide improved worst-case timing guarantees for real-time embedded systems. The proposed controllers significantly differ in terms of arbitration, configuration, and simulation environment, making it difficult to assess the contribution of each approach. To bridge this gap, this article provides the first comprehensive evaluation of state-of-the-art predictable DRAM controllers. We propose a categorization of available controllers, and introduce an analytical performance model based on worst-case latency. We then conduct an extensive evaluation for all state-of-the-art controllers based on a common simulation platform, and discuss findings and recommendations. Danlu Guo, Mohamed Hassan 0002, Rodolfo Pellizzoni, Hiren D. Patel |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2017 | Modeling the Effects of AUTOSAR Overheads on Application Timing and SchedulabilityabstractAUTOSAR (AUTomotive Open System ARchitecture) provides an open and standardized E/E architecture for automobiles. AUTOSAR systems exhibit real-time requirements, i.e., an AUTOSAR application must always be schedulable. In this paper, we propose an overhead-aware method to find schedulable design configurations for an AUTOSAR application. We show how to construct a timing model for the application, discuss how to quantify the overheads of an AUTOSAR stack implementation, and assess their impact on timing and schedulability. We demonstrate the proposed method on an automotive case study and evaluate the effects of different types of overheads using synthetic applications. Manish Chauhan, Rodolfo Pellizzoni, Krzysztof Czarnecki 0001 |
DAC | 2 |
| 2017 | WCET Derivation under Single Core Equivalence with Explicit Memory Budget AssignmentabstractIn the last decade there has been a steady uptrend in the popularity of embedded multi-core platforms. This represents a turning point in the theory and implementation of real-time systems. From a real-time standpoint, however, the extensive sharing of hardware resources (e.g. caches, DRAM subsystem, I/O channels) represents a major source of unpredictability. Budget-based memory regulation (throttling) has been extensively studied to enforce a strict partitioning of the DRAM subsystem’s bandwidth. The common approach to analyze a task under memory bandwidth regulation is to consider the budget of the core where the task is executing, and assume the worst-case about the remaining cores' budgets. In this work, we propose a novel analysis strategy to derive the WCET of a task under memory bandwidth regulation that takes into account the exact distribution of memory budgets to cores. In this sense, the proposed analysis represents a generalization of approaches that consider (i) even budget distribution across cores; and (ii) uneven but unknown (except for the core under analysis) budget assignment. By exploiting the additional piece of information, we show that it is possible to derive a more accurate WCET estimation. Our evaluations highlight that the proposed technique can reduce overestimation by 30% in average, and up to 60%, compared to the state of the art. Renato Mancuso 0001, Rodolfo Pellizzoni, Neriman Tokcan, Marco Caccamo |
ECRTS | 2 |
| 2017 | Contego: An Adaptive Framework for Integrating Security Tasks in Real-Time SystemsabstractEmbedded real-time systems (RTS) are pervasive. Many modern RTS are exposed to unknown security flaws, and threats to RTS are growing in both number and sophistication. However, until recently, cyber-security considerations were an afterthought in the design of such systems. Any security mechanisms integrated into RTS must (a) co-exist with the real-time tasks in the system and (b) operate without impacting the timing and safety constraints of the control logic. We introduce Contego, an approach to integrating security tasks into RTS without affecting temporal requirements. Contego is specifically designed for legacy systems, viz., the real-time control systems in which major alterations of the system parameters for constituent tasks is not always feasible. Contego combines the concept of opportunistic execution with hierarchical scheduling to maintain compatibility with legacy systems while still providing flexibility by allowing security tasks to operate in different modes. We also define a metric to measure the effectiveness of such integration. We evaluate Contego using synthetic workloads as well as with an implementation on a realistic embedded platform (an open-source ARM CPU running real-time Linux). Monowar Hasan, Sibin Mohan, Rodolfo Pellizzoni, Rakesh Bobba |
ECRTS | 3 |
| 2017 | WCET-Driven Dynamic Data Scratchpad Management With Compiler-Directed PrefetchingabstractIn recent years, the real-time community has produced a variety of approaches targeted at managing on-chip memory (scratchpads and caches) in a predictable way. However, to obtain safe WCET bounds, such techniques generally assume that the processor is stalled while waiting to reload the content of the on-chip memory; hence, they are less effective at hiding main memory latency compared to speculation-based techniques, such as hardware prefetching, that are largely used in general-purpose systems. In this work, we introduce a novel compiler-directed prefetching scheme for scratchpad memory that effectively hides the latency of main memory accesses by overlapping data transfers with the program execution. We implement and test an automated program compilation and optimization flow within the LLVM framework, and we show how to obtain improved WCET bounds through static analysis. Muhammad Refaat Soliman, Rodolfo Pellizzoni |
ECRTS | 2 |
| 2017 | HopliteRT: An efficient FPGA NoC for real-time applicationsabstractOverlay NoCs, such as Hoplite, are cheap to implement on an FPGA but provide no bounds on worst-case routing latency of packets traversing the NoC due to deflection routing. In this paper, we show how to adapt Hoplite to enable calculation of precise upper bounds on routing latency by modifying the routing function to prioritize deflections, and by regulating the injection of packets to meet certain throughput and burstiness constraints. We provide an analytical model for computing end-to-end latency in the form of (1) in-flight time in the network Tf, and (2) waiting time at the source node Ts. To bound in-flight time in an m × m NoC, we modify the routing function and switching crossbar richness in the Hoplite router to deliver Tf= ΔΧ + ΔΥ + (AY × m) + 2 where AX and ΔΥ are differences of the source and destination address co-ordinates of the packet. To bound the waiting time at the source, we add a Token Bucket regulator with rate ρiand burstiness σifor each flow fiof node (x, y) to deliver (⌈1/ρi⌉ − 1) + Ts: Ts= ⌈σ(ΓCf)/1−σ(ΓCf);⌉ which depends on the regulator period 1/ρi, burstiness a and the rate ρ of all interfering flows ΓCf. A 64b implementation of our HopliteRT router requires ≈4% fewer LUTs, and similar number of FFs compared to the original Hoplite router. We also need two small counters at each client port for regulating injection. We evaluate our model and RTL implementation across synthetic traffic patterns and observe behavior that conforms with the analytical bounds. Saud Wasly, Rodolfo Pellizzoni, Nachiket Kapre |
FPT | 2 |
| 2017 | A Requests Bundling DRAM Controller for Mixed-Criticality SystemsabstractWe design a novel DRAM controller that bundles and executes memory requests of hard real-time applications in consecutive rounds based on their type to reduce read/write switching delay. At the same time, our controller provides a configurable, guaranteed bandwidth for soft real-time requests. We show that there is a fundamental trade-off between the latency guarantee for hard real-time requests and the bandwidth provided to soft requests. Finally, we compare our approach analytically and experimentally with the current state-of-the-art real-time memory controller for single-rank DRAM devices, which applies type reordering at the level of DRAM commands rather than requests. Our evaluation shows that for tasks exhibiting average row hit ratios, or for which computing a row hit guarantee might be difficult, our controller provides both smaller guaranteed latency and larger bandwidth. Danlu Guo, Rodolfo Pellizzoni |
RTAS | 2 |
| 2017 | A Reliable and Predictable Scratchpad-centric OS for Multi-core Embedded SystemsabstractThe reliable use of multi-core platforms for designing safety-critical systems still represents an open challenge. Recently, the FAA [1] has formally expressed its concern towards the use of multi-core systems in avionics. The sharing of hardware resources introduces non-trivial timing dependencies between logically independent components (e.g. cores); additionally, the increase in size of circuitry, memory resources, and transistor density makes these platforms more susceptible to transient memory (soft) errors. This work addresses the problem of memory soft errors and their recovery at an OS/platform level on commercial multi-core systems. Proposed strategy considers the schedulability impact of recovery procedures on hard real-time workloads. Finally, the implementation of a SPM-centric OS with the proposed OS-level strategies was performed by using a commercially available multi-core platform. The design has been validated and evaluated using a combination of synthetic and realistic (EEMBC) benchmarks. Rohan Tabish, Renato Mancuso 0001, Saud Wasly, Sujit S. Phatak, Rodolfo Pellizzoni, Marco Caccamo |
RTAS | 5 |
| 2017 | PMC: A Requirement-Aware DRAM Controller for Multicore Mixed Criticality SystemsabstractWe propose a novel approach to schedule memory requests in Mixed Criticality Systems (MCS). This approach supports an arbitrary number of criticality levels by enabling the MCS designer to specify memory requirements per task. It retains locality within large-size requests to satisfy memory requirements of all tasks. To achieve this target, we introduce a compact time-division-multiplexing scheduler, and a framework that constructs optimal schedules to manage requests to off-chip memory. We also present a static analysis that guarantees meeting requirements of all tasks. We compare the proposed controller against state-of-the-art memory controllers using both a case study and synthetic experiments. Mohamed Hassan 0002, Hiren D. Patel, Rodolfo Pellizzoni |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2016 | Trading Cores for Memory Bandwidth in Real-Time SystemsabstractFederated scheduling has been proposed for parallel tasks. In this scheduling scheme, each parallel task is a assigned a private set of cores whereas sequential tasks share the remaining set of cores. Since parallel tasks are assigned dedicated cores, they receive no interference from other tasks. However, multicore processors are commonly built with shared main memory. The memory bandwidth is limited and therefore subject to contention. Consequently, parallel tasks can interfere with each other through the shared main memory. In this paper, we propose a novel method that is memory-aware when assigning cores to tasks. Our experimental results show a significant advantage of our method with respect to memory- oblivious methods. Ahmed Alhammad, Rodolfo Pellizzoni |
RTAS | 2 |
| 2016 | Memory Servers for Multicore SystemsabstractIn multicore systems, tasks can be significantly delayed due to contention for access to shared physical resources. Due to the complexity of the underlying hardware arbitration, analyzing resources such as DRAM main memory is challenging. In particular, compositional delay bounds tend to be highly pessimistic, while non- compositional bounds are tighter but prevent independent subsystem development. To address this issue, in this paper we introduce a novel memory server for multicore real-time systems under partitioned fixed-priority scheduling. Similar to a hierarchical server, our memory server regulates the amount of resource (bandwidth) that a group of tasks is allowed to consume, but the server interface is modified to account for the properties of delay analysis. We show how to derive schedulability conditions for each server and for the system as a whole. Our technique can support multiple memory regulation implementations and delay analyses, while significantly improving system schedulability compared to the unregulated case. Rodolfo Pellizzoni, Heechul Yun |
RTAS | 1 |
| 2016 | A Real-Time Scratchpad-Centric OS for Multi-Core Embedded SystemsabstractMulti-core processors have replaced single-core systems in almost every segment of the industry. Unfortunately, their increased complexity often causes a loss of temporal predictability which represents a key requirement for hard real-time systems. Major sources of unpredictability are the shared low level resources, such as the memory hierarchy and the I/O subsystem. In this paper, we approach the problem of shared resource arbitration at an OS-level and propose a novel scratchpad-centric OS design for multi-core platforms. In the proposed OS, the predictable usage of shared resources across multiple cores represents a central design-time goal. Hence, we show (i) how contention-free execution of real-time tasks can be achieved on scratchpad-based architectures, and (ii) how a separation of application logic and I/O perations in the time domain can be enforced. To validate the proposed design, we implemented the proposed OS using a commercial-off-the-shelf (COTS) platform. Experiments show that this novel design delivers predictable temporal behavior to hard real-time tasks, and it improves performance up to 2.1× compared to traditional approaches. Rohan Tabish, Renato Mancuso 0001, Saud Wasly, Ahmed Alhammad, Sujit S. Phatak, Rodolfo Pellizzoni, Marco Caccamo |
RTAS | 6 |
| 2016 | Exploring Opportunistic Execution for Integrating Security into Legacy Hard Real-Time SystemsabstractDue to physical isolation as well as use of proprietary hardware and protocols, traditional real-time systems (RTS) were considered to be invulnerable to security breaches and external attacks. This assumption is being challenged by recent attacks that highlight vulnerabilities in RTS. Besides, a straightforward integration of security mechanisms might compromise the safety and predictability guarantees of such systems. In this paper, we focus on integrating security mechanisms into RTS (especially legacy RTS) and define a metric to measure the effectiveness of such integration. We combine opportunistic execution with hierarchical scheduling to maintain compatibility with legacy systems while still providing flexibility. The proposed approach is shown to increase the security posture of RTS without impacting their temporal (and hence, safety) constraints. Monowar Hasan, Sibin Mohan, Rakesh Bobba, Rodolfo Pellizzoni |
RTSS | 4 |
| 2016 | Integrating security constraints into fixed priority real-time schedulers
Sibin Mohan, Man-Ki Yoon, Rodolfo Pellizzoni, Rakesh Bobba |
Real Time Syst. | 3 |
| 2016 | A composable worst case latency analysis for multi-rank DRAM devices under open row policy
Zheng Pei Wu, Rodolfo Pellizzoni, Danlu Guo |
Real Time Syst. | 2 |
| 2016 | Global Real-Time Memory-Centric Scheduling for Multicore SystemsabstractAs the number of cores increases, more master components can simultaneously access main memory. In real-time systems, this ongoing trend is leading to crippling pessimism when computing the worst-case cache miss time, since a memory request could potentially contend with other requests coming from every other core in the system. CPU-centric scheduling policies, therefore, are no longer sufficient to guarantee schedulability without introducing unacceptable pessimism for memory-intensive task sets. For this reason, we believe a shift is needed towards real-time scheduling approaches that can prevent timing interference from memory contention, while still making efficient use of the multicore platform. Previously, we have demonstrated the practicality of the PREM task model, where each job consists of a sequence of phases, some of which access memory and some of which perform only computation on cached data. In this work, we present the first global memory-centric scheduling policy for memory-intensive task sets whose jobs can be modeled as a sequence of memory-intensive (memory phase) and execution-intensive (execution phase) phases. The proposed policy is parameterizable based on the number of cores which are allowed to concurrently access main memory without saturating it. Building upon results from multicore response-time analysis, we introduce the notion of virtual memory cores as a fundamental technique for performing phase-based response time analysis for memory-intensive task sets. Finally, we use synthetic task set generation to demonstrate that proposed scheduling policy and related schedulability bound do indeed better schedule memory-intensive task sets when compared to state-of-art multicore scheduling. Rodolfo Pellizzoni, Stanley Bak, Heechul Yun, Marco Caccamo |
IEEE Trans. Computers | 2 |
| 2016 | Schedulability Analysis for Memory Bandwidth Regulated Multicore Real-Time SystemsabstractMulticore architecture brings a significant challenge in designing critical real-time systems because of timing variability caused by concurrent accesses to shared memory. We propose a memory bandwidth regulated system architecture and a novel analysis method to address this challenge. In the proposed architecture, each core's memory access rate is regulated in a globally coordinated manner. The architecture allows system designers to control the system to satisfy desired real-time performance. The proposed analysis method provides a way to calculate worst case response time of each real-time task independently from other activities on other cores; it only depends on the task under analysis, the assigned bandwidth, and the number of cores in the system. We believe this independence is critical to enable modular certification of critical real-time systems. We implement the proposed system model on the gem5 architecture simulator. We evaluate the proposed analysis method by comparing the computed runtime with the measured runtime on the modified simulator. We show that the analysis method provides reasonable upper-bounds based on the SPEC2006 benchmark suite. Heechul Yun, Zheng Pei Wu, Rodolfo Pellizzoni, Marco Caccamo, Lui Sha |
IEEE Trans. Computers | 4 |
| 2016 | Memory Bandwidth Management for Efficient Performance Isolation in Multi-Core PlatformsabstractMemory bandwidth in modern multi-core platforms is highly variable for many reasons and it is a big challenge in designing real-time systems as applications are increasingly becoming more memory intensive. In this work, we proposed, designed, and implemented an efficient memory bandwidth reservation system, that we call MemGuard. MemGuard separates memory bandwidth in two parts: guaranteed and best effort. It provides bandwidth reservation for the guaranteed bandwidth for temporal isolation, with efficient reclaiming to maximally utilize the reserved bandwidth. It further improves performance by exploiting the best effort bandwidth after satisfying each core's reserved bandwidth. MemGuard is evaluated with SPEC2006 benchmarks on a real hardware platform, and the results demonstrate that it is able to provide memory performance isolation with minimal impact on overall throughput. Heechul Yun, Rodolfo Pellizzoni, Marco Caccamo, Lui Sha |
IEEE Trans. Computers | 3 |
| 2015 | WCET(m) Estimation in Multi-core Systems Using Single Core EquivalenceabstractMulti-core platforms represent the answer of the industry to the increasing demand for computational capabilities. From a real-time perspective, however, the inherent sharing of resources, such as memory subsystem and I/O channels, creates inter-core timing interference among critical tasks and applications deployed on different cores. As a result, modular per-core certification cannot be performed, meaning that: (1) current industrial engineering processes cannot be reused, (2) software developed and certified for single-core chips cannot be deployed on multi-core platforms as is. In this work, we propose the Single Core Equivalence (SCE) technology: a framework of OS-level techniques designed for commercial (COTS) architectures that exports a set of equivalent single-core virtual machines from a multi-core platform. This allows per-core schedulability results to be calculated in isolation and to hold when multiple cores of the system run in parallel. Thus, SCE allows each core of a multi-core chip to be considered as a conventional single-core chip, ultimately enabling industry to reuse existing software, schedulability analysis methodologies and engineering processes. Renato Mancuso 0001, Rodolfo Pellizzoni, Marco Caccamo, Lui Sha, Heechul Yun |
ECRTS | 2 |
| 2015 | Parallelism-Aware Memory Interference Delay Analysis for COTS Multicore SystemsabstractIn modern Commercial Off-The-Shelf (COTS) mul-ticore systems, each core can generate many parallel memory requests at a time. The processing of these parallel requests in the DRAM controller greatly affects the memory interference delay experienced by running tasks on the platform. In this paper, we present a new parallelism-aware worst-case memory interference delay analysis for COTS multicore systems. The analysis considers a COTS processor that can generate multiple outstanding requests and a COTS DRAM controller that has a separate read and write request buffer, prioritizes reads over writes, and supports out-of-order request processing. Focusing on LLC and DRAM bank partitioned systems, our analysis computes worst-case upper bounds on memory-interference delays, caused by competing memory requests. We validate our analysis on a Gem5 full-system simulator modeling a realistic COTS multicore platform, with a set of carefully designed synthetic benchmarks as well as SPEC2006benchmarks. The evaluation results show that our analysis produces safe upper bounds in all tested benchmarks, while the current state-of-the-art analysis significantly under-estimates the delays. Heechul Yun, Rodolfo Pellizzoni, Prathap Kumar Valsan |
ECRTS | 2 |
| 2015 | Generation of communication schedules using component interfacesabstractWith the growing demand in embedded systems, safety and non-safety critical parts are integrated together although guaranteeing safety is a hard problem to tackle due to the complexity of possible interactions between components involving communication. However, it is sufficiently recognized that separation of computation and communication can reduce the complexity of guaranteeing safety involved in interactions between components. In this work, we propose to use component interfaces derived from periodic resource supplies that can meet the demand of components experiencing bounded delays. The advantage of using interfaces is to provide minimal information of components without requiring the entire task specifications to generate a multi-mode communication schedule. We use integer linear programming to find assignments in generating schedules that are guaranteed to have low average mode-change delay. A video monitoring case study demonstrates the advantages on using our approach in generating communication schedules. Akramul Azim, Rodolfo Pellizzoni, Sebastian Fischmeister |
ETFA | 2 |
| 2015 | Memory efficient global scheduling of real-time tasksabstractCurrent computing architectures are commonly built with multiple cores and a single shared main memory. Even though this architecture increases the overall computation power, main memory can easily become a bottleneck. Simultaneous access to main memory from multiple cores can cause both (1) severe degradation in performance and (2) unpredictable execution time for real-time applications. We propose in this paper to mitigate these two problems by co-scheduling cores as well as the main memory for predictable execution. In particular, we use a DMA component to overlap memory with computation for hiding the memory latency and therefore increasing the system performance. The main contribution of this paper is a novel global co-scheduling algorithm along with its associated schedulability analysis for sporadic hard real-time tasks. We evaluated our system by generating synthetic tasksets based on real benchmark parameters. The results show a significant improvement in system utilization while retaining a predictable system behavior. Ahmed Alhammad, Saud Wasly, Rodolfo Pellizzoni |
RTAS | 3 |
| 2015 | A framework for scheduling DRAM memory accesses for multi-core mixed-time critical systemsabstractMixed-time critical systems are real-time systems that accommodate both hard real-time (HRT) and soft realtime (SRT) tasks. HRT tasks mandate a gurantee on the worstcase latency, while SRT tasks have average-case bandwidth (BW) demands. Memory requests in mixed-time critical systems usually have different transaction sizes based on whether the issuer task is HRT or SRT. For example, HRT tasks often issue requests with a cache line size. On the other side, SRT tasks may issue requests with a size of KBs. Requests from multimedia cores, cores controlling network interfaces and direct memory accesses (DMAs) are obvious examples of these large-size requests. Based on these observations, we promote in this work a new approach to schedule memory requests. This approach retains locality within large-size requests to minimize the worst-case latency, while maintaining the average-case BW as high as required. To achieve this target, we introduce a novel and compact time-division-multiplexing scheduler that is adequate for mixed-time critical systems. We also present a novel framework that constructs optimal offchip DRAM memory controller schedules for multi-core mixedtime critical systems. These schedules are loaded to the memory controller during boot-time. Based on the proposed schedule, we provide a detailed static analysis that guarantees predictability. We compare the proposed controller against state-of-the-art realtime memory controllers using synthetic experiments as well as a practical use-case from multimedia systems. Mohamed Hassan 0002, Hiren D. Patel, Rodolfo Pellizzoni |
RTAS | 3 |
| 2015 | A generalized model for preventing information leakage in hard real-time systemsabstractTraditionally real-time systems and security have been considered as separate domains. Recent attacks on various systems with real-time properties have shown the need for a redesign of such systems to include security as a first class principle. In this paper, we propose a general model for capturing security constraints between tasks in a real-time system. This model is then used in conjunction with real-time scheduling algorithms to prevent the leakage of information via storage channels on implicitly shared resources. We expand upon a mechanism to enforce these constraints viz., cleaning up of shared resource state, and provide schedulability conditions based on fixed priority scheduling with both preemptive and non-preemptive tasks. We perform extensive evaluations, both theoretical and experimental, the latter on a hardware-in-the-loop simulator of an unmanned aerial vehicle (UAV) that executes on a demonstration platform. Rodolfo Pellizzoni, Neda Paryab, Man-Ki Yoon, Stanley Bak, Sibin Mohan, Rakesh Bobba |
RTAS | 1 |
| 2014 | Time-predictable execution of multithreaded applications on multicore systemsabstractIn multicore systems, contention for access to main memory between application threads complicates timing analysis and may lead to pessimistic bounds on execution time. This is particularly problematic for real-time applications, which require provable bounds on worst-case performance. In this work, we employ a predictable execution model to schedule memory accesses performed by application threads without relying on unpredictable hardware arbiters. In addition, we statically schedule application's threads with the objective to minimize the application's makespan. Our experimental evaluation with NAS Parallel Benchmarks on 4-core system indicates that the proposed execution scheme yields an aggregated improvement of 21% over contention execution in which application's threads uncontrollably access main memory. Ahmed Alhammad, Rodolfo Pellizzoni |
DATE | 2 |
| 2014 | Generation of communication schedules for multi-mode distributed real-time applicationsabstractA key problem in designing multi-mode real-time systems is the generation of schedules to reduce the complexities of transforming the model semantics to code. Moreover, distributed multi-mode applications are prone to suffer from delays incurred during mode changes. We therefore aim to generate communication schedules that have low average mode-change delay for multi-mode real-time distributed applications. In this paper, we use optimization constraints associated to timing requirements to generate state-based schedules for multi-mode communication systems, and illustrate the workflow for generating schedules from specifications through a real-time video monitoring case-study. Our experiments in the case-study demonstrate that schedules generated using the proposed method reduce the average mode-change delay in relation to a randomized algorithm and the well-known EDF scheduling algorithm. Akramul Azim, Gonzalo Carvajal, Rodolfo Pellizzoni, Sebastian Fischmeister |
DATE | 3 |
| 2014 | A Rank-Switching, Open-Row DRAM Controller for Time-Predictable SystemsabstractWe introduce ROC, a Rank-switching, Open-row Controller for Double Data Rate Dynamic RAM (DDR DRAM). ROC is optimized for mixed-criticality multicore systems using modern DDR devices: compared to existing real-time memory controllers, it provides significantly lower worst case latency bounds for hard real-time tasks and supports throughput-oriented optimizations for soft real-time applications. The key to improved performance is an innovative rank-switching mechanism which hides the latency of write-read transitions in DRAM devices without requiring unpredictable request reordering. We further employ open row policy to take advantage of the data caching mechanism (row buffering) in each device. ROC provides complete timing isolation between hard and soft tasks and allows for compositional timing analysis over the number of cores and memory ranks in the system. We implemented and synthesized the ROC back end in Verilog RTL, and evaluated its performance on both synthetic tasks and a set of representative benchmarks. Yogen Krish, Zheng Pei Wu, Rodolfo Pellizzoni |
ECRTS | 3 |
| 2014 | Real-Time Systems Security through Scheduler ConstraintsabstractReal-time systems (RTS) were typically considered to be invulnerable to external attacks, mainly due to their use of proprietary hardware and protocols, as well as physical isolation. As a result, RTS and security have traditionally been separate domains. These assumptions are being challenged by a series of recent events that highlight the vulnerabilities in RTS. In this paper we focus on integrating security as a first class principle in the design of RTS: we show that certain security requirements can be specified as real-time scheduling constraints. Using information leakage as a motivating problem, we illustrate our techniques with fixed-priority (FP) real-time schedulers. We evaluate our approach and discuss tradeoffs. Our evaluation shows that many real-time task sets can be scheduled under the proposed constraints without significant performance impact. Sibin Mohan, Man-Ki Yoon, Rodolfo Pellizzoni, Rakesh Bobba |
ECRTS | 3 |
| 2014 | Schedulability analysis of global memory-predictable schedulingabstractThe use of multicore CPUs in real-time systems poses significant challenges in estimating their temporal behavior. A factor that has a large impact on this issue is the contention for access to main memory among multiple cores. To overcome this problem, an execution model called PREM has been previously introduced to co-schedule CPU execution and accesses to main memory without relying on hardware arbiters. In this paper, we provide a global schedulability analysis for this predictable execution model, and we prove its correctness. We also evaluate the effectiveness of the proposed solution with extensive simulations. The results show a significant advantage of the proposed solution when compared to contention execution in which tasks access main memory unpredictably. Ahmed Alhammad, Rodolfo Pellizzoni |
EMSOFT | 2 |
| 2014 | Hiding memory latency using fixed priority schedulingabstractModern embedded platforms contain a variety of physical resources, such as caches, interconnects, main memory, etc., which the processor must access during the execution of a task. We argue that processor task execution and accesses to physical resources should be co-scheduled in real-time systems to predictably hide resource access latency. In particular, in this work we focus on co-scheduling task execution and accesses to main memory to hide DRAM access latency. Since modern systems implement DMA controllers that can be operated independently of processor execution, this allows us to hide memory transfer latency by scheduling DMA transfer in parallel with processor execution. The main contribution of this paper is a dynamic scheduling algorithm for a set of sporadic real-time tasks that efficiently co-schedules processor and DMA execution to hide memory transfer latency. The proposed algorithm can be applied to either uniprocessor or partitioned multiprocessor systems. We demonstrate that we improve processor utilization significantly compared to existing scratchpad and cache management systems. Saud Wasly, Rodolfo Pellizzoni |
RTAS | 2 |
| 2014 | PALLOC: DRAM bank-aware memory allocator for performance isolation on multicore platformsabstractDRAM consists of multiple resources called banks that can be accessed in parallel and independently maintain state information. In Commercial Off-The-Shelf (COTS) multicore platforms, banks are typically shared among all cores, even though programs running on the cores do not share memory space. In this situation, memory performance is highly unpredictable due to contention in the shared banks. In this paper, we propose PALLOC, a DRAM bank-aware memory allocator which exploits the page-based virtual memory system to allocate memory pages of each application to specific banks. With PALLOC, we can dynamically partition banks to avoid bank sharing among cores, thereby improving isolation on COTS multicore platforms without requiring any special hardware support. We performed an extensive set of experiments to investigate the performance impact of DRAM bank partitioning on two COTS multicore platforms with a set of synthetic and SPEC2006 benchmarks. Our evaluation results demonstrate that DRAM bank partitioning significantly improves isolation and real-time performance. Heechul Yun, Renato Mancuso 0001, Zheng Pei Wu, Rodolfo Pellizzoni |
RTAS | 4 |
| 2014 | A formal approach to the WCRT analysis of multicore systems with memory contention under phase-structured task sets
Kai Lampka, Georgia Giannopoulou, Rodolfo Pellizzoni, Nikolay Stoimenov |
Real Time Syst. | 3 |
| 2013 | A Dynamic Scratchpad Memory Unit for Predictable Real-Time Embedded SystemsabstractScratchpad memory is an attractive alternative to caches in real-time embedded systems due to its advantages in terms of timing predictability and power consumption. However, dynamic management of scratchpad content is challenging in multitasking environments. To address this issue, we propose the design of a novel Real-Time Scratchpad Memory Unit (RSMU). Our RSMU can be integrated in existing systems with minimal architectural modifications. Furthermore, scratchpad management is performed at the OS level, requiring no application changes. Compared to existing multitasking scratchpad management schemes, our approach improves schedulability by hiding the latency of memory transfers. We demonstrate and evaluate our system design on an embedded FPGA platform. Saud Wasly, Rodolfo Pellizzoni |
ECRTS | 2 |
| 2013 | ORTAP: An Offset-based response time analysis for a pipelined communication resource modelabstractThis work addresses the challenge of computing worst-case response times of hard real-time applications deployed on multiprocessor systems. In particular, the worst-case response time analysis (WCRTA) focuses on the communication between distributed tasks of hard real-time applications. The proposed WCRTA models the communication as a pipelined communication resource model. This model incorporates the effect of pipelining, and the parallel transmission of data. Applications of such a model include multiprocessor systems that use complex interconnects such as network-on-chips (NoC)s with priorities. In this paper, we present an exponential analysis, and a polynomial analysis, and prove its correctness. As an application, we apply the pipelined communication resource model to priority-aware NoCs, and we compare the proposed analyses against prior analysis techniques. Our experimental evaluation on two instances of 4 × 4 and 8 × 8 NoCs with 512,000 synthetic benchmarks shows 48.3% and 66.7% improvement in schedulability for the two NoC sizes over prior work. Hany Kashif, Sina Gholamian, Rodolfo Pellizzoni, Hiren D. Patel, Sebastian Fischmeister |
IEEE Real-Time and Embedded Technology and Applications Symposium | 3 |
| 2013 | Real-time cache management framework for multi-core architecturesabstractMulti-core architectures are shaking the fundamental assumption that in real-time systems the WCET, used to analyze the schedulability of the complete system, is calculated on individual tasks. This is not even true in an approximate sense in a modern multi-core chip, due to interference caused by hardware resource sharing. In this work we propose (1) a complete framework to analyze and profile task memory access patterns and (2) a novel kernel-level cache management technique to enforce an efficient and deterministic cache allocation of the most frequently accessed memory areas. In this way, we provide a powerful tool to address one of the main sources of interference in a system where the last level of cache is shared among two or more CPUs. The technique has been implemented on commercial hardware and our evaluations show that it can be used to significantly improve the predictability of a given set of critical tasks. Renato Mancuso 0001, Roman Dudko, Emiliano Betti, Marco Cesati, Marco Caccamo, Rodolfo Pellizzoni |
IEEE Real-Time and Embedded Technology and Applications Symposium | 6 |
| 2013 | MemGuard: Memory bandwidth reservation system for efficient performance isolation in multi-core platformsabstractMemory bandwidth in modern multi-core platforms is highly variable for many reasons and is a big challenge in designing real-time systems as applications are increasingly becoming more memory intensive. In this work, we proposed, designed, and implemented an efficient memory bandwidth reservation system, that we call MemGuard. MemGuard distinguishes memory bandwidth as two parts: guaranteed and best effort. It provides bandwidth reservation for the guaranteed bandwidth for temporal isolation, with efficient reclaiming to maximally utilize the reserved bandwidth. It further improves performance by exploiting the best effort bandwidth after satisfying each core's reserved bandwidth. MemGuard is evaluated with SPEC2006 benchmarks on a real hardware platform, and the results demonstrate that it is able to provide memory performance isolation with minimal impact on overall throughput. Heechul Yun, Rodolfo Pellizzoni, Marco Caccamo, Lui Sha |
IEEE Real-Time and Embedded Technology and Applications Symposium | 3 |
| 2013 | Worst Case Analysis of DRAM Latency in Multi-requestor SystemsabstractAs multi-core systems are becoming more popular in real-time embedded systems, strict timing requirements for accessing shared resources must be met. In particular, a detailed latency analysis for Double Data Rate Dynamic RAM (DDR DRAM) is highly desirable. Several researchers have proposed predictable memory controllers to provide guaranteed memory access latency. However, the performance of such controllers sharply decreases as DDR devices become faster and the width of memory buses is increased. In this paper, we present a novel, composable worst case analysis for DDR DRAM that provides improved latency bounds compared to existing works by explicitly modeling the DRAM state. In particular, our approach scales better with increasing number of requestors and memory speed. Benchmark evaluations show up to 62% improvement in worst case task execution time compared to a competing predictable memory controller for a system with 8 requestors. Zheng Pei Wu, Yogen Krish, Rodolfo Pellizzoni |
RTSS | 3 |
| 2013 | Implementation and evaluation of global and partitioned scheduling in a real-time OS
Giovani Gracioli, Antônio Augusto Fröhlich, Rodolfo Pellizzoni, Sebastian Fischmeister |
Real Time Syst. | 3 |
| 2013 | Real-Time I/O Management System with COTS PeripheralsabstractReal-time embedded systems are increasingly being built using commercial-off-the-shelf (COTS) components such as mass-produced peripherals and buses to reduce costs, time-to-market, and increase performance. Unfortunately, COTS-interconnect systems do not usually guarantee timeliness, and might experience severe timing degradation in the presence of high-bandwidth I/O peripherals. Moreover, peripherals do not implement any internal priority-based scheduling mechanism, hence, sharing a device can result in data of high priority tasks being delayed by data of low priority tasks. To address these problems, we designed a real-time I/O management system comprised of 1) real-time bridges with I/O virtualization capabilities, and 2) a peripheral scheduler. The proposed framework is used to transparently put the I/O subsystem of a COTS-based embedded system under the discipline of real-time scheduling, minimizing the timing unpredictability due to the peripherals sharing the bus. We also discuss computing the maximum delay due to buffered I/O data transactions as well as determining the buffer size needed to avoid data loss. Finally, we demonstrate experimentally that our prototype real-time I/O management system successfully exports multiple virtual devices for a single physical device and prioritizes I/O traffic, guaranteeing its timeliness. Emiliano Betti, Stanley Bak, Rodolfo Pellizzoni, Marco Caccamo, Lui Sha |
IEEE Trans. Computers | 3 |
| 2012 | Memory Access Control in Multiprocessor for Real-Time Systems with Mixed CriticalityabstractShared resource access interference, particularly memory and system bus, is a big challenge in designing predictable real-time systems because its worst case behavior can significantly differ. In this paper, we propose a software based memory throttling mechanism to explicitly control the memory interference. We developed analytic solutions to compute proper throttling parameters that satisfy schedulability of critical tasks while minimize performance impact caused by throttling. We implemented the mechanism in Linux kernel and evaluated isolation guarantee and overall performance impact using a set of synthetic and real applications. Heechul Yun, Rodolfo Pellizzoni, Marco Caccamo, Lui Sha |
ECRTS | 3 |
| 2012 | Memory-Aware Scheduling of Multicore Task Sets for Real-Time SystemsabstractReal-time scheduling of memory-intensive applications is a particularly difficult challenge. On a multi-core system, not only is the CPU scheduling an issue, but equally important is the management of mutual interference among tasks caused by simultaneous access to the shared main memory. To confront this problem, we explore real-time schedulers for task sets which adhere to the Predictable Execution Model (PREM). In each PREM-compliant task, execution is divided into phases which retrieve data from main memory, and phases which perform local computation using previously-cached data. In this work, we perform a simulation-based analysis with the goal of determining which schedulers are generally better at scheduling PREM-compliant task sets. We investigate several memory intensive real-time benchmarks from the EEMBC benchmark suite, in order to drive our task set generation parameters. We elaborate on a PREM-complaint task set simulator which we designed specifically to be able to simulate PREM-compliant tasks. The overall best scheduling policy we found, which we call M-LAX, schedules access to memory in a no preemptive fashion according to a least-laxity-first policy. M-LAX outperforms an EDF-based approach, a previously-analyzed TDMA arbitration scheme, and the unscheduled case where tasks interfere when accessing memory. Stanley Bak, Rodolfo Pellizzoni, Marco Caccamo |
RTCSA | 3 |
| 2012 | Memory-centric scheduling for multicore hard real-time systems
Rodolfo Pellizzoni, Stanley Bak, Emiliano Betti, Marco Caccamo |
Real Time Syst. | 2 |
| 2012 | Real-Time Scheduling of Concurrent Transactions in Multidomain Ring BusesabstractWe address the problem of scheduling concurrent periodic real-time transactions on Multidomain Ring Bus (MDRB). The problem is challenging because although the bus allows multiple nonoverlapping transactions to be executed concurrently, the degree of concurrency depends on the topology of the bus and of executed transactions. To solve this problem, first, we propose two novel efficient scheduling algorithms for topographically acyclic transaction sets. The first algorithm is optimal for transaction sets under restrictive assumptions while the second one induces a good sufficient schedulable utilization bound for more general transaction sets. Then, we extend these two algorithms for the scheduling of topographically cyclic transaction sets. Extensive simulations show that the proposed algorithm can schedule transaction sets with high bus utilization and is better than that of related works in most practical settings. The implementation of the algorithms in a real testbed shows that they have relatively low execution-time overhead. Bach Duy Bui, Rodolfo Pellizzoni, Marco Caccamo |
IEEE Trans. Computers | 2 |
| 2012 | Modeling towards incremental early analyzability of networked avionics systems using virtual integrationabstractWith the advance of hardware technology, more features are incrementally added to already existing networked systems. Avionics has a stronger tendency to use preexisting applications due to its complexity and scale. As resource sharing becomes intense among the network and the computing modules, it has become a difficult task for the system designer to make confident architectural decisions even for incremental changes. Providing a tailored environment to model and analyze incremental changes requires a combination of software tools and hardware support. We have built a virtual integration tool called ASIIST which can provide a worst-case end-to-end latency of data that is sent through a network and the internal bus architecture of the end-systems. Also, we have devised a new real-time switching algorithm which guarantees the worst-case network delay of preexisting network traffic under feasible conditions. With the real-time switch support, ASIIST can provide an early modularized analysis of the end-to-end latency to make architectural design choices and incremental changes easier for the user. Min-Young Nam, Kyungtae Kang, Rodolfo Pellizzoni, Kyung-Joon Park, Jung-Eun Kim, Lui Sha |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2011 | A Predictable Execution Model for COTS-Based Embedded SystemsabstractBuilding safety-critical real-time systems out of inexpensive, non-real-time, COTS components is challenging. Although COTS components generally offer high performance, they can occasionally incur significant timing delays. To prevent this, we propose controlling the operating point of each shared resource (like the cache, memory, and interconnection buses) to maintain it below its saturation limit. This is necessary because the low-level arbiters of these shared resources are not typically designed to provide real-time guarantees. In this work, we introduce a novel system execution model, the Predictable Execution Model (PREM), which, in contrast to the standard COTS execution model, coschedules at a high level all active components in the system, such as CPU cores and I/O peripherals. In order to permit predictable, system-wide execution, we argue that real-time embedded applications should be compiled according to a new set of rules dictated by PREM. To experimentally validate our theory, we developed a COTS-based PREM testbed and modified the LLVM Compiler Infrastructure to produce PREM-compatible executables. Rodolfo Pellizzoni, Emiliano Betti, Stanley Bak, John Criswell, Marco Caccamo, Russell Kegley |
IEEE Real-Time and Embedded Technology and Applications Symposium | 1 |
| 2011 | Timing Analysis for Resource Access Interference on Adaptive Resource ArbitersabstractModern multiprocessor and multicore architectures adopt shared resources to meet increased performance requirements. Adaptive arbiters, such as FlexRay, have been adopted to grant access to shared resources. While increasing the performance, timing analysis is more challenging with this kind of arbiter. This paper considers real-time tasks that are composed of super blocks, while super blocks themselves are composed of phases. Phases are characterized by their worst-case computation time on their processing element and their worst-case number of access requests to a shared resource. Resource accesses, such as access to caches or scratchpad memory, are synchronous and cause the processing element to stall until the access is served. Based on dynamic programming, we develop an algorithm that safely derives an upper-bound of the worst-case response time of a phase. The worst-case response time of a task can then be determined for both sequential or time-triggered execution of super blocks. Experimental results are conducted for a real-world application. Andreas Schranzhofer, Rodolfo Pellizzoni, Jian-Jia Chen, Lothar Thiele, Marco Caccamo |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2011 | A Slot-Based Real-Time Scheduling Algorithm for Concurrent Transactions in NoCabstractWe address the problem of scheduling real-time transactions in Network-on-Chip (NoC). In particular, we propose a novel slot-based scheduling algorithm for acyclic transaction sets in NoC. The algorithm induces a competitive sufficient schedulability utilization bound. Since the proposed algorithm is able to exploit the parallelism between non-overlapping transactions, under given assumptions, it performs better than the existing fixed-priority solutions. We evaluate performance through extensive simulations. Furthermore, we discuss some important factors in the implementation of the algorithm and present an implementation in a real system. The measurement shows that the proposed algorithm has relatively low overhead. Bach Duy Bui, Marco Caccamo, Rodolfo Pellizzoni |
RTCSA (1) | 3 |
| 2010 | Worst-case response time analysis of resource access models in multi-core systemsabstractMulti-processor and multi-core systems are becoming increasingly important in time critical systems. Shared resources, such as shared memory or communication buses are used to share data and read sensors. We consider real-time tasks constituted by superblocks, which can be executed sequentially or by a time triggered static schedule. Three models to access shared resources are explored: (1) the dedicated access model, in which accesses happen only in dedicated phases, (2) the general access model, in which accesses could happen at anytime, and (3) the hybrid access model, combining the dedicated and general access model. For resource access based on a Time Division Multiple Access (TDMA) protocol, we analyze the worst-case completion time for a superblock, derive worst-case response times for tasks and obtain the relation of schedulability between different models. We conclude with proposing the dedicated sequential model as the model of choice for time critical resource sharing multi-processor/multi-core systems. Andreas Schranzhofer, Rodolfo Pellizzoni, Jian-Jia Chen, Lothar Thiele, Marco Caccamo |
DAC | 2 |
| 2010 | Worst case delay analysis for memory interference in multicore systemsabstractEmploying COTS components in real-time embedded systems leads to timing challenges. When multiple CPU cores and DMA peripherals run simultaneously, contention for access to main memory can greatly increase a task's WCET. In this paper, we introduce an analysis methodology that computes upper bounds to task delay due to memory contention. First, an arrival curve is derived for each core representing the maximum memory traffic produced by all tasks executed on it. Arrival curves are then combined with a representation of the cache behavior for the task under analysis to generate a delay bound. Based on the computed delay, we show how tasks can be feasibly scheduled according to assigned time slots on each core. Rodolfo Pellizzoni, Andreas Schranzhofer, Jian-Jia Chen, Marco Caccamo, Lothar Thiele |
DATE | 1 |
| 2010 | Real-Time Communication for Multicore Systems with Multi-domain Ring BusesabstractWe address the problem of scheduling real-time data transactions on a multicore processor bus. In particular, to in-crease system predictability and tighten WCET estimation, we propose to employ a software-controllable Multi-Domain Ring Bus (MDRB) architecture. The problem of scheduling periodic real-time transactions on MDRB is challenging because the bus allows multiple non-overlapping transactions to be executed concurrently, and because the degree of concurrency depends on the topology of the bus and of executed transactions. We propose a practical abstraction mechanism for the scheduling problem together with two novel scheduling algorithms. The first algorithm is optimal for transaction sets under restrictive assumptions while the second one induces a competitive sufficient schedulable utilization bound for more general transaction sets. Bach Duy Bui, Rodolfo Pellizzoni, Deepti K. Chivukula, Marco Caccamo |
RTCSA | 2 |
| 2010 | Impact of Peripheral-Processor Interference on WCET Analysis of Real-Time Embedded SystemsabstractThe integration phase of real-time COTS-based systems is challenging. When multiple tasks run concurrently, the interference at the bus level between cache fetching activities and I/O peripheral transactions is significant and causes unpredictable behaviors: experimentally, we show that tasks can have computation time variance up to 46 percent in a typical embedded system. In this work, we present a theoretical framework able to model the interaction between CPU and peripherals contending for shared main memory through the Front Side Bus (FSB). We first show how to compute worst case execution time (WCET) for a task given a trace of its cache activity and given an upper bound function that models peripheral activities. Then, we show how the analysis can be extended to a multitasking environment assuming a restricted-preemption model. Finally, we introduce the novel idea of ¿hardware server¿ as a means of controlling the unpredictable behavior of COTS peripheral components. Rodolfo Pellizzoni, Marco Caccamo |
IEEE Trans. Computers | 1 |
| 2009 | Handling mixed-criticality in SoC-based real-time embedded systemsabstractSystem-on-Chip (SoC) is a promising paradigm to implement safety-critical embedded systems, but it poses significant challenges from a design and verification point of view. In particular, in a mixed-criticality system, low criticality applications must be prevented from interfering with high criticality ones. In this paper, we introduce a new design methodology for SoC that provides strong isolation guarantees to applications with different criticalities. A set of certificates describing the assumed application behavior is extracted from a functional Architectural Analysis and Design Language (AADL) specification. Our tools then automatically generate hardware wrappers that enforce at run-time the behavior described by the certificates. In particular, we employ run-time monitoring to formally check all data communication in the system, and we enforce timing reservations for both computation and communication resources. Verification is greatly simplified because certificates are much simpler than the components used to implement low-criticality applications. The effectiveness of our methodology is proven on a case study consisting of a medical pacemaker. Rodolfo Pellizzoni, Patrick O'Neil Meredith, Min-Young Nam, Mu Sun, Marco Caccamo, Lui Sha |
EMSOFT | 1 |
| 2009 | ASIIST: Application Specific I/O Integration Support Tool for Real-Time Bus Architecture DesignsabstractIn hard real-time systems such as avionics, computer board level designs are typically customized to meet specific reliability and real time requirements. This paper focuses on computer-aided application-specific design of I/O architecture using PCI as an example. We have built a tool (ASIIST) that will enable engineers to explore design spaces at the I/O bus architecture level, performing analysis that incorporates bus protocols, to provide guarantees of real-time properties. Min-Young Nam, Rodolfo Pellizzoni, Lui Sha, Richard M. Bradford |
ICECCS | 2 |
| 2009 | Real-Time Control of I/O COTS Peripherals for Embedded SystemsabstractReal-time embedded systems are increasingly being built using commercial-off-the-shelf (COTS) components such as mass-produced peripherals and buses to reduce costs, time-to-market, and increase performance. Unfortunately, COTS interconnect systems do not usually guarantee timeliness, and might experience severe timing degradation in the presence of high-bandwidth I/O peripherals. To address this problem, we designed a real-time I/O management system comprised of 1) real-time bridges, and 2) a reservation controller. The proposed framework is used to transparently put the I/O subsystem of a COTS-based embedded system under the discipline of real-time scheduling. We also discuss computing a delay bound for I/O data transactions and determining worst-case buffer size. Finally, we demonstrate experimentally that our prototype real-time I/O management system successfully prioritizes I/O traffic and guarantees its timeliness. Stanley Bak, Emiliano Betti, Rodolfo Pellizzoni, Marco Caccamo, Lui Sha |
RTSS | 3 |
| 2009 | Rapid Early-Phase Virtual IntegrationabstractIn complex hard real-time systems with tight constraints on system resources, small changes in one component of a system can cause a cascade of adverse effects on other parts of the system. We address the inherent complexity of making architectural decisions by raising the level of abstraction at which the analysis is performed. Our analysis approach gives the system architect a rigorous method for quickly determining which system architectures should be pursued, and it allows the architect to track and manage the cascading effects of subsystem/component changes in a comprehensive, quantitative manner. The end product is a virtual architecture analysis that systematically incorporates the inherent coupling among interacting system components that share limited system resources. Sibin Mohan, Min-Young Nam, Rodolfo Pellizzoni, Lui Sha, Richard M. Bradford, Shana Fliginger |
RTSS | 3 |
| 2008 | Hybrid Hardware-Software Architecture for Reconfigurable Real-Time SystemsabstractRecent developments in the field of reconfigurable SoC devices (FPGAs) will enable the development of embedded systems where software tasks, running on a CPU, can coexist with hardware tasks. We devised a real-time computing architecture that can integrate hardware and software executions in a transparent manner, and can support real-time QoS adaptation by means of partial reconfiguration of modern FPGA devices. Tasks are allowed to migrate seamlessly from CPU to FPGA and vice versa to support dynamic QoS adaptation and cope with dynamic workloads. In this paper, we discuss the design and implementation of an on-chip infrastructure, OS extensions and task design methodology that enable hardware-software transparency in the presence of relocation. The overall architecture is suitable to schedule real-time workloads and we derive bounds on relocation overhead. Finally, we show the applicability of our design methodology on a concrete task design case. Rodolfo Pellizzoni, Marco Caccamo |
IEEE Real-Time and Embedded Technology and Applications Symposium | 1 |
| 2008 | Coscheduling of CPU and I/O Transactions in COTS-Based Embedded SystemsabstractIntegrating COTS components in critical real-time systems is challenging. In particular, we show that the interference between cache activity and I/O traffic generated by COTS peripherals can unpredictably slow down a real-time task by up to 44%. To solve this issue, we propose a framework comprised of three main components: 1) a COTS-compatible device, the peripheral gate, that controls peripheral access to the system; 2) an analytical technique that computes safe bounds on the I/O-induced task delay; 3) a coscheduling algorithm that maximizes the amount of allowed peripheral traffic while guaranteeing all real-time task constraints. We implemented the complete framework on a COTS-based system using PCI peripherals, and we performed extensive experiments to show its feasibility. Rodolfo Pellizzoni, Bach Duy Bui, Marco Caccamo, Lui Sha |
RTSS | 1 |
| 2008 | Hardware Runtime Monitoring for Dependable COTS-Based Real-Time Embedded SystemsabstractCOTS peripherals are heavily used in the embedded market, but their unpredictability is a threat for high-criticality real-time systems: it is hard or impossible to formally verify COTS components. Instead, we propose to monitor the runtime behavior of COTS peripherals against their assumed specifications. If violations are detected, then an appropriate recovery measure can be taken. Our monitoring solution is decentralized: a monitoring device is plugged in on a peripheral bus and monitors the peripheral behavior by examining read and write transactions on the bus. Provably correct (w.r.t. given specifications) hardware monitors are synthesized from high level specifications, and executed on FPGAs, resulting in zero runtime overhead on the system CPU. The proposed technique, called BusMOP, has been implemented as an instance of a generic runtime verification framework, called MOP, which until now has only been used for software monitoring. We experimented with our technique using a COTS data acquisition board. Rodolfo Pellizzoni, Patrick O'Neil Meredith, Marco Caccamo, Grigore Rosu |
RTSS | 1 |
| 2008 | M-CASH: A real-time resource reclaiming algorithm for multiprocessor platforms
Rodolfo Pellizzoni, Marco Caccamo |
Real Time Syst. | 1 |
| 2007 | Soft Real-Time Chains for Multi-Hop Wireless Ad-Hoc NetworksabstractPrioritized MAC protocols are needed to support soft real-time communication in wireless networks. In this paper, we introduce real-time chain, a new prioritized MAC protocol to support soft real-time data flows in multi-hop wireless ad-hoc networks. By avoiding packet collisions and limiting the effect of priority inversions, real-time chain is able to provide soft real-time and bandwidth guarantees. Furthermore, the use of multiple channels enables high spatial reuse and transmission rates. Finally, our approach can be integrated with a slightly modified version of the IEEE 802.15.4 standard. The protocol has been fully implemented on Crossbow MICAz hardware and its performance has been validated with a large set of both indoor and outdoor experiments Bach Duy Bui, Rodolfo Pellizzoni, Marco Caccamo, Chin F. Cheah, Andrew Tzakis |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2007 | Toward the Predictable Integration of Real-Time COTS Based SystemsabstractThe integration phase of real-time COTS-based systems is often problematic because when multiple tasks run concurrently, the interference at the bus level between cache fetching activities and I/O peripheral transactions is significant and causes unpredictable behaviors: experimentally, tasks can have computation time variance up to 50%. In this work, we present a theoretical framework able to model the interaction between CPU and peripherals contending for shared main memory through the front side bus (FSB). We first show how to compute worst case execution times for a task given a trace of its cache activity and given an upper bound function that models peripheral activities; then, we introduce the novel idea of "hardware server" as a means of controlling the unpredictable behavior of COTS peripheral components. Rodolfo Pellizzoni, Marco Caccamo |
RTSS | 1 |
| 2007 | Holistic analysis of asynchronous real-time transactions with earliest deadline scheduling
Rodolfo Pellizzoni, Giuseppe Lipari |
J. Comput. Syst. Sci. | 1 |
| 2007 | Real-Time Management of Hardware and Software Tasks for FPGA-based Embedded SystemsabstractOperating systems for reconfigurable devices enable the development of embedded systems where software tasks, running on a CPU, can coexist with hardware tasks running on a reconfigurable hardware device (FPGA). In this work, we consider real-time systems subject to dynamic workloads and whose tasks can be computationally intensive. We introduce a novel resource allocation scheme and an online admission control test that achieve high performance and flexibility; in addition, runtime reconfiguration is used to maximize the number of admitted real-time tasks. Moreover, in detail, we first discuss a 1D system architecture and its prototype for a Xilinx Virtex-4 FPGA; then, we concentrate on the online admission control problem. Online task allocation and migration between the CPU and the reconfigurable device are discussed and sufficient feasibility tests are derived for both the commonly used slotted and 1D area models. Finally, the effectiveness of our admission control and relocation strategy is shown through a series of synthetic simulations. Rodolfo Pellizzoni, Marco Caccamo |
IEEE Trans. Computers | 1 |
| 2005 | Improved Schedulability Analysis of Real-Time Transactions with Earliest Deadline SchedulingabstractIn this paper we address the problem of schedulability analysis of distributed real-time transactions under EDF, where each transaction is a chain of precedence constrained tasks. We propose a new efficient methodology and a set of algorithms that explicitly take into account the offsets of the transactions. We show, with an extensive set of simulations, that our methodology provides improved schedulability conditions with respect to existing algorithms. Finally, we apply the methodology to an important class of systems: heterogeneous multiprocessor systems, with a general purpose processor and one or more coprocessors (DSPs). Rodolfo Pellizzoni, Giuseppe Lipari |
IEEE Real-Time and Embedded Technology and Applications Symposium | 1 |
| 2005 | Feasibility Analysis of Real-Time Periodic Tasks with Offsets
Rodolfo Pellizzoni, Giuseppe Lipari |
Real Time Syst. | 1 |
| 2004 | A New Sufficient Feasibility Test for Asynchronous Real-Time Periodic Task Sets
Rodolfo Pellizzoni, Giuseppe Lipari |
ECRTS | 1 |