VLDB 2026 Research / reviewers in the wild / expert
Marco Caccamo
dblp:86/450
· DBLP profile ↗
135ranked-venue papers
10as first author
45since 2021 · last 2026
0000-0003-2328-044XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 65 · 3 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 5 first-author · 7 since 2021Software engineering, systems software and programming languages · 10 · 6 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Computer networks · 5Theory of computation · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AI Inference in the Heat: Thermal-Aware Strict Partitioning for Configurable Real-Time Gang Tasks
Binqi Sun, Jinyang Li 0004, Tomasz Kloda, Tarek F. Abdelzaher, Marco Caccamo |
RTAS | 5 |
| 2026 | ETM2: Empowering Traditional Memory Bandwidth Regulation using ETM
Alexander Züpke, Ashutosh Pradhan, Daniele Ottaviano, Andrea Bastoni, Marco Caccamo |
RTAS | 5 |
| 2025 | Multi-Objective Memory Bandwidth Regulation and Cache Partitioning for Multicore Real-Time Systems
Binqi Sun, Zhihang Wei, Andrea Bastoni, Debayan Roy, Mirco Theile, Tomasz Kloda, Rodolfo Pellizzoni, Marco Caccamo |
ECRTS | 8 |
| 2025 | Arm Dynamiq Shared Unit and Real-Time: An Empirical EvaluationabstractThe increasing complexity of embedded hardware platforms poses significant challenges for real-time workloads. Architectural features such as Intel RDT, Arm QoS, and Arm MPAM are either unavailable on commercial embedded platforms or designed primarily for server environments optimized for average-case performance and might fail to deliver the expected real-time guarantees. Arm DynamIQ Shared Unit (DSU) includes isolation features-among others, hardware per-way cache partitioning-that can improve the real-time guarantees of complex embedded multicore systems and facilitate real-time analysis. However, the DSU also targets average cases, and its real-time capabilities have not yet been evaluated. This paper presents the first comprehensive analysis of three real-world deployments of the Arm DSU on Rockchip RK3568, Rockchip RK3588, and NVIDIA Orin platforms. We integrate support for the DSU at the operating system and hypervisor level and conduct a large-scale evaluation using both synthetic and real-world benchmarks with varying types and intensities of interference. Our results make extensive use of performance counters and indicate that, although effective, the quality of partitioning and isolation provided by the DSU depends on the type and the intensity of the interfering workloads. In addition, we uncover and analyze in detail the correlation between benchmarks and different types and intensities of interference. Ashutosh Pradhan, Daniele Ottaviano, Haozheng Huang, Alexander Züpke, Andrea Bastoni, Marco Caccamo |
RTAS | 7 |
| 2025 | Predictable Memory Bandwidth Regulation for DynamIQ Arm Systems
Ashutosh Pradhan, Daniele Ottaviano, Haozheng Huang, Alexander Züpke, Andrea Bastoni, Marco Caccamo |
RTCSA | 8 |
| 2025 | Work-in-Progress: Toward Real-Time Cross-ISA Execution on the AMD Embedded+ ArchitectureabstractEmerging embedded platforms increasingly rely on heterogeneous processing units to address diverse performance and energy requirements. The recently introduced AMD Embedded+ architecture reflects this trend by interconnecting via PCIe on the same motherboard one AMD x86 host processor with one Arm AArch64+FPGA complex. This implementation is another step forward towards a more compact heterogeneousISA platform designed with embedded applications in mind. While cross-ISA execution has been explored in the past with a focus on performance, programmability, and energy efficiency, its potential for embedded and predictable real-time workloads remains largely unexplored. In this paper, we start exploring such potential by investigating the real-time capabilities of the first commercial platform based on the AMD Embedded+ architecture: the Sapphire Edge+. We (1) outline key research challenges and real-time use-cases, (2) discuss suitable software architectures for the use-cases and highlight associated trade-offs, and (3) report an initial assessment of the potential of such architectures and use-cases via an experimental evaluation of latency and bandwidth on the real hardware. Lukas Neef, Daniele Ottaviano, Denis Hoornaert, Alexander Züpke, Marco Caccamo, Andrea Bastoni |
RTSS | 5 |
| 2025 | Work-in-Progress: A First Practical Look at Arm's MPAM for Real-Time SystemsabstractArm's Memory Partitioning and Monitoring (MPAM) extension introduces standardized mechanisms for partitioning cache and memory bandwidth. From a real-time systems perspective, this can aid in improving predictability in heterogeneous MPSoCs. In this paper, we present the first practical evaluation of MPAM on a COTS platform—the Radxa Orion O6 with the CIX CD8180 SoC. We characterize the SoC's MPAM capabilities and experimentally assess cache portion partitioning and proportional stride memory bandwidth partitioning under controlled interference. Our results show that enabling MPAM features can reduce interference, but their behavior often diverges from expectations based on the specification, with anomalous effects observed across workloads and cores. These findings highlight both the promise of predictability from MPAM for real-time systems and the current challenges arising from optionality, heterogeneity, and limited documentation. We conclude that broader evaluation across future MPAM-enabled SoCs, aided by detailed performance counter analysis, is essential to establish MPAM's practical value for real-time practitioners. Ashutosh Pradhan, Daniele Ottaviano, Alexander Züpke, Andrea Bastoni, Marco Caccamo |
RTSS | 5 |
| 2025 | Work-in-Progress: Learning to Refine Priority Assignment in Fixed-Priority Real-Time SchedulingabstractWe address the problem of priority assignment for global fixed-priority scheduling on multicore real-time systems, where identifying a feasible priority ordering is a combinatorial challenge. We propose a learning-based framework that trains a lightweight policy network via reinforcement learning to refine existing priority assignments toward schedulable solutions. Based on the policy network, we propose an inference-time policy refinement mechanism that improves schedulability without additional training. It combines breadth sampling—generating candidate orderings via stochastic perturbations—with depth refinement, which iteratively enhances promising candidates. A continuous reward function based on a schedulability hazard metric enables effective training. Preliminary experiments show that the proposed method performs better than classical heuristics such as Deadline Monotonic and DkC, demonstrating its potential as an effective learning-assisted approach to real-time scheduling. Binqi Sun, Linghan Fang, Andrea Bastoni, Marco Caccamo |
RTSS | 4 |
| 2025 | Position paper: deep reinforcement learning for real-time resource managementabstractAbstract Many real-time problems can be characterized as combinatorial optimization problems where exact solutions are infeasible at scale. As problem complexity grows, handcrafted heuristics become increasingly difficult to design. Reinforcement learning (RL) has emerged as a promising alternative, enabling the discovery of decision-making policies without requiring explicit supervision. While RL does not guarantee optimality, it provides adaptive heuristics to solve complex problems. This paper explores the potential of RL for real-time resource management, outlining key principles, demonstrating an application to directed acyclic graph (DAG) scheduling, and identifying open challenges for future research. Mirco Theile, Binqi Sun, Marco Caccamo |
Real Time Syst. | 3 |
| 2025 | SAPar: A Surrogate-Assisted DNN Partitioner for Efficient Inferences on Edge TPU PipelinesabstractPipelining deep neural networks (DNNs) across multiple Edge Tensor Processing Units (TPUs) can enhance on-device performance by increasing the capacity for DNN parameters caching and enabling pipeline parallelism. Effective deployment on pipelined Edge TPUs requires a partitioning tool to divide the DNN into segments, each assigned to a different Edge TPU in the pipeline. Achieving balanced workload distribution across these segments is crucial for optimal timing performance. However, workload balancing across Edge TPUs is challenging, as DNN execution time is influenced by proprietary hardware architecture and compiler internals, forming a black-box function inaccessible to partitioning tools. To address this challenge, this article introduces SAPar , a new surrogate-assisted DNN partitioner that integrates a neighborhood search engine with a surrogate-assisted evaluator for effective and efficient DNN partitioning. The neighborhood search engine systematically explores the decision space, guided by knowledge obtained from empirical insights and neighborhood evaluation feedback provided by the surrogate-assisted evaluator. The evaluator cooperatively applies an accurate yet time-consuming latency profiler and an efficient graph transformer-based surrogate model , achieving both precision and scalability. Experiments on real Edge TPU hardware demonstrate that SAPar achieves significantly better pipeline performance than Google’s current profiling-based partitioner with an 8.82× to 110× speedup in partitioning time. Moreover, SAPar reduces the bottleneck latency by 8.93% to 44.15% across five classic DNN models compared with a state-of-the-art reinforcement learning-based partitioner. Binqi Sun, Bohua Zou, Yigong Hu, Tomasz Kloda, Ling Wang 0001, Tarek F. Abdelzaher, Marco Caccamo |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2025 | Response Time Analysis and Optimal Priority Assignment for Global Non-Preemptive Fixed-Priority Rigid Gang SchedulingabstractNon-preemptive rigid gang scheduling combines the efficiency of parallel execution with the reduced overhead of non-preemptive scheduling. This approach is particularly advantageous for parallel hardware accelerators, such as Google's Edge Tensor Processing Unit (TPU), which is widely used for deep neural network (DNN) inference on embedded systems. This paper studies sporadic global non-preemptive fixed-priority (NP-FP) rigid gang scheduling, which is well-suited for DNN applications in Edge TPU pipelines. Each gang task spawns a fixed number of threads that must execute concurrently across distinct processing units. We introduce the first carry-in limitation technique specifically designed for gang task response time analysis, addressing the unique challenges posed by intra-task parallelism. This technique is formulated as a generalized knapsack problem, and we develop both a linear programming relaxation and a dynamic programming approach to solve it under different time complexities. Additionally, we propose the first optimal priority assignment policy for NP-FP gang schedulability tests. Our proposed schedulability analysis and optimal priority assignment policy are evaluated through extensive experiments, including both synthetic task sets and a case study using DNN benchmarks on commercial off-the-shelf Edge TPU accelerators. The results demonstrate that the proposed approaches effectively enhance the state-of-the-art global NP-FP gang schedulability tests, achieving improvements of up to 57.9% for synthetic task sets and 76.7% for Edge TPU benchmarks. Furthermore, we conduct an ablations study to examine the impact of different algorithmic components in the proposed technique, providing valuable insights for future research. Binqi Sun, Tomasz Kloda, Jiyang Chen, Cen Lu, Marco Caccamo |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2024 | Partitioned Scheduling and Parallelism Assignment for Real-Time DNN Inference Tasks on Multi-TPUabstractPipelining on Edge Tensor Processing Units (TPUs) optimizes the deep neural network (DNN) inference by breaking it down into multiple stages processed concurrently on multiple accelerators. Such DNN inference tasks can be modeled as sporadic non-preemptive gangs with execution times that vary with their parallelism levels. This paper proposes a strict partitioning strategy for deploying DNN inferences in real-time systems. The strategy determines tasks' parallelism levels and assigns tasks to disjoint processor partitions. Configuring the tasks in the same partition with a uniform parallelism level avoids scheduling anomalies and enables schedulability verification using well-understood uniprocessor analyses. Evaluation using real-world Edge TPU benchmarks demonstrated that the proposed method achieves a higher schedulability ratio than state-of-the-art gang scheduling techniques. Binqi Sun, Tomasz Kloda, Chu-Ge Wu, Marco Caccamo |
DAC | 4 |
| 2024 | Response Time Analysis for Fixed-Priority Preemptive Uniform Multiprocessor SystemsabstractWe present a response time analysis for global fixed-priority preemptive scheduling of constrained-deadline tasks upon a uniform multiprocessor where each processor can be characterized by a different speed. A fixed-priority scheduler assigns the jobs with the highest priorities to the fastest processors. Since determining whether all tasks can meet their deadlines is generally intractable even with identical processors, we propose two sufficient schedulability tests that calculate upper bounds on the task’s worst-case response time within polynomial and pseudo-polynomial time. The proposed tests leverage the linear programming model to upper bound the interference of the higher-priority tasks. Furthermore, we identify specific conditions and platforms upon which the problem can be solved more efficiently within linear time. These formulations are used to iteratively evaluate and refine possible solutions until a safe upper bound on the task’s worst-case response time is found. Additionally, we demonstrate that, with specific minor modifications, the proposed tests are compatible with Audsley’s optimal priority assignment. Experimental evaluations performed on synthetic task sets show that the proposed approach outperforms the state-of-the-art methods. Binqi Sun, Tomasz Kloda, Marco Caccamo |
ECRTS | 3 |
| 2024 | Physics-Regulated Deep Reinforcement Learning: Invariant EmbeddingsabstractThis paper proposes the Phy-DRL: a physics-regulated deep reinforcement learning (DRL) framework for safety-critical autonomous systems. The Phy-DRL has three distinguished invariant-embedding designs: i) residual action policy (i.e., integrating data-driven-DRL action policy and physics-model-based action policy), ii) automatically constructed safety-embedded reward, and iii) physics-model-guided neural network (NN) editing, including link editing and activation editing. Theoretically, the Phy-DRL exhibits 1) a mathematically provable safety guarantee and 2) strict compliance of critic and actor networks with physics knowledge about the action-value function and action policy. Finally, we evaluate the Phy-DRL on a cart-pole system and a quadruped robot. The experiments validate our theoretical results and demonstrate that Phy-DRL features guaranteed safety compared to purely data-driven DRL and solely model-based design while offering remarkably fewer learning parameters and fast training towards safety guarantee. Hongpeng Cao, Yanbing Mao, Lui Sha, Marco Caccamo |
ICLR | 4 |
| 2024 | Equivariant Ensembles and Regularization for Reinforcement Learning in Map-based Path PlanningabstractIn reinforcement learning (RL), exploiting environmental symmetries can significantly enhance efficiency, robustness, and performance. However, ensuring that the deep RL policy and value networks are respectively equivariant and invariant to exploit these symmetries is a substantial challenge. Related works try to design networks that are equivariant and invariant by construction, limiting them to a very restricted library of components, which in turn hampers the expressiveness of the networks. This paper proposes a method to construct equivariant policies and invariant value functions without specialized neural network components, which we term equivariant ensembles. We further add a regularization term for adding inductive bias during training. In a map-based path planning case study, we show how equivariant ensembles and regularization benefit sample efficiency and performance. Mirco Theile, Hongpeng Cao, Marco Caccamo, Alberto L. Sangiovanni-Vincentelli |
IROS | 3 |
| 2024 | RaceMOP: Mapless Online Path Planning for Multi-Agent Autonomous Racing using Residual Policy LearningabstractThe interactive decision-making in multi-agent autonomous racing offers insights valuable beyond the domain of self-driving cars. Mapless online path planning is particularly of practical appeal but poses a challenge for safely overtaking opponents due to the limited planning horizon. To address this, we introduce RaceMOP, a novel method for mapless online path planning designed for multi-agent racing of F1TENTH cars. Unlike classical planners that rely on predefined racing lines, RaceMOP operates without a map, utilizing only local observations to execute high-speed overtaking maneuvers. Our approach combines an artificial potential field method as a base policy with residual policy learning to enable long-horizon planning. We advance the field by introducing a novel approach for policy fusion with the residual policy directly in probability space. Extensive experiments on twelve simulated racetracks validate that RaceMOP is capable of long-horizon decision-making with robust collision avoidance during overtaking maneuvers. RaceMOP demonstrates superior handling over existing mapless planners and generalizes to unknown racetracks, affirming its potential for broader applications in robotics. Our code is available at http://github.com/raphajaner/racemop. Raphael Trumpp, Ehsan Javanmardi, Jin Nakazato, Manabu Tsukada, Marco Caccamo |
IROS | 5 |
| 2024 | Strict Partitioning for Sporadic Rigid Gang TasksabstractThe rigid gang task model is based on the idea of executing multiple threads simultaneously on a fixed number of processors to increase efficiency and performance. Although there is extensive literature on global rigid gang scheduling, partitioned approaches have several practical advantages (e.g., task isolation and reduced scheduling overheads). In this paper, we propose a new partitioned scheduling strategy for rigid gang tasks, named strict partitioning. The method creates disjoint partitions of tasks and processors to avoid inter-partition interference. Moreover, it tries to assign tasks with similar volumes (i.e., parallelisms) to the same partition so that the intra-partition interference can be reduced. Within each partition, the tasks can be scheduled using any type of scheduler, which allows the use of a less pessimistic schedulability test. Extensive synthetic experiments and a case study based on Edge TPU benchmarks show that strict partitioning achieves better schedulability performance than state-of-the-art global gang schedulability analyses for both preemptive and non-preemptive rigid gang task sets. Binqi Sun, Tomasz Kloda, Marco Caccamo |
RTAS | 3 |
| 2024 | A Containerized Microservice Architecture for a ROS 2 Autonomous Driving Software: An End-to-End Latency EvaluationabstractThe automotive industry is transitioning from traditional ECU-based systems to software-defined vehicles. A central role of this revolution is played by containers, lightweight virtualization technologies that enable the flexible consolidation of complex software applications on a common hardware platform. Despite their widespread adoption, the impact of containerization on fundamental real-time metrics such as end-to-end latency, communication jitter, as well as memory and CPU utilization has remained virtually unexplored. This paper presents a microservice architecture for a real-world autonomous driving application where containers isolate each service. Our comprehensive evaluation shows the benefits in terms of end-to-end latency of such a solution even over standard bare-Linux deployments. Specifically, in the case of the presented microservice architecture, the mean end-to-end latency can be improved by 5–8%. Also, the maximum latencies were significantly reduced using container deployment. Tobias Betz, Long Wen 0003, Fengjunjie Pan, Gemb Kaljavesi, Alexander Züpke, Andrea Bastoni, Marco Caccamo, Alois C. Knoll, Johannes Betz |
RTCSA | 7 |
| 2024 | Coherence-Aided Memory Bandwidth RegulationabstractWith the increasing adoption of PS-PL (Processor System-Programmable Logic) platforms, also known as CPU+FPGA systems, there arises a need for efficient resource management strategies. This work explores memory bandwidth regulation in such systems, leveraging the capabilities of tightly coupled FPGAs to offer elegant, low-overhead solutions with highly flexible regulation policies. We introduce MemCoRe, a novel approach that exploits the FPGA’s interaction with cache coherence interfaces and cross-trigger signals to achieve finegrained spatiotemporal awareness of processor activity and software-free control. By comparing MemCoRe with state-of-theart software-based approaches, namely MemGuard and MemPol, we demonstrate significant improvements in regulation precision and overhead reduction. Key contributions include nanosecondscale memory bandwidth regulation, off-core memory bandwidth accounting, address-aware regulation, low-overhead token-bucket regulation, and asymmetric on-off core throttling. Our evaluation on a Xilinx Zynq UltraScale+ ZCU102 CPU+FPGA platform showcases MemCoRe’s capability to regulate memory bandwidth with nanosecond-scale precision. Overall, MemCoRe presents a promising avenue for efficient memory bandwidth regulation in PS-PL platforms, with strong applicability to real-time systems. Ivan Izhbirdeev, Denis Hoornaert, Weifan Chen 0003, Alexander Züpke, Youssef Hammad, Marco Caccamo, Renato Mancuso 0001 |
RTSS | 6 |
| 2024 | Mcti: mixed-criticality task-based isolationabstractAbstract The ever-increasing demand for high performance in the time-critical, low-power embedded domain drives the adoption of powerful but unpredictable, heterogeneous Systems-on-Chip. On these platforms, the main source of unpredictability—the shared memory subsystem—has been widely studied, and several approaches to mitigate undesired effects have been proposed over the years. Among them, performance-counter-based regulation methods have proved particularly successful. Unfortunately, such regulation methods require precise knowledge of each task’s memory consumption and cannot be extended to isolate mixed-criticality tasks running on the same core as the regulation budget is shared. Moreover, the desirable combination of these methodologies with well-known time-isolation techniques—such as server-based reservations—is still an uncharted territory and lacks a precise characterization of possible benefits and limitations. Recognizing the importance of such consolidation for designing predictable real-time systems, we introduce MCTI (Mixed-Criticality Task-based Isolation) as a first initial step in this direction. MCTI is a hardware/software co-design architecture that aims to improve both CPU and memory isolations among tasks with different criticalities even when they share the same CPU. In order to ascertain the correct behavior and distill the benefits of MCTI, we implemented and tested the proposed prototype architecture on a widely available off-the-shelf platform. The evaluation of our prototype shows that (1) MCTI helps shield critical tasks from concurrent non-critical tasks sharing the same memory budget, with only a limited increase in response time being observed, and (2) critical tasks running under memory stress exhibit an average response time close to that achieved when running without memory stress. Denis Hoornaert, Golsana Ghaemi, Andrea Bastoni, Renato Mancuso 0001, Marco Caccamo, Giulio Corradi |
Real Time Syst. | 5 |
| 2024 | Minimizing cache usage with fixed-priority and earliest deadline first schedulingabstractAbstract Cache partitioning is a technique to reduce interference among tasks running on the processors with shared caches. To make this technique effective, cache segments should be allocated to tasks that will benefit the most from having their data and instructions stored in the cache. The requests for cached data and instructions can be retrieved faster from the cache memory instead of fetching them from the main memory, thereby reducing overall execution time. The existing partitioning schemes for real-time systems divide the available cache among the tasks to guarantee their schedulability as the sole and primary optimization criterion. However, it is also preferable, particularly in systems with power constraints or mixed criticalities where low- and high-criticality workloads are executing alongside, to reduce the total cache usage for real-time tasks. Cache minimization as part of design space exploration can also help in achieving optimal system performance and resource utilization in embedded systems. In this paper, we develop optimization algorithms for cache partitioning that, besides ensuring schedulability, also minimize cache usage. We consider both preemptive and non-preemptive scheduling policies on single-processor systems with fixed- and dynamic-priority scheduling algorithms ( Rate Monotonic ( RM ) and Earliest Deadline First ( EDF ), respectively). For preemptive scheduling, we formulate the problem as an integer quadratically constrained program and propose an efficient heuristic achieving near-optimal solutions. For non-preemptive scheduling, we combine linear and binary search techniques with different fixed-priority schedulability tests and Quick Processor-demand Analysis (QPA) for EDF. Our experiments based on synthetic task sets with parameters from real-world embedded applications show that the proposed heuristic: (i) achieves an average optimality gap of 0.79% within 0.1× run time of a mathematical programming solver and (ii) reduces average cache usage by 39.15% compared to existing cache partitioning approaches. Besides, we find that for large task sets with high utilization, non-preemptive scheduling can use less cache than preemptive to guarantee schedulability. Binqi Sun, Tomasz Kloda, Sergio Arribas García, Giovani Gracioli, Marco Caccamo |
Real Time Syst. | 5 |
| 2024 | MemPol: polling-based microsecond-scale per-core memory bandwidth regulationabstractAbstract In today’s multiprocessor systems-on-a-chip, the shared memory subsystem is a known source of temporal interference. The problem causes logically independent cores to affect each other’s performance, leading to pessimistic worst-case execution time analysis. Memory regulation via throttling is one of the most practical techniques to mitigate interference. Traditional regulation schemes rely on a combination of timer and performance counter interrupts to be delivered and processed on the same cores running real-time workload. Unfortunately, to prevent excessive overhead, regulation can only be enforced at a millisecond-scale granularity. In this work, we present a novel regulation mechanism from outside the cores that monitors performance counters for the application core’s activity in main memory at a microsecond scale. The approach is fully transparent to the applications on the cores, and can be implemented using widely available on-chip debug facilities. The presented mechanism also allows more complex composition of metrics to enact load-aware regulation. For instance, it allows redistributing unused bandwidth between cores while keeping the overall memory bandwidth of all cores below a given threshold. We implement our approach on a host of embedded platforms and conduct an in-depth evaluation on the Xilinx Zynq UltraScale+ ZCU102, NXP i.MX8M and NXP S32G2 platforms using the San Diego Vision Benchmark Suite. Alexander Züpke, Andrea Bastoni, Weifan Chen 0003, Marco Caccamo, Renato Mancuso 0001 |
Real Time Syst. | 4 |
| 2024 | Perception simplex: Verifiable collision avoidance in autonomous vehicles amidst obstacle detection faultsabstractAbstract Advances in deep learning have revolutionized cyber‐physical applications, including the development of autonomous vehicles. However, real‐world collisions involving autonomous control of vehicles have raised significant safety concerns regarding the use of deep neural networks (DNNs) in safety‐critical tasks, particularly perception. The inherent unverifiability of DNNs poses a key challenge in ensuring their safe and reliable operation. In this work, we propose perception simplex ( ), a fault‐tolerant application architecture designed for obstacle detection and collision avoidance. We analyse an existing LiDAR‐based classical obstacle detection algorithm to establish strict bounds on its capabilities and limitations. Such analysis and verification have not been possible for deep learning‐based perception systems yet. By employing verifiable obstacle detection algorithms, identifies obstacle existence detection faults in the output of unverifiable DNN‐based object detectors. When faults with potential collision risks are detected, appropriate corrective actions are initiated. Through extensive analysis and software‐in‐the‐loop simulations, we demonstrate that provides deterministic fault tolerance against obstacle existence detection faults, establishing a robust safety guarantee. Ayoosh Bansal, Hunmin Kim, Simon Yu, Bo Li 0026, Naira Hovakimyan, Marco Caccamo, Lui Sha |
Softw. Test. Verification Reliab. | 6 |
| 2024 | Edge Generation Scheduling for DAG Tasks Using Deep Reinforcement LearningabstractDirected acyclic graph (DAG) tasks are currently adopted in the real-time domain to model complex applications from the automotive, avionics, and industrial domains that implement their functionalities through chains of intercommunicating tasks. This paper studies the problem of scheduling real-time DAG tasks by presenting a novel schedulability test based on the concept oftrivial schedulability. Using this schedulability test, we propose a new DAG scheduling framework (edge generation scheduling—EGS) that attempts to minimize the DAG width by iteratively generating edges while guaranteeing the deadline constraint. We study how to efficiently solve the problem of generating edges by developing a deep reinforcement learning algorithm combined with a graph representation neural network to learn an efficient edge generation policy for EGS. We evaluate the effectiveness of the proposed algorithm by comparing it with state-of-the-art DAG scheduling heuristics and an optimal mixed-integer linear programming baseline. Experimental results show that the proposed algorithm outperforms the state-of-the-art by requiring fewer processors to schedule the same DAG tasks.https://github.com/binqi-sun/egs Binqi Sun, Mirco Theile, Ziyuan Qin 0002, Daniele Bernardini 0002, Debayan Roy, Andrea Bastoni, Marco Caccamo |
IEEE Trans. Computers | 7 |
| 2023 | Towards Safe AI: Sandboxing DNNs-Based Controllers in Stochastic GamesabstractNowadays, AI-based techniques, such as deep neural networks (DNNs), are widely deployed in autonomous systems for complex mission requirements (e.g., motion planning in robotics). However, DNNs-based controllers are typically very complex, and it is very hard to formally verify their correctness, potentially causing severe risks for safety-critical autonomous systems. In this paper, we propose a construction scheme for a so-called Safe-visor architecture to sandbox DNNs-based controllers. Particularly, we consider the construction under a stochastic game framework to provide a system-level safety guarantee which is robust to noises and disturbances. A supervisor is built to check the control inputs provided by a DNNs-based controller and decide whether to accept them. Meanwhile, a safety advisor is running in parallel to provide fallback control inputs in case the DNN-based controller is rejected. We demonstrate the proposed approaches on a quadrotor employing an unverified DNNs-based controller. Bingzhuo Zhong, Hongpeng Cao, Majid Zamani 0001, Marco Caccamo |
AAAI | 4 |
| 2023 | Flexible Gear Assembly with Visual Servoing and Force FeedbackabstractThis paper presents a vision-guided two-stage approach with force feedback to achieve high-precision and flexible gear assembly. The proposed approach integrates YOLO to coarsely localize the target workpiece in a searching phase and deep reinforcement learning (DRL) to complete the insertion. Specifically, DRL addresses the challenge of partial visibility when the on-wrist camera is too close to the workpiece of a small size. Moreover, we use force feedback to improve the robustness of the vision-guided assembly process. To reduce the effort of collecting training data on real robots, we use synthetic RGB images for training YOLO and construct an offline interaction environment leveraging sampled real-world data for training DRL agents. The proposed approach was evaluated in an industrial gear assembly experiment, which requires an assembly clearance of 0.3 mm, demonstrating high robustness and efficiency in gear searching and insertion from arbitrary positions. Junjie Ming, Daniel Bargmann, Hongpeng Cao, Marco Caccamo |
IROS | 4 |
| 2023 | Residual Policy Learning for Vehicle Control of Autonomous Racing CarsabstractThe development of vehicle controllers for autonomous racing is challenging because racing cars operate at their physical driving limit. Prompted by the demand for improved performance, autonomous racing research has seen the proliferation of machine learning-based controllers. While these approaches show competitive performance, their practical applicability is often limited. Residual policy learning promises to mitigate this drawback by combining classical controllers with learned residual controllers. The critical advantage of residual controllers is their high adaptability parallel to the classical controller’s stable behavior. We propose a residual vehicle controller for autonomous racing cars that learns to amend a classical controller for the path-following of racing lines. In an extensive study, performance gains of our approach are evaluated for a simulated car of the F1TENTH autonomous racing series. The evaluation for twelve replicated real-world racetracks shows that the residual controller reduces lap times by an average of 4.55 % compared to a classical controller and even enables lap time gains on unknown racetracks. Raphael Trumpp, Denis Hoornaert, Marco Caccamo |
IV | 3 |
| 2023 | Schedulability Analysis of Non-preemptive Sporadic Gang Tasks on Hardware AcceleratorsabstractNon-preemptive rigid gang scheduling combines the performance benefits of parallel execution with the low overhead of non-preemptive scheduling and rigid task programming model. This approach appears particularly well-suited for parallel hardware accelerators where the context switch and migration overheads are critical and should be avoided. One of the most notable examples today is Google's Edge Tensor Processing Unit (TPU) used for neural network inference on embedded boards. The paper studies sporadic non-preemptive rigid gang scheduling applied to multi-TPU edge AI accelerators. Each gang task spawns a fixed number of threads that must execute simultaneously on distinct processing units. We consider non-preemptive fixed-priority gang (NP-FP-Gang) scheduling and propose the first carry-in limitation for gang task response time analysis. The gang task carry-in limitation differs from conventional sequential tasks due to the intra-task parallelism. We formulate it as a generalized knapsack problem and develop a linear programming relaxation and a dynamic programming approach to solve the problem under different time complexities. The performance of the proposed schedulability analysis is evaluated through randomly generated synthetic task sets and a case study using neural network benchmarks executed on commercial off-the-shelf multi-TPU edge AI accelerators. The evaluation results show that the proposed response time analysis effectively improves the state of-the-art NP-FP-Gang schedulability test even by 85.7% for the Edge TPU benchmarks in particular. Binqi Sun, Tomasz Kloda, Jiyang Chen, Cen Lu, Marco Caccamo |
RTAS | 5 |
| 2023 | MemPol: Policing Core Memory Bandwidth from Outside of the CoresabstractIn today’s multiprocessor systems-on-a-chip (MP- SoC), the shared memory subsystem is a known source of temporal interference. The problem causes logically independent cores to affect each other’s performance, leading to pessimistic worstcase execution time (WCET) analysis. One of the most practical techniques to mitigate interference is memory regulation via throttling. Traditional regulation schemes rely on a combination of timer and performance counter interrupts to be delivered and processed on the same cores running real-time workload. Unfortunately, to prevent excessive overhead, regulation can only be enforced at a millisecond-scale granularity. In this work, we present a novel regulation mechanism from outside the cores that monitors performance counters for the application core’s activity in main memory at a microsecond scale. The approach is fully transparent to the applications on the cores, and can be implemented using widely available onchip debug facilities. The presented mechanism also allows more complex composition of metrics to enact load-aware regulation. For instance, it allows redistributing unused bandwidth between cores while keeping the overall memory bandwidth of all cores below a given threshold. We implement our approach on a host of embedded platforms and carry out an in-depth evaluation on the Xilinx Zynq UltraScale+ZCUl02 platform using the SD-VBS. Alexander Züpke, Andrea Bastoni, Weifan Chen 0003, Marco Caccamo, Renato Mancuso 0001 |
RTAS | 4 |
| 2023 | Co-Optimizing Cache Partitioning and Multi-Core Task Scheduling: Exploit Cache Sensitivity or Not?abstractCache partitioning techniques have been successfully adopted to mitigate interference among concurrently executing real-time tasks on multi-core processors. Considering that the execution time of a cache-sensitive task strongly depends on the cache available for it to use, co-optimizing cache partitioning and task allocation improves the system's schedulability. In this paper, we propose a hybrid multi-layer design space exploration technique to solve this multi-resource management problem. We explore the interplay between cache partitioning and schedulability by systematically interleaving three optimization layers, viz., (i) in the outer layer, we perform a breadth-first search combined with proactive pruning for cache partitioning; (ii) in the middle layer, we exploit a first-fit heuristic for allocating tasks to cores; and (iii) in the inner layer, we use the well-known recurrence relation for the schedulability analysis of non-preemptive fixed-priority (NP-FP) tasks in a uniprocessor setting. Although our focus is on NP-FP scheduling, we evaluate the flexibility of our framework in supporting different scheduling policies (NP-EDF, P-EDF) by plugging in appropriate analysis methods in the inner layer. Experiments show that, compared to the state-of-the-art techniques, the proposed framework can improve the real-time schedulability of NP-FP task sets by an average of 15.2% with a maximum improvement of 233.6% (when tasks are highly cache-sensitive) and a minimum of 1.6% (when cache sensitivity is low). For such task sets, we found that clustering similar- period (or mutually compatible) tasks often leads to higher schedulability (on average 7.6 %) than clustering by cache sensitivity. In our evaluation, the framework also achieves good results for preemptive and dynamic-priority scheduling policies. Binqi Sun, Debayan Roy, Tomasz Kloda, Andrea Bastoni, Rodolfo Pellizzoni, Marco Caccamo |
RTSS | 6 |
| 2023 | X-Stream: Accelerating streaming segments on MPSoCs for real-time applications
Rohan Tabish, Rodolfo Pellizzoni, Renato Mancuso 0001, Giovani Gracioli, Reza Mirosanlou, Marco Caccamo |
J. Syst. Archit. | 6 |
| 2023 | SchedGuard++: Protecting against Schedule Leaks Using Linux Containers on Multi-Core ProcessorsabstractTiming correctness is crucial in a multi-criticality real-time system, such as an autonomous driving system. It has been recently shown that these systems can be vulnerable to timing inference attacks, mainly due to their predictable behavioral patterns. Existing solutions like schedule randomization cannot protect against such attacks, often limited by the system’s real-time nature. This article presents “ SchedGuard++ ”: a temporal protection framework for Linux-based real-time systems that protects against posterior schedule-based attacks by preventing untrusted tasks from executing during specific time intervals. SchedGuard++ supports multi-core platforms and is implemented using Linux containers and a customized Linux kernel real-time scheduler. We provide schedulability analysis assuming the Logical Execution Time (LET) paradigm, which enforces I/O predictability. The proposed response time analysis takes into account the interference from trusted and untrusted tasks and the impact of the protection mechanism. We demonstrate the effectiveness of our system using a realistic radio-controlled rover platform. Not only is “ SchedGuard++ ” able to protect against the posterior schedule-based attacks, but it also ensures that the real-time tasks/containers meet their temporal requirements. Jiyang Chen, Tomasz Kloda, Rohan Tabish, Ayoosh Bansal, Chien-Ying Chen, Bo Liu 0044, Sibin Mohan, Marco Caccamo, Lui Sha |
ACM Trans. Cyber Phys. Syst. | 8 |
| 2023 | Lazy Load Scheduling for Mixed-criticality Applications in Heterogeneous MPSoCsabstractNewly emerging multiprocessor system-on-a-chip (MPSoC) platforms provide hard processing cores with programmable logic (PL) for high-performance computing applications. In this article, we take a deep look into these commercially available heterogeneous platforms and show how to design mixed-criticality applications such that different processing components can be isolated to avoid contention on the shared resources such as last-level cache and main memory. Our approach involves software/hardware co-design to achieve isolation between the different criticality domains. At the hardware level, we use a scratchpad memory (SPM) with dedicated interfaces inside the PL to avoid conflicts in the main memory. At the software level, we employ a hypervisor to support cache-coloring such that conflicts at the shared L2 cache can be avoided. In order to move the tasks in/out of the SPM memory, we rely on a DMA engine and propose a new CPU-DMA co-scheduling policy, called Lazy Load , for which we also derive the response time analysis. The results of a case study on image processing demonstrate that the contention on the shared memory subsystem can be avoided when running with our proposed architecture. Moreover, comprehensive schedulability evaluations show that the newly proposed Lazy Load policy outperforms the existing CPU-DMA scheduling approaches and is effective in mitigating the main memory interference in our proposed architecture. Tomasz Kloda, Giovani Gracioli, Rohan Tabish, Reza Mirosanlou, Renato Mancuso 0001, Rodolfo Pellizzoni, Marco Caccamo |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2022 | Latency analysis of self-suspending task chainsabstractMany cyber-physical systems are offloading computation-heavy programs to hardware accelerators (e.g., GPU and TPU) to reduce execution time. These applications will self-suspend between offloading data to the accelerators and obtaining the returned results. Previous efforts have shown that self-suspending tasks can cause scheduling anomalies, but none has examined inter-task communication. This paper aims to explore self-suspending tasks' data chain latency with periodic activation and asynchronous message passing. We first present the cause for suspension-induced delays and worst-case latency analysis. We then propose a rule for utilizing the hardware co-processors to reduce data chain latency and schedulability analysis. Simulation results show that the proposed strategy can improve overall latency while preserving system schedulability. Tomasz Kloda, Jiyang Chen, Antoine Bertout, Lui Sha, Marco Caccamo |
DATE | 5 |
| 2022 | Reconciling QoS and Concurrency in NVIDIA GPUs via Warp-Level SchedulingabstractThe widespread deployment of NVIDIA GPUs in latency-sensitive systems today requires predictable GPU multi-tasking, which cannot be trivially achieved. The NVIDIA CUDA API allows programmers to easily exploit the processing power provided by these massively parallel accelerators and is one of the major reasons behind their ubiquity. However, NVIDIA GPUs and the CUDA programming model favor throughput instead of latency and timing predictability. Hence, providing real-time and quality-of-service (QoS) properties to GPU applications presents an interesting research challenge. Such a challenge is paramount when considering simultaneous multikernel (SMK) scenarios, wherein kernels are executed concurrently within each streaming multiprocessor (SM). In this work, we explore QoS-based fine-grained multitasking in SMK via job arbitration at the lowest level of the GPU scheduling hierarchy, i.e., between warps. We present QoS-aware warp scheduling (QAWS) and evaluate it against state-of-the-art, kernel-agnostic policies seen in NVIDIA hardware today. Since the NVIDIA ecosystem lacks a mechanism to specify and enforce kernel priority at the warp granularity, we implement and evaluate our proposed warp scheduling policy on GPGPU-Sim. QAWS not only improves the response time of the higher priority tasks but also has comparable or better throughput than the state-of-the-art policies. Jayati Singh, Ignacio Sanudo Olmedo, Nicola Capodieci, Andrea Marongiu, Marco Caccamo |
DATE | 5 |
| 2022 | Memory allocation for low-power real-time embedded microcontroller: a case studyabstractMemory allocation of instructions and data can affect the program execution speed. This paper tests various memory-intensive benchmarks under different memory allocations on a Cortex-M4-based microcontroller and solves the allocation problem using integer linear programming. Zhishen Zhang, Yuwen Shen, Binqi Sun, Tomasz Kloda, Marco Caccamo |
ETFA | 5 |
| 2022 | Poster Abstract: Controller Synthesis for Nonlinear Stochastic Games via Approximate Probabilistic RelationsabstractNo abstract available. Bingzhuo Zhong, Abolfazl Lavaei, Majid Zamani 0001, Marco Caccamo |
HSCC | 4 |
| 2022 | Cloud-Edge Training Architecture for Sim-to-Real Deep Reinforcement LearningabstractDeep reinforcement learning (DRL) is a promising approach to solve complex control tasks by learning policies through interactions with the environment. However, the training of DRL policies requires large amounts of training experiences, making it impractical to learn the policy directly on physical systems. Sim-to-real approaches leverage simulations to pretrain DRL policies and then deploy them in the real world. Unfortunately, the direct real-world deployment of pretrained policies usually suffers from performance deterioration due to the different dynamics, known as the reality gap. Recent sim-to-real methods, such as domain randomization and domain adaptation, focus on improving the robustness of the pretrained agents. Nevertheless, the simulation-trained policies often need to be tuned with real-world data to reach optimal performance, which is challenging due to the high cost of real-world samples. This work proposes a distributed cloud-edge architecture to train DRL agents in the real world in real-time. In the architecture, the inference and training are assigned to the edge and cloud, separating the real-time control loop from the computationally expensive training loop. To overcome the reality gap, our architecture exploits sim-to-real transfer strategies to continue the training of simulation-pretrained agents on a physical system. We demonstrate its applicability on a physical inverted-pendulum control system, analyzing critical parameters. The real-world experiments show that our architecture can adapt the pretrained DRL agents to unseen dynamics consistently and efficiently.11A video showing a real-world training process under the proposed method can be found from https://youtu.be/hMY9-c0SST0. Hongpeng Cao, Mirco Theile, Federico G. Wyrwal, Marco Caccamo |
IROS | 4 |
| 2022 | Verifiable Obstacle DetectionabstractPerception of obstacles remains a critical safety concern for autonomous vehicles. Real-world collisions have shown that the autonomy faults leading to fatal collisions originate from obstacle existence detection. Open source autonomous driving implementations show a perception pipeline with complex interdependent Deep Neural Networks. These networks are not fully verifiable, making them unsuitable for safety-critical tasks. In this work, we present a safety verification of an existing LiDAR based classical obstacle detection algorithm. We establish strict bounds on the capabilities of this obstacle detection algorithm. Given safety standards, such bounds allow for determining LiDAR sensor properties that would reliably satisfy the standards. Such analysis has as yet been unattainable for neural network based perception systems. We provide a rigorous analysis of the obstacle detection system with empirical results based on real-world sensor data. Ayoosh Bansal, Hunmin Kim, Simon Yu, Bo Li 0026, Naira Hovakimyan, Marco Caccamo, Lui Sha |
ISSRE | 6 |
| 2021 | Flexible Cache Partitioning for Multi-Mode Real-Time SystemsabstractCache partitioning is a well-studied technique that mitigates the inter-processor cache interference in multiprocessor systems. The resulting optimization problem involves allocating portions of the cache to individual processors. In multi-mode applications (e.g., flight control system that runs in take-off, cruise, or landing mode), the cache memory requirement can change over time, making runtime cache repartitioning necessary. This paper presents a cache partition allocation framework enabling flexible cache partitioning for multi-mode real-time systems. The main objective is to guarantee timing predictability in the steady states and during mode changes. We evaluate the effectiveness of our approach for multiple embedded benchmarks with different ranges of cache size sensitivity. The results show increased schedulability compared to static partitioning approaches. Ohchul Kwon, Gero Schwäricke, Tomasz Kloda, Denis Hoornaert, Giovani Gracioli, Marco Caccamo |
DATE | 6 |
| 2021 | Timing Debugging for Cyber-Physical SystemsabstractThis paper is concerned with the following question: Given a set of control tasks that are not schedulable, i.e., their required timing properties cannot be satisfied, what should be changed? While the real-time systems literature proposes many different schedulability analysis techniques, it surprisingly provides almost no guidelines on what should be changed to make a task set schedulable, when it is not. We show that when the tasks in question are control tasks, this timing debugging question in the context of cyber-physical systems (CPS) may be answered by exploiting the dynamics of the physical systems that these control tasks are expected to influence. Towards this, we study a very simple setup, viz., when a set of periodic tasks with implicit deadlines is not schedulable, by how much should the periods be changed in order to make the task set schedulable? Among the many ways in which the periods can be modified, our proposed strategy is to change the periods in a manner such that while the task set becomes schedulable, the poles of the closed-loop system experience the minimal shift. Since the poles influence the closed loop dynamics of the system, we thereby ensure that we obtain a system with the desired timing properties whose dynamics is very similar to the dynamics of the original (non-schedulable) system. We formulate this CPS timing debugging strategy as an optimization problem and illustrate it with a concrete example. Debayan Roy, Clara Hobbs, James H. Anderson, Marco Caccamo, Samarjit Chakraborty |
DATE | 4 |
| 2021 | SchedGuard: Protecting against Schedule Leaks Using Linux ContainersabstractReal-time systems have recently been shown to be vulnerable to timing inference attacks, mainly due to their predictable behavioral patterns. Existing solutions such as schedule randomization lack the ability to protect against such attacks, often limited by the system's real-time nature. This paper presents “SchedGuard”: a temporal protection framework for Linux-based hard real-time systems that protects against posterior scheduler side-channel attacks by preventing untrusted tasks from executing during specific time segments. SchedGuard is integrated into the Linux kernel using cgroups, making it amenable to use with container frameworks. We demonstrate the effectiveness of our system using a realistic radio-controlled rover platform and synthetically generated workloads. Not only is SchedGuard able to protect against the attacks mentioned above, but it also ensures that the real-time tasks/containers meet their temporal requirements. Jiyang Chen, Tomasz Kloda, Ayoosh Bansal, Rohan Tabish, Chien-Ying Chen, Bo Liu 0044, Sibin Mohan, Marco Caccamo, Lui Sha |
RTAS | 8 |
| 2021 | Work in Progress: Identifying Unexpected Inter-core Interference Induced by Shared CacheabstractIn modern real-time multicore systems, understanding and adequately managing shared caches is essential to ensure the temporal isolation of critical tasks. Recent research has identified and extensively studied the sources of unpredictability imputable to shared caches, heavily promoting techniques such as cache partitioning and internal resources management. In this article, we highlight the existence of an enigmatic source of inter-core interference: the CPU-brainfreeze. Experiments realized on a development board show that benchmarks (selected from the San-Diego Vision Benchmark Suite) can exhibit up to a 10-fold increase in their execution time. The same experiment shows that for extreme cases, the core cluster can be stalled indefinitely. Denis Hoornaert, Shahin Roozkhosh, Renato Mancuso 0001, Marco Caccamo |
RTAS | 4 |
| 2021 | A Real-Time Virtio-Based Framework for Predictable Inter-VM CommunicationabstractEnsuring real-time properties on current heterogeneous multiprocessor systems on a chip is a challenging task. Furthermore, online artificial intelligent applications –which are routinely deployed on such chips– pose increasing pressure on the memory subsystem that becomes a source of unpredictability. Although techniques have been proposed to restore independent access to memory for concurrently executing virtual machines (VM), providing predictable inter-VM communication remains challenging. In this work, we tackle the problem of predictably transferring data between virtual machines and virtualized hardware resources on multiprocessor systems on chips under consideration of memory interference. We design a "broker-based" real-time communication framework for otherwise isolated virtual machines, provide a virtio-based reference implementation on top of the Jailhouse hypervisor, assess its overheads for FreeRTOS virtual machines, and formally analyze its communication flow schedulability under consideration of the implementation overheads. Furthermore, we define a methodology to assess the maximum DRAM memory saturation empirically, evaluate the framework’s performance and compare it with the theoretical schedulability. Gero Schwäricke, Rohan Tabish, Rodolfo Pellizzoni, Renato Mancuso 0001, Andrea Bastoni, Alexander Züpke, Marco Caccamo |
RTSS | 7 |
| 2021 | An Analyzable Inter-core Communication Framework for High-Performance Multicore Embedded Systems
Rohan Tabish, Jen-Yang Wen, Rodolfo Pellizzoni, Renato Mancuso 0001, Heechul Yun, Marco Caccamo, Lui Sha |
J. Syst. Archit. | 6 |
| 2020 | Fixed-Priority Memory-Centric Scheduler for COTS-Based MultiprocessorsabstractMemory-centric scheduling attempts to guarantee temporal predictability on commercial-off-the-shelf (COTS) multiprocessor systems to exploit their high performance for real-time applications. Several solutions proposed in the real-time literature have hardware requirements that are not easily satisfied by modern COTS platforms, like hardware support for strict memory partitioning or the presence of scratchpads. However, even without said hardware support, it is possible to design an efficient memory-centric scheduler. In this article, we design, implement, and analyze a memory-centric scheduler for deterministic memory management on COTS multiprocessor platforms without any hardware support. Our approach uses fixed-priority scheduling and proposes a global "memory preemption" scheme to boost real-time schedulability. The proposed scheduling protocol is implemented in the Jailhouse hypervisor and Erika real-time kernel. Measurements of the scheduler overhead demonstrate the applicability of the proposed approach, and schedulability experiments show a 20% gain in terms of schedulability when compared to contention-based and static fair-share approaches. Gero Schwäricke, Tomasz Kloda, Giovani Gracioli, Marko Bertogna, Marco Caccamo |
ECRTS | 5 |
| 2020 | UAV Path Planning for Wireless Data Harvesting: A Deep Reinforcement Learning ApproachabstractAutonomous deployment of unmanned aerial vehicles (UAVs) supporting next-generation communication networks requires efficient trajectory planning methods. We propose a new end-to-end reinforcement learning (RL) approach to UAV-enabled data collection from Internet of Things (IoT) devices in an urban environment. An autonomous drone is tasked with gathering data from distributed sensor nodes subject to limited flying time and obstacle avoidance. While previous approaches, learning and non-learning based, must perform expensive recomputations or relearn a behavior when important scenario parameters such as the number of sensors, sensor positions, or maximum flying time, change, we train a double deep Q-network (DDQN) with combined experience replay to learn a UAV control policy that generalizes over changing scenario parameters. By exploiting a multi-layer map of the environment fed through convolutional network layers to the agent, we show that our proposed network architecture enables the agent to make movement decisions for a variety of scenario parameters that balance the data collection goal with flight time efficiency and safety constraints. Considerable advantages in learning efficiency from using a map centered on the UAV's position over a non-centered map are also illustrated. Harald Bayerlein, Mirco Theile, Marco Caccamo, David Gesbert |
GLOBECOM | 3 |
| 2020 | UAV Coverage Path Planning under Varying Power Constraints using Deep Reinforcement LearningabstractCoverage path planning (CPP) is the task of designing a trajectory that enables a mobile agent to travel over every point of an area of interest. We propose a new method to control an unmanned aerial vehicle (UAV) carrying a camera on a CPP mission with random start positions and multiple options for landing positions in an environment containing no-fly zones. While numerous approaches have been proposed to solve similar CPP problems, we leverage end-to-end reinforcement learning (RL) to learn a control policy that generalizes over varying power constraints for the UAV. Despite recent improvements in battery technology, the maximum flying range of small UAVs is still a severe constraint, which is exacerbated by variations in the UAV's power consumption that are hard to predict. By using map-like input channels to feed spatial information through convolutional network layers to the agent, we are able to train a double deep Q-network (DDQN) to make control decisions for the UAV, balancing limited power budget and coverage goal. The proposed method can be applied to a wide variety of environments and harmonizes complex goal structures with system constraints. Mirco Theile, Harald Bayerlein, Richard Nai, David Gesbert, Marco Caccamo |
IROS | 5 |
| 2020 | Latency-Aware Generation of Single-Rate DAGs from Multi-Rate Task SetsabstractModern automotive and avionics embedded systems integrate several functionalities that are subject to complex timing requirements. A typical application in these fields is composed of sensing, computation, and actuation. The ever increasing complexity of heterogeneous sensors implies the adoption of multi-rate task models scheduled onto parallel platforms. Aspects like freshness of data or first reaction to an event are crucial for the performance of the system. The Directed Acyclic Graph (DAG) is a suitable model to express the complexity and the parallelism of these tasks. However, deriving age and reaction timing bounds is not trivial when DAG tasks have multiple rates. In this paper, a method is proposed to convert a multi-rate DAG task-set with timing constraints into a single-rate DAG that optimizes schedulability, age and reaction latency, by inserting suitable synchronization constructs. An experimental evaluation is presented for an autonomous driving benchmark, validating the proposed approach against state-of-the-art solutions. Micaela Verucchi, Mirco Theile, Marco Caccamo, Marko Bertogna |
RTAS | 3 |
| 2020 | GoodSpread: Criticality-Aware Static Scheduling of CPS with Multi-QoS ResourcesabstractIn practice, safety-critical cyber-physical systems (CPS) are often implemented using high quality-of-service (QoS) resources to provide maximum performance in all scenarios. Such implementations are oblivious to the changing criticality levels of CPS based on their physical dynamics (e.g., steady or transient state). Considering that high-QoS resources are constrained for cost-sensitive CPS, such criticality-oblivious implementations are highly inefficient. Towards a tighter dimensioning of these resources, state-of-the-art approaches have considered multi-QoS resources and studied criticality-aware dynamic resource allocation along the lines of mixed-criticality systems. However, these approaches have high implementation overheads. Moreover, in safety-critical domains like automotive and avionics, certification of such dynamic policies is challenging and the implementation platforms typically do not support dynamic reconfiguration. To address these challenges, we present GoodSpread that uses a static scheduling strategy and offers the same performance guarantees while saving resources (more than 50 % in certain cases) compared to the existing dynamic schemes. The main idea here is to spread the high-QoS resources as uniformly as possible over time in order to accommodate the uncertainty of when the criticality level might change. Our proposed strategy studies the physical dynamics to determine the spread factor, i.e., how often the high-QoS resources need to be provisioned. We further propose an extensibility-driven optimization approach to obtain a static schedule that will accommodate future workloads on the remaining resources with maximum flexibility. Debayan Roy, Sumana Ghosh, Qi Zhu 0002, Marco Caccamo, Samarjit Chakraborty |
RTSS | 4 |
| 2020 | Software Fault Tolerance for Cyber-Physical Systems via Full System RestartabstractThe article addresses the issue of reliability of complex embedded control systems in the safety-critical environment. In this article, we propose a novel approach to design controller that (i) guarantees the safety of nonlinear physical systems, (ii) enables safe system restart during runtime, and (iii) allows the use of complex, unverified controllers (e.g., neural networks) that drive the physical systems toward complex specifications. We use abstraction-based controller synthesis approach to design a formally verified controller that provides application and system-level fault tolerance along with safety guarantee. Moreover, our approach is implementable using a commercial-off-the-shelf (COTS) processing unit. To demonstrate the efficacy of our solution and to verify the safety of the system under various types of faults injected in applications and in the underlying real-time operating system (RTOS), we implemented the proposed controller for the inverted pendulum and three degrees-of-freedom (3-DOF) helicopter. Pushpak Jagtap, Fardin Abdi Taghi Abad, Matthias Rungger, Majid Zamani 0001, Marco Caccamo |
ACM Trans. Cyber Phys. Syst. | 5 |
| 2019 | Designing Mixed Criticality Applications on Modern Heterogeneous MPSoC PlatformsabstractMultiprocessor Systems-on-Chip (MPSoC) integrating hard processing cores with programmable logic (PL) are becoming increasingly common. While these platforms have been originally designed for high performance computing applications, their rich feature set can be exploited to efficiently implement mixed criticality domains serving both critical hard real-time tasks, as well as soft real-time tasks. In this paper, we take a deep look at commercially available heterogeneous MPSoCs that incorporate PL and a multicore processor. We show how one can tailor these processors to support a mixed criticality system, where cores are strictly isolated to avoid contention on shared resources such as Last-Level Cache (LLC) and main memory. In order to avoid conflicts in last-level cache, we propose the use of cache coloring, implemented in the Jailhouse hypervisor. In addition, we employ ScratchPad Memory (SPM) inside the PL to support a multi-phase execution model for real-time tasks that avoids conflicts in shared memory. We provide a full-stack, working implementation on a latest-generation MPSoC platform, and show results based on both a set of data intensive tasks, as well as a case study based on an image processing benchmark application. Giovani Gracioli, Rohan Tabish, Renato Mancuso 0001, Reza Mirosanlou, Rodolfo Pellizzoni, Marco Caccamo |
ECRTS | 6 |
| 2019 | Trajectory Estimation for Geo-Fencing Applications on Small-Size Fixed-Wing UAVsabstractThe steadily increasing popularity of Unmanned Aerial Vehicles (UAVs) is creating new opportunities in diverse fields of technology and business. However, this increase of popularity also raises safety concerns. To tackle the primary concern of keeping the UAV inside a designated region, a novel trajectory estimation algorithm for geo-fencing applications is proposed. We derive the Beta-Trajectory that takes into account constraints in curvature as well as constraints in the change of curvature which is bounded by the maximum roll-rate of the aircraft. We incorporate the Beta-Trajectory into a geo-fencing algorithm. By using our open-source uavAP autopilot, the applicability and necessity of accurate trajectory estimation algorithms for geo-fencing applications are shown on small fixed-wing aircraft. The model and algorithm are validated in high-fidelity simulations as well as in real flight testing. Mirco Theile, Simon Yu, Or D. Dantsker, Marco Caccamo |
IROS | 4 |
| 2019 | Segment Streaming for the Three-Phase Execution Model: Design and ImplementationabstractScheduling tasks using the three-phase execution model (load-execute-unload) can effectively reduce the contention on shared resources in real-time systems. Due to system and program constraints, a task is generally segmented and executed over multiple intervals. Several works showed that co-scheduling memory (unload-load) and computation phases can improve the system schedulability by hiding the memory transfer time. However, this is limited to segments of different tasks and hence executing segments of the same task back-to-back is not allowed. In this paper, we propose a new streaming model to allow overlapping the memory and execution phases of segments of the same task. This is accomplished by a segmentation framework implemented within an LLVM-based compiler-level tool along with a Real-Time Operating System (RTOS) API to handle load/unload requests. Memory phases are processed by a DMA engine that loads/unloads the task content into ScratchPad Memory (SPM). We provide a schedulability analysis of the proposed model under fixed priority partitioned scheme and an RTOS implementation of the API on a latest-generation Multiprocessor System-on-Chip (MPSoC). Muhammad Refaat Soliman, Giovani Gracioli, Rohan Tabish, Rodolfo Pellizzoni, Marco Caccamo |
RTSS | 5 |
| 2019 | Preserving Physical Safety Under Cyber AttacksabstractPhysical plants that form the core of the cyber-physical systems (CPSs) often have stringent safety requirements and, recent attacks have shown that cyber intrusions can cause damage to these plant. In this paper, we demonstrate how to ensure the safety of the physical plant even when the platform is compromised. We leverage the fact that due to physical inertia, an adversary cannot destabilize the plant (even with complete control over the software) instantaneously. In fact, it often takes finite (even considerable time). This paper provides the analytical framework that utilizes this property to compute safe operational windows in run-time during which the safety of the plant is guaranteed. To ensure the correctness of the computations in runtime, we discuss two approaches to ensure the integrity of these computations in an untrusted environment: 1) full platform-wide restarts coupled with a root-of-trust timer and 2) utilizing trusted execution environment features available in hardware. We demonstrate our approach using two realistic systems-a 3 degree-of-freedom helicopter and a simulated warehouse temperature management unit and show that our system is robust against multiple emulated attacks-essentially the attackers are not able to compromise the safety of the CPS. Fardin Abdi Taghi Abad, Chien-Ying Chen, Monowar Hasan, Songran Liu, Sibin Mohan, Marco Caccamo |
IEEE Internet Things J. | 6 |
| 2019 | A real-time scratchpad-centric OS with predictable inter/intra-core communication for multi-core embedded systems
Rohan Tabish, Renato Mancuso 0001, Saud Wasly, Rodolfo Pellizzoni, Marco Caccamo |
Real Time Syst. | 5 |
| 2018 | uavEE: A Modular, Power-Aware Emulation Environment for Rapid Prototyping and Testing of UAVsabstractState of the art design and testing of avionics for unmanned aircraft is an iterative process that involves many test flights, interleaved with multiple revisions of the flight management software and hardware. To significantly reduce flight test time and software development costs, we have developed a real-time UAV Emulation Environment (uavEE) using ROS that interfaces with high fidelity simulators to simulate the flight behavior of the aircraft. Our uavEE emulates the avionics hardware by interfacing directly with the embedded hardware used in real flight. The modularity of uavEE allows the integration of countless test scenarios and applications. Furthermore, we present an accurate data driven approach for modeling of propulsion power of fixed-wing UAVs, which is integrated into uavEE. Finally, uavEE and the proposed UAV Power Model have been experimentally validated using a fixed-wing UAV testbed. Mirco Theile, Or D. Dantsker, Richard Nai, Marco Caccamo |
RTCSA | 4 |
| 2017 | WCET Derivation under Single Core Equivalence with Explicit Memory Budget AssignmentabstractIn the last decade there has been a steady uptrend in the popularity of embedded multi-core platforms. This represents a turning point in the theory and implementation of real-time systems. From a real-time standpoint, however, the extensive sharing of hardware resources (e.g. caches, DRAM subsystem, I/O channels) represents a major source of unpredictability. Budget-based memory regulation (throttling) has been extensively studied to enforce a strict partitioning of the DRAM subsystem’s bandwidth. The common approach to analyze a task under memory bandwidth regulation is to consider the budget of the core where the task is executing, and assume the worst-case about the remaining cores' budgets. In this work, we propose a novel analysis strategy to derive the WCET of a task under memory bandwidth regulation that takes into account the exact distribution of memory budgets to cores. In this sense, the proposed analysis represents a generalization of approaches that consider (i) even budget distribution across cores; and (ii) uneven but unknown (except for the core under analysis) budget assignment. By exploiting the additional piece of information, we show that it is possible to derive a more accurate WCET estimation. Our evaluations highlight that the proposed technique can reduce overestimation by 30% in average, and up to 60%, compared to the state of the art. Renato Mancuso 0001, Rodolfo Pellizzoni, Neriman Tokcan, Marco Caccamo |
ECRTS | 4 |
| 2017 | A Reliable and Predictable Scratchpad-centric OS for Multi-core Embedded SystemsabstractThe reliable use of multi-core platforms for designing safety-critical systems still represents an open challenge. Recently, the FAA [1] has formally expressed its concern towards the use of multi-core systems in avionics. The sharing of hardware resources introduces non-trivial timing dependencies between logically independent components (e.g. cores); additionally, the increase in size of circuitry, memory resources, and transistor density makes these platforms more susceptible to transient memory (soft) errors. This work addresses the problem of memory soft errors and their recovery at an OS/platform level on commercial multi-core systems. Proposed strategy considers the schedulability impact of recovery procedures on hard real-time workloads. Finally, the implementation of a SPM-centric OS with the proposed OS-level strategies was performed by using a commercially available multi-core platform. The design has been validated and evaluated using a combination of synthetic and realistic (EEMBC) benchmarks. Rohan Tabish, Renato Mancuso 0001, Saud Wasly, Sujit S. Phatak, Rodolfo Pellizzoni, Marco Caccamo |
RTAS | 6 |
| 2017 | Restart-based fault-tolerance: System design and schedulability analysisabstractEmbedded systems in safety-critical environments are continuously required to deliver more performance and functionality, while expected to provide verified safety guarantees. Nonetheless, platform-wide software verification (required for safety) is often expensive. Therefore, design methods that enable utilization of components such as real-time operating systems (RTOS), without requiring their correctness to guarantee safety, is necessary. In this paper, we propose a design approach to deploy safe-by-design embedded systems. To attain this goal, we rely on a small core of verified software to handle faults in applications and RTOS and recover from them while ensuring that timing constraints of safety-critical tasks are always satisfied. Faults are detected by monitoring the application timing and fault-recovery is achieved via full platform restart and software reload, enabled by the short restart time of embedded systems. Schedulability analysis is used to ensure that the timing constraints of critical plant control tasks are always satisfied in spite of faults and consequent restarts. We derive schedulability results for four restart-tolerant task models. We use a simulator to evaluate and compare the performance of the considered scheduling models. Fardin Abdi Taghi Abad, Renato Mancuso 0001, Rohan Tabish, Marco Caccamo |
RTCSA | 4 |
| 2017 | A scheduling framework for handling integrated modular avionic systems on multicore platformsabstractAlthough multicore chips are quickly replacing uniprocessor ones, safety-critical embedded systems are still developed using single processor architecture. The reasons mainly concern predictability and certification issues. This paper proposes a scheduling framework for handling Integrated Modular Avionics (IMA) on multicore platforms providing predictability as well as flexibility in managing dynamic load conditions and unexpected temporal misbehaviors of multicore. A new computational model is proposed to allow specifying a higher degree of flexibility and minimum performance requirements. Schedulability analysis is derived for providing off-line guarantees of real-time constraints in worst-case scenarios, and an efficient reclaiming mechanism is proposed to improve the average-case performance. Simulation and experimental results are reported to validate the proposed approach. Alessandra Melani, Renato Mancuso 0001, Marco Caccamo, Giorgio C. Buttazzo, Johannes Freitag, Sascha Uhrig |
RTCSA | 3 |
| 2017 | Optimizing resource speed for two-stage real-time tasks
Alessandra Melani, Renato Mancuso 0001, Daniel Cullina, Marco Caccamo, Lothar Thiele |
Real Time Syst. | 4 |
| 2016 | Speed optimization for tasks with two resources
Alessandra Melani, Renato Mancuso 0001, Daniel Cullina, Marco Caccamo, Lothar Thiele |
DATE | 4 |
| 2016 | Reset-based recovery for real-time cyber-physical systems with temporal safety constraintsabstractIn traditional computing systems, software problems are often resolved by platform restarts. This approach, however, cannot be naïvely used in cyber-physical systems (CPS). In fact, in this class of systems, ensuring safety strictly depends on the ability to respect hard real-time constraints. Several adaptations of the Simplex architecture have been proposed to guarantee safety in spite of misbehaving software components. However, the problem of performing recovery into a fully operational state has not been extensively addressed. In this work, we discuss how resets can be used in CPS as an effective strategy to recover from a variety of software faults. Our work extends the Simplex architecture in a number of directions. First, we provide sufficient conditions under which safety is guaranteed in spite of fault-induced resets. Second, we introduce a novel technique to express not only state-dependent safety constraints, as typically done in Simplex, but also time-dependent safety properties. Finally, through a proof-of-concept minimal implementation on a small R/C helicopter and simulation-based system modeling, we show the effectiveness of the proposed recovery strategy under the assumed fault model. Fardin Abdi Taghi Abad, Renato Mancuso 0001, Stanley Bak, Or D. Dantsker, Marco Caccamo |
ETFA | 5 |
| 2016 | A Real-Time Scratchpad-Centric OS for Multi-Core Embedded SystemsabstractMulti-core processors have replaced single-core systems in almost every segment of the industry. Unfortunately, their increased complexity often causes a loss of temporal predictability which represents a key requirement for hard real-time systems. Major sources of unpredictability are the shared low level resources, such as the memory hierarchy and the I/O subsystem. In this paper, we approach the problem of shared resource arbitration at an OS-level and propose a novel scratchpad-centric OS design for multi-core platforms. In the proposed OS, the predictable usage of shared resources across multiple cores represents a central design-time goal. Hence, we show (i) how contention-free execution of real-time tasks can be achieved on scratchpad-based architectures, and (ii) how a separation of application logic and I/O perations in the time domain can be enforced. To validate the proposed design, we implemented the proposed OS using a commercial-off-the-shelf (COTS) platform. Experiments show that this novel design delivers predictable temporal behavior to hard real-time tasks, and it improves performance up to 2.1× compared to traditional approaches. Rohan Tabish, Renato Mancuso 0001, Saud Wasly, Ahmed Alhammad, Sujit S. Phatak, Rodolfo Pellizzoni, Marco Caccamo |
RTAS | 7 |
| 2016 | Guest editorial: special issue on multicore systems
Marco Caccamo, Marko Bertogna |
Real Time Syst. | 1 |
| 2016 | Global Real-Time Memory-Centric Scheduling for Multicore SystemsabstractAs the number of cores increases, more master components can simultaneously access main memory. In real-time systems, this ongoing trend is leading to crippling pessimism when computing the worst-case cache miss time, since a memory request could potentially contend with other requests coming from every other core in the system. CPU-centric scheduling policies, therefore, are no longer sufficient to guarantee schedulability without introducing unacceptable pessimism for memory-intensive task sets. For this reason, we believe a shift is needed towards real-time scheduling approaches that can prevent timing interference from memory contention, while still making efficient use of the multicore platform. Previously, we have demonstrated the practicality of the PREM task model, where each job consists of a sequence of phases, some of which access memory and some of which perform only computation on cached data. In this work, we present the first global memory-centric scheduling policy for memory-intensive task sets whose jobs can be modeled as a sequence of memory-intensive (memory phase) and execution-intensive (execution phase) phases. The proposed policy is parameterizable based on the number of cores which are allowed to concurrently access main memory without saturating it. Building upon results from multicore response-time analysis, we introduce the notion of virtual memory cores as a fundamental technique for performing phase-based response time analysis for memory-intensive task sets. Finally, we use synthetic task set generation to demonstrate that proposed scheduling policy and related schedulability bound do indeed better schedule memory-intensive task sets when compared to state-of-art multicore scheduling. Rodolfo Pellizzoni, Stanley Bak, Heechul Yun, Marco Caccamo |
IEEE Trans. Computers | 5 |
| 2016 | Schedulability Analysis for Memory Bandwidth Regulated Multicore Real-Time SystemsabstractMulticore architecture brings a significant challenge in designing critical real-time systems because of timing variability caused by concurrent accesses to shared memory. We propose a memory bandwidth regulated system architecture and a novel analysis method to address this challenge. In the proposed architecture, each core's memory access rate is regulated in a globally coordinated manner. The architecture allows system designers to control the system to satisfy desired real-time performance. The proposed analysis method provides a way to calculate worst case response time of each real-time task independently from other activities on other cores; it only depends on the task under analysis, the assigned bandwidth, and the number of cores in the system. We believe this independence is critical to enable modular certification of critical real-time systems. We implement the proposed system model on the gem5 architecture simulator. We evaluate the proposed analysis method by comparing the computed runtime with the measured runtime on the modified simulator. We show that the analysis method provides reasonable upper-bounds based on the SPEC2006 benchmark suite. Heechul Yun, Zheng Pei Wu, Rodolfo Pellizzoni, Marco Caccamo, Lui Sha |
IEEE Trans. Computers | 5 |
| 2016 | Memory Bandwidth Management for Efficient Performance Isolation in Multi-Core PlatformsabstractMemory bandwidth in modern multi-core platforms is highly variable for many reasons and it is a big challenge in designing real-time systems as applications are increasingly becoming more memory intensive. In this work, we proposed, designed, and implemented an efficient memory bandwidth reservation system, that we call MemGuard. MemGuard separates memory bandwidth in two parts: guaranteed and best effort. It provides bandwidth reservation for the guaranteed bandwidth for temporal isolation, with efficient reclaiming to maximally utilize the reserved bandwidth. It further improves performance by exploiting the best effort bandwidth after satisfying each core's reserved bandwidth. MemGuard is evaluated with SPEC2006 benchmarks on a real hardware platform, and the results demonstrate that it is able to provide memory performance isolation with minimal impact on overall throughput. Heechul Yun, Rodolfo Pellizzoni, Marco Caccamo, Lui Sha |
IEEE Trans. Computers | 4 |
| 2016 | Real-Time Reachability for Verified Simplex DesignabstractThe Simplex architecture ensures the safe use of an unverifiable complex/smart controller by using it in conjunction with a verified safety controller and verified supervisory controller (switching logic). This architecture enables the safe use of smart, high-performance, untrusted, and complex control algorithms to enable autonomy without requiring the smart controllers to be formally verified or certified. Simplex incorporates a supervisory controller that will take over control from the unverified complex/smart controller if it misbehaves and use a safety controller. The supervisory controller should (1) guarantee that the system never enters an unsafe state (safety), but should also (2) use the complex/smart controller as much as possible (minimize conservatism). The problem of precisely and correctly defining the switching logic of the supervisory controller has previously been considered either using a control-theoretic optimization approach or through an offline hybrid-systems reachability computation. In this work, we show that a combined online/offline approach that uses aspects of the two earlier methods, along with a real-time reachability computation, also maintains safety, but with significantly less conservatism, allowing the complex controller to be used more frequently. We demonstrate the advantages of this unified approach on a saturated inverted pendulum system, in which the verifiable region of attraction is over twice as large compared to the earlier approach. Additionally, to validate the claims that the real-time reachability approach may be implemented on embedded platforms, we have ported and conducted embedded hardware studies using both ARM processors and Atmel AVR microcontrollers. This is the first ever demonstration of a hybrid-systems reachability computation in real time on actual embedded platforms, which required addressing significant technical challenges. Taylor T. Johnson, Stanley Bak, Marco Caccamo, Lui Sha |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2015 | WCET(m) Estimation in Multi-core Systems Using Single Core EquivalenceabstractMulti-core platforms represent the answer of the industry to the increasing demand for computational capabilities. From a real-time perspective, however, the inherent sharing of resources, such as memory subsystem and I/O channels, creates inter-core timing interference among critical tasks and applications deployed on different cores. As a result, modular per-core certification cannot be performed, meaning that: (1) current industrial engineering processes cannot be reused, (2) software developed and certified for single-core chips cannot be deployed on multi-core platforms as is. In this work, we propose the Single Core Equivalence (SCE) technology: a framework of OS-level techniques designed for commercial (COTS) architectures that exports a set of equivalent single-core virtual machines from a multi-core platform. This allows per-core schedulability results to be calculated in isolation and to hold when multiple cores of the system run in parallel. Thus, SCE allows each core of a multi-core chip to be considered as a conventional single-core chip, ultimately enabling industry to reuse existing software, schedulability analysis methodologies and engineering processes. Renato Mancuso 0001, Rodolfo Pellizzoni, Marco Caccamo, Lui Sha, Heechul Yun |
ECRTS | 3 |
| 2015 | Using traffic phase shifting to improve AFDX link utilizationabstractThe Avionic Full-Duplex Switched Ethernet (AFDX) is a data network certified for avionic operations. AFDX closely follows the IEEE 802.3 (Ethernet) standard for packet forwarding. On top of that, bandwidth enforcement using traffic shaping is performed to provide deterministic delivery guarantees. The design of an AFDX network, however, imposes that bandwidth enforcement is performed at a coarse granularity. This, together with the tight requirements on transmission jitter, determines a low utilization of the physical links. In this work, we propose traffic phase shifting (TPS) as a way to increase the granularity of bandwidth assignment to nodes of an AFDX network using logic time synchronization among traffic sources. Specifically, we leverage the periodic nature of real-time traffic and use phase-shifing to prevent link congestion. This in turns allows a more fine-grained bandwidth control via the AFDX protocol. We show that TPS leads to significant improvements in terms of per-link utilization without violating predictability. Renato Mancuso 0001, Andrew V. Louis, Marco Caccamo |
EMSOFT | 3 |
| 2015 | A Memory Access Detection Methodology for Accurate Workload CharacterizationabstractTools for memory access detection are widely used, playing an important role especially in real-time systems. For example, on multi-core platforms, the problem of co-scheduling CPU and memory resources with hard real-time constraints requires a deep understanding of the memory access patterns of the deployed task set. While code execution flow can be analyzed by considering the control-flow graph and reasoning in terms of basic blocks, a similar approach cannot apply to data accesses. In this paper, we propose MadT, a tool that uses a novel mechanism to perform memory access detection of general purpose applications. MadT does not perform binary instrumentation and always executes application code natively on the platform. Hence it can operate entirely in user-space without sand-boxing the task under analysis. Furthermore, MadT provides detailed symbolic information about the accessed memory structures, so it is able to translate the virtual addresses to their original symbolic variable names. Finally, it requires no modifications to application source code. The proposed methodology relies on existing OS-level capabilities. In this paper, we describe how MadT has been implemented on commercial hardware and compare its performance with state-of-the-art software techniques for memory access detection. Marco Cesati, Renato Mancuso 0001, Emiliano Betti, Marco Caccamo |
RTCSA | 4 |
| 2015 | Safety and Progress for Distributed Cyber-Physical Systems with Unreliable CommunicationabstractCyber-physical systems (CPSs) may interact and manipulate objects in the physical world, and therefore formal guarantees about their behavior are strongly desired. Static-time proofs of safety invariants, however, may be intractable for systems with distributed physical-world interactions. This is further complicated when realistic communication models are considered, for which there may not be bounds on message delays, or even when considering that messages will eventually reach their destination. In this work, we address the challenge of proving safety and progress in distributed CPSs communicating over an unreliable communication layer. We show that for this type of communication model, system safety is closely related to the results of a hybrid system’s reachability computation, which can be computed at runtime. However, since computing reachability at runtime may be computationally intensive, we provide an approach that moves significant parts of the computation to design time. This approach is demonstrated with a case study of a simulation of multiple vehicles moving within a shared environment. Stanley Bak, Zhenqi Huang, Fardin Abdi Taghi Abad, Marco Caccamo |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2014 | Light-PREM: Automated software refactoring for predictable execution on COTS embedded systemsabstractAs real-time embedded systems become more complex, there is the need to build them using high performance commercial off-the-shelf (COTS) components. However, tasks can exhibit hard to predict worst case execution times (WCET) when executing on commodity hardware, due to contention among shared physical resources. Past work has introduced the PRedictable Execution Model (PREM) [1] to solve this issue, but unfortunately, the time required to manually refactor existing code according to this model is too high. Light-PREM proposes a novel technique that automates the refactoring process needed to convert legacy software applications to PREM-compliant code. The advantage of Light-PREM is twofold. On one side, it makes the adoption of PREM more attractive from an industrial point of view, because it significantly reduces the amount of work that is needed to generate PREM-compliant code. On the other hand, the proposed methodology is general enough to be used with any embedded software design. Experimental results show that Light-PREM significantly improves the predictability of real-time applications without requiring software engineers to gain a deep understanding about software memory usage. Renato Mancuso 0001, Roman Dudko, Marco Caccamo |
RTCSA | 3 |
| 2014 | A hardware architecture to deploy complex multiprocessor scheduling algorithmsabstractAn increasing demand for high-performance systems has been observed in the domain of both general purpose and real-time systems, pushing the industry towards a pervasive transition to multi-core platforms. Unfortunately, well-known and efficient scheduling results for single-core systems do not scale well to the multi-core domain. This justifies the adoption of more computationally intensive algorithms, but the complexity and computational overhead of these algorithms impact their applicability to real OSes. We propose an architecture to migrate the burden of multi-core scheduling to a dedicated hardware component. We show that it is possible to mitigate the overhead of complex algorithms, while achieving power efficiency and optimizing processors utilization. We develop the idea of “active monitoring” to continuously track the evolution of scheduling parameters as tasks execute on processors. This allows reducing the gap between implementable scheduling techniques and the ideal fluid scheduling model, under the constraints of realistic hardware. Renato Mancuso 0001, Prakalp Srivastava, Deming Chen, Marco Caccamo |
RTCSA | 4 |
| 2014 | Real-Time Reachability for Verified Simplex DesignabstractThe Simplex Architecture ensures the safe use of an unverifiable complex controller by using a verified safety controller and verified switching logic. This architecture enables the safe use of high-performance, untrusted, and complex control algorithms without requiring them to be formally verified. Simplex incorporates a supervisory controller and safety controller that will take over control if the unverified logic misbehaves. The supervisory controller should (1) guarantee the system never enters and unsafe state (safety), but (2) use the complex controller as much as possible (minimize conservatism). The problem of precisely and correctly defining this switching logic has previously been considered either using a control-theoretic optimization approach, or through an offline hybrid systems reach ability computation. In this work, we prove that a combined online/offline approach, which uses aspects of the two earlier methods along with a real-time reach ability computation, also maintains safety, but with significantly less conservatism. We demonstrate the advantages of this unified approach on a saturated inverted pendulum system, where the usable region of attraction is 227% larger than the earlier approach. Stanley Bak, Taylor T. Johnson, Marco Caccamo, Lui Sha |
RTSS | 3 |
| 2013 | Real-time cache management framework for multi-core architecturesabstractMulti-core architectures are shaking the fundamental assumption that in real-time systems the WCET, used to analyze the schedulability of the complete system, is calculated on individual tasks. This is not even true in an approximate sense in a modern multi-core chip, due to interference caused by hardware resource sharing. In this work we propose (1) a complete framework to analyze and profile task memory access patterns and (2) a novel kernel-level cache management technique to enforce an efficient and deterministic cache allocation of the most frequently accessed memory areas. In this way, we provide a powerful tool to address one of the main sources of interference in a system where the last level of cache is shared among two or more CPUs. The technique has been implemented on commercial hardware and our evaluations show that it can be used to significantly improve the predictability of a given set of critical tasks. Renato Mancuso 0001, Roman Dudko, Emiliano Betti, Marco Cesati, Marco Caccamo, Rodolfo Pellizzoni |
IEEE Real-Time and Embedded Technology and Applications Symposium | 5 |
| 2013 | MemGuard: Memory bandwidth reservation system for efficient performance isolation in multi-core platformsabstractMemory bandwidth in modern multi-core platforms is highly variable for many reasons and is a big challenge in designing real-time systems as applications are increasingly becoming more memory intensive. In this work, we proposed, designed, and implemented an efficient memory bandwidth reservation system, that we call MemGuard. MemGuard distinguishes memory bandwidth as two parts: guaranteed and best effort. It provides bandwidth reservation for the guaranteed bandwidth for temporal isolation, with efficient reclaiming to maximally utilize the reserved bandwidth. It further improves performance by exploiting the best effort bandwidth after satisfying each core's reserved bandwidth. MemGuard is evaluated with SPEC2006 benchmarks on a real hardware platform, and the results demonstrate that it is able to provide memory performance isolation with minimal impact on overall throughput. Heechul Yun, Rodolfo Pellizzoni, Marco Caccamo, Lui Sha |
IEEE Real-Time and Embedded Technology and Applications Symposium | 4 |
| 2013 | Using run-time checking to provide safety and progress for distributed cyber-physical systemsabstractCyber-physical systems (CPS) may interact and manipulate objects in the physical world, and therefore ideally would have formal guarantees about their behavior. Performing static-time proofs of safety invariants, however, may be intractable for systems with distributed physical-world interactions. This is further complicated when realistic communication models are considered, for which there may not be bounds on message delays, or even that messages will eventually reach their destination. In this work, we address the challenge of proving safety and progress in distributed CPS communicating over an unreliable communication layer. This is done in two parts. First, we show that system safety can be verified by partially relying upon run-time checks, and that dropping messages if the run-time checks fail will maintain safety. Second, we use a notion of compatible action chains to guarantee system progress, despite unbounded message delays. We demonstrate the effectiveness of our approach on a multi-agent vehicle flocking system, and show that the overhead of the proposed run-time checks is not overbearing. Stanley Bak, Fardin Abdi Taghi Abad, Zhenqi Huang, Marco Caccamo |
RTCSA | 4 |
| 2013 | Real-Time I/O Management System with COTS PeripheralsabstractReal-time embedded systems are increasingly being built using commercial-off-the-shelf (COTS) components such as mass-produced peripherals and buses to reduce costs, time-to-market, and increase performance. Unfortunately, COTS-interconnect systems do not usually guarantee timeliness, and might experience severe timing degradation in the presence of high-bandwidth I/O peripherals. Moreover, peripherals do not implement any internal priority-based scheduling mechanism, hence, sharing a device can result in data of high priority tasks being delayed by data of low priority tasks. To address these problems, we designed a real-time I/O management system comprised of 1) real-time bridges with I/O virtualization capabilities, and 2) a peripheral scheduler. The proposed framework is used to transparently put the I/O subsystem of a COTS-based embedded system under the discipline of real-time scheduling, minimizing the timing unpredictability due to the peripherals sharing the bus. We also discuss computing the maximum delay due to buffered I/O data transactions as well as determining the buffer size needed to avoid data loss. Finally, we demonstrate experimentally that our prototype real-time I/O management system successfully exports multiple virtual devices for a single physical device and prioritizes I/O traffic, guaranteeing its timeliness. Emiliano Betti, Stanley Bak, Rodolfo Pellizzoni, Marco Caccamo, Lui Sha |
IEEE Trans. Computers | 4 |
| 2012 | Memory Access Control in Multiprocessor for Real-Time Systems with Mixed CriticalityabstractShared resource access interference, particularly memory and system bus, is a big challenge in designing predictable real-time systems because its worst case behavior can significantly differ. In this paper, we propose a software based memory throttling mechanism to explicitly control the memory interference. We developed analytic solutions to compute proper throttling parameters that satisfy schedulability of critical tasks while minimize performance impact caused by throttling. We implemented the mechanism in Linux kernel and evaluated isolation guarantee and overall performance impact using a set of synthetic and real applications. Heechul Yun, Rodolfo Pellizzoni, Marco Caccamo, Lui Sha |
ECRTS | 4 |
| 2012 | A Fault Resilient Architecture for Distributed Cyber-Physical SystemsabstractIn this paper we discuss a general approach and architecture for design of distributed cyber-physical systems in order to make them resilient to communication faults. In this approach, each node exploits physical connections between nodes to estimate some of the state parameters of the remote nodes in order to detect the faults and also to maintain stability of system after fault occurrence. Finally, based on this architecture and approach, a fault-resilient decentralized voltage control algorithm is presented and evaluated. Fardin Abdi Taghi Abad, Marco Caccamo, Brett A. Robbins |
RTCSA | 2 |
| 2012 | Memory-Aware Scheduling of Multicore Task Sets for Real-Time SystemsabstractReal-time scheduling of memory-intensive applications is a particularly difficult challenge. On a multi-core system, not only is the CPU scheduling an issue, but equally important is the management of mutual interference among tasks caused by simultaneous access to the shared main memory. To confront this problem, we explore real-time schedulers for task sets which adhere to the Predictable Execution Model (PREM). In each PREM-compliant task, execution is divided into phases which retrieve data from main memory, and phases which perform local computation using previously-cached data. In this work, we perform a simulation-based analysis with the goal of determining which schedulers are generally better at scheduling PREM-compliant task sets. We investigate several memory intensive real-time benchmarks from the EEMBC benchmark suite, in order to drive our task set generation parameters. We elaborate on a PREM-complaint task set simulator which we designed specifically to be able to simulate PREM-compliant tasks. The overall best scheduling policy we found, which we call M-LAX, schedules access to memory in a no preemptive fashion according to a least-laxity-first policy. M-LAX outperforms an EDF-based approach, a previously-analyzed TDMA arbitration scheme, and the unscheduled case where tasks interfere when accessing memory. Stanley Bak, Rodolfo Pellizzoni, Marco Caccamo |
RTCSA | 4 |
| 2012 | Memory-centric scheduling for multicore hard real-time systems
Rodolfo Pellizzoni, Stanley Bak, Emiliano Betti, Marco Caccamo |
Real Time Syst. | 5 |
| 2012 | Real-Time Scheduling of Concurrent Transactions in Multidomain Ring BusesabstractWe address the problem of scheduling concurrent periodic real-time transactions on Multidomain Ring Bus (MDRB). The problem is challenging because although the bus allows multiple nonoverlapping transactions to be executed concurrently, the degree of concurrency depends on the topology of the bus and of executed transactions. To solve this problem, first, we propose two novel efficient scheduling algorithms for topographically acyclic transaction sets. The first algorithm is optimal for transaction sets under restrictive assumptions while the second one induces a good sufficient schedulable utilization bound for more general transaction sets. Then, we extend these two algorithms for the scheduling of topographically cyclic transaction sets. Extensive simulations show that the proposed algorithm can schedule transaction sets with high bus utilization and is better than that of related works in most practical settings. The implementation of the algorithms in a real testbed shows that they have relatively low execution-time overhead. Bach Duy Bui, Rodolfo Pellizzoni, Marco Caccamo |
IEEE Trans. Computers | 3 |
| 2011 | A step towards verification and synthesis from simulink/stateflow modelsabstractThis paper describes a toolkit for synthesizing hybrid supervisory control systems starting from the popular Simulink/Stateflow modeling environment. The toolkit provides a systematic strategy for translating Simulink/Stateflow models to hybrid automata and a discrete abstraction-based algorithm for synthesizing supervisory controllers. Karthik Manamcheri, Sayan Mitra 0001, Stanley Bak, Marco Caccamo |
HSCC | 4 |
| 2011 | A Predictable Execution Model for COTS-Based Embedded SystemsabstractBuilding safety-critical real-time systems out of inexpensive, non-real-time, COTS components is challenging. Although COTS components generally offer high performance, they can occasionally incur significant timing delays. To prevent this, we propose controlling the operating point of each shared resource (like the cache, memory, and interconnection buses) to maintain it below its saturation limit. This is necessary because the low-level arbiters of these shared resources are not typically designed to provide real-time guarantees. In this work, we introduce a novel system execution model, the Predictable Execution Model (PREM), which, in contrast to the standard COTS execution model, coschedules at a high level all active components in the system, such as CPU cores and I/O peripherals. In order to permit predictable, system-wide execution, we argue that real-time embedded applications should be compiled according to a new set of rules dictated by PREM. To experimentally validate our theory, we developed a COTS-based PREM testbed and modified the LLVM Compiler Infrastructure to produce PREM-compatible executables. Rodolfo Pellizzoni, Emiliano Betti, Stanley Bak, John Criswell, Marco Caccamo, Russell Kegley |
IEEE Real-Time and Embedded Technology and Applications Symposium | 6 |
| 2011 | Timing Analysis for Resource Access Interference on Adaptive Resource ArbitersabstractModern multiprocessor and multicore architectures adopt shared resources to meet increased performance requirements. Adaptive arbiters, such as FlexRay, have been adopted to grant access to shared resources. While increasing the performance, timing analysis is more challenging with this kind of arbiter. This paper considers real-time tasks that are composed of super blocks, while super blocks themselves are composed of phases. Phases are characterized by their worst-case computation time on their processing element and their worst-case number of access requests to a shared resource. Resource accesses, such as access to caches or scratchpad memory, are synchronous and cause the processing element to stall until the access is served. Based on dynamic programming, we develop an algorithm that safely derives an upper-bound of the worst-case response time of a phase. The worst-case response time of a task can then be determined for both sequential or time-triggered execution of super blocks. Experimental results are conducted for a real-world application. Andreas Schranzhofer, Rodolfo Pellizzoni, Jian-Jia Chen, Lothar Thiele, Marco Caccamo |
IEEE Real-Time and Embedded Technology and Applications Symposium | 5 |
| 2011 | A Slot-Based Real-Time Scheduling Algorithm for Concurrent Transactions in NoCabstractWe address the problem of scheduling real-time transactions in Network-on-Chip (NoC). In particular, we propose a novel slot-based scheduling algorithm for acyclic transaction sets in NoC. The algorithm induces a competitive sufficient schedulability utilization bound. Since the proposed algorithm is able to exploit the parallelism between non-overlapping transactions, under given assumptions, it performs better than the existing fixed-priority solutions. We evaluate performance through extensive simulations. Furthermore, we discuss some important factors in the implementation of the algorithm and present an implementation in a real system. The measurement shows that the proposed algorithm has relatively low overhead. Bach Duy Bui, Marco Caccamo, Rodolfo Pellizzoni |
RTCSA (1) | 2 |
| 2010 | Worst-case response time analysis of resource access models in multi-core systemsabstractMulti-processor and multi-core systems are becoming increasingly important in time critical systems. Shared resources, such as shared memory or communication buses are used to share data and read sensors. We consider real-time tasks constituted by superblocks, which can be executed sequentially or by a time triggered static schedule. Three models to access shared resources are explored: (1) the dedicated access model, in which accesses happen only in dedicated phases, (2) the general access model, in which accesses could happen at anytime, and (3) the hybrid access model, combining the dedicated and general access model. For resource access based on a Time Division Multiple Access (TDMA) protocol, we analyze the worst-case completion time for a superblock, derive worst-case response times for tasks and obtain the relation of schedulability between different models. We conclude with proposing the dedicated sequential model as the model of choice for time critical resource sharing multi-processor/multi-core systems. Andreas Schranzhofer, Rodolfo Pellizzoni, Jian-Jia Chen, Lothar Thiele, Marco Caccamo |
DAC | 5 |
| 2010 | Worst case delay analysis for memory interference in multicore systemsabstractEmploying COTS components in real-time embedded systems leads to timing challenges. When multiple CPU cores and DMA peripherals run simultaneously, contention for access to main memory can greatly increase a task's WCET. In this paper, we introduce an analysis methodology that computes upper bounds to task delay due to memory contention. First, an arrival curve is derived for each core representing the maximum memory traffic produced by all tasks executed on it. Arrival curves are then combined with a representation of the cache behavior for the task under analysis to generate a delay bound. Based on the computed delay, we show how tasks can be feasibly scheduled according to assigned time slots on each core. Rodolfo Pellizzoni, Andreas Schranzhofer, Jian-Jia Chen, Marco Caccamo, Lothar Thiele |
DATE | 4 |
| 2010 | Preemption Points Placement for Sporadic Task SetsabstractLimited preemption scheduling has been introduced as a viable alternative to non-preemptive and fully preemptive scheduling when reduced blocking times need to coexist with an acceptable context switch overhead. To achieve this goal, preemptions are allowed only at selected points of the code of each task, decreasing the preemption overhead and simplifying the estimation of worst-case execution parameters. Unfortunately, the problem of how to place these preemption points is rather complex and has not been solved. In this paper, a method is presented for the optimal placement of preemption points under simplifying conditions, namely, a fixed preemption overhead at each point. We will prove that if our method is not able to produce a feasible schedule, then no other possible preemption point placement (including non-preemptive and fully preemptive scheduling) can find a schedulable solution. The presented method is general enough to be applicable to both EDF and Fixed Priority scheduling, with limited modifications. Marko Bertogna, Giorgio C. Buttazzo, Mauro Marinoni, Francesco Esposito, Marco Caccamo |
ECRTS | 6 |
| 2010 | Real-Time Communication for Multicore Systems with Multi-domain Ring BusesabstractWe address the problem of scheduling real-time data transactions on a multicore processor bus. In particular, to in-crease system predictability and tighten WCET estimation, we propose to employ a software-controllable Multi-Domain Ring Bus (MDRB) architecture. The problem of scheduling periodic real-time transactions on MDRB is challenging because the bus allows multiple non-overlapping transactions to be executed concurrently, and because the degree of concurrency depends on the topology of the bus and of executed transactions. We propose a practical abstraction mechanism for the scheduling problem together with two novel scheduling algorithms. The first algorithm is optimal for transaction sets under restrictive assumptions while the second one induces a competitive sufficient schedulable utilization bound for more general transaction sets. Bach Duy Bui, Rodolfo Pellizzoni, Deepti K. Chivukula, Marco Caccamo |
RTCSA | 4 |
| 2010 | Impact of Peripheral-Processor Interference on WCET Analysis of Real-Time Embedded SystemsabstractThe integration phase of real-time COTS-based systems is challenging. When multiple tasks run concurrently, the interference at the bus level between cache fetching activities and I/O peripheral transactions is significant and causes unpredictable behaviors: experimentally, we show that tasks can have computation time variance up to 46 percent in a typical embedded system. In this work, we present a theoretical framework able to model the interaction between CPU and peripherals contending for shared main memory through the Front Side Bus (FSB). We first show how to compute worst case execution time (WCET) for a task given a trace of its cache activity and given an upper bound function that models peripheral activities. Then, we show how the analysis can be extended to a multitasking environment assuming a restricted-preemption model. Finally, we introduce the novel idea of ¿hardware server¿ as a means of controlling the unpredictable behavior of COTS peripheral components. Rodolfo Pellizzoni, Marco Caccamo |
IEEE Trans. Computers | 2 |
| 2009 | Handling mixed-criticality in SoC-based real-time embedded systemsabstractSystem-on-Chip (SoC) is a promising paradigm to implement safety-critical embedded systems, but it poses significant challenges from a design and verification point of view. In particular, in a mixed-criticality system, low criticality applications must be prevented from interfering with high criticality ones. In this paper, we introduce a new design methodology for SoC that provides strong isolation guarantees to applications with different criticalities. A set of certificates describing the assumed application behavior is extracted from a functional Architectural Analysis and Design Language (AADL) specification. Our tools then automatically generate hardware wrappers that enforce at run-time the behavior described by the certificates. In particular, we employ run-time monitoring to formally check all data communication in the system, and we enforce timing reservations for both computation and communication resources. Verification is greatly simplified because certificates are much simpler than the components used to implement low-criticality applications. The effectiveness of our methodology is proven on a case study consisting of a medical pacemaker. Rodolfo Pellizzoni, Patrick O'Neil Meredith, Min-Young Nam, Mu Sun, Marco Caccamo, Lui Sha |
EMSOFT | 5 |
| 2009 | The System-Level Simplex Architecture for Improved Real-Time Embedded System SafetyabstractEmbedded systems in safety-critical environments demand safety guarantees while providing many useful services that are too complex to formally verify or fully test. Existing application-level fault-tolerance methods, even if formally verified, leave the system vulnerable to errors in the real-time operating system (RTOS), middleware, and microprocessor. We introduce the system-level simplex architecture, which uses hardware/software co-design to provide fail-operational guarantees for both logical application-level faults, as well as faults in previously dependent layers including the RTOS and microprocessor. We also provide an end-to-end design process for the system-level simplex architecture where the AADL architecture description is automatically constructed and checked and the VHDL hardware code is generated. To show the efficacy of System-Level Simplex design, we apply the approach to both a classic inverted pendulum and a cardiac pacemaker. We perform fault-injection tests on the inverted pendulum design which demonstrate robustness in spite of software controller and operating system faults. For the pacemaker, we contrast the provided safety guarantees with those of a previous-generation pacemaker. Stanley Bak, Deepti K. Chivukula, Olugbemiga Adekunle, Mu Sun, Marco Caccamo, Lui Sha |
IEEE Real-Time and Embedded Technology and Applications Symposium | 5 |
| 2009 | Real-Time Control of I/O COTS Peripherals for Embedded SystemsabstractReal-time embedded systems are increasingly being built using commercial-off-the-shelf (COTS) components such as mass-produced peripherals and buses to reduce costs, time-to-market, and increase performance. Unfortunately, COTS interconnect systems do not usually guarantee timeliness, and might experience severe timing degradation in the presence of high-bandwidth I/O peripherals. To address this problem, we designed a real-time I/O management system comprised of 1) real-time bridges, and 2) a reservation controller. The proposed framework is used to transparently put the I/O subsystem of a COTS-based embedded system under the discipline of real-time scheduling. We also discuss computing a delay bound for I/O data transactions and determining worst-case buffer size. Finally, we demonstrate experimentally that our prototype real-time I/O management system successfully prioritizes I/O traffic and guarantees its timeliness. Stanley Bak, Emiliano Betti, Rodolfo Pellizzoni, Marco Caccamo, Lui Sha |
RTSS | 4 |
| 2008 | Hybrid Hardware-Software Architecture for Reconfigurable Real-Time SystemsabstractRecent developments in the field of reconfigurable SoC devices (FPGAs) will enable the development of embedded systems where software tasks, running on a CPU, can coexist with hardware tasks. We devised a real-time computing architecture that can integrate hardware and software executions in a transparent manner, and can support real-time QoS adaptation by means of partial reconfiguration of modern FPGA devices. Tasks are allowed to migrate seamlessly from CPU to FPGA and vice versa to support dynamic QoS adaptation and cope with dynamic workloads. In this paper, we discuss the design and implementation of an on-chip infrastructure, OS extensions and task design methodology that enable hardware-software transparency in the presence of relocation. The overall architecture is suitable to schedule real-time workloads and we derive bounds on relocation overhead. Finally, we show the applicability of our design methodology on a concrete task design case. Rodolfo Pellizzoni, Marco Caccamo |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2008 | Impact of Cache Partitioning on Multi-tasking Real Time Embedded SystemsabstractCache partitioning techniques have been proposed in the past as a solution for the cache interference problem. Due to qualitative differences with general purpose platforms, real-time embedded systems need to minimize task real-time utilization (function of execution time and period) instead of only minimizing the number of cache misses. In this work, the partitioning problem is presented as an optimization problem whose solution sets the size of each cache partition and assigns tasks to partitions such that system worst-case utilization is minimized thus increasing real-time schedulability. Since the problem is NP-Hard, a genetic algorithm is presented to find a near optimal solution. A case study and experiments show that in a typical real-time embedded system, the proposed algorithm is able to reduce the worst-case utilization by 15% (on average) if compared to the case when the system uses a shared cache or a proportional cache partitioned environment. Bach Duy Bui, Marco Caccamo, Lui Sha, Joseph Martinez |
RTCSA | 2 |
| 2008 | Coscheduling of CPU and I/O Transactions in COTS-Based Embedded SystemsabstractIntegrating COTS components in critical real-time systems is challenging. In particular, we show that the interference between cache activity and I/O traffic generated by COTS peripherals can unpredictably slow down a real-time task by up to 44%. To solve this issue, we propose a framework comprised of three main components: 1) a COTS-compatible device, the peripheral gate, that controls peripheral access to the system; 2) an analytical technique that computes safe bounds on the I/O-induced task delay; 3) a coscheduling algorithm that maximizes the amount of allowed peripheral traffic while guaranteeing all real-time task constraints. We implemented the complete framework on a COTS-based system using PCI peripherals, and we performed extensive experiments to show its feasibility. Rodolfo Pellizzoni, Bach Duy Bui, Marco Caccamo, Lui Sha |
RTSS | 3 |
| 2008 | Hardware Runtime Monitoring for Dependable COTS-Based Real-Time Embedded SystemsabstractCOTS peripherals are heavily used in the embedded market, but their unpredictability is a threat for high-criticality real-time systems: it is hard or impossible to formally verify COTS components. Instead, we propose to monitor the runtime behavior of COTS peripherals against their assumed specifications. If violations are detected, then an appropriate recovery measure can be taken. Our monitoring solution is decentralized: a monitoring device is plugged in on a peripheral bus and monitors the peripheral behavior by examining read and write transactions on the bus. Provably correct (w.r.t. given specifications) hardware monitors are synthesized from high level specifications, and executed on FPGAs, resulting in zero runtime overhead on the system CPU. The proposed technique, called BusMOP, has been implemented as an instance of a generic runtime verification framework, called MOP, which until now has only been used for software monitoring. We experimented with our technique using a COTS data acquisition board. Rodolfo Pellizzoni, Patrick O'Neil Meredith, Marco Caccamo, Grigore Rosu |
RTSS | 3 |
| 2008 | M-CASH: A real-time resource reclaiming algorithm for multiprocessor platforms
Rodolfo Pellizzoni, Marco Caccamo |
Real Time Syst. | 2 |
| 2008 | Sharp Thresholds for Scheduling Recurring Tasks with Distance ConstraintsabstractThe problem of identifying suitable conditions for the schedulability of (nonpreemptive) recurring tasks with deadlines is of great importance to real-time systems. In this paper, motivated by the problem of scheduling radar dwells, we show that scheduling problems of this nature show a sharp threshold behavior with respect to system utilization. Sharp thresholds are associated with phase transitions: When the utilization of a task set is less than a critical value, it can be scheduled almost surely and, when the utilization increases beyond the critical level, almost no task set can be scheduled. We make connections to work on random graphs to prove the sharp threshold behavior in the scheduling problem of interest. Using extensive experiments, we determine the threshold for the radar dwell scheduling problem and use it for performance optimization. The connections to random graph theory suggest new ways for understanding the average-case behavior of scheduling policies. These results emphasize the ease with which performance can be controlled in a variety of real-time systems. Sathish Gopalakrishnan, Marco Caccamo, Lui Sha |
IEEE Trans. Computers | 2 |
| 2007 | Real-time implications of multiple transmission rates in wireless networksabstractWireless networks are increasingly being used for latency-sensitive applications that require data delivery to be timely, efficient and reliable. This trend is primarily driven by the proliferation of wireless networks of real-time data-gathering sensor-actuator devices. This has led to a strong need to bring real-time concerns to the forefront of an integrated research thrust into wireless real-time systems. In this paper, we introduce and analyze a specific instance of the rich set of problems in this domain. We consider a wireless network serving real-time flows in which the underlying physical layer provides multiple transmission rates. Higher rates have more stringent SINR requirements and thus represent a trade-off between raw transmission speed and packet error rate. We adopt a first principles approach to the design of optimal real-time scheduling algorithms for such a multi-rate wireless network. We illustrate the inherent complexities of the problem through examples and obtain provably optimal structural results. We then characterize the optimal policy for an approximate model. Our theoretical analysis provides guidelines for heuristic scheduler design. Our initial work indicates that this is a rich problem domain with the potential for a unifying theory that integrates real-time requirements into multi-rate wireless network design. Vartika Bhandari, Vivek Raghunathan, Bach Duy Bui, Marco Caccamo |
MobiCom | 4 |
| 2007 | Soft Real-Time Chains for Multi-Hop Wireless Ad-Hoc NetworksabstractPrioritized MAC protocols are needed to support soft real-time communication in wireless networks. In this paper, we introduce real-time chain, a new prioritized MAC protocol to support soft real-time data flows in multi-hop wireless ad-hoc networks. By avoiding packet collisions and limiting the effect of priority inversions, real-time chain is able to provide soft real-time and bandwidth guarantees. Furthermore, the use of multiple channels enables high spatial reuse and transmission rates. Finally, our approach can be integrated with a slightly modified version of the IEEE 802.15.4 standard. The protocol has been fully implemented on Crossbow MICAz hardware and its performance has been validated with a large set of both indoor and outdoor experiments Bach Duy Bui, Rodolfo Pellizzoni, Marco Caccamo, Chin F. Cheah, Andrew Tzakis |
IEEE Real-Time and Embedded Technology and Applications Symposium | 3 |
| 2007 | Toward the Predictable Integration of Real-Time COTS Based SystemsabstractThe integration phase of real-time COTS-based systems is often problematic because when multiple tasks run concurrently, the interference at the bus level between cache fetching activities and I/O peripheral transactions is significant and causes unpredictable behaviors: experimentally, tasks can have computation time variance up to 50%. In this work, we present a theoretical framework able to model the interaction between CPU and peripherals contending for shared main memory through the front side bus (FSB). We first show how to compute worst case execution times for a task given a trace of its cache activity and given an upper bound function that models peripheral activities; then, we introduce the novel idea of "hardware server" as a means of controlling the unpredictable behavior of COTS peripheral components. Rodolfo Pellizzoni, Marco Caccamo |
RTSS | 2 |
| 2007 | Real-Time Management of Hardware and Software Tasks for FPGA-based Embedded SystemsabstractOperating systems for reconfigurable devices enable the development of embedded systems where software tasks, running on a CPU, can coexist with hardware tasks running on a reconfigurable hardware device (FPGA). In this work, we consider real-time systems subject to dynamic workloads and whose tasks can be computationally intensive. We introduce a novel resource allocation scheme and an online admission control test that achieve high performance and flexibility; in addition, runtime reconfiguration is used to maximize the number of admitted real-time tasks. Moreover, in detail, we first discuss a 1D system architecture and its prototype for a Xilinx Virtex-4 FPGA; then, we concentrate on the online admission control problem. Online task allocation and migration between the CPU and the reconfigurable device are discussed and sufficient feasibility tests are derived for both the commonly used slotted and 1D area models. Finally, the effectiveness of our admission control and relocation strategy is shown through a series of synthetic simulations. Rodolfo Pellizzoni, Marco Caccamo |
IEEE Trans. Computers | 2 |
| 2007 | Robust implicit EDF: A wireless MAC protocol for collaborative real-time systemsabstractAdvances in wireless technology have brought us closer to extensive deployment of distributed real-time embedded systems connected through a wireless channel. The medium-access control (MAC) layer protocol is critical in providing a real-time guarantee. We have devised a real-time wireless MAC protocol, robust implicit earliest deadline first, or RI-EDF. Packets are transmitted according to EDF scheduling rules, offering a protocol that implicitly avoids contention. In the event of a packet loss or a node failure, every node has the opportunity to recover the schedule based on a static recovery priority, offering a protocol that is robust with no central point of failure. We demonstrate in simulations that RI-EDF provides better goodput and lower packet loss than existing protocols like 802.11 PCF and EDCF. In our implementation and distributed control test-bed, we show that RI-EDF provides better throughput than the TinyOS MAC-layer protocol. Overall, RI-EDF provides predictable temporal behavior with minimal impact on node failures, packet losses, and noise in the channel. Tanya L. Crenshaw, Spencer Hoke, Ajay Tirumala, Marco Caccamo |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2007 | Building Robust Wireless LAN for Industrial Control with the DSSS-CDMA Cell Phone Network ParadigmabstractWireless LAN for industrial control (IC-WLAN) provides many benefits, such as mobility, low deployment cost, and ease of reconfiguration. However, the top concern is robustness of wireless communications. Wireless control loops must be maintained under persistent adverse channel conditions, such as noise, large-scale path loss, fading, and many electromagnetic interference sources in industrial environments. The conventional IEEE 802.11 WLANs, originally designed for high bandwidth instead of high robustness, are therefore inappropriate for IC-WLAN. A solution lies in the direct sequence spread spectrum (DSSS) technology: by deploying the largest possible processing gain (slowest bit rate) that fully exploits the low data rate feature of industrial control, much higher robustness can be achieved. We hereby propose using DSSS-CDMA to build IC-WLAN. We carry out fine-grained physical layer simulations and Monte Carlo comparisons. The results show that DSSS-CDMA IC-WLAN provides much higher robustness than IEEE 802.11/802.15.4 WLAN, so that reliable wireless industrial control loops become feasible. We also show that deploying larger processing gain is preferable, to deploying more intensive convolutional coding. The DSSS-CDMA IC-WLAN scheme also opens up a new problem space for interdisciplinary study, involving real-time scheduling, resource management, communication, networking, and control Qixin Wang 0001, Xue (Steve) Liu, Weiqun Chen, Lui Sha, Marco Caccamo |
IEEE Trans. Mob. Comput. | 5 |
| 2006 | The Dependency Management Framework: A Case Study of the ION CubeSatabstractDue to the complexity and requirements of modern realtime systems, multiple teams must often work concurrently and independently to develop the various components of the system. Since a team typically only knows the dependency relations between the components they wrote and those they directly use, keeping track of system-wide dependency relations is not possible for any individual team. To further complicate matters, dependency relations often change as software components are refined or their interactions modified. Because the robustness of any real-time system hinges on the availability of essential services in spite of faults and failures in useful but non-essential components, keeping track of the constantly evolving dependency relations between the system's components is crucial. If a system's designers cannot ensure that critical services only USE but do not DEPEND ON less critical components, a seemingly minor fault can propagate along complex and unforeseen dependency chains and bring down the entire system. Therefore, automatically tracking and analyzing system-wide dependency relations given only local dependency information is vital for the development of robust real time systems. This paper presents DMF (dependency management framework), a prototype toolkit for dependency management in designing robust real-time systems. We demonstrate the usability and scalability of DMF with a case study of ION CubeSat, the University of Illinois at Urbana-Champaign 's first student-developed satellite. Leon Arber, Lui Sha, Marco Caccamo |
ECRTS | 4 |
| 2006 | Formal Simulation and Analysis of the CASH Scheduling Algorithm in Real-Time Maude
Peter Csaba Ölveczky, Marco Caccamo |
FASE | 2 |
| 2006 | I-Living: An Open System Architecture for Assisted LivingabstractAdvances in networking, sensors, and embedded devices have made it feasible to monitor and provide medical and other assistance to people in their homes. Aging populations will benefit from reduced costs and improved healthcare through assisted living based on these technologies. However, these systems challenge current state-of-the-art techniques for usability, reliability, and security. This is a particular challenge for open and extensible systems that combine software and hardware from many vendors and provide information to diverse clinicians. In this paper we present the I-Living architecture for assisted living that allows independent parties work together in a dependable, secure, and low-cost fashion with predictable properties. Our approach is based on an Assisted Living Service Provider (ALSP) who provides a server that collects and maintains encrypted assisted persons (APs)' records. Our ALSP can be a third party distinct from APs, communication providers, and clinicians; or it can be part of an ISP, hospital or similar enterprise. We have explored the architecture by developing a collection of applications and implementing them in a prototype system. Our system shows the feasibility and opportunity of an open approach to assisted living systems. Qixin Wang 0001, Wook Shin, Xue (Steve) Liu, Zheng Zeng 0001, Cham Oh, Bedoor K. AlShebli, Marco Caccamo, Carl A. Gunter, Elsa L. Gunter, Jennifer C. Hou, Karrie Karahalios, Lui Sha |
SMC | 7 |
| 2006 | Finite-horizon scheduling of radar dwells with online template construction
Sathish Gopalakrishnan, Marco Caccamo, Chi-Sheng Shih 0001, Chang-Gun Lee, Lui Sha |
Real Time Syst. | 2 |
| 2006 | Optimal real-time sampling rate assignment for wireless sensor networksabstractHow to allocate computing and communication resources in a way that maximizes the effectiveness of control and signal processing, has been an important area of research. The characteristic of a multi-hop Real-Time Wireless Sensor Network raises new challenges. First, the constraints are more complicated and a new solution method is needed. Second, a distributed solution is needed to achieve scalability. This article presents solutions to both of the new challenges. The first solution to the optimal rate allocation is a centralized solution that can handle the more general form of constraints as compared with prior research. The second solution is a distributed version for large sensor networks using a pricing scheme. It is capable of incremental adjustment when utility functions change. This article also presents a new sensor device/network backbone architecture---Real-time Independent CHannels (RICH), which can easily realize multi-hop real-time wireless sensor networking. Xue (Steve) Liu, Qixin Wang 0001, Wenbo He 0003, Marco Caccamo, Lui Sha |
ACM Trans. Sens. Networks | 4 |
| 2005 | A Robust Implicit Access Protocol for Real-Time Wireless CollaborationabstractAdvances in wireless technology have brought us closer to extensive deployment of distributed real-time embedded systems connected through a wireless channel. The medium access control (MAC) layer protocol is critical in providing a real-time guarantee. We have devised a real-time wireless MAC protocol which, demonstrated in: both simulations and implementation, provides better throughput than existing protocols and predictable temporal behavior with minimal impact on node failures, packet losses and noise in the channel. Tanya L. Crenshaw, Ajay Tirumala, Spencer Hoke, Marco Caccamo |
ECRTS | 4 |
| 2005 | Spare CASH: Reclaiming Holes to Minimize Aperiodic Response Times in a Firm Real-Time EnvironmentabstractScheduling periodic tasks that allow some instances to be skipped produces spare capacity in the schedule. Only a fraction of this spare capacity is uniformly distributed and can easily be reclaimed for servicing aperiodic requests. The remaining fraction of the spare capacity is non-uniformly distributed, and no existing technique has been able to reclaim it. We present a method for improving the response times of aperiodic tasks by identifying the non-uniform holes in the schedule and adding these holes as extra capacity to the capacity queue of the CASH mechanism. The non-uniform holes can account for a significant portion of spare capacity, and reclaiming this capacity results in considerable improvements to aperiodic response times. Deepu C. Thomas, Sathish Gopalakrishnan, Marco Caccamo, Chang-Gun Lee |
ECRTS | 3 |
| 2005 | Time-Parameterized Sensing Task Model for Real-Time TrackingabstractThis paper proposes a novel task model in which its physical and temporal parameters are specified as time-parameterized functions and their values are finally determined at the actual dispatch time. This model is clearly differentiated from the classical task model where parameters are fixed at the job release time. The new model better suits sensing tasks in tracking applications, since the sensor parameters such as field-of-view and measurement duration can be properly adjusted at the actual sensing time. The new model, however, creates the cyclic dependency between task parameters and scheduling behavior, that is, the task parameters depend on scheduling behavior and the latter in turn depends on the former. This cyclic dependency makes the schedulability check even more difficult. We handle this difficulty by iterative convergence and probabilistic schedulability envelope, which provides an efficient online schedulability check. The experimental study shows that the new model significantly improves the effective capacity of tracking systems without losing track accuracy Min-Young Nam, Chang-Gun Lee, Kanghee Kim, Marco Caccamo |
RTSS | 4 |
| 2005 | Building Robust Wireless LAN for Industrial Control with DSSS-CDMA Cellphone Network ParadigmabstractDeploying wireless LAN for industrial control (IC-WLAN) has many benefits, such as mobility, low deployment cost and ease of reconfiguration. However, the top concern is robustness of wireless communications. Wireless control loops must be maintained under persistent adverse channel conditions, such as noise, large-scale path loss and fading. Many electro-magnetic interference sources in industrial environments, e.g. electric motor and welding, make wireless communication more challenging. The conventional IEEE 802.11 WLANs, which are designed for providing high bandwidth instead of high robustness, are therefore inappropriate for IC-WLAN. On the other hand, if the low data rate feature of industrial control is fully exploited by the state-of-the-art direct sequence spread spectrum (DSSS) technology, much higher robustness can be achieved. We hereby propose using DSSS-CDMA to build IC-WLAN, and exploiting the low data rate feature of industrial control loops for enhanced robustness. We carried out fine-grained physical layer simulations and Monte Carlo comparisons. The results show that DSSS-CDMA IC-WLAN provides much higher robustness than IEEE 802.11 WLAN, so that reliable wireless industrial control loops are made feasible. The DSSS-CDMA IC-WLAN scheme also opens up a new problem space for interdisciplinary study, involving real-time scheduling and resource management, communication, networking and control. In this paper, we study the resource management problems on maximizing robustness and minimizing control utility loss. Analytical resource optimization solutions are given Qixin Wang 0001, Xue (Steve) Liu, Weiqun Chen, Wenbo He 0003, Marco Caccamo |
RTSS | 5 |
| 2005 | Efficient Reclaiming in Reservation-Based Real-Time Systems with Variable Execution TimesabstractWe present a general CPU scheduling methodology for managing overruns in a real-time environment, where tasks may have different criticality, flexible timing constraints, shared resources, and variable execution times. The proposed method enhances, the constant bandwidth server (CBS) by providing two important extensions. First, it includes an efficient bandwidth sharing mechanism that reclaims the unused bandwidth to enhance task responsiveness. It is proven that the reclaiming mechanism does not violate the isolation property of the CBS and can be safely adopted to achieve temporal protection even when resource reservations are not precisely assigned. Second, the proposed method allows the CBS to work in the presence of shared resources. The enhancements achieved by the proposed approach turned out to be very effective with respect to classical CPU reservation schemes. The algorithm complexity is O(ln N), where N is the number of real-time tasks in the system, and its performance has been experimentally evaluated by extensive simulations. Marco Caccamo, Giorgio C. Buttazzo, Deepu C. Thomas |
IEEE Trans. Computers | 1 |
| 2004 | Collaborative Resource Allocation in Wireless Sensor Networks
Simone Giannecchini, Marco Caccamo, Chi-Sheng Shih 0001 |
ECRTS | 2 |
| 2004 | Finite-Horizon Scheduling of Radar Dwells with Online Template ConstructionabstractTiming constraints for radar tasks are usually specified in terms of the minimum and maximum temporal distance between successive radar dwells. We utilize the idea of feasible intervals for dealing with the temporal distance constraints. In order to increase the freedom that the scheduler can offer a high-level resource manager, we introduce a technique for nesting and interleaving dwells online while accounting for the energy constraint that radar systems need to satisfy. Further, in radar systems, the task set changes frequently and we advocate the use of finite horizon scheduling in order to avoid the pessimism that is inherent in schedulers that assume a task executes forever. We also develop the notion of modular schedule update which allows portions of a schedule to be altered without affecting the entire schedule, thereby simplifying the scheduler. Through extensive simulations, we validate our claims of providing greater scheduling flexibility without compromising on performance when compared with earlier work based on templates constructed offline. Sathish Gopalakrishnan, Marco Caccamo, Chi-Sheng Shih 0001, Chang-Gun Lee, Lui Sha |
RTSS | 2 |
| 2004 | Hard Real-Time Communication in Bus-Based NetworksabstractRoute selection is an important aspect of the design of real-time systems in which messages might have to travel over multiple hops to reach their destination and multiple paths exist between a source and a destination. The length of a route affects the ability to meet deadlines and greedy routing might leave certain messages with no feasible route. We consider bus-based networks on which periodic message transmissions need to be scheduled and present a technique for synthesizing routes such that all messages meet their deadlines. Our offline technique enables system designers to configure routes in a large-scale embedded system. In our solution, we allow message fragmentation and utilize multiple paths to satisfy the requirements of each message. The routing problem is NP-complete and our approximation algorithm is based on a linear programming formulation. In our methodology, we deal with both earliest deadline first and rate monotonic scheduling at each bus in the system. Apart from point-to-point messages, we discuss scheduling multicast messages to facilitate the publisher/subscriber model. Finally, we also mention some heuristics for online routing which might be of value in soft real-time systems. Sathish Gopalakrishnan, Lui Sha, Marco Caccamo |
RTSS | 3 |
| 2004 | Real Time Scheduling Theory: A Historical Perspective
Lui Sha, Tarek F. Abdelzaher, Karl-Erik Årzén, Anton Cervin, Theodore P. Baker, Alan Burns 0001, Giorgio C. Buttazzo, Marco Caccamo, John P. Lehoczky, Aloysius K. Mok |
Real Time Syst. | 8 |
| 2003 | The Capacity of Implicit EDF in Wireless Sensor NetworksabstractDistributed networks of wireless sensors/actuators will enable the reliable monitoring and intelligent control of the physical environment accomplishing different tasks ranging from space monitoring and surveillance to homeland security without human intervention. Motivated by the observation that many of these applications are safety critical and have hard real-time requirements, this paper focuses on the problem of providing delay and throughput guarantee to real-time messages in wireless sensor networks. It extends the cellular structure described by M. Caccamo et al. (2002) and analyzes the sensor network capacity when messages are scheduled with the implicit-EDF (earliest deadline first) algorithm. Marco Caccamo, Lynn Y. Zhang |
ECRTS | 1 |
| 2003 | Scheduling Real-Time Dwells Using Tasks with Synthetic PeriodsabstractThis paper addresses the problem of scheduling real-time dwells in multi-function phase array radar systems. To keep track of targets, a radar system must meet its timing and energy constraints. We propose a new task model for radar dwells to accurately characterize their timing parameters. We develop an algorithm of transforming every dwell task as a semi-period task so the dwell task can meet its timing constraint and the interarrival times of the task will not be a constant. We also develop an enhanced template-based scheduling algorithm to schedule such tasks to meet the timing and energy constraints. Simulation results show that this algorithm can significantly improve the resource utilization. Chi-Sheng Shih 0001, Sathish Gopalakrishnan, Phanindra Ganti, Marco Caccamo, Lui Sha |
RTSS | 4 |
| 2002 | An Implicit Prioritized Access Protocol for Wireless Sensor NetworksabstractRecent advances in wireless technology have brought us closer to the vision of pervasive computing where sensors/actuators can be connected through a wireless network. Due to cost constraints and the dynamic nature of sensor networks, it is undesirable to assume the existence of base stations connected by a wired backbone. In this paper, we present a network architecture suitable for sensor networks along with a medium access control protocol based on earliest deadline first. Marco Caccamo, Lynn Y. Zhang, Lui Sha, Giorgio C. Buttazzo |
RTSS | 1 |
| 2002 | Elastic Scheduling for Flexible Workload ManagementabstractAn increasing number of real-time applications related to multimedia and adaptive control systems require greater flexibility than classical real-time theory usually permits. We present a novel scheduling framework in which tasks are treated as springs with given elastic coefficients to better conform to the actual load conditions. Under this model, periodic tasks can intentionally change their execution rate to provide different quality of service and the other tasks can automatically adapt their periods to keep the system underloaded. The proposed model can also be used to handle overload conditions in a more flexible way and to provide a simple and efficient mechanism for controlling a system's performance as a function of the current load. Giorgio C. Buttazzo, Giuseppe Lipari, Marco Caccamo, Luca Abeni |
IEEE Trans. Computers | 3 |
| 2002 | Handling Execution Overruns in Hard Real-Time Control SystemsabstractIn many real-time control applications, the task periods are typically fixed and worst-case execution times are used in schedulability analysis. With the advancement of robotics, flexible visual sensing using cameras has become a popular alternative to the use of embedded sensors. Unfortunately, the execution time of visual tracking varies greatly. In such environments, control tasks have a normally short computation time, but also an occasional long computation time; therefore, the use of worst-case execution time is inefficient for controlling performance optimization. Nevertheless, to maintain the control stability, we still need to guarantee the schedulability of the task set, even if the worst case arises. In this paper, we propose an integrated approach to control performance optimization and task scheduling for control applications where the execution time of each task can vary greatly. We present an innovative approach to overrun management that allows us to fully utilize the processor for optimizing the control performance and yet guaranteeing the schedulability of all tasks under worst-case conditions. Marco Caccamo, Giorgio C. Buttazzo, Lui Sha |
IEEE Trans. Computers | 1 |
| 2001 | Aperiodic Servers with Resource ConstraintsabstractIntegrating soft and hard activities in a real-time environment has been an active area of research both under fixed priority scheduling and dynamic priority scheduling. Most of the existing work, however, has been done under the assumption that soft real-time tasks and hard real-time tasks are independent. The paper presents an efficient method that allows soft realtime aperiodic tasks and hard real-time tasks to share resources. Marco Caccamo, Lui Sha |
RTSS | 1 |
| 2000 | Elastic feedback controlabstractIn many real time control applications, the task periods are typically fixed and worst case execution times are used in schedulability analysis. With the advancement of robotics, flexible visual sensing using cameras has become a popular alternative to the use of embedded sensors. Unfortunately, the execution time of visual tracking varies greatly. In such environments, control tasks have a normally short computation time but also an occasional long computation time; therefore, the use of worst case execution time is inefficient for controlling performance optimization. Nevertheless, to maintain the control stability, we still need to guarantee the task set, even if the worst case arises. We propose an integrated approach to control performance optimization and task scheduling for control applications where the execution time of each task can vary greatly. We create an innovative approach to elastic control that allows us to fully utilize the processor to optimize the control performance and yet guarantee the schedulability of all tasks under worst case conditions. Marco Caccamo, Giorgio C. Buttazzo, Lui Sha |
ECRTS | 1 |
| 2000 | Capacity Sharing for Overrun ControlabstractPresents a general scheduling methodology for managing overruns in a real-time environment, where tasks may have different criticalities and flexible timing constraints. The proposed method achieves isolation among tasks through a resource reservation mechanism which bounds the effects of task interference but which also performs efficient reclamation of the unused computation times in order to relax the utilization constraints imposed by isolation. The enhancements achieved by the proposed approach were found to be very effective with respect to classical reservation schemes. The performance has been evaluated by implementing the algorithm on a real-time kernel. The runtime overhead introduced by the scheduling mechanism has also been investigated with specific experiments, in order for this to be taken into account in the schedulability analysis. However, this overhead was found to be negligible in most practical cases. Marco Caccamo, Giorgio C. Buttazzo, Lui Sha |
RTSS | 1 |
| 1999 | Sharing Resources among Periodic and Aperiodic Tasks with Dynamic DeadlinesabstractIn this paper, we address the problem of scheduling hybrid task sets consisting of hard periodic and soft aperiodic tasks that may share resources in exclusive mode in a dynamic environment, where tasks are scheduled based on their deadlines. Bounded blocking on exclusive resources is achieved by means of a dynamic resource access protocol which also prevents deadlocks and chained blocking. A tunable servicing technique is used to improve aperiodic responsiveness in the presence of resource constraints. The schedulability analysis is also extended to the case in which aperiodic deadlines vary at runtime. The results achieved in this paper can also be used for developing adaptive real-time systems, where task deadlines or periods can change to conform to new load conditions. Marco Caccamo, Giuseppe Lipari, Giorgio C. Buttazzo |
RTSS | 1 |
| 1999 | Minimizing Aperiodic Response Times in a Firm Real-Time EnvironmentabstractIn certain real-time applications, ranging from multimedia to telecommunication systems, timing constraints can be more flexible than scheduling theory usually permits. In this paper, we deal with the problem of scheduling hybrid sets of tasks, consisting of firm periodic tasks (i.e. tasks with deadlines which can occasionally skip one instance) and soft aperiodic requests, which have to be served as soon as possible to achieve good responsiveness. We propose and analyze an algorithm, based on a variant of earliest-deadline-first scheduling, which exploits skips to minimize the response time of aperiodic requests. One of the most interesting features of our algorithm is that it can easily be tuned to balance performance vs. complexity, for adapting it to different application requirements. Extensive simulation experiments show the effectiveness of the proposed approach with respect to existing methods. Schedulability bounds are also derived to perform off-line analysis. Giorgio C. Buttazzo, Marco Caccamo |
IEEE Trans. Software Eng. | 2 |
| 1997 | Exploiting skips in periodic tasks for enhancing aperiodic responsivenessabstractIn certain real-time applications, ranging from multimedia to telecommunication systems, timing constraints can be more flexible than scheduling theory usually permits. For example, in video reception, missing a deadline is acceptable, provided that most deadlines are met. We deal with the problem of scheduling hybrid sets of tasks, consisting of firm periodic tasks (i.e., tasks with deadlines which can occasionally skip one instance) and soft aperiodic requests, which have to be served as soon as possible to minimize their average response time. We propose and analyze an algorithm, based on a variant of earliest deadline first scheduling, which exploits skips to enhance the response time of aperiodic requests. Schedulability bounds are also derived to perform off-line analysis. Marco Caccamo, Giorgio C. Buttazzo |
RTSS | 1 |