Gabriel Parmer

dblp:86/4930 · DBLP profile ↗
← Back
41ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0003-1343-4492ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-authorSecurity and privacy · 2
YearPublicationVenuePosition
2025 Poster: Deadline-aware Serverless SLAs for the Multi-tenant Edge
abstract
The resource-constrained nature of edge computing makes serverless platforms an ideal fit due to their elasticity and cost-effectiveness. However, supporting multi-tenant environments while meeting strict per-request deadlines is challenging, as existing systems fail to reconcile resource reservations with deadline guarantees. In this paper, we introduce DaSLA, a deadline-aware multi-tenant serverless framework that combines techniques like Earliest Deadline First (EDF) scheduling, online Demand Bound Function (DBF) analysis, and a Sliding Window Server (SWS) algorithm to ensure both deadline adherence and tenant resource reservations. This approach enables high-performance, deadline-aware, and multi-tenant serverless computing tailored for the edge.
Emil Abbasov, Gabriel Parmer
SEC2
2025 SledgeScale: Load-Aware Dispatch and Deadline-Driven Scheduling for Scalable, Dense Serverless Computing in Edge Data Centers
abstract
Serverless and Edge Computing ought to be a perfect match to create flexible, efficient, and responsive applications for emerging areas like augmented reality and autonomous vehicles. Unfortunately, current serverless designs incur high overheads—preventing submillisecond execution—and consume large amounts of resources—preventing dense deployment within constrained edge environments. We present a platform to overcome these challenges, while providing strong isolation and performance in multi-tenant edge data centers.
Xiaosu Lyu, Emil Abbasov, Sean McBride, Gabriel Parmer, Timothy Wood 0001
SEC4
2025 SPR: Shielded Processor Reservations with Bounded Management Overhead
abstract
With growing hardware consolidation in modern computational infrastructures, ensuring predictable CPU allocation has become increasingly critical. Processor reservations, usually realized through rate-limiting servers, play an essential role in providing such predictability by precisely controlling when and how long each task may execute. In theory, ratelimiting servers provide strong temporal isolation, meaning that a task's timely access to its guaranteed budget is not contingent on the behavior of any other task in the system. However, in practice, these guarantees are easily undermined by the realities of actual hardware and shortcomings in naïve reservation implementations. When confronted with malicious tasks using the reservation policy itself to attack the very security and rate properties it is meant to uphold, temporal isolation breaks down in current implementations. In response, this paper presents shielded processor reservation (SPR) scheduling, a novel approach that ensures that at most two reservations are processed per scheduler invocation and integrates deferred timer handling, early replenishment processing, and processor access granularity guarantees to provide robust temporal isolation. We implement SPR and existing rate-limiting servers in the Composite operating system and evaluate its performance. The results demonstrate that SPR provides reliable rate-limiting with low overhead while also mitigating vulnerabilities that can be exploited to attack existing reservation systems.
Esma Kökten, Gabriel Parmer, Björn B. Brandenburg
RTAS2
2025 Janus: OS Support for a Secure, Fast Control-Plane
abstract
The emergence of the edge cloud, empowered by advanced wireless technologies such as 5G, is aimed at delivering predictable, interactive services within low milliseconds. Even in the traditional cloud, latency-sensitive services are pervasive. To enable low-latency software, much focus has been on data-plane optimizations for fast processing of requests. Unfortunately, these efforts alone are insufficient: effective, low-latency services also require advances in the control-plane. However, control-plane operations require strong spatial and temporal isolation between tenant computations, thus can result in significant overhead. This paper introduces Janus, an OS abstraction for a flexible control-plane which provides both strong isolation and low latency through the use of pervasive kernel-bypass. Janus leverages Memory Protection Key (MPK) hardware to enable low-cost control operations for protected procedure call (PPC) and thread dispatch - two essential building blocks of spatial and temporal isolation. By transparently improving these fundamental control-plane operations, Janus enables efficient and predictable user-defined and customizable system control policies and mechanisms, while maintaining strong isolation. We evaluate Janus's ability to extensibly define new control mechanisms, increase the efficiency of an existing RTOS, and support low-latency services in a memcached server. Compared to a Linux approach, a specialized latency-sensitive control plane using Janus provides over a 6x improvement in throughput, while providing 99th percentile tail latencies almost 3x lower. Demonstrating the utility of an efficient control plane, Janus improves tail latency in a multi-tenant system by orders of magnitude.
Wenyuan Shao, Evan Stella, Linnea Dierksheide, Phani Kishore Gadepalli, Gabriel Parmer
RTAS6
2024 Byways: High-Performance, Isolated Network Functions for Multi-Tenant Cloud Servers
abstract
Network functions (NFs) have become pervasive in data centers as a means to monitor and transform traffic as it flows between services. Softwarization of the network has further added to the diversity of functions that can be deployed, yet managing the performance, efficiency, tenant-customizability, and security of these functions remains a major challenge. We present Byways, an abstraction that provides facilitates to safely deploy NFs alongside end-host VMs in a multi-tenant cloud environment. Byways guarantee strict isolation between the host system, the network functions, and VM-based cloud applications, while still maintaining high performance. A Byway manages a specific set of services, and an associated NF only processes flows associated with those services, using per-byway resources (e.g., processing time). This separation of end-host traffic across Byways provides strong fault isolation - a failing NF does not impact other services. Byways augment this isolation with per-Byway access rights that restrict a NFs access (e.g., read, write, drop) to the flow, limiting the impact of a faulty NF on even its own services.
Yuan Gao 0041, Gabriel Parmer, Timothy Wood 0001
SoCC3
2024 OmniWasm: Efficient, Granular Fault Isolation and Control-Flow Integrity for Arm Microcontrollers
abstract
Traditional embedded systems often ignore security, instead focusing on simplicity. Unfortunately, increasingly per-vasive network connectivity exposes these systems to malicious input, and the consolidation of code of various qualities and sources onto single systems increases the chance of errant behavior. Microcontrollers do not often have advanced security facilities to help prevent malicious threats - even the use of memory protection to constrain the ill-effects of a compromise is not pervasive. Software techniques to provide strong security properties for application execution often focus on limiting accessible memory through dynamic checks on memory accesses through software fault isolation (SFI), and on ensuring that software cannot suffer from control-flow hijack attacks through control-flow integrity (CFI). This paper introduces OmniWasm, which provides sandboxes in which applications execute that provide both SFI and CFI. OmniWasm focuses on two core contributions. First, it reduces the overhead of these techniques by using a novel application of common, but obscure memory operations to both limit sandboxed memory accesses (SFI) using hardware, and allow memory accesses outside the sandbox to enable CFI. Second, it focuses on enabling embedded soft-ware to be decomposed into multiple sandboxes by providing optimized communication facilities between them that avoid thread context switching overheads and providing single-copy message passing. We show that the overheads of OmniWasm are better than existing software address validation techniques while also providing better memory utilization. We also show that OmniWasm enables significant performance gains for inter-sandbox communication across varying message sizes.
Maorui Bai, Runyu Pan, Gabriel Parmer
RTAS3
2022 The Thundering Herd: Amplifying Kernel Interference to Attack Response Times
abstract
Embedded and real-time systems are increasingly attached to networks. This enables broader coordination beyond the physical system, but also opens the system to attacks. The increasingly complex workloads of these systems include software of varying assurance levels, including that which might be susceptible to compromise by remote attackers. To limit the impact of compromise, μ-kernels focus on maintaining strong memory protection domains between different bodies of software, including system services. They enable limited coordination between processes through Inter-Process Communication (IPC). Real-time systems also require strong temporal guarantees for tasks, and thus need temporal isolation to limit the impact of malicious software. This is challenging as multiple client threads that use IPC to request service from a shared server will impact each other’s response times.To constrain the temporal interference between threads, modern μ-kernels often build priority and budget awareness into the system. Unfortunately, this paper demonstrates that this is more challenging than previously thought. Adding priority awareness to IPC processing can lead to significant interference due to the kernel’s prioritization logic. Adding budget awareness similarly creates opportunities for interference due to the budget tracking and management operations. In both situations, a Thundering Herd of malicious threads can significantly delay the activation of mission-critical tasks. The Thundering Herd effects are evaluated on seL4 and results demonstrate that high-priority threads can be delayed by over 100,000 cycles per malicious thread. This paper reveals a challenging dilemma: the temporal protections μ-kernels add can, themselves, provide means of threatening temporal isolation. Finally, to defend the system, we identify and empirically evaluate possible mitigations, and propose an admission-control test based upon an interference-aware analysis.
Samuel Mergendahl, Samuel Jero, Bryan C. Ward, Juliana Furgala, Gabriel Parmer, Richard Skowyra
RTAS5
2022 SBIs: Application Access to Safe, Baremetal Interrupt Latencies*
abstract
The continued increase in Cyber-Physical System (CPS) complexity and tightening of Size, Weight and Power (SWaP) constraints are driving the need for consolidation of software tasks onto fewer microcontrollers. Many embedded systems, prominently including those in the Internet-of-Things (IoT), use software packages from multiple untrusted sources, while their network interfaces expose new attack surfaces that are not present in traditional off-line devices. Increased consolidation with untrusted code of various assurance levels complicates system design, and requires increased spatial and temporal isolation between the applications. Current microcontroller protection domain designs are limited by their long interrupt latencies to isolated applications, forcing the system designers to place timing-sensitive application code into the kernel interrupt handlers, trading spatial isolation for tightly-bounded temporal predictability.SBI (Secure Baremetal Interrupt) enables zero-software-cost delivery of interrupts to protection domains in a secure manner that maintains isolation. We demonstrate an implementation of SBI using the new hardware-accelerated interrupt delivery features on TrustZone-M-enabled microcontrollers. This implementation reduces interrupt latencies by up to 95%, while maintaining strong spatial and temporal isolation. We believe SBI could significantly enable future real-time systems that require both isolation and high responsiveness.
Runyu Pan, Gabriel Parmer
RTAS2
2022 Edge-RT: OS Support for Controlled Latency in the Multi-Tenant, Real-Time Edge
abstract
Embedded and real-time devices in many domains are increasingly dependent on network connectivity. The ability to offload computations encourages Cost, Size, Weight and Power (C-SWaP) optimizations, while coordination over the network effectively enables systems to sense the environment beyond their own local sensors, and to collaborate globally. The promise is significant: Autonomous Vehicles (AVs) coordinating with each other through infrastructure, factories aggregating data for global optimization, and power-constrained devices leveraging offloaded inference tasks. Low-latency wireless (e.g., 5G) technologies paired with the edge cloud, are further enabling these trends. Unfortunately, computation at the edge poses significant challenges due to the challenging combination of limited resources, required high performance, security due to multi-tenancy, and real-time latency. This paper introduces Edge-RT, a set of OS extensions for the edge designed to meet the end-to-end (packet reception to transmission) deadlines across chains of computations. It supports strong security by executing a chain per-client device, thus isolating tenant and device computations. Despite a practical focus on deadlines and strong isolation, it maintains high system efficiency. To do so, Edge-RT focuses on per-packet deadlines inherited by the computations that operate on it. It introduces mechanisms to avoid per-packet system overheads, while trading only bounded impacts on predictable scheduling. Results show that compared to Linux and EdgeOS, Edge-RT can both maintain higher throughput and meet significantly more deadlines both for systems with bimodal workloads with utilization above 60%, in the presence of malicious tasks, and as the system scales up in clients.
Wenyuan Shao, Bite Ye, Huachuan Wang, Gabriel Parmer, Yuxin Ren 0001
RTSS4
2022 Sharing non-cache-coherent memory with bounded incoherence
abstract
Summary Cache coherence in modern computer architectures enables easier programming by sharing data across multiple processors. Unfortunately, it can also limit scalability due to cache coherency traffic initiated by competing memory accesses. Rack‐scale systems introduce shared memory across a whole rack, but without inter‐node cache coherence. This poses memory management and concurrency control challenges for applications that must explicitly manage cache‐lines. To fully utilize rack‐scale systems for low‐latency and scalable computation, applications need to maintain cached memory accesses in spite of non‐coherency. This paper introduces Bounded Incoherence, a memory consistency model that enables cached access to shared data‐structures in non‐cache‐coherency memory. It ensures that updates to memory on one node are visible within at most a bounded amount of time on all other nodes. We evaluate this memory model on modified PowerGraph graph processing framework, and boost its performance by 30% with eight sockets by enabling cached‐access to data‐structures.
Yuxin Ren 0001, Gabriel Parmer, Dejan S. Milojicic
Concurr. Comput. Pract. Exp.2
2021 Practical Principle of Least Privilege for Secure Embedded Systems
abstract
Many embedded systems have evolved from simple bare-metal control systems to highly complex network-connected systems. These systems increasingly demand rich and feature-full operating-systems (OS) functionalities. Furthermore, the network connectedness offers attack vectors that require stronger security designs. To that end, this paper defines a prototypical RTOS API called Patina that provides services common in featurerich OSes (e.g., Linux) but absent in more trustworthy μ -kernel based systems. Examples of such services include communication channels, timers, event management, and synchronization. Two Patina implementations are presented, one on Composite and the other on seL4, each of which is designed based on the Principle of Least Privilege (PoLP) to increase system security. This paper describes how each of these μ -kernels affect the PoLP based design, as well as discusses security and performance tradeoffs in the two implementations. Results of comprehensive evaluations demonstrate that the performance of the PoLP based implementation of Patina offers comparable or superior performance to Linux, while offering heightened isolation.
Samuel Jero, Juliana Furgala, Runyu Pan, Phani Kishore Gadepalli, Alexandra Clifford, Bite Ye, Roger I. Khazan, Bryan C. Ward, Gabriel Parmer, Richard Skowyra
RTAS9
2021 Precise Cache Profiling for Studying Radiation Effects
abstract
Increased access to space has led to an increase in the usage of commodity processors in radiation environments. These processors are vulnerable to transient faults such as single event upsets that may cause bit-flips in processor components. Caches in particular are vulnerable due to their relatively large area, yet are often omitted from fault injection testing because many processors do not provide direct access to cache contents and they are often not fully modeled by simulators. The performance benefits of caches make disabling them undesirable, and the presence of error correcting codes is insufficient to correct for increasingly common multiple bit upsets. This work explores building a program’s cache profile by collecting cache usage information at an instruction granularity via commonly available on-chip debugging interfaces. The profile provides a tighter bound than cache utilization for cache vulnerability estimates (50% for several benchmarks). This can be applied to reduce the number of fault injections required to characterize behavior by at least two-thirds for the benchmarks we examine. The profile enables future work in hardware fault injection for caches that avoids the biases of existing techniques.
Robert Gifford, Gedare Bloom, Gabriel Parmer, Rahul Simha
ACM Trans. Embed. Comput. Syst.4
2020 Sledge: a Serverless-first, Light-weight Wasm Runtime for the Edge
abstract
Emerging IoT applications with real-time latency constraints require new data processing systems operating at the Edge. Serverless computing offers a new compelling paradigm, where a user can execute a small application without handling the operational issues of server provisioning and resource management. Despite a variety of existing commercial and open source serverless platforms (utilizing VMs and containers), these solutions are too heavy-weight for a resource-constrained Edge systems (due to large memory footprint and high invocation time). Moreover, serverless workloads that focus on per-client, short-running computations are not an ideal fit for existing general purpose computing systems.
Phani Kishore Gadepalli, Sean McBride, Gregor Peach, Ludmila Cherkasova, Gabriel Parmer
Middleware5
2020 Slite: OS Support for Near Zero-Cost, Configurable Scheduling *
abstract
Despite over 35 years of wildly changing requirements and applications, real-time systems have treated the kernel implementation of system scheduling policy as given. New time management policies are either adapted to the kernel, and rarely adopted, or emulated in user-level under the restriction that they must stay within the confines of the underlying system’s mechanisms. This not only hinders the agility of the system to adapt to new requirements, but also harms non-functional properties of the system such as the effective use of parallelism, and the application of isolation to promote security. This paper introduces Slite, a system designed to investigate the possibilities of near zero-cost scheduling of system-level threads at user-level. This system efficiently and predictably enables the user-level implementation of configurable scheduling policies. This capability has wide-ranging impact, and we investigate how it can increase the isolation, thus dependability, of a user-level real-time OS, and how it can provide a real-time parallel runtime with better analytical properties using thread-based – rather than the conventional task-based – scheduling. We believe these benefits motivate a strong reconsideration of the fundamental scheduling structure in future real-time systems.
Phani Kishore Gadepalli, Runyu Pan, Gabriel Parmer
RTAS3
2020 Fine-Grained Isolation for Scalable, Dynamic, Multi-tenant Edge Clouds
Yuxin Ren 0001, Guyue Liu, Vlad Nitu, Wenyuan Shao, Riley Kennedy, Gabriel Parmer, Timothy Wood 0001, Alain Tchana
USENIX ATC6
2020 eWASM: Practical Software Fault Isolation for Reliable Embedded Devices
abstract
As we connect more microcontrollers to the Internet and employ them to control the physical world around us, their reliability and security are increasingly important. Many microcontrollers provide limited facilities for hardware isolation, and real-time OSes offer custom APIs, that require coupling applications into the ecosystem and abstractions of that specific OS to leverage isolation. This article investigates the use of software sandboxing of applications to support isolation for resource-constrained devices. Toward this, we detail the design of eWASM, a processes abstraction that adapts a popular sandbox, Wasm, for microcontrollers. eWASM provides a runtime to constrain memory accesses and control flow, enabled by our aWsm Wasm compiler. We discuss and evaluate its multiple implementations that effectively trade time and space, optimizing for the constraints of embedded systems. This enables popular languages (e.g., C) to be effectively sandboxed by software. We demonstrate performance within 40% of native C on Polybench. We believe this is a practical and compelling result for many IoT domains, and it represents the first compiled sandboxing environment for microcontrollers. We show that restrictions of the current Wasm specification lead to significant memory consumption and provide suggestions for the creation of an embedded-specific Wasm variant.
Gregor Peach, Runyu Pan, Zhuoyi Wu, Gabriel Parmer, Christopher Haster, Ludmila Cherkasova
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 Scalable Data-structures with Hierarchical, Distributed Delegation
abstract
Scaling data-structures up to the increasing number of cores provided by modern systems is challenging. The quest for scalability is complicated by the non-uniform memory accesses (NUMA) of multi-socket machines that often prohibit the effective use of data-structures that span memory localities. Conventional shared memory data-structures using efficient non-blocking or lock-based implementations inevitably suffer from cache-coherency overheads, and non-local memory accesses between sockets. Multi-socket systems are common in cloud hardware, and many products are pushing shared memory systems to greater scales, thus making the ability to scale data-structures all the more pressing.
Yuxin Ren 0001, Gabriel Parmer
Middleware2
2019 Chaos: a System for Criticality-Aware, Multi-Core Coordination
abstract
The incentive to minimize size, weight and power (SWaP) in embedded systems has driven the consolidation both of disparate processors into single multi-core systems, and of software of various functionalities onto shared hardware. These consolidated systems must address a number of challenges that include providing strong isolation of the highly-critical tasks that impact human or equipment safety from the more feature-rich, less trustworthy applications, and the effective use of spare system capacity to increase functionality. The coordination between high and low criticality tasks is particularly challenging, and is common, for example, in autonomous vehicles where controllers, planners, sensor fusion, telemetry processing, cloud communication, and logging all must be orchestrated together. In such a case, they must share the code of the software run-time system that manages resources, and provides communication abstractions. This paper presents the Chaos system that uses devirtualization to extract high-criticality tasks from shared software environments, thus alleviating interference, and runs them in a minimal runtime. To maintain access to more feature-rich software, Chaos provides low-level coordination through proxies that tightly bound the overheads for coordination. We demonstrate Chaos's ability to scalably use multiple cores while maintaining high isolation with controlled inter-criticality coordination. For a sensor/actuation loop in satellite software experiencing inter-core interference, Chaos lowers processing latency by a factor of 2.7, while reducing worst-case by a factor 3.5 over a real-time Linux variant.
Phani Kishore Gadepalli, Gregor Peach, Gabriel Parmer, Joseph Espy, Zach Day
RTAS3
2019 Challenges and Opportunities for Efficient Serverless Computing at the Edge
abstract
Serverless computing frameworks allow users to execute a small application (dedicated to a specific task) without handling operational issues such as server provisioning, resource management, and resource scaling for the increased load. Serverless computing originally emerged as a Cloud computing framework, but might be a perfect match for IoT data processing at the Edge. However, the existing serverless solutions, based on VMs and containers, are too heavy-weight (large memory footprint and high function invocation time) for operating efficiency and elastic scaling at the Edge. Moreover, many novel IoT applications require low-latency data processing and near real-time responses, which makes the current cloud-based serverless solutions unsuitable. Recently, WebAssembly (Wasm) has been proposed as an alternative method for running serverless applications at near-native speeds, while having a small memory footprint and optimized invocation time. In this paper, we discuss some existing serverless solutions, their design details, and unresolved performance challenges for an efficient serverless management at the Edge. We outline our serverless framework, called aWsm, based on the WebAssembly approach, and discuss the opportunities enabled by the aWsm design, including function profiling and SLO-driven performance management of users' functions. Finally, we present an initial assessment of aWsm performance featuring average startup time (12μs to 30μs) and an economical memory footprint (ranging from 10s to 100s of kB) for a subset of MiBench microbenchmarks used as functions.
Phani Kishore Gadepalli, Gregor Peach, Ludmila Cherkasova, Robert C. Aitken, Gabriel Parmer
SRDS5
2019 MxU: Towards Predictable, Flexible, and Efficient Memory Access Control for the Secure IoT
abstract
The advanced functionality requirements of modern embedded and Internet of Things (IoT) devices -- from autonomous vehicles, to city and power-grid management -- are driving an ever-increasing software complexity. At the same time, the pervasive internet connections of these systems necessitate the fundamental design of security into these devices. The isolation of complex features from those that are critical through protection domains is an effective means to constrain the scope of faults and security breaches. Common hardware-provided memory facilities to enforce protection domains through memory access control -- including Memory Management Units (MMUs) usually found in microprocessors, and Memory Protection Units (MPUs) usually found in microcontrollers -- must meet the goals of enabling flexible, efficient and dynamic management of memory , and must enable tight bounds on the worst-case execution of critical code. Unfortunately, current system memory management facilities are ill-prepared to handle this challenge: MMUs that use extensive caches to achieve strong average-case performance suffer from debilitating worst-case and even average-case behavior under hefty interference, while MPUs struggle to provide flexible memory management. This paper details MxU, a memory protection and allocation abstraction that integrates temporal specifications into the memory management subsystem, to enable portable code to achieve both predictable, tightly-bounded execution and dynamic management across both MMU- and MPU-based systems. We implement MxU in the Composite microkernel, and evaluate its flexibility and predictability over two different architectures: a MPU-based Cortex-M7 microcontroller and a MMU-based Cortex-A9 microprocessor using a suite of modern applications including neural network-based inference, SQLite, and a javascript runtime. For MMU-based systems, MxU reduces application TLB stall by up to 68.0%. For MPU-based systems, MxU enables flexible dynamic memory management often with application overheads of 1%, increasing to 6.1% under significant interference.
Runyu Pan, Gabriel Parmer
ACM Trans. Embed. Comput. Syst.2
2018 Predictable Virtualization on Memory Protection Unit-Based Microcontrollers
abstract
With the increasing penetration of embedded systems into the consumer market there is a pressure to have all of inexpensiveness, predictability, reliability, and security. As these systems are often attached to networks and execute complex code from varying sources, reliability and security become essential. To maintain low price and small power budgets, many systems use small microcontrollers with limited memory (on the order of 128KB of SRAM). Unfortunately, the isolation and protection facilities of these systems are often lackluster, making a principled treatment of reliability and security difficult. This paper details a system that provides isolation along the three dimensions of CPU, memory, and I/O on small microcontrollers. A key challenge is providing a effective means of harnessing the limited hardware memory protection facilities of microcontrollers. This is achieved through a combination of a static analysis to make the most of limited hardware protection facilities, and a run-time based on our Composite OS. On this foundation, we build a virtualization infrastructure to execute multiple embedded real-time operating systems predictably. We show that VMs based on FreeRTOS achieve reasonable efficiency and predictability, while easily enabling scaling up to 8 VMs in 512 KB SRAM.
Runyu Pan, Gregor Peach, Yuxin Ren 0001, Gabriel Parmer
RTAS4
2018 Scalable Memory Reclamation for Multi-Core, Real-Time Systems
abstract
A core challenge in best utilizing an increasing number of cores in real-time systems is addressing the problem of efficient and predictable resource sharing. Traditional mechanisms for mutual exclusion, such as locks, limit parallelism due to serialized resource access. Relaxing mutual exclusion, reader-writer locks enable selective parallelism for a subset of accesses, but can suffer from increased implementation overheads. In all such implementations, the costs of cache-coherency alone can be prohibitive for an increasing number of cores. This paper investigates the use of techniques such as Read-Copy Update (RCU) to enable truly parallel access to data-structures. Such techniques optimize for data-structure read-paths, and can completely avoid stores to shared structures, thus avoiding cache-coherency overheads. We show that existing implementations of preemptive RCU aren't designed to provide real-time latencies, and require a potentially unbounded amount of dynamically allocated memory. Thus, we introduce two new implementations that are both predictable and efficient, and a matching analysis that establishes bounds on memory consumption. We additionally provide a schedulability analysis that demonstrates the effectiveness of scalable read-side operations, achieving consistently higher schedulability than existing techniques. We further apply the analysis to provide admission control for a soft real-time application to both achieve higher throughput than existing approaches (up to 40% higher) while limiting 99th percentile read-path latencies (4x lower than existing techniques).
Yuxin Ren 0001, Guyue Liu, Gabriel Parmer, Björn B. Brandenburg
RTAS3
2017 Temporal Capabilities: Access Control for Time
abstract
Embedded systems are increasingly required to handle code of various qualities that must often be isolated, yet predictably share resources. This has motivated the isolation of, for example, mission-critical code from best-effort features using isolation structures such as virtualization. Such systems usually focus on limiting interference between subsystems, which complicates the increasingly common functional dependencies between them. Though isolation must be paramount, the fundamental goal of efficiently sharing hardware motivates a principled mechanism for cooperating between subsystems. This paper introduces Temporal Capabilities (TCaps) which integrate CPU management into a capability-based access-control system and distribute authority for scheduling. In doing so, the controlled temporal coordination between subsystems becomes a first-class concern of the system. By enabling temporal delegations to accompany activations and requests for service, we apply TCaps to a virtualization environment with a shared VM for orchestrating I/O. We show that TCaps, unlike prioritizations and carefully chosen budgets, both meet deadlines for a hard real-time subsystem, and maintain high throughput for a best-effort subsystem.
Phani Kishore Gadepalli, Robert Gifford, Lucas Baier, Michael Kelly, Gabriel Parmer
RTSS5
2016 SuperGlue: IDL-Based, System-Level Fault Tolerance for Embedded Systems
abstract
As the processor feature sizes shrink, mitigating faults in low level system services has become a critical aspect of dependable system design. In this paper we introduce SuperGlue, an interface description language (IDL) and compiler for recovery from transient faults in a component-based operating system. SuperGlue generates code for interface-driven recovery that uses commodity hardware isolation, micro-rebooting, and interface-directed fault recovery to provide predictable and efficient recovery from faults that impact low-level system services. SuperGlue decreases the amount of recovery code system designers need to implement by an order of magnitude, and replaces it with declarative specifications. We evaluate SuperGlue with a fault injection campaign in low-level system components (e.g., memory mapping manager and scheduler). Additionally, we evaluate the performance of SuperGlue in a web-server application. Results show that SuperGlue improves system reliability with only a small performance degradation of 11.84%.
Jiguo Song, Gedare Bloom, Gabriel Parmer
DSN3
2016 Parallel sections: scaling system-level data-structures
abstract
As systems continue to increase the number of cores within cache coherency domains, traditional techniques for enabling parallel computation on data-structures are increasingly strained. A single contended cache-line bouncing between different caches can prohibit continued performance gains with additional cores. New abstractions and mechanisms are required to reassess how data-structure consistency can be provided, while maintaining stable per-core access latencies.
Tim Stamler, Gabriel Parmer
EuroSys3
2016 CBufs: efficient, system-wide memory management and sharing
abstract
Modern systems are composed of many different protection domains separating privilege levels, subsystems, users, clients, and software of differing levels of assurance. System-wide memory management must consider not only allocation to single processes, but also efficient sharing of data across protection domains, and the allocation of memory based on the performance of applications that span multiple protection domains. This paper introduces the CBuf system for the global management of virtual and physical memory, including zero-copy sharing between protection domains. We present the design and implementation of both garbage collection techniques to enable efficient sharing, and policies that balance memory between protection domains specifically to satisfy system and application constraints such as quality of service. We show that a CBuf-enabled webserver achieves over a factor of 2.5 throughput speedup while using less processing time than Apache on Linux, and that the system can intentionally control system throughput through intelligent memory allocation.
Yuxin Ren 0001, Gabriel Parmer, Teo Georgiev, Gedare Bloom
ISMM2
2015 C'Mon: a predictable monitoring infrastructure for system-level latent fault detection and recovery
abstract
Embedded and real-time systems must balance between many often conflicting goals including predictability, high utilization, efficiency, reliability, and SWaP (size, weight, and power). Reliability is particularly difficult to achieve without significantly impacting the other factors. Though reliability solutions exist for application-level, they are invalidated by system-level faults that are particularly difficult to detect and recover from. This paper presents the C'Mon system for predictably and efficiently monitoring system-level execution, and validating that it conforms with the high-level analytical models that underlie the timing guarantees of the system. Latent faults such as timing errors, incorrect scheduler decisions, unbounded priority inversions, or deadlocks are detected, the faulty component is identified, and using previous work in system recovery, the system is brought back to a stable state - all without missing deadlines.
Jiguo Song, Gabriel Parmer
RTAS2
2015 SPeCK: a kernel for scalable predictability
abstract
Multi- and many-core systems are increasingly prevalent in embedded systems. Additionally, isolation requirements between different partitions and criticalities are gaining in importance. This difficult combination is not well addressed by current software systems. Parallel systems require consistency guarantees on shared data-structures often provided by locks that use predictable resource sharing protocols. However, as the number of cores increase, even a single shared cache-line (e.g. for the lock) can cause significant interference. In this paper, we present a clean-slate design of the SPeCK kernel, the next generation of our COMPOSITE OS, that attempts to provide a strong version of scalable predictability - where predictability bounds made on a single core, remain constant with an increase in cores. Results show that, despite using a non-preemptive kernel, it has strong scalable predictability, low average-case overheads, and demonstrates better response-times than a state-of-the-art preemptive system.
Yuxin Ren 0001, Matt Scaperoth, Gabriel Parmer
RTAS4
2014 FJOS: Practical, predictable, and efficient system support for fork/join parallelism
abstract
With the increasing use of multi- and many-core processors in real-time and embedded systems, software's ability to utilize those cores to increase system capability and functionality is important. Of particular interest is intra-task parallelism whereby a single task is able to harness the computational power of multiple cores to do processing of a complexity that is untenable on a single core. This paper introduces the design and implementation of FJOS, a system supporting predictable and efficient fork/join, intra-task parallelism. FJOS is implemented using abstractions that are close to the hardware, and decouples parallelism management, from thread coordination, yielding efficient fast-path operations. Compared to a traditional fork/join implementation, results show that FJOS has less overhead, is more scalable up to 40 cores, and can generally make better use of parallelism. We modify a response-time analysis to integrate system overheads to assess schedulability in a hard real-time environment, and design an effective algorithm for assigning task computation to cores. This assignment more than triples effective system utilization, and when implementation overheads are considered, FJOS maintains high system utilizations, thus providing a strong foundation for predictable, real-time intra-task parallelism.
Gabriel Parmer
RTAS2
2013 Predictable, Efficient System-Level Fault Tolerance in C^3
abstract
Predictable reliability is an increasingly important aspect of embedded and real-time systems. This includes the ability to recover from unknown faults in a manner that maintains system timing guarantees, even when these faults occur within system components. This paper presents the C3system, which is the first system implementation we know of for predictable, system-level fault tolerance that doesn't require physical redundancy. We introduce both the system design, and two timing analyses that enable the predictable recovery from faults in operating system components, and identify recovery inversion as a main impediment to schedulable recovery. C3provides fault-tolerance for low-level system components using a combination of efficient u-reboots, and an interface-driven mechanism to recreate component state. C3introduces on-demand recovery that properly prioritizes aspects of the recovery process to avoid this inversion and not inhibit system timeliness. We compare this system to both eager recovery, and to check pointing of a Para virtualized real-time OS.
Jiguo Song, John Wittrock, Gabriel Parmer
RTSS3
2013 Predictable and configurable component-based scheduling in the Composite OS
abstract
This article presents the design of user-level scheduling hierarchies in the C omposite component-based system. The motivation for this is centered around the design of a system that is both dependable and predictable, and which is configurable to the needs of specific applications. Untrusted application developers can safely develop services and policies, that are isolated in protection domains outside the kernel. To ensure predictability, C omposite enforces timing control over user-space services. Moreover, it must provide a means by which asynchronous events, such as interrupts, are handled in a timely manner without jeopardizing the system. Towards this end, we describe the features of C omposite that allow user-defined scheduling policies to be composed for the purposes of combined interrupt and task management. A significant challenge arises from the need to synchronize access to shared data structures (e.g., scheduling queues), without allowing untrusted code to disable interrupts. Additionally, efficient upcall mechanisms are needed to deliver asynchronous event notifications in accordance with policy-specific priorities, without undue recourse to schedulers. We show how these issues are addressed in C omposite , by comparing several hierarchies of scheduling polices, to manage both tasks and the interrupts on which they depend. Studies show how it is possible to implement guaranteed differentiated services as part of the handling of I/O requests from a network device while diminishing livelock. Microbenchmarks indicate that the costs of implementing and invoking user-level schedulers in C omposite are on par with, or less than, those in other systems, with thread switches more than twice as fast as in Linux.
Gabriel Parmer, Richard West
ACM Trans. Embed. Comput. Syst.1
2012 Shared hardware data structures for hard real-time systems
abstract
Hardware support can reduce the time spent operating on data structures by exploiting circuit-level parallelism. Such hardware data structures (HWDSs) can reduce the latency and jitter of data structure operations, which can benefit real-time systems by reducing worst-case execution times (WCETs). For example, a hardware priority queue (HWPQ) can enqueue and dequeue prioritized items in constant time with low variance; the best software implementations are in logarithmic-time asymptotic complexity for at least one of the enqueue or dequeue operations. The main problems with HWDSs are the limited size of hardware and the complexity of sharing it. In this paper we show that software support can help circumvent the size and sharing limitations of hardware so that applications can benefit from a HWDS. We evaluate our work by showing how the choice of software or hardware affects schedulability of task sets that use multiple priority queues of varying sizes. We model task behavior on two applications that are important in real-time and embedded domains: the grey-weighted distance transform for topology mapping and Dijkstra's algorithm for GPS navigation. Our results indicate that HWDSs can reduce the WCET of applications even when a HWDS is shared by multiple data structures or when data structure sizes exceed HWDS size constraints.
Gedare Bloom, Gabriel Parmer, Bhagirath Narahari, Rahul Simha
EMSOFT2
2012 Increasing Memory Utilization with Transient Memory Scheduling
abstract
In addition to predictability, both reliability and security are increasingly important for embedded systems. To limit the scope of errant behavior in open and mixed criticality systems, a common approach is to raise isolation barriers between software components. However, this decentralizes memory management across all system components. Memory is often cached and quickly accessible in each application. This paper introduces the TMEM system for increasing memory utilization while optimizing for application end-to-end constraints such as meeting deadlines. In addition to the traditional spatial multiplexing of memory, TMEM introduces the predictable temporal multiplexing of memory within caches in a system component, and memory scheduling to continually reallocate memory between components to best benefit the system. We find that TMEM is able to maintain the efficiency of caches, while also lowering both task tardiness and system memory requirements.
Jiguo Song, Gabriel Parmer, Andrew Sweeney, Guru Venkataramani
RTSS3
2012 Mutable Protection Domains: Adapting System Fault Isolation for Reliability and Efficiency
abstract
As software systems are becoming increasingly complex, the likelihood of faults and unexpected behaviors will naturally increase. Today, mobile devices to large-scale servers feature many millions of lines of code. Compile-time checks and offline verification methods are unlikely to capture all system states and control flow interactions of a running system. For this reason, many researchers have developed methods to contain faults at runtime by using software and hardware-based techniques to define protection domains. However, these approaches tend to impose isolation boundaries on software components that are static, and thus remain intact while the system is running. An unfortunate consequence of statically structured protection domains is that they may impose undue overhead on the communication between separate components. This paper proposes a new runtime technique that trades communication cost for fault isolation. We describe Mutable Protection Domains (MPDs) in the context of our Composite operating system. MPD dynamically adapts hardware isolation between interacting software components, depending on observed communication “hot-paths,” with the purpose of maximizing fault isolation where possible. In this sense, MPD naturally tends toward a system of maximal component isolation, while collapsing protection domains where costs are prohibitive. By increasing isolation for low-cost interacting components, MPD limits the scope of impact of future unexpected faults. We demonstrate the utility of MPD using a webserver, and identify different hot-paths for different workloads that dictate adaptations to system structure. Experiments show up to 40 percent improvement in throughput compared to a statically organized system, while maintaining high-fault isolation.
Gabriel Parmer, Richard West
IEEE Trans. Software Eng.1
2011 HiRes: A System for Predictable Hierarchical Resource Management
abstract
This paper presents HiRes, a system structured around predictable, hierarchical resource management (HRM). Applications and different subsystems use customized resource managers that control the allocation and usage of memory, CPU, and I/O. This increased resource management flexibility enables subsystems with different timing constraints to specialize resource management around meeting these requirements. In HiRes, subsystems delegate the management of resources to other subsystems, thus creating the resource management hierarchy. In delegating the control of resources, the subsystem focuses on providing isolation between competing subsystems. To make HRM both predictable and efficient, HiRes ensures that regardless of a subsystem's depth in the hierarchy, the overheads of resource usage and control remain constant. In doing so, HiRes encourages HRM as a fundamental system design technique. Results show that HiRes has competitive performance with existing systems, and that HRM naturally provides both strong isolation guarantees, and flexible and efficient subsystem control over resources.
Gabriel Parmer, Richard West
IEEE Real-Time and Embedded Technology and Applications Symposium1
2011 Execution Stack Management for Hard Real-Time Computation in a Component-Based OS
abstract
In addition to predictability, both reliability and security constraints are increasingly important. Mixed criticality, and open real-time systems execute software of different certification and trust levels. To limit the scope of errant behavior in these systems, a common approach is to raise isolation barriers between software components. However, a thread that executes through multiple components computes on execution stacks spread across each component. As these stacks require backing memory, each component has a finite amount of execution stacks. In this paper, we treat these stacks as shared resources, and investigate the implementation of traditional resource sharing protocols in a real component-based system. We implement multi-resource versions of the Priority Inheritance Protocol (PIP) and Priority Ceiling Protocol (PCP) for these shared stacks and find -- surprisingly -- that neither provide better schedulability characteristics than the other for all system parameterizations. Additionally, we identify the relationship between allocating additional stacks to components, and system schedulability. Given this, we describe and evaluate algorithms to ensure system schedulability while seeking to minimize the amount of memory consumed for stacks.
Jiguo Song, Gabriel Parmer
RTSS3
2011 Application-specific service technologies for commodity operating systems in real-time environments
abstract
In order to eliminate the costs of proprietary systems and special purpose hardware, many real-time and embedded computing platforms are being built on commodity operating systems and generic hardware. Unfortunately, many such systems are ill-suited to the low-latency and predictable timing requirements of real-time applications. This article, therefore, focuses on application-specific service technologies for low-cost commodity operating systems and hardware, so that real-time service guarantees can be met. We describe contrasting methods to deploy first-class services on commodity systems that are dispatched with low latency and execute asynchronously according to bounds on CPU, memory, and I/O device usage. Specifically, we present a “user-level sandboxing” (ULS) mechanism that relies on hardware protection to isolate application-specific services from the core kernel. This approach is compared with a hybrid language and runtime protection scheme, called SafeX , that allows untrusted services to be dynamically linked and loaded into a base kernel. SafeX and ULS have been implemented on commodity Linux systems. Experimental results have shown—that both approaches are capable of reducing service violations (and, hence, better qualities of service) for real-time tasks, compared to traditional user-level methods of service deployment in process-private address spaces. ULS imposes minimal additional overheads on service dispatch latency compared to SafeX, with the advantage that it does not require application-specific services to execute in the trusted kernel domain. As evidence of the potential capabilities of ULS, we show how a user-level networking stack can be implemented to avoid data copying via the kernel and allow packet processing without explicit process scheduling. This improves throughput and reduces jitter.
Richard West, Gabriel Parmer
ACM Trans. Embed. Comput. Syst.2
2008 Predictable Interrupt Management and Scheduling in the Composite Component-Based System
abstract
This paper presents the design of user-level scheduling hierarchies in the composite component-based system. The motivation for this is centered around the design of a system that is both dependable and predictable, and which is configurable to the needs of specific applications. Untrusted application developers can safely develop services and policies, that are isolated in protection domains outside the kernel. To ensure predictability, composite needs to enforce timing control over user-space services. Moreover, it must provide a means by which asynchronous events, such as interrupts, are handled in a timely manner without jeopardizing the system. Towards this end, we describe the features of composite that allow user-defined scheduling policies to be composed for the purposes of combined interrupt and task management. A significant challenge arises from the need to synchronize access to shared data structures (e.g., scheduling queues), without allowing untrusted code to disable interrupts or use atomic instructions that lock the memory bus. Additionally, efficient upcall mechanisms are needed to deliver asynchronous event notifications in accordance with policy-specific priorities, without undue recourse to schedulers. We show how these issues are addressed in Composite, by comparing several hierarchies of scheduling polices, to manage both tasks and the interrupts on which they depend. Studies show how it is possible to implement guaranteed differentiated services as part of the handling of I/O requests from a network device while avoiding livelock. Microbenchmarks indicate that the costs of implementing and invoking user-level schedulers in composite are on par with, or less than, those in other systems, with thread switches more than twice as fast as in Linux.
Gabriel Parmer, Richard West
RTSS1
2007 Hijack: Taking Control of COTS Systems for Real-Time User-Level Services
abstract
This paper focuses on a technique to empower commercial-off-the-shelf (COTS) systems with an execution environment, and corresponding services, to support real-time and embedded applications. By leveraging COTS systems, we are able to reduce the potentially expensive maintenance and development costs of proprietary solutions. We describe a system called "Hijack" that enables user-level services to take control of features such as CPU scheduling, interrupt handling and synchronization. In contrast to other approaches that support real-time tasks within the kernel of commodity systems such as Linux, Hijack provides the basis for predictable thread execution at user-level. No changes to the kernel source code are required to support this approach. Instead, Hijack works by using a combination of kernel module support and an interposed execution environment between traditional process address spaces and the kernel. This technique enables system calls and hardware interrupts to be intercepted with bounded latencies via the kernel module, that passes control to a user-level real-time executive. From within the executive, system-wide services and policies can be deployed to over-ride certain features of the underlying kernel, while still leveraging base kernel services where appropriate. Using this technique, we show how a vanilla Linux system can be hijacked to support predictable service execution using a series of user-defined policies. In particular, we show how to deliver and process asynchronous events with bounded latency, using interposition agents within a Hijack execution environment. Results show that for realtime streaming applications, Hijack is able to receive and process packets with significantly lower loss rates and jitter compared to using alternative application-level processes for the same task
Gabriel Parmer, Richard West
IEEE Real-Time and Embedded Technology and Applications Symposium1
2007 Mutable Protection Domains: Towards a Component-Based System for Dependable and Predictable Computing
abstract
The increasing complexity of software poses significant challenges for real-time and embedded systems beyond those based purely on timeliness. With embedded systems and applications running on everything from mobile phones, PDAs, to automobiles, aircraft and beyond, an emerging challenge is to ensure both the functional and timing correctness of complex software. We argue that static analysis of software is insufficient to verify the safety of all possible control flow interactions. Likewise, a static system structure upon which software can be isolated in separate protection domains, thereby defining immutable boundaries between system and application-level code, is too inflexible to the challenges faced by real-time applications with explicit timing requirements. This paper, therefore, investigates a concept called "mutable protection domains" that supports the notion of hardware-adaptable isolation boundaries between software components. In this way, a system can be dynamically reconfigured to maximize software fault isolation, increasing dependability, while guaranteeing various tasks are executed according to specific time constraints. Using a series of simulations on multidimensional, multiple-choice knapsack problems, we show how various heuristics compare in their ability to rapidly reorganize the fault isolation boundaries of a component- based system, to ensure resource constraints while simultaneously maximizing isolation benefit. Our ssh oneshot algorithm offers a promising approach to address system dynamics, including changing component invocation patterns, changing execution times, and mispredictions in isolation costs due to factors such as caching.
Gabriel Parmer, Richard West
RTSS1
2004 An efficient end-host architecture for cluster communication
abstract
Cluster computing environments built from commodity hardware have provided a cost-effective solution for many scientific and high-performance applications. Likewise, middleware techniques have provided the basis for large-scale applications to communicate and exchange data across the various end-hosts in a distributed system. Unfortunately, middleware services are typically encapsulated in user-level address spaces that suffer from scheduling delays and communication overheads induced by the host kernel. For various high performance distributed computing applications such overheads are unacceptable. This work therefore addresses the problem of providing an efficient end-host architecture to support application-specific communication services at user-level, without the need to explicitly schedule such services or copy data via the kernel. We briefly describe a sandboxing mechanism that allows applications to configure and deploy services at user-level that may execute in the context of any address space. Using Linux as the basis for our approach, we focus specifically on the implementation of a user-space network protocol stack that avoids copying data via the kernel when communicating with the network interface. Our approach enables services to efficiently process and forward data via proxies, or intermediate hosts, in the communication path of high performance data streams. Unlike other user-level networking implementations, our method makes no special hardware requirements. Results show that we achieve a substantial increase in throughput, and a reduction in jitter, over comparable user-space communication methods.
Gabriel Parmer, Richard West
CLUSTER2