Bryan C. Ward

dblp:66/5887 · DBLP profile ↗
← Back
32ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0001-7168-6693ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 7 since 2021Systems, architecture and hardware · 10 · 1 first-author · 5 since 2021Security and privacy · 2 · 1 first-author
YearPublicationVenuePosition
2026 DART: A Real-Time Address-Randomization Defense with Predictable Timing
abstract
Embedded and real-time systems are increasingly connected and deployed in safety and mission-critical environments, making them a persistent target for attacks capable of compromising industrial control systems and other embedded devices. At the same time, these devices often have strict real-time requirements that require predictable worst-case performance. However, many strong and widely deployed software-security defenses are designed and evaluated with respect to average-case performance, a more important metric in enterprise systems. The worst-case performance of such defenses is not well understood and indeed such defenses are less commonly deployed in embedded systems. In particular, one class of commonly deployed defenses in enterprise systems is code randomization, which protects a system by altering the layout of the virtual address space so that attackers cannot easily target specific parts of a vulnerable application, but randomization is often seen as fundamentally counter to real-time predictability. This paper presents DART, a real-time address randomization defense with page-level randomization. DART randomizes code in the virtual address space at page-level granularity under placement constraints that move cache behavior from a runtime OS-allocator property to a statically encoded binary property, allowing for timing analysis. An analysis of DART’s timing behavior on a real-time testbed demonstrates how the design makes layout-induced timing variance bounded and characterizable across the space of layouts produced, supporting predictable execution-time analysis. The resulting layout search space is then analyzed, and a closed-form expression for the randomization entropy induced by DART is derived. Evaluation results across TACLeBench binaries show increased combinatorial entropy with modest numbers of virtual memory pages per cache color, providing a suitable defense that outperforms traditional virtual-memory protections for attacks such as partial-pointer overwriting or more broadly control-flow hijacking.
Patrick Dobranowski, Owen Rice, Ryan Burrow, Nathan Burow, Bryan C. Ward
ECRTS5
2026 CacheFlow: Using Maximum Flow to Bound Cache-Based Preemption Delays
abstract
Cache-related preemption delay (CRPD) analysis bounds the additional execution time caused by cache evictions during preemptions. Tightly bounding CRPDs is challenging as there are many possible preemption patterns that can occur at runtime, and thus there has been continuous work over three decades to refine these bounds. This paper presents CacheFlow, a framework that formulates total CRPD as a maximum-flow problem. In the flow network, nodes and edge capacities can be constructed to model certain eviction patterns that can occur. Therefore, by (safely) removing nodes or edges, or reducing edge capacities, tighter CRPD bounds can be derived. This is demonstrated with different CacheFlow refinements, some of which include insights from prior analyses, as well as a refinement for simply periodic systems. An iterative max-flow formulation is also described to more efficiently integrate the max-flow solving in the context of standard fixed-priority response-time analysis. Experiments on synthetic task systems demonstrate significant schedulability improvements across a range of system configurations, while also having reasonable solving times.
Tiancheng He, Bryan C. Ward
ECRTS2
2026 MIRAGE: MILP-Based Block Grouping for Real-Time Signal Processing
Tiancheng He, Bryan C. Ward
RTAS2
2026 Balancing Security and Schedulability: WCET Evaluation and Security Optimization in CPS
Marion Sudvarg, Ching-Hsiang Chan, Ryan Burrow, Nathan Burow, Cailani Lemieux Mack, Sanjoy Baruah, Ning Zhang 0017, Bryan C. Ward
RTAS9
2026 Introduction to Special Issue on Security and Privacy in Safety-Critical Cyber-Physical Systems - Part 2
Ning Zhang 0017, Bryan C. Ward, Andrew Clark 0001, Ziming Zhao 0001, Aiping Xiong
ACM Trans. Cyber Phys. Syst.2
2026 Introduction to Special Issue on Security and Privacy in Safety-Critical Cyber-Physical Systems - Part 3
Ning Zhang 0017, Bryan C. Ward, Andrew Clark 0001, Ziming Zhao 0001, Aiping Xiong
ACM Trans. Cyber Phys. Syst.2
2025 Resilient Scheduling of Real-Time Cyber-Physical Systems Against Memory-Corruptions
Abdullah Al Arafat, Kurt M. Wilson, Sudharsan Vaidhun, Bryan C. Ward, Zhishan Guo
RTCSA4
2025 Introduction to the Special Issue on Security and Privacy in Safety-Critical Cyber-Physical Systems
abstract
No abstract available.
Ning Zhang 0017, Bryan C. Ward, Andrew Clark 0001, Ziming Zhao 0001, Aiping Xiong
ACM Trans. Cyber Phys. Syst.2
2024 InsectACIDE: Debugger-Based Holistic Asynchronous CFI for Embedded System
abstract
Real-time and embedded systems are predominantly written in C, a language that is notoriously not memory safe. This has led to widespread memory-corruption vulnerabilities in real-time embedded cyber-physical systems (CPS). This is concerning, as such devices are becoming increasingly networked with the Internet of Things (IoT) and other communication technologies (e.g., 5G), rendering them vulnerable to remote attacks. Attackers have demonstrated how memory-corruption vulnerabilities can be used to hijack program control flow to implement arbitrary attacker-controlled logic. One class of defenses that has been developed to prevent such attacks is called control-flow integrity (CFI), which applies checks at control-flow transitions to ensure the target is valid. Unfortunately, attackers have shown how to divert control flow to seemingly valid targets in an invalid and malicious sequence. This paper presents InsectACIDE, the first holistic CFI for embedded and real-time systems that does not require binary instrumentation and that is context sensitive, i.e., it checks that the sequence of control-flow transitions taken is valid, not just individual transitions, thereby detecting such attacks. InsectACIDE is implemented on an embedded Cortex-M processor using the TrustZone trusted execution environment, and holistic context-sensitive CFI is enforced for both applications and the kernel. InsectACIDE uses hardware debugging features on the Cortex-M processor and therefore does not require any kernel or application binary modification. Experimental results show that InsectACIDE incurs significantly less runtime overhead compared to the state-of-the-art holistic CFI solution. Real-time schedulability analysis is presented, along with a schedulability evaluation, to demonstrate the tradeoff between stronger protection and real-time schedulability.
Cailani Lemieux Mack, Xi Tan 0002, Ning Zhang 0017, Ziming Zhao 0001, Sanjoy Baruah, Bryan C. Ward
RTAS7
2024 Job-Level Batching for Software-Defined Radio on Multi-Core
abstract
Conventional wireless communication is built upon hardware-based signal processing. This enables high performance, but is inflexible as the signal-processing algorithms are “baked in” to the hardware. Software-defined radio (SDR) is an emerging solution in which more of the signal-processing logic is implemented in software instead of hardware. This allows for adaptability to spectrum conditions (e.g., jamming or congestion), changes to protocols, and software updates that improve signal-processing logic – features that are beneficial in many consumer and military applications. However, the high sampling rate $(\text{kHz}$ to MHz or faster) of many SDR applications poses significant challenges for real-time scheduling of such workloads. To manage high sampling rates on general-purpose processors, which process samples sequentially instead of in parallel, as can be done using hardware acceleration, samples must be buffered, or “batched” together, to minimize overheads and maximize locality. To address this characteristic of high-frequency signal processing, this paper presents an extension of traditional real-time scheduling models called the marginal cost model, which reflects the fact that when batching many samples, the marginal cost of processing additional samples is often much less than the cost of processing the first sample. Empirical evaluations are presented from the open source GNU Radio SDR framework to validate the marginal cost model. Experiments are then presented that demonstrate the trade-offs between batching and worst-case latency for synthetic SDR workloads. Finally, a case study is presented to demonstrate the utility of the presented model and batching techniques in real-world signal-processing applications.
Abigail Eisenklam, Will Hedgecock, Bryan C. Ward
RTSS3
2024 Partial Context-Sensitive Pointer Integrity for Real-time Embedded Systems
abstract
Safety- and mission-critical cyber-physical systems (CPSs) require temporal correctness to ensure safe physical behavior. This manifests as strict timing requirements, which cannot be missed at runtime. Counter-intuitively, this implies that real-time tasks can be delayed so long as they remain guaranteed to meet their deadlines. This paper explores how extra time in a schedule can be analytically recapitalized for the purpose of applying stronger security protection within individual tasks at compile time. This is achieved through the development of a partial context-sensitive pointer-integrity framework (ParCSPI). In this framework, more fine-grained policies can be enforced, with greater runtime overheads, where so doing does not violate real-time constraints. A whole-system optimization framework based upon a mixed-integer linear programming approach to fixed-priority response-time analysis is used to identify precisely which contexts can be checked within the available system-wide time while maximizing system-wide security. ParCSPI leverages Arm pointer authentication (PA) to encode context-based equivalence classes into the modifiers of the pointer signature and is implemented using a customized program analyzer and LLVM compiler passes. An evaluation of ParCSPI is presented that includes per-task and system-wide overhead and security tradeoffs, as well as a demonstration on a real-world CPS. Empirical results are presented showing that ParCSPI achieves up to 62% pointer-integrity protection with only 10% worst-case execution time (WCET) overhead, and can find optimal security trade-offs in complex real-time task sets as well as approximate them in reasonable time.
Cailani Lemieux Mack, Thidapat Chantem, Sanjoy Baruah, Ning Zhang 0017, Bryan C. Ward
RTSS6
2023 Who's Afraid of Butterflies? A Close Examination of the Butterfly Attack
abstract
The Butterfly Attack, introduced in an RTSS 2019 paper, was billed as a new kind of timing attack against control loops in cyber-physical systems. We conduct a close inspection of the Butterfly Attack in order to identify the root vulnerability that it exploits, and show that an appropriate application of real-time scheduling theory provides an effective countermeasure. We propose improved defenses against this and similar attacks by drawing upon techniques from real-time scheduling theory, control theory, and systems implementation, that are both provably secure and are able to make efficient use of computing resources.
Sanjoy Baruah, Pontus Ekberg, Mehdi Hosseinzadeh 0002, Ao Li 0006, Bryan C. Ward, Ning Zhang 0017
RTSS5
2022 The Thundering Herd: Amplifying Kernel Interference to Attack Response Times
abstract
Embedded and real-time systems are increasingly attached to networks. This enables broader coordination beyond the physical system, but also opens the system to attacks. The increasingly complex workloads of these systems include software of varying assurance levels, including that which might be susceptible to compromise by remote attackers. To limit the impact of compromise, μ-kernels focus on maintaining strong memory protection domains between different bodies of software, including system services. They enable limited coordination between processes through Inter-Process Communication (IPC). Real-time systems also require strong temporal guarantees for tasks, and thus need temporal isolation to limit the impact of malicious software. This is challenging as multiple client threads that use IPC to request service from a shared server will impact each other’s response times.To constrain the temporal interference between threads, modern μ-kernels often build priority and budget awareness into the system. Unfortunately, this paper demonstrates that this is more challenging than previously thought. Adding priority awareness to IPC processing can lead to significant interference due to the kernel’s prioritization logic. Adding budget awareness similarly creates opportunities for interference due to the budget tracking and management operations. In both situations, a Thundering Herd of malicious threads can significantly delay the activation of mission-critical tasks. The Thundering Herd effects are evaluated on seL4 and results demonstrate that high-priority threads can be delayed by over 100,000 cycles per malicious thread. This paper reveals a challenging dilemma: the temporal protections μ-kernels add can, themselves, provide means of threatening temporal isolation. Finally, to defend the system, we identify and empirically evaluate possible mitigations, and propose an admission-control test based upon an interference-aware analysis.
Samuel Mergendahl, Samuel Jero, Bryan C. Ward, Juliana Furgala, Gabriel Parmer, Richard Skowyra
RTAS3
2021 Light Reading: Optimizing Reader/Writer Locking for Read-Dominant Real-Time Workloads
abstract
This paper is directed at reader/writer locking for read-dominant real-time workloads. It is shown that state-of-the-art real-time reader/writer locking protocols are subject to performance limitations when reads dominate, and that existing schedulability analysis fails to leverage the sparsity of writes in this case. A new reader/writer locking-protocol implementation and new inflation-free schedulability analysis are proposed to address these problems. Overhead evaluations of the new implementation show a decrease in overheads of up to 70% over previous implementations, leading to throughput for read operations increasing by up to 450%. Schedulability experiments are presented that show that the analysis results in schedulability improvements of up to 156.8% compared to the existing state-of-the-art approach.
Catherine E. Nemitz, Shai Caspin, James H. Anderson, Bryan C. Ward
ECRTS4
2021 Practical Principle of Least Privilege for Secure Embedded Systems
abstract
Many embedded systems have evolved from simple bare-metal control systems to highly complex network-connected systems. These systems increasingly demand rich and feature-full operating-systems (OS) functionalities. Furthermore, the network connectedness offers attack vectors that require stronger security designs. To that end, this paper defines a prototypical RTOS API called Patina that provides services common in featurerich OSes (e.g., Linux) but absent in more trustworthy μ -kernel based systems. Examples of such services include communication channels, timers, event management, and synchronization. Two Patina implementations are presented, one on Composite and the other on seL4, each of which is designed based on the Principle of Least Privilege (PoLP) to increase system security. This paper describes how each of these μ -kernels affect the PoLP based design, as well as discusses security and performance tradeoffs in the two implementations. Results of comprehensive evaluations demonstrate that the performance of the PoLP based implementation of Patina offers comparable or superior performance to Linux, while offering heightened isolation.
Samuel Jero, Juliana Furgala, Runyu Pan, Phani Kishore Gadepalli, Alexandra Clifford, Bite Ye, Roger I. Khazan, Bryan C. Ward, Gabriel Parmer, Richard Skowyra
RTAS8
2021 TORTIS: Retry-Free Software Transactional Memory for Real-Time Systems
abstract
Software transactional memory (STM) is a synchronization paradigm originally proposed for throughput-oriented computing to facilitate producing performant concurrent code that is free of synchronization bugs. With STM, programmers merely annotate code sections requiring synchronization; the underlying STM framework automatically resolves how synchronization is done. Today, the programming issues that motivated STM are becoming a concern in embedded computing, where ever more sophisticated systems are being produced that require highly parallel implementations. These implementations are often produced by engineers and control experts who may not be well versed in concurrency-related issues. In this context, a real-time STM framework would be useful in ensuring that the synchronization aspects of a system pass real-time certification. However, all prior STM approaches fundamentally rely on retries to resolve conflicts, and such retries can yield high worst-case synchronization costs compared to lock-based approaches. This paper presents a new STM class called Retry-Free Real-Time STM (R2STM), which is designed for worst-case real-time performance. The benefit of a retry-free approach for use in a real-time system is demonstrated by a schedulability study, in which it improved overall schedulability across all considered task systems by an average of 95.3% over a retry-based approach. This paper also presents TORTIS, the first R2STM implementation for real-time systems. Throughput-oriented benchmarks are presented to highlight the tradeoffs between throughput and schedulability for TORTIS.
Claire Nord, Shai Caspin, Catherine E. Nemitz, Howard E. Shrobe, Hamed Okhravi, James H. Anderson, Nathan Burow, Bryan C. Ward
RTSS8
2019 Controller-Oblivious Dynamic Access Control in Software-Defined Networks
abstract
Conventional network access control approaches are static (e.g., user roles in Active Directory), coarse-grained (e.g., 802.1x), or both (e.g., VLANs). Such systems are unable to meaningfully stop or hinder motivated attackers seeking to spread throughout an enterprise network. To address this threat, we present Dynamic Flow Isolation (DFI), a novel architecture for supporting dynamic, fine-grained access control policies enforced in a Software-Defined Network (SDN). These policies can emit and revoke specific access control rules automatically in response to network events like users logging off, letting the network adaptively reduce unnecessary reachability that could be potentially leveraged by attackers. DFI is oblivious to the SDN controller implementation and processes new packets prior to the controller, making DFI's access control resilient to a malicious or faulty controller or its applications. We implemented DFI for OpenFlow networks and demonstrated it on an enterprise SDN testbed with around 100 end hosts and servers. Finally, we evaluated the performance of DFI and how it enables a novel policy, which is otherwise difficult to enforce, that protects against a surrogate of the recent NotPetya malware in an infection scenario. We found that the threat was most limited in its ability to spread using our policy, which automatically restricted network flows over the course of the attack, compared to no access control or a static role-based policy.
Steven R. Gomez, Samuel Jero, Richard Skowyra, Patrick Sullivan, David Bigelow, Zachary Ellenbogen, Bryan C. Ward, Hamed Okhravi, James Landry
DSN8
2019 Control-Flow Integrity for Real-Time Embedded Systems
abstract
Attacks on real-time embedded systems can endanger lives and critical infrastructure. Despite this, techniques for securing embedded systems software have not been widely studied. Many existing security techniques for general-purpose computers rely on assumptions that do not hold in the embedded case. This paper focuses on one such technique, control-flow integrity (CFI), that has been vetted as an effective countermeasure against control-flow hijacking attacks on general-purpose computing systems. Without the process isolation and fine-grained memory protections provided by a general-purpose computer with a rich operating system, CFI cannot provide any security guarantees. This work proposes RECFISH, a system for providing CFI guarantees on ARM Cortex-R devices running minimal real-time operating systems. We provide techniques for protecting runtime structures, isolating processes, and instrumenting compiled ARM binaries with CFI protection. We empirically evaluate RECFISH and its performance implications for real-time systems. Our results suggest RECFISH can be directly applied to binaries without compromising real-time performance; in a test of over six million realistic task systems running FreeRTOS, 85% were still schedulable after adding RECFISH.
Robert J. Walls, Nicholas F. Brown, Thomas Le Baron, Craig A. Shue, Hamed Okhravi, Bryan C. Ward
ECRTS6
2019 The Leakage-Resilience Dilemma
Bryan C. Ward, Richard Skowyra, Chad Spensky, Hamed Okhravi
ESORICS (1)1
2017 Sustainability in Mixed-Criticality Scheduling
abstract
Sustainability is a formalization of the requirement for scheduling algorithms and schedulability tests that a system deemed to be correctly schedulable should remain so if its run-time behavior is better than anticipated. The notion of sustainability is extended to mixed-criticality systems, and sustainability properties are determined for a variety of widely-studied uniprocessor and multi-processor mixed-criticality scheduling algorithms.
Zhishan Guo, Sai Sruti, Bryan C. Ward, Sanjoy Baruah
RTSS3
2017 Attacking the one-out-of-m multicore problem by combining hardware management with mixed-criticality provisioning
Namhoon Kim, Bryan C. Ward, Micaiah Chisholm, James H. Anderson, F. Donelson Smith
Real Time Syst.2
2016 Attacking the One-Out-Of-m Multicore Problem by Combining Hardware Management with Mixed-Criticality Provisioning
abstract
The multicore revolution is having limited impact in safety-critical application domains. A key reason is the "one-out-of-m" problem: when validating real-time constraints on an m-core platform, excessive analysis pessimism can effectively negate the processing capacity of the additional m-1 cores so that only "one core's worth" of capacity is available. Two approaches have been investigated previously to address this problem: mixed-criticality allocation techniques, which provision less-critical software components less pessimistically, and hardware-management techniques, which make the underlying platform itself more predictable. A better way forward may be to combine both approaches, but to show this, fundamentally new criticality-cognizant hardware-management tradeoffs must be explored. Such tradeoffs are investigated herein in the context of a large-scale, overhead-aware schedulability study. This study was guided by extensive trace data obtained by executing benchmark tasks on a new variant of the MC^2 framework that supports configurable criticality-based hardware management. This study shows that the two approaches mentioned above can be much more effective when applied together instead of alone.
Namhoon Kim, Bryan C. Ward, Micaiah Chisholm, Cheng-Yang Fu, James H. Anderson, F. Donelson Smith
RTAS2
2016 Reconciling the Tension Between Hardware Isolation and Data Sharing in Mixed-Criticality, Multicore Systems
abstract
Recent work involving a mixed-criticality framework called MC2 has shown that, by combining hardware-management techniques and criticality-aware task provisioning, capacity loss can be significantly reduced when supporting real-time workloads on multicore platforms. However, as in most other prior research on multicore hardware management, tasks were assumed in that work to not share data. Data sharing is problematic in the context of hardware management because it can violate the isolation properties hardware-management techniques seek to ensure. Clearly, for research on such techniques to have any practical impact, data sharing must be permitted. Towards this goal, this paper presents a new version of MC2 that permits tasks to share data within and across criticality levels through shared memory. Several techniques are presented for mitigating capacity loss due to data sharing. The effectiveness of these techniques is demonstrated by means of a large-scale, overhead-aware schedulability study driven by micro-benchmark data.
Micaiah Chisholm, Namhoon Kim, Bryan C. Ward, Nathan Otterness, James H. Anderson, F. Donelson Smith
RTSS3
2015 Cache Sharing and Isolation Tradeoffs in Multicore Mixed-Criticality Systems
abstract
In mixed-critical applications, tension exists between sharing and isolation with respect to hardware resources: while strong isolation might be required for highly critical tasks, somewhat permissive sharing might be reasonable for less critical tasks to improve throughput or average-case performance. In this paper, this tension is examined as it pertains to shared last-level caches (LLCs) on multicore platforms. In particular, criticality-aware optimization techniques based on linear programming are presented for allocating LLC areas in the context of the previously proposed MC2 (mixed-criticality on multicore) framework. Experiments are also presented that show that these techniques can result in significant schedulability improvements.
Micaiah Chisholm, Bryan C. Ward, Namhoon Kim, James H. Anderson
RTSS2
2015 Relaxing Resource-Sharing Constraints for Improved Hardware Management and Schedulability
abstract
Modern computer architectures, particularly multicore systems, include shared hardware resources such as caches and interconnects that introduce timing-interference channels. Unmanaged access to such resources can adversely affect the execution time of other tasks, and lead to unpredictable execution times and associated analysis pessimism that can entirely negate the benefits of a multicore processor. To mitigate such effects, accesses to shared hardware resources should be managed, for example, by a real-time locking protocol. However, accesses to some hardware resources can be managed with more relaxed sharing constraints than mutual exclusion while still mitigating timing-interference channels. This paper presents two new classes of sharing constraints, preemptive mutual exclusion, and half-protected sharing, which are motivated by the sharing constraints of buses and caches, respectively. Synchronization algorithms are presented for both sharing constraints, where applicable, on both uni-and multi-processor systems. A fundamentally new analysis technique called idleness analysis is presented to account for the effects of blocking in globally scheduled multiprocessor systems. Experimental results suggest that these relaxed synchronization requirements and improved analysis techniques can improve schedulability by up to 250%. Furthermore, idleness analysis can be applied to existing locking protocols to improve schedulability in many cases.
Bryan C. Ward
RTSS1
2014 Multi-resource Real-Time Reader/Writer Locks for Multiprocessors
abstract
A fine-grained locking protocol permits multiple locks to be held simultaneously by the same task. In the case of real-time multiprocessor systems, prior work on such protocols has considered only mutex constraints. This unacceptably limits concurrency in systems in which some resource accesses are read-only. To remedy this situation, a variant of a recently proposed fine-grained protocol called the real-time nested locking protocol (RNLP) is presented that enables concurrent reads. This variant is shown to have worst-case blocking no worse (and often better) than existing coarse-grained real-time reader/writer locking protocols, while allowing for additional parallelism. Experimental evaluations of the proposed protocol are presented that consider both schedulability (i.e., the ability to validate timing constraints) and implementation-related overheads. These evaluations demonstrate that the RNLP (both the mutex and the proposed reader/writer variant) provides improved schedulability over existing coarse-grained locking protocols, and is practically implementable.
Bryan C. Ward, James H. Anderson
IPDPS1
2014 Fair lateness scheduling: reducing maximum lateness in G-EDF-like scheduling
Jeremy P. Erickson, James H. Anderson, Bryan C. Ward
Real Time Syst.3
2013 Outstanding Paper Award: Making Shared Caches More Predictable on Multicore Platforms
abstract
In safety-critical cyber-physical systems, the usage of multicore platforms has been hampered by problems due to interactions across cores through shared hardware. The inability to precisely characterize such interactions can lead to worst-case execution time pessimism that is so great, the extra processing capacity of additional cores is entirely negated. In this paper, several techniques are proposed and analyzed for dealing with such interactions in the context of shared caches. These techniques are applied in a mixed-criticality scheduling framework motivated by the needs of next-generation unmanned air vehicles.
Bryan C. Ward, Jonathan L. Herman, Christopher J. Kenna, James H. Anderson
ECRTS1
2013 GPUSync: A Framework for Real-Time GPU Management
abstract
This paper describes GPUSync, which is a framework for managing graphics processing units (GPUs) in multi-GPU multicore real-time systems. GPUSync was designed with flexibility, predictability, and parallelism in mind. Specifically, it can be applied under either static-or dynamic priority CPU scheduling, can allocate CPUs/GPUs on a partitioned, clustered, or global basis, provides flexible mechanisms for allocating GPUs to tasks, enables task state to be migrated among different GPUs, with the potential of breaking such state into smaller "chunks", provides migration cost predictors that determine when migrations can be effective, enables a single GPU's different engines to be accessed in parallel, properly supports GPU-related interrupt and worker threads according to the sporadic task model, even when GPU drivers are closed-source, and provides budget policing to the extent possible, given that GPU access is non-preemptive. No prior real-time GPU management framework provides a comparable range of features.
Glenn A. Elliott, Bryan C. Ward, James H. Anderson
RTSS2
2012 Supporting Nested Locking in Multiprocessor Real-Time Systems
abstract
This paper presents the first real-time multiprocessor locking protocol that supports fine-grained nested resource requests. This locking protocol relies on a novel technique for ordering the satisfaction of resource requests to ensure a bounded duration of priority inversions for nested requests. This technique can be applied on partitioned, clustered, and globally scheduled systems in which waiting is realized by either spinning or suspending. Furthermore, this technique can be used to construct fine-grained nested locking protocols that are efficient under spin-based, suspension-oblivious or suspension-aware analysis of priority inversions. Locking protocols built upon this technique perform no worse than coarse-grained locking mechanisms, while allowing for increased parallelism in the average case (and, depending upon the task set, better worst-case performance).
Bryan C. Ward, James H. Anderson
ECRTS1
2012 Replica-Request Priority Donation: A Real-Time Progress Mechanism for Global Locking Protocols
abstract
Real-time locking protocols employ progress mechanism(s) to ensure that resource-holding jobs are scheduled. These mechanisms are required to bound the duration of priority-inversion blocking (pi-blocking) for jobs sharing resources. Examples of such progress mechanisms include priority inheritance and priority donation. Unfortunately, some progress mechanisms can cause any job, including those that never request shared resources, to be blocked upon job release. This paper presents a variant of priority donation for globally-scheduled systems that only causes blocking for jobs waiting for shared resources. Additionally, this variant of priority donation is employed to construct a new suspension-based locking protocol called the replica-request donation global locking protocol (R2DGLP), which is asymptotically optimal for both mutex and k-exclusion (i.e., multi-unit) resources. This work is motivated by multicore systems where tasks may share I/O devices (e.g., GPUs) where critical sections can be long. In such applications, progress mechanisms that cause jobs that do not access I/O devices to be blocked to ensure progress can be detrimental from a schedulability perspective.
Bryan C. Ward, Glenn A. Elliott, James H. Anderson
RTCSA1
2008 Enhancing the Credibility of Wireless Network Simulations with Experiment Automation
abstract
The last few years have witnessed a growing consensus around the notion that many papers discussing wireless network simulation are plagued by issues that weaken their scientific value. A number of articles have shown evidence of this crisis of credibility and identified many of its causes. In this paper, we show that the methodology flaws in wireless network simulation can be avoided with the use of a framework for experiment automation. We describe the rationale that drove us to develop tools for component-based simulators intending to guide the experimental process from first to last stages. We conclude that a framework that imposes the right constraints on the experimenter can lead to more credible simulation studies. The framework we present helps the construction of consistent models, the definition of model parameters, the design and the execution of experiments, the analysis of output data, and the preparation of data for the dissemination of results that allow experiments to be reproduced.
L. Felipe Perrone, Christopher J. Kenna, Bryan C. Ward
WiMob3