VLDB 2026 Research / reviewers in the wild / expert
Neil C. Audsley
dblp:06/898
· DBLP profile ↗
69ranked-venue papers
10as first author
14since 2021 · last 2024
0000-0003-3739-6590ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 2 first-authorTheory of computation · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Hopscotch: A Hardware-Software Co-Design for Efficient Cache Resizing on Multi-Core SoCsabstractFollowing the trend of increasing autonomy in real-time systems, multi-core System-on-Chips (SoCs) have enabled devices to better handle the large streams of data and intensive computation required by such autonomous systems. In modern multi-core SoCs, each L1 cache is designed to be tied to an individual processor, and a processor can only access its own L1 cache. This design method ensures the system's average throughput, but also limits the possibility of parallelism, significantly reducing the system's real-time schedulability. To overcome this problem, we present a new system framework for highly-parallel multi-core systems,Hopscotch.Hopscotchintroduces re-sizable L1 cache which is shared between processors in the same computing cluster. At execution,Hopscotchdynamically allocates L1 cache capacity to the tasks executed by the processors, unblocking the available parallelism in the system. Based on the new hardware architecture, we also present a new theoretical model and schedulability analysis providing cache size selection methods and corresponding timing guarantees for the system. As demonstrated in the evaluations,Hopscotcheffectively improves system-level schedulability with negligible extra overhead. Zhe Jiang 0004, Kecheng Yang 0001, Nathan Fisher, Nan Guan, Neil C. Audsley, Zheng Dong 0002 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2023 | A High-Resilience Imprecise Computing Architecture for Mixed-Criticality SystemsabstractConventional mixed-criticality systems (MCS)s are designed to terminate the execution of less critical tasks in exceptional situations so that the timing properties of more critical tasks can be preserved. Such a strategy can be controversial and has proven difficult to implement in practice, as it can lead to hazards and reduced functionality due to the absence of the discarded tasks. To mitigate this issue, the imprecise mixed-critically system model (IMCS) has been proposed. In such a model, instead of completely dropping less-critical tasks, these tasks are executed as much as possible through the use of decreased computation precision. Although IMCS could effectively improve the survivability of the less-critical tasks, it also introduces three key drawbacks - run-time computation errors, real-time performance degradation, and lack of flexibility. In this paper, we present a novel IMCS framework, which can (i) mitigate the computation errors caused by imprecise computation; (ii) achieve real-time performance near to that of a conventional MCS; (iii) enhance system-level throughput; and (iv) provide flexibility for run-time configuration. We describe the design details ofHIART-MCS, and then present the corresponding theoretical analysis and optimisation method for its run-time configuration. Finally,HIART-MCS is evaluated against other MCS frameworks using a variety of experimental metrics. Zhe Jiang 0004, Xiaotian Dai 0001, Alan Burns 0001, Neil C. Audsley, Zonghua Gu 0001, Ian Gray |
IEEE Trans. Computers | 4 |
| 2023 | AXI-IC$^{\mathrm{ RT}}$ RT : Towards a Real-Time AXI-Interconnect for Highly Integrated SoCsabstractIn modern real-time heterogeneous System-on-Chips (SoCs), ensuring the predictability of interconnects is becoming increasingly important. Most of the existing interconnects are mainly designed to achieve high throughput, with their micro-architectures usually based on FIFO queues. The FIFO-based design prevents transaction prioritization based on importance and leads to occurrences of physical priority inversion. Such problems lead to difficulties in ensuring transaction predictability, especially when the system scales to a large number of elements. In this paper, we introduce AXI-Interconnect^{rt} (AXI-IC^{rt}, for short) -- a real-time AXI interconnect for heterogeneous SoCs, which redefines the micro-architecture of interconnects by enabling random accesses of buffered transactions and organizing transactions through compositional scheduling. This hardware-software co-design approach provides predictable and scalable real-time performance for highly integrated SoCs. Zhe Jiang 0004, Kecheng Yang 0001, Nathan Fisher, Ian Gray, Neil C. Audsley, Zheng Dong 0002 |
IEEE Trans. Computers | 5 |
| 2023 | Towards Hard Real-Time and Energy-Efficient Virtualization for Many-Core Embedded SystemsabstractIn safety-critical computing systems, the I/O virtualization must simultaneously satisfy different requirements, including time-predictability, performance, and energy-efficiency. However, these requirements are challenging to achieve due to complex I/O access path and resource management at the system level, lack of support from preemptive scheduling at I/O hardware level, and missing an effective energy management method. In this paper, we propose a new framework, I/O-GUARD, which reconstructs the system architecture of I/O virtualization, bringing a dedicated hardware hypervisor to handle resource management throughout the system. The hypervisor improves system real-time performance by enabling preemptive scheduling in I/O virtualization with both analytical and experimental real-time guarantees. Furthermore, we also present a dedicated energy management unit to adjustI/O-GUARD's dynamic energy using frequency scaling. Associated with that, a frequency identification algorithm is proposed to find the appropriate executing frequency at run-time. As shown in experiments,I/O-GUARDsimultaneously improves the predictability, performance and energy-efficiency compared to the state-of-the-art I/O virtualization. Zhe Jiang 0004, Kecheng Yang 0001, Yunfeng Ma, Nathan Fisher, Neil C. Audsley, Zheng Dong 0002 |
IEEE Trans. Computers | 5 |
| 2022 | BlueScale: a scalable memory architecture for predictable real-time computing on highly integrated SoCsabstractIn real-time embedded computing, time-predictability and performance are required simultaneously by memory transactions. However, with increasingly more elements being integrated into hardware, memory interconnects become a critical stumbling block to satisfying timing correctness, due to lack of hardware and scheduling scalability. In this paper, we propose a new hierarchically distributed memory interconnect, BlueScale, managing memory transactions using identical Scale Elements, which ensures hardware scalability. The Scale Element introduces two nested priority queues, achieving iterative compositional scheduling for memory transactions, guaranteeing transaction tasks' scheduling schedulability. Associated with the new architecture, a theoretical model is established to improve BlueScale's real-time performance. Zhe Jiang 0004, Kecheng Yang 0001, Neil C. Audsley, Nathan Fisher, Weisong Shi, Zheng Dong 0002 |
DAC | 3 |
| 2022 | PSpSys: A time-predictable mixed-criticality system architecture based on ARM TrustZone
Zhe Jiang 0004, Pan Dong, Qingling Zhao, Dizhong Zhu, Yan Zhuang 0013, Neil C. Audsley |
J. Syst. Archit. | 8 |
| 2022 | Towards an energy-efficient quarter-clairvoyant mixed-criticality system
Zhe Jiang 0004, Kecheng Yang 0001, Nathan Fisher, Neil C. Audsley, Zheng Dong 0002 |
J. Syst. Archit. | 4 |
| 2022 | BlueVisor: Time-Predictable Hardware Hypervisor for Many-Core Embedded SystemsabstractWhilst virtualization was once restricted to large-scale computing platforms, and it is now widely deployed on modern embedded computing systems. This has been driven by the availability of hardware support which alleviates the performance penalties incurred by traditional software virtualization technologies. In the domain of hard real-time systems, specialist virtualization technology which respects restricted timing requirements and constraints can be deployed to allow sharing of processors. However, other aspects of the embedded system (I/O, memory, and communication) are harder to analyse. In this paper, we argue that in order to support real-time virtualization on modern embedded systems, additional system-wide hardware support is required. We propose BlueVisor, an analyzable and scalable hardware hypervisor for many-core embedded systems, which enables time-predictable CPU, memory, and I/O virtualization, as well as supporting a fast interrupt handler, and inter-VM communication. We describe the design and implementation of the real-time hypervisor and demonstrate how a BlueVisor-based virtualization system can be leveraged to meet real-time requirements with significant improvement in system performance, and with a low-performance cost when executing different types of software. Zhe Jiang 0004, Pan Dong, Yan Zhuang 0013, Neil C. Audsley, Ian Gray |
IEEE Trans. Computers | 5 |
| 2022 | Toward an Analysable, Scalable, Energy-Efficient I/O Virtualization for Mixed-Criticality SystemsabstractIn mixed-criticality systems (MCSs), timely handling of I/O operations is a key for the system being successfully implemented and appropriately functioned. The I/O system for an MCS must simultaneously enable different features, including isolation/separation, timing-predictability, performance, scalability, and energy-efficiency. Moreover, such an I/O system also requires to manage I/O resource in an adaptive manner to facilitate efficient yet safe resource sharing among components of different criticality levels. Existing approaches cannot achieve all of these requirements simultaneously. This article presents a mixed-criticality I/O management framework, termed MCS-IOV. MCS-IOV is based on hardware-assisted virtualization, which provides temporal and spatial isolation and prohibits fault propagation with limited extra overhead. MCS-IOV extends a real-time I/O virtualization system, by supporting the concept of mixed criticalities and customized interfaces for schedulers, which offers good timing predictability and scalability. Finally, we introduce an energy management framework for MCS-IOV, ensuring the power-efficiency of the design. The MCS-IOV is the first systematical solution that fulfills all the requirements as a mixed-criticality I/O system. Zhe Jiang 0004, Xiaotian Dai 0001, Pan Dong, Neil C. Audsley, Nan Guan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2022 | Bridging the Pragmatic Gaps for Mixed-Criticality Systems in the Automotive IndustryabstractAn increasingly important trend in the design of safety-critical systems is the integration of components with different levels of criticality onto a common hardware platform. Mixed-criticality systems (MCSs) have been well researched in academia, but can be difficult to implement in industrial scenarios as the theoretical models underpinning the research do not sufficiently consider industrial safety practice and safety standards. In this article, we make the first attempt toward the implementation of the MCS theoretical model in industrial settings. To this end, we identify the pragmatic gaps between theory and practice, and then propose a generic industrial MCS architecture, termedP-MCS(Practical-MCS).P-MCSis built upon the conventional theoretical MCS model with additional considerations of industrial safety requirements: 1) runtime safety analysis, determining preserved applications in each system mode and 2) correct partitioning and isolation of different critical elements. We introduce three implementing methods forP-MCS. Corresponding to the new system architecture, we present a theoretical model and schedulability analysis (with consideration of shared resources) to ensure system predictability. Finally, we evaluate and demonstrateP-MCSin terms of system schedulability, overheads, throughput, and predictability, along with a real-world case study. As shown in the evaluation, the considerations of industrial requirements lead to extra overheads and performance reduction inP-MCS. Such weaknesses can be considerably mitigated by hardware assistance and acceleration. Zhe Jiang 0004, Shuai Zhao 0004, Richard Paterson, Nan Guan, Yan Zhuang 0013, Neil C. Audsley |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2021 | Invited: Hardware/Software Co-Synthesis and Co-Optimization for Autonomous SystemsabstractWith ever more complicated functionalities being integrated in modern autonomous systems, traditional design methods may not remain sufficient to deliver trusted and high-performance systems with stringent temporal, safety and cost efficiency requirements. In this paper, we discuss the limitations of the traditional design methods with the above requirements enforced, in which hardware and software design are often considered separately. To tackle these limitations, this paper presents a novel design solution that synthesizes both software-level and hardware-level design. First, we highlight and analyze the interconnections between software-level methods (e.g. priority assignment and task allocation) and hardware design (e.g. cache and memory management), in terms of the resulting system performance, e.g. latency. Second, by applying the identified interconnections, we propose an optimization framework to produce high-quality synthesized solutions of both software and hardware design based on a set of candidate design methods. In addition, we describe potential research directions derived from the work and major challenges that can be investigated jointly by engineers and researchers from embedded systems, system safety and programming languages communities. Wanli Chang 0001, Shuai Zhao 0004, Simon Burton 0001, Haitong Wang, Ting Chen 0002, Neil C. Audsley |
DAC | 7 |
| 2021 | I/O-GUARD: Hardware/Software Co-Design for I/O Virtualization with Guaranteed Real-time PerformanceabstractFor safety-critical| computer systems, time-predictability and performance are usually required simultaneously in I/O virtualization. However, both requirements are challenging to achieve due to complex I/O access path and resource management at system level and lack of support from preemptive scheduling at I/O hardware level. In this paper, we propose a new framework, I/O-GUARD, which reconstructs the system architecture of I/O virtualization, bringing a dedicated hardware hypervisor to handle resource management throughout the system. The hypervisor improves system real-time performance by enabling preemptive scheduling in I/O virtualization with both analytical and experimental real-time guarantees. Specifically, I/O-GUARD is a First-of-Its-Kind framework for multi-/many-core I/O virtualization. Zhe Jiang 0004, Kecheng Yang 0001, Yunfeng Ma, Nathan Fisher, Neil C. Audsley, Zheng Dong 0002 |
DAC | 5 |
| 2021 | Brief Industry Paper: AXI-InterconnectRT: Towards a Real-Time AXI-Interconnect for System-on-ChipsabstractIn modern, real-time heterogeneous systems, ensuring the predictability of interconnects is becoming increasingly important. Existing interconnects are mainly designed to achieve high throughput, with their micro-architectures usually based on FIFO queues. This FIFO-based design prevents prioritization of transactions based on their importance, leading to difficulties in ensuring transaction predictability, especially in a system with a large number of system components. In this paper, we introduce AXI-InterconnectRT, a real-time AXI interconnect for heterogeneous SoCs, which redefines the micro-architecture of interconnects by enabling random accesses of buffered transactions and organizing transactions using dedicated hardware units. With the new micro-architecture, AXI-InterconnectRTcan manage transactions based on their importance, guaranteeing their predictability. Zhe Jiang 0004, Neil C. Audsley, Dayu Shill, Kecheng Yang 0001, Nathan Fisher, Zheng Dong 0002 |
RTAS | 2 |
| 2021 | HIART-MCS: High Resilience and Approximated Computing Architecture for Imprecise Mixed-Criticality SystemsabstractIn mixed-criticality systems (MCSs), less-critical tasks are often terminated to ensure the correct execution of high critical tasks. This strategy could however lead to safety hazards, and largely reduce system functionality due to the absence of the discarded tasks. To overcome this problem, we introduce a high resilience and approximated computing framework for MCS, i.e., HIART-MCS. HIART-MCS introduces a novel processor which supports approximation at the hardware level. Associated with this, we also introduce a new intermediate system mode which allows less-critical tasks to be executed with reduced precision instead of being directly dropped out. Corresponding to the HIART-MCS, we further present a new theoretical model and schedulability analysis providing a timing guarantee for the system, followed by optimisations of the mode switch strategy. As demonstrated in both the theoretical and practical evaluations, HIART-MCS effectively improves the survivability of less-critical tasks with limited sacrifice of the critical tasks and negligible extra overhead. It is notable that HIART-MCS is the first practical framework for imprecise MCSs. Zhe Jiang 0004, Xiaotian Dai 0001, Neil C. Audsley |
RTSS | 3 |
| 2020 | Re-Thinking Mixed-Criticality Architecture for Automotive IndustryabstractMixed-Criticality System (MCS) has been considered widely within academic literature, but is proving difficulty to implement in industry as the theoretical models underpinning the research do not always consider industrial safety standards and practice (e.g., DO-178C, ISO26262, and EN50128). This paper analyses and formalises the mismatches between theoretical models and industrial standards, and presents a generic industrial MCS architecture, termed as Z-MCS. Z-MCS is built upon the conventional theoretical MCS model (i.e., Adaptive Mixed-Criticality), but with additional satisfaction on the industrial safety requirements: i). run-time safety analysis, which determines preserved applications in each system mode; ii). correct partitioning and isolation of different critical elements with temporal, spatial and fault isolation. Furthermore, three implementing methods of Z-MCS are proposed, with a generic schedulability analysis for timing guarantee. Finally, we evaluate and demonstrate Z-MCS in terms of system schedulability and overheads, along with a real-world case study. In addition, this paper is the first attempt for connecting the theoretical MCS model with the industrial context. Zhe Jiang 0004, Shuai Zhao 0004, Pan Dong, Nan Guan, Neil C. Audsley |
ICCD | 7 |
| 2020 | Addressing Resource Contention and Timing Predictability for Multi-Core Architectures with Shared Memory InterconnectsabstractMulti-core architectures are increasingly being used in real-time embedded systems. In general, such systems have more processors than the shared memory modules, potentially causing severe interference over memory accesses. This resource contention could lead to substantial variation on memory access latencies, and thus wide fluctuation in the overall system performance, which is highly undesirable especially for the time-critical applications. In this paper, we address resource contention and timing predictability for multi-core architectures with distributed memory interconnects. We focus on the locally arbitrated interconnect constructed by pipelined multiplexing stages with local arbitration, while the globally arbitrated interconnect employing global scheduling to the same architecture potentially suffers synchronisation issue and requires strict coordination. Our contributions are mainly threefold: (i) We analyse the resource contention across the memory access data path, and report the accurate calculational method to bound the worst-case behaviour. (ii) We compare the average-case behaviour of the locally arbitrated and the globally arbitrated architectures with experiments, demonstrating varying memory latencies caused by the resource sharing issue. (iii) We propose an architectural modification to smooth resource sharing. Evaluations on simulators and FPGA implementations with synthetic memory workload show that the latency variation is significantly reduced, contributing towards timing predictability of multi-core systems. Haitong Wang, Neil C. Audsley, Wanli Chang 0001 |
RTAS | 2 |
| 2020 | Pythia-MCS: Enabling Quarter-Clairvoyance in I/O-Driven Mixed-Criticality SystemsabstractIn mixed-criticality systems, mode switch is a key strategy which dynamically provides a balance between system performance and safety. In conventional MCS frameworks, mode switch is triggered by the over-execution of a task; i.e., a task overruns the less pessimistic worst-case execution time. In cyber-physical systems, the data volume generated by I/O affects and can even dominate task computation time. With this in mind, we introduce a novel MCS architecture, termed Pythia-MCS, which predicts task execution time according to I/O run-time behaviors. With the new feature of future-prediction, the Pythia-MCS provides more timely, but still accurate, mode switch. We also present a new theoretical model (quarter-clairvoyance), which guarantees the timing predictability of the design, and a new schedulability analysis for the Pythia-MCS, which demonstrates improved schedulability compared to conventional MCS frameworks. The Pythia-MCS is the first MCS framework enabling the clairvoyance functionality. Zhe Jiang 0004, Kecheng Yang 0001, Nathan Fisher, Neil C. Audsley, Zheng Dong 0002 |
RTSS | 4 |
| 2020 | Meshed Bluetree: Time-Predictable Multimemory Interconnect for Multicore ArchitecturesabstractMulticore architectures are widely adopted in the emerging real-time applications, such as autonomous vehicles and robotics, where latency is required to be both bounded in the worst case (i.e., time predictability) and low. With the number of processors growing, the conventional memory interconnects, i.e., shared bus, crossbar, and network-on-chip (NoC), suffer high latency due to the increasing logic size of their centralized arbiter, which is deployed for time predictability. In this article, we introduce a novel distributed multimemory interconnect, Meshed Bluetree, and explain its operation. Constructed by coupling a router network with multiple Bluetree-based memory architectures in parallel, Meshed Bluetree allows simultaneous access to multiple memory modules. We present the analysis for the predictable timing behavior of memory access to bound the worst case. The evaluation of FPGA with synthetic memory workloads and real-world benchmarks demonstrates the effectiveness of our work, i.e., as the number of memory modules increases, the latency is reduced with the same scale. This work reports the first time-predictable distributed multimemory interconnect, significantly contributing to multicore real-time systems. Haitong Wang, Neil C. Audsley, Xiaobo Sharon Hu, Wanli Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | MCS-IOV: Real-Time I/O Virtualization for Mixed-Criticality SystemsabstractIn mixed-criticality systems, timely handling of I/O is a key for the system being successfully implemented and functioning appropriately. The criticality levels of functions and sometimes the whole system are often dependent on the state of the I/O. An I/O system for a MCS must provide simultaneously isolation/separation, performance/efficiency and timing-predictability, as well as being able to manage I/O resource in an adaptive manner to facilitate efficient yet safe resource sharing among components of different criticality levels. Existing approaches cannot achieve all of these requirements simultaneously. This paper presents a MCS I/O management framework, termed MCS-IOV. MCS-IOV is based on hardware assisted virtualisation, which provides temporal and spatial isolation and prohibits fault propagation with small extra overhead in performance. MCS-IOV extends a real-time I/O virtualisation system, by supporting the concept of mixed criticalities and customised interfaces for schedulers, which offers good timing-preditability. MCS-IOV supports I/O driven criticality mode switch (the mode switch can be triggered by detection of unexpected I/O behaviors, e.g., a higher I/O utilization than expected) and timely I/O resource reconfiguration up on that. Finally, We evaluated and demonstrate MCS-IOV in different aspects. Zhe Jiang 0004, Neil C. Audsley, Pan Dong, Nan Guan, Xiaotian Dai 0001, Lifeng Wei |
RTSS | 2 |
| 2019 | Many suspensions, many problems: a review of self-suspending tasks in real-time systemsabstractIn general computing systems, a job (process/task) may suspend itself whilst it is waiting for some activity to complete, e.g., an accelerator to return data. In real-time systems, such self-suspension can cause substantial performance/schedulability degradation. This observation, first made in 1988, has led to the investigation of the impact of self-suspension on timing predictability, and many relevant results have been published since. Unfortunately, as it has recently come to light, a number of the existing results are flawed. To provide a correct platform on which future research can be built, this paper reviews the state of the art in the design and analysis of scheduling algorithms and schedulability tests for self-suspending tasks in real-time systems. We provide (1) a systematic description of how self-suspending tasks can be handled in both soft and hard real-time systems; (2) an explanation of the existing misconceptions and their potential remedies; (3) an assessment of the influence of such flawed analyses on partitioned multiprocessor fixed-priority scheduling when tasks synchronize access to shared resources; and (4) a discussion of the computational complexity of analyses for different self-suspension task models. Jian-Jia Chen, Geoffrey Nelissen, Wen-Hung Kevin Huang, Maolin Yang 0004, Björn B. Brandenburg, Konstantinos Bletsas 0001, Cong Liu 0005, Pascal Richard, Frédéric Ridouard, Neil C. Audsley, Ragunathan Rajkumar, Dionisio de Niz, Georg von der Brüggen |
Real Time Syst. | 10 |
| 2019 | BlueIO: A Scalable Real-Time Hardware I/O Virtualization System for Many-core Embedded SystemsabstractIn safety-critical systems, time predictability is vital. This extends to I/O operations that require predictability, timing-accuracy, parallel access, scalability, and isolation. Currently, existing approaches cannot achieve all these requirements at the same time. In this article, we propose a framework of hardware framework for real-time I/O virtualization—termed BlueIO —to meet all these requirements simultaneously. BlueIO integrates the functionalities of I/O virtualization, low-layer I/O drivers, and a clock cycle level timing-accurate I/O controller (using the GPIOCP [36]). BlueIO provides this functionality in the hardware layer, supporting abstract virtualized access to I/O from the software domain. The hardware implementation includes I/O virtualization and I/O drivers, provides isolation and parallel (concurrent) access to I/O operations, and improves I/O performance. Furthermore, the approach includes the previously proposed GPIOCP to guarantee that I/O operations will occur at a specific clock cycle (i.e., be timing-accurate and predictable). In this article, we present a hardware consumption analysis of BlueIO to show that it linearly scales with the number of CPUs and I/O devices, which is evidenced by our implementation in VLSI and FPGA. We also describe the design and implementation of BlueIO and demonstrate how a BlueIO-based system can be exploited to meet real-time requirements with significant improvements in I/O performance and a low running cost on different OSs. Zhe Jiang 0004, Neil C. Audsley, Pan Dong |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | BlueVisor: A Scalable Real-Time Hardware Hypervisor for Many-Core Embedded SystemsabstractVirtualization technology is widespread in real-time embedded systems, resulting from the availability of hardware support. Hardware assistance allows the penalties suffered by traditional software virtualization technologies to be alleviated, e.g., significant software overhead. However, current technologies are not necessarily applicable to real-time systems as they are not designed to satisfy strict timing requirements and constraints. In this paper, we propose a scalable real-time hardware hypervisor for many-core embedded system, named BlueVisor, developed from our previously proposed real-time I/O hypervisor (VCDC), I/O controller (GPIOCP) and memory interconnect (BlueTree), which enables predictable CPU, memory, and I/O virtualization, as well as fast interrupt handler, and inter-VM communication. We propose the design idea and specific implementation of the real-time hypervisor, as well as demonstrate how a BlueVisor-based virtualization system can be adequately exploited to meet the real-time requirements with significant improvements on system performance, while presenting a low performance cost executing different operating systems (OSs). Zhe Jiang 0004, Neil C. Audsley, Pan Dong |
RTAS | 2 |
| 2017 | GPIOCP: Timing-accurate general purpose I/O controller for many-core real-time systemsabstractModern SoC / NoC chips often provide GeneralPurpose I/O (GPIO) pins for connecting devices that are not directly integrated within the chip. Timing accurate control of devices connected to GPIO is often required within embedded real-time systems - ie. I/O operations should occur at exact times, with minimal error, neither being significantly early or late. This is difficult to achieve due to the latencies and contentions present in architecture, between CPU instigating the I/O operation, and the device connected to the GPIO - software drivers, RTOS, buses and bus contentions all introduce significant variable latencies before the command reaches the device. This is compounded in NoC devices utilising a mesh interconnect between CPUs and I/O devices. The contribution of this paper is a resource efficient programmable I/O controller, termed the GPIO Command Processor (GPIOCP), that permits applications to instigate complex sequences of I/O operations at an exact time, so achieving timing-accuracy at a single clock cycle level. Also, I/O operations can be programmed to occur at some point in the future, periodically, or reactively. The GPIOCP is a parallel I/O controller, supporting cycle level timing accuracy across several devices connected to GPIO simultaneously. The GPIOCP exploits the tradeoff between placing using a full sequential CPU to control each GPIO connected device, which achieves some timing accuracy at high resource cost; and poor timing-accuracy achieved where the application CPU controls the device remotely. The GPIOCP has efficient hardware cost compared to CPU approaches, with the additional benefits of total timing accuracy (CPU solutions do not provide this in general) and parallel control of many I/O devices. Zhe Jiang 0004, Neil C. Audsley |
DATE | 2 |
| 2017 | VCDC: The Virtualized Complicated Device ControllerabstractI/O virtualization enables time and space multiplexing of I/O devices, by mapping multiple logical I/O devices upon a smaller number of physical devices. However, due to the existence of additional virtualization layers, requesting an I/O from a guest virtual machine requires complicated sequences of operations. This leads to I/O performance losses, and makes precise timing of I/O operations unpredictable. This paper proposes a hardware I/O virtualization system, termed the Virtualized Complicated Device Controller (VCDC). This I/O system allows user applications to access and operate I/O devices directly from guest VMs, and bypasses the guest OS, the Virtual Machine Monitor (VMM) and low layer I/O drivers. We show that the VCDC efficiently reduces the software overhead and enhances the I/O performance and timing predictability. Furthermore, VCDC also exhibits good scalability that can handle I/O requests from variable number of CPUs in a system. Zhe Jiang 0004, Neil C. Audsley |
ECRTS | 2 |
| 2017 | A Distributed Stream Library for Java 8abstractJava 8 has introduced new capabilities such as lambda expressions and streams which simplify data-parallel computing. However, as a base language for Big Data systems, it still lacks a number of important capabilities such as processing very large datasets and distributing the computation over multiple machines. This paper gives an overview of the Java 8 Streams API and proposes extensions to allow its use in Big Data systems. It also shows how the API can be used to implement a range of standard Big Data paradigms. Finally, it compares performance with that of Hadoop and Spark. Despite being a proof-of-concept implementation, results indicate that it is a lightweight and efficient framework, comparable in performance to Hadoop and Spark, and is up to 5 times faster for the largest input sizes tested. Yu Chan, Andy J. Wellings, Ian Gray, Neil C. Audsley |
IEEE Trans. Big Data | 4 |
| 2017 | A Globally Arbitrated Memory Tree for Mixed-Time-Criticality SystemsabstractEmbedded systems are increasingly based on multi-core platforms to accommodate a growing number of applications, some of which have real-time requirements. Resources, such as off-chip DRAM, are typically shared between the applications using memory interconnects with different arbitration polices to cater to diverse bandwidth and latency requirements. However, traditional centralized interconnects are not scalable as the number of clients increase. Similarly, current distributed interconnects either cannot satisfy the diverse requirements or have decoupled arbitration stages, resulting in larger area, power and worst-case latency. The four main contributions of this article are: 1) a Globally Arbitrated Memory Tree (GAMT) with a distributed architecture that scales well with the number of cores, 2) an RTL-level implementation that can be configured with five arbitration policies (three distinct and two as special cases), 3) the concept of mixed arbitration policies that allows the policy to be selected individually per core, and 4) a worst-case analysis for a mixed arbitration policy that combines TDM and FBSP arbitration.We compare the performance of GAMT with centralized implementations and show that it can run up to four times faster and have over 51 and 37 percent reduction in area and power consumption, respectively, for a given bandwidth. Manil Dev Gomony, Jamie Garside, Benny Akesson, Neil C. Audsley, Kees Goossens |
IEEE Trans. Computers | 4 |
| 2016 | Architecting Time-Critical Big-Data SystemsabstractCurrent infrastructures for developing big-data applications are able to process -via big-data analytics- huge amounts of data, using clusters of machines that collaborate to perform parallel computations. However, current infrastructures were not designed to work with the requirements of time-critical applications; they are more focused on general-purpose applications rather than time-critical ones. Addressing this issue from the perspective of the real-time systems community, this paper considers time-critical big-data. It deals with the definition of a time-critical big-data system from the point of view of requirements, analyzing the specific characteristics of some popular big-data applications. This analysis is complemented by the challenges stemmed from the infrastructures that support the applications, proposing an architecture and offering initial performance patterns that connect application costs with infrastructure performance. Pablo Basanta-Val, Neil C. Audsley, Andy J. Wellings, Ian Gray, Norberto Fernández García |
IEEE Trans. Big Data | 2 |
| 2015 | A generic, scalable and globally arbitrated memory tree for shared DRAM access in real-time systems
Manil Dev Gomony, Jamie Garside, Benny Akesson, Neil C. Audsley, Kees Goossens |
DATE | 4 |
| 2015 | Task allocation for decoding multiple hard real-time video streams on homogeneous NoCsabstractHard-real time video systems require deterministic admission control decisions to maintain high levels of predictability. These decisions can be based on the state-of-the-art schedulability analysis of tasks and flows. However, due to the pessimistic behaviour of the schedulability analysis and the uncertainties in the application, the multi-core system resources are usually under-utilised. In this paper we propose two task allocation techniques that exploit application and platform characteristics in order to increase the number of simultaneous, fully schedulable, video streams handled by the system. The first, more generic technique, uses the worst-case remaining slack of the mapped tasks as a heuristic to determine the task to processing core allocation. The paper also investigates a second technique that maps the heavily communicating, critical path tasks of the applications onto same core to reduce the communication overhead. We compare against other heuristic based dynamic mapping techniques in the literature, and show that an overall improvement of up to 10%-15% can be obtained, in admission rates and system utilisation. Hashan R. Mendis, Neil C. Audsley, Leandro Soares Indrusiak |
INDIN | 2 |
| 2015 | Reducing the Implementation Overheads of IPCP and DFPabstractMost resource control protocols such as IPCP (Immediate Priority Ceiling Protocol) require a kernel system call to implement the necessary control over any shared data. This call can be expensive, involving a potentially slow switch from CPU user-mode to kernel-mode (and back). In this paper we look at two anticipatory schemes (IPCP and DFP - Deadline Floor Protocol) and show how they can be implemented with the minimum number of calls on the kernel. Specifically, no kernel calls are needed when there is no contention, and only one when there is. A standard implementation would need two such calls. The protocols developed are verified by the use of model checking. A prototype implementation is described for POSIX pThreads (thus opening up improvements to a range of programming approaches). Experimental results demonstrate the effectiveness of the scheme, showing average case savings of 86%. H. Almatary, Neil C. Audsley, Alan Burns 0001 |
RTSS | 2 |
| 2015 | Improving the predictability of distributed stream processors
Pablo Basanta-Val, Norberto Fernández García, Andy J. Wellings, Neil C. Audsley |
Future Gener. Comput. Syst. | 4 |
| 2015 | T-CREST: Time-predictable multi-core architecture for embedded systemsabstractReal-time systems need time-predictable platforms to allow static analysis of the worst-case execution time (WCET). Standard multi-core processors are optimized for the average case and are hardly analyzable. Within the T-CREST project we propose novel solutions for time-predictable multi-core architectures that are optimized for the WCET instead of the average-case execution time. The resulting time-predictable resources (processors, interconnect, memory arbiter, and memory controller) and tools (compiler, WCET analysis) are designed to ease WCET analysis and to optimize WCET performance. Compared to other processors the WCET performance is outstanding. The T-CREST platform is evaluated with two industrial use cases. An application from the avionic domain demonstrates that tasks executing on different cores do not interfere with respect to their WCET. A signal processing application from the railway domain shows that the WCET can be reduced for computation-intensive tasks when distributing the tasks on several cores and using the network-on-chip for communication. With three cores the WCET is improved by a factor of 1.8 and with 15 cores by a factor of 5.7. The T-CREST project is the result of a collaborative research and development project executed by eight partners from academia and industry. The European Commission funded T-CREST. Martin Schoeberl, Sahar Abbaspour, Benny Akesson, Neil C. Audsley, Raffaele Capasso, Jamie Garside, Kees Goossens, Sven Goossens, Scott Hansen, Reinhold Heckmann, Stefan Hepp, Benedikt Huber, Alexander Jordan, Evangelia Kasapaki, Jens Knoop, Yonghui Li 0002, Daniel Wiltsche-Prokesch, Wolfgang Puffitsch, Peter P. Puschner, André Rocha, Cláudio Silva 0002, Jens Sparsø, Alessandro Tocchi |
J. Syst. Archit. | 4 |
| 2014 | Explicit reservation of cache memory in a predictable, preemptive multitasking real-time systemabstractWe describe and evaluate explicit reservation of cache memory to reduce the cache-related preemption delay (CRPD) observed when tasks share a cache in a preemptive multitasking hard real-time system. We demonstrate the approach using measurements obtained from a hardware prototype, and present schedulability analyses for systems that share a cache by explicit reservation. These analyses form the basis for a series of experiments to further evaluate the approach. We find that explicit reservation is most useful for larger task sets with high utilization. Some task sets cannot be scheduled with a conventional cache, but are schedulable with explicit reservation. Jack Whitham, Neil C. Audsley, Robert I. Davis 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2012 | Optimal Program Partitioning for Predictable PerformanceabstractScratchpad memory (SPM) provides a predictable and energy efficient way to store program instructions and data. It would be ideal for embedded real-time systems if not for the practical difficulty that most programs have to be modified in source or binary form in order to use it effectively. This modification process is called partitioning, and it splits a large program into sub-units called regions that are small enough to be stored in SPM. Earlier papers on this subject have only considered regions formed around program structures, such as loops, methods and even entire tasks. Region formation and SPM allocation are performed in two separate steps. This is an approximation that does not make best use of SPM. In this paper, we propose a k-partitioning algorithm as a new way to solve the problem. This allows us to carry out region formation and SPM allocation simultaneously. We can generate optimal partitions for programs expressed either as call trees or by a restricted form of control-flow graph (CFG). We show that this approach obtains superior results to the previous two-step approach. We apply our algorithm to various programs and SPM sizes and show that it reduces the execution time cost for executing those programs relative to execution with cache. Jack Whitham, Neil C. Audsley |
ECRTS | 2 |
| 2012 | Challenges in software development for multicore System-on-Chip developmentabstractMultiprocessor Systems-on-Chip (MPSoC)-based platforms are becoming more common in the embedded domain. Such systems are a significant deviation from the homogeneous, uniprocessor architectures that have been traditionally employed by embedded designers, thereby making the software development process to effectively target the platform more challenging. Low-resource embedded systems rely on efficient implementations that are not well supported by traditional solutions based on architecture virtualisation or middleware. Within this paper we examine these challenges and discuss ways in which they can be mitigated. In particular, we focus on the contributions made by two recent approaches based on Model-Driven Engineering (MDE). We also discuss challenges for future research. Ian Gray, Neil C. Audsley |
RSP | 2 |
| 2012 | Developing Predictable Real-Time Embedded Systems Using AnvilJabstractThis paper proposes Anvil J, a novel technology developed to assist the development of software for predictable, embedded applications. In particular, the work focuses on the complexities of programming for heterogeneous embedded systems in an industrial context, in which the need for predictability is an important requirement. Anvil J converts architecturally-neutral Java code into a set of target-specific programs, automatically distributing the input software over the heterogeneous target architecture whilst ensuring preservation of predictability. During translation it generates a low-to zero-overhead runtime that is tailored to the specific combination of input application and target system, thereby ensuring maximum efficiency. Anvil J uses a technique called Compile-Time Virtualisation that allows it to work with existing compilers and removes the need for language extensions which can hinder certification efforts. Ian Gray, Neil C. Audsley |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2012 | Explicit Reservation of Local Memory in a Predictable, Preemptive Multitasking Real-Time SystemabstractThis paper proposes Carousel, a mechanism to manage local memory space, i.e. cache or scratch pad memory (SPM), such that inter-task interference is completely eliminated. The cost of saving and restoring the local memory state across context switches is explicitly handled by the preempting task, rather than being imposed implicitly on preempted tasks. Unlike earlier attempts to eliminate inter-task interference, Carousel allows each task to use as much local memory space as it requires, permitting the approach to scale to large numbers of tasks. Carousel is experimentally evaluated using a simulator. We demonstrate that preemption has no effect on task execution times, and that the Carousel technique compares well to the conventional approach to handling interference, where worst-case interference costs are simply added to the worst-case execution times (WCETs) of lower-priority tasks. Jack Whitham, Neil C. Audsley |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2012 | Investigation of Scratchpad Memory for Preemptive MultitaskingabstractWe present a multitasking scratchpad memory reuse scheme (MSRS) for the dynamic partitioning of scratchpad memory between tasks in a preemptive multitasking system. We specify a means to compute the worst-case response time (WCRT) and schedulability of task sets executed using MSRS. Our scratchpad-related preemption delay (SRPD) is an analog of cache-related preemption delay (CRPD), proposed in previous work as a way to compute the worst-case cost imposed upon a preempted task by preemption in a multitasking system. Unlike CRPD, however, SRPD is independent of the number of tasks and the local memory size. We compare SRPD with CRPD by experiment and determine that neither dominates the other, i.e. either may be better for certain task sets. However, MSRS leads to improved schedulability versus cache when contention for local memory space is high, either because the local memory size is small, or because the task set is large, provided that the cost of loading blocks from external memory to scratchpad is similar to the cost of loading blocks into cache. Jack Whitham, Robert I. Davis 0001, Neil C. Audsley, Sebastian Altmeyer, Claire Maïza |
RTSS | 3 |
| 2011 | Targeting complex embedded architectures by combining the multicore communications API (mcapi) with compile-time virtualisationabstractWithin the domain of embedded systems, hardware architectures are commonly characterised by application-specific heterogeneity. Systems may contain multiple dissimilar processing elements, non-standard memory architectures, and custom hardware elements. The programming of such systems is a considerable challenge, not only because of the need to exploit large degrees of parallelism but also because hardware architectures change from system to system. To solve this problem, this paper proposes the novel combination of a new industry standard for communication across multicore architectures (MCAPI), with a minimal-overhead technique for targeting complex architectures with standard programming languages (Compile-Time Virtualisation). Ian Gray, Neil C. Audsley |
LCTES | 2 |
| 2010 | Investigating Average versus Worst-Case Timing Behavior of Data Caches and Data ScratchpadsabstractThis paper shows that a program using a time-predictable memory system for data storage can achieve a similar worst-case execution time (WCET) to the average-case execution time (ACET) using a conventional heuristic-based memory system including a data cache. This result is useful within any embedded system where time-predictability and performance are both important, particularly hard real-time systems carrying out intensive data processing activities. It is a counter-example to the conventional wisdom that time-predictable means “slow” in comparison to ACET-focused heuristics. To carry out the investigation, 36 “memory access models” are derived from benchmark programs and assumed to be representative of typical code. The models generate LOAD/STORE instructions to exercise a data cache or scratchpad memory management unit (SMMU). The ACET is determined for the data cache and the WCET is determined for the SMMU. After improvements are applied, results show that the SMMU WCET is within 5% of the data cache ACET for 34 models. In 16 of 36 cases, the SMMU WCET is better than the data cache ACET. Jack Whitham, Neil C. Audsley |
ECRTS | 2 |
| 2010 | Studying the Applicability of the Scratchpad Memory Management UnitabstractA combination of a scratchpad and scratchpad memory management unit (SMMU) has been proposed as a way to implement fast and time-predictable memory access operations in programs that use dynamic data structures. A memory access operation is time-predictable if its execution time is known or bounded-this is important within a hard real-time task so that the worst-case execution time (WCET) can be determined. However, the requirement for time-predictability does not remove the conventional requirement for efficiency: operations must be serviced as quickly as possible under worst-case conditions. This paper studies the capabilities of the SMMU when applied to a number of benchmark programs. A new allocation algorithm is proposed to dynamically manage the scratchpad space. In many cases,the SMMU vastly reduces the number of accesses to dynamic data structures stored in external memory along the worst-case execution path (WCEP). Across all the benchmarks,an average of 47% of accesses are rerouted to scratchpad, with nearly 100% for some programs. In previous scratchpad-based work, time-predictability could only be assured for these operations using external memory.The paper also examines situations in which the SMMU does not perform so well, and discusses how these could be addressed. Jack Whitham, Neil C. Audsley |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2010 | Supporting islands of coherency for highly-parallel embedded architectures using compile-time virtualisationabstractAs their complexity grows, the architectures of embedded systems are becoming increasingly parallel. However, the frameworks used to assist development on highly-parallel general-purpose systems (such as CORBA or MPI) are too heavyweight for use on the non-standard architectures of embedded systems. They introduce significant overheads due to the lack of architectural and structural information contained within most programming languages. Specifically, thread migration across irregular architectures can lead to very poor memory access times, and unconstrained cache coherency cannot scale to cope with large systems. Ian Gray, Neil C. Audsley |
SCOPES | 2 |
| 2010 | Time-Predictable Out-of-Order Execution for Hard Real-Time SystemsabstractSuperscalar out-of-order CPU designs can achieve higher performance than simpler in-order designs through exploitation of instruction-level parallelism in software. However, these CPU designs are often considered to be unsuitable for hard real-time systems because of the difficulty of guaranteeing the worst-case execution time (WCET) of software. This paper proposes and evaluates modifications for a superscalar out-of-order CPU core to allow instruction-level parallelism to be exploited without sacrificing time predictability and support for WCET analysis. Experiments using the M5 O3 CPU simulator show that WCETs can be two-four times smaller than those obtained using an idealized in-order CPU design, as instruction-level parallelism is exploited without compromising timing safety. Jack Whitham, Neil C. Audsley |
IEEE Trans. Computers | 2 |
| 2009 | Exposing non-standard architectures to embedded software using compile-time virtualisationabstractThe architectures of embedded systems are often application-specific, containing multiple heterogenous cores, non-uniform memory, on-chip networks and custom hardware elements (e.g. DSP cores). Standard programming languages do not use these many of these features natively because they assume a traditional single processor and a single logical address space abstraction that hides these architectural details. This paper describes Compile-Time Virtualisation, a technique which uses a virtualisation layer to map software onto the target architecture whilst allowing the programmer to control the virtualisation mappings in order to effectively exploit custom architectures. Ian Gray, Neil C. Audsley |
CASES | 2 |
| 2009 | Implementing time-predictable load and store operationsabstractScratchpads have been widely proposed as an alternative to caches for embedded systems. Advantages of scratchpads include reduced energy consumption in comparison to a cache and access latencies that are independent of the preceding memory access pattern. The latter property makes memory accesses time-predictable, which is useful for hard real-time tasks as the worst-case execution time (WCET) must be safely estimated in order to check that the system will meet timing requirements. Jack Whitham, Neil C. Audsley |
EMSOFT | 2 |
| 2009 | Synthesis of the SR programming language for complex FPGAsabstractMost existing approaches to targeting high-level software to FPGAs are based on extensions to C and do not map easily to the features and characteristics of modern FPGAs. These include massive parallelism and a variety of complex IP-blocks (eg. RAMs, DSPs). In this paper we discuss a hardware implementation of SR, a software language with first class concurrency and high-level IPC.We show that the language model can be implemented efficiently on an FPGA, and that it provides a natural means to encapsulate FPGA resources. We compare against a commercial C-based synthesis tool and achieve similar resource usage using a more expressive language. Nick Gasson, Neil C. Audsley |
FPL | 2 |
| 2008 | Using Trace Scratchpads to Reduce Execution Times in Predictable Real-Time ArchitecturesabstractInstruction scratchpads have been previously suggested as a way to reduce the worst case execution time (WCET) of hard real-time programs without introducing the analysis issues posed by caches. Trace scratchpads extend this paradigm with support for instruction level parallelism (ILP) while preserving simplicity of WCET analysis. In this paper, we demonstrate trace scratchpads using the MCGREP-2 CPU architecture. We provide a sample algorithm to automatically reduce the WCET of a program using a trace scratchpad, and compare the results with the use of an instruction scratchpad. We find that the two types of scratchpad are best used together. Instruction scratchpads provide excellent WCET improvements at low cost, but trace scratchpads reduce WCET further by optimizing worst case (WC) paths and exploiting ILP across basic block boundaries. Using our experimental implementation, we have observed WCET improvements over an instruction scratchpad of up to 149% with some Malardalen WCET benchmarks. Jack Whitham, Neil C. Audsley |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2008 | Forming Virtual Traces for WCET Analysis and ReductionabstractIt is notoriously difficult to model superscalar out-of-order CPUs for the purposes of worst-case execution time (WCET) analysis, which can force the use of simpler CPUs in hard real-time systems. To address this problem, it has been suggested that traces could be used to capture the timing properties of a complex CPU operation scheduler as it runs a sequence of basic blocks. In previous work, traces have been implemented using application-specific microcode. This paper proposes restrictions to a dynamic superscalar out-of-order CPU to implement virtual traces. These have the same timing properties as the traces in previous work, but microcode is not used. Instead, CPU modifications implement the same functionality. This allows traces to be used throughout a program because space requirements are minimal. To take advantage of this, a new allocation algorithm is proposed and evaluated for virtual traces. Jack Whitham, Neil C. Audsley |
RTCSA | 2 |
| 2008 | Predictable Out-of-Order Execution Using Virtual TracesabstractThe problem of worst-case execution time (WCET) analysis of complex CPUs is addressed in this paper using a proposed architectural modification. The virtual trace controller (VTC) constrains execution to follow only the paths that have been considered by the WCETanalysis model, allowing the WCET to be determined safely by measurement. Each path has a constant execution time regardless of CPU complexity because the VTC enforces predictable operation.This paper evaluates the VTC using benchmark programs and the M5 simulator.The results show that guaranteed throughput is increased for many programs using the constrained CPU model versus an idealized in-order design, indicating that the VTC can make complex CPU designs operate predictably without reducing throughputto the level of a simple CPU design. Additional results providemore information about the implications of each of the VTC features.Of all the restrictions introduced for predictability,disabling memory forwarding has the greatest effect on the maximum throughput, although conditional branches can also be significant. This paper suggests ways to improve the VTC to increase the guaranteed throughput. Jack Whitham, Neil C. Audsley |
RTSS | 2 |
| 2007 | An Efficient Page Lock/Release OS Mechanism for Out-of-Core Embedded ApplicationsabstractEmbedded system applications are becoming more complex, requiring increased memory. However, additional physical memory increases system cost and power consumption. Virtual memory techniques such as paging, can make use of low-power auxiliary memory, allowing applications increased memory for execution. Currently paging yields poor performance due to page swapping overheads. This paper presents a combined approach of using application hints along with an efficient page lock/release mechanism in the OS to reduce paging overheads. This makes paging a viable solution to support out-of-core embedded real-time applications. The Co-operative Application Specific Paging (CASP) mechanism presented works in conjunction with most existing page replacement policies, providing explicit support for applications via insertion of paging hints in the application source code. Both automatic and manual methods of inserting hints are described and evaluated. The benchmark results of a CASP implementation in the Linux 2.6.16 kernel have shown significant reduction in the number of page-faults (22.3%) and a considerable improvement in application execution times (12.5%). Ameet Patil, Neil C. Audsley |
RTCSA | 2 |
| 2007 | Efficiently Accessing Remote Resources in Distributed Real-Time SystemsabstractThis paper examines the structures that are used by traditional real-time and networked operating systems in order to show that they are poorly suited to providing efficient access to remote devices. It also argues that efficiency can be improved by being more specific and by making better use of the network mediums characteristics. Paul Simon Usher, Neil C. Audsley |
RTCSA | 2 |
| 2007 | A Deterministic Implementation Process for Accurate and Traceable System Timing and Space AnalysisabstractTraditionally, implementations of dependable real-time systems have targeted CPUs, with application level concurrency implemented as pseudo-concurrency on the CPU. For such systems, much research has addressed timing and resource analysis to enable offline guarantees regarding actual worst-case run-time performance. Three major weaknesses exist with the traditional implementation method. Firstly, analysis is post-hoc, after application compilation and worst-case execution time analysis. Secondly, timing analysis is pessimistic and difficult, due to the unpredictable nature of complex CPUs. Thirdly, the compilation process is largely non-traceable, in that it is difficult to relate object code back to source code (which introduces verification difficulties in safety-critical systems). This paper addresses these three problems with an implementation approach and analysis method that: enables timing and space properties to be established directly from source (not after compilation); provides a deterministic and traceable implementation to ease verification; and enables non-pessimistic timing analysis of the implementation as no CPU is utilised. As an exemplar of the approach, the compilation of a standard real-time safety-critical subset of Ada to a circuit (implemented on field programmable gate array) is presented. Neil C. Audsley |
RTCSA | 2 |
| 2006 | Syntax-driven implementation of software programming language control constructs and expressions on FPGAsabstractThis paper considers the efficient parallel implementation of control constructs and expressions written in a common software programming language and synthesised to FPGA platforms. The context of this work are Syntax-Driven Language Specific Processors (SDLSP). An SDLSP for a given software programming language has its architecture defined by the grammar rules of the language itself. Each statement and expression rule in the grammar is implemented on the FPGA, together with sufficient control logic to load program statements sequentially onto the processor, and interface with program store. The instructions executed are a high-level (effectively one-to-one) encoding of the application software program. The advantages of this approach lie in its parallelism and space-efficiency. Syntax-driven language processors take less space than a full CPU on FPGA, and execute statements with a comparable speed; take significantly less space in general than directly compiled approaches (such as Handel-C), although have longer execution times for the same code. Neil C. Audsley |
CASES | 1 |
| 2006 | Towards a File System Interface for Mobile Resources in Networked Embedded SystemsabstractNetworks for real-time embedded systems are a key emerging technology for current and future systems. Such networks need to enable reliable communication without requiring significant resources, and provide an easy programming interface. This paper considers a file-system interface across all resources in a networked embedded system, ie. an application can access local, remote and mobile resources using a file interface. The approach is based on Styx (Dorward et al., 1997), part of the network protocol of the Inferno/Plan 9 OS. The Styx protocol provides file system level abstractions for ease of developing and management at an application layer. To this, we have added limited fault-tolerance and potential mobility for resources. To ensure applicability in a low-resource context, we have defined and implemented a (hardware) Styx IP-core Module1, removing the need for a CPU and software overhead. Neil C. Audsley, Ameet Patil |
ETFA | 1 |
| 2006 | MCGREP - A Predictable Architecture for Embedded Real-Time SystemsabstractReal-time systems design involves many important choices, including that of the processor. The fastest processors achieve performance by utilizing architectural features that make them unpredictable, leading to difficulties proving offline that application process deadlines will be met, in the worst-case. Utilizing slower, more predictable processors, may not provide sufficient instruction throughput to execute all required application processes. This exposes a key trade-off in processor selection for real-time systems: predictability versus instruction throughput. This paper proposes MCGREP, a novel CPU architecture that combines predictability, high instruction throughput and flexibility. MCGREP is entirely microprogrammed, with multiple execution units. Basic operation involves implementation of a conventional set of CPU instructions in microcode - MCGREP then executes object code suitably compiled. Advanced operation allows the application to dynamically load new microcode, enabling new application specific instructions to increase overall performance. MCGREP is implemented upon reconfigurable logic (FPGA) - an increasingly important platform for the embedded RTS. Custom microcode configurations for new instructions are generated from C source. MCGREP is shown to have performance comparable to two popular FPGA softcore CPUs (OpenRISC and Microblaze, the latter a commercial product). Flexibility is demonstrated by implementing an existing instruction set (OpenRISC) in microcode, with application-specific instructions to improve overall performance. As a further demonstration, predictable two-level interrupt and synchronization mechanisms are programmed in microcode Jack Whitham, Neil C. Audsley |
RTSS | 2 |
| 2006 | Optimal priority assignment in the presence of blocking
Konstantinos Bletsas 0001, Neil C. Audsley |
Inf. Process. Lett. | 2 |
| 2005 | Implementing Application Specific RTOS Policies using ReflectionabstractConventionally, a real-time operating system (RTOS) is built without knowing which specific applications is executed upon it. The RTOS is built for the general case, rather than to meet the specific requirements of an application. This paper proposes a generic module-based reflective framework to implement an RTOS that allows applications to dynamically adapt the policies within the RTOS to better meet application-specific requirements. The specific approach taken is to augment a conventional /spl mu/-kernel with a module-based reflective mechanism that allows applications to dynamically change the behaviour of themselves, and the policies of the underlying RTOS. Reflection is used to allow applications and system modules to access key OS data structures to obtain information pertaining to the current system performance and resource management policies (e.g. scheduling). An application is then able to modify or introduce new policies into the RTOS to satisfy its demands. Evaluation of our approach shows a considerable performance gain. Ameet Patil, Neil C. Audsley |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |
| 2005 | Extended Analysis with Reduced Pessimism for Systems with Limited ParallelismabstractUnder limited parallelism, processes competing for a single processor may issue at any time operations on remote co-processors, during which the processor is not idled but granted to other ready processes instead. We reduce the pessimism in existing worst-case response time (WCRT) analysis for such systems by examining temporal patterns of local/remote execution. We extend to multi-CPU variants of the model and offer a WCRT-based feasibility test for symmetric multiprocessor (SMP) systems. Konstantinos Bletsas 0001, Neil C. Audsley |
RTCSA | 2 |
| 2004 | Fixed Priority Timing Analysis of Real-Time Systems with Limited Parallelism
Neil C. Audsley, Konstantinos Bletsas 0001 |
ECRTS | 1 |
| 2004 | Realistic Analysis of Limited Parallel Software / Hardware ImplementationsabstractProposed real-time system implementations combine reconfigurable hardware (for speed-up) with processor-memory architectures. Such hardware can execute many functions in parallel, leading to a limited parallel system where a single software process can execute on the processor at any time, in parallel with a number of functions implemented on the reconfigurable hardware. This approach is not amenable to conventional fixed priority timing analysis, as fundamental assumptions are compromised, namely that of a critical instant. This paper describes generalised fixed priority timing analysis for limited parallel systems, illustrated by an example system utilising field programmable gate arrays as the reconfigurable hardware resource. Neil C. Audsley, Konstantinos Bletsas 0001 |
IEEE Real-Time and Embedded Technology and Applications Symposium | 1 |
| 2002 | Hardware implementation of the Ravenscar Ada tasking profileabstractReal-Time Systems place large demands on the languages used to implement them. Processor based implementation methods do not allow accurate timing analysis of systems due to the complexity of modern processors. FPGAs provide a means to implement a real-time system in a way that allows accurate timing analysis to be performed.Existing hardware implementations of high-level programming languages do not support the needs of real-time systems. This paper presents a hardware implementation of the SPARK Ravenscar subsets of Ada which can be accurately analysed for its timing properties. A method of compiling sequential Ada programs has been described elsewhere [21], and this is expanded to include the compilation of protected objects and tasks. The effect this has on the ability to analyse the timing of the program is then examined. Neil C. Audsley |
CASES | 2 |
| 2001 | Hardware compilation of sequential AdaabstractNormal implementations of real-time systems on conventional processors are becoming much more difficult to prove correct to their timing specification. This is due to the complexity of modern processors (e.g. the worst case execution time of a program becomes hard to calculate in the presence of CPU speed up features such as caches and pipelines).Field Programmable Gate Arrays (FPGAs) provide a way to ease this problem by providing an implementation medium that has a simple timing model. However there is no support for real-time languages on FPGAs.This paper describes a compiler for a sequential subset of Ada95, concentrating upon compilation of subprograms and statements. It is shown how the resulting circuits give simple timing analysis. Extensions to the current compiler are explored to give support for a larger range of types and a predictable subset of the Ada95 concurrency model. Neil C. Audsley |
CASES | 2 |
| 2001 | Predictable and Efficient Virtual Addressing for Safety-Critical Real-Time SystemsabstractConventionally, the use of virtual memory in safety-critical real-time systems has been avoided, one reason being the difficulties it provides to timing analysis. The difficulties arise due to the Memory Management Unit (MMU) on commercial processors being optimised to improve average performance, to the detriment of simple worst-case analysis. However within safety-critical systems, there is a move towards implementations where processes of differing integrity levels are allocated to the same processor. This requires adequate partitioning between processes of different integrity levels. One method for achieving this in the context of commercial processor is via use of the MMU and its support for virtual memory. The focus of this paper is upon the provision of virtual memory for processes of all integrity levels without complicating the timing analysis of safety-critical processes with hard deadlines. Also, for lower integrity processes without hard deadlines, the flexibility of the virtual memory provided does not restrict the process functionality, The virtual memory system proposed is generic and can be implemented on many commercial architectures e.g. PowerPC, ARM and MIPS. This paper details the PowerPC implementation. M. D. Bennett, Neil C. Audsley |
ECRTS | 2 |
| 2001 | On priority assignment in fixed priority scheduling
Neil C. Audsley |
Inf. Process. Lett. | 1 |
| 1998 | On Fixed Priority Scheduling, Offsets and Co-Prime Task Periods
Neil C. Audsley, Alan Burns 0001 |
Inf. Process. Lett. | 1 |
| 1996 | Analysing APEX applicationsabstractThe next generation of civil aircraft may be produced using Integrated Modular Avionics (IMA). A component of IMA is APEX, a standard operating system interface. This supports a two-level scheduling scheme consisting of fixed priority scheduling within a statically generated cyclic schedule. This paper illustrates how APEX applications can be analysed for their response times and shows that there is potential for a large amount of release jitter. Neil C. Audsley, Andy J. Wellings |
RTSS | 1 |
| 1995 | Fixed Priority Pre-emptive Scheduling: An Historical Perspective
Neil C. Audsley, Alan Burns 0001, Robert I. Davis 0001, Ken Tindell, Andy J. Wellings |
Real Time Syst. | 1 |
| 1994 | Mechanisms for Enhancing the Flexibility and Utility of Hard Real-Time SystemsabstractAdaptive and dynamic behaviour is seen as one of the key characteristics of next generation hard real-time systems. Whilst fixed priority pre-emptive scheduling is rapidly becoming a de facto standard in real-time systems engineering, it remains inflexible in its purest form. One method of increasing flexibility is via the incorporation of optional components into processes with hard deadlines. Such components are not guaranteed off-line, but may be accepted at run-time if sufficient spare capacity becomes available. This paper describes new mechanisms which are required to schedule effectively optional components: mechanisms which enable spare capacity to be detected early and on-line guarantees to be given.> Neil C. Audsley, Robert I. Davis 0001, Alan Burns 0001 |
RTSS | 1 |
| 1994 | STRESS: a Simulator for Hard Real-time SystemsabstractAbstract The STRESS environment is a collection of CASE tools for analysing and simulating the behaviour of hard real‐time safety‐critical applications. It is primarily intended as a means by which various scheduling and resource management algorithms can be evaluated, but can also be used to study the general behaviour of applications and real‐time kernels. This paper describes the structure of the STRESS language and its environment, and gives examples of its use. Neil C. Audsley, Alan Burns 0001, Mike F. Richardson, Andy J. Wellings |
Softw. Pract. Exp. | 1 |