EDBT 2026 Demo / reviewers in the wild / expert
Selma Saidi
dblp:95/10843
· DBLP profile ↗
30ranked-venue papers
7as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 6 first-author · 9 since 2021Software engineering, systems software and programming languages · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Global Scheduling of Weakly-Hard Real-Time Tasks using Job-Level Priority ClassesabstractReal-time systems are intrinsic components of many pivotal applications, such as self-driving vehicles, aerospace and defense systems. The trend in these applications is to incorporate multiple tasks onto fewer, more powerful hardware platforms, e.g., multi-core systems, mainly for reducing cost and power consumption. Many real-time tasks, like control tasks, can tolerate occasional deadline misses due to robust algorithms. These tasks can be modeled using the weakly-hard model. Literature shows that leveraging the weakly-hard model can relax the over-provisioning associated with designed real-time systems. However, a wide-range of the research focuses on single-core platforms. Therefore, we strive to extend the state-of-the-art of scheduling weakly-hard real-time tasks to multi-core platforms. We present a global job-level fixed priority scheduling algorithm together with its schedulability analysis. The scheduling algorithm leverages the tolerable continuous deadline misses to assigning priorities to jobs. The proposed analysis extends the Response Time Analysis (RTA) for global scheduling to test the schedulability of tasks. Hence, our analysis scales with the number of tasks and number of cores because, unlike literature, it depends neither on Integer Linear Programming nor reachability trees. Schedulability analyses show that the schedulability ratio is improved by 40% comparing to the global Rate Monotonic (RM) scheduling and up to 60% more than the global EDF scheduling, which are the state-of-the-art schedulers on the RTEMS real-time operating system. Our evaluation on industrial embedded multi-core platform running RTEMS shows that the scheduling overhead of our proposal does not exceed 60 nanosecond. Victor Gabriel Moyano, Zain Alabedin Haj Hammadeh, Selma Saidi, Daniel Lüdtke |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | Teleoperation as a Step Towards Fully Autonomous SystemsabstractIn the foreseeable future, highly automated mobile systems, such as vehicles, robots, UAVs, or trains, will be confronted with difficult situations that require external support. The availability of such external support corresponds to level 4 driving automation and is an essential feature in current robotaxis and automated public transportation. While the first generation of level 4 prototypes relied on safety driver support, commercial systems are gradually moving towards support by teleoperation. Designing teleoperation support for level 4 systems is an end-to-end problem involving two main research and practical challenges, the teleoperation function defining the remote human interface with its scene representation and available control functions, and the real-time communication channel involving wired and wireless segments, which must provide reliable end-to-end data transport. Both challenges are tightly linked, and combined solutions are needed to reach the required safe teleoperation. Solutions can make use of the rich sensing and control system of a level 4 vehicle, which however can only be exploited if the communication channel provides adequate real-time access. Alex Bendrick, Daniel Tappe, Nora Sperling, Rolf Ernst, Andrea Nota, Selma Saidi, Frank Diermeyer |
DATE | 6 |
| 2025 | Obstacle-Aware Proactive Beam Management Based on Real-Time Ray Tracing for mmWave NetworksabstractThe increasing demand for efficient millimeter-wave (mmWave) beam management in wireless communication systems of 5G, 6G and beyond necessitates innovative approaches to mitigate the resource-intensive and latency-prone nature of exhaustive search methods. This paper presents a novel approach leveraging FPGA-accelerated real-time ray tracing for dynamic beam management. The proposed method dynamically reduces the number of beam directions and burst periodicity based on the mobility and density of mobile devices, as well as the presence of obstacles and available propagation paths. By utilizing a digital twin and real-time ray tracing, the system estimates the required beam directions, enabling adaptive beam management that allows to respond to dynamic obstacles and user mobility. The efficacy of the proposed methodology is showcased in an initial vehicular logistics scenario and can be transferred to further scenarios like intelligent transportation systems and others. Within the sample scenario, the resource overhead is reduced to roughly a third, while the utilization of the base station antenna's beamforming gain is significantly improved. Karsten Heimann, Jintong An, Louay Hito, Selma Saidi, Christian Wietfeld |
VTC2025-Spring | 4 |
| 2025 | Towards Efficient Multi-Frame Clustering in Response Time Analysis for Large Object CommunicationabstractIn autonomous systems, growing sizes of application data, primarily related to perception tasks, have to be transmitted over communication infrastructures that provide higher data rates. Knowledge and exploitation of the clustered structure of multi-frame application data have proven to reduce the interference between different real-time critical communication streams, as recently shown for synchronous systems. However, the accompanying increase in frame numbers will foreseeably make the corresponding analysis impractical due to frame-dependent processing times. As a solution, based on the synchronous example mentioned, we demonstrate how to compose multi-frame transmissions to larger frame clusters in the analysis and decouple the analysis complexity from the application data sizes. Based on an automotive and industrial TSN use case, our results show comparable analytical response times and very low computation times at arbitrary data rates and sizes. It thus enables the deployment of efficient configuration methods that arbitrate large data samples transmitted through the network as one whole frame cluster. In addition, we made the developed analysis openly accessible. Jonas Peeck, Rolf Ernst, Selma Saidi |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | Trace-Enabled Timing Model Synthesis for ROS2-based Autonomous ApplicationsabstractAutonomous applications are typically developed over Robot Operating System 2.0 (ROS2) even in time-critical systems like automotive. Recent years have seen increased interest in developing model-based timing analysis and schedule opti-mization approaches for ROS2-based applications. To complement these approaches, we propose a tracing and measurement framework to obtain timing models of ROS2-based applications. It offers a tracer based on extended Berkeley Packet Filter that probes different functions in ROS2 middleware and reads their arguments or return values to reason about the data flow in applications. It combines event traces from ROS2 and the operating system to generate a directed acyclic graph showing ROS2 callbacks, precedence relations between them, and their timing attributes. While being compatible with existing analyses, we also show how to model (i) message synchronization, e.g., in sensor fusion, and (ii) service requests from multiple clients, e.g., in motion planning. Considering that, in real-world scenarios, the application code might be confidential and formal models are unavailable, our framework still enables the application of existing analysis and optimization techniques. We demonstrate our framework's capabilities by synthesizing the timing model of a real-world benchmark implementing LIDAR-based localization in Autoware's Autonomous Valet Parking. Hazem Abaza, Debayan Roy, Shiqing Fan, Selma Saidi, Antonios Motakis |
DATE | 4 |
| 2024 | Time-Bounded Throughput Guarantees Considering Channel Variations in 5GabstractWith the introduction of the 5th generation of mobile communication networks (5G) and the on-going research for the 6th generation (6G), the adoption of mission critical applications is increasing, where flexible and highly reliable communication systems are needed. 5G has offered new and advanced features with a prominent example being the Ultra Reliable and Low Latency Communications (URLLC) features that can provide high reliability targets for critical applications. With this work we provide potential enhancements towards the 6G, that can provide reliability features for Enhanced Mobile Broadband (eMBB) traffic. Our proposal does not rely on any modification of the Radio Access Network (RAN) protocols or features but it provides an easy to implement solution to operators, reducing also the standardization overhead. More specifically, our proposed method can help to improve reliability of services by bounding the reported achievable throughput enabling the application to improve the overall data transmission. Specifically of importance is that our method does not require any Machine Learning (ML) technique alleviating the need for lengthy data collection processes, nor any other life-cycle management aspects. Our results are based on a realistic measurement campaign coming from a deployed network in an industrial environment. To demonstrate the effectiveness of our procedure, we compare it with a deep learning-based method for throughput prediction. Andrea Nota, Selma Saidi, Alex Palaios |
ETFA | 2 |
| 2024 | HLS-Based Approach for Embedded Real-Time Ray Tracing in Wireless CommunicationsabstractWith the development of wireless communication technology, complex and dynamic scenarios pose great challenges to the Quality of Service (QoS) of wireless communication, especially in indoor scenarios. The quality of beam management can be greatly improved if signal ray-tracing module is embedded in wireless devices to handle synthetic multipath transmissions in real time. In this article, a novel reflection path derivation algorithm for ray tracing of signal beams is proposed, which builds the core mechanism of the proposed FPGA accelerator for ray tracing: by decomposing the computation of the entire ray path into mutually independent subproblems associated with the respective planes involved in the reflection and implemented by independent processing element on FPGAs, the parallelization of the entire ray tracing is realized, which significantly improves the convergence speed of the ray tracing; meanwhile, a new high-level synthesis workflow corresponds to the proposed algorithm and hardware architecture is proposed, which opens the door on synthesizing embedded hardware dedicated for robust and real-time wireless communication. After validation, the method proposed in this article can generate FPGA accelerator for real-time ray-tracing effectively, which achieves ray-tracing simulation in milliseconds. Jintong An, Selma Saidi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | RDMA-Based Deterministic Communication Architecture for Autonomous DrivingabstractAutonomous driving is a big challenge for next-generation vehicles and requires multiple computationally-intensive deep neural networks (DNNs) to be implemented on distributed automotive platforms. Distributed software-enabling autonomous functionalities-has strict timing requirements, e.g., low and deterministic end-to-end latency. Such timings rely on the communication technologies used in the automotive platform, as much on the computation performance of CPUs, GPUs, TPUs, and FPGAs. Hence, we advocate the use of Remote Direct Memory Access (RDMA) technology-typically used in data centers-in automotive platforms. As shown by our experiments with real hardware, Soft-RoCE (software implementation of RDMA) offers low latency communication because of minimal CPU involvement and reduced memory copies. Simultaneously, we show that the native implementation of RDMA does not support determinism, i.e., there is a high variation in communication delays in the presence of interfering data packets. To mitigate this issue, we propose a multi-layer communication stack comprising a deterministic scheduler on top of the Soft-RoCE layer. Further, we have developed a C++ library that offers easy-to-use communication interfaces for distributed applications while implementing the proposed architecture. Experiments show that our library (i) reduces the end-to-end latency of distributed object detection by nearly 9% while having an implementation overhead of less than 1.5% and (ii) minimizes the effects of other data traffic on the delay in high-priority communication. Hazem Abaza, Abhinaba Habishyashi, Debayan Roy, Andrea Bastoni, Zain Alabedin Haj Hammadeh, Shiqing Fan, Selma Saidi, Sergey Tverdyshev |
RTCSA | 7 |
| 2022 | Providing Response Times Guarantees for Mixed-Criticality Network Slicing in 5GabstractMission critical applications in domains such as Industry 4.0, autonomous vehicles or Smart Grids are increasingly dependent on flexible, yet highly reliable communication systems. In this context, Fifth Generation of mobile Communication Networks (5G) promises to support mixed-criticality applications on a single unified physical communication network. This is achieved by a novel approach known as network slicing, that promises to fulfil diverging requirements while providing strict separation between network tenants. We focus in this work on hard performance guarantees by formalizing an analytical method for bounding response times in mixed-criticality 5G network slicing. We reduce pessimism considering models on workload variations. Andrea Nota, Selma Saidi, Dennis Overbeck, Fabian Kurtz, Christian Wietfeld |
DATE | 2 |
| 2022 | Contract-Based Quality-of-Service Assurance in Dynamic Distributed SystemsabstractTo offer an infrastructure for autonomous systems offloading parts of their functionality, dynamic distributed systems must be able to satisfy non-functional quality-of-service (QoS) requirements. However, providing hard QoS guarantees without complex global verification that are satisfied even under uncertain conditions is very challenging. In this work, we propose a contract-based QoS assurance for centralized, hierarchical systems, which requires local verification only and has the potential to cope with dynamic changes and uncertainties. Lea Schönberger, Susanne Graf, Selma Saidi, Dirk Ziegenbein, Arne Hamann 0001 |
DATE | 3 |
| 2022 | Context-based Latency Guarantees Considering Channel Degradation in 5G Network SlicingabstractMission critical applications in domains such as Industry 4.0, autonomous vehicles or smart grids are increasingly dependent on flexible, yet highly reliable communication systems. The Fifth Generation of mobile Communication Networks (5G) promises to support critical communications on a single unified physical communication network through a novel approach known as network slicing. We focus in this work on context-based hard performance guarantees by formalizing an analytical method for bounding response times in critical systems. This approach allows to consider different contexts based on models of degradation of channel quality, and avoids a global highly pessimistic worst-case bound computed for worst possible channel conditions. We demonstrate that the proposed method for computing context-based response times guarantees successfully bounds results obtained in realistic mobility scenarios using a machine-learning based 5G simulation framework. Andrea Nota, Selma Saidi, Dennis Overbeck, Fabian Kurtz, Christian Wietfeld |
RTSS | 2 |
| 2021 | The Road towards Predictable Automotive High - Performance PlatformsabstractDue to the trends of centralizing the EIE architecture and new computing-intensive applications, high-performance hardware platforms are currently finding their way into automotive systems. However, the Systems-on-Chip (SoCs) currently available on the market have significant weaknesses when it comes to providing predictable performance for time-critical applications. The main reason for this is that these platforms are optimized for average-case performance. This shortcoming represents one major risk in the development of current and future automotive systems. In this paper we describe how highperformance and predictability could (and should) be reconciled in future HW /SW platforms. We believe that this goal can only be reached via a close collaboration among system suppliers, IP providers, semiconductor companies, and OS/hypervisor vendors. Furthermore, academic input will be needed to solve remaining challenges and to further improve initial solutions. Falk Rehm, Jörg Seitter, Jan-Peter Larsson, Selma Saidi, Giovanni Stea, Raffaele Zippo, Dirk Ziegenbein, Matteo Andreozzi, Arne Hamann 0001 |
DATE | 4 |
| 2020 | Building End-to-End IoT Applications with QoS GuaranteesabstractMany industrial players are currently challenged in building distributed CPS and IoT applications with stringent end-to-end QoS requirements. Examples are Vehicle-to-X applications, Advanced Driver-Assistance Systems (ADAS) or functionalities in the Industrial Internet of Things (IIoT). Currently, there is no comprehensive solution allowing to efficiently program, deploy, and operate such distributed applications. This paper will focus on real-time concerns, in building distributed CPS and IoT systems. Thereby, the focus lies, on the one hand, on mechanisms required inside of the IoT (compute) nodes, and, on the other hand, on communication protocols such as TSN and 5G connecting them. In the authors' view, the required building blocks for a first end-to-end technology stack are available. However, their integration into a holistic framework is missing. Arne Hamann 0001, Selma Saidi, David Ginthör, Christian Wietfeld, Dirk Ziegenbein |
DAC | 2 |
| 2020 | EDA for Autonomous Behavior AssuranceabstractAutonomous systems are self-governed and self-adaptive systems that must additionally comply with high assurance correctness and safety criteria. Such autonomous systems cannot be tested and verified in the traditional design process. While all systems hardware and software components can be implemented as usual, test and verification only cover the autonomous system functionality, but do not include the goal-driven autonomous behavior in all possible circumstances. This autonomous behavior is a primary design target. Thus, autonomous systems pose a number of emerging challenges and opportunities to the field of electronic design automation (EDA). Examples include specification of (evolving) requirements involving components and their interaction, defining different assurance levels for bounded operational environments, synthesis of mechanisms to guideline diagnosis and rigorously monitor systems integration after deployment. Selma Saidi, Dirk Ziegenbein, Jyotirmoy V. Deshmukh, Rolf Ernst |
ICCAD | 1 |
| 2019 | On Analyzing Memory Latency for Embedded CPS PlatformsabstractThe current trend towards automation and connectivity is driving the increased adoption of complex and parallel computational embedded multiprocessors platforms in CPS. These platforms are characterized by a tightly-coupled shared memory system. Data storage and access have then a significant impact on performance and need to be carefully considered to comply with stringent timing constraints often required by CPS. However, verifying such timing requirements becomes more and more challenging due to the increasing complexity of the underlying memory system thereby leading to non-deterministic access latencies. In this paper we present some of the features that need to be considered when bounding shared memory latency in complex systems and discuss how standard system performance analysis methods can be enhanced to consider specific shared-memory hardware and software features like address mapping, locality of accesses and requests interleaving. Selma Saidi |
DSD | 1 |
| 2019 | Code-Inherent Traffic Shaping for Hard Real-Time SystemsabstractModern hard real-time systems evolved from isolated single-core architectures to complex multi-core architectures which are often connected in a distributed manner. With the increasing influence of interconnections in hard real-time systems, the access behavior to shared resources of single tasks or cores becomes a crucial factor for the system’s overall worst-case timing properties. Traffic shaping is a powerful technique to decrease contention in a network and deliver guarantees on network streams. In this paper we present a novel approach to automatically integrate a traffic shaping behavior into the code of a program for different traffic shaping profiles while being as least invasive as possible. As this approach is solely depending on modifying programs on a code-level, it does not rely on any additional hardware or operating system-based functions. We show how different traffic shaping profiles can be implemented into programs using a greedy heuristic and an evolutionary algorithm, as well as their influences on the modified programs. It is demonstrated that the presented approaches can be used to decrease worst-case execution times in multi-core systems and lower buffer requirements in distributed systems. Dominic Oehlert, Selma Saidi, Heiko Falk |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | Compiler-based Extraction of Event Arrival Functions for Real-Time Systems AnalysisabstractEvent arrival functions are commonly required in real-time systems analysis. Yet, event arrival functions are often either modeled based on specifications or generated by using potentially unsafe captured traces. To overcome this shortcoming, we present a compiler-based approach to safely extract event arrival functions. The extraction takes place at the code-level considering a complete coverage of all possible paths in the program and resulting in a cycle accurate event arrival curve. In order to reduce the runtime overhead of the proposed algorithm, we extend our approach with an adjustable level of granularity always providing a safe approximation of the tightest possible event arrival curve. In an evaluation, we demonstrate that the required extraction time can be heavily reduced while maintaining a high precision. Dominic Oehlert, Selma Saidi, Heiko Falk |
ECRTS | 2 |
| 2018 | Exploiting Locality for the Performance Analysis of Shared Memory Systems in MPSoCsabstractThe integration trend and increased required computing power is driving the advent of common embedded consumer devices like MPSoCs platforms in the safety critical domain. MPSoCs often feature a shared tightly-coupled memory system where a careful management of data storage and transfers is a key enabler for performance. However, providing real-time guarantees for these platforms is extremely challenging as they rely on exploiting data locality to improve average latencies in shared-memory architectures. This effect is often disregarded by existing real-time analysis approaches which furthermore often focus solely on a single component of the memory system. In this paper, we propose a framework for the timing analysis of shared memory systems composed of on-chip scratchpad memories, off-chip DRAMs and DMA engines. The analysis captures the effect on the performance of the system of the locality of accesses, their interleaving and granularity. Selma Saidi, Alexander Syring |
RTSS | 1 |
| 2017 | On the Benefits of Multicores for Real-Time SystemsabstractIn real-time and safety-critical systems, the move towards multicores is becoming unavoidable in order to keep pace with the increasing required processing power and to meet the high integration trend while maintaining a reasonable power consumption. However, the benefit expected from multicore platforms may not step up to the mark, and real-time constraints can be easily violated. Indeed, an efficient use of multicore platforms usually requires a suitable parallelization of the software and a proper knowledge of the hardware constraints in order to find the best trade-off between increasing the parallelization benefit and decreasing contentions from accessing shared resources. In this paper, we illustrate this duality and investigate a general approach to automate some deployment decisions in order to maximize the benefit from using multicores while meeting the schedulability constraints of the system. Selma Saidi |
DSD | 1 |
| 2017 | Designing Networks-on-Chip for High Assurance Real-Time SystemsabstractConventional fault-tolerance approaches for Networks-on-Chip (NoCs) cannot be applied to high assurance real-time systems due to their different goals and constraints. These systems impose strict integrity, resilience and real-time requirements. All possible effects of hardware errors must be taken into account and the resulting system must be predictable, even in the presence of errors. In this paper, we present a wormhole-switched NoC with virtual channels for high assurance real-time systems hardened against soft errors. All possible duration and impacts of soft errors are taken into account and the resulting NoC operates with formal guarantees. Experimental evaluation shows that the network is able to provide a predictable behavior even in aggressive environments with very high error rates. Eberle A. Rambo, Christoph Seitz, Selma Saidi, Rolf Ernst |
PRDC | 3 |
| 2017 | Ensuring safety and efficiency in networks-on-chip
Adam Kostrzewa, Selma Saidi, Leonardo Ecco, Rolf Ernst |
Integr. | 2 |
| 2016 | Dynamic admission control for real-time networks-on-chipsabstractNetworks-on-Chip (NoCs) for real-time systems require solutions for safe and predictable sharing of network resources between transmissions with different quality-of service requirementrs. In this work, we present a mechanism for a global and dynamic admission control in NoCs designed for realtime systems. It introduces an overlay network to synchronize transmissions using arbitration units called Resource Managers (RMs), which allows a global and work-conserving scheduling. We present a formal worst-case timing analysis for the proposed mechanism and demonstrate that this solution not only exposes higher performance in simulation but, even more importantly, consistently reaches smaller formally guaranteed worst-case latencies than TDM for realistic levels of system's utilization. Our mechanism does not require modification of routers and therefore can be used together with any architecture utilizing non-blocking routers. Adam Kostrzewa, Selma Saidi, Leonardo Ecco, Rolf Ernst |
ASP-DAC | 2 |
| 2016 | Slack-based resource arbitration for real-time Networks-on-Chip
Adam Kostrzewa, Selma Saidi, Rolf Ernst |
DATE | 2 |
| 2016 | Providing formal latency guarantees for ARQ-based protocols in Networks-on-Chip
Eberle A. Rambo, Selma Saidi, Rolf Ernst |
DATE | 2 |
| 2016 | Safe and dynamic traffic rate control for networks-on-chipsabstractNetworks-on-Chip (NoCs) for real-time systems require solutions for a safe and predictable sharing of resources between transmissions with different quality-of service (QoS) requirements. In this work, we present a mechanism which allows to apply existing wormhole-switched and performance optimized NoCs in safety critical domains, without requiring complex hardware modifications. For this purpose, we introduce a global and dynamic admission control mechanism implemented in the form of an access layer, controlling the rates at which running applications can access the NoC. The mechanism allows to enforce behavioral models for different data streams as well as to dynamically adapt the rates values to the number of currently active applications. We prove this important feature using formal timing analysis. Our approach results in a higher performance and tighter guarantees while simultaneously decreasing hardware (up to 60%) and temporal overhead (up to 80%) when compared with existing solutions. Adam Kostrzewa, Sebastian Tobuschat, Rolf Ernst, Selma Saidi |
NOCS | 4 |
| 2015 | Dynamic Detection and Mitigation of DMA Races in MPSoCsabstractExplicitly managed memories have emerged as a good alternative for multicore processors design in order to reduce energy and performance costs. Memory transfers then rely on Direct Memory Access (DMA) engines which provide a hardware support for accelerating data. However, programming explicit data transfers is very challenging for developers who must manually orchestrate data movements through the memory hierarchy. This is in practice very error-prone and can easily lead to memory inconsistency. In this paper, we propose a runtime approach for monitoring DMA races. The monitor acts as a safeguard for programmers and is able to enforce at runtime a correct behavior w.r.t the semantics of the program execution. We validate the approach using traces extracted from industrial benchmarks and executed on the multiprocessor system-onchip platform STHORM. Our experiments demonstrate that the monitoring algorithm has a low overhead (less than 1.5 KB) of on-chip memory consumption and an overhead of less than 2% of additional execution time. Selma Saidi, Yliès Falcone |
DSD | 1 |
| 2015 | Dynamic Control for Mixed-Critical Networks-on-ChipabstractNetworks-on-Chip (NoCs) for future real-time systems must provide service guarantees for applications with different levels of criticality. In this work, we propose an efficient mechanism for supporting mixed-criticality which combines the global, work-conserving scheduling for the end to end guarantees with the local arbitration in routers. We introduce a dynamic control layer with a central Resource Manager (RM) synchronizing transmissions with a dedicated protocol. The proposed mechanism allows to improve over existing solutions through reducing hardware overhead compared to non-blocking routers with rate control as well as temporal overhead compared to Time-Division Multiplexing (TDM). By using formal analysis, we show that RMs provide efficient service guarantees to all synchronized applications. We validate experimentally, using benchmarks, these guarantees along with the performance of the mechanism and induced overhead. Adam Kostrzewa, Selma Saidi, Rolf Ernst |
RTSS | 2 |
| 2014 | A mixed critical memory controller using bank privatization and fixed priority schedulingabstractMixed critical platforms are those in which applications that have different criticalities, i.e. different levels of importance for system safety, coexist and share resources. Such platforms require a memory controller capable of providing sufficient timing independence for critical applications. Existing real-time memory controllers, however, either do not support mixed criticality or still allow a certain degree of interference between applications. The former issue leads to overly constrained, and hence more expensive, systems. The latter issue forces designers to assume the worst case latency for every individual memory transaction, which can be very conservative when applied to determine the worst-case execution time (WCET) of a task that performs many memory requests. In this paper, we address both issues. The main contributions are: (1) A memory controller that allows a predetermined number of critical and non-critical applications to coexist, while providing an interference-free memory for the former. To achieve that, we treat the memory as a set of independent virtual devices (VDs). Therefore, we also provide (2) a partitioning strategy to properly map mixed critical workloads to VDs. We present experiments that show that our controller allows DRAM sharing with no interference on critical applications and minimal performance overhead on non-critical ones (they perform on average only 15% slower in the shared environment). Leonardo Ecco, Sebastian Tobuschat, Selma Saidi, Rolf Ernst |
RTCSA | 3 |
| 2012 | Optimal 2D Data Partitioning for DMA Transfers on MPSoCsabstractReducing the effects of off-chip memory access latency is a key factor in exploiting efficiently embedded multicore platforms. We consider architectures that admit a multi-core computation fabric, having its own fast and small memory to which the data blocks to be processed are fetched from external memory using a DMA (direct memory access) engine, employing a double- or multiple-buffering scheme to avoid processor idling. In this paper we focus on application programs that process two dimensional data arrays and we determine automatically the size and shape of the portions of the data array which are subject to a single DMA call, based on hardware and applications parameters. When the computation on different array elements are completely independent, the asymmetry of memory structure leads always to prefer one-dimensional horizontal pieces of memory, while when the computation of a data element shares some data with its neighbors, there is a pressure for more "square" shapes to reduce the amount of redundant data transfers. We provide an analytic model for this optimization problem and validate our results by running a mean filter application on the CELL simulator. Selma Saidi, Pranav Tendulkar, Thierry Lepley, Oded Maler |
DSD | 1 |
| 2012 | Optimizing explicit data transfers for data parallel applications on the cell architectureabstractIn this paper we investigate a general approach to automate some deployment decisions for a certain class of applications on multi-core computers. We consider data-parallelizable programs that use the well-known double buffering technique to bring the data from the off-chip slow memory to the local memory of the cores via a DMA (direct memory access) mechanism. Based on the computation time and size of elementary data items as well as DMA characteristics, we derive optimal and near optimal values for the number of blocks that should be clustered in a single DMA command. We then extend the results to the case where a computation for one data item needs some data in its neighborhood. In this setting we characterize the performance of several alternative mechanisms for data sharing. Our models are validated experimentally using a cycle-accurate simulator of the Cell Broadband Engine architecture. Selma Saidi, Pranav Tendulkar, Thierry Lepley, Oded Maler |
ACM Trans. Archit. Code Optim. | 1 |