VLDB 2026 Research / reviewers in the wild / expert
Moris Behnam
dblp:91/459
· DBLP profile ↗
102ranked-venue papers
14as first author
17since 2021 · last 2026
0000-0002-1687-930XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 62 · 6 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 5 first-author · 2 since 2021Software engineering, systems software and programming languages · 10 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Digital twins for essential servicesabstractDigital twins, dynamic digital representations of physical systems, are emerging as transformative tools for enhancing crisis preparedness and resilience in critical societal sectors. By enabling real-time monitoring, simulation, and optimization, these technologies offer actionable insights to support proactive risk mitigation, efficient resource allocation, and continuous improvement of crisis response strategies. This study provides a comprehensive knowledge overview of digital twins, focusing on their applicability and impact in key sectors such as energy, healthcare, and transportation. Specifically, it examines the essential services most suited for digital twin adoption, the role of safety-critical data throughout their life-cycle, and their utility in identifying and mitigating risks within critical infrastructure. We employed a mixed-methods research design, combining systematic and gray literature reviews with expert interviews to integrate academic insights with practical perspectives. The findings reveal significant opportunities for digital twins to enhance operational efficiency, strategic planning, and crisis management. However, practical implementation remains in its infancy, with challenges related to cost, complexity, and limited real-world applications. In addition, this study provides actionable recommendations for stakeholders, emphasizing investment in digital twin technologies, robust data governance, and the development of standardized protocols. Future research directions include exploring applications of DTs in emerging sectors, such as crisis preparedness and societal resilience, advancing artificial intelligence integration, and adopting a system-of-systems perspective to address societal challenges comprehensively. Alessio Bucaioni, Jakob Axelsson, Moris Behnam, Enxhi Ferko |
Future Gener. Comput. Syst. | 3 |
| 2026 | From engineering models to digital twins: Generating AAS from SysML v2 modelsabstractContext: Digital twins serve as virtual representations of systems, enabling capabilities such as intelligent monitoring, real-time control, decision-making, and predictive analytics. The Asset Administration Shell (AAS) is the pivotal Industry 4.0 standard for digital twin engineering. In parallel, the Systems Modeling Language (SysML) has emerged as a modeling standard for systems engineering, providing a formalized and semantically rich approach to system modeling. SysML v2 is its recent evolution. With its growing adoption, multiple models are expected to be widely available, each capturing different facets of the modeled system by leveraging diverse engineering capabilities offered by various tool ecosystems. Problem: Instead of manually re-creating models for digital twinning, existing system models should be leveraged to relieve repetitive modeling tasks. While SysML v2 and AAS are prominent standards in DT engineering, they lack direct integration, necessitating a dedicated approach for their seamless interoperability. Purpose: This paper presents a practical investigation into the conceptual alignment between the SysML v2 and AAS specifications, with a focus on their structural and behavioral modeling aspects. It proposes an implementable approach for mapping SysML v2 to AAS, enabling the automated generation of AAS models from SysML v2 models. Method: To realize this approach, we employ model-driven engineering techniques leveraging the Eclipse Modeling Framework (EMF) and model transformations based on the Query View Transformation (QVT) language. The proposed model transformation incorporates query mechanisms for extracting structured elements, preserving information and structural integrity, and ensuring static semantic consistency at design-time and seamless integration between the two investigated standards. We develop and validate the model transformation following an iterative test-driven development approach using an existing set of 24 SysML v2 examples, sourced from the official SysML v2 repository. Result: We deliver a QVT-based, EMF-compliant transformation that automatically generates AAS submodel templates from SysML v2 models, preserving structural hierarchies and behavioral semantics via dedicated AAS concepts and their extension. Through an iterative, test-driven development process, we validate metamodel conformance, information preservation, and structural integrity. The current mapping addresses design-time concepts, and the implementation supports forward transformation. All conceptual mappings, QVT scripts, and example artifacts are publicly available in a dedicated repository. Enxhi Ferko, Luca Berardinelli, Alessio Bucaioni, Moris Behnam, Manuel Wimmer |
J. Syst. Softw. | 4 |
| 2025 | Replication-Driven Resource Sharing in Real-Time Multicore SystemsabstractResource sharing in multicore architectures remains a significant challenge that limits the performance potential of such systems, especially in industrial domains requiring high performance and strict timing guarantees. In this paper, we introduce a novel synchronization protocol called Multi-Replicas Time-Bounded Consistency (MTC), designed to mitigate the adverse effects of resource sharing among tasks distributed across different cores. MTC achieves this by replicating shared resources while maintaining a bounded level of consistency among replicas within a defined time constraint. We present a response-time analysis of the proposed protocol and demonstrate that MTC can substantially simplify the evaluation of response-time analysis compared to traditional lock-based synchronization methods, offering a potentially more efficient and scalable alternative. Moris Behnam |
ETFA | 1 |
| 2025 | Generative AI Multi-Agent System with Retrieval Augmented Generation for Real-Time Furnace Operation Support in Industrial ManufacturingabstractThis paper presents a prototype AI-powered decision support system that leverages large language models (LLMs) and digital agents to assist furnace-stage operations in transformer manufacturing. The system integrates real furnace sensor data with synthetic manuals and expert insights using a modular architecture that combines Retrieval-Augmented Generation (RAG), structured prompting, multi-agent coordination, and response validation. It processes natural language queries to generate context-aware responses grounded in process data and documentation, demonstrating the potential of domain-adapted generative AI for industrial support. Experiments show a 100% improvement over a baseline LLM, though performance remains 29% below ChatGPT-4o, indicating both promise and areas for future improvement. Irini Provatidis, Victor Cobilean, Kristian Sandström, Moris Behnam |
IECON | 4 |
| 2024 | Predicting Cache Behaviour of Concurrent ApplicationsabstractModern digital solutions are built around a variety of applications. The continuous integration of these applications brings advancements in technology. Therefore, it is essential to understand how these applications will behave when they run together. However, this can be challenging to interpret due to the increasing complexity of the execution details. One such fundamental detail is the utilization of shared cache as it goes hand in hand with the computation capacity of computer systems. Since cache utilization behavior is not simple enough to translate with few assumptions we have investigated if this complex behavior can be predicted with the help of machine learning. We trained the deep neural network with enough examples that represent the cache behavior when applications were running alone and when they were running concurrently on the same core. The Long Short-Term Memory (LSTM) network learns the entire execution period of each application in the training set. As a result, without running two applications together in reality, provided with the L1 cache misses of two applications (running alone), it can predict how the cache will look like if two applications wish to run together. The model returns a time series that reflects the cache behavior in concurrency. Shamoona Imtiaz, Moris Behnam, Gabriele Capannini, Jan Carlson, Marcus Jägemar |
ETFA | 2 |
| 2024 | Hierarchical Resource Orchestration Framework for Real-time ContainersabstractContainer-based virtualization is a promising deployment model in fog and edge computing applications, because it allows a seamless co-existence of virtualized applications in a heterogeneous environment without introducing significant overhead. Certain application domains (e.g., industrial automation, automotive, or aerospace) mandate that applications exhibit a certain degree of temporal predictability. Container-based virtualization cannot be easily used for such applications, since the technology is not designed to support real-time properties and handle temporal disturbances. This article proposes a framework consisting of a static offline and a dynamic online phase for resource allocation and adaptive re-dimensioning of real-time containers. In the offline phase, the optimal initial deployment and dimensioning of containers are decided based on ideal system models. Additionally, to adapt to dynamic variations caused by changing workloads or interferences, the online phase adapts the CPU usage and limits of real-time containers at runtime to improve the real-time behavior of the real-time containerized applications while optimizing resource usage. We implement the framework in a real Linux-based system and show through a series of experiments that the proposed framework is able to adjust and re-distribute computing resources between containers to improve the real-time behavior of containerized applications in the presence of temporal disturbances while optimizing resource usage. Václav Struhár, Silviu S. Craciunas, Mohammad Ashjaei, Moris Behnam, Alessandro Vittorio Papadopoulos |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2023 | Analysing Interoperability in Digital Twin Software Architectures for Manufacturing
Enxhi Ferko, Alessio Bucaioni, Patrizio Pelliccione, Moris Behnam |
ECSA | 4 |
| 2023 | Automatic Clustering of Performance EventsabstractModern hardware and software are becoming increasingly complex due to advancements in digital and smart solutions. This is why industrial systems seek efficient use of resources to confront the challenges caused by the complex resource utilization demand. The demand and utilization of different resources show the particular execution behavior of the applications. One way to get this information is by monitoring performance events and understanding the relationship among them. However, manual analysis of this huge data is tedious and requires experts’ knowledge. This paper focuses on automatically identifying the relationship between different performance events. Therefore, we analyze the data coming from the performance events and identify the points where their behavior changes. Two events are considered related if their values are changing at "approximately" the same time. We have used the Sigmoid function to compute a real-value similarity between two sets (representing two events). The resultant value of similarity is induced as a similarity or distance metric in a traditional clustering algorithm. The proposed solution is applied to 6 different software applications that are widely used in industrial systems to show how different setups including the selection of cost functions can affect the results. Shamoona Imtiaz, Gabriele Capannini, Jan Carlson, Moris Behnam, Marcus Jägemar |
ETFA | 4 |
| 2023 | Resource Adaptation for Real-Time Containers Considering Quality of ControlabstractContainer-based virtualization has become a promising deployment model for industrial applications mainly due to its benefits, such as providing support for co-located applications in heterogeneous environments. However, such facilitation brings challenges, including full temporal isolation among real-time applications and support for time-critical applications. In this paper, we tackle such challenges, in particular when the applications are time-sensitive Control Applications. The literature suggests that flexible timing constraints for Control Applications are beneficial in responding to disturbances and minimizing response deviation. Therefore, we propose a mechanism to support such a runtime adaptation in container-based virtualization. To show the performance of the proposed mechanism, we implement our approach on a Linux-based hierarchical scheduling platform, and we evaluate it for a Control application. Václav Struhár, Mohammad Ashjaei, Moris Behnam, Alessandro Vittorio Papadopoulos, Silviu S. Craciunas |
ETFA | 3 |
| 2023 | Standardisation in Digital Twin Architectures in ManufacturingabstractEngineering digital twins following standardised reference architectures is an upcoming requirement for ensuring their adoption and facilitating their creation, processing, and integration. The ISO 23247 standard proposes a reference architecture for digital twins in manufacturing, including an entity-based reference model and a functional view specified in terms of functional entities. During our experience with projects in the field, we noticed that standards, and in particular the ISO 23247 standard, are not completely followed. In this paper, we analyse to what extent digital twin architectures documented in the literature are aligned with the reference architecture presented in the ISO 23247 standard. We achieved this through a mixed-methods research methodology that includes the analysis of 29 digital twin architectures in the manufacturing domain resulting from a systematic literature review of 140 peer-reviewed studies, a survey with 33 respondents, and four semi-structured, in-depth expert interviews. On the basis of our findings, practitioners and researchers can reflect, discuss, and plan actions for future research and development activities. Enxhi Ferko, Alessio Bucaioni, Patrizio Pelliccione, Moris Behnam |
ICSA | 4 |
| 2021 | LLM-shark - A Tool for Automatic Resource-boundness Analysis and Cache Partitioning SetupabstractWe present LLM-shark, a tool for automatic hardware resource-boundness detection and cache-partitioning. Our tool has three primary objectives: First, it determines the hardware resource-boundness of a given application. Secondly, it estimates the initial cache partition size to ensure that the application performance is conserved and not affected by other processes competing for cache utilization. Thirdly, it continuously monitors that the application performance is maintained over time and, if necessary, change the cache partition size. We demonstrate LLM-shark’s functionality through a series of tests using six different applications, including a set of feature detection algorithms and two synthetic applications. Our tests reveal that it is possible to determine an application’s resource-boundness using a Pearson-correlation scheme implemented in LLM-shark. We propose a scheme to size cache partitions based on the correlation coefficient applications depending on their resource boundness. Jakob Danielsson, Tiberiu Seceleanu, Marcus Jägemar, Moris Behnam, Mikael Sjödin |
COMPSAC | 4 |
| 2021 | Modelling Application Cache Behavior using Regression ModelsabstractIn this paper, we describe the creation of resource usage forecasts for applications with unknown execution characteristics, by evaluating different regression processes, including autoregressive, multivariate adaptive regression splines, exponential smoothing, etc. We utilize Performance Monitor Units (PMU) and generate hardware resource usage models for the L2-cache and the L3-cache using nine different regression processes. The measurement strategy and regression process methodology are general and applicable to any given hardware resource when performance counters are available. We use three benchmark applications: the SIFT feature detection algorithm, a standard matrix multiplication, and a version of Bubblesort. Our evaluation shows that Multi Adaptive Regressive Spline (MARS) models generate the best resource usage forecasts among the considered models, followed by Single Exponential Splines (SES) and Triple Exponential Splines (TES). Jakob Danielsson, Janne Suuronen, Marcus Jägemar, Tiberiu Seceleanu, Moris Behnam, Mikael Sjödin |
COMPSAC | 5 |
| 2021 | Automatic Quality of Service Control in Multi-core Systems using Cache PartitioningabstractIn this paper, we present a last-level cache partitioning controller for multi-core systems. Our objective is to control the Quality of Service (QoS) of applications in multi-core systems by monitoring run-time performance and continuously re-sizing cache partition sizes according to the applications' needs. We discuss two different use-cases; one that promotes application fairness and another one that prioritizes applications according to the system engineers' desired execution behavior. We display the performance drawbacks of maintaining a fair schedule for all system tasks and its performance implications for system applications. We, therefore, implement a second control algorithm that enforces cache partition assignments according to user-defined priorities rather than system fairness. Our experiments reveal that it is possible, with non-instrusive (0.3-0.7% CPU utilization) cache controlling measures, to increase performance according to setpoints and maintain the QoS for specific applications in an over-saturated system. Jakob Danielsson, Tiberiu Seceleanu, Marcus Jägemar, Moris Behnam, Mikael Sjödin |
ETFA | 4 |
| 2021 | Automatic Platform-Independent Monitoring and Ranking of Hardware Resource UtilizationabstractIn this paper, we discuss a method for automatic monitoring of hardware and software events using performance monitoring counters. Computer applications are complex and utilize a broad spectra of the available hardware resources, where multiple performance counters can be of significant interest to understand. The number of performance counters that can be captured simultaneously is, however, small due to hardware limitations in most modern computers. We suggest a platform independent solution to automatically retrieve hardware events from an underlying architecture. Moreover, to mitigate the hardware limitations we propose a mechanism that pinpoints the most relevant performance counters for an application's performance. In our proposal, we utilize the Pearson's correlation coefficient to rank the most relevant performance counters and filter out those that are most relevant and ignore the rest. Shamoona Imtiaz, Jakob Danielsson, Moris Behnam, Gabriele Capannini, Jan Carlson, Marcus Jägemar |
ETFA | 3 |
| 2021 | REACT: Enabling Real-Time Container OrchestrationabstractFog and edge computing offer the flexibility and decentralized architecture benefits of cloud computing without suffering from the latency issues inherent in the cloud. This makes fog computing very attractive in real-time and safety-critical applications, especially if combined with container-based technologies. Whereas different orchestration systems are available to manage the container placement based on their resource demand, no orchestration system is considering real-time requirements for containerized applications. In this paper, we present the architecture and design of a real-time container orchestrator based on Kubernetes. Moreover, this paper defines metrics for the performance evaluation of real-time containers, and describes an initial model for allocating a mixture of real-time and non-real-time containers. We present an initial implementation of our real-time container extension and evaluate its feasibility on Linux-based systems. Václav Struhár, Silviu S. Craciunas, Mohammad Ashjaei, Moris Behnam, Alessandro Vittorio Papadopoulos |
ETFA | 4 |
| 2021 | Offloading Accelerator-intensive Workloads in CPU-GPU Heterogeneous ProcessorsabstractAutonomous vehicular systems require computer vision and intelligent on-board decision making functionalities that include a mix of sequential and parallel workloads. The execution times of the workloads and power consumption in these functionalities can be lowered by utilizing the accelerators (e.g., GPU) instead of running the workloads entirely on the host processing units (CPU). However, allocating all the parallelizable workload to accelerators can create a computation bottleneck in the accelerators that, in turn, can have an adverse effect on schedulability of the systems. This paper presents a novel framework that can allocate the accelerate-intensive workloads to the accelerators as well as to the non-accelerated host processing units. Within the context of this framework, the paper introduces five offloading techniques to mitigate the accelerator-intensive workloads by utilizing excess capacity of non-accelerated processing units under dynamic scheduling in CPU-GPU heterogeneous processors. The proposed techniques are evaluated using simulation experiments. The evaluation results indicate that one of the proposed techniques can achieve up to 16% improvement in schedulability of the task sets compared to the traditional non-offloading technique. Nandinbaatar Tsog, Saad Mubeen, Fredrik Bruhn, Moris Behnam, Mikael Sjödin |
ETFA | 4 |
| 2021 | Selected papers presented at the 26th International Conference on Real-Time and Network Systems (RTNS 2018)
Moris Behnam, Mathieu Jan |
Real Time Syst. | 1 |
| 2020 | Resource Depedency Analysis in Multi-Core SystemsabstractIn this paper, we evaluate different methods for statistical determination of application resource dependency in multi-core systems. We measure the performance counters of an application during run-time and create a system resource usage profile. We then use the resource profile to evaluate the application dependency on the specific resource. We discuss and evaluate two methods to process the data, including moving average filter and partitioning the data into smaller segments in order to interpret data for correlation calculations. Our aim with this study is to evaluate and create a generalizeable methods for automatic determination of resource dependencies. The final outcome of the methods used in this study is the answer to the question: "To what resources is this application dependent on?". The recommendation of this tool will be used in conjunction with our last-level cache partitioning controller (LLC-PC), to make decision if an application should receive last-level cache partition slices. Jakob Danielsson, Tiberiu Seceleanu, Marcus Jägemar, Moris Behnam, Mikael Sjödin |
COMPSAC | 4 |
| 2020 | Schedulability analysis of Time-Sensitive Networks with scheduled traffic and preemption support
Lucia Lo Bello, Mohammad Ashjaei, Gaetano Patti, Moris Behnam |
J. Parallel Distributed Comput. | 4 |
| 2019 | Testing Performance-Isolation in Multi-core SystemsabstractIn this paper we present a methodology to be used for quantifying the level of performance isolation for a multi-core system. We have devised a test that can be applied to breaches of isolation in different computing resources that may be shared between different cores. We use this test to determine the level of isolation gained by using the Jailhouse hypervisor compared to a regular Linux system in terms of CPU isolation, cache isolation and memory bus isolation. Our measurements show that the Jailhouse hypervisor provides performance isolation of local computing resources such as CPU. We have also evaluated if any isolation could be gained for shared computing resources such as the system wide cache and the memory bus controller. Our tests show no measurable difference in partitioning between a regular Linux system and a Jailhouse partitioned system for shared resources. Using the Jailhouse hypervisor provides only a small noticeable overhead when executing multiple shared-resource intensive tasks on multiple cores, which implies that running Jailhouse in a memory saturated system will not be harmful. However, contention still exist in the memory bus and in the system-wide cache. Jakob Danielsson, Tiberiu Seceleanu, Marcus Jägemar, Moris Behnam, Mikael Sjödin |
COMPSAC (1) | 4 |
| 2019 | Run-Time Cache-Partition Controller for Multi-Core SystemsabstractThe current trend in automotive systems is to integrate more software applications into fewer ECU's to decrease the cost and increase efficiency. This means more applications share the same resources which in turn can cause congestion on resources such as such as caches. Shared resource congestion may cause problems for time critical applications due to unpredictable interference among applications. It is possible to reduce the effects of shared resource congestion using cache partitioning techniques, which assign dedicated cache lines to different applications. We propose a cache partition controller called LLC-PC that uses the Palloc page coloring framework to decrease the cache partition sizes for applications during runtime. LLC-PC creates cache partitioning directives for the Palloc tool by evaluating the performance gained from increasing the cache partition size. We have evaluated LLC-PC using 3 different applications, including the SIFT image processing algorithm which is commonly used for feature detection in vision systems. We show that LLC-PC is able to decrease the amount of cache size allocated to applications while maintaining their performance allowing more cache space to be allocated for other applications. Jakob Danielsson, Marcus Jägemar, Moris Behnam, Tiberiu Seceleanu, Mikael Sjödin |
IECON | 3 |
| 2019 | DART: Dynamic Bandwidth Distribution Framework for Virtualized Software Defined NetworksabstractIn this paper we address a network architecture that uses a combination of network virtualization and software defined networking in order to reduce complexity of network management and at the same time support high quality of service. Within this network architecture, we propose a framework to be able to dynamically distribute the network bandwidth to various services such that the network resources are utilized efficiently. In many industrial domains, multiple services may use the same hardware platform for the sake of a better resource utilization. Therefore, bandwidth distribution among the services should be done in an efficient way during runtime. We also develop an admission control in this framework which dynamically coordinates the bandwidth distributions based on requested quality of services. We show the applicability of the proposed framework by implementing it on a common SDN controller. Moreover, we conduct a set of experiments to show the performance of the proposed framework. Václav Struhár, Mohammad Ashjaei, Moris Behnam, Silviu S. Craciunas, Alessandro Vittorio Papadopoulos |
IECON | 3 |
| 2019 | Static Allocation of Parallel Tasks to Improve Schedulability in CPU-GPU Heterogeneous Real-Time SystemsabstractAutonomous driving is one of the main challenges of modern cars. Computer visions and intelligent on-board decision making are crucial in autonomous driving and require heterogeneous processors with high computing capability under low power consumption constraints. The progress of parallel computing using heterogeneous processing units is further supported by software frameworks like OpenCL, OpenMP, CUDA, and C++AMP. These frameworks allow the allocation of parallel computation on different compute resources. This, however, creates a difficulty in allocating the right computation segments to the right processing units in such a way that the complete system meets all its timing requirements. In this paper, we consider pre-runtime static allocations of parallel tasks to perform their execution either sequentially on CPU or in parallel using a GPU. This allows for improving any unbalanced use of GPU accelerators in a heterogeneous environment. By performing several heuristic algorithms, we show that the overuse of accelerators results in a bottle-neck of the entire system execution. The experimental results show that our allocation schemes that target a balanced use of GPU improves the system schedulability up to 90%. Nandinbaatar Tsog, Matthias Becker 0004, Fredrik Bruhn, Moris Behnam, Mikael Sjödin |
IECON | 4 |
| 2018 | Scheduling multi-rate real-time applications on clustered many-core architectures with memory constraintsabstractAccess to shared memory is one of the main challenges for many-core processors. One group of scheduling strategies for such platforms focuses on the division of tasks' access to shared memory and code execution. This allows to orchestrate the access to shared local and off-chip memory in a way such that access contention between different compute cores is avoided by design. In this work, an execution framework is introduced that leverages local memory by statically allocating a subset of tasks to cores. This reduces the access times to shared memory, as off-chip memory access is avoided, and in turn improves the schedulability of such systems. A Constraint Programming (CP) formulation is presented to select the statically allocated tasks and to generate the complete system schedule. Evaluations show that the proposed approach yields an up to 19% higher schedulability ratio than related work, and a case study demonstrates its applicability to industrial problems. Matthias Becker 0004, Saad Mubeen, Dakshina Dasari, Moris Behnam, Thomas Nolte |
ASP-DAC | 4 |
| 2018 | Measurement-Based Evaluation of Data-Parallelism for OpenCV Feature-Detection AlgorithmsabstractWe investigate the effects on the execution time, shared cache usage and speed-up gains when using datapartitioned parallelism for the feature detection algorithms available in the OpenCV library. We use a data set of three different images which are scaled to six different sizes to exercise the different cache memories of our test architectures. Our measurements reveal that the algorithms using the default settings of OpenCV behave very differently when using data-partitioned parallelism. Our investigation shows that the executions of the algorithms SURF, Dense and MSER correlate to L3-cache usage and they are therefore not suitable for data-partitioned parallelism on multicore CPUs. Other algorithms: BRISK, FAST, ORB, HARRIS, GFTT, SimpleBlob and SIFT, do not correlate to L3-cache in the same extent, and they are therefore more suitable for data-partitioned parallelism. Furthermore, the SIFT algorithm provides the most stable speed-up, resulting in an execution between 3 and 3.5 times faster than the original execution time for all image sizes. We also have evaluated the hardware resource usage by measuring the algorithm execution time simultaneously with the L3-cache usage. We have used our measurements to conclude which algorithms are suitable for parallelization on hardware with shared resources. Jakob Danielsson, Marcus Jägemar, Moris Behnam, Mikael Sjödin, Tiberiu Seceleanu |
COMPSAC (1) | 3 |
| 2018 | Fog computing for adaptive human-robot collaboration: work-in-progressabstractFog computing is an emerging technology that enables the design of novel time sensitive industrial applications. This new computing paradigm also opens several new research challenges in different scientific domains, ranging from computer architectures to networks, from robotics to real-time systems. In this paper, we present a use case in the human-robot collaboration domain, and we identify some of the most relevant research challenges. Václav Struhár, Alessandro Vittorio Papadopoulos, Moris Behnam |
EMSOFT | 3 |
| 2018 | Enforcing Quality of Service Through Hardware Resource Aware Process SchedulingabstractHardware manufacturers are forced to improve system performance continuously due to advanced and computationally demanding system functions. Unfortunately - more powerful hardware leads to increased costs. Instead, companies attempt to improve performance by consolidating multiple functions to share the same hardware to exploit existing performance instead. In legacy systems, each function had individual execution environment that guaranteed HW resource isolation and therefore the Quality of Service (QoS). Consolidation of multiple functions increases the risk of shared resource congestion. Current process schedulers focus on time quanta and do not consider shared resources. We present a novel process scheduler that complements current process schedulers by enforcing QoS though Shared Resource Aware (SRA) process scheduling. The SRA scheduler programs the Performance Monitoring Unit (PMU) to generate an overflow interrupt when reaching the assigned process resource quota. The scheduler has the possibility to swap out the process when receiving the interrupt allowing it to enforce the QoS for the scheduled process. We have implemented our scheduling policy as a new scheduling class in Linux. Our experiments show that it efficiently enforces QoS without seriously affect the shared resource usage of other processes executing on the same HW. Marcus Jägemar, Andreas Ermedahl, Sigrid Eldh, Moris Behnam, Björn Lisper |
ETFA | 4 |
| 2018 | Practical Challenges for FSLMabstractThe flexible spin-lock model (FSLM) unifies suspension-based and spin-based resource access protocols for partitioned fixed-priority preemptive scheduling based real-time multi-core platforms. Recent work has been done in defining the protocol for FSLM, providing schedulability analysis, and investigating the practical consequences of the theoretical model. FSLM complies to the AUTOSAR standard for the automotive industry, and prototype implementations of FSLM in the OSEK/VDX-complaint Erika Enterprise Real-Time Operating System have been realized. In this paper, we briefly describe some practical challenges to improve efficiency and generality. S. Muthu N. Balasubramanian, Sara Afshar, Paolo Gai, Moris Behnam, Reinder J. Bril |
RTCSA | 4 |
| 2017 | A tighter recursive calculus to compute the worst case traversal time of real-time traffic over NoCsabstractNetwork-on-Chip (NoC) is a communication subsystem which has been widely utilized in many-core processors and system-on-chips in general. In this paper, we focus on a Round-Robin Arbitration (RRA) based wormhole-switched NoC which is a common architecture used in most of the existing implementations. In order to execute real-time applications on such a NoC based platform, a number of given real-time requirements need to be fulfilled. One of the most typical requirements is schedulability which refers to determining if real-time packets can be delivered within the given time durations. Timing analysis is a common tool to verify the schedulability of a real-time system. Unfortunately, the existing timing analyses of RRA-based NoCs either provide too pessimistic estimates which results in overly allocated resources or require a large amount of processing which limits the applicability in reality. Therefore, in this paper, we present an improved timing analysis, aiming to provide more accurate estimates along with acceptable computation time. From the evaluation results, we can clearly observe the improvement achieved by the proposed timing analysis. Meng Liu 0001, Matthias Becker 0004, Moris Behnam, Thomas Nolte |
ASP-DAC | 3 |
| 2017 | Using segmentation to improve schedulability of RRA-based NoCs with mixed trafficabstractNetwork-on-Chip (NoC) is the interconnect of choice for many-core processors and system-on-chips in general. Most of the existing NoC designs focus on the performance with respect to average throughput, which makes them less applicable for real-time applications especially when applications have hard timing requirements on the worst-case scenarios. In this paper, we focus on a Round-Robin Arbitration (RRA) based wormhole-switched NoC which is a common architecture used in most of the existing implementations. We propose a novel segmentation algorithm targeting RRA-based NoCs in order to improve the schedulability of real-time traffic without modifying the hardware architecture. Additionally, we also address the problem of transmitting both real-time traffic and best-effort traffic in the same NoC. The proposed solutions aim to provide timing guarantees to real-time traffic and achieve low latency for best-effort traffic. According to the evaluation results, the proposed segmentation solution can significantly improve the schedulability of the whole network. Meng Liu 0001, Matthias Becker 0004, Moris Behnam, Thomas Nolte |
ASP-DAC | 3 |
| 2017 | Performance evaluation of network convergence time measurement techniquesabstractIn this paper we evaluate solutions that provide measurements for the network convergence time in switched Ethernet networks when links failures happen. We evaluate three solutions to measure the network convergence time in a faulty situation. Compared to the commercially available solutions, our proposals are cost-effective, portable, and open source. Thus, they are easy to deploy on many testbeds. We show the performance of the solutions by measuring different metrics including jitter, network convergence time and packet loss during the network recovery time. Our measurements indicate that it is possible to accurately measure the network convergence time using a packet sender which does not suffer interference from the overlying operating system. Furthermore, we noticed that the packet sniffer TShark did not suffer from kernel interrupts from overlying operating systems. Jakob Danielsson, Mohammad Ashjaei, Moris Behnam, Thomas Sorensen, Mikael Sjödin, Thomas Nolte |
ETFA | 3 |
| 2017 | Investigating execution-characteristics of feature-detection algorithmsabstractWe discuss how to obtain information of execution characteristics, such as parallelizability and memory utilization, with the final aim to improve the performance and predictability of feature and corner detection algorithms for use in e.g. robotics and autonomous machines. Our aim is to obtain a better understanding of how computer vision algorithms use hardware resources and how to improve the time predictability and execution time of such algorithms when executing on multi-core CPUs. We evaluate a fork-join model applicable to feature detection algorithms and present a method for measuring how well the algorithm performance correlates with hardware resource usage. We have applied our method to the Featured from Accelerated Segment Test (FAST) algorithm. Our characterization of FAST reveals that it is an algorithm with excellent parallelism opportunities, resulting in an almost linear speed-up per core. Our measurements also reveal that the performance of FAST correlates very little with the number of misses in the L1 data cache, L1 instruction cache, data translation lookaside buffer and L2 cache. Thus, the FAST algorithm will not have a negative effect on the execution time when the input data fits in the L2 cache. Jakob Danielsson, Marcus Jägemar, Moris Behnam, Mikael Sjödin |
ETFA | 3 |
| 2017 | A scheduling architecture for enforcing quality of service in multi-process systemsabstractThere is a massive deployment of multi-core CPUs. It requires a significant drive to consolidate multiple services while still achieving high performance on these off-the-shelf CPUs. Each function had earlier an own execution environment, which guaranteed a certain Quality of Service (QoS). Consolidating multiple services can give rise to shared resource congestions, resulting in lower and non-deterministic QoS. We describe a method to increase the overall system performance by assisting the operating system process scheduler to utilize shared resources more efficiently. Our method utilizes hardware- and system-level performance counters to profile the shared resource usage of each process. We also use a big-data approach to analyzing statistics from many nodes. The outcome of the analysis is a decision support model that is utilized by the process scheduler when allocating and scheduling process. Our scheduler can efficiently distribute processes compared to traditional CPU-load based process schedulers by considering the hardware capacity and previous scheduling- and allocation decisions. Marcus Jägemar, Andreas Ermedahl, Sigrid Eldh, Moris Behnam |
ETFA | 4 |
| 2017 | Incorporating implementation overheads in the analysis for the flexible spin-lock modelabstractThe flexible spin-lock model (FSLM) unifies suspension-based and spin-based resource sharing protocols for partitioned fixed-priority preemptive scheduling based real-time multiprocessor platforms. Recent work has been done in defining the protocol for FSLM and providing a schedulability analysis without accounting for the implementation overheads. In this paper, we extend the analysis for FSLM with implementation overheads. Utilizing an initial implementation of FSLM in the OSEK/VDX-compliant Erika Enterprise RTOS on an Altera Nios II platform using 4 soft-core processors, we present an improved implementation. Given the design of the implementation, the overheads are characterized and incorporated in specific terms of the existing analysis. The paper also supplements the analysis with measurement results, enabling an analytical comparison of FSLM with the natively provided multiprocessor stack resource policy (MSRP), which may serve as a guideline for the choice of FSLM or MSRP for a specific application. S. Muthu N. Balasubramanian, Sara Afshar, Paolo Gai, Moris Behnam, Reinder J. Bril |
IECON | 4 |
| 2017 | Buffer-Aware Analysis for Worst-Case Traversal Time of Real-Time Traffic over RRA-based NoCsabstractNetwork-on-Chip (NoC) is a communication subsystem which has been widely utilized in many-core processors and system-on-chips in general. In order to execute time-critical applications on a NoC-based platform, the timing behavior of the network needs to be predicted during system design. One of the most important timing requirements is regarding schedulability, which refers to determining if a real-time packet can be delivered within a specific time duration. To verify the fulfillment of such timing requirement, a proper timing analysis is mandatory. Our work focuses on a Round-Robin Arbitration (RRA) based wormhole-switched NoC, which is a common architecture used in many of the existing implementations. Recursive Calculus (RC) is one of the existing analysis approaches for RRA-based NoCs which has been utilized in many research works. However, RC does not take buffer-effects into account. As a result, while performing RC on most of the existing RRA-based NoC designs, it can produce unsafe estimates which is not acceptable for time-critical systems. In this paper, we identify the optimistic problem of RC, and we propose a Revised Recursive Calculus (RRC) which extends RC by considering buffer-effects as well as supporting packetization. Meng Liu 0001, Matthias Becker 0004, Moris Behnam, Thomas Nolte |
PDP | 3 |
| 2017 | Partitioning and Analysis of the Network-on-Chip on a COTS Many-Core PlatformabstractMany-core processors can provide the computational power required by future complex embedded systems. However, their adoption is not trivial, since several sources of interference on COTS many-core platforms have adverse effects on the resulting performance. One main source of performance degradation is the contention on the Network-on-Chip (NoC), which is used for communication among the compute cores via the off-chip memory. Available analysis techniques for the traversal time of messages on the NoC do not consider many of the architectural features found on COTS platforms. In this work, we target a state-of-the-art many-core processor, the Kalray MPPAR®. A novel partitioning strategy for reducing the contention on the NoC is proposed. Further, we present an analysis technique dedicated to the proposed partitioning strategy, which considers all architectural features of the COTS NoC. Additionally, it is shown how to configure the parameters for flow-regulation on the NoC, such that the Worst-Case Traversal Time (WCTT) is minimal and buffers never overflow. The benefits of our approach are evaluated based on extensive experiments that show that contention is significantly reduced compared to the unconstrained case, while the proposed analysis outperforms a state-of-the-art analysis for the same platform. An industrial case study shows the tightness of the proposed analysis. Matthias Becker 0004, Borislav Nikolic, Dakshina Dasari, Benny Akesson, Vincent Nélis, Moris Behnam, Thomas Nolte |
RTAS | 6 |
| 2017 | A generic framework facilitating early analysis of data propagation delays in multi-rate systems (Invited paper)abstractA majority of multi-rate real-time systems are constrained by a multitude of timing requirements, in addition to the traditional deadlines on well-studied response times. This means, the timing predictability of these systems not only depends on the schedulability of certain task sets but also on the timely propagation of data through the chains of tasks from sensors to actuators. In the automotive industry, four different timing constraints corresponding to various data propagation delays are commonly specified on the systems. This paper identifies and addresses the source of pessimism as well as optimism in the calculations for one such delay, namely the reaction delay, in the state-of-the-art analysis that is already implemented in several industrial tools. Furthermore, a generic framework is proposed to compute all the four end-to-end data propagation delays, complying with the established delay semantics, in a scheduler and hardware-agnostic manner. This allows analysis of the system models already at early development phases, where limited system information is present. The paper further introduces mechanisms to generate job-level dependencies, a partial ordering of jobs, which need to be satisfied by any execution platform in order to meet the data propagation timing requirements. The job-level dependencies are first added to all task chains of the system and then reduced to its minimum required set such that the job order is not affected. Moreover, a necessary schedulability test is provided, allowing for varying the number of CPUs. The experimental evaluations demonstrate the tightness in the reaction delay with the proposed framework as compared to the existing state-of-the-art and practice solutions. Matthias Becker 0004, Saad Mubeen, Dakshina Dasari, Moris Behnam, Thomas Nolte |
RTCSA | 4 |
| 2017 | An optimal spin-lock priority assignment algorithm for real-time multi-core systemsabstractSupport for exclusive access to shared (global) resources is instrumental in the context of embedded real-time multi-core systems, and mechanisms for achieving such access must be deterministic and efficient. There exist two traditional approaches for multiprocessors when a task requests a global resource that is locked by a task on a remote core: a spin-based approach, i.e. non-preemptive busy waiting for the resource to become available, and a suspension-based approach, i.e. the task relinquishes the processor. A suspension-based approach can be viewed as a spin-based approach where the lowest priority on a core is used during spinning, similar to a non-preemptive spin-based approach where the highest priority on a core is used. By taking such a view, we previously provided a general model for spinning, where any arbitrary priority can be used for spinning, i.e. from the lowest to the highest priority on a core. Targeting partitioned fixed-priority preemptive scheduled multiprocessors and spin-based approaches that use a fixed priority for spinning per core for all tasks, we aim at increasing the schedulability of multiprocessor systems by using the spin-lock priority per core as parameter. In this paper, we present (i) a generalization of the traditional worst-case response-time analysis for non-preemptive spin-based approaches addressing an arbitrary but fixed spin-lock priority per core, (ii) an optimal spin-lock priority assignment (OSPA) algorithm per core, i.e. an algorithm that will find a fixed spin-lock priority per core that will make the system schedulable, whenever such an assignment exists and, (iii) comparative evaluations of the OSPA algorithm with the spin-based and suspension-based approaches where OSPA showed up to 38% improvement compared to both approaches. Sara Afshar, Moris Behnam, Reinder J. Bril, Thomas Nolte |
RTCSA | 2 |
| 2017 | End-to-end timing analysis of cause-effect chains in automotive embedded systemsabstractAutomotive embedded systems are subjected to stringent timing requirements that need to be verified. One of the most complex timing requirement in these systems is the data age constraint. This constraint is specified on cause-effect chains and restricts the maximum time for the propagation of data through the chain. Tasks in a cause-effect chain can have different activation patterns and different periods, that introduce over- and under-sampling effects, which additionally aggravate the end-to-end timing analysis of the chain. Furthermore, the level of timing information available at various development stages (from modeling of the software architecture to the software implementation) varies a lot, the complete timing information is available only at the implementation stage. This uncertainty and limited timing information can restrict the end-to-end timing analysis of these chains. In this paper, we present methods to compute end-to-end delays based on different levels of system information. The characteristics of different communication semantics are further taken into account, thereby enabling timing analysis throughout the development process of such heterogeneous software systems. The presented methods are evaluated with extensive experiments. As a proof of concept, an industrial case study demonstrates the applicability of the proposed methods following a state-of-the-practice development process. Matthias Becker 0004, Dakshina Dasari, Saad Mubeen, Moris Behnam, Thomas Nolte |
J. Syst. Archit. | 4 |
| 2017 | Designing end-to-end resource reservations in predictable distributed embedded systemsabstractContemporary distributed embedded systems in many domains have become highly complex due to ever-increasing demand on advanced computer controlled functionality. The resource reservation techniques can be effective in lowering the software complexity, ensuring predictability and allowing flexibility during the development and execution of these systems. This paper proposes a novel end-to-end resource reservation model for distributed embedded systems. In order to support the development of predictable systems using the proposed model, the paper provides a method to design resource reservations and an end-to-end timing analysis. The reservation design can be subjected to different optimization criteria with respect to runtime footprint, overhead or performance. The paper also presents and evaluates a case study to show the usability of the proposed model, reservation design method and end-to-end timing analysis. Mohammad Ashjaei, Nima Khalilzad, Saad Mubeen, Moris Behnam, Ingo Sander, Luís Almeida 0001, Thomas Nolte |
Real Time Syst. | 4 |
| 2017 | Schedulability analysis of Ethernet Audio Video Bridging networks with scheduled traffic supportabstractThe IEEE Audio Video Bridging (AVB) technology is nowadays under consideration in several automation domains, such as, automotive, avionics, and industrial communications. AVB offers several benefits, such as open specifications, the existence of multiple providers of electronic components, and the real-time support, as AVB provides bounded latency to real-time traffic classes. In addition to the above mentioned properties, in the automotive domain, comparing with the existing in-vehicle networks, AVB offers significant advantages in terms of high bandwidth, significant reduction of cabling costs, thickness and weight, while meeting the challenging EMC/EMI requirements. Recently, an improvement of the AVB protocol, called the AVB ST, was proposed in the literature, which allows for supporting scheduled traffic, i.e., a class of time-sensitive traffic that requires time-driven transmission and low latency. In this paper, we present a schedulability analysis for the real-time traffic crossing through the AVB ST network. In addition, we formally prove that, if the bandwidth in the network is allocated according to the AVB standard, the schedulability test based on response time analysis will fail for most cases even if, in reality, these cases are schedulable. In order to provide guarantees based on analysis test a bandwidth over-reservation is required. In this paper, we propose a solution to obtain a minimized bandwidth over-reservation. To the best of our knowledge, this is the first attempt to formally spot the limitation and to propose a solution for overcoming it. The proposed analysis is applied to both the AVB standard and the AVB ST. The analysis results are compared with the results of several simulative assessments, obtained using OMNeT++, on both automotive and industrial case studies. The comparison between the results of the analysis and the simulation ones shows the effectiveness of the analysis proposed in this work. Mohammad Ashjaei, Gaetano Patti, Moris Behnam, Thomas Nolte, Giuliana Alderisi, Lucia Lo Bello |
Real Time Syst. | 3 |
| 2017 | Fixed priority scheduling with pre-emption thresholds and cache-related pre-emption delays: integrated analysis and evaluationabstractCommercial off-the-shelf programmable platforms for real-time systems typically contain a cache to bridge the gap between the processor speed and main memory speed. Because cache-related pre-emption delays (CRPD) can have a significant influence on the computation times of tasks, CRPD have been integrated in the response time analysis for fixed-priority pre-emptive scheduling (FPPS). This paper presents CRPD aware response-time analysis of sporadic tasks with arbitrary deadlines for fixed-priority pre-emption threshold scheduling (FPTS), generalizing earlier work. The analysis is complemented by an optimal (pre-emption) threshold assignment algorithm, assuming the priorities of tasks are given. We further improve upon these results by presenting an algorithm that searches for a layout of tasks in memory that makes a task set schedulable. The paper includes an extensive comparative evaluation of the schedulability ratios of FPPS and FPTS, taking CRPD into account. The practical relevance of our work stems from FPTS support in AUTOSAR, a standardized development model for the automotive industry. [(This paper forms an extended version of Bril et al. (in Proceedings of 35th IEEE real-time systems symposium (RTSS), 2014 ). The main extensions are described in Sect. 1.2 .] Reinder J. Bril, Sebastian Altmeyer, Martijn M. H. P. van den Heuvel, Robert I. Davis 0001, Moris Behnam |
Real Time Syst. | 5 |
| 2017 | Using non-preemptive regions and path modification to improve schedulability of real-time traffic over priority-based NoCsabstractNetwork-on-Chip (NoC) is a preferred communication medium for massively parallel platforms. Fixed-priority based scheduling using virtual-channels is one of the promising solutions to support real-time traffic in on-chip networks. Most of the existing works regarding priority-based NoCs use a flit-level preemptive scheduling. Under such a mechanism, preemptions can only happen between the transmissions of successive flits but not during the transmission of a single flit. In this paper, we present a modified framework where the non-preemptive region of each NoC packet increases from a single flit. Using the proposed approach, the response times of certain traffic flows can be reduced, which can thus improve the schedulability of the whole network. As a result, the utilization of NoCs can be improved by admitting more real-time traffic. Schedulability tests regarding the proposed framework are presented along with the proof of the correctness. Additionally, we also propose a path modification approach on top of the non-preemptive region based method to further improve schedulability. A number of experiments have been performed to evaluate the proposed solutions, where we can observe significant improvement on schedulability compared to the original flit-level preemptive NoCs. Meng Liu 0001, Matthias Becker 0004, Moris Behnam, Thomas Nolte |
Real Time Syst. | 3 |
| 2017 | Guest Editorial Special Section on Communications in Automation-Innovation Drivers and New TrendsabstractThe papers in this special section focus on telecommunication services in automation applicaitons. The evolution of automation systems has always been the result of continuous improvements to state-of-the-art systems, occasionally punctuated with significant paradigm and architectural changes that acted as innovation drivers. This happened with the industrial revolutions, which had in their origin the appearance of disruptive technologies, such as the steam engine and machine tools (first industrial revolution); electrification, mass production, and assembly lines (second industrial revolution); use of microprocessors and software (third industrial revolution). Lucia Lo Bello, Moris Behnam, Paulo Pedreiras, Thilo Sauter |
IEEE Trans. Ind. Informatics | 2 |
| 2016 | Tighter time analysis for real-time traffic in on-chip networks with shared prioritiesabstractThe Network-on-Chip (NoC) is the preferred interconnection medium for massively parallel platforms. Targeting real-time applications, fixed-priority based NoCs with virtual channels have been proposed as a promising solution. In order to verify if specific time requirements can be satisfied, schedulability tests are typically used. Several analysis approaches have been proposed targeting priority-based NoCs. However, due to the approximation considered in the analyses, the results may involve a large amount of pessimism. The applicability of the analyses is thus limited in practice. In this paper, we identify a number of properties of NoCs with shared priorities. An improved time analysis is proposed where pessimism can be significantly reduced for many cases. In order to evaluate the proposed analysis, a number of experiments have been generated along with a case study based on an automotive application. The improvement can be clearly observed from the evaluation results. Meng Liu 0001, Matthias Becker 0004, Moris Behnam, Thomas Nolte |
NOCS | 3 |
| 2016 | End-to-End Resource Reservations in Distributed Embedded SystemsabstractThe resource reservation techniques provide effective means to lower the software complexity, ensure predictability and allow flexibility during the development and execution of complex distributed embedded systems. In this paper we propose a new end-to-end resource reservation model for distributed embedded systems. The model is comprehensive in such a way that it supports end-to-end resource reservations on distributed transactions with various activation patterns that are commonly used in industrial control systems. The model allows resource reservations on processors and real-time network protocols. We also present timing analysis for the distributed embedded systems that are developed using the proposed model. The timing analysis computes the end-to-end response times as well as delays such as data age and reaction delays. The presented analysis also supports real-time networks that can autonomously initiate transmissions. Such networks are not supported by the existing analyses. We also include a case study to show the usability of the model and end-to-end timing analysis with resource reservations. Mohammad Ashjaei, Saad Mubeen, Moris Behnam, Luís Almeida 0001, Thomas Nolte |
RTCSA | 3 |
| 2016 | Synthesizing Job-Level Dependencies for Automotive Multi-rate Effect ChainsabstractToday's automotive embedded systems comprise a multitude of functionalities, many with complex timing requirements. Besides task specific timing requirements, such applications often have timing requirements for the propagation of data through a chain of tasks. An important metric for control applications is the data age, which is addressed in this paper. The analysis of such systems is non-trivial because tasks involved in the data propagation may execute at different periods, which leads to over and undersampling within one chain. This paper presents a novel method to compute worst-and best-case end-to-end latencies for such systems. A second contribution synthesizes job-level dependencies for such task sets in a way that data paths which exceed the age constraint are eliminated. An extensive evaluation is performed on synthetic task sets and the applicability to industrial applications is demonstrated in a case study. Matthias Becker 0004, Dakshina Dasari, Saad Mubeen, Moris Behnam, Thomas Nolte |
RTCSA | 4 |
| 2016 | Scheduling Real-Time Packets with Non-preemptive Regions on Priority-Based NoCsabstractNetwork-on-Chip (NoC) is a preferred communication medium for massively parallel platforms. Fixed-priority based scheduling using virtual-channels is one of the promising solutions to support real-time traffic in on-chip networks. Most of the existing NoC implementations which can support fixed-priority based scheduling use a flit-level preemptive scheduling. Under such a mechanism, preemptions can happen between the transmissions of successive flits. In this paper, we present a modified framework where the non-preemptive region of each NoC packet increases from a single flit. Using the proposed approach, the response times of certain packet flows can be reduced, which can thus improve the schedulability of the whole network. As a result, the utilization of NoCs can be improved by admitting more real-time traffic. Schedulability tests regarding the proposed framework are presented along with the proof of the correctness. Moreover, a number of experiments as well as a case study based on an automotive application have been generated, where we can clearly observe the improvement of our solution compared to the original flit-level preemptive NoC. Meng Liu 0001, Matthias Becker 0004, Moris Behnam, Thomas Nolte |
RTCSA | 3 |
| 2016 | Real-Time Capabilities of HSA Compliant COTS PlatformsabstractDuring recent years, the interest in using heterogeneous computing architecture in industrial applications has increased dramatically. These architectures provide the computational power that makes them attractive for many industrial applications. However, most of these existing heterogeneous architectures suffer from the following limitations: difficulties of heterogeneous parallel programming and high communication cost between the computing units. To overcome these disadvantages, several leading hardware manufacturers have formed the HSA Foundation to develop a new hardware architecture: Heterogeneous System Architecture (HSA). In this paper, we investigate the suitability of using HSA for real-time embedded systems. A preliminary experimental study has been conducted to measure massive computing power and timing predictability of HSA. Nandinbaatar Tsog, Matthias Becker 0004, Marcus Larsson, Fredrik Bruhn, Moris Behnam, Mikael Sjödin |
RTSS | 5 |
| 2016 | Dynamic reconfiguration in HaRTES switched ethernet networksabstractThe ability of reconfiguring a system during runtime is essential for dynamic real-time applications in which resource usage is traded online for quality of service. The HaRTES switch, which is a modified Ethernet switch, holds this ability for the network resource, and at the same time it provides hard real-time support for both periodic and sporadic traffic. Although the HaRTES switch technologically caters this ability, a protocol to actually perform the dynamic reconfiguration is missing in multi-hop HaRTES networks. In this paper we introduce such a protocol that is compatible with the traffic scheduling method used in the architecture. We prove the correctness of the protocol using a model checking technique. Moreover, we conduct a set of simulation experiments to show the performance of the protocol and we also show that the reconfiguration process is terminated within a bounded time. Mohammad Ashjaei, Luís Almeida 0001, Moris Behnam, Thomas Nolte |
WFCS | 4 |
| 2016 | A dependency-graph based priority assignment algorithm for real-time traffic over NoCs with shared virtual-channelsabstractThe Network-on-Chip (NoC) is the on-chip interconnection medium of choice for modern massively parallel processors and System-on-Chip (SoC) in general. Fixed-priority based preemptive scheduling using virtual-channels is a solution to support real-time communications in on-chip networks. Targeting the priority assignment problem in the context of NoCs, heuristic based priority assignment algorithms are more practical, due to the exponentially increased search space as the number of flows goes up. In our previous work, we have proposed a graph-based heuristic priority assignment algorithm (called GHSA) for NoC communications, where we show that taking the dependencies between flows into account can significantly reduce the search space. However, GHSA only works for NoCs with distinct priorities. Routers in such type of platforms may have a large amount of buffer cost when the number of flows is high. The applicability can thus be limited in reality. One solution to reduce the buffer cost is to allow priority sharing of different flows. In this paper, we propose a dependency-graph based priority assignment algorithm (called eGHSA) targeting NoCs with shared virtual-channels. A number of experiments as well as a case study based on an automotive application are generated, which clearly show that eGHSA improves the efficiency compared to the existing solution in the literature. Meng Liu 0001, Matthias Becker 0004, Moris Behnam, Thomas Nolte |
WFCS | 3 |
| 2016 | On providing real-time guarantees in cloud-based platformsabstractCloud technologies are gaining more and more attentions in recent years. Cloud-based service brings benefits in cost, energy efficiency, sharing of resources, increased flexibility, adaptability and evolvability. However, there are a number of associated challenges that need to be properly addressed before applying the cloud technique generally in industries. Providing efficient and predictable computation and communication is one of the important challenges, since many industrial systems (e.g. a control system) have specific timing requirements. Our current work thus focuses on guaranteeing the predictability of a cloud-based service. Virtualization, as one of the key technologies in Cloud Computing, is used to abstract details of resources away from end-services which simplifies the resource sharing. It thus improves the resource utilization and saves budget for end-users. In this preliminary work, we have implemented a distributed system using virtualization techniques (including virtual machines and virtual switches). Additionally, we generate a number of experiments to investigate how QoS policies can help us to provide real-time communication guarantees. Meng Liu 0001, Cezar Chiru, Moris Behnam, Kristian Sandström, Thomas Nolte |
WFCS | 3 |
| 2016 | Towards automated deployment of IEC 61131-3 applications on multi-core systemsabstractThe IEC 61131-3 standard, a widely used standard in the automation industry, defines various programming languages for programmable logic controllers. Today, the open source tools that comply with this standard do not support deployment of the applications on multi-core platforms. In this paper, we introduce a novel multi-step approach that aims to support automatic deployment of the automation control applications, developed using the IEC 61131-3 standard, to multi-core platforms. In the first step, the generated sequential code is partitioned. In the second step, the partitioned code is allocated to tasks while the tasks are mapped to various cores, without violating the dependencies, synchronization and communication constraints in the application. In order to provide a proof of concept, we develop a prototype by extending an existing tool that complies with the standard. We also perform a case study and a preliminary evaluation of the prototype. Saad Mubeen, Matthias Becker 0004, Xiaosha Zhao, Lingjian Gan, Moris Behnam, Thomas Nolte |
WFCS | 5 |
| 2016 | MTU configuration for real-time switched Ethernet networks
Mohammad Ashjaei, Moris Behnam, Luís Almeida 0001, Thomas Nolte |
J. Syst. Archit. | 2 |
| 2016 | SEtSim: A modular simulation tool for switched Ethernet networks
Mohammad Ashjaei, Moris Behnam, Thomas Nolte |
J. Syst. Archit. | 2 |
| 2015 | Investigation on AUTOSAR-Compliant Solutions for Many-Core ArchitecturesabstractAs of today, AUTOSAR is the de facto standard in the automotive industry, providing a common software architecture and development process for automotive applications. While this standard is originally written for singlecore operated Electronic Control Units (ECU), new guidelines and recommendations have been added recently to provide support for multicore architectures. This update came as a response to the steady increase of the number and complexity of the software functions embedded in modern vehicles, which call for the computing power of multicore execution environments. In this paper, we enumerate and analyze the design options and the challenges of porting AUTOSAR-based automotive applications onto multicore platforms. In particular, we investigate those options when considering the emerging many-core architectures that provide a more "scalable" environment than the traditional multicore systems. Such platforms are suitable to enable massive parallel execution, and their design is more suitable for partitioning and isolating the software components. Matthias Becker 0004, Dakshina Dasari, Vincent Nélis, Moris Behnam, Luís Miguel Pinho, Thomas Nolte |
DSD | 4 |
| 2015 | Resource sharing in a hybrid partitioned/global scheduling framework for multiprocessorsabstractFor resource-constrained embedded real-time systems, resource-efficient approaches are very important. Such an approach is presented in this paper, targeting systems where a critical application is partitioned on a multi-core platform and the remaining capacity on each core is provided to a noncritical application using resource reservation techniques. To exploit the potential parallelism of the non-critical application, global scheduling is used for its constituent tasks. Previously, we enabled intra-application resource sharing for such a framework, i.e. each application has its own dedicated set of resources. In this paper, we enable inter-application resource sharing, in particular between the critical application and the non-critical application. This effectively enables resource sharing in a hybrid partitioned/global scheduling framework on multiprocessors. For resource sharing, we use a spin-based synchronization protocol. We derive blocking bounds and extend existing schedulability analysis for such a system. Sara Afshar, Moris Behnam, Reinder J. Bril, Thomas Nolte |
ETFA | 2 |
| 2015 | Compositional analysis for the Multi-Resource ServerabstractThe Multi-Resource Server (MRS) technique has been proposed to enable predictable execution of memory intensive real-time applications on COTS multi-core platforms. It uses resource reservation approaches in the context of CPU-bandwidth and memory-bus bandwidth reservations to bound the interference between the applications running on the same core as well as between the applications running on different cores. In this paper we present a complete composable local and global schedulability analysis for the Multi-Resource Server technique. Based on the proposed analysis, we further provide an experimental study that investigates the behaviour of the MRS and identifies the factors that contribute mostly on the overall system performance. Rafia Inam, Moris Behnam, Thomas Nolte, Mikael Sjödin |
ETFA | 2 |
| 2015 | A many-core based execution framework for IEC 61131-3abstractProgrammable logic controllers are widely used for the control of automation systems. The standard IEC 61131-3 defines the execution model as well as the programming languages for such systems. Nowadays, actuators and sensors connect to the programmable logic controller via automation buses. While such buses, as well as the sensors and actuators, become more and more powerful, a shift away from the current distributed operation of automation systems, close to the field level, becomes possible. Instead, execution of complex control functions can be relocated to more powerful hardware, and technologies. This paper presents an execution framework for IEC 61131-3, based on a many-core processors. The presented execution model exploits the characteristics of the IEC 61131-3 applications as well as the characteristics of the many-core processor, yielding a predictable execution. We present the platform architecture and an algorithm to allocate a number of IEC 61131-3 conform applications. Experimental as well as simulation based evaluation is provided. Matthias Becker 0004, Kristian Sandström, Moris Behnam, Thomas Nolte |
IECON | 3 |
| 2015 | A feedback scheduling framework for component-based soft real-time systemsabstractComponent-based software systems with real-time requirements are often scheduled using processor reservation techniques. Such techniques have mainly evolved around hard real-time systems in which worst-case resource demands are considered for the reservations. In soft real-time systems, reserv- ing the processors based on the worst-case demands results in unnecessary over-allocations. In this paper, targeting soft real-time systems running on multiprocessor platforms, we focus on components for which processor demand varies during run-time. We propose a feedback scheduling framework where processor reservations are used for scheduling components. The reservation bandwidths as well as the reservation periods are adapted using MIMO LQR controllers. We provide an allocation mechanism for distributing components over processors. The proposed framework is implemented in the TrueTime simulation tool for system identification. We use a case study to investigate the performance of our framework in the simulation tool. Finally, the framework is implemented in the Linux kernel for practical evaluations. The evaluation results suggest that the framework can efficiently adapt the reservation parameters during run-time by imposing negligible overhead. Nima Moghaddami Khalilzad, Fanxin Kong, Xue (Steve) Liu, Moris Behnam, Thomas Nolte |
RTAS | 4 |
| 2015 | On Component-Based Software Development for Multiprocessor Real-Time SystemsabstractComponent-based software development provides a modular approach to develop complex software systems. In the context of real-time systems, it is desirable to abstract the timing properties of software components using an interface for each component. The timing properties of the whole system, composed of multiple components, is studied using the component interfaces. In this paper we focus on periodic interface models. In the case of components developed for single processor platforms, for examining the system schedulability, the interfaces can be regarded as periodic tasks. Thus, making it possible to use the conventional schedulability analyses for the system level schedulability test. In the case of components developed for multiprocessors, since interfaces may have utilization larger than 100% of a single processor, it is not possible to directly use the component interfaces for the system schedulability test. Therefore, the interfaces have to be decomposed before performing the system level schedulability test. In this paper, we target the special case of partitioned EDF for scheduling the components integrated on a multiprocessor. Therefore, the system level schedulability test is equivalent to finding a feasible allocation of component interfaces on the multiprocessor. We propose two algorithms for allocating the multiprocessor periodic interfaces. In addition, we propose an orthogonal approach for developing component-based real-time systems on multiprocessors in which components with utilization more than 100% of a single processor are divided into smaller subcomponents before abstracting their interfaces. We show, through extensive evaluations, that our alternative approach significantly reduces the interface overhead. Nima Moghaddami Khalilzad, Moris Behnam, Thomas Nolte |
RTCSA | 2 |
| 2015 | A Stochastic Response Time Analysis for Communications in On-chip NetworksabstractPriority-based wormhole-switching has been proposed as a solution to handle real-time traffic in on-chip networks. In order to support real-time traffic, the predictability of end-to-end delays need to be guaranteed. Several deterministic schedulability analysis approaches for wormhole-switched networks have been proposed. These approaches calculate a single upper-bound of the response time of each Network-on-Chip (NoC) flow, which is suitable for hard real-time applications. However, for many soft real-time applications, the performance does not depend on the worst-case scenario, which means that the calculated single upper-bounds are not sufficient to represent the performance. Therefore, in this paper, we present a stochastic Response Time Analysis (RTA) which can calculate a distribution of the response times of a real-time NoC flow. The estimated distributions can be utilized for multiple purposes, such as calculating deadline miss ratios, and computing upper-bounds regarding different probabilities. A number of simulation-based experiments are generated in order to investigate the pessimism involved in the analysis. Moreover, the processing time of the analysis is also measured from the experiments in order to examine the scalability of the proposed approach. Meng Liu 0001, Moris Behnam, Thomas Nolte |
RTCSA | 2 |
| 2015 | Semi-partitioning under a Blocking-Aware Task AllocationabstractSemi-partitioned scheduling is a resource efficient scheduling approach compared to the conventional multiprocessor scheduling approaches in terms of system utilization and migration overhead. Semi-partitioned scheduling can better utilize processor bandwidth compared to the partitioned scheduling while introducing less overhead compared to the global scheduling. Various techniques have been proposed to schedule tasks in a semi-partitioned environment, however, they have used blocking-agnostic allocation mechanisms in presence of resource sharing protocols. Since, the allocation mechanism can highly affect the system schedulability, in this paper we provide a blocking-aware allocation mechanism for semi-partitioned scheduling framework under a suspension-based resource sharing protocol. We have applied new heuristics for sorting the tasks in the algorithm that shows improvements upon system schedulability. Finally, we present our preliminary results. Sara Afshar, Moris Behnam, Thomas Nolte |
RTSS | 2 |
| 2014 | Evaluation of dynamic reconfiguration architecture in multi-hop switched ethernet networksabstractOn-the-fly adaptability and reconfigurability are recently becoming an interest in real-time communications. To assure a continued real-time behavior, the admission control with a quality-of-service mechanism is required, that screen all adaptation and reconfiguration requests. In the context of switched Ethernet networks, the FTT-SE protocol provides adaptive real-time communication. Recently, we proposed two methods to perform the online reconfiguration in multi-hop FTT-SE architectures. However, the methods lack the experimental evaluation. In this paper, we evaluate both methods in terms of the reconfiguration time. Mohammad Ashjaei, Paulo Pedreiras, Moris Behnam, Luís Almeida 0001, Thomas Nolte |
ETFA | 3 |
| 2014 | Limiting temperature gradients on many-cores by adaptive reallocation of real-time workloadsabstractThe advent of many-core processors came with the increase in computational power needed for future applications. However new challenges arrived at the same time, especially for the real-time community. Each core on such a processor is a heat source and uneven usage can lead to hot spots on the processor, affecting its lifetime and reliability. For real-time systems, it is therefore of paramount importance to keep the temperature differences between the individual cores below critical values, in order to prevent premature failure of the system. We argue that this problem can not be solved by traditional approaches, since the growing number of cores makes them intractable. We rather argue to split the problem in the spacial domain and control the temperature on core level. The cores control their temperature by rearranging the load in a predictable manner during runtime. To achieve this, a feedback controller is implemented on each core. We conclude our work with a simulation based evaluation of the proposed approach comparing its performance against a previously presented algorithm. Matthias Becker 0004, Kristian Sandström, Moris Behnam, Thomas Nolte |
ETFA | 3 |
| 2014 | A server-based approach for overrun management in multi-core real-time systemsabstractThis paper presents a server-based framework for task overrun management in multi-core real-time systems. Unlike most existing scheduling methods which usually assume a single upper bound of the Worst-Case Execution Time (WCET) for each task, our approach targets scenarios with task overruns. The main idea of our framework is to employ Synchronized Deferrable Servers (SDS) to deal with globally scheduled task overruns, while a partitioned scheduling approach is applied on regular task executions. Moreover, we provide a deterministic Worst-Case Response Time (WCRT) analysis focusing on hard timing constraints, along with a probabilistic analysis of Deadline Miss Ratio (DMR) for soft real-time applications. In the evaluation phase, we have implemented two types of experiments evaluating different timing constraints. Meng Liu 0001, Moris Behnam, Shinpei Kato, Thomas Nolte |
ETFA | 2 |
| 2014 | PASA: Framework for partitioning and scheduling automation applications on multicore controllersabstractWith multicore controllers becoming available for industrial automation applications, new tools and algorithms to compute efficient partitioning and scheduling solutions for control applications need to be developed. Optimizing the deployment and the schedule of a set of Function Block Diagrams on a parallel architecture are both NP hard. Additionally, control engineers need help to shift from the single core towards the multicore paradigm. By taking advantage of the parallelism inside the control applications it is effectively possible to decrease the finish times of the applications which enables to decrease their cycle times and improve the quality of service of the controller processes. This paper presents a practical solution to this problem that consists in a framework, called PASA, designed for partitioning and scheduling control applications modeled as function block diagrams. It enables new algorithms tailored to solve these optimization problems. This paper presents an extension of list-based DAG scheduling algorithms designed to compute a deployment and schedule for several control applications with different cycle times. The different variants of this algorithm are compared against each other as well as against some other existing solutions on a set of randomly generated examples. Aurelien Monot, Aneta Vulgarakis Feljan, Moris Behnam |
ETFA | 3 |
| 2014 | The Multi-Resource Server for predictable execution on multi-core platformsabstractIn this paper we present an implementation and demonstration of the Multi-Resource Server (MRS) which enables predictable execution of real-time applications on multi-core platforms. The MRS provides temporal isolation both between tasks running on the same core, as well as, between tasks running on different cores. The latter could, without MRS, interfere with each other due to contention on a shared memory bus. We demonstrate that MRS can be used to “encapsulate” legacy systems and to give them enough resources to fulfill their purpose. In our case study a legacy media-player is integrated with several resource-hungry tasks running at a different core. We show that without MRS the media-player starts to drop frames due to the interference from other tasks; while introduction of MRS alleviates this problem. Another part of our demonstration shows how traditional periodic real-time tasks can be kept schedulable even when tasks with high memory-demand are added to the system. Rafia Inam, Nesredin Mahmud, Moris Behnam, Thomas Nolte, Mikael Sjödin |
RTAS | 3 |
| 2014 | Reduced buffering solution for multi-hop HaRTES switched Ethernet networksabstractIn the context of switched Ethernet networks, multi-hop communication is essential as the networks in industrial applications comprise a high amount of nodes, that is far beyond the capability of a single switch. In this paper, we focus on multi-hop communication using HaRTES switches. The HaRTES switch is a modified Ethernet switch that provides real-time traffic scheduling, dynamic Quality-of-Service and temporal isolation between real-time and non-real-time traffic. Herein, we propose a method, called Reduced Buffering Scheme, to conduct the traffic through multiple HaRTES switches in a multi-hop HaRTES architecture. In order to enable the new scheduling method we propose to modify the HaRTES switch structure. Moreover, we develop a response time analysis for the new method. We also compare the proposed method with a method previously proposed, called Distributed Global Scheduling, based on their traffic response times. We show that, the new method forwards all types of traffic including the highest, the medium and the lowest priority, faster than the previous method in most of the cases. Furthermore, we show that the new method performs even better for larger networks compared with the previous one. Mohammad Ashjaei, Moris Behnam, Paulo Pedreiras, Reinder J. Bril, Luís Almeida 0001, Thomas Nolte |
RTCSA | 2 |
| 2014 | Optimal and fast composition of resource-sharing components in hierarchical real-time systemsabstractIn this paper we consider various flavors of the stack resource policy (SRP) for arbitrating access to shared resources in a hierarchical scheduling framework (HSF) upon a uni-processor. We propose algorithms for exploring and selecting the (local) resource ceilings within components, such that it results in an optimal composition of resource-sharing components in an HSF. Existing methods are non-optimal, because (i) they optimize just the length of global non-preemptive execution of tasks and (ii) they do not detect whether or not resources are shared globally (i.e., between tasks of different components). Lifting these limitations leads to an exponential growth of the design space. This paper contributes a fast three-step methodology which lifts these limitations, i.e., we apply the SRP at each level of the HSF to just the resources being shared. Our algorithm selects those component interfaces (i.e., by fixing the local resource ceilings of each component) that minimize the system load. Martijn M. H. P. van den Heuvel, Moris Behnam, Reinder J. Bril, Johan J. Lukkien, Thomas Nolte |
RTCSA | 2 |
| 2014 | An adaptive server-based scheduling framework with capacity reclaiming and borrowingabstractIn this paper, we present a new reservation based scheduling framework for soft real-time systems using EDF algorithm (called CARB-EDF). This framework has the features of Capacity Adaptation, Reclaiming and Borrowing. This framework can simplify the initial configuration of the system, where the system designer does not need to provide any estimations of task execution times. We also present a Chebyshev's inequality based predictor to estimate task execution times. A number of simulation-based experiments have been implemented. According to the results compared with some related works, our scheduling framework can provide a better performance with acceptable extra scheduling overhead. Meng Liu 0001, Moris Behnam, Shinpei Kato, Thomas Nolte |
RTCSA | 2 |
| 2014 | Integrating Cache-Related Pre-Emption Delays into Analysis of Fixed Priority Scheduling with Pre-Emption ThresholdsabstractCache-related pre-emption delays (CRPD) have been integrated into the schedulability analysis of sporadic tasks with constrained deadlines for fixed-priority pre-emptive scheduling (FPPS). This paper generalizes that work by integrating CRPD into the schedulability analysis of tasks with arbitrary deadlines for fixed-priority pre-emption threshold scheduling (FPTS). The analysis is complemented by an optimal threshold assignment algorithm that minimizes CRPD. The paper includes a comparative evaluation of the schedulability ratios of FPPS and FPTS, for constrained-deadline tasks, taking CRPD into account. Reinder J. Bril, Sebastian Altmeyer, Martijn M. H. P. van den Heuvel, Robert I. Davis 0001, Moris Behnam |
RTSS | 5 |
| 2013 | Implementing a clock synchronization protocol on a multi-master Switched Ethernet networkabstractThe interest to use Switched Ethernet technologies in real-time communication is increasing due to its absence of collisions when transmitting messages. Nevertheless, using COTS switches affect the timeliness guarantee inherent in potentially overflowing internal FIFO queues. In this paper we focus on a solution, called the FTT-SE protocol, which is developed based on a master-slave technique. Recently, an extension of the FTT-SE protocol has been proposed where the transmission of messages are controlled using multiple master nodes. In order to guarantee the correctness of the protocol, the masters should be timely synchronized. Therefore, in this paper we investigate the possibility of using a clock synchronization protocol, based on the IEEE 1588 standard, among master nodes. Moreover, we evaluate the overhead that is imposed by the clock synchronization protocol to the FTT-SE protocol. Finally, we present a formal verification of this solution by means of model checking technique to prove the correctness of the FTT-SE protocol when the clock synchronization protocol is applied. Mohammad Ashjaei, Moris Behnam, Guillermo Rodríguez-Navas, Thomas Nolte |
ETFA | 2 |
| 2013 | Towards implementing multi-resource server on multi-core Linux platformabstractIn this paper we present our ongoing work on implementing the multi-resource server technology in the Linux operating system running on multi-core architectures. The multi-resource server is used to control the access to both CPU and memory bandwidth resources such that the execution of real-time tasks become predictable. We are targeting Legacy applications to be migrated from single to multi-core architectures. We investigate the available techniques and mechanisms that can support our multi-resource servers and we discuss the potential problems that needed to be tackled considering the requirements of legacy applications. Rafia Inam, Joris Slatman, Moris Behnam, Mikael Sjödin, Thomas Nolte |
ETFA | 3 |
| 2013 | Adaptive hierarchical scheduling framework: Configuration and evaluationabstractWe have introduced an adaptive hierarchical scheduling framework as a solution for composing dynamic realtime systems, i.e., systems where the CPU demand of its tasks are subjected to unknown and potentially drastic changes during runtime. The framework consists of a controller which periodically adapts the system to the current load situation. In this paper, we unveil and explore the detailed behavior and performance of such an adaptive framework. Specifically, we investigate the controller configurations enabling efficient control parameters which maximizes performance, and we evaluate the adaptive framework against a traditional static one. Furthermore, we demonstrate the results of our investigation using a practical multimedia case study in which we simulate the timing behavior of video decoding tasks running on our proposed framework. In addition, we compare the results of using our framework with the results of using static resource allocation approach. Nima Moghaddami Khalilzad, Moris Behnam, Thomas Nolte |
ETFA | 2 |
| 2013 | Schedulability analysis of Multi-Frame messages over Controller Area Networks with mixed-queuesabstractThe Controller Area Network (CAN) is one of the most widely utilized real-time communication networks, which has plenty of applications especially in automotive industry. Many works have been proposed regarding the CAN schedulability analysis which is very important for guaranteeing the safety and reliability of hard real-time systems. Most of the existing analysis methods assume a periodic message model, and some of them take sporadic messages into account. However, for some applications, message transmissions may follow a specific pattern instead of repeating the same transmission period by period, where applying the existing methods may include much pessimism. In this paper, we apply the Multi-Frame task model, which is first proposed by Mok and Chen in [1], on messages over a CAN-based system with mixed message queues. In this preliminary work, the analysis is based on some specific assumptions. A more general analysis will be investigated in our later works. Meng Liu 0001, Moris Behnam, Thomas Nolte |
ETFA | 2 |
| 2013 | Resource sharing using the rollback mechanism in hierarchically scheduled real-time open systemsabstractIn this paper we present a new synchronization protocol called RRP (Rollback Resource Policy) which is compatible with hierarchically scheduled open systems and specialized for resources that can be aborted and rolled back. We conduct an extensive event-based simulation and compare RRP against all equivalent existing protocols in hierarchical fixed priority preemptive scheduling; SIRAP (Subsystem Integration and Resource Allocation Policy), OPEN-HSRPnP (open systems version of Hierarchical Stack Resource Policy no Payback) and OPEN-HSRPwP (open systems version of Hierarchical Stack Resource Policy with Payback). Our simulation study shows that RRP has better average-case response-times than the state-of-the-art protocol in open systems, i.e., SIRAP, and that it performs better than OPEN-HSRPnP/OPEN-HSRPwP in terms of schedulability of randomly generated systems. The simulations consider both resources that are compatible with rollback as well as resources incompatible with rollback (only abort), such that the resource-rollback overhead can be evaluated. We also measure CPU overhead costs (in VxWorks) related to the rollback mechanism of tasks and resources. We use the eXtremeDB (embedded real-time) database to measure the resource-rollback overhead1. Mikael Asberg, Thomas Nolte, Moris Behnam |
IEEE Real-Time and Embedded Technology and Applications Symposium | 3 |
| 2013 | Multi-level adaptive hierarchical scheduling framework for composing real-time systemsabstractProcessor partitioning and hierarchical scheduling have been widely used for composing hard real-time systems on a shared hardware platform while preserving the timing requirements of the systems. Due to the safety critical nature of hard real-time systems, a conservative analysis is often used for deriving a sufficient partition size. Applying the exact same analysis for deriving the partition sizes for soft real-time systems result in unnecessary processors overallocation and consequently waste of the CPU resource. In this paper, to address the problem of composing soft and hard real-time systems on a resource constrained shared hardware, we present a multi-level adaptive hierarchical scheduling framework. In our framework, we adapt the processor partition sizes of soft real-time systems according to their need at each time point by on-line monitoring their processor demand. Furthermore, we implement our adaptive framework in the Linux kernel and show the performance of our framework using a case study. Nima Moghaddami Khalilzad, Moris Behnam, Thomas Nolte |
RTCSA | 2 |
| 2013 | Applying the peak over thresholds method on worst-case response time analysis of complex real-time systemsabstractThe predictability of timing behavior is a very important performance issue of a real-time system. As the complexity of modern industrial systems increases, analyzing the timing behaviors of those systems becomes more and more challenging. Most of the existing analysis methods depend on static and detailed information of the systems under analysis. However, sometimes only partial information of a system can be available, or it may require too much effort on obtaining those details, making those analysis methods much less feasible. Moreover, those methods usually focus on some specific system models with unrealistic assumptions, consequently, applying those methods on a complex industrial real-time system may result in overly pessimistic results. Therefore, in this paper, we propose a statistical method to compute Worst-Case Response Times (WCRTs) of complex real-time systems regarding soft timing constraints, which can provide a higher general applicability with less required system information. Our approach employs a Peak Over Thresholds (POT) method, which is a branch of the Extreme Value Theory (EVT). For the evaluation, we have applied this approach on the analysis of message transmission latencies over Controller Area Networks (CAN). Meng Liu 0001, Moris Behnam, Thomas Nolte |
RTCSA | 2 |
| 2012 | On jitter in time partitioned real-time systemsabstractRecent trends towards adopting hypervisors, hierarchical scheduling, and other virtualization technologies that achieve partitioned access to the CPU and other resources impose significant impact with respect to jitter performance in embedded real-time systems. In this paper we make a first step towards characterization, modeling and calculation of this jitter. Kristian Sandström, Thomas Nolte, Moris Behnam, Reinder J. Bril |
ETFA | 3 |
| 2012 | Data Distribution Service for industrial automationabstractThe IEC 61499 is an open standard for the next generation of distributed control and automation. Data Distribution Service for Real-Time Systems (DDS) is a specification of a publish/subscribe middleware for distributed systems, created by the Object Management Group (OMG) to standardize a data-centric publish-subscribe programming model for distributed systems. This paper evaluates the DDS communication performance based on a model built within the IEC 61499 standard and compares it with the traditional socket based solution for communication. According to the test results, the DDS communication has the potential to reduce the complexity and is suggested as a suitable solution for some classes of industrial control systems. Jinsong Yang, Kristian Sandström, Thomas Nolte, Moris Behnam |
ETFA | 4 |
| 2011 | Independently-Developed Real-Time Systems on Multi-cores with Shared ResourcesabstractIn this paper we propose a synchronization protocol for resource sharing among independently-developed real-time systems on multi-core platforms. The systems may use different scheduling policies and they may have their own local priority settings. Each system is allocated on a dedicated processor (core). In the proposed synchronization protocol, each system is abstracted by an interface which abstracts the information needed for supporting global resources. The protocol facilitates the composability of various real-time systems with different scheduling and priority settings on a multi-core platform. We have performed experimental evaluations and compared the performance of our proposed protocol (MSOS) against the two existing synchronization protocols MPCP and FMLP. The results show that the new synchronization protocol enables composability without any significant loss of performance. In fact, in most cases the new protocol performs better than at least one of the other two synchronization protocols. Hence, we believe that the proposed protocol is a viable solution for synchronization among independently-developed real-time systems executing on a multi-core platform. Farhang Nemati, Moris Behnam, Thomas Nolte |
ECRTS | 2 |
| 2011 | Multi-level hierarchical scheduling in ethernet switchesabstractThe complexity of Networked Embedded Systems (NES) has been growing steeply, due to increases both in size and functionality, and is becoming a major development concern. This situation is pushing for paradigm changes in NES design methodologies towards higher composability and flexibility. Component-oriented design technologies, in particular supported by server-based scheduling, seem to be good candidates to provide the needed properties. Moris Behnam, Thomas Nolte, Paulo Pedreiras, Luís Almeida 0001 |
EMSOFT | 2 |
| 2011 | Analysis and optimization of the MTU in real-time communications over Switched EthernetabstractThe Flexible Time-Triggered communication over Switched Ethernet protocol (FTT-SE) was proposed to overcome the limitation of guaranteeing the real-time communication requirements of conventional switches, and at the same time to support reconfiguration of dynamic adaptive systems. The protocol fragments large messages into a sequence of packets that are individually scheduled. The maximum transmission unit (MTU), that restricts the packets size, has a significant effect on the schedulability of the packets. In this paper, we investigate the problem of selecting the optimal MTU size that maximizes the schedulability of real-time messages. We propose two algorithms to find optimal/sub-optimal values of MTU; the first one finds an optimal solution but exhibits high computational complexity, while the second one is sub-optimal but exhibits a lower computational complexity. Finally, we evaluate our proposed algorithms by means of simulation studies and compare their results with the results of assigning MTU to the maximum packet size that the protocol can allow. Moris Behnam, Ricardo Marau, Paulo Pedreiras |
ETFA | 1 |
| 2011 | Towards adaptive hierarchical scheduling of real-time systemsabstractHierarchical scheduling provides a modular framework for integrating, scheduling and guaranteeing timing constraints of compositional real-time systems. In such a scheduling framework, all modules should receive a sufficient portion of the shared CPU to be able to guarantee timing constraints of their internal parts. In dynamic systems i.e., systems where the execution time of tasks are subjected to sudden and drastic changes during run-time, assigning fixed CPU portions to the modules is conducive to either low CPU utilization or numerous task deadline misses. In this paper, in order to address this problem, we propose an adaptive CPU allocation method which dynamically assigns CPU portions to the modules during runtime based on their current CPU demand. Besides, the presented approach is evaluated using a series of different simulations. In addition, we present a method for scheduling modules in situations when the CPU resource is not sufficient for scheduling all modules. We introduce the notion of module (subsystem) criticality, and in an overload situation we distribute the CPU resource based on the criticality of modules. Nima Moghaddami Khalilzad, Thomas Nolte, Moris Behnam, Mikael Asberg |
ETFA | 3 |
| 2011 | Tighter Schedulability Analysis of Synchronization Protocols Based on Overrun without Payback for Hierarchical Scheduling FrameworksabstractIn this paper, we show that both global as well as local schedulability analysis of synchronization protocols based on the stack resource policy (SRP) and overrun without payback for hierarchical scheduling frameworks based on fixed-priority preemptive scheduling (FPPS) are pessimistic. We present tighter global and local schedulability analysis, illustrate the improvements of the new analysis by means of examples, and show that the improved global analysis is both uniform and sustainable. We evaluate the new global and local schedulability analysis based on an extensive simulation study and compare the results with the existing analysis. Moris Behnam, Thomas Nolte, Reinder J. Bril |
ICECCS | 1 |
| 2010 | On optimal real-time subsystem-interface generation in the presence of shared resourcesabstractThe Hierarchical Scheduling Framework (HSF) has been introduced as a design-time framework enabling compositional schedulability analysis of embedded software systems with real-time properties. However, supporting resource sharing in HSF is a major challenge, since it increases the amount of CPU resources required to guarantee schedulability of the hard real time tasks, and it decreases the composability at the system level. In this paper, we focus on a compositional framework called the bounded-delay resource open environment (BROE) server, and we identify key parameters of this framework that have a great effect on how the framework will utilize CPU resources. In addition, we show how to select optimal values for these parameters in order to reduce the required CPU resource. Moris Behnam, Thomas Nolte, Nathan Fisher |
ETFA | 1 |
| 2010 | Partitioning Real-Time Systems on Multiprocessors with Shared Resources
Farhang Nemati, Thomas Nolte, Moris Behnam |
OPODIS | 3 |
| 2010 | Bounding the Number of Self-Blocking Occurrences of SIRAPabstractThis paper presents a new schedulability analysis for hierarchically scheduled real-time systems executing on a single processor using SIRAP, a synchronization protocol for inter subsystem task synchronization. We show that it is possible to bound the number of self-blocking occurrences that should be taken into consideration in the schedulability analysis of subsystems. Correspondingly, we present two novel schedulability analysis approaches with proof of correctness for SIRAP. An evaluation suggests that this new schedulability analysis can decrease the analytical subsystem utilization significantly. Moris Behnam, Thomas Nolte, Reinder J. Bril |
RTSS | 1 |
| 2010 | Overrun Methods and Resource Holding Times for Hierarchical Scheduling of Semi-Independent Real-Time SystemsabstractThe hierarchical scheduling framework (HSF) has been introduced as a design-time framework to enable compositional schedulability analysis of embedded software systems with real-time properties. In this paper, a software system consists of a number of semi-independent components called subsystems. Subsystems are developed independently and later integrated to form a system. To support this design process, in the paper, the proposed methods allow non-intrusive configuration and tuning of subsystem timing-behavior via subsystem interfaces for selecting scheduling parameters. This paper considers three methods to handle overruns due to resource sharing between subsystems in the HSF. For each one of these three overrun methods corresponding scheduling algorithms and associated schedulability analysis are presented together with analysis that shows under what circumstances one or the other is preferred. The analysis is generalized to allow for both fixed priority scheduling (FPS) and earliest deadline first (EDF) scheduling. Also, a further contribution of the paper is the technique of calculating resource-holding times within the framework under different scheduling algorithms; the resource holding times being an important parameter in the global schedulability analysis. Moris Behnam, Thomas Nolte, Mikael Sjödin, Insik Shin |
IEEE Trans. Ind. Informatics | 1 |
| 2009 | Refining SIRAP with a dedicated resource ceiling for self-blockingabstractIn recent years, several synchronization protocols for resource sharing have been presented for use in a Hierarchical Scheduling Framework (HSF). An initial comparative assessment of existing protocols revealed that none of the protocols is superior to the others and that the performance of a protocol heavily depends on system parameters. In this paper, we aim at efficiency improvements of the synchronization protocol SIRAP [5] and its associated schedulability analysis, where efficiency refers to calculated CPU resource needs. The contribution of the paper is threefold. Firstly, we present an improvement of the schedulability analysis for SIRAP, which makes SIRAP more efficient. Secondly, we generalize SIRAP by distinguishing separate resource ceilings for self-blocking and resource access. Using a separate resource ceiling for self-blocking enables a reduction of the interference from lower priority tasks, which can result in efficiency improvements. The efficiency improvement depends on both subsystem characteristics and the value selected for the resource ceiling for self-blocking, however. The third contribution of this paper is therefore an algorithm that given a subsystem selects for each globally shared resource an optimal value in terms of efficiency for its resource ceiling for self-blocking. The efficiency improvement gained by the algorithm compared to the original SIRAP approach is evaluated by means of simulation. Moris Behnam, Thomas Nolte, Reinder J. Bril |
EMSOFT | 1 |
| 2009 | Towards Hierarchical Scheduling in AUTOSARabstractAUTOSAR is a partnership between automotive manufactures and suppliers. It aims at standardizing the automotive software architecture and separating software and hardware. This approach makes software more independent, maintainable, reuseable, etc. Still there is much work to do in order for this standard to be usable. This paper focus on automotive software integration in AUTOSAR, with the use of hierarchical scheduling as an enabling technology. At this point, AUTOSAR components do not have any timing relation with its tasks. This causes an unpredictive runtime behavior which can only be analyzed and verified after integration phase. We discuss how integration can be done in AUTOSAR, with runtime temporal isolation of components. This enable schedulability analysis at the level of components rather than at the level of tasks. Mikael Asberg, Moris Behnam, Farhang Nemati, Thomas Nolte |
ETFA | 2 |
| 2009 | Improved SIRAP Analysis for Synchronization in Hierarchical Scheduled Real-time SystemsabstractWe present our ongoing work on synchronization in hierarchical scheduled real-time systems, where tasks are scheduled using fixed-priority pre-emptive scheduling. In this paper, we show that the original local schedulability analysis of the synchronization protocol SIRAP [4] is very pessimistic when tasks of a subsystem access many global shared resources. The analysis therefore suggests that a subsystem requires more CPU resources than necessary. A new way to perform the schedulability analysis is presented which can make the SIRAP protocol more efficient in terms of calculated CPU resource needs. Moris Behnam, Thomas Nolte, Reinder J. Bril |
ETFA | 1 |
| 2009 | Efficiently Migrating Real-time Systems to Multi-coresabstractPower consumption and thermal problems limit a further increase of speed in single-core processors. Multi-core architectures have therefore received significant interest. However, a shift to multi-core processors is a big challenge for developers of embedded real-time systems, especially considering existing ¿legacy¿ systems which have been developed with uniprocessor assumptions. These systems have been developed and maintained by many developers over many years, and cannot easily be replaced due to the huge development investments they represent. An important issue while migrating to multi-cores is how to distribute tasks among cores to increase performance offered by the multi-core platform. In this paper we propose a partitioning algorithm to efficiently distribute legacy system tasks along with newly developed ones onto different cores. The target of the partitioning is increasing system performance while ensuring correctness. Farhang Nemati, Moris Behnam, Thomas Nolte |
ETFA | 2 |
| 2009 | Investigation of Implementing a Synchronization Protocol under Multiprocessors Hierarchical SchedulingabstractIn the multi-core and multiprocessor domain, there has been considerable work done on scheduling techniques assuming that real-time tasks are independent. In practice a typical real-time system usually share logical resources among tasks. However, synchronization in the multiprocessor area has not received enough attention. In this paper we investigate the possibilities of extending multiprocessor hierarchical scheduling to support an existing synchronization protocol (FMLP) in multiprocessor systems. We discuss problems regarding implementation of the synchronization protocol under the multiprocessor hierarchical scheduling. Farhang Nemati, Moris Behnam, Thomas Nolte, Reinder J. Bril |
ETFA | 2 |
| 2009 | Overrun and Skipping in Hierarchically Scheduled Real-Time SystemsabstractRecently, two SRP-based synchronization protocols for hierarchically scheduled real-time systems based on fixed priority preemptive scheduling (FPPS) have been presented, i.e., HSRP and SIRAP. Preventing depletion of budget during global resource access, the former implements an overrun mechanism, while the later exploits a skipping mechanism. A theoretical comparison of the performance of these mechanisms revealed that none of them was superior to the other, as their performance is heavily dependent on the system's parameters. To better understand the relative strengths and weaknesses of these mechanisms, this paper presents a comparative evaluation of the depletion prevention mechanisms overrun (with or without payback) and skipping. These mechanisms are investigated in detail and the corresponding system load imposed by these mechanisms is explored in a simulation study. The mechanisms are evaluated assuming FPPS and a periodic resource model. The periodic resource model is selected as it supports locality of schedulability analysis, allowing for a truthful comparison of the mechanisms. Given system characteristics, guiding the design of hierarchically scheduled real-time systems, the results of this paper indicate when one mechanism is better than the other and how a system should be configured in order to operate efficiently. Moris Behnam, Thomas Nolte, Mikael Asberg, Reinder J. Bril |
RTCSA | 1 |
| 2009 | A Synchronization Protocol for Temporal Isolation of Software Components in Vehicular SystemsabstractWe present a method that allows for integration of individually developed functions of software components into a predictable real-time system. The method has been designed to provide a lightweight mechanism that gives temporal firewalls between functions, preventing unpredictable side effects during function integration. The method maps well to the AUTOSAR (automotive open system architecture) software component model and can thus be used to facilitate seamless and predictable integration and isolation of AUTOSAR components that have been developed by different manufacturers. Specifically, this paper presents a protocol for synchronization in a hierarchical real-time scheduling framework. Using our protocol, a software component does not need to know, and is not dependent on, the timing behavior of software components belonging to other functions; even though they share mutually exclusive resources. In this paper, we also prove the correctness of our approach and evaluate its efficiency and cost in terms of system load in a vehicular context. Thomas Nolte, Insik Shin, Mikael Sjödin, Moris Behnam |
IEEE Trans. Ind. Informatics | 4 |
| 2008 | An Overrun Method to Support Composition of Semi-independent Real-Time ComponentsabstractEngineers of embedded software systems rely on efficient design techniques and tools along with efficient run-time support. In the design of complex embedded real-time systems, the hierarchical scheduling framework (HSF) has been introduced as a design-time framework enabling compositional schedulability analysis of embedded software systems with real-time properties. Moreover, the HSF provides a run-time framework guaranteeing that these nonfunctional requirements are met. In this paper a system consists of a number of semi- independent components called subsystems, and these subsystems are allowed to share logical resources. The HSF makes sure that the individual subsystems respect their allocated CPU budgets. However, as semi-independent subsystems share logical resources, extra complexity is introduced. Specifically, the contribution of this paper is a novel method to allow for budget overruns; a common scenario when a subsystem utilizes shared logical resources. This proposed method is not only more resource efficient than existing methods, but it is also more appropriate for supporting composability of independently developed real-time subsystems. Moris Behnam, Insik Shin, Thomas Nolte, Mikael Nolin |
COMPSAC | 1 |
| 2008 | Scheduling of semi-independent real-time components: Overrun methods and resource holding timesabstractThe hierarchical scheduling framework (HSF) has been introduced as a design-time framework enabling compositional schedulability analysis of embedded software systems with real-time properties. In this paper a system consists of a number of semi-independent components called subsystems. Subsystems are developed independently and later integrated to form a system. To support this design process, our proposed methods allow non-intrusive configuration and tuning of subsystem timing-behaviour via subsystem interfaces for selecting scheduling parameters. This paper considers two methods to handle overruns due to resource sharing between subsystems in the HSF. We present the scheduling algorithms for overruns and their associated schedulability analysis, together with analysis that shows under what circumstances one or the other overrun method is preferred. Furthermore, we show how to calculate resource-holding times within our framework. Moris Behnam, Insik Shin, Thomas Nolte, Mikael Nolin |
ETFA | 1 |
| 2008 | Synthesis of Optimal Interfaces for Hierarchical Scheduling with ResourcesabstractThis paper presents algorithms that (1) facilitate system-independent synthesis of timing-interfaces for subsystems and (2) system-level selection of interfaces to minimize CPU load. The results presented are developed for hierarchical fixed-priority scheduling of subsystems that may share logical recourses (i.e. semaphores). We show that the use of shared resources results in a tradeoff problem, where resource locking times can be traded for CPU allocation, complicating the problem of finding the optimal interface configuration subject to schedulability. This paper presents a methodology where such a tradeoff can be effectively explored. It first synthesizes a bounded set of interface-candidates for each subsystem, independently of the final system, such that the set contains the interface that minimizes system load for any given system. Then, integrating subsystems into a system, it finds the optimal selection of interfaces. Our algorithms have linear complexity to the number of tasks involved. Thus, our approach is also suitable for adaptable and reconfigurable systems. Insik Shin, Moris Behnam, Thomas Nolte, Mikael Nolin |
RTSS | 2 |
| 2007 | SIRAP: a synchronization protocol for hierarchical resource sharingin real-time open systemsabstractThis paper presents a protocol for resource sharing in a hierarchical real-time scheduling framework. Targeting real-time open systems, the protocol and the scheduling framework significantly reduce the efforts and errors associated with integrating multiple semi-independent subsystems on a single processor. Thus, our proposed techniques facilitate modern software development processes, where subsystems are developed by independent teams (or subcontractors) and at a later stage integrated into a single product. Using our solution, a subsystem need not know, and is not dependent on, the timing behaviour of other subsystems; even though they share mutually exclusive resources. In this paper we also prove the correctness of our approach and evaluate its efficiency. Moris Behnam, Insik Shin, Thomas Nolte, Mikael Nolin |
EMSOFT | 1 |
| 2007 | Real-Time Control and Scheduling Co-Design for Efficient Jitter HandlingabstractIn this paper we propose an integrated approach for control design and real-time scheduling, suitable for both discrete-time and continuous-time controllers. It guarantees system performance by accepting a certain minimum value of jitter for control tasks and feasibly schedules them together with other tasks in the system. Results from comparison with other approaches from real-time and control theory domains underline the effectiveness of our method. Moris Behnam, Damir Isovic |
RTCSA | 1 |