Mikael Sjödin

dblp:37/6301 · DBLP profile ↗
← Back
78ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0001-7586-0409ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 40 · 9 since 2021Software engineering, systems software and programming languages · 21 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Slowdown Modeling under Interference in Spatially Partitioned GPUs
Sahar Mobaiyen, Mikael Sjödin, Saad Mubeen
ISORC2
2025 Experimental Evaluation of a CAN-to-TSN Gateway Implementation
abstract
The increasing complexity of modern embedded systems highlights the limitations of Controller Area Network (CAN) in terms of transmission speed and scalability. The IEEE Time-Sensitive Networking (TSN) task group developed a set of standards to enhance switched Ethernet with high bandwidth, low jitter, and deterministic communication. Despite these advances, CAN will likely co-exist with TSN in, e.g., the automotive industry due to factors such as cost-effectiveness and legacy of CAN. This paper presents an experimental evaluation of a CAN-toTSN gateway implementation, focusing on the impact of different forwarding and scheduling strategies on network performance. We analyze various queuing techniques and scheduling mechanisms in a realistic experimental setup and assess their impact on end-to-end delay and TSN bandwidth utilization. The evaluation results demonstrate that encapsulating only a single CAN frame within a TSN frame effectively minimizes the end-to-end delay of CAN frames, in particular when a high-speed TSN network is used. Furthermore, we perform a comparative evaluation of the Time-Aware Shaper (TAS) and Weighted Round Robin (WRR) mechanisms in the TSN network. Interestingly, WRR leads to lower delays for CAN frames in the TSN network compared to TAS, which we attribute to the lack of synchronization between CAN and TSN.
Aldin Berisa, Benjamin Kraljusic, Nejla Zahirovic, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen
ISORC6
2025 Learning single and compound-protocol automata and checking behavioral equivalences
abstract
Abstract This paper presents a method and a practical implementation that complements traditional conformance testing. We infer a Mealy state machine of the system-under-test using active automata learning. This automaton is checked for bisimulation with a specification automaton modeled after the standard, which provides a strong verdict of conformance or nonconformance. We further present a method to learn models of multiple communication protocols running on the same device using a dispatcher system in conjunction with the same automata learning algorithms. We subsequently use similar checking methods to compare it with separately learned models. This allows for determining whether there is some interference or interaction between those protocols. In the practical execution of the system, we concentrate on lower levels of the Near-Field Communication (NFC, ISO/IEC 14443-3) and the Bluetooth Low-Energy (BLE) protocols. As a by-product, we share some observations of the performance of different learning algorithms and calibrations in the specific setting of ISO/IEC 14443-3, which is the difficulty to learn models of systems that a) consist of two very similar structures and b) timeout very frequently, as well as the role of conformance testing for compound models and speed optimizations for time-sensitive protocols.
Stefan Marksteiner, David Schögler, Marjan Sirjani, Mikael Sjödin
Int. J. Softw. Tools Technol. Transf.4
2024 Automated Passport Control: Mining and Checking Models of Machine Readable Travel Documents
abstract
Passports are part of critical infrastructure for a very long time. They also have been pieces of automatically processable information devices, more recently through the ISO/IEC 14443 (Near-Field Communication – NFC) protocol. For obvious reasons, it is crucial that the information stored on devices are sufficiently protected. The International Civil Aviation Organization (ICAO) specifies exactly what information should be stored on electronic passports (also Machine Readable Travel Documents – MRTDs) and how and under which conditions they can be accessed. We propose a model-based approach for checking the conformance with this specification in an automated and very comprehensive manner: we use automata learning to learn a full model of passport documents and use trace equivalence and primitive model checking techniques to check the conformance with an automaton modeled after the ICAO standard. Since the full behavior is underspecified in the standard, we compare a part of the learned model and apply a primitive checking ruleset to assure proper authentication. The result is an automated (non-interactive), yet very thorough test for compliance, despite the underspecification. This approach can also be used with other applications for which a specification automaton can be modeled and is therefore broadly applicable.
Stefan Marksteiner, Marjan Sirjani, Mikael Sjödin
ARES3
2023 Comparative Evaluation of Various Generations of Controller Area Network Based on Timing Analysis
abstract
This paper performs a comparative evaluation of various generations of Controller Area Network (CAN), including the classical CAN, CAN Flexible Data-Rate (FD), and CAN Extra Long (XL). We utilize response-time analysis for the evaluation. In this regard, we identify that the state of the art lacks the response-time analysis for CAN XL. Hence, we discuss the worst-case transmission times calculations for CAN XL frames and incorporate them to the existing analysis for CAN to support response-time analysis of CAN XL frames. Using the extended analysis, we perform a comparative evaluation of the three generations of CAN by analyzing an automotive industrial use case. In crux, we show that using CAN FD is more advantageous than the classical CAN and CAN XL when using frames with payloads of up to 8 bytes, despite the fact that CAN XL supports higher bit rates. For frames with 12-64 bytes payloads, CAN FD performs better than CAN XL when running at the same bit rate, but CAN XL performs better when running at a higher bit rate. Additionally, we discovered that CAN XL performs better than the classical CAN and CAN FD when the frame payload is over 64 bytes, even if it runs at the same or higher bit rates than CAN FD.
Aldin Berisa, Adis Panjevic, Imran Kovac, Hans Lyngbäck, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen
ETFA7
2023 Investigating and Analyzing CAN-to-TSN Gateway Forwarding Techniques
abstract
Controller Area Network (CAN) and Ethernet network are expected to co-exist in automotive industry as Ethernet provides a high-bandwidth communication, while CAN is a legacy cost-effective solution. Due to the shortcomings of conventional switched Etherent, such as determinism, IEEE Time Sensitive Networking (TSN) task group developed a set of standards to enhance the switched Ethernet technology providing low-jitter and deterministic communication. Considering these two network domains, we investigate various design approaches for a gateway that connects a CAN domain to a TSN domain. We present three gateway forwarding techniques and we develop end-to-end delay analysis methods for them. Via the analysis methods and applying them to synthetic use cases we show that the intuitive existing approach of encapsulating multiple CAN frames into a single Ethernet frame is not necessarily an efficient solution. In fact, we demonstrate several cases where it is preferable to encapsulate only one CAN frame into a TSN frame, in particular when we use a high speed TSN network. The results have a significant impact on developing such gateways as the implementation of the one-to-one frame encapsulation is considerably simpler than other complex gateway-forwarding techniques.
Aldin Berisa, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen
ISORC4
2023 End-to-end Timing Modeling and Analysis of TSN in Component-Based Vehicular Software
abstract
In this paper, we present an end-to-end timing model to capture timing information from software architectures of distributed embedded systems that use network communication based on the Time-Sensitive Networking (TSN) standards. Such a model is required as an input to perform end-to-end timing analysis of these systems. Furthermore, we present a methodology that aims at automated extraction of instances of the end-to-end timing model from component-based software architectures of the systems and the TSN network configurations. As a proof of concept, we implement the proposed end-to-end timing model and the extraction methodology in the Rubus Component Model (RCM) and its tool chain Rubus-ICE that are used in the vehicle industry. We demonstrate the usability of the proposed model and methodology by modeling a vehicular industrial use case and performing its timing analysis.
Bahar Houtan, Mehmet Onur Aybek, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, John Lundbäck, Saad Mubeen
ISORC5
2023 Supporting end-to-end data propagation delay analysis for TSN-based distributed vehicular embedded systems
abstract
In this paper, we identify that the existing end-to-end data propagation delay analysis for distributed embedded systems can calculate pessimistic (over-estimated) analysis results when the nodes are synchronized. This is particularly the case of the Scheduled Traffic (ST) class in Time-sensitive Networking (TSN), which is scheduled offline according to the IEEE 802.1Qbv standard and the nodes are synchronized according to the IEEE 802.1AS standard. We present a comprehensive system model for distributed embedded systems that incorporates all of the above mentioned aspect as well as all traffic classes in TSN. We extend the analysis to support both synchronization and non-synchronization among the ECUs as well as offline schedules on the networks. The extended analysis can now be used to analyze all traffic classes in TSN when the nodes are synchronized without introducing any pessimism in the analysis results. We evaluate the proposed model and the extended analysis on a vehicular industrial use case.
Bahar Houtan, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen
J. Syst. Archit.4
2023 A comprehensive systematic review of integration of time sensitive networking and 5G communication
abstract
Many industrial real-time applications in various domains, e.g., automotive, industrial automation, industrial IoT, and industry 4.0, require ultra-low end-to-end network latency, often in the order of 10 milliseconds or less. The IEEE 802.1 time-sensitive networking (TSN) is a set of standards that supports the required low-latency wired communication with ultra-low jitter. The flexibility of such a wired connection can be increased if it is integrated with a mobile wireless network. The fifth generation of cellular networks (5G) is capable of supporting the required levels of network latency with the Ultra-Reliable Low Latency Communication (URLLC) service. To fully utilize the potential of these two technologies (TSN and 5G) in industrial applications, seamless integration of the TSN wired-based network with the 5G wireless-based network is needed. In this article, we provide a comprehensive and well-structured snapshot of the existing research on TSN-5G integration. In this regard, we present the planning, execution, and analysis results of the systematic review. We also identify the trends, technical characteristics, and potential gaps in the state of the art, thus highlighting future research directions in the integration of TSN and 5G communication technologies. We notice that 73% of the primary studies address the time synchronization in the integration of TSN and 5G technologies, introducing approaches with an accuracy starting from the levels of hundred nanoseconds to one microsecond. Majority of primary studies aim at optimizing communication latency in their approach, which is a key quality attribute in automotive and industrial automation applications today.
Zenepe Satka, Mohammad Ashjaei, Hossein Fotouhi, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen
J. Syst. Archit.5
2023 From low-level programming to full-fledged industrial model-based development: the story of the Rubus Component Model
abstract
Abstract Developing distributed real-time systems is a complex task that has historically entailed specialized handcraft. In this paper, we propose a retrospective on the (r)evolutionary changes that led to the transition from low-level programming to industrial full-fledged model-based development embodied by the Rubus Component Model and its tool-ecosystem. We focus on the needs, challenges, and solutions of a 15-year-long evolution journey of a software development approach that has gone from low-level and manual programming to a highly automated environment offering modeling, analysis, and development of vehicular software systems with multi-criticality for deployment on single- and multi-core platforms.
Alessio Bucaioni, Federico Ciccozzi, Amleto Di Salle, Mikael Sjödin
Softw. Syst. Model.4
2022 TAS: Ternarized Neural Architecture Search for Resource-Constrained Edge Devices
abstract
Ternary Neural Networks (TNNs) compress network weights and activation functions into 2-bit representation resulting in remarkable network compression and energy efficiency. However, there remains a significant gap in accuracy between TNNs and full-precision counterparts. Recent advances in Neural Architectures Search (NAS) promise opportunities in automated optimization for various deep learning tasks. Unfortunately, this area is unexplored for optimizing TNNs. This paper proposes TAS, a framework that drastically reduces the accuracy gap between TNNs and their full-precision counterparts by integrating quantization into the network design. We experienced that directly applying NAS to the ternary domain provides accuracy degradation as the search settings are customized for full-precision networks. To address this problem, we propose (i) a new cell template for ternary networks with maximum gradient propagation; and (ii) a novel learnable quantizer that adaptively relaxes the ternarization mechanism from the distribution of the weights and activation functions. Experimental results reveal that TAS delivers 2.64% higher accuracy and ≃2.8 ×memory saving over competing methods with the same bit-width resolution on the CIFAR-10 dataset. These results suggest that TAS is an effective method that paves the way for the efficient design of the next generation of quantized neural networks.
Mohammad Loni, Mohammad Riazati, Masoud Daneshtalab, Mikael Sjödin
DATE5
2022 End-to-end Timing Model Extraction from TSN-Aware Distributed Vehicle Software
abstract
Extraction of end-to-end timing information from software architectures of vehicular systems to support their timing analysis is a daunting challenge. To address this challenge, this paper presents a systematic method to extract this information from vehicular software architectures that can be distributed over several electronic control units connected by Time-Sensitive Networking (TSN) networks. As a proof of concept, the proposed extraction method is applied to an industrial component model, namely the Rubus Component Model (RCM), and its toolchain. Furthermore, the usability of the proposed method is demonstrated in an industrial use case from the vehicular domain.
Bahar Houtan, Mehmet Onur Aybek, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen
SEAA5
2022 QoS-MAN: A Novel QoS Mapping Algorithm for TSN-5G Flows
abstract
Integrating wired Ethernet networks, such as Time-Sensitive Networks (TSN), to 5G cellular network requires a flow management technique to efficiently map TSN traffic to 5G Quality-of-Service (QoS) flows. The 3GPP Release 16 provides a set of predefined QoS characteristics, such as priority level, packet delay budget, and maximum data burst volume, which can be used for the 5G QoS flows. Within this context, mapping TSN traffic flows to 5G QoS flows in an integrated TSN-5G network is of paramount importance as the mapping can significantly impact on the end-to-end QoS in the integrated network. In this paper, we present a novel and efficient mapping algorithm to map different TSN traffic flows to 5G QoS flows. To the best of our knowledge, this is the first QoS-aware mapping algorithm based on the application constraints used to exchange flows between TSN and 5G network domains. We evaluate the proposed mapping algorithm on synthetic scenarios with random sets of constraints on deadline, jitter, bandwidth, and packet loss rate. The evaluation results show that the proposed mapping algorithm can fulfill over 90% of the applications’ constraints.
Zenepe Satka, Mohammad Ashjaei, Hossein Fotouhi, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen
RTCSA5
2022 Developing a Translation Technique for Converged TSN-5G Communication
abstract
Time Sensitive Networking (TSN) is a set of IEEE standards based on switched Ethernet that aim at meeting high-bandwidth and low-latency requirements in wired communication. TSN implementations typically do not support integration of wireless networks, which limits their applicability to many industrial applications that need both wired and wire-less communication. The development of 5G and its promised Ultra-Reliable and Low-Latency Communication (URLLC) in-tegrated with TSN would offer a promising solution to meet the bandwidth, latency and reliability requirements in these industrial applications. In order to support such an integration, we propose a technique to translate the traffic between TSN and 5G communication technologies. As a proof of concept, we implement the translation technique in a well-known TSN simulator, namely NeSTiNg, that is based on the OMNeT ++ tool. Furthermore, we evaluate the proposed technique using an automotive industrial use case.
Zenepe Satka, David Pantzar, Alexander Magnusson, Mohammad Ashjaei, Hossein Fotouhi, Mikael Sjödin, Masoud Daneshtalab, Saad Mubeen
WFCS6
2022 FastStereoNet: A Fast Neural Architecture Search for Improving the Inference of Disparity Estimation on Resource-Limited Platforms
abstract
Convolutional neural networks (CNNs) provide the best accuracy for disparity estimation. However, CNNs are computationally expensive, making them unfavorable for resource-limited devices with real-time constraints. Recent advances in neural architectures search (NAS) promise opportunities in automated optimization for disparity estimation. However, the main challenge of the NAS methods is the significant amount of computing time to explore a vast search space [e.g.,$1.6\times 10^{29}$] and costly training candidates. To reduce the NAS computational demand, many proxy-based NAS methods have been proposed. Despite their success, most of them are designed for comparatively small-scale learning tasks. In this article, we propose a fast NAS method, called FastStereoNet, to enable resource-aware NAS within an intractably large search space. FastStereoNet automatically searches for hardware-friendly CNN architectures based on late acceptance hill climbing (LAHC), followed by simulated annealing (SA). FastStereoNet also employs a fine-tuning with a transferred weights mechanism to improve the convergence of the search process. The collection of these ideas provides competitive results in terms of search time and strikes a balance between accuracy and efficiency. Compared to the state of the art, FastStereoNet provides$5.25\times $reduction in search time and$44.4\times $reduction in model size. These benefits are attained while yielding a comparable accuracy that enables seamless deployment of disparity estimation on resource-limited devices. Finally, FastStereoNet significantly improves the perception quality of disparity estimation deployed on field-programmable gate array and Intel Neural Compute Stick 2 accelerator in a significantly less onerous manner.
Mohammad Loni, Ali Zoljodi, Amin Majd, Byung Hoon Ahn, Masoud Daneshtalab, Mikael Sjödin, Hadi Esmaeilzadeh
IEEE Trans. Syst. Man Cybern. Syst.6
2021 LLM-shark - A Tool for Automatic Resource-boundness Analysis and Cache Partitioning Setup
abstract
We present LLM-shark, a tool for automatic hardware resource-boundness detection and cache-partitioning. Our tool has three primary objectives: First, it determines the hardware resource-boundness of a given application. Secondly, it estimates the initial cache partition size to ensure that the application performance is conserved and not affected by other processes competing for cache utilization. Thirdly, it continuously monitors that the application performance is maintained over time and, if necessary, change the cache partition size. We demonstrate LLM-shark’s functionality through a series of tests using six different applications, including a set of feature detection algorithms and two synthetic applications. Our tests reveal that it is possible to determine an application’s resource-boundness using a Pearson-correlation scheme implemented in LLM-shark. We propose a scheme to size cache partitions based on the correlation coefficient applications depending on their resource boundness.
Jakob Danielsson, Tiberiu Seceleanu, Marcus Jägemar, Moris Behnam, Mikael Sjödin
COMPSAC5
2021 Modelling Application Cache Behavior using Regression Models
abstract
In this paper, we describe the creation of resource usage forecasts for applications with unknown execution characteristics, by evaluating different regression processes, including autoregressive, multivariate adaptive regression splines, exponential smoothing, etc. We utilize Performance Monitor Units (PMU) and generate hardware resource usage models for the L2-cache and the L3-cache using nine different regression processes. The measurement strategy and regression process methodology are general and applicable to any given hardware resource when performance counters are available. We use three benchmark applications: the SIFT feature detection algorithm, a standard matrix multiplication, and a version of Bubblesort. Our evaluation shows that Multi Adaptive Regressive Spline (MARS) models generate the best resource usage forecasts among the considered models, followed by Single Exponential Splines (SES) and Triple Exponential Splines (TES).
Jakob Danielsson, Janne Suuronen, Marcus Jägemar, Tiberiu Seceleanu, Moris Behnam, Mikael Sjödin
COMPSAC6
2021 Automatic Quality of Service Control in Multi-core Systems using Cache Partitioning
abstract
In this paper, we present a last-level cache partitioning controller for multi-core systems. Our objective is to control the Quality of Service (QoS) of applications in multi-core systems by monitoring run-time performance and continuously re-sizing cache partition sizes according to the applications' needs. We discuss two different use-cases; one that promotes application fairness and another one that prioritizes applications according to the system engineers' desired execution behavior. We display the performance drawbacks of maintaining a fair schedule for all system tasks and its performance implications for system applications. We, therefore, implement a second control algorithm that enforces cache partition assignments according to user-defined priorities rather than system fairness. Our experiments reveal that it is possible, with non-instrusive (0.3-0.7% CPU utilization) cache controlling measures, to increase performance according to setpoints and maintain the QoS for specific applications in an over-saturated system.
Jakob Danielsson, Tiberiu Seceleanu, Marcus Jägemar, Moris Behnam, Mikael Sjödin
ETFA5
2021 Schedulability Analysis of Best-Effort Traffic in TSN Networks
abstract
This paper presents a schedulability analysis for the Best-Effort (BE) traffic class within Time Sensitive Networking (TSN) networks. The presented analysis considers several features in the TSN standards, including the Credit-Based Shaper (CBS), the Time-Aware Shaper (TAS) and the frame preemption. Although the BE class in TSN is primarily used for the traffic with no strict timing requirements, some industrial applications prefer to utilize this class for the non-hard real-time traffic instead of classes that use the CBS. The reason mainly lies in the fact that the complexity of TSN configuration becomes significantly high when the time-triggered traffic via the TAS and other classes via the CBS are used altogether. We demonstrate the applicability of the presented analysis on a vehicular application use case. We show that a network designer can get information on the schedulability of the BE traffic, based on which the network configuration can be further refined with respect to the application requirements.
Bahar Houtan, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, Sara Afshar, Saad Mubeen
ETFA4
2021 Offloading Accelerator-intensive Workloads in CPU-GPU Heterogeneous Processors
abstract
Autonomous vehicular systems require computer vision and intelligent on-board decision making functionalities that include a mix of sequential and parallel workloads. The execution times of the workloads and power consumption in these functionalities can be lowered by utilizing the accelerators (e.g., GPU) instead of running the workloads entirely on the host processing units (CPU). However, allocating all the parallelizable workload to accelerators can create a computation bottleneck in the accelerators that, in turn, can have an adverse effect on schedulability of the systems. This paper presents a novel framework that can allocate the accelerate-intensive workloads to the accelerators as well as to the non-accelerated host processing units. Within the context of this framework, the paper introduces five offloading techniques to mitigate the accelerator-intensive workloads by utilizing excess capacity of non-accelerated processing units under dynamic scheduling in CPU-GPU heterogeneous processors. The proposed techniques are evaluated using simulation experiments. The evaluation results indicate that one of the proposed techniques can achieve up to 16% improvement in schedulability of the task sets compared to the traditional non-offloading technique.
Nandinbaatar Tsog, Saad Mubeen, Fredrik Bruhn, Moris Behnam, Mikael Sjödin
ETFA5
2021 A novel frame preemption model in TSN networks
abstract
This paper identifies a limitation in the frame preemption model in the TSN standard (IEEE 802.1Q-2018), due to which high priority frames can experience significantly long blocking delays, thereby exacerbating their worst-case response times. This limitation can have a considerable impact on the design, analysis and performance of TSN-based systems. To address this limitation, the paper presents a novel and more efficient frame preemption model in the TSN standard that allows over 90% reduction in the maximum blocking delay leading to lower worst-case response times of high priority frames compared to the frame preemption model used in the existing works. The paper also shows that the improvement becomes even more significant in multi-switch TSN networks. In order to evaluate the effects of preemption, the paper performs simulations by enabling and disabling preemptions as well as enabling and disabling the Hold/Release mechanism supported by TSN. Furthermore, the paper performs a comparative evaluation of the two models of frame preemption in TSN using simulations. The evaluation shows that the maximum response times of high priority frames can be significantly reduced with very small impact on the response times of lower priority frames. The paper also shows the improvement in the maximum response times of higher priority frames using an automotive industrial use case that employs a multi-hop TSN network for on-board communication.
Mohammad Ashjaei, Mikael Sjödin, Saad Mubeen
J. Syst. Archit.2
2020 DenseDisp: Resource-Aware Disparity Map Estimation by Compressing Siamese Neural Architecture
abstract
Stereo vision cameras are flexible sensors due to providing heterogeneous information such as color, luminance, disparity map (depth), and shape of the objects. Today, Convolutional Neural Networks (CNNs) present the highest accuracy for the disparity map estimation [1]. However, CNNs require considerable computing capacity to process billions of floating-point operations in a real-time fashion. Besides, commercial stereo cameras produce huge size images (e.g., 10 Megapixels [2]), which impose a new computational cost to the system. The problem will be pronounced if we target resource-limited hardware for the implementation. In this paper, we propose DenseDisp, an automatic framework that designs a Siamese neural architecture for disparity map estimation in a reasonable time. DenseDisp leverages a meta-heuristic multi-objective exploration to discover hardware-friendly architectures by considering accuracy and network FLOPS as the optimization objectives. We explore the design space with four different fitness functions to improve the accuracy-FLOPS trade-off and convergency time of the DenseDisp. According to the experimental results, DenseDisp provides up to 39. 1x compression rate while losing around 5% accuracy compared to the state-of-the-art results.
Mohammad Loni, Ali Zoljodi, Daniel Maier 0002, Amin Majd, Masoud Daneshtalab, Mikael Sjödin, Ben H. H. Juurlink, Reza Akbari
CEC6
2020 Resource Depedency Analysis in Multi-Core Systems
abstract
In this paper, we evaluate different methods for statistical determination of application resource dependency in multi-core systems. We measure the performance counters of an application during run-time and create a system resource usage profile. We then use the resource profile to evaluate the application dependency on the specific resource. We discuss and evaluate two methods to process the data, including moving average filter and partitioning the data into smaller segments in order to interpret data for correlation calculations. Our aim with this study is to evaluate and create a generalizeable methods for automatic determination of resource dependencies. The final outcome of the methods used in this study is the answer to the question: "To what resources is this application dependent on?". The recommendation of this tool will be used in conjunction with our last-level cache partitioning controller (LLC-PC), to make decision if an application should receive last-level cache partition slices.
Jakob Danielsson, Tiberiu Seceleanu, Marcus Jägemar, Moris Behnam, Mikael Sjödin
COMPSAC5
2020 SHiLA: Synthesizing High-Level Assertions for High-Speed Validation of High-Level Designs
abstract
In the past, assertions were mostly used to validate the system through the design and simulation process. Later, a new method known as assertion synthesis was introduced, which enabled the designers to use the assertions for high-speed hardware emulation and safety and reliability insurance after tape-out. Although the synthesis of the assertions at the register transfer level is proposed and implemented in several works, none of them can be adopted for high-level assertions. In this paper, we propose the SHiLA framework and a detailed implementation guide by which assertion synthesis can also be applied to the high-level design processes. The proposed method, which is fully tool independent, is not only an enabler to highspeed assertion-assisted simulation but can also be used in other scenarios that need assertion synthesis, as it has the minimum possible effect on the main design's performance.
Mohammad Riazati, Masoud Daneshtalab, Mikael Sjödin, Björn Lisper
DDECS3
2020 Adjustable self-healing methodology for accelerated functions in heterogeneous systems
abstract
Self-healing is a promising approach for designing reliable digital systems. It refers to the ability of a system to detect faults and automatically fixing them to avoid total failure. With the development of digital systems, heterogeneous systems, in which some parts of the system are executed on the programmable logic, and some other parts run on the processing elements (CPU), are becoming more prevalent. In this work, we propose an adjustable self-healing method that is applicable to heterogeneous systems with accelerated functions and enables the designers to add the self-healing feature to the design. In this method, by manipulating the software codes that are being executed on the processing element, we add the ability to verify the accelerated functions on the programmable logic and heal the possible failures to the system. This is done not only in a straightforward manner but also without being forced to choose a specific reliability-overhead point. The designer will have the option to select the optimum configuration for a desired reliability level. Experimental results on a large design including several accelerated functions are provided and show 42% improvement of reliability by having 27% overhead, as an example of the reliability-overhead point.
Mohammad Riazati, Tara Ghasempouri, Masoud Daneshtalab, Jaan Raik, Mikael Sjödin, Björn Lisper
DSD5
2020 Preface from General Co-Chairs: PDP 2020
abstract
Presents the introductory welcome message from the conference proceedings. May include the conference officers' congratulations to all involved with the conference event and publication of the proceedings record.
Masoud Daneshtalab, Francesco Leporati, Mikael Sjödin
PDP3
2020 Modelling multi-criticality vehicular software systems: evolution of an industrial component model
abstract
Abstract Software in modern vehicles consists of multi-criticality functions, where a function can be safety-critical with stringent real-time requirements, less critical from the vehicle operation perspective, but still with real-time requirements, or not critical at all. Next-generation autonomous vehicles will require higher computational power to run multi-criticality functions and such a power can only be provided by parallel computing platforms such as multi-core architectures. However, current model-based software development solutions and related modelling languages have not been designed to effectively deal with challenges specific of multi-core, such as core-interdependency and controlled allocation of software to hardware. In this paper, we report on the evolution of the Rubus Component Model for the modelling, analysis, and development of vehicular software systems with multi-criticality for deployment on multi-core platforms. Our goal is to provide a lightweight and technology-preserving transition from model-based software development for single-core to multi-core. This is achieved by evolving the Rubus Component Model to capture explicit concepts for multi-core and parallel hardware and for expressing variable criticality of software functions. The paper illustrates these contributions through an industrial application in the vehicular domain.
Alessio Bucaioni, Saad Mubeen, Federico Ciccozzi, Antonio Cicchetti, Mikael Sjödin
Softw. Syst. Model.5
2019 Testing Performance-Isolation in Multi-core Systems
abstract
In this paper we present a methodology to be used for quantifying the level of performance isolation for a multi-core system. We have devised a test that can be applied to breaches of isolation in different computing resources that may be shared between different cores. We use this test to determine the level of isolation gained by using the Jailhouse hypervisor compared to a regular Linux system in terms of CPU isolation, cache isolation and memory bus isolation. Our measurements show that the Jailhouse hypervisor provides performance isolation of local computing resources such as CPU. We have also evaluated if any isolation could be gained for shared computing resources such as the system wide cache and the memory bus controller. Our tests show no measurable difference in partitioning between a regular Linux system and a Jailhouse partitioned system for shared resources. Using the Jailhouse hypervisor provides only a small noticeable overhead when executing multiple shared-resource intensive tasks on multiple cores, which implies that running Jailhouse in a memory saturated system will not be harmful. However, contention still exist in the memory bus and in the system-wide cache.
Jakob Danielsson, Tiberiu Seceleanu, Marcus Jägemar, Moris Behnam, Mikael Sjödin
COMPSAC (1)5
2019 TOT-Net: An Endeavor Toward Optimizing Ternary Neural Networks
abstract
High computation demands and big memory resources are the major implementation challenges of Convolutional Neural Networks (CNNs) especially for low-power and resource-limited embedded devices. Many binarized neural networks are recently proposed to address these issues. Although they have significantly decreased computation and memory footprint, they have suffered from accuracy loss especially for large datasets. In this paper, we propose TOT-Net, a ternarized neural network with [-1, 0, 1] values for both weights and activation functions that has simultaneously achieved a higher level of accuracy and less computational load. In fact, first, TOT-Net introduces a simple bitwise logic for convolution computations to reduce the cost of multiply operations. To improve the accuracy, selecting proper activation function and learning rate are influential, but also difficult. As the second contribution, we propose a novel piece-wise activation function, and optimized learning rate for different datasets. Our findings first reveal that 0.01 is a preferable learning rate for the studied datasets. Third, by using an evolutionary optimization approach, we found novel piece-wise activation functions customized for TOT-Net. According to the experimental results, TOT-Net achieves 2.15%, 8.77%, and 5.7/5.52% better accuracy compared to XNOR-Net on CIFAR-10, CIFAR-100, and ImageNet top-5/top-1 datasets, respectively.
Najmeh Nazari, Mohammad Loni, Mostafa E. Salehi, Masoud Daneshtalab, Mikael Sjödin
DSD5
2019 Holistic Modeling of Time Sensitive Networking in Component-Based Vehicular Embedded Systems
abstract
This paper presents the first holistic modeling approach for Time-Sensitive Networking (TSN) communication that integrates into a model-and component-based software development framework for distributed embedded systems. Based on these new models, we also present an end-to-end timing model for TSN-interconnected distributed embedded systems. Our approach is expressive enough to model the timing information of TSN and the timing behaviour of software that communicates over TSN, hence allowing end-to-end timing analysis. A proof of concept for the proposed approach is provided by implementing it for a component model and tool suite used in the vehicle industry. Moreover, a use case from the vehicle industry is modeled and analyzed with the proposed approach to demonstrate its usability.
Saad Mubeen, Mohammad Ashjaei, Mikael Sjödin
SEAA3
2019 NeuroPower: Designing Energy Efficient Convolutional Neural Network Architecture for Embedded Systems
Mohammad Loni, Ali Zoljodi, Sima Sinaei, Masoud Daneshtalab, Mikael Sjödin
ICANN (1)5
2019 Run-Time Cache-Partition Controller for Multi-Core Systems
abstract
The current trend in automotive systems is to integrate more software applications into fewer ECU's to decrease the cost and increase efficiency. This means more applications share the same resources which in turn can cause congestion on resources such as such as caches. Shared resource congestion may cause problems for time critical applications due to unpredictable interference among applications. It is possible to reduce the effects of shared resource congestion using cache partitioning techniques, which assign dedicated cache lines to different applications. We propose a cache partition controller called LLC-PC that uses the Palloc page coloring framework to decrease the cache partition sizes for applications during runtime. LLC-PC creates cache partitioning directives for the Palloc tool by evaluating the performance gained from increasing the cache partition size. We have evaluated LLC-PC using 3 different applications, including the SIFT image processing algorithm which is commonly used for feature detection in vision systems. We show that LLC-PC is able to decrease the amount of cache size allocated to applications while maintaining their performance allowing more cache space to be allocated for other applications.
Jakob Danielsson, Marcus Jägemar, Moris Behnam, Tiberiu Seceleanu, Mikael Sjödin
IECON5
2019 Static Allocation of Parallel Tasks to Improve Schedulability in CPU-GPU Heterogeneous Real-Time Systems
abstract
Autonomous driving is one of the main challenges of modern cars. Computer visions and intelligent on-board decision making are crucial in autonomous driving and require heterogeneous processors with high computing capability under low power consumption constraints. The progress of parallel computing using heterogeneous processing units is further supported by software frameworks like OpenCL, OpenMP, CUDA, and C++AMP. These frameworks allow the allocation of parallel computation on different compute resources. This, however, creates a difficulty in allocating the right computation segments to the right processing units in such a way that the complete system meets all its timing requirements. In this paper, we consider pre-runtime static allocations of parallel tasks to perform their execution either sequentially on CPU or in parallel using a GPU. This allows for improving any unbalanced use of GPU accelerators in a heterogeneous environment. By performing several heuristic algorithms, we show that the overuse of accelerators results in a bottle-neck of the entire system execution. The experimental results show that our allocation schemes that target a balanced use of GPU improves the system schedulability up to 90%.
Nandinbaatar Tsog, Matthias Becker 0004, Fredrik Bruhn, Moris Behnam, Mikael Sjödin
IECON5
2019 Work in Progress: Investigating the Effects of High Priority Traffic on the Best Effort Traffic in TSN Networks
abstract
This paper investigates the effects of various parameters of high priority traffic classes on the Best Effort (BE) traffic in the networks based on the IEEE Time Sensitive Networking (TSN) standards. In this regard, the paper discusses ongoing work and presents preliminary results using a TSN simulator. The results indicate that several parameters of the high priority traffic such as periods, offsets and preemption modes can have a significant impact on the quality of service (e.g., guaranteed message delivery and message delays) of the BE traffic.
Bahar Houtan, Mohammad Ashjaei, Masoud Daneshtalab, Mikael Sjödin, Saad Mubeen
RTSS4
2019 Supporting timing analysis of vehicular embedded systems through the refinement of timing constraints
abstract
The collective use of several models and tools at various abstraction levels and phases during the development of vehicular distributed embedded systems poses many challenges. Within this context, this paper targets the challenges that are concerned with the unambiguous refinement of timing requirements, constraints and other timing information among various abstraction levels. Such information is required by the end-to-end timing analysis engines to provide pre-run-time verification about the predictability of these systems. The paper proposes an approach to represent and refine such information among various abstraction levels. As a proof of concept, the approach provides a representation of the timing information at the higher levels using the models that are developed with EAST-ADL and Timing Augmented Description Language. The approach then refines the timing information for the lower abstraction levels. The approach exploits the Rubus Component Model at the lower level to represent the timing information that cannot be clearly specified at the higher levels, such as trigger paths in distributed chains. A vehicular-application case study is conducted to show the applicability of the proposed approach.
Saad Mubeen, Thomas Nolte, Mikael Sjödin, John Lundbäck, Kurt-Lennart Lundbäck
Softw. Syst. Model.3
2018 Measurement-Based Evaluation of Data-Parallelism for OpenCV Feature-Detection Algorithms
abstract
We investigate the effects on the execution time, shared cache usage and speed-up gains when using datapartitioned parallelism for the feature detection algorithms available in the OpenCV library. We use a data set of three different images which are scaled to six different sizes to exercise the different cache memories of our test architectures. Our measurements reveal that the algorithms using the default settings of OpenCV behave very differently when using data-partitioned parallelism. Our investigation shows that the executions of the algorithms SURF, Dense and MSER correlate to L3-cache usage and they are therefore not suitable for data-partitioned parallelism on multicore CPUs. Other algorithms: BRISK, FAST, ORB, HARRIS, GFTT, SimpleBlob and SIFT, do not correlate to L3-cache in the same extent, and they are therefore more suitable for data-partitioned parallelism. Furthermore, the SIFT algorithm provides the most stable speed-up, resulting in an execution between 3 and 3.5 times faster than the original execution time for all image sizes. We also have evaluated the hardware resource usage by measuring the algorithm execution time simultaneously with the L3-cache usage. We have used our measurements to conclude which algorithms are suitable for parallelization on hardware with shared resources.
Jakob Danielsson, Marcus Jägemar, Moris Behnam, Mikael Sjödin, Tiberiu Seceleanu
COMPSAC (1)4
2018 ADONN: Adaptive Design of Optimized Deep Neural Networks for Embedded Systems
abstract
Nowadays, many modern applications, e.g. autonomous system, and cloud data services need to capture and process a big amount of raw data at runtime that ultimately necessitates a high-performance computing model. Deep Neural Network (DNN) has already revealed its learning capabilities in runtime data processing for modern applications. However, DNNs are becoming more deep sophisticated models for gaining higher accuracy which require a remarkable computing capacity. Considering high-performance cloud infrastructure as a supplier of required computational throughput is often not feasible. Instead, we intend to find a near-sensor processing solution which will lower the need for network bandwidth and increase privacy and power efficiency, as well as guaranteeing worst-case response-times. Toward this goal, we introduce ADONN framework, which aims to automatically design a highly robust DNN architecture for embedded devices as the closest processing unit to the sensors. ADONN adroitly searches the design space to find improved neural architectures. Our proposed framework takes advantage of a multi-objective evolutionary approach, which exploits a pruned design space inspired by a dense architecture. Unlike recent works that mainly have tried to generate highly accurate networks, ADONN also considers the network size factor as the second objective to build a highly optimized network fitting with limited computational resource budgets while delivers comparable accuracy level. In comparison with the best result on CIFAR-10 dataset, a generated network by ADONN presents up to 26.4 compression rate while loses only 4% accuracy. In addition, ADONN maps the generated DNN on the commodity programmable devices including ARM Processor, High-Performance CPU, GPU, and FPGA.
Mohammad Loni, Masoud Daneshtalab, Mikael Sjödin
DSD3
2017 Technology-Preserving Transition from Single-Core to Multi-core in Modelling Vehicular Systems
Alessio Bucaioni, Saad Mubeen, Federico Ciccozzi, Antonio Cicchetti, Mikael Sjödin
ECMFA5
2017 Performance evaluation of network convergence time measurement techniques
abstract
In this paper we evaluate solutions that provide measurements for the network convergence time in switched Ethernet networks when links failures happen. We evaluate three solutions to measure the network convergence time in a faulty situation. Compared to the commercially available solutions, our proposals are cost-effective, portable, and open source. Thus, they are easy to deploy on many testbeds. We show the performance of the solutions by measuring different metrics including jitter, network convergence time and packet loss during the network recovery time. Our measurements indicate that it is possible to accurately measure the network convergence time using a packet sender which does not suffer interference from the overlying operating system. Furthermore, we noticed that the packet sniffer TShark did not suffer from kernel interrupts from overlying operating systems.
Jakob Danielsson, Mohammad Ashjaei, Moris Behnam, Thomas Sorensen, Mikael Sjödin, Thomas Nolte
ETFA5
2017 Investigating execution-characteristics of feature-detection algorithms
abstract
We discuss how to obtain information of execution characteristics, such as parallelizability and memory utilization, with the final aim to improve the performance and predictability of feature and corner detection algorithms for use in e.g. robotics and autonomous machines. Our aim is to obtain a better understanding of how computer vision algorithms use hardware resources and how to improve the time predictability and execution time of such algorithms when executing on multi-core CPUs. We evaluate a fork-join model applicable to feature detection algorithms and present a method for measuring how well the algorithm performance correlates with hardware resource usage. We have applied our method to the Featured from Accelerated Segment Test (FAST) algorithm. Our characterization of FAST reveals that it is an algorithm with excellent parallelism opportunities, resulting in an almost linear speed-up per core. Our measurements also reveal that the performance of FAST correlates very little with the number of misses in the L1 data cache, L1 instruction cache, data translation lookaside buffer and L2 cache. Thus, the FAST algorithm will not have a negative effect on the execution time when the input data fits in the L2 cache.
Jakob Danielsson, Marcus Jägemar, Moris Behnam, Mikael Sjödin
ETFA4
2016 Handling Uncertainty in Automatically Generated Implementation Models in the Automotive Domain
abstract
Models and model transformations, the two core constituents of Model-Driven Engineering, aid in software development by automating, thus taming, error-proneness of tedious engineering activities. In many cases, the result of these automated activities is an overwhelming amount of information. This is the case of one-to-many model transformations that, e.g. in model-based design-space exploration, can potentially generate a massive amount of candidate models (i.e., solution space) from one single source model. In our scenario, from one design model we generate a set of possible implementation models on which timing analysis is run. The aim is to find the best model from a timing perspective. However, multiple implementation models can have equally good analysis results. Therefore, the engineer is expected to investigate the solution space for making a final decision, using criteria which fall outside the analysis' criteria themselves. Since candidate models can be many and very similar to each other, manually finding differences and commonalities is an impractical and error-prone task. In order to provide the engineer with an expressive representation of models' commonalities and differences, we propose the use of modelling with uncertainty. We achieve this by elevating the solution space to a first-class status, adopting a compact notation capable of representing the solution space by means of a single model with uncertainty. Commonalities and differences are thus represented by means of uncertainty points for the engineer to easily grasp them and consistently make her decision without manually inspecting each model individually.
Alessio Bucaioni, Antonio Cicchetti, Federico Ciccozzi, Saad Mubeen, Alfonso Pierantonio, Mikael Sjödin
SEAA6
2016 Real-Time Capabilities of HSA Compliant COTS Platforms
abstract
During recent years, the interest in using heterogeneous computing architecture in industrial applications has increased dramatically. These architectures provide the computational power that makes them attractive for many industrial applications. However, most of these existing heterogeneous architectures suffer from the following limitations: difficulties of heterogeneous parallel programming and high communication cost between the computing units. To overcome these disadvantages, several leading hardware manufacturers have formed the HSA Foundation to develop a new hardware architecture: Heterogeneous System Architecture (HSA). In this paper, we investigate the suitability of using HSA for real-time embedded systems. A preliminary experimental study has been conducted to measure massive computing power and timing predictability of HSA.
Nandinbaatar Tsog, Matthias Becker 0004, Marcus Larsson, Fredrik Bruhn, Moris Behnam, Mikael Sjödin
RTSS6
2015 Compositional analysis for the Multi-Resource Server
abstract
The Multi-Resource Server (MRS) technique has been proposed to enable predictable execution of memory intensive real-time applications on COTS multi-core platforms. It uses resource reservation approaches in the context of CPU-bandwidth and memory-bus bandwidth reservations to bound the interference between the applications running on the same core as well as between the applications running on different cores. In this paper we present a complete composable local and global schedulability analysis for the Multi-Resource Server technique. Based on the proposed analysis, we further provide an experimental study that investigates the behaviour of the MRS and identifies the factors that contribute mostly on the overall system performance.
Rafia Inam, Moris Behnam, Thomas Nolte, Mikael Sjödin
ETFA4
2015 End-to-End Timing Analysis of Black-Box Models in Legacy Vehicular Distributed Embedded Systems
abstract
A majority of existing techniques and tools, used in the vehicular industry, support the extraction of end-to-end timing models. Such models are used to perform timing analysis of distributed embedded systems at an abstraction level that is close to their implementation. This paper takes a first initiative to provide such a support at a higher level of abstraction. At such a level, the system can be modeled with inter-connected black-box models of nodes whose internal software architectures may not be available. However, most of the design decisions about network communication are available. This represents a typical scenario in the vehicular industry where most of the artifacts are reused from either legacy systems, other projects or previous releases of the vehicle. In this paper we present an approach for the extraction of end-to-end timing models at the highest level of abstraction used in the vehicular domain. Using these models, end-to-end path delay analysis of the systems can be performed at a higher abstraction level and at an early phase during the development. As a proof of concept we implement this technique in an industrial tool suite, Rubus-ICE, that is used for the development of these systems by several international companies. Using the extended tool, we conduct a vehicular-application case study.
Saad Mubeen, Mikael Sjödin, Thomas Nolte, John Lundbäck, Mattias Galnander, Kurt-Lennart Lundbäck
RTCSA2
2015 Integrating mixed transmission and practical limitations with the worst-case response-time analysis for Controller Area Network
Saad Mubeen, Jukka Mäki-Turja, Mikael Sjödin
J. Syst. Softw.3
2014 Combating unpredictability in multicores through the multi-resource server
abstract
In this paper we present challenges that hinder the predictable integration and execution of real-time applications on multicore platforms. We investigate how shared resources, like CPU, memory-bus bandwidth, caches, and memory cause unpredictability and interference. We propose to adapt the traditional server-based scheduling approach on the multicore platforms with additional resource-reservations to control the shared access to such resources and present the multi-resource server as a solution such that the execution of real-time applications becomes predictable.
Rafia Inam, Mikael Sjödin
ETFA2
2014 Response time analysis with offsets for mixed messages in CAN supporting transmission abort requests
abstract
The existing worst-case response-time analysis for Controller Area Network (CAN) does not support mixed messages that are scheduled with offsets in the systems where the CAN controllers implement abortable transmit buffers. Mixed messages are partly periodic and partly sporadic. These messages are implemented by several higher-level protocols based on CAN that are used in the automotive industry. Moreover, most of the CAN controllers implement abortable transmit buffers. We extend the existing analysis with offsets for mixed messages in CAN. The extended analysis is applicable to any higher-level protocol for CAN that uses periodic, sporadic, and mixed transmission of messages where periodic and mixed messages can be scheduled with offsets in the systems that implement abortable transmit buffers in the CAN controllers. The extended analysis also supports gateway nodes in CAN by considering arbitrary jitter and deadlines for the messages. We also perform comparative evaluation of the existing and extended analyses.
Saad Mubeen, Jukka Mäki-Turja, Mikael Sjödin
ETFA3
2014 The Multi-Resource Server for predictable execution on multi-core platforms
abstract
In this paper we present an implementation and demonstration of the Multi-Resource Server (MRS) which enables predictable execution of real-time applications on multi-core platforms. The MRS provides temporal isolation both between tasks running on the same core, as well as, between tasks running on different cores. The latter could, without MRS, interfere with each other due to contention on a shared memory bus. We demonstrate that MRS can be used to “encapsulate” legacy systems and to give them enough resources to fulfill their purpose. In our case study a legacy media-player is integrated with several resource-hungry tasks running at a different core. We show that without MRS the media-player starts to drop frames due to the interference from other tasks; while introduction of MRS alleviates this problem. Another part of our demonstration shows how traditional periodic real-time tasks can be kept schedulable even when tasks with high memory-demand are added to the system.
Rafia Inam, Nesredin Mahmud, Moris Behnam, Thomas Nolte, Mikael Sjödin
RTAS5
2014 Communications-oriented development of component-based vehicular distributed real-time embedded systems
Saad Mubeen, Jukka Mäki-Turja, Mikael Sjödin
J. Syst. Archit.3
2014 MPS-CAN analyzer: Integrated implementation of response-time analyses for Controller Area Network
Saad Mubeen, Jukka Mäki-Turja, Mikael Sjödin
J. Syst. Archit.3
2014 Predictable integration and reuse of executable real-time components
Rafia Inam, Jan Carlson, Mikael Sjödin, Jirí Kuncar
J. Syst. Softw.3
2013 Testing of Timing Properties in Real-Time Systems: Verifying Clock Constraints
abstract
Ensuring that timing constraints in a real-time system are satisfied and met is of utmost importance. There are different static analysis methods that are introduced to statically evaluate the correctness of such systems in terms of timing properties, such as schedulability analysis techniques. Regardless of the fact that some of these techniques might be too pessimistic or hard to apply in practice, there are also situations that can still occur at runtime resulting in the violation of timing properties and thus invalidation of the static analyses' results. Therefore, it is important to be able to test the runtime behavior of a real-time system with respect to its timing properties. In this paper, we introduce an approach for testing the timing properties of real-time systems focusing on their internal clock constraints. For this purpose, test cases are generated from timed automata models that describe the timing behavior of real-time tasks. The ultimate goal is to verify that the actual timing behavior of the system at runtime matches the timed automata models. This is achieved by tracking and time-measuring of state transitions at runtime.
Mehrdad Saadatmand, Mikael Sjödin
APSEC (2)2
2013 Mode-change mechanisms support for hierarchical FreeRTOS implementation
abstract
Multi-mode embedded real-time systems exhibit a specific behaviour for each mode, and upon a mode-change request the task-set and timing interfaces of the system need to be changed. This paper presents the implementation of a MultiMode Adaptive Hierarchical Scheduling Framework (MMAHSF) and provides a generic skeleton (framework) for a two-level adaptive hierarchical scheduling supporting multiple modes and multiple mode-change mechanisms on an open source real-time operating system (FreeRTOS). The MMAHSF enable application-specific implementations of mode-change protocols using a set of predefined mode-change mechanisms. The paper addresses different mode-change mechanisms at both global and local scheduling levels. It presents examples of mode-change protocols that are developed by composing together these mechanisms in multiple ways and provide the initial results of executing these protocols in the MMAHSF implementation on an AVR 32-bit board EVK1100.
Rafia Inam, Mikael Sjödin, Reinder J. Bril
ETFA2
2013 Towards implementing multi-resource server on multi-core Linux platform
abstract
In this paper we present our ongoing work on implementing the multi-resource server technology in the Linux operating system running on multi-core architectures. The multi-resource server is used to control the access to both CPU and memory bandwidth resources such that the execution of real-time tasks become predictable. We are targeting Legacy applications to be migrated from single to multi-core architectures. We investigate the available techniques and mechanisms that can support our multi-resource servers and we discuss the potential problems that needed to be tackled considering the requirements of legacy applications.
Rafia Inam, Joris Slatman, Moris Behnam, Mikael Sjödin, Thomas Nolte
ETFA4
2013 Extending offset-based response-time analysis for mixed messages in Controller Area Network
abstract
The existing offset-based response-time analysis for mixed messages in Controller Area Network (CAN) assumes the jitter and deadline of a message to be smaller or equal to the transmission period. However, practical systems may contain messages whose release jitter and deadlines can be greater than their periods, e.g., in the gateway nodes. We extend the existing response-time analysis for mixed messages in CAN that are scheduled with offsets and have arbitrary jitter and deadlines. Mixed messages are implemented by several higher-level protocols for CAN that are used in the automotive industry. The extended analysis is applicable to any higher-level protocol for CAN that uses periodic, sporadic and mixed transmission modes.
Saad Mubeen, Jukka Mäki-Turja, Mikael Sjödin
ETFA3
2013 Round-trip support for extra-functional property management in model-driven engineering of embedded systems
Federico Ciccozzi, Antonio Cicchetti, Mikael Sjödin
Inf. Softw. Technol.3
2012 Towards Accurate Monitoring of Extra-Functional Properties in Real-Time Embedded Systems
abstract
Management and preservation of Extra-Functional Properties (EFPs) is critical in real-time embedded systems to ensure their correct behavior. Deviation of these properties, such as timing and memory usage, from their acceptable and valid values can impair the functionality of the system. In this regard, monitoring is an important means to investigate the state of the system and identify such violations. The monitoring result can also be used to make adaptation and re-configuration decisions in the system as well. Most of the works related to monitoring EFPs are based on the assumption that monitoring results accurately represent the true state of the system at the monitoring request time point. In some systems this assumption can be safe and valid. However, if in a system the value of an EFP changes frequently, the result of monitoring may not accurately represent the state of the system at the time point when the monitoring request has been issued. The consequences of such inaccuracies can be critical in certain systems and applications. In this paper, we mainly introduce and discuss this practical problem and also provide a solution to improve the monitoring accuracy of EFPs.
Mehrdad Saadatmand, Mikael Sjödin
APSEC2
2012 Enhancing the generation of correct-by-construction code from design models for complex embedded systems
abstract
Modern embedded systems are becoming more and more complex thus demanding for new powerful development mechanisms. Model-driven engineering has been recognised as a promising paradigm for the development of complex systems especially for its capability of abstracting the problem through models and then manipulating them to automatically generate target code. In our previous works, we presented mechanisms for the generation of 100% of the target code from UML models to be run on singlecore platforms. In this work we provide possible solutions to enhance the generation process to entail a more complex set of platform configurations (i.e., multiprocess, multicore) as well as heterogeneous processing units (i.e., CPU, GPU).
Federico Ciccozzi, Mikael Sjödin
ETFA2
2012 Bandwidth measurement using performance counters for predictable multicore software
abstract
Memory contention is one of the largest sources of inter-core interference in statically partitioned multicore systems, and the contention reduces the overall performance of applications and causes unpredictable execution-times. A first step in achieving predictable execution is to accurately measure the amount of consumed memory bandwidth for each application. Such measurements can be used to track down bottlenecks, provide better partitioning among cores, and ultimately be used to arbitrate and police access to the memory bus. We propose to use hardware performance counters to continuously track the memory-bandwidth consumed by different applications executing in parallel. In this paper we describe ongoing efforts exploring suitable performance counters on core-level and on system-on-chip level for the 8-core Freescale P4080 processor. The aim is to accurately and efficiently track consumed memory bandwidth per application; with the final goal to use these measurements to improve predictability of multicore realtime software.
Rafia Inam, Mikael Sjödin, Marcus Jägemar
ETFA2
2012 Worst-case response-time analysis for mixed messages with offsets in Controller Area Network
abstract
The existing response-time analysis for Controller Area Network (CAN) does not support mixed messages that are scheduled with offsets. Mixed messages are implemented by several high-level protocols for CAN that are used in the automotive industry. We extend the existing offset-based analysis which is applicable to any high-level protocol for CAN that uses periodic, sporadic and mixed transmission of messages. Moreover, we implement the extended analysis as a standalone simulator that will be integrated as a plug-in with the existing industrial tool suite (Rubus-ICE). The experiments, that we performed, indicate that it is possible to achieve up to 4.48% improvement in schedulability when mixed messages are scheduled with offsets.
Saad Mubeen, Jukka Mäki-Turja, Mikael Sjödin
ETFA3
2012 Extending response-time analysis of mixed messages in CAN with controllers implementing non-abortable transmit buffers
abstract
The existing response-time analysis for messages in Controller Area Network (CAN) with controllers implementing non-abortable transmit buffers does not support mixed messages that are implemented by several high-level protocols used in the automotive industry. We present the work in progress on the extension of the existing analysis for mixed messages. The extended analysis will be applicable to any high-level protocol for CAN that uses periodic, sporadic and mixed transmission modes and implements non-abortable transmit buffers in CAN controllers.
Saad Mubeen, Jukka Mäki-Turja, Mikael Sjödin
ETFA3
2012 Monitoring capabilities of schedulers in model-driven development of real-time systems
abstract
Model-driven development has the potential to reduce the design complexity of real-time embedded systems by increasing the abstraction level, enabling analysis at earlier phases of development, and automatic generation of code from the models. In this context, capabilities of schedulers as part of the underlying platform play an important role. They can affect the complexity of code generators and how the model is implemented on the platform. Also, the way a scheduler monitors the timing behaviors of tasks and schedules them can facilitate the extraction of runtime information. This information can then be used as feedback to the original model in order to identify parts of the model that may need to be re-designed and modified. This is especially important in order to achieve round-trip support for model-driven development of real-time systems. In this paper, we describe our work in providing such monitoring features by introducing a second layer scheduler on top of the OSE real-time operating system's scheduler. The goal is to extend the monitoring capabilities of the scheduler without modifying the kernel. The approach can also contribute to the predictability of applications by bringing more awareness to the scheduler about the type of real-time tasks (i.e., periodic, sporadic, and aperiodic) that are to be scheduled and the information that should be monitored and logged for each type.
Mehrdad Saadatmand, Mikael Sjödin, Naveed Ul Mustafa
ETFA2
2012 Data management for component-based embedded real-time systems: The database proxy approach
Andreas Hjertström, Dag Nyström, Mikael Sjödin
J. Syst. Softw.3
2011 Support for hierarchical scheduling in FreeRTOS
abstract
This paper presents the implementation of a Hierarchical Scheduling Framework (HSF) on an open source real-time operating system (FreeRTOS) to support the temporal isolation between a number of applications, on a single processor. The goal is to achieve predictable integration and reusability of independently developed components or applications. We present the initial results of the HSF implementation by running it on an AVR 32-bit board EVK1100. The paper addresses the fixed-priority preemptive scheduling at both global and local scheduling levels. It describes the detailed design of HSF with the emphasis of doing minimal changes to the underlying FreeRTOS kernel and keeping its API intact. Finally it provides (and compares) the results for the performance measures of idling and deferrable servers with respect to the overhead of the implementation.
Rafia Inam, Jukka Mäki-Turja, Mikael Sjödin, Seyed M. H. Ashjaei, Sara Afshar
ETFA3
2011 Extending schedulability analysis of Controller Area Network (CAN) for mixed (periodic/sporadic) messages
abstract
The schedulability analysis of Controller Area Network (CAN) developed by the research community is able to compute the response times of CAN messages that are queued for transmission periodically or sporadically. However, there are a few high-level protocols for CAN such as CANopen and Hägglunds Controller Area Network (HCAN) that support the transmission of mixed messages as well. A mixed message can be queued for transmission both periodically and sporadically. Thus, it does not exhibit a periodic activation pattern. The existing analysis of CAN does not support the analysis of mixed messages. We extend the existing analysis to compute the response times of mixed messages. The extended analysis is generally applicable to any high level protocol for CAN that uses any combination of periodic, event and mixed (periodic/event) transmission of messages.
Saad Mubeen, Jukka Mäki-Turja, Mikael Sjödin
ETFA3
2011 Extending response-time analysis of Controller Area Network (CAN) with FIFO queues for mixed messages
abstract
Existing response-time analysis for Controller Area Network (CAN) messages in networks where some nodes implement FIFO queues while others implement priority queues, assumes that at every node, CAN messages are queued for transmission periodically or sporadically. However, there are a few high level protocols for CAN such as CANopen and Hägglunds Controller Area Network (HCAN) that support the transmission of mixed messages as well. A mixed message can be queued for transmission both periodically and sporadically. The existing analysis of CAN with FIFO queues does not support the analysis of mixed messages. We extend the existing response-time analysis of mixed-type CAN messages. The extended analysis can compute the response-times of mixed (periodic/ sporadic) messages in the CAN network where some nodes use FIFO queues while others use priority queues.
Saad Mubeen, Jukka Mäki-Turja, Mikael Sjödin
ETFA3
2011 Towards resource sharing by message passing among real-time components on multi-cores
abstract
In this paper we propose a message passing synchronization protocol for resource sharing among real-time applications on multi-core platforms where each application is allocated on a cluster of cores. In this protocol the resources that are only used within an application (local resources) are handled by shared memory synchronization while the resources shared cross applications (global resources) are accessed by means of message passing. In our protocol the global resources are safely accessed without requiring to lock the resources explicitly. The goal is to avoid resource locking using shared memory, since accessing shared memory in multi-cores is very time consuming, whereas message passing has the potential to be much more efficient in systems with deep memory hierarchies.
Farhang Nemati, Rafia Inam, Thomas Nolte, Mikael Sjödin
ETFA4
2011 Enabling trade-off analysis of NFRs on models of embedded systems
abstract
Satisfaction of Non-Functional Requirements (NFR), is a key factor in successful design of embedded systems. This is mainly due to the constraints and resource limitations in these systems. A design that cannot achieve functionality of the system under these limitations is actually a failure. Therefore, NFRs in design of embedded systems deserve special attention. However, one big issue is that NFRs are interconnected and cannot be considered in isolation; especially that they can have direct impacts on each other such as security and performance. This means that a careful balance and trade-off analysis among NFRs is necessary. In this paper, we focus on this need and identify what information about NFRs is required in order to perform trade-off analysis. We propose and explain our in-progress approach to incorporate this information into system models in order to enable trade-off analysis. Our approach is based on UML profiling method to annotate model elements with necessary information.
Mehrdad Saadatmand, Antonio Cicchetti, Mikael Sjödin
ETFA3
2010 Database Proxies for Component-Based Real-Time Systems
abstract
We introduce the concept of database proxies capable of mitigating the gap between two disjoint productivity-enhancing techniques: Component Based Software Engineering (CBSE) and Real-Time Database Management Systems (RTDBMS). The coexistence of the two techniques is neither obvious nor intuitive since CBSE and RTDBMS promotes opposing design goals, CBSE promotes encapsulation and decoupling of component internals from the component environment, whilst RTDBMS provide mechanisms for efficient and predictable global data sharing. Database proxies decouple components from an underlying database residing in the component framework. This enables components to remain encapsulated and reusable, while providing temporally predictable access to data maintained in a database. We specifically target embedded systems with a subset of functionality with real-time requirements. Our implementation results show that the above benefits do not come at the expense of run-time overheads or less accurate timing predictions.
Andreas Hjertström, Dag Nyström, Mikael Sjödin
ECRTS3
2010 Guest editorial: special issue on the Real-Time and Network Systems (RTNS 2009) conference
Maryline Chetto, Mikael Sjödin
Real Time Syst.2
2010 Overrun Methods and Resource Holding Times for Hierarchical Scheduling of Semi-Independent Real-Time Systems
abstract
The hierarchical scheduling framework (HSF) has been introduced as a design-time framework to enable compositional schedulability analysis of embedded software systems with real-time properties. In this paper, a software system consists of a number of semi-independent components called subsystems. Subsystems are developed independently and later integrated to form a system. To support this design process, in the paper, the proposed methods allow non-intrusive configuration and tuning of subsystem timing-behavior via subsystem interfaces for selecting scheduling parameters. This paper considers three methods to handle overruns due to resource sharing between subsystems in the HSF. For each one of these three overrun methods corresponding scheduling algorithms and associated schedulability analysis are presented together with analysis that shows under what circumstances one or the other is preferred. The analysis is generalized to allow for both fixed priority scheduling (FPS) and earliest deadline first (EDF) scheduling. Also, a further contribution of the paper is the technique of calculating resource-holding times within the framework under different scheduling algorithms; the resource holding times being an important parameter in the global schedulability analysis.
Moris Behnam, Thomas Nolte, Mikael Sjödin, Insik Shin
IEEE Trans. Ind. Informatics3
2009 A Data-entity Approach for Component-based Real-time Embedded Systems Development
abstract
In this paper the data-entity approach for efficient design-time management of run-time data in component-based real-time embedded systems is presented. The approach formalizes the concept of a data entity which enable design-time modeling, management, documentation and analysis of run-time data items. Previous studies on data management for embedded real-time systems show that current data management techniques are not adequate, and therefore impose unnecessary costs and quality problems during system development. It is our conclusion that data management needs to be incorporated as an integral part of the development of the entire system architecture. Therefore, we propose an approach where run-time data is acknowledged as first class objects during development with proper documentation and where properties such as usage, validity and dependency can be modeled. In this way we can increase the knowledge and understanding of the system. The approach also allows analysis of data dependencies, type matching, and redundancy early in the development phase as well as in existing systems.
Andreas Hjertström, Dag Nyström, Mikael Sjödin
ETFA3
2009 A Synchronization Protocol for Temporal Isolation of Software Components in Vehicular Systems
abstract
We present a method that allows for integration of individually developed functions of software components into a predictable real-time system. The method has been designed to provide a lightweight mechanism that gives temporal firewalls between functions, preventing unpredictable side effects during function integration. The method maps well to the AUTOSAR (automotive open system architecture) software component model and can thus be used to facilitate seamless and predictable integration and isolation of AUTOSAR components that have been developed by different manufacturers. Specifically, this paper presents a protocol for synchronization in a hierarchical real-time scheduling framework. Using our protocol, a software component does not need to know, and is not dependent on, the timing behavior of software components belonging to other functions; even though they share mutually exclusive resources. In this paper, we also prove the correctness of our approach and evaluate its efficiency and cost in terms of system load in a vehicular context.
Thomas Nolte, Insik Shin, Mikael Sjödin, Moris Behnam
IEEE Trans. Ind. Informatics3
2003 Server-based scheduling of the CAN bus
abstract
In this paper we present a new share-driven server-based method for scheduling messages sent over the controller area network (CAN). Share-driven methods are useful in many applications, since they provide both fairness and bandwidth isolation among the users of the resource. Our method is the first share-driven scheduling method proposed for CAN. Our server-based scheduling is based on earliest deadline first (EDF), which allows higher utilization of the network than using CAN's native fixed-priority scheduling approach. We use simulation to show the performance and properties of server-based scheduling for CAN. The simulation results show that the bandwidth isolation property is kept, and they show that our method provides a quality-of-service (QoS), where virtually all messages are delivered within a specified time.
Thomas Nolte, Mikael Sjödin, Hans A. Hansson
ETFA (1)2
2003 Worst-case execution-time analysis for embedded real-time systems
Jakob Engblom, Andreas Ermedahl, Mikael Sjödin, Jan Gustafsson, Hans A. Hansson
Int. J. Softw. Tools Technol. Transf.3
2000 Supporting Timing Analysis by Automatic Bounding of Loop Iterations
Christopher A. Healy, Mikael Sjödin, Viresh Rustagi, David B. Whalley, Robert A. van Engelen
Real Time Syst.2
1998 Improved Response-Time Analysis Calculations
abstract
Schedulability analysis of fixed priority preemptive scheduled systems can be performed by calculating the worst-case response-time of the involved processes. The system is deemed schedulable if the calculated response-time for each process is less than its corresponding deadline. It is desirable that the Response-Time Analysis (RTA) can be efficiently performed. This is particularly important in dynamic real-time systems when a fast response is needed to decide whether a new job can be accommodated, or when the RTA is extensively applied, e.g., when used to guide the heuristics in a higher level optimiser. This paper presents a set of methods to improve the efficiency of RTA calculations. The methods are proved correct, in the sense that they give the same results as traditional (non-improved) RTA. We also present an evaluation of the improvements, by applying them to the particularly time-consuming traffic model used in RTA for ATM communication networks. Our evaluation shows that the proposed methods can give an order of magnitude reduction of the execution time of RTA.
Mikael Sjödin, Hans A. Hansson
RTSS1
1997 Response-time guarantees in ATM networks
abstract
We present a method for providing response time guarantees in Asynchronous Transfer Mode (ATM) networks. The method is based on traditional real time CPU Response Time Analysis (RTA), and is intended to be used for admission control of hard real time traffic. The method determines if a new connection can be admitted without violating the strict timing requirements specified for the new as well as old connections. We illustrate the merits of our method by comparing it with Weighted Fair Queuing (WFQ) and the Calculus for Network Delays (CND). Two types of comparisons are made. In the first, we evaluate how well the associated analysis can accommodate different traffic scenarios and loads, and in the second comparison we use simulation to compare observed worst case behaviors with estimates obtained by the analysis. The comparisons clearly indicate that RTA outperforms both WFQ and CND for a set of realistic traffic scenarios.
Andreas Ermedahl, Hans A. Hansson, Mikael Sjödin
RTSS3