Ingo Sander

dblp:21/1257 · DBLP profile ↗
← Back
60ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-4859-3100ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 40 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 33 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Diagrams as a Service
Niklas Rentz, Maximilian Kasperowski, Reinhard von Hanxleden, Klara Modin, Samuel Miksits, Muhammad Afif Ramadhan, Ingo Sander
Diagrams7
2026 Bridging the Abstraction Gap: A Systematic Approach to Rule-Based Transformational Design for Embedded Systems
abstract
Raising the level of abstraction is considered key to addressing the ever-increasing complexity of embedded system design, but it causes additional challenges due to the larger abstraction gap between the initial specification and the final implementation. This article addresses the current lack of systematic design methods by extending existing design-transformation-based approaches and wrapping them into a rule-based transformational design methodology for heterogeneous multi-processor platforms. The methodology cross-fertilizes embedded system design with program transformation techniques while taking into account the interplay of tight constraints and platform heterogeneity inherent in such systems. It advocates step-wise transformations starting from initial requirements to yield a final refined model that is efficient for implementation. To consider the effect of transformations on different properties of the system at each step, the system is specified with a set of requirements, an application model, a platform model, and a set of mapping decisions; referred to as the RAMP view of the system. The RAMP view and its carefully selected underlying unified abstract graph representation lay the foundations for mechanizing and potentially automating design transformations. A pattern matching technique is introduced and a proof-of-concept tool is implemented that automatically detects all possible transformations by matching the patterns defined by the transformation rules to the abstract graph representation of the system model. The underlying graph representation enables complex transformations on different aspects of the design, resulting in an improved design space definition. The design space can be explored by application-platform co-exploration techniques, yielding the most promising sequence of application transformations alongside the best matching platform. The applicability and potential of the proposed methodology are showcased through the design of both an image processing system and a cloud detection system.
Fahimeh Bahrami, Rodolfo Jordão, Ingo Sander, Ingemar Söderquist
ACM Trans. Embed. Comput. Syst.3
2025 Towards Coherent Semantics: A Quantitatively Typed EDSL for Synchronous System Design
abstract
We present SynQ, an embedded DSL (EDSL) targeting synchronous system design with quantitative types. SynQ is designed to facilitate semantically coherent system design processes by language embedding and advanced type systems. The current case study indicates the potential for a seamless system design process.
Ingo Sander
DATE2
2025 Automating Transformation Strategy via Attributed Graphs for Process Network Parallelization
abstract
The increasing complexity of multiprocessor embedded systems demands automated design flows that bridge the gap between high-level specifications and efficient parallel implementations. Automatic parallelization via rule-based model transformations has proven promising, but navigating the vast transformation space remains challenging. We present an automated strategy that systematically guides transformation application, enabling performance-aware parallelization through scalable and targeted space exploration.Our approach employs attributed graphs as a formal and expressive intermediate representation for evaluating and applying model transformations. These graphs enrich nodes and edges with semantic properties (such as process types, execution costs, and communication dependencies), capturing both application structure and platform characteristics in a unified model.We evaluate our strategy on an image processing application using a prototype implementation. The tool autonomously reduces the number of transformations from 203 to 74 and shortens the exploration time from over 17 hours to under 12, while improving performance by filtering out non-beneficial transformations.
Fahimeh Bahrami, Ingo Sander
FDL2
2025 Synchronous System Design with Quantitative Types
Ingo Sander
SETTA2
2025 A transformation strategy for process partitioning in hierarchical concurrent process networks
abstract
Concurrent process networks are a widely used parallel programming model for designing multiprocessor embedded systems, where system functionality is decomposed into processes that communicate via signals. These processes can be mapped onto different processing elements and executed concurrently. While the initial process network is designed to effectively capture high-level parallelism , it may not fully exploit the available parallelism . To enhance concurrency and balance workload distribution , process partitioning transformations are applied, restructuring process networks to expose finer-grained parallelism. The effectiveness of these transformations, however, depends on how well they align with the underlying hardware’s parallel capabilities. A variety of partitioning transformations have been introduced for process networks constructed using higher-order functions in the form of process constructors and data-parallel skeletons . For such networks, algebraic laws of functions provide a principled foundation for defining transformation rules , enabling a systematic and non-ad-hoc approach to process network modification. However, selecting the most suitable transformation to optimize key performance metrics remains an open challenge. To address this, we propose a transformation strategy that systematically identifies the most effective partitioning transformations. Our approach introduces evaluation metrics and analytical models to assess the impact of parametric transformations across different configurations. We validate the proposed strategy through the transformation of two image processing algorithms , demonstrating that our analytical models correctly predict the most suitable transformations for enhancing parallelism and performance.
Fahimeh Bahrami, Ingo Sander
J. Syst. Archit.2
2024 Automatic Parallelization of Embedded Software via Hierarchical Process Network Transformations
abstract
To fully utilize multi-processors, new tools are required to manage software complexity. We present a novel technique that enables automating hierarchical process network transformations to derive optimized parallel applications. Designers leverage a library of process constructors and data-parallel algorithmic skeletons, utilizing the well-defined semantics of a restricted set of operators. This carefully chosen set addresses both temporal and spatial aspects of computation, enabling the automated identification of various parallel patterns. We utilize an augmented version of a meta-modeling framework grounded in system graphs and trait hierarchies to generate an intermediate representation (IR) of the system model to simplify automatic transformations and evaluations. Our augmentation allows for capturing skeletons and hierarchical networks. By meticulously selecting the underlying framework, we alleviate the need for tool integration in our design flow. We validate our approach through a proof-of-concept implementation, where our automated tool applied 193 transformations to fully parallelize an image processing application.
Fahimeh Bahrami, Rodolfo Jordão, Ingo Sander, George Ungureanu
FDL3
2024 A Quantitative Type Approach to Formal Component-Based System Design
abstract
Functional programming languages are recognised for their high abstraction level, high expressiveness, formal semantics, and correspondence to formal logic. However, the utilisation of functional languages in system design is limited because the existence of stateful, black-box components, e.g., simulation models and legacy components, breaks the functional languages’ ground. Existing solutions to this situation, e.g. monads, are sub-optimal due to their ad-hoc and overconstrained nature. To address this challenge, we employ the quantitative type theory (QTT), which combines the dependent and linear (resource) type systems, for component specification. QTT enables stateful components to be used as pure functions with minimised restrictions. To this end, a functional language with QTT can be used for glue specification in a component-based design framework with all its advantages leveraged. The proposed approach is demonstrated by a case study in which a QTT-based RV32I instruction set architecture (ISA) specification in Idris2, a Haskell-like language, is simulated, verified, transformed and implemented in Verilog HDL by utilising properties of pure functions, which confirms the advantages of the proposed approach.
Ingo Sander
FDL2
2024 Multi-objective preference-free exact design space exploration of static DSP on multicore platforms
abstract
A challenge in designing resource-constrained embedded systems for digital signal processing (DSP) is their complexity due to their vast design spaces, where only a fraction of implementations are feasible or optimal. A crucial tool to aid in this challenge is automated design space exploration (DSE). However, no exact, multi-objective, and preference-free DSE approach exists for DSP applications on resource-constrained embedded platforms.We propose a novel DSE solution with these ideal characteristics to perform DSE of analyzable DSP applications for tile-based multiprocessing embedded platforms. Our proposal harmonizes the exactness of constraint programming (CP) and the exploration efficiency of genetic algorithms (GA). Through this synergy, no single-objective reduction strategy or a priori objective preferences is required.We evaluate the proposal through state-of-the-art single-objective case studies and multi-objective case studies inspired by these. The evaluations show that our proposal improves the single-objective state-of-the-art and finds high-quality approximate Pareto-frontiers for the multi-objective case study. Therefore, our proposal is a more performant single-objective DSE solution than the state-of-the-art, and it is the first exact, multi-objective, and preference-free DSE approach for the problem addressed.
Rodolfo Jordão, Fahimeh Bahrami, Yu Yang 0020, Matthias Becker 0004, Ingo Sander, Kathrin Rosvall
FDL5
2024 IDeSyDe: Systematic Design Space Exploration via Design Space Identification
abstract
Design space exploration (DSE) is a key activity in embedded design processes, where a mapping between applications and platforms that meets the process design requirements must be found. Finding such mappings is very challenging due to the complexity of modern embedded platforms and applications. DSE tools aid in this challenge by potentially covering sections of the design space that could be unintuitive to designers, leading to more optimised designs. Despite this potential benefit, DSE tools remain relatively niche in the embedded industry. A significant obstacle hindering their wider adoption is integrating such tools into embedded design processes. We present two contributions that address this integration issue. First, we present the design space identification (DSI) approach for systematically constructing DSE solutions that are modular and tuneable. Modularity means that DSE solutions can be reused to construct other DSE solutions, while tuneability means that the most specific DSE solution is chosen for the target DSE problem. Moreover, DSI enables transparent cooperation between exploration algorithms. Second, we present IDeSyDe, an extensible DSE framework for DSE solutions based on DSI. IDeSyDe allows extensions to be developed in different programming languages in a manner compliant with the DSI approach. We showcase the relevance of these contributions through five different case studies. The case study evaluations showed that non-exploration DSI procedures create overheads, which are marginal compared to the exploration algorithms. Empirically, most evaluations average 2% of the total DSE request. More importantly, the case studies have shown that IDeSyDe indeed provides a modular and incremental framework for constructing DSE solutions. In particular, the last case study required minimal extensions over the previous case studies so that support for a new application type was added to IDeSyDe.
Rodolfo Jordão, Matthias Becker 0004, Ingo Sander
ACM Trans. Design Autom. Electr. Syst.3
2022 A multi-view and programming language agnostic framework for model-driven engineering
abstract
Model-driven engineering (MDE) addresses the complexity of modern-day embedded system design. Multiple MDE frameworks are often integrated into a design process to use each MDE framework’s state-of-the-art tools for increased productivity. However, this integration requires substantial development effort.In this paper, we propose an MDE framework based on a formalism of system graphs and trait hierarchies for programming-language-agnostic integration between tools within our frame-work and with tools of other MDE frameworks. Implementing our framework for each programming language is a one-time development effort.We evaluate our proposal in an MDE design process by developing a Java supporting library and an AMALTHEA connector. Then we perform an MDE industrial avionics case study with both. The evaluation shows that our framework facilitates the integration of different tools and the independent development of different system parts. Therefore, our framework is a reliable MDE framework that lowers the effort of integrating tools to benefit from their combined state-of-the-art.
Rodolfo Jordão, Fahimeh Bahrami, Ingo Sander
FDL4
2021 Formulation of Design Space Exploration Problems by Composable Design Space Identification
abstract
Design space exploration (DSE) is a key activity in embedded system design methodologies and can be supported by well-defined models of computation (MoCs) and predictable platform architectures. The original design model, covering the application models, platform models and design constraints needs to be converted into a form analyzable by computer-aided decision procedures such as mathematical programming or genetic algorithms. This conversion is the process of design space identification (DSI), which becomes very challenging if the design domain comprises several MoCs and platforms. For a systematic solution to this problem, separation of concerns between the design domain and decision domain is of key importance. We propose in this paper a systematic DSI scheme that is (a) composable, as it enables the stepwise and simultaneous extension of both design and decision domain, and (b) tuneable, because it also enables different DSE solving techniques given the same design model. We exemplify this DSI scheme by an illustrative example that demonstrates the mechanisms for composition and tuning. Additionally, we show how different compositions can lead to the same decision model as an important property of this DSI scheme.
Rodolfo Jordão, Ingo Sander, Matthias Becker 0004
DATE2
2021 An automated parallel simulation flow for cyber-physical system design
Seyed-Hosein Attarzadeh-Niaki, Ingo Sander
Integr.2
2021 ForSyDe-Atom: Taming Complexity in Cyber Physical System Design with Layers
abstract
We present ForSyDe-Atom, a formal framework intended as an entry point for disciplined design of complex cyber-physical systems. This framework provides a set of rules for combining several domain-specific languages as structured, enclosing layers to orthogonalize the many aspects of system behavior, yet study their interaction in tandem . We define four layers: one for capturing timed interactions in heterogeneous systems, one for structured parallelism, one for modeling uncertainty, and one for describing component properties. This framework enables a systematic exploitation of design properties in a design flow by facilitating the stepwise projection of certain layers of interest, the isolated analysis and refinement on projections, and the seamless reconstruction of a system model by virtue of orthogonalization. We demonstrate the capabilities of this approach by providing a compact yet expressive model of an active electronically scanned array antenna and signal processing chain, simulate it, validate its conformity with the design specifications, refine it, synthesize a sub-system to VHDL and sequential code, and co-simulate the generated artifacts.
George Ungureanu, José Edil G. de Medeiros, Timmy Sundström, Ingemar Söderquist, Anders Ahlander, Ingo Sander
ACM Trans. Embed. Comput. Syst.6
2020 Exploiting Dataflow Models for Parallel Simulation of Discrete Timed Systems
abstract
The shift towards parallel computing witnessed since the turn of this century has forced us to rethink traditional software design paradigms to better utilize resources. Yet, the simulation of time-aware systems remains a challenging topic due to the inherent semantics of time and causality whose consistency needs to be controlled, traditionally in form of a global event queue, limiting the potential for parallel exploitation. We propose a rehash of this problem by tackling it from a different modeling perspective, one which is able to express concurrency more naturally, i.e. dataflow (DF) models of computation (MoCs). By abstracting time aspects as an algebra hosted on a pure DF MoC, we are able to apply recent results from MoC theory not only for the purpose of describing deterministic behaviors for distributed timed systems, but also to overcome the existing limitations of timed execution in order to increase a simulation model's performance. We use a well-known example of a deadlock-prone distributed discrete event system as a driver to introduce the modeling concepts and show their potential for parallelism.
George Ungureanu, Rodolfo Jordão, Ingo Sander
FDL3
2019 Formal Design, Co-Simulation and Validation of a Radar Signal Processing System
abstract
With the ever increasing complexity in safety-critical and performance-demanding application domains such as automotive and avionics, the costs of designing, producing and especially testing systems does not scale well for the next generation of applications. One example is the active electronically scanned array (AESA) antenna signal processing chain, which is currently out-of-reach from consumer products but rather part of a few exclusive hi-tech appliances. To cope with the associated complexity of such systems, we propose a design flow starting from a high-level formal modeling language which captures and exposes important design properties to enable their systematic exploitation for the purpose of simulation, analysis and synthesis towards cost-efficient implementations. We demonstrate the capabilities of this approach by providing a compact yet expressive description of the AESA signal processing chain, generate automatic test-cases to verify the conformity of model with design specifications, synthesize a part of it to VHDL and co-simulate the generated artifact to validate its correctness.
George Ungureanu, Timmy Sundström, Anders Ahlander, Ingo Sander, Ingemar Söderquist
FDL4
2019 Modeling and Simulation of Dynamic Applications Using Scenario-Aware Dataflow
abstract
The tradeoff between analyzability and expressiveness is a key factor when choosing a suitable dataflow model of computation (MoC) for designing, modeling, and simulating applications considering a formal base. A large number of techniques and analysis tools exist for static dataflow models, such as synchronous dataflow. However, they cannot express the dynamic behavior required for more dynamic applications in signal streaming or to model runtime reconfigurable systems. On the other hand, dynamic dataflow models like Kahn process networks sacrifice analyzability for expressiveness. Scenario-aware dataflow (SADF) is an excellent tradeoff providing sufficient expressiveness for dynamic systems, while still giving access to powerful analysis methods. In spite of an increasing interest in SADF methods, there is a lack of formally-defined functional models for describing and simulating SADF systems. This article overcomes the current situation by introducing a functional model for the SADF MoC, as well as a set of abstract operations for simulating it. We present the first modeling and simulation tool for SADF so far, implemented as an open source library in the functional framework ForSyDe. We demonstrate the capabilities of the functional model through a comprehensive tutorial-style example of a RISC processor described as an SADF application, and a traditional streaming application where we model an MPEG-4 simple profile decoder. We also present a couple of alternative approaches for functionally modeling SADF on different languages and paradigms. One of such approaches is used in a performance comparison with our functional model using the MPEG-4 simple profile decoder as a test case. As a result, our proposed model presented a good tradeoff between execution time and implementation succinctness. Finally, we discuss the potential of our formal model as a frontend for formal system design flows regarding dynamic applications.
Ricardo Bonna, Denis Silva Loubach, George Ungureanu, Ingo Sander
ACM Trans. Design Autom. Electr. Syst.4
2018 An algebra for modeling continuous time systems
abstract
Advancements on analog integrated design have led to new possibilities for complex systems combining both continuous and discrete time modules on a signal processing chain. However, this also increases the complexity any design flow needs to address in order to describe a synergy between the two domains, as the interactions between them should be better understood. We believe that a common language for describing continuous and discrete time computations is beneficial for such a goal and a step towards it is to gain insight and describe more fundamental building blocks. In this work we present an algebra based on the General Purpose Analog Computer, a theoretical model of computation recently updated as a continuous time equivalent of the Turing Machine.
José Edil G. de Medeiros, George Ungureanu, Ingo Sander
DATE3
2018 Bridging discrete and continuous time models with atoms
abstract
Recent trends in replacing traditionally digital components with analog counterparts in order to overcome physical limitations have led to an increasing need for rigorous modeling and simulation of hybrid systems. Combining the two domains under the same set of semantics is not straightforward and often leads to chaotic and non-deterministic behavior due to the lack of a common understanding of aspects concerning time. We propose an algebra of primitive interactions between continuous and discrete aspects of systems which enables their description within two orthogonal layers of computation. We show its benefits from the perspective of modeling and simulation, through the example of an RC oscillator modeled in a formal framework implementing this algebra.
George Ungureanu, José Edil G. de Medeiros, Ingo Sander
DATE3
2018 Exploring Power and Throughput for Dataflow Applications on Predictable NoC Multiprocessors
abstract
System level optimization for multiple mixed-criticality applications on shared networked multiprocessor platforms is extremely challenging. Substantial complexity arises from the interdependence between the multiple subproblems of mapping, scheduling and platform configuration under the consideration of several, potentially orthogonal, performance metrics and constraints. Instead of using heuristic algorithms and problem decomposition, novel unified design space exploration (DSE) approaches based on Constraint Programming (CP) have in the recent years shown promising results. The work in this paper takes advantage of the modularity of CP models, in order to support heterogeneous multiprocessor Network-on-Chip (NoC) with Temporally Disjoint Networks (TDNs) aware message injection. The DSE supports a range of design criteria, in particular the optimization and satisfaction of power and throughput. In addition, the DSE now provides a valid configuration for the TDNs that guarantees the performance required to fulfil the design goals. The experiments show the capability of the approach to find low-power and high-throughput designs, and validate a resulting design on a physical TDN-based NoC implementation.
Kathrin Rosvall, Tage Mohammadat, George Ungureanu, Johnny Öberg, Ingo Sander
DSD5
2018 Flexible and Tradeoff-Aware Constraint-Based Design Space Exploration for Streaming Applications on Heterogeneous Platforms
abstract
Due to its complexity, the problem of mapping and scheduling streaming applications on heterogeneous MPSoCs under real-time and performance constraints has traditionally been tackled by incomplete heuristic algorithms. In recent years, approaches based on Constraint Programming (CP) have shown promising results as complete methods for finding optimal mappings, in particular concerning throughput. However, so far none of the available CP approaches consider the tradeoff between throughput and buffer requirements or throughput and power consumption. This article integrates tradeoff awareness into the CP model and introduces a two-step solving approach that utilizes the advantages of heuristics, while still keeping the completeness property of CP. With a number of experiments considering several streaming applications and different platform models, the article illustrates not only the efficiency of the presented model but also its suitability for solving different problems with various combinations of performance constraints.
Kathrin Rosvall, Ingo Sander
ACM Trans. Design Autom. Electr. Syst.2
2017 Automatic construction of models for analytic system-level design space exploration problems
abstract
Due to the variety of application models and also the target platforms used in embedded electronic system design, it is challenging to formulate a generic and extensible analytic design-space exploration (DSE) framework. Current approaches support a restricted class of application and platform models and are difficult to extend. This paper proposes a framework for automatic construction of system-level DSE problem models based on a coherent, constraint-based representation of system functionality, flexible target platforms, and binding policies. Heterogeneous semantics is captured using constraints on logical clocks. The applicability of this method is demonstrated by constructing DSE problem models from different combinations of application and platforms models. Time-triggered and untimed models of the system functionality and heterogeneous target platforms are used for this purpose. Another potential advantage of this approach is that constructed models can be solved using a variety of standard and ad-hoc solvers and search heuristics.
Seyed-Hosein Attarzadeh-Niaki, Ingo Sander
DATE2
2017 A layered formal framework for modeling of cyber-physical systems
abstract
Designing cyber-physical systems is highly challenging due to its manifold interdependent aspects such as composition, timing, synchronization and behavior. Several formal models exist for description and analysis of these aspects, but they focus mainly on a single or only a few system properties. We propose a formal composable framework which tackles these concerns in isolation, while capturing interaction between them as a single layered model. This yields a holistic, fine-grained, hierarchical and structured view of a cyber-physical system. We demonstrate the various benefits for modeling, analysis and synthesis through a typical example.
George Ungureanu, Ingo Sander
DATE2
2017 Designing end-to-end resource reservations in predictable distributed embedded systems
abstract
Contemporary distributed embedded systems in many domains have become highly complex due to ever-increasing demand on advanced computer controlled functionality. The resource reservation techniques can be effective in lowering the software complexity, ensuring predictability and allowing flexibility during the development and execution of these systems. This paper proposes a novel end-to-end resource reservation model for distributed embedded systems. In order to support the development of predictable systems using the proposed model, the paper provides a method to design resource reservations and an end-to-end timing analysis. The reservation design can be subjected to different optimization criteria with respect to runtime footprint, overhead or performance. The paper also presents and evaluates a case study to show the usability of the proposed model, reservation design method and end-to-end timing analysis.
Mohammad Ashjaei, Nima Khalilzad, Saad Mubeen, Moris Behnam, Ingo Sander, Luís Almeida 0001, Thomas Nolte
Real Time Syst.5
2016 CONTREX: Design of Embedded Mixed-Criticality CONTRol Systems under Consideration of EXtra-Functional Properties
abstract
The increasing processing power of today's HW/SW platforms leads to the integration of more and more functions in a single device. Additional design challenges arise when these functions share computing resources and belong to different criticality levels. The paper presents the CONTREX European project and its preliminary results. CONTREX complements current activities in the area of predictable computing platforms and segregation mechanisms with techniques to consider the extra-functional properties, i.e., timing constraints, power, and temperature. CONTREX enables energy efficient and cost aware design through analysis and optimization of these properties with regard to application demands at different criticality levels.
Ralph Görgen, Kim Grüttner, Fernando Herrera, Pablo Peñil, Julio L. Medina, Eugenio Villar, Gianluca Palermo, William Fornaciari, Carlo Brandolese, Davide Gadioli, Sara Bocchio, Luca Ceva, Paolo Azzoni, Massimo Poncino, Sara Vinco, Enrico Macii, Salvatore Cusenza, John M. Favaro, Raúl Valencia, Ingo Sander, Kathrin Rosvall, Davide Quaglia
DSD20
2016 SAFEPOWER Project: Architecture for Safe and Power-Efficient Mixed-Criticality Systems
abstract
With the ever increasing industrial demand for bigger, faster and more efficient systems, a growing number of cores is integrated on a single chip. Additionally, their performance is further maximized by simultaneously executing as many processes as possible not regarding their criticality. Even safety critical domains like railway and avionics apply these paradigms under strict certification regulations. As the number of cores is continuously expanding, the importance of cost-effectiveness grows. One way to increase the cost-efficiency of such System on Chip (SoC) is to enhance the way the SoC handles its power resources. By increasing the power efficiency, the reliability of the SoC is raised, because the lifetime of the battery lengthens. Secondly, by having less energy consumed, the emitted heat is reduced in the SoC which translates into fewer cooling devices. Though energy efficiency has been thoroughly researched, there is no application of those power saving methods in safety critical domains yet. The EU project SAFEPOWER1 targets this research gap and aims to introduce certifiable methods to improve the power efficiency of mixed-criticality real-time systems (MCRTES). This paper will introduce the requirements that a power efficient SoC has to meet and the challenges such a SoC has to overcome.
Alina Lenz, Mikel Azkarate-askatsua, Javier Coronel, Alfons Crespo, Simon Davidmann, Juan Carlos Diaz Garcia, Nera González Romero, Kim Grüttner, Roman Obermaisser, Johnny Öberg, Jon Pérez 0001, Ingo Sander, Ingemar Söderquist
DSD12
2016 A modular design space exploration framework for multiprocessor real-time systems
abstract
Embedded system designers often face a large number of design alternatives when designing complex systems. A designer must select an alternative which satisfies application constraints (e.g. timing requirements) while optimizing system level objectives such as overall energy consumption. The size of design space is often very large giving rise to the need for systematic Design Space Exploration (DSE) methods. In this paper we address the DSE problem for real-time applications that belong to two different domains: (i) streaming applications modeled using the synchronous dataflow graphs; (ii) feedback control tasks modeled using the periodic task model. We consider a heterogeneous multiprocessor platform in which processors communicate through a predictable bus architecture. We present our DSE tool in which the DSE problem is modeled as a constraint satisfaction problem, and it is solved using a constraint programming solver. This approach provides a modular framework in which different constraints such as deadline, throughput and energy consumption can easily be plugged depending on the system being designed.
Nima Khalilzad, Kathrin Rosvall, Ingo Sander
FDL3
2015 Customization of OpenCL applications for efficient task mapping under heterogeneous platform constraints
Edoardo Paone, Francesco Robino, Gianluca Palermo, Vittorio Zaccaria, Ingo Sander, Cristina Silvano
DATE5
2014 A constraint-based design space exploration framework for real-time applications on MPSoCs
abstract
Design space exploration (DSE) is a critical step in the design process of real-time multiprocessor systems. Combining a formal base in form of SDF graphs with predictable platforms providing guaranteed QoS, the paper proposes a flexible and extendable DSE framework that can provide performance guarantees for multiple applications implemented on a shared platform. The DSE framework is formulated in a declarative style as interprocess communication-aware constraint programming (CP) model. Apart from mapping and scheduling of application graphs, the model supports design constraints on several cost and performance metrics, as e.g. memory consumption and achievable throughput. Using constraints with different compliance level, the framework introduces support for mixed criticality in the CP model. The potential of the approach is demonstrated by means of experiments using a Sobel filter, a SUSAN filter, a RASTA-PLP application and a JPEG encoder.
Kathrin Rosvall, Ingo Sander
DATE2
2014 Synthesizing code for GPGPUs from abstract formal models
abstract
Today multiple frameworks exist for elevating the task of writing programs for GPGPUs, which are massively dataparallel execution platforms. These are needed as writing correct and high-performing applications for GPGPUs is notoriously difficult due to the intricacies of the underlying architecture. However, the existing frameworks lack a formal foundation that makes them difficult to use together with formal verification, testing, and design space exploration. We present in this paper a novel software synthesis tool - called f2cc - which is capable of generating efficient GPGPU code from abstract formal models based on the synchronous model of computation. These models can be built using high-level modeling methodologies that hide low-level architecture details from the developer. The correctness of the tool has been experimentally validated on models derived from two applications. The experiments also demonstrate that the synthesized GPGPU code yielded a 28 x speedup when executed on a graphics card with 96 cores and compared against a sequential version that uses only the CPU.
Gabriel Hjort Blindell, Christian Menne, Ingo Sander
FDL3
2014 An extensible infrastructure for modeling and time analysis of predictable embedded systems
abstract
Efficient design of predictable systems on top of multiprocessor-based architectures is challenging. It demands an integration effort to support system models relying on Models-of- Computation (MoC) theory, supporting real-time (RT) analysis and electronic system-level (ESL) design techniques. This paper presents a SystemC-based framework for modelling and time analysis of predictable embedded systems which aims such an integration. The framework has features for system-level design and research of predictable systems. Moreover, the framework is extensible, to enable experts from different communities to explore and assess their contributions, e.g. new schedulers, schedulability analyses, and predictable platform components, without having to rely on a physical platform.
Fernando Herrera, Ingo Sander
FDL2
2013 The RecoBlock SoC platform: a flexible array of reusable run-time-reconfigurable IP-blocks
abstract
Run-time reconfigurable (RTR) FPGAs combine the flexibility of software with the high efficiency of hardware. Still, their potential cannot be fully exploited due to increased complexity of the design process. Consequently, to enable an efficient design flow, we devise a set of prerequisites to increase the flexibility and reusability of current FPGA-based RTR architectures. We apply these principles to design and implement the RecoBlock SoC platform, which main characterization is (1) a RTR plug-and-play IP-Core whose functionality is configured at run-time; (2) flexible inter-block communication configured via software, and (3) built-in buffers to support data-driven streams and inter-process communications. We illustrate the potential of our platform by a tutorial case study using an adaptive streaming application to investigate different combinations of reconfigurable arrays and schedules. The experiments underline the benefits of the platform and shows resource utilization.
Byron Navas, Ingo Sander, Johnny Öberg
DATE2
2013 An automated parallel simulation flow for heterogeneous embedded systems
abstract
Simulation of complex embedded and cyber-physical systems requires exploitation of the computation power of available parallel architectures. Current simulation environments either do not address this parallelism or use separate models for parallel simulation and for analysis and synthesis, which might lead to model mismatches. We extend a formal modeling framework targeting heterogeneous systems with elements that enable parallel simulations. An automated flow is then proposed that starting from a serial executable specification generates an efficient MPI-based parallel simulation model by using a constraint-based method. The proposed flow generates parallel models with acceptable speedups for a representative example.
Seyed-Hosein Attarzadeh-Niaki, Ingo Sander
DATE2
2013 Towards a Modelling and Design Framework for Mixed-Criticality SoCs and Systems-of-Systems
abstract
Mixed-criticality system (MCS) design is an emerging discipline, which has been identified as a core foundational concept in fields such as cyber-physical systems. The hard real-time design community has pioneered the contributions to MCS design, extending scheduling theory to consider mixed-criticalities and the impact of on-chip and off-chip communication infrastructures. However, the development of MCS design methodologies capable to provide safe and efficient solutions for complex applications and platforms in an acceptable design time demands a more interdisciplinary approach. This paper is a first step towards such an approach in the development of MCS design methodologies. The paper first identifies main design disciplines to be involved in MCS design, both at SoC and system-of-systems (SoS) scales. Then, the paper proposes a core ontology for modelling a mixed-criticality system at both SoC scale (MCSoC) and SoS scale (MCSoS). Finally, the paper introduces a set of aspects required for MCS design which have been identified as open and challenging attending the overviewed state-of-the-art.
Fernando Herrera, Seyed-Hosein Attarzadeh-Niaki, Ingo Sander
DSD3
2013 Combining analytical and simulation-based design space exploration for time-critical systems
Fernando Herrera, Ingo Sander
FDL2
2013 Rapid virtual prototyping of real-time systems using predictable platform characterizations
Seyed-Hosein Attarzadeh-Niaki, Marcus Mikulcak, Ingo Sander
FDL3
2012 Integrating virtual platforms into a heterogeneous MoC-based modeling framework
Gilmar S. Beserra, Seyed-Hosein Attarzadeh-Niaki, Ingo Sander
FDL3
2012 Formal heterogeneous system modeling with SystemC
Seyed-Hosein Attarzadeh-Niaki, Mikkel Koefoed Jakobsen, Tero Sulonen, Ingo Sander
FDL4
2012 Performance Analysis of Reconfigurations in Adaptive Real-Time Streaming Applications
abstract
We propose a performance analysis framework for adaptive real-time synchronous data flow streaming applications on runtime reconfigurable FPGAs. As the main contribution, we present a constraint based approach to capture both streaming application execution semantics and the varying design concerns during reconfigurations. With our event models constructed as cumulative functions on data streams, we exploit a novel compile-time analysis framework based on iterative timing phases. Finally, we implement our framework on a public domain constraint solver, and illustrate its capabilities in the analysis of design trade-offs due to reconfigurations with experiments.
Jun Zhu 0011, Ingo Sander, Axel Jantsch
ACM Trans. Embed. Comput. Syst.2
2011 Predicting bus contention effects on energy and performance in multi-processor SoCs
abstract
We present a high-level method for rapidly and accurately predicting bus contention effects on energy and performance in multi-processor SoCs. Unlike most other approaches, which rely on Transaction-Level Modeling (TLM), we infer the information we need directly from executing the algorithmic specification, without needing to build any high-level architectural model. This results in higher estimation speed and allows us to maintain our prediction results within ∼2% of gate-level estimation accuracy.
Sandro Penolazzi, Ingo Sander, Ahmed Hemani
DATE2
2011 Semi-formal refinement of heterogeneous embedded systems by foreign model integration
Seyed-Hosein Attarzadeh-Niaki, Ingo Sander
FDL2
2010 Constrained global scheduling of streaming applications on MPSoCs
abstract
We present a global scheduling framework for synchronous data flow (SDF) streaming applications on MPSoCs, based on optimized computation and contention-free routing. The global scheduling of processors computing and communication transactions are formulated as constraint based problem, to avoid the scheduling overhead in TDMA-like heuristic schemes. A public domain constraint solver is exploited to solve the NP-complete scheduling efficiently, together with problem specific constraint modeling techniques. Experimental results show that the proposed framework can achieve a high predictable application throughput with minimized buffer cost. For instance, for applications in communication domain, higher throughput (up to 87%) has been observed with less buffer cost, compared to scenarios considering the heuristic scheduling overhead.
Jun Zhu 0011, Ingo Sander, Axel Jantsch
ASP-DAC2
2010 Predicting energy and performance overhead of Real-Time Operating Systems
abstract
We present a high-level method for rapidly and accurately estimating energy and performance overhead of Real-Time Operating Systems. Unlike most other approaches, which rely on Transaction-Level Modeling (TLM), we infer the information we need directly from executing the algorithmic specification, without needing to build any high-level architectural model. We distinguish two main components in our approach: first, an accurate one-time pre-characterization of the main RTOS functionalities in terms of energy and cycles; second, the development of an algorithm to rapidly predict the occurrences of such RTOS functionalities. Finally, we demonstrate the feasibility of our approach by comparing it against gate level for accuracy and against TLM for speed. We obtain a worst-case energy error of 12% against a mean speedup of 36X.
Sandro Penolazzi, Ingo Sander, Ahmed Hemani
DATE2
2010 Pareto efficient design for reconfigurable streaming applications on CPU/FPGAs
abstract
We present a Pareto efficient design method for multi-dimensional optimization of run-time reconfigurable streaming applications on CPU/FPGA platforms, which automatically allocates applications with optimized buffer requirement and software/hardware implementation cost. At the same time, application performance is guaranteed with sustainable throughput during run-time reconfigurations. As the main contribution, we formulate the constraint based application allocation, scheduling, and reconfiguration analysis, and propose a design Pareto-point calculation flow. A public domain solver - Gecode is used in solutions finding. The capability of our method has been exemplified by two cases studies on applications from media and communication domains.
Jun Zhu 0011, Ingo Sander, Axel Jantsch
DATE2
2010 HetMoC: Heterogeneous Modelling in SystemC
Jun Zhu 0011, Ingo Sander, Axel Jantsch
FDL2
2009 Buffer minimization of real-time streaming applications scheduling on hybrid CPU/FPGA architectures
abstract
We address the problem of real-time streaming applications scheduling on hybrid CPU/FPGA architectures. The main contribution is a two-step approach to minimize the buffer requirement for streaming applications with throughput guarantees. A novel declarative way of constraint based scheduling for real-time hybrid SW/HW systems is proposed, while the application throughput is guaranteed by periodic phases in execution. We use a voice-band modem application to exemplify the scheduling capabilities of our method. The experimental results show the advantages of our techniques in both less buffer requirement and higher throughput guarantees compared to the traditional PAPS method.
Jun Zhu 0011, Ingo Sander, Axel Jantsch
DATE2
2009 High-level estimation and trade-off analysis for adaptive real-time systems
abstract
We propose a novel design estimation method for adaptive streaming applications to be implemented on a partially reconfigurable FPGA. Based on experimental results we enable accurate design cost estimates at an early design stage. Given the size and computation time of a set of configurations, which can be derived through logic synthesis, our method gives estimates for configuration parameters, such as bitstream sizes, computation and reconfiguration times. To fulfil the system's throughput requirements, the required FIFO buffer sizes are then calculated using a hybrid analysis approach based on integer linear programming and simulation. Finally, we are able to calculate the total design cost as the sum of the costs for the FPGA area, the required configuration memory and the FIFO buffers. We demonstrate our method by analysing non-obvious trade-offs for a static and dynamic implementation of adaptivity.
Ingo Sander, Jun Zhu 0011, Axel Jantsch, Andreas Herrholz, Philipp A. Hartmann, Wolfgang Nebel
IPDPS1
2008 Energy efficient streaming applications with guaranteed throughput on MPSoCs
abstract
In this paper we present a design space exploration flow to achieve energy efficiency for streaming applications on MPSoCs while meeting the specified throughput constraints. The public domain simulators Sim-Panalyzer and Cacti are used to estimate the energy dissipations of the parameterized architectural components. As the main contributions, we schedule the streaming applications on a multi-clock synchronous modeling framework, guarantee the application timing properties by throughput analysis, and customize both processor voltage-frequency levels and memory sizes in the design space to optimize the application pipeline parallelism for energy efficiency. Two widely used heuristic algorithms (i.e., greedy and taboo search) are used during the design optimization process. Our experiments show an energy reduction of 21% without any loss in application throughput compared with an ad-hoc approach.
Jun Zhu 0011, Ingo Sander, Axel Jantsch
EMSOFT2
2008 Application and Verification of Local Nonsemantic-Preserving Transformations in System Design
abstract
Due to the increasing abstraction gap between the initial system model and a final implementation, the verification of the respective models against each other is a formidable task. This paper addresses the verification problem by proposing a stepwise application of combined refinement and verification activities in the context of synchronous model of computation. An implementation model is developed from the system model by applying predefined design transformations which are as follows: 1) semantic preserving or 2) nonsemantic preserving. Nonsemantic-preserving transformations introduce lower level implementation details, which are necessary to yield an efficient implementation. Our approach divides the verification tasks into two activities: 1) the local correctness of a refined block is checked by using formal verification tools and predefined properties, which are developed for each nonsemantic-preserving transformation, and 2) the global influence of the refinement to the entire system is studied through static analysis. We illustrate the design refinement and verification approach with three transformations: 1) a communication refinement mapping a synchronous channel to an asynchronous one including a handshake mechanism; 2) a computation refinement, which introduces resource sharing in a combinational computation block; and 3) a synchronization demanding refinement, where an algorithm analyzes the influence of a local refinement to the temporal properties of the entire system and restores the system's correct temporal behavior if necessary.
Tarvo Raudvere, Ingo Sander, Axel Jantsch
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 The ANDRES Project: Analysis and Design of Run-Time Reconfigurable, Heterogeneous Systems
abstract
Today's heterogeneous embedded systems combine components from different domains, such as software, analogue hardware and digital hardware. The design and implementation of these systems is still a complex and error-prone task due to the different Models of Computations (MoCs), design languages and tools associated with each of the domains. Though making such systems adaptive is technologically feasible, most of the current design methodologies do not explicitely support adaptive architectures. This paper present the ANDRES project. The main objective of ANDRES is the development of a seamless design flow for adaptive heterogeneous embedded systems (AHES) based on the modelling language SystemC. Using domain-specific modelling extensions and libraries, ANDRES will provide means to efficiently use and exploit adaptivity in embedded system design. The design flow is completed by a methodology tools for automatic hardware and software synthesis for adaptive architectures.
Andreas Herrholz, Frank Oppenheimer, Philipp A. Hartmann, Andreas Schallenberg, Wolfgang Nebel, Christoph Grimm 0001, Markus Damm, Jan Haase 0001, Florian Brame, Fernando Herrera, Eugenio Villar, Ingo Sander, Axel Jantsch, Anne-Marie Fouilliart, Marcos Martínez
FPL12
2007 A synchronization algorithm for local temporal refinements in perfectly synchronous models with nested feedback loops
abstract
Due to the abstract and simple computation and communication mechanism in the synchronous computational model it is easy to simulate synchronous systems and to apply formal verification methods. In synchronous models, a local temporal refinement that increases the delay in a single computation block may affect the functionality of the entire model. To preserve the system's functionality after temporal refinements we provide a synchronization algorithm that applies also to models with nested feedback loops. The algorithm adds pure delay elements to the model in order to balance the delay caused by refinement and to assure concurrent data arrival at computation blocks. It is done so that the refined model stays latency equivalent to the original model. The advantages of our approach are that (a) we remain fully within the synchronous model of computation, (b) we preserve the functionality of the existing computation blocks, and (c) we do not require additional computation resources, wrapper circuits or schedulers.
Tarvo Raudvere, Ingo Sander, Axel Jantsch
ACM Great Lakes Symposium on VLSI2
2006 Towards Performance-Oriented Pattern-Based Refinement of Synchronous Models onto NoC Communication
abstract
We present a performance-oriented refinement approach that refines a perfectly synchronous communicationmodel onto Network-on-Chip (NoC) communication. We first identify four basic forms of NoC process interaction patterns at the process level, namely, producer-consumer, peers, client-server, and multicast. We propose a threestep top-down refinement method: channel refinement, protocol refinement and channel mapping. For the producer-consumer pattern, we describe it in detail. In channel refinement, we deal with interfacing multiple clock domains and use a stochastic process to model channel delay and jitter. In protocol refinement, we show how to refine communication towards application requirements such as reliability and throughput. In channel mapping, we discuss channel convergence and channel merge arising from channel overlapping. All the refinements have been conducted and validated as an integral design phase towards implementation in ForSyDe, a formal system-level design methodology based on a synchronous model of computation.
Zhonghai Lu, Ingo Sander, Axel Jantsch
DSD2
2006 Flexible Bus and NoC Performance Analysis with Configurable Synthetic Workloads
abstract
We present a flexible method for bus and network on chip performance analysis, which is based on the adaptation of workload models to resemble various applications. Our analysis method assists in the selection of a communication infrastructure early in the design process. The method uses (1) synthetic workload models which are similar to timed Petri nets and (2) the b-model for self-similar workloads. This allows the exploration of larger portions of the design space than possible with traditional stochastic models. The method is illustrated with tutorial examples where both a NoC and a bus based platform are analyzed
Rikard Thid, Ingo Sander, Axel Jantsch
DSD2
2005 Feasibility analysis of messages for on-chip networks using wormhole routing
abstract
The feasibility of a message in a network concerns if its timing property can be satisfied without jeopardizing any messages already in the network to meet their timing properties. We present a novel feasibility analysis for real-time (RT) and nonreal-time (NT) messages in wormhole-routed networks on chip. For RT messages, we formulate a contention tree that captures contentions in the network. For coexisting RT and NT messages, we propose a simple bandwidth partitioning method that allows us to analyze their feasibility independently.
Zhonghai Lu, Axel Jantsch, Ingo Sander
ASP-DAC3
2005 Refinement of Perfectly Synchronous Communication Model
Zhonghai Lu, Ingo Sander, Axel Jantsch
FDL2
2005 System level verification of digital signal processing applications based on the polynomial abstraction technique
abstract
Polynomial abstraction has been developed for data abstraction of sequential circuits, where the functionality can be expressed as polynomials. The method, based on the fundamental theorem of algebra, abstracts a possibly infinite domain of input values, into a much smaller and finite one, whose size is calculated according to the degree of the respective polynomial. The abstract model preserves the system's control and data properties, which can be verified by model checking. Experiments show that our approach does not only allow an automatic verification, but also gives considerably better results than existing methods.
Tarvo Raudvere, Ashish Kumar Singh, Ingo Sander, Axel Jantsch
ICCAD3
2004 Polynomial Abstraction for Verification of Sequentially Implemented Combinational Circuits
abstract
Today's integrated circuits with increasing complexity cause the well known state space explosion problem in verification tools. In order to handle this problem a much simpler abstract model of the design has to be created for verification. We introduce the polynomial abstraction technique, which efficiently simplifies the verification task of sequential design blocks whose functionality can be expressed as a polynomial. Through our technique, the domains of possible values of data input signals can be reduced. This is done in such a way that the abstract model is still valid for model checking of the design functionality in terms of the system's control and data properties. We incorporate polynomial abstraction into the ForSyDe methodology, for the verification of clock domain design refinements.
Tarvo Raudvere, Ashish Kumar Singh, Ingo Sander, Axel Jantsch
DATE3
2004 System modeling and transformational design refinement in ForSyDe [formal system design]
abstract
The scope of the formal system design (ForSyDe) methodology is high-level modeling and refinement of systems-on-a-chip and embedded systems. Starting with a formal specification model, that captures the functionality of the system at a high abstraction level, it provides formal design-transformation methods for a transparent refinement process of the system model into an implementation model that is optimized for synthesis. The main contribution of this paper is the ForSyDe modeling technique and the formal treatment of transformational design refinement. We introduce process constructors, that cleanly separate the computation part of a process from the synchronization and communication part. We develop the characteristic function for each process type and use it to define semantic preserving and design decision transformations. In a study of a digital equalizer example, we illustrate the modeling and refinement process and focus in particular on refinement of the clock domain, communication refinement, and resource sharing.
Ingo Sander, Axel Jantsch
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2003 Development and Application of Design Transformations in ForSyDe
Ingo Sander, Axel Jantsch, Zhonghai Lu
DATE1
2002 Transformation based communication and clock domain refinement for system design
abstract
The ForSyDe methodology has been developed for system level design. In this paper we present formal transformation methods for the refinement of an abstract and formal system model into an implementation model. The methodology defines two classes of design transformations: (1) semantic-preserving transformations and (2) design decisions. In particular we present and illustrate communication and clock domain refinement by way of a digital equalizer system.
Ingo Sander, Axel Jantsch
DAC1