Rainer Dömer

dblp:35/2352 · DBLP profile ↗
← Back
52ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0001-9586-4861ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 45 · 5 first-author · 3 since 2021Software engineering, systems software and programming languages · 16 · 4 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Mixed-Level Modeling and Evaluation of a Cache-less Grid of Processing Cells
abstract
Modern processors experience memory contention when the speed of their computational units exceeds the rate at which new data is available to be processed. This phenomenon is well known as the memory wall and is a great challenge in computer engineering. The reason for this phenomenon is the unequal growth rate in memory access speeds compared to processor clock rates. To mitigate the memory bottleneck in classic computer architectures, a scalable parallel computing platform—the Grid of Processing Cells (GPC)—has been proposed. To evaluate its effectiveness, the GPC is modeled at the instruction level and functional level using SystemC TLM-2.0, with a focus on memory contention. Individual GPC cells can be switched between the two abstraction levels. Our mixed-level system model enables fast and accurate simulations. We test multiple streaming applications on the GPC, and analyze software-based optimization methods and their effects on the GPC, at both abstraction levels. The performance is then compared against the traditional shared memory processor architecture. Experimental results show improved execution times on the GPC primarily due to a large decrease in main memory contention.
Vivek Govindasamy, Rainer Dömer
ACM Trans. Embed. Comput. Syst.2
2025 A Quantitative Guide to Navigate Speed/Accuracy Tradeoffs in System Level Design of RISC-V Processor Grids
abstract
Even when following a well-structured top-down design methodology, system designers regularly face obstacles and pitfalls posed by speed/accuracy tradeoffs in modeling and simulation of complex hardware and software systems. Across the abstraction levels, modeling details grow exponentially, while simulation speed decreases by multiple orders of magnitude. To quantify these effects, we systematically generate, simulate, and evaluate grid-based systems-on-chip in a top-down open-source based tool flow. We map two parallel software applications onto a scalable grid of RISC-V processors and successively refine and validate the models at lower abstraction levels, namely TLM, ISS, RTL, and FPGA. Our comprehensive experimental evaluation over five abstraction levels quantifies the speed-accuracy tradeoffs in simulator build and run times. In addition to its educational value, our work can guide the system designer on an efficient path to a cycle-accurate software simulation on fully constructed hardware.
Lars Luchterhandt, Vivek Govindasamy, Christoph Scheytt, Wolfgang Müller 0003, Rainer Dömer
FDL6
2024 BusyMap, an Efficient Data Structure to Observe Interconnect Contention in SystemC TLM-2.0
abstract
For designing embedded computer architectures that meet desired performance constraints at low cost, fast and accurate simulation models are needed early in the design flow. To identify and avoid bottlenecks early at the system level, observing the contention of shared resources is critical. In this paper, we propose and evaluate a novel data structure called BusyMap that accurately reflects contention at system busses or similar inter-connect components. BusyMap is an efficient data structure that allows the system designer to accurately model and easily observe contention in IEEE loosely-timed TLM-2.0. In contrast to prior state-of-the-art, our model fully supports temporal decoupling and multiple levels of interconnect. Our experiments demonstrate the effectiveness of BusyMap with results showing high accuracy at high -speed System C simulation.
Emad Malekzadeh Arasteh, Vivek Govindasamy, Rainer Dömer
DATE3
2024 Fast Loosely-Timed Deep Neural Network Models with Accurate Memory Contention
abstract
The emergence of data-intensive applications, such as Deep Neural Networks (DNN), exacerbates the well-known memory bottleneck in computer systems and demands early attention in the design flow. Electronic System-Level (ESL) design using SystemC Transaction Level Modeling (TLM) enables effective performance estimation, design space exploration (DSE), and gradual refinement. However, memory contention is often only detectable after detailed TLM-2.0 approximately-timed or cycle-accurate RTL models are developed. A memory bottleneck detected at such a late stage can severely limit the available design choices or even require costly redesign. In this work, we propose a novel TLM-2.0 loosely-timed contention-aware (LT-CA) modeling style that offers high-speed simulation close to traditional loosely-timed (LT) models, yet shows the same accuracy for memory contention as low-level approximately-timed (AT) models. Thus, our proposed LT-CA modeling breaks the speed/accuracy tradeoff between regular LT and AT models and offers fast and accurate observation and visualization of memory contention. Our extensible SystemC model generator automatically produces desired TLM-1 and TLM-2.0 models from a DNN architecture description for design space exploration focusing on memory contention. We demonstrate our approach with a real-world industry-strength DNN application, GoogLeNet. The experimental results show that the proposed LT-CA modeling is 46× faster in simulation than equivalent AT models with an average error of less than 1% in simulated time. Early detection of memory contentions also suggests that local memories close to computing cores can eliminate memory contention in such applications.
Emad Malekzadeh Arasteh, Rainer Dömer
ACM Trans. Embed. Comput. Syst.2
2023 Instruction-Level Modeling and Evaluation of a Cache-Less Grid of Processing Cells
abstract
While processor speeds continue to show performance increases, memory access speeds remain significantly slower. One solution is to develop novel computer architectures, specifically designed to address the memory wall such as the cache-less Grid of Processing Cells (GPC). In this work, we model the GPC architecture using SystemC TLM-2.0 at the instruction-level with high accuracy, while retaining the characteristic high simulation speeds of transaction-level modeling. To compare the performance of our modeled GPC with existing architectures, we also model a single core and an 8-core SMP with optimized caches and run a bare-metal Canny application on the three architectures. We show promising performance evaluation results in architectures without caches.
Vivek Govindasamy, Rainer Dömer
FDL2
2021 Improving Parallelism in System Level Models by Assessing PDES Performance
abstract
For effective embedded system design, transaction level modeling (TLM) must explicitly expose any available parallelism in the application. Traditional TLM in SystemC utilizes channels for communication and synchronization between concurrent modules, whereas modern TLM-2.0 emphasizes address-accurate communication via explicit interconnect and memories. In both modeling styles, the choice of synchronization mechanisms has a significant impact on the available parallelism in the model which can be exploited by parallel discrete event simulation (PDES). In this work, we propose and analyze a set of non-invasive standard-compliant modeling techniques to increase parallelism in IEEE SystemC TLM-1 and TLM-2.0 models. We measure the performance of aggressive out-of-order PDES in the Recoding Infrastructure for SystemC (RISC) and analyze the parallelism in the models. Our case study on six modeling styles of a state-of-art deep neural network (DNN), namely the GoogLeNet image classification algorithm, demonstrates the impact of varying synchronization mechanisms with simulator run time reduced by 38% compared to a synchronous parallel reference model on a 16-core host machine. Our study also suggests that increased parallel simulation performance indicates better models with higher amounts of parallelism exposed.
Emad Malekzadeh Arasteh, Rainer Dömer
FDL2
2020 Event Delivery using Prediction for Faster Parallel SystemC Simulation
abstract
Out-of-order Parallel Discrete Event Simulation (OoO PDES) is an advanced simulation approach that efficiently verifies and validates SystemC models. To preserve the simulation semantics, OoO PDES performs a conservative event delivery strategy which often postpones the execution of waiting threads due to unknown future behaviors of the model. In this paper, based on predicted behaviors of threads, we introduce a novel event delivery strategy that allows waiting threads to resume execution earlier, resulting in significantly increased simulation speed. Experimental results show that the proposed approach increases the OoO PDES simulation speed by up to 4.9x compared to the original one on a 4-core machine.
Zhongqi Cheng, Emad Malekzadeh Arasteh, Rainer Dömer
ASP-DAC3
2020 Lazy Event Prediction using Defining Trees and Schedule Bypass for Out-of-Order PDES
abstract
Out-of-order parallel discrete event simulation (PDES) has been shown to be very effective in speeding up system design by utilizing parallel processors on multi- and many-core hosts. As the number of threads in the design model grows larger, however, the original scheduling approach does not scale. In this work, we analyze the out-of-order scheduler and identify a bottleneck with quadratic complexity in event prediction. We propose a more efficient lazy strategy based on defining trees and a schedule bypass with O(m log2m) complexity which shows sustained and improved performance gains in simulation of SystemC models with many processes. For models containing over 1000 processes, experimental results show simulation run time speedups of up to 90× using lazy event prediction against the original out-of-order PDES approach.
Daniel Mendoza, Zhongqi Cheng, Emad Malekzadeh Arasteh, Rainer Dömer
DATE4
2019 Analyzing Variable Entanglement for Parallel Simulation of SystemC TLM-2.0 Models
abstract
The SystemC TLM-2.0 standard is widely used in modern electronic system level design for better interoperability and higher simulation speed. However, TLM-2.0 has been identified as an obstacle for parallel SystemC simulation due to the disappearance of channels. Without a containment construct, simulation threads are permitted to directly access data of other modules and that makes it difficult to synchronize such accesses as required by the SystemC execution semantics. In this paper, we propose a compile time approach to statically analyze potential conflicts among threads in SystemC TLM-2.0 loosely- and approximately-timed models. We introduce a new Socket Call Path technique which provides the compiler with socket binding information for precise static analysis. We also propose an algorithm to analyze entangled variable pairs. Experimental results show that our approach is able to support automatically safe parallel simulation of SystemC models with TLM-2.0 Blocking Transport Interface, Direct Memory Interface and Non-blocking Transport Interface, resulting in impressive simulation speeds.
Zhongqi Cheng, Rainer Dömer
ACM Trans. Embed. Comput. Syst.2
2018 Port call path sensitive conflict analysis for instance-aware parallel SystemC simulation
abstract
Many parallel SystemC approaches expect a thread safe and conflict free model from the designer. Alternatively, an advanced compiler can identify and avoid possible parallel access conflicts. While manual conflict resolution can theoretically be more precise, it is impractical for real world applications because of the inherent complexities. Here automatic compiler-based analysis is preferred which provides conservative conflict avoidance with minimal false positives. This paper introduces a novel compiler technique called Port Call Path analysis that greatly reduces the amount of false positive conflicts resulting in significantly increased simulation speed. Experimental results show that the new analysis reduces the amount of false conflicts by up to 98% and, on a 4-core processor, speeds up the simulation up to 3× for a NoC particle simulator and 3.5× for a bitcoin miner SystemC model.
Tim Schmidt, Zhongqi Cheng, Rainer Dömer
DATE3
2018 SystemC Coding Guideline for Faster Out-of-order Parallel Discrete Event Simulation
abstract
IEEE SystemC is one of the most popular standards for system level design. With the Recoding Infrastructure for SystemC (RISC), a SystemC model can be executed at segment level in parallel. Although the parallel simulation is generally faster than its sequential counterpart, any data conflict among segments reduces the simulation speed significantly. In this paper, we propose for RISC users a coding guideline that increases the granularity of segments, so that the level of parallelism in the design increases and higher simulation speed becomes possible. Our experimental results show that a maximum speedup of over 6.0x is achieved on an 8-core processor, which is 1.7 times faster than parallel simulation without the coding guideline.
Zhongqi Cheng, Tim Schmidt, Rainer Dömer
FDL3
2017 Hybrid analysis of SystemC models for fast and accurate parallel simulation
abstract
Parallel SystemC approaches expect a thread-safe and race-condition-free model from the designer or use a compiler which identifies the race conditions. However, they have strong limitations for real world examples. Two major obstacles remain: a) all the source code must be available and b) the entire design must be statically analyzable. In this paper, we propose a solution for a fast and fully accurate parallel SystemC simulation which overcomes these two obstacles a) and b). We propose a hybrid approach which includes both static and dynamic analysis of the design model. We also handle library calls in the compiler analysis where the source code of the library functions is not available. Our experiments demonstrate a 100% accurate execution and a speedup of 6.39× for a Network-on-Chip particle simulator.
Tim Schmidt, Guantao Liu, Rainer Dömer
ASP-DAC3
2017 Exploiting Thread and Data Level Parallelism for Ultimate Parallel SystemC Simulation
abstract
Most parallel SystemC approaches have two limitations: (a) the user must manually separate all parallel threads to avoid data corruption due to race conditions, and (b) available hardware vector units are not utilized. In this paper, we present an advanced compiler infrastructure for automatic parallelization of SystemC models at the thread-level. In addition, our infrastructure exploits opportunities for data-level parallelization. Our experimental results show a nearly linear speedup of NxM, where N and M denote the thread and data-level factors, respectively. In turn, a 4-core multi-processor achieves a speedup of up to 8.8x, and a 60-core Xeon Phi processor reaches up to 212x.
Tim Schmidt, Guantao Liu, Rainer Dömer
DAC3
2016 Automatic Generation of Thread Communication Graphs from SystemC Source Code
abstract
In an ideal top-down system design flow, graphical diagrams are designed before an executable specification in a System Level Description Language (SLDL) is derived. Such initial charts typically also serve as visual documentation of the textual specification and aid in maintaining the model. In the absence of graphical charts, e.g. in case of legacy or 3rd party code, a textual SLDL model is hard to comprehend for any unfamiliar designer. Here, we propose to automatically extract graphical charts from given SystemC code to ease the understanding of the source code with a visual representation. Specifically, we extract the communication flow between the threads from the design model by use of an automatic SystemC compiler infrastructure that statically analyzes the code and generates custom Thread Communication Graphs (TCG) similar to message sequence charts. Our experimental results on embedded applications demonstrate that our novel static analysis can quickly extract accurate TCG that are highly useful for designers in becoming familiar with new source code.
Tim Schmidt, Guantao Liu, Rainer Dömer
SCOPES3
2015 Communication protocol analysis of transaction-level models using Satisfiability Modulo Theories
abstract
A critical aspect in SoC design is the correctness of communication between system blocks. In this work, we present a novel approach to formally verify various aspects of communication models, including timing constraints and liveness. Our approach automatically extracts timing relations and constraints from the design and builds a Satisfiability Modulo Theories (SMT) model whose assertions are then formally verified along with properties of interest input by the designer. Our method also addresses the complexity growth with a hierarchical approach. We demonstrate our approach on models communicating over industry standard bus protocol AMBA AHB and CAN bus. Our results show that the generated assertions can be solved within resonable time.
Rainer Dömer
ASP-DAC2
2015 Optimizing thread-to-core mapping on manycore platforms with distributed Tag Directories
abstract
With the increasing demand for parallel computing power, manycore platforms are attracting more and more attention due to their potential to improve performance and scalability of parallel applications. However, as the core count increases, core-to-core communication becomes expensive. For manycore architectures using directory-based cache coherence protocols, the core-to-core communication latency depends not only on the physical placement on the chip, but also on the location of the distributed cache tag directory. In this paper, we first define the concept of core distance for multicore and manycore architectures. Using a ping-pong spin-lock benchmark, we quantify the core distance on a ring-network platform and propose an approach to optimize thread-to-core mapping in order to minimize on-chip communication overhead. In our experiments, our approach speeds up communication-intensive benchmarks by more than 25% on average over the Linux default mapping strategy.
Guantao Liu, Tim Schmidt, Rainer Dömer, Ajit Dingankar, Desmond Kirkpatrick
ASP-DAC3
2015 May-happen-in-parallel analysis of ESL models using UPPAAL model checking
Rainer Dömer
DATE2
2014 May-happen-in-parallel analysis based on segment graphs for safe ESL models
abstract
A well-defined system-level model contains explicit parallelism and should be free from parallel access conflicts to shared variables. However, safe parallelism is difficult to achieve since risky shared variables are often hidden deep in the design and are not exposed through simulation. In this paper, we propose a new static analysis approach based on segment graphs that identifies a tight set of potential access conflicts in segments that may-happen-in-parallel (MHP). Our experimental results show that the analysis is complete, accurate and fast to reveal dangerous shared variables in several embedded application models. Compared to earlier work, our approach significantly reduces the number of false conflict reports and thus saves the designer time.
Weiwei Chen 0001, Xu Han 0002, Rainer Dömer
DATE3
2014 Powermonitor: a versatile API for automated power-aware ESL design
abstract
System Level Description Languages (SLDL) emerged a decade ago for high-level System-on-Chip (SoC) design and efficient design space exploration. Initially performance and area constraints were the major concerns. Nowadays the shrinking of transistor size has brought power and temperature to top of the list of designer concerns. Although SLDLs are welldefined for functional and timing modeling, the dimension for power- and temperature-aware modeling is missing and needs to be added manually by the user in an ad-hoc fashion. In this paper we introduce; PowerMonitor as an Application Programming Interface (API) capable of (a) power annotation in the executable model, (b) power-aware simulation, and (c) monitoring and graphically representing power dissipation at the system level with flexible granularity and compositions.
Yasaman Samei Syahkal, Rainer Dömer
FDL2
2014 Automated estimation of power consumption for rapid system level design
abstract
This paper describes an early power estimation method for Electronic System Level(ESL) design, which provides a scalable API to support automated power profiling and analysis at the early stages of the design process. The proposed framework utilizes a high-level power modeling mechanism along with an automated profiler to extract energy activity from the simulated system model. These two features are integrated into PowerMeter, a framework that automatically annotates power meters as well as energy and performance functions into the executable model. This integrated profiling helps the designer to rapidly explore the design space, trading off performance against power cost in order to make best design decisions. Our approach also provides the designer with the ability to quantify the effect of revisions in the ESL design models, in terms of both power and performance. Despite the high abstraction level, our results show that the PowerMeter delivers rapid estimates with high fidelity and at minimal cost.
Yasaman Samei Syahkal, Rainer Dömer
IPCCC2
2014 Out-of-Order Parallel Discrete Event Simulation for Transaction Level Models
abstract
The validation of system models at the transaction-level typically relies on discrete event (DE) simulation. In order to reduce simulation time, parallel discrete event simulation (PDES) can be used by utilizing multiple cores available on today's host PCs. However, the total order of time imposed by regular DE simulators becomes a bottleneck that severely limits the benefits of parallel simulation. In this paper, we present a new out-of-order (OoO) PDES technique for simulating transaction-level models on multicore hosts. By localizing the simulation time to individual threads and carefully handling events at different times, a system model can be simulated following a partial order of time without loss of accuracy. Subject to advanced static analysis at compile time and table-based decisions at run time, threads can be issued early, reducing the idle time of available cores. Our proposed OoO PDES technique shows high performance gains in simulation speed with only a small increase in compile time. Using six embedded application examples, we also show the speed trade-off for multicore PDES based on different multithreading libraries.
Weiwei Chen 0001, Xu Han 0002, Guantao Liu, Rainer Dömer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2013 Optimized out-of-order parallel discrete event simulation using predictions
abstract
Parallel Discrete Event Simulation (PDES) enables efficient validation of ESL models on multi-core simulation hosts. Out-of-order PDES is an advanced scheduling technique which allows multiple threads to run in parallel even in different simulation cycles. To maintain simulation semantics and timing accuracy, the compiler performs complex static conflict analysis so that the scheduler can make quick and safe decisions at run time and issue threads early. Often, however, out-of-order scheduling is prevented because of the unknown future behavior of the threads. In this paper, we extend the analysis in order to predict the future of candidate threads. Looking ahead of the current simulation state allows the scheduler to issue more threads in parallel, resulting in significantly reduced simulator run time. Our experimental results show simulation speedup up to 1.92x with only negligible increase in compile time.
Weiwei Chen 0001, Rainer Dömer
DATE2
2012 An optimizing compiler for out-of-order parallel ESL simulation exploiting instance isolation
abstract
Electronic system-level (ESL) design relies on fast discrete event (DE) simulation for the validation of design models written in system-level description languages (SLDLs). An advanced technique to speedup ESL validation is out-of-order parallel DE simulation which allows multiple threads to run early and in parallel on multi-core hosts. To avoid data hazards and ensure timing accuracy, this technique requires the compiler to statically analyze the design model for potential data access conflicts. In this paper, we propose a compiler optimization that improves the data conflict analysis by exploiting instance isolation. The reduction in the number of conflicts increases the available parallelism and results in significantly reduced simulation time. Our experimental results show up to 90% gain in simulation speed for less than 6% increase in compilation time.
Weiwei Chen 0001, Rainer Dömer
ASP-DAC2
2012 Parallel discrete event simulation of Transaction Level Models
abstract
Describing Multi-Processor Systems-on-Chip (MPSoC) at the abstract Electronic System Level (ESL) is one task, validating them efficiently is another. Here, fast and accurate system-level simulation is critical. Recently, Parallel Discrete Event Simulation (PDES) has gained significant attraction again as it promises to utilize the existing parallelism in today's multicore CPU hosts. This paper discusses the parallel simulation of Transaction-Level Models (TLMs) described in System-Level Description Languages (SLDLs), such as SystemC and SpecC.We review how PDES exploits the explicit parallelism in the ESL design models and uses the parallel processing units available on multicore host PCs to significantly reduce the simulation time. We show experimental results for two highly parallel benchmarks as well as for two actual embedded applications.
Rainer Dömer, Weiwei Chen 0001, Xu Han 0002
ASP-DAC1
2012 Out-of-order parallel simulation for ESL design
abstract
At the Electronic System Level (ESL), design validation often relies on discrete event (DE) simulation. Recently, parallel simulators have been proposed which increase simulation speed by using multiple cores available on today's PCs. However, the total order of time in DE simulation is a bottleneck that severely limits the benefits of parallel simulation. This paper presents a new out-of-order simulator for multi-core parallel DE simulation of hardware/software designs at any abstraction level. By localizing the simulation time and carefully handling events at different times, a system model can be simulated following a partial order of time. Subject to automatic static data analysis at compile time and table-based decisions at run time, threads can be issued early which reduces the idle time of available cores. Our experiments show high performance gains in simulation speed with only a small increase of compile time.
Weiwei Chen 0001, Xu Han 0002, Rainer Dömer
DATE3
2012 Computer-Aided Recoding to Create Structured and Analyzable System Models
abstract
In embedded system design, the quality of the input model has a direct bearing on the effectiveness of the system exploration and synthesis tools. Given a well-written system model, tools today are effective in generating working implementations. However, readily available C reference code is not conducive for immediate system synthesis as it lacks needed features for automatic analysis and synthesis. Among others, the lack of proper structure and the presence of intractable pointers in the reference code are factors that seriously hamper the effectiveness of system design tools. To overcome these deficiencies, we aim to automate the conversion of flat C code into a well-structured system model by applying automated source code transformations. We present a set of computer-aided recoding operations that enable the system designer to mitigate pointer problems and quickly create the necessary structural hierarchy so that the design model becomes easily analyzable and synthesizable. Utilizing the designer’s knowledge, our interactive recoding transformations aid the designer in efficiently creating well-structured system models for rapid design space exploration and successful synthesis. Our estimated and measured experimental results show significant productivity gains through a substantial reduction of the model creation time.
Pramod Chandraiah, Rainer Dömer
ACM Trans. Embed. Comput. Syst.2
2011 Multi-core parallel simulation of System-level Description Languages
abstract
The validation of transaction level models described in System-level Description Languages (SLDLs) often relies on extensive simulation. However, traditional Discrete Event (DE) simulation of SLDLs is cooperative and cannot utilize the available parallelism in modern multi-core CPU hosts. In this work, we study the SLDL execution semantics of concurrent threads and present a multi-core parallel simulation approach which automatically protects communication between concurrent threads so that parallel simulation on multi-core hosts becomes possible. We demonstrate significant speed-up in simulation time of several system models, including a H.264 video decoder and a JPEG encoder.
Rainer Dömer, Weiwei Chen 0001, Xu Han 0002, Andreas Gerstlauer
ASP-DAC1
2010 A fast heuristic scheduling algorithm for periodic ConcurrenC models
abstract
Embedded system design usually starts from an executable specification model described in a C-based System Level Description Language (SLDL), such as SystemC or SpecC. In this paper, we identify a subset of well-defined C-based design models, called periodic ConcurrenC models, that can be statically scheduled, resulting in significant higher simulation and execution speed. We propose a novel heuristic scheduling algorithm that not only is faster than classic matrix-based synchronous dataflow (SDF) scheduling approaches, but also reduces the model execution time by an order of magnitude over the default discrete event simulation.
Weiwei Chen 0001, Rainer Dömer
ASP-DAC2
2010 Computer-aided recoding for multi-core systems
abstract
The design of embedded computing systems faces a serious productivity gap due to the increasing complexity of their hardware and software components. One solution to address this problem is the modeling at higher levels of abstraction. However, manually writing proper executable system models is challenging, error-prone, and very time-consuming. We aim to automate critical coding tasks in the creation of system models. This paper outlines a novel modeling technique called computer-aided recoding which automates the process of writing abstract models of embedded systems by use of advanced computer-aided design (CAD) techniques. Using an interactive, designer-controlled approach with automated source code transformations, our computer-aided recoding technique derives an executable parallel system model directly from available sequential reference code. Specifically, we describe three sets of source code transformations that create structural hierarchy, expose potential parallelism, and create explicit communication and synchronization. As a result, system modeling is significantly streamlined. Our experimental results demonstrate the shortened design time and higher productivity.
Rainer Dömer
ASP-DAC1
2010 System-level development of embedded software
abstract
Embedded software plays an increasingly important role in implementing modern embedded systems. Development of embedded software, and of hardware-dependent software in particular, is challenging due to the tight integration with the underlying hardware architecture. In this paper, we describe our system-level design approach that allows designers to develop software in form of a platform-agnostic specification. Our design environment enables exploration of different architectural alternatives and subsequently generates the software implementation. It generates the application code, communication drivers, and an adaptation to a chosen RTOS. It completes the process by producing the final target binary for each processor. Our experimental results demonstrate the automatic generation of the binaries for five control and media oriented applications.
Gunar Schirner, Andreas Gerstlauer, Rainer Dömer
ASP-DAC3
2010 Panel Session - Who Is Closing the embedded software design gap?
Wolfgang Ecker, Pierre Bricaud, Rainer Dömer, Yossi Veller, Stefan Heinen, Jürgen Mössinger, Andreas von Schwerin
DATE3
2010 Fast and accurate processor models for efficient MPSoC design
abstract
With growing system complexity and ever-increasing software content, the development of embedded software for upcoming MPSoC architectures is a tremendous challenge. Traditional ISS-based validation becomes infeasible due to the large complexity. Addressing the need for flexible and fast simulating models, we introduce in this article our approach of abstract processor modeling in the context of multiprocessor architectures. We combine modeling of computation on processors with an abstract RTOS and accurate interrupt handling into a versatile, multifaceted processor model with several levels of features. Our processor models are utilized in a framework allowing designers to develop a system in a top-down manner using automatic model generation and compilation down to a given MPSoC architecture. During generation, instances of our processor models are integrated into a system model combining software, hardware, and bus communication. The generated system model serves for rapid design space exploration and a fast and accurate system validation. Our experimental results show the benefits of our processor modeling using an actual multiprocessor mobile phone baseband platform. Our abstract models of this complex system reach a simulation speed of 300MCycles/s within a high accuracy of less than 3% error. In addition, our results quantify the speed/accuracy trade-off at varying abstraction levels of our models to guide future processor model designers.
Gunar Schirner, Andreas Gerstlauer, Rainer Dömer
ACM Trans. Design Autom. Electr. Syst.3
2009 Introduction to hardware-dependent software design hardware-dependent software for multi- and many-core embedded systems
abstract
Due to the rapidly increasing software content in embedded systems, hardware-dependent software (HdS) has become a critical topic in system design. In this paper, we provide a brief overview on the topic of HdS, discuss the issues and complexities involved in the design of HdS, and motivate the need for special attention to HdS in research and development.
Rainer Dömer, Andreas Gerstlauer, Wolfgang Müller 0003
ASP-DAC1
2009 Programming MPSoC platforms: Road works ahead!
Rainer Leupers, András Vajda, Marco Bekooij, Soonhoi Ha, Rainer Dömer, Achim Nohl
DATE5
2008 Automatic re-coding of reference code into structured and analyzable SoC models
abstract
The quality of the input system model has a direct bearing on the effectiveness of the system exploration and synthesis tools. Given a well-structured system model, tools today are effective in generating efficient implementations. However, readily available reference C codes are not conducive for system synthesis as they lack the necessary structure and analyzability needed by the design flow. Usually reference C code is manually converted into a SoC model by applying necessary transformations. The type of transformations depends on the underlying design flow and tools. Proper structural hierarchy is one essential feature needed for architectural exploration. In this paper, we provide automatic C code transformations to encapsulate functions and insert structural hierarchy to create well-structured and analyzable SoC models. Our automatic transformations, combined with interactive application of the designer's knowledge and experience, enable faster creation of structural hierarchy in C models and hence result in significant reduction of the overall design time.
Pramod Chandraiah, Rainer Dömer
ASP-DAC2
2008 Automatic generation of hardware dependent software for MPSoCs from abstract system specifications
abstract
Increasing software content in embedded systems and SoCs drives the demand to automatically synthesize software binaries from abstract models. This is especially critical for Hardware dependent Software (HdS) due to the tight coupling. In this paper, we present our approach to automatically synthesize HdS from an abstract system model. We synthesize driver code, interrupt handlers and startup code. We furthermore automatically adjust the application to use RTOS services. We target traditional RTOS-based multi-tasking solutions, as well as a pure interrupt-based implementation (without any RTOS). Our experimental results show the automatic generation of final binary images for six real-life target applications and demonstrate significant productivity gains due to automation. Our HdS synthesis is an enabler for efficient MPSoC development and rapid design space exploration.
Gunar Schirner, Andreas Gerstlauer, Rainer Dömer
ASP-DAC3
2008 Introducing Preemptive Scheduling in Abstract RTOS Models using Result Oriented Modeling
abstract
With the increasing SW content of modern SoC designs, modeling and development of Hardware Dependent Software (HDS) become critical. Previous work addressed this by introducing abstract RTOS modeling, which exposes dynamic scheduling effects early in the system design flow. However, such models insufficiently capture preemption. In particular, the accuracy of preemption depends on the granularity of the timing annotation. For an accurately modeled interrupt response time, very fine-grained timing annotation is necessary, which contradicts the RTOS abstraction idea and is detrimental to simulation performance. In this paper, we eliminate the granularity dependency by applying the Result Oriented Modeling (ROM) technique previously used only for communication modeling. Our ROM approach allows precise preemptive scheduling, while retaining all the benefits of abstract RTOS modeling. Our experimental results demonstrate tremendous improvements. While the traditional model simulated an interrupt response time with a severe inaccuracy (12x longer in average and 40x longer for 96thpercentile), our ROM- based model was accurate within 8% (average and 50thpercentile) using identical timing annotations.
Gunar Schirner, Rainer Dömer
DATE2
2008 Code and Data Structure Partitioning for Parallel and Flexible MPSoC Specification Using Designer-Controlled Recoding
abstract
MultiProcessor systems-on-chip (MPSoCs) are increasingly being used to build efficient and cost-effective embedded systems that meet the necessary real-time requirements. However, programming heterogeneous MPSoCs is highly challenging. The existing automatic parallelizing techniques, although effective on homogeneous shared-memory architectures, are insufficient for MPSoCs, which are typically characterized by heterogeneous processing elements and memory architectures. The lack of effective automatic techniques for recoding and parallelization requires designers to manually partition the code and the data structures in the reference application to generate a parallel and flexible specification model. Such manual algorithm partitioning by the designer is time consuming and error prone. In this paper, we motivate the need for automation in system specification and present a novel designer-controlled approach to recode applications written in a C-based System-Level Description Language. We present six automated source code transformations that, under the control of the designer, automatically partition and reorganize code and data structures to create a parallel and flexible abstract specification model that can be mapped onto a heterogeneous MPSoC using a top-down system-level design flow. Our experimental results show significant productivity gains and quality improvements in the end design.
Pramod Chandraiah, Rainer Dömer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2008 Quantitative analysis of the speed/accuracy trade-off in transaction level modeling
abstract
The increasing complexity of embedded systems requires modeling at higher levels of abstraction. Transaction level modeling (TLM) has been proposed to abstract communication for high-speed system simulation and rapid design space exploration. Although being widely accepted for its high performance and efficiency, TLM often exhibits a significant loss in model accuracy. In this article, we systematically analyze and quantify the speed/accuracy trade-off in TLM. To this end, we provide a classification of TLM abstraction levels based on model granularity and define appropriate metrics and test setups to quantitatively measure and compare the performance and accuracy of such models. Addressing several classes of embedded communication protocols, we apply our analysis to three common bus architectures, the industry-standard AMBA advanced high-performance bus (AHB) as an on-chip parallel bus, the controller area network (CAN) as an off-chip serial bus, and the Motorola ColdFire Master Bus as an example for a custom embedded processor bus. Based on the analysis of these individual busses, we then generalize our results for a broader conclusion. The general TLM trade-off offers gains of up to four orders of magnitude in simulation speed, generally however, at the price of low accuracy. We conclude further that model granularity is the key to efficient TLM abstraction, and we identify conditions for accuracy of abstract models. As a result, this article provides general guidelines that allow the system designer to navigate the TLM trade-off effectively and choose the most suitable model for the given application with fast and accurate results.
Gunar Schirner, Rainer Dömer
ACM Trans. Embed. Comput. Syst.2
2008 An Interactive Design Environment for C-Based High-Level Synthesis of RTL Processors
abstract
Abstract—Much effort in register transfer level (RTL) design has been devoted to developing “push-button ” types of tools. However, given the highly complex nature, and lack of control on RTL design, push-button type synthesis is not accepted by many designers. Interactive design with assistance of algorithms and tools can be more effective if it provides control to the steps of synthesis. In this paper, we propose an interactive RTL design environment which enables designers to control the design steps and to integrate hardware components into a system. Our design environment is targeting a generic RTL processor architecture and supporting pipelining, multicycling, and chaining. Tasks in the RTL design process include clock definition, component allocation, scheduling, binding, and validation. In our interactive environment, the user can control the design process at every stage, observe the effects of design decisions, and manually override synthesis decisions at will. We present a set of experimental results that demonstrate the benefits of our approach. Our combination of automated tools and interactive control by the designer results in quickly generated RTL designs with better performance than fully-automatic results, comparable to fully manually optimized designs. Index Terms—Embedded systems, high level synthesis, interactive design environment, register transfer level (RTL) processor, system-on-chip (SoC). I.
Dongwan Shin, Andreas Gerstlauer, Rainer Dömer, Daniel Gajski
IEEE Trans. Very Large Scale Integr. Syst.3
2007 Creating Explicit Communication in SoC Models Using Interactive Re-Coding
abstract
Communication exploration has become a critical step during SoC design. Researchers in the CAD community have proposed fast and efficient techniques for comprehensive design space exploration to expedite this critical design step. Although these advances have been helpful in reducing the design time significantly, the overall design time of the system is still a bottleneck. All these techniques assume the availability of an initial SoC input model with explicit communication, whose quality significantly impacts the effectiveness of the communication exploration techniques. Today, these initial models need to be manually written by engineers, which is tedious, error-prone and time consuming. In fact, our studies on industrial-size examples have shown that about 50% of the communication exploration time is spent on coding and re-coding of the initial specification model. In this paper, we propose an efficient interactive approach to explicit communication creation by automating some of the common coding tasks in specification models for communication exploration. Our results show significant savings in designer time.
Pramod Chandraiah, Junyu Peng, Rainer Dömer
ASP-DAC3
2007 Abstract, Multifaceted Modeling of Embedded Processors for System Level Design
abstract
Embedded software is playing an increasing role in todays SoC designs. It allows a flexible adaptation to evolving standards and to customer specific demands. As software emerges more and more as a design bottleneck, early, fast, and accurate simulation of software becomes crucial. Therefore, an efficient modeling of programmable processors at high levels of abstraction is required. In this article, we focus on abstraction of computation and describe our abstract modeling of embedded processors. We combine the computation modeling with task scheduling support and accurate interrupt handling into a versatile, multi-faceted processor model with varying levels of features. Incorporating the abstract processor model into a communication model, we achieve fast co-simulation of a complete custom target architecture for a system level design exploration. We demonstrate the effectiveness of our approach using an industrial strength telecommunication example executing on a Motorola DSP architecture. Our results indicate the tremendous value of abstract processor modeling. Different feature levels achieve a simulation speedup of up to 6600 times with an error of less than 8% over a ISS based simulation. On the other hand, our full featured model exhibits a 3% error in simulated timing with a 1800 times speedup.
Gunar Schirner, Andreas Gerstlauer, Rainer Dömer
ASP-DAC3
2007 Designer-Controlled Generation of Parallel and Flexible Heterogeneous MPSoC Specification
abstract
Programming multi-processor systems-on-chip (MPSoC) involves partitioning and mapping of sequential reference code onto multiple parallel processing elements. The immense potential available through MPSoC architectures depends heavily on the effectiveness of this programming. Existing automatic parallelizing techniques, though effective on shared memory architectures, are insufficient for MPSoCs, which are typically characterized by heterogeneous processing elements and memory architectures. The lack of effective automatic techniques requires designers to manually partition the code and the data structures in the reference application to generate a parallel and flexible specification. Manual creation of this model is time consuming and error prone.
Pramod Chandraiah, Rainer Dömer
DAC2
2007 Automatic Layer-Based Generation of System-On-Chip Bus Communication Models
abstract
With growing market pressures and rising system complexities, automated system-level communication design with efficient design space exploration capabilities is becoming increasingly important. At the same time, customized network-oriented communication architectures become necessary in enabling a high-performance communication among the system components. To this end, corresponding communication design flows that are supported by efficient design automation techniques need to be developed. In this paper, we present a system-level design environment for the generation of bus-based system-on-chip architectures. Our approach supports a two-stage design flow using automated model refinement toward custom heterogeneous communication networks. Starting from an abstract specification of the desired communication channels, our environment automatically generates tailored network models at various levels of abstraction. At its core, an automatic layer-based refinement approach is utilized. We have applied our approach to a set of industrial-strength examples with a wide range of target architectures. Our experimental results show significant productivity gains over a traditional communication design, allowing early and rapid design space exploration.
Andreas Gerstlauer, Dongwan Shin, Junyu Peng, Rainer Dömer, Daniel Gajski
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2007 Result-Oriented Modeling - A Novel Technique for Fast and Accurate TLM
abstract
Efficient communication modeling is a critical task in system-on-chip design and exploration. In particular, fast and accurate communication is needed to predict the performance of a system. Recently, transaction level modeling is used to speed up communication simulation at the cost of accuracy. This paper proposes a novel modeling technique, called result-oriented modeling (ROM), which removes the inaccuracy drawback of transaction level models (TLMs) in many cases. Using ROM, simulation models yield nearly the same speed as their traditional TLM counterparts, yet are still 100% accurate in timing. ROM utilizes the fact that internal states in the communication channel are not observable by the caller. Hence, ROM omits the internal states entirely and optimistically predicts the end result. Retroactively, the outcome of the prediction is checked, and if necessary, corrective measures are taken to maintain the accuracy of the model. We have applied the ROM concept to two examples: the industry standard AMBA AHB and the controller area network. To validate the proposed ROM approach, we have analyzed the models in detail for performance and accuracy. Our experimental results show the clear advantages of the ROM concept. For both bus systems, ROM achieves 100% accuracy and highest speeds. In essence, ROM eliminates the TLM tradeoff for a wide range of platforms. It frees the system designer from having multiple models for different purposes and extends the TLM idea to applications that require timing accurate simulation, such as real-time communication.
Gunar Schirner, Rainer Dömer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2006 Quantitative analysis of transaction level models for the AMBA bus
abstract
The increasing complexity of embedded systems pushes system designers to higher levels of abstraction. Transaction level modeling (TLM) has been proposed to model communication in systems in an abstract manner Although being widely accepted, TLMs have not been analyzed for their loss in accuracy. This paper will analyze and quantify the speed-accuracy tradeoff of TLM using a case study on AMBA, an industry bus standard. It shows the results of modeling the advanced high-performance bus (AHB) of AMBA using a set of models at different abstraction levels. The analysis of the simulation speed shows improvements of two orders of magnitude for each TLM abstraction, while the timing in the model remains accurate for many applications. As a result, the paper will classify the different models towards their applicability in typical modeling situations, allowing the system designer to achieve fast and accurate simulation of communication
Gunar Schirner, Rainer Dömer
DATE2
2006 A Flexible, Syntax Independent Representation (SIR) for System Level Design Models
abstract
System level design (SLD) is widely seen as a solution for bridging the gap between chip complexity and design productivity of systems on chip (SoC). SLD relieves the designer from detailed manual implementation by raising the level of abstraction in design models. There are many different modeling approaches to SLD. With the abundance of design languages and supporting tools, a designer can create a multitude of models of the same system that are difficult to compare and evaluate against each other. Also, it is not unusual for the same system model to comply to specification guidelines in one simulation tool, and exceed them in another. This paper presents an approach to circumvent this problem by describing a design model using a syntax independent representation (SIR). Such representation of the model complies to modeling guidelines shared by various SLD languages but is not restricted by their syntactic compositions. Further, the same structure can represent models on different levels of abstraction. Our multidimensional and multipurpose SIR structure can be projected onto planes of modeling languages including SpecC and SystemC, as demonstrated by experimental results shown in this work
Ines Viskic, Rainer Dömer
DSD2
2006 Fast and accurate transaction level models using result oriented modeling
abstract
Efficient communication modeling is a critical task in SoC design and exploration. In particular, fast and accurate communication is needed to predict the performance of a system. Recently, Transaction Level Modeling (TLM) is used to speedup communication simulation at the cost of accuracy. This paper proposes a novel modeling technique called Result Oriented Modeling (ROM) which removes the accuracy drawback of TLM. Using ROM, models yield the same speed as their TLM counterparts, yet still are 100 % accurate in timing. ROM utilizes the fact that internal states in the communication channel are not observable by the caller. Hence, ROM omits the internal states entirely and optimistically predicts the end result. Retroactively, the outcome is checked and, if necessary, corrective measures are taken to maintain the accuracy of the model. In this paper, we apply ROM to the AMBA AHB bus architecture. Our experimental results show that ROM exhibits the same high simulation performance as traditional TLM, yet it retains the same accuracy as the bus functional model. Thus, the proposed ROM approach eliminates the speed/accuracy tradeoff exhibited by traditional TLM. 1.
Gunar Schirner, Rainer Dömer
ICCAD2
2005 System-level communication modeling for network-on-chip synthesis
abstract
As we are entering the network-on-chip era and system communication is becoming a dominating factor, communication abstraction and synthesis are becoming the integral part of system design flows. The key to the success of any design flow are well-defined abstraction levels and models, which enable automation of early validation, synthesis and verification. In this paper, we define system communication abstraction layers and corresponding design models that support successive, stepwise refinement from abstract message-passing down to a cycle-accurate, bus-functional implementation. Experimental results show the benefits of our definitions and design flow.
Andreas Gerstlauer, Dongwan Shin, Rainer Dömer, Daniel Gajski
ASP-DAC3
2004 Embedded software generation from system level design languages
Haobo Yu, Rainer Dömer, Daniel Gajski
ASP-DAC2
2000 Reuse and protection of intellectual property in the SpecC system
abstract
No abstract available.
Rainer Dömer, Daniel Gajski
ASP-DAC1
1997 Built-in chaining: introducing complex components into architectural synthesis
abstract
Extends the set of library components which are usually considered in architectural synthesis by components with built-in chaining (BIC). For such components, the result of some internally computed arithmetic function is made available as an argument to some other function through a local connection. These components can be used to implement chaining in a data-path in a single component. Components with BIC are combinatorial circuits. They correspond to "complex gates" in logic synthesis. If compared to implementations with several components, components with BIC usually provide a denser layout, reduced power consumption and a shorter delay time. Multiplier/accumulators are the most prominent example of such components. Such components require new approaches for library mapping in architectural synthesis. In this paper, we describe an integer programming (IP) based approach taken in our OSCAR (Optimum Simultaneous sCheduling, Allocation and Resource assignment) synthesis system.
Peter Marwedel, Birger Landwehr, Rainer Dömer
ASP-DAC3