VLDB 2026 Research / reviewers in the wild / expert
Daniel Mueller-Gritschneder
dblp:96/8188 · also Daniel Mueller 0001, Daniel Müller-Gritschneder
· DBLP profile ↗
61ranked-venue papers
9as first author
24since 2021 · last 2025
0000-0003-0903-631XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 50 · 8 first-author · 15 since 2021Software engineering, systems software and programming languages · 25 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SuperFast: Fast Supernet Training Using Initial KnowledgeabstractOnce-for-all based neural architecture search (NAS) proposes to train a supernet once and extract specialized subnets from it for efficient deployment. This decoupling between training and search enables easy multi-target deployment without retraining. Nevertheless, the initial training cost has remained extremely high, with SOTA approaches like ElasticViT and NASViT taking more than 72 and 83 GPU days respectively. While other approaches have tried to accelerate the training by warming up the largest model in the search space, we argue that this is suboptimal, and knowledge is easier scaled upward than downward. Hence, we propose SuperFast, a simple, plug and play workflow, that (I.) pretrains a subnet of the supernet search space, and (II.) distributes its knowledge within the supernet before the training. SuperFast offers a substantial acceleration in the supernet training, resulting in a significantly better accuracy vs. training-cost trade-off. Using SuperFast on both ElasticViT and NASViT supernets achieves the baseline’s accuracy $1.4 \times$ and $\mathbf{1. 8} \times$ faster on the ImageNet dataset. Moreover, for a given time budget, SuperFast improves accuracy vs. latency trade-offs for subnets, gaining 4.0 p.p. for the $20-50 \mathrm{~ms}$ range on Pixel 6. Code available in https://github.com/MoritzTho/SuperFast. Moritz Thoma, Emad Aghajanzadeh, Shambhavi Balamuthu Sampath, Pierpaolo Morì, Nael Fasfous, Alexander Frickenstein, Manoj Rohit Vemparala, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DAC | 8 |
| 2025 | Rapid Fault Injection Simulation by Hash-Based Differential Fault Effect Equivalence ChecksabstractAssessing a computational system's resilience to hardware faults is essential for safety and security-related systems. Fault Injection (FI) simulation is a valuable tool that can increase confidence in computational systems and guide hardware and software design decisions in the early stages of development. However, simulating hardware at low levels of abstraction, such as Register Transfer Level (RTL), is costly, and minimizing the effort required for large-scale FI campaigns is a significant objective. This work introduces Hash-based Differential Fault Effect Equivalence Checks to automatically terminate experiments early based on predicting their outcome. We achieve this by matching observed fault effects to ones already encountered in previous experiments. We generate these hashes from differentials computed by repurposing existing fast boot checkpoints from a state-of-the-art acceleration method. By integrating these approaches in an automated manner, we can accelerate a large-scale FI simulation of a CPU at RTL. We reduce the average simulation time by a factor of up to 25 compared to a factor of around 2 to 5 for state-of-the-art techniques. While maintaining 100 % accuracy, we can recover the faulty state through the stored differentials. Johannes Geier, Leonidas Kontopoulos, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DATE | 3 |
| 2025 | HiFi-SAGE: High Fidelity GraphSAGE-Based Latency Estimators for DNN OptimizationabstractAs deep neural networks (DNNs) are increasingly deployed on resource-constrained edge devices, optimizing and compressing them for real-time performance becomes crucial. Traditional hardware-aware DNN search methods often rely on inaccurate proxy metrics, expensive latency lookup tables, or slow hardware-in-the-Iloop (HIL) evaluations. To address this, quasi-generalized latency estimators, typically meta-learning-based, were proposed to replace HIL evaluations and accelerate the search. These come with a one-time data collection and training cost and can adapt to new hardware with few measurements. However, they still have some drawbacks: (1) They increase complexity by trying to generalize across a range of diverse hardware types; (2) They depend on handcrafted hardware descriptors, which may fail to capture hardware characteristics; (3) They often perform poorly on new, unseen hardware that significantly differs from their initial training set. To overcome these challenges, this paper turns to the more straightforward platform-specific estimators that do not require hardware descriptors and can be easily trained on any hardware. We introduce HiFi-SAGE, a high fidelity GraphSAGE-based platform-specific latency estimator. When trained from scratch on only 100 latency measurements, our novel dual-head estimator design surpasses the state-of-the-art (SoTA) on the 10% error bound metric by up to 17.4 p.p. while achieving an impressive fidelity score of 99% on the diverse LatBench dataset. We demonstrate that applying HiFi-SAGE to a genetic algorithm-based DNN compression search, achieved a Pareto front comparable to real HIL feedback with a mean absolute percentage error (MAPE) of 2.54%, 2.48%, and 4.16%, for InceptionV3, DenseNet169, and ResNet50 respectively. Compared to existing platform-specific works, the lower number of latency measurements and higher fidelity scores positions HiFi-SAGE as an attractive alternative to replace expensive HIL setups. Code is available at: https://github.com/shamvbs/HiFi-SAGE * Shambhavi Balamuthu Sampath, Leon Hecht, Moritz Thoma, Lukas Frickenstein, Pierpaolo Morì, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Walter Stechele, Daniel Mueller-Gritschneder, Claudio Passerone |
DATE | 10 |
| 2025 | Automated Graph-level Passes for TinyML Fault ToleranceabstractDeploying Machine Learning (ML) applications on Microcontroller Unit (MCU)-type devices, known as TinyML, poses significant challenges due to constrained resources. Consequently, this restricts the integration of fault tolerance mechanisms. Many redundancy techniques consume substantial amounts of already limited resources. To address this, Algorithm-Based Fault Tolerance (ABFT) methods for ML workloads focus on resource-intensive neural network operators at the kernel-level. However, this abstraction level introduces challenges, as these operators are often encapsulated within vendor-specific proprietary libraries. This work introduces automated and universal graph-level Data Flow Graph (DFG) passes supporting common kernel-level ABFT methods, along with a sophisticated redundancy mechanism called Dual Module Redundancy Island (DMRland). These DFG passes are implemented in the Tensor Virtual Machine (TVM) compiler framework and evaluated on a CPU-only execution platform with the MLPerf™Tiny benchmark. Our experimental results demonstrate that the proposed approach achieves competitive fault resilience in comparison to kernel-level ABFT methods, resulting in a reduction by two orders of magnitude for misclassification with combined ABFT and DMRland methods. Further, the run-time overhead (RTO) remains, on average, 19% lower than for comparable kernel-based implementations. Johannes Kappes, Johannes Geier, Phillipp van Kempen, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
IJCNN | 4 |
| 2025 | DRIP: DRop unImportant data Points - Enhancing Machine Learning Efficiency with Grad-CAM-Based Streaming Data Prioritization for On-Device TrainingabstractSelecting data points for model training is critical in machine learning. Effective selection methods can reduce the labeling effort, optimize on-device training for embedded systems with limited data storage, and enhance the model performance. This paper introduces a novel algorithm that uses Grad-CAM to make online decisions about retaining or discarding data points. Optimized for embedded devices, the algorithm computes a unique DRIP Score to quantify the importance of each data point. This enables dynamic decision-making on whether a data point should be stored for potential retraining or discarded without compromising model performance. Experimental evaluations on four benchmark datasets demonstrate that our approach can match or even surpass the accuracy of models trained on the entire dataset, while achieving storage savings of up to 39%. To our knowledge, this is the first algorithm to make online decisions about data point retention without requiring access to the entire dataset. Marcus Rüb, Daniel Konegen, Axel Sikora, Daniel Mueller-Gritschneder |
IJCNN | 4 |
| 2024 | MATAR: Multi-Quantization-Aware Training for Accurate and Fast Hardware RetargetingabstractQuantization of deep neural networks (DNNs) reduces their memory footprint and simplifies their hardware arithmetic logic, enabling efficient inference on edge devices. Different hardware targets can support different forms of quantization, e.g. full 8-bit, or 8/4/2-bit mixed-precision combinations, or fully-flexible bit-serial solutions. This makes standard quantization-aware training (QAT) of a DNN for different targets challenging, as there needs to be careful consideration of the supported quantization-levels of each target at training time. In this paper, we propose a generalized QAT solution that results in a DNN which can be retargeted to different hardware, without any retraining or prior knowledge of the hardware's supported quantization policy. First, we present the novel training scheme which makes the model aware of multiple quantization strategies. Then we demonstrate the retargeting capabilities of the resulting DNN by using a genetic algorithm to search for layer-wise, mixed-precision solutions that maximize performance and/or accuracy on the hardware target, without the need of fine-tuning. By making the DNN agnostic of the final hardware target, our method allows DNNs to be distributed to many users on different hardware platforms, without the need for sharing the training loop or dataset of the DNN developers, nor detailing the hardware capabilities ahead of time by the end-users of the efficient quantized solution. Models trained with our approach can generalize on multiple quantization policies with minimal accuracy degradation compared to target-specific quantization counterparts. Pierpaolo Morì, Moritz Thoma, Lukas Frickenstein, Shambhavi Balamuthu Sampath, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Walter Stechele, Daniel Mueller-Gritschneder, Claudio Passerone |
DATE | 9 |
| 2024 | Seal5: Semi-Automated LLVM Support for RISC-V ISA Extensions Including AutovectorizationabstractThe RISC-V instruction set architecture (ISA) is popular for its extensibility, allowing easy integration of custom vendor-defined instructions tailored to specific applications. However, a quick exploration of instruction candidates fails due to the lack of tools to auto-generate embedded software toolchain support. In particular, exploiting SIMD instructions to accelerate typical DSP and machine learning workloads needs specialized integration. This work establishes a semi-automated flow to generate LLVM compiler support for custom instructions based on a C-style ISA description language. The implemented Seal5 tool is capable of generating support for functionalities ranging from baseline assembler-level support, over builtin functions to compiler code generation patterns for scalar as well as vector instructions, while requiring no deeper compiler know-how. This paper focuses primarily on a novel pattern generator approach for the optimized code generation for SIMD instructions, including support for autovectorization. The auto generated LLVM toolchain reduces development times drastically while performing similarly or better compared to the existing, manually implemented Core-V reference LLVM toolchain on a wide variety of benchmarks. Seal5 further allows the addition of compiler code generation support for the Core-VSIMD instructions, which is not yet available in the reference toolchain. Additionally, Seal5 facilitates a quick exploration of custom instruction candidates as demonstrated for a cryptography extension. Philipp van Kempen, Mathis Salmen, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DSD | 3 |
| 2024 | EGIC: Enhanced Low-Bit-Rate Generative Image Compression Guided by Semantic Segmentation
Nikolai Körber, Eduard Kromer, Andreas Siebert, Sascha Hauke, Daniel Mueller-Gritschneder, Björn W. Schuller |
ECCV (35) | 5 |
| 2024 | Advancing On-Device Neural Network Training with TinyPropv2: Dynamic, Sparse, and Efficient BackpropagationabstractThis study introduces EmbeddedTrain, an innovative algorithm optimized for on-device learning in deep neural networks, specifically designed for low-power microcontroller units. EmbeddedTrain refines sparse backpropagation by dynamically adjusting the level of sparity, including the ability to selectively skip training steps. This feature significantly lowers computational effort without substantially compromising accuracy. Our comprehensive evaluation across diverse datasets—CIFAR 10, CIFAR100, Flower, Food, Speech Command, MNIST, HAR, and DCASE2020—reveals that EmbeddedTrain achieves near-parity with full training methods, with an average accuracy drop of only around 1% in most cases. For instance, against full training, EmbeddedTrain’s accuracy drop is minimal, for example, only 0.82% on CIFAR 10 and 1.07% on CIFAR100. In terms of computational effort, EmbeddedTrain shows a marked reduction, requiring as little as 10% of the computational effort needed for full training in some scenarios, and consistently outperforms other sparse training methodologies. These findings underscore EmbeddedTrain’s capacity to efficiently manage computational resources while maintaining high accuracy, positioning it as an advantageous solution for advanced embedded device applications in the IoT ecosystem. Marcus Rüb, Axel Sikora, Daniel Mueller-Gritschneder |
IJCNN | 3 |
| 2024 | Wino Vidi Vici: Conquering Numerical Instability of 8-bit Winograd Convolution for Accurate Inference Acceleration on EdgeabstractWinograd-based convolution can reduce the total number of operations needed for convolutional neural network (CNN) inference on edge devices. Most edge hardware accelerators use low-precision, 8-bit integer arithmetic units to improve energy efficiency and latency. This makes CNN quantization a critical step before deploying the model on such an edge device. To extract the benefits of fast Winograd-based convolution and efficient integer quantization, the two approaches must be combined. Research has shown that the transform required to execute convolutions in the Winograd domain results in numerical instability and severe accuracy degradation when combined with quantization, making the two techniques incompatible on edge hardware. This paper proposes a novel training scheme to achieve efficient Winograd-accelerated, quantized CNNs. 8-bit quantization is applied to all the intermediate results of the Winograd convolution without sacrificing task-related accuracy. This is achieved by introducing clipping factors in the intermediate quantization stages as well as using the complex numerical system to improve the transform. We achieve 2.8× and 2.1× reduction in MAC operations on ResNet-20-CIFAR-10 and ResNet-18-ImageNet, respectively, with no accuracy degradation. Pierpaolo Morì, Lukas Frickenstein, Shambhavi Balamuthu Sampath, Moritz Thoma, Nael Fasfous, Manoj Rohit Vemparala, Alexander Frickenstein, Christian Unger, Walter Stechele, Daniel Mueller-Gritschneder, Claudio Passerone |
WACV | 10 |
| 2024 | Flexible Generation of Fast and Accurate Software Performance Simulators From Compact Processor DescriptionsabstractTo find optimal solutions for modern embedded systems, designers frequently rely on the software performance simulators. These simulators combine an abstract functional description of a processor with a nonfunctional timing model to accurately estimate the processor’s timing while maintaining high simulation speeds. However, current performance simulators either inflexibly target specific processors or sacrifice accuracy or simulation speed. This article presents a new approach to the software performance simulation, combining flexibility with highly accurate estimates and high simulation speed. A code generator converts a compact structural description of the target processor’s pipeline into sets of timing constraints, describing the processor’s instruction execution. Based on these, it generates corresponding scheduling functions and timing variables, representing the availability of the modeled pipeline. The performance estimator uses these components to approximate the processor’s timing based on an instruction trace provided by an instruction set simulator. Results for the state-of-the-art CV32E40P and CVA6 RISC-V processors show an average relative error of 0.0015% and 3.88%, respectively, over a large set of benchmarks. Our approach reaches an average simulation speed of 24 and 15 million instructions per second (MIPS), respectively. Conrad Foik, Robert Kunzelmann, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | CompaSeC: A Compiler-Assisted Security Countermeasure to Address Instruction Skip Fault Attacks on RISC-VabstractFault-injection attacks are a risk for any computing system executing security-relevant tasks, such as a secure boot process. While hardware-based countermeasures to these invasive attacks have been found to be a suitable option, they have to be implemented via hardware extensions and are thus not available in most Commonly used Off-The-Shelf (COTS) components. Software Implemented Hardware Fault Tolerance (SIHFT) is therefore the only valid option to enhance a COTS system's resilience against fault attacks. Established SIHFT techniques usually target the detection of random hardware errors for functional safety and not targeted attacks. Using the example of a secure boot system running on a RISC-V processor, in this work we first show that when the software is hardened by these existing techniques from the safety domain, the number of vulnerabilities in the boot process to single, double, triple, and quadruple instruction skips cannot be fully closed. We extend these techniques to the security domain and propose Compiler-assisted Security Countermeasure (CompaSeC). We demonstrate that CompaSeC can close all vulnerabilities for the studied secure boot system. To further reduce performance and memory overheads we additionally propose a method for CompaSeC to selectively harden individual vulnerable functions without compromising the security against the considered instruction skip faults. Johannes Geier, Lukas Auer, Daniel Mueller-Gritschneder, Uzair Sharif, Ulf Schlichtmann |
ASP-DAC | 3 |
| 2023 | vRTLmod: An LLVM based Open-source Tool to Enable Fault Injection in Verilator RTL SimulationsabstractWith an ever increasing utilization of open-source hardware, the demand for improved productivity in its various development phases rises. This demand can only be met by open-source tools accompanying these open-hardware projects to cover their crucial tasks in design, test, and verification. One of these required tasks is fault injection analysis. At early stages of development, it can help to identify vulnerabilities of the hardware designs. In this work, we propose the open-source tool vRTLmod which enables low overhead and scaling fault injection simulations at Register Tranfer Level (RTL). Johannes Geier, Daniel Mueller-Gritschneder |
CF | 2 |
| 2023 | Extended Abstract: Monitoring-based Thermal Management for Mixed-Criticality SystemsabstractWith a rapidly growing number of functions in embedded real-time systems, it becomes inevitable to integrate tasks of different safety integrity levels (SILs) into one mixed-criticality system. Here, it is important to not only isolate shared architectural resources, as tasks executing on different cores may also interfere via the processor's thermal manager. In order to prevent a scenario where best-effort tasks cause deadline violations for critical tasks, we propose a thermal management strategy that guarantees a sufficient thermal isolation between tasks of different SILs, and simultaneously reduces the run-time of best-effort tasks by up to 45 % compared to the state of the art without incurring any real-time violations for critical tasks. Marcel Mettler, Martin Rapp, Heba Khdr, Daniel Mueller-Gritschneder, Jörg Henkel, Ulf Schlichtmann |
DATE | 4 |
| 2023 | Efficient Software-Implemented HW Fault Tolerance for TinyML Inference in Safety-critical ApplicationsabstractTinyML research has mainly focused on optimizing neural network inference in terms of latency, code-size and energy-use for efficient execution on low-power micro-controller units (MCUs). However, distinctive design challenges emerge in safety-critical applications, for example in small unmanned autonomous vehicles such as drones, due to the susceptibility of off-the-shelf MCU devices to soft-errors. We propose three new techniques to protect TinyML inference against random soft errors with the target to reduce run-time overhead: one for protecting fully-connected layers; one adaptation of existing algorithmic fault tolerance techniques to depth-wise convolutions; and an efficient technique to protect the so-called epilogues within TinyML layers. Integrating these layer-wise methods, we derive a full-inference hardening solution for TinyML that achieves run-time efficient soft-error resilience. We evaluate our proposed solution on MLPerf-Tiny benchmarks. Our experimental results show that competitive resilience can be achieved compared with currently available methods, while reducing run-time overheads by ~120% for one fully-connected neural network (NN); ~20% for the two CNNs with depth-wise convolutions; and ~2% for standard CNN. Additionally, we propose selective hardening which reduces the incurred run-time overhead further by ~2x for the studied CNNs by focusing exclusively on avoiding mispredictions. Uzair Sharif, Daniel Mueller-Gritschneder, Rafael Stahl, Ulf Schlichtmann |
DATE | 2 |
| 2023 | Memory Latency Distribution-Driven Regulation for Temporal Isolation in MPSoCs
Ahsan Saeed, Denis Hoornaert, Dakshina Dasari, Dirk Ziegenbein, Daniel Mueller-Gritschneder, Ulf Schlichtmann, Andreas Gerstlauer, Renato Mancuso 0001 |
ECRTS | 5 |
| 2022 | The Scale4Edge RISC-V EcosystemabstractThis paper introduces the project Scale4Edge. The project is focused on enabling an effective RISC-V ecosystem for optimization of edge applications. We describe the basic components of this ecosystem and introduce the envisioned demonstrators, which will be used in their evaluation. Wolfgang Ecker, Peer Adelt, Wolfgang Müller 0003, Reinhold Heckmann, Milos Krstic, Vladimir Herdt, Rolf Drechsler, Gerhard Angst, Ralf Wimmer 0001, Andreas Mauderer, Rafael Stahl, Karsten Emrich, Daniel Mueller-Gritschneder, Bernd Becker 0001, Philipp M. Scholl, Eyck Jentzsch, Jan Schlamelcher, Kim Grüttner, Paul Palomero Bernardo, Oliver Bringmann 0001, Brindusa Mihaela Damian-Kosterhon, Julian Oppermann, Andreas Koch 0001, Jörg Bormann, Johannes Partzsch, Christian Mayr 0001, Wolfgang Kunz |
DATE | 13 |
| 2022 | CorePerfDSL: A Flexible Processor Description Language for Software Performance SimulationabstractInstruction set simulators (ISSs) model the functional behavior of embedded processors for early software development. While they offer high simulation speeds, an ISS usually does not model the timing behavior of the processor accurately. Existing software performance simulators are typically either specific to a certain microarchitecture or their description language mixes functional and microarchitectural aspects.In this paper, we introduce CorePerfDSL, an architecture description language (ADL) specifically designed to model the timing behavior of processor microarchitectures for software performance estimation. CorePerfDSL is clearly separated from any functional description of the modelled processor by a generic trace definition. As such, it is well suited to generate performance simulators that can be paired with an existing ISS, which supplies an execution trace. In addition, CorePerfDSL provides a high degree of flexibility, supporting the fast generation of models for various microarchitecture variants, which can be used for rapid architectural exploration. We demonstrate the flexibility of CorePerfDSL by describing several variants of a single-issue five-stage RISC-V microarchitecture and estimate their performances for a software benchmark program. Conrad Foik, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
FDL | 2 |
| 2022 | Universal Safety Format: Automated Safety Software GenerationabstractThe development of safety-critical software requires a significant additional effort compared to standard software. Safety mechanisms, e.g., for mitigating hardware errors, have to be designed and integrated into the functional code. This results not only in substantial implementation overhead, but also reduces the overall maintainability of the software. In this paper, we present the Universal Safety Format (USF), which enables a model-driven approach that complies with the separation of concerns principle. Software safety mechanisms are specified as patterns via a domain-agnostic transformation language, separated from the functional software. Various domain-specific tools apply these safety patterns to domain-specific artifacts, such as code or software architecture models. This enables the reuse of safety patterns in multiple designs as well as in a single design to artifacts from different domains. Frederik Haxel, Alexander Viehl, Michael Benkel, Bjoern Beyreuther, Klaus Birken, Rolf Schmedes, Kim Grüttner, Daniel Mueller-Gritschneder |
MODELSWARD | 8 |
| 2022 | Memory Utilization-Based Dynamic Bandwidth Regulation for Temporal Isolation in Multi-CoresabstractTemporal isolation is one of the key challenges for co-running mixed-criticality applications on Commercial Off-The-Shelf (COTS) multi-core platforms. In particular, the main memory subsystem is one of the most prominent causes of interference and loss of isolation. Existing mechanisms for memory bandwidth regulation are limited to conservative bandwidth reservation, use pessimistic worst-case execution time (WCET) estimations or require dedicated hardware that is not feasible in COTS multi-core platforms.In this paper, we propose a novel mechanism for memory interference control that uses feedback-based control to dynamically regulate memory accesses of individual cores in a multicore platform. Our mechanism directly regulates the source of interference by leveraging information about memory utilization, acquired from existing hardware performance counters provided by modern COTS-based memory controllers. The proposed solution is implemented on Linux as a loadable kernel module. The results of evaluating our approach with real and synthetic benchmarks on a COTS multi-core (NXP S32V234) platform demonstrate that it is able to provide temporal isolation with up to 4x and 2x more overall throughput for non-real-time applications compared to static and dynamic memory bandwidth-based regulation approaches, respectively, while maintaining guarantees for applications running on the real-time core. Ahsan Saeed, Dakshina Dasari, Dirk Ziegenbein, Varun Rajasekaran, Falk Rehm, Michael Pressler, Arne Hamann 0001, Daniel Mueller-Gritschneder, Andreas Gerstlauer, Ulf Schlichtmann |
RTAS | 8 |
| 2022 | An FPGA-based Approach to Evaluate Thermal and Resource Management Strategies of Many-core ProcessorsabstractThe continuous technology scaling of integrated circuits results in increasingly higher power densities and operating temperatures. Hence, modern many-core processors require sophisticated thermal and resource management strategies to mitigate these undesirable side effects. A simulation-based evaluation of these strategies is limited by the accuracy of the underlying processor model and the simulation speed. Therefore, we present, for the first time, an field-programmable gate array (FPGA)-based evaluation approach to test and compare thermal and resource management strategies using the combination of benchmark generation, FPGA-based application-specific integrated circuit (ASIC) emulation, and run-time monitoring. The proposed benchmark generation method enables an evaluation of run-time management strategies for applications with various run-time characteristics. Furthermore, the ASIC emulation platform features a novel distributed temperature emulator design, whose overhead scales linearly with the number of integrated cores, and a novel dynamic voltage frequency scaling emulator design, which precisely models the timing and energy overhead of voltage and frequency transitions. In our evaluations, we demonstrate the proposed approach for a tiled many-core processor with 80 cores on four Virtex-7 FPGAs. Additionally, we present the suitability of the platform to evaluate state-of-the-art run-time management techniques with a case study. Marcel Mettler, Martin Rapp, Heba Khdr, Daniel Mueller-Gritschneder, Jörg Henkel, Ulf Schlichtmann |
ACM Trans. Archit. Code Optim. | 4 |
| 2021 | Tiny Generative Image Compression for Bandwidth-Constrained Sensor ApplicationsabstractDeep image compression algorithms based on Generative Adversarial Networks (GANs) are a promising direction to address the strict communication bandwidth limitations commonly encountered in IoT sensor networks (e.g. Low Power Wide Area Networks). However, current methods do not consider that the sensor nodes, which perform the image encoding, usually only offer very limited computation and memory capabilities, e.g. a resource-constrained tiny device such as a micro-controller. In this paper, we propose the first tiny generative image compression method specifically designed for image compression on micro-controllers. We base our encoder on the well-known MobileNetV2 network architecture, while keeping the decoder side fixed. To cope with the resulting asymmetric design of the compression pipeline, we investigate the impact of different training strategies (end-to-end, knowledge distillation) and integer quantization techniques (post-training, quantization-aware training) on the GAN-training stability. On the Cityscapes dataset, we achieve a compression performance that is very close to the state-of-the-art, while requiring 99% less SRAM size, 97% smaller flash storage and 87% less multiply-add operations. Our findings suggest that tiny generative image compression is particularly well suited for application-specific domains. Nikolai Körber, Andreas Siebert, Sascha Hauke, Daniel Mueller-Gritschneder |
ICMLA | 4 |
| 2021 | A Distributed Hardware Monitoring System for Runtime Verification on Multi-Tile MPSoCsabstractExhaustive verification techniques do not scale with the complexity of today’s multi-tile Multi-processor Systems-on-chip (MPSoCs). Hence, runtime verification (RV) has emerged as a complementary method, which verifies the correct behavior of applications executed on the MPSoC during runtime. In this article, we propose a decentralized monitoring architecture for large-scale multi-tile MPSoCs. In order to minimize performance and power overhead for RV, we propose a lightweight and non-intrusive hardware solution. It features a new specialized tracing interconnect that distributes and sorts detected events according to their timestamps. Each tile monitor has a consistent view on a globally sorted trace of events on which the behavior of the target application can be verified using logical and timing requirements. Furthermore, we propose an integer linear programming-based algorithm for the assignment of requirements to monitors to exploit the local resources best. The monitoring architecture is demonstrated for a four-tiled MPSoC with 20 cores implemented on a Virtex-7 field-programmable gate array (FPGA). Marcel Mettler, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ACM Trans. Archit. Code Optim. | 2 |
| 2021 | REPAIR: Control Flow Protection based on Register Pairing Updates for SW-Implemented HW Fault ToleranceabstractSafety-critical embedded systems may either use specialized hardware or rely on Software-Implemented Hardware Fault Tolerance (SIHFT) to meet soft error resilience requirements. SIHFT has the advantage that it can be used with low-cost, off-the-shelf components such as standard Micro-Controller Units. For this, SIHFT methods apply redundancy in software computation and special checker codes to detect transient errors, so called soft errors, that either corrupt the data flow or the control flow of the software and may lead to Silent Data Corruption (SDC). So far, this is done by applying separate SIHFT methods for the data and control flow protection, which leads to large overheads in computation time. This work in contrast presents REPAIR, a method that exploits the checks of the SIHFT data flow protection to also detect control flow errors as well, thereby, yielding higher SDC resilience with less computational overhead. For this, the data flow protection methods entail duplicating the computation with subsequent checks placed strategically throughout the program. These checks assure that the two redundant computation paths, which work on two different parts of the register file, yield the same result. By updating the pairing between the registers used in the primary computation path and the registers in the duplicated computation path using the REPAIR method, these checks also fail with high coverage when a control flow error, which leads to an illegal jumps, occurs. Extensive RTL fault injection simulations are carried out to accurately quantify soft error resilience while evaluating Mibench programs along with an embedded case-study running on an OpenRISC processor. Our method performs slightly better on average in terms of soft error resilience compared to the best state-of-the-art method but requiring significantly lower overheads. These results show that REPAIR is a valuable addition to the set of known SIHFT methods. Uzair Sharif, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2020 | Investigating the Inherent Soft Error Resilience of Embedded Applications by Full-System SimulationabstractIt has long been acknowledged that some applications feature inherent resilience against soft errors, e.g., the impact of soft errors on multimedia applications is often non-visible to humans. In this paper we investigate the inherent resilience of two typical embedded applications using a case study of a control system and a robot arm. Both studies were enabled by our mixed-mode fault injection simulator ETISS-ML, which allows RTL-accurate fault injection while being able to simulate very long scenarios, e.g. robot movements of several seconds. Our results indicate that full simulation of the embedded system and its environment are required to classify whether the system can tolerate the impact of a soft error. This is due to the fact that it is hard to predict the impact of a certain output deviation without investigating the change in the system behavior taking into account the control loop. Based on this classification method we hope to be able to exploit this resilience for lowering the cost of error detection mechanisms in future research. Uzair Sharif, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ASP-DAC | 2 |
| 2020 | Machine Learning Approaches for Efficient Design Space Exploration of Application-Specific NoCsabstractIn many Multi-Processor Systems-on-Chip (MPSoCs), traffic between cores is unbalanced. This motivates the use of an application-specific Network-on-Chip (NoC) that is customized and can provide a high performance at low cost in terms of power and area. However, finding an optimized application-specific NoC architecture is a challenging task due to the huge design space. This article proposes to apply machine learning approaches for this task. Using graph rewriting, the NoC Design Space Exploration (DSE) is modelled as a Markov Decision Process (MDP). Monte Carlo Tree Search (MCTS), a technique from reinforcement learning, is used as search heuristic. Our experimental results show that—with the same cost function and exploration budget—MCTS finds superior NoC architectures compared to Simulated Annealing (SA) and a Genetic Algorithm (GA). However, the NoC DSE process suffers from the high computation time due to expensive cycle-accurate SystemC simulations for latency estimation. This article therefore additionally proposes to replace latency simulation by fast latency estimation using a Recurrent Neural Network (RNN). The designed RNN is sufficiently general for latency estimation on arbitrary NoC architectures. Our experiments show that compared to SystemC simulation, the RNN-based latency estimation offers a similar speed-up as the widely used Queuing Theory (QT). Yet, in terms of estimation accuracy and fidelity, the RNN is superior to QT, especially for high-traffic scenarios. When replacing SystemC simulations with the RNN estimation, the obtained solution quality decreases only slightly, whereas it suffers significantly when QT is used. Marcel Mettler, Daniel Mueller-Gritschneder, Thomas Wild, Andreas Herkersdorf, Ulf Schlichtmann |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2019 | SeRoHAL: generation of selectively robust hardware abstraction layers for efficient protection of mixed-criticality systemsabstractA major challenge in mixed-criticality system design is to ensure safe behavior under the influence of hardware errors while complying with cost and performance constraints. SeRoHAL generates hardware abstraction layers with software-based safety mechanisms to handle errors in peripheral interfaces. To reduce performance and memory overheads, SeRoHAL can select protection mechanisms, depending on the criticality of the hardware accesses. Petra R. Kleeberger, Juana Rivera, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ASP-DAC | 3 |
| 2019 | Cross-Layer Resilience: Challenges, Insights, and the Road AheadabstractResilience to errors in the underlying hardware is a key design objective for a large class of computing systems, from embedded systems all the way to the cloud. Sources of hardware errors include radiation, circuit aging, variability induced by manufacturing and operating conditions, manufacturing test escapes, and early-life failures. Many publications have suggested that cross-layer resilience, where multiple error resilience techniques from different layers of the system stack cooperate to achieve cost-effective resilience, is essential for designing cost-effective resilient digital systems. This paper presents a comprehensive overview of cross-layer resilience by addressing fundamental cross-layer resilience questions, by summarizing insights derived from recent advances in cross-layer resilience research, and by discussing future cross-layer resilience challenges. Eric Cheng, Daniel Mueller-Gritschneder, Jacob A. Abraham, Pradip Bose, Alper Buyuktosunoglu, Deming Chen, Hyungmin Cho, Yanjing Li, Uzair Sharif, Kevin Skadron, Mircea R. Stan, Ulf Schlichtmann, Subhasish Mitra |
DAC | 2 |
| 2019 | ACCESS: HW/SW Co-Equivalence Checking for Firmware OptimizationabstractCustomizing embedded computing platforms to specific application domains often necessitates optimizing the firmware and/or the HW/SW interface under tight resource constraints. Such optimizations frequently alter the communication between the firmware and the peripheral devices, possibly compromising functional correctness of the input/output behavior of the embedded system. This paper proposes a formal HW/SW co-equivalence checking technique for verifying correct I/O behavior of peripherals under a modified firmware. We demonstrate the great promise of our approach on RTL implementations of several open-source peripherals. In our experiments we successfully prove or disprove correctness of firmware optimizations for an industrial driver software. In addition, we also found a subtle bug in one of the peripherals and several undocumented preconditions for correct device behavior. Michael Schwarz 0010, Raphael Stahl, Daniel Mueller-Gritschneder, Ulf Schlichtmann, Dominik Stoffel, Wolfgang Kunz |
DAC | 3 |
| 2019 | Towards Reliable and Secure Post-Quantum Co-Processors based on RISC-VabstractIncreasingly complex and powerful Systems-on-Chips (SoCs), connected through a 5G network, form the basis of the Internet-of-Things (IoT). These technologies will drive the digitalization in all domains, e.g. industry automation, automotive, avionics, and healthcare. A major requirement for all above domains is the long-term (10 to 30 years) secure communication between the SoCs and the cloud over public 5G networks. The foreseeable breakthrough of quantum computers represents a risk for all communication. In order to prepare for such an event, SoCs must integrate secure quantum-computer-resistant cryptography which is reliable and protected against SW and HW attacks. Empowering SoCs with such strong security poses a challenging problem due to limited resources, tight performance requirements and long-term life-cycles. While current works are focused on efficient implementations of post-quantum cryptography, implementation-security and reliability aspects for SoCs are still largely unexplored. To this end, we present three contributions. First, we present a RISC-V co-processor for post-quantum security, able to support lattice-based cryptography. Second, we use HW/SW co-design techniques to accelerate the NTT transformation and hash generation. Third, we perform the fault analysis of the implementation. We show that our coprocessor achieves high reliability and security capabilities while preserving good performance. Tim Fritzmann, Uzair Sharif, Daniel Mueller-Gritschneder, Cezar Reinbrecht, Ulf Schlichtmann, Martha Johanna Sepúlveda |
DATE | 3 |
| 2019 | SRAM Design Exploration with Integrated Application-Aware Aging AnalysisabstractOn-Chip SRAMs are an integral part of safety-critical System-on-Chips. At the same time however, they are also most susceptible to reliability threats such as Bias Temperature Instability (BTI), originating from the continuous trend of technology shrinking. BTI leads to a significant performance degradation, especially in the Sense Amplifiers (SAs) of SRAMs, where failures are fatal, since the data of a whole column is destroyed. As BTI strongly depends on the workload of an application, the aging rates of SAs in a memory array differ significantly and the incorporation of workload information into aging simulations is vital. Especially in safety-critical systems precise estimation of application specific reliability requirements to predict the memory lifetime is a key concern. In this paper we present a workload-aware aging analysis for On-Chip SRAMs that incorporates the workload of real applications executed on a processor. According to this workload, we predict the performance degradation of the SAs in the memory. We integrate this aging analysis into an aging-aware SRAM design exploration framework that generates and characterizes memories of different array granularity to select the most reliable memory architecture for the intended application. We show that this technique can mitigate SA degradation significantly depending on the environmental conditions and the application workload. Alexandra Listl, Daniel Mueller-Gritschneder, Ulf Schlichtmann, Sani R. Nassif |
DATE | 2 |
| 2019 | Analysis of Dissipative Losses in Modular Reconfigurable Energy Storage Systems Using SystemC TLM and SystemC-AMSabstractBattery storage systems are becoming more popular in the automotive industry as well as in stationary applications. To fulfill the requirements in terms of power and energy, the literature is increasingly discussing electrically reconfigurable interconnection topologies. However, these topologies use switching elements on the cell and module level that exhibit an electric resistance due to their design and hence generate undesirable dissipative losses. In this article, we propose a new analysis and optimization framework to examine and minimize the losses in such topologies. For this purpose, we develop a SystemC model to investigate static and dynamic load scenarios, e.g., from the automotive domain. The model uses SystemC TLM for the digital subsystem, SystemC-AMS for the mixed-signal subsystem, and host-compiled simulation for the microcontroller executing the embedded software. Here, we analyze the impact of the dissipative losses on the system efficiency that depend on the modularization level, implying the number of serial and parallel switching elements. Our analysis clearly shows that in reconfigurable topologies, the modularization level has a significant influence on the losses, which in our automotive example covers several orders of magnitude. For the topologies we have investigated, the highest efficiency can be reached when a parallel-only modularization is aspired and the number of serial switching elements is minimized. It is also shown that the losses of the state-of-the-art topology with one battery pack protection switch are almost as high as in a smart cell approach in which each energy storage cell has its own switching element. However, due to the high number of switching elements, this results in a reduction of energy density and increases the system costs, showing that this is a multi-criteria optimization problem. Thomas Zimmermann 0006, Mathias Mora, Sebastian Steinhorst, Daniel Mueller-Gritschneder, Andreas Jossen |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2018 | Design and validation of fault-tolerant embedded controllersabstractEmbedded control systems are an important and often safety-critical class of applications that need to operate reliably even in the presence of faults. We show that intermittent fault scenarios caused by wear-out effects due to a higher density and a smaller geometry of the embedded electronic components may become a reliability concern for real-time embedded control applications. To mitigate the effects of such intermittent faults, we propose a novel fault-tolerant controller design method such that the resulting controllers ensure closed loop stability (i.e., guarantee safety) with only possibly degraded performance under such fault scenarios. In order to measure the amortized performance offered by the software implementations of such fault-tolerant controllers, we provide a program analysis methodology that statically estimates the quality of control guaranteed by the C code implementation of the fault-tolerant control law. This combination of fault-tolerant controller design followed by performance feedback computed using a formal analysis is illustrated with a case study from the automotive domain. Saurav Kumar Ghosh, Soumyajit Dey, Dip Goswami, Daniel Mueller-Gritschneder, Samarjit Chakraborty |
DATE | 4 |
| 2018 | ETISS-ML: A multi-level instruction set simulator with RTL-level fault injection support for the evaluation of cross-layer resiliency techniquesabstractETISS is an instruction set simulator (ISS) for Virtual Prototypes (VPs) modeled with SystemC/TLM. In this paper, we propose the extension ETISS-ML, which enables a multi-level simulation that switches between ISS-level and register transfer level (RTL) to accurately evaluate the impact of soft errors in the pipeline of a RISC processor. ETISS-ML achieves close-to-RTL-accurate fault injection simulation results with close-to-ISS simulation performance with a speed up gain up to 100x compared to RTL. For this, we propose an approach to dynamically determine the length of the RTL simulation period. The high simulation performance of ETISS-ML enables an ultra-efficient and accurate evaluation of cross-layer resiliency techniques for embedded applications, which requires running a large number of fault injections for long simulation scenarios. This is demonstrated on a case study of a Microcontroller Unit (MCU) executing a control algorithm for adaptive cruise control. Daniel Mueller-Gritschneder, Martin Dittrich, Josef Weinzierl, Eric Cheng, Subhasish Mitra, Ulf Schlichtmann |
DATE | 1 |
| 2018 | Automated Redirection of Hardware Accesses for Host-Compiled Software SimulationabstractFor host-compiled software simulation it is required that accesses from the target software to memory-mapped hardware are identified, so that they can be redirected to a virtual prototype. This is straight-forward if the software uses a hardware abstraction layer as interface. If such an interface is not used by existing or third-party source code, the rewriting of the code for host-compiled simulation involves a considerable manual effort. In this paper, we present a method to automate this process with the help of a symbolic execution engine. With our approach the time to adjust software for host-compilation is significantly reduced. We show that the most memory-mapped hardware accesses are correctly rewritten in a real-world application by comparing recorded access traces on a virtual prototype. Additionally, a test suite has been developed to cover edge-cases. Rafael Stahl, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
FDL | 2 |
| 2018 | Wavefront-MCTS: multi-objective design space exploration of NoC architectures based on Monte Carlo tree searchabstractApplication-specific MPSoCs profit immensely from a custom-fit Network-on-Chip (NoC) architecture in terms of network performance and power consumption. In this paper we suggest a new approach to explore application-specific NoC architectures. In contrast to other heuristics, our approach uses a set of network modifications defined with graph rewriting rules to model the design space exploration as a Markov Decision Process (MDP). The MDP can be efficiently explored using the Monte Carlo Tree Search (MCTS) heuristics. We formulate a weighted sum reward function to compute a single solution with a good trade-off between power and latency or a set of max reward functions to compute the complete Pareto front between the two objectives. The Wavefront feature adds additional efficiency when computing the Pareto front by exchanging solutions between parallel MCTS optimization processes. Comparison with other popular search heuristics demonstrates a higher efficiency of MCTS-based heuristics for several test cases. Additionally, the Wavefront-MCTS heuristics allows complete tracability and control by the designer to enable an interactive design space exploration process. Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ICCAD | 2 |
| 2018 | Performance and accuracy in soft-error resilience evaluation using the multi-level processor simulator ETISS-MLabstractSoft errors are a major safety concern in many devices, e.g., in automotive, industrial, control or medical applications. Ideally, safety-critical systems should be resilient against the impact of soft errors, but at a low cost. This requires to evaluate the soft error resilience, which is typically done by extensive fault injection. In this paper, we present ETISS-ML, a multi-level processor simulator, which manages to achieve both accuracy and performance for fault simulation by intelligently switching the level of abstraction between an Instruction Set Simulator (ISS) and an RTL simulator. For a given software testcase and fault scenario, the software is first executed in ISS-mode until shortly before the fault injection. Then ETISS-ML switches to RTL-mode for accurate fault simulation. Whenever the impact of the fault is propagated completely out of the processor's micro-architecture, the simulation can switch back to ISS-mode. This paper describes the methods needed to preserve accuracy during both of these switches. Experimental results show that ETISS-ML obtains near to ISS performance with RTL accuracy. It is also shown that ETISS-ML can be used as the processor model in SystemC / TLM virtual prototypes (VPs) and, hence, allows to investigate the impact of soft errors at system level. Daniel Mueller-Gritschneder, Uzair Sharif, Ulf Schlichtmann |
ICCAD | 1 |
| 2018 | Emulation of an ASIC Power, Temperature and Aging Monitor System for FPGA PrototypingabstractTechnology scaling has enabled the fabrication of Multi-Processor Systems-on-Chips (MPSoCs), which satisfy the ever growing demand for performance, while continuously reducing the chip size. Thus, scaling has also led to new challenges such as increasing power densities, which critically influence the chip temperatures and accelerate device degradation due to aging. Runtime power management can be utilized to counter these reliability threats to increase the lifetime of a system. For the development of runtime power management strategies monitoring data for power, temperature and aging is required. In this paper we propose a real-time power, temperature and aging monitor system (eTAPMon) for FPGA prototypes of MPSoCs. The monitor system emulates data characterized from the target ASIC design. The emulation approach models the behavior of ASIC power monitors based on an instruction-level energy model, temperature monitors based on a linear regression model obtained from thermal offline simulations and aging monitors based on a critical path model to compute the decreasing timing margin due to aging. An accelerated aging emulation is possible to predict aged ASIC behavior. Hence, this FPGA emulation enables the early evaluation of runtime power management strategies. Alexandra Listl, Daniel Mueller-Gritschneder, Fabian Kluge, Ulf Schlichtmann |
IOLTS | 2 |
| 2018 | Efficient Fault Injection for Embedded Systems: As Fast as Possible but as Accurate as NecessaryabstractWhen used for safety-critical applications, embedded systems must behave safely at all times - even in the presence of random hardware faults. To ensure this, fault effect simulation by simulation-based fault injection is an integral part of embedded system development. The high complexity of embedded systems results in low simulation performance if all details of the system are simulated. Not simulating all details, i.e. increasing the simulation abstraction level, speeds up fault injection but can result in less accuracy in predicting the fault impacts on the system behavior. To achieve high accuracy and high simulation performance at the same time, we avoid simulation of details unrelated to the injected fault. For this, we divide the set of faults that can occur in an embedded system into three subsets. For each subset, we select the fault injection abstraction level of the embedded processor model that is as accurate as necessary but as fast as possible. The considered levels are host-compiled simulation, instruction set simulation and register transfer level simulation. For additional speed-up, the abstraction level can be switched during the fault injection simulation between register transfer and instruction set level. The fault set for host-compiled simulation can be reduced by static program analysis. Our results show that adapting the abstraction level to the fault set achieves high performance of the fault injection simulation. Petra R. Maier, Uzair Sharif, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
IOLTS | 3 |
| 2018 | Fault Injection for Test-Driven Development of Robust SoC FirmwareabstractRobustness against errors in hardware must be considered from the very beginning of safety-critical system-on-chip firmware design. Therefore, we present fault injection for test-driven development (TDD) of robust firmware. As TDD is based on instant feedback to the designer, fault injection must execute within few minutes. In contrast to state-of-the-art approaches, we avoid long simulation scenarios and runtimes by injecting faults at the unit level and utilizing host-compiled simulation. Further, three static bit-level analyses of firmware source code and hardware specification reduce the fault set significantly. This accelerates fault injection by several orders of magnitude and enables robustness-aware TDD. Petra R. Maier, Veit Kleeberger, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2018 | Graph-Grammar-Based IP-Integration (GRIP) - An EDA Tool for Software-Defined SoCsabstractIn modern system-on-chip (SoC) designs, IP-reuse is considered a driving force to increase productivity. To support various designs, a huge amount of Intellectual Property (IP) hardware blocks have been developed. The integration of those IPs into an SoC may require significant effort—up to days or weeks depending on experience and complexity. This article presents a novel approach to significantly reduce the design effort to bring-up a working SoC design by automatic IP integration as part of a library-based Software-defined SoC flow. In detail, the IP-supplier prepares a HW-accelerated software library (HASL) for the SoC architect, who wants to use the IP in an SoC design. As a key point of our approach, integration knowledge is encoded in the library as a set of integration rules. These rules are defined in the machine-readable standardized IP-XACT format by the IP supplier, who has a good knowledge of the IP’s hardware details. The library preparation step on the IP supplier’s side is also partly automated in the proposed flow, including a partial generation of configurable HW drivers, schedulers, and the software library functions. For the SoC architect, we have developed the graph-grammar-based IP-integration (GRIP) tool. The software application is developed using the functions supplied in the HASL. According to the calls to the HASL functions, the GRIP tool automatically integrates IP-blocks using the rule information supplied with the library and runs a full Design Space Exploration. For this, the SoC architecture and rules are transformed into the graph domain to apply graph rewriting methods. The GRIP tool is model-driven and based on the Eclipse Modeling Framework. With code generation techniques, SoC candidate architectures can be transformed to hardware descriptions for the target platform. The HW/SW interfaces between SW library functions and IP blocks can be automatically generated for bare-metal or Linux-based applications. The approach is demonstrated with two case-studies on the Xilinx Zynq-based ZedBoard evaluation board using a HASL for computer vision. It can yield 10×-150× performance improvement for the bare-metal application versions and 4×--7× performance improvement for the Linux-based application versions, when executed on an optimized HW-accelerated SoC architecture compared to a non HW-accelerated SoC. The effort for IP integration is comparable to using a software library, hence, providing a significant advantage over a manual IP integration. Munish Jassi, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2017 | The extendable translating instruction set simulator (ETISS) interlinked with an MDA framework for fast RISC prototypingabstractThis paper describes the Extendable Translating Instruction Set Simulator (ETISS). In addition to binary translation, ETISS features a plugin mechanism that allows to quickly include new functionality into the translation stage, the simulation loop, during accesses to the memory or whenever an interrupt is received. ETISS targets to become an advanced industrial-strength ISS with special focus on virtual prototypes (VPs) written in SystemC/TLM. In this paper, we will show examples of ETISS Plugins which include tracing tools, SystemC interfaces, closey-coupled peripherals or triggers for fault injection. A major drawback of developing a new binary translator such as ETISS is its lack of support for a variety of instruction set architectures (ISAs). At the moment ETISS supports the open-source OpenRISC orlk and partly RISC-V ISAs. Yet, in order to overcome this problem, we developed a toolchain to generate the binary translation stage for different ISAs following the MDA concept based on meta-modeling and code generation. It is planned to make ETISS available as an open-source tool to the research community. Daniel Mueller-Gritschneder, Keerthikumara Devarajegowda, Martin Dittrich, Wolfgang Ecker, Marc Greim, Ulf Schlichtmann |
RSP | 1 |
| 2016 | Hardware-Accelerated Software Library Drivers Generation for IP-Centric SoC DesignsabstractIn recent years, the semiconductor industry has been witnessing an increasing reuse of hardware IPs for System-on-Chip (SoC) designs and embedded computing systems on FPGA platforms with hard-core processors. The IP-reuse comes with an increasing complexity at the hardware-software (HW-SW) interface. The efforts required to access the HW through the increasingly complex HW-SW interface diminishes the potential IP-reuse productivity gain. In our work, we are proposing hierarchical drivers for accessing IP-subsystems and its generation for enabling easier SW application adaptation to HW-changes and faster design space exploration (DSE) on a targeted HW-accelerated SW libraries. At the lowest level, closest to the HW, is the hardware abstraction layer (HAL), these are the platform-specific register-access drivers. At the next layer are the drivers to access the registers and bit-fields of each IP component of the IP-library. Next are the IP-subsystems drivers. At the top-layer, closest to the SW, is the simple scheduler with SW interface library that provides access functions to the SW application. The drivers generator uses the HW knowledge of IPs and IP-subsystems encoded in IP-XACT for generating the drivers for both operating system (OS) and non-OS based applications. For the OS-based applications, user-space drivers are generated, as well as device tree source (DTS) and drivers mapping in the kernel-space. In a case study, we have validated our methodology while performing DSE for a video processing application targeted to an IP-library, both as non-OS and with OS on Xilinx Zynq-based FPGA. Munish Jassi, Uzair Sharif, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ACM Great Lakes Symposium on VLSI | 3 |
| 2015 | GRIP: grammar-based IP integration and packaging for acceleration-rich SoC designsabstractIncreased hardware IP reuse is required to meet the productivity demands for the future complex Systems-on-Chip (SoC). Nowadays, IP integration is enabled using standardized meta-data formats such as IP-XACT. We present a new concept called grammar-based IP integration and packaging (GRIP), which additionally encodes design integration knowledge into a set of graph re-writing rules using standard IP-XACT. These GRIP rules are packaged into a domain-specific library of IP blocks. The library can be supplied by an IP provider along to an SoC architect. An integration tool can automatically use the GRIP rules to search the design space using the integration knowledge of the IP provider. The tool generates all design alternatives with different trade-offs for the SoC architect. We demonstrate the GRIP approach on a computer vision IP library for FPGA-based SoCs. Eighteen functional design alternatives are automatically generated within a few hours using IP integration knowledge encoded by the GRIP rules. Munish Jassi, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DAC | 2 |
| 2015 | The next generation of virtual prototyping: ultra-fast yet accurate simulation of HW/SW systems
Oliver Bringmann 0001, Wolfgang Ecker, Andreas Gerstlauer, Ajay Goyal, Daniel Mueller-Gritschneder, Prasanth Sasidharan |
DATE | 5 |
| 2014 | Safety Evaluation of Automotive Electronics Using Virtual Prototypes: State of the Art and Research ChallengesabstractIntelligent automotive electronics significantly improved driving safety in the last decades. With the increasing complexity of automotive systems, dependability of the electronic components themselves and of their interaction must be assured to avoid any risk to driving safety due to unexpected failures caused by internal or external faults. Jan-Hendrik Oetjens, Nico Bannow, Markus Becker 0001, Oliver Bringmann 0001, Andreas Burger, Moomen Chaari, Samarjit Chakraborty, Rolf Drechsler, Wolfgang Ecker, Kim Grüttner, Thomas Kruse, Christoph Kuznik, Hoang Minh Le 0001, Andreas Mauderer, Wolfgang Müller 0003, Daniel Mueller-Gritschneder, Frank Poppen, Hendrik Post, Sebastian Reiter 0003, Wolfgang Rosenstiel, S. Roth, Ulf Schlichtmann, Andreas von Schwerin, Bogdan-Andrei Tabacaru, Alexander Viehl |
DAC | 16 |
| 2014 | Deterministic Synthesis of Hybrid Application-Specific Network-on-Chip TopologiesabstractNetworks-on-Chip (NoCs) enable cost-efficient and effective communication between the processing elements inside modern systems-on-chip (SoCs). NoCs with regular topologies such as meshes, tori, rings, and trees are well suited for general-purpose many core SoCs. These topologies might prove suboptimal for SoCs with predefined application characteristics and traffic patterns. Such SoCs benefit from application-specific NoC topologies, designed and optimized according to the application characteristics. This paper proposes a synthesis approach for creating hybrid, application-specific NoCs from an input floorplan and a set of use cases, describing the applications running on the SoC. The method considers latency, port count, and link length constraints. It produces hybrid topologies that utilize both NoC routers and shared buses. Furthermore, the proposed approach can insert intermediate relay routers that act as bridges or repeaters and help to reduce the cost further. Finally, the approach creates a deadlock-free routing of the communication flows by either finding deadlock-free paths or by inserting virtual channels. The benefits of the proposed method are demonstrated by comparing it to state-of-the-art approaches on a generic and an industrial SoC examples. Vladimir Todorov, Daniel Mueller-Gritschneder, Helmut Reinig, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | Memory access reconstruction based on memory allocation mechanism for source-level simulation of embedded softwareabstractTo date, there still lacks a way to accurately simulate data memory accesses in source-level simulation (SLS) of host-compiled embedded SW. The difficulty lies in that the accessed addresses for the load and store instructions can not be statically determined. Without knowing those addresses, the source code can not be annotated appropriately for data cache simulation. In this paper, we show an approach that is capable of resolving the accessed memory addresses based on the memory allocation mechanism. Applying this approach, the source code can be annotated to perform precise data cache simulation. The novelty of our methodology is that it is the first of its kind to take the memory allocation mechanism into account and thus can handle all the stack, data, heap and text sections. Moreover, a method is also proposed to handle pointer dereferences. In experiments, SLS with our approach yields almost identical cache miss rate and pattern when compared to the reference simulation. Kun Lu 0005, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
ASP-DAC | 2 |
| 2013 | Fast cache simulation for host-compiled simulation of embedded softwareabstractHost-compiled simulation has been proposed for software performance estimation, because of its high simulation speed. However, the simulation speed may be significantly lowered due to the cache simulation overhead. In this paper, we propose an approach that can reduce much of the cache simulation overhead, while still calculating cache misses precisely. For instruction cache, we statically analyze possible cache conflicts and perform cache conflicts aware annotation for host-compiled simulation. Within loops, the conflicts are dynamically captured by tagging the basic blocks instead of performing the expensive cache simulation. In this way, a vast majority of the cache accesses can be saved from simulation. For data cache, aggregated cache simulation is used for a large data block. Further, the data locality can be bound by considering the data allocation principle of a program. Experiments show that our approach improves the speed of host-compiled simulation by one order of magnitude, while providing the cache miss numbers with high accuracy. Kun Lu 0005, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DATE | 2 |
| 2013 | Analytical timing estimation for temporally decoupled TLMs considering resource conflictsabstractTransaction level models (TLMs) can use temporal decoupling to increase the simulation speed. However, there is a lack of modeling support to time the temporally decoupled TLMs. In this paper, we propose a timing estimation mechanism for TLMs with temporal decoupling. This mechanism features an analytical model and novel delay formulas. Concepts such as resource usage and availability are used to derive the delay formulas. Based on them, a fast scheduling algorithm resolves resource conflicts and dynamically determines the timing of concurrent transaction sequences. Experiments show that the delay estimation formulas are capable of capturing the timing effects of resource conflicts. At the same time, the overhead of the scheduling algorithm is very low, hence the simulation speed remains high. Kun Lu 0005, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DATE | 2 |
| 2013 | A virtual prototyping platform for real-time systems with a case study for a two-wheeled robotabstractIn today's real-time system design, a virtual prototype can help to increase both the design speed and quality. Developing a virtual prototyping platform requires realistic modeling of the HW system, accurate simulation of the real-time SW, and integration with a reactive real-time environment. Such a VP simulation platform is often difficult to develop. In this paper, we propose a case-study of autonomous two-wheeled robot to show how to develop a virtual prototyping platform rapidly in SystemC/TLM to adequately aid in the design of this instable system with hard real-time constraints. Our approach is an integration of four major model components. Firstly, an accurate physical model of the robot is provided. Secondly, a virtual world is modeled in Java that offers a 3D environment for the robot to move in. Thirdly, the embedded control SW is developed. Finally, the overall HW system is modeled in SystemC at transaction level. This HW model wraps the physical model, interacts with the virtual world, and simulates the real-time SW by integrating an Instruction Set Simulator of the embedded CPU. By integrating these components into a platform, designers can efficiently optimize the embedded SW architecture, explore the design space and check real-time conditions for different system parameters such as buffer sizes, CPU frequency or cache configurations. Daniel Mueller-Gritschneder, Kun Lu 0005, Erik Wallander, Marc Greim, Ulf Schlichtmann |
DATE | 1 |
| 2013 | A spectral clustering approach to application-specific network-on-chip synthesisabstractModern System-on-Chip (SoC) design relies heavily on efficient interconnects like Networks-on-Chip (NoCs). They provide an effective, flexible and cost efficient way of communication exchange between the individual processing elements of the SoC. Therefore, the choice of topology and design of the NoC itself plays a crucial role in the performance of the system. Depending on the field of application, standard topologies like meshes, fat-trees, and tori might be suboptimal in terms of power consumption, latency and area. This calls for a custom topology design methodology, which is based on the requirements imposed by the application, function and the use-cases of the SoC in question. This work proposes a fast approach, which uses spectral clustering and cluster ensembles to partition the system using normalized cuts and insert the necessary routers. Then, by using delay-constrained minimum spanning trees, links between the individual routers are created, such that any present latency constraints are satisfied at minimum cost. Results from applying the methodology to a smartphone SoC are presented. Vladimir Todorov, Daniel Mueller-Gritschneder, Helmut Reinig, Ulf Schlichtmann |
DATE | 2 |
| 2013 | A greedy approach for latency-bounded deadlock-free routing path allocation for application-specific NoCsabstractCustom network-on-chip (NoC) structures have improved power and area metrics compared to regular NoC topologies for application-specific systems-on-a-chip (SoCs). The synthesis of an application-specific NoC is a combinatorial problem. This paper presents a novel heuristic for solving the routing path allocation step. Its main advantages are the support of realistic nonlinear cost estimation and the ability to handle latency constraints, which guarantee high performance of processing elements sensitive to communication delays. Additionally, the method generates deadlock-free routing by avoiding cycles in the channel dependency graph. The NoC is constructed sequentially in a greedy manner by selecting the routing path for each communication flow in such a way that the additional NoC HW resources are kept minimal. The routing path is found using a binary search cheapest bounded path (BSCBP) algorithm. The method is highly efficient and provides a NoC routing path allocation for a smart phone SoC with 25 processing elements and 96 flows in less than a minute. Pritpal S. Multani, Daniel Mueller-Gritschneder, Vladimir Todorov, Ulf Schlichtmann |
NOCS | 3 |
| 2012 | Accurately timed transaction level models for virtual prototyping at high abstraction levelabstractTransaction level modeling (TLM) improves the simulation performance by raising the abstraction level. In the TLM 2.0 standard based on OSCI SystemC, a single transaction can transfer a large data block. Due to such high abstraction, a great amount of information becomes invisible and thus timing accuracy can be degraded heavily. We present a methodology to accurately time such block transactions and achieve high simulation performance at the same time. First, before abstraction, a profiling process is performed on an instruction set simulator (ISS). Driver functions that implement the transfer of the data blocks are simulated. Several techniques are employed to trace the exact start and end of the driver functions as well as HW usages. Thus, a profile library of those driver functions can be constructed. Then, the application programs are host-compiled and use a single transaction to transfer a data block. A strategy is presented that efficiently estimates the timing of block transactions based on the profile library. It is the first method that takes into account caching effects that influence the timing of block transactions. Moreover, it ensures overall timing accuracy when integrated in other SW timing tools for full system simulation. Experimental results show that the block transactions are accurately timed, with average error less than 1%. At the same time, the simulation gain can be up to three orders of magnitude. Kun Lu 0005, Daniel Mueller-Gritschneder, Ulf Schlichtmann |
DATE | 2 |
| 2012 | Automated construction of a cycle-approximate transaction level model of a memory controllerabstractTransaction level (TL) models are key to early design exploration, performance estimation and virtual prototyping. Their speed and accuracy enable early and rapid System-on-Chip (SoC) design evaluation and software development. Most devices have only register transfer level (RTL) models that are too complex for SoC simulation. Abstracting these models to TL ones, however, is a challenging task, especially when the RTL description is too obscure or not accessible. This work presents a methodology for automatically creating a TL model of an RTL memory controller component. The device is treated as a black box and a multitude of simulations is used to obtain results, showing its timing behavior. The results are classified into conditional probability distributions, which are reused within a TL model to approximate the RTL timing behavior. The presented method is very fast and highly accurate. The resulting TL model executes approximately 1200 times faster, with a maximum measured average timing offset error of 7.66%. Vladimir Todorov, Daniel Mueller-Gritschneder, Helmut Reinig, Ulf Schlichtmann |
DATE | 2 |
| 2011 | Control-Flow-Driven Source Level Timing Annotation for Embedded Software Models on Transaction LevelabstractInstrumented software models feature a combination of software functionality as well as timing information to model execution times on embedded processors. They aim to replace instruction set simulators in virtual prototypes (VP) of embedded systems to improve simulation efficiency. In this work, a novel control flow mapping algorithm is presented to automatically generate timing annotations for instrumented software models. The method is based on the analysis of loop and control dependency properties of basic code blocks in the binary and source code control flow. With these properties, the method can find suitable positions to annotate the timing delay statements of binary code basic blocks into the source code. It shows high accuracy even in the case that the binary code is optimized during compilation. The paper also presents the novel idea of adding timing control statements into the source code to improve timing accuracy. The error in runtime estimation was found to be below 6\% for standard test programs. A case study for a VP shows a gain in simulation efficiency of three orders of magnitude compared to an ISS based model. Daniel Mueller-Gritschneder, Kun Lu 0005, Ulf Schlichtmann |
DSD | 1 |
| 2010 | Computation of yield-optimized Pareto fronts for analog integrated circuit specificationsabstractFor any analog integrated circuit, a simultaneous analysis of the performance trade-offs and impact of variability can be conducted by computing the Pareto front of the realizable specifications. The resulting Specification Pareto front shows the most ambitious specification combinations for a given minimum parametric yield. Recent Pareto optimization approaches compute a so-called yield-aware specification Pareto front by applying a two-step approach. First, the Pareto front is calculated for nominal conditions. Then, a subsequent analysis of the impact of variability is conducted. In the first part of this work, it is shown that such a two-step approach fails to generate the most ambitious realizable specification bounds for mismatch-sensitive performances. In the second part of this work, a novel single-step approach to compute yield-optimized specification Pareto fronts is presented. Its optimization objectives are the realizable specification bounds themselves. Experimental results show that for mismatch-sensitive performances the resulting yield-optimized specification Pareto front is superior to the yield-aware specification Pareto front. Daniel Mueller-Gritschneder, Helmut E. Graeb |
DATE | 1 |
| 2007 | Trade-off design of analog circuits using goal attainment and "Wave Front" sequential quadratic programmingabstractOne of the main tasks in analog design is the sizing of the circuit parameters, such as transistor lengths and widths, in order to obtain optimal circuit performances, such as high gain or low power consumption. In most cases one performance can only be optimized at cost of others, therefore a sizing must aim at an optimal trade-off between the important circuit performances. This paper presents a new deterministic method to calculate the complete range of performance trade-offs, the so-called Pareto-optimal front, of a given circuit topology. Known deterministic methods solve a set of constrained multi-objective optimization problems independently of each other. The presented method minimizes a set of goal attainment (GA) optimization problems simultaneously. In a parallel algorithm, the individual GA optimization processes compare and exchange their iterative solutions. This leads to a significant improvement in the efficiency and quality of analog trade-off design Daniel Mueller-Gritschneder, Helmut E. Graeb, Ulf Schlichtmann |
DATE | 1 |
| 2006 | A CPPLL hierarchical optimization methodology considering jitter, power and locking timeabstractIn this paper, a hierarchical optimization methodology for charge pump phase-locked loops (CPPLLs) is proposed. It has the following features: 1) A comprehensive and efficient behavioral modeling of the PLL enables fast simulations and includes the important PLL performances jitter, power and locking time, as well as stability constraints for the nonlinear locking process and the linear lock-in state; 2) Behavioral modeling of the PLL building blocks addresses as behavioral-level parameters: current and jitter of the charge pump (CP), gain, current and jitter of the voltage controlled oscillator (VCO), as well as R, C's of the loop filter (LF). It enables a proper propagation of PLL specifications down to the circuit-level design parameters; 3) An accurate and efficient performance space exploration technique on circuit level provides the feasible regions of the behavioral-level parameters of the building block by multidimensional Pareto-optimal fronts. This enables a first-time-successful top-down optimization process. Experimental results show the efficacy and efficiency of the presented method. The methodology can be applied to other large-scale analog circuits. Daniel Mueller-Gritschneder, Helmut E. Graeb, Ulf Schlichtmann |
DAC | 2 |
| 2006 | Fast evaluation of analog circuit structures by polytopal approximationsabstractIn this paper we present a method for the fast evaluation of circuit structures. It is part of a methodology for the structural synthesis of analog circuits which generates a large number of different circuit structures. Goal of the presented methods is to find circuit structures, which fit best the design goals. Based on implicit analog circuit specifications, as well as explicit performance specifications given by the designer, the presented method approximates the feasible region of parameters by a polytope. This polytopal approximation of the performance capabilities can be calculated and visualized or the feasibility of the specification can be tested by linear programming. The method has been validated on a set of operational amplifier structures Daniel Mueller-Gritschneder, Guido Stehr, Helmut E. Graeb, Ulf Schlichtmann |
ISCAS | 1 |
| 2005 | Deterministic approaches to analog performance space exploration (PSE)abstractPerformance space exploration (PSE) determines the range of feasible performance values of a circuit block for a given topology and technology. In this paper, we present two deterministic approaches for PSE. One approximates the feasible performance space based on linearized circuit models and is suitable for investigating a large number of performances. The other one computes discretizations of the Pareto front of competing performances. In addition, a motivation and application of PSE using a hierarchical design example is presented. Daniel Mueller-Gritschneder, Guido Stehr, Helmut E. Graeb, Ulf Schlichtmann |
DAC | 1 |