VLDB 2026 Research / reviewers in the wild / expert
Stefan Wildermann
dblp:24/5422
· DBLP profile ↗
53ranked-venue papers
7as first author
21since 2021 · last 2025
0000-0002-4324-2187ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 5 first-author · 18 since 2021Software engineering, systems software and programming languages · 8 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorTheory of computation · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Response Range Optimization for Run-Time Requirement Enforcement on MPSoCsabstractEmbedded system applications normally come with a set of nonfunctional requirements on execution properties (e.g., latency), expressed by a corridor of permissible values. These requirements should be guaranteed during each program execution on a given MPSoC platform. This can be achieved using a reactive control loop based on a requirement response, with an enforcer finite state machine (FSM) controlling the properties to be enforced, e.g., by adapting the number of cores allocated to a program or by scaling the voltage/frequency mode of active processors. A finer-grained control can be achieved using response ranges, which allow an enforcer to react based on the amount of violation of a requirement. But as the search space of enforcer FSMs to be explored by design space exploration (DSE) can be quite huge when jointly exploring transition relations together with response ranges of the transitions, we propose two heuristics for generating suitable response ranges prior to performing a DSE of proper enforcement FSMs. Our evaluation shows that the two proposed heuristics can generate efficient enforcement FSMs within a substantially smaller number of iterations (respectively time) compared to the case of using DSE to explore the joint space of transition relations and response ranges. Khalil Esper, Stefan Wildermann, Jürgen Teich |
ASP-DAC | 2 |
| 2025 | WiP Paper: Utility-Aware Transmission of Sensor Data on Energy-Harvesting IoT Gateways
Pierre-Louis Sixdenier, Jebacyril Arockiaraj, Stefan Wildermann, Jürgen Teich |
EWSN | 3 |
| 2024 | Range-Based Run-time Requirement Enforcement of Non-Functional Properties on MPSoCsabstractEmbedded system applications normally come with a set of non-functional requirements defined over properties (e.g., latency), expressed as a corridor of correct values via a lower and an upper bound per requirement. These requirements should be guaranteed during each execution of an application program on a given MPSoC platform. This can be achieved using a reactive control loop, where an enforcer controls a set of properties to be enforced, e.g., by adapting the number of cores allocated to a program or by scaling the voltage/frequency mode of active processors. An enforcement strategy may react on a requirement response differently, depending on (a) satisfying a requirement, or violating (b) a lower bound or (c) an upper bound. A better strategy might be to differentiate the reaction taken according to the amount of violation of a lower or upper bound, thus to react in a finer granular way. In this paper, we propose a design space exploration (DSE) method called Co-explore that automatically partitions the requirement corridors into so-called response ranges (i.e., sub-corridors) such that formulated verification goals of simultaneously generated enforcer FSMs, e.g., the number of consecutive violations of a requirement, are optimized. The evaluation shows that the explored enforcement FSMs can achieve higher probabilities of meeting a given set of requirements compared to reacting solely based on the ternary information (a), (b), or (c). Khalil Esper, Stefan Wildermann, Jürgen Teich |
DATE | 2 |
| 2024 | ABACUS: ASIP-Based Avro Schema-Customizable Parser Acceleration on FPGAsabstractBig Data applications frequently process data streams encoded in semi-structured data formats such as JSON, Protobuf, or Avro. Parsing these data formats into a representation to then be processed by a CPU frequently takes up a major share of the processing time. As a remedy, JSON and Avro FPGA accelerators have been introduced that can parse the data directly in the data path, offloading this workload from the CPU without requiring any additional data movement. However, these accelerators are schema-specific circuits that require time-consuming resynthesis processes for schema adaptations. This is particularly critical in Big Data applications, where multiple schemas may be in use simultaneously. As a remedy, we present an application-specific instruction set processor (ASIP) architecture for parsing Avro data on FPGAs. An instruction program controls the ASIP to parse a specific schema. Any schema change therefore only requires the loading of a new instruction sequence into an instruction memory. It is also shown that this approach is more resource-efficient than related work, as functional units only need to be instantiated once for each Avro data type. Our experimental evaluation shows that we can achieve a throughput of 707–818 MB/s per kLUT which is about 7 to 14 times higher than the throughput per LUT achieved in related work. Tobias Hahn, Daniel Schüll, Stefan Wildermann, Jürgen Teich |
DDECS | 3 |
| 2024 | JSON-CooP: A JSON Decompression/Parsing Co-Design for FPGAsabstractBig Data applications frequently involve the processing of data streams encoded in semi-structured data formats such as JSON. A major challenge here is that the parsing of such data formats is usually highly complex. Accelerating JSON parsing on FPGAs has therefore become a focus of recent research. However, as JSON data is highly sparse, compression is frequently applied before writing records to storage or transmitting them over the network. Consequently, the data must be decompressed before parsing it into a suitable format for further processing.While existing work has addressed decompression and parsing of semi-structured data separately, we propose a co-design for JSON decompression/parsing. This co-design includes a compression scheme tailored specifically for JSON, along with an FPGA parser architecture capable of operating directly on compressed input data. Additionally, our design adopts a lazy decompression approach to only decompress projected attributes, thereby significantly reducing the workload on the decompressor. Our experimental evaluation shows that this co-design exploiting several synergies results in higher resource efficiency compared to existing approaches for JSON parsing on FPGAs, while also enabling the decompression of data. For high compression factors, we observe a 1.8 x increase in parsed JSON tuples per second and LUT compared to the most efficient related approach. Moreover, our presented compression scheme offers higher compression factors for JSON data than related work while providing similar performance. Tobias Hahn, Stefan Wildermann, Jürgen Teich |
FPL | 2 |
| 2024 | A Scenario-Based DVFS-Aware Hybrid Application Mapping Methodology for MPSoCsabstractSound techniques for mapping soft real-time applications to resources are indispensable for meeting the application deadlines and minimizing objectives such as energy consumption, particularly on heterogeneous MPSoC architectures. For applications with input-dependent workload variations, static mappings are not able to sufficiently cope with the run-time variation, which can lead to deadline misses or unnecessary energy consumption. As a remedy, hybrid application mapping (HAM) techniques combine a design-time optimization with run-time management that adapts the mappings dynamically to the changes of the arriving input. This paper focuses on scenario-based HAM techniques. Here, the application input space is systematically clustered such that data inside the same scenario exhibit similar characteristics concerning workload when being processed under the same operating points. This static clustering of the input space into data scenarios has proven to be a good abstraction layer for simplifying the design and employment of high-quality run-time managers. However, existing state-of-the-art scenario-based HAM approaches neglect or underutilize the synergistic interplay between mapping selection and the usage of dynamic voltage/frequency scaling (DVFS) when adapting to workload variation. By combining mapping and DVFS selection, variations in the input can be either compensated by a complete re-mapping of the application, evoking a potential high reconfiguration overhead or by just changing the DVFS settings of the resources, offering a low-overhead adaptation alternative and thus significantly reducing the necessary overhead compared to DVFS-agnostic HAM. Furthermore, DVFS enables a fine-grained adaptation of a mapped application to the input data variation, e.g., by slowing down tasks with no impact on the end-to-end latency for the current input using low-frequency DVFS settings. It is shown that this combined approach can save even more energy than a pure mapping adaptation scheme, especially in the presence of data scenarios. In particular, scenario-based design operates as a catalyst for eliciting the synergies between a combined DVFS and mapping optimization and the peculiarities inside a data scenario, i.e., exploiting the commonalities inside a data scenario by perfectly tailored DVFS settings and task mapping. In this scope, this paper proposes two supplementary scenario-based DVFS-aware HAM approaches that consistently outperform existing state-of-the-art mapping approaches in terms of the number of deadline misses and energy consumption as we demonstrate in an empirical study on the basis of four different applications and three different architectures. It is also shown that these benefits still apply to target architectures with increasing mapping migration overheads, thwarting frequent mapping reconfigurations. Jan Spieck, Stefan Wildermann, Jürgen Teich |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2024 | Design, Calibration, and Evaluation of Real-time Waveform Matching on an FPGA-based Digitizer at 10 GS/sabstractDigitizing side-channel signals at high sampling rates produces huge amounts of data, while side-channel analysis techniques only need those specific trace segments containing Cryptographic Operations (COs). For detecting these segments, waveform-matching techniques have been established comparing the signal with a template of the CO’s characteristic pattern. Real-time waveform matching requires highly parallel implementations as achieved by hardware design but also reconfigurability as provided by Field-Programmable Gate Arrays (FPGAs) to adapt the matching hardware to a specific CO pattern. However, currently proposed designs process the samples from analog-to-digital converters sequentially and can only process low sampling rates due to the limited clock speed of FPGAs. In this article, we present a parallel waveform-matching architecture capable of performing high-speed waveform matching on a high-end FPGA-based digitizer. We also present a workflow for calibrating the waveform-matching system to the specific pattern of the CO in the presence of hardware restrictions provided by the FPGA hardware. Our implementation enables waveform matching at 10 GS/s, offering a speedup of 50× compared to the fastest state-of-the-art implementation known to us. We demonstrate how to apply the technique for attacking the widespread XTS-AES algorithm using waveform matching to recover the encrypted tweak even in the presence of so-called systemic noise. Jens Trautmann 0001, Paul Krüger, Andreas Becher, Stefan Wildermann, Jürgen Teich |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2023 | Special Session - Non-Volatile Memories: Challenges and Opportunities for Embedded System Architectures with Focus on Machine Learning ApplicationsabstractThis paper explores the challenges and opportunities of integrating non-volatile memories (NVMs) into embedded systems for machine learning. NVMs offer advantages such as increased memory density, lower power consumption, non-volatility, and compute-in-memory capabilities. The paper focuses on integrating NVMs into embedded systems, particularly in intermittent computing, where systems operate during periods of available energy. NVM technologies bring persistence closer to the CPU core, enabling efficient designs for energy-constrained scenarios. Next, computation in resistive NVMs is explored, highlighting its potential for accelerating machine learning algorithms. However, challenges related to reliability and device non-idealities need to be addressed. The paper also discusses memory-centric machine learning, leveraging NVMs to overcome the memory wall challenge. By optimizing memory layouts and utilizing probabilistic decision tree execution and neural network sparsity, NVM-based systems can improve cache behavior and reduce unnecessary computations. In conclusion, the paper emphasizes the need for further research and optimization for the widespread adoption of NVMs in embedded systems presenting relevant challenges, especially for machine learning applications. Jörg Henkel, Lokesh Siddhu, Lars Bauer, Jürgen Teich, Stefan Wildermann, Mehdi Baradaran Tahoori, Mahta Mayahinia, Jerónimo Castrillón, Asif Ali Khan, Hamid Farzaneh, João Paulo C. de Lima, Jian-Jia Chen, Christian Hakert, Kuan-Hsun Chen, Chia-Lin Yang, Hsiang-Yun Cheng |
CASES | 5 |
| 2023 | SPEAR-JSON: Selective Parsing of JSON to Enable Accelerated Stream Processing on FPGAsabstractBig Data applications frequently involve the processing of data streams encoded in semi-structured data formats such as JSON. A major challenge here is that the parsing of such data formats is usually highly complex. Accelerating JSON parsing on FPGAs has therefore become a focus of recent research. FPGA accelerators were presented which serve as a co-processor for a CPU to convert JSON into a format that is easier for the CPU to process, e.g., Apache Arrow. However, in case the parsed data should be further processed on the FPGA, such solutions are insufficient as the format created is unsuitable for further processing on FPGAs and, above all, because the accelerators have an immense resource requirement. In this paper, we present a novel FPGA parser architecture that is able to interpret JSON data to selectively extract attributes based on a query expression into a format suitable for stream processing on FPGAs. Furthermore, it is shown how the sparsity of JSON can be used to implement a resource-efficient design, only requiring few FPGA resources. This leaves the major share of resources free for accelerating subsequent processing steps of a given application. Our experimental evaluation shows that we can achieve a throughput of 51.1 MB/s per kLUT which is about 3.8 times higher than the throughput per LUT achievable on the most efficient related approach. Tobias Hahn, Stefan Wildermann, Jürgen Teich |
FPL | 2 |
| 2023 | Hybrid Genetic Reinforcement Learning for Generating Run-Time Requirement Enforcers
Jan Spieck, Pierre-Louis Sixdenier, Khalil Esper, Stefan Wildermann, Jürgen Teich |
MEMOCODE | 4 |
| 2023 | Automatic Synthesis of FSMs for Enforcing Non-functional Requirements on MPSoCs Using Multi-objective Evolutionary AlgorithmsabstractEmbedded system applications often require guarantees regarding non-functional properties when executed on a given MPSoC platform. Examples of such requirements include real-time, energy, or safety properties on corresponding programs. One option to implement the enforcement of such requirements is by a reactive control loop, where an enforcer decides based on a system response (feedback) how to control the system, e.g., by adapting the number of cores allocated to a program or by scaling the voltage/frequency mode of involved processors. Typically, a violation of a requirement must either never happen in case of strict enforcement, or only happen temporally (in case of so-called loose enforcement). However, it is a challenge to design enforcers for which it is possible to give formal guarantees with respect to requirements, especially in the presence of typically largely varying environmental input (workload) per execution. Technically, an enforcement strategy can be formally modeled by a finite state machine (FSM) and the uncertain environment determining the workload by a discrete-time Markov chain. It has been shown in previous work that this formalization allows the formal verification of temporal properties (verification goals) regarding the fulfillment of requirements for a given enforcement strategy. In this article, we consider the so-far-unsolved problem of design space exploration and automatic synthesis of enforcement automata that maximize a number of deterministic and probabilistic verification goals formulated on a given set of non-functional requirements. For the design space exploration (DSE), an approach based on multi-objective evolutionary algorithms is proposed in which enforcement automata are encoded as genes of states and state transition conditions. For each individual, the verification goals are evaluated using probabilistic model checking. At the end, the DSE returns a set of efficient FSMs in terms of probabilities of meeting given requirements. As experimental results, we present three use cases while considering requirements on latency and energy consumption. Khalil Esper, Stefan Wildermann, Jürgen Teich |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2023 | A Learning-based Methodology for Scenario-aware Mapping of Soft Real-time Applications onto Heterogeneous MPSoCsabstractSoft real-time streaming applications often process input data that evoke varying workloads for their tasks. This may lead to high energy consumption or deadline misses in case their mapping onto a heterogeneous MPSoC target architecture is not adapted, e.g., when tasks with high execution times for the current input are assigned to resources of low computational power. To handle the vast variety of different input data, we propose to cluster data with similar execution characteristics into so-called data scenarios for which we determine specialized mappings by performing a scenario-aware design space exploration (DSE). A runtime manager (RTM) uses these mappings to adapt the execution of the running applications to their upcoming input by first identifying their best-suited scenarios. Subsequently, the RTM selects mappings considering their identified scenarios, which minimize the total number of deadline misses and the consumed energy. We embed the RTM into hybrid application mapping (HAM); ergo, performing time-consuming optimizations offline. In this article, we propose a novel data-scenario-aware HAM methodology that can cope with multiple applications and comprises two novel scenario-based mapping selection algorithms: Inter-Application Resource Mediation Mapping introduces barely any runtime overhead. Adaptive multi-app mapping selection is highly adaptive to changes in the application workload but imposes a small runtime overhead. Our HAM approach is fully automated and uses machine-learning techniques to learn the selection of suitable mappings from training data sequences at design time. Experiments on three differently complex target architectures show that our proposed approach consistently outperforms existing state-of-the-art solutions regarding the number of deadline misses and consumed energy. Jan Spieck, Stefan Wildermann, Jürgen Teich |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2022 | Raw Filtering of JSON Data on FPGAsabstractMany Big Data applications include the processing of data streams on semi-structured data formats such as JSON. A disadvantage of such formats is that an application may spend a significant amount of processing time just on unselectively parsing all data. To relax this issue, the concept of raw filtering is proposed with the idea to remove data from a stream prior to the costly parsing stage. However, as accurate filtering of raw data is often only possible after the data has been parsed, raw filters are designed to be approximate in the sense of allowing false-positives in order to be implemented efficiently. Contrary to previously proposed CPU-based raw filtering techniques that are restricted to string matching, we present FPGA-based primitives for filtering strings, numbers and also number ranges. In addition, a primitive respecting the basic structure of JSON data is proposed that can be used to further increase the accuracy of introduced raw filters. The proposed raw filter primitives are designed to allow for their composition according to a given filter expression of a query. Thus, complex raw filters can be created for FPGAs which enable a drastical decrease in the amount of generated false-positives, particularly for IoT workload. As there exists a trade-off between accuracy and resource consumption, we evaluate primitives as well as composed raw filters using different queries from the RiotBench benchmark. Our results show that up to 94.3% of the raw data can be filtered without producing any observed false-positives using only a few hundred LUTs. Tobias Hahn, Andreas Becher, Stefan Wildermann, Jürgen Teich |
DATE | 3 |
| 2022 | Characterization of Side Channels on FPGA-based Off-The-Shelf Boards against Automated AttacksabstractFPGAs offer fast and reliable near-data processing and are therefore suitable candidates for implementing IoT and edge computing systems. As they are usually deployed in exposed locations, they are vulnerable to physical attacks, especially Side-Channel Analysis (SCA).In this paper, we characterize side-channels and how they can be exploited for SCA on FPGA-based off-the-shelf boards, i.e. without having to make any modifications to the board, hardware, or software. The basic requirement for any kind of SCA is that the individual Cryptographic Operations (COs) in the side-channel traces can be detected.To this end, we apply a SCA for semi-automatic CO detection that can be generically applied off-the-shelf to a wide variety of boards. Additionally, we introduce a new metric called Signal of COs to Noise Ratio (SCONR), that allows to quantify the pronouncedness of COs versus noise in a side channel. We then evaluate side channels measured on three different boards containing Xilinx 7 series FPGAs. We further investigate the influence of other sources of noise and how much they affect the attackability of a system.Our results show that FPGAs have a high vulnerability to SCA in general and that even noise from an operating system will not hinder the recording and finding of COs in an automated fashion as long as there are no countermeasures in place. Finally, SCONR converges after fewer recorded traces and gives a clearer indication whether a side channel is susceptible to this type of automated attack than leakage assessment techniques such as TVLA. Jens Trautmann 0001, Jürgen Teich, Stefan Wildermann |
FCCM | 3 |
| 2022 | Real-Time Waveform Matching with a Digitizer at 10 GS/sabstractSide-Channel Analysis (SCA) requires the detection of the specific time frame within which Cryptographic Operations (COs) take place in the side-channel signal. In laboratory conditions with full control over the Device under Test (DuT), dedicated trigger signals can be implemented to indicate the start and end of COs. For real-world scenarios, waveform-matching techniques have been established which compare the side-channel signal with a template of the CO's pattern in real time to detect the CO in the side channel. State-of-the-art approaches are implemented on Field-Programmable Gate Arrays (FPGAs). However, current waveform-matching designs process the samples from Analog-to-Digital Converters (ADCs) sequentially and can only work with low sampling rates due to the limited clock speed of FPGAs. This makes it increasingly difficult to apply existing techniques on modern DuTs that operate with clock speeds in the GHz range. In this paper, we present a parallel waveform-matching architecture that is capable of performing waveform matching at the speed of fast ADCs. We implement the proposed architecture in a high-end FPGA-based digitizer and deploy it to detect AES COs from the side channel of a single-board computer operating at 1 GHz. Our implementation allows for waveform matching at 10 GS/s with high accuracy, thus offering a speedup of 50× compared to the fastest state-of-the-art implementation known to us. Jens Trautmann 0001, Nikolaos Patsiatzis, Andreas Becher, Jürgen Teich, Stefan Wildermann |
FPL | 5 |
| 2022 | Auto-Tuning of Raw Filters for FPGAsabstractMany Big Data applications include the processing of data streams on semi-structured data formats such as JSON. A disadvantage of these formats, however, is that applications may require a significant portion of their processing time to unselectively parse all data. As a remedy, so-called raw filters have been introduced in the past, aiming to reduce the data load before the costly parsing stage. Since filtering unparsed data can also become very costly, raw filters can be designed to filter data approximately, in the sense that they allow false positives to occur, in order to be implemented efficiently. While previously proposed CPU-based solutions are restricted to just string filtering, FPGA approaches have recently been proposed with much more expressive raw filters, allowing also to capture numbers and structural relationships. Yet, as a consequence of the variety of filter possibilities as well as the limited amount of resources available on FPGAs, the selection of optimal filters before their deployment has been identified as a complex problem resulting in the potential need to select less expressive filters in order to consume fewer resources. Many Big Data applications (e.g., stream processing) operate on incoming real-time data over long, potentially unlimited time periods. As a consequence, the conditions for which such a filter is optimized can change over time after its deployment. In this realm, this paper presents a new methodology which automatically adapts the hardware accelerator for raw filtering by means of dynamic hardware reconfiguration. Data is sampled on-the-fly during operation and used by an optimizer-in-the-loop to select and generate a raw filter with optimized selectivity for these data samples. As the optimizer has to take into account the resource costs of the hardware accelerator, we introduce models to estimate the resource costs in order to avoid performing a full synthesis. The filter selection problem can thus be solved within a few minutes with results close to the accurate resource cost estimation. If the selectivity of a query changes over time, such as seasonal differences in the analysis of IoT data, the system can auto-tune its filter to adapt to the situation. Depending on the query and the variability of inherent data changes, significant improvements in the amount of filtered data are presented, resulting in a significant parsing speedup in comparison to a state-of-the-art non-adaptive approach. Tobias Hahn, Stefan Wildermann, Jürgen Teich |
FPL | 2 |
| 2022 | On Transferring Application Mapping Knowledge Between Differing MPSoC ArchitecturesabstractThe mapping of soft real-time applications onto a heterogeneous MPSoC target architecture is of high importance for meeting deadlines and minimizing secondary objectives like the consumed energy on the platform. In particular, for applications with input-dependent workload variation, hybrid application mapping (HAM) has crystallized itself as the state-of-the-art mapping approach that combines time-intensive design space exploration with lightweight run-time management. However, one general problem of HAM is that the explored mappings and models are highly specific to a certain target architecture. If the architecture is modified or changed (e.g., due to hardware upgrades, downgrades, or architectural degradation), the design-time optimization has to be repeated once again, which might take up to multiple weeks. As a remedy, this article proposes a twofold mapping transfer methodology that speeds up the design-time optimization for a novel or modified target architecture based on the knowledge we gained from the optimization for the original source architecture. First, we describe a greedy mapping transfer heuristic that provides feasible mappings for the new architecture in negligible optimization time. Second, we present a mapping refinement heuristic that improves these mappings even further while needing only a fraction of the optimization time of state-of-the-art approaches. As we show in the evaluation section, our approach can drastically speed up the convergence of the optimization for the novel architecture even if the sizes of the original and novel target architecture or the characteristics of the used resources differ significantly. Note that our approach assumes that the source and target architectures are both tile-based Network-on-Chip (NoC) meshes. Jan Spieck, Stefan Wildermann, Jürgen Teich |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Design and Evaluation of a Tunable PUF Architecture for FPGAsabstractFPGA-based Physical Unclonable Functions (PUF) have emerged as a viable alternative to permanent key storage by turning effects of inaccuracies during the manufacturing process of a chip into a unique, FPGA-intrinsic secret. However, many fixed PUF designs may suffer from unsatisfactory statistical properties in terms of uniqueness, uniformity, and robustness. Moreover, a PUF signature may alter over time due to aging or changing operating conditions, rendering a PUF insecure in the worst case. As a remedy, we propose CHOICE , a novel class of FPGA-based PUF designs with tunable uniqueness and reliability characteristics. By the use of addressable shift registers available on an FPGA, we show that a wide configuration space for adjusting a device-specific PUF response is obtained without any sacrifice of randomness. In particular, we demonstrate the concept of address-tunable propagation delays, whereby we are able to increase or decrease the probability of obtaining “ 1 ”s in the PUF response. Experimental evaluations on a group of six 28 nm Xilinx Artix-7 FPGAs show that CHOICE PUFs provide a large range of configurations to allow a fine-tuning to an average uniqueness between 49% and 51%, while simultaneously achieving bit error rates below 1.5%, thus outperforming state-of-the-art PUF designs. Moreover, with only a single FPGA slice per PUF bit, CHOICE is one of the smallest PUF designs currently available for FPGAs. It is well-known that signal propagation delays are affected by temperature, as the operating temperature impacts the internal currents of transistors that ultimately make up the circuit. We therefore comprehensively investigate how temperature variations affect the PUF response and demonstrate how the tunability of CHOICE enables us to determine configurations that show a high robustness to such variations. As a case study, we present a cryptographic key generation scheme based on CHOICE PUF responses as device-intrinsic secret and investigate the design objectives resource costs, performance, and temperature robustness to show the practicability of our approach. Franz-Josef Streit, Paul Krüger, Andreas Becher, Stefan Wildermann, Jürgen Teich |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2021 | Approximate Logic Synthesis of Very Large Boolean NetworksabstractFor very large Boolean circuits most approximate logic synthesis techniques successively apply local approximation transformations affecting only a portion of the whole design. Hence, allowing such transformations to be implemented in polynomial time and to gain better control of the introduced error. A key issue here is to derive efficient techniques for selecting from all the possible portions of the design those more likely to yield better trade-offs between hardware resources and quality. Due to the likelihood of error masking growing with increasing circuit complexity, we expect the likelihood of a local transformation reaching-or being observable at-the primary outputs to decrease at a similar rate. Comparatively, the closer a portion undergoing a local transformation is to the primary outputs, the more likely the error introduced can be observed at the primary outputs. Based on this observation, this paper proposes a novel methodology for the selection of portions-or sub-functions-of Boolean circuits-represented by Boolean networks-for approximation according to their degree of connectivity with other portions of the design. Our selection criterion is based on that a Boolean sub-function shall be a better candidate for approximation when it drives many other sub-functions, especially those being driven by many other sub-functions. We introduce, integrate, and compare our connectivity-based selection methodology with a state-of-the-art approximate logic synthesis framework. Experimental results show that our selection technique yields better trade-offs between hardware resources and accuracy of the resulting approximated circuits. Moreover, our technique is efficient and can speed up the design space exploration of the aforementioned framework. Jorge Echavarria, Stefan Wildermann, Jürgen Teich |
DATE | 2 |
| 2021 | Choice - A Tunable PUF-Design for FPGAsabstractFPGA-based Physical Unclonable Functions (PUFs) have emerged as a viable alternative to permanent key storage by turning inaccuracies during the manufacturing process of a chip into a unique, FPGA-intrinsic secret. However, many fixed PUF designs may suffer from unsatisfactory statistical properties in terms of uniqueness, uniformity, and robustness. Moreover, a PUF signature may alter over time due to aging or changing operating conditions, rendering a PUF insecure in the worst case. As a remedy, we propose CHOICE, a novel class of FPGA-based PUF designs with tunable uniqueness and reliability characteristics. By the use of addressable shift registers available on an FPGA, we show that a wide configuration space for adjusting a device-specific PUF response is obtained without any sacrifice of randomness. In particular, we demonstrate the concept of address-tunable propagation delays, whereby we are able to increase or decrease the probability of obtaining 1’s in the PUF response. Experimental evaluations on a group of six 28 nm Xilinx Artix-7 FPGAs show that CHOICE PUFs provide a large range of configurations to allow a fine-tuning to an average uniqueness between 49% and 51%, while simultaneously achieving bit error rates below 1.5%, thus outperforming state-of-the-art PUF designs. Moreover, with only a single FPGA slice per PUF bit, CHOICE is one of the smallest PUF designs currently available for FPGAs. Franz-Josef Streit, Paul Krüger, Andreas Becher, Jens Trautmann 0001, Stefan Wildermann, Jürgen Teich |
FPL | 5 |
| 2021 | Enforcement FSMs: specification and verification of non-functional properties of program executions on MPSoCsabstractMany embedded system applications impose hard real-time, energy or safety requirements on corresponding programs typically concurrently executed on a given MPSoC target platform. Even when mutually isolating applications in space or time, the enforcement of such properties, e.g., by adjusting the number of processors allocated to a program or by scaling the voltage/frequency mode of involved processors, is a difficult problem to solve, particularly in view of typically largely varying environmental input (workload) per execution. In this paper, we formalize the related control problem using finite state machine models for the uncertain environment determining the workload, the system response (feedback), as well as the enforcer strategy. The contributions of this paper are as follows: a) Rather than trace-based simulation, the uncertain environment is modeled by a discrete-time Markov chain (DTMC) as a random process to characterize possible input sequences an application may experience. b) A number of important verification goals to analyze different enforcer FSMs are formulated in PCTL for the resulting stochastic verification problem, i.e., the likelihood of violating a timing or energy constraint, or the expected number of steps for a system to return to a given execution time corridor. c) Applying stochastic model checking, i.e., PRISM to analyze and compare enforcer FSMs in these properties, and finally d) proposing an approach for reducing the environment DTMC by partitioning equivalent environmental states (i.e., input states leading to an equal system response in each MPSoC mode) such that verification times can be reduced by orders of magnitude to just a few ms for real-world examples. Khalil Esper, Stefan Wildermann, Jürgen Teich |
MEMOCODE | 2 |
| 2020 | Run-Time Enforcement of Non-Functional Application Requirements in Heterogeneous Many-Core SystemsabstractFor many embedded applications, non-functional requirements such as safety, reliability, and execution time must be guaranteed in tight bounds on a given multi-core platform. Here, jitter in non-functional program execution qualities is caused either by outer influences such as faults injected by the environment, but can be induced also from the system management software itself, including thread-to-core mapping, scheduling and power management. A second huge source of variability typically stems from data-dependent workloads. In this paper, we classify and present techniques to enforce nonfunctional execution properties on multi-core platforms. Based on a static design space exploration and analysis of influences of variability of non-functional properties, enforcement strategies are generated to guide the execution of periodically executed applications in given requirement corridors. Using the case study of a complex image streaming application, we show that by controlling DVFS settings of cores proactively, not only tight execution times, but also reliability requirements may be enforced dynamically while trying to minimize energy consumption. Jürgen Teich, Behnaz Pourmohseni, Oliver Keszöcze, Jan Spieck, Stefan Wildermann |
ASP-DAC | 5 |
| 2020 | Probabilistic Error Propagation through Approximated Boolean NetworksabstractMost approximate logic synthesis techniques successively apply local approximate transformations to Boolean circuits. Naturally, an efficient, robust, and scalable error estimation technique is due. This paper addresses this problem by propagating error probabilities within a network of circuits, each circuit being described by an approximated Boolean function. We specifically tackle error rate, that is, the likelihood of a logic network evaluating to an erroneous output. Our simulation-free error rate estimation technique is fully accurate when there are no mutual dependencies among signals in the Boolean network-also known as fanout-reconvergence-and shows a neglectable inaccuracy lying within 1% with respect to exhaustively simulated values for benchmark designs including signal correlations. Moreover, our methodology is capable of computing the error rate in the order of milliseconds for every tested benchmark, allowing the proposed error analysis to be applied during design space exploration. For comparison, we finally applied our methodology to a state-of-the-art approximate logic synthesis framework showing its superiority in terms of quality and runtime. Jorge Echavarria, Stefan Wildermann, Oliver Keszöcze, Jürgen Teich |
DAC | 2 |
| 2020 | Scenario-Based Soft Real-Time Hybrid Application Mapping for MPSoCsabstractFor soft real-time applications, a fixed mapping to a heterogeneous MPSoC architecture can lead to high energy consumption and even deadline misses if tasks have input-dependent execution times. Here, specialized mappings are required that, e.g., map tasks with high execution times for the current input to resources with high computational power as they else may cause deadline misses. However, optimizing mappings for both energy and latency at run time is too compute-intensive. As a remedy, we propose a hybrid application mapping technique suited for blackbox applications, i.e., no information about functional behaviors is available. It is based on clustering input data evoking similar workloads into so-called workload scenarios. At design time, we optimize the scenario distribution and their associated mappings regarding energy consumption and latency by an iterative design space exploration. At run time, a machine-learning-based runtime manager first identifies the scenario of the current input by monitoring its non-functional execution properties. Based on these identified scenarios, a mapping for subsequent data processing is selected so that missed deadlines and the energy are minimized. Evaluations performed based on two dynamic applications show that the proposed hybrid application mapping procedure consistently outperforms state-of-the-art mapping approaches with regard to both deadline misses and energy consumption. Jan Spieck, Stefan Wildermann, Jürgen Teich |
DAC | 2 |
| 2020 | SQL Query Processing Using an Integrated FPGA-based Near-Data Accelerator in ReProVide
Lekshmi B. G., Andreas Becher, Klaus Meyer-Wegener, Stefan Wildermann, Jürgen Teich |
EDBT | 4 |
| 2019 | Isolation-Aware Timing Analysis and Design Space Exploration for Predictable and Composable Many-Core SystemsabstractComposable many-core systems enable the independent development and analysis of applications which will be executed on a shared platform where the mix of concurrently executed applications may change dynamically at run time. For each individual application, an off-line DSE is performed to compute several mapping alternatives on the platform, offering Pareto-optimal trade-offs in terms of real-time guarantees, resource usage, etc. At run time, one mapping is then chosen to launch the application on demand. In this context, to enable an independent analysis of each individual application at design time, so-called inter-application isolation schemes are applied which specify temporal/spatial isolation policies between applications. State-of-the-art composable many-core systems are developed based on a fixed isolation scheme that is exclusively applied to every resource in every mapping of every application and use a timing analysis tailored to that isolation scheme to derive timing guarantees for each mapping. A fixed isolation scheme, however, heavily restricts the explored space of solutions and can, therefore, lead to suboptimality. Lifting this restriction necessitates a timing analysis that is applicable to mappings with an arbitrary mix of isolation schemes on different resources. To address this issue, in this paper, we (a) present an isolation-aware timing analysis that - unlike existing analyses - can handle multiple isolation schemes in combination within one mapping and delivers safe yet tight timing bounds by identifying and excluding interference scenarios that can never happen under the given combination of isolation schemes. Based on the timing analysis, we (b) present a DSE which explores the choices of isolation scheme per resource within each mapping and uses the proposed timing analysis for timing verification. Experimental results demonstrate that, for a variety of real-time applications and many-core platforms, the proposed approach achieves an improvement of up to 67% in the quality of delivered mappings compared to approaches based on a fixed isolation scheme. Behnaz Pourmohseni, Fedor Smirnov, Stefan Wildermann, Jürgen Teich |
ECRTS | 3 |
| 2019 | Thermally Composable Hybrid Application Mapping for Real-Time Applications in Heterogeneous Many-Core SystemsabstractModern embedded many-core systems host, among others, real-time applications which must be dynamically launched at run time. To this end, Hybrid Application Mapping (HAM) methodologies combine design-time analysis with runtime mapping techniques to enable dynamic application mapping with performance guarantees, e.g., w.r.t. real-time constraints. They rely on composability to derive the required performance guarantees in an isolated analysis of individual applications at design time. The ongoing process technology downsizing, however, has given rise to an increased on-chip temperature, so that the thermal integrity of the platform must be monitored and enforced at run time by means of Dynamic Thermal Management (DTM) techniques which use countermeasures e.g. DVFS and power gating. This, however, violates composability, as the thermally unsafe behavior of one application may trigger DTM countermeasures that affect other applications running in the thermally affected region which, in turn, may lead to the violation of their real-time constraints. As a remedy, this paper proposes, for the first time, a thermally composable HAM methodology that enforces thermal safety proactively at the launch time of applications and, thereby, prevents DTM interferences which react to thermal violations. To that end, we present (a) a novel thermal-safety analysis that can be integrated into the design-time analysis of HAM and (b) a set of thermal-safety admission checks that can be used at run time when launching an application. By establishing thermal composability among running applications, the proposed HAM approach enables providing thermally safe real-time guarantees for dynamically mapped applications in many-core systems. Experimental results for a variety of hard real-time applications on multiple heterogeneous many-core architectures demonstrate the efficiency and effectiveness of the proposed methodology. Behnaz Pourmohseni, Fedor Smirnov, Heba Khdr, Stefan Wildermann, Jürgen Teich, Jörg Henkel |
RTSS | 4 |
| 2019 | Hard real-time application mapping reconfiguration for NoC-based many-core systems
Behnaz Pourmohseni, Stefan Wildermann, Michael Glaß, Jürgen Teich |
Real Time Syst. | 2 |
| 2019 | Compilation of Dataflow Applications for Multi-Cores using Adaptive Multi-Objective OptimizationabstractState-of-the-art system synthesis techniques employ meta-heuristic optimization techniques for Design Space Exploration (DSE) to tailor application execution, e.g., defined by a dataflow graph, for a given target platform. Unfortunately, the performance evaluation of each implementation candidate is computationally very expensive, in particular on recent multi-core platforms, as this involves compilation to and extensive evaluation on the target hardware. Applying heuristics for performance evaluation on the one hand allows for a reduction of the exploration time but on the other hand may deteriorate the convergence of the optimization technique toward performance-optimal solutions with respect to the target platform. To address this problem, we propose DSE strategies that are able to dynamically trade off between (i) approximating heuristics to guide the exploration and (ii) accurate performance evaluation, i.e., compilation of the application and subsequent performance measurement on the target platform. Technically, this is achieved by introducing a set of additional, but easily computable guiding objective functions, and varying the set of objective functions that are evaluated during the DSE adaptively. One major advantage of these guiding objectives is that they are generically applicable for dataflow models without having to apply any configuration techniques to tailor their parameters to the specific use case. We show this for synthetic benchmarks as well as a real-world control application. Moreover, the experimental results demonstrate that our proposed adaptive DSE strategies clearly outperform a state-of-the-art DSE approach known from literature in terms of the quality of the gained implementations as well as exploration times. Amongst others, we show a case for a two-core implementation where after about 3 hours of exploration time one of our proposed adaptive DSE strategies already obtains a 60% higher performance value than obtained by the state-of-the-art approach. Even when the state-of-the-art approach is given a total exploration time of more than 2 weeks to optimize this value, the proposed adaptive DSE strategy features a 20% higher performance value after a total exploration time of about 4 days. Tobias Schwarzer, Joachim Falk, Martín Letras, Christian Heidorn, Stefan Wildermann, Jürgen Teich |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2018 | Architecture decomposition in system synthesis of heterogeneous many-core systemsabstractDetermining feasible application mappings for Design Space Exploration (DSE) and run-time embedding is a challenge for modern many-core systems. The underlying NP-complete system-synthesis problem faces tremendously complex problem instances due to the hundreds of heterogeneous processing elements, their communication infrastructure, and the resulting number of mapping possibilities. Thus, we propose to employ a search-space splitting (SSS) technique using architecture decomposition to increase the performance of existing design-time and run-time synthesis approaches. The technique first restricts the search for application embeddings to selected sub-architectures at substantially reduced complexity; therefore, the complete architecture needs to be searched only in case no embedding is found on any sub-system. Furthermore, we introduce a basic learning mechanism to detect promising sub-architectures and subsequently restrict the search to those. We exemplify the SSS for a SAT-based and a problem-specific backtracking-based system synthesis as part of DSE for NoC-based many-core systems. Experimental results show drastically reduced execution times (≈ 15--50 x on a 24×24 architecture) and an enhanced quality of the embedding, since less mappings (≈20--40 x, compared to the non-decomposing procedures) need to be discarded due to a timeout. Valentina Richthammer, Tobias Schwarzer, Stefan Wildermann, Jürgen Teich, Michael Glaß |
DAC | 3 |
| 2018 | Optimistic regular expression matching on FPGAs for near-data processingabstractRegular expressions (regex) are the main means to search for specific patterns in the vast amount of stored textual information. As a consequence, different designs of hardware accelerators have been proposed that enable memory-bound regex processing. Here, the regular expression to be evaluated is translated to a non-deterministic (NFA) or deterministic finite automaton (DFA) which is then mapped onto the hardware design. The available hardware resources of the design imply the maximum size (in terms of amount of states and transitions) of the supported automata. However, regular expressions may be arbitrarily complex. Andreas Becher, Stefan Wildermann, Jürgen Teich |
DaMoN | 2 |
| 2018 | AConFPGA: A Multiple-Output Boolean Function Approximation DSE Technique Targeting FPGAsabstractNew relaxed quality standards laid down by approximate computing enrich the design pool with architectures dissipating less power, consuming fewer resources or with smaller latencies. In LUT-based FPGA logic approximation, the number of LUTs and latency associated to a design can be optimized by allowing the approximation of circuit results. In this paper, we present techniques for automatic design space exploration (DSE) of Boolean function falsifications and the ability and impact to reduce resources usage as well as the length of critical paths on LUT-based FPGAs. Our experiments give evidence that resource reductions of about 20% are easily achievable for error rates amounting to less than 0.05% w.r.t. accurate designs. Jorge Echavarria, Stefan Wildermann, Jürgen Teich |
FPT | 2 |
| 2018 | Design space exploration of multi-output logic function approximationsabstractApproximate Computing has emerged as a design paradigm that allows to decrease hardware costs by reducing the accuracy of the computation for applications that are robust against such errors. In Boolean logic approximation, the number of terms and literals of a logic function can be reduced by allowing to produce erroneous outputs for some input combinations. This paper proposes a novel methodology for the approximation of multi-output logic functions. Related work on multi-output logic approximation minimizes each output function separately. In this paper, we show that thereby a huge optimization potential is lost. As a remedy, our methodology considers the effect on all output functions when introducing errors thus exploiting the cross-function minimization potential. Moreover, our approach is integrated into a design space exploration technique to obtain not only a single solution but a Pareto-set of designs with different trade-offs between hardware costs (terms and literals) and error (number of minterms that have been falsified). Experimental results show our technique is very efficient in exploring Pareto-optimal fronts. For some benchmarks, the number of terms could be reduced from an accurate function implementation by up to 15% and literals by up to 19% with degrees of inaccuracy around 0.1% w.r.t. accurate designs. Moreover, we show that the Pareto-fronts obtained by our methodology dominate the results obtained when applying related work. Jorge Echavarria, Stefan Wildermann, Jürgen Teich |
ICCAD | 2 |
| 2018 | Dynamic resource management for heterogeneous many-coresabstractWith the advent of many-core systems, use cases of embedded systems have become more dynamic: Plenty of applications are concurrently executed, but may dynamically be exchanged and modified even after deployment. Moreover, resources may temporally or permanently become unavailable because of thermal aspects, dynamic power management, or the occurrence of faults. This poses new challenges for reaching objectives like timeliness for real-time or performance for best-effort program execution and maximizing system utilization. In this work, we first focus on dynamic management schemes for reliability/aging optimization under thermal constraints. The reliability of on-chip systems in the current and upcoming technology nodes is continuously degrading with every new generation because transistor scaling is approaching its fundamental limits. Protecting systems against degradation effects such as circuits' aging comes with considerable losses in efficiency. We demonstrate in this work why sustaining reliability while maximizing the utilization of available resources and hence avoiding efficiency loss is quite challenging – this holds even more when thermal constraints come into play. Then, we discuss techniques for run-time management of multiple applications which sustain real-time properties. Our solution relies on hybrid application mapping denoting the combination of design-time analysis with run-time application mapping. We present a method for Real-time Mapping Reconfiguration (RMR) which enables the Run-Time Manager (RM) to execute realtime applications even in the presence of dynamic thermal- and reliability-aware resource management. This paper is paper of the ICCAD 2018 Special Session on “Managing Heterogeneous Many-cores for High-Performance and Energy-Efficiency”. The other two papers of this Special sessions are [1] and [2]. Jörg Henkel, Jürgen Teich, Stefan Wildermann, Hussam Amrouch |
ICCAD | 3 |
| 2018 | Symmetry-Eliminating Design Space Exploration for Hybrid Application Mapping on Many-Core ArchitecturesabstractLarge scale many-core systems are able to execute concurrently changing mixes of different parallel applications. Hybrid application mapping combines the strengths of design-time exploration/analysis of resource constellations for task-to-core mappings with the flexibility of choosing concrete mappings at run time. However, state-of-the-art design space exploration (DSE) techniques so far ignore the problem of symmetries in modern heterogeneous architectures: not only recurring patterns in the architecture but the mapping of tasks to instances of the same processor type may unnecessarily increase the search space by redundant, symmetrical implementations which typically affects the quality of the DSE. As a remedy, we propose a novel meta-heuristic DSE approach that eliminates architectural symmetries by abstracting the problem to a clustering of tasks and their mapping to processor types. However, we demonstrate that simple task clustering and type mappings may again introduce encoding symmetries in our search space. Thus, we present a formulation of the task clustering and type mapping as a 0-1 integer linear program (ILP) which eliminates all architectural as well as encoding symmetries from the search space. We also contribute a formal feasibility check to ensure that only implementations with at least one feasible concrete mapping are considered. To further improve the search process for feasible solutions, we apply satisfiability modulo theories-like learning techniques: from each infeasible implementation, we extract conditions why the implementation is infeasible and enrich our 0-1 ILP by additional constraints continuously during the DSE. Experimental results show that a DSE equipped with the novel symmetry-eliminating search space and the proposed learning techniques clearly outperforms a state-of-the-art approach known from literature in terms of the quality of the gained implementation classes. Tobias Schwarzer, Andreas Weichslgartner, Michael Glaß, Stefan Wildermann, Peter Brand, Jürgen Teich |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | A Design-Time/Run-Time Application Mapping Methodology for Predictable Execution Time in MPSoCsabstractExecuting multiple applications on a single MPSoC brings the major challenge of satisfying multiple quality requirements regarding real-time, energy, and so on. Hybrid application mapping denotes the combination of design-time analysis with run-time application mapping. In this article, we present such a methodology, which comprises a design space exploration coupled with a formal performance analysis. This results in several resource reservation configurations, optimized for multiple objectives, with verified real-time guarantees for each individual application. The Pareto-optimal configurations are handed over to run-time management, which searches for a suitable mapping according to this information. To provide any real-time guarantees, the performance analysis needs to be composable and the influence of the applications on each other has to be bounded. We achieve this either by spatial or a novel temporal isolation for tasks and by exploiting composable networks-on-chip (NoCs). With the proposed temporal isolation, tasks of different applications can be mapped to the same resource, while, with spatial isolation, one computing resource can be exclusively used by only one application. The experiments reveal that the success rate in finding feasible application mappings can be increased by the proposed temporal isolation by up to 30% and energy consumption can be reduced compared to spatial isolation. Andreas Weichslgartner, Stefan Wildermann, Deepak Gangadharan, Michael Glaß, Jürgen Teich |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2017 | Exploiting Predictability in Dynamic Network Communication for Power-Efficient Data Transmission in LTE Radio SystemsabstractIn embedded systems powered by batteries, power is undoubtedly a critical resource making power management an important topic in the design phase. Even though power management is a heavily researched topic, most approaches focus on improving the way the power manager reacts to outside control events. In this paper, we propose techniques that not only react but rather try to predict these outside control events in advance, thus, broadening the capabilities of any employed power manager by allowing for superior transition decisions and even saving redundant calculations. We present results on employing a predictive power management system that couples a classic dynamic power manager with a machine learning subsystem in the context of a mobile device in a Long Term Evolution (LTE) system, with emphasis on evaluating the potential of saving power as well as the handling of the induced prediction uncertainty. First, we examine the LTE communication protocol and showcase certain control data that has to be received periodically, but may contain no information for the receiver. Finally, we show a proof-of-concept based on real LTE traces and hardware simulation, that prediction of this information can be leveraged to allow for a far superior decision process compared to a non-predicting system. Here, we achieve a theoretical best case power saving of 15 % for an idealized prediction with 100 % accuracy and no additional power consumption. Peter Brand, Jonathan Ah Sue, Johannes Brendel, Joachim Falk, Ralph Hasholzner, Jürgen Teich, Stefan Wildermann |
SCOPES | 7 |
| 2017 | Automatic Conversion of Simulink Models to SysteMoC Actor NetworksabstractSimulink has gained a lot of acceptance due to its intuitive through block-based algorithm design, simulation, and rapid prototyping capabilities for signal processing as well as control applications. However, automatic code generation for heterogeneous architectures is currently not supported by Simulink. In the literature, there exist automatic translation toolchains for generation of C or C++ code from Simulink models, which then are used for implementation or validation purposes. But few of them approach the generation of models that can be used in well-established Electronic System Level (ESL) design methodologies and tools. In order to address this issue, we present a methodology to extract an executable specification based on Data Flow Graphs (DFGs) from a given Simulink model. Such a specification can then be used by ESL tools to perform a Design Space Exploration (DSE) and generate code for hardware/software partitions directly from the ESL model. In a case study from signal processing, we validate the equivalence of the results of the simulation in Simulink and the results obtained by simulation of the DFG fully automatically generated from the Simulink model in the SystemC-based actor language SysteMoC. Martín Letras, Joachim Falk, Stefan Wildermann, Jürgen Teich |
SCOPES | 3 |
| 2017 | Self-Adaptive FPGA-Based Image Processing Filters Using Approximate ArithmeticsabstractApproximate Computing aims at trading off computational accuracy against improvements regarding performance, resource utilization and power consumption by making use of the capability of many applications to tolerate a certain loss of quality. A key issue is the dependency of the impact of approximation on the input data as well as user preferences and environmental conditions. In this context, we therefore investigate the concept of self-adaptive image processing that is able to autonomously adapt 2D-convolution filter operators of different accuracy degrees by means of partial reconfiguration on Field-Programmable-Gate-Arrays (FPGAs). Experimental evaluation shows that the dynamic system is able to better exploit a given error tolerance than any static approximation technique due to its responsiveness to changes in input data. Additionally, it provides a user control knob to select the desired output quality via the metric threshold at runtime. Jutta Pirkl, Andreas Becher, Jorge Echavarria, Jürgen Teich, Stefan Wildermann |
SCOPES | 5 |
| 2016 | A LUT-Based Approximate AdderabstractIn this paper, we propose a novel approximate adder structure for LUT-based FPGA technology. Compared with a full featured accurate carry-ripple adder, the longest path is significantly shortened which enables the clocking with an increased clock frequency. By using the proposed adder structure, the throughput of an FPGA-based implementation can be significantly increased. On the other hand, the resulting average error can be reduced compared to similar approaches for ASIC implementations. Andreas Becher, Jorge Echavarria, Daniel Ziener, Stefan Wildermann, Jürgen Teich |
FCCM | 4 |
| 2016 | FAU: Fast and error-optimized approximate adder units on LUT-Based FPGAsabstractDuring the design of embedded systems, many design decisions have to be made to trade off between conflicting objectives such as cost, performance, and power. Approximate computing allows to optimize each objective, yet for the sake of accuracy. This means that a functional flaw is allowed to produce an error as long as this is small enough to maintain a feasible operation of the system or guarantee a certain accuracy of the results. In this paper, we propose a new technique for approximate addition optimized for LUT-Based FPGAs with segmented carry chains. Our optimized adder structure is able to a) best exploit artifacts of LUT-Based FPGAs such as unused inputs and b) provide a smaller average error than previously proposed approximate adder structures, as well as c) a reduced critical path delay than dedicated accurate logic in modern FPGAs. We present a novel stochastic error calculus that is able to take into account also non-uniform input distributions and present a detailed comparison of approximate adder structures proposed in literature with our novel LUT-Based approximate arithmetic structure. Jorge Echavarria, Stefan Wildermann, Andreas Becher, Jürgen Teich, Daniel Ziener |
FPT | 2 |
| 2016 | Design-Time/Run-Time Mapping of Security-Critical Applications in Heterogeneous MPSoCsabstractDifferent applications concurrently running on modern MPSoCs can interfere with each other when they use shared resources. This interference can cause side channels, i.e., sources of unintended information flow between applications. To prevent such side channels, we propose a hybrid mapping methodology that attempts to ensure spatial isolation, i.e., a mutually-exclusive allocation of resources to applications in the MPSoC. At design time and as a first step, we compute compact and connected application mappings (called shapes). In a second step, run-time management uses this information to map multiple spatially segregated shapes to the architecture. We present and evaluate a (fast) heuristic and an (exact) SAT-based mapper, demonstrating the viability of the approach. Andreas Weichslgartner, Stefan Wildermann, Johannes Götzfried, Felix C. Freiling, Michael Glaß, Jürgen Teich |
SCOPES | 2 |
| 2014 | Multi-objective distributed run-time resource management for many-coresabstractDynamic usage scenarios of many-core systems require sophisticated run-time resource management that can deal with multiple often conflicting application and system objectives. This paper proposes an approach based on nonlinear programming techniques that is able to trade off between objectives while respecting targets regarding their values. We propose a distributed application embedding for dealing with soft system-wide constraints as well as a centralized one for strict constraints. The experiments show that both approaches may significantly outperform related heuristics. Stefan Wildermann, Michael Glaß, Jürgen Teich |
DATE | 1 |
| 2013 | Game-theoretic analysis of decentralized core allocation schemes on many-core systemsabstractMany-core architectures used in embedded systems will contain hundreds of processors in the near future. Already now, it is necessary to study how to manage such systems when dynamically scheduling applications with different phases of parallelism and resource demands. A recent research area called invasive computing proposes a decentralized workload management scheme of such systems: applications may dynamically claim additional processors during execution and release these again, respectively. In this paper, we study how to apply the concepts of invasive computing for realizing decentralized core allocation schemes in homogeneous many-core systems with the goal of maximizing the average speedup of running applications at any point in time. A theoretical analysis based on game theory shows that it is possible to define a core allocation scheme that uses local information exchange between applications only, but is still able to provably converge to optimal results. The experimental evaluation demonstrates that this allocation scheme reduces the overhead in terms of exchanged messages by up to 61.4% and even the convergence time by up to 13.4% compared to an allocation scheme where all applications exchange information globally with each other. Stefan Wildermann, Tobias Ziermann, Jürgen Teich |
DATE | 1 |
| 2012 | Distributed self-organizing bandwidth allocation for priority-based bus communicationabstractSUMMARY The raising complexity in distributed embedded systems makes it necessary that the communication of such systems organizes itself automatically. In this paper, we tackle the problem of sharing bandwidth on priority‐based buses. Based on a game theoretic model, reinforcement learning algorithms are proposed that use simple local rules to establish bandwidth sharing. The algorithms require little computational effort and no additional communication. Extensive experiments show that the proposed algorithms establish the desired properties without global knowledge ortextita priori information. It is proven that communication nodes using these algorithms can co‐exist with nodes using other scheduling techniques. Finally, we propose a procedure that helps to set the learning parameters according to the desired behavior. Copyright © 2011 John Wiley & Sons, Ltd. Tobias Ziermann, Stefan Wildermann, Nina Mühleis, Jürgen Teich |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | Unifying Partitioning and Placement for SAT-Based Exploration of Heterogeneous Reconfigurable SoCsabstractHeterogeneous reconfigurable SoCs provide more flexibility, maintainability, and reusability than hardwired SoCs. Designing such systems is a complex task, since early decisions, as design partitioning, influence the subsequent design steps, such as placement of partially reconfigurable modules In this paper, we investigate a symbolic design space exploration (DSE) approach for this kind of SoCs, where we transform the problem of finding a feasible implementation to a Boolean satisfiability problem (SAT). We present three encoding variants which unify partitioning and placement to overcome the drawbacks of their separation. In particular, we will show that the runtime of DSE can be speeded up when we perform a preprocessing mechanism that identifies those partitionings which inevitably lead to infeasibility, and then incorporate this information into the symbolic encoding for calculating feasible placements. Our experiments show the effectiveness of our SAT-based approach and compare the presented encoding variants. Stefan Wildermann, Jürgen Teich, Daniel Ziener |
FPL | 1 |
| 2011 | Operational mode exploration for reconfigurable systems with multiple applicationsabstractModern embedded systems incorporate multiple applications that run on the same execution platform. However, due to limited resources and other constraints, not all the combinations of applications may run concurrently. This paper tackles the problem of determining which combinations of applications can run on a given hardware architecture without violating given constraints, thus creating feasible operational modes of the system. The architecture itself may include standard processors for software implementation of applications and dynamically (partially) reconfigurable hardware resources, which enable dynamic sharing of the resource in different operational modes. The paper describes the models, theoretical results, and the mode exploration algorithm to perform this task. Here, the specification is symbolically encoded so that the feasibility of modes can be tested by applying a SAT solver. In the experiment, we demonstrate how to apply our approach to build a self-organizing smart camera framework. Stefan Wildermann, Felix Reimann, Jürgen Teich, Zoran A. Salcic |
FPT | 1 |
| 2011 | Dynamic decentralized mapping of tree-structured applications on NoC architecturesabstractThis paper presents a novel application-driven and resource-aware mapping methodology for tree-structured streaming applications onto NoCs. This includes strategies for mapping the source of streaming applications (seed point selection), as well as embedding strategies so that each process autonomously embeds its own succeeding tasks. The proposed embedding strategies only consider the local view of neighboring cells on the NoC which allows to significantly reduce computation and monitoring overhead. Our vision is that this approach facilitates self-organizing embedded systems that provide the flexibility and fault-tolerance required in future silicon technologies. The results provided in this paper show that our local and decentralized algorithms can compete with previously presented global and centralized algorithms. Andreas Weichslgartner, Stefan Wildermann, Jürgen Teich |
NOCS | 2 |
| 2010 | Self-organizing Computer Vision for Robust Object Tracking in Smart Cameras
Stefan Wildermann, Andreas Oetken, Jürgen Teich, Zoran A. Salcic |
ATC | 1 |
| 2010 | A Bus-Based SoC Architecture for Flexible Module Placement on Reconfigurable FPGAsabstractThis paper proposes an FPGA-based System-on-Chip (SoC) architecture with support for dynamic runtime reconfiguration. The SoC is divided into two parts, the static embedded CPU sub-system and the dynamically reconfigurable part. An additional bus system connects the embedded CPU sub-system with modules within the dynamic area, offering a flexible way to communicate among all SoC components. This makes it possible to implement a reconfigurable design with support for free module placement. An enhanced memory access method is included for high-speed access to an external memory. The dynamic part includes a streaming technology which implements a direct connection between reconfigurable modules. The paper describes the architecture and shows the advantages in a smart camera case study. Andreas Oetken, Stefan Wildermann, Jürgen Teich, Dirk Koch |
FPL | 2 |
| 2009 | CAN+: A new backward-compatible Controller Area Network (CAN) protocol with up to 16× higher data ratesabstractAs the number of electronic components in automobiles steadily increases, the demand for higher communication bandwidth also rises dramatically. Instead of installing new wiring harnesses and new bus structures, it would be useful, if already available structures could be used, but driven at higher data rates. In this paper, we a) propose an extension of the well-known Controller Area Network (CAN) called CAN+ with which the target rate of 1Mbit/s can be increased up to 16 times. Moreover, b) existing CAN hardware and devices not dedicated to these boosted data rates can still be used without interferences on communication. The major idea is a change of the protocol. In particular, we exploit the fact that data could be sent in time slots, where CAN-conform nodes don't listen. Finally, c) an implementation of this type of overclocking scheme on an FPGA is provided to prove the feasibility and the impressive throughput gains. Tobias Ziermann, Stefan Wildermann, Jürgen Teich |
DATE | 2 |
| 2009 | Self-organizing multi-cue fusion for FPGA-based embedded imagingabstractSelf-organization is a natural concept that helps complex systems to adapt themselves autonomically to their environment. In this paper, we present a self-organizing framework for multi-cue fusion in embedded imaging. This means that several simple image filters are used in combination to lead to a more robust system behavior. Human motion tracking serves as a show case. The system adapts to changes in the environment while tracking a person. Besides this, system customization can be simplified. The designer just has to select a desired set of image filters for a given task. The system then finds the appropriate parameters, e.g., the weighting of different cues. With the option of partial re-configuration, FPGAs support this type of customization. An FPGA-based prototype implementation demonstrates the feasibility of this approach. Tracking and adaptation work in real-time with 25 FPS and a resolution of 640 times 480. Stefan Wildermann, Gregor Walla, Tobias Ziermann, Jürgen Teich |
FPL | 1 |
| 2008 | A Sequential Learning Resource Allocation Network for Image Processing ApplicationsabstractOnline adaptation is a key requirement for image processing applications when used in dynamic environments. In contrast to batch learning, where retraining is required each time a new observation occurs, sequential learning algorithms offer the ability to iteratively adapt the existing classifier. In this paper, we present a neural network architecture and a fast online learning algorithm that allow to use the class of resource allocation networks for such adaptive image processing applications. The network is based on receptive fields that are processed by RBF sub-nets. The learning algorithm builds such networks online by adding new units to the sub-nets each time novel input data is observed. For this, we define a global and a local novelty criterion. Experimental results show that the proposed network outperforms existing RAN algorithms when used for face detection and recognition and is competitive with existing classifiers. Stefan Wildermann, Jürgen Teich |
HIS | 1 |