VLDB 2026 Research / reviewers in the wild / expert
Lars Bauer
dblp:55/2100
· DBLP profile ↗
77ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0003-0253-4594ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 71 · 9 first-author · 19 since 2021Software engineering, systems software and programming languages · 21 · 2 first-author · 4 since 2021Computer networks · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIQARA: Mixed-Criticality Queue-based Architecture for Reconfigurable Accelerator PlatformsabstractCoexistence of safety-critical control functions and besteffort computations in mixed-criticality systems poses a challenge in resource allocation and scheduling, as high-criticality jobs must adhere to strict timing guarantees, while lower-criticality jobs should make effective use of available resources without compromising the system’s safety and predictability. This paper introduces MIQARA1, a mixed-criticality queue-based architecture designed for reconfigurable accelerator platforms. MIQARA efficiently combines software-programmable CPUs with reconfigurable hardware, utilizing a dynamic job pipeline, token-based dependency tracking, and out-of-order scheduling to optimize resource utilization. At the same time, MIQARA has been designed to satisfy real-time constraints. MIQARA is evaluated on four FPGA platforms: the Zed Board, DipForty board, ZCU102 board, all of which have ARM CPUs implemented on chip, and Arty A7 with a RISC-V soft-core processor, representing systems that rely on soft CPUs. Results demonstrate substantial performance gains, particularly in terms of execution speed, flexibility, and adaptability to mixed-criticality workloads. The integration of features such as a streaming network further illustrates MIQARA’s scalability to complex data-intensive applications, making it a compelling solution for embedded mixed-criticality systems. MIQARA requires a hardware overhead of 17.8% and achieves a speedup of up to 4×. Hassan Nassar, Martin Rapp, Lars Bauer, Mostafa Elshimy, Zeynep Demirdag, Jörg Henkel |
DATE | 3 |
| 2025 | Through Fabric: A Cross-world Thermal Covert Channel on TEE-enhanced FPGA-MPSoC SystemsabstractThe ever-evolving computing landscape gets more complex in every moment and the need for heterogeneous compute systems becomes more relevant. As the usability of such systems grew, finding methods for securing them became more relevant. Commercial vendors already introduced Trusted Execution Environments (TEEs) for those systems. TEEs serve the need for isolation, where sensitive data are processed in a secure world, and non-trusted applications are executed in the normal world. In this paper, we introduce Through Fabric: a novel attack against TEE-enhanced FPGA-MPSoCs. We show that existing benign hardware accelerators can be manipulated from the secure world to implement a temperature-based covert channel. We successfully run this attack on a commercial FPGA-MPSoC within the OP-TEE environment without additional access rights. We use an open-source implementation of AES for the accelerator and we reach a transmission speed of 2 bits per second with bit error rate of 1.9% and packet error rate of 4.3%. We are the first to show that a TEE can be bypassed on FPGA-MPSoCs via temperature-based covert channel communication. Hassan Nassar, Jeferson González-Gómez, Varun Manjunath, Lars Bauer, Jörg Henkel |
ASP-DAC | 4 |
| 2025 | Hardware/Software Co-Analysis for Worst Case Execution Time BoundsabstractEnsuring that safety-critical systems meet timing constraints is crucial to avoid disastrous failures. To verify that timing requirements are met, a worst-case execution time (WCET) bound is computed. However, traditional WCET tools require a predefined timing model for each target processor, which is not available when using custom instruction set extensions. We introduce a novel approach based on hardware-software coanalysis that employs an instrumented hardware description of the target processor, removing the requirement for a separate timing model. We demonstrate this approach by extending the FemtoRV32 Individua RISC-V processor with a custom instruction set extension and show that it accurately models the timing behavior of the resulting system. Can Joshua Lehmann, Lars Bauer, Hassan Nassar, Heba Khdr, Jörg Henkel |
DATE | 2 |
| 2024 | HBMorphic: FHE Acceleration via HBM-Enabled Recursive Karatsuba Multiplier on FPGAabstractCloud computing offers advantages such as seamless scalability and speedup of computation. Nevertheless, these benefits come with notable tradeoffs, e.g., processing sensitive data without compromising security. Fully Homomorphic Encryption (FHE) solves this by processing of encrypted data. In this work, we develop an FHE hardware accelerator that uses a custom control interface to maximally utilize the bandwidth of HBM, following the memory access patterns of FHE. Hassan Nassar, Lars Bauer, Jörg Henkel |
FCCM | 2 |
| 2024 | Covert-Hammer: Coordinating Power-Hammering on Multi-tenant FPGAs via Covert ChannelsabstractWith the rise of AI, end of Moore's law, and the digitization of public services, the demand for accelerated computing is growing. To address this demand, major cloud service providers like Amazon Web Services, Microsoft Azure, and Google Cloud Platform have incorporated FPGA instances into their infrastructure with efficient and adaptable resource allocation models. Interest is increasing in multi-tenant FPGAs, which enable multiple users to utilize FPGA resources concurrently, while the FPGA can be split into smaller sections, one per tenant. Nevertheless, it introduces significant security vulnerabilities. For instance, by configuring a malicious circuit in one tenant's section of the FPGA, attacks that cause faults or crash the entire FPGA become feasible, affecting other tenants. By splitting an FPGA into smaller fractions, a single tenant has less potential to cause catastrophic outcomes. However, in this paper, we propose another threat, which is to perform an attack where several malicious tenants coordinate an attack using an unintended covert channel. We practically verify this possibility and introduce such a synchronized and coordinated voltage drop attack from multiple malicious tenants. For synchronization, the malicious tenants use a voltage-based covert channel. Our results show that the communication is robust reaching less than 1% packet error rate and that the attack is successful and avoids state-of-the-art countermeasures. Hassan Nassar, Philipp Machauer, Dennis Gnad, Lars Bauer, Mehdi Baradaran Tahoori, Jörg Henkel |
FPGA | 4 |
| 2024 | Co-Designing NVM-based Systems for Machine Learning and In-memory Search ApplicationsabstractWith the rapid development of the Internet of Things, machine learning applications on edge devices with limited resources face challenges due to large data scales and irregular memory access patterns. Non-volatile memory (NVM) technologies provide promising solutions by offering larger capacity, low leakage power, and data persistence. In this paper, we discuss the potential of NVM technology in enhancing machine learning applications by improving energy efficiency and reducing latency through in-memory computation and different NVM write modes. The insights from this analysis provide valuable guidance to device researchers and system architects working to develop highperformance systems for machine learning and accelerators in large-scale search applications using NVMs. Jörg Henkel, Lokesh Siddhu, Hassan Nassar, Lars Bauer, Jian-Jia Chen, Christian Hakert, Tristan Taylan Seidl, Kuan-Hsun Chen, Xiaobo Sharon Hu, Mengyuan Li 0001, Chia-Lin Yang, Ming-Liang Wei |
ICCAD | 4 |
| 2024 | DoS-FPGA: Denial of Service on Cloud FPGAs via Coordinated Power HammeringabstractThe adoption of FPGA instances by major cloud service providers (CSPs) reflects the growing demand for accelerated and heterogeneous computing across various applications, e.g., AI. To improve the efficiency, utilization and virtualization, multi-tenant FPGAs allow multiple users to utilize FPGA resources concurrently, with each FPGA partition assigned to a separate tenant. However, this introduces significant security vulnerabilities, such as the potential for attacks by configuring a malicious circuit in one tenant's FPGA partition. One notable vulnerability is disrupting the FPGA's power distribution network, leading to faults or even crashing the entire FPGA, affecting other tenants. Usually, such an attack requires a considerable amount of resources. A naive solution would be splitting an FPGA into smaller fractions to reduce the potential for successful Power-Hammering by individual tenants and enhance the security. However, our paper demonstrates that even with smaller fractions per tenant, attacks can still occur. We propose the threat of coordinated attacks, where malicious tenants use an unintended covert channel between them. We practically validate this threat in a real cloud computing environment by introducing a synchronized and coordinated power-hammering attack from multiple malicious tenants. These tenants synchronize their actions using a voltage-based covert channel. Our results reveal the success of the attack, surpassing state-of-the-art countermeasures and detection mechanisms with a success rate exceeding 90%, compared to 30% for uncoordinated attacks. Hassan Nassar, Philipp Machauer, Lars Bauer, Dennis Gnad, Mehdi Baradaran Tahoori, Jörg Henkel |
ICCAD | 3 |
| 2024 | Balancing Security and Efficiency: System-Informed Mitigation of Power-Based Covert ChannelsabstractAs the digital landscape continues to evolve, the security of computing systems has become a critical concern. Power-based covert channels (e.g., thermal covert channel s (TCCs)), a form of communication that exploits the system resources to transmit information in a hidden or unintended manner, have been recently studied as an effective mechanism to leak information between malicious entities via the modulation of CPU power. To this end, dynamic voltage and frequency scaling (DVFS) has been widely used as a countermeasure to mitigate TCCs by directly affecting the communication between the actors. Although this technique has proven effective in neutralizing such attacks, it introduces significant performance and energy penalties, that are particularly detrimental to energy-constrained embedded systems. In this article, we propose different system-informed countermeasures to power-based covert channels from the heuristic and machine learning (ML) domains. Our proposed techniques leverage task migration and DVFS to jointly mitigate the channels and maximize energy efficiency. Our extensive experimental evaluation on two commercial platforms: 1) the NVIDIA Jetson TX2 and 2) Jetson Orin shows that our approach significantly improves the overall energy efficiency of the system compared to the state-of-the-art solution while nullifying the attack at all times. Jeferson González-Gómez, Mohammed Bakr Sikal, Heba Khdr, Lars Bauer, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | Meta-Scanner: Detecting Fault Attacks via Scanning FPGA Designs MetadataabstractWith the rise of the big data, processing in the cloud has become more significant. One method of accelerating applications in the cloud is to use field programmable gate arrays (FPGAs) to provide the needed acceleration for the user-specific applications. Multitenant FPGAs are a solution to increase efficiency. In this case, multiple cloud users upload their accelerator designs to the same FPGA fabric to use them in the cloud. However, multitenant FPGAs are vulnerable to low-level denial-of-service attacks that induce excessive voltage drops using the legitimate configurations. Through such attacks, the availability of the cloud resources to the nonmalicious tenants can be hugely impacted, leading to downtime and thus financial losses to the cloud service provider. In this article, we propose a tool for the offline classification to identify which FPGA designs can be malicious during operation by analysing the metadata of the bitstream generation step. We generate and test 475 FPGA designs that include 38% malicious designs. We identify and extract five relevant features out of the metadata provided from the bitstream generation step. Using ten-fold cross-validation to train a random forest classifier, we achieve an average accuracy of 97.9%. This significantly surpasses the conservative comparison with the state-of-the-art approaches, which stands at 84.0%, as our approach detects stealthy attacks undetectable by the existing methods. Hassan Nassar, Jonas Krautter, Lars Bauer, Dennis Gnad, Mehdi Baradaran Tahoori, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Special Session - Non-Volatile Memories: Challenges and Opportunities for Embedded System Architectures with Focus on Machine Learning ApplicationsabstractThis paper explores the challenges and opportunities of integrating non-volatile memories (NVMs) into embedded systems for machine learning. NVMs offer advantages such as increased memory density, lower power consumption, non-volatility, and compute-in-memory capabilities. The paper focuses on integrating NVMs into embedded systems, particularly in intermittent computing, where systems operate during periods of available energy. NVM technologies bring persistence closer to the CPU core, enabling efficient designs for energy-constrained scenarios. Next, computation in resistive NVMs is explored, highlighting its potential for accelerating machine learning algorithms. However, challenges related to reliability and device non-idealities need to be addressed. The paper also discusses memory-centric machine learning, leveraging NVMs to overcome the memory wall challenge. By optimizing memory layouts and utilizing probabilistic decision tree execution and neural network sparsity, NVM-based systems can improve cache behavior and reduce unnecessary computations. In conclusion, the paper emphasizes the need for further research and optimization for the widespread adoption of NVMs in embedded systems presenting relevant challenges, especially for machine learning applications. Jörg Henkel, Lokesh Siddhu, Lars Bauer, Jürgen Teich, Stefan Wildermann, Mehdi Baradaran Tahoori, Mahta Mayahinia, Jerónimo Castrillón, Asif Ali Khan, Hamid Farzaneh, João Paulo C. de Lima, Jian-Jia Chen, Christian Hakert, Kuan-Hsun Chen, Chia-Lin Yang, Hsiang-Yun Cheng |
CASES | 3 |
| 2023 | Smart Detection of Obfuscated Thermal Covert Channel Attacks in Many-core ProcessorsabstractIn thermal covert channel (TCC) attacks, malicious applications seek to leak private information in a stealthy and hard-to-detect manner. State-of-the-art approaches for TCC detection employ the Discrete Fourier Transform (DFT) combined with heuristics to identify possible channels. However, as we demonstrate in this paper, these approaches are limited when detecting short-duration attacks, where an attacker intentionally halts the transmission for a time interval to avoid the detection. In order to overcome this limitation of the state-of-the-art solutions, we propose the first detection method for short-duration TCC attacks. Our solution, Dotecca, is a machine learning-based technique that employs short windows of time-domain measurements instead of the DFT to detect TCCs. To evaluate our solution, we introduce a new obfuscated short-duration attack that disguises as a regular application from the perspective of a DFT spectrum. Our experiments show that the new obfuscated attack is able to remain undetected even under advanced DFT-based state-of-the-art detection approaches, reducing their detection accuracy to about 18 %. In contrast, our smart detection approach is able to detect state-of-the-art and new obfuscated attacks with an accuracy of 99 %. Moreover, our solution reduces the overhead of the DFT-based state-of-the-art solution by more than 14 ×. Jeferson González-Gómez, Mohammed Bakr Sikal, Heba Khdr, Lars Bauer, Jörg Henkel |
DAC | 4 |
| 2023 | Late Breaking Results: Configurable Ring Oscillators as a Side-Channel CountermeasureabstractSide-channel attacks are a threat to computing devices. In this work, we propose a novel countermeasure against power analysis side-channel attacks. This countermeasure uses ring oscillators with runtime-configurable chain lengths to generate noise to hide the effects of the secret intermediate values on the device’s power consumption. We develop our countermeasure to be compatible with a state-of-the-art of side-channel-attack detection mechanism. Therefore, our solution does not incur any extra area overhead as it uses a subset of the circuit needed for detection. We evaluate our countermeasure using the test vector leakage assessment test (TVLA test). When our countermeasure is active no side-channel leakage could be detected. Hassan Nassar, Simon Pankner, Lars Bauer, Jörg Henkel |
DAC | 3 |
| 2023 | The First Concept and Real-world Deployment of a GPU-based Thermal Covert Channel: Attack and CountermeasuresabstractThermal covert channel (TCC) attacks have been studied as a threat to CPU-based systems over recent years. In this paper, we propose a new type of TCC attack that for the first time leverages the Graphics Processing Unit (GPU) of a system to create a stealthy communication channel between two malicious applications. We evaluate our new attack on two different real-world platforms: a GPU-dedicated general computing platform and a GPU-integrated embedded platform. Our results are the first to show that a GPU-based thermal covert channel attack is possible. From our experiments, we obtain a transmission rate of up to 8.75 bps with a very low error rate of less than 2 % for a 12-bit packet size, which is comparable to CPU-based TCCs in the state of the art. Moreover, we show how existing state-of-the-art countermeasures for TCCs need to be extended to tackle the new GPU-based attack at the cost of added overhead. To reduce this overhead, we propose our own DVFS-based countermeasure which mitigates the attack, while causing$2\times$less performance loss than the state-of-the-art countermeasure on a set of compute-intensive GPU benchmark applications. Jeferson González-Gómez, Kevin Cordero-Zuñiga, Lars Bauer, Jörg Henkel |
DATE | 3 |
| 2023 | Memory Carousel: LLVM-Based Bitwise Wear Leveling for Nonvolatile Main MemoryabstractEmerging non-volatile memory yields, alongside many advantages, technical shortcomings, such as reduced cell lifetime. Although many wear-leveling approaches exist to extend the lifetime of such memories, usually a trade-off for the granularity of wear-leveling has to be made. Due to iterative write schemes (repeatedly sense and write), wear-out of memory in certain systems is directly dependent on the written bit value and thus can be highly imbalanced, requiring dedicated bit-wise wear-leveling. Such a bit-wise wear-leveling so far has only be proposed together with a special hardware support. However, if no dedicated hardware solutions are available, especially for commercial off-the-shelf systems with non-volatile memories, a software solution can be crucial for the system lifetime. In this work, we propose entirely software-based bit-wise wearleveling, where the position of bits within CPU words in main memory is rotated on a regular basis. We leverage the LLVM intermediate representation to adjust load and store operations of the application with a custom compiler pass. Experimental evaluation shows that the lifetime by applying local rotation within the CPU word can be extended by a factor of up to 21×. We also show that our method can incorporate with coarser-grained wear-leveling, e.g. on block granularity and assist achievement of higher lifetime improvements. Nils Hölscher, Christian Hakert, Hassan Nassar, Kuan-Hsun Chen, Lars Bauer, Jian-Jia Chen, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | ANV-PUF: Machine-Learning-Resilient NVM-Based Arbiter PUFabstractPhysical Unclonable Functions (PUFs) have been widely considered an attractive security primitive. They use the deviations in the fabrication process to have unique responses from each device. Due to their nature, they serve as a DNA-like identity of the device. But PUFs have also been targeted for attacks. It has been proven that machine learning (ML) can be used to effectively model a PUF design and predict its behavior, leading to leakage of the internal secrets. To combat such attacks, several designs have been proposed to make it harder to model PUFs. One design direction is to use Non-Volatile Memory (NVM) as the building block of the PUF. NVM typically are multi-level cells, i.e, they have several internal states, which makes it harder to model them. However, the current state of the art of NVM-based PUFs is limited to ‘weak PUFs’, i.e., the number of outputs grows only linearly with the number of inputs, which limits the number of possible secret values that can be stored using the PUF. To overcome this limitation, in this work we design the Arbiter Non-Volatile PUF (ANV-PUF) that is exponential in the number of inputs and that is resilient against ML-based modeling. The concept is based on the famous delay-based Arbiter PUF (which is not resilient against modeling attacks) while using NVM as a building block instead of switches. Hence, we replace the switch delays (which are easy to model via ML) with the multi-level property of NVM (which is hard to model via ML). Consequently, our design has the exponential output characteristics of the Arbiter PUF and the resilience against attacks from the NVM-based PUFs. Our results show that the resilience to ML modeling, uniqueness, and uniformity are all in the ideal range of 50%. Thus, in contrast to the state-of-the-art, ANV-PUF is able to be resilient to attacks, while having an exponential number of outputs. Hassan Nassar, Lars Bauer, Jörg Henkel |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | Cache-Based Side-Channel Attack Mitigation for Many-Core Distributed Systems via Dynamic Task MigrationabstractSide-channel attacks (SCA) are a serious threat to cryptographic systems due to mostly unavoidable information leakage. Cache-based SCAs take advantage of cache inherent timing properties on shared memory systems to extract security-critical information. In this paper, we present a novel approach to mitigate cache-based SCAs on distributed many-core systems, based on a resource management technique. Our solution leverages dynamic task migration as a mechanism to ensure a secure execution scenario for security-critical applications. Additionally, we propose a resource-management-based mechanism to ensure a secure execution when migration is not possible due to a lack of available resources. We evaluate our solution in terms of gained security and performance impact using the Sniper simulator for different configurations. Results show that our technique effectively develops resilience against SCAs, while causing a low performance slowdown (1.6% on average, 9% worst case). For all tested benchmarks, our worst-case performance slowdown is 20% less than a state-of-the-art countermeasure. Moreover, our solution utilizes less than 1 ms of system run-time overhead for a 64 core platform with 100% utilization. Jeferson González-Gómez, Lars Bauer, Jörg Henkel |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | ARMOR: A Reliable and Mobility-Aware RPL for Mobile Internet of Things InfrastructuresabstractMobile portable embedded devices are becoming an integral part of our daily activities in the vision of Internet of Things (IoT). Nevertheless, due to lack of mobility support in the IPv6 routing protocol for low-power and lossy networks (RPLs), which is standardized for multihop IoT infrastructures, providing reliable communications in terms of packet delivery ratio (PDR) in mobile IoT applications has become significantly challenging. While several studies tried to enhance the adaptability of RPL to network dynamics, their utilized routing metrics have prevented them from establishing long-lasting reliable paths. Furthermore, the stochastic parent replacement policy in the standard version of RPL has intensified this challenge. Aside from this, due to the existing tradeoff between reliability and power efficiency, most of the existing approaches have only concentrated on one of these concerns without paying attention to the other one. To address these issues, this article introduces ARMOR, a routing mechanism built upon RPL, which employs a novel mobility-aware routing metric, i.e., time to reside (TTR), and a corresponding parent replacement policy. According to the motion characteristics of the mobile objects, TTR provides an estimation of how long the nodes will be in the transmission range of each other. This enables ARMOR to select nodes, which provide longer connection period and consequently higher reliability. In comparison with the state of the art, while keeping the power consumption constant, ARMOR significantly improves the amount of PDR in the network by up to$2.5\times $, while it enhances the reliability against the original version of this protocol by up to$4.2\times $. Ali Asghar Mohammad Salehi, Bardia Safaei 0001, Amir Mahdi Hosseini Monazzah, Lars Bauer, Jörg Henkel, Alireza Ejlali |
IEEE Internet Things J. | 4 |
| 2022 | CaPUF: Cascaded PUF Structure for Machine Learning ResiliencyabstractWith the rise of the Internet of Things (IoT), resource-constrained and power-constrained devices attract more attention. The need for lightweight solutions as alternatives to resource-intensive applications became more urgent. Moreover, as the number of connected devices grew, authenticating them became more challenging. Traditionally, this would be performed by using hash functions and secure memory to store a key, which both come at a high cost. physical unclonable functions (PUFs) emerged as a suitable lightweight alternative to hash functions to authenticate the devices. Using the inherent minute differences between integrated circuits (ICs), they can generate IC-specific responses for input challenges coming from a so-called verifier. Through the years, machine learning (ML) has been used to attack PUFs by modeling them and accurately predicting their response to a given challenge. This stimulated research on ML-resilient PUFs. This resilience came with the significant area and challenge-to-response delay overheads. In this work, we introduce the novel cascaded PUF (CaPUF) and show that it is resilient against state-of-the-art ML-based attacks, i.e., logistic regression (LR) and support vector machines (SVMs). These attacks could not achieve accuracy better than 52% against our CaPUF, which is only as good as flipping a coin. Additionally, our CaPUF requires 89% less area compared to state-of-the-art ML-resilient PUFs. Hassan Nassar, Lars Bauer, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Software-Managed Read and Write Wear-Leveling for Non-Volatile Main MemoryabstractIn-memory wear-leveling has become an important research field for emerging non-volatile main memories over the past years. Many approaches in the literature perform wear-leveling by making use of special hardware. Since most non-volatile memories only wear out from write accesses, the proposed approaches in the literature also usually try to spread write accesses widely over the entire memory space. Some non-volatile memories, however, also wear out from read accesses, because every read causes a consecutive write access. Software-based solutions only operate from the application or kernel level, where read and write accesses are realized with different instructions and semantics. Therefore different mechanisms are required to handle reads and writes on the software level. First, we design a method to approximate read and write accesses to the memory to allow aging aware coarse-grained wear-leveling in the absence of special hardware, providing the age information. Second, we provide specific solutions to resolve access hot-spots within the compiled program code (text segment) and on the application stack. In our evaluation, we estimate the cell age by counting the total amount of accesses per cell. The results show that employing all our methods improves the memory lifetime by up to a factor of 955×. Christian Hakert, Kuan-Hsun Chen, Horst Schirmeier, Lars Bauer, Paul R. Genssler, Georg von der Brüggen, Hussam Amrouch, Jörg Henkel, Jian-Jia Chen |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2021 | TiVaPRoMi: Time-Varying Probabilistic Row-Hammer MitigationabstractRow-Hammering is a challenge for computing systems that use DRAM. It can cause bit flips in a DRAM row by accessing its neighboring rows. Several mitigation techniques on memory controller level were already suggested. The techniques are in two categories: The first category uses static probabilities, which leads to a performance penalty due to a high number of extra row activations. The second category is based on so-called Tabled Counters, which have large hardware requirements and are mostly infeasible to implement. We introduce a novel Row-Hammer mitigation technique that uses time-varying probabilities combined with a relatively small history table. Our technique reduces the number of extra row activations compared to static probabilistic techniques and it demands less storage than Tabled Counters techniques. Compared to state of the art, our technique offers a good compromise that has 9× - 27× reduced storage requirement than Tabled Counters and 6× - 12× fewer activations than probabilistic techniques. Hassan Nassar, Lars Bauer, Jörg Henkel |
DATE | 2 |
| 2021 | LoopBreaker: Disabling Interconnects to Mitigate Voltage-Based Attacks in Multi-Tenant FPGAsabstractFPGAs are being offered in the cloud as accelerator resources that can be shared among multiple users (i.e. tenants). Recently, various approaches have shown that fault attacks launched from one tenant region to another are possible, leading to timing faults or crashes of the FPGA. It is, therefore, important that malicious tenants are limited in their ability to cause such security problems. So far, the existing countermeasures against such attacks check the configuration bitstreams before they are reconfigured. Such offline approaches have various practical limitations, e.g. they may force the tenants to unveil their design secrets. In this paper, we present LoopBreaker, a novel runtime solution that can disable the entire activity of a malicious tenant region, in order to rapidly stop a potential attack before it results in a crash (i.e. Denial-of-Service). We implemented and tested multiple attack types and found that realistic attacks demand at least 12–26 µs to be successful. A partial reconfiguration to overwrite the malicious tenant region demands 200 µs in our realworld implementation, which is too slow to prevent the attack from leading to a crash. Instead, our proposed LoopBreaker method only needs 1.5 µs to stop a malicious tenant, which makes it the first online approach that can successfully stop challenging voltage drop-based attacks from causing a crash. Hassan Nassar, Hanna AlZughbi, Dennis Gnad, Lars Bauer, Mehdi Baradaran Tahoori, Jörg Henkel |
ICCAD | 4 |
| 2020 | Hierarchical Classification for Constrained IoT Devices: A Case Study on Human Activity RecognitionabstractThe massive number of Internet-of-Things (IoT) devices generates a hard-to-manage volume of data. Cloud-centric processing approaches for the IoT data suffer from high and unpredictable network latency, which causes poor experience in real-time IoT applications, such as healthcare. To address this issue, in edge computing, the data inference starts from the data source (i.e., the IoT devices). However, the constrained computational capabilities of the IoT device and the power-hungry data transmission demand a tradeoff between onboard processing and computation offloading. Hence, the IoT information inference requires efficient and lightweight techniques that are tailored for this tradeoff and respect the constrained resources on IoT devices, such as wearables. This article presents a hierarchical classification approach that decomposes the problem into three classifiers in two hierarchy layers. In the first layer, a lightweight classifier executes directly on the IoT device and decides whether to offload the computation to the gateway or to perform it onboard. The second layer comprises a lightweight classifier on the IoT device (can only distinguish a subset of classes) and a complex classifier on the gateway (to distinguish the remaining classes). The experimental results (using a real-world data set for human activity recognition and implemented on a wearable IoT device) show higher accuracy (92% on average) than a nonhierarchical classifier (87% on average). The execution time and power measurements on the IoT device show $3\times $ energy saving for the classification. Farzad Samie, Lars Bauer, Jörg Henkel |
IEEE Internet Things J. | 2 |
| 2020 | Fast Operation Mode Selection for Highly Efficient IoT Edge DevicesabstractIn the emerging paradigm of edge computing (EC) for Internet of Things (IoT), data processing is pushed to the edge of the IoT network (e.g., gateways and embedded IoT devices). IoT devices must support multiple operation modes in order to adapt to varying runtime situations, like preserving energy at low battery, while still maintaining some crucial functionality, etc. Adapting the optimal operation mode is a challenge for edge devices given the limited resources at the edge of the network (both bandwidth and processing power of the shared gateway), various constraints (e.g., battery lifetime), etc. This paper proposes a fast and low-overhead scheme to determine and adapt the operation mode of edge devices at runtime and orchestrate devices in a way that the efficiency of IoT devices is optimized with respect to the gateway's resource constraints. The proposed scheme breaks the optimization problem into several smaller ones (i.e., subproblems) whose solutions are aggregated to find the final solution. We present a novel memoization technique that determines the solution to a range of subproblems based on subproblems that are already solved. In addition, we present a novel pruning technique that reduces the search space and consequently reduces both memory and execution time overhead. The experimental results show up to 50% reduction in memory overhead and 14× reduction in execution time overhead compared to the state-of-the-art solution which is a major step toward efficient EC for IoT. Farzad Samie, Vasileios Tsoutsouras, Dimosthenis Masouros, Lars Bauer, Dimitrios Soudris, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | DMRM: Distributed Market-Based Resource Management of Edge Computing SystemsabstractResource management is a key technique for efficiently operating devices in Internet of Things (IoT). In this paper, we propose DMRM, a new algorithm based on economic and pricing models for dynamic resource management of IoT networks under CPU, memory, bandwidth and latency constraints. We use a supply and demand model, smart data pricing and perceived valued pricing, implementing a marketplace where IoT devices and Gateways buy and sell computing and communication resources necessary for task execution. Our new market-based algorithm is compared to relevant approaches showing that it not only reaches near-optimal results, but also, its scalable, distributed nature leads to three orders of magnitude lower execution requirements compared to centralized approaches. Manolis Katsaragakis, Dimosthenis Masouros, Vasileios Tsoutsouras, Farzad Samie, Lars Bauer, Jörg Henkel, Dimitrios Soudris |
DATE | 5 |
| 2019 | WCET Guarantees for Opportunistic Runtime ReconfigurationabstractTime-critical systems need to be analyzable for timing guarantees. There is an increasing demand for predictable performance that modern processor architectures fail to provide since they focus on average-case performance only. Recent work has demonstrated that runtime reconfiguration of hardware accelerators via an FPGA is a viable way to achieve high performance for optimized worst-case execution time (WCET) guarantees. Since execution of the worst-case path is highly improbable, configuring accelerators for this path costs reconfigurable area that could better be used to accelerate more probable paths. This work presents the first approach that comprises (1) an online average-case execution time (ACET) optimization while (2) maintaining the optimized WCET guarantee utilizing reconfigurable accelerators. We achieve this by a new design-time technique which determines the runtime slack bounds that allow speculative reconfiguration of accelerators that benefit the ACET. Combined with an online slack monitoring approach that introduces negligible overheads by using a performance counter, we show a runtime reduction of up to 10.4% for a complex and real-world application on top of an already-optimized WCET guarantee. Marvin Damschen, Lars Bauer, Jörg Henkel |
ICCAD | 2 |
| 2019 | From Cloud Down to Things: An Overview of Machine Learning in Internet of ThingsabstractWith the numerous Internet of Things (IoT) devices, the cloud-centric data processing fails to meet the requirement of all IoT applications. The limited computation and communication capacity of the cloud necessitate the edge computing, i.e., starting the IoT data processing at the edge and transforming the connected devices to intelligent devices. Machine learning (ML) the key means for information inference, should extend to the cloud-to-things continuum too. This paper reviews the role of ML in IoT from the cloud down to embedded devices. Different usages of ML for application data processing and management tasks are studied. The state-of-the-art usages of ML in IoT are categorized according to their application domain, input data type, exploited ML techniques, and where they belong in the cloud-to-things continuum. The challenges and research trends toward efficient ML on the IoT edge are discussed. Moreover, the publications on the “ML in IoT” are retrieved and analyzed systematically using ML classification techniques. Then, the growing topics and application domains are identified. Farzad Samie, Lars Bauer, Jörg Henkel |
IEEE Internet Things J. | 2 |
| 2019 | Oops: Optimizing Operation-mode Selection for IoT Edge DevicesabstractThe massive increase of IoT devices and their collected data raises the question of how to analyze all that data. Edge computing provides a suitable compromise, but the question remains: How much processing should be done locally vs. offloaded to other devices? The diverse application requirements and limited resources at the edge extend the challenges. We propose Oops , an optimization framework to adapt the resource management at runtime distributedly. It orchestrates the IoT devices and adapts their operation mode with respect to their constraints and the gateway’s limited shared resources. Oops reduces runtime overhead significantly while increasing user utility compared to state-of-the-art. Farzad Samie, Vasileios Tsoutsouras, Lars Bauer, Sotirios Xydis, Dimitrios Soudris, Jörg Henkel |
ACM Trans. Internet Techn. | 3 |
| 2018 | Highly efficient and accurate seizure prediction on constrained IoT devicesabstractIn this paper we present an efficient and accurate algorithm for epileptic seizure prediction on low-power and portable IoT devices. State-of-the-art algorithms suffer from two issues: computation intensive features and large internal memory requirement, which make them inapplicable for constrained devices. We reduce the memory requirement of our algorithm by reducing the size of data segments (i.e. the window of input stream data on which the processing is performed), and the number of required EEG channels. To respect the limitations of the processing capability, we reduce the complexity of our exploited features by only considering the simple features, which also contributes to reducing the memory requirements. Then, we provide new relevant features to compensate the information loss due to the simplifications (i.e. less number of channels, simpler features, shorter segment, etc.). We measured the energy consumption (12.41 mJ) and execution time (565 ms) for processing each segment (i.e. 5.12 seconds of EEG data) on a low-power MSP432 device. Even though the state-of-art does not fit to IoT devices, we evaluate the classification performance and show that our algorithm achieves the highest AUC score (0.79) for the held-out data and outperforms the state-of-the-art. Farzad Samie, Sebastian Paul, Lars Bauer, Jörg Henkel |
DATE | 3 |
| 2018 | Distributed Trade-Based Edge Device Management in Multi-Gateway IoTabstractThe Internet-of-Things (IoT) envisions an infrastructure of ubiquitous networked smart devices offering advanced monitoring and control services. The current art in IoT architectures utilizes gateways to enable application-specific connectivity to IoT devices. In typical configurations, IoT gateways are shared among several IoT edge devices. Given the limited available bandwidth and processing capabilities of an IoT gateway, the service quality (SQ) of connected IoT edge devices must be adjusted over time not only to fulfill the needs of individual IoT device users but also to tolerate the SQ needs of the other IoT edge devices sharing the same gateway. However, having multiple gateways introduces an interdependent problem, the binding, i.e., which IoT device shall connect to which gateway. In this article, we jointly address the binding and allocation problems of IoT edge devices in a multigateway system under the constraints of available bandwidth, processing power, and battery lifetime. We propose a distributed trade-based mechanism in which after an initial setup, gateways negotiate and trade the IoT edge devices to increase the overall SQ. We evaluate the efficiency of the proposed approach with a case study and through extensive experimentation over different IoT system configurations regarding the number and type of the employed IoT edge devices. Experiments show that our solution improves the overall SQ by up to 56% compared to an unsupervised system. Our solution also achieves up to 24.6% improvement on overall SQ compared to the state-of-the-art SQ management scheme, while they both meet the battery lifetime constraints of the IoT devices. Farzad Samie, Vasileios Tsoutsouras, Lars Bauer, Sotirios Xydis, Dimitrios Soudris, Jörg Henkel |
ACM Trans. Cyber Phys. Syst. | 3 |
| 2018 | Preemption of the Partial Reconfiguration Process to Enable Real-Time Computing With FPGAsabstractTo improve computing performance in real-time applications, modern embedded platforms comprise hardware accelerators that speed up the task’s most compute-intensive parts. A recent trend in the design of real-time embedded systems is to integrate field-programmable gate arrays (FPGA) that are reconfigured with different accelerators at runtime, to cope with dynamic workloads that are subject to timing constraints. One of the major limitations when dealing with partial FPGA reconfiguration in real-time systems is that the reconfiguration port can only perform one reconfiguration at a time: if a high-priority task issues a reconfiguration request while the reconfiguration port is already occupied by a lower-priority task, the high-priority task has to wait until the current reconfiguration is completed (a phenomenon known as priority inversion ), unless the current reconfiguration is aborted (introducing unbounded delays in low-priority tasks, a phenomenon known as starvation ). This article shows how priority inversion and starvation can be solved by making the reconfiguration process preemptive —that is, allowing it to be interrupted at any time and resumed at a later time without restarting it from scratch. Such a feature is crucial for the design of runtime reconfigurable real-time systems but not yet available in today’s platforms. Furthermore, the trade-off of achieving a guaranteed bound on the reconfiguration delay for low-priority tasks and the maximum delay induced for high-priority tasks when preempting an ongoing reconfiguration has been identified and analyzed. Experimental results on the Xilinx Zynq-7000 platform show that the proposed implementation of preemptive reconfiguration introduces a low runtime overhead, thus effectively solving priority inversion and starvation. Enrico Rossi, Marvin Damschen, Lars Bauer, Giorgio C. Buttazzo, Jörg Henkel |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2017 | Ultra-low power and dependability for IoT devices (Invited paper for IoT technologies)abstractRecent advances in technologies have allowed the design of small-size low-power and low-cost devices that can be connected to the Internet, enabling the emerging paradigm of Internet-of-things (IoT). IoT covers an ever-increasing range of applications, e.g., health-care monitoring, smart homes and buildings, etc. In this invited paper, we discuss and summarize the IoT paradigm with a special focus on energy consumption and methodologies for its minimization. Furthermore, we also discuss about reliability in the context of IoT devices. In all, this paper attempts to be a starting point for readers interested in developing energy-efficient IoT devices. Jörg Henkel, Santiago Pagani, Hussam Amrouch, Lars Bauer, Farzad Samie |
DATE | 4 |
| 2017 | Aging Resilience and Fault Tolerance in Runtime Reconfigurable ArchitecturesabstractRuntime reconfigurable architectures based on Field-Programmable Gate Arrays (FPGAs) allow areaand power-efficient acceleration of complex applications. However, being manufactured in latest semiconductor process technologies, FPGAs are increasingly prone to aging effects, which reduce the reliability and lifetime of such systems. Aging mitigation and fault tolerance techniques for the reconfigurable fabric become essential to realize dependable reconfigurable architectures. This article presents an accelerator diversification method that creates multiple configurations for runtime reconfigurable accelerators that are diversified in their usage of Configurable Logic Blocks (CLBs). In particular, it creates a minimal number of configurations such that all single-CLB and some multi-CLB faults can be tolerated. For each fault we ensure that there is at least one configuration that does not use that CLB. Second, a novel runtime accelerator placement algorithm is presented that exploits the diversity in resource usage of these configurations to balance the stress imposed by executions of the accelerators on the reconfigurable fabric. By tracking the stress due to accelerator usage at runtime, the stress is balanced both within a reconfigurable region as well as over all reconfigurable regions of the system. The accelerator placement algorithm also considers faulty CLBs in the regions and selects the appropriate configuration such that the system maintains a high performance in presence of multiple permanent faults. Experimental results demonstrate that our methods deliver up to 3.7× higher performance in presence of faults at marginal runtime costs and 1.6× higher MTTF than state-ofthe-art aging mitigation methods. Hongyan Zhang 0004, Lars Bauer, Michael A. Kochte, Eric Schneider, Hans-Joachim Wunderlich, Jörg Henkel |
IEEE Trans. Computers | 2 |
| 2017 | Timing Analysis of Tasks on Runtime Reconfigurable ProcessorsabstractReal-time embedded systems need to be analyzable for timing guarantees. Despite significant scientific advances, however, timing analysis lags years behind current microarchitectures with out-of-order scheduling pipelines, several hardware threads, and multiple (shared) cache layers. To satisfy the increasing performance demands, analyzable performance features are required. We propose a novel timing analysis approach to introduce runtime reconfigurable instruction set processors as one way to escape the scarcity of analyzable performance while preserving the flexibility of the system. We introduce extensions to the state-of-the-art Integer linear programming (ILP)-based program path analysis for computing precise worst case time bounds in the presence of the widely used technique to continue processor execution during reconfiguration by emulating not yet reconfigured custom instructions (CIs) in software. We identify and safely bound a timing anomaly of runtime reconfiguration, where executing faster than worst case time during reconfiguration extends the execution time of the whole program. Stalling the processor during reconfiguration (easier to analyze but not state-of-the-art for reconfigurable processors) is not required in our approach. Finally, we show the precision of our analysis on a complex multimedia application with multiple reconfigurable CIs for several hardware parameters and give advice on how to deal with reconfiguration delay under timing guarantees. Marvin Damschen, Lars Bauer, Jörg Henkel |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Efficient Partial Online Synthesis of Special Instructions for Reconfigurable ProcessorsabstractReconfigurable processors with fine-grained runtime-reconfigurable fabrics are used to speed up applications from different domains. Such a reconfigurable fabric allows loading of application-specific accelerators, where multiple accelerators can be combined using a coarse-grained runtimereconfigurable μProgram to speed up complex computationally intensive kernels. To allow a large degree of adaptivity in the reconfigurable fabric, as it is required by, e.g., multitasking systems, the μProgram for a kernel should not be generated at compile time, as it would constrain the adaptivity of the system. To enable flexible and efficient use of the reconfigurable fabric, we propose the necessary algorithms for runtime: 1) accelerator placement (i.e., deciding where on the fabric an accelerator should be reconfigured at runtime); 2) μProgram generation; and 3) μProgram caching. Accelerator synthesis and implementation are done at compile time to reduce runtime overhead in generating accelerators. We evaluate the proposed algorithms using different application scenarios and demonstrate the proposed concepts on an field-programmable gate array-based prototype of a reconfigurable processor. In comparison with state-of-the-art reconfigurable processors that generate μPrograms at compile time, we obtain an average speedup of 1.29× (up to 1.84×). Artjom Grudnitsky, Lars Bauer, Jörg Henkel |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Resource budgeting for reliability in reconfigurable architecturesabstractSRAM-based reconfigurable architectures are susceptible to soft-errors. Accelerators in the reconfigurable fabric need to be protected by fault tolerance techniques such as modular redundancy and scrubbing. However, blindly applying these techniques to all accelerators leads to suboptimal performance due to overprotection. Hongyan Zhang 0004, Lars Bauer, Jörg Henkel |
DAC | 2 |
| 2016 | Extending the WCET Problem to Optimize for Runtime-Reconfigurable ProcessorsabstractThe correctness of a real-time system does not depend on the correctness of its calculations alone but also on the non-functional requirement of adhering to deadlines. Guaranteeing these deadlines by static timing analysis, however, is practically infeasible for current microarchitectures with out-of-order scheduling pipelines, several hardware threads, and multiple (shared) cache layers. Novel timing-analyzable features are required to sustain the strongly increasing demand for processing power in real-time systems. Recent advances in timing analysis have shown that runtime-reconfigurable instruction set processors are one way to escape the scarcity of analyzable processing power while preserving the flexibility of the system. When moving calculations from software to hardware by means of reconfigurable custom instructions (CIs)—additional to a considerable speedup—the overestimation of a task’s worst-case execution time (WCET) can be reduced. CIs typically implement functionality that corresponds to several hundred instructions on the central processing unit (CPU) pipeline. While analyzing instructions for worst-case latency may introduce pessimism, the latency of CIs—executed on the reconfigurable fabric—is precisely known. In this work, we introduce the problem of selecting reconfigurable CIs to optimize the WCET of an application. We model this problem as an extension to state-of-the-art integer linear programming (ILP)-based program path analysis. This way, we enable optimization based on accurate WCET estimates with integration of information about global program flow, for example, infeasible paths. We present an optimal solution with effective techniques to prune the search space and a greedy heuristic that performs a maximum number of steps linear in the number of partitions of reconfigurable area available. Finally, we show the effectiveness of optimizing the WCET on a reconfigurable processor by evaluating a complex multimedia application with multiple reconfigurable CIs for several hardware parameters. Marvin Damschen, Lars Bauer, Jörg Henkel |
ACM Trans. Archit. Code Optim. | 2 |
| 2016 | Architecting On-Chip DRAM Cache for Simultaneous Miss Rate and Latency ReductionabstractOn-chip dynamic random access memory (DRAM) cache has been recently employed in the memory hierarchy to mitigate the widening latency gap between high-speed cores and off-chip memory. Two important parameters are the DRAM cache miss rate (D$-MR) and the DRAM cache hit latency (D$-HL), as they strongly influence the performance. These parameters depend upon the DRAM set mapping policy. Recently proposed DRAM set mapping policies are predominantly optimized for either D$-MR or D$-HL. We propose novel DRAM set mapping policies that simultaneously reduce D$-MR (via high associativity) and D$-HL (via improved row buffer hit rates). To further improve the D$-HL, we propose a small and low latency DRAM Tag cache (DTC) structure that can quickly determine whether an access to the DRAM cache will be a hit or a miss. The performance of the proposed DTC depends upon the DTC hit rate. To increase it, we present a novel DTC insertion policy that also increases the DTC hit rate. We investigate the latency and miss rate tradeoffs when designing a DRAM cache hierarchy and analyze the effects of different policies on the overall performance. We evaluate our policies on a wide variety of workloads and compare its performance with three recent proposals for on-chip DRAM caches. For a 16-core system, our set mapping policy along with our DTC and its adaptive DTC insertion policy improve the harmonic mean instruction per cycle throughput by 25.4%, 15.5%, and 7.3% compared to state-of-the-art, while requiring 55% less storage overhead for DRAM cache hit/miss prediction. Fazal Hameed, Lars Bauer, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2015 | Adaptive on-the-fly application performance modeling for many cores
Sebastian Kobbe, Lars Bauer, Jörg Henkel |
DATE | 2 |
| 2015 | Online binding of applications to multiple clock domains in shared FPGA-based systems
Farzad Samie, Lars Bauer, Chih-Ming Hsieh, Jörg Henkel |
DATE | 2 |
| 2015 | STRAP: Stress-Aware Placement for Aging Mitigation in Runtime Reconfigurable ArchitecturesabstractAging effects in nano-scale CMOS circuits impair the reliability and Mean Time to Failure (MTTF) of embedded systems. Especially for FPGAs that are manufactured in the latest technology node, aging is amajor concern. We introduce the first cross-layer aging-aware placement method for accelerators in FPGA-based runtime reconfigurable architectures. It optimizes stress distribution by accelerator placement at runtime, i.e. to which reconfigurable region an accelerator shall be reconfigured. Additionally, it optimizes logic placement at synthesis time to diversify the resource usage of individual accelerators, i.e. which CLBs of a reconfigurable region shall be used by an accelerator. Both layers together balance the intra- and inter-region stress induced by the application workload at negligible performance cost. Experimental results show significant reduction of maximum stress of up to 64% and 35%, which leads to up to 177% and 14% MTTF improvement relative to state-of-the-art methods w.r.t. HCI and BTI aging, respectively. Hongyan Zhang 0004, Michael A. Kochte, Eric Schneider, Lars Bauer, Hans-Joachim Wunderlich, Jörg Henkel |
ICCAD | 4 |
| 2015 | Resource-awareness on heterogeneous MPSoCs for image processing
Johny Paul, Walter Stechele, Benjamin Oechslein, Christoph Erhardt, Jens Schedel, Daniel Lohmann, Wolfgang Schröder-Preikschat, Manfred Kröhnert, Tamim Asfour, Éricles Sousa, Vahid Lari, Frank Hannig, Jürgen Teich, Artjom Grudnitsky, Lars Bauer, Jörg Henkel |
J. Syst. Archit. | 15 |
| 2015 | Multicast FullHD H.264 Intra Video Encoder ArchitectureabstractHigh throughput demands have resulted in enormous increase in complexity of multicast video applications, which require multiple video encoders to simultaneously compress individual views. In this paper, we present an approach to encode independent videos using H.264 intra encoder on a single hardware platform, where the hardware resources are shared by independent encoders in a time-multiplexed manner. In addition to lowering the latency introduced by multicasting, we address the strong sequential data dependencies within the encoder. At 25 frames/s, 150 MHz prototype of the proposed encoder and multiple video capture/display on a mid-range field programmable gate array is also presented. Muhammad Usman Karim Khan, Muhammad Shafique 0001, Lars Bauer, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2014 | COREFAB: Concurrent reconfigurable fabric utilization in heterogeneous multi-core systemsabstractApplication-specific accelerators may provide considerable speedup in single-core systems with a runtime-reconfigurable fabric (for simplicity called "fabric" in the following). A reconfigurable core, i.e. processor core pipeline coupled to a fabric, can be integrated along with regular general purpose processor cores (GPPs) into a reconfigurable multi-core system with widely improved system performance. As most applications only use a fraction of the available fabric at a time, making the fabric usable by the GPPs (in addition to the reconfigurable core) in such a multi-core system is desirable. Existing work focused on algorithms that decide the amount of fabric that is assigned to each core in a multi-core system. However, when multiple cores access the fabric simultaneously, they are either limited to serialized fabric access or, when parallel access is supported, the size of the fabric share assigned to a core is inflexible and tends to be over- or undersized for the running application, thereby not efficiently utilizing the fabric. We propose a novel approach that allows GPPs to access the fabric of the reconfigurable core and that enables concurrent fabric utilization on-the-fly through merging fabric accesses from different cores at run-time. Compared to state-of-the art, our approach improves performance of the GPPs in a reconfigurable multi-core system by 1.3x on average, without reducing the performance of the reconfigurable core. Artjom Grudnitsky, Lars Bauer, Jörg Henkel |
CASES | 2 |
| 2014 | Automatic custom instruction identification in memory streaming algorithmsabstractApplication-specific instruction set processors (ASIPs) extend the instruction set of a general purpose processor by dedicated custom instructions (CIs). In the last decade, reconfigurable processors advanced this concept towards run-time reconfiguration to increase the efficiency and adaptivity. Compiler support for automatic identification and implementation of ASIP CIs exists commercially and on research platforms, but these compilers do not support CIs with memory accesses, as ASIP CIs typically work on register file data. While being acceptable for ASIPs, this imposes a limitation for reconfigurable processors as they achieve their performance by exploiting data-level parallelism. Consequently, we propose a novel approach to CI identification for runtime reconfigurable processors with support for memory operations in contrast to previous works that explicitly exclude them. Our algorithm extracts memory access patterns which allows us to abstract from single memory operations and merge accesses to optimally utilize the available memory bandwidth. We implemented our algorithm in a state-of-the-art compiler framework. Martin Haaß, Lars Bauer, Jörg Henkel |
CASES | 2 |
| 2014 | Reducing Latency in an SRAM/DRAM Cache Hierarchy via a Novel Tag-Cache ArchitectureabstractMemory speed has become a major performance bottleneck as more and more cores are integrated on a multi-core chip. The widening latency gap between high speed cores and memory has led to the evolution of multi-level SRAM/DRAM cache hierarchies that exploit the latency benefits of smaller caches (e.g. private L1 and L2 SRAM caches) and the capacity benefits of larger caches (e.g. shared L3 SRAM and shared L4 DRAM cache). The main problem of employing large L3/L4 caches is their high tag lookup latency. To solve this problem, we introduce the novel concept of small and low latency SRAM/DRAM Tag-Cache structures that can quickly determine whether an access to the large L3/L4 caches will be a hit or a miss. The performance of the proposed Tag-Cache architecture depends upon the Tag-Cache hit rate and to improve it we propose a novel Tag-Cache insertion policy and a DRAM row buffer mapping policy that reduce the latency of memory requests. For a 16-core system, this improves the average harmonic mean instruction per cycle throughput of latency sensitive applications by 13.3% compared to state-of-the-art. Fazal Hameed, Lars Bauer, Jörg Henkel |
DAC | 2 |
| 2014 | Multi-Layer Dependability: From Microarchitecture to Application LevelabstractWe show in this paper that multi-layer dependability is an indispensable way to cope with the increasing amount of technology-induced dependability problems that threaten to proceed further scaling. We introduce the definition of multi-layer dependability and present our design flow within this paradigm that seamlessly integrates techniques starting at circuit layer all the way up to application layer and thereby accounting for ASIC-based architectures as well as for reconfigurable-based architectures. At the end, we give evidence that the paradigm of multi-layer dependability bears a large potential for significantly increasing dependability at reasonable effort. Jörg Henkel, Lars Bauer, Hongyan Zhang 0004, Semeen Rehman, Muhammad Shafique 0001 |
DAC | 2 |
| 2014 | GUARD: GUAranteed Reliability in Dynamically Reconfigurable SystemsabstractSoft errors are a reliability threat for reconfigurable systems implemented with SRAM-based FPGAs. They can be handled through fault tolerance techniques like scrubbing and modular redundancy. However, selecting these techniques statically at design or compile time tends to be pessimistic and prohibits optimal adaptation to changing soft error rate at runtime. Hongyan Zhang 0004, Michael A. Kochte, Michael E. Imhof, Lars Bauer, Hans-Joachim Wunderlich, Jörg Henkel |
DAC | 4 |
| 2014 | MORP: makespan optimization for processors with an embedded reconfigurable fabricabstractProcessors with an embedded runtime reconfigurable fabric have been explored in academia and industry started production of commercial platforms (e.g. Xilinx Zynq-7000). While providing significant performance and efficiency, the comparatively long reconfiguration time limits these advantages when applications request reconfigurations frequently. In multi-tasking systems frequent task switches lead to frequent reconfigurations and thus are a major hurdle for further performance increases. Sophisticated task scheduling is a very effective means to reduce the negative impact of these reconfiguration requests. In this paper, we propose an online approach for combined task scheduling and re-distribution of reconfigurable fabric between tasks in order to reduce the makespan, i.e. the completion time of a taskset that executes on a runtime reconfigurable processor. Evaluating multiple tasksets comprised of multimedia applications, our proposed approach achieves makespans that are on average only 2.8% worse than those achieved by a theoretical optimal scheduling that assumes zero-overhead reconfiguration time. In comparison, scheduling approaches deployed in state-of-the-art reconfigurable processors achieve makespans 14%-20% worse than optimal. As our approach is a purely software-side mechanism, a multitude of reconfigurable platforms aimed at multi-tasking can benefit from it. Artjom Grudnitsky, Lars Bauer, Jörg Henkel |
FPGA | 2 |
| 2014 | Adaptive Energy Management for Dynamically Reconfigurable ProcessorsabstractWe present an adaptive energy management system for dynamically reconfigurable processors that chooses an energy-minimizing set of custom instructions (CIs) and then power-gates the temporarily unused subset of CIs. It requires a comprehensive power model to estimate the power consumption of different CIs at run time. We deploy our new energy management in two state-of-the-art reconfigurable processors (RISPP and Molen) and perform an elaborative evaluation of energy savings under various area and performance constraints for different technology nodes. We demonstrate the energy benefits by comparing it to state-of-the-art power-gating techniques for FPGAs. The work is implemented as a prototype on a Xilinx FPGA platform. Muhammad Shafique 0001, Lars Bauer, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | Simultaneously optimizing DRAM cache hit latency and miss rate via novel set mapping policiesabstractTwo key parameters that determine the performance of a DRAM cache based multi-core system are DRAM cache hit latency (HL) and DRAM cache miss rate (MR), as they strongly influence the average DRAM cache access latency. Recently proposed DRAM set mapping policies are either optimized for HL or for MR. None of these policies provides a good HL and MR at the same time. This paper presents a novel DRAM set mapping policy that simultaneously targets both parameters with the goal of achieving the best of both to reduce the overall DRAM cache access latency. For a 16-core system, our proposed set mapping policy reduces the average DRAM cache access latency (depends upon HL and MR) compared to state-of-the-art DRAM set mapping policies that are optimized for either HL or MR by 29.3% and 12.1%, respectively. Fazal Hameed, Lars Bauer, Jörg Henkel |
CASES | 2 |
| 2013 | Hardware acceleration for programs in SSA formabstractRegister allocation is one of the most time-consuming parts of the compilation process. Depending on the quality of the register allocation, a large amount of shuffle code to move values between registers is generated. In this paper, we propose a processor architecture extension to provide register file permutations by which the shuffle code can be implemented more efficiently. We present compiler support to utilize this extension, an evaluation regarding performance and compilation time using the SPEC CINT2000 benchmark, as well as an analysis of area and frequency overhead of our architecture implementation. We find that using our extension, the number of executed instructions is reduced by up to 5.1 % while the compilation time is unaffected. Manuel Mohr, Artjom Grudnitsky, Tobias Modschiedler, Lars Bauer, Sebastian Hack, Jörg Henkel |
CASES | 4 |
| 2013 | Reliable on-chip systems in the nano-era: lessons learnt and future trendsabstractReliability concerns due to technology scaling have been a major focus of researchers and designers for several technology nodes. Therefore, many new techniques for enhancing and optimizing reliability have emerged particularly within the last five to ten years. This perspective paper introduces the most prominent reliability concerns from today's points of view and roughly recapitulates the progress in the community so far. The focus of this paper is on perspective trends from the industrial as well as academic points of view that suggest a way for coping with reliability challenges in upcoming technology nodes. Jörg Henkel, Lars Bauer, Nikil Dutt, Puneet Gupta 0001, Sani R. Nassif, Muhammad Shafique 0001, Mehdi Baradaran Tahoori, Norbert Wehn |
DAC | 2 |
| 2013 | Adaptive cache management for a combined SRAM and DRAM cache hierarchy for multi-coresabstractOn-chip DRAM caches may alleviate the memory bandwidth problem in future multi-core architectures through reducing off-chip accesses via increased cache capacity. For memory intensive applications, recent research has demonstrated the benefits of introducing high capacity on-chip L4-DRAM as Last-Level-Cache between L3-SRAM and off-chip memory. These multi-core cache hierarchies attempt to exploit the latency benefits of L3-SRAM and capacity benefits of L4-DRAM caches. However, not taking into consideration the cache access patterns of complex applications can cause inter-core DRAM interference and inter-core cache contention. In this paper, we contest to re-architect existing cache hierarchies by proposing a hybrid cache architecture, where the Last-Level-Cache is a combination of SRAM and DRAM caches. We propose an adaptive DRAM placement policy in response to the diverse requirements of complex applications with different cache access behaviors. It reduces inter-core DRAM interference and inter-core cache contention in SRAM/DRAM-based hybrid cache architectures: increasing the harmonic mean instruction-per-cycle throughput by 23.3% (max. 56%) and 13.3% (max. 35.1%) compared to state-of-the-art. Fazal Hameed, Lars Bauer, Jörg Henkel |
DATE | 2 |
| 2013 | An H.264 Quad-FullHD low-latency intra video encoderabstractVideo applications are moving from Full-HD capability (1920×1080) to even higher resolutions such as Quad-FullHD (3840×2160). The H.264 Intra-mode can be used by embedded devices to trade off the better encoding efficiency of H.264 temporal prediction (Inter-mode) against savings in area and power as well as saving the massive computational overhead of the sub-pixel motion estimation by using only spatial prediction (Intra-mode). Still, the H.264 Intra-mode requires a large computational effort and imposes severe challenges when targeting Quad-FullHD 25 fps real-time video encoding at moderate operating frequencies (we target 150 MHz) and limited area budget. Therefore, in this work we address the strong sequential data dependencies within H.264 Intra-mode that restrict the parallelism and inhibit high resolution encoding by a) decoupling of DC and AC transform paths, b) cycle-budget aware mode prediction scheduling while c) being area efficient. Using our proposed techniques, Quad-FullHD (3840×2160) 28 fps video encoding is achieved at 150 MHz, making our architecture applicable for high definition recording. Muhammad Usman Karim Khan, Jan Micha Borrmann, Lars Bauer, Muhammad Shafique 0001, Jörg Henkel |
DATE | 3 |
| 2013 | Module diversification: Fault tolerance and aging mitigation for runtime reconfigurable architecturesabstractRuntime reconfigurable architectures based on Field-Programmable Gate Arrays (FPGAs) are attractive for realizing complex applications. However, being manufactured in latest semiconductor process technologies, FPGAs are increasingly prone to aging effects, which reduce the reliability of such systems and must be tackled by aging mitigation and application of fault tolerance techniques. This paper presents module diversification, a novel design method that creates different configurations for runtime reconfigurable modules. Our method provides fault tolerance by creating the minimal number of configurations such that for any faulty Configurable Logic Block (CLB) there is at least one configuration that does not use that CLB. Additionally, we determine the fraction of time that each configuration should be used to balance the stress and to mitigate the aging process in FPGA-based runtime reconfigurable systems. The generated configurations significantly improve reliability by fault-tolerance and aging mitigation. Hongyan Zhang 0004, Lars Bauer, Michael A. Kochte, Eric Schneider, Claus Braun, Michael E. Imhof, Hans-Joachim Wunderlich, Jörg Henkel |
ITC | 2 |
| 2013 | Test Strategies for Reliable Runtime Reconfigurable ArchitecturesabstractField-programmable gate array (FPGA)-based reconfigurable systems allow the online adaptation to dynamically changing runtime requirements. The reliability of FPGAs, being manufactured in latest technologies, is threatened by soft errors, as well as aging effects and latent defects. To ensure reliable reconfiguration, it is mandatory to guarantee the correct operation of the reconfigurable fabric. This can be achieved by periodic or on-demand online testing. This paper presents a reliable system architecture for runtime-reconfigurable systems, which integrates two nonconcurrent online test strategies: preconfiguration online tests (PRET) and postconfiguration online tests (PORT). The PRET checks that the reconfigurable hardware is free of faults by periodic or on-demand tests. The PORT has two objectives: It tests reconfigured hardware units after reconfiguration to check that the configuration process completed correctly and it validates the expected functionality. During operation, PORT is used to periodically check the reconfigured hardware units for malfunctions in the programmable logic. Altogether, this paper presents PRET, PORT, and the system integration of such test schemes into a runtime-reconfigurable system, including the resource management and test scheduling. Experimental results show that the integration of online testing in reconfigurable systems incurs only minimum impact on performance while delivering high fault coverage and low test latency. Lars Bauer, Claus Braun, Michael E. Imhof, Michael A. Kochte, Eric Schneider, Hongyan Zhang 0004, Jörg Henkel, Hans-Joachim Wunderlich |
IEEE Trans. Computers | 1 |
| 2012 | Invasive manycore architecturesabstractThis paper introduces a scalable hardware and software platform applicable for demonstrating the benefits of the invasive computing paradigm. The hardware architecture consists of a heterogeneous, tile-based manycore structure while the software architecture comprises a multi-agent management layer underpinned by distributed runtime and OS services. The necessity for invasive-specific hardware assist functions is analytically shown and their integration into the overall manycore environment is described. Jörg Henkel, Andreas Herkersdorf, Lars Bauer, Thomas Wild, Michael Hübner 0001, Ravi Kumar Pujari, Artjom Grudnitsky, Jan Heisswolf, Aurang Zaib, Benjamin Vogel, Vahid Lari, Sebastian Kobbe |
ASP-DAC | 3 |
| 2012 | Partial online-synthesis for mixed-grained reconfigurable architecturesabstractProcessor architectures with Fine-Grained Reconfigurable Accelerators (FGRAs) allow for a high degree of adaptivity to address varying application requirements. When processing computation intensive kernels, multiple FGRAs may be used to execute a complex function. In order to exploit the adaptivity of a fine-grained reconfigurable fabric, a runtime system should decide when and which FGRAs to reconfigure with respect to application requirements. To enable this adaptivity, a flexible infrastructure is required that allows combining FGRAs to execute complex functions. We propose a mixed-grained reconfigurable architecture composed from a Coarse-Grained Reconfigurable Infrastructure (CGRI) that connects the FGRAs. At runtime we synthesize CGRI configurations that depend on decisions of the runtime system, e.g. which FGRAs shall be reconfigured. Synthesis and place & route of the FGRAs are done at compile time for performance reasons. Combined, this results in a partial online synthesis for mixed grained reconfigurable architectures, which allows maintaining a low runtime overhead while exploiting the inherent adaptivity of the reconfigurable fabric. In this work we focus on the crucial parts of synthesizing the configurations for the CGRI at runtime, propose algorithms, and compare their performance/overhead trade-offs for different application scenarios. We are the first to exploit the increased adaptivity of FGRAs that are connected by a CGRI, by using our partial online synthesis. In comparison to a state-of-the-art reconfigurable architecture that synthesizes the configurations for the CGRI at compile time we obtain an average speedup of 1.79x. Artjom Grudnitsky, Lars Bauer, Jörg Henkel |
DATE | 2 |
| 2012 | Dynamic cache management in multi-core architectures through run-time adaptationabstractNon-Uniform Cache Access (NUCA) architectures provide a potential solution to reduce the average latency for the last-level-cache (LLC), where the cache is organized into per-core local and remote partitions. Recent research has demonstrated the benefits of cooperative cache sharing among local and remote partitions. However, ignoring cache access patterns of concurrently executing applications sharing the local and remote partitions can cause inter-partition contention that reduces the overall instruction throughput. We propose a dynamic cache management scheme for LLC in NUCA-based architectures, which reduces inter-partition contention. Our proposed scheme provides efficient cache sharing by adapting migration, insertion, and promotion policies in response to the dynamic requirements of the individual applications with different cache access behaviors. Our adaptive cache management scheme allows individual cores to steal cache capacity from remote partitions to achieve better resource utilization. On average, our proposed scheme increases the performance (instructions per cycle) by 28% (minimum 8.4%, maximum 75%) compared to a private LLC organization. Fazal Hameed, Lars Bauer, Jörg Henkel |
DATE | 2 |
| 2012 | PATS: A Performance Aware Task Scheduler for Runtime Reconfigurable ProcessorsabstractMulti-tasking is one of the main requirements for complex embedded systems to fulfill user expectations (e.g. flexibility of the system), increase the resource utilization, and thus increase the system efficiency. In general, the flexibility and efficiency can be increased by incorporating a fine-grained reconfigurable fabric (e.g. an embedded FPGA) that is coupled with a general-purpose processor and accelerates the computationally intensive kernels. This work focuses on reconfigurable processors that use a reconfigurable fabric to implement Special Instructions (SIs) that are invoked by the processor and process data-dominant parts. For each SI the decision whether it is executed in hardware or emulated in software can be changed dynamically at runtime. In this paper, we present our novel Performance Aware Task Scheduler (PATS) that decides the task schedule at runtime while considering the specific system state of the reconfigurable processor. For instance, if a task t has to emulate several SI executions in software because reconfiguring the corresponding hardware implementations is not completed yet, then it might be more efficient to schedule other tasks first, depending on the soft-deadlines of the tasks, until the reconfigurations of that task t are completed. In comparison to other task schedulers (earliest deadline first, rate monotonic scheduling, and round robin), PATS achieves on average a 1.45x better system tardiness (i.e., the sum of cycles by which tasks miss their deadlines). Additionally, PATS reduces the make span (i.e. the time when all tasks have completed all of their jobs) on average by 1.17x (up to 1.58x). Especially in challenging multi-tasking scenarios with tight deadlines or a small reconfigurable fabric PATS performs significantly better than other task schedulers do. Lars Bauer, Artjom Grudnitsky, Muhammad Shafique 0001, Jörg Henkel |
FCCM | 1 |
| 2012 | Transparent structural online test for reconfigurable systemsabstractFPGA-based reconfigurable systems allow the online adaptation to dynamically changing runtime requirements. However, the reliability of modern FPGAs is threatened by latent defects and aging effects. Hence, it is mandatory to ensure the reliable operation of the FPGA's reconfigurable fabric. This can be achieved by periodic or on-demand online testing. In this paper, a system-integrated, transparent structural online test method for runtime reconfigurable systems is proposed. The required tests are scheduled like functional workloads, and thorough optimizations of the test overhead reduce the performance impact. The proposed scheme has been implemented on a reconfigurable system. The results demonstrate that thorough testing of the reconfigurable fabric can be achieved at negligible performance impact on the application. Mohamed Abdelfattah, Lars Bauer, Claus Braun, Michael E. Imhof, Michael A. Kochte, Hongyan Zhang 0004, Jörg Henkel, Hans-Joachim Wunderlich |
IOLTS | 2 |
| 2011 | mRTS: Run-time system for reconfigurable processors with multi-grained instruction-set extensionsabstractWe present a run-time system for a multi-grained reconfigurable processor in order to provide a dynamic trade-off between performance and available area budgets for both fine- as well as coarse-grained reconfigurable fabrics as part of one reconfigurable processor. Our run-time system is the first implementation of its kind that dynamically selects and steers a performance-maximizing multi-grained instruction set under run-time varying constraints. It achieves a performance improvement of more than 2× compared to state-of-the-art run-time systems for multi-grained architectures. To elaborate the benefits of our approach further, we also compare it with offline- and online-optimal instruction-set selection schemes. Waheed Ahmed, Muhammad Shafique 0001, Lars Bauer, Jörg Henkel |
DATE | 3 |
| 2011 | Minority-Game-based resource allocation for run-time reconfigurable multi-core processorsabstractA novel policy for allocating reconfigurable fabric resources in multi-core processors is presented. We deploy a Minority-Game to maximize the efficient use of the reconfigurable fabric while meeting performance constraints of individual tasks running on the cores. As we will show, the Minority Game ensures a fair allocation of resources, e.g., no single core will monopolize the reconfigurable fabric. Rather, all cores receive a “fair” share of the fabric, i.e., their tasks would miss their performance constraints by approximately the same margin, thus ensuring an overall graceful degradation. The policy is implemented on a Virtex-4 FPGA and evaluated for diverse applications ranging from security to multimedia domains. Our results show that the Minority-Game policy achieves on average 2× higher application performance and a 5× improved efficiency of resource utilization compared to state-of-the-art. Muhammad Shafique 0001, Lars Bauer, Waheed Ahmed, Jörg Henkel |
DATE | 2 |
| 2011 | Run-Time Resource Allocation for Simultaneous Multi-tasking in Multi-core Reconfigurable ProcessorsabstractState-of-the-art multi-core reconfigurable processors do not exploit the full potential of simultaneous multi-tasking with run-time adaptive reconfigurable fabric allocation. We propose a novel run-time system for simultaneous multi-tasking in a multi-core reconfigurable processor that adaptively allocates the mixed-grained reconfigurable fabric resource at run time among different tasks considering their performance constraints. Our scheme employs the novel concept of refined task-criticality (based on the functional-block-level performance constraints) considering the computational properties of dependent tasks and their inherent potential for acceleration. Our scheme dynamically compensates the deadline misses at the functional block level. It thereby reduces the potential task-level deadline misses under competing scenarios. With the help of a secure video conferencing application (with 4 dependent tasks of diverse computational properties), we demonstrate that our scheme reduces the deadline misses by (on average) 6× under given performance constraints, when compared to state-of-the-art reconfigurable processors. Waheed Ahmed, Muhammad Shafique 0001, Lars Bauer, Manuel Hammerich, Jörg Henkel, Jürgen Becker 0001 |
FCCM | 3 |
| 2010 | KAHRISMA: A novel Hypermorphic Reconfigurable-Instruction-Set Multi-grained-Array architectureabstractFacing the requirements of next generation applications, current approaches of embedded systems design will soon hit the limit where they may no longer perform efficiently. The unpredictable nature and diverse processing behavior of future applications requires to transgress the barrier of tailor-made, application-/domain-specific embedded system designs. As a consequence, next generation architectures for embedded systems have to react much more flexible to unforeseeable run-time scenarios. In this paper we present our innovative processor architecture concept KAHRISMA (KArlsruhe's Hypermorphic Reconfigurable-Instruction-Set Multi-grained-Array). It tightly integrates coarse- and fine-grained run-time reconfigurable fabrics that can incorporate to realize hardware acceleration for computationally complex algorithms. Furthermore, the fabrics can be combined to realize different Instruction Set Architectures that may execute in parallel. With the help of an encrypted H.264 en-/decoding case study we demonstrate that our novel KAHRISMA architecture will deliver the required flexibility to design future-proof embedded systems that are not limited to a certain computational domain. Ralf König 0001, Lars Bauer, Timo Stripf, Muhammad Shafique 0001, Waheed Ahmed, Jürgen Becker 0001, Jörg Henkel |
DATE | 2 |
| 2010 | enBudget: A Run-Time Adaptive Predictive Energy-Budgeting scheme for energy-aware Motion Estimation in H.264/MPEG-4 AVC video encoderabstractThe limited energy resources in portable multimedia devices require the reduction of encoding complexity. The complex Motion Estimation (ME) scheme of H.264/MPEG-4 AVC accounts for a major part of the encoder energy. In this paper we present a Run-Time Adaptive Predictive Energy Budgeting (enBudget) scheme for energy-aware ME that predicts the energy budget for different video frames and different Macroblocks (MBs) in an adaptive manner considering the run-time changing scenarios of available energy, video frame characteristics, and user-defined coding constraints while keeping a good video quality. It assigns different Energy-Quality Classes to different video frames and fine-tunes at MB level depending upon the predictive energy quota in order to cope with above-mentioned run-time unpredictable scenarios. Compared to UMHexagonS, EPZS, and FastME, our enBudget scheme for energy-aware ME achieves an energy saving of up to 93%, 90%, 88% (average 88%, 77%, 66%), respectively. It suffers from an average Peak Signal to Noise Ratio (PSNR) loss of 0.29 dB compared to Full Search. We also demonstrate that enBudget is equally beneficial to various other state-of-the-art fast adaptive MEs (e.g.). We have evaluated our scheme for ASIC and various FPGAs. Muhammad Shafique 0001, Lars Bauer, Jörg Henkel |
DATE | 2 |
| 2010 | Selective instruction set muting for energy-aware adaptive processorsabstractWe propose a new way to save energy in adaptive processors. According to an execution context the custom instruction set of an adaptive processor is selectively 'muted' at run time and thus the energy efficiency is significantly increased. Implemented are multiple so-called 'muting modes' each leading to particular leakage energy savings. A key challenge of this work is to determine which of the muting modes are beneficial for which part of the custom instruction set in a specific execution context. We demonstrate the feasibility by means of an H.264 video encoder (although not limited to that) for various technology nodes. The complex and unpredictable processing behavior of an H.264 encoder represents thereby a real-world scenario. Our results show on average more than 30% energy savings compared to state-of-the-art. We claim that adaptive processors (and reconfigurable computing in general) would be far more energy efficient if FPGA vendors would provide a basic infrastructure that is necessary to exert our novel technique. Muhammad Shafique 0001, Lars Bauer, Jörg Henkel |
ICCAD | 2 |
| 2009 | Cross-architectural design space exploration tool for reconfigurable processorsabstractProcessors that deploy fine-grained reconfigurable fabrics to implement application-specific accelerators on-demand obtained significant attention within the last decade. They trade-off the flexibility of general-purpose processors with the performance of application-specific circuits without tailoring the processor towards a specific application domain like Application Specific Instruction Set Processors (ASIPs). Vast amounts of reconfigurable processors have been proposed, differing in multifarious architectural decisions. However, it has always been an open question, which of the proposed concepts is more efficient in certain application and/or parameter scenarios. Various reconfigurable processors were investigated in certain scenarios, but never before a systematic design space exploration across diverse reconfigurable processor concepts has been conducted with the aim to aid a designer of a reconfigurable processor. We have developed a first-of-its-kind comprehensive design space exploration tool that allows to systematically explore diverse reconfigurable processors and architectural parameters. Our tool allows presenting the first cross-architectural design space exploration of multiple fine-grained reconfigurable processors on a fair comparable basis. After categorizing fine-grained reconfigurable processors and their relevant parameters, we present our tool and an in-depth analysis of reconfigurable processors within different relevant scenarios. Lars Bauer, Muhammad Shafique 0001, Jörg Henkel |
DATE | 1 |
| 2009 | A parallel approach for high performance hardware design of intra prediction in H.264/AVC Video CodecabstractThe H.264/AVC Intra Frame Codec (i.e. all frames are coded as I-frames) targets high-resolution/high-end encoding applications (e.g. digital cinema and high quality archiving etc.), providing much better compression efficiency at lower computational complexity compared to MJPEG2000. Moreover, in case of video coding of very high motion scenes, the number of Intra Macroblocks is dominant. Intra Prediction is a compute intensive and memory-critical part that consumes 80% of the computation time of the entire Intra Compression process when executing the H.264 encoder on MIPS processor. We therefore present a novel hardware for H.264 Intra Prediction that processes all the prediction modes in parallel inside one integrated module (i.e. mode-level parallelism) enabling us to exploit the full space of optimization. It exhibits a group-based write-back scheme to reduce the memory transfers in order to facilitate the fast mode-decision schemes. Our Luma 4times4 hardware is 3.6times, 5.2times, and 5.5times faster than state-of-the-art approaches, QS0, respectively. Our results show that processing Luma 16times16, Chroma 8times8, and Luma 4times4 with the proposed approach is 7.2times, 6.5times, and 1.8times faster (while giving an energy saving of 60%, 80%, and 74%) when compared with Dedicated Module Approach (each prediction mode is processed with its independent hardware module i.e. a typical ASIC style for Intra Prediction). We get an area saving of 58% for Luma 4times4 hardware. Muhammad Shafique 0001, Lars Bauer, Jörg Henkel |
DATE | 2 |
| 2009 | RISPP: A run-time adaptive reconfigurable embedded processorabstractProcessors that deploy reconfigurable fabrics to implement application-specific accelerators on-demand obtained significant attention within the last decade. They trade-off the flexibility of general-purpose processors with the performance of application-specific circuits without tailoring the processor towards a specific application domain like application specific instruction set processors (ASIPs). However, even though they reconfigure parts of the hardware at run time, the decisions which accelerators shall be reconfigured at which time are typically determined at compile time. Therefore, it is conceptually not possible to react to dynamically changing situations like varying dynamic control flow (e.g. due to changed input data), changing task priorities / performance constraints, and changing availability of reconfigurable hardware (may be reassigned to another task). This work presents the novel rotating instruction set processing platform (RISPP) with its run-time system that enables dynamic adaptations in the above highlighted scenarios efficiently. Therefore, the presented approach is suitable for long-term as well as frequent adaptation requirements. Lars Bauer, Muhammad Shafique 0001, Jörg Henkel |
FPL | 1 |
| 2009 | REMiS: Run-time energy minimization scheme in a reconfigurable processor with dynamic power-gated instruction setabstractReconfigurable processors provide a means to flexible and energy-aware computing. In this paper, we present a new scheme for runtime energy minimization (REMiS) as part of a dynamically reconfigurable processor that is exposed to run-time varying constraints like performance and footprint (i.e. amount of reconfigurable fabric). The scheme chooses an energy-minimizing set of so-called Special Instructions (considering leakage, dynamic, and reconfiguration energy) and then 'power-gates' a temporarily unused subset of the Special Instruction set. We provide a comprehensive evaluation for different technologies (ranging from 65 nm to 150 nm) and thereby show that our scheme is technology independent, i.e. it is beneficial for various technologies alike. By means of an H.264 video encoder we demonstrate that for certain performance constraints our scheme (applied to our in-house reconfigurable processor) achieves an allover energy saving of up to 40.8% (avg. 24.8%) compared to a performance-maximizing scheme. We also demonstrate that our scheme is equally beneficial to various other state-of-the-art reconfigurable processor architectures like Molen [9] where it achieves energy savings of up to 48.7% (avg. 28.93%) at 65 nm. We have employed an H.264 encoder within this paper as an application in order to demonstrate the strengths of our scheme, since the H.264's complexity and run-time unpredictability present a challenging scenario for state-of-the-art architectures. Muhammad Shafique 0001, Lars Bauer, Jörg Henkel |
ICCAD | 2 |
| 2008 | Run-time instruction set selection in a transmutable embedded processorabstractWe are presenting a new concept of an application-specific processor that is capable of transmuting its instruction set according to non-predictive application behavior during run-time. In those scenarios, current (extensible) embedded processors are less efficient since they are not run-time adaptive. We have identified the instruction set selection to be a critical step to perform at run time and hence we focus this paper on that crucial part. Our paradigm conducts as many steps as possible at compile/design time and as little as necessary at run time with the constraint to provide a sufficient flexibility to react to non-predictive application behavior efficiently. We provide an in-depth analysis of our scheme and achieve a speed-up of up to 7.19x (average: 3.63x) compared to state-of-the-art adaptive approaches (like [19]). As an application, we have employed a whole H.264 video encoder though our scheme is by principle applicable to many other embedded applications. Our results are evaluated by an implementation of the instruction set selection for our transmutable processor on an FPGA platform. Lars Bauer, Muhammad Shafique 0001, Jörg Henkel |
DAC | 1 |
| 2008 | Run-time System for an Extensible Embedded Processor with Dynamic Instruction SetabstractOne of the upcoming challenges in embedded processing is to incorporate an increasing amount of adaptivity in order to respond to the multifarious constraints induced by today's embedded systems that feature complex and diverse application behaviors. We present a novel concept (evaluated with a hardware prototype) that moves traditional design-time jobs to run time in order to increase efficiency (in this paper we focus on performance). Adaptivity is achieved dynamically through what we call special instructions (Sis) which may change during run time according to non-predictable application behavior. The new contribution of this paper is the principal component that actually makes the entire embedded processor work efficiently, namely the "special instruction scheduler". It determines during run time 'when' and 'how' Special Instructions are composed and executed. We achieve a 2.38times performance increase over a reconfigurable processor system with dynamic instruction set (Molen). Our whole platform consists of a toolchain including estimation and simulation tools plus a running hardware prototype. Throughout this paper, we discuss the functionality by means of an H.264 video encoder in detail even though the concept is not limited to this application. Lars Bauer, Muhammad Shafique 0001, Stephanie Kreutz, Jörg Henkel |
DATE | 1 |
| 2008 | A computation- and communication- infrastructure for modular special instructions in a dynamically reconfigurable processorabstractProcessors with a reconfigurable instruction set combine the performance of dedicated application accelerators with a flexibility that goes beyond that of traditional application specific instruction set processors (ASIPs). The latter are optimized for certain application domains and thus typically do not provide a high performance and/or efficiency when deployed in other domains. State-of-the-art reconfigurable processors on the other side still use the concept of monolithic Special Instructions (SIs, i.e. the application accelerators). In our work, we instead present modular SIs as a hierarchy of elementary data paths and different SI implementations that facilitate a high flexibility and performance. This is a novel concept that achieves a speedup of 26.6x compared to a general purpose processor and 1.24x compared to a state-of-the-art reconfigurable processor (that is statically optimized for the predetermined benchmark situation) when executing an H.264 video encoder. We introduce a novel infrastructure for computation and communication that actually enables the implementation of modular SIs and offers various parameters to match specific requirements. The infrastructure is implemented and tested on an FPGA-based prototype to demonstrate its feasibility. Lars Bauer, Muhammad Shafique 0001, Jörg Henkel |
FPL | 1 |
| 2008 | 3-tier dynamically adaptive power-aware motion estimator for h.264/AVC video encodingabstractThe limitation of energy in portable communication/entertainment devices necessitates the reduction of video encoding complexity. The H.264/AVC video coding standard is one of the latest video codecs and features a complex Motion Estimation scheme that accounts for a major part of the encoder energy [2]. We therefore present a power-aware Motion Estimator for H.264 that adapts at run time according to the available energy level. We perform a set of adaptations at different Processing Stages of Motion Estimation. Our results show that in case of CIF videos (typically used in portable devices; but our approach is equally applicable to other video resolutions too), we achieve an average energy reduction of 52 times and 27 times as compared to UMHexagonS [12] and EPZS [14] respectively. This energy saving comes at the cost of an average loss of only 0.39 dB in Peak Signal to Noise Ratio (PSNR: an objective quality measure) and 23% increase in area (synthesized for 90nm technology). Muhammad Shafique 0001, Lars Bauer, Jörg Henkel |
ISLPED | 2 |
| 2008 | Efficient Resource Utilization for an Extensible Processor Through Dynamic Instruction Set AdaptationabstractState-of-the-art application-specific instruction set processors (ASIPs) allow the designer to define individual prefabrication customizations, thus improving the degree of specialization towards the actual application requirements, e.g., the computational hot spots. However, only a subset of hot spots can be targeted to keep the ASIP within a reasonable size. We propose a modular special instruction composition with multiple implementation possibilities per special instruction, compile-time embedded instructions to trigger a run-time adaptation of the instruction set, and a run-time system that dynamically selects an appropriate variation of the instruction set, i.e., a situation-dependent beneficial implementation for each special instruction. We thereby achieve a better efficiency of resource usage of up to 3.0 times (average 1.4 times) compared with current state-of-the-art ASIPs, resulting in a 3.1 times (average 1.4 times) improved application performance (compared with a general purpose processor up to 25.7 times and average 17.6 times). Lars Bauer, Muhammad Shafique 0001, Jörg Henkel |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2007 | RISPP: Rotating Instruction Set Processing PlatformabstractAdaptation in embedded processing is key in order to address efficiency. The concept of extensible embedded processors works well if a few a-priori known hot spots exist. However, they are far less efficient if many and possible at-design-time-unknown hot spots need to be dealt with. Our RISPP approach advances the extensible processor concept by providing flexibility through runtime adaptation by what we call "instruction rotation". It allows sharing resources in a highly flexible scheme of compatible components (called Atoms and Molecules). As a result, we achieve high speed-ups at moderate additional hardware. Furthermore, we can dynamically tradeoff between area and speed-up through runtime adaptation. We present the main components of our platform and discuss by means of an H.264 video codec. Lars Bauer, Muhammad Shafique 0001, Simon Kramer 0002, Jörg Henkel |
DAC | 1 |