Jörg Henkel

dblp:h/JorgHenkel · DBLP profile ↗
← Back
436ranked-venue papers
27as first author
102since 2021 · last 2026
0000-0001-9602-2922ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 387 · 27 first-author · 89 since 2021Software engineering, systems software and programming languages · 106 · 4 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 2 since 2021Computer networks · 13 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 13Security and privacy · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Efficient Federated Learning with Low-Rank Updates under Homomorphic Encryption
abstract
Federated Learning has been widely adopted for its ability to collaboratively train models without exposing raw data. However, the server-side aggregation process may still leak sensitive information about client data. Homomorphic Encryption enables privacy-preserving aggregation, but it introduces substantial communication overhead for clients and high computational costs for the server. To address these challenges, we propose HEAL-FL, a federated learning framework that is based on low-rank shared basis vectors across clients. Instead of transmitting full encrypted model updates, clients send only encrypted low-rank coefficients, thereby reducing both communication costs and server-side aggregation overhead. Furthermore, HEAL-FL incorporates a communication-efficient basis update scheme that relies exclusively on homomorphic addition at the server. Our evaluation across various homomorphic encryption schemes shows that HEAL-FL reduces client communication and server aggregation costs, leading to improved efficiency of Federated Learning systems. Notably, these savings translate into up to a significant reduction of 38.6% in total training time compared to conventional homomorphic FedAvg with full model parameter transmission, demonstrating the practical benefits of our approach.
Mohamed Aboelenien Ahmed, Mohamed Alsharkawy, Hassan Nassar, Heba Khdr, Jeferson González-Gómez, Jörg Henkel
DATE6
2026 TrustSeed: Lightweight Attestation Protocol for Ensuring LLM Integrity
abstract
Over the last couple of years, large language models have increasingly been integrated into many computing applications. For privacy preservation, they are now deployed on edge devices. However, these deployments are vulnerable to bit flip attacks and backdoor attacks that compromise the integrity of the model. Traditional remote attestation techniques fail to detect such manipulations due to the large model size and the stealthiness of the attacks.In this paper, we present TrustSeed, a lightweight functional attestation protocol that uses a single inference to ensure large language models’ integrity. TrustSeed verifies integrity by applying deterministic, seed-based modifications to model weights within a Trusted Execution Environment and comparing the last intermediate activations and output distribution against a golden reference on the verifier. This approach prevents precomputed or forged responses, ensuring freshness and unpredictability in each attestation round. Our analysis shows that output distribution and last intermediate activations are effective indicators of integrity. We test TrustSeed against bit-flip, data poisoning, and weight poisoning attacks, reliably detecting even single-bit alterations. Extensive evaluations on edge platforms and an HPC system demonstrate minimal overhead and up to 127× faster attestation compared to state-of-the-art full-model hashing.
Mohamed Alsharkawy, Mohamed Aboelenien Ahmed, Hassan Nassar, Jeferson González-Gómez, Heba Khdr, Osama Abboud, Xun Xiao, Jörg Henkel
DATE8
2026 GLEAM: A Graph-Learning Enhanced Adaptive Metaheuristic for Power-Aware Scheduling on Heterogeneous Cyber-Physical Systems
abstract
The increasing complexity of embedded and Cyber-Physical Systems (CPS) has accelerated the adoption of heterogeneous multi-core architectures, which combine performance and energy efficiency. However, scheduling dependent tasks on such platforms introduces significant challenges due to strict real-time constraints, high energy consumption, and the NP-hard nature of task mapping. This paper proposes a novel hybrid scheduling framework to jointly optimize energy efficiency and timeliness for Directed Acyclic Graph (DAG) applications. The framework operates in three tiers: first, a Genetic Algorithm (GA) performs a global search to determine near-optimal task-to-core mappings; second, a Dynamic Voltage and Frequency Scaling (DVFS) manager is integrated into the GA’s fitness function to accurately capture energy-performance trade-offs; and third, a Graph Neural Network (GNN) is trained to imitate the GA+DVFS policy, enabling fast and high-quality online scheduling decisions. Experimental results demonstrate that the proposed approach achieves a balanced trade-off between power consumption and deadline satisfaction, while the GNN significantly accelerates scheduling without compromising solution quality. Our GLEAM method reduced energy consumption on average by 49.08% and improved the makespan on average by 27.03% compared to baseline methods.
Amir Hossein Ansari, Mohsen Ansari, Sepideh Safari, Alireza Ejlali, Jörg Henkel
DATE5
2026 MIQARA: Mixed-Criticality Queue-based Architecture for Reconfigurable Accelerator Platforms
abstract
Coexistence of safety-critical control functions and besteffort computations in mixed-criticality systems poses a challenge in resource allocation and scheduling, as high-criticality jobs must adhere to strict timing guarantees, while lower-criticality jobs should make effective use of available resources without compromising the system’s safety and predictability. This paper introduces MIQARA1, a mixed-criticality queue-based architecture designed for reconfigurable accelerator platforms. MIQARA efficiently combines software-programmable CPUs with reconfigurable hardware, utilizing a dynamic job pipeline, token-based dependency tracking, and out-of-order scheduling to optimize resource utilization. At the same time, MIQARA has been designed to satisfy real-time constraints. MIQARA is evaluated on four FPGA platforms: the Zed Board, DipForty board, ZCU102 board, all of which have ARM CPUs implemented on chip, and Arty A7 with a RISC-V soft-core processor, representing systems that rely on soft CPUs. Results demonstrate substantial performance gains, particularly in terms of execution speed, flexibility, and adaptability to mixed-criticality workloads. The integration of features such as a streaming network further illustrates MIQARA’s scalability to complex data-intensive applications, making it a compelling solution for embedded mixed-criticality systems. MIQARA requires a hardware overhead of 17.8% and achieves a speedup of up to 4×.
Hassan Nassar, Martin Rapp, Lars Bauer, Mostafa Elshimy, Zeynep Demirdag, Jörg Henkel
DATE6
2026 Multi-Partner Project: A Holistic and Open-Source Approach to Efficient, Secure and Reliable AI Hardware Deployment in DI-EDAI
abstract
Artificial Intelligence (AI) has demonstrated strong capabilities across various domains over the past decade. Edge and specifically mission-critical applications, such as automotive and aerospace, require both high performance and efficiency without compromises in security and reliability. This stems from tightly constrained power consumption, failures that can have catastrophic consequences and devices that may be physically accessible to malicious actors. AI algorithm deployment to hardware also presents significant barriers, requiring specialized knowledge and expensive development tools. The DI-EDAI project aims to offer a holistic approach for connecting high-level AI algorithms with hardware implementations while tackling the aforementioned issues. Unlike other approaches that address individual aspects of the AI deployment flow, we investigate solutions across multiple layers of the design stack. Through our work we develop efficient hardware, map AI algorithms to hardware while simultaneously ensuring security and reliability. Furthermore, we leverage AI-techniques to assist with Electronic Design Automation (EDA) workflows for design optimization, verification and implementation. Our open source approach aims to reduce entry barriers, promote transparency and education, and spark innovation. This paper presents the current state of the DI-EDAI project at midterm, highlighting our latest contributions, identifying limitations in existing state-of-the-art approaches, and outlining ongoing work to address these gaps.
Georgios Sotiropoulos, Felix Frombach, Julian Höfer, Tanja Harbaum, Jürgen Becker 0001, Henrik Iver Thorøe, Vincent Meyers, Mehdi Baradaran Tahoori, Zeynep Demirdag, Mohammed Bakr Sikal, Hassan Nassar, Heba Khdr, Jörg Henkel, Christopher Wolters, Philipp van Kempen, Johannes Geier, Ulf Schlichtmann, Batuhan Sesli, Muhammad Sabih, Jakob Wittmann, Frank Hannig, Jürgen Teich, Lukas Steiner, Norbert Wehn, Mohamed Shelkamy Ali, Philipp Schmitz, Wolfgang Kunz, Stefan Koegler, Georg Sigl
DATE13
2026 RIFLE: Robust Distillation-based FL for Deep Model Deployment on Resource-Constrained IoT Networks
Pouria Arefijamal, Mahdi Ahmadlou, Bardia Safaei 0001, Jörg Henkel
ICC4
2026 HAMLET: Heterogeneous Adaptive Mapping and Low-Energy Task Scheduling for Heterogeneous Multicore IoT Devices
abstract
The rapid growth of Internet of Things (IoT) deployments and the increasing integration of multiple cores on a single chip have made managing power consumption and ensuring thermal safety critical challenges for multicore IoT devices. This paper proposes HAMLET, a heterogeneous adaptive mapping and low-energy task scheduling framework that jointly manages energy consumption, timing constraints, thermal safety, and reliability in multicore IoT platforms through a multi-objective genetic optimization approach. In the offline phase, task-to-core mapping and replication decisions are optimized to minimize peak power and energy consumption while preserving schedulability and reliability targets. Moreover, predictive temperature control is achieved in the runtime phase using a Long Short-Term Memory (LSTM) model, which anticipates core temperature evolution and enables proactive control actions. This approach ensures that tasks are allocated to processing cores to prevent thermal violations while maintaining system-level reliability. Furthermore, DVFS and DPM mechanisms are adaptively applied at runtime based on predicted thermal states to reduce energy consumption and avoid thermal emergencies. Experimental results demonstrate that the proposed method achieves up to 84.39% reduction in peak power and up to 81.98% reduction in energy consumption, while improving schedulability by up to 20.7% compared to state-of-the-art techniques, confirming its effectiveness for energy- and reliability-aware multicore IoT systems.
Amir Hossein Ansari, Mohsen Ansari, Alireza Ejlali, Jörg Henkel
IEEE Internet Things J.4
2026 SIREN: Multiobjective Game-Theoretic Scheduler Based on Memory-Driven Gray Wolf Optimization in Fog-Cloud Computing
abstract
Fog-cloud task scheduling faces the dual challenge of maintaining critical IoT workloads despite node failures while adhering to strict energy budgets. We present SIREN, a game-theoretic framework that treats fog nodes as strategic players, embedding reliability benefits and DVFS-aware energy costs directly into their payoffs. By searching the joint strategy space with a Memory-Driven Grey Wolf Optimizer (MDGWO), SIREN adapts placements, selective replication, and frequency settings to workload dynamics. Extensive evaluations on the Alibaba 2018 and Google 2011 cluster traces and on a latency-critical healthcare application demonstrate that SIREN converges to near-Nash schedules that minimize energy while maximizing reliability. Results confirm that SIREN delivers (i) 100% task success rates in critical healthcare scenarios, (ii) 2.08×–4.24× lower worst-case energy consumption than leading baselines, and (iii) a 3.9×–5.8× reduction in network usage, establishing a new benchmark for resilient, energy-efficient fog computing.
Abolfazl Younesi, Mohsen Ansari, Alireza Ejlali, MohammadAmin Fazli, Muhammad Shafique 0001, Jörg Henkel
IEEE Internet Things J.6
2026 ARDiS: A Portable and Unified Resource Management Framework in Real Hardware Systems
abstract
Designing efficient RM strategies is a cornerstone of modern computing, driving innovations in performance optimization, energy efficiency, and security. While simulators have long been the go-to tools for RM research, they fail to balance accuracy and practicality: high-fidelity simulators are excruciatingly slow, and low-fidelity ones compromise on reliability. Real hardware offers unparalleled precision and accuracy but remains underutilized due to significant barriers, including fragmented implementations, lack of portability, and prohibitive development overhead. We present ARDiS, the first open-source 1 and portable framework to provide a unified, architecture-agnostic platform for running system-level resource management (RM) techniques directly on real hardware. ARDiS eliminates the need to “reinvent the wheel,” enabling researchers to design, implement, and evaluate sophisticated RM strategies—including machine learning-based approaches—with minimal effort and maximum reproducibility. To demonstrate its versatility, we evaluate ARDiS on two real-world hardware platforms: a server-grade heterogeneous processor (Intel i9-12900) and a resource-constrained embedded system (NVIDIA Jetson TX2). Through extensive experimentation, we validate the ability of ARDiS to deliver accurate, scalable, and reproducible results across diverse platforms and application domains. By lowering the barriers to hardware-based RM research, ARDiS empowers the design automation community to explore new frontiers in system-level optimization and innovation.
Mohammed Bakr Sikal, Jeferson González-Gómez, Andreas Noebel, Heba Khdr, Jörg Henkel
ACM Trans. Design Autom. Electr. Syst.5
2026 Pilot: Power-Aware Hybrid Fault Tolerance in Multi-Core Embedded Systems
abstract
With the advancement of technology size and the integration of multiple cores on a single chip, the probability of fault occurrence has increased. These faults can be transient or permanent, requiring techniques to manage both types. Hybrid fault tolerance techniques have emerged as effective solutions to handle both types. In this paper, we propose a power-aware hybrid fault tolerance (called Pilot). Our approach utilizes checkpointing with rollback-recovery and primary/backup techniques, tolerating two kinds of faults. Moreover, in real-time embedded systems, power consumption is a critical constraint that must be managed. To do this, we exploit the Thermal Safe Power (TSP) constraint for each processing core. Based on this constraint and the utilization of each core, tasks are mapped and scheduled, while guaranteeing the timing constraints. Our experimental results demonstrate that our proposed methods can meet the reliability target by tolerating the optimal number of fault occurrences in each task while reducing power consumption. Our proposed methods are compared to state-of-the-art techniques in terms of schedulability, power consumption, Quality of Service (QoS), energy consumption, and reliability. The peak power and energy consumption are reduced on average by 34.2% and 15.9%, respectively, the QoS is improved on average to 28.7%, and the schedulability is improved on average to 14.6% while satisfying the system reliability target.
Amir Hossein Ansari, Moein Esnaashari, Sepideh Safari, Mohsen Ansari, Alireza Ejlali, Jörg Henkel
IEEE Trans. Parallel Distributed Syst.6
2025 Through Fabric: A Cross-world Thermal Covert Channel on TEE-enhanced FPGA-MPSoC Systems
abstract
The ever-evolving computing landscape gets more complex in every moment and the need for heterogeneous compute systems becomes more relevant. As the usability of such systems grew, finding methods for securing them became more relevant. Commercial vendors already introduced Trusted Execution Environments (TEEs) for those systems. TEEs serve the need for isolation, where sensitive data are processed in a secure world, and non-trusted applications are executed in the normal world. In this paper, we introduce Through Fabric: a novel attack against TEE-enhanced FPGA-MPSoCs. We show that existing benign hardware accelerators can be manipulated from the secure world to implement a temperature-based covert channel. We successfully run this attack on a commercial FPGA-MPSoC within the OP-TEE environment without additional access rights. We use an open-source implementation of AES for the accelerator and we reach a transmission speed of 2 bits per second with bit error rate of 1.9% and packet error rate of 4.3%. We are the first to show that a TEE can be bypassed on FPGA-MPSoCs via temperature-based covert channel communication.
Hassan Nassar, Jeferson González-Gómez, Varun Manjunath, Lars Bauer, Jörg Henkel
ASP-DAC5
2025 Special Session - Hardware-Software Co-Design for Machine Learning Systems Made Open-Source
abstract
Chip technologies are crucial for the digital transformation of industry and society. Machine Learning (ML) and Artificial Intelligence (AI) are increasingly shaping both daily life and industrial applications, with AI hardware playing a vital role in enabling efficient and scalable ML deployment. However, significant challenges remain in bridging the gap between ML algorithm development and hardware implementation, particularly for edge ML applications where efficiency, power constraints, and adaptability are critical. In such resource-constrained environments, hardware-software co-design becomes essential to achieve the necessary trade-offs between performance, energy efficiency, and system responsiveness. One of the key bottlenecks in ML hardware development is the lack of seamless integration between ML toolchains and electronic design automation (EDA) tools for hardware synthesis and mapping. Current solutions often require extensive manual optimization and costly proprietary software, limiting accessibility and innovation. Open-source tools can play a transformative role in democratizing ML hardware design, fostering collaboration, and addressing the growing shortage of skilled professionals. This paper covers key aspects of hardware-software co-design for ML systems, such as ML algorithms, hardware design, compiler technologies and system security, with a focus on open-source solutions. We highlight the critical need for open-source toolchains that connect ML model development with hardware synthesis and optimization and present solutions for custom hardware, as well as FPGA accelerators.
Mehdi Baradaran Tahoori, Vincent Meyers, Mahboobe Sadeghipourrudsari, Huashuangyang Xu, Jürgen Becker 0001, Tanja Harbaum, Felix Frombach, Julian Höfer, Georgios Sotiropoulos, Jörg Henkel, Zeynep Demirdag, Heba Khdr, Hassan Nassar, Ulf Schlichtmann, Johannes Geier, Philipp van Kempen, Georg Sigl, Stefan Koegler, Matthias Probst, Jürgen Teich, Frank Hannig, Muhammad Sabih, Batuhan Sesli, Norbert Wehn, Lukas Steiner, Wolfgang Kunz, Mohamed Shelkamy Ali
CODES+ISSS10
2025 Late Breaking Results: Decentralized Voting-Based Attestation for IoT Devices
abstract
Remote Attestation (RA) has become a valuable security service for Internet of Things (IoT) devices, as the security of these devices is often not prioritized during the manufacturing process. However, traditional RA schemes suffer from a single point of failure because they rely on a trusted verifier. To address this issue, we propose a voting-based blockchain attestation protocol that provides a reliable solution by eliminating the single point of failure through distributed verification across all nodes. In addition, it offers a traceable and immutable public history of the attestation results, which can be verified by external auditors at any time. Finally, we verify our proposed protocol on three NVIDIA Jetson embedded devices hosting up to 15 attestation nodes.
Mohamed Alsharkawy, Eren Sönmez, Jeferson González-Gómez, Hassan Nassar, Jörg Henkel
DAC5
2025 Late Breaking Results: The Hidden Risks of Activation Duration in PLPUFs
abstract
The security of Internet of Things (IoT) devices is crucial to protect the vast amounts of data exposed due to their widespread adoption. Authentication is one of the key aspects of IoT security, but it becomes increasingly challenging, especially for resource-constrained devices that require lightweight and efficient solutions. Physical Unclonable Functions (PUFs) have emerged as a promising lightweight solution by using the unique physical properties of Integrated Circuits (ICs). Pseudo Liner Feedback Shift Register PUF (PLPUF) is one of the state-of-the-art implementations known for its flexibility in altering the challenge-response space by changing the activation duration. In this work, we demonstrate that selecting an appropriate activation duration for PLPUF is critical, as improper choices can compromise security. By analyzing the linear dependency between the responses of different PLPUF pairs, our results reveal that predictability can reach up to $96 \%$ when an unsuitable activation duration is chosen.
Mohamed Alsharkawy, Jan Zwerschke, Hassan Nassar, Jeferson González-Gómez, Jörg Henkel
DAC5
2025 Centralized Training and Decentralized Control through the Actor-Critic Paradigm for Highly Optimized Multicores
abstract
While distributed, neural-network-based resource controllers represent the state of the art for their ability to cope with the ever-expanding decision space, such approaches suffer from several limitations, like conflicting control decisions and partial observability. These effects can significantly impair the controllers’ learning capabilities and the stability of their control policies, causing substantial performance losses. We are the first to solve this problem employing a centralized training and decentralized control regime to mitigate the aforementioned limitations. Specifically, we design a centralized neural network (critic) that evaluates the behavior of multiple decentralized neural controllers (actors) in a system-wide context. The objective of our proposed technique is to maximize the performance under a temperature constraint through dynamic voltage frequency scaling. The evaluation of our technique shows its superiority over the state of the art, yielding average (peak) performance improvements of 20% (34%), which we consider a breakthrough as the gains are measured on a real-world platform.
Benedikt Dietrich, Heba Khdr, Jörg Henkel
DAC3
2025 Contention-Aware Forecasting of Energy Efficiency through Sequence-Based Models in Modern Heterogeneous Processors
abstract
We present EffiCast, the first methodology for contentionaware energy efficiency forecasting in clustered heterogeneous processors using sequence-based models. Through extensive experimental analysis of energy efficiency sensitivities across core types, voltage/frequency (V/f) levels, application phases, and resource contention scenarios, EffiCast uncovers key factors driving energy efficiency variability in modern heterogeneous processors. Leveraging structured data generation and advanced LSTM- and Transformer-based models, EffiCast achieves unprecedented accuracy while outperforming state-of-the-art predictive techniques. Deployed on a real heterogenous processor with Intel’s oneDNN acceleration, EffiCast delivers inference latencies as low as 1.82 ms per sequence, enabling seamless integration into proactive resource management frameworks. With the ability to forecast future system states under dynamic workloads, EffiCast sets a new standard for energy efficiency optimization in energy-constrained application domains.
Mohammed Bakr Sikal, Jeferson González-Gómez, Heba Khdr, Jörg Henkel
DAC4
2025 Federated Reinforcement Learning for Optimizing the Power Efficiency of Edge Devices
abstract
Reinforcement learning (RL) holds great promise for adaptively optimizing microprocessor performance under power constraints. It allows for online learning of application characteristics at runtime and enables adjustment to varying system dynamics such as changes in the workload, user preferences or ambient conditions. However, online policy optimization remains resource-intensive, with high computational demand and requiring many samples to converge, making it challenging to deploy to edge devices. In this work, we overcome both of these obstacles and present federated power control using dynamic voltage and frequency scaling (DVFS). Our technique leverages federated RL and enables multiple independent power controllers running on separate devices to collaboratively train a shared DVFS policy, consolidating experience from a multitude of different applications, while ensuring that no privacy-sensitive information leaves the devices. This leads to faster convergence and to increased robustness of the learned policies. We show that our federated power control achieves 57 % average performance improvements over a policy that is only trained on local data. Compared to a state-of-the-art collaborative power control, our technique leads to 22 % better performance on average for the running applications under the same power constraint.
Benedikt Dietrich, Rasmus Müller-Both, Heba Khdr, Jörg Henkel
DATE4
2025 Hardware/Software Co-Analysis for Worst Case Execution Time Bounds
abstract
Ensuring that safety-critical systems meet timing constraints is crucial to avoid disastrous failures. To verify that timing requirements are met, a worst-case execution time (WCET) bound is computed. However, traditional WCET tools require a predefined timing model for each target processor, which is not available when using custom instruction set extensions. We introduce a novel approach based on hardware-software coanalysis that employs an instrumented hardware description of the target processor, removing the requirement for a separate timing model. We demonstrate this approach by extending the FemtoRV32 Individua RISC-V processor with a custom instruction set extension and show that it accurately models the timing behavior of the resulting system.
Can Joshua Lehmann, Lars Bauer, Hassan Nassar, Heba Khdr, Jörg Henkel
DATE5
2025 REAP-NVM: Resilient Endurance-Aware NVM-Based PUF Against Learning-Based Attacks
abstract
NVM-based PUFs offer secure authentication and cryptographic applications by exploiting NVMs' MLC to generate diverse, ML-attack-resistant responses. Yet, frequent writes degrade these PUFs, lowering reliability and lifespan. This paper presents a model to assess endurance effects on NVM PUFs, guiding the creation of more robust PUFs. Our novel NVM PUF design enhances endurance by evenly distributing writes, thus mitigating cell stress, achieving a 62x improvement over current solutions while preserving security against learning-based attacks.
Hassan Nassar, Ming-Liang Wei, Chia-Lin Yang, Jörg Henkel, Kuan-Hsun Chen
DATE4
2025 Multi-Partner Project: Open-Source Design Tools for Co-Development of AI Algorithms and AI Chips: (Initial Stage)
abstract
Chip technologies are crucial for the digital transformation of industry and society. Artificial Intelligence (AI) is playing an increasingly important role in both our daily lives and in industry. The development of advanced AI chip designs, essential for the successful deployment of AI, is of critical importance for innovation and competitiveness. However, challenges arise from the complexity of hardware development, expensive access to state-of-the-art design tools, and a global shortage of hardware experts. In addition to cost optimization, computational power, and energy consumption, security and trustworthiness are becoming increasingly important. This project aims to address these challenges in AI chip design by enabling efficient hardware development. We are developing a seamless transition between software-based AI model development and optimization, and efficient hardware implementation, while considering security, trustworthiness, and energy efficiency. An open-source approach plays a key role, facilitating access for small and medium-sized enterprises (SMEs) and expanding the community involved in AI chip design to help mitigate the shortage of skilled professionals.
Mehdi Baradaran Tahoori, Jürgen Becker 0001, Jörg Henkel, Wolfgang Kunz, Ulf Schlichtmann, Georg Sigl, Jürgen Teich, Norbert Wehn
DATE3
2025 FLARE: Fault Attack Leveraging Address Reconfiguration Exploits in Multi-Tenant FPGAs
Jayeeta Chaudhuri, Hassan Nassar, Dennis Gnad, Jörg Henkel, Mehdi Baradaran Tahoori, Krishnendu Chakrabarty
ETS4
2025 Invited Paper: Hardware-Software Co-Design for Highly Optimized, Customized, and Reliable AI Systems
abstract
Over the past decade, AI has been rapidly integrated into our daily life, coming in every shape and size and working across systems from big clouds to IoT. As a result, AI systems are increasingly requiring enhancements in model efficiency, hardware acceleration, and memory systems to satisfy stringent constraints on efficiency, reliability, and security. However, advancing across these fronts is challenging as compute demand outpaces Moore’s-law efficiency, hardening into an AI compute wall and an AI energy wall. Breaking through requires a unified AI co-design loop that co-optimizes algorithms and hardware, including efficient AI-to-hardware mapping, so that ongoing goals (accuracy, sparsity, latency) align with concrete hardware choices (precision modes, interconnects, memory hierarchies) and AI-specific execution and memory-reuse patterns. This paper details the principal co-design challenges, presents complementary strategies, and outlines a practical roadmap toward highly optimized, efficient, reliable, and secure AI systems.
Jörg Henkel, Mehdi Baradaran Tahoori, Heba Khdr, Hassan Nassar, Vincent Meyers, Deming Chen, Selin Yildirim, Yingbing Huang, Nirmal Saxena, Saurabh Hukerikar, Srivi Dhruvanarayan
ICCAD1
2025 R2T-Tiny: Runtime-Reconfigurable Throughput-Optimized TinyML for Hybrid Inference Acceleration on FPGA SoCs
abstract
The emergence of tiny machine learning (TinyML) has represented a paradigm shift toward energy efficient and on-device inference, with TinyML research primarily focusing on low-cost and energy-efficiency microcontroller units (MCUs). However, the low computational capabilities of MCUs greatly limit the performance of TinyML applications, especially in the case of throughput driven tasks. Field programmable gate arrays (FPGAs), are therefore a promising alternative to MCUs, offering high parallelism and low energy requirements. FPGA-based accelerators typically follow one of two directions: high-throughput resource-intensive pipelined designs, or low-throughput sequential systolic arrays. Incorporating both approaches into a hybrid streaming/sequential accelerator strikes a balance between resource efficiency and throughput, but incurs resource contention within a tiny FPGA. In this work, we address the aforementioned limitations and propose a throughput-driven hybrid acceleration methodology for TinyML. We introduce R2T-Tiny, an adaptive framework that brings layer-wise customizability into throughput-driven inference. By leveraging runtime partial reconfiguration, R2T-Tiny dynamically adjusts the accelerator type and applies tailored approximations per layer, achieving high throughput while adhering to the tight resource constraints of tiny embedded FPGAs. Our comprehensive evaluation on popular TinyML benchmarks showcases the capabilities of our framework in achieving high throughput inference on the PYNQ-Z2 FPGA board, increasing throughput by an average of 1.6x across 3 popular deep neural networks (DNNs) in the TinyML domain, when compared to DNNDK, a systolic array based accelerator from Xilinx, while incurring less than 1% accuracy loss.
Georgios Mentzos, Valentin Alexander Frey, Konstantinos Balaskas, Georgios Zervakis 0001, Jörg Henkel
ICCAD5
2025 Approximate Multiplier Mapping for Unfairness Mitigation in Energy-Efficient DNNs
abstract
Embedded devices struggle with the heavy computational demands of extensive neural network models, a problem partially addressed by integrating accelerators with numerous multiply-accumulate units. However, this solution increases energy consumption. While using approximate circuits in accelerators can lower energy usage, it compromises accuracy and raises concerns about maintaining fair inference across diverse populations, particularly in the medical field. This work leverages approximate multipliers in deep neural networks to retain fairness and reduce energy consumption while keeping the inference accuracy within strict thresholds.
Ourania Spantidi, Georgios Zervakis 0001, Jörg Henkel, Iraklis Anagnostopoulos
ISCAS3
2025 EDEN: Energy-aware Dynamic Genetic and Neural Network-based Path Predictive Routing and Clustering for Mobile SD-IoT Networks
abstract
The proliferation of battery-equipped smart devices in Internet of Things (IoT) applications has underscored the critical need for energy-efficient communication and computation solutions to extend lifetime. Meanwhile, the complex nature of mobile IoT networks’ communications poses significant challenges. Software-Defined Networking (SDN) offers minimized device-related overheads associated with processing and computations by centralizing energy-intensive tasks. The employed central controller in SDN can effectively follow, and manage the continuous topological alterations in dynamic mobile environments, thereby mitigating energy overhead imposed on individual IoT devices. On the other hand, clustering can further improve energy efficiency in IoT networks by reducing the number of transmissions, aggregating data efficiently, balancing the load among nodes, and optimizing routing paths. Accordingly, this paper introduces EDEN; an energy-aware SDN-based routing and clustering approach for mobile IoT networks to reduce energy consumption, increase network lifetime, and enhance reliability in the network in terms of Packet Delivery Ratio (PDR). To determine the optimal number of clusters and ensure a balanced distribution, EDEN utilizes a dynamic genetic algorithm to adaptively determine mutation and crossover rates. In selecting cluster heads, EDEN incorporates multiple objective function parameters, including node centrality, remaining energy, and distance, to optimize energy efficiency. Furthermore, EDEN employs a path prediction algorithm based on the LSTM Neural Network (NN) to forecast the trajectory of the mobile nodes to maintain cluster stability and reduce the frequency of re-clustering, thereby enhancing both energy efficiency and reliability. Extensive simulations in the NS3 environment demonstrate the effectiveness of the proposed solution, showing improvements in energy consumption by at least 32% while improving PDR by more than 98% compared to the state-of-the-art.
Negar Javadzadeh No, Hossein Taghizadeh, Mohammad Parsa Sedighi, Bardia Safaei 0001, Jörg Henkel
IWCMC5
2025 Enabling Printed Multilayer Perceptrons Realization via Area-Aware Neural Minimization
abstract
Printed Electronics (PE) set up a new path for the realization of ultra low-cost circuits that can be deployed in every-day consumer goods and disposables. In addition, PE satisfy requirements such as porosity, flexibility, and conformity. However, the large feature sizes in PE and limited device counts incur high restrictions and increased area and power overheads, prohibiting the realization of complex circuits. As a result, although printed Machine Learning (ML) circuits could open new horizons and bring “intelligence” in such domains, the implementation of complex classifiers, as required in target applications, is hardly feasible. In this paper, we aim to address this and focus on the design of battery-powered printed Multilayer Perceptrons (MLPs). To that end, we exploit fully-customized circuit (bespoke) implementations, enabled in PE, and propose a hardware-aware neural minimization framework dedicated for such customized MLP circuits. Our evaluation demonstrates that, for up to 3% accuracy loss, our co-design methodology enables, for the first time, battery-powered operation of complex printed MLPs.
Argyris Kokkinis, Georgios Zervakis 0001, Kostas Siozios, Mehdi Baradaran Tahoori, Jörg Henkel
IEEE Trans. Computers5
2025 Energy-Aware Heterogeneous Federated Learning via Approximate DNN Accelerators
abstract
In Federated Learning (FL), devices that participate in the training usually have heterogeneous resources, i.e., energy availability. In current deployments of FL, devices that do not fulfill certain hardware requirements are often dropped from the collaborative training. However, dropping devices in FL can degrade training accuracy and introduce bias or unfairness. Several works have tackled this problem on an algorithm level, e.g., by letting constrained devices train a subset of the server neural network (NN) model. However, it has been observed that these techniques are not effective w.r.t. accuracy. Importantly, they make simplistic assumptions about devices’ resources via indirect metrics, such as multiply accumulate (MAC) operations or peak memory requirements. We observe that memory access costs (that are currently not considered in simplistic metrics) have a significant impact on the energy consumption. In this work, for the first time, we consider on-device accelerator design for FL with heterogeneous devices. We utilize compressed arithmetic formats and approximate computing, targeting to satisfy limited energy budgets. Using a hardware-aware energy model, we observe that, contrary to the state of the art’s moderate energy reduction, our technique allows for lowering the energy requirements (by$4\times $) while maintaining higher accuracy.
Kilian Pfeiffer, Konstantinos Balaskas, Kostas Siozios, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 DPReF: Decentralized Key Generation Using Physical-Related Functions
abstract
Physical Unclonable Functions (PUFs) serve as a lightweight source to generate cryptographic keys utilizing the inherent physical device properties, making them particularly suitable for resource-constrained environments such as Internet of Things (IoT) devices. Recently, Physical-Related Functions (PReFs) extended PUFs to enable multiple devices to generate similar keys without the need to exchange or store them, improving security. However, state-of-the-art PReF implementations rely on a Trusted Third Party (TTP) to identify relative challenges, introducing a potential vulnerability if the TTP is compromised. In this work, we propose the first decentralized PReF protocol, removing reliance on the TTP and mitigating associated security risks. The proposed protocol allows relative challenges to be identified directly between devices in a decentralized manner. Additionally, we formalize a mathematical model to estimate the minimum number of devices required to build a network, based on the sizes of the PUF and the shared Challenge-Response Pair (CRP).. We demonstrate the generality of our model by verifying it across different types of state-of-the-art PUFs (Arbiter-based Non-Volatile Memory PUF (ANV-PUF) and Pseudo Linear Feedback Shift Register PUF (PLPUF).). We establish a 128 bit cryptographic key using the proposed protocol that matches the state-of-the-art but in a decentralized manner. Moreover, we prove that our protocol can be used to construct hardware-assisted attestation networks using ANV-PUF and PLPUF implementations with a shared secret of 16 bit that allows for both integrity and identity verification.
Mohamed Alsharkawy, Hassan Nassar, Jeferson González-Gómez, Xun Xiao, Osama Abboud, Jörg Henkel
ACM Trans. Embed. Comput. Syst.6
2025 Timekeepers: ML-Driven SDF Analysis for Power-Wasters Detection in FPGAs
abstract
As the integration of FPGAs into cloud computing platforms accelerates, the risk of fault injection attacks - especially through power-wasting designs - becomes increasingly critical. Malicious tenants can upload FPGA designs that, under specific input stimuli, generate excessive power consumption, jeopardizing the integrity of the shared power delivery network (PDN) and enabling denial-of-service or side-channel attacks. Traditional detection techniques relying on netlist and bitstream analysis struggle with generalization and can be evaded through circuit obfuscation and seemingly benign designs. In contrast to these netlist-based approaches, we introduce Timekeepers, a novel detection method that utilizes Standard Delay Format (SDF) timing data combined with machine learning to detect anomalous power behavior in synthesized FPGA designs. Our method trains a decision tree classifier on SDF files generated from both benign and malicious designs, focusing on timing characteristics such as propagation delays and setup/hold violations to identify power wasters at the primitive level. By abstracting away from circuit connectivity and emphasizing timing patterns, our framework is both scalable and robust across different FPGA architectures. The classifier independently evaluates each FPGA component and aggregates the results using a threshold-based voting system to improve detection granularity and reduce false positives. Timekeepers achieves 99.6% accuracy and demonstrates superior performance compared to state-of-the-art solutions. Furthermore, our approach is platform-agnostic and does not require access to netlists or bitstreams, preserving intellectual property confidentiality while enhancing pre-deployment security checks.
Mohamed Fathy, Hassan Nassar, Mohamed Abdelghany, Jörg Henkel
ACM Trans. Embed. Comput. Syst.4
2025 DIST: Distributed Learning-Based Energy-Efficient and Reliable Task Scheduling and Resource Allocation in Fog Computing
abstract
This paper presents DIST, a novel distributed reinforcement learning-based (DRL) framework for energyefficient and reliable task scheduling and resource allocation in fog computing, low-latency computing solutions driven by the rapid deployment of IoT devices, and time-sensitive applications. DIST is built based on a novel distributed Q-learning to enable fog nodes to learn an optimal strategy to balance energy consumption, task execution time, and system reliability. The main novelty includes a cooperative Dynamic Voltage and Frequency Scaling-enabled task scheduling policy that dynamically adjusts node energy level to ensure power consumption reduction without sacrificing deadline adherence or reliability. The results demonstrate that DIST reduces energy consumption by up to 52.26%, realizes 38% higher success rates, and reduces task wait times by up to 46.77%, compared with state-of-the-art algorithms.
Elyas Oustad, Abolfazl Younesi, Mohsen Ansari, Sepideh Safari, Mohammad Arman Soleimani, Jörg Henkel, Alireza Ejlali
IEEE Trans. Serv. Comput.6
2024 Hacking the Fabric: Targeting Partial Reconfiguration for Fault Injection in FPGA Fabrics
abstract
FPGAs are now ubiquitous in cloud computing infrastructures and reconfigurable system-on-chip, particularly for AI acceleration. Major cloud service providers such as Amazon and Microsoft are increasingly incorporating FPGAs for specialized compute-intensive tasks within their data centers. The availability of FPGAs in cloud data centers has opened up new opportunities for users to improve application performance by implementing customizable hardware accelerators directly on the FPGA fabric. However, the virtualization and sharing of FPGA resources among multiple users open up new security risks and threats. We present a novel fault attack methodology capable of causing persistent fault injections in partial bitstreams during the process of FPGA reconfiguration. This attack leverages powerwasters and is timed to inject faults into bitstreams as they are being loaded onto the FPGA through the reconfiguration manager, without needing to remain active throughout the entire reconfiguration process. Our experiments, conducted on a Pynq FPGA setup, demonstrate the feasibility of this attack on various partial application bitstreams, such as a neural network accelerator unit and a signal processing accelerator unit.
Jayeeta Chaudhuri, Hassan Nassar, Dennis Gnad, Jörg Henkel, Mehdi Baradaran Tahoori, Krishnendu Chakrabarty
ATS4
2024 Multi-Agent Reinforcement Learning for Thermally-Restricted Performance Optimization on Manycores
abstract
The problem of performance maximization under a thermal constraint has been tackled by means of dynamic voltage and frequency scaling (DVFS) in many system-level optimization techniques. State-of-the-art ones have exploited Su-pervised Learning (SL) to develop models that predict power and performance characteristics of applications and temperature of the cores. Such predictions enable proactive and efficient optimization decisions that exploit performance potentials under a temperature constraint. SL- based models are built at design time based on training data generated considering specific environment settings, i.e., processor architecture, cooling system, ambient temperature, etc. Hence, these models cannot adapt at runtime to different environment settings. In contrast, Reinforcement Learning (RL) employs an agent that explores and learns the environment at runtime, and hence can adapt to its potential changes. Nonetheless, using an RL agent to perform optimization on manycores is challenging because of the inherent large state/action spaces that might hinder the agent's ability to converge. To get the advantages of RL while tackling this challenge, we employ for the first time multi -agent RL to perform thermally-restricted performance optimization for manycores through DVFS. We investigated two RL algorithms-Table-based Q-Learning (TQL) and Deep Q-Learning (DQL)-and demonstrated that the latter outperforms the former. Compared to the state of the art, our DQL delivers a significant performance improvement of 34.96% on average, while also guaranteeing thermally -safe operation on the manycore. Our evaluation reveals the runtime adaptability of our DQL to varying workloads and ambient temperatures.
Heba Khdr, Mustafa Enes Batur, Kanran Zhou, Mohammed Bakr Sikal, Jörg Henkel
DATE5
2024 HBMorphic: FHE Acceleration via HBM-Enabled Recursive Karatsuba Multiplier on FPGA
abstract
Cloud computing offers advantages such as seamless scalability and speedup of computation. Nevertheless, these benefits come with notable tradeoffs, e.g., processing sensitive data without compromising security. Fully Homomorphic Encryption (FHE) solves this by processing of encrypted data. In this work, we develop an FHE hardware accelerator that uses a custom control interface to maximally utilize the bandwidth of HBM, following the memory access patterns of FHE.
Hassan Nassar, Lars Bauer, Jörg Henkel
FCCM3
2024 Covert-Hammer: Coordinating Power-Hammering on Multi-tenant FPGAs via Covert Channels
abstract
With the rise of AI, end of Moore's law, and the digitization of public services, the demand for accelerated computing is growing. To address this demand, major cloud service providers like Amazon Web Services, Microsoft Azure, and Google Cloud Platform have incorporated FPGA instances into their infrastructure with efficient and adaptable resource allocation models. Interest is increasing in multi-tenant FPGAs, which enable multiple users to utilize FPGA resources concurrently, while the FPGA can be split into smaller sections, one per tenant. Nevertheless, it introduces significant security vulnerabilities. For instance, by configuring a malicious circuit in one tenant's section of the FPGA, attacks that cause faults or crash the entire FPGA become feasible, affecting other tenants. By splitting an FPGA into smaller fractions, a single tenant has less potential to cause catastrophic outcomes. However, in this paper, we propose another threat, which is to perform an attack where several malicious tenants coordinate an attack using an unintended covert channel. We practically verify this possibility and introduce such a synchronized and coordinated voltage drop attack from multiple malicious tenants. For synchronization, the malicious tenants use a voltage-based covert channel. Our results show that the communication is robust reaching less than 1% packet error rate and that the attack is successful and avoids state-of-the-art countermeasures.
Hassan Nassar, Philipp Machauer, Dennis Gnad, Lars Bauer, Mehdi Baradaran Tahoori, Jörg Henkel
FPGA6
2024 Co-Designing NVM-based Systems for Machine Learning and In-memory Search Applications
abstract
With the rapid development of the Internet of Things, machine learning applications on edge devices with limited resources face challenges due to large data scales and irregular memory access patterns. Non-volatile memory (NVM) technologies provide promising solutions by offering larger capacity, low leakage power, and data persistence. In this paper, we discuss the potential of NVM technology in enhancing machine learning applications by improving energy efficiency and reducing latency through in-memory computation and different NVM write modes. The insights from this analysis provide valuable guidance to device researchers and system architects working to develop highperformance systems for machine learning and accelerators in large-scale search applications using NVMs.
Jörg Henkel, Lokesh Siddhu, Hassan Nassar, Lars Bauer, Jian-Jia Chen, Christian Hakert, Tristan Taylan Seidl, Kuan-Hsun Chen, Xiaobo Sharon Hu, Mengyuan Li 0001, Chia-Lin Yang, Ming-Liang Wei
ICCAD1
2024 DoS-FPGA: Denial of Service on Cloud FPGAs via Coordinated Power Hammering
abstract
The adoption of FPGA instances by major cloud service providers (CSPs) reflects the growing demand for accelerated and heterogeneous computing across various applications, e.g., AI. To improve the efficiency, utilization and virtualization, multi-tenant FPGAs allow multiple users to utilize FPGA resources concurrently, with each FPGA partition assigned to a separate tenant. However, this introduces significant security vulnerabilities, such as the potential for attacks by configuring a malicious circuit in one tenant's FPGA partition. One notable vulnerability is disrupting the FPGA's power distribution network, leading to faults or even crashing the entire FPGA, affecting other tenants. Usually, such an attack requires a considerable amount of resources. A naive solution would be splitting an FPGA into smaller fractions to reduce the potential for successful Power-Hammering by individual tenants and enhance the security. However, our paper demonstrates that even with smaller fractions per tenant, attacks can still occur. We propose the threat of coordinated attacks, where malicious tenants use an unintended covert channel between them. We practically validate this threat in a real cloud computing environment by introducing a synchronized and coordinated power-hammering attack from multiple malicious tenants. These tenants synchronize their actions using a voltage-based covert channel. Our results reveal the success of the attack, surpassing state-of-the-art countermeasures and detection mechanisms with a success rate exceeding 90%, compared to 30% for uncoordinated attacks.
Hassan Nassar, Philipp Machauer, Lars Bauer, Dennis Gnad, Mehdi Baradaran Tahoori, Jörg Henkel
ICCAD6
2024 Balancing Security and Efficiency: System-Informed Mitigation of Power-Based Covert Channels
abstract
As the digital landscape continues to evolve, the security of computing systems has become a critical concern. Power-based covert channels (e.g., thermal covert channel s (TCCs)), a form of communication that exploits the system resources to transmit information in a hidden or unintended manner, have been recently studied as an effective mechanism to leak information between malicious entities via the modulation of CPU power. To this end, dynamic voltage and frequency scaling (DVFS) has been widely used as a countermeasure to mitigate TCCs by directly affecting the communication between the actors. Although this technique has proven effective in neutralizing such attacks, it introduces significant performance and energy penalties, that are particularly detrimental to energy-constrained embedded systems. In this article, we propose different system-informed countermeasures to power-based covert channels from the heuristic and machine learning (ML) domains. Our proposed techniques leverage task migration and DVFS to jointly mitigate the channels and maximize energy efficiency. Our extensive experimental evaluation on two commercial platforms: 1) the NVIDIA Jetson TX2 and 2) Jetson Orin shows that our approach significantly improves the overall energy efficiency of the system compared to the state-of-the-art solution while nullifying the attack at all times.
Jeferson González-Gómez, Mohammed Bakr Sikal, Heba Khdr, Lars Bauer, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 Meta-Scanner: Detecting Fault Attacks via Scanning FPGA Designs Metadata
abstract
With the rise of the big data, processing in the cloud has become more significant. One method of accelerating applications in the cloud is to use field programmable gate arrays (FPGAs) to provide the needed acceleration for the user-specific applications. Multitenant FPGAs are a solution to increase efficiency. In this case, multiple cloud users upload their accelerator designs to the same FPGA fabric to use them in the cloud. However, multitenant FPGAs are vulnerable to low-level denial-of-service attacks that induce excessive voltage drops using the legitimate configurations. Through such attacks, the availability of the cloud resources to the nonmalicious tenants can be hugely impacted, leading to downtime and thus financial losses to the cloud service provider. In this article, we propose a tool for the offline classification to identify which FPGA designs can be malicious during operation by analysing the metadata of the bitstream generation step. We generate and test 475 FPGA designs that include 38% malicious designs. We identify and extract five relevant features out of the metadata provided from the bitstream generation step. Using ten-fold cross-validation to train a random forest classifier, we achieve an average accuracy of 97.9%. This significantly surpasses the conservative comparison with the state-of-the-art approaches, which stands at 84.0%, as our approach detects stealthy attacks undetectable by the existing methods.
Hassan Nassar, Jonas Krautter, Lars Bauer, Dennis Gnad, Mehdi Baradaran Tahoori, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2024 ML-Based Thermal and Cache Contention Alleviation on Clustered Manycores With 3-D HBM
abstract
Enabled by the recent advancements in 2.5D/3-D integration and packaging, the integration of clustered manycore processors with high-bandwidth memory (HBM) is gaining prominence to satisfy the increasing memory bandwidth demands. Although this integration can offer significant performance gains, it is still limited by cache contention in the final-level cache on the clusters and by the thermal issues in the 3-D HBM. While the existing state-of-the-art resource management techniques have tackled these issues in isolation, we argue that the cache contention and the temperature of both the manycore and the HBM must be considered jointly to harness the full performance potential of such modern architectures. To cover this gap in the literature, we present MTCM, the first resource management technique that considers the cache contention in maximizing the system performance, while maintaining the thermal safety across both the manycore and the HBM stack. Enabled by our accurate, yet lightweight, neural network models, our proposed task migration and dynamic voltage and frequency scaling policies can accurately predict the impact of runtime decisions on the performance and temperature of both the subsystems. Our extensive evaluation experiments reveal a significant performance improvement over existing state of the art by up to$1\times $, while maintaining thermal safety of both the manycore and the HBM.
Mohammed Bakr Sikal, Heba Khdr, Lokesh Siddhu, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 NPU-Accelerated Imitation Learning for Thermal Optimization of QoS-Constrained Heterogeneous Multi-Cores
abstract
Thermal optimization of a heterogeneous clustered multi-core processor under user-defined QoS targets requires application migration and DVFS. However, selecting the core to execute each application and the VF levels of each cluster is a complex problem because (1) the diverse characteristics and QoS targets of applications require different optimizations, and (2) per-cluster DVFS requires a global optimization considering all running applications. State-of-the-art resource management for power or temperature minimization either relies on measurements that are commonly not available (such as power) or fails to consider all the dimensions of the optimization (e.g., by using simplified analytical models). To solve this, ML methods can be employed. In particular, IL leverages the optimality of an oracle policy, yet at low run-time overhead, by training a model from oracle demonstrations. We are the first to employ IL for temperature minimization under QoS targets. We tackle the complexity by training NN at design time and accelerate the run-time NN inference using NPU. While such NN accelerators are becoming increasingly widespread, they are so far only used to accelerate user applications. In contrast, we use for the first time an existing accelerator on a real platform to accelerate NN-based resource management. To show the superiority of IL compared to RL in our targeted problem, we also develop multi-agent RL-based management. Our evaluation on a HiKey 970 board with an Arm big.LITTLE CPU and NPU shows that IL achieves significant temperature reductions at a negligible run-time overhead. We compare TOP-IL against several techniques. Compared to ondemand Linux governor, TOP-IL reduces the average temperature by up to 17 ˆC at minimal QoS violations for both techniques. Compared to the RL policy, our TOP-IL achieves 63 % to 89 % fewer QoS violations while resulting similar average temperatures. Moreover, TOP-IL outperforms the RL policy in terms of stability. We additionally show that our IL-based technique also generalizes to different software (unseen applications) and even hardware (different cooling) than used for training.
Martin Rapp, Heba Khdr, Nikita Krohmer, Jörg Henkel
ACM Trans. Design Autom. Electr. Syst.4
2023 Special Session - Non-Volatile Memories: Challenges and Opportunities for Embedded System Architectures with Focus on Machine Learning Applications
abstract
This paper explores the challenges and opportunities of integrating non-volatile memories (NVMs) into embedded systems for machine learning. NVMs offer advantages such as increased memory density, lower power consumption, non-volatility, and compute-in-memory capabilities. The paper focuses on integrating NVMs into embedded systems, particularly in intermittent computing, where systems operate during periods of available energy. NVM technologies bring persistence closer to the CPU core, enabling efficient designs for energy-constrained scenarios. Next, computation in resistive NVMs is explored, highlighting its potential for accelerating machine learning algorithms. However, challenges related to reliability and device non-idealities need to be addressed. The paper also discusses memory-centric machine learning, leveraging NVMs to overcome the memory wall challenge. By optimizing memory layouts and utilizing probabilistic decision tree execution and neural network sparsity, NVM-based systems can improve cache behavior and reduce unnecessary computations. In conclusion, the paper emphasizes the need for further research and optimization for the widespread adoption of NVMs in embedded systems presenting relevant challenges, especially for machine learning applications.
Jörg Henkel, Lokesh Siddhu, Lars Bauer, Jürgen Teich, Stefan Wildermann, Mehdi Baradaran Tahoori, Mahta Mayahinia, Jerónimo Castrillón, Asif Ali Khan, Hamid Farzaneh, João Paulo C. de Lima, Jian-Jia Chen, Christian Hakert, Kuan-Hsun Chen, Chia-Lin Yang, Hsiang-Yun Cheng
CASES1
2023 Smart Detection of Obfuscated Thermal Covert Channel Attacks in Many-core Processors
abstract
In thermal covert channel (TCC) attacks, malicious applications seek to leak private information in a stealthy and hard-to-detect manner. State-of-the-art approaches for TCC detection employ the Discrete Fourier Transform (DFT) combined with heuristics to identify possible channels. However, as we demonstrate in this paper, these approaches are limited when detecting short-duration attacks, where an attacker intentionally halts the transmission for a time interval to avoid the detection. In order to overcome this limitation of the state-of-the-art solutions, we propose the first detection method for short-duration TCC attacks. Our solution, Dotecca, is a machine learning-based technique that employs short windows of time-domain measurements instead of the DFT to detect TCCs. To evaluate our solution, we introduce a new obfuscated short-duration attack that disguises as a regular application from the perspective of a DFT spectrum. Our experiments show that the new obfuscated attack is able to remain undetected even under advanced DFT-based state-of-the-art detection approaches, reducing their detection accuracy to about 18 %. In contrast, our smart detection approach is able to detect state-of-the-art and new obfuscated attacks with an accuracy of 99 %. Moreover, our solution reduces the overhead of the DFT-based state-of-the-art solution by more than 14 ×.
Jeferson González-Gómez, Mohammed Bakr Sikal, Heba Khdr, Lars Bauer, Jörg Henkel
DAC5
2023 Late Breaking Results: Configurable Ring Oscillators as a Side-Channel Countermeasure
abstract
Side-channel attacks are a threat to computing devices. In this work, we propose a novel countermeasure against power analysis side-channel attacks. This countermeasure uses ring oscillators with runtime-configurable chain lengths to generate noise to hide the effects of the secret intermediate values on the device’s power consumption. We develop our countermeasure to be compatible with a state-of-the-art of side-channel-attack detection mechanism. Therefore, our solution does not incur any extra area overhead as it uses a subset of the circuit needed for detection. We evaluate our countermeasure using the test vector leakage assessment test (TVLA test). When our countermeasure is active no side-channel leakage could be detected.
Hassan Nassar, Simon Pankner, Lars Bauer, Jörg Henkel
DAC4
2023 Machine Learning-based Thermally-Safe Cache Contention Mitigation in Clustered Manycores
abstract
We present the first technique that mitigates cache contention under thermal constraints in clustered manycores. We show by means of extensive experiments that significant performance gains in this scenario can be achieved. The background is that concurrently-running applications on manycore clusters compete for the shared cache, slowing down their execution. In addition, heavy parallel computations on physically-close cores increase temperatures to non-sustainable levels, which in turn triggers a throttle down of voltage/frequency levels and hence performance is compromised. These problems are not unknown, but as our analysis shows, tackling them independently is sub-optimal. We introduce the first task migration technique that jointly mitigates cache contention while enforcing the thermal constraint at the same time. It works in conjunction with cluster-level dynamic voltage and frequency scaling. Our technique needs to predict the impact of task migration on performance considering cache contention. Since it is impossible to derive an analytical model for cache contention that is both sufficiently accurate and practically feasible to implement, we employ an accurate, yet lightweight neural network (NN) model. As a result, we can operate the manycore system at higher performance while safely staying within thermal constraints. We report a significant step forward in this paper and unveil new potentials for performance optimization.
Mohammed Bakr Sikal, Heba Khdr, Martin Rapp, Jörg Henkel
DAC4
2023 The First Concept and Real-world Deployment of a GPU-based Thermal Covert Channel: Attack and Countermeasures
abstract
Thermal covert channel (TCC) attacks have been studied as a threat to CPU-based systems over recent years. In this paper, we propose a new type of TCC attack that for the first time leverages the Graphics Processing Unit (GPU) of a system to create a stealthy communication channel between two malicious applications. We evaluate our new attack on two different real-world platforms: a GPU-dedicated general computing platform and a GPU-integrated embedded platform. Our results are the first to show that a GPU-based thermal covert channel attack is possible. From our experiments, we obtain a transmission rate of up to 8.75 bps with a very low error rate of less than 2 % for a 12-bit packet size, which is comparable to CPU-based TCCs in the state of the art. Moreover, we show how existing state-of-the-art countermeasures for TCCs need to be extended to tackle the new GPU-based attack at the cost of added overhead. To reduce this overhead, we propose our own DVFS-based countermeasure which mitigates the attack, while causing$2\times$less performance loss than the state-of-the-art countermeasure on a set of compute-intensive GPU benchmark applications.
Jeferson González-Gómez, Kevin Cordero-Zuñiga, Lars Bauer, Jörg Henkel
DATE4
2023 Hardware-Aware Automated Neural Minimization for Printed Multilayer Perceptrons
abstract
The demand of many application domains for flexibility, stretchability, and porosity cannot be typically met by the silicon VLSI technologies. Printed Electronics (PE) has been introduced as a candidate solution that can satisfy those requirements and enable the integration of smart devices on consumer goods at ultra low-cost enabling also in situ and on-demand fabrication. However, the large features sizes in PE constraint those efforts and prohibit the design of complex ML circuits due to area and power limitations. Though, classification is mainly the core task in printed applications. In this work, we examine, for the first time, the impact of neural minimization techniques, in conjunction with bespoke circuit implementations, on the area-efficiency of printed Multilayer Perceptron classifiers. Results show that for up to 5 % accuracy loss up to 8× area reduction can be achieved.
Argyris Kokkinis, Georgios Zervakis 0001, Kostas Siozios, Mehdi Baradaran Tahoori, Jörg Henkel
DATE5
2023 Extended Abstract: Monitoring-based Thermal Management for Mixed-Criticality Systems
abstract
With a rapidly growing number of functions in embedded real-time systems, it becomes inevitable to integrate tasks of different safety integrity levels (SILs) into one mixed-criticality system. Here, it is important to not only isolate shared architectural resources, as tasks executing on different cores may also interfere via the processor's thermal manager. In order to prevent a scenario where best-effort tasks cause deadline violations for critical tasks, we propose a thermal management strategy that guarantees a sufficient thermal isolation between tasks of different SILs, and simultaneously reduces the run-time of best-effort tasks by up to 45 % compared to the state of the art without incurring any real-time violations for critical tasks.
Marcel Mettler, Martin Rapp, Heba Khdr, Daniel Mueller-Gritschneder, Jörg Henkel, Ulf Schlichtmann
DATE5
2023 Aggregating Capacity in FL through Successive Layer Training for Computationally-Constrained Devices
abstract
Federated learning (FL) is usually performed on resource-constrained edge devices, e.g., with limited memory for the computation. If the required memory to train a model exceeds this limit, the device will be excluded from the training. This can lead to a lower accuracy as valuable data and computation resources are excluded from training, also causing bias and unfairness. The FL training process should be adjusted to such constraints. The state-of-the-art techniques propose training subsets of the FL model at constrained devices, reducing their resource requirements for training. However, these techniques largely limit the co-adaptation among parameters of the model and are highly inefficient, as we show: it is actually better to train a smaller (less accurate) model by the system where all the devices can train the model end-to-end than applying such techniques. We propose a new method that enables successive freezing and training of the parameters of the FL model at devices, reducing the training’s resource requirements at the devices while still allowing enough co-adaptation between parameters. We show through extensive experimental evaluation that our technique greatly improves the accuracy of the trained model (by 52.4 p.p. ) compared with the state of the art, efficiently aggregating the computation capacity available on distributed devices.
Kilian Pfeiffer, Ramin Khalili, Jörg Henkel
NeurIPS3
2023 ATLAS: Aging-Aware Task Replication for Multicore Safety-Critical Systems
abstract
A major requirement of safety-critical systems is high reliability at low power consumption. Dynamic voltage and frequency (v/f) scaling (DVFS) techniques are widely exploited to reduce power consumption. However, DVFS through downscaling v/f levels has a negative impact on the reliability of the tasks running on the cores, and through upscaling v/f levels has circuitlevel aging effects. To achieve high reliability in multicore safetycritical systems, task replication as a fault-tolerant technique is an established way to deal with the negative effect of downscaling v/f levels, but it may accelerate aging effects due to elevating the on-chip temperatures. In this paper, we propose an aging-aware task replication (called ATLAS) method that solves the problem of satisfying the desired reliability target for a set of periodic hard real-time tasks which are executed on a multicore system. The proposed method satisfies the reliability target of the tasks through updating the required number of replicas for each task at different years. We replicate the tasks through our proposed formulas such that the reliability target is satisfied. However, task replication increases the temperature of the system and accelerates aging. To decelerate aging, we attempt to reduce the temperature while mapping and scheduling the tasks. We have also developed a modified demand bound function (DBF) for our aging-aware task replication method to verify scheduling the realtime tasks. Compared to the existing state-of-the-art techniques, experimental results for safety-critical applications on different configurations of multicore systems demonstrate the efficiency and effectiveness of our proposed method. Experiments show that our proposed method improves schedulability on average by 16.1% and reduces the temperature on average by 7.4°C compared to state-of-the-art methods while meeting the system reliability target.
Mohsen Ansari, Sepideh Safari, Amir Yeganeh-Khaksar, Roozbeh Siyadatzadeh, Pourya Gohari-Nazari, Heba Khdr, Muhammad Shafique 0001, Jörg Henkel, Alireza Ejlali
RTAS8
2023 ReLIEF: A Reinforcement-Learning-Based Real-Time Task Assignment Strategy in Emerging Fault-Tolerant Fog Computing
abstract
Due to the real-time requirements in several IoT applications, fog computing has emerged to overcome the long latency and other constraints of cloud computing. Due to the high probability of packet loss, energy limitation of IoT devices, and the external disturbances that may frequently occur on the fog infrastructure, the timing constraints of real-time tasks may be compromised. Therefore, the reliability of executing real-time tasks has always been a significant challenge in fog computing. In addition to the correct execution of the tasks, it is also important to execute them before their deadlines according to their real-time classification. State-of-the-art methods generally focus on the delay or functionality of tasks in fog computing systems. However, those methods do not widely focus on the reliability of tasks with real-time constraints in dynamic environments. In this article, a novel primary backup task assignment strategy based on machine learning (ReLIEF) is proposed to improve the reliability of fog-based IoT systems. To identify suitable nodes for the execution of the primary and backup tasks, ReLIEF employs a reinforcement learning (RL) approach, which has an outstanding performance in dynamic environments by establishing a balance between communication delay and workload on each fog device. Based on the simulations, our newly proposed technique has been able to reduce the amount of task dropping rate by up to 84% against the state of the art. Moreover, it is capable of balancing the workload distribution while increasing the reliability of the system by nearly 72% compared with its counterparts.
Roozbeh Siyadatzadeh, Fatemeh Mehrafrooz, Mohsen Ansari, Bardia Safaei 0001, Muhammad Shafique 0001, Jörg Henkel, Alireza Ejlali
IEEE Internet Things J.6
2023 Co-Design of Approximate Multilayer Perceptron for Ultra-Resource Constrained Printed Circuits
abstract
Printed Electronics (PE) exhibits on-demand, extremely low-cost hardware due to its additive manufacturing process, enabling machine learning (ML) applications for domains that feature ultra-low cost, conformity, and non-toxicity requirements that silicon-based systems cannot deliver. Nevertheless, large feature sizes in PE prohibit the realization of complex printed ML circuits. In this work, we present, for the first time, an automated printed-aware software/hardware co-design framework that exploits approximate computing principles to enable ultra-resource constrained printed multilayer perceptrons (MLPs). Our evaluation demonstrates that, compared to the state-of-the-art baseline, our circuits feature on average 6x (5.7x) lower area (power) and less than 1% accuracy loss.
Giorgos Armeniakos, Georgios Zervakis 0001, Dimitrios Soudris, Mehdi Baradaran Tahoori, Jörg Henkel
IEEE Trans. Computers5
2023 Massively Parallel Circuit Setup in GPU-SPICE
abstract
SPICE simulations are the industry standard to analyze circuits for decades. However, they are computationally complex as each circuit is simulated at the transistor-level where individual transistor is modeled with dozens of sophisticated equations. This limits the practicality of SPICE simulations to relatively small circuits. However, this is in a direct conflict with the ever-increasing demands of circuit designers in which SPICE simulations for large circuits (e.g., DSPs, AES, etc.) at full accuracy are inevitably required to fulfill new industrial standards like automotive safety ISO 26262 with tool confidence level 1. To accelerate SPICE simulation without sacrificing accuracy, state-of-the-art approaches have started to employ GPUs to parallelize the LU-factorization and device linearization phases. Instead of focusing on these phases, this article demonstrates for the first time that when large circuits come into play, a new and equally important performance bottleneck emerges at the circuit setup phase. Speeding up the circuit setup phase in SPICE is our key focus in this paper. Our two implementations demonstrate that our GPU-based circuit setup reduces the analysis time from 4.5 days to merely 89 seconds for a 256-bit multiplier, which consists of more than 1M transistors. Our achieved speedup is 4396x compared to the baseline (open-source NGSPICE) and more than 2x compared to commercial (HSPICE and Spectre) SPICE circuit setup.
Victor M. van Santen, Fu Lam Florian Diep, Jörg Henkel, Hussam Amrouch
IEEE Trans. Computers3
2023 Model-to-Circuit Cross-Approximation For Printed Machine Learning Classifiers
abstract
Printed electronics (PEs) promises on-demand fabrication, low nonrecurring engineering costs, and subcent fabrication costs. It also allows for high customization that would be infeasible in silicon, and bespoke architectures prevail to improve the efficiency of emerging PE machine learning (ML) applications. Nevertheless, large feature sizes in PE prohibit the realization of complex ML models in PE, even with bespoke architectures. In this work, we present an automated, cross-layer approximation framework tailored to bespoke architectures that enable complex ML models, such as multilayer perceptrons (MLPs) and support vector machines (SVMs), in PE. Our framework adopts cooperatively a hardware-driven coefficient approximation of the ML model at algorithmic level, a netlist pruning at logic level, and a voltage overscaling at the circuit level. Extensive experimental evaluation on 12 MLPs and 12 SVMs and more than 6000 approximate and exact designs demonstrates that our model-to-circuit cross-approximation delivers power and area optimal designs that, compared to the state-of-the-art exact designs, feature on average 51% and 66% area and power reduction, respectively, for less than 5% accuracy loss. Finally, we demonstrate that our framework enables 80% of the examined classifiers to be battery-powered with almost identical accuracy with the exact designs, paving thus the way toward smart complex printed applications.
Giorgos Armeniakos, Georgios Zervakis 0001, Dimitrios Soudris, Mehdi Baradaran Tahoori, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 AdaPT: Fast Emulation of Approximate DNN Accelerators in PyTorch
abstract
Current state-of-the-art employs approximate multipliers to address the highly increased power demands of deep neural network (DNN) accelerators. However, evaluating the accuracy of approximate DNNs is cumbersome due to the lack of adequate support for approximate arithmetic in DNN frameworks. We address this inefficiency by presenting AdaPT, a fast emulation framework that extends PyTorch to support approximate inference as well as approximation-aware retraining. AdaPT can be seamlessly deployed and is compatible with the most DNNs. We evaluate the framework on several DNN models and application fields, including CNNs, LSTMs, and GANs for a number of approximate multipliers with distinct bitwidth values. The results show substantial error recovery from approximate retraining and reduced inference time up to$53.9 \times $with respect to the baseline approximate implementation.
Dimitrios Danopoulos, Georgios Zervakis 0001, Kostas Siozios, Dimitrios Soudris, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 Memory Carousel: LLVM-Based Bitwise Wear Leveling for Nonvolatile Main Memory
abstract
Emerging non-volatile memory yields, alongside many advantages, technical shortcomings, such as reduced cell lifetime. Although many wear-leveling approaches exist to extend the lifetime of such memories, usually a trade-off for the granularity of wear-leveling has to be made. Due to iterative write schemes (repeatedly sense and write), wear-out of memory in certain systems is directly dependent on the written bit value and thus can be highly imbalanced, requiring dedicated bit-wise wear-leveling. Such a bit-wise wear-leveling so far has only be proposed together with a special hardware support. However, if no dedicated hardware solutions are available, especially for commercial off-the-shelf systems with non-volatile memories, a software solution can be crucial for the system lifetime. In this work, we propose entirely software-based bit-wise wearleveling, where the position of bits within CPU words in main memory is rotated on a regular basis. We leverage the LLVM intermediate representation to adjust load and store operations of the application with a custom compiler pass. Experimental evaluation shows that the lifetime by applying local rotation within the CPU word can be extended by a factor of up to 21×. We also show that our method can incorporate with coarser-grained wear-leveling, e.g. on block granularity and assist achievement of higher lifetime improvements.
Nils Hölscher, Christian Hakert, Hassan Nassar, Kuan-Hsun Chen, Lars Bauer, Jian-Jia Chen, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2023 Golden-Free Robust Age Estimation to Triage Recycled ICs
abstract
Nondestructive golden-free detection of recycled/counterfeit integrated circuits (ICs) is the focus of this article. This is achieved by estimating the functional/operational age of the IC. The age estimation method is based on exploiting short-term aging effects in advanced transistor technologies to induce bit errors at the IC’s output. Gate-level simulations are used to capture the impact of workload on short-term aging. In advanced technology nodes, including bulk CMOS at 45 nm or below and FinFET, combining transistor aging with ultrafast voltage scaling magnifies the effects of aging-induced degradation at high voltage when voltage scales to a lower level, causing short-term aging-based timing violations. These timing violations create bit errors at IC outputs. We employ the bit error patterns to build a machine learning (ML)-based nonlinear regression model to estimate the IC’s age. Our study confirms that short-term aging-induced output bit error patterns can be used to estimate long-term age of an IC. If the IC’s age is beyond a predefined threshold, it can be marked as recycled. Although this article considers the FinFET technology, the method applies to bulk CMOS advanced nodes at 45 nm or below. We model IC-to-IC variations taking into account the voltage scaling. We demonstrate the approach on two cryptographic ICs and the method accurately estimates the long-term age of an IC, facilitating recycled IC detection.
Virinchi Roy Surabhi, Prashanth Krishnamurthy, Hussam Amrouch, Jörg Henkel, Ramesh Karri, Farshad Khorrami
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 ANV-PUF: Machine-Learning-Resilient NVM-Based Arbiter PUF
abstract
Physical Unclonable Functions (PUFs) have been widely considered an attractive security primitive. They use the deviations in the fabrication process to have unique responses from each device. Due to their nature, they serve as a DNA-like identity of the device. But PUFs have also been targeted for attacks. It has been proven that machine learning (ML) can be used to effectively model a PUF design and predict its behavior, leading to leakage of the internal secrets. To combat such attacks, several designs have been proposed to make it harder to model PUFs. One design direction is to use Non-Volatile Memory (NVM) as the building block of the PUF. NVM typically are multi-level cells, i.e, they have several internal states, which makes it harder to model them. However, the current state of the art of NVM-based PUFs is limited to ‘weak PUFs’, i.e., the number of outputs grows only linearly with the number of inputs, which limits the number of possible secret values that can be stored using the PUF. To overcome this limitation, in this work we design the Arbiter Non-Volatile PUF (ANV-PUF) that is exponential in the number of inputs and that is resilient against ML-based modeling. The concept is based on the famous delay-based Arbiter PUF (which is not resilient against modeling attacks) while using NVM as a building block instead of switches. Hence, we replace the switch delays (which are easy to model via ML) with the multi-level property of NVM (which is hard to model via ML). Consequently, our design has the exponential output characteristics of the Arbiter PUF and the resilience against attacks from the NVM-based PUFs. Our results show that the resilience to ML modeling, uniqueness, and uniformity are all in the ideal range of 50%. Thus, in contrast to the state-of-the-art, ANV-PUF is able to be resilient to attacks, while having an exponential number of outputs.
Hassan Nassar, Lars Bauer, Jörg Henkel
ACM Trans. Embed. Comput. Syst.3
2023 Cache-Based Side-Channel Attack Mitigation for Many-Core Distributed Systems via Dynamic Task Migration
abstract
Side-channel attacks (SCA) are a serious threat to cryptographic systems due to mostly unavoidable information leakage. Cache-based SCAs take advantage of cache inherent timing properties on shared memory systems to extract security-critical information. In this paper, we present a novel approach to mitigate cache-based SCAs on distributed many-core systems, based on a resource management technique. Our solution leverages dynamic task migration as a mechanism to ensure a secure execution scenario for security-critical applications. Additionally, we propose a resource-management-based mechanism to ensure a secure execution when migration is not possible due to a lack of available resources. We evaluate our solution in terms of gained security and performance impact using the Sniper simulator for different configurations. Results show that our technique effectively develops resilience against SCAs, while causing a low performance slowdown (1.6% on average, 9% worst case). For all tested benchmarks, our worst-case performance slowdown is 20% less than a state-of-the-art countermeasure. Moreover, our solution utilizes less than 1 ms of system run-time overhead for a 64 core platform with 100% utilization.
Jeferson González-Gómez, Lars Bauer, Jörg Henkel
IEEE Trans. Inf. Forensics Secur.3
2022 DISTREAL: Distributed Resource-Aware Learning in Heterogeneous Systems
abstract
We study the problem of distributed training of neural networks (NNs) on devices with heterogeneous, limited, and time-varying availability of computational resources. We present an adaptive, resource-aware, on-device learning mechanism, DISTREAL, which is able to fully and efficiently utilize the available resources on devices in a distributed manner, increasing the convergence speed. This is achieved with a dropout mechanism that dynamically adjusts the computational complexity of training an NN by randomly dropping filters of convolutional layers of the model. Our main contribution is the introduction of a design space exploration (DSE) technique, which finds Pareto-optimal per-layer dropout vectors with respect to resource requirements and convergence speed of the training. Applying this technique, each device is able to dynamically select the dropout vector that fits its available resource without requiring any assistance from the server. We implement our solution in a federated learning (FL) system, where the availability of computational resources varies both between devices and over time, and show through extensive evaluation that we are able to significantly increase the convergence speed over the state of the art without compromising on the final accuracy.
Martin Rapp, Ramin Khalili, Kilian Pfeiffer, Jörg Henkel
AAAI4
2022 Cross-Layer Approximation For Printed Machine Learning Circuits
abstract
Printed electronics (PE) feature low non-recurring engineering costs and low per unit-area fabrication costs, enabling thus extremely low-cost and on-demand hardware. Such low-cost fabrication allows for high customization that would be infeasible in silicon, and bespoke architectures prevail to improve the efficiency of emerging PE machine learning (ML) applications. However, even with bespoke architectures, the large feature sizes in PE constraint the complexity of the ML models that can be implemented. In this work, we bring together, for the first time, approximate computing and PE design targeting to enable complex ML models, such as Multi-Layer Perceptrons (MLPs) and Support Vector Machines (SVMs), in PE. To this end, we propose and implement a cross-layer approximation, tailored for bespoke ML architectures. At the algorithmic level we apply a hardware-driven coefficient approximation of the ML model and at the circuit level we apply a netlist pruning through a full search exploration. In our extensive experimental evaluation we consider 14 MLPs and SVMs and evaluate more than 4300 approximate and exact designs. Our results demonstrate that our cross approximation delivers Pareto optimal designs that, compared to the state-of-the-art exact designs, feature 47% and 44% average area and power reduction, respectively, and less than 1% accuracy loss.
Giorgos Armeniakos, Georgios Zervakis 0001, Dimitrios Soudris, Mehdi Baradaran Tahoori, Jörg Henkel
DATE5
2022 NPU-Accelerated Imitation Learning for Thermal- and QoS-Aware Optimization of Heterogeneous Multi-Cores
abstract
Task migration and dynamic voltage and frequency scaling (DVFS) are indispensable means in thermal optimization of a heterogeneous clustered multi-core processor under user-defined quality of service (QoS) targets. However, selecting the core to execute each application and the voltage/frequency (V/f) levels of each cluster is a complex problem because 1) the diverse characteristics and QoS targets of applications require different optimizations, and 2) V/f levels are often shared between cores on a cluster, which requires a global optimization considering all running applications. State-of-the-art techniques for power or temperature minimization either rely on measurements that are often not available (such as power) or fail to consider all the dimensions of the problem (e.g., by using simplified analytical models). Imitation learning (IL) enables to use the optimality of an oracle policy, yet at low run-time overhead, by training a model from oracle demonstrations. We are the first to employ IL for temperature minimization under QoS targets. We tackle the complexity by using a neural network (NN) model and accelerate the NN inference using a neural processing unit (NPU). While such NN accelerators are becoming increasingly widespread on end devices, they are so far only used to accelerate user applications. In contrast, we use an accelerator on a real platform to accelerate NN-based resource management. Our evaluation on a HiKey970 board with an Arm big.LITTLE CPU and an NPU shows significant temperature reductions at a negligible overhead while satisfying OoS targets.
Martin Rapp, Nikita Krohmer, Heba Khdr, Jörg Henkel
DATE4
2022 Thermal- and Cache-Aware Resource Management based on ML- Driven Cache Contention Prediction
abstract
While on-chip many-core systems enable a large number of applications to run in parallel, the increased overall performance may come at the cost of complicating the performance constraints of individual applications due to contention on shared resources. For instance, the competition for last-level cache by concurrently-running applications may lead to slowing down the execution and to potentially violating individual performance constraints. Clustered many-cores reduce cache contention at chip level by sharing caches only at cluster level. To reduce cache con-tention within a cluster, state-of-the art techniques aim to co-map a memory-intensive application with a compute-intensive application onto one cluster. However, compute-intensive applications typ-ically consume high power, and therefore, executing another application in their nearby cores may lead to high temperatures. Hence, there is a trade-off between cache contention and temperature. This paper is the first to consider this trade-off through a novel thermal- and cache-aware resource management technique. We build a neural network (NN)-based model to predict the slowdown of the application execution induced by cache contention feeding our resource management technique that then optimizes the application mapping and selects the voltage/frequency levels of the clus-ters to compensate for the potential contention-induced slowdown. Thereby, it meets the performance constraints, while minimizing temperature. Compared to the state of the art, our technique significantly reduces the temperature by 30% on average, while satisfying performance constraints of all individual applications.
Mohammed Bakr Sikal, Heba Khdr, Martin Rapp, Jörg Henkel
DATE4
2022 Approximate Computing and the Efficient Machine Learning Expedition
abstract
Approximate computing (AxC) has been long accepted as a design alternative for efficient system implementation at the cost of relaxed accuracy requirements. Despite the AxC research activities in various application domains, AxC thrived the past decade when it was applied in Machine Learning (ML). The by definition approximate notion of ML models but also the increased computational overheads associated with ML applications-that were effectively mitigated by corresponding approximations-led to a perfect matching and a fruitful synergy. AxC for AI/ML has transcended beyond academic prototypes. In this work, we enlighten the synergistic nature of AxC and ML and elucidate the impact of AxC in designing efficient ML systems. To that end, we present an overview and taxonomy of AxC for ML and use two descriptive application scenarios to demonstrate how AxC boosts the efficiency of ML systems.
Jörg Henkel, Hai Li 0001, Anand Raghunathan, Mehdi Baradaran Tahoori, Swagath Venkataramani, Xiaoxuan Yang 0001, Georgios Zervakis 0001
ICCAD1
2022 ARMOR: A Reliable and Mobility-Aware RPL for Mobile Internet of Things Infrastructures
abstract
Mobile portable embedded devices are becoming an integral part of our daily activities in the vision of Internet of Things (IoT). Nevertheless, due to lack of mobility support in the IPv6 routing protocol for low-power and lossy networks (RPLs), which is standardized for multihop IoT infrastructures, providing reliable communications in terms of packet delivery ratio (PDR) in mobile IoT applications has become significantly challenging. While several studies tried to enhance the adaptability of RPL to network dynamics, their utilized routing metrics have prevented them from establishing long-lasting reliable paths. Furthermore, the stochastic parent replacement policy in the standard version of RPL has intensified this challenge. Aside from this, due to the existing tradeoff between reliability and power efficiency, most of the existing approaches have only concentrated on one of these concerns without paying attention to the other one. To address these issues, this article introduces ARMOR, a routing mechanism built upon RPL, which employs a novel mobility-aware routing metric, i.e., time to reside (TTR), and a corresponding parent replacement policy. According to the motion characteristics of the mobile objects, TTR provides an estimation of how long the nodes will be in the transmission range of each other. This enables ARMOR to select nodes, which provide longer connection period and consequently higher reliability. In comparison with the state of the art, while keeping the power consumption constant, ARMOR significantly improves the amount of PDR in the network by up to$2.5\times $, while it enhances the reliability against the original version of this protocol by up to$4.2\times $.
Ali Asghar Mohammad Salehi, Bardia Safaei 0001, Amir Mahdi Hosseini Monazzah, Lars Bauer, Jörg Henkel, Alireza Ejlali
IEEE Internet Things J.5
2022 An FPGA-based Approach to Evaluate Thermal and Resource Management Strategies of Many-core Processors
abstract
The continuous technology scaling of integrated circuits results in increasingly higher power densities and operating temperatures. Hence, modern many-core processors require sophisticated thermal and resource management strategies to mitigate these undesirable side effects. A simulation-based evaluation of these strategies is limited by the accuracy of the underlying processor model and the simulation speed. Therefore, we present, for the first time, an field-programmable gate array (FPGA)-based evaluation approach to test and compare thermal and resource management strategies using the combination of benchmark generation, FPGA-based application-specific integrated circuit (ASIC) emulation, and run-time monitoring. The proposed benchmark generation method enables an evaluation of run-time management strategies for applications with various run-time characteristics. Furthermore, the ASIC emulation platform features a novel distributed temperature emulator design, whose overhead scales linearly with the number of integrated cores, and a novel dynamic voltage frequency scaling emulator design, which precisely models the timing and energy overhead of voltage and frequency transitions. In our evaluations, we demonstrate the proposed approach for a tiled many-core processor with 80 cores on four Virtex-7 FPGAs. Additionally, we present the suitability of the platform to evaluate state-of-the-art run-time management techniques with a case study.
Marcel Mettler, Martin Rapp, Heba Khdr, Daniel Mueller-Gritschneder, Jörg Henkel, Ulf Schlichtmann
ACM Trans. Archit. Code Optim.5
2022 CoMeT: An Integrated Interval Thermal Simulation Toolchain for 2D, 2.5D, and 3D Processor-Memory Systems
abstract
Processing cores and the accompanying main memory working in tandem enable modern processors. Dissipating heat produced from computation remains a significant problem for processors. Therefore, the thermal management of processors continues to be an active subject of research. Most thermal management research is performed using simulations, given the challenges in measuring temperatures in real processors. Fast yet accurate interval thermal simulation toolchains remain the research tool of choice to study thermal management in processors at the system level. However, the existing toolchains focus on the thermal management of cores in the processors, since they exhibit much higher power densities than memory. The memory bandwidth limitations associated with 2D processors lead to high-density 2.5D and 3D packaging technology: 2.5D packaging technology places cores and memory on the same package; 3D packaging technology takes it further by stacking layers of memory on the top of cores themselves. These new packagings significantly increase the power density of the processors, making them prone to overheating. Therefore, mitigating thermal issues in high-density processors (packaged with stacked memory) becomes even more pressing. However, given the lack of thermal modeling for memories in existing interval thermal simulation toolchains, they are unsuitable for studying thermal management for high-density processors. To address this issue, we present the first integrated Core and Memory interval Thermal (CoMeT) simulation toolchain.CoMeTcomprehensively supports thermal simulation of high- and low-density processors corresponding to four different core-memory (integration) configurations—off-chip DDR memory, off-chip 3D memory, 2.5D, and 3D.CoMeTsupports several novel features that facilitate overlying system research.CoMeTadds only an additional ~5% simulation-time overhead compared to an equivalent state-of-the-art core-only toolchain. The source code ofCoMeThas been made open for public use under theMITlicense.
Lokesh Siddhu, Rajesh Kedia, Shailja Pandey, Martin Rapp, Anuj Pathania, Jörg Henkel, Preeti Ranjan Panda
ACM Trans. Archit. Code Optim.6
2022 On the Reliability of FeFET On-Chip Memory
abstract
Ferroelectric Field-Effect Transistor (FeFET) is a promising future technology for non-volatile on-chip memories. It is rapidly attracting an ever-increasing attention from industry. The key advantage of FeFETs is full compatibility with the existing CMOS fabrication process beside their very low power consumption. To enable ultra-dense memories, 1-FeFET AND Arrays were proposed in which a memory cell is formed from merely a single FeFET. All access transistors, which are traditionally needed to operate memory cells, are removed. However, this imposes a new challenge ofindirect write disturbances. Neighboring memory cells are indirectly degraded whenever adirect write operationoccurs to a particular FeFET cell. Only recently the impact of such indirect disturbances on the FeFET reliability was experimentally investigated at device (i.e., transistor) level. However, to explore and properly judge the feasibility of 1-FeFET AND Arrays for on-chip memories, investigating only the reliability of individual cells is indeed insufficient. Bridging the gap between the device level and system (i.e., chip) level is inevitable. In the presence of indirect disturbances, the position of a write access within the array plays a key role, which is governed by the running workloads. In addition, whether the write operation flips the previously stored value or not also plays an important role with regards to reliability. Hence, running workloads, which determine not only the position of the memory cells to be written but also the values written to them, plays an essential role in determining 1-FeFET AND Array reliability over time. Therefore, studying the reliability of FeFETs only at the device level (as done in state of the art) is insufficient. In this work, we investigate, for the first time, the reliability of FeFET memories from device to system level. To achieve that, we develop a unified model capturing the impact of bothindirectdisturbances anddirectwrites on the reliability of FeFET cells. Our study at system level then employs the unified model in the context of application workloads. We investigate different array sizes, write voltages, write methods and a wide range of workloads using the example of CPU caches as an example of on-chip memory. We demonstrate that indirect write disturbances are the dominate effect degrading the reliability of FeFET memories. For most cells, it contributes over 90 percent to the overall induced degradation. This provides guidelines for researchers at both device and circuit level to optimize the FeFET reliability further while considering thehiddenimpact of indirect write disturbances.
Paul R. Genssler, Victor M. van Santen, Jörg Henkel, Hussam Amrouch
IEEE Trans. Computers3
2022 A Framework for Crossing Temperature-Induced Timing Errors Underlying Hardware Accelerators to the Algorithm and Application Layers
abstract
Temperature rising is an unavoidable effect on VLSI and has always been a critical issue in any system-on-chip – especially when targeting compute-intensive applications. This effect increases the delay in hardware accelerators, resulting in timing errors due to unsustainable clock frequency, whose impact must be carefully evaluated on design time to measure the performance degradation of the hardware accelerator. Further, a hardware operating at a higher temperature accelerates device aging, which incurs in more timing errors. This issue is usually addressed with the inclusion of timing guardbands that compensate for the deleterious effects of temperature, ensuring the hardware accelerator works within a reliable zone, i.e., without any timing errors caused by temperature effects at runtime. However, guardbands directly result in considerable performance and efficiency losses because the circuit will be clocked at a frequency lower than its full potential. Accelerators on edge devices often dismiss such guardbands to explore the full potential of the designed circuits, posing an enormous design challenge as this approach requires a careful evaluation of the impact of timing errors on the quality of the target applications. Many algorithms, such as in multimedia and machine learning applications, are capable of tolerating hardware errors. Yet, these algorithms have a dynamic behavior (i.e., closed-loop) where a timing error can be propagated, affecting subsequent steps. Measuring the degradation-induced errors in these applications is very challenging given that an accurate gate-level simulation to investigate degradation-induced timing errors needs to be coupled dynamically with a system-level simulator to unveil how induced errors in the underlying hardware ultimately impact the algorithm execution in the hardware accelerator.This is the first work to achieve this goal. State-of-the-art works have studied accelerators under timing-errors when removing (or narrowing) guardbands. However, their approach was suitableonly for open-loop hardware accelerators which are entirely agnostic of complex interactions of the algorithms. Unlike prior work, this paper investigates temperature- and aging-induced timing-errors in the joint accelerator-algorithm interactions and their runtime impacts. Our framework investigates aging effects across the different layers starting from transistor physics all the way up to the algorithm layer. The hardware accelerator employed as a case study in this work is the sum of absolute differences (SAD), which is the most compute-intensive accelerator on commercial video encoder for mobile applications. Our results demonstrate the runtime behavior impacts of three advanced block-matching algorithms of the video encoder in a joint operation by a SAD accelerator under timing-errors induced by temperature and aging effects considering a 14nm FinFET technology.
Guilherme Paim, Hussam Amrouch, Leandro M. G. Rocha, Brunno Abreu, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel
IEEE Trans. Computers7
2022 Impact of NCFET Technology on Eliminating the Cooling Cost and Boosting the Efficiency of Google TPU
abstract
Recent breakthroughs in Neural Networks (NNs) led to significant accuracy improvements of several machine learning applications such as image classification and voice recognition. However, this accuracy improvement comes at the cost of an immense increase in computation demands. NNs became one of the most common and computationally intensive workloads in today's datacenters. To address these computational demands, Google announced in 2016 the Tensor Processing Unit (TPU), an advanced custom ASIC accelerator for NN inference. Two new TPU versions (v2 and v3) followed in 2017 and 2018 that support also training. Google TPUv3 packs an immense processing power ($\mathrm{90TFLOPS}$per chip) in a tiny and condensed area, leading to very high on-chip power densities and thus excessive temperature. In this article, superlattice thermoelectric cooling, which is one of the emerging on-chip cooling, is considered as an advanced cooling example for Google TPU and we investigate the impact of Negative Capacitance FET (NCFET), which is one of the recent emerging technologies, on the cooling and efficiency of TPU. Through full-chip design, of the computational core of the TPU, based on$14\mathrm{nm}$Intel FinFET technology and multiphysics temperature simulations, we demonstrate that NCFET can significantly minimize the required cooling-cost. More than 4000 NCFET configurations are evaluated in order to traverse the entire design space defined by the thickness of the ferroelectric layer of NCFET, the operating voltage, cooling, and the operating frequency, in addition to all possible FinFET's configurations. Moreover, our experimental evaluation shows that by eliminating the cooling cost, NCFET delivers 2.8x higher efficiency compared to the conventional FinFET baseline.
Sami Salamin, Georgios Zervakis 0001, Florian Klemme, Hammam Kattan, Yogesh Singh Chauhan, Jörg Henkel, Hussam Amrouch
IEEE Trans. Computers6
2022 Trojan Detection in Embedded Systems With FinFET Technology
abstract
This study considers detecting Trojans in circuits using FinFET technology non-destructively, when a golden Integrated Circuit (IC) is unavailable. The method employs short-term aging effects in FinFET transistors and circuit overclocking to induce bit errors at the circuit outputs in conjunction with Machine Learning (ML) tools learning Trojan-free behavior. Short-term aging causes delays along multiple paths in the IC to vary dynamically, causing bit errors at circuit outputs. Overclocking enhances this in FinFET but is not necessary for bulk CMOS technology. We use bit error patterns at the output of the circuit to detect Trojans using an ML classifier trained on simulations of the Trojan-free circuit. The study shows efficacy of the method by using dynamic short-term aging-aware standard cell libraries with FinFET technology that are modeled by considering the dynamic short-term aging of each cell. Trojan detection is robust to chip-to-chip variations. We apply the technique on fourteen Trust-Hub Trojans. Our method detects Trojans with$>$95% accuracy. Trojan detection in FinFET technology is more challenging than in bulk CMOS because the voltage range for switching from a high to low value is smaller. Therefore we use overclocking.
Virinchi Roy Surabhi, Prashanth Krishnamurthy, Hussam Amrouch, Jörg Henkel, Ramesh Karri, Farshad Khorrami
IEEE Trans. Computers4
2022 FeFET-Based Binarized Neural Networks Under Temperature-Dependent Bit Errors
abstract
Ferroelectric FET (FeFET) is a highly promising emerging non-volatile memory (NVM) technology, especially for binarized neural network (BNN) inference on the low-power edge. The reliability of such devices, however, inherently depends on temperature. Hence, changes in temperature during run time manifest themselves as changes in bit error rates. In this work, we reveal the temperature-dependent bit error model of FeFET memories, evaluate its effect on BNN accuracy, and propose countermeasures. We begin on the transistor level and accurately model the impact of temperature on bit error rates of FeFET. This analysis reveals temperature-dependent asymmetric bit error rates. Afterwards, on the application level, we evaluate the impact of the temperature-dependent bit errors on the accuracy of BNNs. Under such bit errors, the BNN accuracy drops to unacceptable levels when no countermeasures are employed. We propose two countermeasures: (1) Training BNNs for bit error tolerance by injecting bit flips into the BNN data, and (2) applying a bit error rate assignment algorithm (BERA) which operates in a layer-wise manner and does not inject bit flips during training. In experiments, the BNNs, to which the countermeasures are applied to, effectively tolerate temperature-dependent bit errors for the entire range of operating temperature.
Mikail Yayla, Sebastian Buschjäger, Aniket Gupta, Jian-Jia Chen, Jörg Henkel, Katharina Morik, Kuan-Hsun Chen, Hussam Amrouch
IEEE Trans. Computers5
2022 Thermal-Aware Design for Approximate DNN Accelerators
abstract
Recent breakthroughs in Neural Networks (NNs) have made DNN accelerators ubiquitous and led to an ever-increasing quest on adopting them from Cloud to edge computing. However, state-of-the-art DNN accelerators pack immense computational power in a relatively confined area, inducing significant on-chip power densities that lead to intolerable thermal bottlenecks. Existing state of the art focuses on using approximate multipliers only to trade-off efficiency with inference accuracy. In this work, we present a thermal-aware approximate DNN accelerator design in which we additionally trade-off approximation with temperature effects towards designing DNN accelerators that satisfy tight temperature constraints. Using commercial multi-physics tool flows for heat simulations, we demonstrate how our thermal-aware approximate design reduces the temperature from 139$^{\circ }$C, in an accurate circuit, down to 79$^{\circ }$C. This enables DNN accelerators to fulfill tight thermal constraints, while still maximizing the performance and reducing the energy by around 75% with a negligible accuracy loss of merely 0.44% on average for a wide range of NN models. Furthermore, using physics-based transistor aging models, we demonstrate how reductions in voltage and temperature obtained by our approximate design considerably improve the circuit’s reliability. Our approximate design exhibits around 40% less aging-induced degradation compared to the baseline design.
Georgios Zervakis 0001, Iraklis Anagnostopoulos, Sami Salamin, Ourania Spantidi, Isai Roman-Ballesteros, Jörg Henkel, Hussam Amrouch
IEEE Trans. Computers6
2022 CaPUF: Cascaded PUF Structure for Machine Learning Resiliency
abstract
With the rise of the Internet of Things (IoT), resource-constrained and power-constrained devices attract more attention. The need for lightweight solutions as alternatives to resource-intensive applications became more urgent. Moreover, as the number of connected devices grew, authenticating them became more challenging. Traditionally, this would be performed by using hash functions and secure memory to store a key, which both come at a high cost. physical unclonable functions (PUFs) emerged as a suitable lightweight alternative to hash functions to authenticate the devices. Using the inherent minute differences between integrated circuits (ICs), they can generate IC-specific responses for input challenges coming from a so-called verifier. Through the years, machine learning (ML) has been used to attack PUFs by modeling them and accurately predicting their response to a given challenge. This stimulated research on ML-resilient PUFs. This resilience came with the significant area and challenge-to-response delay overheads. In this work, we introduce the novel cascaded PUF (CaPUF) and show that it is resilient against state-of-the-art ML-based attacks, i.e., logistic regression (LR) and support vector machines (SVMs). These attacks could not achieve accuracy better than 52% against our CaPUF, which is only as good as flipping a coin. Additionally, our CaPUF requires 89% less area compared to state-of-the-art ML-resilient PUFs.
Hassan Nassar, Lars Bauer, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 MLCAD: A Survey of Research in Machine Learning for CAD Keynote Paper
abstract
Due to the increasing size of integrated circuits (ICs), their design and optimization phases (i.e., computer-aided design, CAD) grow increasingly complex. At design time, a large design space needs to be explored to find an implementation that fulfills all specifications and then optimizes metrics like energy, area, delay, reliability, etc. At run time, a large configuration space needs to be searched to find the best set of parameters (e.g., voltage/frequency) to further optimize the system. Both spaces are infeasible for exhaustive search typically leading to heuristic optimization algorithms that find some tradeoff between design quality and computational overhead. Machine learning (ML) can build powerful models that have successfully been employed in related domains. In this survey, we categorize how ML may be used and is used for design-time and run-time optimization and exploration strategies of ICs. A metastudy of published techniques unveils areas in CAD that are well explored and underexplored with ML, as well as trends in the employed ML algorithms. We present a comprehensive categorization and summary of the state of the art on ML for CAD. Finally, we summarize the remaining challenges and promising open research directions.
Martin Rapp, Hussam Amrouch, Yibo Lin, Bei Yu 0001, David Z. Pan, Marilyn Wolf, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2022 Energy-Efficient DNN Inference on Approximate Accelerators Through Formal Property Exploration
abstract
Deep neural networks (DNNs) are being heavily utilized in modern applications, putting energy-constraint devices to the test. To bypass high energy consumption issues, approximate computing has been employed in DNN accelerators to balance out the accuracy-energy reduction trade-off. However, the approximation-induced accuracy loss can be very high and drastically degrade the performance of the DNN. Therefore, there is a need for a fine-grain mechanism that would assign specific DNN operations to approximation to maintain acceptable DNN accuracy, while achieving low energy consumption. We present an automated framework for weight-to-approximation mapping through formal property exploration for approximate DNN accelerators. At the MAC unit level, our experimental evaluation surpassed already energy-efficient mappings by more than$\times 2$in terms of energy gains, while supporting a fine-grain control over the introduced approximation.
Ourania Spantidi, Georgios Zervakis 0001, Iraklis Anagnostopoulos, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Variability-Aware Approximate Circuit Synthesis via Genetic Optimization
abstract
One of the major barriers that CMOS devices face at nanometer scale is increasing parameter variation due to manufacturing imperfections. Process variations severely inhibit the reliable operation of circuits, as the operational frequency at the nominal process corner is insufficient to suppress timing violations across the entire variability spectrum. To avoid variability-induced timing errors, previous efforts impose pessimistic and performance-degrading timing guardbands atop the operating frequency. In this work, we employ approximate computing principles and propose a circuit-agnostic automated framework for generating variability-aware approximate circuits that eliminate process-induced timing guardbands. Variability effects are accurately portrayed with the creation of variation-aware standard cell libraries, fully compatible with standard EDA tools. The underlying transistors are fully calibrated against industrial measurements from Intel 14nm FinFET in which both electrical characteristics of transistors and variability effects are accurately captured. In this work, we explore the design space of approximate variability-aware designs to automatically generate circuits of reduced variability and increased performance without the need for timing guardbands. Experimental results show that by introducing negligible functional error of merely$\boldsymbol {5.3 \times 10^{-3}}$, our variability-aware approximate circuits can be reliably operated under process variations without sacrificing the application performance.
Konstantinos Balaskas, Florian Klemme, Georgios Zervakis 0001, Kostas Siozios, Hussam Amrouch, Jörg Henkel
IEEE Trans. Circuits Syst. I Regul. Pap.6
2022 Bridging the Gap Between Voltage Over-Scaling and Joint Hardware Accelerator-Algorithm Closed-Loop
abstract
Voltage over-scaling (VOS) optimizes energy while causing timing errors due to an unsustainable clock frequency. Many algorithms, such as in multimedia and machine learning applications, are capable of tolerating such errors. VOS has never been investigated in hardware accelerators running closed-loop algorithms. As the errors impact most decisions and actions in the subsequent steps, closed-loops dynamically change the execution flow. Timing errors should be evaluated by an accurate gate-level simulation, but a large gap still remains: how these timing errors propagate from the underlying hardware all the way up to the entire algorithm run, where they just may degrade the performance and quality of service of the application at stake? This paper tackles this issue showing a framework for VOS investigation, embracing any kind of application. Our framework simulates the VOS-induced timing errors at gate-level, dynamically linking the hardware result with the algorithm and vice versa during the evolution of the runtime of the application. The state-of-the-art VOS literature for video encoding application fails to assess the ultimate impacts of VOS-induced timing errors, as current works open the encoding loops. Unlike those, our work investigates the ultimate impact of a hardware accelerator dynamically carrying through to the video encoder all VOS-induced timing errors and preserving the full compliance to the standard. We employ a parallel sum of absolute differences (SAD) hardware accelerator as a case study. We assess the performance of the overall encoder under varying timing guardbands. Next, it is demonstrated that, under VOS, the ultimate impact in compression efficiency is related to the video’s motion intensity. Additionally, the advantages of timing guardband controlled reduction are clearly quantified in our results by virtue of the framework. Reducing at maximum 9.5% the clock frequency, energy savings (up to 16.5% in energy/operation) are achieved in SAD for video compression.
Guilherme Paim, Hussam Amrouch, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel
IEEE Trans. Circuits Syst. Video Technol.5
2022 Towards a New Thermal Monitoring Based Framework for Embedded CPS Device Security
abstract
This article introduces a thermal side channel as a proxy for the behavior of embedded processors to detect changes in the behavior in a cyber-physical system. Such changes may be due to software/hardware attacks and altered processors. Since control system processes are periodic computations, the thermal side channels exhibit a temporal pattern. This enables the detection of altered code and changed device characteristics. We present a machine learning approach to estimate the activity of the embedded device from the time sequence of thermal images and show that deviations from expected behavior can be detected. The approach is validated on a multi-core processor running a periodic computational code. The infrared imager collects thermal imagery from the processor, which is cooled from the backside. Instead of an external imager, one can deploy a finite number of on-chip temperature sensors. This article shows that integrating on-chip temperature sensors allows robust real-time monitoring of the processor behavior. Finally, we offer a machine learning approach to optimally place the on-chip sensors to aid detection.
Naman Patel, Prashanth Krishnamurthy, Hussam Amrouch, Jörg Henkel, Michael Shamouilian, Ramesh Karri, Farshad Khorrami
IEEE Trans. Dependable Secur. Comput.4
2022 Software-Managed Read and Write Wear-Leveling for Non-Volatile Main Memory
abstract
In-memory wear-leveling has become an important research field for emerging non-volatile main memories over the past years. Many approaches in the literature perform wear-leveling by making use of special hardware. Since most non-volatile memories only wear out from write accesses, the proposed approaches in the literature also usually try to spread write accesses widely over the entire memory space. Some non-volatile memories, however, also wear out from read accesses, because every read causes a consecutive write access. Software-based solutions only operate from the application or kernel level, where read and write accesses are realized with different instructions and semantics. Therefore different mechanisms are required to handle reads and writes on the software level. First, we design a method to approximate read and write accesses to the memory to allow aging aware coarse-grained wear-leveling in the absence of special hardware, providing the age information. Second, we provide specific solutions to resolve access hot-spots within the compiled program code (text segment) and on the application stack. In our evaluation, we estimate the cell age by counting the total amount of accesses per cell. The results show that employing all our methods improves the memory lifetime by up to a factor of 955×.
Christian Hakert, Kuan-Hsun Chen, Horst Schirmeier, Lars Bauer, Paul R. Genssler, Georg von der Brüggen, Hussam Amrouch, Jörg Henkel, Jian-Jia Chen
ACM Trans. Embed. Comput. Syst.8
2022 Introduction and Evaluation of Attachability for Mobile IoT Routing Protocols With Markov Chain Analysis
abstract
Reliability of routing mechanisms in wireless networks is typically measured with Packet Delivery Ratio (PDR). Basically, PDR is reported with an optimistic assumption that the topology is fully constructed, and the nodes have started their packet transmission. This is despite the fact that prior to being able to transmit packets, nodes must first join the network, and then try to keep connected as much as possible. This is a key factor in the overall reliability provided by the routing protocols, especially in mobile IoT applications, where disconnections occur frequently. Nevertheless, there is a lack of appropriate metrics, which could evaluate the routing mechanisms from this perspective. Accordingly, this paper introduces attachability; a new metric for evaluating the capability of routing protocols in assisting the mobile or stationary nodes in joining, and maintaining their connections to the network. Our newly proposed metric is calculated via Markov chain analysis along with the sample frequency-based estimating technique. To evaluate attachability, we have simulated a mobile IoT infrastructure, and conducted a comprehensive set of experiments on different versions of the IPv6 Routing Protocol for Low-power and lossy networks (RPL). Based on our observations, attachability is significantly dependent on the employed metrics and path selection policies in the routing mechanisms. Among the three different versions of RPL, including the original version (ORPL), which is standardized for stationary IoT applications, and two mobility-aware versions, i.e., MARPL, and OMARPL, OMARPL showed up to 42%, and 10% of improvement in terms of attachability against ORPL, and MARPL, respectively.
Bardia Safaei 0001, Hossein Taghizade, Amir Mahdi Hosseini Monazzah, Kimia Talaei Khoosani, Parham Sadeghi, Ali Asghar Mohammad Salehi, Jörg Henkel, Alireza Ejlali
IEEE Trans. Netw. Serv. Manag.7
2022 Power-Aware Checkpointing for Multicore Embedded Systems
abstract
Increasing the number of cores integrated on a single chip offers a great potential for the implementation of fault-tolerant techniques to achieve high reliability in real-time embedded systems. Checkpointing with rollback-recovery is a well-established technique to tolerate transient faults in multicore platforms. To consider the worst-case fault occurrence scenario, checkpointing technique requires to re-execute some parts of the tasks, and that might lead to simultaneous execution of task parts with high power consumptions, which eventually might result in a peak power increase beyond the thermal design power (TDP). Exceeding TDP can elevate on-chip temperatures beyond safe limits, and thereby triggering countermeasures that throttle down the voltage and frequency levels or power gate the cores. Such countermeasures might lead to violating task deadlines and degrading the system's reliability. To avoid such severe scenarios, it is inevitable to consider the impact of applying fault-tolerant techniques on the power consumption and prevent violating the power constraint of the chip, i.e., TDP. This paper presents for the first time, a peak-power-aware checkpointing (PPAC) technique that tolerates a given number of faults,k, while at the same time meets the power constraint in hard real-time embedded systems. To do this, our proposed technique (PPAC) adjusts the timing of the checkpoints, which have lower power consumption than the tasks to the execution time points that have power spikes beyond TDP. Moreover, PPAC exploits the available slack times on the cores to delay the execution of some tasks to avoid the remaining power spikes beyond TDP, which could not be mitigated by solely adjusting checkpoints. To evaluate our technique, we extend the state-of-the-art system-level simulator, gem5, with the state-of-the-art checkpointing module in Linux. Our experimental results show that our proposed technique is able to tolerate a given number of faults without exceeding the timing and power constraints in hard real-time embedded systems. The resulting peak power reduction achieved by our technique compared to state-of-the-art techniques is an average of 23%. Moreover, our technique employs the Dynamic Power Management (DPM) during the slack times resulting at runtime in the case of fault-free scenarios, which provides energy savings with an average of 17.28% and up to 61.1%.
Mohsen Ansari, Sepideh Safari, Heba Khdr, Pourya Gohari-Nazari, Jörg Henkel, Alireza Ejlali, Shaahin Hessabi
IEEE Trans. Parallel Distributed Syst.5
2022 TherMa-MiCs: Thermal-Aware Scheduling for Fault-Tolerant Mixed-Criticality Systems
abstract
Multicore platforms are becoming the dominant trend in designing Mixed-Criticality Systems (MCSs), which integrate applications of different levels of criticality into the same platform. A well-known MCS is the dual-criticality system that is composed of low-criticality and high-criticality tasks. The availability of multiple cores on a single chip provides opportunities to employ fault-tolerant techniques, such as N-Modular Redundancy (NMR), to ensure the reliability of MCSs. However, applying fault-tolerant techniques will increase the power consumption on the chip, and thereby on-chip temperatures might increase beyond safe limits. To prevent thermal emergencies, urgent countermeasures, like Dynamic Voltage and Frequency Scaling (DVFS) or Dynamic Power Management (DPM) will be triggered to cool down the chip. Such countermeasures, however, might not only lead to suspending low-criticality tasks, but also it might lead to violating timing constraints of high-criticality tasks. In order to prevent such severe scenarios, it is indispensable to consider a temperature constraint within the scheduling process of fault-tolerant MCSs. Therefore, this paper presents, for the first time, a thermal-aware scheduling scheme for fault-tolerant MCSs, named TherMa-MiCs. In particular, TherMa-MiCs, satisfies the temperature constraint jointly with the timing constraints of the high-criticality tasks, while attempting to maximize the QoS of low-criticality tasks under the predefined constraints. At the same time, a reliability target is satisfied by employing the well-known N-Modular Redundancy (NMR) fault-tolerant technique. Experimental results show that our proposed scheme meets the temperature and timing constraints, while at the same time, improving the QoS of low-criticality tasks, with an average of 44%.
Sepideh Safari, Heba Khdr, Pourya Gohari-Nazari, Mohsen Ansari, Shaahin Hessabi, Jörg Henkel
IEEE Trans. Parallel Distributed Syst.6
2022 FN-CACTI: Advanced CACTI for FinFET and NC-FinFET Technologies
abstract
Cache memories are an indispensable component of many processor-based systems and contribute significantly to the overall area, power consumption, and delay. This leads to an important role played by modeling tools for estimating the area, power consumption, and access time of cache memories. However, existing modeling tools such as CACTI and its various extensions have been primarily designed using data from various projections. For the first time, we propose an entire flow for obtaining/calibrating the transistor characteristics from a commercial technology and use these characteristics within CACTI. We also improve the modeling approach to make them more fine-grained and follow recent manufacturing trends suitable for FinFET technology. Further, for the first time, we extend CACTI to support negative capacitance fin field effect transistor (NC-FinFET), an emerging technology depicting negative capacitance whose current and capacitive characteristics are very different compared to those of the FinFET. We use the proposed tool (FN-CACTI) to identify NC-FinFET-based caches to be significantly more energy-efficient than corresponding FinFET-based caches. We also study an application of FN-CACTI to determine optimal voltages corresponding to the lowest energy consumption for NC-FinFET and FinFET-based caches of various sizes.
Divya Praneetha Ravipati, Rajesh Kedia, Victor M. van Santen, Jörg Henkel, Preeti Ranjan Panda, Hussam Amrouch
IEEE Trans. Very Large Scale Integr. Syst.4
2021 Approximate Computing for ML: State-of-the-art, Challenges and Visions
abstract
In this paper, we present our state-of-the-art approximate techniques that cover the main pillars of approximate computing research. Our analysis considers both static and reconfigurable approximation techniques as well as operation-specific approximate components (e.g., multipliers) and generalized approximate highlevel synthesis approaches. As our application target, we discuss the improvements that such techniques bring on machine learning and neural networks. In addition to the conventionally analyzed performance and energy gains, we also evaluate the improvements that approximate computing brings in the operating temperature.
Georgios Zervakis 0001, Hassaan Saadat, Hussam Amrouch, Andreas Gerstlauer, Sri Parameswaran, Jörg Henkel
ASP-DAC6
2021 Multiple approximate instances in neural processing units for energy-efficient circuit synthesis: work-in-progress
abstract
We present an architectural approach toward energy-efficient synthesis of circuits used in neural processing units. Neural network applications are shown to tolerate varying operand precisions between different inputs, accuracy targets, their phases, and learning methods, without significantly impacting the classification accuracy. Using multiple instances of systolic arrays at different precisions, we show that significant energy gains are possible beyond the conventional approach, using the same circuit for all precisions.
Tanfer Alan, Jorge Castro-Godínez, Jörg Henkel
CASES3
2021 SmartBoost: Lightweight ML-Driven Boosting for Thermally-Constrained Many-Core Processors
abstract
Dynamic voltage and frequency scaling (DVFS)-based boosting is indispensable for optimizing the performance of thermally-constrained many-core processors. State-of-the-art techniques employ the voltage/frequency (V/D sensitivity of the performance of an application as a boosting metric. This paper demonstrates that this leads to suboptimal boosting decisions because the sensitivities of power and temperature also play a profound impact and need to be included within the optimization. Therefore, we introduce a novel boosting metric that integrates all relevant metrics: the application-dependent V/f sensitivities of performance and power, and the core-dependent sensitivity of the temperature. This new boosting metric is derived at run-time using machine learning via a neural network (NN) model, which accurately estimates the V/f sensitivities of performance and power of a priori unknown applications with diverse and time-varying characteristics. This new metric enables to build a smart, yet lightweight, boosting technique to maximize the performance under a temperature constraint. The experimental results demonstrate a 21 % average improvement of the system performance over the state-of-the-art at a negligible run-time overhead of 0.8 %.
Martin Rapp, Mohammed Bakr Sikal, Heba Khdr, Jörg Henkel
DAC4
2021 Control Variate Approximation for DNN Accelerators
abstract
In this work, we introduce a control variate approximation technique for low error approximate Deep Neural Network (DNN) accelerators. The control variate technique is used in Monte Carlo methods to achieve variance reduction. Our approach significantly decreases the induced error due to approximate multiplications in DNN inference, without requiring time-exhaustive retraining compared to state-of-the-art. Leveraging our control variate method, we use highly approximated multipliers to generate power-optimized DNN accelerators. Our experimental evaluation on six DNNs, for Cifar-10 and Cifar100 datasets, demonstrates that, compared to the accurate design, our control variate approximation achieves same performance and 24% power reduction for a merely 0.16% accuracy loss.
Georgios Zervakis 0001, Ourania Spantidi, Iraklis Anagnostopoulos, Hussam Amrouch, Jörg Henkel
DAC5
2021 TiVaPRoMi: Time-Varying Probabilistic Row-Hammer Mitigation
abstract
Row-Hammering is a challenge for computing systems that use DRAM. It can cause bit flips in a DRAM row by accessing its neighboring rows. Several mitigation techniques on memory controller level were already suggested. The techniques are in two categories: The first category uses static probabilities, which leads to a performance penalty due to a high number of extra row activations. The second category is based on so-called Tabled Counters, which have large hardware requirements and are mostly infeasible to implement. We introduce a novel Row-Hammer mitigation technique that uses time-varying probabilities combined with a relatively small history table. Our technique reduces the number of extra row activations compared to static probabilistic techniques and it demands less storage than Tabled Counters techniques. Compared to state of the art, our technique offers a good compromise that has 9× - 27× reduced storage requirement than Tabled Counters and 6× - 12× fewer activations than probabilistic techniques.
Hassan Nassar, Lars Bauer, Jörg Henkel
DATE3
2021 Long Short-Term Memory Neural Network-based Power Forecasting of Multi-Core Processors
abstract
We propose a novel technique to forecast the power consumption of processor cores at run-time. Power consumption varies strongly with different running applications and within their execution phases. Accurately forecasting future power changes is highly relevant for proactive power/thermal management. While forecasting power is straightforward for known or periodic workloads, the challenge for general unknown workloads at different voltage/frequency (v/n-levels is still unsolved. Our technique is based on a long short-term memory (LSTM) recurrent neural network (RNN) to forecast the average power consumption for both the next 1ms and 10ms periods. The runtime inputs for the LSTM RNN are current and past power information as well as performance counter readings. An LSTM RNN enables this forecasting due to its ability to preserve the history of power and performance counters. Our LSTM RNN needs to be trained only once at design-time while adapting during run-time to different system behavior through its internal memory. We demonstrate that our approach accurately forecasts power for unseen applications at different v/f-levels. The experimental results shows that the forecasts of our LSTM RNN provide 43% lower worst case error for the 1ms forecasts and 38% for the 10ms forecasts. comnared to the state of the art.
Mark Sagi, Martin Rapp, Heba Khdr, Yizhe Zhang 0005, Nael Fasfous, Nguyen Anh Vu Doan, Thomas Wild, Jörg Henkel, Andreas Herkersdorf
DATE8
2021 Reliability-Aware Quantization for Anti-Aging NPUs
Sami Salamin, Georgios Zervakis 0001, Ourania Spantidi, Iraklis Anagnostopoulos, Jörg Henkel, Hussam Amrouch
DATE5
2021 FeFET and NCFET for Future Neural Networks: Visions and Opportunities
abstract
The goal of this special session paper is to introduce and discuss different emerging technologies for logic circuitry and memory as well as new lightweight architectures for neural networks. We demonstrate how the ever-increasing complexity in Artificial Intelligent (AI) applications, resulting in an immense increase in the computational power, necessitates inevitably employing innovations starting from the underlying devices all the way up to the architectures. Two different promising emerging technologies will be presented: (i) Negative Capacitance Field-Effect Transistor (NCFET) as a new beyond-CMOS technology with advantages for offering low power and/or higher accuracy for neural network inference. (ii) Ferroelectric FET (FeFET) as a novel non-volatile, area-efficient and ultra-low power memory device. In addition, we demonstrate how Binarized Neural Networks (BNNs) offer a promising alternative for traditional Deep Neural Networks (DNNs) due to its lightweight hardware implementation. Finally, we present the challenges from combining FeFET-based NVM with NNs and summarize our perspectives for future NNs and the vital role that emerging technologies may play.
Mikail Yayla, Kuan-Hsun Chen, Georgios Zervakis 0001, Jörg Henkel, Jian-Jia Chen, Hussam Amrouch
DATE4
2021 LoopBreaker: Disabling Interconnects to Mitigate Voltage-Based Attacks in Multi-Tenant FPGAs
abstract
FPGAs are being offered in the cloud as accelerator resources that can be shared among multiple users (i.e. tenants). Recently, various approaches have shown that fault attacks launched from one tenant region to another are possible, leading to timing faults or crashes of the FPGA. It is, therefore, important that malicious tenants are limited in their ability to cause such security problems. So far, the existing countermeasures against such attacks check the configuration bitstreams before they are reconfigured. Such offline approaches have various practical limitations, e.g. they may force the tenants to unveil their design secrets. In this paper, we present LoopBreaker, a novel runtime solution that can disable the entire activity of a malicious tenant region, in order to rapidly stop a potential attack before it results in a crash (i.e. Denial-of-Service). We implemented and tested multiple attack types and found that realistic attacks demand at least 12–26 µs to be successful. A partial reconfiguration to overwrite the malicious tenant region demands 200 µs in our realworld implementation, which is too slow to prevent the attack from leading to a crash. Instead, our proposed LoopBreaker method only needs 1.5 µs to stop a malicious tenant, which makes it the first online approach that can successfully stop challenging voltage drop-based attacks from causing a crash.
Hassan Nassar, Hanna AlZughbi, Dennis Gnad, Lars Bauer, Mehdi Baradaran Tahoori, Jörg Henkel
ICCAD6
2021 Positive/Negative Approximate Multipliers for DNN Accelerators
abstract
Recent Deep Neural Networks (DNNs) manage to deliver superhuman accuracy levels on many AI tasks. DNN accelerators are becoming integral components of modern systems-on-chips. DNNs perform millions of arithmetic operations per inference and DNN accelerators integrate thousands of multiply-accumulate units leading to increased energy requirements. To lower the energy consumption of DNN accelerators, approximate computing principles are employed. However, complex DNNs can be increasingly sensitive to approximation. In this work, we present a dynamically configurable approximate multiplier that supports three operation modes, i.e., exact, positive error, and negative error. In addition, we propose a filter-oriented approximation method to map the weights to the appropriate modes of the approximate multiplier. Our mapping algorithm balances the positive with the negative errors due to the approximate multiplications, aiming at maximizing the energy reduction while minimizing the overall convolution error. We evaluate our approach on multiple DNNs and datasets against state-of-the-art approaches, where our method achieves 18.33% energy gains on average across 7 NNs on 4 different datasets for a maximum accuracy drop of only 1%.
Ourania Spantidi, Georgios Zervakis 0001, Iraklis Anagnostopoulos, Hussam Amrouch, Jörg Henkel
ICCAD5
2021 Reliability-Driven Voltage Optimization for NCFET-based SRAM Memory Banks
abstract
Negative Capacitance Field-Effect Transistors (NCFET) are promising significant power reductions while maintaining performance due to their internal voltage amplification. However, the addition of the ferroelectric layer also introduces a higher gate capacitance, which has to be charged and discharged resulting in higher power consumption. This results in trade-offs when employing NC-FinFET with respect to the thickness of the ferroelectric layer and their operating voltage on power, performance and reliability in circuits. This design-space is currently not explored, as existing research focused on a transistor-to-transistor comparison to show the superiority of NC-FinFET at the same voltage. In this work, we evaluate NC-FinFET employment in a full SRAM memory array (including write driver, sense amplifier, pre-charging, etc.) to obtain circuit delay, read and hold power and reliability metrics. This work shows, that solely evaluating SRAM cells results in inaccurate delay and power estimations compared to a full SRAM array. We explore iso-voltage and iso-performance NC-FinFET operation. Additionally, we explore two new operation modes: operating NC-FinFET within the same overall power consumption (iso-power) and operating at the same noise margins (iso-reliability). This exploration shows, for the first time, how ferroelectric layer thickness plays a role on reliability as a 4 nm layer features a 47% loss compared to FinFET. Lastly, we obtain the activity of a register file in a processor simulator to obtain the ultimate impact on power and energy consumption of employing NC-FinFET in a microprocessor.
Victor M. van Santen, Simon Thomann, Yogesh S. Chauchan, Jörg Henkel, Hussam Amrouch
VTS4
2021 Neural Network-Based Performance Prediction for Task Migration on S-NUCA Many-Cores
abstract
The performance of a task running on a many-core with distributed shared last-level cache (LLC) strongly depends on two parameters: the power budget needed to guarantee thermally-safe operation and the LLC latency. The task's thread-to-core mapping determines both the parameters and needs to make a trade-off because both cannot be simultaneously optimal. Arrival and departure of tasks on a many-core deployed in an open system can change its state significantly in terms of available cores and power budgets. Task migrations can thereupon be used as a tool to keep the many-core operating at peak performance. Furthermore, the relative impacts of power budget and LLC latency on a task's performance may change with its different execution phases mandating its migration on-the-fly. We propose the first run-time algorithmPCMigthat increases the performance of a many-core with distributed shared LLC by migrating tasks based on their phases and the many-core's state.PCMigis based on a model that predicts the performance impact of migrations. We propose a performance prediction model based on a lightweight neural network (NN). To serve as a reference, we also propose an analytical model of the many-core that operates on CPI stacks. We demonstrate an NN-based model achieves a higher prediction accuracy at a lower overhead than an analytical model.PCMigis based on the NN prediction model and results in an up to 7.3 percent increase in performance under a thermal constraint for mixed workloads compared to architecture-aware state-of-the-art (up to 20 percent increase for individual applications). This is achieved with a run-time overhead of less than 0.5 percent.
Martin Rapp, Anuj Pathania, Tulika Mitra, Jörg Henkel
IEEE Trans. Computers4
2021 Power-Efficient Heterogeneous Many-Core Design With NCFET Technology
abstract
Multi-/many-core, homogeneous or heterogeneous architectures, using the existing CMOS technology are inevitably approaching the limit of attainable power efficiency due to the fundamental limits in scaling. Negative Capacitance Field-Effect Transistor (NCFET) is rapidly emerging as an alternative technology that promises a multi-fold increase in the power efficiency of transistors, yet is compatible with the existing CMOS fabrication process. NCFET incorporates a ferroelectric (FE) layer within the transistor's gate stack, which exhibits a negative capacitance effect amplifying the internal voltage. NCFET has been in detail studied in both physics and devices/circuits communities where its superiority has been demonstrated in semiconductor measurements. However, the full promise of NCFET remains unmodeled and unquantified unless the research is further continued to the microarchitecture and system levels. This article, for the first time, explores system- and application-level benefits of NCFET-based multi-/many-core designs in terms of performance and power-efficiency compared to state-of-the-art FinFET-based designs. This exploration is done first through analytical modeling in which we extend Amdahl's law for NCFET multi-/many-cores, and then through quantitative modeling. The latter is achieved through RTL- and system-level simulations of NCFET-based multi-cores. The analytical modeling shows that a novel type of technology-based heterogeneity in which cores with the same microarchitecture but different FE thickness are combined is highly beneficial. Our exploration shows that this novel heterogeneity increases the power-efficiency by up to 3.5× over homogeneous systems and even achieves 8.3% better performance and 20% higher power-efficiency than conventional heterogeneity in the microarchitecture without having to cope with the complexity of managing different microarchitectures.
Sami Salamin, Martin Rapp, Anuj Pathania, Arka Maity, Jörg Henkel, Tulika Mitra, Hussam Amrouch
IEEE Trans. Computers5
2021 Post-Silicon Heat-Source Identification and Machine-Learning-Based Thermal Modeling Using Infrared Thermal Imaging
abstract
In this article, we present a novel post-silicon approach to locating the dominant heat sources on commercial multicore processors using heatmaps measured via an infrared (IR) thermal imaging setup. To locate the heat sources, 2-D spatial Laplacian transformation is performed on the heatmaps followed by K-means clustering to find the dominant power/heat-source clusters. This is an exclusively post-silicon approach that does not require any knowledge of the underlying design of the commercial chips other than the information that is publicly available. Since the identified clusters are the thermally vulnerable areas on the die, we then propose a machine-learning-based framework to deriving a thermal model capable of estimating their temperatures during online use. Our approach involves collecting transient temperature data of the aforementioned heat sources and synchronized high-level performance metrics from the chip, and training a long-short-term-memory (LSTM) neural network (NN) that uses the performance metrics as inputs to estimate the temperatures of the identified heat sources in real time. Since the model is meant for real-time use, we explore methods of reducing the performance overhead and inference time of the model. This includes a novel power correlation-based approach to identifying the thermally irrelevant performance metrics and eliminating them in order to reduce the input dimensionality of the model, and an analysis on network sizing to determine the ideal NN configuration for the problem at hand. The model is trained and tested exclusively using measured thermal data from commercial multicore processors. The experimental results from two Intel multicore processors (i5-3337U and i7-8650U) show that the proposed approach achieves very high accuracy (root-mean-square error: 0.55 °C-0.93 °C) in estimating the temperatures of all the identified heat sources on the chip.
Sheriff Sadiqbatcha, Hengyang Zhao, Hussam Amrouch, Jörg Henkel, Sheldon X.-D. Tan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2021 Automated Design Approximation to Overcome Circuit Aging
abstract
Transistor aging phenomena manifest themselves as degradations in the main electrical characteristics of transistors. Over time, they result in a significant increase of cell propagation delay, leading to errors due to timing violations, since the operating frequency becomes unsustainable as the circuit ages. Conventional techniques employ timing guardbands to mitigate aging-induced delay increase, which leads to considerable performance losses from the beginning of the circuit’s lifetime. Leveraging the inherent error resilience of a vast number of application domains, approximate computing was recently introduced as an aging mitigation mechanism. In this work, we present the first automated framework for generatingaging-aware approximate circuits. Our framework, by applying directed gate-level netlist approximation, induces a small functional error and recovers the delay degradation due to aging. As a result, our optimized circuits eliminate aging-induced timing errors. Experimental evaluation over a variety of arithmetic circuits and image processing benchmarks demonstrates that for an average error of merely$5\times 10^{-3}$, our framework completely eliminates aging-induced timing guardbands. Compared to the respective baseline circuits without timing guardbands (i.e., iso-performance evaluation), the error of the circuits generated by our framework is$1208\times $smaller.
Konstantinos Balaskas, Georgios Zervakis 0001, Hussam Amrouch, Jörg Henkel, Kostas Siozios
IEEE Trans. Circuits Syst. I Regul. Pap.4
2021 On the Resiliency of NCFET Circuits Against Voltage Over-Scaling
abstract
Approximate computing is established as a design alternative to improve the energy requirements of a vast number of applications, leveraging their intrinsic error tolerance. Voltage over-scaling (VOS) is one of the most energy-efficient approximation techniques, but its exploitation is still limited due to the large errors it induces. In this work, we investigate, for the first time, the resiliency of negative capacitance transistor (NCFET) technology to VOS in comparison to conventional CMOS technology. Our work reveals that circuits implemented using the NCFET technology exhibit much less timing errors under VOS due to the inherent voltage amplification provided by the ferroelectric layer. NCFET is one of the very promising emerging technologies that is rapidly evolving for low-power circuit as it enables the transistors to switch faster without the need to increase the voltage. We demonstrate how NCFET technology allows circuit designers to effectively employ VOS to boost the efficiency of their approximate circuits, while still keeping the induced errors marginal. Our analysis shows that the VOS-resilience of NCFET circuits enables maximizing the voltage decrease and thus, NCFET based VOS approximate circuits achieve from 1.83× up to 2.78× higher energy reduction compared to the corresponding FinFET circuits for the same error bounds.
Guilherme Paim, Georgios Zervakis 0001, Girish Pahwa, Yogesh Singh Chauhan, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel, Hussam Amrouch
IEEE Trans. Circuits Syst. I Regul. Pap.7
2021 PROTON: Post-Synthesis Ferroelectric Thickness Optimization for NCFET Circuits
abstract
For the first time, we demonstrate an optimization technique to synthesize circuits in the Negative Capacitance FET (NCFET) technology. NCFET is a rapidly emerging technology to replace the currently employed CMOS technology due to its profound ability to overcome the fundamental limit in scaling along with its full compatibility with the existing fabrication process. This is achieved by replacing the traditional transistor gate dielectric with a ferroelectric layer that manifests itself as a Negative Capacitance (NC), which magnifies the electric field. As a result, NCFET-based circuits can operate at a higher clock frequency without the need to increase the operating voltage. NC breaks one of the fundamental laws in physics in which the total capacitance of two capacitors connected in series becomes larger–instead of smaller in ordinary capacitors– than each of them. This could lead to sub-optimal netlists, suffering from significant increase in dynamic power and IR-drops. To suppress that, we employ the relation between delay decrease and capacitance increase of gates w.r.t ferroelectric thickness. Our technique takes an optimized netlist, obtained from commercial EDA tools, and then selectively determines the optimal ferroelectric thickness for each gate in the netlist, so that the maximum performance provided by NCFET is still achieved while the dynamic power is considerably decreased (45% on average),i.e., no trade-offs. Particularly, our technique enables the full exploitation of the performance benefits originating by NCFET, at a significantly lower (power) cost. Compared to state of the art, our technique decreases the energy-delay-product of circuits by 25% on average and reduces the deleterious effects of IR-drop by 56%. Hence, efficiency and reliability of circuits are improved without any loss in the obtained performance from NCFET.
Sami Salamin, Georgios Zervakis 0001, Yogesh Singh Chauhan, Jörg Henkel, Hussam Amrouch
IEEE Trans. Circuits Syst. I Regul. Pap.4
2021 Cross-Layer Approximate Hardware Synthesis for Runtime Configurable Accuracy
abstract
Approximate computing trades off computation accuracy against energy efficiency. The extent of approximation tolerance, however, significantly varies with a change in input characteristics and applications. We propose a novel cross-layer approach for the synthesis of runtime accuracy-configurable hardware that minimizes energy consumption at area expense. To that end, first, we explore instantiating multiple hardware blocks in the architecture with different fixed approximation levels. These blocks can be selected dynamically and thus allow to configure the accuracy during runtime. They benefit from having fewer transistors and also synthesis relaxations in contrast to state-of-the-art gating mechanisms that only switch off a group of paths of the circuit. Our cross-layer approach combines instantiating such blocks in the architecture with area-efficient gating mechanisms that reduce toggling activity, creating a fine-grained design-time knob on energy versus area. We present a systematic methodology to explore this joint design space and find energy-area optimal solutions as a function of required accuracies, their utilization in the workload, together with hardware parameters: dynamic power savings, area of the hardware block, and leakage of the technology. Examining total energy savings for a range of circuits under different workloads and accuracy tolerances shows that our method finds Pareto-optimal solutions providing up to 32% and 60% energy savings compared to state-of-the-art accuracy-configurable gating mechanism and an exact hardware block, respectively, at 2× area cost.
Tanfer Alan, Andreas Gerstlauer, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.3
2021 Automatic Floorplanning and Standalone Generation of Bitstream-Level IP Cores
abstract
Partially reconfigurable designs on field-programmable gate array (FPGA) bring an opportunity for developers to license third-party intellectual property (IP) cores. There are multiple IP licensing models that can be used by the FPGA IP market. Their focus is mainly on feasibility and security; however, two major challenges have been ignored by almost all of them. First, both academic or industrial tools do not provide a flow to generate IPs in a standalone environment. Second, these tools only offer manual floorplanning of the IPs, which is both time and performance inefficient. In this work, we present a framework, that can be used by multiple parties to generate different parts of a design independently, that are compatible with each other. It also provides automatic floorplanning based on mixed-integer linear programming (MILP) that considers the distribution of heterogeneous resources in modern FPGAs, with efficient resource utilization as the main objective. The proposed floorplanning is evaluated with benchmarks from the related work. Furthermore, a use case of internal and open-source designs is used for the validation and evaluation of the independent IP generation and the floorplanner.
Nadir Khan, Jorge Castro-Godínez, Shixiang Xue, Jörg Henkel, Jürgen Becker 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2020 NCFET to Rescue Technology Scaling: Opportunities and Challenges
abstract
Negative Capacitance Field Effect Transistor (NCFET) is one of the promising emerging technologies that may overcome the fundamental limits of conventional CMOS technology. NCFET features a ferroelectric (FE) layer within the transistor's gate, which internally amplifies the voltage, allowing NCFET to operate at a lower voltage while sustaining performance at considerable energy savings. In this work, we raise awareness that n- and p-NCFET transistors are asymmetrically affected by the FE layer and show, for the first time, how this asymmetry results in unbalanced circuit performance (e.g., longer fall than rise propagation delay, reduced noise margins). As NCFET are meant to maintain performance while reducing power, we present a solution by scaling the number of fins in n-NCFET to regain symmetry. We optimize iteratively in conjunction with supply voltage scaling to find the minimal energy consumption while maintaining performance. In our first case study, we achieve at least 34% lower power consumption and thus 34% higher energy efficiency as the circuit exhibits identical propagation delay. However, our second case study reveals that NCFETs can consume 3× more power and energy than the FinFET design. In summary, not considering the asymmetry and replacing FinFET with current-matched NCFET results in unreliable circuits (timing violations). This work exemplifies how the power and energy consumption of a NCFET circuit might surpass that of a FinFET, if circuits are designed considering asymmetry and circuit metric matching.
Hussam Amrouch, Victor M. van Santen, Girish Pahwa, Yogesh Singh Chauhan, Jörg Henkel
ASP-DAC5
2020 Machine Learning Based Online Full-Chip Heatmap Estimation
abstract
Runtime power and thermal control is crucial in any modern processor. However, these control schemes require accurate real-time temperature information, ideally of the entire die area, in order to be effective. On-chip temperature sensors alone cannot provide the full-chip temperature information since the number of sensors that are typically available is very limited due to their high area and power overheads. Furthermore, as we will demonstrate, the peak locations within hot-spots are not stationary and are very workload dependent, making it difficult to rely on fixed temperature sensors alone. Therefore, we propose a novel approach to real-time estimation of fullchip transient heatmaps for commercial processors based on machine learning. The model derived in this work supplements the temperature data sensed from the existing on-chip sensors, allowing for the development of more robust runtime power and thermal control schemes that can take advantage of the additional thermal information that is otherwise not available. The new approach involves offline acquisition of accurate spatial and temporal heatmaps using an infrared thermal imaging setup while nominal working conditions are maintained on the chip. To build the dynamic thermal model, we apply LongShort-Term-Memory (LSTM) neutral networks with system-level variables such as chip frequency, instruction counts, and other performance metrics as inputs. To reduce the dimensionality of the model, 2D spatial discrete cosine transformation (DCT) is first performed on the heatmaps so that they can be expressed with just their dominant DCT frequencies. Our study shows that only 6×6 DCT coefficients are required to maintain sufficient accuracy across a variety of workloads. Experimental results show that the proposed approach can estimate the full-chip heatmaps with less than 1.4°C root-mean-square-error and take only ~19ms for each inference which suits well for real-time use.
Sheriff Sadiqbatcha, Hussam Amrouch, Jörg Henkel, Sheldon X.-D. Tan
ASP-DAC5
2020 Impact of Self-Heating on Performance, Power and Reliability in FinFET Technology
abstract
Self-heating is one of the biggest threats to reliability in current and advanced CMOS technologies like FinFET and Nanowire, respectively. Encapsulating the channel with the gate dielectric improved electrostatics, but also thermally insulates the channel resulting in elevated channel temperatures as the generated heat is trapped within the channel. Elevated channel temperatures lowers the performance, increases leakage power and degrades the reliability of circuits. Self-heating becomes worse in each new transistor structure (from planar transistor to FinFET to Nanowire) due to the ever-increasing thermal resistance of the transistor. This leads to elevated temperatures, which must be carefully considered while designing circuits. Otherwise, reliability cannot be ensured. This work presents a self-heating study to illustrate how self-heating matters in digital circuits. It also explores the impact of running workloads in SRAM arrays, such as register files in CPUs, and how self-heating effects in SRAM cells can be mitigated.
Victor M. van Santen, Paul R. Genssler, Om Prakash 0007, Simon Thomann, Jörg Henkel, Hussam Amrouch
ASP-DAC5
2020 Towards Quality-Driven Approximate Software Generation for Accurate Hardware: Work-in-Progress
abstract
Many existing processor-based systems, especially using off-the-shelf components, cannot afford hardware modifications to embrace different approximate computing techniques proposed by the research community. In that case, it is mainly at the software level where error resiliency can be exploited efficiently. Although a multitude of approximate techniques can be applied at software level, they have been presented in isolation and little has been done to report their combined usage on different types of applications amenable to approximations. We present here AxSWGen, an automated quality-driven methodology to jointly explore and apply multiple approximate techniques to error-tolerant sections of applications. AxSWGen is implemented using LLVM compiler infrastructure. We present results of automated approximate software generated with AxSWGen and executed on a RISC-V processor (SiFive HiFive1 board), achieving up to 50% energy reduction for a 5% image degradation for an approximate Gaussian filter.
Jorge Castro-Godínez, Muhammad Shafique 0001, Jörg Henkel
CASES3
2020 Run-Time Accuracy Reconfigurable Stochastic Computing for Dynamic Reliability and Power Management: Work-in-Progress
abstract
In this paper, we propose a novel accuracy-reconfigurable stochastic computing (ARSC) framework for dynamic reliability and power management. Different than the existing stochastic computing works, where the accuracy versus power/energy trade-off is carried out in the design time, the new ARSC design can change accuracy or bit-width of the data in the run-time so that it can accommodate the long-term aging effects by slowing the system clock frequency at the cost of accuracy while maintaining the throughput of the computing. We validate the ARSC concept on a discrete cosine transformation (DCT) and inverse DCT designs for image compressing/decompressing applications, which are implemented on Xilinx Spartan-6 family XC6SLX45 platform. Experimental results show that the new design can easily mitigate the long-term aging induced effects by accuracy trade-off while maintaining the throughput of the whole computing process using simple frequency scaling. We further show that one-bit precision loss for the input data, which translated to 3.44dB of the accuracy loss in term of Peak Signal to Noise Ratio (PSNR) for images, we can sufficiently compensate the NBTI induced aging effects in 10 years while maintaining the pre-aging computing throughput of 7.19 frames per second. At the same time, we can save 74% power consumption by 10.67dB of accuracy loss. The proposed ARSC computing framework also allows much aggressive frequency scaling, which can lead to order of magnitude power savings compared to the traditional dynamic voltage and frequency scaling (DVFS) techniques.
Shuyuan Yu, Han Zhou 0002, Shaoyi Peng, Hussam Amrouch, Jörg Henkel, Sheldon X.-D. Tan
CASES5
2020 Runtime Accuracy-Configurable Approximate Hardware Synthesis Using Logic Gating and Relaxation
abstract
Approximate computing trades off computation accuracy against energy efficiency. Algorithms from several modern application domains such as decision making and computer vision are tolerant to approximations while still meeting their requirements. The extent of approximation tolerance, however, significantly varies with a change in input characteristics and applications.We propose a novel hybrid approach for the synthesis of runtime accuracy configurable hardware that minimizes energy consumption at area expense. To that end, first we explore instantiating multiple hardware blocks with different fixed approximation levels. These blocks can be selected dynamically and thus allow to configure the accuracy during runtime. They benefit from having fewer transistors and also synthesis relaxations in contrast to state-of-the-art gating mechanisms which only switch off a group of logic. Our hybrid approach combines instantiating such blocks with area-efficient gating mechanisms that reduce toggling activity, creating a fine-grained design-time knob on energy vs. area. Examining total energy savings for a Sobel Filter under different workloads and accuracy tolerances show that our method finds Pareto-optimal solutions providing up to 16% and 44% energy savings compared to state-of-the-art accuracy-configurable gating mechanism and an exact hardware block, respectively, at 2x area cost.
Tanfer Alan, Andreas Gerstlauer, Jörg Henkel
DATE3
2020 Impact of NBTI Aging on Self-Heating in Nanowire FET
abstract
This is the first work that investigates the impact of Negative Bias Temperature Instability (NBTI) on the Self-Heating (SH) phenomenon in Silicon Nanowire Field-Effect Transistors (SiNW-FETs). We investigate the individual as well as joint impact of NBTI and SH on pSiNW-FETs and demonstrate that NBTI-induced traps mitigate SH effects due to reduced current densities. Our Technology CAD (TCAD)-based SiNW-FET device is calibrated against experimental data. It accounts for thermodynamic and hydrodynamic effects in 3-D nano structures for accurate modeling of carrier transport mechanisms. Our analysis focuses on how lattice temperature, thermal resistance and thermal capacitance of pSiNW-FETs are affected due to NBTI, demonstrating that accurate self-heating modeling necessitates considering the effects that NBTI aging has over time. Hence, NBTI and SH effects need to be jointly and not individually modeled. Our evaluation shows that an individual modeling of NBTI and SH effects leads to a noticeable overestimation of the overall induced delay increase in circuits due to the impact of NBTI traps on SH mitigation. Hence, it is necessary to model NBTI and SH effects jointly in order to estimate efficient (i.e. small, yet sufficient) timing guardbands that protect circuits against timing violations, which will occur at runtime due to delay increases induced by aging and self-heating.
Om Prakash 0007, Hussam Amrouch, Sanjeev Manhas 0001, Jörg Henkel
DATE4
2020 Energy Optimization in NCFET-based Processors
abstract
Energy consumption is a key optimization goal for all modern processors. Negative Capacitance Field-Effect Transistors (NCFETs) are a leading emerging technology that promises outstanding performance in addition to better energy efficiency. Thickness of the additional ferroelectric layer, frequency, and voltage are the key parameters in NCFET technology that impact the power and frequency of processors. However, their joint impact on energy optimization has not been investigated yet.In this work, we are the first to demonstrate that conventional (i.e., NCFET-unaware) dynamic voltage/frequency scaling (DVFS) techniques to minimize energy are sub-optimal when applied to NCFET-based processors. We further demonstrate that state-of-the-art NCFET-aware voltage scaling for power minimization is also sub-optimal when it comes to energy. This work provides the first NCFET-aware DVFS technique that optimizes the processor's energy through optimal runtime frequency/voltage selection. In NCFETs, energy-optimal frequency and voltage are dependent on the workload and technology parameters. Our NCFET-aware DVFS technique considers these effects to perform optimal voltage/frequency selection at runtime depending on workload characteristics. Results show up to 90 % energy savings compared to conventional DVFS techniques. Compared to state-of-the-art NCFET-aware power management, our technique provides up to 72 % energy savings along with 3.7x higher performance.
Sami Salamin, Martin Rapp, Hussam Amrouch, Andreas Gerstlauer, Jörg Henkel
DATE5
2020 AxHLS: Design Space Exploration and High-Level Synthesis of Approximate Accelerators using Approximate Functional Units and Analytical Models
abstract
With the emergence of approximate computing as a design paradigm, many approximate functional units have been proposed, particularly approximate adders and multipliers. These circuits compromise the accuracy of their results within a tolerable limit to reduce the required computational effort and energy requirements. However, for an ongoing number of such approximate circuits reported in the literature, selecting those that minimize the required resources for designing and generating an approximate accelerator from a high-level specification, while satisfying a defined accuracy constraint, is a joint high-level synthesis (HLS) and design space exploration (DSE) challenge. In this paper, we propose a novel automated framework for HLS of approximate accelerators using a given library of approximate functional units. Since repetitive circuit synthesis and gate-level simulations require a significant amount of time, to enable our framework, we present AxME, a set of analytical models for estimating the required computational resources when using approximate adders and multipliers in approximate designs. We propose DSEwam, a DSE methodology for error-tolerant applications, in which analytical models, such as AxME, are used to estimate resources needed and the accuracy of approximate designs. Furthermore, we integrate DSEwam into an HLS tool to automatically generate Pareto-optimal, or near Pareto-optimal, approximate accelerators from C language descriptions, for a given error threshold and minimization goal. We release our DSE framework as an open-source contribution, which will significantly boost the research and development in the field of automatic generation of approximate accelerators.
Jorge Castro-Godínez, Julián Mateus-Vargas, Muhammad Shafique 0001, Jörg Henkel
ICCAD4
2020 Cell Library Characterization using Machine Learning for Design Technology Co-Optimization
abstract
To explore the full potential of any circuit and ensure its functionality at run-time, cell libraries beyond the typical PVT corners are needed. This holds even more for emerging technologies like Negative Capacitance (NC)-FinFET, where research in finding the optimal set of transistor parameters is still in its infancy. Design Technology Co-Optimization (DTCO) tackles bridging the large existing gap between device physics and the figures of merit of circuits. In this paper, we propose a Machine Learning (ML) approach to rapidly generate full cell libraries on demand. This enables the designer to perform extensive design space exploration and fully automated Design Technology Co-Optimization while lowering the barrier of accessibility. We demonstrate library prediction with an R2 score of around 98% for individual values and Static Timing Analysis (STA) reports. Experimental results show that our DTCO approach overestimates the achievable improvement by around 5%, nevertheless improving upon the baseline configuration.
Florian Klemme, Yogesh Singh Chauhan, Jörg Henkel, Hussam Amrouch
ICCAD3
2020 Modeling Emerging Technologies using Machine Learning: Challenges and Opportunities
abstract
Compact models of transistors act as the link between semiconductor technology and circuit design via circuit simulations. Unfortunately, compact model development and calibration is a challenging and time-intensive task, hindering rapid prototyping of a circuit (via circuit simulations) in emerging technologies. Moreover, foundries want to protect their confidential technology details to prevent reverse engineering. Hence, they limit access to compact transistor models of commercial technologies (e.g., with Non-Disclosure-Agreements). In this work, we propose Machine Learning (ML) to bridge the gap between early device measurements and later occurring compact model development. Our approach employs a Neural Network (NN) that captures the electrical response of a conventional FinFET transistor without knowledge of semiconductor physics. Additionally, our approach can be applied to emerging technologies, using Negative Capacitance FinFET (NC-FinFET) as an example for a (challenging to model) emerging technology. Inherently, the black-box nature of ML approaches keeps technology manufacturing details confidential. Furthermore, we show how using solely R2 score as our fitness function is insufficient and instead propose fitness based on key electrical characteristics or transistors like threshold voltage. Our NN-based transistor modeling can infer FinFET and NC-FinFET with an R2 score larger than 0.99 and transistor characteristics within 5% of experimental data.
Florian Klemme, Jannik Prinz, Victor M. van Santen, Jörg Henkel, Hussam Amrouch
ICCAD4
2020 Hierarchical Classification for Constrained IoT Devices: A Case Study on Human Activity Recognition
abstract
The massive number of Internet-of-Things (IoT) devices generates a hard-to-manage volume of data. Cloud-centric processing approaches for the IoT data suffer from high and unpredictable network latency, which causes poor experience in real-time IoT applications, such as healthcare. To address this issue, in edge computing, the data inference starts from the data source (i.e., the IoT devices). However, the constrained computational capabilities of the IoT device and the power-hungry data transmission demand a tradeoff between onboard processing and computation offloading. Hence, the IoT information inference requires efficient and lightweight techniques that are tailored for this tradeoff and respect the constrained resources on IoT devices, such as wearables. This article presents a hierarchical classification approach that decomposes the problem into three classifiers in two hierarchy layers. In the first layer, a lightweight classifier executes directly on the IoT device and decides whether to offload the computation to the gateway or to perform it onboard. The second layer comprises a lightweight classifier on the IoT device (can only distinguish a subset of classes) and a complex classifier on the gateway (to distinguish the remaining classes). The experimental results (using a real-world data set for human activity recognition and implemented on a wearable IoT device) show higher accuracy (92% on average) than a nonhierarchical classifier (87% on average). The execution time and power measurements on the IoT device show $3\times $ energy saving for the classification.
Farzad Samie, Lars Bauer, Jörg Henkel
IEEE Internet Things J.3
2020 Power- and Cache-Aware Task Mapping with Dynamic Power Budgeting for Many-Cores
abstract
Two factors primarily affect the performance of multi-threaded tasks on many-core processors with logically-shared and physically-distributed Last-Level Cache (LLC): the LLC latencies of threads running on different cores and the per-core power budgets that aim to guarantee thermally safe operation. Two knobs affect these factors: First, the mapping of threads to cores affects both the LLC latencies and the power budgets. Second, dynamic power budgeting refines the power budgets during task execution. A mapping that spatially distributes threads across the many-core increases the power budgets, but unfortunately also increases the LLC latencies. Contrarily, mapping all threads near the center of the many-core minimizes the LLC latencies, but unfortunately also decreases the power budgets. Consequently, both metrics cannot be simultaneously optimal, which leads to a Pareto-optimization for task mapping that has formerly not been exploited. Dynamic power budgeting reallocates the power budgets according to the tasks' execution phases. This results in a dynamically changing non-uniform power budget, which further increases the performance. We are the first to present a run-time algorithm PCGov combining task-agnostic task mapping and task-aware dynamic power budgeting for many-cores with shared distributed LLC. PCGov yields up to 21 percent lower response time and 13 percent lower energy consumption compared to the state-of-the-art, with a low overhead of less than 0.5 percent.
Martin Rapp, Mark Sagi, Anuj Pathania, Andreas Herkersdorf, Jörg Henkel
IEEE Trans. Computers5
2020 NPU Thermal Management
abstract
Neural processing units (NPUs) are becoming an integral part in all modern computing systems due to their substantial role in accelerating neural networks (NNs). The significant improvements in cost-energy-performance stem from the massive array of multiply accumulate (MAC) units that remarkably boosts the throughput of NN inference. In this work, we are the first to investigate the thermal challenges that NPUs bring, revealing how MAC arrays, which form the heart of any NPU, impose serious thermal bottlenecks to on-chip systems due to their excessive power densities. For the first time, we explore: 1) the effectiveness of precision scaling and frequency scaling (FS) in temperature reductions and 2) how advanced on-chip cooling using superlattice thin-film thermoelectric (TE) open doors for new tradeoffs between temperature, throughput, cooling cost, and inference accuracy in NPU chips. Our work unveils that hybrid thermal management, which composes different means to reduce the NPU temperature, is a key. To achieve that, we propose and implement PFS-TE technique that couples precision and FS together with superlattice TE cooling for effective NPU thermal management. Using commercial signoff tools, we obtain accurate power and timing analysis of MAC arrays after a full-chip design is performed based on 14-nm Intel FinFET technology. Then, multiphysics simulations using finite-element methods are carried out for accurate heat simulations in the presence and absence of on-chip cooling. Afterward, comprehensive design-space exploration is presented to demonstrate the Pareto frontier and the existing tradeoffs between temperature reductions, power overheads due to cooling, throughput, and inference accuracy. Using a wide range of NNs trained for image classification, experimental results demonstrate that our novel NPU thermal management increases the inference efficiency (TOPS/Joule) by 1.33×, 1.87×, and 2× under different temperature constraints; 105 °C, 85 °C, and 70 °C, respectively, while the average accuracy drops merely from 89.0% to 85.5%.
Hussam Amrouch, Georgios Zervakis 0001, Sami Salamin, Hammam Kattan, Iraklis Anagnostopoulos, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2020 Toward Model Checking-Driven Fair Comparison of Dynamic Thermal Management Techniques Under Multithreaded Workloads
abstract
Dynamic thermal management (DTM) techniques are being widely used for attenuation of thermal hot spots in many-core systems. Conventionally, DTM techniques are analyzed using simulation and emulation methods, which are in-exhaustive due to their inherent limitations and cannot provide for a comprehensive comparison between DTM techniques owing to the wide range of corresponding design parameters. In order to handle the above discrepancies, we propose to use model checking, a state-space-based formal method, to model, evaluate, and compare DTM techniques across various functional and performance parameters. The suggested framework includes a modeling flow and a set of generic modules that realistically model many-core and DTM parameters like temperature, power, application, intercore communication and task migration, etc. For analysis purpose, the framework provides a common ground for comparing DTM techniques by formalizing DTM principles and performance parameters as a set of logical properties. These properties are verified for different task load configurations, e.g., multithreaded, malleable, and the applications which do not support migration. We analyze state-of-the-art central (c-) and distributed (d-) DTM techniques to demonstrate the generality and efficacy of our approach. Our formal analysis shows that the state-of-the-art cDTM technique performs better than dDTM in terms of achieving thermal stability, task migration, and communication overhead. We believe that conventional analysis methods do not facilitate such an exhaustive comparison among the DTM techniques.
Syed Ali Asadullah Bukhari, Faiq Khalid, Osman Hasan, Muhammad Shafique 0001, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2020 Machine Learning for Power, Energy, and Thermal Management on Multicore Processors: A Survey
abstract
Due to the high integration density and roadblock of voltage scaling, modern multicore processors experience higher power densities than previous technology scaling nodes. When unattended, this issue might lead to temperature hot spots, that in turn may cause nonuniform aging, accelerate chip failure, impair reliability, and reduce the performance of the system. This paper presents an overview of several research efforts that propose to use machine learning (ML) techniques for power and thermal management on single-core and multicore processors. Traditional power and thermal management techniques rely on a certain a-priori knowledge of the chip's thermal model, as well as information of the workloads/applications to be executed (e.g., transient and average power consumption). Nevertheless, these a-priori information is not always available, and even if it is, it cannot reflect the spatial and temporal uncertainties and variations that come from the environment, the hardware, or from the workloads/applications. Contrarily, techniques based on ML can potentially adapt to varying system conditions and workloads, learning from past events in order to improve themselves as the environment changes, resulting in improved management decisions.
Santiago Pagani, Sai Manoj Pudukotai Dinakarrao, Axel Jantsch, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 A Lightweight Nonlinear Methodology to Accurately Model Multicore Processor Power
abstract
Many power management algorithms demand accurate and fine-grained runtime estimations of dynamic core power. In the absence of fine-grained power sensors, model-based estimations are needed. Such power models commonly approximate the switching activity of logic gates using performance counters while assuming a linear performance counter/power relation at a fixed frequency and voltage. It has been shown that this relation cannot be captured accurately enough with purely linear models and that well-established nonlinear modeling techniques, e.g., polynomial modeling, easily overfit the underlying performance/power relations. Although neural-network-based modeling has shown to accurately capture nonlinear relations, it has a large training and inference overhead which is too high for fine-grained models on core-level and estimation rates in the range of 1-10 kHz. We propose a methodology for nonlinear transformation of specific performance counters to increase power modeling accuracy at constant frequency and voltage with a relatively low overhead for both model generation and run-time application over a linear model. Furthermore, we use least-angle regression (LARS) to determine a ranking of the performance counter inputs for use in linear and nonlinear modeling and show that the transformed performance counters are better suited for power modeling. The generated dynamic power model consisting of a nonlinear transformation block and a linear regression block reduces relative estimation error on average by 4% and in worst-case scenarios by 7% compared to state-of-the-art fine-grained linear power models. Compared to a state-of-the-art polynomial regression model our proposed approach reduces the relative estimation error by 10% in worst-case scenarios.
Mark Sagi, Nguyen Anh Vu Doan, Martin Rapp, Thomas Wild, Jörg Henkel, Andreas Herkersdorf
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2020 Dynamic Power and Energy Management for NCFET-Based Processors
abstract
Power and energy consumption are the key optimization goals in all modern processors. Negative capacitance field-effect transistors (NCFETs) are a leading emerging technology that promises outstanding performance in addition to better energy efficiency. The thickness of the added ferroelectric layer as well as frequency and voltage are the key parameters that impact the power and energy of NCFET-based processors in addition to the characteristics of runtime workloads. Unlike existing CMOS technologies, operating NCFET-based processors at a higher frequency than the required minimum can result in power/energy minimization. The optimal operating point, however, strongly depends on dynamic workload characteristics and technology parameters. In this work, we propose and implement the first NCFET-aware power and energy management approach that minimizes the processor's power and energy through optimal voltage/frequency selection under different runtime scenarios. Such an NCFET-aware approach does not result in any tradeoff between power/energy and performance. Instead, it can achieve higher performance while minimizing energy. A comprehensive, simulation-based evaluation of our runtime management under realistic workloads demonstrates up to 58% energy saving with 2.1× higher performance, and 46% power saving compared to conventional NCFET-unaware management techniques, over the total execution of a benchmark. Compared to state-of-the-art NCFET-aware management techniques, our technique provides up to 49% energy saving and 32% power saving.
Sami Salamin, Martin Rapp, Jörg Henkel, Andreas Gerstlauer, Hussam Amrouch
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Fast Operation Mode Selection for Highly Efficient IoT Edge Devices
abstract
In the emerging paradigm of edge computing (EC) for Internet of Things (IoT), data processing is pushed to the edge of the IoT network (e.g., gateways and embedded IoT devices). IoT devices must support multiple operation modes in order to adapt to varying runtime situations, like preserving energy at low battery, while still maintaining some crucial functionality, etc. Adapting the optimal operation mode is a challenge for edge devices given the limited resources at the edge of the network (both bandwidth and processing power of the shared gateway), various constraints (e.g., battery lifetime), etc. This paper proposes a fast and low-overhead scheme to determine and adapt the operation mode of edge devices at runtime and orchestrate devices in a way that the efficiency of IoT devices is optimized with respect to the gateway's resource constraints. The proposed scheme breaks the optimization problem into several smaller ones (i.e., subproblems) whose solutions are aggregated to find the final solution. We present a novel memoization technique that determines the solution to a range of subproblems based on subproblems that are already solved. In addition, we present a novel pruning technique that reduces the search space and consequently reduces both memory and execution time overhead. The experimental results show up to 50% reduction in memory overhead and 14× reduction in execution time overhead compared to the state-of-the-art solution which is a major step toward efficient EC for IoT.
Farzad Samie, Vasileios Tsoutsouras, Dimosthenis Masouros, Lars Bauer, Dimitrios Soudris, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2020 Exposing Hardware Trojans in Embedded Platforms via Short-Term Aging
abstract
We demonstrate a novel technique that employs transistor short-term aging effects in integrated circuits (ICs) to detect hardware Trojans in embedded systems. In advanced technology nodes (≤ 45 nm), voltage scaling in combination with short-term aging opens doors for short-term degradations. The induced short-term degradations result in dynamic variation of delays along various paths within the IC. Aging degradation generated under fast voltage switching from high to low results in bit errors at the circuit output. Our experiments use short-term aging-aware standard cell libraries to show the effectiveness of short-term aging to detect hardware Trojans. We extract a rich set of features that capture bit error patterns at the outputs of the IC. We use a one class SVM-based classifier that uses these features to learn the distribution of bit errors at the outputs of a clean IC. We discern the deviation in the pattern of bit errors due to a Trojan in the IC from the baseline distribution. To reiterate, the method uses the model of a clean IC. Furthermore, it is robust against chip-to-chip variations. We illustrate the technique on six Trojans from Trust-Hub spanning two cryptographic chips and an embedded PIC microcontroller. Our approach detects Trojans with an accuracy ≥ 95%. It is easier to detect Trojans in an optimized-netlist circuit as more paths are close to the critical path. Even when the circuit is not optimized (i.e., when very few paths are close to the critical path), short-term aging plus mild overclocking can detect Trojans with high accuracy.
Virinchi Roy Surabhi, Prashanth Krishnamurthy, Hussam Amrouch, Jörg Henkel, Ramesh Karri, Farshad Khorrami
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 A Cross-Layer Gate-Level-to-Application Co-Simulation for Design Space Exploration of Approximate Circuits in HEVC Video Encoders
abstract
A cross-layer design space exploration (DSE) method based on a proposed co-simulation technique is presented herein. The proposed method is demonstrated evaluating the impacts on both coding efficiency and power dissipation of applying distinct approximate logic operators in a sum of absolute differences (SAD) kernel that accelerates an H.265/HEVC (high-efficiency video coding) encoder. The proposed method simulates the gate-level circuit dynamically inside the application, with realistic results of the impact of the adder-tree approximate logic implementation on both quality and encoder bit-rate results. A comprehensive DSE is shown herein, with 13 types of 6 classes of approximate adders in the SAD accelerator hardware blocks. Over 3,000 logic variants of approximations at gate-level were developed. Actual video sequences as inputs to the x265 software encoder are co-simulated, to dynamically capture the video motion-estimation (ME) behavior in the presence of logic approximations. While the prior art that only estimates the impact of the approximate logic on power, area, and quality on static designs with statistical assumptions, which are agnostic to the actual algorithm data-dependent behavior in the application, our method explores accurately the trade-off between power dissipation and coding efficiency dynamically over the entire HEVC encoding. Our approach shows that the lower-part-or and error-tolerant adder I approximate adders, as well as truncation-to-zero deliver better compression-power trade-offs, with substantial differences from the static analysis.
Guilherme Paim, Leandro M. G. Rocha, Hussam Amrouch, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel
IEEE Trans. Circuits Syst. Video Technol.6
2020 Introduction to the Special Issue on Machine Learning for CAD
abstract
No abstract available.
Jörg Henkel, Hussam Amrouch, Marilyn Wolf
ACM Trans. Design Autom. Electr. Syst.1
2020 Combinatorial Auctions for Temperature-Constrained Resource Management in Manycores
abstract
Although manycore processors have plenty of cores, not all of them may run simultaneously at full speed and even some of them might need to be power-gated in order to keep the chip within safe temperature limits. Hence, a resource management technique, that allocates cores to application aiming at maximizing the system performance, will not be able to achieve its goal without taking into account the on-chip temperature and its impact on the availability of the chip's resources. However, considering a temperature constraint by the resource management will further increase its complexity, especially in manycores, and thus implementing it in a centralized scheme might lead to a computation bottleneck and a single point of failure. To avoid such scenarios, it is inevitable to distribute the computation required by the resource management technique throughout the chip. In this article, we propose a distributed resource management technique that considers temperature as an essential factor in allocating cores to applications and determining the power states of these cores and their voltage/frequency levels, while taking into account the performance models of the applications in order to maximize the overall system performance under a temperature constraint. Our proposed technique employs, for the first time, combinatorial auctions within an agent system to achieve the targeted goal in a distributed manner. The experimental evaluations show that our proposed technique achieves significant performance improvements with an average of 41% compared to several distributed resource management techniques.
Heba Khdr, Muhammad Shafique 0001, Santiago Pagani, Andreas Herkersdorf, Jörg Henkel
IEEE Trans. Parallel Distributed Syst.5
2019 Performance, Power and Cooling Trade-Offs with NCFET-based Many-Cores
abstract
Negative Capacitance Field-Effect Transistor (NCFET) is an emerging technology that incorporates a ferroelectric layer within the transistor gate stack to overcome the fundamental limit of sub-threshold swing in transistors. Even though physics-based NCFET models have been recently proposed, system-level NCFET models do not exist and research is still in its infancy. In this work, we are the first to investigate the impact of NCFET on performance, energy and cooling costs in many-core processors. Our proposed methodology starts from accurate physics models all the way up to the system level, where the performance and power of a many-core are widely affected. Our new methodology and system-level models allow, for the first time, the exploration of the novel trade-offs between performance gains and power losses that NCFET now offers to system-level designers. We demonstrate that an optimal ferroelectric thickness does exist. In addition, we reveal that current state-of-the-art power management techniques fail when NCFET (with a thick ferroelectric layer) comes into play.
Martin Rapp, Sami Salamin, Hussam Amrouch, Girish Pahwa, Yogesh Singh Chauhan, Jörg Henkel
DAC6
2019 Smart Thermal Management for Heterogeneous Multicores
abstract
Heterogeneous multicores have attracted a major focus in recent years, as they provide many possibilities for performance improvements. However, due to the discontinuation of Dennard scaling, on-chip power densities are continuously increasing along with technology scaling, and hence on-chip temperatures are elevated. Therefore, several thermal management techniques have emerged to keep the temperature of the chip within safe limits. These techniques, however, lead to performance losses which ultimately erase a big portion of the expected performance gains from the heterogeneous multicores. Thus, it is indispensable to deploy thermal management techniques that are able to make efficient decisions which satisfy temperature constraints while at the same time maximizing the performance. This paper presents smart thermal management techniques for heterogeneous multicores that exploit relevant information about several heterogeneity parameters at the chip level and at the application level to increase thermal efficiency1. Compared to the state of the art, the presented techniques are able to obtain significant performance improvements under the same thermal constraint.This paper is part of the DATE 2019 special session on "Smart Resource Management and Design Space Exploration for Heterogeneous Processors". The other two papers of this special session are [1] and [2].
Jörg Henkel, Heba Khdr, Martin Rapp
DATE1
2019 A Fine-Grained Soft Error Resilient Architecture under Power Considerations
abstract
Besides the limited power budgets and the dark-silicon issue, soft error is one of the most critical reliability issues in computing systems fabricated using nano-scale devices. During the execution, different applications have varying performance, power/energy consumption and vulnerability properties. Different trade-offs can be devised to provide required resiliency within the allowed power constraints. To exploit this behavior, we propose a novel soft error resilient architecture and the corresponding run-time system that enables power-aware fine-grained resiliency for different processor components. It selectively determines the reliability state of various components, such that the overall application reliability is improved under a given power budget. Our architecture saves power up to 16% and reliability degradation up to 11% compared to state-of-the-art techniques.
Muhammad Shafique 0001, Jörg Henkel
DATE3
2019 Thermal-Awareness in a Soft Error Tolerant Architecture
abstract
It is crucial to provide soft error reliability in a power-efficient manner such that the maximum chip temperature remains within the safe operating limits. Different execution phases of an application have diverse performance, power, temperature and vulnerability behavior that can be leveraged to fulfill the resiliency requirements within the allowed thermal constraints. We propose a soft error tolerant architecture with fine-grained redundancy for different architectural components, such that their reliable operations can be activated selectively at fine-granularity to maximize the reliability under a given thermal constraint. When compared with state-of-the-art, our temperature-aware fine-grained reliability manager provides up to 30% reliability within the thermal budget.
Muhammad Shafique 0001, Jörg Henkel
DATE3
2019 DMRM: Distributed Market-Based Resource Management of Edge Computing Systems
abstract
Resource management is a key technique for efficiently operating devices in Internet of Things (IoT). In this paper, we propose DMRM, a new algorithm based on economic and pricing models for dynamic resource management of IoT networks under CPU, memory, bandwidth and latency constraints. We use a supply and demand model, smart data pricing and perceived valued pricing, implementing a marketplace where IoT devices and Gateways buy and sell computing and communication resources necessary for task execution. Our new market-based algorithm is compared to relevant approaches showing that it not only reaches near-optimal results, but also, its scalable, distributed nature leads to three orders of magnitude lower execution requirements compared to centralized approaches.
Manolis Katsaragakis, Dimosthenis Masouros, Vasileios Tsoutsouras, Farzad Samie, Lars Bauer, Jörg Henkel, Dimitrios Soudris
DATE6
2019 Prediction-Based Task Migration on S-NUCA Many-Cores
abstract
Performance of a task running on a many-core with distributed shared Last-Level Cache (LLC) strongly depends on two factors: the power budget needed to guarantee thermally safe operation and the LLC latency. The task's thread-to-core mapping determines both the factors. Arrival and departure of tasks on a many-core deployed in an open system can change its state significantly in terms of available cores and power budget. Task migrations can thereupon be used as a tool to keep the many-core operating at the peak performance. Furthermore, the relative impacts of power budget and LLC latency on a task's performance can change with its different execution phases mandating its migration on-the-fly.We propose the first run-time algorithm PCMig that increases the performance of a many-core with distributed shared LLC by migrating tasks based on their phases and the many-core's state. PCMig is based on a performance-prediction model that predicts the performance impact of migrations. PCMig results in up to 16 % reduction in the average response time compared to the state-of-the-art.
Martin Rapp, Anuj Pathania, Tulika Mitra, Jörg Henkel
DATE4
2019 Hot Spot Identification and System Parameterized Thermal Modeling for Multi-Core Processors Through Infrared Thermal Imaging
abstract
Accurate thermal models suitable for system level dynamic thermal, power and reliability regulation and management are vital for many commercial multi-core processors. However, developing such accurate thermal models and identifying the related thermal-power relevant spatial locations for commercial processors is a challenging task due to the lack of information and available tools. Existing tools such as HotSpot-like thermal models may suffer from inaccuracy or inefficiency for online applications, primarily because most rely on parameters that cannot be precisely quantified, such as power-traces, while others are numerical methods not suitable for runtime use. In this work, we propose a novel approach to automatically detecting the major heat-sources on a commercial multi-core microprocessor using an infrared thermal imaging setup. Our approach involves a number of steps including 2D discrete cosine transformation filter for noise reduction on the measured thermal maps, and Laplacian transformation followed by K-mean clustering for heat-source identification. Since the identified heat-sources are the thermally vulnerable areas of the die, we propose a novel approach to deriving a thermal model capable of predicting their temperatures during runtime. We apply Long-Short-Term-Memory (LSTM) networks to build a dynamic thermal model which uses system-level variables such as chip frequency, voltage and instruction count as inputs. The model is trained and tested exclusively using measured thermal data from a commercial multi-core processor. Experimental results show that the proposed thermal model achieves very high accuracy (root-mean-square-error: 2.04°C to 2.57° C) in predicting the temperature of all the identified heat-sources on the chip.
Sheriff Sadiqbatcha, Hengyang Zhao, Hussam Amrouch, Jörg Henkel, Sheldon X.-D. Tan
DATE4
2019 Selecting the Optimal Energy Point in Near-Threshold Computing
abstract
Near-Threshold Computing (NTC) has recently emerged as an attractive paradigm as it allows devices to operate close to their optimal energy point (OEP). This work demonstrates, for the first time, that determining where the OEP of a processor exists is challenging because standard cells, forming the processor's netlist, unevenly profit w.r.t power and also unevenly degrade w.r.t delay when the voltage approaches the near-threshold region. To precisely explore, at design time, where OEP is, we create voltage-aware cell libraries that enable designers to seamlessly employ the standard tool flows, even they were not designed for that purpose, to perform voltage-aware timing and power analysis. Besides determining where the OEP is, we also demonstrate how providing logic synthesis tool flows with voltage-aware cell libraries results in a 35% higher performance at NTC. In addition, we investigate how the performance loss at NTC can be compensated through parallelized computing demonstrating, for the first time, that the OEP moves far from NTC as the number of cores increases. Our proposed methodology enables designers to select the maximum number of cores along with the optimal operating voltage jointly in which a specific power budget is fulfilled. Finally, we show how voltage-aware design for parallelized NTC provides [40%-50%] performance increase compared to traditional (i.e., voltage-unaware design) parallelized NTC.
Sami Salamin, Hussam Amrouch, Jörg Henkel
DATE3
2019 WCET Guarantees for Opportunistic Runtime Reconfiguration
abstract
Time-critical systems need to be analyzable for timing guarantees. There is an increasing demand for predictable performance that modern processor architectures fail to provide since they focus on average-case performance only. Recent work has demonstrated that runtime reconfiguration of hardware accelerators via an FPGA is a viable way to achieve high performance for optimized worst-case execution time (WCET) guarantees. Since execution of the worst-case path is highly improbable, configuring accelerators for this path costs reconfigurable area that could better be used to accelerate more probable paths. This work presents the first approach that comprises (1) an online average-case execution time (ACET) optimization while (2) maintaining the optimized WCET guarantee utilizing reconfigurable accelerators. We achieve this by a new design-time technique which determines the runtime slack bounds that allow speculative reconfiguration of accelerators that benefit the ACET. Combined with an online slack monitoring approach that introduces negligible overheads by using a performance counter, we show a runtime reduction of up to 10.4% for a complex and real-world application on top of an already-optimized WCET guarantee.
Marvin Damschen, Lars Bauer, Jörg Henkel
ICCAD3
2019 The Impact of Emerging Technologies on Architectures and System-level Management: Invited Paper
abstract
The goal of this work is to introduce and discuss different kinds of emerging technologies for logic circuitry and memory with respect to the key question of how they will impact future system-on-chip architectures and system-level management techniques. It is obvious that emerging technologies should have an impact there in order to fully exploit their technological advantages but also in order to deal with any disadvantages they might come with. In this special session paper, three promising emerging technologies are presented: (i) Negative Capacitance Field-Effect Transistor (NCFET) as a new CMOS technology with advantages primarily for low-power design, (ii) Ferroelectric FET (FeFET) as a non-volatile, area-efficient and low-power combined logic and memory as well as (iii) a Phase-Change Memory (PCM) and Resistive RAM (ReRAM) offering a large potential for tackling the memory wall problem in the von Neumann architecture. Our analysis demonstrates that not only new computing paradigms are promoted by these new technologies, it will also be seen that the trade-offs between the classical design parameters of low power, performance etc. will shift and hence emerging technologies will offer new Pareto points in the design space of future on-chip architectures. In that context, this work is unique as it bridges the gap between the technology side and system/architecture-level side to draw a vision of new technologies and their impact on architectures and system-level management.
Jörg Henkel, Hussam Amrouch, Martin Rapp, Sami Salamin, Dayane Reis, Xunzhao Yin, Michael T. Niemier, Cheng Zhuo, Xiaobo Sharon Hu, Hsiang-Yun Cheng, Chia-Lin Yang
ICCAD1
2019 Reliability Challenges with Self-Heating and Aging in FinFET Technology
abstract
The introduction of FinFET technology as an effective solution to continue technology scaling has pushed self-heating effects to the forefront of reliability challenges, especially at the 14nm technology node and below. Due to limited silicon volume for heat dissipation, elevated temperatures across the transistors channel can be generated during operation. This results in a considerable degradation of the key properties of transistors like decreased drain and increased leakage current. In addition, excessive temperatures considerably accelerate aging phenomena in transistors such as Bias Temperature Instability (BTI) and Hot Carrier Injection (HCI), which shorten the lifetime of circuits. In this work, we discuss how self-heating effects in FinFET transistors can prolong the delay of circuits leading to reliability problems. We evaluate self-heating in an entire SRAM block consisting of SRAM cells, pre-charging circuit, sense amplifiers and an output latch. When it comes to reliability and lifetime, we demonstrate how self-heating effects can result in larger aging-induced degradations which, in turn, enforce designers to include wider and wider safety margins to sustain reliability. Lastly, we provide an outlook of self-heating and reliability concerns in Negative Capacitance Field Effect Transistors (NCFET).
Hussam Amrouch, Victor M. van Santen, Om Prakash 0007, Hammam Kattan, Sami Salamin, Simon Thomann, Jörg Henkel
IOLTS7
2019 Aging Gracefully with Approximation
abstract
This paper presents a design methodology to turn aging-induced chip slowdown into approximation without adding reliability guardband or increasing supply voltage. It guarantees always-best quality while the system is under aging. It is based on run-time monitoring of critical path delay. If the delay increases due to aging, the proposed approach curtails the critical path at the cost of precision reduction. We evaluate our approach at the component level as well as microarchitecture level. The evaluation results show that the approach reduces the dynamic and static power consumptions by 19.8% and 10.2%, respectively, with minimal area overhead and quality degradation.
Heesu Kim, Hussam Amrouch, Jörg Henkel, Andreas Gerstlauer, Kiyoung Choi
ISCAS4
2019 NCFET-Aware Voltage Scaling
abstract
Negative Capacitance Field-Effect Transistor (NCFET) has recently attracted significant attention. In the NCFET technology with a thick ferroelectric layer, voltage reduction increases the leakage power, rather than decreases, due to the negative Drain-Induced Barrier Lowering (DIBL) effect. This work is the first to demonstrate the far-reaching consequences of such an inverse dependency w.r.t. the existing power management techniques. Moreover, this work is the first to demonstrate that state-of-the-art Dynamic Voltage Scaling (DVS) techniques are sub-optimal for NCFET. Our investigation revealed that the optimal voltage at which the total power is minimized is not necessarily at the point of the minimum voltage required to fulfill the performance constraint (as in traditional DVS). Hence, an NCFET-aware DVS is key for high energy efficiency. In this work, we therefore propose the first NCFET-aware DVS technique that selects the optimal voltage to minimize the power following the dynamics of workloads. Our experimental results of a multi-core system demonstrate that NCFET-aware DVS results in 20% on average, and up to 27% energy saving while still fulfilling the same performance constraint (i.e., no trade-offs) compared to traditional NCFET-unaware DVS techniques.
Sami Salamin, Martin Rapp, Hussam Amrouch, Girish Pahwa, Yogesh Singh Chauhan, Jörg Henkel
ISLPED6
2019 Thermally Composable Hybrid Application Mapping for Real-Time Applications in Heterogeneous Many-Core Systems
abstract
Modern embedded many-core systems host, among others, real-time applications which must be dynamically launched at run time. To this end, Hybrid Application Mapping (HAM) methodologies combine design-time analysis with runtime mapping techniques to enable dynamic application mapping with performance guarantees, e.g., w.r.t. real-time constraints. They rely on composability to derive the required performance guarantees in an isolated analysis of individual applications at design time. The ongoing process technology downsizing, however, has given rise to an increased on-chip temperature, so that the thermal integrity of the platform must be monitored and enforced at run time by means of Dynamic Thermal Management (DTM) techniques which use countermeasures e.g. DVFS and power gating. This, however, violates composability, as the thermally unsafe behavior of one application may trigger DTM countermeasures that affect other applications running in the thermally affected region which, in turn, may lead to the violation of their real-time constraints. As a remedy, this paper proposes, for the first time, a thermally composable HAM methodology that enforces thermal safety proactively at the launch time of applications and, thereby, prevents DTM interferences which react to thermal violations. To that end, we present (a) a novel thermal-safety analysis that can be integrated into the design-time analysis of HAM and (b) a set of thermal-safety admission checks that can be used at run time when launching an application. By establishing thermal composability among running applications, the proposed HAM approach enables providing thermally safe real-time guarantees for dynamically mapped applications in many-core systems. Experimental results for a variety of hard real-time applications on multiple heterogeneous many-core architectures demonstrate the efficiency and effectiveness of the proposed methodology.
Behnaz Pourmohseni, Fedor Smirnov, Heba Khdr, Stefan Wildermann, Jürgen Teich, Jörg Henkel
RTSS6
2019 From Cloud Down to Things: An Overview of Machine Learning in Internet of Things
abstract
With the numerous Internet of Things (IoT) devices, the cloud-centric data processing fails to meet the requirement of all IoT applications. The limited computation and communication capacity of the cloud necessitate the edge computing, i.e., starting the IoT data processing at the edge and transforming the connected devices to intelligent devices. Machine learning (ML) the key means for information inference, should extend to the cloud-to-things continuum too. This paper reviews the role of ML in IoT from the cloud down to embedded devices. Different usages of ML for application data processing and management tasks are studied. The state-of-the-art usages of ML in IoT are categorized according to their application domain, input data type, exploited ML techniques, and where they belong in the cloud-to-things continuum. The challenges and research trends toward efficient ML on the IoT edge are discussed. Moreover, the publications on the “ML in IoT” are retrieved and analyzed systematically using ML classification techniques. Then, the growing topics and application domains are identified.
Farzad Samie, Lars Bauer, Jörg Henkel
IEEE Internet Things J.3
2019 Application and Thermal-reliability-aware Reinforcement Learning Based Multi-core Power Management
abstract
Power management through dynamic voltage and frequency scaling (DVFS) is one of the most widely adopted techniques. However, it impacts application reliability (due to soft errors, circuit aging, and deadline misses). However, increased power density impacts the thermal reliability of the chip, sometimes leading to permanent failure. To balance both application- and thermal-reliability along with achieving power savings and maintaining performance, we propose application- and thermal-reliability-aware reinforcement learning–based multi-core power management in this work. The proposed power management scheme employs a reinforcement learner to consider the power savings and variations in the application and thermal reliability caused by DVFS. To overcome the computational overhead, the power management decisions are determined at the application-level rather than per-core or system-level granularity. Experimental evaluation of proposed multi-core power management on a microprocessor with up to 32 cores, running PARSEC applications, was done to demonstrate the applicability and efficiency of the proposed technique. Compared to the existing state-of-the-art techniques, the proposed technique enables an average energy savings of up to ∼20%, up to 4.926°C temperature reduction without degradation in the application- and thermal-reliability.
Sai Manoj Pudukotai Dinakarrao, Anand Haridass, Muhammad Shafique 0001, Jörg Henkel, Houman Homayoun
ACM J. Emerg. Technol. Comput. Syst.5
2019 On the Efficiency of Voltage Overscaling under Temperature and Aging Effects
abstract
Voltage overscaling has received extensive attention in the last decade as an attractive paradigm for systems in which resulting timing errors and thus a loss in accuracy can be accepted in exchange for an increase in energy efficiency. At the same time, the delay of a circuit is, in turn, and in addition to voltage, also subject to temperature and aging. Existing work has largely studied voltage overscaling in isolation. This ignores interdependencies with temperature and aging, which can lead to wrong or misleading conclusions. In this work, we are the first to model the combined impact of voltage, temperature and aging on the delay of circuits towards investigating the actual existing trade-offs between efficiency and accuracy provided by voltage overscaling. We show that analyzing voltage in isolation overestimates timing errors and thus underestimates the voltage scaling potential. We further develop an approach that leverages interdependencies to optimize energy, delay and accuracy trade-offs. We precisely translate the individual and combined impact of voltage-, temperature-, and aging-induced delay increase into corresponding probability of error (Perror). This reveals that the same amount of timing increase results in different error probabilities depending on the origin (i.e., voltage, temperature or aging). For the same timing increase, voltage reductions result in the smallest-Perror compared to temperature or aging, while also reducing temperature- and aging-induced delay increases themselves. This allows voltage reduction to be employed as an effective means to minimize delay, reduce energy and thus maximize efficiency under a given upper bound on error probability. We apply our approach to multipliers in GPUs exploring the trade-off between efficiency and accuracy. We demonstrate how only accounting for voltage scaling alone leads to a considerably larger Perror (74% on average) than in reality. Our investigation also shows that for the same Perror constraint, optimizing for combined voltage, temperature and aging effects results, on average, in 116% better energydelay product (EDP) compared to state of the art.
Hussam Amrouch, Seyed Borna Ehsani, Andreas Gerstlauer, Jörg Henkel
IEEE Trans. Computers4
2019 Dynamic Guardband Selection: Thermal-Aware Optimization for Unreliable Multi-Core Systems
abstract
Circuit aging has become the major reliability concern in current and upcoming technology nodes. For instance, Bias Temperature Instability (BTI) leads to an increase in the threshold voltage of a transistor. That, in turn, may prolong the critical path delay of the processor and eventually may lead to timing errors. In order to avoid aging-induced timing errors, designers employ guardbands either with respect to voltage or frequency. State-of-the-art techniques determine a guardband type at the circuit level at design time irrespective from the running workload at the system level. Our investigation revealed that generated temperatures by a running workload have the potential to play a key role in determining the appropriate guardband type with respect to system performance. Therefore, we propose a paradigm shift in designing guardbands: to select the guardband types on-the-fly with respect to the workload-induced temperatures aiming at optimizing for performance under temperature and reliability constraints. Moreover, different guardband types for different cores can be selected simultaneously when multiple applications with diverse properties suggest this to be useful. Our dynamic guardband selection allows for a higher performance compared to techniques that employ a fixed (at design time) guardband type throughout.
Heba Khdr, Hussam Amrouch, Jörg Henkel
IEEE Trans. Computers3
2019 Hybrid Scratchpad Video Memory Architecture for Energy-Efficient Parallel HEVC
abstract
A hybrid scratchpad video memory (Hy-SVM) for energy-efficient tiles-parallelized high-efficiency video coding (HEVC) is presented here. The key ideas behind the Hy-SVM include: application-specific design and management; combined multiple levels of private and shared memories that jointly exploit intra-tile and inter-tiles data reuse; scratchpad memories (SPMs) as on-chip data storage; SRAM; and STT-RAM hybrid design. We propose a design methodology for the Hy-SVM that leverages application-specific properties to properly define the SPMs parameters. The inter-tiles data reuse potential of parallel HEVC is exploited by our run-time overlap prediction scheme, which identifies the redundant memory access behavior by analyzing monitored past frames encoding. Based on the predicted overlap characteristics, the Hy-SVM integrates memory access management units to control the access dynamics to the private/shared SPM levels. Furthermore, adaptive access management units (APMUs) can strongly reduce on-chip energy consumption due to the predicted overlap formation. The experimental results demonstrate the Hy-SVM overall energy savings of 11%-64% (4-tile) and 8%-46% (8-tile) when compared with related works. From the external memory perspective, the Hy-SVM can improve data reuse, resulting in 14%-59% of off-chip energy consumption (compared with no inter-tiles data reuse scenarios). In addition, our APMU contributes by reducing on-chip energy consumption of the Hy-SVM by 58%, on average. Thus, compared with related works, the Hy-SVM presents the lowest on-chip energy consumption. Moreover, the overhead of implementing our management units insignificantly affects the performance- and energy-efficiency of the Hy-SVM.
Felipe Sampaio, Bruno Zatt, Muhammad Shafique 0001, Jörg Henkel, Sergio Bampi
IEEE Trans. Circuits Syst. Video Technol.4
2019 ECAx: Balancing Error Correction Costs in Approximate Accelerators
abstract
Approximate computing has emerged as a design paradigm amenable to error-tolerant applications. It enables trading the quality of results for efficiency improvement in terms of delay, power, and energy consumption under user-provided tolerable quality degradation. Approximate accelerators have been proposed to expedite frequently executing code sections of error-resilient applications while meeting a defined quality level. However, these accelerators may produce unacceptable errors at run time if the input data changes or dynamic adjustments are made for a defined output quality constraint. State-of-the-art approaches in approximate computing address this issue by correctly re-computing those accelerator invocations that produce unacceptable errors; this is achieved by using the host processor or an alternate exact accelerator, which is activated on-demand. Nevertheless, such approaches can nullify the benefits of approximate computing, especially when input data variations are high at run time and errors due to approximations are above a tolerable threshold. As a robust and general solution to this problem, we propose ECAx, a novel methodology to explore low-overhead error correction in approximate accelerators by selectively correcting most significant errors, in terms of their magnitude, without losing the gains of approximations. We particularly consider the case of approximate accelerators built with approximate functional units such as approximate adders. Our novel methodology reduces the required exact re-computations on the host processor, achieving up to 20% performance gain compared to state-of-the-art approaches.
Jorge Castro-Godínez, Muhammad Shafique 0001, Jörg Henkel
ACM Trans. Embed. Comput. Syst.3
2019 Oops: Optimizing Operation-mode Selection for IoT Edge Devices
abstract
The massive increase of IoT devices and their collected data raises the question of how to analyze all that data. Edge computing provides a suitable compromise, but the question remains: How much processing should be done locally vs. offloaded to other devices? The diverse application requirements and limited resources at the edge extend the challenges. We propose Oops , an optimization framework to adapt the resource management at runtime distributedly. It orchestrates the IoT devices and adapts their operation mode with respect to their constraints and the gateway’s limited shared resources. Oops reduces runtime overhead significantly while increasing user utility compared to state-of-the-art.
Farzad Samie, Vasileios Tsoutsouras, Lars Bauer, Sotirios Xydis, Dimitrios Soudris, Jörg Henkel
ACM Trans. Internet Techn.6
2019 Estimating and Mitigating Aging Effects in Routing Network of FPGAs
abstract
In this paper, we present a comprehensive analysis of the impact of aging on the interconnection network of field-programmable gate arrays (FPGAs) and propose novel approaches to mitigate the aging effects on the routing network. We first show the insignificant impact of aging on data integrity of FPGAs, i.e., static noise margin and soft error rate of the configuration cells, as well as we show the negligible impact of the mentioned degradations on the FPGA performance. As such, we focus on the performance degradation of datapath transistors. In this regard, we propose a routing accompanied by a placement algorithm that prevents constant stress on transistors by evenly distributing the stress through the interconnection resources. By observing the impact of the signal probability on the aging of routing buffers, we enhance the synthesis flow as well as augment the proposed routing algorithm to converge the signal probabilities toward aging-friendly values. Experimental results over a set of industrial benchmarks and commerciallike FPGA architecture indicate the effectiveness of the proposed method with 64.3% reduction of stress duration in multiplexers and up to 45.2% improvement of the degradation of buffers. Altogether, the proposed method reduces the timing guardband by from 14.1% to 31.7%, depending on the FPGA routing architecture.
Behnam Khaleghi, Behzad Omidi, Hussam Amrouch, Jörg Henkel, Hossein Asadi 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2019 Modeling the Interdependences Between Voltage Fluctuation and BTI Aging
abstract
With technology scaling, the susceptibility of circuits to different reliability degradations is steadily increasing. Aging in transistors due to bias temperature instability (BTI) and voltage fluctuation in the power delivery network of circuits due to IR-drops are the most prominent. In this paper, we are reporting for the first time that there are interdependences between voltage fluctuation and BTI aging that are nonnegligible. Modeling and investigating the joint impact of voltage fluctuation and BTI aging on the delay of circuits, while remaining compatible with the existing standard design flow, is indispensable in order to answer the vital question, “what is an efficient (i.e., small, yet sufficient) timing guardband to sustain the reliability of circuit for the projected lifetime?” This is, concisely, the key goal of this paper. Achieving that would not be possible without employing a physics-based BTI model that precisely describes the underlying generation and recovery mechanisms of defects under arbitrary stress waveforms. For this purpose, our model is validated against varied semiconductor measurements covering a wide range of voltage, temperature, frequency, and duty cycle conditions. To bring reliability awareness to existing EDA tool flows, we create standard cell libraries that contain the delay information of cells under the joint impact of aging and IR-drop. Our libraries can be directly deployed within the standard design flow because they are compatible with existing commercial tools (e.g., Synopsys and Cadence). Hence, designers can leverage the mature algorithms of these tools to accurately estimate the required timing guardbands for any circuit despite its complexity. Our investigation demonstrates that considering aging and IR-drop effects independently, as done in the state of the art, leads to employing insufficient and thus unreliable guardbands because of the nonnegligible (on average 15% and up to 25%) underestimations. Importantly, considering interdependences between aging and IR-drop does not only allow correct guardband estimations, but it also results in employing more efficient guardbands.
Sami Salamin, Victor M. van Santen, Hussam Amrouch, Narendra Parihar, Souvik Mahapatra, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.6
2018 Aging-constrained performance optimization for multi cores
abstract
Circuit aging has become a dire design concern and hence it is considered a primary design constraint. Current practice to cope with this problem is to apply (too) conservative means.
Heba Khdr, Hussam Amrouch, Jörg Henkel
DAC3
2018 QoS-aware stochastic power management for many-cores
abstract
A many-core processor can execute hundreds of multi-threaded tasks in parallel on its 100s - 1000s of processing cores. When deployed in a Quality of Service (QoS)-based system, the many-core must execute a task at a target QoS. The amount of processing required by the task for the QoS varies over the task's lifetime. Accordingly, Dynamic Voltage and Frequency Scaling (DVFS) allows the many-core to deliver precise amount of processing required to meet the task QoS guarantee while conserving power. Still, a global control is necessitated to ensure that the many-core overall does not exceed its power budget.
Anuj Pathania, Heba Khdr, Muhammad Shafique 0001, Tulika Mitra, Jörg Henkel
DAC5
2018 Compiler-driven error analysis for designing approximate accelerators
abstract
Approximate Computing has emerged as a design paradigm suitable to applications with inherent error resilience. This paradigm aims to reduce the associated computing costs (such as execution time, area, or energy) of exact calculations by reducing the quality of their results. Several approximate arithmetic circuits have been proposed, which can be used to implement hardware blocks such as approximate accelerators. However, to satisfy quality constraints in these accelerators, it is imperative to assess how the errors introduced by approximate circuits propagate through other exact and approximate computations, and finally accumulate at the output. This is, in particular, crucial to enable high-level synthesis of approximate accelerators. This work proposes a compiler-driven error analysis methodology to evaluate the behavior of errors generated from approximate adders in the design of approximate accelerators. We present CEDA, a tool to perform a static analysis of the error propagation. This tool uses #pragma-based annotated C/C++ source code as input. With these annotations, exact additions are replaced by approximate ones during the code analysis to estimate the error at the output. The error estimations produced by our tool are comparable to those obtained through simulations.
Jorge Castro-Godínez, Sven Esser, Muhammad Shafique 0001, Santiago Pagani, Jörg Henkel
DATE5
2018 Task scheduling for many-cores with S-NUCA caches
abstract
A many-core processor may comprise a large number of processing cores on a single chip. The many-core's last-level shared cache can potentially be physically distributed alongside the cores in the form of cache banks connected through a Network on Chip (NoC). Static Non-Uniform Cache Access (S-NUCA) memory address mapping policy provides a scalable mechanism for providing the cores quick access to the entire last-level cache. By design, S-NUCA introduces a unique topology-based performance heterogeneity and we introduce a scheduler that can exploit it. The proposed scheduler improves performance of the many-core by 9.93% in comparison to a state-of-the-art generic many-core scheduler with minimal run-time overheads.
Anuj Pathania, Jörg Henkel
DATE2
2018 Highly efficient and accurate seizure prediction on constrained IoT devices
abstract
In this paper we present an efficient and accurate algorithm for epileptic seizure prediction on low-power and portable IoT devices. State-of-the-art algorithms suffer from two issues: computation intensive features and large internal memory requirement, which make them inapplicable for constrained devices. We reduce the memory requirement of our algorithm by reducing the size of data segments (i.e. the window of input stream data on which the processing is performed), and the number of required EEG channels. To respect the limitations of the processing capability, we reduce the complexity of our exploited features by only considering the simple features, which also contributes to reducing the memory requirements. Then, we provide new relevant features to compensate the information loss due to the simplifications (i.e. less number of channels, simpler features, shorter segment, etc.). We measured the energy consumption (12.41 mJ) and execution time (565 ms) for processing each segment (i.e. 5.12 seconds of EEG data) on a low-power MSP432 device. Even though the state-of-art does not fit to IoT devices, we evaluate the classification performance and show that our algorithm achieves the highest AUC score (0.79) for the held-out data and outperforms the state-of-the-art.
Farzad Samie, Sebastian Paul, Lars Bauer, Jörg Henkel
DATE4
2018 Estimating and optimizing BTI aging effects: from physics to CAD
abstract
Transistor aging due to Bias Temperature Instability (BTI) is a crucial degradation that affects the reliability of circuits over time. Aging-aware circuit design flows do virtually not exist yet and even research is in its infancy. In this work, we demonstrate how the deleterious effects BTI-induced degradations can be modeled from physics, where they do occur, all the way up to the system level, where they finally take place and affect the delay and power of circuits. To achieve that, degradation-aware cell libraries, that properly capture the impact of BTI not only on the delay of standard cells but also on their static and dynamic power, are created. Unlike state of the art, which solely models the impact of BTI on the threshold voltage of transistors $(V_{th})$ , we are the first to model the other key transistor parameters degraded by BTI like carrier mobility ( $\mu$ ), sub-threshold slope ( $SS$ ), and gate-drain capacitance $(C_{gd})$ . Our cell libraries are compatible with existing commercial CAD tools. Employing the mature algorithms in such tools, enables designers – after importing our cell libraries – to accurately estimate the overall impact of aging on changing the delay and/or power of any circuit, despite its complexity. We demonstrate that $\Delta V_{th}$ alone (as done in state of the art) is insufficient to correctly model the impact of BTI either on delay or power of circuits. On the one hand, neglecting BTl-induced $\mu$ and $C_{gd}$ degradations leads to underestimating the impact that BTI has on increasing the delay of circuits. Hence, designers will employ narrower timing guardbands in which reliability of circuits during lifetime cannot be sustained. On the other hand, neglecting BTI-induced $SS$ degradation leads to overestimating the impact that BTI has on static power reduction. Hence, the potential benefit of circuits from BTI will be exaggerated.
Hussam Amrouch, Victor M. van Santen, Jörg Henkel
ICCAD3
2018 Dynamic resource management for heterogeneous many-cores
abstract
With the advent of many-core systems, use cases of embedded systems have become more dynamic: Plenty of applications are concurrently executed, but may dynamically be exchanged and modified even after deployment. Moreover, resources may temporally or permanently become unavailable because of thermal aspects, dynamic power management, or the occurrence of faults. This poses new challenges for reaching objectives like timeliness for real-time or performance for best-effort program execution and maximizing system utilization. In this work, we first focus on dynamic management schemes for reliability/aging optimization under thermal constraints. The reliability of on-chip systems in the current and upcoming technology nodes is continuously degrading with every new generation because transistor scaling is approaching its fundamental limits. Protecting systems against degradation effects such as circuits' aging comes with considerable losses in efficiency. We demonstrate in this work why sustaining reliability while maximizing the utilization of available resources and hence avoiding efficiency loss is quite challenging – this holds even more when thermal constraints come into play. Then, we discuss techniques for run-time management of multiple applications which sustain real-time properties. Our solution relies on hybrid application mapping denoting the combination of design-time analysis with run-time application mapping. We present a method for Real-time Mapping Reconfiguration (RMR) which enables the Run-Time Manager (RM) to execute realtime applications even in the presence of dynamic thermal- and reliability-aware resource management. This paper is paper of the ICCAD 2018 Special Session on “Managing Heterogeneous Many-cores for High-Performance and Energy-Efficiency”. The other two papers of this Special sessions are [1] and [2].
Jörg Henkel, Jürgen Teich, Stefan Wildermann, Hussam Amrouch
ICCAD1
2018 Trading Off Temperature Guardbands via Adaptive Approximations
abstract
Runtime circuit delay variations due to degradation effects like temperature are traditionally protected against using worst-case timing guardbands. Such an approach leads to a permanent performance overhead even though effects may only be transient. Recently, approximate computing has been proposed as a technique to trade off quality for various metrics. Existing approaches, however, do not target reductions in circuit delays and guardbands, or have only been applied statically. In this paper, we propose a novel design paradigm in which adaptive approximations are employed to dynamically trade off transient, degradation-induced variations in circuit delays and associated worst-case timing guardbands for permanent performance improvements with minimal quality loss. A key challenge is to design circuits that exhibit a significant delay profile across approximation levels while maintaining a high base performance. To achieve that, we introduce and implement two approaches for synthesizing arbitrary dynamic quality-versus delay-configurable circuits at fine temporal and spatial granularities while exploring associated area, speed and quality trade-offs. We apply our approach specifically to temperature variations and guardbands. Results for an IDCT image decoding example show up to 21% speedup with less than 2% area and energy impact compared to traditional guardbanding while maintaining a worst-case transient PSNR of at least 39dB.
Behzad Boroujerdian, Hussam Amrouch, Jörg Henkel, Andreas Gerstlauer
ICCD3
2018 Reliability Estimations of Large Circuits in Massively-Parallel GPU-SPICE
abstract
SPICE simulations for reliability have special requirements. We present GPU-SPICE to serve these special requirements. First, our GPU-SPICE employs the massive parallelism found in GPUs to enable circuit simulations beyond 200K transistors. This is necessary to study reliability in microarchitecture components (e.g., multipliers, adders), as reliability estimations require full analogue SPICE simulations (instead of STA or other heuristics). Secondly, our GPU-SPICE can update transistor parameters during the circuit simulation, a feature necessary to model reliability degradation, which constantly reacts to circuit activity (e.g., Bias Temperature Instability reacting to Vgschanges by increasing/decreasing ΔVthin each transistor). Lastly, our GPU-SPICE is open-source software, this ensures that it easily can be employed, adapted and extended by other researchers. Due to the massive parallelism in a GPU and performance optimizations (convergence criteria, CUDA memory management, etc.), our GPU-SPICE is up to 218x faster than its single-threaded baseline NGSPICE.
Victor M. van Santen, Hussam Amrouch, Jörg Henkel
IOLTS3
2018 Pareto-Optimal Power- and Cache-Aware Task Mapping for Many-Cores with Distributed Shared Last-Level Cache
abstract
Two factors primarily affect performance of multi-threaded tasks on many-core processors with both shared and physically distributed Last-Level Cache (LLC): the power budget associated with a certain task mapping that aims to guarantee thermally safe operation and the non-uniform LLC access latency of threads running on different cores. Spatially distributing threads across the many-core increases the power budget, but unfortunately also increases the associated LLC latency. On the other side, mapping more threads to cores near the center of the many-core decreases the LLC latency, but unfortunately also decreases the power budget. Consequently, both metrics (LLC latency and power budget) cannot be simultaneously optimal, which leads to a Pareto-optimization that has formerly not been exploited. We are the first to present a run-time task mapping algorithm called PCMap that exploits this trade-off. Our approach results in up to 8.6% reduction in the average task response time accompanied by a reduction of up to 8.5% in the energy consumption compared to the state-of-the-art.
Martin Rapp, Anuj Pathania, Jörg Henkel
ISLPED3
2018 Recent advances in EM and BTI induced reliability modeling, analysis and optimization (invited)
Sheldon X.-D. Tan, Hussam Amrouch, Taeyoung Kim 0001, Zeyu Sun 0001, Chase Cook, Jörg Henkel
Integr.6
2018 Aging-Aware Boosting
abstract
DVFS-based boosting techniques have been widely employed by commercial multi-core processors, due to their superiority in improving the performance. Boosting, however, is particularly stressing circuits and hence it significantly contributes to an accelerated aging process. Circuit aging has become a real reliability concern because it leads to an increase in transistor threshold voltage that may cause timing errors as a result of higher delays in critical paths. Thus, high performance is desirable but it shortens the circuit lifetime through aging leaving a choice to trade-off. Besides well-known long-term aging effects, recent research also reported short-term aging effects. Our claim is that DVFS-based boosting techniques should consider both long- and short-term aging effects. This can be circumvented by wider timing guardbands. But that would be more expensive. The goal of this work is therefore to analyze and optimize boosting under specific consideration of long-term and short-term aging effects. As a result of our findings, we propose the first comprehensive aging-aware, yet efficient boosting technique. The employed aging-aware cell libraries in this work are publicly available at http://ces.itec.kit.edu/dependable-hardware.php.
Heba Khdr, Hussam Amrouch, Jörg Henkel
IEEE Trans. Computers3
2018 SlackHammer: Logic Synthesis for Graceful Errors Under Frequency Scaling
abstract
We present a novel systematic logic synthesis methodology, that assesses potential delay improvements in noncritical paths for any circuit. It synthesizes them with tighter constraints toward minimizing the number of near-critical paths and, as a result, reducing the probability of timing violations when frequency is overscaled. We demonstrate that our methodology reduces the number of near-critical paths by up to 93% and offer favorable accuracy and performance tradeoffs, up to 15× reduction in error rate and 7× reduction in mean relative error under timing speculations when compared to traditional synthesis methods. Additionally, when used together with precision scaling in cross-layer approximation techniques, it facilitates a further 27% frequency increase over an increase achievable with traditional synthesis methods. The area and power overheads of experimented circuits are up to 14% and 12%, respectively. Our methodology is compatible with traditional Electronic design automation flows. It inherits the rich feature set of existing tools and leverages their entire scope of optimizations.
Tanfer Alan, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2018 Co-Scheduling on Fused CPU-GPU Architectures With Shared Last Level Caches
abstract
Fused CPU-GPU architectures integrate a CPU and general-purpose GPU on a single die. Recent fused architectures even share the last level cache (LLC) between CPU and GPU. This enables hardware-supported byte-level coherency. Thus, CPU and GPU can execute computational kernels collaboratively, but novel methods to co-schedule work are required. This paper contributes three dynamic co-scheduling methods. Two of our methods implement workers that autonomously acquire work from a common set of independent work items (similar to bag- f-tasks scheduling). The third method, host-side profiling, uses a fraction of the total work of a kernel to determine a ratio of how to distribute work to CPU and GPU based on profiling. The resulting ratio is used for the following executions of the same kernel. Our methods are realized using OpenCL 2.0, which introduces fine-grained shared virtual memory (SVM) to allocate coherent memory between CPU and GPU. We port the Rodinia Benchmark Suite, a standard suite for heterogeneous computing, to fine-grained SVM and fused CPU-GPU architectures (Rodinia-SVM). We evaluate the overhead of fine-grained SVM and analyze the suitability of OpenCL 2.0's new features for co-scheduling. Our host-side profiling method performs competitively to the optimal choice of executing kernels either on CPU or GPU (hypothetical xor-Oracle). On average, it achieves 97% of xor-Oracle's performance and a 1.43× speedup over using the GPU alone (standard in Rodinia). We show, however, that in most cases it is not beneficial to split the work of a kernel between CPU and GPU compared to exclusively running it on the most suitable single compute device. For a fixed amount of work per device, cache-related stalls can increase by up to 1.75× when both devices are used in parallel instead of exclusively while cache misses remain the same. Thus, not the cost of cache conflicts, but inefficient cache coherence is a major performance bottleneck for current fused CPU-GPU Intel architectures with shared LLC.
Marvin Damschen, Frank Mueller 0001, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2018 Distributed Trade-Based Edge Device Management in Multi-Gateway IoT
abstract
The Internet-of-Things (IoT) envisions an infrastructure of ubiquitous networked smart devices offering advanced monitoring and control services. The current art in IoT architectures utilizes gateways to enable application-specific connectivity to IoT devices. In typical configurations, IoT gateways are shared among several IoT edge devices. Given the limited available bandwidth and processing capabilities of an IoT gateway, the service quality (SQ) of connected IoT edge devices must be adjusted over time not only to fulfill the needs of individual IoT device users but also to tolerate the SQ needs of the other IoT edge devices sharing the same gateway. However, having multiple gateways introduces an interdependent problem, the binding, i.e., which IoT device shall connect to which gateway. In this article, we jointly address the binding and allocation problems of IoT edge devices in a multigateway system under the constraints of available bandwidth, processing power, and battery lifetime. We propose a distributed trade-based mechanism in which after an initial setup, gateways negotiate and trade the IoT edge devices to increase the overall SQ. We evaluate the efficiency of the proposed approach with a case study and through extensive experimentation over different IoT system configurations regarding the number and type of the employed IoT edge devices. Experiments show that our solution improves the overall SQ by up to 56% compared to an unsupervised system. Our solution also achieves up to 24.6% improvement on overall SQ compared to the state-of-the-art SQ management scheme, while they both meet the battery lifetime constraints of the IoT devices.
Farzad Samie, Vasileios Tsoutsouras, Lars Bauer, Sotirios Xydis, Dimitrios Soudris, Jörg Henkel
ACM Trans. Cyber Phys. Syst.6
2018 Guest Editorial for the Special Issue of ESWEEK 2016
abstract
Embedded Systems Week (ESWEEK) is the premier event covering major aspects of hardware and software in the design and architecture of embedded and cyber-physical systems.It brings together three leading conferences (CASES, CODES+ISSS, and EMSOFT), two symposia (ESTIMedia and RSP), workshops, and tutorials.With 363 registered attendees, ESWEEK 2016 was well attended, showing the large interest in the field.ESWEEK 2016 was held in Pittsburgh October 2-7, following the tradition of ESWEEK to rotate between Europe, America, and Asia.At the core of ESWEEK, from Monday to Wednesday, are the three conferences CASES, Codes+ISSS, and EMSOFT, which received 63 technical paper submissions (acceptance ratio 29%), 80 (26%), and 98 (26%), respectively.In addition to the technical paper sessions, special sessions continued to be an important part of the conferences, as experts gave overviews of the newest embedded systems trends.New in 2016 was a focus on the Internet of Things (IoT): The IoT Day (part of Codes+ISSS) presented newest trends in IoT as a mix of technical papers, special sessions, and invited speakers from an embedded-systems point of view.For the first time, the review process of the conferences was conducted in a journal-like, two-stage, peer-reviewed process, with the opportunity for minor/major revision before final decision.This step was in preparation for ESWEEK to move to a journal-integrated publication model starting with 2017.Highlights of ESWEEK were the three keynote presentations: The Monday keynote by Prof. Srini Devadas from MIT emphasized the importance of Secure Hardware Platforms for the IoT.The Tuesday keynote by Louis K. Scheffer from the Howard Hughes Medical Institute presented new paradigms for how to design software and hardware in Learning from Life.Finally, the Wednesday keynote by Kaushik Roy from Purdue University presented the newest trends in approximate computing, a new paradigm that promises to increase computing efficiency.Out of the 241 received technical papers, nine outstanding contributions have been nominated as Best Paper Candidates.Finally, the best paper committees selected the best paper for each conference.The authors of eight papers nominated as Best Paper Candidates for the three ESWEEK conferences have submitted extended versions of their work to this special issue.We have three papers from CASES ("CaffePresso: Accelerating Convolutional Networks on Embedded SoCs," winner of the Best Paper Award; "LOCUS: Low-Power Customizable Many-Core Architecture for Wearables"; and "D-PUF: An Intrinsically Reconfigurable DRAM PUF for Device Authentication and Random Number Generation"), two papers from CODES+ISSS ("Improving Write Performance and Extending Endurance of Object-Based NAND Flash Devices," winner of the Best Paper Award; and "Fault Injection for Test-Driven Development of Robust SoC Firmware"), and three papers from EMSOFT ("Underminer: A Framework for Identifying Nonconverging Behaviors in Black Box System Models," winner of the Best Paper Award; "Simulation-Driven Reachability Using Matrix Measures"; and "Predictable Shared Cache Management for Multicore Real-Time Virtualization").The CASES papers address three extremely relevant application areas of current and future embedded systems: machine learning, Internet of Things, and support for security.The first paper,
Petru Eles, Jörg Henkel
ACM Trans. Embed. Comput. Syst.2
2018 Preemption of the Partial Reconfiguration Process to Enable Real-Time Computing With FPGAs
abstract
To improve computing performance in real-time applications, modern embedded platforms comprise hardware accelerators that speed up the task’s most compute-intensive parts. A recent trend in the design of real-time embedded systems is to integrate field-programmable gate arrays (FPGA) that are reconfigured with different accelerators at runtime, to cope with dynamic workloads that are subject to timing constraints. One of the major limitations when dealing with partial FPGA reconfiguration in real-time systems is that the reconfiguration port can only perform one reconfiguration at a time: if a high-priority task issues a reconfiguration request while the reconfiguration port is already occupied by a lower-priority task, the high-priority task has to wait until the current reconfiguration is completed (a phenomenon known as priority inversion ), unless the current reconfiguration is aborted (introducing unbounded delays in low-priority tasks, a phenomenon known as starvation ). This article shows how priority inversion and starvation can be solved by making the reconfiguration process preemptive —that is, allowing it to be interrupted at any time and resumed at a later time without restarting it from scratch. Such a feature is crucial for the design of runtime reconfigurable real-time systems but not yet available in today’s platforms. Furthermore, the trade-off of achieving a guaranteed bound on the reconfiguration delay for low-priority tasks and the maximum delay induced for high-priority tasks when preempting an ongoing reconfiguration has been identified and analyzed. Experimental results on the Xilinx Zynq-7000 platform show that the proposed implementation of preemptive reconfiguration introduces a low runtime overhead, thus effectively solving priority inversion and starvation.
Enrico Rossi, Marvin Damschen, Lars Bauer, Giorgio C. Buttazzo, Jörg Henkel
ACM Trans. Reconfigurable Technol. Syst.5
2017 Containing guardbands
abstract
Reliability concerns may overtake conventional design constraints such as cost and performance because transistors in deep nano-CMOS era are increasingly susceptible to degradation effects. This made reliability become unsustainably expensive due to need for wider and wider guardbands (i.e. safety margins). It is in fact the time to reverse this trend: instead of widening guardbands, it is inevitable to contain them. In this work, we summarize three novel means to achieve this goal. Since the causes are of physical origin, it cannot be excluded that degradation effects influence (i.e. amplify or cancel) each other. Hence, we first investigate the interdependencies of degradation effects demonstrating that they should jointly and not separately be modeled towards designing smaller, yet sufficient guardbands. Then, we show how aging-aware logic synthesis based on our so-called degradation-aware cell libraries, enables designers to employ mature optimization algorithms available in the commercial synthesis tools to obtain more resilient circuits in which guardbands are inherently contained. Finally, instantaneous transistors aging is a recent discovery that bears a large potential for reliability optimization since it is hardly explored until now. Though aging in general has been extensively studied in last decade, investigating the impact of instantaneous aging on circuits' reliability is still in its infancy. In fact, this is a paradigm shift in aging from sole long-term reliability degradation, as in the traditional view, to short-term reliability degradation. We demonstrate how employing our physics-based aging models results in considerably smaller guardbands due to the high certainty compared to empirical aging models.
Hussam Amrouch, Jörg Henkel
ASP-DAC2
2017 Emerging (un-)reliability based security threats and mitigations for embedded systems: special session
abstract
This paper addresses two reliability-based security threats and mitigations for embedded systems namely, aging and thermal side channels. Device aging can be used as a hardware attack vector by using voltage scaling or specially crafted instruction sequences to violate embedded processor guard bands. Short-term aging effects can be utilized to cause transient degradation of the embedded device without leaving any trace of the attack. (Thermal) side channels can be used as an attack vector and as a defense. Specifically, thermal side channels are an effective and secure way to remotely monitor code execution on an embedded processor and/or to possibly leak information. Although various algorithmic means to detect anomaly are available, machine learning tools are effective for anomaly detection. We will show such utilization of deep learning networks in conjunction with thermal side channels to detect code injection/modification representing anomaly.
Hussam Amrouch, Prashanth Krishnamurthy, Naman Patel, Jörg Henkel, Ramesh Karri, Farshad Khorrami
CASES4
2017 Towards Aging-Induced Approximations
abstract
In recent technology nodes, wide guardbands are needed to overcome reliability degradations due to aging. Such guardbands manifest as reduced efficiency and performance. Existing approaches to reduce guardbands trade off aging impact for increased circuit overhead. By contrast, the goal of this work is to completely remove guardbands through exploring, for the first time, application of approximate computing principles in the context of aging. As a result of naively narrowing or removing guardbands, timing errors start to appear as transistors age. We demonstrate that even in circuits that may tolerate errors, aging can be catastrophic due to unacceptable quality loss. Furthermore, quantifying such aging-induced quality loss necessitates expensive (often infeasible) gate-level simulations of the complete design. We show how nondeterministic aging-induced timing errors can be converted into deterministic and controlled approximations instead. We first translate the required guardband over time into an equivalent reduction in precision for individual RTL components. We then demonstrate how, based on pre-characterization of RTL components, we can quantify aging-induced approximation at the whole microarchitecture level without the need for further gate-level simulations. Results show that a 3 bit reduction in precision is sufficient to sustain 10 years of operation under worst-case aging in the context of an image processing circuit. This corresponds to an acceptable PSNR reduction of merely 8 dB, while at the same time increasing area and energy efficiency by 13%.
Hussam Amrouch, Behnam Khaleghi, Andreas Gerstlauer, Jörg Henkel
DAC4
2017 Soft error-aware architectural exploration for designing reliability adaptive cache hierarchies in multi-cores
abstract
Mainstream multi-core processors employ large multilevel on-chip caches making them highly susceptible to soft errors. We demonstrate that designing a reliable cache hierarchy requires understanding the vulnerability interdependencies across different cache levels. This involves vulnerability analyses depending upon the parameters of different cache levels (partition size, line size, etc.) and the corresponding cache access patterns for different applications. This paper presents a novel soft error-aware cache architectural space exploration methodology and vulnerability analysis of multi-level caches considering their vulnerability interdependencies. Our technique significantly reduces exploration time while providing reliability-efficient cache configurations. We also show applicability/benefits for ECC-protected caches under multi-bit fault scenarios.
Arun Subramaniyan 0001, Semeen Rehman, Muhammad Shafique 0001, Akash Kumar 0001, Jörg Henkel
DATE5
2017 Optimizing temperature guardbands
abstract
We introduce the first temperature guardbands optimization based on thermal-aware logic synthesis and thermal-aware timing analysis. The optimized guardbands are obtained solely due to using our so-called thermal-aware cell libraries together with existing tool flows and not due to sacrificing timing constraints (i.e. no trade-offs). We demonstrate that temperature guardbands can be optimized at design time through thermal-aware logic synthesis in which more resilient circuits against worst-case temperatures are obtained. Our static guardband optimization leads to 18% smaller guardbands on average. We also demonstrate that thermal-aware timing analysis enables designers to accurately estimate the required guardbands for a wide range of temperatures without over/under-estimations. Therefore, temperature guardbands can be optimized at operation time through employing the small, yet sufficient guardband that corresponds to the current temperature rather than employing throughout a conservative guardband that corresponds to the worst-case temperature. Our adaptive guardband optimization results, on average, in a 22% higher performance along with 9 2% less energy. Neither thermal-aware logic synthesis nor thermal-aware timing analysis would be possible without our thermal-aware cell libraries. They are compatible with use of existing commercial tools. Hence, they allow designers, for the first time, to automatically consider thermal concerns within their design tool flows even if they were not designed for that purpose.
Hussam Amrouch, Behnam Khaleghi, Jörg Henkel
DATE3
2017 CAnDy-TM: Comparative analysis of dynamic thermal management in many-cores using model checking
abstract
Dynamic thermal management (DTM) techniques based on task migration provide a promising solution to mitigate thermal emergencies and thereby ensuring safe operation and reliability of Many-Core systems. These techniques can be classified as central or distributed on the basis of a central DTM controller for the whole system or individual DTM controllers for each core or set of cores in the system, respectively. However, having a trustworthy comparison between central (c-) and distributed (d-) DTM techniques to find out the most suitable one for a given system is quite challenging. This is primarily due to the systemic difference between cDTM and dDTM controllers, and the inherent non-exhaustiveness of simulation and emulation methods conventionally used for DTM analysis. In this paper, we present a novel methodology called CAnDy-TM (stands for Comparative Analysis of Dynamic Thermal Management) that employs Model Checking to perform formal comparative analysis for cDTM and dDTM techniques. We identify a set of generic functional and performance properties to provide a common ground for their comparison. We demonstrate the usability and benefits of our methodology by comparing state-of-the-art cDTM and dDTM techniques, and illustrate which technique is good w.r.t. thermal stability and other task migration parameters. Such an analysis helps in selecting the most appropriate DTM for a given chip.
Syed Ali Asadullah Bukhari, Faiq Khalid, Osman Hasan, Muhammad Shafique 0001, Jörg Henkel
DATE5
2017 Ultra-low power and dependability for IoT devices (Invited paper for IoT technologies)
abstract
Recent advances in technologies have allowed the design of small-size low-power and low-cost devices that can be connected to the Internet, enabling the emerging paradigm of Internet-of-things (IoT). IoT covers an ever-increasing range of applications, e.g., health-care monitoring, smart homes and buildings, etc. In this invited paper, we discuss and summarize the IoT paradigm with a special focus on energy consumption and methodologies for its minimization. Furthermore, we also discuss about reliability in the context of IoT devices. In all, this paper attempts to be a starting point for readers interested in developing energy-efficient IoT devices.
Jörg Henkel, Santiago Pagani, Hussam Amrouch, Lars Bauer, Farzad Samie
DATE1
2017 Scalable probabilistic power budgeting for many-cores
abstract
Many-core processors exhibit hundreds to thousands of cores, which can execute lots of multi-threaded tasks in parallel. Restrictive power dissipation capacity of a many-core prevents all its executing tasks from operating at their peak performance together. Furthermore, the ability of a task to exploit part of the power budget allocated to it depends upon its current execution phase. This mandates careful rationing of the power budget amongst the tasks for full exploitation of the many-core. Past research proposed power budgeting techniques that redistribute power budget amongst tasks based on up-to-date information about their current phases. This phase information needs to be constantly propagated throughout the system and processed, inhibiting scalability. In this work, we propose a novel probabilistic technique for power budgeting which requires no exchange of phase information yet provides mathematical guarantees on judicial use of the TDP. The proposed probabilistic technique reduces the power budgeting overheads by 97.13% in comparison to a non-probabilistic approach, while providing almost equal performance on simulated thousand-core system.
Anuj Pathania, Heba Khdr, Muhammad Shafique 0001, Tulika Mitra, Jörg Henkel
DATE5
2017 Theorem proving based Formal Verification of Distributed Dynamic Thermal Management schemes
Muhammad Usama Sardar, Osman Hasan, Muhammad Shafique 0001, Jörg Henkel
J. Parallel Distributed Comput.4
2017 FAMe-TM: Formal analysis methodology for task migration algorithms in Many-Core systems
Syed Ali Asadullah Bukhari, Faiq Khalid, Osman Hasan, Muhammad Shafique 0001, Jörg Henkel
Sci. Comput. Program.5
2017 Defragmentation of Tasks in Many-Core Architecture
abstract
Many-cores can execute multiple multithreaded tasks in parallel. A task performs most efficiently when it is executed over a spatially connected and compact subset of cores so that performance loss due to communication overhead imposed by the task’s threads spread across the allocated cores is minimal. Over a span of time, unallocated cores can get scattered all over the many-core, creating fragments in the task mapping. These fragments can prevent efficient contiguous mapping of incoming new tasks leading to loss of performance. This problem can be alleviated by using a task defragmenter, which consolidates smaller fragments into larger fragments wherein the incoming tasks can be efficiently executed. Optimal defragmentation of a many-core is an NP-hard problem in the general case. Therefore, we simplify the original problem to a problem that can be solved optimally in polynomial time. In this work, we introduce a concept of exponentially separable mapping (ESM), which defines a set of task mapping constraints on a many-core. We prove that an ESM enforcing many-core can be defragmented optimally in polynomial time.
Anuj Pathania, Vanchinathan Venkataramani, Muhammad Shafique 0001, Tulika Mitra, Jörg Henkel
ACM Trans. Archit. Code Optim.5
2017 Power Density-Aware Resource Management for Heterogeneous Tiled Multicores
abstract
Increasing power densities have led to the dark silicon era, for which heterogeneous multicores with different power and performance characteristics are promising architectures. This paper focuses on maximizing the overall system performance under a critical temperature constraint for heterogeneous tiled multicores, where all cores or accelerators inside a tile share the same voltage and frequency levels. For such architectures, we present a resource management technique that introduces power density as a novel system level constraint, in orderto avoid thermal violations. The proposed technique then assigns applications to tiles by choosing their degree of parallelism and the voltage/frequency levels of each tile, such that the power density constraint is satisfied. Moreover, our technique provides runtime adaptation of the power density constraint according to the characteristics of the executed applications, and reacting to workload changes at runtime. Thus, the available thermal headroom is exploited to maximize the overall system performance.
Heba Khdr, Santiago Pagani, Éricles Sousa, Vahid Lari, Anuj Pathania, Frank Hannig, Muhammad Shafique 0001, Jürgen Teich, Jörg Henkel
IEEE Trans. Computers9
2017 Fine-Grained Checkpoint Recovery for Application-Specific Instruction-Set Processors
abstract
Checkpoint recovery (CR) is a classic fault-tolerance technique, which enables computing systems to execute correctly even when affected by transient faults. Although a number of software and hardware based approaches for CR does exist, these approaches usually are either too large, too slow, or require extensive modifications to the software and the caching/memory schemes. In this paper, we propose a novel CR approach, which is based on re-engineering the instruction set of a target processor. We take the base instruction set and augment the native micro-operations, i.e., an architectural description language (ADL), with additional microoperations to perform checkpointing at the granularity of basic blocks. The recovery mechanism is realized by three custom instructions, which can undo the corruptions caused by transient faults during instruction execution, including the values of general-purpose registers, data memory, and special-purpose registers (PC, status registers, etc.), which were incorrectly modified. Our checkpoint storage is sized according to the application program executed. The experimental results show that our approach degrades the system performance by just 0.76 percent when there is no fault, and introduces an area overhead of 44 percent on average and 79 percent in the worst case. During the fault injection test with the benchmark applications, the recovery took just 62 clock cycles (worst case).
Tuo Li 0001, Muhammad Shafique 0001, Jude Angelo Ambrose, Jörg Henkel, Sri Parameswaran
IEEE Trans. Computers4
2017 Probabilistic Error Modeling for Approximate Adders
abstract
Approximate adders are widely being advocated as a means to achieve performance gain in error resilient applications. In this paper, a generic methodology for analytical modeling of probability of occurrence of error and the Probability Mass Function (PMF) of error value in a selected class of approximate adders is presented, which can serve as performance metrics for the comparative analysis of various adders and their configurations. The proposed model is applicable to approximate adders that comprise of subadder units of uniform as well as non-uniform lengths. Using a systematic methodology, we derive closed form expressions for the probability of error for a number of state-of-the-art high-performance approximate adders. The probabilistic analysis is carried out for arbitrary input distributions. It can be used to study the dependence of error statistics in an adder's output on its configuration and input distribution. Moreover, it is shown that by building upon the proposed error model, we can estimate the probability of error in circuits with multiple approximate adders. We also demonstrate that, using the proposed analysis, the comparative performance of different approximate adders can be correctly predicted in practical applications of image processing.
Sana Mazahir, Osman Hasan, Rehan Hafiz, Muhammad Shafique 0001, Jörg Henkel
IEEE Trans. Computers5
2017 Thermal Safe Power (TSP): Efficient Power Budgeting for Heterogeneous Manycore Systems in Dark Silicon
abstract
Chip manufacturers provide the Thermal Design Power (TDP) for a specific chip. The cooling solution is designed to dissipate this power level. But because TDP is not necessarily the maximum power that can be applied, chips are operated with Dynamic Thermal Management (DTM) techniques. To avoid excessive triggers of DTM, usually, system designers also use TDP as power constraint. However, using a single and constant value as power constraint, e.g., TDP, can result in significant performance losses in homogeneous and heterogeneous manycore systems. Having better power budgeting techniques is a major step towards dealing with the dark silicon problem. This paper presents a new power budget concept, called Thermal Safe Power (TSP), which is an abstraction that provides safe power and power density constraints as a function of the number of simultaneously active cores. Executing cores at any power consumption below TSP ensures that DTM is not triggered. TSP can be computed offline for the worst cases, or online for a particular mapping of cores. TSP can also serve as a fundamental tool for guiding task partitioning and core mapping decisions, specially when core heterogeneity or timing guarantees are involved. Moreover, TSP results in dark silicon estimations which are less pessimistic than estimations using constant power budgets.
Santiago Pagani, Heba Khdr, Jian-Jia Chen, Muhammad Shafique 0001, Minming Li, Jörg Henkel
IEEE Trans. Computers6
2017 Application-Guided Power-Efficient Fault Tolerance for H.264 Context Adaptive Variable Length Coding
abstract
This paper presents a fault-tolerance technique for H.264's Context-Adaptive Variable Length Coding (CAVLC) on unreliable computing hardware. The application-specific knowledge is leveraged at both algorithm and architecture levels to protect the CAVLC process (especially context adaptation and coding tables) in a reliable yet power-efficient manner. Specifically, the statistical analysis of coding syntax and video content properties are exploited for: (1) selective redundancy of coefficient/header data of video bitstreams; (2) partitioning the coding tables into various sub-tables to reduce the power overhead of fault tolerance; and (3) run-time power management of memory parts storing the sub-tables and their parity computations. Experimental results demonstrate that leveraging application-specific knowledge reduces area and performance overhead by 2x compared to a double-parity table protection technique. For functional verification and area comparison, the complete H.264 CAVLC architecture is prototyped on a Xilinx Virtex-5 FPGA (though not limited to it).
Muhammad Shafique 0001, Semeen Rehman, Florian Kriebel, Muhammad Usman Karim Khan, Bruno Zatt, Arun Subramaniyan 0001, Bruno Boessio Vizzotto, Jörg Henkel
IEEE Trans. Computers8
2017 Aging Resilience and Fault Tolerance in Runtime Reconfigurable Architectures
abstract
Runtime reconfigurable architectures based on Field-Programmable Gate Arrays (FPGAs) allow areaand power-efficient acceleration of complex applications. However, being manufactured in latest semiconductor process technologies, FPGAs are increasingly prone to aging effects, which reduce the reliability and lifetime of such systems. Aging mitigation and fault tolerance techniques for the reconfigurable fabric become essential to realize dependable reconfigurable architectures. This article presents an accelerator diversification method that creates multiple configurations for runtime reconfigurable accelerators that are diversified in their usage of Configurable Logic Blocks (CLBs). In particular, it creates a minimal number of configurations such that all single-CLB and some multi-CLB faults can be tolerated. For each fault we ensure that there is at least one configuration that does not use that CLB. Second, a novel runtime accelerator placement algorithm is presented that exploits the diversity in resource usage of these configurations to balance the stress imposed by executions of the accelerators on the reconfigurable fabric. By tracking the stress due to accelerator usage at runtime, the stress is balanced both within a reconfigurable region as well as over all reconfigurable regions of the system. The accelerator placement algorithm also considers faulty CLBs in the regions and selects the appropriate configuration such that the system maintains a high performance in presence of multiple permanent faults. Experimental results demonstrate that our methods deliver up to 3.7× higher performance in presence of faults at marginal runtime costs and 1.6× higher MTTF than state-ofthe-art aging mitigation methods.
Hongyan Zhang 0004, Lars Bauer, Michael A. Kochte, Eric Schneider, Hans-Joachim Wunderlich, Jörg Henkel
IEEE Trans. Computers6
2017 Optimal Greedy Algorithm for Many-Core Scheduling
abstract
In this paper, we propose an optimal greedy algorithm for the problem of run-time many-core scheduling. The previously best known centralized optimal algorithm proposed for the problem is based on dynamic programming. A dynamic programming-based scheduler has high overheads which grow fast with increase in both the number of cores in the many-cores as well as number of tasks independently executing on them. We show in this paper that the inherent concavity of extractable instructions per cycle in tasks with increase in number of allocated cores allows for an alternative greedy algorithm. The proposed algorithm significantly reduces the run-time scheduling overheads, while maintaining theoretical optimality. In practice, it reduces the problem solving time 10 000x to provide near-optimal solutions.
Anuj Pathania, Vanchinathan Venkataramani, Muhammad Shafique 0001, Tulika Mitra, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2017 Energy Efficiency for Clustered Heterogeneous Multicores
abstract
Heterogeneous multicore systems clustered in multiple Voltage Frequency Islands (VFIs) are the next-generation solution for power and energy efficient computing systems. Due to the heterogeneity, the power consumption and execution time of a task changes not only with Dynamic Voltage and Frequency Scaling (DVFS), but also according to the task-to-island assignment, presenting major challenges for power management and energy minimization techniques. This paper focuses on energy minimization of periodic real-time tasks (or performance-constrained tasks) on such systems, in which the cores in an island are homogeneous and share the same voltage and frequency, but different islands have different types and numbers of cores and can be executed at other voltages and frequencies. We present an efficient algorithm to minimize the total energy consumption while satisfying the timing constraints of all tasks. Our technique consists of the coordinated selection of the voltage and frequency levels for each island, together with a task partitioning strategy that considers the energy consumption of the task executing on different islands and at different frequencies, as well as the impact of the frequency and the underlying core architecture to the resulting execution time. Every task is then mapped to the most energy efficient island for the selected voltage and frequency levels, and to a core inside the island such that the workloads of the cores in a VFI are balanced. We experimentally evaluate our technique and compare it to state-of-the-art solutions, resulting in average in 25 percent less energy consumption (and up to 87 percent for some cases), while guaranteeing that all tasks meet their deadlines.
Santiago Pagani, Anuj Pathania, Muhammad Shafique 0001, Jian-Jia Chen, Jörg Henkel
IEEE Trans. Parallel Distributed Syst.5
2017 Timing Analysis of Tasks on Runtime Reconfigurable Processors
abstract
Real-time embedded systems need to be analyzable for timing guarantees. Despite significant scientific advances, however, timing analysis lags years behind current microarchitectures with out-of-order scheduling pipelines, several hardware threads, and multiple (shared) cache layers. To satisfy the increasing performance demands, analyzable performance features are required. We propose a novel timing analysis approach to introduce runtime reconfigurable instruction set processors as one way to escape the scarcity of analyzable performance while preserving the flexibility of the system. We introduce extensions to the state-of-the-art Integer linear programming (ILP)-based program path analysis for computing precise worst case time bounds in the presence of the widely used technique to continue processor execution during reconfiguration by emulating not yet reconfigured custom instructions (CIs) in software. We identify and safely bound a timing anomaly of runtime reconfiguration, where executing faster than worst case time during reconfiguration extends the execution time of the whole program. Stalling the processor during reconfiguration (easier to analyze but not state-of-the-art for reconfigurable processors) is not required in our approach. Finally, we show the precision of our analysis on a complex multimedia application with multiple reconfigurable CIs for several hardware parameters and give advice on how to deal with reconfiguration delay under timing guarantees.
Marvin Damschen, Lars Bauer, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.3
2017 Efficient Partial Online Synthesis of Special Instructions for Reconfigurable Processors
abstract
Reconfigurable processors with fine-grained runtime-reconfigurable fabrics are used to speed up applications from different domains. Such a reconfigurable fabric allows loading of application-specific accelerators, where multiple accelerators can be combined using a coarse-grained runtimereconfigurable μProgram to speed up complex computationally intensive kernels. To allow a large degree of adaptivity in the reconfigurable fabric, as it is required by, e.g., multitasking systems, the μProgram for a kernel should not be generated at compile time, as it would constrain the adaptivity of the system. To enable flexible and efficient use of the reconfigurable fabric, we propose the necessary algorithms for runtime: 1) accelerator placement (i.e., deciding where on the fabric an accelerator should be reconfigured at runtime); 2) μProgram generation; and 3) μProgram caching. Accelerator synthesis and implementation are done at compile time to reduce runtime overhead in generating accelerators. We evaluate the proposed algorithms using different application scenarios and demonstrate the proposed concepts on an field-programmable gate array-based prototype of a reconfigurable processor. In comparison with state-of-the-art reconfigurable processors that generate μPrograms at compile time, we obtain an average speedup of 1.29× (up to 1.84×).
Artjom Grudnitsky, Lars Bauer, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.3
2016 Power and thermal management in massive multicore chips: theoretical foundation meets architectural innovation and resource allocation
abstract
Continuing progress and integration levels in silicon technologies make possible complete end-user systems consisting of extremely high number of cores on a single chip targeting either embedded or high-performance computing. However, without new paradigms of energy- and thermally-efficient designs, producing information and communication systems capable of meeting the computing, storage and communication demands of the emerging applications will be unlikely. The broad topic of power and thermal management of massive multicore chips is actively being pursued by a number of researchers worldwide, from a variety of different perspectives, ranging from workload modeling to efficient on-chip network infrastructure design to resource allocation. Successful solutions will likely adopt and encompass elements from all or at least several levels of abstraction. Starting from these ideas, we consider a holistic approach in establishing the Power-Thermal-Performance (PTP) trade-offs of massive multicore processors by considering three inter-related but varying angles, viz., on-chip traffic modeling, novel Networks-on-Chip (NoC) architecture and resource allocation/mapping
Paul Bogdan, Partha Pratim Pande, Hussam Amrouch, Muhammad Shafique 0001, Jörg Henkel
CASES5
2016 Reliability-aware design to suppress aging
abstract
Due to aging, circuit reliability has become extraordinary challenging. Reliability-aware circuit design flows do virtually not exist and even research is in its infancy. In this paper, we propose to bring aging awareness to EDA tool flows based on so-called degradation-aware cell libraries. These libraries include detailed delay information of gates/cells under the impact that aging has on both threshold voltage (Vth) and carrier mobility (μ) of transistors. This is unlike state of the art which considers Vth only. We show how ignoring μ degradation leads to underestimating guard-bands by 19% on average. Our investigation revealed that the impact of aging is strongly dependent on the operating conditions of gates (i.e. input signal slew and output load capacitance), and not solely on the duty cycle of transistors. Neglecting this fact results in employing insufficient guard-bands and thus not sustaining reliability during lifetime.
Hussam Amrouch, Behnam Khaleghi, Andreas Gerstlauer, Jörg Henkel
DAC4
2016 ageOpt-RMT: compiler-driven variation-aware aging optimization for redundant multithreading
abstract
Reliability optimization in the nano-era needs to account for multiple reliability concerns. Redundant Multithreading (RMT) has emerged as a promising technique to mitigate soft-errors in multi-cores. Since variation- and aging-unawareness during RMT may increase aging of slow cores due to high utilization or lead to unbalanced aging under varying workload scenarios, we propose to leverage variations in vulnerability and duty cycle by means of multiple compiled versions. We perform variation-aware task mapping to proactively reduce the aging of slower cores and thereby maintaining the minimum processing capabilities. Afterwards, we perform an aging-aware activation/deactivation of RMT considering tasks' variable resilience properties and select appropriate reliable versions for the mapped tasks. Experimental results demonstrate that compared to state-of-the-art aging-unaware RMT techniques, our ageOpt-RMT provides improved and balanced aging profiles by 2x on average.
Florian Kriebel, Semeen Rehman, Muhammad Shafique 0001, Jörg Henkel
DAC4
2016 An area-efficient consolidated configurable error correction for approximate hardware accelerators
abstract
Approximate adders are widely being advocated for developing hardware accelerators to perform complex arithmetic operations. Most of the state-of-the-art accuracy configurable approximate adders utilize some integrated Error Detection and Correction (EDC) circuitry. Consequently, the accumulated area overhead due to the EDC (integrated within individual adders) is significant. In this paper, we propose a low-cost Consolidated Error Correction (CEC) unit, that essentially corrects the accumulated error at the accelerator output. The proposed CEC is based on a mathematical model of approximation error. We integrate our CEC unit in approximate hardware accelerators deployed in different applications to demonstrate its area savings and speed enhancement compared to state-of-the-art.
Sana Mazahir, Osman Hasan, Rehan Hafiz, Muhammad Shafique 0001, Jörg Henkel
DAC5
2016 Distributed scheduling for many-cores using cooperative game theory
abstract
Many-cores are envisaged to include hundreds of processing cores etched on to a single die and will execute tens of multi-threaded tasks in parallel to exploit their massive parallel processing potential. A task can be sped up by assigning it to more than one core. Moreover, processing requirements of tasks are in a constant state of flux and some of the cores assigned to a task entering a low processing requirement phase can be transferred to a task entering high requirement phase, maximizing overall performance of the system.
Anuj Pathania, Vanchinathan Venkataramani, Muhammad Shafique 0001, Tulika Mitra, Jörg Henkel
DAC5
2016 Improving mobile gaming performance through cooperative CPU-GPU thermal management
abstract
State-of-the-art thermal management techniques independently throttle the frequencies of high-performance multi-core CPU and powerful graphics processing units (GPU) on heterogeneous multiprocessor system-on-chips deployed in latest mobile devices. For graphics-intensive gaming applications, this approach is inadequate because both the CPU and the GPU contribute towards the overall application performance (frames per second or FPS) as well as the on-chip temperature. The lack of coordination between CPU and GPU induces recurrent frequency throttling to maintain on-chip temperature below the permissible limit. This leads to significantly degraded application performance and large variation in temperature over time. We propose a control-theory based dynamic thermal management technique that cooperatively scales CPU and GPU frequencies to meet the thermal constraint while achieving high performance for mobile gaming. Experimental results with six popular Android games on a commercial mobile platform show an average 19% performance improvement and over 90% reduction in temperature variance compared to the original Linux approach.
Alok Prakash, Hussam Amrouch, Muhammad Shafique 0001, Tulika Mitra, Jörg Henkel
DAC5
2016 Designing guardbands for instantaneous aging effects
abstract
Bias Temperature Instability (BTI) is one of the key causes of reliability degradations of nano-CMOS circuits. While the long-term impact of BTI has been studied since years, the short-term implications of BTI on circuits are unexplored. In fact, in physics short-term BTI effects, i.e. instantaneous (i.e. sub μs) frequency dependent processes, have been recently reported. In order to design circuits with guardbands that are safe for long-term and instantaneous effects, new aging models are required. We are presenting the first approach that in fact considers both long-term as well as instantaneous BTI effects. It can be employed for complex circuits at the micro-architecture level. Designing guardbands based upon our physical BTI model reduces the guardbands by 41% and thus allows for the development of more cost-effective yet reliable designs. We also revisit existing state-of-the-art aging mitigation techniques to investigate how they can be properly adapted to additionally account for instantaneous aging effects. Along with our BTI model this further reduces the guardbands by up to 59%.
Victor M. van Santen, Hussam Amrouch, Javier Martín-Martínez, Montserrat Nafría, Jörg Henkel
DAC5
2016 Invited - Cross-layer approximate computing: from logic to architectures
abstract
We present a survey of approximate techniques and discuss concepts for building power-/energy-efficient computing components reaching from approximate accelerators to arithmetic blocks (like adders and multipliers). We provide a systematical understanding of how to generate and explore the design space of approximate components, which enables a wide-range of power/energy, performance, area and output quality tradeoffs, and a high degree of design flexibility to facilitate their design. To enable cross-layer approximate computing, bridging the gap between the logic layer (i.e. arithmetic blocks) and the architecture layer (and even considering the software layers) is crucial. Towards this end, this paper introduces open-source libraries of low-power and high-performance approximate components. The elementary approximate arithmetic blocks (adder and multiplier) are used to develop multi-bit approximate arithmetic blocks and accelerators. An analysis of data-driven resilience and error propagation is discussed. The approximate computing components are a first steps towards a systematic approach to introduce approximate computing paradigms at all levels of abstractions.
Muhammad Shafique 0001, Rehan Hafiz, Semeen Rehman, Walaa El-Harouni, Jörg Henkel
DAC5
2016 Resource budgeting for reliability in reconfigurable architectures
abstract
SRAM-based reconfigurable architectures are susceptible to soft-errors. Accelerators in the reconfigurable fabric need to be protected by fault tolerance techniques such as modular redundancy and scrubbing. However, blindly applying these techniques to all accelerators leads to suboptimal performance due to overprotection.
Hongyan Zhang 0004, Lars Bauer, Jörg Henkel
DAC3
2016 Towards performance and reliability-efficient computing in the dark silicon era
Jörg Henkel, Santiago Pagani, Heba Khdr, Florian Kriebel, Semeen Rehman, Muhammad Shafique 0001
DATE1
2016 Formal probabilistic analysis of distributed resource management schemes in on-chip systems
Shafaq Iqtedar, Osman Hasan, Muhammad Shafique 0001, Jörg Henkel
DATE4
2016 Power-efficient load-balancing on heterogeneous computing platforms
Muhammad Usman Karim Khan, Muhammad Shafique 0001, Apratim Gupta, Thomas Schumann, Jörg Henkel
DATE5
2016 Thermal optimization using adaptive approximate computing for video coding
Daniel Palomino 0001, Muhammad Shafique 0001, Altamiro Amadeu Susin, Jörg Henkel
DATE4
2016 Distributed fair scheduling for many-cores
Anuj Pathania, Vanchinathan Venkataramani, Muhammad Shafique 0001, Tulika Mitra, Jörg Henkel
DATE5
2016 Aging-aware voltage scaling
Victor M. van Santen, Hussam Amrouch, Narendra Parihar, Souvik Mahapatra, Jörg Henkel
DATE5
2016 Stress-aware routing to mitigate aging effects in SRAM-based FPGAs
abstract
Continuous shrinking of transistor size to provide high computation capability along with low power consumption has been accompanied by reliability degradations due to e.g., aging phenomenon. In this regard, with huge number of configuration bits, Field-Programmable Gate Arrays (FPGAs) are more susceptible to aging since aging not only degrades the performance, it may additionally result in corrupting the configuration cells and thus causing permanent circuit malfunctioning. While several works have investigated the aging effects in Look-Up Tables (LUTs), the routing fabric of these devices is seldom studied - even though it contributes to the majority of FPGAs' resources and configuration bits. Furthermore, there is a high prospect that errors in its state to propagate to the device outputs. In this paper, we first investigate aging effects in the routing fabric of FPGAs with respect to performance and reliability degradations. Based on this investigation, we enhance the conventional routing algorithm to mitigate the impact of aging by increasing the recovery time (i.e., the mechanism used to heal aging-induced defects) of transistors used in the routing resources. We examine our proposed method as reduction in stress time and required guardband to protect against aging in the routing fabric, as well as in improving the FPGA's lifetime. Our experiments show that the proposed method reduces the average stress time and aging-induced delay of routing resources by 41% and 18.3%, respectively. This, in turn, leads to improving the device lifetime by 130% compared to baseline routing. The proposed method can be applied by simple amending of conventional routing algorithms. Thus, it incurs negligible delay overhead.
Behnam Khaleghi, Behzad Omidi, Hussam Amrouch, Jörg Henkel, Hossein Asadi 0001
FPL4
2016 Architectural-space exploration of approximate multipliers
abstract
This paper presents an architectural-space exploration methodology for designing approximate multipliers. Unlike state-of-the-art, our methodology generates various design points by adapting three key parameters: (1) different types of elementary approximate multiply modules, (2) different types of elementary adder modules for summing the partial products, and (3) selection of bits for approximation in a wide-bit multiplier design. Generation and exploration of such a design space enables a wide-range of multipliers with varying approximation levels, each exhibiting distinct area, power, and output quality, and thereby facilitates approximate computing at higher abstraction levels. We synthesized our designs using Synopsys Design Compiler with a TSMC 45nm technology library and verified using ModelSim gate-level simulations. Power and quality evaluations for various designs are performed using PrimeTime and behavioral models, respectively. The selected designs are then deployed in a JPEG application. For reproducibility and to facilitate further research and development at higher abstraction layers, we have released the RTL and behavioral models of these approximate multipliers and adders as an open-source library at https://sourceforge.net/projects/lpaclib/.
Semeen Rehman, Walaa El-Harouni, Muhammad Shafique 0001, Akash Kumar 0001, Jörg Henkel
ICCAD5
2016 An Energy-Efficient Middleware for Computation Offloading in Real-Time Embedded Systems
abstract
Embedded systems have limited resources, such as computation capabilities and battery life. The Dynamic Voltage and Frequency Scaling (DVFS) technique is used to save energy by running the processor of the embedded system at low voltage and frequency levels. However, this prolongs the execution time, which may cause potential deadline misses for real-time tasks. In this paper, we propose a general-purpose middleware to reduce the energy consumption in embedded systems without violating the real-time constraints. The algorithms in the middleware adopt the computation offloading concept to reduce the workload on the processor of the embedded system by sending the computation-intensive tasks to a powerful server. The algorithms are further combined with the DVFS technique to find the running frequency (or speed) such that the energy consumption is minimized and the real-time constraints are satisfied. The evaluation shows that our approach reduces the average energy consumption down to nearly 60%, compared to executing all the tasks locally at the maximum processor speed.
Anas Toma, Santiago Pagani, Jian-Jia Chen, Wolfgang Karl, Jörg Henkel
RTCSA5
2016 Cross-Layer Reliability Modeling and Optimization: Compiler and Run-Time System Interactions
abstract
This paper presents a cross-layer reliability modeling and optimization approach that leverages multiple software layers like compiler and run-time system to improve the overall reliability considering unreliable or partially-reliable hardware. In order to bridge the gap between hardware and software to achieve high efficiency, our technique incorporates the knowledge from hardware layers during reliability modeling and design of optimization techniques. We demonstrate how different software layers operate synergistically to achieve a high degree of reliability.
Muhammad Shafique 0001, Semeen Rehman, Florian Kriebel, Jörg Henkel
SCOPES4
2016 Extending the WCET Problem to Optimize for Runtime-Reconfigurable Processors
abstract
The correctness of a real-time system does not depend on the correctness of its calculations alone but also on the non-functional requirement of adhering to deadlines. Guaranteeing these deadlines by static timing analysis, however, is practically infeasible for current microarchitectures with out-of-order scheduling pipelines, several hardware threads, and multiple (shared) cache layers. Novel timing-analyzable features are required to sustain the strongly increasing demand for processing power in real-time systems. Recent advances in timing analysis have shown that runtime-reconfigurable instruction set processors are one way to escape the scarcity of analyzable processing power while preserving the flexibility of the system. When moving calculations from software to hardware by means of reconfigurable custom instructions (CIs)—additional to a considerable speedup—the overestimation of a task’s worst-case execution time (WCET) can be reduced. CIs typically implement functionality that corresponds to several hundred instructions on the central processing unit (CPU) pipeline. While analyzing instructions for worst-case latency may introduce pessimism, the latency of CIs—executed on the reconfigurable fabric—is precisely known. In this work, we introduce the problem of selecting reconfigurable CIs to optimize the WCET of an application. We model this problem as an extension to state-of-the-art integer linear programming (ILP)-based program path analysis. This way, we enable optimization based on accurate WCET estimates with integration of information about global program flow, for example, infeasible paths. We present an optimal solution with effective techniques to prune the search space and a greedy heuristic that performs a maximum number of steps linear in the number of partitions of reconfigurable area available. Finally, we show the effectiveness of optimizing the WCET on a reconfigurable processor by evaluating a complex multimedia application with multiple reconfigurable CIs for several hardware parameters.
Marvin Damschen, Lars Bauer, Jörg Henkel
ACM Trans. Archit. Code Optim.3
2016 Task Mapping for Redundant Multithreading in Multi-Cores with Reliability and Performance Heterogeneity
abstract
Due to the architectural design, process variations and aging, individual cores in many-core systems exhibit heterogeneous performance. In many-core systems, a commonly adopted soft error mitigation technique is Redundant Multithreading (RMT) that achieves error detection and recovery through redundant thread execution on different cores for an application. However,task mappingandthe task execution mode(i.e., whether a task executes in a reliable mode with RMT or unreliable mode without RMT) need to be considered for achieving resource-efficient reliability. This paper explores how to efficiently assign the tasks onto different cores with heterogeneous performance properties and determine the execution modes of tasks in order to achieve high reliability and satisfy the tolerance of timeliness. We demonstrate that the task mapping problem under heterogeneous performance can be solved by employing Hungarian Algorithm as subroutine to efficiently assign the tasks onto the cores to optimize the system reliability with polynomial time complexity. To obtain the efficient task execution modes, we also propose an iterative mode adaptation technique and guarantee the tolerable timing constraint. Our results illustrate that compared to state-of-the-art, the proposed approaches achieve up to$80$percent reliability improvement (on average$20$percent) under different scenarios of chip frequency variation maps.
Kuan-Hsun Chen, Jian-Jia Chen, Florian Kriebel, Semeen Rehman, Muhammad Shafique 0001, Jörg Henkel
IEEE Trans. Computers6
2016 Cross-Layer Software Dependability on Unreliable Hardware
abstract
To enable reliable embedded systems, it is imperative to leverage the compiler and system software for joint optimization of functional correctness (i.e., vulnerability indexes) and timing correctness (i.e., deadline misses). This paper considers the optimization of the reliability-timing (RT) penalty, defined as a linear combination of the vulnerability and deadline misses. We propose a cross-layer approach to achieve reliable code generation and execution at compilation and system software layers for embedded systems. This is enabled by the concept of generating multiple versions for given application functions, with diverse performance and reliability tradeoffs, by exploiting different reliability-guided compilation options. As the execution time of a function is not fixed, the selection of the versions depends upon the execution behavior of the previous functions. Based on the reliability and execution time profiling of these versions, our reliability-driven system software decides the prioritization of the functions for determining their execution order and employs dynamic version selection to dynamically select a suitable version of a function. Specifically, our scheme builds a schedule table offline to optimize the RT penalty, and uses this table at run time to select suitable versions for the subsequent functions. A complex real-world application of “secure video and audio processing” composed of various functions is evaluated for reliable code generation and execution.
Semeen Rehman, Kuan-Hsun Chen, Florian Kriebel, Anas Toma, Muhammad Shafique 0001, Jian-Jia Chen, Jörg Henkel
IEEE Trans. Computers7
2016 Scalable Power Management for On-Chip Systems with Malleable Applications
abstract
We present a scalable Dynamic Power Management (DPM) scheme where malleable applications may change their degree of parallelism at run time depending upon the workload and performance constraints. We employ a per-application predictive power manager that autonomously controls the power states of the cores with the goal of energy efficiency. Furthermore, our DPM allows the applications to lend their idle cores for a short time period to expedite other critical applications. In this way, it allows for application-level scalability, while aiming at the overall system energy optimization. Compared to state-of-the-art centralized and distributed power management approaches, we achieve up to 58 percent (average ≈15-20 percent) ED2P reduction.
Muhammad Shafique 0001, Anton Ivanov, Benjamin Vogel, Jörg Henkel
IEEE Trans. Computers4
2016 Content-Aware Low-Power Configurable Aging Mitigation for SRAM Memories
abstract
Aging through Negative Bias Temperature Instability (NBTI) significantly jeopardizes reliability of SRAM-based memories. We propose a content-aware microarchitectural-level technique for mitigating aging of these SRAM-based memories, by altering the input and output data. The goal is to achieve cost-effective lifetime improvement through low-power aging balancing of all memory cells. For a configurable design, we perform power, area, and aging analysis of different aging balancing circuits. This analysis is leveraged to design a novel aging resilient memory architecture. To curtail the power overhead while still achieving a balanced aging, our architecture employs an anti-aging controller that leverages the data characteristics to take spatio-temporal aging balancing decisions. It dynamically selects: (1) which aging balancing circuit to activate, (2) at what time instant the circuit should be activated, and (3) on which SRAM cells aging balancing should be applied. This is achieved by identifying different configuration parameters, which can be adjusted at run time to balance the aging of SRAM memories. Our experiments demonstrate significant aging improvements at a low power overhead. In addition, we perform sensitivity analysis of different parameters of our architecture to demonstrate power vs. reliability tradeoffs under different run-time scenarios.
Muhammad Shafique 0001, Muhammad Usman Karim Khan, Jörg Henkel
IEEE Trans. Computers3
2016 Architecting On-Chip DRAM Cache for Simultaneous Miss Rate and Latency Reduction
abstract
On-chip dynamic random access memory (DRAM) cache has been recently employed in the memory hierarchy to mitigate the widening latency gap between high-speed cores and off-chip memory. Two important parameters are the DRAM cache miss rate (D$-MR) and the DRAM cache hit latency (D$-HL), as they strongly influence the performance. These parameters depend upon the DRAM set mapping policy. Recently proposed DRAM set mapping policies are predominantly optimized for either D$-MR or D$-HL. We propose novel DRAM set mapping policies that simultaneously reduce D$-MR (via high associativity) and D$-HL (via improved row buffer hit rates). To further improve the D$-HL, we propose a small and low latency DRAM Tag cache (DTC) structure that can quickly determine whether an access to the DRAM cache will be a hit or a miss. The performance of the proposed DTC depends upon the DTC hit rate. To increase it, we present a novel DTC insertion policy that also increases the DTC hit rate. We investigate the latency and miss rate tradeoffs when designing a DRAM cache hierarchy and analyze the effects of different policies on the overall performance. We evaluate our policies on a wide variety of workloads and compare its performance with three recent proposals for on-chip DRAM caches. For a 16-core system, our set mapping policy along with our DTC and its adaptive DTC insertion policy improve the harmonic mean instruction per cycle throughput by 25.4%, 15.5%, and 7.3% compared to state-of-the-art, while requiring 55% less storage overhead for DRAM cache hit/miss prediction.
Fazal Hameed, Lars Bauer, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2016 Resource and Throughput Aware Execution Trace Analysis for Efficient Run-Time Mapping on MPSoCs
abstract
There have been several efforts on run-time mapping of applications on multiprocessor-systems-on-chip. These traditional efforts perform either on-the-fly processing or use design-time analyzed results. However, on-the-fly processing often leads to low-quality mappings, and design-time analysis becomes computationally costly for large-size problems and require huge storage for large number of applications. In this paper, we present a novel run-time mapping approach, where identification of an efficient mapping for a use-case is done by the online execution trace analysis of the active applications. The trace analysis facilitates for fast identification of the mapping while optimizing for the system resource usage and throughput of the active applications, leading to reduced energy consumption as well. By rapidly identifying the efficient mapping at run-time, the proposed approach overcomes the mappings' exploration time bottleneck for large-size problems and their storage overhead problem when compared to the traditional approaches. Our experiments show that on average the exploration time to identify the mapping is reduced $14 {\times }$ when compared to state-of-the-art approaches and storage overhead is reduced by 92%. Additionally, energy and resource savings are achieved along with identification of high-quality mapping.
Amit Kumar Singh 0002, Muhammad Shafique 0001, Akash Kumar 0001, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2016 Reliability-Aware Adaptations for Shared Last-Level Caches in Multi-Cores
abstract
On account of their large footprint, on-chip last-level caches in multi-core systems are one of the most vulnerable components to soft errors. However, vulnerability to soft errors highly depends on the configuration and parameters of the last-level cache, especially when executing different applications concurrently. In this article we propose a novel reliability-aware reconfigurable last-level cache architecture (R 2 Cache) and cache vulnerability model for multi-cores. R 2 Cache supports various reliability-wise efficient cache configurations (i.e., cache parameter selection and cache partitioning) for different concurrently executing applications. The proposed vulnerability model takes into account the vulnerability of both the data and tag arrays as well as the active cache area for applications in different execution phases. To enable runtime adaptations, we introduce a lightweight online vulnerability predictor that exploits the knowledge of performance metrics like number of L2 misses to accurately estimate the cache vulnerability to soft errors. Based on the predicted vulnerabilities of different concurrently executing applications in the current execution epoch, our runtime reliability manager reconfigures the cache such that, for the next execution epoch, the total vulnerability for all concurrently executing applications is minimized under user-provided tolerable performance/energy overheads. In scenarios where single-bit error correction for cache lines may be afforded, vulnerability-aware reconfigurations can be leveraged to increase the reliability of the last-level cache against multi-bit errors. Compared to state-of-the-art vulnerability-minimizing and reconfigurable caches, the proposed architecture provides 35.27% and 23.42% vulnerability savings, respectively, when averaged across numerous experiments, while reducing the vulnerability by more than 65% and 60%, respectively, for selected applications and application phases.
Florian Kriebel, Semeen Rehman, Arun Subramaniyan 0001, Segnon Jean Bruno Ahandagbe, Muhammad Shafique 0001, Jörg Henkel
ACM Trans. Embed. Comput. Syst.6
2016 SPMPool: Runtime SPM Management for Memory-Intensive Applications in Embedded Many-Cores
Hossein Tajik, Bryan Donyanavard, Nikil Dutt, Janmartin Jahn, Jörg Henkel
ACM Trans. Embed. Comput. Syst.5
2016 Power-Efficient Workload Balancing for Video Applications
abstract
High workload and throughput requirements of image and video processing applications can be sustained on a many-core system. However, inefficient parallelization and processing assignments to the cores result in reduced system efficiency. Eliminating them necessitates a power-efficient and balanced workload distribution among the cores. This paper addresses these challenges by introducing a novel workload balancing and adaptation scheme. Our scheme accounts for the application characteristics and the underlying hardware, and the variation of load. Automatic selection of the number of cores and distribution of workload to each core depends on the throughput requirements, available number of cores, allowable voltage-frequency settings, and data content. Moreover, runtime derivation and fine-tuning of the workload-dependent frequency estimation models of each core are achieved using a closed-loop feedback mechanism. Furthermore, we propose an optional feedback control-based workload-tuning scheme that can further reduce the total power consumption. A case study of an advanced multithreaded video application demonstrates up to~42% power savings (average~39%) with negligible video quality degradation, using our proposed power-efficient workload-balancing and tuning.
Muhammad Usman Karim Khan, Muhammad Shafique 0001, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.3
2016 Two-State Checkpointing for Energy-Efficient Fault Tolerance in Hard Real-Time Systems
abstract
Checkpointing with rollback recovery is a well-established technique to tolerate transient faults. However, it incurs significant time and energy overheads, which go wasted in fault-free execution states and may not even be feasible in hard real-time systems. This paper presents a low-overhead two-state checkpointing (TsCp) scheme for fault-tolerant hard real-time systems. It differentiates between the fault-free and faulty execution states and leverages two types of checkpoint intervals for these two different states. The first type is nonuniform intervals that are used while no fault has occurred. These intervals are determined based on postponing checkpoint insertions in fault-free states, with the aim of decreasing the number of checkpoint insertions. The second type is uniform intervals that are used from the time when the first fault occurs. They are determined so as to minimize execution time for faulty states, leaving more time available for energy management in fault-free states. Experimental evaluation on an embedded processor (LEON3) and an emerging nonvolatile memory technology (ReRAM) illustrates that TsCp significantly reduces the number of checkpoints (62% on average) compared with previous works, while preserving fault tolerance. This results in 14% and 13% reduced execution time and energy consumption, respectively. Furthermore, we combine TsCp with dynamic voltage scaling (DVS) and achieve up to 26% (21% on average) energy saving compared with the state-of-the-art techniques.
Mohammad Khavari Tavana, Semeen Rehman, Muhammad Shafique 0001, Alireza Ejlali, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.6
2016 Analysis and Mapping for Thermal and Energy Efficiency of 3-D Video Processing on 3-D Multicore Processors
abstract
Three-dimensional video processing has high computation requirements and multicore processors realized in 3-D integrated circuits (ICs) provide promising high performance computing platforms. However, the conventional approaches to accelerate the computations involved in 3-D video processing do not exploit the high performance potential of 3-D ICs. In this paper, we propose an application-driven methodology that performs efficient mapping of 3-D video applications' components on 3-D multicores to achieve high performance (throughput). The methodology involves an extensive application analysis to exploit the spatial and temporal correlation available in 3-D neighborhood. Afterward, it leverages the correlation and thermal properties of different 3-D views to perform an efficient mapping of 3-D video processing on cores available at different layers of 3-D IC. The goal is to optimize energy consumption and peak temperature while meeting the throughput requirement. Experiments show 76% reduction in communication energy along with reduction in peak temperature when compared with approaches exploiting architecture characteristics only.
Amit Kumar Singh 0002, Muhammad Shafique 0001, Akash Kumar 0001, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.4
2015 ADAPT: An adaptive manycore methodology for software pipelined applications
abstract
Future on-chip manycore systems are expected to have hundreds of cores, and to be used for a number of applications to amortize their fabrication costs. In this paper, we examine how software pipelines, which are useful for streaming/multimedia applications, can be efficiently executed on a manycore system with shared memory. The goal is to balance the stages of the pipeline under workload and resource variations. This paper presents ADAPT, a method to quickly detect bottleneck stages and add cores (workers) to those bottleneck stages at run-time. Further, if there are no idle workers, then a shuffling of workers across stages is performed to improve/maintain throughput. ADAPT is implemented in a 48-core system which is built using a commercial core and tool suite. For a variety of applications, ADAPT takes less than 2 µs for one run-time adaptation, and achieves up to 2.1× the throughput of a state-of-the-art method (which is modified and implemented in the same system for a fair comparison). These results illustrate the applicability of ADAPT for fine-grained run-time management of manycore systems to achieve high throughput for software pipelines.
Xi Zhang 0003, Haris Javaid, Muhammad Shafique 0001, Jude Angelo Ambrose, Jörg Henkel, Sri Parameswaran
ASP-DAC5
2015 Approximation-aware Multi-Level Cells STT-RAM cache architecture
abstract
Current manycore processors exhibit large on-chip last-level caches that may reach sizes of 32MB - 128MB and incur high power/energy consumption. The emerging Multi-Level Cells (MLC) STT-RAM memory technology improves the capacity and energy efficiency issues of large-sized memory banks. However, MLC STT-RAM incurs non-negligible protection overhead to ensure reliable operations when compared to the Single-Level Cells (SLC) STT-RAM. In this paper, we propose an approximation-aware MLC STT-RAM cache architecture, which is partially-protected to restrict the reliability overhead and in turn leverages variable resilience characteristics of different applications for adaptively curtailing the protection overhead under a given error tolerance level. It thereby improves the energy-efficiency of the cache while meeting the reliability requirements. Our cache architecture is equipped with a latency-aware hardware module for double-error correction. To achieve high energy efficiency, approximation-aware read and write policies are proposed that perform approximate storage management while tolerating some errors bounded within the user-provided tolerance level. The architecture also facilitates runtime control on the quality of applications' results. We perform a case study on the next-generation advanced video encoding that exhibit memory-intensive functional blocks with varying resilience properties and support for parallelism. Experimental results demonstrate that our approximation-aware MLC STT-RAM based cache architecture can improve the energy efficiency compared to state-of-the-art fully-protected caches (7%-19%, on average), while incurring minimal quality penalties in the output (-0.219% to -0.426%, on average). Furthermore, our architecture supports complete error protection coverage for all cache data when processing non-resilient application. The hardware overhead to implement our approximation-aware management negligibly affects the energy efficiency (0.15%-1.3% of overhead) and the access latency (only 0.02%-1.56% of overhead).
Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel
CASES5
2015 SuperNet: multimode interconnect architecture for manycore chips
abstract
Designers of the on-chip interconnect for manycore chips are faced with the dilemma of meeting performance, power and reliability requirements for different operational scenarios. In this paper, we propose a multimode on-chip interconnect called SuperNet. This interconnect can be configured to run in three different modes: energy efficient mode; performance mode; and, reliability mode. Our proposed interconnect is based on two parallel multi-vt optimized packet switched network-on-chip (NoC) meshes. We describe the circuit design techniques and architectural modifications required to realize such a multimode interconnect. Our evaluation with diverse set of applications show that the energy efficient mode can save on average 40% NoC power, whereas the performance mode can improve the core IPC by up to 13% on selected high MPKI applications. The reliability mode provides protection against soft errors in the router's data path through byte oriented SECDED codes that can correct up to 8 bit errors and detect up to 16 bit errors in a 64 bit flit, whereas the router's control path is protected through DMR lock step execution.
Haseeb Bokhari, Haris Javaid, Muhammad Shafique 0001, Jörg Henkel, Sri Parameswaran
DAC4
2015 Hayat: harnessing dark silicon and variability for aging deceleration and balancing
abstract
Elevated power densities result in the so-called Dark Silicon constraint that prohibits simultaneous activation of all the cores in an on-chip system (in the full performance mode) to respect the safe thermal limits, thus enforcing a significant amount of on-chip resources to stay 'dark' (i.e., power-gated). In this paper, we show that how Dark Silicon together with the manufacturing process induced variability can be harnessed to mitigate reliability threats in the nano-era. In particular, we propose a run-time system Hayat* that harnesses Dark Silicon to decelerate and/or balance temperature-dependent aging, while also considering variability in order to improve the overall system performance for a given lifetime. Experimental evaluation across a range of chips to account for process variations illustrates that our Hayat system can provide a significant aging/performance improvement and decelerates the chip aging by 6 months -- 5 years (depending upon the required lifetime constraint) compared to state-of-the-art techniques.
Dennis Gnad, Muhammad Shafique 0001, Florian Kriebel, Semeen Rehman, Duo Sun, Jörg Henkel
DAC6
2015 New trends in dark silicon
abstract
This paper presents new trends in dark silicon reflecting, among others, the deployment of FinFETs in recent technology nodes and the impact of voltage/frquency scaling, which lead to new less-conservative predictions. The focus is on dark silicon from a thermal perspective: we show that it is not simply the chip's total power budget, e.g., the Thermal Design Power (TDP), that leads to the dark silicon problem, but instead it is the power density and related thermal effects. We therefore propose to use Thermal Safe Power (TSP) as a more efficient power budget. It is also shown that sophisticated spatio-temporal mapping decisions result in improved thermal profiles with reduced peak temperatures. Moreover, we discuss the implications of Near-Threshold Computing (NTC) and employment of Boosting techniques in dark silicon systems.
Jörg Henkel, Heba Khdr, Santiago Pagani, Muhammad Shafique 0001
DAC1
2015 Thermal constrained resource management for mixed ILP-TLP workloads in dark silicon chips
abstract
In dark silicon chips, a significant amount of on-chip resources cannot be simultaneously powered on and need to stay dark, i.e., power gated, in order to avoid thermal emergencies. This paper presents a resource management technique, called DsRem, that selects the number of active cores jointly with their voltage/frequency (v/f) levels, considering the high Instruction Level Parallelism (ILP) or Thread Level Parallelism (TLP) nature of different applications, in order to maximize the overall system performance. DsRem leverages the positioning of dark cores, to efficiently dissipate the heat generated by the active cores. This facilitates increasing the v/f level of the active cores, which leads to further performance improvement. Compared to state-of-the-art thermal-aware task application mapping, DsRem achieves up to 46% performance gain, while avoiding any thermal emergencies. Additionally, DsRem outperforms the boosting technique with 26%.
Heba Khdr, Santiago Pagani, Muhammad Shafique 0001, Jörg Henkel
DAC4
2015 A low latency generic accuracy configurable adder
abstract
High performance approximate adders typically comprise of multiple smaller sub-adders, carry prediction units and error correction units. In this paper, we present a low-latency generic accuracy configurable adder to support variable approximation modes. It provides a higher number of potential configurations compared to state-of-the-art, thus enabling a high degree of design flexibility and trade-off between performance and output quality. An error correction unit is integrated to provide accurate results for cases where high accuracy is required. Furthermore, an associated scheme for error probability estimation allows convenient comparison of different approximate adder configurations without requiring the need to numerically simulate the adder. Our experimental results validate the developed error model and also the lower latency of our generic accuracy configurable adder over state-of-the-art approximate adders. For functional verification and prototyping, we have used a Xilinx Virtex-6 FPGA. Our adder model and synthesizable RTL are made open-source.
Muhammad Shafique 0001, Rehan Hafiz, Jörg Henkel
DAC4
2015 EnAAM: energy-efficient anti-aging for on-chip video memories
abstract
Negative Biased Temperature Instability-induced aging has emerged as one of the critical reliability threats in the nano-era. In this paper, we propose a microarchitectural-level technique for mitigating aging of on-chip SRAM-based memories. The goal is to achieve balanced aging of all memory cells at minimal energy overhead, leading to a longer lifetime. For a configurable and energy-efficient design, we perform aging and energy analysis of different aging balancing circuits. This analysis is leveraged to design a novel energy-efficient anti-aging memory architecture. Our architecture employs a configurable anti-aging controller that leverages the data characteristics to dynamically select (1) at what time instant the aging balancing circuit should be activated, and (2) on which SRAM cells aging balancing should be applied. Our experiments demonstrate significant aging improvements while providing up to 30% energy savings.
Muhammad Shafique 0001, Muhammad Usman Karim Khan, Adnan Orcun Tüfek, Jörg Henkel
DAC4
2015 Malleable NoC: dark silicon inspired adaptable Network-on-Chip
Haseeb Bokhari, Haris Javaid, Muhammad Shafique 0001, Jörg Henkel, Sri Parameswaran
DATE4
2015 A deblocking filter hardware architecture for the high efficiency video coding standard
Cláudio Machado Diniz, Muhammad Shafique 0001, Felipe Vogel Dalcin, Sergio Bampi, Jörg Henkel
DATE5
2015 Formal probabilistic analysis of distributed dynamic thermal management
Shafaq Iqtedar, Osman Hasan, Muhammad Shafique 0001, Jörg Henkel
DATE4
2015 Power-efficient accelerator allocation in adaptive dark silicon many-core systems
Muhammad Usman Karim Khan, Muhammad Shafique 0001, Jörg Henkel
DATE3
2015 Adaptive on-the-fly application performance modeling for many cores
Sebastian Kobbe, Lars Bauer, Jörg Henkel
DATE3
2015 ACSEM: accuracy-configurable fast soft error masking analysis in combinatorial circuits
Florian Kriebel, Semeen Rehman, Duo Sun, Pau Vilimelis Aceituno, Muhammad Shafique 0001, Jörg Henkel
DATE6
2015 MatEx: efficient transient and peak temperature computation for compact thermal models
Santiago Pagani, Jian-Jia Chen, Muhammad Shafique 0001, Jörg Henkel
DATE4
2015 Online binding of applications to multiple clock domains in shared FPGA-based systems
Farzad Samie, Lars Bauer, Chih-Ming Hsieh, Jörg Henkel
DATE4
2015 Variability-aware dark silicon management in on-chip many-core systems
Muhammad Shafique 0001, Dennis Gnad, Siddharth Garg, Jörg Henkel
DATE4
2015 E-pipeline: elastic hardware/software pipelines on a many-core fabric
Xi Zhang 0003, Haris Javaid, Muhammad Shafique 0001, Jorgen Peddersen, Jörg Henkel, Sri Parameswaran
DATE5
2015 Mitigating the Power Density and Temperature Problems in the Nano-Era
abstract
This paper introduces the power-density and temperature induced issues in the modern on-chip systems. In particular, the emerging Dark Silicon problem is discussed along with critical research challenges. Afterwards, an overview of key research efforts and concepts is presented that leverage dark silicon for performance and reliability optimization. In case temperature constraints are violated, an efficient dynamic thermal management technique is employed.
Muhammad Shafique 0001, Jörg Henkel
ICCAD2
2015 STRAP: Stress-Aware Placement for Aging Mitigation in Runtime Reconfigurable Architectures
abstract
Aging effects in nano-scale CMOS circuits impair the reliability and Mean Time to Failure (MTTF) of embedded systems. Especially for FPGAs that are manufactured in the latest technology node, aging is amajor concern. We introduce the first cross-layer aging-aware placement method for accelerators in FPGA-based runtime reconfigurable architectures. It optimizes stress distribution by accelerator placement at runtime, i.e. to which reconfigurable region an accelerator shall be reconfigured. Additionally, it optimizes logic placement at synthesis time to diversify the resource usage of individual accelerators, i.e. which CLBs of a reconfigurable region shall be used by an accelerator. Both layers together balance the intra- and inter-region stress induced by the application workload at negligible performance cost. Experimental results show significant reduction of maximum stress of up to 64% and 35%, which leads to up to 177% and 14% MTTF improvement relative to state-of-the-art methods w.r.t. HCI and BTI aging, respectively.
Hongyan Zhang 0004, Michael A. Kochte, Eric Schneider, Lars Bauer, Hans-Joachim Wunderlich, Jörg Henkel
ICCAD6
2015 Energy-efficient multimedia systems for high efficiency video coding
abstract
This paper presents a cross-layer approach for analyzing and designing energy-efficient advanced multimedia systems with next-generation High-Efficiency Video Coding (HEVC) standard. Our approach leverages both algorithmic and architectural layers of system design abstractions in order to achieve a high power/energy efficiency. We present an analysis and design of an HEVC-based multimedia system while leveraging algorithmic-architectural collaborative optimizations and video content properties to achieve high energy efficiency.
Jörg Henkel, Muhammad Usman Karim Khan, Muhammad Shafique 0001
ISCAS1
2015 Lucid infrared thermography of thermally-constrained processors
abstract
Thermal analysis is a prerequisite for developing reliability increasing techniques for thermally-constrained processors, i.e. processors with a high power density. For that purpose, infrared (IR) camera measurement setups have been deployed with the purpose to provide direct feedback of the impact that thermal mitigation techniques have. To obtain lucid IR images1, the IR-opaque cooling must be removed and hence, an alternative IR-transparent cooling needs to be provided to protect the chip. To this end, the majority of state-of-the-art employs an IR coolant liquid to prevent the chip from overheating. The problem is that several aspects like thermal convection may interfere with the measured IR radiations resulting in equivocal IR images. Thus, they decrease the accuracy in a way that leads to incorrectly estimating reliability. Solving this prominent problem, we introduce an IR-transparent cooling that cools the chip from its rear side allowing the camera to perspicuously capture the IR emissions as no additional layer in between impedes the radiation. It maintains the on-chip temperatures within a safe range equivalent to the original heat sink-based cooling. We demonstrate how state-of-the-art inaccurate thermal analysis results in incorrectly estimating reliability. Our setup is the most accurate, least intrusive one that has been both proposed and actually applied to state-of-the-art multi-cores (Intel 45nm dual-core and 22nm octa-core).
Hussam Amrouch, Jörg Henkel
ISLPED2
2015 Hierarchical power budgeting for Dark Silicon chips
abstract
The emerging Dark Silicon limitation has led the application designers to carefully consider the available Thermal Design Power (TDP) budgets, hardware resources, and software characteristics. In this paper, we propose a hierarchical scheme for distributing the resources and TDP budget among concurrently executing applications with multi-threaded workloads under throughput constraints. Afterwards, the application-level TDP budget is partitioned among its threads depending upon their workloads, which can then be fine-tuned at run time considering workload variations. We evaluate our scheme for the next-generation, multi-threaded, High Efficiency Video Codec and demonstrate that up to 30.86% higher throughput is achieved compared to the state-of-the-art.
Muhammad Usman Karim Khan, Muhammad Shafique 0001, Jörg Henkel
ISLPED3
2015 Power management for mobile games on asymmetric multi-cores
abstract
Gaming on mobile platforms is highly power hungry and rapidly drains the limited-capacity battery. In multi-threaded gaming, each thread has different processing requirements and even a single slow thread may lead to Quality of Service (QoS) violations. Further, modern mobile platforms are equipped with asymmetric multi-core processors, so that different cores exhibit diverse power and performance properties. These asymmetric cores along with different Dynamic Power Management (DPM) techniques enable a high degree of power efficiency in mobile gaming. The default Linux power manager (i.e. “Governor”) of asymmetric multi-cores performs power-wise inefficient for mobile games as it over allocates resources for processing threads by being oblivious to the QoS. The state-of-the-art Governor for mobile gaming does not account for multi-threaded gaming workloads, which are mainstream in mobile gaming. In this work, we present a power-performance characterization of multi-threaded mobile games by executing them on a real-world mobile platform with an asymmetric multi-core. This analysis is leveraged to propose a QoS-aware Governor running a lightweight online heuristic that holistically accounts for thread-to-core mapping and DPM. This solution, when integrated into the platform's Operating System (OS), provides 12% improved power efficiency on average.
Anuj Pathania, Santiago Pagani, Muhammad Shafique 0001, Jörg Henkel
ISLPED4
2015 DRVS: Power-efficient reliability management through Dynamic Redundancy and Voltage Scaling under variations
abstract
Many-core processors facilitate coarse-grained reliability by exploiting available cores for redundant multithreading. However, ensuring high reliability with reduced power consumption necessitates joint considerations of variations in vulnerability, performance and power properties of software as well as the underlying hardware. In this paper, we propose a power-efficient reliability management system for many-core processors. It exploits various basic redundancy techniques (like, dual and triple modular redundancy) operating in different voltage-frequency levels, each offering distinct reliability, performance and power properties. Our system performs Dynamic Redundancy and Voltage Scaling (DRVS) considering process variations in hardware, and diversities in software vulnerability and execution time properties. Experiments show that DRVS system provides significant reliability improvements while providing up to 60% reduced power consumption compared to state-of-the-art techniques.
Mohammad Khavari Tavana, Semeen Rehman, Florian Kriebel, Muhammad Shafique 0001, Alireza Ejlali, Jörg Henkel
ISLPED7
2015 Dark Silicon: From Computation to Communication
abstract
In the emerging Dark Silicon era, not all parts of an on-chip system (i.e., cores, Network-on-Chip, and memory resources) can be simultaneously powered-on at the full speed. This paper aims at exposing dark silicon challenges to the NOCS community with an overview of some of the early research efforts that are attempting to shape the design and run-time management of future generation heterogeneous dark silicon processors. The goal is to cover both the computation and communication perspectives. In particular, we exploit computation and communication heterogeneity at multiple levels of system abstractions to design and manage dark silicon processors. The available dark silicon is leveraged to improve power/energy, performance, and reliability efficiency.
Jörg Henkel, Haseeb Bokhari, Siddharth Garg, Muhammad Usman Karim Khan, Heba Khdr, Florian Kriebel, Ümit Y. Ogras, Sri Parameswaran, Muhammad Shafique 0001
NOCS1
2015 Probabilistic Formal Verification Methodology for Decentralized Thermal Management in On-Chip Systems
abstract
Just like any other algorithm, Dynamic Thermal Management (DTM) schemes for multi-core architectures are susceptible to errors. Moreover, due to the wide spread usage and safety-critical nature of these schemes, there is a key demand for robust verification of these schemes before deployment. Traditional analysis techniques, like simulation and emulation, are inherently incomplete and therefore they cannot guarantee a complete absence of bugs. In this paper, we present a generic formal verification methodology, based of probabilistic model checking, for verifying decentralized DTM schemes. The paper provides a general modelling approach for developing a Markovian model of any decentralized DTM scheme. Moreover, we identify a set of generic probabilistic properties that can be of an interest to DTM scheme designers. For illustration purposes, the proposed methodology is used to verify Thermal Aware Agent Based Power Economy (TAPE), which is a state-of-the-art decentralized DTM scheme.
Shafaq Iqtedar, Osman Hasan, Muhammad Shafique 0001, Jörg Henkel
WETICE4
2015 Resource-awareness on heterogeneous MPSoCs for image processing
Johny Paul, Walter Stechele, Benjamin Oechslein, Christoph Erhardt, Jens Schedel, Daniel Lohmann, Wolfgang Schröder-Preikschat, Manfred Kröhnert, Tamim Asfour, Éricles Sousa, Vahid Lari, Frank Hannig, Jürgen Teich, Artjom Grudnitsky, Lars Bauer, Jörg Henkel
J. Syst. Archit.16
2015 A Reconfigurable Hardware Architecture for Fractional Pixel Interpolation in High Efficiency Video Coding
abstract
We present a novel reconfigurable hardware architecture for interpolation filtering in high efficient video coding that adapts to run-time changes of the number of interpolation filter calls and thereby provides a high potential of energy efficiency. It employs a picture-based prediction scheme to estimate the number of interpolation filter calls at run-time by monitoring the group of pictures history based on video coding structure knowledge. Reconfigurable acceleration engines are developed that can adapt to different filter types. Dynamic composition of different instances of these engines enables different implementation versions with area versus throughput tradeoff. A run-time selection scheme determines the best implementation version for each picture based on the throughput requirements. Compared to state-of-the-art, our architecture reduces resource usage by 57% while supporting various throughputs and video resolutions.
Cláudio Machado Diniz, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2015 Multicast FullHD H.264 Intra Video Encoder Architecture
abstract
High throughput demands have resulted in enormous increase in complexity of multicast video applications, which require multiple video encoders to simultaneously compress individual views. In this paper, we present an approach to encode independent videos using H.264 intra encoder on a single hardware platform, where the hardware resources are shared by independent encoders in a time-multiplexed manner. In addition to lowering the latency introduced by multicasting, we address the strong sequential data dependencies within the encoder. At 25 frames/s, 150 MHz prototype of the proposed encoder and multiple video capture/display on a mid-range field programmable gate array is also presented.
Muhammad Usman Karim Khan, Muhammad Shafique 0001, Lars Bauer, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2015 Energy and Peak Power Efficiency Analysis for the Single Voltage Approximation (SVA) Scheme
abstract
Energy efficiency is an important issue in computing systems and operating within a safe power budget is a necessary constraint. This paper presents a simple and practical solution both for energy minimization and peak power reduction, called Single Voltage Approximation (SVA) scheme, for periodic real-time tasks on multicore systems with a shared supply voltage in a voltage island. SVA is inspired by the Single Frequency Approximation (SFA) scheme, in which all the cores in the island run at a single voltage and frequency such that all tasks can meet their deadlines. In SVA, all the cores in the island are also executed at the same single voltage as in SFA. However, the frequency of each core is individually chosen, such that the tasks in each core can meet their deadlines, but without running at unnecessarily high frequencies. Thus, all the cores are executing tasks all the time and there is no need for any Dynamic Power Management (DPM) technique for reducing the energy consumption for idling. For task partitioning, SVA is combined with the Double Largest Task First (DLTF) partitioning scheme. Most importantly, this paper provides comprehensive analysis for combining DLTF and SVA, deriving its worst-case behavior both for energy minimization and peak power reduction, compared against the optimal solutions. Our analysis shows that, depending on the hardware, the energy consumption by combining DLTF and SVA is at most 1.95 (2.21, 2.42, and 2.59, respectively), compared to the optimal solutions, when the voltage island has up to 4 (8, 16, and 32, respectively) cores, which outperforms the worst-case factors of SFA when the cores fail to sleep efficiently. For peak power reduction, due to running at slower frequencies, combining DLTF and SVA always outperforms SFA, both in average and corner cases. Finally, we extend our analysis considering multicore systems with discrete voltage and frequency pairs and multiple voltage islands.
Santiago Pagani, Jian-Jia Chen, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2014 Low power design of the next-generation High Efficiency Video Coding
abstract
This paper provides a comprehensive analysis of the computational complexity, power consumption, temperature, and memory access behavior for the next-generation High Efficiency Video Coding (HEVC) standard. We highlight the associated design challenges and present several low-power algorithmic and architectural techniques for developing power-efficient HEVC-based multimedia system. We explore the interplay between the algorithms and architectures to provide high power efficiency while leveraging the application-specific knowledge and video content characteristics.
Muhammad Shafique 0001, Jörg Henkel
ASP-DAC2
2014 COREFAB: Concurrent reconfigurable fabric utilization in heterogeneous multi-core systems
abstract
Application-specific accelerators may provide considerable speedup in single-core systems with a runtime-reconfigurable fabric (for simplicity called "fabric" in the following). A reconfigurable core, i.e. processor core pipeline coupled to a fabric, can be integrated along with regular general purpose processor cores (GPPs) into a reconfigurable multi-core system with widely improved system performance. As most applications only use a fraction of the available fabric at a time, making the fabric usable by the GPPs (in addition to the reconfigurable core) in such a multi-core system is desirable. Existing work focused on algorithms that decide the amount of fabric that is assigned to each core in a multi-core system. However, when multiple cores access the fabric simultaneously, they are either limited to serialized fabric access or, when parallel access is supported, the size of the fabric share assigned to a core is inflexible and tends to be over- or undersized for the running application, thereby not efficiently utilizing the fabric. We propose a novel approach that allows GPPs to access the fabric of the reconfigurable core and that enables concurrent fabric utilization on-the-fly through merging fabric accesses from different cores at run-time. Compared to state-of-the art, our approach improves performance of the GPPs in a reconfigurable multi-core system by 1.3x on average, without reducing the performance of the reconfigurable core.
Artjom Grudnitsky, Lars Bauer, Jörg Henkel
CASES3
2014 Automatic custom instruction identification in memory streaming algorithms
abstract
Application-specific instruction set processors (ASIPs) extend the instruction set of a general purpose processor by dedicated custom instructions (CIs). In the last decade, reconfigurable processors advanced this concept towards run-time reconfiguration to increase the efficiency and adaptivity. Compiler support for automatic identification and implementation of ASIP CIs exists commercially and on research platforms, but these compilers do not support CIs with memory accesses, as ASIP CIs typically work on register file data. While being acceptable for ASIPs, this imposes a limitation for reconfigurable processors as they achieve their performance by exploiting data-level parallelism. Consequently, we propose a novel approach to CI identification for runtime reconfigurable processors with support for memory operations in contrast to previous works that explicitly exclude them. Our algorithm extracts memory access patterns which allows us to abstract from single memory operations and merge accesses to optimally utilize the available memory bandwidth. We implemented our algorithm in a state-of-the-art compiler framework.
Martin Haaß, Lars Bauer, Jörg Henkel
CASES3
2014 darkNoC: Designing Energy-Efficient Network-on-Chip with Multi-Vt Cells for Dark Silicon
abstract
In this paper, we propose a novel NoC architecture, called darkNoC, where multiple layers of architecturally identical, but physically different routers are integrated, leveraging the extra transistors available due to dark silicon. Each layer is separately optimized for a particular voltage-frequency range by the adroit use of multi-Vt circuit optimization. At a given time, only one of the network layers is illuminated while all the other network layers are dark. We provide architectural support for seamless integration of multiple network layers, and a fast inter-layer switching mechanism without dropping in-network packets. Our experiments on a 4 × 4 mesh with multi-programmed real application workloads show that darkNoC improves energy-delay product by up to 56% compared to a traditional single layer NoC with state-of-the-art DVFS. This illustrates darkNoC can be used as an energy-efficient communication fabric in future dark silicon chips.
Haseeb Bokhari, Haris Javaid, Muhammad Shafique 0001, Jörg Henkel, Sri Parameswaran
DAC4
2014 Reducing Latency in an SRAM/DRAM Cache Hierarchy via a Novel Tag-Cache Architecture
abstract
Memory speed has become a major performance bottleneck as more and more cores are integrated on a multi-core chip. The widening latency gap between high speed cores and memory has led to the evolution of multi-level SRAM/DRAM cache hierarchies that exploit the latency benefits of smaller caches (e.g. private L1 and L2 SRAM caches) and the capacity benefits of larger caches (e.g. shared L3 SRAM and shared L4 DRAM cache). The main problem of employing large L3/L4 caches is their high tag lookup latency. To solve this problem, we introduce the novel concept of small and low latency SRAM/DRAM Tag-Cache structures that can quickly determine whether an access to the large L3/L4 caches will be a hit or a miss. The performance of the proposed Tag-Cache architecture depends upon the Tag-Cache hit rate and to improve it we propose a novel Tag-Cache insertion policy and a DRAM row buffer mapping policy that reduce the latency of memory requests. For a 16-core system, this improves the average harmonic mean instruction per cycle throughput of latency sensitive applications by 13.3% compared to state-of-the-art.
Fazal Hameed, Lars Bauer, Jörg Henkel
DAC3
2014 CAP: Communication Aware Programming
abstract
Networks on Chip (NoC) come along with increased complexity from the implementation and management perspective. This leads to higher energy consumption and programming complexity of NoC architectures.
Jan Heisswolf, Aurang Zaib, Andreas Zwinkau, Sebastian Kobbe, Andreas Weichslgartner, Jürgen Teich, Jörg Henkel, Gregor Snelting, Andreas Herkersdorf, Jürgen Becker 0001
DAC7
2014 Multi-Layer Dependability: From Microarchitecture to Application Level
abstract
We show in this paper that multi-layer dependability is an indispensable way to cope with the increasing amount of technology-induced dependability problems that threaten to proceed further scaling. We introduce the definition of multi-layer dependability and present our design flow within this paradigm that seamlessly integrates techniques starting at circuit layer all the way up to application layer and thereby accounting for ASIC-based architectures as well as for reconfigurable-based architectures. At the end, we give evidence that the paradigm of multi-layer dependability bears a large potential for significantly increasing dependability at reasonable effort.
Jörg Henkel, Lars Bauer, Hongyan Zhang 0004, Semeen Rehman, Muhammad Shafique 0001
DAC1
2014 ASER: Adaptive Soft Error Resilience for Reliability-Heterogeneous Processors in the Dark Silicon Era
abstract
The Dark Silicon provides opportunities to realize Reliability-Heterogeneous Processors with ISA compatible cores having different levels of protection against reliability threats (like soft errors). This paper presents design-time customization of Reliability-Heterogeneous Processors given a set of applications and area constraints. A run-time system adaptively manages the soft error resilience under a given thermal design power (TDP) budget. We synthesize an embedded processor with different levels of protection and present area and power results for a 45nm technology. We illustrate the benefits of adaptive soft error resilience by comparing it with four different state-of-the-art approaches where we achieve 58%-96% overall system reliability improvements under a tight TDP constraint (corresponding to a 65% dark area).
Florian Kriebel, Semeen Rehman, Duo Sun, Muhammad Shafique 0001, Jörg Henkel
DAC5
2014 dTune: Leveraging Reliable Code Generation for Adaptive Dependability Tuning under Process Variation and Aging-Induced Effects
abstract
Designing dependable on-chip manycore systems is subjected to consideration of multiple reliability threats, i.e. soft errors, aging, process variation, etc. In this paper, we introduce a novel adaptive Dependability Tuning (dTune) scheme for many-core processors. It leverages the knowledge of varying vulnerability and error masking properties of different applications along with multiple compiled versions (each offering distinct reliability and performance properties). Our dTune system dynamically tunes the dependability mode at the hardware level through hybrid Redundant Multithreading tuning and at the software level through selection of reliable code version under given performance constraints. It jointly accounts for soft errors and cores' performance variations due to design-time process variation and/or run-time aging-induced performance degradation. We compare our dTune system with four different state-of-the-art techniques and achieve on average 44% and up to 63% improved task reliability for different chip configurations, different variability maps, and different aging years.
Semeen Rehman, Florian Kriebel, Duo Sun, Muhammad Shafique 0001, Jörg Henkel
DAC5
2014 The EDA Challenges in the Dark Silicon Era: Temperature, Reliability, and Variability Perspectives
abstract
Technology scaling has resulted in smaller and faster transistors in successive technology generations. However, transistor power consumption no longer scales commensurately with integration density and, consequently, it is projected that in future technology nodes it will only be possible to simultaneously power on a fraction of cores on a multi-core chip in order to stay within the power budget. The part of the chip that is powered off is referred to as dark silicon and brings new challenges as well as opportunities for the design community, particularly in the context of the interaction of dark silicon with thermal, reliability and variability concerns. In this perspectives paper we describe these new challenges and opportunities, and provide preliminary experimental evidence in their support.
Muhammad Shafique 0001, Siddharth Garg, Jörg Henkel, Diana Marculescu
DAC3
2014 GUARD: GUAranteed Reliability in Dynamically Reconfigurable Systems
abstract
Soft errors are a reliability threat for reconfigurable systems implemented with SRAM-based FPGAs. They can be handled through fault tolerance techniques like scrubbing and modular redundancy. However, selecting these techniques statically at design or compile time tends to be pessimistic and prohibits optimal adaptation to changing soft error rate at runtime.
Hongyan Zhang 0004, Michael A. Kochte, Michael E. Imhof, Lars Bauer, Hans-Joachim Wunderlich, Jörg Henkel
DAC6
2014 Software architecture of High Efficiency Video Coding for many-core systems with power-efficient workload balancing
abstract
The High Efficiency Video Coding (HEVC) standard aims at providing ~50% better compression compared to its predecessor (H.264) at the cost of high computational complexity. To enable HEVC video encoding in real-time scenarios, special coding support for parallelization is provided in HEVC that can be exploited by many-core systems. In this work, we present a HEVC software architecture where a video frame is adaptively divided into independent video frame regions (i.e. so-called video tiles) which are processed concurrently on multiple cores. By balancing the workload of each video tile mapped to a particular core, the total power consumption of a system is reduced (through dynamically scaling the operating frequency) under a given frame-rate constraint. We also exploit user tolerance to further curtail the HEVC workload with insignificant video quality degradation. Experimental results illustrate that the proposed approach results in ~43% power savings on a many-core system.
Muhammad Usman Karim Khan, Muhammad Shafique 0001, Jörg Henkel
DATE3
2014 mDTM: Multi-objective dynamic thermal management for on-chip systems
abstract
Thermal hot spots and unbalanced temperatures between cores on chip can cause either degradation in performance or may have a severe impact on reliability, or both. In this paper, we propose mDTM, a proactive dynamic thermal management technique for on-chip systems. It employs multi-objective management for migrating tasks in order to both prevent the system from hitting an undesirable thermal threshold and to balance the temperatures between the cores. Our evaluation on the Intel SCC platform shows that mDTM can successfully avoid a given thermal threshold and reduce spatial thermal variation by 22%. Compared to state-of-the-art, our mDTM achieves up to 58% performance gain. Additionally, we deploy an FPGA and IR camera based setup to analyze the effectiveness of our technique.
Heba Khdr, Thomas Ebi, Muhammad Shafique 0001, Hussam Amrouch, Jörg Henkel
DATE5
2014 hevcDTM: Application-driven Dynamic Thermal Management for High Efficiency Video Coding
abstract
This paper presents an application-driven algorithm for Dynamic Thermal Management (DTM) for the High Efficiency Video Coding (HEVC). For efficient design of such a DTM policy, we perform an offline thermal analysis of an HEVC encoder and demonstrate the impact of different video sequences and different coding configurations on the processor temperature. Our thermal analysis is leveraged to develop an efficient application-driven DTM policy that performs temperature-aware coding along with an application-driven control of DTM knobs (e.g., frequency scaling) in order to meet the temperature constraints while still providing high video quality (i.e. PSNR loss <; 0.01dB). For accurate thermal analysis and evaluation, we deploy an infrared camera-based thermal measurement setup that, on the contrary to state-of-the-art setups, does not require adding any extra layer on top of the measured chip, thus allowing the camera to accurately capture the infrared emissions from the die.
Daniel Palomino 0001, Muhammad Shafique 0001, Hussam Amrouch, Altamiro Amadeu Susin, Jörg Henkel
DATE5
2014 Compiler-driven dynamic reliability management for on-chip systems under variabilities
abstract
This paper presents a novel Dynamic Reliability Management System (DyReMS) for on-chip systems that performs resilience-driven resource allocation and mapping. It accounts for both the tasks' resilience properties and heterogeneous error recovery features of different cores. DyReMS also chooses a reliable task version (out of multiple reliability-aware transformed options) depending upon the reliability level of the allocated core. In case of error detection, rollbacks are performed. Our system provides 70%-87% improved task reliability compared to a timing reliability-optimizing core assignment, i.e. minimizing the probability of deadline misses (with EDF scheduling).
Semeen Rehman, Florian Kriebel, Muhammad Shafique 0001, Jörg Henkel
DATE4
2014 dSVM: Energy-efficient distributed Scratchpad Video Memory Architecture for the next-generation High Efficiency Video Coding
abstract
An energy-efficient distributed Scratchpad Video Memory Architecture (dSVM) for the next-generation parallel High Efficiency Video Coding is presented. Our dSVM combines private and overlapping (shared) Scratchpad Memories (SPMs) to support data reuse within and across different cores concurrently executing multiple parallel HEVC threads. We developed a statistical method to size and design the organization of the SPMs along with a supporting memory reading policy for energy efficiency. The key is to leverage the HEVC and video content knowledge. Furthermore, we integrate an adaptive power management policy for SPMs to manage the power states of different memory parts at run time depending upon the varying video content properties. Our experimental results illustrate that our dSVM architecture reduces the overall memory energy consumption by up to 51%-61% compared to parallelized state-of-the-art solutions [11]. The dSVM external memory energy savings increase with an increasing number of parallel HEVC threads and size of search window. Moreover, our SPM power management reacts to the current video properties and achieves up to 54% on-chip leakage energy savings.
Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel
DATE5
2014 Dependable task and communication migration in tiled manycore system-on-chip
abstract
Power densities and thermal hotspots are a major concern for the dependability of future multi-processor systemon- chip. They can lead to transient faults affecting the functionality in the short term and can cause permanent damage of a device. The dependability problem can be tackled on different layers such as technology hardening or application awareness. This work is based on an approach that addresses the issue for tile-based manycore system-on-chip on software and architecture layer. An agent-based system management employs task migration to react to thermal hotspots and pro-actively avoid them. The inter-task communication plays an important role as communication channels need to be migrated accordingly. The presented work focuses on the issue of communication migration and is based on the idea of handling it transparently to the task migration. Network-on-chip protection switching techniques have been introduced before and in this paper we evaluate the potential and bottlenecks of such methods in a realistic platform.
Stefan Wallentowitz, Stefan Rosch, Thomas Wild, Andreas Herkersdorf, Volker Wenzel, Jörg Henkel
FDL6
2014 MORP: makespan optimization for processors with an embedded reconfigurable fabric
abstract
Processors with an embedded runtime reconfigurable fabric have been explored in academia and industry started production of commercial platforms (e.g. Xilinx Zynq-7000). While providing significant performance and efficiency, the comparatively long reconfiguration time limits these advantages when applications request reconfigurations frequently. In multi-tasking systems frequent task switches lead to frequent reconfigurations and thus are a major hurdle for further performance increases. Sophisticated task scheduling is a very effective means to reduce the negative impact of these reconfiguration requests. In this paper, we propose an online approach for combined task scheduling and re-distribution of reconfigurable fabric between tasks in order to reduce the makespan, i.e. the completion time of a taskset that executes on a runtime reconfigurable processor. Evaluating multiple tasksets comprised of multimedia applications, our proposed approach achieves makespans that are on average only 2.8% worse than those achieved by a theoretical optimal scheduling that assumes zero-overhead reconfiguration time. In comparison, scheduling approaches deployed in state-of-the-art reconfigurable processors achieve makespans 14%-20% worse than optimal. As our approach is a purely software-side mechanism, a multitude of reconfigurable platforms aimed at multi-tasking can benefit from it.
Artjom Grudnitsky, Lars Bauer, Jörg Henkel
FPGA3
2014 Run-time accelerator binding for tile-based mixed-grained reconfigurable architectures
abstract
Run-time mixed-grained reconfigurable architectures emerged as an efficient solution to deal with the heterogeneous and at-design-time unpredictable nature of advanced applications. Due to interconnection limitations, the reconfigurable elements are grouped into tiles communicating through an on-chip network. State-of-the-art run-time accelerator binding schemes, i.e., mapping the accelerators to elements in the physical reconfigurable array, do not deal with such tile-based architectures. We propose a new scheme for run-time accelerator binding into our tile-based mixed-grained reconfigurable architecture. By means of an advanced video encoding application, we illustrate that our scheme reduces the inter-tile communication overhead by up to 44% (avg. 23%).
Cláudio Machado Diniz, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel
FPL4
2014 Towards interdependencies of aging mechanisms
Hussam Amrouch, Victor M. van Santen, Thomas Ebi, Volker Wenzel, Jörg Henkel
ICCAD5
2014 Energy-efficient architecture for advanced video memory
abstract
An energy-efficient hybrid on-chip video memory architecture (enHyV) is presented that combines private and shared memories using a hybrid design (i.e., SRAM and emerging STT-RAM). The key is to leverage the application-specific properties to efficiently design and manage the enHyV. To increase STT-RAM lifetime, we propose a data management technique that alleviates the bit-toggling write occurrences. An adaptive power management is also proposed for static-energy savings. Experimental results illustrate that enHyV reduces on-chip static memory energy compared to SRAM-only version of enHyV and to state-of-art AMBER hybrid video memory [9] by 66%-75% and 55%-76%, respectively. Furthermore, negligible external memory energy consumption is required for reference frames communication (98% lower than state-of-the-art Level C+ technique [18]). Our data management significantly improves the enHyV STT-RAM lifetime, achieving 0.83 of normalized lifetime (near to the optimal case). Our hybrid memory design and management incur low overhead in terms of latency and dynamic energy.
Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel
ICCAD5
2014 Fast hierarchical intra angular mode selection for high efficiency video coding
abstract
In this work, we address the complexity of the most time consuming module of High Efficiency Video Coding (HEVC) Intra-encoding, i.e. the Intra prediction generation. We reduce the computational complexity by estimating the candidates list (most probable Intra prediction modes), using content- and priority-driven gradient detection and hierarchically gathering the results of previous computations. This complexity reduction scheme is adaptive and can be controlled, depending upon the requirements, e.g. frame rate and video quality. On average, our scheme is capable of delivering 44% more time savings than the state-of-the-art scheme for fast Intra mode estimation.
Muhammad Usman Karim Khan, Muhammad Shafique 0001, Jörg Henkel
ICIP3
2014 Power efficient and workload balanced tiling for parallelized high efficiency video coding
abstract
The increased workload of the High Efficiency Video Coding (HEVC) and processing of high resolution videos require parallelization of the encoding/decoding process. However, to efficiently utilize the hardware resources and power budgets in a many-core processor, workload balanced parallelization of HEVC encoding is of high importance. Further, minimizing the number of active cores for processing the given HEVC encoding workload is required to decrease the power consumption. In order to address the above challenges, this work presents a HEVC parallelization technique to adaptively determine the Tile partitioning while accounting for the compute capabilities of the underlying processing cores. Afterwards, it determines a mapping of Tiled-HEVC processing on different cores such that the number of compute cores is minimized, and hence reducing the power consumption. Experimental results demonstrate that in addition to reducing the total compute cores, our technique provides up to 14.4% power savings compared to state-of-the-art uniform Tile partitioning approach.
Muhammad Shafique 0001, Muhammad Usman Karim Khan, Jörg Henkel
ICIP3
2014 Peak Power Management for scheduling real-time tasks on heterogeneous many-core systems
abstract
The number and diversity of cores in on-chip systems is increasing rapidly. However, due to the Thermal Design Power (TDP) constraint, it is not possible to continuously operate all cores at the same time. Exceeding the TDP constraint may activate the Dynamic Thermal Management (DTM) to ensure thermal stability. Such hardware based closed-loop safeguards pose a big challenge in using many-core chips for real-time tasks. Managing the worst-case peak power usage of a chip can help toward resolving this issue. We present a scheme to minimize the peak power usage for frame-based and periodic real-time tasks on many-core processors by scheduling the sleep cycles for each active core and introduce the concept of a sufficient test for peak power consumption for task feasibility. We consider both inter-task and inter-core diversity in terms of power usage and present computationally efficient algorithms for peak power minimization for these cases, i.e., a special case of “homogeneous tasks on homogeneous cores” to the general case of “heterogeneous tasks on heterogeneous cores”. We evaluate our solution through extensive simulations using the 48-core SCC platform and gem5 architecture simulator. Our simulation results show the efficacy of our scheme.
Waqaas Munawar, Heba Khdr, Santiago Pagani, Muhammad Shafique 0001, Jian-Jia Chen, Jörg Henkel
ICPADS6
2014 TONE: adaptive temperature optimization for the next generation video encoders
abstract
This paper presents an adaptive temperature optimization technique for the next generation video encoders. It exploits both application-specific knowledge (i.e. video encoding configurations) and video content properties in order to efficiently manage the temperature of advanced video coding systems at the software layer. For designing an efficient technique, we perform an extensive offline analysis to understand the impact of different video properties and configurations on the CPU thermal profiles when processing the next generation video encoder. Our temperature optimization technique performs an application-level prediction of the temperature trend followed by an application-level thermal management policy. The policy dynamically manages the temperature by performing an adaptive encoder configuration selection while providing minimum penalties in terms of bit rate and video quality. The experimental results show that our policy meets temperature constraints with negligible encoding performance loss. Moreover, when compared to state-of-the-art techniques, our policy provides a relatively reduced video quality loss while still meeting the temperature constraints.
Daniel Palomino 0001, Muhammad Shafique 0001, Altamiro Amadeu Susin, Jörg Henkel
ISLPED4
2014 Content-driven memory pressure balancing and video memory power management for parallel high efficiency video coding
abstract
We present a novel content-driven memory pressure balancing and video memory power management scheme for parallel High Efficiency Video Coding (HEVC). The key is to leverage the application-specific knowledge to balance the (instant) access pressure on Scratchpad-based Video Memories (SVMs) for parallelized video processing. Our scheme accurately predicts the memory requirements of each processing core based on monitored memory usage and leverages this knowledge to perform a categorization of different video regions. Afterwards, it employs an adaptive policy for memory pressure balancing by rescheduling encoding of different video blocks based on their categories. This balancing also facilitates our scheme to perform efficient power-gating of unused parts of SVMs. Experimental results show that our scheme reduces the variations in the memory pressure by 37%-83% when compared to the traditional raster scan processing for 4- and 16-core parallelized HEVC encoder. Our content-driven power management saves 56% (on average) of SVM leakage energy.
Felipe Sampaio, Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel
ISLPED5
2014 Messages from the conference chairs
abstract
Welcome to Chongqing, China, and the 20th IEEE International Conference on Embedded and Real- Time Computing Systems and Applications (RTCSA 2014). RTCSA has been a long-running technical conference sponsored by IEEE. The objective of the conference is to bring together academic researchers and industry developers for intensive discussion of recent advancing in the field of embedded systems, real-time systems, system design practice and emerging applications.
Edwin H.-M. Sha, Jörg Henkel, Kaijie Wu 0001, Tarek F. Abdelzaher, Hojung Cha
RTCSA2
2014 High-speed enoding/decoding technique for reliable data transmission in wireless sensor networks
abstract
Reliability has become one of the most vital requirements in wireless sensor networks (WSNs). One efficient way to increase the reliability is by using erasure coding techniques such as Reed-Solomon/Cauchy Reed-Solomon. The problem is the time required for encoding and decoding the data words, especially when large amount of data need to be sent like in complex WSN applications. In this paper, we propose HSC, a high-speed coding technique for reliable data transmission and we show that it reduces the number of XORs required for encoding by more than 94% (and consequently the encoding time) in comparison to the number of XORs required by Cauchy Reed-Solomon encoding. In addition, we present a new transmission scheme that works along with HSC to add redundancy on any size of a data load.
M. Sammer Srouji, Talal Bonny, Jörg Henkel
SECON3
2014 RESI: Register-Embedded Self-Immunity for Reliability Enhancement
abstract
Technology scaling in the nano-CMOS era has reached a point where coping with the failures produced by soft errors has become one of the key challenges when it comes to reliability. Akin to the fact that a register file is accessed more frequently than any other architectural component, register file protection is imperative to obstruct errors from propagating throughout a computing system. Furthermore, negative bias temperature instability (NBTI) has emerged as a major concern due to its negative impact on the lifetime of pMOS devices. Indeed, many of the pMOS transistors most affected by NBTI are in the register files as they are implemented as SRAM, which are particularly vulnerable due to their small structure size. Based on our observation that some register bits are not continuously used to represent a value stored in a register, we present a technique that exploits unused bits to improve the register file immunity against soft errors and mitigate NBTI effects. We show that our technique can reduce, on average, the register file vulnerability against multiple bit upsets by 97% (up to 100%), resulting in a high system fault coverage under various scenarios, while consuming less power and still occupying a similar area footprint compared to protecting the register file against single bit upsets (SBUs) only. The achieved result is 63% better compared to the state-of-the-art in register file protection. To compare and quantify the effect of our technique, we observe its impact on the processor's temperature using an infrared thermal camera and show that, due to consuming less power per area, our technique also operates at a lower temperature compared to protecting the register file against SBUs only. Finally, we investigate how our technique additionally moderates the stress induced by NBTI in register file SRAM cells.
Hussam Amrouch, Thomas Ebi, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2014 Energy-Efficient Adaptive Pipelined MPSoCs for Multimedia Applications
abstract
Pipelined MPSoCs provide a high throughput implementation platform for multimedia applications. They are typically balanced at design-time considering worst-case scenarios so that a given throughput can be fulfilled at all times. Such worst-case pipelined MPSoCs lack runtime adaptability and result in inefficient resource utilization and high power/energy consumption under a dynamic workload. In this paper, we propose a novel adaptive architecture and a distributed runtime processor manager to enable runtime adaptation in pipelined MPSoCs. The proposed architecture consists of main processors and auxiliary processors, where a main processor uses differing number of auxiliary processors considering runtime workload variations. The runtime processor manager uses a combination of application's execution and knowledge, and offline profiling and statistical information to proactively predict the auxiliary processors that should be used by a main processor. The idle auxiliary processors are then deactivated using clock- or power-gating. Each main processor with a pool of auxiliary processors has its own runtime manager, which is independent of the other main processors, enabling a distributed runtime manager. Our experiments with an H.264 video encoder for HD720p resolution at 30 frames/s show that the adaptive pipelined MPSoC consumed up to 29% less energy (computed using processors and caches) than a worst-case pipelined MPSoC, while delivering a minimum of 28.75 frames/s. Our results show that adaptive pipelined MPSoCs can emerge as an energy-efficient implementation platform for advanced multimedia applications.
Haris Javaid, Muhammad Shafique 0001, Jörg Henkel, Sri Parameswaran
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2014 Reliability-Driven Software Transformations for Unreliable Hardware
abstract
We propose multiple reliability-driven software transformations targeting unreliable hardware. These transformations reduce the executions of critical instructions and spatial/temporal vulnerabilities of different instructions with respect to different processor components. The goal is to lower the application's susceptibility toward failures. Compared to performance-optimized compilation, our method incurs 60% lower application failures, averaged over various fault injection scenarios and fault rates.
Semeen Rehman, Florian Kriebel, Muhammad Shafique 0001, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2014 Adaptive Energy Management for Dynamically Reconfigurable Processors
abstract
We present an adaptive energy management system for dynamically reconfigurable processors that chooses an energy-minimizing set of custom instructions (CIs) and then power-gates the temporarily unused subset of CIs. It requires a comprehensive power model to estimate the power consumption of different CIs at run time. We deploy our new energy management in two state-of-the-art reconfigurable processors (RISPP and Molen) and perform an elaborative evaluation of energy savings under various area and performance constraints for different technology nodes. We demonstrate the energy benefits by comparing it to state-of-the-art power-gating techniques for FPGAs. The work is implemented as a prototype on a Xilinx FPGA platform.
Muhammad Shafique 0001, Lars Bauer, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2013 Thermal management for dependable on-chip systems
abstract
Dependability has become a growing concern in the nano-CMOS era due to elevated temperatures and an increased susceptibility to temperature of the small structures. We present an overview of temperature-related effects that threaten dependability and a methodology for reducing the dependability concerns through thermal management utilizing the concept of aging budgeting.
Jörg Henkel, Thomas Ebi, Hussam Amrouch, Heba Khdr
ASP-DAC1
2013 Simultaneously optimizing DRAM cache hit latency and miss rate via novel set mapping policies
abstract
Two key parameters that determine the performance of a DRAM cache based multi-core system are DRAM cache hit latency (HL) and DRAM cache miss rate (MR), as they strongly influence the average DRAM cache access latency. Recently proposed DRAM set mapping policies are either optimized for HL or for MR. None of these policies provides a good HL and MR at the same time. This paper presents a novel DRAM set mapping policy that simultaneously targets both parameters with the goal of achieving the best of both to reduce the overall DRAM cache access latency. For a 16-core system, our proposed set mapping policy reduces the average DRAM cache access latency (depends upon HL and MR) compared to state-of-the-art DRAM set mapping policies that are optimized for either HL or MR by 29.3% and 12.1%, respectively.
Fazal Hameed, Lars Bauer, Jörg Henkel
CASES3
2013 Hardware acceleration for programs in SSA form
abstract
Register allocation is one of the most time-consuming parts of the compilation process. Depending on the quality of the register allocation, a large amount of shuffle code to move values between registers is generated. In this paper, we propose a processor architecture extension to provide register file permutations by which the shuffle code can be implemented more efficiently. We present compiler support to utilize this extension, an evaluation regarding performance and compilation time using the SPEC CINT2000 benchmark, as well as an analysis of area and frequency overhead of our architecture implementation. We find that using our extension, the number of executed instructions is reduced by up to 5.1 % while the compilation time is unaffected.
Manuel Mohr, Artjom Grudnitsky, Tobias Modschiedler, Lars Bauer, Sebastian Hack, Jörg Henkel
CASES6
2013 Reliable on-chip systems in the nano-era: lessons learnt and future trends
abstract
Reliability concerns due to technology scaling have been a major focus of researchers and designers for several technology nodes. Therefore, many new techniques for enhancing and optimizing reliability have emerged particularly within the last five to ten years. This perspective paper introduces the most prominent reliability concerns from today's points of view and roughly recapitulates the progress in the community so far. The focus of this paper is on perspective trends from the industrial as well as academic points of view that suggest a way for coping with reliability challenges in upcoming technology nodes.
Jörg Henkel, Lars Bauer, Nikil Dutt, Puneet Gupta 0001, Sani R. Nassif, Muhammad Shafique 0001, Mehdi Baradaran Tahoori, Norbert Wehn
DAC1
2013 Optimizations for configuring and mapping software pipelines in many core systems
abstract
Efficiently utilizing the computational resources of many core systems is one of the most prominent challenges. The problem worsens when resource requirements vary unpredictably and applications may be started/stopped at any time. To address this challenge, we propose two schemes that calculate and adapt task mappings at runtime: a centralized, optimal mapping scheme and a distributed, hierarchical mapping scheme that trades optimality for a high degree of scalability. Experiments on Intel's 48-core Single-Chip Cloud Computer and in a many core simulator show that a significant improvement in system performance can be achieved over current state-of-the-art.
Janmartin Jahn, Santiago Pagani, Sebastian Kobbe, Jian-Jia Chen, Jörg Henkel
DAC5
2013 RASTER: runtime adaptive spatial/temporal error resiliency for embedded processors
abstract
Applying error recovery monotonously can either compromise the real-time constraint, or worsen the power/energy envelope. Neither of these violations can be realistically accepted in embedded system design, which expects ultra efficient realization of a given application. In this paper, we propose a HW/SW methodology that exploits both application specific characteristics and Spatial/Temporal redundancy. Our methodology combines design-time and runtime optimizations, to enable the resultant embedded processor to perform runtime adaptive error recovery operations, precisely targeting the reliability-wise critical instruction executions. The proposed error recovery functionality can dynamically 1) evaluate the reliability cost economy (in terms of execution-time and dynamic power), 2) determine the most profitable scheme, and 3) adapt to the corresponding error recovery scheme, which is composed of spatial and temporal redundancy based error recovery operations. The experimental results have shown that our methodology at best can achieve fifty times greater reliability while maintaining the execution time and power deadlines, when compared to the state of the art.
Tuo Li 0001, Muhammad Shafique 0001, Jude Angelo Ambrose, Semeen Rehman, Jörg Henkel, Sri Parameswaran
DAC5
2013 Exploiting program-level masking and error propagation for constrained reliability optimization
abstract
Since embedded systems design involves stringent design constraints, designing a system for reliability requires optimization under tolerable overhead constraints. This paper presents a novel reliability-driven compilation scheme for software program reliability optimization under tolerable overhead constraints. Our scheme exploits program-level error masking and propagation properties to perform reliability-driven prioritization of instructions and selective protection during compilation. To enable this, we develop statistical models for estimating error masking and propagation probabilities. Our scheme provides significant improvement in reliability efficiency (avg. 30%-60%) compared to state-of-the-art program-level protection schemes.
Muhammad Shafique 0001, Semeen Rehman, Pau Vilimelis Aceituno, Jörg Henkel
DAC4
2013 Mapping on multi/many-core systems: survey of current and emerging trends
abstract
The reliance on multi/many-core systems to satisfy the high performance requirement of complex embedded software applications is increasing. This necessitates the need to realize efficient mapping methodologies for such complex computing platforms. This paper provides an extensive survey and categorization of state-of-the-art mapping methodologies and highlights the emerging trends for multi/many-core systems. The methodologies aim at optimizing system's resource usage, performance, power consumption, temperature distribution and reliability for varying application models. The methodologies perform design-time and run-time optimization for static and dynamic workload scenarios, respectively. These optimizations are necessary to fulfill the end-user demands. Comparison of the methodologies based on their optimization aim has been provided. The trend followed by the methodologies and open research challenges have also been discussed.
Amit Kumar Singh 0002, Muhammad Shafique 0001, Akash Kumar 0001, Jörg Henkel
DAC4
2013 Adaptive cache management for a combined SRAM and DRAM cache hierarchy for multi-cores
abstract
On-chip DRAM caches may alleviate the memory bandwidth problem in future multi-core architectures through reducing off-chip accesses via increased cache capacity. For memory intensive applications, recent research has demonstrated the benefits of introducing high capacity on-chip L4-DRAM as Last-Level-Cache between L3-SRAM and off-chip memory. These multi-core cache hierarchies attempt to exploit the latency benefits of L3-SRAM and capacity benefits of L4-DRAM caches. However, not taking into consideration the cache access patterns of complex applications can cause inter-core DRAM interference and inter-core cache contention. In this paper, we contest to re-architect existing cache hierarchies by proposing a hybrid cache architecture, where the Last-Level-Cache is a combination of SRAM and DRAM caches. We propose an adaptive DRAM placement policy in response to the diverse requirements of complex applications with different cache access behaviors. It reduces inter-core DRAM interference and inter-core cache contention in SRAM/DRAM-based hybrid cache architectures: increasing the harmonic mean instruction-per-cycle throughput by 23.3% (max. 56%) and 13.3% (max. 35.1%) compared to state-of-the-art.
Fazal Hameed, Lars Bauer, Jörg Henkel
DATE3
2013 DANCE: distributed application-aware node configuration engine in shared reconfigurable sensor networks
abstract
Wireless sensor networks (WSNs) are often primarily tailored to single applications to achieve one specific mission. Considering that the same physical phenomenon can be used by multiple applications, the benefit of sharing the WSN infrastructure is obvious in terms of development and deployment cost. However, allocating the tasks to the WSNs to meet the requirements of all applications while keeping the energy efficiency is very challenging. Introducing reconfigurable nodes in the shared sensor networks can improve the performance, the energy efficiency and the flexibility but it increases the system complexity. In this paper, we propose a biologically inspired node configuration scheme in shared reconfigurable sensor network named DANCE, which can adapt to the changing environment and efficiently utilize WSN resources. Our experiments show that our scheme reduces the energy consumption by up to 76%.
Chih-Ming Hsieh, Zhonglei Wang, Jörg Henkel
DATE3
2013 Pipelets: self-organizing software pipelines for many-core architectures
abstract
We present the novel concept of Pipelets: self-organizing stages of software pipelines that monitor their computational demands and communication patterns and interact to optimize the performance of the application they belong to. They enable dynamic task remapping and exploit application-specific properties. Our experiments show that they improve performance by up to 31.2% compared to state-of-the-art when resource demands of applications alter at runtime as is the case for many complex applications.
Janmartin Jahn, Jörg Henkel
DATE2
2013 An H.264 Quad-FullHD low-latency intra video encoder
abstract
Video applications are moving from Full-HD capability (1920×1080) to even higher resolutions such as Quad-FullHD (3840×2160). The H.264 Intra-mode can be used by embedded devices to trade off the better encoding efficiency of H.264 temporal prediction (Inter-mode) against savings in area and power as well as saving the massive computational overhead of the sub-pixel motion estimation by using only spatial prediction (Intra-mode). Still, the H.264 Intra-mode requires a large computational effort and imposes severe challenges when targeting Quad-FullHD 25 fps real-time video encoding at moderate operating frequencies (we target 150 MHz) and limited area budget. Therefore, in this work we address the strong sequential data dependencies within H.264 Intra-mode that restrict the parallelism and inhibit high resolution encoding by a) decoupling of DC and AC transform paths, b) cycle-budget aware mode prediction scheduling while c) being area efficient. Using our proposed techniques, Quad-FullHD (3840×2160) 28 fps video encoding is achieved at 150 MHz, making our architecture applicable for high definition recording.
Muhammad Usman Karim Khan, Jan Micha Borrmann, Lars Bauer, Muhammad Shafique 0001, Jörg Henkel
DATE5
2013 Hardware-software collaborative complexity reduction scheme for the emerging HEVC intra encoder
abstract
High Efficiency Video Coding (HEVC/H.265) is an emerging standard for video compression that provides almost double compression efficiency at the cost of major computational complexity increase as compared to current industry-standard Advanced Video Coding (AVC/H.264). This work proposes a collaborative hardware and software scheme for complexity reduction in an HEVC Intra encoding system, with run-time adaptivity. Our scheme leverages video content properties which drive the complexity management layer (software) to generate a highly probable coding configuration. The intra prediction size and direction are estimated for the prediction unit which provides reduced computational-complexity. At the hardware layer, specialized coprocessors with enhanced reusability are employed as accelerators. Additionally, depending upon the video properties, the software layer administers the energy management of the hardware coprocessors. Experimental results show that a complexity reduction of up to 60 % and the energy reduction up to 42 % are achieved.
Muhammad Usman Karim Khan, Muhammad Shafique 0001, Mateus Grellert, Jörg Henkel
DATE4
2013 CSER: HW/SW configurable soft-error resiliency for application specific instruction-set processors
abstract
Soft error has been identified as one of the major challenges to CMOS technology based computing systems. To mitigate this problem, error recovery is a key component, which usually accounts for a substantial cost, since they must introduce redundancies in either time or space. Consequently, using state-of-art recovery techniques could heavily worsen the design constraint, which is fairly stringent for embedded system design. In this paper, we propose a HW/SW methodology that generates the processor, which performs finely configured error recovery functionality targeting the given design constraints (e.g., performance, area and power). Our methodology employs three application-specific optimization heuristics, which generate the optimized composition and configuration based on the two primitive error recovery techniques. The resultant processor is composed of selected primitive techniques at corresponding instruction execution, and configured to perform error recovery at run-time accordingly to the scheme determined at design time. The experiment results have shown that our methodology can at best achieve nine times reliability while maintaining the given constraints, in comparison to the state of the art.
Tuo Li 0001, Muhammad Shafique 0001, Semeen Rehman, Swarnalatha Radhakrishnan, Roshan G. Ragel, Jude Angelo Ambrose, Jörg Henkel, Sri Parameswaran
DATE7
2013 Leveraging variable function resilience for selective software reliability on unreliable hardware
abstract
State-of-the-art reliability optimizing schemes deploy spatial or temporal redundancy for the complete functionality. This introduces significant performance/area overhead which is often prohibitive within the stringent design constraints of embedded systems. This paper presents a novel scheme for selective software reliability optimization constraint under user-provided tolerable performance overhead constraint. To enable this scheme, statistical models for quantifying software resilience and error masking properties at function and instruction level are proposed. These models leverage a whole new range of reliability optimization. Given a tolerable performance overhead, our scheme selectively protects the reliability-wise most important instructions based on their masking probability, vulnerability, and redundancy overhead. Compared to state-of-the-art [7], our scheme provides a 4.84X improved reliability at 50% tolerable performance overhead constraint.
Semeen Rehman, Muhammad Shafique 0001, Pau Vilimelis Aceituno, Florian Kriebel, Jian-Jia Chen, Jörg Henkel
DATE6
2013 Energy-efficient memory hierarchy for motion and disparity estimation in multiview video coding
abstract
This work presents an energy-efficient memory hierarchy for Motion and Disparity Estimation on Multiview Video Coding employing a Reference Frames-Centered Data Reuse (RCDR) scheme. In RCDR the reference search window becomes the center of the motion/disparity estimation processing flow and calls for processing all blocks requesting its data. By doing so, RCDR avoids multiple search window retransmissions leading to reduced number of external memory accesses, thus memory energy reduction. To deal with out-of-order processing and further reduce external memory traffic, a statistics-based partial results compressor is developed. The on-chip video memory energy is reduced by employing a statistical power gating scheme and candidate blocks reordering. Experimental results show that our reference-centered memory hierarchy outperforms the state-of-the-art [7][13] by providing reduction of up to 71% for external memory energy, 88% on-chip memory static energy, and 65% on-chip memory dynamic energy.
Felipe Sampaio, Bruno Zatt, Muhammad Shafique 0001, Luciano Volcan Agostini, Sergio Bampi, Jörg Henkel
DATE6
2013 Self-adaptive hybrid dynamic power management for many-core systems
abstract
We present a self-adaptive, hybrid Dynamic Power Management (DPM) scheme for many-core systems that targets concurrently executing applications with what we call “expanding” and “shrinking” resource allocations as, for example, in [13]–[15] [27]. To avoid frequent allocation and de-allocation, it enables applications to temporarily reserve their resources and to perform local power management decisions. The expand-to-shrink time periods and resource demands are predicted on-the-fly based on the application-specific knowledge and the monitored system information. Experimental results demonstrate up to 15%-40% Energy-Delay2Product reduction of our scheme compared to state-of-the-art power management schemes like [4][8]. Self-adaptive local power-management decisions make our scheme scalable for large-scaled many-core systems as illustrated by numerous experiments.
Muhammad Shafique 0001, Benjamin Vogel, Jörg Henkel
DATE3
2013 Fast and accurate cache modeling in source-level simulation of embedded software
abstract
Recently, source-level software models are increasingly used for software simulation in TLM (Transaction Level Modeling)-based virtual prototypes of multicore systems. A source-level model is generated by annotating timing information into application source code and allows for very fast software simulation. Accurate cache simulation is a key issue in multicore systems design because the memory subsystem accounts for a large portion of system performance. However, cache simulation at source level faces two major problems: (1) as target data addresses cannot be statically resolved during source code instrumentation, accurate data cache simulation is very difficult at source level, and (2) cache simulation brings large overhead in simulation performance and therefore cancels the gain of source level simulation. In this paper, we present a novel approach for accurate data cache simulation at source level. In addition, we also propose a cache modeling method to accelerate both instruction and data cache simulation. Our experiments show that simulation with the fast cache model achieves 450.7 MIPS (million simulated instructions per second) on a standard x86 laptop, 2.3x speedup compared with a standard cache model. The source-level models with cache simulation achieve accuracy comparable to an Instruction Set Simulator (ISS). We also use a complex multimedia application to demonstrate the efficiency of the proposed approach for multicore systems design.
Zhonglei Wang, Jörg Henkel
DATE2
2013 Stress balancing to mitigate NBTI effects in register files
abstract
Negative Bias Temperature Instability (NBTI) is considered one of the major reliability concerns of transistors in current and upcoming technology nodes and a main cause of their diminished lifetime. We propose a new means to mitigate the effects of NBTI on SRAM-based register files, which are particularly vulnerable due to their small structure size and are under continuous voltage stress for prolonged intervals. The conducted results from our technology simulator demonstrate the severity of NBTI effects on the SRAM cells - especially when process variation is taken into account. Based on the presented analysis, we show that NBTI stress in different registers needs to be tackled using different strategies corresponding to their access patterns. To this end, we propose to selectively increase the resilience of individual registers against NBTI. Our technique balances the gate voltage stress of the two PMOS transistors of an SRAM cell such that both are under stress for approximately the same amount of time during operation - thereby minimizing the deleterious effects of NBTI. We present mitigation implementations in both hardware and in software along with the incurred overhead. Through a wide range of applications we can show that our technique reduces the NBTI-induced reliability degradation by 35% on average. This is 22% better than current State-of-the-Art.
Hussam Amrouch, Thomas Ebi, Jörg Henkel
DSN3
2013 Accurate Thermal-Profile Estimation and Validation for FPGA-Mapped Circuits
abstract
Accurate thermal profile estimation for FPGA, at design time, is necessary to avoid unexpected thermal hot-spots in the circuit before deploying the FPGA to the in-field operation. Both accurate dynamic and leakage power values are needed for the thermal profile estimation and they can be estimated using the FPGA vendor's tools. However these report leakage power as a single value for the whole chip, and no details are given in literature or the FPGA toolset about its distribution across the FPGA chip for the thermal simulation. To cope with this problem, we present a method for properly distributing the leakage power across the FPGA chip. The method uses a temperature-leakage loop estimation model for distributing and adapting the leakage power for more accurate thermal simulation. Furthermore, to accurately calibrate the presented method and its model and also to validate the resulting thermal profiles, we utilize an infrared thermal camera, which measures the emissions from the backside of a Virtex-5 FPGA chip. The results of testing several designs, with different sizes and frequencies, show that our approach can achieve accurate thermal-profile estimation when compared to the camera measurements, with average absolute estimation error of around 1°C across the chip.
Abdulazim Amouri, Hussam Amrouch, Thomas Ebi, Jörg Henkel, Mehdi Baradaran Tahoori
FCCM4
2013 Analyzing the thermal hotspots in FPGA-based embedded systems
abstract
The rapid push towards the minimization of the feature sizes of the process technology nodes in the nano-CMOS era has significantly increased the power densities and made State-of-the-Art FPGAs vulnerable to diverse problems induced by excessive temperatures. As a result, there is a prominent need to accurately study the FPGA thermal characteristics. Our experimental setup employs a thermal camera that captures the infrared emissions from the silicon wafer of an FPGA die allowing us to evaluate the accuracy of the methods conventionally used for thermal analysis, such as thermal simulations. Based on our observation that the memory interface is the thermal hotspot of an FPGA-based embedded system, we demonstrate that the cache plays a dominant thermal role by reducing memory accesses; we carefully examine the influence of various cache parameters on the FPGA temperature and propose a model linking these quantities, with an average maximum error of only 0.64°C.
Hussam Amrouch, Thomas Ebi, Josef Schneider, Sri Parameswaran, Jörg Henkel
FPL5
2013 DHASER: dynamic heterogeneous adaptation for soft-error resiliency in ASIP-based multi-core systems
abstract
Soft error has become a major adverse effect in CMOS based electronic systems. Mitigating soft error requires enhancing the underlying system with error recovery functionality, which typically leads to considerable design cost overhead, in terms of performance, power and area. For embedded systems, where stringent design constraints apply, such cost must be properly bounded. In this paper, we propose a HW/SW methodology DHASER, which enables efficient error recovery functionality for embedded ASIP-based multi-core systems. DHASER consists of three main parts: task level correctness (TLC) analysis, TLC-based processor/core customization, and runtime reliability-aware task management mechanism. It enables each individual ASIP-based processing core to dynamically adapt its specific error recovery functionality according to the corresponding task's characteristics (i.e., soft error vulnerability and execution time deadline). The goal is to optimize the overall system reliability while considering performance/throughput. The experimental results have shown that DHASER can significantly improve the reliability of the system, with little cost overhead, in comparison to the state-of-art counterparts.
Tuo Li 0001, Muhammad Shafique 0001, Semeen Rehman, Jude Angelo Ambrose, Jörg Henkel, Sri Parameswaran
ICCAD5
2013 ISOMER: integrated selection, partitioning, and placement methodology for reconfigurable architectures
abstract
Quality system design on dynamic partially reconfigurable platform needs exploration of a vast and multidimensional design space for (1) selection among implementation variants of hardware accelerators, (2) partitioning the reconfigurable fabric, and (3) their placement on the reconfigurable fabric partitions. This paper presents a novel methodology ISOMER for integrated solution of selection, partitioning and placement for performance optimization. Architecture under consideration is a general purpose processor coupled with reconfigurable fabric that can be partitioned in multi-sized partially reconfigurable bins. Our methodology determines performance-efficient partitioning and usage of reconfigurable fabric. Extensive evaluation illustrates that our methodology is scalable and outperforms state-of-the-art techniques for non-partially reconfigurable architectures.
Rana Muhammad Bilal, Rehan Hafiz, Muhammad Shafique 0001, Saad Shoaib, Asim Munawar, Jörg Henkel
ICCAD6
2013 Formal verification of distributed dynamic thermal management
abstract
Simulation is the state-of-the-art analysis technique for distributed thermal management schemes. Due to the numerous parameters involved and the distributed nature of these schemes, such non-exhaustive verification may fail to catch functional bugs in the algorithm or may report misleading performance characteristics. To overcome these limitations, we propose a methodology to perform formal verification of distributed dynamic thermal management for many-core systems. The proposed methodology is based on the SPIN model checker and the Lamport timestamps algorithm. Our methodology allows specification and verification of both functional and timing properties in a distributed many-core system. In order to illustrate the applicability and benefits of our methodology, we perform a case study on a state-of-the-art agent-based distributed thermal management scheme.
Osman Hasan, Thomas Ebi, Muhammad Shafique 0001, Jörg Henkel
ICCAD5
2013 MOMA: mapping of memory-intensive software-pipelined applications for systems with multiple memory controllers
abstract
In many-core systems, the efficient deployment of computational and other resources is key in order to achieve a high throughput. Current state-of-the-art task mapping schemes balance the computational load among cores while avoiding congestions within the communication links. The problem is that a large number of cores running many memory-intensive tasks may congest memory controllers because their number and bandwidth is constrained. To avoid a high throughput degradation that could result from congested memory controllers, the mapping of tasks must be sensitized to the limited bandwidth of off-chip memory. Designing efficient and effective algorithms to optimize the throughput by jointly considering the load of memory controllers, computation, and communication is very challenging. In this paper, we address this problem by distributing cores among applications and then heuristically map tasks such that the load of the memory controllers is sufficiently balanced. Our heuristic also minimizes the effect of decreased throughput resulting from mapping communicating tasks to cores that belong to different controllers. Our experiments encourage us in that we can reduce the saturation of memory controllers and significantly increase the system throughput compared to employing several state-of-the-art task mapping schemes.
Janmartin Jahn, Santiago Pagani, Jian-Jia Chen, Jörg Henkel
ICCAD4
2013 AMBER: adaptive energy management for on-chip hybrid video memories
abstract
The ever increasing leakage power of memories in a system has motivated researches for exploiting unconventional memory architectures. Non-Volatile Memory (NVM) used in conjunction with the conventional on-chip SRAMs has given birth to the hybrid memory paradigm, which can be intelligently exploited to reduce the energy consumption while tackling the high read and write latencies of NVMs. We present a novel scheme AMBER that aims at minimizing the total memory energy consumption of a video processing system by leveraging the application-specific properties and distinct latency and power properties of different memory types. AMBER also features architectural support for data-fetching from external memory and adaptively filling the different on-chip memories. We employ AMBER in the next-generation High Efficiency Video Coding (HEVC) standard to minimize the energy consumption of the new complex motion prediction process. Experimental results demonstrate that our AMBER scheme achieves significant energy savings (average 43%) for the on-chip memory.
Muhammad Usman Karim Khan, Muhammad Shafique 0001, Jörg Henkel
ICCAD3
2013 Agent-based distributed power management for kilo-core processors
abstract
Power management for Kilo-core processors have become an intricate problem due to the scalability issues and mixed-workloads of massively multi-threaded applications. This paper highlights the power related issues in Kilo-core processors and presents two emerging trends towards agent-based distributed and self-adaptive power management for Kilo-core processors. Agent-based power management allows applications to autonomously control the power states of their resources while operate efficiently as a whole to improve the overall system's energy efficiency. The first approach based on our concept of virtual power gating that allows applications to temporarily reserve their resources to locally optimize for power efficiency. The second approach is game-theoretic power management to achieve fair resource allocations while maximizing the energy efficiency. We present results for scalability and energy efficiency.
Muhammad Shafique 0001, Jörg Henkel
ICCAD2
2013 High-throughput interpolation hardware architecture with coarse-grained reconfigurable datapaths for HEVC
abstract
Fractional-pel interpolation for motion estimation and motion compensation is one of the key computational hotspots in the new High Efficient Video Coding (HEVC) standard. This work presents a high-throughput interpolation hardware architecture to improve performance of HEVC encoding and decoding. It employs two acceleration engines for luma and chroma filtering, each with 12-pel-parallel coarse-grained reconfigurable interpolation datapaths. An adaptive scheduling scheme manages the operation of these interpolation datapaths in different ways depending upon the prediction unit (PU) size and the execution scenario (i.e. motion estimation or motion compensation). We have implemented our hardware architecture in 150 nm technology. Compared to state-of-the-art techniques [12], our architecture required 49% less hardware area, while processing QFHD (3840×2160) resolution @ 30 fps.
Cláudio Machado Diniz, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel
ICIP4
2013 An adaptive workload management scheme for HEVC encoding
abstract
Managing the complexity of the emerging HEVC standard is a matter of academic and industrial research since its earlier versions. The sophisticated and computation-intensive tools involved in the encoding process must be leveraged if real-time applications are considered. In this paper, we propose a workload management scheme for dynamically controlling the computational complexity of HEVC, under user-defined operation frequency and target FPS. Our scheme receives these two parameters as input and aims to meet the target FPS by adjusting different encoding parameters during execution time. Experiments demonstrate that our scheme successfully meets the target FPS while introducing negligible rate-distortion losses. A comparison with state-of-the-art shows that our scheme is capable of achieving a time reduction of up to 43% for Full HD sequences, with a maximum loss of 0.03 dB in Y-PSNR and a 3.5% increase in bitrate.
Mateus Grellert, Muhammad Shafique 0001, Muhammad Usman Karim Khan, Luciano Volcan Agostini, Júlio C. B. de Mattos, Jörg Henkel
ICIP6
2013 An adaptive complexity reduction scheme with fast prediction unit decision for HEVC intra encoding
abstract
The next-generation High Efficiency Video Coding (HEVC) standard aims at providing double compression compared to the state-of-the-art H.264/AVC standard. However, this improved compression efficiency accompanies high computational complexity, which is primarily due to the recursive nature of the Coding Tree Unit structure and complex Rate-Distortion (RD) Optimization of Prediction Units (PU). In this paper, we propose a content-driven adaptive complexity reduction scheme that employs an a priori PU size selection algorithm. We adaptively combine smaller PUs into larger PUs depending upon the video frame contents under consideration. Our scheme eliminates the need for recursive or iterative RD cost comparisons to select the best PU size for an HEVC intra encoder. Experimental results demonstrate that the proposed scheme provides significant time savings (on average 43.74 %) for HEVC intra encoding, while incurring an insignificant video quality loss (on average -0.048 dB BD-PSNR).
Muhammad Usman Karim Khan, Muhammad Shafique 0001, Jörg Henkel
ICIP3
2013 Content-adaptive reference frame compression based on intra-frame prediction for multiview video coding
abstract
This paper presents a content-adaptive reference frame compression scheme to alleviate the large overhead of external memory communication during the Motion and Disparity Estimation process in Multiview Video Coding (MVC). Our scheme is based on a simplified intra-prediction process to reduce the spatial redundancy of the reference samples. The intra-prediction residue is compressed by a path composed of non-linear quantization and Huffman-based entropy encoder. Four different quantization strengths and Huffman tables were statistically defined. They are dynamically selected according to a content adaptation strategy, which classifies the original blocks based on their spatial homogeneity. Experimental results show that the proposed content-adaptive compression scheme is able to reduce the external memory accesses by up to 63% along with negligible losses in the MVC encoder rate-distortion performance. Compared to the best available related work [12] our content-adaptive reference frame compression achieves 39% reduced external memory accesses, while still providing a BD-PSNR increase of 0.03dB.
Felipe Sampaio, Bruno Zatt, Muhammad Shafique 0001, Luciano Volcan Agostini, Jörg Henkel, Sergio Bampi
ICIP5
2013 Content-driven adaptive computation offloading for energy-aware hybrid distributed video coding
abstract
Hybrid Distributed Video Coding (Hybrid-DVC) is an emerging compression paradigm for resource-constrained embedded devices, like wireless video sensor nodes. This paper presents a low-overhead energy-aware Hybrid-DVC with an adaptive computation offloading scheme. Our scheme offloads the workload of multiple video encoding devices concurrently to a high-end decoder device to minimize the overall system energy under constraints of compute capabilities. It leverages video content knowledge to determine regions in the video frames that are offloaded to the remote decoder device. It jointly minimizes the transmission and communication energy and provides a high video quality. Our scheme provides up to 39% energy savings compared to state-of-the-art Hybrid-DVC schemes.
Muhammad Shafique 0001, Muhammad Usman Karim Khan, Jörg Henkel
ISLPED3
2013 Module diversification: Fault tolerance and aging mitigation for runtime reconfigurable architectures
abstract
Runtime reconfigurable architectures based on Field-Programmable Gate Arrays (FPGAs) are attractive for realizing complex applications. However, being manufactured in latest semiconductor process technologies, FPGAs are increasingly prone to aging effects, which reduce the reliability of such systems and must be tackled by aging mitigation and application of fault tolerance techniques. This paper presents module diversification, a novel design method that creates different configurations for runtime reconfigurable modules. Our method provides fault tolerance by creating the minimal number of configurations such that for any faulty Configurable Logic Block (CLB) there is at least one configuration that does not use that CLB. Additionally, we determine the fraction of time that each configuration should be used to balance the stress and to mitigate the aging process in FPGA-based runtime reconfigurable systems. The generated configurations significantly improve reliability by fault-tolerance and aging mitigation.
Hongyan Zhang 0004, Lars Bauer, Michael A. Kochte, Eric Schneider, Claus Braun, Michael E. Imhof, Hans-Joachim Wunderlich, Jörg Henkel
ITC8
2013 Fast HEVC intra mode decision algorithm based on new evaluation order in the Coding Tree Block
abstract
This paper presents a fast mode decision algorithm for the HEVC intra prediction. A new evaluation order in the Coding Tree Block (CTB) allows the use of modes from low level PUs to be used as reference to the current PU decision. In this paper we use this idea to develop a fast intra mode decision algorithm that can be configured to run in two different complexity modes, relaxed and aggressive. Experimental results have shown that our algorithm achieved encoding time savings of almost 60% with negligible loss in the compression efficiency when compared to the full RDO based decision. Besides, our mode decision algorithm presented the best result in terms of time saving per compression efficiency when compared with all related works.
Daniel Palomino 0001, Eduardo Cavichioli, Altamiro Amadeu Susin, Luciano Volcan Agostini, Muhammad Shafique 0001, Jörg Henkel
PCS6
2013 Reliable code generation and execution on unreliable hardware under joint functional and timing reliability considerations
abstract
To enable reliable embedded systems, it is imperative to leverage the compiler and system software for joint optimization of functional correctness, i.e., vulnerability indexes, and timing correctness, i.e., the deadline misses. This paper considers the optimization of the Reliability-Timing (RT) penalty, defined as a linear combination of the vulnerability indexes (reliability penalties) and the deadline misses. We propose a multi-layer approach to achieve reliable code generation and execution at compilation and system software layers for embedded systems. This is enabled by the concept of generating multiple versions, for given application functions, with diverse performance and reliability tradeoffs, by exploiting different reliability-guided compilation options. Based on the reliability and execution time profiling of these versions, our reliability-driven system software employs dynamic version selections to dynamically select a suitable version of a function according to the execution behavior of the previous functions. Specifically, our scheme builds a schedule table offline to optimize the RT penalty, and uses this table at run time to select suitable versions for the subsequent functions properly. A complex real-world application of “secure video and audio processing” composed of various functions is evaluated for reliable code generation and execution. The reliability analysis and evaluation is performed on a reliability-aware processor simulator.
Semeen Rehman, Anas Toma, Florian Kriebel, Muhammad Shafique 0001, Jian-Jia Chen, Jörg Henkel
IEEE Real-Time and Embedded Technology and Applications Symposium6
2013 Test Strategies for Reliable Runtime Reconfigurable Architectures
abstract
Field-programmable gate array (FPGA)-based reconfigurable systems allow the online adaptation to dynamically changing runtime requirements. The reliability of FPGAs, being manufactured in latest technologies, is threatened by soft errors, as well as aging effects and latent defects. To ensure reliable reconfiguration, it is mandatory to guarantee the correct operation of the reconfigurable fabric. This can be achieved by periodic or on-demand online testing. This paper presents a reliable system architecture for runtime-reconfigurable systems, which integrates two nonconcurrent online test strategies: preconfiguration online tests (PRET) and postconfiguration online tests (PORT). The PRET checks that the reconfigurable hardware is free of faults by periodic or on-demand tests. The PORT has two objectives: It tests reconfigured hardware units after reconfiguration to check that the configuration process completed correctly and it validates the expected functionality. During operation, PORT is used to periodically check the reconfigured hardware units for malfunctions in the programmable logic. Altogether, this paper presents PRET, PORT, and the system integration of such test schemes into a runtime-reconfigurable system, including the resource management and test scheduling. Experimental results show that the integration of online testing in reconfigurable systems incurs only minimum impact on performance while delivering high fault coverage and low test latency.
Lars Bauer, Claus Braun, Michael E. Imhof, Michael A. Kochte, Eric Schneider, Hongyan Zhang 0004, Jörg Henkel, Hans-Joachim Wunderlich
IEEE Trans. Computers7
2013 Model Predictive Hierarchical Rate Control With Markov Decision Process for Multiview Video Coding
abstract
This paper presents a novel hierarchical rate control (HRC) for the Multiview Video Coding standard targeting improved bandwidth usage and high video quality. The HRC is designed to jointly address the rate control at both frame level and basic unit (BU) level. The proposed scheme is able to exploit the bitrate distribution correlation with neighboring frames to efficiently predict the future bitrate behavior by employing a model predictive control that defines a proper control action through quantization parameter (QP) adaptation. To provide a fine-grained tuning, the QP is further adapted within each frame by a Markov decision process implemented at BU level able to take into consideration a map of the regions of interest. A coupled frame/BU level feedback is performed in order to guarantee the system consistency. Experimental results show the superiority of our HRC compared to state-of-the-art solutions in terms of bitrate allocation accuracy and rate distortion while delivering smooth video quality at frame and BU levels.
Bruno Boessio Vizzotto, Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel
IEEE Trans. Circuits Syst. Video Technol.5
2013 Guest Editorial Special Section on Power-Aware Design for Embedded Systems
abstract
The papers in this special section present state-of-the-art power aware design for embedded systems. The articles examine recent developments used to address such topics as power consumption, power design, power requirements, energy management, and system performance.
Jian-Jia Chen, Jörg Henkel, Xiaobo Sharon Hu
IEEE Trans. Ind. Informatics2
2012 Invasive manycore architectures
abstract
This paper introduces a scalable hardware and software platform applicable for demonstrating the benefits of the invasive computing paradigm. The hardware architecture consists of a heterogeneous, tile-based manycore structure while the software architecture comprises a multi-agent management layer underpinned by distributed runtime and OS services. The necessity for invasive-specific hardware assist functions is analytically shown and their integration into the overall manycore environment is described.
Jörg Henkel, Andreas Herkersdorf, Lars Bauer, Thomas Wild, Michael Hübner 0001, Ravi Kumar Pujari, Artjom Grudnitsky, Jan Heisswolf, Aurang Zaib, Benjamin Vogel, Vahid Lari, Sebastian Kobbe
ASP-DAC1
2012 RAISE: Reliability-Aware Instruction SchEduling for unreliable hardware
abstract
A compile-time Reliability-Aware Instruction SchEduling (RAISE) scheme is presented, which takes into account the spatial and temporal vulnerabilities of different processor resources (pipeline, register file, etc.) used during the execution of different instructions. It reduces the software program's susceptibility towards failures by minimizing the occupancy cycles of critical instructions inside the pipeline stages in addition to reducing the vulnerable periods of their operands. To facilitate RAISE, a novel technique for static reliability estimation during compilation is presented (i.e. before instructions scheduling). Compared to state-of-the-art reliability-aware instruction schedulers, our scheme provides up to 32.7% reduced software program failures over three different fault rates.
Semeen Rehman, Muhammad Shafique 0001, Florian Kriebel, Jörg Henkel
ASP-DAC4
2012 Instruction scheduling for reliability-aware compilation
abstract
An instruction scheduling technique is presented that targets at improving the reliability of a software program given a user-provided tolerable performance overhead. A look-ahead-based heuristic schedules instructions by evaluating the reliability of dependent instructions while reducing the impact of spatial and temporal vulnerabilities of various processor components. Our reliability-driven instruction scheduler (implemented into the GCC compiler) provides on average a 22% reduction of program failures compared to state-of-the-art.
Semeen Rehman, Muhammad Shafique 0001, Jörg Henkel
DAC3
2012 Adaptive power management of on-chip video memory for multiview video coding
abstract
An adaptive power management of on-chip video memory for Multiview Video Coding is presented. It leverages texture, motion and disparity properties of objects and their correlations in the 3D-neighborhood. It groups different Macroblocks of a frame and predicts the highly-probable motion/disparity search direction in order to power-gate idle memory regions. Exploited are the statistical properties of Macroblock groups to predict idle sectors. Our approach achieves on average 32% and 61% energy reduction (averaged over various video sequences) compared to state-of-the-art DSW [7] and Level C [12], respectively. The Motion/Disparity Estimation architecture with video memory and power management scheme is implemented using an ASIC flow (IBM-65nm Low-Power technology) and it processes 4-view [email protected]
Muhammad Shafique 0001, Bruno Zatt, Fabio Leandro Walter, Sergio Bampi, Jörg Henkel
DAC5
2012 Partial online-synthesis for mixed-grained reconfigurable architectures
abstract
Processor architectures with Fine-Grained Reconfigurable Accelerators (FGRAs) allow for a high degree of adaptivity to address varying application requirements. When processing computation intensive kernels, multiple FGRAs may be used to execute a complex function. In order to exploit the adaptivity of a fine-grained reconfigurable fabric, a runtime system should decide when and which FGRAs to reconfigure with respect to application requirements. To enable this adaptivity, a flexible infrastructure is required that allows combining FGRAs to execute complex functions. We propose a mixed-grained reconfigurable architecture composed from a Coarse-Grained Reconfigurable Infrastructure (CGRI) that connects the FGRAs. At runtime we synthesize CGRI configurations that depend on decisions of the runtime system, e.g. which FGRAs shall be reconfigured. Synthesis and place & route of the FGRAs are done at compile time for performance reasons. Combined, this results in a partial online synthesis for mixed grained reconfigurable architectures, which allows maintaining a low runtime overhead while exploiting the inherent adaptivity of the reconfigurable fabric. In this work we focus on the crucial parts of synthesizing the configurations for the CGRI at runtime, propose algorithms, and compare their performance/overhead trade-offs for different application scenarios. We are the first to exploit the increased adaptivity of FGRAs that are connected by a CGRI, by using our partial online synthesis. In comparison to a state-of-the-art reconfigurable architecture that synthesizes the configurations for the CGRI at compile time we obtain an average speedup of 1.79x.
Artjom Grudnitsky, Lars Bauer, Jörg Henkel
DATE3
2012 Dynamic cache management in multi-core architectures through run-time adaptation
abstract
Non-Uniform Cache Access (NUCA) architectures provide a potential solution to reduce the average latency for the last-level-cache (LLC), where the cache is organized into per-core local and remote partitions. Recent research has demonstrated the benefits of cooperative cache sharing among local and remote partitions. However, ignoring cache access patterns of concurrently executing applications sharing the local and remote partitions can cause inter-partition contention that reduces the overall instruction throughput. We propose a dynamic cache management scheme for LLC in NUCA-based architectures, which reduces inter-partition contention. Our proposed scheme provides efficient cache sharing by adapting migration, insertion, and promotion policies in response to the dynamic requirements of the individual applications with different cache access behaviors. Our adaptive cache management scheme allows individual cores to steal cache capacity from remote partitions to achieve better resource utilization. On average, our proposed scheme increases the performance (instructions per cycle) by 28% (minimum 8.4%, maximum 75%) compared to a private LLC organization.
Fazal Hameed, Lars Bauer, Jörg Henkel
DATE3
2012 Power-efficient error-resiliency for H.264/AVC Context-Adaptive Variable Length Coding
abstract
Technology scaling has led to unreliable computing hardware due to high susceptibility against soft errors. In this paper, we propose an error-resilient architecture for Context-Adaptive Variable Length Coding (CAVLC) in H.264/AVC. Due to its context-adaptive nature and intricate control flow CAVLC is very sensitive to soft errors. An error during the CAVLC process (especially during the context adaptation or in VLC tables) may result in severe mismatch between encoder and decoder. The primary goal in our error-resilient CAVLC architecture is to protect codeword/codelength tables and context adaptation in a reliable yet power efficient manner. For reducing the power over-head, the tables are partitioned in various sub-tables each protected with variable-sized parity. Moreover, for further power reduction, our approach incorporates state-retentive power-gating of different sub-tables at run time depending upon the statistical distribution of syntax elements. Compared to the unprotected case, our scheme provides a video quality improvement of 18dB (averaged over various fault injection cases and video sequences) at the cost of a 35% area overhead and 45% performance overhead due to the error-detection logic. However, partitioned sub-tables increase the potential for power-gating, thus bring a leakage energy saving of 58%. Compared to state-of-the-art table protection, our scheme provides 2x reduced area and performance overhead. For function-al verification and area comparison, the architecture is prototyped on a Xilinx Virtex-5 FPGA, though not limited to it. For the soft errors experiments, evaluation of error-resiliency and power efficiency, we have developed a fault injection and simulation setup.
Muhammad Shafique 0001, Bruno Zatt, Semeen Rehman, Florian Kriebel, Jörg Henkel
DATE5
2012 Accurate source-level simulation of embedded software with respect to compiler optimizations
abstract
Source code instrumentation is a widely used method to generate fast software simulation models by annotating timing information into application source code. Source-level simulation models can be easily integrated into SystemC based simulation environment for fast simulation of complex multiprocessor systems. The accurate back-annotation of the timing information relies on the mapping between source code and binary code. The compiler optimizations might make it hard to get accurate mapping information. This paper addresses the mapping problems caused by complex compiler optimizations, which are the main source of simulation errors. To obtain accurate mapping information, we propose a method called fine-grained flow mapping that establishes a mapping between sequences of control flow of source code and binary code. In case that the code structure of a program is heavily altered by compiler optimizations, we propose to replace the altered part of the source code with functionally-equivalent IR-level code which has an optimized structure, leading to Partly Optimized Source Code (POSC). Then the flow mapping can be established between the POSC and the binary code and the timing information is back-annotated to the POSC. Our experiments demonstrate the accuracy and speed of simulation models generated by our approach.
Zhonglei Wang, Jörg Henkel
DATE2
2012 Dependable embedded systems: The German research foundation DFG priority program SPP 1500
abstract
When migrating to future technology nodes, dependability becomes a major design problem as variability, aging and susceptibility to soft errors increase. The purpose of this program is to research cross-layer solutions that address the physical problems at system-level i.e. at hardware-level, operating system level, application level etc. The goals and an overview of the DFG SPP 1500 research program are presented.
Jörg Henkel, Oliver Bringmann 0001, Andreas Herkersdorf, Wolfgang Rosenstiel, Norbert Wehn
ETS1
2012 PATS: A Performance Aware Task Scheduler for Runtime Reconfigurable Processors
abstract
Multi-tasking is one of the main requirements for complex embedded systems to fulfill user expectations (e.g. flexibility of the system), increase the resource utilization, and thus increase the system efficiency. In general, the flexibility and efficiency can be increased by incorporating a fine-grained reconfigurable fabric (e.g. an embedded FPGA) that is coupled with a general-purpose processor and accelerates the computationally intensive kernels. This work focuses on reconfigurable processors that use a reconfigurable fabric to implement Special Instructions (SIs) that are invoked by the processor and process data-dominant parts. For each SI the decision whether it is executed in hardware or emulated in software can be changed dynamically at runtime. In this paper, we present our novel Performance Aware Task Scheduler (PATS) that decides the task schedule at runtime while considering the specific system state of the reconfigurable processor. For instance, if a task t has to emulate several SI executions in software because reconfiguring the corresponding hardware implementations is not completed yet, then it might be more efficient to schedule other tasks first, depending on the soft-deadlines of the tasks, until the reconfigurations of that task t are completed. In comparison to other task schedulers (earliest deadline first, rate monotonic scheduling, and round robin), PATS achieves on average a 1.45x better system tardiness (i.e., the sum of cycles by which tasks miss their deadlines). Additionally, PATS reduces the make span (i.e. the time when all tasks have completed all of their jobs) on average by 1.17x (up to 1.58x). Especially in challenging multi-tasking scenarios with tight deadlines or a small reconfigurable fabric PATS performs significantly better than other task schedulers do.
Lars Bauer, Artjom Grudnitsky, Muhammad Shafique 0001, Jörg Henkel
FCCM4
2012 A Model Predictive Controller for Frame-Level Rate Control in Multiview Video Coding
abstract
In this work, we present a novel frame-level Rate Control algorithm for Multiview Video Coding encoder that adopts the Model Predictive Control technique in order to provide low bitrate fluctuation and high video quality. Our Model Predictive Rate Control (MPRC) predicts the bitrate for a frame by employing (i) inter-view inter-GOP (Group of Pictures) phase-based bitrate prediction, and (ii) temporal (intra-GOP) target bitrate linear weighting. Moreover, the MPRC also defines an optimal control action through frame-level QP value selection. Experimental results demonstrate that our MPRC bitrate prediction incurs a Mean Bit Estimation Error (MBEE) of 1.13% compared to 2.46% provided by single view-based Rate Control and 1.61% provided by the state-of-the-art MVC Rate Control. Our solution also provides on average 0.876dB BD-PSNR increase and 28.92% BD-Bitrate reduction while providing smoother quality and bitrate variations when compared to state-of-the-art.
Bruno Boessio Vizzotto, Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel
ICME5
2012 A Reconfigurable Hardware Accelerated Platform for Clustered Wireless Sensor Networks
abstract
With the advent of the FPGA technology, a reconfigurable platform becomes possible to enhance the capability and adaptability of wireless sensor nodes while reducing the energy consumption. In this paper, we investigate the application of reconfigurable platform in wireless sensor networks. We focus on using hardware accelerated lossless compression to reduce the energy consumption of the data aggregation in clustered networks. Our study is based on real-world measurement on our integrated reconfigurable sensor node and the measured data is then used in the network simulation with a proper channel model. Our experiments show that the use of hardware acceleration on our platform can reduce the energy consumption by up to 35%. In addition, with proper parameters of the network, the run-time reconfiguration becomes beneficial and this opens a lot of opportunities for the optimizations according to the field situations.
Chih-Ming Hsieh, Zhonglei Wang, Jörg Henkel
ICPADS3
2012 Transparent structural online test for reconfigurable systems
abstract
FPGA-based reconfigurable systems allow the online adaptation to dynamically changing runtime requirements. However, the reliability of modern FPGAs is threatened by latent defects and aging effects. Hence, it is mandatory to ensure the reliable operation of the FPGA's reconfigurable fabric. This can be achieved by periodic or on-demand online testing. In this paper, a system-integrated, transparent structural online test method for runtime reconfigurable systems is proposed. The required tests are scheduled like functional workloads, and thorough optimizations of the test overhead reduce the performance impact. The proposed scheme has been implemented on a reconfigurable system. The results demonstrate that thorough testing of the reconfigurable fabric can be achieved at negligible performance impact on the application.
Mohamed Abdelfattah, Lars Bauer, Claus Braun, Michael E. Imhof, Michael A. Kochte, Hongyan Zhang 0004, Jörg Henkel, Hans-Joachim Wunderlich
IOLTS7
2012 An adaptive data gathering strategy for target tracking in cluster-based wireless sensor networks
Zhonglei Wang, Jörg Henkel
ISCC3
2012 An adaptive data gathering strategy for target tracking in cluster-based wireless sensor networks
abstract
In a typical cluster-based sensor network, the data is usually gathered and fused on cluster heads. In target tracking applications, a target is often detected by sensor nodes in multiple clusters, leading to redundant data transmissions through multiple paths from the cluster heads to the data sink. To reduce such redundant data transmissions and thus to save energy, this paper proposes an adaptive data gathering strategy, called ADGS. Our novel idea is to adaptively select one node with the most residual energy and the least communication cost from the active nodes around the target. This node is responsible for gathering and aggregating the data from the other active nodes and is therefore called Aggregation Node (AN). The aggregated data is then transmitted only from the AN to the sink. Our experiments demonstrate that the proposed approach achieves a significant reduction in power consumption for data transmission and prolongs the network lifetime by 857.6% and 85.8% compared to two state-of-the-art data gathering approaches.
Zhonglei Wang, Jörg Henkel
ISCC3
2012 A complexity reduction scheme with adaptive search direction and mode elimination for multiview video coding
abstract
A novel complexity reduction scheme for Multiview Video Coding (MVC) is presented that adaptively eliminates the less-probable Motion or Disparity Estimation (ME, DE) search directions and less-probable block coding modes. Based on their texture difference w.r.t. the current Macroblock, matching neighbors in the 3D-neighborhood (spatial, temporal, view domains) are identified. Our scheme employs a multi-level decision process to predict the more-probable ME/DE search direction based on the texture and (motion/disparity) activity classification of the matching neighbors. For a predicted search direction, more-probable block coding modes are predicted depending upon the texture classification and RD-Cost of the current Macroblock. Quantization Parameter based thresholds are formulated using an offline statistical analysis of texture, motion/disparity, and RD-Cost properties. Our scheme achieves a complexity reduction of up to 81% and 40% compared to the exhaustive Rate-Distortion-Optimized Mode Decision and state-of-the-art, respectively, at the cost of an average BD-PSNR loss of 0.03 dB.
Muhammad Shafique 0001, Bruno Zatt, Jörg Henkel
PCS3
2012 ECO/ee: Energy-aware Collaborative Organic execution environment for wireless sensor networks
abstract
This paper presents an energy-aware execution environment, called Energy-aware Collaborative Organic Execution Environment (ECO/ee), for Wireless Sensor Networks (WSN). ECO/ee provides an energy-aware and autonomous data routing scheme. It can dynamically adapt to the user's queries and efficiently determine the energy-abundant delivery paths. The fundamental concepts are inspired by the forage and task allocation mechanisms of honey bees as well as the interaction between the honey bee colony and the bee keeper. The experiments show that our approach outperforms the state-of-the-art in terms of energy efficiency and adaptability to the user requirements.
Chih-Ming Hsieh, Zhonglei Wang, Jörg Henkel
WCNC3
2012 AdNoC: Runtime Adaptive Network-on-Chip Architecture
abstract
Networsk-on-chip (NoCs) have emerged as a promising on-chip interconnect for future multi/many-core architectures as NoCs are able to scale communication links with the growing number of cores. State-of-the-art NoC designs rely mainly on a static network configuration using fixed routing algorithms and buffer placements. These approaches are not effective in dealing with hard-to-predict system behavior, for instance due to user behavior or varying workloads, since in order for static NoCs to cover these scenarios, they would have to be designed for worst case scenarios. In this paper, we address these problems with a runtime adaptive network-on-chip (AdNoC). Focusing on the architecture-level adaptation, we present an adaptive route allocation algorithm which provides a required level of QoS (guaranteed bandwidth) coupled with an adaptive buffer assignment scheme which reassigns buffer blocks on-demand. Furthermore, the adaptivity requires a comprehensive, hardly intrusive, runtime observability infrastructure, i.e., using monitoring components, in order to gather data on the system state. The area overhead introduced by the adaptive scheme can be traded off against the flexibility gained. Moreover, the area overhead is also reduced by resource multiplexing due to the on-demand buffer assignment at each output port (we achieved on an average 42% buffer saving in our experiments). We demonstrate the advantage by using various digital media applications and compare our approach to the state-of-the-art static NoC architectures e.g., Xpipe, QNoC, and Æthereal.
Mohammad Abdullah Al Faruque, Thomas Ebi, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.3
2011 SEAL: soft error aware low power scheduling by Monte Carlo state space under the influence of stochastic spatial and temporal dependencies
abstract
A processor's performance and power consumption are tied; an increased performance demands more power, and vice versa. An optimal tradeoff can only be achieved by an improved prediction of the task execution times, prior to an efficient scheduling. Moreover, since the processor's soft error rate is a function of its operating voltage, it is also linked to the performance-power trade-off. The situation is further complicated for the case of multicore architectures where the tasks are to be mapped on separate cores (processing elements). This paper proposes a joint State-Space model to achieve improved task execution time estimation, leading to better scheduling for optimizing the trade-off, particularly in the context of multicore soft real-time systems. It does not assume any `a priori' knowledge about the task graph or its properties, and is independent of the underlying architecture. It learns the system dynamics over time. The state-space solution is formulated using a recursive implementation of the online Monte Carlo Method. Having obtained the estimates of the execution times, they are compensated for the soft error according to a given soft error rate. At the beginning of each scheduling interval, the low power EDF scheduling decision is carried out to execute the tasks. The proposed method (SEAL) achieves 29% better energy savings compared to state-of-the-art, while the deadline misses are under 7% without the loss of system failure probability. The results obtained clearly show the advantage in terms of energy savings.
Nabeel Iqbal, Muhammad Adnan Siddique, Jörg Henkel
DAC3
2011 Low-power adaptive pipelined MPSoCs for multimedia: an H.264 video encoder case study
abstract
Pipelined MPSoCs provide a high throughput implementation platform for multimedia applications, with reduced design time and improved flexibility. Typically a pipelined MPSoC is balanced at design-time using worst-case parameters. Where there is a widely varying workload, such designs consume exorbitant amount of power. In this paper, we propose a novel adaptive pipelined MPSoC architecture that adapts itself to varying workloads. Our architecture consists of Main Processors and Auxiliary Processors with a distributed run-time balancing approach, where each Main Processor, independent of other Main Processors, decides for itself the number of required Auxiliary Processors at run-time depending on its varying workload. The proposed run-time balancing approach is based on off-line statistical information along with workload prediction and run-time monitoring of current and previous workloads' execution times. We exploited the adaptability of our architecture through a case study on an H.264 video encoder supporting HD720p at 30 fps, where clock- and power-gating were used to deactivate idle Auxiliary Processors during low workload periods. The results show that an adaptive pipelined MPSoC provides energy savings of up to 34% and 40% for clock- and power-gating based deactivation of Auxiliary Processors respectively with a minimum throughput of 29 fps when compared to a design-time balanced pipelined MPSoC.
Haris Javaid, Muhammad Shafique 0001, Sri Parameswaran, Jörg Henkel
DAC4
2011 Run-time adaptive energy-aware motion and disparity estimation in multiview video coding
abstract
This paper presents a novel run-time adaptive energy-aware Motion and Disparity Estimation (ME, DE) architecture for Multiview Video Coding (MVC). It incorporates efficient memory access and data prefetching techniques for jointly reducing the on/off-chip memory energy consumption. A dynamically expanding search window is constructed at run time to reduce the off-chip memory accesses. Considering the multi-stage processing nature of advanced fast ME/DE schemes, a reduced-sized multi-bank on-chip memory is employed which can be power-gated depending upon the video properties. As a result, when tested for various video sequence, our approach provides a dynamic energy reduction of 82--96% for the off-chip memory and a leakage energy reduction of 57--75% for the on-chip memory compared to the Level-C and Level-C+ [7] prefetching techniques (which are the prominent data reuse and prefetching techniques in ME for video coding). The proposed ME/DE architecture is synthesized using a 65nm IBM low power technology. Compared to state-of-the-art MVC ME/DE hardware [14], our architecture provides 66% and 72% reduction in the area and power consumption, respectively. Moreover, our scheme achieves 30fps ME/DE 4-view HD1080p encoding with a power consumption of 74mW.
Bruno Zatt, Muhammad Shafique 0001, Felipe Sampaio, Luciano Volcan Agostini, Sergio Bampi, Jörg Henkel
DAC6
2011 mRTS: Run-time system for reconfigurable processors with multi-grained instruction-set extensions
abstract
We present a run-time system for a multi-grained reconfigurable processor in order to provide a dynamic trade-off between performance and available area budgets for both fine- as well as coarse-grained reconfigurable fabrics as part of one reconfigurable processor. Our run-time system is the first implementation of its kind that dynamically selects and steers a performance-maximizing multi-grained instruction set under run-time varying constraints. It achieves a performance improvement of more than 2× compared to state-of-the-art run-time systems for multi-grained architectures. To elaborate the benefits of our approach further, we also compare it with offline- and online-optimal instruction-set selection schemes.
Waheed Ahmed, Muhammad Shafique 0001, Lars Bauer, Jörg Henkel
DATE4
2011 Dynamic thermal management in 3D multi-core architecture through run-time adaptation
abstract
3D multi-core architectures are seen to provide increased transistor density, reduced power consumption, and improved performance through wire length reduction. However, 3D suffers from increased power density, which exacerbates thermal hotspots. In this paper, we present a novel 3D multi-core architecture that reduces processor activity on the die distant to the heat sink and a core-level dynamic thermal management technique based on the architectural adaptation, e.g. dynamically adapting core-resources depending on diverse application requirements and thermal behavior. The proposed thermal management technique synergistically combines the benefits of the architectural adaptation supported by our 3D multi-core architecture with dynamic voltage and frequency scaling. Our proposed technique provides 19.4% (maximum 24.4%, minimum 15.5%) improvement in the instruction throughput compared to the state-of-the-art thermal management techniques [4, 5] applied to the thermal-aware 3D processor architecture without considering run-time adaptation [10].
Fazal Hameed, Mohammad Abdullah Al Faruque, Jörg Henkel
DATE3
2011 CARAT: Context-aware runtime adaptive task migration for multi core architectures
abstract
Multi core architectures that are built to reap performance and energy efficiency benefits from the parallel execution of applications often employ runtime adaptive techniques in order to achieve, among others, load balancing, dynamic thermal management, and to enhance the reliability of a system. Typically, such runtime adaptation in the system level requires the ability to quickly and consistently migrate a task from one core to another. For distributed memory architectures, the policy for transferring the task context between source and destination cores is of vital importance to the performance and to the successful operation of the system. As its performance is negatively correlated with the communication overhead, energy consumption and the dissipated heat, task migration needs to be runtime adaptive to account for the system load, chip temperature, or battery capacity. This work presents a novel context-aware runtime adaptive task migration mechanism (CARAT) that reduces the task migration latency by 93.12%, 97.03% and 100% compared to three state-of-the-art mechanisms and allows to control the maximum migration delay and the performance overhead tradeoff at runtime. This novel mechanism is built on an in-depth analysis of the memory access behavior of several multi-media and robotic embedded-systems applications.
Janmartin Jahn, Mohammad Abdullah Al Faruque, Jörg Henkel
DATE3
2011 Minority-Game-based resource allocation for run-time reconfigurable multi-core processors
abstract
A novel policy for allocating reconfigurable fabric resources in multi-core processors is presented. We deploy a Minority-Game to maximize the efficient use of the reconfigurable fabric while meeting performance constraints of individual tasks running on the cores. As we will show, the Minority Game ensures a fair allocation of resources, e.g., no single core will monopolize the reconfigurable fabric. Rather, all cores receive a “fair” share of the fabric, i.e., their tasks would miss their performance constraints by approximately the same margin, thus ensuring an overall graceful degradation. The policy is implemented on a Virtex-4 FPGA and evaluated for diverse applications ranging from security to multimedia domains. Our results show that the Minority-Game policy achieves on average 2× higher application performance and a 5× improved efficiency of resource utilization compared to state-of-the-art.
Muhammad Shafique 0001, Lars Bauer, Waheed Ahmed, Jörg Henkel
DATE4
2011 Multi-level pipelined parallel hardware architecture for high throughput motion and disparity estimation in Multiview Video Coding
abstract
This paper presents a novel motion and disparity estimation (ME, DE) scheme in Multiview Video Coding (MVC) that addresses the high throughput challenge jointly at the algorithm and hardware levels. Our scheme is composed of a fast ME/DE algorithm and a multi-level pipelined parallel hardware architecture. The proposed fast ME/DE algorithm exploits the correlation available in the 3D-neighborhood (spatial, temporal, and view). It eliminates the search step for different frames by prioritizing and evaluating the neighborhood predictors. It thereby reduces the coding computations by up to 83% with 0.1 dB quality loss. The proposed hardware architecture further improves the throughput by using parallel ME/DE modules with a shared array of SAD (Sum of Absolute Differences) accelerators and by exploiting the four levels of parallelism inherent to the MVC prediction structure (view, frame, reference frame, and macroblock levels). A multi-level pipeline schedule is introduced to reduce the pipeline stalls. The proposed architecture is implemented for a Xilinx Virtex-6 FPGA and as an ASIC with an IBM 65nm low power technology. It is compared to state-of-the-art at both algorithm and hardware levels. Our scheme achieves a real-time (30fps) ME/DE in 4-view High Definition (HD1080p) encoding with a low power consumption of 81 mW.
Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel
DATE4
2011 Run-Time Resource Allocation for Simultaneous Multi-tasking in Multi-core Reconfigurable Processors
abstract
State-of-the-art multi-core reconfigurable processors do not exploit the full potential of simultaneous multi-tasking with run-time adaptive reconfigurable fabric allocation. We propose a novel run-time system for simultaneous multi-tasking in a multi-core reconfigurable processor that adaptively allocates the mixed-grained reconfigurable fabric resource at run time among different tasks considering their performance constraints. Our scheme employs the novel concept of refined task-criticality (based on the functional-block-level performance constraints) considering the computational properties of dependent tasks and their inherent potential for acceleration. Our scheme dynamically compensates the deadline misses at the functional block level. It thereby reduces the potential task-level deadline misses under competing scenarios. With the help of a secure video conferencing application (with 4 dependent tasks of diverse computational properties), we demonstrate that our scheme reduces the deadline misses by (on average) 6× under given performance constraints, when compared to state-of-the-art reconfigurable processors.
Waheed Ahmed, Muhammad Shafique 0001, Lars Bauer, Manuel Hammerich, Jörg Henkel, Jürgen Becker 0001
FCCM5
2011 System-level application-aware dynamic power management in adaptive pipelined MPSoCs for multimedia
abstract
System-level dynamic power management (DPM) schemes in Multiprocessor System on Chips (MPSoCs) exploit the idleness of processors to reduce the energy consumption by putting idle processors to low-power states. In the presence of multiple low-power states, the challenge is to predict the duration of the idle period with high accuracy so that the most beneficial power state can be selected for the idle processor. In this work, we propose a novel dynamic power management scheme for adaptive pipelined MPSoCs, suitable for multimedia applications. We leverage application knowledge in the form of future workload prediction to forecast the duration of idle periods. The predicted duration is then used to select an appropriate power state for the idle processor. We proposed five heuristics as part of the DPM and compared their effectiveness using an MPSoC implementation of the H.264 video encoder supporting HD720p at 30 fps. The results show that one of the application prediction based heuristic (MAMAPBH) predicted the most beneficial power states for idle processors with less than 3% error when compared to an optimal solution. In terms of energy savings, MAMAPBH was always within 1% of the energy savings of the optimal solution. When compared with a naive approach (where only one of the possible power states is used for all the idle processors), MAMAPBH achieved up to 40% more energy savings with only 0.5% degradation in throughput. These results signify the importance of leveraging application knowledge at system-level for dynamic power management schemes.
Haris Javaid, Muhammad Shafique 0001, Jörg Henkel, Sri Parameswaran
ICCAD3
2011 A low-power memory architecture with application-aware power management for motion & disparity estimation in Multiview Video Coding
abstract
A low-power architecture for an on-chip multi-banked video memory for motion and disparity estimation in Multiview Video Coding is proposed. The memory organization (size, banks, sectors, etc.) is driven by an extensive analysis of memory-usage behavior for various 3D-video sequences. Considering a multiple-sleep state model, an application-aware power management scheme is employed to reduce the leakage energy of the on-chip memory. The knowledge of motion and disparity estimation algorithm in conjunction with video properties are considered to predict the memory requirements of each Macroblock. A cost function is evaluated to determine an appropriate sleep mode for the idle memory sectors, while considering the wakeup overhead (latency and energy). The complete motion and disparity estimation architecture is implemented in a 65nm low power IBM technology. The experiments (for various test video sequences) demonstrate that our architecture provides up to 80% leakage energy reduction compared to state-of-the-art. Our scheme processes motion and disparity estimation of four HD1080p views encoding at 30fps with a power consumption of 57mW.
Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel
ICCAD4
2011 Revc: Computationally Reliable Video Coding on unreliable hardware platforms: A case study on error-tolerant H.264/AVC CAVLC entropy coding
abstract
Advancement in technology has complemented the increased complexity of state-of-the-art video coding standards by providing high-performance multimedia platforms. This has resulted in nonnegligible computational reliability issues, especially with respect to soft errors. This paper motivates the dire need towards computationally Reliable Video Coding (ReVC) on modern hardware platforms. We highlight the issues with the help of a fault injection analysis during the CAVLC processing. We present a case study on error-tolerant CAVLC, where error tolerance is achieved at various levels of abstraction, ranging from algorithm to hardware. Application-specific knowledge of CAVLC is used to detect potential errors at the algorithm level, while hardware means are provided to protect the tables for CAVLC. Experiments demonstrate that our approach significantly improves the video quality under high fault rates. Furthermore, we demonstrate that when considering the application-specific knowledge, our proposed solution incurs minimal area and performance overhead compared to redundancy-based reliability methods.
Semeen Rehman, Muhammad Shafique 0001, Florian Kriebel, Jörg Henkel
ICIP4
2011 A high-throughput parallel hardware architecture for H.264/AVC CAVLC encoding
abstract
This paper presents a high-throughput hardware architecture for H.264/AVC CAVLC encoding. Our scheme eliminates the pipeline stage of computing the coefficient statistics (as adopted by state-of- the-art hardware architectures) with a pre-processing stage during the quantization in order to avoid the extra looping logic in CAVLC. This provides significant performance improvement compared to state-of-the-art (saving of 16 cycles per 4×4 sub-block compared to [2]). Furthermore, our hardware architecture employs parallel processing of Trailing Ones (which is one of the inherently sequential steps in CAVLC) and encodes levels and runs in parallel in the same pipeline stage. An intelligent bitstream writing logic generates the compliant bitstream. Compared to state-of-the-art, our proposed hardware architecture requires 72% reduced area and achieves 2× higher throughput, while processing HD1080p@30fps.
Muhammad Shafique 0001, Adnan Orcun Tüfek, Jörg Henkel
ICIP3
2011 A multi-level dynamic complexity reduction scheme for multiview video coding
abstract
In this paper, we propose a novel scheme for dynamically reducing the computational complexity of MVC. Our scheme exploits the coding mode correlation available in the 3D-neighborhood (i.e., spatial, temporal, and view) along with the rate-distortion proper- ties of the neighboring Macroblocks. Our scheme incorporates a multi-level mode decision process based on a mode-ranking mechanism that categorizes more-probable and less-probable coding modes. In order to react to the changing bitrates, our scheme deploys Quantization Parameter based threshold equations which are formulated using an offline statistical analysis. Compared to the exhaustive Rate-Distortion-Optimized Mode Decision (RDO-MD), our scheme achieves a complexity reduction of up to 80% (68% on average) with an average PSNR loss of 0.075 dB. Compared to state-of-the-art fast RDO-MD, our scheme achieves a complexity reduction of up to 34% with an average PSNR gain of 0.007 dB.
Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel
ICIP4
2011 RDTS: A Reliable Erasure-Coding Based Data Transfer Scheme for Wireless Sensor Networks
abstract
Information redundancy using erasure coding is an efficient way to increase the reliability of data transmission in communication systems. In Wireless Sensor Networks (WSNs), erasure encoding and decoding are performed on the source node and sink node, respectively, and a large amount of redundant data is generated according to the quality of the whole path and transmitted through multiple hops. In this paper, we propose a reliable data transfer scheme, RDTS, where erasure coding is performed in a hop-by-hop manner, which means that each intermediate node is able to perform erasure coding and adaptively calculates the number of redundant packets for the next hop. Usually, only a small amount of redundant data is needed for reliable transmission over a single hop. Therefore, using RDTS, the network load caused by redundant data is significantly reduced and also well balanced, leading to a longer network lifetime. In addition, hop-by-hop coding has also the advantage of low coding overhead. We further reduce the coding time by proposing a partial coding scheme. Our experimental results show that RDTS achieves up to 69.7% less network load and 153.8% longer lifetime, and meanwhile, the coding overhead is reduced by up to 78.1%, compared with a state-of-the-art erasure-coding based approach.
M. Sammer Srouji, Zhonglei Wang, Jörg Henkel
ICPADS3
2010 RMOT: Recursion in model order for task execution time estimation in a software pipeline
abstract
This paper addresses the problem of execution time estimation for tasks in a software pipeline independent of the application structure or the underlying architecture. A regression model is developed to obtain the estimates from previously observed data. To improve the quality of the estimates execution times of predecessor task in a software pipeline is exploited. Since the Model order (number of past observations required to obtain optimal estimate) cannot be determined at design time and to circumvent this, we propose means to dynamically update the order and hence obtain a critical-fit model without resorting to analytical benchmarking or calibration runs. The estimation scheme comprises of two estimation methods, namely `Wiener-Hopf' and Order-recursive estimation. The selection of the estimation method is automatic and depends on the required quality of the estimate against a user selectable threshold. In order recursion, new model order is obtained in conjunction to estimates, so order recursion solve the system both for order and estimate simultaneously. We experimented on two multicore platforms using H.264 decoder, a control dominant, computationally demanding application. Results show that estimates obtained by our method are up to 39% better in case of the first task in the software pipeline. The estimate quality improves significantly for the task with predecessor(s) in pipeline and comparison shows up to 54% improvement in estimation results.
Nabeel Iqbal, Muhammad Adnan Siddique, Jörg Henkel
DATE3
2010 DAGS: Distribution agnostic sequential Monte Carlo scheme for task execution time estimation
abstract
This paper addresses the problem of stochastic task execution time estimation agnostic to the process distributions. The proposed method is orthogonal to the application structure and underlying architecture. We build the time varying state space model of the task execution time. In the case of software pipelined tasks, to refine the estimate quality, the state-space is modeled as Multiple Input Single Output (MISO) system by taking into account the current execution time of the predecessor task. To obtain nearly Bayesian estimates, irrespective of the process distribution, the sequential Monte Carlo method is applied which form the recursive solution to reduce the overheads and comprises of time update and correction steps. We experimented on three different platforms, including multicore, using the time parallelized H.264 decoder: a control dominant computationally demanding application and AES encoder: a pure data flow application. Results show that estimates obtained by our method are superior in quality and are up to 68% better in comparison to others.
Nabeel Iqbal, Muhammad Adnan Siddique, Jörg Henkel
DATE3
2010 KAHRISMA: A novel Hypermorphic Reconfigurable-Instruction-Set Multi-grained-Array architecture
abstract
Facing the requirements of next generation applications, current approaches of embedded systems design will soon hit the limit where they may no longer perform efficiently. The unpredictable nature and diverse processing behavior of future applications requires to transgress the barrier of tailor-made, application-/domain-specific embedded system designs. As a consequence, next generation architectures for embedded systems have to react much more flexible to unforeseeable run-time scenarios. In this paper we present our innovative processor architecture concept KAHRISMA (KArlsruhe's Hypermorphic Reconfigurable-Instruction-Set Multi-grained-Array). It tightly integrates coarse- and fine-grained run-time reconfigurable fabrics that can incorporate to realize hardware acceleration for computationally complex algorithms. Furthermore, the fabrics can be combined to realize different Instruction Set Architectures that may execute in parallel. With the help of an encrypted H.264 en-/decoding case study we demonstrate that our novel KAHRISMA architecture will deliver the required flexibility to design future-proof embedded systems that are not limited to a certain computational domain.
Ralf König 0001, Lars Bauer, Timo Stripf, Muhammad Shafique 0001, Waheed Ahmed, Jürgen Becker 0001, Jörg Henkel
DATE7
2010 enBudget: A Run-Time Adaptive Predictive Energy-Budgeting scheme for energy-aware Motion Estimation in H.264/MPEG-4 AVC video encoder
abstract
The limited energy resources in portable multimedia devices require the reduction of encoding complexity. The complex Motion Estimation (ME) scheme of H.264/MPEG-4 AVC accounts for a major part of the encoder energy. In this paper we present a Run-Time Adaptive Predictive Energy Budgeting (enBudget) scheme for energy-aware ME that predicts the energy budget for different video frames and different Macroblocks (MBs) in an adaptive manner considering the run-time changing scenarios of available energy, video frame characteristics, and user-defined coding constraints while keeping a good video quality. It assigns different Energy-Quality Classes to different video frames and fine-tunes at MB level depending upon the predictive energy quota in order to cope with above-mentioned run-time unpredictable scenarios. Compared to UMHexagonS, EPZS, and FastME, our enBudget scheme for energy-aware ME achieves an energy saving of up to 93%, 90%, 88% (average 88%, 77%, 66%), respectively. It suffers from an average Peak Signal to Noise Ratio (PSNR) loss of 0.29 dB compared to Full Search. We also demonstrate that enBudget is equally beneficial to various other state-of-the-art fast adaptive MEs (e.g.). We have evaluated our scheme for ASIC and various FPGAs.
Muhammad Shafique 0001, Lars Bauer, Jörg Henkel
DATE3
2010 An HVS-based Adaptive Computational Complexity Reduction Scheme for H.264/AVC video encoder using Prognostic Early Mode Exclusion
abstract
The H.264/AVC video encoder standard significantly improves the compression efficiency by using variable block-sized Inter (P) and Intra (I) Macroblock (MB) coding modes. In this paper, we propose a novel Human Visual System based Adaptive Computational Complexity Reduction Scheme (ACCoReS). It performs Prognostic Early Mode Exclusion and a Hierarchical Fast Mode Prediction to exclude as many I-MB and P-MB coding modes as possible (up to 73%) even before the actual Rate Distortion Optimized Mode Decision (RDO-MD) and Motion Estimation while keeping a good quality. In the best case, ACCoReS processes exactly one MB Type and one corresponding near-optimal coding mode, such that the complete RDO-MD process is skipped. Experimental results show that compared to state-of-the-art approaches, ACCoReS achieves a speedup of up to 9.14× (average 3×) with an average PSNR loss of 0.66 dB. Compared to exhaustive RDO-MD, our ACCoReS provides a performance improvement of up to 19× (average 10×) for an average 3% PSNR loss.
Muhammad Shafique 0001, Bastian Molkenthin, Jörg Henkel
DATE3
2010 SETS: Stochastic execution time scheduling for multicore systems by joint state space and Monte Carlo
abstract
The advent of multicore platforms has renewed the interest in scheduling techniques for real-time systems. Historically, `scheduling decisions' are implemented considering fixed task execution times, as for the case of Worst Case Execution Time (WCET). The limitations of scheduling considering WCET manifest in terms of under-utilization of resources for large application classes. In the realm of multicore systems, the notion of WCET is hardly meaningful due to the large set of factors influencing it. Within soft real-time systems, a more realistic modeling approach would be to consider tasks featuring varying execution times (i.e. stochastic). This paper addresses the problem of stochastic task execution time scheduling that is agnostic to statistical properties of the execution time. Our proposed method is orthogonal to any number of linear acyclic task graphs and their underlying architecture. The joint estimation of execution time and the associated parameters, relying on the interdependence of parallel tasks, help build a `nonlinear Non-Gaussian state space' model. To obtain nearly Bayesian estimates, irrespective of the execution time characteristics, a recursive solution of the state space model is found by means of the Monte Carlo method. The recursive solution reduces the computational and memory overhead and adapts statistical properties of execution times at run time. Finally, the variable laxity EDF scheduler schedules the tasks considering the predicted execution times. We show that variable execution time scheduling improves the utilization of resources and ensures the quality of service. Our proposed new solution does not require any a priori knowledge of any kind and eliminates the fundamental constraints associated with the estimation of execution times. Results clearly show the advantage of the proposed method as it achieves 76% better task utilization, 68% more task scheduling and deadline miss reduction by 53% compared to current state-of-the-art methods.
Nabeel Iqbal, Jörg Henkel
ICCAD2
2010 Selective instruction set muting for energy-aware adaptive processors
abstract
We propose a new way to save energy in adaptive processors. According to an execution context the custom instruction set of an adaptive processor is selectively 'muted' at run time and thus the energy efficiency is significantly increased. Implemented are multiple so-called 'muting modes' each leading to particular leakage energy savings. A key challenge of this work is to determine which of the muting modes are beneficial for which part of the custom instruction set in a specific execution context. We demonstrate the feasibility by means of an H.264 video encoder (although not limited to that) for various technology nodes. The complex and unpredictable processing behavior of an H.264 encoder represents thereby a real-world scenario. Our results show on average more than 30% energy savings compared to state-of-the-art. We claim that adaptive processors (and reconfigurable computing in general) would be far more energy efficient if FPGA vendors would provide a basic infrastructure that is necessary to exert our novel technique.
Muhammad Shafique 0001, Lars Bauer, Jörg Henkel
ICCAD3
2010 Power-aware complexity-scalable multiview video coding for mobile devices
abstract
We propose a novel power-aware scheme for complexity-scalable multiview video coding on mobile devices. Our scheme exploits the asymmetric view quality which is based on the binocular suppression theory. Our scheme employs different quality-complexity classes (QCCs) and adapts at run time depending upon the current battery state. It thereby enables a run-time tradeoff between complexity and video quality. The experimental results show that our scheme is superior to state-of-the-art and it provides an up to 87% complexity reduction while keeping the PSNR close to the exhaustive mode decision. We have demonstrated the power-aware adaptivity between different QCCs using a laptop with battery charging and discharging scenarios.
Muhammad Shafique 0001, Bruno Zatt, Sergio Bampi, Jörg Henkel
PCS4
2010 An adaptive early skip mode decision scheme for multiview video coding
abstract
In this work a novel scheme is proposed for adaptive early SKIP mode decision in the multiview video coding based on mode correlation in the 3D-neighborhood, variance, and ratedistortion properties. Our scheme employs an adaptive thresholding mechanism in order to react to the changing values of Quantization Parameter (QP). Experimental results demonstrate that our scheme provides a consistent time saving over a wide range of QP values. Compared to the exhaustive mode decision, our scheme provides a significant reduction in the encoding complexity (up to 77%) at the cost of a small PSNR loss (0.172 dB in average). Compared to state-of-the-art, our scheme provides an average 2x higher complexity reduction with a relatively higher PSNR value (avg. 0.2 dB).
Bruno Zatt, Muhammad Shafique 0001, Sergio Bampi, Jörg Henkel
PCS4
2010 Huffman-based code compression techniques for embedded processors
abstract
The size of embedded software is increasing at a rapid pace. It is often challenging and time consuming to fit an amount of required software functionality within a given hardware resource budget. Code compression is a means to alleviate the problem by providing substantial savings in terms of code size. In this article we introduce a novel and efficient hardware-supported compression technique that is based on Huffman Coding. Our technique reduces the size of the generated decoding table, which takes a large portion of the memory. It combines our previous techniques, Instruction Splitting Technique and Instruction Re-encoding Technique into new one called Combined Compression Technique to improve the final compression ratio by taking advantage of both previous techniques. The instruction Splitting Technique is instruction set architecture (ISA)-independent. It splits the instructions into portions of varying size (called patterns) before Huffman coding is applied. This technique improves the final compression ratio by more than 20% compared to other known schemes based on Huffman Coding. The average compression ratios achieved using this technique are 48% and 50% for ARM and MIPS, respectively. The Instruction Re-encoding Technique is ISA-dependent. It investigates the benefits of reencoding unused bits (we call them reencodable bits) in the instruction format for a specific application to improve the compression ratio. Reencoding those bits can reduce the size of decoding tables by up to 40%. Using this technique, we improve the final compression ratios in comparison to the first technique to 46% and 45% for ARM and MIPS, respectively (including all overhead that incurs). The Combined Compression Technique improves the compression ratio to 45% and 42% for ARM and MIPS, respectively. In our compression technique, we have conducted evaluations using a representative set of applications and we have applied each technique to two major embedded processor architectures, namely ARM and MIPS.
Talal Bonny, Jörg Henkel
ACM Trans. Design Autom. Electr. Syst.2
2010 Call for papers ACM transactions on design automation of electronic systems (TODAES) special section on low-power electronics and design
abstract
No abstract available.
Naehyuck Chang, Jörg Henkel
ACM Trans. Design Autom. Electr. Syst.2
2010 Guest Editorial: Current Trends in Low-Power Design
abstract
Low-power consumption of semiconductor devices, circuits, or systems is not only a constraint but a goal of the design process.Initially, low-power design mainly focused on dynamic power consumption.Later, leakage and standby power consumption became important as semiconductor scaled.Recently, lowpower design has expanded its focus to thermal management and green computing.The range of low-power systems now includes power management of large-scale data centers, Grid-scale energy generation, and storage systems as well.The International Symposium on Low-Power Electronics and Design (ISLPED) focuses on recent advances in all aspects of low-power electronics and design, ranging from process and circuit technologies, simulation and synthesis tools, to system-level design and optimization.This special section invited extended versions of distinguished papers at ISLPED 2009.Besides
Naehyuck Chang, Jörg Henkel
ACM Trans. Design Autom. Electr. Syst.2
2009 LICT: left-uncompressed instructions compression technique to improve the decoding performance of VLIW processors
abstract
Compressing program code compiled for VLIW processors to reduce the amount of memory is a necessary means to decrease costs. The main disadvantage of any code compression technique is the system performance penalty because of the extra time required to decode the compressed instructions during run time. In this paper we improve the performance of decoding compressed instructions by using our novel compression technique (LICT: Left-uncompressed Instruction Technique) which can be used in conjunction with any compression algorithm. Furthermore, we adapt a new code compression approach called Burrows-Wheeler (BW) [9] which has been used before in data compression. It significantly reduces the code size compared to state-of-the-art approaches for VLIW processors. Using our LICT in conjunction with the BW algorithm improves the performance explicitly (2.5x) with little impact on the compression ratio (only 3% compression ratio loss).
Talal Bonny, Jörg Henkel
DAC2
2009 Cross-architectural design space exploration tool for reconfigurable processors
abstract
Processors that deploy fine-grained reconfigurable fabrics to implement application-specific accelerators on-demand obtained significant attention within the last decade. They trade-off the flexibility of general-purpose processors with the performance of application-specific circuits without tailoring the processor towards a specific application domain like Application Specific Instruction Set Processors (ASIPs). Vast amounts of reconfigurable processors have been proposed, differing in multifarious architectural decisions. However, it has always been an open question, which of the proposed concepts is more efficient in certain application and/or parameter scenarios. Various reconfigurable processors were investigated in certain scenarios, but never before a systematic design space exploration across diverse reconfigurable processor concepts has been conducted with the aim to aid a designer of a reconfigurable processor. We have developed a first-of-its-kind comprehensive design space exploration tool that allows to systematically explore diverse reconfigurable processors and architectural parameters. Our tool allows presenting the first cross-architectural design space exploration of multiple fine-grained reconfigurable processors on a fair comparable basis. After categorizing fine-grained reconfigurable processors and their relevant parameters, we present our tool and an in-depth analysis of reconfigurable processors within different relevant scenarios.
Lars Bauer, Muhammad Shafique 0001, Jörg Henkel
DATE3
2009 Configurable links for runtime adaptive on-chip communication
abstract
Reliability concerns associated with upcoming technology nodes coupled with unpredictable system scenarios resulting from increasingly complex systems require considering runtime adaptivity in all possible parts of future on-chip systems. We are presenting a novel configurable link which can change its supported bandwidth on-demand at runtime (2X-Links) for an adaptive on-chip communication architecture. We have evaluated our results using real-time multi-media and the E3S application benchmark suits. Our 2X-Links provide a higher throughput of up to 36%, with an average throughput increase of 21.3%, compared to the Normal-Full-Duplex-Links [12], [14], [17], [20] and keep performance-related guarantees with as low as 50% of the Normal-Full-Duplex-Links capacity. Our simulation shows when some links fail, the NoC with 2X-Links can recover from these faults with an average probability of 82.2% whereas these faults would be fatal for the Normal-Full-Duplex-Links.
Mohammad Abdullah Al Faruque, Thomas Ebi, Jörg Henkel
DATE3
2009 Efficient constant-time entropy decoding for H.264
abstract
Diverse approaches to parallel implementation of H.264 have been proposed; however, they all share a common problem. The entropy decoder in H.264 remains mapped on a single processing element (PE). Due to the inherently sequential and context-adaptive nature of the entropy decoder, it cannot be parallelized. This renders a bottleneck to the performance of the entire decoding process. Depending on the type of the processing core and the video bit-rate, the performance of the entire decoding process is subject to the process of entropy decoding. It is, therefore, needful to research and implement new algorithmic solutions to compensate for this bottleneck, and thereby make optimal use of parallel implementation of H.264 decoder on mainstream multi-core systems. This paper presents a new CAVLC decoding method which is de-rived by constructing custom CAVLC decoding tables using dasiatable groupingpsila. Compared to the conventional dasiasequential table look-uppsila method, which requires multiple memory accesses. Our proposed method accesses the custom tables only once for the decoding of any symbol. Moreover, in our proposed method, the symbol decoding time does not depend on the symbol length and it is constant for each symbol, resulting in a nearly linear increase in computational complexity with increase in video fidelity as compared to an non linear increase in earlier proposed methods. Experimental results show that our proposed algorithm features up to 7times higher performance and 83% less memory accesses compared to conventional methods. We compare to three commonly used, state-of-the-art CAVLC algorithms, such as table look-up by sequential search, table look-up by binary search, and ldquoMoon's methodrdquo.
Nabeel Iqbal, Jörg Henkel
DATE2
2009 A parallel approach for high performance hardware design of intra prediction in H.264/AVC Video Codec
abstract
The H.264/AVC Intra Frame Codec (i.e. all frames are coded as I-frames) targets high-resolution/high-end encoding applications (e.g. digital cinema and high quality archiving etc.), providing much better compression efficiency at lower computational complexity compared to MJPEG2000. Moreover, in case of video coding of very high motion scenes, the number of Intra Macroblocks is dominant. Intra Prediction is a compute intensive and memory-critical part that consumes 80% of the computation time of the entire Intra Compression process when executing the H.264 encoder on MIPS processor. We therefore present a novel hardware for H.264 Intra Prediction that processes all the prediction modes in parallel inside one integrated module (i.e. mode-level parallelism) enabling us to exploit the full space of optimization. It exhibits a group-based write-back scheme to reduce the memory transfers in order to facilitate the fast mode-decision schemes. Our Luma 4times4 hardware is 3.6times, 5.2times, and 5.5times faster than state-of-the-art approaches, QS0, respectively. Our results show that processing Luma 16times16, Chroma 8times8, and Luma 4times4 with the proposed approach is 7.2times, 6.5times, and 1.8times faster (while giving an energy saving of 60%, 80%, and 74%) when compared with Dedicated Module Approach (each prediction mode is processed with its independent hardware module i.e. a typical ASIC style for Intra Prediction). We get an area saving of 58% for Luma 4times4 hardware.
Muhammad Shafique 0001, Lars Bauer, Jörg Henkel
DATE3
2009 RISPP: A run-time adaptive reconfigurable embedded processor
abstract
Processors that deploy reconfigurable fabrics to implement application-specific accelerators on-demand obtained significant attention within the last decade. They trade-off the flexibility of general-purpose processors with the performance of application-specific circuits without tailoring the processor towards a specific application domain like application specific instruction set processors (ASIPs). However, even though they reconfigure parts of the hardware at run time, the decisions which accelerators shall be reconfigured at which time are typically determined at compile time. Therefore, it is conceptually not possible to react to dynamically changing situations like varying dynamic control flow (e.g. due to changed input data), changing task priorities / performance constraints, and changing availability of reconfigurable hardware (may be reassigned to another task). This work presents the novel rotating instruction set processing platform (RISPP) with its run-time system that enables dynamic adaptations in the above highlighted scenarios efficiently. Therefore, the presented approach is suitable for long-term as well as frequent adaptation requirements.
Lars Bauer, Muhammad Shafique 0001, Jörg Henkel
FPL3
2009 TAPE: Thermal-aware agent-based power econom multi/many-core architectures
abstract
A growing challenge in embedded system design is coping with increasing power densities resulting from packing more and more transistors onto a small die area, which in turn transform into thermal hotspots. In the current late silicon era silicon structures have become more susceptible to transient faults and aging effects resulting from these thermal hotspots. In this paper we present an agent-based power distribution approach (TAPE) which aims to balance the power consumption of a multi/many-core architecture in a pro-active manner. By further taking the system's thermal state into consideration when distributing the power throughout the chip, TAPE is able to noticeably reduce the peak temperature. In our simulation we provide a fair comparison with the state-of-the-art approaches HRTM [19] and PDTM [9] using the MiBench benchmark suite [18]. When running multiple applications simultaneously on a multi/many-core architecture, we are able to achieve an 11.23% decrease in peak temperature compared to the approach that uses no thermal management [14]. At the same time we reduce the execution time (i.e. we increase the performance of the applications) by 44.2% and reduce the energy consumption by 44.4% compared to PDTM [9]. We also show that our approach exhibits higher scalability, requiring 11.9 times less communication overhead in an architecture with 96 cores compared to the state-of-the-art approaches.
Thomas Ebi, Mohammad Abdullah Al Faruque, Jörg Henkel
ICCAD3
2009 REMiS: Run-time energy minimization scheme in a reconfigurable processor with dynamic power-gated instruction set
abstract
Reconfigurable processors provide a means to flexible and energy-aware computing. In this paper, we present a new scheme for runtime energy minimization (REMiS) as part of a dynamically reconfigurable processor that is exposed to run-time varying constraints like performance and footprint (i.e. amount of reconfigurable fabric). The scheme chooses an energy-minimizing set of so-called Special Instructions (considering leakage, dynamic, and reconfiguration energy) and then 'power-gates' a temporarily unused subset of the Special Instruction set. We provide a comprehensive evaluation for different technologies (ranging from 65 nm to 150 nm) and thereby show that our scheme is technology independent, i.e. it is beneficial for various technologies alike. By means of an H.264 video encoder we demonstrate that for certain performance constraints our scheme (applied to our in-house reconfigurable processor) achieves an allover energy saving of up to 40.8% (avg. 24.8%) compared to a performance-maximizing scheme. We also demonstrate that our scheme is equally beneficial to various other state-of-the-art reconfigurable processor architectures like Molen [9] where it achieves energy savings of up to 48.7% (avg. 28.93%) at 65 nm. We have employed an H.264 encoder within this paper as an application in order to demonstrate the strengths of our scheme, since the H.264's complexity and run-time unpredictability present a challenging scenario for state-of-the-art architectures.
Muhammad Shafique 0001, Lars Bauer, Jörg Henkel
ICCAD3
2009 Non-linear rate control for H.264/AVC video encoder with multiple picture types using image-statistics and motion-based Macroblock Prioritization
abstract
A rate control (RC) algorithm is a primitive block of video encoders that fulfills the bandwidth and buffer constraints for given channel and application properties. State-of-the-art RC schemes perform inefficient in terms of buffer and quality smoothness when handling varying rate distortion characteristics of different picture types (I, P, B) and different MBs in one picture (e.g. bright, textured, static/moving MBs) while posing a high computational overhead. In this paper, we propose a novel RC scheme that covers GOP, picture/ slice, and basic unit levels. It treats different picture types (I, P, B) in a non-linear fashion with consideration of whether they are referenced or non-referenced pictures. Our novel RC scheme prioritizes Macroblocks depending upon their spatial and temporal characteristics (considering eye-catching regions) for refined Quantization Parameter allocation. Compared to RC-Mode-3 (i.e. the latest RC Mode in JM reference software), our RC achieves up to 77.8% and 72.4% reduced buffer-and quality fluctuations, respectively. Compared to RC-Mode-0, our RC provides 2.97dB (i.e. 7.2%) better PSNR for the mixed Susie sequence. Moreover, our proposed RC is 16.6× faster than the RC-Mode-0 when executing on Intel Core2Duo T5500 (1.66 GHz).
Muhammad Shafique 0001, Bastian Molkenthin, Jörg Henkel
ICIP3
2008 Block cache for embedded systems
abstract
On chip memories provide fast and energy efficient storage for code and data in comparison to caches or external memories. We present techniques and algorithms that allow for an automated use of on chip memory for code blocks of instructions which are dynamically scheduled at runtime to increase performance and reduce power consumption.
Dominic Hillenbrand, Jörg Henkel
ASP-DAC2
2008 Run-time instruction set selection in a transmutable embedded processor
abstract
We are presenting a new concept of an application-specific processor that is capable of transmuting its instruction set according to non-predictive application behavior during run-time. In those scenarios, current (extensible) embedded processors are less efficient since they are not run-time adaptive. We have identified the instruction set selection to be a critical step to perform at run time and hence we focus this paper on that crucial part. Our paradigm conducts as many steps as possible at compile/design time and as little as necessary at run time with the constraint to provide a sufficient flexibility to react to non-predictive application behavior efficiently. We provide an in-depth analysis of our scheme and achieve a speed-up of up to 7.19x (average: 3.63x) compared to state-of-the-art adaptive approaches (like [19]). As an application, we have employed a whole H.264 video encoder though our scheme is by principle applicable to many other embedded applications. Our results are evaluated by an implementation of the instruction set selection for our transmutable processor on an FPGA platform.
Lars Bauer, Muhammad Shafique 0001, Jörg Henkel
DAC3
2008 ADAM: run-time agent-based distributed application mapping for on-chip communication
abstract
Design-time decisions can often only cover certain scenarios and fail in efficiency when hard-to-predict system scenarios occur. This drives the development of run-time adaptive systems. To the best of our knowledge, we are presenting the first scheme for a run-time application mapping in a distributed manner using agents targeting for adaptive NoC-based heterogeneous multi-processor systems. Our approach reduces the overall traffic produced to collect the current state of the system (monitoring-traffic), needed for runtime mapping, compared to a centralized mapping scheme. In our experiment, we obtain 10.7 times lower monitoring traffic compared to the centralized mapping scheme proposed in [8] for a 64 x 64 NoC. Our proposed scheme also requires less execution cycles compared to a non-clustered centralized approach. We achieve on an average 7.1 times lower computational effort for the mapping algorithm compared to the simple nearest-neighbor (NN) heuristics proposed in [6] in a 64 x 32 NoC. We demonstrate the advantage of our scheme by means of a robot application and a set of multimedia applications and compare it to the state-of-the-art run-time mapping schemes proposed in [6, 8, 19].
Mohammad Abdullah Al Faruque, Rudolf Krist, Jörg Henkel
DAC3
2008 Run-time System for an Extensible Embedded Processor with Dynamic Instruction Set
abstract
One of the upcoming challenges in embedded processing is to incorporate an increasing amount of adaptivity in order to respond to the multifarious constraints induced by today's embedded systems that feature complex and diverse application behaviors. We present a novel concept (evaluated with a hardware prototype) that moves traditional design-time jobs to run time in order to increase efficiency (in this paper we focus on performance). Adaptivity is achieved dynamically through what we call special instructions (Sis) which may change during run time according to non-predictable application behavior. The new contribution of this paper is the principal component that actually makes the entire embedded processor work efficiently, namely the "special instruction scheduler". It determines during run time 'when' and 'how' Special Instructions are composed and executed. We achieve a 2.38times performance increase over a reconfigurable processor system with dynamic instruction set (Molen). Our whole platform consists of a toolchain including estimation and simulation tools plus a running hardware prototype. Throughout this paper, we discuss the functionality by means of an H.264 video encoder in detail even though the concept is not limited to this application.
Lars Bauer, Muhammad Shafique 0001, Stephanie Kreutz, Jörg Henkel
DATE4
2008 Instruction Re-encoding Facilitating Dense Embedded Code
abstract
Reducing the code size of embedded applications is one of the important constraint in embedded system design. Code compression can provide substantial savings in terms of size. In this paper, we introduce a novel and efficient hardware-supported approach. Our approach investigates the benefits of re-encoding the unused bits (we call them re-encodable bits) in the instruction format for a specific application to improve the compression ratio. Re-encoding those bits may reduce the size of decoding table by more than 37%. We achieve compression ratios as low as 44% (including all overhead that incurs). We have conducted evaluations using a representative set of applications and have applied it to two major embedded processors, namely MIPS and ARM.
Talal Bonny, Jörg Henkel
DATE2
2008 Minimizing Virtual Channel Buffer for Routers in On-chip Communication Architectures
abstract
We present a novel methodology for design space exploration using a two-steps scheme to optimize the number of virtual channel buffers (buffers take the premier share of the router in a NoC) used to implement logical channels multiplexed across the physical channel in a router output port for QoS supported on-chip communication. In the first step, the number of virtual channels is minimized during the mapping of tasks to the NoC at the design time of a system on chip (SoC)for which we use a swarm intelligence-based ant colony optimization (ACO) algorithm. In the second step, a probabilistic approach based on the traffic model of the application is used to further minimize the number of virtual channels. We achieve on average 90.2% reduction in the number of virtual channels compared to a fixed state- of-the-art (i.e. QNoC) allocation for the E3S embedded application benchmark suit. The reduction depends on the designer and the QoS parameter, and it is dependent on the specific application driven traffic model. We demonstrate our design space exploration by means of a complete robot application and also extend our exploration by evaluating the E3S embedded application benchmark suit.
Mohammad Abdullah Al Faruque, Jörg Henkel
DATE2
2008 A computation- and communication- infrastructure for modular special instructions in a dynamically reconfigurable processor
abstract
Processors with a reconfigurable instruction set combine the performance of dedicated application accelerators with a flexibility that goes beyond that of traditional application specific instruction set processors (ASIPs). The latter are optimized for certain application domains and thus typically do not provide a high performance and/or efficiency when deployed in other domains. State-of-the-art reconfigurable processors on the other side still use the concept of monolithic Special Instructions (SIs, i.e. the application accelerators). In our work, we instead present modular SIs as a hierarchy of elementary data paths and different SI implementations that facilitate a high flexibility and performance. This is a novel concept that achieves a speedup of 26.6x compared to a general purpose processor and 1.24x compared to a state-of-the-art reconfigurable processor (that is statically optimized for the predetermined benchmark situation) when executing an H.264 video encoder. We introduce a novel infrastructure for computation and communication that actually enables the implementation of modular SIs and offers various parameters to match specific requirements. The infrastructure is implemented and tested on an FPGA-based prototype to demonstrate its feasibility.
Lars Bauer, Muhammad Shafique 0001, Jörg Henkel
FPL3
2008 FBT: filled buffer technique to reduce code size for VLIW processors
abstract
VLIW processors provide higher performance and better efficiency etc. than RISC processors in specific domains like multimedia applications etc. A disadvantage is the bloated code size of the compiled application code. Therefore, reducing the application code size is a design key issue for VLIW processors. In this paper we adapt a hardware-supported approach called ldquoDeflaterdquo which has been used before in data compression. It can significantly reduce the code size compared to state-of-the-art approaches for VLIW processors as we will show within this work. In fact, we enhance the ldquoDeflaterdquo algorithm by using a new technique called Filled Buffer Technique which can be applied to any Lempel-Ziv family algorithms to improve compression ratio in average by more than 13% compared to the sole ldquoDeflaterdquo algorithm. Using our Filled Buffer Technique in conjunction with ldquoV2Frdquo improves the compression ratio by 10%. We have conducted evaluations using a representative set of benchmarks (from Mediabench and Mibench) and have applied our scheme to two VLIW processors, namely TMS320C62x and TMS320C64x. We achieved allover compression ratios as low as 44% using the ldquoDeflaterdquo algorithm (61% and 56% in average for TMS320C62x and TMS320C64x, respectively).
Talal Bonny, Jörg Henkel
ICCAD2
2008 ROAdNoC: runtime observability for an adaptive network on chip architecture
abstract
Hard-to-predict system behavior and/or reliability issues resulting from migrating to new technology nodes requires considering runtime adaptivity in future on-chip systems. Run-time observability is a prerequisite for runtime adaptivity as it is providing necessary system information gathered on-the-fly. We are presenting the first comprehensive runtime observability infrastructure for an adaptive network on chip architecture which is flexible (e.g. in choosing the routing path), hardly intrusive, and requires little additional overhead (around 0.7% of the total link bandwidth). The hardware overhead is negligible, too, and is in fact less than the hardware savings due to resource multiplexing capabilities that are achieved through runtime observability/adaptivity. As an example, our on-demand buffer assignment scheme increases the buffer utilization and decreases the overall buffer requirements by an average of 42% (the buffer area amounts to about 60% of the entire router area [19]) in our case study analysis compared to a fixed buffer assignment scheme [7]. Our runtime observability on an average also increases the connection success rate by 62% compared to the case without runtime observability for the applications from the E3S benchmark suite [6]. We show the advantages obtained through runtime observability and compare with state-of-the art communication-centric designs.
Mohammad Abdullah Al Faruque, Thomas Ebi, Jörg Henkel
ICCAD3
2008 3-tier dynamically adaptive power-aware motion estimator for h.264/AVC video encoding
abstract
The limitation of energy in portable communication/entertainment devices necessitates the reduction of video encoding complexity. The H.264/AVC video coding standard is one of the latest video codecs and features a complex Motion Estimation scheme that accounts for a major part of the encoder energy [2]. We therefore present a power-aware Motion Estimator for H.264 that adapts at run time according to the available energy level. We perform a set of adaptations at different Processing Stages of Motion Estimation. Our results show that in case of CIF videos (typically used in portable devices; but our approach is equally applicable to other video resolutions too), we achieve an average energy reduction of 52 times and 27 times as compared to UMHexagonS [12] and EPZS [14] respectively. This energy saving comes at the cost of an average loss of only 0.39 dB in Peak Signal to Noise Ratio (PSNR: an objective quality measure) and 23% increase in area (synthesized for 90nm technology).
Muhammad Shafique 0001, Lars Bauer, Jörg Henkel
ISLPED3
2008 Efficient Resource Utilization for an Extensible Processor Through Dynamic Instruction Set Adaptation
abstract
State-of-the-art application-specific instruction set processors (ASIPs) allow the designer to define individual prefabrication customizations, thus improving the degree of specialization towards the actual application requirements, e.g., the computational hot spots. However, only a subset of hot spots can be targeted to keep the ASIP within a reasonable size. We propose a modular special instruction composition with multiple implementation possibilities per special instruction, compile-time embedded instructions to trigger a run-time adaptation of the instruction set, and a run-time system that dynamically selects an appropriate variation of the instruction set, i.e., a situation-dependent beneficial implementation for each special instruction. We thereby achieve a better efficiency of resource usage of up to 3.0 times (average 1.4 times) compared with current state-of-the-art ASIPs, resulting in a 3.1 times (average 1.4 times) improved application performance (compared with a general purpose processor up to 25.7 times and average 17.6 times).
Lars Bauer, Muhammad Shafique 0001, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.3
2008 Efficient Code Compression for Embedded Processors
abstract
Code density is of increasing concern in embedded system design since it reduces the need for the scarce resource memory and also implicitly improves further important design parameters like power consumption and performance. In this paper we introduce a novel, hardware-supported approach. Besides the code, also thelookuptables(LUTs) are compressed, that can become significant in size if the application is large and/or high compression is desired. Our scheme optimizes the number and size of generated LUTs to improve the compression ratio. To show the efficiency of our approach, we apply it to two compression schemes: ldquodictionary-basedrdquo and ldquostatisticalrdquo. We achieve an average compression ratio of 48% (already including the overhead of the LUTs). Thereby, our scheme is orthogonal to approaches that take particularities of a certain instruction set architecture into account. We have conducted evaluations using a representative set of applications and have applied it to three major embedded processor architectures, namely ARM, MIPS, and PowerPC.
Talal Bonny, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.2
2008 Guest Editorial Special Section on Low-Power Electronics and Design
abstract
The eight papers in this special section were published in the ACM/IEEE International Symposium on Low Power Electronics and Design (ISLPED), held in Tegernsee, Germany, October 4-6, 2006. The papers cover various areas related to low-power electronics and design, ranging from circuit and design technology to system level power modeling and optimization. The papers are summarized here.
Diana Marculescu, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.2
2007 Transaction Specific Virtual Channel Allocation in QoS Supported On-chip Communication
abstract
We propose a scheme to reduce the number of virtual channel buffers used to implement logical channels multiplexed across the physical channel in a router output port for QoS supported on-chip communication. The number of virtual channels is minimized during the mapping of tasks to the NoC during the design time of a System on Chip (SoC) and a swarm intelligence-based Ant Colony Optimization (ACO) algorithm is used. We achieve on average 88% reduction in the number of virtual channels compared to a fixed allocation for the E3S embedded application benchmark suit and a collection of existing real-time embedded applications.
Mohammad Abdullah Al Faruque, Jörg Henkel
ASAP2
2007 RISPP: Rotating Instruction Set Processing Platform
abstract
Adaptation in embedded processing is key in order to address efficiency. The concept of extensible embedded processors works well if a few a-priori known hot spots exist. However, they are far less efficient if many and possible at-design-time-unknown hot spots need to be dealt with. Our RISPP approach advances the extensible processor concept by providing flexibility through runtime adaptation by what we call "instruction rotation". It allows sharing resources in a highly flexible scheme of compatible components (called Atoms and Molecules). As a result, we achieve high speed-ups at moderate additional hardware. Furthermore, we can dynamically tradeoff between area and speed-up through runtime adaptation. We present the main components of our platform and discuss by means of an H.264 video codec.
Lars Bauer, Muhammad Shafique 0001, Simon Kramer 0002, Jörg Henkel
DAC4
2007 Instruction Splitting for Efficient Code Compression
abstract
The size of embedded software is rising at a rapid pace. It is often challenging and time consuming to fit an amount of required software functionality within a given hardware resource budget. Code compression is a means to alleviate the problem. In this paper we introduce a novel and efficient hardware-supported approach. Our scheme reduces the size of the generated decoding table by splitting instructions into portions of varying size (called patterns) before Huffman Coding compression is applied. It improves the final compression ratio (including all overhead that incurs) by more than 20% compared to known schemes based on Huffman Coding. We achieve allover compression ratios as low as 44%. Thereby, our scheme is orthogonal to approaches that take particularities of a certain instruction set architectures into account. We have conducted evaluations using a representative set of applications and have applied it to two major embedded processors, namely ARM and MIPS.
Talal Bonny, Jörg Henkel
DAC2
2007 Efficient code density through look-up table compression
abstract
Code density is a major requirement in embedded system design since it not only reduces the need for the scarce resource memory but also implicitly improves further important design parameters like power consumption and performance. Within this paper we introduce a novel and efficient hardware-supported approach that belongs to the group of statistical compression schemes as it is based on canonical Huffman coding. In particular, our scheme is the first to also compress the necessary Look-up Tables that can become significant in size if the application is large and/or high compression is desired. Our scheme optimizes the number of generated look-up tables to improve the compression ratio. In average, we achieve compression ratios as low as 49% (already including the overhead of the lookup tables). Thereby, our scheme is entirely orthogonal to approaches that take particularities of a certain instruction set architecture into account. We have conducted evaluations using a representative set of applications and have applied it to three major embedded processor architectures, namely ARM, MIPS and PowerPC
Talal Bonny, Jörg Henkel
DATE2
2007 Instruction trace compression for rapid instruction cache simulation
abstract
Modern application specific instruction set processors (ASIPs) have customizable caches, where the size, associativity and line size can all be customized to suit a particular application. To find the best cache size suited for a particular embedded system, the applications) is/are executed, traces obtained, and caches simulated. Typically, program trace files can range from a few megabytes to several gigabytes. Simulation of cache performance using large program trace files is a time consuming process. In this paper, a novel instruction cache simulation methodology that can operate directly on a compressed program trace file without the need for decompression is presented. This feature allowed our simulation methodology to have an average speed up of 9.67 times compared to the existing state of the art tool (Dinero IV cache simulator), for a range of applications from the Mediabench suite
Andhi Janapsatya, Aleksandar Ignjatovic, Sri Parameswaran, Jörg Henkel
DATE4
2007 Run-time adaptive on-chip communication scheme
abstract
During run-time varying workloads and/or constraints in embedded systems require run-time adaptivity to provide a high degree of efficiency during any operation mode/scenario. Design time decisions can often only cover certain scenarios and fail in efficiency when hard-to-predict system scenarios occur. We are presenting the first approach of an adaptive on-chip communication scheme. It provides an adaptive routing/path allocation algorithm to meet a required level of QoS (guaranteed bandwidth). In our architecture adaptive runtime links are established by re-assigning buffer blocks ondemand. This adaptive buffer allocation scheme increases the buffer utilization and decreases the overall buffer use on an average of 42% in our case study analysis compared to a fixed buffer assignment strategy. The area overhead introduced by the adaptive scheme can be traded-off against the flexibility in order to select an available path and on-demand buffer allocation. We demonstrate the advantage by using various real world digital media applications and compare our approach to the state-ofthe- art static on-chip communication schemes.
Mohammad Abdullah Al Faruque, Thomas Ebi, Jörg Henkel
ICCAD3
2006 Using Lin-Kernighan algorithm for look-up table compression to improve code density
abstract
The presented work uses code compression to improve the design efficiency of an embedded system. In particular, we present a method and architecture for compressing the so-called Look-up Tables that are necessary for the de-compression process. No other work has yet focused on minimizing the Look-up Tables that, as we show, have a significant impact on the total overhead of a hardware-based decompression scheme. We introduce a novel and very efficient hardware-supported approach based on Canonical Huffman Coding. Using the Lin-Kernighan algorithm we reduce the Look-up Table size by up to 45%. As a result, we achieve all-over compression ratios as low as 45% (already including the overhead of the Look-up Tables). Thereby, our scheme is entirely orthogonal to approaches that take particularities of a certain instruction set architecture into account, meaning that compression could be further improved. Factoring in the orthogonality, our scheme is the basis for not-yet-achieved efficiency in hardware-supported compression schemes. We have conducted evaluations using a representative set (in terms of size and application domain) of applications and have applied it to three major embedded processor architectures, namely ARM, MIPS and PowerPC. The hardware evaluation shows no performance penalty.
Talal Bonny, Jörg Henkel
ACM Great Lakes Symposium on VLSI2
2006 A design methodology for application-specific networks-on-chip
abstract
With the help of HW/SW codesign, system-on-chip (SoC) can effectively reduce cost, improve reliability, and produce versatile products. The growing complexity of SoC designs makes on-chip communication subsystem design as important as computation subsystem design. While a number of codesign methodologies have been proposed for on-chip computation subsystems, many works are needed for on-chip communication subsystems. This paper proposes application-specific networks-on-chip (ASNoC) and its design methodology. ASNoC is used for two high-performance SoC applications. The methodology (1) can automatically generate optimized ASNoC for different applications, (2) can generate a corresponding distributed shared memory along with an ASNoC, (3) can use both recorded and statistical communication traces for cycle-accurate performance analysis, (4) is based on standardized network component library and floorplan to estimate power and area, (5) adapts an industrial-grade network modeling and simulation environment, OPNET, which makes the methodology ready to use, and (6) can be easily integrated into current HW/SW codesign flow. Using the methodology, ASNoC is generated for a H.264 HDTV decoder SoC and Smart Camera SoC. ASNoC and 2D mesh networks-on-chip are compared in performance, power, and area in detail. The comparison results show that ASNoC provide substantial improvements in power, performance, and cost compared to 2D mesh networks-on-chip. In the H.264 HDTV decoder SoC, ASNoC uses 39% less power, 59% less silicon area, 74% less metal area, 63% less switch capacity, and 69% less interconnection capacity to achieve 2X performance compared to 2D mesh networks-on-chip.
Jiang Xu 0001, Marilyn Wolf, Jörg Henkel, Srimat T. Chakradhar
ACM Trans. Embed. Comput. Syst.3
2006 Distance-based recent use (DRU): an enhancement to instruction cache replacement policies for transition energy reduction
abstract
According to the International Technology Roadmap for Semiconductors (ITRS), the minimum feature size for microprocessors will shrink to 40 nm by 2010. Leakage currents in devices fabricated at these dimensions have been shown to be so dominant that design methodologies driven by power budgets will face challenges in reducing static power in addition to active power. An effective solution to tackle static power is to transition devices to a low-static-power sleep mode using special circuit-level techniques. However, these transitions come with energy costs, and as these techniques are perfected, and devices transition more often to sleep state, the relative contribution of transition energy to total energy will increase. To deal with the transition overhead, often used techniques are history-based and concentrate only on recognizing when to transition, but do not provide for reducing total transitions without adversely effecting the total sleep time of the devices. In this paper, we study transition-overhead reduction in associative instruction caches. We take advantage of the fact that many programs, particularly those for multimedia applications, spend most of their time in loops and most execution is near-sequential (high spatial locality). We present a technique called DRU (Distance-based Recent Use), which constrains near-sequential fetches to a single bank from the set of associative banks. Evaluation of DRU for different replacement policies in a system-level environment using Mediabench's applications and with various processor architectures (including SPARC and MIPS) have shown energy savings between 20%-28% with negligible hardware and timing overheads.
Praveen Kalla, Xiaobo Sharon Hu, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.3
2005 Battery-aware instruction generation for embedded processors
abstract
Automatic instruction generation is an efficient method to satisfy growing performance and meet design constraints for application specific instruction-set processors. A typical approach for instruction generation is to combine a large group of primitive instructions into a single extensible instruction for maximizing speedups. However, this approach often leads to large power dissipation and discharge current, posing a challenge to battery-powered products. In this paper, we propose a battery-aware automatic tool to design extensible instructions which minimizes power dissipation distribution by separating an instruction into multiple instructions. We verify our automatic tool using 50 different code segments, and five large real-world applications. Our tool reduces energy consumption by a further 5.8% on average (up to 17.7%) compared to extensible instructions generated by previous approaches. For real-world applications, energy consumption is reduced by 6.6% on average (up to 16.53%) as well as an increase in performance for most cases. The automatic instruction generation tool is integrated into our application specific instruction-set processor tool suite.
Newton Cheung, Sri Parameswaran, Jörg Henkel
ASP-DAC3
2005 A flexible framework for communication evaluation in SoC design
abstract
We present SoCExplore, a framework for fast communication-centric design space exploration of complex SoCs with network-based interconnects. Speed-up in exploration is achieved through abstraction of computation as a high-level trace, and accuracy is maintained through cycle-accurate interconnect simulation. The flexibility offered allows for fast partition/mapping and interconnect design space exploration. Error analysis of such frameworks is non-trivial and is presented for the first time. As a case study, a speed-up of 94% over architectural simulation is reported for the MPEG application.
Praveen Kalla, Xiaobo Sharon Hu, Jörg Henkel
ASP-DAC3
2005 H.264 HDTV Decoder Using Application-Specific Networks-On-Chip
abstract
This paper studied an H. 264 HDTV decoder on two multiprocessor system-on-chip architectures. Two types of networks-on-chip, the RAW network and the application specific networks-on-chip, were used. Regular-topology networks-on-chip (mesh, torus, and fat tree) have been proposed. However, we showed in this paper that the application-specific networks-on-chip provided substantial improvements in power, performance, and cost compared to regular-topology networks-on-chip. We measured the power, performance, area, total switch and link capacity, and switch and link utilization based on floorplans and circuit designs. Measurement results showed th at the application-specific networks-on-chip was both faster in absolute terms and more efficient. The application-specific networks-on-chip used 39% less power, 59% less silicon area, 74% less metal area, 63% less switch capacity, and 69% less link capacity to achieve 2X performance compared to the RAW network.
Jiang Xu 0001, Marilyn Wolf, Jörg Henkel, Srimat T. Chakradhar
ICME3
2005 Approximate arithmetic coding for bus transition reduction in low power designs
abstract
We present a method for reducing the power consumption of compressed-code systems by selectively inverting bits that are transmitted on the bus. By incorporating bus inversion into code compression/decompression, we reduce power consumption with no cost in hardware or power relative to code compression without inversion. Inverting has to be done carefully to ensure that the codes can still be decoded. As an additional challenge, compression will generally increase bit-toggling as it removes redundancies from the code transmitted. Therefore, we need to find the right balance between compression ratio and bit-toggling reduction. This paper presents a suitable algorithm that will combine approximate compression techniques with bit-toggling reduction and will explore the various tradeoffs. We take advantage of the approximations introduced to modify codes and reduce bit-toggling, while maintaining compression performance and decoding speed. An interesting result that is derived from our work is that high compression ratios do not necessarily result in the lowest power consumption. By using our method, bus-related power consumption has been reduced by as much as 35% compared to a system with no compression, and as much as 14% compared to a compressed-code system. Bit-toggling reduction does not impose any additional hardware costs other than the decompression engine. We also present a detailed analysis on how bus widths affect bit-toggling when transmitting compressed code, and we show experimental results on ARM, MIPS, and SPARC code. We finally compare our work with Bus Invert and show results that are superior except for the random data case where Bus Invert performs better.
Haris Lekatsas, Jörg Henkel, Marilyn Wolf
IEEE Trans. Very Large Scale Integr. Syst.2
2005 Instruction code mapping for performance increase and energy reduction in embedded computer systems
abstract
In this paper, we present a novel and fast constructive technique that relocates the instruction code in such a manner into the main memory that the cache is utilized more efficiently. The technique is applied as a preprocessing step, i.e., before the code is executed. Our technique is applicable in embedded systems where the number and characteristics of tasks running on the system is known a priori. The technique does not impose any computational overhead to the system. As a result of applying our technique to a variety of real-world applications we observed through simulation a significant drop of cache misses. Furthermore, the energy consumption of the whole system (CPU, caches, buses, main memory) is reduced by up to 65%. These benefits could be achieved by a slightly increased main memory size of about 13% on average.
Sri Parameswaran, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.2
2004 MINCE: Matching INstructions Using Combinational Equivalence for Extensible Processor
abstract
Designing custom-extensible instructions for extensible processors is a computationally complex task because of the large design space. The task of automatically matching candidate instructions in an application (e.g. written in a high-level language) to a pre-designed library of extensible instructions is especially challenging. Previous approaches have focused on identifying extensible instructions (e.g. through profiling), synthesizing extensible instructions, estimating expected performance gains etc. In this paper we introduce our approach of automatically matching extensible instructions as this key step is missing in automating the entire design flow of an ASIP with extensible instruction capabilities. Since matching using simulation is practically infeasible (simulation time), and traditional pattern matching approaches would not yield reliable results (ambiguity related to a functionally equivalent code that can be represented in many different ways), we adopt combinational equivalence checking. Our MINCE tool as part of our ASIP design flow consists of a translator, a filtering algorithm and a combinational equivalence checking tool. We report matching times of extensible instructions that are 7.3x faster on average (using Mediabench applications) compared to the best known approaches to the problem (partial simulations). In all our experiments MINCE matched correctly and the outcome of the matching step yielded an average speedup of the application of 2.47x. As a summary, our work represents a key step towards automating the whole design flow of an ASIP with extensible instruction capabilities.
Newton Cheung, Sri Parameswaran, Jörg Henkel, Jeremy Chan
DATE3