EDBT 2026 Demo / reviewers in the wild / expert
Heba Khdr
dblp:129/2048
· DBLP profile ↗
42ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0003-0245-2062ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 8 first-author · 25 since 2021Software engineering, systems software and programming languages · 14 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Federated Learning with Low-Rank Updates under Homomorphic EncryptionabstractFederated Learning has been widely adopted for its ability to collaboratively train models without exposing raw data. However, the server-side aggregation process may still leak sensitive information about client data. Homomorphic Encryption enables privacy-preserving aggregation, but it introduces substantial communication overhead for clients and high computational costs for the server. To address these challenges, we propose HEAL-FL, a federated learning framework that is based on low-rank shared basis vectors across clients. Instead of transmitting full encrypted model updates, clients send only encrypted low-rank coefficients, thereby reducing both communication costs and server-side aggregation overhead. Furthermore, HEAL-FL incorporates a communication-efficient basis update scheme that relies exclusively on homomorphic addition at the server. Our evaluation across various homomorphic encryption schemes shows that HEAL-FL reduces client communication and server aggregation costs, leading to improved efficiency of Federated Learning systems. Notably, these savings translate into up to a significant reduction of 38.6% in total training time compared to conventional homomorphic FedAvg with full model parameter transmission, demonstrating the practical benefits of our approach. Mohamed Aboelenien Ahmed, Mohamed Alsharkawy, Hassan Nassar, Heba Khdr, Jeferson González-Gómez, Jörg Henkel |
DATE | 4 |
| 2026 | TrustSeed: Lightweight Attestation Protocol for Ensuring LLM IntegrityabstractOver the last couple of years, large language models have increasingly been integrated into many computing applications. For privacy preservation, they are now deployed on edge devices. However, these deployments are vulnerable to bit flip attacks and backdoor attacks that compromise the integrity of the model. Traditional remote attestation techniques fail to detect such manipulations due to the large model size and the stealthiness of the attacks.In this paper, we present TrustSeed, a lightweight functional attestation protocol that uses a single inference to ensure large language models’ integrity. TrustSeed verifies integrity by applying deterministic, seed-based modifications to model weights within a Trusted Execution Environment and comparing the last intermediate activations and output distribution against a golden reference on the verifier. This approach prevents precomputed or forged responses, ensuring freshness and unpredictability in each attestation round. Our analysis shows that output distribution and last intermediate activations are effective indicators of integrity. We test TrustSeed against bit-flip, data poisoning, and weight poisoning attacks, reliably detecting even single-bit alterations. Extensive evaluations on edge platforms and an HPC system demonstrate minimal overhead and up to 127× faster attestation compared to state-of-the-art full-model hashing. Mohamed Alsharkawy, Mohamed Aboelenien Ahmed, Hassan Nassar, Jeferson González-Gómez, Heba Khdr, Osama Abboud, Xun Xiao, Jörg Henkel |
DATE | 5 |
| 2026 | Multi-Partner Project: A Holistic and Open-Source Approach to Efficient, Secure and Reliable AI Hardware Deployment in DI-EDAIabstractArtificial Intelligence (AI) has demonstrated strong capabilities across various domains over the past decade. Edge and specifically mission-critical applications, such as automotive and aerospace, require both high performance and efficiency without compromises in security and reliability. This stems from tightly constrained power consumption, failures that can have catastrophic consequences and devices that may be physically accessible to malicious actors. AI algorithm deployment to hardware also presents significant barriers, requiring specialized knowledge and expensive development tools. The DI-EDAI project aims to offer a holistic approach for connecting high-level AI algorithms with hardware implementations while tackling the aforementioned issues. Unlike other approaches that address individual aspects of the AI deployment flow, we investigate solutions across multiple layers of the design stack. Through our work we develop efficient hardware, map AI algorithms to hardware while simultaneously ensuring security and reliability. Furthermore, we leverage AI-techniques to assist with Electronic Design Automation (EDA) workflows for design optimization, verification and implementation. Our open source approach aims to reduce entry barriers, promote transparency and education, and spark innovation. This paper presents the current state of the DI-EDAI project at midterm, highlighting our latest contributions, identifying limitations in existing state-of-the-art approaches, and outlining ongoing work to address these gaps. Georgios Sotiropoulos, Felix Frombach, Julian Höfer, Tanja Harbaum, Jürgen Becker 0001, Henrik Iver Thorøe, Vincent Meyers, Mehdi Baradaran Tahoori, Zeynep Demirdag, Mohammed Bakr Sikal, Hassan Nassar, Heba Khdr, Jörg Henkel, Christopher Wolters, Philipp van Kempen, Johannes Geier, Ulf Schlichtmann, Batuhan Sesli, Muhammad Sabih, Jakob Wittmann, Frank Hannig, Jürgen Teich, Lukas Steiner, Norbert Wehn, Mohamed Shelkamy Ali, Philipp Schmitz, Wolfgang Kunz, Stefan Koegler, Georg Sigl |
DATE | 12 |
| 2026 | ARDiS: A Portable and Unified Resource Management Framework in Real Hardware SystemsabstractDesigning efficient RM strategies is a cornerstone of modern computing, driving innovations in performance optimization, energy efficiency, and security. While simulators have long been the go-to tools for RM research, they fail to balance accuracy and practicality: high-fidelity simulators are excruciatingly slow, and low-fidelity ones compromise on reliability. Real hardware offers unparalleled precision and accuracy but remains underutilized due to significant barriers, including fragmented implementations, lack of portability, and prohibitive development overhead. We present ARDiS, the first open-source 1 and portable framework to provide a unified, architecture-agnostic platform for running system-level resource management (RM) techniques directly on real hardware. ARDiS eliminates the need to “reinvent the wheel,” enabling researchers to design, implement, and evaluate sophisticated RM strategies—including machine learning-based approaches—with minimal effort and maximum reproducibility. To demonstrate its versatility, we evaluate ARDiS on two real-world hardware platforms: a server-grade heterogeneous processor (Intel i9-12900) and a resource-constrained embedded system (NVIDIA Jetson TX2). Through extensive experimentation, we validate the ability of ARDiS to deliver accurate, scalable, and reproducible results across diverse platforms and application domains. By lowering the barriers to hardware-based RM research, ARDiS empowers the design automation community to explore new frontiers in system-level optimization and innovation. Mohammed Bakr Sikal, Jeferson González-Gómez, Andreas Noebel, Heba Khdr, Jörg Henkel |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2025 | Special Session - Hardware-Software Co-Design for Machine Learning Systems Made Open-SourceabstractChip technologies are crucial for the digital transformation of industry and society. Machine Learning (ML) and Artificial Intelligence (AI) are increasingly shaping both daily life and industrial applications, with AI hardware playing a vital role in enabling efficient and scalable ML deployment. However, significant challenges remain in bridging the gap between ML algorithm development and hardware implementation, particularly for edge ML applications where efficiency, power constraints, and adaptability are critical. In such resource-constrained environments, hardware-software co-design becomes essential to achieve the necessary trade-offs between performance, energy efficiency, and system responsiveness. One of the key bottlenecks in ML hardware development is the lack of seamless integration between ML toolchains and electronic design automation (EDA) tools for hardware synthesis and mapping. Current solutions often require extensive manual optimization and costly proprietary software, limiting accessibility and innovation. Open-source tools can play a transformative role in democratizing ML hardware design, fostering collaboration, and addressing the growing shortage of skilled professionals. This paper covers key aspects of hardware-software co-design for ML systems, such as ML algorithms, hardware design, compiler technologies and system security, with a focus on open-source solutions. We highlight the critical need for open-source toolchains that connect ML model development with hardware synthesis and optimization and present solutions for custom hardware, as well as FPGA accelerators. Mehdi Baradaran Tahoori, Vincent Meyers, Mahboobe Sadeghipourrudsari, Huashuangyang Xu, Jürgen Becker 0001, Tanja Harbaum, Felix Frombach, Julian Höfer, Georgios Sotiropoulos, Jörg Henkel, Zeynep Demirdag, Heba Khdr, Hassan Nassar, Ulf Schlichtmann, Johannes Geier, Philipp van Kempen, Georg Sigl, Stefan Koegler, Matthias Probst, Jürgen Teich, Frank Hannig, Muhammad Sabih, Batuhan Sesli, Norbert Wehn, Lukas Steiner, Wolfgang Kunz, Mohamed Shelkamy Ali |
CODES+ISSS | 12 |
| 2025 | Centralized Training and Decentralized Control through the Actor-Critic Paradigm for Highly Optimized MulticoresabstractWhile distributed, neural-network-based resource controllers represent the state of the art for their ability to cope with the ever-expanding decision space, such approaches suffer from several limitations, like conflicting control decisions and partial observability. These effects can significantly impair the controllers’ learning capabilities and the stability of their control policies, causing substantial performance losses. We are the first to solve this problem employing a centralized training and decentralized control regime to mitigate the aforementioned limitations. Specifically, we design a centralized neural network (critic) that evaluates the behavior of multiple decentralized neural controllers (actors) in a system-wide context. The objective of our proposed technique is to maximize the performance under a temperature constraint through dynamic voltage frequency scaling. The evaluation of our technique shows its superiority over the state of the art, yielding average (peak) performance improvements of 20% (34%), which we consider a breakthrough as the gains are measured on a real-world platform. Benedikt Dietrich, Heba Khdr, Jörg Henkel |
DAC | 2 |
| 2025 | Contention-Aware Forecasting of Energy Efficiency through Sequence-Based Models in Modern Heterogeneous ProcessorsabstractWe present EffiCast, the first methodology for contentionaware energy efficiency forecasting in clustered heterogeneous processors using sequence-based models. Through extensive experimental analysis of energy efficiency sensitivities across core types, voltage/frequency (V/f) levels, application phases, and resource contention scenarios, EffiCast uncovers key factors driving energy efficiency variability in modern heterogeneous processors. Leveraging structured data generation and advanced LSTM- and Transformer-based models, EffiCast achieves unprecedented accuracy while outperforming state-of-the-art predictive techniques. Deployed on a real heterogenous processor with Intel’s oneDNN acceleration, EffiCast delivers inference latencies as low as 1.82 ms per sequence, enabling seamless integration into proactive resource management frameworks. With the ability to forecast future system states under dynamic workloads, EffiCast sets a new standard for energy efficiency optimization in energy-constrained application domains. Mohammed Bakr Sikal, Jeferson González-Gómez, Heba Khdr, Jörg Henkel |
DAC | 3 |
| 2025 | Federated Reinforcement Learning for Optimizing the Power Efficiency of Edge DevicesabstractReinforcement learning (RL) holds great promise for adaptively optimizing microprocessor performance under power constraints. It allows for online learning of application characteristics at runtime and enables adjustment to varying system dynamics such as changes in the workload, user preferences or ambient conditions. However, online policy optimization remains resource-intensive, with high computational demand and requiring many samples to converge, making it challenging to deploy to edge devices. In this work, we overcome both of these obstacles and present federated power control using dynamic voltage and frequency scaling (DVFS). Our technique leverages federated RL and enables multiple independent power controllers running on separate devices to collaboratively train a shared DVFS policy, consolidating experience from a multitude of different applications, while ensuring that no privacy-sensitive information leaves the devices. This leads to faster convergence and to increased robustness of the learned policies. We show that our federated power control achieves 57 % average performance improvements over a policy that is only trained on local data. Compared to a state-of-the-art collaborative power control, our technique leads to 22 % better performance on average for the running applications under the same power constraint. Benedikt Dietrich, Rasmus Müller-Both, Heba Khdr, Jörg Henkel |
DATE | 3 |
| 2025 | Hardware/Software Co-Analysis for Worst Case Execution Time BoundsabstractEnsuring that safety-critical systems meet timing constraints is crucial to avoid disastrous failures. To verify that timing requirements are met, a worst-case execution time (WCET) bound is computed. However, traditional WCET tools require a predefined timing model for each target processor, which is not available when using custom instruction set extensions. We introduce a novel approach based on hardware-software coanalysis that employs an instrumented hardware description of the target processor, removing the requirement for a separate timing model. We demonstrate this approach by extending the FemtoRV32 Individua RISC-V processor with a custom instruction set extension and show that it accurately models the timing behavior of the resulting system. Can Joshua Lehmann, Lars Bauer, Hassan Nassar, Heba Khdr, Jörg Henkel |
DATE | 4 |
| 2025 | Invited Paper: Hardware-Software Co-Design for Highly Optimized, Customized, and Reliable AI SystemsabstractOver the past decade, AI has been rapidly integrated into our daily life, coming in every shape and size and working across systems from big clouds to IoT. As a result, AI systems are increasingly requiring enhancements in model efficiency, hardware acceleration, and memory systems to satisfy stringent constraints on efficiency, reliability, and security. However, advancing across these fronts is challenging as compute demand outpaces Moore’s-law efficiency, hardening into an AI compute wall and an AI energy wall. Breaking through requires a unified AI co-design loop that co-optimizes algorithms and hardware, including efficient AI-to-hardware mapping, so that ongoing goals (accuracy, sparsity, latency) align with concrete hardware choices (precision modes, interconnects, memory hierarchies) and AI-specific execution and memory-reuse patterns. This paper details the principal co-design challenges, presents complementary strategies, and outlines a practical roadmap toward highly optimized, efficient, reliable, and secure AI systems. Jörg Henkel, Mehdi Baradaran Tahoori, Heba Khdr, Hassan Nassar, Vincent Meyers, Deming Chen, Selin Yildirim, Yingbing Huang, Nirmal Saxena, Saurabh Hukerikar, Srivi Dhruvanarayan |
ICCAD | 3 |
| 2024 | Multi-Agent Reinforcement Learning for Thermally-Restricted Performance Optimization on ManycoresabstractThe problem of performance maximization under a thermal constraint has been tackled by means of dynamic voltage and frequency scaling (DVFS) in many system-level optimization techniques. State-of-the-art ones have exploited Su-pervised Learning (SL) to develop models that predict power and performance characteristics of applications and temperature of the cores. Such predictions enable proactive and efficient optimization decisions that exploit performance potentials under a temperature constraint. SL- based models are built at design time based on training data generated considering specific environment settings, i.e., processor architecture, cooling system, ambient temperature, etc. Hence, these models cannot adapt at runtime to different environment settings. In contrast, Reinforcement Learning (RL) employs an agent that explores and learns the environment at runtime, and hence can adapt to its potential changes. Nonetheless, using an RL agent to perform optimization on manycores is challenging because of the inherent large state/action spaces that might hinder the agent's ability to converge. To get the advantages of RL while tackling this challenge, we employ for the first time multi -agent RL to perform thermally-restricted performance optimization for manycores through DVFS. We investigated two RL algorithms-Table-based Q-Learning (TQL) and Deep Q-Learning (DQL)-and demonstrated that the latter outperforms the former. Compared to the state of the art, our DQL delivers a significant performance improvement of 34.96% on average, while also guaranteeing thermally -safe operation on the manycore. Our evaluation reveals the runtime adaptability of our DQL to varying workloads and ambient temperatures. Heba Khdr, Mustafa Enes Batur, Kanran Zhou, Mohammed Bakr Sikal, Jörg Henkel |
DATE | 1 |
| 2024 | Balancing Security and Efficiency: System-Informed Mitigation of Power-Based Covert ChannelsabstractAs the digital landscape continues to evolve, the security of computing systems has become a critical concern. Power-based covert channels (e.g., thermal covert channel s (TCCs)), a form of communication that exploits the system resources to transmit information in a hidden or unintended manner, have been recently studied as an effective mechanism to leak information between malicious entities via the modulation of CPU power. To this end, dynamic voltage and frequency scaling (DVFS) has been widely used as a countermeasure to mitigate TCCs by directly affecting the communication between the actors. Although this technique has proven effective in neutralizing such attacks, it introduces significant performance and energy penalties, that are particularly detrimental to energy-constrained embedded systems. In this article, we propose different system-informed countermeasures to power-based covert channels from the heuristic and machine learning (ML) domains. Our proposed techniques leverage task migration and DVFS to jointly mitigate the channels and maximize energy efficiency. Our extensive experimental evaluation on two commercial platforms: 1) the NVIDIA Jetson TX2 and 2) Jetson Orin shows that our approach significantly improves the overall energy efficiency of the system compared to the state-of-the-art solution while nullifying the attack at all times. Jeferson González-Gómez, Mohammed Bakr Sikal, Heba Khdr, Lars Bauer, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | ML-Based Thermal and Cache Contention Alleviation on Clustered Manycores With 3-D HBMabstractEnabled by the recent advancements in 2.5D/3-D integration and packaging, the integration of clustered manycore processors with high-bandwidth memory (HBM) is gaining prominence to satisfy the increasing memory bandwidth demands. Although this integration can offer significant performance gains, it is still limited by cache contention in the final-level cache on the clusters and by the thermal issues in the 3-D HBM. While the existing state-of-the-art resource management techniques have tackled these issues in isolation, we argue that the cache contention and the temperature of both the manycore and the HBM must be considered jointly to harness the full performance potential of such modern architectures. To cover this gap in the literature, we present MTCM, the first resource management technique that considers the cache contention in maximizing the system performance, while maintaining the thermal safety across both the manycore and the HBM stack. Enabled by our accurate, yet lightweight, neural network models, our proposed task migration and dynamic voltage and frequency scaling policies can accurately predict the impact of runtime decisions on the performance and temperature of both the subsystems. Our extensive evaluation experiments reveal a significant performance improvement over existing state of the art by up to$1\times $, while maintaining thermal safety of both the manycore and the HBM. Mohammed Bakr Sikal, Heba Khdr, Lokesh Siddhu, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | NPU-Accelerated Imitation Learning for Thermal Optimization of QoS-Constrained Heterogeneous Multi-CoresabstractThermal optimization of a heterogeneous clustered multi-core processor under user-defined QoS targets requires application migration and DVFS. However, selecting the core to execute each application and the VF levels of each cluster is a complex problem because (1) the diverse characteristics and QoS targets of applications require different optimizations, and (2) per-cluster DVFS requires a global optimization considering all running applications. State-of-the-art resource management for power or temperature minimization either relies on measurements that are commonly not available (such as power) or fails to consider all the dimensions of the optimization (e.g., by using simplified analytical models). To solve this, ML methods can be employed. In particular, IL leverages the optimality of an oracle policy, yet at low run-time overhead, by training a model from oracle demonstrations. We are the first to employ IL for temperature minimization under QoS targets. We tackle the complexity by training NN at design time and accelerate the run-time NN inference using NPU. While such NN accelerators are becoming increasingly widespread, they are so far only used to accelerate user applications. In contrast, we use for the first time an existing accelerator on a real platform to accelerate NN-based resource management. To show the superiority of IL compared to RL in our targeted problem, we also develop multi-agent RL-based management. Our evaluation on a HiKey 970 board with an Arm big.LITTLE CPU and NPU shows that IL achieves significant temperature reductions at a negligible run-time overhead. We compare TOP-IL against several techniques. Compared to ondemand Linux governor, TOP-IL reduces the average temperature by up to 17 ˆC at minimal QoS violations for both techniques. Compared to the RL policy, our TOP-IL achieves 63 % to 89 % fewer QoS violations while resulting similar average temperatures. Moreover, TOP-IL outperforms the RL policy in terms of stability. We additionally show that our IL-based technique also generalizes to different software (unseen applications) and even hardware (different cooling) than used for training. Martin Rapp, Heba Khdr, Nikita Krohmer, Jörg Henkel |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2023 | Smart Detection of Obfuscated Thermal Covert Channel Attacks in Many-core ProcessorsabstractIn thermal covert channel (TCC) attacks, malicious applications seek to leak private information in a stealthy and hard-to-detect manner. State-of-the-art approaches for TCC detection employ the Discrete Fourier Transform (DFT) combined with heuristics to identify possible channels. However, as we demonstrate in this paper, these approaches are limited when detecting short-duration attacks, where an attacker intentionally halts the transmission for a time interval to avoid the detection. In order to overcome this limitation of the state-of-the-art solutions, we propose the first detection method for short-duration TCC attacks. Our solution, Dotecca, is a machine learning-based technique that employs short windows of time-domain measurements instead of the DFT to detect TCCs. To evaluate our solution, we introduce a new obfuscated short-duration attack that disguises as a regular application from the perspective of a DFT spectrum. Our experiments show that the new obfuscated attack is able to remain undetected even under advanced DFT-based state-of-the-art detection approaches, reducing their detection accuracy to about 18 %. In contrast, our smart detection approach is able to detect state-of-the-art and new obfuscated attacks with an accuracy of 99 %. Moreover, our solution reduces the overhead of the DFT-based state-of-the-art solution by more than 14 ×. Jeferson González-Gómez, Mohammed Bakr Sikal, Heba Khdr, Lars Bauer, Jörg Henkel |
DAC | 3 |
| 2023 | Machine Learning-based Thermally-Safe Cache Contention Mitigation in Clustered ManycoresabstractWe present the first technique that mitigates cache contention under thermal constraints in clustered manycores. We show by means of extensive experiments that significant performance gains in this scenario can be achieved. The background is that concurrently-running applications on manycore clusters compete for the shared cache, slowing down their execution. In addition, heavy parallel computations on physically-close cores increase temperatures to non-sustainable levels, which in turn triggers a throttle down of voltage/frequency levels and hence performance is compromised. These problems are not unknown, but as our analysis shows, tackling them independently is sub-optimal. We introduce the first task migration technique that jointly mitigates cache contention while enforcing the thermal constraint at the same time. It works in conjunction with cluster-level dynamic voltage and frequency scaling. Our technique needs to predict the impact of task migration on performance considering cache contention. Since it is impossible to derive an analytical model for cache contention that is both sufficiently accurate and practically feasible to implement, we employ an accurate, yet lightweight neural network (NN) model. As a result, we can operate the manycore system at higher performance while safely staying within thermal constraints. We report a significant step forward in this paper and unveil new potentials for performance optimization. Mohammed Bakr Sikal, Heba Khdr, Martin Rapp, Jörg Henkel |
DAC | 2 |
| 2023 | Extended Abstract: Monitoring-based Thermal Management for Mixed-Criticality SystemsabstractWith a rapidly growing number of functions in embedded real-time systems, it becomes inevitable to integrate tasks of different safety integrity levels (SILs) into one mixed-criticality system. Here, it is important to not only isolate shared architectural resources, as tasks executing on different cores may also interfere via the processor's thermal manager. In order to prevent a scenario where best-effort tasks cause deadline violations for critical tasks, we propose a thermal management strategy that guarantees a sufficient thermal isolation between tasks of different SILs, and simultaneously reduces the run-time of best-effort tasks by up to 45 % compared to the state of the art without incurring any real-time violations for critical tasks. Marcel Mettler, Martin Rapp, Heba Khdr, Daniel Mueller-Gritschneder, Jörg Henkel, Ulf Schlichtmann |
DATE | 3 |
| 2023 | ATLAS: Aging-Aware Task Replication for Multicore Safety-Critical SystemsabstractA major requirement of safety-critical systems is high reliability at low power consumption. Dynamic voltage and frequency (v/f) scaling (DVFS) techniques are widely exploited to reduce power consumption. However, DVFS through downscaling v/f levels has a negative impact on the reliability of the tasks running on the cores, and through upscaling v/f levels has circuitlevel aging effects. To achieve high reliability in multicore safetycritical systems, task replication as a fault-tolerant technique is an established way to deal with the negative effect of downscaling v/f levels, but it may accelerate aging effects due to elevating the on-chip temperatures. In this paper, we propose an aging-aware task replication (called ATLAS) method that solves the problem of satisfying the desired reliability target for a set of periodic hard real-time tasks which are executed on a multicore system. The proposed method satisfies the reliability target of the tasks through updating the required number of replicas for each task at different years. We replicate the tasks through our proposed formulas such that the reliability target is satisfied. However, task replication increases the temperature of the system and accelerates aging. To decelerate aging, we attempt to reduce the temperature while mapping and scheduling the tasks. We have also developed a modified demand bound function (DBF) for our aging-aware task replication method to verify scheduling the realtime tasks. Compared to the existing state-of-the-art techniques, experimental results for safety-critical applications on different configurations of multicore systems demonstrate the efficiency and effectiveness of our proposed method. Experiments show that our proposed method improves schedulability on average by 16.1% and reduces the temperature on average by 7.4°C compared to state-of-the-art methods while meeting the system reliability target. Mohsen Ansari, Sepideh Safari, Amir Yeganeh-Khaksar, Roozbeh Siyadatzadeh, Pourya Gohari-Nazari, Heba Khdr, Muhammad Shafique 0001, Jörg Henkel, Alireza Ejlali |
RTAS | 6 |
| 2022 | NPU-Accelerated Imitation Learning for Thermal- and QoS-Aware Optimization of Heterogeneous Multi-CoresabstractTask migration and dynamic voltage and frequency scaling (DVFS) are indispensable means in thermal optimization of a heterogeneous clustered multi-core processor under user-defined quality of service (QoS) targets. However, selecting the core to execute each application and the voltage/frequency (V/f) levels of each cluster is a complex problem because 1) the diverse characteristics and QoS targets of applications require different optimizations, and 2) V/f levels are often shared between cores on a cluster, which requires a global optimization considering all running applications. State-of-the-art techniques for power or temperature minimization either rely on measurements that are often not available (such as power) or fail to consider all the dimensions of the problem (e.g., by using simplified analytical models). Imitation learning (IL) enables to use the optimality of an oracle policy, yet at low run-time overhead, by training a model from oracle demonstrations. We are the first to employ IL for temperature minimization under QoS targets. We tackle the complexity by using a neural network (NN) model and accelerate the NN inference using a neural processing unit (NPU). While such NN accelerators are becoming increasingly widespread on end devices, they are so far only used to accelerate user applications. In contrast, we use an accelerator on a real platform to accelerate NN-based resource management. Our evaluation on a HiKey970 board with an Arm big.LITTLE CPU and an NPU shows significant temperature reductions at a negligible overhead while satisfying OoS targets. Martin Rapp, Nikita Krohmer, Heba Khdr, Jörg Henkel |
DATE | 3 |
| 2022 | Thermal- and Cache-Aware Resource Management based on ML- Driven Cache Contention PredictionabstractWhile on-chip many-core systems enable a large number of applications to run in parallel, the increased overall performance may come at the cost of complicating the performance constraints of individual applications due to contention on shared resources. For instance, the competition for last-level cache by concurrently-running applications may lead to slowing down the execution and to potentially violating individual performance constraints. Clustered many-cores reduce cache contention at chip level by sharing caches only at cluster level. To reduce cache con-tention within a cluster, state-of-the art techniques aim to co-map a memory-intensive application with a compute-intensive application onto one cluster. However, compute-intensive applications typ-ically consume high power, and therefore, executing another application in their nearby cores may lead to high temperatures. Hence, there is a trade-off between cache contention and temperature. This paper is the first to consider this trade-off through a novel thermal- and cache-aware resource management technique. We build a neural network (NN)-based model to predict the slowdown of the application execution induced by cache contention feeding our resource management technique that then optimizes the application mapping and selects the voltage/frequency levels of the clus-ters to compensate for the potential contention-induced slowdown. Thereby, it meets the performance constraints, while minimizing temperature. Compared to the state of the art, our technique significantly reduces the temperature by 30% on average, while satisfying performance constraints of all individual applications. Mohammed Bakr Sikal, Heba Khdr, Martin Rapp, Jörg Henkel |
DATE | 2 |
| 2022 | An FPGA-based Approach to Evaluate Thermal and Resource Management Strategies of Many-core ProcessorsabstractThe continuous technology scaling of integrated circuits results in increasingly higher power densities and operating temperatures. Hence, modern many-core processors require sophisticated thermal and resource management strategies to mitigate these undesirable side effects. A simulation-based evaluation of these strategies is limited by the accuracy of the underlying processor model and the simulation speed. Therefore, we present, for the first time, an field-programmable gate array (FPGA)-based evaluation approach to test and compare thermal and resource management strategies using the combination of benchmark generation, FPGA-based application-specific integrated circuit (ASIC) emulation, and run-time monitoring. The proposed benchmark generation method enables an evaluation of run-time management strategies for applications with various run-time characteristics. Furthermore, the ASIC emulation platform features a novel distributed temperature emulator design, whose overhead scales linearly with the number of integrated cores, and a novel dynamic voltage frequency scaling emulator design, which precisely models the timing and energy overhead of voltage and frequency transitions. In our evaluations, we demonstrate the proposed approach for a tiled many-core processor with 80 cores on four Virtex-7 FPGAs. Additionally, we present the suitability of the platform to evaluate state-of-the-art run-time management techniques with a case study. Marcel Mettler, Martin Rapp, Heba Khdr, Daniel Mueller-Gritschneder, Jörg Henkel, Ulf Schlichtmann |
ACM Trans. Archit. Code Optim. | 3 |
| 2022 | Power-Aware Checkpointing for Multicore Embedded SystemsabstractIncreasing the number of cores integrated on a single chip offers a great potential for the implementation of fault-tolerant techniques to achieve high reliability in real-time embedded systems. Checkpointing with rollback-recovery is a well-established technique to tolerate transient faults in multicore platforms. To consider the worst-case fault occurrence scenario, checkpointing technique requires to re-execute some parts of the tasks, and that might lead to simultaneous execution of task parts with high power consumptions, which eventually might result in a peak power increase beyond the thermal design power (TDP). Exceeding TDP can elevate on-chip temperatures beyond safe limits, and thereby triggering countermeasures that throttle down the voltage and frequency levels or power gate the cores. Such countermeasures might lead to violating task deadlines and degrading the system's reliability. To avoid such severe scenarios, it is inevitable to consider the impact of applying fault-tolerant techniques on the power consumption and prevent violating the power constraint of the chip, i.e., TDP. This paper presents for the first time, a peak-power-aware checkpointing (PPAC) technique that tolerates a given number of faults,k, while at the same time meets the power constraint in hard real-time embedded systems. To do this, our proposed technique (PPAC) adjusts the timing of the checkpoints, which have lower power consumption than the tasks to the execution time points that have power spikes beyond TDP. Moreover, PPAC exploits the available slack times on the cores to delay the execution of some tasks to avoid the remaining power spikes beyond TDP, which could not be mitigated by solely adjusting checkpoints. To evaluate our technique, we extend the state-of-the-art system-level simulator, gem5, with the state-of-the-art checkpointing module in Linux. Our experimental results show that our proposed technique is able to tolerate a given number of faults without exceeding the timing and power constraints in hard real-time embedded systems. The resulting peak power reduction achieved by our technique compared to state-of-the-art techniques is an average of 23%. Moreover, our technique employs the Dynamic Power Management (DPM) during the slack times resulting at runtime in the case of fault-free scenarios, which provides energy savings with an average of 17.28% and up to 61.1%. Mohsen Ansari, Sepideh Safari, Heba Khdr, Pourya Gohari-Nazari, Jörg Henkel, Alireza Ejlali, Shaahin Hessabi |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | TherMa-MiCs: Thermal-Aware Scheduling for Fault-Tolerant Mixed-Criticality SystemsabstractMulticore platforms are becoming the dominant trend in designing Mixed-Criticality Systems (MCSs), which integrate applications of different levels of criticality into the same platform. A well-known MCS is the dual-criticality system that is composed of low-criticality and high-criticality tasks. The availability of multiple cores on a single chip provides opportunities to employ fault-tolerant techniques, such as N-Modular Redundancy (NMR), to ensure the reliability of MCSs. However, applying fault-tolerant techniques will increase the power consumption on the chip, and thereby on-chip temperatures might increase beyond safe limits. To prevent thermal emergencies, urgent countermeasures, like Dynamic Voltage and Frequency Scaling (DVFS) or Dynamic Power Management (DPM) will be triggered to cool down the chip. Such countermeasures, however, might not only lead to suspending low-criticality tasks, but also it might lead to violating timing constraints of high-criticality tasks. In order to prevent such severe scenarios, it is indispensable to consider a temperature constraint within the scheduling process of fault-tolerant MCSs. Therefore, this paper presents, for the first time, a thermal-aware scheduling scheme for fault-tolerant MCSs, named TherMa-MiCs. In particular, TherMa-MiCs, satisfies the temperature constraint jointly with the timing constraints of the high-criticality tasks, while attempting to maximize the QoS of low-criticality tasks under the predefined constraints. At the same time, a reliability target is satisfied by employing the well-known N-Modular Redundancy (NMR) fault-tolerant technique. Experimental results show that our proposed scheme meets the temperature and timing constraints, while at the same time, improving the QoS of low-criticality tasks, with an average of 44%. Sepideh Safari, Heba Khdr, Pourya Gohari-Nazari, Mohsen Ansari, Shaahin Hessabi, Jörg Henkel |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | SmartBoost: Lightweight ML-Driven Boosting for Thermally-Constrained Many-Core ProcessorsabstractDynamic voltage and frequency scaling (DVFS)-based boosting is indispensable for optimizing the performance of thermally-constrained many-core processors. State-of-the-art techniques employ the voltage/frequency (V/D sensitivity of the performance of an application as a boosting metric. This paper demonstrates that this leads to suboptimal boosting decisions because the sensitivities of power and temperature also play a profound impact and need to be included within the optimization. Therefore, we introduce a novel boosting metric that integrates all relevant metrics: the application-dependent V/f sensitivities of performance and power, and the core-dependent sensitivity of the temperature. This new boosting metric is derived at run-time using machine learning via a neural network (NN) model, which accurately estimates the V/f sensitivities of performance and power of a priori unknown applications with diverse and time-varying characteristics. This new metric enables to build a smart, yet lightweight, boosting technique to maximize the performance under a temperature constraint. The experimental results demonstrate a 21 % average improvement of the system performance over the state-of-the-art at a negligible run-time overhead of 0.8 %. Martin Rapp, Mohammed Bakr Sikal, Heba Khdr, Jörg Henkel |
DAC | 3 |
| 2021 | Long Short-Term Memory Neural Network-based Power Forecasting of Multi-Core ProcessorsabstractWe propose a novel technique to forecast the power consumption of processor cores at run-time. Power consumption varies strongly with different running applications and within their execution phases. Accurately forecasting future power changes is highly relevant for proactive power/thermal management. While forecasting power is straightforward for known or periodic workloads, the challenge for general unknown workloads at different voltage/frequency (v/n-levels is still unsolved. Our technique is based on a long short-term memory (LSTM) recurrent neural network (RNN) to forecast the average power consumption for both the next 1ms and 10ms periods. The runtime inputs for the LSTM RNN are current and past power information as well as performance counter readings. An LSTM RNN enables this forecasting due to its ability to preserve the history of power and performance counters. Our LSTM RNN needs to be trained only once at design-time while adapting during run-time to different system behavior through its internal memory. We demonstrate that our approach accurately forecasts power for unseen applications at different v/f-levels. The experimental results shows that the forecasts of our LSTM RNN provide 43% lower worst case error for the 1ms forecasts and 38% for the 10ms forecasts. comnared to the state of the art. Mark Sagi, Martin Rapp, Heba Khdr, Yizhe Zhang 0005, Nael Fasfous, Nguyen Anh Vu Doan, Thomas Wild, Jörg Henkel, Andreas Herkersdorf |
DATE | 3 |
| 2020 | Combinatorial Auctions for Temperature-Constrained Resource Management in ManycoresabstractAlthough manycore processors have plenty of cores, not all of them may run simultaneously at full speed and even some of them might need to be power-gated in order to keep the chip within safe temperature limits. Hence, a resource management technique, that allocates cores to application aiming at maximizing the system performance, will not be able to achieve its goal without taking into account the on-chip temperature and its impact on the availability of the chip's resources. However, considering a temperature constraint by the resource management will further increase its complexity, especially in manycores, and thus implementing it in a centralized scheme might lead to a computation bottleneck and a single point of failure. To avoid such scenarios, it is inevitable to distribute the computation required by the resource management technique throughout the chip. In this article, we propose a distributed resource management technique that considers temperature as an essential factor in allocating cores to applications and determining the power states of these cores and their voltage/frequency levels, while taking into account the performance models of the applications in order to maximize the overall system performance under a temperature constraint. Our proposed technique employs, for the first time, combinatorial auctions within an agent system to achieve the targeted goal in a distributed manner. The experimental evaluations show that our proposed technique achieves significant performance improvements with an average of 41% compared to several distributed resource management techniques. Heba Khdr, Muhammad Shafique 0001, Santiago Pagani, Andreas Herkersdorf, Jörg Henkel |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | Smart Thermal Management for Heterogeneous MulticoresabstractHeterogeneous multicores have attracted a major focus in recent years, as they provide many possibilities for performance improvements. However, due to the discontinuation of Dennard scaling, on-chip power densities are continuously increasing along with technology scaling, and hence on-chip temperatures are elevated. Therefore, several thermal management techniques have emerged to keep the temperature of the chip within safe limits. These techniques, however, lead to performance losses which ultimately erase a big portion of the expected performance gains from the heterogeneous multicores. Thus, it is indispensable to deploy thermal management techniques that are able to make efficient decisions which satisfy temperature constraints while at the same time maximizing the performance. This paper presents smart thermal management techniques for heterogeneous multicores that exploit relevant information about several heterogeneity parameters at the chip level and at the application level to increase thermal efficiency1. Compared to the state of the art, the presented techniques are able to obtain significant performance improvements under the same thermal constraint.This paper is part of the DATE 2019 special session on "Smart Resource Management and Design Space Exploration for Heterogeneous Processors". The other two papers of this special session are [1] and [2]. Jörg Henkel, Heba Khdr, Martin Rapp |
DATE | 2 |
| 2019 | Thermally Composable Hybrid Application Mapping for Real-Time Applications in Heterogeneous Many-Core SystemsabstractModern embedded many-core systems host, among others, real-time applications which must be dynamically launched at run time. To this end, Hybrid Application Mapping (HAM) methodologies combine design-time analysis with runtime mapping techniques to enable dynamic application mapping with performance guarantees, e.g., w.r.t. real-time constraints. They rely on composability to derive the required performance guarantees in an isolated analysis of individual applications at design time. The ongoing process technology downsizing, however, has given rise to an increased on-chip temperature, so that the thermal integrity of the platform must be monitored and enforced at run time by means of Dynamic Thermal Management (DTM) techniques which use countermeasures e.g. DVFS and power gating. This, however, violates composability, as the thermally unsafe behavior of one application may trigger DTM countermeasures that affect other applications running in the thermally affected region which, in turn, may lead to the violation of their real-time constraints. As a remedy, this paper proposes, for the first time, a thermally composable HAM methodology that enforces thermal safety proactively at the launch time of applications and, thereby, prevents DTM interferences which react to thermal violations. To that end, we present (a) a novel thermal-safety analysis that can be integrated into the design-time analysis of HAM and (b) a set of thermal-safety admission checks that can be used at run time when launching an application. By establishing thermal composability among running applications, the proposed HAM approach enables providing thermally safe real-time guarantees for dynamically mapped applications in many-core systems. Experimental results for a variety of hard real-time applications on multiple heterogeneous many-core architectures demonstrate the efficiency and effectiveness of the proposed methodology. Behnaz Pourmohseni, Fedor Smirnov, Heba Khdr, Stefan Wildermann, Jürgen Teich, Jörg Henkel |
RTSS | 3 |
| 2019 | Dynamic Guardband Selection: Thermal-Aware Optimization for Unreliable Multi-Core SystemsabstractCircuit aging has become the major reliability concern in current and upcoming technology nodes. For instance, Bias Temperature Instability (BTI) leads to an increase in the threshold voltage of a transistor. That, in turn, may prolong the critical path delay of the processor and eventually may lead to timing errors. In order to avoid aging-induced timing errors, designers employ guardbands either with respect to voltage or frequency. State-of-the-art techniques determine a guardband type at the circuit level at design time irrespective from the running workload at the system level. Our investigation revealed that generated temperatures by a running workload have the potential to play a key role in determining the appropriate guardband type with respect to system performance. Therefore, we propose a paradigm shift in designing guardbands: to select the guardband types on-the-fly with respect to the workload-induced temperatures aiming at optimizing for performance under temperature and reliability constraints. Moreover, different guardband types for different cores can be selected simultaneously when multiple applications with diverse properties suggest this to be useful. Our dynamic guardband selection allows for a higher performance compared to techniques that employ a fixed (at design time) guardband type throughout. Heba Khdr, Hussam Amrouch, Jörg Henkel |
IEEE Trans. Computers | 1 |
| 2018 | Aging-constrained performance optimization for multi coresabstractCircuit aging has become a dire design concern and hence it is considered a primary design constraint. Current practice to cope with this problem is to apply (too) conservative means. Heba Khdr, Hussam Amrouch, Jörg Henkel |
DAC | 1 |
| 2018 | QoS-aware stochastic power management for many-coresabstractA many-core processor can execute hundreds of multi-threaded tasks in parallel on its 100s - 1000s of processing cores. When deployed in a Quality of Service (QoS)-based system, the many-core must execute a task at a target QoS. The amount of processing required by the task for the QoS varies over the task's lifetime. Accordingly, Dynamic Voltage and Frequency Scaling (DVFS) allows the many-core to deliver precise amount of processing required to meet the task QoS guarantee while conserving power. Still, a global control is necessitated to ensure that the many-core overall does not exceed its power budget. Anuj Pathania, Heba Khdr, Muhammad Shafique 0001, Tulika Mitra, Jörg Henkel |
DAC | 2 |
| 2018 | Aging-Aware BoostingabstractDVFS-based boosting techniques have been widely employed by commercial multi-core processors, due to their superiority in improving the performance. Boosting, however, is particularly stressing circuits and hence it significantly contributes to an accelerated aging process. Circuit aging has become a real reliability concern because it leads to an increase in transistor threshold voltage that may cause timing errors as a result of higher delays in critical paths. Thus, high performance is desirable but it shortens the circuit lifetime through aging leaving a choice to trade-off. Besides well-known long-term aging effects, recent research also reported short-term aging effects. Our claim is that DVFS-based boosting techniques should consider both long- and short-term aging effects. This can be circumvented by wider timing guardbands. But that would be more expensive. The goal of this work is therefore to analyze and optimize boosting under specific consideration of long-term and short-term aging effects. As a result of our findings, we propose the first comprehensive aging-aware, yet efficient boosting technique. The employed aging-aware cell libraries in this work are publicly available at http://ces.itec.kit.edu/dependable-hardware.php. Heba Khdr, Hussam Amrouch, Jörg Henkel |
IEEE Trans. Computers | 1 |
| 2017 | Scalable probabilistic power budgeting for many-coresabstractMany-core processors exhibit hundreds to thousands of cores, which can execute lots of multi-threaded tasks in parallel. Restrictive power dissipation capacity of a many-core prevents all its executing tasks from operating at their peak performance together. Furthermore, the ability of a task to exploit part of the power budget allocated to it depends upon its current execution phase. This mandates careful rationing of the power budget amongst the tasks for full exploitation of the many-core. Past research proposed power budgeting techniques that redistribute power budget amongst tasks based on up-to-date information about their current phases. This phase information needs to be constantly propagated throughout the system and processed, inhibiting scalability. In this work, we propose a novel probabilistic technique for power budgeting which requires no exchange of phase information yet provides mathematical guarantees on judicial use of the TDP. The proposed probabilistic technique reduces the power budgeting overheads by 97.13% in comparison to a non-probabilistic approach, while providing almost equal performance on simulated thousand-core system. Anuj Pathania, Heba Khdr, Muhammad Shafique 0001, Tulika Mitra, Jörg Henkel |
DATE | 2 |
| 2017 | Power Density-Aware Resource Management for Heterogeneous Tiled MulticoresabstractIncreasing power densities have led to the dark silicon era, for which heterogeneous multicores with different power and performance characteristics are promising architectures. This paper focuses on maximizing the overall system performance under a critical temperature constraint for heterogeneous tiled multicores, where all cores or accelerators inside a tile share the same voltage and frequency levels. For such architectures, we present a resource management technique that introduces power density as a novel system level constraint, in orderto avoid thermal violations. The proposed technique then assigns applications to tiles by choosing their degree of parallelism and the voltage/frequency levels of each tile, such that the power density constraint is satisfied. Moreover, our technique provides runtime adaptation of the power density constraint according to the characteristics of the executed applications, and reacting to workload changes at runtime. Thus, the available thermal headroom is exploited to maximize the overall system performance. Heba Khdr, Santiago Pagani, Éricles Sousa, Vahid Lari, Anuj Pathania, Frank Hannig, Muhammad Shafique 0001, Jürgen Teich, Jörg Henkel |
IEEE Trans. Computers | 1 |
| 2017 | Thermal Safe Power (TSP): Efficient Power Budgeting for Heterogeneous Manycore Systems in Dark SiliconabstractChip manufacturers provide the Thermal Design Power (TDP) for a specific chip. The cooling solution is designed to dissipate this power level. But because TDP is not necessarily the maximum power that can be applied, chips are operated with Dynamic Thermal Management (DTM) techniques. To avoid excessive triggers of DTM, usually, system designers also use TDP as power constraint. However, using a single and constant value as power constraint, e.g., TDP, can result in significant performance losses in homogeneous and heterogeneous manycore systems. Having better power budgeting techniques is a major step towards dealing with the dark silicon problem. This paper presents a new power budget concept, called Thermal Safe Power (TSP), which is an abstraction that provides safe power and power density constraints as a function of the number of simultaneously active cores. Executing cores at any power consumption below TSP ensures that DTM is not triggered. TSP can be computed offline for the worst cases, or online for a particular mapping of cores. TSP can also serve as a fundamental tool for guiding task partitioning and core mapping decisions, specially when core heterogeneity or timing guarantees are involved. Moreover, TSP results in dark silicon estimations which are less pessimistic than estimations using constant power budgets. Santiago Pagani, Heba Khdr, Jian-Jia Chen, Muhammad Shafique 0001, Minming Li, Jörg Henkel |
IEEE Trans. Computers | 2 |
| 2016 | Towards performance and reliability-efficient computing in the dark silicon era
Jörg Henkel, Santiago Pagani, Heba Khdr, Florian Kriebel, Semeen Rehman, Muhammad Shafique 0001 |
DATE | 3 |
| 2015 | New trends in dark siliconabstractThis paper presents new trends in dark silicon reflecting, among others, the deployment of FinFETs in recent technology nodes and the impact of voltage/frquency scaling, which lead to new less-conservative predictions. The focus is on dark silicon from a thermal perspective: we show that it is not simply the chip's total power budget, e.g., the Thermal Design Power (TDP), that leads to the dark silicon problem, but instead it is the power density and related thermal effects. We therefore propose to use Thermal Safe Power (TSP) as a more efficient power budget. It is also shown that sophisticated spatio-temporal mapping decisions result in improved thermal profiles with reduced peak temperatures. Moreover, we discuss the implications of Near-Threshold Computing (NTC) and employment of Boosting techniques in dark silicon systems. Jörg Henkel, Heba Khdr, Santiago Pagani, Muhammad Shafique 0001 |
DAC | 2 |
| 2015 | Thermal constrained resource management for mixed ILP-TLP workloads in dark silicon chipsabstractIn dark silicon chips, a significant amount of on-chip resources cannot be simultaneously powered on and need to stay dark, i.e., power gated, in order to avoid thermal emergencies. This paper presents a resource management technique, called DsRem, that selects the number of active cores jointly with their voltage/frequency (v/f) levels, considering the high Instruction Level Parallelism (ILP) or Thread Level Parallelism (TLP) nature of different applications, in order to maximize the overall system performance. DsRem leverages the positioning of dark cores, to efficiently dissipate the heat generated by the active cores. This facilitates increasing the v/f level of the active cores, which leads to further performance improvement. Compared to state-of-the-art thermal-aware task application mapping, DsRem achieves up to 46% performance gain, while avoiding any thermal emergencies. Additionally, DsRem outperforms the boosting technique with 26%. Heba Khdr, Santiago Pagani, Muhammad Shafique 0001, Jörg Henkel |
DAC | 1 |
| 2015 | Dark Silicon: From Computation to CommunicationabstractIn the emerging Dark Silicon era, not all parts of an on-chip system (i.e., cores, Network-on-Chip, and memory resources) can be simultaneously powered-on at the full speed. This paper aims at exposing dark silicon challenges to the NOCS community with an overview of some of the early research efforts that are attempting to shape the design and run-time management of future generation heterogeneous dark silicon processors. The goal is to cover both the computation and communication perspectives. In particular, we exploit computation and communication heterogeneity at multiple levels of system abstractions to design and manage dark silicon processors. The available dark silicon is leveraged to improve power/energy, performance, and reliability efficiency. Jörg Henkel, Haseeb Bokhari, Siddharth Garg, Muhammad Usman Karim Khan, Heba Khdr, Florian Kriebel, Ümit Y. Ogras, Sri Parameswaran, Muhammad Shafique 0001 |
NOCS | 5 |
| 2014 | mDTM: Multi-objective dynamic thermal management for on-chip systemsabstractThermal hot spots and unbalanced temperatures between cores on chip can cause either degradation in performance or may have a severe impact on reliability, or both. In this paper, we propose mDTM, a proactive dynamic thermal management technique for on-chip systems. It employs multi-objective management for migrating tasks in order to both prevent the system from hitting an undesirable thermal threshold and to balance the temperatures between the cores. Our evaluation on the Intel SCC platform shows that mDTM can successfully avoid a given thermal threshold and reduce spatial thermal variation by 22%. Compared to state-of-the-art, our mDTM achieves up to 58% performance gain. Additionally, we deploy an FPGA and IR camera based setup to analyze the effectiveness of our technique. Heba Khdr, Thomas Ebi, Muhammad Shafique 0001, Hussam Amrouch, Jörg Henkel |
DATE | 1 |
| 2014 | Peak Power Management for scheduling real-time tasks on heterogeneous many-core systemsabstractThe number and diversity of cores in on-chip systems is increasing rapidly. However, due to the Thermal Design Power (TDP) constraint, it is not possible to continuously operate all cores at the same time. Exceeding the TDP constraint may activate the Dynamic Thermal Management (DTM) to ensure thermal stability. Such hardware based closed-loop safeguards pose a big challenge in using many-core chips for real-time tasks. Managing the worst-case peak power usage of a chip can help toward resolving this issue. We present a scheme to minimize the peak power usage for frame-based and periodic real-time tasks on many-core processors by scheduling the sleep cycles for each active core and introduce the concept of a sufficient test for peak power consumption for task feasibility. We consider both inter-task and inter-core diversity in terms of power usage and present computationally efficient algorithms for peak power minimization for these cases, i.e., a special case of “homogeneous tasks on homogeneous cores” to the general case of “heterogeneous tasks on heterogeneous cores”. We evaluate our solution through extensive simulations using the 48-core SCC platform and gem5 architecture simulator. Our simulation results show the efficacy of our scheme. Waqaas Munawar, Heba Khdr, Santiago Pagani, Muhammad Shafique 0001, Jian-Jia Chen, Jörg Henkel |
ICPADS | 2 |
| 2013 | Thermal management for dependable on-chip systemsabstractDependability has become a growing concern in the nano-CMOS era due to elevated temperatures and an increased susceptibility to temperature of the small structures. We present an overview of temperature-related effects that threaten dependability and a methodology for reducing the dependability concerns through thermal management utilizing the concept of aging budgeting. Jörg Henkel, Thomas Ebi, Hussam Amrouch, Heba Khdr |
ASP-DAC | 4 |