EDBT 2026 Demo / reviewers in the wild / expert
Martin Rapp
dblp:223/9604
· DBLP profile ↗
25ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0002-5989-2950ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 9 first-author · 13 since 2021Software engineering, systems software and programming languages · 8 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIQARA: Mixed-Criticality Queue-based Architecture for Reconfigurable Accelerator PlatformsabstractCoexistence of safety-critical control functions and besteffort computations in mixed-criticality systems poses a challenge in resource allocation and scheduling, as high-criticality jobs must adhere to strict timing guarantees, while lower-criticality jobs should make effective use of available resources without compromising the system’s safety and predictability. This paper introduces MIQARA1, a mixed-criticality queue-based architecture designed for reconfigurable accelerator platforms. MIQARA efficiently combines software-programmable CPUs with reconfigurable hardware, utilizing a dynamic job pipeline, token-based dependency tracking, and out-of-order scheduling to optimize resource utilization. At the same time, MIQARA has been designed to satisfy real-time constraints. MIQARA is evaluated on four FPGA platforms: the Zed Board, DipForty board, ZCU102 board, all of which have ARM CPUs implemented on chip, and Arty A7 with a RISC-V soft-core processor, representing systems that rely on soft CPUs. Results demonstrate substantial performance gains, particularly in terms of execution speed, flexibility, and adaptability to mixed-criticality workloads. The integration of features such as a streaming network further illustrates MIQARA’s scalability to complex data-intensive applications, making it a compelling solution for embedded mixed-criticality systems. MIQARA requires a hardware overhead of 17.8% and achieves a speedup of up to 4×. Hassan Nassar, Martin Rapp, Lars Bauer, Mostafa Elshimy, Zeynep Demirdag, Jörg Henkel |
DATE | 2 |
| 2024 | NPU-Accelerated Imitation Learning for Thermal Optimization of QoS-Constrained Heterogeneous Multi-CoresabstractThermal optimization of a heterogeneous clustered multi-core processor under user-defined QoS targets requires application migration and DVFS. However, selecting the core to execute each application and the VF levels of each cluster is a complex problem because (1) the diverse characteristics and QoS targets of applications require different optimizations, and (2) per-cluster DVFS requires a global optimization considering all running applications. State-of-the-art resource management for power or temperature minimization either relies on measurements that are commonly not available (such as power) or fails to consider all the dimensions of the optimization (e.g., by using simplified analytical models). To solve this, ML methods can be employed. In particular, IL leverages the optimality of an oracle policy, yet at low run-time overhead, by training a model from oracle demonstrations. We are the first to employ IL for temperature minimization under QoS targets. We tackle the complexity by training NN at design time and accelerate the run-time NN inference using NPU. While such NN accelerators are becoming increasingly widespread, they are so far only used to accelerate user applications. In contrast, we use for the first time an existing accelerator on a real platform to accelerate NN-based resource management. To show the superiority of IL compared to RL in our targeted problem, we also develop multi-agent RL-based management. Our evaluation on a HiKey 970 board with an Arm big.LITTLE CPU and NPU shows that IL achieves significant temperature reductions at a negligible run-time overhead. We compare TOP-IL against several techniques. Compared to ondemand Linux governor, TOP-IL reduces the average temperature by up to 17 ˆC at minimal QoS violations for both techniques. Compared to the RL policy, our TOP-IL achieves 63 % to 89 % fewer QoS violations while resulting similar average temperatures. Moreover, TOP-IL outperforms the RL policy in terms of stability. We additionally show that our IL-based technique also generalizes to different software (unseen applications) and even hardware (different cooling) than used for training. Martin Rapp, Heba Khdr, Nikita Krohmer, Jörg Henkel |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2023 | Machine Learning-based Thermally-Safe Cache Contention Mitigation in Clustered ManycoresabstractWe present the first technique that mitigates cache contention under thermal constraints in clustered manycores. We show by means of extensive experiments that significant performance gains in this scenario can be achieved. The background is that concurrently-running applications on manycore clusters compete for the shared cache, slowing down their execution. In addition, heavy parallel computations on physically-close cores increase temperatures to non-sustainable levels, which in turn triggers a throttle down of voltage/frequency levels and hence performance is compromised. These problems are not unknown, but as our analysis shows, tackling them independently is sub-optimal. We introduce the first task migration technique that jointly mitigates cache contention while enforcing the thermal constraint at the same time. It works in conjunction with cluster-level dynamic voltage and frequency scaling. Our technique needs to predict the impact of task migration on performance considering cache contention. Since it is impossible to derive an analytical model for cache contention that is both sufficiently accurate and practically feasible to implement, we employ an accurate, yet lightweight neural network (NN) model. As a result, we can operate the manycore system at higher performance while safely staying within thermal constraints. We report a significant step forward in this paper and unveil new potentials for performance optimization. Mohammed Bakr Sikal, Heba Khdr, Martin Rapp, Jörg Henkel |
DAC | 3 |
| 2023 | Extended Abstract: Monitoring-based Thermal Management for Mixed-Criticality SystemsabstractWith a rapidly growing number of functions in embedded real-time systems, it becomes inevitable to integrate tasks of different safety integrity levels (SILs) into one mixed-criticality system. Here, it is important to not only isolate shared architectural resources, as tasks executing on different cores may also interfere via the processor's thermal manager. In order to prevent a scenario where best-effort tasks cause deadline violations for critical tasks, we propose a thermal management strategy that guarantees a sufficient thermal isolation between tasks of different SILs, and simultaneously reduces the run-time of best-effort tasks by up to 45 % compared to the state of the art without incurring any real-time violations for critical tasks. Marcel Mettler, Martin Rapp, Heba Khdr, Daniel Mueller-Gritschneder, Jörg Henkel, Ulf Schlichtmann |
DATE | 2 |
| 2023 | Speed-Oblivious Online Scheduling: Knowing (Precise) Speeds is not NecessaryabstractWe consider online scheduling on unrelated (heterogeneous) machines in a speed-oblivious setting, where an algorithm is unaware of the exact job-dependent processing speeds. We show strong impossibility results for clairvoyant and non-clairvoyant algorithms and overcome them in models inspired by practical settings: (i) we provide competitive learning-augmented algorithms, assuming that (possibly erroneous) predictions on the speeds are given, and (ii) we provide competitive algorithms for the speed-ordered model, where a single global order of machines according to their unknown job-dependent speeds is known. We prove strong theoretical guarantees and evaluate our findings on a representative heterogeneous multi-core processor. These seem to be the first empirical results for scheduling algorithms with predictions that are evaluated in a non-synthetic hardware environment. Alexander Lindermayr, Nicole Megow, Martin Rapp |
ICML | 3 |
| 2022 | DISTREAL: Distributed Resource-Aware Learning in Heterogeneous SystemsabstractWe study the problem of distributed training of neural networks (NNs) on devices with heterogeneous, limited, and time-varying availability of computational resources. We present an adaptive, resource-aware, on-device learning mechanism, DISTREAL, which is able to fully and efficiently utilize the available resources on devices in a distributed manner, increasing the convergence speed. This is achieved with a dropout mechanism that dynamically adjusts the computational complexity of training an NN by randomly dropping filters of convolutional layers of the model. Our main contribution is the introduction of a design space exploration (DSE) technique, which finds Pareto-optimal per-layer dropout vectors with respect to resource requirements and convergence speed of the training. Applying this technique, each device is able to dynamically select the dropout vector that fits its available resource without requiring any assistance from the server. We implement our solution in a federated learning (FL) system, where the availability of computational resources varies both between devices and over time, and show through extensive evaluation that we are able to significantly increase the convergence speed over the state of the art without compromising on the final accuracy. Martin Rapp, Ramin Khalili, Kilian Pfeiffer, Jörg Henkel |
AAAI | 1 |
| 2022 | NPU-Accelerated Imitation Learning for Thermal- and QoS-Aware Optimization of Heterogeneous Multi-CoresabstractTask migration and dynamic voltage and frequency scaling (DVFS) are indispensable means in thermal optimization of a heterogeneous clustered multi-core processor under user-defined quality of service (QoS) targets. However, selecting the core to execute each application and the voltage/frequency (V/f) levels of each cluster is a complex problem because 1) the diverse characteristics and QoS targets of applications require different optimizations, and 2) V/f levels are often shared between cores on a cluster, which requires a global optimization considering all running applications. State-of-the-art techniques for power or temperature minimization either rely on measurements that are often not available (such as power) or fail to consider all the dimensions of the problem (e.g., by using simplified analytical models). Imitation learning (IL) enables to use the optimality of an oracle policy, yet at low run-time overhead, by training a model from oracle demonstrations. We are the first to employ IL for temperature minimization under QoS targets. We tackle the complexity by using a neural network (NN) model and accelerate the NN inference using a neural processing unit (NPU). While such NN accelerators are becoming increasingly widespread on end devices, they are so far only used to accelerate user applications. In contrast, we use an accelerator on a real platform to accelerate NN-based resource management. Our evaluation on a HiKey970 board with an Arm big.LITTLE CPU and an NPU shows significant temperature reductions at a negligible overhead while satisfying OoS targets. Martin Rapp, Nikita Krohmer, Heba Khdr, Jörg Henkel |
DATE | 1 |
| 2022 | Thermal- and Cache-Aware Resource Management based on ML- Driven Cache Contention PredictionabstractWhile on-chip many-core systems enable a large number of applications to run in parallel, the increased overall performance may come at the cost of complicating the performance constraints of individual applications due to contention on shared resources. For instance, the competition for last-level cache by concurrently-running applications may lead to slowing down the execution and to potentially violating individual performance constraints. Clustered many-cores reduce cache contention at chip level by sharing caches only at cluster level. To reduce cache con-tention within a cluster, state-of-the art techniques aim to co-map a memory-intensive application with a compute-intensive application onto one cluster. However, compute-intensive applications typ-ically consume high power, and therefore, executing another application in their nearby cores may lead to high temperatures. Hence, there is a trade-off between cache contention and temperature. This paper is the first to consider this trade-off through a novel thermal- and cache-aware resource management technique. We build a neural network (NN)-based model to predict the slowdown of the application execution induced by cache contention feeding our resource management technique that then optimizes the application mapping and selects the voltage/frequency levels of the clus-ters to compensate for the potential contention-induced slowdown. Thereby, it meets the performance constraints, while minimizing temperature. Compared to the state of the art, our technique significantly reduces the temperature by 30% on average, while satisfying performance constraints of all individual applications. Mohammed Bakr Sikal, Heba Khdr, Martin Rapp, Jörg Henkel |
DATE | 3 |
| 2022 | An FPGA-based Approach to Evaluate Thermal and Resource Management Strategies of Many-core ProcessorsabstractThe continuous technology scaling of integrated circuits results in increasingly higher power densities and operating temperatures. Hence, modern many-core processors require sophisticated thermal and resource management strategies to mitigate these undesirable side effects. A simulation-based evaluation of these strategies is limited by the accuracy of the underlying processor model and the simulation speed. Therefore, we present, for the first time, an field-programmable gate array (FPGA)-based evaluation approach to test and compare thermal and resource management strategies using the combination of benchmark generation, FPGA-based application-specific integrated circuit (ASIC) emulation, and run-time monitoring. The proposed benchmark generation method enables an evaluation of run-time management strategies for applications with various run-time characteristics. Furthermore, the ASIC emulation platform features a novel distributed temperature emulator design, whose overhead scales linearly with the number of integrated cores, and a novel dynamic voltage frequency scaling emulator design, which precisely models the timing and energy overhead of voltage and frequency transitions. In our evaluations, we demonstrate the proposed approach for a tiled many-core processor with 80 cores on four Virtex-7 FPGAs. Additionally, we present the suitability of the platform to evaluate state-of-the-art run-time management techniques with a case study. Marcel Mettler, Martin Rapp, Heba Khdr, Daniel Mueller-Gritschneder, Jörg Henkel, Ulf Schlichtmann |
ACM Trans. Archit. Code Optim. | 2 |
| 2022 | CoMeT: An Integrated Interval Thermal Simulation Toolchain for 2D, 2.5D, and 3D Processor-Memory SystemsabstractProcessing cores and the accompanying main memory working in tandem enable modern processors. Dissipating heat produced from computation remains a significant problem for processors. Therefore, the thermal management of processors continues to be an active subject of research. Most thermal management research is performed using simulations, given the challenges in measuring temperatures in real processors. Fast yet accurate interval thermal simulation toolchains remain the research tool of choice to study thermal management in processors at the system level. However, the existing toolchains focus on the thermal management of cores in the processors, since they exhibit much higher power densities than memory. The memory bandwidth limitations associated with 2D processors lead to high-density 2.5D and 3D packaging technology: 2.5D packaging technology places cores and memory on the same package; 3D packaging technology takes it further by stacking layers of memory on the top of cores themselves. These new packagings significantly increase the power density of the processors, making them prone to overheating. Therefore, mitigating thermal issues in high-density processors (packaged with stacked memory) becomes even more pressing. However, given the lack of thermal modeling for memories in existing interval thermal simulation toolchains, they are unsuitable for studying thermal management for high-density processors. To address this issue, we present the first integrated Core and Memory interval Thermal (CoMeT) simulation toolchain.CoMeTcomprehensively supports thermal simulation of high- and low-density processors corresponding to four different core-memory (integration) configurations—off-chip DDR memory, off-chip 3D memory, 2.5D, and 3D.CoMeTsupports several novel features that facilitate overlying system research.CoMeTadds only an additional ~5% simulation-time overhead compared to an equivalent state-of-the-art core-only toolchain. The source code ofCoMeThas been made open for public use under theMITlicense. Lokesh Siddhu, Rajesh Kedia, Shailja Pandey, Martin Rapp, Anuj Pathania, Jörg Henkel, Preeti Ranjan Panda |
ACM Trans. Archit. Code Optim. | 4 |
| 2022 | MLCAD: A Survey of Research in Machine Learning for CAD Keynote PaperabstractDue to the increasing size of integrated circuits (ICs), their design and optimization phases (i.e., computer-aided design, CAD) grow increasingly complex. At design time, a large design space needs to be explored to find an implementation that fulfills all specifications and then optimizes metrics like energy, area, delay, reliability, etc. At run time, a large configuration space needs to be searched to find the best set of parameters (e.g., voltage/frequency) to further optimize the system. Both spaces are infeasible for exhaustive search typically leading to heuristic optimization algorithms that find some tradeoff between design quality and computational overhead. Machine learning (ML) can build powerful models that have successfully been employed in related domains. In this survey, we categorize how ML may be used and is used for design-time and run-time optimization and exploration strategies of ICs. A metastudy of published techniques unveils areas in CAD that are well explored and underexplored with ML, as well as trends in the employed ML algorithms. We present a comprehensive categorization and summary of the state of the art on ML for CAD. Finally, we summarize the remaining challenges and promising open research directions. Martin Rapp, Hussam Amrouch, Yibo Lin, Bei Yu 0001, David Z. Pan, Marilyn Wolf, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | SmartBoost: Lightweight ML-Driven Boosting for Thermally-Constrained Many-Core ProcessorsabstractDynamic voltage and frequency scaling (DVFS)-based boosting is indispensable for optimizing the performance of thermally-constrained many-core processors. State-of-the-art techniques employ the voltage/frequency (V/D sensitivity of the performance of an application as a boosting metric. This paper demonstrates that this leads to suboptimal boosting decisions because the sensitivities of power and temperature also play a profound impact and need to be included within the optimization. Therefore, we introduce a novel boosting metric that integrates all relevant metrics: the application-dependent V/f sensitivities of performance and power, and the core-dependent sensitivity of the temperature. This new boosting metric is derived at run-time using machine learning via a neural network (NN) model, which accurately estimates the V/f sensitivities of performance and power of a priori unknown applications with diverse and time-varying characteristics. This new metric enables to build a smart, yet lightweight, boosting technique to maximize the performance under a temperature constraint. The experimental results demonstrate a 21 % average improvement of the system performance over the state-of-the-art at a negligible run-time overhead of 0.8 %. Martin Rapp, Mohammed Bakr Sikal, Heba Khdr, Jörg Henkel |
DAC | 1 |
| 2021 | Long Short-Term Memory Neural Network-based Power Forecasting of Multi-Core ProcessorsabstractWe propose a novel technique to forecast the power consumption of processor cores at run-time. Power consumption varies strongly with different running applications and within their execution phases. Accurately forecasting future power changes is highly relevant for proactive power/thermal management. While forecasting power is straightforward for known or periodic workloads, the challenge for general unknown workloads at different voltage/frequency (v/n-levels is still unsolved. Our technique is based on a long short-term memory (LSTM) recurrent neural network (RNN) to forecast the average power consumption for both the next 1ms and 10ms periods. The runtime inputs for the LSTM RNN are current and past power information as well as performance counter readings. An LSTM RNN enables this forecasting due to its ability to preserve the history of power and performance counters. Our LSTM RNN needs to be trained only once at design-time while adapting during run-time to different system behavior through its internal memory. We demonstrate that our approach accurately forecasts power for unseen applications at different v/f-levels. The experimental results shows that the forecasts of our LSTM RNN provide 43% lower worst case error for the 1ms forecasts and 38% for the 10ms forecasts. comnared to the state of the art. Mark Sagi, Martin Rapp, Heba Khdr, Yizhe Zhang 0005, Nael Fasfous, Nguyen Anh Vu Doan, Thomas Wild, Jörg Henkel, Andreas Herkersdorf |
DATE | 2 |
| 2021 | Neural Network-Based Performance Prediction for Task Migration on S-NUCA Many-CoresabstractThe performance of a task running on a many-core with distributed shared last-level cache (LLC) strongly depends on two parameters: the power budget needed to guarantee thermally-safe operation and the LLC latency. The task's thread-to-core mapping determines both the parameters and needs to make a trade-off because both cannot be simultaneously optimal. Arrival and departure of tasks on a many-core deployed in an open system can change its state significantly in terms of available cores and power budgets. Task migrations can thereupon be used as a tool to keep the many-core operating at peak performance. Furthermore, the relative impacts of power budget and LLC latency on a task's performance may change with its different execution phases mandating its migration on-the-fly. We propose the first run-time algorithmPCMigthat increases the performance of a many-core with distributed shared LLC by migrating tasks based on their phases and the many-core's state.PCMigis based on a model that predicts the performance impact of migrations. We propose a performance prediction model based on a lightweight neural network (NN). To serve as a reference, we also propose an analytical model of the many-core that operates on CPI stacks. We demonstrate an NN-based model achieves a higher prediction accuracy at a lower overhead than an analytical model.PCMigis based on the NN prediction model and results in an up to 7.3 percent increase in performance under a thermal constraint for mixed workloads compared to architecture-aware state-of-the-art (up to 20 percent increase for individual applications). This is achieved with a run-time overhead of less than 0.5 percent. Martin Rapp, Anuj Pathania, Tulika Mitra, Jörg Henkel |
IEEE Trans. Computers | 1 |
| 2021 | Power-Efficient Heterogeneous Many-Core Design With NCFET TechnologyabstractMulti-/many-core, homogeneous or heterogeneous architectures, using the existing CMOS technology are inevitably approaching the limit of attainable power efficiency due to the fundamental limits in scaling. Negative Capacitance Field-Effect Transistor (NCFET) is rapidly emerging as an alternative technology that promises a multi-fold increase in the power efficiency of transistors, yet is compatible with the existing CMOS fabrication process. NCFET incorporates a ferroelectric (FE) layer within the transistor's gate stack, which exhibits a negative capacitance effect amplifying the internal voltage. NCFET has been in detail studied in both physics and devices/circuits communities where its superiority has been demonstrated in semiconductor measurements. However, the full promise of NCFET remains unmodeled and unquantified unless the research is further continued to the microarchitecture and system levels. This article, for the first time, explores system- and application-level benefits of NCFET-based multi-/many-core designs in terms of performance and power-efficiency compared to state-of-the-art FinFET-based designs. This exploration is done first through analytical modeling in which we extend Amdahl's law for NCFET multi-/many-cores, and then through quantitative modeling. The latter is achieved through RTL- and system-level simulations of NCFET-based multi-cores. The analytical modeling shows that a novel type of technology-based heterogeneity in which cores with the same microarchitecture but different FE thickness are combined is highly beneficial. Our exploration shows that this novel heterogeneity increases the power-efficiency by up to 3.5× over homogeneous systems and even achieves 8.3% better performance and 20% higher power-efficiency than conventional heterogeneity in the microarchitecture without having to cope with the complexity of managing different microarchitectures. Sami Salamin, Martin Rapp, Anuj Pathania, Arka Maity, Jörg Henkel, Tulika Mitra, Hussam Amrouch |
IEEE Trans. Computers | 2 |
| 2020 | Energy Optimization in NCFET-based ProcessorsabstractEnergy consumption is a key optimization goal for all modern processors. Negative Capacitance Field-Effect Transistors (NCFETs) are a leading emerging technology that promises outstanding performance in addition to better energy efficiency. Thickness of the additional ferroelectric layer, frequency, and voltage are the key parameters in NCFET technology that impact the power and frequency of processors. However, their joint impact on energy optimization has not been investigated yet.In this work, we are the first to demonstrate that conventional (i.e., NCFET-unaware) dynamic voltage/frequency scaling (DVFS) techniques to minimize energy are sub-optimal when applied to NCFET-based processors. We further demonstrate that state-of-the-art NCFET-aware voltage scaling for power minimization is also sub-optimal when it comes to energy. This work provides the first NCFET-aware DVFS technique that optimizes the processor's energy through optimal runtime frequency/voltage selection. In NCFETs, energy-optimal frequency and voltage are dependent on the workload and technology parameters. Our NCFET-aware DVFS technique considers these effects to perform optimal voltage/frequency selection at runtime depending on workload characteristics. Results show up to 90 % energy savings compared to conventional DVFS techniques. Compared to state-of-the-art NCFET-aware power management, our technique provides up to 72 % energy savings along with 3.7x higher performance. Sami Salamin, Martin Rapp, Hussam Amrouch, Andreas Gerstlauer, Jörg Henkel |
DATE | 2 |
| 2020 | Power- and Cache-Aware Task Mapping with Dynamic Power Budgeting for Many-CoresabstractTwo factors primarily affect the performance of multi-threaded tasks on many-core processors with logically-shared and physically-distributed Last-Level Cache (LLC): the LLC latencies of threads running on different cores and the per-core power budgets that aim to guarantee thermally safe operation. Two knobs affect these factors: First, the mapping of threads to cores affects both the LLC latencies and the power budgets. Second, dynamic power budgeting refines the power budgets during task execution. A mapping that spatially distributes threads across the many-core increases the power budgets, but unfortunately also increases the LLC latencies. Contrarily, mapping all threads near the center of the many-core minimizes the LLC latencies, but unfortunately also decreases the power budgets. Consequently, both metrics cannot be simultaneously optimal, which leads to a Pareto-optimization for task mapping that has formerly not been exploited. Dynamic power budgeting reallocates the power budgets according to the tasks' execution phases. This results in a dynamically changing non-uniform power budget, which further increases the performance. We are the first to present a run-time algorithm PCGov combining task-agnostic task mapping and task-aware dynamic power budgeting for many-cores with shared distributed LLC. PCGov yields up to 21 percent lower response time and 13 percent lower energy consumption compared to the state-of-the-art, with a low overhead of less than 0.5 percent. Martin Rapp, Mark Sagi, Anuj Pathania, Andreas Herkersdorf, Jörg Henkel |
IEEE Trans. Computers | 1 |
| 2020 | A Lightweight Nonlinear Methodology to Accurately Model Multicore Processor PowerabstractMany power management algorithms demand accurate and fine-grained runtime estimations of dynamic core power. In the absence of fine-grained power sensors, model-based estimations are needed. Such power models commonly approximate the switching activity of logic gates using performance counters while assuming a linear performance counter/power relation at a fixed frequency and voltage. It has been shown that this relation cannot be captured accurately enough with purely linear models and that well-established nonlinear modeling techniques, e.g., polynomial modeling, easily overfit the underlying performance/power relations. Although neural-network-based modeling has shown to accurately capture nonlinear relations, it has a large training and inference overhead which is too high for fine-grained models on core-level and estimation rates in the range of 1-10 kHz. We propose a methodology for nonlinear transformation of specific performance counters to increase power modeling accuracy at constant frequency and voltage with a relatively low overhead for both model generation and run-time application over a linear model. Furthermore, we use least-angle regression (LARS) to determine a ranking of the performance counter inputs for use in linear and nonlinear modeling and show that the transformed performance counters are better suited for power modeling. The generated dynamic power model consisting of a nonlinear transformation block and a linear regression block reduces relative estimation error on average by 4% and in worst-case scenarios by 7% compared to state-of-the-art fine-grained linear power models. Compared to a state-of-the-art polynomial regression model our proposed approach reduces the relative estimation error by 10% in worst-case scenarios. Mark Sagi, Nguyen Anh Vu Doan, Martin Rapp, Thomas Wild, Jörg Henkel, Andreas Herkersdorf |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Dynamic Power and Energy Management for NCFET-Based ProcessorsabstractPower and energy consumption are the key optimization goals in all modern processors. Negative capacitance field-effect transistors (NCFETs) are a leading emerging technology that promises outstanding performance in addition to better energy efficiency. The thickness of the added ferroelectric layer as well as frequency and voltage are the key parameters that impact the power and energy of NCFET-based processors in addition to the characteristics of runtime workloads. Unlike existing CMOS technologies, operating NCFET-based processors at a higher frequency than the required minimum can result in power/energy minimization. The optimal operating point, however, strongly depends on dynamic workload characteristics and technology parameters. In this work, we propose and implement the first NCFET-aware power and energy management approach that minimizes the processor's power and energy through optimal voltage/frequency selection under different runtime scenarios. Such an NCFET-aware approach does not result in any tradeoff between power/energy and performance. Instead, it can achieve higher performance while minimizing energy. A comprehensive, simulation-based evaluation of our runtime management under realistic workloads demonstrates up to 58% energy saving with 2.1× higher performance, and 46% power saving compared to conventional NCFET-unaware management techniques, over the total execution of a benchmark. Compared to state-of-the-art NCFET-aware management techniques, our technique provides up to 49% energy saving and 32% power saving. Sami Salamin, Martin Rapp, Jörg Henkel, Andreas Gerstlauer, Hussam Amrouch |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Performance, Power and Cooling Trade-Offs with NCFET-based Many-CoresabstractNegative Capacitance Field-Effect Transistor (NCFET) is an emerging technology that incorporates a ferroelectric layer within the transistor gate stack to overcome the fundamental limit of sub-threshold swing in transistors. Even though physics-based NCFET models have been recently proposed, system-level NCFET models do not exist and research is still in its infancy. In this work, we are the first to investigate the impact of NCFET on performance, energy and cooling costs in many-core processors. Our proposed methodology starts from accurate physics models all the way up to the system level, where the performance and power of a many-core are widely affected. Our new methodology and system-level models allow, for the first time, the exploration of the novel trade-offs between performance gains and power losses that NCFET now offers to system-level designers. We demonstrate that an optimal ferroelectric thickness does exist. In addition, we reveal that current state-of-the-art power management techniques fail when NCFET (with a thick ferroelectric layer) comes into play. Martin Rapp, Sami Salamin, Hussam Amrouch, Girish Pahwa, Yogesh Singh Chauhan, Jörg Henkel |
DAC | 1 |
| 2019 | Smart Thermal Management for Heterogeneous MulticoresabstractHeterogeneous multicores have attracted a major focus in recent years, as they provide many possibilities for performance improvements. However, due to the discontinuation of Dennard scaling, on-chip power densities are continuously increasing along with technology scaling, and hence on-chip temperatures are elevated. Therefore, several thermal management techniques have emerged to keep the temperature of the chip within safe limits. These techniques, however, lead to performance losses which ultimately erase a big portion of the expected performance gains from the heterogeneous multicores. Thus, it is indispensable to deploy thermal management techniques that are able to make efficient decisions which satisfy temperature constraints while at the same time maximizing the performance. This paper presents smart thermal management techniques for heterogeneous multicores that exploit relevant information about several heterogeneity parameters at the chip level and at the application level to increase thermal efficiency1. Compared to the state of the art, the presented techniques are able to obtain significant performance improvements under the same thermal constraint.This paper is part of the DATE 2019 special session on "Smart Resource Management and Design Space Exploration for Heterogeneous Processors". The other two papers of this special session are [1] and [2]. Jörg Henkel, Heba Khdr, Martin Rapp |
DATE | 3 |
| 2019 | Prediction-Based Task Migration on S-NUCA Many-CoresabstractPerformance of a task running on a many-core with distributed shared Last-Level Cache (LLC) strongly depends on two factors: the power budget needed to guarantee thermally safe operation and the LLC latency. The task's thread-to-core mapping determines both the factors. Arrival and departure of tasks on a many-core deployed in an open system can change its state significantly in terms of available cores and power budget. Task migrations can thereupon be used as a tool to keep the many-core operating at the peak performance. Furthermore, the relative impacts of power budget and LLC latency on a task's performance can change with its different execution phases mandating its migration on-the-fly.We propose the first run-time algorithm PCMig that increases the performance of a many-core with distributed shared LLC by migrating tasks based on their phases and the many-core's state. PCMig is based on a performance-prediction model that predicts the performance impact of migrations. PCMig results in up to 16 % reduction in the average response time compared to the state-of-the-art. Martin Rapp, Anuj Pathania, Tulika Mitra, Jörg Henkel |
DATE | 1 |
| 2019 | The Impact of Emerging Technologies on Architectures and System-level Management: Invited PaperabstractThe goal of this work is to introduce and discuss different kinds of emerging technologies for logic circuitry and memory with respect to the key question of how they will impact future system-on-chip architectures and system-level management techniques. It is obvious that emerging technologies should have an impact there in order to fully exploit their technological advantages but also in order to deal with any disadvantages they might come with. In this special session paper, three promising emerging technologies are presented: (i) Negative Capacitance Field-Effect Transistor (NCFET) as a new CMOS technology with advantages primarily for low-power design, (ii) Ferroelectric FET (FeFET) as a non-volatile, area-efficient and low-power combined logic and memory as well as (iii) a Phase-Change Memory (PCM) and Resistive RAM (ReRAM) offering a large potential for tackling the memory wall problem in the von Neumann architecture. Our analysis demonstrates that not only new computing paradigms are promoted by these new technologies, it will also be seen that the trade-offs between the classical design parameters of low power, performance etc. will shift and hence emerging technologies will offer new Pareto points in the design space of future on-chip architectures. In that context, this work is unique as it bridges the gap between the technology side and system/architecture-level side to draw a vision of new technologies and their impact on architectures and system-level management. Jörg Henkel, Hussam Amrouch, Martin Rapp, Sami Salamin, Dayane Reis, Xunzhao Yin, Michael T. Niemier, Cheng Zhuo, Xiaobo Sharon Hu, Hsiang-Yun Cheng, Chia-Lin Yang |
ICCAD | 3 |
| 2019 | NCFET-Aware Voltage ScalingabstractNegative Capacitance Field-Effect Transistor (NCFET) has recently attracted significant attention. In the NCFET technology with a thick ferroelectric layer, voltage reduction increases the leakage power, rather than decreases, due to the negative Drain-Induced Barrier Lowering (DIBL) effect. This work is the first to demonstrate the far-reaching consequences of such an inverse dependency w.r.t. the existing power management techniques. Moreover, this work is the first to demonstrate that state-of-the-art Dynamic Voltage Scaling (DVS) techniques are sub-optimal for NCFET. Our investigation revealed that the optimal voltage at which the total power is minimized is not necessarily at the point of the minimum voltage required to fulfill the performance constraint (as in traditional DVS). Hence, an NCFET-aware DVS is key for high energy efficiency. In this work, we therefore propose the first NCFET-aware DVS technique that selects the optimal voltage to minimize the power following the dynamics of workloads. Our experimental results of a multi-core system demonstrate that NCFET-aware DVS results in 20% on average, and up to 27% energy saving while still fulfilling the same performance constraint (i.e., no trade-offs) compared to traditional NCFET-unaware DVS techniques. Sami Salamin, Martin Rapp, Hussam Amrouch, Girish Pahwa, Yogesh Singh Chauhan, Jörg Henkel |
ISLPED | 2 |
| 2018 | Pareto-Optimal Power- and Cache-Aware Task Mapping for Many-Cores with Distributed Shared Last-Level CacheabstractTwo factors primarily affect performance of multi-threaded tasks on many-core processors with both shared and physically distributed Last-Level Cache (LLC): the power budget associated with a certain task mapping that aims to guarantee thermally safe operation and the non-uniform LLC access latency of threads running on different cores. Spatially distributing threads across the many-core increases the power budget, but unfortunately also increases the associated LLC latency. On the other side, mapping more threads to cores near the center of the many-core decreases the LLC latency, but unfortunately also decreases the power budget. Consequently, both metrics (LLC latency and power budget) cannot be simultaneously optimal, which leads to a Pareto-optimization that has formerly not been exploited. We are the first to present a run-time task mapping algorithm called PCMap that exploits this trade-off. Our approach results in up to 8.6% reduction in the average task response time accompanied by a reduction of up to 8.5% in the energy consumption compared to the state-of-the-art. Martin Rapp, Anuj Pathania, Jörg Henkel |
ISLPED | 1 |