VLDB 2026 Research / reviewers in the wild / expert
Raid Ayoub
dblp:91/7550 · also Raid Zuhair Ayoub
· DBLP profile ↗
40ranked-venue papers
12as first author
10since 2021 · last 2026
0000-0002-1175-2983ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 12 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 1 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accurate Analytical Modeling for NoCs with Hybrid Arbitration under High Traffic InjectionabstractAnalytical performance modeling of Networks-on-Chip (NoC) are important for fast design space exploration and quick pre-silicon evaluation. Existing NoC performance analysis techniques assume certain micro-architectural details (e.g., a particular arbitration technique) to be homogeneous across the entire NoC. However, emerging NoC architectures may have hybrid arbitration across the NoC to ensure high throughput. Moreover, existing analytical models estimating performance of NoCs with finite buffers fail to analyze the performance of the NoC accurately under high traffic injection which occur in several modern-day server as well as client applications. In this work, we propose a performance analysis technique for NoCs with hybrid arbitration under high traffic injection. We propose a novel transformation to accurately compute the waiting time of the queues under hybrid arbitration. We also develop a technique to compute the effective arrival statistics to the queues when the desired injection rate is high. Thorough experimental evaluation with a wide range of injection rates at the queues of an industrial NoC show that our proposed analytical model incurs only 7% error on average and 3 orders of speed-up with respect to cycle-accurate simulation under high traffic injection. Rahul Tripathy, Mohammad Majharul Islam, Riad Akram, Raid Ayoub, Sumit K. Mandal |
DATE | 4 |
| 2024 | Multi-Objective Software-Hardware Co-Optimization for HD-PIM via Noise-Aware Bayesian OptimizationabstractIn hardware accelerator design, software-hardware co-optimization requires intricate trade-offs and tight integration between software algorithms and hardware design to optimize performance, power efficiency, and area (PPA) while ensuring high accuracy. Furthermore, the inherent non-ideality in some emerging hardware technologies poses extra challenges to the co-optimization problem. This paper proposes a novel software-hardware co-optimization framework for hyperdimensional (HD) computing accelerators with emerging ReRAM-based processing in-memory (PIM) technologies, which have shown superior performance and energy efficiency over conventional machine learning accelerators. We first comprehensively characterize the non-trivial trade-offs between design parameters in HD-PIM and PPA and accuracy metrics in HD-PIM. Then, we develop a multi-objective noise-aware Bayesian optimization algorithm to find the Pareto set (optimal trade-offs between metrics) of the HD-PIM design. Our methodology uniquely addresses the stochastic nature of ReRAM by integrating error characteristics into the optimization process, thereby enhancing the quality of the generated designs. Experimental results show that our configurations achieve up to 4.28% accuracy improvement, 35.38% power reduction, 49x timing improvement, and 10% area reduction over a non-optimized design. Chien-Yi Yang, Minxuan Zhou, Flavio Ponzina, Suraj Sathya Prakash, Raid Ayoub, Pietro Mercati, Mahesh Subedar, Tajana Rosing |
ICCAD | 5 |
| 2023 | A Lightweight Congestion Control Technique for NoCs with Deflection RoutingabstractNetwork-on-Chip (NoC) congestion builds up during heavy traffic load and leads to wasted link bandwidth, crippling the system performance. We propose a lightweight machine learning-based technique that helps predict congestion in the net-work by collecting features related to traffic at each destination and labelling it using a novel time reversal approach. The labelled data is used to design a low overhead and an explainable decision tree model used at runtime congestion control. Experimental evaluations with synthetic and real traffic on industrial$\boldsymbol{6\times 6}$NoC show that the proposed approach increases fairness and memory read bandwidth by up to 114% with respect to existing congestion control technique while incurring less than 0.01% of overhead. Shruti Yadav Narayana, Sumit K. Mandal, Raid Ayoub, Michael Kishinevsky, Ümit Y. Ogras |
DATE | 3 |
| 2023 | Uncertainty-Aware Online Learning for Dynamic Power Management in Large Manycore SystemsabstractLarge-scale manycore System-on-Chips (SoCs) need to satisfy the conflicting objectives of maximizing performance and minimizing energy consumption for dynamically changing applications. In this paper, we consider the problem of dynamic power management (DPM) in large manycore SoCs for unseen applications at runtime. We employ a machine learning (ML) based DPM policy, which selects the voltage/frequency (V/F) levels for different cluster of cores as a function of the application features such as core computation, inter-core traffic etc. We propose a novel uncertainty-aware online learning framework to learn the DPM policy, which can adapt to unseen applications at runtime. It relies on two key ideas. First, an entropy-based uncertainty measure is used to distinguish between seen and unseen system states. Second, we employ conformal prediction to compute uncertain V/F sets for unseen system states. We perform bounded-search over the uncertain V/F configurations using power/performance models to identify the best V/F configurations to minimize the energy-delay product (EDP) and create supervised examples for online learning. Our experiments on 64-core system show that the EDP is reduced by up to 50 % and 60 % when compared to existing online-imitation learning and reinforcement learning methods, respectively. Gaurav Narang, Raid Ayoub, Michael Kishinevsky, Janardhan Rao Doppa, Partha Pratim Pande |
ISLPED | 2 |
| 2023 | Dynamic Reliability Management of Multigateway IoT Edge Computing SystemsabstractThe emerging paradigm of edge computing envisions to overcome the shortcomings of cloud-centric Internet of Things (IoT) by providing data processing and storage capabilities closer to the source of data. Accordingly, IoT edge devices, with the increasing demand of computation workloads on them, are prone to failures more than ever. Hard failures in hardware due to aging and reliability degradation are particularly important since they are irrecoverable, requiring maintenance for the replacement of defective parts, at high costs. In this article, we propose a novel dynamic reliability management (DRM) technique for multigateway IoT edge computing systems to mitigate degradation and defer early hard failures. Taking advantage of the edge computing architecture, we utilize gateways for computation offloading with the primary goal of maximizing the battery lifetime of edge devices, while satisfying the Quality of Service (QoS) and reliability requirements. We present a two-level management scheme, which work together to 1) choose the offloading rates of edge devices; 2) assign edge devices to gateways; and 3) decide multihop data flow routes and rates in the network. The offloading rates are selected by a hierarchical multitimescale distributed controller. We assign edge devices by solving a bottleneck generalized assignment problem (BGAP) and compute optimal flows in a fully distributed fashion, leveraging the subgradient method. Our results, based on real measurements and trace-driven simulation, demonstrate that the proposed scheme can achieve a similar battery lifetime and better QoS compared to the state-of-the-art approaches while satisfying reliability requirements, where other approaches fail by a large margin. Kazim Ergun, Raid Ayoub, Pietro Mercati, Tajana Rosing |
IEEE Internet Things J. | 2 |
| 2023 | Dynamic Power Management in Large Manycore Systems: A Learning-to-Search FrameworkabstractThe complexity of manycore System-on-chips (SoCs) is growing faster than our ability to manage them to reduce the overall energy consumption. Further, as SoC design moves toward three-dimensional (3D) architectures, the core's power density increases leading to unacceptable high peak chip temperatures. In this article, we consider the optimization problem of dynamic power management (DPM) in manycore SoCs for an allowable performance penalty (say, 5%) and admissible peak chip temperature. We employ a machine learning– (ML) based DPM policy, which selects the voltage/frequency levels for different cluster of cores as a function of the application workload features such as core computation and inter-core traffic, and so on. We propose a novel learning-to-search (L2S) framework to automatically identify an optimized sequence of DPM decisions from a large combinatorial space for joint energy-thermal optimization for one or more given applications. The optimized DPM decisions are given to a supervised learning algorithm to train a DPM policy, which mimics the corresponding decision-making behavior. Our experiments on two different manycore architectures designed using wireless interconnect and monolithic 3D demonstrate that principles behind the L2S framework are applicable for more than one configuration. Moreover, L2S-based DPM policies achieve up to 30% energy-delay product savings and reduce the peak chip temperature by up to 17 °C compared to the state-of-the-art ML methods for an allowable performance overhead of only 5%. Gaurav Narang, Aryan Deshwal, Raid Ayoub, Michael Kishinevsky, Janardhan Rao Doppa, Partha Pratim Pande |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | Fast Performance Analysis for NoCs With Weighted Round-Robin Arbitration and Finite BuffersabstractWeighted round-robin (WRR) arbitration provides global fairness in networks-on-chip (NoCs) as opposed to the commonly used round-robin and priority-based arbitration techniques. However, the large number of weights explodes the design space and exacerbates performance (latency-throughput) tuning. Therefore, fast and accurate performance analysis techniques for NoCs are crucial for accelerating design space exploration and accurate pre-silicon evaluation. This article presents the first comprehensive performance analysis technique for NoCs with WRR arbitration and finite buffers. It can handle bursty traffic and is scalable to large NoC sizes. The proposed technique first estimates the probability that a queue is full and uses this result to compute the modified service time and queuing delay. Thorough experimental evaluations with synthetic traffic and real applications show that the proposed analytical model is always more than 10% accurate compared to cycle-accurate simulations. Moreover, the proposed performance analysis technique is five orders of magnitude faster than cycle-accurate simulations for a$16\times16$mesh NoC. Sumit K. Mandal, Shruti Yadav Narayana, Raid Ayoub, Michael Kishinevsky, Ahmed Abousamra, Ümit Y. Ogras |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | Reinforcement learning based reliability-aware routing in IoT networks
Kazim Ergun, Raid Ayoub, Pietro Mercati, Tajana Rosing |
Ad Hoc Networks | 2 |
| 2021 | Energy and QoS-Aware Dynamic Reliability Management of IoT Edge Computing SystemsabstractThe Internet of Things (IoT) systems, as any electronic or mechanical system, are prone to failures. Hard failures in hardware due to aging and degradation are particularly important since they are irrecoverable, requiring maintenance for the replacement of defective parts, at high costs. In this paper, we propose a novel dynamic reliability management (DRM) technique for IoT edge computing systems to satisfy the Quality of Service (QoS) and reliability requirements while maximizing the remaining energy of the edge device batteries. We formulate a state-space optimal control problem with a battery energy objective, QoS, and terminal reliability constraints. We decompose the problem into low-overhead subproblems and solve it employing a hierarchical and multi-timescale control approach, distributed over the edge devices and the gateway. Our results, based on real measurements and trace-driven simulation demonstrate that the proposed scheme can achieve a similar battery lifetime compared to the state-of-the-art approaches while satisfying reliability requirements, where other approaches fail to do so. Kazim Ergun, Raid Ayoub, Pietro Mercati, Dancheng Liu, Tajana Rosing |
ASP-DAC | 2 |
| 2021 | Theoretical Analysis and Evaluation of NoCs with Weighted Round-Robin ArbitrationabstractFast and accurate performance analysis techniques are essential in early design space exploration and pre-silicon evaluations, including software eco-system development. In particular, on-chip communication continues to play an increasingly important role as the many-core processors scale up. This paper presents the first performance analysis technique that targets networks-on-chip (NoCs) that employ weighted round-robin (WRR) arbitration. Besides fairness, WRR arbitration provides flexibility in allocating bandwidth proportionally to the importance of the traffic classes, unlike basic round-robin and priority-based arbitration. The proposed approach first estimates the effective service time of the packets in the queue due to WRR arbitration. Then, it uses the effective service time to compute the average waiting time of the packets. Next, we incorporate a decomposition technique to extend the analytical model to handle NoC of any size. The proposed approach achieves less than 5% error while executing real applications and 10% error under challenging synthetic traffic with different burstiness levels. Sumit K. Mandal, Jie Tong, Raid Ayoub, Michael Kishinevsky, Ahmed Abousamra, Ümit Y. Ogras |
ICCAD | 3 |
| 2020 | Online Adaptive Learning for Runtime Resource Management of Heterogeneous SoCsabstractDynamic resource management has become one of the major areas of research in modern computer and communication system design due to lower power consumption and higher performance demands. The number of integrated cores, level of heterogeneity and amount of control knobs increase steadily. As a result, the system complexity is increasing faster than our ability to optimize and dynamically manage the resources. Moreover, offline approaches are sub-optimal due to workload variations and large volume of new applications unknown at design time. This paper first reviews recent online learning techniques for predicting system performance, power, and temperature. Then, we describe the use of predictive models for online control using two modern approaches: imitation learning (IL) and an explicit nonlinear model predictive control (NMPC). Evaluations on a commercial mobile platform with 16 benchmarks show that the IL approach successfully adapts the control policy to unknown applications. The explicit NMPC provides 25% energy savings compared to a state-of-the-art algorithm for multi-variable power management of modern GPU sub-systems. Sumit K. Mandal, Ümit Y. Ogras, Janardhan Rao Doppa, Raid Ayoub, Michael Kishinevsky, Partha Pratim Pande |
DAC | 4 |
| 2020 | Performance Analysis of Priority-Aware NoCs with Deflection Routing under Traffic CongestionabstractPriority-aware networks-on-chip (NoCs) are used in industry to achieve predictable latency under different workload conditions. These NoCs incorporate deflection routing to minimize queuing resources within routers and achieve low latency during low traffic load. However, deflected packets can exacerbate congestion during high traffic load since they consume the NoC bandwidth. State-of-the-art analytical models for priority-aware NoCs ignore deflected traffic despite its significant latency impact during congestion. This paper proposes a novel analytical approach to estimate end-to-end latency of priority-aware NoCs with deflection routing under bursty and heavy traffic scenarios. Experimental evaluations show that the proposed technique outperforms alternative approaches and estimates the average latency for real applications with less than 8% error compared to cycle-accurate simulations. Sumit K. Mandal, Anish Krishnakumar, Raid Ayoub, Michael Kishinevsky, Ümit Y. Ogras |
ICCAD | 3 |
| 2019 | Dynamic Optimization of Battery Health in IoT NetworksabstractThe reliability and maintainability of the Internet of Things (IoT) devices become highly important as the number of "things" grows rapidly. The majority of the IoT devices have batteries which age, degrade, and eventually require maintenance. Existing work focuses on ensuring that batteries have sufficient amount of stored charge to operate until they can recharge, but does not consider battery degradation. This leads to high replacement and maintenance costs in large IoT networks. In this paper, we formulate the problem of minimizing battery degradation to improve the lifetime of IoT networks and solve it with Model Predictive Control (MPC) leveraging models for battery dynamics and State of Health (SoH). The battery SoH is modeled using a realistic non-linear model while taking ambient temperature into account. We demonstrate that our solution can improve network lifetime up to 68.5% compared to conventional energy consumption focused algorithms, which use simple linear battery models. The proposed approach achieves near-optimal performance in terms of preserving battery health, staying within 8.7% SoH with respect to an ideal oracle solution on average. Kazim Ergun, Raid Ayoub, Pietro Mercati, Tajana Rosing |
ICCD | 2 |
| 2019 | Analytical Performance Models for NoCs with Multiple Priority Traffic ClassesabstractNetworks-on-chip (NoCs) have become the standard for interconnect solutions in industrial designs ranging from client CPUs to many-core chip-multiprocessors. Since NoCs play a vital role in system performance and power consumption, pre-silicon evaluation environments include cycle-accurate NoC simulators. Long simulations increase the execution time of evaluation frameworks, which are already notoriously slow, and prohibit design-space exploration. Existing analytical NoC models, which assume fair arbitration, cannot replace these simulations since industrial NoCs typically employ priority schedulers and multiple priority classes. To address this limitation, we propose a systematic approach to construct priority-aware analytical performance models using micro-architecture specifications and input traffic. Our approach decomposes the given NoC into individual queues with modified service time to enable accurate and scalable latency computations. Specifically, we introduce novel transformations along with an algorithm that iteratively applies these transformations to decompose the queuing system. Experimental evaluations using real architectures and applications show high accuracy of 97% and up to 2.5× speedup in full-system simulation. Sumit K. Mandal, Raid Ayoub, Michael Kishinevsky, Ümit Y. Ogras |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | STAFF: online learning with stabilized adaptive forgetting factor and feature selection algorithmabstractDynamic resource management techniques rely on power consumption and performance models to optimize the operating frequency and utilization of processing elements, such as CPU and GPU. Despite the importance of these decisions, many existing approaches rely on fixed power and performance models that are learned offline. However, offline models cannot guarantee accuracy when workloads differ significantly from the training available at design time. This paper presents an online learning framework (STAFF) that constructs adaptive run-time models for stationary and non-stationary workloads. STAFF is the first framework that (1) guarantees stability while quickly adapting to workload changes, (2) performs online feature selection with linear complexity, and (3) adapts to new model coefficients by employing adaptively varying forgetting factor, all at the same time. Experiments on an Intel® Coreh™ i5 6th generation platform demonstrate up to 6× improvement in the performance prediction accuracy compared to existing techniques. Ujjwal Gupta, Manoj Babu, Raid Ayoub, Michael Kishinevsky, Francesco Paterna, Ümit Y. Ogras |
DAC | 3 |
| 2018 | An Online Learning Methodology for Performance Modeling of Graphics ProcessorsabstractApproximately 18 percent of the 3.2 million smartphone applications rely on integrated graphics processing units (GPUs) to achieve competitive performance. Graphics performance, typically measured in frames per second, is a strong function of the GPU frequency, which in turn has a significant impact on mobile processor power consumption. Consequently, dynamic power management algorithms have to assess the performance sensitivity to the frequency accurately to choose the operating frequency of the GPU effectively. Since the impact of GPU frequency on performance varies rapidly over time, there is a need for online performance models that can adapt to varying workloads. This paper presents a light-weight adaptive runtime performance model that predicts the frame processing time of graphics workloads at runtime without apriori characterization. We employ this model to estimate the frame time sensitivity to the GPU frequency, i.e., the partial derivative of the frame time with respect to the GPU frequency. The proposed model does not rely on any parameter learned offline. Our experiments on commercial platforms with common GPU benchmarks show that the mean absolute percentage error in frame time and frame time sensitivity prediction are 4.2 and 6.7 percent, respectively. Ujjwal Gupta, Manoj Babu, Raid Ayoub, Michael Kishinevsky, Francesco Paterna, Suat Gumussoy, Ümit Y. Ogras |
IEEE Trans. Computers | 3 |
| 2017 | Multi-variable Dynamic Power Management for the GPU SubsystemabstractIn this work, we present a control-theoretic algorithm to improve the energy efficiency of the GPU targeting deadline-driven graphics applications. Our algorithm dynamically controls multiple power knobs within the GPU (DVFS and number of active slices) that have different control time granularities. We developed a multi-rate predictive control to overcome the time granularity constraints in the control variables and reduce runtime overhead. To enable predictive control, we developed runtime analytical predictive models for performance and power of the GPU, that take input from hardware counters and temperature sensor readings. We evaluated our approach on the latest generation of Intel Core i5 platform. Our experimental results demonstrate significant average GPU energy savings of 25% compared to the state-of-the-art algorithm at negligible performance overhead. Pietro Mercati, Raid Ayoub, Michael Kishinevsky, Eric Samson, Marc Beuchat, Francesco Paterna, Tajana Rosing |
DAC | 2 |
| 2017 | User-aware Frame Rate Management in Android SmartphonesabstractFrame rate has a direct impact on the energy consumption of smartphones: the higher the frame rate, the higher the power consumption. Hence, reducing display refreshes will reduce the power consumption. However, it is risky to manipulate frame rate drastically as it can deteriorate user satisfaction with the device. In this work, we introduce a screen management system that controls the frame rate on smartphone displays based on a model that detects user dissatisfaction due to display refreshes. This approach is based on understanding when higher frame rates are necessary, and providing lower frame rates —thus, saving power— if the lower rate is predicted not to cause user dissatisfaction. According to the results of our first user survey with 20 participants, individuals show highly varying requirements: while some users require high frame rates for the highest satisfaction, others are equally satisfied with lower frame rates. Based on this observation, we develop a system that predicts user dissatisfaction on the runtime and either increases or decreases the maximum frame rate setting. For user dissatisfaction predictions, we have compared two different approaches: (1) static model, which uses dissatisfaction characteristics of a fixed group of people, and (2) user-specific model, which is learning only from the specific user. Our second set of experiments with 20 participants shows that users report 32% less dissatisfaction and 4% more dissatisfaction than the default Android system with user-specific and static systems, respectively. These experiments also show that, compared to the default scheme, our mechanisms reduce the power consumption of the phone by 7.2% and 1.8% on average with the user-specific and static models, respectively. Begum Egilmez, Matthew Schuchhardt, Gokhan Memik, Raid Ayoub, Niranjan Soundararajan, Michael Kishinevsky |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2016 | Adaptive performance prediction for integrated GPUsabstractIntegrated GPUs have become an indispensable component of mobile processors due to the increasing popularity of graphics applications. The GPU frequency is a key factor both in application throughput and mobile processor power consumption under graphics workloads. Therefore, dynamic power management algorithms have to assess the performance sensitivity to the GPU frequency accurately. Since the impact of the GPU frequency on performance varies rapidly over time, there is a need for online performance models that can adapt to varying workloads. This paper presents a light-weight adaptive runtime performance model that predicts the frame processing time. We use this model to estimate the frame time sensitivity to the GPU frequency. Our experiments on a mobile platform running common GPU benchmarks show that the mean absolute percentage error in frame time and frame time sensitivity prediction are 3.8% and 3.9%, respectively. Ujjwal Gupta, Joseph Campbell, Ümit Y. Ogras, Raid Ayoub, Michael Kishinevsky, Francesco Paterna, Suat Gumussoy |
ICCAD | 4 |
| 2015 | Optimizing mobile display brightness by leveraging human visual perceptionabstractModern smartphones and tablets are battery-constrained by their mobility; this constraint is heavily factored into any design decision made on the device. Furthermore, the display is one of the most power-consuming subsystems. Adaptive display brightness systems attempt to address this high display power consumption by setting the brightness depending on the surrounding ambient light levels. Matthew Schuchhardt, Susmit Jha, Raid Ayoub, Michael Kishinevsky, Gokhan Memik |
CASES | 3 |
| 2015 | A control-theoretic approach for energy efficient CPU-GPU subsystem in mobile platformsabstractThis paper presents a control-theoretic approach to optimize the energy consumption of integrated CPU and GPU subsystems for graphic applications. It achieves this via a dynamic management of the CPU and GPU frequencies. To this end, we first model the interaction between the GPU and CPU as a queuing system. Second, we formulate a Multi-Input-Multi-Output state-space closed loop control to ensure robustness and stability. We evaluated this control on an Intel Baytrail-based Android platform. Experimental evaluations show energy savings of 17.4% in the CPU-GPU subsystem with a low performance impact of 0.9%. David Kadjo, Raid Ayoub, Michael Kishinevsky, Paul Gratz |
DAC | 2 |
| 2014 | CAPED: Context-aware personalized display brightness for mobile devicesabstractThe display remains the primary user interface on many computing devices, ranging from traditional devices such as desktops and laptops, to the more pervasive devices such as smartphones and smartwatches. Thus, the overall user experience with these computing devices is greatly determined by the display subsystem. Ideal display brightness is critical to good user experience, but actually predicting the ideal brightness level which would most satisfy the user is a challenge. Finding the right screen brightness is even more challenging on mobile devices (which is the focus of this work), as the screen tends to be one of the most power consuming components. Currently, the control of display brightness is usually done through a simplistic, static one-size-fits-all model which chooses a fixed brightness level for a given ambient light condition. Matthew Schuchhardt, Susmit Jha, Raid Ayoub, Michael Kishinevsky, Gokhan Memik |
CASES | 3 |
| 2013 | Dynamic voltage and frequency scaling for shared resources in multicore processor designsabstractAs the core count in processor chips grows, so do the on-die, shared resources such as on-chip communication fabric and shared cache, which are of paramount importance for chip performance and power. This paper presents a method for dynamic voltage/frequency scaling of networks-on-chip and last level caches in multicore processor designs, where the shared resources form a single voltage/frequency domain. Several new techniques for monitoring and control are developed, and validated through full system simulations on the PARSEC benchmarks. These techniques reduce energy-delay product by 56% compared to a state-of-the-art prior work. Zheng Xu 0006, Paul Gratz, Jiang Hu 0001, Michael Kishinevsky, Ümit Y. Ogras, Raid Ayoub |
DAC | 8 |
| 2013 | Temperature aware thread block scheduling in GPGPUsabstractIn this paper, we present a first general purpose GPU thermal management design that consists of both hardware architecture and OS scheduler changes. Our techniques schedule thread blocks from multiple computational kernels in spatial, temporal, and spatio-temporal ways depending on the thermal state of the system. We can reduce the computation slowdown by 60% on average relative to the state of the art techniques while meeting the thermal constraints. We also extend our work to multi GPGPU cards and show improvements of 44% on average relative to existing technique. Rajib Nath, Raid Ayoub, Tajana Rosing |
DAC | 2 |
| 2013 | Managing mobile platform powerabstractPower consumption has been one of the major design considerations for more than a decade [6]. Hence, energy efficient techniques have been widely studied to harness the processing power within available power and thermal budgets [3][4][10]. With the proliferation of smart mobile devices, the criticality of energy efficiency is multiplied. On one hand, increasing computational power as well as sensing, storage, and communication capabilities open up wide range of power-hungry application domains. On the other hand, the battery life rises as one of the major concerns of the end user [9]. Furthermore, these fanless devices are subject to tight surface, or skin, temperature constraints which limit the peak power consumption, since the skin temperature directly affects the user experience (UX). As a result, power management techniques crafted specifically for smart mobile devices become necessary. In this paper, we review three differentiating aspects for managing the power of smart mobile devices. More specifically, we emphasize the importance of platform view, user experience and platform level optimization. Ümit Y. Ogras, Raid Ayoub, Michael Kishinevsky, David Kadjo |
ICCAD | 2 |
| 2013 | Power gating with block migration in chip-multiprocessor last-level cachesabstractWe propose a novel technique to significantly reduce the leakage energy of last level caches while mitigating any significant performance impact. In general, cache blocks are not ordered by their temporal locality within the sets; hence, simply power gating off a partition of the cache, as done in previous studies, may lead to considerable performance degradation. We propose a solution that migrates the high temporal locality blocks to facilitate power gating, where blocks likely to be used in the future are migrated from the partition being shutdown to the live partition at a negligible performance impact and hardware overhead. Our detailed simulations show energy savings of 66% at low performance degradation of 2.16%. David Kadjo, Paul Gratz, Jiang Hu 0001, Raid Ayoub |
ICCD | 5 |
| 2013 | CoMETC: Coordinated management of energy/thermal/cooling in serversabstractWe introduce a Coordinated Management of Energy, Thermal, and Cooling (CoMETC) technique to minimize cooling and memory energy of server machines. State-of-the-art solutions decouple the optimization of cooling energy costs and energy consumption of CPU and memory subsystems. This results in suboptimal solutions due to thermal dependencies between CPU and memory and the nonlinearity in energy costs of cooling. In contrast, we develop a unified solution that integrates energy, thermal, and cooling management for CPU and memory subsystems to maximize energy savings. CoMETC reduces the operational energy of the memory by clustering active memory pages to a subset of memory modules while accounting for thermal and cooling aspects. At the same time, CoMETC removes hotspots between and within the CPU sockets and reduces the effects of thermal coupling with memory in order to minimize cooling energy costs. We design CoMETC using a control-theoretic approach to guarantee meeting these objectives. We introduce a formal thermal and cooling model to be used for online decisions inside CoMETC. Our experimental results show that CoMETC achieves average cooling and memory energy savings of 58% compared to state-of-the-art techniques at a performance overhead of less than 0.3%. Raid Ayoub, Rajib Nath, Tajana Rosing |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2012 | TempoMP: Integrated prediction and management of temperature in heterogeneous MPSoCsabstractHeterogeneous Multi-Processor Systems on a Chip (MPSoCs) are more complex from a thermal perspective compared to the homogeneous MPSoCs because of their inherent imbalance in power density. In this work we develop TempoMP, a new technique for thermal management of heterogeneous MPSoCs which leverages multi-parametric optimization along with our novel thermal predictor, Tempo. TempoMP is able to deliver locally optimal dynamic thermal management decisions to meet thermal constraints while minimizing power and maximizing performance. It leverages our Tempo predictor which, unlike the previous techniques, can estimate the impact of future power state changes at negligible overhead. Our experiments show that compared to the state of the art, Tempo can reduce the maximum prediction error by up to an order of magnitude. Our experiments with heterogeneous MPSoCs also show that TempoMP meets thermal constraints while reducing the average task lateness by 2.5X and energy-lateness product by 5X compared to the state of the art techniques. Shervin Sharifi, Raid Ayoub, Tajana Rosing |
DATE | 2 |
| 2012 | JETC: Joint energy thermal and cooling management for memory and CPU subsystems in serversabstractIn this work we propose a joint energy, thermal and cooling management technique (JETC) that significantly reduces per server cooling and memory energy costs. Our analysis shows that decoupling the optimization of cooling energy of CPU & memory and the optimization of memory energy leads to suboptimal solutions due to thermal dependencies between CPU and memory and non-linearity in cooling energy. This motivates us to develop a holistic solution that integrates the energy, thermal and cooling management to maximize energy savings with negligible performance hit. JETC considers thermal and power states of CPU & memory, thermal coupling between them and fan speed to arrive at energy efficient decisions. It has CPU and memory actuators to implement its decisions. The memory actuator reduces the energy of memory by performing cooling aware clustering of memory pages to a subset of memory modules. The CPU actuator saves cooling energy by reducing the hot spots between and within the CPU sockets and minimizing the effects of thermal coupling. Our experimental results show that employing JETC results in 50.7% average energy reduction in cooling and memory subsystems with less than 0.3% performance overhead. Raid Ayoub, Rajib Nath, Tajana Rosing |
HPCA | 1 |
| 2011 | OS-level power minimization under tight performance constraints in general purpose systems
Raid Ayoub, Ümit Y. Ogras, Eugene Gorbatov, Yanqin Jin, Timothy Kam, Paul Diefenbaugh, Tajana Rosing |
ISLPED | 1 |
| 2011 | Temperature Aware Dynamic Workload Scheduling in Multisocket CPU ServersabstractIn this paper, we propose a multitier approach for significantly lowering the cooling costs associated with fan subsystems without compromising the system performance. Our technique manages the fan speed by intelligently allocating the workload at the core level as well as at the CPU socket level. At the core level we propose a proactive dynamic thermal management scheme. We introduce a new predictor that utilizes the band-limited property of the temperature frequency spectrum. A big advantage of our predictor is that it does not require the costly training phase and still maintains high accuracy. At the socket level, we use control theoretic approach to develop a stable scheduler that reduces the cooling costs further by providing a better thermal distribution. Our thermal management scheme incorporates runtime workload characterization to perform efficient thermally aware scheduling. The experimental results show that our approach delivers an average cooling energy savings of 80% compared to the state of the art techniques. The reported results also show that our formal technique maintains stability while heuristic solutions fail in this aspect. Raid Ayoub, Krishnam Raju Indukuri, Tajana Rosing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2010 | Cool and save: cooling aware dynamic workload scheduling in multi-socket CPU systemsabstractTraditionally CPU workload scheduling and fan control in multi-socket systems have been designed separately leading to less efficient solutions. In this paper we present Cool and Save, a cooling aware dynamic workload management strategy that is significantly more energy efficient than state-of-the art solutions in multi-socket CPU systems because it performs workload scheduling in tandem with controlling socket fan speeds. Our experimental results indicate that applying our scheme gives average fan energy savings of 73% concurrently with reducing the maximum fan speed by 53%, thus leading to lower vibrations and noise levels. Raid Ayoub, Tajana Rosing |
ASP-DAC | 1 |
| 2010 | GentleCool: Cooling aware proactive workload scheduling in multi-machine systemsabstractIn state of the art systems, workload scheduling and server fan speed operate independently leading to cooling inefficiencies. In this work we propose GentleCool, a proactive multi-tier approach for significantly lowering the fan cooling costs without compromising the performance. Our technique manages the fan speed through intelligently allocating the workload across different machines. The experimental results show our approach delivers average cooling energy savings of 72% and improves the mean time between failures (MTBF) of the fans by 2.3× compared to the state of the art. Raid Ayoub, Shervin Sharifi, Tajana Rosing |
DATE | 1 |
| 2010 | Performance and energy efficient cache migrationapproach for thermal management in embedded systemsabstractIn this paper we propose an approach for performance and power aware warm start for the data cache during core level migration events that originate from overheating. We utilize the concept of reuse in the references to eliminate unnecessary information from being migrated. Furthermore, we exploit the temperature predictability to trigger the cache migration slightly before the actual thermal limit to allow sufficient time for the extraction of the reuse information and transfer of it to the destination cache while the execution is resuming normally. The suggested hardware not only is cost efficient but is also programmable so as to maintain flexibility in targeting application particularities. The experimental results we provide confirm the applicability of this approach. Raid Ayoub, Alex Orailoglu |
ACM Great Lakes Symposium on VLSI | 1 |
| 2010 | Energy efficient proactive thermal management in memory subsystemabstractEnergy management of memory subsystem is challenging due to performance and thermal constraints. Big energy gains can be obtained by clustering memory accesses, however this also leads to a higher need for cooling due to larger temperatures in active areas of memory. Our solution to memory thermal management problem is based on proactive thermal management that intelligently allocates workload pages to few memory units and powers down rest of the memory. Our experimental results show that this approach improves energy savings by 43% and reduces performance overhead by 85% with respect to the state of the art polices. Raid Ayoub, Krishnam Raju Indukuri, Tajana Rosing |
ISLPED | 1 |
| 2009 | Filtering Global History: Power and Performance Efficient Branch PredictorabstractIn this paper we present an Application Customizable Branch Predictor, ACBP, that delivers efficiency in energy savings and performance without compromising prediction accuracy. The idea of our technique is to filter unnecessary global history information within the global history register to minimize the predictor size while maintaining prediction accuracy. We suggest in this work an efficient algorithm to capture the beneficial correlations. A cost-efficient and programmable hardware architecture is presented. Extensive experimental analysis confirms significant improvements in power savings and latency, ranging up to 84% and 30%,respectively. Raid Ayoub, Alex Orailoglu |
ASAP | 1 |
| 2009 | PDRAM: a hybrid PRAM and DRAM main memory systemabstractIn this paper, we propose PDRAM, a novel energy efficient main memory architecture based on phase change random access memory (PRAM) and DRAM. The paper explores the challenges involved in incorporating PRAM into the main memory hierarchy of computing systems, and proposes a low overhead hybrid hardware-software solution for managing it. Our experimental results indicate that our solution is able to achieve average energy savings of 30% at negligible overhead over conventional memory architectures. Gaurav Dhiman 0001, Raid Ayoub, Tajana Rosing |
DAC | 2 |
| 2009 | Predict and act: dynamic thermal management for multi-core processorsabstractIn this paper, we propose a proactive dynamic thermal management scheme for chip multiprocessors that run multi-threaded workloads. We introduce a new predictor that utilizes the band-limited property of the temperature frequency spectrum. A big advantage of our predictor is that it does not require the costly training phase like ARMA [7]. Our thermal management scheme incorporates temperature prediction information and runtime workload characterization to perform efficient thermally aware scheduling. Our results show that applying our algorithm considerably improves the average system temperature, hottest core temperature, product MTTF and performance by 6 °C, 8 °C, 41% and 72% respectively. Raid Ayoub, Tajana Rosing |
ISLPED | 1 |
| 2007 | Power efficient register file update approach for embedded processorsabstractIn this paper we present an approach for a low power register file in the domain of embedded processors. The suggested approach obtains power savings through tackling the unnecessary writes to register files for short live registers. Writes to register files are essentially redundant when an instruction manages to forward its results to all of its dependents through forwarding hardware. As the percentage of registers that exhibit short liveness is shown to be significant, tackling unnecessary writes contributes to delivering appreciable power savings. In this work we show that tackling the unnecessary writes could be attained efficiently through a register based encoding scheme. The suggested encoding scheme exploits application-specific information and renames all or most of the short live registers to a small subset of the registers that are prespecified during the hardware design. The renaming process is performed at the compiler level. Power savings can be obtained through precluding the set of prespecified registers from writing to the register file. We suggest in this paper efficient algorithms for the purpose of renaming, one algorithm to perform the renaming in the cases of no register pressure and another one for the cases of register pressure. In the cases of register pressure, some of the prespecified registers may need to be turned into normal registers, a process that is managed through the use of reprogrammable hardware support. Although the cases of register pressure could impact power savings, the detailed analysis we outline shows that the size of the prespecified registers subset is typically small which makes register pressure an infrequent event. Experimental analysis on numerical and DSP codes indicates appreciable improvements in power savings. Raid Ayoub, Alex Orailoglu |
ICCD | 1 |
| 2005 | A unified transformational approach for reductions in fault vulnerability, power, and crosstalk noise & delay on processor busesabstractIn this paper we propose a coding scheme for general-purpose applications that can reduce power dissipation, crosstalk noise and crosstalk delay on the bus lines while simultaneously detecting errors at run time. The reduction in power dissipation can be achieved through reducing the bus switching activity. Not only is the switching activity in individual lines reduced but so is the coupling activity across the adjacent lines, the major contributor to the overall power dissipation in deep submicron technology. Detailed analysis of crosstalk noise and delay shows that eliminating certain patterns of transitions and reducing the infeasible ones in terms of crosstalk noise and power dissipation is a feasible strategy for alleviating these problems. We propose an encoding technique consisting of the use of predefined patterns of transitions, one for each possible combination of input data, to generate the codewords. The restriction to the predefined patterns of transitions enables fast encoding and low hardware overhead. This work presents an extensive analysis of the consequent reduction in crosstalk and power. SPICE derived experimental results show a reduction in worst case crosstalk delay and noise, ranging up to 24% and 10% respectively. Extensive experimental results for various applications show significant reduction in power dissipation ranging up to 44% for switching activity on the bus lines and up to 25% for coupling activity. The results also show a drastic reduction ranging up to 98% in the number of patterns that are most likely to produce crosstalk errors. Raid Ayoub, Alex Orailoglu |
ASP-DAC | 1 |