Anil Kanduri

dblp:144/4346 · DBLP profile ↗
← Back
27ranked-venue papers
6as first author
16since 2021 · last 2025
0000-0003-3188-8703ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 6 first-author · 15 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
YearPublicationVenuePosition
2025 HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge Platforms
abstract
Edge inference techniques partition and distribute Deep Neural Network (DNN) inference tasks among multiple edge nodes for low latency inference, without considering the core-level heterogeneity of edge nodes. Further, default DNN inference frameworks also do not fully utilize the resources of heterogeneous edge nodes, resulting in higher inference latency. In this work, we propose a hierarchical DNN partitioning strategy (HiDP) for distributed inference on heterogeneous edge nodes. Our strategy hierarchically partitions DNN workloads at both global and local levels by considering the core-level heterogeneity of edge nodes. We evaluated our proposed HiDP strategy against relevant distributed inference techniques over widely used DNN models on commercial edge devices. On average our strategy achieved 38% lower latency, 46% lower energy, and 56% higher throughput in comparison with other relevant approaches,
Zain Taufique, Aman Vyas, Antonio Miele, Pasi Liljeberg, Anil Kanduri
DATE5
2025 Invited Paper: Mindful AI for Pervasive Health and Wellbeing (PHW)
abstract
Emerging AI-driven pervasive health and wellbeing (PHW) services (e.g., personalized health assistants and mobile health applications) face critical challenges in handling noisy/intermittent sensory data, integrating cross-modal insights, and stringent energy and compute constraints. We present Mindful AI, a cognitive-inspired framework designed to enable adaptive, resilient, and efficient PHW services in real-world conditions. Our dual-mode intelligence—Automatic (System 1) and Reflective (System 2)—selectively directs system attention toward the most relevant sensing and compute contexts, unifying bottom-up stimuli (driven by input quality, inference demands and model confidence, and resource availability) with top-down insights (reflecting user demands, system goals/constraints, and contextual information). Our framework distills and orchestrates insights across sensing, communication, and computation through hybrid attention toward bottom-up and top-down insights that support cross-layer sense-compute co-optimization to achieve resilient, low-latency, and energy-efficient PHW services. We evaluate our approach on multi-tier device-edge-cloud platforms, using real-world case studies in pain assessment, stress monitoring, and human activity recognition to demonstrate adaptation to real-world uncertainties (e.g., sensor degradation, context drift, network variability), while maintaining strict QoS, accuracy, and latency guarantees.
Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt
ICCAD2
2025 Twill: Scheduling Compound AI Systems on Heterogeneous Mobile Edge Platforms
abstract
Compound AI (cAI) systems chain multiple AI models to solve complex problems. cAI systems are typically composed of deep neural networks (DNNs), transformers, and large language models (LLMs), exhibiting a high degree of computational diversity and dynamic workload variation. Deploying cAI services on mobile edge platforms poses a significant challenge in scheduling concurrent DNN-transformer inference tasks, which arrive dynamically in an unknown sequence. Existing mobile edge AI inference strategies manage multi-DNN or transformer-only workloads, relying on design-time profiling, and cannot handle concurrent inference of DNNs and transformers required by cAI systems. In this work, we address the challenge of scheduling cAI systems on heterogeneous mobile edge platforms. We present Twill, a run-time framework to handle concurrent inference requests of cAI workloads through task affinity-aware cluster mapping and migration, priority-aware task freezing/unfreezing, and Dynamic Voltage/Frequency Scaling (DVFS), while minimizing inference latency within power budgets. We implement and deploy our Twill framework on the Nvidia Jetson Orin NX platform. We evaluate Twill against state-of-the-art edge AI inference techniques over contemporary DNNs and LLMs, reducing inference latency by 54% on average, while honoring power budgets.
Zain Taufique, Aman Vyas, Antonio Miele, Pasi Liljeberg, Anil Kanduri
ICCAD5
2025 Exploiting Approximation for Run-time Resource Management of Embedded HMPs
abstract
Run-time resource management (RTM) of multi-programmed workloads on heterogeneous multi-core platforms is challenging due to (i) fixed power budget of the device, (ii) variable performance requirements of the workloads, and (iii) unknown arrival of the applications. Existing RTM solutions lack power-performance coordination, resulting in performance degradation during power actuation or power violations during performance provisioning. Exploiting inherent error-resilience of the applications can address the performance loss incurred in power actuation, by combining run-time approximation with traditional power knobs (including Dynamic Voltage/Frequency Scaling, Task Migration, Degree of Parallelism, and CPU Quota ). In this work, we present an accuracy-aware resource management framework that jointly actuates run-time approximation and traditional power knobs for efficient power-performance management of multi-programmed and multi-threaded workloads running on heterogeneous mobile platforms. Our strategy configures the accuracy of the applications at run-time to exploit accuracy-performance trade-offs, by considering system-wide power-performance dynamics. We use heuristic estimation models to jointly enforce accuracy configuration and traditional power knobs settings at run-time. We evaluated our framework on real-world embedded mobile platforms, including Odroid XU3 and Asus Tinker Edge R boards to demonstrate the efficiency of our proposed approach across multiple workload scenarios. Our approach achieved 25% lower performance violations against the state-of-the-art run-time resource management policies at the cost of 2.2% accuracy loss across six applications.
Zain Taufique, Anil Kanduri, Antonio Miele, Amir-Mohammad Rahmani, Cristiana Bolchini, Nikil Dutt, Pasi Liljeberg
ACM Trans. Embed. Comput. Syst.2
2024 Adaptive Workload Distribution for Accuracy-aware DNN Inference on Collaborative Edge Platforms
abstract
DNN inference can be accelerated by distributing the workload among a cluster of collaborative edge nodes. Heterogeneity among edge devices and accuracy-performance trade-offs of DNN models present a complex exploration space while catering to the inference performance requirements. In this work, we propose adaptive workload distribution for DNN inference, jointly considering node-level heterogeneity of edge devices, and application-specific accuracy and performance requirements. Our proposed approach combinatorially optimizes heterogeneity-aware workload partitioning and dynamic accuracy configuration of DNN models to ensure performance and accuracy guarantees. We tested our approach on an edge cluster of Odroid XU4, Raspberry Pi4, and Jetson Nano boards and achieved an average gain of 41.52% in performance and 5.2% in output accuracy as compared to state-of-the-art workload distribution strategies.
Zain Taufique, Antonio Miele, Pasi Liljeberg, Anil Kanduri
ASPDAC4
2024 ECG Unveiled: Analysis of Client Re-Identification Risks in Real-World ECG Datasets
abstract
While ECG data is crucial for diagnosing and monitoring heart conditions, it also contains unique biometric information that poses significant privacy risks. Existing ECG re-identification studies rely on exhaustive analysis of numerous deep learning features, confining to ad-hoc explainability towards clinicians decision making. In this work, we delve into explainability of ECG re-identification risks using transparent machine learning models. We use SHapley Additive exPlanations (SHAP) analysis to identify and explain the key features contributing to re-identification risks. We conduct an empirical analysis of identity re-identification risks using ECG data from five diverse real-world datasets, encompassing 223 participants. By employing transparent machine learning models, we reveal the diversity among different ECG features in contributing towards re-identification of individuals with an accuracy of 0.76 for gender, 0.67 for age group, and 0.82 for participant ID re-identification. Our approach provides valuable insights for clinical experts and guides the development of effective privacy-preserving mechanisms. Further, our findings emphasize the necessity for robust privacy measures in real-world health applications and offer detailed, actionable insights for enhancing data anonymization techniques.
Anil Kanduri, Seyed Amir Hossein Aqajari, Salar Jafarlou, Sanaz R. Mousavi, Pasi Liljeberg, Shaista Malik, Amir-Mohammad Rahmani
BSN2
2024 Work-in-Progress: Context and Noise Aware Resilience for Autonomous Driving Applications
abstract
Autonomous Vehicles (AVs) often use noise prone sensory data from cameras and LiDAR for perception. In specific noisy scenarios, different object detection models exhibit non-intuitive and varying degrees of resilience, necessitating adaptive model selection. In this work, we develop a context and noise aware framework for run-time adaptive configuration of objection models for high accuracy and low latency inference. We combine driving scene context and input data noise to prioritize among input modalities, followed by selection and configuration of most resilient object detection model appropriate for the context. Our evaluation for 2D object detection on nuScenes dataset provided average 1.83x speedup in latency compared to baseline while preserving average prediction confidence.
Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt
CODES+ISSS2
2024 SEAL: Sensing Efficient Active Learning on Wearables through Context-awareness
abstract
In this paper, we introduce SEAL, a co-optimization framework designed to enhance both sensing and querying strategies in wearable devices for mHealth applications. Employing Reinforcement Learning (RL), SEAL strategically utilizes user contextual information and the machine learning model's confidence levels to make efficient decisions. This innovative approach is particularly significant in addressing the challenge of battery drain due to continuous physiological signal sensing, such as Photoplethysmography (PPG). Our framework demonstrates its effectiveness in a stress monitoring application, achieving a substantial reduction of 76% in the volume of PPG signals collected, while only experiencing a minor 6% decrease in user-labeled data quality. This balance showcases SEAL's potential in optimizing data collection in a way that is considerate of both device constraints and data integrity.
Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt
DATE3
2024 Tango: Low Latency Multi-DNN Inference on Heterogeneous Edge Platforms
abstract
There is an increasing demand to run DNN applications on edge platforms for low-latency inference. Executing multi-DNN workloads with diverse compute and latency requirements on resource-constrained heterogeneous edge platforms poses a significant scheduling challenge. In this work, we present Tango framework for orchestrating multi-DNN inference on heterogeneous edge platforms. Our approach uses a Proximal Policy-based Reinforcement Learning agent to jointly optimize cluster selection, accuracy configuration, and frequency scaling to minimize inference latency with a tolerable accuracy loss. We implemented the proposed Tango framework as a portable middleware and deployed it on real hardware of the Jetson TX edge platform. Our evaluation against relevant multi-DNN scheduling strategies demonstrates 61 % lower latency and 48.4 % lower energy consumption at a maximum accuracy loss of 1.59 %.
Zain Taufique, Aman Vyas, Antonio Miele, Pasi Liljeberg, Anil Kanduri
ICCD5
2024 EA^2: Energy Efficient Adaptive Active Learning for Smart Wearables
abstract
Mobile Health (mHealth) applications rely on supervised Machine Learning (ML) algorithms, requiring end-user-labeled data for the training phase. The gold standard for obtaining such labeled data is by sending queries to users and gathering responses for the corresponding label, which was conventionally done through triggering questions sent at random. Active Learning (AL) methods use intelligent query-sending policies by incorporating users' contextual information to maximize the response rate and informativeness of the collected labeled data. However, wearable devices' substantial battery drainage associated with the sensing of physiological signals underscores the need for developing an efficient sensing policy in addition to a query-sending policy. In this work, we present a co-optimization framework for both sensing and querying strategies within wearable devices, leveraging contextual information and ML model's prediction confidence. We designed a Reinforcement Learning (RL) agent to quantify different contextual parameters combined with model confidence to determine sensing and querying decisions. Our evaluation of an exemplar stress monitoring application showed a 76% reduction in sensing and data transmission energy consumption, with only a 6% drop in user-labeled data.
Hamidreza Alikhani, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt
ISLPED3
2024 Adaptive approximate computing in edge AI and IoT applications: A review
abstract
Recent advancements in hardware and software systems have been driven by the deployment of emerging smart health and mobility applications. These developments have modernized the traditional approaches by replacing conventional computing systems with cyber-physical and intelligent systems combining the Internet of Things (IoT) with Edge Artificial Intelligence. Despite the many advantages and opportunities of these systems within various application domains, the scarcity of energy, extensive computing needs, and limited communication must be considered when orchestrating their deployment. Inducing savings in these directions is central to the Approximate Computing (AxC) paradigm, in which the accuracy of some operations is traded off with energy, latency, and/or communication reductions. Unfortunately, the dynamics of the environments in which AxC-equipped IoT systems operate have been paid little attention. We bridge this gap by surveying adaptive AxC techniques applied to three emerging application domains, namely autonomous driving, smart sensing and wearables, and positioning, paying special attention to hardware acceleration. We discuss the challenges of such applications, how adaptive AxC can aid their deployment, and which savings it can bring based on traits of the data and devices involved. Insights arising thereof may serve as inspiration to researchers, engineers, and students active within the considered domains.
Hans Jakob Damsgaard, Antoine Grenier, Dewant Katare, Zain Taufique, Salar Shakibhamedan, Tiago Troccoli, Georgios Chatzitsompanis, Anil Kanduri, Aleksandr Ometov, Aaron Yi Ding, Nima Taherinejad, Georgios Karakonstantis, Roger F. Woods, Jari Nurmi
J. Syst. Archit.8
2022 AMSER: Adaptive Multimodal Sensing for Energy Efficient and Resilient eHealth Systems
abstract
eHealth systems deliver critical digital healthcare and wellness services for users by continuously monitoring physiological and contextual data. eHealth applications use multi-modal machine learning kernels to analyze data from different sensor modalities and automate decision-making. Noisy inputs and motion artifacts during sensory data acquisition affect the i) prediction accuracy and resilience of eHealth services and ii) energy efficiency in processing garbage data. Monitoring raw sensory inputs to identify and drop data and features from noisy modalities can improve prediction accuracy and energy efficiency. We propose a closed-loop monitoring and control framework for multi-modal eHealth applications, AMSER, that can mitigate garbage-in garbage-out by i) monitoring input modalities, ii) analyzing raw input to selectively drop noisy data and features, and iii) choosing appropriate machine learning models that fit the configured data and feature vector - to improve prediction accuracy and energy efficiency. We evaluate our AMSER approach using multi-modal eHealth applications of pain assessment and stress monitoring over different levels and types of noisy components incurred via different sensor modalities. Our approach achieves up to 22% improvement in prediction accuracy and 5.6× energy consumption reduction in the sensing phase against the state-of-the-art multi-modal monitoring application.
Emad Kasaeyan Naeini, Sina Shahhosseini, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani, Nikil Dutt
DATE3
2022 Concurrent Application Bias Scheduling for Energy Efficiency of Heterogeneous Multi-Core Platforms
abstract
Minimizing energy consumption of concurrent applications on heterogeneous multi-core platforms is challenging given the diversity in energy-performance profiles of both the applications and hardware. Adaptive learning techniques made the exhaustive Pareto-optimal space exploration practically feasible to identify an energy efficient configuration. Existing approaches consider a single application's characteristic for optimizing energy consumption. However, an optimal configuration for a given application in isolation may not be optimal when other applications are run concurrently. Approaches that consider concurrent application scenarios overlook the weight of total energy consumption per application, restricting them from prioritizing among applications. We address this limitation by considering the mutual effect of concurrent applications on system wide energy consumption to adapt resource configuration at run-time. We characterize each application's power-performance profile as a weighted bias through off-line profiling. We infer this model combined with an on-line predictive strategy to make resource allocation decisions for minimizing energy consumption while honoring performance requirements. The proposed strategy is implemented as a user-space process and evaluated on a heterogeneous hardware platform of Odroid XU3 over the Rodinia benchmark suite. Experimental results show up to 61 percent of energy saving compared to the standard baseline of Linux governors and up to 27 percent of energy gain compared to state-of-the-art adaptive learning-based resource management techniques.
Elham Shamsa, Anil Kanduri, Pasi Liljeberg, Amir-Mohammad Rahmani
IEEE Trans. Computers2
2022 Online Learning for Orchestration of Inference in Multi-user End-edge-cloud Networks
abstract
Deep-learning-based intelligent services have become prevalent in cyber-physical applications, including smart cities and health-care. Deploying deep-learning-based intelligence near the end-user enhances privacy protection, responsiveness, and reliability. Resource-constrained end-devices must be carefully managed to meet the latency and energy requirements of computationally intensive deep learning services. Collaborative end-edge-cloud computing for deep learning provides a range of performance and efficiency that can address application requirements through computation offloading. The decision to offload computation is a communication-computation co-optimization problem that varies with both system parameters (e.g., network condition) and workload characteristics (e.g., inputs). However, deep learning model optimization provides another source of tradeoff between latency and model accuracy. An end-to-end decision-making solution that considers such computation-communication problem is required to synergistically find the optimal offloading policy and model for deep learning services. To this end, we propose a reinforcement-learning-based computation offloading solution that learns optimal offloading policy considering deep learning model selection techniques to minimize response time while providing sufficient accuracy. We demonstrate the effectiveness of our solution for edge devices in an end-edge-cloud system and evaluate with a real-setup implementation using multiple AWS and ARM core configurations. Our solution provides 35% speedup in the average response time compared to the state-of-the-art with less than 0.9% accuracy reduction, demonstrating the promise of our online learning framework for orchestrating DL inference in end-edge-cloud systems.
Sina Shahhosseini, Dongjoo Seo, Anil Kanduri, Sung-Soo Lim, Bryan Donyanavard, Amir-Mohammad Rahmani, Nikil Dutt
ACM Trans. Embed. Comput. Syst.3
2021 Energy-Performance Co-Management of Mixed-Sensitivity Workloads on Heterogeneous Multi-core Systems
abstract
Satisfying performance of complex workload scenarios with respect to energy consumption on Heterogeneous Multi-core Platforms (HMPs) is challenging when considering i) the increasing variety of applications, and ii) the large space of resource management configurations. Existing run-time resource management approaches use online and offline learning to handle such complexity. However, they focus on one type of application, neglecting concurrent execution of mixed sensitivity workloads. In this work, we propose an energy-performance co-management method which prioritizes mixed type of applications at run-time, and searches in the configuration space to find the optimal configuration for each application which satisfies the performance requirements while saving energy. We evaluate our approach on a real Odroid XU3 platform over mixed-sensitivity embedded workloads. Experimental results show our approach provides 54% lower performance violation with 50% higher energy saving compared to the existing approaches.
Elham Shamsa, Anil Kanduri, Amir-Mohammad Rahmani, Pasi Liljeberg
ASP-DAC2
2021 UBAR: User- and Battery-aware Resource Management for Smartphones
abstract
Smartphone users require high Battery Cycle Life (BCL) and high Quality of Experience (QoE) during their usage. These two objectives can be conflicting based on the user preference at run-time. Finding the best trade-off between QoE and BCL requires an intelligent resource management approach that considers and learns user preference at run-time. Current approaches focus on one of these two objectives and neglect the other, limiting their efficiency in meeting users’ needs. In this article, we present UBAR, User- and Battery-aware Resource management, which considers dynamic workload, user preference, and user plug-in/out pattern at run-time to provide a suitable trade-off between BCL and QoE. UBAR personalizes this trade-off by learning the user’s habits and using that to satisfy QoE, while considering battery temperature and State of Charge (SOC) pattern to maximize BCL. The evaluation results show that UBAR achieves 10% to 40% improvement compared to the existing state-of-the-art approaches.
Elham Shamsa, Alma Pröbstl, Nima Taherinejad, Anil Kanduri, Samarjit Chakraborty, Amir-Mohammad Rahmani, Pasi Liljeberg
ACM Trans. Embed. Comput. Syst.4
2019 Goal-Driven Autonomy for Efficient On-chip Resource Management: Transforming Objectives to Goals
abstract
Run-time resource allocation of heterogeneous multi-core systems is challenging with varying workloads and limited power and energy budgets. User interaction within these systems changes the performance requirements, often conflicting with concurrent applications' objective and system constraints. Current resource allocation approaches focus on optimizing fixed objective, ignoring the variation in system and applications' objective at run-time. For an efficient resource allocation, the system has to operate autonomously by formulating a hierarchy of goals. We present goal-driven autonomy (GDA) for on-chip resource allocation decisions, which allows systems to generate and prioritize goals in response to the workload and system dynamic variation. We implemented a proof-of-concept resource management framework that integrates the proposed goal management control to meet power, performance and user requirements simultaneously. Experimental results on an Exynos platform containing ARM's big.LITTLE-based heterogeneous multi-processor (HMP) show the effectiveness of GDA in efficient resource allocation in comparison with existing fixed objective policies.
Elham Shamsa, Anil Kanduri, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt
DATE2
2018 Approximation-aware coordinated power/performance management for heterogeneous multi-cores
abstract
Run-time resource management of heterogeneous multi-core systems is challenging due to i) dynamic workloads, that often result in ii) conflicting knob actuation decisions, which potentially iii) compromise on performance for thermal safety. We present a runtime resource management strategy for performance guarantees under power constraints using functionally approximate kernels that exploit accuracy-performance trade-offs within error resilient applications. Our controller integrates approximation with power knobs - DVFS, CPU quota, task migration - in coordinated manner to make performance-aware decisions on power management under variable workloads. Experimental results on Odroid XU3 show the effectiveness of this strategy in meeting performance requirements without power violations compared to existing solutions.
Anil Kanduri, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Cristiana Bolchini, Nikil Dutt
DAC1
2018 Trends in On-chip Dynamic Resource Management
abstract
The Complexity of emerging multi/many-core architectures and diversity of modern workloads demands coordinated dynamic resource management methods. We introduce a classification for these methods capturing the utilized resources and metrics. In this work, we use this classification to survey the key efforts in dynamic resource management. We first cover heuristic and optimization methods used to manage resources such as power, energy, temperature, Quality-of-Service (QoS) and reliability of the system. We then identify some of the machine learning based methods used in tuning architectural parameters in computer systems. In many cases, resource managers need to enforce design constraints during runtime with a certain level of guarantee. Hence, we also study the trend in deploying formal control theoretic approaches in order to achieve efficient and robust dynamic resource management.
Kasra Moazzemi, Anil Kanduri, David Juhasz, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt
DSD2
2018 Approximation for Run-time Power Management
abstract
Performance and energy efficiency of multi-core and many-core systems are restricted by increasing power densities and/or limited energy resources. Maximizing performance while minimizing power and energy consumption becomes challenging with emerging workloads. Approximate computing is an alternative solution that offers the required performance and energy gains, leveraging inherent error resilience of specific application domains. Dynamic power management using approximation as another knob can maximize performance and energy efficiency within fixed power budgets. Disciplined tuning of approximation along with other traditional power knobs requires efficient runtime resource management techniques. We present our strategy for using approximation as another knob for tuning the performance loss incurred in power actuation in many-core systems, which is also portable for heterogeneous multi-core systems.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg
ISCAS1
2018 adBoost: Thermal Aware Performance Boosting Through Dark Silicon Patterning
abstract
Increasing power densities of many-core systems leaves a fraction of on-chip resources inactive, referred to as dark silicon. Efficient management of critical interlinked parameters - power, performance and temperature can improve resource utilization and mitigate dark silicon. In this paper, we present a run-time resource management system for thermal aware performance boosting using a dark silicon aware run-time application mapping strategy. The mapping policy patterns inactive cores among active cores for relatively lower and even distribution of operating temperatures. This provides enough thermal headroom for boosting the frequency of active cores upon performance surges and allows sustained boosting periods, improving the performance further. We design a controller for thermal aware performance boosting that decides on efficient allocation utilization of power budget and thermal headroom obtained from patterning. Our strategy yields up to 37 percent better throughput, 29 percent lower waiting time and up to 2 x longer boosting periods, in comparison with other state-of-the-art run-time mapping policies.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Muhammad Shafique 0001, Axel Jantsch, Pasi Liljeberg
IEEE Trans. Computers1
2017 Accuracy-Aware Power Management for Many-Core Systems Running Error-Resilient Applications
abstract
Power capping techniques based on dynamic voltage and frequency scaling (DVFS) and power gating (PG) are oriented toward power actuation, compromising on performance and energy. Inherent error resilience of emerging application domains, such as Internet-of-Things (IoT) and machine learning, provides opportunities for energy and performance gains. Leveraging accuracy-performance tradeoffs in such applications, we propose approximation (APPX) as another knob for closelooped power management, to complement power knobs with performance and energy gains. We design a power management framework, APPEND+, that can switch between accurate and approximate modes of execution subject to system throughput requirements. APPEND+ considers the sensitivity of the application to error to make disciplined alteration between levels of APPX such that performance is maximized while error is minimized. We implement a power management scheme that uses APPX, DVFS, and PG knobs hierarchically. We evaluated our proposed approach over machine learning and signal processing applications along with two case studies on IoT-early warning score system and fall detection. APPEND+ yields 1.9× higher throughput, improved latency up to five times, better performance per energy, and dark silicon mitigation compared with the state-of-the-art power management techniques over a set of applications ranging from high to no error resilience.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Approximation knob: power capping meets energy efficiency
abstract
Power Capping techniques are used to restrict power consumption of computer systems to a thermally safe limit. Current many-core systems employ dynamic voltage and frequency scaling (DVFS), power gating (PG) and scheduling methods as actuators for power capping. These knobs arc oriented towards power actuation, while the need for performance and energy savings are increasing in the dark silicon era. To address this, we propose approximation (APPX) as another knob for close-looped power management, lending performance and energy efficiency to existing power capping techniques. We use approximation in a pro-active way for long-term performance-energy objectives, complementing the short-term reactive power objectives. We implement an approximation-enabled power management framework, APPEND, that dynamically chooses an application with appropriate level of approximation from a set of variable accuracy implementations. Subject to the system dynamics, our power manager chooses an effective combination of knobs - APPX, DVFS and PG, in a hierarchical way to ensure power capping with performance and energy gains. Our proposed approach yields 1.5× higher throughput, improved latency upto 5×, better performance per energy and dark silicon mitigation compared to state-of-the-art power management techniques over a set of applications ranging from high to no error resilience.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt, Hannu Tenhunen
ICCAD1
2015 Dark silicon aware runtime mapping for many-core systems: A patterning approach
abstract
Limitation on power budget in many-core systems leaves a fraction of on-chip resources inactive, referred to as dark silicon. In such systems, an efficient run-time application mapping approach can considerably enhance resource utilization and mitigate the dark silicon phenomenon. In this paper, we propose a dark silicon aware runtime application mapping approach that patterns active cores alongside the inactive cores in order to evenly distribute power density across the chip. This approach leverages dark silicon to balance the temperature of active cores to provide higher power budget and better resource utilization, within a safe peak operating temperature. In contrast with exhaustive search based mapping approach, our agile heuristic approach has a negligible runtime overhead. Our patterning strategy yields a surplus power budget of up to 17% along with an improved throughput of up to 21% in comparison with other state-of-the-art run-time mapping strategies, while the surplus budget is as high as 40% compared to worst case scenarios.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen
ICCD1
2015 Dynamic power management for many-core platforms in the dark silicon era: A multi-objective control approach
abstract
Power management of NoC-based many-core systems with runtime application mapping becomes more challenging in the dark silicon era. It necessitates a multi-objective control approach to consider an upper limit on total power consumption, dynamic behaviour of workloads, processing elements utilization, per-core power consumption, and load on network-on-chip. In this paper, we propose a multi-objective dynamic power management method that simultaneously considers all of these parameters. Fine-grained voltage and frequency scaling, including near-threshold operation, and per-core power gating are utilized to optimize the performance. In addition, a disturbance rejecter is designed that proactively scales down activity in running applications when a new application commences execution, to prevent sharp power budget violations. Simulations of dynamic workloads and mixed time-critical application profiles show that our method is effective in honoring the power budget while considerably boosting the system throughput and reducing power budget violation, compared to the state-of-the-art power management policies.
Amir-Mohammad Rahmani, M. H. Haghbayan, Anil Kanduri, Awet Yemane Weldezion, Pasi Liljeberg, Juha Plosila, Axel Jantsch, Hannu Tenhunen
ISLPED3
2015 MapPro: Proactive Runtime Mapping for Dynamic Workloads by Quantifying Ripple Effect of Applications on Networks-on-Chip
abstract
Increasing dynamic workloads running on NoC-based many-core systems necessitates efficient runtime mapping strategies. With an unpredictable nature of application profiles, selecting a rational region to map an incoming application is an NP-hard problem in view of minimizing congestion and maximizing performance. In this paper, we propose a proactive region selection strategy which prioritizes nodes that offer lower congestion and dispersion. Our proposed strategy, MapPro, quantitatively represents the propagated impact of spatial availability and dispersion on the network with every new mapped application. This allows us to identify a suitable region to accommodate an incoming application that results in minimal congestion and dispersion. We cluster the network into squares of different radii to suit applications of different sizes and proactively select a suitable square for a new application, eliminating the overhead caused with typical reactive mapping approaches. We evaluated our proposed strategy over different traffic patterns and observed gains of up to 41% in energy efficiency, 28% in congestion and 21% dispersion when compared to the state-of-the-art region selection methods.
M. H. Haghbayan, Anil Kanduri, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen
NOCS2
2014 Mixed-Criticality Run-Time Task Mapping for NoC-Based Many-Core Systems
abstract
Contiguous processor allocation improves both the network and the application performance, by decreasing the congestion probability among communication of different applications. Consequently, the average, standard deviation and worst-case latency of the network is decreased significantly. This makes the contiguous allocation a good solution for time-critical applications with bounded deadlines. On the other hand, non-contiguous allocation will increase the system throughput significantly. Isolated nodes are utilized and more applications can finish their job in a time unit. However, this will lead to poor network metrics, unsuitable for real-time applications. In this work, we combine these two approaches in order to manage workloads with mixed-critical characteristics. Real-time applications are mapped contiguously, while non-critical applications are allowed to get dispersed over the available system nodes. Results show over 50% improvement in worst-case latency and 100 times improvement in deadline misses.
Mohammad Fattah, Amir-Mohammad Rahmani, Thomas Canhao Xu, Anil Kanduri, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
PDP4