M. H. Haghbayan

dblp:24/9056 · also M. Hashem Haghbayan, Mohammad Hashem Haghbayan · DBLP profile ↗
← Back
38ranked-venue papers
12as first author
14since 2021 · last 2026
0000-0001-6583-4418ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 12 first-author · 5 since 2021Artificial intelligence and machine learning · 12 · 9 since 2021Software engineering, systems software and programming languages · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive Simulation Platform In NVIDIA Isaac Sim For Self-Aware Heterogeneous Robots
abstract
This paper presents an adaptive simulation framework, implemented in NVIDIA Isaac Sim, for self-aware heterogeneous robotic systems. The proposed platform enables each robot to maintain an internal predictive model of its own embodiment and its interaction with the surrounding environment, while explicitly accounting for both mechanical actuation and onboard computational processes. By treating mechanical and computational energy as internal state variables, the framework enables runtime estimation of the energetic consequences of robot actions. This capability is demonstrated across diverse robot morphologies, including a mobile humanoid, a fixed-base manipulator, and an aerial robot operating in a shared warehouse environment. Experimental results show that internal computational processes constitute a substantial portion of the overall energy budget and are tightly coupled with mechanical action, highlighting the importance of incorporating internal resource state into self-regulatory decision-making. The framework provides a reusable platform for studying resource-aware self-regulation and energy-informed decision-making, with extensions toward multi-robot and swarm systems.
Afrooz Naseri, Juha Plosila, M. H. Haghbayan
ECMS3
2025 Simulation Of Mechanical And Computational Power Consumption In Mobile Robots
abstract
This paper presents a simulation model to estimate the instantaneous power consumption of a mobile robot by taking into account both its mechanical and computational components. The simulation model is adaptable to be tuned based on the level of accuracy needed for estimating the power consumption for the robot and the simulation time penalty. This makes a multi-fidelity power estimation tool for the robot with the capability to run-time changing the fidelity according to environmental conditions and internal computational capabilities. Such multi-fidelity energy prediction is suitable for run-time predictive decision making in a wide range of usages such as training process in model-based Reinforcement Learning (RL) as well as decision making in Model Predictive Control (MPC). The experimental results show that the simulation accurately estimates energy consumption at different fidelity levels. Higher-fidelity models closely match real-world measurements, while lower-fidelity models trade some accuracy for faster predictions. Therefore, higher estimation precision comes at the cost of increased computation.
Afrooz Naseri, Sajad Shahsavari, Juha Plosila, M. H. Haghbayan
ECMS4
2025 An Open-Source Framework For CFD-Based Digital Twins: A Case Study On Storm Water Management
abstract
Digital Twins (DTs) are increasingly applied for optimization of operations in logistics, healthcare, smart cities, and beyond. However, implementing high-fidelity DTs remains challenging in computationally intensive domains such as Computational Fluid Dynamics (CFD). While simplified models can facilitate real-time operation, they often lack physical fidelity. This article presents an open-source scalable software framework along with a case study of CFD-based digital twining on stormwater management. The presented framework enables online execution of CFD-based models by containerizing and integrating them into OpenShift platform, providing a two-way communication channel for simulation parameters and results. The framework is capable of dynamically scaling computing resources to run computationally-intensive CFD-models. In the case study, we present a novel CFD simulation model of a bioretention cell intended to reduce runoff volumes of urban stormwater. The simulation model, implemented in OpenFOAM, is then integrated into the presented software framework to create the DT. The framework source code, simulation model and the DT are made publicly available to promote future research.
Sajad Shahsavari, Ashvinkumar Chaudhari, Eero Immonen, M. H. Haghbayan
ECMS4
2025 Runtime Energy-Efficient Control Policy for Mobile Robots with Computing Workload and Battery Awareness
abstract
Energy efficiency is a fundamental goal in robotic control. Various components within a robot, such as mechanical systems, computational units, and sensors, consume energy, all powered by the battery unit. Each component features several actuators and individual controllers that optimize energy usage locally, often without regard to one another. In this paper, we highlight a significant phenomenon indicating a considerable dependency between the mechanical and computational parts of the robot as energy consumers and the battery state of charge (SOC) as the energy provider. We demonstrate that as the battery SOC fluctuates, the behavior of energy consumption also varies, necessitating a unified controller with awareness of this relationship. Motivated by this observation, we propose a battery-aware co-optimization strategy for the mechanical and computational units, leveraging configuration space exploration to optimize the motor speed and the CPU frequency under different environmental conditions and battery SOC levels. Experimental results demonstrate the effectiveness of our approach in extending the operational lifetime of a robot under varying battery SOC and workload conditions, enhancing the energy efficiency of a case study rover by up to 53.93% w.r.t. selected baselines and similar past approaches.
M. H. Haghbayan, Abdul Malik, Antonio Miele, Juha Plosila
IROS2
2025 A Coordinated Approach to Control Mechanical and Computing Resources in Mobile Robots
abstract
Energy management of mechanical and cyber parts in mobile robots consists of two processes operating concurrently at runtime. Both the two processes can significantly improve the robots' battery lifetime and further extend mission time. In each process, information on energy consumption of one of the two parts is captured and analyzed to manipulate various mechanical/computational actuators in a robot, such as motor speed and CPU voltage/frequency. In this article, we show that considering management of mechanical and computational segments separately does not necessarily result in an energy-optimal solution due to their co-dependence; as a consequence, a runtime co-management scheme is required. We propose a proactive energy optimization methodology in which dynamically trained internal models are utilized to predict the future energy consumption for the mechanical and computational parts of a mobile robot, and based on that, the optimal mechanical speed and CPU voltage/frequency are determined at runtime. The experimental results on a ground wheeled robot show up to 36.34% reduction in the overall energy consumption compared to the state-of-the-art methods.
Sajad Shahsavari, M. H. Haghbayan, Antonio Miele, Eero Immonen, Juha Plosila
IEEE Trans. Robotics2
2023 A Coupled Battery State-of-Charge and Voltage Model for Optimal Control Applications
abstract
Optimal control of electric vehicle (EV) batteries for maximal energy efficiency, safety and lifespan requires that the Battery Management System (BMS) has accurate real-time information on both the battery State-of-Charge (SoC) and its dynamics, i.e. long-term and short-term energy supply capacity, at all times. However, these quantities cannot be measured directly from the battery, and, in practice, only SoC estimation is typically carried out. In this article, we propose a novel parametric algebraic voltage model coupled to the well-known Manwell-McGowan dynamic Kinetic Battery Model (KiBaM), which is able to predict both battery SoC dynamics and its electrical response. Numerical simulations, based on laboratory measurements, are presented for prismatic Lithium-Titanate Oxide (LTO) battery cells. Such cells are prime candidates for modern heavy offroad EV applications.
Masoomeh Karami, Sajad Shahsavari, Eero Immonen, M. H. Haghbayan, Juha Plosila
DATE4
2023 A Light-Weight Model For Run-Time Battery SOC-SOH Estimation While Considering Aging
abstract
Batteries are becoming one important part to power varieties of devices including electro-mechanical robots and vehicles. Understanding the behaviour of the battery and its state of charge can help the control systems to significantly improve the decision-making and risk management at run-time, after the device starts its operation. Currently, there is an increased interest in tracking battery dynamics as a function of health in both academia and industry. In this paper, we propose a light-weight approach for modeling the state of charge of lithium-ion (Li-ion) batteries during the life-time of the system. We also consider the battery capacity of charge degradation over its usage. To do that, we use electrical equivalent circuit model (EECM) modeling as the basis for modeling the battery and add the aging model to it to consider the effect of battery usage in the long term. Experimental results show that our proposed technique successfully estimates the battery state of charge at different states of health for the National Aeronautics and Space Administration (NASA) randomized usage battery dataset in comparison with the state-of-the-art. The obtained estimation error in the worst case is 2.2%.
Mohsen Heydarzadeh, Eero Immonen, M. H. Haghbayan, Juha Plosila
ECMS3
2023 Run-Time Resource Management in CMPs Handling Multiple Aging Mechanisms
abstract
Run-time resource management is fundamental for efficient execution of workloads on Chip Multiprocessors. Application- and system-level requirements (e.g., on performance versus power versus lifetime reliability) are generally conflicting each other, and any decision on resource assignment, such as core allocation or frequency tuning, may positively affect some of them while penalizing some others. Resource assignment decisions can be perceived in few instants of time on performance and power consumption, but not on lifetime reliability. In fact, this latter changes very slowly based on the accumulation of effects of various decisions over a long time horizon. Moreover, aging mechanisms are various and have different causes; most of them, such as Electromigration (EM), are subject to temperature levels, while Thermal Cycling (TC) is caused mainly by temperature variations (both amplitude and frequency). Mitigating only EM may negatively affect TC and vice versa. We propose a resource orchestration strategy to balance the performance and power consumption constraints in the short-term and EM and TC aging in the long-term. Experimental results show that the proposed approach improves the average Mean Time To Failure at least by 17% and 20% w.r.t. EM and TC, respectively, while providing same performance level of the nominal counterpart and guaranteeing the power budget.
M. H. Haghbayan, Antonio Miele, Onur Mutlu, Juha Plosila
IEEE Trans. Computers1
2022 How To Run A World Record? A Reinforcement Learning Approach
abstract
Finding the optimal distribution of exerted effort by an athlete in competitive sports has been widely investigated in the fields of sport science, applied mathematics and optimal control. In this article, we propose a reinforcement learning-based solution to the optimal control problem in the running race application. Well-known mathematical model of Keller is used for numerically simulating the dynamics in runner's energy storage and motion. A feed-forward neural network is employed as the probabilistic controller model in continuous action space which transforms the current state (position, velocity and available energy) of the runner to the predicted optimal propulsive force that the runner should apply in the next time step. A logarithmic barrier reward function is designed to evaluate performance of simulated races as a continuous smooth function of runner's position and time. The neural network parameters, then, are identified by maximizing the expected reward using on-policy actor-critic policy-gradient RL algorithm. We trained the controller model for three race lengths: 400, 1500 and 10000 meters and found the force and velocity profiles that produce a near-optimal solution for the runner's problem. Results conform with Keller's theoretical findings with relative percent error of 0.59% and are comparable to real world records with relative percent error of 2.38%, while the same error for Keller's findings is 2.82%.
Sajad Shahsavari, Eero Immonen, Masoomeh Karami, M. H. Haghbayan, Juha Plosila
ECMS4
2021 Capacity Loss Estimation For Li-Ion Batteries Based On A Semi-Empirical Model
abstract
Understanding battery capacity degradation is instrumental for designing modern electric vehicles. In this paper, a Semi-Empirical Model for predicting the Capacity Loss of Lithium-ion batteries during Cycling and Calendar Aging is developed. In order to redict the Capacity Loss with a high accuracy, battery operation data from different test conditions and different Lithium-ion batteries chemistries were obtained from literature for parameter optimization (fitting). The obtained models were then compared to experimental data for validation. Our results show that the average error between the estimated Capacity Loss and measured Capacity Loss is less than 1.5% during Cycling Aging, and less than 2% during Calendar Aging. An electric mining dumper, with simulated duty cycle data, is considered as an application example.
Mohammed Rabah, Eero Immonen, Sajad Shahsavari, M. H. Haghbayan, Kirill Murashko, Paula Immonen
ECMS4
2021 MCX ? An Open-Source Framework For Digital Twins
abstract
This article describes ModelConductor-eXtended (MCX), which is an open-source software architecture for digital twins. The MCX framework facilitates co-execution of, and asynchronous data communication between, physical systems and their digital simulation models. MCX supports running FMUs (simulation models packaged according to the FMI specification) as well as machine learning models and customized models. We propose extensions to the previously published ModelConductor framework for higher performance and better scalability. The extensions include decoupling of the queue and the model computation module, utilization of a standard data transmission protocol and implementation of the facility to run time-consuming simulation models in a time synchronous manner. Additionally, three new validation case studies are presented. A performance evaluation shows that the extensions improve the average response time almost 4 times in three specific experiments.
Sajad Shahsavari, Eero Immonen, Mohammed Rabah, M. H. Haghbayan, Juha Plosila
ECMS4
2021 Hierarchical Fault Simulation of Deep Neural Networks on Multi-Core Systems
abstract
In this paper, a hierarchical fault simulation technique for neural networks is proposed, supporting both permanent and temporary faults. In the proposed technique, different levels of hierarchy are used, forming a mixed-level simulation environment. In such an environment, the pre-synthesis behavioral specification of the network and the post-synthesis gate-level model are co-simulated. To accelerate the fault simulation process, faults are injected in the gate-level specification of the selected neurons while the behavioral model in different levels of abstraction is used to simulate the remaining neurons. Further speedup is obtained through event-driven simulation and parallelization. Experimental results confirm the time efficiency of the proposed fault simulation technique.
Masoomeh Karami, M. H. Haghbayan, Masoumeh Ebrahimi, Antonio Miele, Hannu Tenhunen, Juha Plosila
ETS2
2021 Energy-Efficient Mobile Robot Control via Run-time Monitoring of Environmental Complexity and Computing Workload
abstract
We propose an energy-efficient controller to minimize the energy consumption of a mobile robot by dynamically manipulating the mechanical and computational actuators of the robot. The mobile robot performs real-time vision-based applications based on an event-based camera. The actuators of the controller are CPU voltage/frequency for the computation part and motor voltage for the mechanical part. We show that independently considering speed control of the robot and voltage/frequency control of the CPU does not necessarily result in an energy-efficient solution. In fact, to obtain the highest efficiency, the computation and mechanical parts should be controlled together in synergy. We propose a fast hill-climbing optimization algorithm to allow the controller to find the best CPU/motor configuration at run-time and whenever the mobile robot is facing a new environment during its travel. Experimental results on a robot with Brushless DC Motors, Jetson TX2 board as the computing unit, and a DAVIS-346 event-based camera show that the proposed control algorithm can save battery energy by an average of 50.5%, 41%, and 30%, in low-complexity, medium-complexity, and high-complexity environments, over baselines.
Sherif Abdelmonem Sayed Mohamed, M. H. Haghbayan, Antonio Miele, Onur Mutlu, Juha Plosila
IROS2
2021 High-Performance Parallel Fault Simulation for Multi-Core Systems
abstract
Fault simulation is a time-consuming process that requires customized methods and techniques to accelerate it. Multi-threading and Multi-core approaches are two promising techniques that can be exploited to accelerate the fault simulation process by using different parts of the hardware at the same time. However, an efficient parallelization is obtained only by the refinement of software with respect to the hardware platform. In this paper, a parallel multi-thread fault simulation technique is proposed to accelerate the simulation process on multi-core platforms. In this approach, the gate input values are independently assigned to each thread. Each input value carries the information of several parallel simulation processes. This provides a multithread parallel fault simulation environment. The experimental results show that the proposed technique can efficiently use the hardware platform. In a single-core platform, the proposed technique can reduce the time by 25% while in a dual-core increasing the thread approximately halves the execution time.
Masoomeh Karami, M. H. Haghbayan, Masoumeh Ebrahimi, Hamid Nejatollahi, Hannu Tenhunen, Juha Plosila
PDP2
2020 Thermal-Cycling-aware Dynamic Reliability Management in Many-Core System-on-Chip
abstract
Dynamic Reliability Management (DRM) is a common approach to mitigate aging and wear-out effects in multi- /many-core systems. State-of-the-art DRM approaches apply finegrained control on resource management to increase/balance the chip reliability while considering other system constraints, e.g., performance, and power budget. Such approaches, acting on various knobs such as workload mapping and scheduling, Dynamic Voltage/Frequency Scaling (DVFS) and Per-Core Power Gating (PCPG), demonstrated to work properly with the various aging mechanisms, such as electromigration, and Negative-Bias Temperature Instability (NBTI). However, we claim that they do not suffice for thermal cycling. Thus, we here propose a novel thermal-cycling-aware DRM approach for shared-memory many-core systems running multi-threaded applications. The approach applies a fine-grained control capable at reducing both temperature levels and variations. The experimental evaluations demonstrated that the proposed approach is able to achieve 39% longer lifetime than past approaches.
M. H. Haghbayan, Antonio Miele, Zhuo Zou, Hannu Tenhunen, Juha Plosila
DATE1
2020 Navigation System For Landing A Swarm Of Autonomous Drones On A Movable Surface
Anam Tahir, Jari M. Böling, M. H. Haghbayan, Juha Plosila
ECMS3
2020 Dynamic Resource-Aware Corner Detection for Bio-Inspired Vision Sensors
abstract
Event-based cameras are vision devices that transmit only brightness changes with low latency and ultra-low power consumption. Such characteristics make event-based cameras attractive in the field of localization and object tracking in resource-constrained systems. Since the number of generated events in such cameras is huge, the selection and filtering of the incoming events are beneficial from both increasing the accuracy of the features and reducing the computational load. In this paper, we present an algorithm to detect asynchronous corners form a stream of events in real-time on embedded systems. The algorithm is called the Three Layer Filtering-Harris or TLF-Harris algorithm. The algorithm is based on an events' filtering strategy whose purpose is 1) to increase the accuracy by deliberately eliminating some incoming events, i.e., noise and 2) to improve the real-time performance of the system, i.e., preserving a constant throughput in terms of input events per second, by discarding unnecessary events with a limited accuracy loss. An approximation of the Harris algorithm, in turn, is used to exploit its high-quality detection capability with a low-complexity implementation to enable seamless real-time performance on embedded computing platforms. The proposed algorithm is capable of selecting the best corner candidate among neighbors and achieves an average execution time savings of 59% compared with the conventional Harris score. Moreover, our approach outperforms the competing methods, such as eFAST, eHarris, and FA-Harris, in terms of real-time performance, and surpasses Arc* in terms of accuracy.
Sherif Abdelmonem Sayed Mohamed, Jawad Naveed Yasin, M. H. Haghbayan, Antonio Miele, Jukka Heikkonen, Hannu Tenhunen, Juha Plosila
ICPR3
2019 Monocular visual odometry based on hybrid parameterization
abstract
Visual odometry (VO) is one of the most challenging techniques in computer vision for autonomous vehicle/vessels. In VO, the camera pose that also represents the robot pose in ego-motion is estimated analyzing the features and pixels extracted from the camera images. Different VO techniques mainly provide different trade-offs among the resources that are being considered for odometry, such as camera resolution, computation/communication capacity, power/energy consumption, and accuracy. In this paper, a hybrid technique is proposed for camera pose estimation by combining odometry based on triangulation using the long-term period of direct-based odometry and the short-term period of inverse depth mapping. Experimental results based on the EuRoC data set shows that the proposed technique significantly outperforms the traditional direct-based pose estimation method for Micro Aerial Vehicle (MAV), keeping its potential negative effect on performance negligible.
Sherif Abdelmonem Sayed Mohamed, M. H. Haghbayan, Jukka Heikkonen, Hannu Tenhunen, Juha Plosila
ICMV2
2018 Object Detection Based on Multi-sensor Proposal Fusion in Maritime Environment
abstract
In this paper, we propose an effective object detection framework based on proposal fusion of multiple sensors such as infrared camera, RGB cameras, radar and LiDAR. Our framework first applies the Selective Search (SS) method on RGB image data to extract possible candidate proposals which likely contain the objects of interest. Then it uses the information from other sensors in order to reduce the number of generated proposals by SS and find more dense proposals. Finally, the class of objects within the final proposals are identified by Convolutional Neural Network (CNN). Experimental results on real dataset demonstrate that our framework can precisely detect meaningful object regions using a smaller number of proposals than other object proposals methods. Further, our framework can achieve reliable object detection and classification results in maritime environments.
Fahimeh Farahnakian, M. H. Haghbayan, Jonne Poikonen, Markus Laurinen, Paavo Nevalainen, Jukka Heikkonen
ICMLA2
2018 Approximation for Run-time Power Management
abstract
Performance and energy efficiency of multi-core and many-core systems are restricted by increasing power densities and/or limited energy resources. Maximizing performance while minimizing power and energy consumption becomes challenging with emerging workloads. Approximate computing is an alternative solution that offers the required performance and energy gains, leveraging inherent error resilience of specific application domains. Dynamic power management using approximation as another knob can maximize performance and energy efficiency within fixed power budgets. Disciplined tuning of approximation along with other traditional power knobs requires efficient runtime resource management techniques. We present our strategy for using approximation as another knob for tuning the performance loss incurred in power actuation in many-core systems, which is also portable for heterogeneous multi-core systems.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg
ISCAS2
2018 adBoost: Thermal Aware Performance Boosting Through Dark Silicon Patterning
abstract
Increasing power densities of many-core systems leaves a fraction of on-chip resources inactive, referred to as dark silicon. Efficient management of critical interlinked parameters - power, performance and temperature can improve resource utilization and mitigate dark silicon. In this paper, we present a run-time resource management system for thermal aware performance boosting using a dark silicon aware run-time application mapping strategy. The mapping policy patterns inactive cores among active cores for relatively lower and even distribution of operating temperatures. This provides enough thermal headroom for boosting the frequency of active cores upon performance surges and allows sustained boosting periods, improving the performance further. We design a controller for thermal aware performance boosting that decides on efficient allocation utilization of power budget and thermal headroom obtained from patterning. Our strategy yields up to 37 percent better throughput, 29 percent lower waiting time and up to 2 x longer boosting periods, in comparison with other state-of-the-art run-time mapping policies.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Muhammad Shafique 0001, Axel Jantsch, Pasi Liljeberg
IEEE Trans. Computers2
2017 Performance/Reliability-Aware Resource Management for Many-Cores in Dark Silicon Era
abstract
Aggressive technology scaling has enabled the fabrication of many-core architectures while triggering challenges such as limited power budget and increased reliability issues, like aging phenomena. Dynamic power management and runtime mapping strategies can be utilized in such systems to achieve optimal performance while satisfying power constraints. However, lifetime reliability is generally neglected. We propose a novel lifetime reliability/performance-aware resource co-management approach for many-core architectures in the dark silicon era. The approach is based on a two-layered architecture, composed of a long-term runtime reliability controller and a short-term runtime mapping and resource management unit. The former evaluates the cores' aging status w.r.t. a target reference specified by the designer, and performs recovery actions on highly stressed cores by means of power capping. The aging status is utilized in runtime application mapping to maximize system performance while fulfilling reliability requirements and honoring the power budget. Experimental evaluation demonstrates the effectiveness of the proposed strategy, which outperforms most recent state-of-the-art contributions.
M. H. Haghbayan, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen
IEEE Trans. Computers1
2017 Accuracy-Aware Power Management for Many-Core Systems Running Error-Resilient Applications
abstract
Power capping techniques based on dynamic voltage and frequency scaling (DVFS) and power gating (PG) are oriented toward power actuation, compromising on performance and energy. Inherent error resilience of emerging application domains, such as Internet-of-Things (IoT) and machine learning, provides opportunities for energy and performance gains. Leveraging accuracy-performance tradeoffs in such applications, we propose approximation (APPX) as another knob for closelooped power management, to complement power knobs with performance and energy gains. We design a power management framework, APPEND+, that can switch between accurate and approximate modes of execution subject to system throughput requirements. APPEND+ considers the sensitivity of the application to error to make disciplined alteration between levels of APPX such that performance is maximized while error is minimized. We implement a power management scheme that uses APPX, DVFS, and PG knobs hierarchically. We evaluated our proposed approach over machine learning and signal processing applications along with two case studies on IoT-early warning score system and fall detection. APPEND+ yields 1.9× higher throughput, improved latency up to five times, better performance per energy, and dark silicon mitigation compared with the state-of-the-art power management techniques over a set of applications ranging from high to no error resilience.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Reliability-Aware Runtime Power Management for Many-Core Systems in the Dark Silicon Era
abstract
Power management of networked many-core systems with runtime application mapping becomes more challenging in the dark silicon era. It necessitates considering network characteristics at runtime to achieve better performance while honoring the peak power upper bound. On the other hand, power management has a direct effect on chip temperature, which is the main driver of the aging effects. Therefore, alongside performance fulfillment, the controlling mechanism must also consider the current cores' reliability in its actuator manipulation to enhance the overall system lifetime in the long term. In this paper, we propose a multiobjective dynamic power management technique that uses current power consumption and other network characteristics including the reliability of the cores as the feedback while utilizing fine-grained voltage and frequency scaling and per-core power gating as the actuators. In addition, disturbance rejecter and reliability balancer are designed to help the controller to better smooth power consumption in the short term and reliability in the long term, respectively. Simulations of dynamic workloads and mixed criticality application profiles show that our method not only is effective in honoring the power budget while considerably boosting the system throughput, but also increases the overall system lifetime by minimizing aging effects by means of power consumption balancing.
Amir-Mohammad Rahmani, M. H. Haghbayan, Antonio Miele, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.2
2016 A lifetime-aware runtime mapping approach for many-core systems in the dark silicon era
M. H. Haghbayan, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen
DATE1
2016 Approximation knob: power capping meets energy efficiency
abstract
Power Capping techniques are used to restrict power consumption of computer systems to a thermally safe limit. Current many-core systems employ dynamic voltage and frequency scaling (DVFS), power gating (PG) and scheduling methods as actuators for power capping. These knobs arc oriented towards power actuation, while the need for performance and energy savings are increasing in the dark silicon era. To address this, we propose approximation (APPX) as another knob for close-looped power management, lending performance and energy efficiency to existing power capping techniques. We use approximation in a pro-active way for long-term performance-energy objectives, complementing the short-term reactive power objectives. We implement an approximation-enabled power management framework, APPEND, that dynamically chooses an application with appropriate level of approximation from a set of variable accuracy implementations. Subject to the system dynamics, our power manager chooses an effective combination of knobs - APPX, DVFS and PG, in a hierarchical way to ensure power capping with performance and energy gains. Our proposed approach yields 1.5× higher throughput, improved latency upto 5×, better performance per energy and dark silicon mitigation compared to state-of-the-art power management techniques over a set of applications ranging from high to no error resilience.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt, Hannu Tenhunen
ICCAD2
2016 A dynamic specification to automatically debug and correct various divider circuits
M. H. Haghbayan, Bijan Alizadeh
Integr.1
2016 A Power-Aware Approach for Online Test Scheduling in Many-Core Architectures
abstract
Aggressive technology scaling triggers novel challenges to the design of multi-/many-core systems, such as limited power budget and increased reliability issues. Today's many-core systems employ dynamic power management and runtime mapping strategies trying to offer optimal performance while fulfilling power constraints. On the other hand, due to the reliability challenges, online testing techniques are becoming a necessity in current and near future technologies. However, state-of-the-art techniques are not aware of the other power/performance requirements. This paper proposes a power-aware non-intrusive online testing approach for many-core systems. The approach schedules software based self-test routines on the various cores during their idle periods, while honoring the power budget and limiting delays in the workload execution. A test criticality metric, based on a device aging model, is used to select cores to be tested at a time. Moreover, power and reliability issues related to the testing at different voltage and frequency levels are also handled. Extensive experimental results reveal that the proposed approach can i) efficiently test the cores within the available power budget causing a negligible performance penalty, ii) adapt the test frequency to the current cores' aging status, and iii) cover available voltage and frequency levels during the testing.
M. H. Haghbayan, Amir-Mohammad Rahmani, Antonio Miele, Mohammad Fattah, Juha Plosila, Pasi Liljeberg, Hannu Tenhunen
IEEE Trans. Computers1
2015 Power-aware online testing of manycore systems in the dark silicon era
M. H. Haghbayan, Amir-Mohammad Rahmani, Mohammad Fattah, Pasi Liljeberg, Juha Plosila, Zainalabedin Navabi, Hannu Tenhunen
DATE1
2015 Dark silicon aware runtime mapping for many-core systems: A patterning approach
abstract
Limitation on power budget in many-core systems leaves a fraction of on-chip resources inactive, referred to as dark silicon. In such systems, an efficient run-time application mapping approach can considerably enhance resource utilization and mitigate the dark silicon phenomenon. In this paper, we propose a dark silicon aware runtime application mapping approach that patterns active cores alongside the inactive cores in order to evenly distribute power density across the chip. This approach leverages dark silicon to balance the temperature of active cores to provide higher power budget and better resource utilization, within a safe peak operating temperature. In contrast with exhaustive search based mapping approach, our agile heuristic approach has a negligible runtime overhead. Our patterning strategy yields a surplus power budget of up to 17% along with an improved throughput of up to 21% in comparison with other state-of-the-art run-time mapping strategies, while the surplus budget is as high as 40% compared to worst case scenarios.
Anil Kanduri, M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen
ICCD2
2015 Dynamic power management for many-core platforms in the dark silicon era: A multi-objective control approach
abstract
Power management of NoC-based many-core systems with runtime application mapping becomes more challenging in the dark silicon era. It necessitates a multi-objective control approach to consider an upper limit on total power consumption, dynamic behaviour of workloads, processing elements utilization, per-core power consumption, and load on network-on-chip. In this paper, we propose a multi-objective dynamic power management method that simultaneously considers all of these parameters. Fine-grained voltage and frequency scaling, including near-threshold operation, and per-core power gating are utilized to optimize the performance. In addition, a disturbance rejecter is designed that proactively scales down activity in running applications when a new application commences execution, to prevent sharp power budget violations. Simulations of dynamic workloads and mixed time-critical application profiles show that our method is effective in honoring the power budget while considerably boosting the system throughput and reducing power budget violation, compared to the state-of-the-art power management policies.
Amir-Mohammad Rahmani, M. H. Haghbayan, Anil Kanduri, Awet Yemane Weldezion, Pasi Liljeberg, Juha Plosila, Axel Jantsch, Hannu Tenhunen
ISLPED2
2015 MapPro: Proactive Runtime Mapping for Dynamic Workloads by Quantifying Ripple Effect of Applications on Networks-on-Chip
abstract
Increasing dynamic workloads running on NoC-based many-core systems necessitates efficient runtime mapping strategies. With an unpredictable nature of application profiles, selecting a rational region to map an incoming application is an NP-hard problem in view of minimizing congestion and maximizing performance. In this paper, we propose a proactive region selection strategy which prioritizes nodes that offer lower congestion and dispersion. Our proposed strategy, MapPro, quantitatively represents the propagated impact of spatial availability and dispersion on the network with every new mapped application. This allows us to identify a suitable region to accommodate an incoming application that results in minimal congestion and dispersion. We cluster the network into squares of different radii to suit applications of different sizes and proactively select a suitable square for a new application, eliminating the overhead caused with typical reactive mapping approaches. We evaluated our proposed strategy over different traffic patterns and observed gains of up to 41% in energy efficiency, 28% in congestion and 21% dispersion when compared to the state-of-the-art region selection methods.
M. H. Haghbayan, Anil Kanduri, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen
NOCS1
2014 Online testing of many-core systems in the Dark Silicon era
abstract
As the dark silicon era is about to embrace, it is not anymore possible to attain commensurate performance benefits by increasing the number of transistors due to thermal design power. Dark Silicon issue stresses that a fraction of silicon chip being able to switch in full frequency is dropping and designers will soon face the growing underutilization inherent in future technologies. On the other hand, by reducing the transistor size, susceptibility to internal defects drastically increases and large ranges of defects such as aging or transient faults will be shown up more frequently. In this paper, we propose an online test scheduling algorithm using software based self-test for dark silicon era to test dark cores while considering thermal design power of the system. As the dark area of the system is dynamic and reshapes at a runtime, the tested cores can be used by other applications in the near future. Empirical results show the effectiveness of the proposed algorithm in terms of applicability and fault coverage with a negligible negative impact on the system throughput.
M. H. Haghbayan, Amir-Mohammad Rahmani, Pasi Liljeberg, Juha Plosila, Hannu Tenhunen
DDECS1
2014 Dark silicon aware power management for manycore systems under dynamic workloads
abstract
Dark Silicon denotes the phenomenon that, due to thermal and power constraints, the fraction of transistors that can operate at full frequency is decreasing with each technology generation. We propose a PID (Proportional Integral Derivative) controller based dynamic power management method that considers an upper bound on power consumption (called the Thermal Design Power (TDP)). To avoid violation of the TDP constraint for manycore systems running highly dynamic workloads, it provides fine-grained DVFS (Dynamic Voltage and Frequency Scaling) including near-threshold operation. In addition, the method distinguishes applications with hard Real-Time, soft Real-Time and no Real-Time constraints and treats them with appropriate priorities. In simulations with dynamic workloads mixed-critical application profiles, we show that the method is effective in honoring the TDP bound and it can boost system throughput by over 43% compared to a naive TDP scheduling policy.
M. H. Haghbayan, Amir-Mohammad Rahmani, Awet Yemane Weldezion, Pasi Liljeberg, Juha Plosila, Axel Jantsch, Hannu Tenhunen
ICCD1
2013 Graph based fault model definition for bus testing
abstract
In this paper we present a new fault model for testing standard On-Chip buses using a graph model. This method will be optimized for speed of testing. Using AMBA-AHB as the experimental result, the proposed fault model shows efficiency in comparison with corresponding stuck-at fault testing.
Elmira Karimi, M. H. Haghbayan, Adele Maleki, Mahmoud Tabandeh
VLSI-SoC2
2012 Power constraint testing for multi-clock domain SoCs using concurrent hybrid BIST
abstract
This paper presents a novel approach for selecting optimal pseudo random and deterministic test patterns and minimizing test time for multi-clock domain SoCs based on a hybrid BIST architecture for each core. For test scheduling, a concurrent method considering peak power upper bound is used. A test scheduling graph is presented for modeling concurrent hybrid BIST test scheduling. Furthermore, a heuristic is proposed for selecting cores to be tested concurrently and the order of applying sequence of test patterns to each core. Experimental results show that the proposed heuristics for both selecting groups of cores to be tested concurrently during the SoC test process, and determining the amount of deterministic and pseudo random test patterns for each core, give us an optimized method for multi clock domain SoC testing compared with the existing methods.
M. H. Haghbayan, Saeed Safari, Zainalabedin Navabi
DDECS1
2011 Online Test Macro Scheduling and Assignment in MPSoC Design
abstract
Due to unreliability of the cores in embedded systems in deep sub-micron technologies, a method for testing cores in the field is needed. In this paper an online method for testing cores of embedded designs is presented. The proposed task scheduling method runs the test routine ASAP periodically considering the real time constraints. A software test routine based on a proposed method will be generated and a task scheduling process including the test task (for each core) and other existing applications of the embedded system will be presented. A software based checksum is issued for online test result analysis that shortens the memory usage of the test process. Experimental results show that this method improves the test application time (TAT) and fault coverage (in proportion to TAT) as compared with the existing methods.
Behnam Khodabandeloo, Seyyed Alireza Hoseini, Sajjad Taheri, M. H. Haghbayan, Mahmood Reza Babaei, Zainalabedin Navabi
Asian Test Symposium4
2010 Test Pattern Selection and Compaction for Sequential Circuits in an HDL Environment
abstract
In this paper we are revisiting the issue of sequential circuit test generation, and use a selective random pattern test generation method implemented in an HDL environment. The method uses a statistical expectation graph and states of the sequential circuit for selecting the appropriate test vectors to achieve better fault coverage and a more compact test set. To further reduce the size of the generated test set, a static compaction method, which is also implemented in an HDL environment, is used after the test generation process. The experimental results show that selecting good test patterns among random test patterns, not only can be implemented dynamically in an HDL design environment, but also results in a better fault coverage and shorter test pattern length in comparison with some traditional deterministic methods. In addition, it will be shown that static test set compaction methods can considerably reduce the test length of test patterns for sequential designs obtained by our proposed method.
M. H. Haghbayan, Sara Karamati, Fatemeh Javaheri, Zainalabedin Navabi
Asian Test Symposium1