Mohsen Ansari

dblp:174/1559 · DBLP profile ↗
← Back
25ranked-venue papers
8as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 7 first-author · 10 since 2021Computer networks · 7 · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 GLEAM: A Graph-Learning Enhanced Adaptive Metaheuristic for Power-Aware Scheduling on Heterogeneous Cyber-Physical Systems
abstract
The increasing complexity of embedded and Cyber-Physical Systems (CPS) has accelerated the adoption of heterogeneous multi-core architectures, which combine performance and energy efficiency. However, scheduling dependent tasks on such platforms introduces significant challenges due to strict real-time constraints, high energy consumption, and the NP-hard nature of task mapping. This paper proposes a novel hybrid scheduling framework to jointly optimize energy efficiency and timeliness for Directed Acyclic Graph (DAG) applications. The framework operates in three tiers: first, a Genetic Algorithm (GA) performs a global search to determine near-optimal task-to-core mappings; second, a Dynamic Voltage and Frequency Scaling (DVFS) manager is integrated into the GA’s fitness function to accurately capture energy-performance trade-offs; and third, a Graph Neural Network (GNN) is trained to imitate the GA+DVFS policy, enabling fast and high-quality online scheduling decisions. Experimental results demonstrate that the proposed approach achieves a balanced trade-off between power consumption and deadline satisfaction, while the GNN significantly accelerates scheduling without compromising solution quality. Our GLEAM method reduced energy consumption on average by 49.08% and improved the makespan on average by 27.03% compared to baseline methods.
Amir Hossein Ansari, Mohsen Ansari, Sepideh Safari, Alireza Ejlali, Jörg Henkel
DATE2
2026 HAMLET: Heterogeneous Adaptive Mapping and Low-Energy Task Scheduling for Heterogeneous Multicore IoT Devices
abstract
The rapid growth of Internet of Things (IoT) deployments and the increasing integration of multiple cores on a single chip have made managing power consumption and ensuring thermal safety critical challenges for multicore IoT devices. This paper proposes HAMLET, a heterogeneous adaptive mapping and low-energy task scheduling framework that jointly manages energy consumption, timing constraints, thermal safety, and reliability in multicore IoT platforms through a multi-objective genetic optimization approach. In the offline phase, task-to-core mapping and replication decisions are optimized to minimize peak power and energy consumption while preserving schedulability and reliability targets. Moreover, predictive temperature control is achieved in the runtime phase using a Long Short-Term Memory (LSTM) model, which anticipates core temperature evolution and enables proactive control actions. This approach ensures that tasks are allocated to processing cores to prevent thermal violations while maintaining system-level reliability. Furthermore, DVFS and DPM mechanisms are adaptively applied at runtime based on predicted thermal states to reduce energy consumption and avoid thermal emergencies. Experimental results demonstrate that the proposed method achieves up to 84.39% reduction in peak power and up to 81.98% reduction in energy consumption, while improving schedulability by up to 20.7% compared to state-of-the-art techniques, confirming its effectiveness for energy- and reliability-aware multicore IoT systems.
Amir Hossein Ansari, Mohsen Ansari, Alireza Ejlali, Jörg Henkel
IEEE Internet Things J.2
2026 MOFCO: Mobility- and Migration-Aware Task Offloading in Three-Layer Fog Computing Environments
abstract
Task offloading in three-layer fog computing environments presents a critical challenge due to user equipment (UE) mobility, which frequently triggers costly service migrations and degrades overall system performance. This paper addresses this problem by proposing MOFCO, a novel Mobility- and Migration-aware Task Offloading algorithm for Fog Computing environments. The proposed method formulates task offloading and resource allocation as a Mixed-Integer Nonlinear Programming (MINLP) problem and employs a heuristic-aided evolutionary game theory approach to solve it efficiently. To evaluate MOFCO, we simulate mobile users using SUMO, providing realistic mobility patterns. Experimental results show that MOFCO reduces system cost—defined as a combination of latency and energy consumption—by an average of 23% and up to 67% in certain scenarios compared to state-of-the-art methods.
Soheil Mahdizadeh, Elyas Oustad, Mohsen Ansari
IEEE Internet Things J.3
2026 SIREN: Multiobjective Game-Theoretic Scheduler Based on Memory-Driven Gray Wolf Optimization in Fog-Cloud Computing
abstract
Fog-cloud task scheduling faces the dual challenge of maintaining critical IoT workloads despite node failures while adhering to strict energy budgets. We present SIREN, a game-theoretic framework that treats fog nodes as strategic players, embedding reliability benefits and DVFS-aware energy costs directly into their payoffs. By searching the joint strategy space with a Memory-Driven Grey Wolf Optimizer (MDGWO), SIREN adapts placements, selective replication, and frequency settings to workload dynamics. Extensive evaluations on the Alibaba 2018 and Google 2011 cluster traces and on a latency-critical healthcare application demonstrate that SIREN converges to near-Nash schedules that minimize energy while maximizing reliability. Results confirm that SIREN delivers (i) 100% task success rates in critical healthcare scenarios, (ii) 2.08×–4.24× lower worst-case energy consumption than leading baselines, and (iii) a 3.9×–5.8× reduction in network usage, establishing a new benchmark for resilient, energy-efficient fog computing.
Abolfazl Younesi, Mohsen Ansari, Alireza Ejlali, MohammadAmin Fazli, Muhammad Shafique 0001, Jörg Henkel
IEEE Internet Things J.2
2026 REEG: Reinforcement Learning-Based Gang Scheduling in Multicore Embedded Systems
abstract
In modern cyber-physical systems, the increasing complexity and parallel execution demands of applications necessitate adopting advanced task scheduling and task orchestration models to optimize system performance and resource allocation in multicore systems. In this brief, we present a novel task model and its corresponding RL-based algorithm, named Reinforcement Learning-Based Energy-Efficient Gang Scheduling in Multicore Cyber-Physical Systems (REEG) in which each node of a Directed Acyclic Graph (DAG) is characterized as a gang task, requiring simultaneous execution across multiple cores, thus offering a precise and scalable framework for optimizing the performance of complex, highly parallel workloads in modern multicore systems. However, significant energy optimization challenges exist due to the complex dependencies and simultaneous core usage. We propose a framework based on reinforcement learning (RL) that dynamically modifies core allocation and execution techniques to reduce energy consumption and maintain computational efficiency without compromising quality of service (QoS) or performance. This approach reduces energy consumption while maintaining performance and QoS, ensuring efficient computation. Our experimental results demonstrate that the RL-based method achieves notable energy reductions and improves system efficiency compared to the state-of-the-art method. On average, our proposed method (REEG) achieves 28.19% less energy consumption and 60.75% more QoS compared to the state-of-the-art methods.
Majid Hajilou, Mohsen Ansari
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2026 MOSAIC: Mobility-Oriented Scheduling and Intelligent Resource Allocation for IoT
abstract
The relentless growth of mobile Internet of Things (IoT) devices has shifted computation toward a distributed computing continuum, spanning edge, fog, and cloud layers, where energy efficiency, low latency, and dynamic node mobility are critical yet often conflicting goals. Existing scheduling frameworks struggle to balance these demands under real-world conditions, especially as device movement and heterogeneous workloads increase system complexity. We present MOSAIC, a mobility-aware scheduling and resource management framework designed to optimize performance in dynamic IoT environments. Our approach introduces three key innovations. First, a refined five-tier architecture extends the traditional edge-fog-cloud hierarchy by adding proximity, local, and regional mobility layers, enabling computation to follow mobile users more effectively and reducing unnecessary network traffic. Second, MOSAIC integrates a preemption-aware dynamic scheduler with an Adaptive-$\lambda$reinforcement learning-based resource manager that adapts based on workload changes and mobility patterns, prioritizing energy-efficient edge execution while meeting strict deadlines. Third, the framework utilizes real-world mobility traces, including Levy-Walk, Random-Walk, and Geolife, to drive reconfiguration and improve decision accuracy. We evaluate MOSAIC through a large-scale deployment across three geographically distributed regions of the Grid'5000 testbed, using realistic workflows and mixed periodic/DAG task loads. Our results show that, compared to state-of-the-art schedulers, MOSAIC reduces energy consumption by 35.9%–×1.5, lowers latency by 42.8%–×4.9, and shortens makespan by 22.6%–×7.2, all while maintaining 100% deadline satisfaction across diverse mobility scenarios.
Abolfazl Younesi, Mehrab Toghani, Sepideh Safari, Mohsen Ansari, Thomas Fahringer
IEEE Trans. Mob. Comput.4
2026 Pilot: Power-Aware Hybrid Fault Tolerance in Multi-Core Embedded Systems
abstract
With the advancement of technology size and the integration of multiple cores on a single chip, the probability of fault occurrence has increased. These faults can be transient or permanent, requiring techniques to manage both types. Hybrid fault tolerance techniques have emerged as effective solutions to handle both types. In this paper, we propose a power-aware hybrid fault tolerance (called Pilot). Our approach utilizes checkpointing with rollback-recovery and primary/backup techniques, tolerating two kinds of faults. Moreover, in real-time embedded systems, power consumption is a critical constraint that must be managed. To do this, we exploit the Thermal Safe Power (TSP) constraint for each processing core. Based on this constraint and the utilization of each core, tasks are mapped and scheduled, while guaranteeing the timing constraints. Our experimental results demonstrate that our proposed methods can meet the reliability target by tolerating the optimal number of fault occurrences in each task while reducing power consumption. Our proposed methods are compared to state-of-the-art techniques in terms of schedulability, power consumption, Quality of Service (QoS), energy consumption, and reliability. The peak power and energy consumption are reduced on average by 34.2% and 15.9%, respectively, the QoS is improved on average to 28.7%, and the schedulability is improved on average to 14.6% while satisfying the system reliability target.
Amir Hossein Ansari, Moein Esnaashari, Sepideh Safari, Mohsen Ansari, Alireza Ejlali, Jörg Henkel
IEEE Trans. Parallel Distributed Syst.4
2025 Work-in-Progress: LEETMIC: Reinforcement Learning-Based Energy-Efficient Task Scheduling in Multicore Cyber-Physical Systems
abstract
Energy efficiency is a critical design constraint in multicore Cyber-Physical Systems (CPS). Using energy management methods can violate timing constraints; hence, designing scheduling policies that can adapt to the dynamic and unpredictable nature of aperiodic real-time tasks remains a significant challenge. This paper introduces LEETMIC, a novel deep reinforcement learning framework for energy-aware real-time scheduling. To manage a variable number of active jobs, LEETMIC utilizes a learned policy network to determine the scheduling priority and Dynamic Voltage and Frequency Scaling (DVFS) level for each task individually, based on a combination of the task's local attributes and a summary of the global system state. The policy is trained using Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm to co-optimize for both task schedulability and energy consumption. The trained policy neural network is a compact multi-layer perceptron (MLP), making it suitable for online deployment in embedded systems. Experimental results demonstrate that, compared to the Global Earliest Deadline First (GEDF) scheduler, LEETMIC achieves a similar success ratio while significantly reducing energy consumption by 41.5% on average (up to 70%).
Erfan Bagheri Soula, Moein Esnaashari, Sepideh Safari, Mohsen Ansari
RTSS4
2025 DReaM: Deep Reinforcement Learning for Joint Reliability and Power Management in Multicore IoT Devices
abstract
In multicore Internet of Things (IoT) devices, high reliability, low power/energy design, and real-time computing are three important design requirements. The multiple cores in these systems enable us to exploit different fault-tolerance techniques for improving system reliability. However, these techniques impose power/energy and timing overheads on the system. This paper proposes DReaM, a power- and reliability-aware approach that utilizes Deep Reinforcement Learning (DRL) to minimize power consumption in multicore systems while maintaining the desired reliability level and meeting timing constraints. DReaM controls the power consumption by adjusting the voltage and frequency levels for each task and choosing the most suitable fault-tolerant techniques between Primary/Backup (P/B), Triple Modular Redundancy (TMR), and N-Modular Redundancy (NMR). Evaluations using an ARM-based processor and applications from the MiBench benchmark suite were conducted under different fault injection scenarios. The experimental results show that DReaM reduces power consumption on average by 19.86% (up to 64.1%), with improved Quality of Service (QoS) by an average of 11.02% (up to 77%) while maintaining the required system reliability.
Moein Esnaashari, Mohsen Ansari, Alireza Ejlali
IEEE Internet Things J.2
2025 RASOUL: A Reliability-Aware Task Allocation Strategy to Improve Success Rate and Energy Saving in Mobile-Edge Computing
abstract
With the emergence of the mobile-edge computing (MEC) paradigm, many challenges related to the cloud computing paradigm were addressed, including bandwidth consumption, high latency, nonreal-time computing, etc. This makes the MEC paradigm a promising technology for real-time applications, such as healthcare, automotive vehicles, smart cities, virtual reality, etc. Since MEC environments operate on wireless networks, one of the open problems in the MEC paradigm is task allocation strategies. Reliable task allocation in such environments directly impacts the network’s Quality of Service (QoS). This article proposes RASOUL, a learning-based task allocation strategy to improve the overall task success ratio and energy consumption in MEC environments. We have formulated the reliability-aware task allocation problem to improve task success ratio and reduce energy consumption and conducted experiments on the proposed method to evaluate its performance compared to state-of-the-art methods. Experimental results demonstrate that RASOUL can achieve an average improvement in QoS of 10.1% (up to 16.3%) and 18.7% (up to 22.1%) in task success ratio while reducing energy consumption by an average of 38.1% compared to other task allocation strategies.
Amir Mahdi Rasouli, Moein Esnaashari, Mohsen Ansari
IEEE Internet Things J.3
2025 DIST: Distributed Learning-Based Energy-Efficient and Reliable Task Scheduling and Resource Allocation in Fog Computing
abstract
This paper presents DIST, a novel distributed reinforcement learning-based (DRL) framework for energyefficient and reliable task scheduling and resource allocation in fog computing, low-latency computing solutions driven by the rapid deployment of IoT devices, and time-sensitive applications. DIST is built based on a novel distributed Q-learning to enable fog nodes to learn an optimal strategy to balance energy consumption, task execution time, and system reliability. The main novelty includes a cooperative Dynamic Voltage and Frequency Scaling-enabled task scheduling policy that dynamically adjusts node energy level to ensure power consumption reduction without sacrificing deadline adherence or reliability. The results demonstrate that DIST reduces energy consumption by up to 52.26%, realizes 38% higher success rates, and reduces task wait times by up to 46.77%, compared with state-of-the-art algorithms.
Elyas Oustad, Abolfazl Younesi, Mohsen Ansari, Sepideh Safari, Mohammad Arman Soleimani, Jörg Henkel, Alireza Ejlali
IEEE Trans. Serv. Comput.3
2025 MoTiCPS: Energy Optimization on Multi-Objective Task Scheduling in IoT-Integrated Cyber-Physical Systems
abstract
Fog computing enhances cyber-physical systems (CPS) by processing data closer to the network edge. However, the performance of fog nodes is critical for maintaining system responsiveness and quality of service (QoS). This paper introduces MoTiCPS, a novel task scheduling and resource allocation method built on the Osprey Optimization Algorithm (OOA). MoTiCPS improves task reliability and balances resource use across edge devices, optimizing fog node performance under real-time constraints. Simulation results show that MoTiCPS increases task success rates by 32% and reduces energy consumption by 39%, significantly outperforming benchmark methods. These improvements highlight MoTiCPS's potential to enhance the efficiency and scalability of CPSs in various application domains.
Abolfazl Younesi, Elyas Oustad, Mohammad Abolnejadian, Mohsen Ansari, Alireza Ejlali
IEEE Trans. Sustain. Comput.4
2023 ATLAS: Aging-Aware Task Replication for Multicore Safety-Critical Systems
abstract
A major requirement of safety-critical systems is high reliability at low power consumption. Dynamic voltage and frequency (v/f) scaling (DVFS) techniques are widely exploited to reduce power consumption. However, DVFS through downscaling v/f levels has a negative impact on the reliability of the tasks running on the cores, and through upscaling v/f levels has circuitlevel aging effects. To achieve high reliability in multicore safetycritical systems, task replication as a fault-tolerant technique is an established way to deal with the negative effect of downscaling v/f levels, but it may accelerate aging effects due to elevating the on-chip temperatures. In this paper, we propose an aging-aware task replication (called ATLAS) method that solves the problem of satisfying the desired reliability target for a set of periodic hard real-time tasks which are executed on a multicore system. The proposed method satisfies the reliability target of the tasks through updating the required number of replicas for each task at different years. We replicate the tasks through our proposed formulas such that the reliability target is satisfied. However, task replication increases the temperature of the system and accelerates aging. To decelerate aging, we attempt to reduce the temperature while mapping and scheduling the tasks. We have also developed a modified demand bound function (DBF) for our aging-aware task replication method to verify scheduling the realtime tasks. Compared to the existing state-of-the-art techniques, experimental results for safety-critical applications on different configurations of multicore systems demonstrate the efficiency and effectiveness of our proposed method. Experiments show that our proposed method improves schedulability on average by 16.1% and reduces the temperature on average by 7.4°C compared to state-of-the-art methods while meeting the system reliability target.
Mohsen Ansari, Sepideh Safari, Amir Yeganeh-Khaksar, Roozbeh Siyadatzadeh, Pourya Gohari-Nazari, Heba Khdr, Muhammad Shafique 0001, Jörg Henkel, Alireza Ejlali
RTAS1
2023 ReLIEF: A Reinforcement-Learning-Based Real-Time Task Assignment Strategy in Emerging Fault-Tolerant Fog Computing
abstract
Due to the real-time requirements in several IoT applications, fog computing has emerged to overcome the long latency and other constraints of cloud computing. Due to the high probability of packet loss, energy limitation of IoT devices, and the external disturbances that may frequently occur on the fog infrastructure, the timing constraints of real-time tasks may be compromised. Therefore, the reliability of executing real-time tasks has always been a significant challenge in fog computing. In addition to the correct execution of the tasks, it is also important to execute them before their deadlines according to their real-time classification. State-of-the-art methods generally focus on the delay or functionality of tasks in fog computing systems. However, those methods do not widely focus on the reliability of tasks with real-time constraints in dynamic environments. In this article, a novel primary backup task assignment strategy based on machine learning (ReLIEF) is proposed to improve the reliability of fog-based IoT systems. To identify suitable nodes for the execution of the primary and backup tasks, ReLIEF employs a reinforcement learning (RL) approach, which has an outstanding performance in dynamic environments by establishing a balance between communication delay and workload on each fog device. Based on the simulations, our newly proposed technique has been able to reduce the amount of task dropping rate by up to 84% against the state of the art. Moreover, it is capable of balancing the workload distribution while increasing the reliability of the system by nearly 72% compared with its counterparts.
Roozbeh Siyadatzadeh, Fatemeh Mehrafrooz, Mohsen Ansari, Bardia Safaei 0001, Muhammad Shafique 0001, Jörg Henkel, Alireza Ejlali
IEEE Internet Things J.3
2023 Power-Efficient and Aging-Aware Primary/Backup Technique for Heterogeneous Embedded Systems
abstract
One of the essential requirements of embedded systems is a guaranteed level of reliability. In this regard, fault-tolerance techniques are broadly applied to these systems to enhance reliability. However, fault-tolerance techniques may increase power consumption due to their inherent redundancy. For this purpose, power management techniques are applied, along with fault-tolerance techniques, which generally prolong the system lifespan by decreasing the temperature and leading to an aging rate reduction. Yet, some power management techniques, such as Dynamic voltage and frequency scaling (DVFS), increase the transient fault rate and timing error. For this reason, heterogeneous multicore platforms have received much attention due to their ability to make a trade-off between power consumption and performance. Still, it is more complicated to map and schedule tasks in a heterogeneous multicore system. In this paper, for the first time, we propose a power management method for a heterogeneous multicore system that reduces power consumption and tolerates both transient and permanent faults through primary/backup technique while considering core-level power constraint, real-time constraint, and aging effect. Experimental evaluations demonstrate the efficiency of our proposed method in terms of reducing power consumption compared to the state-of-the-art schemes, together with guaranteeing reliability and considering the aging effect.
Mohsen Ansari, Sepideh Safari, Nezam Rohbani, Alireza Ejlali, Bashir M. Al-Hashimi
IEEE Trans. Sustain. Comput.1
2023 Passive Primary/Backup-Based Scheduling for Simultaneous Power and Reliability Management on Heterogeneous Embedded Systems
abstract
In addition to meeting the real-time constraint, power/energy efficiency and high reliability are two vital objectives for real-time embedded systems. Recently, heterogeneous multicore systems have been considered an appropriate solution for achieving joint power/energy efficiency and high reliability. However, power/energy and reliability are two conflict requirements due to the inherent redundancy of fault-tolerance techniques. Also, because of the heterogeneity of the system, the execution of the tasks, especially real-time tasks, in the heterogeneous system is more complicated than the homogeneous system. The proposed method in this paper employs a passive primary/backup technique to preserve the reliability requirement of the system at a satisfactory level and reduces power/energy consumption in heterogeneous multicore systems by considering real-time and peak power constraints. The proposed method attempts to map the primary and backup tasks in a mixed manner to benefit from the execution of the tasks in different core types and schedules the backup tasks after finishing the primary tasks to remove the overlap between the execution of the primary and backup tasks. Compared to the existing state-of-the-art methods, experimental results demonstrate our proposed method's power efficiency and effectiveness in terms of schedulability.
Sina Yari-Karin, Roozbeh Siyadatzadeh, Mohsen Ansari, Alireza Ejlali
IEEE Trans. Sustain. Comput.3
2022 Convolutional Deep Kernel Method for Land Cover Mapping from Hyperspectral Imagery
abstract
In recent years, kernel-based methods and Deep Learning (DL) models have become the two most successful Remote Sensing (RS) analysis techniques for various Earth observations, particularly hyperspectral images. However, kernel-based methods are generally considered shallow models and intrinsically inconsistent with end-to-end learning. On the other hand, end-to-end learning is one of DL models' essential features as it seems to be responsible for their proven higher performances. Nevertheless, kernel methods are based on rigid mathematical theory and can efficiently cope with high-dimensional data. This paper proposed a hybrid deep kernel model to benefit from both kernel-based methods and DL models. This novel deep kernel model, namely Convolutional Kernel Network (CKN), was applied to two benchmark hyperspectral image datasets. Moreover, the proposed hybrid method was compared to Support Vector Machine (SVM) classifiers with various kernel functions. The experimental results indicated that the CKN's outperforms SVM.
Mohsen Ansari, Weimin Huang 0001, Saeid Homayouni, Saeid Niazmardi, Abdolreza Safari
IGARSS1
2022 Power-Aware Checkpointing for Multicore Embedded Systems
abstract
Increasing the number of cores integrated on a single chip offers a great potential for the implementation of fault-tolerant techniques to achieve high reliability in real-time embedded systems. Checkpointing with rollback-recovery is a well-established technique to tolerate transient faults in multicore platforms. To consider the worst-case fault occurrence scenario, checkpointing technique requires to re-execute some parts of the tasks, and that might lead to simultaneous execution of task parts with high power consumptions, which eventually might result in a peak power increase beyond the thermal design power (TDP). Exceeding TDP can elevate on-chip temperatures beyond safe limits, and thereby triggering countermeasures that throttle down the voltage and frequency levels or power gate the cores. Such countermeasures might lead to violating task deadlines and degrading the system's reliability. To avoid such severe scenarios, it is inevitable to consider the impact of applying fault-tolerant techniques on the power consumption and prevent violating the power constraint of the chip, i.e., TDP. This paper presents for the first time, a peak-power-aware checkpointing (PPAC) technique that tolerates a given number of faults,k, while at the same time meets the power constraint in hard real-time embedded systems. To do this, our proposed technique (PPAC) adjusts the timing of the checkpoints, which have lower power consumption than the tasks to the execution time points that have power spikes beyond TDP. Moreover, PPAC exploits the available slack times on the cores to delay the execution of some tasks to avoid the remaining power spikes beyond TDP, which could not be mitigated by solely adjusting checkpoints. To evaluate our technique, we extend the state-of-the-art system-level simulator, gem5, with the state-of-the-art checkpointing module in Linux. Our experimental results show that our proposed technique is able to tolerate a given number of faults without exceeding the timing and power constraints in hard real-time embedded systems. The resulting peak power reduction achieved by our technique compared to state-of-the-art techniques is an average of 23%. Moreover, our technique employs the Dynamic Power Management (DPM) during the slack times resulting at runtime in the case of fault-free scenarios, which provides energy savings with an average of 17.28% and up to 61.1%.
Mohsen Ansari, Sepideh Safari, Heba Khdr, Pourya Gohari-Nazari, Jörg Henkel, Alireza Ejlali, Shaahin Hessabi
IEEE Trans. Parallel Distributed Syst.1
2022 TherMa-MiCs: Thermal-Aware Scheduling for Fault-Tolerant Mixed-Criticality Systems
abstract
Multicore platforms are becoming the dominant trend in designing Mixed-Criticality Systems (MCSs), which integrate applications of different levels of criticality into the same platform. A well-known MCS is the dual-criticality system that is composed of low-criticality and high-criticality tasks. The availability of multiple cores on a single chip provides opportunities to employ fault-tolerant techniques, such as N-Modular Redundancy (NMR), to ensure the reliability of MCSs. However, applying fault-tolerant techniques will increase the power consumption on the chip, and thereby on-chip temperatures might increase beyond safe limits. To prevent thermal emergencies, urgent countermeasures, like Dynamic Voltage and Frequency Scaling (DVFS) or Dynamic Power Management (DPM) will be triggered to cool down the chip. Such countermeasures, however, might not only lead to suspending low-criticality tasks, but also it might lead to violating timing constraints of high-criticality tasks. In order to prevent such severe scenarios, it is indispensable to consider a temperature constraint within the scheduling process of fault-tolerant MCSs. Therefore, this paper presents, for the first time, a thermal-aware scheduling scheme for fault-tolerant MCSs, named TherMa-MiCs. In particular, TherMa-MiCs, satisfies the temperature constraint jointly with the timing constraints of the high-criticality tasks, while attempting to maximize the QoS of low-criticality tasks under the predefined constraints. At the same time, a reliability target is satisfied by employing the well-known N-Modular Redundancy (NMR) fault-tolerant technique. Experimental results show that our proposed scheme meets the temperature and timing constraints, while at the same time, improving the QoS of low-criticality tasks, with an average of 44%.
Sepideh Safari, Heba Khdr, Pourya Gohari-Nazari, Mohsen Ansari, Shaahin Hessabi, Jörg Henkel
IEEE Trans. Parallel Distributed Syst.4
2021 READY: Reliability- and Deadline-Aware Power-Budgeting for Heterogeneous Multicore Systems
abstract
Tackling the dark silicon problem in a heterogeneous multicore system, the temperature constraints across the system should be addressed carefully by assigning a proper set of tasks to a pool of the heterogeneous cores during the run-time. When such a system is utilized in a reliable/real-time application, the reliability/timing constraints of the application should also be augmented to the temperature constraints and make the tasks mapping problem more and more complex. To solve the mapping problem in such a situation, we propose READY; an online reliability- and deadline-aware mapping and scheduling algorithm for heterogeneous multicore systems. READY utilizes an adaptive power constraint (as a metric for temperature measurement) that is updated according to the number and position of the active cores on the chip. READY, first, attempts to meet the reliability target of the system by improving the reliability of each task. Then, it performs the mapping and scheduling of the tasks on cores of different islands, so that the peak power and timing constraints are met. The simulation results illustrate that while READY guarantees the timing constraints and meets reliability targets, it improves the peak-power-aware system schedulability (chip performance) by 23.77% (up to 40.69%).
Javad Saber-Latibari, Mohsen Ansari, Pourya Gohari-Nazari, Sina Yari-Karin, Amir Mahdi Hosseini Monazzah, Alireza Ejlali
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Peak-Power-Aware Energy Management for Periodic Real-Time Applications
abstract
Two main objectives in designing real-time embedded systems are high reliability and low power consumption. Hardware replication (e.g., standby-sparing) can provide high reliability while keeping the power consumption under control. In this paper, we consider a standby-sparing system where the main tasks on primary cores are scheduled by our proposed peak-power-aware earliest-deadline-first policy while the backup tasks on spare cores are scheduled by our proposed peak-power-aware earliest-deadline-late policy to meet the chip thermal design power (TDP) constraint. These policies provide the best opportunity to shift the task executions as much as possible to minimize execution overlaps between main and backup tasks that consume high power consumption. Since TDP is the maximum amount of power generated by a chip that the cooling component is designed to dissipate under any workload, the total power consumption should not be higher than the TDP constraint. When a task finishes successfully a larger portion of its corresponding copy task can be canceled, resulting in a significant amount of peak/average power reduction. To achieve further peak/average power reduction, we use dynamic voltage and frequency scaling and dynamic power management (DPM). The main reason of using DPM is that, once the first copy of each task has finished successfully, its corresponding copy task is terminated, and if there is no more task for execution, the core goes to a low-power mode. We evaluated our scheme under various system configurations. Experiments show that our scheme provides up to 47.6% (on average by 28.2%) peak power reduction compared to four state-of-the-art techniques.
Mohsen Ansari, Amir Yeganeh-Khaksar, Sepideh Safari, Alireza Ejlali
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 Simultaneous Management of Peak-Power and Reliability in Heterogeneous Multicore Embedded Systems
abstract
Analysis of reliability, power, and performance at hardware and software levels due to heterogeneity is a crucial requirement for heterogeneous multicore embedded systems. Escalating power densities have led to thermal issues for heterogeneous multicore embedded systems. This paper proposes a peak-power-aware reliability management scheme to meet power constraints through distributing power density on the whole chip such that reliability targets are satisfied. In this paper, we consider peak power consumption as a system-level power constraint to prevent system failure. To balance the power consumption, we also employ a Dynamic Frequency Scaling (DFS) method to further reduce peak power consumption and satisfy thermal constraints on the chip. We illustrate the benefits of our scheme by comparing it with state-of-the-art schemes, resulting in average in 26.5 percent less peak power consumption (up to 54.3 percent).
Mohsen Ansari, Javad Saber-Latibari, Mostafa Pasandideh, Alireza Ejlali
IEEE Trans. Parallel Distributed Syst.1
2019 Peak Power Management to Meet Thermal Design Power in Fault-Tolerant Embedded Systems
abstract
Multicore platforms provide a great opportunity for implementation of fault-tolerance techniques to achieve high reliability in real-time embedded systems. Passive redundancy is well-suited for multicore platforms and a well-established technique to tolerate transient and permanent faults. However, it incurs significant power overheads, which go wasted in fault-free execution scenarios. Meanwhile, due to the Thermal Design Power (TDP) constraint, in some cases, it is not feasible to simultaneously power on all cores on a multicore platform. Since TDP is the maximum sustainable power that a chip can consume, violating TDP makes some cores automatically restart or significantly reduce their performance to prevent a permanent damage. This may affect timeliness of the system, and hence, designers face a challenge in deciding how to use multicore platforms in real-time embedded systems. In this paper, at first, we study how the use of passive redundancy (especially for Triple Modular redundancy) can violate TDP on multicore platforms. Then, we propose a scheme for scheduling real-time tasks in multicore systems to conquer the peak power problem in NMR systems. This is because in multicore embedded systems an efficient solution for meeting the TDP constraint is reducing the peak power consumption. The proposed scheme tries to remove overlaps of the peak power of concurrently executing tasks to keep the maximum power consumption below the chip TDP. In the proposed scheme, we devised a policy called PPA-LTF to manage peak power consumption. This policy prevents tasks execution that consume higher power according to the tasks’ power traces. Our experimental results show that our scheme provides up to 50 percent (on average by 39 percent) peak power reduction compared to state-of-the-art schemes.
Mohsen Ansari, Sepideh Safari, Amir Yeganeh-Khaksar, Alireza Ejlali
IEEE Trans. Parallel Distributed Syst.1
2019 On the Scheduling of Energy-Aware Fault-Tolerant Mixed-Criticality Multicore Systems with Service Guarantee Exploration
abstract
Advancement of Cyber-Physical Systems has attracted attention to Mixed-Criticality Systems (MCSs), both in research and in industrial designs. As multicore platforms are becoming the dominant trend in MCSs, joint energy and reliability management is a crucial issue. In addition, providing guaranteed service level for low-criticality tasks in critical mode is of great importance. To address these problems, we propose “LETR-MC” scheme that simultaneously supports certification, energy management, fault-tolerance, and guaranteed service level in mixed-criticality multicore systems. In this paper, we exploit task-replication to not only satisfy reliability requirements, but also to improve the QoS of low-criticality tasks in overrun situation. Our proposed LETR-MC scheme determines the number of replicas, and reduces the execution time overlap between the primary tasks and replicas. Moreover, instead of ignoring low-criticality tasks or selectively executing them without any guaranteed service level in overrun mode, it mathematically explores the minimum achievable service guarantee for each low-criticality task in different execution modes, i.e., normal, fault-occurrence, overrun and critical operation modes. We develop novel unified demand bound functions (DBF), along with a DVFS method based on the proposed DBF analysis. Our experimental results show that LETR-MC provides up to 59 percent (24 percent on average) energy saving, and significantly improves the service levels of low-criticality tasks compared to the state-of-the-art schemes.
Sepideh Safari, Mohsen Ansari, Ghazal Ershadi, Shaahin Hessabi
IEEE Trans. Parallel Distributed Syst.2
2016 Parallel HTTP for Video Streaming in Wireless Networks
abstract
To stream video using HTTP, a client device sequentially requests and receives chunks of the video file from the server over a TCP connection. It is well-known that TCP performs poorly in networks with high latency and packet loss such as wireless networks. On mobile devices, in particular, using a single TCP connection for video streaming is not efficient, and thus, the user may not receive the highest video quality possible. In this paper, we design and analyze a system called ParS that uses parallel TCP connections to stream video on mobile devices. Our system uses parallel connections to fetch each chunk of the video file using HTTP range requests. We present measurement results to characterize the performance of ParS under various network conditions in terms of latency, loss rate and bandwidth. Given the limited communication and computational resources of mobile devices, we then focus on determining the minimum number of TCP connections required to achieve high utilization of the wireless bandwidth.
Mohsen Ansari, Majid Ghaderi
MASCOTS1