Sepideh Safari

dblp:174/1719 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0003-4645-8255ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GLEAM: A Graph-Learning Enhanced Adaptive Metaheuristic for Power-Aware Scheduling on Heterogeneous Cyber-Physical Systems
abstract
The increasing complexity of embedded and Cyber-Physical Systems (CPS) has accelerated the adoption of heterogeneous multi-core architectures, which combine performance and energy efficiency. However, scheduling dependent tasks on such platforms introduces significant challenges due to strict real-time constraints, high energy consumption, and the NP-hard nature of task mapping. This paper proposes a novel hybrid scheduling framework to jointly optimize energy efficiency and timeliness for Directed Acyclic Graph (DAG) applications. The framework operates in three tiers: first, a Genetic Algorithm (GA) performs a global search to determine near-optimal task-to-core mappings; second, a Dynamic Voltage and Frequency Scaling (DVFS) manager is integrated into the GA’s fitness function to accurately capture energy-performance trade-offs; and third, a Graph Neural Network (GNN) is trained to imitate the GA+DVFS policy, enabling fast and high-quality online scheduling decisions. Experimental results demonstrate that the proposed approach achieves a balanced trade-off between power consumption and deadline satisfaction, while the GNN significantly accelerates scheduling without compromising solution quality. Our GLEAM method reduced energy consumption on average by 49.08% and improved the makespan on average by 27.03% compared to baseline methods.
Amir Hossein Ansari, Mohsen Ansari, Sepideh Safari, Alireza Ejlali, Jörg Henkel
DATE3
2026 MOSAIC: Mobility-Oriented Scheduling and Intelligent Resource Allocation for IoT
abstract
The relentless growth of mobile Internet of Things (IoT) devices has shifted computation toward a distributed computing continuum, spanning edge, fog, and cloud layers, where energy efficiency, low latency, and dynamic node mobility are critical yet often conflicting goals. Existing scheduling frameworks struggle to balance these demands under real-world conditions, especially as device movement and heterogeneous workloads increase system complexity. We present MOSAIC, a mobility-aware scheduling and resource management framework designed to optimize performance in dynamic IoT environments. Our approach introduces three key innovations. First, a refined five-tier architecture extends the traditional edge-fog-cloud hierarchy by adding proximity, local, and regional mobility layers, enabling computation to follow mobile users more effectively and reducing unnecessary network traffic. Second, MOSAIC integrates a preemption-aware dynamic scheduler with an Adaptive-$\lambda$reinforcement learning-based resource manager that adapts based on workload changes and mobility patterns, prioritizing energy-efficient edge execution while meeting strict deadlines. Third, the framework utilizes real-world mobility traces, including Levy-Walk, Random-Walk, and Geolife, to drive reconfiguration and improve decision accuracy. We evaluate MOSAIC through a large-scale deployment across three geographically distributed regions of the Grid'5000 testbed, using realistic workflows and mixed periodic/DAG task loads. Our results show that, compared to state-of-the-art schedulers, MOSAIC reduces energy consumption by 35.9%–×1.5, lowers latency by 42.8%–×4.9, and shortens makespan by 22.6%–×7.2, all while maintaining 100% deadline satisfaction across diverse mobility scenarios.
Abolfazl Younesi, Mehrab Toghani, Sepideh Safari, Mohsen Ansari, Thomas Fahringer
IEEE Trans. Mob. Comput.3
2026 Pilot: Power-Aware Hybrid Fault Tolerance in Multi-Core Embedded Systems
abstract
With the advancement of technology size and the integration of multiple cores on a single chip, the probability of fault occurrence has increased. These faults can be transient or permanent, requiring techniques to manage both types. Hybrid fault tolerance techniques have emerged as effective solutions to handle both types. In this paper, we propose a power-aware hybrid fault tolerance (called Pilot). Our approach utilizes checkpointing with rollback-recovery and primary/backup techniques, tolerating two kinds of faults. Moreover, in real-time embedded systems, power consumption is a critical constraint that must be managed. To do this, we exploit the Thermal Safe Power (TSP) constraint for each processing core. Based on this constraint and the utilization of each core, tasks are mapped and scheduled, while guaranteeing the timing constraints. Our experimental results demonstrate that our proposed methods can meet the reliability target by tolerating the optimal number of fault occurrences in each task while reducing power consumption. Our proposed methods are compared to state-of-the-art techniques in terms of schedulability, power consumption, Quality of Service (QoS), energy consumption, and reliability. The peak power and energy consumption are reduced on average by 34.2% and 15.9%, respectively, the QoS is improved on average to 28.7%, and the schedulability is improved on average to 14.6% while satisfying the system reliability target.
Amir Hossein Ansari, Moein Esnaashari, Sepideh Safari, Mohsen Ansari, Alireza Ejlali, Jörg Henkel
IEEE Trans. Parallel Distributed Syst.3
2025 Work-in-Progress: LEETMIC: Reinforcement Learning-Based Energy-Efficient Task Scheduling in Multicore Cyber-Physical Systems
abstract
Energy efficiency is a critical design constraint in multicore Cyber-Physical Systems (CPS). Using energy management methods can violate timing constraints; hence, designing scheduling policies that can adapt to the dynamic and unpredictable nature of aperiodic real-time tasks remains a significant challenge. This paper introduces LEETMIC, a novel deep reinforcement learning framework for energy-aware real-time scheduling. To manage a variable number of active jobs, LEETMIC utilizes a learned policy network to determine the scheduling priority and Dynamic Voltage and Frequency Scaling (DVFS) level for each task individually, based on a combination of the task's local attributes and a summary of the global system state. The policy is trained using Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm to co-optimize for both task schedulability and energy consumption. The trained policy neural network is a compact multi-layer perceptron (MLP), making it suitable for online deployment in embedded systems. Experimental results demonstrate that, compared to the Global Earliest Deadline First (GEDF) scheduler, LEETMIC achieves a similar success ratio while significantly reducing energy consumption by 41.5% on average (up to 70%).
Erfan Bagheri Soula, Moein Esnaashari, Sepideh Safari, Mohsen Ansari
RTSS3
2025 LEC-MiCs: Low-Energy Checkpointing in Mixed-Criticality Multicore Systems
abstract
With the advent of multicore platforms in designing Mixed-Criticality Systems (MCSs), simultaneous management of reliability and energy while guaranteeing an acceptable service level for low-criticality tasks is a crucial challenge. To ensure the reliability of the MCSs against transient faults, fault-tolerant techniques are employed which will increase energy consumption. To mitigate the energy overhead, the Dynamic Voltage and Frequency Scaling (DVFS) technique will be exploited. However, this technique might lead to violating the timing constraints of high-criticality tasks. Therefore, this article presents, for the first time, the low-energy checkpointing technique to guarantee the reliability of multiple preemptive periodic mixed-criticality tasks in a multicore platform. In contrast to the previous works in checkpointing technique which consider a specific number of faults that all the tasks in the system should tolerate, in this article, the number of tolerable faults for each execution section of a task and in each voltage and frequency level is determined through proposed formulas to meet the reliability target based on safety standards. Then, our proposed method determines the number of checkpoints and their non-uniform intervals for the normal and overrun sections of each task to reduce energy consumption, respectively. Moreover, the unified demand bound function (DBF) analysis is proposed for analyzing the schedulability of the task set, where each high-criticality task meets its timing and reliability constraints, and low-criticality tasks execute based on their derived guaranteed periods in each operational mode of the system. Experimental results show that our proposed scheme meets the timing and reliability constraints while at the same time, improving the Quality of Service (QoS) of low-criticality tasks and managing energy consumption with an average of 29.49% and 32.78%, respectively.
Sepideh Safari, Shayan Shokri, Shaahin Hessabi, Pejman Lotfi-Kamran
ACM Trans. Cyber Phys. Syst.1
2025 DIST: Distributed Learning-Based Energy-Efficient and Reliable Task Scheduling and Resource Allocation in Fog Computing
abstract
This paper presents DIST, a novel distributed reinforcement learning-based (DRL) framework for energyefficient and reliable task scheduling and resource allocation in fog computing, low-latency computing solutions driven by the rapid deployment of IoT devices, and time-sensitive applications. DIST is built based on a novel distributed Q-learning to enable fog nodes to learn an optimal strategy to balance energy consumption, task execution time, and system reliability. The main novelty includes a cooperative Dynamic Voltage and Frequency Scaling-enabled task scheduling policy that dynamically adjusts node energy level to ensure power consumption reduction without sacrificing deadline adherence or reliability. The results demonstrate that DIST reduces energy consumption by up to 52.26%, realizes 38% higher success rates, and reduces task wait times by up to 46.77%, compared with state-of-the-art algorithms.
Elyas Oustad, Abolfazl Younesi, Mohsen Ansari, Sepideh Safari, Mohammad Arman Soleimani, Jörg Henkel, Alireza Ejlali
IEEE Trans. Serv. Comput.4
2025 Energy-Aware Fault-Tolerant Mapping of Mixed-Criticality Tasks on Heterogeneous Multicores
abstract
Due to the different timing requirements of tasks in different criticality modes, it is challenging for Mixed-Criticality Systems (MCSs) designers to use time-redundant fault-tolerant techniques. Checkpointing with rollback recovery can be an effective option for ensuring the reliability of tasks in such systems. Despite the benefits of checkpointing it can impose significant energy overheads. A common approach to overcome this energy consumption is using Dynamic Voltage and Frequency Scaling (DVFS). However, DVFS can lead to missed deadlines for high-criticality tasks. In addition, there is an increasing trend of deploying tasks on heterogeneous multicore platforms. These platforms offer promising opportunities to meet the design requirements of mixed-criticality applications more effectively. In this paper, we propose an energy-aware checkpointing scheme for mixed-criticality tasks on heterogeneous multicore platforms. First, we calculate the timing demand of each task according to DVFS and the time overhead of checkpointing to perform the EY schedulability test in different operational modes of the system. Then we focus on energy-aware mapping of tasks to heterogeneous platforms. Finally, we use DVFS to mitigate the energy overhead of checkpoints. Our experiments show that our scheme can improve schedulability by an average of 16% compared to Little Island First (LIF) mapping, while the energy does not change appreciably. Moreover, it consumes up to 36% (20% on average) less energy compared to Big Island First (BLF) mapping.
Amir Hassan Safizadeh, Sepideh Safari, Shayan Shokri, Shaahin Hessabi
IEEE Trans. Sustain. Comput.2
2023 ATLAS: Aging-Aware Task Replication for Multicore Safety-Critical Systems
abstract
A major requirement of safety-critical systems is high reliability at low power consumption. Dynamic voltage and frequency (v/f) scaling (DVFS) techniques are widely exploited to reduce power consumption. However, DVFS through downscaling v/f levels has a negative impact on the reliability of the tasks running on the cores, and through upscaling v/f levels has circuitlevel aging effects. To achieve high reliability in multicore safetycritical systems, task replication as a fault-tolerant technique is an established way to deal with the negative effect of downscaling v/f levels, but it may accelerate aging effects due to elevating the on-chip temperatures. In this paper, we propose an aging-aware task replication (called ATLAS) method that solves the problem of satisfying the desired reliability target for a set of periodic hard real-time tasks which are executed on a multicore system. The proposed method satisfies the reliability target of the tasks through updating the required number of replicas for each task at different years. We replicate the tasks through our proposed formulas such that the reliability target is satisfied. However, task replication increases the temperature of the system and accelerates aging. To decelerate aging, we attempt to reduce the temperature while mapping and scheduling the tasks. We have also developed a modified demand bound function (DBF) for our aging-aware task replication method to verify scheduling the realtime tasks. Compared to the existing state-of-the-art techniques, experimental results for safety-critical applications on different configurations of multicore systems demonstrate the efficiency and effectiveness of our proposed method. Experiments show that our proposed method improves schedulability on average by 16.1% and reduces the temperature on average by 7.4°C compared to state-of-the-art methods while meeting the system reliability target.
Mohsen Ansari, Sepideh Safari, Amir Yeganeh-Khaksar, Roozbeh Siyadatzadeh, Pourya Gohari-Nazari, Heba Khdr, Muhammad Shafique 0001, Jörg Henkel, Alireza Ejlali
RTAS2
2023 Power-Efficient and Aging-Aware Primary/Backup Technique for Heterogeneous Embedded Systems
abstract
One of the essential requirements of embedded systems is a guaranteed level of reliability. In this regard, fault-tolerance techniques are broadly applied to these systems to enhance reliability. However, fault-tolerance techniques may increase power consumption due to their inherent redundancy. For this purpose, power management techniques are applied, along with fault-tolerance techniques, which generally prolong the system lifespan by decreasing the temperature and leading to an aging rate reduction. Yet, some power management techniques, such as Dynamic voltage and frequency scaling (DVFS), increase the transient fault rate and timing error. For this reason, heterogeneous multicore platforms have received much attention due to their ability to make a trade-off between power consumption and performance. Still, it is more complicated to map and schedule tasks in a heterogeneous multicore system. In this paper, for the first time, we propose a power management method for a heterogeneous multicore system that reduces power consumption and tolerates both transient and permanent faults through primary/backup technique while considering core-level power constraint, real-time constraint, and aging effect. Experimental evaluations demonstrate the efficiency of our proposed method in terms of reducing power consumption compared to the state-of-the-art schemes, together with guaranteeing reliability and considering the aging effect.
Mohsen Ansari, Sepideh Safari, Nezam Rohbani, Alireza Ejlali, Bashir M. Al-Hashimi
IEEE Trans. Sustain. Comput.2
2022 Power-Aware Checkpointing for Multicore Embedded Systems
abstract
Increasing the number of cores integrated on a single chip offers a great potential for the implementation of fault-tolerant techniques to achieve high reliability in real-time embedded systems. Checkpointing with rollback-recovery is a well-established technique to tolerate transient faults in multicore platforms. To consider the worst-case fault occurrence scenario, checkpointing technique requires to re-execute some parts of the tasks, and that might lead to simultaneous execution of task parts with high power consumptions, which eventually might result in a peak power increase beyond the thermal design power (TDP). Exceeding TDP can elevate on-chip temperatures beyond safe limits, and thereby triggering countermeasures that throttle down the voltage and frequency levels or power gate the cores. Such countermeasures might lead to violating task deadlines and degrading the system's reliability. To avoid such severe scenarios, it is inevitable to consider the impact of applying fault-tolerant techniques on the power consumption and prevent violating the power constraint of the chip, i.e., TDP. This paper presents for the first time, a peak-power-aware checkpointing (PPAC) technique that tolerates a given number of faults,k, while at the same time meets the power constraint in hard real-time embedded systems. To do this, our proposed technique (PPAC) adjusts the timing of the checkpoints, which have lower power consumption than the tasks to the execution time points that have power spikes beyond TDP. Moreover, PPAC exploits the available slack times on the cores to delay the execution of some tasks to avoid the remaining power spikes beyond TDP, which could not be mitigated by solely adjusting checkpoints. To evaluate our technique, we extend the state-of-the-art system-level simulator, gem5, with the state-of-the-art checkpointing module in Linux. Our experimental results show that our proposed technique is able to tolerate a given number of faults without exceeding the timing and power constraints in hard real-time embedded systems. The resulting peak power reduction achieved by our technique compared to state-of-the-art techniques is an average of 23%. Moreover, our technique employs the Dynamic Power Management (DPM) during the slack times resulting at runtime in the case of fault-free scenarios, which provides energy savings with an average of 17.28% and up to 61.1%.
Mohsen Ansari, Sepideh Safari, Heba Khdr, Pourya Gohari-Nazari, Jörg Henkel, Alireza Ejlali, Shaahin Hessabi
IEEE Trans. Parallel Distributed Syst.2
2022 TherMa-MiCs: Thermal-Aware Scheduling for Fault-Tolerant Mixed-Criticality Systems
abstract
Multicore platforms are becoming the dominant trend in designing Mixed-Criticality Systems (MCSs), which integrate applications of different levels of criticality into the same platform. A well-known MCS is the dual-criticality system that is composed of low-criticality and high-criticality tasks. The availability of multiple cores on a single chip provides opportunities to employ fault-tolerant techniques, such as N-Modular Redundancy (NMR), to ensure the reliability of MCSs. However, applying fault-tolerant techniques will increase the power consumption on the chip, and thereby on-chip temperatures might increase beyond safe limits. To prevent thermal emergencies, urgent countermeasures, like Dynamic Voltage and Frequency Scaling (DVFS) or Dynamic Power Management (DPM) will be triggered to cool down the chip. Such countermeasures, however, might not only lead to suspending low-criticality tasks, but also it might lead to violating timing constraints of high-criticality tasks. In order to prevent such severe scenarios, it is indispensable to consider a temperature constraint within the scheduling process of fault-tolerant MCSs. Therefore, this paper presents, for the first time, a thermal-aware scheduling scheme for fault-tolerant MCSs, named TherMa-MiCs. In particular, TherMa-MiCs, satisfies the temperature constraint jointly with the timing constraints of the high-criticality tasks, while attempting to maximize the QoS of low-criticality tasks under the predefined constraints. At the same time, a reliability target is satisfied by employing the well-known N-Modular Redundancy (NMR) fault-tolerant technique. Experimental results show that our proposed scheme meets the temperature and timing constraints, while at the same time, improving the QoS of low-criticality tasks, with an average of 44%.
Sepideh Safari, Heba Khdr, Pourya Gohari-Nazari, Mohsen Ansari, Shaahin Hessabi, Jörg Henkel
IEEE Trans. Parallel Distributed Syst.1
2021 REALISM: Reliability-aware energy management in multi-level mixed-criticality systems with service level degradation
Hoora Sobhani, Sepideh Safari, Javad Saber-Latibari, Shaahin Hessabi
J. Syst. Archit.2
2020 Peak-Power-Aware Energy Management for Periodic Real-Time Applications
abstract
Two main objectives in designing real-time embedded systems are high reliability and low power consumption. Hardware replication (e.g., standby-sparing) can provide high reliability while keeping the power consumption under control. In this paper, we consider a standby-sparing system where the main tasks on primary cores are scheduled by our proposed peak-power-aware earliest-deadline-first policy while the backup tasks on spare cores are scheduled by our proposed peak-power-aware earliest-deadline-late policy to meet the chip thermal design power (TDP) constraint. These policies provide the best opportunity to shift the task executions as much as possible to minimize execution overlaps between main and backup tasks that consume high power consumption. Since TDP is the maximum amount of power generated by a chip that the cooling component is designed to dissipate under any workload, the total power consumption should not be higher than the TDP constraint. When a task finishes successfully a larger portion of its corresponding copy task can be canceled, resulting in a significant amount of peak/average power reduction. To achieve further peak/average power reduction, we use dynamic voltage and frequency scaling and dynamic power management (DPM). The main reason of using DPM is that, once the first copy of each task has finished successfully, its corresponding copy task is terminated, and if there is no more task for execution, the core goes to a low-power mode. We evaluated our scheme under various system configurations. Experiments show that our scheme provides up to 47.6% (on average by 28.2%) peak power reduction compared to four state-of-the-art techniques.
Mohsen Ansari, Amir Yeganeh-Khaksar, Sepideh Safari, Alireza Ejlali
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 LESS-MICS: A Low Energy Standby-Sparing Scheme for Mixed-Criticality Systems
abstract
Multicore platforms are becoming the dominant trend in mixed-criticality systems (MCSs). Multicores provide great opportunities to realize task-level redundancy for reliability enhancement. However, they may experience limited utility in battery-powered mixed-criticality embedded systems. Hence, joint energy and reliability management is a crucial issue in designing MCSs. In this article, we propose the low energy standby-sparing mechanism in mixed-criticality system (LESS-MICS) scheme, which uses the inherent redundancy of multicores to apply the standby-sparing technique for fault-tolerance. Also, by using the inherent redundancy, the LESS-MICS scheme proposes the Parallelism and Reduction policy that can be applied to any graph traverse algorithm to enhance the schedulability of graph-based mixed-criticality tasks, as well as joint energy and reliability management, and guarantying an acceptable service level for low-criticality tasks in overrun mode. To achieve further energy reduction, we minimize energy through convex optimization, and also propose energy management heuristics which use dynamic voltage and frequency scaling and dynamic power management. We evaluated our scheme under various system configurations. Experiments show that our scheme provides, on average, 24.2% energy reduction compared to state-of-the-art techniques while preserving an acceptable QoS level.
Sepideh Safari, Shaahin Hessabi, Ghazal Ershadi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Peak Power Management to Meet Thermal Design Power in Fault-Tolerant Embedded Systems
abstract
Multicore platforms provide a great opportunity for implementation of fault-tolerance techniques to achieve high reliability in real-time embedded systems. Passive redundancy is well-suited for multicore platforms and a well-established technique to tolerate transient and permanent faults. However, it incurs significant power overheads, which go wasted in fault-free execution scenarios. Meanwhile, due to the Thermal Design Power (TDP) constraint, in some cases, it is not feasible to simultaneously power on all cores on a multicore platform. Since TDP is the maximum sustainable power that a chip can consume, violating TDP makes some cores automatically restart or significantly reduce their performance to prevent a permanent damage. This may affect timeliness of the system, and hence, designers face a challenge in deciding how to use multicore platforms in real-time embedded systems. In this paper, at first, we study how the use of passive redundancy (especially for Triple Modular redundancy) can violate TDP on multicore platforms. Then, we propose a scheme for scheduling real-time tasks in multicore systems to conquer the peak power problem in NMR systems. This is because in multicore embedded systems an efficient solution for meeting the TDP constraint is reducing the peak power consumption. The proposed scheme tries to remove overlaps of the peak power of concurrently executing tasks to keep the maximum power consumption below the chip TDP. In the proposed scheme, we devised a policy called PPA-LTF to manage peak power consumption. This policy prevents tasks execution that consume higher power according to the tasks’ power traces. Our experimental results show that our scheme provides up to 50 percent (on average by 39 percent) peak power reduction compared to state-of-the-art schemes.
Mohsen Ansari, Sepideh Safari, Amir Yeganeh-Khaksar, Alireza Ejlali
IEEE Trans. Parallel Distributed Syst.2
2019 On the Scheduling of Energy-Aware Fault-Tolerant Mixed-Criticality Multicore Systems with Service Guarantee Exploration
abstract
Advancement of Cyber-Physical Systems has attracted attention to Mixed-Criticality Systems (MCSs), both in research and in industrial designs. As multicore platforms are becoming the dominant trend in MCSs, joint energy and reliability management is a crucial issue. In addition, providing guaranteed service level for low-criticality tasks in critical mode is of great importance. To address these problems, we propose “LETR-MC” scheme that simultaneously supports certification, energy management, fault-tolerance, and guaranteed service level in mixed-criticality multicore systems. In this paper, we exploit task-replication to not only satisfy reliability requirements, but also to improve the QoS of low-criticality tasks in overrun situation. Our proposed LETR-MC scheme determines the number of replicas, and reduces the execution time overlap between the primary tasks and replicas. Moreover, instead of ignoring low-criticality tasks or selectively executing them without any guaranteed service level in overrun mode, it mathematically explores the minimum achievable service guarantee for each low-criticality task in different execution modes, i.e., normal, fault-occurrence, overrun and critical operation modes. We develop novel unified demand bound functions (DBF), along with a DVFS method based on the proposed DBF analysis. Our experimental results show that LETR-MC provides up to 59 percent (24 percent on average) energy saving, and significantly improves the service levels of low-criticality tasks compared to the state-of-the-art schemes.
Sepideh Safari, Mohsen Ansari, Ghazal Ershadi, Shaahin Hessabi
IEEE Trans. Parallel Distributed Syst.1