EDBT 2026 Demo / reviewers in the wild / expert
Shaahin Hessabi
dblp:18/731
· DBLP profile ↗
48ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0003-3193-2567ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 1 first-author · 7 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 2Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Provider Caching in Multi-Tier Fog Networks
Ferdous Sharifi, Shaahin Hessabi, Young Choon Lee |
ICFEC | 2 |
| 2026 | A comprehensive survey on multi-GPU systems
Atiyeh Gheibi-Fetrat, Arad Maleki, Sahand Zoufan, Masoud Mohammadi-Lak, Amirsaeed Ahmadi-Tonekaboni, Mahdi Alinejad, Komeil Yahyazadeh, Mohammad Alizadeh, Negar Akbarzadeh, Sina Darabi-moghaddam, Shaahin Hessabi, Hamid Sarbazi-Azad |
Parallel Comput. | 11 |
| 2025 | Poster: Unified Fog Node Utilization for Multiple Content Providers through Cluster-Based Cooperative CachingabstractEfficient content caching is crucial for video streaming services, enhancing user experience and conserving network bandwidth. While traditional Content Delivery Networks (CDNs) address this need to a certain extent, fog computing is emerging as a complementary solution. In particular, this new computing paradigm utilizes nodes positioned between users and the cloud continuum (fog nodes). However, the limited capacity of fog nodes poses a challenge. Previous studies have addressed this challenge by focusing on cooperative caching, considering factors like popularity and user location, yet often overlooked the shared use of fog node storage by multiple content providers (MCPs). This paper introduces CCoFog Caching (CCo-Fog), a cluster-based cooperative content caching strategy that allocates fog node storage capacity among content providers using user clustering and the multi-tier feature of fog networks. It also proposes a content placement algorithm that considers popularity to determine the number of content copies in the network. Evaluation with real-world data shows that CCo-Fog significantly improves latency by 54% and hit ratio by1|6% compared to existing strategies. Ferdous Sharifi, Young Choon Lee, Shaahin Hessabi |
WoWMoM | 3 |
| 2025 | LEC-MiCs: Low-Energy Checkpointing in Mixed-Criticality Multicore SystemsabstractWith the advent of multicore platforms in designing Mixed-Criticality Systems (MCSs), simultaneous management of reliability and energy while guaranteeing an acceptable service level for low-criticality tasks is a crucial challenge. To ensure the reliability of the MCSs against transient faults, fault-tolerant techniques are employed which will increase energy consumption. To mitigate the energy overhead, the Dynamic Voltage and Frequency Scaling (DVFS) technique will be exploited. However, this technique might lead to violating the timing constraints of high-criticality tasks. Therefore, this article presents, for the first time, the low-energy checkpointing technique to guarantee the reliability of multiple preemptive periodic mixed-criticality tasks in a multicore platform. In contrast to the previous works in checkpointing technique which consider a specific number of faults that all the tasks in the system should tolerate, in this article, the number of tolerable faults for each execution section of a task and in each voltage and frequency level is determined through proposed formulas to meet the reliability target based on safety standards. Then, our proposed method determines the number of checkpoints and their non-uniform intervals for the normal and overrun sections of each task to reduce energy consumption, respectively. Moreover, the unified demand bound function (DBF) analysis is proposed for analyzing the schedulability of the task set, where each high-criticality task meets its timing and reliability constraints, and low-criticality tasks execute based on their derived guaranteed periods in each operational mode of the system. Experimental results show that our proposed scheme meets the timing and reliability constraints while at the same time, improving the Quality of Service (QoS) of low-criticality tasks and managing energy consumption with an average of 29.49% and 32.78%, respectively. Sepideh Safari, Shayan Shokri, Shaahin Hessabi, Pejman Lotfi-Kamran |
ACM Trans. Cyber Phys. Syst. | 3 |
| 2025 | Energy-Aware Fault-Tolerant Mapping of Mixed-Criticality Tasks on Heterogeneous MulticoresabstractDue to the different timing requirements of tasks in different criticality modes, it is challenging for Mixed-Criticality Systems (MCSs) designers to use time-redundant fault-tolerant techniques. Checkpointing with rollback recovery can be an effective option for ensuring the reliability of tasks in such systems. Despite the benefits of checkpointing it can impose significant energy overheads. A common approach to overcome this energy consumption is using Dynamic Voltage and Frequency Scaling (DVFS). However, DVFS can lead to missed deadlines for high-criticality tasks. In addition, there is an increasing trend of deploying tasks on heterogeneous multicore platforms. These platforms offer promising opportunities to meet the design requirements of mixed-criticality applications more effectively. In this paper, we propose an energy-aware checkpointing scheme for mixed-criticality tasks on heterogeneous multicore platforms. First, we calculate the timing demand of each task according to DVFS and the time overhead of checkpointing to perform the EY schedulability test in different operational modes of the system. Then we focus on energy-aware mapping of tasks to heterogeneous platforms. Finally, we use DVFS to mitigate the energy overhead of checkpoints. Our experiments show that our scheme can improve schedulability by an average of 16% compared to Little Island First (LIF) mapping, while the energy does not change appreciably. Moreover, it consumes up to 36% (20% on average) less energy compared to Big Island First (BLF) mapping. Amir Hassan Safizadeh, Sepideh Safari, Shayan Shokri, Shaahin Hessabi |
IEEE Trans. Sustain. Comput. | 4 |
| 2025 | Synapse: Synergizing Approximate STT-MRAM and CNN Features for Energy-Efficient AcceleratorsabstractConvolutional Neural Networks (CNNs) require a lot of data and parameters for accuracy. Therefore, CNN accelerators need an efficient and extensive on-chip memory. Nevertheless, traditional memory technologies have limitations for on-chip memory expansion. Spin-Transfer Torque Magnetic Random-Access Memory (STT-MRAM) is a promising Non-Volatile Memory for on-chip memory due to its negligible static power consumption and high density. As STT-MRAMs come with high write overheads, we propose methods to alleviate these drawbacks. By taking advantage of CNN error tolerance, our methods leverage the approximate write technique for STT-MRAM to achieve this goal. First, considering the difference in importance of intra-layer data, we propose a heuristic algorithm to determine the tolerable error probabilities of CNN data to minimize the write energy of STT-MRAM while satisfying accuracy constraints. In the second step, we utilize the asymmetric overheads of the STT-MRAM writing process alongside the asymmetric error tolerance of CNN models against errors of writing logical ’0’ and ’1’ to improve the proposed method. Under the maximum 1% accuracy drop constraint, the proposed method reduces on-chip memory total energy by 36.7% on average. Pouya Toutounchian, Shaahin Hessabi |
IEEE Trans. Sustain. Comput. | 2 |
| 2024 | Cross-core Data Sharing for Energy-efficient GPUsabstractGraphics Processing Units (GPUs) are the accelerator of choice in a variety of application domains, because they can accelerate massively parallel workloads and can be easily programmed using general-purpose programming frameworks such as CUDA and OpenCL. Each Streaming Multiprocessor (SM) contains an L1 data cache (L1D) to exploit the locality in data accesses. L1D misses are costly for GPUs for two reasons. First, L1D misses consume a lot of energy as they need to access the L2 cache (L2) via an on-chip network and the off-chip DRAM in case of L2 misses. Second, L1D misses impose performance overhead if the GPU does not have enough active warps to hide the long memory access latency. We observe that threads running on different SMs share 55% of the data they read from the memory. Unfortunately, as the L1Ds are in the non-coherent memory domain, each SM independently fetches data from the L2 or the off-chip memory into its L1D, even though the data may be currently available in the L1D of another SM. Our goal is to service L1D read misses via other SMs, as much as possible, to cut down costly accesses to the L2 or the off-chip DRAM. To this end, we propose a new data-sharing mechanism, called Cross-Core Data Sharing (CCDS) . CCDS employs a predictor to estimate whether the required cache block exists in another SM. If the block is predicted to exist in another SM’s L1D, then CCDS fetches the data from the L1D that contain the block. Our experiments on a suite of 26 workloads show that CCDS improves average energy and performance by 1.30× and 1.20×, respectively, compared to the baseline GPU. Compared to the state-of-the-art data-sharing mechanism, CCDS improves average energy and performance by 1.37× and 1.11×, respectively. Hajar Falahati, Mohammad Sadrosadati, Qiumin Xu, Juan Gómez-Luna, Banafsheh S. Latibari, Hyeran Jeon, Shaahin Hessabi, Hamid Sarbazi-Azad, Onur Mutlu, Murali Annavaram, Massoud Pedram |
ACM Trans. Archit. Code Optim. | 7 |
| 2024 | A Robust Heterogeneous Offloading Setup Using Adversarial TrainingabstractDeep Neural Networks (DNNs) are very resource-demanding at inference time. Hence, one needs to be able to offload the model execution on the cloud as a solution. The problem is that we should use the same model on both resource-constrained devices and cloud sides. On the other hand, adversarial robustness is one of the main issues in many real-world applications, such as autonomous driving, where one desires model stability under imperceptible but adversarial input perturbations. However, adversarial training (AT) requires access to the actual model architecture and weights during the training. In our setup, two different deep models (suitable for each side) are broken into several blocks. Then, we select a combination of blocks to perform the computation according to the constraints in the inference time, and each block is executed on its respective side. Moreover, we propose a novel modified AT method that can virtually train all the mentioned blocks collectively. Rigorous evaluations of our method on CIFAR-10 and CIFAR-100 show that the proposed AT is effective in making the models robust under various offloading scenarios. Furthermore, we show that the more blocks of the large network are present in the selected model, the higher the final accuracy. To the best of our knowledge, our method is the first one, in which a heterogeneous offloading scheme under adversarial robustness is investigated. Mahdi Amiri, Mohammad H. Rohban, Shaahin Hessabi |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | Mobility-Aware Fog Offloading
Ferdous Sharifi, Ali Rasaii, Melika Honarmand, Shaahin Hessabi, Young Choon Lee |
APNOMS | 4 |
| 2022 | Power-Aware Checkpointing for Multicore Embedded SystemsabstractIncreasing the number of cores integrated on a single chip offers a great potential for the implementation of fault-tolerant techniques to achieve high reliability in real-time embedded systems. Checkpointing with rollback-recovery is a well-established technique to tolerate transient faults in multicore platforms. To consider the worst-case fault occurrence scenario, checkpointing technique requires to re-execute some parts of the tasks, and that might lead to simultaneous execution of task parts with high power consumptions, which eventually might result in a peak power increase beyond the thermal design power (TDP). Exceeding TDP can elevate on-chip temperatures beyond safe limits, and thereby triggering countermeasures that throttle down the voltage and frequency levels or power gate the cores. Such countermeasures might lead to violating task deadlines and degrading the system's reliability. To avoid such severe scenarios, it is inevitable to consider the impact of applying fault-tolerant techniques on the power consumption and prevent violating the power constraint of the chip, i.e., TDP. This paper presents for the first time, a peak-power-aware checkpointing (PPAC) technique that tolerates a given number of faults,k, while at the same time meets the power constraint in hard real-time embedded systems. To do this, our proposed technique (PPAC) adjusts the timing of the checkpoints, which have lower power consumption than the tasks to the execution time points that have power spikes beyond TDP. Moreover, PPAC exploits the available slack times on the cores to delay the execution of some tasks to avoid the remaining power spikes beyond TDP, which could not be mitigated by solely adjusting checkpoints. To evaluate our technique, we extend the state-of-the-art system-level simulator, gem5, with the state-of-the-art checkpointing module in Linux. Our experimental results show that our proposed technique is able to tolerate a given number of faults without exceeding the timing and power constraints in hard real-time embedded systems. The resulting peak power reduction achieved by our technique compared to state-of-the-art techniques is an average of 23%. Moreover, our technique employs the Dynamic Power Management (DPM) during the slack times resulting at runtime in the case of fault-free scenarios, which provides energy savings with an average of 17.28% and up to 61.1%. Mohsen Ansari, Sepideh Safari, Heba Khdr, Pourya Gohari-Nazari, Jörg Henkel, Alireza Ejlali, Shaahin Hessabi |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2022 | TherMa-MiCs: Thermal-Aware Scheduling for Fault-Tolerant Mixed-Criticality SystemsabstractMulticore platforms are becoming the dominant trend in designing Mixed-Criticality Systems (MCSs), which integrate applications of different levels of criticality into the same platform. A well-known MCS is the dual-criticality system that is composed of low-criticality and high-criticality tasks. The availability of multiple cores on a single chip provides opportunities to employ fault-tolerant techniques, such as N-Modular Redundancy (NMR), to ensure the reliability of MCSs. However, applying fault-tolerant techniques will increase the power consumption on the chip, and thereby on-chip temperatures might increase beyond safe limits. To prevent thermal emergencies, urgent countermeasures, like Dynamic Voltage and Frequency Scaling (DVFS) or Dynamic Power Management (DPM) will be triggered to cool down the chip. Such countermeasures, however, might not only lead to suspending low-criticality tasks, but also it might lead to violating timing constraints of high-criticality tasks. In order to prevent such severe scenarios, it is indispensable to consider a temperature constraint within the scheduling process of fault-tolerant MCSs. Therefore, this paper presents, for the first time, a thermal-aware scheduling scheme for fault-tolerant MCSs, named TherMa-MiCs. In particular, TherMa-MiCs, satisfies the temperature constraint jointly with the timing constraints of the high-criticality tasks, while attempting to maximize the QoS of low-criticality tasks under the predefined constraints. At the same time, a reliability target is satisfied by employing the well-known N-Modular Redundancy (NMR) fault-tolerant technique. Experimental results show that our proposed scheme meets the temperature and timing constraints, while at the same time, improving the QoS of low-criticality tasks, with an average of 44%. Sepideh Safari, Heba Khdr, Pourya Gohari-Nazari, Mohsen Ansari, Shaahin Hessabi, Jörg Henkel |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | REALISM: Reliability-aware energy management in multi-level mixed-criticality systems with service level degradation
Hoora Sobhani, Sepideh Safari, Javad Saber-Latibari, Shaahin Hessabi |
J. Syst. Archit. | 4 |
| 2020 | LESS-MICS: A Low Energy Standby-Sparing Scheme for Mixed-Criticality SystemsabstractMulticore platforms are becoming the dominant trend in mixed-criticality systems (MCSs). Multicores provide great opportunities to realize task-level redundancy for reliability enhancement. However, they may experience limited utility in battery-powered mixed-criticality embedded systems. Hence, joint energy and reliability management is a crucial issue in designing MCSs. In this article, we propose the low energy standby-sparing mechanism in mixed-criticality system (LESS-MICS) scheme, which uses the inherent redundancy of multicores to apply the standby-sparing technique for fault-tolerance. Also, by using the inherent redundancy, the LESS-MICS scheme proposes the Parallelism and Reduction policy that can be applied to any graph traverse algorithm to enhance the schedulability of graph-based mixed-criticality tasks, as well as joint energy and reliability management, and guarantying an acceptable service level for low-criticality tasks in overrun mode. To achieve further energy reduction, we minimize energy through convex optimization, and also propose energy management heuristics which use dynamic voltage and frequency scaling and dynamic power management. We evaluated our scheme under various system configurations. Experiments show that our scheme provides, on average, 24.2% energy reduction compared to state-of-the-art techniques while preserving an acceptable QoS level. Sepideh Safari, Shaahin Hessabi, Ghazal Ershadi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Toward On-chip Network Security Using Runtime Isolation MappingabstractMany-cores execute a large number of diverse applications concurrently. Inter-application interference can lead to a security threat as timing channel attack in the on-chip network. A non-interference communication in the shared on-chip network is a dominant necessity for secure many-core platforms to leverage the concepts of the cloud and embedded system-on-chip. The current non-interference techniques are limited to static scheduling and need router modification at micro-architecture level. Mapping of applications can effectively determine the interference among applications in on-chip network. In this work, we explore non-interference approaches through run-time mapping at software and application level. We map the same group of applications in isolated domain(s) to meet non-interference flows. Through run-time mapping, we can maximize utilization of the system without leaking information. The proposed run-time mapping policy requires no router modification in contrast to the best known competing schemes, and the performance degradation is, on average, 16% compared to the state-of-the-art baselines. Mohammad Sadegh Sadeghi, Siavash Bayat Sarmadi, Shaahin Hessabi |
ACM Trans. Archit. Code Optim. | 3 |
| 2019 | On the Scheduling of Energy-Aware Fault-Tolerant Mixed-Criticality Multicore Systems with Service Guarantee ExplorationabstractAdvancement of Cyber-Physical Systems has attracted attention to Mixed-Criticality Systems (MCSs), both in research and in industrial designs. As multicore platforms are becoming the dominant trend in MCSs, joint energy and reliability management is a crucial issue. In addition, providing guaranteed service level for low-criticality tasks in critical mode is of great importance. To address these problems, we propose “LETR-MC” scheme that simultaneously supports certification, energy management, fault-tolerance, and guaranteed service level in mixed-criticality multicore systems. In this paper, we exploit task-replication to not only satisfy reliability requirements, but also to improve the QoS of low-criticality tasks in overrun situation. Our proposed LETR-MC scheme determines the number of replicas, and reduces the execution time overlap between the primary tasks and replicas. Moreover, instead of ignoring low-criticality tasks or selectively executing them without any guaranteed service level in overrun mode, it mathematically explores the minimum achievable service guarantee for each low-criticality task in different execution modes, i.e., normal, fault-occurrence, overrun and critical operation modes. We develop novel unified demand bound functions (DBF), along with a DVFS method based on the proposed DBF analysis. Our experimental results show that LETR-MC provides up to 59 percent (24 percent on average) energy saving, and significantly improves the service levels of low-criticality tasks compared to the state-of-the-art schemes. Sepideh Safari, Mohsen Ansari, Ghazal Ershadi, Shaahin Hessabi |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2018 | SPONGE: A Scalable Pivot-based On/Off Gating Engine for Reducing Static Power in NoC RoutersabstractDue to high aggregate idle time of Networks-on-Chip (NoCs) routers in practical applications, power-gating techniques have been proposed to combat the ever-increasing ratio of static power. Nevertheless, the sporadic packet arrivals compromise the effectiveness of power-gating by incurring significant latency and energy overhead. In this paper, we propose a Scalable Pivot-based On/Off Gating Engine (SPONGE) which efficiently manages power-gating decisions and routing mechanism by adaptively selecting a small set of powered-on columns of routers and keeping the others in power-gated state. To this end, a router architecture augmented with a novel routing algorithm is proposed in which a packet can traverse powered-off routers without waking them up, and can only turn in predetermined powered-on routers. Experimental results on SPLASH-2 benchmarks demonstrate that, compared to the conventional power-gating method, SPONGE on average not only improves static power consumption by 81.7%, it also improves average packet latency by 63%. Hossein Farrokhbakht, Hadi Mardani Kamali, Natalie D. Enright Jerger, Shaahin Hessabi |
ISLPED | 4 |
| 2018 | DuCNoC: A High-Throughput FPGA-Based NoC Simulator Using Dual-Clock Lightweight Router Micro-ArchitectureabstractOn-chip interconnections play an important role in multi/many-processor systems-on-chip (MPSoCs). In order to achieve efficient optimization, each specific application must utilize a specific architecture, and consequently a specific interconnection network. For design space exploration and finding the best NoC solution for each specific application, a fast and flexible NoC simulator is necessary, especially for large design spaces. In this paper, we present an FPGA-based NoC co-simulator, which is able to be configured via software. In our proposed NoC simulator, entitledDuCNoC, we implement aDual-Clockrouter micro-architecture, which demonstrates 75x$-$350x speed-up against BOOKSIM. Additionally, we implement a two-layer configurable global interconnection in our proposed architecture to (1) reduce virtualization time overhead, (2) make an efficient trade-off between the resource utilization and simulation time of the whole simulator, and especially (3) provide the capability of simulating irregular topologies. Migration of some important sub-modules like traffic generators (TGs) and traffic receptors (TRs) to software side, and implementing a dual-clock context switching in virtualization are other major features of DuCNoC. Thanks to its dual-clock router micro-architecture, as well as TGs and TRs migration to software side, DuCNoC can simulate a 100-node (10$\times$10) non-virtualized or a 2048-node virtualized mesh network on Xilinx Zynq-7000. Hadi Mardani Kamali, Kimia Zamiri Azar, Shaahin Hessabi |
IEEE Trans. Computers | 3 |
| 2017 | SMART: A Scalable Mapping And Routing Technique for Power-Gating in NoC RoutersabstractReducing the size of the technology increases leakage power in Network-on-Chip (NoC) routers drastically. Power-gating, particularly in NoC routers, is one of the most efficient approaches for alleviating the leakage power. Although applying power-gating techniques alleviates NoC power consumption due to high proportion of idleness in NoC routers, since the timing behavior of packets is irregular, even in low injection rates, performance overhead in power-gated routers is significant. In this paper, we present SMART, a Scalable Mapping And Routing Technique, with virtually no area overhead on the network. It improves the irregularity of the timing behavior of packets in order to mitigate leakage power and lighten the imposed performance overhead. SMART employs a special deterministic routing algorithm, which reduces number of packets encounter power-gated routers. It establishes a dedicated path between each source-destination pair to maximize using powered-on routers, which roughly halves the number of wake-ups. Additionally, in order to maximize the efficiency of the proposed routing algorithm, SMART provides an exclusive mapping for each communication task graph. In proposed mapping, all cores should be arranged with a special layout suited for the proposed routing, which helps us to minimize the number of hops. Furthermore, we modify the predictor of conventional power-gating technique to reduce energy overhead of inconsistent wake-ups. Experimental results on SPLASH-2 benchmarks indicate that the proposed technique can save 21.9% of static power, and reduce the latency overhead by 42.9% compared with the conventional power-gating technique. Hossein Farrokhbakht, Hadi Mardani Kamali, Shaahin Hessabi |
NOCS | 3 |
| 2017 | Topology exploration of a thermally resilient wavelength-based ONoC
Melika Tinati, Roshanak Karimi, Somayyeh Koohi, Shaahin Hessabi |
J. Parallel Distributed Comput. | 4 |
| 2016 | AdapNoC: A fast and flexible FPGA-based NoC simulatorabstractNetwork on Chip (NoC) is the most common interconnection platform for multiprocessor systems-on-chips (MPSoCs). In order to explore the design space of this platform, we need a high-speed, cycle-accurate, and flexible simulation tool. In this paper, we present AdapNoC, a configurable cycle-accurate FPGA-based NoC simulator, which can be configured via software. A wide range of parameters are configurable in FPGA side of the proposed simulator, and the software side is implemented on an embedded soft-core processor. We transfer some parts of simulator, such as Traffic Generators (TGs) and Traffic Receptors (TRs), to software side without any degradation in simulation speed. Moreover, we implement a dual-clock architecture as an innovation in virtualization methodology, which is also capable to share idle time-slots, which helps not only simulate bigger NoCs, but also reduce simulation time drastically. Also, by employing a traffic aggregator architecture, AdapNoC provides table-based adaptive routing algorithm as a configurable parameter in router microarchitecture. We evaluate simulation time of AdapNoC by using Xilinx Virtex-6 XC6VLX240T, and demonstrate 53x–180x speed-up against BOOKSIM. Also, due to our proposed virtualization, and TGs and TRs migration to software side, we can implement a 64-node non-virtualized or a 1024-node virtualized mesh network in only %72 of Xilinx Virtex-6 XC6VLX240T resources. Hadi Mardani Kamali, Shaahin Hessabi |
FPL | 2 |
| 2016 | TooT: an efficient and scalable power-gating method for NoC routersabstractWith the advent in technology and shrinking the transistor size down to nano scale, static power may become the dominant power component in Networks-on-Chip (NoCs). Powergating is an efficient technique to reduce the static power of under-utilized resources in different types of circuits. For NoC, routers are promising candidates for power gating, since they present high idle time. However, routers in a NoC are not usually idle for long consecutive cycles due to distribution of resources in NoC and its communication-based nature, even in low network utilizations. Therefore, power-gating loses its efficiency due to performance and power overhead of the packets that encounter powered-off routers. In this paper, we propose Turn-on on Turn (TooT) which reduces the number of wake-ups by leveraging the characteristics of deterministic routing algorithms and mesh topology. In the proposed method, we avoid powering a router on when it forwards a straight packet or ejects a packet, i.e., a router is powered on only when either a packet turns through it or its associated node injects a packet. Experimental results on PARSEC benchmarks demonstrate that, compared with the conventional power-gating, the proposed method improves static power and performance by 57.9% and 35.3%, respectively, at the cost of a negligible area overhead. Hossein Farrokhbakht, Mohammadkazem Taram, Behnam Khaleghi, Shaahin Hessabi |
NOCS | 4 |
| 2015 | Low Energy yet Reliable Data Communication Scheme for Network-on-ChipabstractIn this paper, a low energy yet reliable communication scheme for network-on-chip is suggested. To reduce the communication energy consumption, we invoke low-swing signals for transmitting data, as well as data encoding techniques, for minimizing both self and coupling switching capacitance activity factors. To maintain the communication reliability of communication at low-voltage swing, an error control coding (ECC) technique is exploited. The decision about end-to-end or hop-to-hop ECC schemes and the proper number of detectable errors are determined through high-level mathematical analysis on the energy and reliability characteristics of the techniques. Based on the analysis, the extended single error correction double-error detecting end-to-end coding technique with three bits of error detection is used in the network layer. For minimization of the self and coupling switching capacitance activity factors, the odd, even, full invert scheme is employed in the data link layer. This coding has an inherent error detection probability for the flits, which is exploited in the suggested technique. The efficiency of the scheme is studied by using both synthetic and real traffic scenarios. The study reveals savings of up to 43% and 58%, for power dissipation and energy consumption, respectively, without any significant performance degradation and overhead in the network interface. Nima Jafarzadeh, Maurizio Palesi, Saeedeh Eskandari, Shaahin Hessabi, Ali Afzali-Kusha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2015 | Power-efficient prefetching on GPGPUs
Hajar Falahati, Shaahin Hessabi, Mania Abdi, Amirali Baniasadi |
J. Supercomput. | 2 |
| 2014 | QuT: A low-power optical Network-on-ChipabstractTo enable the adoption of optical Networks-on-Chip (NoCs) and allow them to scale to large systems, they must be designed to consume less power and energy. Therefore, optical NoCs must use a small number of wavelengths, avoid excessive insertion loss and reduce the number of microring resonators. We propose the Quartern Topology (QuT), a novel low-power all-optical NoC. We also propose a deterministic wavelength routing algorithm based on Wavelength Division Multiplexing that allows us to reduce the number of wavelengths and microring resonators in optical routers. The key advantages of QuT network are simplicity and lower power consumption. We compare QuT against three alternative all-optical NoCs: optical Spidergon, λ-router and Corona under different synthetic traffic patterns. QuT demonstrates good scalability with significantly lower power and competitive latency. Our optical topology reduces power by 23%, 86.3% and 52.7% compared with 128-node optical Spidergon, λ-router and Corona, respectively. Parisa Khadem Hamedani, Natalie D. Enright Jerger, Shaahin Hessabi |
NOCS | 3 |
| 2014 | All-Optical Wavelength-Routed Architecture for a Power-Efficient Network on ChipabstractIn this paper, we propose a new architecture for nanophotonic Networks on Chip (NoC), named 2D-HERT, which consists of optical data and control planes. The proposed data plane is built upon a new topology and all-optical switches that passively route optical data streams based on their wavelengths. Utilizing wavelength routing method, the proposed deterministic routing algorithm, and Wavelength Division Multiplexing (WDM) technique, the proposed data plane eliminates the need for optical resource reservation at the intermediate nodes. For resolving end-point contention, we propose an all-optical request-grant arbitration architecture which reduces optical losses compared to the alternative arbitration schemes. By performing a series of simulations, we study the efficiency of the proposed architecture, its power and energy consumption, and the data transmission delay. Moreover, we compare the proposed architecture with electrical NoCs and alternative optical NoCs under various synthetic traffic patterns. Averaged across different traffic patterns, 2D-HERT reduces data transmission delay by 24, 15, 18, 4, and 70 percent and achieves average per-packet power reduction of 58, 47, 52, 45, and 95 percent over Phastlane, Firefly, Corona architecture, $(\lambda)$-router, and electrical Torus, respectively. Somayyeh Koohi, Shaahin Hessabi |
IEEE Trans. Computers | 2 |
| 2014 | Towards a scalable, low-power all-optical architecture for networks-on-chipabstractThis article proposes a scalable wavelength-routed optical Network on Chip (NoC) based on the Spidergon topology, named Power-efficient Scalable Wavelength-routed Network-on-chip (PeSWaN). The key idea of the proposed all-optical architecture is the utilization of per-receiver wavelengths in the data network to prevent network contention and the adoption of per-sender wavelengths in the control network to avoid end-point contention. By performing a series of simulations, we study the efficiency of the proposed architecture, its power and energy consumption, and the data transmission delay. Moreover, we compare the proposed architecture with electrical NoCs and alternative ONoC architectures under various traffic patterns. Somayyeh Koohi, Yawei Yin, Shaahin Hessabi, S. J. Ben Yoo |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2012 | ONC3: All-Optical NoC Based on Cube-Connected Cycles with Quasi-DOR AlgorithmabstractThis paper proposes a nanophotonic Network-on-Chip architecture based on the traditional Cube-Connected Cycles topology (CCC), which is named as ONC3. We also suggest a contention-free quasi-Dimension-Order-Routing algorithm for the proposed structure. Compared to the previous 2D layouts, our novel scheme lessens the crosstalk parameter of the insertion loss and consequently, the power consumption. Besides, the router structure is area-efficient. On the other hand, optical destination checking supersedes electrical resource reservation, with utilizing passive wavelength routing method and Wavelength Division Multiplexing scheme, simultaneously. The efficiency of the proposed architecture, in terms of the essential NoC parameters such as delay and power consumption, is investigated through simulation-based experiments and is compared with electrical CCC and Torus topologies, as well as the existing ONoCs such as the optical crossbar and ?-router in various synthetic traffic patterns. Meisam Abdollahi, Mohammad Khavari Tavana, Somayyeh Koohi, Shaahin Hessabi |
DSD | 4 |
| 2012 | Exploration of Temperature Constraints for Thermal Aware Mapping of 3D Networks on ChipabstractThis paper proposes three ILP-based static thermal-aware mapping algorithms for 3D Networks on Chip (NoC) to explore the thermal constraints and their effects on temperature and performance. Through complexity analysis, we show that the first algorithm, an optimal one, is not suitable for 3D NoC. Therefore, we develop two approximation algorithms and analyze their algorithmic complexities to show their proficiency. As the simulation results show, the mapping algorithms that employ direct thermal calculation to minimize the temperature reduce the peak temperature by up to 24% and 22%, for the benchmarks that have the highest communication rate and largest number of tasks, respectively. This comes at the price of a higher power-delay product. This exploration shows that considering power balancing early in the mapping algorithms does not affect the chip temperature. Moreover, it shows that considering the explicit performance constraint in the thermal mapping has no major effect on performance. Parisa Khadem Hamedani, Shaahin Hessabi, Hamid Sarbazi-Azad, Natalie D. Enright Jerger |
PDP | 2 |
| 2012 | Scalable architecture for a contention-free optical network on-chip
Somayyeh Koohi, Shaahin Hessabi |
J. Parallel Distributed Comput. | 2 |
| 2011 | All-optical wavelength-routed NoC based on a novel hierarchical topologyabstractThis paper proposes a novel topology for optical Network on Chip (NoC) architectures with the key advantages of regularity, vertex symmetry, scalability to large scale networks, constant node degree, and simplicity. Moreover, we propose a minimal deterministic routing algorithm for the proposed topology which leads to small and simple photonic routers. Built upon our novel network topology, we present a scalable all-optical NoC, referred to as 2D-HERT, which offers passive routing of optical data streams based on their wavelengths. Utilizing wavelength routing method along with Wavelength Division Multiplexing technique, our proposed optical NoC eliminates the need for electrical resource reservation. We compare performance of the proposed architecture against electrical NoCs and alternative all-optical on-chip architectures under various synthetic traffic patterns. Averaging through different traffic patterns, achieves average perpacket power reduction of 53%, 45%, and 95% over optical crossbar, λ-router, and electrical Torus, respectively. Somayyeh Koohi, Meisam Abdollahi, Shaahin Hessabi |
NOCS | 3 |
| 2011 | Hierarchical opto-electrical on-chip network for future multiprocessor architectures
Somayyeh Koohi, Shaahin Hessabi |
J. Syst. Archit. | 2 |
| 2010 | Scalable Architecture for Wavelength-Switched Optical NoC with Multicasting CapabilityabstractThis paper proposes a novel all-optical router as a building block for a scalable wavelength-switched optical NoC. The proposed optical router, named as AOR, performs passive routing of optical data streams based on their wavelengths. Utilizing wavelength routing method, AOR eliminates the need for electrical resource reservation and the corresponding latency and area overheads. Taking advantage of Wavelength Division Multiplexing (WDM) technique, the proposed architecture is capable of data multicasting, concurrent with unicast data transmission, with high bandwidth and low power dissipation, without imposing noticeable area and latency overheads. Comparing AOR against previously proposed optical routers, we deduce that the proposed router architecture reduces optical insertion loss, electrical power consumption, and number of micro rings, and also improves scalability of the on-chip network. Somayyeh Koohi, Alireza Shafaei, Shaahin Hessabi |
DSD | 3 |
| 2010 | Hierarchical on-Chip Routing of Optical Packets in Large Scale MPSoCsabstractIn this paper, we extract analytical models for data transmission delay, power consumption, and energy dissipation of optical and traditional NoCs. Utilizing extracted models, we compare optical NoC with electrical one for varying values of link length and degree of multiplexing and calculate lower bound limit on the optical link length below which optical on-chip network loses its efficiency. Based on this constraint, we propose a novel hierarchical on-chip network architecture, named as H2NoC, which benefits from optical transmissions in large scale SoCs and overcomes the scalability problem resulted from lower bound limit on the optical link length. Performing a series of simulation-based experiments, we study efficiency of H2NoC along with its power and energy consumption and data transmission delay. Through experimental results, we show that despite slight delay increment in H2NoC compared to non-hierarchical ONoC, later architecture reduces power and energy dissipation of the network. Somayyeh Koohi, Shaahin Hessabi |
PDP | 2 |
| 2009 | Low Power Encoding in NoCs Based on Coupling Transition AvoidanceabstractCoupling capacitances between adjacent wires in on-chip interconnects significantly affect the amount of power consumption in Ultra-Deep-Submicron technologies. On the other hand, the propagation delay across global on chip interconnects has increasingly become a limiting factor in high-speed design. Crosstalk between adjacent links on the bus contributes a significant portion of this delay. Crosstalk noise also affects the integrity of signals. Decreasing the coupling transitions can improve the side effects of crosstalk noise. We propose an algorithm to minimize the coupling activity transition. We also introduce a new solution to fit the proposed algorithm for network-on-chip (NoC) architecture. The experimental results show that the proposed algorithm reduces the power consumption of NoCs up to 27% in an 8-bit bus. Meysam Taassori, Shaahin Hessabi |
DSD | 2 |
| 2009 | Contention-free on-chip routing of optical packetsabstractWe propose a new architecture for on-chip routing of optical packets. The proposed infrastructure, referred to as CONoC, facilitates the development of an all-optical on-chip network and alleviates the role of electrical NoCs. As the first step for designing an all-optical NoC, CONoC resolves packet congestions optically and does not use electrical methods. Utilizing wavelength routing method, wavelength division multiplexing, and path reconfiguration capability in CONoC leads to a contention-free architecture. This architectural advantage along with simple and small photonic router architecture results in simple electrical transactions, reduced setup latency, and high transmission capacity. Moreover, we discuss the proper topology for on-chip optical interconnects. Performing a series of simulation-based experiments, we study the efficiency of CONoC along with its power and energy consumption and data transmission delay. Somayyeh Koohi, Shaahin Hessabi |
NOCS | 2 |
| 2008 | PERMAP: A performance-aware mapping for application-specific SoCsabstractFuture system-on-chip (SoC) designs will need efficient on-chip communication architectures that can provide efficient and scalable data transport among the intellectual properties (IPs). Designing and optimizing SoCs is an increasingly difficult task due to the size and complexity of the SoC design space, high cost of detailed simulation, and several constraints that the design must satisfy. For efficient design of SoCs, an efficient mapping of IPs onto networks-on-chip (NoCs) is highly desirable. Towards this end, we have presented PERMAP, a performance-aware mapping algorithm which maps the IPs onto a generic NoC architecture such that the average communication delay is minimized. This is accomplished by a performance analytical model which can be used for any arbitrary network topology with wormhole routing. The algorithm is used for mapping a video application onto a tile-based NoC and experimental results show that PERMAP is fast and robust. Abbas Eslami Kiasari, Shaahin Hessabi, Hamid Sarbazi-Azad |
ASAP | 2 |
| 2008 | Caspian: A Tunable Performance Model for Multi-core Systems
Abbas Eslami Kiasari, Hamid Sarbazi-Azad, Shaahin Hessabi |
Euro-Par | 3 |
| 2008 | A Markovian Performance Model for Networks-on-ChipabstractNetwork-on-chip (NoC) has been proposed as a solution for addressing the design challenges of future high-performance nanoscale architectures. Thus, it is of crucial importance for a designer to have access to last methods for evaluating the performance of on-chip networks. To this end, we present a Markovian model for evaluating the latency and energy consumption of on-chip networks. We compute the average delay due to path contention, virtual channel and crossbar switch arbitration using a queuing-based approach, which can capture the blocking phenomena of wormhole switching quite accurately. The model is then used to estimate the power consumption of all routers in NoCs. The performance results from the analytical models are validated with those obtained from a synthesizable VHDL-based cycle accurate simulator. Comparison with simulation results indicate that the proposed analytical model is quite accurate and can be used as an efficient design tool by SoC designers. Abbas Eslami Kiasari, Dara Rahmati, Hamid Sarbazi-Azad, Shaahin Hessabi |
PDP | 4 |
| 2007 | An On-Line BIST Technique for Delay Fault Detection in CMOS CircuitsabstractThis paper presents a simulation-based study of the delay fault testing in CMOS logic circuits. A novel built-in self-test (BIST) technique is presented for detecting delay faults in this logic family. This scheme does not need test-pattern generation, and thus can be used for robust on-line testing. Simulation results for area, delay, and power overheads are presented. Elham K. Moghaddam, Shaahin Hessabi |
ATS | 2 |
| 2007 | An Empirical Investigation of Mesh and Torus NoC Topologies Under Different Routing Algorithms and Traffic ModelsabstractNoC is an efficient on-chip communication architecture for SoC architectures. It enables integration of a large number of computational and storage blocks on a single chip. NoCs have tackled the SoCs disadvantages and are scalable. In this paper, we compare two popular NoC topologies, i.e., mesh and torus, in terms of different figures of merit e.g., latency, power consumption, and power/throughput ratio under different routing algorithms and two common traffic models, uniform and hotspot. To the best of our knowledge, this is the first effort in comparing mesh and torus topologies under different routing algorithms and traffic models with respect to their performance and power consumption. Mohammad Mirza-Aghatabar, Somayyeh Koohi, Shaahin Hessabi, Massoud Pedram |
DSD | 3 |
| 2007 | An On-Line BIST Technique for Stuck-Open Fault Detection in CMOS CircuitsabstractThis paper presents a simulation-based study of the stuck-open fault testing in CMOS logic circuits. A novel built-in self-test (BIST) technique is presented for detecting stuck-open faults in these logic families. This scheme does not need test-pattern generation, and thus can be used for robust on-line testing. Simulation results for area, delay, and power overheads are presented. Elham K. Moghaddam, Shaahin Hessabi |
DSD | 2 |
| 2007 | Implementation of a jpeg object-oriented ASIP: a case study on a system-level design methodologyabstractIn this paper, we present a JPEG decoder implemented in our ODYSSEY design methodology. We start with an object-oriented JPEG decoder model. The total operation from modeling to implementation is done automatically by our EDA tool-set in about 10 hours. The resultant system is a JPEG decoder ASIP whose hardware part is implemented on FPGA logic blocks and software part runs on a MicroBlaze processor. This ASIP can be extended by software routines to implement the motion JPEG or MPEG2 decoding algorithms. We implemented our system on ML402 FPGA-based prototype board. Experimental results show that our ASIP implementation is comparable to other approaches while our approach enables quick and easy development of an ASIP using our EDA tool-set and effectively reduces time-to-market. Naser MohammadZadeh, Morteza NajafVand, Shaahin Hessabi, Maziar Goudarzi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2007 | Using on-chip networks to implement polymorphism in the co-design of object-oriented embedded systems
Maziar Goudarzi, Naser MohammadZadeh, Shaahin Hessabi |
J. Comput. Syst. Sci. | 3 |
| 2006 | DotGrid: A .NET-based Infrastructure for Global Grid Computing
Alireza Poshtkohi, Ali Haj Abutalebi, Leila Mahmoudi Ayough, Shaahin Hessabi |
CCGRID | 4 |
| 2006 | A performance and power analysis of WK-Recursive and Mesh Networks for Network-on-ChipsabstractNetwork-on-chip (NoC) has been proposed as an attractive alternative to traditional dedicated wires to achieve high performance and modularity. Power efficiency is one of the most important concerns in NoC architecture design. The choice of network topology is important in designing a low-power and high-performance NoC. In this paper, we propose the use of the WK-recursive networks to be used as the underlying topology in NoC. We have implemented VHDL hardware model of mesh and WK-recursive topologies and measured the latency results using simulation with these implementation. We also propose a novel approach in high level power modeling based on latency for these topologies and show that the power consumption of WK-recursive topology is less than that of the equivalent mesh on a chip. Dara Rahmati, Abbas Eslami Kiasari, Shaahin Hessabi, Hamid Sarbazi-Azad |
ICCD | 3 |
| 2004 | Overhead-Free Polymorphism in Network-on-Chip Implementation of Object-Oriented ModelsabstractWe unify virtual-method despatch (polymorphism implementation) and network packet-routing operations; virtual-method calls correspond to network packets, and network addresses are allocated such that routing the packet corresponds to dispatching the call. As the run-time routing structure is inherent in network-on-chip platforms, this unification implements polymorphism for free. Maziar Goudarzi, Shaahin Hessabi, Alan Mycroft |
DATE | 2 |
| 2003 | Object-Oriented ASIP Design and Synthesis
Maziar Goudarzi, Shaahin Hessabi, Alan Mycroft |
FDL | 2 |
| 1995 | Differential BiCMOS logic circuits: fault characterization and design-for-testabilityabstractMerged Current Switch Logic (MCSL) and Differential Cascode Voltage Switch Logic (DCVSL) are two common structures for differential BiCMOS logic family, that have several potential applications in high-speed VLSI circuits. This paper studies the fault characterization of these BiCMOS circuits. The impact of each possible single defect on the behavior of the circuits is analyzed by simulation. A new class of faults which is unique to differential circuits is identified and its testability is assessed. We propose a design-for-testability method that facilitates testing of this class of faults. Two different realizations for this method are introduced. The impact of this circuit modification on the behavior of the circuit in normal mode is investigated.> Shaahin Hessabi, Mohamed Y. Osman, Mohamed I. Elmasry |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |