Siamak Mohammadi

dblp:00/1192 · DBLP profile ↗
← Back
43ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-1515-7281ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 30 · 5 since 2021Software engineering, systems software and programming languages · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Adaptive urgency-based real-time task scheduling in ADAS systems
Mahdi Seyfipoor, Sayyed Muhammad Jaffry, Siamak Mohammadi
Sci. Comput. Program.3
2025 Efficient quantized transformer for atrial fibrillation detection in cross-domain datasets
Maedeh H. Toosi, Mahdi Mohammadi-nasab, Siamak Mohammadi, Mostafa E. Salehi
Eng. Appl. Artif. Intell.3
2025 SRLL: Improving Security and Reliability with User-Defined Constraint-Aware Logic Locking
abstract
As chip fabrication costs rise, designers have shifted to a fabless and outsourced development model which opens up the possibility for IP piracy. To address these challenges, logic locking methods modify designs to limit functionality to authorized users that present a valid secret key. However, existing techniques often face limitations in resilience against advanced attacks and do not provide solutions to achieve user-defined constraints and goals. In this article, we propose SRLL, a user-defined constraint-aware logic locking technique that aims to improve the security and reliability of hardware designs. SRLL bridges the gap between exact and approximate attacks and allows the user to balance the resiliency against satisfiability-based, machine-learning-based, and constant propagation attacks while securing design constraints provided by the user. To enable this, we limit the locking functions to the non-critical path components and insert key gates at specific nodes, introducing a new set of critical parameters specifically designed to prevent target attacks. Finally, we obfuscate the netlist to hide inserted key gates and locking functions. Results show that SRLL maintains strong resiliency by exponentially increasing the required number of distinguishing input patterns, the complexity of finding these patterns, and adding sufficient structural complexity to the design. We evaluate SRLL using ISCAS ’85, MCNC ’91, and ITC ’99 benchmarks, demonstrating resiliency with low overhead against modern attacks, including SAT, AppSAT, OMLA, SAIL, and SCOPE.
Mona Hashemi, Siamak Mohammadi, Trevor E. Carlson
ACM J. Emerg. Technol. Comput. Syst.2
2023 Model Checking of Hyperledger Fabric Smart Contracts
abstract
Conducting interactions between shared-purpose organizations that are not entirely trustworthy of each other without centralized oversight is an idea that emerged with the advent of private blockchains such as Hyperledger Fabric and its smart contracts. It is critical to check contracts to ensure their proper functionality, as organizations may collaborate with competitors. Due to the new architecture of Hyperledger Fabric, tools in this area are limited. To formally verify the source code of contracts, we mapped Fabric contract concepts into the Rebeca modeling language. Rebeca is an actor-based language that enables the modeling of concurrent and distributed systems and is supported by a model checking tool, Afra. We have identified vulnerabilities such as deadlock and starvation by examining the desired properties. Using the model checking approach, we could debug the code and hence benefit from speeding up the transactions, creating fewer extra blocks, requiring less storage space to store the ledger, and avoiding wasting computing resources.
Elmira Ebrahimi, Ehsan Khamespanah, Marjan Sirjani, Siamak Mohammadi
ETFA4
2022 High-level Modeling and Verification Platform for Elastic Circuits with Process Variation Considerations
abstract
In addition to the advantages of asynchronous circuits, compatibility with synchronous EDA tools is another strength point of synchronous elastic circuits. Synchronous elastic circuits face some challenges, such as process variations that can compromise its performance and functionality, and the multitude of available implementations based on elastic elements’ combinations, meaning that choosing the best combination could not be simple. In this paper, a novel method is introduced to model and verify synchronous elastic circuits in the presence of variations. The model is based on xMAS, which is a new formal modeling paradigm to synthesize, test, and verify circuits and networks. In this method, various elastic elements are modeled and available in the form of a library in xMAS, so the designer can build complicated elastic circuits by combining different elastic elements. Additionally, by translating a high-level xMAS model into a SAN statistical model and using its capabilities, elements’ internal delays will be embedded, which makes the high-level modeling and elastic circuits’ high-resolution time analysis available. Based on the obtained results, elastic circuits are highly capable of tolerating variations. However, this phenomenon could lead to a maximum of 2.35% error in synchronization control units and data in these circuits.
Meysam Zaeemi, Siamak Mohammadi
ACM J. Emerg. Technol. Comput. Syst.2
2022 Infrastructure Aware Heterogeneous-Workloads Scheduling for Data Center Energy Cost Minimization
abstract
A huge amount of energy consumption, the cost of this usage and environmental effects have become serious issues for commercial cloud providers. Solar energy is a promising clean energy source, to provide some portion of the Internet data center's (IDC's) energy usage which can reduce environmental effects and total energy costs. Moreover, due to the high energy consumption of the cooling system, considering cooling power in job scheduling can provide efficient solutions to reduce total energy consumption. In this article, we investigate the problem of minimizing the energy cost of an IDC and propose an algorithm which schedules heterogeneous IDC workloads, by considering available renewable energy, cooling subsystem, and electricity rate structure. We evaluate the effectiveness and feasibility of our algorithm using real and synthetic workload traces. The simulation results illustrate how our proposed solution reduces the data center's energy cost by up to 46 percent compared to previous solutions. Moreover, results show that our solution is capable of reducing energy cost of data centers under different weather conditions, and rate structures.
Kawsar Haghshenas, Somayyeh Taheri, Maziar Goudarzi, Siamak Mohammadi
IEEE Trans. Cloud Comput.4
2022 MAGNETIC: Multi-Agent Machine Learning-Based Approach for Energy Efficient Dynamic Consolidation in Data Centers
abstract
Improving the energy efficiency of data centers while guaranteeing Quality of Service (QoS), together with detecting performance variability of servers caused by either hardware or software failures, are two of the major challenges for efficient resource management of large-scale cloud infrastructures. Previous works in the area of dynamic Virtual Machine (VM) consolidation are mostly focused on addressing the energy challenge, but fall short in proposing comprehensive, scalable, and low-overhead approaches that jointly tackle energy efficiency and performance variability. Moreover, they usually assume over-simplistic power models, and fail to accurately consider all the delay and power costs associated with VM migration and host power mode transition. These assumptions are no longer valid in modern servers executing heterogeneous workloads and lead to unrealistic or inefficient results. In this paper, we propose a centralized-distributed low-overhead failure-aware dynamic VM consolidation strategy to minimize energy consumption in large-scale data centers. Our approach selects the most adequate power mode and frequency of each host during runtime using a distributed multi-agent Machine Learning (ML) based strategy, and migrates the VMs accordingly using a centralized heuristic. Our Multi-AGent machine learNing-based approach for Energy efficienT dynamIc Consolidation (MAGNETIC) is implemented in a modified version of the CloudSim simulator, and considers the energy and delay overheads associated with host power mode transition and VM migration, and is evaluated using power traces collected from various workloads running in real servers and resource utilization logs from cloud data center infrastructures. Results show how our strategy reduces data center energy consumption by up to 15 percent compared to other works in the state-of-the-art (SoA), guaranteeing the same QoS and reducing the number of VM migrations and host power mode transitions by up to 86 and 90 percent, respectively. Moreover, it shows better scalability than all other approaches, taking less than 0.7 percent time overhead to execute for a data center with 1,500 VMs. Finally, our solution is capable of detecting host performance variability due to failures, automatically migrating VMs from failing hosts and draining them from workload.
Kawsar Haghshenas, Ali Pahlevan, Marina Zapater, Siamak Mohammadi, David Atienza 0001
IEEE Trans. Serv. Comput.4
2021 THAMON: Thermal-aware High-performance Application Mapping onto Opto-electrical network-on-chip
Meisam Abdollahi, Yasaman Firouzabadi, Fatemeh Dehghani, Siamak Mohammadi
J. Syst. Archit.4
2020 Developing Safe Smart Contracts
abstract
Blockchain is a shared, distributed ledger on which transactions are digitally recorded and linked together. Smart Contracts are programs running on Blockchain and are used to perform transactions in a distributed environment without need for any trusted third party. Since smart contracts are used to transfer assets between contractual parties, their safety and security are crucial and badly written and insecure contracts may result in catastrophe. Actor-based programming is known to solve several problems in building distributed software systems. Moreover, formal verification is a solid technique for developing dependable systems. In this paper, we show how the actor model can be used for modeling, analysis and synthesis of smart contracts. We propose Smart Rebeca as an extension of the actor-based language Rebeca, and use the model checking toolset Afra for verification of smart contracts. We implement a synthesizer to synthesize Solidity programs that run on the Ethereum platform from Smart Rebeca models. We examine the challenges and opportunities of our approach in modeling, formal verification, and synthesis of smart contracts using actors.
Sajjad Rezaei, Ehsan Khamespanah, Marjan Sirjani, Ali Sedaghatbaf, Siamak Mohammadi
COMPSAC5
2020 Vulnerability assessment of fault-tolerant optical network-on-chips
Meisam Abdollahi, Siamak Mohammadi
J. Parallel Distributed Comput.2
2020 Prediction-based underutilized and destination host selection approaches for energy-efficient dynamic VM consolidation in data centers
Kawsar Haghshenas, Siamak Mohammadi
J. Supercomput.2
2019 CMV: Clustered Majority Voting Reliability-Aware Task Scheduling for Multicore Real-Time Systems
abstract
This paper proposes a novel reliability-aware hard real-time task scheduling method for multicore systems along with a quantitative reliability model. The proposed method uses a heuristic clustered replication to maintain the desired reliability threshold with both minimum replication overhead and latency increase. It also minimizes intercore communication overhead of tasks. Both single and multiple soft errors are considered in this method. Simulation results show that the efficiency of our proposed approach improves with larger network-on-chip sizes, higher reliability thresholds, and higher number of tolerating errors. The proposed method achieves near optimal replica overhead (up to$\text{7.3}\%$higher than optimal replica overhead) with up to$\text{2500}\%$time complexity improvement compared to exhaustive exploration. Experimental results also show that the feasibility of the proposed method is higher than the conventional replication method up to$\text{9.3}\%$. All experiments are performed on both synthetic random task graphs and PARSEC real application benchmarks. Obtained task mapping solutions with communication volume reduction and near optimal replica overhead impose negligible latency increase (up to$\text{6.3}\%$) in comparison with the space exploration approach.
Alireza Namazi, Saeed Safari, Siamak Mohammadi
IEEE Trans. Reliab.3
2018 Exploration of approximate multipliers design space using carry propagation free compressors
abstract
Many emerging application domains, such as machine learning, can tolerate limited amounts of arithmetic inaccuracy. When designing custom compute accelerators for these domains, hardware designers can explore tradeoffs that sacrifice accuracy in order to reduce area, delay, and/or power consumption. This paper explores the design space of approximate multipliers using a family of approximate compressors as building blocks for the partial product reduction tree. We present a tool that allows the user to specify an allowable level of error tolerance, and returns the minimum area, delay, or power approximate multiplier that provides that level of accuracy. Our experimental results indicate that our proposed compressors generate more accurate and more efficient approximate multipliers than existing state-of-the-art techniques.
Sina Boroumand, Hadi Parandeh-Afshar, Philip Brisk, Siamak Mohammadi
ASP-DAC4
2018 A Majority-Based Reliability-Aware Task Mapping in High-Performance Homogenous NoC Architectures
abstract
This article presents a new reliability-aware task mapping approach in a many-core platform at design time for applications with DAG-based task graphs. The main goal is to devise a task mapping which meets a predefined reliability threshold considering a minimized performance degradation. The proposed approach uses a majority-voting replication technique to fulfill error-masking capability. A quantitative reliability model is also proposed for the platform. Our platform is a homogenous many-core architecture with mesh-based interconnection using traditional deterministic XY routing algorithm. Our iterative approach is applicable to an unlimited number of system fault types. All parts of the platform, including cores, links, and routers, are assumed to be prone to failures. We used the MNLP optimization technique to find the optimal mapping of the presented task graph. Experimental results show that our suggested task mappings not only comply with predefined reliability thresholds but also achieve notable time complexity reduction with respect to exhaustive space exploration.
Alireza Namazi, Meisam Abdollahi, Saeed Safari, Siamak Mohammadi
ACM Trans. Embed. Comput. Syst.4
2018 Hypervisor and Neighbors' Noise: Performance Degradation in Virtualized Environments
abstract
Users expect isolated performance from rented virtual machines (VMs) in an infrastructure as a service (IaaS) cloud environment. However, this is not happening in todays’ systems because basically VMs are running in a shared environment. In this paper, we study performance degradation in a virtualized environment similar to IaaS clouds using Parsec 2.1 benchmarks. We consider slowdowns caused by hypervisor—hypervisor's noise—as well as co-located VMs—neighbors’ noise. Previous researches did not consider multi-virtual CPU (vCPU) VMs in an overcommitted environments similar to IaaS clouds. Our target system consists of multiple multi-processor VMs running on a commodity chip-multiprocessor by a hypervisor. This configuration is widespread in todays’ IaaS clouds like Amazon EC2. We find that performance degradation in a virtualized environment could be up to$16\times$which is far more than previous findings. Beside shared resources of memory sub-system, blindness of hypervisor's scheduler have large impact on the slowdown and this is contrary to recent researches that mostly blame last-level cache (LLC) contention for performance degradation. After investigating the causes of performance degradation, we provide some ideas that motivate researchers to reduce performance degradation through hardware and software techniques. We also mention some hints that help organizations to see if their applications are ready for the cloud.
Seyed Hossein Nikounia, Siamak Mohammadi
IEEE Trans. Serv. Comput.2
2017 LORAP: Low-Overhead Power and Reliability-Aware Task Mapping Based on Instruction Footprint for Real-Time Applications
abstract
This paper presents a novel power and reliability-aware task mapping approach in many-core platforms for hard realtime applications which is called LORAP. The LORAP contrives a task mapping scenario to meet the predefined reliability threshold ensuring minimum power consumption overhead. It drastically decreases task mapping time complexity and uses slack time of running applications to apply a heuristic DVFS to reduce the power consumption overhead. The quantitative reliability modeling of this paper uses effective failure rate based on instruction footprints of the task which is obtained using a new low complexity AVF calculations algorithm. Proposed novel 4-step iterative approach is applicable to an unlimited number of fault types. All parts of the platform including cores, links, and routers are assumed to be prone to failures. The Mixed Non-Linear Programming (MNLP) optimization technique is used to find the task mapping solution. Results show that LORAP reaches the solution with up to 1005% higher than the exhaustive approach and also it only imposes power consumption up to 18.2% higher than optimal solution.
Alireza Namazi, Meisam Abdollahi, Saeed Safari, Siamak Mohammadi
DSD4
2017 Cache Energy Management through Dynamic Reconfiguration Approach in Opto-Electrical NoC
abstract
Multi/Many-core architectures will be the popular platform for future system design. Recent investigations show that the hybrid optical-electrical interconnection network can be an appropriate alternative to the traditional electrical NoC. Undoubtedly, memory wall is one of the most important challenges of multi/many-core systems which can somehow be alleviated thanks to hierarchical memory structure. Cache subsystem plays an essential role in increasing the efficiency of the memory structure. In this paper, after exploring the effect of cache subsystem's parameters in a many-core platform with opto-electrical interconnect, we propose a mechanism to increase the cache energy efficiency. The simulation results of the proposed approach show that after applying our method, the energy efficiency parameter improves by 17% and 23% in SPLASH2 and PARSEC benchmarks, respectively.
Saba Jamilan, Meisam Abdollahi, Siamak Mohammadi
PDP3
2017 A self-organized load balancing mechanism for cloud computing
abstract
Summary The growth in computer and networking technologies over the past decades established cloud computing as a new paradigm in information technology. The cloud computing promises to deliver cost‐effective services by running workloads in a large scale data center consisting of thousands of virtualized servers. The main challenge with a cloud platform is its unpredictable performance. A possible solution to this challenge could be load balancing mechanism that aims to distribute the workload across the servers of the data center effectively. In this paper, we present a distributed and scalable load balancing mechanism for cloud computing using game theory. The mechanism is self‐organized and depends only on the local information for the load balancing. We proved that our mechanism converges and its inefficiency is bounded. Simulation results show that the generated placement of workload on servers provides an efficient, scalable, and reliable load balancing scheme for the cloud data center. Copyright © 2016 John Wiley & Sons, Ltd.
Hadi Khani, Nasser Yazdani, Siamak Mohammadi
Concurr. Comput. Pract. Exp.3
2016 A Majority-Based Reliability-Aware Task-Mapping in High-Performance Homogenous NoC Architectures
abstract
This paper presents a new reliability-aware task mapping approach in a many-core platform at design time for applications with DAG-based task graphs. The main goal of this approach is to devise a task mapping scenario which meets the predefined reliability threshold ensuring minimum performance degradation. The proposed approach uses majority-voting replication technique to fulfill error-masking capability. A quantitative reliability model is also proposed for the platform. Our platform is a homogenous many-core architecture with mesh-based interconnection using traditional deterministic XY routing algorithm. The novel 3-step iterative approach is applicable to an unlimited number of fault types. All parts of the platform including cores, links and routers are assumed to be prone to failures. We used the MNLP optimization technique to find the optimal mapping of presented task graph. Experimental results show that our suggested task mapping approach not only comply with the predefined reliability threshold but also notable time complexity reduction is achieved with respect to exhaustive space exploration.
Alireza Namazi, Meisam Abdollahi, Saeed Safari, Siamak Mohammadi
DSD4
2016 Clustering Effects on the Design of Opto-Electrical Network-on-Chip
abstract
Emerging nanoscale silicon-photonics with its advances in fabrication and integration of on-chip CMOS-compatible optical elements are good news for system designers. Optical Network-on-Chips (ONoCs) could be the next generation of NoCs. On the other hand, hybrid opto-electrical networks may provide higher bandwidth, lower latency and better power dissipation when considering both optical and electrical characteristics on multicore platforms. The cluster-based technique locally connects processing cores through electrical interconnect, while the clusters themselves are connected together through an optical waveguide. The experimental results show that in most benchmark applications, the cluster size of 4 proves to be an appropriate size for optimizing the energy-delay product (EDP) parameter.
Meisam Abdollahi, Alireza Namazi, Siamak Mohammadi
PDP3
2016 Statistical analysis of asynchronous pipelines in presence of process variation using formal models
Mahdi Mosaffa, Siamak Mohammadi, Saeed Safari
Integr.2
2015 A Low-Overhead, Fully-Distributed, Guaranteed-Delivery Routing Algorithm for Faulty Network-on-Chips
abstract
This paper introduces a new, practical routing algorithm, Maze-routing, to tolerate faults in network-on-chips. The algorithm is the first to provide all of the following properties at the same time: 1) fully-distributed with no centralized component, 2) guaranteed delivery (it guarantees to deliver packets when a path exists between nodes, or otherwise indicate that destination is unreachable, while being deadlock and livelock free), 3) low area cost, 4) low reconfiguration overhead upon a fault. To achieve all these properties, we propose Maze-routing, a new variant of face routing in on-chip networks and make use of deflections in routing. Our evaluations show that Maze-routing has 16X less area overhead than other algorithms that provide guaranteed delivery. Our Maze-routing algorithm is also high performance: for example, when up to 5 links are broken, it provides 50% higher saturation throughput compared to the state-of-the-art.
Mohammad Fattah, Antti Airola, Rachata Ausavarungnirun, Nima Mirzaei, Pasi Liljeberg, Juha Plosila, Siamak Mohammadi, Tapio Pahikkala, Onur Mutlu, Hannu Tenhunen
NOCS7
2015 A Clustered GALS NoC Architecture with Communication-Aware Mapping
abstract
As processors migrate to multi- and many-core architectures, the role of the communication network becomes more important. Efficient communication architecture can drastically improve overall system performance. Taking into account the application behavior can facilitate system-level solutions that manage the communication cost. To address this issue, we propose a Clustered Globally Asynchronous Locally Synchronous Network-on-Chip (C-GALS NoC) communication architecture. C-GALS NoC is composed of local, synchronous clusters and a global asynchronous network. Additionally, we propose a cluster based communication-aware mapping algorithm (CAM) for mapping the application tasks to the C-GALS NoC, while minimizing the communication cost. The synergy of the C-GLAS NoC and the CAM algorithm results in a system-level mechanism that, according to our results, provides up to 2x and 3x, in performance and power improvement, respectively, in comparison with a regular GALS NoC. Finally, we demonstrate that C-GALS NoC is standard-cell compatible by synthesizing it using Design Compiler.
Kazem Cheshmi, Siamak Mohammadi, Daniel Versick, Djamshid Tavangarian, Jelena Trajkovic
PDP2
2015 Variation-aware approaches with power improvement in digital circuits
Mohammad Mirzaei, Mahdi Mosaffa, Siamak Mohammadi
Integr.3
2015 Architecture Support for Tightly-Coupled Multi-Core Clusters with Shared-Memory HW Accelerators
abstract
Coupling processors with acceleration hardware is an effective manner to improve energy efficiency of embedded systems. Many-core is nowadays a dominating design paradigm for SoCs, which opens new challenges and opportunities for designing HW blocks. Exploring acceleration solutions that naturally fit into well-established parallel programming models and that can be incrementally added on top of existing parallel applications is thus extremely important. In this paper we focus on tightly-coupled multi-core cluster architectures, representative of the basic building block of the most recent many-cores, and we enhance it with dedicated HW processing units (HWPU). We propose an architecture where the HWPUs share the same L1 data memory through which processors also communicate, implementing azero-copycommunication model. High-level synthesis (HLS) tools are used to generate HW blocks, then a custom wrapper interfaces the latter to the tightly coupled cluster. We validate our proposal on RTL models, running both synthetic workload and real applications. Experimental results demonstrate that on average our solution provides nearly identical performance to traditional private-memory coarse-grained accelerators, but it achieves up to 32 percent better performance/area/watt and it requires only minimal modifications to legacy parallel codes.
Masoud Dehyadegari, Andrea Marongiu, Mohammad Reza Kakoee, Siamak Mohammadi, Nasser Yazdani, Luca Benini
IEEE Trans. Computers4
2015 Gem5v: a modified gem5 for simulating virtualized systems
Seyed Hossein Nikounia, Siamak Mohammadi
J. Supercomput.2
2013 Power and Variability Improvement of an Asynchronous Router Using Stacking and Dual-Vth Approaches
abstract
Below 45nm technology, process variation causes the occurrence of unpredictable characteristics in fabricated transistors. In this paper, a platform has been developed to examine Die-to-Die process and environment variations impacts on power and delay of network-on-chip routers. As a benchmark an asynchronous router will be considered. To reduce power, Power Delay Product (PDP) and variability of this router, three approaches, namely Suitable Sizing, Stacking and Dual-Vth are proposed. By using Suitable Sizing and applying Stacking approaches on input ports of a particular router configuration, power and PDP are reduced by 41.37% and 39.31%, respectively, for 3.63% delay increase only. Simultaneous use of Dual-Vth and Suitable Sizing approaches in one of the router configurations causes the reduction of power and PDP by 27.74% and 26.54%, respectively, for a delay overhead of 1.74%. Our proposed approaches reduce the router variability to some parameters variation such as Vdd, Vth, and PMOS and NMOS transistors length.
Mohammad Mirzaei, Mahdi Mosaffa, Siamak Mohammadi, Jelena Trajkovic
DSD3
2013 Modeling symmetrical independent gate FinFET using predictive technology model
abstract
Predicting MOSFET models plays a pivotal role in circuit design and its optimization. Independent Gate FinFETs (IGFinFET) are interesting for designers as they are more flexible than Common Multi-Gate FinFETs (CMGFinFET) in digital circuit design. In this work, we implement a model for symmetrical IGFinFET using CMGFinFET model based on Multi-Gate Predictive Technology Model (PTM-MG). This model has been developed from TCAD IGFinFET, based on previously published experimental results of CMG-FinFET. Different basic gates in SG (shorted gate), LP (low power), IG (low area), and IG/LP modes have been designed using the implemented model. For LP, IG, and IG/LP NAND gates, the leakage power is reduced by 89%, 26%, and 67%, respectively in comparison to SG. To show that our model does not have any convergence problem for large circuits, we used ISCAS'85 benchmark suite. The results show that for independent gate in high performance PTM-MG library, on average we can save up to 24% in the number of transistors and lower the total power by 42%.
Mohammad Yousef Zarei, Reza Asadpour, Siamak Mohammadi, Ali Afzali-Kusha, Razi Seyyedi
ACM Great Lakes Symposium on VLSI3
2013 Quota setting router architecture for quality of service in GALS NoC
abstract
Network on Chip (NoC) is a new communication paradigm for emerging multi- and many-core architectures. Despite major benefits, like scalability and power efficiency, it suffers from lack of guaranteed bounded latency. Many contemporary applications, like multimedia and real-time applications, require such a guarantee. The growth of these applications in embedded systems emphasizes the need for guaranteed services in NoCs. Additionally, increasing numbers of cores in NoCs highlights the clock distribution issue. Globally asynchronous locally synchronous (GALS) NoC architectures propose to solve this issue through using asynchronous routers to connect synchronous blocks. This paper presents a novel approach for guaranteed service in a GALS NoC by using router with set port quota. We propose a novel router architecture which facilitates guaranteed latency for accessing shared media. Our simulations show up to 39% improvement in latency, with a negligible (up to 5%) power overhead.
Kazem Cheshmi, Mohammadreza Soltaniyeh, Siamak Mohammadi, Jelena Trajkovic
RSP3
2013 Distributed fair DRAM scheduling in network-on-chips architecture
Masoud Dehyadegari, Siamak Mohammadi, Nasser Yazdani
J. Syst. Archit.2
2011 Mutant Fault Injection in Functional Properties of a Model to Improve Coverage Metrics
abstract
This paper proposes integrating mutation analysis into model checking to improve coverage metrics of digital circuits. In contrast to traditional mutation testing where mutant faults are generated and injected into the code description of the model, we apply a series of newly defined mutation operators directly to the model properties rather than to the model code. We claim that any mutant properties that are generated from the initial properties and validated by the model checker should be considered as new properties that have been missed during the initial verification procedure. Therefore, adding these newly identified properties to the existing list of properties improves the coverage metric of the formal verification and consequently lead to a more reliable design. Preliminary simulation results of applying this approach to a 4x4 Booth-Multiplier with 6 and 8 initial properties, demonstrates a 40% and 45% coverage improvement respectively compared to the initial coverage metric.
Ali Abbasinasab, Mahdi Mohammadi, Siamak Mohammadi, Svetlana N. Yanushkevich, Michael Smith 0002
DSD3
2011 Designing Robust Asynchronous Circuits Based on FinFET Technology
abstract
Double-gate FinFETs have proved to be a promising alternative for deep sub-micron bulk CMOS. In this paper, we have investigated the feasibility of FinFET transistors in asynchronous design which has gained much attention for its advantages such as absence of clock distribution, process variation aware performance, and robustness. Excellent short-channel characteristic, low leakage power, threshold voltage control and the potential of designing area-efficient circuits are the motivation to employ FinFET transistor in asynchronous circuit design. We have designed three novel FinFET-based asynchronous static C-elements which differ in front gate and back gate connections. They are evaluated in terms of leakage and dynamic power, area, and delay characteristics and compared against bulk CMOS C-element in 32nm technology. With technology scaling, vulnerability of combinational logic to soft errors exponentially increases. In this paper we also examine these C-elements nodes sensitivity against soft errors and propose a robust logic. We show that our proposed robustness method increases robustness of the most sensitive node in Shorted gate (SG) and Low power (LP) C-elements 60 times. A dual rail Muller pipeline has been designed with each kind to evaluate our C-elements and compare them to bulk MOSFET pipeline. Compared to SG, simulation results show that Independent gate (IG) and LP modes are most efficient in area and leakage power respectively and in terms of robustness SG and LP modes show better robustness than IG mode.
Fataneh Jafari, Mahdi Mosaffa, Siamak Mohammadi
DSD3
2010 A fault-tolerant and congestion-aware routing algorithm for Networks-on-Chip
abstract
This paper presents a fault-tolerant routing algorithm for mesh-based Networks-on-Chip (NoC) with faulty links. It is a distributed, adaptive and congestion-aware routing algorithm where only two virtual channels are used for both adaptiveness and fault-tolerance. The proposed routing method has a multilevel fault-tolerance capability and therefore it is capable to tolerate more faulty links in more complicated faulty situations with additional hardware costs. The network performance, fault-tolerance capability and hardware overhead are evaluated through appropriate simulations. The experimental results show that the overall reliability of a Network-on-Chip is significantly enhanced against multiple link failures or partially faulty routers with only a small hardware overhead.
Mojtaba Valinataj, Siamak Mohammadi, Juha Plosila, Pasi Liljeberg
DDECS2
2010 A fault-aware, reconfigurable and adaptive routing algorithm for NoC applications
abstract
This paper presents a very low cost routing method to tolerate faulty links and routers in mesh-based Networks-on-Chip (NoC). With reconfigurability, this new algorithm supports irregular topologies caused by faulty components in a network. It concurrently uses both fault and congestion information to route the packets by utilizing only two virtual channels for both fault-tolerance and adaptivity. This method has a multi-level fault-tolerance capability and therefore it is capable to tolerate more faulty components with additional costs. Its performance and overhead are evaluated through appropriate simulations and syntheses. The experimental results show that a significant reliability improvement is achieved against multiple component failures with only a few percent area and power overheads.
Mojtaba Valinataj, Siamak Mohammadi
VLSI-SoC2
2009 An efficent dynamic multicast routing protocol for distributing traffic in NOCs
abstract
Nowadays, in MPSoCs and NoCs, multicast protocol is significantly used for many parallel applications such as cache coherency in distributed shared-memory architectures, clock synchronization, replication, or barrier synchronization. Among several multicast schemes proposed in on chip interconnection networks, path-based multicast scheme has been proven to be more efficient than the tree-based, and unicast-based. In this paper a low distance path-based multicast scheme is proposed. The proposed method takes advantage of the network partitioning, and utilizing of an efficient destination ordering algorithm. The results in performance, and power consumption show that the proposed method outstands the previous on chip path-based multicasting algorithms.
Masoumeh Ebrahimi, Masoud Daneshtalab, Mohammad Hossein Neishaburi, Siamak Mohammadi, Ali Afzali-Kusha, Juha Plosila, Hannu Tenhunen
DATE4
2009 A Hazard-Free Delay-Insensitive 4-phase On-Chip Link Using MVCM Signaling
abstract
In this paper, we introduce a 3 valued MVCM 4-phase link, where cores at each end of the link use 4-phase dual-rail protocol. The dual-rail N-bit data are encoded onto N + 1 wires on the link, thus reducing the number of interconnects between cores and improving power and crosstalk features. We show that it is impractical to encode a 2-phase dual-rail asynchronous data bit onto one wire using MVCM signaling, which was used in previous works, as it generates an unavoidable hazard at the receiver end. The main advantage of our design is that it does not generate any hazards. To evaluate our claim, we use a simple transmitter and receiver implemented in 130 nm technology. Results show a hazard-free communication over different link lengths in contrast to previous works.
Mohammad Fattah, Soodeh Aghli Moghaddam, Siamak Mohammadi
DSD3
2008 Architectural Synthesis with Control Data Flow Extraction toward an Asynchronous CAD Tool
abstract
Asynchronous digital design approach liberates VLSI systems from clock signal and offers potential for low power and high performance design methods. Due to lack of commercial CAD tools, asynchronous circuit design has not been regarded with favor. To alleviate the situation, a SystemC library is developed as an extension to the existing SystemC language to enable asynchronous circuit description at the highest level of abstraction. A tool has been developed which extracts optimized control and data flow graphs from the high level description. Also novel architectural asynchronous synthesis algorithms were proposed to generate optimized asynchronous circuit from the extracted data-flow graphs. The proposed library enables the modeling and designing of efficient asynchronous circuits at a high level without having to deal with details of asynchronous implementation. Extracted structures are produced in well-defined form that can easily be used for synthesis purposes, verification or test generation. And finally proposed synthesis tool produces asynchronous circuits with minimum required resources. Results are given by using some high-level synthesis benchmark circuits.
Morteza Damavandpeyma, Siamak Mohammadi
DSD2
2008 Generating RTL Synthesizable Code from Behavioral Testbenches for Hardware-Accelerated Verification
abstract
Hardware Accelerated Simulation is widely used in validation of complicated hardware designs. The process of designing a circuit consists of writing the HDL code, and writing and applying the testbenches to the design. Unfortunately, testbenches are often not synthesizable and cannot be used in hardware accelerated simulation. In this paper we propose a method to convert an existing non-synthesizable testbench to a synthesizable one, and apply it to some case studies to show its effectiveness in the hardware accelerated simulation.
Mohammad Reza Kakoee, Mohammad Riazati, Siamak Mohammadi
DSD3
2007 A Superior Low Complexity Rate Control Algorithm
abstract
In this paper, a new low complexity Rate-Distortion optimization algorithm has been proposed. The proposed method can be employed with non-convex curves as well as convex curves. The new technique has been used in a hardware implementation of a JPEG2000 encoder. Simulation results indicate that the proposed algorithm is less sensitive to the shape of R-D curves in comparison with current algorithms. Compared to the exact method, our performance degradation is less than 0.43 dB. The low complexity of this algorithm makes it suitable for real time applications and in applications like Digital Cinema that have to process a large number of input data.
Alireza Aminlou, Maryam Homayouni, Mohammad Hossein Neishaburi, Siamak Mohammadi
AICCSA4
2007 System Level Voltage Scheduling Technique Using UML-RT Model
abstract
In this paper, we present optimized methodology for Intra-task voltage scheduling. Our proposed method gets data flow and control flow of application that represents coloration between different parts of the application at the early stage of design using UML-RT model and decides to schedule processor's voltage. By applying this technique on JPEG encoder system experimental results show reduction in energy consumption by 18-54 % over common Intra-DVS algorithm.
Mohammad Hossein Neishaburi, Masoud Daneshtalab, Majid Nabi, Siamak Mohammadi
AICCSA4
2007 Optimized Assignment Coverage Computation in Formal Verification of Digital Systems
abstract
Model checking thoroughly verifies the design correctness with respect to a specification. When the verification process succeeds, we can only postulate the correctness of the design relative to the given specification. How far can we affirm the verified design implements all the behavior of the desired system? With this regard we need to estimate the completeness of the properties by using some coverage metrics. In this paper, we have proposed a new metric called assignment coverage and an optimized method to overcome the intensive computations required for the multiple transformations among the abstract layers in the verification tool. The proposed coverage computation method provides adequate information to complete the set of properties. Finally, we have applied the proposed metric to some verification benchmark to reveal the effectiveness of this metric in finding undetected coverage holes.
Majid Nabi, Hamid Shojaei, Siamak Mohammadi, Zainalabedin Navabi
ATS3
2007 Functional Test-Case Generation by a Control Transaction Graph for TLM Verification
abstract
Transaction level modeling allows exploring several SoC design architectures leading to better performance and easier verification of the final product. Test cases play an important role in determining the quality of a design. Inadequate test-cases may cause bugs to remain after verification. Although TLM expedites the verification of a hardware design, the problem of having high coverage test cases remains unsettled at this level of abstraction. In this paper, first, in order to generate test-cases for a TL model we present a Control-Transaction Graph (CTG) describing the behavior of a TL Model. A Control Graph is a control flow graph of a module in the design and Transactions represent the interactions such as synchronization between the modules. Second, we define dependent paths (DePaths) on the CTG as test-cases for a transaction level model. The generated DePaths can find some communication errors in simulation and detect unreachable statements concerning interactions. We also give coverage metrics for a TL model to measure the quality of the generated test-cases. Finally, we apply our method on the SystemC model of AMBA-AHB bus as a case study and generate testcases based on the CTG of this model.
Mohammad Reza Kakoee, Mohammad Hossein Neishaburi, Siamak Mohammadi
DSD3
1994 A new scheme for massively parallel image analysis
abstract
This paper presents a new reconfigurable architecture for image analysis: the associative mesh. This architecture provides powerful computational primitives that can apply an associative operator over the connex sets of a graph. These primitives can be easily and efficiently realised in hardware by means of asynchronous operations and are adapted to a large number of image analysis primitives. As an example, the implementation of distance transforms is described.
Alain Mérigot, Didier Dulac, Siamak Mohammadi
ICPR (3)3