EDBT 2026 Demo / reviewers in the wild / expert
Siamak Mohammadi
dblp:00/1192
· DBLP profile ↗
43ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-1515-7281ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 5 since 2021Software engineering, systems software and programming languages · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive urgency-based real-time task scheduling in ADAS systems
Mahdi Seyfipoor, Sayyed Muhammad Jaffry, Siamak Mohammadi |
Sci. Comput. Program. | 3 |
| 2025 | Efficient quantized transformer for atrial fibrillation detection in cross-domain datasets
Maedeh H. Toosi, Mahdi Mohammadi-nasab, Siamak Mohammadi, Mostafa E. Salehi |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | SRLL: Improving Security and Reliability with User-Defined Constraint-Aware Logic LockingabstractAs chip fabrication costs rise, designers have shifted to a fabless and outsourced development model which opens up the possibility for IP piracy. To address these challenges, logic locking methods modify designs to limit functionality to authorized users that present a valid secret key. However, existing techniques often face limitations in resilience against advanced attacks and do not provide solutions to achieve user-defined constraints and goals. In this article, we propose SRLL, a user-defined constraint-aware logic locking technique that aims to improve the security and reliability of hardware designs. SRLL bridges the gap between exact and approximate attacks and allows the user to balance the resiliency against satisfiability-based, machine-learning-based, and constant propagation attacks while securing design constraints provided by the user. To enable this, we limit the locking functions to the non-critical path components and insert key gates at specific nodes, introducing a new set of critical parameters specifically designed to prevent target attacks. Finally, we obfuscate the netlist to hide inserted key gates and locking functions. Results show that SRLL maintains strong resiliency by exponentially increasing the required number of distinguishing input patterns, the complexity of finding these patterns, and adding sufficient structural complexity to the design. We evaluate SRLL using ISCAS ’85, MCNC ’91, and ITC ’99 benchmarks, demonstrating resiliency with low overhead against modern attacks, including SAT, AppSAT, OMLA, SAIL, and SCOPE. Mona Hashemi, Siamak Mohammadi, Trevor E. Carlson |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2023 | Model Checking of Hyperledger Fabric Smart ContractsabstractConducting interactions between shared-purpose organizations that are not entirely trustworthy of each other without centralized oversight is an idea that emerged with the advent of private blockchains such as Hyperledger Fabric and its smart contracts. It is critical to check contracts to ensure their proper functionality, as organizations may collaborate with competitors. Due to the new architecture of Hyperledger Fabric, tools in this area are limited. To formally verify the source code of contracts, we mapped Fabric contract concepts into the Rebeca modeling language. Rebeca is an actor-based language that enables the modeling of concurrent and distributed systems and is supported by a model checking tool, Afra. We have identified vulnerabilities such as deadlock and starvation by examining the desired properties. Using the model checking approach, we could debug the code and hence benefit from speeding up the transactions, creating fewer extra blocks, requiring less storage space to store the ledger, and avoiding wasting computing resources. Elmira Ebrahimi, Ehsan Khamespanah, Marjan Sirjani, Siamak Mohammadi |
ETFA | 4 |
| 2022 | High-level Modeling and Verification Platform for Elastic Circuits with Process Variation ConsiderationsabstractIn addition to the advantages of asynchronous circuits, compatibility with synchronous EDA tools is another strength point of synchronous elastic circuits. Synchronous elastic circuits face some challenges, such as process variations that can compromise its performance and functionality, and the multitude of available implementations based on elastic elements’ combinations, meaning that choosing the best combination could not be simple. In this paper, a novel method is introduced to model and verify synchronous elastic circuits in the presence of variations. The model is based on xMAS, which is a new formal modeling paradigm to synthesize, test, and verify circuits and networks. In this method, various elastic elements are modeled and available in the form of a library in xMAS, so the designer can build complicated elastic circuits by combining different elastic elements. Additionally, by translating a high-level xMAS model into a SAN statistical model and using its capabilities, elements’ internal delays will be embedded, which makes the high-level modeling and elastic circuits’ high-resolution time analysis available. Based on the obtained results, elastic circuits are highly capable of tolerating variations. However, this phenomenon could lead to a maximum of 2.35% error in synchronization control units and data in these circuits. Meysam Zaeemi, Siamak Mohammadi |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2022 | Infrastructure Aware Heterogeneous-Workloads Scheduling for Data Center Energy Cost MinimizationabstractA huge amount of energy consumption, the cost of this usage and environmental effects have become serious issues for commercial cloud providers. Solar energy is a promising clean energy source, to provide some portion of the Internet data center's (IDC's) energy usage which can reduce environmental effects and total energy costs. Moreover, due to the high energy consumption of the cooling system, considering cooling power in job scheduling can provide efficient solutions to reduce total energy consumption. In this article, we investigate the problem of minimizing the energy cost of an IDC and propose an algorithm which schedules heterogeneous IDC workloads, by considering available renewable energy, cooling subsystem, and electricity rate structure. We evaluate the effectiveness and feasibility of our algorithm using real and synthetic workload traces. The simulation results illustrate how our proposed solution reduces the data center's energy cost by up to 46 percent compared to previous solutions. Moreover, results show that our solution is capable of reducing energy cost of data centers under different weather conditions, and rate structures. Kawsar Haghshenas, Somayyeh Taheri, Maziar Goudarzi, Siamak Mohammadi |
IEEE Trans. Cloud Comput. | 4 |
| 2022 | MAGNETIC: Multi-Agent Machine Learning-Based Approach for Energy Efficient Dynamic Consolidation in Data CentersabstractImproving the energy efficiency of data centers while guaranteeing Quality of Service (QoS), together with detecting performance variability of servers caused by either hardware or software failures, are two of the major challenges for efficient resource management of large-scale cloud infrastructures. Previous works in the area of dynamic Virtual Machine (VM) consolidation are mostly focused on addressing the energy challenge, but fall short in proposing comprehensive, scalable, and low-overhead approaches that jointly tackle energy efficiency and performance variability. Moreover, they usually assume over-simplistic power models, and fail to accurately consider all the delay and power costs associated with VM migration and host power mode transition. These assumptions are no longer valid in modern servers executing heterogeneous workloads and lead to unrealistic or inefficient results. In this paper, we propose a centralized-distributed low-overhead failure-aware dynamic VM consolidation strategy to minimize energy consumption in large-scale data centers. Our approach selects the most adequate power mode and frequency of each host during runtime using a distributed multi-agent Machine Learning (ML) based strategy, and migrates the VMs accordingly using a centralized heuristic. Our Multi-AGent machine learNing-based approach for Energy efficienT dynamIc Consolidation (MAGNETIC) is implemented in a modified version of the CloudSim simulator, and considers the energy and delay overheads associated with host power mode transition and VM migration, and is evaluated using power traces collected from various workloads running in real servers and resource utilization logs from cloud data center infrastructures. Results show how our strategy reduces data center energy consumption by up to 15 percent compared to other works in the state-of-the-art (SoA), guaranteeing the same QoS and reducing the number of VM migrations and host power mode transitions by up to 86 and 90 percent, respectively. Moreover, it shows better scalability than all other approaches, taking less than 0.7 percent time overhead to execute for a data center with 1,500 VMs. Finally, our solution is capable of detecting host performance variability due to failures, automatically migrating VMs from failing hosts and draining them from workload. Kawsar Haghshenas, Ali Pahlevan, Marina Zapater, Siamak Mohammadi, David Atienza 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2021 | THAMON: Thermal-aware High-performance Application Mapping onto Opto-electrical network-on-chip
Meisam Abdollahi, Yasaman Firouzabadi, Fatemeh Dehghani, Siamak Mohammadi |
J. Syst. Archit. | 4 |
| 2020 | Developing Safe Smart ContractsabstractBlockchain is a shared, distributed ledger on which transactions are digitally recorded and linked together. Smart Contracts are programs running on Blockchain and are used to perform transactions in a distributed environment without need for any trusted third party. Since smart contracts are used to transfer assets between contractual parties, their safety and security are crucial and badly written and insecure contracts may result in catastrophe. Actor-based programming is known to solve several problems in building distributed software systems. Moreover, formal verification is a solid technique for developing dependable systems. In this paper, we show how the actor model can be used for modeling, analysis and synthesis of smart contracts. We propose Smart Rebeca as an extension of the actor-based language Rebeca, and use the model checking toolset Afra for verification of smart contracts. We implement a synthesizer to synthesize Solidity programs that run on the Ethereum platform from Smart Rebeca models. We examine the challenges and opportunities of our approach in modeling, formal verification, and synthesis of smart contracts using actors. Sajjad Rezaei, Ehsan Khamespanah, Marjan Sirjani, Ali Sedaghatbaf, Siamak Mohammadi |
COMPSAC | 5 |
| 2020 | Vulnerability assessment of fault-tolerant optical network-on-chips
Meisam Abdollahi, Siamak Mohammadi |
J. Parallel Distributed Comput. | 2 |
| 2020 | Prediction-based underutilized and destination host selection approaches for energy-efficient dynamic VM consolidation in data centers
Kawsar Haghshenas, Siamak Mohammadi |
J. Supercomput. | 2 |
| 2019 | CMV: Clustered Majority Voting Reliability-Aware Task Scheduling for Multicore Real-Time SystemsabstractThis paper proposes a novel reliability-aware hard real-time task scheduling method for multicore systems along with a quantitative reliability model. The proposed method uses a heuristic clustered replication to maintain the desired reliability threshold with both minimum replication overhead and latency increase. It also minimizes intercore communication overhead of tasks. Both single and multiple soft errors are considered in this method. Simulation results show that the efficiency of our proposed approach improves with larger network-on-chip sizes, higher reliability thresholds, and higher number of tolerating errors. The proposed method achieves near optimal replica overhead (up to$\text{7.3}\%$higher than optimal replica overhead) with up to$\text{2500}\%$time complexity improvement compared to exhaustive exploration. Experimental results also show that the feasibility of the proposed method is higher than the conventional replication method up to$\text{9.3}\%$. All experiments are performed on both synthetic random task graphs and PARSEC real application benchmarks. Obtained task mapping solutions with communication volume reduction and near optimal replica overhead impose negligible latency increase (up to$\text{6.3}\%$) in comparison with the space exploration approach. Alireza Namazi, Saeed Safari, Siamak Mohammadi |
IEEE Trans. Reliab. | 3 |
| 2018 | Exploration of approximate multipliers design space using carry propagation free compressorsabstractMany emerging application domains, such as machine learning, can tolerate limited amounts of arithmetic inaccuracy. When designing custom compute accelerators for these domains, hardware designers can explore tradeoffs that sacrifice accuracy in order to reduce area, delay, and/or power consumption. This paper explores the design space of approximate multipliers using a family of approximate compressors as building blocks for the partial product reduction tree. We present a tool that allows the user to specify an allowable level of error tolerance, and returns the minimum area, delay, or power approximate multiplier that provides that level of accuracy. Our experimental results indicate that our proposed compressors generate more accurate and more efficient approximate multipliers than existing state-of-the-art techniques. Sina Boroumand, Hadi Parandeh-Afshar, Philip Brisk, Siamak Mohammadi |
ASP-DAC | 4 |
| 2018 | A Majority-Based Reliability-Aware Task Mapping in High-Performance Homogenous NoC ArchitecturesabstractThis article presents a new reliability-aware task mapping approach in a many-core platform at design time for applications with DAG-based task graphs. The main goal is to devise a task mapping which meets a predefined reliability threshold considering a minimized performance degradation. The proposed approach uses a majority-voting replication technique to fulfill error-masking capability. A quantitative reliability model is also proposed for the platform. Our platform is a homogenous many-core architecture with mesh-based interconnection using traditional deterministic XY routing algorithm. Our iterative approach is applicable to an unlimited number of system fault types. All parts of the platform, including cores, links, and routers, are assumed to be prone to failures. We used the MNLP optimization technique to find the optimal mapping of the presented task graph. Experimental results show that our suggested task mappings not only comply with predefined reliability thresholds but also achieve notable time complexity reduction with respect to exhaustive space exploration. Alireza Namazi, Meisam Abdollahi, Saeed Safari, Siamak Mohammadi |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2018 | Hypervisor and Neighbors' Noise: Performance Degradation in Virtualized EnvironmentsabstractUsers expect isolated performance from rented virtual machines (VMs) in an infrastructure as a service (IaaS) cloud environment. However, this is not happening in todays’ systems because basically VMs are running in a shared environment. In this paper, we study performance degradation in a virtualized environment similar to IaaS clouds using Parsec 2.1 benchmarks. We consider slowdowns caused by hypervisor—hypervisor's noise—as well as co-located VMs—neighbors’ noise. Previous researches did not consider multi-virtual CPU (vCPU) VMs in an overcommitted environments similar to IaaS clouds. Our target system consists of multiple multi-processor VMs running on a commodity chip-multiprocessor by a hypervisor. This configuration is widespread in todays’ IaaS clouds like Amazon EC2. We find that performance degradation in a virtualized environment could be up to$16\times$which is far more than previous findings. Beside shared resources of memory sub-system, blindness of hypervisor's scheduler have large impact on the slowdown and this is contrary to recent researches that mostly blame last-level cache (LLC) contention for performance degradation. After investigating the causes of performance degradation, we provide some ideas that motivate researchers to reduce performance degradation through hardware and software techniques. We also mention some hints that help organizations to see if their applications are ready for the cloud. Seyed Hossein Nikounia, Siamak Mohammadi |
IEEE Trans. Serv. Comput. | 2 |
| 2017 | LORAP: Low-Overhead Power and Reliability-Aware Task Mapping Based on Instruction Footprint for Real-Time ApplicationsabstractThis paper presents a novel power and reliability-aware task mapping approach in many-core platforms for hard realtime applications which is called LORAP. The LORAP contrives a task mapping scenario to meet the predefined reliability threshold ensuring minimum power consumption overhead. It drastically decreases task mapping time complexity and uses slack time of running applications to apply a heuristic DVFS to reduce the power consumption overhead. The quantitative reliability modeling of this paper uses effective failure rate based on instruction footprints of the task which is obtained using a new low complexity AVF calculations algorithm. Proposed novel 4-step iterative approach is applicable to an unlimited number of fault types. All parts of the platform including cores, links, and routers are assumed to be prone to failures. The Mixed Non-Linear Programming (MNLP) optimization technique is used to find the task mapping solution. Results show that LORAP reaches the solution with up to 1005% higher than the exhaustive approach and also it only imposes power consumption up to 18.2% higher than optimal solution. Alireza Namazi, Meisam Abdollahi, Saeed Safari, Siamak Mohammadi |
DSD | 4 |
| 2017 | Cache Energy Management through Dynamic Reconfiguration Approach in Opto-Electrical NoCabstractMulti/Many-core architectures will be the popular platform for future system design. Recent investigations show that the hybrid optical-electrical interconnection network can be an appropriate alternative to the traditional electrical NoC. Undoubtedly, memory wall is one of the most important challenges of multi/many-core systems which can somehow be alleviated thanks to hierarchical memory structure. Cache subsystem plays an essential role in increasing the efficiency of the memory structure. In this paper, after exploring the effect of cache subsystem's parameters in a many-core platform with opto-electrical interconnect, we propose a mechanism to increase the cache energy efficiency. The simulation results of the proposed approach show that after applying our method, the energy efficiency parameter improves by 17% and 23% in SPLASH2 and PARSEC benchmarks, respectively. Saba Jamilan, Meisam Abdollahi, Siamak Mohammadi |
PDP | 3 |
| 2017 | A self-organized load balancing mechanism for cloud computingabstractSummary The growth in computer and networking technologies over the past decades established cloud computing as a new paradigm in information technology. The cloud computing promises to deliver cost‐effective services by running workloads in a large scale data center consisting of thousands of virtualized servers. The main challenge with a cloud platform is its unpredictable performance. A possible solution to this challenge could be load balancing mechanism that aims to distribute the workload across the servers of the data center effectively. In this paper, we present a distributed and scalable load balancing mechanism for cloud computing using game theory. The mechanism is self‐organized and depends only on the local information for the load balancing. We proved that our mechanism converges and its inefficiency is bounded. Simulation results show that the generated placement of workload on servers provides an efficient, scalable, and reliable load balancing scheme for the cloud data center. Copyright © 2016 John Wiley & Sons, Ltd. Hadi Khani, Nasser Yazdani, Siamak Mohammadi |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | A Majority-Based Reliability-Aware Task-Mapping in High-Performance Homogenous NoC ArchitecturesabstractThis paper presents a new reliability-aware task mapping approach in a many-core platform at design time for applications with DAG-based task graphs. The main goal of this approach is to devise a task mapping scenario which meets the predefined reliability threshold ensuring minimum performance degradation. The proposed approach uses majority-voting replication technique to fulfill error-masking capability. A quantitative reliability model is also proposed for the platform. Our platform is a homogenous many-core architecture with mesh-based interconnection using traditional deterministic XY routing algorithm. The novel 3-step iterative approach is applicable to an unlimited number of fault types. All parts of the platform including cores, links and routers are assumed to be prone to failures. We used the MNLP optimization technique to find the optimal mapping of presented task graph. Experimental results show that our suggested task mapping approach not only comply with the predefined reliability threshold but also notable time complexity reduction is achieved with respect to exhaustive space exploration. Alireza Namazi, Meisam Abdollahi, Saeed Safari, Siamak Mohammadi |
DSD | 4 |
| 2016 | Clustering Effects on the Design of Opto-Electrical Network-on-ChipabstractEmerging nanoscale silicon-photonics with its advances in fabrication and integration of on-chip CMOS-compatible optical elements are good news for system designers. Optical Network-on-Chips (ONoCs) could be the next generation of NoCs. On the other hand, hybrid opto-electrical networks may provide higher bandwidth, lower latency and better power dissipation when considering both optical and electrical characteristics on multicore platforms. The cluster-based technique locally connects processing cores through electrical interconnect, while the clusters themselves are connected together through an optical waveguide. The experimental results show that in most benchmark applications, the cluster size of 4 proves to be an appropriate size for optimizing the energy-delay product (EDP) parameter. Meisam Abdollahi, Alireza Namazi, Siamak Mohammadi |
PDP | 3 |
| 2016 | Statistical analysis of asynchronous pipelines in presence of process variation using formal models
Mahdi Mosaffa, Siamak Mohammadi, Saeed Safari |
Integr. | 2 |
| 2015 | A Low-Overhead, Fully-Distributed, Guaranteed-Delivery Routing Algorithm for Faulty Network-on-ChipsabstractThis paper introduces a new, practical routing algorithm, Maze-routing, to tolerate faults in network-on-chips. The algorithm is the first to provide all of the following properties at the same time: 1) fully-distributed with no centralized component, 2) guaranteed delivery (it guarantees to deliver packets when a path exists between nodes, or otherwise indicate that destination is unreachable, while being deadlock and livelock free), 3) low area cost, 4) low reconfiguration overhead upon a fault. To achieve all these properties, we propose Maze-routing, a new variant of face routing in on-chip networks and make use of deflections in routing. Our evaluations show that Maze-routing has 16X less area overhead than other algorithms that provide guaranteed delivery. Our Maze-routing algorithm is also high performance: for example, when up to 5 links are broken, it provides 50% higher saturation throughput compared to the state-of-the-art. Mohammad Fattah, Antti Airola, Rachata Ausavarungnirun, Nima Mirzaei, Pasi Liljeberg, Juha Plosila, Siamak Mohammadi, Tapio Pahikkala, Onur Mutlu, Hannu Tenhunen |
NOCS | 7 |
| 2015 | A Clustered GALS NoC Architecture with Communication-Aware MappingabstractAs processors migrate to multi- and many-core architectures, the role of the communication network becomes more important. Efficient communication architecture can drastically improve overall system performance. Taking into account the application behavior can facilitate system-level solutions that manage the communication cost. To address this issue, we propose a Clustered Globally Asynchronous Locally Synchronous Network-on-Chip (C-GALS NoC) communication architecture. C-GALS NoC is composed of local, synchronous clusters and a global asynchronous network. Additionally, we propose a cluster based communication-aware mapping algorithm (CAM) for mapping the application tasks to the C-GALS NoC, while minimizing the communication cost. The synergy of the C-GLAS NoC and the CAM algorithm results in a system-level mechanism that, according to our results, provides up to 2x and 3x, in performance and power improvement, respectively, in comparison with a regular GALS NoC. Finally, we demonstrate that C-GALS NoC is standard-cell compatible by synthesizing it using Design Compiler. Kazem Cheshmi, Siamak Mohammadi, Daniel Versick, Djamshid Tavangarian, Jelena Trajkovic |
PDP | 2 |
| 2015 | Variation-aware approaches with power improvement in digital circuits
Mohammad Mirzaei, Mahdi Mosaffa, Siamak Mohammadi |
Integr. | 3 |
| 2015 | Architecture Support for Tightly-Coupled Multi-Core Clusters with Shared-Memory HW AcceleratorsabstractCoupling processors with acceleration hardware is an effective manner to improve energy efficiency of embedded systems. Many-core is nowadays a dominating design paradigm for SoCs, which opens new challenges and opportunities for designing HW blocks. Exploring acceleration solutions that naturally fit into well-established parallel programming models and that can be incrementally added on top of existing parallel applications is thus extremely important. In this paper we focus on tightly-coupled multi-core cluster architectures, representative of the basic building block of the most recent many-cores, and we enhance it with dedicated HW processing units (HWPU). We propose an architecture where the HWPUs share the same L1 data memory through which processors also communicate, implementing azero-copycommunication model. High-level synthesis (HLS) tools are used to generate HW blocks, then a custom wrapper interfaces the latter to the tightly coupled cluster. We validate our proposal on RTL models, running both synthetic workload and real applications. Experimental results demonstrate that on average our solution provides nearly identical performance to traditional private-memory coarse-grained accelerators, but it achieves up to 32 percent better performance/area/watt and it requires only minimal modifications to legacy parallel codes. Masoud Dehyadegari, Andrea Marongiu, Mohammad Reza Kakoee, Siamak Mohammadi, Nasser Yazdani, Luca Benini |
IEEE Trans. Computers | 4 |
| 2015 | Gem5v: a modified gem5 for simulating virtualized systems
Seyed Hossein Nikounia, Siamak Mohammadi |
J. Supercomput. | 2 |
| 2013 | Power and Variability Improvement of an Asynchronous Router Using Stacking and Dual-Vth ApproachesabstractBelow 45nm technology, process variation causes the occurrence of unpredictable characteristics in fabricated transistors. In this paper, a platform has been developed to examine Die-to-Die process and environment variations impacts on power and delay of network-on-chip routers. As a benchmark an asynchronous router will be considered. To reduce power, Power Delay Product (PDP) and variability of this router, three approaches, namely Suitable Sizing, Stacking and Dual-Vth are proposed. By using Suitable Sizing and applying Stacking approaches on input ports of a particular router configuration, power and PDP are reduced by 41.37% and 39.31%, respectively, for 3.63% delay increase only. Simultaneous use of Dual-Vth and Suitable Sizing approaches in one of the router configurations causes the reduction of power and PDP by 27.74% and 26.54%, respectively, for a delay overhead of 1.74%. Our proposed approaches reduce the router variability to some parameters variation such as Vdd, Vth, and PMOS and NMOS transistors length. Mohammad Mirzaei, Mahdi Mosaffa, Siamak Mohammadi, Jelena Trajkovic |
DSD | 3 |
| 2013 | Modeling symmetrical independent gate FinFET using predictive technology modelabstractPredicting MOSFET models plays a pivotal role in circuit design and its optimization. Independent Gate FinFETs (IGFinFET) are interesting for designers as they are more flexible than Common Multi-Gate FinFETs (CMGFinFET) in digital circuit design. In this work, we implement a model for symmetrical IGFinFET using CMGFinFET model based on Multi-Gate Predictive Technology Model (PTM-MG). This model has been developed from TCAD IGFinFET, based on previously published experimental results of CMG-FinFET. Different basic gates in SG (shorted gate), LP (low power), IG (low area), and IG/LP modes have been designed using the implemented model. For LP, IG, and IG/LP NAND gates, the leakage power is reduced by 89%, 26%, and 67%, respectively in comparison to SG. To show that our model does not have any convergence problem for large circuits, we used ISCAS'85 benchmark suite. The results show that for independent gate in high performance PTM-MG library, on average we can save up to 24% in the number of transistors and lower the total power by 42%. Mohammad Yousef Zarei, Reza Asadpour, Siamak Mohammadi, Ali Afzali-Kusha, Razi Seyyedi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2013 | Quota setting router architecture for quality of service in GALS NoCabstractNetwork on Chip (NoC) is a new communication paradigm for emerging multi- and many-core architectures. Despite major benefits, like scalability and power efficiency, it suffers from lack of guaranteed bounded latency. Many contemporary applications, like multimedia and real-time applications, require such a guarantee. The growth of these applications in embedded systems emphasizes the need for guaranteed services in NoCs. Additionally, increasing numbers of cores in NoCs highlights the clock distribution issue. Globally asynchronous locally synchronous (GALS) NoC architectures propose to solve this issue through using asynchronous routers to connect synchronous blocks. This paper presents a novel approach for guaranteed service in a GALS NoC by using router with set port quota. We propose a novel router architecture which facilitates guaranteed latency for accessing shared media. Our simulations show up to 39% improvement in latency, with a negligible (up to 5%) power overhead. Kazem Cheshmi, Mohammadreza Soltaniyeh, Siamak Mohammadi, Jelena Trajkovic |
RSP | 3 |
| 2013 | Distributed fair DRAM scheduling in network-on-chips architecture
Masoud Dehyadegari, Siamak Mohammadi, Nasser Yazdani |
J. Syst. Archit. | 2 |
| 2011 | Mutant Fault Injection in Functional Properties of a Model to Improve Coverage MetricsabstractThis paper proposes integrating mutation analysis into model checking to improve coverage metrics of digital circuits. In contrast to traditional mutation testing where mutant faults are generated and injected into the code description of the model, we apply a series of newly defined mutation operators directly to the model properties rather than to the model code. We claim that any mutant properties that are generated from the initial properties and validated by the model checker should be considered as new properties that have been missed during the initial verification procedure. Therefore, adding these newly identified properties to the existing list of properties improves the coverage metric of the formal verification and consequently lead to a more reliable design. Preliminary simulation results of applying this approach to a 4x4 Booth-Multiplier with 6 and 8 initial properties, demonstrates a 40% and 45% coverage improvement respectively compared to the initial coverage metric. Ali Abbasinasab, Mahdi Mohammadi, Siamak Mohammadi, Svetlana N. Yanushkevich, Michael Smith 0002 |
DSD | 3 |
| 2011 | Designing Robust Asynchronous Circuits Based on FinFET TechnologyabstractDouble-gate FinFETs have proved to be a promising alternative for deep sub-micron bulk CMOS. In this paper, we have investigated the feasibility of FinFET transistors in asynchronous design which has gained much attention for its advantages such as absence of clock distribution, process variation aware performance, and robustness. Excellent short-channel characteristic, low leakage power, threshold voltage control and the potential of designing area-efficient circuits are the motivation to employ FinFET transistor in asynchronous circuit design. We have designed three novel FinFET-based asynchronous static C-elements which differ in front gate and back gate connections. They are evaluated in terms of leakage and dynamic power, area, and delay characteristics and compared against bulk CMOS C-element in 32nm technology. With technology scaling, vulnerability of combinational logic to soft errors exponentially increases. In this paper we also examine these C-elements nodes sensitivity against soft errors and propose a robust logic. We show that our proposed robustness method increases robustness of the most sensitive node in Shorted gate (SG) and Low power (LP) C-elements 60 times. A dual rail Muller pipeline has been designed with each kind to evaluate our C-elements and compare them to bulk MOSFET pipeline. Compared to SG, simulation results show that Independent gate (IG) and LP modes are most efficient in area and leakage power respectively and in terms of robustness SG and LP modes show better robustness than IG mode. Fataneh Jafari, Mahdi Mosaffa, Siamak Mohammadi |
DSD | 3 |
| 2010 | A fault-tolerant and congestion-aware routing algorithm for Networks-on-ChipabstractThis paper presents a fault-tolerant routing algorithm for mesh-based Networks-on-Chip (NoC) with faulty links. It is a distributed, adaptive and congestion-aware routing algorithm where only two virtual channels are used for both adaptiveness and fault-tolerance. The proposed routing method has a multilevel fault-tolerance capability and therefore it is capable to tolerate more faulty links in more complicated faulty situations with additional hardware costs. The network performance, fault-tolerance capability and hardware overhead are evaluated through appropriate simulations. The experimental results show that the overall reliability of a Network-on-Chip is significantly enhanced against multiple link failures or partially faulty routers with only a small hardware overhead. Mojtaba Valinataj, Siamak Mohammadi, Juha Plosila, Pasi Liljeberg |
DDECS | 2 |
| 2010 | A fault-aware, reconfigurable and adaptive routing algorithm for NoC applicationsabstractThis paper presents a very low cost routing method to tolerate faulty links and routers in mesh-based Networks-on-Chip (NoC). With reconfigurability, this new algorithm supports irregular topologies caused by faulty components in a network. It concurrently uses both fault and congestion information to route the packets by utilizing only two virtual channels for both fault-tolerance and adaptivity. This method has a multi-level fault-tolerance capability and therefore it is capable to tolerate more faulty components with additional costs. Its performance and overhead are evaluated through appropriate simulations and syntheses. The experimental results show that a significant reliability improvement is achieved against multiple component failures with only a few percent area and power overheads. Mojtaba Valinataj, Siamak Mohammadi |
VLSI-SoC | 2 |
| 2009 | An efficent dynamic multicast routing protocol for distributing traffic in NOCsabstractNowadays, in MPSoCs and NoCs, multicast protocol is significantly used for many parallel applications such as cache coherency in distributed shared-memory architectures, clock synchronization, replication, or barrier synchronization. Among several multicast schemes proposed in on chip interconnection networks, path-based multicast scheme has been proven to be more efficient than the tree-based, and unicast-based. In this paper a low distance path-based multicast scheme is proposed. The proposed method takes advantage of the network partitioning, and utilizing of an efficient destination ordering algorithm. The results in performance, and power consumption show that the proposed method outstands the previous on chip path-based multicasting algorithms. Masoumeh Ebrahimi, Masoud Daneshtalab, Mohammad Hossein Neishaburi, Siamak Mohammadi, Ali Afzali-Kusha, Juha Plosila, Hannu Tenhunen |
DATE | 4 |
| 2009 | A Hazard-Free Delay-Insensitive 4-phase On-Chip Link Using MVCM SignalingabstractIn this paper, we introduce a 3 valued MVCM 4-phase link, where cores at each end of the link use 4-phase dual-rail protocol. The dual-rail N-bit data are encoded onto N + 1 wires on the link, thus reducing the number of interconnects between cores and improving power and crosstalk features. We show that it is impractical to encode a 2-phase dual-rail asynchronous data bit onto one wire using MVCM signaling, which was used in previous works, as it generates an unavoidable hazard at the receiver end. The main advantage of our design is that it does not generate any hazards. To evaluate our claim, we use a simple transmitter and receiver implemented in 130 nm technology. Results show a hazard-free communication over different link lengths in contrast to previous works. Mohammad Fattah, Soodeh Aghli Moghaddam, Siamak Mohammadi |
DSD | 3 |
| 2008 | Architectural Synthesis with Control Data Flow Extraction toward an Asynchronous CAD ToolabstractAsynchronous digital design approach liberates VLSI systems from clock signal and offers potential for low power and high performance design methods. Due to lack of commercial CAD tools, asynchronous circuit design has not been regarded with favor. To alleviate the situation, a SystemC library is developed as an extension to the existing SystemC language to enable asynchronous circuit description at the highest level of abstraction. A tool has been developed which extracts optimized control and data flow graphs from the high level description. Also novel architectural asynchronous synthesis algorithms were proposed to generate optimized asynchronous circuit from the extracted data-flow graphs. The proposed library enables the modeling and designing of efficient asynchronous circuits at a high level without having to deal with details of asynchronous implementation. Extracted structures are produced in well-defined form that can easily be used for synthesis purposes, verification or test generation. And finally proposed synthesis tool produces asynchronous circuits with minimum required resources. Results are given by using some high-level synthesis benchmark circuits. Morteza Damavandpeyma, Siamak Mohammadi |
DSD | 2 |
| 2008 | Generating RTL Synthesizable Code from Behavioral Testbenches for Hardware-Accelerated VerificationabstractHardware Accelerated Simulation is widely used in validation of complicated hardware designs. The process of designing a circuit consists of writing the HDL code, and writing and applying the testbenches to the design. Unfortunately, testbenches are often not synthesizable and cannot be used in hardware accelerated simulation. In this paper we propose a method to convert an existing non-synthesizable testbench to a synthesizable one, and apply it to some case studies to show its effectiveness in the hardware accelerated simulation. Mohammad Reza Kakoee, Mohammad Riazati, Siamak Mohammadi |
DSD | 3 |
| 2007 | A Superior Low Complexity Rate Control AlgorithmabstractIn this paper, a new low complexity Rate-Distortion optimization algorithm has been proposed. The proposed method can be employed with non-convex curves as well as convex curves. The new technique has been used in a hardware implementation of a JPEG2000 encoder. Simulation results indicate that the proposed algorithm is less sensitive to the shape of R-D curves in comparison with current algorithms. Compared to the exact method, our performance degradation is less than 0.43 dB. The low complexity of this algorithm makes it suitable for real time applications and in applications like Digital Cinema that have to process a large number of input data. Alireza Aminlou, Maryam Homayouni, Mohammad Hossein Neishaburi, Siamak Mohammadi |
AICCSA | 4 |
| 2007 | System Level Voltage Scheduling Technique Using UML-RT ModelabstractIn this paper, we present optimized methodology for Intra-task voltage scheduling. Our proposed method gets data flow and control flow of application that represents coloration between different parts of the application at the early stage of design using UML-RT model and decides to schedule processor's voltage. By applying this technique on JPEG encoder system experimental results show reduction in energy consumption by 18-54 % over common Intra-DVS algorithm. Mohammad Hossein Neishaburi, Masoud Daneshtalab, Majid Nabi, Siamak Mohammadi |
AICCSA | 4 |
| 2007 | Optimized Assignment Coverage Computation in Formal Verification of Digital SystemsabstractModel checking thoroughly verifies the design correctness with respect to a specification. When the verification process succeeds, we can only postulate the correctness of the design relative to the given specification. How far can we affirm the verified design implements all the behavior of the desired system? With this regard we need to estimate the completeness of the properties by using some coverage metrics. In this paper, we have proposed a new metric called assignment coverage and an optimized method to overcome the intensive computations required for the multiple transformations among the abstract layers in the verification tool. The proposed coverage computation method provides adequate information to complete the set of properties. Finally, we have applied the proposed metric to some verification benchmark to reveal the effectiveness of this metric in finding undetected coverage holes. Majid Nabi, Hamid Shojaei, Siamak Mohammadi, Zainalabedin Navabi |
ATS | 3 |
| 2007 | Functional Test-Case Generation by a Control Transaction Graph for TLM VerificationabstractTransaction level modeling allows exploring several SoC design architectures leading to better performance and easier verification of the final product. Test cases play an important role in determining the quality of a design. Inadequate test-cases may cause bugs to remain after verification. Although TLM expedites the verification of a hardware design, the problem of having high coverage test cases remains unsettled at this level of abstraction. In this paper, first, in order to generate test-cases for a TL model we present a Control-Transaction Graph (CTG) describing the behavior of a TL Model. A Control Graph is a control flow graph of a module in the design and Transactions represent the interactions such as synchronization between the modules. Second, we define dependent paths (DePaths) on the CTG as test-cases for a transaction level model. The generated DePaths can find some communication errors in simulation and detect unreachable statements concerning interactions. We also give coverage metrics for a TL model to measure the quality of the generated test-cases. Finally, we apply our method on the SystemC model of AMBA-AHB bus as a case study and generate testcases based on the CTG of this model. Mohammad Reza Kakoee, Mohammad Hossein Neishaburi, Siamak Mohammadi |
DSD | 3 |
| 1994 | A new scheme for massively parallel image analysisabstractThis paper presents a new reconfigurable architecture for image analysis: the associative mesh. This architecture provides powerful computational primitives that can apply an associative operator over the connex sets of a graph. These primitives can be easily and efficiently realised in hardware by means of asynchronous operations and are adapted to a large number of image analysis primitives. As an example, the implementation of distance transforms is described. Alain Mérigot, Didier Dulac, Siamak Mohammadi |
ICPR (3) | 3 |