Maziar Goudarzi

dblp:40/3527 · DBLP profile ↗
← Back
36ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0002-1272-4589ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 6 first-author · 6 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 InterStellar 2.0: Fine-grained stream-guided HW/SW co-design for multi-channel DRAM performance steering
abstract
The gap between processor speed and memory latency limits system scalability, especially in data-intensive and artificial intelligence workloads where memory-level parallelism and bandwidth efficiency are critical. Prior work ( InterStellar ) showed that hardware/software (HW/SW) co-design can expose program-level access streams to the memory system, enabling more informed memory-controller (MC) scheduling. This work extends that approach to high-bandwidth, multi-channel platforms. We present InterStellar 2.0 , a scalable HW/SW co-design that: (1) supports multi-channel dynamic random-access memory (DRAM) by partitioning stream batches across channels, allowing each channel to operate independently without cross-channel coordination; and (2) introduces fine-grained stream descriptors so software can distinguish distinct access patterns, even within the same data structure. These capabilities improve DRAM locality management and allow the MC to issue future requests efficiently. We evaluate InterStellar 2.0 on an 8-core RISC-V platform across DRAM configurations from 1 to 32 channels. At 32 channels, InterStellar 2.0 improves performance by up to 2.92 × and increases memory bandwidth by up to 2.83 × over a commercial off-the-shelf (COTS) controller. Fine-grained stream tracking alone improves performance by up to 1 . 35 × . Overall, InterStellar 2.0 shows that stream-aware HW/SW co-design is practical, compatible, and scalable for multi-channel memory systems without ISA changes or inter-channel communication.
Abdelrhman Mohamed Abotaleb, Maziar Goudarzi, Tomasz S. Czajkowski, Mohamed Hassan 0002
J. Syst. Archit.2
2026 Enhancing decision-making for software architects: selecting appropriate architectural patterns based on quality attribute requirements
Maryam Gholami, Jafar Habibi, Maziar Goudarzi
Sci. Comput. Program.3
2025 A TRRIP Down Memory Lane: Temperature-Based Re-Reference Interval Prediction For Instruction Caching
abstract
Modern mobile CPU software pose challenges for conventional instruction cache replacement policies due to their complex runtime behavior causing high reuse distance between executions of the same instruction.Mobile code commonly suffers from large amounts of stalls in the CPU frontend and thus starvation of the rest of the CPU resources.Complexity of these applications and their code footprint are projected to grow at a rate faster than available on-chip memory due to power and area constraints, making conventional hardware-centric methods for managing instruction caches to be inadequate.We present a novel software-hardware co-design approach called TRRIP (Temperature-based Re-Reference Interval Prediction) that enables the compiler to analyze, classify, and transform code based on "temperature" (hot/cold), and to provide the hardware with a summary of code temperature information through a well-defined OS interface based on using code page attributes.TRRIP's lightweight hardware extension employs code temperature attributes to optimize the instruction cache replacement policy resulting in the eviction rate reduction of hot code.TRRIP is designed to be practical and adoptable in real mobile systems that have strict feature requirements on both the software and hardware components.TRRIP can reduce the L2 MPKI for instructions by 26.5% resulting in geomean speedup of 3.9%, on top of RRIP cache replacement running mobile code already optimized using PGO.
Henry Kao, Nikhil Sreekumar, Prabhdeep Singh Soni, Ali Sedaghati, Fang Su, Maziar Goudarzi
MICRO7
2024 RoDMap: A Reserve-on-Demand Mapper for Spatially-Configured Coarse-Grained Reconfigurable Arrays
abstract
We propose, implement, and evaluate a novel approach for mapping dataflow graphs (DFGs) onto spatially configured Coarse-Grained Reconfigurable Arrays (CGRAs). The approach tackles mapping failure due to the congestion that arises when more than one routing path uses the same CGRA link. Heuristics are used to identify congestion patterns and “reserve” CGRA processing elements (PEs) around the congestion by preventing them from being used for DFG nodes in a mapping re-attempt. The reserved PEs effectively increases routing resources around the congestion, thereby increasing the likelihood of mapping success. This approach is referred to as reserve-on-demand mapping since PEs are reserved only when congestion exists and is driven by its patterns.
Kyle Zhao Bin Chen, Tarek S. Abdelrahman, Tomasz S. Czajkowski, Maziar Goudarzi
ICPP5
2024 Variant Parallelism: Lightweight Deep Convolutional Models for Distributed Inference on IoT Devices
abstract
Two major techniques are commonly used to meet real-time inference limitations when distributing models across resource-constrained IoT devices: 1) model parallelism (MP) and 2) class parallelism (CP). In MP, transmitting bulky intermediate data (orders of magnitude larger than input) between devices imposes huge communication overhead. Although CP solves this problem, it has limitations on the number of submodels. In addition, both solutions are fault intolerant, an issue when deployed on edge devices. We propose variant parallelism (VP), an ensemble-based deep learning distribution method where different variants of a main model are generated and can be deployed on separate machines. We design a family of lighter models around the original model, and train them simultaneously to improve accuracy over single models. Our experimental results on six common mid-sized object recognition data sets demonstrate that our models can have$5.8\times $–$7.1\times $fewer parameters,$4.3\times $–$31\times $fewer multiply accumulations (MACs), and$2.5\times $–$13.2\times $less response time on atomic inputs compared to MobileNetV2 while achieving comparable or higher accuracy. Our technique easily generates several variants of the base architecture. Each variant returns only$\boldsymbol {2k}$outputs$\boldsymbol {1 \leq k \leq ({\#classes}/{2})}$, representing$\boldsymbol {Top{-} k}$classes, instead of tons of floating point values required in MP. Since each variant provides a full-class prediction, our approach maintains higher availability compared with MP and CP in presence of failure.
Navidreza Asadi, Maziar Goudarzi
IEEE Internet Things J.2
2023 H-Storm: A Hybrid CPU-FPGA Architecture to Accelerate Apache Storm
Hamid Nasiri, Armin Darjani, Nima Kavand, Maziar Goudarzi
J. Grid Comput.4
2022 Infrastructure Aware Heterogeneous-Workloads Scheduling for Data Center Energy Cost Minimization
abstract
A huge amount of energy consumption, the cost of this usage and environmental effects have become serious issues for commercial cloud providers. Solar energy is a promising clean energy source, to provide some portion of the Internet data center's (IDC's) energy usage which can reduce environmental effects and total energy costs. Moreover, due to the high energy consumption of the cooling system, considering cooling power in job scheduling can provide efficient solutions to reduce total energy consumption. In this article, we investigate the problem of minimizing the energy cost of an IDC and propose an algorithm which schedules heterogeneous IDC workloads, by considering available renewable energy, cooling subsystem, and electricity rate structure. We evaluate the effectiveness and feasibility of our algorithm using real and synthetic workload traces. The simulation results illustrate how our proposed solution reduces the data center's energy cost by up to 46 percent compared to previous solutions. Moreover, results show that our solution is capable of reducing energy cost of data centers under different weather conditions, and rate structures.
Kawsar Haghshenas, Somayyeh Taheri, Maziar Goudarzi, Siamak Mohammadi
IEEE Trans. Cloud Comput.3
2021 SVNN: an efficient PacBio-specific pipeline for structural variations calling using neural networks
abstract
BACKGROUND: Once aligned, long-reads can be a useful source of information to identify the type and position of structural variations. However, due to the high sequencing error of long reads, long-read structural variation detection methods are far from precise in low-coverage cases. To be accurate, they need to use high-coverage data, which in turn, results in an extremely time-consuming pipeline, especially in the alignment phase. Therefore, it is of utmost importance to have a structural variation calling pipeline which is both fast and precise for low-coverage data. RESULTS: In this paper, we present SVNN, a fast yet accurate, structural variation calling pipeline for PacBio long-reads that takes raw reads as the input and detects structural variants of size larger than 50 bp. Our pipeline utilizes state-of-the-art long-read aligners, namely NGMLR and Minimap2, and structural variation callers, videlicet Sniffle and SVIM. We found that by using a neural network, we can extract features from Minimap2 output to detect a subset of reads that provide useful information for structural variation detection. By only mapping this subset with NGMLR, which is far slower than Minimap2 but better serves downstream structural variation detection, we can increase the sensitivity in an efficient way. As a result of using multiple tools intelligently, SVNN achieves up to 20 percentage points of sensitivity improvement in comparison with state-of-the-art methods and is three times faster than a naive combination of state-of-the-art tools to achieve almost the same accuracy. CONCLUSION: Since prohibitive costs of using high-coverage data have impeded long-read applications, with SVNN, we provide the users with a much faster structural variation detection platform for PacBio reads with high precision and sensitivity in low-coverage scenarios.
Shaya Akbarinejad, Mostafa Hadadian Nejad Yousefi, Maziar Goudarzi
BMC Bioinform.3
2021 Profit Maximization of Big Data Jobs in Cloud Using Stochastic Optimization
abstract
Reserved instances offered by cloud providers make it possible to reserve resources and computing capacity for a specific period of time. One should pay for all the hours of that time interval; in exchange, the hourly rate is significantly lower than on-demand instances. Reserved Instances can significantly reduce the monetary cost of resources needed to process big data applications in cloud. However, purchases of these instances are non-refundable, and hence, one should be able to estimate the required resources prior to purchase to avoid over-payment. It becomes important especially when the results obtained by big data job has monetary value, such as business intelligence applications. But, estimating the resource demand of big data processing jobs is hard because of numerous factors that affect them such as data locality, data skew, stragglers, internal settings of big data processing framework, interference among instances, instances availability, etc. To maximize the profit of processing such big data jobs in cloud considering fluctuating nature of their resource demand, as well as reserved instances limitations, we propose Reserved Instances Stochastic Allocation (RISA) approach. Using historical traces of resource demand of big data jobs submitted by user, RISA leverages stochastic optimization to determine the amount of resources needed to be reserved for that user to maximize the profit. Our evaluation using real-world traces shows that RISA can increase the net profit by up to 10x, compared to previous approaches. RISA can also find solutions as close as 2 percent to the best possible solution.
Seyed Morteza Nabavinejad, Maziar Goudarzi
IEEE Trans. Cloud Comput.2
2019 IMOS: improved Meta-aligner and Minimap2 On Spark
abstract
BACKGROUND: Long reads provide valuable information regarding the sequence composition of genomes. Long reads are usually very noisy which renders their alignments on the reference genome a daunting task. It may take days to process datasets enough to sequence a human genome on a single node. Hence, it is of primary importance to have an aligner which can operate on distributed clusters of computers with high performance in accuracy and speed. RESULTS: In this paper, we presented IMOS, an aligner for mapping noisy long reads to the reference genome. It can be used on a single node as well as on distributed nodes. In its single-node mode, IMOS is an Improved version of Meta-aligner (IM) enhancing both its accuracy and speed. IM is up to 6x faster than the original Meta-aligner. It is also implemented to run IM and Minimap2 on Apache Spark for deploying on a cluster of nodes. Moreover, multi-node IMOS is faster than SparkBWA while executing both IM (1.5x) and Minimap2 (25x). CONCLUSION: In this paper, we purposed an architecture for mapping long reads to a reference. Due to its implementation, IMOS speed can increase almost linearly with respect to the number of nodes in a cluster. Also, it is a multi-platform application able to operate on Linux, Windows, and macOS.
Mostafa Hadadian Nejad Yousefi, Maziar Goudarzi, Abolfazl S. Motahari
BMC Bioinform.2
2019 Heterogeneous Architectures for Big Data Batch Processing in MapReduce Paradigm
abstract
The amount of digital data produced worldwide is exponentially growing. While the source of this data, collectively known as Big Data, varies from among mobile services to cyber physical systems and beyond, the invariant is their increasingly rapid growth for the foreseeable future. Immense incentives exist, from marketing campaigns to forensics and to research in social sciences, that motivate processing increasingly bigger data so as to extract information and knowledge for the betterment of processes and benefits. Consequently, the need for more efficient computing systems tailored to such big data applications is increasingly intensified. Such custom architectures would expectedly embrace heterogeneity to better match each phase of the computation. In this paper we review state of the art as well as envisioned future large-scale computing architectures customized for batch processing of big data applications in the MapReduce paradigm. We also provide our view of current important trends relevant to such systems, and their impacts on future architectures and architectural features expected to address the needs of tomorrow big data processing in this paradigm.
Maziar Goudarzi
IEEE Trans. Big Data1
2019 Faster MapReduce Computation on Clouds Through Better Performance Estimation
abstract
Processing Big Data in cloud is on the increase. An important issue for efficient execution of Big Data processing jobs on a cloud platform is selecting the best fitting virtual machine (VM) configuration(s) among the miscellany of choices that cloud providers offer. Wise selection of VM configurations can lead to better performance, cost and energy consumption. Therefore, it is crucial to explore the available configurations and opt for the best ones that well suit each MapReduce application. Profiling the given application on all the configurations is costly, time and energy consuming. An alternative is to run the application on a subset of configurations (sample configurations) and estimate its performance on other configurations based on the obtained values by sample configurations. We show that the choice of these sample configurations highly affects accuracy of later estimations. Our Smart Configuration Selection (SCS) scheme chooses better representatives from among all configurations by once-off analysis of given performance figures of the benchmarks so as to increase the accuracy of estimations of missing values, and consequently, to more accurately choose the configuration providing the highest performance. The results show that the SCS choice of sample configurations is very close to the best choice, and can reduce estimation error to 11.58 percent from the original 19.72 percent of random configuration selection. More importantly, using SCS estimations in a makespan minimization algorithm improves the execution time by up to 36.03 percent compared with random sample selection.
Seyed Morteza Nabavinejad, Maziar Goudarzi
IEEE Trans. Cloud Comput.2
2019 SAIR: significance-aware approach to improve QoR of big data processing in case of budget constraint
Hossein Ahmadvand, Maziar Goudarzi
J. Supercomput.2
2018 QoR-aware power capping for approximate big data processing
abstract
To limit the peak power consumption of a cluster, a centralized power capping system typically assigns power caps to the individual servers, which are then enforced using local capping controllers. Consequently, the performance and throughput of the servers are affected, and the runtime of jobs is extended as a result. We observe that servers in big data processing clusters often execute big data applications that have different tolerance for approximate results. To mitigate the impact of power capping, we propose a new power-Capping aware resource manager for Approximate Big data processing (CAB) that takes into consideration the minimum Quality-of-Result (QoR) of the jobs. We use industry-standard feedback power capping controllers to enforce a power cap quickly, while, simultaneously modifying the resource allocations to various jobs based on their progress rate, target minimum QoR, and the power cap such that the impact of capping on runtime is minimized. Based on the applied cap and the progress rates of jobs, CAB dynamically allocates the computing resources (i.e., number of cores and memory) to the jobs to mitigate the impact of capping on the finish time. We implement CAB in Hadoop-2.7.3 and evaluate its improvement over other methods on a state-of-the-art 28-core Xeon server. We demonstrate that CAB minimizes the impact of power capping on runtime by up to 39.4% while meeting the minimum QoR constraints.
Seyed Morteza Nabavinejad, Xin Zhan, Maziar Goudarzi, Sherief Reda
DATE4
2018 A Task-Based Greedy Scheduling Algorithm for Minimizing Energy of MapReduce Jobs
Mostafa Hadadian Nejad Yousefi, Maziar Goudarzi
J. Grid Comput.2
2017 On Reliability-Aware Server Consolidation in Cloud Datacenters
abstract
In the past few years, datacenter (DC) energy consumption has become an important issue in technology world. Server consolidation using virtualization and virtual machine (VM) live migration allows cloud DCs to improve resource utilization and hence energy efficiency. In order to save energy, consolidation techniques try to turn off the idle servers, while because of workload fluctuations, these offline servers should be turned on to support the increased resource demands. These repeated on-off cycles could affect the hardware reliability and wear-and-tear of servers and as a result, increase the maintenance and replacement costs. In this paper we propose a holistic mathematical model for reliability-aware server consolidation with the objective of minimizing total DC costs including energyand reliability costs. In fact, we try to minimize the number of active PMs and racks, in a reliability-aware manner. We formulate the problem as a Mixed Integer Linear Programming (MILP) model which is in form of NP-complete. Finally, we evaluate the performance of our approach in different scenarios using extensive numerical MATLAB simulations.
Amir Varasteh, Farzad Tashtarian, Maziar Goudarzi
ISPDC3
2016 Energy efficiency in cloud-based MapReduce applications through better performance estimation
Seyed Morteza Nabavinejad, Maziar Goudarzi
DATE2
2016 The Memory Challenge in Reduce Phase of MapReduce Applications
abstract
MapReduce has become a popular paradigm for Big Data processing. Each MapReduce Application has two phases: Map and Reduce. Each phase consist of several tasks in a defaulted sequence of processes. It is common place to determine the number of Map tasks equal to the number of data blocks in the input data. However, there is no specific rule for determining the number of Reduce tasks based on the amount of intermediate data generated by Map tasks or the specifications of machines that execute the tasks. Since the Reduce tasks bring the data into memory for processing, this may lead to inefficient execution of application and even application failure because of memory shortage or temporary consumption. In this work, we first evaluate this challenge and show its problematic significance. To address this challenge, we propose a Mnemonic approach. Mnemonic leverages a profiling mechanism to detect the application behavior regarding intermediate data generation. It first decides the amount of memory to be dedicated to each Reduce slot. Then it determines the number of Reduce tasks based on the gathered information through profiling and the decided size of memory for Reduce slots. Experimental results using PUMA benchmark suit indicates that our proposed memory-aware approach can 1) completely remove the likelihood of application failure due to out of memory error and 2) decrease the execution time of Reduce phase up to 58.27, 79.36, and 88.79 percent compared with Memory Oblivious, Fine Grain 1, and Fine Grain 2 approaches, respectively.
Seyed Morteza Nabavinejad, Maziar Goudarzi, Shirin Mozaffari
IEEE Trans. Big Data2
2015 TABEMS: Tariff-Aware Building Energy Management System for Sustainability through Better Use of Electricity
abstract
Smarter use of the renewable energy produced by solar panels reduces the return time of the investment necessary for their installation. This improvement consequently motivates more households to use solar panels so as to not only help protect the environment, but also better use the expensive energy. The difference in tariff prices at different hours of the day is one such opportunity for smarter use of solar electricity: we propose and implement a real-time strategy to more economically use the produced solar electrical energy by forecasting future demand of a few days ahead and by using that energy at the most economical time. Evaluation of the proposed technique in an educational building showed that this scheme improves financial advantage of solar panels by 41% compared with the direct connection of production of solar panels to the grid, or using the stored solar energy completely unawares, hence it can reduce the return time of investment by the same amount. Moreover, since our technique reduces power usage from the utility grid at peak tariff hours, it is one way to move toward a uniform consumption at the suppliers’ level that leads to better use and higher quality and stability.
Hamed Javidi, Maziar Goudarzi
Comput. J.2
2015 Energy-Aware Scheduling for Precedence-Constrained Parallel Virtual Machines in Virtualized Data Centers
Vahid Ebrahimirad, Maziar Goudarzi, Aboozar Rajabi
J. Grid Comput.2
2014 Simultaneous hardware and time redundancy with online task scheduling for low energy highly reliable standby-sparing system
abstract
Standby-sparing is one of the common techniques in order to design fault-tolerant safety-critical systems where the high level of reliability is needed. Recently, the minimization of energy consumption in embedded systems has attracted a lot of concerns. Simultaneous considering of high reliability and low energy consumption by DVS is a challenging problem in designing such a system, since using DVS has been shown to reduce the reliability profoundly. In this article, we have studied different schemes of standby-sparing systems from the energy consumption and reliability point of view. Moreover, we propose a new standby-sparing scheme which addresses both reliability and energy consumption jointly together. This scheme uses a simple energy management coupled with an online task scheduler which tries to dispatch those ready tasks which are expected to lead to high reliability and low energy consumption in the system. The effectiveness of the proposed scheme has been shown on TGFF under stochastic workloads. The results show 52% improvement on energy saving compared to the conventional hot standby-sparing system. Moreover, two orders of magnitude higher reliability is obtained on average, while preserving the same level of energy saving as compared to the state-of-the-art low-energy standby-sparing system (LESS).
Mohammad Khavari Tavana, Nasibeh Teimouri, Meisam Abdollahi, Maziar Goudarzi
ACM Trans. Embed. Comput. Syst.4
2014 Power reduction in HPC data centers: a joint server placement and chassis consolidation approach
Ali Pahlavan, Mahmoud Momtazpour, Maziar Goudarzi
J. Supercomput.3
2012 Accurate Estimation of Leakage Power Variability in Sub-micrometer CMOS Circuits
abstract
Leakage power has already become the major contributor to the total on-chip power consumption, rendering its estimation a necessary step in the IC design flow. The problem is further exacerbated with the increasing uncertainty in the manufacturing process known as process variability. We develop a method to estimate the variation of leakage power in the presence of both intra-die and inter-die process variability. Various complicating issues of leakage prediction such as spatial correlation of process parameters, the effect of different input states of gates on the leakage, and DIBL and stack effects are taken into account while we model the simultaneous variability of the two most critical process parameters, threshold voltage and effective channel length. Our subthreshold leakage current model is shown to fit closely on the HSPICE Monte Carlo simulation data with an average coefficient of determination (R2) value of 0.9984 for all the cells of a standard library. We also demonstrate the adjustability of this model to wider ranges of variation and its extendability to future technology scalings. We show that our framework imposes little timing penalty on the system design flow and is applicable to real design cases. The procedures explained in this paper are part of VAREX, an academic variability modeling framework for estimation of the effect of process variation on power consumption and performance of Multiprocessor SoCs.
Omid Assare, Mahmoud Momtazpour, Maziar Goudarzi
DSD3
2012 Variation-aware Server Placement and Task Assignment for Data Center Power Minimization
abstract
Size and number of data centers are fast growing all over the world and their increasing total power consumption is a worldwide concern. Moreover, increase in the amount of process variation in nanometer technologies and its effect on total power consumption of servers has made it inevitable to move toward variation-aware power reduction strategies. This paper formulates a variation-aware joint server placement and task assignment method using Integer Linear Programming (ILP) to minimize total power consumption of data centers. We first determine the optimum placement of servers in the data center racks based on total power consumption of each server and the data center recirculation model obtained by Computational Fluid Dynamics (CFD) simulations. Then, we dynamically consolidate the ON servers in chassis and racks such that the use of power-greedy servers is minimized. Experimental results reveal up to 14.85% and an average of 8.92% power saving at different server utilization rates with respect to conventional methods.
Ali Pahlavan, Mahmoud Momtazpour, Maziar Goudarzi
ISPA3
2011 Simultaneous variation-aware architecture exploration and task scheduling for MPSoC energy minimization
abstract
In nanometer-scale process technologies, the effects of process variations are observed in Multiprocessor System-on-Chips (MPSoC) in terms of variations in frequencies and leakage powers among the processors on the same chip as well as across different chips of the same design. Traditionally, worst-case values are assumed for these parameters and then a deterministic optimization technique is applied to the MPSoC application under design. We show that such worst-case-based approaches are not optimal with the increasing variation observed at system-level, and instead, statistical approaches should be employed. We consider the problem of simultaneously choosing MPSoC architecture and task allocation for energy optimization under a given performance constraint. Our experimental results on E3S benchmark suite show that the proposed statistical optimization technique can achieve 33.7% improvement on average over conventional worst-case-based techniques and up to 21.7 % improvement over best previously proposed statistical analysis technique.
Mahmoud Momtazpour, Mahboobeh Ghorbani, Maziar Goudarzi, Esmaeil Sanaei
ACM Great Lakes Symposium on VLSI3
2010 SRAM Leakage Reduction by Row/Column Redundancy Under Random Within-Die Delay Variation
abstract
Share of leakage in total power consumption of static RAM (SRAM) memories is increasing with technology scaling. Reverse body biasing increases threshold voltage (Vth), which exponentially reduces subthreshold leakage, but it increases SRAM access delay. Traditionally, when all cells of an SRAM block used to have almost the same delay, within-die variations are increasingly widening the delay distribution of cells even within a single SRAM block, and hence, most of these cells are substantially faster than the delay set for the entire block. Consequently, after the reverse body biasing and the resulting delay rise, only a small number of cells violate the original delay of the SRAM block; we propose to replace them with sufficient number of spare rows/columns of SRAM. Our experiments show that the leakage can be reduced by up to 40% in a 90-nm predictive technology by adding less than ten spare columns to an 8-kB SRAM array for a negligible penalty in delay, dynamic power, and area in the presence of 3% uncorrelated random delay variation.
Maziar Goudarzi, Tohru Ishihara
IEEE Trans. Very Large Scale Integr. Syst.1
2008 Instruction cache leakage reduction by changing register operands and using asymmetric sram cells
abstract
Share of leakage in cache memories is increasing with technology scaling. Studies show that most stored bits in instruction caches are zero, and hence, asymmetric SRAM cells which dissipate less leakage when storing 0, effectively reduce leakage with negligible performance penalty. We show that by carefully choosing register operands of instructions, it is possible to further increase the number of 0 bits, and hence, increase leakage savings in instruction cache. This compiler technique is performed off-line and introduces absolutely no delay penalty since processor registers are all the same. Experimental results of our benchmarks show up to 33% (averaging 30.35%) improvement in leakage.
Maziar Goudarzi, Tohru Ishihara
ACM Great Lakes Symposium on VLSI1
2008 Variation-Aware Software Techniques for Cache Leakage Reduction Using Value-Dependence of SRAM Leakage Due to Within-Die Process Variation
Maziar Goudarzi, Tohru Ishihara, Hamid Noori
HiPEAC1
2008 Row/column redundancy to reduce SRAM leakage in presence of random within-die delay variation
abstract
Traditionally, spare rows/columns have been used in two ways: either to replace too leaky cells to reduce leakage, or to substitute faulty cells to improve yield. In contrast, we first choose a higher threshold voltage (Vth) and/or gate-oxide thickness (Tox) for SRAM transistors at design time to reduce leakage, and then substitute the resulting too slow cells by spare rows/columns. We show that due to within-die delay variation of SRAM cells only a few cells violate target timing at higher Vth or Tox; we carefully choose the Vth and Tox values such that the original memory timing-yield remains intact for a negligible extra delay. On a commercial 90nm process assuming 3% variation in SRAM cell delay, we obtained 47% leakage reduction by adding only 5 redundant columns at negligible area, dynamic power and delay costs.
Maziar Goudarzi, Tohru Ishihara
ISLPED1
2007 A Software Technique to Improve Yield of Processor Chips in Presence of Ultra-Leaky SRAM Cells Caused by Process Variation
abstract
Exceptionally leaky transistors are increasingly more frequent in nano-scale technologies due to lower threshold voltage and its increased variation. Such leaky transistors may even change position with changes in the operating voltage and temperature, and hence, redundancy at circuit-level is not sufficient to tolerate such threats to yield. We show that in SRAM cells this leakage depends on the cell value and propose a first software-based runtime technique that suppresses such abnormal leakages by storing safe values in the corresponding cache lines before going to standby mode. Analysis shows the performance penalty is, in the worst case, linearly dependent to the number of so-cured cache lines while the energy saving linearly increases by the time spent in standby mode. Analysis and experimental results on commercial processors confirm that the technique is viable if the standby duration is more than a small fraction of a second.
Maziar Goudarzi, Tohru Ishihara, Hiroto Yasuura
ASP-DAC1
2007 Interactive presentation: Generating and executing multi-exit custom instructions for an adaptive extensible processor
abstract
To improve the performance of embedded processors, an effective technique is collapsing critical computation subgraphs as application-specific instruction set extensions and executing them on custom functional units. The problems of this approach are immense cost and long time of designing. To address these issues, an adaptive extensible processor was proposed in which custom instructions (CIs) are generated and added after chip-fabrication. To support this feature, custom functional units are replaced by a reconfigurable matrix of functional units with the capability of conditional execution. Unlike previous proposed CIs, it can include multiple exits. Experimental results show that multi-exit CIs enhance the performance by 46% in average compared to CIs limited to one basic block. A maximum speedup of 2.89 compared to a 4-issue in-order RISC processor, and a speedup of 1.66 in average, was achieved on MiBench benchmark suite
Hamid Noori, Farhad Mehdipour, Kazuaki J. Murakami, Koji Inoue, Maziar Goudarzi
DATE5
2007 Implementation of a jpeg object-oriented ASIP: a case study on a system-level design methodology
abstract
In this paper, we present a JPEG decoder implemented in our ODYSSEY design methodology. We start with an object-oriented JPEG decoder model. The total operation from modeling to implementation is done automatically by our EDA tool-set in about 10 hours. The resultant system is a JPEG decoder ASIP whose hardware part is implemented on FPGA logic blocks and software part runs on a MicroBlaze processor. This ASIP can be extended by software routines to implement the motion JPEG or MPEG2 decoding algorithms. We implemented our system on ML402 FPGA-based prototype board. Experimental results show that our ASIP implementation is comparable to other approaches while our approach enables quick and easy development of an ASIP using our EDA tool-set and effectively reduces time-to-market.
Naser MohammadZadeh, Morteza NajafVand, Shaahin Hessabi, Maziar Goudarzi
ACM Great Lakes Symposium on VLSI4
2007 The effect of temperature on cache size tuning for low energy embedded systems
abstract
Energy consumption is a major concern in embedded computing systems. Several studies have shown that cache memories account for about 40% or more of the total energy consumed in these systems. In older technology nodes, active power was the primary contributor to total power dissipation of a CMOS design. However, with the scaling of feature sizes, the share of leakage in total power consumption of digital systems continues to grow. Temperature is a factor which exponentially increases the leakage current. In this paper, we show the effects of temperature on the selection of optimal cache size for low energy embedded systems. Our results show that for a given application, the optimal cache size selection is affected by the temperature. Our experiments have been done for 100nm technology. Our study reveals that the cache size selection for different temperatures depends on the rate at which cache miss increases when reducing the cache size. When the miss rate increases sharply the optimal point is the same for all examined temperatures, however when it becomes smoother, the optimal point for different temperatures begin to get farther.
Hamid Noori, Maziar Goudarzi, Koji Inoue, Kazuaki J. Murakami
ACM Great Lakes Symposium on VLSI2
2007 Using on-chip networks to implement polymorphism in the co-design of object-oriented embedded systems
Maziar Goudarzi, Naser MohammadZadeh, Shaahin Hessabi
J. Comput. Syst. Sci.1
2004 Overhead-Free Polymorphism in Network-on-Chip Implementation of Object-Oriented Models
abstract
We unify virtual-method despatch (polymorphism implementation) and network packet-routing operations; virtual-method calls correspond to network packets, and network addresses are allocated such that routing the packet corresponds to dispatching the call. As the run-time routing structure is inherent in network-on-chip platforms, this unification implements polymorphism for free.
Maziar Goudarzi, Shaahin Hessabi, Alan Mycroft
DATE1
2003 Object-Oriented ASIP Design and Synthesis
Maziar Goudarzi, Shaahin Hessabi, Alan Mycroft
FDL1