EDBT 2026 Demo / reviewers in the wild / expert
Ehsan Atoofian
dblp:26/1095
· DBLP profile ↗
35ranked-venue papers
22as first author
9since 2021 · last 2024
0000-0002-1662-5334ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 17 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PCTC: Hardware and Software Co-design for Pruned Capsule Networks on Tensor Cores
Mohammad Hafezan, Reza Jahadi, Ehsan Atoofian |
Euro-Par (2) | 3 |
| 2024 | Transient Fault Detection in Tensor Cores for Modern GPUsabstractDeep neural networks (DNNs) have emerged as an effective solution for many machine learning applications. However, the great success comes with the cost of excessive computation. The Volta graphics processing unit (GPU) from NVIDIA introduced a specialized hardware unit called tensor core (TC) aiming at meeting the growing computation demand needed by DNNs. Most previous studies on TCs have focused on performance improvement through the utilization of the TC's high degree of parallelism. However, as DNNs are deployed into security-sensitive applications such as autonomous driving, the reliability of TCs is as important as performance. In this work, we exploit the unique architectural characteristics of TCs and propose a simple and implementation-efficient hardware technique called fault detection in tensor core (FDTC) to detect transient faults in TCs. In particular, FDTC exploits the zero-valued weights that stem from network pruning as well as sparse activations arising from the common ReLU operator to verify tensor operations. The high level of sparsity in tensors allows FDTC to run original and verifying products simultaneously, leading to zero performance penalty. For applications with a low sparsity rate, FDTC relies on temporal redundancy to re-execute effectual products. FDTC schedules the execution of verifying products only when multipliers are idle. Our experimental results reveal that FDTC offers 100% fault coverage with no performance penalty and small energy overhead in TCs. Mohammad Hafezan, Ehsan Atoofian |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | PTTS: Power-aware tensor cores using two-sided sparsity
Ehsan Atoofian |
J. Parallel Distributed Comput. | 1 |
| 2022 | Practical approximate quantum multipliers for NISQ devicesabstractThe landscape of quantum computing has changed with the introduction of quantum computers with dozens of qubits. While these computers enable deployment of quantum algorithms into real quantum machines, they suffer from decoherent errors which cause incorrect computation results. One of the factors that impacts noise level in quantum computers is the complexity of circuits. In particular, depth and T-count of quantum circuits are keys to the fidelity of quantum programs. Sohrab Sajadimanesh, Jean Paul Latyr Faye, Ehsan Atoofian |
CF | 3 |
| 2022 | Increasing Robustness against Adversarial Attacks through Ensemble of Approximate MultipliersabstractOver the past few years, deep neural networks (DNNs) have been used to solve a wide range of real-life problems. However, DNNs are vulnerable to adversarial attacks where carefully crafted input perturbations can mislead a well-trained DNN to produce false results. As DNNs are being deployed into security-sensitive applications such as autonomous driving, adversarial attacks may lead to catastrophic consequences. In this work, we propose ensemble of approximate multipliers (EAM) where DNNs with different approximate multipliers are used as a new approach to boost robustness against adversarial attacks. A DNN equipped with an approximate multiplier is an effective method to enhance resiliency of DNNs. However, the degree of robustness in a DNN varies with the type of approximate multiplier. Depending on the level of approximation, the accuracy of the DNN varies for different adversarial attacks. We exploit this variability and propose mixing DNNs with different types of approximate multipliers. Our proposed technique does not require changing the architecture of a model nor memory hierarchy. We only use additional approximate units within a multiplier. We evaluate EAM across different DNNs and under a variety of adversarial attacks. Our evaluations reveal that EAM increases robustness by a large margin compared to an exact model while maintaining accuracy on benign inputs. Ehsan Atoofian |
NAS | 1 |
| 2022 | NISQ-Friendly Non-Linear Activation Functions for Quantum Neural NetworksabstractThe current generation of quantum computers calls for quantum algorithms that require a limited number of quantum gates and are resilient to noises. A suitable design strategy is variational circuits where parameters of circuits are determined through training, an approach that conforms with characteristics of machine learning applications. In this paper, we propose a low-depth and implementation-efficient non-linear activation function for quantum neural networks (QNNs). The building block of a quantum circuit is quantum gate which is a unitary operation. Thus, building a non-linear component out of quantum gates is challenging. While the majority of prior works used measurement as the source of non-linearity, this method has limited ability in classifying datasets. We propose a quantum circuit for the popular Rectified Linear Unit (ReLU) activation function. Our proposed circuit is based on low-cost quantum gates that can be synthesized into primitive gates in contemporary quantum computers. We exploit QNNs that rely on quantum rotation to define decision boundaries for classification problems. In addition, we use controlled quantum gates to detect correlation in data through entanglement of qubits. Our evaluations reveal that QNNs equipped with our proposed quantum ReLU perform well on standard benchmark datasets while requiring dramatically fewer number of epochs for training compared with classical neural networks. In addition, we run QNNs with different number of quantum layers on an IBM quantum computer and show that our proposed circuits are practical and generate meaningful results on real quantum computers. Sohrab Sajadimanesh, Jean Paul Latyr Faye, Ehsan Atoofian |
NAS | 3 |
| 2021 | Sparsity-aware Power Gating for Tensor CoresabstractThis paper introduces an architectural technique that reduces energy of Tensor Cores in GPGPUs. Over the past few years, deep neural networks (DNNs) have become the compelling solution for many applications such as image classification, speech recognition, and natural language processing. Various hardware frameworks have been proposed to accelerate DNNs. In particular, Tensor Cores in NVIDIA GPGPUs offer significant speedup compared with previous GPGPU architectures. However, the great success comes at the cost of excessive energy. Value-based optimization techniques have been utilized to accelerate DNNs. In particular, several studies exploited sparse values to skip unnecessary computations. However, the majority of these studies focused on acceleration of DNNs rather than energy saving. In this work, we exploit power gating to reduce energy of Tensor Cores. We show that blindly applying power gating to multipliers results in significant performance loss due to timing overhead of power gating. In order to mitigate performance penalty of power gating, we propose sparsity-aware power gating (SPG) that monitors inputs of multipliers and turns them off only if inputs remain sparse for long intervals. We further improve SPG by introducing an adaptive technique that dynamically changes power gating policy based on frequency of changes in inputs of multipliers. Our experimental results show that our proposed technique can achieve 21% energy saving in Tensor Cores with negligible impact on performance while maintaining accuracy. Ehsan Atoofian |
SBAC-PAD | 1 |
| 2021 | Reducing Energy in GPGPUs through Approximate Trivial BypassingabstractGeneral-purpose computing using graphics processing units (GPGPUs) is an attractive option for acceleration of applications with massively data-parallel tasks. While performance of modern GPGPUs is increasing rapidly, the power consumption of these devices is becoming a major concern. In particular, execution units and register file are among the top three most power-hungry components in GPGPUs. In this work, we exploit trivial instructions to reduce power consumption in GPGPUs. Trivial instructions are those instructions that do not need computations, i.e., multiplication by one. We found that, during the course of a program's execution, a GPGPU executes many trivial instructions. Execution of these instructions wastes power unnecessarily. In this work, we propose trivial bypassing which skips execution of trivial instructions and avoids unnecessary allocation of resources for trivial instructions. By power gating execution units and skipping trivial computing, trivial bypassing reduces both static and dynamic power. Also, trivial bypassing reduces dynamic energy of register file by avoiding access to register file for source and/or destination operands of trivial instructions. While trivial bypassing reduces energy of GPGPUs, it has detrimental impact on performance as a power-gated execution unit requires several cycles to resume its normal operation. Conventional warp schedulers are oblivious to the status of execution units. We propose a new warp scheduler that prioritizes warps based on availability of execution units. We also propose a set of new power management techniques to reduce performance penalty of power gating, further. To increase energy saving of trivial bypassing, we also propose approximating operands of instructions. We offer a set of new techniques to approximate both integer and floating-point instructions and increase the pool of trivial instructions. Our evaluations using a diverse set of benchmarks reveal that our proposed techniques are able to reduce energy of execution units by 11.2% and dynamic energy of register file by 12.2% with minimal performance and quality degradation. Ehsan Atoofian, Zayan Shaikh, Ali Jannesari |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2021 | Adaptive Computation Reuse for Energy-Efficient Training of Deep Neural NetworksabstractIn recent years, Deep Neural Networks (DNNs) have been deployed into a diverse set of applications from voice recognition to scene generation mostly due to their high-accuracy. DNNs are known to be computationally intensive applications, requiring a significant power budget. There have been a large number of investigations into energy-efficiency of DNNs. However, most of them primarily focused on inference while training of DNNs has received little attention. This work proposes an adaptive technique to identify and avoid redundant computations during the training of DNNs. Elements of activations exhibit a high degree of similarity, causing inputs and outputs of layers of neural networks to perform redundant computations. Based on this observation, we propose Adaptive Computation Reuse for Tensor Cores (ACRTC) where results of previous arithmetic operations are used to avoid redundant computations. ACRTC is an architectural technique, which enables accelerators to take advantage of similarity in input operands and speedup the training process while also increasing energy-efficiency. ACRTC dynamically adjusts the strength of computation reuse based on the tolerance of precision relaxation in different training phases. Over a wide range of neural network topologies, ACRTC accelerates training by 33% and saves energy by 32% with negligible impact on accuracy. Jason Servais, Ehsan Atoofian |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2020 | Approximate trivial instructionsabstractApproximate computing has the potential to improve performance and energy efficiency in high-performance processors. This work focuses on the impact of approximating conventionally non-trivial instructions to trivial instructions. Instructions which do not need to be processed due to the nature of their operands, such as division by 1 or addition with 0 are trivial instructions. By approximating instructions which results in an acceptable level of accuracy in programs' outputs, we can increase the number of trivial instructions and enhance power and performance of trivial bypassing. To approximate integer values, we mask the least significant bits (LSBs) of instructions' operands. The number of masked bits is under the control of programmers. To approximate floating-point values, we propose two different schemes. The first scheme sets a threshold and approximates the values that lie within the threshold region. A 32- or 64-bit comparator, depending on the operand size, is used for comparison between the operand and the threshold. Thus, instructions which would have used the expensive floating-point units are bypassed and only a comparator and a few gates are used instead. The second scheme reduces cost of approximation by replacing full-blown comparators with smaller ones and performing inexact comparisons between the operand and the threshold. Our evaluations using a diverse set of benchmarks reveal that precise comparison and trivial bypassing improve energy-delay by 21% and 13%, respectively while the inexact approximation improves energy-delay by 22%. Zayan Shaikh, Ehsan Atoofian |
CF | 2 |
| 2020 | Energy Efficient On-Demand Dynamic Branch Prediction ModelsabstractThe branch predictor unit (BPU) is among the main energy consuming components in out-of-order (OoO) processors. For integer applications, we find 16 percent of the processor energy is consumed by the BPU. BPU is accessed in parallel with the instruction cache before it is known if a fetch group contains control instructions. We find 85 percent of BPU lookups are done for non-branch operations, and of the remaining lookups, 42 percent are done for highly biased branches that can be predicted statically with high accuracy. We evaluate two variants of a branch prediction model that combines dynamic and static branch prediction to achieve energy improvements for power-constrained applications. These models, named on-demand branch prediction (ODBP) and path-based on-demand branch prediction (ODBP-PATH), are two novel prediction techniques that eliminate unnecessary BPU lookups using compiler generated hints to identify instructions that can be more accurately predicted statically. ODBP-PATH is an implementation of ODBP that combines static and dynamic branch prediction based on the program path of execution. For a 4-wide OoO processor, ODBP-PATH delivers 11 percent average energy-delay (ED) product improvement, and 9 percent core average energy saving on the SPEC Int 2006 benchmarks. Milad Mohammadi, Song Han 0003, Ehsan Atoofian, Amirali Baniasadi, Tor M. Aamodt, William J. Dally |
IEEE Trans. Computers | 3 |
| 2020 | Approximate Cache in GPGPUsabstractThere is a growing number of application domains ranging from multimedia to machine learning where a certain level of inexactness can be tolerated. For these applications, approximate computing is an effective technique that trades off some loss in output data integrity for energy and/or performance gains. In this article, we present the approximate cache, which approximates similar values and saves energy in the L2 cache of general-purpose graphics processing units (GPGPUs). The L2 cache is a critical component in memory hierarchy of GPGPUs, as it accommodates data of thousands of simultaneously executing threads. Simply increasing the size of the L2 cache is not a viable solution to keep up with the growing size of data in many-core applications. This work is motivated by the observation that threads within a warp write values into memory that are arithmetically similar. We exploit this property and propose a low-cost and implementation-efficient hardware to trade off accuracy for energy. The approximate cache identifies similar values during the runtime and allows only one thread writes into the cache in the event of similarity. Since the approximate cache is able to pack more data in a smaller space, it enables downsizing of the data array with negligible impact on cache misses and lower-level memory. The approximate cache reduces both dynamic and static energy. By storing data of a thread into a cache block, each memory instruction requires accessing fewer cache cells, thus reducing dynamic energy. In addition, the approximate cache increases frequency of bank idleness. By power gating idle banks, static energy is reduced. Our evaluations reveal that the approximate cache reduces energy by 52% with minimal quality degradation while maintaining performance of a diverse set of GPGPU applications. Ehsan Atoofian |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2018 | Improving performance of transactional memory through machine learningabstractSummary Transactional memory (TM) is a programming paradigm that facilitates parallel programming for multi‐core processors. In the last few years, some chip manufacturers provided hardware support for TM to reduce runtime overhead of Software Transactional Memory (STM). In this work, we offer two optimization techniques for TMs. The first technique focuses on Restricted Transactional Memory (RTM) in Intel's Haswell processor and shows that while in some applications, RTM improves performance over STM, in some others, it falls behind STM. We exploit this variability and propose an adaptive technique that switches between RTM and STM, statically. The second technique focuses on the overhead of TM and enhances the speed of the adaptive system. In particular, we focus on the size of transactions and improve performance by changing the transaction size. Optimizing the transaction size manually is a time‐consuming process and requires significant software engineering effort. We use a combination of Linear Regression (LR) and decision tree to decide on the transaction size, automatically. We evaluate our optimization techniques using a set of benchmarks from NAS, DiscoPoP, and STAMP benchmark suites. Our experimental results reveal that our optimization techniques are able to improve the performance of TM programs by 9% and energy‐delay by 15%, on average. Yang Xiao 0001, Thireshan Jeyakumaran, Ehsan Atoofian, Ali Jannesari |
Concurr. Comput. Pract. Exp. | 3 |
| 2018 | Data-type specific cache compression in GPGPUs
Ehsan Atoofian, Sean Rea |
J. Supercomput. | 1 |
| 2016 | Compressed L1 data cache and L2 cache in GPGPUsabstractGeneral-Purpose Graphics Processing Units (GPGPUs) exploit several levels of caches to hide latency of memory and provide data for thousands of simultaneously executing threads. L1 data cache and L2 cache are critical to performance of GPGPUs as an L1 data cache should provide data for all threads within the corresponding Streaming Multiprocessor (SM) and the L2 cache should service memory requests of all threads across all SMs. In this paper, we exploit compression to increase effective capacity of the L1 data cache and the L2 cache and improve performance and energy of GPGPUs. Our work is motivated by the observation that many cache blocks accommodate values with low dynamic range, i.e. the differences between values within a cache block are small. Removing redundancy of cache values through compression reduces the effective cache block width, thereby enabling storing more blocks into a cache and improving performance. Also, we exploit opportunities provided by cache compression to reduce energy. Cache compression reduces effective size of cache blocks and this reduces dynamic power as a subset of a full-blown cache block is accessed for compressed data. Furthermore, cache compression increases the number of idle banks. Hence, static power can be reduced by power gating idle banks. Evaluation results reveal that on average, cache compression improves performance by 10.1% and reduces energy of caches by 8%. Ehsan Atoofian |
ASAP | 1 |
| 2016 | Improving Performance of Transactional Applications through Adaptive Transactional MemoryabstractTransactional memory (TM) has become progressively widespread especially with hardware transactional memory implementation becoming increasingly available. In this paper, we focus on Restricted Transactional Memory (RTM) in Intel's Haswell processor and show that performance of RTM varies across applications. While RTM enhances performance of some applications relative to software transactional memory (STM), in some others, it degrades performance. We exploit this variability and present an adaptive system which is a static approach that switches between HTM and STM in transaction granularity. By incorporating a decision tree prediction module, we are able to predict the optimum TM system for a given transaction based on its characteristics. Our adaptive system supports both HTM and STM with the aim of increasing an application's performance. We show that our adaptive system has an average overall speedup of 20.82% over both TM systems. Thireshan Jeyakumaran, Ehsan Atoofian, Yang Xiao 0001, Zhen Li 0005, Ali Jannesari |
PDP | 2 |
| 2015 | Reducing shift penalty in Domain Wall Memory through register localityabstractGeneral-purpose graphics processing units (GPGPUs) have the ability to execute hundreds to thousands of threads simultaneously. Extreme multithreading requires a large register file to hold state of executing threads and facilitate context switching. As feature size reduces, power consumption in the large register file becomes a major concern. In this work, we exploit Domain Wall Memory (DWM) which is a spin-based memory to reduce power consumption in register file. DWM is a promising technology and offers non-volatility, high energy efficiency, and high density by storing several bits into the domains of a ferromagnetic wire. However, despite of favourable properties of DMW over SRAM technology, DWM poses a unique challenge that the bits must be accessed serially through shift operations, leading to variable and potentially higher access latencies. To address this challenge, we propose a new predictive shift policy. In this policy, we exploit register locality across threads and predict source and destination operands of instructions. We record history of registers accessed by instructions and shift magnetic domains of DWM tracks for subsequent instructions, speculatively. Over a wide range of applications from NVIDIA CUDA SDK, ISPASS, and Rodinia, our predictive scheme achieves dramatic energy saving over an SRAM register file while changing performance negligibly. Ehsan Atoofian |
CASES | 1 |
| 2015 | Automatic Optimization of Software Transactional Memory Through Linear Regression and Decision Tree
Yang Xiao 0001, Zhen Li 0005, Ehsan Atoofian, Ali Jannesari |
ICA3PP (4) | 3 |
| 2015 | Shift-aware racetrack memoryabstractIn this work, we exploit racetrack memory for L2cache in GPGPUs. Racetrack memory is a memory technology in which several bits of data are packed into the domains of a ferromagnetic wire. While racetrack memory reduces power consumption compared to CMOS technology, it increases access time. To read or write a memory bit, memory cells should be shifted serially until the requested bit reaches to an access port. This increases latency of L2cache and hurts performance. To address this challenge, we use address predictors to pre-shift memory cells ahead of time. Our evaluations using a set of GPGPU applications reveal that our speculative approach is effective and is able to reduce performance penalty of racetrack memory. Ehsan Atoofian, Ahsan Saghir |
ICCD | 1 |
| 2014 | Improving Power of Cache and Register File through Critical Path InstructionsabstractAs feature size shrinks, power becomes one of the limiting factors in design of modern processors. Cache and register-file are the two power hungry components in processors, consuming more than one third of total processors' power budget. In this work, we propose a new architecture for cache and register-file which exploits critical path instructions to reduce power consumption. In this architecture, we have cache and register-file cells operating at two different voltage levels and we change the structure of the cells so that they dynamically switch between nominal and reduced supply voltages. Those cells that are accessed frequently by critical instructions are assigned to use nominal supply voltage to preserve performance. On the other side, the cells that are rarely accessed by critical instructions are assigned to low supply voltage to reduce power consumption. To reduce performance impact of voltage switching, we monitor critical instructions within long intervals and adjust the voltage of cells only when the intervals are elapsed. Our simulation results reveal that our optimization technique results in significant power saving with negligible effect on performance. KuangLun Chen, Ehsan Atoofian, Ali Manzak |
DSD | 2 |
| 2014 | Power-Aware L1 and L2 Caches for GPGPUs
Ehsan Atoofian, Ali Manzak |
Euro-Par | 1 |
| 2013 | VGTS: Variable Granularity Transactional Snoop
Ehsan Atoofian |
Euro-Par | 1 |
| 2013 | Read-Write Lock Allocation in Software Transactional MemoryabstractTransactional Memory (TM) is a promising programming model for managing concurrent accesses to the shared memory locations. Time-based Software Transactional Memories (STMs) exploit a global clock to maintain consistency of transactions and validate transactional data. One of the shortcomings of this technique is that the global clock becomes bottleneck as the number of transactions increases. In this paper, we introduce two optimization techniques to overcome the overhead of the global clock. The first technique is Read-Write Lock Allocation (RWLA) which does not exploit any central data structure to maintain consistency of transactions. This method improves performance of STMs only if transactions commit successfully. However, in the event of frequent conflicts, RWLA increases cost of abort and degrades performance. Our second optimization technique is an adaptive technique which dynamically selects either baseline scheme or RWLA. Our experimental results reveal that our adaptive technique is effective and is able to improve performance of transactional applications up to 66%. Amir Ghanbari Bavarsad, Ehsan Atoofian |
ICPP | 2 |
| 2013 | Consistency Check through O-GEHL PredictorsabstractTransactional Memory (TM) is a promising paradigm to facilitate parallel programming for multicore processors. In Software implementation of TMs (STMs), transactions rely on a global clock to maintain consistency of transactional data. While this method is simple to implement, it results in significant timing overhead if transactions commit frequently. The alternative approach is Thread Local Clock (TLC) which exploits decentralized local variables to maintain consistency in transactions. However, TLC may increase false aborts and degrade performance of STMs. In this paper, we introduce Adaptive Clock (AC) which dynamically selects one of the two validation techniques based on probability of conflicts. AC is a speculative approach and relies on O-GEHL predictors to speculate future conflicts. We have incorporated AC into TL2 and compared the performance of the new implementation with the original STM using Stamp v0.9.10 benchmark suite. Our results reveal that AC is effective and improves performance of transactional applications up to 33%. Ehsan Atoofian |
PDP | 1 |
| 2013 | ARV-ALA: Improving performance of software transactional memory through adaptive read and write policies
Ehsan Atoofian, Amirali Baniasadi, Yvonne Coady |
Sci. Comput. Program. | 1 |
| 2013 | Improving performance of software transactional memory through contention locality
Ehsan Atoofian |
J. Supercomput. | 1 |
| 2012 | Maintaining Consistency in Software Transactional Memory through Dynamic Versioning Tuning
Ehsan Atoofian, Amir Ghanbari Bavarsad |
ICA3PP (2) | 1 |
| 2012 | TRT: Transactional Read TrackingabstractMany recent Software Transactional Memory (STM) algorithms exploit a global clock to maintain consistency among transactions. While this method is simple to implement it results in contention over the global clock, especially when transactions commit frequently. In this paper, we introduce Transactional Read Tracking (TRT) which does not exploit any central data structure to maintain consistency of transactions. TRT tracks transactional read and write operations and aborts transactions if they conflict over a shared memory location. As such, TRT maintains consistency and eliminates the overhead of the global clock. Our experimental results reveal that this method is effective and is able to improve performance of transactional applications up to 63%. Amir Ghanbari Bavarsad, Ehsan Atoofian |
PDCAT | 2 |
| 2012 | ArTA: Adaptive Granularity in Transactional ApplicationsabstractSoftware Transactional Memory (STM) is a programming paradigm which simplifies parallel programming for multi-core processors. A key requirement in STMs is the mechanism to track memory accesses and detect conflicts among speculative transactions. Current STMs exploit a fixed-size tracking scheme to detect conflicts, i.e. at the word level. However, the choice of access granularity significantly affects the performance of STMs. While a coarse-grained access tracking increases false conflicts a fine-grained scheme may increase overhead of STMs due to the cost of lock acquisitions. In order to mitigate the disadvantages of a fixed-size access tracking, we propose adaptive granularity in transactional applications (ArTA) to change the granularity of STMs dynamically and in runtime. ArTA is a speculative approach and relies on history of transactions to select access granularity for shared data structures. We have incorporated ArTA into TL2 and compared the performance of the new implementation with the original STM using Stamp v0.9.10 benchmark suite. Our results reveal that ArTA improves performance of applications up to 42%. Ehsan Atoofian |
PDP | 1 |
| 2008 | Using supplier locality in power-aware interconnects and caches in chip multiprocessors
Ehsan Atoofian, Amirali Baniasadi |
J. Syst. Archit. | 1 |
| 2007 | A Power-Aware Prediction-Based Cache Coherence Protocol for Chip MultiprocessorsabstractSnoopy cache coherence protocols broadcast requests to all nodes, reducing the latency of cache to cache transfer misses at the expense of increasing interconnect power. We propose speculative supplier identification (SSI) to reduce power dissipation in binary tree interconnects in snoopy cache coherence implementations. In SSI, instead of broadcasting a request to all processors, we send the request to the node more likely to have the missing data. We reduce power as we limit access only to the interconnect components between the requestor and the supplier node. We evaluate SSI using shared memory applications. We show that SSI reduces interconnect power by 23% in a 4-way multiprocessor. This comes with negligible performance cost and hardware overhead. SSI does not change existing coherence protocols and is completely transparent to software and the operating system. Ehsan Atoofian, Amirali Baniasadi |
IPDPS | 1 |
| 2007 | Speculative trivialization point advancing in high-performance processors
Ehsan Atoofian, Amirali Baniasadi |
J. Syst. Archit. | 1 |
| 2006 | A Test Approach for Look-Up Table Based FPGAs
Ehsan Atoofian, Zainalabedin Navabi |
J. Comput. Sci. Technol. | 1 |
| 2003 | A BIST Architecture for FPGA Look-Up Table Testing Reduces ReconfigurationsabstractThis paper describes a test architecture for minimum number of test configurations for test of FPGA (field programmable gate array) LUTs (look up tables). Our test architecture includes a TPG (test pattern generator) that is tested while it is generating test data for LE (logic elements) that form our CUT (circuit under test). This scheme eliminates the need for switching LEs between CUT, TPG and ORA (output response analyzer) and having to perform many reconfigurations of the FPGA. An external ORA locates faults of the FPGA under test. In addition to the LUTs, we are also presenting a scheme for testing other parts of the LEs. Compared with other methods, our method uses the least number of reconfigurations of an FPGA for its LUT testing. Ehsan Atoofian, Zainalabedin Navabi |
Asian Test Symposium | 1 |
| 2003 | A Low Power BIST Architecture for FPGA Look-Up Table Testing
Ehsan Atoofian, Zainalabedin Navabi |
VLSI-SOC | 1 |