EDBT 2026 Demo / reviewers in the wild / expert
Ali Afzali-Kusha
dblp:07/1409
· DBLP profile ↗
95ranked-venue papers
1as first author
16since 2021 · last 2026
0000-0001-8614-2007ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 85 · 12 since 2021Software engineering, systems software and programming languages · 13 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-authorArtificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AFMIS: An approximate floating-point multiplier based on input segmentation
Asma Naseri Rad, Shaghayegh Vahdat, Ali Afzali-Kusha, Massoud Pedram |
Future Gener. Comput. Syst. | 3 |
| 2026 | Reliable yet high-speed memristor-based mixed-signal coarse-grained reconfigurable architecture for CNN inference acceleration
Reza Kazerooni-Zand, Ali Afzali-Kusha, Mehdi Kamal |
Neurocomputing | 2 |
| 2026 | On the use of approximate computing for improving the robustness of DNNs against adversarial attacks
Sahand Divsalar, Fatemeh Arezoomand, Shaghayegh Vahdat, Ali Afzali-Kusha, Massoud Pedram |
J. Supercomput. | 4 |
| 2023 | ReMeCo: Reliable Memristor-Based in-Memory Neuromorphic ComputationabstractMemristor-based in-memory neuromorphic computing systems promise a highly efficient implementation of vector-matrix multiplications, commonly used in artificial neural networks (ANNs). However, the immature fabrication process of memristors and circuit level limitations, i.e., stuck-at-fault (SAF), IR-drop, and device-to-device (D2D) variation, degrade the reliability of these platforms and thus impede their wide deployment. In this paper, we present ReMeCo, a redundancy-based reliability improvement framework. It addresses the non-idealities while constraining the induced overhead. It achieves this by performing a sensitivity analysis on ANN. With the acquired insight, ReMeCo avoids the redundant calculation of least sensitive neurons and layers. ReMeCo uses a heuristic approach to find the balance between recovered accuracy and imposed overhead. ReMeCo further decreases hardware redundancy by exploiting the bit-slicing technique. In addition, the framework employs the ensemble averaging method at the output of every ANN layer to incorporate the redundant neurons. The efficacy of the ReMeCo is assessed using two well-known ANN models, i.e., LeNet, and AlexNet, running the MNIST and CIFAR10 datasets. Our results show 98.5% accuracy recovery with roughly 4% redundancy which is more than 20× lower than the state-of-the-art. Ali BanaGozar, Seyed Hossein Hashemi Shadmehri, Sander Stuijk, Mehdi Kamal, Ali Afzali-Kusha, Henk Corporaal |
ASP-DAC | 5 |
| 2023 | Federated learning by employing knowledge distillation on edge devices with limited hardware resources
Ehsan Tanghatari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
Neurocomputing | 3 |
| 2023 | A2P-MANN: Adaptive Attention Inference Hops Pruned Memory-Augmented Neural NetworksabstractIn this work, to limit the number of required attention inference hops in memory-augmented neural networks, we propose an online adaptive approach called [Formula: see text]-memory-augmented neural network (MANN). By exploiting a small neural network classifier, an adequate number of attention inference hops for the input query are determined. The technique results in the elimination of a large number of unnecessary computations in extracting the correct answer. In addition, to further lower computations in [Formula: see text]-MANN, we suggest pruning weights of the final fully connected (FC) layers. To this end, two pruning approaches, one with negligible accuracy loss and the other with controllable loss on the final accuracy, are developed. The efficacy of the technique is assessed by applying it to two different MANN structures and two question answering (QA) datasets. The analytical assessment reveals, for the two benchmarks, on average, 50% fewer computations compared to the corresponding baseline MANNs at the cost of less than 1% accuracy loss. In addition, when used along with the previously published zero-skipping technique, a computation count reduction of approximately 70% is achieved. Finally, when the proposed approach (without zero skipping) is implemented on the CPU and GPU platforms, on average, a runtime reduction of 43% is achieved. Mohsen Ahmadzadeh, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Accuracy Configurable Adders with Negligible Delay Overhead in Exact Operating ModeabstractIn this paper, two accuracy configurable adders capable of operating in approximate and exact modes are proposed. In the adders, which include a block-based carry propagate and a parallel prefix structure, the carry chains are cut off in the approximate mode limiting the carry chain depth to two blocks. In the case of parallel prefix adder, we propose a special carry generate tree equipped with a power gating means. In both of the proposed structures, the critical paths of the adders are not increased in the exact operating mode. Thus, the main objective of proposing these approximate adder structures is to present an accuracy configurable adder structure whose delay in the exact mode is almost the same as an exact adder. The efficacies of the proposed accuracy configurable adders are compared with some state-of-the-art adder structures using a 15nm CMOS technology. In addition, their efficacies are evaluated in two error-resilient applications. These studies show that the proposed carry-propagate adder has 22% (51%) lower energy consumption (error rate) compared to the best prior works. Also, the proposed parallel prefix adder provides, on average, 20% lower energy consumption compared to the exact parallel prefix adders. Farhad Ebrahimi-Azandaryani, Omid Akbari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2023 | Memristive-based Mixed-signal CGRA for Accelerating Deep Neural Network InferenceabstractIn this paper, a mixed-signal coarse-grained reconfigurable architecture (CGRA) for accelerating inference in deep neural networks (DNNs) is presented. It is based on performing dot-product computations using analog computing to achieve a considerable speed improvement. Other computations are performed digitally. In the proposed structure (called MX-CGRA), analog tiles consisting of memristor crossbars are employed. To reduce the overhead of converting the data between analog and digital domains, we utilize a proper interface between the analog and digital tiles. In addition, the structure benefits from an efficient memory hierarchy where the data is moved as close as possible to the computing fabric. Moreover, to fully utilize the tiles, we define a set of micro instructions to configure the analog and digital domains. Corresponding context words used in the CGRA are determined by these instructions (generated by a companion compiler tool). The efficacy of the MX-CGRA is assessed by modeling the execution of state-of-the-art DNN architectures on this structure. The architectures are used to classify images of the ImageNet dataset. Simulation results show that, compared to the previous mixed-signal DNN accelerators, on average, a higher throughput of 2.35 × is achieved. Reza Kazerooni-Zand, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2022 | SySCIM: SystemC-AMS Simulation of Memristive Computation In-MemoryabstractComputation-in-memory (CIM) is one of the most appealing computing paradigms, especially for implementing artificial neural networks. Non-volatile memories like ReRAMs, PCMs, etc., have proven to be promising candidates for the realization of CIM processors. However, these devices and their driving circuits are subject to non-idealities. This paper presents a comprehensive platform, named SysCIM, for simulating memristor-based CIM systems. SySCIM considers the impact of the non-idealities of the CIM components, including memristor device, memristor crossbar (interconnects), analog-to-digital converter, and transimpedance amplifier, on the vector-matrix multiplication performed by the CIM unit. The CIM modules are described in SystemC and SystemC-AMS to reach a higher simulation speed while maintaining high simulation accuracy. Experiments under different crossbar sizes show SySCIM performs simulations up to 117 x faster than HSPICE with less than 4% accuracy loss. The modular design of SySCIM provides researchers with an easy design-space exploration tool to investigate the effects of various non-idealities. Seyed Hossein Hashemi Shadmehri, Ali BanaGozar, Mehdi Kamal, Sander Stuijk, Ali Afzali-Kusha, Massoud Pedram, Henk Corporaal |
DATE | 5 |
| 2022 | Distributing DNN training over IoT edge devices based on transfer learning
Ehsan Tanghatari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
Neurocomputing | 3 |
| 2022 | An Adaptive Memory-Side Encryption Method for Improving Security and Lifetime of PCM-Based Main MemoryabstractIn this article, we present a main memory system for improving the lifetime and security of phase-change main memories. Storing encrypted data increases the bit-flip rates in memory cells, which adversely affects the lifetime of the phase-change memory cells. Thus, to improve the lifetime and security, the proposed system reduces the bit-flip rates by introducing two techniques. The first technique is a memory-side encryption which provides security against DIMM stealing attacks. To prevent unauthorized accesses, in this technique, the encrypted data are not saved in the main memory. As the second technique, we suggest an adaptive partial encryption approach, which makes use of behavior tracking of the application in the CPU side to minimize the latency overhead of the first technique. Additionally, it prevents the loss of data against application-based attacks. This technique uses a recurrent neural network (RNN) to do sequence classification and detect malicious applications. In addition, an auxiliary method, called periodic encryption (PE), which overcomes the security loss in some applications induced by the low accuracy of the employed neural network, is presented. The efficacy of the proposed method is evaluated using gem5 simulator and some benchmarks. Compared to DEUCE and Crypto-Comp methods, the results for the lifetime evaluation show an average bit-flip rate reduction of 11%. In addition, the security improvements against the DIMM stealing and application-based attacks are about 100% and 92.5%, respectively. Morteza Soltani, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Posit Process Element for Using in Energy-Efficient DNN AcceleratorsabstractIn this work, we present an energy-efficient posit processing element (PE) for utilization in array-based deep neural network (DNN) accelerators along with an approximation method for further reducing the energy consumption of the unit. The posit arithmetic used in the proposed PE provides high precision for the considered data widths even when approximation is used for operations. Using some modification/simplification approaches and proposing a speculative posit adder (SPA) unit, we reduce the complexity of the employed posit multiply–accumulator (MAC) in the proposed PE. The effectiveness of the proposed PE is studied using a 45-nm CMOS technology. The results reveal$3.5\times $and 92% improvements in the delay and energy consumption, respectively, compared to those of the state-of-the-art posit PE. To assess the efficacy of the proposed PE, we have modeled an 8-bit DNN accelerator and employed it for the implementation of some DNN architectures. The results indicate that the proposed PE and its approximate one provide, on average, 19.3% and 29.6% lower energy consumptions compared to that of the latest prior work when providing 10.6% and 5.8% higher accuracies, respectively. Mohamadreza Zolfagharinejad, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2021 | Loading-Aware Reliability Improvement of Ultra-Low Power Memristive Neural NetworksabstractIn this paper, a method for offline training of inverter-based memristive neural networks (IM-NNs), called ERIM, is presented. In this method, the output voltage of the inverter is modeled very accurately by considering the loading effect of the memristive crossbar. To properly choose the size of each inverter, its output load and the required slope of its voltage transfer characteristic (VTC) for an acceptable level of resiliency to the circuit element non-idealities are taken into account. The efficacy of ERIM is investigated by comparing its accuracy to those of two recently proposed offline training methods for IM-NNs (RIM and PHAX). The study is performed using IRIS, BCW, MNIST, and Fashion MNIST datasets. Simulation results show that 72% (56%) reduction in average energy consumption of the trained networks is achieved compared to RIM (PHAX) thanks to proper sizing of the inverters. In addition, due to the higher accuracy of the NN mathematical model, ERIM results in significant improvements in the match between the results of high-level modeling and HSPICE simulations while exhibiting lower sensitivity to circuit element variations. Shaghayegh Vahdat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Reliability Enhancement of Inverter-Based Memristor Crossbar Neural Networks Using Mathematical Analysis of Circuit Non-IdealitiesabstractIn this paper, the sensitivity of the neural network (NN) outputs to device parameter uncertainties (non-idealities) in inverter-based memristor (IM) crossbar neuromorphic circuits is mathematically modeled and verified using exhaustive circuit and system-level simulations. The NN sensitivity is obtained by modeling the sensitivity of theIMneuron output to the non-idealities of its circuit elements. The analysis reveals a higher sensitivity of the output voltage of theIMneuron to the non-idealities of the inverters compared to the conductance variation of the memristors. Among the inverter non-idealities, horizontal shift of the inverters voltage transfer characteristic (VTC) shows the highest impact on the output voltage of the neuron. To reduce the accuracy loss due to the variations, a training approach which includes a sensitivity term in the cost function of the training phase, is suggested. The achievable improvements through the said NN training approach are evaluated. In the evaluation, the California Housing, MNIST, and Fashion MNIST datasets are employed. The results show up to 50% reduction in the NN output variations in the presence of circuit elements’ non-idealities. Shaghayegh Vahdat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | An Energy-Efficient Inference Method in Convolutional Neural Networks Based on Dynamic Adjustment of the Pruning LevelabstractIn this article, we present a low-energy inference method for convolutional neural networks in image classification applications. The lower energy consumption is achieved by using a highly pruned (lower-energy) network if the resulting network can provide a correct output. More specifically, the proposed inference method makes use of two pruned neural networks (NNs), namely mildly and aggressively pruned networks, which are both designed offline. In the system, a third NN makes use of the input data for the online selection of the appropriate pruned network. The third network, for its feature extraction, employs the same convolutional layers as those of the aggressively pruned NN, thereby reducing the overhead of the online management. There is some accuracy loss induced by the proposed method where, for a given level of accuracy, the energy gain of the proposed method is considerably larger than the case of employing any one pruning level. The proposed method is independent of both the pruning method and the network architecture. The efficacy of the proposed inference method is assessed on Eyeriss hardware accelerator platform for some of the state-of-the-art NN architectures. Our studies show that this method may provide, on average, 70% energy reduction compared to the original NN at the cost of about 3% accuracy loss on the CIFAR-10 dataset. Mohammad Ali Maleki, Alireza Nabipour-Meybodi, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2021 | OPTIMA: An Approach for Online Management of Cache Approximation Levels in Approximate Processing SystemsabstractIn this article, we present an approach for adjusting the approximation levels of the cache memories in the memory hierarchy of an approximate processing system. The technique, which is called online management of cache approximation level (OPTIMA), adjusts the approximation levels of the caches under a predefined accuracy constraint. OPTIMA may also be employed for multicore processors, which comprise cores with private and shared caches running applications with different error constraints. To reduce the energy consumption, OPTIMA determines the proper approximation level of each cache memory using heuristic algorithms in two main steps. In the first step, the approximate levels are adjusted to maximize the power efficiency by dropping the application accuracy to a level that still meets a desirable minimum output quality. In the second step, output accuracy variations due to input pattern changes are compensated by fine tuning. We suggest two algorithms (with different adjustment speeds of approximate levels) for the first step and another algorithm for the second step. To assess the efficacy of OPTIMA, we integrate it in the gem5 simulator and simulate some multiprocessor configurations by running eight approximate benchmarks. The results show that the proposed approach provides up to 44% power consumption reduction in the memory hierarchy. Roohollah Yarmand, Mehdi Kamal, Ali Afzali-Kusha, Pooria Esmaeli, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | Circuit-Level Techniques for Logic and Memory Blocks in Approximate Computing SystemsxabstractThis article presents an overview of circuit-level techniques used for approximate computing (AC), including both computation and data storage units. After providing some background concept and methodology review, this article proceeds to provide a detailed review of prior art in circuit-level approximation techniques for data path and memory. The focus is on identifying key circuit-level approximation techniques that are applicable to the computational blocks in general and for both volatile and nonvolatile memory circuit technologies. Emphasis is also placed on the error metrics used to assess the output quality of approximate compute and memory units and whether the accuracy setting is dynamically reconfigurable. This article is concluded with a summary of the key distinguishing features of the reviewed prior art. Saba Amanollahi, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
Proc. IEEE | 3 |
| 2020 | X-CGRA: An Energy-Efficient Approximate Coarse-Grained Reconfigurable ArchitectureabstractIn this article, we present an energy-efficient approximate CGRA (X-CGRA). Instead of conventional exact arithmetic units, it employs configurable approximate adders and multipliers in the so-called quality-scalable processing elements (QSPEs). Furthermore, the structure and functionality of the other architectural components, like context memory, are modified based on the quality-scalable operating modes of the QSPEs. The quality reconfigurability of the X-CGRA makes it amenable for both error-resilient and nonresilient applications. To map the applications on the X-CGRA, a mapping technique is proposed that efficiently utilizes the QSPEs and selects appropriate approximation modes in order to lower the energy consumption while satisfying a user-defined quality constraint. We evaluate the efficacy of our X-CGRA for several benchmark applications from different domains, including image/video processing, signal processing, and scientific computations. Different sizes of X-CGRA are synthesized using a 15-nm FinFET technology. For these benchmarks, the results indicate energy consumption reduction of up to $3.21\times $ compared to those of a typical exact CGRA, at the cost of 4% quality loss. Omid Akbari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram, Muhammad Shafique 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | POLAR: A Pipelined/Overlapped FPGA-Based LSTM AcceleratorabstractIn this brief, a low resource utilization field-programmable gate array (FPGA)-based long short-term memory (LSTM) network architecture for accelerating the inference phase is presented. The architecture has low-power and high-speed features that are achieved through overlapping the timing of the operations and pipelining the datapath. Moreover, this architecture requires negligible internal memory size for storing the intermediate data leading to low resource utilization and simple routing, which provides lower interconnect delay (higher operating frequency). A designer may adjust the resource utilization (as well as the latency) of the proposed architecture readily at the register-transfer level (RTL) design by adjusting the amount of parallelization. This makes the process of mapping the architecture onto different types of FPGAs, subject to defined constraints, a simple one. The efficacy of the proposed architecture is assessed by implementing an LSTM network on different types of FPGAs. Compared with the recent works, the proposed architecture provides up to about 1.6x , 43.6x , 21.9x , and 114.5x improvements in frequency, power efficiency, GOP/s, and GOP/s/W, respectively. Finally, our proposed architecture operates at 17.64 GOP/s, which is 2.31 faster than the best previously reported results. Erfan Bank Tavakoli, Seyed Abolfazl Ghasemzadeh, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | RandShift: An Energy-Efficient Fault-Tolerant Method in Secure Nonvolatile Main MemoryabstractIn this article, we present a simple, yet energy- and area-efficient method for tolerating the stuck-at faults caused by an endurance issue in secure-resistive main memories. In the proposed method, by employing the random characteristics of the encrypted data encoded by the Advanced Encryption Standard (AES) as well as a rotational shift operation, a large number of memory locations with stuck-at faults could be employed for correctly storing the data. Due to the simple hardware implementation of the proposed method, its energy consumption is considerably smaller than that of other recently proposed methods. The technique may be employed along with other error correction methods, including the error correction code (ECC) and the error correction pointer (ECP). To assess the efficacy of the proposed method, it is implemented in a phase-change memory (PCM)based main memory system and compared with three error tolerating methods. The results reveal that for a stuck-at fault occurrence rate of 10-2and with the uncorrected bit error rate of 2 × 10-3, the proposed method achieves 82% energy reduction compared to the state-of-the-art method. More generally, using a simulation analysis technique, we show that the fault coverage of the proposed method is similar to that of the state-of-the-art method. Morteza Soltani, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | Interstice: Inverter-Based Memristive Neural Networks Discretization for Function Approximation ApplicationsabstractIn this article, the accuracy of inverter-based memristive neural networks (NNs) for function approximation applications is improved under the presence of process variations. The improvement is achieved by using a design approach, called INTERSTICE (Inverter-based Memristive Neural Networks Dis cretization for Function Approximation Applications), which discretizes the output values by employing a classifier. More precisely, in the INTERSTICE approach, the output range is divided into K subranges where each subrange is considered as a class. To train the classifier, the training samples are labeled where each label shows belonging to a specific class. To evaluate the efficacy of the design technique, some function approximation applications such as BlackScholes, FFT, K-means, and Sobel are considered. Compared to PHAX, a recently published inverter-based memristive NN, INTERSTICE provides lower mean squared error (MSE) values in the presence of memristor and transistor variations. More specifically, the improvements in the mean of MSE (μMSE) are in the range of 40%-80% when considering 10% variations in the memristor resistance and transistor parameters. In addition, for most of the benchmarks, INTERSTICE improves the μMSE values of the nominal case (the case where all circuit elements are ideal) compared to PHAX. As another advantage compared to the PHAX, in INTERSTICE, digital outputs can be generated based on the selected classes which eliminates the need for an analog-to-digital converter at the output port connected to the digital part of the system. Finally, achieving lower μMSE values using fewer memristors and consuming lower energy is also attainable with this design approach. Shaghayegh Vahdat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | DART: A Framework for Determining Approximation Levels in an Approximable Memory HierarchyabstractIn this article, we propose a framework for determining approximation levels of approximable memories in a memory hierarchy for executing error resilient applications. The framework aims at optimizing the configuration for employing approximate memories in a computing system. It is based on considering data footprints at different memory hierarchy levels and an expected output quality to determine the amount of approximations at each memory hierarchy level. The problem of finding a suitable memory approximation configuration is performed using a branch-and-bound algorithm considering all possible memory approximation arrangements. The best configuration leading to the lowest power consumption when meeting the expected output quality is selected. The efficacy of the proposed framework for two memory hierarchies with different cache topologies is evaluated by comparing energy consumptions of approximate memories with those of the exact memory units in the memory hierarchy under different output accuracy level targets. For example, with 28 dB as a peak signal to noise ratio (PSNR) constraint, the study, which is performed for four image processing applications, indicates up to 54% and 22% power consumption improvements for the SRAM cache and the DRAM memory, respectively. Roohollah Yarmand, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | ACHILLES: Accuracy-Aware High-Level Synthesis Considering Online Quality ManagementabstractIn this paper, we present an accuracy-aware design framework [called accuracy-aware high-level synthesis (Achilles)], which synthesizes a high-level description of an input application with the objective of minimizing the energy consumption of the synthesized circuit. The proposed framework includes two main parts of Achilles and light-weight predictor selection. The framework leverages light-weight error predictors (i.e., machine learning-based classifiers) to achieve more energy reduction by dynamically managing the output quality level (exact or approximate) of the synthesized circuit. To synthesize the input application, first, we exploit a heuristic algorithm to determine the quality level required for each operation in the data flow graph (DFG) representation of the input application. Next, for synthesizing the input application, we propose an effective Achilles algorithm which utilizes the flexibility of the available multiquality arithmetic units in a high-level cell library to synthesize the datapath. To improve the efficiency, the process starts by iteratively reducing the number of functional units required for synthesizing the DFG. Then, a proper light-weight error predictor satisfying the user expected quality is chosen from the available predictors in the framework. Based on the quality requirements, three different quality management modes are considered. The efficacy of the proposed framework is assessed for benchmarks from image and signal processing as well as robotics domains. The study of these benchmarks indicates that Achilles may reduce the energy consumption up to 51% (36% on average), up to 72% (51% on average), and up to 57% (33% on average) in threshold, average, and hybrid modes, respectively, for the studied cases. Moreover, the results show that relative coverage of large errors may be increased from 21% to 55% by employing synthetic minority oversampling technique method. Shayan Tabatabaei Nikkhah, Mahdi Zahedi, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | TOSAM: An Energy-Efficient Truncation- and Rounding-Based Scalable Approximate MultiplierabstractA scalable approximate multiplier, called truncation- and rounding-based scalable approximate multiplier (TOSAM) is presented, which reduces the number of partial products by truncating each of the input operands based on their leading one-bit position. In the proposed design, multiplication is performed by shift, add, and small fixed-width multiplication operations resulting in large improvements in the energy consumption and area occupation compared to those of the exact multiplier. To improve the total accuracy, input operands of the multiplication part are rounded to the nearest odd number. Because input operands are truncated based on their leading one-bit positions, the accuracy becomes weakly dependent on the width of the input operands and the multiplier becomes scalable. Higher improvements in design parameters (e.g., area and energy consumption) can be achieved as the input operand widths increase. To evaluate the efficiency of the proposed approximate multiplier, its design parameters are compared with those of an exact multiplier and some other recently proposed approximate multipliers. Results reveal that the proposed approximate multiplier with a mean absolute relative error in the range of 11%-0.3% improves delay, area, and energy consumption up to 41%, 90%, and 98%, respectively, compared to those of the exact multiplier. It also outperforms other approximate multipliers in terms of speed, area, and energy consumption. The proposed approximate multiplier has an almost Gaussian error distribution with a near-zero mean value. We exploit it in the structure of a JPEG encoder, sharpening, and classification applications. The results indicate that the quality degradation of the output is negligible. In addition, we suggest an accuracy configurable TOSAM where the energy consumption of the multiplication operation can be adjusted based on the minimum required accuracy. Shaghayegh Vahdat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | PX-CGRA: Polymorphic approximate coarse-grained reconfigurable architectureabstractCoarse-Grained Reconfigurable Architectures (CGRAs) provide tradeoff between the energy-efficiency of Application Specific Integrated Circuits (ASICs) and the flexibility of General Purpose Processors (GPPs). State-of-the-art CGRAs only support exact architectures and precise application executions. However, a majority of the streaming applications such as multimedia and digital signal processing, which are amenable to CGRAs, are inherently error resilient. Therefore, these applications can greatly benefit from the emerging trend of Approximate Computing that leverages this error-resiliency to provide higher energy efficiency proportional to the tolerable accuracy loss (can even be constrained). This paper, for the first time, introduces the novel concept of Polymorphic Approximate CGRA (PX-CGRA) that employs heterogeneous tiles of Polymorphic-Approximated ALU Clusters (PACs) connected in a 2-D mesh style connection. These PACs can implement different approximate modes as well as accurate modes depending upon their selected configuration as per the run-time requirements of executing applications. For designing an efficient PX-CGRA, we propose a bottom-up design flow. In addition, the flow of application mapping on PX-CGRA is discussed including accuracy-level mapping, scheduling, and binding steps. To comprehensively evaluate the efficacy of the proposed CGRA, the complete PX-CGRA architecture in different sizes as well as with different PACs configurations are synthesized using a 15-nm FinFET technology. Our results show up to 15%-45% energy efficiency improvement for 5%-35% output quality degradation, respectively, when compared to the state-of-the-art exact-mode CGRA. Our proposed architecture and design methodology enable a new era of accuracy-configurable CGRAs to provide significant energy gains. Omid Akbari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram, Muhammad Shafique 0001 |
DATE | 3 |
| 2018 | An Energy-Efficient, Yet Highly-Accurate, Approximate Non-Iterative DividerabstractIn1 this paper, we present a highly accurate and energy efficient non-iterative divider, which uses multiplication as its main building block. In this structure, the division operation is performed by first reforming both dividend and divisor inputs, and then multiplying the rounded value of the scaled dividend by the reciprocal of the rounded value of the scaled divisor. Precisely, the interval representing the fractional value of the scaled divisor is partitioned into non-overlapping sub-intervals, and the reciprocal of the scaled divisor is then approximated with a linear function in each of these sub-intervals. The efficacy of the proposed divider structure is assessed by comparing its design parameters and accuracy with state-of-the-art, non-iterative approximate dividers as well as exact dividers in 45nm digital CMOS technology. Circuit simulation results show that the mean absolute relative error of the proposed structure for doing 1 32-bit division is less than 0.2%, while the proposed structure has significantly lower energy consumption than the exact divider. Finally, the effectiveness of the proposed divider in one image processing application is reported and discussed. Marzieh Vaeztourshizi, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
ISLPED | 3 |
| 2018 | Lifetime improvement by exploiting aggressive voltage scaling during runtime of error-resilient applications
Farzaneh Nakhaee, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram, Sied Mehdi Fakhraie, Hamed Dorosti |
Integr. | 3 |
| 2018 | An Ultra Low-Power Memristive Neuromorphic Circuit for Internet of Things Smart SensorsabstractIn this paper, we propose an ultra low-power analog neuromorphic circuit to be trained to process sensory data in the Internet of Things smart sensors where low-power and are efficient computing is required. To reduce the operating voltage of the circuit while maintaining the performance, we focus on designing a memristive neuromorphic circuit without employing operational amplifiers. Therefore, we use the CMOS inverters as the neurons in our memristive neuromorphic circuit. We also propose ultra low-power mixed-signal input/output interfaces to make the circuit connectable to other digital components such as embedded processor. To assess the efficacy of the proposed circuit and its interfaces which include memristive neural network based A/D and D/A converters, HSPICE simulations are utilized. The results indicate that at the operating voltage of ±0.25 V, at least 108× (278×) reduction in the power consumption of the output (input) interface compared to that of the conventional structures is achieved. Additionally, the effectiveness of the neuromorphic circuit enhanced by the proposed interfaces is evaluated under some applications such as image recognition, human behavior analysis, and air quality predictions. The results of the study reveal that the designed neuromorphic circuits, along with the proposed A/D and D/A converters, provide an average power saving (speedup) of 2960× (37×) over the ASIC implementation in a 90-nm CMOS technology. Arash Fayyazi, Mohammad Ansari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Internet Things J. | 4 |
| 2018 | PHAX: Physical Characteristics Aware Ex-Situ Training Framework for Inverter-Based Memristive Neuromorphic CircuitsabstractIn this paper, we propose a training framework for an inverter-based memristive neuromorphic hardware. The framework, which is called PHAX, is a physical characteristics aware one relying on anex-situtraining approach. The considered neuromorphic circuit is highly energy efficient hybrid CMOS-memristive implementation of neuromorphic circuits. To solve the problem of high sensitivity of the training to the mismatches between the high-level mathematical modeling of the neurons and the corresponding physical characteristics, an approach for analytical yet accurate modeling of the memristive crossbar and neuron circuits is suggested. The approach, which is based on SPICE simulations, models the inverter-based neurons using a hyperbolic tangent function. To increase the training efficacy, the backpropagation training algorithm is modified by considering some constraints based on the physical characteristics of the memristive circuit. This modification along with the accurate back-annotation of the physical characteristics considerably improve the effectiveness of theex-situtraining method of the neuromorphic circuit. The results of this paper show an average reduction of 1805× in the training runtime compared to that of thein-situtraining approach. Furthermore, the results of applying the approach on the kernels of some applications such as image recognition, image processing, and financial analysis reveal that the designed neuromorphic circuits provide an average power saving (speed up) of 1478× (5.2×) over the ASIC implementation in a 90-nm CMOS technology. Mohammad Ansari, Arash Fayyazi, Ali BanaGozar, Mohammad Ali Maleki, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2018 | TheSPoT: Thermal Stress-Aware Power and Temperature Management for Multiprocessor Systems-on-ChipabstractThermal stress including temperature gradients in time and space, as well as thermal cycling, influences lifetime reliability and performance of modern multiprocessor systems-on-chip (MPSoCs). Conventional power and temperature management techniques considering the peak temperature/power consumption do not provide a comprehensive solution to avoid high spatial and temporal thermal variations. This work presents TheSPoT, a novel multilevel thermal stress-aware power and thermal management approach for MPSoCs. At the top level, core consolidation and deconsolidation is performed based on peak temperature, thermal stress, and power consumption constraints. These constraints are also used at the next level, where operating frequencies are determined. At this level, we obtain optimal core frequencies by solving a convex optimization problem. However, thereafter, to reduce the runtime overhead in large MPSoCs, we alternatively propose to use a fast heuristic algorithm. The efficacy of the proposed approaches in reducing the thermal cycles and temporal/spatial temperature gradients is evaluated by comparing the results with the state-of-the-art methods. The evaluation performed on 4-core, 8-core, and 16-core MPSoCs, using PARSEC benchmarks, reveals a considerable reduction in thermal stress. For the 8-core MPSoC case study, on average, for the proposed heuristic(optimal) approach, the mean time to failure improved by 47(35)% compared to the state-of-the-art techniques with only 6(4)% performance degradation. Also, our simulations show that TheSPoT is more efficient in thermal stress reduction when more heterogeneous workloads are used. Arman Iranfar, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram, David Atienza 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | An Efficient False Path-Aware Heuristic Critical Path Selection Method with High Coverage of the Process Variation SpaceabstractIn this article, we present a critical path selection method that efficiently finds true (sensitizable) critical paths of a circuit in the presence of process variations. The method, which is based on the viability analysis, tries to select the least number of true critical paths that cover all of circuit critical gates. Critical gates are those that make a path critical with a probability higher than a predefined threshold value. Selecting fewer critical paths leads to less computation time for the algorithm and shorter test time of fabricated chips. For this purpose, an efficient Statistical Static Timing Analysis– (SSTA) based technique is suggested. This technique tries to find circuit-critical gates whose process parameter variations cover a major part of the process space. Improving the process space coverage using fewer paths is achieved by considering both spatial (proximity of gates) and structural (having common gates) correlations in the analysis of choosing the critical paths. In the selection process, paths with low similarities in their characteristics are preferred. In addition, only true paths whose delays affect the maximum delay of the circuit are included. The selected paths can be used in the test process of the fabricated chips to determine if the chip meets its timing requirements. Also, a modified viability analysis that incorporates statistical computations is used in the SSTA. The efficacy of the proposed method is evaluated by comparing its results for combinational and sequential ISCAS benchmarks with those obtained by exhaustive search. Results indicate although, on average, only 4.38% of all the critical paths found by the exhaustive search are selected by the proposed method, the maximum probability of criticality for the paths that are not considered in our method is, on average, less than 4%. Sheis Abolma'ali, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2018 | Approximate Reverse Carry Propagate Adder for Energy-Efficient DSP Applications
Masoud Pashaeifar, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Robust neuromorphic computing in the presence of process variationabstractIn this paper, an approach for increasing the sustainability of inverter-based memristive neuromorphic circuits in the presence of process variation is presented. The approach works based on extracting the impact of process variations on the neurons characteristics during the test phase through a proposed algorithm. In this method, first, some combinations of inputs and weights (based on the neuromorphic circuit structure) are injected into the circuit and the features of the neurons are determined. Next, these features which are back-annotated, are utilized in an efficient ex-situ training approach to determine the proper weights of the neurons. The approach provides a considerable improvement in the output accuracy. To evaluate the effectiveness of the proposed approach, some approximate applications are studied using 90nm CMOS technology. The results of the study reveal that using this framework provides, on average, 17X higher output accuracy compared to the cases that the impact of the process variation is not considered at all. Ali BanaGozar, Mohammad Ali Maleki, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
DATE | 4 |
| 2017 | TruncApp: A truncation-based approximate divider for energy efficient DSP applicationsabstractIn this paper, we present a high speed yet energy efficient approximate divider where the division operation is performed by multiplying the dividend by the inverse of the divisor. In this structure, truncated value of the dividend is multiplied exactly (approximately) by the approximate inverse value of divisor. To assess the efficacy of the proposed divider, its design parameters are extracted and compared to those of a number of prior art dividers in a 45nm CMOS technology. Results reveal that this structure provides 66% and 52% improvements in the area and energy consumption, respectively, compared to the most advanced prior art approximate divider. In addition, delay and energy consumption of the division operation are reduced about 94.4% and 99.93%, respectively, compared to those of an exact SRT radix-4 divider. Finally, the efficacy of the proposed divider in image processing application is studied. Shaghayegh Vahdat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram, Zainalabedin Navabi |
DATE | 3 |
| 2017 | CL-CPA: A hybrid carry-lookahead/carry-propagate adder for low-power or high-performance operation mode
Milad Bahadori, Mehdi Kamal, Ali Afzali-Kusha, Yasmin Afsharnezhad, Elham Zahraie Salehi |
Integr. | 3 |
| 2017 | Hybrid TFET-MOSFET circuit: A solution to design soft-error resilient ultra-low power digital circuit
Maede Hemmat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
Integr. | 3 |
| 2017 | Efficient Critical Path Identification Based on Viability Analysis Method Considering Process VariationsabstractIn this brief, we propose an effective adaptation of viability analysis in statistical static timing analysis. The adaption benefits well from a dynamic programming implementation of the viability function. For a rapid identification of statistical longest true paths, the technique makes use of a fast preprocessing step identifying the gates with a small probability of being viable in the circuit, and a number of simple optimization techniques. This makes the approach fast without lowering its accuracy. The efficacy of the proposed statistical timing analysis is assessed using ISCAS benchmark circuits and carry skip adders. The results show that the proposed technique leads to, on average, 18× higher speed compared to those of the state-of-the-art technique. This improvement is achieved at the cost of -1.7% precision lost compared to that of the Monte-Carlo method. Sheis Abolma'ali, Nika Mansouri-Ghiasi, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | Dual-Quality 4: 2 Compressors for Utilizing in Dynamic Accuracy Configurable MultipliersabstractIn this paper, we propose four 4:2 compressors, which have the flexibility of switching between the exact and approximate operating modes. In the approximate mode, these dual-quality compressors provide higher speeds and lower power consumptions at the cost of lower accuracy. Each of these compressors has its own level of accuracy in the approximate mode as well as different delays and power dissipations in the approximate and exact modes. Using these compressors in the structures of parallel multipliers provides configurable multipliers whose accuracies (as well as their powers and speeds) may change dynamically during the runtime. The efficiencies of these compressors in a 32-bit Dadda multiplier are evaluated in a 45-nm standard CMOS technology by comparing their parameters with those of the state-of-the-art approximate multipliers. The results of comparison indicate, on average, 46% and 68% lower delay and power consumption in the approximate mode. Also, the effectiveness of these compressors is assessed in some image processing applications. Omid Akbari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | RoBA Multiplier: A Rounding-Based Approximate Multiplier for High-Speed yet Energy-Efficient Digital Signal ProcessingabstractIn this paper, we propose an approximate multiplier that is high speed yet energy efficient. The approach is to round the operands to the nearest exponent of two. This way the computational intensive part of the multiplication is omitted improving speed and energy consumption at the price of a small error. The proposed approach is applicable to both signed and unsigned multiplications. We propose three hardware implementations of the approximate multiplier that includes one for the unsigned and two for the signed operations. The efficiency of the proposed multiplier is evaluated by comparing its performance with those of some approximate and accurate multipliers using different design parameters. In addition, the efficacy of the proposed approximate multiplier is studied in two image processing applications, i.e., image sharpening and smoothing. Reza Zendegani, Mehdi Kamal, Milad Bahadori, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | SEERAD: A high speed yet energy-efficient rounding-based approximate divider
Reza Zendegani, Mehdi Kamal, Arash Fayyazi, Ali Afzali-Kusha, Saeed Safari, Massoud Pedram |
DATE | 4 |
| 2016 | Hybrid TFET-MOSFET circuits: An approach to design reliable ultra-low power circuits in the presence of process variationabstractIn this work, to increase the timing yield of Tunnel Field Effect Transistor (TFET) circuits in the presence of the process variation, we propose to use MOSFET-based gates instead of some TFET-based gates in the TFET circuits. This hybridization approach originates from the fact that TFETs are more sensitive to process variation, when compared to conventional MOSFETs. First, we investigate the impact of process variations on Homojunction InAs TFETs by extracting the distributions of electrical parameters such as threshold voltage. Then, a hybrid TFET-MOSFET circuit design approach for increasing the reliability of the TFET circuits is introduced. The power consumptions of hybrid circuits are considerably smaller than the corresponding ones realized using CMOS circuits. In the proposed hybrid approach, the circuit is basically implemented in TFET to reduce the power and energy consumption while the gates whose their variations may lead to the timing violation, are implemented using MOSFET-based gates. The decision on replacing the TFET-based gates by their corresponding MOSFET-based gates during the hybrid design is made through a heuristic algorithm. The proposed algorithm considers the sensitivity of each TFET-based gate to the process variation. To assess the efficacy of the proposed approach, the proposed algorithm is applied to some circuits of the ISCAS'85 and ISCAS'89 benchmark packages. The results show that the reliability of the TFET-MOSFET-based circuits are up to 74% larger than that of the pure TFET-based circuits. Furthermore, the energy and leakage power consumptions of the proposed hybrid circuits are up to 56% and 80%, respectively, smaller than those of the pure MOSFET-based design. Maede Hemmat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
VLSI-SoC | 3 |
| 2016 | Power and energy reduction of racetrack-based caches by exploiting shared shift operationsabstractIn this paper, we propose a technique for reducing the power and energy consumptions of the racetrack-based caches. The technique uses a mapping method from the logical cache lines to the physical domains of the nanowires. The mapping method exploits the fact that, in a nanowire with several access heads, the shift operations are shared by the heads on that nanowire. Utilizing this inherent sharing, fewer nanowires are shifted to make a cache line available for both the read and write accesses. By using this method, the cache sets are shifted separately, which results in increase in the number of average shift operations. Thanks to the sharing of the shift operations among multiple heads, the total power and energy consumption of the shift operations are reduced. The effectiveness of the proposed technique is studied using the PARSEC benchmark package. The study shows that the power, energy consumption, energy-delay-product, and energy-delay-squared-product of L2 caches are reduced, on average, by 53%, 44%, 32%, 17%, respectively, compared to the state-of-the-art mapping methods. Seyed Saber Nabavi Larimi, Mehdi Kamal, Ali Afzali-Kusha, Hamid Mahmoodi |
VLSI-SoC | 3 |
| 2016 | A comparative study on performance and reliability of 32-bit binary adders
Milad Bahadori, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
Integr. | 3 |
| 2016 | All-Region Statistical Model for Delay Variation Based on Log-Skew-Normal DistributionabstractIn this paper, we propose a single probability density function for the distributions of the delay in the presence of the process variation for different regions of operation. The delay variation model is inspired by considering the analytical current models for each operating region. Based on these models, we suggest using the log-skew-normal distribution for modeling the delay variation for a wide range of supply voltages from the subthreshold to above-threshold regions. To assess the accuracy of the proposed delay distribution, the mean, standard deviation, skewness, 99th percentile, and yield of the proposed distribution are compared with those of the normal and log-normal distributions using the Monte Carlo (MC) simulations for different circuits in both bulk and FinFET technologies. The results show a higher accuracy for the proposed distribution in all regions of operation. Also, the proposed model enables us to obtain the 3σ yield of the distribution using up to 3.4 times less MC simulation time. Hadi Ahmadi Balef, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Yield and Speedup Improvements in Extensible Processors by Allocating Extra Cycles to Some Custom Instructions
Mehdi Kamal, Ali Afzali-Kusha, Saeed Safari, Massoud Pedram |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2016 | High-Speed and Energy-Efficient Carry Skip Adder Operating Under a Wide Range of Supply Voltage LevelsabstractIn this paper, we present a carry skip adder (CSKA) structure that has a higher speed yet lower energy consumption compared with the conventional one. The speed enhancement is achieved by applying concatenation and incrementation schemes to improve the efficiency of the conventional CSKA (Conv-CSKA) structure. In addition, instead of utilizing multiplexer logic, the proposed structure makes use of AND-OR-Invert (AOI) and OR-AND-Invert (OAI) compound gates for the skip logic. The structure may be realized with both fixed stage size and variable stage size styles, wherein the latter further improves the speed and energy parameters of the adder. Finally, a hybrid variable latency extension of the proposed structure, which lowers the power consumption without considerably impacting the speed, is presented. This extension utilizes a modified parallel structure for increasing the slack time, and hence, enabling further voltage reduction. The proposed structures are assessed by comparing their speed, power, and energy parameters with those of other adders using a 45-nm static CMOS technology for a wide range of supply voltages. The results that are obtained using HSPICE simulations reveal, on average, 44% and 38% improvements in the delay and energy, respectively, compared with those of the Conv-CSKA. In addition, the power-delay product was the lowest among the structures considered in this paper, while its energy-delay product was almost the same as that of the Kogge-Stone parallel prefix adder with considerably smaller area and power consumption. Simulations on the proposed hybrid variable latency CSKA reveal reduction in the power consumption compared with the latest works in this field while having a reasonably high speed. Milad Bahadori, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | A thermal stress-aware algorithm for power and temperature management of MPSoCs
Mehdi Kamal, Arman Iranfar, Ali Afzali-Kusha, Massoud Pedram |
DATE | 3 |
| 2015 | A heuristic machine learning-based algorithm for power and thermal management of heterogeneous MPSoCsabstractIn this work, we propose a power and thermal management algorithm based on machine learning to control the thermal stresses and power consumption of the heterogeneous MPSoCs. The objectives of the proposed algorithm are increasing the performance and decreasing the spatial and temporal temperature gradients along with the thermal cycling under the power and temperature constraints. Our proposed power and thermal management method is based on a heuristic approach to speed up the convergence of the machine learning algorithm which makes it applicable for general purpose processors. Adopting Q-Learning as the machine learning algorithm, the heuristic approach aids to limit the learning space by suggesting the most appropriate actions to the agent in each decision epoch. The heuristic algorithm employs the current and previous states of the machine learning, as well as the amount of the temperature stress and power consumption of each core to determine the appropriate action for each core, independently. The proposed algorithm is evaluated on 4-core, 8-core and 16-core homogeneous and heterogeneous MPSoCs for some benchmarks in the Splash2 benchmark package. The results reveal a faster convergence of machine learning and more thermal stresses reduction. Arman Iranfar, Soheil Nazar Shahsavani, Mehdi Kamal, Ali Afzali-Kusha |
ISLPED | 4 |
| 2015 | A near-threshold 7T SRAM cell with high write and read margins and low write time for sub-20 nm FinFET technologies
Mohammad Ansari, Hassan Afzali-Kusha, Behzad Ebrahimi, Zainalabedin Navabi, Ali Afzali-Kusha, Massoud Pedram |
Integr. | 5 |
| 2015 | Design of NBTI-resilient extensible processors
Mehdi Kamal, Ali Afzali-Kusha, Saeed Safari, Massoud Pedram |
Integr. | 2 |
| 2015 | Low Energy yet Reliable Data Communication Scheme for Network-on-ChipabstractIn this paper, a low energy yet reliable communication scheme for network-on-chip is suggested. To reduce the communication energy consumption, we invoke low-swing signals for transmitting data, as well as data encoding techniques, for minimizing both self and coupling switching capacitance activity factors. To maintain the communication reliability of communication at low-voltage swing, an error control coding (ECC) technique is exploited. The decision about end-to-end or hop-to-hop ECC schemes and the proper number of detectable errors are determined through high-level mathematical analysis on the energy and reliability characteristics of the techniques. Based on the analysis, the extended single error correction double-error detecting end-to-end coding technique with three bits of error detection is used in the network layer. For minimization of the self and coupling switching capacitance activity factors, the odd, even, full invert scheme is employed in the data link layer. This coding has an inherent error detection probability for the flits, which is exploited in the suggested technique. The efficiency of the scheme is studied by using both synthetic and real traffic scenarios. The study reveals savings of up to 43% and 58%, for power dissipation and energy consumption, respectively, without any significant performance degradation and overhead in the network interface. Nima Jafarzadeh, Maurizio Palesi, Saeedeh Eskandari, Shaahin Hessabi, Ali Afzali-Kusha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2015 | OPLE: A Heuristic Custom Instruction Selection Algorithm Based on Partitioning and Local Exploration of Application Dataflow GraphsabstractIn this article, a heuristic custom instruction (CI) selection algorithm is presented. The proposed algorithm, which is called OPLE for “Optimization based on Partitioning and Local Exploration,” uses a combination of greedy and optimal optimization methods. It searches for the near-optimal solution by reducing the search space based on partitioning the identified CI set. The partitioning of the identified set guarantees the success of the algorithm independent of the size of the identified set. First, the algorithm finds the near-optimal CIs from the candidate CIs for each part. Next, the suggested CIs from different parts are combined to determine the final selected CI set. To improve the set of the selected CIs, the solution is evolved by calling the algorithm iteratively. The efficacy of the algorithm is assessed by comparing its performance to those of optimal and nonoptimal methods. A comparative study is performed for a number of benchmarks under different area budgets and I/O constraints. The results reveal higher speedups for the OPLE algorithm, especially for larger identified candidate sets and/or small area budgets compared to those of the nonoptimal solutions. Compared to the nonoptimal techniques, the proposed algorithm provides 30% higher speedup improvement on average. The maximum improvement is 117%. The results also demonstrate that in many cases OPLE is able to find the optimal solution. Mehdi Kamal, Ali Afzali-Kusha, Saeed Safari, Massoud Pedram |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2015 | Dynamic Flip-Flop Conversion: A Time-Borrowing Method for Performance Improvement of Low-Power Digital Circuits Prone to VariationsabstractDynamic flip-flop (FF) conversion is a method of time borrowing (TB) for improving the performance of digital systems prone to variations. The first type of this method (Type A), which was previously presented, suffers from a large inefficient transparency window. In this brief, we present an improved structure for this method (Type B) that mitigates this problem by automatically closing the window after the arrival of late data at timing critical FFs. This method was compared with soft edge FF and dynamic clock stretching through simulations on different ITC'99 benchmarks. We defined a parameter called improvement efficiency, which is the ratio of the timing yield improvement to the power overhead of TB. According to the simulation results, the efficiency of Type A is on average 269% more than the best results of other methods when considering only the setup time violations. But, when taking both the setup time and the hold time violations into account, Type B is on average 46% more efficient than the best results of other methods. The simulations also show that the yield improvement of this method increases in higher clock frequencies and it remains the most efficient method when reducing the voltage down to near-threshold region. Mehrzad Nejat, Bijan Alizadeh, Ali Afzali-Kusha |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Improving efficiency of extensible processors by using approximate custom instructionsabstractIn this paper, we propose to move the conventional extensible processor design flow to the approximate computing domain to gain more speedup. In this domain, the instruction set architecture (ISA) design flow selects both exact and approximate custom instructions (CIs). The proposed approach could be used for the applications where imprecise results may be tolerated. In the CI identification phase of the flow, the CIs which do not satisfy the maximum propagation delay but can provide approximate results also may be included in the CI candidate set. Next, in the selection phase, we propose a merit function which selects CIs with higher cycle savings and small error rates. The efficacy of the proposed approximate design flow is investigated using the case studies of the discrete cosine transform (DCT) and inverse DCT (iDCT) of the MPEG2 application. Also, the impact of the process variation on the impreciseness of the results is investigated. Mehdi Kamal, Amin Ghasemazar, Ali Afzali-Kusha, Massoud Pedram |
DATE | 3 |
| 2014 | Dynamic Flip-Flop conversion to tolerate process variation in low power circuitsabstractA novel time borrowing method called dynamic Flip-Flop conversion is presented in this paper. A timing violation predictor detects the violations halfway in the critical path and dynamically converts the critical Flip-Flop to a latch. This way, time borrowing benefits of latches are utilized in a Flip-Flop based design which is more adaptable with Computer-Aided-Design tools. The overhead of this method is smaller than that of similar methods due to the elimination of delay elements. According to the post-synthesis simulations and Monte-Carlo analysis of Spice simulations on some ITC'99 benchmark circuits, the power overhead of the proposed method is about 15% and 19% smaller than that of Soft-Edge-Flip-Flop and Dynamic-Clock-Stretching circuits respectively in a simple case of about 40% yield improvement. This overhead would be relatively even smaller for higher performance and yield improvements. Mehrzad Nejat, Bijan Alizadeh, Ali Afzali-Kusha |
DATE | 3 |
| 2014 | Impact of Process Variations on Speedup and Maximum Achievable Frequency of Extensible ProcessorsabstractIn this article, we investigate the impact of process variations on the speedup and maximum frequency of the extended ISA processor. First, without considering process variations, a custom functional unit (CFU) is designed based on nominal timing parameters, then the timing variations of critical paths of the extensible processor, including the baseline processor and the CFU, are investigated by considering both systematic and random variations. Next, the maximum frequency of the extensible processor and the speed enhancement factor of the extended ISA for different benchmarks are investigated. Results show that timing variation could reduce the speedup of the extensible processor. However, this reduction is highly dependent on the baseline processor and the CFU structures. Additionally, the impact of process variations in the worst-case design approach is studied. Results show that the speedup of the extensible processor is reduced more than in the case when custom instructions (CIs) are selected without considering process variations. To study the impact of each variation type, speedup variations due to random and systematic variations are investigated separately. The study reveals that random variation has a similar effect on the CFU and the baseline processor, while the impact of systematic variation on the baseline processor is greater than the CFU. Mehdi Kamal, Ali Afzali-Kusha, Saeed Safari, Massoud Pedram |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2014 | Data Encoding Techniques for Reducing Energy Consumption in Network-on-ChipabstractAs technology shrinks, the power dissipated by the links of a network-on-chip (NoC) starts to compete with the power dissipated by the other elements of the communication subsystem, namely, the routers and the network interfaces (NIs). In this paper, we present a set of data encoding schemes aimed at reducing the power dissipated by the links of an NoC. The proposed schemes are general and transparent with respect to the underlying NoC fabric (i.e., their application does not require any modification of the routers and link architecture). Experiments carried out on both synthetic and real traffic scenarios show the effectiveness of the proposed schemes, which allow to save up to 51% of power dissipation and 14% of energy consumption without any significant performance degradation and with less than 15% area overhead in the NI. Nima Jafarzadeh, Maurizio Palesi, Ahmad Khademzadeh, Ali Afzali-Kusha |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2013 | An efficient network on-chip architecture based on isolating local and non-local communicationsabstractIn this paper, we propose a scheme for reducing the latency of packets transmitted via on-chip interconnect network in MultiProcessor Systems on Chips (MPSoCs). In this scheme, the network architecture separates the packets transmitted to near destinations from those transmitted to distant ones by using two network layers. These two layers are realized by dividing the channel width among the cores. The optimum ratio for the channel width division is a function of relative significances of the two types of communications. Simulation results indicate that for non-uniform traffic constituting of more than 30 percent local traffic, the proposed network, on average provides 64% and 70% improvement over the conventional one in terms of average network latency and Energy-Delay product (EDP), respectively. Also, for uniform and NED traffic patterns, by adjusting the number of hops between local nodes to include approximately 55 percent of total communications in local ones, the proposed architecture provides the latency reduction of 50%. Vahideh Akhlaghi, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
DATE | 3 |
| 2013 | Modeling symmetrical independent gate FinFET using predictive technology modelabstractPredicting MOSFET models plays a pivotal role in circuit design and its optimization. Independent Gate FinFETs (IGFinFET) are interesting for designers as they are more flexible than Common Multi-Gate FinFETs (CMGFinFET) in digital circuit design. In this work, we implement a model for symmetrical IGFinFET using CMGFinFET model based on Multi-Gate Predictive Technology Model (PTM-MG). This model has been developed from TCAD IGFinFET, based on previously published experimental results of CMG-FinFET. Different basic gates in SG (shorted gate), LP (low power), IG (low area), and IG/LP modes have been designed using the implemented model. For LP, IG, and IG/LP NAND gates, the leakage power is reduced by 89%, 26%, and 67%, respectively in comparison to SG. To show that our model does not have any convergence problem for large circuits, we used ISCAS'85 benchmark suite. The results show that for independent gate in high performance PTM-MG library, on average we can save up to 24% in the number of transistors and lower the total power by 42%. Mohammad Yousef Zarei, Reza Asadpour, Siamak Mohammadi, Ali Afzali-Kusha, Razi Seyyedi |
ACM Great Lakes Symposium on VLSI | 4 |
| 2012 | An architecture-level approach for mitigating the impact of process variations on extensible processorsabstractIn this paper, we present an architecture-level approach to mitigate the impact of process variations on extended instruction set architectures (ISAs). The proposed architecture adds one extra cycle to execute custom instructions (CIs) that violate the maximum allowed propagation delay due to the process variations. Using this method, the parametric yield of manufactured chips will greatly improve. The cost is an increase in the cycle latency of some of the CIs, and hence, a slight performance degradation for the extensible processor architectures. To minimize the performance penalty of the proposed approach, we introduce a new merit function for selecting the CIs during the selection phase of the ISA extension design flow. To evaluate the efficacy of the new selection method, we compare the extended ISAs obtained by this method with those selected based on the worst-case delay. Simulation results reveal that a speedup improvement of about 18% may be obtained by the proposed selection method. Also, by using the proposed merit function, the proposed architecture can improve the speedup about 20.7%. Mehdi Kamal, Ali Afzali-Kusha, Saeed Safari, Massoud Pedram |
DATE | 2 |
| 2012 | An efficient reliability simulation flow for evaluating the hot carrier injection effect in CMOS VLSI circuitsabstractHot carrier injection (HCI) effect is one of the major reliability concerns in VLSI circuits. This paper presents a scalable reliability simulation flow, including a logic cell characterization method and an efficient full chip simulation method, to analyze the HCI-induced transistor aging with a fast run time and high accuracy. The transistor-level HCI effect is modeled based on the Reaction-Diffusion (R-D) framework. The gate-level HCI impact characterization method combines HSpice simulation and piecewise linear curve fitting. The proposed characterization method reveals that the HCI effect on some transistors is much more significant than the others according to the logic cell structure. Additionally, during the circuit simulation, pertinent transitions are identified and all cells in the circuit are classified into two groups: critical and non-critical. The proposed method reduces the simulation time while maintaining high accuracy by applying fine granularity simulation time steps to the critical cells and coarse granularity ones to the non-critical cells in the circuit. Mehdi Kamal, Qing Xie 0001, Massoud Pedram, Ali Afzali-Kusha, Saeed Safari |
ICCD | 4 |
| 2012 | An accurate analytical I-V model for sub-90-nm MOSFETs and its application to read static noise margin modelingabstractWe propose an accurate model to describe the I–V characteristics of a sub-90-nm metal-oxide-semiconductor field-effect transistor (MOSFET) in the linear and saturation regions for fast analytical calculation of the current. The model is based on the BSIM3v3 model. Instead of using constant threshold voltage and early voltage, as is assumed in the BSIM3v3 model, we define these voltages as functions of the gate-source voltage. The accuracy of the model is verified by comparison with HSPICE for the 90-, 65-, 45-, and 32-nm CMOS technologies. The model shows better accuracy than the n th-power and BSIM3v3 models. Then, we use the proposed I–V model to calculate the read static noise margin (SNM) of nano-scale conventional 6T static random-access memory (SRAM) cells with high accuracy. We calculate the read SNM by approximating the inverter transfer voltage characteristic of the cell in the regions where vertices of the maximum square of the butterfly curves are placed. The results for the SNM are also in excellent agreement with those of the HSPICE simulation for 90-, 65-, 45-, and 32-nm technologies. Verification in the presence of process variations and negative bias temperature instability (NBTI) shows that the model can accurately predict the minimum supply voltage required for a target yield. Behrouz Afzal, Behzad Ebrahimi, Ali Afzali-Kusha, Massoud Pedram |
J. Zhejiang Univ. Sci. C | 3 |
| 2012 | High-performance low-leakage regions of nano-scaled CMOS digital gates under variations of threshold voltage and mobilityabstractWe propose a modeling methodology for both leakage power consumption and delay of basic CMOS digital gates in the presence of threshold voltage and mobility variations. The key parameters in determining the leakage and delay are OFF and ON currents, respectively, which are both affected by the variation of the threshold voltage. Additionally, the current is a strong function of mobility. The proposed methodology relies on a proper modeling of the threshold voltage and mobility variations, which may be induced by any source. Using this model, in the plane of threshold voltage and mobility, we determine regions for different combinations of performance (speed) and leakage. Based on these regions, we discuss the trade-off between leakage and delay where the leakage-delay-product is the optimization objective. To assess the accuracy of the proposed model, we compare its predictions with those of HSPICE simulations for both basic digital gates and ISCAS85 benchmark circuits in 45-, 65-, and 90-nm technologies. Hossein Aghababa, Behjat Forouzandeh, Ali Afzali-Kusha |
J. Zhejiang Univ. Sci. C | 3 |
| 2011 | Timing variation-aware custom instruction extension techniqueabstractIn this paper, we propose a technique for custom instruction (CI) extension considering process variations. It bridges the gap between the high level custom instruction extension and chip fabrication in nanotechnologies. In the proposed method, instead of using the conventional static timing analysis (STA), statistical static timing analysis (SSTA) which in turn results in a probabilistic approach to identifying and selecting different parts of the CI extension is utilized. More precisely, we use the delay Probability Density Function (PDF) of the CIs in identification and selection phases of the CI extension. In the identification phase, the delay of each CI is modeled by PDF whereas the performance yield is added as a constraint. Additionally, in the selection phase, the merit function of the conventional approaches is modified to increase the performance gain of the selected CIs at the price of slightly sacrificing the design yield. Also, to make the approach computationally more efficient, we propose a method for reducing the modeling time of the PDF of the CIs by reducing the number of candidate CIs before extracting the PDF. Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
DATE | 2 |
| 2011 | Statistical Design Optimization of FinFET SRAM Using Back-Gate VoltageabstractIn this paper, an optimal approach for the design of 6-T FinFET-based SRAM cells is proposed. The approach considers the statistical distributions of gate length and silicon thickness and their corresponding statistical correlations due to process variations. In this method, a back-gate voltage is used as the optimization knob. With the help of particle swarm optimization (PSO), the back-gate voltages that maximize the yield of the SRAM array against read, write, and access time failures are found. It will be shown that, with this method, a very high yield is achieved. Behzad Ebrahimi, Masoud Rostami, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Dynamic Voltage and Frequency Scheduling for Embedded Processors Considering Power/Performance TradeoffsabstractAn adaptive method to perform dynamic voltage and frequency scheduling (DVFS) for minimizing the energy consumption of microprocessor chips is presented. Instead of using a fixed update interval, the proposed DVFS system makes use of adaptive update intervals for optimal frequency and voltage scheduling. The optimization enables the system to rapidly track the workload changes so as to meet soft real-time deadlines. The technique, which can be realized with very simple hardware, is completely transparent to the application. The results of applying the method to some real application workloads demonstrate considerable power savings and fewer frequency updates compared to DVFS systems based on fixed update intervals. Mostafa E. Salehi, Mehrzad Samadi, Mehrdad Najibi, Ali Afzali-Kusha, Massoud Pedram, Sied Mehdi Fakhraie |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2010 | Statistical delay modeling of read operation of SRAMs due to channel length variationabstractIn this paper, we present a statistical modeling for the transition time of the static random-access memories during the read operation in the presence of the channel length variation. To model the I-V characteristics of the transistors, the a-power law model which is a simple analytical MOS model is used. To increase the accuracy, the effects of the short channel lengths as well as the drain bias are included in the modeling. The statistical analytical modeling is achieved by taking the partial derivations of the transition time expression. To assess the accuracy of the technique, HSPICE Monte-Carlo simulations have been used for a 65nm CMOS technology. The comparison, which is performed for different correlation coefficients, shows a very good accuracy for the model which is evaluated at substantially lower runtime. Hossein Aghababa, Mahmoud Zangeneh, Ali Afzali-Kusha, Behjat Forouzandeh |
ISCAS | 3 |
| 2010 | EDXY - A low cost congestion-aware routing algorithm for network-on-chips
Pejman Lotfi-Kamran, Amir-Mohammad Rahmani, Masoud Daneshtalab, Ali Afzali-Kusha, Zainalabedin Navabi |
J. Syst. Archit. | 4 |
| 2010 | A low-power and low-energy flexible GF(p) elliptic-curve cryptography processorabstractWe investigate the use of two integer inversion algorithms, a modified Montgomery modulo inverse and a Fermat’s Little Theorem based inversion, in a prime-field affine-coordinate elliptic-curve crypto-processor. To perform this, we present a low-power/energy GF( p ) affine-coordinate elliptic-curve cryptography (ECC) processor design with a simplified architecture and complete flexibility in terms of the field and curve parameters. The design can use either of the inversion algorithms. Based on the implementations of this design for 168-, 192-, and 224-bit prime fields using a standard 0.13 μm CMOS technology, we compare the efficiency of the algorithms in terms of power/energy consumption, area, and calculation time. The results show that while the Fermat’s theorem approach is not appropriate for the affine-coordinate ECC processors due to its long computation time, the Montgomery modulo inverse algorithm is a good candidate for low-energy implementations. The results also show that the 168-bit ECC processor based on the Montgomery modulo inverse completes one scalar multiplication in only 0.4 s at a 1 MHz clock frequency consuming only 12.92 μJ, which is lower than the reported values for similar designs. Hamid Reza Ahmadi, Ali Afzali-Kusha |
J. Zhejiang Univ. Sci. C | 2 |
| 2009 | An efficent dynamic multicast routing protocol for distributing traffic in NOCsabstractNowadays, in MPSoCs and NoCs, multicast protocol is significantly used for many parallel applications such as cache coherency in distributed shared-memory architectures, clock synchronization, replication, or barrier synchronization. Among several multicast schemes proposed in on chip interconnection networks, path-based multicast scheme has been proven to be more efficient than the tree-based, and unicast-based. In this paper a low distance path-based multicast scheme is proposed. The proposed method takes advantage of the network partitioning, and utilizing of an efficient destination ordering algorithm. The results in performance, and power consumption show that the proposed method outstands the previous on chip path-based multicasting algorithms. Masoumeh Ebrahimi, Masoud Daneshtalab, Mohammad Hossein Neishaburi, Siamak Mohammadi, Ali Afzali-Kusha, Juha Plosila, Hannu Tenhunen |
DATE | 5 |
| 2009 | Low-Power Low-Energy Prime-Field ECC Processor Based on Montgomery Modular Inverse AlgorithmabstractIn this paper, we present a fast low-power low-energy standard public-key cryptography processor for use in power/energy-limited applications. The proposed prime-field elliptic-curve cryptography hardware uses a modified Montgomery modular inverse algorithm to minimize the total calculation time and is completely flexible in terms of the field and curve parameters. The power consumption is minimized by simplifying the architecture and circuit implementation. To assess the power/energy and timing efficiency of the design, we have implemented the processor for 192-bit prime fields using a standard 0.13 ¿m CMOS technology. The simulation results show that the processor consumes only 39.3 ¿W/MHz which is lower than the power consumption reported for similar designs. Our proposed hardware completes one 192-bit scalar multiplication in 0.525s at a frequency of 1 MHz, consuming only 20.63 ¿J. With these specifications, the proposed processor may be used in many applications of wireless sensors and RFID tags. Hamid Reza Ahmadi, Ali Afzali-Kusha |
DSD | 2 |
| 2009 | Very Low-power Flexible GF(p) Elliptic-curve Crypto-processor for Non-time-critical ApplicationsabstractIn this paper, we present a very low power prime-field elliptic-curve cryptography (ECC) processor for use in power-limited applications. The proposed ECC processor is flexible in terms of field and curve parameters and the power consumption is minimized by simplifying the architecture of the processor and also trading off speed for power. To assess the power efficiency of the design, we have implemented the processor for 168-bit and 192-bit fields using a standard 0.13 mum CMOS technology. The simulation results show that the 168-bit and 192-bit processors consume only 23.1 muW and 26.3 muW at a 1 MHz clock respectively, which is considerably lower than the power consumptions reported for similar designs. At the standard frequency of 13.56 MHz, the hardware completes one 168-bit scalar multiplication in 0.72 seconds and one 192-bit scalar multiplication in 1.15 seconds. With these specifications, the proposed ECC processor may be used in non-time-critical applications of contact-less smart cards as well as RFID tags and wireless sensors. Hamid Reza Ahmadi, Ali Afzali-Kusha |
ISCAS | 2 |
| 2009 | BZ-FAD: A Low-Power Low-Area Multiplier Based on Shift-and-Add ArchitectureabstractIn this paper, a low-power structure called bypass zero, feed A directly (BZ-FAD) for shift-and-add multipliers is proposed. The architecture considerably lowers the switching activity of conventional multipliers. The modifications to the multiplier which multipliesAbyBinclude the removal of the shifting theBregister, direct feeding ofAto the adder, bypassing the adder whenever possible, using a ring counter instead of a binary counter and removal of the partial product shift. The architecture makes use of a low-power ring counter proposed in this work. Simulation results for 32-bit radix-2 multipliers show that the BZ-FAD architecture lowers the total switching activity up to 76% and power consumption up to 30% when compared to the conventional architecture. The proposed multiplier can be used for low-power applications where the speed is not a primary design parameter. M. Mottaghi-Dastjerdi, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Design and Analysis of Two Low-Power SRAM Cell StructuresabstractIn this paper, two static random access memory (SRAM) cells that reduce the static power dissipation due to gate and subthreshold leakage currents are presented. The first cell structure results in reduced gate voltages for the NMOS pass transistors, and thus lowers the gate leakage current. It reduces the subthreshold leakage current by increasing the ground level during the idle (inactive) mode. The second cell structure makes use of PMOS pass transistors to lower the gate leakage current. In addition, dual threshold voltage technology with forward body biasing is utilized with this structure to reduce the subthreshold leakage while maintaining performance. Compared to a conventional SRAM cell, the first cell structure decreases the total gate leakage current by 66% and the idle power by 58% and increases the access time by approximately 2% while the second cell structure reduces the total gate leakage current by 27% and the idle power by 37% with no access time degradation. G. Razavipour, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | Wavelet-based dynamic power management for nonstationary service requestsabstractIn this article, a wavelet-based dynamic power management policy (WBDPM) is proposed. In this approach, the workload source (service requester) is modeled by a nonstationary time series which, in turn, represented by a nondecimated Haar wavelet as its basis. The proposed approach is robust and has the ability to minimize energy dissipation under different performance constraints. To assess the accuracy of the model, the algorithm was implemented for data extracted from the hard disks of computers. Prediction results of this approach for the case of a nonstationary service requester exhibit accuracies of more than 95%. Ali Abbasian, Safar Hatami, Ali Afzali-Kusha, Massoud Pedram |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2007 | A High-Speed and Low-Power Voltage Controlled Oscillator in 0.18-µm CMOS ProcessabstractIn this paper, we propose a new voltage controlled oscillator (VCO) with a high oscillation frequency yet low power consumption. The oscillator which is a single stage circuit has a low phase noise due to reduced noise sources. To evaluate the performance parameters, the oscillator was simulated in a 0.18-μm standard CMOS process. The results show that the oscillation frequency of VCO may vary between 4.66-5.9 GHz. Also, the phase noise of the VCO at oscillation frequency of 5.6 GHz is -99.7 dBc at 1 MHz offset frequency. Also, the power consumption was 4.8 mW at the same oscillation frequency. Mostafa Savadi Oskooei, Ali Afzali-Kusha, Seyed Mojtaba Atarodi |
ISCAS | 2 |
| 2007 | Clock Delayed Domino Logic With Efficient Variable Threshold Voltage KeeperabstractIn this paper, efficient clock delayed domino logic with variable strength voltage keeper is proposed. The variable strength of the keeper is achieved through applying two different body biases to the keeper. The circuits used to generate the body biases are called capacitive body bias generator and cross-coupled capacitive body bias generator. Compared to a previous work, the body bias generator circuits presented in this paper are simpler and do not require double or triple power supply while consuming less area and power. To show the efficiency of the proposed technique, the implementation of a carry generator circuit by the proposed techniques and the previous work are compared. The simulation results for standard CMOS technologies of 0.18 mum and 70 nm show considerable improvements in terms of power and power delay product. In addition, the proposed technique shows much less temperature dependence when compared to that of previous work Amir Amirabadi, Ali Afzali-Kusha, Y. Mortazavi, Mehrdad Nourani |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | NoC Hot Spot minimization Using AntNet Dynamic Routing AlgorithmabstractIn this paper, a routing model for minimizing hot spots in the network on chip (NOC) is presented. The model makes use of AntNet routing algorithm which is based on Ant colony. Using this algorithm, which we call AntNet routing algorithm, heavy packet traffics are distributed on the chip minimizing the occurrence of hot spots. To evaluate the efficiency of the scheme, the proposed algorithm was compared to the XY, Odd- Even, and DyAD routing models. The simulation results show that in realistic (Transpose) traffic as well as in heavy packet traffic, the proposed model has less average delay and peak power compared to the other routing models. In addition, the maximum temperature in the proposed algorithm is less than those of the other routing algorithms. Masoud Daneshtalab, Ashkan Sobhani, Ali Afzali-Kusha, Omid Fatemi, Zainalabedin Navabi |
ASAP | 3 |
| 2006 | Double edge triggered Feedback Flip-Flop in sub 100NM technologyabstractIn this paper, a new flip-flop called double-edge triggered feedback flip-flop (DFFF) is proposed. The dynamic power consumption of DFFF is reduced by avoiding unnecessary internal node transition. The subthreshold current in the flip-flops is very low compared to other structures. Reducing the number of transistor in the stack and increasing the number of charge path leads to higher operational speed compared to others flip-flops. The simulation results show an improvement of 44% in the speed and 45% in the static leakage power S. H. Rasouli, Amir Amirabadi, A. Seyedi, Ali Afzali-Kusha |
ASP-DAC | 4 |
| 2006 | Dynamic voltage and frequency management based on variable update intervals for frequency settingabstractAn efficient adaptive method to perform dynamic voltage and frequency management (DVFM) for minimizing the energy consumption of microprocessor chips is presented. Instead of using a fixed update interval, the proposed DVFM system makes use of adaptive update intervals for optimal frequency and voltage scheduling. The optimization enables the system to rapidly track the workload changes so as to meet soft real-time deadlines. The method, which is based on introducing the concept of an effective deadline, utilizes the correlation between consecutive values of the workload. In practice because the frequency and voltage update rates are dynamically set based on variable update interval lengths, voltage fluctuations on the power network are also minimized. The technique, which may be implemented by simple hardware and is completely transparent from the application, leads to power savings of up to 60% for highly correlated workloads compared to DVFM systems based on fixed update intervals. Mehrdad Najibi, Mostafa E. Salehi, Ali Afzali-Kusha, Massoud Pedram, Sied Mehdi Fakhraie, Hossein Pedram |
ICCAD | 3 |
| 2006 | Low power and high performance clock delayed domino logic using saturated keeperabstractIn this work, domino logic with a saturated keeper technique is proposed. The circuit, which is used to implement the technique, is as simple as the utilized NOT gate in standard domino. By using the simple structure, we can obtain better performance, noise immunity, and lower power consumption. The simulation results for a 70 nm CMOS technology show an improvement between 7% and 62.5% in delay and 9% and 14% in power consumption, over its previous suggestions Amir Amirabadi, A. Chehelcheraghi, S. H. Rasouli, A. Seyedi, Ali Afzali-Kusha |
ISCAS | 5 |
| 2006 | Power efficient sequential multiplication using pre-computationabstractA pre-computation based technique to lower the power consumption of sequential multipliers is presented. This technique also speeds up the multiplication by reducing the number of clock ticks required to complete a multiplication. The proposed technique may be applied to different sequential multiplication schemes. The benchmark data is extracted from typical DSP applications to show the efficiency of the proposed technique in the domain of DSP computations in which the low power computing is of rapidly increasing importance. The results show an average of 25% reduction in the switching activity and 30% reduction in the clock tick count, compared to sequential multipliers without this technique Nima Honarmand, M. Reza Javaheri, Naser Sedaghati, Ali Afzali-Kusha |
ISCAS | 4 |
| 2006 | High performance circuit techniques for dynamic OR gatesabstractIn this paper, two methods for high fan-in dynamic OR gates are proposed. The methods are called high-speed low-swing OR gate (HSLS-OR) and low-power selective evaluate OR gate (LPSE-OR). HSLS-OR contains separate parallel NMOS logic trees in which one controls the evaluation phase of the other ones. This leads to a low voltage swing in the dynamic capacitive nodes. In LPSE-OR, the NMOS logic tree is divided into several successive parts to prevent using strong keepers. Unnecessary parts are disabled in the evaluation phase for saving the power. (16, 32, and 64)-bit HSLS-OR and (32 and 64)-bit LPSE-OR gates are simulated using HSPICE in 65 nm bulk CMOS technology. Compared to the previous works, the new circuits show 34-48% better power delay product (PDP) Bahman Kheradmand Boroujeni, Fatemeh Aezinia, Ali Afzali-Kusha |
ISCAS | 3 |
| 2006 | A compact low power mixed-signal equalizer for gigabit Ethernet applicationsabstractIn this paper we propose a novel structure of a discrete-time mixed-signal linear equalizer designed for analog front end of Gigabit Ethernet receivers. The circuit is an FIR filter which involves 6 taps based on a coefficient-rotating structure. Here, a simple structure is used for merging digital to analog conversion of the filter's coefficients and multipliers needed for 6 taps. This structure results in high speed and low power dissipation as well as less A/D converter complexity. Simulated in a 0.18 /spl mu/m CMOS technology, this equalizer operates at 125 MHz while dissipating 10 mw from a 1.8 V power supply. Saeid Mehrmanesh, Behzad Eghbalkhah, Saeed Saeedi, Ali Afzali-Kusha, Seyed Mojtaba Atarodi |
ISCAS | 4 |
| 2006 | A very high performance address BUS encoderabstractThis paper presents a very fast and low-power address bus encoder which is less dependent on address bus width. Its encoding scheme is the same as TO-XOR encoding method but its encoder and decoder architectures are much faster. Both the analytical analysis and the simulation results show that the delay of proposed architecture is about one third of delay of optimized TO-XOR encoder/decoder Hadi Parandeh-Afshar, Ali Afzali-Kusha, Ali Khaki-Firooz |
ISCAS | 2 |
| 2006 | WL-VC SRAM: a low leakage memory circuit for deep sub-micron designabstractIn this paper, a static random access memory (SRAM) cell that reduces the gate leakage power with low access latency is proposed. The technique reduces the gate leakage current both in the zero and in the one states. The efficiency of the design is evaluated by simulating the circuit in a 45-nm CMOS technology. Compared to the conventional SRAM cell, the proposed design reduces the total gate leakage current around 58% for an oxide thickness of 1.4nm. The increase in the area of the proposed cell is minimal compared to the conventional SRAM. The read access time of this SRAM is only 5.6% slower than that of the conventional SRAM. G. Razavipour, A. Motamedi, Ali Afzali-Kusha |
ISCAS | 3 |
| 2006 | Low-power multiplier with static decision for input manipulationabstractIn this paper, a method to reduce the power consumption of low-power multipliers is proposed. A decision logic module decides about complementing the input before sending it to the multiplier core. The decision is based on minimizing the switching activity and hence the power consumption. The decision logic proposed in this work, unlike the conventional dynamic decision logic modules, uses a small look up table to make its decisions statically once for all the possible combinations on new input and old input. This feature has lead to the reduction of the power, delay, and complexity of the logic compared to its predecessors while the power saving is negligibly reduced Mohammad Riazati, Ashkan Sobhani, M. Mottaghi-Dastjerdi, Ali Afzali-Kusha, Ali Khaki-Firooz |
ISCAS | 4 |
| 2006 | Low-power and low-latency cluster topology for local traffic NoCsabstractIn this paper, we introduce a topology for network on chips that is named cluster-mesh (CM) topology. This architecture reduces dynamic and static power consumption in NoCs and can reduce latency of communications in low traffic or local traffic applications. With cluster-mesh topology, area reduction in routers is about 44% and in links we can save more than 50% in area too. The dynamic power in this architecture is reduced more than 20%. The idea of clustering may be applied to some other topologies such as Torus and Octagon Mohsen Saneei, Ali Afzali-Kusha, Zainalabedin Navabi |
ISCAS | 2 |
| 2006 | Low power low leakage clock gated static pulsed flip-flopabstractIn this paper, a low power low leakage flip-flop called clock gated static pulsed flip-flop (CGSPFF) is proposed. The dynamic power consumption in CGSPFF is reduced by avoiding unnecessary input pulse transitions with clock gating. Two transistors in the main block of the flip-flop are eliminated to achieve low leakage power as well. Using the new clock pulse generator leads to a higher operational speed and lower power consumption compared to the previously proposed flip-flops. The results of the simulation show that the PDP of the proposed flip-flop is reduced by at least 58.3% A. S. Seyedi, S. H. Rasouli, Amir Amirabadi, Ali Afzali-Kusha |
ISCAS | 4 |
| 2006 | Substrate Noise Coupling in SoC Design: Modeling, Avoidance, and ValidationabstractIssues related to substrate noise in system-on-chip design are described including the physical phenomena responsible for its creation, coupling transmission mechanisms and media, parameters affecting coupling strength, and its impact on mixed-signal integrated circuits. Design guidelines and best practices to minimize the generation, transmission, and reception of substrate noise are outlined, and different modeling approaches and computer simulation methods used in quantifying the noise coupling phenomena are presented. Finally, experiments that validate the modeling approaches and mitigation techniques are reviewed Ali Afzali-Kusha, Makoto Nagata, Nishath K. Verghese, David J. Allstot |
Proc. IEEE | 1 |
| 2005 | Sign bit reduction encoding for low power applicationsabstractThis paper proposes a low power technique, called SBR (Sign Bit Reduction) which may reduce the switching activity in multipliers as well as data buses. Utilizing the multipliers based on this scheme, the dynamic power consumption of some digital systems such as digital filters based on CMOS logic system can be reduced considerably compared to those based on 2's complement implementation. To verify the efficacy of the SBR, a 16-bit multiplier was implemented by this scheme. The results for voice data show an average of 29% to 35% switching reduction compared to the 2's complement implementation. For 16-bit random data, this scheme decreases the switching of 16-bit multipliers by an average of 21%. Finally, the application of the technique to a 16-bit data bus leads up to 14.5% switching reduction on average. Mohsen Saneei, Ali Afzali-Kusha, Zainalabedin Navabi |
DAC | 2 |
| 2005 | Simultaneous Reduction of Dynamic and Static Power in Scan StructuresabstractPower dissipation during test is a major challenge in testing integrated circuits. Dynamic power has been the dominant part of power dissipation in CMOS circuits, however, in future technologies the static portion of power dissipation will outreach the dynamic portion. This paper proposes an efficient technique to reduce both dynamic and static power dissipation in scan structures. Scan cell outputs which are not on the critical path(s) are multiplexed to fixed values during scan mode. These constant values and primary inputs are selected such that the transitions occurring on nonmultiplexed scan cells are suppressed and the leakage current during scan mode is decreased. A method for finding these vectors is also proposed. The effectiveness of this technique is proved by experiments performed on ISCAS89 benchmark circuits. Shervin Sharifi, Javid Jaffari, Mohammad Hosseinabady, Ali Afzali-Kusha, Zainalabedin Navabi |
DATE | 4 |
| 2005 | Optimization of the VT control method for low-power ultra-thin double-gate SOI logic circuits
Davood Shahrjerdi, Bahman Hekmatshoar, Ali Khaki-Firooz, Ali Afzali-Kusha |
Integr. | 4 |
| 2004 | Leakage current reduction by new technique in standby modeabstractIn this paper, a new approach for reducing the subthreshold leakage current of digital circuits is proposed. It does not use a multi-threshold process technique which is more expensive. The technique which makes to use of a new variable supply voltage oscillator, combines the ideas of both Standby Leakage Control Using Transistor Stacks (SRB) and variable threshold (VT) methods. In this technique compared to the Leakage Control Using Transistor Stacks method, the subthreshold current decreases three times. In addition, leakage current monitor (LCM) and its related circuits which are used in VT techniques are not required here. This makes the proposed technique more power and area efficient. Amir Amirabadi, Javid Jaffari, Ali Afzali-Kusha, Mehrdad Nourani, Ali Khaki-Firooz |
ACM Great Lakes Symposium on VLSI | 3 |
| 2004 | Accurate and efficient modeling of SOI MOSFET with technology independent neural networksabstractThis paper presents neural network (NN) approaches for modeling the I-V characteristics of silicon-on-insulator MOSFETs. The modeling approach is technology independent, fast, and accurate, which makes it suitable for circuit simulators. In the model, two different NN architectures, namely, multilayer perceptron and generalized radial basis function, are used and compared. To increase the training efficiency of the NN, both modular and region partitioning methods have been proposed and utilized. In addition, two approaches for obtaining the transconductance and output conductance of the device are discussed. The first approach makes use of an NN for the conductances, while the second uses the numerical differentiation of the I-V results. To confirm the accuracy of the model, the drain-current characteristics as well as conductances obtained by the model are compared to the simulation data for the points where the NNs are not trained. The comparison shows excellent agreements with relative errors of around 1% over a wide range of drain and gate voltages as well as channel lengths and widths. Safar Hatami, M. Yaser Azizi, Hamid-Reza Bahrami 0002, Davoud Motavalizadeh-Naeini, Ali Afzali-Kusha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |