EDBT 2026 Demo / reviewers in the wild / expert
Mehdi Kamal
dblp:06/3034
· DBLP profile ↗
70ranked-venue papers
12as first author
24since 2021 · last 2026
0000-0001-7098-6440ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 60 · 12 first-author · 18 since 2021Software engineering, systems software and programming languages · 10 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 3 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MARCO: Hardware-Aware Neural Architecture Search for Edge Devices with Multi-Agent Reinforcement Learning and Conformal FilteringabstractWe present MARCO (Multi-Agent Reinforcement learning with Conformal Optimization), a hardware-aware neural architecture search (NAS) framework for resource-constrained edge devices. MARCO combines multi-agent reinforcement learning (MARL) with Conformal Prediction (CP) to efficiently explore architectures under strict memory and latency budgets. Unlike once-for-all (OFA) supernets that require expensive pretraining, MARCO separates the NAS task into a Hardware Configuration Agent and a Quantization Agent, coordinated via a centralized-critic, decentralized-execution (CTDE) paradigm. A calibrated CP surrogate model offers distribution-free guarantees to filter low-reward candidates before costly training or simulation, significantly accelerating the search. Experiments on MNIST, CIFAR-10, and CIFAR-100 show MARCO achieves $3-4 \times$ faster search than OFA while maintaining accuracy within 0.3% and reducing latency. Validation on the MAX78000 confirms simulator fidelity with less than 5% error. Arya Fayyazi, Mehdi Kamal, Massoud Pedram |
ASP-DAC | 2 |
| 2026 | FAIR-SIGHT: Fairness Assurance in Image Recognition via Simultaneous Conformal Thresholding and Dynamic Output RepairabstractWe present FAIR-SIGHT, a post-hoc framework that enforces statistical fairness in computer vision models without retraining or access to internal parameters. The method computes a fairness-aware non-conformity score combining prediction error and demographic disparity and uses conformal prediction to calibrate a threshold that guarantees a user-specified violation rate under finite-sample, distribution-free conditions. Inputs exceeding this threshold are automatically repaired through lightweight adjustments (such as logit shifts for classification or confidence scaling for detection) to reduce group disparities while preserving accuracy. Experiments on CelebA, UTKFace, GeoDE, and COCO show over 30% reductions in DPD/EOD and detection gaps with negligible utility loss and only ∼0.5ms/img overhead, establishing FAIR-SIGHT as a practical, scalable solution for bias mitigation in black-box vision systems. Arya Fayyazi, Mehdi Kamal, Massoud Pedram |
WACV | 2 |
| 2026 | Reliable yet high-speed memristor-based mixed-signal coarse-grained reconfigurable architecture for CNN inference acceleration
Reza Kazerooni-Zand, Ali Afzali-Kusha, Mehdi Kamal |
Neurocomputing | 3 |
| 2026 | Guest Editorial: Special Section on International Symposium on Low Power Electronics and Design (ISLPED) 2025
Srividhya Venkataraman, Gregory K. Chen, Deliang Fan, Mehdi Kamal |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | ICD2S: A Hybrid Ising-Classical-Machines Data-Driven QUBO Solver MethodabstractWe present a heuristic algorithm designed to solve Quadratic Unconstrained Binary Optimization (QUBO) problems efficiently. The algorithm, referred to as IC-D2S, leverages a hybrid approach using Ising and classical machines to address very large problem sizes. Considering the practical limitation on the size of the Ising machine (IM), our algorithm partitions the QUBO problem into a collection of QUBO subproblems (called subQUBOs) and utilizes the IM to solve each subQUBO. Our proposed heuristic algorithm uses a set of control parameters to generate the subQUBOs and explore the search space. Also, it utilizes an annealer based on cosine waveform and applies a mutation operator at each step of the search to diversify the solution space and facilitate the process of finding the global minimum of the problem. We have evaluated the effectiveness of our IC-D2S algorithm on three large-sized problem sets and compared its efficiency in finding the (near-)optimal solution with three QUBO solvers. One of the solvers is a software-based algorithm (D2TS), while the other one (D-Wave) employs a similar approach to ours, utilizing both classical and Ising machines. The results demonstrate that for large-sized problems (≥ 5000) the proposed algorithm identifies superior solutions. Additionally, for smaller-sized problems (= 2500), IC-D2S efficiently finds the optimal solution in a significantly faster manner. Armin Abdollahi, Mehdi Kamal, Massoud Pedram |
ASP-DAC | 2 |
| 2025 | Dynamic Co-Optimization Compiler: Leveraging Multi-Agent Reinforcement Learning for Enhanced DNN Accelerator PerformanceabstractThis paper introduces a novel Dynamic Co-Optimization Compiler (DCOC), which employs an adaptive Multi-Agent Reinforcement Learning (MARL) framework to enhance the efficiency of mapping machine learning (ML) models, particularly Deep Neural Networks (DNNs), onto diverse hardware platforms. DCOC incorporates three specialized actor-critic agents within MARL, each dedicated to different optimization facets: one for hardware and two for software. This cooperative strategy results in an integrated hardware/software co-optimization approach, improving the precision and speed of DNN deployments. By focusing on high-confidence configurations, DCOC effectively reduces the search space, achieving remarkable performance over existing methods. Our results demonstrate that DCOC enhances throughput by up to 37.95% while reducing optimization time by up to 42.2% across various DNN models, outperforming current state-of-the-art frameworks. Arya Fayyazi, Mehdi Kamal, Massoud Pedram |
ASP-DAC | 2 |
| 2025 | FACTER: Fairness-Aware Conformal Thresholding and Prompt Engineering for Enabling Fair LLM-Based Recommender SystemsabstractWe propose FACTER, a fairness-aware framework for LLM-based recommendation systems that integrates conformal prediction with dynamic prompt engineering. By introducing an adaptive semantic variance threshold and a violation-triggered mechanism, FACTER automatically tightens fairness constraints whenever biased patterns emerge. We further develop an adversarial prompt generator that leverages historical violations to reduce repeated demographic biases without retraining the LLM. Empirical results on MovieLens and Amazon show that FACTER substantially reduces fairness violations (up to 95.5%) while maintaining strong recommendation accuracy, revealing semantic variance as a potent proxy of bias. Arya Fayyazi, Mehdi Kamal, Massoud Pedram |
ICML | 2 |
| 2025 | An Analog Multiplier Utilizing an Unconventional Bit-Weighting Scheme with Application to Neural Network QuantizationabstractThis paper introduces a dynamic precision analog multiplier architecture for analog mixed-signal machine learning (ML) accelerators. The proposed architecture, built on the C2C ladder structure, enables runtime adjustment of the multiplier bit weights from the least significant bit (LSB) to the most significant bit (MSB), allowing for flexible implementation of mixed-precision ML models. By adjusting the bit position weights of learnable neural network weights, the representable numbers can be tailored to the required precision for a given node or filter in the ML model. To determine the optimal bit weights for our proposed multiplier, we propose a non-uniform quantization-aware training algorithm that trains the multiplier bit weights and fine-tunes the pre-trained ML model weights to utilize the suggested multiplication engine efficiently. We implemented the proposed multiplier in 12nm technology and evaluated its performance on low-bit-width BERT and GPT2 models for eight natural language processing (NLP) tasks. The results show that the 32×32 crossbar of the proposed 4-bit multiplier achieves 100 TOPS/W energy efficiency. Moreover, our proposed analog multiplier achieves scores comparable to those of FP32 models when using 4-bit models, demonstrating its efficacy. Mehdi Kamal, Massoud Pedram |
ISLPED | 1 |
| 2025 | Enhancing Low-Precision Deep Learning: A Posit8 Framework for Energy Efficient DNN TrainingabstractThis paper introduces an innovative framework for training low-precision deep neural networks (DNNs) using the 8-bit Posit (Posit8) number system. The framework utilizes a ‘fake’ quantization strategy where all computations are performed in high precision (e.g., 32-bit floating-point, FP32) while weights, activations, loss, and gradients are quantized to Posit8. This approach significantly reduces off-chip memory requirements during training, lowering energy consumption by decreasing the data bandwidth between computing engines and off-chip memory. Our training framework dynamically adjusts the exponent size within the Posit number system to effectively balance dynamic range and precision needs. Additionally, we incorporate a tensor-wise scaling technique to mitigate precision loss associated with the reduced representation bandwidth of Posit8. We also propose a specialized rounding mechanism for the quantization process from FP32 to Posit8, aimed at minimizing accuracy degradation in low-precision training. To evaluate the effectiveness of our approach, we implemented it in PyTorch and conducted experiments using several benchmark neural networks on the CIFAR-10, CIFAR-100, and Tiny ImageNet datasets. In addition, we compared its performance against an FP8 training framework. Experiment results indicate an accuracy drop of approximately 0.86% compared to high-precision training, with energy consumption reduced by 1.7x. At the same time, our framework achieves an average improvement of up to ~1.1% (and as much as 2.28%) in accuracy over models trained with FP8 quantization with ~2% more energy consumption. Dongyang Wu, Mehdi Kamal, Massoud Pedram |
ISLPED | 2 |
| 2024 | X-IMM: Mixed-Signal Iterative Montgomery Modular MultiplicationabstractIn this paper, we present a mixed-signal implementation of iterative Montgomery multiplication algorithm (called X-IMM) for using in large arithmetic word size (LAWS) computations. LAWS is mainly utilized in security applications such as lattice-based cryptography, where the width of the input operands may be equal to or larger than 1,024 bits. The proposed architecture is based on the iterative implementation of the Montgomery multiplication (MM) algorithm, where some critical parts of the multiplication are computed in the analog domain by mapping them on the memristor crossbar. Using a memristor crossbar reduces the area usage and latency of the modular multiplication unit compared to its fully digital implementation. The devised mixed-signal MM implementation is scalable by cascading the smaller X-IMMs to support dynamically adjustable larger operand sizes at runtime. The effectiveness of the proposed MM structure is assessed in the 45nm technology and the comparative studies show that the proposed 1,024-bit Radix-4 (Radix-16) Montgomery multiplication architecture provides about 13% (22%) higher GOPS/mm2 compared to the state-of-the-art digital ASIC implementations of the iterative MM. Also, owing to analog computing, the proposed structure reduces energy consumption considerably as well. Mehdi Kamal, Massoud Pedram |
ISLPED | 1 |
| 2024 | Low-Precision Mixed-Computation Models for Inference on EdgeabstractThis article presents a mixed-computation neural network processing approach for edge applications that incorporates low-precision (low-width) Posit and low-precision fixed point (FixP) number systems. This mixed-computation approach uses 4-bit Posit (Posit4), which has higher precision around 0, for representing weights with high sensitivity, while it uses 4-bit FixP (FixP4) for representing other weights. A heuristic for analyzing the importance and the quantization error of the weights is presented to assign the proper number system to different weights. In addition, a gradient approximation for Posit representation is introduced to improve the quality of weight updates in the backpropagation process. Due to the high energy consumption of the fully Posit-based computations, neural network operations are carried out in FixP or Posit/FixP. An efficient hardware implementation of an MAC operation with a first Posit operand and FixP for a second operand and accumulator is presented. The efficacy of the proposed low-precision mixed-computation approach is extensively assessed on vision and language models. The results show that on average, the accuracy of the mixed-computation is about 1.5% higher than that of FixP with a cost of 0.19% energy overhead. Seyedarmin Azizi, Mahdi Nazemi, Mehdi Kamal, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | ReMeCo: Reliable Memristor-Based in-Memory Neuromorphic ComputationabstractMemristor-based in-memory neuromorphic computing systems promise a highly efficient implementation of vector-matrix multiplications, commonly used in artificial neural networks (ANNs). However, the immature fabrication process of memristors and circuit level limitations, i.e., stuck-at-fault (SAF), IR-drop, and device-to-device (D2D) variation, degrade the reliability of these platforms and thus impede their wide deployment. In this paper, we present ReMeCo, a redundancy-based reliability improvement framework. It addresses the non-idealities while constraining the induced overhead. It achieves this by performing a sensitivity analysis on ANN. With the acquired insight, ReMeCo avoids the redundant calculation of least sensitive neurons and layers. ReMeCo uses a heuristic approach to find the balance between recovered accuracy and imposed overhead. ReMeCo further decreases hardware redundancy by exploiting the bit-slicing technique. In addition, the framework employs the ensemble averaging method at the output of every ANN layer to incorporate the redundant neurons. The efficacy of the ReMeCo is assessed using two well-known ANN models, i.e., LeNet, and AlexNet, running the MNIST and CIFAR10 datasets. Our results show 98.5% accuracy recovery with roughly 4% redundancy which is more than 20× lower than the state-of-the-art. Ali BanaGozar, Seyed Hossein Hashemi Shadmehri, Sander Stuijk, Mehdi Kamal, Ali Afzali-Kusha, Henk Corporaal |
ASP-DAC | 4 |
| 2023 | Federated learning by employing knowledge distillation on edge devices with limited hardware resources
Ehsan Tanghatari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
Neurocomputing | 2 |
| 2023 | A2P-MANN: Adaptive Attention Inference Hops Pruned Memory-Augmented Neural NetworksabstractIn this work, to limit the number of required attention inference hops in memory-augmented neural networks, we propose an online adaptive approach called [Formula: see text]-memory-augmented neural network (MANN). By exploiting a small neural network classifier, an adequate number of attention inference hops for the input query are determined. The technique results in the elimination of a large number of unnecessary computations in extracting the correct answer. In addition, to further lower computations in [Formula: see text]-MANN, we suggest pruning weights of the final fully connected (FC) layers. To this end, two pruning approaches, one with negligible accuracy loss and the other with controllable loss on the final accuracy, are developed. The efficacy of the technique is assessed by applying it to two different MANN structures and two question answering (QA) datasets. The analytical assessment reveals, for the two benchmarks, on average, 50% fewer computations compared to the corresponding baseline MANNs at the cost of less than 1% accuracy loss. In addition, when used along with the previously published zero-skipping technique, a computation count reduction of approximately 70% is achieved. Finally, when the proposed approach (without zero skipping) is implemented on the CPU and GPU platforms, on average, a runtime reduction of 43% is achieved. Mohsen Ahmadzadeh, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Accuracy Configurable Adders with Negligible Delay Overhead in Exact Operating ModeabstractIn this paper, two accuracy configurable adders capable of operating in approximate and exact modes are proposed. In the adders, which include a block-based carry propagate and a parallel prefix structure, the carry chains are cut off in the approximate mode limiting the carry chain depth to two blocks. In the case of parallel prefix adder, we propose a special carry generate tree equipped with a power gating means. In both of the proposed structures, the critical paths of the adders are not increased in the exact operating mode. Thus, the main objective of proposing these approximate adder structures is to present an accuracy configurable adder structure whose delay in the exact mode is almost the same as an exact adder. The efficacies of the proposed accuracy configurable adders are compared with some state-of-the-art adder structures using a 15nm CMOS technology. In addition, their efficacies are evaluated in two error-resilient applications. These studies show that the proposed carry-propagate adder has 22% (51%) lower energy consumption (error rate) compared to the best prior works. Also, the proposed parallel prefix adder provides, on average, 20% lower energy consumption compared to the exact parallel prefix adders. Farhad Ebrahimi-Azandaryani, Omid Akbari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | Memristive-based Mixed-signal CGRA for Accelerating Deep Neural Network InferenceabstractIn this paper, a mixed-signal coarse-grained reconfigurable architecture (CGRA) for accelerating inference in deep neural networks (DNNs) is presented. It is based on performing dot-product computations using analog computing to achieve a considerable speed improvement. Other computations are performed digitally. In the proposed structure (called MX-CGRA), analog tiles consisting of memristor crossbars are employed. To reduce the overhead of converting the data between analog and digital domains, we utilize a proper interface between the analog and digital tiles. In addition, the structure benefits from an efficient memory hierarchy where the data is moved as close as possible to the computing fabric. Moreover, to fully utilize the tiles, we define a set of micro instructions to configure the analog and digital domains. Corresponding context words used in the CGRA are determined by these instructions (generated by a companion compiler tool). The efficacy of the MX-CGRA is assessed by modeling the execution of state-of-the-art DNN architectures on this structure. The architectures are used to classify images of the ImageNet dataset. Simulation results show that, compared to the previous mixed-signal DNN accelerators, on average, a higher throughput of 2.35 × is achieved. Reza Kazerooni-Zand, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2022 | SySCIM: SystemC-AMS Simulation of Memristive Computation In-MemoryabstractComputation-in-memory (CIM) is one of the most appealing computing paradigms, especially for implementing artificial neural networks. Non-volatile memories like ReRAMs, PCMs, etc., have proven to be promising candidates for the realization of CIM processors. However, these devices and their driving circuits are subject to non-idealities. This paper presents a comprehensive platform, named SysCIM, for simulating memristor-based CIM systems. SySCIM considers the impact of the non-idealities of the CIM components, including memristor device, memristor crossbar (interconnects), analog-to-digital converter, and transimpedance amplifier, on the vector-matrix multiplication performed by the CIM unit. The CIM modules are described in SystemC and SystemC-AMS to reach a higher simulation speed while maintaining high simulation accuracy. Experiments under different crossbar sizes show SySCIM performs simulations up to 117 x faster than HSPICE with less than 4% accuracy loss. The modular design of SySCIM provides researchers with an easy design-space exploration tool to investigate the effects of various non-idealities. Seyed Hossein Hashemi Shadmehri, Ali BanaGozar, Mehdi Kamal, Sander Stuijk, Ali Afzali-Kusha, Massoud Pedram, Henk Corporaal |
DATE | 3 |
| 2022 | Distributing DNN training over IoT edge devices based on transfer learning
Ehsan Tanghatari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
Neurocomputing | 2 |
| 2022 | An Adaptive Memory-Side Encryption Method for Improving Security and Lifetime of PCM-Based Main MemoryabstractIn this article, we present a main memory system for improving the lifetime and security of phase-change main memories. Storing encrypted data increases the bit-flip rates in memory cells, which adversely affects the lifetime of the phase-change memory cells. Thus, to improve the lifetime and security, the proposed system reduces the bit-flip rates by introducing two techniques. The first technique is a memory-side encryption which provides security against DIMM stealing attacks. To prevent unauthorized accesses, in this technique, the encrypted data are not saved in the main memory. As the second technique, we suggest an adaptive partial encryption approach, which makes use of behavior tracking of the application in the CPU side to minimize the latency overhead of the first technique. Additionally, it prevents the loss of data against application-based attacks. This technique uses a recurrent neural network (RNN) to do sequence classification and detect malicious applications. In addition, an auxiliary method, called periodic encryption (PE), which overcomes the security loss in some applications induced by the low accuracy of the employed neural network, is presented. The efficacy of the proposed method is evaluated using gem5 simulator and some benchmarks. Compared to DEUCE and Crypto-Comp methods, the results for the lifetime evaluation show an average bit-flip rate reduction of 11%. In addition, the security improvements against the DIMM stealing and application-based attacks are about 100% and 92.5%, respectively. Morteza Soltani, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Posit Process Element for Using in Energy-Efficient DNN AcceleratorsabstractIn this work, we present an energy-efficient posit processing element (PE) for utilization in array-based deep neural network (DNN) accelerators along with an approximation method for further reducing the energy consumption of the unit. The posit arithmetic used in the proposed PE provides high precision for the considered data widths even when approximation is used for operations. Using some modification/simplification approaches and proposing a speculative posit adder (SPA) unit, we reduce the complexity of the employed posit multiply–accumulator (MAC) in the proposed PE. The effectiveness of the proposed PE is studied using a 45-nm CMOS technology. The results reveal$3.5\times $and 92% improvements in the delay and energy consumption, respectively, compared to those of the state-of-the-art posit PE. To assess the efficacy of the proposed PE, we have modeled an 8-bit DNN accelerator and employed it for the implementation of some DNN architectures. The results indicate that the proposed PE and its approximate one provide, on average, 19.3% and 29.6% lower energy consumptions compared to that of the latest prior work when providing 10.6% and 5.8% higher accuracies, respectively. Mohamadreza Zolfagharinejad, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | Loading-Aware Reliability Improvement of Ultra-Low Power Memristive Neural NetworksabstractIn this paper, a method for offline training of inverter-based memristive neural networks (IM-NNs), called ERIM, is presented. In this method, the output voltage of the inverter is modeled very accurately by considering the loading effect of the memristive crossbar. To properly choose the size of each inverter, its output load and the required slope of its voltage transfer characteristic (VTC) for an acceptable level of resiliency to the circuit element non-idealities are taken into account. The efficacy of ERIM is investigated by comparing its accuracy to those of two recently proposed offline training methods for IM-NNs (RIM and PHAX). The study is performed using IRIS, BCW, MNIST, and Fashion MNIST datasets. Simulation results show that 72% (56%) reduction in average energy consumption of the trained networks is achieved compared to RIM (PHAX) thanks to proper sizing of the inverters. In addition, due to the higher accuracy of the NN mathematical model, ERIM results in significant improvements in the match between the results of high-level modeling and HSPICE simulations while exhibiting lower sensitivity to circuit element variations. Shaghayegh Vahdat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | Reliability Enhancement of Inverter-Based Memristor Crossbar Neural Networks Using Mathematical Analysis of Circuit Non-IdealitiesabstractIn this paper, the sensitivity of the neural network (NN) outputs to device parameter uncertainties (non-idealities) in inverter-based memristor (IM) crossbar neuromorphic circuits is mathematically modeled and verified using exhaustive circuit and system-level simulations. The NN sensitivity is obtained by modeling the sensitivity of theIMneuron output to the non-idealities of its circuit elements. The analysis reveals a higher sensitivity of the output voltage of theIMneuron to the non-idealities of the inverters compared to the conductance variation of the memristors. Among the inverter non-idealities, horizontal shift of the inverters voltage transfer characteristic (VTC) shows the highest impact on the output voltage of the neuron. To reduce the accuracy loss due to the variations, a training approach which includes a sensitivity term in the cost function of the training phase, is suggested. The achievable improvements through the said NN training approach are evaluated. In the evaluation, the California Housing, MNIST, and Fashion MNIST datasets are employed. The results show up to 50% reduction in the NN output variations in the presence of circuit elements’ non-idealities. Shaghayegh Vahdat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | An Energy-Efficient Inference Method in Convolutional Neural Networks Based on Dynamic Adjustment of the Pruning LevelabstractIn this article, we present a low-energy inference method for convolutional neural networks in image classification applications. The lower energy consumption is achieved by using a highly pruned (lower-energy) network if the resulting network can provide a correct output. More specifically, the proposed inference method makes use of two pruned neural networks (NNs), namely mildly and aggressively pruned networks, which are both designed offline. In the system, a third NN makes use of the input data for the online selection of the appropriate pruned network. The third network, for its feature extraction, employs the same convolutional layers as those of the aggressively pruned NN, thereby reducing the overhead of the online management. There is some accuracy loss induced by the proposed method where, for a given level of accuracy, the energy gain of the proposed method is considerably larger than the case of employing any one pruning level. The proposed method is independent of both the pruning method and the network architecture. The efficacy of the proposed inference method is assessed on Eyeriss hardware accelerator platform for some of the state-of-the-art NN architectures. Our studies show that this method may provide, on average, 70% energy reduction compared to the original NN at the cost of about 3% accuracy loss on the CIFAR-10 dataset. Mohammad Ali Maleki, Alireza Nabipour-Meybodi, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2021 | OPTIMA: An Approach for Online Management of Cache Approximation Levels in Approximate Processing SystemsabstractIn this article, we present an approach for adjusting the approximation levels of the cache memories in the memory hierarchy of an approximate processing system. The technique, which is called online management of cache approximation level (OPTIMA), adjusts the approximation levels of the caches under a predefined accuracy constraint. OPTIMA may also be employed for multicore processors, which comprise cores with private and shared caches running applications with different error constraints. To reduce the energy consumption, OPTIMA determines the proper approximation level of each cache memory using heuristic algorithms in two main steps. In the first step, the approximate levels are adjusted to maximize the power efficiency by dropping the application accuracy to a level that still meets a desirable minimum output quality. In the second step, output accuracy variations due to input pattern changes are compensated by fine tuning. We suggest two algorithms (with different adjustment speeds of approximate levels) for the first step and another algorithm for the second step. To assess the efficacy of OPTIMA, we integrate it in the gem5 simulator and simulate some multiprocessor configurations by running eight approximate benchmarks. The results show that the proposed approach provides up to 44% power consumption reduction in the memory hierarchy. Roohollah Yarmand, Mehdi Kamal, Ali Afzali-Kusha, Pooria Esmaeli, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Circuit-Level Techniques for Logic and Memory Blocks in Approximate Computing SystemsxabstractThis article presents an overview of circuit-level techniques used for approximate computing (AC), including both computation and data storage units. After providing some background concept and methodology review, this article proceeds to provide a detailed review of prior art in circuit-level approximation techniques for data path and memory. The focus is on identifying key circuit-level approximation techniques that are applicable to the computational blocks in general and for both volatile and nonvolatile memory circuit technologies. Emphasis is also placed on the error metrics used to assess the output quality of approximate compute and memory units and whether the accuracy setting is dynamically reconfigurable. This article is concluded with a summary of the key distinguishing features of the reviewed prior art. Saba Amanollahi, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
Proc. IEEE | 2 |
| 2020 | X-CGRA: An Energy-Efficient Approximate Coarse-Grained Reconfigurable ArchitectureabstractIn this article, we present an energy-efficient approximate CGRA (X-CGRA). Instead of conventional exact arithmetic units, it employs configurable approximate adders and multipliers in the so-called quality-scalable processing elements (QSPEs). Furthermore, the structure and functionality of the other architectural components, like context memory, are modified based on the quality-scalable operating modes of the QSPEs. The quality reconfigurability of the X-CGRA makes it amenable for both error-resilient and nonresilient applications. To map the applications on the X-CGRA, a mapping technique is proposed that efficiently utilizes the QSPEs and selects appropriate approximation modes in order to lower the energy consumption while satisfying a user-defined quality constraint. We evaluate the efficacy of our X-CGRA for several benchmark applications from different domains, including image/video processing, signal processing, and scientific computations. Different sizes of X-CGRA are synthesized using a 15-nm FinFET technology. For these benchmarks, the results indicate energy consumption reduction of up to $3.21\times $ compared to those of a typical exact CGRA, at the cost of 4% quality loss. Omid Akbari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram, Muhammad Shafique 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Design Exploration of Energy-Efficient Accuracy-Configurable Dadda Multipliers With Improved Lifetime Based on Voltage OverscalingabstractThis article investigates an energy-efficient accuracy-configurable Dadda (X-Dadda) multiplier. The structure employs the voltage overscaling and approximate width setting as the approximation knobs for improving the energy consumption as well as the reliability and lifetime of the multiplier. While the former may be set in the design time as well as the runtime, the latter may only be invoked in the design time. For a given accuracy level, the partial product columns and the overscaled voltage for optimizing the energy are determined. Normally, to have the error within a tolerable limit, the voltage overscaled columns are those at lower bit significances which have higher switching activities. The structure makes use of a low number of level shifters for a low-overhead realization. The approximate columns which start from the first column are contiguous. To further improve the efficiency of the multiplier, four-bit truncation of the multiplier output is also suggested. The efficiency of the X-Dadda structure is investigated using a 15-nm FinFET technology. The results indicate that, for example, when the approximate mode with the mean relative error distance (MRED) of 0.11 is considered, up to 43% energy saving is achieved. In addition, for this case, the Bias temperature instability (BTI)-induced delay degradation of the multiplier decreases up to 9.9% compared to 50% in the case of the exact mode. Also, the impact of process variations on the accuracy of the X-Dadda is studied. Finally, the efficacy of the X-Dadda multiplier, when used in neural networks for image classification and image-processing applications, is assessed. Hassan Afzali-Kusha, Marzieh Vaeztourshizi, Mehdi Kamal, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | POLAR: A Pipelined/Overlapped FPGA-Based LSTM AcceleratorabstractIn this brief, a low resource utilization field-programmable gate array (FPGA)-based long short-term memory (LSTM) network architecture for accelerating the inference phase is presented. The architecture has low-power and high-speed features that are achieved through overlapping the timing of the operations and pipelining the datapath. Moreover, this architecture requires negligible internal memory size for storing the intermediate data leading to low resource utilization and simple routing, which provides lower interconnect delay (higher operating frequency). A designer may adjust the resource utilization (as well as the latency) of the proposed architecture readily at the register-transfer level (RTL) design by adjusting the amount of parallelization. This makes the process of mapping the architecture onto different types of FPGAs, subject to defined constraints, a simple one. The efficacy of the proposed architecture is assessed by implementing an LSTM network on different types of FPGAs. Compared with the recent works, the proposed architecture provides up to about 1.6x , 43.6x , 21.9x , and 114.5x improvements in frequency, power efficiency, GOP/s, and GOP/s/W, respectively. Finally, our proposed architecture operates at 17.64 GOP/s, which is 2.31 faster than the best previously reported results. Erfan Bank Tavakoli, Seyed Abolfazl Ghasemzadeh, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | RandShift: An Energy-Efficient Fault-Tolerant Method in Secure Nonvolatile Main MemoryabstractIn this article, we present a simple, yet energy- and area-efficient method for tolerating the stuck-at faults caused by an endurance issue in secure-resistive main memories. In the proposed method, by employing the random characteristics of the encrypted data encoded by the Advanced Encryption Standard (AES) as well as a rotational shift operation, a large number of memory locations with stuck-at faults could be employed for correctly storing the data. Due to the simple hardware implementation of the proposed method, its energy consumption is considerably smaller than that of other recently proposed methods. The technique may be employed along with other error correction methods, including the error correction code (ECC) and the error correction pointer (ECP). To assess the efficacy of the proposed method, it is implemented in a phase-change memory (PCM)based main memory system and compared with three error tolerating methods. The results reveal that for a stuck-at fault occurrence rate of 10-2and with the uncorrected bit error rate of 2 × 10-3, the proposed method achieves 82% energy reduction compared to the state-of-the-art method. More generally, using a simulation analysis technique, we show that the fault coverage of the proposed method is similar to that of the state-of-the-art method. Morteza Soltani, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Interstice: Inverter-Based Memristive Neural Networks Discretization for Function Approximation ApplicationsabstractIn this article, the accuracy of inverter-based memristive neural networks (NNs) for function approximation applications is improved under the presence of process variations. The improvement is achieved by using a design approach, called INTERSTICE (Inverter-based Memristive Neural Networks Dis cretization for Function Approximation Applications), which discretizes the output values by employing a classifier. More precisely, in the INTERSTICE approach, the output range is divided into K subranges where each subrange is considered as a class. To train the classifier, the training samples are labeled where each label shows belonging to a specific class. To evaluate the efficacy of the design technique, some function approximation applications such as BlackScholes, FFT, K-means, and Sobel are considered. Compared to PHAX, a recently published inverter-based memristive NN, INTERSTICE provides lower mean squared error (MSE) values in the presence of memristor and transistor variations. More specifically, the improvements in the mean of MSE (μMSE) are in the range of 40%-80% when considering 10% variations in the memristor resistance and transistor parameters. In addition, for most of the benchmarks, INTERSTICE improves the μMSE values of the nominal case (the case where all circuit elements are ideal) compared to PHAX. As another advantage compared to the PHAX, in INTERSTICE, digital outputs can be generated based on the selected classes which eliminates the need for an analog-to-digital converter at the output port connected to the digital part of the system. Finally, achieving lower μMSE values using fewer memristors and consuming lower energy is also attainable with this design approach. Shaghayegh Vahdat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | DART: A Framework for Determining Approximation Levels in an Approximable Memory HierarchyabstractIn this article, we propose a framework for determining approximation levels of approximable memories in a memory hierarchy for executing error resilient applications. The framework aims at optimizing the configuration for employing approximate memories in a computing system. It is based on considering data footprints at different memory hierarchy levels and an expected output quality to determine the amount of approximations at each memory hierarchy level. The problem of finding a suitable memory approximation configuration is performed using a branch-and-bound algorithm considering all possible memory approximation arrangements. The best configuration leading to the lowest power consumption when meeting the expected output quality is selected. The efficacy of the proposed framework for two memory hierarchies with different cache topologies is evaluated by comparing energy consumptions of approximate memories with those of the exact memory units in the memory hierarchy under different output accuracy level targets. For example, with 28 dB as a peak signal to noise ratio (PSNR) constraint, the study, which is performed for four image processing applications, indicates up to 54% and 22% power consumption improvements for the SRAM cache and the DRAM memory, respectively. Roohollah Yarmand, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | ACHILLES: Accuracy-Aware High-Level Synthesis Considering Online Quality ManagementabstractIn this paper, we present an accuracy-aware design framework [called accuracy-aware high-level synthesis (Achilles)], which synthesizes a high-level description of an input application with the objective of minimizing the energy consumption of the synthesized circuit. The proposed framework includes two main parts of Achilles and light-weight predictor selection. The framework leverages light-weight error predictors (i.e., machine learning-based classifiers) to achieve more energy reduction by dynamically managing the output quality level (exact or approximate) of the synthesized circuit. To synthesize the input application, first, we exploit a heuristic algorithm to determine the quality level required for each operation in the data flow graph (DFG) representation of the input application. Next, for synthesizing the input application, we propose an effective Achilles algorithm which utilizes the flexibility of the available multiquality arithmetic units in a high-level cell library to synthesize the datapath. To improve the efficiency, the process starts by iteratively reducing the number of functional units required for synthesizing the DFG. Then, a proper light-weight error predictor satisfying the user expected quality is chosen from the available predictors in the framework. Based on the quality requirements, three different quality management modes are considered. The efficacy of the proposed framework is assessed for benchmarks from image and signal processing as well as robotics domains. The study of these benchmarks indicates that Achilles may reduce the energy consumption up to 51% (36% on average), up to 72% (51% on average), and up to 57% (33% on average) in threshold, average, and hybrid modes, respectively, for the studied cases. Moreover, the results show that relative coverage of large errors may be increased from 21% to 55% by employing synthetic minority oversampling technique method. Shayan Tabatabaei Nikkhah, Mahdi Zahedi, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | TOSAM: An Energy-Efficient Truncation- and Rounding-Based Scalable Approximate MultiplierabstractA scalable approximate multiplier, called truncation- and rounding-based scalable approximate multiplier (TOSAM) is presented, which reduces the number of partial products by truncating each of the input operands based on their leading one-bit position. In the proposed design, multiplication is performed by shift, add, and small fixed-width multiplication operations resulting in large improvements in the energy consumption and area occupation compared to those of the exact multiplier. To improve the total accuracy, input operands of the multiplication part are rounded to the nearest odd number. Because input operands are truncated based on their leading one-bit positions, the accuracy becomes weakly dependent on the width of the input operands and the multiplier becomes scalable. Higher improvements in design parameters (e.g., area and energy consumption) can be achieved as the input operand widths increase. To evaluate the efficiency of the proposed approximate multiplier, its design parameters are compared with those of an exact multiplier and some other recently proposed approximate multipliers. Results reveal that the proposed approximate multiplier with a mean absolute relative error in the range of 11%-0.3% improves delay, area, and energy consumption up to 41%, 90%, and 98%, respectively, compared to those of the exact multiplier. It also outperforms other approximate multipliers in terms of speed, area, and energy consumption. The proposed approximate multiplier has an almost Gaussian error distribution with a near-zero mean value. We exploit it in the structure of a JPEG encoder, sharpening, and classification applications. The results indicate that the quality degradation of the output is negligible. In addition, we suggest an accuracy configurable TOSAM where the energy consumption of the multiplication operation can be adjusted based on the minimum required accuracy. Shaghayegh Vahdat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | PX-CGRA: Polymorphic approximate coarse-grained reconfigurable architectureabstractCoarse-Grained Reconfigurable Architectures (CGRAs) provide tradeoff between the energy-efficiency of Application Specific Integrated Circuits (ASICs) and the flexibility of General Purpose Processors (GPPs). State-of-the-art CGRAs only support exact architectures and precise application executions. However, a majority of the streaming applications such as multimedia and digital signal processing, which are amenable to CGRAs, are inherently error resilient. Therefore, these applications can greatly benefit from the emerging trend of Approximate Computing that leverages this error-resiliency to provide higher energy efficiency proportional to the tolerable accuracy loss (can even be constrained). This paper, for the first time, introduces the novel concept of Polymorphic Approximate CGRA (PX-CGRA) that employs heterogeneous tiles of Polymorphic-Approximated ALU Clusters (PACs) connected in a 2-D mesh style connection. These PACs can implement different approximate modes as well as accurate modes depending upon their selected configuration as per the run-time requirements of executing applications. For designing an efficient PX-CGRA, we propose a bottom-up design flow. In addition, the flow of application mapping on PX-CGRA is discussed including accuracy-level mapping, scheduling, and binding steps. To comprehensively evaluate the efficacy of the proposed CGRA, the complete PX-CGRA architecture in different sizes as well as with different PACs configurations are synthesized using a 15-nm FinFET technology. Our results show up to 15%-45% energy efficiency improvement for 5%-35% output quality degradation, respectively, when compared to the state-of-the-art exact-mode CGRA. Our proposed architecture and design methodology enable a new era of accuracy-configurable CGRAs to provide significant energy gains. Omid Akbari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram, Muhammad Shafique 0001 |
DATE | 2 |
| 2018 | Energy Consumption and Lifetime Improvement of Coarse-Grained Reconfigurable Architectures Targeting Low-Power Error-Tolerant ApplicationsabstractIn this work, the application of a voltage over-scaling (VOS) technique for improving the lifetime and reliability of coarse-grained reconfigurable architectures (GCRAs) is presented. The proposed technique, which may be applied to CGRAs used as accelerators for low-power, error-tolerant applications, reduces the (strongly voltage-dependent) wearout effects and the energy consumption of processing elements (PEs) whenever the error impact on the output quality degradation can be tolerated. This provides us with the ability to lessen the wearout and reduce energy consumption of PEs when accuracy requirement for the results is rather low. Multiple degrees of computational accuracy can be achieved by using different overscaled voltage levels for the PEs. The efficacy of the proposed technique is studied by considering the bias temperature instability. The study is performed for two error-resilient applications. The CGRAs are implemented with 15nm FinFET operating at a nominal supply voltage of 0.8V. In addition, supply voltages of 0.75, 0.7, 0.65, and 0.6V are considered as overscaled voltage levels for this technology. Based on the quality constraint requirements of the benchmarks, optimum overscaled voltage levels for various PEs are determined and utilized. The approach may provide considerable lifetime and energy consumption improvements over those of the conventional exact and approximate computation approaches. Hassan Afzali-Kusha, Omid Akbari, Mehdi Kamal, Massoud Pedram |
ACM Great Lakes Symposium on VLSI | 3 |
| 2018 | An Energy-Efficient, Yet Highly-Accurate, Approximate Non-Iterative DividerabstractIn1 this paper, we present a highly accurate and energy efficient non-iterative divider, which uses multiplication as its main building block. In this structure, the division operation is performed by first reforming both dividend and divisor inputs, and then multiplying the rounded value of the scaled dividend by the reciprocal of the rounded value of the scaled divisor. Precisely, the interval representing the fractional value of the scaled divisor is partitioned into non-overlapping sub-intervals, and the reciprocal of the scaled divisor is then approximated with a linear function in each of these sub-intervals. The efficacy of the proposed divider structure is assessed by comparing its design parameters and accuracy with state-of-the-art, non-iterative approximate dividers as well as exact dividers in 45nm digital CMOS technology. Circuit simulation results show that the mean absolute relative error of the proposed structure for doing 1 32-bit division is less than 0.2%, while the proposed structure has significantly lower energy consumption than the exact divider. Finally, the effectiveness of the proposed divider in one image processing application is reported and discussed. Marzieh Vaeztourshizi, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
ISLPED | 2 |
| 2018 | Lifetime improvement by exploiting aggressive voltage scaling during runtime of error-resilient applications
Farzaneh Nakhaee, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram, Sied Mehdi Fakhraie, Hamed Dorosti |
Integr. | 2 |
| 2018 | An Ultra Low-Power Memristive Neuromorphic Circuit for Internet of Things Smart SensorsabstractIn this paper, we propose an ultra low-power analog neuromorphic circuit to be trained to process sensory data in the Internet of Things smart sensors where low-power and are efficient computing is required. To reduce the operating voltage of the circuit while maintaining the performance, we focus on designing a memristive neuromorphic circuit without employing operational amplifiers. Therefore, we use the CMOS inverters as the neurons in our memristive neuromorphic circuit. We also propose ultra low-power mixed-signal input/output interfaces to make the circuit connectable to other digital components such as embedded processor. To assess the efficacy of the proposed circuit and its interfaces which include memristive neural network based A/D and D/A converters, HSPICE simulations are utilized. The results indicate that at the operating voltage of ±0.25 V, at least 108× (278×) reduction in the power consumption of the output (input) interface compared to that of the conventional structures is achieved. Additionally, the effectiveness of the neuromorphic circuit enhanced by the proposed interfaces is evaluated under some applications such as image recognition, human behavior analysis, and air quality predictions. The results of the study reveal that the designed neuromorphic circuits, along with the proposed A/D and D/A converters, provide an average power saving (speedup) of 2960× (37×) over the ASIC implementation in a 90-nm CMOS technology. Arash Fayyazi, Mohammad Ansari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Internet Things J. | 3 |
| 2018 | PHAX: Physical Characteristics Aware Ex-Situ Training Framework for Inverter-Based Memristive Neuromorphic CircuitsabstractIn this paper, we propose a training framework for an inverter-based memristive neuromorphic hardware. The framework, which is called PHAX, is a physical characteristics aware one relying on anex-situtraining approach. The considered neuromorphic circuit is highly energy efficient hybrid CMOS-memristive implementation of neuromorphic circuits. To solve the problem of high sensitivity of the training to the mismatches between the high-level mathematical modeling of the neurons and the corresponding physical characteristics, an approach for analytical yet accurate modeling of the memristive crossbar and neuron circuits is suggested. The approach, which is based on SPICE simulations, models the inverter-based neurons using a hyperbolic tangent function. To increase the training efficacy, the backpropagation training algorithm is modified by considering some constraints based on the physical characteristics of the memristive circuit. This modification along with the accurate back-annotation of the physical characteristics considerably improve the effectiveness of theex-situtraining method of the neuromorphic circuit. The results of this paper show an average reduction of 1805× in the training runtime compared to that of thein-situtraining approach. Furthermore, the results of applying the approach on the kernels of some applications such as image recognition, image processing, and financial analysis reveal that the designed neuromorphic circuits provide an average power saving (speed up) of 1478× (5.2×) over the ASIC implementation in a 90-nm CMOS technology. Mohammad Ansari, Arash Fayyazi, Ali BanaGozar, Mohammad Ali Maleki, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2018 | TheSPoT: Thermal Stress-Aware Power and Temperature Management for Multiprocessor Systems-on-ChipabstractThermal stress including temperature gradients in time and space, as well as thermal cycling, influences lifetime reliability and performance of modern multiprocessor systems-on-chip (MPSoCs). Conventional power and temperature management techniques considering the peak temperature/power consumption do not provide a comprehensive solution to avoid high spatial and temporal thermal variations. This work presents TheSPoT, a novel multilevel thermal stress-aware power and thermal management approach for MPSoCs. At the top level, core consolidation and deconsolidation is performed based on peak temperature, thermal stress, and power consumption constraints. These constraints are also used at the next level, where operating frequencies are determined. At this level, we obtain optimal core frequencies by solving a convex optimization problem. However, thereafter, to reduce the runtime overhead in large MPSoCs, we alternatively propose to use a fast heuristic algorithm. The efficacy of the proposed approaches in reducing the thermal cycles and temporal/spatial temperature gradients is evaluated by comparing the results with the state-of-the-art methods. The evaluation performed on 4-core, 8-core, and 16-core MPSoCs, using PARSEC benchmarks, reveals a considerable reduction in thermal stress. For the 8-core MPSoC case study, on average, for the proposed heuristic(optimal) approach, the mean time to failure improved by 47(35)% compared to the state-of-the-art techniques with only 6(4)% performance degradation. Also, our simulations show that TheSPoT is more efficient in thermal stress reduction when more heterogeneous workloads are used. Arman Iranfar, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram, David Atienza 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | An Efficient False Path-Aware Heuristic Critical Path Selection Method with High Coverage of the Process Variation SpaceabstractIn this article, we present a critical path selection method that efficiently finds true (sensitizable) critical paths of a circuit in the presence of process variations. The method, which is based on the viability analysis, tries to select the least number of true critical paths that cover all of circuit critical gates. Critical gates are those that make a path critical with a probability higher than a predefined threshold value. Selecting fewer critical paths leads to less computation time for the algorithm and shorter test time of fabricated chips. For this purpose, an efficient Statistical Static Timing Analysis– (SSTA) based technique is suggested. This technique tries to find circuit-critical gates whose process parameter variations cover a major part of the process space. Improving the process space coverage using fewer paths is achieved by considering both spatial (proximity of gates) and structural (having common gates) correlations in the analysis of choosing the critical paths. In the selection process, paths with low similarities in their characteristics are preferred. In addition, only true paths whose delays affect the maximum delay of the circuit are included. The selected paths can be used in the test process of the fabricated chips to determine if the chip meets its timing requirements. Also, a modified viability analysis that incorporates statistical computations is used in the SSTA. The efficacy of the proposed method is evaluated by comparing its results for combinational and sequential ISCAS benchmarks with those obtained by exhaustive search. Results indicate although, on average, only 4.38% of all the critical paths found by the exhaustive search are selected by the proposed method, the maximum probability of criticality for the paths that are not considered in our method is, on average, less than 4%. Sheis Abolma'ali, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2018 | Approximate Reverse Carry Propagate Adder for Energy-Efficient DSP Applications
Masoud Pashaeifar, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Robust neuromorphic computing in the presence of process variationabstractIn this paper, an approach for increasing the sustainability of inverter-based memristive neuromorphic circuits in the presence of process variation is presented. The approach works based on extracting the impact of process variations on the neurons characteristics during the test phase through a proposed algorithm. In this method, first, some combinations of inputs and weights (based on the neuromorphic circuit structure) are injected into the circuit and the features of the neurons are determined. Next, these features which are back-annotated, are utilized in an efficient ex-situ training approach to determine the proper weights of the neurons. The approach provides a considerable improvement in the output accuracy. To evaluate the effectiveness of the proposed approach, some approximate applications are studied using 90nm CMOS technology. The results of the study reveal that using this framework provides, on average, 17X higher output accuracy compared to the cases that the impact of the process variation is not considered at all. Ali BanaGozar, Mohammad Ali Maleki, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
DATE | 3 |
| 2017 | TruncApp: A truncation-based approximate divider for energy efficient DSP applicationsabstractIn this paper, we present a high speed yet energy efficient approximate divider where the division operation is performed by multiplying the dividend by the inverse of the divisor. In this structure, truncated value of the dividend is multiplied exactly (approximately) by the approximate inverse value of divisor. To assess the efficacy of the proposed divider, its design parameters are extracted and compared to those of a number of prior art dividers in a 45nm CMOS technology. Results reveal that this structure provides 66% and 52% improvements in the area and energy consumption, respectively, compared to the most advanced prior art approximate divider. In addition, delay and energy consumption of the division operation are reduced about 94.4% and 99.93%, respectively, compared to those of an exact SRT radix-4 divider. Finally, the efficacy of the proposed divider in image processing application is studied. Shaghayegh Vahdat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram, Zainalabedin Navabi |
DATE | 2 |
| 2017 | CL-CPA: A hybrid carry-lookahead/carry-propagate adder for low-power or high-performance operation mode
Milad Bahadori, Mehdi Kamal, Ali Afzali-Kusha, Yasmin Afsharnezhad, Elham Zahraie Salehi |
Integr. | 2 |
| 2017 | Hybrid TFET-MOSFET circuit: A solution to design soft-error resilient ultra-low power digital circuit
Maede Hemmat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
Integr. | 2 |
| 2017 | Efficient Critical Path Identification Based on Viability Analysis Method Considering Process VariationsabstractIn this brief, we propose an effective adaptation of viability analysis in statistical static timing analysis. The adaption benefits well from a dynamic programming implementation of the viability function. For a rapid identification of statistical longest true paths, the technique makes use of a fast preprocessing step identifying the gates with a small probability of being viable in the circuit, and a number of simple optimization techniques. This makes the approach fast without lowering its accuracy. The efficacy of the proposed statistical timing analysis is assessed using ISCAS benchmark circuits and carry skip adders. The results show that the proposed technique leads to, on average, 18× higher speed compared to those of the state-of-the-art technique. This improvement is achieved at the cost of -1.7% precision lost compared to that of the Monte-Carlo method. Sheis Abolma'ali, Nika Mansouri-Ghiasi, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Dual-Quality 4: 2 Compressors for Utilizing in Dynamic Accuracy Configurable MultipliersabstractIn this paper, we propose four 4:2 compressors, which have the flexibility of switching between the exact and approximate operating modes. In the approximate mode, these dual-quality compressors provide higher speeds and lower power consumptions at the cost of lower accuracy. Each of these compressors has its own level of accuracy in the approximate mode as well as different delays and power dissipations in the approximate and exact modes. Using these compressors in the structures of parallel multipliers provides configurable multipliers whose accuracies (as well as their powers and speeds) may change dynamically during the runtime. The efficiencies of these compressors in a 32-bit Dadda multiplier are evaluated in a 45-nm standard CMOS technology by comparing their parameters with those of the state-of-the-art approximate multipliers. The results of comparison indicate, on average, 46% and 68% lower delay and power consumption in the approximate mode. Also, the effectiveness of these compressors is assessed in some image processing applications. Omid Akbari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | RoBA Multiplier: A Rounding-Based Approximate Multiplier for High-Speed yet Energy-Efficient Digital Signal ProcessingabstractIn this paper, we propose an approximate multiplier that is high speed yet energy efficient. The approach is to round the operands to the nearest exponent of two. This way the computational intensive part of the multiplication is omitted improving speed and energy consumption at the price of a small error. The proposed approach is applicable to both signed and unsigned multiplications. We propose three hardware implementations of the approximate multiplier that includes one for the unsigned and two for the signed operations. The efficiency of the proposed multiplier is evaluated by comparing its performance with those of some approximate and accurate multipliers using different design parameters. In addition, the efficacy of the proposed approximate multiplier is studied in two image processing applications, i.e., image sharpening and smoothing. Reza Zendegani, Mehdi Kamal, Milad Bahadori, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | SEERAD: A high speed yet energy-efficient rounding-based approximate divider
Reza Zendegani, Mehdi Kamal, Arash Fayyazi, Ali Afzali-Kusha, Saeed Safari, Massoud Pedram |
DATE | 2 |
| 2016 | Hybrid TFET-MOSFET circuits: An approach to design reliable ultra-low power circuits in the presence of process variationabstractIn this work, to increase the timing yield of Tunnel Field Effect Transistor (TFET) circuits in the presence of the process variation, we propose to use MOSFET-based gates instead of some TFET-based gates in the TFET circuits. This hybridization approach originates from the fact that TFETs are more sensitive to process variation, when compared to conventional MOSFETs. First, we investigate the impact of process variations on Homojunction InAs TFETs by extracting the distributions of electrical parameters such as threshold voltage. Then, a hybrid TFET-MOSFET circuit design approach for increasing the reliability of the TFET circuits is introduced. The power consumptions of hybrid circuits are considerably smaller than the corresponding ones realized using CMOS circuits. In the proposed hybrid approach, the circuit is basically implemented in TFET to reduce the power and energy consumption while the gates whose their variations may lead to the timing violation, are implemented using MOSFET-based gates. The decision on replacing the TFET-based gates by their corresponding MOSFET-based gates during the hybrid design is made through a heuristic algorithm. The proposed algorithm considers the sensitivity of each TFET-based gate to the process variation. To assess the efficacy of the proposed approach, the proposed algorithm is applied to some circuits of the ISCAS'85 and ISCAS'89 benchmark packages. The results show that the reliability of the TFET-MOSFET-based circuits are up to 74% larger than that of the pure TFET-based circuits. Furthermore, the energy and leakage power consumptions of the proposed hybrid circuits are up to 56% and 80%, respectively, smaller than those of the pure MOSFET-based design. Maede Hemmat, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
VLSI-SoC | 2 |
| 2016 | Power and energy reduction of racetrack-based caches by exploiting shared shift operationsabstractIn this paper, we propose a technique for reducing the power and energy consumptions of the racetrack-based caches. The technique uses a mapping method from the logical cache lines to the physical domains of the nanowires. The mapping method exploits the fact that, in a nanowire with several access heads, the shift operations are shared by the heads on that nanowire. Utilizing this inherent sharing, fewer nanowires are shifted to make a cache line available for both the read and write accesses. By using this method, the cache sets are shifted separately, which results in increase in the number of average shift operations. Thanks to the sharing of the shift operations among multiple heads, the total power and energy consumption of the shift operations are reduced. The effectiveness of the proposed technique is studied using the PARSEC benchmark package. The study shows that the power, energy consumption, energy-delay-product, and energy-delay-squared-product of L2 caches are reduced, on average, by 53%, 44%, 32%, 17%, respectively, compared to the state-of-the-art mapping methods. Seyed Saber Nabavi Larimi, Mehdi Kamal, Ali Afzali-Kusha, Hamid Mahmoodi |
VLSI-SoC | 2 |
| 2016 | A comparative study on performance and reliability of 32-bit binary adders
Milad Bahadori, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
Integr. | 2 |
| 2016 | All-Region Statistical Model for Delay Variation Based on Log-Skew-Normal DistributionabstractIn this paper, we propose a single probability density function for the distributions of the delay in the presence of the process variation for different regions of operation. The delay variation model is inspired by considering the analytical current models for each operating region. Based on these models, we suggest using the log-skew-normal distribution for modeling the delay variation for a wide range of supply voltages from the subthreshold to above-threshold regions. To assess the accuracy of the proposed delay distribution, the mean, standard deviation, skewness, 99th percentile, and yield of the proposed distribution are compared with those of the normal and log-normal distributions using the Monte Carlo (MC) simulations for different circuits in both bulk and FinFET technologies. The results show a higher accuracy for the proposed distribution in all regions of operation. Also, the proposed model enables us to obtain the 3σ yield of the distribution using up to 3.4 times less MC simulation time. Hadi Ahmadi Balef, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2016 | Yield and Speedup Improvements in Extensible Processors by Allocating Extra Cycles to Some Custom Instructions
Mehdi Kamal, Ali Afzali-Kusha, Saeed Safari, Massoud Pedram |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2016 | High-Speed and Energy-Efficient Carry Skip Adder Operating Under a Wide Range of Supply Voltage LevelsabstractIn this paper, we present a carry skip adder (CSKA) structure that has a higher speed yet lower energy consumption compared with the conventional one. The speed enhancement is achieved by applying concatenation and incrementation schemes to improve the efficiency of the conventional CSKA (Conv-CSKA) structure. In addition, instead of utilizing multiplexer logic, the proposed structure makes use of AND-OR-Invert (AOI) and OR-AND-Invert (OAI) compound gates for the skip logic. The structure may be realized with both fixed stage size and variable stage size styles, wherein the latter further improves the speed and energy parameters of the adder. Finally, a hybrid variable latency extension of the proposed structure, which lowers the power consumption without considerably impacting the speed, is presented. This extension utilizes a modified parallel structure for increasing the slack time, and hence, enabling further voltage reduction. The proposed structures are assessed by comparing their speed, power, and energy parameters with those of other adders using a 45-nm static CMOS technology for a wide range of supply voltages. The results that are obtained using HSPICE simulations reveal, on average, 44% and 38% improvements in the delay and energy, respectively, compared with those of the Conv-CSKA. In addition, the power-delay product was the lowest among the structures considered in this paper, while its energy-delay product was almost the same as that of the Kogge-Stone parallel prefix adder with considerably smaller area and power consumption. Simulations on the proposed hybrid variable latency CSKA reveal reduction in the power consumption compared with the latest works in this field while having a reasonably high speed. Milad Bahadori, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | A thermal stress-aware algorithm for power and temperature management of MPSoCs
Mehdi Kamal, Arman Iranfar, Ali Afzali-Kusha, Massoud Pedram |
DATE | 1 |
| 2015 | A heuristic machine learning-based algorithm for power and thermal management of heterogeneous MPSoCsabstractIn this work, we propose a power and thermal management algorithm based on machine learning to control the thermal stresses and power consumption of the heterogeneous MPSoCs. The objectives of the proposed algorithm are increasing the performance and decreasing the spatial and temporal temperature gradients along with the thermal cycling under the power and temperature constraints. Our proposed power and thermal management method is based on a heuristic approach to speed up the convergence of the machine learning algorithm which makes it applicable for general purpose processors. Adopting Q-Learning as the machine learning algorithm, the heuristic approach aids to limit the learning space by suggesting the most appropriate actions to the agent in each decision epoch. The heuristic algorithm employs the current and previous states of the machine learning, as well as the amount of the temperature stress and power consumption of each core to determine the appropriate action for each core, independently. The proposed algorithm is evaluated on 4-core, 8-core and 16-core homogeneous and heterogeneous MPSoCs for some benchmarks in the Splash2 benchmark package. The results reveal a faster convergence of machine learning and more thermal stresses reduction. Arman Iranfar, Soheil Nazar Shahsavani, Mehdi Kamal, Ali Afzali-Kusha |
ISLPED | 3 |
| 2015 | Design of NBTI-resilient extensible processors
Mehdi Kamal, Ali Afzali-Kusha, Saeed Safari, Massoud Pedram |
Integr. | 1 |
| 2015 | OPLE: A Heuristic Custom Instruction Selection Algorithm Based on Partitioning and Local Exploration of Application Dataflow GraphsabstractIn this article, a heuristic custom instruction (CI) selection algorithm is presented. The proposed algorithm, which is called OPLE for “Optimization based on Partitioning and Local Exploration,” uses a combination of greedy and optimal optimization methods. It searches for the near-optimal solution by reducing the search space based on partitioning the identified CI set. The partitioning of the identified set guarantees the success of the algorithm independent of the size of the identified set. First, the algorithm finds the near-optimal CIs from the candidate CIs for each part. Next, the suggested CIs from different parts are combined to determine the final selected CI set. To improve the set of the selected CIs, the solution is evolved by calling the algorithm iteratively. The efficacy of the algorithm is assessed by comparing its performance to those of optimal and nonoptimal methods. A comparative study is performed for a number of benchmarks under different area budgets and I/O constraints. The results reveal higher speedups for the OPLE algorithm, especially for larger identified candidate sets and/or small area budgets compared to those of the nonoptimal solutions. Compared to the nonoptimal techniques, the proposed algorithm provides 30% higher speedup improvement on average. The maximum improvement is 117%. The results also demonstrate that in many cases OPLE is able to find the optimal solution. Mehdi Kamal, Ali Afzali-Kusha, Saeed Safari, Massoud Pedram |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2014 | Improving efficiency of extensible processors by using approximate custom instructionsabstractIn this paper, we propose to move the conventional extensible processor design flow to the approximate computing domain to gain more speedup. In this domain, the instruction set architecture (ISA) design flow selects both exact and approximate custom instructions (CIs). The proposed approach could be used for the applications where imprecise results may be tolerated. In the CI identification phase of the flow, the CIs which do not satisfy the maximum propagation delay but can provide approximate results also may be included in the CI candidate set. Next, in the selection phase, we propose a merit function which selects CIs with higher cycle savings and small error rates. The efficacy of the proposed approximate design flow is investigated using the case studies of the discrete cosine transform (DCT) and inverse DCT (iDCT) of the MPEG2 application. Also, the impact of the process variation on the impreciseness of the results is investigated. Mehdi Kamal, Amin Ghasemazar, Ali Afzali-Kusha, Massoud Pedram |
DATE | 1 |
| 2014 | Impact of Process Variations on Speedup and Maximum Achievable Frequency of Extensible ProcessorsabstractIn this article, we investigate the impact of process variations on the speedup and maximum frequency of the extended ISA processor. First, without considering process variations, a custom functional unit (CFU) is designed based on nominal timing parameters, then the timing variations of critical paths of the extensible processor, including the baseline processor and the CFU, are investigated by considering both systematic and random variations. Next, the maximum frequency of the extensible processor and the speed enhancement factor of the extended ISA for different benchmarks are investigated. Results show that timing variation could reduce the speedup of the extensible processor. However, this reduction is highly dependent on the baseline processor and the CFU structures. Additionally, the impact of process variations in the worst-case design approach is studied. Results show that the speedup of the extensible processor is reduced more than in the case when custom instructions (CIs) are selected without considering process variations. To study the impact of each variation type, speedup variations due to random and systematic variations are investigated separately. The study reveals that random variation has a similar effect on the CFU and the baseline processor, while the impact of systematic variation on the baseline processor is greater than the CFU. Mehdi Kamal, Ali Afzali-Kusha, Saeed Safari, Massoud Pedram |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2013 | An efficient network on-chip architecture based on isolating local and non-local communicationsabstractIn this paper, we propose a scheme for reducing the latency of packets transmitted via on-chip interconnect network in MultiProcessor Systems on Chips (MPSoCs). In this scheme, the network architecture separates the packets transmitted to near destinations from those transmitted to distant ones by using two network layers. These two layers are realized by dividing the channel width among the cores. The optimum ratio for the channel width division is a function of relative significances of the two types of communications. Simulation results indicate that for non-uniform traffic constituting of more than 30 percent local traffic, the proposed network, on average provides 64% and 70% improvement over the conventional one in terms of average network latency and Energy-Delay product (EDP), respectively. Also, for uniform and NED traffic patterns, by adjusting the number of hops between local nodes to include approximately 55 percent of total communications in local ones, the proposed architecture provides the latency reduction of 50%. Vahideh Akhlaghi, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
DATE | 2 |
| 2012 | An architecture-level approach for mitigating the impact of process variations on extensible processorsabstractIn this paper, we present an architecture-level approach to mitigate the impact of process variations on extended instruction set architectures (ISAs). The proposed architecture adds one extra cycle to execute custom instructions (CIs) that violate the maximum allowed propagation delay due to the process variations. Using this method, the parametric yield of manufactured chips will greatly improve. The cost is an increase in the cycle latency of some of the CIs, and hence, a slight performance degradation for the extensible processor architectures. To minimize the performance penalty of the proposed approach, we introduce a new merit function for selecting the CIs during the selection phase of the ISA extension design flow. To evaluate the efficacy of the new selection method, we compare the extended ISAs obtained by this method with those selected based on the worst-case delay. Simulation results reveal that a speedup improvement of about 18% may be obtained by the proposed selection method. Also, by using the proposed merit function, the proposed architecture can improve the speedup about 20.7%. Mehdi Kamal, Ali Afzali-Kusha, Saeed Safari, Massoud Pedram |
DATE | 1 |
| 2012 | An efficient reliability simulation flow for evaluating the hot carrier injection effect in CMOS VLSI circuitsabstractHot carrier injection (HCI) effect is one of the major reliability concerns in VLSI circuits. This paper presents a scalable reliability simulation flow, including a logic cell characterization method and an efficient full chip simulation method, to analyze the HCI-induced transistor aging with a fast run time and high accuracy. The transistor-level HCI effect is modeled based on the Reaction-Diffusion (R-D) framework. The gate-level HCI impact characterization method combines HSpice simulation and piecewise linear curve fitting. The proposed characterization method reveals that the HCI effect on some transistors is much more significant than the others according to the logic cell structure. Additionally, during the circuit simulation, pertinent transitions are identified and all cells in the circuit are classified into two groups: critical and non-critical. The proposed method reduces the simulation time while maintaining high accuracy by applying fine granularity simulation time steps to the critical cells and coarse granularity ones to the non-critical cells in the circuit. Mehdi Kamal, Qing Xie 0001, Massoud Pedram, Ali Afzali-Kusha, Saeed Safari |
ICCD | 1 |
| 2011 | Timing variation-aware custom instruction extension techniqueabstractIn this paper, we propose a technique for custom instruction (CI) extension considering process variations. It bridges the gap between the high level custom instruction extension and chip fabrication in nanotechnologies. In the proposed method, instead of using the conventional static timing analysis (STA), statistical static timing analysis (SSTA) which in turn results in a probabilistic approach to identifying and selecting different parts of the CI extension is utilized. More precisely, we use the delay Probability Density Function (PDF) of the CIs in identification and selection phases of the CI extension. In the identification phase, the delay of each CI is modeled by PDF whereas the performance yield is added as a constraint. Additionally, in the selection phase, the merit function of the conventional approaches is modified to increase the performance gain of the selected CIs at the price of slightly sacrificing the design yield. Also, to make the approach computationally more efficient, we propose a method for reducing the modeling time of the PDF of the CIs by reducing the number of candidate CIs before extracting the PDF. Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
DATE | 1 |
| 2011 | Securing Embedded Processors against Power Analysis Based Side Channel Attacks Using Reconfigurable ArchitectureabstractPower analysis based side channel attacks are significant security risk in embedded applications. Reconfigurable architecture has already been proposed as a security improvement method for run time monitoring systems or implementing critical parts of cryptographic applications. Here we propose reconfigurable architecture as a hardware countermeasure against power analysis based side channel attacks. We augment an embedded processor with a reconfigurable functional unit (RFU). By random execution of custom instructions on the RFU we mask power analysis based side channel attacks. Moreover we devised an automatic design flow to generate RFU and its configuration bits from cryptographic algorithms' source code. Obfuscation of base processor power traces as our primary goal is complied with RFUs covering in average about 30% of object code. We also report the power traces for our secure processor and the overall timing and area overhead. The correlation coefficient is calculated for AES and SHA cryptographic algorithms and experimental results show our method produces power traces close to random traces. Our approach is completely generic and can be used for any cryptographic application. Compared to previous methods, our work costs no runtime overhead and an average of 27% area overhead. Sahar Abbaspour Seyyedi, Mehdi Kamal, Hamid Noori, Saeed Safari |
EUC | 2 |
| 2010 | Dual-purpose custom instruction identification algorithm based on Particle Swarm OptimizationabstractExtending instruction set architecture (ISA) of embedded processors is an effective way to enhance performance and energy efficiency. The typical approaches for identifying custom instructions (CIs) limit the maximum number of input and output (I/O) operands to the available register file port. Recently, there are several work that explore CI candidates without imposing a limit on the number of input and output operands. In this paper, we present a new algorithm based on Particle Swarm Optimization (PSO) to identify CIs within a given data flow graph (DFG) and evaluate it for both categories of CI identification approaches (with and without I/O constrains). By novel evolving strategy, we enhance the quality of the results in our partitioning algorithm. Experimental results show that in most cases CI identification with I/O constraints based on PSO finds better or the same CIs in terms of performance compared to genetic algorithm (GA)[1] and ISEGEN [2] (96% and 90%, respectively). Comparing our proposed algorithm with [12] and [13] reveals that ours has a shorter run-time several order of magnitudes for large DFGs and is independent of the number of forbidden nodes. Moreover, we propose a modified version of PSO called Wrapper PSO that is up to 100× and 500× faster than GA and ISEGEN in large DFGs, respectively. Mehdi Kamal, Neda Kazemian Amiri, Arezoo Kamran, Seyyed Alireza Hoseini, Masoud Dehyadegari, Hamid Noori |
ASAP | 1 |
| 2007 | HW/SW partitioning using discrete particle swarmabstractHardware/Software partitioning is one of the most important issues of codesign of embedded systems, since the costs and delays of the final results of a design will strongly depend on partitioning. We present an algorithm based on Particle Swarm Optimization to perform the hardware/software partitioning of a given task graph for minimum cost subject to timing constraint. By novel evolving strategy, we enhance the efficiency and result's quality of our partitioning algorithm in an acceptable run-time. Also, we compare our results with those of Genetic Algorithm on different task graphs. Experimental results show the algorithm's effectiveness in achieving the optimal solution of the HW/SW partitioning problem even in large task graphs. Amin Farmahini Farahani, Mehdi Kamal, Sied Mehdi Fakhraie, Saeed Safari |
ACM Great Lakes Symposium on VLSI | 2 |
| 2006 | SOPC-Based Parallel Genetic AlgorithmabstractThe ever-growing complexity of the modern chips is forcing fundamental changes in the way systems are designed. System-on-a-Programmable-Chip (SOPC) concept is bringing a major revolution in the design of integrated circuits, due to the fact that it makes unprecedented levels of in-field integration possible. Genetic Algorithm (GA) is a powerful function optimizer that is used successfully to solve problems in many different disciplines. A major drawback of GA is that it needs huge computation time for sequential execution on PCs. Therefore, the hardware implementation of GA has been the focus of some recent studies. Parallel GA (PGA) is particularly important for efficient hardware implementation and promise substantial gains in performance and results. In this paper, a SOPC-based PGA framework is proposed. Our proposed framework can be used in real-time applications. We have implemented our proposed system on an Altera ® Stratix Development Kit and we compare its performance with the corresponding software simulation. The results obtained indicate a speedup of up to 50 times in the elapsed computation time. Mehdi Salmani Jelodar, Mehdi Kamal, Sied Mehdi Fakhraie, Majid Nili Ahmadabadi |
IEEE Congress on Evolutionary Computation | 2 |