EDBT 2026 Demo / reviewers in the wild / expert
Omid Akbari
dblp:197/6177
· DBLP profile ↗
9ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0003-4022-663XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | HPR-Mul: An Area and Energy-Efficient High-Precision Redundancy Multiplier by Approximate ComputingabstractFor critical applications that require a higher level of reliability, the triple modular redundancy (TMR) scheme is usually employed to implement fault-tolerant arithmetic units. However, this method imposes a significant area and power/energy overhead. Also, the majority-based voter in the typical TMR designs is highly sensitive to soft errors and the design diversity of the triplicated module, which may result in an error for a small difference between the output of the TMR modules. However, a wide range of applications deployed in critical systems are inherently error-resilient, that is, they can tolerate some inexact results at their output while having a given level of reliability. In this article, we propose a high precision redundancy multiplier (HPR-Mul) that relies on the principles of approximate computing to achieve higher energy efficiency and lower area, as well as resolve the aforementioned challenges of the typical TMR schemes, while retaining the required level of reliability. The HPR-Mul is composed of full precision (FP) and two reduced precision (RP) multipliers, along with a simple voter to determine the output. Unlike the state-of-the-art RP redundancy multipliers (RPR-Muls) that require a complex voter, the voter of the proposed HPR-Mul is designed based on mathematical formulas resulting in a simpler structure. Furthermore, we use the intermediate signals of the FP multiplier as the inputs of the RP multipliers, which significantly enhance the accuracy of the HPR-Mul. The efficiency of the proposed HPR-Mul is evaluated in a 15-nm FinFET technology, where the results show up to 70% and 69% lower power consumption and area, respectively, compared to the typical TMR-based multipliers. Also, the HPR-Mul outperforms the state-of-the-art RPR-Mul by achieving up to 84% higher soft error tolerance. Moreover, by employing the HPR-Mul in different image processing applications, up to 13% higher output image quality is achieved in comparison with the state-of-the-art RPR multipliers. Jafar Vafaei, Omid Akbari |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | Design Exploration of Fault-Tolerant Deep Neural Networks Using Posit Number Representation SystemabstractThe applications of deep neural networks (DNNs) in different safety-critical systems (such as autonomous vehicles and robotics) are experiencing emerging growth due to their high accuracy and potential for solving complex problems. However, a single failure in the hardware performing these DNN models can lead to irreparable results. Thus, improving the resilience of these models to transient faults (i.e., soft errors) has been of great interest in recent years. However, the traditional hardware redundancy techniques (such as TMR) are not cost-efficient due to their high resource overheads. In this article, we explored the potential of leveraging the posit number representation system for composing the DNN models, to achieve higher fault tolerance compared to the conventional fixed-point and IEEE 754 32-bit floating-point (FLOAT) number representation-based DNNs, without incurring significant overheads of the traditional hardware redundancy techniques. Posit numbers are composed of four fields, including a sign bit, regime value, exponent, and fraction bits, where the regime value does not exist in the FLOAT numbers. Our explorations are performed at the model, layer, and bitwise levels to determine the most appropriate posit format (e.g., the bit width of different posit fields) for composing the fault-tolerant DNN models. We studied the different posit, fixed-point, and FLOAT-based LeNet-5 and ResNet-50 networks trained with the MNIST and CIFAR-10 datasets, respectively. We then proposed a hardware-level method for error detection and correction of posit-based DNN models. Based on the results, the fault tolerance of the posit-based DNNs outperforms the FLOAT-based models, at each of the three investigated levels, where a 32-bit posit-based DNN achieved up to 15% more classification accuracy than the FLOAT-based one, for the LeNet-5 network. Also, in comparison with the fixed-point models, the posit-based networks showed up to 23% higher accuracy. Moreover, the enhanced 8-bit posit-based DNN that employed the proposed error detection and correction method results in, up to 11% and 41% higher classification accuracy than the unprotected posit and conventional FLOAT-based DNNs, respectively. Morteza Yousefloo, Omid Akbari |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | BEAD: Bounded error approximate adder with carry and sum speculations
Afshin Khaksari, Omid Akbari, Behzad Ebrahimi |
Integr. | 2 |
| 2023 | Accuracy Configurable Adders with Negligible Delay Overhead in Exact Operating ModeabstractIn this paper, two accuracy configurable adders capable of operating in approximate and exact modes are proposed. In the adders, which include a block-based carry propagate and a parallel prefix structure, the carry chains are cut off in the approximate mode limiting the carry chain depth to two blocks. In the case of parallel prefix adder, we propose a special carry generate tree equipped with a power gating means. In both of the proposed structures, the critical paths of the adders are not increased in the exact operating mode. Thus, the main objective of proposing these approximate adder structures is to present an accuracy configurable adder structure whose delay in the exact mode is almost the same as an exact adder. The efficacies of the proposed accuracy configurable adders are compared with some state-of-the-art adder structures using a 15nm CMOS technology. In addition, their efficacies are evaluated in two error-resilient applications. These studies show that the proposed carry-propagate adder has 22% (51%) lower energy consumption (error rate) compared to the best prior works. Also, the proposed parallel prefix adder provides, on average, 20% lower energy consumption compared to the exact parallel prefix adders. Farhad Ebrahimi-Azandaryani, Omid Akbari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2023 | X-Rel: Energy-Efficient and Low-Overhead Approximate Reliability Framework for Error-Tolerant Applications Deployed in Critical SystemsabstractTriple modular redundancy (TMR) is one of the most common techniques in fault-tolerant systems, in which the output is determined by a majority voter. However, the design diversity of replicated modules and/or soft errors that are more likely to happen in the nanoscale era may affect the majority voting scheme. Besides, the significant overheads of the TMR scheme may limit its usage in energy consumption and area-constrained critical systems. However, for most inherently error-resilient applications such as image processing and vision deployed in critical systems (such as autonomous vehicles and robotics), achieving a given level of reliability has more priority than precise results. Therefore, these applications can benefit from the approximate computing paradigm to achieve higher energy efficiency and a lower area. This article proposes an energy-efficient approximate reliability (X-Rel) framework to overcome the aforementioned challenges of the TMR systems and get the full potential of approximate computing without sacrificing the desired reliability constraint and output quality. The X-Rel framework relies on relaxing the precision of the voter based on a systematical error bounding method that leverages user-defined quality and reliability constraints. Afterward, the size of the achieved voter is used to approximate the TMR modules such that the overall area and energy consumption are minimized. The effectiveness of employing the proposed X-Rel technique in a TMR structure, for different quality constraints as well as with various reliability bounds, is evaluated in a 15-nm FinFET technology. The results of the X-Rel voter show delay, area, and energy consumption reductions of up to 86%, 87%, and 98%, respectively, when compared to those of the state-of-the-art approximate TMR voters. Also, the effectiveness of the proposed X-Rel-based TMR structure is assessed in four benchmark applications from different domains. For these benchmarks, the results show$1.59\times $,$2.35\times $, and$3.39\times $energy-delay-area-product (EDAP) reduction for less than 1%, 5%, and 10% output quality degradations, respectively. Finally, an image processing application is benchmarked to evaluate the X-Rel framework efficacy in the presence of errors, where the results show up to a$4.78\times $higher output image quality in comparison with the typical TMR voters. Jafar Vafaei, Omid Akbari, Muhammad Shafique 0001, Christian Hochberger |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | X-CGRA: An Energy-Efficient Approximate Coarse-Grained Reconfigurable ArchitectureabstractIn this article, we present an energy-efficient approximate CGRA (X-CGRA). Instead of conventional exact arithmetic units, it employs configurable approximate adders and multipliers in the so-called quality-scalable processing elements (QSPEs). Furthermore, the structure and functionality of the other architectural components, like context memory, are modified based on the quality-scalable operating modes of the QSPEs. The quality reconfigurability of the X-CGRA makes it amenable for both error-resilient and nonresilient applications. To map the applications on the X-CGRA, a mapping technique is proposed that efficiently utilizes the QSPEs and selects appropriate approximation modes in order to lower the energy consumption while satisfying a user-defined quality constraint. We evaluate the efficacy of our X-CGRA for several benchmark applications from different domains, including image/video processing, signal processing, and scientific computations. Different sizes of X-CGRA are synthesized using a 15-nm FinFET technology. For these benchmarks, the results indicate energy consumption reduction of up to $3.21\times $ compared to those of a typical exact CGRA, at the cost of 4% quality loss. Omid Akbari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram, Muhammad Shafique 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | PX-CGRA: Polymorphic approximate coarse-grained reconfigurable architectureabstractCoarse-Grained Reconfigurable Architectures (CGRAs) provide tradeoff between the energy-efficiency of Application Specific Integrated Circuits (ASICs) and the flexibility of General Purpose Processors (GPPs). State-of-the-art CGRAs only support exact architectures and precise application executions. However, a majority of the streaming applications such as multimedia and digital signal processing, which are amenable to CGRAs, are inherently error resilient. Therefore, these applications can greatly benefit from the emerging trend of Approximate Computing that leverages this error-resiliency to provide higher energy efficiency proportional to the tolerable accuracy loss (can even be constrained). This paper, for the first time, introduces the novel concept of Polymorphic Approximate CGRA (PX-CGRA) that employs heterogeneous tiles of Polymorphic-Approximated ALU Clusters (PACs) connected in a 2-D mesh style connection. These PACs can implement different approximate modes as well as accurate modes depending upon their selected configuration as per the run-time requirements of executing applications. For designing an efficient PX-CGRA, we propose a bottom-up design flow. In addition, the flow of application mapping on PX-CGRA is discussed including accuracy-level mapping, scheduling, and binding steps. To comprehensively evaluate the efficacy of the proposed CGRA, the complete PX-CGRA architecture in different sizes as well as with different PACs configurations are synthesized using a 15-nm FinFET technology. Our results show up to 15%-45% energy efficiency improvement for 5%-35% output quality degradation, respectively, when compared to the state-of-the-art exact-mode CGRA. Our proposed architecture and design methodology enable a new era of accuracy-configurable CGRAs to provide significant energy gains. Omid Akbari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram, Muhammad Shafique 0001 |
DATE | 1 |
| 2018 | Energy Consumption and Lifetime Improvement of Coarse-Grained Reconfigurable Architectures Targeting Low-Power Error-Tolerant ApplicationsabstractIn this work, the application of a voltage over-scaling (VOS) technique for improving the lifetime and reliability of coarse-grained reconfigurable architectures (GCRAs) is presented. The proposed technique, which may be applied to CGRAs used as accelerators for low-power, error-tolerant applications, reduces the (strongly voltage-dependent) wearout effects and the energy consumption of processing elements (PEs) whenever the error impact on the output quality degradation can be tolerated. This provides us with the ability to lessen the wearout and reduce energy consumption of PEs when accuracy requirement for the results is rather low. Multiple degrees of computational accuracy can be achieved by using different overscaled voltage levels for the PEs. The efficacy of the proposed technique is studied by considering the bias temperature instability. The study is performed for two error-resilient applications. The CGRAs are implemented with 15nm FinFET operating at a nominal supply voltage of 0.8V. In addition, supply voltages of 0.75, 0.7, 0.65, and 0.6V are considered as overscaled voltage levels for this technology. Based on the quality constraint requirements of the benchmarks, optimum overscaled voltage levels for various PEs are determined and utilized. The approach may provide considerable lifetime and energy consumption improvements over those of the conventional exact and approximate computation approaches. Hassan Afzali-Kusha, Omid Akbari, Mehdi Kamal, Massoud Pedram |
ACM Great Lakes Symposium on VLSI | 2 |
| 2017 | Dual-Quality 4: 2 Compressors for Utilizing in Dynamic Accuracy Configurable MultipliersabstractIn this paper, we propose four 4:2 compressors, which have the flexibility of switching between the exact and approximate operating modes. In the approximate mode, these dual-quality compressors provide higher speeds and lower power consumptions at the cost of lower accuracy. Each of these compressors has its own level of accuracy in the approximate mode as well as different delays and power dissipations in the approximate and exact modes. Using these compressors in the structures of parallel multipliers provides configurable multipliers whose accuracies (as well as their powers and speeds) may change dynamically during the runtime. The efficiencies of these compressors in a 32-bit Dadda multiplier are evaluated in a 45-nm standard CMOS technology by comparing their parameters with those of the state-of-the-art approximate multipliers. The results of comparison indicate, on average, 46% and 68% lower delay and power consumption in the approximate mode. Also, the effectiveness of these compressors is assessed in some image processing applications. Omid Akbari, Mehdi Kamal, Ali Afzali-Kusha, Massoud Pedram |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |