EDBT 2026 Demo / reviewers in the wild / expert
Zahra Ebrahimi
dblp:153/0637
· DBLP profile ↗
10ranked-venue papers
8as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 8 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | X-DINC: Toward Cross-Layer ApproXimation for theDistributed and In-Network ACceleration of Multi-Kernel ApplicationsabstractWith the rapid evolution of programmable network devices and the urge for energy-efficient and sustainable computing, network infrastructures are mutating toward a computing pipeline, providing In-Network Computing (INC) capability. Despite the initial success in offloading single/small kernels to the network devices, deploying multi-kernel applications remains challenging due to limited memory, computing resources, and lack of support for Floating Point (FP) and complex operations. To tackle these challenges, we present a cross-layer approximation and distribution methodology (X-DINC) that exploits the error resilience of applications. X-DINC utilizes a chain of techniques to facilitate kernel deployment and distribution across heterogeneous devices in INC environments. First, we identify approximation and optimization opportunities in data acquisition and computation phases of multi-kernel applications. Second, we simplify complex arithmetic operations to cope with the computation limitations of the programmable network switches. Third, we perform application-level sensitivity analysis to measure the trade-off between performance gain and Quality of Results (QoR) loss when approximating individual kernels via various techniques. Finally, a greedy heuristic swiftly generates Pareto/near-Pareto mixed-precision configurations that maximize the performance gain while maintaining the user-defined QoR. X-DINC is prototyped on a Virtex-7 Field Programmable Gate Array (FPGA) and evaluated using the Blind Source Separation (BSS) application on industrial audio dataset. Results show that X-DINC performs separation up to 35% faster with up to 88% lower Area-Delay Product (ADP) compared to an Accurate-Centralized approach, when distributed across 2 to 7 network nodes, while maintaining audio quality within an acceptable range of 15–20 dB. Zahra Ebrahimi, Maryam Eslami, Xun Xiao, Akash Kumar 0001 |
Future Gener. Comput. Syst. | 1 |
| 2024 | GREEN: An Approximate SIMD/MIMD CGRA for Energy-Efficient Processing at the EdgeabstractThe rapid evolution of compute-intensive programs from bio-signal to image-, and video-processing has motivated moving toward Coarse Grained Reconfigurable Architectures (CGRAs), having high parallelism capability with post-fabrication datapath versatility. To enhance energy-efficiency of such error-resilient applications, State-of-the-Art (SoA) CGRAs exploit approximation techniques, while maintaining an acceptable accuracy for the final Quality of Result (QoR). However, such CGRAs suffer from overheads of utilizing separate Add/Mul/Div units. We propose GREEN as an energy-efficient CGRA, which enables synergistic effects of a chain of approximation and optimization techniques in various levels of abstraction, from application-, to architecture-, to circuit-level, in a cross-layer hierarchy. Enabling this, GREEN offers different levels of energy-accuracy trade-offs through the flexibility of its small Processing Elements (PEs), each of which can support various functionalities and precision-adaptability in a Single Instruction, Multiple Data (SIMD) or Multiple Instruction, Multiple Data (MIMD) manner. Experimental results obtained with Synopsys Design Compiler and Cadence Innovus at 45 nm CMOS technology node demonstrate the efficiency of the proposed SISD/SIMD/MIMD CGRA over the accurate and SoA counterparts. In particular, the MIMD mode of GREEN enables up to 6.6× higher throughput while dissipating 21% less energy than the accurate counterpart. Moreover, the end-to-end evaluation of GREEN variants on eight single-and multi-kernel applications from classification, bio-signal (ECG/EEG), and image/video processing domains demonstrates significant performance improvement, compared to the accurate CGRA. In particular, GREEN-MIMD not only speed-ups the ECG QRS detection by 49% and consumes 43% less area and 66% less energy than the accurate CGRA, but also maintains the heartbeat detection accuracy at 100%. GREEN implementations is available at https://cfaed.tu-dresden.de/pd-downloads. Zahra Ebrahimi, Akash Kumar 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | RAPID: Approximate Pipelined Soft Multipliers and Dividers for High Throughput and Energy EfficiencyabstractThe rapid updates in error-resilient applications along with their quest for high throughput has motivated designing fast approximate functional units for field-programmable gate arrays (FPGAs). Studies have proposed various imprecise functional techniques, albeit posed with three shortcomings: first, most existing inexact multipliers and dividers are specialized for application-specific integrated circuit (ASIC) platforms. Therefore, due to the architectural differences of underlying building blocks in FPGA and ASIC, ASIC-customized designs have not yielded comparable improvements when directly synthesized and ported to FPGAs. Second, state-of-the-art (SoA) approximate units are substituted, mostly in a single kernel of a multikernel application. Moreover, the end-to-end assessment is adopted on the quality of results (QoR), but not on the overall gained performance. Finally, the existing imprecise components are not designed to support a pipelined approach, which could boost the operating frequency/throughput of, e.g., division-included applications. In this article, we propose RAPID, the first pipelined approximate multiplier and divider architectures, customized for FPGAs. The proposed units efficiently utilize 6-input look-up tables (6-LUTs) and fast carry chains to implement Mitchell’s approximate algorithms. Our novel error-refinement scheme not only has negligible overhead over the baseline Mitchell’s approach but also boosts its accuracy to 99.4% for arbitrary size of multiplication and division. Experimental results obtained with Xilinx Vivado demonstrate the efficiency of the proposed pipelined and nonpipelined RAPID multipliers and dividers over accurate counterparts. In particular, the 4-stage pipelined architecture of a 32-bit RAPID multiplier (divider) enables$3.3\times $($5.1\times $) higher throughput,$2.3\times $($6.8\times $) higher throughput/Watt, and 52% (31%) savings of look-up tables (LUTs), over their 4-stage pipelined, accurate Intellectual Property (IP) counterparts. Moreover, the end-to-end evaluations of nonpipelined RAPID, deployed in three multikernel applications in the domains of biosignal processing, image processing, and moving object tracking for unmanned aerial vehicles (UAVs) indicate up to 35%, 33%, and 45% improvements in area, latency, and area-delay-product (ADP), respectively, over accurate kernels, with negligible loss in QoR. To springboard future research in reconfigurable and approximate computing communities, our implementations will be available and opensourced athttps://cfaed.tu-dresden.de/pd-downloads. Zahra Ebrahimi, Muhammad Zaid, Mark Wijtvliet, Akash Kumar 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Plasticine: A Cross-layer Approximation Methodology for Multi-kernel Applications through Minimally Biased, High-throughput, and Energy-efficient SIMD Soft Multiplier-dividerabstractThe rapid evolution of error-resilient programs intertwined with their quest for high throughput has motivated the use of Single Instruction, Multiple Data (SIMD) components in Field-Programmable Gate Arrays (FPGAs). Particularly, to exploit the error-resiliency of such applications, Cross-layer approximation paradigm has recently gained traction, the ultimate goal of which is to efficiently exploit approximation potentials across layers of abstraction. From circuit- to application-level, valuable studies have proposed various approximation techniques, albeit linked to four drawbacks: First, most of approximate multipliers and dividers operate only in SISD mode. Second, imprecise units are often substituted, merely in a single kernel of a multi-kernel application, with an end-to-end analysis in Quality of Results (QoR) and not in the gained performance. Third, state-of-the-art (SoA) strategies neglect the fact that each kernel contributes differently to the end-to-end QoR and performance metrics. Therefore, they lack in adopting a generic methodology for adjusting the approximation knobs to maximize performance gains for a user-defined quality constraint. Finally, multi-level techniques lack in being efficiently supported, from application-, to architecture-, to circuit-level, in a cohesive cross-layer hierarchy. In this article, we propose Plasticine , a cross-layer methodology for multi-kernel applications, which addresses the aforementioned challenges by efficiently utilizing the synergistic effects of a chain of techniques across layers of abstraction. To this end, we propose an application sensitivity analysis and a heuristic that tailor the precision at constituent kernels of the application by finding the most tolerable degree of approximations for each of consecutive kernels, while also satisfying the ultimate user-defined QoR. The chain of approximations is also effectively enabled in a cross-layer hierarchy, from application- to architecture- to circuit-level, through the plasticity of SIMD multiplier-dividers, each supporting dynamic precision variability along with hybrid functionality. The end-to-end evaluations of Plasticine on three multi-kernel applications employed in bio-signal processing, image processing, and moving object tracking for Unmanned Air Vehicles (UAV) demonstrate 41%–64%, 39%–62%, and 70%–86% improvements in area, latency, and Area-Delay-Product (ADP), respectively, over 32-bit fixed precision, with negligible loss in QoR. To springboard future research in reconfigurable and approximate computing communities, our implementations will be available and open-sourced at https://cfaed.tu-dresden.de/pd-downloads. Zahra Ebrahimi, Dennis Klar, Mohammad Aasim Ekhtiyar, Akash Kumar 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2021 | BioCare: An Energy-Efficient CGRA for Bio-Signal Processing at the EdgeabstractCoarse Grained Reconfigurable Architectures (CGRAs) have proved to be viable platforms for health monitoring applications. Targeting energy-efficiency, state-of-the-art (SoA) CGRAs are augmented with approximation techniques, while still maintain acceptable accuracy at final Quality of Result (QoR). However, such CGRAs suffer from overheads of collecting separate Add/Mul/Div units. We propose BioCare as an area- and energy-efficient CGRA for health-monitoring edge devices, which exploits the synergistic effects of multiple approximations across HW/SW stack. BioCare offers different levels of energy-accuracy trade-off through the plasticity of its small PEs, each can support precision-adaptability with a Single Instruction, Multiple Data (SIMD) manner. BioCare demonstrates its superiority over SoAs, by achieving up to 32% and 67% area- and energy-savings, with 3.6 χ higher throughput. In addition to analysis on multiple kernels, evaluations on a multi-kernel ECG application shows that BioCare speed-ups the QRS detection latency by 61%, with 0% loss in accuracy. Our implementations will be available at https://cfaed.tu-dresden.de/pd-downloads. Zahra Ebrahimi, Akash Kumar 0001 |
ISCAS | 1 |
| 2020 | LeAp: Leading-one Detection-based Softcore Approximate Multipliers with Tunable AccuracyabstractApproximate multipliers are ubiquitously used in diverse applications by exploiting circuit simplification, mainly specialized for Application-Specific Integrated Circuit (ASIC) platforms. However, the intrinsic architectural specifications of Field-Programmable Gate Arrays (FPGAs) prohibited comparable resource gains when directly applying these techniques. LeAp is an area-, throughput-, and energy-efficient approximate multiplier for FPGAs which efficiently utilizes 6-input Look-up Tables (6-LUTs) and fast carry chains in its novel approximate log calculator to implement Mitchell's algorithm. Moreover, three novel error-refinement schemes with negligible area overhead and independent from multiplier-size, have boosted accuracy to>99%. Experimental results obtained from Vivado, Artificial Neural Network (ANN) and image processing applications indicate superiority of proposed multiplier over accurate and state-of-the-art approximate counterparts. In particular, LeAp outperforms the 32x32 accurate multiplier by achieving 69.7%, 14.7%, 42.1%, and 37.1% improvement in area, throughput, power, and energy, respectively. The library of RTL and behavioral implementations will be open-sourced at https://cfaed.tu-dresden.de/pd-downloads. Zahra Ebrahimi, Salim Ullah, Akash Kumar 0001 |
ASP-DAC | 1 |
| 2020 | SIMDive: Approximate SIMD Soft Multiplier-Divider for FPGAs with Tunable AccuracyabstractThe ever-increasing quest for data-level parallelism and variable precision in ubiquitous multimedia and Deep Neural Network (DNN) applications has motivated the use of Single Instruction, Multiple Data (SIMD) architectures. To alleviate energy as their main resource constraint, approximate computing has re-emerged, albeit mainly specialized for their Application-Specific Integrated Circuit (ASIC) implementations. This paper, presents for the first time, an SIMD architecture based on novel multiplier and divider with tunable accuracy, targeted for Field-Programmable Gate Arrays (FPGAs). The proposed hybrid architecture implements Mitchell's algorithms and supports precision variability from 8 to 32 bits. Experimental results obtained from Vivado, multimedia and DNN applications indicate superiority of proposed architecture (both in SISD and SIMD) over accurate and state-of-the-art approximate counterparts. In particular, the proposed SISD divider outperforms the accurate Intellectual Property (IP) divider provided by Xilinx with 4x higher speed and 4.6x less energy and tolerating only 0.8% error. Moreover, the proposed SIMD multiplier-divider supersede accurate SIMD multiplier by achieving up to 26%, 45%, 36%, and 56% improvement in area, throughput, power, and energy, respectively. Zahra Ebrahimi, Salim Ullah, Akash Kumar 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2019 | An Efficient SRAM-Based Reconfigurable Architecture for Embedded ProcessorsabstractNowadays, embedded processors are widely used in wide range of domains from low-power to safety-critical applications. By providing prominent features such as variant peripheral support and flexibility to partial or major design modifications, field-programmable gate arrays (FPGAs) are commonly used to implement either an entire embedded system or a hardware description language-based processor, known as soft-core processor. FPGA-based designs, however, suffer from high power consumption, large die area, and low performance that hinders common use of soft-core processors in low-power embedded systems. In this paper, we present an efficient reconfigurable architecture to implement soft-core embedded processors in SRAM-based FPGAs by using characteristics such as low utilization and fragmented accessibility of comprising units. To this end, we integrate the low utilized functional units into efficiently designed look-up table (LUT)-based reconfigurable units (RUs). To further improve the efficiency of the proposed architecture, we used a set of efficient configurable hard logics that implement frequent Boolean functions while the other functions will still be employed by LUTs. We have evaluated effectiveness of the proposed architecture by implementing the Berkeley RISC-V processor and running MiBench benchmarks. We have also examined the applicability of the proposed architecture on an alternative open-source processor (i.e., LEON2) and a digital signal processing core. Experimental results show that the proposed architecture as compared to the conventional LUT-based soft-core processors improves area footprint, static power, energy consumption, and total execution time by 30.7%, 32.5%, 36.9%, and 6.3%, respectively. Sajjad Tamimi, Zahra Ebrahimi, Behnam Khaleghi, Hossein Asadi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | PEAF: A Power-Efficient Architecture for SRAM-Based FPGAs Using Reconfigurable Hard Logic Design in Dark Silicon EraabstractSignificant increase of static power in nano-CMOS era and, subsequently, the end of Dennard scaling has put a Power Wall to further integration of CMOS technology in Field-Programmable Gate Arrays (FPGAs). An efficient solution to cope with this obstacle is power gating inactive fractions of a single die, resulting in Dark Silicon. Previous studies employing power gating on SRAM-based FPGAs have primarily focused on using large-input Look-up Tables (LUTs). The architectures proposed in such studies inherently suffer from poor logic utilization which limits the benefits of power gating techniques. This paper proposes a Power-Efficient Architecture for FPGAs (PEAF) based on combination of Reconfigurable Hard Logics (RHLs) and a small-input LUT. In the proposed architecture, we selectively turn off unused RHLs and/or LUTs within each logic block by employing a reconfigurable controller. By mapping a majority of logic functions to simple-design RHLs, PEAF is able to significantly improve power efficiency without deteriorating the performance. Experimental results over a comprehensive set of benchmarks (MCNC, IWLS'05, and VTR) demonstrate that compared with baseline four-LUT architecture, PEAF reduces the total static power and Power-Delay-Product (PDP), on average, by 24.5 and 21.7 percent, respectively. This is while the overall system performance is also improved by 1.8 percent. PEAF increases total area by 18.9 percent, however, it still occupies 22.1 percent less area footprint than the six-LUT architecture with 31.5 percent improvement in PDP. Zahra Ebrahimi, Behnam Khaleghi, Hossein Asadi 0001 |
IEEE Trans. Computers | 1 |
| 2014 | Towards dark silicon era in FPGAs using complementary hard logic designabstractWhile the transistor density continues to grow exponentially in Field-Programmable Gate Arrays (FPGAs), the increased leakage current of CMOS transistors act as a power wall for the aggressive integration of transistors in a single die. One recently trend to alleviate the power wall in FPGAs is to turn off inactive regions of the silicon die, referred to as dark silicon. This paper presents a reconfigurable architecture to enable effective fine-grained power gating of unused Logic Blocks (LBs) in FPGAs. In the proposed architecture, the traditional soft logic is replaced with Mega Cells (MCs), each consists of a set of complementary Generic Reconfigurable Hard Logic (GRHL) and a conventional Look-Up Table (LUT). Both GRHL cells and LUTs can be power gated and turned off by controlling configuration bits. In the proposed MC, only one cell is active and the others are turned off. Experimental results on MCNC benchmark suite reveal that the proposed architecture reduces the critical path delay, power, and Power Delay Product (PDP) of LBs up to 5.3%, 30.4%, and 28.8% as compared to the equivalent LUT-based architecture. Ali Ahari, Behnam Khaleghi, Zahra Ebrahimi, Hossein Asadi 0001, Mehdi Baradaran Tahoori |
FPL | 3 |