EDBT 2026 Demo / reviewers in the wild / expert
Omar Eldash
dblp:201/5997 · also Omar K. Eldash
· DBLP profile ↗
6ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-3651-5132ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Analog In-Memory Computing Enhanced FPGA for High-Throughput and Energy-Efficient AccelerationabstractThe ever-growing demand for AI computing, coupled with slowing performance gains in chip manufacturing, has heightened the role of FPGA-based accelerators. FPGAs enable the implementation of application-customized parallel dataflows due to their reconfigurability, achieving high energy efficiency. However, the bit-level routing fabric on FPGAs often results in high overheads because large amounts of data must be shuttled between compute blocks and memory blocks on the FPGA. We propose enhancing FPGAs with in-memory computing macros, specifically analog Dot Product Engines based on non-volatile RRAM devices. Using the Verilog to Routing (VTR) framework, we simulate a novel 40 nm, 26.2 mm × 26.2 mm architecture and employ a custom event-driven simulator to evaluate its performance. Our design achieves 25.5 ×103TOPS/W, an average ×31.4 throughput improvement and an average ×9,380 energy efficiency improvement when compared to state-of-the-art FPGA implementations of AI models. Archit Gajjar, Omar Eldash, Aishwarya Natarajan, Xia Sheng, Giacomo Pedretti, Aman Arora 0001, Paolo Faraboschi, Jim Ignowski, Luca Buonanno |
FCCM | 3 |
| 2025 | Enhancing FPGAs with Analog In-Memory Computing MacrosabstractWhile the AI computing needs are ever-increasing and the innovation in models generates tens of new architectures yearly, the performance gain from improvements in chip manufacturing has slowed down. Within this context, FPGA-based accelerators play a fundamental role. FPGAs are the backbone of specialized architectures, their reconfigurability being the key differentiation that enables an effective design space exploration. At the same time, to overcome the limitations induced by the memory bottleneck, the computing architectures community has proposed the in-memory computing paradigm: storage and computations are both performed in non-volatile memory devices. Archit Gajjar, Omar Eldash, Aishwarya Natarajan, Rand Jean, Xia Sheng, Giacomo Pedretti, Paolo Faraboschi, Jim Ignowski, Luca Buonanno |
FPGA | 3 |
| 2025 | RACE-IT: A Reconfigurable Analog Computing Engine for In-Memory Transformer AccelerationabstractTransformer models represent the cutting edge of Deep Neural Networks (DNNs) and excel in a wide range of machine learning tasks. However, processing these models demands significant computational resources and results in a substantial memory footprint. While In-memory Computing (IMC) offers promise for accelerating Vector-Matrix Multiplications (VMMs) with high computational parallelism and minimal data movement, employing it for other crucial DNN operators remains a formidable task. This challenge is exacerbated by the extensive use of complex activation functions, Softmax, and data-dependent matrix multiplications (DMMuls) within Transformer models. To address this challenge, we introduce a Reconfigurable Analog Computing Engine (RACE) by enhancing Analog Content Addressable Memories (ACAMs) to support broader operations. Based on the RACE, we propose the RACE-IT accelerator (meaning RACE for In-memory Transformers) to enable efficient analog-domain execution of all core operations of Transformer models. Given the flexibility of our proposed RACE in supporting arbitrary computations, RACE-IT is well-suited for adapting to emerging and non-traditional DNN architectures without requiring hardware modifications. We compare RACE-IT with various accelerators. Results show that RACE-IT increases performance by 453× and 15×, and reduces energy by 354× and 122× over the state-of-the-art GPUs and existing Transformer-specific IMC accelerators, respectively. Aishwarya Natarajan, Luca Buonanno, Archit Gajjar, Ron M. Roth, Sergey Serebryakov, John Moon, Omar Eldash, Jim Ignowski, Giacomo Pedretti |
ICCD | 8 |
| 2024 | CAMSHAP: Accelerating Machine Learning Model Explainability with Analog CAMabstractThe recent success of machine learning (ML) models has led to increasing demands for model explanations - why a result was given - along with model predictions. Tree-based ML models are considered more explainable than deep neural networks and higher performers in several domains. However, algorithms computing model explanations are irregular and scale poorly with model size. While many custom accelerators for training and inference have been proposed, little attention has been paid to accelerating model explanations. This lack of explanatory capability has limited the use of these models for real-time decision-making systems in critical fields such as healthcare, autonomous operation and cybersecurity. John Moon, Giacomo Pedretti, Pedro Bruel, Sergey Serebryakov, Omar Eldash, Luca Buonanno, Catherine Graves, Paolo Faraboschi, Jim Ignowski |
ICCAD | 5 |
| 2022 | Designing Novel AAD Pooling in Hardware for a Convolutional Neural Network AcceleratorabstractConvolutional neural network (CNN) hardware accelerators for specialized Internet of Things (IoT) requiring high accuracy is an emerging research topic. The pooling module in a CNN pipeline impacts both the speed and accuracy of a classification task. This work proposes the design and hardware implementation of a novel pooling method absolute average deviation (AAD) for CNN accelerator. AAD utilizes the spatial locality of pixels using vertical and horizontal deviations to achieve higher accuracy, lower area, and lower power consumption than mixed pooling without increasing the computational complexity. AAD is tested on four different datasets: EEG, ImageNet, Common Objects in Context (COCO), United States Postal Service (USPS), and multiple CNN structures: CNN, VGG16, VGG19, ResNet, and DenseNet. In hardware, AAD is implemented using Very High Speed Integrated Circuit (VHSIC) Hardware Description Language (VHDL) on Altera Arria10 GX field-programmable gate array (FPGA) and 45-nm technology using Synopsys Design Compiler. The area and power consumption are found to be 244.46 nm2and 0.31 mW, respectively. AAD achieves 98% accuracy with lower computational and hardware costs compared to mixed pooling, making it an ideal pooling mechanism for an IoT CNN accelerator. Kasem Khalil, Omar Eldash, Ashok Kumar 0001, Magdy A. Bayoumi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | A Cost-Effective Self-Healing Approach for Reliable Hardware SystemsabstractIn this paper, self-healing concept for hardware systems is investigated and a new approach is proposed. Hardware systems have been proposing imitations to biological organisms in the way they offer healing and recovery abilities. Digital systems with inspired homogeneous architecture have improved capabilities to compensate for any faults. Self-healing is defined by the ability of a system to detect faults or failures and fix them. One of the main problems in current self-healing approaches is area overhead and scalability for complex structures considering they are based on redundancy and spare blocks. This paper proposes a different approach for self-healing based on embryonic structures without a need for spare cells. The area overhead is lower compared to other approaches relying on spare cells. The proposed approach relies on time multiplexing two functions in one cell within one clock cycle. The reliability of the proposed technique is studied and compared to conventional system with different failure rates. This approach is capable of healing up to 50% of the cells where each cell can cover another neighbor failed cell at most. The area overhead is 9% for the proposed approach which is much lower compared to other approaches using spare cell. The proposed approach is applied to investigate two case studies; ALU array, and neural network. Kasem Khalil, Omar Eldash, Magdy A. Bayoumi |
ISCAS | 2 |