EDBT 2026 Demo / reviewers in the wild / expert
Ahmed Kamaleldin
dblp:206/8227
· DBLP profile ↗
9ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-7446-7741ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A RISC-V Coprocessor to Accelerate Structured Sparse-Dense Matrix Multiplications
Rohan Krishna Vijayaraghavan, Ahmed Kamaleldin, Pruthul Pradeep Andy, Diana Göhringer |
ISCAS | 2 |
| 2026 | MoSim: A Modular Simulation Framework for FPGA-based Systolic CNN Accelerators
Zuwen Ou, Guangjin Li, Ahmed Kamaleldin, Diana Göhringer |
ISCAS | 4 |
| 2026 | ProCon-V: A Programmable Tightly Coupled Convolution Accelerator Based on RISC-V Custom Instructions for Edge DevicesabstractThe increasing use of deep neural network (DNN) based applications on edge devices has driven the need for innovative computing solutions that can meet the high computational requirements of DNN workloads while ensuring low power consumption. Therefore, enhancing the general-purpose (GP) computing units of edge devices by integrating DNN-specific functions and operations into existing instruction set architectures (ISAs) provides the flexibility to support a wide variety of neural network models, thereby improving computing efficiency. This article proposes a programmable tightly coupled accelerator for convolution operations called ProCon-V. It is based on RISCV custom instructions, extending the standard ISA to support 2-D convolution operations. The accelerator is tightly coupled with a RISC-V-based pipeline through an open-source extension interface [21], offloading convolution operations from the pipeline’s main execution unit. The ProCon-V computing engine is based on a scalable systolic array architecture for partial-sum computation, with hardware-supported Im2Col transformation that supports multi-channel convolution layers with various configurations. Moreover, the tightly coupled accelerator enables the execution of convolution layers through a set of custom instructions handled by the RISC-V core without requiring extra modifications to the core pipeline. ProCon-V is implemented and evaluated on an AMD/Xilinx Virtex Ultrascale+ FPGA featuring dual clock domains for the RISC-V core and the convolution accelerator. Evaluation results demonstrate that the proposed accelerator features a maximum computing throughput of 174.4 GOPS for INT8-based operations with an energy efficiency of 181.7 GOPS/W. Ahmed Kamaleldin, Habib Aouinti, Diana Göhringer |
IEEE Trans. Computers | 1 |
| 2025 | Towards an Energy-Efficient RISC-V Core Architecture with Dynamic Dual-Issue and Clock GatingabstractThe escalating demand for energy-efficient embedded systems necessitates innovative approaches to mitigate dynamic power consumption in general-purpose cores. This paper introduces a threshold-based, workload-adaptive clockgating technique for a dual-issue RISC-V core, dynamically disabling the underutilized second-issue datapath to reduce dynamic power consumption. The method employs a threshold-based clock-gating controller that dynamically activates the secondissue pipeline according to the target workload and desired energy/performance trade-offs. By integrating latch- and flip-flop-based clock-gating mechanisms into the open-source VeeR EH1 dual-issue core, coarse-grained control over the secondissue pipeline is achieved, enhancing energy efficiency with minimal impact on computing performance. For prototyping and evaluation, an FPGA implementation targeting AMD/Xilinx FPGA devices, along with CoreMark benchmark evaluation, has been conducted, demonstrating 5% improvement in energy efficiency and a dynamic power savings of up to 15%, making the approach particularly suitable for power-constrained embedded systems. Ahmad Othman, Hueseyin Ege Pamuk, Ahmed Kamaleldin, Diana Göhringer |
DSD | 3 |
| 2025 | Towards Instruction-Controlled In-Pipeline GEMM Acceleration in a Dual-Issue RISC-V Core for Edge ApplicationsabstractGEMM (General Matrix Multiplication) is a fundamental operation in deep learning (DL), serving as the key computing kernel for neural network layers, such as convolutional and fully connected layers. As the deployment of DL models expands beyond high-performance computing to resource-constrained embedded systems, there is a growing demand for efficient GEMM implementations that can meet strict power budgets and real-time processing requirements. This paper introduces a dual-issue in-order RISC-V core with an inpipeline GEMM accelerator based on the VeeR EH1 core [13], eliminating control signals and data transfer overhead inherent to coprocessors and standalone accelerators. By integrating GEMM execution directly into the pipeline, the design leverages core utilization and enables compiler-time preloading of operands, streamlining memory access and reducing dynamic branching, thereby minimizing branch mispredictions. Further, pipelined load/store operations mitigate memory access stalls and cache misses, thereby enhancing the pipeline throughput. Speedup, measured using Verilator simulation, demonstrates an average of$\sim 5 \times$over the reference VeeR EH1 baseline core. By integrating GEMM acceleration within the pipeline, this work bridges the gap between domain-specific accelerators and general-purpose cores, delivering scalable computing performance. Ahmad Othman, Darmen Ilyas, Ahmed Kamaleldin, Diana Göhringer |
FPL | 3 |
| 2022 | A Hybrid Memory/Accelerator Tile Architecture for FPGA-based RISC-V Manycore SystemsabstractMulti/manycore Systems-on-Chip are increasingly adopted for heterogeneous systems, providing a high degree of computing scalability and energy efficiency. However, the steady increase in heterogeneous tiles number leads to an expansion in resource usage and design cost. Therefore, reusability and modularity of the tile architecture to support different types of compute or memory units are key elements to reduce resource usage. Meanwhile, with the proliferation of RISC-V instruction set architecture, the modularity and reusability of compute tiles have been increased. In this work, we present a modular and reusable memory/accelerator tile architecture that supports two modes of operations as a memory or an accelerator tile. The proposed tile architecture is suitable to be integrated into a NoC based manycore architecture along with RISC-V based compute tiles. The hybrid tile features a shared non-coherent scratchpad memory that can be accessed directly by RISC-V compute tiles through NoC or by the local hardware accelerator logic inside the tile. Tile mode configuration and data transfer over the NoC are managed through control messages issued by RISC-V compute tiles based on running application requirements. Moreover, the proposed tile supports the flexibility to change the local hardware accelerator functionality at run-time using dynamic and partial reconfiguration. For evaluation, two manycore configurations are developed including 4 and 8 RISC-V compute tiles with 4 cores per tile. Several use cases based on signal processing kernels and hardware accelerators are used for performance evaluation in terms of memory transfer latency and computing time for two manycore configurations. Maximum data transfer throughput of 500 MB/s is achieved between the proposed hybrid tile and a single RISC-V compute tile. The proposed tile architecture is implemented and evaluated on a Xilinx Virtex Ultrascale+ FPGA. Ahmed Kamaleldin, Diana Göhringer |
FPL | 1 |
| 2022 | An Agile Tile-based Platform for Adaptive Heterogeneous Many-Core SystemsabstractComputing heterogeneity is a crucial demand for today's systems-on-chip requirements. Current many-core computing architectures feature a scalable number of heterogeneous compute units supporting a wide range of application domains. However, supporting both heterogeneity and computing scalability brings significant design challenges related to on-chip communication between heterogeneous components and run-time management. This leads to growing design time, development cost, and lack of hardware modularity and re-usability. This PhD work aims to develop and design a modular and adaptive hardware platform for realizing different types and taxonomies of heterogeneous many-core systems targeting FPGAs reusing the same hardware components. The proposed platform is based on a modular and scalable tile-based architecture supporting heterogeneous instruction set architectures (ISAs), seamless integration of custom hardware accelerators and several memory hierarchies. In this paper, the proposed tile-based platform, preliminary results, and evaluation are presented targeting FPGAs. Finally, planned and future works are highlighted. Ahmed Kamaleldin, Diana Göhringer |
FPT | 1 |
| 2021 | Design For Agility: A Modular Reconfigurable Platform for Heterogeneous Many-Core ArchitecturesabstractReconfigurable many-core computing platforms are gaining increasing attention for cloud and edge computing because of their high degree of scalability as well as flexibility. Heterogeneous many-core architectures provide more computing capabilities for domain-specific and general-purpose applications. However, bringing heterogeneous and custom computing elements together increases on-chip communication and run-time management complexities. This leads to growing design time and development cost, in addition to lack of platform re-usability. The scope of this PhD work is the design of a modifiable and modular hardware platform that provides a high degree of agility to change types or specifications of computing elements at design and run-time to achieve the best performance for different application demands using the same platform components. Different acceleration strategies and memory hierarchies are supported. In this paper, the proposed platform, preliminary results, and evaluation are presented targeting FPGAs. Finally, planned and future works are highlighted. Ahmed Kamaleldin, Diana Göhringer |
FPL | 1 |
| 2017 | Design guidelines for the high-speed dynamic partial reconfiguration based software defined radio implementations on Xilinx Zynq FPGAabstractReconfigurability of Field Programmable Gate Array (FPGA) makes it one of the most promising approaches in the implementation of the Software Defined Radio (SDR). FPGA Dynamic Partial Reconfiguration (DPR) feature emphasizes that approach by allowing the implemented SDR system to switch between multiple communications standards in runtime reusing the same FPGA hardware resources. Reconfiguration time is a significant parameter in DPR designs especially when a fast switching is required in real time system like SDR. In this paper, different designs of Partial Reconfiguration (PR) controllers are studied and evaluated according to their impact to improve the reconfiguration time of DPR-based SDR implementation. A multi-standard convolutional encoder design is implemented using DPR with different PR controllers as a case study. The design is implemented and tested on Xilinx Zynq evaluation board “ZC702”. This comparative study provides important design insights and recommendations to the DPR-based SDR designers to help them select the best PR controller based on their system throughput requirement and power budget. Ahmed Kamaleldin, Ahmed M. Soliman, Ahmed Nagy, Youssef Gamal, Ahmed Shalash, Yehea I. Ismail, Hassan Mostafa |
ISCAS | 1 |