EDBT 2026 Demo / reviewers in the wild / expert
Veronia Iskandar
dblp:303/8273
· DBLP profile ↗
6ranked-venue papers
5as first author
6since 2021 · last 2023
0000-0001-9420-5450ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Investigating the Impact of Non-Volatile Memories on Energy-Efficiency of Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-Grained Reconfigurable Architectures (CGRAs) are promising solutions to achieve more performance with the end of Moore's law. CGRAs can provide flexibility as well as near-ASIC energy efficiency. Since the advent of IoT and battery-powered edge devices, energy efficiency is becoming increasingly important. Memory accesses contribute to about 50% of overall energy consumption of the CGRAs. Interesting features of emerging non-volatile memories (eNVMs) like low power consumption and high density have grown the attentions. In this work, the effect of eNVMs on energy efficiency of CGRAs have been investigated. The analysis using Polybench benchmark suite shows that STT-MRAM and PCM can result in a 94 % and 85 % reduction of energy consumption of memory accesses respectively compared to SRAM. Moreover, total access latency can also be improved by 60% and 49% in STT-MRAM and PCM. Ensieh Aliagha, Veronia Iskandar, Stephan Enseleit, Diana Göhringer |
DSD | 2 |
| 2023 | Auto-DOK: Compiler-Assisted Automatic Detection of Offload Kernels for FPGA-HBM ArchitecturesabstractThe bandwidth improvement provided by high-bandwidth memory (HBM), and the capability of FPGAs to customize the processing and memory hierarchy, results in a considerable performance increase for memory-intensive work-loads such as graph processing, sorting, machine learning, and database analytics. Modern systems integrating 3D-stacked DRAM memory can be leveraged to realize the Near-Memory Computing (NMC) paradigm by offloading some computations to accelerators placed near the HBM. Although numerous studies have investigated efficient accelerators for FPGA-HBM platforms, researchers have not proposed a systematic way for identifying which application kernels are suitable for execution near the HBM. In this article, we propose compiler support for recognizing offloading candidates without any burden on programmers. Auto-DOK analyzes an application code based on criteria derived from the hardware design goals of FPGA-HBM platforms, and automatically identifies kernels suitable for offloading. We evaluate Auto-DOK on benchmarks ranging from microbenchmarks to real-world kernels. Our results show that Auto-DOK can correctly identify kernels and input sizes suitable for execution near the HBM, and prevents slowdown caused by incorrect offloading decisions for other workloads. Moreover, Auto-DOK operates at compile time with negligible overhead and without the need for expensive profiling. Veronia Iskandar, Mohamed Abdelghany, Diana Göhringer |
DSD | 1 |
| 2023 | Compiler-Assisted Kernel Selection for FPGA-based Near-Memory Computing PlatformsabstractThe speed of modern computing systems has improved significantly, thanks to advances in CMOS technology. However, the memory bandwidth of DRAM has not kept pace with these improvements in terms of latency and energy consumption, which is known as the memory wall [1]. FPGAs with high-bandwidth memory (HBM) provide significantly improved performance on memory-intensive tasks, such as graph processing and machine learning. By leveraging 3D-stacked DRAM memory on FPGAs, it is possible to realize the Near-Memory Computing (NMC) paradigm, which involves offloading some kernels to be processed close to the memory. While there have been many studies on NMC accelerators, there is no established method for determining which application kernels are suitable for execution near the HBM. To fully realize the potential of FPGA-HBM architectures, it is important to identify offloading candidates without relying on programmers' knowledge. However, this is a non-trivial task due to the complexity of modern applications. To address this issue, we propose a compiler-assisted tool-flow for the automatic selection of kernels to be offloaded. Veronia Iskandar, Mohamed Abdelghany, Diana Göhringer |
FCCM | 1 |
| 2023 | Performance Estimation and Prototyping of Reconfigurable Near-Memory Computing SystemsabstractThe concept of near-memory computing (NMC) has emerged as a promising solution to address the memory wall challenges faced by future computing architectures. By utilizing modern systems that integrate 3D-stacked DRAM memory, the NMC paradigm minimizes unnecessary data movement between the memory subsystem and the CPU. FPGA vendors have incorporated 3D-stacked memories into their products to meet the increasing bandwidth requirements of memory-intensive applications, enabling FPGAs to compete with GPU solutions in terms of speed and energy efficiency. Recent NMC proposals focus on different data processing workloads, including graph processing and machine learning. This work addresses the research questions of how to leverage the full bandwidth of 3D-stacked high-bandwidth memory and how to facilitate the adoption of the near-memory computing paradigm. Veronia Iskandar, Mohamed Abdelghany, Diana Göhringer |
FPL | 1 |
| 2023 | Near-memory Computing on FPGAs with 3D-stacked Memories: Applications, Architectures, and OptimizationsabstractThe near-memory computing (NMC) paradigm has transpired as a promising method for overcoming the memory wall challenges of future computing architectures. Modern systems integrating 3D-stacked DRAM memory can be leveraged to prevent unnecessary data movement between the main memory and the CPU. FPGA vendors have started introducing 3D memories to their products in an effort to remain competitive on bandwidth requirements of modern memory-intensive applications. Recent NMC proposals target various types of data processing workloads such as graph processing, MapReduce, sorting, machine learning, and database analytics. In this article, we conduct a literature survey on previous proposals of NMC systems on FPGAs integrated with 3D memories. By leveraging the high bandwidth offered from such memories together with specifically designed hardware, FPGA architectures have become a competitor to GPU solutions in terms of speed and energy efficiency. Various FPGA-based NMC designs have been proposed with software and hardware optimization methods to achieve high performance and energy efficiency. Our review investigates various aspects of NMC designs such as platforms, architectures, workloads, and tools. We identify the key challenges and open issues with future research directions. Veronia Iskandar, Mohamed Abdelghany, Diana Göhringer |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2021 | Near-Data-Processing Architectures Performance Estimation and Ranking using Machine Learning PredictorsabstractThe near-data processing (NDP) paradigm has emerged as a promising solution for the memory wall challenges of future computing architectures. Modern 3D-stacked DRAM systems can be exploited to prevent unnecessary data movement between the main memory and the CPU. To date, no standardized simulation frameworks or benchmarks are available for the systematic evaluation of NDP systems. Identifying which type of high-performance 3D memory is suitable to use in an NDP system remains a challenge. This is mainly due to the fact that understanding the interactions between modern workloads and the memory subsystem is not a trivial task. Each memory type has its advantages and drawbacks. Additionally, memory access patterns vary greatly across applications. As a result, the performance of a given application on a given memory type is difficult to intuitively predict. There is no specific memory type that can effectively provide high performance for all applications.In this work, we propose a machine learning framework that can efficiently decide which NDP system is suitable for an application. The framework relies on performance prediction based on an input set of application characteristics. For each NDP system we are examining, we build a machine learning model that can accurately predict performance of previously unseen applications on this system. Our models are on average 200x faster than architectural simulation. They can accurately predict performance with coefficients of determination ranging between 0.88 and 0.92, and root mean square errors ranging between 0.08 and 0.19. Veronia Iskandar, Mohamed Abdelghany, Diana Göhringer |
DSD | 1 |