EDBT 2026 Demo / reviewers in the wild / expert
Masudul Hassan Quraishi
dblp:279/3087
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2024
0000-0001-6939-1669ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Reconfigurable computing and FPGAs · 56% Distributed systems · 44% | |
| Computer networks
1 paper |
Edge and fog computing · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs
FPGA virtualization |
0.5 | 1 | 2021 | A Survey of System Architectures and Techniques for FPGA Virtualization · IEEE Trans. Parallel Distributed Syst. 2021 |
Distributed systems
resource sharing |
0.5 | 1 | 2021 | A Survey of System Architectures and Techniques for FPGA Virtualization · IEEE Trans. Parallel Distributed Syst. 2021 |
Reconfigurable computing and FPGAs › cloud FPGA
cloud FPGA acceleration |
0.1 | 1 | 2021 | A Survey of System Architectures and Techniques for FPGA Virtualization · IEEE Trans. Parallel Distributed Syst. 2021 |
Methods — techniques the papers use, named apart from their topics
survey · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | FSpGEMM: A Framework for Accelerating Sparse General Matrix-Matrix Multiplication Using Gustavson's Algorithm on FPGAsabstractGeneral sparse matrix–matrix multiplication (SpGEMM) is integral to many high-performance computing (HPC) and machine learning applications. However, prior field-programmable gate array (FPGA)-based SpGEMM accelerators either use the inner product algorithm with wasted and costly operations or Gustavson’s algorithm with a cache-based hardware architecture suffering from long-latency cache miss penalties and limited to embedded devices. In this work, we propose framework for accelerating SpGEMM (FSpGEMM), an OpenCL-based SpGEMM framework for accelerating Gustvason’s algorithm that includes an FPGA kernel implementing a throughput-optimized and scalable hardware architecture compatible with high-bandwidth memory (HBM) or traditional DDR-based memory. In addition, to address the irregular memory access patterns incurred by Gustavson’s algorithm, we propose a new buffering scheme tailored to Gustavson’s algorithm enabled by a new compressed sparse vector (CSV) format for representing sparse matrices and a row reordering technique as a preprocessing step to improve data reuse, and consequently, resource utilization. The proposed framework includes a host program implementing preprocessing functions for reordering input matrices and storing them in the proposed CSV format for further use. We implemented FSpGEMM using Intel FPGA SDK for OpenCL and experimented with a benchmark of sparse matrices selected from the SuiteSparse Matrix Collection on a Bittware 520N-MX FPGA board. The results show that the reordering technique improves the performance on average by 20.3% compared with the baseline. Finally, FSpGEMM outperforms the state-of-the-art (SOTA) FPGA implementation by an average of$2.23\times $in terms of execution cycles with the same benchmark and memory system configuration for a fair comparison. Erfan Bank Tavakoli, Michael Riera, Masudul Hassan Quraishi, Fengbo Ren |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2021 | FSCHOL: An OpenCL-based HPC Framework for Accelerating Sparse Cholesky Factorization on FPGAsabstractThe proposed FSCHOL framework consists of an FPGA kernel implementing a throughput-optimized hardware architecture for accelerating the supernodal multifrontal algorithm for sparse Cholesky factorization and a host program implementing a novel scheduling algorithm for finding the optimal execution order of supernodes computations for an elimination tree on the FPGA to eliminate the need for off-chip memory access for storing intermediate results. Moreover, the proposed scheduling algorithm minimizes on-chip memory requirements for buffering intermediate results by resolving the dependency of parent nodes in an elimination tree through temporal parallelism. Experiment results for factorizing a set of sparse matrices in various sizes from SuiteSparse Matrix Collection show that the proposed FSCHOL implemented on an Intel Stratix 10 GX FPGA development board achieves on average 5.5× and 9.7× higher performance and 10.3× and 24.7× lower energy consumption than implementations of CHOLMOD on an Intel Xeon E5-2637 CPU and an NVIDIA V100 GPU, respectively. Erfan Bank Tavakoli, Michael Riera, Masudul Hassan Quraishi, Fengbo Ren |
SBAC-PAD | 3 |
| 2021 | A Survey of System Architectures and Techniques for FPGA VirtualizationabstractFPGA accelerators are gaining increasing attention in both cloud and edge computing because of their hardware flexibility, high computational throughput, and low power consumption. However, the design flow of FPGAs often requires specific knowledge of the underlying hardware, which hinders the wide adoption of FPGAs by application developers. Therefore, the virtualization of FPGAs becomes extremely important to create a useful abstraction of the hardware suitable for application developers. Such abstraction also enables the sharing of FPGA resources among multiple users and accelerator applications, which is important because, traditionally, FPGAs have been mostly used in single-user, single-embedded-application scenarios. There are many works in the field of FPGA virtualization covering different aspects and targeting different application areas. In this article, we review the system architectures used in the literature for FPGA virtualization. In addition, we identify the primary objectives of FPGA virtualization, based on which we summarize the techniques for realizing FPGA virtualization. This article helps researchers to efficiently learn about FPGA virtualization research by providing a comprehensive review of the existing literature. Masudul Hassan Quraishi, Erfan Bank Tavakoli, Fengbo Ren |
IEEE Trans. Parallel Distributed Syst. | 1 |