EDBT 2026 Demo / reviewers in the wild / expert
Dimitrios Gourounas
dblp:238/5525
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-5011-9646ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HighWave: Large-Scale High-Bandwidth Wave Simulations on FPGAsabstractWave simulations are an important task in high-performance computing across various fields. Wave problems represented as partial differential equations (PDEs) are mas-sively parallel and thus well-suited for acceleration. FPGAs are an attractive solution due to their reconfigurability, allowing application-specific optimizations of various wave equations. However, memory bandwidth eventually becomes a bottleneck, where effective utilization of High-Bandwidth-Memory (HBM) or similar technology is essential in achieving scalable performance. GPUs rely on significant hardware resources to manage memory traffic. By contrast, FPGAs can specialize memory systems and optimize access patterns to achieve higher utilization on a lower power budget, but this requires careful design optimizations. This work introduces HighWave, a scalable FPGA-based spatial architecture that efficiently leverages HBM in acceleration of wave simulations. We focus on wave solvers that use discontinu-ous Galerkin schemes with Gauss-Lobatto-Legendre quadrature points, characterized by low arithmetic intensity and complex access patterns. HighWave integrates configurable processing cores and applies memory optimizations that maximize data reuse, minimize access irregularities, and ensure load balancing, thereby achieving high HBM bandwidth utilization. Unlike prior works for acceleration of wave simulations on FPGAs that lack HBM integration, HighWave provides better scalability for cluster-level deployment and supports a wide range of problems. We evaluate HighWave on the simulation of elastic and acoustic wave equations using an Intel Stratix 10 MX FPGA. Despite the high irregularity of data access patterns, HighWave achieves up to 72.2% utilization of peak HBM bandwidth. When normalizing for equivalent peak bandwidth, it delivers up to 48% higher performance than state-of-the-art GPU implementations on an Nvidia V100, with up to 2.84x higher energy efficiency. Dimitrios Gourounas, Austin G. James, Bagus Hanindhito, Arash Fathi, Lizy Kurian John, Andreas Gerstlauer |
FCCM | 1 |
| 2024 | Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAsabstractFPGAs are a promising platform for accelerating Deep Learning (DL) applications, due to their high performance, low power consumption, and reconfigurability. Recently, the leading FPGA vendors have enhanced their architectures to more efficiently support the computational demands of DL workloads. However, the two most prominent AI-optimized FPGAs, i.e., AMD/Xilinx Versal ACAP and Intel Stratix 10 NX, employ significantly different architectural approaches. This paper presents novel systematic frameworks to optimize the performance of General Matrix Multiplication (GEMM), a fundamental operation in DL workloads, by exploiting the unique and distinct architectural characteristics of each FPGA. Our evaluation on GEMM workloads for int8 precision shows up to 77 and 68 TOPs (int8) throughput, with up to 0.94 and 1.35 TOPs/W energy efficiency for Versal VC1902 and Stratix 10 NX, respectively. This work provides insights and guidelines for optimizing GEMM-based applications on both platforms, while also delving into their programmability trade-offs and associated challenges. Endri Taka, Dimitrios Gourounas, Andreas Gerstlauer, Diana Marculescu, Aman Arora 0001 |
FCCM | 2 |
| 2023 | FAWS: FPGA Acceleration of Large-Scale Wave SimulationsabstractEfficiently solving large-scale scientific problems described by partial differential equations (PDEs) is a critical task in high-performance computing. Application-specific hardware design is a well-known solution, but the wide range of kernels makes it infeasible to provision supercomputers with accelerators for all applications. This makes reconfigurable platforms a promising direction. In this work, we focus on wave simulations using discontinuous Galerkin solvers, as an important class of PDE applications. Existing work using FPGAs is limited to accelerating specific kernels or small problems that fit into FPGA BRAM. We present FAWS, a generic and configurable architecture for large-scale accelerated wave simulation problems running on FPGAs out of DRAM. FAWS exploits fine- and coarse-grain parallelism using a scalable array of application-specific processing cores, and incorporates novel dataflow optimizations, including prefetching, kernel fusion, and memory layout optimizations to minimize data transfers and maximize DRAM bandwidth utilization. We demonstrate FAWS on the simulation of elastic wave equations. Results show that a single FPGA core achieves 2.2x higher performance than 24 Xeon cores with 18.64x better energy efficiency, when given ~1.94x less peak DRAM bandwidth. Scaling to the same peak DRAM bandwidth, an FPGA is 4.27x and 1.96x faster than 24 CPU cores and an Nvidia P100 GPU, with 31.33x and 5.29x better efficiency, respectively. Dimitrios Gourounas, Bagus Hanindhito, Arash Fathi, Dimitar Trenev, Lizy Kurian John, Andreas Gerstlauer |
ASAP | 1 |
| 2023 | LAWS: Large-Scale Accelerated Wave Simulations on FPGAsabstractComputing numerical solution to large-scale scientific computing problems described by partial differential equations is a common task in high-performance computing. Improving their performance and efficiency is critical to exa-scale computing. Application-specific hardware design is a well-known solution, but the wide range of kernels makes it infeasible to provision supercomputers with accelerators for all applications. This makes reconfigurable platforms a promising direction. In this work, we focus on wave simulations using discontinuous Galerkin solvers, as an important class of applications. Existing work using FPGAs is limited to accelerating specific kernels or small problems that fit into FPGA BRAM. We present LAWS, a generic and configurable architecture for large-scale accelerated wave simulation problems running on FPGAs out of DRAM. LAWS exploits fine- and coarse-grain parallelism using a scalable array of application-specific cores, and incorporates novel dataflow optimizations, including prefetching, kernel fusion, and memory layout optimizations to minimize data transfers and maximize DRAM bandwidth utilization. We further accompany LAWS with an analytical performance model that allows for scaling across technology trends and architecture configurations. We demonstrate LAWS on the simulation of elastic wave equations. Results show that a single FPGA core achieves 69% higher performance than 24 Xeon cores with 13.27x better energy efficiency, when given 1.94x less peak DRAM bandwidth. Scaling to the same peak DRAM bandwidth shows that an FPGA is 3.27x and 1.5x faster than 24 CPU cores and an Nvidia P100 GPU, with 22.3x and 4.53x better efficiency, respectively. Dimitrios Gourounas, Bagus Hanindhito, Arash Fathi, Dimitar Trenev, Lizy Kurian John, Andreas Gerstlauer |
FPGA | 1 |
| 2022 | GAPS: GPU-acceleration of PDE solvers for wave simulationabstractLarge-scale simulations of wave-type equations have many industrial applications, such as in oil and gas exploration. Realistic simulations, which involve a vast amount of data, are often performed on multiple nodes of an HPC cluster. Using GPUs for these simulations is attractive due to considerable parallelizability of the algorithms. Many industry-relevant simulations have characteristics in their physics or geometry that can be exploited to improve computational efficiency. Furthermore, the choice of simulation algorithm impacts computational efficiency significantly. Bagus Hanindhito, Dimitrios Gourounas, Arash Fathi, Dimitar Trenev, Andreas Gerstlauer, Lizy Kurian John |
ICS | 2 |
| 2021 | Wave-PIM: Accelerating Wave Simulation Using Processing-in-MemoryabstractWave simulations are used in many applications: medical imaging, oil and gas exploration, earthquake hazard mitigation, and defense systems, among others. Most of these applications require repeated solutions of the wave equation on supercomputers. Minimizing time to solution and energy consumption are very beneficial in this domain. Data movement overhead is one of the key bottlenecks that affect energy consumption. Bagus Hanindhito, Ruihao Li 0002, Dimitrios Gourounas, Arash Fathi, Karan Govil, Dimitar Trenev, Andreas Gerstlauer, Lizy Kurian John |
ICPP | 3 |