EDBT 2026 Demo / reviewers in the wild / expert
Bengisu Elis
dblp:247/5657
· DBLP profile ↗
6ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0002-0781-8206ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | POSTER: Performance Comparison of GPU Programming Models Using HeCBench BenchmarksabstractGPUs play an important role in High-Performance Computing.The choice of GPU programming models plays a crucial role in achieving portability and performance.High-level programming models, such as SYCL and OpenMP offloading, have emerged, offering unified abstractions that enable developers to target multiple architectures with a single, maintainable codebase.However, achieving consistent performance across different models remains a significant challenge due to variations in abstraction levels, compiler optimizations, and runtime behavior.We present a profiling-based methodology for systematically comparing GPU programming models on NVIDIA and AMD GPUs.We apply our methodology to over 150 benchmarks of HeCBench, demonstrating its effectiveness in identifying performance issues in OpenMP, SYCL, HIP and CUDA implementations for AMD and NVIDIA GPUs. CCS Concepts• General and reference Jakob Schäffeler, Bengisu Elis, Amir Raoofy, Josef Weidendorfer, Martin Schulz 0001 |
CF | 2 |
| 2024 | A Portable Tool to Compare Performance Profiles from GPU Offloading Programming ModelsabstractGPUs are growingly dominating the High-Performance Computing ecosystem, and therefore, the ease of their programming is getting increasingly important. Standard and high-level offloading methods, like OpenMP offloading and OpenACC, facilitate portable and efficient offloading across different GPU platforms. However, pinpointing and troubleshooting performance variations among different models, implementations, or architectures poses a challenge due to varying abstraction levels and profilers employed. Therefore, to tackle this problem and to unwind the performance issues related to various offloading abstractions and models that are entangled together in practice, in this work, we introduce a portable tool to enable the comparison of performance profiles acquired from various offloading models and GPU platforms. For this, the tool first processes the collected profiles by different profilers to extract key performance indicatory metrics. For ease of comparison, the tool utilizes plots depicting the metrics of all target variants for relative comparison. Moreover, we demonstrate the tool's capabilities by discussing specific issues discovered by using the tool when comparing OpenMP offloading and CUDA implementations of Babelstream. Jakob Schäffeler, Bengisu Elis, Amir Raoofy, Josef Weidendorfer, Martin Schulz 0001 |
CF | 2 |
| 2024 | A Mechanism to Generate Interception Based Tools for HPC Libraries
Bengisu Elis, David Böhme, Olga Pearce, Martin Schulz 0001 |
Euro-Par (1) | 1 |
| 2024 | Non-Blocking GPU-CPU Notifications to Enable More GPU-CPU ParallelismabstractGPUs are increasingly popular in HPC systems, and more applications are adopting GPUs each day. However, the control synchronization of GPUs with CPUs is suboptimal and only possible after GPU kernel termination points, resulting in serialized host and device tasks. In this paper, we propose a novel CPU-GPU notification method that enables non-blocking in-kernel control synchronization of device and host tasks in combination with persistent GPU kernels. Using this notification method, we increase the overlap of CPU and GPU execution and with that parallelism. We present the concept and structure of the proposed notification mechanism together with in-kernel GPU-CPU control synchronization, using halo-exchange as an example. We analyze the performance of the halo-exchange pattern using our new notification method, as well as the interference between CPU and GPU operations due to the execution overlap. Finally, we verify our results using a performance model covering the halo-exchange pattern with the new notification method. Bengisu Elis, Olga Pearce, David Böhme, Jason Burmark, Martin Schulz 0001 |
HPC Asia | 1 |
| 2020 | QMPI: A next generation MPI profiling interface for modern HPC platforms
Bengisu Elis, Dai Yang, Olga Pearce, Kathryn Mohror, Martin Schulz 0001 |
Parallel Comput. | 1 |
| 2019 | QMPI: a next generation MPI profiling interface for modern HPC platformsabstractAs we approach exascale and start planning for beyond, the rising complexity of systems and applications demands new monitoring, analysis, and optimization approaches. This requires close coordination with the parallel programming system used, which for HPC in most cases includes MPI, the Message Passing Interface. While MPI provides comprehensive tool support in the form of the MPI Profiling interface, PMPI, which has inspired a generation of tools, it is not sufficient for the new arising challenges. In particular, it does not support modern software design principles nor the composition of multiple monitoring solutions from multiple agents or sources. We approach these gaps and present QMPI, as a possible successor to PMPI. In this paper, we present the use cases and requirements that drive its development, offer a prototype design and implementation, and demonstrate its effectiveness and low overhead. Bengisu Elis, Dai Yang, Martin Schulz 0001 |
EuroMPI | 1 |